Compare commits

..

1 Commits

Author SHA1 Message Date
Dai Ha 2bce7e37b6 CB-622: rename the opencode mount key bridged -> fleetd
Unit C could not make this change. The tracked opencode.json is neutralized
by the worktree overlay (GitWorktrees.WORKTREE_HOSTILE_CONFIGS), so every
worker sees a stub instead of the real file and correctly reported that it
held no mount key. The lead has to make this one by hand, the same way
bridged.yaml is handled.

Takes effect for an opencode lead on its next session start.
2026-08-22 21:53:07 +02:00
395 changed files with 19800 additions and 77102 deletions
+5 -5
View File
@@ -1,15 +1,15 @@
{
"name": "fleetd",
"description": "Tooling for orchestrating a fleet of delegated coding agents through the fleetd MCP gateway.",
"name": "claude-bridge",
"description": "Tooling for orchestrating a fleet of delegated coding agents through the bridged MCP gateway.",
"owner": {
"name": "LTMS"
},
"plugins": [
{
"name": "fleet",
"name": "claude-bridge",
"source": "./plugin",
"description": "Mount the fleetd MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.2.0",
"description": "Make a project bridge-ready: mount the bridged MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.1.0",
"author": {
"name": "LTMS"
}
+2 -2
View File
@@ -3,7 +3,7 @@ name: architect
description: Refine work into clear, independent units before implementation.
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You are an architect in this fleet. You refine work before anyone builds it: scope,
acceptance criteria, risks, and a unit split. You read the repo and write analysis.
@@ -13,7 +13,7 @@ A design task is worked by two architects. Design alone first, then exchange and
say plainly where you disagree. Do not concede just to agree.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
investigate it further. Use `bridge_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
+2 -2
View File
@@ -3,14 +3,14 @@ name: dev
description: Implement one assigned unit, test it, and open a pull request.
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You implement the one unit you were given and nothing else. Work in your assigned
git worktree and branch. Never check out, rebase onto, or push to `main`. Confirm
the worktree root and branch before you edit. Use only paths under that root.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
investigate it further. Use `bridge_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
-23
View File
@@ -1,23 +0,0 @@
---
name: hunter
description: Sweep one assigned scope for defects and report ranked findings without changes.
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
You sweep the assigned package or scope for real defects. Read the full assigned scope before you
judge it. Report several ranked findings when the evidence supports them. Change nothing: do not
edit code, commit, push, or open a pull request.
You may run the build or tests to check a finding. Read the complete output and report the real
result. Do not hide failures with a pipe. State only checks you actually ran. The primary's IDE
tools are not yours. A mounted forge tool may use a blocked credential and fail by design.
Do only the assigned scope. Note anything outside it in one line and do not investigate it further.
Use `fleet_ask{question}` only when a decision belongs to the lead, such as an unclear requirement
or two defensible fixes. Do not ask about something you can decide by reading more code.
Your handoff must name the files you read, each ranked finding or `NO FINDINGS`, the checks you ran,
and any caveat for review.
The launcher provides the required bridge reply instructions for every member.
+2 -2
View File
@@ -3,7 +3,7 @@ name: reviewer
description: Review one assigned scope and report the most important real issue.
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You review the diff you were given. Report bugs, risks, and missing tests. You do
not change code.
@@ -12,7 +12,7 @@ Read the whole assigned scope before judging it. Review only that scope. If you
something outside it, note it in one line and do not investigate it further. Do not
run the build. The owner makes changes and runs checks.
Use `fleet_ask{question}` only when a decision belongs to the lead, such as an
Use `bridge_ask{question}` only when a decision belongs to the lead, such as an
unclear requirement or two defensible fixes. Do not ask about something you can
decide by reading more code.
-274
View File
@@ -1,274 +0,0 @@
---
name: fleets-status
description: Report the status of every fleet that shares one LavinMQ instance. Use for local daemon health, broker-wide fleet presence, and cross-host lead coordination checks.
---
# Status of every fleet on the shared LavinMQ instance
**The headline: always report what is missing.** This skill starts with the local fleet, then adds
broker-wide facts when its read-only credential exists. A missing fleet must appear as `unknown` or
`not reachable`, with the reason and the fix. Never leave it out.
The known topology has one LavinMQ instance on `10.10.20.13` (`fleet01`). AMQP uses port `5672`,
and the management API uses port `15672`. The Mac fleet owns vhost `/mac`. The fleet01 fleet owns
vhost `/fleet01`.
## 1. Protect credentials before any probe
**Hard rule — never print `LAVINMQ_URI`.** It is an AMQP URI with its password inline. It only
resolves in a login shell because `${SHARED_ENV}/tools/secrets.sh` supplies it. A non-login shell
can make every broker probe look empty.
- Never run `echo "$LAVINMQ_URI"`.
- Never put `${LAVINMQ_URI:-something}` in output. That form expands to the secret value when set.
- Parse the user, host, and password into shell or Python variables. Use them without printing them.
- Prefer `resolves` or `does not resolve` over any part of the value.
- Every command that can read `LAVINMQ_URI` must send all output through this redaction before it
reaches the report:
```bash
sed -E 's#://[^@]*@#://<redacted>@#g'
```
**The `g` flag is not optional.** Without it `sed` replaces only the first match on each line, so a
line carrying two URIs leaks the second one. `scripts/redeploy-fleetd.sh --check` prints lines like
that. Checked on 2026-08-27: without `g`, `amqp://u1:p1@h1/mac and http://u2:p2@h2:15672/api`
redacts the first pair and prints `u2:p2` in the clear.
Keep `pipefail` on when applying that filter. Otherwise the filter can hide a failed probe. Apply
the same no-print rule to the management password below, even though it is not in an AMQP URI.
## 2. Tier 1 — this fleet (always run)
Start here even when the broker tier is blocked. Work from the local fleetd checkout.
First run the read-only deployment check. It already checks the daemon process, deployed jar versus
the checkout `HEAD`, launchd state, and whether each configured token resolves in a login shell.
Do not copy those checks into new shell code. The script reads `LAVINMQ_URI`, so redact all output:
```bash
set -o pipefail
scripts/redeploy-fleetd.sh --check 2>&1 \
| sed -E 's#://[^@]*@#://<redacted>@#g'
git rev-parse HEAD
```
Treat jar drift as a top-level warning. A merge is not a deployment. State the running jar result
as `matches HEAD`, `drift`, or `unknown`; do not turn an unclear timestamp into a match.
Report the process identifier (PID) and uptime too:
```bash
PIDS="$(pgrep -f 'target/fleetd.jar' || true)"
if [ -z "$PIDS" ]; then
printf '%s\n' 'fleetd: not running'
else
for PID in $PIDS; do
ps -p "$PID" -o pid=,etime=,lstart=,command=
done
fi
```
Read the full health response. Keep the HTTP status because `503` means fleetd is running but herdr
is not reachable. Report both `herdr.version` and `herdr.protocol` when present:
```bash
curl -sS --max-time 5 -w '\nHTTP %{http_code}\n' http://127.0.0.1:8765/healthz
```
Call `fleet_whoami`, then call `fleet_list`. Preserve its sections in the report:
- `leads`, including which row is this lead;
- `members`, including state, role, profile, branch, and worktree when present;
- every per-profile `capacity` row, including `maxLoad`, `live`, `free`, and quarantine facts;
- the exact `healthCoverage` value.
Do not describe an empty `members` list as an empty fleet. It says only that no members are spawned.
Also do not hide a profile with `free: 0`; say whether load or credential quarantine caused it.
Show only WARN, ERROR, and SEVERE lines after the last `fleetd listening` line. This anchor stops an
old incident from looking current:
```bash
python3 - <<'PY'
from pathlib import Path
import re
path = Path("fleetd/fleetd.out")
if not path.exists():
print("cannot check current WARN/ERROR: fleetd/fleetd.out does not exist")
else:
lines = path.read_text(errors="replace").splitlines()
starts = [i for i, line in enumerate(lines) if "fleetd listening" in line]
if not starts:
print("cannot anchor WARN/ERROR: no 'fleetd listening' line exists")
else:
current = lines[starts[-1]:]
alerts = [line for line in current if re.search(r"\b(?:WARN|ERROR|SEVERE)\b", line)]
print(f"current WARN/ERROR/SEVERE count: {len(alerts)}")
for line in alerts[-50:]:
print(line)
PY
```
**What this tier cannot see:** it proves facts only about the Mac daemon at `127.0.0.1:8765`.
It cannot show the fleet01 daemon, broker queue depth, or broker consumers. The fleet01 REST service
at `10.10.20.13:8765` is not reachable from the Mac. Say this in the report rather than omitting
fleet01.
**But fleet01 IS reachable over SSH — checked 2026-08-28.** An older version of this line said SSH
was denied. That is true only for the user `dai.ha`. The host alias `fleet01` maps to user `ltms`,
and `ssh fleet01` works with key auth:
```bash
ssh -o BatchMode=yes -o ConnectTimeout=6 fleet01 'echo $(id -un)@$(hostname)'
```
So fleet01's daemon PID, uptime, jar and `/healthz` **can** be reported — over SSH, not over REST.
Do that rather than writing `not reachable`. `ltms` also has passwordless sudo there.
## 3. Tier 2 — the shared broker (run when management access exists)
**This tier is blocked today.** The AMQP user in `LAVINMQ_URI` can connect on port `5672`, but gets
HTTP `401` from the management API on port `15672`. An AMQP connection does not grant monitoring
access.
The operator must create a separate, read-only LavinMQ management user with the `monitoring` tag.
It needs access to inspect both `/mac` and `/fleet01`. Store its values as
`LAVINMQ_MANAGEMENT_USER` and `LAVINMQ_MANAGEMENT_PASSWORD` in
`${SHARED_ENV}/tools/secrets.sh`. Do not reuse or print the AMQP URI. Full multi-fleet status stays
blocked until this user exists.
When both variables resolve, run this from a login shell. It calls `GET /api/overview`,
`GET /api/vhosts`, `GET /api/queues`, and `GET /api/connections`. It prints selected status fields,
but never the user, password, Authorization header, or AMQP URI:
```bash
zsh -lc 'python3 - "$@"' -- <<'PY'
import base64
import json
import os
import sys
import urllib.error
import urllib.request
base = "http://10.10.20.13:15672"
user = os.environ.get("LAVINMQ_MANAGEMENT_USER", "")
password = os.environ.get("LAVINMQ_MANAGEMENT_PASSWORD", "")
if not user or not password:
print("broker tier: BLOCKED — management credential does not resolve in a login shell")
sys.exit(0)
token = base64.b64encode(f"{user}:{password}".encode()).decode()
def get(path):
request = urllib.request.Request(
base + path,
headers={"Authorization": "Basic " + token, "Accept": "application/json"},
)
with urllib.request.urlopen(request, timeout=5) as response:
return json.load(response)
try:
overview = get("/api/overview")
vhosts = get("/api/vhosts")
queues = get("/api/queues")
connections = get("/api/connections")
except urllib.error.HTTPError as error:
print(f"broker tier: BLOCKED — management API returned HTTP {error.code}")
sys.exit(0)
except Exception as error:
print(f"broker tier: BLOCKED — management API is not reachable: {type(error).__name__}")
sys.exit(0)
fleet_names = {"/mac": "Mac fleet", "/fleet01": "fleet01 fleet"}
print(json.dumps({
"overview": {
"lavinmq_version": overview.get("lavinmq_version"),
"rabbitmq_version": overview.get("rabbitmq_version"),
"queue_totals": overview.get("queue_totals", {}),
"object_totals": overview.get("object_totals", {}),
},
"fleets": [
{
"fleet": fleet_names.get(vhost.get("name"), "UNKNOWN FLEET"),
"vhost": vhost.get("name"),
"queues": [
{
"name": queue.get("name"),
"messages": queue.get("messages", 0),
"messages_ready": queue.get("messages_ready", 0),
"messages_unacknowledged": queue.get("messages_unacknowledged", 0),
"consumers": queue.get("consumers", 0),
}
for queue in queues if queue.get("vhost") == vhost.get("name")
],
"connections": [
{
"name": connection.get("name"),
"peer_host": connection.get("peer_host"),
"state": connection.get("state"),
}
for connection in connections if connection.get("vhost") == vhost.get("name")
],
}
for vhost in vhosts
],
}, indent=2, sort_keys=True))
PY
```
Map `/mac` to the Mac fleet and `/fleet01` to the fleet01 fleet. Keep any other vhost in the
report as `UNKNOWN FLEET`; do not drop it. For each vhost, total the ready, unacknowledged, and all
messages. Report every queue's consumer count and each live connection.
**A vhost with queues but zero consumers means that fleet's daemon is down while its durable state
survives. Call this out as a top-level warning.** This is the main reason to use the management API
instead of calling each remote daemon.
**What this tier cannot see:** without the new `monitoring` credential it cannot enumerate any
vhost, queue, depth, consumer, or connection. With the credential it still cannot report fleet01's
daemon PID, uptime, jar revision, `/healthz`, herdr version, or member capacity. Those need reachable
fleet01 REST or SSH access, which the Mac does not have today.
## 4. Tier 3 — cross-fleet lead coordination
Use the queue data from Tier 2. Select queues whose names match `lead.<coordId>.inbox`. Report each
queue's vhost, depth, consumer count, and the `coordId` between the prefix and suffix.
- A lead inbox with a consumer shows that a lead mailbox is live on that vhost.
- A durable lead inbox with zero consumers shows saved coordination state, but no live receiver.
- No lead inbox is not proof that coordination is disabled. The daemon may be down before declaring
its queue, or this account may not be allowed to see the vhost.
This Mac fleet currently sets both `broker.uriEnv` and `coordinator.uriEnv` to the same variable,
`LAVINMQ_URI`. Therefore its coordinator connects to `/mac`. Cross-host `fleet_send{coordId}` routes
only when both leads share the same coordinator vhost. If the fleet01 lead uses `/fleet01` for its
coordinator, the leads cannot see each other and the send will not route.
**Open question:** the fleet01 coordinator vhost has not been checked. Surface this question in
every report until a live `lead.<coordId>.inbox` consumer or fleet01's config proves the answer. Do
not claim that fleet01 uses `/fleet01` just because its member queues do.
Also compare these broker facts with the `leads` rows from local `fleet_list`. A missing remote lead
is `not visible from this coordinator`, not `down`, unless the broker consumer facts prove it.
**What this tier cannot see:** without Tier 2 management access it cannot list lead inboxes or their
consumers. Even with that access, a stopped fleet01 daemon leaves only durable queue history. That
history cannot prove which coordinator URI its current config would use after restart.
## 5. Report all fleets
Use one row per known or discovered fleet. Include blocked rows.
| Fleet | Daemon | Deployment | Herdr | Members/capacity | Queues/consumers | Lead coordination | Cannot check |
|---|---|---|---|---|---|---|---|
| Mac (`/mac`) | PID + uptime | jar vs `HEAD` | health + version + protocol | `fleet_list` + `healthCoverage` | facts or blocked reason | inbox facts or open question | exact missing facts |
| fleet01 (`/fleet01`) | reachable/down/unknown | value or `not reachable` | value or `not reachable` | value or `not reachable` | facts or blocked reason | inbox facts plus coordinator-vhost question | exact missing facts and fix |
Add rows for unknown vhosts. End with three short sections: `Current warnings`, `Checks that were
blocked`, and `Operator action`. Until the management user exists, `Operator action` must say:
> Create a read-only LavinMQ management user with the `monitoring` tag and access to `/mac` and
> `/fleet01`. Put its user and password in `${SHARED_ENV}/tools/secrets.sh` as
> `LAVINMQ_MANAGEMENT_USER` and `LAVINMQ_MANAGEMENT_PASSWORD`.
-226
View File
@@ -1,226 +0,0 @@
---
name: handover
description: Procedure for an outgoing lead to write the handover file that a fresh lead session inherits. Load this when your context is filling up and you are about to be replaced, whether you hand off by hand or fleetd does it for you. The file is the new lead's only inheritance — follow it exactly.
---
# Handover — write the file the next lead depends on
A lead session fills up its context and has to be replaced by a fresh one. The outgoing lead
writes a handover file, and the new session reads that file and carries on.
**There are two ways to hand off, and the file is the same either way.**
- **By hand.** You write the file, then tell the operator where it is. The operator starts the new
session and points it at the file. This always works.
- **With `fleet_handover`** (fleetd #480, merged 2026-09-11). You ask fleetd to do the swap: it
checks the file, clears your pane, and tells the fresh session to read it. This needs
`leadRollover:` in `fleetd.yaml`; without it every action answers a clean refusal naming
`NOT_CONFIGURED`, and you fall back to the manual path. Section 11 below is the procedure.
Nothing else in this skill changes between the two. Only who performs the swap changes.
**The new lead's only inheritance is that file.** It does not see your conversation, your plan,
or your screen. If the file is thin or wrong, the new lead re-derives what you already knew, and
that wastes hours. Writing a good handover file is real work. It is not paperwork you rush
through at the end of a session.
This skill is the procedure for writing it. Every rule below earned its place because a past
handover got it wrong.
## 1. Confirm you are the right session to write this
Run `fleet_whoami` first. It must answer `primary`. Only a primary (lead) session writes a
handover file. A worker's job ends with its own pull request, not a fleet-wide handoff.
## 2. Every number needs a command, run in this turn
A number is a claim: a count, a commit hash, a process id, a percentage, a queue depth. Before
you write one, run the command that produces it — now, in this turn, against the live state.
Never take a number from:
- earlier in your own conversation — the state has moved since then,
- a peer lead's report — that is their measurement, not yours,
- your own memory of an earlier session.
Put the command, or its real output, next to the number. That lets the next lead re-run it and
check it still matches. If you cannot measure something yourself, say so instead of guessing:
"the fleet01 lead reports 91 commits behind; I have not checked this myself."
## 3. Say what you measured and what you did not
Mark every claim as one of two things:
- **"I checked this myself, in the code or on this host, at `<time>`."**
- **"I did not check this myself; `<who>` reported it."**
Never present someone else's measurement as your own. This matters most for cross-host claims —
a peer lead's daemon, a worker's report, or something the operator said earlier that you cannot
re-verify from here.
## 4. Record open decisions, and who owns them
List three things:
- what the operator actually asked for, in their own words where you have them,
- what is still unanswered,
- any question you decided yourself instead of asking, with your reason.
Write the decision so it cannot be mistaken for the operator's instruction. Say plainly: "the
operator never answered X; I decided Y, because Z." Without this, the next lead either silently
reopens a closed question or assumes the operator chose something they never did.
## 5. Record live hazards
List anything that will break if the next lead does the obvious thing next. This includes:
- unpushed commits or unmerged branches,
- a build, a spawn, or a redeploy still running,
- code merged to `main` but not yet redeployed to the live daemon,
- any trap that looks safe and is not — say what goes wrong and why, not only that something is
"tricky."
## 6. Record what is explicitly not owed
List work that is finished, and work that another party has said they do not want touched. Name
who said so and when. Without this line, the next lead re-does closed work or reopens a question
a peer already declined to revisit.
## 7. Open the file with three re-measurement commands
The file's own first section must give the next lead three concrete commands to run before
acting on anything else in the file:
1. confirm role — for example `fleet_whoami`,
2. confirm the state of the working tree — for example `git status` and
`git rev-list origin/main..HEAD`,
3. read the live fleet — for example `fleet_list`.
Record what each command answered when you wrote the file, and tell the reader to run it again
rather than trust your answer. The point of this section is that the reader checks live state
before acting on any claim in the rest of the file, including yours.
## 8. Stamp the file with time and commit
Near the top of the file, write:
- the date and time you wrote it,
- the commit the tree was on (`git rev-parse HEAD`),
- whether the tree was clean (`git status`).
Without this, nobody can tell how old the file is, or which code it describes.
## 9. State plainly that the file goes stale fast
Say near the top: **re-measure anything you act on.** The file goes stale the moment anyone
merges a branch, spawns a member, or restarts the daemon. Everything in the file is a snapshot
of one moment, not a live fact.
## 10. What to leave out
Do not include:
- narration of how the session felt, or how hard something was,
- anything the repo already records — code structure, git history, or a rule already written in
`CLAUDE.md`. Point at it instead of repeating it,
- advice that is only true for the session that is ending — a half-open terminal, a local
variable, a train of thought with no state behind it.
A handover file is a record of state and decisions. It is not a diary.
## 11. Using `fleet_handover` (only if `leadRollover:` is configured)
**Run the three steps in this order. The order is not a style choice — the wrong order is
refused.**
1. **`fleet_handover{action: "open", reason: "<why now>"}`.** It returns a `token` and the
`handoverPath` you must write to. Nothing has happened to your pane yet.
**Write to exactly that path, and do not resolve it yourself.** It is always absolute, even when
the operator configured a relative `handoverPath`: fleetd resolves a relative one against your
own workspace before it hands it to you. The daemon and your pane can run in different
directories, so a path you resolve yourself can point at a different file from the one the daemon
will check.
**Check that the path is ignored by git before you write to it (#491).** A relative
`handoverPath` resolves inside YOUR workspace, which is usually a repository — and usually not
the `fleetd` one, so an ignore rule added to `fleetd` does not protect it. Run
`grep -n handover <your workspace>/.gitignore`. No output means the file you are about to write
will show up as untracked content in that repo. The file is a snapshot of live state and must
never be committed, so tell the operator rather than committing it or silently editing their
`.gitignore`.
2. **Write the handover file at that path**, following sections 1–10 above.
3. **Ask the operator, then `fleet_handover{action: "confirm", token, operatorConfirmed: true}`.**
Why that order: `confirm` refuses with `HANDOVER_STALE` unless the file was modified **after** the
`open` request. That check stops a leftover file from an earlier session being accepted as this
one's handover. So writing the file first and then calling `open` — the obvious order — always
fails.
`{action: "cancel", token}` drops a pending request without rolling.
**Things that will surprise you:**
- **`accepted` does not mean your pane has been cleared.** It means every gate passed and the roll
is scheduled to run once your current turn ends. Say your goodbye in the same turn — you will not
get another one.
- **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved
from your connection, so you can only ever roll yourself.
- **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you
are confident. Ask, wait for the answer, then pass what they said. `requireOperatorConfirm`
defaults to `true` and this is the only thing standing between a judgement call and a wiped
session.
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
- **The bootstrap prompt has never yet landed, and the fix is unproven (fleetd #489).** The first
real rollover, on 2026-09-12, joined `/clear` and the bootstrap text into one line and Claude Code
refused it as `Unknown command: /clearFresh`. The pane was never cleared and no context was lost,
so the failure was safe — the roll simply did nothing. PR #490 fixed the cause and is deployed,
but no roll has bootstrapped a fresh session end to end yet. **Assume it may still fail, and tell
the operator so before you confirm.** The recovery is the same either way: the file is already
written, so the operator starts a session and points it at the file. That is why you write the
file before you confirm, and never the other way round.
## Writing style
Write in plain English. Use everyday words, one idea per sentence, and active voice. Keep every
class, method, file, flag, and config key exactly as it appears in the code — replacing a
precise term with a vague one makes the sentence wrong, not simpler. Explain an abbreviation the
first time you use it.
If a diagram genuinely helps, put it in the `.md` file as a fenced ` ```mermaid ` block with no
hardcoded colors, so it stays readable on light and dark backgrounds. Quote any label that has
brackets, colons, or slashes.
## Template
```markdown
# Handover — <fleet name> lead session, <date and time>
Written at commit `<output of git rev-parse HEAD>`. Tree was <clean, or dirty: `<git status
summary>`>. Re-measure anything you act on — this file goes stale the moment anyone merges,
spawns, or restarts.
## 0. Do these three things first
1. Confirm your role: `fleet_whoami` — must answer `primary`. (Answered `<result>` at `<time>`.)
2. Confirm tree state: `git status`, `git rev-list origin/main..HEAD`. (`<result>` at `<time>`.)
3. Read the live fleet: `fleet_list`. (`<result>` at `<time>`.)
## 1. What the operator asked for
<the live instructions, in their words where you have them; what is still open; any decision
you made yourself, and why>
## 2. Open decisions, and who owns them
<one line per decision: who owns it, what is unanswered>
## 3. Live hazards
<one entry per hazard: what breaks, and why, if the next lead does the obvious thing>
## 4. What is not owed
<finished work, and work another party has declined; name who said so and when>
```
-102
View File
@@ -1,102 +0,0 @@
---
name: hunter
description: Defect-hunt procedure for a fleetd worker — sweep an assigned package for real bugs and report several ranked findings without fixing anything. Load this when the lead asks you to hunt or audit a scope rather than review one diff. Do NOT load `reviewer` for this; the two want different output.
---
# Hunter worker — procedure
The turn contract (one `fleet_reply`, `fleet_ask` for the lead's decisions, honest reporting,
never merge) is in **`CLAUDE.md` → Bridge communication → Worker** and already applies.
**This skill is not `reviewer`.** `reviewer` judges one diff and reports the *single* most
important issue in about 90 words. A hunt sweeps a whole package and reports *several* findings
in a long structured form. Loading both gives you two contradictory output contracts, and the
usual result is a worker that writes a good report into its terminal and ends the turn without
sending it. Load exactly one.
## 0. Read this before you read code: how the report gets home
Your terminal reaches nobody. The lead sees **only** the text inside your `fleet_reply` call.
A long report is exactly the case where this goes wrong, so plan for it:
- **Write the report into the `fleet_reply` argument itself.** Do not compose it in your terminal
and then summarise it into the call.
- If the report is long, **send it anyway** — one `fleet_reply` with everything.
- If you end the turn without replying, the bridge scrapes your pane instead. That scrape carries
at most the last 4000 characters, and on a hunt it usually captures the tail of the lead's own
brief rather than your findings. The lead then has nothing and has to ask you again.
## 1. Change nothing
A hunt is read-only. Do not edit a production file, do not "quickly fix" what you find, and do
not run a formatter. You may run the build and tests to *check* a claim, and you should say so
when you did.
## 2. Read the whole scope first
Read every file in the assigned package before you judge any of it. A defect that a caller
elsewhere in the same package makes unreachable is not a defect, and you cannot know that from
one file.
Stay inside the scope. If a defect there depends on a class outside it, read that class to
confirm — but the defect itself must live in the scope you were given.
## 3. The bar — this matters more than the count
**Name the path into the bad state.** Say which caller, in which state, reaches it. A defect on
paper is not a reachable defect. If you cannot name that path, keep the finding but mark it
`unproven` and say exactly what you could not check. Do not drop it, and do not dress it up.
**Say which direction the harm goes.** Data loss, privilege escalation and silent wrong answers
are worth reporting even when the window is narrow. A finding whose worst outcome is a worse log
line is not worth a block.
Two workers once ran the same scope: the one that applied the direction-of-harm filter found ten
real defects, the one that did not found none. Fewer findings the lead can act on beat many the
lead has to triage.
## 4. Shapes that have produced real merged fixes here
Read for these first:
1. **A one-way gate.** A guard added after an incident closes only the direction that incident
came from. Do not only ask what closes the gate — ask **which states still open it**.
2. **A value read once, then used later to authorise something destructive**, after something
else has had a chance to change it.
3. **A failure downgraded to a value that looks like a legitimate result** — `-1`, `null`, an
empty list, `false` — which a caller then trusts.
4. **A lock held for one half of a read-modify-write and not the other**, or two collections
updated under different locks.
5. **A comment or javadoc stating an invariant the code no longer keeps.** Comments are
load-bearing in this repo; a stale one has already caused a bug.
## 5. What you cannot check, and must not claim you did
- `fleetd/fleetd.yaml` is gitignored and **absent from your worktree**. You cannot read it. If a
finding depends on live configuration, name the key and say you could not check it.
- `.mcp.json`, `opencode.json` and `.autoenv` in your worktree are neutralised stubs, not the
repo's real files.
- The `wiki/` submodule pointer is months old. Do not cite it.
Reporting a fact you took from the lead's brief as something you measured yourself is a false
report, even when the fact is correct. Say where each fact came from.
## 6. The report — what goes in `fleet_reply`
One block per finding, most severe first:
```
FINDING N — <one line>
file:line
Path in: <which caller, in which state, reaches this>
Direction: <data loss | escalation | silent wrong answer | outage | ...>
Window/trigger: <when it actually happens>
Confidence: <confirmed by reading | unproven — say what you could not check>
Why nothing else catches it: <the guard or test you checked, and why it misses>
```
End with one line naming every file you read, so the lead knows the denominator.
**Nothing clears the bar?** Reply `NO FINDINGS`, name the files you read, and say what you ruled
out. A clean sweep is a valid result; an invented defect is worse than none.
+6 -19
View File
@@ -1,11 +1,11 @@
---
name: implementer
description: Implementer-role procedure for a fleetd worker — verify your worktree, implement the scope, commit, push, open your own PR, and hand off the PR URL. Load this when the lead delegates you an implementation task over fleetd.
description: Implementer-role procedure for a bridged worker — verify your worktree, implement the scope, commit, push, open your own PR, and hand off the PR URL. Load this when the lead delegates you an implementation task over bridged.
---
# Implementer worker — procedure
The turn contract (one `fleet_reply`, `fleet_ask` for the lead's decisions, honest reporting,
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
never merge, never commit `.mcp.json` or `wiki/`) is in **`CLAUDE.md` → Bridge communication →
Worker** and already applies. This skill is only the *implement-and-hand-off procedure*.
@@ -39,19 +39,6 @@ a worker made all 59 of its edits in the primary's tree and never noticed.
test "$(git rev-parse --show-toplevel)" = "$PWD" || cd "$(git rev-parse --show-toplevel)"
```
**Never run `git stash` (or `git stash pop`/`apply`/`drop`).** Your worktree is isolated, but the
stash is **not**: `refs/stash` is one stack shared by the primary's checkout and every other
worker's worktree of this repo. Measured on 2026-09-04 — `git stash list` from a worker's worktree
and from the primary's tree returned byte-identical output. So a `git stash` you run can be popped
into someone else's tree, and a `git stash pop` you run can drop **another worker's** uncommitted
edits on top of yours. This has already happened here: two workers were running in parallel and one
of them had its in-progress edit silently overwritten by the other's stash.
The branch is your isolation, so use it instead. To set work aside, commit it on your own branch
(`git commit -m "wip: ..."`) and carry on; to try something and back out, use
`git diff > /tmp/<your-branch>.patch` then `git checkout -- <file>`. Both stay inside your worktree.
If you find a stash entry you did not create, leave it alone and say so in your report.
## 2. Implement
- Implement exactly the scope the lead named. Keep the diff focused; note anything out of scope
@@ -62,7 +49,7 @@ If you find a stash entry you did not create, leave it alone and say so in your
your worktree*:
```bash
cd "$(git rev-parse --show-toplevel)/fleetd" && mvn clean install
cd "$(git rev-parse --show-toplevel)/bridged" && mvn clean install
echo "exit=$?"
```
@@ -98,7 +85,7 @@ host (`GITEA_HOST`) into your env for exactly this — the token can create a PR
merge**.
```bash
API="${GITEA_HOST%/}/api/v1/repos/fleet/fleetd/pulls"
API="${GITEA_HOST%/}/api/v1/repos/lms/claude-bridge/pulls"
BRANCH="$(git branch --show-current)"
curl -sS -X POST "$API" \
-H "Authorization: token ${GITEA_TOKEN}" \
@@ -116,7 +103,7 @@ fix it if the cause is yours (e.g. branch not pushed yet), and report the failur
inventing a URL. If `GITEA_TOKEN` is unset your profile was not granted PR-create: push the branch
and report its name so the lead opens the PR.
## 6. Hand off — what goes in `fleet_reply`
## 6. Hand off — what goes in `bridge_reply`
The reply is the entire handoff; the lead cannot see your terminal.
@@ -143,7 +130,7 @@ sequenceDiagram
I->>G: git push -u origin HEAD
I->>G: POST /pulls (GITEA_TOKEN) — open PR to main
G-->>I: html_url
I->>L: fleet_reply(PR url, branch, files, tests)
I->>L: bridge_reply(PR url, branch, files, tests)
Note over L,G: the lead reviews the PR and merges on green — you never merge
```
+1 -1
View File
@@ -53,7 +53,7 @@ Map each entry from `.mcp.json`:
"$schema": "https://opencode.ai/config.json",
"instructions": ["CLAUDE.md"],
"mcp": {
"fleetd": { "type": "remote", "url": "http://127.0.0.1:8765/mcp", "enabled": true },
"bridged": { "type": "remote", "url": "http://127.0.0.1:8765/mcp", "enabled": true },
"context7": { "type": "remote", "url": "https://example.dev/mcp", "enabled": true,
"headers": { "Authorization": "Bearer {env:CONTEXT7_TOKEN}" } },
"gitea": { "type": "local", "command": ["gitea-mcp", "-t", "stdio"], "enabled": true,
-72
View File
@@ -1,72 +0,0 @@
---
name: redeploy-fleetd
description: Rebuild and restart the live fleetd daemon after a merge (lead / primary only). Load this before redeploying — it holds the script, the drain step, the permission grant, and the five checks that have each gone wrong here before. Workers must never do this.
---
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-fleetd.sh --check # report state, change nothing
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
```
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
+4 -9
View File
@@ -1,20 +1,15 @@
---
name: reviewer
description: Reviewer-role procedure for a fleetd worker — how to work a review scope and the exact shape of the finding to report. Load this when the lead delegates you a code review over fleetd.
description: Reviewer-role procedure for a bridged worker — how to work a review scope and the exact shape of the finding to report. Load this when the lead delegates you a code review over bridged.
---
# Reviewer worker — procedure
The turn contract (one `fleet_reply`, `fleet_ask` for the lead's decisions, honest reporting,
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
never merge) is in **`CLAUDE.md` → Bridge communication → Worker** and already applies. This
skill is only the *review procedure*: how to work the scope, and the exact shape of what you
send back.
**Wrong skill for a sweep.** This one reviews *one* diff or scope and reports the *single* most
important issue. If the lead asked you to hunt or audit a whole package for several defects, load
`hunter` instead and ignore this file — the two want different output, and following both is how a
worker ends its turn with a good report that never gets sent.
## 1. Read the whole scope before you judge
The delegation names your scope — a file, a diff, a PR, a function. **Read all of it first.**
@@ -29,13 +24,13 @@ wrong.
covers what it was given.
- Do **not** edit files or run the build. You review; the owner acts.
## 3. Reach for `fleet_ask` only for a genuine fork
## 3. Reach for `bridge_ask` only for a genuine fork
Ambiguous requirement, a missing acceptance criterion, "intended or a bug?", or two defensible
fixes with different consequences — those are the lead's call, and guessing produces a
confident-but-wrong finding. Anything you could settle by reading more code is yours to settle.
## 4. The finding — what goes in `fleet_reply`
## 4. The finding — what goes in `bridge_reply`
Report the **single most important** real issue in the scope, in these four lines, under
~90 words:
+12 -37
View File
@@ -14,7 +14,7 @@ jobs:
# does not depend on the wiki repo being reachable.
- uses: actions/checkout@v4
# The runner image ships an older default-jdk; fleetd sets maven.compiler.release=25, so
# The runner image ships an older default-jdk; bridged sets maven.compiler.release=25, so
# provision the JDK explicitly rather than apt-installing whatever "default" means today.
- name: Set up JDK 25
uses: actions/setup-java@v4
@@ -32,17 +32,13 @@ jobs:
mvn -version
- name: Build and test
working-directory: fleetd
working-directory: bridged
# This IS the mock-socket surface CB-503 asks for: the pom's `default-excludes` profile
# already sets excludedGroups=contract, so the @Tag("contract") tests — which need a live
# herdr socket and a RabbitMQ container — are excluded without any flag here. Everything
# that runs does so against the fake UDS herdr and fake ccs/claude stubs.
run: mvn -B clean install
- name: javadoc reference lint
working-directory: fleetd
run: mvn -B -DskipTests javadoc:javadoc -Ddoclint=reference
# Deliberately NOT actions/upload-artifact: this Gitea instance presents as GHES, and
# @actions/artifact v2+ (i.e. upload-artifact@v4) refuses to run there —
# "GHESNotSupportedError ... not currently supported on GHES", which red-Xes an otherwise
@@ -50,7 +46,7 @@ jobs:
# the log instead, where they are actually readable.
- name: Failing test output
if: failure()
working-directory: fleetd
working-directory: bridged
run: |
for f in target/surefire-reports/*.txt; do
[ -f "$f" ] || continue
@@ -58,29 +54,12 @@ jobs:
done
exit 0
# fleetd #550 — nothing ran scripts/test-redeploy-fleetd.sh in CI before this, on any platform,
# so it had run only on macOS by hand and two Linux-only bugs (this issue's items 1 and 2)
# survived undetected: shasum is a macOS-only tool (it ships with Perl; GNU coreutils, i.e. every
# mainstream Linux distro including this runner's ubuntu-latest, does not have it and ships
# sha256sum instead). The gate here is the step's own exit code, nothing else: a `run:` step in
# Gitea/GitHub Actions already fails the job on a non-zero exit with no extra scripting needed,
# so this deliberately does NOT grep the output for a `FAIL:` count. That is the #550 item-2
# lesson one level up — a suite that dies before it runs a single test prints zero FAIL lines,
# which is exactly what a clean pass also prints, so counting FAIL lines can never be the gate.
shell-tests:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: redeploy-fleetd.sh shell suite
run: bash scripts/test-redeploy-fleetd.sh
# CB-521 — actually run the AMQP contract test in CI, against a REAL broker. The broker is a
# RabbitMQ SERVICE CONTAINER, not Testcontainers-with-Docker: the runner image has no Docker, so
# AmqpReplyInboxContractTest reads AMQP_URI (set below to the service's network alias) and binds
# straight to it — no Docker, no skipped tests. This separation (build job hermetic and
# Docker-free; contract job broker-provided) is deliberate — see the default-excludes/contract
# profiles in fleetd/pom.xml. `setup-java` provides the JDK only; Maven is installed separately,
# profiles in bridged/pom.xml. `setup-java` provides the JDK only; Maven is installed separately,
# exactly as in the build job above.
contract:
runs-on: ubuntu-latest
@@ -90,6 +69,8 @@ jobs:
env:
RABBITMQ_DEFAULT_USER: guest
RABBITMQ_DEFAULT_PASS: guest
ports:
- 5672:5672
env:
# Service containers are reachable from the job by their network alias on their internal port.
AMQP_URI: amqp://guest:guest@rabbitmq:5672
@@ -108,22 +89,16 @@ jobs:
apt-get update && apt-get install -y --no-install-recommends maven
mvn -version
# The `contract` profile clears the default-excludes group, so `-Dgroups=contract` runs every
# @Tag("contract") test and nothing from the unit suite the `build` job already covered — a
# tag selects the whole group, so a test added to it later runs here automatically. A prior
# version of this step pinned `-Dtest=AmqpReplyInboxContractTest` by class name instead: that
# silently excluded every other contract test (including the herdr ones) from CI, and nobody
# noticed until the herdr protocol drifted out from under a test that never ran here
# (fleetd #449). If this runner has no herdr socket, the herdr-backed tests in the group
# skip on their own `assumeTrue` and only the broker-backed ones actually run — check the
# step output rather than assuming which.
# The `contract` profile clears the default-excludes group, so the @Tag("contract") AMQP test
# runs against the RabbitMQ service container (AMQP_URI). Pinned to the one contract test to
# avoid re-running the unit suite already covered by the `build` job.
- name: Contract tests
working-directory: fleetd
run: mvn -B -Pcontract test -Dgroups=contract
working-directory: bridged
run: mvn -B -Pcontract test -Dtest=AmqpReplyInboxContractTest
- name: Failing test output
if: failure()
working-directory: fleetd
working-directory: bridged
run: |
for f in target/surefire-reports/*.txt; do
[ -f "$f" ] || continue
+4 -10
View File
@@ -14,14 +14,8 @@
.env
.envrc
# Daemon runtime artefacts. fleetd appends its log wherever it is launched from, so both the
# repo root and fleetd/ collect one; neither belongs in git.
fleetd.out
fleetd/fleetd.out
# Daemon runtime artefacts. bridged appends its log wherever it is launched from, so both the
# repo root and bridged/ collect one; neither belongs in git.
bridged.out
bridged/bridged.out
logs/
# fleetd #480: the lead rollover handover file. `leadRollover.handoverPath` points here, and the
# outgoing lead rewrites it on every rollover. It is a snapshot of one moment's live state —
# unpushed branches, running builds, open questions — so it is stale the moment it is written and
# has no business in git history.
.handover/
+1 -1
View File
@@ -1,3 +1,3 @@
[submodule "wiki"]
path = wiki
url = ssh://git@git.ltms.dev:2224/fleet/fleetd.wiki.git
url = ssh://git@git.ltms.dev:2224/lms/claude-bridge.wiki.git
+2 -2
View File
@@ -3,7 +3,7 @@ description: Refine work into clear, independent units before implementation.
mode: primary
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You are an architect in this fleet. You refine work before anyone builds it: scope,
acceptance criteria, risks, and a unit split. You read the repo and write analysis.
@@ -13,7 +13,7 @@ A design task is worked by two architects. Design alone first, then exchange and
say plainly where you disagree. Do not concede just to agree.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
investigate it further. Use `bridge_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
+2 -2
View File
@@ -3,14 +3,14 @@ description: Implement one assigned unit, test it, and open a pull request.
mode: primary
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You implement the one unit you were given and nothing else. Work in your assigned
git worktree and branch. Never check out, rebase onto, or push to `main`. Confirm
the worktree root and branch before you edit. Use only paths under that root.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
investigate it further. Use `bridge_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
+2 -2
View File
@@ -3,7 +3,7 @@ description: Review one assigned scope and report the most important real issue.
mode: primary
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
You review the diff you were given. Report bugs, risks, and missing tests. You do
not change code.
@@ -12,7 +12,7 @@ Read the whole assigned scope before judging it. Review only that scope. If you
something outside it, note it in one line and do not investigate it further. Do not
run the build. The owner makes changes and runs checks.
Use `fleet_ask{question}` only when a decision belongs to the lead, such as an
Use `bridge_ask{question}` only when a decision belongs to the lead, such as an
unclear requirement or two defensible fixes. Do not ask about something you can
decide by reading more code.
+122 -211
View File
@@ -4,43 +4,35 @@
> **Canonical block.** Everything down to §Layering is the portable bridge charter, copied verbatim
> into every project that mounts the bridge MCP. Keep it byte-identical with the template in the
> wiki ([Use Cases](https://git.ltms.dev/fleet/fleetd/wiki/7-Use-Cases) → *The portable
> wiki ([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable
> CLAUDE.md block*); improvements go to the template first, then out to each project. Anything
> specific to *this* repo lives under §Project addendum below, never inline above it.
>
> **Anything you measure in an addendum is perishable.** Date it, give the command that
> re-measures it and what each outcome means, and tell the reader to delete the section once
> it stops reproducing. The four parts work together: deciding what would falsify a claim is
> the expensive step, and a reader in the middle of another task will not pay it, so a bare
> "verify before relying on this" costs the same space and does nothing. The case this is for
> is a note that goes stale as a live restriction — it will tell a future session it cannot do
> the thing at the moment doing it becomes the job.
If no `fleet_*` MCP tools are mounted in this session, this section does not apply — skip it.
If no `bridge_*` MCP tools are mounted in this session, this section does not apply — skip it.
`fleetd` is the **sole communication gateway** between agents here. The orchestrating session (the
`bridged` is the **sole communication gateway** between agents here. The orchestrating session (the
**primary**) and every delegated peer (a **member**) mount the *same* MCP server and talk only
through its `fleet_*` tools. No session addresses a peer, a broker, or the network directly.
through its `bridge_*` tools. No session addresses a peer, a broker, or the network directly.
### Which role am I? — settle this before acting
**Every role reads this file.** A member runs in a git worktree of this same repo, so it inherits
this `CLAUDE.md` verbatim, and every rule below is role-conditional.
**Call `fleet_whoami`.** It returns `primary`, `worker`, or `architect`, resolved by the daemon from
**Call `bridge_whoami`.** It returns `primary`, `worker`, or `architect`, resolved by the daemon from
your connection — unforgeable, and the same resolution its authorization gate uses. A worker also
carries its `sessionId`, `profile`, `worktree` and `branch`; an architect carries the slot name it
was bound to. Don't infer what you can ask.
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
fires: the reply charter in your system prompt (*"You are a spawned member in the
claude-bridge fleet"*) ⇒ **spawned member**; fleet tools prefixed `mcp__fleet__*` ⇒ **spawned
fires: the reply charter in your system prompt (*"You are an off-subscription worker in the
claude-bridge fleet"*) ⇒ **spawned member**; bridge tools prefixed `mcp__bridge__*` ⇒ **spawned
member** (the launcher fixes that mount name; a primary's mount is named by whoever wrote its
`.mcp.json`, so it varies — and a member spawned before CB-632 still says `mcp__bridge__*`); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
`.mcp.json`, so it varies); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect —
only `fleet_whoami` does. **Still unsure ⇒ act as a worker**, the most restricted member role. The
only `bridge_whoami` does. **Still unsure ⇒ act as a worker**, the most restricted member role. The
two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate —
loud and self-correcting — while a member acting as the primary ends its turn with no `fleet_reply`,
loud and self-correcting — while a member acting as the primary ends its turn with no `bridge_reply`,
and the sender silently receives nothing. Fail toward the recoverable error.
### Invariants — both roles, no exceptions
@@ -49,7 +41,7 @@ and the sender silently receives nothing. Fail toward the recoverable error.
stays on subscription; only the bridge puts a member off it, at spawn. Mounting the bridge must
never move a session across that boundary.
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `fleet_*` call is silently discarded.
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead or architect**;
reply/ask are only-as-itself — any peer may answer for its own pane, and for no other. A call
@@ -58,18 +50,17 @@ and the sender silently receives nothing. Fail toward the recoverable error.
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
deliverable, and a send waits on that gate for ~60s and then fails without ever reaching its pane.
5. **Never move a fleet session, pane or peer except through the bridge.** The bridge owns policy;
the multiplexer owns PTYs. Any route that changes fleet state without the bridge's checks
bypasses every rule above — the `herdr` CLI and its socket are the usual example.
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
### Primary (lead) — run this on every task, in order
**Delegate by default — that is the job.** With the bridge mounted you are an orchestrator on a
metered subscription, and workers are cheap, parallel, and disposable. The default answer to "who
does this?" is **a worker**, not you. Reach for `fleet_send` before you reach for `Edit`. The steps
does this?" is **a worker**, not you. Reach for `bridge_send` before you reach for `Edit`. The steps
below are the procedure — run them in order, every task, not only the big ones.
0. **Know your role** — `fleet_whoami`, once per session, before anything else.
0. **Know your role** — `bridge_whoami`, once per session, before anything else.
1. **Split.** Write the unit list. Every unit carries: scope · the files or PR in question ·
acceptance criteria · exactly what to report back. A unit with no acceptance criteria is not
ready to delegate — refine it or keep it.
@@ -79,37 +70,22 @@ below are the procedure — run them in order, every task, not only the big ones
delegate. The keep-list is closed: the conversation with the user, decomposition and planning,
the final judgment call, verification, merges, and anything that depends on context only you
hold. Nothing else is yours by default.
3. **Spawn every delegated unit first** — `fleet_spawn{profile, worktree:true, ticket}`, one per
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model, cost and
LIVENESS, not in tier, so the default is rarely what you want. The default is whatever the
daemon reports, and on a host where it sits on an exhausted or withdrawn credential every
unqualified spawn fails — sometimes loudly, sometimes as a member that spawns fine and then
produces nothing. `fleet_profiles` reports the default; check it once per session.
4. **Then send them all** — `fleet_send{sessionId, content, wait:false}`. Line 1 of every brief is
3. **Spawn every delegated unit first** — `bridge_spawn{profile, worktree:true, ticket}`, one per
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model and cost, not in
tier, so the default is rarely what you want.
4. **Then send them all** — `bridge_send{sessionId, content, wait:false}`. Line 1 of every brief is
`Load the <name> skill.` naming the worker's playbook; those skills are opt-in and that line is
what makes them reliable. Where the project ships no such skill, spell the procedure out in the
brief instead. The brief is self-contained — the worker sees your message and the repo, nothing
of your context, your plan, or your screen.
5. **Collect** — `fleet_poll{ticket}` → `fleet_ack{target, msgId}`. Answer a worker's `fleet_ask`
with `fleet_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`fleet_status`, never by reading its terminal; it also reports an open question and the `turnId`
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{target, msgId}`. Answer a worker's `bridge_ask`
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`bridge_status`, never by reading its terminal; it also reports an open question and the `turnId`
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
**A correction cannot reach a busy member.** A `fleet_send` to a working member is *accepted* and
returns a ticket, and is then never delivered — measured here three times in one session, and the
member was released still executing a brief that had been retracted twice. The receipt is true and
it is a fact about the *mailbox*; what you needed was a fact about the *pane*. **A push delivery
needs the recipient free at send time; a pull channel needs only that they look before acting.** So
put every correction on the **ticket**, which they can read whenever they look, and send as the
notification. That obliges you, not them: **all corrections go to the ticket, and the brief is
write-once.** The member cannot check which source is newer — it just always prefers the ticket —
so the day you revise a brief in place instead of commenting, it obeys your rule and does the wrong
thing. A *first* brief for a unit not yet running is not a correction, and may be the whole spec.
6. **Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
forge MCP server it appears to have holds a blocked credential and fails every call, and a piped
command (`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a
fact. Its injected repo-scoped `GITEA_TOKEN` is a different credential and does work, so a worker
reporting that it opened its own PR is reporting something it really can do.
forge tools it appears to have hold a blocked credential and fail, and a piped command
(`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a fact.
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
@@ -117,59 +93,33 @@ below are the procedure — run them in order, every task, not only the big ones
read it yourself.
8. **Adjudicate, merge, tear down — yours alone.** Read the diff yourself: fully if it is small,
targeted at the reported findings and the risky paths if it is large. Reviewer findings direct
your attention; they never substitute for it. Then merge, then `fleet_stop{paneId}`.
**If the forge refuses you the merge** — a protected branch, a token without the grant — the
adjudication is still yours. Read the diff, decide, and hand the operator a merge-ready queue
with the refusal quoted. Never report a PR as merged, and never call one "ready to merge"
without having read the diff yourself. A refusal is exactly when that shortcut is tempting,
because no action is left that forces you to look, and taking it turns this step into
forwarding a reviewer's verdict — which is delegating the merge by proxy, two lines above.
**Test a refusal; do not read it off a permissions field.** A protected branch holds its merge
rights separately from the repository permissions, so that field can say yes while the merge is
refused, and still say no after a grant makes it work. Probe instead, with a request that cannot
succeed on its merits, so a rejection can only mean the refusal. Treat a transport failure as a
third answer that proves nothing: a timeout, a DNS error or a bad URL is not a refusal, and
counting it as one makes you sure of something you never measured.
your attention; they never substitute for it. Then merge, then `bridge_stop{paneId}`.
**Steps 3 and 4 are separate on purpose** — spawning and sending in one loop is how parallel work
silently becomes serial, and it is the most common way this layer is wasted. For the same reason,
prefer `wait:false` + `fleet_poll` for anything non-trivial: a blocking `fleet_send` is capped by
prefer `wait:false` + `bridge_poll` for anything non-trivial: a blocking `bridge_send` is capped by
*your own* MCP client call timeout (~60s), well below the task's real runtime.
**Delegating does not delegate responsibility.** Workers open PRs; you are the gate. Never delegate
the merge — and merging on a reviewer's word is delegating it by proxy.
**When a decision blocks you, consult architects — not the operator.** Spawn one or more architect
members, give them the question and the evidence you have, and act on what they agree. They are
authorized to settle it, not only to advise. Architects first form independent positions, then
compare them. If they still disagree after that comparison, they return both positions and their
checked evidence; the lead decides. Go to the operator only for an action the fleet has no
authority to take, such as spending money, granting access, or making a promise to someone else.
**Then write the decision on the ticket.** Taking the operator out of the loop also removes the
signal they used to get, because that signal was the block itself — work stopped, so they found
out. A ticket comment replaces it, and it reaches them whether or not they are at a terminal when
you decide.
| Intent | Tool |
|---|---|
| Confirm your own role | `fleet_whoami` |
| See backends available | `fleet_profiles` |
| Start a member | `fleet_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` → `sessionId` + `paneId` |
| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) + `loopHealth` (`RUNNING`, `STALLED`, or `STOPPED` for `statusPoller` and `sessionReaper`) · one peer's state: `fleet_status{sessionId}` |
| Delegate (blocking) | `fleet_send{sessionId, content}` |
| Delegate (long task) | `fleet_send{sessionId, content, wait:false}` → ticket → `fleet_poll{ticket}` |
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}` if they are on this host. **Not** `fleet_reply`: it has no peer route and the publish is refused |
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
| Tear down a member | `fleet_stop{paneId}` |
| Replace your OWN lead session when its context is full | `fleet_handover{action:"open", reason?}` → write the handover file it names → `fleet_handover{action:"confirm", token, operatorConfirmed}`. Primary-only. **In that order**: the file must be modified *after* `open`, or `confirm` refuses it as stale. There is no terminal parameter — the pane is always your own, so you can never roll another lead. `{action:"cancel", token}` drops a pending request |
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) · one peer's state: `bridge_status{sessionId}` |
| Delegate (blocking) | `bridge_send{sessionId, content}` |
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
| Answer a member's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** | `bridge_send{sessionId: <their terminal>, content}` — `bridge_list` → `leads` reports it. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `bridge_reply{content}` — the one case a lead replies |
| Collect a held reply | `bridge_poll{target}` · then `bridge_ack{target, msgId}` |
| Tear down a member | `bridge_stop{paneId}` |
### Lead ↔ lead — coordinate, never delegate
`fleet_list` returns `leads` alongside `members`; your own row carries `self: true`. Every other row
`bridge_list` returns `leads` alongside `members`; your own row carries `self: true`. Every other row
is a peer — an orchestrator with its own context, its own members, and its own judgment. An empty
`members` array means no members are spawned; it says nothing about peers.
@@ -189,56 +139,39 @@ The traffic between leads is coordination and nothing else:
3. **Verify a peer exactly as you verify yourself.** Peer status buys nothing: check the claim
against the code, and re-run the build. A peer's correction gets the same treatment — right or
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
**N observations are N data points only if they differ in the axis you are trusting.** This cuts
both ways. N *failures* blamed on one cause are one data point when the cases share what you are
not varying. N *agreeing measurements* are also one data point when they share an instrument —
two hosts, two operators and the same formula is one formula, not two confirmations.
4. **Ask a peer to read your project addendum.** Your addendum is instruction surface: every future
session on your host obeys it, and a wrong one is obeyed just as faithfully as a right one. The
author is the worst reader of their own qualifier placement — measured here, one addendum carried
two defects and a non-author found both. If you have no peer, at least re-read it asking "which
sentence goes false first, and would a reader reach the caveat before acting?"
Being messaged by a peer does not make you its worker: answer the way you would open —
`fleet_send{coordId}` for another daemon, `fleet_send{sessionId}` on this host — and push back on
the substance if it is wrong. `fleet_reply` resolves a member's blocked `fleet_send`; a peer's
coord-id message is durable and non-blocking, so there is nothing for it to resolve. A peer that
simply complies has thrown away the reason there are two of you.
Being messaged by a peer does not make you its worker: answer with `bridge_reply`, and push back on
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
you.
### Member (worker or architect) — the turn contract
1. **Load the playbook skill the lead named** before doing anything else.
2. **Do the assigned scope only.** Note anything you spot outside it in one line; don't go hunt it.
3. **`fleet_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
3. **`bridge_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
defensible fixes, "bug or intended?"). It blocks and you resume the *same* turn with the answer.
Don't ask what you could decide yourself.
4. **Re-read the ticket before you act on anything you were told earlier**, and again before you
commit. A message reaches you only while you are free to receive it; the ticket is there whenever
you look. **If a ticket comment contradicts your brief, the ticket comment is newer and it wins.**
5. **End the turn with exactly one `fleet_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `fleet_reply` ⇒ the sender gets nothing and the exchange stalls.
4. **End the turn with exactly one `bridge_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `bridge_reply` ⇒ the sender gets nothing and the exchange stalls.
Do **not** lean on the completion fallback to carry your answer for you: when you end a turn
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
lead with its end cut off.
6. **Report honestly.** State only what you actually ran and its real output, including failures,
5. **Report honestly.** State only what you actually ran and its real output, including failures,
and never claim the result of a check you had no way to run. **Measure your own tools; do not
assume them.** What you mount depends on your backend: an opencode member gets the bridge and
nothing else, while a Claude Code member also inherits the operator's user-scope MCP servers,
which the bridge never chose for you. Two rules follow. The primary's IDE tooling is still not
yours, whatever you see. And **a mounted tool is not a working tool** — the forge MCP server you
may find there holds a deliberately blocked credential and fails every call, by design. That is
not your only forge route, and the two must not be confused: the repo-scoped `GITEA_TOKEN` the
daemon injects into your environment does work, and using it to open your own PR is part of the
job. A blocked MCP tool is never a reason to skip that step.
7. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
yours, whatever you see. And **a mounted tool is not a working tool** — the forge server you may
find there holds a deliberately blocked credential and fails every call, by design.
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
project marks as not-yours-to-commit.
### Where each rule lives (don't duplicate — extend the right layer)
| Layer | Scope | Reaches |
|---|---|---|
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `fleet_reply`* | every spawned member, at launch, every peer kind — never a lead |
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every spawned member, at launch, every peer kind — never a lead |
| **this section** | protocol + orchestration policy | primary **and** every member that reads the repo — tracked in git, so worktrees inherit it |
| role agent definition files | role contract and per-job procedure | a member whose launcher binds its role to the matching file in its worktree |
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a member told to load one |
@@ -251,90 +184,78 @@ must obey belongs in the charter, not here.
## Project addendum — claude-bridge (not part of the canonical block)
- **This repo is the bridge.** The daemon is `fleetd`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/FleetMcp` (tools), `auth/Authz` (the role table),
- **This repo is the bridge.** The daemon is `bridged`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/BridgeMcp` (tools), `auth/Authz` (the role table),
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
- **Herdr socket tests (measured 2026-09-10).** In this repo, herdr is a subject under test. A
worker assigned to herdr code, and the lead, may let a test open the herdr socket directly in a
throwaway workspace that the test tears down. This only covers
`fleetd/src/test/java/dev/ltms/fleet/herdr/AgentControlContractTest.java`,
`fleetd/src/test/java/dev/ltms/fleet/herdr/HerdrContractTest.java`,
`fleetd/src/test/java/dev/ltms/fleet/herdr/PaneLocatorContractTest.java`, and
`fleetd/src/test/java/dev/ltms/fleet/herdr/WorkspacePlacementContractTest.java`. It is not a
general licence. Using the herdr CLI or socket to move a real fleet session, pane, or peer stays
banned. That is the control plane that invariant 5 protects. Re-measure with
`grep -rl 'UnixSocketHerdrClient.connect()' fleetd/src/test/java --include='*.java'`. A non-empty
result means tests still open the socket and this note still applies. An empty result means nobody
does this any more; delete this section. Canonical invariant 5 restatement is tracked in #458 and
is not part of this change.
- **`fleet_profiles`/`fleet_list` report two separate outage states, and they are not the same
thing.** *Quarantined* (CB-578) means the backend told us it is out of capacity — a long,
1800s-default cooldown. *Cooling off* (fleetd #201/#227) means a profile's credential threw two
distinct backend errors (a non-exhaustion failure such as an HTTP 5xx) within 60 seconds — a
short, fixed 60s cooldown, not configurable per profile. Each check runs independently, so a
profile can show both at once. In the JSON: a cooling profile carries `credentialId` and
`coolingOffForSeconds`; a quarantined profile carries `quarantinedForSeconds`; a profile hit by
both carries all three fields, and either state alone already sets that profile's `free` to `0`.
A `fleet_spawn` naming a cooling-off profile is refused before it ever reaches the backend
adapter, with a message naming the credential and the remaining seconds ("cooling off after
repeated backend errors") — distinct wording from a quarantine refusal, so don't conflate the
two when reading a spawn failure.
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR),
`reviewer` (one diff → one structured finding) and `hunter` (sweep a package → several ranked
findings, change nothing). Name exactly one in every delegation. **`reviewer` and `hunter` are
not interchangeable** — `reviewer` caps the answer at one finding in about 90 words, so naming
it for a multi-finding sweep hands the worker two contradictory output contracts. That has
already cost three workers' turns: each wrote a good report to its terminal and ended the turn
with no `fleet_reply`, and the scrape returned the tail of the brief instead.
Spawn `implementer` with role `dev`, `reviewer` with role `reviewer`, and `hunter` with role
`hunter`.
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
`reviewer` (scoped review → one structured finding). Name one in every delegation.
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
`port-to-opencode` (make an OpenCode session a participant in this workspace),
`fleets-status` (report every fleet that shares one LavinMQ instance),
`redeploy-fleetd` (rebuild and restart the live daemon after a merge) and
`handover` (write the file a fresh lead session inherits when the outgoing one hands off,
fleetd #480).
- **This repo is also a Claude Code marketplace, and ships a plugin.** `.claude-plugin/marketplace.json`
points at `plugin/`, which carries the MCP mount and the `setup` skill
(`/claude-bridge:setup` — make any project bridge-ready). It was added in CB-527 and then went
unmentioned by every instruction file, so it drifted and a later session planned it from scratch
(#362). **Read `plugin/` before designing anything about onboarding a project.** Two limits are
structural, not bugs: a plugin cannot carry the role agent files, because
`ClaudeCodeLauncher.java:371` requires `<cwd>/.claude/agents/<role>.md` in the member's own
worktree; and a plugin cannot deliver anything to members at all, because
`ClaudeCodeLauncher.java:285` exports `CLAUDE_CONFIG_DIR` and every Claude profile here sets it,
so a member never reads the operator's plugin store. **The plugin is the lead-side surface;
member-facing assets travel in the worktree.**
`port-to-opencode` (make an OpenCode session a participant in this workspace).
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **A provisioned worktree neutralizes `.mcp.json`, `opencode.json` and `.autoenv`** — the repo's
committed copies would otherwise mount the primary's IDE and forge servers (fleetd #134). The
worktree's copy of each is a stub, **not** the repo's real file, so a worker that reads one and
reports what it found is reporting on the stub. The daemon logs a per-spawn summary, but the
worker cannot see that log. From inside its own worktree a worker — or a lead debugging one —
reads the list with `git config --worktree --get-all fleet.neutralizedConfig`, and the
consequence with `git config --worktree --get fleet.neutralizedConfigNote`. Never brief a worker
to edit one of these files: the edit cannot be committed, and it will not tell you so.
- **Flows and the error model** — rendezvous, `fleet_ask`, detached delivery, the turn-done
fallback and status gating — are diagrammed in `docs/MCP-Contract.md`. That page is now flows
only: its pre-build tool catalogue, parameter tables and REST paths were deleted rather than
corrected, because a hand-maintained second copy of the tool surface is what drifted for a month
while this line pointed every session at it (CB-609 / #114). **The live MCP schema is the tool
reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if
that page names a `fleet_*` tool the server does not register. The flows are kept out of this
file because this file loads into every session's context.
- **Flows and the error model** — rendezvous, `bridge_ask`, detached delivery, turn-done fallback —
are diagrammed in `docs/MCP-Contract.md` **§6 only**. The rest of that page is a pre-build design
doc whose tool names, parameter names and REST paths never caught up with the code, so do not use
it as the tool reference (CB-609). Section 6 is kept out of this file because this file loads into
every session's context.
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. The lead **may and
should** redeploy rather than hand the job back to the operator. **Workers must never do this** — a
worker has no business restarting the daemon it is talking through, and stopping it kills the
worker's own channel mid-turn.
**A merge is not a deployment.** The running `bridged` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
**Load the `redeploy-fleetd` skill before you redeploy.** It holds `scripts/redeploy-fleetd.sh`
and its flags, the drain step, the operator's permission grant, and the five checks that have each
gone wrong here before. Do not hand-roll the steps from memory.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-bridged.sh --check # report state, change nothing
scripts/redeploy-bridged.sh # build, confirm drain, restart, verify
scripts/redeploy-bridged.sh --yes # skip the drain prompt (fleet already checked)
```
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `bridge_list`, then `bridge_stop` each member, and collect anything
you still want with `bridge_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `bridge_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `bridged listening` line at the end of
`bridged/bridged.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-bridged.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
### The prompt is part of the product — update it with the code (mandatory)
@@ -347,15 +268,15 @@ Before you call any work done, check the row that matches what you touched:
| You changed… | Re-read and update… |
|---|---|
| a `fleet_*` tool — added, removed, renamed, or its params/semantics | the primary's intent→tool table; any rule that names that tool |
| a `bridge_*` tool — added, removed, renamed, or its params/semantics | the primary's intent→tool table; any rule that names that tool |
| `Authz` / the role table | invariant 3, and the primary-only vs worker-only claims |
| `ConnectionIdentity` / how a caller is resolved | the `fleet_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__fleet__*`), and the layering table's top row |
| `ConnectionIdentity` / how a caller is resolved | the `bridge_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__bridge__*`), and the layering table's top row |
| the injector / status gating | invariant 4 |
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
| a new peer kind (non-Claude adapter) | what that peer can read — anything it must obey belongs in its charter, not in the block |
| **anything an operator can use, configure, or observe** — an MCP tool, a `fleetd.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
| **anything an operator can use, configure, or observe** — an MCP tool, a `bridged.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
That last row is not bookkeeping. Chapters 1–10 answer *how is this built* and *why this way*;
none of them has a home for *what can it do and how do I turn it on*, so for twenty tickets a
@@ -366,7 +287,7 @@ is a Roadmap line. A change that touches none of the three earns no entry, and t
outcome rather than an omission.
Then **propagate**: the block in this file and the template in the wiki
([Use Cases](https://git.ltms.dev/fleet/fleetd/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
block*) must stay byte-identical, and other projects carrying the block need the same edit. Verify
rather than trust:
@@ -382,16 +303,6 @@ print("in sync:", w[i:w.index("\n```\n", i) + 1] == block)
PY
```
**Only the lead can run that check (measured 2026-09-10).** A member's provisioned worktree has
`wiki/` uninitialized, so the script dies with `FileNotFoundError: wiki/7-Use-Cases.md`. Measured
in three worker worktrees: `git submodule status` printed a leading `-` and `wiki/` held 0
entries; the primary's own clone printed a leading `+` and the file was there. So never make this
check a member's acceptance criterion — it is unsatisfiable for them, and a brief that asks for it
is asking a worker to invent a pass. A member told to check it must say it could not run it, and
must never report it as passed. The lead runs it in the main clone before merging. Re-measure with
`git submodule status` in a member's worktree: a leading `-` means this still applies; once it
prints a commit with no `-`, delete this paragraph.
## IDE MCP tools & validation workflow (enforced)
> **Primary only.** Workers have no IDE MCP mount — if you are a worker, skip this section and
@@ -399,10 +310,10 @@ prints a commit with no `-`, delete this paragraph.
Two IDE MCP servers are connected: **intellij-index** (semantic code intelligence) and
**jetbrains** (file problems, reformat, debugger). IntelliJ has multiple projects open; our
module is **`fleetd`**. Always pass these to IDE MCP tools:
module is **`bridged`**. Always pass these to IDE MCP tools:
- `project_path` = `/Users/dai.ha/LTMS/claude-bridge/fleetd`
- IDE paths are relative to `fleetd/` (e.g. `src/main/java/dev/ltms/fleet/...`)
- `project_path` = `/Users/dai.ha/LTMS/claude-bridge/bridged`
- IDE paths are relative to `bridged/` (e.g. `src/main/java/dev/ltms/bridged/...`)
### After editing any file — mandatory
@@ -416,7 +327,7 @@ module is **`fleetd`**. Always pass these to IDE MCP tools:
whole-project gate before declaring work done or committing.
**Whenever dependencies change (or a `pom.xml` edit), validate CVEs with
`jetbrains get_file_problems{filePath: "fleetd/pom.xml"}`** — its Mend.io check reflects the
`jetbrains get_file_problems{filePath: "bridged/pom.xml"}`** — its Mend.io check reflects the
dependencies on disk. (Note: `ide_diagnostics` / intellij-index does NOT re-resolve dependencies
after a pom edit without a full Maven reimport, so it reports stale CVE results — don't trust it
for this.) Treat a CVE warning like any other: bump to a patched version and confirm
+27 -29
View File
@@ -9,23 +9,23 @@ Sibling of [`crush-bridge`](https://git.ltms.dev/systems/vms) (which drives a he
process*, so it inherits `CLAUDE.md`, hooks, skills, and MCP — just pointed at a
cheaper/local model.
## Leading approach — herdr-centric message server (`fleetd`)
## Leading approach — herdr-centric message server (`bridged`)
A small always-on message server, **`fleetd`**, controls
A small always-on message server, **`bridged`**, controls
[herdr](https://herdr.dev) (an agent multiplexer) over its Unix-socket API and exposes a
clean 2-way messaging API as an **MCP server that both the primary and the workers mount** —
one unified Claude setup and the **sole communication gateway** (REST/SSE stays for non-Claude
clients; any broker is `fleetd`-internal, below the gateway).
herdr owns the PTYs, multiplexing, persistence, and **agent-status events**; `fleetd` owns
clients; any broker is `bridged`-internal, below the gateway).
herdr owns the PTYs, multiplexing, persistence, and **agent-status events**; `bridged` owns
policy (subscription boundary, session lifecycle, status-gated delivery) and the client
contract. A Claude member launches with `ANTHROPIC_BASE_URL` pointed at the gateway,
`https://llm.ltms.dev/anthropic`, plus a bearer token; the lead stays env-clean and calls
`fleetd`'s MCP tools. See the wiki's **[13 User Guide](wiki/13-User-Guide.md)** to run it.
`bridged`'s MCP tools. See the wiki's **[13 User Guide](wiki/13-User-Guide.md)** to run it.
```mermaid
flowchart LR
OPUS["Opus — primary<br/>(Claude Code, env CLEAN)<br/>MCP client"]
subgraph BD["fleetd — standalone daemon (not a claude process)"]
subgraph BD["bridged — standalone daemon (not a claude process)"]
SRV["SERVER face<br/>MCP · REST/SSE · policy"]
CLI["CLIENT face<br/>status-gated injector · herdr socket"]
SRV --> CLI
@@ -34,8 +34,8 @@ flowchart LR
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
M["llm.ltms.dev<br/>(the one gateway)"]
OPUS -->|"MCP fleet_send (blocks)"| SRV
W -.->|"MCP fleet_reply"| SRV
OPUS -->|"MCP bridge_send (blocks)"| SRV
W -.->|"MCP bridge_reply"| SRV
CLI -->|"Unix socket<br/>send_text · events.subscribe"| HERDR
HERDR -->|"drives PTY"| W
W -->|"inference"| M
@@ -47,26 +47,24 @@ flowchart LR
```
- **Subscription boundary:** the *primary* never sets `ANTHROPIC_BASE_URL` (stays on
Pro/Max). Only the *secondary* process is off-subscription — and `fleetd` itself is a
Pro/Max). Only the *secondary* process is off-subscription — and `bridged` itself is a
plain daemon (no Anthropic quota), so it may poll/subscribe freely.
- **One gateway (unified MCP setup):** `fleetd` is the **sole communication path** for every
- **One gateway (unified MCP setup):** `bridged` is the **sole communication path** for every
Claude session. Primary and workers each mount it as an MCP server (one `claude mcp add`
line, same on both) and talk over MCP tools — `fleet_send` / `fleet_reply` /
`fleet_status` (with `fleet_ask` planned for the blocked-worker path). **No Claude session
ever addresses a broker, a peer, or the network
directly**; any queue is `fleetd`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so
mounting the bridge is subscription-safe by construction.
**Tool naming:** the tools are `fleet_*` (renamed from `bridge_*` in CB-622). The old
`bridge_*` names were removed in CB-634 — only `fleet_*` answers now.
- **How the primary consumes a reply:** a single **blocking MCP call** (`fleet_send`);
`fleetd` holds it open until the worker calls `fleet_reply` or its turn hits
line, same on both) and talk over MCP tools — `bridge_send` / `bridge_reply` /
`bridge_status` (with `bridge_ask` planned for the blocked-worker path). **No Claude session
ever addresses a broker, a peer, or the network
directly**; any queue is `bridged`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so
mounting the bridge is subscription-safe by construction.
- **How the primary consumes a reply:** a single **blocking MCP call** (`bridge_send`);
`bridged` holds it open until the worker calls `bridge_reply` or its turn hits
`agent_status=done`, then returns the reply as the tool result. No cross-turn busy-poll, so
no quota burn. SSE is an optional side-channel for humans/dashboards watching status.
- **Worker → primary** rides `fleetd`'s **MCP rendezvous** — the reply resolves the primary's
blocking call (or, for detached work, `fleetd` **injects the primary's idle pane** when it's
- **Worker → primary** rides `bridged`'s **MCP rendezvous** — the reply resolves the primary's
blocking call (or, for detached work, `bridged` **injects the primary's idle pane** when it's
ready), so *no keystroke-into-primary and no broker are involved, even single-host*. The one
exception: a split-host primary that isn't a herdr pane wakes via its own `Stop`-hook, which
polls **`fleetd`** (never a broker). See the wiki for the two topologies.
polls **`bridged`** (never a broker). See the wiki for the two topologies.
- **Different model per process** sidesteps Claude Code's lack of per-subagent provider
routing — the worker isn't a subagent, it's its own configured process.
- **AgentAPI** ([`coder/agentapi`](https://github.com/coder/agentapi)) is retained only as a
@@ -75,11 +73,11 @@ flowchart LR
## Docs
Full design, setup, and operations live in the **[wiki](https://git.ltms.dev/fleet/fleetd/wiki)**,
Full design, setup, and operations live in the **[wiki](https://git.ltms.dev/lms/claude-bridge/wiki)**,
vendored here as a submodule under [`wiki/`](./wiki):
```bash
git clone --recurse-submodules ssh://git@git.ltms.dev:2224/fleet/fleetd.git
git clone --recurse-submodules ssh://git@git.ltms.dev:2224/lms/claude-bridge.git
# or, after a plain clone:
git submodule update --init
```
@@ -89,7 +87,7 @@ Gitea wiki.
## Status
🟢 **Implemented & dogfooded** — the herdr-centric **`fleetd`** message server is built and in
🟢 **Implemented & dogfooded** — the herdr-centric **`bridged`** message server is built and in
real use: an Opus primary delegates tasks to off-subscription workers that reply through the
bridge (code reviews delegated this way have produced committed bug fixes). Selected as the
primary approach 2026-07-11, superseding the AgentAPI plan (2026-07-08); AgentAPI retained as a
@@ -100,9 +98,9 @@ tests run separately via `mvn test -Pcontract`):
- **Core gateway** — herdr socket client (contract-tested vs live 0.7.0); guard-checked worker
spawn with `ANTHROPIC_BASE_URL` injected only into the worker's env; status-gated injector;
blocking `fleet_send` with reply rendezvous; MCP server as a thin adapter over the REST core.
- **MCP tools** — `fleet_send` / `fleet_reply` / `fleet_status` (messaging) and `fleet_spawn`
/ `fleet_list` / `fleet_stop` / `fleet_profiles` / `fleet_poll` (fleet). Caller identity is
blocking `bridge_send` with reply rendezvous; MCP server as a thin adapter over the REST core.
- **MCP tools** — `bridge_send` / `bridge_reply` / `bridge_status` (messaging) and `bridge_spawn`
/ `bridge_list` / `bridge_stop` / `bridge_profiles` / `bridge_poll` (fleet). Caller identity is
connection-based (loopback peer PID → herdr pane), so the same mount serves primary and workers.
- **Delivery reliability** — completion fallback (a confirmed `working→idle` turn resolves a
send); async fire-and-poll (beats the caller's MCP call timeout for long tasks); and failure
@@ -110,7 +108,7 @@ tests run separately via `mvn test -Pcontract`):
- **Fleet** — multiple worker profiles, each with an independent base_url guard check; workers
inherit the primary's working directory (never `$HOME`); a readiness gate holds delivery until
a worker's Claude has connected the bridge MCP (no paste lost into its boot window).
- **Blocked-worker path** — `fleet_ask` reverse rendezvous: a worker pauses its delegated turn to
- **Blocked-worker path** — `bridge_ask` reverse rendezvous: a worker pauses its delegated turn to
ask the primary and resumes the *same* turn with the answer (CB-205).
- **Session lifecycle** — session manager with spawn/reuse/recycle, `idle_ttl` reaper, `context_cap`,
and graceful drain on shutdown (CB-301/CB-303); per-worker git worktrees on their own branch with
+14
View File
@@ -0,0 +1,14 @@
# Build output
target/
dependency-reduced-pom.xml
# Local runtime config (copy from bridged.example.yaml)
bridged.yaml
# CB-505 audit trail + daemon stdout/stderr — runtime records, never source
logs/
# Editor / OS
*.iml
.idea/
.DS_Store
@@ -1,10 +1,10 @@
# fleetd configuration (example). Copy to fleetd.yaml and adjust.
# bridged configuration (example). Copy to bridged.yaml and adjust.
#
# fleetd is the sole gateway between primary/worker Claude sessions and herdr.
# bridged is the sole gateway between primary/worker Claude sessions and herdr.
# It is NOT a Claude process and must never carry ANTHROPIC_BASE_URL.
# REST + MCP listen address. Keep it on loopback unless you also switch auth.mode to `token`
# below — fleetd REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
# below — bridged REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
bind:
host: 127.0.0.1
port: 8765
@@ -21,7 +21,7 @@ bind:
# bind — the daemon fails fast otherwise, because "unauthenticated ⇒
# primary" on a reachable port would hand spawn/stop/send to anyone.
# tokenEnv → host env var holding the token (never the literal value). Default
# FLEETD_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
# BRIDGED_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
#
# TLS is deliberately NOT terminated in the daemon (CB-501 D3): run a reverse proxy in front and
# let it own certificate lifecycle, e.g.
@@ -29,13 +29,13 @@ bind:
# The broker link gets TLS from its own URI (amqps://…) — see `broker` below.
# auth:
# mode: token
# tokenEnv: FLEETD_API_TOKEN
# tokenEnv: BRIDGED_API_TOKEN
# Optional pinned primary terminal (CB-307). Names the herdr pane the PRIMARY itself runs in:
# a caller whose connection maps to this pane resolves as the primary (no credential needed —
# the pane mapping is as unforgeable as a worker's), and reply nudges are pushed to it.
# REQUIRED when the primary runs inside a herdr pane — without it the pane match reads the
# primary as a worker and refuses spawn/send/stop. Get the id from fleet_whoami; re-pin if
# primary as a worker and refuses spawn/send/stop. Get the id from bridge_whoami; re-pin if
# the primary moves panes.
# primary:
# terminal: term_0123456789abcd
@@ -48,13 +48,13 @@ bind:
# every orchestration call. List each lead's pane here and all of them resolve as leads.
#
# tab → the ONLY field identity depends on (CB-579); the exact label of the tab hosting the lead.
# Label the tab yourself, or let fleetd label one it launches — see `fleet.leaders:` below.
# kind/model → descriptive; they document what runs in the pane and are echoed by fleet_whoami
# Label the tab yourself, or let bridged label one it launches — see `fleet.leaders:` below.
# kind/model → descriptive; they document what runs in the pane and are echoed by bridge_whoami
#
# A lead's tab must already carry its label (or be launched by fleetd, which labels it) — there is
# A lead's tab must already carry its label (or be launched by bridged, which labels it) — there is
# no terminal id to paste in and nothing to re-pin when the session restarts: the tab survives, so
# the same label resolves the same lead again on the next scan.
# `fleet_whoami` reports `{"role":"primary","leader":"<name>"}`; role stays "primary" because a lead
# `bridge_whoami` reports `{"role":"primary","leader":"<name>"}`; role stays "primary" because a lead
# IS a primary for authorization, so nothing that keys on the role breaks.
#
# KEEP `primary:` when adding leads: it still addresses the CB-307 push loop, which needs a single
@@ -63,13 +63,13 @@ bind:
# Leads are configured under `fleet.leaders:` — see THE FLEET further down.
#
# Two things stop the tab-name convention from becoming a way to claim leadership: the configured
# member spaces are excluded from the scan, so nothing fleetd places can land in a matching tab;
# member spaces are excluded from the scan, so nothing bridged places can land in a matching tab;
# and startup REFUSES a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel`
# override, also matches — so the two namespaces cannot overlap by accident. The label is a NAME,
# never a capability: what a pane may do is decided by the role the daemon resolves for it.
# CB-551: IDLE-LEAD HEARTBEAT — nudge the single lead back to work when it has been continuously
# idle (no open fleet_send driving it) past the quiet period. The fleet is one lead + architects +
# idle (no open bridge_send driving it) past the quiet period. The fleet is one lead + architects +
# workers, so a lead that stalls is a single point of failure; the ReplyPushLoop only nudges when a
# reply lands, and this timer catches the gap where nothing lands and the lead just sits idle.
#
@@ -77,113 +77,44 @@ bind:
# a lead turn nobody asked for), so upgrading the daemon must never switch it on for you. Absent
# block = feature off, exactly as before.
#
# Four knobs, each with a default that errs on the side of not burning context:
# Three knobs, each with a default that errs on the side of not burning context:
# idleAfterSeconds: 300 # how long the lead must stay idle before the FIRST nudge (default 300 —
# # absorbs normal post-turn pauses; re-prompting every pause burns context)
# backoffMs: 60000 # re-check cadence / spacing between nudges past the quiet period (default 60000)
# quietNudgeCap: 3 # cap on consecutive nudges that find NOTHING pending, then it stops
# # until real state appears (default 3 — never nag an empty fleet forever)
# contextHighNudge: false # fleetd #609 — when true, an idle lead whose OWN Claude Code context
# # reads HIGH (see LeadContextGauge; fleet_list's context row) gets a text
# # notice appended to its nudge telling it to consider fleet_handover. Text
# # only — it never rolls a pane by itself, and only the operator can approve
# # a roll. Fires once per HIGH stretch (a later OK reading re-arms it), and
# # never spends the quietNudgeCap budget. Default false/absent = off, same
# # as every other knob here — an upgraded daemon must not start telling
# # leads to hand over on its own.
# leadHeartbeat:
# idleAfterSeconds: 300
# backoffMs: 60000
# quietNudgeCap: 3
# contextHighNudge: false
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
# own pane and bootstrap a fresh session against that file.
#
# Opt-in on purpose — it clears the lead's own pane on request, so upgrading the daemon must never
# acquire that ability for you. Absent block = feature off, and nothing is constructed at all. Even
# once present, nothing but an explicit confirm() call — one that passes every check — can ever
# cause a /clear: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so
# it cannot clear the pane inline (that pane is still WORKING); instead it schedules a one-shot
# continuation that waits for the SAME confirm() call's turn to end, then does the actual work. See
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order (fleetd #480 correction).
#
# handoverPath: REQUIRED when this block is present — where the handover file a fresh lead session
# reads must live. No default (an operator-specific path); a present block with no
# handoverPath refuses to start. May be relative: it then resolves against the
# CALLING lead's own fleet.leaders.<name>.cwd (falling back to the daemon's own
# working directory when that lead has none configured) — never against whatever
# directory the daemon process happens to have been started in. An absolute path is
# used unchanged. Prefer an absolute path if the daemon and the lead's pane might not
# share a working directory (fleetd #480 follow-up).
# requireOperatorConfirm: true # default true — confirm() refuses unless the caller also passes
# # operatorConfirmed: true
# maxDocAgeSeconds: 3600 # default 3600 — refuse a handover file older than this
# turnSettleSeconds: 20 # default 20 — how long the deferred roll waits for the CALLING
# # lead's own turn to end (its pane to report injectable again)
# # before sending /clear at all. If this elapses, /clear is NEVER
# # sent — a lead that never goes idle is still doing real work.
# clearSettleSeconds: 20 # default 20 — how long to wait for the pane to become injectable
# # again AFTER /clear before giving up (never sends bootstrapText
# # if this elapses). A separate, second wait from turnSettleSeconds.
# bootstrapText: "..." # default names the RESOLVED (absolute) handoverPath — sent to
# # the lead once its pane settles after /clear
# leadRollover:
# handoverPath: /path/to/handover.md
# requireOperatorConfirm: true
# maxDocAgeSeconds: 3600
# turnSettleSeconds: 20
# clearSettleSeconds: 20
# bootstrapText: "Fresh lead session: read the handover file and carry on."
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
# per tick.
# intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes
# Math.max(15, intervalSeconds), so a lower value is silently raised, not
# rejected.
# workingSuspectAfterSeconds → age before a BUSY member is suspected of a stall (default 600).
# ENFORCED floor of 300: a lower value is silently raised.
# paneProbeIntervalSeconds → accepted and parsed, but NOT YET READ by anything. Setting it changes
# nothing right now. It exists so a later build can start honouring it without
# another config-shape change.
# notifications.mode → "webhook" flips what fleet_list REPORTS (healthCoverage: "full" instead
# of "detection-only") — it does NOT make fleetd send any webhook call; no
# workingSuspectAfterSeconds, paneProbeIntervalSeconds → accepted and parsed, but NOT YET READ by
# anything — the dormant monitor only consumes intervalSeconds today (CB-573
# shipped ahead of the evidence publishers these two knobs are for). Setting
# them changes nothing right now, and no minimum is enforced on either, because
# nothing reads them to enforce one. They exist so a later build can start
# honouring them without another config-shape change.
# notifications.mode → "webhook" flips what bridge_list REPORTS (healthCoverage: "full" instead
# of "detection-only") — it does NOT make bridged send any webhook call; no
# delivery mechanism is implemented yet. Any other value, or omitting the
# block, reports "detection-only".
# health:
# enabled: true
# intervalSeconds: 30 # floor 15
# workingSuspectAfterSeconds: 600 # floor 300 — how long BUSY with no activity means STALL_SUSPECTED
# paneProbeIntervalSeconds: 60 # parsed, but nothing reads it yet — changing it changes nothing
# intervalSeconds: 30
# workingSuspectAfterSeconds: 600
# paneProbeIntervalSeconds: 60
# notifications:
# mode: disabled
# Idle-sleep guard: while at least one member is live, hold an OS-level assertion against idle
# sleep (macOS only — a `caffeinate -i` child; a no-op elsewhere or if caffeinate is missing), so
# an unattended host does not idle-sleep out from under a member's long turn. Unlike health/
# configReload above, this is ON BY DEFAULT — omitting the block entirely leaves it enabled, the
# same as `enabled: true`. Uncomment only to turn it off:
# idleSleepGuard:
# enabled: false
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
# Optional socket for member panes. Omit this to use herdrSocket for both leads and members.
# memberHerdrSocket: /Users/member/.config/herdr/herdr.sock
# fleetd #213: the login shell the member OS user (memberHerdrSocket above) actually runs. ONLY
# read when memberHerdrSocket is set — fleetd's own $SHELL says nothing about a pane running
# under a different OS user, and there is no channel to ask herdr for that user's shell, so this
# must be told rather than guessed. Absent, blank, or anything not ending in "zsh" is treated the
# same as "not zsh": the memberCredentials.policy: allow-list ZDOTDIR scrub (see worktreeGroup
# below) is skipped in favour of the weaker CB-596 sentinel overlay — a degraded control, never a
# refusal to spawn. When memberHerdrSocket is absent this key is never consulted at all.
# memberLoginShell: /bin/zsh
# How member sessions are spawned. Define one or more named profiles (backends) under
# `profiles`; each key is the profile name (also the ccs profile). A profile says only WHICH
# BACKEND — model, CLI adapter, credentials, cost. It says nothing about what a member spawned on
@@ -193,34 +124,8 @@ herdrSocket: ~/.config/herdr/herdr.sock
# Shared knobs (placement/workspace) can be repeated per profile; they usually match.
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
# Use `pane` for the legacy behaviour (split the focused tab).
# mcpUrl → fleetd mounts the bridge MCP (--mcp-config, inline) + reply charter
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
# (--append-system-prompt) as launch flags; nothing is written to the profile.
# ideMcpUrl → opt-in (CB-634), default off. When set, fleetd mounts the IDE Index MCP as a
# second inline server named `intellij`, and adds an IDE charter that pins every
# ide_* call to the member's own worktree. A URL, not a boolean — host and port
# are host-specific. Set it only on a host where the IDE actually runs.
# ideProjectDir → repo-relative module dir the IDE opens and the overlay pins (CB-634). Only read
# when ideMcpUrl is set. This repo's Maven pom lives in `fleetd/`, not at the
# worktree root, so opening the root imports no module and ide_* resolves nothing;
# set this to `fleetd`. Omit for a repo whose project is the worktree root.
# ideOpenCommand → host command that opens ideProjectDir in the IDE at spawn (CB-634 auto-open).
# Only read when ideMcpUrl is set. `{dir}` is replaced with the absolute module
# dir and the command runs through `/bin/sh -c`, so set env inline if needed —
# e.g. `env DISPLAY=:10.0 idea {dir}`. Best-effort: a failure is logged, never
# fails the spawn. Omit to open the member's module by hand. There is no close
# half yet — an opened module stays open until the operator closes it.
# autoCompactWindow → opt-in, default off. A bounded token window that forces a spawned member to
# compact its context instead of running on the backend's own default and dying
# mid-turn (losing its fleet_reply — the whole point of the turn — with it).
# Validated at config load to [100000, 1000000] — the band Claude Code's own
# --autocompact flag accepts.
# CROSS-BACKEND SEMANTICS DIFFER: on claude-code this is a launch-time
# `--autocompact <tokens>` flag — the member compacts AT this window. opencode
# has no equivalent flag (it only forces `compaction.auto: true`, unconditionally,
# already), so this is instead applied as the model's `limit.context` in the
# generated opencode.json — the member compacts WITHIN this window, not exactly
# at it — and only when this profile's `model:` is in `provider/model` form; if it
# isn't, fleetd logs a WARN naming the profile rather than silently doing nothing.
# tokenEnv → host env var holding the worker's auth token (value never stored in config);
# omit for a backend that needs no token (e.g. a local ollama).
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
@@ -230,12 +135,7 @@ herdrSocket: ~/.config/herdr/herdr.sock
# skills/MCP/hooks. Omit to leave the worker on the host default.
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
# worktree sees the same local config (CB-301-ext). Omit for the default set:
# [.env] only (CB-148). .envrc is left out of the default on purpose: it is
# executable shell that direnv runs on every cd, so copying it carries
# behaviour into the worker, not just values, unlike .env. An operator who
# wants it copied can still write parityOverlay: [.env, .envrc] explicitly.
# (.claude/settings.local.json is NOT in the default — it
# pre-approves IDE/tool grants a member must not hold ambiently; CB-525/CB-634.)
# [.claude/settings.local.json, .env, .envrc].
#
# Do NOT add .mcp.json (CB-525). A worker's tools are whatever its launcher
# mounts — the bridge, and nothing else. Replicating the primary's MCP config
@@ -243,7 +143,7 @@ herdrSocket: ~/.config/herdr/herdr.sock
# checkout, so its navigation returned paths OUTSIDE its own worktree: one
# worker made all 59 of its edits in the primary tree while compiling its
# worktree, and every build it ran was of code that did not contain them.
# fleetd neutralizes a provisioned worktree's .mcp.json for this reason;
# bridged neutralizes a provisioned worktree's .mcp.json for this reason;
# listing it here would copy the primary's back over that.
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
@@ -252,17 +152,13 @@ herdrSocket: ~/.config/herdr/herdr.sock
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
# exhaustedPattern → regex matched against a completion-fallback scrape (CB-578 stage A) to
# classify a turn that ended with no fleet_reply as the backend having
# classify a turn that ended with no bridge_reply as the backend having
# refused on a subscription usage limit, rather than a real answer. Opt-in —
# omit and this profile's completion fallback behaves exactly as before.
# Every backend words its refusal differently, so this is config, never a
# vendor string baked into fleetd itself.
# HOT (fleetd #446): read live, cached by profile name, at every
# completion-fallback check AND by fleet_profiles' exhaustionDetectionArmed —
# editing it and reloading arms or disarms usage-limit detection for this
# profile with no daemon restart. (Before fleetd #446 this was DEFERRED,
# compiled once into a startup pattern map like model/baseUrl/argv still are —
# see errorPattern below, which is still deferred that way on purpose.)
# vendor string baked into bridged itself.
# DEFERRED: compiled once into a startup pattern map — editing it needs a
# daemon restart, same as this profile's model/baseUrl/argv.
# credentialId → CB-578 stage B: the credential this profile quarantines WITH when a
# BACKEND_EXHAUSTED classification fires. Two profiles that set the SAME
# credentialId share one quarantine — the case this exists for is two models
@@ -272,52 +168,18 @@ herdrSocket: ~/.config/herdr/herdr.sock
# profile quarantines alone, under its own name, exactly as if the field did
# not exist. Cooldown length is the top-level quarantineCooldownSeconds below.
# HOT: read live at every spawn/exhaustion check — no restart needed.
# errorPattern → fleetd #201 / #227: regex matched against a completion-fallback scrape to
# classify a turn that ended with no fleet_reply as a BACKEND ERROR — a
# credential outage or a provider 5xx — rather than a real answer or a
# usage-limit exhaustion (exhaustedPattern above always wins when a line
# matches both). Opt-in. Omit it and this profile falls back to fleetd's
# built-in legacy pattern `(?i)\bAPI Error\s*:` — classification still
# happens, just without a profile-specific match; every backend words its
# failure differently, so a hardcoded sentence would only ever match one
# of them.
# DEFERRED: compiled once into a startup pattern map — editing it needs a
# daemon restart. Unlike exhaustedPattern above (made hot by fleetd #446),
# errorPattern was scoped out of that ticket on purpose and stays deferred.
# # errorPattern: "503 Service Unavailable" # opt-in: classify a backend outage
#
# What happens once a match fires (BackendOutagePolicy, credentialId-keyed,
# SEPARATE from the CB-578 stage B quarantine above and never merged with it):
# - threshold 2 — TWO DISTINCT TARGETS (never raw events) on the same
# effective credential inside a 60-second window start an "incident" and a
# 60-second cool-off for that credential. One member repeating the same
# classified line twice never cools anything off — a real outage hits
# every target on that credential, so requiring a second, independent
# target loses nothing against the case this guards against, while
# protecting against a heuristic misfire on one flaky member.
# - a fresh error while a credential is already cooling off is ignored
# outright: it neither extends the 60s deadline nor starts a new incident.
# - `fleet_list`/`fleet_profiles` report a cooling credential with
# `coolingOffForSeconds` (never `quarantinedForSeconds`, unless CB-578
# exhaustion quarantine is ALSO independently active for the same
# credential — the two checks can both fire at once). A spawn onto a
# cooling profile is refused with a message naming the credential and
# remaining seconds — "cooling off", never "exhausted", so an operator can
# tell a short transient fault from a spent subscription at a glance.
# - the lead gets ONE nudge per incident (not one per affected target), via
# the same push loop that already delivers ticket/question reminders.
# env → extra environment for this profile's workers, as a literal key/value map
# (CB-511). Use it to give workers a toolchain.
#
# A worker's environment does NOT come from your shell. fleetd hands herdr an
# A worker's environment does NOT come from your shell. bridged hands herdr an
# explicit env map and herdr merges it into ITS OWN process env — so before
# CB-511 a worker inherited whatever PATH the herdr server happened to be
# started with, which on a long-lived herdr can predate your toolchain entirely
# and leave workers unable to run `mvn` or `java` at all.
# fleetd now propagates ITS OWN PATH to every worker by default; set `env:`
# bridged now propagates ITS OWN PATH to every worker by default; set `env:`
# only to override that or add more (JAVA_HOME, …). Since the default is the
# daemon's PATH, make sure the daemon is started with a good one — see the PATH
# lines in deploy/dev.ltms.fleet.plist and deploy/fleetd.service.
# lines in deploy/dev.ltms.bridged.plist and deploy/bridged.service.
#
# Adapter-owned variables always win over `env:`: ANTHROPIC_BASE_URL and the
# rest of the ANTHROPIC_*/CLAUDE_* wiring are applied after it, so an `env:`
@@ -330,20 +192,20 @@ profiles:
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
model: coder
placement: tab
workspace: fleetd-workers
workspace: bridged-workers
# tabLabel: an optional per-profile override; the fleet template usually covers it
mcpUrl: http://127.0.0.1:8765/mcp
tokenEnv: FLEETD_WORKER_TOKEN
tokenEnv: BRIDGED_WORKER_TOKEN
argv: ["ccs", "gx10"]
# weight: relative selection weight for automatic placement (weighted, round-robin, and
# fixed's fallback walk). Absent defaults to 1.0. An explicit 0 or negative value means
# "never auto-select this profile" (CB-554) — it stays reachable via an explicit
# `fleet_spawn{profile:"gx10"}`, which bypasses placement entirely; only automatic
# `bridge_spawn{profile:"gx10"}`, which bypasses placement entirely; only automatic
# selection skips it.
weight: 0.5
# maxLoad: max live workers on this profile. Omit for unlimited. An explicit 0 (CB-585) caps
# the profile at zero live members — it is excluded from automatic placement and an explicit
# `fleet_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile
# `bridge_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile
# is named directly. Negative is refused at config load — there is no sane meaning for it.
maxLoad: 2
# subscription: true
@@ -362,64 +224,24 @@ profiles:
# subscription path no guard would vet the URL, so allowing both would be a way around the
# guard rather than a configuration.
#
# GOTCHA 1 — it is invisible to the startup secret check. `Fleetd.reportRequiredSecrets`
# GOTCHA 1 — it is invisible to the startup secret check. `Bridged.reportRequiredSecrets`
# skips subscription profiles on purpose (they need no token), so a boot log that reports
# every secret as fine says nothing about these profiles.
#
# GOTCHA 2 — `maxLoad` is the ONLY throttle you have here. There is no metering, no budget
# and no refusal on cost; the cap on live members is the single thing standing between a
# fan-out and your monthly limit. Set it deliberately and keep it small.
#
# GOTCHA 3 (fleetd #176, corrected by fleetd #257) — `maxLoad` counts members, never the lead
# itself. The lead is a live `claude` session on this SAME account (a lead is never moved
# off-subscription, whatever its own profile says), so it already holds one seat before any
# member spawns. If a lead's `fleet.leaders.<name>.profile` names THIS profile — or ANY OTHER
# `subscription: true` profile that shares this one's account (see THE SENTINEL, just below,
# next to `credentialId:`) — `fleet_list` reports that seat count under `leadSeats`; see
# `profile:` under THE FLEET below. `free` itself is NEVER reduced by `leadSeats`: `free` means
# "what the real placement gate (`CompositePeerLauncher#enforceMaxLoad`) will actually grant a
# fresh `fleet_spawn` right now", and that gate only ever compares live members against
# `maxLoad` — it has no notion of the lead's own seat. An earlier cut of this feature
# subtracted `leadSeats` from `free` on the theory it made `free` describe the true ceiling on
# the account, but no backend seat ceiling shared with the lead has ever actually been
# measured, and the subtraction just made `free` disagree with the one thing it is supposed to
# describe — the fleetd #257 fix. `maxLoad: 3` means 3 member slots, full stop; a lead sharing
# the account is a fact you can see in `leadSeats`, not a reason `free` undercounts spawns that
# will, in practice, succeed.
#
# THE SENTINEL (fleetd #176 stage 2, correcting an inert stage 1 fix): every `subscription:
# true` profile that leaves `credentialId` unset shares ONE implicit account-wide credential
# id with every other such profile on this host — because a subscription profile doesn't
# authenticate with a credential of its own, it authenticates as the operator's own Claude
# login, and there is exactly one of those. So on a typical host, `opus` (the lead's profile)
# and `sonnet` (the members' profile) are linked automatically, with NOTHING to set here — that
# is what makes GOTCHA 3 above work without also writing matching `credentialId:` values on
# both. This linkage is not just cosmetic: it is the same key `BackendQuarantine`/cool-off use,
# so a usage-limit hit on `opus` now quarantines `sonnet` too (and vice versa) — correct, since
# they are one Claude account, but worth knowing before you wonder why an unrelated-looking
# profile went quarantined.
#
# WHEN TO OVERRIDE — set explicit, DIFFERENT `credentialId:` values on two `subscription: true`
# profiles only when they are genuinely two separate Claude logins on the same host (a real,
# if unusual, setup). An explicit `credentialId` always wins over the sentinel, so this is the
# one way to keep two subscription profiles from being treated as one account for lead-seat
# counting AND for quarantine/cool-off grouping alike.
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
# credentialId: shared-openai # opt-in: quarantine together with every other profile sharing this id (CB-578)
# errorPattern: "503 Service Unavailable" # opt-in: classify a backend outage (fleetd #201/#227) — see the key doc above
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
# parityOverlay: [".env"] # the default; add ".envrc" explicitly if you want it copied too (CB-148) — never add .mcp.json or .claude/settings.local.json — see above
# ideMcpUrl: http://127.0.0.1:29170/index-mcp/streamable-http # opt-in (CB-634): IDE code intelligence, pinned to the worktree
# ideProjectDir: fleetd # CB-634: module dir the IDE opens + the overlay pins (this repo's pom is in fleetd/)
# ideOpenCommand: env DISPLAY=:10.0 idea {dir} # CB-634 auto-open: opens {dir} in the IDE at spawn; omit to open by hand
# autoCompactWindow: 250000 # opt-in: bound member context; claude-code compacts AT this, opencode within it (model limit.context)
# parityOverlay: [".claude/settings.local.json", ".env", ".envrc"] # never add .mcp.json — see above
gx11: # a second backend, so `placement: weighted` has a choice
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
placement: tab
workspace: fleetd-workers
workspace: bridged-workers
# tabLabel: an optional per-profile override; the fleet template usually covers it
mcpUrl: http://127.0.0.1:8765/mcp
argv: ["ccs", "gx11"]
@@ -437,15 +259,15 @@ profiles:
# The bridge MCP + reply charter mount via a generated OPENCODE_CONFIG file, and the model is a
# `provider/model` selector. Placement, tabs, cwd, and the readiness gate are shared with Claude.
#
# Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → fleet_send →
# structured fleet_reply → teardown). The `opencode/*-free` models run on opencode's own gateway
# Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → bridge_send →
# structured bridge_reply → teardown). The `opencode/*-free` models run on opencode's own gateway
# and need NO credentials — check `opencode models` for the current free list, since the names
# change. That also makes the worker off-subscription by construction.
# opencode-free:
# kind: opencode
# model: opencode/north-mini-code-free # `provider/model` selector, injected as `-m`
# placement: tab
# workspace: fleetd-workers
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
@@ -470,7 +292,7 @@ profiles:
# baseUrl: http://127.0.0.1:8000
# model: local-vllm/deepseek-v4-flash
# placement: tab
# workspace: fleetd-workers
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
@@ -502,29 +324,14 @@ placement: weighted
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
# 1800 (30 minutes) when omitted or non-positive.
#
# fleetd #466: this is now only the BASE of an escalating backoff, not a flat retry rate. A
# credential quarantined again within one base cooldown of the previous quarantine ending (still
# reporting exhausted — e.g. a weekly subscription limit that hasn't reset) backs off further:
# cooldown doubles each such time, capped at 12x this value (~6 hours at the 1800s default). A
# quarantine that starts after a base-cooldown's worth of quiet resets back to this value. Not
# configurable per se — the multiplier and ceiling are constants in BackendQuarantine, not new
# YAML keys; see its class doc for the exact formula and why there is no automatic probe to clear
# it early (the operator's own design constraint — a probe spends the quota it's measuring).
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
# restart. Editing this needs a daemon restart to take effect.
#
# This does NOT govern the fleetd #201 / #227 backend-error cool-off documented under errorPattern
# above — that mechanism is a separate, shorter-lived, NOT-configurable policy (threshold 2 distinct
# targets, 60-second window, 60-second cool-off), on purpose: it exists to survive a brief transient
# fault, not to replace this 30-minute exhaustion quarantine. Do not conflate the two when reading
# fleet_list/fleet_profiles — coolingOffForSeconds and quarantinedForSeconds are independent facts.
# quarantineCooldownSeconds: 1800
# Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an
# upgraded fleetd keeps the old behaviour: the file is read once at boot and never again.
# enabled → turn the watch on. fleetd checks the file's modified time on a timer and
# upgraded bridged keeps the old behaviour: the file is read once at boot and never again.
# enabled → turn the watch on. bridged checks the file's modified time on a timer and
# reloads when it moves.
# intervalSeconds → how often to check (default 10). One `stat` per tick, so this is cheap.
#
@@ -532,11 +339,10 @@ placement: weighted
# when the reload happens — not about how important the key is:
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad
# / credentialId / exhaustedPattern. Those are hot because the placement policy (and,
# for credentialId, the CB-578 stage B quarantine check; for exhaustedPattern, fleetd
# #446's LiveExhaustedPatterns) reads them through a supplier — being config is not by
# itself enough to make a key hot.
# EXCEPT `fleet.leaders`: Fleetd.main reads it once at startup to build the lead tab
# / credentialId. Those are hot because the placement policy (and, for credentialId,
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
# not by itself enough to make a key hot.
# EXCEPT `fleet.leaders`: Bridged.main reads it once at startup to build the lead tab
# scanner and launcher, and neither is rebuilt on reload. A changed/added/removed
# `fleet.leaders` entry is silently accepted — the reload reports "config reloaded"
# with nothing in the deferred list — but has NO effect until you restart. Treat it
@@ -547,10 +353,8 @@ placement: weighted
# stage B — baked once into the quarantine tracker built at startup), ADDING or
# REMOVING a profile (a new backend needs its own launcher, and launchers are built
# once), AND an existing profile's launch settings — model, baseUrl, argv, env,
# configDir, mcpUrl, tabLabel, errorPattern (fleetd #201 / #227 — compiled once into
# a startup pattern map; exhaustedPattern used to be compiled the same way until
# fleetd #446 made it hot — see above). The launcher takes a copy of `profiles:` at
# startup and resolves every spawn out of that copy, so those never
# configDir, mcpUrl, tabLabel, exhaustedPattern. The launcher takes a copy of
# `profiles:` at startup and resolves every spawn out of that copy, so those never
# reach a launch until you restart. The reload logs them by name rather than
# pretending they applied.
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
@@ -569,7 +373,7 @@ placement: weighted
# older keys: `leaders:`, `members:`, `leadScan:` and `defaultProfile:`.
#
# A member is anything a lead spawns, and every member has two INDEPENDENT attributes:
# role — which contract: architect, dev, hunter or reviewer. It picks the launch charter, the role
# role — which contract: architect, dev or reviewer. It picks the launch charter, the role
# file, the playbook skill and the authz row.
# profile — which backend: one of the `profiles:` keys above (model, CLI adapter, cost).
# They vary on their own. A reviewer may run on the same profile as the dev whose diff it reads,
@@ -581,13 +385,13 @@ placement: weighted
#
# Each pool lists the profiles that role MAY run on — these are pools, not identities. That is also
# what replaced `defaultProfile:`: an unqualified spawn names a role, and that role's pool supplies
# the candidates, in definition order. A dev, hunter and reviewer staying anonymous is exactly
# compatible with being listed here; the entry key just names the entry.
# the candidates, in definition order. A dev and a reviewer staying anonymous is exactly compatible
# with being listed here; the entry key just names the entry.
fleet:
# Optional launch-charter text, keyed only by the singular role wire names: architect, dev,
# hunter, reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put
# secrets here: a later launch step writes this text to a world-readable temp file, and ${ENV}
# interpolation is deliberately not supported.
# reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put secrets
# here: a later launch step writes this text to a world-readable temp file, and ${ENV} interpolation
# is deliberately not supported.
charters:
architect: |-
You are an architect in this fleet. You refine work before anyone builds it:
@@ -598,9 +402,6 @@ fleet:
dev: |-
You implement the one unit you were given, and nothing else. You test it,
commit it, and open your own pull request. You never merge.
hunter: |-
You sweep the assigned scope for real defects. You may run the build or tests
to check a finding. You change nothing, and report several ranked findings.
reviewer: |-
You review the diff you were given. You report bugs, risks and missing tests.
You do not change code.
@@ -614,23 +415,6 @@ fleet:
# recognised: give it a `profile:` and the daemon launches the shortfall when fewer than
# `instances` are live. Omit `profile:` and it is recognise-only, as before.
#
# `profile:` has a SECOND job as of fleetd #176, even for a recognise-only lead you never want
# auto-launched: it is also how fleetd learns which account this lead's own session shares. A
# `subscription: true` profile bills the operator's Claude account, and the lead itself is always
# a live `claude` session on that same account — `maxLoad` never counted that seat. If a lead
# entry here names a profile that shares a worker profile's account, `fleet_list` reports the
# lead's live seat(s) on that worker profile under `leadSeats` — informational only, as of fleetd
# #257 it is NEVER subtracted from `free` (see GOTCHA 3, next to `maxLoad:`, in THE WORKERS above,
# for why). "Shares the account" is decided by matching `effectiveCredentialId()`, which (fleetd
# #176 stage 2 — see THE SENTINEL, next to `credentialId:`, in THE WORKERS above) means: an
# explicit, matching `credentialId:` on both, OR — the common case, needing NO extra config — both
# being `subscription: true` with `credentialId` left unset, since those all share one implicit
# account-wide id. A lead on `opus` and workers on `sonnet` link automatically this way; they do
# NOT need the same profile name. Setting `profile:` on an already-running, recognise-only lead is
# safe — the daemon only launches the SHORTFALL below `instances`, so naming a profile here does
# not, by itself, start anything. Omit it and fleetd has no way to derive the sharing — there is
# no other reliable signal on the daemon's side — so that lead's seat never appears in `leadSeats`.
#
# `tab:` (CB-579) is REQUIRED and is the only field identity depends on — the exact label of the
# tab hosting the lead, matched case-insensitively. Label the tab yourself and put that same
# string here, and the pane is recognised on the next rescan. Reopen the tab later, or the session
@@ -662,8 +446,8 @@ fleet:
# workspace: leads # where a launched lead's tab is created (default "leads").
# # MUST NOT be a member workspace — those are excluded from the
# # scan, so a lead placed in one is never found again.
# cwd: /path/to/repo # the launched lead's working directory (default: fleetd's own)
# kind: claude # descriptive; reported by fleet_whoami
# cwd: /path/to/repo # the launched lead's working directory (default: bridged's own)
# kind: claude # descriptive; reported by bridge_whoami
# gpt-sol-5.6:
# tab: "lead: gpt-sol-5.6"
# kind: opencode
@@ -678,9 +462,6 @@ fleet:
developers:
gx10:
profile: gx10
# hunters:
# gx10:
# profile: gx10 # a hunt may run checks, but never changes code
# reviewers:
# gx10:
# profile: gx10 # the same backend may serve two roles; that is the point
@@ -694,7 +475,7 @@ guard:
# Member credential policy (CB-596, gitea issue #82). A herdr pane runs a LOGIN shell, and that
# shell re-sources the operator's own secret store — so a spawned member inherits every credential
# the operator's shell holds, not just the ones fleetd means to give it. Measured on this host:
# the operator's shell holds, not just the ones bridged means to give it. Measured on this host:
# 31 credential names, all set, with only ONE (GITEA_ACCESS_TOKEN) blocked before this — and that
# block was a single name hardcoded in HerdrPeerLauncher.java, not driven by this file. This block
# replaces that hardcoded shadow with a config-driven list of names.
@@ -724,66 +505,20 @@ guard:
# through either: the daemon logs a WARN naming any credential-shaped env var it finds on neither
# list (never its value), so a secret added to the store later does not go unnoticed forever.
#
# policy → "deny-by-default" (the default; also accepted spelled "deny-list") overlays each
# known-but-not-allowed name BEFORE the pane's login shell runs — real protection only
# where that shell does not re-export the name (see ROUND-2 CORRECTION above). An
# unrecognized value refuses to start, naming it.
# policy → "allow-list" (CB-633) moves the control to a per-spawn ZDOTDIR directory the daemon
# generates and passes through tab.create's env map. Each generated startup file sources
# its ~/ counterpart FIRST and then runs the scrub, so the scrub happens after the
# operator's whole chain and no sourced file can undo it.
# The scrub is sourced from BOTH the generated .zshrc and the generated .zlogin, because
# herdr does not open the same kind of shell everywhere: macOS panes run a LOGIN zsh (so
# .zlogin runs), Linux panes run a plain interactive zsh (so .zlogin never runs at all).
# A scrub in .zlogin alone would be a control that silently does nothing on Linux.
# The allow-list is DERIVED, never typed:
# every profile's tokenEnv/gitTokenEnv/gitHostEnv values and env-map keys, plus an
# infrastructure set (PATH HOME SHELL TERM LANG LC_* TMPDIR USER LOGNAME PWD SHLVL EDITOR
# PAGER JAVA_HOME XDG_* ZDOTDIR), plus whatever keys this spawn's own env overlay carries.
# Adding a profile can therefore only widen the list, never break another spawn's scrub.
# Under this policy `known`/`allow` below become REPORTING ONLY — they feed the gap WARN,
# they are no longer a control. If the member's login shell is NOT zsh, the daemon logs a
# loud WARN saying protection is off and falls back to deny-by-default's overlay.
# Each pane writes a scrub-report.txt naming how many variables it kept of how many it
# saw; the daemon logs that "allowed N of M" line when the pane stops. If the report is
# MISSING the daemon logs a WARN instead — the scrub cannot then be confirmed to have
# run, and a silently-dead control is exactly what this policy exists to prevent.
# allow → credential names a member legitimately needs. Under deny-by-default, left OUT of the
# pane's env overlay entirely, so the value the pane's own (login) shell exports passes
# through untouched. Under allow-list: reporting only.
# known → every credential name the operator's store is known to export. Under deny-by-default,
# every name here NOT also in `allow` is overlaid with a non-secret sentinel value before
# the pane's login shell runs — real protection only for names that shell does not itself
# re-export (see the ROUND-2 CORRECTION note above). Under allow-list: reporting only.
# sshAgentEnv → whether SSH_AUTH_SOCK may pass through under allow-list ("inherit") or is omitted
# from the member environment ("omit", the default). Omitting it only omits the
# inherited ssh-agent path. It discourages automatic use of the operator's agent.
# It does not deny same-user access to that socket. It also does not block SSH keys that
# are readable on disk. Git over SSH may still work from inside a member. Keep the block:
# it is correct and costs nothing, but it is not a control. A member runs as the same OS
# user as the lead. Inside one uid, ordinary Unix permissions provide no meaningful
# confidentiality boundary. A real boundary needs a different OS user or OS-level
# confinement, such as a container or VM. That is the open question in fleetd #184.
#
# Still do not set this to "inherit" casually. SSH_AUTH_SOCK is a live handle to YOUR
# ssh-agent, so a member holding it can sign with EVERY key the agent holds. It sits in
# no secret file and looks like no credential, which is why it slipped past three
# earlier tickets (gitea #110). Blocking it does not contain a member, but allowing it
# hands one a signing capability for no gain — the block costs nothing, so keep it.
#
# Both halves of this are measured, not argued. 2026-08-28: a member with
# SSH_AUTH_SOCK blanked pushed to the forge over SSH successfully, because `ssh -G`
# resolves an IdentityFile outside ~/.ssh that is readable and has no passphrase. An
# earlier version of this comment claimed blocking the socket BREAKS git over SSH. It
# does not. That claim came from looking only in ~/.ssh, which holds nothing but four
# `Include` lines — looking in one place and concluding about the whole host.
# policy → only "deny-by-default" exists today (an operator-authored deny-list was deliberately
# rejected — see above). An unrecognized value refuses to start, naming it.
# allow → credential names a member legitimately needs. Left OUT of the pane's env overlay
# entirely, so the value the pane's own (login) shell exports passes through untouched.
# known → every credential name the operator's store is known to export. Every name here NOT
# also in `allow` is overlaid with a non-secret sentinel value before the pane's login
# shell runs — real protection only for names the login shell does not itself re-export
# (see the ROUND-2 CORRECTION note above for the ones it does).
#
# HOT-RELOADABLE the same way `fleet:` is (CB-559): read fresh on every spawn, so editing this list
# and reloading config (or restarting) changes what the NEXT spawn inherits; already-running members
# are unaffected either way.
# memberCredentials:
# policy: deny-by-default # or "deny-list", or "allow-list" (CB-633) — see above
# sshAgentEnv: omit # allow-list only; see the sshAgentEnv note above
# policy: deny-by-default
# allow:
# - AI_GATEWAY_TOKEN # named in a profile's tokenEnv (local/gx) — a member reaching the
# # gateway is by design, not a leak
@@ -835,46 +570,6 @@ guard:
# to a sibling directory of the repo root.
# worktreeRoot: /Users/me/src/.bridged-worktrees
# Worktree group sharing (fleetd #185 stage 3). OPTIONAL, off by default. Names an OS group
# that a provisioned worktree's repo is made group-writable for (git config
# core.sharedRepository group, plus a one-time chgrp/chmod/setgid fix-up), so a member spawned
# under a DIFFERENT OS user (see memberHerdrSocket) can write its own worktree, its
# per-worktree git metadata, and its own commit objects — without it, every file GitWorktrees
# creates is owned by fleetd's own uid and unwritable by another user.
# CAUTION: this isolates credentials, not the repository — a member in the group can still
# write the operator's git objects and refs in the shared repo. The operator running fleetd
# must already be a member of the named group, or every provisioning spawn fails loudly.
#
# fleetd #213: this is also the ONE group the memberCredentials.policy: allow-list ZDOTDIR scrub
# reuses when memberHerdrSocket is set — deliberately not a second config key. Under
# memberHerdrSocket, the scrub directory is generated under worktreeRoot (never java.io.tmpdir,
# which the member OS user cannot reach) and shared read-only with this group. If worktreeGroup
# is unset while memberHerdrSocket is set, the scrub cannot be guaranteed reachable by the member,
# so fleetd falls back to the weaker CB-596 sentinel overlay instead (a WARN names the gap).
# worktreeGroup: fleet-workers
# fleetd #362: a directory of skill folders (each a subdirectory holding a SKILL.md, the same
# shape as this repo's own .claude/skills/) copied into every PROVISIONED worktree's
# .claude/skills/, so a member spawned against ANY repo — not only one that already ships its own
# copy — can load a bridge skill (e.g. implementer). Unset (the default): no worktree is touched
# beyond today's behaviour. A skill folder the target repo already carries under
# .claude/skills/<name> is never overwritten — the repo's own copy always wins. Best-effort like
# worktreeGroup above: a missing/unreadable directory here is logged and skipped, never a failed
# spawn. Every non-hidden subdirectory of this directory is copied wholesale, with no per-file
# allowlist — don't park scratch files or drafts alongside the real skill folders, they will be
# copied into every provisioned worktree too.
#
# fleetd #393: which member KINDS actually consume this once it is copied. kind: claude-code —
# the Claude Code CLI discovers .claude/skills/ on its own; nothing else is needed. kind: opencode
# — opencode has no such discovery, so OpenCodeLauncher reads whatever landed under
# .claude/skills/ and appends each seeded skill's SKILL.md to the generated instructions[] file
# (opencode's only channel for static guidance text; unlike Claude Code's Skill tool, the content
# is always part of the system prompt, not loaded on demand). Both kinds are covered as of #393 —
# earlier builds copied the files for every kind but only claude-code could read them, and the
# seeding log said "N of M" regardless. Check the per-spawn launcher log (not just the seeding
# log) to see what a given member actually got.
# memberSkills: /path/to/fleetd/checkout/.claude/skills
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
@@ -898,43 +593,14 @@ guard:
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
# uriEnv → CB-151: name of a host env var holding the AMQP URI, preferred over `uri` (wins
# whenever set). The URI carries `user:pass@` inline, so naming a variable keeps the
# password out of fleetd.yaml — same pattern as auth.tokenEnv/Profile.tokenEnv. A
# uriEnv that resolves to an unset or blank variable is treated as NOT configured and
# the daemon falls back to the in-memory inbox, warning loudly.
# prefetch → CB-527: consumer basicQos, capping how many unacked messages the inbox holds
# in-heap per owned target (the rest sits on the broker's durable queue instead of
# growing the JVM heap). Default 32 when omitted.
# broker:
# uriEnv: LAVINMQ_URI
# uri: amqp://guest:guest@127.0.0.1:5672
# prefetch: 32
# Shared cross-host LEADER coordination broker. OMIT this block to leave lead-to-lead messaging
# off entirely (config-only in this ticket — nothing here wires it into a live LeadMailbox yet).
# This is a SEPARATE AMQP vhost from `broker:` above: member/worker inboxes always stay on the
# per-fleet `broker:` vhost, and this vhost carries only leader-to-leader traffic, so two fleets
# whose members must never see each other can still share one coordination vhost for their leads.
# uriEnv → name of a host env var holding the coordination AMQP URI, same convention as
# broker.uriEnv (keeps the credential out of fleetd.yaml). Wins over `uri` when set.
# selfId → this daemon's own lead coord-id — the name its mailbox is owned under
# (lead.<selfId>.inbox), e.g. "mac-opus" or "fleet01-lead". Must be globally unique
# across every daemon sharing this vhost.
# prefetch → consumer basicQos, capping how many unacked messages the mailbox holds in-heap.
# Default 32 when omitted.
# peers → fleetd #361: the coord-ids of the OTHER daemons on this vhost, declared by the
# operator (the daemon never guesses). fleet_list reports each one's live reachability
# (a passive queue check, never a presence protocol) alongside this daemon's own
# mailbox state. Omit, or leave empty, for a daemon with no known peers yet — an
# undeclared peer can still reach you and be reached by fleet_send, it just will not
# show up as a row in fleet_list.
# coordinator:
# uriEnv: LEAD_COORD_URI
# selfId: mac-opus
# prefetch: 32
# peers: [fleet01-lead]
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open fleet_send,
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
# pane — status-gated (only when injectable, never mid-turn) and bounded. Ack = drain: the loop
# stops as soon as the primary's inbox is empty.
@@ -945,11 +611,11 @@ guard:
#
# REQUIRED (CB-522) if the primary itself runs inside a herdr pane. Caller
# identity resolves a loopback PID to its herdr pane, and PaneLocator scans
# EVERY pane — not just fleetd-spawned ones — so such a primary is otherwise
# EVERY pane — not just bridged-spawned ones — so such a primary is otherwise
# classified as a WORKER and refused SPAWN/SEND/STOP. That failure is
# self-locking: the learned terminal is populated by the very orchestration
# calls being refused, so only this pinned value can break the cycle. Read the
# id off fleet_whoami (it reports the current terminal even while
# id off bridge_whoami (it reports the current terminal even while
# misclassified) and re-pin whenever the primary moves panes.
# pushReminders → max nudges before giving up (default 5)
# pushBackoffMs → delay between nudges in ms (default 15000)
@@ -957,32 +623,3 @@ guard:
# terminal: term_65619bd6174568
# pushReminders: 5
# pushBackoffMs: 15000
# Central allow-list of models any profiles: entry may name. Nothing checked a profile's model:
# value before this block existed — it was a free-form string handed straight to the backend
# adapter, and a withdrawn or misspelled name failed silently instead of at config load (opencode
# falls back to a default model rather than erroring on an unknown -m).
#
# Absent, or present with an empty allow:, is OFF: no profile's model: is checked, exactly like
# before this block existed. fleetd.yaml is gitignored on every host, so an upgrade must not force
# every operator to enumerate their models before the daemon will start.
#
# The list is the authority; profiles: is checked against it, never the reverse — adding or
# editing a profiles: entry cannot, by itself, widen what is permitted here.
#
# Enforcement is at CONFIG LOAD only (a bad model: fails the daemon at startup, naming both the
# model and the profile). There is no spawn-time enforcement, no runtime on/off switch, and no
# interaction with BackendQuarantine — those are separate, later units.
#
# allow → the permitted models. Each entry is its own block (not a bare string) so a later unit
# can add an on/off state or a load limit per model without changing this shape.
# model → the model id exactly as a profiles: entry's model: field would write it. One flat,
# opaque-string namespace: a bare Claude id (claude-sonnet-5) and an opencode
# provider-prefixed id (openai/gpt-5.6-terra) both fit here unchanged — the check is a
# plain string match, never a parse of the provider prefix or a branch on kind:.
# models:
# allow:
# - model: claude-sonnet-5
# - model: claude-opus-5
# - model: openai/gpt-5.6-terra
# - model: amazon.nova-pro-v1:0
@@ -5,10 +5,10 @@
## Why this exists
Stage 2 gave a worker→primary reply a **durable place to wait** when no `fleet_send` is
Stage 2 gave a worker→primary reply a **durable place to wait** when no `bridge_send` is
open: it lands in `agent.<target>.inbox` on the broker and survives a daemon bounce. But
delivery is still **pull** — the primary only sees the reply if it happens to call
`fleet_poll(target)` / `GET /sessions/{id}/replies`. A reply can sit indefinitely while
`bridge_poll(target)` / `GET /sessions/{id}/replies`. A reply can sit indefinitely while
the primary works on something else.
This layer makes delivery **active**: the bridge *pushes* a nudge to the primary the moment
@@ -26,11 +26,11 @@ pointed at the primary's pane instead.
```mermaid
flowchart LR
W["worker"] -->|"fleet_reply (no open send)"| MS["MessageService.reply"]
W["worker"] -->|"bridge_reply (no open send)"| MS["MessageService.reply"]
MS -->|"inbox.publish"| INBOX[("agent.&lt;target&gt;.inbox<br/>(durable, LavinMQ)")]
MS -->|"notify"| LOOP["ReplyPushLoop"]
LOOP -->|"status-gated inject"| PANE["primary's herdr pane"]
PANE -->|"primary drains"| DRAIN["fleet_poll(target)<br/>= peek + ack"]
PANE -->|"primary drains"| DRAIN["bridge_poll(target)<br/>= peek + ack"]
DRAIN -->|"inbox now empty"| LOOP
LOOP -.->|"still non-empty →<br/>re-inject on backoff"| PANE
classDef store fill:#2c5282,stroke:#1a365d,color:#ffffff;
@@ -51,13 +51,13 @@ the caller runs in a herdr pane on this host. Today it's discarded for the prima
(`presence.markPresent` is a no-op on it).
**Plan:** a single-slot `PrimaryRegistry` (thread-safe) holding the primary's `terminal_id`.
Populate it from the **orchestration-side** MCP tools — `fleet_send`, `fleet_spawn`,
`fleet_poll`, `fleet_list`, `fleet_status`, `fleet_profiles` — capturing
Populate it from the **orchestration-side** MCP tools — `bridge_send`, `bridge_spawn`,
`bridge_poll`, `bridge_list`, `bridge_status`, `bridge_profiles` — capturing
`callerTerminal(exchange)` when it is (a) non-null and (b) **not** a registered worker
session in `SessionManager`. That caller is, by construction, the primary. Worker-side tools
(`fleet_reply`, `fleet_ask`) never set it.
(`bridge_reply`, `bridge_ask`) never set it.
- **Config override / pin:** a `primary: { terminal: "<id>" }` block in `FleetConfig`
- **Config override / pin:** a `primary: { terminal: "<id>" }` block in `BridgedConfig`
(nested record, same shape as `Broker`). Lets an operator pin it, or supply it when
derivation can't (see degrade case).
- **Degrade:** if the primary is off-host or in a non-herdr terminal, `terminalForPid`
@@ -71,13 +71,13 @@ A `ReplyPushLoop` component, notified at the single no-waiter call site
(`MessageService.reply` → the `inbox.publish` branch, `MessageService.java:192`).
- **Inject a nudge, not the payload.** The injected turn tells the primary *to drain*
(e.g. "Worker `<target>` returned a reply — run `fleet_poll(target=<target>)` to collect
(e.g. "Worker `<target>` returned a reply — run `bridge_poll(target=<target>)` to collect
it"), it does **not** carry the reply text. Rationale: replies can be large/multiline and
terminal injection would mangle them; the drain response is the clean transport. Keeps the
push idempotent — re-nudging is harmless.
- **Ack = drain.** The primary draining (`drainReplies` = peek + ack) is the acknowledgement.
The loop's **stop condition is `inbox.peek(target).isEmpty()`** — the reply is gone from the
inbox because it was acked. No new `fleet_ack` tool needed for v1 (see Increment 3).
inbox because it was acked. No new `bridge_ack` tool needed for v1 (see Increment 3).
- **Status-gated injection (mechanism (b), chosen).** A dedicated lightweight scheduled loop,
**not** the worker `Injector`. It injects via `AgentControl.send(primaryTerminal, nudge)`
(the same herdr `agent.send` = `pane send-text` + submit that delivers to workers) only when
@@ -93,10 +93,10 @@ A `ReplyPushLoop` component, notified at the single no-waiter call site
the reply remains in the durable inbox and the next natural poll (or a later worker reply's
nudge) still surfaces it. Bounded so the bridge never spams the primary.
### Increment 3 — optional per-`msgId` `fleet_ack` tool (deferred)
### Increment 3 — optional per-`msgId` `bridge_ack` tool (deferred)
Drain-as-ack is coarse: it clears *all* pending replies for a target at once. If finer
control is ever needed (ack one reply, leave others held), add a `fleet_ack(msgId)` tool
control is ever needed (ack one reply, leave others held), add a `bridge_ack(msgId)` tool
mapping to `inbox.ack(target, msgId)` — the port already supports per-`msgId` ack. Not built
in v1; the stop-on-empty loop is sufficient.
@@ -106,7 +106,7 @@ in v1; the stop-on-empty loop is sufficient.
resolved terminal is non-null **and not a registered worker session**, seen on an
orchestration-side tool. This never mislabels a worker (workers are in `SessionManager`)
and needs no new env var or argument (identity stays connection-derived, per the existing
`FleetMcp` invariant).
`BridgeMcp` invariant).
2. **Readiness-gate mismatch → dedicated loop.** The existing `Injector` gates delivery on
`ready.test(target)` = `WorkerPresence` (the *worker's* MCP connected). The primary is not
@@ -132,7 +132,7 @@ boundary**. The bridge is signalling the primary that it has mail — not drivin
non-null terminal AND not a registered session" predicate; the loop's stop-on-empty and
bounded-reminder logic with an injected clock + a fake injector (no real herdr).
- **Live dogfood (primary-side):** with the daemon on the broker jar + a real worker,
delegate a task, let the worker reply after the `fleet_send` window closes, and observe the
delegate a task, let the worker reply after the `bridge_send` window closes, and observe the
bridge inject a drain nudge into *this* primary pane; confirm draining stops the reminders;
confirm an unreachable primary (registry empty) degrades to pull with no loss.
+7 -37
View File
@@ -5,17 +5,17 @@
<modelVersion>4.0.0</modelVersion>
<groupId>dev.ltms</groupId>
<artifactId>fleetd</artifactId>
<artifactId>bridged</artifactId>
<version>1.0.0</version>
<packaging>jar</packaging>
<name>fleetd</name>
<description>fleet message server: sole gateway between a lead session, its members, and herdr</description>
<name>bridged</name>
<description>claude-bridge message server: sole gateway between primary/worker Claude sessions and herdr</description>
<properties>
<maven.compiler.release>25</maven.compiler.release>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
<mainClass>dev.ltms.fleet.Fleetd</mainClass>
<mainClass>dev.ltms.bridged.Bridged</mainClass>
<jackson.version>2.19.0</jackson.version>
<javalin.version>6.7.0</javalin.version>
@@ -28,8 +28,6 @@
<testcontainers.version>1.20.4</testcontainers.version>
<commons-compress.version>1.27.1</commons-compress.version>
<commons-lang3.version>3.18.0</commons-lang3.version>
<sqlite-jdbc.version>3.53.4.0</sqlite-jdbc.version>
<archunit.version>1.5.0</archunit.version>
</properties>
<!--
@@ -46,12 +44,6 @@
3.0-rc5; bumping Jackson 3 to the patched 3.2.x breaks the SDK (annotation mismatch).
Only the loopback /mcp endpoint parses this JSON, from trusted local Claude clients.
The 11.0.23 -> 11.0.25 bump did clear jetty CVE-2024-8184 (5.9) and CVE-2024-6763.
fleetd #206: org.xerial:sqlite-jdbc 3.53.4.0 (added for OpenCodeSessionDiscovery) — the
only known advisory against this artifact is CVE-2023-32697 (RCE via an attacker-controlled
JDBC URL), fixed in 3.41.2.2; 3.53.4.0 is well past that fix and OSV.dev reports no open
advisory against it. Checked via the OSV.dev API (no Mend.io/JetBrains IDE MCP mount
available from this worktree) on 2026-08-31.
-->
<!-- Force the latest patched Jetty 11.x across all Javalin-pulled Jetty modules (no version
@@ -114,7 +106,7 @@
</dependency>
<!-- MCP server: the SERVER face. Streamable-HTTP servlet mounted on Javalin's Jetty at
/mcp, exposing fleet_send/fleet_reply/fleet_status as thin adapters over REST. -->
/mcp, exposing bridge_send/bridge_reply/bridge_status as thin adapters over REST. -->
<dependency>
<groupId>io.modelcontextprotocol.sdk</groupId>
<artifactId>mcp</artifactId>
@@ -131,17 +123,6 @@
<version>${amqp.version}</version>
</dependency>
<!-- fleetd #206: opencode moved its session store from a JSON tree to SQLite
(opencode.db). This is the JDBC driver OpenCodeSessionDiscovery uses to read it
read-only. Ships bundled native libraries (linux/mac/windows, several archs), so it
is a heavier jar than most deps here — see the pom's dependency-security note below
for the size/CVE tradeoff actually measured. -->
<dependency>
<groupId>org.xerial</groupId>
<artifactId>sqlite-jdbc</artifactId>
<version>${sqlite-jdbc.version}</version>
</dependency>
<!-- Logging -->
<dependency>
<groupId>org.slf4j</groupId>
@@ -177,21 +158,10 @@
<version>${testcontainers.version}</version>
<scope>test</scope>
</dependency>
<!-- fleetd #131: package-boundary and cycle enforcement (PackageCyclesTest). -->
<dependency>
<groupId>com.tngtech.archunit</groupId>
<artifactId>archunit-junit5</artifactId>
<version>${archunit.version}</version>
<scope>test</scope>
</dependency>
</dependencies>
<build>
<!-- CB-634: the cutover renamed the module dir (bridged/ -> fleetd/), the jar, and the
launchd plist together. The installed plist names fleetd/target/fleetd.jar and
KeepAlive is armed, so this name, the plist, and the wrapper must move as one. -->
<finalName>fleetd</finalName>
<finalName>bridged</finalName>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
@@ -233,7 +203,7 @@
</configuration>
</plugin>
<!-- Runnable fat jar: java -jar target/fleetd.jar -->
<!-- Runnable fat jar: java -jar target/bridged.jar -->
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-shade-plugin</artifactId>
@@ -0,0 +1,683 @@
package dev.ltms.bridged;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.config.ConfigRef;
import dev.ltms.bridged.config.ConfigWatcher;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.LeadTabScanner;
import dev.ltms.bridged.lead.LeadLauncher;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.ExhaustedPatternLookup;
import dev.ltms.bridged.inject.ExhaustionSink;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.auth.MemberRegistry;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.mcp.BridgeMcp;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.health.FleetHealthMonitor;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
import dev.ltms.bridged.msg.AmqpReplyInbox;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.msg.LeadHeartbeatLoop;
import dev.ltms.bridged.msg.ReplyPushLoop;
import dev.ltms.bridged.rest.BridgedApp;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.session.SessionReaper;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.member.CompositePeerLauncher;
import dev.ltms.bridged.member.HerdrPeerLauncher;
import dev.ltms.bridged.member.OpenCodeLauncher;
import dev.ltms.bridged.placement.BackendQuarantine;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.Executors;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.function.Supplier;
import java.util.regex.Pattern;
import java.util.stream.Collectors;
/**
* {@code bridged} entry point. Wires the real herdr socket client to the REST app and
* starts listening. Before anything else it asserts its own environment is clean —
* {@code bridged} is not a Claude process and must never carry a base_url.
*/
public final class Bridged {
private static final Logger log = LoggerFactory.getLogger(Bridged.class);
/** CB-504: how long to wait at startup for herdr's socket before serving degraded. */
private static final long HERDR_WAIT_SECONDS = 30;
private static final long HERDR_WAIT_POLL_MILLIS = 500;
static void main(String[] args) {
Path configPath = Path.of(args.length > 0 ? args[0] : "bridged.yaml");
BridgedConfig cfg = BridgedConfig.load(configPath);
// CB-594: report which secret env vars the config actually needs, by name, before anything
// else can fail on a silently-empty one. A daemon started without a login shell (launchd)
// boots fine either way — this is the only thing that says so out loud.
reportRequiredSecrets(cfg);
// CB-596: an absent (or empty) memberCredentials: block blocks NOTHING — no credential
// name is hardcoded any more to fall back on. Say so loudly, the same way a missing
// secret is reported above, so upgrading past this commit never silently drops CB-592's
// protection.
reportMemberCredentialsGap(cfg);
// CB-559: `cfg` stays the startup snapshot — every validation and every piece of one-time
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
// contract; adding a reader here does not make a key reloadable by itself.
ConfigRef config = new ConfigRef(configPath, cfg);
// The primary/host env that launched bridged must not be tainted.
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
guard.assertPrimaryClean(System.getenv());
// CB-501: refuse to start if the bind is wider than the auth mode can defend. Under
// loopback-trust, "not a known worker" means "the primary" — sound only because the OS
// refuses remote connections to a loopback socket. This throws rather than warns so the
// dangerous configuration cannot be reached by ignoring a log line.
cfg.validateAuthExposure();
cfg.validateLeadTabPrefixes();
// CB-542: a subscription:true profile whose env: reseats ANTHROPIC_BASE_URL/AUTH_TOKEN would
// reach an unguarded endpoint (the launcher skips SubscriptionGuard for it). Refuse at load.
cfg.validateSubscriptionProfiles();
cfg.validateCharters();
// CB-548: every architect slot must name a configured workers: profile — the strong-model
// backend the future spawn lifecycle would read. A stale reference dies here, not later.
cfg.validateMembers();
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
: UnixSocketHerdrClient.defaultSocketPath();
UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect(socket, new com.fasterxml.jackson.databind.ObjectMapper());
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
Map<String, BridgedConfig.Profile> claudeProfiles = new LinkedHashMap<>();
Map<String, BridgedConfig.Profile> opencodeProfiles = new LinkedHashMap<>();
cfg.profiles().forEach((name, w) -> {
if (w.isOpenCode()) {
opencodeProfiles.put(name, w);
} else {
claudeProfiles.put(name, w);
}
});
List<HerdrPeerLauncher> adapters = new ArrayList<>();
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
// unless opencode is the only kind configured.
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
claudeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet(),
() -> config.get().memberCredentials()));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(new OpenCodeLauncher(agents, spaces,
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet(),
() -> config.get().memberCredentials()));
}
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
// The cooldown is deferred (see BridgedConfig#quarantineCooldownSeconds): it is read once
// here, at startup, and a config reload only changes it for a daemon restart.
BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
PeerLauncher workers = new CompositePeerLauncher(
adapters,
cfg.effectiveDefaultProfile(),
config,
profileName -> liveCountRef.get().apply(profileName),
quarantine);
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
// with /healthz reporting "degraded" is strictly more useful than exiting.
boolean herdrUp = awaitHerdr(herdr);
if (herdrUp) {
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
// with the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
} else {
log.warn("herdr did not answer within {}s — starting anyway; /healthz will report "
+ "degraded until it comes up. Orphaned worker panes (if any) were NOT reaped.",
HERDR_WAIT_SECONDS);
}
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
// CB-303 part 2: context cap is opt-in and disabled (0) when absent/null.
int contextCap = 0;
if (cfg.lifecycle() != null && cfg.lifecycle().contextCap() != null
&& cfg.lifecycle().contextCap() > 0) {
contextCap = cfg.lifecycle().contextCap();
}
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()),
System::nanoTime, contextCap, clearAfterTurn);
liveCountRef.set(profileName -> (int) sessions.roster().stream()
.filter(s -> profileName.equals(s.profile()))
.count());
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
final SessionReaper reaper;
if (cfg.lifecycle() != null
&& cfg.lifecycle().idleTtlSeconds() != null
&& cfg.lifecycle().idleTtlSeconds() > 0) {
reaper = new SessionReaper(sessions, cfg.lifecycle().idleTtlSeconds());
reaper.start();
} else {
reaper = null;
}
// CB-530: every pane the config names as a lead, merged from `leaders:` and the legacy
// singular pin. PrimaryRegistry below still tracks ONE terminal — it addresses the push
// loop's nudges, which need a single destination — so it keeps the legacy pin.
Map<String, String> leadTerminals = cfg.leaderTerminals();
if (leadTerminals.size() > 1) {
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
}
// CB-531: on top of the legacy primary.terminal pin, discover leads by the tab labels the
// operator writes. CB-557 moved the settings onto the lead they describe, so scanning is on
// whenever a `fleet.leaders:` entry exists — with no leads configured the supplier is a
// constant and never touches herdr, exactly as a missing `leadScan:` block used to behave.
// CB-579: each lead now names its own exact `tab:` label, so one scanner discovers every
// configured lead regardless of how differently their tabs are labelled — the old
// single-shared-tabPrefix limitation (and its warning) is gone.
final Supplier<Map<String, String>> leads;
var leaders = cfg.fleet().leaders();
if (!leaders.isEmpty()) {
Set<String> memberSpaces = cfg.profiles().values().stream()
.map(BridgedConfig.Profile::workspace)
.filter(Objects::nonNull)
.collect(Collectors.toSet());
Map<String, String> tabToName = new LinkedHashMap<>();
leaders.forEach((name, leader) -> {
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
tabToName.put(leader.tab(), name);
}
});
// One shared rescan cadence: still taken from the first entry, as before — it is an
// operational cadence, not identity, so there is no correctness reason to give every
// lead its own scanner.
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
leads = new LeadTabScanner(herdr, tabToName, memberSpaces,
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), System::nanoTime);
log.info("lead scan: tabs {} host a lead (rescan every {}s, member spaces {} excluded)",
tabToName.keySet(), scanIntervalSeconds, memberSpaces);
} else {
leads = () -> leadTerminals;
}
// CB-558: start any declared lead that is not already running. After the scanner is built,
// because both read the same tab labels and the ordering makes that dependency visible; and
// only when herdr answered, because the launcher's whole safety property is that it can
// count live leads first — it must never guess and risk a second orchestrator.
if (herdrUp && !leaders.isEmpty()) {
int launched = new LeadLauncher(agents, spaces, cfg).ensureLeads();
if (launched > 0) {
log.info("lead auto-launch: {} lead(s) started", launched);
}
}
// CB-548: config-declared architect slots. Config supplies only the stable name → profile
// map; the terminal → slot binding is owned by the registry and is empty at startup, so no
// pane resolves to an architect until the later spawn lifecycle binds one. The registry is
// what CallerResolver resolves against and what that lifecycle will read profiles from;
// nothing here spawns a slot.
MemberRegistry members = new MemberRegistry(cfg.fleet());
sessions.setMemberLifecycle(members);
if (!members.slots().isEmpty()) {
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
+ "spawn lifecycle binds a live terminal to it)",
members.slots().size(), members.slots().keySet());
}
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
Rendezvous rendezvous = new Rendezvous();
// CB-578 stage A: classify a completion-fallback scrape that matches a profile's configured
// usage-limit refusal as BACKEND_EXHAUSTED rather than handing it back as a real answer.
// Compiled once at startup, keyed by profile name; a profile with no exhaustedPattern is
// simply absent here, so its workers keep today's completion-fallback behaviour unchanged.
Map<String, Pattern> exhaustedPatternsByProfile = new LinkedHashMap<>();
cfg.profiles().forEach((name, profile) -> {
if (profile.hasExhaustedPattern()) {
exhaustedPatternsByProfile.put(name, Pattern.compile(profile.exhaustedPattern()));
}
});
ExhaustedPatternLookup exhaustedPatterns = target -> sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(session -> exhaustedPatternsByProfile.get(session.profile()))
.orElse(null);
log.info("backend-exhausted classification (CB-578 stage A): {}",
CompletionResolver.coverage(cfg.profiles().keySet(), exhaustedPatternsByProfile.keySet()));
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
// the profile config live off `config`, so a credentialId edit is hot: no restart needed.
ExhaustionSink exhaustionSink = (target, reason) -> sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(MemberSession::profile)
.map(profileName -> config.get().profiles().get(profileName))
.ifPresent(profile -> {
String credentialId = profile.effectiveCredentialId();
quarantine.quarantine(credentialId);
log.warn("credential '{}' quarantined for {}s (profile '{}' classified "
+ "BACKEND_EXHAUSTED): {}", credentialId,
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
});
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns, exhaustionSink);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
MemberPresence presence = sessions.asPresence();
TurnListener turnListener = new TurnListener() {
@Override
public void onTurnComplete(String target) {
completion.onTurnComplete(target);
sessions.onTurnComplete(target);
}
@Override
public boolean hasPostTurnAction(String target) {
return sessions.hasPostTurnAction(target);
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
completion.resolveBeforePostAction(target);
return sessions.onTurnCompleteWithPostAction(target);
}
@Override
public void onDelivered(String target, dev.ltms.bridged.msg.TurnToken token) {
completion.onDelivered(target, token);
sessions.onDelivered(target, token);
}
@Override
public void onTurnFailed(String target) {
completion.onTurnFailed(target);
sessions.onTurnFailed(target);
}
@Override
public void onTurnFailed(String target, String reason) {
completion.onTurnFailed(target, reason);
sessions.onTurnFailed(target);
}
};
Injector injector = new Injector(agents, turnListener, deliverableTo(presence, leads),
presence::forget);
StatusPoller poller = new StatusPoller(agents, injector, Injector.POLL_INTERVAL_MILLIS);
poller.start();
// CB-307: reply inbox. A broker: block (with a uri) selects the AMQP-backed durable adapter;
// absent, bridged stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
// connection, so keep the reference to close it in the ordered shutdown hook.
final ReplyInbox replyInbox;
if (cfg.broker() != null && cfg.broker().isConfigured()) {
replyInbox = AmqpReplyInbox.open(cfg.broker().uri(), cfg.broker().prefetchOrDefault());
log.info("reply inbox: AMQP broker (durable) at {} (prefetch={})",
cfg.broker().uri(), cfg.broker().prefetchOrDefault());
} else {
replyInbox = new InMemoryReplyInbox();
log.info("reply inbox: in-memory (soft-state)");
}
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
// The pin also feeds CallerResolver below: a primary running inside a herdr pane would
// otherwise resolve as a worker and be refused every orchestration tool.
String pinnedPrimaryTerminal = cfg.primary() != null ? cfg.primary().terminal() : null;
PrimaryRegistry primaryRegistry = new PrimaryRegistry(pinnedPrimaryTerminal);
// CB-532: `primary.terminal` is superseded and no longer needed for either of its jobs —
// identity comes from `leaders:`/`leadScan:`, and reply nudges now follow the delegating
// lead. Say so once at startup rather than leaving a redundant pin to look load-bearing.
if (pinnedPrimaryTerminal != null && !pinnedPrimaryTerminal.isBlank()) {
log.warn("primary.terminal is DEPRECATED (CB-532) and can be deleted: identity now comes "
+ "from leaders:/leadScan:, and reply nudges follow the lead that delegated. "
+ "It still works, and is still the fallback nudge destination when a restart "
+ "has lost the delegation map. Its pushReminders/pushBackoffMs stay valid.");
}
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
// open bridge_send. Uses its own lightweight scheduled executor, separate from the injector.
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
var pushScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-push-").unstarted(r));
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
// CB-512: the push loop takes it too, so nudge outcomes (delivered|exhausted) are counted.
Metrics metrics = BridgedMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, agents, replyInbox,
pushScheduler, maxReminders, backoffMs, metrics);
// CB-551: idle-lead heartbeat. Opt-in; absent `leadHeartbeat:` this is never constructed, so
// an upgraded daemon cannot silently start spending subscription on nudging an idle lead.
// It has its own single-thread scheduler and holds its own scheduler shutdown via close().
final LeadHeartbeatLoop heartbeat;
var heartbeatScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-heartbeat-").unstarted(r));
if (cfg.leadHeartbeat() != null) {
var hb = cfg.leadHeartbeat();
heartbeat = new LeadHeartbeatLoop(primaryRegistry, agents, replyInbox, sessions::roster,
pushLoop, heartbeatScheduler, System::nanoTime,
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
metrics);
heartbeat.start();
} else {
heartbeat = null;
heartbeatScheduler.shutdownNow();
}
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
final FleetHealthMonitor healthMonitor;
var healthScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-health-").unstarted(r));
if (cfg.health() != null && cfg.health().isEnabled()) {
// CB-580: a member found GONE/NEVER_READY must fail whatever ticket is waiting on it,
// through the same idempotent target-wide operation CB-516 already uses on release.
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
System::nanoTime, cfg.health().intervalOrDefault(), messages::abandon);
String coverage = FleetHealthMonitor.coverage(true,
cfg.health().notifications() != null && cfg.health().notifications().configured());
if ("detection-only".equals(coverage)) {
log.warn("fleet health: {} (no notification sink configured)", coverage);
} else {
log.info("fleet health: {}", coverage);
}
healthMonitor.start();
} else {
healthMonitor = null;
healthScheduler.shutdownNow();
}
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
sessions.onAcquire(replyInbox::own);
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
// reached /metrics — the delegation was unresolvable and nothing said so.
sessions.onRelease(detail -> {
// CB-578 stage C, acceptance criterion 10: a failed ticket's detail should tell a lead
// where to re-dispatch onto the same tree, not just that the worker vanished.
String reason = "the worker session was released before it replied";
if (detail.worktreePath() != null) {
reason += "; worktree=" + detail.worktreePath() + " branch=" + detail.branch()
+ " snapshot=" + (detail.snapshotRef() != null ? detail.snapshotRef() : "none");
}
// CB-584 (issue #65 criterion 5): also name the agent session, so a lead can resume the
// member's conversation instead of only re-dispatching a fresh one onto the same files.
if (detail.agentSessionId() != null) {
reason += " agentSessionId=" + detail.agentSessionId();
}
messages.abandon(detail.terminalId(), reason);
replyInbox.release(detail.terminalId());
primaryRegistry.forgetDelegation(detail.terminalId()); // CB-532: don't leak the lead binding
});
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
ConnectionIdentity identity = new ConnectionIdentity(
new PaneLocator(herdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
final CallerResolver callers;
if (cfg.auth().tokenMode()) {
String token = System.getenv(cfg.auth().tokenEnv());
if (token == null || token.isBlank()) {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting bridged");
}
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
BridgeMcp mcp = new BridgeMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics, new BridgeMcp.CapacitySource(profile -> liveCountRef.get().apply(profile),
profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.maxLoad();
}, () -> config.get().profiles().keySet(), System::nanoTime),
new BridgeMcp.HealthCoverageSource(() -> {
var health = config.get().health();
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
health != null && health.notifications() != null && health.notifications().configured());
}),
new BridgeMcp.QuarantineSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, quarantine));
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed, so an
// upgraded daemon behaves exactly as before — the file is read once at boot and never again.
final ConfigWatcher configWatcher;
if (cfg.configReload() != null && cfg.configReload().isEnabled()) {
configWatcher = new ConfigWatcher(config, cfg.configReload().intervalSeconds());
configWatcher.start();
} else {
configWatcher = null;
}
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
// last. This replaces the earlier independent hooks that could race and close herdr early.
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
poller.stop();
messages.close();
pushLoop.close();
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
if (healthMonitor != null) healthMonitor.stop();
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
mcp.close();
if (reaper != null) reaper.stop();
// Release the broker connection last among message resources (no-op for the in-memory inbox).
if (replyInbox instanceof AutoCloseable closeable) {
try {
closeable.close();
} catch (Exception e) {
log.debug("reply inbox close: {}", e.toString());
}
}
herdr.close();
}));
Javalin app = new BridgedApp(herdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics).build();
app.start(cfg.bind().host(), cfg.bind().port());
log.info("bridged listening on {}:{}, herdr socket {}",
cfg.bind().host(), cfg.bind().port(), socket);
}
/**
* The {@link Injector}'s readiness gate (CB-534): a target is deliverable if it is a spawned
* member whose agent has connected the bridge MCP, <em>or</em> a lead.
*
* <p>The gate exists for one reason — to hold a delivery out of a <em>spawned</em> member's boot
* window, where herdr already reports {@code idle} but the TUI would drop an injected paste. That
* hazard is a property of spawning. A lead is never spawned: the operator started it and named it
* (or labelled its tab) only once it was up, so there is no boot window to guard.
*
* <p>A lead is also never enrolled in {@link MemberPresence} — {@code BridgeMcp} marks presence
* for every spawned member (worker and architect), deliberately, since that map doubles as the
* member roster's availability signal and a lead counted there would show up as an available
* member. So without the second disjunct a lead is permanently un-deliverable: every
* lead→lead send sat on the gate for {@code READINESS_GRACE_POLLS} (~60s) and then failed
* having never been typed into the pane.
*
* <p>The lead set is read through the supplier on each call rather than snapshotted, so a lead
* discovered by {@code leadScan} after startup becomes deliverable without a restart.
*/
static Predicate<String> deliverableTo(MemberPresence presence, Supplier<Map<String, String>> leads) {
return target -> presence.isPresent(target) || leads.get().containsKey(target);
}
/**
* CB-594: which env vars the loaded config actually needs, and why — every non-{@code
* subscription} profile's {@code tokenEnv} (a subscription profile never reads one, see
* {@link BridgedConfig.Profile#isSubscription()}), plus every profile's {@code gitTokenEnv}
* where set (opt-in). Derived from the config, not hard-coded, so a new profile is covered for
* free. A var required by more than one profile is one entry naming every profile that needs
* it. Deliberately excludes {@code auth.tokenEnv}: that one is already enforced loudly, by a
* startup throw in {@code main()} — about 370 lines <em>below</em> this method's call site
* ({@link #reportRequiredSecrets(BridgedConfig)}), not a few lines above it. That throw only
* fires when {@code auth.mode: token} is configured; under the default loopback-trust mode it
* never runs, and {@code auth.tokenEnv} is simply not required.
*
* <p>Package-private and pure (no I/O, no logging) so the derivation is unit-testable without
* capturing log output; {@link #reportRequiredSecrets(BridgedConfig)} is the logging caller.
*/
static Map<String, List<String>> requiredSecretEnvVars(BridgedConfig cfg) {
Map<String, List<String>> requiredBy = new LinkedHashMap<>();
cfg.profiles().forEach((name, profile) -> {
if (!profile.isSubscription()) {
requiredBy.computeIfAbsent(profile.tokenEnv(), _ -> new ArrayList<>())
.add("profile '" + name + "' tokenEnv");
}
if (profile.hasGitToken()) {
requiredBy.computeIfAbsent(profile.gitTokenEnv(), _ -> new ArrayList<>())
.add("profile '" + name + "' gitTokenEnv");
}
});
return requiredBy;
}
/**
* CB-594: log, by name only, which required env vars (see {@link #requiredSecretEnvVars}) are
* set in the daemon's own process environment — the environment every profile's {@code
* tokenEnv}/{@code gitTokenEnv} is read from at spawn time (see
* {@code HerdrPeerLauncher.resolveEnv}). Never logs a value, a prefix, or a length.
*
* <p>A missing entry only warns — it must never refuse to start. A daemon that boots and says
* what is wrong is strictly more useful than one that will not boot at all.
*/
private static void reportRequiredSecrets(BridgedConfig cfg) {
Map<String, List<String>> requiredBy = requiredSecretEnvVars(cfg);
if (requiredBy.isEmpty()) {
log.info("startup secrets: no profile references a token env var — nothing to check");
return;
}
Map<String, String> env = System.getenv();
requiredBy.forEach((varName, sources) -> {
String value = env.get(varName);
if (value != null && !value.isBlank()) {
log.info("startup secret {}: set ({})", varName, String.join(", ", sources));
} else {
log.warn("startup secret {}: MISSING ({}) — the daemon will start anyway, and this "
+ "failure stays invisible until a worker actually needs it. Fix "
+ "${SHARED_ENV}/tools/secrets.sh and restart bridged from a LOGIN "
+ "shell (see scripts/redeploy-bridged.sh).",
varName, String.join(", ", sources));
}
});
}
/**
* CB-596: {@code known:} empty (block absent entirely, or present but empty) means {@link
* BridgedConfig.MemberCredentials#blockedSet()} is empty too — every member pane inherits the
* operator's whole secret store, unblocked, exactly the defect this ticket fixes. Unlike a
* missing token ({@link #reportRequiredSecrets}), there is no name to point at: the point is
* that the block itself is missing. Warn once at startup and say what to add; never refuse to
* start over it — see {@link #reportRequiredSecrets} for why a daemon that boots and says
* what is wrong beats one that will not boot at all.
*
* <p>Package-private so the test can capture the log directly, the same way {@link
* #requiredSecretEnvVars} is exposed for {@link #reportRequiredSecrets}'s own test.
*/
static void reportMemberCredentialsGap(BridgedConfig cfg) {
BridgedConfig.MemberCredentials creds = cfg.memberCredentials();
if (creds != null && !creds.known().isEmpty()) {
log.info("memberCredentials: {} known name(s), {} allowed — blocking {} on every spawn",
creds.known().size(), creds.allow().size(), creds.blockedSet().size());
return;
}
log.warn("memberCredentials: absent or empty — the daemon will start anyway, and every "
+ "member pane inherits the operator's WHOLE secret store, unblocked (CB-592's "
+ "protection is lost). Add a memberCredentials: block (policy/allow/known) to "
+ "bridged.yaml — see bridged.example.yaml — and restart.");
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
*
* @return true if herdr answered, false if it never did
*/
private static boolean awaitHerdr(HerdrClient herdr) {
long deadline = System.nanoTime() + HERDR_WAIT_SECONDS * 1_000_000_000L;
boolean waited = false;
while (true) {
try {
herdr.call("ping");
if (waited) {
log.info("herdr is up");
}
return true;
} catch (HerdrException e) {
if (System.nanoTime() >= deadline) {
return false;
}
if (!waited) {
log.info("waiting up to {}s for the herdr socket…", HERDR_WAIT_SECONDS);
waited = true;
}
try {
Thread.sleep(HERDR_WAIT_POLL_MILLIS);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
return false;
}
}
}
}
private Bridged() {
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -1,9 +1,9 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
/**
* The authorization table (CB-505), stated once and enforced on both entry paths.
*
* <p>Most of these rules are already true de facto — {@code FleetMcp} derives a worker's identity
* <p>Most of these rules are already true de facto — {@code BridgeMcp} derives a worker's identity
* from the connection rather than reading it from an argument, so a worker has never been able to
* reply <em>as</em> another worker over MCP. What was missing is that the REST surface trusted the
* session id in the URL path, and neither surface checked role at all. This class makes the
@@ -30,27 +30,8 @@ public final class Authz {
DRAIN,
/** Read-only observation: status, roster, profiles, task polling. */
READ,
/**
* Read (never ack) this daemon's own held lead-to-lead coordination mail (fleetd #421).
*
* <p>Deliberately <strong>not</strong> folded into {@link #READ}. {@code READ}'s grant
* rests on "the roster carries no secrets" (see its case below) — a lead-to-lead body is
* not the roster; it is where leads discuss host shapes, credentials and unmerged work.
* Mapping this to {@code READ} would let any worker read every peer lead's mail in full
* and would silently falsify that comment for every other {@code READ} caller.
*/
COORD_READ,
/** Scrape the metrics endpoint. */
METRICS,
/**
* Drive the lead-rollover executor ({@code fleet_handover}: open/confirm/cancel a
* self-replace, fleetd #480 Unit C). Primary-only, same as {@link #SPAWN}/{@link #STOP}/
* {@link #DRAIN} — and, unlike those, the terminal it acts on is never even an argument:
* {@code LeadRollover#open}/{@code #confirm} are always called with the CALLER's own
* connection-resolved terminal (see {@code dev.ltms.fleet.lead.LeadRollover}'s class
* javadoc, fleetd #480 correction 2), so a primary can only ever roll itself.
*/
HANDOVER
METRICS
}
/**
@@ -69,7 +50,7 @@ public final class Authz {
// deliberately does NOT get these (CB-548), so it cannot tear down or stand up workers
// even though it coordinates them; and a worker driving any of these would be a worker
// escalating into the orchestrator role.
case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary();
case SPAWN, STOP, DRAIN -> caller.isPrimary();
// Delivering a turn is open to the primary and the architect: an architect delegates
// to workers (that is the role's point) but still has no lifecycle rights. A worker is
@@ -87,11 +68,6 @@ public final class Authz {
// Observation is open to every authenticated role: a worker legitimately polls its own
// status, and the roster carries no secrets.
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
// fleetd #421: reading held lead-to-lead mail is the primary's alone. An architect
// holds READ today (CB-548), so "not primary" must mean not-architect here too — this
// is coordination between leads, not observation of the roster.
case COORD_READ -> caller.isPrimary();
};
}
@@ -1,7 +1,7 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.peer.MemberRole;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
@@ -12,7 +12,7 @@ import java.util.function.Supplier;
/**
* Resolves every caller to a {@link Principal}, for both entry paths into the core (CB-501).
*
* <p>There are two of them and they are not layered the way the docs suggest: {@code FleetMcp}
* <p>There are two of them and they are not layered the way the docs suggest: {@code BridgeMcp}
* calls the service layer directly and is mounted as a raw servlet (so it never passes through a
* Javalin filter), while the REST routes historically resolved no identity at all. Both now
* delegate here, so the authorization rules are stated once instead of drifting apart.
@@ -59,7 +59,7 @@ public final class CallerResolver {
* <p>Like {@link #leadTerminals}, a supplier rather than a fixed map, so a binding injected
* after startup — when the later spawn lifecycle establishes a live architect session, or an
* operator pins one — takes effect without a restart. Consulted per resolve; today's wiring
* in {@code Fleetd} reads a constant from config, which is the degenerate live case.
* in {@code Bridged} reads a constant from config, which is the degenerate live case.
*/
private final Supplier<Map<String, String>> architectTerminals;
private final Function<String, MemberRole> memberSlotRoles;
@@ -221,8 +221,7 @@ public final class CallerResolver {
// The config/live binding names this pane as an architect slot's own. Same
// unforgeable pane mapping; the live binding, never a request argument, decides.
// Check the slot role too: this defence in depth prevents a bad lifecycle bind from
// escalating a dev, hunter or reviewer into an architect. Checked before
// the worker fallback.
// escalating a dev or reviewer into an architect. Checked before the worker fallback.
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
}
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
@@ -236,24 +235,8 @@ public final class CallerResolver {
// loopback-trust: same-host callers that are not workers are the primary. A non-loopback
// caller is anonymous even here — and startup refuses that combination anyway
// (FleetConfig.validateAuthExposure), so this is defence in depth, not the control.
//
// fleetd #317: "not a worker" must not be conflated with "identity unresolved". The real
// primary is a real process — its pid resolves (c.resolved()), it just owns no herdr pane.
// A caller whose peer-PID lookup failed (LsofPeerPidLookup's -1 sentinel — on any failure,
// silently including "lsof found no match") has no such pid, and PaneLocator's own javadoc
// already names what happens if that case is handed the primary role: a worker→primary
// escalation. So an unresolved caller is refused (ANONYMOUS — the same clean, already-tested
// "authenticated as nothing" outcome used everywhere else in this method), never promoted.
//
// fleetd #505: the OTHER way a real pid can wrongly reach here with a null terminal — not a
// failed lsof lookup, but a herdr error partway through PaneLocator's pane scan. c.resolved()
// says nothing about that; it only tests the lsof sentinel (by design — see
// ConnectionIdentity.Caller#resolved). c.scanComplete() is the separate signal: a scan that
// could not check every pane must not be read as "checked everywhere, no match" — the pane it
// could not check might have been the caller's own. So both must hold before this promotes.
return isLoopback(remoteAddr) && c.resolved() && c.scanComplete()
? Principal.primary(c.pid()) : Principal.anonymous();
// (BridgedConfig.validateAuthExposure), so this is defence in depth, not the control.
return isLoopback(remoteAddr) ? Principal.primary(c.pid()) : Principal.anonymous();
}
private boolean presentedTokenMatches(String authorizationHeader) {
@@ -279,15 +262,11 @@ public final class CallerResolver {
return token.isEmpty() ? null : token;
}
/**
* fleetd #305: delegates to {@link ConnectionIdentity#isLoopback}. This used to be a second,
* independent copy of the same rule, and the two drifted: this one accepted all of
* {@code 127.0.0.0/8}, {@code ConnectionIdentity}'s accepted only {@code 127.0.0.1}. A caller
* from {@code 127.0.0.2} therefore had its identity skipped (so it had no terminal) and was
* then read as loopback here — which under loopback-trust is the primary. Sharing the inputs
* would not have prevented that; only sharing the computation does.
*/
private static boolean isLoopback(String remoteAddr) {
return ConnectionIdentity.isLoopback(remoteAddr);
if (remoteAddr == null) {
return false;
}
return remoteAddr.equals("127.0.0.1") || remoteAddr.equals("::1")
|| remoteAddr.equals("0:0:0:0:0:0:0:1") || remoteAddr.startsWith("127.");
}
}
@@ -0,0 +1,21 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.peer.MemberRole;
/** Optional session lifecycle hook for live member-slot bindings. */
public interface MemberLifecycle {
MemberLifecycle NONE = new MemberLifecycle() {
@Override
public void acquired(MemberRole role, String profile, String terminal) {
}
@Override
public void released(String terminal) {
}
};
void acquired(MemberRole role, String profile, String terminal);
void released(String terminal);
}
@@ -0,0 +1,232 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.peer.MemberRole;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.HashMap;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Objects;
/**
* The architect-slot registry (CB-548): every gateway-local architect name and the strong-model
* profile it points at, plus the <em>live</em> bindings from a live architect's herdr terminal to
* its slot.
*
* <p>Two halves, split by who owns each:
* <ul>
* <li><b>slots</b> — configured once, keyed by the gateway-local unique name; each carries the
* {@code profile} reference the spawn lifecycle reads when it stands the slot up. A read-only
* snapshot taken at construction.</li>
* <li><b>terminal bindings</b> — owned by this registry and initially <em>empty</em>. Config
* declares no architect terminal, so at startup every slot is idle and nothing resolves to an
* architect; a session only becomes one when the spawn lifecycle {@linkplain #bind(String,
* String) binds} its terminal to a slot. {@link CallerResolver} reads this through
* {@link #snapshot()} to turn a pane into an {@link Role#ARCHITECT}.</li>
* </ul>
*
* <p>Spawning/lifecycle is deliberately a separate unit: this class only owns the bindings and
* exposes the map the resolver resolves against plus the profile lookup lifecycle will call.
* Nothing here creates or manages an architect session.
*/
public final class MemberRegistry implements MemberLifecycle {
private static final Logger log = LoggerFactory.getLogger(MemberRegistry.class);
/**
* One flattened {@code fleet:} entry.
*
* <p>Flattened because a slot name is unique only <em>within</em> its pool — {@code sonnet} may
* legitimately be both a developer and a reviewer — while a terminal binds to exactly one thing.
* The qualified {@link #key()} is what that binding uses.
*
* @param name the slot's key inside its pool
* @param role the pool it came from
* @param profile the backend it runs on
*/
public record Entry(String name, MemberRole role, String profile) {
/** {@code "architect:opus"} — unique across pools, unlike {@link #name()}. */
public String key() {
return role.wireName() + ":" + name;
}
}
private final Map<String, Entry> slots;
/** Live {@code terminal_id → qualified slot key}; guarded by {@code this}. */
private final Map<String, String> terminalToSlot = new HashMap<>();
/** Flatten every role pool in {@code fleet} into one registry. Leaders are not members. */
public MemberRegistry(BridgedConfig.Fleet fleet) {
Map<String, Entry> flat = new LinkedHashMap<>();
if (fleet != null) {
for (MemberRole role : MemberRole.values()) {
fleet.pool(role).forEach((name, slot) -> {
if (slot != null) {
Entry e = new Entry(name, role, slot.profile());
flat.put(e.key(), e);
}
});
}
}
this.slots = Collections.unmodifiableMap(flat);
}
/** The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot. */
public Map<String, Entry> slots() {
return slots;
}
/** The slots belonging to {@code role}, in definition order. */
public Map<String, Entry> slotsFor(MemberRole role) {
Map<String, Entry> out = new LinkedHashMap<>();
slots.forEach((key, e) -> {
if (e.role() == role) {
out.put(key, e);
}
});
return Collections.unmodifiableMap(out);
}
/**
* An immutable copy of the live {@code terminal_id → slot name} bindings.
*
* <p>Passed to {@link CallerResolver} as the source of architect identity, and what
* {@code bridge_whoami}/the roster will read to say which slot a pane hosts. Empty until the
* spawn lifecycle binds a slot.
*/
public Map<String, String> snapshot() {
synchronized (terminalToSlot) {
return Map.copyOf(terminalToSlot);
}
}
/** The slot a live terminal is bound to, or {@code null} if it is not an architect slot. */
public String slotForTerminal(String terminal) {
if (terminal == null) {
return null;
}
synchronized (terminalToSlot) {
return terminalToSlot.get(terminal);
}
}
/**
* The strong-model profile a slot runs under — what the spawn lifecycle reads.
*
* @return the slot's configured {@code profile}, or {@code null} if the slot is unknown or
* declares none
*/
public String profileForSlot(String slotName) {
Entry e = slots.get(slotName);
return (e == null || e.profile() == null) ? null : e.profile();
}
/** The role a qualified slot key belongs to, or {@code null} when the key is unknown. */
public MemberRole roleForSlot(String slotName) {
Entry e = slots.get(slotName);
return e == null ? null : e.role();
}
/** The unqualified configured name for a slot, or {@code null} if it is unknown. */
public String nameForSlot(String slotName) {
Entry e = slots.get(slotName);
return e == null ? null : e.name();
}
/** True when {@code slotName} is a configured architect slot. */
public boolean isSlot(String slotName) {
return slots.containsKey(slotName);
}
/**
* Bind {@code terminal} to {@code slot} (CB-548).
*
* <p>The spawn lifecycle calls this when it stands a slot up. The bind is atomic and preserves
* the two cardinality invariants: a terminal may occupy at most one slot, and a slot may host at
* most one terminal. Binding the same terminal to the same slot again is a harmless no-op.
*
* @param slot a configured slot name, or the bind is refused
* @param terminal the pane that will act as this architect
* @return {@code true} if the binding is now {@code terminal → slot}; {@code false} if it was
* refused — an unknown slot, a terminal already bound to a different slot, or a slot
* already hosting a different terminal
*/
public boolean bind(String slot, String terminal) {
if (slot == null || terminal == null || terminal.isBlank()) {
return false;
}
synchronized (terminalToSlot) {
if (!isSlot(slot)) {
return false; // unknown slot — nothing to bind to
}
String existingSlot = terminalToSlot.get(terminal);
if (existingSlot != null) {
return slot.equals(existingSlot); // already this slot (idempotent) or a different one
}
if (terminalToSlot.containsValue(slot)) {
return false; // slot already hosts a terminal — no second one
}
terminalToSlot.put(terminal, slot);
return true;
}
}
/**
* Compare-safe unbind of {@code expectedTerminal} from {@code slot} (CB-548).
*
* <p>The spawn lifecycle calls this when it tears a slot down. Only the exact binding
* {@code expectedTerminal → slot} is removed; if that terminal was since rebound to a different
* slot (or the slot to a different terminal), the call is a no-op returning {@code false} — a
* stale unbind must never remove a replacement.
*
* @param slot the slot the caller believes the terminal is bound to
* @param expectedTerminal the terminal it expects to be bound there
* @return {@code true} if {@code expectedTerminal → slot} was removed; {@code false} if nothing
* was (no such binding, or the binding had already moved)
*/
public boolean unbind(String slot, String expectedTerminal) {
if (slot == null || expectedTerminal == null) {
return false;
}
synchronized (terminalToSlot) {
String current = terminalToSlot.get(expectedTerminal);
if (current == null || !slot.equals(current)) {
return false; // absent, or a replacement/moved binding — leave it in place
}
terminalToSlot.remove(expectedTerminal);
return true;
}
}
/**
* Bind only architect sessions to a free slot with the resolved profile.
*
* <p>The role check is lifecycle policy. {@link CallerResolver} repeats it when resolving a
* binding, so a later lifecycle regression cannot turn a worker into an architect.
*/
@Override
public void acquired(MemberRole role, String profile, String terminal) {
if (role != MemberRole.ARCHITECT || terminal == null || terminal.isBlank()) {
return;
}
// slotsFor preserves definition order, so duplicate-profile slots use the first free one.
for (Entry entry : slotsFor(MemberRole.ARCHITECT).values()) {
if (Objects.equals(profile, entry.profile()) && bind(entry.key(), terminal)) {
return;
}
}
log.info("member slot: no free architect slot for profile={}; session remains a worker", profile);
}
/** Unbind a released terminal using the compare-safe registry operation. */
@Override
public void released(String terminal) {
String slot = slotForTerminal(terminal);
if (slot != null) {
unbind(slot, terminal);
}
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
/**
* A resolved caller: its {@link Role}, and — for a worker — the herdr {@code terminal_id} that
@@ -39,14 +39,14 @@ public record Principal(Role role, String terminal, long pid, String name) {
*
* <p>Carries {@link Role#PRIMARY}: a lead <em>is</em> a primary as far as authorization goes,
* so every existing {@code isPrimary()} gate keeps working unchanged and the role table needed
* no new entry. The name is reporting only — it lets {@code fleet_whoami} say <em>which</em>
* no new entry. The name is reporting only — it lets {@code bridge_whoami} say <em>which</em>
* lead is asking once more than one is configured.
*
* <p><strong>CB-532: a lead now carries the terminal it was matched by.</strong> Under CB-530 it
* deliberately did not, because {@code terminal} meant "which worker pane" everywhere and a
* non-null one would have enrolled the lead in the worker presence map. That reading was what
* made a lead unaddressable: {@link #ownsSession} could never be true for it, so
* {@code fleet_reply} was refused and one lead could send to another but never be answered.
* {@code bridge_reply} was refused and one lead could send to another but never be answered.
* The terminal now means "which pane is this caller", the presence map keys on
* {@link #isSpawnedMember()} instead, and a lead is a peer that can both send and receive.
*/
@@ -63,7 +63,7 @@ public record Principal(Role role, String terminal, long pid, String name) {
* An architect (CB-548), identified by the slot it occupies and the pane bound to it.
*
* <p>Carries {@link Role#ARCHITECT}. {@code slotName} is reporting only — it lets
* {@code fleet_whoami} say <em>which</em> architect slot is asking, and it is the key the
* {@code bridge_whoami} say <em>which</em> architect slot is asking, and it is the key the
* (future) spawn lifecycle reads a profile back from. Identity is the {@code terminal}: like a
* worker's it comes from the connection and the live terminal→slot binding, so
* {@code ownsSession} works exactly as it does for a worker — an architect acts as its own
@@ -1,4 +1,4 @@
package dev.ltms.fleet.auth;
package dev.ltms.bridged.auth;
/**
* What a caller is allowed to be on the bus (CB-501).
@@ -0,0 +1,291 @@
package dev.ltms.bridged.config;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Supplier;
/**
* The daemon's live configuration, re-readable without a restart (CB-559).
*
* <p>Consumers hold this, not a {@link BridgedConfig}, and read through {@link #get()} at the point
* of use. A component that captures {@code ref.get()} into a field at construction has opted out of
* reload — which is sometimes right (see <em>deferred</em> below), but it must then be a deliberate
* choice rather than an accident of where the field was initialised.
*
* <h2>Not every key can change under a running daemon</h2>
* Keys fall into three classes, and the difference is about what already exists when the reload
* happens — not about how important the key is.
*
* <ul>
* <li><strong>Hot</strong> — re-read per use, so a reload takes effect on the next spawn:
* {@code fleet:} (every role pool, {@code charters}, and {@code tabLabel}),
* {@code placement:}, and an existing profile's {@code weight} / {@code maxLoad}. Those
* three are read through a supplier on {@code CompositePeerLauncher}, which is what makes
* them hot — not the fact that they are config. <strong>This does NOT include
* {@code fleet.leaders}</strong>: {@code Bridged.main} reads {@code cfg.fleet().leaders()}
* once at startup to build the {@code LeadTabScanner} and the {@code LeadLauncher}, and
* neither is reconstructed on reload — so a lead added, removed, or re-{@code tab}'d under
* {@code fleet.leaders} needs a restart, the same as any deferred key below.</li>
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code quarantineCooldownSeconds}
* (CB-578 stage B — baked once into the {@code BackendQuarantine} built at startup),
* {@code guard:}, {@code worktreeRoot:}, adding or removing a profile (a new backend needs its own launcher,
* which is constructed once), <em>and an existing profile's launch settings</em> —
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl},
* {@code exhaustedPattern} (CB-578 stage A — compiled once into {@code Bridged.main}'s
* pattern map at startup), and the rest. {@code credentialId} (CB-578 stage B) is NOT on
* this list — it is read live off the config supplier at every quarantine check and
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so it is hot instead.
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} at construction and resolves
* each spawn out of that copy, so those never reach a launch until the daemon restarts. A
* reload logs these rather than pretending they applied.</li>
* <li><strong>Cold</strong> — cannot change at all under a running daemon: {@code bind:},
* {@code herdrSocket:}, {@code broker:} and {@code auth:}. The socket is bound, the broker
* connection is open, and the auth mode decides who may reach the port that is already
* listening.</li>
* </ul>
*
* <p><strong>A cold change refuses the whole reload.</strong> Not the hot half applied and the cold
* half warned about: that would leave the running daemon in a state matching no file on disk, which
* is the worst thing a reload can do to an operator debugging one. Refusing keeps the invariant that
* the live config is always some version of the file, and the message names the keys that must
* change through a restart.
*
* <p>A reload that fails to parse or fails validation is also refused, and the previous config keeps
* running. A config file being edited is normally read once mid-save; degrading a working daemon
* because it caught a half-written file would be a bad trade.
*/
public final class ConfigRef implements Supplier<BridgedConfig> {
private static final Logger log = LoggerFactory.getLogger(ConfigRef.class);
/** Keys that cannot change under a running daemon — see the class doc. */
private static final Set<String> COLD_KEYS =
Set.of("bind", "herdrSocket", "broker", "auth");
private final Path path;
private final AtomicReference<BridgedConfig> current;
public ConfigRef(Path path, BridgedConfig initial) {
this.path = path;
this.current = new AtomicReference<>(Objects.requireNonNull(initial, "initial config"));
}
/** A fixed reference that never reloads — for tests and for wiring built from a config in code. */
public static ConfigRef fixed(BridgedConfig cfg) {
return new ConfigRef(null, cfg);
}
/** The live configuration. Read this per use; do not cache it in a field. */
@Override
public BridgedConfig get() {
return current.get();
}
/** The file this ref reloads from, or {@code null} for a {@link #fixed} ref. */
public Path path() {
return path;
}
/**
* What a reload attempt did.
*
* @param applied true when the new config is now live
* @param coldKeys cold keys whose value changed, which is why an unapplied reload was refused
* @param deferred keys that changed and were accepted, but whose effect waits for a restart
* @param error the parse or validation failure that refused the reload, else {@code null}
*/
public record Outcome(boolean applied, List<String> coldKeys, List<String> deferred,
String error) {
public Outcome {
coldKeys = List.copyOf(coldKeys);
deferred = List.copyOf(deferred);
}
static Outcome refusedCold(List<String> keys) {
return new Outcome(false, keys, List.of(), null);
}
static Outcome failed(String error) {
return new Outcome(false, List.of(), List.of(), error);
}
/** A one-line summary for the operator — the reason, not just the verdict. */
public String summary() {
if (error != null) {
return "config reload refused — " + error;
}
if (!applied) {
return "config reload refused — these keys cannot change under a running daemon: "
+ String.join(", ", coldKeys) + ". Restart bridged to apply them.";
}
if (!deferred.isEmpty()) {
return "config reloaded; these changes need a restart to take effect: "
+ String.join(", ", deferred);
}
return "config reloaded";
}
}
/**
* Re-read the file, validate it, and swap it in when nothing cold changed.
*
* <p>Never throws: a reload is a best-effort operation on a daemon that is already serving, and
* a bad edit must not take it down. Every failure path leaves the previous config live and is
* reported through the returned {@link Outcome}.
*/
public Outcome reload() {
if (path == null) {
return Outcome.failed("this config was built in code and has no file to reload from");
}
BridgedConfig old = current.get();
BridgedConfig fresh;
try {
fresh = BridgedConfig.load(path);
// The same gate startup runs. A config that would have refused to boot must not be able
// to slip in through a reload — that is how a daemon ends up in a state it could never
// have started in, which is the hardest kind to debug.
fresh.validateAuthExposure();
fresh.validateLeadTabPrefixes();
fresh.validateSubscriptionProfiles();
fresh.validateCharters();
fresh.validateMembers();
} catch (RuntimeException e) {
String msg = e.getMessage() == null ? e.toString() : e.getMessage();
log.warn("config reload from {} refused, keeping the running config: {}", path, msg);
return Outcome.failed(msg);
}
List<String> cold = changedColdKeys(old, fresh);
if (!cold.isEmpty()) {
Outcome out = Outcome.refusedCold(cold);
log.warn(out.summary());
return out;
}
List<String> deferred = changedDeferredKeys(old, fresh);
current.set(fresh);
Outcome out = new Outcome(true, List.of(), deferred, null);
log.info(out.summary());
return out;
}
/** Cold keys whose value differs between the running config and the candidate. */
private static List<String> changedColdKeys(BridgedConfig old, BridgedConfig fresh) {
List<String> changed = new ArrayList<>();
if (!Objects.equals(old.bind(), fresh.bind())) {
changed.add("bind");
}
if (!Objects.equals(old.herdrSocket(), fresh.herdrSocket())) {
changed.add("herdrSocket");
}
if (!Objects.equals(old.broker(), fresh.broker())) {
changed.add("broker");
}
if (!Objects.equals(old.auth(), fresh.auth())) {
changed.add("auth");
}
// Kept in step with COLD_KEYS so the doc and the code cannot drift apart silently.
assert COLD_KEYS.containsAll(changed) : "a cold key was reported that COLD_KEYS omits";
return changed;
}
/** Changed keys that were accepted but whose effect waits for a restart. */
private static List<String> changedDeferredKeys(BridgedConfig old, BridgedConfig fresh) {
List<String> changed = new ArrayList<>();
if (!Objects.equals(old.lifecycle(), fresh.lifecycle())) {
changed.add("lifecycle");
}
if (!Objects.equals(old.leadHeartbeat(), fresh.leadHeartbeat())) {
changed.add("leadHeartbeat");
}
if (!Objects.equals(old.guard(), fresh.guard())) {
changed.add("guard");
}
if (!Objects.equals(old.worktreeRoot(), fresh.worktreeRoot())) {
changed.add("worktreeRoot");
}
if (!Objects.equals(old.spawnReadyTimeoutMs(), fresh.spawnReadyTimeoutMs())
|| !Objects.equals(old.spawnReadyPollMs(), fresh.spawnReadyPollMs())) {
changed.add("spawnReady*");
}
// CB-578 stage B: baked once into the BackendQuarantine built at startup — a running
// quarantine keeps its original cooldown regardless, and a new cooldown only applies to a
// quarantine that starts after a restart.
if (!Objects.equals(old.quarantineCooldownSeconds(), fresh.quarantineCooldownSeconds())) {
changed.add("quarantineCooldownSeconds");
}
Map<String, BridgedConfig.Profile> before =
old.profiles() == null ? Map.of() : old.profiles();
Map<String, BridgedConfig.Profile> after =
fresh.profiles() == null ? Map.of() : fresh.profiles();
// Adding or removing a profile is deferred: a new backend needs its own launcher, and
// launchers are built once at startup.
if (!before.keySet().equals(after.keySet())) {
Set<String> diff = new LinkedHashSet<>(before.keySet());
diff.addAll(after.keySet());
diff.removeIf(p -> before.containsKey(p) && after.containsKey(p));
changed.add("profiles (added/removed: " + String.join(", ", diff) + ")");
}
// An EXISTING profile's launch settings are deferred too, and this is easy to get wrong:
// `HerdrPeerLauncher` takes `Map.copyOf(profiles)` at construction and `spawn` resolves the
// profile out of that snapshot, so a reloaded model/baseUrl/argv/env never reaches a launch.
// Only weight and maxLoad are genuinely hot, because placement reads them through the
// supplier on the composite rather than from the adapter's copy. Without this check a
// changed model would report "config reloaded" and silently do nothing — the worst outcome
// a reload can produce, because the operator has no reason to doubt it.
List<String> relaunch = new ArrayList<>();
before.forEach((name, was) -> {
BridgedConfig.Profile now = after.get(name);
if (now != null && !sameLaunchSettings(was, now)) {
relaunch.add(name);
}
});
if (!relaunch.isEmpty()) {
changed.add("profiles." + String.join("/", relaunch) + " launch settings "
+ "(model, baseUrl, argv, env, …) — the launcher holds a startup snapshot");
}
return changed;
}
/**
* Whether two versions of a profile would launch a peer identically. Compares every component
* the launcher reads at spawn; {@code weight}, {@code maxLoad} and {@code credentialId} are
* excluded because those are read live (by the placement policy and, for credentialId, by
* {@code CompositePeerLauncher}/the CB-578 stage B exhaustion sink) and really do take effect on
* the next spawn.
*/
private static boolean sameLaunchSettings(BridgedConfig.Profile a, BridgedConfig.Profile b) {
return Objects.equals(a.baseUrl(), b.baseUrl())
&& Objects.equals(a.model(), b.model())
&& Objects.equals(a.configDir(), b.configDir())
&& Objects.equals(a.tokenEnv(), b.tokenEnv())
&& Objects.equals(a.argv(), b.argv())
&& Objects.equals(a.placement(), b.placement())
&& Objects.equals(a.workspace(), b.workspace())
&& Objects.equals(a.tabLabel(), b.tabLabel())
&& Objects.equals(a.mcpUrl(), b.mcpUrl())
&& Objects.equals(a.cwd(), b.cwd())
&& Objects.equals(a.parityOverlay(), b.parityOverlay())
&& Objects.equals(a.gitTokenEnv(), b.gitTokenEnv())
&& Objects.equals(a.gitHostEnv(), b.gitHostEnv())
&& Objects.equals(a.kind(), b.kind())
&& Objects.equals(a.env(), b.env())
&& Objects.equals(a.subscription(), b.subscription())
// CB-578 stage B: exhaustedPattern is compiled once into Bridged.main's pattern map
// at startup (see ExhaustedPatternLookup wiring) — a reload never re-reads it, so a
// changed pattern must be reported as deferred, exactly like model/baseUrl/argv.
&& Objects.equals(a.exhaustedPattern(), b.exhaustedPattern());
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.config;
package dev.ltms.bridged.config;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -11,7 +11,7 @@ import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
/**
* Polls {@code fleetd.yaml}'s modified time and asks {@link ConfigRef} to reload when it moves
* Polls {@code bridged.yaml}'s modified time and asks {@link ConfigRef} to reload when it moves
* (CB-559). Opt-in through {@code configReload.enabled}.
*
* <p><strong>Why polling and not a filesystem watch.</strong> {@code WatchService} on macOS has no
@@ -1,4 +1,4 @@
package dev.ltms.fleet.guard;
package dev.ltms.bridged.guard;
/** Thrown when the subscription boundary would be violated. Never swallow this. */
public class GuardException extends RuntimeException {
@@ -1,4 +1,4 @@
package dev.ltms.fleet.guard;
package dev.ltms.bridged.guard;
import java.net.URI;
import java.net.URISyntaxException;
@@ -16,7 +16,7 @@ import java.util.Set;
* means its traffic would leave the subscription. That is a hard stop.</li>
* </ul>
*
* Both checks throw {@link GuardException} on violation. {@code fleetd} calls
* Both checks throw {@link GuardException} on violation. {@code bridged} calls
* {@link #assertWorker} before spawning a worker and {@link #assertPrimaryClean}
* against its own environment at startup.
*/
@@ -1,7 +1,7 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.session.MemberSession;
/**
* Pure classifier. Collection and repair are deliberately outside this package.
@@ -0,0 +1,151 @@
package dev.ltms.bridged.health;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.session.MemberSession;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.HashMap;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.function.BiConsumer;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/** Slow whole-fleet evidence collection. It is deliberately separate from the delivery poller. */
public final class FleetHealthMonitor {
private static final Logger log = LoggerFactory.getLogger(FleetHealthMonitor.class);
/** Bounded attempts to run {@link #failTarget} for one transition. Never retried tick-to-tick (CB-580). */
static final int MAX_FAIL_TARGET_ATTEMPTS = 3;
private final AgentControl agents;
private final Supplier<List<MemberSession>> roster;
private final MessageService messages;
private final ScheduledExecutorService scheduler;
private final LongSupplier clock;
private final long intervalSeconds;
private final BiConsumer<String, String> failTarget;
private final Map<String, HealthPrior> priors = new HashMap<>();
private final Map<String, HealthState> states = new HashMap<>();
// These facts need the evidence publishers introduced by later M4 units. They are not negatives.
private static final boolean NOT_YET_OBSERVED = false;
/**
* @param failTarget CB-568's idempotent target-wide failure operation (e.g. {@code messages::abandon}),
* invoked once when a member transitions into a terminal health state. Required —
* there is deliberately no defaulting overload; a caller that does not want the
* fail-tickets-on-terminal-health behavior must pass an explicit inert value (see
* {@code TestTurnTokens.inert} / {@code BridgeMcp.CapacitySource.none()} for the pattern).
*/
public FleetHealthMonitor(AgentControl agents, Supplier<List<MemberSession>> roster, MessageService messages,
ScheduledExecutorService scheduler, LongSupplier clock, long intervalSeconds,
BiConsumer<String, String> failTarget) {
this.agents = agents;
this.roster = roster;
this.messages = messages;
this.scheduler = scheduler;
this.clock = clock;
this.intervalSeconds = intervalSeconds;
this.failTarget = Objects.requireNonNull(failTarget, "failTarget");
}
/** Pure per-member decision seam. */
static HealthDecision decide(HealthSnapshot snapshot, HealthPrior prior, long nowNanos) {
return FleetHealth.decide(snapshot, prior, nowNanos);
}
public void start() { scheduler.schedule(this::tick, intervalSeconds, TimeUnit.SECONDS); }
public void stop() { scheduler.shutdownNow(); }
// Package-private so tests can run one tick without waiting.
void tick() {
try {
List<Agent> agentsNow = agents.list(); // Exactly one list call for this complete observation.
List<MemberSession> rosterNow = roster.get(); // One in-memory roster snapshot for this tick.
Map<String, Agent> live = new HashMap<>();
for (Agent agent : agentsNow) live.put(agent.terminalId(), agent);
HashSet<String> current = new HashSet<>();
for (MemberSession session : rosterNow) {
current.add(session.terminalId());
Agent agent = live.get(session.terminalId());
AgentStatus status = agent == null ? AgentStatus.UNKNOWN : agent.status();
boolean accepted = messages.hasAcceptedDelivery(session.terminalId());
HealthSnapshot snapshot = new HealthSnapshot(session.state(), status, accepted, NOT_YET_OBSERVED,
messages.hasInboxMessage(session.terminalId()), agent != null, NOT_YET_OBSERVED,
NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED);
HealthDecision decision = decide(snapshot, priors.getOrDefault(session.terminalId(), HealthPrior.NONE),
clock.getAsLong());
priors.put(session.terminalId(), decision.prior());
reportTransition(session.terminalId(), decision.state());
}
priors.keySet().retainAll(current);
states.keySet().retainAll(current);
} catch (Throwable error) {
// A list failure is health evidence, and must never kill the monitor's only scheduler task.
log.warn("fleet health collection failed; will retry next tick", error);
} finally {
if (!scheduler.isShutdown()) {
scheduler.schedule(this::tick, intervalSeconds, TimeUnit.SECONDS);
}
}
}
void reportTransition(String target, HealthState next) {
HealthState previous = states.put(target, next);
if (previous == next) return;
if (fault(next)) {
log.warn("fleet health member={} state={} previous={}", target, next, previous);
} else if (previous != null && fault(previous)) {
log.info("fleet health member={} recovered state={} previous={}", target, next, previous);
}
// CB-580: a member entering GONE/NEVER_READY must not leave its waiting tickets pending
// forever. Fire exactly once per transition — never on a tick where the state is unchanged,
// which is what made the rejected commit call abandon() once per tick for as long as a
// member stayed terminal.
if (terminal(next)) {
failTerminalTarget(target, next);
}
}
private void failTerminalTarget(String target, HealthState state) {
String reason = "fleet health: member reached terminal state " + state.name();
RuntimeException last = null;
for (int attempt = 1; attempt <= MAX_FAIL_TARGET_ATTEMPTS; attempt++) {
try {
failTarget.accept(target, reason);
return;
} catch (RuntimeException error) {
last = error;
log.warn("fleet health: failTarget attempt {}/{} failed for member={} state={}",
attempt, MAX_FAIL_TARGET_ATTEMPTS, target, state, error);
}
}
log.warn("fleet health: giving up on failTarget for member={} state={} after {} attempts",
target, state, MAX_FAIL_TARGET_ATTEMPTS, last);
}
private static boolean terminal(HealthState state) {
return state == HealthState.GONE || state == HealthState.NEVER_READY;
}
private static boolean fault(HealthState state) {
return switch (state) {
case NEVER_READY, GONE, TURN_BOUNDARY_LOST, ERROR_ON_SCREEN, STALL_SUSPECTED,
MUTE, REPLY_STRANDED, DELEGATION_ORPHANED, CONTROL_LINK_DOWN -> true;
default -> false;
};
}
public static String coverage(boolean enabled, boolean notificationConfigured) {
return !enabled ? "off" : notificationConfigured ? "full" : "detection-only";
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
/** Classification plus the private fact that the next pure decision needs. */
public record HealthDecision(HealthState state, HealthPrior prior) { }
@@ -1,4 +1,4 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
/** Private cross-tick observation. It is deliberately not a reported health value. */
public record HealthPrior(boolean busyButDone) {
@@ -1,7 +1,7 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.session.MemberSession;
/** Read-only facts from one fleet collection tick. */
public record HealthSnapshot(MemberSession.State sessionState, AgentStatus liveStatus,
@@ -1,4 +1,4 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
/** Health classifications reported for a member. */
public enum HealthState {
@@ -1,12 +1,12 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.bridged.msg.Rendezvous;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
/**
* Counts turns that ended via the completion fallback instead of {@code fleet_reply}.
* Counts turns that ended via the completion fallback instead of {@code bridge_reply}.
* MUTE is an observation by target and profile, not a classifier state and never suppresses faults.
*/
public final class MuteCounter {
@@ -1,4 +1,4 @@
package dev.ltms.fleet.health;
package dev.ltms.bridged.health;
import java.util.ArrayList;
import java.util.HashMap;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
@@ -38,11 +38,6 @@ public final class AgentControl {
this.herdr = herdr;
}
/** The herdr daemon this control object sends its agent calls to. */
public HerdrClient herdr() {
return herdr;
}
/** One agent-targeted call, translating a terminal id to its pane id (retrying once fresh). */
private JsonNode agentCall(String method, String target, Map<String, Object> extra) {
String resolved = resolveTarget(target);
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
/**
* A herdr agent's lifecycle state, as reported by {@code agent_status}. Drives the
@@ -32,7 +32,7 @@ public enum AgentStatus {
};
}
/** Whether {@code fleetd} may inject a message now without stepping on a live turn. */
/** Whether {@code bridged} may inject a message now without stepping on a live turn. */
public boolean injectable() {
return this == IDLE || this == BLOCKED || this == DONE;
}
@@ -1,17 +1,17 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
/**
* Client face onto the herdr daemon (protocol 19, herdr 0.8.0).
* Client face onto the herdr daemon (protocol 14, herdr 0.7.0).
*
* <p>This is the ONLY thing in {@code fleetd} that speaks to herdr. Every method
* <p>This is the ONLY thing in {@code bridged} that speaks to herdr. Every method
* maps to a herdr JSON-RPC call over its Unix domain socket. Requests are
* newline-delimited JSON with a <em>string</em> id; responses carry either a
* {@code result} object (whose {@code type} field discriminates the payload) or an
* {@code error} object.
*
* <p>Higher layers ({@code fleetd}'s policy brain, REST endpoints, MCP adapters)
* <p>Higher layers ({@code bridged}'s policy brain, REST endpoints, MCP adapters)
* depend on this interface, not on the socket. Tests substitute a fake; the
* {@code contract}-tagged suite exercises the real implementation against a running
* herdr to catch protocol drift.
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.JsonNode;
@@ -8,7 +8,7 @@ import com.fasterxml.jackson.databind.node.ObjectNode;
import java.nio.charset.StandardCharsets;
/**
* Wire codec for herdr's newline-delimited JSON-RPC (protocol 19).
* Wire codec for herdr's newline-delimited JSON-RPC (protocol 14).
*
* <p>Split out from the socket so the framing rules — the ones that actually bit us
* during the spike (id MUST be a string; response carries {@code result} or
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
/**
* Raised when a herdr call fails: transport error, or an {@code error} envelope
@@ -1,11 +1,10 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.Locale;
import java.util.Map;
@@ -15,7 +14,7 @@ import java.util.function.Supplier;
/**
* Discovers which panes host a lead by scanning herdr for tabs the operator labelled by convention
* (CB-531), and hands {@link dev.ltms.fleet.auth.CallerResolver} the resulting
* (CB-531), and hands {@link dev.ltms.bridged.auth.CallerResolver} the resulting
* {@code terminal_id → lead name} map.
*
* <p><strong>Why scan at all.</strong> A lead is never spawned — a human opens a tab and starts an
@@ -40,54 +39,31 @@ import java.util.function.Supplier;
* <li>Worker spaces are excluded wholesale ({@code excludedWorkspaceLabels}), so a worker cannot
* become a lead by being placed — as a split, say — inside a matching tab.</li>
* <li>A worker cannot rename a tab: {@code tab.rename} is reachable only through
* {@link WorkspaceControl}, which no {@code fleet_*} tool exposes. The label is writable by
* {@link WorkspaceControl}, which no {@code bridge_*} tool exposes. The label is writable by
* the human at the terminal and by nobody the bridge is defending against.</li>
* <li>The label is a <em>name</em>, not a capability. What a pane may do is decided by
* {@code Authz} against the role {@code CallerResolver} returns; a tab that calls itself a
* lead still cannot act as one unless the daemon's own registry agrees.</li>
* </ol>
*
* <p><strong>CB-558 — fleetd now writes lead labels too.</strong> This class used to be able to say
* that fleetd never renames a lead tab, so the label was always the human's own writing and there
* <p><strong>CB-558 — bridged now writes lead labels too.</strong> This class used to be able to say
* that bridged never renames a lead tab, so the label was always the human's own writing and there
* was no round-trip from the daemon's rename back into its next decision.
* {@code dev.ltms.fleet.lead.LeadLauncher} ends that: an auto-launched lead is labelled by the
* daemon and found again by this scan. The trust direction above is unaffected — fleetd writing a
* {@code dev.ltms.bridged.lead.LeadLauncher} ends that: an auto-launched lead is labelled by the
* daemon and found again by this scan. The trust direction above is unaffected — bridged writing a
* name for a lead it just started is not a pane promoting itself — but <em>staleness</em> becomes
* real: a label left behind by a session that has since died would read as a live lead forever.
* This scanner does not solve that (its job is naming, and a stale name costs nothing here); the
* launcher does, by requiring a running agent in the tab before it counts the lead as live. If you
* ever make a decision that <em>removes</em> something based on this map, add the same check.
* The remaining hazard is an <em>operator</em> one — a worker {@code tabLabel} template that
* happens to start with the same prefix would promote the whole fleet — and that is refused at
* startup by {@code FleetConfig.validateLeadTabPrefixes} rather than documented here.
*
* <p><strong>fleetd #359 — the staleness check this class used to skip.</strong> This used to say
* "a stale name costs nothing here" and leave liveness to {@code LeadLauncher}, on the theory that
* naming and removing are different decisions. That was wrong: {@code dev.ltms.fleet.msg.LeadCoordLoop}
* makes exactly the kind of removal decision the old javadoc warned about, by reading this map to
* pick which pane a peer message goes into — and a stale entry there is not free. On a host where
* {@code fleetd} had restarted more than once, a labelled-but-dead tab from a previous life was
* reported right alongside the live one; {@code LeadCoordLoop.resolveLocalLead()} saw more than one
* candidate and refused to guess (safe), but the fix the daemon's own WARN suggests — name a lead
* after {@code coordinator.selfId} — stops being safe once two tabs can share a label: step 1 of
* that resolution picks whichever matching entry it finds first, which can be the dead one, and
* typing a peer's message into a dead shell does not fail — it is silently gone instead of merely
* held. {@link #scan()} now cross-checks every labelled tab against {@code agent.list} (the same
* signal {@code LeadLauncher.countLeads} already trusts for the same purpose) and drops any tab
* with no agent running in it, so a dead tab is never in the map for a caller to pick at all.
* startup by {@code BridgedConfig.validateLeadTabPrefixes} rather than documented here.
*
* <p><strong>Caching.</strong> {@link #get()} is on the request path (every resolve), so the scan
* is TTL-cached and a stale-but-valid map is preferred to a herdr round-trip. A failed scan (herdr
* throws) keeps the previous answer instead of emptying it — a herdr hiccup must not silently
* demote a live lead mid-session.
*
* <p><strong>fleetd #359 review, finding 2 — a successful-but-wrong scan is the same hazard.</strong>
* The catch above only fires when a call throws. It does nothing for a call that returns 200 with an
* incomplete answer — exactly what the ticket's own evidence showed {@code agent.list} can do. Once
* this class started trusting that signal, an empty read would otherwise get cached as fact and
* silently drop a lead {@code CallerResolver} had, until then, correctly resolved — turning it into a
* {@code Role.WORKER}, which refuses every orchestration call. So a terminal this class already
* reported as live is not dropped the first time {@code agent.list} loses it: {@link #scan()} grants
* it one grace scan (see {@code gracedTerminals}) and only drops it if a <em>later</em> scan still
* finds no agent. A terminal never reported live before gets no grace — that would weaken the
* original #359 fix itself, which this class's own test suite already pins.
* is TTL-cached and a stale-but-valid map is preferred to a herdr round-trip. A failed scan keeps
* the previous answer instead of emptying it — a herdr hiccup must not silently demote a live lead
* mid-session.
*/
public final class LeadTabScanner implements Supplier<Map<String, String>> {
@@ -103,15 +79,6 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
private long scannedAtNanos;
private boolean everScanned;
/**
* Terminals currently on their one grace scan: {@code cached} reported them live, the most
* recent {@link #scan()} found no agent for them, and they were re-included anyway. Cleared for
* a terminal the instant it is seen live again; a terminal still here on the <em>next</em> scan
* is finally dropped. Scoped separately from {@link #cached} so a graced terminal cannot renew
* its own grace forever just by staying in the exposed map (fleetd #359 review, finding 2).
*/
private Set<String> gracedTerminals = Set.of();
/**
* @param herdr the herdr client to query ({@code workspace.list},
* {@code tab.list}, {@code pane.list} — all read-only)
@@ -174,7 +141,7 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
return cached;
}
/** One full pass: labelled tabs → live agents in them → those panes' terminals. */
/** One full pass: labelled tabs → their panes → those panes' terminals. */
private Map<String, String> scan() {
Map<String, String> nameByTab = new LinkedHashMap<>();
for (JsonNode w : herdr.call("workspace.list").path("workspaces")) {
@@ -191,47 +158,17 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
}
}
if (nameByTab.isEmpty()) {
gracedTerminals = Set.of();
return Map.of();
}
// fleetd #359: a labelled tab is only a lead when herdr also reports a running agent in
// it — the same liveness signal LeadLauncher.countLeads trusts for the identical purpose.
// Without this, a tab left behind by a session that has since died reads as live forever.
Set<String> tabsWithAgent = new HashSet<>();
for (JsonNode a : herdr.call("agent.list").path("agents")) {
String tabId = a.path("tab_id").asText(null);
if (tabId != null) {
tabsWithAgent.add(tabId);
}
}
Map<String, String> byTerminal = new LinkedHashMap<>();
Set<String> stillGraced = new HashSet<>();
// One pane.list for every tab: panes carry tab_id, so the join is local.
for (JsonNode p : herdr.call("pane.list", Map.of()).path("panes")) {
String tabId = p.path("tab_id").asText(null);
String name = nameByTab.get(tabId);
String terminal = p.path("terminal_id").asText(null);
if (name == null || terminal == null || terminal.isBlank()) {
continue;
}
if (tabsWithAgent.contains(tabId)) {
byTerminal.put(terminal, name);
continue;
}
// No agent reported for this tab, but its tab/pane are still here — this is the
// ambiguous case review finding 2 named: a successful agent.list that came back short
// does not prove the lead is dead. Grant one grace scan to a terminal we had already
// reported as live; a terminal we never reported live gets none, so the original #359
// fix (a genuinely dead tab is never reported) is unaffected for the common case.
if (cached.containsKey(terminal) && !gracedTerminals.contains(terminal)) {
byTerminal.put(terminal, name);
stillGraced.add(terminal);
if (!nameByTab.isEmpty()) {
// One pane.list for every tab: panes carry tab_id, so the join is local.
for (JsonNode p : herdr.call("pane.list", Map.of()).path("panes")) {
String name = nameByTab.get(p.path("tab_id").asText(null));
String terminal = p.path("terminal_id").asText(null);
if (name != null && terminal != null && !terminal.isBlank()) {
byTerminal.put(terminal, name);
}
}
}
gracedTerminals = stillGraced;
return Collections.unmodifiableMap(byTerminal);
}
@@ -240,14 +177,12 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
*
* <p>Exact match (case-insensitive, ends stripped) against {@link #tabToName} — no prefix
* stripping, so an operator's {@code "lead: something-else"} tab is never mistaken for a
* configured lead just because it shares a prefix. The match strips a trailing
* {@link PendingCloseMarker} first, so a tab {@code LeadLauncher} has flagged as maybe-dead but
* not yet closed keeps resolving normally while that reconcile is pending.
* configured lead just because it shares a prefix.
*/
private String leadNameOf(String label) {
if (label == null) {
return null;
}
return tabToName.get(PendingCloseMarker.strip(label).toLowerCase(Locale.ROOT));
return tabToName.get(label.strip().toLowerCase(Locale.ROOT));
}
}
@@ -0,0 +1,59 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import java.util.Map;
/**
* Resolves which herdr pane a process belongs to — the herdr half of connection-based MCP
* identity (CB-105). Given the PID that opened an MCP connection, {@link #terminalForPid} finds
* the agent pane whose process tree contains it, so {@code bridged} can tell <em>which worker</em>
* is calling without the worker sending anything spoofable.
*
* <p>herdr owns the PID→pane truth: {@code pane.process_info} reports each pane's {@code shell_pid}
* and foreground process PIDs. This scans agent panes; a spawn-time {@code pid→terminal} cache is
* the obvious optimization once wired into {@code ClaudeCodeLauncher}.
*/
public final class PaneLocator {
private final HerdrClient herdr;
public PaneLocator(HerdrClient herdr) {
this.herdr = herdr;
}
/**
* The {@code terminal_id} of the agent pane whose process tree contains {@code pid}, or
* {@code null} if no agent pane owns it (e.g. the caller is the primary, or off-host).
*/
public String terminalForPid(long pid) {
if (pid <= 0) {
return null;
}
for (JsonNode pane : herdr.call("pane.list", Map.of()).path("panes")) {
String paneId = pane.path("pane_id").asText(null);
if (paneId != null && paneOwnsPid(paneId, pid)) {
return pane.path("terminal_id").asText(null);
}
}
return null;
}
private boolean paneOwnsPid(String paneId, long pid) {
JsonNode info;
try {
info = herdr.call("pane.process_info", Map.of("pane_id", paneId)).path("process_info");
} catch (HerdrException e) {
return false; // pane vanished mid-scan — just skip it
}
if (info.path("shell_pid").asLong(-1) == pid) {
return true;
}
for (JsonNode p : info.path("foreground_processes")) {
if (p.path("pid").asLong(-1) == pid) {
return true;
}
}
return false;
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
@@ -7,7 +7,7 @@ import com.fasterxml.jackson.databind.JsonNode;
* dedicated worker space so they never split or clutter the user's real work spaces.
*
* @param workspaceId herdr's stable id (e.g. {@code "w4"})
* @param label display label shown in herdr's UI (e.g. {@code "fleetd-workers"})
* @param label display label shown in herdr's UI (e.g. {@code "bridged-workers"})
* @param activeTabId the workspace's currently-focused tab, or {@code null}
*/
public record Workspace(String workspaceId, String label, String activeTabId) {
@@ -1,4 +1,4 @@
package dev.ltms.fleet.herdr;
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.slf4j.Logger;
@@ -0,0 +1,372 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.TurnToken;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.Objects;
import java.util.Set;
import java.util.TreeSet;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.regex.Pattern;
/**
* The CB-106 completion fallback: bridges the {@link Injector}'s turn-completion signal to the
* {@link Rendezvous} so a blocking {@code bridge_send} resolves even when the worker finishes its
* task without ever calling {@code bridge_reply} — the common case for a real delegated coding task.
*
* <p>On a confirmed {@code working → idle} boundary it scrapes the worker's recent transcript and
* resolves the awaiting send with that tail (a {@link Rendezvous.Kind#COMPLETION} resolution, so the
* caller can tell a scrape from a structured reply). It scrapes only when a send is actually waiting
* — a fleet worker's own turns, or a send that already timed out, cost no herdr traffic. An explicit
* {@code bridge_reply} that raced in first wins; {@link Rendezvous#resolveCompletion} is then a no-op.
*
* <p>It also handles the CB-109 stall signal ({@link #onTurnFailed}): a worker that ran a turn then
* wedged in an {@code unknown} state resolves the send as a failure (with the error screen as
* context) rather than leaving it to time out.
*
* <p>The scrape is cleaned to the last {@code ⏺} assistant block (stripping TUI chrome) and guarded
* against misattribution (CB-115): the pane content is baselined on delivery ({@link #onDelivered}),
* and a completion whose scrape is unchanged from that baseline — the previous turn's wind-down
* sampled as this turn's boundary on a rapid back-to-back send — is suppressed rather than resolving
* the send with a stale answer.
*
* <p><strong>Waiter-specific resolution (CB-116).</strong> On delivery we also capture the exact
* {@link Rendezvous} waiter this turn belongs to, and the completion/failure fallbacks resolve
* <em>that</em> waiter — never "whatever send is waiting now". A completion fallback runs on a virtual
* thread and can land after the worker's {@code bridge_reply} already resolved the turn and the
* <em>next</em> send opened its own waiter on the same session; resolving the current waiter would
* then deliver turn N's stale scrape as turn N+1's answer. Targeting the captured waiter makes a late
* completion a harmless no-op (its waiter is already done) instead of a cross-turn stale reply.
*
* <p>Wired as the {@link Injector}'s {@link TurnListener}; the handlers hand off to a virtual thread
* so the scrape's herdr round-trip never stalls the status poller. The captured waiter is read on the
* poller thread (before any next-turn delivery can overwrite it) and passed into the virtual thread.
*/
public final class CompletionResolver implements TurnListener {
private static final Logger log = LoggerFactory.getLogger(CompletionResolver.class);
/**
* herdr {@code agent.read} source for the completion scrape. {@code recent} returns the tail of
* the transcript (the worker's last output), which is what a delegator wants when the worker
* didn't structure a reply.
*/
static final String SCRAPE_SOURCE = "recent";
/** Cap the scraped tail so a long transcript can't return an unbounded blob. */
static final int MAX_SCRAPE_CHARS = 4000;
private static final String CLIPPED_PANE_TAIL_MARKER =
"[Pane tail clipped: member did not call bridge_reply.]";
private final AgentControl agents;
private final Rendezvous rendezvous;
private final ExhaustedPatternLookup exhaustedPatterns;
private final ExhaustionSink exhaustionSink;
/**
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
* delivering send opened, plus the assistant block present when it was delivered.
*
* <p>The {@code waiter} is what makes a late fallback safe (CB-116): we resolve it, not "whoever
* is waiting now", so a completion that fires after the next send has opened its own waiter is a
* no-op rather than a cross-turn stale reply. The {@code baseline} is the CB-115 staleness
* reference: a completion scrape equal to it means the worker produced no new output (the previous
* turn's wind-down sampled as this boundary), so it is suppressed. Overwritten on each delivery;
* cleared when the turn resolves. Package-private so tests can capture and replay a specific turn.
*/
record InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline) {
}
private final ConcurrentHashMap<String, InFlight> inFlight = new ConcurrentHashMap<>();
/**
* @param exhaustedPatterns CB-578 stage A: per-target lookup for a profile's configured
* usage-limit refusal pattern. Required — there is deliberately no
* defaulting overload; a caller that does not want the classification
* must pass an explicit inert value ({@link ExhaustedPatternLookup#none()}).
* @param exhaustionSink CB-578 stage B: notified when a {@code BACKEND_EXHAUSTED}
* classification actually resolves a waiter. Required for the same
* reason as {@code exhaustedPatterns} — pass {@link ExhaustionSink#none()}
* to opt out.
*/
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
ExhaustionSink exhaustionSink) {
this.agents = agents;
this.rendezvous = rendezvous;
this.exhaustedPatterns = Objects.requireNonNull(exhaustedPatterns, "exhaustedPatterns");
this.exhaustionSink = Objects.requireNonNull(exhaustionSink, "exhaustionSink");
}
@Override
public void onDelivered(String target, TurnToken token) {
// Capture the exact waiter this turn belongs to (CB-116) and snapshot the pane's pre-turn
// content — what it shows *before* the just-delivered turn produces output — as the staleness
// reference (CB-115). Done synchronously (like the delivering send itself) so both are in
// place before this turn's completion can fire.
captureBaseline(target, token);
}
/** Capture the in-flight turn: its waiter and pre-turn baseline (the testable core of {@link #onDelivered}). */
void captureBaseline(String target, TurnToken token) {
CompletableFuture<Rendezvous.Resolution> waiter = token.waiter();
if (waiter == null) {
inFlight.remove(target); // no send is waiting on this delivery — nothing to resolve later
return;
}
String baseline;
try {
// Clip to the same cap resolve() applies to the tail (line ~134): the CB-115 misattribution
// guard compares baseline.equals(tail), so both sides must be the same capped representation.
// An unclipped baseline vs a clipped tail would never match for a >MAX_SCRAPE_CHARS block,
// defeating the guard and letting a stale completion resolve the send.
baseline = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
} catch (RuntimeException e) {
baseline = null; // fail open: no baseline ⇒ no suppression
log.debug("delivery baseline for {} failed: {}", target, e.getMessage());
}
inFlight.put(target, new InFlight(waiter, baseline));
}
/** The turn currently baselined for {@code target}, or {@code null} — a test hook for the captureBaseline path. */
InFlight inFlight(String target) {
return inFlight.get(target);
}
@Override
public void onTurnComplete(String target) {
// Read the in-flight turn on the poller thread — before any next-turn delivery can overwrite
// it — then off-load the scrape (a herdr round-trip we must not block polling on) to a vthread.
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("completion-" + target).start(() -> resolve(target, turn));
}
/**
* Resolve the completed turn before adapter housekeeping can erase its rendered output. This is
* intentionally synchronous and used only when a post-turn context reset is enabled; the normal
* path remains off-loaded so polling is not blocked by a scrape.
*/
public void resolveBeforePostAction(String target) {
resolve(target, inFlight.get(target));
}
@Override
public void onTurnFailed(String target) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn, null));
}
@Override
public void onTurnFailed(String target, String reason) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn, reason));
}
/** Synchronous resolve (the unit-testable core of {@link #onTurnComplete}). */
void resolve(String target, InFlight turn) {
CompletableFuture<Rendezvous.Resolution> waiter = turn == null ? null : turn.waiter();
if (waiter == null || waiter.isDone()) {
// Nobody is blocked on THIS turn (it had no send, or its bridge_reply already won). Skip
// the scrape; resolving the current waiter here would be the CB-116 cross-turn stale reply.
inFlight.remove(target, turn);
return;
}
String tail;
String assistantBlock = null;
int originalLength = 0;
boolean clipped = false;
boolean scrapeFailed = false;
try {
assistantBlock = lastAssistantBlock(agents.read(target, SCRAPE_SOURCE));
originalLength = assistantBlock.strip().length();
clipped = originalLength > MAX_SCRAPE_CHARS;
tail = clip(assistantBlock);
} catch (RuntimeException e) {
// The worker finished but we couldn't read its screen — still resolve the send so the
// caller unblocks; an empty tail beats hanging until the caller's timeout.
log.warn("completion scrape for {} failed; resolving with an empty tail: {}",
target, e.getMessage());
tail = "";
scrapeFailed = true;
}
// Misattribution guard (CB-115): if the scrape is byte-identical to the pane content at
// delivery, this turn produced no new output — the boundary belongs to the previous turn's
// wind-down (common on rapid back-to-back sends). Suppress rather than resolve the send with
// a stale answer; the real bridge_reply (or a later genuine completion) resolves it instead.
// A scrape that failed to read is exempt — an empty tail there is "couldn't see", not "no change".
String baseline = turn.baseline();
if (!scrapeFailed && baseline != null && baseline.equals(tail)) {
log.debug("suppressing misattributed completion for {} (no output change since delivery)",
target);
return; // keep the in-flight record: a later genuine completion still needs it
}
// CB-578 stage A: a turn that ended with no bridge_reply AND whose scrape matches the
// backend's configured usage-limit pattern is a refusal, not an answer. Classify it as
// BACKEND_EXHAUSTED rather than handing the caller a scrape that reads like a real reply.
if (!scrapeFailed) {
Pattern exhausted = exhaustedPatterns.patternFor(target);
String matchedLine = exhausted == null ? null : firstMatchingLine(assistantBlock, exhausted);
if (matchedLine != null) {
String reason = "backend exhausted (usage limit): " + matchedLine;
if (rendezvous.resolveExhausted(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("completion for {} classified BACKEND_EXHAUSTED (no bridge_reply; scrape "
+ "matched the profile's exhausted pattern): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
}
return;
}
}
String completion = clipped ? tail + "\n" + CLIPPED_PANE_TAIL_MARKER : tail;
if (rendezvous.resolveCompletion(waiter, completion)) {
inFlight.remove(target, turn);
if (clipped) {
log.warn("completion scrape for {} clipped from {} chars to the {} char cap; "
+ "member did not call bridge_reply, so the pane tail is partial",
target, originalLength, MAX_SCRAPE_CHARS);
}
log.debug("resolved send to {} via turn-completion fallback ({} chars scraped)",
target, tail.length());
}
}
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
void fail(String target, InFlight turn) {
fail(target, turn, null);
}
/** Synchronous fail with an optional reason supplied by a dropped worker queue. */
void fail(String target, InFlight turn, String explicitReason) {
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
// fall back to the currently-registered waiter (unambiguous — that send never completed, so
// no next turn exists to confuse it with).
CompletableFuture<Rendezvous.Resolution> waiter =
turn != null ? turn.waiter() : rendezvous.currentWaiter(target);
if (waiter == null || waiter.isDone()) {
inFlight.remove(target, turn); // nobody blocked on this worker — nothing to fail
return;
}
String reason = explicitReason;
if (reason == null || reason.isBlank()) {
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
}
}
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("failing send to {} via turn-stall fallback: {}", target, reason);
}
}
/**
* The first line of {@code text} matching {@code pattern}, stripped — the CB-578 stage A
* evidence carried in a {@code BACKEND_EXHAUSTED} reason so the operator sees the real refusal
* text, never a generic label. {@code null} if no line matches.
*/
static String firstMatchingLine(String text, Pattern pattern) {
if (text == null || text.isEmpty()) return null;
for (String line : text.split("\n", -1)) {
if (pattern.matcher(line).find()) {
return line.strip();
}
}
return null;
}
/**
* Coverage summary for the CB-578 stage A exhausted-pattern classification, logged at startup
* the way {@link dev.ltms.bridged.health.FleetHealthMonitor#coverage} is — so an operator can
* see whether the classification is on, and for which profiles, without reading every
* profile's config by hand.
*
* @param allProfiles every configured profile name
* @param configuredProfiles the subset of {@code allProfiles} that carry an exhausted pattern
*/
public static String coverage(Set<String> allProfiles, Set<String> configuredProfiles) {
if (configuredProfiles.isEmpty()) {
return "off (no profile has an exhaustedPattern configured; profiles: " + sorted(allProfiles) + ")";
}
Set<String> unconfigured = new TreeSet<>(allProfiles);
unconfigured.removeAll(configuredProfiles);
return unconfigured.isEmpty()
? "full (all profiles configured: " + sorted(allProfiles) + ")"
: "partial (configured: " + sorted(configuredProfiles) + "; not configured: " + sorted(unconfigured) + ")";
}
private static List<String> sorted(Set<String> names) {
return names.stream().sorted().toList();
}
private static String clip(String s) {
if (s == null) return "";
String trimmed = s.strip();
return trimmed.length() <= MAX_SCRAPE_CHARS
? trimmed
: trimmed.substring(trimmed.length() - MAX_SCRAPE_CHARS);
}
/**
* Extract the last assistant message from a raw Claude Code pane scrape (CB-115). Claude Code
* prefixes each assistant turn with {@code ⏺}; the delegator wants that answer, not the TUI
* chrome around it. Take everything from the final {@code ⏺} onward and stop at the <em>first</em>
* hard interface boundary below it — the spinner/status line, input box, {@code ❯} prompt (which
* may echo the <em>next</em> turn's text), footer, or tips/warnings. Stopping at the first
* boundary (rather than trimming only trailing chrome) is what keeps a following turn's echoed
* prompt out of this reply. Blank lines are not boundaries, so a multi-paragraph answer survives;
* trailing blanks are trimmed at the end. With no {@code ⏺} marker (an unusual render) the whole
* text is scanned the same way, so we never lose the reply.
*
* <p>Package-private and pure so it is unit-testable without herdr.
*/
static String lastAssistantBlock(String raw) {
if (raw == null || raw.isBlank()) return "";
int marker = raw.lastIndexOf('⏺');
String block = marker >= 0 ? raw.substring(marker + 1) : raw;
StringBuilder out = new StringBuilder();
int kept = 0;
for (String line : block.split("\n", -1)) {
if (isBoundary(line)) break; // first TUI boundary ends the assistant message
if (kept++ > 0) out.append('\n');
out.append(line);
}
return out.toString().strip();
}
/**
* A hard TUI boundary line that marks the end of an assistant message and the start of interface
* chrome (input box, prompt, spinner, footer, tips/warnings). Blank lines are <em>not</em>
* boundaries — an answer may contain them — so they are kept and trimmed only if trailing.
*/
private static boolean isBoundary(String line) {
String t = line.strip();
if (t.isEmpty()) return false;
// A horizontal rule / all box-drawing separators (e.g. "──────").
if (t.chars().allMatch(c -> c == '─' || c == '—' || c == '━' || c == '═' || c == '-')) {
return true;
}
String lower = t.toLowerCase();
return t.startsWith("╭") || t.startsWith("│") || t.startsWith("╰") || t.startsWith("┌")
|| t.startsWith("└") || t.startsWith("❯") || t.startsWith("⏵")
|| t.startsWith("⎿") || t.startsWith("⚠")
// Status/spinner lines Claude Code renders below a settled or in-flight turn,
// e.g. "✻ Baked for 21s", "✶ Forming…".
|| t.startsWith("✻") || t.startsWith("✳") || t.startsWith("✽") || t.startsWith("·")
|| t.startsWith("●") || t.startsWith("◐") || t.startsWith("✢") || t.startsWith("✶")
|| lower.contains("auto mode") || lower.contains("for shortcuts")
|| lower.contains("esc to interrupt") || lower.contains("bypass permissions");
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import java.util.regex.Pattern;
@@ -0,0 +1,31 @@
package dev.ltms.bridged.inject;
/**
* Notified when {@link CompletionResolver} actually delivers a {@code BACKEND_EXHAUSTED}
* classification to a waiting send (CB-578 stage B) — never on a race that lost (see
* {@link CompletionResolver#resolve}, which only calls this after
* {@code Rendezvous.resolveExhausted} returns {@code true}).
*
* <p>{@link CompletionResolver} knows only {@code target} (a herdr terminal id); it has no notion of
* profiles or credentials, so mapping {@code target} to whatever should be quarantined is entirely
* the sink's job — see {@code Bridged.main}'s wiring, which resolves target → session → profile →
* {@code effectiveCredentialId()} and calls {@code BackendQuarantine.quarantine} on it.
*/
@FunctionalInterface
public interface ExhaustionSink {
/**
* @param target the herdr terminal id whose turn was classified {@code BACKEND_EXHAUSTED}
* @param reason the matched-line reason carried by the classification
*/
void onExhausted(String target, String reason);
/**
* Inert sink — nothing happens on exhaustion. The explicit stand-in a caller (or a test not
* exercising this feature) passes instead of a defaulting overload, exactly like
* {@link ExhaustedPatternLookup#none()}.
*/
static ExhaustionSink none() {
return (target, reason) -> { };
}
}
@@ -0,0 +1,418 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.msg.TurnToken;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayDeque;
import java.util.ArrayList;
import java.util.Deque;
import java.util.List;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Consumer;
import java.util.function.Predicate;
import java.util.stream.Collectors;
/**
* The status-gated injector (CB-103): the single writer that delivers a message into a
* worker only when it is safe — {@code idle} or {@code blocked}, never mid-turn.
*
* <p>Per target it holds a FIFO queue and delivers <strong>at most one message per turn</strong>:
* after a send it waits for the worker to pick the message up (go {@code working}) before
* delivering the next, so two rapid deliveries never interleave into one turn. Each
* status-check-then-send for a target is serialized on the target's monitor, closing the
* TOCTOU window between "is it idle?" and "send" — one writer per worker.
*
* <p><strong>Polling caveat.</strong> This is driven by {@link #onStatus} sampling (a
* {@code StatusPoller}), not by reliable status <em>edges</em>. A turn can begin and end
* entirely between two polls, so the {@code working} pickup may never be sampled. To avoid
* wedging a queue forever, an awaited pickup is released after {@link #PICKUP_GRACE_POLLS}
* consecutive injectable samples (the worker has plainly moved on). Only {@code working} — not
* a transient {@code unknown} — counts as a real pickup, so a detection glitch can't prematurely
* release the latch. Perfectly reliable turn boundaries require a herdr {@code events.subscribe}
* stream; that is the intended upgrade and would replace only the sampling, not this queue.
*
* <p><strong>Turn completion (CB-106).</strong> Beyond delivery, the injector reports when a
* delegated turn <em>finishes</em>: after a delivery is picked up (a real {@code working} sample),
* the next injectable sample is a confirmed {@code working → idle} boundary and fires
* {@link TurnListener#onTurnComplete}. Completion is only ever synthesized from a <em>confirmed</em>
* turn — the pickup-grace path (a turn too fast to sample) unwedges the queue but does not fire
* completion, since without a sampled {@code working} there is no trustworthy "the worker just
* finished the task" signal to act on.
*/
public final class Injector {
private static final Logger log = LoggerFactory.getLogger(Injector.class);
/**
* How many consecutive injectable samples (with no {@code working} in between) after a send
* before we assume the turn completed unobserved and release the pickup latch. At the
* default 250ms poll interval this is a ~2s grace — far longer than a worker takes to start
* a turn, so it only fires on a genuinely missed pickup edge.
*/
private static final int PICKUP_GRACE_POLLS = 8;
/**
* How many consecutive {@code unknown} samples while a delegation is outstanding before we
* declare it stalled and fire {@link TurnListener#onTurnFailed} (CB-109). A worker wedged in a
* state herdr can't classify (e.g. an API-error screen) stays {@code unknown} indefinitely and
* would otherwise never resolve; any {@code working}/{@code idle} sample resets the streak, so a
* transient detection glitch cannot trip it. At the 250ms poll interval this is ~30s — far longer
* than any real detection blip, and still vastly better than the async send's timeout.
*/
private static final int TURN_STALL_GRACE_POLLS = 120;
/**
* How many consecutive injectable samples a queued-but-undelivered message may wait on the
* {@link #ready} gate before we give up and fail it (CB-114). The gate holds a message out of a
* worker's boot window (herdr reports {@code idle} while its Claude is still starting), but a
* worker whose Claude crashes during boot — or never connects the bridge MCP — stays "idle and
* not ready" forever: {@link #ready} never accepts it, the message is never delivered, and the
* target would be polled indefinitely with its caller's future never completing. After this
* grace the queued messages are failed and the target released. At the 250ms poll interval this
* is ~60s — deliberately longer than {@link #TURN_STALL_GRACE_POLLS}, since a first boot (spawn
* + model load + MCP connect) legitimately takes longer than an in-turn detection blip.
*/
private static final int READINESS_GRACE_POLLS = 240;
/**
* The single source for the injector poll cadence — how often the {@link StatusPoller} drives
* {@link #onStatus} at. {@code Bridged} passes this to every {@link StatusPoller} it constructs,
* and this class reads it to state the readiness grace in seconds on the CB-562 expiry log
* instead of hardcoding "60s". One constant, so a cadence change cannot silently desync a log
* that claims a grace duration.
*/
public static final long POLL_INTERVAL_MILLIS = 250;
private final AgentControl agents;
private final TurnListener turnListener;
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
private final Consumer<String> forget; // CB-114: clear a gone worker's readiness/presence
private final ConcurrentHashMap<String, Target> targets = new ConcurrentHashMap<>();
/** Delivery only; completion signalling is a no-op and every target is treated as available. */
public Injector(AgentControl agents) {
this(agents, TurnListener.NOOP);
}
/** Delivery plus turn-completion signalling (CB-106); every target is treated as available. */
public Injector(AgentControl agents, TurnListener turnListener) {
this(agents, turnListener, _ -> true);
}
/**
* Delivery, completion signalling (CB-106), and a readiness gate (CB-113): a message is delivered
* only when {@code ready} accepts the target — i.e. the worker's Claude has connected the bridge
* MCP. This holds the first delivery out of the worker's boot window, where herdr already reports
* {@code idle} but the TUI would drop an injected paste.
*/
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready) {
this(agents, turnListener, ready, _ -> {
});
}
/**
* Delivery, completion signalling (CB-106), a readiness gate (CB-113), and readiness cleanup
* (CB-114): {@code forget} is invoked with a target when its worker is gone — dropped
* (pane crash) or timed out on the readiness gate — so its stale presence/readiness is cleared
* and does not linger past the worker's life.
*/
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
this.agents = agents;
this.turnListener = turnListener;
this.ready = ready;
this.forget = forget;
}
/** A pending message and the future that completes when it has been delivered. */
private record Pending(String text, TurnToken token, CompletableFuture<Void> delivered) {
}
/** Per-worker delivery state, guarded by its own monitor (single writer per worker). */
private static final class Target {
final Deque<Pending> queue = new ArrayDeque<>();
boolean awaitingPickup; // sent a message, waiting for the worker to pick it up
int injectableSincePickup; // consecutive injectable samples while awaitingPickup
boolean awaitingCompletion; // a delivered message's turn is not yet known-complete
boolean turnObserved; // saw a real `working` sample since that delivery (turn ran)
int unknownSinceTurn; // consecutive `unknown` samples while a delegation is outstanding (CB-109)
int notReadySincePoll; // consecutive injectable samples a queued message waited on the readiness gate (CB-114)
boolean postTurnPending; // completion observed; adapter housekeeping has not started yet
boolean awaitingPostTurnPickup;
boolean postTurnObserved;
int injectableSincePostTurnPickup;
synchronized void add(Pending p) {
queue.add(p);
}
}
/**
* Queue {@code text} for delivery to {@code target} (a herdr {@code terminal_id}). Returns
* immediately with a future that completes when the message is actually sent — the worker
* may be mid-turn, in which case delivery waits for the next injectable window.
*
* <p>Uses an atomic map update so a concurrent {@link #drop} cannot slip between "find the
* target" and "queue the message" and orphan it in a target it just removed.
*/
public CompletableFuture<Void> enqueue(String target, String text, TurnToken token) {
CompletableFuture<Void> delivered = new CompletableFuture<>();
Pending p = new Pending(text, token, delivered);
targets.compute(target, (_, existing) -> {
Target t = (existing != null) ? existing : new Target();
t.add(p); // synchronized on the Target monitor — atomic with a concurrent drop
return t;
});
return delivered;
}
/**
* Feed a fresh status observation for {@code target}. Delivers the head of the queue iff the
* worker is injectable and no earlier message is still awaiting pickup. Serialized per target
* so the check and the send cannot race another delivery to the same worker; the delivered
* future is completed <em>after</em> the monitor is released so a caller's continuation never
* runs on the poller thread while it holds the lock.
*/
public void onStatus(String target, AgentStatus status) {
Target t = targets.get(target);
if (t == null) return;
Pending sent = null;
RuntimeException sendError = null;
boolean turnCompleted = false;
boolean turnFailed = false;
boolean resubmit = false;
boolean startPostTurn = false;
List<Pending> notReady = null; // queued messages failed because the worker never became ready
synchronized (t) {
if (status == AgentStatus.WORKING) {
if (t.awaitingPostTurnPickup) {
t.awaitingPostTurnPickup = false;
t.injectableSincePostTurnPickup = 0;
t.postTurnObserved = true;
}
// Definitive pickup: the worker is busy on our last message, and (if a delivery is
// outstanding) a real turn is now confirmed to be running.
t.awaitingPickup = false;
t.injectableSincePickup = 0;
t.unknownSinceTurn = 0;
t.notReadySincePoll = 0;
if (t.awaitingCompletion) t.turnObserved = true;
} else if (status.injectable()) { // IDLE or BLOCKED
t.unknownSinceTurn = 0;
if (t.awaitingPostTurnPickup) {
if (++t.injectableSincePostTurnPickup >= PICKUP_GRACE_POLLS) {
t.awaitingPostTurnPickup = false;
t.injectableSincePostTurnPickup = 0;
} else {
resubmit = true;
}
} else if (t.postTurnObserved) {
t.postTurnObserved = false;
}
if (t.awaitingPickup) {
if (++t.injectableSincePickup >= PICKUP_GRACE_POLLS) {
// Pickup edge was never sampled (turn faster than the poll, or status lag).
// Release the latch rather than wedge — and give up on synthesizing a
// completion for this message, since without a confirmed `working` we cannot
// trust that a task-processing turn actually ran.
t.awaitingPickup = false;
t.injectableSincePickup = 0;
t.awaitingCompletion = false;
t.turnObserved = false;
} else {
// Delivered but still idle → the worker hasn't picked it up; the submit
// keystroke likely raced the paste (esp. right as the TUI became ready).
// Re-nudge Enter (CB-113) until the worker starts (WORKING) or the grace ends.
resubmit = true;
}
}
if (!t.awaitingPickup) {
// A confirmed turn (a `working` sample was seen) that has now returned to idle is
// a trustworthy `working → idle` completion boundary.
if (t.awaitingCompletion && t.turnObserved) {
t.awaitingCompletion = false;
t.turnObserved = false;
turnCompleted = true;
if (turnListener.hasPostTurnAction(target)) {
t.postTurnPending = true;
startPostTurn = true;
}
}
// Deliver the next queued message only once the prior turn is fully settled, so a
// completion is never confused with the pickup of the following message — and only
// once the worker is available (CB-113), so we never paste into its boot window.
if (!t.awaitingCompletion && !t.postTurnPending
&& !t.awaitingPostTurnPickup && !t.postTurnObserved) {
Pending p = t.queue.peek();
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
try {
agents.send(target, p.text());
t.queue.poll();
t.awaitingPickup = true;
t.awaitingCompletion = true;
t.turnObserved = false;
t.injectableSincePickup = 0;
sent = p;
} catch (RuntimeException e) {
// Delivery failed at herdr; drop the poisoned message and surface it
// rather than blocking the queue behind it.
t.queue.poll();
sent = p;
sendError = e;
}
} else if (p != null && ++t.notReadySincePoll >= READINESS_GRACE_POLLS) {
// The worker has been idle-but-not-ready for the whole grace: its Claude
// never connected the bridge MCP (crashed during boot, or wedged on a
// startup prompt). The readiness gate would hold this message forever, so
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
log.warn("readiness grace for {} expired after {} polls ({}s): target never "
+ "became deliverable, so failing {} queued message(s) that never "
+ "reached its pane",
target, READINESS_GRACE_POLLS,
READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000, notReady.size());
t.queue.clear();
t.notReadySincePoll = 0;
}
}
}
} else {
// UNKNOWN (or any other non-injectable, non-working): not a safe window nor a
// reliable pickup signal, so we never deliver or release the pickup latch here. But
// an outstanding delegation whose worker has gone unresponsive — stuck in a state
// herdr can't classify (CB-109) — will never yield a working→idle boundary. After a
// sustained streak, declare it failed so the awaiting send resolves rather than
// riding out the async timeout. (This also frees a delivery that wedged before it
// was ever picked up, which the injectable-only pickup grace could never release.)
if (t.awaitingCompletion && ++t.unknownSinceTurn >= TURN_STALL_GRACE_POLLS) {
t.awaitingPickup = false;
t.awaitingCompletion = false;
t.turnObserved = false;
t.unknownSinceTurn = 0;
turnFailed = true;
}
}
// Reclaim the entry once the worker is fully quiescent (nothing queued, no pickup or
// completion awaited), so the map cannot grow without bound across short-lived workers.
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion
&& !t.postTurnPending && !t.awaitingPostTurnPickup && !t.postTurnObserved) {
targets.remove(target, t);
}
}
// Fire listeners / herdr calls after releasing the monitor so nothing runs on the poller
// thread while it holds the target lock.
if (resubmit) {
try {
agents.submit(target); // nudge a raced Enter so the pending paste submits
} catch (RuntimeException e) {
log.debug("resubmit to {} failed (will retry next poll): {}", target, e.getMessage());
}
}
if (notReady != null) {
// Worker never became available: forget its (never-set) readiness, unblock every queued
// caller, and route the awaiting send through the same failure path as a stalled turn so
// a blocking or async waiter resolves WORKER_FAILED rather than riding out the timeout.
forget.accept(target);
RuntimeException cause = new IllegalStateException(
target + " never became available (no bridge MCP connection within the boot window)");
for (Pending p : notReady) {
p.delivered().completeExceptionally(cause);
}
turnListener.onTurnFailed(target);
}
if (turnCompleted) {
if (startPostTurn) {
boolean started = turnListener.onTurnCompleteWithPostAction(target);
synchronized (t) {
t.postTurnPending = false;
if (started) {
t.awaitingPostTurnPickup = true;
t.injectableSincePostTurnPickup = 0;
}
if (t.queue.isEmpty() && !t.awaitingPostTurnPickup) {
targets.remove(target, t);
}
}
} else {
turnListener.onTurnComplete(target);
}
}
if (turnFailed) {
turnListener.onTurnFailed(target);
}
if (sent != null) {
if (sendError != null) {
log.warn("inject to {} failed, dropped message: {}", target, sendError.getMessage());
sent.delivered().completeExceptionally(sendError);
} else {
// Baseline the pane's pre-turn content so a misattributed completion (no new output)
// can't resolve this send with the previous turn's stale answer (CB-115).
turnListener.onDelivered(target, sent.token());
sent.delivered().complete(null);
}
}
}
/**
* Targets the poller must keep sampling: those with a queued message, an awaited pickup, or an
* awaited turn completion (so the {@code working → idle} boundary is observed).
*/
public Set<String> activeTargets() {
return targets.entrySet().stream()
.filter(e -> {
synchronized (e.getValue()) {
Target t = e.getValue();
return !t.queue.isEmpty() || t.awaitingPickup || t.awaitingCompletion
|| t.postTurnPending || t.awaitingPostTurnPickup || t.postTurnObserved;
}
})
.map(java.util.Map.Entry::getKey)
.collect(Collectors.toSet());
}
/**
* Forget a target whose worker is gone, failing every still-queued message so awaiting callers
* unblock instead of hanging forever. If a message had already been <em>delivered</em> but its
* turn was not yet resolved (CB-110 — the worker vanished mid-turn, e.g. its pane crashed), fire
* {@link TurnListener#onTurnFailed} for it: a delivered message is no longer in the queue, so
* failing queued waiters alone would leave that send's rendezvous hanging until the async
* timeout. Futures and listeners are completed after the monitor is released.
*/
public void drop(String target, Throwable cause) {
Target t = targets.remove(target);
if (t == null) return;
List<Pending> pending;
boolean hadDeliveredTurn;
synchronized (t) {
pending = new ArrayList<>(t.queue);
t.queue.clear();
hadDeliveredTurn = t.awaitingCompletion;
t.awaitingCompletion = false;
t.awaitingPickup = false;
t.postTurnPending = false;
t.awaitingPostTurnPickup = false;
t.postTurnObserved = false;
}
log.warn("{} is gone, dropping its queue: {} message(s) failed{}; cause: {}", target,
pending.size(),
hadDeliveredTurn ? " (including one turn already in flight whose completion was never confirmed)" : "",
cause.getMessage());
forget.accept(target); // the worker is gone — clear its readiness/presence too (CB-114)
for (Pending p : pending) {
p.delivered().completeExceptionally(cause);
}
// A queued send has no in-flight record, while a delivered turn does. CompletionResolver
// handles both forms and resolves its waiter at most once.
turnListener.onTurnFailed(target, cause.getMessage());
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import java.util.concurrent.ConcurrentHashMap;
import java.util.Set;
@@ -0,0 +1,93 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.HerdrException;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Set;
/**
* Drives the {@link Injector} by sampling each active worker's {@code agent_status} and
* feeding it in. A single virtual-thread loop polls only targets that have work outstanding,
* so an idle bridge does no herdr traffic.
*
* <p>This is the Stage-1 gate signal. It can later be replaced (or fronted) by a herdr
* {@code events.subscribe} stream without touching the {@link Injector} — the injector only
* consumes {@code onStatus} calls, however they are produced.
*/
public final class StatusPoller {
private static final Logger log = LoggerFactory.getLogger(StatusPoller.class);
private final AgentControl agents;
private final Injector injector;
private final StatusRefiner refiner;
private final long intervalMillis;
private volatile boolean running;
private Thread thread;
public StatusPoller(AgentControl agents, Injector injector, long intervalMillis) {
this(agents, injector, new StatusRefiner(agents), intervalMillis);
}
public StatusPoller(AgentControl agents, Injector injector, StatusRefiner refiner,
long intervalMillis) {
this.agents = agents;
this.injector = injector;
this.refiner = refiner;
this.intervalMillis = intervalMillis;
}
/** Start the polling loop on a virtual thread. Idempotent. */
public synchronized void start() {
if (running) return;
running = true;
thread = Thread.ofVirtual().name("status-poller").start(this::loop);
log.info("status poller started (interval {}ms)", intervalMillis);
}
private void loop() {
while (running) {
Set<String> active = injector.activeTargets();
for (String target : active) {
if (!running) return;
try {
// herdr's agent_status can misreport a settled worker as `unknown`; refine it
// against the pane content before it drives delivery/completion (CB-115).
AgentStatus status = refiner.refine(target, agents.status(target));
injector.onStatus(target, status);
} catch (HerdrException e) {
// The worker's agent is gone — stop trying and unblock its waiters.
if (e.code() != null && e.code().endsWith("_not_found")) {
log.debug("target {} gone; dropping its queue", target);
injector.drop(target, e);
} else {
log.debug("status poll for {} failed (will retry): {}", target, e.getMessage());
}
} catch (RuntimeException e) {
// Never let one target's unexpected error (e.g. an odd agent.get shape) kill
// the single poller thread and stall injection for every worker.
log.warn("unexpected error polling {}; skipping this round", target, e);
}
}
sleep();
}
}
private void sleep() {
try {
Thread.sleep(intervalMillis);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
running = false;
}
}
/** Stop the polling loop. Idempotent. */
public synchronized void stop() {
running = false;
if (thread != null) thread.interrupt();
}
}
@@ -1,7 +1,7 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -41,32 +41,15 @@ public final class StatusRefiner {
}
/**
* Return a trustworthy status for {@code target}, reading its pane through this refiner's own
* {@link AgentControl}. Equivalent to {@link #refine(String, AgentStatus, AgentControl)} with
* that control — kept for callers that only ever talk to one herdr daemon.
* Return a trustworthy status for {@code target}. Any non-{@code UNKNOWN} {@code raw} is returned
* unchanged; an {@code UNKNOWN} triggers a pane read and content classification. A read failure
* leaves it {@code UNKNOWN} (the safe default: no delivery, and the stall path still applies).
*/
public AgentStatus refine(String target, AgentStatus raw) {
return refine(target, raw, agents);
}
/**
* Return a trustworthy status for {@code target}. Any non-{@code UNKNOWN} {@code raw} is returned
* unchanged; an {@code UNKNOWN} triggers a pane read (through {@code control}) and content
* classification. A read failure leaves it {@code UNKNOWN} (the safe default: no delivery, and
* the stall path still applies).
*
* <p>CB-185: {@code control} must be the {@link AgentControl} for the <em>same</em> daemon the
* raw status was sampled from — a router splits lead and member targets across two herdr
* daemons, and reading a lead's pane through the member client (or vice versa) fails to find
* the pane and leaves the target wedged at {@code UNKNOWN} forever. Callers that route per
* target (e.g. {@code StatusPoller}) must pass that target's control explicitly rather than
* relying on the control fixed at construction.
*/
public AgentStatus refine(String target, AgentStatus raw, AgentControl control) {
if (raw != AgentStatus.UNKNOWN) return raw;
String pane;
try {
pane = control.read(target, PROBE_SOURCE);
pane = agents.read(target, PROBE_SOURCE);
} catch (RuntimeException e) {
log.debug("status refine read for {} failed; leaving UNKNOWN: {}", target, e.getMessage());
return AgentStatus.UNKNOWN;
@@ -1,12 +1,12 @@
package dev.ltms.fleet.inject;
package dev.ltms.bridged.inject;
import dev.ltms.fleet.msg.TurnToken;
import dev.ltms.bridged.msg.TurnToken;
/**
* Notified when a worker's delegated turn is observed to complete — a confirmed
* {@code WORKING → IDLE} transition after a delivery. This is the CB-106 completion signal the
* {@code CompletionResolver} uses to resolve a blocked send whose worker never called
* {@code fleet_reply}. Kept as a seam so the {@link Injector} needs no dependency on the message
* {@code bridge_reply}. Kept as a seam so the {@link Injector} needs no dependency on the message
* layer and stays unit-testable with a capturing fake.
*/
@FunctionalInterface
@@ -1,24 +1,22 @@
package dev.ltms.fleet.lead;
package dev.ltms.bridged.lead;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.PendingCloseMarker;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import dev.ltms.fleet.peer.PeerLauncher;
import java.util.stream.Collectors;
/**
* Starts the leads {@code fleet.leaders:} declares, when none is already running (CB-558).
@@ -28,7 +26,7 @@ import dev.ltms.fleet.peer.PeerLauncher;
* must never receive:
* <ol>
* <li>it appends the <em>reply charter</em> — "you are an off-subscription worker … end every turn
* with {@code fleet_reply}". A lead is the orchestrator; telling it that it is a worker is
* with {@code bridge_reply}". A lead is the orchestrator; telling it that it is a worker is
* exactly backwards.</li>
* <li>it registers the session with {@code SessionManager}, which subjects it to the idle reaper,
* the context cap and the shutdown drain. An idle lead is the normal state of a lead, so the
@@ -42,38 +40,13 @@ import dev.ltms.fleet.peer.PeerLauncher;
* broken later, quietly.
*
* <p><strong>Liveness, and the label round-trip.</strong> {@code LeadTabScanner} used to be able to
* promise that fleetd never writes a lead label. That is no longer true — an auto-launched lead is
* promise that bridged never writes a lead label. That is no longer true — an auto-launched lead is
* labelled by this class, and the scanner reads that label back. The risk this opens is not
* privilege escalation (the tab label never granted anything a pane could take for itself; see that
* class's javadoc), but <em>staleness</em>: a label left behind by a crashed session would otherwise
* read as a live lead forever, and the lead would never be relaunched. So a lead counts as live only
* when herdr also reports a <em>running agent</em> in that tab — see {@link #countLeads}. A labelled
* when herdr also reports a <em>running agent</em> in that tab — see {@link #liveLeads}. A labelled
* tab with no agent in it is not a lead.
*
* <p><strong>fleetd #359 — a stale label used to pile up, not just mislead.</strong> Finding "not
* live" here used to mean only one thing: launch another. The old label was left exactly where it
* was, so a daemon that restarted enough times — or hit one herdr read that missed a genuinely
* running agent — accumulated one more identically-labelled dead tab per occurrence, and
* {@code LeadTabScanner} (before its own #359 fix) reported every one of them as a lead.
*
* <p><strong>Review finding 1 — closing on one reading is worse than the bug.</strong> The first
* version of this fix closed a name's dead tabs the moment a single {@link #countLeads} reading
* called them dead. The ticket's own live evidence rules that out: on a real host, {@code
* agent.list} was seen reporting "0 live" for a tab that a plain {@code ps} confirmed was running a
* real session. Closing on that reading would have destroyed the operator's actual lead — a worse
* failure than the extra tab it replaces. So {@link #ensureLeads()} now needs the same dead reading
* <em>twice</em>, one restart apart, before it closes anything: the first time a labelled tab reads
* dead, it is only flagged ({@link PendingCloseMarker}), left running, untouched; it is closed only
* if a <em>later</em>, independently-connected reconcile still finds it dead while the flag is
* still there. A transient miss self-heals — the next reconcile sees the agent again and clears the
* flag (see {@code toUnflag} below) — so the worst case for a single bad reading is one extra tab
* surviving one more restart, never a live session destroyed. Two alternatives were considered and
* rejected: corroborating {@code agent.list} against a second, truly independent signal was dropped
* because nothing else herdr exposes proves "is a process attached to this pane" any better — a
* second call to the same unreliable source is not independent evidence; capping the close to "all
* but the most recent dead tab" was dropped because "most recent" has no reliable ordering across
* tab ids and would leave the true failure mode (a name that is <em>never</em> reconfirmed) growing
* by one tab per bad reading forever, which is the exact defect this ticket exists to fix.
*/
public final class LeadLauncher {
@@ -81,14 +54,14 @@ public final class LeadLauncher {
private final AgentControl agents;
private final WorkspaceControl spaces;
private final FleetConfig cfg;
private final BridgedConfig cfg;
/**
* @param agents herdr agent control (start, list)
* @param spaces workspace / tab control (ensure, create, label, list)
* @param cfg the loaded config — {@code fleet.leaders}, {@code profiles} and each lead's tab
*/
public LeadLauncher(AgentControl agents, WorkspaceControl spaces, FleetConfig cfg) {
public LeadLauncher(AgentControl agents, WorkspaceControl spaces, BridgedConfig cfg) {
this.agents = agents;
this.spaces = spaces;
this.cfg = cfg;
@@ -102,14 +75,14 @@ public final class LeadLauncher {
* the remaining leads are still attempted.
*/
public int ensureLeads() {
Map<String, FleetConfig.Leader> leaders = cfg.fleet().leaders();
Map<String, BridgedConfig.Leader> leaders = cfg.fleet().leaders();
if (leaders.isEmpty()) {
return 0;
}
Map<String, LeadCount> live;
Map<String, Integer> live;
try {
live = countLeads(leaders);
live = liveLeads(leaders);
} catch (HerdrException e) {
// Counting is the whole safety mechanism against double-spawning. If we cannot count, we
// must not guess — spawning a second orchestrator is worse than starting none.
@@ -118,56 +91,12 @@ public final class LeadLauncher {
}
int started = 0;
for (Map.Entry<String, FleetConfig.Leader> e : leaders.entrySet()) {
for (Map.Entry<String, BridgedConfig.Leader> e : leaders.entrySet()) {
String name = e.getKey();
FleetConfig.Leader lead = e.getValue();
LeadCount state = live.getOrDefault(name, LeadCount.NONE);
int running = state.running();
BridgedConfig.Leader lead = e.getValue();
int running = live.getOrDefault(name, 0);
int wanted = lead.instances();
// fleetd #359 review finding 1: a tab already flagged pending-close, still labelled for
// this lead, and STILL hosting no agent on this separate reconcile — two independent
// readings agree, so close it. A tab found dead for the first time is only flagged below,
// never closed on the spot.
for (String tabId : state.toClose()) {
log.info("lead '{}': closing tab {} — flagged pending-close on a previous reconcile "
+ "and still no agent running in it", name, tabId);
try {
spaces.closeTab(tabId);
} catch (RuntimeException cleanup) {
log.warn("could not close stale tab {} for lead '{}': {}",
tabId, name, cleanup.getMessage());
}
}
// A tab labelled for this lead, with no agent running in it, seen dead for the first
// time — flag it rather than closing it. One reading of `agent.list` is not enough
// evidence to destroy a tab that might genuinely be live (see the class javadoc).
for (String tabId : state.toFlag()) {
String flagged = lead.tabLabel() + PendingCloseMarker.SUFFIX;
log.info("lead '{}': tab {} has no agent running in it this reconcile — flagging it "
+ "'{}' rather than closing; it is only closed if a later reconcile still "
+ "finds it dead", name, tabId, flagged);
try {
spaces.renameTab(tabId, flagged);
} catch (RuntimeException cleanup) {
log.warn("could not flag stale tab {} for lead '{}': {}",
tabId, name, cleanup.getMessage());
}
}
// A previously-flagged tab that is running an agent again — the miss that flagged it was
// transient. Clear the flag so a future, unrelated miss starts its own two-reading count
// rather than closing on the strength of this one's already-spent flag.
for (String tabId : state.toUnflag()) {
log.info("lead '{}': tab {} is running an agent again — clearing its pending-close flag",
name, tabId);
try {
spaces.renameTab(tabId, lead.tabLabel());
} catch (RuntimeException cleanup) {
log.warn("could not clear the pending-close flag on tab {} for lead '{}': {}",
tabId, name, cleanup.getMessage());
}
}
if (running >= wanted) {
log.info("lead '{}': {} live, {} wanted — nothing to start", name, running, wanted);
continue;
@@ -176,12 +105,12 @@ public final class LeadLauncher {
// A lead with a `tab:` but no `profile:` is recognise-only by design: the operator
// opens it by hand. Say so once rather than looking like a silent failure.
log.info("lead '{}' is not live, and names no profile — it can be recognised but not "
+ "launched. Add `profile:` under fleet.leaders.{} to have fleetd start it.",
+ "launched. Add `profile:` under fleet.leaders.{} to have bridged start it.",
name, name);
continue;
}
FleetConfig.Profile profile = cfg.profiles().get(lead.profile());
BridgedConfig.Profile profile = cfg.profiles().get(lead.profile());
if (profile == null) {
log.warn("lead '{}' names profile '{}', which is not configured — not launching",
name, lead.profile());
@@ -198,105 +127,59 @@ public final class LeadLauncher {
}
/**
* How many live leads exist per configured name, and which of that name's labelled tabs are
* <em>not</em> live: a running agent in a tab labelled with that lead's exact {@code tab}
* (CB-579). Member workspaces are excluded, exactly as the scanner excludes them: a member must
* not be counted as a lead because it happens to sit in a matching tab.
* How many live leads exist per configured name: a running agent in a tab labelled with that
* lead's exact {@code tab} (CB-579). Member workspaces are excluded, exactly as the scanner
* excludes them: a member must not be counted as a lead because it happens to sit in a matching
* tab.
*
* <p>There used to be a second path here — a running agent on the terminal a
* {@code fleet.leaders.<name>.terminal} pin named, for a lead opened and pinned by hand. That
* pin is retired: {@code tab} is now the only field identity depends on, and {@link Agent}
* already carries {@link Agent#tabId()} directly, so a hand-opened lead is found the same way an
* auto-launched one is — by labelling its tab to match.
*
* <p>fleetd #359 review finding 1: a labelled tab with nothing running in it is split into
* {@code toClose} (already flagged pending-close by a previous reconcile, and still dead — two
* independent readings agree) and {@code toFlag} (dead for the first time — not enough evidence
* to close yet). {@code toUnflag} is the reverse: a tab flagged pending-close that is running an
* agent again, so the flag it carries no longer means anything and {@link #ensureLeads()} clears
* it.
*/
private record LeadCount(int running, List<String> toClose, List<String> toFlag, List<String> toUnflag) {
static final LeadCount NONE = new LeadCount(0, List.of(), List.of(), List.of());
}
private Map<String, Integer> liveLeads(Map<String, BridgedConfig.Leader> leaders) {
Set<String> memberSpaces = cfg.profiles().values().stream()
.map(BridgedConfig.Profile::workspace)
.filter(w -> w != null && !w.isBlank())
.collect(Collectors.toSet());
private Map<String, LeadCount> countLeads(Map<String, FleetConfig.Leader> leaders) {
// A lead and the members share ONE workspace now (the operator asked for a single "session"
// with many tabs), so a workspace can no longer be excluded wholesale — the lead lives in the
// member workspace by design. The sole discriminator is the exact tab label: a lead carries
// its configured `fleet.leaders.<name>.tab` ("lead: opus"), while a member carries its
// profile's `worker: {profile} #{n}` template. These never collide, so an exact-label match
// separates them without needing to know which workspace anyone is in.
// tabId → the lead name its label declares.
Map<String, String> nameByTab = new LinkedHashMap<>();
Set<String> flaggedTabIds = new LinkedHashSet<>();
for (Workspace ws : spaces.listWorkspaces()) {
if (ws.workspaceId() == null) {
if (ws.workspaceId() == null || memberSpaces.contains(ws.label())) {
continue;
}
for (Tab tab : spaces.listTabs(ws.workspaceId())) {
String declared = leadNameOf(tab.label(), leaders);
if (declared != null && tab.tabId() != null) {
nameByTab.put(tab.tabId(), declared);
if (PendingCloseMarker.isFlagged(tab.label())) {
flaggedTabIds.add(tab.tabId());
}
}
}
}
Set<String> liveTabIds = new LinkedHashSet<>();
Map<String, Integer> counts = new LinkedHashMap<>();
for (Agent a : agents.list()) {
String name = nameByTab.get(a.tabId());
if (name != null) {
counts.merge(name, 1, Integer::sum);
liveTabIds.add(a.tabId());
}
}
Map<String, List<String>> toCloseByName = new LinkedHashMap<>();
Map<String, List<String>> toFlagByName = new LinkedHashMap<>();
Map<String, List<String>> toUnflagByName = new LinkedHashMap<>();
nameByTab.forEach((tabId, name) -> {
boolean live = liveTabIds.contains(tabId);
boolean flagged = flaggedTabIds.contains(tabId);
if (live) {
if (flagged) {
toUnflagByName.computeIfAbsent(name, k -> new ArrayList<>()).add(tabId);
}
return;
}
if (flagged) {
toCloseByName.computeIfAbsent(name, k -> new ArrayList<>()).add(tabId);
} else {
toFlagByName.computeIfAbsent(name, k -> new ArrayList<>()).add(tabId);
}
});
Map<String, LeadCount> out = new LinkedHashMap<>();
for (String name : leaders.keySet()) {
out.put(name, new LeadCount(counts.getOrDefault(name, 0),
toCloseByName.getOrDefault(name, List.of()),
toFlagByName.getOrDefault(name, List.of()),
toUnflagByName.getOrDefault(name, List.of())));
}
return out;
return counts;
}
/**
* The configured lead a tab label names, or {@code null} for a label that names none.
*
* <p>Matched exactly (case-insensitively) against each lead's configured {@code tab}, so an
* operator's {@code "lead: something-else"} tab is not mistaken for a configured lead. A
* trailing {@link PendingCloseMarker} is stripped first, so a tab this class flagged on a
* previous reconcile is still recognised as the same lead's tab on this one.
* operator's {@code "lead: something-else"} tab is not mistaken for a configured lead.
*/
private String leadNameOf(String label, Map<String, FleetConfig.Leader> leaders) {
private String leadNameOf(String label, Map<String, BridgedConfig.Leader> leaders) {
if (label == null) {
return null;
}
String l = PendingCloseMarker.strip(label);
for (Map.Entry<String, FleetConfig.Leader> e : leaders.entrySet()) {
String l = label.strip();
for (Map.Entry<String, BridgedConfig.Leader> e : leaders.entrySet()) {
String tab = e.getValue().tabLabel();
if (tab != null && l.equalsIgnoreCase(tab.strip())) {
return e.getKey();
@@ -306,7 +189,7 @@ public final class LeadLauncher {
}
/** Start one lead. Returns false (having logged) rather than throwing on any failure. */
private boolean launch(String name, FleetConfig.Leader lead, FleetConfig.Profile profile) {
private boolean launch(String name, BridgedConfig.Leader lead, BridgedConfig.Profile profile) {
String label = lead.tabLabel();
String cwd = (lead.cwd() == null || lead.cwd().isBlank())
? System.getProperty("user.dir") : lead.cwd();
@@ -353,7 +236,7 @@ public final class LeadLauncher {
* The herdr agent kind for this profile — the same value the matching member adapter passes, so
* herdr resolves the same executable for a lead as it does for a member on that backend.
*/
private static String herdrKind(FleetConfig.Profile profile) {
private static String herdrKind(BridgedConfig.Profile profile) {
return profile.isOpenCode() ? "opencode" : "claude";
}
@@ -365,12 +248,11 @@ public final class LeadLauncher {
* any other primary. This is the single most important difference from the member launchers;
* do not "unify" it back.
*/
private List<String> leadArgv(FleetConfig.Profile profile) {
private List<String> leadArgv(BridgedConfig.Profile profile) {
List<String> argv = new ArrayList<>(profile.argv());
if (profile.hasMcp()) {
argv.add("--mcp-config");
argv.add("{\"mcpServers\":{\"" + PeerLauncher.MCP_MOUNT_NAME
+ "\":{\"type\":\"http\",\"url\":\""
argv.add("{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ profile.mcpUrl() + "\"}}}");
}
// Appended last, for the same reason the member launcher does it (CB-533): the argv is
@@ -392,7 +274,7 @@ public final class LeadLauncher {
* definition, so there is no configuration under which pointing it elsewhere is correct, and a
* profile that carries them (a member profile reused as a lead's backend) must not leak them in.
*/
private Map<String, String> leadEnv(FleetConfig.Profile profile) {
private Map<String, String> leadEnv(BridgedConfig.Profile profile) {
Map<String, String> out = new LinkedHashMap<>();
if (profile.env() != null) {
out.putAll(profile.env());
@@ -1,4 +1,4 @@
package dev.ltms.fleet.logging;
package dev.ltms.bridged.logging;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,66 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.herdr.PaneLocator;
/**
* Resolves <em>who is calling</em> an MCP tool from the connection alone — the anti-spoofing
* identity model of the MCP contract. It ties the connection's loopback peer PID (from the OS)
* to a herdr agent pane (from herdr), yielding the caller's worker {@code terminal_id}. A caller
* that maps to no worker pane — the primary, or an off-host client — resolves to {@code null}.
*
* <p>Both sources are authoritative and unforgeable: the OS reports the real connecting PID, and
* herdr owns the PID→pane mapping. A worker cannot claim to be another worker, nor the primary.
* Single-host only (the herd shares the {@code bridged} host); the token path is the split-host
* fallback.
*/
public final class ConnectionIdentity {
private final PaneLocator panes;
private final PeerPidLookup pids;
private final ProcessCwdLookup cwds;
/** Identity only (no cwd resolution — {@link #cwdForPid} returns {@code null}). */
public ConnectionIdentity(PaneLocator panes, PeerPidLookup pids) {
this(panes, pids, _ -> null);
}
/** Identity plus cwd resolution (CB-112 — inherit the primary's directory on spawn). */
public ConnectionIdentity(PaneLocator panes, PeerPidLookup pids, ProcessCwdLookup cwds) {
this.panes = panes;
this.pids = pids;
this.cwds = cwds;
}
/**
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
* primary / an off-host client) and its {@code pid} (or {@code -1} if not resolvable).
*/
public record Caller(String terminal, long pid) {
}
/** Resolve the caller's terminal and PID from one peer-PID lookup. */
public Caller resolve(String remoteAddr, int remotePort) {
if (!isLoopback(remoteAddr)) {
return new Caller(null, -1); // only same-host callers can be workers
}
long pid = pids.pidForLocalPort(remotePort);
return new Caller(panes.terminalForPid(pid), pid);
}
/**
* The calling worker's {@code terminal_id}, or {@code null} if the caller is not a known
* on-host worker (treat as the primary).
*/
public String callerTerminal(String remoteAddr, int remotePort) {
return resolve(remoteAddr, remotePort).terminal();
}
/** The working directory of {@code pid} (the primary's cwd on an MCP spawn), or {@code null}. */
public String cwdForPid(long pid) {
return pid > 0 ? cwds.cwdForPid(pid) : null;
}
private static boolean isLoopback(String addr) {
return "127.0.0.1".equals(addr) || "::1".equals(addr) || "0:0:0:0:0:0:0:1".equals(addr);
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -41,14 +41,6 @@ public final class LsofPeerPidLookup implements PeerPidLookup {
if (!p.waitFor(2, TimeUnit.SECONDS)) {
p.destroyForcibly();
}
if (found < 0) {
// fleetd #317: this is the silent path — lsof ran clean and simply reported no
// matching process (e.g. queried before the OS socket table settles). Previously
// this logged nothing at all, which is exactly why the escalation went unnoticed;
// the exception path below already logs. A caller now refused because of this is
// still refused (never promoted) — this line only makes the refusal diagnosable.
log.debug("lsof peer-pid lookup for port {} found no matching process", port);
}
return found;
} catch (Exception e) {
log.debug("lsof peer-pid lookup for port {} failed: {}", port, e.getMessage());
@@ -1,4 +1,4 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
/**
* Resolves the OS PID that owns a loopback TCP source port — the OS half of connection-based MCP
@@ -1,4 +1,4 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -11,7 +11,7 @@ import java.util.concurrent.atomic.AtomicReference;
* Single-slot, thread-safe registry for the primary's herdr {@code terminal_id}.
*
* <p>Populated from the caller terminal of orchestration-side MCP tools
* ({@code fleet_send}, {@code fleet_spawn}) — tools that only the primary calls.
* ({@code bridge_send}, {@code bridge_spawn}) — tools that only the primary calls.
* A pinned terminal (from config) seeds the registry at construction and makes
* subsequent {@link #record(String)} calls no-ops.
*
@@ -29,7 +29,7 @@ public final class PrimaryRegistry {
/**
* CB-532: worker terminal → the lead that delegated to it. The single slot above answers "who is
* THE primary", a question with no correct answer once two leads orchestrate the same fleet:
* whichever called {@code fleet_send} first captured every nudge, including nudges for the
* whichever called {@code bridge_send} first captured every nudge, including nudges for the
* other lead's delegations. This map answers the question that actually matters — "who is
* waiting on THIS worker" — and is what lets {@code primary.terminal} be retired.
*/
@@ -71,7 +71,7 @@ public final class PrimaryRegistry {
*
* <p>Called from the {@code MessageService} accepted-delivery hook — only after a send has won
* the session's send lock and queued delivery — where both halves are known (CB-548). It is
* deliberately <em>not</em> called at {@code fleet_send} request time: a concurrent sender that
* deliberately <em>not</em> called at {@code bridge_send} request time: a concurrent sender that
* times out {@code BUSY} must not steal a live delegation's reply routing without ever owning
* the turn. Last writer wins — if a second lead's later send is accepted, replies follow the
* lead that most recently delegated to it, which is the one waiting.
@@ -1,4 +1,4 @@
package dev.ltms.fleet.mcp;
package dev.ltms.bridged.mcp;
/**
* Resolves a process's current working directory from its PID — the OS half of CB-112's
@@ -0,0 +1,427 @@
package dev.ltms.bridged.member;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>Claude Code</strong> — the safe path from a
* delegation request to a running off-subscription Claude.
*
* <p>Everything transport-related (tab/pane placement, the CB-306 spawn-readiness gate, unique
* naming, CB-117 orphan reap, teardown, listing, cwd resolution) lives in the base. This class
* supplies only the two Claude-specific seams:
* <ul>
* <li>the {@code claude} name prefix (so reap matches {@code claude-*} panes, never another
* adapter's), and</li>
* <li>{@link #buildLaunch}, which encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em> touching
* herdr, and mount the bridge MCP + reply charter as inline launch flags. A worker's base_url
* lives in the env map handed to herdr and nowhere else; {@code bridged}'s own environment is
* never mutated, and nothing is written to the worker's profile.</li>
* </ul>
*/
public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "claude";
private static final Logger log = LoggerFactory.getLogger(ClaudeCodeLauncher.class);
private final SubscriptionGuard guard;
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so
* existing deployments and tests keep the legacy non-blocking spawn semantics.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, guard, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300));
}
/**
* Production constructor with spawn-ready gate enabled. The gate polls {@code agents.status()}
* until the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs, spawnReadyPollMs, null);
}
/**
* Production constructor carrying the fleet-wide tab-label template (CB-557). The template comes
* from {@code fleet.tabLabel}, which a profile cannot know because it names the member's
* <em>role</em>; a profile may still override it with its own {@code tabLabel}.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<BridgedConfig.Fleet> fleet) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
fleet);
}
/**
* Production constructor, plus the CB-596 {@code memberCredentials} policy supplier.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
fleet, memberCredentials);
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
* {@code sleeper} is never called when the gate is disabled ({@code spawnReadyTimeoutMs == 0}).
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param guard subscription-boundary guard (checked before spawning)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (e.g. {@code () -> Thread.sleep(pollMs)}); it
* already encodes the poll interval, so the 8th positional argument
* (poll ms) is accepted for API symmetry but otherwise unused here
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, null);
}
/**
* Full testability constructor, plus the fleet-wide tab-label template (CB-557).
*
* @param fleet live fleet config, read once for each spawn
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<BridgedConfig.Fleet> fleet) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet);
this.guard = guard;
}
/**
* Full testability constructor, plus the CB-596 {@code memberCredentials} policy supplier.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials);
this.guard = guard;
}
/**
* Full testability constructor, plus an injectable host-env-names source for the CB-596
* criterion-4 gap detector. Test seam only — every production call site leaves this at the
* default (the real {@code System.getenv()} key set) via the constructor above.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials,
Supplier<Set<String>> hostEnvNames) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials, hostEnvNames);
this.guard = guard;
}
/**
* {@inheritDoc}
*
* <p>The spawn sequence encodes the subscription boundary: assert the profile's base_url is on
* the allowlist <em>before</em> any herdr call, then build the worker env with
* {@code ANTHROPIC_*}, the parity-neutral git-forge grant, and the bridge MCP + reply charter
* mounted as inline launch flags. When the request carries session identity (CB-547a) it is
* applied here — see {@link #applySessionIdentity}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
// CB-539: a profile may deliberately opt into the subscription (subscription: true) when no
// off-subscription endpoint exists for it — e.g. `sonnet` on `ccs`. That profile gets no
// ANTHROPIC_BASE_URL/AUTH_TOKEN (there is nothing to point them at) and the guard's base_url
// requirement is skipped FOR IT ONLY. Every other profile keeps the hard boundary below.
boolean onSubscription = cfg.isSubscription();
String baseUrl = cfg.baseUrl();
if (onSubscription) {
// NO SILENT CONTRADICTION: subscription:true + a baseUrl state opposite intents; refuse
// loudly rather than pick a winner.
if (baseUrl != null && !baseUrl.isBlank()) {
throw new IllegalStateException("profile '" + cfg.profile()
+ "' sets both subscription: true and a baseUrl ('" + baseUrl + "') — the two "
+ "are contradictory: a subscription profile must not point at an endpoint. "
+ "Drop baseUrl, or drop subscription: true.");
}
// Visible without anyone going looking for it: this worker bills the subscription.
log.warn("spawning profile '{}' on the Claude subscription (subscription: true) — this "
+ "worker WILL bill the operator's subscription", cfg.profile());
} else {
guard.assertWorker(baseUrl); // hard stop before we spawn anything
}
Map<String, String> workerEnv = baseEnv(cfg);
if (onSubscription) {
// CB-542 belt-and-braces: on the subscription path no guard vets these two keys, and the
// profile's env: is layered in by baseEnv — so strip any that rode in there. Config load
// already rejects this (loudly, naming the profile); this makes the boundary hold even
// for a profile built in code that never passed through that validation.
workerEnv.remove("ANTHROPIC_BASE_URL");
workerEnv.remove("ANTHROPIC_AUTH_TOKEN");
} else {
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", env.apply(cfg.tokenEnv()));
}
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
applyGitToken(workerEnv, cfg);
// CB-547a: Claude Code can MINT its own session id, so bridged chooses it — a fresh spawn
// gets a UUID we pass as --session-id and return from agentSessionId(), so the resume
// handle is known BEFORE the agent has written anything; a resume spawn adopts its prior
// id via -r and passes no --session-id (the two conflict). Both are injected before the
// model flag so --model keeps outranking the operator's own argv.
// mutableArgv: argvWithBridge may hand back the profile's own (immutable) List.of when it
// has neither MCP nor a charter — session flags must be added into a list we own.
List<String> argv = mutableArgv(argvWithBridge(cfg, spec));
String agentSessionId = applySessionIdentity(argv, spec.sessionName(), spec.resumeSessionId());
return new Launch(workerEnv, argvWithModel(argv, cfg), agentSessionId);
}
/**
* Add the Claude-specific session-identity flags to {@code argv} and return the peer's OWN
* session id — the resume handle. A resume request passes the prior id via {@code -r} and
* returns that id; a fresh named session mints a new UUID, passes it via {@code --session-id},
* and returns the mint. The bridge's logical name rides along as {@code -n} when present. When
* <em>no</em> identity is requested (sessionName and resumeSessionId both blank) this adds
* nothing and returns {@code null}, keeping the legacy no-identity launch byte-identical.
*/
private static String applySessionIdentity(List<String> argv, String sessionName, String resumeSessionId) {
boolean resuming = resumeSessionId != null && !resumeSessionId.isBlank();
boolean named = sessionName != null && !sessionName.isBlank();
if (!resuming && !named) {
return null; // no identity requested — keep the legacy launch byte-identical
}
if (named) {
argv.add("-n");
argv.add(sessionName);
}
if (resuming) {
argv.add("-r");
argv.add(resumeSessionId);
return resumeSessionId;
}
String minted = UUID.randomUUID().toString();
argv.add("--session-id");
argv.add(minted);
return minted;
}
/**
* The launch argv, plus an inline {@code --mcp-config} when {@code worker.mcpUrl} is set, the
* CB-617 charter flags, and {@code --agent <role>} when the role has an agent-definition file
* under the worker's cwd. Neither touches the profile's config; all are pure command-line flags.
* This inline-flag mount is Claude Code specific — other adapters mount MCP and instructions
* their own way.
*
* <p>CB-617: the role charter is operator-authored and often multi-line, so it can never be a
* single inline argv element — herdr refuses to shell-encode a multi-line argument
* ({@code invalid_agent_argument}). It is written to a temp file instead and mounted with
* {@code --append-system-prompt-file}, which this host confirms Claude Code accepts for a
* multi-line file.
*
* <p>CB-618: Claude Code refuses to start when BOTH {@code --append-system-prompt} and
* {@code --append-system-prompt-file} are on the command line ("Cannot use both ... Please use
* only one"), so the two charters can never travel on separate flags. When both are present they
* are concatenated into the one file, role charter first and reply charter last — last is where
* the reply rule must sit, because it is the rule that must survive. When only the reply charter
* is present it keeps its proven inline {@code --append-system-prompt} delivery, which is also
* the only form that reaches a member with no repo checkout.
*/
private List<String> argvWithBridge(BridgedConfig.Profile cfg, LaunchSpec spec) {
String roleCharter = nonBlank(spec.roleCharter());
String replyCharter = nonBlank(spec.replyCharter());
Path agentFile = agentDefinitionFile(spec.cwd(), spec.role(), ".claude", "agents");
if (!cfg.hasMcp() && roleCharter == null && replyCharter == null && agentFile == null) {
return cfg.argv();
}
List<String> argv = mutableArgv(cfg.argv());
if (cfg.hasMcp()) {
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ cfg.mcpUrl() + "\"}}}";
argv.add("--mcp-config");
argv.add(mcpJson);
}
if (roleCharter != null) {
String combined = replyCharter == null ? roleCharter : roleCharter + "\n\n" + replyCharter;
argv.add("--append-system-prompt-file");
argv.add(writeCharterFile(combined).toString());
} else if (replyCharter != null) {
argv.add("--append-system-prompt");
argv.add(replyCharter);
}
if (agentFile != null) {
argv.add("--agent");
argv.add(spec.role().wireName());
}
return argv;
}
/** {@code s}, or {@code null} when {@code s} is null/blank — the charter-presence test used above. */
private static String nonBlank(String s) {
return (s == null || s.isBlank()) ? null : s;
}
/**
* Write the role charter to a fresh temp file so it can be mounted with
* {@code --append-system-prompt-file} instead of riding inline in argv (CB-617). Best-effort
* cleaned via {@code deleteOnExit} — the same disposable-worker-config cleanup
* {@link OpenCodeLauncher#writeConfig} already uses for its charter file, since the process that
* reads this file (the spawned peer) outlives this JVM call and there is no spawn-scoped teardown
* hook to delete it synchronously.
*/
private static Path writeCharterFile(String charterText) {
try {
Path file = Files.createTempFile("bridged-role-charter-", ".md");
Files.writeString(file, charterText);
file.toFile().deleteOnExit();
return file;
} catch (IOException e) {
throw new UncheckedIOException("cannot write role charter temp file", e);
}
}
/**
* Pin the model on the command line as well as in {@code ANTHROPIC_MODEL} (CB-533).
*
* <p>The env var alone is not a reliable pin for this adapter, because the argv is usually a
* launcher rather than {@code claude} itself — {@code ["ccs", "<profile>"]} — and {@code ccs}
* exports its profile's own model family ({@code ANTHROPIC_MODEL}, {@code DEFAULT_OPUS/SONNET/
* HAIKU}, {@code CLAUDE_CODE_SUBAGENT_MODEL}) over whatever it inherited. A worker profile that
* set {@code model:} therefore got silently overruled by its own launcher. Claude Code's
* {@code --model} flag outranks the environment, and {@code ccs <profile> [claude-args...]}
* passes trailing arguments through, so the flag survives the wrapper.
*
* <p>Appended last so it also outranks anything in the operator's own {@code argv}. Profiles
* that deliberately leave {@code model:} unset (letting {@code ccs} own model selection, as
* {@code gx10} does) are untouched — this adds nothing when there is nothing to add. This is
* the {@code kind: claude} counterpart of the opencode adapter's {@code -m provider/model}.
*/
private static List<String> argvWithModel(List<String> argv, BridgedConfig.Profile cfg) {
if (cfg.model() == null || cfg.model().isBlank()) {
return argv;
}
List<String> withModel = mutableArgv(argv);
withModel.add("--model");
withModel.add(cfg.model());
return withModel;
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE,
Capability.CONTEXT_RESET, Capability.ORPHAN_REAP,
Capability.SESSION_NAME, Capability.SESSION_RESUME);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
@Override
public boolean clearContext(String id) {
String target = agentTarget(id);
if (target == null) {
return false;
}
// This deliberately bypasses Injector: /clear is housekeeping, not a delegated turn.
agents().send(target, "/clear");
return true;
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Profile::hasGitToken);
}
// --- CB-117 reap predicate (Claude prefix), kept for direct unit testing -------------------
/**
* Whether {@code name} is a Claude Code bridge worker started by a <em>different</em> process
* than {@code currentNonce}. A thin {@code claude}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -0,0 +1,510 @@
package dev.ltms.bridged.member;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.BackendQuarantine;
import dev.ltms.bridged.placement.PlacementCandidate;
import dev.ltms.bridged.placement.PlacementContext;
import dev.ltms.bridged.placement.PlacementException;
import dev.ltms.bridged.placement.PlacementPolicies;
import dev.ltms.bridged.placement.PlacementPolicy;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.Collections;
import java.util.EnumSet;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Function;
import java.util.function.Supplier;
import java.util.stream.Collectors;
/**
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
* router in front of one {@link HerdrPeerLauncher} per peer {@code kind} (Claude Code, opencode, …).
* It owns no transport of its own; it dispatches each SPI call to the delegate that owns the profile
* involved, and fans the fleet-wide queries (list/reap/caps/profiles) across all delegates.
*
* <p>Routing rules:
* <ul>
* <li><strong>By profile</strong> — {@link #spawn}, {@link #effectiveCwd}, {@link #parityOverlay}
* resolve the profile (a null/blank name → the global {@link #defaultProfile}) and delegate to
* the single adapter that declares it. Profiles partition cleanly across adapters: the
* constructor rejects a name claimed by two.</li>
* <li><strong>By pane id</strong> — {@link #stop} routes to the adapter that spawned that pane
* (recorded at spawn time). A pane the composite never spawned (only real for a caller that
* hand-rolls an id) falls back to the first delegate; teardown is pane-id addressed and
* tab cleanup is single-occupant guarded, so it is safe either way.</li>
* <li><strong>Fleet-wide</strong> — {@link #reapOrphanWorkers} and {@link #capabilities} fan out
* and combine. {@link #list} is deduplicated by pane id because every herdr-backed delegate
* shares one herdr connection and so reports the same global agent set.</li>
* </ul>
*
* <p>CB-518: an unqualified spawn is routed through a {@link PlacementPolicy}. The default
* {@code fixed} policy reproduces the historical default-profile behaviour; {@code weighted} uses
* smooth weighted round-robin with {@code maxLoad} gating. If a chosen profile fails with
* {@link PeerUnreachableException}, the composite advances to the next available candidate and
* retries, bounded by the number of candidates.
*/
public final class CompositePeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(CompositePeerLauncher.class);
private final List<HerdrPeerLauncher> delegates;
private final Map<String, HerdrPeerLauncher> byProfile;
private final String defaultProfile;
/** paneId → the delegate that spawned it, so {@link #stop} tears down through the right adapter. */
private final Map<String, HerdrPeerLauncher> spawnedBy = new ConcurrentHashMap<>();
private final Function<String, Integer> liveCount;
/**
* CB-559: the placement inputs are read <em>per spawn</em>, not captured at construction, so a
* config reload changes where the next member lands without a restart. These are the hot keys —
* role pools, an existing profile's weight/maxLoad, and the placement policy. What cannot change
* this way is the set of adapters ({@link #byProfile}), because a new backend needs a launcher
* and launchers are built once; {@code ConfigRef} classifies that as deferred and says so.
*/
private final Supplier<Map<String, BridgedConfig.Profile>> profileConfigs;
private final Supplier<PlacementPolicy> placementPolicy;
/** CB-578 stage B: credential cooldown, checked before an explicit spawn and filtered into placement. */
private final BackendQuarantine quarantine;
/**
* CB-557: the role pools an unqualified spawn draws its candidates from. A supplier that yields
* {@code null}, and an empty pool for a role, both fall back to every configured profile — the
* pre-CB-557 behaviour.
*/
private final Supplier<BridgedConfig.Fleet> fleet;
/**
* Backward-compatible constructor: fixed placement, no live-counting. Use this for tests and
* simple wiring; it preserves the pre-CB-518 behaviour exactly.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to (may be null)
* @throws IllegalArgumentException if {@code delegates} is empty or two adapters claim one profile
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates, String defaultProfile) {
this(delegates, defaultProfile, Map.of(), PlacementPolicies.fixed(), _ -> 0);
}
/**
* Production constructor with a placement policy and live-worker counter. Quarantine (CB-578
* stage B) is off for this constructor — {@link BackendQuarantine#none()} — since it predates
* the feature and existing callers of this exact overload never exercised it; use the 7-arg
* overload below to wire a real {@link BackendQuarantine}.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to under {@code fixed} policy
* @param profileConfigs all configured worker profiles (used for candidate weights/caps)
* @param placementPolicy which policy governs unqualified spawns
* @param liveCount live worker count per profile (must never return {@code null})
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount) {
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, null, BackendQuarantine.none());
}
/**
* Production constructor with role pools (CB-557). An unqualified spawn draws its candidates from
* {@code fleet.<role>} instead of from every configured profile, so a reviewer is placed on a
* reviewer backend and never on, say, the architect-only one. Quarantine is off for this
* constructor too, for the same reason as the 5-arg overload above.
*
* @param fleet the configured role pools; {@code null} ⇒ every profile is a candidate for every
* role, which is the pre-CB-557 behaviour
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
BridgedConfig.Fleet fleet) {
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet, BackendQuarantine.none());
}
/**
* Production constructor with role pools and quarantine (CB-578 stage B). The full-featured
* non-reloading form; {@link #CompositePeerLauncher(List, String, Supplier, Function, BackendQuarantine)}
* is what {@code Bridged.main} actually wires up.
*
* @param quarantine required — pass {@link BackendQuarantine#none()} for a caller that does not
* want the feature, never a defaulting overload (CB-578 stage B's own rule).
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
BridgedConfig.Fleet fleet,
BackendQuarantine quarantine) {
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
// differ from one JVM run to the next.
this(delegates, defaultProfile,
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
constant(placementPolicy), liveCount, constant(fleet), quarantine);
}
/**
* Production constructor that re-reads its placement inputs per spawn (CB-559), so a config
* reload retargets the next member without a restart.
*
* @param config the live configuration — read at every spawn, never captured
* @param quarantine required — CB-578 stage B; pass {@link BackendQuarantine#none()} to opt out
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Supplier<BridgedConfig> config,
Function<String, Integer> liveCount,
BackendQuarantine quarantine) {
this(delegates, defaultProfile,
() -> config.get().profiles(),
() -> PlacementPolicies.fromName(config.get().placement()),
liveCount,
() -> config.get().fleet(),
quarantine);
}
/** The all-suppliers form every other constructor funnels into. */
private CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Supplier<Map<String, BridgedConfig.Profile>> profileConfigs,
Supplier<PlacementPolicy> placementPolicy,
Function<String, Integer> liveCount,
Supplier<BridgedConfig.Fleet> fleet,
BackendQuarantine quarantine) {
this.fleet = fleet;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
if (delegates.isEmpty()) {
throw new IllegalArgumentException("at least one peer adapter must be configured");
}
this.delegates = List.copyOf(delegates);
this.defaultProfile = defaultProfile;
this.profileConfigs = profileConfigs;
this.placementPolicy = placementPolicy;
this.liveCount = liveCount;
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
HerdrPeerLauncher prev = index.putIfAbsent(profile, d);
if (prev != null) {
throw new IllegalArgumentException(
"worker profile '" + profile + "' is claimed by two peer adapters");
}
}
}
// Order-preserving for the same reason, and because profiles() is user-visible (bridge_profiles).
this.byProfile = Collections.unmodifiableMap(index);
}
/** A supplier of a value fixed at construction — how the non-reloading constructors funnel in. */
private static <T> Supplier<T> constant(T value) {
return () -> value;
}
/**
* The currently-configured profiles, never null.
*
* <p>Read fresh on every call so a reload is visible; a caller that needs two consistent reads
* takes one local, as {@link #poolFor} does.
*/
private Map<String, BridgedConfig.Profile> profiles0() {
Map<String, BridgedConfig.Profile> m = profileConfigs.get();
return m == null ? Map.of() : m;
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
private HerdrPeerLauncher route(String profileName) {
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (resolved == null) {
// No profile and no default configured — hand to the first delegate so it raises the
// same "no default" error it would on its own; keeps the SPI contract single-sourced.
return delegates.getFirst();
}
HerdrPeerLauncher d = byProfile.get(resolved);
if (d == null) {
throw new IllegalArgumentException("unknown worker profile: " + resolved);
}
return d;
}
@Override
public PeerHandle spawn(SpawnRequest req) {
String requestedProfile = req.profileName();
if (requestedProfile != null && !requestedProfile.isBlank()) {
// An explicit profile bypasses the placement policy, but not the capacity cap: maxLoad
// is documented as an unconditional limit on this profile (BridgedConfig.Profile), and
// the charter makes explicit-profile spawns the normal path — so skipping the check
// here would leave the cap dead config in real operation.
HerdrPeerLauncher d = route(requestedProfile);
enforceNotQuarantined(requestedProfile);
enforceMaxLoad(requestedProfile);
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
}
// CB-557: an unqualified spawn is placed inside the pool of the role it asked for, not across
// the whole profile list. An EXPLICIT profile (above) is left alone on purpose — it is the
// operator overriding, and refusing it would break `bridge_spawn{profile:"opus"}`, which
// carries no role and so would be judged against the dev pool it was never meant for.
List<PlacementCandidate> candidates = candidates(req.role());
String roleDefault = defaultProfileFor(req.role());
Set<String> unreachable = new HashSet<>();
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
// call, so re-deriving it per retry would only cost work, never change the answer.
Set<String> quarantined = quarantinedProfiles(candidates);
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable, quarantined);
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
for (int attempt = 0; attempt < maxAttempts; attempt++) {
// Deliberately uncaught: when no candidate is left (all at cap, or all unreachable) the
// policy already throws a clear message. Catching it to rethrow a generic
// PeerUnreachableException would replace a precise diagnosis with a vague one.
PlacementCandidate chosen = placementPolicy.get().select(ctx);
HerdrPeerLauncher d = byProfile.get(chosen.profile());
if (d == null) {
// A configured profile with no adapter is a wiring bug; fail fast.
unreachable.add(chosen.profile());
continue;
}
// CB-547a: route the chosen profile but keep the caller's session identity — dropping it
// here would silently sever the resume handle on every policy-routed spawn. CB-557: the
// role rides along for the same reason, or a routed spawn would be labelled as a dev.
SpawnRequest routedReq = new SpawnRequest(chosen.profile(), req.requestedCwd(), req.callerCwd(),
req.sessionName(), req.resumeSessionId(), req.role());
try {
PeerHandle handle = d.spawn(routedReq);
spawnedBy.put(handle.id(), d);
return handle;
} catch (PeerUnreachableException e) {
log.warn("spawn on profile {} unreachable, will retry next candidate if any: {}",
chosen.profile(), e.getMessage());
unreachable.add(chosen.profile());
// Update the context for the next selection so the policy excludes this profile.
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable, quarantined);
}
}
throw new PeerUnreachableException(
"no reachable worker profile available after trying " + unreachable.size()
+ " candidate(s): " + String.join(", ", unreachable));
}
/**
* Refuse an explicit-profile spawn when the profile is at its {@code maxLoad} cap.
*
* <p>maxLoad is a documented, unconditional capacity limit (see {@code BridgedConfig.Profile#maxLoad}),
* and the charter makes explicit-profile spawns the normal path — so enforcing it only in placement
* ({@code PlacementPolicyUtil}, package-private, hence not linked) would leave the cap dead config
* on every call that names a profile. Same rule as placement: {@code live >= cap} is at capacity.
*
* <p>Deliberately no fallback to another profile: the caller named {@code profile} for a cost/model
* reason, and silently re-routing a paid-tier (subscription) request elsewhere is worse than
* refusing it. A caller that wants placement should omit the profile and let the policy pick.
*
* <p>Known TOCTOU limitation — documented, not fixed. {@link #liveCount} is read outside any lock and
* {@code SessionManager} registers a session only after {@code launcher.spawn} returns, so two
* genuinely concurrent spawns can both pass this check. The race already exists on the placement
* path. Closing it needs slot reservation in the registry; serializing spawn here would block on
* the readiness gate and is a far worse trade.
*
* @param profile the profile the caller explicitly named
* @throws PlacementException when the profile is at capacity
*/
/**
* Refuse an explicit-profile spawn whose credential is quarantined (CB-578 stage B): a prior
* {@code BACKEND_EXHAUSTED} classification on this profile, or on another profile sharing its
* {@code credentialId}, is still on cooldown.
*
* <p>Checked before {@link #enforceMaxLoad}, and for the same reason that check exists: an
* explicit profile bypasses the placement policy's filtering entirely, so without this the cap
* (there, quarantine here) would be dead config on the very path the charter calls normal.
* Deliberately no fallback to another profile, matching {@link #enforceMaxLoad}'s own reasoning —
* the caller named this profile for a cost/model reason.
*
* @throws PlacementException naming the profile, its credential, and the remaining cooldown
*/
private void enforceNotQuarantined(String profile) {
String credentialId = credentialIdFor(profile);
quarantine.remainingSeconds(credentialId).ifPresent(remaining -> {
throw new PlacementException("worker profile '" + profile + "' is quarantined "
+ "(credential '" + credentialId + "' exhausted; ~" + remaining
+ "s remaining) — refusing spawn");
});
}
/** {@code profile}'s credential group (CB-578 stage B), or the profile's own name if unconfigured. */
private String credentialIdFor(String profile) {
BridgedConfig.Profile cfg = profiles0().get(profile);
return cfg == null ? profile : cfg.effectiveCredentialId();
}
/** The subset of {@code candidates} whose credential is currently quarantined (CB-578 stage B). */
private Set<String> quarantinedProfiles(List<PlacementCandidate> candidates) {
return candidates.stream()
.map(PlacementCandidate::profile)
.filter(p -> quarantine.isQuarantined(credentialIdFor(p)))
.collect(Collectors.toSet());
}
private void enforceMaxLoad(String profile) {
// Absent config, or a config whose maxLoad normalized to null (ABSENT ⇒ unlimited at load),
// means no cap — never cap what wasn't configured. Note "non-positive ⇒ unlimited" was true
// until CB-585: an explicit `maxLoad: 0` now survives as 0 and is a real cap of zero, so the
// check below refuses every spawn on that profile, and a negative value is refused at config
// load rather than normalized away.
BridgedConfig.Profile cfg = profiles0().get(profile);
Integer cap = (cfg == null) ? null : cfg.maxLoad();
if (cap == null) {
return;
}
int live = liveCount.apply(profile);
if (live >= cap) {
throw new PlacementException("worker profile '" + profile + "' is at maxLoad: " + live
+ " live >= " + cap + " cap; refusing spawn — no fallback to another profile");
}
}
/**
* The profile names {@code role} may be placed on, in definition order.
*
* <p>An empty or absent pool means "unconstrained", not "nothing allowed": a config that declares
* no pool for a role must keep spawning, so it falls back to every configured profile. Names in a
* pool that no adapter declares are dropped here rather than thrown — config load already rejects
* a pool entry with no profile, so a survivor is a profile this particular composite does not own.
*/
private List<String> poolFor(MemberRole role) {
Map<String, BridgedConfig.Profile> configured = profiles0();
BridgedConfig.Fleet f = fleet.get();
List<String> pool = (f == null) ? List.of() : f.profilesFor(role);
List<String> known = pool.stream().filter(configured::containsKey).toList();
return known.isEmpty() ? List.copyOf(configured.keySet()) : known;
}
/** The profile an unqualified spawn for {@code role} falls back to under {@code fixed} placement. */
private String defaultProfileFor(MemberRole role) {
List<String> pool = poolFor(role);
return pool.isEmpty() ? defaultProfile : pool.getFirst();
}
/** Build the candidate list from {@code role}'s pool, in definition order. */
private List<PlacementCandidate> candidates(MemberRole role) {
List<PlacementCandidate> out = new ArrayList<>();
for (String name : poolFor(role)) {
BridgedConfig.Profile w = profiles0().get(name);
if (w != null) {
out.add(new PlacementCandidate(name, null, w.weight(), w.maxLoad()));
}
}
return out;
}
@Override
public String effectiveCwd(SpawnRequest req) {
return route(req.profileName()).effectiveCwd(req);
}
@Override
public List<String> parityOverlay(String profileName) {
return route(profileName).parityOverlay(profileName);
}
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.remove(id);
if (d == null) {
log.debug("stop({}) — no recorded owner, routing to the first adapter (pane-addressed)", id);
d = delegates.getFirst();
}
d.stop(id);
}
@Override
public boolean clearContext(String id) {
HerdrPeerLauncher delegate = spawnedBy.get(id);
if (delegate == null) {
log.debug("clearContext({}) ignored — no recorded owning adapter", id);
return false;
}
return delegate.clearContext(id);
}
@Override
public Set<String> profiles() {
return byProfile.keySet();
}
@Override
public String defaultProfile() {
return defaultProfile;
}
/**
* {@inheritDoc}
*
* <p>Routes to the specific delegate {@code profileName} resolves to, not the fleet-wide
* union {@link #capabilities()} returns — the whole reason this method exists (CB-584): in a
* mixed fleet, one adapter's capability must never be read as every profile's.
*/
@Override
public Set<Capability> capabilitiesFor(String profileName) {
return route(profileName).capabilities();
}
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
@Override
public List<Agent> list() {
Map<String, Agent> byPane = new LinkedHashMap<>();
for (HerdrPeerLauncher d : delegates) {
for (Agent a : d.list()) {
if (a.paneId() != null) {
byPane.putIfAbsent(a.paneId(), a);
}
}
}
return List.copyOf(byPane.values());
}
@Override
public int reapOrphanWorkers() {
int reaped = 0;
for (HerdrPeerLauncher d : delegates) {
reaped += d.reapOrphanWorkers();
}
return reaped;
}
/** The union of every adapter's capabilities — a capability any adapter offers, the fleet offers. */
@Override
public Set<Capability> capabilities() {
EnumSet<Capability> caps = EnumSet.noneOf(Capability.class);
for (HerdrPeerLauncher d : delegates) {
caps.addAll(d.capabilities());
}
return Set.copyOf(caps);
}
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,512 @@
package dev.ltms.bridged.member;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.node.ObjectNode;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.SpawnRequest;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>opencode</strong> — an open-source,
* provider-agnostic terminal coding agent. Its whole reason for existing is to prove the
* {@code PeerLauncher} SPI is genuinely provider-neutral: opencode shares none of Claude Code's
* private launch seams, yet reuses every line of shared transport in the base (tab/pane placement,
* the CB-306 readiness gate, unique naming + CB-117 reap, teardown, listing, cwd).
*
* <p>The divergences from {@link ClaudeCodeLauncher}, all confined to {@link #buildLaunch}:
* <ul>
* <li><strong>No subscription boundary.</strong> opencode carries no {@code ANTHROPIC_BASE_URL}
* and there is no {@link dev.ltms.bridged.guard.SubscriptionGuard} — the guard is a
* Claude-private concern, not part of the SPI. opencode reads the operator's own provider
* credentials from its global {@code auth.json}; the bridge injects none.</li>
* <li><strong>File-based MCP mount + instructions.</strong> opencode has no inline
* {@code --mcp-config}/{@code --append-system-prompt}. Instead the bridge writes an ephemeral
* {@code opencode.json} that declares the bridge as a {@code remote} MCP server and lists a
* member-charter file under {@code instructions}, then points the worker at it with
* {@code OPENCODE_CONFIG}. This is the one place the launcher touches disk — Claude never did.</li>
* <li><strong>Model as a flag.</strong> the {@code provider/model} selector is passed as
* {@code -m}, not an env var.</li>
* <li><strong>{@code opencode} name prefix</strong> so reap matches {@code opencode-*} panes and
* never another adapter's.</li>
* </ul>
*/
public final class OpenCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "opencode";
/** Writer for the generated {@code opencode.json}. */
private static final ObjectMapper JSON = new ObjectMapper();
/** Root under which per-spawn opencode config dirs are created (injectable for tests). */
private final Path configRoot;
/**
* Session discovery against opencode's on-disk storage ({@link OpenCodeSessionDiscovery}) —
* the one seam that knows opencode's private session-file layout. Its root is injectable for
* tests so they never touch the operator's real {@code ~/.local/share/opencode}.
*/
private final OpenCodeSessionDiscovery discovery;
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so it
* matches the legacy non-blocking spawn semantics. Config dirs are created under the JVM temp dir.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300),
defaultConfigRoot(), defaultDiscoveryRoot());
}
/**
* Production constructor with the spawn-ready gate enabled. Polls {@code agents.status()} until
* the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, spawnReadyPollMs, null);
}
/**
* Production constructor carrying the fleet-wide tab-label template (CB-557). The template comes
* from {@code fleet.tabLabel}, which a profile cannot know because it names the member's
* <em>role</em>; a profile may still override it with its own {@code tabLabel}.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<BridgedConfig.Fleet> fleet) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot(), defaultDiscoveryRoot(), fleet);
}
/**
* Production constructor, plus the CB-596 {@code memberCredentials} policy supplier.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot(), defaultDiscoveryRoot(), fleet, memberCredentials);
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply a
* fake clock ({@code nowMillis}), poll-loop wait ({@code sleeper}), and a temp {@code configRoot}
* they can inspect the generated {@code opencode.json}/charter under.
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (encodes the poll interval; never called when the
* gate is disabled)
* @param configRoot existing directory under which per-spawn config dirs are created
* @param discoveryRoot opencode's on-disk storage root to scan for session records
* (injectable for tests; opencode's layout is matched at
* {@link OpenCodeSessionDiscovery})
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Path configRoot, Path discoveryRoot) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
nowMillis, sleeper, configRoot, discoveryRoot, null);
}
/**
* Full testability constructor, plus the fleet-wide tab-label template (CB-557).
*
* @param fleet live fleet config, read once for each spawn
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Path configRoot, Path discoveryRoot,
Supplier<BridgedConfig.Fleet> fleet) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet);
this.configRoot = configRoot;
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
}
/**
* Full testability constructor, plus the CB-596 {@code memberCredentials} policy supplier.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Path configRoot, Path discoveryRoot,
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials);
this.configRoot = configRoot;
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
}
private static Path defaultConfigRoot() {
return Path.of(System.getProperty("java.io.tmpdir"));
}
/** The default opencode storage root: {@code ~/.local/share/opencode} (the XDG data dir). */
private static Path defaultDiscoveryRoot() {
return Path.of(System.getProperty("user.home"), ".local", "share", "opencode");
}
/**
* {@inheritDoc}
*
* <p>Builds the opencode launch: no {@code ANTHROPIC_*} and no guard (opencode reads its own
* provider credentials); when the profile mounts the bridge MCP or has a member charter,
* generate an ephemeral {@code opencode.json} (remote MCP server + member-charter instructions)
* and point the worker at it via {@code OPENCODE_CONFIG}; carry the parity-neutral git-forge
* grant; and select the model with {@code -m}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
Map<String, String> workerEnv = baseEnv(cfg);
// A config file is needed for the bridge MCP mount, a member charter, or a pinned endpoint (CB-508).
if (cfg.hasMcp() || spec.charter() != null || hasCustomProvider(cfg)) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg, spec.charter()).toString());
}
applyGitToken(workerEnv, cfg);
List<String> argv = argvWithResume(argvWithModel(argvWithAuto(cfg), cfg), spec.resumeSessionId());
return new Launch(workerEnv, argvWithAgent(argv, spec));
}
/**
* The launch argv plus, when the role has an agent-definition file under the worker's cwd,
* opencode's {@code --agent <role>} flag (CB-617). A role with no such file gets nothing added —
* the member must still spawn.
*/
private List<String> argvWithAgent(List<String> argv, LaunchSpec spec) {
Path agentFile = agentDefinitionFile(spec.cwd(), spec.role(), ".opencode", "agent");
if (agentFile == null) {
return argv;
}
List<String> withAgent = mutableArgv(argv);
withAgent.add("--agent");
withAgent.add(spec.role().wireName());
return withAgent;
}
/**
* True when this profile pins its own OpenAI-compatible endpoint (CB-508) rather than using
* whatever provider opencode resolves by default.
*
* <p>Note this reuses {@code baseUrl}, the same field the Claude adapter injects as
* {@code ANTHROPIC_BASE_URL} — but it does <em>not</em> go through {@code SubscriptionGuard}.
* That asymmetry is deliberate and safe: the guard exists to stop a worker borrowing the
* primary's Anthropic subscription, and an opencode process has no Anthropic credential path
* at all. Pointing it at a local vLLM cannot leak the subscription.
*/
private static boolean hasCustomProvider(BridgedConfig.Profile cfg) {
return cfg.baseUrl() != null && !cfg.baseUrl().isBlank();
}
/**
* The launch argv plus the unconditional {@code --auto} flag, which auto-approves the
* permissions opencode does not explicitly deny. It is unconditional, not a preference: a
* spawned peer has no human at its pane — the bridge spawned it — so one that stops at an
* approval prompt is a wedged agent, indistinguishable from a legitimate mid-turn wait and
* unable to end its turn with {@code bridge_reply}. opencode's own help calls this
* "dangerous!", but the blast radius here is already bounded by design: a worker runs in its
* own git worktree on its own branch, is off-subscription, and cannot merge — the lead is the
* gate.
*/
private List<String> argvWithAuto(BridgedConfig.Profile cfg) {
List<String> argv = mutableArgv(cfg.argv());
argv.add("--auto");
return argv;
}
/**
* The launch argv plus, on a resumed spawn, opencode's {@code -s <id>} flag to continue a prior
* conversation by its session id. {@code -s, --session <id>} resumes an existing session; on a
* fresh spawn (no resume target) no flag is added, letting opencode start a brand-new session.
* The id comes from the base launch spec.
*/
private List<String> argvWithResume(List<String> argv, String id) {
if (id == null || id.isBlank()) {
return argv;
}
List<String> withResume = mutableArgv(argv);
withResume.add("-s");
withResume.add(id);
return withResume;
}
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
private List<String> argvWithModel(List<String> argv, BridgedConfig.Profile cfg) {
if (cfg.model() != null && !cfg.model().isBlank()) {
argv.add("-m");
argv.add(cfg.model());
}
return argv;
}
/**
* Write an ephemeral {@code opencode.json} (and the member-charter file it references) into a
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
*/
private Path writeConfig(BridgedConfig.Profile cfg, String charterText) {
try {
Path dir = Files.createTempDirectory(configRoot, "bridged-opencode-");
dir.toFile().deleteOnExit();
ObjectNode root = JSON.createObjectNode();
root.put("$schema", "https://opencode.ai/config.json");
// CB-523, opencode side: a worker that runs out of context dies mid-turn, and its reply
// — the entire point of the turn — is lost with it. Auto-compaction is therefore not an
// operator preference for a bridged worker, it is a condition of the turn contract.
//
// Stated deliberately even though it is redundant today: OPENCODE_CONFIG is MERGED over
// ~/.config/opencode/config.json rather than replacing it, so a worker already inherits
// an `auto: true` set at home. We do not want that inheritance to be what the guarantee
// rests on — the home file is outside this repo, differs per machine, and is not ours.
//
// Know the cost before removing it: this key WINS over the home config (verified — an
// OPENCODE_CONFIG value overrides the home value, it does not defer to it), so an
// operator who sets `compaction.auto: false` at home cannot turn it off for bridged
// workers. That is the intended trade for peers we spawn and whose turns we must land;
// if per-profile control is ever wanted, add a profile knob rather than dropping this.
root.putObject("compaction").put("auto", true);
if (charterText != null) {
Path charter = dir.resolve("member-charter.md");
Files.writeString(charter, charterText);
charter.toFile().deleteOnExit();
root.putArray("instructions").add(charter.toAbsolutePath().toString());
}
if (cfg.hasMcp()) {
ObjectNode bridge = root.putObject("mcp").putObject("bridge");
bridge.put("type", "remote");
bridge.put("url", cfg.mcpUrl());
bridge.put("enabled", true);
}
if (hasCustomProvider(cfg)) {
addCustomProvider(root, cfg);
}
Path cfgFile = dir.resolve("opencode.json");
// Built with Jackson rather than string concatenation: the provider block is nested and
// carries operator-supplied values (URL, model id, api key), so escaping must be real.
Files.writeString(cfgFile, JSON.writerWithDefaultPrettyPrinter().writeValueAsString(root));
cfgFile.toFile().deleteOnExit();
return cfgFile;
} catch (IOException e) {
throw new UncheckedIOException(
"cannot write opencode config for profile " + cfg.profile(), e);
}
}
/**
* Declare a custom OpenAI-compatible provider so the worker talks to a pinned endpoint (a local
* vLLM, say) instead of opencode's default gateway (CB-508).
*
* <p>The provider id comes from the {@code provider/model} selector in {@code model:}, so one
* field drives both the declaration and the {@code -m} flag and they cannot drift apart.
*/
private void addCustomProvider(ObjectNode root, BridgedConfig.Profile cfg) {
String[] parts = splitModelSelector(cfg);
String providerId = parts[0];
String modelId = parts[1];
ObjectNode provider = root.putObject("provider").putObject(providerId);
provider.put("npm", "@ai-sdk/openai-compatible");
provider.put("name", providerId + " (bridged)");
ObjectNode options = provider.putObject("options");
options.put("baseURL", openAiBaseUrl(cfg.baseUrl()));
// vLLM and friends usually ignore the key, but the AI SDK still requires a non-empty one.
String token = resolveEnv(cfg.tokenEnv());
options.put("apiKey", (token == null || token.isBlank()) ? "bridged-local-noauth" : token);
provider.putObject("models").putObject(modelId).put("name", modelId);
}
/**
* Split {@code model:} into its {@code provider} and {@code model} halves. A pinned endpoint
* needs both, so a bare model name is rejected loudly rather than silently falling back to the
* default gateway — a worker quietly talking to the wrong endpoint is the failure this avoids.
*/
private static String[] splitModelSelector(BridgedConfig.Profile cfg) {
String model = cfg.model();
int slash = model == null ? -1 : model.indexOf('/');
if (model == null || model.isBlank() || slash <= 0 || slash == model.length() - 1) {
throw new IllegalArgumentException(
"profile " + cfg.profile() + " sets baseUrl (a pinned opencode endpoint) so"
+ " model: must be \"<provider>/<model>\", e.g."
+ " \"local-vllm/deepseek-v4-flash\"; got "
+ (model == null ? "null" : '"' + model + '"'));
}
return new String[]{model.substring(0, slash), model.substring(slash + 1)};
}
/**
* The OpenAI-compatible base URL for {@code baseUrl}. A bare {@code host:port} gets {@code /v1}
* appended (where these servers put the API); a URL that already carries a path is taken as-is,
* so an endpoint mounted somewhere unusual is still reachable.
*/
private static String openAiBaseUrl(String baseUrl) {
String trimmed = baseUrl.trim();
while (trimmed.endsWith("/")) {
trimmed = trimmed.substring(0, trimmed.length() - 1);
}
int schemeEnd = trimmed.indexOf("://");
String afterScheme = schemeEnd < 0 ? trimmed : trimmed.substring(schemeEnd + 3);
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
}
/** Add lazy on-disk session discovery to the base handle. */
@Override
public PeerHandle spawn(SpawnRequest req) {
PeerHandle inner = super.spawn(req);
return new SessionAwareHandle(inner, discovery, effectiveCwd(req));
}
/**
* A {@link PeerHandle} that delegates everything to the base's worker handle but resolves
* {@link #agentSessionId()} lazily through opencode session discovery. Delegate-only, so the
* base's id/terminalId/profile semantics (CB-519's host-unique routing key, herdr coordinates)
* are untouched — only the opencode-specific identity answer is added. {@code sessionName()}
* stays null: opencode has no display-name seam, so the logical name lives only in the bridge's
* roster (see the SESSION_NAME capability).
*/
private static final class SessionAwareHandle implements PeerHandle {
private final PeerHandle delegate;
private final OpenCodeSessionDiscovery discovery;
private final String cwd;
SessionAwareHandle(PeerHandle delegate, OpenCodeSessionDiscovery discovery, String cwd) {
this.delegate = delegate;
this.discovery = discovery;
this.cwd = cwd;
}
@Override
public String id() {
return delegate.id();
}
@Override
public String terminalId() {
return delegate.terminalId();
}
@Override
public String profile() {
return delegate.profile();
}
@Override
public String sessionName() {
return delegate.sessionName();
}
@Override
public String agentSessionId() {
// Lazy + retried, never a spawn-time blocker: opencode writes the session record only
// when the session is first persisted, so null here is the correct interim answer and
// the caller re-calls later (each call re-scans, picking up a record that has since
// appeared).
return discovery.sessionIdForDirectory(cwd);
}
@Override
public CharterReceipt charterReceipt() {
return delegate.charterReceipt();
}
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE,
Capability.ORPHAN_REAP, Capability.SESSION_RESUME);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
// Deliberately NOT SESSION_NAME: opencode has no display-name flag, so the bridge's logical
// name can't surface in the peer's own UI — declaring the capability would hide that
// asymmetry rather than make it honest. For opencode the name lives only in the bridge's
// roster (see PeerHandle.sessionName() returning null).
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Profile::hasGitToken);
}
// --- CB-117 reap predicate (opencode prefix), kept for direct unit testing -----------------
/**
* Whether {@code name} is an opencode bridge worker started by a <em>different</em> process than
* {@code currentNonce}. A thin {@code opencode}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -0,0 +1,122 @@
package dev.ltms.bridged.member;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.stream.Stream;
/**
* Resolves the opencode session id for a bridged worker from opencode's on-disk storage — the
* only place this adapter touches opencode's private layout, and deliberately the <em>only</em>
* class that does.
*
* <p><strong>Why this is isolated behind one seam.</strong> The layout is version-coupled and not a
* stable contract: opencode writes one JSON file per session under
* {@code <storageRoot>/session/<projectID>/<ses_*.json>}, and each record carries a
* {@code "version"} field (e.g. {@code "1.1.31"}), so the exact directory shape, file naming, and
* field names can move between opencode releases. opencode also ships a headless HTTP server that
* may supersede file scanning entirely. Everything this adapter knows about that private storage —
* its shape, naming, and field names — lives here, so a layout change, or a switch to the HTTP
* server, changes exactly one class and nothing in {@link OpenCodeLauncher}.
*
* <p>The determinism that makes this useful is structural, not a guess: every bridged worker runs
* in its own unique git worktree, so the record's {@code directory} (its project root) equals the
* worker's cwd identifies <em>its</em> session unambiguously. We match on {@code directory} rather
* than diffing {@code opencode session list} before/after — that races under concurrent spawns, and
* the CLI listing does not even show the directory.
*
* <p>All reads are best-effort and never throw: a missing or unreadable storage root, a record that
* fails to parse, or a directory with no record yet all yield {@code null}, and the caller (the
* session handle) treats that as "identity not resolved yet" and retries later.
*/
final class OpenCodeSessionDiscovery {
private final Path storageRoot; // e.g. ~/.local/share/opencode (injectable for tests)
private final ObjectMapper json;
OpenCodeSessionDiscovery(Path storageRoot) {
this.storageRoot = storageRoot;
this.json = new ObjectMapper();
}
/**
* The opencode session id whose record references {@code directory} (the worker's cwd), or
* {@code null} when no record matches yet. When several records share the directory — e.g.
* repeated spawns into the same worktree — the <em>most recently modified</em> one wins: it is
* the session the pane most likely corresponds to.
*
* <p>Never throws: a missing {@code storageRoot}, an unreadable/malformed record, or a
* directory that has not been persisted yet all resolve to {@code null} rather than failing a
* spawn. A bridged worker's session record is written lazily (when the session is first
* persisted), so {@code null} here is the normal answer right after the pane is ready, and the
* caller retries later.
*
* @param directory the worker's cwd, as resolved for this spawn
* @return the matching session id, or {@code null} if none is known yet
*/
String sessionIdForDirectory(String directory) {
if (directory == null || directory.isBlank()) {
return null;
}
Path sessionRoot = storageRoot.resolve("session");
if (!Files.isDirectory(sessionRoot)) {
return null;
}
String best = null;
long bestMtime = Long.MIN_VALUE;
try (Stream<Path> projectDirs = Files.list(sessionRoot)) {
for (Path projectDir : projectDirs.filter(Files::isDirectory).toList()) {
try (Stream<Path> records = Files.list(projectDir)) {
for (Path record : records.toList()) {
String id = matchId(record, directory);
if (id == null) {
continue;
}
long mtime = lastModifiedEpochMillis(record);
if (mtime > bestMtime) {
bestMtime = mtime;
best = id;
}
}
} catch (IOException ignored) {
// one project dir unreadable — skip it; another may still match
}
}
} catch (IOException ignored) {
// storage root vanished or became unreadable — "no session known yet"
return null;
}
return best;
}
/**
* The record's session id when it references {@code directory}, else {@code null}. A record
* that is not JSON, lacks {@code id}/{@code directory}, or points at a different directory is
* simply not our session; a malformed one is skipped, never fatal.
*/
private String matchId(Path record, String directory) {
try {
JsonNode node = json.readTree(record.toFile());
JsonNode id = node == null ? null : node.get("id");
JsonNode dir = node == null ? null : node.get("directory");
if (id == null || dir == null || !directory.equals(dir.asText())) {
return null;
}
return id.asText();
} catch (IOException e) {
return null;
}
}
/** The record's last-modified epoch ms, or {@code Long.MIN_VALUE} if unreadable (never wins). */
private static long lastModifiedEpochMillis(Path record) {
try {
return Files.getLastModifiedTime(record).toMillis();
} catch (IOException e) {
return Long.MIN_VALUE;
}
}
}
@@ -1,8 +1,8 @@
package dev.ltms.fleet.metrics;
package dev.ltms.bridged.metrics;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.MemberSession;
import java.util.LinkedHashMap;
import java.util.Map;
@@ -13,33 +13,33 @@ import java.util.Map;
*
* <p>The set is deliberately small: each series maps to a failure mode this project has actually
* hit, not to whatever was easy to count. The two worth watching in practice are
* {@code fleet_sends_total{outcome="completion_fallback"}} — a rising share means turn detection
* {@code bridged_sends_total{outcome="completion_fallback"}} — a rising share means turn detection
* is degrading, the CB-115/116/118 failure family — and
* {@code fleet_push_nudges_total{outcome="exhausted"}}, which means the primary stopped draining
* {@code bridged_push_nudges_total{outcome="exhausted"}}, which means the primary stopped draining
* its inbox and CB-307's active push gave up.
*/
public final class FleetMetrics {
public final class BridgedMetrics {
/** Counter: delegated sends by terminal outcome. */
public static final String SENDS = "fleet_sends_total";
public static final String SENDS = "bridged_sends_total";
/** Counter: worker replies by the path that carried them (rendezvous vs stranded-to-inbox). */
public static final String REPLIES = "fleet_replies_total";
public static final String REPLIES = "bridged_replies_total";
/** Counter: push-loop nudges to the primary, by outcome. */
public static final String PUSH_NUDGES = "fleet_push_nudges_total";
public static final String PUSH_NUDGES = "bridged_push_nudges_total";
/** Counter: idle-lead heartbeat nudges to the lead, by outcome (CB-551). */
public static final String HEARTBEAT_NUDGES = "fleet_lead_heartbeat_nudges_total";
public static final String HEARTBEAT_NUDGES = "bridged_lead_heartbeat_nudges_total";
/** Counter: spawn attempts by peer kind and outcome. */
public static final String SPAWNS = "fleet_spawns_total";
public static final String SPAWNS = "bridged_spawns_total";
/** Counter: herdr socket calls by method and outcome. */
public static final String HERDR_CALLS = "fleet_herdr_calls_total";
public static final String HERDR_CALLS = "bridged_herdr_calls_total";
/** Counter: rejected requests by reason (CB-501). */
public static final String AUTH_FAILURES = "fleet_auth_failures_total";
public static final String AUTH_FAILURES = "bridged_auth_failures_total";
/** Gauge: session census by lifecycle state. */
public static final String SESSIONS = "fleet_sessions";
public static final String SESSIONS = "bridged_sessions";
/** Gauge: undrained replies held per target. */
public static final String INBOX_DEPTH = "fleet_inbox_depth";
public static final String INBOX_DEPTH = "bridged_inbox_depth";
private FleetMetrics() {
private BridgedMetrics() {
}
/**
@@ -56,13 +56,10 @@ public final class FleetMetrics {
m.describe(REPLIES, "counter",
"Worker replies by delivery path (rendezvous=resolved an open send, inbox=stranded and held).");
m.describe(PUSH_NUDGES, "counter",
"CB-307 push-loop nudges to the primary (sent|exhausted). fleetd #365: \"sent\" means "
+ "the herdr paste-and-submit call succeeded, not that the pane read it — this "
+ "layer has no read-receipt concept.");
"CB-307 push-loop nudges to the primary (delivered|exhausted).");
m.describe(HEARTBEAT_NUDGES, "counter",
"CB-551 idle-lead heartbeat nudges (sent|failed|exhausted). Quiet-cap exhaustion "
+ "means the lead idled with nothing pending and was told to stand down. "
+ "fleetd #365: \"sent\" means the herdr call succeeded, not that the lead read it.");
"CB-551 idle-lead heartbeat nudges (delivered|failed|exhausted). Quiet-cap exhaustion "
+ "means the lead idled with nothing pending and was told to stand down.");
m.describe(SPAWNS, "counter",
"Worker spawn attempts by peer kind and outcome (ready|timeout|guard_rejected).");
m.describe(HERDR_CALLS, "counter",
@@ -1,4 +1,4 @@
package dev.ltms.fleet.metrics;
package dev.ltms.bridged.metrics;
import java.util.Map;
import java.util.NavigableMap;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import com.rabbitmq.client.AMQP;
import com.rabbitmq.client.Channel;
@@ -8,7 +8,6 @@ import com.rabbitmq.client.DeliverCallback;
import com.rabbitmq.client.Recoverable;
import com.rabbitmq.client.RecoveryListener;
import com.rabbitmq.client.Return;
import com.rabbitmq.client.impl.DefaultExceptionHandler;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -23,7 +22,6 @@ import java.util.concurrent.ConcurrentSkipListMap;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
import java.util.concurrent.atomic.AtomicReference;
/**
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
@@ -31,9 +29,8 @@ import java.util.concurrent.atomic.AtomicReference;
*
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target has a durable
* queue {@code agent.<target>.inbox}. The gateway that owns the target starts a manual-ack consumer
* ({@link #own}) that pulls persistent messages, up to its prefetch window, off that queue into an
* in-memory <em>held</em> map (keyed by {@code msgId}) but does <em>not</em> ack them.
* {@link #peek} returns that snapshot;
* ({@link #own}) that pulls persistent messages off that queue into an in-memory <em>held</em> map
* (keyed by {@code msgId}) but does <em>not</em> ack them. {@link #peek} returns that snapshot;
* {@link #ack} acks the broker delivery-tag and drops the entry. Because messages stay unacked until
* the owning gateway actually drains them, a crash (or a {@code java -jar} bounce) before caller-ack
* leaves them on the broker — it redelivers on reconnect. That is the durability the in-memory
@@ -95,23 +92,8 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
private final Channel channel;
/** All channel operations (publish/declare/ack/cancel) serialize on this — a Channel is not thread-safe. */
private final Object channelLock = new Object();
/**
* target → (msgId → held delivery). Per-target map is guarded by synchronizing on itself.
*
* <p><strong>CB-318 tombstone.</strong> The value {@link #RELEASED} is a reserved sentinel: it
* marks a target whose {@link #release} has already run, so {@link #deliverCallback} can tell a
* delivery landing after release() apart from a fresh target it has never seen. See both methods'
* javadoc for why a plain {@code held.remove(target)} is not enough.
*/
/** target → (msgId → held delivery). Per-target map is guarded by synchronizing on itself. */
private final ConcurrentHashMap<String, LinkedHashMap<String, Held>> held = new ConcurrentHashMap<>();
/**
* CB-318 sentinel stored in {@link #held} for a target whose {@link #release} has already run.
* Never mutated — every read site compares it by reference ({@code ==}) before touching it as a
* map, because it is a single object shared across every released target and calling a mutator on
* it would corrupt state for all of them.
*/
private static final LinkedHashMap<String, Held> RELEASED = new LinkedHashMap<>();
/** Targets whose queue is declared and consumer is running, mapped to their broker consumer tag. */
private final ConcurrentHashMap<String, String> consumerTags = new ConcurrentHashMap<>();
@@ -154,22 +136,17 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
/** As {@link #open(String)}, with an explicit consumer prefetch (CB-527: caps the held backlog per target). */
public static AmqpReplyInbox open(String uri, int prefetch) {
try {
return new AmqpReplyInbox(connectionFactory(uri).newConnection(AmqpConnectionFailureLogger.REPLY_INBOX), prefetch);
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"), prefetch);
} catch (Exception e) {
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
}
}
static ConnectionFactory connectionFactory(String uri) throws Exception {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
factory.setExceptionHandler(new AmqpConnectionFailureLogger(AmqpConnectionFailureLogger.REPLY_INBOX, log));
return factory;
}
/** Wrap an already-open connection with {@link #DEFAULT_PREFETCH} (injection seam for the contract test). */
AmqpReplyInbox(Connection connection) {
this(connection, DEFAULT_PREFETCH);
@@ -223,12 +200,6 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
String tag = channel.basicConsume(queue, false, deliverCallback(target), _ -> { });
consumerTags.put(target, tag);
// CB-318: drop a stale RELEASED tombstone from a prior ownership of this same target
// string, so a delivery under this fresh consumer is held normally instead of being
// nacked forever by deliverCallback's RELEASED check. Safe to do here, still under
// channelLock: no delivery for the consumer tag just registered above can reach
// deliverCallback before this basicConsume call returns.
held.remove(target, RELEASED);
log.debug("AMQP inbox owns queue {} for target {}", queue, target);
} catch (IOException e) {
throw new IllegalStateException("cannot own queue " + queue, e);
@@ -236,103 +207,18 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
}
/**
* Release ownership of {@code target}: cancel its consumer, then nack-with-requeue every
* delivery still held for it instead of just dropping the local record.
*
* <p><strong>Cancelling a consumer does not requeue its in-flight deliveries.</strong> In AMQP,
* a delivery that was pushed to a consumer stays unacked, attached to the still-open
* {@link #channel}, until that channel or the connection closes — {@code basicCancel} alone does
* neither. So before this method existed with a requeue step, it dropped {@link #held}'s entries
* for {@code target} while the broker still considered them outstanding: never acked, never
* nacked, never requeued, and no longer reachable by {@link #peek} — permanently invisible. This
* is unlike {@link RecoveryListener#handleRecovery(Recoverable)} and {@link #close()}, whose bare {@code held.clear()} is
* correct because each has already made the broker requeue (a real connection drop, or
* {@code channel.close()} respectively) before clearing local state.
*
* <p><strong>Order: cancel first, then nack.</strong> A delivery tag stays valid for
* {@code basicNack} on this channel regardless of whether its consumer is still attached — only
* a channel/connection close invalidates it — so cancelling {@code target}'s consumer first does
* not risk the tags. Doing it the other way round does: nacking a delivery with {@code requeue}
* while its consumer is still active hands the message straight back to that <em>same</em>
* consumer the instant a prefetch slot frees up (confirmed against a real broker — see
* {@code AmqpReplyInboxContractTest.releaseCancelsConsumerAndRequeuesHeldDeliveryForRecovery}),
* which races this method's own {@code held.remove(target)}: the redelivery can land after the
* clear and leave a stale entry behind, so {@link #peek} is no longer reliably empty right after
* {@link #release}. Cancelling first closes that consumer, so the requeued message goes back to
* the queue for whichever consumer picks it up next (a later {@link #own}), not this one.
*
* <p><strong>Failure of the requeue is best-effort, not fatal.</strong> {@link #release} runs
* during teardown ({@code Fleetd} calls it right after {@code MessageService.abandon}), and a
* throw here would abort cleanups the caller depends on — the same argument fleetd #293 settled
* for {@code HerdrPeerLauncher.stop()}'s tab-close step. So a failed {@code basicNack} is logged
* at WARN, naming the target and delivery tag that leaked, and release proceeds; the delivery
* stays unacked on the broker rather than being silently dropped, so it is still recoverable by a
* later connection drop even though this release did not manage to requeue it immediately. A
* failed {@code basicCancel} still throws, unchanged from before this fix — that failure means
* the consumer may still be attached, so best-effort requeue is not attempted underneath it.
*
* <p><strong>CB-318: {@code held.remove(target)} alone leaves a second window open.</strong> The
* bullet above already explains why cancelling first does not save a tag from going stale — but
* that only accounts for a delivery landing before this method starts touching {@link #held}.
* {@code basicCancel} stops <em>new</em> dispatches; it does not flush one already handed to the
* consumer work pool. So a delivery can still land on that pool's thread and reach
* {@link #deliverCallback} at any point during, or after, this method's body — and a plain
* {@code held.remove(target)} does nothing to stop it: {@code deliverCallback}'s
* {@code computeIfAbsent} finds the key gone and happily creates a brand-new map under it, which
* this method — already past its {@code remove} — never looks at again. That entry then sits
* delivered-but-unacked on {@link #channel} until the whole inbox closes: never requeued, never
* redelivered, and {@link #peek} is never called again for a target nothing owns any more.
*
* <p>The fix is {@link #held}{@code .compute(target, ...)} instead of {@code remove}: it takes
* whatever was held (to nack, same as before) and, in the same atomic step, leaves the
* {@link #RELEASED} tombstone behind instead of an absent key. {@code computeIfAbsent} and
* {@code compute} calls for the same key are mutually exclusive in {@link ConcurrentHashMap} —
* whichever of this call and a concurrent {@code deliverCallback} runs first is fully visible to
* the other, with no gap between them. So a delivery that loses the race sees a real map here and
* gets nacked by the loop below, same as always; a delivery that wins the race (runs first) is
* itself nacked by that same loop, once it settles into {@code held}. A delivery that arrives once
* this method has stored {@link #RELEASED} finds it via {@code computeIfAbsent} and refuses itself
* — see {@link #deliverCallback}. Either way nothing is silently retained forever, satisfying the
* ticket's invariant against dropping a message. This closes the window rather than merely
* narrowing it — correctness does not depend on how much time elapses between the swap and this
* method returning.
*/
@Override
public void release(String target) {
synchronized (channelLock) {
String tag = consumerTags.remove(target);
if (tag != null) {
try {
channel.basicCancel(tag);
} catch (IOException e) {
throw new IllegalStateException("cannot cancel consumer for " + target, e);
}
held.remove(target); // stale delivery tags must not survive release
if (tag == null) {
return;
}
AtomicReference<LinkedHashMap<String, Held>> previouslyHeld = new AtomicReference<>();
held.compute(target, (_, v) -> {
previouslyHeld.set(v);
return RELEASED;
});
var perTarget = previouslyHeld.get();
if (perTarget != null && perTarget != RELEASED) {
synchronized (perTarget) {
for (Held h : perTarget.values()) {
try {
channel.basicNack(h.deliveryTag(), false, true); // requeue, don't drop
} catch (IOException | RuntimeException e) {
// Caught broadly (not just IOException) for the same reason #293 catches
// RuntimeException in HerdrPeerLauncher.stop(): best-effort teardown must
// not be guarded only against the expected failure and bare against any
// other. The message stays unacked on the broker either way — not lost,
// just not proactively requeued — until a connection drop frees it.
log.warn("release({}): could not requeue held delivery (msgId={}, tag={})"
+ " back to the broker — it stays unacked until a connection"
+ " drop frees it: {}",
target, h.message().msgId(), h.deliveryTag(), e.getMessage());
}
}
}
try {
channel.basicCancel(tag);
} catch (IOException e) {
throw new IllegalStateException("cannot cancel consumer for " + target, e);
}
}
}
@@ -385,7 +271,7 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
@Override
public List<InboxMessage> peek(String target) {
var perTarget = held.get(target);
if (perTarget == null || perTarget == RELEASED) {
if (perTarget == null) {
return List.of();
}
synchronized (perTarget) {
@@ -394,17 +280,17 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
@Override
public boolean ack(String target, String msgId) {
public void ack(String target, String msgId) {
var perTarget = held.get(target);
if (perTarget == null || perTarget == RELEASED) {
return false;
if (perTarget == null) {
return;
}
Held h;
synchronized (perTarget) {
h = perTarget.remove(msgId);
}
if (h == null) {
return false; // never held (or already acked) — no-op
return; // never held (or already acked) — no-op
}
try {
synchronized (channelLock) {
@@ -418,7 +304,6 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
throw new IllegalStateException("cannot ack reply " + msgId + " on " + queueName(target), e);
}
return true;
}
private DeliverCallback deliverCallback(String target) {
@@ -430,20 +315,6 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
String content = new String(delivery.getBody(), StandardCharsets.UTF_8);
var perTarget = held.computeIfAbsent(target, _ -> new LinkedHashMap<>());
if (perTarget == RELEASED) {
// CB-318: release() already ran for this target and left the RELEASED tombstone in
// held (see release()'s javadoc) — computeIfAbsent() is guaranteed to see it rather
// than recreate a fresh map, because ConcurrentHashMap serializes compute/
// computeIfAbsent calls for the same key against each other. Refuse the delivery
// instead of holding it somewhere release() will never look at again: requeue it, the
// same way release() nacks its own held entries, so a later owner (or a connection
// drop) can still recover it. This does not need channelLock across a broker round
// trip — basicNack, like the duplicate-ack case just below, does not wait for one.
synchronized (channelLock) {
channel.basicNack(tag, false, true);
}
return;
}
boolean duplicate;
synchronized (perTarget) {
if (perTarget.containsKey(msgId)) {
@@ -578,44 +449,3 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
}
}
/**
* Keeps RabbitMQ's forgiving exception behaviour while adding the connection identity that its
* default logger drops. Package-private so both AMQP connections use the same two names.
*/
final class AmqpConnectionFailureLogger extends DefaultExceptionHandler {
static final String REPLY_INBOX = "fleetd-reply-inbox";
static final String LEAD_MAILBOX = "fleetd-lead-mailbox";
private final String connectionName;
private final Logger logger;
AmqpConnectionFailureLogger(String connectionName, Logger logger) {
this.connectionName = connectionName;
this.logger = logger;
}
String connectionName() {
return connectionName;
}
@Override
protected void log(String message, Throwable cause) {
if (isSocketClosedOrConnectionReset(cause)) {
logger.warn("AMQP connection {}: {} (Exception message: {})", connectionName, message, cause.getMessage());
} else {
logger.error("AMQP connection {}: {}", connectionName, message, cause);
}
}
private static boolean isSocketClosedOrConnectionReset(Throwable cause) {
// Deliberate copy of ForgivingExceptionHandler's private static helper; check it on amqp-client upgrades.
if (!(cause instanceof IOException)) {
return false;
}
return "Connection reset".equals(cause.getMessage())
|| "Socket closed".equals(cause.getMessage())
|| "Connection reset by peer".equals(cause.getMessage());
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import java.util.LinkedHashMap;
import java.util.List;
@@ -16,7 +16,7 @@ import java.util.concurrent.ConcurrentHashMap;
* non-broker path stays interchangeable.
*
* <p><strong>This is soft-state, NOT persistence.</strong> Lost on a {@code java -jar} bounce — that
* is correct and consistent with "fleetd stays soft-state." The Stage-2 AMQP adapter replaces this.
* is correct and consistent with "bridged stays soft-state." The Stage-2 AMQP adapter replaces this.
*/
public final class InMemoryReplyInbox implements ReplyInbox {
@@ -59,17 +59,16 @@ public final class InMemoryReplyInbox implements ReplyInbox {
}
@Override
public boolean ack(String target, String msgId) {
public void ack(String target, String msgId) {
if (!owned.contains(target)) {
return false;
return;
}
var perTarget = store.get(target);
if (perTarget == null) {
return false;
}
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
return perTarget.remove(msgId) != null;
if (perTarget != null) {
//noinspection SynchronizationOnLocalVariableOrMethodParameter
synchronized (perTarget) {
perTarget.remove(msgId);
}
}
}
}
@@ -1,12 +1,11 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.session.MemberSession;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -14,13 +13,12 @@ import java.util.ArrayList;
import java.util.List;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* CB-551: an opt-in heartbeat that nudges the single idle lead back to work once it has been
* continuously idle past a quiet period with no open {@code fleet_send} driving it.
* continuously idle past a quiet period with no open {@code bridge_send} driving it.
*
* <p>Why this exists: the fleet is ONE lead + architects + workers, so an idle, stalled lead is a
* single point of failure for the fleet's progress. {@link ReplyPushLoop} nudges the lead only when
@@ -46,15 +44,6 @@ import java.util.function.Supplier;
* loop stands down, so two competing injections never start two turns in the same pane
* (constraint 6).</li>
* </ol>
*
* <p><b>fleetd #609 — context-high notice.</b> Optionally ({@code contextHighNudge}, opt-in like the
* loop itself), a tick that finds the lead's own {@link LeadContextGauge} reading at {@link
* LeadContextGauge.State#HIGH} appends a text notice to whatever nudge it sends, telling the lead to
* consider {@code fleet_handover}. This is text only — it never rolls a pane itself. It fires once per
* HIGH stretch (a latch, cleared only by a later {@code OK} reading — {@code UNKNOWN} neither sets nor
* clears it, since "I could not look" must not be read as "it got better"), and it never spends the
* quiet-nudge budget: an idle, quiet, HIGH-context lead is exactly the case {@link Action#QUIET_DONE}
* would otherwise swallow, and it is the one case most worth interrupting the quiet cap for.
*/
public final class LeadHeartbeatLoop {
@@ -74,15 +63,11 @@ public final class LeadHeartbeatLoop {
private final long backoffMs;
private final int quietNudgeCap;
private final Metrics metrics; // CB-512 pattern: nullable — no registry in unit tests
private final LeadContextSource contextSource; // fleetd #609
private final boolean contextHighNudge; // fleetd #609: opt-in, like the loop itself
/** When the current idle stretch began (nanos), or {@link #NOT_IDLE}. Single scheduler thread only. */
private long idleSinceNanos = NOT_IDLE;
/** Consecutive nudges that found no pending fleet state. Single scheduler thread only. */
private int quietCount = 0;
/** fleetd #609: latched "already told this lead about this HIGH stretch". Single scheduler thread only. */
private boolean contextNotified = false;
/** Constructor with an injectable clock and no metric registry (unit tests, or wiring that opts out). */
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
@@ -90,7 +75,7 @@ public final class LeadHeartbeatLoop {
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap) {
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
idleAfterNanos, backoffMs, quietNudgeCap, null, LeadContextSource.none(), false);
idleAfterNanos, backoffMs, quietNudgeCap, null);
}
/** As above, with a metric registry (the CB-512 pattern) so nudge outcomes are counted. */
@@ -98,20 +83,6 @@ public final class LeadHeartbeatLoop {
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics) {
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
idleAfterNanos, backoffMs, quietNudgeCap, metrics, LeadContextSource.none(), false);
}
/**
* fleetd #609: as above, plus the lead's own context source and whether a HIGH reading should
* append a hand-over notice to the loop's nudge. Pass {@link LeadContextSource#none()} and
* {@code false} to keep the pre-#609 behaviour exactly (both existing public constructors do).
*/
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics,
LeadContextSource contextSource, boolean contextHighNudge) {
this.primaryRegistry = primaryRegistry;
this.agents = agents;
this.inbox = inbox;
@@ -123,32 +94,12 @@ public final class LeadHeartbeatLoop {
this.backoffMs = backoffMs;
this.quietNudgeCap = quietNudgeCap;
this.metrics = metrics;
this.contextSource = contextSource;
this.contextHighNudge = contextHighNudge;
}
/**
* fleetd #609: one lead's own context reading, keyed by its terminal id — the same injected-source
* idiom {@code FleetMcp.LeadSeatSource}/{@code FleetMcp.LeadConfigDirSource} already use.
*/
public record LeadContextSource(Function<String, LeadContextGauge.Reading> readingFor) {
/** Inert source — every lead reads UNKNOWN, so the context notice can never fire. */
public static LeadContextSource none() {
return new LeadContextSource(_ -> LeadContextGauge.Reading.unknown());
}
}
/**
* Count one nudge outcome when a registry is wired; a no-op in unit tests.
*
* <p>fleetd #365: the {@code "sent"} outcome (renamed from {@code "delivered"}) records only
* that {@link #injectNudge} — a one-way herdr {@code agent.prompt} paste-and-submit — returned
* without throwing, not that the lead's pane actually read or acted on the text. This layer has
* no read-receipt concept, so "sent" is the honest word for what this call can ever establish.
*/
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
private void countNudge(String outcome) {
if (metrics != null) {
metrics.inc(FleetMetrics.HEARTBEAT_NUDGES, "outcome", outcome);
metrics.inc(BridgedMetrics.HEARTBEAT_NUDGES, "outcome", outcome);
}
}
@@ -168,11 +119,8 @@ public final class LeadHeartbeatLoop {
STAND_DOWN
}
/**
* The outcome of one decision: the action, the state to persist for the next tick, and (fleetd
* #609) whether the lead has now been told about the current HIGH context stretch.
*/
record Decision(Action action, Long idleSinceNanos, int quietCount, boolean contextNotified) {}
/** The outcome of one decision: the action plus the state to persist for the next tick. */
record Decision(Action action, Long idleSinceNanos, int quietCount) {}
/**
* Pure decision function: given the current loop state and fleet/lead facts, return what to do
@@ -187,47 +135,34 @@ public final class LeadHeartbeatLoop {
* @param pushLoopActive whether {@link ReplyPushLoop} is currently nudging some target (constraint 6)
* @param leadKnown whether a lead terminal is known to nudge at all
* @param fleet a snapshot of the pending fleet state (constraint 5)
* @param context fleetd #609: the lead's own {@link LeadContextGauge} reading for this tick
* @param contextNotified fleetd #609: whether the lead has already been told about the current HIGH
* stretch — a latch, carried forward by {@link #applyDecision}
* @return the action to take and the state to persist
*/
Decision decide(long nowNanos, Long idleSinceNanos, int quietCount, AgentStatus status,
boolean pushLoopActive, boolean leadKnown, FleetState fleet,
LeadContextGauge.State context, boolean contextNotified) {
// fleetd #609: re-arm the latch only on a positive OK reading. UNKNOWN means "I could not
// look", not "it got better" — re-arming on UNKNOWN would let a flapping gauge (a transcript
// read that misses one tick) nudge a full lead again on every recovery, defeating the "once
// per HIGH stretch" promise. Computed once, up front, so every gate below carries it forward
// unchanged unless it is the gate that actually discharges it.
boolean latch = context == LeadContextGauge.State.OK ? false : contextNotified;
boolean contextHigh = contextHighNudge && context == LeadContextGauge.State.HIGH;
boolean pushLoopActive, boolean leadKnown, FleetState fleet) {
// Constraint 6: while ReplyPushLoop is actively nudging the lead, injecting a second,
// competing prompt into the same pane would start a second turn — racing loops multiply
// turns and context burn. Stand aside, and treat the active push as real state (re-arm the
// quiet counter), because the reply that drove it is exactly the kind of new state that
// should reset the cap. The context latch is untouched: standing down must not spend the
// one notice this stretch gets.
// should reset the cap.
if (pushLoopActive) {
return new Decision(Action.STAND_DOWN, idleSinceNanos, 0, latch);
return new Decision(Action.STAND_DOWN, idleSinceNanos, 0);
}
// Constraint 2: a WORKING lead is making progress and must NOT be touched; an unreadable
// status (read failure, or the agent is gone) is safest treated the same way — never inject
// into a state we cannot read. Either way, reset the idle window and the quiet counter: the
// lead was / may be active, so the next idle stretch must count its own quiet period fresh.
if (status == null || !status.injectable()) {
return new Decision(Action.LEAD_BUSY, null, 0, latch);
return new Decision(Action.LEAD_BUSY, null, 0);
}
if (idleSinceNanos == null) {
// The lead just became injectable — record the start of an idle stretch and wait out the
// debounce quiet period before ever nudging (constraint 3).
return new Decision(Action.WAIT_IDLE, nowNanos, quietCount, latch);
return new Decision(Action.WAIT_IDLE, nowNanos, quietCount);
}
if (nowNanos - idleSinceNanos < idleAfterNanos) {
// Still within the quiet period: the lead that just finished a turn sits momentarily idle
// and must not be re-prompted into every natural pause.
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount, latch);
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
}
// Past the quiet period with an injectable lead: it is a genuine candidate for a nudge. Two
// gating facts decide whether and how:
@@ -235,33 +170,23 @@ public final class LeadHeartbeatLoop {
// No lead terminal is known yet (e.g. the registry has not learned one) — there is nobody
// to nudge. Keep waiting; the window stays open so discovery re-arms it without a fresh
// quiet period.
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount, latch);
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
}
if (fleet.hasPending()) {
// Real fleet state is waiting — a worker reply or a DONE session. This is new state, so
// it resets the quiet counter (constraint 4) and the lead is nudged to go collect it. The
// nudge text carries the context notice too when contextHigh — see injectNudge/contextNotice
// — so this route discharges the same duty and must set the latch.
return new Decision(Action.INJECT, idleSinceNanos, 0, latch || contextHigh);
}
if (contextHigh && !latch) {
// fleetd #609: the lead is idle, its context is full, and nothing is pending. This is the
// one case the quiet cap would otherwise swallow, and it is exactly when the lead most
// needs to hear it. Fire once per HIGH stretch, and do NOT spend the quiet budget on it:
// this is an event notice, not a "are you still there" nudge.
return new Decision(Action.INJECT, idleSinceNanos, quietCount, true);
// it resets the quiet counter (constraint 4) and the lead is nudged to go collect it.
return new Decision(Action.INJECT, idleSinceNanos, 0);
}
if (quietCount < quietNudgeCap) {
// Nothing is pending, but the cap is not exhausted: nudge anyway, telling the lead
// exactly that nothing is waiting so it can choose to stand down rather than hunt
// (constraint 5). Count it toward the consecutive-quiet cap. This nudge also carries the
// context notice when contextHigh (already latched above, or being latched now).
return new Decision(Action.INJECT, idleSinceNanos, quietCount + 1, latch || contextHigh);
// (constraint 5). Count it toward the consecutive-quiet cap.
return new Decision(Action.INJECT, idleSinceNanos, quietCount + 1);
}
// Nothing pending and the cap is exhausted: stop nudging until real state appears again
// (constraint 4). The loop still ticks on backoff so a genuinely new reply or session change
// re-arms it — QUIET_DONE stops injection, not observation.
return new Decision(Action.QUIET_DONE, idleSinceNanos, quietCount, latch);
return new Decision(Action.QUIET_DONE, idleSinceNanos, quietCount);
}
// --- loop ----------------------------------------------------------------------------------
@@ -276,152 +201,57 @@ public final class LeadHeartbeatLoop {
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
}
/** One loop tick, every {@link #backoffMs} — the thin scheduler around {@link #decide}. Package-private
* (mirroring {@link ReplyPushLoop#tick(String)}) so tests can drive it directly with a fake clock and a
* fake {@link AgentControl} instead of racing the scheduler thread. */
void tick() {
/** One loop tick, every {@link #backoffMs} — the thin scheduler around {@link #decide}. */
private void tick() {
boolean leadKnown = primaryRegistry.primaryTerminal().isPresent();
FleetState fleet = snapshot(inbox, roster);
AgentStatus status = AgentStatus.UNKNOWN;
LeadContextGauge.Reading reading = LeadContextGauge.Reading.unknown();
if (leadKnown) {
String leadTerminal = primaryRegistry.primaryTerminal().orElseThrow();
try {
status = agents.status(leadTerminal);
status = agents.status(primaryRegistry.primaryTerminal().orElseThrow());
} catch (RuntimeException e) {
// A failed status read degrades to "unknown" — decide() treats that like a busy lead
// and never injects into a state it cannot read. Retry on the next backoff.
log.debug("idle-heartbeat: status check failed for lead, will retry: {}", e.toString());
}
// fleetd #609: read the lead's own context regardless of status — decide() still gates on
// status first (constraint 2), so this is harmless work on a WORKING lead and lets the
// latch state stay accurate for whenever the lead does go idle.
reading = contextSource.readingFor().apply(leadTerminal);
}
Decision d = decide(clock.getAsLong(),
idleSinceNanos == NOT_IDLE ? null : idleSinceNanos,
quietCount, status, pushLoop.isActive(), leadKnown, fleet,
reading.state(), contextNotified);
quietCount, status, pushLoop.isActive(), leadKnown, fleet);
applyDecision(d);
switch (d.action()) {
case INJECT -> injectNudge(d, fleet, reading);
case QUIET_DONE -> {
countNudge("exhausted");
contextNotified = d.contextNotified();
}
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN -> contextNotified = d.contextNotified();
case INJECT -> injectNudge(fleet);
case QUIET_DONE -> countNudge("exhausted");
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN -> { /* nothing to inject, nothing to count */ }
}
scheduleNext();
}
/**
* Persist the idle/quiet state a decision returned, so the next tick starts from it.
*
* <p>fleetd #609 review: the context latch ({@link #contextNotified}) is deliberately <em>not</em>
* set here any more. Setting it from the decision unconditionally — before {@link #injectNudge} even
* tries to send — is exactly the review's blocker: a decision to notify is not the same fact as "the
* notice reached the pane". Every branch of {@link #tick} now assigns {@link #contextNotified} itself,
* once it knows whether a send happened and whether it carried the notice (see {@link #injectNudge}).
*/
/** Persist the state a decision returned, so the next tick starts from it. */
private void applyDecision(Decision d) {
idleSinceNanos = d.idleSinceNanos() == null ? NOT_IDLE : d.idleSinceNanos();
quietCount = d.quietCount();
}
/**
* Send the nudge to the known lead, with the fleetd #609 context notice appended when it applies, and
* persist the context latch based on what actually happened this tick — not merely what {@code d}
* chose to attempt.
*/
private void injectNudge(Decision d, FleetState fleet, LeadContextGauge.Reading reading) {
// fleetd #609 review: build the notice from the latch as it stood BEFORE this tick's decision —
// d.contextNotified() is the value to persist once delivery is confirmed, not the value the text
// itself should be built from. Otherwise a HIGH stretch that is still latched would never see the
// notice at all, defeating the very check this fixes.
String notice = contextNotice(contextHighNudge, reading, contextNotified);
/** Send the nudge to the known lead. */
private void injectNudge(FleetState fleet) {
var lead = primaryRegistry.primaryTerminal();
boolean sent = lead.isPresent() && trySend(lead.get(), fleet.nudgeText() + notice, notice);
// The latch becomes true only when all three hold: decide() chose to notify, a notice was
// actually included in the text, and the send reached the pane without throwing. Whenever no
// notice was attempted (disabled, not HIGH, or already latched), nothing was promised to the lead
// this tick, so apply the decision's own carried-forward value unconditionally — that is how the
// OK-only re-arm rule and STAND_DOWN's "don't burn the notice" rule keep working through this path
// too. A lead that disappeared between the decision and the send (lead.isEmpty()) is treated the
// same as a failed send: nothing reached the pane, so the latch must not be set.
contextNotified = notice.isEmpty() ? d.contextNotified() : sent;
}
/**
* Attempt one herdr send and count its outcome. Returns whether {@code agents.send} returned without
* throwing — the caller ({@link #injectNudge}) needs this to decide whether the fleetd #609 context
* latch may be persisted as set.
*/
private boolean trySend(String leadTerminal, String text, String notice) {
if (lead.isEmpty()) {
return; // the lead disappeared between the decision and the injection
}
String leadTerminal = lead.get();
try {
agents.send(leadTerminal, text);
agents.send(leadTerminal, fleet.nudgeText());
log.debug("idle-heartbeat: nudge sent to lead {} (quiet nudges so far in this stretch: {})",
leadTerminal, quietCount);
// fleetd #609: a nudge that carries the context notice is counted under its own outcome so
// it is visible in /metrics — one count per nudge either way, never two.
countNudge(notice.isEmpty() ? "sent" : "sent_context");
return true;
countNudge("delivered");
} catch (RuntimeException e) {
log.warn("idle-heartbeat: failed to nudge lead {}: {}", leadTerminal, e.toString());
countNudge("failed");
return false;
}
}
/**
* fleetd #609: the text appended to a nudge when the lead's own context is full — {@code ""}
* whenever the notice does not apply, so callers can unconditionally append this without an extra
* branch. Wording stays plain (CEFR B1) and honest that only the operator approves a roll — this
* loop only ever prints text, it never calls {@code fleet_handover} itself.
*
* @param enabled the {@code leadHeartbeat.contextHighNudge} config flag
* @param reading the lead's current {@link LeadContextGauge} reading
* @return the notice text (starting with a leading space, to append directly after {@link
* FleetState#nudgeText()}), or {@code ""} when disabled or the state is not {@code HIGH}
*/
static String contextNotice(boolean enabled, LeadContextGauge.Reading reading) {
return contextNotice(enabled, reading, false);
}
/**
* fleetd #609 review: as {@link #contextNotice(boolean, LeadContextGauge.Reading)}, but also gated on
* {@code alreadyNotified} — the context latch as it stood <em>before</em> the current tick's decision.
* Without this gate, every pending-driven {@code INJECT} that lands while the context stays {@code
* HIGH} would re-append the full notice on top of an already-latched stretch, making the notice's own
* closing sentence ("You will not be told again until your context reads ok.") false. {@link
* #injectNudge} is the only caller that passes a non-default {@code alreadyNotified}.
*
* @param alreadyNotified whether the lead has already been told about the current HIGH stretch
*/
static String contextNotice(boolean enabled, LeadContextGauge.Reading reading, boolean alreadyNotified) {
if (!enabled || alreadyNotified || reading.state() != LeadContextGauge.State.HIGH) {
return "";
}
StringBuilder sb = new StringBuilder(" Your own context is nearly full");
String compactionWord = reading.compactions() == 1 ? "compaction" : "compactions";
if (reading.tokens() != null) {
sb.append(": ").append(reading.tokens()).append(" tokens used, ")
.append(reading.compactions()).append(' ').append(compactionWord).append(" so far.");
} else {
// A HIGH reading always carries a non-null token count today: LeadContextGauge only
// reaches HIGH by comparing a number against HIGH_THRESHOLD_TOKENS. That invariant
// lives in another class and nothing asserts it, so this branch does not rely on it —
// it drops the token clause rather than printing "null tokens".
sb.append(" (").append(reading.compactions()).append(' ').append(compactionWord)
.append(" so far).");
}
sb.append(" A fresh session would work better. To hand over: call fleet_handover(action=\"open\"), "
+ "write the file it names, ask the operator, then call fleet_handover(action=\"confirm\", "
+ "token, operatorConfirmed). Only the operator can approve the roll. You will not be told "
+ "again until your context reads ok.");
return sb.toString();
}
/** Schedule the next tick on the scheduler thread pool. */
private void scheduleNext() {
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
@@ -460,7 +290,7 @@ public final class LeadHeartbeatLoop {
/** The nudge body, phrased for the two cases the heartbeat distinguishes. */
String nudgeText() {
StringBuilder sb = new StringBuilder(
"Heartbeat: you are idle and no fleet_send is waiting on you.");
"Heartbeat: you are idle and no bridge_send is waiting on you.");
if (hasPending()) {
sb.append(" The fleet has state to collect: ").append(pendingDetail());
} else {
@@ -479,10 +309,10 @@ public final class LeadHeartbeatLoop {
.append(" pending collection");
if (!replyTargets.isEmpty()) {
// Render each as the exact command so the lead can act without parsing: the nearest
// analogue to ReplyPushLoop's fleet_poll(target=...) nudge.
// analogue to ReplyPushLoop's bridge_poll(target=...) nudge.
sb.append(" (")
.append(String.join(", ",
replyTargets.stream().map(t -> "fleet_poll(target=" + t + ")").toList()))
replyTargets.stream().map(t -> "bridge_poll(target=" + t + ")").toList()))
.append(")");
}
sb.append(", ").append(doneSessions).append(" DONE session").append(doneSessions == 1 ? "" : "s")
@@ -0,0 +1,831 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.UUID;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.CompletionException;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
import java.util.concurrent.atomic.AtomicLong;
import java.util.concurrent.locks.ReentrantLock;
import java.util.function.LongSupplier;
/**
* The blocking delegation feature (CB-104): deliver {@code content} into a worker and block until
* the worker returns a <em>structured reply</em> via {@code bridge_reply} (the {@link Rendezvous}),
* then hand that reply back. Delivery is the {@link Injector}'s job (the background poller sends it
* when the worker is injectable); this service never drives the injector or scrapes the terminal —
* completion is the worker's explicit reply, not a guess about {@code agent_status}.
*
* <p>Sends are serialized per session so exactly one reply can be outstanding per worker, which is
* what lets a reply map unambiguously to its send (no cross-talk between concurrent callers).
*
* <p>If the worker never replies within the timeout, the caller gets a typed "still working" /
* "queued" outcome — the message may still be mid-flight. A finished-but-unreplied turn is caught
* by the CB-106 completion fallback (see {@link Rendezvous#resolveCompletion}).
*
* <p><strong>Async fire-and-poll (CB-107).</strong> A caller's MCP client caps a blocking call at
* ~60s, but a real delegated task runs for minutes. {@link #sendAsync} therefore runs the same
* blocking {@link #send} on a background virtual thread and hands back a <em>ticket</em> the caller
* polls with {@link #poll}. The blocking and async paths share one code path (and the same per-target
* serialization), so async inherits the reply + completion resolution behaviour for free.
*/
public final class MessageService {
private static final Logger log = LoggerFactory.getLogger(MessageService.class);
/**
* The window a fire-and-poll send waits for resolution — generous, since no caller is blocked on
* it; a real delegated task resolves (reply or completion) well within this, and only a genuinely
* hung worker rides it out.
*/
private static final long ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L;
/**
* How long a finished (terminal) ticket is retained for polling before it is pruned. Package-
* private (not {@code private}) so a test can advance an injected clock past it deterministically
* instead of duplicating the magic number or sleeping for real.
*/
static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
/** Outcome of a blocking send. */
public enum Outcome {
/** The worker called {@code bridge_reply}; {@code text} holds the structured answer. */
REPLIED,
/**
* The worker's delegated turn finished without a {@code bridge_reply} (CB-106 fallback);
* {@code text} is the scraped transcript tail rather than a structured answer.
*/
COMPLETED_UNREPLIED,
/**
* The worker ran the turn then wedged in an unrecoverable state (CB-109); {@code text} is the
* failure context (e.g. the error screen). Terminal, but not a successful completion.
*/
WORKER_FAILED,
/**
* The turn finished without a {@code bridge_reply} and the scrape matched the backend's
* configured usage-limit refusal pattern (CB-578 stage A); {@code text} is the reason,
* carrying the matched line. The worker's pane is healthy — only its account is refusing —
* so this is never reported as a completed reply, and is kept distinct from
* {@link #WORKER_FAILED} (a wedged worker) and a session simply going {@code GONE}.
*/
BACKEND_EXHAUSTED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the
* question and {@code turnId} correlates the answer. Not terminal — the primary answers with
* {@link #answer(String, String, long)} and the turn resumes.
*/
QUESTION,
/** Timed out after the message was delivered — the worker is still working. */
TIMED_OUT_WORKING,
/** Timed out before delivery — the message is still queued for the worker. */
TIMED_OUT_QUEUED,
/** Another send to this session was in flight for the whole window. */
BUSY,
/**
* An answer ({@link #answer(String, String, long)}) referenced a {@code turnId} that is no
* longer open — the worker's {@code bridge_ask} already timed out or was answered.
*/
STALE_TURN
}
/**
* @param outcome how the send ended (or paused)
* @param text the worker's answer when {@link #completed()} (a structured {@code bridge_reply}
* for {@link Outcome#REPLIED}, a scraped transcript tail for
* {@link Outcome#COMPLETED_UNREPLIED}), or the question for {@link Outcome#QUESTION},
* else {@code null}
* @param turnId correlation id for a {@link Outcome#QUESTION} (answered via
* {@link #answer(String, String, long)}), else {@code null}
*/
public record Reply(Outcome outcome, String text, String turnId) {
/** A reply with no correlation id (the common terminal outcomes). */
public Reply(Outcome outcome, String text) {
this(outcome, text, null);
}
/** Whether the worker's turn actually finished with an answer (replied or scraped). */
public boolean completed() {
return outcome == Outcome.REPLIED || outcome == Outcome.COMPLETED_UNREPLIED;
}
}
/** How a worker's {@code bridge_ask} (CB-205) resolved. */
public enum AskOutcome {
/** The primary answered; {@link AskResult#answer} carries it. */
ANSWERED,
/** No delegation was open to surface the question to — the worker has no one to ask. */
NO_WAITER,
/** The primary did not answer within the window. */
TIMED_OUT
}
/** The outcome of a worker's {@code bridge_ask}: how it resolved and (if answered) the answer. */
public record AskResult(AskOutcome outcome, String answer) {
}
/** Lifecycle phase of an async delegation ticket. */
public enum Phase {
/** Delegated and in flight — queued for the worker or being worked. */
PENDING,
/** The worker is paused in {@code bridge_ask}; {@link TaskView#reply} and {@link TaskView#turnId} identify it. */
ASKING,
/** The worker's turn finished; {@link TaskView#reply} holds the answer. */
DONE,
/** The delegation could not complete (timed out, worker gone, or busy). */
FAILED
}
/**
* A poll snapshot of an async delegation.
*
* @param reply the answer when {@link #phase} is {@link Phase#DONE}, or the question when
* {@link #phase} is {@link Phase#ASKING}; otherwise {@code null}
* @param replySource {@code "reply"} (structured {@code bridge_reply}) or {@code "transcript"}
* (completion scrape) when {@link Phase#DONE}, else {@code null}
* @param detail a human note (live worker status while pending, ask state, or failure reason)
* @param turnId correlation id for an {@link Phase#ASKING} ticket, else {@code null}
*/
public record TaskView(String ticket, Phase phase, String reply, String replySource, String detail,
String turnId) {
}
/** An in-flight or finished async delegation, keyed by its ticket. */
private static final class Task {
private final String ticket;
private final String target;
private final CompletableFuture<Reply> future = new CompletableFuture<>();
private final long createdNanos;
private volatile Reply question;
private volatile String turnId;
private Task(String ticket, String target, long createdNanos) {
this.ticket = ticket;
this.target = target;
this.createdNanos = createdNanos;
}
}
/**
* A worker session's currently-open {@code bridge_ask} question, surfaced so {@code bridge_status}
* can show it without the caller needing the ticket first (CB-582). Only covers async
* (fire-and-poll) delegations, which track the question on their {@link Task}; a blocking
* ({@code wait:true}) send already hands the question straight back to its own caller, so there is
* nothing hidden left for {@code bridge_status} to surface in that case.
*/
public record PendingAsk(String ticket, String question, String turnId) {
}
private final AgentControl agents;
private final Injector injector;
private final Rendezvous rendezvous;
private final ReplyInbox inbox;
private final ReplyPushLoop pushLoop;
private final Metrics metrics; // CB-502: nullable — no registry in unit tests
// CB-588: injectable so pruneTerminalTickets' 10-minute TICKET_TTL_NANOS can be exercised in a
// test without a real wait — same seam SessionManager already uses for its idle reaper (nowNanos).
private final LongSupplier nowNanos;
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
/** Async task that owns each exact forward rendezvous waiter. */
private final ConcurrentHashMap<CompletableFuture<Rendezvous.Resolution>, Task> asyncTasksByWaiter =
new ConcurrentHashMap<>();
/** Async tickets paused on a specific {@code bridge_ask} turn. */
private final ConcurrentHashMap<String, Task> asyncTasksByTurn = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
Thread.ofVirtual().name("bridge-async-", 0).factory());
/**
* Create with an explicit {@link ReplyInbox} and optional {@link ReplyPushLoop}.
*
* @param pushLoop nullable — when non-null, the push loop is notified on the no-waiter reply
* branch ({@link #reply}) so it can nudge the primary to drain the inbox,
* (CB-588) whenever an async ticket started by {@link #sendAsync} reaches a
* terminal phase, whenever {@link #poll} hands a terminal ticket to its caller,
* and (CB-582) whenever an async ticket's worker pauses mid-turn in
* {@code bridge_ask} or that pause ends (answered or lapsed)
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop) {
this(agents, injector, rendezvous, inbox, pushLoop, null);
}
/**
* As above, with a metric registry (CB-502). Instrumenting here rather than at the REST and MCP
* edges means both surfaces are counted by one piece of code and cannot drift.
*
* @param metrics nullable — when null, nothing is recorded
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop, Metrics metrics) {
this(agents, injector, rendezvous, inbox, pushLoop, metrics, System::nanoTime);
}
/** Test constructor with an injectable clock (CB-588: exercise the ticket-prune TTL without a real wait). */
MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox,
ReplyPushLoop pushLoop, Metrics metrics, LongSupplier nowNanos) {
this.agents = agents;
this.injector = injector;
this.rendezvous = rendezvous;
this.inbox = inbox;
this.pushLoop = pushLoop;
this.metrics = metrics;
this.nowNanos = nowNanos;
}
/** Create with an explicit {@link ReplyInbox} and no push loop. */
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox) {
this(agents, injector, rendezvous, inbox, null);
}
/** Backward-compatible constructor that uses a default {@link InMemoryReplyInbox}. */
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous) {
this(agents, injector, rendezvous, new InMemoryReplyInbox());
}
/** Current lifecycle status of a worker (the {@code GET /sessions/{id}/status} surface). */
public AgentStatus status(String target) {
return agents.status(target);
}
/** Read-only delegation fact for fleet views. */
public boolean hasAcceptedDelivery(String target) {
return rendezvous.isWaiting(target);
}
/** Read-only inbox fact for fleet views. */
public boolean hasInboxMessage(String target) {
return !inbox.peek(target).isEmpty();
}
/**
* Route a worker's explicit {@code bridge_reply}: resolve an open send, or queue it in the
* inbox if no send is currently open. Unlike the bare {@link Rendezvous#resolve}, a no-waiter
* result is <em>not</em> a failure — the reply is held for later drain.
*
* <p><strong>Do NOT use this for mid-turn questions.</strong> {@code bridge_ask} /
* {@link Rendezvous#resolveQuestion} must keep today's {@code NO_WAITER} behaviour — questions
* are interactive and must never be queued.
*
* @return always {@code true} — the reply either resolved a live send or was queued
*/
public boolean reply(String session, String content) {
if (rendezvous.resolve(session, content)) {
count(BridgedMetrics.REPLIES, "path", "rendezvous");
return true; // a live send took it — unchanged fast path
}
inbox.publish(session, UUID.randomUUID().toString(), content);
// A rising inbox share is the signal CB-307 exists to make visible: the worker finished but
// nobody was waiting, so delivery now depends on the push loop and a drain.
count(BridgedMetrics.REPLIES, "path", "inbox");
if (pushLoop != null) {
pushLoop.onReplyQueued(session);
}
return true; // held, not lost
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
if (metrics != null) {
metrics.inc(name, labels);
}
}
/** Count a send's terminal outcome and pass the reply through unchanged. */
private Reply recorded(Reply r) {
String label = sendOutcomeLabel(r.outcome());
if (label != null) {
count(BridgedMetrics.SENDS, "outcome", label);
}
return r;
}
/** Map a terminal send outcome to its metric label, or {@code null} for non-terminal ones. */
private static String sendOutcomeLabel(Outcome o) {
return switch (o) {
case REPLIED -> "replied";
case COMPLETED_UNREPLIED -> "completion_fallback";
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> "timeout";
case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted";
case STALE_TURN, QUESTION -> null; // not a completed delegation
};
}
/**
* Abandon any send still waiting on {@code target} because its session has gone away (CB-516).
*
* <p>Without this, tearing a worker down left its rendezvous waiter open: a blocking
* {@code bridge_send} kept blocking, and an async one kept reporting {@code PENDING} until
* {@link #ASYNC_TIMEOUT_MS} — thirty minutes — even though the worker provably no longer
* existed and the delegation could never complete. Worse, {@code poll} already had the evidence
* (it calls {@code liveStatus} to build its detail string and gets back {@code "unknown"}) and
* reported {@code PENDING} anyway.
*
* <p>Resolving the waiter as a failure — rather than letting it time out — also means the
* outcome is counted, so a torn-down delegation stops being invisible to {@code /metrics}.
*
* @return true if a live waiter was failed
*/
public boolean abandon(String target, String reason) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
boolean failed = waiter != null && !waiter.isDone() && rendezvous.resolveFailure(waiter, reason);
boolean asyncFailed = false;
for (Task task : tasks.values()) {
if (target.equals(task.target) && task.question == null
&& task.future.complete(new Reply(Outcome.WORKER_FAILED, reason))) {
asyncFailed = true;
}
}
if (failed) {
log.warn("abandoning the blocked send to {}: {}", target, reason);
}
return failed || asyncFailed;
}
/**
* Acknowledge a specific reply by {@code msgId} for {@code target}. Removes it from the inbox
* so that a subsequent drain or peek no longer returns it.
*/
public void ackReply(String target, String msgId) {
inbox.ack(target, msgId);
}
/**
* Drain (peek + ack) all pending inbox replies for {@code target}.
*
* <p><strong>The ack happens here, before the caller has the messages</strong> — before the MCP
* or REST response carrying them has been written, and long before the client has processed
* them. That ordering is what the two adapters disagree about, so do not read this method as
* "at-least-once" without qualifying which inbox is behind it (CB-529):
*
* <ul>
* <li>{@code InMemoryReplyInbox} — the ack only drops an entry from a local map. The messages
* are already in the returned list, so nothing can be lost after this point.
* <li>{@code AmqpReplyInbox} — the ack is a broker-side {@code basicAck}. Once it lands the
* broker has forgotten the message. If the daemon dies while writing the response, the
* reply is gone from the broker <em>and</em> the client never received it. Re-polling
* cannot recover it, because there is nothing left to re-deliver.
* </ul>
*
* <p>So the loss window is the response write, and it is a genuine loss rather than a
* redelivery. This is accepted, not overlooked: the alternative — ack on the next poll — turns
* every normal drain into a double delivery, which costs more than the window it closes. A
* caller that needs certainty re-polls; that is idempotent for every case except this one.
*
* <p>Any change here must be checked against <em>both</em> adapters. The previous version of
* this javadoc claimed "the ack is local", which was true when only the in-memory inbox existed
* and silently became false when the AMQP adapter landed.
*
* @return the drained messages, newest last (FIFO); empty list if none
*/
public List<ReplyInbox.InboxMessage> drainReplies(String target) {
var messages = inbox.peek(target);
for (var msg : messages) {
inbox.ack(target, msg.msgId());
}
return messages;
}
/**
* Deliver {@code content} to {@code target} (a herdr {@code terminal_id}) and block until the
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
*/
public Reply send(String target, String content, long timeoutMillis) {
return send(target, content, timeoutMillis, null);
}
/**
* As {@link #send(String, String, long)}, but with an accepted-delivery hook.
*
* <p>{@code onAccepted} is invoked exactly once, once this send has won {@code target}'s send
* lock and so become the <em>accepted target turn</em> — it runs <em>before</em> delivery is
* queued, so a throwing hook fails the send cleanly (the waiter it already opened is closed and
* nothing is left queued). It is <em>not</em> invoked when the send is {@link Outcome#BUSY}
* (lock never taken). A caller uses this to record that <em>it</em> now owns the delegation's
* reply routing (CB-548: {@code PrimaryRegistry} delegator ownership) — recording only on
* acceptance means a concurrent sender that times out {@code BUSY} can never steal ownership it
* never earned. {@code null} disables the hook.
*/
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted) {
return send(target, content, timeoutMillis, onAccepted, null);
}
/** Run a send, optionally stopping an async task that teardown already failed before acceptance. */
private Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, Task task) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null); // another send held the session the whole window
}
try {
if (task != null && task.future.isDone()) {
return task.future.getNow(null);
}
if (hasAsyncQuestion(target)) {
return new Reply(Outcome.BUSY, null); // the worker's current turn is paused for its lead
}
// Open the waiter BEFORE queueing delivery (CB-548). A fast reply — the worker already
// injectable the instant we enqueue — otherwise arrives before the waiter is registered
// and orphans into the inbox while this send blocks to the timeout (the enqueue-before-
// open race). Opening first also means a throwing onAccepted (fired before enqueue) or an
// enqueue failure is safely closed by the finally below: nothing is left queued, and the
// failed send leaves no stale waiter behind.
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
try {
if (task != null) {
asyncTasksByWaiter.put(reply, task);
}
TurnToken token = new TurnToken(target, reply);
// The send has won the lock; the accepted-delivery hook records delegator ownership
// here (CB-548). It runs BEFORE enqueue so a throwing hook — onAccepted is now a
// public callback — fails the send without queuing a message that would orphan.
if (onAccepted != null) {
onAccepted.run();
}
CompletableFuture<Void> delivered = injector.enqueue(target, content, token);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
return recorded(new Reply(
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
}
} finally {
asyncTasksByWaiter.remove(reply);
rendezvous.close(target, reply);
}
} finally {
lock.unlock();
}
}
/**
* A worker's mid-turn question (CB-205 reverse rendezvous): surface {@code question} to the
* primary by resolving its open blocking {@code bridge_send}, then block this (worker) call until
* the primary answers via {@link #answer} or {@code timeoutMillis} elapses. Identity is the
* worker's own session — it does not address the primary.
*
* <p>Returns {@link AskOutcome#NO_WAITER} when no delegation is open to surface the question to
* (nothing to answer it), {@link AskOutcome#ANSWERED} with the primary's answer, or
* {@link AskOutcome#TIMED_OUT} if the primary stayed silent. The worker resumes its turn either
* way — an answered ask hands back the answer; an unanswered one leaves it to proceed alone.
*/
public AskResult ask(String workerSession, String question, long timeoutMillis) {
Rendezvous.AskTicket ticket = rendezvous.openAsk(workerSession);
// Only the freshly-opening caller surfaces the question; a coalesced duplicate simply blocks on
// the shared answer future that the fresh owner is already responsible for.
if (ticket.fresh()) {
// Register the reverse waiter first, then surface the question — so the answer, which can
// arrive the instant the primary reacts, always finds an open waiter to resolve.
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(workerSession);
Task task = markAsyncQuestion(waiter, question, ticket.turnId());
if (!rendezvous.resolveQuestion(workerSession, question, ticket.turnId())) {
if (task != null) {
clearAsyncQuestion(ticket.turnId(), true);
}
rendezvous.closeAsk(ticket.turnId());
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
}
// CB-582: the question just became visible via bridge_poll (Phase.ASKING) for an async
// (wait:false) delegation — nudge the lead's own pane the same way a terminal ticket does
// (CB-588), since the lead's normal poll cadence is minutes away and the reverse-rendezvous
// window (~55s, see BridgeMcp/BridgedApp) is far shorter. A blocking (wait:true) send has
// no Task and gets the question directly in its own reply, so task == null there — nothing
// to nudge.
if (task != null && pushLoop != null) {
pushLoop.onQuestionOpened(task.ticket, workerSession, ticket.turnId(), question);
}
}
try {
String answer = ticket.answer().get(timeoutMillis, TimeUnit.MILLISECONDS);
return new AskResult(AskOutcome.ANSWERED, answer);
} catch (TimeoutException e) {
log.debug("bridge_ask from {} went unanswered in {}ms", workerSession, timeoutMillis);
clearAsyncQuestion(ticket.turnId(), true);
return new AskResult(AskOutcome.TIMED_OUT, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting the primary's answer for " + workerSession, e);
} finally {
// Only the fresh owner tears down the shared turn; a duplicate must leave it open.
if (ticket.fresh()) {
rendezvous.closeAsk(ticket.turnId());
// CB-582: tear the push loop's copy down at the same point, not only on the three
// paths that call clearAsyncQuestion. The answer future can complete exceptionally
// (ExecutionException) or the thread be interrupted, and both leave this method by
// throwing — the question would stay pending forever, keep being named in nudges
// until its own cap, and never be removed from the map. Already-closed is a no-op.
if (pushLoop != null) {
pushLoop.questionClosed(ticket.turnId());
}
}
}
}
/**
* The primary's answer to a worker's {@code bridge_ask} (CB-205): resolve the worker's blocked
* question identified by {@code turnId}, then — like a fresh {@link #send} — block for the worker's
* eventual {@code bridge_reply} as it finishes the resumed turn. The worker session is derived from
* {@code turnId}, never a caller argument.
*
* <p>Unlike {@link #send} this does not re-inject through the {@link Injector}: the worker is
* mid-turn (already picked up), so the answer flows back through its own open {@code bridge_ask}
* call, not a new status-gated delivery. The forward waiter is opened <em>before</em> the worker
* is unblocked so a reply that lands the instant it resumes is not lost.
*/
public Reply answer(String turnId, String content, long timeoutMillis) {
String workerSession = rendezvous.askSession(turnId);
if (workerSession == null) {
return new Reply(Outcome.STALE_TURN, null); // the ask lapsed (timed out or already answered)
}
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(workerSession, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null);
}
try {
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession);
if (!rendezvous.answerAsk(turnId, content)) {
rendezvous.close(workerSession, reply);
return new Reply(Outcome.STALE_TURN, null); // lapsed between the lookup and the unblock
}
clearAsyncQuestion(turnId, false);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
Reply result = new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
finishAsyncTask(turnId, result);
return result;
} catch (TimeoutException e) {
// The worker resumed but hasn't replied yet — no completion fallback arms an answered
// turn (it never re-entered the injector), so a silent worker rides out the window.
return new Reply(Outcome.TIMED_OUT_WORKING, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + workerSession, e);
} finally {
rendezvous.close(workerSession, reply);
}
} finally {
lock.unlock();
}
}
/**
* Fire-and-poll variant of {@link #send}: deliver {@code content} to {@code target} on a
* background virtual thread and return immediately with a ticket to {@link #poll}. This is how a
* long task is delegated without tripping the caller's MCP client call timeout.
*
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content) {
return sendAsync(target, content, null);
}
/**
* As {@link #sendAsync(String, String)}, with the accepted-delivery hook of
* {@link #send(String, String, long, Runnable)} — the running {@code send} invokes {@code onAccepted}
* the moment it becomes the accepted target turn, so async flooding records delegator ownership
* exactly as the blocking path does (CB-548).
*
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content, Runnable onAccepted) {
String ticket = "task-" + ticketSeq.incrementAndGet();
Task task = new Task(ticket, target, nowNanos.getAsLong());
tasks.put(ticket, task);
if (pushLoop != null) {
// CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a
// worker paused in bridge_ask leaves it running, per finishAsyncTask's own contract — so
// this fires exactly once, from whichever path completes it: finishAsyncTask(task, result)
// below on any non-QUESTION outcome of send() — a worker's bridge_reply, the CB-106
// completion fallback, a CB-109 wedge, TIMED_OUT, BUSY, or BACKEND_EXHAUSTED — the same
// finishAsyncTask reached via answer()'s finishAsyncTask(turnId, result) once a QUESTION
// is resolved, completeExceptionally(t) just below when send() itself throws, or a CB-516
// abandon() on teardown. Without this, MessageService.reply's rendezvous fast path (the
// one an async ticket always takes) never told the push loop anything happened — see the
// class javadoc on sendAsync/CB-107.
task.future.whenComplete((reply, ex) -> {
boolean failed = ex != null || reply == null || !reply.completed();
pushLoop.onTicketTerminal(ticket, target, failed);
});
}
asyncExecutor.submit(() -> {
try {
Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task);
if (result.outcome() == Outcome.QUESTION) {
// Keep the accepted owner until answer() finishes it. markAsyncQuestion may run
// just after resolveQuestion wakes this thread.
} else {
finishAsyncTask(task, result);
}
} catch (Throwable t) {
task.future.completeExceptionally(t);
}
});
pruneTerminalTickets();
log.debug("async send {} -> {}", ticket, target);
return ticket;
}
/**
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket;
* otherwise a {@link Phase#PENDING} view (with the live worker status as detail), a
* {@link Phase#DONE} view carrying the reply, or a {@link Phase#FAILED} view with the reason.
*/
public TaskView poll(String ticket) {
Task task = tasks.get(ticket);
if (task == null) {
return null;
}
CompletableFuture<Reply> f = task.future;
if (!f.isDone()) {
Reply question = task.question;
if (question != null) {
return new TaskView(ticket, Phase.ASKING, question.text(), null,
"worker is waiting for your answer", question.turnId());
}
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target), null);
}
// CB-588: the ticket is terminal and being handed to the caller right here — tell the push
// loop it is collected so a later tick's nudge never names a ticket the lead already has.
if (pushLoop != null) {
pushLoop.ticketCollected(ticket);
}
Reply r;
try {
r = f.getNow(null);
} catch (CompletionException | java.util.concurrent.CancellationException e) {
Throwable cause = (e instanceof CompletionException ce && ce.getCause() != null) ? ce.getCause() : e;
return new TaskView(ticket, Phase.FAILED, null, null, cause.getMessage(), null);
}
if (r.completed()) {
String source = r.outcome() == Outcome.REPLIED ? "reply" : "transcript";
return new TaskView(ticket, Phase.DONE, r.text(), source, null, null);
}
// A wedged worker (CB-109) or a backend-exhausted classification (CB-578 stage A) carries
// the real cause as its reason; the timeout/busy outcomes carry none, so fall back to the
// outcome name.
boolean carriesReason = r.outcome() == Outcome.WORKER_FAILED || r.outcome() == Outcome.BACKEND_EXHAUSTED;
String detail = carriesReason && r.text() != null
? r.text()
: "no reply — " + r.outcome().name().toLowerCase();
return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
}
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
private String liveStatus(String target) {
try {
return agents.status(target).name().toLowerCase();
} catch (RuntimeException e) {
return "unknown";
}
}
/**
* Drop finished tickets older than the TTL so {@link #tasks} cannot grow without bound.
*
* <p>{@code tasks} is the sole authority on whether a ticket still exists — {@link #poll} returns
* {@code null} the instant a ticket is gone from here, before it ever reaches the terminal branch
* that calls {@link ReplyPushLoop#ticketCollected}. Without telling the push loop about a prune
* too, its own {@code pendingTickets} entry would outlive the ticket it names: an unpolled ticket
* (or one the reminder cap already gave up on) is pruned here but never collected there, so it
* lingers in {@code pendingTickets} forever and rides along on every later nudge to the same lead
* — naming a ticket {@code bridge_poll} can no longer find (CB-588 follow-up).
*/
private void pruneTerminalTickets() {
long cutoff = nowNanos.getAsLong() - TICKET_TTL_NANOS;
tasks.entrySet().removeIf(e -> {
Task t = e.getValue();
boolean expired = t.future.isDone() && t.createdNanos < cutoff;
if (expired && pushLoop != null) {
pushLoop.ticketCollected(e.getKey());
}
return expired;
});
}
/** Record the active question for an async ticket; blocking sends have no entry and stay unchanged. */
private Task markAsyncQuestion(CompletableFuture<Rendezvous.Resolution> waiter, String text, String turnId) {
Task task = waiter == null ? null : asyncTasksByWaiter.get(waiter);
if (task != null) {
task.question = new Reply(Outcome.QUESTION, text, turnId);
task.turnId = turnId;
asyncTasksByTurn.put(turnId, task);
}
return task;
}
/** Clear an answered or lapsed question, but only when it matches the ticket's current turn. */
private void clearAsyncQuestion(String turnId, boolean forgetTurn) {
// CB-582: tell the push loop first — like ticketCollected, a removal for a turnId it never
// nudged about (or already dropped) is a harmless no-op, so this is safe to call unconditionally
// rather than threading the guard below through it.
if (pushLoop != null) {
pushLoop.questionClosed(turnId);
}
Task task = asyncTasksByTurn.get(turnId);
if (task != null && turnId.equals(task.turnId)) {
task.question = null;
if (forgetTurn) {
asyncTasksByTurn.remove(turnId, task);
task.turnId = null;
}
}
}
/** Complete and detach an async ticket after its worker's actual terminal reply. */
private void finishAsyncTask(Task task, Reply result) {
task.future.complete(result);
if (task.turnId != null) {
asyncTasksByTurn.remove(task.turnId, task);
}
}
/** Complete the async ticket correlated to a specific answered turn. */
private void finishAsyncTask(String turnId, Reply result) {
Task task = asyncTasksByTurn.get(turnId);
if (task != null) {
finishAsyncTask(task, result);
}
}
/** A new send must not open a waiter while an async ticket owns this worker's paused turn. */
private boolean hasAsyncQuestion(String target) {
return asyncTasksByTurn.values().stream().anyMatch(task -> target.equals(task.target));
}
/**
* The question {@code workerSession} is currently paused on via {@code bridge_ask}, if any
* (CB-582) — {@code bridge_status} uses this to show a pending question without the caller
* needing the ticket. {@code null} when the session has no open async question (including a
* session mid a <em>blocking</em> {@code bridge_ask}, which has no {@link Task} to look up — see
* {@link PendingAsk}).
*/
public PendingAsk pendingAsk(String workerSession) {
for (Task task : tasks.values()) {
Reply q = task.question;
if (q != null && workerSession.equals(task.target)) {
return new PendingAsk(task.ticket, q.text(), q.turnId());
}
}
return null;
}
/** Release the async executor. */
public void close() {
asyncExecutor.shutdown();
}
/** Map a rendezvous {@link Rendezvous.Kind} onto its send {@link Outcome} (shared by send/answer). */
private static Outcome outcomeOf(Rendezvous.Kind kind) {
return switch (kind) {
case REPLY -> Outcome.REPLIED;
case COMPLETION -> Outcome.COMPLETED_UNREPLIED;
case FAILED -> Outcome.WORKER_FAILED;
case BACKEND_EXHAUSTED -> Outcome.BACKEND_EXHAUSTED;
case QUESTION -> Outcome.QUESTION;
};
}
private static boolean tryLock(ReentrantLock lock, long millis) {
try {
return lock.tryLock(Math.max(0, millis), TimeUnit.MILLISECONDS);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting the session send lock", e);
}
}
private static long remainingMillis(long deadlineNanos) {
return (deadlineNanos - System.nanoTime()) / 1_000_000L;
}
}
@@ -1,13 +1,13 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicLong;
/**
* The reply rendezvous: where a blocking {@code fleet_send} awaits how the worker's delegated turn
* The reply rendezvous: where a blocking {@code bridge_send} awaits how the worker's delegated turn
* ends. The sending (primary) request thread {@link #open}s a waiter; it is resolved either by the
* worker's explicit {@code fleet_reply} ({@link #resolve}, arriving on a different thread via
* worker's explicit {@code bridge_reply} ({@link #resolve}, arriving on a different thread via
* {@code POST /sessions/{id}/reply}) or — the CB-106 fallback — by the injector observing the
* worker's delegated turn return to idle without a reply ({@link #resolveCompletion}).
*
@@ -27,14 +27,14 @@ public final class Rendezvous {
/** How a delegated turn ended (or paused). */
public enum Kind {
/** The worker called {@code fleet_reply} with a structured answer. */
/** The worker called {@code bridge_reply} with a structured answer. */
REPLY,
/** The worker's turn finished without a {@code fleet_reply}; {@code text} is a scrape. */
/** The worker's turn finished without a {@code bridge_reply}; {@code text} is a scrape. */
COMPLETION,
/** The worker ran the turn then wedged (CB-109); {@code text} is the failure context. */
FAILED,
/**
* The turn finished without a {@code fleet_reply}, and the scrape matched the backend's
* The turn finished without a {@code bridge_reply}, and the scrape matched the backend's
* configured usage-limit refusal pattern (CB-578 stage A); {@code text} is the reason,
* carrying the matched line. The pane is healthy — only the account is refusing — so this
* is kept separate from a session simply going {@code GONE}.
@@ -43,7 +43,7 @@ public final class Rendezvous {
/**
* The worker paused mid-turn to ask the primary a question (CB-205 reverse rendezvous);
* {@code text} is the question and {@code turnId} correlates the primary's answer back to
* the worker's blocked {@code fleet_ask}. Not terminal — the turn resumes after the answer.
* the worker's blocked {@code bridge_ask}. Not terminal — the turn resumes after the answer.
*/
QUESTION
}
@@ -75,7 +75,7 @@ public final class Rendezvous {
/** Reverse rendezvous (CB-205): worker questions awaiting the primary's answer, keyed by {@code turnId}. */
private final ConcurrentHashMap<String, AskWaiter> asks = new ConcurrentHashMap<>();
private final AtomicLong askSeq = new AtomicLong();
/** Per-session index of the currently-open ask, so duplicate fleet_ask calls coalesce onto one turn. */
/** Per-session index of the currently-open ask, so duplicate bridge_ask calls coalesce onto one turn. */
private final ConcurrentHashMap<String, String> openAsksBySession = new ConcurrentHashMap<>();
/**
@@ -132,7 +132,7 @@ public final class Rendezvous {
return complete(session, new Resolution(Kind.REPLY, content));
}
// --- reverse rendezvous (CB-205 fleet_ask) ------------------------------------------------
// --- reverse rendezvous (CB-205 bridge_ask) ------------------------------------------------
/**
* Open a reverse-rendezvous waiter for a worker's mid-turn question. If {@code session} already has
@@ -166,7 +166,7 @@ public final class Rendezvous {
}
/**
* Surface a worker's mid-turn {@code question} by resolving the primary's open {@code fleet_send}
* Surface a worker's mid-turn {@code question} by resolving the primary's open {@code bridge_send}
* with a {@link Kind#QUESTION} carrying {@code turnId}. Same session-keyed semantics as
* {@link #resolve}: the one outstanding send for {@code session} unblocks with the question.
*
@@ -184,7 +184,7 @@ public final class Rendezvous {
}
/**
* Resolve a worker's blocked {@code fleet_ask} with the primary's {@code answer}, unblocking it
* Resolve a worker's blocked {@code bridge_ask} with the primary's {@code answer}, unblocking it
* to resume its turn.
*
* @return {@code true} if the ask was still open and got the answer; {@code false} if the
@@ -195,7 +195,7 @@ public final class Rendezvous {
return w != null && w.answer().complete(answer);
}
/** Drop a reverse-rendezvous turn once its {@code fleet_ask} has resolved (answered or lapsed). */
/** Drop a reverse-rendezvous turn once its {@code bridge_ask} has resolved (answered or lapsed). */
public void closeAsk(String turnId) {
AskWaiter w = asks.get(turnId);
if (w == null) {
@@ -209,9 +209,9 @@ public final class Rendezvous {
/**
* Resolve a specific captured {@code waiter} as a completion (the delegated turn finished with no
* {@code fleet_reply}); {@code text} is the scraped transcript tail. The waiter is the one
* {@code bridge_reply}); {@code text} is the scraped transcript tail. The waiter is the one
* captured when this turn was delivered, so a late completion for turn N cannot land on turn N+1's
* send (CB-116). A no-op if that waiter was already resolved — a raced {@code fleet_reply} wins.
* send (CB-116). A no-op if that waiter was already resolved — a raced {@code bridge_reply} wins.
*
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
*/
@@ -233,7 +233,7 @@ public final class Rendezvous {
/**
* Resolve a specific captured {@code waiter} as {@link Kind#BACKEND_EXHAUSTED} (CB-578 stage A):
* the turn finished with no {@code fleet_reply} and the scrape matched the backend's configured
* the turn finished with no {@code bridge_reply} and the scrape matched the backend's configured
* usage-limit pattern; {@code reason} carries the matched line. Like
* {@link #resolveCompletion(CompletableFuture, String)} it targets the exact captured send
* (CB-116). A no-op if that waiter was already resolved — first resolution wins.
@@ -1,4 +1,4 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import java.util.List;
@@ -46,13 +46,6 @@ public interface ReplyInbox {
/** Non-destructive snapshot of pending replies for {@code target} (FIFO), empty list if none. */
List<InboxMessage> peek(String target);
/**
* Remove the reply {@code msgId} for {@code target} once the primary has taken it.
*
* @return {@code true} if an entry was actually removed, {@code false} if there was nothing to
* remove (unknown {@code target}, unowned {@code target}, or a {@code msgId} not held for
* it). A {@code false} is not an error — acking a {@code target} this daemon does not own is
* part of the normal contract, not a failure.
*/
boolean ack(String target, String msgId);
/** Remove the reply {@code msgId} for {@code target} once the primary has taken it. No-op if absent. */
void ack(String target, String msgId);
}
@@ -1,20 +1,16 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.Collection;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
@@ -23,9 +19,9 @@ import java.util.stream.Collectors;
/**
* A status-gated push loop that nudges a lead's own herdr pane when it has uncollected work
* waiting: a worker reply queued with no live {@code fleet_send} to resolve it (CB-307), an
* async delegation ticket ({@code fleet_send(wait:false)}) that reached a terminal phase
* (CB-588), or an async ticket's worker pausing mid-turn in {@code fleet_ask} to await an answer
* waiting: a worker reply queued with no live {@code bridge_send} to resolve it (CB-307), an
* async delegation ticket ({@code bridge_send(wait:false)}) that reached a terminal phase
* (CB-588), or an async ticket's worker pausing mid-turn in {@code bridge_ask} to await an answer
* (CB-582).
*
* <p><strong>CB-590: one schedule per lead.</strong> All three kinds of work are triggered
@@ -44,7 +40,7 @@ import java.util.stream.Collectors;
* still holds an unacked message ({@link #pendingReplies}), tickets not yet collected
* ({@link #pendingTickets}), and open questions not yet answered or lapsed
* ({@link #pendingQuestions}) — and sends at most one combined nudge per tick
* ({@link #injectNudge(String, int, int, int, int, int)}). Work that arrives while the lead is busy is
* ({@link #injectNudge(String, int, int, int)}). Work that arrives while the lead is busy is
* never lost: it is re-read fresh on every tick until the lead is injectable or its own reminder
* cap ({@link #maxReminders}) is reached — each source spends from its own budget, so one source
* exhausting its cap does not stop nudges about the others (post-CB-590 regression fix; see
@@ -54,30 +50,24 @@ import java.util.stream.Collectors;
public final class ReplyPushLoop {
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
static final String NUDGE_FORMAT = "Worker %s returned a reply — run fleet_poll(target=%s) to collect it";
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
/** Coalesced form, several uncollected replies for the same lead. */
static final String REPLIES_NUDGE_FORMAT =
"%d workers returned replies — run fleet_poll(target=...) for each to collect them: %s";
"%d workers returned replies — run bridge_poll(target=...) for each to collect them: %s";
/** Singular form, one uncollected ticket. */
static final String TICKET_NUDGE_FORMAT =
"Ticket %s finished%s — run fleet_poll(ticket=%s) to collect it";
"Ticket %s finished%s — run bridge_poll(ticket=%s) to collect it";
/** Coalesced form, several uncollected tickets for the same lead. */
static final String TICKETS_NUDGE_FORMAT =
"%d tickets finished%s — run fleet_poll(ticket=...) for each to collect them: %s";
/** Singular form, one worker paused mid-turn in fleet_ask (CB-582) — names the answer call directly. */
"%d tickets finished%s — run bridge_poll(ticket=...) for each to collect them: %s";
/** Singular form, one worker paused mid-turn in bridge_ask (CB-582) — names the answer call directly. */
static final String QUESTION_NUDGE_FORMAT =
"Worker %s asked a question (ticket %s) — answer it with fleet_send(turnId=\"%s\", "
"Worker %s asked a question (ticket %s) — answer it with bridge_send(turnId=\"%s\", "
+ "content=...) to resume its turn:\n%s";
/** Coalesced form, several open questions for the same lead. */
static final String QUESTIONS_NUDGE_FORMAT =
"%d workers are paused on a question — run fleet_poll(ticket=...) for each, then answer "
+ "with fleet_send(turnId=..., content=...): %s";
static final String BACKEND_INCIDENT_NUDGE_FORMAT =
"Backend credential %s is cooling for %d remaining seconds; affected profiles: %s; "
+ "affected workers: %s. Run fleet_list to see more.";
static final String BACKEND_TARGET_UNMAPPED_NUDGE_FORMAT =
"A backend error on worker %s could not be mapped to a credential (%s) — no cool-off "
+ "was applied. Run fleet_list to check that worker.";
"%d workers are paused on a question — run bridge_poll(ticket=...) for each, then answer "
+ "with bridge_send(turnId=..., content=...): %s";
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
@@ -97,19 +87,12 @@ public final class ReplyPushLoop {
/** Tickets that have gone terminal but not yet been polled, keyed by ticket. */
private final ConcurrentHashMap<String, PendingTicket> pendingTickets = new ConcurrentHashMap<>();
/**
* Open {@code fleet_ask} questions not yet answered or lapsed, keyed by {@code turnId}
* Open {@code bridge_ask} questions not yet answered or lapsed, keyed by {@code turnId}
* (CB-582). A question's own nudge count is tracked the same per-item way as
* {@link #pendingTickets} (CB-598): a fresh question keeps its source eligible regardless of
* how depleted an older, still-open question's count is.
*/
private final ConcurrentHashMap<String, PendingQuestion> pendingQuestions = new ConcurrentHashMap<>();
/** Backend incidents awaiting one successful delivery, keyed by incident id and owning lead. */
private final ConcurrentHashMap<IncidentLead, PendingIncident> pendingIncidents = new ConcurrentHashMap<>();
/** Incident/lead pairs already delivered. They make repeated incident reports one-shot. */
private final Set<IncidentLead> deliveredIncidents = ConcurrentHashMap.newKeySet();
private final ConcurrentHashMap<UnmappedTargetLead, PendingUnmappedTarget> pendingUnmappedTargets =
new ConcurrentHashMap<>();
private final Set<UnmappedTargetLead> deliveredUnmappedTargets = ConcurrentHashMap.newKeySet();
/** CB-590: leads with an active combined reminder schedule (replies and/or tickets and/or questions). */
private final ConcurrentHashMap<String, Boolean> activeLeads = new ConcurrentHashMap<>();
@@ -132,18 +115,10 @@ public final class ReplyPushLoop {
this.metrics = metrics;
}
/**
* Count one nudge outcome when a registry is wired; a no-op in unit tests.
*
* <p>fleetd #365: the {@code "sent"} outcome (renamed from {@code "delivered"}) records only
* that {@code agents.send} — a one-way herdr {@code agent.prompt} paste-and-submit — returned
* without throwing. Nothing in this loop, or anywhere downstream of it, confirms the pane
* actually read or acted on the text; there is no read-receipt concept at this layer. "Sent"
* says exactly that; "delivered" claimed more than this call can ever establish.
*/
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
private void countNudge(String outcome) {
if (metrics != null) {
metrics.inc(FleetMetrics.PUSH_NUDGES, "outcome", outcome);
metrics.inc(BridgedMetrics.PUSH_NUDGES, "outcome", outcome);
}
}
@@ -155,7 +130,7 @@ public final class ReplyPushLoop {
/**
* Reply targets still pending for {@code lead} — registered via {@link #onReplyQueued} and
* whose inbox still holds an unacked message. A target whose inbox has since drained (acked,
* or collected via a live {@code fleet_send} rendezvous instead) is dropped from
* or collected via a live {@code bridge_send} rendezvous instead) is dropped from
* {@link #pendingReplies} here rather than lingering forever; there is no explicit "reply
* collected" callback the way {@link #ticketCollected} exists for tickets, so the inbox itself
* is the only signal.
@@ -203,20 +178,7 @@ public final class ReplyPushLoop {
* tracked per question, not per lead per source).
*/
private record PendingQuestion(String turnId, String ticket, String target, String lead,
String question, int nudgeCount) {
}
private record IncidentLead(String incidentId, String lead) {
}
private record PendingIncident(IncidentLead key, String credential, List<String> profiles,
List<String> targets, int remainingCoolOffSeconds, int nudgeCount) {
}
private record UnmappedTargetLead(String target, String reason, String lead) {
}
private record PendingUnmappedTarget(UnmappedTargetLead key, int nudgeCount) {
String question, int nudgeCount) {
}
/** Questions still open for {@code lead}, snapshotted fresh for one tick. */
@@ -230,44 +192,6 @@ public final class ReplyPushLoop {
.collect(Collectors.toUnmodifiableSet());
}
/**
* Test seam only (fleetd #418): carries no production behaviour, and nothing in this class
* calls it. Exposes {@link #pendingQuestionTurnIdsFor} — the exact state {@link #decide} reads
* to decide whether a question keeps a lead's schedule alive.
*
* <p>{@code MessageService.ask()} does three things in order before a question is fully open to
* this loop: it flips the ticket's {@code poll()} phase to {@code Phase.ASKING}, then resolves
* the reverse-rendezvous waiter, then calls {@link #onQuestionOpened}, which is what actually
* populates {@link #pendingQuestions}. A test that barriers on {@code Phase.ASKING} observes only
* the first of those three steps — under load the asker thread can be descheduled between steps
* one and three, so the barrier releases before this method's underlying map is populated, and
* {@link #decide} correctly reports nothing pending yet. A test that must order itself after the
* state {@link #decide} actually reads waits on this instead of on the phase.
*
* @return an unmodifiable snapshot; empty for a lead with no open questions
*/
Set<String> pendingQuestionTurnIdsForTest(String lead) {
return pendingQuestionTurnIdsFor(lead);
}
private List<PendingIncident> pendingIncidentsFor(String lead) {
return pendingIncidents.values().stream().filter(i -> lead.equals(i.key().lead())).toList();
}
private Set<IncidentLead> pendingIncidentKeysFor(String lead) {
return pendingIncidentsFor(lead).stream().map(PendingIncident::key)
.collect(Collectors.toUnmodifiableSet());
}
private List<PendingUnmappedTarget> pendingUnmappedTargetsFor(String lead) {
return pendingUnmappedTargets.values().stream().filter(i -> lead.equals(i.key().lead())).toList();
}
private Set<UnmappedTargetLead> pendingUnmappedTargetKeysFor(String lead) {
return pendingUnmappedTargetsFor(lead).stream().map(PendingUnmappedTarget::key)
.collect(Collectors.toUnmodifiableSet());
}
/**
* The reply-source reminder count {@link #decide} should see for {@code lead} on this tick:
* the <em>minimum</em> nudge count among the reply targets currently pending for it (CB-598).
@@ -312,15 +236,6 @@ public final class ReplyPushLoop {
return min == Integer.MAX_VALUE ? 0 : min;
}
private int minIncidentNudgeCountFor(String lead) {
return pendingIncidentsFor(lead).stream().mapToInt(PendingIncident::nudgeCount).min().orElse(0);
}
private int minUnmappedTargetNudgeCountFor(String lead) {
return pendingUnmappedTargetsFor(lead).stream().mapToInt(PendingUnmappedTarget::nudgeCount)
.min().orElse(0);
}
/**
* Pure decision function: examine everything pending for {@code lead} — reply targets and
* tickets alike — and return what the loop should do.
@@ -357,35 +272,17 @@ public final class ReplyPushLoop {
* @param questionReminderCount the lowest nudge count among questions open for this lead
*/
Action decide(String lead, int replyReminderCount, int ticketReminderCount, int questionReminderCount) {
return decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount,
minIncidentNudgeCountFor(lead));
}
/** As above, with backend incidents as a fourth, independently bounded source. */
Action decide(String lead, int replyReminderCount, int ticketReminderCount, int questionReminderCount,
int incidentReminderCount) {
return decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount,
incidentReminderCount, minUnmappedTargetNudgeCountFor(lead));
}
/** As above, with unmapped backend targets as a fifth, independently bounded source. */
Action decide(String lead, int replyReminderCount, int ticketReminderCount, int questionReminderCount,
int incidentReminderCount, int unmappedTargetReminderCount) {
boolean hasReplyWork = !pendingReplyTargetsFor(lead).isEmpty();
boolean hasTicketWork = !pendingTicketIdsFor(lead).isEmpty();
boolean hasQuestionWork = !pendingQuestionTurnIdsFor(lead).isEmpty();
boolean hasIncidentWork = !pendingIncidentKeysFor(lead).isEmpty();
boolean hasUnmappedTargetWork = !pendingUnmappedTargetKeysFor(lead).isEmpty();
if (!hasReplyWork && !hasTicketWork && !hasQuestionWork && !hasIncidentWork && !hasUnmappedTargetWork) {
if (!hasReplyWork && !hasTicketWork && !hasQuestionWork) {
log.debug("push: nothing pending for lead {}, stopping reminder", lead);
return Action.STOP;
}
boolean replyEligible = hasReplyWork && replyReminderCount < maxReminders;
boolean ticketEligible = hasTicketWork && ticketReminderCount < maxReminders;
boolean questionEligible = hasQuestionWork && questionReminderCount < maxReminders;
boolean incidentEligible = hasIncidentWork && incidentReminderCount < maxReminders;
boolean unmappedTargetEligible = hasUnmappedTargetWork && unmappedTargetReminderCount < maxReminders;
if (!replyEligible && !ticketEligible && !questionEligible && !incidentEligible && !unmappedTargetEligible) {
if (!replyEligible && !ticketEligible && !questionEligible) {
log.debug("push: reminder cap ({}) reached for lead {} on every source with pending work, stopping",
maxReminders, lead);
countNudge("exhausted");
@@ -405,80 +302,6 @@ public final class ReplyPushLoop {
return Action.WAIT_BUSY;
}
/**
* Resolve who to nudge about {@code target}, the way every public entry point below wants it:
* {@link PrimaryRegistry#nudgeTargetFor}, but only after checking the delegating lead it names
* is still actually there (fleetd #368).
*
* <p><strong>The bug this closes.</strong> {@code PrimaryRegistry.forgetDelegation} is wired to
* exactly one event — a worker's release — because that is the only teardown the daemon already
* observes for a session in this map. Nothing removes a binding when the LEAD half goes away: a
* lead that is closed, crashes, or is relaunched leaves {@code leadByTarget} entries pointing at
* a terminal that no longer exists. {@code nudgeTargetFor} falls back to the single known
* primary only when the map holds nothing for {@code target} — a stale non-null entry beats the
* fallback every time, which is exactly backwards: the fallback's own javadoc argues it is safe
* precisely in the case a stale entry now hides.
*
* <p><strong>The fix.</strong> Before trusting a recorded delegation, probe the lead the same
* way {@link #decide} already does every tick ({@code agents.status}) — cheap, since it is a
* local herdr round-trip, and it is the same signal {@code AgentControl.paneByTerminal} already
* trusts to tell a genuinely dead target from a live one. A lead that fails the probe is treated
* as if it had never been recorded: the stale entry is forgotten (self-healing, exactly like
* {@code AgentControl.paneByTerminal} already does on {@code agent_not_found}) and resolution is
* retried, which now reaches the fallback {@code nudgeTargetFor} was built to reach — the same
* empty-map state its javadoc already argues is correct.
*
* <p><strong>fleetd #368 review — only a positive "gone" reading forgets the binding.</strong>
* The first version of this method treated <em>any</em> {@code RuntimeException} from the probe
* as death, which is the #359 mistake repeated: a transient socket blip or a codec error on a
* perfectly live lead would silently and permanently unbind it, with no re-record ever coming.
* That is destructive on one bad reading, exactly what #359 shipped a two-reading guard to avoid
* for the analogous lead-tab-liveness question. {@link #isLive} now matches
* {@code AgentControl.agentCall}'s own narrower rule (see its {@code agent_not_found} check): only
* that specific, affirmative "herdr has no such agent" signal counts as gone. Every other failure
* — timeout, transport error, a decode error — is treated as still live and the binding is left
* alone, because guessing wrong here is unrecoverable while guessing "live" merely costs one more
* retry on the next tick, which {@link #decide} already tolerates.
*/
private Optional<String> resolveLiveLead(String target) {
Optional<String> lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty() || isLive(lead.get())) {
return lead;
}
log.debug("push: lead {} delegated to for {} is no longer live, forgetting the stale binding "
+ "and falling back", lead.get(), target);
primaryRegistry.forgetDelegation(target);
return primaryRegistry.nudgeTargetFor(target);
}
/**
* Whether {@code lead} should still be trusted: {@code false} only when herdr affirmatively
* reports the terminal gone ({@code agent_not_found}), never on a merely inconclusive failure.
*
* <p>fleetd #368 review: an earlier version returned {@code false} for any {@code RuntimeException},
* which made a transient herdr hiccup on a live lead indistinguishable from the lead actually
* being dead — and the caller's response to {@code false} ({@code forgetDelegation}) is
* destructive and permanent. Narrowed to the one code {@code AgentControl.agentCall} itself
* already treats as a genuine, resolvable absence (see its {@code agent_not_found} handling) —
* every other {@code RuntimeException} is treated as "still live" and the binding survives to be
* probed again next time, which costs nothing worse than one more retry.
*/
private boolean isLive(String lead) {
try {
agents.status(lead);
return true;
} catch (RuntimeException e) {
boolean gone = e instanceof HerdrException he && "agent_not_found".equals(he.code());
if (gone) {
log.debug("push: lead {} no longer exists ({})", lead, e.toString());
} else {
log.debug("push: liveness check for lead {} was inconclusive ({}); treating as live "
+ "rather than risk destroying a live binding", lead, e.toString());
}
return !gone;
}
}
// --- public entrypoints ----------------------------------------------------------------------
/**
@@ -489,7 +312,7 @@ public final class ReplyPushLoop {
* backstop until a lead is recorded.
*/
public void onReplyQueued(String target) {
var lead = resolveLiveLead(target);
var lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}, skipping reminder", target);
return;
@@ -500,7 +323,7 @@ public final class ReplyPushLoop {
}
/**
* Called when an async delegation ticket ({@code fleet_send(wait:false)}, CB-107) reaches a
* Called when an async delegation ticket ({@code bridge_send(wait:false)}, CB-107) reaches a
* terminal phase — DONE or a failure. Unlike {@link #onReplyQueued}, which nudges about the
* durable-inbox no-waiter path, this covers the path {@code MessageService.reply} takes when a
* fire-and-poll send's own rendezvous waiter resolves the reply directly: that path returns
@@ -517,7 +340,7 @@ public final class ReplyPushLoop {
* @param failed whether the ticket ended in a failure phase rather than {@code DONE}
*/
public void onTicketTerminal(String ticket, String target, boolean failed) {
var lead = resolveLiveLead(target);
var lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on ticket {} (target {}), skipping nudge",
ticket, target);
@@ -529,7 +352,7 @@ public final class ReplyPushLoop {
}
/**
* Called when a ticket's terminal state has been collected via {@code fleet_poll}. Removes it
* Called when a ticket's terminal state has been collected via {@code bridge_poll}. Removes it
* from the pending set so a scheduled tick — and any nudge it sends — never names a ticket the
* lead already has. A ticket that was never pending (unknown ticket, or one nudged with no push
* loop configured) is a no-op.
@@ -539,20 +362,20 @@ public final class ReplyPushLoop {
}
/**
* Called when an async ticket's worker pauses mid-turn in {@code fleet_ask} (CB-582): the
* question is now visible via {@code fleet_poll} (Phase.ASKING), but the reverse-rendezvous
* window it opened with (~55s default, see {@code FleetMcp}/{@code FleetApp}) is far shorter
* Called when an async ticket's worker pauses mid-turn in {@code bridge_ask} (CB-582): the
* question is now visible via {@code bridge_poll} (Phase.ASKING), but the reverse-rendezvous
* window it opened with (~55s default, see {@code BridgeMcp}/{@code BridgedApp}) is far shorter
* than a lead's normal minutes-long poll cadence — exactly the gap this closes. Resolves the
* delegating lead the same way {@link #onTicketTerminal} does and coalesces onto the same
* per-lead schedule (CB-590).
*
* @param ticket the async ticket the question belongs to (for {@code fleet_poll})
* @param ticket the async ticket the question belongs to (for {@code bridge_poll})
* @param target the worker session that asked
* @param turnId correlation id the lead answers with ({@code fleet_send turnId=...})
* @param turnId correlation id the lead answers with ({@code bridge_send turnId=...})
* @param question the question text
*/
public void onQuestionOpened(String ticket, String target, String turnId, String question) {
var lead = resolveLiveLead(target);
var lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}'s question (turnId {}), skipping nudge",
target, turnId);
@@ -563,7 +386,7 @@ public final class ReplyPushLoop {
}
/**
* Called when a worker's {@code fleet_ask} resolves — answered or lapsed unanswered — so a
* Called when a worker's {@code bridge_ask} resolves — answered or lapsed unanswered — so a
* scheduled tick never nudges about a question the lead already handled. A {@code turnId} that
* was never pending (never nudged, or already closed) is a no-op.
*/
@@ -571,48 +394,6 @@ public final class ReplyPushLoop {
pendingQuestions.remove(turnId);
}
/**
* Queue a one-shot backend credential outage notice for every distinct lead that owns an
* affected worker. A successful injection records its {@code (incidentId, lead)} key, so a
* repeat report never reminds that lead again.
*/
public void onBackendIncident(String incidentId, Collection<String> targets, String credential,
Collection<String> profiles, int remainingCoolOffSeconds) {
Map<String, List<String>> targetsByLead = new ConcurrentHashMap<>();
for (String target : targets) {
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.warn("push: backend incident {} has no known lead for target {}", incidentId, target);
continue;
}
targetsByLead.computeIfAbsent(lead.get(), _ -> new ArrayList<>()).add(target);
}
List<String> profileNames = profiles.stream().sorted().toList();
for (var entry : targetsByLead.entrySet()) {
IncidentLead key = new IncidentLead(incidentId, entry.getKey());
if (deliveredIncidents.contains(key)) continue;
pendingIncidents.putIfAbsent(key, new PendingIncident(key, credential, profileNames,
entry.getValue().stream().sorted().toList(), remainingCoolOffSeconds, 0));
startOrCoalesce(entry.getKey());
}
}
/**
* Tell the owning lead that a classified backend error could not be tied to a credential.
* Without an owning lead, emit a warning because no control can act on the target.
*/
public void onBackendTargetUnmapped(String target, String reason) {
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.warn("push: backend target {} could not map to a credential: {}", target, reason);
return;
}
UnmappedTargetLead key = new UnmappedTargetLead(target, reason, lead.get());
if (deliveredUnmappedTargets.contains(key)) return;
pendingUnmappedTargets.putIfAbsent(key, new PendingUnmappedTarget(key, 0));
startOrCoalesce(lead.get());
}
// --- the schedule ----------------------------------------------------------------------------
/** Start a reminder schedule for {@code lead}, or join the one already running. */
@@ -644,25 +425,18 @@ public final class ReplyPushLoop {
Set<String> repliesBefore = pendingReplyTargetsFor(lead);
Set<String> ticketsBefore = pendingTicketIdsFor(lead);
Set<String> questionsBefore = pendingQuestionTurnIdsFor(lead);
Set<IncidentLead> incidentsBefore = pendingIncidentKeysFor(lead);
Set<UnmappedTargetLead> unmappedTargetsBefore = pendingUnmappedTargetKeysFor(lead);
int replyReminderCount = minReplyNudgeCountFor(lead);
int ticketReminderCount = minTicketNudgeCountFor(lead);
int questionReminderCount = minQuestionNudgeCountFor(lead);
int incidentReminderCount = minIncidentNudgeCountFor(lead);
int unmappedTargetReminderCount = minUnmappedTargetNudgeCountFor(lead);
var action = decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount,
incidentReminderCount, unmappedTargetReminderCount);
var action = decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount);
switch (action) {
case INJECT -> {
injectNudge(lead, replyReminderCount, ticketReminderCount, questionReminderCount,
incidentReminderCount, unmappedTargetReminderCount);
injectNudge(lead, replyReminderCount, ticketReminderCount, questionReminderCount);
scheduleNext(lead);
}
// Re-check after the configured backoff; the lead may become injectable soon.
case WAIT_BUSY -> scheduleNext(lead);
case STOP -> stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore, incidentsBefore,
unmappedTargetsBefore);
case STOP -> stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore);
}
}
@@ -706,25 +480,11 @@ public final class ReplyPushLoop {
* decision-to-release window reclaims the schedule slot exactly like a raced-in reply or ticket.
*/
void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore,
Set<String> questionsBefore) {
stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore, pendingIncidentKeysFor(lead));
}
private void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore,
Set<String> questionsBefore, Set<IncidentLead> incidentsBefore) {
stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore, incidentsBefore,
pendingUnmappedTargetKeysFor(lead));
}
private void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore,
Set<String> questionsBefore, Set<IncidentLead> incidentsBefore,
Set<UnmappedTargetLead> unmappedTargetsBefore) {
Set<String> questionsBefore) {
activeLeads.remove(lead);
boolean racedIn = pendingReplyTargetsFor(lead).stream().anyMatch(t -> !repliesBefore.contains(t))
|| pendingTicketIdsFor(lead).stream().anyMatch(t -> !ticketsBefore.contains(t))
|| pendingQuestionTurnIdsFor(lead).stream().anyMatch(t -> !questionsBefore.contains(t))
|| pendingIncidentKeysFor(lead).stream().anyMatch(i -> !incidentsBefore.contains(i))
|| pendingUnmappedTargetKeysFor(lead).stream().anyMatch(i -> !unmappedTargetsBefore.contains(i));
|| pendingQuestionTurnIdsFor(lead).stream().anyMatch(t -> !questionsBefore.contains(t));
if (racedIn && activeLeads.putIfAbsent(lead, Boolean.TRUE) == null) {
log.debug("push: new work for lead {} raced the reminder loop's stop — restarting", lead);
scheduleNext(lead);
@@ -735,22 +495,18 @@ public final class ReplyPushLoop {
/** Send one combined nudge covering everything currently pending for {@code lead}. */
private void injectNudge(String lead, int replyReminderCount, int ticketReminderCount,
int questionReminderCount, int incidentReminderCount,
int unmappedTargetReminderCount) {
int questionReminderCount) {
// Re-read rather than threading it down from decide(): a reply can drain, a ticket be
// collected, or a question be answered (or another arrive), between the decision and the
// injection.
Set<String> replyTargets = pendingReplyTargetsFor(lead);
List<PendingTicket> tickets = pendingTicketsFor(lead);
List<PendingQuestion> questions = pendingQuestionsFor(lead);
List<PendingIncident> incidents = pendingIncidentsFor(lead);
List<PendingUnmappedTarget> unmappedTargets = pendingUnmappedTargetsFor(lead);
if (replyTargets.isEmpty() && tickets.isEmpty() && questions.isEmpty() && incidents.isEmpty()
&& unmappedTargets.isEmpty()) {
if (replyTargets.isEmpty() && tickets.isEmpty() && questions.isEmpty()) {
log.debug("push: pending work for lead {} drained before the nudge could be sent", lead);
return;
}
String nudge = formatNudge(replyTargets, tickets, questions, incidents, unmappedTargets);
String nudge = formatNudge(replyTargets, tickets, questions);
try {
agents.send(lead, nudge);
log.debug("push: nudge sent to lead {} (reply {}/{}, ticket {}/{}, question {}/{}; "
@@ -758,17 +514,7 @@ public final class ReplyPushLoop {
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
questionReminderCount + 1, maxReminders,
replyTargets.size(), tickets.size(), questions.size());
countNudge("sent");
for (PendingIncident incident : incidents) {
if (pendingIncidents.remove(incident.key(), incident)) {
deliveredIncidents.add(incident.key());
}
}
for (PendingUnmappedTarget unmappedTarget : unmappedTargets) {
if (pendingUnmappedTargets.remove(unmappedTarget.key(), unmappedTarget)) {
deliveredUnmappedTargets.add(unmappedTarget.key());
}
}
countNudge("delivered");
} catch (RuntimeException e) {
log.warn("push: failed to nudge lead {} (reply {}/{}, ticket {}/{}, question {}/{}): {}",
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
@@ -780,13 +526,12 @@ public final class ReplyPushLoop {
// item already at or over the cap keeps riding along in the text (still pending, still
// named) but its extra bumps here are inert: decide() already treats it as ineligible once
// its count reaches maxReminders.
bumpNudgeCounts(replyTargets, tickets, questions, incidents, unmappedTargets);
bumpNudgeCounts(replyTargets, tickets, questions);
}
/** Record that every one of these items was just named in a sent (or attempted) nudge. */
private void bumpNudgeCounts(Set<String> replyTargets, List<PendingTicket> tickets,
List<PendingQuestion> questions, List<PendingIncident> incidents,
List<PendingUnmappedTarget> unmappedTargets) {
List<PendingQuestion> questions) {
for (String target : replyTargets) {
pendingReplies.computeIfPresent(target, (t, e) -> new ReplyEntry(e.lead(), e.nudgeCount() + 1));
}
@@ -799,14 +544,6 @@ public final class ReplyPushLoop {
new PendingQuestion(e.turnId(), e.ticket(), e.target(), e.lead(), e.question(),
e.nudgeCount() + 1));
}
for (PendingIncident incident : incidents) {
pendingIncidents.computeIfPresent(incident.key(), (id, e) -> new PendingIncident(e.key(),
e.credential(), e.profiles(), e.targets(), e.remainingCoolOffSeconds(), e.nudgeCount() + 1));
}
for (PendingUnmappedTarget unmappedTarget : unmappedTargets) {
pendingUnmappedTargets.computeIfPresent(unmappedTarget.key(), (id, e) ->
new PendingUnmappedTarget(e.key(), e.nudgeCount() + 1));
}
}
/** Schedule the next tick on the scheduler thread pool. */
@@ -819,8 +556,7 @@ public final class ReplyPushLoop {
/** Render everything pending for one lead as a single nudge line. */
private static String formatNudge(Set<String> replyTargets, List<PendingTicket> tickets,
List<PendingQuestion> questions, List<PendingIncident> incidents,
List<PendingUnmappedTarget> unmappedTargets) {
List<PendingQuestion> questions) {
List<String> parts = new ArrayList<>();
if (!replyTargets.isEmpty()) {
parts.add(formatRepliesNudge(replyTargets));
@@ -831,12 +567,6 @@ public final class ReplyPushLoop {
if (!questions.isEmpty()) {
parts.add(formatQuestionsNudge(questions));
}
if (!incidents.isEmpty()) {
parts.addAll(incidents.stream().map(ReplyPushLoop::formatBackendIncidentNudge).toList());
}
if (!unmappedTargets.isEmpty()) {
parts.addAll(unmappedTargets.stream().map(ReplyPushLoop::formatUnmappedTargetNudge).toList());
}
return String.join(" | ", parts);
}
@@ -876,16 +606,6 @@ public final class ReplyPushLoop {
return QUESTIONS_NUDGE_FORMAT.formatted(pending.size(), ids);
}
private static String formatBackendIncidentNudge(PendingIncident incident) {
return BACKEND_INCIDENT_NUDGE_FORMAT.formatted(incident.credential(), incident.remainingCoolOffSeconds(),
String.join(", ", incident.profiles()), String.join(", ", incident.targets()));
}
private static String formatUnmappedTargetNudge(PendingUnmappedTarget unmappedTarget) {
return BACKEND_TARGET_UNMAPPED_NUDGE_FORMAT.formatted(unmappedTarget.key().target(),
unmappedTarget.key().reason());
}
// --- lifecycle -----------------------------------------------------------------------------
/**
@@ -907,10 +627,6 @@ public final class ReplyPushLoop {
pendingReplies.clear();
pendingTickets.clear();
pendingQuestions.clear();
pendingIncidents.clear();
deliveredIncidents.clear();
pendingUnmappedTargets.clear();
deliveredUnmappedTargets.clear();
}
/** @see #stop() */
@@ -1,4 +1,4 @@
package dev.ltms.fleet.msg;
package dev.ltms.bridged.msg;
import java.util.concurrent.CompletableFuture;
@@ -9,19 +9,12 @@ import java.util.concurrent.CompletableFuture;
public final class TurnToken {
private final String target;
private final CompletableFuture<Rendezvous.Resolution> waiter;
private final String injectedText;
public TurnToken(String target, CompletableFuture<Rendezvous.Resolution> waiter) {
this(target, waiter, null);
}
public TurnToken(String target, CompletableFuture<Rendezvous.Resolution> waiter, String injectedText) {
this.target = target;
this.waiter = waiter;
this.injectedText = injectedText;
}
public String target() { return target; }
public CompletableFuture<Rendezvous.Resolution> waiter() { return waiter; }
public String injectedText() { return injectedText; }
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.peer;
package dev.ltms.bridged.peer;
/**
* Declared capabilities of a {@link PeerLauncher}. The protocol is the union across all
@@ -9,7 +9,7 @@ package dev.ltms.fleet.peer;
public enum Capability {
/**
* The peer supports {@code fleet_ask} rendezvous — pausing its delegated turn to ask
* The peer supports {@code bridge_ask} rendezvous — pausing its delegated turn to ask
* the primary a question, then resuming once answered. All Claude Code peers support this.
*/
MID_TURN_ASK,
@@ -1,4 +1,4 @@
package dev.ltms.fleet.peer;
package dev.ltms.bridged.peer;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
@@ -1,4 +1,4 @@
package dev.ltms.fleet.peer;
package dev.ltms.bridged.peer;
import java.util.Locale;
@@ -29,8 +29,8 @@ public enum MemberRole {
* <p>Reads the repo and writes analysis. Never commits code and never opens a pull request —
* an architect that starts implementing has stopped doing the job that makes it useful.
*
* <p>Architects are the one member kind with live slot binding, because a lead addresses the
* same slots across many tickets and needs a stable name for them.
* <p>Architects are the one member kind declared in config, because a lead addresses the same
* slots across many tickets and needs a stable name for them.
*/
ARCHITECT,
@@ -43,14 +43,6 @@ public enum MemberRole {
*/
DEV,
/**
* Sweeps an assigned package for defects and reports several ranked findings.
*
* <p>Never changes code, commits, or opens a pull request. A hunt gathers evidence, which can
* include running the build, but leaves every fix to a later implementation unit.
*/
HUNTER,
/**
* Reviews a diff it did not write and reports one structured finding.
*
@@ -67,7 +59,7 @@ public enum MemberRole {
/**
* The {@code fleet:} block that holds this role's pool — {@code architects},
* {@code developers}, {@code hunters}, {@code reviewers}.
* {@code developers}, {@code reviewers}.
*
* <p>Plural, and not always the wire name: the pool of things a {@code dev} may run on reads
* naturally as {@code developers:}. The wire name stays the singular {@code dev}, because that
@@ -77,7 +69,6 @@ public enum MemberRole {
return switch (this) {
case ARCHITECT -> "architects";
case DEV -> "developers";
case HUNTER -> "hunters";
case REVIEWER -> "reviewers";
};
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.peer;
package dev.ltms.bridged.peer;
/**
* An opaque handle returned by {@link PeerLauncher#spawn(SpawnRequest)}. The core routes on
@@ -0,0 +1,107 @@
package dev.ltms.bridged.peer;
import java.util.List;
import java.util.Set;
/**
* SPI for materializing a connected peer — the only way the bridge core creates or tears down
* a peer process. Every launcher is a first-party, in-tree adapter selected by (future) profile
* config; today's single adapter is the {@code ClaudeCodeLauncher} / Claude Code over herdr.
*
* <p>The core delegates spawn and teardown to this interface without knowing how the peer is set
* up. Environment variables, CLI flags, subscription guards, transport (herdr tab/pane) layout,
* and naming conventions are all adapter-private — the core sees only the returned
* {@link PeerHandle} whose {@code id()} is the registry/routing key.
*
* <p>The interface is a superset of what {@code SessionManager} and {@code Bridged.main} call
* on the concrete launcher today.
*/
public interface PeerLauncher {
/**
* The set of {@link Capability capabilities} this launcher declares. A peer whose profile
* opts into a git-forge token should include {@link Capability#SELF_PR}; the base set for
* the Claude Code herdr adapter is always {@code MID_TURN_ASK, WORKTREE, ORPHAN_REAP}.
*/
Set<Capability> capabilities();
/**
* The capabilities of the adapter that {@code profileName} resolves to (null/blank → the
* default profile, the same resolution {@link #spawn} uses). Distinct from {@link
* #capabilities()}, which unions every configured adapter: a caller that must know whether
* <em>this</em> profile's backend supports a capability — e.g. {@link Capability#SESSION_RESUME}
* before honoring {@link SpawnRequest#resumeSessionId()} — needs the per-profile answer, not
* the fleet-wide union, or a mixed fleet could OK a resume that lands on a non-supporting
* adapter (CB-584).
*
* @throws IllegalArgumentException if the profile is unknown and no default is configured
*/
Set<Capability> capabilitiesFor(String profileName);
/**
* {@code profileName}/requestedCwd null/blank → default resolution. Returns after the peer
* process is live (env + argv + placement complete). Never returns {@code null}.
*
* @param req the spawn parameters (profile, requested cwd, caller cwd)
* @return a handle whose {@link PeerHandle#id()} is the registry/routing key
* @throws IllegalArgumentException if the profile is unknown and no default is configured
*/
PeerHandle spawn(SpawnRequest req);
/**
* The configured worker profile names — the set of names {@code spawn(profileName)} accepts.
*/
Set<String> profiles();
/**
* The profile a no-argument {@link #spawn(SpawnRequest)} uses, or {@code null} if none is configured.
*/
String defaultProfile();
/**
* Resolve the effective working directory for a spawn {@code req} without actually spawning.
* Resolution order: requestedCwd → profile cwd → callerCwd → daemon cwd.
*
* @return the resolved absolute path, never null/blank
*/
String effectiveCwd(SpawnRequest req);
/**
* The parity-overlay file list for {@code profileName} (default list when unset). Used by
* worktree provisioning to copy config files into the isolated checkout before spawning.
*/
List<String> parityOverlay(String profileName);
/**
* The set of all agents this launcher currently tracks, transport-specific. Each element
* exposes at minimum a pane-like {@code id()} matching this launcher's {@link PeerHandle}
* scheme, plus transport-level status. Callers merge this set with the session registry to
* build a live roster view.
*/
List<?> list();
/**
* Reap orphaned peers left behind by a prior daemon process. Only peers whose naming scheme
* matches this launcher's and whose nonce differs from the current process are eligible.
* Best-effort: a failure to list or to stop any one peer is logged and never aborts startup.
*
* @return the number of orphaned peers reaped
*/
int reapOrphanWorkers();
/**
* Tear a peer down by its registry/routing key ({@link PeerHandle#id()}). Tolerates an
* already-gone peer. Also cleans up launcher-private resources (e.g. empty dedicated tabs)
* when safe to do so.
*/
void stop(String id);
/**
* Discard the context of the peer identified by {@code id}. Implementations must bypass normal
* bridge delivery/turn accounting. Unsupported peer kinds return {@code false} without sending
* a guessed command.
*
* @return {@code true} when a reset was sent and its status transition must settle before reuse
*/
boolean clearContext(String id);
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.peer;
package dev.ltms.bridged.peer;
/**
* Thrown when a {@link PeerLauncher} starts a peer process but the peer
@@ -1,4 +1,4 @@
package dev.ltms.fleet.peer;
package dev.ltms.bridged.peer;
/**
* What a caller asks for when spawning a peer.
@@ -29,9 +29,4 @@ public record SpawnRequest(String profileName, String requestedCwd, String calle
String sessionName, String resumeSessionId) {
this(profileName, requestedCwd, callerCwd, sessionName, resumeSessionId, null);
}
/** Return a copy of this request with {@code profileName} replaced by {@code profile}. */
public SpawnRequest withProfile(String profile) {
return new SpawnRequest(profile, requestedCwd, callerCwd, sessionName, resumeSessionId, role);
}
}
@@ -0,0 +1,115 @@
package dev.ltms.bridged.placement;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Objects;
import java.util.OptionalLong;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.LongSupplier;
/**
* Where a credential (not a profile — see {@code BridgedConfig.Profile#effectiveCredentialId()})
* sits out a cooldown after a {@code BACKEND_EXHAUSTED} classification (CB-578 stage B), so a fresh
* spawn does not walk straight back onto the account that just refused on a usage limit.
*
* <p>Keyed by credential id, never by profile name: two profiles sharing one credential (e.g. two
* models on the same OpenAI account) share one quarantine — {@link #quarantine} one credential id
* and every profile whose {@code effectiveCredentialId()} equals it is quarantined too, without this
* class knowing anything about profiles at all. That mapping is the caller's job (see
* {@code CompositePeerLauncher} and {@code dev.ltms.bridged.inject.ExhaustionSink}).
*
* <p>The clock is injected ({@link LongSupplier}, conventionally {@code System::nanoTime} like
* {@code FleetHealthMonitor}), never read inline, so a quarantine's expiry is testable without a
* real sleep.
*/
public final class BackendQuarantine {
private final ConcurrentHashMap<String, Long> quarantinedUntilNanos = new ConcurrentHashMap<>();
private final LongSupplier nowNanos;
private final long cooldownNanos;
/** True only for {@link #none()}. See {@link #quarantine} for why this exists. */
private final boolean inert;
/**
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos how long a fresh {@link #quarantine} call blocks the credential for;
* must be positive
*/
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos) {
this(nowNanos, cooldownNanos, false);
}
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, boolean inert) {
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
if (cooldownNanos <= 0) {
throw new IllegalArgumentException("cooldownNanos must be positive: " + cooldownNanos);
}
this.cooldownNanos = cooldownNanos;
this.inert = inert;
}
/**
* Inert quarantine — {@link #quarantine} does nothing on this instance, so nothing is ever
* quarantined. The explicit stand-in a caller (or a test not exercising this feature) passes
* instead of a defaulting overload, exactly like {@code ExhaustedPatternLookup.none()}.
*/
public static BackendQuarantine none() {
return new BackendQuarantine(() -> 0L, 1, true);
}
/**
* Quarantine {@code credentialId} for the configured cooldown, starting now. A repeat call while
* already quarantined restarts the cooldown at full length — a fresh refusal is fresh evidence the
* account is still exhausted, not a reason to let an earlier, shorter wait stand.
*
* <p>On {@link #none()} this is a no-op. It has to be: that instance holds a clock frozen at 0,
* so recording a deadline would produce a quarantine that never expires — a credential locked out
* for the life of the daemon. Two production {@code CompositePeerLauncher} constructors default to
* {@code none()}, so the failure would be silent and permanent. An inert stand-in must omit the
* fact, never invent one.
*/
public void quarantine(String credentialId) {
Objects.requireNonNull(credentialId, "credentialId");
if (inert) {
return;
}
quarantinedUntilNanos.put(credentialId, nowNanos.getAsLong() + cooldownNanos);
}
/** Whether {@code credentialId} is quarantined right now. */
public boolean isQuarantined(String credentialId) {
return remainingNanos(credentialId) > 0;
}
/** Seconds left on {@code credentialId}'s quarantine, or empty when it is not quarantined. */
public OptionalLong remainingSeconds(String credentialId) {
long remaining = remainingNanos(credentialId);
return remaining > 0 ? OptionalLong.of(toSecondsRoundedUp(remaining)) : OptionalLong.empty();
}
/**
* Every currently-quarantined credential id and its remaining seconds (CB-578 stage B fleet
* reporting) — expired entries are never included. Not pruned from the backing map here: it stays
* small (bounded by the number of distinct credentials ever exhausted) and a lazily-stale entry is
* harmless, since every read already checks the deadline.
*/
public Map<String, Long> activeRemainingSeconds() {
Map<String, Long> out = new LinkedHashMap<>();
quarantinedUntilNanos.forEach((credentialId, deadline) -> {
long remaining = deadline - nowNanos.getAsLong();
if (remaining > 0) {
out.put(credentialId, toSecondsRoundedUp(remaining));
}
});
return out;
}
private long remainingNanos(String credentialId) {
Long deadline = quarantinedUntilNanos.get(credentialId);
return deadline == null ? 0L : deadline - nowNanos.getAsLong();
}
private static long toSecondsRoundedUp(long nanos) {
return (nanos + 999_999_999L) / 1_000_000_000L;
}
}
@@ -0,0 +1,66 @@
package dev.ltms.bridged.placement;
/**
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps and
* reachability so that a pre-existing config behaves identically after upgrade.
*
* <p>Two exceptions walk past the default instead of returning it unconditionally:
* <ul>
* <li>Quarantine (CB-578 stage B): a quarantined default is a credential that just refused on
* a usage limit, not a transient capacity or reachability concern.
* <li>Weight 0 (CB-554): {@code fixed} is still automatic selection, so a profile the operator
* marked "never auto-select me" ({@code weight <= 0}) must be skipped here exactly as
* {@code weighted}/{@code round-robin} skip it — an explicit {@code bridge_spawn} naming
* the profile is unaffected, only this automatic fallback walk.
* </ul>
* A fleet where nothing is ever quarantined or weight-0 never exercises either path, so today's
* behaviour is unchanged.
*/
final class FixedPlacementPolicy implements PlacementPolicy {
@Override
public PlacementCandidate select(PlacementContext ctx) {
String d = ctx.defaultProfile();
if (d != null && !d.isBlank() && !ctx.quarantined().contains(d) && !weightExcluded(ctx, d)) {
return new PlacementCandidate(d, null, 1.0f, null);
}
for (PlacementCandidate c : ctx.candidates()) {
if (!ctx.quarantined().contains(c.profile()) && !c.excluded()) {
return new PlacementCandidate(c.profile(), null, c.weight(), c.maxLoad());
}
}
if (d != null && !d.isBlank()) {
boolean dQuarantined = ctx.quarantined().contains(d);
boolean dWeightExcluded = weightExcluded(ctx, d);
if (dQuarantined && dWeightExcluded) {
throw new PlacementException("worker profile '" + d + "' is quarantined (backend "
+ "exhausted) and has weight 0 (excluded from automatic selection), and no "
+ "available candidate remains");
}
if (dWeightExcluded) {
throw new PlacementException("worker profile '" + d + "' has weight 0 (excluded "
+ "from automatic selection) and no available candidate remains");
}
if (dQuarantined) {
throw new PlacementException("worker profile '" + d + "' is quarantined (backend "
+ "exhausted) and no un-quarantined candidate is available");
}
}
if (!ctx.candidates().isEmpty()) {
throw new PlacementException(
"all worker profiles are excluded from automatic selection (quarantined or weight-0)");
}
throw new PlacementException("no worker profiles configured");
}
/** Whether {@code profile} carries {@code weight <= 0} (CB-554) among {@code ctx}'s candidates. */
private static boolean weightExcluded(PlacementContext ctx, String profile) {
for (PlacementCandidate c : ctx.candidates()) {
if (c.profile().equals(profile)) {
return c.excluded();
}
}
return false;
}
}
@@ -1,4 +1,4 @@
package dev.ltms.fleet.placement;
package dev.ltms.bridged.placement;
/**
* A profile (and, in CB-308, a host) that can be chosen by a {@link PlacementPolicy}.
@@ -21,12 +21,12 @@ public record PlacementCandidate(String profile, String host, float weight, Inte
/**
* True when this candidate carries an explicit {@code weight <= 0} (CB-554) and must be
* skipped by every automatic policy — the same way a quarantined or unreachable candidate is
* skipped. {@code FleetConfig.Profile}'s compact constructor already normalises "absent" to
* skipped. {@code BridgedConfig.Profile}'s compact constructor already normalises "absent" to
* {@code 1.0} and "negative" to {@code 0.0}, so this is a plain threshold check here; it does
* not need to distinguish "explicit 0" from "absent" itself.
*
* <p>Exclusion is about <em>automatic</em> selection only — an explicit
* {@code fleet_spawn{profile:"..."}} bypasses placement entirely and is unaffected.
* {@code bridge_spawn{profile:"..."}} bypasses placement entirely and is unaffected.
*/
public boolean excluded() {
return weight <= 0.0f;
@@ -0,0 +1,23 @@
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.Set;
import java.util.function.Function;
/**
* Everything a {@link PlacementPolicy} needs to make one selection.
*
* @param defaultProfile profile a {@code fixed} policy should return (may be {@code null})
* @param candidates every configured candidate; the policy filters out those at cap or unreachable
* @param liveCount current live worker count per profile (from the session registry)
* @param unreachable profiles already known to have failed in this spawn attempt
* @param quarantined profiles whose credential is currently quarantined (CB-578 stage B) — a
* {@code BACKEND_EXHAUSTED} classification put it, or a profile it shares a
* credential with, on cooldown. Filtered the same way as {@code unreachable}.
*/
public record PlacementContext(String defaultProfile,
List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable,
Set<String> quarantined) {
}

Some files were not shown because too many files have changed in this diff Show More