Compare commits
79 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 7930a31b94 | |||
| 850fb12807 | |||
| b14b66ab03 | |||
| 15ff6bcde5 | |||
| 7822772905 | |||
| 0efe1567c0 | |||
| 837fed7690 | |||
| aa4ee64a34 | |||
| bb750cdba3 | |||
| d56c77b368 | |||
| 83e2ff06cf | |||
| f5deaafd06 | |||
| fe46311266 | |||
| 08968bb1b7 | |||
| 27bbd11f06 | |||
| 48d7841fbf | |||
| cc1df11f69 | |||
| fdfd4ac491 | |||
| d5dd5639ae | |||
| 7a120b3256 | |||
| 863d477966 | |||
| cec48832be | |||
| a36b7ccd7c | |||
| 32bf324a1e | |||
| 16de9df000 | |||
| 65c9deb4d1 | |||
| 28ae27b8e1 | |||
| 81a0cf4710 | |||
| 8a837a2830 | |||
| 0a2b3a4a56 | |||
| 613ece92dc | |||
| 8d4206c2b5 | |||
| 03286a589b | |||
| 88b9503c3b | |||
| c553d795d8 | |||
| 3ba6d6784c | |||
| 78ca24dc3f | |||
| 4fa6553db5 | |||
| 2124e043ce | |||
| e4f3620acb | |||
| 032a59a34d | |||
| e689090024 | |||
| 1cc34888fd | |||
| 0331ecd5d3 | |||
| 831a918c30 | |||
| 3db5277ae8 | |||
| 6939e0cbbc | |||
| f0095bf8b2 | |||
| 5206679efd | |||
| ac044e7573 | |||
| 6ebad2a91f | |||
| 3d10ed385c | |||
| 6d0c94dbdb | |||
| e01563a550 | |||
| 30e3225a3c | |||
| ef186a1516 | |||
| 5d5b3bdc76 | |||
| 2f8c98dac9 | |||
| 2db7189067 | |||
| 94476ac109 | |||
| 5fe02b7c98 | |||
| a1052f4fd1 | |||
| 8c9904a7c4 | |||
| 23af5dfdff | |||
| 3a10f6ad17 | |||
| a3842c873d | |||
| c29c3f063d | |||
| 758d62a396 | |||
| 6725642274 | |||
| a1f4dc365a | |||
| 165b62ee20 | |||
| 72d3481de3 | |||
| 29ccb747ad | |||
| 4749e27453 | |||
| e501d39988 | |||
| 6b6cf25862 | |||
| 180de840eb | |||
| 129dd4a838 | |||
| e09cac6f1f |
@@ -6,6 +6,14 @@
|
|||||||
# Settings backups inherit the env block — and secrets with it.
|
# Settings backups inherit the env block — and secrets with it.
|
||||||
.claude/settings.local.json.bak*
|
.claude/settings.local.json.bak*
|
||||||
|
|
||||||
|
# The default profile parityOverlay copies these primary→worktree, so they appear in EVERY worker
|
||||||
|
# worktree. Two reasons they must be ignored. They hold environment values, which is reason enough.
|
||||||
|
# And since CB-576 a release preserves any worktree that `git status --porcelain` calls dirty —
|
||||||
|
# untracked files included, deliberately. An untracked overlay file would therefore make every
|
||||||
|
# COMPLETED release preserve its worktree, and worktrees would pile up with no error to notice.
|
||||||
|
.env
|
||||||
|
.envrc
|
||||||
|
|
||||||
# Daemon runtime artefacts. bridged appends its log wherever it is launched from, so both the
|
# Daemon runtime artefacts. bridged appends its log wherever it is launched from, so both the
|
||||||
# repo root and bridged/ collect one; neither belongs in git.
|
# repo root and bridged/ collect one; neither belongs in git.
|
||||||
bridged.out
|
bridged.out
|
||||||
|
|||||||
@@ -80,10 +80,12 @@ below are the procedure — run them in order, every task, not only the big ones
|
|||||||
of your context, your plan, or your screen.
|
of your context, your plan, or your screen.
|
||||||
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{ticket, msgId}`. Answer a worker's `bridge_ask`
|
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{ticket, msgId}`. Answer a worker's `bridge_ask`
|
||||||
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
|
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
|
||||||
`bridge_status`, never by reading its terminal.
|
`bridge_status`, never by reading its terminal; it also reports an open question and the `turnId`
|
||||||
6. **Verify yourself.** Re-run the build and the checks. A worker mounts only the bridge MCP and
|
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
|
||||||
cannot run your other tooling, and a piped command (`… | tail`) hides failures behind a zero
|
brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
|
||||||
exit — never promote a worker's "clean" to a fact.
|
6. **Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
|
||||||
|
forge tools it appears to have hold a blocked credential and fail, and a piped command
|
||||||
|
(`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a fact.
|
||||||
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
|
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
|
||||||
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
|
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
|
||||||
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
|
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
|
||||||
@@ -105,8 +107,8 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
|
|||||||
|---|---|
|
|---|---|
|
||||||
| Confirm your own role | `bridge_whoami` |
|
| Confirm your own role | `bridge_whoami` |
|
||||||
| See backends available | `bridge_profiles` |
|
| See backends available | `bridge_profiles` |
|
||||||
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
|
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` → `sessionId` + `paneId` |
|
||||||
| See the fleet | `bridge_list` → `leads` (your peers) + `members` · one peer's state: `bridge_status{sessionId}` |
|
| See the fleet | `bridge_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) · one peer's state: `bridge_status{sessionId}` |
|
||||||
| Delegate (blocking) | `bridge_send{sessionId, content}` |
|
| Delegate (blocking) | `bridge_send{sessionId, content}` |
|
||||||
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
|
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
|
||||||
| Answer a member's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
|
| Answer a member's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
|
||||||
@@ -155,9 +157,13 @@ you.
|
|||||||
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
|
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
|
||||||
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
|
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
|
||||||
lead with its end cut off.
|
lead with its end cut off.
|
||||||
5. **Report honestly.** State only what you actually ran and its real output, including failures.
|
5. **Report honestly.** State only what you actually ran and its real output, including failures,
|
||||||
You mount **only** the bridge MCP — the primary's other servers (IDE, forge, docs) are not yours,
|
and never claim the result of a check you had no way to run. **Measure your own tools; do not
|
||||||
so never claim the result of a check you had no way to run.
|
assume them.** What you mount depends on your backend: an opencode member gets the bridge and
|
||||||
|
nothing else, while a Claude Code member also inherits the operator's user-scope MCP servers,
|
||||||
|
which the bridge never chose for you. Two rules follow. The primary's IDE tooling is still not
|
||||||
|
yours, whatever you see. And **a mounted tool is not a working tool** — the forge server you may
|
||||||
|
find there holds a deliberately blocked credential and fails every call, by design.
|
||||||
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
|
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
|
||||||
project marks as not-yours-to-commit.
|
project marks as not-yours-to-commit.
|
||||||
|
|
||||||
@@ -189,6 +195,64 @@ charter, not here.
|
|||||||
are diagrammed in `docs/MCP-Contract.md` §6, kept out of this file because it loads into every
|
are diagrammed in `docs/MCP-Contract.md` §6, kept out of this file because it loads into every
|
||||||
session's context.
|
session's context.
|
||||||
|
|
||||||
|
### Redeploying the daemon — the lead may do this (primary only)
|
||||||
|
|
||||||
|
**A merge is not a deployment.** The running `bridged` holds the jar it was started with, so a
|
||||||
|
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
|
||||||
|
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
|
||||||
|
rather than hand the job back to the operator.
|
||||||
|
|
||||||
|
Workers must never do this. A worker has no business restarting the daemon it is talking through,
|
||||||
|
and stopping it kills the worker's own channel mid-turn.
|
||||||
|
|
||||||
|
**Use the script — do not hand-roll the steps.**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
scripts/redeploy-bridged.sh --check # report state, change nothing
|
||||||
|
scripts/redeploy-bridged.sh # build, confirm drain, restart, verify
|
||||||
|
scripts/redeploy-bridged.sh --yes # skip the drain prompt (fleet already checked)
|
||||||
|
```
|
||||||
|
|
||||||
|
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
|
||||||
|
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
|
||||||
|
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
|
||||||
|
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
|
||||||
|
|
||||||
|
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
|
||||||
|
if the script is unavailable or a step fails, this is what it was protecting you from.
|
||||||
|
|
||||||
|
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
|
||||||
|
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
|
||||||
|
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
|
||||||
|
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
|
||||||
|
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
|
||||||
|
whether the name resolves without ever printing the value.
|
||||||
|
2. **Drain live members first.** `bridge_list`, then `bridge_stop` each member, and collect anything
|
||||||
|
you still want with `bridge_poll` before you kill anything. A restart drops in-flight tickets and
|
||||||
|
rendezvous, and a member's report is not recoverable once its ticket is gone.
|
||||||
|
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
|
||||||
|
it. The startup log names which keys it accepted and which it deferred — read those lines rather
|
||||||
|
than assuming.
|
||||||
|
4. **Re-check identity afterwards.** Call `bridge_whoami` and confirm it still answers `primary`. The
|
||||||
|
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
|
||||||
|
demoted to worker, which refuses every orchestration call.
|
||||||
|
5. **Prove the new jar is the one running.** Confirm a *fresh* `bridged listening` line at the end of
|
||||||
|
`bridged/bridged.out`, dated after the restart. An old daemon that never died looks identical from
|
||||||
|
the outside.
|
||||||
|
|
||||||
|
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
|
||||||
|
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
|
||||||
|
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
|
||||||
|
start every time. The rule lives in the operator's Claude Code settings:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{ "permissions": { "allow": ["Bash(scripts/redeploy-bridged.sh:*)"] } }
|
||||||
|
```
|
||||||
|
|
||||||
|
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
|
||||||
|
running the stop and start as separate commands — that is exactly the approval the script replaced.
|
||||||
|
Say what you were going to run and why, and let the operator decide.
|
||||||
|
|
||||||
### The prompt is part of the product — update it with the code (mandatory)
|
### The prompt is part of the product — update it with the code (mandatory)
|
||||||
|
|
||||||
This repo *is* the bridge, so the canonical block above is not documentation about someone else's
|
This repo *is* the bridge, so the canonical block above is not documentation about someone else's
|
||||||
|
|||||||
+189
-18
@@ -88,15 +88,28 @@ bind:
|
|||||||
# backoffMs: 60000
|
# backoffMs: 60000
|
||||||
# quietNudgeCap: 3
|
# quietNudgeCap: 3
|
||||||
|
|
||||||
# Fleet health detection is dormant unless enabled. It reads one whole-fleet agent list per tick.
|
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
|
||||||
# It can run without a webhook; bridge_list then reports healthCoverage: detection-only.
|
# per tick.
|
||||||
|
# intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes
|
||||||
|
# Math.max(15, intervalSeconds), so a lower value is silently raised, not
|
||||||
|
# rejected.
|
||||||
|
# workingSuspectAfterSeconds, paneProbeIntervalSeconds → accepted and parsed, but NOT YET READ by
|
||||||
|
# anything — the dormant monitor only consumes intervalSeconds today (CB-573
|
||||||
|
# shipped ahead of the evidence publishers these two knobs are for). Setting
|
||||||
|
# them changes nothing right now, and no minimum is enforced on either, because
|
||||||
|
# nothing reads them to enforce one. They exist so a later build can start
|
||||||
|
# honouring them without another config-shape change.
|
||||||
|
# notifications.mode → "webhook" flips what bridge_list REPORTS (healthCoverage: "full" instead
|
||||||
|
# of "detection-only") — it does NOT make bridged send any webhook call; no
|
||||||
|
# delivery mechanism is implemented yet. Any other value, or omitting the
|
||||||
|
# block, reports "detection-only".
|
||||||
# health:
|
# health:
|
||||||
# enabled: true
|
# enabled: true
|
||||||
# intervalSeconds: 30 # minimum 15
|
# intervalSeconds: 30
|
||||||
# workingSuspectAfterSeconds: 600 # minimum 300
|
# workingSuspectAfterSeconds: 600
|
||||||
# paneProbeIntervalSeconds: 60 # minimum 60
|
# paneProbeIntervalSeconds: 60
|
||||||
# notifications:
|
# notifications:
|
||||||
# mode: disabled # disabled (default) or webhook
|
# mode: disabled
|
||||||
|
|
||||||
# herdr Unix socket. Omit to use the client default
|
# herdr Unix socket. Omit to use the client default
|
||||||
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
|
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
|
||||||
@@ -144,6 +157,17 @@ herdrSocket: ~/.config/herdr/herdr.sock
|
|||||||
# omit and this profile's completion fallback behaves exactly as before.
|
# omit and this profile's completion fallback behaves exactly as before.
|
||||||
# Every backend words its refusal differently, so this is config, never a
|
# Every backend words its refusal differently, so this is config, never a
|
||||||
# vendor string baked into bridged itself.
|
# vendor string baked into bridged itself.
|
||||||
|
# DEFERRED: compiled once into a startup pattern map — editing it needs a
|
||||||
|
# daemon restart, same as this profile's model/baseUrl/argv.
|
||||||
|
# credentialId → CB-578 stage B: the credential this profile quarantines WITH when a
|
||||||
|
# BACKEND_EXHAUSTED classification fires. Two profiles that set the SAME
|
||||||
|
# credentialId share one quarantine — the case this exists for is two models
|
||||||
|
# on one account (e.g. sol and terra both billing one OpenAI credential): an
|
||||||
|
# exhaustion on either one must lock out both, or the fleet just walks onto
|
||||||
|
# the same dead account under the sibling's name. Opt-in — omit and this
|
||||||
|
# profile quarantines alone, under its own name, exactly as if the field did
|
||||||
|
# not exist. Cooldown length is the top-level quarantineCooldownSeconds below.
|
||||||
|
# HOT: read live at every spawn/exhaustion check — no restart needed.
|
||||||
# env → extra environment for this profile's workers, as a literal key/value map
|
# env → extra environment for this profile's workers, as a literal key/value map
|
||||||
# (CB-511). Use it to give workers a toolchain.
|
# (CB-511). Use it to give workers a toolchain.
|
||||||
#
|
#
|
||||||
@@ -173,11 +197,21 @@ profiles:
|
|||||||
mcpUrl: http://127.0.0.1:8765/mcp
|
mcpUrl: http://127.0.0.1:8765/mcp
|
||||||
tokenEnv: BRIDGED_WORKER_TOKEN
|
tokenEnv: BRIDGED_WORKER_TOKEN
|
||||||
argv: ["ccs", "gx10"]
|
argv: ["ccs", "gx10"]
|
||||||
weight: 0.5 # relative selection weight for placement: weighted
|
# weight: relative selection weight for automatic placement (weighted, round-robin, and
|
||||||
maxLoad: 2 # max live workers on this profile (omit for unlimited)
|
# fixed's fallback walk). Absent defaults to 1.0. An explicit 0 or negative value means
|
||||||
|
# "never auto-select this profile" (CB-554) — it stays reachable via an explicit
|
||||||
|
# `bridge_spawn{profile:"gx10"}`, which bypasses placement entirely; only automatic
|
||||||
|
# selection skips it.
|
||||||
|
weight: 0.5
|
||||||
|
# maxLoad: max live workers on this profile. Omit for unlimited. An explicit 0 (CB-585) caps
|
||||||
|
# the profile at zero live members — it is excluded from automatic placement and an explicit
|
||||||
|
# `bridge_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile
|
||||||
|
# is named directly. Negative is refused at config load — there is no sane meaning for it.
|
||||||
|
maxLoad: 2
|
||||||
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
|
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
|
||||||
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
|
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
|
||||||
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
|
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
|
||||||
|
# credentialId: shared-openai # opt-in: quarantine together with every other profile sharing this id (CB-578)
|
||||||
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
|
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
|
||||||
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
|
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
|
||||||
# parityOverlay: [".claude/settings.local.json", ".env", ".envrc"] # never add .mcp.json — see above
|
# parityOverlay: [".claude/settings.local.json", ".env", ".envrc"] # never add .mcp.json — see above
|
||||||
@@ -241,8 +275,37 @@ profiles:
|
|||||||
# argv: ["opencode"]
|
# argv: ["opencode"]
|
||||||
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
|
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
|
||||||
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
|
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
|
||||||
|
#
|
||||||
|
# `weighted` IS NOT "cheapest first" — read this before you set weights (CB-589).
|
||||||
|
# It is smooth weighted round-robin: it spreads spawns across EVERY profile that has a free slot,
|
||||||
|
# in weight ratio. It has no idea which profile costs money. So with local:10 / paid:2 you do not
|
||||||
|
# get "use local, overflow to paid" — you get roughly one spawn in six going to the paid profile
|
||||||
|
# while the local box still has a free slot.
|
||||||
|
#
|
||||||
|
# There is a sharper second effect. The policy's running score map lives for the daemon's whole
|
||||||
|
# life. While a profile is at maxLoad it is filtered out and its score FREEZES, so the paid
|
||||||
|
# profiles keep accumulating against it. When the local slot frees up it returns with a stale
|
||||||
|
# score and can LOSE the next pick — a paid spawn while the free box sits idle.
|
||||||
|
#
|
||||||
|
# Until a real cost-first policy exists, the workaround is to make the ratio decisive rather than
|
||||||
|
# proportional: give the free profile a weight so large that it wins every pick it is eligible
|
||||||
|
# for, and paid profiles only ever take genuine overflow. On this host that is local weight 100
|
||||||
|
# against paid weights of ~1.
|
||||||
|
#
|
||||||
|
# The gotcha with that workaround: it expresses a PREFERENCE ORDER through a RATIO knob. Add a
|
||||||
|
# future profile at weight 150 and it silently outranks the free box, with nothing to warn you.
|
||||||
|
# Re-check the weights whenever you add a profile.
|
||||||
placement: weighted
|
placement: weighted
|
||||||
|
|
||||||
|
# How long a credential sits out after a BACKEND_EXHAUSTED classification (CB-578 stage B), in
|
||||||
|
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
|
||||||
|
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
|
||||||
|
# 1800 (30 minutes) when omitted or non-positive.
|
||||||
|
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
|
||||||
|
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
|
||||||
|
# restart. Editing this needs a daemon restart to take effect.
|
||||||
|
# quarantineCooldownSeconds: 1800
|
||||||
|
|
||||||
# Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an
|
# Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an
|
||||||
# upgraded bridged keeps the old behaviour: the file is read once at boot and never again.
|
# upgraded bridged keeps the old behaviour: the file is read once at boot and never again.
|
||||||
# enabled → turn the watch on. bridged checks the file's modified time on a timer and
|
# enabled → turn the watch on. bridged checks the file's modified time on a timer and
|
||||||
@@ -252,17 +315,25 @@ placement: weighted
|
|||||||
# Not every key can move under a running daemon, and the difference is about what already exists
|
# Not every key can move under a running daemon, and the difference is about what already exists
|
||||||
# when the reload happens — not about how important the key is:
|
# when the reload happens — not about how important the key is:
|
||||||
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
|
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
|
||||||
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad. Those are
|
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad
|
||||||
# hot because the placement policy reads them through a supplier — being config is
|
# / credentialId. Those are hot because the placement policy (and, for credentialId,
|
||||||
|
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
|
||||||
# not by itself enough to make a key hot.
|
# not by itself enough to make a key hot.
|
||||||
|
# EXCEPT `fleet.leaders`: Bridged.main reads it once at startup to build the lead tab
|
||||||
|
# scanner and launcher, and neither is rebuilt on reload. A changed/added/removed
|
||||||
|
# `fleet.leaders` entry is silently accepted — the reload reports "config reloaded"
|
||||||
|
# with nothing in the deferred list — but has NO effect until you restart. Treat it
|
||||||
|
# as deferred in practice, even though today's reload output does not say so.
|
||||||
# DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value
|
# DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value
|
||||||
# until you restart: `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`,
|
# until you restart: `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`,
|
||||||
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, ADDING or REMOVING a profile (a new
|
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, `quarantineCooldownSeconds` (CB-578
|
||||||
# backend needs its own launcher, and launchers are built once), AND an existing
|
# stage B — baked once into the quarantine tracker built at startup), ADDING or
|
||||||
# profile's launch settings — model, baseUrl, argv, env, configDir, mcpUrl, tabLabel.
|
# REMOVING a profile (a new backend needs its own launcher, and launchers are built
|
||||||
# The launcher takes a copy of `profiles:` at startup and resolves every spawn out of
|
# once), AND an existing profile's launch settings — model, baseUrl, argv, env,
|
||||||
# that copy, so those never reach a launch until you restart. The reload logs them by
|
# configDir, mcpUrl, tabLabel, exhaustedPattern. The launcher takes a copy of
|
||||||
# name rather than pretending they applied.
|
# `profiles:` at startup and resolves every spawn out of that copy, so those never
|
||||||
|
# reach a launch until you restart. The reload logs them by name rather than
|
||||||
|
# pretending they applied.
|
||||||
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
|
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
|
||||||
# bound, the broker connection is open, and the auth mode decides who may reach the
|
# bound, the broker connection is open, and the auth mode decides who may reach the
|
||||||
# port that is already listening.
|
# port that is already listening.
|
||||||
@@ -335,6 +406,12 @@ fleet:
|
|||||||
# An auto-launched lead is NOT a member: it gets no worker reply charter, is never registered with
|
# An auto-launched lead is NOT a member: it gets no worker reply charter, is never registered with
|
||||||
# the session lifecycle (the idle reaper would kill your orchestrator), and stays on the
|
# the session lifecycle (the idle reaper would kill your orchestrator), and stays on the
|
||||||
# subscription — ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says.
|
# subscription — ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says.
|
||||||
|
#
|
||||||
|
# GET THE `tab:` VALUE RIGHT. A pane that does not match any configured `tab:` (a typo, a renamed
|
||||||
|
# tab, a pane no entry names at all) is not recognised as a lead — it resolves as an ordinary
|
||||||
|
# WORKER instead, silently, and every orchestration call it makes (spawn/stop/send/drain) is
|
||||||
|
# refused. There is no error at startup for this: an unmatched pane is simply not a lead. If your
|
||||||
|
# primary suddenly can't spawn or send, check this section first.
|
||||||
# leaders:
|
# leaders:
|
||||||
# opus-5.0:
|
# opus-5.0:
|
||||||
# profile: opus # omit to never create this lead, only recognise it
|
# profile: opus # omit to never create this lead, only recognise it
|
||||||
@@ -373,6 +450,91 @@ guard:
|
|||||||
- gx00.gw
|
- gx00.gw
|
||||||
- gx01.gw
|
- gx01.gw
|
||||||
|
|
||||||
|
# Member credential policy (CB-596, gitea issue #82). A herdr pane runs a LOGIN shell, and that
|
||||||
|
# shell re-sources the operator's own secret store — so a spawned member inherits every credential
|
||||||
|
# the operator's shell holds, not just the ones bridged means to give it. Measured on this host:
|
||||||
|
# 31 credential names, all set, with only ONE (GITEA_ACCESS_TOKEN) blocked before this — and that
|
||||||
|
# block was a single name hardcoded in HerdrPeerLauncher.java, not driven by this file. This block
|
||||||
|
# replaces that hardcoded shadow with a config-driven list of names.
|
||||||
|
#
|
||||||
|
# ROUND-2 CORRECTION, measured live: the pane-creation env overlay below (applied at tab.create /
|
||||||
|
# pane.split, BEFORE the pane's login shell runs) does NOT survive that login shell for any name
|
||||||
|
# secrets.sh actually exports — the shell re-exports it afterwards and overwrites the sentinel.
|
||||||
|
# Proof: GITEA_ACCESS_TOKEN comes back blocked only because secrets.sh itself carries a guarded
|
||||||
|
# export (`[ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...`) — that guard, not this
|
||||||
|
# file, is what wins. No other name in `known` below has a matching guard in secrets.sh yet (1
|
||||||
|
# guard measured against 33 export lines there). So today this block's overlay is REAL protection
|
||||||
|
# only for a name secrets.sh does not export, or a peer kind whose pane never runs a login shell —
|
||||||
|
# for everything secrets.sh exports and guards, the guard in secrets.sh (out of scope for this
|
||||||
|
# ticket) is what actually blocks it, not this list. An exec-time fix (winning after the login
|
||||||
|
# shell finishes, before the agent process starts) was attempted and found to have no seam in the
|
||||||
|
# current herdr protocol — AgentControl.start takes a fixed `kind` (herdr resolves the executable)
|
||||||
|
# plus trailing CLI args for that binary, not an arbitrary argv or an env map; only tab.create /
|
||||||
|
# pane.split accept `env`, and that is this same pane-creation overlay. See gitea #82 for the open
|
||||||
|
# design question this leaves.
|
||||||
|
#
|
||||||
|
# DENY-BY-DEFAULT, NOT A DENY-LIST. A deny-list (name the bad ones, let everything else through) is
|
||||||
|
# silently wrong the moment the operator's store gains a new secret — nothing would ever report it.
|
||||||
|
# Deny-by-default inverts that: `known` bounds the blast radius to names actually enumerated below,
|
||||||
|
# and EVERY one of them is blocked UNLESS it is also in `allow`. Omitting this block entirely (the
|
||||||
|
# shipped default) blocks NOTHING — unlike most optional blocks in this file, absence here is a real
|
||||||
|
# gap, not a safe "feature off". A name that is neither `known` nor `allow`-ed is not silently let
|
||||||
|
# through either: the daemon logs a WARN naming any credential-shaped env var it finds on neither
|
||||||
|
# list (never its value), so a secret added to the store later does not go unnoticed forever.
|
||||||
|
#
|
||||||
|
# policy → only "deny-by-default" exists today (an operator-authored deny-list was deliberately
|
||||||
|
# rejected — see above). An unrecognized value refuses to start, naming it.
|
||||||
|
# allow → credential names a member legitimately needs. Left OUT of the pane's env overlay
|
||||||
|
# entirely, so the value the pane's own (login) shell exports passes through untouched.
|
||||||
|
# known → every credential name the operator's store is known to export. Every name here NOT
|
||||||
|
# also in `allow` is overlaid with a non-secret sentinel value before the pane's login
|
||||||
|
# shell runs — real protection only for names the login shell does not itself re-export
|
||||||
|
# (see the ROUND-2 CORRECTION note above for the ones it does).
|
||||||
|
#
|
||||||
|
# HOT-RELOADABLE the same way `fleet:` is (CB-559): read fresh on every spawn, so editing this list
|
||||||
|
# and reloading config (or restarting) changes what the NEXT spawn inherits; already-running members
|
||||||
|
# are unaffected either way.
|
||||||
|
# memberCredentials:
|
||||||
|
# policy: deny-by-default
|
||||||
|
# allow:
|
||||||
|
# - AI_GATEWAY_TOKEN # named in a profile's tokenEnv (local/gx) — a member reaching the
|
||||||
|
# # gateway is by design, not a leak
|
||||||
|
# - WORKER_GITEA_TOKEN # the repo-scoped forge token a member needs to open its own PR (CB-302)
|
||||||
|
# - CONTEXT7_TOKEN # already decided as allowed by CB-593
|
||||||
|
# - GITEA_HOST # not a credential — a hostname, paired with the forge token above
|
||||||
|
# known:
|
||||||
|
# - AI_GATEWAY_TOKEN
|
||||||
|
# - BESZEL_ADMIN_EMAIL
|
||||||
|
# - BESZEL_ADMIN_PASSWORD
|
||||||
|
# - BESZEL_HUB_URL
|
||||||
|
# - BESZEL_KEY
|
||||||
|
# - BESZEL_UNIVERSAL_TOKEN
|
||||||
|
# - BRAIN_MCP_TOKEN
|
||||||
|
# - CF_ACCOUNT_ID
|
||||||
|
# - CF_API_TOKEN
|
||||||
|
# - CF_USER_TOKEN
|
||||||
|
# - CONFLUENCE_API_TOKEN
|
||||||
|
# - CONFLUENCE_USERNAME
|
||||||
|
# - CONTEXT7_TOKEN
|
||||||
|
# - GITEA_ACCESS_TOKEN
|
||||||
|
# - GITEA_HOST
|
||||||
|
# - GITLAB_OAUTH_CLIENT_SECRET
|
||||||
|
# - GITLAB_PERSONAL_ACCESS_TOKEN
|
||||||
|
# - GRAFANA_ADMIN_PASSWORD
|
||||||
|
# - GRAFANA_ADMIN_USER
|
||||||
|
# - HASS_TOKEN
|
||||||
|
# - HW_PASSWORD
|
||||||
|
# - HW_USER
|
||||||
|
# - LTMS_API_KEY
|
||||||
|
# - MEMORY_MCP_TOKEN
|
||||||
|
# - METRICS_PUSH_TOKEN
|
||||||
|
# - OPENCODE_AUTOMODE_MODEL
|
||||||
|
# - TELEGRAM_BOT_TOKEN
|
||||||
|
# - TELEGRAM_CHAT_ID
|
||||||
|
# - TS_API_KEY
|
||||||
|
# - TS_AUTHKEY
|
||||||
|
# - WORKER_GITEA_TOKEN
|
||||||
|
|
||||||
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
|
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
|
||||||
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
|
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
|
||||||
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
|
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
|
||||||
@@ -390,10 +552,15 @@ guard:
|
|||||||
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
|
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
|
||||||
# contextCap → force-release a session after this many delegated turns
|
# contextCap → force-release a session after this many delegated turns
|
||||||
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
|
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
|
||||||
|
# clearAfterTurn → whether a reusable worker discards its conversation context after every
|
||||||
|
# completed delegated turn (default false). Works for claude-code workers
|
||||||
|
# only — any other peer kind (e.g. opencode) logs "context reset is
|
||||||
|
# unsupported for peer kind …" once and the reset is a no-op.
|
||||||
# lifecycle:
|
# lifecycle:
|
||||||
# idleTtlSeconds: 300
|
# idleTtlSeconds: 300
|
||||||
# contextCap: 10
|
# contextCap: 10
|
||||||
# drainTimeoutSeconds: 5
|
# drainTimeoutSeconds: 5
|
||||||
|
# clearAfterTurn: false
|
||||||
|
|
||||||
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
|
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
|
||||||
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
|
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
|
||||||
@@ -401,10 +568,14 @@ guard:
|
|||||||
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
|
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
|
||||||
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
|
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
|
||||||
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
|
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
|
||||||
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
|
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
|
||||||
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
|
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
|
||||||
|
# prefetch → CB-527: consumer basicQos, capping how many unacked messages the inbox holds
|
||||||
|
# in-heap per owned target (the rest sits on the broker's durable queue instead of
|
||||||
|
# growing the JVM heap). Default 32 when omitted.
|
||||||
# broker:
|
# broker:
|
||||||
# uri: amqp://guest:guest@127.0.0.1:5672
|
# uri: amqp://guest:guest@127.0.0.1:5672
|
||||||
|
# prefetch: 32
|
||||||
|
|
||||||
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
|
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
|
||||||
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
|
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
|
||||||
|
|||||||
@@ -14,6 +14,7 @@ import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
|
|||||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||||
import dev.ltms.bridged.inject.CompletionResolver;
|
import dev.ltms.bridged.inject.CompletionResolver;
|
||||||
import dev.ltms.bridged.inject.ExhaustedPatternLookup;
|
import dev.ltms.bridged.inject.ExhaustedPatternLookup;
|
||||||
|
import dev.ltms.bridged.inject.ExhaustionSink;
|
||||||
import dev.ltms.bridged.inject.Injector;
|
import dev.ltms.bridged.inject.Injector;
|
||||||
import dev.ltms.bridged.inject.StatusPoller;
|
import dev.ltms.bridged.inject.StatusPoller;
|
||||||
import dev.ltms.bridged.inject.TurnListener;
|
import dev.ltms.bridged.inject.TurnListener;
|
||||||
@@ -37,6 +38,7 @@ import dev.ltms.bridged.msg.LeadHeartbeatLoop;
|
|||||||
import dev.ltms.bridged.msg.ReplyPushLoop;
|
import dev.ltms.bridged.msg.ReplyPushLoop;
|
||||||
import dev.ltms.bridged.rest.BridgedApp;
|
import dev.ltms.bridged.rest.BridgedApp;
|
||||||
import dev.ltms.bridged.session.GitWorktrees;
|
import dev.ltms.bridged.session.GitWorktrees;
|
||||||
|
import dev.ltms.bridged.session.MemberSession;
|
||||||
import dev.ltms.bridged.session.SessionManager;
|
import dev.ltms.bridged.session.SessionManager;
|
||||||
import dev.ltms.bridged.peer.PeerLauncher;
|
import dev.ltms.bridged.peer.PeerLauncher;
|
||||||
import dev.ltms.bridged.session.SessionReaper;
|
import dev.ltms.bridged.session.SessionReaper;
|
||||||
@@ -44,6 +46,7 @@ import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
|||||||
import dev.ltms.bridged.member.CompositePeerLauncher;
|
import dev.ltms.bridged.member.CompositePeerLauncher;
|
||||||
import dev.ltms.bridged.member.HerdrPeerLauncher;
|
import dev.ltms.bridged.member.HerdrPeerLauncher;
|
||||||
import dev.ltms.bridged.member.OpenCodeLauncher;
|
import dev.ltms.bridged.member.OpenCodeLauncher;
|
||||||
|
import dev.ltms.bridged.placement.BackendQuarantine;
|
||||||
import io.javalin.Javalin;
|
import io.javalin.Javalin;
|
||||||
import org.slf4j.Logger;
|
import org.slf4j.Logger;
|
||||||
import org.slf4j.LoggerFactory;
|
import org.slf4j.LoggerFactory;
|
||||||
@@ -80,6 +83,15 @@ public final class Bridged {
|
|||||||
static void main(String[] args) {
|
static void main(String[] args) {
|
||||||
Path configPath = Path.of(args.length > 0 ? args[0] : "bridged.yaml");
|
Path configPath = Path.of(args.length > 0 ? args[0] : "bridged.yaml");
|
||||||
BridgedConfig cfg = BridgedConfig.load(configPath);
|
BridgedConfig cfg = BridgedConfig.load(configPath);
|
||||||
|
// CB-594: report which secret env vars the config actually needs, by name, before anything
|
||||||
|
// else can fail on a silently-empty one. A daemon started without a login shell (launchd)
|
||||||
|
// boots fine either way — this is the only thing that says so out loud.
|
||||||
|
reportRequiredSecrets(cfg);
|
||||||
|
// CB-596: an absent (or empty) memberCredentials: block blocks NOTHING — no credential
|
||||||
|
// name is hardcoded any more to fall back on. Say so loudly, the same way a missing
|
||||||
|
// secret is reported above, so upgrading past this commit never silently drops CB-592's
|
||||||
|
// protection.
|
||||||
|
reportMemberCredentialsGap(cfg);
|
||||||
// CB-559: `cfg` stays the startup snapshot — every validation and every piece of one-time
|
// CB-559: `cfg` stays the startup snapshot — every validation and every piece of one-time
|
||||||
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
|
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
|
||||||
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
|
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
|
||||||
@@ -132,20 +144,29 @@ public final class Bridged {
|
|||||||
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
|
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
|
||||||
claudeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
|
claudeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
|
||||||
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
|
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
|
||||||
() -> config.get().fleet()));
|
() -> config.get().fleet(),
|
||||||
|
() -> config.get().memberCredentials()));
|
||||||
}
|
}
|
||||||
if (!opencodeProfiles.isEmpty()) {
|
if (!opencodeProfiles.isEmpty()) {
|
||||||
adapters.add(new OpenCodeLauncher(agents, spaces,
|
adapters.add(new OpenCodeLauncher(agents, spaces,
|
||||||
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
|
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
|
||||||
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
|
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
|
||||||
() -> config.get().fleet()));
|
() -> config.get().fleet(),
|
||||||
|
() -> config.get().memberCredentials()));
|
||||||
}
|
}
|
||||||
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
|
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
|
||||||
|
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
|
||||||
|
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
|
||||||
|
// The cooldown is deferred (see BridgedConfig#quarantineCooldownSeconds): it is read once
|
||||||
|
// here, at startup, and a config reload only changes it for a daemon restart.
|
||||||
|
BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,
|
||||||
|
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
|
||||||
PeerLauncher workers = new CompositePeerLauncher(
|
PeerLauncher workers = new CompositePeerLauncher(
|
||||||
adapters,
|
adapters,
|
||||||
cfg.effectiveDefaultProfile(),
|
cfg.effectiveDefaultProfile(),
|
||||||
config,
|
config,
|
||||||
profileName -> liveCountRef.get().apply(profileName));
|
profileName -> liveCountRef.get().apply(profileName),
|
||||||
|
quarantine);
|
||||||
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
|
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
|
||||||
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
|
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
|
||||||
// the first thing that actually talks to herdr, so without this wait a boot-order race
|
// the first thing that actually talks to herdr, so without this wait a boot-order race
|
||||||
@@ -272,7 +293,23 @@ public final class Bridged {
|
|||||||
.orElse(null);
|
.orElse(null);
|
||||||
log.info("backend-exhausted classification (CB-578 stage A): {}",
|
log.info("backend-exhausted classification (CB-578 stage A): {}",
|
||||||
CompletionResolver.coverage(cfg.profiles().keySet(), exhaustedPatternsByProfile.keySet()));
|
CompletionResolver.coverage(cfg.profiles().keySet(), exhaustedPatternsByProfile.keySet()));
|
||||||
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns);
|
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
|
||||||
|
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
|
||||||
|
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
|
||||||
|
// the profile config live off `config`, so a credentialId edit is hot: no restart needed.
|
||||||
|
ExhaustionSink exhaustionSink = (target, reason) -> sessions.roster().stream()
|
||||||
|
.filter(session -> target.equals(session.terminalId()))
|
||||||
|
.findFirst()
|
||||||
|
.map(MemberSession::profile)
|
||||||
|
.map(profileName -> config.get().profiles().get(profileName))
|
||||||
|
.ifPresent(profile -> {
|
||||||
|
String credentialId = profile.effectiveCredentialId();
|
||||||
|
quarantine.quarantine(credentialId);
|
||||||
|
log.warn("credential '{}' quarantined for {}s (profile '{}' classified "
|
||||||
|
+ "BACKEND_EXHAUSTED): {}", credentialId,
|
||||||
|
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
|
||||||
|
});
|
||||||
|
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns, exhaustionSink);
|
||||||
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
|
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
|
||||||
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
|
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
|
||||||
MemberPresence presence = sessions.asPresence();
|
MemberPresence presence = sessions.asPresence();
|
||||||
@@ -322,8 +359,9 @@ public final class Bridged {
|
|||||||
// connection, so keep the reference to close it in the ordered shutdown hook.
|
// connection, so keep the reference to close it in the ordered shutdown hook.
|
||||||
final ReplyInbox replyInbox;
|
final ReplyInbox replyInbox;
|
||||||
if (cfg.broker() != null && cfg.broker().isConfigured()) {
|
if (cfg.broker() != null && cfg.broker().isConfigured()) {
|
||||||
replyInbox = AmqpReplyInbox.open(cfg.broker().uri());
|
replyInbox = AmqpReplyInbox.open(cfg.broker().uri(), cfg.broker().prefetchOrDefault());
|
||||||
log.info("reply inbox: AMQP broker (durable) at {}", cfg.broker().uri());
|
log.info("reply inbox: AMQP broker (durable) at {} (prefetch={})",
|
||||||
|
cfg.broker().uri(), cfg.broker().prefetchOrDefault());
|
||||||
} else {
|
} else {
|
||||||
replyInbox = new InMemoryReplyInbox();
|
replyInbox = new InMemoryReplyInbox();
|
||||||
log.info("reply inbox: in-memory (soft-state)");
|
log.info("reply inbox: in-memory (soft-state)");
|
||||||
@@ -402,10 +440,22 @@ public final class Bridged {
|
|||||||
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
|
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
|
||||||
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
|
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
|
||||||
// reached /metrics — the delegation was unresolvable and nothing said so.
|
// reached /metrics — the delegation was unresolvable and nothing said so.
|
||||||
sessions.onRelease(terminal -> {
|
sessions.onRelease(detail -> {
|
||||||
messages.abandon(terminal, "the worker session was released before it replied");
|
// CB-578 stage C, acceptance criterion 10: a failed ticket's detail should tell a lead
|
||||||
replyInbox.release(terminal);
|
// where to re-dispatch onto the same tree, not just that the worker vanished.
|
||||||
primaryRegistry.forgetDelegation(terminal); // CB-532: don't leak the lead binding
|
String reason = "the worker session was released before it replied";
|
||||||
|
if (detail.worktreePath() != null) {
|
||||||
|
reason += "; worktree=" + detail.worktreePath() + " branch=" + detail.branch()
|
||||||
|
+ " snapshot=" + (detail.snapshotRef() != null ? detail.snapshotRef() : "none");
|
||||||
|
}
|
||||||
|
// CB-584 (issue #65 criterion 5): also name the agent session, so a lead can resume the
|
||||||
|
// member's conversation instead of only re-dispatching a fresh one onto the same files.
|
||||||
|
if (detail.agentSessionId() != null) {
|
||||||
|
reason += " agentSessionId=" + detail.agentSessionId();
|
||||||
|
}
|
||||||
|
messages.abandon(detail.terminalId(), reason);
|
||||||
|
replyInbox.release(detail.terminalId());
|
||||||
|
primaryRegistry.forgetDelegation(detail.terminalId()); // CB-532: don't leak the lead binding
|
||||||
});
|
});
|
||||||
|
|
||||||
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
|
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
|
||||||
@@ -440,7 +490,11 @@ public final class Bridged {
|
|||||||
var health = config.get().health();
|
var health = config.get().health();
|
||||||
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
|
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
|
||||||
health != null && health.notifications() != null && health.notifications().configured());
|
health != null && health.notifications() != null && health.notifications().configured());
|
||||||
}));
|
}),
|
||||||
|
new BridgeMcp.QuarantineSource(profile -> {
|
||||||
|
var configured = config.get().profiles().get(profile);
|
||||||
|
return configured == null ? null : configured.effectiveCredentialId();
|
||||||
|
}, quarantine));
|
||||||
|
|
||||||
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed, so an
|
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed, so an
|
||||||
// upgraded daemon behaves exactly as before — the file is read once at boot and never again.
|
// upgraded daemon behaves exactly as before — the file is read once at boot and never again.
|
||||||
@@ -506,6 +560,91 @@ public final class Bridged {
|
|||||||
return target -> presence.isPresent(target) || leads.get().containsKey(target);
|
return target -> presence.isPresent(target) || leads.get().containsKey(target);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-594: which env vars the loaded config actually needs, and why — every non-{@code
|
||||||
|
* subscription} profile's {@code tokenEnv} (a subscription profile never reads one, see
|
||||||
|
* {@link BridgedConfig.Profile#isSubscription()}), plus every profile's {@code gitTokenEnv}
|
||||||
|
* where set (opt-in). Derived from the config, not hard-coded, so a new profile is covered for
|
||||||
|
* free. A var required by more than one profile is one entry naming every profile that needs
|
||||||
|
* it. Deliberately excludes {@code auth.tokenEnv}: that one is already enforced loudly, by a
|
||||||
|
* startup throw in {@code main()} — about 370 lines <em>below</em> this method's call site
|
||||||
|
* ({@link #reportRequiredSecrets(BridgedConfig)}), not a few lines above it. That throw only
|
||||||
|
* fires when {@code auth.mode: token} is configured; under the default loopback-trust mode it
|
||||||
|
* never runs, and {@code auth.tokenEnv} is simply not required.
|
||||||
|
*
|
||||||
|
* <p>Package-private and pure (no I/O, no logging) so the derivation is unit-testable without
|
||||||
|
* capturing log output; {@link #reportRequiredSecrets(BridgedConfig)} is the logging caller.
|
||||||
|
*/
|
||||||
|
static Map<String, List<String>> requiredSecretEnvVars(BridgedConfig cfg) {
|
||||||
|
Map<String, List<String>> requiredBy = new LinkedHashMap<>();
|
||||||
|
cfg.profiles().forEach((name, profile) -> {
|
||||||
|
if (!profile.isSubscription()) {
|
||||||
|
requiredBy.computeIfAbsent(profile.tokenEnv(), _ -> new ArrayList<>())
|
||||||
|
.add("profile '" + name + "' tokenEnv");
|
||||||
|
}
|
||||||
|
if (profile.hasGitToken()) {
|
||||||
|
requiredBy.computeIfAbsent(profile.gitTokenEnv(), _ -> new ArrayList<>())
|
||||||
|
.add("profile '" + name + "' gitTokenEnv");
|
||||||
|
}
|
||||||
|
});
|
||||||
|
return requiredBy;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-594: log, by name only, which required env vars (see {@link #requiredSecretEnvVars}) are
|
||||||
|
* set in the daemon's own process environment — the environment every profile's {@code
|
||||||
|
* tokenEnv}/{@code gitTokenEnv} is read from at spawn time (see
|
||||||
|
* {@code HerdrPeerLauncher.resolveEnv}). Never logs a value, a prefix, or a length.
|
||||||
|
*
|
||||||
|
* <p>A missing entry only warns — it must never refuse to start. A daemon that boots and says
|
||||||
|
* what is wrong is strictly more useful than one that will not boot at all.
|
||||||
|
*/
|
||||||
|
private static void reportRequiredSecrets(BridgedConfig cfg) {
|
||||||
|
Map<String, List<String>> requiredBy = requiredSecretEnvVars(cfg);
|
||||||
|
if (requiredBy.isEmpty()) {
|
||||||
|
log.info("startup secrets: no profile references a token env var — nothing to check");
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
Map<String, String> env = System.getenv();
|
||||||
|
requiredBy.forEach((varName, sources) -> {
|
||||||
|
String value = env.get(varName);
|
||||||
|
if (value != null && !value.isBlank()) {
|
||||||
|
log.info("startup secret {}: set ({})", varName, String.join(", ", sources));
|
||||||
|
} else {
|
||||||
|
log.warn("startup secret {}: MISSING ({}) — the daemon will start anyway, and this "
|
||||||
|
+ "failure stays invisible until a worker actually needs it. Fix "
|
||||||
|
+ "${SHARED_ENV}/tools/secrets.sh and restart bridged from a LOGIN "
|
||||||
|
+ "shell (see scripts/redeploy-bridged.sh).",
|
||||||
|
varName, String.join(", ", sources));
|
||||||
|
}
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: {@code known:} empty (block absent entirely, or present but empty) means {@link
|
||||||
|
* BridgedConfig.MemberCredentials#blockedSet()} is empty too — every member pane inherits the
|
||||||
|
* operator's whole secret store, unblocked, exactly the defect this ticket fixes. Unlike a
|
||||||
|
* missing token ({@link #reportRequiredSecrets}), there is no name to point at: the point is
|
||||||
|
* that the block itself is missing. Warn once at startup and say what to add; never refuse to
|
||||||
|
* start over it — see {@link #reportRequiredSecrets} for why a daemon that boots and says
|
||||||
|
* what is wrong beats one that will not boot at all.
|
||||||
|
*
|
||||||
|
* <p>Package-private so the test can capture the log directly, the same way {@link
|
||||||
|
* #requiredSecretEnvVars} is exposed for {@link #reportRequiredSecrets}'s own test.
|
||||||
|
*/
|
||||||
|
static void reportMemberCredentialsGap(BridgedConfig cfg) {
|
||||||
|
BridgedConfig.MemberCredentials creds = cfg.memberCredentials();
|
||||||
|
if (creds != null && !creds.known().isEmpty()) {
|
||||||
|
log.info("memberCredentials: {} known name(s), {} allowed — blocking {} on every spawn",
|
||||||
|
creds.known().size(), creds.allow().size(), creds.blockedSet().size());
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
log.warn("memberCredentials: absent or empty — the daemon will start anyway, and every "
|
||||||
|
+ "member pane inherits the operator's WHOLE secret store, unblocked (CB-592's "
|
||||||
|
+ "protection is lost). Add a memberCredentials: block (policy/allow/known) to "
|
||||||
|
+ "bridged.yaml — see bridged.example.yaml — and restart.");
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
|
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
|
||||||
*
|
*
|
||||||
|
|||||||
@@ -5,7 +5,9 @@ import com.fasterxml.jackson.core.JsonParser;
|
|||||||
import com.fasterxml.jackson.core.JsonToken;
|
import com.fasterxml.jackson.core.JsonToken;
|
||||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
import com.fasterxml.jackson.dataformat.yaml.YAMLFactory;
|
import com.fasterxml.jackson.dataformat.yaml.YAMLFactory;
|
||||||
|
import dev.ltms.bridged.msg.AmqpReplyInbox;
|
||||||
import dev.ltms.bridged.peer.MemberRole;
|
import dev.ltms.bridged.peer.MemberRole;
|
||||||
|
import dev.ltms.bridged.placement.PlacementPolicies;
|
||||||
import org.slf4j.Logger;
|
import org.slf4j.Logger;
|
||||||
import org.slf4j.LoggerFactory;
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
@@ -58,6 +60,15 @@ import java.util.Set;
|
|||||||
* {@code fixed} (default), {@code round-robin}, or {@code weighted}
|
* {@code fixed} (default), {@code round-robin}, or {@code weighted}
|
||||||
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
|
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
|
||||||
* historical behaviour), CB-501
|
* historical behaviour), CB-501
|
||||||
|
* @param quarantineCooldownSeconds how long a credential stays quarantined after a
|
||||||
|
* {@code BACKEND_EXHAUSTED} classification (CB-578 stage B); {@code null}/{@code
|
||||||
|
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Baked once into the
|
||||||
|
* {@code BackendQuarantine} built at startup, so it is DEFERRED: changing it
|
||||||
|
* needs a restart, and a quarantine already running keeps whatever cooldown was
|
||||||
|
* live when it started.
|
||||||
|
* @param memberCredentials deny-by-default policy (CB-596) for which of the operator's own host
|
||||||
|
* credentials a spawned member's pane inherits. {@code null} (the block
|
||||||
|
* omitted) blocks nothing — see {@link MemberCredentials}.
|
||||||
*/
|
*/
|
||||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||||
public record BridgedConfig(
|
public record BridgedConfig(
|
||||||
@@ -76,7 +87,23 @@ public record BridgedConfig(
|
|||||||
Health health,
|
Health health,
|
||||||
String placement,
|
String placement,
|
||||||
Auth auth,
|
Auth auth,
|
||||||
ConfigReload configReload) {
|
ConfigReload configReload,
|
||||||
|
Integer quarantineCooldownSeconds,
|
||||||
|
MemberCredentials memberCredentials) {
|
||||||
|
|
||||||
|
/** Back-compat form before the CB-596 {@code memberCredentials:} block was added. */
|
||||||
|
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
|
||||||
|
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
|
||||||
|
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
|
||||||
|
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
|
||||||
|
ConfigReload configReload, Integer quarantineCooldownSeconds) {
|
||||||
|
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
|
||||||
|
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
|
||||||
|
configReload, quarantineCooldownSeconds, null);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Default cooldown (CB-578 stage B) when {@code quarantineCooldownSeconds} is absent/non-positive. */
|
||||||
|
public static final int DEFAULT_QUARANTINE_COOLDOWN_SECONDS = 1800;
|
||||||
|
|
||||||
/** Back-compat 14-arg form — no {@code configReload:} block, so file watching stays off. */
|
/** Back-compat 14-arg form — no {@code configReload:} block, so file watching stays off. */
|
||||||
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
|
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
|
||||||
@@ -84,7 +111,7 @@ public record BridgedConfig(
|
|||||||
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
|
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
|
||||||
LeadHeartbeat leadHeartbeat, String placement, Auth auth) {
|
LeadHeartbeat leadHeartbeat, String placement, Auth auth) {
|
||||||
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
|
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
|
||||||
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, null);
|
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, null, null);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Back-compat form before the optional {@code health:} block was added. */
|
/** Back-compat form before the optional {@code health:} block was added. */
|
||||||
@@ -93,7 +120,18 @@ public record BridgedConfig(
|
|||||||
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
|
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
|
||||||
LeadHeartbeat leadHeartbeat, String placement, Auth auth, ConfigReload configReload) {
|
LeadHeartbeat leadHeartbeat, String placement, Auth auth, ConfigReload configReload) {
|
||||||
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
|
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
|
||||||
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, configReload);
|
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, configReload, null);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Back-compat form before the CB-578 stage B {@code quarantineCooldownSeconds} field was added. */
|
||||||
|
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
|
||||||
|
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
|
||||||
|
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
|
||||||
|
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
|
||||||
|
ConfigReload configReload) {
|
||||||
|
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
|
||||||
|
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
|
||||||
|
configReload, null);
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -151,19 +189,58 @@ public record BridgedConfig(
|
|||||||
* @param cwd fixed working directory for this profile's workers (CB-112 "told otherwise");
|
* @param cwd fixed working directory for this profile's workers (CB-112 "told otherwise");
|
||||||
* {@code null}/blank → inherit the primary's cwd, else the daemon's
|
* {@code null}/blank → inherit the primary's cwd, else the daemon's
|
||||||
* @param parityOverlay repo-relative paths copied primary→worktree for config parity; null/empty
|
* @param parityOverlay repo-relative paths copied primary→worktree for config parity; null/empty
|
||||||
* defaults to a sensible set of local config files
|
* defaults to a sensible set of local config files.
|
||||||
|
* <p><strong>Every overlay path must be gitignored or tracked-and-skipped.</strong>
|
||||||
|
* CB-576 made {@code release()} preserve a worktree that {@code git status
|
||||||
|
* --porcelain} reports as dirty, and it deliberately counts untracked files —
|
||||||
|
* the work lost in CB-576 was a new file nobody had added. So an overlay path
|
||||||
|
* that is neither gitignored nor tracked lands in every worktree as an
|
||||||
|
* untracked file, makes every {@code COMPLETED} release preserve, and
|
||||||
|
* worktrees then accumulate with no error anywhere.
|
||||||
|
* <p>Checked on 2026-08-15 (CB-581): inert as configured. Tracked overlay
|
||||||
|
* files carry {@code --skip-worktree} so {@code --porcelain} cannot see them,
|
||||||
|
* {@code bridged.yaml} is gitignored, and the default pair {@code .env} /
|
||||||
|
* {@code .envrc} does not exist in this repo. Note the default applies to
|
||||||
|
* <em>every</em> profile, so creating either file at the repo root is enough
|
||||||
|
* to make it live. Add a new overlay path to {@code .gitignore} in the same
|
||||||
|
* change that adds it here.
|
||||||
|
* <p>CB-578 stage C raised the stakes: a preserved dirty worktree is now also
|
||||||
|
* committed to {@code refs/wip/<branch>} via {@code git add -A}. The overlay
|
||||||
|
* carries the primary's own environment files into the worktree, so a
|
||||||
|
* non-gitignored overlay path no longer merely sits there as an untracked
|
||||||
|
* file — it gets committed into a git object that survives the worktree's
|
||||||
|
* removal. {@code add -A} respects {@code .gitignore}, which is exactly what
|
||||||
|
* keeps that from happening, so the gitignore rule above is now what stops a
|
||||||
|
* secret from reaching a durable commit.
|
||||||
* @param gitTokenEnv name of the host env var holding the git-forge API token; when set, its
|
* @param gitTokenEnv name of the host env var holding the git-forge API token; when set, its
|
||||||
* value is injected as {@code GITEA_TOKEN} so the worker can open its own PR
|
* value is injected as {@code GITEA_TOKEN} so the worker can open its own PR
|
||||||
* at checkpoint (CB-302). {@code null}/blank ⇒ no token is injected
|
* at checkpoint (CB-302). {@code null}/blank ⇒ no token is injected
|
||||||
* (minimal-grant default — push over SSH stays free, PR-create is opt-in)
|
* (minimal-grant default — push over SSH stays free, PR-create is opt-in)
|
||||||
* @param gitHostEnv name of the host env var holding the forge host (default {@code GITEA_HOST});
|
* @param gitHostEnv name of the host env var holding the forge host (default {@code GITEA_HOST});
|
||||||
* injected as {@code GITEA_HOST} <em>only</em> when {@code gitTokenEnv} is set
|
* injected as {@code GITEA_HOST} <em>only</em> when {@code gitTokenEnv} is set
|
||||||
* @param weight relative selection weight for {@code placement: weighted}. Absent or
|
* @param weight relative selection weight for automatic placement ({@code weighted},
|
||||||
* non-positive ⇒ 1.0. Weights are normalised by the policy, so they need
|
* {@code round-robin}, and {@code fixed}'s fallback walk). Absent ⇒ 1.0.
|
||||||
* not sum to 1.0.
|
* An explicit {@code 0} or negative value (CB-554) means "never
|
||||||
* @param maxLoad max live workers allowed on this profile at one time; absent or
|
* auto-select this profile": it is excluded from every automatic policy's
|
||||||
* non-positive ⇒ unlimited. Live means any session the registry still owns
|
* pool the same way a quarantined candidate is (see
|
||||||
* (acquired and not yet released), in any state.
|
* {@code PlacementPolicyUtil}). This does not make the profile
|
||||||
|
* unreachable — an explicit {@code bridge_spawn{profile:"..."}} bypasses
|
||||||
|
* placement entirely and still resolves it. Weights among the remaining
|
||||||
|
* (non-excluded) candidates need not sum to 1.0; only their ratios matter.
|
||||||
|
* @param maxLoad max live workers allowed on this profile at one time. Absent
|
||||||
|
* (unset/{@code null}) ⇒ unlimited — most profiles rely on this. An
|
||||||
|
* explicit {@code 0} (CB-585) means "cap this profile at zero live
|
||||||
|
* members": it is excluded from every automatic policy's candidate pool
|
||||||
|
* the same way a {@code weight <= 0} profile is (see
|
||||||
|
* {@code PlacementPolicyUtil.available()}, which already treats "at cap"
|
||||||
|
* and "excluded" alike), and an explicit
|
||||||
|
* {@code bridge_spawn{profile:"..."}} against it is refused too (see
|
||||||
|
* {@code CompositePeerLauncher.enforceMaxLoad}) — a cap is a capacity
|
||||||
|
* statement that does not stop being true just because the profile was
|
||||||
|
* named directly. A negative value has no sane meaning (there is no
|
||||||
|
* "excluded" to degrade to below zero) and is refused at config load
|
||||||
|
* instead, naming the profile and the key. Live means any session the
|
||||||
|
* registry still owns (acquired and not yet released), in any state.
|
||||||
* @param kind which peer launcher spawns this profile: {@code "claude-code"} (default —
|
* @param kind which peer launcher spawns this profile: {@code "claude-code"} (default —
|
||||||
* the {@link dev.ltms.bridged.member.ClaudeCodeLauncher}) or {@code "opencode"}.
|
* the {@link dev.ltms.bridged.member.ClaudeCodeLauncher}) or {@code "opencode"}.
|
||||||
* The {@code CompositePeerLauncher} routes {@code spawn}/reap by this value, so
|
* The {@code CompositePeerLauncher} routes {@code spawn}/reap by this value, so
|
||||||
@@ -187,6 +264,17 @@ public record BridgedConfig(
|
|||||||
* {@code null}/blank ⇒ the classification never fires for this profile and
|
* {@code null}/blank ⇒ the classification never fires for this profile and
|
||||||
* today's completion-fallback behaviour is unchanged. Every backend words
|
* today's completion-fallback behaviour is unchanged. Every backend words
|
||||||
* its refusal differently, so this is config, never a vendor string in code.
|
* its refusal differently, so this is config, never a vendor string in code.
|
||||||
|
* @param credentialId the shared account this profile authenticates as (CB-578 stage B). Two
|
||||||
|
* or more profiles setting the <em>same</em> non-blank value are quarantined
|
||||||
|
* together by one {@code exhaustedPattern} classification on any one of
|
||||||
|
* them — the case a single OpenAI (or any other) credential backing several
|
||||||
|
* profiles ({@code sol}, {@code terra}, ...) needs, so a spawn does not walk
|
||||||
|
* straight onto the other profile sharing the same exhausted account.
|
||||||
|
* {@code null}/blank ⇒ this profile's {@link #effectiveCredentialId()} is
|
||||||
|
* its own name, so it quarantines alone — today's behaviour for every
|
||||||
|
* profile that does not opt in. Read live off the current config, so it is
|
||||||
|
* HOT: a change takes effect on the next exhaustion classification / spawn,
|
||||||
|
* no restart needed.
|
||||||
*/
|
*/
|
||||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||||
public record Profile(String profile, String baseUrl, String model,
|
public record Profile(String profile, String baseUrl, String model,
|
||||||
@@ -200,7 +288,8 @@ public record BridgedConfig(
|
|||||||
Float weight,
|
Float weight,
|
||||||
Integer maxLoad,
|
Integer maxLoad,
|
||||||
Boolean subscription,
|
Boolean subscription,
|
||||||
String exhaustedPattern) {
|
String exhaustedPattern,
|
||||||
|
String credentialId) {
|
||||||
|
|
||||||
/** Peer kind spawned by {@link dev.ltms.bridged.member.ClaudeCodeLauncher} (the default). */
|
/** Peer kind spawned by {@link dev.ltms.bridged.member.ClaudeCodeLauncher} (the default). */
|
||||||
public static final String KIND_CLAUDE_CODE = "claude-code";
|
public static final String KIND_CLAUDE_CODE = "claude-code";
|
||||||
@@ -240,12 +329,25 @@ public record BridgedConfig(
|
|||||||
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
|
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
|
||||||
gitHostEnv = (gitHostEnv == null || gitHostEnv.isBlank()) ? "GITEA_HOST" : gitHostEnv;
|
gitHostEnv = (gitHostEnv == null || gitHostEnv.isBlank()) ? "GITEA_HOST" : gitHostEnv;
|
||||||
env = (env == null) ? Map.of() : Map.copyOf(env);
|
env = (env == null) ? Map.of() : Map.copyOf(env);
|
||||||
weight = (weight == null || weight <= 0.0f) ? 1.0f : weight;
|
// CB-554: absent still means 1.0, but an explicit non-positive value must survive as
|
||||||
maxLoad = (maxLoad == null || maxLoad <= 0) ? null : maxLoad;
|
// "excluded from automatic selection" (PlacementCandidate.excluded(), weight <= 0), not
|
||||||
|
// get coerced back up to 1.0 — that coercion was the bug (weight: 0 looked like "never
|
||||||
|
// pick me" and actually meant "pick me as often as anyone else").
|
||||||
|
weight = (weight == null) ? 1.0f : Math.max(weight, 0.0f);
|
||||||
|
// CB-585: absent still means unlimited (null), but an explicit maxLoad: 0 must survive
|
||||||
|
// as "capped at zero," not get coerced back up to null/unlimited — that coercion was
|
||||||
|
// the bug (maxLoad: 0 read as "never run anything here" and behaved as the opposite,
|
||||||
|
// the one throttle a subscription: true profile has against the operator's own paid
|
||||||
|
// plan). A negative value is refused earlier, at config load (rejectNegativeMaxLoad),
|
||||||
|
// so it never reaches this constructor and needs no clamping here.
|
||||||
subscription = (subscription != null && subscription) ? Boolean.TRUE : Boolean.FALSE;
|
subscription = (subscription != null && subscription) ? Boolean.TRUE : Boolean.FALSE;
|
||||||
// exhaustedPattern stays null when unset/blank (opt-in) — no defaulting, no vendor
|
// exhaustedPattern stays null when unset/blank (opt-in) — no defaulting, no vendor
|
||||||
// wording: an unconfigured profile keeps today's completion-fallback behaviour exactly.
|
// wording: an unconfigured profile keeps today's completion-fallback behaviour exactly.
|
||||||
exhaustedPattern = (exhaustedPattern == null || exhaustedPattern.isBlank()) ? null : exhaustedPattern;
|
exhaustedPattern = (exhaustedPattern == null || exhaustedPattern.isBlank()) ? null : exhaustedPattern;
|
||||||
|
// credentialId stays null when unset/blank — effectiveCredentialId() is where the
|
||||||
|
// "quarantines alone" fallback actually lives, so today's behaviour needs no defaulting
|
||||||
|
// here at all.
|
||||||
|
credentialId = (credentialId == null || credentialId.isBlank()) ? null : credentialId;
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -290,7 +392,7 @@ public record BridgedConfig(
|
|||||||
public Profile withProfile(String p) {
|
public Profile withProfile(String p) {
|
||||||
return new Profile(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
|
return new Profile(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
|
||||||
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, subscription,
|
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, subscription,
|
||||||
exhaustedPattern);
|
exhaustedPattern, credentialId);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** True when this profile is served by the Claude Code adapter (the default kind). */
|
/** True when this profile is served by the Claude Code adapter (the default kind). */
|
||||||
@@ -326,11 +428,38 @@ public record BridgedConfig(
|
|||||||
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, null, null);
|
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, null, null);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Backward-compatible constructor without the CB-578 stage B {@code credentialId} — the
|
||||||
|
* profile quarantines alone under its own name (see {@link #effectiveCredentialId()}). Keeps
|
||||||
|
* pre-stage-B call sites (and any YAML that omits the key) compiling and behaving identically.
|
||||||
|
*/
|
||||||
|
public Profile(String profile, String baseUrl, String model,
|
||||||
|
String configDir, String tokenEnv, List<String> argv,
|
||||||
|
String placement, String workspace, String tabLabel, String mcpUrl,
|
||||||
|
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
|
||||||
|
String kind, Map<String, String> env, Float weight, Integer maxLoad,
|
||||||
|
Boolean subscription, String exhaustedPattern) {
|
||||||
|
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
|
||||||
|
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad,
|
||||||
|
subscription, exhaustedPattern, null);
|
||||||
|
}
|
||||||
|
|
||||||
/** True when this profile's CB-578 stage A backend-exhausted classification is configured. */
|
/** True when this profile's CB-578 stage A backend-exhausted classification is configured. */
|
||||||
public boolean hasExhaustedPattern() {
|
public boolean hasExhaustedPattern() {
|
||||||
return exhaustedPattern != null;
|
return exhaustedPattern != null;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The credential group this profile quarantines with (CB-578 stage B): the configured
|
||||||
|
* {@link #credentialId} when set, else this profile's own name — so an unconfigured profile
|
||||||
|
* quarantines alone, exactly as it did before this field existed. Two profiles that set the
|
||||||
|
* same non-blank {@code credentialId} share one quarantine: a {@code BACKEND_EXHAUSTED}
|
||||||
|
* classification on either one quarantines both.
|
||||||
|
*/
|
||||||
|
public String effectiveCredentialId() {
|
||||||
|
return (credentialId == null || credentialId.isBlank()) ? profile : credentialId;
|
||||||
|
}
|
||||||
|
|
||||||
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
|
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
|
||||||
public boolean hasGitToken() {
|
public boolean hasGitToken() {
|
||||||
return gitTokenEnv != null && !gitTokenEnv.isBlank();
|
return gitTokenEnv != null && !gitTokenEnv.isBlank();
|
||||||
@@ -420,16 +549,24 @@ public record BridgedConfig(
|
|||||||
* stays soft-state. Production default is LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1
|
* stays soft-state. Production default is LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1
|
||||||
* and is a URI-only swap.
|
* and is a URI-only swap.
|
||||||
*
|
*
|
||||||
* @param uri AMQP connection URI, e.g. {@code amqp://guest:guest@127.0.0.1:5672/}. Blank/{@code null}
|
* @param uri AMQP connection URI, e.g. {@code amqp://guest:guest@127.0.0.1:5672/}. Blank/
|
||||||
* ⇒ the broker block is treated as absent (in-memory adapter).
|
* {@code null} ⇒ the broker block is treated as absent (in-memory adapter).
|
||||||
|
* @param prefetch CB-527: the consumer's {@code basicQos} prefetch count, bounding how many
|
||||||
|
* unacked messages the AMQP inbox holds in-heap per owned target. {@code null}/
|
||||||
|
* non-positive ⇒ {@link AmqpReplyInbox#DEFAULT_PREFETCH}.
|
||||||
*/
|
*/
|
||||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||||
public record Broker(String uri) {
|
public record Broker(String uri, Integer prefetch) {
|
||||||
|
|
||||||
/** True when a usable broker URI is configured (an empty block does not enable AMQP). */
|
/** True when a usable broker URI is configured (an empty block does not enable AMQP). */
|
||||||
public boolean isConfigured() {
|
public boolean isConfigured() {
|
||||||
return uri != null && !uri.isBlank();
|
return uri != null && !uri.isBlank();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** The prefetch to use, defaulting to {@link AmqpReplyInbox#DEFAULT_PREFETCH} when unset. */
|
||||||
|
public int prefetchOrDefault() {
|
||||||
|
return (prefetch != null && prefetch > 0) ? prefetch : AmqpReplyInbox.DEFAULT_PREFETCH;
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -804,6 +941,58 @@ public record BridgedConfig(
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: which of the operator's own host credentials a spawned member's pane may inherit.
|
||||||
|
*
|
||||||
|
* <p>A herdr pane runs a login shell that re-sources the operator's own secret store, so a
|
||||||
|
* member inherits every credential the operator's shell holds — measured at 31 names on this
|
||||||
|
* host, of which only one ({@code GITEA_ACCESS_TOKEN}) used to be blocked, and that block was a
|
||||||
|
* single name hardcoded in {@link HerdrPeerLauncher} rather than driven by config (gitea issue
|
||||||
|
* #82). This record replaces that hardcoded shadow with a config-driven one.
|
||||||
|
*
|
||||||
|
* <p><b>deny-by-default, not a deny-list.</b> A deny-list (block these specific names, let
|
||||||
|
* everything else through) is silently wrong the moment a new secret is added to the operator's
|
||||||
|
* store — nothing would ever report it. Deny-by-default inverts that: {@link #known} bounds the
|
||||||
|
* blast radius to names the operator has actually enumerated, and every one of them is blocked
|
||||||
|
* UNLESS it is also in {@link #allow}. A name that shows up in neither list is not silently
|
||||||
|
* allowed — see {@code HerdrPeerLauncher}'s gap detector, which logs it.
|
||||||
|
*
|
||||||
|
* @param policy how the block is computed. Only {@link #POLICY_DENY_BY_DEFAULT} is understood
|
||||||
|
* today; {@code null}/blank defaults to it. An operator's own deny-list is
|
||||||
|
* deliberately not supported — see above.
|
||||||
|
* @param allow credential names a member legitimately needs (e.g. the gateway token it reaches
|
||||||
|
* the LLM through, the repo-scoped forge token it opens its own PR with). Every
|
||||||
|
* name here is left unmentioned in the pane's env overlay, so the value the pane's
|
||||||
|
* own (login) shell exports passes through untouched.
|
||||||
|
* @param known every credential name the operator's store is known to export. Every name here
|
||||||
|
* that is NOT also in {@link #allow} is overlaid with a non-secret sentinel value,
|
||||||
|
* shadowing whatever the pane's login shell would otherwise export for it.
|
||||||
|
*/
|
||||||
|
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||||
|
public record MemberCredentials(String policy, List<String> allow, List<String> known) {
|
||||||
|
|
||||||
|
/** The only policy this build understands: block every {@code known} name not in {@code allow}. */
|
||||||
|
public static final String POLICY_DENY_BY_DEFAULT = "deny-by-default";
|
||||||
|
|
||||||
|
public MemberCredentials {
|
||||||
|
policy = (policy == null || policy.isBlank()) ? POLICY_DENY_BY_DEFAULT : policy.toLowerCase();
|
||||||
|
allow = allow == null ? List.of() : List.copyOf(allow);
|
||||||
|
known = known == null ? List.of() : List.copyOf(known);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** {@link #allow} as a set, for membership checks. */
|
||||||
|
public Set<String> allowSet() {
|
||||||
|
return Set.copyOf(allow);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** {@link #known} minus {@link #allow} — the names a spawn must shadow. */
|
||||||
|
public Set<String> blockedSet() {
|
||||||
|
Set<String> blocked = new java.util.LinkedHashSet<>(known);
|
||||||
|
blocked.removeAll(allowSet());
|
||||||
|
return blocked;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* The candidate profiles an unqualified spawn of {@code role} chooses between, in definition
|
* The candidate profiles an unqualified spawn of {@code role} chooses between, in definition
|
||||||
* order (CB-557).
|
* order (CB-557).
|
||||||
@@ -846,11 +1035,16 @@ public record BridgedConfig(
|
|||||||
/**
|
/**
|
||||||
* Top-level keys this version understands. Used only to warn about the rest — see
|
* Top-level keys this version understands. Used only to warn about the rest — see
|
||||||
* {@link #warnUnknownTopLevelKeys}. Keep in step with the record components.
|
* {@link #warnUnknownTopLevelKeys}. Keep in step with the record components.
|
||||||
|
*
|
||||||
|
* <p>Package-private (not {@code private}) so a test can assert every key here is documented in
|
||||||
|
* {@code bridged.example.yaml} — the only committed description of the config schema, since
|
||||||
|
* {@code bridged.yaml} itself is gitignored.
|
||||||
*/
|
*/
|
||||||
private static final Set<String> KNOWN_TOP_LEVEL_KEYS = Set.of(
|
static final Set<String> KNOWN_TOP_LEVEL_KEYS = Set.of(
|
||||||
"bind", "herdrSocket", "profiles", "guard", "worktreeRoot",
|
"bind", "herdrSocket", "profiles", "guard", "worktreeRoot",
|
||||||
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
|
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
|
||||||
"leadHeartbeat", "health", "placement", "auth", "configReload");
|
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds",
|
||||||
|
"memberCredentials");
|
||||||
|
|
||||||
/** Load and validate config from {@code path}. */
|
/** Load and validate config from {@code path}. */
|
||||||
public static BridgedConfig load(Path path) {
|
public static BridgedConfig load(Path path) {
|
||||||
@@ -860,7 +1054,16 @@ public record BridgedConfig(
|
|||||||
rejectLeaderTerminalKey(yaml);
|
rejectLeaderTerminalKey(yaml);
|
||||||
warnUnknownTopLevelKeys(yaml, path);
|
warnUnknownTopLevelKeys(yaml, path);
|
||||||
rejectDuplicateMemberSlots(yaml);
|
rejectDuplicateMemberSlots(yaml);
|
||||||
|
rejectNegativeMaxLoad(yaml);
|
||||||
|
rejectUnknownKind(yaml);
|
||||||
|
rejectUnknownAuthMode(yaml);
|
||||||
|
rejectUnknownPlacement(yaml);
|
||||||
|
rejectUnknownMemberCredentialsPolicy(yaml);
|
||||||
BridgedConfig cfg = YAML.readValue(yaml, BridgedConfig.class);
|
BridgedConfig cfg = YAML.readValue(yaml, BridgedConfig.class);
|
||||||
|
// CB-606: validated here, eagerly, using PlacementPolicies.fromName as the single source
|
||||||
|
// of truth — not lazily at first spawn (see CompositePeerLauncher's placementPolicy
|
||||||
|
// Supplier), where a bad name would still start a daemon that looks healthy.
|
||||||
|
rejectUnknownPlacementPolicy(cfg.placement());
|
||||||
return cfg.withDefaults();
|
return cfg.withDefaults();
|
||||||
} catch (IOException e) {
|
} catch (IOException e) {
|
||||||
throw new UncheckedIOException("cannot read bridged config at " + path, e);
|
throw new UncheckedIOException("cannot read bridged config at " + path, e);
|
||||||
@@ -1119,6 +1322,236 @@ public record BridgedConfig(
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reject a profile whose {@code maxLoad:} is negative, naming both the profile and the key.
|
||||||
|
*
|
||||||
|
* <p>Unlike {@code weight} (CB-554), where a negative value degrades to the same "excluded"
|
||||||
|
* meaning as an explicit {@code 0}, {@code maxLoad} has nowhere lower to degrade to — {@code 0}
|
||||||
|
* already means "capped at zero live members" (CB-585), the strictest cap there is. Silently
|
||||||
|
* normalising a negative value to something else is exactly the shape of bug this ticket fixes
|
||||||
|
* for {@code 0}, so it is refused instead, loud and specific, rather than guessed at.
|
||||||
|
*
|
||||||
|
* @param yaml the raw config text
|
||||||
|
* @throws IllegalStateException when any profile's {@code maxLoad} is negative
|
||||||
|
*/
|
||||||
|
static void rejectNegativeMaxLoad(String yaml) {
|
||||||
|
Map<?, ?> raw;
|
||||||
|
try {
|
||||||
|
raw = YAML.readValue(yaml, Map.class);
|
||||||
|
} catch (IOException | IllegalArgumentException e) {
|
||||||
|
return; // a malformed file is reported by the real parse, not here
|
||||||
|
}
|
||||||
|
if (raw == null || !(raw.get("profiles") instanceof Map<?, ?> profiles)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
List<String> bad = profiles.entrySet().stream()
|
||||||
|
.filter(e -> e.getValue() instanceof Map<?, ?> p
|
||||||
|
&& p.get("maxLoad") instanceof Number n && n.doubleValue() < 0)
|
||||||
|
.map(e -> String.valueOf(e.getKey()))
|
||||||
|
.sorted()
|
||||||
|
.toList();
|
||||||
|
if (!bad.isEmpty()) {
|
||||||
|
throw new IllegalStateException("refusing to start: profile(s) [" + String.join(", ", bad)
|
||||||
|
+ "] set a negative maxLoad — maxLoad caps live workers at a non-negative count;"
|
||||||
|
+ " use 0 to cap a profile at zero live members, or omit the key for unlimited.");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** The peer kinds this build has an adapter for — {@link Profile#kind()}'s only valid values. */
|
||||||
|
private static final Set<String> KNOWN_KINDS = Set.of(Profile.KIND_CLAUDE_CODE, Profile.KIND_OPENCODE);
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reject a profile whose {@code kind:} is not one of {@link #KNOWN_KINDS} (CB-604), naming the
|
||||||
|
* profile, the value it set, and the accepted set.
|
||||||
|
*
|
||||||
|
* <p>{@link Profile}'s compact constructor only lower-cases {@code kind} and compares it against
|
||||||
|
* {@code KIND_OPENCODE} — anything else, including a typo like {@code opencod}, silently falls
|
||||||
|
* into the claude-code bucket ({@link dev.ltms.bridged.member.CompositePeerLauncher} routes by
|
||||||
|
* exact adapter claim, not by membership in a known set). With {@code argv:} also unset, the argv
|
||||||
|
* default special-cases only the exact string {@code "claude-code"}, so the launch command falls
|
||||||
|
* back to {@code List.of(kind)} — the daemon then tries to run a program literally named after the
|
||||||
|
* typo. {@code CompositePeerLauncher}'s constructor already treats a profile claimed by two
|
||||||
|
* adapters as fatal (CB-402); an unrecognized kind is the same class of adapter-routing mistake
|
||||||
|
* and gets the same treatment here, at config load, rather than surfacing later as a failed spawn.
|
||||||
|
*
|
||||||
|
* @param yaml the raw config text
|
||||||
|
* @throws IllegalStateException when any profile's {@code kind} is a non-blank value not in
|
||||||
|
* {@link #KNOWN_KINDS} (case-insensitive)
|
||||||
|
*/
|
||||||
|
static void rejectUnknownKind(String yaml) {
|
||||||
|
Map<?, ?> raw;
|
||||||
|
try {
|
||||||
|
raw = YAML.readValue(yaml, Map.class);
|
||||||
|
} catch (IOException | IllegalArgumentException e) {
|
||||||
|
return; // a malformed file is reported by the real parse, not here
|
||||||
|
}
|
||||||
|
if (raw == null || !(raw.get("profiles") instanceof Map<?, ?> profiles)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
List<String> bad = profiles.entrySet().stream()
|
||||||
|
.filter(e -> e.getValue() instanceof Map<?, ?> p
|
||||||
|
&& p.get("kind") instanceof String k && !k.isBlank()
|
||||||
|
&& !KNOWN_KINDS.contains(k.toLowerCase()))
|
||||||
|
.map(e -> String.valueOf(e.getKey()) + "=" + ((Map<?, ?>) e.getValue()).get("kind"))
|
||||||
|
.sorted()
|
||||||
|
.toList();
|
||||||
|
if (!bad.isEmpty()) {
|
||||||
|
throw new IllegalStateException("refusing to start: profile(s) [" + String.join(", ", bad)
|
||||||
|
+ "] set an unrecognized kind — accepted values are "
|
||||||
|
+ String.join(", ", KNOWN_KINDS.stream().sorted().toList())
|
||||||
|
+ " (case-insensitive); an unrecognized kind would otherwise fall back to the"
|
||||||
|
+ " claude-code adapter and try to launch a program named after the typo.");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** The auth modes this build understands — {@link Auth#mode()}'s only valid values. */
|
||||||
|
private static final Set<String> KNOWN_AUTH_MODES = Set.of(Auth.MODE_LOOPBACK_TRUST, Auth.MODE_TOKEN);
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reject an {@code auth.mode} that is not one of {@link #KNOWN_AUTH_MODES} (CB-606), naming the
|
||||||
|
* value and the accepted set.
|
||||||
|
*
|
||||||
|
* <p>{@link Auth}'s compact constructor only lower-cases {@code mode}, and
|
||||||
|
* {@link Auth#tokenMode()} only compares the result against {@code MODE_TOKEN} — anything else,
|
||||||
|
* including a typo like {@code toekn}, silently behaves as {@code loopback-trust}. That fallback
|
||||||
|
* is otherwise checked only by {@link #validateAuthExposure()}, and only when the bind is
|
||||||
|
* non-loopback: on a loopback bind (the common case) the typo is invisible end to end — the
|
||||||
|
* daemon starts cleanly and authenticates nobody while the operator believes {@code token} mode
|
||||||
|
* is active. Refuse it here, unconditionally, at config load, rather than let it hide behind the
|
||||||
|
* bind check.
|
||||||
|
*
|
||||||
|
* @param yaml the raw config text
|
||||||
|
* @throws IllegalStateException when {@code auth.mode} is a non-blank value not in
|
||||||
|
* {@link #KNOWN_AUTH_MODES} (case-insensitive)
|
||||||
|
*/
|
||||||
|
static void rejectUnknownAuthMode(String yaml) {
|
||||||
|
Map<?, ?> raw;
|
||||||
|
try {
|
||||||
|
raw = YAML.readValue(yaml, Map.class);
|
||||||
|
} catch (IOException | IllegalArgumentException e) {
|
||||||
|
return; // a malformed file is reported by the real parse, not here
|
||||||
|
}
|
||||||
|
if (raw == null || !(raw.get("auth") instanceof Map<?, ?> auth)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (!(auth.get("mode") instanceof String mode) || mode.isBlank()
|
||||||
|
|| KNOWN_AUTH_MODES.contains(mode.toLowerCase())) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
throw new IllegalStateException("refusing to start: auth.mode=" + mode
|
||||||
|
+ " is not recognized — accepted values are "
|
||||||
|
+ String.join(", ", KNOWN_AUTH_MODES.stream().sorted().toList())
|
||||||
|
+ " (case-insensitive); an unrecognized mode would otherwise silently fall back to"
|
||||||
|
+ " loopback-trust, which authenticates nobody.");
|
||||||
|
}
|
||||||
|
|
||||||
|
/** The per-profile placements this build understands — {@link Profile#placement()}'s only valid values. */
|
||||||
|
private static final Set<String> KNOWN_PLACEMENTS = Set.of("tab", "pane");
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reject a profile whose {@code placement:} is not one of {@link #KNOWN_PLACEMENTS} (CB-606),
|
||||||
|
* naming the profile, the value it set, and the accepted set.
|
||||||
|
*
|
||||||
|
* <p>{@link Profile}'s compact constructor only lower-cases {@code placement}, and
|
||||||
|
* {@link Profile#tabPlacement()} only compares the result against {@code "tab"} — anything else,
|
||||||
|
* including a typo like {@code tabb}, silently falls back to the legacy pane placement with no
|
||||||
|
* signal anywhere.
|
||||||
|
*
|
||||||
|
* @param yaml the raw config text
|
||||||
|
* @throws IllegalStateException when any profile's {@code placement} is a non-blank value not in
|
||||||
|
* {@link #KNOWN_PLACEMENTS} (case-insensitive)
|
||||||
|
*/
|
||||||
|
static void rejectUnknownPlacement(String yaml) {
|
||||||
|
Map<?, ?> raw;
|
||||||
|
try {
|
||||||
|
raw = YAML.readValue(yaml, Map.class);
|
||||||
|
} catch (IOException | IllegalArgumentException e) {
|
||||||
|
return; // a malformed file is reported by the real parse, not here
|
||||||
|
}
|
||||||
|
if (raw == null || !(raw.get("profiles") instanceof Map<?, ?> profiles)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
List<String> bad = profiles.entrySet().stream()
|
||||||
|
.filter(e -> e.getValue() instanceof Map<?, ?> p
|
||||||
|
&& p.get("placement") instanceof String pl && !pl.isBlank()
|
||||||
|
&& !KNOWN_PLACEMENTS.contains(pl.toLowerCase()))
|
||||||
|
.map(e -> String.valueOf(e.getKey()) + "=" + ((Map<?, ?>) e.getValue()).get("placement"))
|
||||||
|
.sorted()
|
||||||
|
.toList();
|
||||||
|
if (!bad.isEmpty()) {
|
||||||
|
throw new IllegalStateException("refusing to start: profile(s) [" + String.join(", ", bad)
|
||||||
|
+ "] set an unrecognized placement — accepted values are "
|
||||||
|
+ String.join(", ", KNOWN_PLACEMENTS.stream().sorted().toList())
|
||||||
|
+ " (case-insensitive); an unrecognized placement would otherwise fall back to"
|
||||||
|
+ " legacy pane placement with no signal anywhere.");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** The member-credential policies this build understands — {@link MemberCredentials#policy()}'s only valid value. */
|
||||||
|
private static final Set<String> KNOWN_MEMBER_CREDENTIALS_POLICIES =
|
||||||
|
Set.of(MemberCredentials.POLICY_DENY_BY_DEFAULT);
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reject a {@code memberCredentials.policy} that is not {@link #KNOWN_MEMBER_CREDENTIALS_POLICIES}
|
||||||
|
* (CB-596), naming the value and the accepted set.
|
||||||
|
*
|
||||||
|
* <p>{@link MemberCredentials}'s compact constructor only lower-cases {@code policy} and defaults
|
||||||
|
* a blank one to {@link MemberCredentials#POLICY_DENY_BY_DEFAULT} — nothing rejects an actual
|
||||||
|
* typo like {@code deny-by-defualt}. There is only one policy today, so such a typo would
|
||||||
|
* currently behave identically to the real value by accident; the day a second policy exists
|
||||||
|
* that accident becomes a silent behavior change. Refuse it now, at config load, following the
|
||||||
|
* same pattern as {@link #rejectUnknownAuthMode} and {@link #rejectUnknownPlacement}.
|
||||||
|
*
|
||||||
|
* @param yaml the raw config text
|
||||||
|
* @throws IllegalStateException when {@code memberCredentials.policy} is a non-blank value not in
|
||||||
|
* {@link #KNOWN_MEMBER_CREDENTIALS_POLICIES} (case-insensitive)
|
||||||
|
*/
|
||||||
|
static void rejectUnknownMemberCredentialsPolicy(String yaml) {
|
||||||
|
Map<?, ?> raw;
|
||||||
|
try {
|
||||||
|
raw = YAML.readValue(yaml, Map.class);
|
||||||
|
} catch (IOException | IllegalArgumentException e) {
|
||||||
|
return; // a malformed file is reported by the real parse, not here
|
||||||
|
}
|
||||||
|
if (raw == null || !(raw.get("memberCredentials") instanceof Map<?, ?> mc)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (!(mc.get("policy") instanceof String policy) || policy.isBlank()
|
||||||
|
|| KNOWN_MEMBER_CREDENTIALS_POLICIES.contains(policy.toLowerCase())) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
throw new IllegalStateException("refusing to start: memberCredentials.policy=" + policy
|
||||||
|
+ " is not recognized — accepted values are "
|
||||||
|
+ String.join(", ", KNOWN_MEMBER_CREDENTIALS_POLICIES.stream().sorted().toList())
|
||||||
|
+ " (case-insensitive).");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reject a top-level {@code placement:} policy name {@link PlacementPolicies#fromName} does not
|
||||||
|
* recognize (CB-606), at config load rather than lazily at first spawn.
|
||||||
|
*
|
||||||
|
* <p>{@code CompositePeerLauncher} only calls {@link PlacementPolicies#fromName} per spawn,
|
||||||
|
* through a {@code Supplier} that re-reads live config (CB-559, so a hot-reloaded placement
|
||||||
|
* policy takes effect without a restart) — so a bad name still starts a daemon that looks
|
||||||
|
* healthy and fails only the first time something spawns without naming a profile. Every other
|
||||||
|
* field this class validates fails here, at load; this one gets the same treatment, calling
|
||||||
|
* {@link PlacementPolicies#fromName} itself as the single source of truth for what is valid
|
||||||
|
* rather than duplicating its accepted set.
|
||||||
|
*
|
||||||
|
* @param placement the raw, possibly null/blank {@code placement} value as parsed (before
|
||||||
|
* {@link #withDefaults()} runs); {@code fromName} itself treats null/blank as
|
||||||
|
* {@code fixed}, so this call changes no default
|
||||||
|
* @throws IllegalStateException when {@code placement} is a name {@link PlacementPolicies} does
|
||||||
|
* not recognize
|
||||||
|
*/
|
||||||
|
private static void rejectUnknownPlacementPolicy(String placement) {
|
||||||
|
try {
|
||||||
|
PlacementPolicies.fromName(placement);
|
||||||
|
} catch (IllegalArgumentException e) {
|
||||||
|
throw new IllegalStateException("refusing to start: " + e.getMessage(), e);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
static List<String> unknownTopLevelKeys(String yaml) {
|
static List<String> unknownTopLevelKeys(String yaml) {
|
||||||
Map<?, ?> raw;
|
Map<?, ?> raw;
|
||||||
try {
|
try {
|
||||||
@@ -1157,8 +1590,23 @@ public record BridgedConfig(
|
|||||||
// configReload is left as-is: null is "off", and ConfigReload's own compact constructor
|
// configReload is left as-is: null is "off", and ConfigReload's own compact constructor
|
||||||
// defaults the fields of a block that IS present. Defaulting it here would start watching
|
// defaults the fields of a block that IS present. Defaulting it here would start watching
|
||||||
// the file for every config that never asked to be watched.
|
// the file for every config that never asked to be watched.
|
||||||
|
// quarantineCooldownSeconds IS defaulted, unlike leadHeartbeat/configReload above: it has no
|
||||||
|
// separate on/off switch of its own (BackendQuarantine only ever quarantines a credential
|
||||||
|
// after a BACKEND_EXHAUSTED classification, which stays opt-in via exhaustedPattern), so a
|
||||||
|
// config that never mentions it should still get a sane cooldown rather than a null one.
|
||||||
|
Integer quarantineCooldown = (quarantineCooldownSeconds != null && quarantineCooldownSeconds > 0)
|
||||||
|
? quarantineCooldownSeconds : DEFAULT_QUARANTINE_COOLDOWN_SECONDS;
|
||||||
|
// memberCredentials IS defaulted, like guard/lifecycle/auth above, so no reader ever sees a
|
||||||
|
// null. CB-596: an empty MemberCredentials (empty known, empty allow) blocks NOTHING — unlike
|
||||||
|
// guard/lifecycle, an absent block is not a safe "feature off" default here, it is a gap. It
|
||||||
|
// is deliberately not pre-populated with a Java-side name list (that would just reintroduce
|
||||||
|
// the hardcoded-list defect this record replaces); the block must be configured in
|
||||||
|
// bridged.yaml to protect anything. See bridged.example.yaml's memberCredentials: comment.
|
||||||
|
MemberCredentials mc = memberCredentials != null ? memberCredentials
|
||||||
|
: new MemberCredentials(null, List.of(), List.of());
|
||||||
return new BridgedConfig(b, herdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
|
return new BridgedConfig(b, herdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
|
||||||
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload);
|
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
|
||||||
|
quarantineCooldown, mc);
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
|
|||||||
@@ -37,10 +37,15 @@ import java.util.function.Supplier;
|
|||||||
* {@code fleet.leaders} needs a restart, the same as any deferred key below.</li>
|
* {@code fleet.leaders} needs a restart, the same as any deferred key below.</li>
|
||||||
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
|
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
|
||||||
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
|
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
|
||||||
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code guard:},
|
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code quarantineCooldownSeconds}
|
||||||
* {@code worktreeRoot:}, adding or removing a profile (a new backend needs its own launcher,
|
* (CB-578 stage B — baked once into the {@code BackendQuarantine} built at startup),
|
||||||
|
* {@code guard:}, {@code worktreeRoot:}, adding or removing a profile (a new backend needs its own launcher,
|
||||||
* which is constructed once), <em>and an existing profile's launch settings</em> —
|
* which is constructed once), <em>and an existing profile's launch settings</em> —
|
||||||
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl} and the rest.
|
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl},
|
||||||
|
* {@code exhaustedPattern} (CB-578 stage A — compiled once into {@code Bridged.main}'s
|
||||||
|
* pattern map at startup), and the rest. {@code credentialId} (CB-578 stage B) is NOT on
|
||||||
|
* this list — it is read live off the config supplier at every quarantine check and
|
||||||
|
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so it is hot instead.
|
||||||
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} at construction and resolves
|
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} at construction and resolves
|
||||||
* each spawn out of that copy, so those never reach a launch until the daemon restarts. A
|
* each spawn out of that copy, so those never reach a launch until the daemon restarts. A
|
||||||
* reload logs these rather than pretending they applied.</li>
|
* reload logs these rather than pretending they applied.</li>
|
||||||
@@ -215,6 +220,12 @@ public final class ConfigRef implements Supplier<BridgedConfig> {
|
|||||||
|| !Objects.equals(old.spawnReadyPollMs(), fresh.spawnReadyPollMs())) {
|
|| !Objects.equals(old.spawnReadyPollMs(), fresh.spawnReadyPollMs())) {
|
||||||
changed.add("spawnReady*");
|
changed.add("spawnReady*");
|
||||||
}
|
}
|
||||||
|
// CB-578 stage B: baked once into the BackendQuarantine built at startup — a running
|
||||||
|
// quarantine keeps its original cooldown regardless, and a new cooldown only applies to a
|
||||||
|
// quarantine that starts after a restart.
|
||||||
|
if (!Objects.equals(old.quarantineCooldownSeconds(), fresh.quarantineCooldownSeconds())) {
|
||||||
|
changed.add("quarantineCooldownSeconds");
|
||||||
|
}
|
||||||
Map<String, BridgedConfig.Profile> before =
|
Map<String, BridgedConfig.Profile> before =
|
||||||
old.profiles() == null ? Map.of() : old.profiles();
|
old.profiles() == null ? Map.of() : old.profiles();
|
||||||
Map<String, BridgedConfig.Profile> after =
|
Map<String, BridgedConfig.Profile> after =
|
||||||
@@ -250,8 +261,10 @@ public final class ConfigRef implements Supplier<BridgedConfig> {
|
|||||||
|
|
||||||
/**
|
/**
|
||||||
* Whether two versions of a profile would launch a peer identically. Compares every component
|
* Whether two versions of a profile would launch a peer identically. Compares every component
|
||||||
* the launcher reads at spawn; {@code weight} and {@code maxLoad} are excluded because those are
|
* the launcher reads at spawn; {@code weight}, {@code maxLoad} and {@code credentialId} are
|
||||||
* read live by the placement policy and really do take effect on the next spawn.
|
* excluded because those are read live (by the placement policy and, for credentialId, by
|
||||||
|
* {@code CompositePeerLauncher}/the CB-578 stage B exhaustion sink) and really do take effect on
|
||||||
|
* the next spawn.
|
||||||
*/
|
*/
|
||||||
private static boolean sameLaunchSettings(BridgedConfig.Profile a, BridgedConfig.Profile b) {
|
private static boolean sameLaunchSettings(BridgedConfig.Profile a, BridgedConfig.Profile b) {
|
||||||
return Objects.equals(a.baseUrl(), b.baseUrl())
|
return Objects.equals(a.baseUrl(), b.baseUrl())
|
||||||
@@ -269,6 +282,10 @@ public final class ConfigRef implements Supplier<BridgedConfig> {
|
|||||||
&& Objects.equals(a.gitHostEnv(), b.gitHostEnv())
|
&& Objects.equals(a.gitHostEnv(), b.gitHostEnv())
|
||||||
&& Objects.equals(a.kind(), b.kind())
|
&& Objects.equals(a.kind(), b.kind())
|
||||||
&& Objects.equals(a.env(), b.env())
|
&& Objects.equals(a.env(), b.env())
|
||||||
&& Objects.equals(a.subscription(), b.subscription());
|
&& Objects.equals(a.subscription(), b.subscription())
|
||||||
|
// CB-578 stage B: exhaustedPattern is compiled once into Bridged.main's pattern map
|
||||||
|
// at startup (see ExhaustedPatternLookup wiring) — a reload never re-reads it, so a
|
||||||
|
// changed pattern must be reported as deferred, exactly like model/baseUrl/argv.
|
||||||
|
&& Objects.equals(a.exhaustedPattern(), b.exhaustedPattern());
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -67,6 +67,7 @@ public final class CompletionResolver implements TurnListener {
|
|||||||
private final AgentControl agents;
|
private final AgentControl agents;
|
||||||
private final Rendezvous rendezvous;
|
private final Rendezvous rendezvous;
|
||||||
private final ExhaustedPatternLookup exhaustedPatterns;
|
private final ExhaustedPatternLookup exhaustedPatterns;
|
||||||
|
private final ExhaustionSink exhaustionSink;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
|
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
|
||||||
@@ -89,11 +90,17 @@ public final class CompletionResolver implements TurnListener {
|
|||||||
* usage-limit refusal pattern. Required — there is deliberately no
|
* usage-limit refusal pattern. Required — there is deliberately no
|
||||||
* defaulting overload; a caller that does not want the classification
|
* defaulting overload; a caller that does not want the classification
|
||||||
* must pass an explicit inert value ({@link ExhaustedPatternLookup#none()}).
|
* must pass an explicit inert value ({@link ExhaustedPatternLookup#none()}).
|
||||||
|
* @param exhaustionSink CB-578 stage B: notified when a {@code BACKEND_EXHAUSTED}
|
||||||
|
* classification actually resolves a waiter. Required for the same
|
||||||
|
* reason as {@code exhaustedPatterns} — pass {@link ExhaustionSink#none()}
|
||||||
|
* to opt out.
|
||||||
*/
|
*/
|
||||||
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns) {
|
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
|
||||||
|
ExhaustionSink exhaustionSink) {
|
||||||
this.agents = agents;
|
this.agents = agents;
|
||||||
this.rendezvous = rendezvous;
|
this.rendezvous = rendezvous;
|
||||||
this.exhaustedPatterns = Objects.requireNonNull(exhaustedPatterns, "exhaustedPatterns");
|
this.exhaustedPatterns = Objects.requireNonNull(exhaustedPatterns, "exhaustedPatterns");
|
||||||
|
this.exhaustionSink = Objects.requireNonNull(exhaustionSink, "exhaustionSink");
|
||||||
}
|
}
|
||||||
|
|
||||||
@Override
|
@Override
|
||||||
@@ -210,6 +217,9 @@ public final class CompletionResolver implements TurnListener {
|
|||||||
inFlight.remove(target, turn);
|
inFlight.remove(target, turn);
|
||||||
log.warn("completion for {} classified BACKEND_EXHAUSTED (no bridge_reply; scrape "
|
log.warn("completion for {} classified BACKEND_EXHAUSTED (no bridge_reply; scrape "
|
||||||
+ "matched the profile's exhausted pattern): {}", target, reason);
|
+ "matched the profile's exhausted pattern): {}", target, reason);
|
||||||
|
// CB-578 stage B: only on the resolution that actually won the race — a late
|
||||||
|
// duplicate must never quarantine a credential twice for one refusal.
|
||||||
|
exhaustionSink.onExhausted(target, reason);
|
||||||
}
|
}
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,31 @@
|
|||||||
|
package dev.ltms.bridged.inject;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Notified when {@link CompletionResolver} actually delivers a {@code BACKEND_EXHAUSTED}
|
||||||
|
* classification to a waiting send (CB-578 stage B) — never on a race that lost (see
|
||||||
|
* {@link CompletionResolver#resolve}, which only calls this after
|
||||||
|
* {@code Rendezvous.resolveExhausted} returns {@code true}).
|
||||||
|
*
|
||||||
|
* <p>{@link CompletionResolver} knows only {@code target} (a herdr terminal id); it has no notion of
|
||||||
|
* profiles or credentials, so mapping {@code target} to whatever should be quarantined is entirely
|
||||||
|
* the sink's job — see {@code Bridged.main}'s wiring, which resolves target → session → profile →
|
||||||
|
* {@code effectiveCredentialId()} and calls {@code BackendQuarantine.quarantine} on it.
|
||||||
|
*/
|
||||||
|
@FunctionalInterface
|
||||||
|
public interface ExhaustionSink {
|
||||||
|
|
||||||
|
/**
|
||||||
|
* @param target the herdr terminal id whose turn was classified {@code BACKEND_EXHAUSTED}
|
||||||
|
* @param reason the matched-line reason carried by the classification
|
||||||
|
*/
|
||||||
|
void onExhausted(String target, String reason);
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Inert sink — nothing happens on exhaustion. The explicit stand-in a caller (or a test not
|
||||||
|
* exercising this feature) passes instead of a defaulting overload, exactly like
|
||||||
|
* {@link ExhaustedPatternLookup#none()}.
|
||||||
|
*/
|
||||||
|
static ExhaustionSink none() {
|
||||||
|
return (target, reason) -> { };
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -14,6 +14,8 @@ import dev.ltms.bridged.herdr.HerdrException;
|
|||||||
import dev.ltms.bridged.msg.MessageService;
|
import dev.ltms.bridged.msg.MessageService;
|
||||||
import dev.ltms.bridged.msg.Rendezvous;
|
import dev.ltms.bridged.msg.Rendezvous;
|
||||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||||
|
import dev.ltms.bridged.placement.BackendQuarantine;
|
||||||
|
import dev.ltms.bridged.placement.PlacementException;
|
||||||
import dev.ltms.bridged.session.SessionManager;
|
import dev.ltms.bridged.session.SessionManager;
|
||||||
import dev.ltms.bridged.session.MemberSession;
|
import dev.ltms.bridged.session.MemberSession;
|
||||||
import dev.ltms.bridged.session.WorktreeRequest;
|
import dev.ltms.bridged.session.WorktreeRequest;
|
||||||
@@ -33,6 +35,7 @@ import jakarta.servlet.http.HttpServlet;
|
|||||||
import java.util.LinkedHashMap;
|
import java.util.LinkedHashMap;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
|
import java.util.Objects;
|
||||||
import java.util.Set;
|
import java.util.Set;
|
||||||
import java.util.function.Function;
|
import java.util.function.Function;
|
||||||
import java.util.function.LongSupplier;
|
import java.util.function.LongSupplier;
|
||||||
@@ -78,6 +81,7 @@ public final class BridgeMcp {
|
|||||||
private final Metrics metrics; // CB-502: null → auth failures not counted
|
private final Metrics metrics; // CB-502: null → auth failures not counted
|
||||||
private final CapacitySource capacity;
|
private final CapacitySource capacity;
|
||||||
private final HealthCoverageSource healthCoverage;
|
private final HealthCoverageSource healthCoverage;
|
||||||
|
private final QuarantineSource quarantine;
|
||||||
|
|
||||||
/** Capacity facts used by {@code bridge_list}; production must supply the placement live count. */
|
/** Capacity facts used by {@code bridge_list}; production must supply the placement live count. */
|
||||||
public record CapacitySource(Function<String, Integer> liveCount, Function<String, Integer> maxLoad,
|
public record CapacitySource(Function<String, Integer> liveCount, Function<String, Integer> maxLoad,
|
||||||
@@ -90,17 +94,30 @@ public final class BridgeMcp {
|
|||||||
/** Coverage is supplied by the health wiring, not inferred from a missing dependency. */
|
/** Coverage is supplied by the health wiring, not inferred from a missing dependency. */
|
||||||
public record HealthCoverageSource(Supplier<String> value) { }
|
public record HealthCoverageSource(Supplier<String> value) { }
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-578 stage B quarantine facts used by {@code bridge_profiles}: a profile → credential id
|
||||||
|
* lookup, plus the shared {@link BackendQuarantine} to read remaining cooldowns off.
|
||||||
|
*/
|
||||||
|
public record QuarantineSource(Function<String, String> credentialIdFor, BackendQuarantine quarantine) {
|
||||||
|
/** Inert source — no profile is ever reported quarantined. Explicit stand-in, not a default. */
|
||||||
|
public static QuarantineSource none() { return new QuarantineSource(_ -> null, BackendQuarantine.none()); }
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
|
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
|
||||||
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
|
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
|
||||||
* Jetty's context handler and never passes through Javalin's {@code before}
|
* Jetty's context handler and never passes through Javalin's {@code before}
|
||||||
* filter, so the REST guard does not cover it.
|
* filter, so the REST guard does not cover it.
|
||||||
* @param metrics registry for auth-failure counting; may be {@code null}
|
* @param metrics registry for auth-failure counting; may be {@code null}
|
||||||
|
* @param quarantine CB-578 stage B facts for {@code bridge_profiles}; required — pass
|
||||||
|
* {@link QuarantineSource#none()} for a caller that does not want the feature
|
||||||
*/
|
*/
|
||||||
public BridgeMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
|
public BridgeMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
|
||||||
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
|
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
|
||||||
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage) {
|
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||||
|
QuarantineSource quarantine) {
|
||||||
this.capacity = capacity;
|
this.capacity = capacity;
|
||||||
|
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
|
||||||
this.healthCoverage = healthCoverage;
|
this.healthCoverage = healthCoverage;
|
||||||
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
|
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
|
||||||
this.transport = HttpServletStreamableServerTransportProvider.builder()
|
this.transport = HttpServletStreamableServerTransportProvider.builder()
|
||||||
@@ -210,12 +227,13 @@ public final class BridgeMcp {
|
|||||||
// CB-301-ext: optional isolated worktree for parallel implementers.
|
// CB-301-ext: optional isolated worktree for parallel implementers.
|
||||||
String callerCwd = identity.cwdForPid(callerPid(exchange));
|
String callerCwd = identity.cwdForPid(callerPid(exchange));
|
||||||
return spawn(sessions, str(a, "profile"), str(a, "role"), str(a, "cwd"), callerCwd,
|
return spawn(sessions, str(a, "profile"), str(a, "role"), str(a, "cwd"), callerCwd,
|
||||||
callerTerminal(exchange), worktreeRequest(a));
|
callerTerminal(exchange), worktreeRequest(a),
|
||||||
|
str(a, "sessionName"), str(a, "resumeSessionId"));
|
||||||
})
|
})
|
||||||
.toolCall(listTool(), (exchange, _) -> {
|
.toolCall(listTool(), (exchange, _) -> {
|
||||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
|
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
|
||||||
if (denied != null) return denied;
|
if (denied != null) return denied;
|
||||||
return listFleet(workers, sessions, messages, capacity, healthCoverage,
|
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine,
|
||||||
callers == null ? Map.of() : callers.leads(),
|
callers == null ? Map.of() : callers.leads(),
|
||||||
callerTerminal(exchange));
|
callerTerminal(exchange));
|
||||||
})
|
})
|
||||||
@@ -228,7 +246,7 @@ public final class BridgeMcp {
|
|||||||
.toolCall(profilesTool(), (exchange, _) -> {
|
.toolCall(profilesTool(), (exchange, _) -> {
|
||||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
|
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
|
||||||
if (denied != null) return denied;
|
if (denied != null) return denied;
|
||||||
return profiles(workers);
|
return profiles(workers, quarantine);
|
||||||
})
|
})
|
||||||
.toolCall(whoamiTool(), (exchange, _) -> {
|
.toolCall(whoamiTool(), (exchange, _) -> {
|
||||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
|
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
|
||||||
@@ -559,13 +577,25 @@ public final class BridgeMcp {
|
|||||||
return text("acknowledged " + msgId);
|
return text("acknowledged " + msgId);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** {@code bridge_status}: the live lifecycle status of a worker session. */
|
/**
|
||||||
|
* {@code bridge_status}: the live lifecycle status of a worker session, plus — when the worker
|
||||||
|
* is paused mid-turn in an async {@code bridge_ask} (CB-582) — the open question and how to
|
||||||
|
* answer it, so a lead on its normal poll cadence does not need the ticket to notice.
|
||||||
|
*/
|
||||||
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
|
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
|
||||||
if (isBlank(sessionId)) {
|
if (isBlank(sessionId)) {
|
||||||
return error("sessionId is required");
|
return error("sessionId is required");
|
||||||
}
|
}
|
||||||
try {
|
try {
|
||||||
return text(messages.status(sessionId).name().toLowerCase());
|
String base = messages.status(sessionId).name().toLowerCase();
|
||||||
|
MessageService.PendingAsk ask = messages.pendingAsk(sessionId);
|
||||||
|
if (ask == null) {
|
||||||
|
return text(base);
|
||||||
|
}
|
||||||
|
return text(base + "\n\n[question — worker is waiting for your answer]\n" + ask.question()
|
||||||
|
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + ask.turnId()
|
||||||
|
+ "\" and content set to your answer; the worker resumes the same turn."
|
||||||
|
+ " (ticket " + ask.ticket() + ")");
|
||||||
} catch (HerdrException e) {
|
} catch (HerdrException e) {
|
||||||
return error("herdr error for session " + sessionId + ": " + e.getMessage());
|
return error("herdr error for session " + sessionId + ": " + e.getMessage());
|
||||||
}
|
}
|
||||||
@@ -640,7 +670,7 @@ public final class BridgeMcp {
|
|||||||
|
|
||||||
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
|
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
|
||||||
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile) {
|
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile) {
|
||||||
return spawn(sessions, profile, null, null, null, null, null);
|
return spawn(sessions, profile, null, null, null, null, null, null, null);
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -650,13 +680,18 @@ public final class BridgeMcp {
|
|||||||
* {@code callerCwd} (the primary's directory), else the daemon's.
|
* {@code callerCwd} (the primary's directory), else the daemon's.
|
||||||
* CB-301: the session is registered with {@code ownerTerminal} as its owner.
|
* CB-301: the session is registered with {@code ownerTerminal} as its owner.
|
||||||
* CB-301-ext: {@code worktreeRequest} non-null provisions an isolated git worktree.
|
* CB-301-ext: {@code worktreeRequest} non-null provisions an isolated git worktree.
|
||||||
|
* CB-584: {@code sessionName}/{@code resumeSessionId} request agent session identity — a resumed
|
||||||
|
* conversation requires an explicit {@code profile} whose adapter declares
|
||||||
|
* {@link dev.ltms.bridged.peer.Capability#SESSION_RESUME}, or the spawn is refused rather than
|
||||||
|
* silently starting a cold session.
|
||||||
*
|
*
|
||||||
* <p>{@code role} and {@code profile} are independent: the role picks the contract, the profile
|
* <p>{@code role} and {@code profile} are independent: the role picks the contract, the profile
|
||||||
* picks the backend. A reviewer on the same profile as the dev it reviews is a normal spawn.
|
* picks the backend. A reviewer on the same profile as the dev it reviews is a normal spawn.
|
||||||
*/
|
*/
|
||||||
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile, String role,
|
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile, String role,
|
||||||
String requestedCwd, String callerCwd,
|
String requestedCwd, String callerCwd,
|
||||||
String ownerTerminal, WorktreeRequest worktreeRequest) {
|
String ownerTerminal, WorktreeRequest worktreeRequest,
|
||||||
|
String sessionName, String resumeSessionId) {
|
||||||
MemberRole memberRole;
|
MemberRole memberRole;
|
||||||
try {
|
try {
|
||||||
memberRole = isBlank(role) ? MemberRole.DEV : MemberRole.parse(role);
|
memberRole = isBlank(role) ? MemberRole.DEV : MemberRole.parse(role);
|
||||||
@@ -665,12 +700,17 @@ public final class BridgeMcp {
|
|||||||
}
|
}
|
||||||
try {
|
try {
|
||||||
MemberSession member = sessions.acquire(isBlank(profile) ? null : profile, memberRole,
|
MemberSession member = sessions.acquire(isBlank(profile) ? null : profile, memberRole,
|
||||||
requestedCwd, callerCwd, ownerTerminal, worktreeRequest);
|
requestedCwd, callerCwd, ownerTerminal, worktreeRequest,
|
||||||
|
isBlank(sessionName) ? null : sessionName, isBlank(resumeSessionId) ? null : resumeSessionId);
|
||||||
return text(json(memberView(member)));
|
return text(json(memberView(member)));
|
||||||
} catch (GuardException e) {
|
} catch (GuardException e) {
|
||||||
return error("subscription boundary: " + e.getMessage());
|
return error("subscription boundary: " + e.getMessage());
|
||||||
|
} catch (PlacementException e) {
|
||||||
|
// CB-599: no candidate had capacity (maxLoad, quarantine, or all-exhausted) — distinct
|
||||||
|
// from "profile does not exist" below.
|
||||||
|
return error("no capacity: " + e.getMessage());
|
||||||
} catch (IllegalArgumentException e) {
|
} catch (IllegalArgumentException e) {
|
||||||
return error(e.getMessage()); // unknown / no-default profile
|
return error(e.getMessage()); // unknown / no-default profile, or a refused resumeSessionId
|
||||||
} catch (PeerUnreachableException e) {
|
} catch (PeerUnreachableException e) {
|
||||||
return error("spawn timed out — worker pane never reached injectable state: " + e.getMessage());
|
return error("spawn timed out — worker pane never reached injectable state: " + e.getMessage());
|
||||||
} catch (HerdrException e) {
|
} catch (HerdrException e) {
|
||||||
@@ -706,11 +746,33 @@ public final class BridgeMcp {
|
|||||||
return null;
|
return null;
|
||||||
}
|
}
|
||||||
|
|
||||||
/** {@code bridge_profiles}: the configured worker profiles and the default. */
|
/**
|
||||||
static McpSchema.CallToolResult profiles(PeerLauncher workers) {
|
* {@code bridge_profiles}: the configured worker profiles, the default, and — CB-578 stage B —
|
||||||
return text(json(Map.of(
|
* which of them are currently quarantined (backend exhausted) and for how much longer. The
|
||||||
"profiles", workers.profiles(),
|
* {@code quarantined} key is present only when at least one profile is, so a fleet where nothing
|
||||||
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile())));
|
* has ever been quarantined gets exactly the pre-stage-B shape.
|
||||||
|
*/
|
||||||
|
static McpSchema.CallToolResult profiles(PeerLauncher workers, QuarantineSource quarantine) {
|
||||||
|
Map<String, Object> result = new LinkedHashMap<>();
|
||||||
|
result.put("profiles", workers.profiles());
|
||||||
|
result.put("default", workers.defaultProfile() == null ? "" : workers.defaultProfile());
|
||||||
|
Map<String, Object> quarantined = new LinkedHashMap<>();
|
||||||
|
for (String profile : workers.profiles()) {
|
||||||
|
String credentialId = quarantine.credentialIdFor().apply(profile);
|
||||||
|
if (credentialId == null) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
|
||||||
|
Map<String, Object> row = new LinkedHashMap<>();
|
||||||
|
row.put("credentialId", credentialId);
|
||||||
|
row.put("quarantinedForSeconds", remaining);
|
||||||
|
quarantined.put(profile, row);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
if (!quarantined.isEmpty()) {
|
||||||
|
result.put("quarantined", quarantined);
|
||||||
|
}
|
||||||
|
return text(json(result));
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -729,17 +791,22 @@ public final class BridgeMcp {
|
|||||||
* caller's own row is flagged {@code "self": true}: a peer needs to tell its own pane apart from
|
* caller's own row is flagged {@code "self": true}: a peer needs to tell its own pane apart from
|
||||||
* a peer's, and the alternative is every lead calling {@code bridge_whoami} to subtract itself.
|
* a peer's, and the alternative is every lead calling {@code bridge_whoami} to subtract itself.
|
||||||
*
|
*
|
||||||
|
* <p>CB-583: the {@code capacity} rows reuse {@code quarantine} (the same {@link QuarantineSource}
|
||||||
|
* {@code bridge_profiles} reads) so the two surfaces cannot disagree about which profile is
|
||||||
|
* quarantined — see {@link #capacityView}.
|
||||||
|
*
|
||||||
* @param leads terminal_id → lead name, live from the resolver
|
* @param leads terminal_id → lead name, live from the resolver
|
||||||
* @param selfTerm the calling pane's terminal id, or blank for a caller with no pane
|
* @param selfTerm the calling pane's terminal id, or blank for a caller with no pane
|
||||||
*/
|
*/
|
||||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions,
|
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions,
|
||||||
Map<String, String> leads, String selfTerm) {
|
Map<String, String> leads, String selfTerm) {
|
||||||
return listFleet(workers, sessions, null, CapacitySource.none(), new HealthCoverageSource(() -> "off"), leads, selfTerm);
|
return listFleet(workers, sessions, null, CapacitySource.none(), new HealthCoverageSource(() -> "off"),
|
||||||
|
QuarantineSource.none(), leads, selfTerm);
|
||||||
}
|
}
|
||||||
|
|
||||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||||
Map<String, String> leads, String selfTerm) {
|
QuarantineSource quarantine, Map<String, String> leads, String selfTerm) {
|
||||||
try {
|
try {
|
||||||
Map<String, Agent> live = workers.list().stream()
|
Map<String, Agent> live = workers.list().stream()
|
||||||
.map(Agent.class::cast)
|
.map(Agent.class::cast)
|
||||||
@@ -760,7 +827,7 @@ public final class BridgeMcp {
|
|||||||
result.put("healthCoverage", healthCoverage.value().get());
|
result.put("healthCoverage", healthCoverage.value().get());
|
||||||
if (capacity.available()) result.put("capacity", profiles.stream()
|
if (capacity.available()) result.put("capacity", profiles.stream()
|
||||||
.map(profile -> capacityView(profile, capacity.liveCount(), capacity.maxLoad(), roster, messages,
|
.map(profile -> capacityView(profile, capacity.liveCount(), capacity.maxLoad(), roster, messages,
|
||||||
capacity.clock().getAsLong())).toList());
|
capacity.clock().getAsLong(), quarantine)).toList());
|
||||||
return text(json(result));
|
return text(json(result));
|
||||||
} catch (HerdrException e) {
|
} catch (HerdrException e) {
|
||||||
return error("herdr error listing the fleet: " + e.getMessage());
|
return error("herdr error listing the fleet: " + e.getMessage());
|
||||||
@@ -785,9 +852,18 @@ public final class BridgeMcp {
|
|||||||
return row;
|
return row;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-583: {@code free} alone cannot tell a lead "busy, will free up" from "refusing, and
|
||||||
|
* nothing changes for N seconds" — those need different decisions. So a quarantined profile
|
||||||
|
* forces {@code free} to 0, whatever its {@code maxLoad}/{@code live} say, and the row carries
|
||||||
|
* the same {@code credentialId}/{@code quarantinedForSeconds} facts {@code bridge_profiles}
|
||||||
|
* reports, reusing {@link QuarantineSource} rather than a second lookup. Both new keys are
|
||||||
|
* added only when the profile is actually quarantined, so an ordinary fleet's rows are
|
||||||
|
* byte-identical to before this change.
|
||||||
|
*/
|
||||||
private static Map<String, Object> capacityView(String profile, Function<String, Integer> liveCount,
|
private static Map<String, Object> capacityView(String profile, Function<String, Integer> liveCount,
|
||||||
Function<String, Integer> maxLoad, List<MemberSession> roster,
|
Function<String, Integer> maxLoad, List<MemberSession> roster,
|
||||||
MessageService messages, long nowNanos) {
|
MessageService messages, long nowNanos, QuarantineSource quarantine) {
|
||||||
Integer cap = maxLoad.apply(profile);
|
Integer cap = maxLoad.apply(profile);
|
||||||
int live = liveCount.apply(profile);
|
int live = liveCount.apply(profile);
|
||||||
int reclaimable = (int) roster.stream().filter(s -> profile.equals(s.profile()))
|
int reclaimable = (int) roster.stream().filter(s -> profile.equals(s.profile()))
|
||||||
@@ -797,6 +873,14 @@ public final class BridgeMcp {
|
|||||||
Map<String, Object> row = new LinkedHashMap<>();
|
Map<String, Object> row = new LinkedHashMap<>();
|
||||||
row.put("profile", profile); row.put("maxLoad", cap); row.put("live", live);
|
row.put("profile", profile); row.put("maxLoad", cap); row.put("live", live);
|
||||||
row.put("free", cap == null ? null : Math.max(0, cap - live)); row.put("reclaimable", reclaimable);
|
row.put("free", cap == null ? null : Math.max(0, cap - live)); row.put("reclaimable", reclaimable);
|
||||||
|
String credentialId = quarantine.credentialIdFor().apply(profile);
|
||||||
|
if (credentialId != null) {
|
||||||
|
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
|
||||||
|
row.put("free", 0);
|
||||||
|
row.put("credentialId", credentialId);
|
||||||
|
row.put("quarantinedForSeconds", remaining);
|
||||||
|
});
|
||||||
|
}
|
||||||
return row;
|
return row;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -927,20 +1011,30 @@ public final class BridgeMcp {
|
|||||||
+ "for the default. The two are independent: a reviewer may run on the same profile "
|
+ "for the default. The two are independent: a reviewer may run on the same profile "
|
||||||
+ "as the dev it reviews. The member opens your current directory by default; pass "
|
+ "as the dev it reviews. The member opens your current directory by default; pass "
|
||||||
+ "cwd to pin a different one. Pass worktree:true (with ticket) or "
|
+ "cwd to pin a different one. Pass worktree:true (with ticket) or "
|
||||||
+ "worktree:<ticket-slug> to provision an isolated git worktree. Returns the member's "
|
+ "worktree:<ticket-slug> to provision an isolated git worktree. Pass resumeSessionId "
|
||||||
+ "sessionId (use with bridge_send) and paneId (use with bridge_stop).",
|
+ "to relaunch onto a prior conversation instead of starting cold — this requires an "
|
||||||
|
+ "explicit profile whose backend supports it (bridge_list shows agentSessionId for "
|
||||||
|
+ "resumable members), and is refused otherwise rather than silently starting fresh. "
|
||||||
|
+ "sessionName gives the member a display name in its own UI when the backend supports "
|
||||||
|
+ "one. Returns the member's sessionId (use with bridge_send) and paneId (use with "
|
||||||
|
+ "bridge_stop).",
|
||||||
objectSchema(Map.of(
|
objectSchema(Map.of(
|
||||||
"role", stringProp("What the member is for: architect, dev or reviewer (default dev)"),
|
"role", stringProp("What the member is for: architect, dev or reviewer (default dev)"),
|
||||||
"profile", stringProp("Which backend to run it on (omit for the default profile)"),
|
"profile", stringProp("Which backend to run it on (omit for the default profile)"),
|
||||||
"cwd", stringProp("Working directory for the member (omit to inherit yours)"),
|
"cwd", stringProp("Working directory for the member (omit to inherit yours)"),
|
||||||
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
|
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
|
||||||
"ticket", stringProp("Ticket slug when worktree:true")),
|
"ticket", stringProp("Ticket slug when worktree:true"),
|
||||||
|
"sessionName", stringProp("Logical display name for the member's own session, when its backend supports one"),
|
||||||
|
"resumeSessionId", stringProp("A prior member's agentSessionId (from bridge_list) to resume — requires an explicit profile that supports it")),
|
||||||
List.of()));
|
List.of()));
|
||||||
}
|
}
|
||||||
|
|
||||||
private static McpSchema.Tool profilesTool() {
|
private static McpSchema.Tool profilesTool() {
|
||||||
return tool("bridge_profiles",
|
return tool("bridge_profiles",
|
||||||
"List the configured worker profiles (backends) and which one bridge_spawn uses by default.",
|
"List the configured worker profiles (backends) and which one bridge_spawn uses by "
|
||||||
|
+ "default. A 'quarantined' map is present when a backend-exhausted refusal put "
|
||||||
|
+ "a profile's credential on cooldown — bridge_spawn onto it is refused until "
|
||||||
|
+ "quarantinedForSeconds elapses; a profile sharing that credential is listed too.",
|
||||||
objectSchema(Map.of(), List.of()));
|
objectSchema(Map.of(), List.of()));
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -952,8 +1046,15 @@ public final class BridgeMcp {
|
|||||||
+ "discover a peer lead without being told its address. 'members' are the "
|
+ "discover a peer lead without being told its address. 'members' are the "
|
||||||
+ "sessions delegated to — each with sessionId, paneId, role (architect/dev/"
|
+ "sessions delegated to — each with sessionId, paneId, role (architect/dev/"
|
||||||
+ "reviewer), profile (the backend it runs on), state, optional "
|
+ "reviewer), profile (the backend it runs on), state, optional "
|
||||||
+ "worktree/branch/owner, and live herdr status. An empty 'members' "
|
+ "worktree/branch/owner/agentSessionId (the id to pass as bridge_spawn's "
|
||||||
+ "means no members are spawned; it says nothing about peers.",
|
+ "resumeSessionId to relaunch onto that same conversation, when the backend "
|
||||||
|
+ "supports it), and live herdr status. An empty 'members' "
|
||||||
|
+ "means no members are spawned; it says nothing about peers. When capacity "
|
||||||
|
+ "facts are configured, a 'capacity' row per profile also reports free: 0 for "
|
||||||
|
+ "a quarantined profile's credential (see bridge_profiles), whatever its "
|
||||||
|
+ "maxLoad/live — with credentialId and quarantinedForSeconds naming the "
|
||||||
|
+ "quarantine, so 'free: 0, busy' can be told apart from 'free: 0, refusing "
|
||||||
|
+ "for N seconds'.",
|
||||||
objectSchema(Map.of(), List.of()));
|
objectSchema(Map.of(), List.of()));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -83,6 +83,21 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
|||||||
fleet);
|
fleet);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Production constructor, plus the CB-596 {@code memberCredentials} policy supplier.
|
||||||
|
*/
|
||||||
|
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||||
|
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||||
|
Function<String, String> env,
|
||||||
|
long spawnReadyTimeoutMs, long spawnReadyPollMs,
|
||||||
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
|
||||||
|
this(agents, spaces, guard, profiles, defaultProfile, env,
|
||||||
|
spawnReadyTimeoutMs,
|
||||||
|
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
|
||||||
|
fleet, memberCredentials);
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
|
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
|
||||||
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
|
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
|
||||||
@@ -125,6 +140,39 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
|||||||
this.guard = guard;
|
this.guard = guard;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Full testability constructor, plus the CB-596 {@code memberCredentials} policy supplier.
|
||||||
|
*/
|
||||||
|
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||||
|
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||||
|
Function<String, String> env,
|
||||||
|
long spawnReadyTimeoutMs,
|
||||||
|
LongSupplier nowMillis, Runnable sleeper,
|
||||||
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
|
||||||
|
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
|
||||||
|
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials);
|
||||||
|
this.guard = guard;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Full testability constructor, plus an injectable host-env-names source for the CB-596
|
||||||
|
* criterion-4 gap detector. Test seam only — every production call site leaves this at the
|
||||||
|
* default (the real {@code System.getenv()} key set) via the constructor above.
|
||||||
|
*/
|
||||||
|
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||||
|
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||||
|
Function<String, String> env,
|
||||||
|
long spawnReadyTimeoutMs,
|
||||||
|
LongSupplier nowMillis, Runnable sleeper,
|
||||||
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
Supplier<BridgedConfig.MemberCredentials> memberCredentials,
|
||||||
|
Supplier<Set<String>> hostEnvNames) {
|
||||||
|
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
|
||||||
|
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials, hostEnvNames);
|
||||||
|
this.guard = guard;
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* {@inheritDoc}
|
* {@inheritDoc}
|
||||||
*
|
*
|
||||||
|
|||||||
@@ -8,6 +8,7 @@ import dev.ltms.bridged.peer.PeerHandle;
|
|||||||
import dev.ltms.bridged.peer.PeerLauncher;
|
import dev.ltms.bridged.peer.PeerLauncher;
|
||||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||||
import dev.ltms.bridged.peer.SpawnRequest;
|
import dev.ltms.bridged.peer.SpawnRequest;
|
||||||
|
import dev.ltms.bridged.placement.BackendQuarantine;
|
||||||
import dev.ltms.bridged.placement.PlacementCandidate;
|
import dev.ltms.bridged.placement.PlacementCandidate;
|
||||||
import dev.ltms.bridged.placement.PlacementContext;
|
import dev.ltms.bridged.placement.PlacementContext;
|
||||||
import dev.ltms.bridged.placement.PlacementException;
|
import dev.ltms.bridged.placement.PlacementException;
|
||||||
@@ -23,10 +24,12 @@ import java.util.HashSet;
|
|||||||
import java.util.LinkedHashMap;
|
import java.util.LinkedHashMap;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
|
import java.util.Objects;
|
||||||
import java.util.Set;
|
import java.util.Set;
|
||||||
import java.util.concurrent.ConcurrentHashMap;
|
import java.util.concurrent.ConcurrentHashMap;
|
||||||
import java.util.function.Function;
|
import java.util.function.Function;
|
||||||
import java.util.function.Supplier;
|
import java.util.function.Supplier;
|
||||||
|
import java.util.stream.Collectors;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
|
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
|
||||||
@@ -78,6 +81,9 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
|||||||
private final Supplier<Map<String, BridgedConfig.Profile>> profileConfigs;
|
private final Supplier<Map<String, BridgedConfig.Profile>> profileConfigs;
|
||||||
private final Supplier<PlacementPolicy> placementPolicy;
|
private final Supplier<PlacementPolicy> placementPolicy;
|
||||||
|
|
||||||
|
/** CB-578 stage B: credential cooldown, checked before an explicit spawn and filtered into placement. */
|
||||||
|
private final BackendQuarantine quarantine;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* CB-557: the role pools an unqualified spawn draws its candidates from. A supplier that yields
|
* CB-557: the role pools an unqualified spawn draws its candidates from. A supplier that yields
|
||||||
* {@code null}, and an empty pool for a role, both fall back to every configured profile — the
|
* {@code null}, and an empty pool for a role, both fall back to every configured profile — the
|
||||||
@@ -99,7 +105,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Production constructor with a placement policy and live-worker counter.
|
* Production constructor with a placement policy and live-worker counter. Quarantine (CB-578
|
||||||
|
* stage B) is off for this constructor — {@link BackendQuarantine#none()} — since it predates
|
||||||
|
* the feature and existing callers of this exact overload never exercised it; use the 7-arg
|
||||||
|
* overload below to wire a real {@link BackendQuarantine}.
|
||||||
*
|
*
|
||||||
* @param delegates one adapter per configured peer kind; must be non-empty and declare
|
* @param delegates one adapter per configured peer kind; must be non-empty and declare
|
||||||
* disjoint profile-name sets
|
* disjoint profile-name sets
|
||||||
@@ -113,13 +122,14 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
|||||||
Map<String, BridgedConfig.Profile> profileConfigs,
|
Map<String, BridgedConfig.Profile> profileConfigs,
|
||||||
PlacementPolicy placementPolicy,
|
PlacementPolicy placementPolicy,
|
||||||
Function<String, Integer> liveCount) {
|
Function<String, Integer> liveCount) {
|
||||||
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, null);
|
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, null, BackendQuarantine.none());
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Production constructor with role pools (CB-557). An unqualified spawn draws its candidates from
|
* Production constructor with role pools (CB-557). An unqualified spawn draws its candidates from
|
||||||
* {@code fleet.<role>} instead of from every configured profile, so a reviewer is placed on a
|
* {@code fleet.<role>} instead of from every configured profile, so a reviewer is placed on a
|
||||||
* reviewer backend and never on, say, the architect-only one.
|
* reviewer backend and never on, say, the architect-only one. Quarantine is off for this
|
||||||
|
* constructor too, for the same reason as the 5-arg overload above.
|
||||||
*
|
*
|
||||||
* @param fleet the configured role pools; {@code null} ⇒ every profile is a candidate for every
|
* @param fleet the configured role pools; {@code null} ⇒ every profile is a candidate for every
|
||||||
* role, which is the pre-CB-557 behaviour
|
* role, which is the pre-CB-557 behaviour
|
||||||
@@ -130,29 +140,50 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
|||||||
PlacementPolicy placementPolicy,
|
PlacementPolicy placementPolicy,
|
||||||
Function<String, Integer> liveCount,
|
Function<String, Integer> liveCount,
|
||||||
BridgedConfig.Fleet fleet) {
|
BridgedConfig.Fleet fleet) {
|
||||||
|
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet, BackendQuarantine.none());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Production constructor with role pools and quarantine (CB-578 stage B). The full-featured
|
||||||
|
* non-reloading form; {@link #CompositePeerLauncher(List, String, Supplier, Function, BackendQuarantine)}
|
||||||
|
* is what {@code Bridged.main} actually wires up.
|
||||||
|
*
|
||||||
|
* @param quarantine required — pass {@link BackendQuarantine#none()} for a caller that does not
|
||||||
|
* want the feature, never a defaulting overload (CB-578 stage B's own rule).
|
||||||
|
*/
|
||||||
|
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
|
||||||
|
String defaultProfile,
|
||||||
|
Map<String, BridgedConfig.Profile> profileConfigs,
|
||||||
|
PlacementPolicy placementPolicy,
|
||||||
|
Function<String, Integer> liveCount,
|
||||||
|
BridgedConfig.Fleet fleet,
|
||||||
|
BackendQuarantine quarantine) {
|
||||||
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
|
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
|
||||||
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
|
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
|
||||||
// differ from one JVM run to the next.
|
// differ from one JVM run to the next.
|
||||||
this(delegates, defaultProfile,
|
this(delegates, defaultProfile,
|
||||||
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
|
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
|
||||||
constant(placementPolicy), liveCount, constant(fleet));
|
constant(placementPolicy), liveCount, constant(fleet), quarantine);
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Production constructor that re-reads its placement inputs per spawn (CB-559), so a config
|
* Production constructor that re-reads its placement inputs per spawn (CB-559), so a config
|
||||||
* reload retargets the next member without a restart.
|
* reload retargets the next member without a restart.
|
||||||
*
|
*
|
||||||
* @param config the live configuration — read at every spawn, never captured
|
* @param config the live configuration — read at every spawn, never captured
|
||||||
|
* @param quarantine required — CB-578 stage B; pass {@link BackendQuarantine#none()} to opt out
|
||||||
*/
|
*/
|
||||||
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
|
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
|
||||||
String defaultProfile,
|
String defaultProfile,
|
||||||
Supplier<BridgedConfig> config,
|
Supplier<BridgedConfig> config,
|
||||||
Function<String, Integer> liveCount) {
|
Function<String, Integer> liveCount,
|
||||||
|
BackendQuarantine quarantine) {
|
||||||
this(delegates, defaultProfile,
|
this(delegates, defaultProfile,
|
||||||
() -> config.get().profiles(),
|
() -> config.get().profiles(),
|
||||||
() -> PlacementPolicies.fromName(config.get().placement()),
|
() -> PlacementPolicies.fromName(config.get().placement()),
|
||||||
liveCount,
|
liveCount,
|
||||||
() -> config.get().fleet());
|
() -> config.get().fleet(),
|
||||||
|
quarantine);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** The all-suppliers form every other constructor funnels into. */
|
/** The all-suppliers form every other constructor funnels into. */
|
||||||
@@ -161,8 +192,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
|||||||
Supplier<Map<String, BridgedConfig.Profile>> profileConfigs,
|
Supplier<Map<String, BridgedConfig.Profile>> profileConfigs,
|
||||||
Supplier<PlacementPolicy> placementPolicy,
|
Supplier<PlacementPolicy> placementPolicy,
|
||||||
Function<String, Integer> liveCount,
|
Function<String, Integer> liveCount,
|
||||||
Supplier<BridgedConfig.Fleet> fleet) {
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
BackendQuarantine quarantine) {
|
||||||
this.fleet = fleet;
|
this.fleet = fleet;
|
||||||
|
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
|
||||||
if (delegates.isEmpty()) {
|
if (delegates.isEmpty()) {
|
||||||
throw new IllegalArgumentException("at least one peer adapter must be configured");
|
throw new IllegalArgumentException("at least one peer adapter must be configured");
|
||||||
}
|
}
|
||||||
@@ -225,6 +258,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
|||||||
// the charter makes explicit-profile spawns the normal path — so skipping the check
|
// the charter makes explicit-profile spawns the normal path — so skipping the check
|
||||||
// here would leave the cap dead config in real operation.
|
// here would leave the cap dead config in real operation.
|
||||||
HerdrPeerLauncher d = route(requestedProfile);
|
HerdrPeerLauncher d = route(requestedProfile);
|
||||||
|
enforceNotQuarantined(requestedProfile);
|
||||||
enforceMaxLoad(requestedProfile);
|
enforceMaxLoad(requestedProfile);
|
||||||
PeerHandle handle = d.spawn(req);
|
PeerHandle handle = d.spawn(req);
|
||||||
spawnedBy.put(handle.id(), d);
|
spawnedBy.put(handle.id(), d);
|
||||||
@@ -238,7 +272,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
|||||||
List<PlacementCandidate> candidates = candidates(req.role());
|
List<PlacementCandidate> candidates = candidates(req.role());
|
||||||
String roleDefault = defaultProfileFor(req.role());
|
String roleDefault = defaultProfileFor(req.role());
|
||||||
Set<String> unreachable = new HashSet<>();
|
Set<String> unreachable = new HashSet<>();
|
||||||
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable);
|
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
|
||||||
|
// call, so re-deriving it per retry would only cost work, never change the answer.
|
||||||
|
Set<String> quarantined = quarantinedProfiles(candidates);
|
||||||
|
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable, quarantined);
|
||||||
|
|
||||||
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
|
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
|
||||||
for (int attempt = 0; attempt < maxAttempts; attempt++) {
|
for (int attempt = 0; attempt < maxAttempts; attempt++) {
|
||||||
@@ -268,7 +305,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
|||||||
chosen.profile(), e.getMessage());
|
chosen.profile(), e.getMessage());
|
||||||
unreachable.add(chosen.profile());
|
unreachable.add(chosen.profile());
|
||||||
// Update the context for the next selection so the policy excludes this profile.
|
// Update the context for the next selection so the policy excludes this profile.
|
||||||
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable);
|
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable, quarantined);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -298,9 +335,48 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
|||||||
* @param profile the profile the caller explicitly named
|
* @param profile the profile the caller explicitly named
|
||||||
* @throws PlacementException when the profile is at capacity
|
* @throws PlacementException when the profile is at capacity
|
||||||
*/
|
*/
|
||||||
|
/**
|
||||||
|
* Refuse an explicit-profile spawn whose credential is quarantined (CB-578 stage B): a prior
|
||||||
|
* {@code BACKEND_EXHAUSTED} classification on this profile, or on another profile sharing its
|
||||||
|
* {@code credentialId}, is still on cooldown.
|
||||||
|
*
|
||||||
|
* <p>Checked before {@link #enforceMaxLoad}, and for the same reason that check exists: an
|
||||||
|
* explicit profile bypasses the placement policy's filtering entirely, so without this the cap
|
||||||
|
* (there, quarantine here) would be dead config on the very path the charter calls normal.
|
||||||
|
* Deliberately no fallback to another profile, matching {@link #enforceMaxLoad}'s own reasoning —
|
||||||
|
* the caller named this profile for a cost/model reason.
|
||||||
|
*
|
||||||
|
* @throws PlacementException naming the profile, its credential, and the remaining cooldown
|
||||||
|
*/
|
||||||
|
private void enforceNotQuarantined(String profile) {
|
||||||
|
String credentialId = credentialIdFor(profile);
|
||||||
|
quarantine.remainingSeconds(credentialId).ifPresent(remaining -> {
|
||||||
|
throw new PlacementException("worker profile '" + profile + "' is quarantined "
|
||||||
|
+ "(credential '" + credentialId + "' exhausted; ~" + remaining
|
||||||
|
+ "s remaining) — refusing spawn");
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
/** {@code profile}'s credential group (CB-578 stage B), or the profile's own name if unconfigured. */
|
||||||
|
private String credentialIdFor(String profile) {
|
||||||
|
BridgedConfig.Profile cfg = profiles0().get(profile);
|
||||||
|
return cfg == null ? profile : cfg.effectiveCredentialId();
|
||||||
|
}
|
||||||
|
|
||||||
|
/** The subset of {@code candidates} whose credential is currently quarantined (CB-578 stage B). */
|
||||||
|
private Set<String> quarantinedProfiles(List<PlacementCandidate> candidates) {
|
||||||
|
return candidates.stream()
|
||||||
|
.map(PlacementCandidate::profile)
|
||||||
|
.filter(p -> quarantine.isQuarantined(credentialIdFor(p)))
|
||||||
|
.collect(Collectors.toSet());
|
||||||
|
}
|
||||||
|
|
||||||
private void enforceMaxLoad(String profile) {
|
private void enforceMaxLoad(String profile) {
|
||||||
// Absent config, or a config whose maxLoad normalized to null (non-positive ⇒ unlimited at
|
// Absent config, or a config whose maxLoad normalized to null (ABSENT ⇒ unlimited at load),
|
||||||
// load), means no cap — never cap what wasn't configured.
|
// means no cap — never cap what wasn't configured. Note "non-positive ⇒ unlimited" was true
|
||||||
|
// until CB-585: an explicit `maxLoad: 0` now survives as 0 and is a real cap of zero, so the
|
||||||
|
// check below refuses every spawn on that profile, and a negative value is refused at config
|
||||||
|
// load rather than normalized away.
|
||||||
BridgedConfig.Profile cfg = profiles0().get(profile);
|
BridgedConfig.Profile cfg = profiles0().get(profile);
|
||||||
Integer cap = (cfg == null) ? null : cfg.maxLoad();
|
Integer cap = (cfg == null) ? null : cfg.maxLoad();
|
||||||
if (cap == null) {
|
if (cap == null) {
|
||||||
@@ -387,6 +463,18 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
|||||||
return defaultProfile;
|
return defaultProfile;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* {@inheritDoc}
|
||||||
|
*
|
||||||
|
* <p>Routes to the specific delegate {@code profileName} resolves to, not the fleet-wide
|
||||||
|
* union {@link #capabilities()} returns — the whole reason this method exists (CB-584): in a
|
||||||
|
* mixed fleet, one adapter's capability must never be read as every profile's.
|
||||||
|
*/
|
||||||
|
@Override
|
||||||
|
public Set<Capability> capabilitiesFor(String profileName) {
|
||||||
|
return route(profileName).capabilities();
|
||||||
|
}
|
||||||
|
|
||||||
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
|
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
|
||||||
@Override
|
@Override
|
||||||
public List<Agent> list() {
|
public List<Agent> list() {
|
||||||
|
|||||||
@@ -7,6 +7,7 @@ import dev.ltms.bridged.herdr.HerdrException;
|
|||||||
import dev.ltms.bridged.herdr.Tab;
|
import dev.ltms.bridged.herdr.Tab;
|
||||||
import dev.ltms.bridged.herdr.Workspace;
|
import dev.ltms.bridged.herdr.Workspace;
|
||||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||||
|
import dev.ltms.bridged.peer.Capability;
|
||||||
import dev.ltms.bridged.peer.CharterReceipt;
|
import dev.ltms.bridged.peer.CharterReceipt;
|
||||||
import dev.ltms.bridged.peer.MemberRole;
|
import dev.ltms.bridged.peer.MemberRole;
|
||||||
import dev.ltms.bridged.peer.PeerHandle;
|
import dev.ltms.bridged.peer.PeerHandle;
|
||||||
@@ -19,6 +20,7 @@ import org.slf4j.LoggerFactory;
|
|||||||
import java.security.SecureRandom;
|
import java.security.SecureRandom;
|
||||||
import java.util.ArrayList;
|
import java.util.ArrayList;
|
||||||
import java.util.Collection;
|
import java.util.Collection;
|
||||||
|
import java.util.HashSet;
|
||||||
import java.util.LinkedHashMap;
|
import java.util.LinkedHashMap;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
@@ -89,6 +91,27 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
*/
|
*/
|
||||||
private final Supplier<BridgedConfig.Fleet> fleet;
|
private final Supplier<BridgedConfig.Fleet> fleet;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: the live {@code memberCredentials:} policy, read once per spawn (same hot-reload shape
|
||||||
|
* as {@link #fleet}). {@code null} — either the supplier itself, or what it returns — means no
|
||||||
|
* policy is configured and {@link #applyMemberCredentialPolicy} shadows nothing.
|
||||||
|
*/
|
||||||
|
private final Supplier<BridgedConfig.MemberCredentials> memberCredentials;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Enumerates the daemon's own process environment variable NAMES ONLY, never values — the CB-596
|
||||||
|
* criterion-4 gap detector's data source (see {@link #logCredentialGap}). Injectable for tests;
|
||||||
|
* production always resolves to the real {@code System.getenv()} key set.
|
||||||
|
*
|
||||||
|
* <p>Deliberately the daemon's own environment, not the spawned pane's: nothing in the herdr
|
||||||
|
* client surface lets the daemon read back an arbitrary command's output from a pane before the
|
||||||
|
* peer starts in it, so there is no channel to inspect the pane's environment directly. The
|
||||||
|
* daemon's own process is started the same way (a login shell sourcing the same secret store —
|
||||||
|
* see CB-592's investigation of {@code secrets.sh}), so on a single-host deployment its env
|
||||||
|
* mirrors what the pane's login shell is about to export.
|
||||||
|
*/
|
||||||
|
private final Supplier<Set<String>> hostEnvNames;
|
||||||
|
|
||||||
/** The final instruction always requires a bridge reply when the bridge MCP is mounted. */
|
/** The final instruction always requires a bridge reply when the bridge MCP is mounted. */
|
||||||
protected static final String REPLY_CHARTER =
|
protected static final String REPLY_CHARTER =
|
||||||
"You are a spawned member in the claude-bridge fleet. Every message you receive arrives "
|
"You are a spawned member in the claude-bridge fleet. Every message you receive arrives "
|
||||||
@@ -165,6 +188,41 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
long spawnReadyTimeoutMs,
|
long spawnReadyTimeoutMs,
|
||||||
LongSupplier nowMillis, Runnable sleeper,
|
LongSupplier nowMillis, Runnable sleeper,
|
||||||
Supplier<BridgedConfig.Fleet> fleet) {
|
Supplier<BridgedConfig.Fleet> fleet) {
|
||||||
|
this(namePrefix, agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
|
||||||
|
nowMillis, sleeper, fleet, null);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* As above, plus the live {@code memberCredentials} policy (CB-596).
|
||||||
|
*
|
||||||
|
* @param memberCredentials live member-credential policy, read once per spawn; {@code null} ⇒
|
||||||
|
* no policy configured, so a spawn shadows nothing. A separate
|
||||||
|
* constructor rather than a new parameter on the one above, so every
|
||||||
|
* existing call site keeps the pre-CB-596 default without an edit.
|
||||||
|
*/
|
||||||
|
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
|
||||||
|
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||||
|
Function<String, String> env,
|
||||||
|
long spawnReadyTimeoutMs,
|
||||||
|
LongSupplier nowMillis, Runnable sleeper,
|
||||||
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
|
||||||
|
this(namePrefix, agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
|
||||||
|
nowMillis, sleeper, fleet, memberCredentials, null);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* As above, plus an injectable {@link #hostEnvNames} source for the CB-596 gap detector. Test
|
||||||
|
* seam only — every production call site leaves this {@code null} and gets the real host env.
|
||||||
|
*/
|
||||||
|
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
|
||||||
|
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||||
|
Function<String, String> env,
|
||||||
|
long spawnReadyTimeoutMs,
|
||||||
|
LongSupplier nowMillis, Runnable sleeper,
|
||||||
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
Supplier<BridgedConfig.MemberCredentials> memberCredentials,
|
||||||
|
Supplier<Set<String>> hostEnvNames) {
|
||||||
this.fleet = fleet;
|
this.fleet = fleet;
|
||||||
this.namePrefix = namePrefix;
|
this.namePrefix = namePrefix;
|
||||||
this.agents = agents;
|
this.agents = agents;
|
||||||
@@ -175,6 +233,8 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
this.spawnReadyTimeoutMs = spawnReadyTimeoutMs;
|
this.spawnReadyTimeoutMs = spawnReadyTimeoutMs;
|
||||||
this.nowMillis = nowMillis;
|
this.nowMillis = nowMillis;
|
||||||
this.sleeper = sleeper;
|
this.sleeper = sleeper;
|
||||||
|
this.memberCredentials = memberCredentials;
|
||||||
|
this.hostEnvNames = hostEnvNames != null ? hostEnvNames : () -> System.getenv().keySet();
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- adapter seams -------------------------------------------------------------------------
|
// --- adapter seams -------------------------------------------------------------------------
|
||||||
@@ -248,6 +308,19 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
return defaultProfile;
|
return defaultProfile;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* {@inheritDoc}
|
||||||
|
*
|
||||||
|
* <p>One {@link HerdrPeerLauncher} instance always serves exactly one adapter kind, so every
|
||||||
|
* profile it owns shares that adapter's {@link #capabilities()} — {@code profileName} only
|
||||||
|
* needs validating (throwing on an unknown profile, same as {@link #spawn}), not routing.
|
||||||
|
*/
|
||||||
|
@Override
|
||||||
|
public Set<Capability> capabilitiesFor(String profileName) {
|
||||||
|
requireProfile(profileName);
|
||||||
|
return capabilities();
|
||||||
|
}
|
||||||
|
|
||||||
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
|
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
|
||||||
protected Collection<BridgedConfig.Profile> profileConfigs() {
|
protected Collection<BridgedConfig.Profile> profileConfigs() {
|
||||||
return profiles.values();
|
return profiles.values();
|
||||||
@@ -752,10 +825,62 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/** A fresh mutable env map — the conventional starting point for {@link #buildLaunch}. */
|
/**
|
||||||
|
* CB-596: overlay value that shadows any host credential a herdr pane otherwise inherits from
|
||||||
|
* herdr's own login-shell process environment (gitea issue #82, superseding CB-592's single
|
||||||
|
* hardcoded {@code GITEA_ACCESS_TOKEN} name — see {@link #applyMemberCredentialPolicy}). herdr
|
||||||
|
* spawns a pane from its <em>own</em> process environment and layers our map on top —
|
||||||
|
* {@link dev.ltms.bridged.herdr.WorkspaceControl#createTab} and {@code #splitPane} send only
|
||||||
|
* the keys we put in that map, so any key we never mention passes straight through from
|
||||||
|
* herdr's own shell, admin credentials included.
|
||||||
|
*
|
||||||
|
* <p>Deliberately a non-blank sentinel, not {@code ""}. Whether an empty-string overlay value
|
||||||
|
* overrides an inherited variable or is skipped as blank could not be settled by reading this
|
||||||
|
* codebase — herdr's server-side merge is an external process, not something in this repo.
|
||||||
|
* A non-blank replacement sidesteps that ambiguity entirely: {@link #baseEnv}'s own {@code
|
||||||
|
* PATH} seeding already depends on the overlay reliably replacing an inherited value (see its
|
||||||
|
* javadoc), and that is only demonstrated for a non-blank value, so this reuses the same,
|
||||||
|
* proven-reliable shape rather than the unverified one.
|
||||||
|
*
|
||||||
|
* <p><b>MEASURED ON A LIVE PANE, 2026-08-15 (CB-592): this sentinel alone does NOT hold.</b> The
|
||||||
|
* overlay itself works — {@code GITEA_TOKEN} is injected here, is exported by no shell file, and
|
||||||
|
* does reach the pane. The sentinel loses one step later. A herdr pane runs a <em>login</em>
|
||||||
|
* shell, {@code ~/.zprofile} sources {@code ${SHARED_ENV}/tools/secrets.sh}, and that file does a
|
||||||
|
* plain unconditional {@code export GITEA_ACCESS_TOKEN=...}. A login shell overwrites a value
|
||||||
|
* already in the environment, so the real admin token is put back over this sentinel before the
|
||||||
|
* member process ever starts. That defeat applies to <em>every</em> name {@code secrets.sh}
|
||||||
|
* exports, and no launcher-side overlay can win against it.
|
||||||
|
*
|
||||||
|
* <p>So this constant is not the control on its own — {@link #MEMBER_MARKER} is the other half.
|
||||||
|
* Keeping the sentinel is still worth it: it is correct for any peer kind whose pane does not
|
||||||
|
* start a login shell, and it makes the intent explicit at the one place every adapter passes.
|
||||||
|
*/
|
||||||
|
private static final String BLOCKED_CREDENTIAL_SENTINEL =
|
||||||
|
"blocked-by-bridged-cb596-see-gitea-issue-82";
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-592: marks a pane as a bridged member so a shell startup file can decline to export
|
||||||
|
* operator-only credentials into it (gitea issue #77).
|
||||||
|
*
|
||||||
|
* <p>This name is deliberately one that {@code secrets.sh} never exports, which is exactly why
|
||||||
|
* it survives the login shell that wipes {@link #BLOCKED_GITEA_ACCESS_TOKEN}. The mechanism is
|
||||||
|
* measured, not assumed: {@code GITEA_TOKEN} is injected the same way, is absent from a login
|
||||||
|
* shell of its own, and was observed set inside a live member pane.
|
||||||
|
*
|
||||||
|
* <p>It is a no-op until the operator guards the export, which is a one-line change in a file
|
||||||
|
* this repo does not own and must not edit unasked:
|
||||||
|
*
|
||||||
|
* <pre>{@code
|
||||||
|
* [ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...
|
||||||
|
* }</pre>
|
||||||
|
*
|
||||||
|
* <p>Setting the marker now costs nothing and means that edit is the whole remaining fix.
|
||||||
|
*/
|
||||||
|
static final String MEMBER_MARKER = "BRIDGED_MEMBER";
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Seed a worker's environment (CB-511): the daemon's own {@code PATH}, then the profile's
|
* Seed a worker's environment (CB-511): the daemon's own {@code PATH}, then the profile's
|
||||||
* {@code env:} entries.
|
* {@code env:} entries, then the CB-592 admin-token shadow.
|
||||||
*
|
*
|
||||||
* <p>Why this exists: bridged passes herdr an explicit env map, and herdr merges it into
|
* <p>Why this exists: bridged passes herdr an explicit env map, and herdr merges it into
|
||||||
* <em>its own</em> process environment. So before this, a worker inherited whatever PATH the
|
* <em>its own</em> process environment. So before this, a worker inherited whatever PATH the
|
||||||
@@ -769,6 +894,12 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
* overriding {@code ANTHROPIC_BASE_URL} and slipping past {@link
|
* overriding {@code ANTHROPIC_BASE_URL} and slipping past {@link
|
||||||
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
|
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
|
||||||
* {@code baseUrl} and nothing else.
|
* {@code baseUrl} and nothing else.
|
||||||
|
*
|
||||||
|
* <p>The CB-596 credential shadow and the CB-592 marker are put in <em>last</em>, after the
|
||||||
|
* profile's own {@code env:}, so no profile — present or future — can restore a blocked
|
||||||
|
* credential, or hide that the pane is a member, by naming either in config. This is the one
|
||||||
|
* place both are applied: every {@code buildLaunch} in every adapter calls this first, so a new
|
||||||
|
* profile, and a peer kind not yet written, gets them for free.
|
||||||
*/
|
*/
|
||||||
protected Map<String, String> baseEnv(BridgedConfig.Profile cfg) {
|
protected Map<String, String> baseEnv(BridgedConfig.Profile cfg) {
|
||||||
Map<String, String> workerEnv = new LinkedHashMap<>();
|
Map<String, String> workerEnv = new LinkedHashMap<>();
|
||||||
@@ -779,9 +910,70 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
if (cfg != null && cfg.env() != null) {
|
if (cfg != null && cfg.env() != null) {
|
||||||
workerEnv.putAll(cfg.env());
|
workerEnv.putAll(cfg.env());
|
||||||
}
|
}
|
||||||
|
applyMemberCredentialPolicy(workerEnv);
|
||||||
|
workerEnv.put(MEMBER_MARKER, "1");
|
||||||
return workerEnv;
|
return workerEnv;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: shadow every configured {@code memberCredentials.known} name that is not also
|
||||||
|
* {@code allow}-ed, replacing CB-592's single hardcoded {@code GITEA_ACCESS_TOKEN} name (gitea
|
||||||
|
* issue #82). An allow-listed name is deliberately left unmentioned here — see {@link
|
||||||
|
* #BLOCKED_CREDENTIAL_SENTINEL}'s javadoc for why an overlay entry is the only way to shadow an
|
||||||
|
* inherited value, which is exactly why an allowed name must get NO entry: any entry at all,
|
||||||
|
* blank or not, risks overriding the real value the pane needs.
|
||||||
|
*
|
||||||
|
* <p>No {@code memberCredentials} configured — the supplier is {@code null}, or it resolves to
|
||||||
|
* one whose {@code known} list is empty — shadows nothing. This is a real, config-driven gap
|
||||||
|
* (see {@link BridgedConfig.MemberCredentials}'s javadoc), not a safe default: deny-by-default
|
||||||
|
* only defends names the operator has actually enumerated in {@code known}.
|
||||||
|
*/
|
||||||
|
private void applyMemberCredentialPolicy(Map<String, String> workerEnv) {
|
||||||
|
BridgedConfig.MemberCredentials creds = memberCredentials == null ? null : memberCredentials.get();
|
||||||
|
if (creds == null) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
for (String name : creds.blockedSet()) {
|
||||||
|
workerEnv.put(name, BLOCKED_CREDENTIAL_SENTINEL);
|
||||||
|
}
|
||||||
|
logCredentialGap(creds);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Credential-shaped env var name heuristic for {@link #logCredentialGap} — case-insensitive. */
|
||||||
|
private static final Pattern CREDENTIAL_SHAPED_NAME =
|
||||||
|
Pattern.compile("(?i).*(TOKEN|SECRET|_KEY|APIKEY|PASSWORD|CREDENTIAL|AUTH).*");
|
||||||
|
|
||||||
|
/** Guards {@link #logCredentialGap} to one WARN per launcher instance, not one per spawn. */
|
||||||
|
private final AtomicBoolean credentialGapLogged = new AtomicBoolean();
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596 criterion 4: a credential-shaped host env var name on neither {@code known} nor
|
||||||
|
* {@code allow} is not silently allowed — it is reported. {@link #hostEnvNames} enumerates the
|
||||||
|
* daemon's own environment (see that field's javadoc for why the daemon's env is read rather
|
||||||
|
* than the spawned pane's, which the daemon has no channel to inspect at spawn time); this logs
|
||||||
|
* every such NAME, at WARN, at most once per launcher instance — never a value, a prefix of a
|
||||||
|
* value, or a hash of a value, so the log itself cannot leak anything.
|
||||||
|
*/
|
||||||
|
private void logCredentialGap(BridgedConfig.MemberCredentials creds) {
|
||||||
|
Set<String> covered = new HashSet<>(creds.known());
|
||||||
|
covered.addAll(creds.allow());
|
||||||
|
List<String> gap = hostEnvNames.get().stream()
|
||||||
|
.filter(name -> CREDENTIAL_SHAPED_NAME.matcher(name).matches())
|
||||||
|
.filter(name -> !covered.contains(name))
|
||||||
|
.sorted()
|
||||||
|
.toList();
|
||||||
|
if (gap.isEmpty()) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (credentialGapLogged.compareAndSet(false, true)) {
|
||||||
|
log.warn("memberCredentials gap: {} credential-shaped env var name(s) are on neither "
|
||||||
|
+ "known: nor allow: — every member pane inherits them UNBLOCKED — {}. "
|
||||||
|
+ "Add each to memberCredentials.known (blocked by default) or .allow "
|
||||||
|
+ "(if a member legitimately needs it).",
|
||||||
|
gap.size(), gap);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/** Defensive copy of {@code argv} plus room to append launch flags. */
|
/** Defensive copy of {@code argv} plus room to append launch flags. */
|
||||||
protected static List<String> mutableArgv(List<String> argv) {
|
protected static List<String> mutableArgv(List<String> argv) {
|
||||||
return new ArrayList<>(argv);
|
return new ArrayList<>(argv);
|
||||||
|
|||||||
@@ -104,6 +104,20 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
|||||||
defaultConfigRoot(), defaultDiscoveryRoot(), fleet);
|
defaultConfigRoot(), defaultDiscoveryRoot(), fleet);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Production constructor, plus the CB-596 {@code memberCredentials} policy supplier.
|
||||||
|
*/
|
||||||
|
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
|
||||||
|
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||||
|
Function<String, String> env,
|
||||||
|
long spawnReadyTimeoutMs, long spawnReadyPollMs,
|
||||||
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
|
||||||
|
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
|
||||||
|
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
|
||||||
|
defaultConfigRoot(), defaultDiscoveryRoot(), fleet, memberCredentials);
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply a
|
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply a
|
||||||
* fake clock ({@code nowMillis}), poll-loop wait ({@code sleeper}), and a temp {@code configRoot}
|
* fake clock ({@code nowMillis}), poll-loop wait ({@code sleeper}), and a temp {@code configRoot}
|
||||||
@@ -151,6 +165,23 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
|||||||
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
|
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Full testability constructor, plus the CB-596 {@code memberCredentials} policy supplier.
|
||||||
|
*/
|
||||||
|
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
|
||||||
|
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||||
|
Function<String, String> env,
|
||||||
|
long spawnReadyTimeoutMs,
|
||||||
|
LongSupplier nowMillis, Runnable sleeper,
|
||||||
|
Path configRoot, Path discoveryRoot,
|
||||||
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
|
||||||
|
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
|
||||||
|
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials);
|
||||||
|
this.configRoot = configRoot;
|
||||||
|
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
|
||||||
|
}
|
||||||
|
|
||||||
private static Path defaultConfigRoot() {
|
private static Path defaultConfigRoot() {
|
||||||
return Path.of(System.getProperty("java.io.tmpdir"));
|
return Path.of(System.getProperty("java.io.tmpdir"));
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -7,6 +7,7 @@ import com.rabbitmq.client.ConnectionFactory;
|
|||||||
import com.rabbitmq.client.DeliverCallback;
|
import com.rabbitmq.client.DeliverCallback;
|
||||||
import com.rabbitmq.client.Recoverable;
|
import com.rabbitmq.client.Recoverable;
|
||||||
import com.rabbitmq.client.RecoveryListener;
|
import com.rabbitmq.client.RecoveryListener;
|
||||||
|
import com.rabbitmq.client.Return;
|
||||||
import org.slf4j.Logger;
|
import org.slf4j.Logger;
|
||||||
import org.slf4j.LoggerFactory;
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
@@ -14,7 +15,13 @@ import java.io.IOException;
|
|||||||
import java.nio.charset.StandardCharsets;
|
import java.nio.charset.StandardCharsets;
|
||||||
import java.util.LinkedHashMap;
|
import java.util.LinkedHashMap;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
|
import java.util.NavigableMap;
|
||||||
|
import java.util.concurrent.CompletableFuture;
|
||||||
import java.util.concurrent.ConcurrentHashMap;
|
import java.util.concurrent.ConcurrentHashMap;
|
||||||
|
import java.util.concurrent.ConcurrentSkipListMap;
|
||||||
|
import java.util.concurrent.ExecutionException;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.TimeoutException;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
|
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
|
||||||
@@ -29,6 +36,27 @@ import java.util.concurrent.ConcurrentHashMap;
|
|||||||
* leaves them on the broker — it redelivers on reconnect. That is the durability the in-memory
|
* leaves them on the broker — it redelivers on reconnect. That is the durability the in-memory
|
||||||
* adapter cannot give, with the port contract preserved.
|
* adapter cannot give, with the port contract preserved.
|
||||||
*
|
*
|
||||||
|
* <p><strong>Prefetch bounds the held backlog (CB-527).</strong> The consumer channel calls
|
||||||
|
* {@code basicQos} with a configurable prefetch count ({@link #DEFAULT_PREFETCH} unless the caller
|
||||||
|
* passes another value to {@link #open(String, int)}) before starting any consumer. Without a bound,
|
||||||
|
* the broker pushes its entire queue into {@link #held} the instant a target is {@link #own owned},
|
||||||
|
* so an undrained primary grows the JVM heap without limit and any queue-level control
|
||||||
|
* ({@code x-max-length}, per-message TTL) never fires because the queue never actually holds a
|
||||||
|
* backlog. Prefetch keeps the backlog where it is visible — on the broker — until the owner drains it.
|
||||||
|
*
|
||||||
|
* <p><strong>Publishes require a confirmed, routable delivery (CB-528).</strong> {@link #publish}
|
||||||
|
* runs on a channel separate from the consume/ack channel ({@link #channel}), so a slow or blocked
|
||||||
|
* publish confirm can never hold {@link #channelLock} and stall an ack — the ack path never waits on
|
||||||
|
* a publish confirm. That publish channel is in publisher-confirm mode and every publish sets the
|
||||||
|
* {@code mandatory} flag, so an unroutable publish (queue not declared, e.g. the owner never called
|
||||||
|
* {@link #own}) is returned by the broker instead of silently dropped. The broker sends the
|
||||||
|
* <em>return</em> for an unroutable message before the <em>confirm</em> that covers it — the ack/nack
|
||||||
|
* callback checks the returned-set at confirm time rather than assuming an ack means routed — so
|
||||||
|
* "confirmed" here means "durably queued", not merely "accepted by the broker". A returned or nacked
|
||||||
|
* (or un-confirmed within the timeout) publish surfaces as an {@link IllegalStateException} on the
|
||||||
|
* caller's thread; the caller — {@link MessageService#reply} — must not report success for a
|
||||||
|
* black-holed reply.
|
||||||
|
*
|
||||||
* <p><strong>Ownership is explicit.</strong> {@link #own} declares the queue and starts the consumer;
|
* <p><strong>Ownership is explicit.</strong> {@link #own} declares the queue and starts the consumer;
|
||||||
* {@link #release} cancels it. {@link #publish} sends to the queue but does <em>not</em> imply ownership
|
* {@link #release} cancels it. {@link #publish} sends to the queue but does <em>not</em> imply ownership
|
||||||
* and does not attach a consumer. This split is required by CB-308 federation, where one gateway may
|
* and does not attach a consumer. This split is required by CB-308 federation, where one gateway may
|
||||||
@@ -54,6 +82,12 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
|||||||
private static final String QUEUE_PREFIX = "agent.";
|
private static final String QUEUE_PREFIX = "agent.";
|
||||||
private static final String QUEUE_SUFFIX = ".inbox";
|
private static final String QUEUE_SUFFIX = ".inbox";
|
||||||
|
|
||||||
|
/** CB-527: the prefetch used when a caller does not pass an explicit value to {@link #open(String, int)}. */
|
||||||
|
public static final int DEFAULT_PREFETCH = 32;
|
||||||
|
|
||||||
|
/** How long {@link #publish} waits for its publisher confirm before failing the call (CB-528). */
|
||||||
|
private static final long CONFIRM_TIMEOUT_MS = 10_000L;
|
||||||
|
|
||||||
private final Connection connection;
|
private final Connection connection;
|
||||||
private final Channel channel;
|
private final Channel channel;
|
||||||
/** All channel operations (publish/declare/ack/cancel) serialize on this — a Channel is not thread-safe. */
|
/** All channel operations (publish/declare/ack/cancel) serialize on this — a Channel is not thread-safe. */
|
||||||
@@ -63,39 +97,87 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
|||||||
/** Targets whose queue is declared and consumer is running, mapped to their broker consumer tag. */
|
/** Targets whose queue is declared and consumer is running, mapped to their broker consumer tag. */
|
||||||
private final ConcurrentHashMap<String, String> consumerTags = new ConcurrentHashMap<>();
|
private final ConcurrentHashMap<String, String> consumerTags = new ConcurrentHashMap<>();
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-528: a dedicated channel for {@link #publish}, kept separate from {@link #channel} (consume
|
||||||
|
* + ack) so a publish confirm round trip never blocks under {@link #channelLock} and stalls an ack.
|
||||||
|
*/
|
||||||
|
private final Channel publishChannel;
|
||||||
|
private final Object publishChannelLock = new Object();
|
||||||
|
/** In-flight publishes awaiting their confirm, keyed by the publish channel's sequence number. */
|
||||||
|
private final ConcurrentSkipListMap<Long, Pending> pendingBySeq = new ConcurrentSkipListMap<>();
|
||||||
|
/**
|
||||||
|
* The same in-flight publishes, keyed by {@code msgId} — a broker {@code Return} carries no delivery
|
||||||
|
* tag. Assumes {@code msgId} is unique per in-flight publish: a second {@link #publish} for a
|
||||||
|
* {@code msgId} still awaiting its confirm would overwrite this entry and misdirect
|
||||||
|
* {@link #onReturn}'s lookup. Not reachable today — {@code MessageService.reply} generates a fresh
|
||||||
|
* {@code UUID} per call — so no guard is added for it.
|
||||||
|
*/
|
||||||
|
private final ConcurrentHashMap<String, Pending> pendingByMsgId = new ConcurrentHashMap<>();
|
||||||
|
|
||||||
/** A message pulled off the broker but not yet acked: its delivery-tag plus the port payload. */
|
/** A message pulled off the broker but not yet acked: its delivery-tag plus the port payload. */
|
||||||
private record Held(long deliveryTag, InboxMessage message) {}
|
private record Held(long deliveryTag, InboxMessage message) {}
|
||||||
|
|
||||||
/** Connect to {@code uri} (e.g. {@code amqp://guest:guest@127.0.0.1:5672/}) and open the inbox. */
|
/** A publish awaiting its confirm; {@link #returned} records whether the broker already returned it. */
|
||||||
|
private static final class Pending {
|
||||||
|
final String msgId;
|
||||||
|
final CompletableFuture<Void> confirmed = new CompletableFuture<>();
|
||||||
|
volatile boolean returned;
|
||||||
|
|
||||||
|
Pending(String msgId) {
|
||||||
|
this.msgId = msgId;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Connect to {@code uri} (e.g. {@code amqp://guest:guest@127.0.0.1:5672/}) with {@link #DEFAULT_PREFETCH}. */
|
||||||
public static AmqpReplyInbox open(String uri) {
|
public static AmqpReplyInbox open(String uri) {
|
||||||
|
return open(uri, DEFAULT_PREFETCH);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** As {@link #open(String)}, with an explicit consumer prefetch (CB-527: caps the held backlog per target). */
|
||||||
|
public static AmqpReplyInbox open(String uri, int prefetch) {
|
||||||
try {
|
try {
|
||||||
ConnectionFactory factory = new ConnectionFactory();
|
ConnectionFactory factory = new ConnectionFactory();
|
||||||
factory.setUri(uri);
|
factory.setUri(uri);
|
||||||
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
|
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
|
||||||
factory.setAutomaticRecoveryEnabled(true);
|
factory.setAutomaticRecoveryEnabled(true);
|
||||||
factory.setTopologyRecoveryEnabled(true);
|
factory.setTopologyRecoveryEnabled(true);
|
||||||
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"));
|
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"), prefetch);
|
||||||
} catch (Exception e) {
|
} catch (Exception e) {
|
||||||
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
|
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Wrap an already-open connection (injection seam for the contract test). */
|
/** Wrap an already-open connection with {@link #DEFAULT_PREFETCH} (injection seam for the contract test). */
|
||||||
AmqpReplyInbox(Connection connection) {
|
AmqpReplyInbox(Connection connection) {
|
||||||
|
this(connection, DEFAULT_PREFETCH);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** As above, with an explicit prefetch (injection seam for the contract test). */
|
||||||
|
AmqpReplyInbox(Connection connection, int prefetch) {
|
||||||
this.connection = connection;
|
this.connection = connection;
|
||||||
try {
|
try {
|
||||||
this.channel = connection.createChannel();
|
this.channel = connection.createChannel();
|
||||||
|
// CB-527: bound the held backlog per owned target — must be set before any own()/basicConsume.
|
||||||
|
this.channel.basicQos(prefetch);
|
||||||
|
this.publishChannel = connection.createChannel();
|
||||||
|
this.publishChannel.confirmSelect();
|
||||||
|
this.publishChannel.addReturnListener(this::onReturn);
|
||||||
|
this.publishChannel.addConfirmListener(this::onAck, this::onNack);
|
||||||
} catch (IOException e) {
|
} catch (IOException e) {
|
||||||
throw new IllegalStateException("cannot open AMQP channel", e);
|
throw new IllegalStateException("cannot open AMQP channel", e);
|
||||||
}
|
}
|
||||||
// On automatic recovery the broker redelivers unacked messages with FRESH delivery-tags; the
|
// On automatic recovery the broker redelivers unacked messages with FRESH delivery-tags; the
|
||||||
// tags we were holding are now stale. Drop the held snapshot so the re-attached consumer
|
// tags we were holding are now stale. Drop the held snapshot so the re-attached consumer
|
||||||
// repopulates it with valid tags (dedup by msgId still prevents any double-queue).
|
// repopulates it with valid tags (dedup by msgId still prevents any double-queue). Any publish
|
||||||
|
// confirm still in flight when the connection dropped is equally stale — its sequence number
|
||||||
|
// meant nothing on the old channel and means nothing on the recovered one, so fail it now
|
||||||
|
// rather than let it silently ride out CONFIRM_TIMEOUT_MS.
|
||||||
if (connection instanceof Recoverable recoverable) {
|
if (connection instanceof Recoverable recoverable) {
|
||||||
recoverable.addRecoveryListener(new RecoveryListener() {
|
recoverable.addRecoveryListener(new RecoveryListener() {
|
||||||
@Override
|
@Override
|
||||||
public void handleRecovery(Recoverable recoverable) {
|
public void handleRecovery(Recoverable recoverable) {
|
||||||
held.clear();
|
held.clear();
|
||||||
|
failPendingPublishesOnRecovery();
|
||||||
log.info("AMQP connection recovered; cleared held replies for fresh redelivery");
|
log.info("AMQP connection recovered; cleared held replies for fresh redelivery");
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -141,6 +223,12 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Publish {@code content} and block until the broker's publisher confirm for it lands (CB-528).
|
||||||
|
* Throws {@link IllegalStateException} if the message is returned as unroutable, nacked, or not
|
||||||
|
* confirmed within {@link #CONFIRM_TIMEOUT_MS} — the caller must treat that as a failed publish,
|
||||||
|
* not a lost-and-forgotten one.
|
||||||
|
*/
|
||||||
@Override
|
@Override
|
||||||
public void publish(String target, String msgId, String content) {
|
public void publish(String target, String msgId, String content) {
|
||||||
AMQP.BasicProperties props = new AMQP.BasicProperties.Builder()
|
AMQP.BasicProperties props = new AMQP.BasicProperties.Builder()
|
||||||
@@ -148,12 +236,35 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
|||||||
.deliveryMode(2) // persistent — survives a broker restart
|
.deliveryMode(2) // persistent — survives a broker restart
|
||||||
.contentType("text/plain")
|
.contentType("text/plain")
|
||||||
.build();
|
.build();
|
||||||
try {
|
Pending pending = new Pending(msgId);
|
||||||
synchronized (channelLock) {
|
long seq;
|
||||||
channel.basicPublish("", queueName(target), props, content.getBytes(StandardCharsets.UTF_8));
|
synchronized (publishChannelLock) {
|
||||||
|
seq = publishChannel.getNextPublishSeqNo();
|
||||||
|
pendingBySeq.put(seq, pending);
|
||||||
|
pendingByMsgId.put(msgId, pending);
|
||||||
|
try {
|
||||||
|
publishChannel.basicPublish("", queueName(target), true, props,
|
||||||
|
content.getBytes(StandardCharsets.UTF_8));
|
||||||
|
} catch (IOException e) {
|
||||||
|
pendingBySeq.remove(seq, pending);
|
||||||
|
pendingByMsgId.remove(msgId, pending);
|
||||||
|
throw new IllegalStateException("cannot publish reply to " + queueName(target), e);
|
||||||
}
|
}
|
||||||
} catch (IOException e) {
|
}
|
||||||
throw new IllegalStateException("cannot publish reply to " + queueName(target), e);
|
try {
|
||||||
|
pending.confirmed.get(CONFIRM_TIMEOUT_MS, TimeUnit.MILLISECONDS);
|
||||||
|
} catch (ExecutionException e) {
|
||||||
|
Throwable cause = e.getCause();
|
||||||
|
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
|
||||||
|
} catch (TimeoutException e) {
|
||||||
|
throw new IllegalStateException("publish confirm for reply " + msgId + " to " + queueName(target)
|
||||||
|
+ " timed out after " + CONFIRM_TIMEOUT_MS + "ms — broker may be unreachable or overloaded", e);
|
||||||
|
} catch (InterruptedException e) {
|
||||||
|
Thread.currentThread().interrupt();
|
||||||
|
throw new IllegalStateException("interrupted awaiting publish confirm for " + msgId, e);
|
||||||
|
} finally {
|
||||||
|
pendingBySeq.remove(seq, pending);
|
||||||
|
pendingByMsgId.remove(msgId, pending);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -222,17 +333,115 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
|||||||
};
|
};
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** Broker return for an unroutable {@code mandatory} publish — arrives BEFORE its confirm (CB-528). */
|
||||||
|
private void onReturn(Return r) {
|
||||||
|
String msgId = r.getProperties() == null ? null : r.getProperties().getMessageId();
|
||||||
|
Pending pending = msgId == null ? null : pendingByMsgId.get(msgId);
|
||||||
|
if (pending != null) {
|
||||||
|
pending.returned = true;
|
||||||
|
} else {
|
||||||
|
log.warn("AMQP return for reply {} (routingKey={}, {} {}) with no matching in-flight publish"
|
||||||
|
+ " — already resolved by a prior confirm", msgId, r.getRoutingKey(), r.getReplyCode(),
|
||||||
|
r.getReplyText());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private void onAck(long seq, boolean multiple) {
|
||||||
|
resolveConfirm(seq, multiple, true);
|
||||||
|
}
|
||||||
|
|
||||||
|
private void onNack(long seq, boolean multiple) {
|
||||||
|
resolveConfirm(seq, multiple, false);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Resolve every pending publish covered by this confirm (a single seq, or — {@code multiple} —
|
||||||
|
* every seq up to and including it). Checks {@link Pending#returned} at confirm time: since the
|
||||||
|
* broker's return for an unroutable message always precedes its confirm, an ack that arrives after
|
||||||
|
* a return means "confirmed but never routed", not "durably queued".
|
||||||
|
*/
|
||||||
|
private void resolveConfirm(long seq, boolean multiple, boolean ack) {
|
||||||
|
NavigableMap<Long, Pending> covered = multiple
|
||||||
|
? pendingBySeq.headMap(seq, true)
|
||||||
|
: pendingBySeq.subMap(seq, true, seq, true);
|
||||||
|
for (var it = covered.entrySet().iterator(); it.hasNext(); ) {
|
||||||
|
Pending pending = it.next().getValue();
|
||||||
|
it.remove();
|
||||||
|
pendingByMsgId.remove(pending.msgId, pending);
|
||||||
|
if (ack && !pending.returned) {
|
||||||
|
pending.confirmed.complete(null);
|
||||||
|
} else if (ack) {
|
||||||
|
pending.confirmed.completeExceptionally(new IllegalStateException(
|
||||||
|
"reply " + pending.msgId + " was returned as unroutable (queue not declared/owned)"));
|
||||||
|
} else {
|
||||||
|
pending.confirmed.completeExceptionally(new IllegalStateException(
|
||||||
|
"broker nacked publish of reply " + pending.msgId));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Fail every publish still awaiting its confirm — their sequence numbers are stale after recovery.
|
||||||
|
* Guarded by {@link #publishChannelLock}, the same lock {@link #publish} holds while it takes its
|
||||||
|
* sequence number and registers its {@link Pending}: without it, a {@link #publish} that starts
|
||||||
|
* after the connection has already recovered (so it publishes — and will be confirmed — on the
|
||||||
|
* <em>new</em> channel) can register between this sweep's iteration and its clear, and this sweep
|
||||||
|
* then fails a publish that actually succeeded. {@link #publish} only holds the lock for the
|
||||||
|
* seq/map-put/{@code basicPublish} — it awaits the confirm outside it — so this sweep can only ever
|
||||||
|
* wait for an in-flight {@code basicPublish} call to return, never for a broker round trip. No
|
||||||
|
* deadlock.
|
||||||
|
*
|
||||||
|
* <p>Package-private (rather than {@code private}) only so the unit test can drive it directly
|
||||||
|
* against a concurrent {@link #publish} without a live broker reconnect.
|
||||||
|
*/
|
||||||
|
void failPendingPublishesOnRecovery() {
|
||||||
|
synchronized (publishChannelLock) {
|
||||||
|
for (var it = pendingBySeq.entrySet().iterator(); it.hasNext(); ) {
|
||||||
|
Pending pending = it.next().getValue();
|
||||||
|
it.remove();
|
||||||
|
pendingByMsgId.remove(pending.msgId, pending);
|
||||||
|
pending.confirmed.completeExceptionally(new IllegalStateException(
|
||||||
|
"AMQP connection recovered mid-publish; confirm status of reply " + pending.msgId
|
||||||
|
+ " is unknown"));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Fail every publish still awaiting its confirm with a clear, immediate error instead of leaving it
|
||||||
|
* to time out after {@link #CONFIRM_TIMEOUT_MS} once the channels are closed underneath it. Guarded
|
||||||
|
* by {@link #publishChannelLock} for the same reason as {@link #failPendingPublishesOnRecovery}.
|
||||||
|
*/
|
||||||
|
private void failPendingPublishesOnClose() {
|
||||||
|
synchronized (publishChannelLock) {
|
||||||
|
for (var it = pendingBySeq.entrySet().iterator(); it.hasNext(); ) {
|
||||||
|
Pending pending = it.next().getValue();
|
||||||
|
it.remove();
|
||||||
|
pendingByMsgId.remove(pending.msgId, pending);
|
||||||
|
pending.confirmed.completeExceptionally(new IllegalStateException(
|
||||||
|
"AMQP reply inbox closed while publish of reply " + pending.msgId
|
||||||
|
+ " was still awaiting its confirm"));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
private static String queueName(String target) {
|
private static String queueName(String target) {
|
||||||
return QUEUE_PREFIX + target + QUEUE_SUFFIX;
|
return QUEUE_PREFIX + target + QUEUE_SUFFIX;
|
||||||
}
|
}
|
||||||
|
|
||||||
@Override
|
@Override
|
||||||
public void close() {
|
public void close() {
|
||||||
|
failPendingPublishesOnClose();
|
||||||
try {
|
try {
|
||||||
channel.close();
|
channel.close();
|
||||||
} catch (Exception e) {
|
} catch (Exception e) {
|
||||||
log.debug("AMQP channel close: {}", e.toString());
|
log.debug("AMQP channel close: {}", e.toString());
|
||||||
}
|
}
|
||||||
|
try {
|
||||||
|
publishChannel.close();
|
||||||
|
} catch (Exception e) {
|
||||||
|
log.debug("AMQP publish channel close: {}", e.toString());
|
||||||
|
}
|
||||||
try {
|
try {
|
||||||
connection.close();
|
connection.close();
|
||||||
} catch (Exception e) {
|
} catch (Exception e) {
|
||||||
|
|||||||
@@ -20,6 +20,7 @@ import java.util.concurrent.TimeUnit;
|
|||||||
import java.util.concurrent.TimeoutException;
|
import java.util.concurrent.TimeoutException;
|
||||||
import java.util.concurrent.atomic.AtomicLong;
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
import java.util.concurrent.locks.ReentrantLock;
|
import java.util.concurrent.locks.ReentrantLock;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* The blocking delegation feature (CB-104): deliver {@code content} into a worker and block until
|
* The blocking delegation feature (CB-104): deliver {@code content} into a worker and block until
|
||||||
@@ -52,8 +53,12 @@ public final class MessageService {
|
|||||||
*/
|
*/
|
||||||
private static final long ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L;
|
private static final long ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L;
|
||||||
|
|
||||||
/** How long a finished (terminal) ticket is retained for polling before it is pruned. */
|
/**
|
||||||
private static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
|
* How long a finished (terminal) ticket is retained for polling before it is pruned. Package-
|
||||||
|
* private (not {@code private}) so a test can advance an injected clock past it deterministically
|
||||||
|
* instead of duplicating the magic number or sleeping for real.
|
||||||
|
*/
|
||||||
|
static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
|
||||||
|
|
||||||
/** Outcome of a blocking send. */
|
/** Outcome of a blocking send. */
|
||||||
public enum Outcome {
|
public enum Outcome {
|
||||||
@@ -159,23 +164,39 @@ public final class MessageService {
|
|||||||
|
|
||||||
/** An in-flight or finished async delegation, keyed by its ticket. */
|
/** An in-flight or finished async delegation, keyed by its ticket. */
|
||||||
private static final class Task {
|
private static final class Task {
|
||||||
|
private final String ticket;
|
||||||
private final String target;
|
private final String target;
|
||||||
private final CompletableFuture<Reply> future = new CompletableFuture<>();
|
private final CompletableFuture<Reply> future = new CompletableFuture<>();
|
||||||
private final long createdNanos = System.nanoTime();
|
private final long createdNanos;
|
||||||
private volatile Reply question;
|
private volatile Reply question;
|
||||||
private volatile String turnId;
|
private volatile String turnId;
|
||||||
|
|
||||||
private Task(String target) {
|
private Task(String ticket, String target, long createdNanos) {
|
||||||
|
this.ticket = ticket;
|
||||||
this.target = target;
|
this.target = target;
|
||||||
|
this.createdNanos = createdNanos;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* A worker session's currently-open {@code bridge_ask} question, surfaced so {@code bridge_status}
|
||||||
|
* can show it without the caller needing the ticket first (CB-582). Only covers async
|
||||||
|
* (fire-and-poll) delegations, which track the question on their {@link Task}; a blocking
|
||||||
|
* ({@code wait:true}) send already hands the question straight back to its own caller, so there is
|
||||||
|
* nothing hidden left for {@code bridge_status} to surface in that case.
|
||||||
|
*/
|
||||||
|
public record PendingAsk(String ticket, String question, String turnId) {
|
||||||
|
}
|
||||||
|
|
||||||
private final AgentControl agents;
|
private final AgentControl agents;
|
||||||
private final Injector injector;
|
private final Injector injector;
|
||||||
private final Rendezvous rendezvous;
|
private final Rendezvous rendezvous;
|
||||||
private final ReplyInbox inbox;
|
private final ReplyInbox inbox;
|
||||||
private final ReplyPushLoop pushLoop;
|
private final ReplyPushLoop pushLoop;
|
||||||
private final Metrics metrics; // CB-502: nullable — no registry in unit tests
|
private final Metrics metrics; // CB-502: nullable — no registry in unit tests
|
||||||
|
// CB-588: injectable so pruneTerminalTickets' 10-minute TICKET_TTL_NANOS can be exercised in a
|
||||||
|
// test without a real wait — same seam SessionManager already uses for its idle reaper (nowNanos).
|
||||||
|
private final LongSupplier nowNanos;
|
||||||
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
|
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
|
||||||
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
|
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
|
||||||
/** Async task that owns each exact forward rendezvous waiter. */
|
/** Async task that owns each exact forward rendezvous waiter. */
|
||||||
@@ -191,7 +212,11 @@ public final class MessageService {
|
|||||||
* Create with an explicit {@link ReplyInbox} and optional {@link ReplyPushLoop}.
|
* Create with an explicit {@link ReplyInbox} and optional {@link ReplyPushLoop}.
|
||||||
*
|
*
|
||||||
* @param pushLoop nullable — when non-null, the push loop is notified on the no-waiter reply
|
* @param pushLoop nullable — when non-null, the push loop is notified on the no-waiter reply
|
||||||
* branch ({@link #reply}) so it can nudge the primary to drain the inbox
|
* branch ({@link #reply}) so it can nudge the primary to drain the inbox,
|
||||||
|
* (CB-588) whenever an async ticket started by {@link #sendAsync} reaches a
|
||||||
|
* terminal phase, whenever {@link #poll} hands a terminal ticket to its caller,
|
||||||
|
* and (CB-582) whenever an async ticket's worker pauses mid-turn in
|
||||||
|
* {@code bridge_ask} or that pause ends (answered or lapsed)
|
||||||
*/
|
*/
|
||||||
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
|
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
|
||||||
ReplyInbox inbox, ReplyPushLoop pushLoop) {
|
ReplyInbox inbox, ReplyPushLoop pushLoop) {
|
||||||
@@ -206,12 +231,19 @@ public final class MessageService {
|
|||||||
*/
|
*/
|
||||||
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
|
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
|
||||||
ReplyInbox inbox, ReplyPushLoop pushLoop, Metrics metrics) {
|
ReplyInbox inbox, ReplyPushLoop pushLoop, Metrics metrics) {
|
||||||
|
this(agents, injector, rendezvous, inbox, pushLoop, metrics, System::nanoTime);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Test constructor with an injectable clock (CB-588: exercise the ticket-prune TTL without a real wait). */
|
||||||
|
MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox,
|
||||||
|
ReplyPushLoop pushLoop, Metrics metrics, LongSupplier nowNanos) {
|
||||||
this.agents = agents;
|
this.agents = agents;
|
||||||
this.injector = injector;
|
this.injector = injector;
|
||||||
this.rendezvous = rendezvous;
|
this.rendezvous = rendezvous;
|
||||||
this.inbox = inbox;
|
this.inbox = inbox;
|
||||||
this.pushLoop = pushLoop;
|
this.pushLoop = pushLoop;
|
||||||
this.metrics = metrics;
|
this.metrics = metrics;
|
||||||
|
this.nowNanos = nowNanos;
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Create with an explicit {@link ReplyInbox} and no push loop. */
|
/** Create with an explicit {@link ReplyInbox} and no push loop. */
|
||||||
@@ -333,9 +365,30 @@ public final class MessageService {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Drain (peek + ack) all pending inbox replies for {@code target}. At-least-once: returns the
|
* Drain (peek + ack) all pending inbox replies for {@code target}.
|
||||||
* messages and acknowledges them; an in-flight failure between returning and the caller
|
*
|
||||||
* processing them re-surfaces them on a subsequent drain (the ack is local).
|
* <p><strong>The ack happens here, before the caller has the messages</strong> — before the MCP
|
||||||
|
* or REST response carrying them has been written, and long before the client has processed
|
||||||
|
* them. That ordering is what the two adapters disagree about, so do not read this method as
|
||||||
|
* "at-least-once" without qualifying which inbox is behind it (CB-529):
|
||||||
|
*
|
||||||
|
* <ul>
|
||||||
|
* <li>{@code InMemoryReplyInbox} — the ack only drops an entry from a local map. The messages
|
||||||
|
* are already in the returned list, so nothing can be lost after this point.
|
||||||
|
* <li>{@code AmqpReplyInbox} — the ack is a broker-side {@code basicAck}. Once it lands the
|
||||||
|
* broker has forgotten the message. If the daemon dies while writing the response, the
|
||||||
|
* reply is gone from the broker <em>and</em> the client never received it. Re-polling
|
||||||
|
* cannot recover it, because there is nothing left to re-deliver.
|
||||||
|
* </ul>
|
||||||
|
*
|
||||||
|
* <p>So the loss window is the response write, and it is a genuine loss rather than a
|
||||||
|
* redelivery. This is accepted, not overlooked: the alternative — ack on the next poll — turns
|
||||||
|
* every normal drain into a double delivery, which costs more than the window it closes. A
|
||||||
|
* caller that needs certainty re-polls; that is idempotent for every case except this one.
|
||||||
|
*
|
||||||
|
* <p>Any change here must be checked against <em>both</em> adapters. The previous version of
|
||||||
|
* this javadoc claimed "the ack is local", which was true when only the in-memory inbox existed
|
||||||
|
* and silently became false when the AMQP adapter landed.
|
||||||
*
|
*
|
||||||
* @return the drained messages, newest last (FIFO); empty list if none
|
* @return the drained messages, newest last (FIFO); empty list if none
|
||||||
*/
|
*/
|
||||||
@@ -456,6 +509,15 @@ public final class MessageService {
|
|||||||
rendezvous.closeAsk(ticket.turnId());
|
rendezvous.closeAsk(ticket.turnId());
|
||||||
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
|
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
|
||||||
}
|
}
|
||||||
|
// CB-582: the question just became visible via bridge_poll (Phase.ASKING) for an async
|
||||||
|
// (wait:false) delegation — nudge the lead's own pane the same way a terminal ticket does
|
||||||
|
// (CB-588), since the lead's normal poll cadence is minutes away and the reverse-rendezvous
|
||||||
|
// window (~55s, see BridgeMcp/BridgedApp) is far shorter. A blocking (wait:true) send has
|
||||||
|
// no Task and gets the question directly in its own reply, so task == null there — nothing
|
||||||
|
// to nudge.
|
||||||
|
if (task != null && pushLoop != null) {
|
||||||
|
pushLoop.onQuestionOpened(task.ticket, workerSession, ticket.turnId(), question);
|
||||||
|
}
|
||||||
}
|
}
|
||||||
try {
|
try {
|
||||||
String answer = ticket.answer().get(timeoutMillis, TimeUnit.MILLISECONDS);
|
String answer = ticket.answer().get(timeoutMillis, TimeUnit.MILLISECONDS);
|
||||||
@@ -474,6 +536,14 @@ public final class MessageService {
|
|||||||
// Only the fresh owner tears down the shared turn; a duplicate must leave it open.
|
// Only the fresh owner tears down the shared turn; a duplicate must leave it open.
|
||||||
if (ticket.fresh()) {
|
if (ticket.fresh()) {
|
||||||
rendezvous.closeAsk(ticket.turnId());
|
rendezvous.closeAsk(ticket.turnId());
|
||||||
|
// CB-582: tear the push loop's copy down at the same point, not only on the three
|
||||||
|
// paths that call clearAsyncQuestion. The answer future can complete exceptionally
|
||||||
|
// (ExecutionException) or the thread be interrupted, and both leave this method by
|
||||||
|
// throwing — the question would stay pending forever, keep being named in nudges
|
||||||
|
// until its own cap, and never be removed from the map. Already-closed is a no-op.
|
||||||
|
if (pushLoop != null) {
|
||||||
|
pushLoop.questionClosed(ticket.turnId());
|
||||||
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -550,8 +620,24 @@ public final class MessageService {
|
|||||||
*/
|
*/
|
||||||
public String sendAsync(String target, String content, Runnable onAccepted) {
|
public String sendAsync(String target, String content, Runnable onAccepted) {
|
||||||
String ticket = "task-" + ticketSeq.incrementAndGet();
|
String ticket = "task-" + ticketSeq.incrementAndGet();
|
||||||
Task task = new Task(target);
|
Task task = new Task(ticket, target, nowNanos.getAsLong());
|
||||||
tasks.put(ticket, task);
|
tasks.put(ticket, task);
|
||||||
|
if (pushLoop != null) {
|
||||||
|
// CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a
|
||||||
|
// worker paused in bridge_ask leaves it running, per finishAsyncTask's own contract — so
|
||||||
|
// this fires exactly once, from whichever path completes it: finishAsyncTask(task, result)
|
||||||
|
// below on any non-QUESTION outcome of send() — a worker's bridge_reply, the CB-106
|
||||||
|
// completion fallback, a CB-109 wedge, TIMED_OUT, BUSY, or BACKEND_EXHAUSTED — the same
|
||||||
|
// finishAsyncTask reached via answer()'s finishAsyncTask(turnId, result) once a QUESTION
|
||||||
|
// is resolved, completeExceptionally(t) just below when send() itself throws, or a CB-516
|
||||||
|
// abandon() on teardown. Without this, MessageService.reply's rendezvous fast path (the
|
||||||
|
// one an async ticket always takes) never told the push loop anything happened — see the
|
||||||
|
// class javadoc on sendAsync/CB-107.
|
||||||
|
task.future.whenComplete((reply, ex) -> {
|
||||||
|
boolean failed = ex != null || reply == null || !reply.completed();
|
||||||
|
pushLoop.onTicketTerminal(ticket, target, failed);
|
||||||
|
});
|
||||||
|
}
|
||||||
asyncExecutor.submit(() -> {
|
asyncExecutor.submit(() -> {
|
||||||
try {
|
try {
|
||||||
Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task);
|
Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task);
|
||||||
@@ -589,6 +675,11 @@ public final class MessageService {
|
|||||||
}
|
}
|
||||||
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target), null);
|
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target), null);
|
||||||
}
|
}
|
||||||
|
// CB-588: the ticket is terminal and being handed to the caller right here — tell the push
|
||||||
|
// loop it is collected so a later tick's nudge never names a ticket the lead already has.
|
||||||
|
if (pushLoop != null) {
|
||||||
|
pushLoop.ticketCollected(ticket);
|
||||||
|
}
|
||||||
Reply r;
|
Reply r;
|
||||||
try {
|
try {
|
||||||
r = f.getNow(null);
|
r = f.getNow(null);
|
||||||
@@ -619,10 +710,27 @@ public final class MessageService {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Drop finished tickets older than the TTL so the registry cannot grow without bound. */
|
/**
|
||||||
|
* Drop finished tickets older than the TTL so {@link #tasks} cannot grow without bound.
|
||||||
|
*
|
||||||
|
* <p>{@code tasks} is the sole authority on whether a ticket still exists — {@link #poll} returns
|
||||||
|
* {@code null} the instant a ticket is gone from here, before it ever reaches the terminal branch
|
||||||
|
* that calls {@link ReplyPushLoop#ticketCollected}. Without telling the push loop about a prune
|
||||||
|
* too, its own {@code pendingTickets} entry would outlive the ticket it names: an unpolled ticket
|
||||||
|
* (or one the reminder cap already gave up on) is pruned here but never collected there, so it
|
||||||
|
* lingers in {@code pendingTickets} forever and rides along on every later nudge to the same lead
|
||||||
|
* — naming a ticket {@code bridge_poll} can no longer find (CB-588 follow-up).
|
||||||
|
*/
|
||||||
private void pruneTerminalTickets() {
|
private void pruneTerminalTickets() {
|
||||||
long cutoff = System.nanoTime() - TICKET_TTL_NANOS;
|
long cutoff = nowNanos.getAsLong() - TICKET_TTL_NANOS;
|
||||||
tasks.values().removeIf(t -> t.future.isDone() && t.createdNanos < cutoff);
|
tasks.entrySet().removeIf(e -> {
|
||||||
|
Task t = e.getValue();
|
||||||
|
boolean expired = t.future.isDone() && t.createdNanos < cutoff;
|
||||||
|
if (expired && pushLoop != null) {
|
||||||
|
pushLoop.ticketCollected(e.getKey());
|
||||||
|
}
|
||||||
|
return expired;
|
||||||
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Record the active question for an async ticket; blocking sends have no entry and stay unchanged. */
|
/** Record the active question for an async ticket; blocking sends have no entry and stay unchanged. */
|
||||||
@@ -638,6 +746,12 @@ public final class MessageService {
|
|||||||
|
|
||||||
/** Clear an answered or lapsed question, but only when it matches the ticket's current turn. */
|
/** Clear an answered or lapsed question, but only when it matches the ticket's current turn. */
|
||||||
private void clearAsyncQuestion(String turnId, boolean forgetTurn) {
|
private void clearAsyncQuestion(String turnId, boolean forgetTurn) {
|
||||||
|
// CB-582: tell the push loop first — like ticketCollected, a removal for a turnId it never
|
||||||
|
// nudged about (or already dropped) is a harmless no-op, so this is safe to call unconditionally
|
||||||
|
// rather than threading the guard below through it.
|
||||||
|
if (pushLoop != null) {
|
||||||
|
pushLoop.questionClosed(turnId);
|
||||||
|
}
|
||||||
Task task = asyncTasksByTurn.get(turnId);
|
Task task = asyncTasksByTurn.get(turnId);
|
||||||
if (task != null && turnId.equals(task.turnId)) {
|
if (task != null && turnId.equals(task.turnId)) {
|
||||||
task.question = null;
|
task.question = null;
|
||||||
@@ -669,6 +783,23 @@ public final class MessageService {
|
|||||||
return asyncTasksByTurn.values().stream().anyMatch(task -> target.equals(task.target));
|
return asyncTasksByTurn.values().stream().anyMatch(task -> target.equals(task.target));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The question {@code workerSession} is currently paused on via {@code bridge_ask}, if any
|
||||||
|
* (CB-582) — {@code bridge_status} uses this to show a pending question without the caller
|
||||||
|
* needing the ticket. {@code null} when the session has no open async question (including a
|
||||||
|
* session mid a <em>blocking</em> {@code bridge_ask}, which has no {@link Task} to look up — see
|
||||||
|
* {@link PendingAsk}).
|
||||||
|
*/
|
||||||
|
public PendingAsk pendingAsk(String workerSession) {
|
||||||
|
for (Task task : tasks.values()) {
|
||||||
|
Reply q = task.question;
|
||||||
|
if (q != null && workerSession.equals(task.target)) {
|
||||||
|
return new PendingAsk(task.ticket, q.text(), q.turnId());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
|
||||||
/** Release the async executor. */
|
/** Release the async executor. */
|
||||||
public void close() {
|
public void close() {
|
||||||
asyncExecutor.shutdown();
|
asyncExecutor.shutdown();
|
||||||
|
|||||||
@@ -8,26 +8,66 @@ import dev.ltms.bridged.metrics.Metrics;
|
|||||||
import org.slf4j.Logger;
|
import org.slf4j.Logger;
|
||||||
import org.slf4j.LoggerFactory;
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
|
import java.util.ArrayList;
|
||||||
|
import java.util.HashSet;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Set;
|
||||||
import java.util.concurrent.ConcurrentHashMap;
|
import java.util.concurrent.ConcurrentHashMap;
|
||||||
import java.util.concurrent.ScheduledExecutorService;
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
import java.util.concurrent.TimeUnit;
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.stream.Collectors;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Mechanism (b) of CB-307: a dedicated, status-gated push loop that nudges the primary's own
|
* A status-gated push loop that nudges a lead's own herdr pane when it has uncollected work
|
||||||
* herdr pane when a worker reply lands with no live {@code bridge_send} to resolve it.
|
* waiting: a worker reply queued with no live {@code bridge_send} to resolve it (CB-307), an
|
||||||
|
* async delegation ticket ({@code bridge_send(wait:false)}) that reached a terminal phase
|
||||||
|
* (CB-588), or an async ticket's worker pausing mid-turn in {@code bridge_ask} to await an answer
|
||||||
|
* (CB-582).
|
||||||
*
|
*
|
||||||
* <p>The loop is triggered by {@link #onReplyQueued(String)} (called from
|
* <p><strong>CB-590: one schedule per lead.</strong> All three kinds of work are triggered
|
||||||
* {@link MessageService#reply} after the durable inbox publish). It checks four conditions
|
* through their own entry point — {@link #onReplyQueued(String)},
|
||||||
* at each tick via {@link #decide(String, int)}, then either injects a drain nudge,
|
* {@link #onTicketTerminal(String, String, boolean)}, and
|
||||||
* waits for the primary to become injectable, or stops reminding.
|
* {@link #onQuestionOpened(String, String, String, String)} — but each resolves the lead that
|
||||||
|
* should be nudged and coalesces onto a single per-lead reminder schedule, tracked in
|
||||||
|
* {@link #activeLeads}. Earlier this was two independent schedules (one keyed by worker target
|
||||||
|
* for replies, one keyed by lead for tickets) that could both decide to inject into the same pane
|
||||||
|
* in the same window — a race, not routine behaviour, but the expensive kind: it interrupts the
|
||||||
|
* lead's live turn twice. Collapsing to one schedule per lead makes that structurally impossible:
|
||||||
|
* at most one scheduled tick chain is ever live for a given lead (guarded by {@link #activeLeads}'
|
||||||
|
* compare-and-set), so at most one {@code agents.send} to that lead's pane is ever in flight.
|
||||||
*
|
*
|
||||||
* <p>Bounded: at most {@link #maxReminders} nudges per target, with a configurable backoff
|
* <p>Each tick examines <em>everything</em> pending for that lead — reply targets whose inbox
|
||||||
* between them. The reply is never lost — the durable inbox is the backstop.
|
* still holds an unacked message ({@link #pendingReplies}), tickets not yet collected
|
||||||
|
* ({@link #pendingTickets}), and open questions not yet answered or lapsed
|
||||||
|
* ({@link #pendingQuestions}) — and sends at most one combined nudge per tick
|
||||||
|
* ({@link #injectNudge(String, int, int, int)}). Work that arrives while the lead is busy is
|
||||||
|
* never lost: it is re-read fresh on every tick until the lead is injectable or its own reminder
|
||||||
|
* cap ({@link #maxReminders}) is reached — each source spends from its own budget, so one source
|
||||||
|
* exhausting its cap does not stop nudges about the others (post-CB-590 regression fix; see
|
||||||
|
* {@link #decide}) — whichever the durable inbox / pending set doesn't already answer via
|
||||||
|
* {@code STOP}.
|
||||||
*/
|
*/
|
||||||
public final class ReplyPushLoop {
|
public final class ReplyPushLoop {
|
||||||
|
|
||||||
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
|
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
|
||||||
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
|
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
|
||||||
|
/** Coalesced form, several uncollected replies for the same lead. */
|
||||||
|
static final String REPLIES_NUDGE_FORMAT =
|
||||||
|
"%d workers returned replies — run bridge_poll(target=...) for each to collect them: %s";
|
||||||
|
/** Singular form, one uncollected ticket. */
|
||||||
|
static final String TICKET_NUDGE_FORMAT =
|
||||||
|
"Ticket %s finished%s — run bridge_poll(ticket=%s) to collect it";
|
||||||
|
/** Coalesced form, several uncollected tickets for the same lead. */
|
||||||
|
static final String TICKETS_NUDGE_FORMAT =
|
||||||
|
"%d tickets finished%s — run bridge_poll(ticket=...) for each to collect them: %s";
|
||||||
|
/** Singular form, one worker paused mid-turn in bridge_ask (CB-582) — names the answer call directly. */
|
||||||
|
static final String QUESTION_NUDGE_FORMAT =
|
||||||
|
"Worker %s asked a question (ticket %s) — answer it with bridge_send(turnId=\"%s\", "
|
||||||
|
+ "content=...) to resume its turn:\n%s";
|
||||||
|
/** Coalesced form, several open questions for the same lead. */
|
||||||
|
static final String QUESTIONS_NUDGE_FORMAT =
|
||||||
|
"%d workers are paused on a question — run bridge_poll(ticket=...) for each, then answer "
|
||||||
|
+ "with bridge_send(turnId=..., content=...): %s";
|
||||||
|
|
||||||
private final PrimaryRegistry primaryRegistry;
|
private final PrimaryRegistry primaryRegistry;
|
||||||
private final AgentControl agents;
|
private final AgentControl agents;
|
||||||
@@ -37,8 +77,24 @@ public final class ReplyPushLoop {
|
|||||||
private final long backoffMs;
|
private final long backoffMs;
|
||||||
private final Metrics metrics; // CB-512: nullable — no registry in unit tests
|
private final Metrics metrics; // CB-512: nullable — no registry in unit tests
|
||||||
|
|
||||||
/** Track targets that have an active schedule. */
|
/**
|
||||||
private final ConcurrentHashMap<String, Boolean> activeTargets = new ConcurrentHashMap<>();
|
* Worker targets with a reply queued, keyed by target. Each entry carries its own nudge
|
||||||
|
* count (CB-598) rather than sharing one counter per lead per source: a target's count only
|
||||||
|
* ever reflects nudges that actually named that target, so a target that joins while the
|
||||||
|
* schedule is already deep into another target's reminders still reads as fresh.
|
||||||
|
*/
|
||||||
|
private final ConcurrentHashMap<String, ReplyEntry> pendingReplies = new ConcurrentHashMap<>();
|
||||||
|
/** Tickets that have gone terminal but not yet been polled, keyed by ticket. */
|
||||||
|
private final ConcurrentHashMap<String, PendingTicket> pendingTickets = new ConcurrentHashMap<>();
|
||||||
|
/**
|
||||||
|
* Open {@code bridge_ask} questions not yet answered or lapsed, keyed by {@code turnId}
|
||||||
|
* (CB-582). A question's own nudge count is tracked the same per-item way as
|
||||||
|
* {@link #pendingTickets} (CB-598): a fresh question keeps its source eligible regardless of
|
||||||
|
* how depleted an older, still-open question's count is.
|
||||||
|
*/
|
||||||
|
private final ConcurrentHashMap<String, PendingQuestion> pendingQuestions = new ConcurrentHashMap<>();
|
||||||
|
/** CB-590: leads with an active combined reminder schedule (replies and/or tickets and/or questions). */
|
||||||
|
private final ConcurrentHashMap<String, Boolean> activeLeads = new ConcurrentHashMap<>();
|
||||||
|
|
||||||
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
||||||
ScheduledExecutorService scheduler,
|
ScheduledExecutorService scheduler,
|
||||||
@@ -66,130 +122,511 @@ public final class ReplyPushLoop {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- decision logic (package-private for unit-testing) -------------------------------------
|
// --- pending-work lookups (package-private for unit-testing) -------------------------------
|
||||||
|
|
||||||
/** The action the loop should take for a target at the given reminder count. */
|
/** The action the loop should take for a lead at the given reminder count. */
|
||||||
enum Action { INJECT, WAIT_BUSY, STOP }
|
enum Action { INJECT, WAIT_BUSY, STOP }
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Pure decision function: examine the current state and return what the loop should do.
|
* Reply targets still pending for {@code lead} — registered via {@link #onReplyQueued} and
|
||||||
|
* whose inbox still holds an unacked message. A target whose inbox has since drained (acked,
|
||||||
|
* or collected via a live {@code bridge_send} rendezvous instead) is dropped from
|
||||||
|
* {@link #pendingReplies} here rather than lingering forever; there is no explicit "reply
|
||||||
|
* collected" callback the way {@link #ticketCollected} exists for tickets, so the inbox itself
|
||||||
|
* is the only signal.
|
||||||
|
*/
|
||||||
|
private Set<String> pendingReplyTargetsFor(String lead) {
|
||||||
|
Set<String> result = new HashSet<>();
|
||||||
|
for (var entry : pendingReplies.entrySet()) {
|
||||||
|
String target = entry.getKey();
|
||||||
|
ReplyEntry owning = entry.getValue();
|
||||||
|
if (!lead.equals(owning.lead())) continue;
|
||||||
|
if (inbox.peek(target).isEmpty()) {
|
||||||
|
pendingReplies.remove(target, owning);
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
result.add(target);
|
||||||
|
}
|
||||||
|
return result;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** A pending reply target: which lead to nudge, and how many nudges have named it so far. */
|
||||||
|
private record ReplyEntry(String lead, int nudgeCount) {
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* A ticket awaiting collection: which lead to nudge, whether it ended in failure, and how
|
||||||
|
* many nudges have named it so far (CB-598 — tracked per ticket, not per lead per source).
|
||||||
|
*/
|
||||||
|
private record PendingTicket(String ticket, String lead, boolean failed, int nudgeCount) {
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Tickets still pending for {@code lead}, snapshotted fresh for one tick. */
|
||||||
|
private List<PendingTicket> pendingTicketsFor(String lead) {
|
||||||
|
return pendingTickets.values().stream().filter(t -> lead.equals(t.lead())).toList();
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Ticket ids still pending for {@code lead} — a plain snapshot for race comparison. */
|
||||||
|
private Set<String> pendingTicketIdsFor(String lead) {
|
||||||
|
return pendingTicketsFor(lead).stream().map(PendingTicket::ticket)
|
||||||
|
.collect(Collectors.toUnmodifiableSet());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* An open question awaiting the lead's answer: which ticket it belongs to, which worker asked,
|
||||||
|
* which lead to nudge, the question text, and how many nudges have named it so far (CB-598 —
|
||||||
|
* tracked per question, not per lead per source).
|
||||||
|
*/
|
||||||
|
private record PendingQuestion(String turnId, String ticket, String target, String lead,
|
||||||
|
String question, int nudgeCount) {
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Questions still open for {@code lead}, snapshotted fresh for one tick. */
|
||||||
|
private List<PendingQuestion> pendingQuestionsFor(String lead) {
|
||||||
|
return pendingQuestions.values().stream().filter(q -> lead.equals(q.lead())).toList();
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Question turnIds still open for {@code lead} — a plain snapshot for race comparison. */
|
||||||
|
private Set<String> pendingQuestionTurnIdsFor(String lead) {
|
||||||
|
return pendingQuestionsFor(lead).stream().map(PendingQuestion::turnId)
|
||||||
|
.collect(Collectors.toUnmodifiableSet());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The reply-source reminder count {@link #decide} should see for {@code lead} on this tick:
|
||||||
|
* the <em>minimum</em> nudge count among the reply targets currently pending for it (CB-598).
|
||||||
*
|
*
|
||||||
* @param target the worker session (target terminal id)
|
* <p>Before this, the count passed to {@code decide} was a single counter carried forward
|
||||||
* @param reminderCount how many nudges have been sent so far for this target
|
* across scheduled ticks ({@code scheduleNext(lead, count + 1, ...)}), incremented whenever
|
||||||
|
* the source had <em>any</em> pending work — not tied to which target that work was. A target
|
||||||
|
* that joined while an older target's count was already near the cap inherited that count on
|
||||||
|
* its very next tick, even though no nudge had ever named it. Taking the minimum over what is
|
||||||
|
* actually pending now means a fresh target (count 0) keeps the source eligible regardless of
|
||||||
|
* how many times an older, still-undrained target has already been nudged; that older target
|
||||||
|
* keeps riding along in the combined nudge text without spending any more of its own budget
|
||||||
|
* (see {@link #bumpNudgeCounts}). Returns 0 when nothing is pending — {@link #decide} never
|
||||||
|
* consults the count in that case, since {@code hasReplyWork} is false.
|
||||||
|
*/
|
||||||
|
private int minReplyNudgeCountFor(String lead) {
|
||||||
|
int min = Integer.MAX_VALUE;
|
||||||
|
for (String target : pendingReplyTargetsFor(lead)) {
|
||||||
|
ReplyEntry entry = pendingReplies.get(target);
|
||||||
|
if (entry != null) {
|
||||||
|
min = Math.min(min, entry.nudgeCount());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return min == Integer.MAX_VALUE ? 0 : min;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** As {@link #minReplyNudgeCountFor}, for the ticket source. */
|
||||||
|
private int minTicketNudgeCountFor(String lead) {
|
||||||
|
int min = Integer.MAX_VALUE;
|
||||||
|
for (PendingTicket ticket : pendingTicketsFor(lead)) {
|
||||||
|
min = Math.min(min, ticket.nudgeCount());
|
||||||
|
}
|
||||||
|
return min == Integer.MAX_VALUE ? 0 : min;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** As {@link #minReplyNudgeCountFor}, for the question source (CB-582). */
|
||||||
|
private int minQuestionNudgeCountFor(String lead) {
|
||||||
|
int min = Integer.MAX_VALUE;
|
||||||
|
for (PendingQuestion q : pendingQuestionsFor(lead)) {
|
||||||
|
min = Math.min(min, q.nudgeCount());
|
||||||
|
}
|
||||||
|
return min == Integer.MAX_VALUE ? 0 : min;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Pure decision function: examine everything pending for {@code lead} — reply targets and
|
||||||
|
* tickets alike — and return what the loop should do.
|
||||||
|
*
|
||||||
|
* <p><strong>CB-590-fix: one schedule, two budgets.</strong> The single per-lead schedule
|
||||||
|
* (CB-590) still ticks once for both sources, but each source is capped independently —
|
||||||
|
* {@code replyReminderCount} against a reply target still pending, {@code ticketReminderCount}
|
||||||
|
* against a ticket still pending. A busy reply stream that exhausts its own cap must not stop
|
||||||
|
* the loop from nudging about a ticket that still has budget left, and vice versa: either
|
||||||
|
* source being eligible (has pending work AND is under its own cap) is enough for
|
||||||
|
* {@link Action#INJECT}. Only when neither source has eligible work does the loop
|
||||||
|
* {@link Action#STOP}.
|
||||||
|
*
|
||||||
|
* <p><strong>CB-598: the counts are per-item, not per-tick.</strong> {@link #tick} no longer
|
||||||
|
* carries these counts forward across scheduled calls — it recomputes them fresh every tick via
|
||||||
|
* {@link #minReplyNudgeCountFor} / {@link #minTicketNudgeCountFor}, so this function itself did
|
||||||
|
* not need to change; only what its caller feeds it did.
|
||||||
|
*
|
||||||
|
* @param lead the lead terminal to nudge
|
||||||
|
* @param replyReminderCount the lowest nudge count among reply targets pending for this lead
|
||||||
|
* @param ticketReminderCount the lowest nudge count among tickets pending for this lead
|
||||||
* @return the action the caller should take
|
* @return the action the caller should take
|
||||||
*/
|
*/
|
||||||
Action decide(String target, int reminderCount) {
|
Action decide(String lead, int replyReminderCount, int ticketReminderCount) {
|
||||||
// CB-532: the destination is per-delegation — the lead that sent this worker its work, not
|
return decide(lead, replyReminderCount, ticketReminderCount, minQuestionNudgeCountFor(lead));
|
||||||
// "the primary". With two leads orchestrating one fleet the singular question has no right
|
}
|
||||||
// answer, and answering it anyway interrupted whichever lead happened to call bridge_send
|
|
||||||
// first with results it never asked for.
|
/**
|
||||||
var nudgeTarget = primaryRegistry.nudgeTargetFor(target);
|
* As {@link #decide(String, int, int)}, with the question source (CB-582) folded in on the
|
||||||
if (nudgeTarget.isEmpty()) {
|
* same footing as replies and tickets: its own eligibility (has open questions AND under its
|
||||||
log.debug("push: no lead is known to be waiting on {}, stopping reminder", target);
|
* own {@link #maxReminders} budget) is enough on its own to {@link Action#INJECT}, exactly like
|
||||||
|
* the other two.
|
||||||
|
*
|
||||||
|
* @param questionReminderCount the lowest nudge count among questions open for this lead
|
||||||
|
*/
|
||||||
|
Action decide(String lead, int replyReminderCount, int ticketReminderCount, int questionReminderCount) {
|
||||||
|
boolean hasReplyWork = !pendingReplyTargetsFor(lead).isEmpty();
|
||||||
|
boolean hasTicketWork = !pendingTicketIdsFor(lead).isEmpty();
|
||||||
|
boolean hasQuestionWork = !pendingQuestionTurnIdsFor(lead).isEmpty();
|
||||||
|
if (!hasReplyWork && !hasTicketWork && !hasQuestionWork) {
|
||||||
|
log.debug("push: nothing pending for lead {}, stopping reminder", lead);
|
||||||
return Action.STOP;
|
return Action.STOP;
|
||||||
}
|
}
|
||||||
if (inbox.peek(target).isEmpty()) {
|
boolean replyEligible = hasReplyWork && replyReminderCount < maxReminders;
|
||||||
log.debug("push: inbox empty for {}, stopping reminder", target);
|
boolean ticketEligible = hasTicketWork && ticketReminderCount < maxReminders;
|
||||||
return Action.STOP;
|
boolean questionEligible = hasQuestionWork && questionReminderCount < maxReminders;
|
||||||
}
|
if (!replyEligible && !ticketEligible && !questionEligible) {
|
||||||
if (reminderCount >= maxReminders) {
|
log.debug("push: reminder cap ({}) reached for lead {} on every source with pending work, stopping",
|
||||||
log.debug("push: reminder cap ({}) reached for {}, stopping", maxReminders, target);
|
maxReminders, lead);
|
||||||
countNudge("exhausted");
|
countNudge("exhausted");
|
||||||
return Action.STOP;
|
return Action.STOP;
|
||||||
}
|
}
|
||||||
String leadTerminal = nudgeTarget.get();
|
|
||||||
AgentStatus status;
|
AgentStatus status;
|
||||||
try {
|
try {
|
||||||
status = agents.status(leadTerminal);
|
status = agents.status(lead);
|
||||||
} catch (RuntimeException e) {
|
} catch (RuntimeException e) {
|
||||||
log.debug("push: status check failed for lead {}, will retry", leadTerminal, e);
|
log.debug("push: status check failed for lead {}, will retry", lead, e);
|
||||||
return Action.WAIT_BUSY;
|
return Action.WAIT_BUSY;
|
||||||
}
|
}
|
||||||
if (status.injectable()) {
|
if (status.injectable()) {
|
||||||
return Action.INJECT;
|
return Action.INJECT;
|
||||||
}
|
}
|
||||||
log.debug("push: lead {} is {} (not injectable), waiting", leadTerminal, status);
|
log.debug("push: lead {} is {} (not injectable), waiting", lead, status);
|
||||||
return Action.WAIT_BUSY;
|
return Action.WAIT_BUSY;
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- public entrypoint ---------------------------------------------------------------------
|
// --- public entrypoints ----------------------------------------------------------------------
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Called when a reply is queued for {@code target}. Idempotent per target: a second call while
|
* Called when a reply is queued for {@code target}. Resolves the lead delegating to
|
||||||
* a schedule is active is a no-op. The schedule nudges the primary, then schedules a follow-up
|
* {@code target} (CB-532) and coalesces onto that lead's single reminder schedule — starting
|
||||||
* check (reminder on backoff, or re-check on WAIT_BUSY), until the inbox is empty or the cap
|
* one if none is active, joining an already-active one otherwise. A no-op if no lead is known
|
||||||
* is reached.
|
* to be waiting on {@code target}: there is nobody to nudge yet, and the durable inbox is the
|
||||||
|
* backstop until a lead is recorded.
|
||||||
*/
|
*/
|
||||||
public void onReplyQueued(String target) {
|
public void onReplyQueued(String target) {
|
||||||
if (activeTargets.putIfAbsent(target, Boolean.TRUE) != null) {
|
|
||||||
log.debug("push: already active for {}, ignoring duplicate trigger", target);
|
|
||||||
return; // already scheduled
|
|
||||||
}
|
|
||||||
log.debug("push: starting reminder loop for {}", target);
|
|
||||||
scheduleNext(target, 0);
|
|
||||||
}
|
|
||||||
|
|
||||||
/** Execute one loop tick — called on the scheduler thread. */
|
|
||||||
private void tick(String target, int reminderCount) {
|
|
||||||
var action = decide(target, reminderCount);
|
|
||||||
switch (action) {
|
|
||||||
case INJECT -> {
|
|
||||||
injectNudge(target, reminderCount);
|
|
||||||
scheduleNext(target, reminderCount + 1);
|
|
||||||
}
|
|
||||||
// Re-check after the configured backoff; the primary may become injectable soon.
|
|
||||||
case WAIT_BUSY -> scheduleNext(target, reminderCount);
|
|
||||||
case STOP -> {
|
|
||||||
activeTargets.remove(target);
|
|
||||||
log.debug("push: reminder loop ended for {}", target);
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
/** Send the nudge and log the event. */
|
|
||||||
private void injectNudge(String target, int reminderCount) {
|
|
||||||
// Re-read rather than threading it down from decide(): the delegating lead can change
|
|
||||||
// between the decision and the injection, and the nudge should follow the current one.
|
|
||||||
var lead = primaryRegistry.nudgeTargetFor(target);
|
var lead = primaryRegistry.nudgeTargetFor(target);
|
||||||
if (lead.isEmpty()) {
|
if (lead.isEmpty()) {
|
||||||
log.debug("push: lead for {} disappeared before the nudge could be sent", target);
|
log.debug("push: no lead is known to be waiting on {}, skipping reminder", target);
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
String leadTerminal = lead.get();
|
pendingReplies.compute(target, (t, existing) ->
|
||||||
String nudge = NUDGE_FORMAT.formatted(target, target);
|
new ReplyEntry(lead.get(), existing == null ? 0 : existing.nudgeCount()));
|
||||||
|
startOrCoalesce(lead.get());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Called when an async delegation ticket ({@code bridge_send(wait:false)}, CB-107) reaches a
|
||||||
|
* terminal phase — DONE or a failure. Unlike {@link #onReplyQueued}, which nudges about the
|
||||||
|
* durable-inbox no-waiter path, this covers the path {@code MessageService.reply} takes when a
|
||||||
|
* fire-and-poll send's own rendezvous waiter resolves the reply directly: that path returns
|
||||||
|
* before {@link #onReplyQueued} is ever called, so without this entry point a ticket finishing
|
||||||
|
* that way never nudged anyone (CB-588 / gitea #72).
|
||||||
|
*
|
||||||
|
* <p>Resolves the delegating lead the same way {@link #onReplyQueued} does and coalesces onto
|
||||||
|
* the same per-lead schedule (CB-590) — several tickets, or a ticket and a reply, finishing
|
||||||
|
* for the same lead while its schedule is already active all ride the existing schedule's next
|
||||||
|
* tick rather than firing a nudge each.
|
||||||
|
*
|
||||||
|
* @param ticket the ticket to nudge about
|
||||||
|
* @param target the worker session the ticket was sent to — resolves which lead delegated it
|
||||||
|
* @param failed whether the ticket ended in a failure phase rather than {@code DONE}
|
||||||
|
*/
|
||||||
|
public void onTicketTerminal(String ticket, String target, boolean failed) {
|
||||||
|
var lead = primaryRegistry.nudgeTargetFor(target);
|
||||||
|
if (lead.isEmpty()) {
|
||||||
|
log.debug("push: no lead is known to be waiting on ticket {} (target {}), skipping nudge",
|
||||||
|
ticket, target);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
pendingTickets.compute(ticket, (id, existing) ->
|
||||||
|
new PendingTicket(ticket, lead.get(), failed, existing == null ? 0 : existing.nudgeCount()));
|
||||||
|
startOrCoalesce(lead.get());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Called when a ticket's terminal state has been collected via {@code bridge_poll}. Removes it
|
||||||
|
* from the pending set so a scheduled tick — and any nudge it sends — never names a ticket the
|
||||||
|
* lead already has. A ticket that was never pending (unknown ticket, or one nudged with no push
|
||||||
|
* loop configured) is a no-op.
|
||||||
|
*/
|
||||||
|
public void ticketCollected(String ticket) {
|
||||||
|
pendingTickets.remove(ticket);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Called when an async ticket's worker pauses mid-turn in {@code bridge_ask} (CB-582): the
|
||||||
|
* question is now visible via {@code bridge_poll} (Phase.ASKING), but the reverse-rendezvous
|
||||||
|
* window it opened with (~55s default, see {@code BridgeMcp}/{@code BridgedApp}) is far shorter
|
||||||
|
* than a lead's normal minutes-long poll cadence — exactly the gap this closes. Resolves the
|
||||||
|
* delegating lead the same way {@link #onTicketTerminal} does and coalesces onto the same
|
||||||
|
* per-lead schedule (CB-590).
|
||||||
|
*
|
||||||
|
* @param ticket the async ticket the question belongs to (for {@code bridge_poll})
|
||||||
|
* @param target the worker session that asked
|
||||||
|
* @param turnId correlation id the lead answers with ({@code bridge_send turnId=...})
|
||||||
|
* @param question the question text
|
||||||
|
*/
|
||||||
|
public void onQuestionOpened(String ticket, String target, String turnId, String question) {
|
||||||
|
var lead = primaryRegistry.nudgeTargetFor(target);
|
||||||
|
if (lead.isEmpty()) {
|
||||||
|
log.debug("push: no lead is known to be waiting on {}'s question (turnId {}), skipping nudge",
|
||||||
|
target, turnId);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
pendingQuestions.put(turnId, new PendingQuestion(turnId, ticket, target, lead.get(), question, 0));
|
||||||
|
startOrCoalesce(lead.get());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Called when a worker's {@code bridge_ask} resolves — answered or lapsed unanswered — so a
|
||||||
|
* scheduled tick never nudges about a question the lead already handled. A {@code turnId} that
|
||||||
|
* was never pending (never nudged, or already closed) is a no-op.
|
||||||
|
*/
|
||||||
|
public void questionClosed(String turnId) {
|
||||||
|
pendingQuestions.remove(turnId);
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- the schedule ----------------------------------------------------------------------------
|
||||||
|
|
||||||
|
/** Start a reminder schedule for {@code lead}, or join the one already running. */
|
||||||
|
private void startOrCoalesce(String lead) {
|
||||||
|
if (activeLeads.putIfAbsent(lead, Boolean.TRUE) != null) {
|
||||||
|
log.debug("push: reminder loop already active for lead {}, work coalesced in", lead);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
log.debug("push: starting reminder loop for lead {}", lead);
|
||||||
|
scheduleNext(lead);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Execute one loop tick — called on the scheduler thread (or directly by a test; package-private
|
||||||
|
* for the same reason as {@link #stopOrRestart}).
|
||||||
|
*
|
||||||
|
* <p><strong>CB-598.</strong> The reminder counts fed into {@link #decide} are recomputed fresh
|
||||||
|
* every tick from what is actually pending right now ({@link #minReplyNudgeCountFor} /
|
||||||
|
* {@link #minTicketNudgeCountFor}), rather than carried forward as running counters across
|
||||||
|
* scheduled calls. A counter carried forward has no memory of which item it was counting for:
|
||||||
|
* a target or ticket that joined mid-backoff — after the previous tick fired but before this one
|
||||||
|
* did — is already sitting in {@code repliesBefore} / {@code ticketsBefore} below by the time this
|
||||||
|
* tick takes its snapshot, indistinguishable at that point from backlog the cap is meant to
|
||||||
|
* silence. Recomputing from the per-item counts fixes that: a newly-joined item's own count is
|
||||||
|
* still 0, so it keeps its source eligible regardless of how depleted an older, still-undrained
|
||||||
|
* item's count is.
|
||||||
|
*/
|
||||||
|
void tick(String lead) {
|
||||||
|
Set<String> repliesBefore = pendingReplyTargetsFor(lead);
|
||||||
|
Set<String> ticketsBefore = pendingTicketIdsFor(lead);
|
||||||
|
Set<String> questionsBefore = pendingQuestionTurnIdsFor(lead);
|
||||||
|
int replyReminderCount = minReplyNudgeCountFor(lead);
|
||||||
|
int ticketReminderCount = minTicketNudgeCountFor(lead);
|
||||||
|
int questionReminderCount = minQuestionNudgeCountFor(lead);
|
||||||
|
var action = decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount);
|
||||||
|
switch (action) {
|
||||||
|
case INJECT -> {
|
||||||
|
injectNudge(lead, replyReminderCount, ticketReminderCount, questionReminderCount);
|
||||||
|
scheduleNext(lead);
|
||||||
|
}
|
||||||
|
// Re-check after the configured backoff; the lead may become injectable soon.
|
||||||
|
case WAIT_BUSY -> scheduleNext(lead);
|
||||||
|
case STOP -> stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Release {@code lead}'s active-schedule slot, then restart it only if work landed that
|
||||||
|
* {@code repliesBefore} / {@code ticketsBefore} — the snapshots taken just before this tick's
|
||||||
|
* decision — did not already account for. {@link #onReplyQueued} / {@link #onTicketTerminal}
|
||||||
|
* read {@link #activeLeads} to decide whether to coalesce onto an existing schedule or start
|
||||||
|
* one, so work that lands between {@link #decide} returning {@link Action#STOP} and this
|
||||||
|
* removal running sees the (soon-to-be-stale) slot as occupied, coalesces onto a schedule that
|
||||||
|
* is about to die, and gets no nudge scheduled at all — a lost nudge, exactly what CB-588 (and
|
||||||
|
* now CB-590) exist to remove (originally found in review, gitea PR #73, for the ticket-only
|
||||||
|
* loop; carried forward here for the unified one).
|
||||||
|
*
|
||||||
|
* <p>Restarting on ANY non-empty pending set would be wrong: when STOP is reached because the
|
||||||
|
* reminder cap was hit rather than the backlog draining, the same never-collected work is
|
||||||
|
* expected to still be sitting there — that is the cap doing its job — and restarting would
|
||||||
|
* nudge about it forever, defeating the bound. Diffing the current pending sets against the
|
||||||
|
* "before" snapshots tells the two cases apart: an item present before this tick's decision is
|
||||||
|
* stale backlog, not a race; only an item absent from the "before" snapshot can only have
|
||||||
|
* arrived during the decision-to-release window, which is exactly the race this method closes.
|
||||||
|
*
|
||||||
|
* <p>Package-private so a test can drive the interleaving directly rather than trying to force a
|
||||||
|
* genuine thread race.
|
||||||
|
*
|
||||||
|
* <p>Terminates rather than spinning: this method restarts the schedule at most once per call,
|
||||||
|
* and a fresh {@link #onReplyQueued} / {@link #onTicketTerminal} racing the recheck below still
|
||||||
|
* terminates in one of two ways — either it observes the slot already vacated (by the
|
||||||
|
* {@code activeLeads.remove} above, which happens-before this recheck in program order) and
|
||||||
|
* claims it itself, or it lands first and this recheck then observes its work in
|
||||||
|
* {@link #pendingReplies} / {@link #pendingTickets} and reclaims the slot instead. Exactly one
|
||||||
|
* side always wins; neither can miss the other, so this never loops on its own account.
|
||||||
|
*/
|
||||||
|
void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore) {
|
||||||
|
stopOrRestart(lead, repliesBefore, ticketsBefore, pendingQuestionTurnIdsFor(lead));
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* As {@link #stopOrRestart(String, Set, Set)}, with the question source's (CB-582) own "before"
|
||||||
|
* snapshot folded into the same race check: a question that raced in during the
|
||||||
|
* decision-to-release window reclaims the schedule slot exactly like a raced-in reply or ticket.
|
||||||
|
*/
|
||||||
|
void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore,
|
||||||
|
Set<String> questionsBefore) {
|
||||||
|
activeLeads.remove(lead);
|
||||||
|
boolean racedIn = pendingReplyTargetsFor(lead).stream().anyMatch(t -> !repliesBefore.contains(t))
|
||||||
|
|| pendingTicketIdsFor(lead).stream().anyMatch(t -> !ticketsBefore.contains(t))
|
||||||
|
|| pendingQuestionTurnIdsFor(lead).stream().anyMatch(t -> !questionsBefore.contains(t));
|
||||||
|
if (racedIn && activeLeads.putIfAbsent(lead, Boolean.TRUE) == null) {
|
||||||
|
log.debug("push: new work for lead {} raced the reminder loop's stop — restarting", lead);
|
||||||
|
scheduleNext(lead);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
log.debug("push: reminder loop ended for lead {}", lead);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Send one combined nudge covering everything currently pending for {@code lead}. */
|
||||||
|
private void injectNudge(String lead, int replyReminderCount, int ticketReminderCount,
|
||||||
|
int questionReminderCount) {
|
||||||
|
// Re-read rather than threading it down from decide(): a reply can drain, a ticket be
|
||||||
|
// collected, or a question be answered (or another arrive), between the decision and the
|
||||||
|
// injection.
|
||||||
|
Set<String> replyTargets = pendingReplyTargetsFor(lead);
|
||||||
|
List<PendingTicket> tickets = pendingTicketsFor(lead);
|
||||||
|
List<PendingQuestion> questions = pendingQuestionsFor(lead);
|
||||||
|
if (replyTargets.isEmpty() && tickets.isEmpty() && questions.isEmpty()) {
|
||||||
|
log.debug("push: pending work for lead {} drained before the nudge could be sent", lead);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
String nudge = formatNudge(replyTargets, tickets, questions);
|
||||||
try {
|
try {
|
||||||
agents.send(leadTerminal, nudge);
|
agents.send(lead, nudge);
|
||||||
log.debug("push: nudge {}/{} sent to lead {} for target {}",
|
log.debug("push: nudge sent to lead {} (reply {}/{}, ticket {}/{}, question {}/{}; "
|
||||||
reminderCount + 1, maxReminders, leadTerminal, target);
|
+ "{} reply target(s), {} ticket(s), {} question(s))",
|
||||||
|
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
|
||||||
|
questionReminderCount + 1, maxReminders,
|
||||||
|
replyTargets.size(), tickets.size(), questions.size());
|
||||||
countNudge("delivered");
|
countNudge("delivered");
|
||||||
} catch (RuntimeException e) {
|
} catch (RuntimeException e) {
|
||||||
log.warn("push: failed to nudge lead {} for target {} (reminder {}/{}): {}",
|
log.warn("push: failed to nudge lead {} (reply {}/{}, ticket {}/{}, question {}/{}): {}",
|
||||||
leadTerminal, target, reminderCount + 1, maxReminders, e.toString());
|
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
|
||||||
|
questionReminderCount + 1, maxReminders, e.toString());
|
||||||
|
}
|
||||||
|
// Bump every item actually named in this nudge, not just whatever the shared source-level
|
||||||
|
// eligibility used to gate (CB-598) — each item's own count is what the next tick's
|
||||||
|
// minReplyNudgeCountFor / minTicketNudgeCountFor / minQuestionNudgeCountFor will read. An
|
||||||
|
// item already at or over the cap keeps riding along in the text (still pending, still
|
||||||
|
// named) but its extra bumps here are inert: decide() already treats it as ineligible once
|
||||||
|
// its count reaches maxReminders.
|
||||||
|
bumpNudgeCounts(replyTargets, tickets, questions);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Record that every one of these items was just named in a sent (or attempted) nudge. */
|
||||||
|
private void bumpNudgeCounts(Set<String> replyTargets, List<PendingTicket> tickets,
|
||||||
|
List<PendingQuestion> questions) {
|
||||||
|
for (String target : replyTargets) {
|
||||||
|
pendingReplies.computeIfPresent(target, (t, e) -> new ReplyEntry(e.lead(), e.nudgeCount() + 1));
|
||||||
|
}
|
||||||
|
for (PendingTicket ticket : tickets) {
|
||||||
|
pendingTickets.computeIfPresent(ticket.ticket(),
|
||||||
|
(id, e) -> new PendingTicket(e.ticket(), e.lead(), e.failed(), e.nudgeCount() + 1));
|
||||||
|
}
|
||||||
|
for (PendingQuestion question : questions) {
|
||||||
|
pendingQuestions.computeIfPresent(question.turnId(), (id, e) ->
|
||||||
|
new PendingQuestion(e.turnId(), e.ticket(), e.target(), e.lead(), e.question(),
|
||||||
|
e.nudgeCount() + 1));
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Schedule the next tick on the scheduler thread pool. */
|
/** Schedule the next tick on the scheduler thread pool. */
|
||||||
private void scheduleNext(String target, int nextReminderCount) {
|
private void scheduleNext(String lead) {
|
||||||
scheduler.schedule(() -> tick(target, nextReminderCount), backoffMs, TimeUnit.MILLISECONDS);
|
scheduler.schedule(() -> tick(lead),
|
||||||
|
backoffMs, TimeUnit.MILLISECONDS);
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- nudge formatting ------------------------------------------------------------------------
|
||||||
|
|
||||||
|
/** Render everything pending for one lead as a single nudge line. */
|
||||||
|
private static String formatNudge(Set<String> replyTargets, List<PendingTicket> tickets,
|
||||||
|
List<PendingQuestion> questions) {
|
||||||
|
List<String> parts = new ArrayList<>();
|
||||||
|
if (!replyTargets.isEmpty()) {
|
||||||
|
parts.add(formatRepliesNudge(replyTargets));
|
||||||
|
}
|
||||||
|
if (!tickets.isEmpty()) {
|
||||||
|
parts.add(formatTicketsNudge(tickets));
|
||||||
|
}
|
||||||
|
if (!questions.isEmpty()) {
|
||||||
|
parts.add(formatQuestionsNudge(questions));
|
||||||
|
}
|
||||||
|
return String.join(" | ", parts);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Render one or several pending reply targets. */
|
||||||
|
private static String formatRepliesNudge(Set<String> targets) {
|
||||||
|
if (targets.size() == 1) {
|
||||||
|
String target = targets.iterator().next();
|
||||||
|
return NUDGE_FORMAT.formatted(target, target);
|
||||||
|
}
|
||||||
|
String ids = String.join(", ", targets);
|
||||||
|
return REPLIES_NUDGE_FORMAT.formatted(targets.size(), ids);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Render one or several pending tickets. */
|
||||||
|
private static String formatTicketsNudge(List<PendingTicket> pending) {
|
||||||
|
if (pending.size() == 1) {
|
||||||
|
PendingTicket t = pending.get(0);
|
||||||
|
return TICKET_NUDGE_FORMAT.formatted(t.ticket(), t.failed() ? " (FAILED)" : "", t.ticket());
|
||||||
|
}
|
||||||
|
long failedCount = pending.stream().filter(PendingTicket::failed).count();
|
||||||
|
String ids = pending.stream()
|
||||||
|
.map(t -> t.failed() ? t.ticket() + " (FAILED)" : t.ticket())
|
||||||
|
.collect(Collectors.joining(", "));
|
||||||
|
String failedNote = failedCount > 0 ? " (%d failed)".formatted(failedCount) : "";
|
||||||
|
return TICKETS_NUDGE_FORMAT.formatted(pending.size(), failedNote, ids);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Render one or several open questions (CB-582). */
|
||||||
|
private static String formatQuestionsNudge(List<PendingQuestion> pending) {
|
||||||
|
if (pending.size() == 1) {
|
||||||
|
PendingQuestion q = pending.get(0);
|
||||||
|
return QUESTION_NUDGE_FORMAT.formatted(q.target(), q.ticket(), q.turnId(), q.question());
|
||||||
|
}
|
||||||
|
String ids = pending.stream()
|
||||||
|
.map(q -> q.ticket() + " (turnId=" + q.turnId() + ")")
|
||||||
|
.collect(Collectors.joining(", "));
|
||||||
|
return QUESTIONS_NUDGE_FORMAT.formatted(pending.size(), ids);
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- lifecycle -----------------------------------------------------------------------------
|
// --- lifecycle -----------------------------------------------------------------------------
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Whether any reminder loop is currently active for some target (CB-551). The idle-lead heartbeat
|
* Whether any reminder loop is currently active for some lead (CB-551). The idle-lead heartbeat
|
||||||
* uses this to stand aside: while the push loop is actively nudging the lead, a concurrent
|
* uses this to stand aside: while the push loop is actively nudging a lead, a concurrent
|
||||||
* heartbeat injection would start a second competing turn in the same pane — racing loops multiply
|
* heartbeat injection would start a second competing turn in the same pane — racing loops
|
||||||
* turns and context burn. "Active" means a schedule exists in {@link #activeTargets}; the set is
|
* multiply turns and context burn. "Active" means a schedule exists in {@link #activeLeads},
|
||||||
* bounded by what has been triggered, not by any persistent state.
|
* which now covers reply-queued (CB-307), ticket-terminal (CB-588), and question-open (CB-582)
|
||||||
|
* work (CB-590) — bounded by what has been triggered, not by any persistent state.
|
||||||
*/
|
*/
|
||||||
public boolean isActive() {
|
public boolean isActive() {
|
||||||
return !activeTargets.isEmpty();
|
return !activeLeads.isEmpty();
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Shut down the scheduler. Outstanding reminders are cancelled. */
|
/** Shut down the scheduler. Outstanding reminders are cancelled. */
|
||||||
public void stop() {
|
public void stop() {
|
||||||
scheduler.shutdownNow();
|
scheduler.shutdownNow();
|
||||||
activeTargets.clear();
|
activeLeads.clear();
|
||||||
|
pendingReplies.clear();
|
||||||
|
pendingTickets.clear();
|
||||||
|
pendingQuestions.clear();
|
||||||
}
|
}
|
||||||
|
|
||||||
/** @see #stop() */
|
/** @see #stop() */
|
||||||
|
|||||||
@@ -62,11 +62,15 @@ public interface PeerHandle {
|
|||||||
* non-null for a spawn that requested session identity, because it knows the id before the
|
* non-null for a spawn that requested session identity, because it knows the id before the
|
||||||
* peer has written anything.
|
* peer has written anything.
|
||||||
*
|
*
|
||||||
|
* <p>Deliberately not a {@code default} (CB-584, the same fix CB-571 made for
|
||||||
|
* {@link #charterReceipt()} one method below): a decorator that forgets to override this
|
||||||
|
* silently answers {@code null} for a question it has no basis to answer, and the gap surfaces
|
||||||
|
* only as a resume that quietly starts a cold session, not a compile error. Every
|
||||||
|
* implementation must answer explicitly.
|
||||||
|
*
|
||||||
* @return the peer's own session id, or {@code null} when not determinable
|
* @return the peer's own session id, or {@code null} when not determinable
|
||||||
*/
|
*/
|
||||||
default String agentSessionId() {
|
String agentSessionId();
|
||||||
return null;
|
|
||||||
}
|
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* The charter receipt (CB-571) for this peer's launch — the fingerprint of the exact charter
|
* The charter receipt (CB-571) for this peer's launch — the fingerprint of the exact charter
|
||||||
|
|||||||
@@ -25,6 +25,19 @@ public interface PeerLauncher {
|
|||||||
*/
|
*/
|
||||||
Set<Capability> capabilities();
|
Set<Capability> capabilities();
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The capabilities of the adapter that {@code profileName} resolves to (null/blank → the
|
||||||
|
* default profile, the same resolution {@link #spawn} uses). Distinct from {@link
|
||||||
|
* #capabilities()}, which unions every configured adapter: a caller that must know whether
|
||||||
|
* <em>this</em> profile's backend supports a capability — e.g. {@link Capability#SESSION_RESUME}
|
||||||
|
* before honoring {@link SpawnRequest#resumeSessionId()} — needs the per-profile answer, not
|
||||||
|
* the fleet-wide union, or a mixed fleet could OK a resume that lands on a non-supporting
|
||||||
|
* adapter (CB-584).
|
||||||
|
*
|
||||||
|
* @throws IllegalArgumentException if the profile is unknown and no default is configured
|
||||||
|
*/
|
||||||
|
Set<Capability> capabilitiesFor(String profileName);
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* {@code profileName}/requestedCwd null/blank → default resolution. Returns after the peer
|
* {@code profileName}/requestedCwd null/blank → default resolution. Returns after the peer
|
||||||
* process is live (env + argv + placement complete). Never returns {@code null}.
|
* process is live (env + argv + placement complete). Never returns {@code null}.
|
||||||
|
|||||||
@@ -0,0 +1,115 @@
|
|||||||
|
package dev.ltms.bridged.placement;
|
||||||
|
|
||||||
|
import java.util.LinkedHashMap;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.Objects;
|
||||||
|
import java.util.OptionalLong;
|
||||||
|
import java.util.concurrent.ConcurrentHashMap;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Where a credential (not a profile — see {@code BridgedConfig.Profile#effectiveCredentialId()})
|
||||||
|
* sits out a cooldown after a {@code BACKEND_EXHAUSTED} classification (CB-578 stage B), so a fresh
|
||||||
|
* spawn does not walk straight back onto the account that just refused on a usage limit.
|
||||||
|
*
|
||||||
|
* <p>Keyed by credential id, never by profile name: two profiles sharing one credential (e.g. two
|
||||||
|
* models on the same OpenAI account) share one quarantine — {@link #quarantine} one credential id
|
||||||
|
* and every profile whose {@code effectiveCredentialId()} equals it is quarantined too, without this
|
||||||
|
* class knowing anything about profiles at all. That mapping is the caller's job (see
|
||||||
|
* {@code CompositePeerLauncher} and {@code dev.ltms.bridged.inject.ExhaustionSink}).
|
||||||
|
*
|
||||||
|
* <p>The clock is injected ({@link LongSupplier}, conventionally {@code System::nanoTime} like
|
||||||
|
* {@code FleetHealthMonitor}), never read inline, so a quarantine's expiry is testable without a
|
||||||
|
* real sleep.
|
||||||
|
*/
|
||||||
|
public final class BackendQuarantine {
|
||||||
|
|
||||||
|
private final ConcurrentHashMap<String, Long> quarantinedUntilNanos = new ConcurrentHashMap<>();
|
||||||
|
private final LongSupplier nowNanos;
|
||||||
|
private final long cooldownNanos;
|
||||||
|
/** True only for {@link #none()}. See {@link #quarantine} for why this exists. */
|
||||||
|
private final boolean inert;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* @param nowNanos monotonic clock, injected for testability
|
||||||
|
* @param cooldownNanos how long a fresh {@link #quarantine} call blocks the credential for;
|
||||||
|
* must be positive
|
||||||
|
*/
|
||||||
|
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos) {
|
||||||
|
this(nowNanos, cooldownNanos, false);
|
||||||
|
}
|
||||||
|
|
||||||
|
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, boolean inert) {
|
||||||
|
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
|
||||||
|
if (cooldownNanos <= 0) {
|
||||||
|
throw new IllegalArgumentException("cooldownNanos must be positive: " + cooldownNanos);
|
||||||
|
}
|
||||||
|
this.cooldownNanos = cooldownNanos;
|
||||||
|
this.inert = inert;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Inert quarantine — {@link #quarantine} does nothing on this instance, so nothing is ever
|
||||||
|
* quarantined. The explicit stand-in a caller (or a test not exercising this feature) passes
|
||||||
|
* instead of a defaulting overload, exactly like {@code ExhaustedPatternLookup.none()}.
|
||||||
|
*/
|
||||||
|
public static BackendQuarantine none() {
|
||||||
|
return new BackendQuarantine(() -> 0L, 1, true);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Quarantine {@code credentialId} for the configured cooldown, starting now. A repeat call while
|
||||||
|
* already quarantined restarts the cooldown at full length — a fresh refusal is fresh evidence the
|
||||||
|
* account is still exhausted, not a reason to let an earlier, shorter wait stand.
|
||||||
|
*
|
||||||
|
* <p>On {@link #none()} this is a no-op. It has to be: that instance holds a clock frozen at 0,
|
||||||
|
* so recording a deadline would produce a quarantine that never expires — a credential locked out
|
||||||
|
* for the life of the daemon. Two production {@code CompositePeerLauncher} constructors default to
|
||||||
|
* {@code none()}, so the failure would be silent and permanent. An inert stand-in must omit the
|
||||||
|
* fact, never invent one.
|
||||||
|
*/
|
||||||
|
public void quarantine(String credentialId) {
|
||||||
|
Objects.requireNonNull(credentialId, "credentialId");
|
||||||
|
if (inert) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
quarantinedUntilNanos.put(credentialId, nowNanos.getAsLong() + cooldownNanos);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Whether {@code credentialId} is quarantined right now. */
|
||||||
|
public boolean isQuarantined(String credentialId) {
|
||||||
|
return remainingNanos(credentialId) > 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Seconds left on {@code credentialId}'s quarantine, or empty when it is not quarantined. */
|
||||||
|
public OptionalLong remainingSeconds(String credentialId) {
|
||||||
|
long remaining = remainingNanos(credentialId);
|
||||||
|
return remaining > 0 ? OptionalLong.of(toSecondsRoundedUp(remaining)) : OptionalLong.empty();
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Every currently-quarantined credential id and its remaining seconds (CB-578 stage B fleet
|
||||||
|
* reporting) — expired entries are never included. Not pruned from the backing map here: it stays
|
||||||
|
* small (bounded by the number of distinct credentials ever exhausted) and a lazily-stale entry is
|
||||||
|
* harmless, since every read already checks the deadline.
|
||||||
|
*/
|
||||||
|
public Map<String, Long> activeRemainingSeconds() {
|
||||||
|
Map<String, Long> out = new LinkedHashMap<>();
|
||||||
|
quarantinedUntilNanos.forEach((credentialId, deadline) -> {
|
||||||
|
long remaining = deadline - nowNanos.getAsLong();
|
||||||
|
if (remaining > 0) {
|
||||||
|
out.put(credentialId, toSecondsRoundedUp(remaining));
|
||||||
|
}
|
||||||
|
});
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
private long remainingNanos(String credentialId) {
|
||||||
|
Long deadline = quarantinedUntilNanos.get(credentialId);
|
||||||
|
return deadline == null ? 0L : deadline - nowNanos.getAsLong();
|
||||||
|
}
|
||||||
|
|
||||||
|
private static long toSecondsRoundedUp(long nanos) {
|
||||||
|
return (nanos + 999_999_999L) / 1_000_000_000L;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -4,19 +4,63 @@ package dev.ltms.bridged.placement;
|
|||||||
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
|
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
|
||||||
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps and
|
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps and
|
||||||
* reachability so that a pre-existing config behaves identically after upgrade.
|
* reachability so that a pre-existing config behaves identically after upgrade.
|
||||||
|
*
|
||||||
|
* <p>Two exceptions walk past the default instead of returning it unconditionally:
|
||||||
|
* <ul>
|
||||||
|
* <li>Quarantine (CB-578 stage B): a quarantined default is a credential that just refused on
|
||||||
|
* a usage limit, not a transient capacity or reachability concern.
|
||||||
|
* <li>Weight 0 (CB-554): {@code fixed} is still automatic selection, so a profile the operator
|
||||||
|
* marked "never auto-select me" ({@code weight <= 0}) must be skipped here exactly as
|
||||||
|
* {@code weighted}/{@code round-robin} skip it — an explicit {@code bridge_spawn} naming
|
||||||
|
* the profile is unaffected, only this automatic fallback walk.
|
||||||
|
* </ul>
|
||||||
|
* A fleet where nothing is ever quarantined or weight-0 never exercises either path, so today's
|
||||||
|
* behaviour is unchanged.
|
||||||
*/
|
*/
|
||||||
final class FixedPlacementPolicy implements PlacementPolicy {
|
final class FixedPlacementPolicy implements PlacementPolicy {
|
||||||
|
|
||||||
@Override
|
@Override
|
||||||
public PlacementCandidate select(PlacementContext ctx) {
|
public PlacementCandidate select(PlacementContext ctx) {
|
||||||
String d = ctx.defaultProfile();
|
String d = ctx.defaultProfile();
|
||||||
if (d != null && !d.isBlank()) {
|
if (d != null && !d.isBlank() && !ctx.quarantined().contains(d) && !weightExcluded(ctx, d)) {
|
||||||
return new PlacementCandidate(d, null, 1.0f, null);
|
return new PlacementCandidate(d, null, 1.0f, null);
|
||||||
}
|
}
|
||||||
|
for (PlacementCandidate c : ctx.candidates()) {
|
||||||
|
if (!ctx.quarantined().contains(c.profile()) && !c.excluded()) {
|
||||||
|
return new PlacementCandidate(c.profile(), null, c.weight(), c.maxLoad());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (d != null && !d.isBlank()) {
|
||||||
|
boolean dQuarantined = ctx.quarantined().contains(d);
|
||||||
|
boolean dWeightExcluded = weightExcluded(ctx, d);
|
||||||
|
if (dQuarantined && dWeightExcluded) {
|
||||||
|
throw new PlacementException("worker profile '" + d + "' is quarantined (backend "
|
||||||
|
+ "exhausted) and has weight 0 (excluded from automatic selection), and no "
|
||||||
|
+ "available candidate remains");
|
||||||
|
}
|
||||||
|
if (dWeightExcluded) {
|
||||||
|
throw new PlacementException("worker profile '" + d + "' has weight 0 (excluded "
|
||||||
|
+ "from automatic selection) and no available candidate remains");
|
||||||
|
}
|
||||||
|
if (dQuarantined) {
|
||||||
|
throw new PlacementException("worker profile '" + d + "' is quarantined (backend "
|
||||||
|
+ "exhausted) and no un-quarantined candidate is available");
|
||||||
|
}
|
||||||
|
}
|
||||||
if (!ctx.candidates().isEmpty()) {
|
if (!ctx.candidates().isEmpty()) {
|
||||||
PlacementCandidate first = ctx.candidates().getFirst();
|
throw new PlacementException(
|
||||||
return new PlacementCandidate(first.profile(), null, first.weight(), first.maxLoad());
|
"all worker profiles are excluded from automatic selection (quarantined or weight-0)");
|
||||||
}
|
}
|
||||||
throw new PlacementException("no worker profiles configured");
|
throw new PlacementException("no worker profiles configured");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** Whether {@code profile} carries {@code weight <= 0} (CB-554) among {@code ctx}'s candidates. */
|
||||||
|
private static boolean weightExcluded(PlacementContext ctx, String profile) {
|
||||||
|
for (PlacementCandidate c : ctx.candidates()) {
|
||||||
|
if (c.profile().equals(profile)) {
|
||||||
|
return c.excluded();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return false;
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -17,4 +17,18 @@ public record PlacementCandidate(String profile, String host, float weight, Inte
|
|||||||
public static PlacementCandidate profile(String profile) {
|
public static PlacementCandidate profile(String profile) {
|
||||||
return new PlacementCandidate(profile, null, 1.0f, null);
|
return new PlacementCandidate(profile, null, 1.0f, null);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* True when this candidate carries an explicit {@code weight <= 0} (CB-554) and must be
|
||||||
|
* skipped by every automatic policy — the same way a quarantined or unreachable candidate is
|
||||||
|
* skipped. {@code BridgedConfig.Profile}'s compact constructor already normalises "absent" to
|
||||||
|
* {@code 1.0} and "negative" to {@code 0.0}, so this is a plain threshold check here; it does
|
||||||
|
* not need to distinguish "explicit 0" from "absent" itself.
|
||||||
|
*
|
||||||
|
* <p>Exclusion is about <em>automatic</em> selection only — an explicit
|
||||||
|
* {@code bridge_spawn{profile:"..."}} bypasses placement entirely and is unaffected.
|
||||||
|
*/
|
||||||
|
public boolean excluded() {
|
||||||
|
return weight <= 0.0f;
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -11,9 +11,13 @@ import java.util.function.Function;
|
|||||||
* @param candidates every configured candidate; the policy filters out those at cap or unreachable
|
* @param candidates every configured candidate; the policy filters out those at cap or unreachable
|
||||||
* @param liveCount current live worker count per profile (from the session registry)
|
* @param liveCount current live worker count per profile (from the session registry)
|
||||||
* @param unreachable profiles already known to have failed in this spawn attempt
|
* @param unreachable profiles already known to have failed in this spawn attempt
|
||||||
|
* @param quarantined profiles whose credential is currently quarantined (CB-578 stage B) — a
|
||||||
|
* {@code BACKEND_EXHAUSTED} classification put it, or a profile it shares a
|
||||||
|
* credential with, on cooldown. Filtered the same way as {@code unreachable}.
|
||||||
*/
|
*/
|
||||||
public record PlacementContext(String defaultProfile,
|
public record PlacementContext(String defaultProfile,
|
||||||
List<PlacementCandidate> candidates,
|
List<PlacementCandidate> candidates,
|
||||||
Function<String, Integer> liveCount,
|
Function<String, Integer> liveCount,
|
||||||
Set<String> unreachable) {
|
Set<String> unreachable,
|
||||||
|
Set<String> quarantined) {
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -12,13 +12,16 @@ final class PlacementPolicyUtil {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Candidates that are not known-unreachable and have not reached their maxLoad.
|
* Candidates that are not weight-excluded (CB-554: explicit {@code weight <= 0}, checked
|
||||||
|
* first because it is a static config choice rather than transient state), not
|
||||||
|
* known-unreachable, not quarantined (CB-578 stage B), and have not reached their maxLoad.
|
||||||
* A {@code null} maxLoad means unlimited.
|
* A {@code null} maxLoad means unlimited.
|
||||||
*/
|
*/
|
||||||
static List<PlacementCandidate> available(PlacementContext ctx) {
|
static List<PlacementCandidate> available(PlacementContext ctx) {
|
||||||
List<PlacementCandidate> out = new ArrayList<>();
|
List<PlacementCandidate> out = new ArrayList<>();
|
||||||
for (PlacementCandidate c : ctx.candidates()) {
|
for (PlacementCandidate c : ctx.candidates()) {
|
||||||
if (ctx.unreachable().contains(c.profile())) {
|
if (c.excluded() || ctx.unreachable().contains(c.profile())
|
||||||
|
|| ctx.quarantined().contains(c.profile())) {
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
Integer cap = c.maxLoad();
|
Integer cap = c.maxLoad();
|
||||||
@@ -34,15 +37,23 @@ final class PlacementPolicyUtil {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Build a clear exception describing why every candidate was dropped: all at capacity,
|
* Build a clear exception describing why every candidate was dropped: all weight-0, all
|
||||||
* all unreachable, or a mix.
|
* quarantined, all at capacity, all unreachable, or a mix. Each candidate is counted into
|
||||||
|
* exactly one bucket (weight-excluded takes priority) so a candidate excluded for more than
|
||||||
|
* one reason is never double-counted.
|
||||||
*/
|
*/
|
||||||
static PlacementException emptyException(PlacementContext ctx) {
|
static PlacementException emptyException(PlacementContext ctx) {
|
||||||
|
int weightExcluded = 0;
|
||||||
int atCap = 0;
|
int atCap = 0;
|
||||||
int unreachable = 0;
|
int unreachable = 0;
|
||||||
|
int quarantined = 0;
|
||||||
for (PlacementCandidate c : ctx.candidates()) {
|
for (PlacementCandidate c : ctx.candidates()) {
|
||||||
Integer cap = c.maxLoad();
|
Integer cap = c.maxLoad();
|
||||||
if (ctx.unreachable().contains(c.profile())) {
|
if (c.excluded()) {
|
||||||
|
weightExcluded++;
|
||||||
|
} else if (ctx.quarantined().contains(c.profile())) {
|
||||||
|
quarantined++;
|
||||||
|
} else if (ctx.unreachable().contains(c.profile())) {
|
||||||
unreachable++;
|
unreachable++;
|
||||||
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
|
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
|
||||||
atCap++;
|
atCap++;
|
||||||
@@ -53,6 +64,13 @@ final class PlacementPolicyUtil {
|
|||||||
if (total == 0) {
|
if (total == 0) {
|
||||||
return new PlacementException("no worker profiles configured");
|
return new PlacementException("no worker profiles configured");
|
||||||
}
|
}
|
||||||
|
if (weightExcluded == total) {
|
||||||
|
return new PlacementException(
|
||||||
|
"all worker profiles have weight 0 (excluded from automatic selection)");
|
||||||
|
}
|
||||||
|
if (quarantined == total) {
|
||||||
|
return new PlacementException("all worker profiles are quarantined (backend exhausted)");
|
||||||
|
}
|
||||||
if (atCap == total) {
|
if (atCap == total) {
|
||||||
return new PlacementException("all worker profiles are at maxLoad");
|
return new PlacementException("all worker profiles are at maxLoad");
|
||||||
}
|
}
|
||||||
@@ -60,6 +78,8 @@ final class PlacementPolicyUtil {
|
|||||||
return new PlacementException("all worker profiles are unreachable");
|
return new PlacementException("all worker profiles are unreachable");
|
||||||
}
|
}
|
||||||
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
|
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
|
||||||
+ unreachable + " unreachable, " + (total - atCap - unreachable) + " remaining");
|
+ unreachable + " unreachable, " + quarantined + " quarantined, "
|
||||||
|
+ weightExcluded + " weight-0, "
|
||||||
|
+ (total - atCap - unreachable - quarantined - weightExcluded) + " remaining");
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -13,6 +13,7 @@ import dev.ltms.bridged.herdr.HerdrClient;
|
|||||||
import dev.ltms.bridged.herdr.HerdrException;
|
import dev.ltms.bridged.herdr.HerdrException;
|
||||||
import dev.ltms.bridged.inject.MemberPresence;
|
import dev.ltms.bridged.inject.MemberPresence;
|
||||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||||
|
import dev.ltms.bridged.placement.PlacementException;
|
||||||
import dev.ltms.bridged.msg.MessageService;
|
import dev.ltms.bridged.msg.MessageService;
|
||||||
import dev.ltms.bridged.session.SessionManager;
|
import dev.ltms.bridged.session.SessionManager;
|
||||||
import dev.ltms.bridged.peer.MemberRole;
|
import dev.ltms.bridged.peer.MemberRole;
|
||||||
@@ -234,7 +235,14 @@ public final class BridgedApp {
|
|||||||
List<Map<String, Object>> out = sessions.roster().stream()
|
List<Map<String, Object>> out = sessions.roster().stream()
|
||||||
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
|
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
|
||||||
.toList();
|
.toList();
|
||||||
ctx.status(200).json(Map.of("workers", out));
|
Map<String, Object> body = new LinkedHashMap<>();
|
||||||
|
body.put("workers", out);
|
||||||
|
// CB-586: operator visibility for the refs/wip snapshot store without shelling into the
|
||||||
|
// repo — how many snapshot refs exist and roughly what they cost. Present only once a
|
||||||
|
// worktree session has established the repo, so a never-snapshotted fleet reports nothing.
|
||||||
|
sessions.wipRefs().ifPresent(st -> body.put("wipRefs",
|
||||||
|
Map.of("count", st.count(), "costBytes", st.costBytes())));
|
||||||
|
ctx.status(200).json(body);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** The configured worker profiles and which one a no-argument spawn uses. */
|
/** The configured worker profiles and which one a no-argument spawn uses. */
|
||||||
@@ -292,6 +300,11 @@ public final class BridgedApp {
|
|||||||
ctx.status(201).json(view(member));
|
ctx.status(201).json(view(member));
|
||||||
} catch (GuardException e) {
|
} catch (GuardException e) {
|
||||||
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
|
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
|
||||||
|
} catch (PlacementException e) {
|
||||||
|
// CB-599: no candidate had capacity (maxLoad, quarantine, or all-exhausted) — a benign,
|
||||||
|
// likely-transient refusal, distinct from "profile does not exist" below. 503: the
|
||||||
|
// request was valid and will likely succeed later.
|
||||||
|
ctx.status(503).json(Map.of("error", "no_capacity", "detail", e.getMessage()));
|
||||||
} catch (IllegalArgumentException e) {
|
} catch (IllegalArgumentException e) {
|
||||||
ctx.status(400).json(Map.of("error", "unknown_profile", "detail", e.getMessage()));
|
ctx.status(400).json(Map.of("error", "unknown_profile", "detail", e.getMessage()));
|
||||||
} catch (PeerUnreachableException e) {
|
} catch (PeerUnreachableException e) {
|
||||||
@@ -503,10 +516,20 @@ public final class BridgedApp {
|
|||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
try {
|
try {
|
||||||
ctx.status(200).json(Map.of(
|
Map<String, Object> body = new LinkedHashMap<>();
|
||||||
"sessionId", id,
|
body.put("sessionId", id);
|
||||||
"status", messages.status(id).name().toLowerCase(),
|
body.put("status", messages.status(id).name().toLowerCase());
|
||||||
"ready", presence.isPresent(id)));
|
body.put("ready", presence.isPresent(id));
|
||||||
|
// CB-582: a worker paused mid-turn in an async bridge_ask is otherwise invisible to a
|
||||||
|
// status poll — surface the open question and how to answer it, same as bridge_poll's
|
||||||
|
// Phase.ASKING view.
|
||||||
|
MessageService.PendingAsk ask = messages.pendingAsk(id);
|
||||||
|
if (ask != null) {
|
||||||
|
body.put("question", ask.question());
|
||||||
|
body.put("turnId", ask.turnId());
|
||||||
|
body.put("ticket", ask.ticket());
|
||||||
|
}
|
||||||
|
ctx.status(200).json(body);
|
||||||
} catch (HerdrException e) {
|
} catch (HerdrException e) {
|
||||||
herdrError(ctx, e);
|
herdrError(ctx, e);
|
||||||
}
|
}
|
||||||
@@ -532,6 +555,11 @@ public final class BridgedApp {
|
|||||||
if (v.detail() != null) {
|
if (v.detail() != null) {
|
||||||
body.put("detail", v.detail());
|
body.put("detail", v.detail());
|
||||||
}
|
}
|
||||||
|
// CB-582: Phase.ASKING carries the question in v.reply() (handled above) and its answer-
|
||||||
|
// correlation id here — a REST caller polling this ticket otherwise has no way to answer it.
|
||||||
|
if (v.turnId() != null) {
|
||||||
|
body.put("turnId", v.turnId());
|
||||||
|
}
|
||||||
ctx.status(200).json(body);
|
ctx.status(200).json(body);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -12,7 +12,12 @@ import java.nio.file.Files;
|
|||||||
import java.nio.file.Path;
|
import java.nio.file.Path;
|
||||||
import java.nio.file.StandardCopyOption;
|
import java.nio.file.StandardCopyOption;
|
||||||
import java.security.SecureRandom;
|
import java.security.SecureRandom;
|
||||||
|
import java.util.ArrayList;
|
||||||
|
import java.util.HashSet;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.Optional;
|
||||||
|
import java.util.Set;
|
||||||
import java.util.concurrent.TimeUnit;
|
import java.util.concurrent.TimeUnit;
|
||||||
import java.util.concurrent.atomic.AtomicLong;
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
import java.util.stream.Collectors;
|
import java.util.stream.Collectors;
|
||||||
@@ -217,6 +222,209 @@ public final class GitWorktrees implements Worktrees {
|
|||||||
return Path.of(out.trim()).toAbsolutePath().normalize().toString();
|
return Path.of(out.trim()).toAbsolutePath().normalize().toString();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-578 stage C, fixed by CB-587. Stages into a <em>temporary</em> index (never the worktree's
|
||||||
|
* real one, which the worker may still be writing to) — but that temp index is first <em>seeded</em>
|
||||||
|
* from the worktree's real one, rather than starting empty:
|
||||||
|
*
|
||||||
|
* <pre>
|
||||||
|
* cp $(git -C worktree rev-parse --git-path index) <temp>
|
||||||
|
* GIT_INDEX_FILE=<temp> git -C worktree add -A
|
||||||
|
* tree=$(GIT_INDEX_FILE=<temp> git -C worktree write-tree)
|
||||||
|
* commit=$(git -C worktree commit-tree $tree -p HEAD -m message)
|
||||||
|
* git -C worktree update-ref refs/wip/branch $commit
|
||||||
|
* </pre>
|
||||||
|
*
|
||||||
|
* A fresh empty index carries none of the real index's {@code --skip-worktree} /
|
||||||
|
* {@code --assume-unchanged} bits, so {@code add -A} into it stages a skip-worktree file's local
|
||||||
|
* on-disk content even though {@code git status --porcelain} correctly hides that file (CB-587).
|
||||||
|
* Seeding from the real index preserves those bits, so {@code add -A} then skips exactly what
|
||||||
|
* {@code git status} skips. {@code add -A} (never {@code -f}) also still respects
|
||||||
|
* {@code .gitignore} exactly as it would in the real index — a gitignored file staying ignored is
|
||||||
|
* what keeps secrets and local config out of the snapshot's tree. The temporary index file is
|
||||||
|
* removed afterwards regardless of outcome; the worker's real index is never opened for writing.
|
||||||
|
*/
|
||||||
|
@Override
|
||||||
|
public Optional<String> snapshot(String worktreePath, String branch, String message) {
|
||||||
|
if (!Files.exists(Path.of(worktreePath))) {
|
||||||
|
log.debug("worktree {} already gone — nothing to snapshot", worktreePath);
|
||||||
|
return Optional.empty();
|
||||||
|
}
|
||||||
|
Path tempIndex;
|
||||||
|
try {
|
||||||
|
tempIndex = Files.createTempFile("bridged-wip-index-", ".tmp");
|
||||||
|
} catch (IOException e) {
|
||||||
|
throw new WorktreeException("cannot create a temporary index for snapshot: " + e.getMessage(), e);
|
||||||
|
}
|
||||||
|
Map<String, String> indexEnv = Map.of("GIT_INDEX_FILE", tempIndex.toAbsolutePath().toString());
|
||||||
|
try {
|
||||||
|
Path realIndex = resolveRealIndex(worktreePath);
|
||||||
|
try {
|
||||||
|
Files.copy(realIndex, tempIndex, StandardCopyOption.REPLACE_EXISTING);
|
||||||
|
} catch (IOException e) {
|
||||||
|
throw new WorktreeException("cannot copy the worktree's real index (" + realIndex
|
||||||
|
+ ") into the temporary snapshot index: " + e.getMessage(), e);
|
||||||
|
}
|
||||||
|
exec(indexEnv, "git", "-C", worktreePath, "add", "-A");
|
||||||
|
String tree = exec(indexEnv, "git", "-C", worktreePath, "write-tree").trim();
|
||||||
|
String commit = exec("git", "-C", worktreePath, "commit-tree", tree, "-p", "HEAD", "-m", message).trim();
|
||||||
|
exec("git", "-C", worktreePath, "update-ref", "refs/wip/" + branch, commit);
|
||||||
|
log.info("snapshotted worktree {} to refs/wip/{} commit={}", worktreePath, branch, commit);
|
||||||
|
return Optional.of(commit);
|
||||||
|
} finally {
|
||||||
|
try {
|
||||||
|
Files.deleteIfExists(tempIndex);
|
||||||
|
} catch (IOException e) {
|
||||||
|
log.debug("could not delete temporary snapshot index {}: {}", tempIndex, e.getMessage());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Resolve the path of {@code worktreePath}'s real index. Never assume {@code <worktree>/.git/index}:
|
||||||
|
* in a linked worktree {@code .git} is a <em>file</em> pointing at the main repo's
|
||||||
|
* {@code worktrees/<name>/} directory, and that is where the real per-worktree index lives.
|
||||||
|
* {@code git rev-parse --git-path index} resolves this correctly for both a linked worktree and
|
||||||
|
* the main checkout. Throws {@link WorktreeException} — same as every other failure in this
|
||||||
|
* class — if the command fails or the resolved path does not exist, rather than silently
|
||||||
|
* snapshotting from an empty index.
|
||||||
|
*/
|
||||||
|
private Path resolveRealIndex(String worktreePath) {
|
||||||
|
String out = exec("git", "-C", worktreePath, "rev-parse", "--git-path", "index").trim();
|
||||||
|
Path index = Path.of(out);
|
||||||
|
if (!index.isAbsolute()) {
|
||||||
|
index = Path.of(worktreePath).resolve(index).normalize();
|
||||||
|
}
|
||||||
|
if (!Files.exists(index)) {
|
||||||
|
throw new WorktreeException("worktree's real index not found at resolved path " + index
|
||||||
|
+ " (git rev-parse --git-path index reported '" + out + "')");
|
||||||
|
}
|
||||||
|
return index;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* One {@code refs/wip/<branch>} snapshot ref as read by {@link #listWipRefs}: its full ref name,
|
||||||
|
* the snapshot commit's sha, and that commit's committer time in unix millis (the age of the
|
||||||
|
* snapshot — a snapshot is written once and never rewritten, so the commit date is the ref's).
|
||||||
|
*/
|
||||||
|
private record WipRef(String refName, String sha, long committerMillis) {
|
||||||
|
String branch() {
|
||||||
|
return refName.substring("refs/wip/".length());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public WipRefStats wipRefs(String repoRoot) {
|
||||||
|
List<WipRef> refs = listWipRefs(repoRoot);
|
||||||
|
long costBytes = 0;
|
||||||
|
for (WipRef ref : refs) {
|
||||||
|
costBytes += treeSize(repoRoot, ref.sha());
|
||||||
|
}
|
||||||
|
return new WipRefStats(refs.size(), costBytes);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
|
||||||
|
// The rule is documented on Worktrees#pruneWipRefs: delete only a snapshot whose tree
|
||||||
|
// content is already reachable from main AND that is older than minAgeMillis. Reachability
|
||||||
|
// is the floor that keeps a worker's last copy; the age floor keeps a just-written snapshot
|
||||||
|
// from being swept while a lead may still be looking at it.
|
||||||
|
List<WipRef> refs = listWipRefs(repoRoot);
|
||||||
|
if (refs.isEmpty()) {
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
long nowMillis = System.currentTimeMillis();
|
||||||
|
// Resolve what main carries once per sweep, not once per ref.
|
||||||
|
Set<String> mainObjects = reachableObjectsFromMain(repoRoot);
|
||||||
|
int deleted = 0;
|
||||||
|
for (WipRef ref : refs) {
|
||||||
|
long ageMillis = nowMillis - ref.committerMillis();
|
||||||
|
if (ageMillis <= minAgeMillis) {
|
||||||
|
continue; // too recent — never swept, even if it looks recoverable (CB-586)
|
||||||
|
}
|
||||||
|
String tree = exec("git", "-C", repoRoot, "rev-parse", ref.sha() + "^{tree}").trim();
|
||||||
|
if (!mainObjects.contains(tree)) {
|
||||||
|
// Last copy of the snapshot's content — the worker's work exists nowhere else.
|
||||||
|
// Never delete automatically (CB-586 criterion 2).
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
exec("git", "-C", repoRoot, "update-ref", "-d", ref.refName());
|
||||||
|
deleted++;
|
||||||
|
log.info("pruned snapshot ref refs/wip/{} commit={} (age {}h): its tree is already "
|
||||||
|
+ "reachable from main, so the work is preserved; recover from reflog via "
|
||||||
|
+ "git update-ref refs/wip/{} {}",
|
||||||
|
ref.branch(), ref.sha(), TimeUnit.MILLISECONDS.toHours(ageMillis),
|
||||||
|
ref.branch(), ref.sha());
|
||||||
|
}
|
||||||
|
return deleted;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Every {@code refs/wip/*} ref (see {@link WipRef}). The committer date is read as a unix
|
||||||
|
* count of seconds and converted to millis. {@code %00} (NUL) separates the fields because a
|
||||||
|
* branch name may contain spaces.
|
||||||
|
*/
|
||||||
|
private List<WipRef> listWipRefs(String repoRoot) {
|
||||||
|
String out = exec("git", "-C", repoRoot, "for-each-ref",
|
||||||
|
"--format=%(refname)%00%(objectname)%00%(committerdate:unix)", "refs/wip/");
|
||||||
|
List<WipRef> refs = new ArrayList<>();
|
||||||
|
for (String line : out.split("\\R")) {
|
||||||
|
if (line.isBlank()) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
String[] parts = line.split("\u0000", -1);
|
||||||
|
if (parts.length == 3 && !parts[1].isBlank()) {
|
||||||
|
refs.add(new WipRef(parts[0], parts[1], Long.parseLong(parts[2]) * 1000L));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return refs;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The set of object shas reachable from {@code main}, or an empty set when {@code main} cannot
|
||||||
|
* be resolved. An empty set is the safe direction: the retention sweep then concludes nothing
|
||||||
|
* is recoverable, so it deletes nothing — a repo with no {@code main} must never cause a
|
||||||
|
* worker's last copy of a snapshot to be dropped on a reachability misreading.
|
||||||
|
*/
|
||||||
|
private Set<String> reachableObjectsFromMain(String repoRoot) {
|
||||||
|
if (exitCode("git", "-C", repoRoot, "rev-parse", "--verify", "main") != 0) {
|
||||||
|
log.debug("refs/wip retention: no 'main' ref in {} — treating nothing as reachable", repoRoot);
|
||||||
|
return Set.of();
|
||||||
|
}
|
||||||
|
String out = exec("git", "-C", repoRoot, "rev-list", "--objects", "main");
|
||||||
|
Set<String> objects = new HashSet<>();
|
||||||
|
for (String line : out.split("\\R")) {
|
||||||
|
if (line.isBlank()) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
int sp = line.indexOf(' ');
|
||||||
|
objects.add(sp < 0 ? line : line.substring(0, sp));
|
||||||
|
}
|
||||||
|
return objects;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Approximate cost of a snapshot: the sum of every blob's size in its committed tree. */
|
||||||
|
private long treeSize(String repoRoot, String sha) {
|
||||||
|
String out = exec("git", "-C", repoRoot, "ls-tree", "-r", "-l", sha);
|
||||||
|
long total = 0;
|
||||||
|
for (String line : out.split("\\R")) {
|
||||||
|
if (line.isBlank()) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
// ls-tree -l row: "<mode> <type> <object> <size>\t<path>"; the size is only numeric for
|
||||||
|
// blobs (trees read "-"), so gate on the type token and take the 4th whitespace field.
|
||||||
|
String[] parts = line.split("\\s+");
|
||||||
|
if (parts.length >= 4 && "blob".equals(parts[1])) {
|
||||||
|
try {
|
||||||
|
total += Long.parseLong(parts[3]);
|
||||||
|
} catch (NumberFormatException ignored) {
|
||||||
|
// a '-' size (or any anomaly) contributes nothing to the rough figure
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return total;
|
||||||
|
}
|
||||||
|
|
||||||
/** Resolve the directory that will hold per-session worktree checkouts. */
|
/** Resolve the directory that will hold per-session worktree checkouts. */
|
||||||
private Path resolveRoot(String repoRoot) {
|
private Path resolveRoot(String repoRoot) {
|
||||||
if (configuredRoot != null && !configuredRoot.isBlank()) {
|
if (configuredRoot != null && !configuredRoot.isBlank()) {
|
||||||
@@ -239,11 +447,20 @@ public final class GitWorktrees implements Worktrees {
|
|||||||
* stdout and stderr (merged by redirectErrorStream).
|
* stdout and stderr (merged by redirectErrorStream).
|
||||||
*/
|
*/
|
||||||
private String exec(String... command) {
|
private String exec(String... command) {
|
||||||
|
return exec(Map.of(), command);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Same as {@link #exec(String...)}, with extra environment variables set on the child process. */
|
||||||
|
private String exec(Map<String, String> extraEnv, String... command) {
|
||||||
String out;
|
String out;
|
||||||
int code;
|
int code;
|
||||||
Process p;
|
Process p;
|
||||||
try {
|
try {
|
||||||
p = new ProcessBuilder(command).redirectErrorStream(true).start();
|
ProcessBuilder pb = new ProcessBuilder(command).redirectErrorStream(true);
|
||||||
|
if (extraEnv != null && !extraEnv.isEmpty()) {
|
||||||
|
pb.environment().putAll(extraEnv);
|
||||||
|
}
|
||||||
|
p = pb.start();
|
||||||
} catch (IOException e) {
|
} catch (IOException e) {
|
||||||
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
|
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -26,6 +26,11 @@ import dev.ltms.bridged.peer.MemberRole;
|
|||||||
* @param state current lifecycle state in the one-shot FSM
|
* @param state current lifecycle state in the one-shot FSM
|
||||||
* @param charterReceipt the fingerprint (CB-571) of the charter bytes this member was started
|
* @param charterReceipt the fingerprint (CB-571) of the charter bytes this member was started
|
||||||
* with; {@code null} for a session whose launcher recorded none
|
* with; {@code null} for a session whose launcher recorded none
|
||||||
|
* @param agentSessionId the peer's OWN session id (CB-584) — the handle a later
|
||||||
|
* {@code resumeSessionId} spawn would pass back to resume this exact
|
||||||
|
* conversation; {@code null} when the launcher could not determine one
|
||||||
|
* (an adapter that declines {@code Capability.SESSION_RESUME}, or one
|
||||||
|
* that resolves it lazily and has not yet)
|
||||||
*/
|
*/
|
||||||
public record MemberSession(
|
public record MemberSession(
|
||||||
String paneId,
|
String paneId,
|
||||||
@@ -40,7 +45,8 @@ public record MemberSession(
|
|||||||
State state,
|
State state,
|
||||||
String worktree,
|
String worktree,
|
||||||
String branch,
|
String branch,
|
||||||
CharterReceipt charterReceipt) {
|
CharterReceipt charterReceipt,
|
||||||
|
String agentSessionId) {
|
||||||
|
|
||||||
/** One-shot worker lifecycle states. */
|
/** One-shot worker lifecycle states. */
|
||||||
public enum State {
|
public enum State {
|
||||||
@@ -53,33 +59,34 @@ public record MemberSession(
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Backward-compatible shape: a session with no charter receipt (a test or a launcher before
|
* Backward-compatible shape: a session with no charter receipt and no agent session id (a test
|
||||||
* CB-571). A separate constructor rather than a new parameter on the canonical one, so existing
|
* or a launcher before CB-571 / CB-584). A separate constructor rather than a new parameter on
|
||||||
* call sites that have nothing to record keep compiling unchanged.
|
* the canonical one, so existing call sites that have nothing to record keep compiling
|
||||||
|
* unchanged.
|
||||||
*/
|
*/
|
||||||
public MemberSession(String paneId, String terminalId, String profile, MemberRole role,
|
public MemberSession(String paneId, String terminalId, String profile, MemberRole role,
|
||||||
String cwd, String ownerTerminal, long spawnedAtNanos,
|
String cwd, String ownerTerminal, long spawnedAtNanos,
|
||||||
long lastActivityAtNanos, int turnCount, State state,
|
long lastActivityAtNanos, int turnCount, State state,
|
||||||
String worktree, String branch) {
|
String worktree, String branch) {
|
||||||
this(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
this(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
||||||
lastActivityAtNanos, turnCount, state, worktree, branch, null);
|
lastActivityAtNanos, turnCount, state, worktree, branch, null, null);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Return a copy of this session in {@code state}. */
|
/** Return a copy of this session in {@code state}. */
|
||||||
public MemberSession withState(State state) {
|
public MemberSession withState(State state) {
|
||||||
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
||||||
lastActivityAtNanos, turnCount, state, worktree, branch, charterReceipt);
|
lastActivityAtNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
|
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
|
||||||
public MemberSession withActivity(long nowNanos) {
|
public MemberSession withActivity(long nowNanos) {
|
||||||
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
||||||
nowNanos, turnCount, state, worktree, branch, charterReceipt);
|
nowNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
|
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
|
||||||
public MemberSession bumpTurn(long nowNanos) {
|
public MemberSession bumpTurn(long nowNanos) {
|
||||||
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
||||||
nowNanos, turnCount + 1, state, worktree, branch, charterReceipt);
|
nowNanos, turnCount + 1, state, worktree, branch, charterReceipt, agentSessionId);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -5,6 +5,7 @@ import dev.ltms.bridged.herdr.Agent;
|
|||||||
import dev.ltms.bridged.inject.TurnListener;
|
import dev.ltms.bridged.inject.TurnListener;
|
||||||
import dev.ltms.bridged.inject.MemberPresence;
|
import dev.ltms.bridged.inject.MemberPresence;
|
||||||
import dev.ltms.bridged.msg.TurnToken;
|
import dev.ltms.bridged.msg.TurnToken;
|
||||||
|
import dev.ltms.bridged.peer.Capability;
|
||||||
import dev.ltms.bridged.peer.MemberRole;
|
import dev.ltms.bridged.peer.MemberRole;
|
||||||
import dev.ltms.bridged.peer.PeerHandle;
|
import dev.ltms.bridged.peer.PeerHandle;
|
||||||
import dev.ltms.bridged.peer.PeerLauncher;
|
import dev.ltms.bridged.peer.PeerLauncher;
|
||||||
@@ -17,6 +18,7 @@ import java.util.LinkedHashMap;
|
|||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
import java.util.Optional;
|
import java.util.Optional;
|
||||||
|
import java.util.Set;
|
||||||
import java.util.concurrent.ConcurrentHashMap;
|
import java.util.concurrent.ConcurrentHashMap;
|
||||||
import java.util.concurrent.TimeUnit;
|
import java.util.concurrent.TimeUnit;
|
||||||
import java.util.concurrent.atomic.AtomicLong;
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
@@ -51,11 +53,19 @@ public final class SessionManager implements TurnListener {
|
|||||||
private final int contextCap;
|
private final int contextCap;
|
||||||
private final boolean clearAfterTurn;
|
private final boolean clearAfterTurn;
|
||||||
private volatile MemberLifecycle memberLifecycle = MemberLifecycle.NONE;
|
private volatile MemberLifecycle memberLifecycle = MemberLifecycle.NONE;
|
||||||
|
/**
|
||||||
|
* CB-586: the repo root the fleet actually works in, remembered the first time a worktree
|
||||||
|
* session is spawned (worktrees are checkouts of it). {@code refs/wip/*} live there, and this
|
||||||
|
* single cached value is what the snapshot retention sweep and the operator-visible census run
|
||||||
|
* against. The daemon is bridged into one project at a time, so "the first worktree's repo" is
|
||||||
|
* the repo; {@code null} until any worktree is spawned, meaning nothing to sweep or measure.
|
||||||
|
*/
|
||||||
|
private volatile String fleetRepoRoot;
|
||||||
|
|
||||||
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
|
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
|
||||||
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
|
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||||
/** CB-516: notified with a terminalId on every release; no-op until wired. */
|
/** CB-516: notified with a {@link ReleaseDetail} on every release; no-op until wired. */
|
||||||
private final List<Consumer<String>> releaseListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
|
private final List<Consumer<ReleaseDetail>> releaseListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||||
|
|
||||||
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
|
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
|
||||||
public SessionManager(PeerLauncher launcher) {
|
public SessionManager(PeerLauncher launcher) {
|
||||||
@@ -126,21 +136,49 @@ public final class SessionManager implements TurnListener {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Spawn a member, optionally inside a fresh git worktree. When {@code wt} is non-null the
|
* Spawn a member, optionally inside a fresh git worktree, with no session identity requested.
|
||||||
* worktree is provisioned, parity-overlaid, and its path becomes the member's cwd. On any
|
* Equivalent to {@link #acquire(String, MemberRole, String, String, String, WorktreeRequest,
|
||||||
* failure before registration the worktree is removed so no dangling checkout is left.
|
* String, String)} with both trailing args {@code null}.
|
||||||
*
|
*
|
||||||
* @param profile which backend to run on — a {@code profiles:} key
|
* @param profile which backend to run on — a {@code profiles:} key
|
||||||
* @param role which contract the member runs under; never {@code null}
|
* @param role which contract the member runs under; never {@code null}
|
||||||
*/
|
*/
|
||||||
public MemberSession acquire(String profile, MemberRole role, String requestedCwd, String callerCwd,
|
public MemberSession acquire(String profile, MemberRole role, String requestedCwd, String callerCwd,
|
||||||
String ownerTerminal, WorktreeRequest wt) {
|
String ownerTerminal, WorktreeRequest wt) {
|
||||||
|
return acquire(profile, role, requestedCwd, callerCwd, ownerTerminal, wt, null, null);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Spawn a member, optionally inside a fresh git worktree, optionally onto a chosen or resumed
|
||||||
|
* agent session (CB-547a / CB-584). When {@code wt} is non-null the worktree is provisioned,
|
||||||
|
* parity-overlaid, and its path becomes the member's cwd. On any failure before registration the
|
||||||
|
* worktree is removed so no dangling checkout is left.
|
||||||
|
*
|
||||||
|
* <p>A non-blank {@code resumeSessionId} requires an explicit {@code profile}: a resumed
|
||||||
|
* conversation is tied to the specific backend that started it, so an unqualified spawn (whose
|
||||||
|
* backend a placement policy picks at spawn time) has no safe candidate to check the capability
|
||||||
|
* against. It also requires that profile's adapter to declare {@link Capability#SESSION_RESUME};
|
||||||
|
* refusing rather than silently delivering a cold session on a non-supporting adapter is the
|
||||||
|
* whole point of checking before spawn, not after (CB-584).
|
||||||
|
*
|
||||||
|
* @param profile which backend to run on — a {@code profiles:} key
|
||||||
|
* @param role which contract the member runs under; never {@code null}
|
||||||
|
* @param sessionName the bridge's logical name for the session, or {@code null}
|
||||||
|
* @param resumeSessionId the prior agent session to resume, or {@code null} for a fresh one
|
||||||
|
* @throws IllegalArgumentException if {@code resumeSessionId} is set with no explicit profile,
|
||||||
|
* or the resolved profile's adapter lacks
|
||||||
|
* {@link Capability#SESSION_RESUME}
|
||||||
|
*/
|
||||||
|
public MemberSession acquire(String profile, MemberRole role, String requestedCwd, String callerCwd,
|
||||||
|
String ownerTerminal, WorktreeRequest wt,
|
||||||
|
String sessionName, String resumeSessionId) {
|
||||||
MemberRole memberRole = (role == null) ? MemberRole.DEV : role;
|
MemberRole memberRole = (role == null) ? MemberRole.DEV : role;
|
||||||
|
requireResumeCapability(profile, resumeSessionId);
|
||||||
if (wt == null) {
|
if (wt == null) {
|
||||||
// CB-557: the role must ride on the SpawnRequest, not stay a local. The launcher needs it
|
// CB-557: the role must ride on the SpawnRequest, not stay a local. The launcher needs it
|
||||||
// to pick the profile out of that role's pool and to label the tab; a role kept only on
|
// to pick the profile out of that role's pool and to label the tab; a role kept only on
|
||||||
// the MemberSession is recorded after the spawn it was supposed to steer.
|
// the MemberSession is recorded after the spawn it was supposed to steer.
|
||||||
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd, null, null, memberRole);
|
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd, sessionName, resumeSessionId, memberRole);
|
||||||
PeerHandle handle;
|
PeerHandle handle;
|
||||||
try {
|
try {
|
||||||
handle = launcher.spawn(req);
|
handle = launcher.spawn(req);
|
||||||
@@ -164,7 +202,8 @@ public final class SessionManager implements TurnListener {
|
|||||||
MemberSession.State.SPAWNING,
|
MemberSession.State.SPAWNING,
|
||||||
null,
|
null,
|
||||||
null,
|
null,
|
||||||
handle.charterReceipt());
|
handle.charterReceipt(),
|
||||||
|
handle.agentSessionId());
|
||||||
registry.put(handle.id(), session);
|
registry.put(handle.id(), session);
|
||||||
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
|
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
|
||||||
log.debug("acquired session id={} terminal={} profile={} owner={}",
|
log.debug("acquired session id={} terminal={} profile={} owner={}",
|
||||||
@@ -172,7 +211,30 @@ public final class SessionManager implements TurnListener {
|
|||||||
notifyAcquired(session.terminalId());
|
notifyAcquired(session.terminalId());
|
||||||
return session;
|
return session;
|
||||||
}
|
}
|
||||||
return acquireWithWorktree(profile, memberRole, requestedCwd, callerCwd, ownerTerminal, wt);
|
return acquireWithWorktree(profile, memberRole, requestedCwd, callerCwd, ownerTerminal, wt,
|
||||||
|
sessionName, resumeSessionId);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Refuse a {@code resumeSessionId} that either names no explicit profile or names one whose
|
||||||
|
* adapter does not declare {@link Capability#SESSION_RESUME}. A no-op when
|
||||||
|
* {@code resumeSessionId} is blank — the ordinary, no-identity spawn path.
|
||||||
|
*/
|
||||||
|
private void requireResumeCapability(String profile, String resumeSessionId) {
|
||||||
|
if (resumeSessionId == null || resumeSessionId.isBlank()) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (profile == null || profile.isBlank()) {
|
||||||
|
throw new IllegalArgumentException("resumeSessionId requires an explicit profile — "
|
||||||
|
+ "a resumed conversation is tied to the backend that started it, so it cannot "
|
||||||
|
+ "be left to placement to pick");
|
||||||
|
}
|
||||||
|
Set<Capability> caps = launcher.capabilitiesFor(profile);
|
||||||
|
if (!caps.contains(Capability.SESSION_RESUME)) {
|
||||||
|
throw new IllegalArgumentException("worker profile '" + profile + "' does not declare "
|
||||||
|
+ "Capability.SESSION_RESUME — refusing resumeSessionId rather than silently "
|
||||||
|
+ "starting a cold session");
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Tear a worker down by pane id and remove it from the registry. Idempotent. */
|
/** Tear a worker down by pane id and remove it from the registry. Idempotent. */
|
||||||
@@ -196,14 +258,16 @@ public final class SessionManager implements TurnListener {
|
|||||||
private void release(String paneId, ReleaseCause cause) {
|
private void release(String paneId, ReleaseCause cause) {
|
||||||
MemberSession removed = registry.remove(paneId);
|
MemberSession removed = registry.remove(paneId);
|
||||||
boolean preserveWorktree = cause == ReleaseCause.SHUTDOWN;
|
boolean preserveWorktree = cause == ReleaseCause.SHUTDOWN;
|
||||||
|
String snapshotRef = null;
|
||||||
if (removed != null) {
|
if (removed != null) {
|
||||||
try {
|
try {
|
||||||
memberLifecycle.released(removed.terminalId());
|
memberLifecycle.released(removed.terminalId());
|
||||||
log.debug("releasing session pane={} terminal={} state={} cause={}",
|
log.debug("releasing session pane={} terminal={} state={} cause={}",
|
||||||
removed.paneId(), removed.terminalId(), removed.state(), cause);
|
removed.paneId(), removed.terminalId(), removed.state(), cause);
|
||||||
|
boolean dirty = removed.worktree() != null && worktrees.hasUncommitted(removed.worktree());
|
||||||
if (preserveWorktree && removed.worktree() != null) {
|
if (preserveWorktree && removed.worktree() != null) {
|
||||||
logPreservedForShutdown(removed);
|
logPreservedForShutdown(removed);
|
||||||
} else if (removed.worktree() != null && worktrees.hasUncommitted(removed.worktree())) {
|
} else if (dirty) {
|
||||||
// CB-576: a release that would otherwise remove the worktree finds it holding
|
// CB-576: a release that would otherwise remove the worktree finds it holding
|
||||||
// uncommitted work the bridge cannot see. A worker that ends a turn without
|
// uncommitted work the bridge cannot see. A worker that ends a turn without
|
||||||
// committing (normally because it stopped to ask a question or refused the turn)
|
// committing (normally because it stopped to ask a question or refused the turn)
|
||||||
@@ -214,6 +278,13 @@ public final class SessionManager implements TurnListener {
|
|||||||
+ "the worktree holds uncommitted changes that --force remove would destroy",
|
+ "the worktree holds uncommitted changes that --force remove would destroy",
|
||||||
cause, removed.worktree(), removed.paneId(), removed.terminalId());
|
cause, removed.worktree(), removed.paneId(), removed.terminalId());
|
||||||
}
|
}
|
||||||
|
if (dirty) {
|
||||||
|
// CB-578 stage C: preserving on disk is not saving — the directory is one
|
||||||
|
// `worktree remove --force`, or an operator tidying up, away from gone. Commit
|
||||||
|
// its full state to a ref before the preserve-or-remove decision above can be
|
||||||
|
// undone by anything else, regardless of why this release fired.
|
||||||
|
snapshotRef = trySnapshot(removed, cause);
|
||||||
|
}
|
||||||
} catch (RuntimeException e) {
|
} catch (RuntimeException e) {
|
||||||
// CB-581: hasUncommitted shells out to `git status` and can throw on a non-zero
|
// CB-581: hasUncommitted shells out to `git status` and can throw on a non-zero
|
||||||
// exit. We can no longer tell whether the worktree holds uncommitted work, so fail
|
// exit. We can no longer tell whether the worktree holds uncommitted work, so fail
|
||||||
@@ -228,8 +299,12 @@ public final class SessionManager implements TurnListener {
|
|||||||
// CB-516/CB-581: a send still waiting on this worker can never be answered now, no
|
// CB-516/CB-581: a send still waiting on this worker can never be answered now, no
|
||||||
// matter what happened above. Tell the listener BEFORE the pane is torn down, so a
|
// matter what happened above. Tell the listener BEFORE the pane is torn down, so a
|
||||||
// blocked caller fails fast with a real reason instead of sitting on a rendezvous
|
// blocked caller fails fast with a real reason instead of sitting on a rendezvous
|
||||||
// nothing will ever resolve.
|
// nothing will ever resolve. CB-578 stage C: carry the worktree/branch/snapshot ref
|
||||||
notifyReleased(removed.terminalId());
|
// too, so a failed ticket's detail can point a lead at the same tree to re-dispatch.
|
||||||
|
// CB-584 (issue #65 criterion 5): carry agentSessionId alongside them, so a lead can
|
||||||
|
// also resume the member's conversation, not just re-dispatch onto its files.
|
||||||
|
notifyReleased(new ReleaseDetail(removed.terminalId(), removed.worktree(),
|
||||||
|
removed.branch(), snapshotRef, removed.agentSessionId()));
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
// CB-581: the pane must always stop, even if the dirty check above threw. A session removed
|
// CB-581: the pane must always stop, even if the dirty check above threw. A session removed
|
||||||
@@ -241,6 +316,52 @@ public final class SessionManager implements TurnListener {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Best-effort snapshot of a dirty worktree into {@code refs/wip/<branch>} (CB-578 stage C). A
|
||||||
|
* failure here must never escalate: the caller has already decided to preserve the worktree
|
||||||
|
* regardless of whether this succeeds, so the only cost of a failed snapshot is a WARN and a
|
||||||
|
* missing ref — never a lost pane stop or a lost release notification.
|
||||||
|
*/
|
||||||
|
private String trySnapshot(MemberSession session, ReleaseCause cause) {
|
||||||
|
if (session.worktree() == null || session.branch() == null) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
try {
|
||||||
|
Optional<String> ref = worktrees.snapshot(session.worktree(), session.branch(),
|
||||||
|
snapshotMessage(session, cause));
|
||||||
|
ref.ifPresent(sha -> log.info(
|
||||||
|
"snapshotted dirty worktree {} to refs/wip/{} commit={} for pane={} terminal={}",
|
||||||
|
session.worktree(), session.branch(), sha, session.paneId(), session.terminalId()));
|
||||||
|
return ref.orElse(null);
|
||||||
|
} catch (RuntimeException e) {
|
||||||
|
log.warn("snapshot of dirty worktree {} failed for pane={} terminal={} branch={}: the "
|
||||||
|
+ "worktree is still preserved on disk, just not committed to refs/wip/{}: {}",
|
||||||
|
session.worktree(), session.paneId(), session.terminalId(), session.branch(),
|
||||||
|
session.branch(), e.toString());
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Commit message for a CB-578 stage C snapshot — names the member so an operator can tell runs apart. */
|
||||||
|
private String snapshotMessage(MemberSession session, ReleaseCause cause) {
|
||||||
|
return "CB-578 stage C: snapshot of a released worker\n\n"
|
||||||
|
+ "terminal: " + session.terminalId() + "\n"
|
||||||
|
+ "profile: " + session.profile() + "\n"
|
||||||
|
+ "branch: " + session.branch() + "\n"
|
||||||
|
+ "cause: " + cause;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Facts about a released session that a listener needs beyond the bare terminal id — enough
|
||||||
|
* for a caller to point a lead at where to re-dispatch onto the same tree after a failed
|
||||||
|
* release (CB-578 stage C, acceptance criterion 10). {@code worktreePath} and {@code branch}
|
||||||
|
* are {@code null} for a shared-tree session; {@code snapshotRef} is {@code null} unless this
|
||||||
|
* release snapshotted a dirty worktree into {@code refs/wip/<branch>}.
|
||||||
|
*/
|
||||||
|
public record ReleaseDetail(String terminalId, String worktreePath, String branch, String snapshotRef,
|
||||||
|
String agentSessionId) {
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Why a session is being released — governs whether its worktree is preserved or removed.
|
* Why a session is being released — governs whether its worktree is preserved or removed.
|
||||||
* Worktree removal is reserved for the one case that is genuinely finished; everything else
|
* Worktree removal is reserved for the one case that is genuinely finished; everything else
|
||||||
@@ -289,7 +410,7 @@ public final class SessionManager implements TurnListener {
|
|||||||
* presence view this manager exposes). Wiring it at construction would require breaking that
|
* presence view this manager exposes). Wiring it at construction would require breaking that
|
||||||
* cycle for one callback.
|
* cycle for one callback.
|
||||||
*/
|
*/
|
||||||
public void onRelease(Consumer<String> listener) {
|
public void onRelease(Consumer<ReleaseDetail> listener) {
|
||||||
if (listener != null) {
|
if (listener != null) {
|
||||||
releaseListeners.add(listener);
|
releaseListeners.add(listener);
|
||||||
}
|
}
|
||||||
@@ -315,21 +436,22 @@ public final class SessionManager implements TurnListener {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/** A listener failure must never prevent the teardown it is reacting to. */
|
/** A listener failure must never prevent the teardown it is reacting to. */
|
||||||
private void notifyReleased(String terminalId) {
|
private void notifyReleased(ReleaseDetail detail) {
|
||||||
if (terminalId == null) {
|
if (detail.terminalId() == null) {
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
for (Consumer<String> listener : releaseListeners) {
|
for (Consumer<ReleaseDetail> listener : releaseListeners) {
|
||||||
try {
|
try {
|
||||||
listener.accept(terminalId);
|
listener.accept(detail);
|
||||||
} catch (RuntimeException e) {
|
} catch (RuntimeException e) {
|
||||||
log.warn("release listener failed for terminal {}: {}", terminalId, e.toString());
|
log.warn("release listener failed for terminal {}: {}", detail.terminalId(), e.toString());
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
private MemberSession acquireWithWorktree(String profile, MemberRole memberRole, String requestedCwd, String callerCwd,
|
private MemberSession acquireWithWorktree(String profile, MemberRole memberRole, String requestedCwd, String callerCwd,
|
||||||
String ownerTerminal, WorktreeRequest wt) {
|
String ownerTerminal, WorktreeRequest wt,
|
||||||
|
String sessionName, String resumeSessionId) {
|
||||||
String preResolvedProfile = (profile == null || profile.isBlank())
|
String preResolvedProfile = (profile == null || profile.isBlank())
|
||||||
? launcher.defaultProfile() : profile;
|
? launcher.defaultProfile() : profile;
|
||||||
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
|
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
|
||||||
@@ -339,13 +461,18 @@ public final class SessionManager implements TurnListener {
|
|||||||
// The non-worktree path always used this chain; only this branch was missed.
|
// The non-worktree path always used this chain; only this branch was missed.
|
||||||
String repoRoot = worktrees.repoRoot(
|
String repoRoot = worktrees.repoRoot(
|
||||||
launcher.effectiveCwd(new SpawnRequest(preResolvedProfile, requestedCwd, callerCwd)));
|
launcher.effectiveCwd(new SpawnRequest(preResolvedProfile, requestedCwd, callerCwd)));
|
||||||
|
if (fleetRepoRoot == null) {
|
||||||
|
// CB-586: remember the repo whose worktrees the fleet spawns — its refs/wip/* are the
|
||||||
|
// snapshot store the retention sweep and the operator census operate on.
|
||||||
|
fleetRepoRoot = repoRoot;
|
||||||
|
}
|
||||||
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
|
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
|
||||||
String path = null;
|
String path = null;
|
||||||
PeerHandle handle;
|
PeerHandle handle;
|
||||||
try {
|
try {
|
||||||
path = worktrees.add(repoRoot, branch, wt.baseRef());
|
path = worktrees.add(repoRoot, branch, wt.baseRef());
|
||||||
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(preResolvedProfile));
|
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(preResolvedProfile));
|
||||||
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, null, null, memberRole));
|
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, sessionName, resumeSessionId, memberRole));
|
||||||
} catch (RuntimeException e) {
|
} catch (RuntimeException e) {
|
||||||
log.warn("spawn failed for profile={} role={} branch={} path={}: {}",
|
log.warn("spawn failed for profile={} role={} branch={} path={}: {}",
|
||||||
preResolvedProfile, memberRole, branch, path, e.getMessage());
|
preResolvedProfile, memberRole, branch, path, e.getMessage());
|
||||||
@@ -374,7 +501,8 @@ public final class SessionManager implements TurnListener {
|
|||||||
MemberSession.State.SPAWNING,
|
MemberSession.State.SPAWNING,
|
||||||
path,
|
path,
|
||||||
branch,
|
branch,
|
||||||
handle.charterReceipt());
|
handle.charterReceipt(),
|
||||||
|
handle.agentSessionId());
|
||||||
registry.put(handle.id(), session);
|
registry.put(handle.id(), session);
|
||||||
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
|
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
|
||||||
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
|
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
|
||||||
@@ -440,6 +568,12 @@ public final class SessionManager implements TurnListener {
|
|||||||
if (session.ownerTerminal() != null) {
|
if (session.ownerTerminal() != null) {
|
||||||
m.put("owner", session.ownerTerminal());
|
m.put("owner", session.ownerTerminal());
|
||||||
}
|
}
|
||||||
|
// CB-584: which conversation this member holds — the id a later resumeSessionId spawn would
|
||||||
|
// pass back. Absent for an adapter that declines Capability.SESSION_RESUME, or one that
|
||||||
|
// resolves it lazily and has not yet.
|
||||||
|
if (session.agentSessionId() != null) {
|
||||||
|
m.put("agentSessionId", session.agentSessionId());
|
||||||
|
}
|
||||||
// CB-571: which charter this member was started with — never the charter text itself. The
|
// CB-571: which charter this member was started with — never the charter text itself. The
|
||||||
// digest lets a lead tell at a glance whether all members got the same charter; the source
|
// digest lets a lead tell at a glance whether all members got the same charter; the source
|
||||||
// records whether a role charter was configured ("fleet.charters.<role>") or only the reply
|
// records whether a role charter was configured ("fleet.charters.<role>") or only the reply
|
||||||
@@ -622,6 +756,26 @@ public final class SessionManager implements TurnListener {
|
|||||||
return registry.size();
|
return registry.size();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586: the operator-visible census of {@code refs/wip/*} in the repo the fleet works in —
|
||||||
|
* how many snapshot refs exist and roughly what they cost. Empty (no repo known) until at
|
||||||
|
* least one worktree session has been spawned, exactly so a fleet that has never snapshotted
|
||||||
|
* anything surfaces nothing new, as it did before CB-586.
|
||||||
|
*/
|
||||||
|
public Optional<Worktrees.WipRefStats> wipRefs() {
|
||||||
|
String repo = fleetRepoRoot;
|
||||||
|
return repo == null ? Optional.empty() : Optional.of(worktrees.wipRefs(repo));
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586: run the snapshot retention sweep in the fleet's repo (a no-op until a worktree has
|
||||||
|
* been spawned, which establishes the repo). Returns how many {@code refs/wip/*} it deleted.
|
||||||
|
*/
|
||||||
|
public int sweepWipRefs(long minAgeMillis) {
|
||||||
|
String repo = fleetRepoRoot;
|
||||||
|
return repo == null ? 0 : worktrees.pruneWipRefs(repo, minAgeMillis);
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* The registered session owning {@code terminalId}, or {@code null} if none does.
|
* The registered session owning {@code terminalId}, or {@code null} if none does.
|
||||||
*
|
*
|
||||||
|
|||||||
@@ -14,12 +14,29 @@ public final class SessionReaper {
|
|||||||
|
|
||||||
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
|
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
|
||||||
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
|
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
|
||||||
|
/** CB-586: the refs/wip age floor — never sweep a snapshot younger than 24h (the CB-586 rule). */
|
||||||
|
private static final long WIP_MIN_AGE_MILLIS = TimeUnit.HOURS.toMillis(24);
|
||||||
|
/**
|
||||||
|
* CB-586: how often the retention sweep runs. Given the 24h age floor, running it every few
|
||||||
|
* hours means a ref is dropped within hours of becoming eligible, never within minutes.
|
||||||
|
*/
|
||||||
|
private static final long WIP_SWEEP_INTERVAL_NANOS = TimeUnit.HOURS.toNanos(6);
|
||||||
|
|
||||||
private final SessionManager sessions;
|
private final SessionManager sessions;
|
||||||
private final long idleTtlNanos;
|
private final long idleTtlNanos;
|
||||||
private final long intervalMillis;
|
private final long intervalMillis;
|
||||||
private volatile boolean running;
|
private volatile boolean running;
|
||||||
private Thread thread;
|
private Thread thread;
|
||||||
|
/**
|
||||||
|
* When the retention sweep last ran, and whether it ever has. The flag is not a convenience:
|
||||||
|
* a "never yet" sentinel value cannot be compared by subtraction. {@code Long.MIN_VALUE} was
|
||||||
|
* the obvious choice and it silently overflows — {@code System.nanoTime()} is positive on this
|
||||||
|
* platform, so {@code now - Long.MIN_VALUE} wraps to a large negative number, the interval gate
|
||||||
|
* reads it as "swept moments ago", and it returns before ever assigning the field. The sweep
|
||||||
|
* then never runs at all, for the life of the process, with nothing in the log to say so.
|
||||||
|
*/
|
||||||
|
private volatile boolean sweptOnce;
|
||||||
|
private volatile long lastWipSweepNanos;
|
||||||
|
|
||||||
/** Construct a reaper with the default 5-second polling interval. */
|
/** Construct a reaper with the default 5-second polling interval. */
|
||||||
public SessionReaper(SessionManager sessions, long idleTtlSeconds) {
|
public SessionReaper(SessionManager sessions, long idleTtlSeconds) {
|
||||||
@@ -49,10 +66,39 @@ public final class SessionReaper {
|
|||||||
} catch (RuntimeException e) {
|
} catch (RuntimeException e) {
|
||||||
log.warn("session reaper iteration failed; continuing", e);
|
log.warn("session reaper iteration failed; continuing", e);
|
||||||
}
|
}
|
||||||
|
maybeSweepWipRefs();
|
||||||
sleep();
|
sleep();
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586: run the refs/wip retention sweep on a slow cadence (hours, not the per-iteration
|
||||||
|
* millisecond loop). Best-effort — a failure must never take the idle-reap loop down with it.
|
||||||
|
*/
|
||||||
|
private void maybeSweepWipRefs() {
|
||||||
|
long now = System.nanoTime();
|
||||||
|
// The first pass always sweeps: a restart is a fine moment to sweep, the 24h age floor
|
||||||
|
// makes it safe, and it means the feature is observable right after a redeploy instead of
|
||||||
|
// six hours later. Only after that does the interval gate apply, and by then both operands
|
||||||
|
// come from nanoTime, so the subtraction is well-defined.
|
||||||
|
if (sweptOnce && now - lastWipSweepNanos < WIP_SWEEP_INTERVAL_NANOS) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
try {
|
||||||
|
int deleted = sessions.sweepWipRefs(WIP_MIN_AGE_MILLIS);
|
||||||
|
if (deleted > 0) {
|
||||||
|
log.info("refs/wip retention sweep deleted {} snapshot ref(s) older than 24h whose "
|
||||||
|
+ "content was already reachable from main", deleted);
|
||||||
|
}
|
||||||
|
} catch (RuntimeException e) {
|
||||||
|
log.warn("refs/wip retention sweep failed; continuing", e);
|
||||||
|
}
|
||||||
|
// Set even when the sweep threw, so a broken repo is retried on the slow cadence rather
|
||||||
|
// than hammering git on every 5-second iteration.
|
||||||
|
lastWipSweepNanos = now;
|
||||||
|
sweptOnce = true;
|
||||||
|
}
|
||||||
|
|
||||||
private void sleep() {
|
private void sleep() {
|
||||||
try {
|
try {
|
||||||
Thread.sleep(intervalMillis);
|
Thread.sleep(intervalMillis);
|
||||||
|
|||||||
@@ -1,6 +1,7 @@
|
|||||||
package dev.ltms.bridged.session;
|
package dev.ltms.bridged.session;
|
||||||
|
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
|
import java.util.Optional;
|
||||||
|
|
||||||
/** Seam between {@link SessionManager} and git worktree operations. Tests use a recording fake. */
|
/** Seam between {@link SessionManager} and git worktree operations. Tests use a recording fake. */
|
||||||
public interface Worktrees {
|
public interface Worktrees {
|
||||||
@@ -27,4 +28,74 @@ public interface Worktrees {
|
|||||||
|
|
||||||
/** git -C <cwd> rev-parse --show-toplevel — the repo root that owns cwd. */
|
/** git -C <cwd> rev-parse --show-toplevel — the repo root that owns cwd. */
|
||||||
String repoRoot(String cwd);
|
String repoRoot(String cwd);
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Commit the worktree's full on-disk state — tracked and untracked, respecting
|
||||||
|
* {@code .gitignore} — to {@code refs/wip/<branch>}, so a release that would otherwise leave
|
||||||
|
* the work as loose, unprotected files has a durable git object to fall back on (CB-578 stage
|
||||||
|
* C). Built on a <em>temporary</em> index: the worker's own index, working tree, and HEAD are
|
||||||
|
* never touched, since the worker may still be mid-write. The commit is parented on the
|
||||||
|
* worktree's current HEAD.
|
||||||
|
*
|
||||||
|
* <p>Never writes under {@code refs/heads/} — the ref must not appear in {@code git branch},
|
||||||
|
* must not be pushed by default, and must not be swept by a later {@code git branch -d}.
|
||||||
|
*
|
||||||
|
* <p>Callers are expected to have already confirmed {@link #hasUncommitted} before reaching
|
||||||
|
* for this; it always stages and commits whatever {@code git add -A} finds, so calling it on
|
||||||
|
* a clean worktree still produces a (harmless, tree-identical-to-HEAD) commit rather than
|
||||||
|
* detecting cleanliness itself.
|
||||||
|
*
|
||||||
|
* @param worktreePath absolute path of the worktree to snapshot
|
||||||
|
* @param branch the worktree's own branch — keys {@code refs/wip/<branch>}
|
||||||
|
* @param message the commit message; should name the member, its branch, and the release
|
||||||
|
* cause so an operator can tell which run produced it
|
||||||
|
* @return the created commit's sha, or {@link Optional#empty()} if {@code worktreePath} does
|
||||||
|
* not exist (mirrors {@link #remove} and {@link #hasUncommitted}'s already-gone
|
||||||
|
* tolerance — a worktree that is gone holds nothing to snapshot)
|
||||||
|
*/
|
||||||
|
Optional<String> snapshot(String worktreePath, String branch, String message);
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586: how many {@code refs/wip/*} snapshot refs exist in {@code repoRoot} and roughly what
|
||||||
|
* they cost. This is the operator-visible surface for the snapshot growth CB-578 stage C left
|
||||||
|
* behind — counts of refs alone hide that each one pins a whole tree for {@code git gc}.
|
||||||
|
*
|
||||||
|
* @param repoRoot the repository to scan
|
||||||
|
* @return count of snapshot refs, and {@code costBytes} = the approximate total working-tree
|
||||||
|
* size of every snapshot's committed content (summed per ref, so shared objects are
|
||||||
|
* counted once per ref that carries them)
|
||||||
|
*/
|
||||||
|
WipRefStats wipRefs(String repoRoot);
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586: run the {@code refs/wip/*} retention sweep and return how many refs it deleted.
|
||||||
|
*
|
||||||
|
* <p>The retention rule is <em>reachability plus an age floor</em>. A snapshot ref is deleted
|
||||||
|
* only when <strong>both</strong> hold:
|
||||||
|
* <ol>
|
||||||
|
* <li>its commit's <em>tree content</em> is already reachable from {@code main} — the work
|
||||||
|
* the snapshot preserved has been recovered, so dropping the ref loses nothing; and</li>
|
||||||
|
* <li>the ref is older than {@code minAgeMillis} — a very recent snapshot is never swept
|
||||||
|
* while a lead may still be looking at it.</li>
|
||||||
|
* </ol>
|
||||||
|
*
|
||||||
|
* <p>Reachability is the safety property. A snapshot exists precisely because the work was not
|
||||||
|
* committed anywhere else, so a snapshot whose content is <em>not</em> reachable from
|
||||||
|
* {@code main} is the <strong>last copy</strong> of a worker's work and must never be deleted
|
||||||
|
* automatically — that is the failure CB-576 and CB-578 stage C were built to stop. Age alone
|
||||||
|
* must never drive a deletion, because age-based sweeping is exactly how the last copy gets
|
||||||
|
* destroyed. (Both numbers and the rule are CB-586's decision; this method only implements it.)
|
||||||
|
*
|
||||||
|
* <p>Every deletion logs the ref name and the commit sha, so an operator who finds they lost
|
||||||
|
* the wrong thing can still recover it from git's reflog.
|
||||||
|
*
|
||||||
|
* @param repoRoot the repository whose {@code refs/wip/*} to sweep
|
||||||
|
* @param minAgeMillis the age floor; a ref younger than this is never touched
|
||||||
|
* @return the number of snapshot refs deleted
|
||||||
|
*/
|
||||||
|
int pruneWipRefs(String repoRoot, long minAgeMillis);
|
||||||
|
|
||||||
|
/** CB-586: the operator-visible census of {@code refs/wip/*} in one repository. */
|
||||||
|
record WipRefStats(int count, long costBytes) {
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,108 @@
|
|||||||
|
package dev.ltms.bridged;
|
||||||
|
|
||||||
|
import ch.qos.logback.classic.Logger;
|
||||||
|
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||||
|
import ch.qos.logback.core.read.ListAppender;
|
||||||
|
import dev.ltms.bridged.config.BridgedConfig;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: an absent (or empty) {@code memberCredentials:} block blocks nothing — no name is
|
||||||
|
* hardcoded any more to fall back on, so the daemon must say so out loud at startup rather than
|
||||||
|
* silently dropping CB-592's protection. Mirrors {@link RequiredSecretEnvVarsTest}'s pattern for
|
||||||
|
* the CB-594 startup-secrets report, capturing the real log via a {@link ListAppender}.
|
||||||
|
*/
|
||||||
|
class MemberCredentialsGapReportTest {
|
||||||
|
|
||||||
|
private static BridgedConfig load(Path dir, String yaml) throws Exception {
|
||||||
|
Path f = dir.resolve("bridged.yaml");
|
||||||
|
Files.writeString(f, yaml);
|
||||||
|
return BridgedConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static ListAppender<ILoggingEvent> attach() {
|
||||||
|
Logger logger = (Logger) LoggerFactory.getLogger(Bridged.class);
|
||||||
|
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||||
|
appender.start();
|
||||||
|
logger.addAppender(appender);
|
||||||
|
return appender;
|
||||||
|
}
|
||||||
|
|
||||||
|
private static void detach(ListAppender<ILoggingEvent> appender) {
|
||||||
|
((Logger) LoggerFactory.getLogger(Bridged.class)).detachAppender(appender);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void anAbsentBlockWarnsThatEveryMemberInheritsTheWholeStore(@TempDir Path dir) throws Exception {
|
||||||
|
BridgedConfig cfg = load(dir, "bind:\n host: 127.0.0.1\n port: 8765\n");
|
||||||
|
|
||||||
|
ListAppender<ILoggingEvent> appender = attach();
|
||||||
|
try {
|
||||||
|
Bridged.reportMemberCredentialsGap(cfg);
|
||||||
|
} finally {
|
||||||
|
detach(appender);
|
||||||
|
}
|
||||||
|
|
||||||
|
assertTrue(appender.list.stream().anyMatch(e ->
|
||||||
|
e.getLevel() == ch.qos.logback.classic.Level.WARN
|
||||||
|
&& e.getFormattedMessage().contains("memberCredentials")
|
||||||
|
&& e.getFormattedMessage().contains("WHOLE secret store")),
|
||||||
|
"an absent block must WARN that protection is lost, not stay silent");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void anEmptyKnownListWarnsTheSameAsAbsent(@TempDir Path dir) throws Exception {
|
||||||
|
BridgedConfig cfg = load(dir, """
|
||||||
|
memberCredentials:
|
||||||
|
policy: deny-by-default
|
||||||
|
""");
|
||||||
|
|
||||||
|
ListAppender<ILoggingEvent> appender = attach();
|
||||||
|
try {
|
||||||
|
Bridged.reportMemberCredentialsGap(cfg);
|
||||||
|
} finally {
|
||||||
|
detach(appender);
|
||||||
|
}
|
||||||
|
|
||||||
|
assertTrue(appender.list.stream().anyMatch(e ->
|
||||||
|
e.getLevel() == ch.qos.logback.classic.Level.WARN
|
||||||
|
&& e.getFormattedMessage().contains("memberCredentials")),
|
||||||
|
"policy: with no known: names still blocks nothing and must warn the same way");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aPopulatedKnownListLogsInfoNotWarn(@TempDir Path dir) throws Exception {
|
||||||
|
BridgedConfig cfg = load(dir, """
|
||||||
|
memberCredentials:
|
||||||
|
policy: deny-by-default
|
||||||
|
allow: [AI_GATEWAY_TOKEN]
|
||||||
|
known: [AI_GATEWAY_TOKEN, GITEA_ACCESS_TOKEN]
|
||||||
|
""");
|
||||||
|
|
||||||
|
// logback-test.xml pins dev.ltms.bridged to WARN (see its own comment); raise it here so
|
||||||
|
// the INFO line this test asserts on actually reaches the appender, and restore after.
|
||||||
|
Logger logger = (Logger) LoggerFactory.getLogger(Bridged.class);
|
||||||
|
ch.qos.logback.classic.Level original = logger.getLevel();
|
||||||
|
logger.setLevel(ch.qos.logback.classic.Level.INFO);
|
||||||
|
ListAppender<ILoggingEvent> appender = attach();
|
||||||
|
try {
|
||||||
|
Bridged.reportMemberCredentialsGap(cfg);
|
||||||
|
} finally {
|
||||||
|
detach(appender);
|
||||||
|
logger.setLevel(original);
|
||||||
|
}
|
||||||
|
|
||||||
|
assertFalse(appender.list.stream().anyMatch(e -> e.getLevel() == ch.qos.logback.classic.Level.WARN),
|
||||||
|
"a configured, non-empty known: list must not warn — the block is doing its job");
|
||||||
|
assertTrue(appender.list.stream().anyMatch(e -> e.getFormattedMessage().contains("blocking 1")),
|
||||||
|
"the INFO line should say how many names are actually blocked (known minus allow)");
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,114 @@
|
|||||||
|
package dev.ltms.bridged;
|
||||||
|
|
||||||
|
import dev.ltms.bridged.config.BridgedConfig;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-594: {@link Bridged#requiredSecretEnvVars(BridgedConfig)} is what decides what the startup
|
||||||
|
* secret report checks — it must derive that set from the config, not a hand-written list, or a
|
||||||
|
* new profile's token silently stops being reported.
|
||||||
|
*/
|
||||||
|
class RequiredSecretEnvVarsTest {
|
||||||
|
|
||||||
|
private static BridgedConfig load(Path dir, String yaml) throws Exception {
|
||||||
|
Path f = dir.resolve("bridged.yaml");
|
||||||
|
Files.writeString(f, yaml);
|
||||||
|
return BridgedConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void collectsATokenEnvPerNonSubscriptionProfile(@TempDir Path dir) throws Exception {
|
||||||
|
BridgedConfig cfg = load(dir, """
|
||||||
|
profiles:
|
||||||
|
local:
|
||||||
|
baseUrl: http://gx00.gw:8000
|
||||||
|
tokenEnv: AI_GATEWAY_TOKEN
|
||||||
|
""");
|
||||||
|
|
||||||
|
Map<String, List<String>> required = Bridged.requiredSecretEnvVars(cfg);
|
||||||
|
|
||||||
|
assertTrue(required.containsKey("AI_GATEWAY_TOKEN"));
|
||||||
|
assertEquals(List.of("profile 'local' tokenEnv"), required.get("AI_GATEWAY_TOKEN"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aSubscriptionProfileNeedsNoTokenEnv(@TempDir Path dir) throws Exception {
|
||||||
|
BridgedConfig cfg = load(dir, """
|
||||||
|
profiles:
|
||||||
|
opus:
|
||||||
|
subscription: true
|
||||||
|
model: claude-opus-5
|
||||||
|
""");
|
||||||
|
|
||||||
|
assertTrue(Bridged.requiredSecretEnvVars(cfg).isEmpty(),
|
||||||
|
"subscription: true never reads ANTHROPIC_AUTH_TOKEN — see Profile#isSubscription");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void gitTokenEnvIsOptInAndCollectedWhenSet(@TempDir Path dir) throws Exception {
|
||||||
|
BridgedConfig cfg = load(dir, """
|
||||||
|
profiles:
|
||||||
|
local:
|
||||||
|
baseUrl: http://gx00.gw:8000
|
||||||
|
tokenEnv: AI_GATEWAY_TOKEN
|
||||||
|
gitTokenEnv: WORKER_GITEA_TOKEN
|
||||||
|
""");
|
||||||
|
|
||||||
|
Map<String, List<String>> required = Bridged.requiredSecretEnvVars(cfg);
|
||||||
|
|
||||||
|
assertTrue(required.containsKey("WORKER_GITEA_TOKEN"));
|
||||||
|
assertEquals(List.of("profile 'local' gitTokenEnv"), required.get("WORKER_GITEA_TOKEN"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void noGitTokenEnvMeansNothingIsRequiredForIt(@TempDir Path dir) throws Exception {
|
||||||
|
BridgedConfig cfg = load(dir, """
|
||||||
|
profiles:
|
||||||
|
local:
|
||||||
|
baseUrl: http://gx00.gw:8000
|
||||||
|
tokenEnv: AI_GATEWAY_TOKEN
|
||||||
|
""");
|
||||||
|
|
||||||
|
assertFalse(Bridged.requiredSecretEnvVars(cfg).containsKey("WORKER_GITEA_TOKEN"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aVarSharedByTwoProfilesIsReportedOnceNamingBoth(@TempDir Path dir) throws Exception {
|
||||||
|
BridgedConfig cfg = load(dir, """
|
||||||
|
profiles:
|
||||||
|
local:
|
||||||
|
baseUrl: http://gx00.gw:8000
|
||||||
|
tokenEnv: AI_GATEWAY_TOKEN
|
||||||
|
gitTokenEnv: WORKER_GITEA_TOKEN
|
||||||
|
gx:
|
||||||
|
kind: opencode
|
||||||
|
baseUrl: https://llm.ltms.dev/v1
|
||||||
|
tokenEnv: AI_GATEWAY_TOKEN
|
||||||
|
gitTokenEnv: WORKER_GITEA_TOKEN
|
||||||
|
""");
|
||||||
|
|
||||||
|
Map<String, List<String>> required = Bridged.requiredSecretEnvVars(cfg);
|
||||||
|
|
||||||
|
assertEquals(List.of("profile 'local' tokenEnv", "profile 'gx' tokenEnv"),
|
||||||
|
required.get("AI_GATEWAY_TOKEN"));
|
||||||
|
assertEquals(List.of("profile 'local' gitTokenEnv", "profile 'gx' gitTokenEnv"),
|
||||||
|
required.get("WORKER_GITEA_TOKEN"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void noProfilesMeansNothingIsRequired(@TempDir Path dir) throws Exception {
|
||||||
|
BridgedConfig cfg = load(dir, "bind:\n host: 127.0.0.1\n port: 8765\n");
|
||||||
|
|
||||||
|
assertTrue(Bridged.requiredSecretEnvVars(cfg).isEmpty());
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -10,6 +10,7 @@ import java.nio.file.Path;
|
|||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
import java.util.Set;
|
import java.util.Set;
|
||||||
|
import java.util.regex.Pattern;
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.*;
|
import static org.junit.jupiter.api.Assertions.*;
|
||||||
|
|
||||||
@@ -1039,6 +1040,44 @@ class BridgedConfigTest {
|
|||||||
"an opencode worker with no argv defaults to the opencode binary, never claude");
|
"an opencode worker with no argv defaults to the opencode binary, never claude");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-604: an unrecognized {@code kind:} used to silently fall into the claude-code bucket — not
|
||||||
|
* matching {@code "opencode"} was the only check. With {@code argv:} also unset that meant the
|
||||||
|
* daemon tried to launch a program literally named after the typo.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void unknownKindIsRefusedAtLoadNamingTheValueAndTheAcceptedSet(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("kind-typo.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
profiles:
|
||||||
|
gemini:
|
||||||
|
kind: opencod
|
||||||
|
model: google/gemini-2.5-pro
|
||||||
|
""");
|
||||||
|
|
||||||
|
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
|
||||||
|
assertTrue(e.getMessage().contains("gemini"), "error names the profile: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("opencod"), "error names the bad value: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("claude-code") && e.getMessage().contains("opencode"),
|
||||||
|
"error names the accepted set: " + e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void blankKindStillDefaultsToClaudeCode(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("kind-blank.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
profiles:
|
||||||
|
gx10:
|
||||||
|
baseUrl: http://gx10.gw:8000
|
||||||
|
kind: ""
|
||||||
|
argv: ["ccs", "gx10"]
|
||||||
|
""");
|
||||||
|
|
||||||
|
BridgedConfig cfg = BridgedConfig.load(f);
|
||||||
|
assertEquals(BridgedConfig.Profile.KIND_CLAUDE_CODE, cfg.profiles().get("gx10").kind(),
|
||||||
|
"a blank kind: is documented to behave exactly like an absent one");
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void authDefaultsToLoopbackTrustSoExistingConfigsBehaveAsBefore(@TempDir Path dir) throws Exception {
|
void authDefaultsToLoopbackTrustSoExistingConfigsBehaveAsBefore(@TempDir Path dir) throws Exception {
|
||||||
Path f = dir.resolve("no-auth-block.yaml");
|
Path f = dir.resolve("no-auth-block.yaml");
|
||||||
@@ -1157,6 +1196,10 @@ class BridgedConfigTest {
|
|||||||
idleAfterSeconds: 600
|
idleAfterSeconds: 600
|
||||||
backoffMs: 45000
|
backoffMs: 45000
|
||||||
quietNudgeCap: 5
|
quietNudgeCap: 5
|
||||||
|
memberCredentials:
|
||||||
|
policy: deny-by-default
|
||||||
|
allow: [AI_GATEWAY_TOKEN]
|
||||||
|
known: [AI_GATEWAY_TOKEN, GITEA_ACCESS_TOKEN]
|
||||||
""");
|
""");
|
||||||
|
|
||||||
BridgedConfig cfg = BridgedConfig.load(f);
|
BridgedConfig cfg = BridgedConfig.load(f);
|
||||||
@@ -1187,6 +1230,84 @@ class BridgedConfigTest {
|
|||||||
assertEquals(600, cfg.leadHeartbeat().idleAfterSeconds(), "leadHeartbeat binds at the top level");
|
assertEquals(600, cfg.leadHeartbeat().idleAfterSeconds(), "leadHeartbeat binds at the top level");
|
||||||
assertEquals(45_000L, cfg.leadHeartbeat().backoffMs());
|
assertEquals(45_000L, cfg.leadHeartbeat().backoffMs());
|
||||||
assertEquals(5, cfg.leadHeartbeat().quietNudgeCap());
|
assertEquals(5, cfg.leadHeartbeat().quietNudgeCap());
|
||||||
|
|
||||||
|
assertEquals(BridgedConfig.MemberCredentials.POLICY_DENY_BY_DEFAULT, cfg.memberCredentials().policy(),
|
||||||
|
"memberCredentials binds at the top level");
|
||||||
|
assertEquals(Set.of("AI_GATEWAY_TOKEN"), cfg.memberCredentials().allowSet());
|
||||||
|
assertEquals(Set.of("GITEA_ACCESS_TOKEN"), cfg.memberCredentials().blockedSet());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* A top-level key {@code BridgedConfig} reads but that appears nowhere in
|
||||||
|
* {@code bridged.example.yaml} — live or commented — is invisible drift: {@code bridged.yaml}
|
||||||
|
* is gitignored, so the example is the ONLY committed description of the config schema, and
|
||||||
|
* neither {@link #shippedExampleConfigParses} (example → code: does the example still parse)
|
||||||
|
* nor {@link #everyOptionalKnobDocumentedInTheExampleBinds} (a hand-maintained list of keys
|
||||||
|
* that must bind) can catch a brand-new key nobody added to either.
|
||||||
|
*
|
||||||
|
* <p>This test compares the OTHER direction: every key in {@link BridgedConfig#KNOWN_TOP_LEVEL_KEYS}
|
||||||
|
* (the parser's own accepted set, which backs the unknown-key WARN) must appear as a top-level
|
||||||
|
* key in the example text, live or commented-out — see {@link #topLevelKeyDocumented}.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void everyKnownTopLevelKeyIsDocumentedInTheExample() throws Exception {
|
||||||
|
Path example = Path.of("bridged.example.yaml");
|
||||||
|
assertTrue(Files.exists(example), "bridged.example.yaml must ship next to the pom");
|
||||||
|
String text = Files.readString(example);
|
||||||
|
|
||||||
|
List<String> undocumented = BridgedConfig.KNOWN_TOP_LEVEL_KEYS.stream()
|
||||||
|
.filter(key -> !topLevelKeyDocumented(text, key))
|
||||||
|
.sorted()
|
||||||
|
.toList();
|
||||||
|
|
||||||
|
assertTrue(undocumented.isEmpty(), () -> "key(s) " + undocumented
|
||||||
|
+ " are read by BridgedConfig but appear nowhere in bridged.example.yaml — "
|
||||||
|
+ "document each one there, commented out if optional. bridged.yaml is "
|
||||||
|
+ "gitignored, so this file is the only committed description of the config "
|
||||||
|
+ "schema an operator or a worker can see.");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Most of {@code bridged.example.yaml} is deliberately commented out — optional sections are
|
||||||
|
* documented as commented blocks so the shipped file stays a working minimal config. A key
|
||||||
|
* documented ONLY as a comment must still count as documented; parsing the file as YAML and
|
||||||
|
* reading its live key set (as an earlier attempt at this guard did) gets this wrong, because
|
||||||
|
* every commented section then looks entirely absent.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void commentedOnlyTopLevelKeyCountsAsDocumented() {
|
||||||
|
String yaml = """
|
||||||
|
bind:
|
||||||
|
port: 8765
|
||||||
|
# broker:
|
||||||
|
# uri: amqp://guest:guest@127.0.0.1:5672
|
||||||
|
""";
|
||||||
|
assertTrue(topLevelKeyDocumented(yaml, "broker"),
|
||||||
|
"a key documented only inside a commented-out block must still count as documented");
|
||||||
|
}
|
||||||
|
|
||||||
|
/** A key that appears in neither a live nor a commented top-level line must NOT count. */
|
||||||
|
@Test
|
||||||
|
void absentTopLevelKeyIsNotDocumented() {
|
||||||
|
String yaml = """
|
||||||
|
bind:
|
||||||
|
port: 8765
|
||||||
|
""";
|
||||||
|
assertFalse(topLevelKeyDocumented(yaml, "broker"),
|
||||||
|
"a key mentioned nowhere in the example must not be reported as documented");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* True when {@code key} appears as a top-level YAML key in {@code yaml} — either live
|
||||||
|
* ({@code key:} at column 0) or commented out ({@code # key:}, also at column 0, with only
|
||||||
|
* whitespace between the {@code #} and the key). Anchoring on column 0 is what keeps this a
|
||||||
|
* top-level check: an indented occurrence (a nested field, or prose inside a comment that
|
||||||
|
* happens to end in a colon) never matches, because {@code ^} requires the key's own first
|
||||||
|
* character — or the sole leading {@code #} — to sit at the very start of the line.
|
||||||
|
*/
|
||||||
|
private static boolean topLevelKeyDocumented(String yaml, String key) {
|
||||||
|
Pattern p = Pattern.compile("(?m)^(?:#\\s*)?" + Pattern.quote(key) + ":");
|
||||||
|
return p.matcher(yaml).find();
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -1221,6 +1342,232 @@ class BridgedConfigTest {
|
|||||||
assertNull(w.maxLoad(), "absent maxLoad defaults to unlimited (null)");
|
assertNull(w.maxLoad(), "absent maxLoad defaults to unlimited (null)");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void explicitWeightZeroStaysZeroInsteadOfCoercingToOne(@TempDir Path dir) throws Exception {
|
||||||
|
// CB-554: weight: 0 used to be normalised to 1.0 by this same compact constructor, which
|
||||||
|
// made it read as "never auto-select" while behaving as "select like anyone else."
|
||||||
|
Path f = dir.resolve("weight-zero.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
port: 8080
|
||||||
|
profiles:
|
||||||
|
opus:
|
||||||
|
baseUrl: http://gx10.gw:8000
|
||||||
|
weight: 0
|
||||||
|
""");
|
||||||
|
|
||||||
|
BridgedConfig cfg = BridgedConfig.load(f);
|
||||||
|
BridgedConfig.Profile w = cfg.profiles().get("opus");
|
||||||
|
assertEquals(0.0f, w.weight(), 0.0001f, "an explicit weight: 0 must stay 0, not coerce to 1.0");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void negativeWeightNormalisesToZeroNotOne(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("weight-negative.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
port: 8080
|
||||||
|
profiles:
|
||||||
|
gx10:
|
||||||
|
baseUrl: http://gx10.gw:8000
|
||||||
|
weight: -3.0
|
||||||
|
""");
|
||||||
|
|
||||||
|
BridgedConfig cfg = BridgedConfig.load(f);
|
||||||
|
BridgedConfig.Profile w = cfg.profiles().get("gx10");
|
||||||
|
assertEquals(0.0f, w.weight(), 0.0001f, "a negative weight behaves as excluded (0), not as an error and not as 1.0");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void explicitMaxLoadZeroStaysZeroInsteadOfCoercingToUnlimited(@TempDir Path dir) throws Exception {
|
||||||
|
// CB-585: maxLoad: 0 used to be normalised to null (unlimited) by this same compact
|
||||||
|
// constructor, so an operator writing it to mean "never run anything here" got the exact
|
||||||
|
// opposite. On a subscription: true profile, maxLoad is the only throttle against the
|
||||||
|
// operator's own paid plan.
|
||||||
|
Path f = dir.resolve("maxload-zero.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
port: 8080
|
||||||
|
profiles:
|
||||||
|
opus:
|
||||||
|
baseUrl: http://gx10.gw:8000
|
||||||
|
maxLoad: 0
|
||||||
|
""");
|
||||||
|
|
||||||
|
BridgedConfig cfg = BridgedConfig.load(f);
|
||||||
|
BridgedConfig.Profile w = cfg.profiles().get("opus");
|
||||||
|
assertEquals(0, w.maxLoad(), "an explicit maxLoad: 0 must stay 0, not coerce to unlimited (null)");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void negativeMaxLoadIsRefusedAtLoadNamingTheProfileAndKey(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("maxload-negative.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
port: 8080
|
||||||
|
profiles:
|
||||||
|
opus:
|
||||||
|
baseUrl: http://gx10.gw:8000
|
||||||
|
maxLoad: -2
|
||||||
|
""");
|
||||||
|
|
||||||
|
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
|
||||||
|
assertTrue(e.getMessage().contains("opus"), "error names the profile: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("maxLoad"), "error names the key: " + e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-606: an unrecognized {@code auth.mode} used to silently fall back to
|
||||||
|
* {@code loopback-trust} — {@link BridgedConfig.Auth#tokenMode()} only checked equality
|
||||||
|
* against {@code "token"}. On a loopback bind {@link BridgedConfig#validateAuthExposure()}
|
||||||
|
* never runs (it only fires for a non-loopback bind), so the typo was completely invisible:
|
||||||
|
* the daemon started cleanly and authenticated nobody while the operator believed token mode
|
||||||
|
* was active.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void unknownAuthModeIsRefusedAtLoadNamingTheValueAndTheAcceptedSet(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("auth-mode-typo.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
auth:
|
||||||
|
mode: toekn
|
||||||
|
""");
|
||||||
|
|
||||||
|
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
|
||||||
|
assertTrue(e.getMessage().contains("toekn"), "error names the bad value: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("loopback-trust") && e.getMessage().contains("token"),
|
||||||
|
"error names the accepted set: " + e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-606: an unrecognized per-profile {@code placement:} used to silently fall back to legacy
|
||||||
|
* pane placement — {@link BridgedConfig.Profile#tabPlacement()} only checked equality against
|
||||||
|
* {@code "tab"}.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void unknownProfilePlacementIsRefusedAtLoadNamingTheProfileAndTheAcceptedSet(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("placement-typo.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
profiles:
|
||||||
|
gx10:
|
||||||
|
baseUrl: http://gx10.gw:8000
|
||||||
|
placement: tabb
|
||||||
|
""");
|
||||||
|
|
||||||
|
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
|
||||||
|
assertTrue(e.getMessage().contains("gx10"), "error names the profile: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("tabb"), "error names the bad value: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("tab") && e.getMessage().contains("pane"),
|
||||||
|
"error names the accepted set: " + e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: {@code memberCredentials.policy} is validated the same way {@code auth.mode} and
|
||||||
|
* per-profile {@code placement} are (CB-606's pattern) — a typo must not silently behave as the
|
||||||
|
* one real policy, because the day a second policy exists that silent fallback becomes a real
|
||||||
|
* behavior change instead of a happy accident.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void unknownMemberCredentialsPolicyIsRefusedAtLoadNamingTheValueAndTheAcceptedSet(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("member-credentials-policy-typo.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
memberCredentials:
|
||||||
|
policy: deny-by-defualt
|
||||||
|
known:
|
||||||
|
- GITEA_ACCESS_TOKEN
|
||||||
|
""");
|
||||||
|
|
||||||
|
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
|
||||||
|
assertTrue(e.getMessage().contains("deny-by-defualt"), "error names the bad value: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("deny-by-default"), "error names the accepted set: " + e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
/** {@code allow}/{@code known} bind and {@link BridgedConfig.MemberCredentials#blockedSet()} is known minus allow. */
|
||||||
|
@Test
|
||||||
|
void memberCredentialsBindsAllowAndKnownAndComputesBlockedSet(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("member-credentials.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
port: 8080
|
||||||
|
memberCredentials:
|
||||||
|
policy: deny-by-default
|
||||||
|
allow:
|
||||||
|
- AI_GATEWAY_TOKEN
|
||||||
|
- WORKER_GITEA_TOKEN
|
||||||
|
known:
|
||||||
|
- AI_GATEWAY_TOKEN
|
||||||
|
- WORKER_GITEA_TOKEN
|
||||||
|
- GITEA_ACCESS_TOKEN
|
||||||
|
- GITLAB_PERSONAL_ACCESS_TOKEN
|
||||||
|
""");
|
||||||
|
|
||||||
|
BridgedConfig.MemberCredentials mc = BridgedConfig.load(f).memberCredentials();
|
||||||
|
assertEquals(BridgedConfig.MemberCredentials.POLICY_DENY_BY_DEFAULT, mc.policy());
|
||||||
|
assertEquals(Set.of("AI_GATEWAY_TOKEN", "WORKER_GITEA_TOKEN"), mc.allowSet());
|
||||||
|
assertEquals(Set.of("GITEA_ACCESS_TOKEN", "GITLAB_PERSONAL_ACCESS_TOKEN"), mc.blockedSet(),
|
||||||
|
"blockedSet is known minus allow");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: omitting {@code memberCredentials:} entirely must NOT crash a reader that assumes a
|
||||||
|
* non-null block (the same "fill in nested defaults" contract every other structural field
|
||||||
|
* gets — see {@link BridgedConfig#withDefaults()}), but it also must not pretend anything is
|
||||||
|
* blocked: an empty {@code known} list blocks nothing, and that is a real gap the operator must
|
||||||
|
* close by configuring this block, not a safe default.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void absentMemberCredentialsDefaultsToAnEmptyNonNullBlock(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("member-credentials-absent.yaml");
|
||||||
|
Files.writeString(f, "bind:\n port: 8080\n");
|
||||||
|
|
||||||
|
BridgedConfig.MemberCredentials mc = BridgedConfig.load(f).memberCredentials();
|
||||||
|
assertNotNull(mc, "withDefaults() must never leave this null");
|
||||||
|
assertTrue(mc.known().isEmpty(), "no known list configured — nothing is blocked");
|
||||||
|
assertTrue(mc.allowSet().isEmpty());
|
||||||
|
assertTrue(mc.blockedSet().isEmpty());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void absentProfilePlacementDefaultsToTab(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("placement-absent.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
profiles:
|
||||||
|
gx10:
|
||||||
|
baseUrl: http://gx10.gw:8000
|
||||||
|
""");
|
||||||
|
|
||||||
|
BridgedConfig cfg = BridgedConfig.load(f);
|
||||||
|
BridgedConfig.Profile w = cfg.profiles().get("gx10");
|
||||||
|
assertTrue(w.tabPlacement(), "an absent placement must keep defaulting to tab");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-606: the top-level {@code placement:} policy name WAS validated, but only lazily, by
|
||||||
|
* {@code PlacementPolicies.fromName} through {@code CompositePeerLauncher}'s per-spawn
|
||||||
|
* {@code Supplier} — so a bad name still started a daemon that looked healthy and failed only
|
||||||
|
* the first time something spawned without naming a profile. This must now fail at load.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void unknownTopLevelPlacementPolicyIsRefusedAtLoadNotLazilyAtFirstSpawn(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("placement-policy-typo.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
port: 8080
|
||||||
|
placement: weightd
|
||||||
|
""");
|
||||||
|
|
||||||
|
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
|
||||||
|
assertTrue(e.getMessage().contains("weightd"), "error names the bad value: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("fixed") && e.getMessage().contains("round-robin")
|
||||||
|
&& e.getMessage().contains("weighted"),
|
||||||
|
"error names the accepted set: " + e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void subscriptionFlagBindsAndDefaultsFalse(@TempDir Path dir) throws Exception {
|
void subscriptionFlagBindsAndDefaultsFalse(@TempDir Path dir) throws Exception {
|
||||||
Path f = dir.resolve("subscription.yaml");
|
Path f = dir.resolve("subscription.yaml");
|
||||||
|
|||||||
@@ -338,6 +338,54 @@ class ConfigRefTest {
|
|||||||
assertEquals("deepseek-v4-flash", ref.get().profiles().get("sonnet").model());
|
assertEquals("deepseek-v4-flash", ref.get().profiles().get("sonnet").model());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-578 stage B: exhaustedPattern is compiled once into Bridged.main's pattern map at startup
|
||||||
|
* (see ExhaustedPatternLookup), so a reload never re-reads it — a changed pattern must be
|
||||||
|
* reported deferred exactly like model/baseUrl, not silently claimed as applied.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void changingAProfilesExhaustedPatternIsReportedAsDeferred(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("bridged.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
herdrSocket: ~/.config/herdr/herdr.sock
|
||||||
|
profiles:
|
||||||
|
sonnet:
|
||||||
|
baseUrl: http://gx00.gw:8000
|
||||||
|
model: sonnet
|
||||||
|
exhaustedPattern: "usage limit has been reached"
|
||||||
|
guard:
|
||||||
|
offSubscriptionHosts:
|
||||||
|
- gx00.gw
|
||||||
|
""");
|
||||||
|
ConfigRef ref = refFor(f);
|
||||||
|
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
herdrSocket: ~/.config/herdr/herdr.sock
|
||||||
|
profiles:
|
||||||
|
sonnet:
|
||||||
|
baseUrl: http://gx00.gw:8000
|
||||||
|
model: sonnet
|
||||||
|
exhaustedPattern: "rate limit exceeded"
|
||||||
|
guard:
|
||||||
|
offSubscriptionHosts:
|
||||||
|
- gx00.gw
|
||||||
|
""");
|
||||||
|
ConfigRef.Outcome out = ref.reload();
|
||||||
|
|
||||||
|
assertTrue(out.applied());
|
||||||
|
assertEquals(1, out.deferred().size(), out.deferred().toString());
|
||||||
|
assertTrue(out.deferred().getFirst().contains("sonnet"), out.deferred().toString());
|
||||||
|
assertTrue(out.deferred().getFirst().contains("launch settings"), out.deferred().toString());
|
||||||
|
// The snapshot still carries the new value — a restart is what makes it take effect.
|
||||||
|
assertEquals("rate limit exceeded", ref.get().profiles().get("sonnet").exhaustedPattern());
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void aFixedRefHasNoFileAndRefusesToReload() {
|
void aFixedRefHasNoFileAndRefusesToReload() {
|
||||||
BridgedConfig cfg = new BridgedConfig(null, null, null, null, null, null,
|
BridgedConfig cfg = new BridgedConfig(null, null, null, null, null, null,
|
||||||
|
|||||||
@@ -7,6 +7,7 @@ import java.util.ArrayList;
|
|||||||
import java.util.LinkedHashMap;
|
import java.util.LinkedHashMap;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.CopyOnWriteArrayList;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Recording fake {@link HerdrClient} for unit/acceptance tests. Returns canned frames
|
* Recording fake {@link HerdrClient} for unit/acceptance tests. Returns canned frames
|
||||||
@@ -22,7 +23,13 @@ public final class FakeHerdr implements HerdrClient {
|
|||||||
public static final long WORKER_PID = 4242;
|
public static final long WORKER_PID = 4242;
|
||||||
|
|
||||||
private final ObjectMapper mapper = new ObjectMapper();
|
private final ObjectMapper mapper = new ObjectMapper();
|
||||||
public final List<Call> calls = new ArrayList<>();
|
/**
|
||||||
|
* Thread-safe on purpose. Background loops — {@link dev.ltms.bridged.msg.ReplyPushLoop} and the
|
||||||
|
* lead heartbeat — call this fake from their own scheduler threads while a test polls
|
||||||
|
* {@link #called} from the test thread. A plain {@code ArrayList} threw
|
||||||
|
* {@code ConcurrentModificationException} out of {@code called()} when a nudge landed mid-stream.
|
||||||
|
*/
|
||||||
|
public final List<Call> calls = new CopyOnWriteArrayList<>();
|
||||||
private boolean healthy = true;
|
private boolean healthy = true;
|
||||||
private final List<String> extraWorkspaces = new ArrayList<>();
|
private final List<String> extraWorkspaces = new ArrayList<>();
|
||||||
private final List<String> extraAgents = new ArrayList<>();
|
private final List<String> extraAgents = new ArrayList<>();
|
||||||
|
|||||||
@@ -26,7 +26,7 @@ class CompletionResolverTest {
|
|||||||
void skipsTheScrapeWhenNoSendIsWaiting() {
|
void skipsTheScrapeWhenNoSendIsWaiting() {
|
||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
|
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
|
|
||||||
resolver.resolve("term_a", null); // no in-flight turn captured for this target
|
resolver.resolve("term_a", null); // no in-flight turn captured for this target
|
||||||
|
|
||||||
@@ -38,7 +38,7 @@ class CompletionResolverTest {
|
|||||||
void failSkipsTheScrapeWhenNoSendIsWaiting() {
|
void failSkipsTheScrapeWhenNoSendIsWaiting() {
|
||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
|
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
|
|
||||||
resolver.fail("term_a", null); // no in-flight turn, and no registered waiter to fall back to
|
resolver.fail("term_a", null); // no in-flight turn, and no registered waiter to fall back to
|
||||||
|
|
||||||
@@ -50,7 +50,7 @@ class CompletionResolverTest {
|
|||||||
void captureBaselineSkipsTheReadWhenNoSendIsWaiting() {
|
void captureBaselineSkipsTheReadWhenNoSendIsWaiting() {
|
||||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ X\n❯ ");
|
FakeHerdr herdr = new FakeHerdr().readText("⏺ X\n❯ ");
|
||||||
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
|
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
|
|
||||||
resolver.captureBaseline("term_a", TestTurnTokens.inert("term_a")); // no send to attribute a later completion to
|
resolver.captureBaseline("term_a", TestTurnTokens.inert("term_a")); // no send to attribute a later completion to
|
||||||
|
|
||||||
@@ -139,7 +139,7 @@ class CompletionResolverTest {
|
|||||||
// send must NOT be resolved with the stale answer.
|
// send must NOT be resolved with the stale answer.
|
||||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ 391\n❯ ");
|
FakeHerdr herdr = new FakeHerdr().readText("⏺ 391\n❯ ");
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
|
|
||||||
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
|
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
|
||||||
// The turn as captured at delivery: its waiter, and the previous turn's answer still on screen.
|
// The turn as captured at delivery: its waiter, and the previous turn's answer still on screen.
|
||||||
@@ -154,7 +154,7 @@ class CompletionResolverTest {
|
|||||||
void resolvesACompletionWhoseScrapeChangedSinceDelivery() {
|
void resolvesACompletionWhoseScrapeChangedSinceDelivery() {
|
||||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ No, 391 = 17 × 23.\n❯ "); // the worker's real answer
|
FakeHerdr herdr = new FakeHerdr().readText("⏺ No, 391 = 17 × 23.\n❯ "); // the worker's real answer
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
|
|
||||||
var waiter = rendezvous.open("term_a");
|
var waiter = rendezvous.open("term_a");
|
||||||
// Delivery baseline was the previous turn's "391"; the scrape now differs → resolve.
|
// Delivery baseline was the previous turn's "391"; the scrape now differs → resolve.
|
||||||
@@ -171,7 +171,7 @@ class CompletionResolverTest {
|
|||||||
String block = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 1) + "\n❯ ";
|
String block = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 1) + "\n❯ ";
|
||||||
FakeHerdr herdr = new FakeHerdr().readText(block);
|
FakeHerdr herdr = new FakeHerdr().readText(block);
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
|
|
||||||
var waiter = rendezvous.open("term_a");
|
var waiter = rendezvous.open("term_a");
|
||||||
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
||||||
@@ -185,7 +185,7 @@ class CompletionResolverTest {
|
|||||||
void leavesAnUnclippedCompletionPaneTailUnmarked() {
|
void leavesAnUnclippedCompletionPaneTailUnmarked() {
|
||||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ complete report\n❯ ");
|
FakeHerdr herdr = new FakeHerdr().readText("⏺ complete report\n❯ ");
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
|
|
||||||
var waiter = rendezvous.open("term_a");
|
var waiter = rendezvous.open("term_a");
|
||||||
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
||||||
@@ -197,7 +197,7 @@ class CompletionResolverTest {
|
|||||||
void resolvesSynchronouslyBeforePostTurnContextClearing() {
|
void resolvesSynchronouslyBeforePostTurnContextClearing() {
|
||||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ previous answer\n❯ ");
|
FakeHerdr herdr = new FakeHerdr().readText("⏺ previous answer\n❯ ");
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
var waiter = rendezvous.open("term_a");
|
var waiter = rendezvous.open("term_a");
|
||||||
resolver.captureBaseline("term_a", new TurnToken("term_a", waiter));
|
resolver.captureBaseline("term_a", new TurnToken("term_a", waiter));
|
||||||
herdr.readText("⏺ answer that /clear would erase\n❯ ");
|
herdr.readText("⏺ answer that /clear would erase\n❯ ");
|
||||||
@@ -219,7 +219,7 @@ class CompletionResolverTest {
|
|||||||
String longBlock = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 500) + "\n❯ ";
|
String longBlock = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 500) + "\n❯ ";
|
||||||
FakeHerdr herdr = new FakeHerdr().readText(longBlock);
|
FakeHerdr herdr = new FakeHerdr().readText(longBlock);
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
|
|
||||||
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
|
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
|
||||||
resolver.captureBaseline("term_a", new TurnToken("term_a", waiter)); // baseline is the clipped >cap block
|
resolver.captureBaseline("term_a", new TurnToken("term_a", waiter)); // baseline is the clipped >cap block
|
||||||
@@ -239,7 +239,7 @@ class CompletionResolverTest {
|
|||||||
// No delivery baseline (e.g. the pre-turn read failed) ⇒ never suppress; the completion resolves.
|
// No delivery baseline (e.g. the pre-turn read failed) ⇒ never suppress; the completion resolves.
|
||||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ hello\n❯ ");
|
FakeHerdr herdr = new FakeHerdr().readText("⏺ hello\n❯ ");
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
|
|
||||||
var waiter = rendezvous.open("term_a");
|
var waiter = rendezvous.open("term_a");
|
||||||
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
||||||
@@ -257,7 +257,7 @@ class CompletionResolverTest {
|
|||||||
// byte-identical guard would wrongly match the empty tail and suppress.
|
// byte-identical guard would wrongly match the empty tail and suppress.
|
||||||
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.read throws HerdrException
|
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.read throws HerdrException
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
|
|
||||||
var waiter = rendezvous.open("term_a");
|
var waiter = rendezvous.open("term_a");
|
||||||
var turn = new CompletionResolver.InFlight(waiter, ""); // empty pane baselined at delivery
|
var turn = new CompletionResolver.InFlight(waiter, ""); // empty pane baselined at delivery
|
||||||
@@ -277,7 +277,7 @@ class CompletionResolverTest {
|
|||||||
// fail must not overwrite that value, and must not even scrape the worker — nobody needs it.
|
// fail must not overwrite that value, and must not even scrape the worker — nobody needs it.
|
||||||
FakeHerdr herdr = new FakeHerdr().readText("an error screen");
|
FakeHerdr herdr = new FakeHerdr().readText("an error screen");
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
|
|
||||||
var waiter = rendezvous.open("term_a");
|
var waiter = rendezvous.open("term_a");
|
||||||
var turn = new CompletionResolver.InFlight(waiter, null);
|
var turn = new CompletionResolver.InFlight(waiter, null);
|
||||||
@@ -298,7 +298,7 @@ class CompletionResolverTest {
|
|||||||
// fail falls back to the waiter currently registered on the Rendezvous and fails it.
|
// fail falls back to the waiter currently registered on the Rendezvous and fails it.
|
||||||
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
|
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
|
|
||||||
var waiter = rendezvous.open("term_a"); // send registered, but no captureBaseline ever ran
|
var waiter = rendezvous.open("term_a"); // send registered, but no captureBaseline ever ran
|
||||||
resolver.fail("term_a", null); // no in-flight turn → fall back to the registered waiter
|
resolver.fail("term_a", null); // no in-flight turn → fall back to the registered waiter
|
||||||
@@ -324,7 +324,7 @@ class CompletionResolverTest {
|
|||||||
try {
|
try {
|
||||||
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
|
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
var waiter = rendezvous.open("term_a");
|
var waiter = rendezvous.open("term_a");
|
||||||
|
|
||||||
resolver.fail("term_a", null);
|
resolver.fail("term_a", null);
|
||||||
@@ -352,7 +352,7 @@ class CompletionResolverTest {
|
|||||||
// scrape to turn N+1; targeting turn N's captured waiter makes the late completion a no-op.
|
// scrape to turn N+1; targeting turn N's captured waiter makes the late completion a no-op.
|
||||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ turn N answer\n❯ ");
|
FakeHerdr herdr = new FakeHerdr().readText("⏺ turn N answer\n❯ ");
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
|
|
||||||
var waiterN = rendezvous.open("term_a"); // turn N's send
|
var waiterN = rendezvous.open("term_a"); // turn N's send
|
||||||
// The turn as the injector captured it at delivery (waiter + pre-turn baseline).
|
// The turn as the injector captured it at delivery (waiter + pre-turn baseline).
|
||||||
@@ -382,7 +382,7 @@ class CompletionResolverTest {
|
|||||||
FakeHerdr herdr = new FakeHerdr().readText(block);
|
FakeHerdr herdr = new FakeHerdr().readText(block);
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
|
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns);
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, ExhaustionSink.none());
|
||||||
|
|
||||||
var waiter = rendezvous.open("term_a");
|
var waiter = rendezvous.open("term_a");
|
||||||
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
||||||
@@ -398,7 +398,7 @@ class CompletionResolverTest {
|
|||||||
FakeHerdr herdr = new FakeHerdr().readText(block);
|
FakeHerdr herdr = new FakeHerdr().readText(block);
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
|
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns);
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, ExhaustionSink.none());
|
||||||
|
|
||||||
var waiter = rendezvous.open("term_a");
|
var waiter = rendezvous.open("term_a");
|
||||||
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
||||||
@@ -407,12 +407,53 @@ class CompletionResolverTest {
|
|||||||
waiter.getNow(null).text(), "the reason names the real cause and carries the matched line");
|
waiter.getNow(null).text(), "the reason names the real cause and carries the matched line");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aWinningBackendExhaustedClassificationNotifiesTheExhaustionSink() {
|
||||||
|
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
|
||||||
|
FakeHerdr herdr = new FakeHerdr().readText(block);
|
||||||
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
|
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
|
||||||
|
java.util.List<String> notified = new java.util.ArrayList<>();
|
||||||
|
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
|
||||||
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
|
||||||
|
|
||||||
|
var waiter = rendezvous.open("term_a");
|
||||||
|
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
||||||
|
|
||||||
|
assertEquals(1, notified.size(), "the sink is notified exactly once for the winning classification");
|
||||||
|
assertTrue(notified.get(0).startsWith("term_a: "), "the sink is told which target exhausted");
|
||||||
|
assertTrue(notified.get(0).contains("The usage limit has been reached"),
|
||||||
|
"the sink is told the matched reason: " + notified.get(0));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aLosingBackendExhaustedClassificationNeverNotifiesTheExhaustionSink() {
|
||||||
|
// The waiter was already resolved (e.g. by the worker's own reply) before this scrape landed —
|
||||||
|
// resolveExhausted loses the race and must return false, so the sink must not fire either.
|
||||||
|
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
|
||||||
|
FakeHerdr herdr = new FakeHerdr().readText(block);
|
||||||
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
|
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
|
||||||
|
java.util.List<String> notified = new java.util.ArrayList<>();
|
||||||
|
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
|
||||||
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
|
||||||
|
|
||||||
|
var waiter = rendezvous.open("term_a");
|
||||||
|
var turn = new CompletionResolver.InFlight(waiter, null);
|
||||||
|
assertTrue(rendezvous.resolveCompletion(waiter, "already replied"));
|
||||||
|
|
||||||
|
resolver.resolve("term_a", turn);
|
||||||
|
|
||||||
|
assertTrue(notified.isEmpty(), "a classification that loses the race must not quarantine anything");
|
||||||
|
assertEquals("already replied", waiter.getNow(null).text(), "the earlier resolution stands untouched");
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void aNonMatchingScrapeResolvesAsAnOrdinaryCompletion() {
|
void aNonMatchingScrapeResolvesAsAnOrdinaryCompletion() {
|
||||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ complete report\n❯ ");
|
FakeHerdr herdr = new FakeHerdr().readText("⏺ complete report\n❯ ");
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
|
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
|
||||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns);
|
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, ExhaustionSink.none());
|
||||||
|
|
||||||
var waiter = rendezvous.open("term_a");
|
var waiter = rendezvous.open("term_a");
|
||||||
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
||||||
@@ -430,7 +471,7 @@ class CompletionResolverTest {
|
|||||||
FakeHerdr herdr = new FakeHerdr().readText(block);
|
FakeHerdr herdr = new FakeHerdr().readText(block);
|
||||||
Rendezvous rendezvous = new Rendezvous();
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
CompletionResolver resolver =
|
CompletionResolver resolver =
|
||||||
new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none());
|
new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
|
|
||||||
var waiter = rendezvous.open("term_a");
|
var waiter = rendezvous.open("term_a");
|
||||||
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
||||||
|
|||||||
@@ -72,7 +72,8 @@ class BridgeMcpAuthzTest {
|
|||||||
new PrimaryRegistry(null),
|
new PrimaryRegistry(null),
|
||||||
enforce ? CallerResolver.withLeadsAndMembers(identity, false, null,
|
enforce ? CallerResolver.withLeadsAndMembers(identity, false, null,
|
||||||
Map::of, new MemberRegistry(null)) : null,
|
Map::of, new MemberRegistry(null)) : null,
|
||||||
metrics, BridgeMcp.CapacitySource.none(), new BridgeMcp.HealthCoverageSource(() -> "off"));
|
metrics, BridgeMcp.CapacitySource.none(), new BridgeMcp.HealthCoverageSource(() -> "off"),
|
||||||
|
BridgeMcp.QuarantineSource.none());
|
||||||
return mcp;
|
return mcp;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -16,11 +16,15 @@ import dev.ltms.bridged.inject.MemberPresence;
|
|||||||
import dev.ltms.bridged.session.MemberSession;
|
import dev.ltms.bridged.session.MemberSession;
|
||||||
import dev.ltms.bridged.session.WorktreeRequest;
|
import dev.ltms.bridged.session.WorktreeRequest;
|
||||||
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
||||||
|
import dev.ltms.bridged.member.CompositePeerLauncher;
|
||||||
|
import dev.ltms.bridged.placement.BackendQuarantine;
|
||||||
|
import dev.ltms.bridged.placement.PlacementPolicies;
|
||||||
import io.modelcontextprotocol.spec.McpSchema;
|
import io.modelcontextprotocol.spec.McpSchema;
|
||||||
import dev.ltms.bridged.msg.InMemoryReplyInbox;
|
import dev.ltms.bridged.msg.InMemoryReplyInbox;
|
||||||
import org.junit.jupiter.api.BeforeEach;
|
import org.junit.jupiter.api.BeforeEach;
|
||||||
import org.junit.jupiter.api.Test;
|
import org.junit.jupiter.api.Test;
|
||||||
|
|
||||||
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
import java.util.Set;
|
import java.util.Set;
|
||||||
import java.util.concurrent.CompletableFuture;
|
import java.util.concurrent.CompletableFuture;
|
||||||
@@ -422,11 +426,41 @@ class BridgeMcpTest {
|
|||||||
assertTrue(textOf(res).contains("unknown worker profile"), textOf(res));
|
assertTrue(textOf(res).contains("unknown worker profile"), textOf(res));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-599: a profile at its {@code maxLoad} cap must surface a readable reason on the MCP
|
||||||
|
* surface too, not merely flip {@code isError} with an opaque or absent message.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void spawnAtMaxLoadSurfacesTheCapacityReason() {
|
||||||
|
FakeHerdr h = new FakeHerdr();
|
||||||
|
BridgedConfig.Profile wcfg = new BridgedConfig.Profile(
|
||||||
|
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
|
||||||
|
"tab", "bridged-workers", "worker: {profile} #{n}", null,
|
||||||
|
null, null, null, null, null, null, null, 0, null, null, null);
|
||||||
|
Map<String, BridgedConfig.Profile> profiles = Map.of(wcfg.profile(), wcfg);
|
||||||
|
ClaudeCodeLauncher delegate = new ClaudeCodeLauncher(
|
||||||
|
new AgentControl(h), new WorkspaceControl(h), new SubscriptionGuard(Set.of("gx00.gw")),
|
||||||
|
profiles, wcfg.profile(), k -> "BRIDGED_WORKER_TOKEN".equals(k) ? "tok" : null);
|
||||||
|
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||||
|
List.of(delegate), wcfg.profile(), profiles, PlacementPolicies.fixed(), _ -> 0);
|
||||||
|
SessionManager sm = new SessionManager(composite);
|
||||||
|
|
||||||
|
McpSchema.CallToolResult res = BridgeMcp.spawn(sm, "ltms-local");
|
||||||
|
|
||||||
|
assertTrue(res.isError());
|
||||||
|
String text = textOf(res);
|
||||||
|
assertTrue(text.contains("no capacity"), "surfaces a capacity reason, not a bare error: " + text);
|
||||||
|
assertTrue(text.contains("ltms-local"), "names the profile: " + text);
|
||||||
|
assertTrue(text.contains("maxLoad"), "explains the refusal: " + text);
|
||||||
|
assertFalse(h.called("agent.start"), "at cap, the spawn is refused before any herdr call");
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void spawnPassesTheRequestedCwdToTheWorker() {
|
void spawnPassesTheRequestedCwdToTheWorker() {
|
||||||
FakeHerdr h = new FakeHerdr();
|
FakeHerdr h = new FakeHerdr();
|
||||||
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
||||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, "/req/dir", null, null, null);
|
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, "/req/dir", null, null, null,
|
||||||
|
null, null);
|
||||||
assertNotEquals(Boolean.TRUE, res.isError());
|
assertNotEquals(Boolean.TRUE, res.isError());
|
||||||
// Protocol 19: the requested cwd roots the worker's pane at creation (tab.create).
|
// Protocol 19: the requested cwd roots the worker's pane at creation (tab.create).
|
||||||
@SuppressWarnings("unchecked")
|
@SuppressWarnings("unchecked")
|
||||||
@@ -437,11 +471,29 @@ class BridgeMcpTest {
|
|||||||
@Test
|
@Test
|
||||||
void profilesListsConfiguredProfilesAndDefault() {
|
void profilesListsConfiguredProfilesAndDefault() {
|
||||||
FakeHerdr h = new FakeHerdr();
|
FakeHerdr h = new FakeHerdr();
|
||||||
McpSchema.CallToolResult res = BridgeMcp.profiles(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
McpSchema.CallToolResult res = BridgeMcp.profiles(
|
||||||
|
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), BridgeMcp.QuarantineSource.none());
|
||||||
assertNotEquals(Boolean.TRUE, res.isError());
|
assertNotEquals(Boolean.TRUE, res.isError());
|
||||||
String out = textOf(res);
|
String out = textOf(res);
|
||||||
assertTrue(out.contains("ltms-local"), out);
|
assertTrue(out.contains("ltms-local"), out);
|
||||||
assertTrue(out.contains("\"default\":\"ltms-local\""), out);
|
assertTrue(out.contains("\"default\":\"ltms-local\""), out);
|
||||||
|
assertFalse(out.contains("quarantined"), "no profile is quarantined, so the key is omitted: " + out);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void profilesReportsAQuarantinedCredential() {
|
||||||
|
FakeHerdr h = new FakeHerdr();
|
||||||
|
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||||
|
quarantine.quarantine("shared-openai");
|
||||||
|
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(
|
||||||
|
profile -> "ltms-local".equals(profile) ? "shared-openai" : null, quarantine);
|
||||||
|
McpSchema.CallToolResult res = BridgeMcp.profiles(
|
||||||
|
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), source);
|
||||||
|
assertNotEquals(Boolean.TRUE, res.isError());
|
||||||
|
String out = textOf(res);
|
||||||
|
assertTrue(out.contains("\"quarantined\""), out);
|
||||||
|
assertTrue(out.contains("shared-openai"), out);
|
||||||
|
assertTrue(out.contains("\"quarantinedForSeconds\":1800"), out);
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -474,11 +526,14 @@ class BridgeMcpTest {
|
|||||||
sessions.acquire("ltms-local", null, null, null);
|
sessions.acquire("ltms-local", null, null, null);
|
||||||
McpSchema.CallToolResult res = BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
McpSchema.CallToolResult res = BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||||
sessions, null, new BridgeMcp.CapacitySource(profile -> 2, profile -> 2,
|
sessions, null, new BridgeMcp.CapacitySource(profile -> 2, profile -> 2,
|
||||||
() -> Set.of("ltms-local"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"), Map.of(), "");
|
() -> Set.of("ltms-local"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
|
||||||
|
BridgeMcp.QuarantineSource.none(), Map.of(), "");
|
||||||
String out = textOf(res);
|
String out = textOf(res);
|
||||||
assertTrue(out.contains("\"maxLoad\":2"), out);
|
assertTrue(out.contains("\"maxLoad\":2"), out);
|
||||||
assertTrue(out.contains("\"live\":2"), out);
|
assertTrue(out.contains("\"live\":2"), out);
|
||||||
assertTrue(out.contains("\"free\":0"), out);
|
assertTrue(out.contains("\"free\":0"), out);
|
||||||
|
assertFalse(out.contains("credentialId"), "nothing is quarantined, so no new key appears: " + out);
|
||||||
|
assertFalse(out.contains("quarantinedForSeconds"), out);
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -487,7 +542,8 @@ class BridgeMcpTest {
|
|||||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||||
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||||
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
|
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
|
||||||
() -> Set.of("terra"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"), Map.of(), ""));
|
() -> Set.of("terra"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
|
||||||
|
BridgeMcp.QuarantineSource.none(), Map.of(), ""));
|
||||||
assertTrue(out.contains("\"profile\":\"terra\""), out);
|
assertTrue(out.contains("\"profile\":\"terra\""), out);
|
||||||
assertTrue(out.contains("\"live\":0"), out);
|
assertTrue(out.contains("\"live\":0"), out);
|
||||||
assertTrue(out.contains("\"free\":2"), out);
|
assertTrue(out.contains("\"free\":2"), out);
|
||||||
@@ -499,10 +555,90 @@ class BridgeMcpTest {
|
|||||||
FakeHerdr h = new FakeHerdr();
|
FakeHerdr h = new FakeHerdr();
|
||||||
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||||
new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw"))), null,
|
new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw"))), null,
|
||||||
BridgeMcp.CapacitySource.none(), new BridgeMcp.HealthCoverageSource(() -> "off"), Map.of(), ""));
|
BridgeMcp.CapacitySource.none(), new BridgeMcp.HealthCoverageSource(() -> "off"),
|
||||||
|
BridgeMcp.QuarantineSource.none(), Map.of(), ""));
|
||||||
assertFalse(out.contains("\"capacity\":"), out);
|
assertFalse(out.contains("\"capacity\":"), out);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void quarantinedProfileReportsZeroFreeRegardlessOfMaxLoadAndLive() {
|
||||||
|
FakeHerdr h = new FakeHerdr();
|
||||||
|
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||||
|
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(20));
|
||||||
|
quarantine.quarantine("shared-openai");
|
||||||
|
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(
|
||||||
|
profile -> "terra".equals(profile) ? "shared-openai" : null, quarantine);
|
||||||
|
|
||||||
|
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||||
|
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
|
||||||
|
() -> Set.of("terra"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
|
||||||
|
source, Map.of(), ""));
|
||||||
|
|
||||||
|
assertTrue(out.contains("\"profile\":\"terra\""), out);
|
||||||
|
assertTrue(out.contains("\"maxLoad\":2"), out);
|
||||||
|
assertTrue(out.contains("\"live\":0"), out);
|
||||||
|
assertTrue(out.contains("\"free\":0"), out);
|
||||||
|
assertTrue(out.contains("\"credentialId\":\"shared-openai\""), out);
|
||||||
|
assertTrue(out.contains("\"quarantinedForSeconds\":1200"), out);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void everyProfileSharingTheQuarantinedCredentialReportsZeroFree() {
|
||||||
|
FakeHerdr h = new FakeHerdr();
|
||||||
|
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||||
|
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(20));
|
||||||
|
quarantine.quarantine("shared-openai");
|
||||||
|
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(_ -> "shared-openai", quarantine);
|
||||||
|
|
||||||
|
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||||
|
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
|
||||||
|
() -> Set.of("terra", "sol"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
|
||||||
|
source, Map.of(), ""));
|
||||||
|
|
||||||
|
assertEquals(2, out.split("\"free\":0", -1).length - 1, out);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void listAndProfilesAgreeOnWhatIsQuarantined() {
|
||||||
|
FakeHerdr h = new FakeHerdr();
|
||||||
|
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||||
|
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(20));
|
||||||
|
quarantine.quarantine("shared-openai");
|
||||||
|
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(
|
||||||
|
profile -> "ltms-local".equals(profile) ? "shared-openai" : null, quarantine);
|
||||||
|
var workers = workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||||
|
|
||||||
|
String listOut = textOf(BridgeMcp.listFleet(workers, sessions, null,
|
||||||
|
new BridgeMcp.CapacitySource(profile -> 0, profile -> 2, () -> Set.of("ltms-local"), () -> 0),
|
||||||
|
new BridgeMcp.HealthCoverageSource(() -> "off"), source, Map.of(), ""));
|
||||||
|
String profilesOut = textOf(BridgeMcp.profiles(workers, source));
|
||||||
|
|
||||||
|
assertTrue(listOut.contains("\"free\":0"), listOut);
|
||||||
|
assertTrue(listOut.contains("\"credentialId\":\"shared-openai\""), listOut);
|
||||||
|
assertTrue(profilesOut.contains("\"quarantined\""), profilesOut);
|
||||||
|
assertTrue(profilesOut.contains("shared-openai"), profilesOut);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void expiredQuarantineLeavesTheCapacityRowOrdinary() {
|
||||||
|
FakeHerdr h = new FakeHerdr();
|
||||||
|
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||||
|
long[] clockNanos = {0L};
|
||||||
|
BackendQuarantine quarantine = new BackendQuarantine(() -> clockNanos[0], TimeUnit.MINUTES.toNanos(20));
|
||||||
|
quarantine.quarantine("shared-openai");
|
||||||
|
clockNanos[0] = TimeUnit.MINUTES.toNanos(21); // clock advances past the cooldown
|
||||||
|
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(
|
||||||
|
profile -> "terra".equals(profile) ? "shared-openai" : null, quarantine);
|
||||||
|
|
||||||
|
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||||
|
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
|
||||||
|
() -> Set.of("terra"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
|
||||||
|
source, Map.of(), ""));
|
||||||
|
|
||||||
|
assertTrue(out.contains("\"free\":2"), out);
|
||||||
|
assertFalse(out.contains("quarantinedForSeconds"), out);
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void listReportsLeadsAndFlagsTheCallersOwnRow() {
|
void listReportsLeadsAndFlagsTheCallersOwnRow() {
|
||||||
FakeHerdr h = new FakeHerdr();
|
FakeHerdr h = new FakeHerdr();
|
||||||
@@ -637,6 +773,52 @@ class BridgeMcpTest {
|
|||||||
assertEquals("blocked", textOf(res));
|
assertEquals("blocked", textOf(res));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-582: a lead polling {@code bridge_status} on its normal cadence — not {@code bridge_poll}
|
||||||
|
* — must also see a worker's open async {@code bridge_ask} question, since the reverse-rendezvous
|
||||||
|
* window it opened with is far shorter than that cadence.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void statusReportsAnOpenQuestionWhenTheWorkerIsMidAsk() throws Exception {
|
||||||
|
String ticket = messages.sendAsync(T, "task that asks");
|
||||||
|
long deadline = System.currentTimeMillis() + 3000;
|
||||||
|
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||||
|
Thread.sleep(5);
|
||||||
|
}
|
||||||
|
assertTrue(rendezvous.isWaiting(T), "sendAsync should have opened its rendezvous waiter");
|
||||||
|
|
||||||
|
CompletableFuture<MessageService.AskResult> ask =
|
||||||
|
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||||
|
|
||||||
|
MessageService.TaskView asking;
|
||||||
|
deadline = System.currentTimeMillis() + 3000;
|
||||||
|
do {
|
||||||
|
asking = messages.poll(ticket);
|
||||||
|
Thread.sleep(5);
|
||||||
|
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
|
||||||
|
assertEquals(MessageService.Phase.ASKING, asking.phase());
|
||||||
|
|
||||||
|
McpSchema.CallToolResult res = BridgeMcp.status(messages, T);
|
||||||
|
assertNotEquals(Boolean.TRUE, res.isError());
|
||||||
|
String out = textOf(res);
|
||||||
|
assertTrue(out.startsWith("idle"), "the live status must still lead the text: " + out);
|
||||||
|
assertTrue(out.contains("which config file?"), "the question text must be shown: " + out);
|
||||||
|
assertTrue(out.contains("turnId=\"" + asking.turnId() + "\""), "the turnId must be shown: " + out);
|
||||||
|
assertTrue(out.contains("(ticket " + ticket + ")"), "the ticket must be shown: " + out);
|
||||||
|
|
||||||
|
// Clean up the still-open ask so the background thread does not linger past the test.
|
||||||
|
String turnId = asking.turnId();
|
||||||
|
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||||
|
() -> messages.answer(turnId, "config.yaml", 5000));
|
||||||
|
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||||
|
deadline = System.currentTimeMillis() + 3000;
|
||||||
|
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||||
|
Thread.sleep(5);
|
||||||
|
}
|
||||||
|
assertTrue(rendezvous.resolve(T, "done"));
|
||||||
|
answer.get(5, TimeUnit.SECONDS);
|
||||||
|
}
|
||||||
|
|
||||||
// --- bridge_whoami: the caller's own identity, so an agent never has to guess its role -------
|
// --- bridge_whoami: the caller's own identity, so an agent never has to guess its role -------
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -783,7 +965,8 @@ class BridgeMcpTest {
|
|||||||
void spawnDefaultsTheRoleToDev() {
|
void spawnDefaultsTheRoleToDev() {
|
||||||
FakeHerdr h = new FakeHerdr();
|
FakeHerdr h = new FakeHerdr();
|
||||||
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
||||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, null, null, null, null);
|
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, null, null, null, null,
|
||||||
|
null, null);
|
||||||
|
|
||||||
assertNotEquals(Boolean.TRUE, res.isError());
|
assertNotEquals(Boolean.TRUE, res.isError());
|
||||||
assertTrue(textOf(res).contains("\"role\":\"dev\""), textOf(res));
|
assertTrue(textOf(res).contains("\"role\":\"dev\""), textOf(res));
|
||||||
@@ -794,7 +977,7 @@ class BridgeMcpTest {
|
|||||||
FakeHerdr h = new FakeHerdr();
|
FakeHerdr h = new FakeHerdr();
|
||||||
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
||||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "architect",
|
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "architect",
|
||||||
null, null, null, null);
|
null, null, null, null, null, null);
|
||||||
|
|
||||||
assertNotEquals(Boolean.TRUE, res.isError());
|
assertNotEquals(Boolean.TRUE, res.isError());
|
||||||
assertTrue(textOf(res).contains("\"role\":\"architect\""), textOf(res));
|
assertTrue(textOf(res).contains("\"role\":\"architect\""), textOf(res));
|
||||||
@@ -805,9 +988,36 @@ class BridgeMcpTest {
|
|||||||
FakeHerdr h = new FakeHerdr();
|
FakeHerdr h = new FakeHerdr();
|
||||||
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
||||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "worker",
|
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "worker",
|
||||||
null, null, null, null);
|
null, null, null, null, null, null);
|
||||||
|
|
||||||
assertEquals(Boolean.TRUE, res.isError());
|
assertEquals(Boolean.TRUE, res.isError());
|
||||||
assertTrue(textOf(res).contains("architect, dev, reviewer"), textOf(res));
|
assertTrue(textOf(res).contains("architect, dev, reviewer"), textOf(res));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ── CB-584: bridge_spawn accepts sessionName/resumeSessionId; roster shows agentSessionId ──
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void spawnWithResumeSessionIdPutsTheIdOnTheRoster() {
|
||||||
|
FakeHerdr h = new FakeHerdr();
|
||||||
|
SessionManager sessions = sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||||
|
|
||||||
|
McpSchema.CallToolResult res = BridgeMcp.spawn(sessions, "ltms-local", null, null, null, null, null,
|
||||||
|
null, "cb-resume-99");
|
||||||
|
assertNotEquals(Boolean.TRUE, res.isError(), textOf(res));
|
||||||
|
|
||||||
|
String listOut = textOf(BridgeMcp.listFleet(
|
||||||
|
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, Map.of(), ""));
|
||||||
|
assertTrue(listOut.contains("\"agentSessionId\":\"cb-resume-99\""), listOut);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void spawnRefusesResumeSessionIdWithoutAnExplicitProfile() {
|
||||||
|
FakeHerdr h = new FakeHerdr();
|
||||||
|
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
||||||
|
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null,
|
||||||
|
null, null, null, null, null, "cb-resume-99");
|
||||||
|
|
||||||
|
assertEquals(Boolean.TRUE, res.isError());
|
||||||
|
assertTrue(textOf(res).contains("explicit profile"), textOf(res));
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,5 +1,8 @@
|
|||||||
package dev.ltms.bridged.member;
|
package dev.ltms.bridged.member;
|
||||||
|
|
||||||
|
import ch.qos.logback.classic.Logger;
|
||||||
|
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||||
|
import ch.qos.logback.core.read.ListAppender;
|
||||||
import dev.ltms.bridged.config.BridgedConfig;
|
import dev.ltms.bridged.config.BridgedConfig;
|
||||||
import dev.ltms.bridged.guard.GuardException;
|
import dev.ltms.bridged.guard.GuardException;
|
||||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||||
@@ -12,6 +15,7 @@ import dev.ltms.bridged.peer.PeerHandle;
|
|||||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||||
import dev.ltms.bridged.peer.SpawnRequest;
|
import dev.ltms.bridged.peer.SpawnRequest;
|
||||||
import org.junit.jupiter.api.Test;
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
@@ -676,6 +680,174 @@ class ClaudeCodeLauncherTest {
|
|||||||
"the guard-checked baseUrl must win over any env: entry, or the boundary is bypassable");
|
"the guard-checked baseUrl must win over any env: entry, or the boundary is bypassable");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- CB-596: config-driven member-credential policy (replaces CB-592's hardcoded single name) --
|
||||||
|
|
||||||
|
/** A representative {@code memberCredentials} — 4 allowed, 4 blocked, matching the real ticket shape. */
|
||||||
|
private static final BridgedConfig.MemberCredentials TEST_MEMBER_CREDENTIALS = new BridgedConfig.MemberCredentials(
|
||||||
|
null,
|
||||||
|
List.of("AI_GATEWAY_TOKEN", "WORKER_GITEA_TOKEN", "CONTEXT7_TOKEN", "GITEA_HOST"),
|
||||||
|
List.of("AI_GATEWAY_TOKEN", "WORKER_GITEA_TOKEN", "CONTEXT7_TOKEN", "GITEA_HOST",
|
||||||
|
"GITEA_ACCESS_TOKEN", "GITLAB_PERSONAL_ACCESS_TOKEN", "TS_AUTHKEY", "HASS_TOKEN"));
|
||||||
|
|
||||||
|
private ClaudeCodeLauncher serviceWithCredentials(FakeHerdr herdr, BridgedConfig.MemberCredentials creds) {
|
||||||
|
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||||
|
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||||
|
List.of("claude"), "tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
|
||||||
|
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||||
|
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||||
|
_ -> null, 0, System::currentTimeMillis, () -> {}, null, () -> creds);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* herdr's env map is an overlay onto its own (login-shell) process environment, so a worker
|
||||||
|
* inherits whatever the daemon's shell carries — including the operator's own credentials — for
|
||||||
|
* every key {@code baseEnv} does not explicitly shadow. This pins that every {@code known} name
|
||||||
|
* NOT also {@code allow}-ed gets an explicit (non-blank) sentinel overlay, whatever the profile
|
||||||
|
* is. Asserted against what tab.create's params actually carry, not an internal map built in the
|
||||||
|
* test (gitea #82).
|
||||||
|
*
|
||||||
|
* <p>Scope, measured on a live pane 2026-08-15 (CB-592): this pins what the launcher SENDS, and
|
||||||
|
* that is all a unit test can pin. It does not prove the value survives, and for names the
|
||||||
|
* operator's secrets.sh also exports it does not: the pane runs a login shell that puts the real
|
||||||
|
* value back over this sentinel unless the export is guarded on BRIDGED_MEMBER — see
|
||||||
|
* everySpawnMarksThePaneAsAMember below.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void everyKnownNameNotAllowedIsShadowedWithTheSentinel() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
serviceWithCredentials(herdr, TEST_MEMBER_CREDENTIALS).spawn();
|
||||||
|
|
||||||
|
Map<String, String> env = startEnv(herdr);
|
||||||
|
for (String blocked : List.of("GITEA_ACCESS_TOKEN", "GITLAB_PERSONAL_ACCESS_TOKEN", "TS_AUTHKEY", "HASS_TOKEN")) {
|
||||||
|
String shadowed = env.get(blocked);
|
||||||
|
assertNotNull(shadowed, blocked + " must be explicitly overlaid, not left unmentioned");
|
||||||
|
assertFalse(shadowed.isBlank(), blocked + "'s overlay value must be non-blank");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* An {@code allow}-ed name must get NO overlay entry at all — any entry, blank or not, risks
|
||||||
|
* overriding the real value the pane needs, and the whole point of {@code allow} is that the
|
||||||
|
* pane's own inherited value passes through untouched.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void everyAllowedNameGetsNoOverlayEntrySoTheRealValuePassesThrough() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
serviceWithCredentials(herdr, TEST_MEMBER_CREDENTIALS).spawn();
|
||||||
|
|
||||||
|
Map<String, String> env = startEnv(herdr);
|
||||||
|
for (String allowed : List.of("AI_GATEWAY_TOKEN", "WORKER_GITEA_TOKEN", "CONTEXT7_TOKEN", "GITEA_HOST")) {
|
||||||
|
assertFalse(env.containsKey(allowed),
|
||||||
|
allowed + " is allow-listed — the launcher must not mention it at all");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* No {@code memberCredentials} configured (the pre-CB-596 constructor overloads still used
|
||||||
|
* throughout this file, and the shape a fresh {@code bridged.yaml} with no memberCredentials:
|
||||||
|
* block resolves to) blocks NOTHING. This documents the transitional gap rather than hiding it —
|
||||||
|
* see {@code HerdrPeerLauncher#applyMemberCredentialPolicy}'s javadoc.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void noMemberCredentialsConfiguredBlocksNothing() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
service(herdr, List.of("claude"), null).spawn();
|
||||||
|
|
||||||
|
assertNull(startEnv(herdr).get("GITEA_ACCESS_TOKEN"),
|
||||||
|
"with no memberCredentials configured, nothing is shadowed — config must supply the policy");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* No profile — present or future — may restore a blocked name by naming it in {@code env:}.
|
||||||
|
* The shadow is applied after the profile's own env in {@link HerdrPeerLauncher#baseEnv}
|
||||||
|
* precisely so this can never happen; this test pins that ordering.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void aProfileEnvEntryCannotRestoreABlockedName() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||||
|
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||||
|
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
|
||||||
|
null, Map.of("GITEA_ACCESS_TOKEN", "admin-secret-from-profile-config"), null, null);
|
||||||
|
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||||
|
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||||
|
_ -> null, 0, System::currentTimeMillis, () -> {}, null, () -> TEST_MEMBER_CREDENTIALS)
|
||||||
|
.spawn(cfg.profile(), null, null);
|
||||||
|
|
||||||
|
assertNotEquals("admin-secret-from-profile-config", startEnv(herdr).get("GITEA_ACCESS_TOKEN"),
|
||||||
|
"a profile's own env: must not be able to smuggle a blocked name back in");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596 criterion 4: a credential-shaped host env var name on neither {@code known} nor
|
||||||
|
* {@code allow} is not silently allowed — it must be reported (never its value). This pins the
|
||||||
|
* WARN naming the gap, using an injected host-env-names source rather than the real
|
||||||
|
* {@code System.getenv()} so the test is deterministic.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void aCredentialShapedNameOnNeitherListIsLoggedAsAGap() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||||
|
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||||
|
List.of("claude"), "tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
|
||||||
|
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||||
|
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||||
|
_ -> null, 0, System::currentTimeMillis, () -> {}, null, () -> TEST_MEMBER_CREDENTIALS,
|
||||||
|
() -> Set.of("PATH", "HOME", "AI_GATEWAY_TOKEN", "A_BRAND_NEW_SECRET_TOKEN"));
|
||||||
|
|
||||||
|
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
|
||||||
|
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||||
|
appender.start();
|
||||||
|
logger.addAppender(appender);
|
||||||
|
try {
|
||||||
|
svc.spawn();
|
||||||
|
} finally {
|
||||||
|
logger.detachAppender(appender);
|
||||||
|
}
|
||||||
|
|
||||||
|
assertTrue(appender.list.stream().anyMatch(e ->
|
||||||
|
e.getFormattedMessage().contains("memberCredentials gap")
|
||||||
|
&& e.getFormattedMessage().contains("A_BRAND_NEW_SECRET_TOKEN")),
|
||||||
|
"the gap must name the unrecognized credential-shaped var, never a value");
|
||||||
|
assertFalse(appender.list.stream().anyMatch(e -> e.getFormattedMessage().contains("PATH")),
|
||||||
|
"PATH/HOME are not credential-shaped and must not be reported as a gap");
|
||||||
|
assertFalse(appender.list.stream().anyMatch(e -> e.getFormattedMessage().contains("AI_GATEWAY_TOKEN")),
|
||||||
|
"a name already on allow: is covered, not a gap");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The half of CB-592 that can actually survive the pane's login shell. BRIDGED_MEMBER is a name
|
||||||
|
* secrets.sh never exports, so nothing overwrites it — measured: GITEA_TOKEN is injected the
|
||||||
|
* same way, is absent from a login shell of its own, and was observed set inside a live member
|
||||||
|
* pane. It lets the operator guard the admin export with
|
||||||
|
* `[ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...`, which is the whole fix.
|
||||||
|
* Pinned here so a refactor cannot drop the marker and quietly un-guard every member (#77).
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void everySpawnMarksThePaneAsAMember() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
service(herdr, List.of("claude"), null).spawn();
|
||||||
|
|
||||||
|
assertEquals("1", startEnv(herdr).get("BRIDGED_MEMBER"),
|
||||||
|
"every member pane must be marked, or a shell file cannot tell it apart from the operator's");
|
||||||
|
}
|
||||||
|
|
||||||
|
/** A profile must not be able to hide that its pane is a member, for the same reason as above. */
|
||||||
|
@Test
|
||||||
|
void aProfileEnvEntryCannotClearTheMemberMarker() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||||
|
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||||
|
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
|
||||||
|
null, Map.of("BRIDGED_MEMBER", ""), null, null);
|
||||||
|
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||||
|
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||||
|
_ -> null).spawn();
|
||||||
|
|
||||||
|
assertEquals("1", startEnv(herdr).get("BRIDGED_MEMBER"),
|
||||||
|
"a profile's own env: must not be able to unmark its pane");
|
||||||
|
}
|
||||||
|
|
||||||
// ── CB-533: the model is pinned on the command line, not only in the environment ────────────
|
// ── CB-533: the model is pinned on the command line, not only in the environment ────────────
|
||||||
|
|
||||||
/** A launcher for a profile identical but for its {@code model:} — the only variable here. */
|
/** A launcher for a profile identical but for its {@code model:} — the only variable here. */
|
||||||
|
|||||||
@@ -16,6 +16,7 @@ import dev.ltms.bridged.peer.PeerHandle;
|
|||||||
import dev.ltms.bridged.peer.PeerLauncher;
|
import dev.ltms.bridged.peer.PeerLauncher;
|
||||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||||
import dev.ltms.bridged.peer.SpawnRequest;
|
import dev.ltms.bridged.peer.SpawnRequest;
|
||||||
|
import dev.ltms.bridged.placement.BackendQuarantine;
|
||||||
import dev.ltms.bridged.placement.PlacementException;
|
import dev.ltms.bridged.placement.PlacementException;
|
||||||
import dev.ltms.bridged.placement.PlacementPolicies;
|
import dev.ltms.bridged.placement.PlacementPolicies;
|
||||||
import org.junit.jupiter.api.Test;
|
import org.junit.jupiter.api.Test;
|
||||||
@@ -27,6 +28,8 @@ import java.util.LinkedHashMap;
|
|||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
import java.util.Set;
|
import java.util.Set;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.*;
|
import static org.junit.jupiter.api.Assertions.*;
|
||||||
|
|
||||||
@@ -94,6 +97,7 @@ class CompositePeerLauncherTest {
|
|||||||
@Override public String id() { return "pane-" + p; }
|
@Override public String id() { return "pane-" + p; }
|
||||||
@Override public String terminalId() { return "term-" + p; }
|
@Override public String terminalId() { return "term-" + p; }
|
||||||
@Override public String profile() { return p; }
|
@Override public String profile() { return p; }
|
||||||
|
@Override public String agentSessionId() { return null; }
|
||||||
@Override public CharterReceipt charterReceipt() { return null; }
|
@Override public CharterReceipt charterReceipt() { return null; }
|
||||||
};
|
};
|
||||||
}
|
}
|
||||||
@@ -128,6 +132,13 @@ class CompositePeerLauncherTest {
|
|||||||
weight, maxLoad);
|
weight, maxLoad);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private static BridgedConfig.Profile stubWorker(String profile, String credentialId) {
|
||||||
|
return new BridgedConfig.Profile(profile, "http://gx00.gw:8000", "coder",
|
||||||
|
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers",
|
||||||
|
"w #{n}", null, null, null, null, null, null, null, null, null,
|
||||||
|
null, null, credentialId);
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* An <em>order-preserving</em> profile map. Never {@code Map.of} here: its iteration order is
|
* An <em>order-preserving</em> profile map. Never {@code Map.of} here: its iteration order is
|
||||||
* salted per JVM run, and the weighted policy breaks an exact-weight tie on candidate order —
|
* salted per JVM run, and the weighted policy breaks an exact-weight tie on candidate order —
|
||||||
@@ -326,6 +337,42 @@ class CompositePeerLauncherTest {
|
|||||||
assertEquals(10, b);
|
assertEquals(10, b);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void weightedPolicySkipsWeightZeroProfileOnUnqualifiedSpawn() {
|
||||||
|
// CB-554: weight: 0 must exclude a profile from automatic placement, not coerce to 1.0.
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||||
|
"a", stubWorker("a", 0.0f, null),
|
||||||
|
"b", stubWorker("b", 1.0f, null));
|
||||||
|
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
|
||||||
|
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||||
|
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), _ -> 0);
|
||||||
|
|
||||||
|
for (int i = 0; i < 5; i++) {
|
||||||
|
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
|
||||||
|
assertEquals("b", h.profile(), "profile a has weight 0, so every unqualified spawn must land on b");
|
||||||
|
}
|
||||||
|
assertEquals(0, adapter.spawnCount("a"), "a is never chosen automatically");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void explicitSpawnStillSucceedsOnWeightZeroProfile() {
|
||||||
|
// CB-554: weight: 0 excludes a profile from AUTOMATIC selection only — an explicit
|
||||||
|
// bridge_spawn{profile:"a"} must still work exactly as today (e.g. `opus` on the
|
||||||
|
// operator's own subscription, kept weight-0 so it is never picked automatically).
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||||
|
"a", stubWorker("a", 0.0f, null),
|
||||||
|
"b", stubWorker("b", 1.0f, null));
|
||||||
|
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "b", Set.of());
|
||||||
|
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||||
|
List.of(adapter), "b", profiles, PlacementPolicies.weighted(), _ -> 0);
|
||||||
|
|
||||||
|
PeerHandle h = composite.spawn(new SpawnRequest("a", null, null));
|
||||||
|
assertEquals("a", h.profile(), "naming a weight-0 profile explicitly bypasses placement and still spawns it");
|
||||||
|
assertEquals(1, adapter.spawnCount("a"));
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void failoverRetriesNextCandidateWhenProfileIsUnreachable() {
|
void failoverRetriesNextCandidateWhenProfileIsUnreachable() {
|
||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
@@ -427,6 +474,27 @@ class CompositePeerLauncherTest {
|
|||||||
assertEquals("a", h.profile(), "a profile with no maxLoad is never capped, however many live workers");
|
assertEquals("a", h.profile(), "a profile with no maxLoad is never capped, however many live workers");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void explicitSpawnOnMaxLoadZeroProfileIsRefusedEvenWithZeroLiveWorkers() {
|
||||||
|
// CB-585: before the fix, maxLoad: 0 normalised to null (unlimited) in the compact
|
||||||
|
// constructor, so this exact case — naming a zero-cap profile explicitly, with nothing
|
||||||
|
// live on it yet — would have spawned instead of refusing. "at most zero members" must
|
||||||
|
// hold even when the profile is named directly, not only against automatic placement.
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||||
|
"a", stubWorker("a", 1.0f, 0),
|
||||||
|
"b", stubWorker("b"));
|
||||||
|
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
|
||||||
|
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||||
|
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), _ -> 0);
|
||||||
|
|
||||||
|
PlacementException e = assertThrows(PlacementException.class,
|
||||||
|
() -> composite.spawn(new SpawnRequest("a", null, null)));
|
||||||
|
assertTrue(e.getMessage().contains("'a'"), "message names the profile: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("0 cap"), "message names the zero cap: " + e.getMessage());
|
||||||
|
assertEquals(0, adapter.spawnCount("a"), "a maxLoad: 0 profile accepts no explicit spawn");
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void emptyCandidateSetThrowsClearException() {
|
void emptyCandidateSetThrowsClearException() {
|
||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
@@ -554,4 +622,105 @@ class CompositePeerLauncherTest {
|
|||||||
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)));
|
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)));
|
||||||
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
|
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ── CB-578 stage B: a BACKEND_EXHAUSTED classification quarantines the credential ──────────
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void explicitSpawnOntoAQuarantinedProfileIsRefusedNamingTheCredentialAndRemainingTime() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||||
|
"sol", stubWorker("sol", "shared-openai"),
|
||||||
|
"terra", stubWorker("terra", "shared-openai"));
|
||||||
|
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
|
||||||
|
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||||
|
quarantine.quarantine("shared-openai");
|
||||||
|
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||||
|
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
|
||||||
|
|
||||||
|
PlacementException e = assertThrows(PlacementException.class,
|
||||||
|
() -> composite.spawn(new SpawnRequest("sol", null, null)));
|
||||||
|
assertTrue(e.getMessage().contains("sol"), "message names the profile: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("shared-openai"), "message names the credential: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("1800"), "message names roughly when it lifts: " + e.getMessage());
|
||||||
|
assertEquals(0, adapter.spawnCount("sol"), "the quarantined profile is never delegated to");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The part CB-578 stage B calls out as easy to get wrong: sol and terra are two different
|
||||||
|
* profiles sharing one OpenAI credential. Quarantining because of an exhaustion classified on
|
||||||
|
* ONE of them must lock out the other too, or the fleet just walks onto the same dead account
|
||||||
|
* under the sibling's name.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void twoProfilesSharingACredentialAreBothQuarantinedByOneExhaustionEvent() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||||
|
"sol", stubWorker("sol", "shared-openai"),
|
||||||
|
"terra", stubWorker("terra", "shared-openai"));
|
||||||
|
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
|
||||||
|
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||||
|
// Only "sol" was classified BACKEND_EXHAUSTED — but the two profiles share one credential.
|
||||||
|
quarantine.quarantine("shared-openai");
|
||||||
|
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||||
|
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
|
||||||
|
|
||||||
|
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("sol", null, null)),
|
||||||
|
"sol was the one classified exhausted");
|
||||||
|
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("terra", null, null)),
|
||||||
|
"terra shares sol's credential, so it must be locked out too");
|
||||||
|
assertEquals(0, adapter.spawnCount("sol"));
|
||||||
|
assertEquals(0, adapter.spawnCount("terra"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void placementSkipsAQuarantinedProfileAndRoutesToAnUnquarantinedOne() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||||
|
"sol", stubWorker("sol", "shared-openai"),
|
||||||
|
"b", stubWorker("b"));
|
||||||
|
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
|
||||||
|
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||||
|
quarantine.quarantine("shared-openai");
|
||||||
|
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||||
|
PlacementPolicies.weighted(), _ -> 0, null, quarantine);
|
||||||
|
|
||||||
|
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
|
||||||
|
assertEquals("b", h.profile(), "sol is quarantined, so an unqualified spawn must land on b");
|
||||||
|
assertEquals(0, adapter.spawnCount("sol"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aQuarantineLiftsOnTheInjectedClockAndTheProfileBecomesSpawnableAgain() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||||
|
"sol", stubWorker("sol", "shared-openai"),
|
||||||
|
"b", stubWorker("b"));
|
||||||
|
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
|
||||||
|
AtomicLong nowNanos = new AtomicLong(0L);
|
||||||
|
BackendQuarantine quarantine = new BackendQuarantine(nowNanos::get, TimeUnit.MINUTES.toNanos(30));
|
||||||
|
quarantine.quarantine("shared-openai");
|
||||||
|
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||||
|
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
|
||||||
|
|
||||||
|
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("sol", null, null)),
|
||||||
|
"still inside the cooldown");
|
||||||
|
|
||||||
|
nowNanos.set(TimeUnit.MINUTES.toNanos(31));
|
||||||
|
|
||||||
|
PeerHandle h = composite.spawn(new SpawnRequest("sol", null, null));
|
||||||
|
assertEquals("sol", h.profile(), "the cooldown expired on the injected clock — sol is spawnable again");
|
||||||
|
assertEquals(1, adapter.spawnCount("sol"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aFleetWithNoExhaustedPatternAnywhereBehavesExactlyAsBeforeQuarantineExisted() {
|
||||||
|
// BackendQuarantine.none() is the inert stand-in every constructor already defaults to when
|
||||||
|
// no quarantine is wired — the 2/5/6-arg constructors used throughout this file all exercise
|
||||||
|
// it. This test pins that an explicit .none() also never refuses a spawn, for any profile.
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
PeerLauncher composite = composite(herdr);
|
||||||
|
|
||||||
|
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("claude", null, null)));
|
||||||
|
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("gemini", null, null)));
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -52,6 +52,14 @@ class OpenCodeLauncherTest {
|
|||||||
0, System::currentTimeMillis, () -> { }, configRoot, configRoot, fleet);
|
0, System::currentTimeMillis, () -> { }, configRoot, configRoot, fleet);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private static OpenCodeLauncher serviceWithCredentials(FakeHerdr herdr, Path configRoot,
|
||||||
|
BridgedConfig.Profile cfg,
|
||||||
|
BridgedConfig.MemberCredentials creds) {
|
||||||
|
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||||
|
Map.of(cfg.profile(), cfg), cfg.profile(), k -> "GITEA_ACCESS_TOKEN".equals(k) ? "tok" : null,
|
||||||
|
0, System::currentTimeMillis, () -> { }, configRoot, configRoot, null, () -> creds);
|
||||||
|
}
|
||||||
|
|
||||||
@SuppressWarnings("unchecked")
|
@SuppressWarnings("unchecked")
|
||||||
private static Map<String, Object> lastStart(FakeHerdr herdr) {
|
private static Map<String, Object> lastStart(FakeHerdr herdr) {
|
||||||
return (Map<String, Object>) herdr.lastCall("agent.start").params();
|
return (Map<String, Object>) herdr.lastCall("agent.start").params();
|
||||||
@@ -208,6 +216,36 @@ class OpenCodeLauncherTest {
|
|||||||
"a git-token profile gets the peer-neutral GITEA_TOKEN grant, same as Claude");
|
"a git-token profile gets the peer-neutral GITEA_TOKEN grant, same as Claude");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: the config-driven shadow lives in {@link HerdrPeerLauncher#baseEnv}, shared by every
|
||||||
|
* adapter — this pins that the opencode path gets it too, not just Claude's. See the matching
|
||||||
|
* tests in {@code ClaudeCodeLauncherTest} for the full rationale (gitea #82, superseding CB-592's
|
||||||
|
* single hardcoded name).
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void aKnownNameNotAllowedIsShadowedWithTheSentinel(@TempDir Path root) {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
BridgedConfig.MemberCredentials creds = new BridgedConfig.MemberCredentials(
|
||||||
|
null, List.of("AI_GATEWAY_TOKEN"), List.of("AI_GATEWAY_TOKEN", "GITEA_ACCESS_TOKEN"));
|
||||||
|
serviceWithCredentials(herdr, root, opencodeCfg(null, null, null), creds).spawn();
|
||||||
|
|
||||||
|
String shadowed = startEnv(herdr).get("GITEA_ACCESS_TOKEN");
|
||||||
|
assertNotNull(shadowed, "GITEA_ACCESS_TOKEN must be explicitly overlaid, not left unmentioned");
|
||||||
|
assertFalse(shadowed.isBlank(), "a blank overlay value's override behaviour is unverified — must be non-blank");
|
||||||
|
assertFalse(startEnv(herdr).containsKey("AI_GATEWAY_TOKEN"),
|
||||||
|
"an allow-listed name must get no overlay entry at all");
|
||||||
|
}
|
||||||
|
|
||||||
|
/** No {@code memberCredentials} configured (the pre-CB-596 constructor overloads) blocks nothing. */
|
||||||
|
@Test
|
||||||
|
void noMemberCredentialsConfiguredBlocksNothing(@TempDir Path root) {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
service(herdr, root, opencodeCfg(null, null, null)).spawn();
|
||||||
|
|
||||||
|
assertNull(startEnv(herdr).get("GITEA_ACCESS_TOKEN"),
|
||||||
|
"with no memberCredentials configured, nothing is shadowed — config must supply the policy");
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void capabilitiesDeclareOrphanReapAndMcpAskAndConditionalSelfPr(@TempDir Path root) {
|
void capabilitiesDeclareOrphanReapAndMcpAskAndConditionalSelfPr(@TempDir Path root) {
|
||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
|||||||
@@ -16,6 +16,7 @@ import java.util.List;
|
|||||||
import java.util.concurrent.TimeUnit;
|
import java.util.concurrent.TimeUnit;
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -166,6 +167,88 @@ class AmqpReplyInboxContractTest {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void prefetchBoundsTheHeldBacklog() throws Exception {
|
||||||
|
String target = "worker-prefetch-" + System.nanoTime();
|
||||||
|
int prefetch = 4;
|
||||||
|
int published = 10;
|
||||||
|
try (AmqpReplyInbox inbox = new AmqpReplyInbox(newConnection(), prefetch);
|
||||||
|
Connection inspect = newConnection()) {
|
||||||
|
inbox.own(target);
|
||||||
|
for (int i = 0; i < published; i++) {
|
||||||
|
inbox.publish(target, "m" + i, "payload " + i);
|
||||||
|
}
|
||||||
|
awaitHeldAtLeast(inbox, target, prefetch);
|
||||||
|
|
||||||
|
long depth;
|
||||||
|
try (Channel ch = inspect.createChannel()) {
|
||||||
|
depth = ch.queueDeclarePassive(queueName(target)).getMessageCount();
|
||||||
|
}
|
||||||
|
assertTrue(depth >= published - prefetch,
|
||||||
|
"broker should still hold at least " + (published - prefetch)
|
||||||
|
+ " undelivered messages behind a prefetch of " + prefetch + ", saw " + depth);
|
||||||
|
|
||||||
|
drainUntilEmpty(inbox, target, inspect);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void unroutablePublishReportsFailureNotSilentSuccess() throws Exception {
|
||||||
|
String target = "worker-unroutable-" + System.nanoTime();
|
||||||
|
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
|
||||||
|
// Deliberately never own(target): the queue is never declared, so the default-exchange
|
||||||
|
// route to agent.<target>.inbox does not exist and the broker must return the publish.
|
||||||
|
IllegalStateException ex = assertThrows(IllegalStateException.class,
|
||||||
|
() -> inbox.publish(target, "m1", "nobody home"));
|
||||||
|
assertTrue(ex.getMessage() != null && ex.getMessage().toLowerCase().contains("unroutable"),
|
||||||
|
"expected an unroutable-publish failure, got: " + ex.getMessage());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void confirmedPublishDeliversNormally() throws Exception {
|
||||||
|
String target = "worker-confirm-" + System.nanoTime();
|
||||||
|
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
|
||||||
|
inbox.own(target);
|
||||||
|
inbox.publish(target, "m1", "confirmed delivery"); // must return normally: routed and confirmed
|
||||||
|
|
||||||
|
List<ReplyInbox.InboxMessage> got = awaitPeek(inbox, target);
|
||||||
|
assertEquals(1, got.size());
|
||||||
|
assertEquals("confirmed delivery", got.getFirst().content());
|
||||||
|
inbox.ack(target, "m1");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Poll peek until at least {@code n} replies for {@code target} are held, or ~10s elapse. */
|
||||||
|
@SuppressWarnings("BusyWait")
|
||||||
|
private static void awaitHeldAtLeast(AmqpReplyInbox inbox, String target, int n) throws InterruptedException {
|
||||||
|
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(10);
|
||||||
|
while (inbox.peek(target).size() < n && System.nanoTime() < deadline) {
|
||||||
|
Thread.sleep(50);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Repeatedly ack whatever is currently held (freeing prefetch slots for the next delivery) until
|
||||||
|
* both the local snapshot and the broker's own queue depth are empty, or ~10s elapse.
|
||||||
|
*/
|
||||||
|
@SuppressWarnings("BusyWait")
|
||||||
|
private static void drainUntilEmpty(AmqpReplyInbox inbox, String target, Connection inspect) throws Exception {
|
||||||
|
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(10);
|
||||||
|
while (System.nanoTime() < deadline) {
|
||||||
|
for (ReplyInbox.InboxMessage msg : inbox.peek(target)) {
|
||||||
|
inbox.ack(target, msg.msgId());
|
||||||
|
}
|
||||||
|
try (Channel ch = inspect.createChannel()) {
|
||||||
|
if (ch.queueDeclarePassive(queueName(target)).getMessageCount() == 0 && inbox.peek(target).isEmpty()) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
Thread.sleep(50);
|
||||||
|
}
|
||||||
|
throw new AssertionError("did not drain " + target + " to empty within the deadline");
|
||||||
|
}
|
||||||
|
|
||||||
/** Poll peek (broker delivery is async) until a reply for {@code target} appears or ~10s elapse. */
|
/** Poll peek (broker delivery is async) until a reply for {@code target} appears or ~10s elapse. */
|
||||||
@SuppressWarnings("BusyWait") // deliberate poll for async broker delivery, bounded by the deadline
|
@SuppressWarnings("BusyWait") // deliberate poll for async broker delivery, bounded by the deadline
|
||||||
private static List<ReplyInbox.InboxMessage> awaitPeek(AmqpReplyInbox inbox, String target)
|
private static List<ReplyInbox.InboxMessage> awaitPeek(AmqpReplyInbox inbox, String target)
|
||||||
|
|||||||
@@ -0,0 +1,288 @@
|
|||||||
|
package dev.ltms.bridged.msg;
|
||||||
|
|
||||||
|
import com.rabbitmq.client.AMQP;
|
||||||
|
import com.rabbitmq.client.Channel;
|
||||||
|
import com.rabbitmq.client.ConfirmCallback;
|
||||||
|
import com.rabbitmq.client.Connection;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.Timeout;
|
||||||
|
|
||||||
|
import java.lang.reflect.InvocationHandler;
|
||||||
|
import java.lang.reflect.Proxy;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.concurrent.CopyOnWriteArrayList;
|
||||||
|
import java.util.concurrent.CountDownLatch;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.atomic.AtomicInteger;
|
||||||
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
import java.util.concurrent.atomic.AtomicReference;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-528 follow-up: {@link AmqpReplyInbox#failPendingPublishesOnRecovery()} must not fail a publish
|
||||||
|
* that registers concurrently with (and was not yet visible when) the recovery sweep began — that
|
||||||
|
* would report a publish that actually succeeded as failed, and {@code MessageService.reply} retries
|
||||||
|
* with a fresh {@code msgId}, so the reply is delivered twice. A live broker reconnect cannot be
|
||||||
|
* forced reliably, so this drives {@link AmqpReplyInbox#failPendingPublishesOnRecovery} and a real
|
||||||
|
* {@link AmqpReplyInbox#publish} against each other directly, against fake AMQP channels built with
|
||||||
|
* {@link Proxy} (no mocking library is on the classpath).
|
||||||
|
*
|
||||||
|
* <p>Also covers Finding 2 (CB-528 follow-up): {@link AmqpReplyInbox#close()} must fail an in-flight
|
||||||
|
* publish promptly instead of leaving it to idle out the 10s confirm timeout.
|
||||||
|
*/
|
||||||
|
class AmqpReplyInboxRecoveryRaceTest {
|
||||||
|
|
||||||
|
/** Large enough that thousands of entries are still unprocessed by the time the very first one
|
||||||
|
* is observed as failed (see {@code sweepIsHoldingTheLock} below) — that gap is what makes the
|
||||||
|
* head start deterministic instead of a coin flip. 100,000 gave the same guarantee but made the
|
||||||
|
* test far more expensive than the guarantee needs; the ordering no longer depends on a timing
|
||||||
|
* window sized to the full backlog; just to the tail of it. */
|
||||||
|
private static final int STALE_PUBLISHES = 2_000;
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@Timeout(30)
|
||||||
|
void recoverySweepDoesNotFailAPublishThatRegistersWhileItIsRunning() throws Exception {
|
||||||
|
AtomicLong seqCounter = new AtomicLong();
|
||||||
|
List<Long> seqOrder = new CopyOnWriteArrayList<>();
|
||||||
|
List<String> msgIdOrder = new CopyOnWriteArrayList<>();
|
||||||
|
AtomicReference<ConfirmCallback> ackCallback = new AtomicReference<>();
|
||||||
|
AtomicReference<ConfirmCallback> nackCallback = new AtomicReference<>();
|
||||||
|
|
||||||
|
Channel publishChannel = fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback);
|
||||||
|
Channel consumeChannel = fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback);
|
||||||
|
Connection connection = fakeConnection(consumeChannel, publishChannel);
|
||||||
|
|
||||||
|
AmqpReplyInbox inbox = new AmqpReplyInbox(connection, AmqpReplyInbox.DEFAULT_PREFETCH);
|
||||||
|
|
||||||
|
// STALE_PUBLISHES in-flight publishes that never get confirmed — they sit in pendingBySeq /
|
||||||
|
// pendingByMsgId exactly like publishes whose confirm never arrived before a connection drop.
|
||||||
|
// Virtual threads make this many concurrent blocking publish() calls cheap.
|
||||||
|
CountDownLatch staleStarted = new CountDownLatch(STALE_PUBLISHES);
|
||||||
|
// Counted down by the FIRST stale publish thread to observe its own failure. That can only
|
||||||
|
// happen from inside failPendingPublishesOnRecovery() — nothing else in this test ever
|
||||||
|
// completes a stale Pending exceptionally (no nack/return is simulated for any "stale-*"
|
||||||
|
// msgId) — so seeing it fire is direct, observable proof the sweep is inside its loop, not a
|
||||||
|
// timing guess. It replaces the old fixed Thread.sleep(5) head start.
|
||||||
|
CountDownLatch sweepIsHoldingTheLock = new CountDownLatch(1);
|
||||||
|
for (int i = 0; i < STALE_PUBLISHES; i++) {
|
||||||
|
String msgId = "stale-" + i;
|
||||||
|
Thread.ofVirtual().start(() -> {
|
||||||
|
staleStarted.countDown();
|
||||||
|
try {
|
||||||
|
inbox.publish("worker-stale", msgId, "x");
|
||||||
|
} catch (IllegalStateException expected) {
|
||||||
|
sweepIsHoldingTheLock.countDown();
|
||||||
|
}
|
||||||
|
});
|
||||||
|
}
|
||||||
|
staleStarted.await();
|
||||||
|
// Let the registrations (the synchronized put into pendingBySeq/pendingByMsgId) actually land
|
||||||
|
// for all of them before the sweep starts, so the sweep begins with a large, real backlog.
|
||||||
|
Thread.sleep(300);
|
||||||
|
|
||||||
|
AtomicReference<Throwable> sweepError = new AtomicReference<>();
|
||||||
|
Thread sweepThread = new Thread(() -> {
|
||||||
|
try {
|
||||||
|
inbox.failPendingPublishesOnRecovery();
|
||||||
|
} catch (Throwable t) {
|
||||||
|
sweepError.set(t);
|
||||||
|
}
|
||||||
|
}, "recovery-sweep");
|
||||||
|
sweepThread.start();
|
||||||
|
|
||||||
|
// Deterministic head start: block until the sweep has actually failed one of the stale
|
||||||
|
// publishes. failPendingPublishesOnRecovery() (once guarded, as it is on main) holds
|
||||||
|
// publishChannelLock for its ENTIRE loop, not just per entry — so this failure proves the
|
||||||
|
// sweep is, at this instant, still holding that lock. With STALE_PUBLISHES this large, the
|
||||||
|
// remaining ~1,999 entries give an enormous margin between "first failure observed" and "sweep
|
||||||
|
// releases the lock": there is no window left for "fresh" to slip in before the sweep starts,
|
||||||
|
// or to win the lock ahead of it — see the case-2 note below. This also means Case 1 (the sweep
|
||||||
|
// is already inside its loop, holding the lock, when "fresh" tries to register) is now
|
||||||
|
// guaranteed by construction rather than merely likely under a fixed sleep.
|
||||||
|
assertTrue(sweepIsHoldingTheLock.await(20, TimeUnit.SECONDS),
|
||||||
|
"the sweep never failed a single stale publish — it may not have started");
|
||||||
|
|
||||||
|
// Case 2 ("fresh" wins publishChannelLock before the sweep even starts, so it genuinely
|
||||||
|
// published on the stale channel and the sweep correctly fails it) is impossible by
|
||||||
|
// construction in this test: freshThread.start() below is reached only after
|
||||||
|
// sweepIsHoldingTheLock has counted down, which can only happen once
|
||||||
|
// failPendingPublishesOnRecovery() is already running and has already failed a stale entry.
|
||||||
|
// There is no code path that lets "fresh" start before the sweep starts. That case is real
|
||||||
|
// and correct production behaviour (see AmqpReplyInbox#failPendingPublishesOnRecovery's
|
||||||
|
// javadoc), it is just not reachable from this deterministic ordering, so it does not need a
|
||||||
|
// separate assertion here.
|
||||||
|
|
||||||
|
// This is the exact interleaving CB-528's follow-up describes: "the still-running recovery
|
||||||
|
// sweep" racing a publish that registers while it is mid-flight.
|
||||||
|
AtomicReference<Throwable> publishError = new AtomicReference<>();
|
||||||
|
Thread freshThread = new Thread(() -> {
|
||||||
|
try {
|
||||||
|
inbox.publish("worker-fresh", "fresh", "hello");
|
||||||
|
} catch (Throwable t) {
|
||||||
|
publishError.set(t);
|
||||||
|
}
|
||||||
|
}, "fresh-publish");
|
||||||
|
freshThread.start();
|
||||||
|
|
||||||
|
sweepThread.join(20_000);
|
||||||
|
assertNull(sweepError.get(), "sweep threw: " + sweepError.get());
|
||||||
|
|
||||||
|
// Simulate the broker's real confirm for "fresh" now that the sweep is done, so a correct
|
||||||
|
// implementation's publish() returns normally instead of idling out CONFIRM_TIMEOUT_MS. Poll
|
||||||
|
// for the registration rather than checking once: freshThread may still be contending for
|
||||||
|
// publishChannelLock (behind the sweep's own hold on it, and possibly other stale threads
|
||||||
|
// still unwinding) even though the sweep itself has already finished.
|
||||||
|
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(9);
|
||||||
|
int idx = -1;
|
||||||
|
while (idx < 0 && System.nanoTime() < deadline) {
|
||||||
|
idx = msgIdOrder.indexOf("fresh");
|
||||||
|
if (idx < 0) {
|
||||||
|
Thread.sleep(20);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (idx >= 0 && ackCallback.get() != null) {
|
||||||
|
ackCallback.get().handle(seqOrder.get(idx), false);
|
||||||
|
}
|
||||||
|
freshThread.join(15_000);
|
||||||
|
|
||||||
|
assertNull(publishError.get(),
|
||||||
|
"a publish that registered while the recovery sweep was running must not be failed by "
|
||||||
|
+ "it, but got: " + publishError.get());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@Timeout(15)
|
||||||
|
void closeFailsInFlightPublishPromptlyInsteadOfWaitingOutTheConfirmTimeout() throws Exception {
|
||||||
|
AtomicLong seqCounter = new AtomicLong();
|
||||||
|
List<Long> seqOrder = new CopyOnWriteArrayList<>();
|
||||||
|
List<String> msgIdOrder = new CopyOnWriteArrayList<>();
|
||||||
|
AtomicReference<ConfirmCallback> ackCallback = new AtomicReference<>();
|
||||||
|
AtomicReference<ConfirmCallback> nackCallback = new AtomicReference<>();
|
||||||
|
|
||||||
|
Channel publishChannel = fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback);
|
||||||
|
Channel consumeChannel = fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback);
|
||||||
|
Connection connection = fakeConnection(consumeChannel, publishChannel);
|
||||||
|
|
||||||
|
AmqpReplyInbox inbox = new AmqpReplyInbox(connection, AmqpReplyInbox.DEFAULT_PREFETCH);
|
||||||
|
|
||||||
|
CountDownLatch publishReturned = new CountDownLatch(1);
|
||||||
|
AtomicReference<Throwable> publishError = new AtomicReference<>();
|
||||||
|
AtomicLong elapsedMillis = new AtomicLong();
|
||||||
|
Thread publishThread = new Thread(() -> {
|
||||||
|
long start = System.nanoTime();
|
||||||
|
try {
|
||||||
|
inbox.publish("worker-close", "never-confirmed", "x");
|
||||||
|
} catch (Throwable t) {
|
||||||
|
publishError.set(t);
|
||||||
|
} finally {
|
||||||
|
elapsedMillis.set((System.nanoTime() - start) / 1_000_000);
|
||||||
|
publishReturned.countDown();
|
||||||
|
}
|
||||||
|
}, "publish-during-close");
|
||||||
|
publishThread.start();
|
||||||
|
|
||||||
|
Thread.sleep(200); // let publish() register before close() runs
|
||||||
|
inbox.close();
|
||||||
|
|
||||||
|
assertTrue(publishReturned.await(5, TimeUnit.SECONDS), "publish() did not return after close()");
|
||||||
|
assertNotNull(publishError.get(), "a publish in flight when close() runs must fail, not hang");
|
||||||
|
assertTrue(publishError.get().getMessage() != null
|
||||||
|
&& publishError.get().getMessage().toLowerCase().contains("closed"),
|
||||||
|
"expected a clear closed-inbox message, got: " + publishError.get());
|
||||||
|
assertTrue(elapsedMillis.get() < 5_000,
|
||||||
|
"close() should fail the in-flight publish promptly, not wait out the confirm timeout — took "
|
||||||
|
+ elapsedMillis.get() + "ms");
|
||||||
|
}
|
||||||
|
|
||||||
|
/** A {@link Proxy}-backed {@link Channel}: only the calls {@link AmqpReplyInbox} actually makes
|
||||||
|
* are meaningfully implemented; everything else returns a harmless default. */
|
||||||
|
private static Channel fakeChannel(AtomicLong seqCounter, List<Long> seqOrder, List<String> msgIdOrder,
|
||||||
|
AtomicReference<ConfirmCallback> ackCallback,
|
||||||
|
AtomicReference<ConfirmCallback> nackCallback) {
|
||||||
|
InvocationHandler handler = (proxy, method, args) -> {
|
||||||
|
String name = method.getName();
|
||||||
|
if (name.equals("getNextPublishSeqNo")) {
|
||||||
|
long value = seqCounter.incrementAndGet();
|
||||||
|
seqOrder.add(value);
|
||||||
|
return value;
|
||||||
|
}
|
||||||
|
if (name.equals("basicPublish")) {
|
||||||
|
AMQP.BasicProperties props = (AMQP.BasicProperties) args[3];
|
||||||
|
msgIdOrder.add(props.getMessageId());
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
if (name.equals("addConfirmListener")) {
|
||||||
|
ackCallback.set((ConfirmCallback) args[0]);
|
||||||
|
nackCallback.set((ConfirmCallback) args[1]);
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
if (name.equals("equals")) {
|
||||||
|
return proxy == args[0];
|
||||||
|
}
|
||||||
|
if (name.equals("hashCode")) {
|
||||||
|
return System.identityHashCode(proxy);
|
||||||
|
}
|
||||||
|
if (name.equals("toString")) {
|
||||||
|
return "FakeChannel";
|
||||||
|
}
|
||||||
|
return defaultValue(method.getReturnType());
|
||||||
|
};
|
||||||
|
return (Channel) Proxy.newProxyInstance(AmqpReplyInboxRecoveryRaceTest.class.getClassLoader(),
|
||||||
|
new Class<?>[] {Channel.class}, handler);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** A {@link Proxy}-backed {@link Connection} handing out {@code first} then {@code second} from
|
||||||
|
* successive {@code createChannel()} calls, matching {@link AmqpReplyInbox}'s constructor. */
|
||||||
|
private static Connection fakeConnection(Channel first, Channel second) {
|
||||||
|
AtomicInteger calls = new AtomicInteger();
|
||||||
|
InvocationHandler handler = (proxy, method, args) -> {
|
||||||
|
String name = method.getName();
|
||||||
|
if (name.equals("createChannel") && (args == null || args.length == 0)) {
|
||||||
|
return calls.getAndIncrement() == 0 ? first : second;
|
||||||
|
}
|
||||||
|
if (name.equals("equals")) {
|
||||||
|
return proxy == args[0];
|
||||||
|
}
|
||||||
|
if (name.equals("hashCode")) {
|
||||||
|
return System.identityHashCode(proxy);
|
||||||
|
}
|
||||||
|
if (name.equals("toString")) {
|
||||||
|
return "FakeConnection";
|
||||||
|
}
|
||||||
|
return defaultValue(method.getReturnType());
|
||||||
|
};
|
||||||
|
return (Connection) Proxy.newProxyInstance(AmqpReplyInboxRecoveryRaceTest.class.getClassLoader(),
|
||||||
|
new Class<?>[] {Connection.class}, handler);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static Object defaultValue(Class<?> type) {
|
||||||
|
if (!type.isPrimitive() || type == void.class) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
if (type == boolean.class) {
|
||||||
|
return Boolean.FALSE;
|
||||||
|
}
|
||||||
|
if (type == long.class) {
|
||||||
|
return 0L;
|
||||||
|
}
|
||||||
|
if (type == short.class) {
|
||||||
|
return (short) 0;
|
||||||
|
}
|
||||||
|
if (type == byte.class) {
|
||||||
|
return (byte) 0;
|
||||||
|
}
|
||||||
|
if (type == char.class) {
|
||||||
|
return (char) 0;
|
||||||
|
}
|
||||||
|
if (type == double.class) {
|
||||||
|
return 0.0d;
|
||||||
|
}
|
||||||
|
if (type == float.class) {
|
||||||
|
return 0.0f;
|
||||||
|
}
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -6,6 +6,7 @@ import dev.ltms.bridged.herdr.FakeHerdr;
|
|||||||
import dev.ltms.bridged.herdr.HerdrException;
|
import dev.ltms.bridged.herdr.HerdrException;
|
||||||
import dev.ltms.bridged.inject.CompletionResolver;
|
import dev.ltms.bridged.inject.CompletionResolver;
|
||||||
import dev.ltms.bridged.inject.ExhaustedPatternLookup;
|
import dev.ltms.bridged.inject.ExhaustedPatternLookup;
|
||||||
|
import dev.ltms.bridged.inject.ExhaustionSink;
|
||||||
import dev.ltms.bridged.mcp.PrimaryRegistry;
|
import dev.ltms.bridged.mcp.PrimaryRegistry;
|
||||||
import dev.ltms.bridged.inject.Injector;
|
import dev.ltms.bridged.inject.Injector;
|
||||||
import org.junit.jupiter.api.BeforeEach;
|
import org.junit.jupiter.api.BeforeEach;
|
||||||
@@ -33,7 +34,7 @@ class MessageServiceTest {
|
|||||||
private final AgentControl agents = new AgentControl(herdr);
|
private final AgentControl agents = new AgentControl(herdr);
|
||||||
private final Rendezvous rendezvous = new Rendezvous();
|
private final Rendezvous rendezvous = new Rendezvous();
|
||||||
private final CompletionResolver completion =
|
private final CompletionResolver completion =
|
||||||
new CompletionResolver(agents, rendezvous, ExhaustedPatternLookup.none());
|
new CompletionResolver(agents, rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||||
private final Injector injector = new Injector(agents, completion);
|
private final Injector injector = new Injector(agents, completion);
|
||||||
private final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
private final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||||
private final MessageService messages = new MessageService(agents, injector, rendezvous, inbox);
|
private final MessageService messages = new MessageService(agents, injector, rendezvous, inbox);
|
||||||
@@ -713,6 +714,355 @@ class MessageServiceTest {
|
|||||||
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
|
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- CB-582: bridge_status pendingAsk() ------------------------------------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void pendingAskReturnsNullWhenNoQuestionIsOpen() throws Exception {
|
||||||
|
assertNull(messages.pendingAsk(T), "no async ticket at all -> no pending ask");
|
||||||
|
|
||||||
|
String ticket = messages.sendAsync(T, "long task");
|
||||||
|
awaitWaiting();
|
||||||
|
assertNull(messages.pendingAsk(T), "a plain pending delegation is not a question");
|
||||||
|
|
||||||
|
injectDelivery();
|
||||||
|
assertTrue(rendezvous.resolve(T, "done"));
|
||||||
|
awaitTicketPhase(ticket, MessageService.Phase.DONE);
|
||||||
|
assertNull(messages.pendingAsk(T), "a finished ticket carries no open question either");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void pendingAskReturnsTheOpenQuestionForAnAsyncTicket() throws Exception {
|
||||||
|
String ticket = messages.sendAsync(T, "task that asks");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
|
||||||
|
CompletableFuture<MessageService.AskResult> ask =
|
||||||
|
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||||
|
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
|
||||||
|
|
||||||
|
MessageService.PendingAsk pending = messages.pendingAsk(T);
|
||||||
|
assertNotNull(pending, "bridge_status should see the open question");
|
||||||
|
assertEquals(ticket, pending.ticket());
|
||||||
|
assertEquals("which config file?", pending.question());
|
||||||
|
assertEquals(asking.turnId(), pending.turnId());
|
||||||
|
|
||||||
|
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||||
|
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
|
||||||
|
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||||
|
assertNull(messages.pendingAsk(T), "an answered question is no longer pending");
|
||||||
|
awaitWaiting();
|
||||||
|
assertTrue(rendezvous.resolve(T, "done"));
|
||||||
|
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- CB-588: async ticket terminal nudges ---------------------------------------------------
|
||||||
|
//
|
||||||
|
// MessageService.reply's rendezvous fast path is exactly what an async ticket always takes
|
||||||
|
// (sendAsync registers a rendezvous waiter — see asyncTasksByWaiter), so it never reached
|
||||||
|
// ReplyPushLoop.onReplyQueued. These prove the ticket reaches ReplyPushLoop through the new
|
||||||
|
// onTicketTerminal entry point instead, with no bridge_poll from the lead first.
|
||||||
|
|
||||||
|
private static final String LEAD = "term_lead";
|
||||||
|
|
||||||
|
/** A MessageService wired to a real ReplyPushLoop pointed at a private herdr fake for the lead. */
|
||||||
|
private record PushWiring(MessageService service, FakeHerdr leadHerdr,
|
||||||
|
java.util.concurrent.ScheduledExecutorService scheduler) implements AutoCloseable {
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
scheduler.shutdownNow();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private PushWiring wireWithPushLoop(int maxReminders, long backoffMs) {
|
||||||
|
return wireWithPushLoop(maxReminders, backoffMs, System::nanoTime);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** As above, with an injectable clock (CB-588 follow-up: exercise pruneTerminalTickets' TTL). */
|
||||||
|
private PushWiring wireWithPushLoop(int maxReminders, long backoffMs, java.util.function.LongSupplier nowNanos) {
|
||||||
|
PrimaryRegistry registry = new PrimaryRegistry(null);
|
||||||
|
registry.recordDelegation(T, LEAD);
|
||||||
|
FakeHerdr leadHerdr = new FakeHerdr();
|
||||||
|
AgentControl leadAgents = new AgentControl(leadHerdr);
|
||||||
|
var scheduler = java.util.concurrent.Executors.newSingleThreadScheduledExecutor();
|
||||||
|
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, leadAgents, inbox, scheduler, maxReminders, backoffMs);
|
||||||
|
MessageService service = new MessageService(agents, injector, rendezvous, inbox, pushLoop, null, nowNanos);
|
||||||
|
return new PushWiring(service, leadHerdr, scheduler);
|
||||||
|
}
|
||||||
|
|
||||||
|
private void awaitNudge(FakeHerdr leadHerdr) throws InterruptedException {
|
||||||
|
long deadline = System.currentTimeMillis() + 3000;
|
||||||
|
while (!leadHerdr.called("agent.prompt") && System.currentTimeMillis() < deadline) {
|
||||||
|
Thread.sleep(10);
|
||||||
|
}
|
||||||
|
assertTrue(leadHerdr.called("agent.prompt"), "expected a nudge in the lead's pane");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void anAsyncTicketThatFinishesNudgesTheLeadWithNoPriorPollCall() throws Exception {
|
||||||
|
try (var wiring = wireWithPushLoop(1, 50)) {
|
||||||
|
String ticket = wiring.service().sendAsync(T, "long task");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
assertTrue(rendezvous.resolve(T, "async result"));
|
||||||
|
|
||||||
|
awaitNudge(wiring.leadHerdr());
|
||||||
|
String nudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
|
||||||
|
assertTrue(nudge.contains(ticket), "the nudge should name the ticket: " + nudge);
|
||||||
|
assertTrue(nudge.contains("bridge_poll(ticket="),
|
||||||
|
"the nudge should name the exact ticket-collecting call: " + nudge);
|
||||||
|
assertFalse(nudge.toUpperCase().contains("FAILED"),
|
||||||
|
"a successfully-replied ticket's nudge must not say it failed: " + nudge);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aFailedAsyncTicketAlsoNudgesAndSaysSo() throws Exception {
|
||||||
|
try (var wiring = wireWithPushLoop(1, 50)) {
|
||||||
|
wiring.service().sendAsync(T, "long task");
|
||||||
|
awaitWaiting();
|
||||||
|
wiring.service().abandon(T, "session released"); // a terminal failure phase
|
||||||
|
|
||||||
|
awaitNudge(wiring.leadHerdr());
|
||||||
|
String nudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
|
||||||
|
assertTrue(nudge.toUpperCase().contains("FAILED"), "a failed ticket's nudge must say so: " + nudge);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void anAlreadyCollectedTicketProducesNoNudge() throws Exception {
|
||||||
|
try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: poll before the first tick fires
|
||||||
|
String ticket = wiring.service().sendAsync(T, "long task");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
assertTrue(rendezvous.resolve(T, "async result"));
|
||||||
|
|
||||||
|
MessageService.TaskView view = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.DONE);
|
||||||
|
assertEquals(MessageService.Phase.DONE, view.phase());
|
||||||
|
|
||||||
|
Thread.sleep(400); // let the scheduled tick run — it must find nothing pending
|
||||||
|
assertFalse(wiring.leadHerdr().called("agent.prompt"),
|
||||||
|
"a ticket the lead already polled must never be nudged");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void severalAsyncTicketsFinishingTogetherProduceOneCoalescedNudge() throws Exception {
|
||||||
|
try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: both tickets land before the tick fires
|
||||||
|
String first = wiring.service().sendAsync(T, "first task");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
assertTrue(rendezvous.resolve(T, "first done"));
|
||||||
|
// Settle without polling: poll() itself marks a ticket collected (that's the point of
|
||||||
|
// anAlreadyCollectedTicketProducesNoNudge above) — using it here to detect completion
|
||||||
|
// would collect the ticket before the coalescing this test checks ever gets a chance.
|
||||||
|
Thread.sleep(100);
|
||||||
|
|
||||||
|
String second = wiring.service().sendAsync(T, "second task");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
assertTrue(rendezvous.resolve(T, "second done"));
|
||||||
|
Thread.sleep(100);
|
||||||
|
|
||||||
|
awaitNudge(wiring.leadHerdr());
|
||||||
|
Thread.sleep(200); // settle — nothing more should arrive beyond the one coalesced nudge
|
||||||
|
long nudgeCount = wiring.leadHerdr().calls.stream()
|
||||||
|
.filter(c -> c.method().equals("agent.prompt")).count();
|
||||||
|
assertEquals(1, nudgeCount, "two tickets finishing together must produce ONE nudge, not two");
|
||||||
|
String nudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
|
||||||
|
assertTrue(nudge.contains(first) && nudge.contains(second),
|
||||||
|
"the coalesced nudge should name both tickets: " + nudge);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- CB-582: bridge_ask question-open nudges --------------------------------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void anAsyncTicketThatPausesOnAQuestionNudgesTheLeadWithNoPriorPollCall() throws Exception {
|
||||||
|
try (var wiring = wireWithPushLoop(5, 50)) {
|
||||||
|
String ticket = wiring.service().sendAsync(T, "task that asks");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
|
||||||
|
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
|
||||||
|
() -> wiring.service().ask(T, "which config file?", 5000));
|
||||||
|
MessageService.TaskView asking = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.ASKING);
|
||||||
|
|
||||||
|
awaitNudge(wiring.leadHerdr());
|
||||||
|
String nudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
|
||||||
|
assertTrue(nudge.contains(ticket), "the nudge should name the ticket: " + nudge);
|
||||||
|
assertTrue(nudge.contains(asking.turnId()), "the nudge should name the turnId: " + nudge);
|
||||||
|
assertTrue(nudge.contains("bridge_send(turnId="),
|
||||||
|
"the nudge should name the exact answer call: " + nudge);
|
||||||
|
assertTrue(nudge.contains("which config file?"), "the nudge should include the question: " + nudge);
|
||||||
|
|
||||||
|
// Clean up the still-open ask so the background thread does not linger past the test.
|
||||||
|
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||||
|
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
|
||||||
|
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||||
|
awaitWaiting();
|
||||||
|
assertTrue(rendezvous.resolve(T, "done"));
|
||||||
|
answer.get(5, TimeUnit.SECONDS);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void answeringAQuestionStopsFurtherNudgesAboutIt() throws Exception {
|
||||||
|
try (var wiring = wireWithPushLoop(5, 50)) {
|
||||||
|
String ticket = wiring.service().sendAsync(T, "task that asks");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
|
||||||
|
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
|
||||||
|
() -> wiring.service().ask(T, "which config file?", 5000));
|
||||||
|
MessageService.TaskView asking = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.ASKING);
|
||||||
|
|
||||||
|
awaitNudge(wiring.leadHerdr());
|
||||||
|
long callsBeforeAnswer = wiring.leadHerdr().calls.stream()
|
||||||
|
.filter(c -> c.method().equals("agent.prompt")).count();
|
||||||
|
|
||||||
|
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||||
|
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
|
||||||
|
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||||
|
awaitWaiting();
|
||||||
|
assertTrue(rendezvous.resolve(T, "done"));
|
||||||
|
answer.get(5, TimeUnit.SECONDS);
|
||||||
|
|
||||||
|
// Let several more ticks (and the ticket's own now-legitimate terminal nudge) fire —
|
||||||
|
// none of them may still name the question's turnId, which is closed.
|
||||||
|
Thread.sleep(300);
|
||||||
|
boolean anyNamesClosedQuestion = wiring.leadHerdr().calls.stream()
|
||||||
|
.filter(c -> c.method().equals("agent.prompt"))
|
||||||
|
.skip(callsBeforeAnswer)
|
||||||
|
.anyMatch(c -> c.params().toString().contains(asking.turnId()));
|
||||||
|
assertFalse(anyNamesClosedQuestion,
|
||||||
|
"no nudge sent after the answer may still name the now-closed turnId " + asking.turnId());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void anAskThatLeavesByThrowingStillClosesItsQuestion() throws Exception {
|
||||||
|
// CB-582 follow-up. ask() calls clearAsyncQuestion on three paths — no-waiter, timed out,
|
||||||
|
// and (from answer()) answered — but it can also leave by *throwing*: an interrupt while
|
||||||
|
// blocked on the answer, or an ExecutionException from the answer future. Those paths run
|
||||||
|
// only the finally block, so before the fix the push loop kept the question pending for
|
||||||
|
// good: named in every nudge until its own cap, then never removed from the map at all.
|
||||||
|
PrimaryRegistry registry = new PrimaryRegistry(null);
|
||||||
|
registry.recordDelegation(T, LEAD);
|
||||||
|
var scheduler = java.util.concurrent.Executors.newSingleThreadScheduledExecutor();
|
||||||
|
// A backoff far longer than the test: the schedule is started but no tick ever fires, so
|
||||||
|
// decide() is read directly and nothing here depends on timing.
|
||||||
|
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, new AgentControl(new FakeHerdr()), inbox,
|
||||||
|
scheduler, 5, 60_000);
|
||||||
|
MessageService service = new MessageService(agents, injector, rendezvous, inbox, pushLoop);
|
||||||
|
try {
|
||||||
|
String ticket = service.sendAsync(T, "task that asks");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
|
||||||
|
Thread asker = new Thread(() -> assertThrows(IllegalStateException.class,
|
||||||
|
() -> service.ask(T, "which config file?", 30_000)));
|
||||||
|
asker.start();
|
||||||
|
awaitTicketPhaseOn(service, ticket, MessageService.Phase.ASKING);
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, pushLoop.decide(LEAD, 0, 0, 0),
|
||||||
|
"the open question should be the one thing keeping this lead's schedule alive");
|
||||||
|
|
||||||
|
asker.interrupt();
|
||||||
|
asker.join(5000);
|
||||||
|
assertFalse(asker.isAlive(), "the interrupted ask should have left ask() by throwing");
|
||||||
|
|
||||||
|
assertEquals(ReplyPushLoop.Action.STOP, pushLoop.decide(LEAD, 0, 0, 0),
|
||||||
|
"an ask that threw must still close its question, or the loop nudges about it for good");
|
||||||
|
} finally {
|
||||||
|
service.close();
|
||||||
|
scheduler.shutdownNow();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aFleetWithNoPushLoopConfiguredBehavesExactlyAsToday() throws Exception {
|
||||||
|
// `messages` (the shared field) uses the no-pushLoop constructor — poll() must not throw,
|
||||||
|
// and no nudge mechanism exists to fire regardless of how the ticket resolves.
|
||||||
|
String ticket = messages.sendAsync(T, "long task");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
assertTrue(rendezvous.resolve(T, "async result"));
|
||||||
|
|
||||||
|
MessageService.TaskView view = awaitTicketPhase(ticket, MessageService.Phase.DONE);
|
||||||
|
assertEquals(MessageService.Phase.DONE, view.phase());
|
||||||
|
assertEquals("async result", view.reply());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-588 follow-up: {@code tasks} is the sole authority on whether a ticket exists, and
|
||||||
|
* {@code pruneTerminalTickets} drops entries from it once {@link MessageService#TICKET_TTL_NANOS}
|
||||||
|
* elapses. Before this test, that prune never told {@code ReplyPushLoop} — its own
|
||||||
|
* {@code pendingTickets} entry for a pruned, never-collected ticket had no remover at all, so it
|
||||||
|
* rode along on every later nudge to the same lead, naming a ticket {@code bridge_poll} could no
|
||||||
|
* longer find. Uses the injectable clock (mirroring {@code SessionManager}'s {@code nowNanos} seam
|
||||||
|
* for its idle reaper) to cross the 10-minute TTL without a real wait.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void aPrunedTicketIsReclaimedFromThePushLoopNotLeakedForever() throws Exception {
|
||||||
|
java.util.concurrent.atomic.AtomicLong clock = new java.util.concurrent.atomic.AtomicLong(1_000_000_000L);
|
||||||
|
// maxReminders=1 + a short backoff: the stale ticket gets its one legitimate reminder, then
|
||||||
|
// decideTickets hits the cap and STOPs — activeLeads drops the lead, but (before the fix)
|
||||||
|
// pendingTickets never drops the ticket. That is the exact "cap already STOPped" branch of
|
||||||
|
// the bug report, reached deterministically rather than by timing it against a live tick.
|
||||||
|
try (var wiring = wireWithPushLoop(1, 50, clock::get)) {
|
||||||
|
String stale = wiring.service().sendAsync(T, "first task");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
assertTrue(rendezvous.resolve(T, "stale result"));
|
||||||
|
|
||||||
|
// Let the reminder loop fire its one nudge and hit the cap (STOP removes it from
|
||||||
|
// activeLeads; pendingTickets is untouched either way — that asymmetry is the bug).
|
||||||
|
awaitNudge(wiring.leadHerdr());
|
||||||
|
Thread.sleep(300);
|
||||||
|
assertTrue(wiring.leadHerdr().lastCall("agent.prompt").params().toString().contains(stale),
|
||||||
|
"sanity: the stale ticket's own reminder must have fired first");
|
||||||
|
|
||||||
|
// Cross the TTL — a real clock would need 10 minutes; the injected one does it instantly.
|
||||||
|
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
|
||||||
|
|
||||||
|
// A second, unrelated ticket to the same target/lead reaches sendAsync, which prunes.
|
||||||
|
String fresh = wiring.service().sendAsync(T, "second task");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
assertTrue(rendezvous.resolve(T, "fresh result"));
|
||||||
|
|
||||||
|
// The fresh ticket restarts the (now-dormant) reminder loop with its own nudge.
|
||||||
|
long before = wiring.leadHerdr().calls.stream().filter(c -> c.method().equals("agent.prompt")).count();
|
||||||
|
long deadline = System.currentTimeMillis() + 3000;
|
||||||
|
while (wiring.leadHerdr().calls.stream().filter(c -> c.method().equals("agent.prompt")).count() <= before
|
||||||
|
&& System.currentTimeMillis() < deadline) {
|
||||||
|
Thread.sleep(10);
|
||||||
|
}
|
||||||
|
String latestNudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
|
||||||
|
assertTrue(latestNudge.contains(fresh), "the fresh ticket's nudge must still arrive: " + latestNudge);
|
||||||
|
assertFalse(latestNudge.contains(stale),
|
||||||
|
"a pruned ticket must never be named in a later nudge — it is gone and bridge_poll "
|
||||||
|
+ "on it would return nothing: " + latestNudge);
|
||||||
|
|
||||||
|
// And bridge_poll(ticket=stale) really does return nothing now — the nudge would have lied.
|
||||||
|
assertNull(wiring.service().poll(stale), "the pruned ticket must actually be gone, not just unmentioned");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private MessageService.TaskView awaitTicketPhaseOn(MessageService svc, String ticket,
|
||||||
|
MessageService.Phase phase) throws Exception {
|
||||||
|
long deadline = System.currentTimeMillis() + 3000;
|
||||||
|
MessageService.TaskView view;
|
||||||
|
do {
|
||||||
|
view = svc.poll(ticket);
|
||||||
|
if (view.phase() == phase) {
|
||||||
|
return view;
|
||||||
|
}
|
||||||
|
Thread.sleep(5);
|
||||||
|
} while (System.currentTimeMillis() < deadline);
|
||||||
|
assertEquals(phase, view.phase());
|
||||||
|
return view;
|
||||||
|
}
|
||||||
|
|
||||||
private void assertFailedTicket(String ticket, String reason) throws Exception {
|
private void assertFailedTicket(String ticket, String reason) throws Exception {
|
||||||
MessageService.TaskView view = awaitTicketPhase(ticket, MessageService.Phase.FAILED);
|
MessageService.TaskView view = awaitTicketPhase(ticket, MessageService.Phase.FAILED);
|
||||||
assertEquals(reason, view.detail());
|
assertEquals(reason, view.detail());
|
||||||
|
|||||||
@@ -15,10 +15,12 @@ import java.util.ArrayList;
|
|||||||
import java.util.Collections;
|
import java.util.Collections;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
|
import java.util.Set;
|
||||||
import java.util.concurrent.CountDownLatch;
|
import java.util.concurrent.CountDownLatch;
|
||||||
import java.util.concurrent.Executors;
|
import java.util.concurrent.Executors;
|
||||||
import java.util.concurrent.ScheduledExecutorService;
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
import java.util.concurrent.TimeUnit;
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.atomic.AtomicInteger;
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.*;
|
import static org.junit.jupiter.api.Assertions.*;
|
||||||
|
|
||||||
@@ -26,6 +28,12 @@ import static org.junit.jupiter.api.Assertions.*;
|
|||||||
* Unit tests for {@link ReplyPushLoop}: decision logic, nudge injection, idempotency,
|
* Unit tests for {@link ReplyPushLoop}: decision logic, nudge injection, idempotency,
|
||||||
* bounded reminders, and stop conditions.
|
* bounded reminders, and stop conditions.
|
||||||
*
|
*
|
||||||
|
* <p>CB-590 collapsed the CB-307 reply-nudge schedule and the CB-588 ticket-nudge schedule into
|
||||||
|
* one schedule per lead ({@link ReplyPushLoop#decide}), so most tests below register pending work
|
||||||
|
* through the public entry points ({@code onReplyQueued} / {@code onTicketTerminal}) before
|
||||||
|
* exercising {@code decide} directly, mirroring how the two entry points now share one decision
|
||||||
|
* function keyed by the lead terminal rather than by worker target.
|
||||||
|
*
|
||||||
* <p>Uses a {@link RecordingHerdrClient} that synchronizes access to its call list so the
|
* <p>Uses a {@link RecordingHerdrClient} that synchronizes access to its call list so the
|
||||||
* scheduler thread and test thread never have memory ordering issues. The {@code decide()}
|
* scheduler thread and test thread never have memory ordering issues. The {@code decide()}
|
||||||
* tests use a simple client with no concurrency concern.
|
* tests use a simple client with no concurrency concern.
|
||||||
@@ -34,6 +42,7 @@ class ReplyPushLoopTest {
|
|||||||
|
|
||||||
private static final String PRIMARY = "term_primary";
|
private static final String PRIMARY = "term_primary";
|
||||||
private static final String WORKER = "term_worker";
|
private static final String WORKER = "term_worker";
|
||||||
|
private static final String WORKER2 = "term_worker2";
|
||||||
private static final ObjectMapper MAPPER = new ObjectMapper();
|
private static final ObjectMapper MAPPER = new ObjectMapper();
|
||||||
|
|
||||||
private PrimaryRegistry registry;
|
private PrimaryRegistry registry;
|
||||||
@@ -57,38 +66,50 @@ class ReplyPushLoopTest {
|
|||||||
// --- decide() logic ------------------------------------------------------------------------
|
// --- decide() logic ------------------------------------------------------------------------
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideWithoutPrimaryIsStop() {
|
void onReplyQueuedWithNoKnownLeadNeverStartsASchedule() throws Exception {
|
||||||
agents = agentWithStatus("idle");
|
var rec = recordingClient();
|
||||||
var loop = new ReplyPushLoop(
|
agents = new AgentControl(rec);
|
||||||
new PrimaryRegistry(null), agents, inbox, scheduler, 5, 100);
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(WORKER, 0));
|
var loop = new ReplyPushLoop(new PrimaryRegistry(null), agents, inbox, scheduler, 5, 50);
|
||||||
|
|
||||||
|
loop.onReplyQueued(WORKER); // no lead known -> never registered, never scheduled
|
||||||
|
|
||||||
|
assertFalse(loop.isActive(), "no lead known means nothing to nudge yet");
|
||||||
|
Thread.sleep(150);
|
||||||
|
assertEquals(0, rec.sendCount(), "must not nudge when no lead is known to be waiting");
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideWithEmptyInboxIsStop() {
|
void decideWithNothingPendingIsStop() {
|
||||||
agents = agentWithStatus("idle");
|
agents = agentWithStatus("idle");
|
||||||
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
|
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(PRIMARY, 0, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideAtCapIsStop() {
|
void decideAtCapIsStop() {
|
||||||
agents = agentWithStatus("idle");
|
agents = agentWithStatus("idle");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100).decide(WORKER, 2));
|
var loop = loop(2, 100);
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 2, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideUnderCapWithInjectablePrimaryIsInject() {
|
void decideUnderCapWithInjectablePrimaryIsInject() {
|
||||||
agents = agentWithStatus("idle");
|
agents = agentWithStatus("idle");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
|
var loop = loop();
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideUnderCapWithBlockedPrimaryIsInject() {
|
void decideUnderCapWithBlockedPrimaryIsInject() {
|
||||||
agents = agentWithStatus("blocked");
|
agents = agentWithStatus("blocked");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
|
var loop = loop();
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0),
|
||||||
"BLOCKED is injectable");
|
"BLOCKED is injectable");
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -96,7 +117,9 @@ class ReplyPushLoopTest {
|
|||||||
void decideUnderCapWithDonePrimaryIsInject() {
|
void decideUnderCapWithDonePrimaryIsInject() {
|
||||||
agents = agentWithStatus("done");
|
agents = agentWithStatus("done");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
|
var loop = loop();
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0),
|
||||||
"DONE is injectable");
|
"DONE is injectable");
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -104,23 +127,29 @@ class ReplyPushLoopTest {
|
|||||||
void decideUnderCapWithBusyPrimaryIsWaitBusy() {
|
void decideUnderCapWithBusyPrimaryIsWaitBusy() {
|
||||||
agents = agentWithStatus("working");
|
agents = agentWithStatus("working");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
|
var loop = loop();
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decide(PRIMARY, 0, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideUnderCapWithUnknownPrimaryIsWaitBusy() {
|
void decideUnderCapWithUnknownPrimaryIsWaitBusy() {
|
||||||
agents = agentWithStatus("unknown");
|
agents = agentWithStatus("unknown");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
|
var loop = loop();
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decide(PRIMARY, 0, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideStopsAfterInboxIsEmptied() {
|
void decideStopsAfterInboxIsEmptied() {
|
||||||
agents = agentWithStatus("idle");
|
agents = agentWithStatus("idle");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
|
var loop = loop();
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0));
|
||||||
inbox.ack(WORKER, "m1");
|
inbox.ack(WORKER, "m1");
|
||||||
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- onReplyQueued integration -------------------------------------------------------------
|
// --- onReplyQueued integration -------------------------------------------------------------
|
||||||
@@ -200,6 +229,560 @@ class ReplyPushLoopTest {
|
|||||||
assertTrue(nudge.contains("bridge_poll(target=term_worker)"));
|
assertTrue(nudge.contains("bridge_poll(target=term_worker)"));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void repliesNudgeFormatIsCorrect() {
|
||||||
|
String multi = ReplyPushLoop.REPLIES_NUDGE_FORMAT.formatted(2, "term_worker1, term_worker2");
|
||||||
|
assertTrue(multi.contains("2 workers"));
|
||||||
|
assertTrue(multi.contains("bridge_poll(target=...)"));
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- CB-588: async ticket terminal nudges — decide() logic on tickets -----------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void decideTicketsWithNothingPendingIsStop() {
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(PRIMARY, 0, 0));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void decideTicketsAtCapIsStop() {
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
var loop = loop(2, 100_000);
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 2));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void decideTicketsUnderCapWithInjectableLeadIsInject() {
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
var loop = loop(5, 100_000);
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void decideTicketsUnderCapWithBusyLeadIsWaitBusy() {
|
||||||
|
agents = agentWithStatus("working");
|
||||||
|
var loop = loop(5, 100_000);
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decide(PRIMARY, 0, 0));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void decideTicketsWithoutAKnownLeadIsStop() {
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
var loop = new ReplyPushLoop(new PrimaryRegistry(null), agents, inbox, scheduler, 5, 100_000);
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false); // no lead known -> never registered as pending
|
||||||
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 0));
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- CB-588: onTicketTerminal integration ---------------------------------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void onTicketTerminalCausesExactlyOneNudgeNamingTheTicketAndThePollCall() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
|
||||||
|
loop(1, 50).onTicketTerminal("task-1", WORKER, false);
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one ticket nudge should have been sent");
|
||||||
|
assertEquals(1, rec.sendCount());
|
||||||
|
String nudge = rec.sentParams().getFirst().getValue().toString();
|
||||||
|
assertTrue(nudge.contains("task-1"), "nudge should name the ticket");
|
||||||
|
assertTrue(nudge.contains("bridge_poll(ticket="), "nudge should name the exact ticket-poll call");
|
||||||
|
assertFalse(nudge.contains("bridge_poll(target="), "a ticket-only nudge must not tell the lead to run the target-poll call");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aFailedTicketNudgeSaysItWasAFailure() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
|
||||||
|
loop(1, 50).onTicketTerminal("task-1", WORKER, true);
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS));
|
||||||
|
String nudge = rec.sentParams().getFirst().getValue().toString();
|
||||||
|
assertTrue(nudge.toUpperCase().contains("FAILED"), "a failed ticket's nudge must say so: " + nudge);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void severalTicketsFinishingTogetherProduceOneCoalescedNudge() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
var loop = loop(1, 300); // backoff wide enough that both onTicketTerminal calls land first
|
||||||
|
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
loop.onTicketTerminal("task-2", WORKER, true); // arrives while the schedule is already active
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one coalesced nudge should have been sent");
|
||||||
|
assertEquals(1, rec.sendCount(), "two tickets finishing together must produce ONE nudge, not two");
|
||||||
|
String nudge = rec.sentParams().getFirst().getValue().toString();
|
||||||
|
assertTrue(nudge.contains("task-1") && nudge.contains("task-2"),
|
||||||
|
"the coalesced nudge should name both tickets: " + nudge);
|
||||||
|
assertTrue(nudge.contains("2 tickets"), "the coalesced nudge should name the count: " + nudge);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void onTicketTerminalIsIdempotentPerLeadWhileActive() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
rec.sendLatch = new CountDownLatch(1);
|
||||||
|
|
||||||
|
var loop = loop(1, 100);
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false); // duplicate — should not start a second schedule
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS));
|
||||||
|
Thread.sleep(200);
|
||||||
|
assertEquals(1, rec.sendCount(), "a duplicate onTicketTerminal for the same lead must not double-nudge");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aTicketAlreadyCollectedProducesNoNudge() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
var loop = loop(1, 100);
|
||||||
|
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
loop.ticketCollected("task-1"); // the lead polled before the first tick fired
|
||||||
|
|
||||||
|
Thread.sleep(300); // let the scheduled tick run
|
||||||
|
assertEquals(0, rec.sendCount(), "an already-collected ticket must never be nudged");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void ticketNudgesSendUpToCapThenStop() throws Exception {
|
||||||
|
int cap = 2;
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
rec.sendLatch = new CountDownLatch(cap);
|
||||||
|
|
||||||
|
loop(cap, 50).onTicketTerminal("task-1", WORKER, false);
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(5, TimeUnit.SECONDS), cap + " ticket nudges should have fired");
|
||||||
|
Thread.sleep(300);
|
||||||
|
assertEquals(cap, rec.sendCount(), "exactly " + cap + " ticket nudges (cap=" + cap + ")");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void isActiveReflectsALiveTicketScheduleForTheHeartbeatStandDown() {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
ReplyPushLoop loop = loop(1, 100_000); // long backoff so the tick cannot fire mid-test
|
||||||
|
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
assertTrue(loop.isActive(), "an active ticket-reminder schedule must also stand off the heartbeat");
|
||||||
|
|
||||||
|
loop.stop();
|
||||||
|
assertFalse(loop.isActive(), "stopping clears the active ticket schedule too");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aTicketStillPendingWhenTheLoopStopsIsNotStranded() {
|
||||||
|
// Regression for the race a reviewer found in gitea PR #73: onTicketTerminal's
|
||||||
|
// activeLeads.putIfAbsent can see the lead's slot as still occupied a moment before
|
||||||
|
// decide's STOP releases it, so the ticket coalesces onto a schedule that is about to
|
||||||
|
// die and nothing ever nudges about it. Forcing that exact thread interleaving is not
|
||||||
|
// reliable, so this drives stopOrRestart — the STOP path's own release-and-recheck —
|
||||||
|
// directly, arranging the state it must not lose a ticket in: a ticket pending for the lead
|
||||||
|
// that was NOT part of the pre-decision snapshot (ticketsBefore=empty), standing in for one
|
||||||
|
// that races in during the decision-to-release window.
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
ReplyPushLoop loop = loop(1, 100_000); // long backoff — no natural tick fires during this test
|
||||||
|
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false); // pendingTickets={task-1}; activeLeads={PRIMARY}
|
||||||
|
|
||||||
|
// Stand in for the scheduler thread reaching decide==STOP for this lead — with nothing
|
||||||
|
// pending at decide time — while task-1 races in before the release below runs.
|
||||||
|
loop.stopOrRestart(PRIMARY, Set.of(), Set.of());
|
||||||
|
|
||||||
|
assertTrue(loop.isActive(), "a ticket that raced the loop's stop must reclaim the schedule "
|
||||||
|
+ "slot, not be stranded with no schedule left to ever nudge about it");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aStaleUncollectedTicketAtCapDoesNotRestartTheLoop() {
|
||||||
|
// The other direction of the same fix: restarting on ANY non-empty pending set would be
|
||||||
|
// wrong. When STOP is reached because the reminder cap was hit, the same never-collected
|
||||||
|
// ticket is expected to still be there — that is the cap doing its job (acceptance criterion
|
||||||
|
// #5: no spin / nudges stay bounded). task-1 here was already accounted for at decide time
|
||||||
|
// (it is in ticketsBefore), so it must not restart the loop just because it is still sitting
|
||||||
|
// there.
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
ReplyPushLoop loop = loop(1, 100_000);
|
||||||
|
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false); // pendingTickets={task-1}; activeLeads={PRIMARY}
|
||||||
|
|
||||||
|
loop.stopOrRestart(PRIMARY, Set.of(), Set.of("task-1"));
|
||||||
|
|
||||||
|
assertFalse(loop.isActive(), "a stale ticket already accounted for at decide time must not "
|
||||||
|
+ "restart the loop — that would defeat the reminder cap");
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- mirror of the two stopOrRestart tests above, for the reply arm ------------------------
|
||||||
|
//
|
||||||
|
// Both tests above only ever passed Set.of() for repliesBefore, so racedIn's reply branch
|
||||||
|
// (`pendingReplyTargetsFor(lead).stream().anyMatch(t -> !repliesBefore.contains(t))`) was
|
||||||
|
// never exercised by anything other than an always-empty snapshot. The reviewer flagged this:
|
||||||
|
// racedIn is symmetric in the code, and only half of it was pinned by a test.
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aReplyStillPendingWhenTheLoopStopsIsNotStranded() {
|
||||||
|
// Mirrors aTicketStillPendingWhenTheLoopStopsIsNotStranded: a reply target that raced in
|
||||||
|
// during the decision-to-release window (absent from the "before" snapshot) must reclaim
|
||||||
|
// the schedule slot rather than being stranded with no schedule left to nudge about it.
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
ReplyPushLoop loop = loop(1, 100_000); // long backoff — no natural tick fires during this test
|
||||||
|
|
||||||
|
loop.onReplyQueued(WORKER); // pendingReplies={term_worker}; activeLeads={PRIMARY}
|
||||||
|
|
||||||
|
loop.stopOrRestart(PRIMARY, Set.of(), Set.of());
|
||||||
|
|
||||||
|
assertTrue(loop.isActive(), "a reply that raced the loop's stop must reclaim the schedule "
|
||||||
|
+ "slot, not be stranded with no schedule left to ever nudge about it");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aStaleUncollectedReplyAtCapDoesNotRestartTheLoop() {
|
||||||
|
// Mirrors aStaleUncollectedTicketAtCapDoesNotRestartTheLoop: a reply target already
|
||||||
|
// accounted for at decide time (present in repliesBefore) must not restart the loop —
|
||||||
|
// that is the reminder cap doing its job, not a race.
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
ReplyPushLoop loop = loop(1, 100_000);
|
||||||
|
|
||||||
|
loop.onReplyQueued(WORKER); // pendingReplies={term_worker}; activeLeads={PRIMARY}
|
||||||
|
|
||||||
|
loop.stopOrRestart(PRIMARY, Set.of(WORKER), Set.of());
|
||||||
|
|
||||||
|
assertFalse(loop.isActive(), "a stale reply target already accounted for at decide time "
|
||||||
|
+ "must not restart the loop — that would defeat the reminder cap");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void ticketNudgeFormatIsCorrect() {
|
||||||
|
String single = ReplyPushLoop.TICKET_NUDGE_FORMAT.formatted("task-1", "", "task-1");
|
||||||
|
assertTrue(single.contains("Ticket task-1"));
|
||||||
|
assertTrue(single.contains("bridge_poll(ticket=task-1)"));
|
||||||
|
|
||||||
|
String multi = ReplyPushLoop.TICKETS_NUDGE_FORMAT.formatted(2, "", "task-1, task-2");
|
||||||
|
assertTrue(multi.contains("2 tickets"));
|
||||||
|
assertTrue(multi.contains("bridge_poll(ticket=...)"));
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- CB-307 nudge path is unchanged (regression) --------------------------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void inboxNudgeStillUsesTheOriginalTargetPollCall() {
|
||||||
|
String nudge = ReplyPushLoop.NUDGE_FORMAT.formatted(WORKER, WORKER);
|
||||||
|
assertTrue(nudge.contains("bridge_poll(target=" + WORKER + ")"),
|
||||||
|
"CB-588/CB-590 must not change the CB-307 inbox nudge's call shape");
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- CB-590: one schedule per lead — no overlap, no lost nudges -----------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void replyAndTicketForTheSameLeadCoalesceIntoOneSendNeverOverlapping() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
var loop = loop(1, 300); // backoff wide enough that both entry points land before the first tick
|
||||||
|
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one combined nudge should have been sent");
|
||||||
|
Thread.sleep(300);
|
||||||
|
assertEquals(1, rec.sendCount(),
|
||||||
|
"a reply and a ticket for the same lead must coalesce onto ONE schedule — "
|
||||||
|
+ "two nudge injections into the same lead pane must never overlap");
|
||||||
|
String nudge = rec.sentParams().getFirst().getValue().toString();
|
||||||
|
assertTrue(nudge.contains("bridge_poll(target=" + WORKER + ")"),
|
||||||
|
"the combined nudge must still mention the reply: " + nudge);
|
||||||
|
assertTrue(nudge.contains("task-1"), "the combined nudge must still mention the ticket: " + nudge);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aReplyQueuedWhileTheLeadIsBusyIsNotLostWhenATicketArrivesToo() throws Exception {
|
||||||
|
// The lead is busy for its first two status checks, then becomes injectable. A reply is
|
||||||
|
// queued while busy; a ticket for the same lead arrives before the lead frees up. Neither
|
||||||
|
// may be dropped — deferred is fine, lost is not (acceptance criterion #2).
|
||||||
|
var rec = new BusyThenIdleHerdrClient(2);
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
var loop = loop(1, 50); // cap=1: WAIT_BUSY doesn't count against it, so exactly one send once injectable
|
||||||
|
|
||||||
|
loop.onReplyQueued(WORKER); // schedule starts, first tick(s) WAIT_BUSY
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false); // coalesces onto the same waiting schedule
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
|
||||||
|
"once the lead becomes injectable, the deferred work must still be nudged");
|
||||||
|
Thread.sleep(200);
|
||||||
|
assertEquals(1, rec.sendCount(), "exactly one nudge once injectable — reply and ticket coalesced");
|
||||||
|
String nudge = rec.sentParams().getFirst().getValue().toString();
|
||||||
|
assertTrue(nudge.contains("bridge_poll(target=" + WORKER + ")"), "the reply must not be dropped: " + nudge);
|
||||||
|
assertTrue(nudge.contains("task-1"), "the ticket must not be dropped: " + nudge);
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- CB-590 follow-up: per-source reminder budgets — the regression this round exists for ---
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void oneExhaustedSourceDoesNotBlockANudgeForTheOtherSource() {
|
||||||
|
// The live trace this ticket was filed from: an undrained reply target got nudged up to
|
||||||
|
// its cap (5 reminders), then a ticket for the SAME lead went terminal shortly before the
|
||||||
|
// next scheduled tick — so it coalesced onto the still-active schedule (arriving BEFORE
|
||||||
|
// that tick's "before" snapshot, not during the decision-to-release race stopOrRestart
|
||||||
|
// guards). With CB-590's single shared reminder counter, that tick's decide() saw
|
||||||
|
// reminderCount already at the cap and returned STOP regardless of the ticket, and because
|
||||||
|
// the ticket was already present in that tick's "before" snapshot, stopOrRestart's
|
||||||
|
// racedIn check (proven correct on its own above) did not save it either — it is not a
|
||||||
|
// race, it looks like ordinary stale backlog. The ticket was then stranded: pending
|
||||||
|
// forever with no live schedule, never named in any nudge.
|
||||||
|
//
|
||||||
|
// Fixed by giving each source its own counter. Here the reply source is AT its cap (2/2)
|
||||||
|
// and the ticket source has NEVER been nudged (0/2) — decide() must still return INJECT,
|
||||||
|
// because the ticket is still eligible on its own budget.
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
var loop = loop(2, 100_000); // huge backoff — this test drives decide()/isActive() directly
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 2, 0),
|
||||||
|
"the reply source is exhausted (2/2), but the ticket source has never been "
|
||||||
|
+ "nudged (0/2) — the lead must still be injected so the ticket is not "
|
||||||
|
+ "lost, exactly the CB-590 follow-up regression");
|
||||||
|
|
||||||
|
// isActive() (criterion #5): a real tick that takes the INJECT branch above never calls
|
||||||
|
// stopOrRestart, so the schedule started by onTicketTerminal above stays live — the
|
||||||
|
// ticket is not left stranded with isActive()==false while it is still pending.
|
||||||
|
assertTrue(loop.isActive(), "the schedule must stay active while the ticket source still "
|
||||||
|
+ "has budget left, even though the reply source sharing it is exhausted");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void bothSourcesExhaustedIsStillStop() {
|
||||||
|
// The flip side: per-source budgets must not turn into unbounded nudging. When BOTH
|
||||||
|
// sources are at their cap, decide() must still STOP — a per-source budget is still a
|
||||||
|
// budget.
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
var loop = loop(2, 100_000);
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
|
||||||
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 2, 2),
|
||||||
|
"both the reply and the ticket source are at their own cap — must still stop");
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- CB-598: work arriving during a backoff must not read as stale backlog -------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aTargetArrivingDuringTheBackoffGetsNudgedDespiteAnAlreadyCappedSibling() {
|
||||||
|
// The bug: reminder counts used to be a single counter per lead per source, carried
|
||||||
|
// forward across scheduled ticks (scheduleNext(lead, count + 1, ...)) rather than tracked
|
||||||
|
// per pending item. WORKER gets nudged once here, which — with cap=1 — exhausts the
|
||||||
|
// shared reply-source counter for this lead. WORKER2 then queues a reply for the SAME
|
||||||
|
// lead "during the backoff": while the schedule from WORKER's tick is still active, before
|
||||||
|
// the next tick's own start-of-tick snapshot runs. At that next tick, the OLD code passed
|
||||||
|
// the already-exhausted shared counter into decide() regardless of WORKER2 never having
|
||||||
|
// been named in any nudge, and — because WORKER2 was already present in that tick's
|
||||||
|
// "before" snapshot — stopOrRestart's race check (proven correct on its own elsewhere in
|
||||||
|
// this file) does not save it either: it looks like ordinary stale backlog, not a race.
|
||||||
|
// WORKER2 was then stranded forever with no live schedule and no nudge ever naming it.
|
||||||
|
//
|
||||||
|
// tick() is driven directly (package-private, same reasoning as stopOrRestart being
|
||||||
|
// directly testable) so the exact interleaving is deterministic instead of racing the
|
||||||
|
// scheduler thread over a real ~15s backoff.
|
||||||
|
//
|
||||||
|
// Before the fix, this test fails on the second assertEquals: rec.sendCount() stays at 1
|
||||||
|
// (decide() returns STOP on the second tick(), so injectNudge is never called a second
|
||||||
|
// time) and the "must still get one" assertion never even runs.
|
||||||
|
int cap = 1;
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
inbox.own(WORKER2);
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
var loop = loop(cap, 100_000); // huge backoff — nothing fires on its own; we drive tick()
|
||||||
|
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
loop.tick(PRIMARY); // first tick: nudges WORKER alone; WORKER's own count reaches the cap
|
||||||
|
assertEquals(1, rec.sendCount(), "the first tick should nudge about WORKER");
|
||||||
|
|
||||||
|
// WORKER2 "arrives during the backoff": queued for the same lead while the schedule from
|
||||||
|
// the tick above is still active (activeLeads still holds PRIMARY), before the next tick
|
||||||
|
// (simulated below) takes its own start-of-tick snapshot.
|
||||||
|
inbox.publish(WORKER2, "m2", "hello2");
|
||||||
|
loop.onReplyQueued(WORKER2);
|
||||||
|
|
||||||
|
loop.tick(PRIMARY); // the tick that would fire once that backoff elapsed
|
||||||
|
|
||||||
|
assertEquals(2, rec.sendCount(),
|
||||||
|
"WORKER2 was never named in any nudge yet and must still get one, even though "
|
||||||
|
+ "WORKER's own reminder count is already at the cap");
|
||||||
|
String secondNudge = rec.sentParams().get(1).getValue().toString();
|
||||||
|
assertTrue(secondNudge.contains(WORKER2), "the never-named target must be named: " + secondNudge);
|
||||||
|
|
||||||
|
// Criterion #3: isActive() must reflect that this lead still had a live nudge to give —
|
||||||
|
// the second tick took the INJECT branch, so the schedule stayed live rather than being
|
||||||
|
// torn down under WORKER2.
|
||||||
|
assertTrue(loop.isActive(), "the schedule must stay active after nudging the fresh target");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aTicketArrivingDuringTheBackoffGetsNudgedDespiteAnAlreadyCappedSibling() {
|
||||||
|
// Mirrors the reply-side test above for the ticket source.
|
||||||
|
int cap = 1;
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
var loop = loop(cap, 100_000);
|
||||||
|
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
loop.tick(PRIMARY); // first tick: nudges task-1 alone; its count reaches the cap
|
||||||
|
assertEquals(1, rec.sendCount(), "the first tick should nudge about task-1");
|
||||||
|
|
||||||
|
loop.onTicketTerminal("task-2", WORKER, false); // arrives during the backoff, same lead
|
||||||
|
loop.tick(PRIMARY);
|
||||||
|
|
||||||
|
assertEquals(2, rec.sendCount(),
|
||||||
|
"task-2 was never named in any nudge yet and must still get one, even though "
|
||||||
|
+ "task-1's reminder count is already at the cap");
|
||||||
|
String secondNudge = rec.sentParams().get(1).getValue().toString();
|
||||||
|
assertTrue(secondNudge.contains("task-2"), "the never-named ticket must be named: " + secondNudge);
|
||||||
|
assertTrue(loop.isActive(), "the schedule must stay active after nudging the fresh ticket");
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- CB-582: bridge_ask question-open nudges -------------------------------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void onQuestionOpenedWithNoKnownLeadNeverStartsASchedule() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
var loop = new ReplyPushLoop(new PrimaryRegistry(null), agents, inbox, scheduler, 5, 50);
|
||||||
|
|
||||||
|
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
|
||||||
|
|
||||||
|
assertFalse(loop.isActive(), "no lead known means nothing to nudge yet");
|
||||||
|
Thread.sleep(150);
|
||||||
|
assertEquals(0, rec.sendCount(), "must not nudge when no lead is known to be waiting");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void decideQuestionsWithNothingPendingIsStop() {
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(PRIMARY, 0, 0, 0));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void decideQuestionsAtCapIsStop() {
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
var loop = loop(2, 100_000);
|
||||||
|
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
|
||||||
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 0, 2));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void decideQuestionsUnderCapWithInjectableLeadIsInject() {
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
var loop = loop(5, 100_000);
|
||||||
|
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0, 0));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void onQuestionOpenedCausesExactlyOneNudgeNamingTheTicketAndTurnId() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
|
||||||
|
loop(1, 50).onQuestionOpened("task-1", WORKER, "term_worker#1", "which config file?");
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one question nudge should have been sent");
|
||||||
|
assertEquals(1, rec.sendCount());
|
||||||
|
String nudge = rec.sentParams().getFirst().getValue().toString();
|
||||||
|
assertTrue(nudge.contains("task-1"), "nudge should name the ticket: " + nudge);
|
||||||
|
assertTrue(nudge.contains("term_worker#1"), "nudge should name the turnId: " + nudge);
|
||||||
|
assertTrue(nudge.contains("bridge_send(turnId="), "nudge should name the exact answer call: " + nudge);
|
||||||
|
assertTrue(nudge.contains("which config file?"), "nudge should include the question text: " + nudge);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void questionClosedPreventsFurtherNudging() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
var loop = loop(1, 100);
|
||||||
|
|
||||||
|
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
|
||||||
|
loop.questionClosed("term_worker#1"); // answered/lapsed before the first tick fired
|
||||||
|
|
||||||
|
Thread.sleep(300); // let the scheduled tick run
|
||||||
|
assertEquals(0, rec.sendCount(), "an already-closed question must never be nudged");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void questionNudgesSendUpToCapThenStop() throws Exception {
|
||||||
|
int cap = 2;
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
rec.sendLatch = new CountDownLatch(cap);
|
||||||
|
|
||||||
|
loop(cap, 50).onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(5, TimeUnit.SECONDS), cap + " question nudges should have fired");
|
||||||
|
Thread.sleep(300);
|
||||||
|
assertEquals(cap, rec.sendCount(), "exactly " + cap + " question nudges (cap=" + cap + ")");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void questionAndTicketForTheSameLeadCoalesceIntoOneSend() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
var loop = loop(1, 300); // backoff wide enough that both entry points land before the first tick
|
||||||
|
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
loop.onQuestionOpened("task-2", WORKER, "term_worker#1", "which config?");
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one combined nudge should have been sent");
|
||||||
|
Thread.sleep(300);
|
||||||
|
assertEquals(1, rec.sendCount(),
|
||||||
|
"a ticket and a question for the same lead must coalesce onto ONE schedule");
|
||||||
|
String nudge = rec.sentParams().getFirst().getValue().toString();
|
||||||
|
assertTrue(nudge.contains("task-1"), "the combined nudge must still mention the ticket: " + nudge);
|
||||||
|
assertTrue(nudge.contains("term_worker#1"), "the combined nudge must still mention the question: " + nudge);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void oneExhaustedQuestionSourceDoesNotBlockANudgeForTheOtherSources() {
|
||||||
|
// Mirrors oneExhaustedSourceDoesNotBlockANudgeForTheOtherSource for the question source:
|
||||||
|
// the question source is at its cap (2/2), but the ticket source has never been nudged
|
||||||
|
// (0/2) — decide() must still INJECT so the ticket is not stranded.
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
var loop = loop(2, 100_000);
|
||||||
|
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
|
||||||
|
loop.onTicketTerminal("task-2", WORKER, false);
|
||||||
|
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0, 2),
|
||||||
|
"the question source is exhausted (2/2), but the ticket source has never been "
|
||||||
|
+ "nudged (0/2) — the lead must still be injected so the ticket is not lost");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void questionNudgeFormatIsCorrect() {
|
||||||
|
String single = ReplyPushLoop.QUESTION_NUDGE_FORMAT.formatted(
|
||||||
|
WORKER, "task-1", "term_worker#1", "which config?");
|
||||||
|
assertTrue(single.contains("Worker term_worker"));
|
||||||
|
assertTrue(single.contains("bridge_send(turnId=\"term_worker#1\""));
|
||||||
|
assertTrue(single.contains("which config?"));
|
||||||
|
|
||||||
|
String multi = ReplyPushLoop.QUESTIONS_NUDGE_FORMAT.formatted(2, "task-1 (turnId=t1), task-2 (turnId=t2)");
|
||||||
|
assertTrue(multi.contains("2 workers"));
|
||||||
|
assertTrue(multi.contains("bridge_poll(ticket=...)"));
|
||||||
|
}
|
||||||
|
|
||||||
// --- metrics (CB-512) ----------------------------------------------------------------------
|
// --- metrics (CB-512) ----------------------------------------------------------------------
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -225,14 +808,43 @@ class ReplyPushLoopTest {
|
|||||||
agents = agentWithStatus("idle");
|
agents = agentWithStatus("idle");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
Metrics metrics = new Metrics();
|
Metrics metrics = new Metrics();
|
||||||
|
var loop = loop(2, 100, metrics);
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
|
||||||
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100, metrics).decide(WORKER, 2));
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 2, 0));
|
||||||
|
|
||||||
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
|
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
|
||||||
"hitting the reminder cap must count as exhausted");
|
"hitting the reminder cap must count as exhausted");
|
||||||
assertEquals(0, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"));
|
assertEquals(0, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void successfulTicketNudgeIncrementsDelivered() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
Metrics metrics = new Metrics();
|
||||||
|
|
||||||
|
loop(1, 50, metrics).onTicketTerminal("task-1", WORKER, false);
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one ticket nudge should have been sent");
|
||||||
|
Thread.sleep(200);
|
||||||
|
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"),
|
||||||
|
"a successfully sent ticket nudge must count as delivered, same metric as CB-307");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void ticketReminderCapIncrementsExhausted() {
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
Metrics metrics = new Metrics();
|
||||||
|
var loop = loop(2, 100_000, metrics);
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
|
||||||
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 2));
|
||||||
|
|
||||||
|
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
|
||||||
|
"hitting the ticket reminder cap must count as exhausted");
|
||||||
|
}
|
||||||
|
|
||||||
// --- helpers -------------------------------------------------------------------------------
|
// --- helpers -------------------------------------------------------------------------------
|
||||||
|
|
||||||
private ReplyPushLoop loop() {
|
private ReplyPushLoop loop() {
|
||||||
@@ -315,4 +927,49 @@ class ReplyPushLoopTest {
|
|||||||
private static RecordingHerdrClient recordingClient() {
|
private static RecordingHerdrClient recordingClient() {
|
||||||
return new RecordingHerdrClient();
|
return new RecordingHerdrClient();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Thread-safe fake that reports {@code working} (not injectable) for its first
|
||||||
|
* {@code busyChecks} status calls, then {@code idle} forever after — used to prove work queued
|
||||||
|
* while the lead is busy is deferred, not dropped, once it becomes injectable.
|
||||||
|
*/
|
||||||
|
private static final class BusyThenIdleHerdrClient implements HerdrClient {
|
||||||
|
private final List<Map.Entry<String, Object>> calls =
|
||||||
|
Collections.synchronizedList(new ArrayList<>());
|
||||||
|
private final AtomicInteger statusChecks = new AtomicInteger();
|
||||||
|
private final int busyChecks;
|
||||||
|
volatile CountDownLatch sendLatch = new CountDownLatch(1);
|
||||||
|
|
||||||
|
BusyThenIdleHerdrClient(int busyChecks) {
|
||||||
|
this.busyChecks = busyChecks;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public JsonNode call(String method, Object params) {
|
||||||
|
if ("agent.get".equals(method)) {
|
||||||
|
String status = statusChecks.getAndIncrement() < busyChecks ? "working" : "idle";
|
||||||
|
return MAPPER.createObjectNode()
|
||||||
|
.set("agent", MAPPER.createObjectNode()
|
||||||
|
.put("terminal_id", PRIMARY)
|
||||||
|
.put("agent_status", status));
|
||||||
|
}
|
||||||
|
if ("agent.prompt".equals(method)) {
|
||||||
|
calls.add(Map.entry(method, params));
|
||||||
|
sendLatch.countDown();
|
||||||
|
}
|
||||||
|
return MAPPER.createObjectNode();
|
||||||
|
}
|
||||||
|
|
||||||
|
long sendCount() {
|
||||||
|
return calls.size();
|
||||||
|
}
|
||||||
|
|
||||||
|
List<Map.Entry<String, Object>> sentParams() {
|
||||||
|
return List.copyOf(calls);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
}
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,119 @@
|
|||||||
|
package dev.ltms.bridged.placement;
|
||||||
|
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.OptionalLong;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-578 stage B: the credential-keyed quarantine tracker itself, isolated from placement/spawn
|
||||||
|
* wiring (that's {@code CompositePeerLauncherTest}). The clock is a plain {@link AtomicLong} of
|
||||||
|
* nanos so expiry is exercised without a real sleep.
|
||||||
|
*/
|
||||||
|
class BackendQuarantineTest {
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aFreshCredentialIsNotQuarantined() {
|
||||||
|
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||||
|
assertFalse(q.isQuarantined("shared-openai"));
|
||||||
|
assertEquals(OptionalLong.empty(), q.remainingSeconds("shared-openai"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void quarantineBlocksTheCredentialForTheFullCooldown() {
|
||||||
|
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||||
|
q.quarantine("shared-openai");
|
||||||
|
|
||||||
|
assertTrue(q.isQuarantined("shared-openai"));
|
||||||
|
assertEquals(OptionalLong.of(1800L), q.remainingSeconds("shared-openai"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void onlyTheQuarantinedCredentialIsAffected() {
|
||||||
|
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||||
|
q.quarantine("shared-openai");
|
||||||
|
|
||||||
|
assertFalse(q.isQuarantined("some-other-credential"),
|
||||||
|
"an unrelated credential must not be swept into the quarantine");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void expiresOnTheInjectedClock() {
|
||||||
|
AtomicLong now = new AtomicLong(0L);
|
||||||
|
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.MINUTES.toNanos(30));
|
||||||
|
q.quarantine("shared-openai");
|
||||||
|
assertTrue(q.isQuarantined("shared-openai"));
|
||||||
|
|
||||||
|
now.set(TimeUnit.MINUTES.toNanos(29));
|
||||||
|
assertTrue(q.isQuarantined("shared-openai"), "still inside the cooldown");
|
||||||
|
|
||||||
|
now.set(TimeUnit.MINUTES.toNanos(31));
|
||||||
|
assertFalse(q.isQuarantined("shared-openai"), "the cooldown has elapsed on the injected clock");
|
||||||
|
assertEquals(OptionalLong.empty(), q.remainingSeconds("shared-openai"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aRepeatQuarantineCallRestartsTheCooldownAtFullLength() {
|
||||||
|
AtomicLong now = new AtomicLong(0L);
|
||||||
|
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.MINUTES.toNanos(30));
|
||||||
|
q.quarantine("shared-openai");
|
||||||
|
|
||||||
|
now.set(TimeUnit.MINUTES.toNanos(20));
|
||||||
|
q.quarantine("shared-openai");
|
||||||
|
|
||||||
|
now.set(TimeUnit.MINUTES.toNanos(45)); // 25 min after the second call, 45 after the first
|
||||||
|
assertTrue(q.isQuarantined("shared-openai"),
|
||||||
|
"a fresh exhaustion resets the cooldown to full length, not the earlier shorter wait");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void activeRemainingSecondsListsOnlyStillQuarantinedCredentials() {
|
||||||
|
AtomicLong now = new AtomicLong(0L);
|
||||||
|
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.MINUTES.toNanos(30));
|
||||||
|
q.quarantine("shared-openai");
|
||||||
|
q.quarantine("another-credential");
|
||||||
|
|
||||||
|
now.set(TimeUnit.MINUTES.toNanos(31));
|
||||||
|
q.quarantine("shared-openai"); // re-quarantined after the first one expired
|
||||||
|
|
||||||
|
Map<String, Long> active = q.activeRemainingSeconds();
|
||||||
|
assertEquals(Map.of("shared-openai", 1800L), active,
|
||||||
|
"the expired credential is dropped; the re-quarantined one is reported");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void noneReportsNothingQuarantinedWhenNeverToldTo() {
|
||||||
|
BackendQuarantine q = BackendQuarantine.none();
|
||||||
|
|
||||||
|
assertFalse(q.isQuarantined("anything"));
|
||||||
|
assertTrue(q.activeRemainingSeconds().isEmpty());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void noneIgnoresAQuarantineCallInsteadOfLockingTheCredentialForever() {
|
||||||
|
// .none() holds a clock frozen at 0, so if #quarantine recorded a deadline the credential
|
||||||
|
// would never expire — locked out for the life of the daemon. Two production
|
||||||
|
// CompositePeerLauncher constructors default to none(), so that failure would be silent and
|
||||||
|
// permanent. An inert stand-in must omit the fact, never invent one.
|
||||||
|
BackendQuarantine q = BackendQuarantine.none();
|
||||||
|
|
||||||
|
q.quarantine("shared-openai");
|
||||||
|
|
||||||
|
assertFalse(q.isQuarantined("shared-openai"), "none() must not quarantine anything");
|
||||||
|
assertTrue(q.remainingSeconds("shared-openai").isEmpty());
|
||||||
|
assertTrue(q.activeRemainingSeconds().isEmpty());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aNonPositiveCooldownIsRejected() {
|
||||||
|
assertThrows(IllegalArgumentException.class, () -> new BackendQuarantine(() -> 0L, 0L));
|
||||||
|
assertThrows(IllegalArgumentException.class, () -> new BackendQuarantine(() -> 0L, -1L));
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -24,7 +24,13 @@ class PlacementPolicyTest {
|
|||||||
private static PlacementContext ctx(List<PlacementCandidate> candidates,
|
private static PlacementContext ctx(List<PlacementCandidate> candidates,
|
||||||
Function<String, Integer> liveCount,
|
Function<String, Integer> liveCount,
|
||||||
Set<String> unreachable) {
|
Set<String> unreachable) {
|
||||||
return new PlacementContext("b", candidates, liveCount, unreachable);
|
return ctx(candidates, liveCount, unreachable, Set.of());
|
||||||
|
}
|
||||||
|
|
||||||
|
private static PlacementContext ctx(List<PlacementCandidate> candidates,
|
||||||
|
Function<String, Integer> liveCount,
|
||||||
|
Set<String> unreachable, Set<String> quarantined) {
|
||||||
|
return new PlacementContext("b", candidates, liveCount, unreachable, quarantined);
|
||||||
}
|
}
|
||||||
|
|
||||||
private static PlacementContext ctx(List<PlacementCandidate> candidates,
|
private static PlacementContext ctx(List<PlacementCandidate> candidates,
|
||||||
@@ -46,17 +52,37 @@ class PlacementPolicyTest {
|
|||||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||||
PlacementContext ctx = new PlacementContext(null,
|
PlacementContext ctx = new PlacementContext(null,
|
||||||
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
||||||
noSessions(), Set.of());
|
noSessions(), Set.of(), Set.of());
|
||||||
assertEquals("a", policy.select(ctx).profile());
|
assertEquals("a", policy.select(ctx).profile());
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void fixedThrowsWhenNoProfilesAndNoDefault() {
|
void fixedThrowsWhenNoProfilesAndNoDefault() {
|
||||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||||
PlacementContext ctx = new PlacementContext(null, List.of(), noSessions(), Set.of());
|
PlacementContext ctx = new PlacementContext(null, List.of(), noSessions(), Set.of(), Set.of());
|
||||||
assertThrows(PlacementException.class, () -> policy.select(ctx));
|
assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void fixedSkipsQuarantinedDefault() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||||
|
PlacementContext ctx = new PlacementContext("b",
|
||||||
|
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
||||||
|
noSessions(), Set.of(), Set.of("b"));
|
||||||
|
assertEquals("a", policy.select(ctx).profile(),
|
||||||
|
"the default 'b' is quarantined, so fixed falls through to the first un-quarantined candidate");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void fixedThrowsWhenDefaultAndEveryCandidateQuarantined() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||||
|
PlacementContext ctx = new PlacementContext("b",
|
||||||
|
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
||||||
|
noSessions(), Set.of(), Set.of("a", "b"));
|
||||||
|
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||||
|
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void roundRobinCyclesThroughAvailableProfiles() {
|
void roundRobinCyclesThroughAvailableProfiles() {
|
||||||
PlacementPolicy policy = PlacementPolicies.roundRobin();
|
PlacementPolicy policy = PlacementPolicies.roundRobin();
|
||||||
@@ -82,6 +108,18 @@ class PlacementPolicyTest {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void roundRobinSkipsQuarantinedProfiles() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.roundRobin();
|
||||||
|
List<PlacementCandidate> candidates = List.of(
|
||||||
|
PlacementCandidate.profile("a"),
|
||||||
|
PlacementCandidate.profile("b"));
|
||||||
|
PlacementContext ctx = ctx(candidates, noSessions(), Set.of(), Set.of("a"));
|
||||||
|
for (int i = 0; i < 5; i++) {
|
||||||
|
assertEquals("b", policy.select(ctx).profile(), "a is quarantined, so every pick lands on b");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void roundRobinThrowsWhenAllAtMaxLoad() {
|
void roundRobinThrowsWhenAllAtMaxLoad() {
|
||||||
PlacementPolicy policy = PlacementPolicies.roundRobin();
|
PlacementPolicy policy = PlacementPolicies.roundRobin();
|
||||||
@@ -150,6 +188,29 @@ class PlacementPolicyTest {
|
|||||||
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
|
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void weightedSkipsQuarantinedProfile() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.weighted();
|
||||||
|
List<PlacementCandidate> candidates = List.of(
|
||||||
|
PlacementCandidate.profile("a", 1.0f, null),
|
||||||
|
PlacementCandidate.profile("b", 1.0f, null));
|
||||||
|
PlacementContext ctx = ctx(candidates, noSessions(), Set.of(), Set.of("a"));
|
||||||
|
for (int i = 0; i < 5; i++) {
|
||||||
|
assertEquals("b", policy.select(ctx).profile(), "a is quarantined, so every pick lands on b");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void weightedThrowsWhenAllQuarantined() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.weighted();
|
||||||
|
List<PlacementCandidate> candidates = List.of(
|
||||||
|
PlacementCandidate.profile("a"),
|
||||||
|
PlacementCandidate.profile("b"));
|
||||||
|
PlacementException e = assertThrows(PlacementException.class,
|
||||||
|
() -> policy.select(ctx(candidates, noSessions(), Set.of(), Set.of("a", "b"))));
|
||||||
|
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void weightedThrowsWhenAllUnreachable() {
|
void weightedThrowsWhenAllUnreachable() {
|
||||||
PlacementPolicy policy = PlacementPolicies.weighted();
|
PlacementPolicy policy = PlacementPolicies.weighted();
|
||||||
@@ -176,8 +237,134 @@ class PlacementPolicyTest {
|
|||||||
assertTrue(e.getMessage().contains("1 unreachable"), e.getMessage());
|
assertTrue(e.getMessage().contains("1 unreachable"), e.getMessage());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- CB-554: weight <= 0 excludes a candidate from automatic selection ------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void weightedSkipsWeightZeroProfile() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.weighted();
|
||||||
|
List<PlacementCandidate> candidates = List.of(
|
||||||
|
PlacementCandidate.profile("a", 0.0f, null),
|
||||||
|
PlacementCandidate.profile("b", 1.0f, null));
|
||||||
|
for (int i = 0; i < 5; i++) {
|
||||||
|
assertEquals("b", policy.select(ctx(candidates, noSessions())).profile(),
|
||||||
|
"a has weight 0, so every pick lands on b");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void weightedThrowsWhenAllWeightZero() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.weighted();
|
||||||
|
List<PlacementCandidate> candidates = List.of(
|
||||||
|
PlacementCandidate.profile("a", 0.0f, null),
|
||||||
|
PlacementCandidate.profile("b", 0.0f, null));
|
||||||
|
PlacementException e = assertThrows(PlacementException.class,
|
||||||
|
() -> policy.select(ctx(candidates, noSessions())));
|
||||||
|
assertTrue(e.getMessage().contains("weight 0"), e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void weightedTreatsNegativeWeightAsExcludedNotError() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.weighted();
|
||||||
|
List<PlacementCandidate> candidates = List.of(
|
||||||
|
PlacementCandidate.profile("a", -1.0f, null),
|
||||||
|
PlacementCandidate.profile("b", 1.0f, null));
|
||||||
|
for (int i = 0; i < 5; i++) {
|
||||||
|
assertEquals("b", policy.select(ctx(candidates, noSessions())).profile(),
|
||||||
|
"a has a negative weight, so it is excluded like weight 0 — never picked, never an error");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void weightZeroAndQuarantinedIsNotDoubleCounted() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.weighted();
|
||||||
|
List<PlacementCandidate> candidates = List.of(
|
||||||
|
PlacementCandidate.profile("a", 0.0f, null),
|
||||||
|
PlacementCandidate.profile("b", 1.0f, null),
|
||||||
|
PlacementCandidate.profile("c", 1.0f, null));
|
||||||
|
// a is BOTH weight-0 and quarantined; b and c are quarantined only. If a were counted
|
||||||
|
// into both buckets, "quarantined" would read 3 (or the arithmetic would go negative)
|
||||||
|
// instead of the true count of 2 candidates actually quarantined.
|
||||||
|
PlacementContext mixedCtx = ctx(candidates, noSessions(), Set.of(), Set.of("a", "b", "c"));
|
||||||
|
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(mixedCtx));
|
||||||
|
assertTrue(e.getMessage().contains("1 weight-0"), e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("2 quarantined"), e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("0 remaining"), e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void roundRobinSkipsWeightZeroProfiles() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.roundRobin();
|
||||||
|
List<PlacementCandidate> candidates = List.of(
|
||||||
|
PlacementCandidate.profile("a", 0.0f, null),
|
||||||
|
PlacementCandidate.profile("b", 1.0f, null));
|
||||||
|
for (int i = 0; i < 5; i++) {
|
||||||
|
assertEquals("b", policy.select(ctx(candidates, noSessions())).profile(),
|
||||||
|
"a has weight 0, so every pick lands on b");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void roundRobinThrowsWhenAllWeightZero() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.roundRobin();
|
||||||
|
List<PlacementCandidate> candidates = List.of(
|
||||||
|
PlacementCandidate.profile("a", 0.0f, null),
|
||||||
|
PlacementCandidate.profile("b", 0.0f, null));
|
||||||
|
PlacementException e = assertThrows(PlacementException.class,
|
||||||
|
() -> policy.select(ctx(candidates, noSessions())));
|
||||||
|
assertTrue(e.getMessage().contains("weight 0"), e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void fixedSkipsWeightZeroDefault() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||||
|
PlacementContext ctx = new PlacementContext("b",
|
||||||
|
List.of(PlacementCandidate.profile("a", 1.0f, null),
|
||||||
|
PlacementCandidate.profile("b", 0.0f, null)),
|
||||||
|
noSessions(), Set.of(), Set.of());
|
||||||
|
assertEquals("a", policy.select(ctx).profile(),
|
||||||
|
"the default 'b' has weight 0, so fixed falls through to the first available candidate");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void fixedThrowsWhenDefaultAndEveryCandidateWeightZero() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||||
|
PlacementContext ctx = new PlacementContext("b",
|
||||||
|
List.of(PlacementCandidate.profile("a", 0.0f, null),
|
||||||
|
PlacementCandidate.profile("b", 0.0f, null)),
|
||||||
|
noSessions(), Set.of(), Set.of());
|
||||||
|
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||||
|
assertTrue(e.getMessage().contains("weight 0"), e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void unknownPolicyNameThrows() {
|
void unknownPolicyNameThrows() {
|
||||||
assertThrows(IllegalArgumentException.class, () -> PlacementPolicies.fromName("random"));
|
assertThrows(IllegalArgumentException.class, () -> PlacementPolicies.fromName("random"));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- CB-585: maxLoad 0 caps a profile at zero live members, even with nothing live yet -----
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void weightedSkipsMaxLoadZeroProfileEvenWithZeroLiveWorkers() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.weighted();
|
||||||
|
List<PlacementCandidate> candidates = List.of(
|
||||||
|
PlacementCandidate.profile("a", 1.0f, 0),
|
||||||
|
PlacementCandidate.profile("b", 1.0f, null));
|
||||||
|
Function<String, Integer> liveCount = _ -> 0;
|
||||||
|
for (int i = 0; i < 5; i++) {
|
||||||
|
assertEquals("b", policy.select(ctx(candidates, liveCount)).profile(),
|
||||||
|
"a has maxLoad 0, so it is already at its cap with nobody live — every pick lands on b");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void weightedThrowsWhenAllMaxLoadZero() {
|
||||||
|
PlacementPolicy policy = PlacementPolicies.weighted();
|
||||||
|
List<PlacementCandidate> candidates = List.of(
|
||||||
|
PlacementCandidate.profile("a", 1.0f, 0),
|
||||||
|
PlacementCandidate.profile("b", 1.0f, 0));
|
||||||
|
Function<String, Integer> liveCount = _ -> 0;
|
||||||
|
PlacementException e = assertThrows(PlacementException.class,
|
||||||
|
() -> policy.select(ctx(candidates, liveCount)));
|
||||||
|
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -18,6 +18,8 @@ import dev.ltms.bridged.session.GitWorktrees;
|
|||||||
import dev.ltms.bridged.session.SessionManager;
|
import dev.ltms.bridged.session.SessionManager;
|
||||||
import dev.ltms.bridged.session.Worktrees;
|
import dev.ltms.bridged.session.Worktrees;
|
||||||
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
||||||
|
import dev.ltms.bridged.member.CompositePeerLauncher;
|
||||||
|
import dev.ltms.bridged.placement.PlacementPolicies;
|
||||||
import io.javalin.Javalin;
|
import io.javalin.Javalin;
|
||||||
import org.junit.jupiter.api.AfterEach;
|
import org.junit.jupiter.api.AfterEach;
|
||||||
import org.junit.jupiter.api.Test;
|
import org.junit.jupiter.api.Test;
|
||||||
@@ -240,6 +242,43 @@ class BridgedAppTest {
|
|||||||
assertFalse(herdr.called("agent.start"), "an unknown profile must not spawn anything");
|
assertFalse(herdr.called("agent.start"), "an unknown profile must not spawn anything");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-599: a profile at its {@code maxLoad} cap must not surface as a bare 500 — the caller
|
||||||
|
* needs a structured, readable reason, distinct from "unknown_profile".
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void spawnAtMaxLoadIs503WithTheCapacityReasonNotABare500() throws Exception {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
BridgedConfig.Profile wcfg = new BridgedConfig.Profile(
|
||||||
|
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
|
||||||
|
"tab", "bridged-workers", "worker: {profile} #{n}", null,
|
||||||
|
null, null, null, null, null, null, null, 0, null, null, null);
|
||||||
|
Map<String, BridgedConfig.Profile> profiles = Map.of(wcfg.profile(), wcfg);
|
||||||
|
ClaudeCodeLauncher delegate = new ClaudeCodeLauncher(
|
||||||
|
new AgentControl(herdr), new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
|
||||||
|
profiles, wcfg.profile(), k -> "BRIDGED_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
|
||||||
|
CompositePeerLauncher workers = new CompositePeerLauncher(
|
||||||
|
List.of(delegate), wcfg.profile(), profiles, PlacementPolicies.fixed(), _ -> 0);
|
||||||
|
SessionManager sessions = new SessionManager(workers, new GitWorktrees());
|
||||||
|
this.presence = sessions.asPresence();
|
||||||
|
Injector injector = new Injector(new AgentControl(herdr));
|
||||||
|
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||||
|
sessions.onAcquire(inbox::own);
|
||||||
|
MessageService messages = new MessageService(new AgentControl(herdr), injector, new Rendezvous(), inbox);
|
||||||
|
app = new BridgedApp(herdr, workers, sessions, messages, this.presence, null).build().start("127.0.0.1", 0);
|
||||||
|
int port = app.port();
|
||||||
|
|
||||||
|
HttpResponse<String> res = req(port, "POST", "/members?profile=ltms-local");
|
||||||
|
|
||||||
|
assertEquals(503, res.statusCode(), res.body());
|
||||||
|
JsonNode body = mapper.readTree(res.body());
|
||||||
|
assertEquals("no_capacity", body.get("error").asText());
|
||||||
|
String detail = body.get("detail").asText();
|
||||||
|
assertTrue(detail.contains("ltms-local"), "detail names the profile: " + detail);
|
||||||
|
assertTrue(detail.contains("maxLoad"), "detail explains the refusal: " + detail);
|
||||||
|
assertFalse(herdr.called("agent.start"), "at cap, the spawn is refused before any herdr call");
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void spawnWorkerReusesExistingWorkerSpace() throws Exception {
|
void spawnWorkerReusesExistingWorkerSpace() throws Exception {
|
||||||
// A space labelled "bridged-workers" already exists → no second workspace.create.
|
// A space labelled "bridged-workers" already exists → no second workspace.create.
|
||||||
@@ -445,6 +484,67 @@ class BridgedAppTest {
|
|||||||
assertTrue(mapper.readTree(req(port, "GET", "/sessions/term_a/status").body()).get("ready").asBoolean());
|
assertTrue(mapper.readTree(req(port, "GET", "/sessions/term_a/status").body()).get("ready").asBoolean());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-582: a lead polling {@code GET /sessions/{id}/status} on its normal cadence — not the
|
||||||
|
* ticket-scoped {@code /tasks/{ticket}} — must also see a worker's open async {@code bridge_ask}
|
||||||
|
* question, since the reverse-rendezvous window it opened with is far shorter than that cadence.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void sessionStatusReportsAnOpenQuestionWhenTheWorkerIsMidAsk() throws Exception {
|
||||||
|
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // poller delivers the injection
|
||||||
|
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||||
|
|
||||||
|
HttpResponse<String> accepted = postMessage(port, "{\"content\":\"do it\",\"wait\":false}");
|
||||||
|
assertEquals(202, accepted.statusCode());
|
||||||
|
String ticket = mapper.readTree(accepted.body()).get("ticket").asText();
|
||||||
|
|
||||||
|
Thread.sleep(200); // let the background async send open its rendezvous waiter
|
||||||
|
|
||||||
|
var ask = java.util.concurrent.CompletableFuture.supplyAsync(() -> {
|
||||||
|
try {
|
||||||
|
return postJson(port, "/sessions/term_a/ask",
|
||||||
|
"{\"question\":\"which config file?\",\"timeoutMs\":5000}");
|
||||||
|
} catch (Exception e) {
|
||||||
|
throw new RuntimeException(e);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
JsonNode task;
|
||||||
|
long deadline = System.currentTimeMillis() + 3000;
|
||||||
|
do {
|
||||||
|
task = mapper.readTree(req(port, "GET", "/tasks/" + ticket).body());
|
||||||
|
if ("asking".equals(task.path("phase").asText())) break;
|
||||||
|
//noinspection BusyWait
|
||||||
|
Thread.sleep(10);
|
||||||
|
} while (System.currentTimeMillis() < deadline);
|
||||||
|
assertEquals("asking", task.get("phase").asText());
|
||||||
|
String turnId = task.get("turnId").asText();
|
||||||
|
|
||||||
|
HttpResponse<String> status = req(port, "GET", "/sessions/term_a/status");
|
||||||
|
assertEquals(200, status.statusCode());
|
||||||
|
JsonNode body = mapper.readTree(status.body());
|
||||||
|
assertEquals("idle", body.get("status").asText(), "the live status must still be reported");
|
||||||
|
assertEquals("which config file?", body.get("question").asText());
|
||||||
|
assertEquals(turnId, body.get("turnId").asText());
|
||||||
|
assertEquals(ticket, body.get("ticket").asText());
|
||||||
|
|
||||||
|
// Answer it — via the same /message route bridge_send uses, keyed by turnId — so the
|
||||||
|
// background ask thread does not linger past the test.
|
||||||
|
var answer = java.util.concurrent.CompletableFuture.supplyAsync(() -> {
|
||||||
|
try {
|
||||||
|
return postJson(port, "/sessions/term_a/message",
|
||||||
|
"{\"content\":\"config.yaml\",\"turnId\":\"" + turnId + "\",\"timeoutMs\":4000}");
|
||||||
|
} catch (Exception e) {
|
||||||
|
throw new RuntimeException(e);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
HttpResponse<String> askResponse = ask.get(6, java.util.concurrent.TimeUnit.SECONDS);
|
||||||
|
assertEquals(200, askResponse.statusCode());
|
||||||
|
assertEquals("config.yaml", mapper.readTree(askResponse.body()).get("answer").asText());
|
||||||
|
postJson(port, "/sessions/term_a/reply", "{\"content\":\"done\"}");
|
||||||
|
answer.get(6, java.util.concurrent.TimeUnit.SECONDS);
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void stopWorkerInPanePlacementClosesOnlyThePane() throws Exception {
|
void stopWorkerInPanePlacementClosesOnlyThePane() throws Exception {
|
||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
|||||||
@@ -2,9 +2,11 @@ package dev.ltms.bridged.session;
|
|||||||
|
|
||||||
import java.util.Collections;
|
import java.util.Collections;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
|
import java.util.Optional;
|
||||||
import java.util.Set;
|
import java.util.Set;
|
||||||
import java.util.concurrent.ConcurrentHashMap;
|
import java.util.concurrent.ConcurrentHashMap;
|
||||||
import java.util.concurrent.CopyOnWriteArrayList;
|
import java.util.concurrent.CopyOnWriteArrayList;
|
||||||
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
|
||||||
/** Recording fake {@link Worktrees} for CB-301-ext acceptance tests (no live git). */
|
/** Recording fake {@link Worktrees} for CB-301-ext acceptance tests (no live git). */
|
||||||
public final class FakeWorktrees implements Worktrees {
|
public final class FakeWorktrees implements Worktrees {
|
||||||
@@ -22,16 +24,28 @@ public final class FakeWorktrees implements Worktrees {
|
|||||||
public record RepoRootCall(String cwd) {
|
public record RepoRootCall(String cwd) {
|
||||||
}
|
}
|
||||||
|
|
||||||
|
public record SnapshotCall(String worktreePath, String branch, String message) {
|
||||||
|
}
|
||||||
|
|
||||||
|
public record PruneCall(String repoRoot, long minAgeMillis) {
|
||||||
|
}
|
||||||
|
|
||||||
private final List<AddCall> addCalls = new CopyOnWriteArrayList<>();
|
private final List<AddCall> addCalls = new CopyOnWriteArrayList<>();
|
||||||
private final List<RemoveCall> removeCalls = new CopyOnWriteArrayList<>();
|
private final List<RemoveCall> removeCalls = new CopyOnWriteArrayList<>();
|
||||||
private final List<OverlayCall> overlayCalls = new CopyOnWriteArrayList<>();
|
private final List<OverlayCall> overlayCalls = new CopyOnWriteArrayList<>();
|
||||||
private final List<RepoRootCall> repoRootCalls = new CopyOnWriteArrayList<>();
|
private final List<RepoRootCall> repoRootCalls = new CopyOnWriteArrayList<>();
|
||||||
|
private final List<SnapshotCall> snapshotCalls = new CopyOnWriteArrayList<>();
|
||||||
|
private final List<PruneCall> pruneCalls = new CopyOnWriteArrayList<>();
|
||||||
private final Set<String> existingPaths = ConcurrentHashMap.newKeySet();
|
private final Set<String> existingPaths = ConcurrentHashMap.newKeySet();
|
||||||
private final Set<String> trackedPaths = ConcurrentHashMap.newKeySet();
|
private final Set<String> trackedPaths = ConcurrentHashMap.newKeySet();
|
||||||
|
private final AtomicLong snapshotSeq = new AtomicLong();
|
||||||
private volatile RuntimeException addFailure;
|
private volatile RuntimeException addFailure;
|
||||||
|
private volatile RuntimeException snapshotFailure;
|
||||||
private volatile boolean dirty = false;
|
private volatile boolean dirty = false;
|
||||||
private volatile String repoRoot = "/repo";
|
private volatile String repoRoot = "/repo";
|
||||||
private volatile String prefix = "/worktrees";
|
private volatile String prefix = "/worktrees";
|
||||||
|
private volatile WipRefStats wipRefs = new WipRefStats(0, 0L);
|
||||||
|
private volatile int pruneResult = 0;
|
||||||
|
|
||||||
public FakeWorktrees withRepoRoot(String root) {
|
public FakeWorktrees withRepoRoot(String root) {
|
||||||
this.repoRoot = root;
|
this.repoRoot = root;
|
||||||
@@ -68,6 +82,24 @@ public final class FakeWorktrees implements Worktrees {
|
|||||||
return this;
|
return this;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** Make subsequent {@link #snapshot} calls throw (simulates a failing git snapshot). */
|
||||||
|
public FakeWorktrees failSnapshot(String message) {
|
||||||
|
this.snapshotFailure = new WorktreeException(message);
|
||||||
|
return this;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Configure the value returned by {@link #wipRefs}. */
|
||||||
|
public FakeWorktrees withWipRefs(WipRefStats stats) {
|
||||||
|
this.wipRefs = stats;
|
||||||
|
return this;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Configure the value returned by {@link #pruneWipRefs}. */
|
||||||
|
public FakeWorktrees withPruneResult(int deleted) {
|
||||||
|
this.pruneResult = deleted;
|
||||||
|
return this;
|
||||||
|
}
|
||||||
|
|
||||||
@Override
|
@Override
|
||||||
public String add(String repoRoot, String branch, String baseRef) {
|
public String add(String repoRoot, String branch, String baseRef) {
|
||||||
addCalls.add(new AddCall(repoRoot, branch, baseRef));
|
addCalls.add(new AddCall(repoRoot, branch, baseRef));
|
||||||
@@ -112,6 +144,26 @@ public final class FakeWorktrees implements Worktrees {
|
|||||||
return repoRoot;
|
return repoRoot;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Optional<String> snapshot(String worktreePath, String branch, String message) {
|
||||||
|
snapshotCalls.add(new SnapshotCall(worktreePath, branch, message));
|
||||||
|
if (snapshotFailure != null) {
|
||||||
|
throw snapshotFailure;
|
||||||
|
}
|
||||||
|
return Optional.of("wip" + snapshotSeq.incrementAndGet());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public WipRefStats wipRefs(String repoRoot) {
|
||||||
|
return wipRefs;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
|
||||||
|
pruneCalls.add(new PruneCall(repoRoot, minAgeMillis));
|
||||||
|
return pruneResult;
|
||||||
|
}
|
||||||
|
|
||||||
public List<AddCall> addCalls() {
|
public List<AddCall> addCalls() {
|
||||||
return List.copyOf(addCalls);
|
return List.copyOf(addCalls);
|
||||||
}
|
}
|
||||||
@@ -139,4 +191,16 @@ public final class FakeWorktrees implements Worktrees {
|
|||||||
public OverlayCall lastOverlay() {
|
public OverlayCall lastOverlay() {
|
||||||
return overlayCalls.isEmpty() ? null : overlayCalls.getLast();
|
return overlayCalls.isEmpty() ? null : overlayCalls.getLast();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
public List<SnapshotCall> snapshotCalls() {
|
||||||
|
return List.copyOf(snapshotCalls);
|
||||||
|
}
|
||||||
|
|
||||||
|
public List<PruneCall> pruneCalls() {
|
||||||
|
return List.copyOf(pruneCalls);
|
||||||
|
}
|
||||||
|
|
||||||
|
public SnapshotCall lastSnapshot() {
|
||||||
|
return snapshotCalls.isEmpty() ? null : snapshotCalls.getLast();
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -3,9 +3,13 @@ package dev.ltms.bridged.session;
|
|||||||
import org.junit.jupiter.api.Test;
|
import org.junit.jupiter.api.Test;
|
||||||
import org.junit.jupiter.api.io.TempDir;
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.charset.StandardCharsets;
|
||||||
import java.nio.file.Files;
|
import java.nio.file.Files;
|
||||||
import java.nio.file.Path;
|
import java.nio.file.Path;
|
||||||
|
import java.util.HashSet;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
|
import java.util.Optional;
|
||||||
|
import java.util.Set;
|
||||||
import java.util.concurrent.TimeUnit;
|
import java.util.concurrent.TimeUnit;
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.*;
|
import static org.junit.jupiter.api.Assertions.*;
|
||||||
@@ -68,6 +72,120 @@ class GitWorktreesTest {
|
|||||||
return out;
|
return out;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** Every pending change in {@code cwd} — the whole-tree porcelain status, unlike {@link #status}. */
|
||||||
|
private static String fullStatus(Path cwd) throws Exception {
|
||||||
|
Process p = new ProcessBuilder("git", "status", "--porcelain")
|
||||||
|
.directory(cwd.toFile()).redirectErrorStream(true).start();
|
||||||
|
String out = new String(p.getInputStream().readAllBytes());
|
||||||
|
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git status timed out");
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
private static String revParse(Path cwd, String ref) throws Exception {
|
||||||
|
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "rev-parse", ref)
|
||||||
|
.redirectErrorStream(true).start();
|
||||||
|
String out = new String(p.getInputStream().readAllBytes()).trim();
|
||||||
|
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git rev-parse timed out");
|
||||||
|
assertEquals(0, p.exitValue(), "git rev-parse " + ref + " failed:\n" + out);
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** The recursive file list of a commit's tree — used to check what a snapshot actually committed. */
|
||||||
|
private static String lsTree(Path cwd, String ref) throws Exception {
|
||||||
|
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "ls-tree", "-r", "--name-only", ref)
|
||||||
|
.redirectErrorStream(true).start();
|
||||||
|
String out = new String(p.getInputStream().readAllBytes());
|
||||||
|
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git ls-tree timed out");
|
||||||
|
assertEquals(0, p.exitValue(), "git ls-tree " + ref + " failed:\n" + out);
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** The set of paths in {@code git diff --name-only from..to} — used to check exactly what a
|
||||||
|
* snapshot's tree changed relative to its parent, the same shape {@code git status --porcelain}
|
||||||
|
* reports for the worktree it was taken from. */
|
||||||
|
private static Set<String> diffNameOnly(Path cwd, String from, String to) throws Exception {
|
||||||
|
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "diff", "--name-only", from, to)
|
||||||
|
.redirectErrorStream(true).start();
|
||||||
|
String out = new String(p.getInputStream().readAllBytes());
|
||||||
|
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git diff timed out");
|
||||||
|
assertEquals(0, p.exitValue(), "git diff " + from + ".." + to + " failed:\n" + out);
|
||||||
|
Set<String> paths = new HashSet<>();
|
||||||
|
for (String line : out.split("\\R")) {
|
||||||
|
if (!line.isBlank()) {
|
||||||
|
paths.add(line.trim());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return paths;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** The set of paths a {@code git status --porcelain} listing names, stripping the two-char status
|
||||||
|
* code prefix each line carries. */
|
||||||
|
private static Set<String> porcelainPaths(String porcelain) {
|
||||||
|
Set<String> paths = new HashSet<>();
|
||||||
|
for (String line : porcelain.split("\\R")) {
|
||||||
|
if (!line.isBlank()) {
|
||||||
|
paths.add(line.substring(3).trim());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return paths;
|
||||||
|
}
|
||||||
|
|
||||||
|
private static String forEachRef(Path cwd, String pattern) throws Exception {
|
||||||
|
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "for-each-ref", pattern)
|
||||||
|
.redirectErrorStream(true).start();
|
||||||
|
String out = new String(p.getInputStream().readAllBytes());
|
||||||
|
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git for-each-ref timed out");
|
||||||
|
assertEquals(0, p.exitValue(), "git for-each-ref " + pattern + " failed:\n" + out);
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Write {@code content} as a blob into the object database; returns its sha. */
|
||||||
|
private static String blobOf(Path cwd, String content) throws Exception {
|
||||||
|
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "hash-object", "-w", "--stdin")
|
||||||
|
.redirectErrorStream(true).start();
|
||||||
|
p.getOutputStream().write(content.getBytes(StandardCharsets.UTF_8));
|
||||||
|
p.getOutputStream().close();
|
||||||
|
String out = new String(p.getInputStream().readAllBytes()).trim();
|
||||||
|
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git hash-object timed out");
|
||||||
|
assertEquals(0, p.exitValue(), "git hash-object failed:\n" + out);
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Build a single-file tree object from {@code blob}; returns the tree's sha. */
|
||||||
|
private static String treeOf(Path cwd, String path, String blob) throws Exception {
|
||||||
|
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "mktree")
|
||||||
|
.redirectErrorStream(true).start();
|
||||||
|
p.getOutputStream().write(("100644 blob " + blob + "\t" + path + "\n").getBytes(StandardCharsets.UTF_8));
|
||||||
|
p.getOutputStream().close();
|
||||||
|
String out = new String(p.getInputStream().readAllBytes()).trim();
|
||||||
|
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git mktree timed out");
|
||||||
|
assertEquals(0, p.exitValue(), "git mktree failed:\n" + out);
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** {@code git commit-tree} rooted at {@code tree} with a chosen committer date; returns the sha. */
|
||||||
|
private static String commitTree(Path cwd, String tree, String parent, String committerDate,
|
||||||
|
String message) throws Exception {
|
||||||
|
ProcessBuilder pb = new ProcessBuilder("git", "-C", cwd.toString(), "commit-tree",
|
||||||
|
tree, "-p", parent, "-m", message);
|
||||||
|
pb.environment().put("GIT_COMMITTER_DATE", committerDate);
|
||||||
|
Process p = pb.redirectErrorStream(true).start();
|
||||||
|
String out = new String(p.getInputStream().readAllBytes()).trim();
|
||||||
|
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git commit-tree timed out");
|
||||||
|
assertEquals(0, p.exitValue(), "git commit-tree failed:\n" + out);
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** {@code git update-ref <ref> <sha>} — create the snapshot ref directly. */
|
||||||
|
private static void updateRef(Path cwd, String ref, String sha) throws Exception {
|
||||||
|
git(cwd, "update-ref", ref, sha);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** True when {@code ref} exists in the repo (for-each-ref on a missing ref is empty, not an error). */
|
||||||
|
private static boolean refExists(Path cwd, String ref) throws Exception {
|
||||||
|
return !forEachRef(cwd, ref).trim().isEmpty();
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* The heart of CB-525: a provisioned worktree must not inherit the primary's MCP servers. Without
|
* The heart of CB-525: a provisioned worktree must not inherit the primary's MCP servers. Without
|
||||||
* the isolation step the checked-out {@code .mcp.json} carries them in, and a worker navigating
|
* the isolation step the checked-out {@code .mcp.json} carries them in, and a worker navigating
|
||||||
@@ -245,4 +363,263 @@ class GitWorktreesTest {
|
|||||||
assertEquals("", status(Path.of(wt), "opencode.json"), "opencode.json still shows as modified");
|
assertEquals("", status(Path.of(wt), "opencode.json"), "opencode.json still shows as modified");
|
||||||
assertEquals("", status(Path.of(wt), ".autoenv"), ".autoenv still shows as modified");
|
assertEquals("", status(Path.of(wt), ".autoenv"), ".autoenv still shows as modified");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-578 stage C, acceptance criterion 1. A dirty worktree — a tracked edit plus a brand-new
|
||||||
|
* untracked file, exactly the shape lost in CB-576 — must land in {@code refs/wip/<branch>}'s
|
||||||
|
* tree, and that ref must live outside {@code refs/heads} so it never shows up in
|
||||||
|
* {@code git branch} or gets swept by a branch cleanup.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void dirtySnapshotCreatesARefWhoseTreeContainsUntrackedAndTrackedChanges(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = initRepo(tmp.resolve("repo"));
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
String branch = "cb-578-c-a";
|
||||||
|
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
|
||||||
|
Files.writeString(Path.of(wt).resolve("untracked.txt"), "draft that was never added\n");
|
||||||
|
Files.writeString(Path.of(wt).resolve("README.md"), "edited tracked file\n");
|
||||||
|
|
||||||
|
Optional<String> ref = gitWorktrees.snapshot(wt, branch, "test snapshot");
|
||||||
|
|
||||||
|
assertTrue(ref.isPresent(), "a dirty worktree snapshot returns a commit sha");
|
||||||
|
String tree = lsTree(repo, "refs/wip/" + branch);
|
||||||
|
assertTrue(tree.contains("untracked.txt"), "snapshot tree must include the untracked file:\n" + tree);
|
||||||
|
assertTrue(tree.contains("README.md"), "snapshot tree must include the tracked edit:\n" + tree);
|
||||||
|
String heads = forEachRef(repo, "refs/heads");
|
||||||
|
assertFalse(heads.contains("refs/wip/"), "the snapshot ref must not live under refs/heads:\n" + heads);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-578 stage C, acceptance criterion 2. Building the commit through a temporary
|
||||||
|
* {@code GIT_INDEX_FILE} must leave the worker's own index, working tree, and HEAD exactly as
|
||||||
|
* they were — the worker may still be mid-write, and staging into the real index would corrupt
|
||||||
|
* that.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void snapshotDoesNotTouchTheWorkersOwnIndexWorkingTreeOrHead(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = initRepo(tmp.resolve("repo"));
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
String branch = "cb-578-c-b";
|
||||||
|
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
|
||||||
|
Files.writeString(Path.of(wt).resolve("untracked.txt"), "draft that was never added\n");
|
||||||
|
Files.writeString(Path.of(wt).resolve("README.md"), "edited tracked file\n");
|
||||||
|
String headBefore = revParse(Path.of(wt), "HEAD");
|
||||||
|
String statusBefore = fullStatus(Path.of(wt));
|
||||||
|
|
||||||
|
gitWorktrees.snapshot(wt, branch, "test snapshot");
|
||||||
|
|
||||||
|
assertEquals(headBefore, revParse(Path.of(wt), "HEAD"), "snapshot must not move HEAD");
|
||||||
|
assertEquals(statusBefore, fullStatus(Path.of(wt)),
|
||||||
|
"snapshot must not change the worker's own index or working tree status");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-578 stage C, acceptance criterion 3. {@code add -A} (never {@code -f}) respects
|
||||||
|
* {@code .gitignore}, and that is the only thing keeping a gitignored file (secrets, local
|
||||||
|
* config) out of a snapshot commit — an untracked file that is NOT ignored must still be
|
||||||
|
* included, so this isn't just "untracked files are dropped".
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void gitignoredFileIsExcludedButOtherUntrackedFilesAreNot(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = tmp.resolve("repo");
|
||||||
|
Files.createDirectories(repo);
|
||||||
|
git(repo, "init", "-q", "-b", "main");
|
||||||
|
git(repo, "config", "user.email", "test@example.invalid");
|
||||||
|
git(repo, "config", "user.name", "Test");
|
||||||
|
Files.writeString(repo.resolve(".gitignore"), ".env\n");
|
||||||
|
Files.writeString(repo.resolve("README.md"), "seed\n");
|
||||||
|
git(repo, "add", ".gitignore", "README.md");
|
||||||
|
git(repo, "commit", "-q", "-m", "seed");
|
||||||
|
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
String branch = "cb-578-c-c";
|
||||||
|
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
|
||||||
|
Files.writeString(Path.of(wt).resolve(".env"), "SECRET=shh\n");
|
||||||
|
Files.writeString(Path.of(wt).resolve("untracked.txt"), "draft that was never added\n");
|
||||||
|
|
||||||
|
Optional<String> ref = gitWorktrees.snapshot(wt, branch, "test snapshot");
|
||||||
|
|
||||||
|
assertTrue(ref.isPresent());
|
||||||
|
String tree = lsTree(repo, "refs/wip/" + branch);
|
||||||
|
assertFalse(tree.contains(".env"), "a gitignored file must never enter the snapshot:\n" + tree);
|
||||||
|
assertTrue(tree.contains("untracked.txt"),
|
||||||
|
"a non-ignored untracked file must still be included:\n" + tree);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** CB-578 stage C. A clean worktree still produces a valid, if tree-identical, commit — the caller
|
||||||
|
* (SessionManager) is the one that decides not to call this on a clean worktree. */
|
||||||
|
@Test
|
||||||
|
void snapshotOfAMissingWorktreeReturnsEmptyWithoutThrowing(@TempDir Path tmp) {
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
String gone = tmp.resolve("wts").resolve("does-not-exist").toString();
|
||||||
|
|
||||||
|
assertTrue(gitWorktrees.snapshot(gone, "some-branch", "msg").isEmpty(),
|
||||||
|
"a missing worktree has nothing to snapshot, and must not throw");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-587, acceptance criteria 1 and 2. A file with a REAL {@code --skip-worktree} bit set in the
|
||||||
|
* worktree's own index must never enter the snapshot, even though its on-disk content has locally
|
||||||
|
* diverged from what is committed — that is exactly the divergence {@code --skip-worktree} exists
|
||||||
|
* to hide from {@code git status}, and a snapshot built from a fresh empty temp index (the bug)
|
||||||
|
* stages that local content anyway because the fresh index carries none of the real index's flags.
|
||||||
|
* The snapshot's diff against its parent must list exactly what {@code git status --porcelain}
|
||||||
|
* reports for the worktree — no more, no less.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void dirtySnapshotHonoursARealSkipWorktreeBit(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = tmp.resolve("repo");
|
||||||
|
Files.createDirectories(repo);
|
||||||
|
git(repo, "init", "-q", "-b", "main");
|
||||||
|
git(repo, "config", "user.email", "test@example.invalid");
|
||||||
|
git(repo, "config", "user.name", "Test");
|
||||||
|
Files.writeString(repo.resolve("protected.cfg"), "committed-value\n");
|
||||||
|
Files.writeString(repo.resolve("README.md"), "seed\n");
|
||||||
|
git(repo, "add", "protected.cfg", "README.md");
|
||||||
|
git(repo, "commit", "-q", "-m", "seed");
|
||||||
|
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
String branch = "cb-587-a";
|
||||||
|
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
|
||||||
|
|
||||||
|
git(Path.of(wt), "update-index", "--skip-worktree", "protected.cfg");
|
||||||
|
Files.writeString(Path.of(wt).resolve("protected.cfg"), "locally-diverged-never-commit\n");
|
||||||
|
Files.writeString(Path.of(wt).resolve("untracked.txt"), "draft that was never added\n");
|
||||||
|
Files.writeString(Path.of(wt).resolve("README.md"), "edited tracked file\n");
|
||||||
|
|
||||||
|
String porcelain = fullStatus(Path.of(wt));
|
||||||
|
assertFalse(porcelain.contains("protected.cfg"),
|
||||||
|
"test setup invalid — protected.cfg must not show in git status once skip-worktree is set:\n"
|
||||||
|
+ porcelain);
|
||||||
|
|
||||||
|
Optional<String> ref = gitWorktrees.snapshot(wt, branch, "test snapshot");
|
||||||
|
|
||||||
|
assertTrue(ref.isPresent(), "a dirty worktree snapshot returns a commit sha");
|
||||||
|
Set<String> diffPaths = diffNameOnly(repo, "HEAD", "refs/wip/" + branch);
|
||||||
|
assertFalse(diffPaths.contains("protected.cfg"),
|
||||||
|
"a --skip-worktree file's local drift leaked into the snapshot:\n" + diffPaths);
|
||||||
|
assertEquals(porcelainPaths(porcelain), diffPaths,
|
||||||
|
"snapshot diff must list exactly what git status --porcelain reports, no more, no less");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-587, acceptance criterion 6. If the worktree's real index cannot be resolved/read, snapshot
|
||||||
|
* must fail loudly (throw) rather than silently falling back to an empty temp index and producing
|
||||||
|
* a wrong snapshot. SessionManager's caller already catches and WARNs on any exception here — this
|
||||||
|
* only needs to confirm the failure is not swallowed inside snapshot() itself.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void snapshotThrowsWhenTheWorktreesRealIndexCannotBeResolved(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = initRepo(tmp.resolve("repo"));
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
String branch = "cb-587-b";
|
||||||
|
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
|
||||||
|
|
||||||
|
// Break git's ability to resolve the worktree's real index by removing the linked worktree's
|
||||||
|
// `.git` file (which normally points at the main repo's worktrees/<name>/ directory).
|
||||||
|
Files.delete(Path.of(wt).resolve(".git"));
|
||||||
|
|
||||||
|
assertThrows(WorktreeException.class, () -> gitWorktrees.snapshot(wt, branch, "test snapshot"),
|
||||||
|
"an unresolvable real index must fail loudly, not silently snapshot from an empty index");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586, criterion 2. A snapshot whose content is NOT reachable from {@code main} is the last
|
||||||
|
* copy of a worker's work, and must never be deleted automatically — even when it is old and
|
||||||
|
* even when the caller passes a zero age floor. Uses the real snapshot path on a dirty worktree,
|
||||||
|
* so the unreachable tree is exactly the shape CB-576/CB-578 stage C exist to protect.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void anUnreachableSnapshotIsNeverPruned(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = initRepo(tmp.resolve("repo"));
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
String branch = "cb-586-unreachable";
|
||||||
|
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
|
||||||
|
Files.writeString(Path.of(wt).resolve("worker-draft.txt"), "work that exists nowhere else\n");
|
||||||
|
|
||||||
|
Optional<String> ref = gitWorktrees.snapshot(wt, branch, "snapshot with unreachable content");
|
||||||
|
assertTrue(ref.isPresent());
|
||||||
|
|
||||||
|
// Age floor 0 makes age a non-issue: only reachability can save it — and it must.
|
||||||
|
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), 0),
|
||||||
|
"the unreachable snapshot is the last copy and must not be pruned");
|
||||||
|
assertTrue(refExists(repo, "refs/wip/" + branch),
|
||||||
|
"an unreachable snapshot must survive the sweep");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586, criterion 1 (the reachable half). A snapshot whose tree content IS already reachable
|
||||||
|
* from {@code main} and which is older than the age floor is pure duplication — the work is
|
||||||
|
* recovered — so it must be pruned.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void aReachableSnapshotOlderThanTheFloorIsPruned(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = initRepo(tmp.resolve("repo"));
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
|
||||||
|
// A snapshot whose tree is exactly main's current tree: fully reachable from main.
|
||||||
|
String mainTree = revParse(repo, "main^{tree}");
|
||||||
|
String old = commitTree(repo, mainTree, revParse(repo, "HEAD"), "2020-01-01T00:00:00", "snapshot");
|
||||||
|
updateRef(repo, "refs/wip/recovered", old);
|
||||||
|
|
||||||
|
assertEquals(1, gitWorktrees.pruneWipRefs(repo.toString(), TimeUnit.HOURS.toMillis(24)),
|
||||||
|
"an old, main-reachable snapshot must be pruned");
|
||||||
|
assertFalse(refExists(repo, "refs/wip/recovered"),
|
||||||
|
"the reachable snapshot's ref must be gone after the sweep");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586, the age floor. A snapshot whose content IS reachable from {@code main} but which is
|
||||||
|
* younger than the age floor must not be swept — a lead may still be looking at it.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void aReachableButRecentSnapshotIsNotPruned(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = initRepo(tmp.resolve("repo"));
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
|
||||||
|
// Reachable from main, but committed "now" — a fresh snapshot. The 24h floor must protect it.
|
||||||
|
String mainTree = revParse(repo, "main^{tree}");
|
||||||
|
String fresh = commitTree(repo, mainTree, revParse(repo, "HEAD"),
|
||||||
|
"2038-01-01T00:00:00", "snapshot just taken");
|
||||||
|
updateRef(repo, "refs/wip/fresh", fresh);
|
||||||
|
|
||||||
|
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), TimeUnit.HOURS.toMillis(24)),
|
||||||
|
"a recent snapshot must be kept even when reachable");
|
||||||
|
assertTrue(refExists(repo, "refs/wip/fresh"),
|
||||||
|
"the recent reachable snapshot must survive the sweep");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586, criterion 5. A fleet that has never snapshotted anything has no {@code refs/wip/*},
|
||||||
|
* so a sweep is a no-op and the census reports none — identical to before CB-586 existed.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void aFleetWithNoSnapshotsPrunesNothingAndReportsNothing(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = initRepo(tmp.resolve("repo"));
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
|
||||||
|
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), 0),
|
||||||
|
"no snapshot refs means nothing to prune");
|
||||||
|
Worktrees.WipRefStats stats = gitWorktrees.wipRefs(repo.toString());
|
||||||
|
assertEquals(0, stats.count(), "a never-snapshotted fleet has zero refs/wip refs");
|
||||||
|
assertEquals(0L, stats.costBytes(), "a never-snapshotted fleet costs zero bytes");
|
||||||
|
}
|
||||||
|
|
||||||
|
/** CB-586, criterion 4: the census reports how many refs exist and roughly what they cost. */
|
||||||
|
@Test
|
||||||
|
void wipRefsReportsCountAndCost(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = initRepo(tmp.resolve("repo"));
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
|
||||||
|
String blob = blobOf(repo, "a recoverable snapshot's worth of content");
|
||||||
|
String tree = treeOf(repo, "snapshot.txt", blob);
|
||||||
|
updateRef(repo, "refs/wip/one", commitTree(repo, tree, revParse(repo, "HEAD"),
|
||||||
|
"2020-01-01T00:00:00", "snapshot"));
|
||||||
|
updateRef(repo, "refs/wip/two", commitTree(repo, tree, revParse(repo, "HEAD"),
|
||||||
|
"2020-01-02T00:00:00", "snapshot"));
|
||||||
|
|
||||||
|
Worktrees.WipRefStats stats = gitWorktrees.wipRefs(repo.toString());
|
||||||
|
assertEquals(2, stats.count(), "two snapshot refs are reported");
|
||||||
|
assertTrue(stats.costBytes() > 0, "the cost of the snapshots is a positive byte count");
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -11,9 +11,13 @@ import dev.ltms.bridged.herdr.FakeHerdr;
|
|||||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||||
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
||||||
import dev.ltms.bridged.msg.TestTurnTokens;
|
import dev.ltms.bridged.msg.TestTurnTokens;
|
||||||
|
import dev.ltms.bridged.peer.Capability;
|
||||||
import dev.ltms.bridged.peer.CharterReceipt;
|
import dev.ltms.bridged.peer.CharterReceipt;
|
||||||
import dev.ltms.bridged.peer.MemberRole;
|
import dev.ltms.bridged.peer.MemberRole;
|
||||||
|
import dev.ltms.bridged.peer.PeerHandle;
|
||||||
|
import dev.ltms.bridged.peer.PeerLauncher;
|
||||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||||
|
import dev.ltms.bridged.peer.SpawnRequest;
|
||||||
import org.junit.jupiter.api.Test;
|
import org.junit.jupiter.api.Test;
|
||||||
import org.slf4j.LoggerFactory;
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
@@ -68,9 +72,12 @@ class SessionManagerTest {
|
|||||||
*/
|
*/
|
||||||
private static final class RecordingWorktrees implements Worktrees {
|
private static final class RecordingWorktrees implements Worktrees {
|
||||||
private final List<String> removeCalls = new java.util.ArrayList<>();
|
private final List<String> removeCalls = new java.util.ArrayList<>();
|
||||||
|
private final List<String> snapshotCalls = new java.util.ArrayList<>();
|
||||||
private final java.util.Set<String> failRemoveFor = new java.util.HashSet<>();
|
private final java.util.Set<String> failRemoveFor = new java.util.HashSet<>();
|
||||||
private volatile boolean dirty = false;
|
private volatile boolean dirty = false;
|
||||||
private volatile RuntimeException hasUncommittedFailure;
|
private volatile RuntimeException hasUncommittedFailure;
|
||||||
|
private volatile RuntimeException snapshotFailure;
|
||||||
|
private final java.util.concurrent.atomic.AtomicLong snapshotSeq = new java.util.concurrent.atomic.AtomicLong();
|
||||||
|
|
||||||
RecordingWorktrees dirty(boolean dirty) {
|
RecordingWorktrees dirty(boolean dirty) {
|
||||||
this.dirty = dirty;
|
this.dirty = dirty;
|
||||||
@@ -87,6 +94,11 @@ class SessionManagerTest {
|
|||||||
return this;
|
return this;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
RecordingWorktrees failSnapshotWith(RuntimeException e) {
|
||||||
|
this.snapshotFailure = e;
|
||||||
|
return this;
|
||||||
|
}
|
||||||
|
|
||||||
@Override
|
@Override
|
||||||
public String add(String repoRoot, String branch, String baseRef) {
|
public String add(String repoRoot, String branch, String baseRef) {
|
||||||
return "/wt/" + branch.replace('/', '_');
|
return "/wt/" + branch.replace('/', '_');
|
||||||
@@ -117,9 +129,32 @@ class SessionManagerTest {
|
|||||||
return "/repo";
|
return "/repo";
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public java.util.Optional<String> snapshot(String worktreePath, String branch, String message) {
|
||||||
|
snapshotCalls.add(worktreePath);
|
||||||
|
if (snapshotFailure != null) {
|
||||||
|
throw snapshotFailure;
|
||||||
|
}
|
||||||
|
return java.util.Optional.of("wip" + snapshotSeq.incrementAndGet());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public WipRefStats wipRefs(String repoRoot) {
|
||||||
|
return new WipRefStats(0, 0L);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
|
||||||
List<String> removeCalls() {
|
List<String> removeCalls() {
|
||||||
return List.copyOf(removeCalls);
|
return List.copyOf(removeCalls);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
List<String> snapshotCalls() {
|
||||||
|
return List.copyOf(snapshotCalls);
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock, int contextCap) {
|
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock, int contextCap) {
|
||||||
@@ -175,7 +210,7 @@ class SessionManagerTest {
|
|||||||
// see which charter a member got, without ever carrying the charter prose itself (CB-571).
|
// see which charter a member got, without ever carrying the charter prose itself (CB-571).
|
||||||
MemberSession s = new MemberSession("p1", "term1", "prof", MemberRole.DEV, "/cwd", null,
|
MemberSession s = new MemberSession("p1", "term1", "prof", MemberRole.DEV, "/cwd", null,
|
||||||
0, 0, 0, MemberSession.State.READY, null, null,
|
0, 0, 0, MemberSession.State.READY, null, null,
|
||||||
CharterReceipt.compose(MemberRole.DEV, "prof", "role charter", "role charter\n\nreply"));
|
CharterReceipt.compose(MemberRole.DEV, "prof", "role charter", "role charter\n\nreply"), null);
|
||||||
|
|
||||||
Map<String, Object> view = SessionManager.rosterView(s, null);
|
Map<String, Object> view = SessionManager.rosterView(s, null);
|
||||||
|
|
||||||
@@ -567,7 +602,7 @@ class SessionManagerTest {
|
|||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
SessionManager sessions = sessionManager(herdr);
|
SessionManager sessions = sessionManager(herdr);
|
||||||
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
|
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||||
sessions.onRelease(released::add);
|
sessions.onRelease(detail -> released.add(detail.terminalId()));
|
||||||
|
|
||||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller", null);
|
MemberSession s = sessions.acquire("ltms-local", null, "/caller", null);
|
||||||
sessions.release(s.paneId());
|
sessions.release(s.paneId());
|
||||||
@@ -581,7 +616,7 @@ class SessionManagerTest {
|
|||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
SessionManager sessions = sessionManager(herdr);
|
SessionManager sessions = sessionManager(herdr);
|
||||||
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
|
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||||
sessions.onRelease(released::add);
|
sessions.onRelease(detail -> released.add(detail.terminalId()));
|
||||||
|
|
||||||
sessions.release("w9:p404"); // idempotent teardown of something already gone
|
sessions.release("w9:p404"); // idempotent teardown of something already gone
|
||||||
|
|
||||||
@@ -661,7 +696,7 @@ class SessionManagerTest {
|
|||||||
RecordingWorktrees worktrees = new RecordingWorktrees();
|
RecordingWorktrees worktrees = new RecordingWorktrees();
|
||||||
SessionManager sessions = sessionManager(herdr, worktrees);
|
SessionManager sessions = sessionManager(herdr, worktrees);
|
||||||
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
|
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||||
sessions.onRelease(released::add);
|
sessions.onRelease(detail -> released.add(detail.terminalId()));
|
||||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||||
new WorktreeRequest("cb-581c", null));
|
new WorktreeRequest("cb-581c", null));
|
||||||
worktrees.failHasUncommittedWith(new WorktreeException("git status exited 128"));
|
worktrees.failHasUncommittedWith(new WorktreeException("git status exited 128"));
|
||||||
@@ -769,4 +804,127 @@ class SessionManagerTest {
|
|||||||
|
|
||||||
assertTrue(worktrees.removeCalls().isEmpty(), "SHUTDOWN drain still preserves the worktree");
|
assertTrue(worktrees.removeCalls().isEmpty(), "SHUTDOWN drain still preserves the worktree");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ── CB-584: agentSessionId recorded at acquire, and gated by Capability.SESSION_RESUME ─────
|
||||||
|
|
||||||
|
/**
|
||||||
|
* A minimal {@link PeerLauncher} whose capabilities exclude SESSION_RESUME — the "does not
|
||||||
|
* support it" case. {@link #requireResumeCapability} throws before ever reaching {@link #spawn},
|
||||||
|
* so every method beyond {@link #capabilitiesFor} is unreachable in these tests and left unimplemented.
|
||||||
|
*/
|
||||||
|
private static final class NoResumeLauncher implements PeerLauncher {
|
||||||
|
@Override
|
||||||
|
public Set<Capability> capabilities() {
|
||||||
|
return Set.of(Capability.WORKTREE); // deliberately no SESSION_RESUME
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Set<Capability> capabilitiesFor(String profileName) {
|
||||||
|
return capabilities();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public PeerHandle spawn(SpawnRequest req) {
|
||||||
|
throw new UnsupportedOperationException("not reachable — the capability check refuses first");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Set<String> profiles() {
|
||||||
|
return Set.of("stub-profile");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public String defaultProfile() {
|
||||||
|
return "stub-profile";
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public String effectiveCwd(SpawnRequest req) {
|
||||||
|
throw new UnsupportedOperationException("not reachable — the capability check refuses first");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public List<String> parityOverlay(String profileName) {
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public List<?> list() {
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public int reapOrphanWorkers() {
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void stop(String id) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public boolean clearContext(String id) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void acquireRecordsTheAgentSessionIdFromTheHandle() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
SessionManager sessions = sessionManager(herdr);
|
||||||
|
|
||||||
|
MemberSession s = sessions.acquire("ltms-local", MemberRole.DEV, null, null, null, null,
|
||||||
|
"my-session-name", null);
|
||||||
|
|
||||||
|
assertNotNull(s.agentSessionId(), "a spawn that asked for session identity gets one back");
|
||||||
|
Map<String, Object> view = SessionManager.rosterView(s, null);
|
||||||
|
assertEquals(s.agentSessionId(), view.get("agentSessionId"),
|
||||||
|
"the roster exposes the same id the session recorded");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void acquireWithNeitherSessionFieldLeavesAgentSessionIdNull() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
SessionManager sessions = sessionManager(herdr);
|
||||||
|
|
||||||
|
MemberSession s = sessions.acquire("ltms-local", null, null, null);
|
||||||
|
|
||||||
|
assertNull(s.agentSessionId(), "no identity requested — unchanged from before CB-584");
|
||||||
|
assertFalse(SessionManager.rosterView(s, null).containsKey("agentSessionId"),
|
||||||
|
"a null id is omitted from the roster, like charterSha256 for a receipt-less session");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void resumeSessionIdResumesTheSameConversationOnASupportingAdapter() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
SessionManager sessions = sessionManager(herdr);
|
||||||
|
|
||||||
|
MemberSession s = sessions.acquire("ltms-local", MemberRole.DEV, null, null, null, null,
|
||||||
|
null, "cb-resume-77");
|
||||||
|
|
||||||
|
assertEquals("cb-resume-77", s.agentSessionId(),
|
||||||
|
"a resume adopts the prior id as its own agentSessionId");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void resumeSessionIdWithoutAnExplicitProfileIsRefused() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
SessionManager sessions = sessionManager(herdr);
|
||||||
|
|
||||||
|
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, () ->
|
||||||
|
sessions.acquire(null, MemberRole.DEV, null, null, null, null, null, "cb-resume-1"));
|
||||||
|
|
||||||
|
assertTrue(e.getMessage().contains("explicit profile"), e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void resumeSessionIdOnAnAdapterWithoutTheCapabilityIsRefusedNamingIt() {
|
||||||
|
SessionManager sessions = new SessionManager(new NoResumeLauncher());
|
||||||
|
|
||||||
|
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, () ->
|
||||||
|
sessions.acquire("stub-profile", MemberRole.DEV, null, null, null, null, null, "cb-resume-1"));
|
||||||
|
|
||||||
|
assertTrue(e.getMessage().contains("SESSION_RESUME"), e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("stub-profile"), e.getMessage());
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -11,10 +11,12 @@ import org.junit.jupiter.api.Test;
|
|||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
import java.util.Set;
|
import java.util.Set;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
import java.util.concurrent.atomic.AtomicLong;
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
|
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
|
||||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -90,6 +92,59 @@ class SessionReaperTest {
|
|||||||
return ticks.get() >= n;
|
return ticks.get() >= n;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586: the retention sweep must actually reach the seam when the loop runs.
|
||||||
|
*
|
||||||
|
* <p>The sweep's own tests call {@code SessionManager.sweepWipRefs} directly, which walks
|
||||||
|
* around the reaper's interval gate — and the gate is where it broke. {@code lastWipSweepNanos}
|
||||||
|
* started at {@code Long.MIN_VALUE}, so {@code now - lastWipSweepNanos} overflowed to a large
|
||||||
|
* negative number, the gate read that as "swept moments ago", and it returned <em>before</em>
|
||||||
|
* the assignment that would have fixed the field. The sweep never ran once, for the life of
|
||||||
|
* the process, and every direct-call test still passed.
|
||||||
|
*
|
||||||
|
* <p>So this asserts through the loop: spawn a worktree session (which is what tells the
|
||||||
|
* manager which repo holds {@code refs/wip/*}), start the reaper, and require a real prune call
|
||||||
|
* to arrive at the fake.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void theLoopActuallyRunsTheWipRetentionSweep() throws InterruptedException {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||||
|
SessionManager sessions = new SessionManager(launcher(herdr), worktrees);
|
||||||
|
// Without a worktree session the repo is unknown and the sweep is a legitimate no-op, so
|
||||||
|
// this step is what makes the assertion below meaningful rather than vacuous.
|
||||||
|
sessions.acquire("ltms-local", null, "/caller/proj", null, new WorktreeRequest("cb-586", null));
|
||||||
|
assertTrue(worktrees.pruneCalls().isEmpty(), "nothing has swept before the reaper starts");
|
||||||
|
|
||||||
|
SessionReaper reaper = new SessionReaper(sessions, IDLE_TTL_SECONDS, SHORT_INTERVAL_MILLIS);
|
||||||
|
reaper.start();
|
||||||
|
try {
|
||||||
|
long deadline = System.currentTimeMillis() + 3000;
|
||||||
|
while (worktrees.pruneCalls().isEmpty() && System.currentTimeMillis() < deadline) {
|
||||||
|
Thread.sleep(10);
|
||||||
|
}
|
||||||
|
} finally {
|
||||||
|
reaper.stop();
|
||||||
|
}
|
||||||
|
|
||||||
|
assertFalse(worktrees.pruneCalls().isEmpty(),
|
||||||
|
"the reaper loop must run the refs/wip retention sweep; it never reached the seam");
|
||||||
|
assertEquals("/repo", worktrees.pruneCalls().getFirst().repoRoot(),
|
||||||
|
"the sweep must target the repo the fleet's worktrees came from");
|
||||||
|
assertEquals(TimeUnit.HOURS.toMillis(24), worktrees.pruneCalls().getFirst().minAgeMillis(),
|
||||||
|
"the 24h age floor is the safety rule and must reach the seam intact");
|
||||||
|
}
|
||||||
|
|
||||||
|
/** As {@link #launcher()}, on a caller-supplied herdr so the test can inspect it. */
|
||||||
|
private static ClaudeCodeLauncher launcher(FakeHerdr herdr) {
|
||||||
|
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||||
|
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||||
|
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||||
|
"worker: {profile} #{n}", null, null, null);
|
||||||
|
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||||
|
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void stopIsIdempotentAndSafeBeforeStart() {
|
void stopIsIdempotentAndSafeBeforeStart() {
|
||||||
SessionReaper reaper = reaper();
|
SessionReaper reaper = reaper();
|
||||||
|
|||||||
@@ -362,4 +362,141 @@ class WorktreeSessionManagerTest {
|
|||||||
"the profile's configured cwd must reach repoRoot, not be ignored");
|
"the profile's configured cwd must reach repoRoot, not be ignored");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- CB-578 stage C: snapshot a dirty worktree into git before the preserve-or-remove decision ---
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void releaseOfACleanWorktreeNeverSnapshots() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||||
|
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||||
|
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||||
|
new WorktreeRequest("cb-578c-a", null));
|
||||||
|
|
||||||
|
sessions.release(s.paneId());
|
||||||
|
|
||||||
|
assertTrue(worktrees.snapshotCalls().isEmpty(),
|
||||||
|
"acceptance criterion 4: a clean release must create no snapshot ref or commit");
|
||||||
|
assertEquals(1, worktrees.removeCalls().size(), "a clean worktree is still removed as before");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void releaseOfADirtyWorktreeSnapshotsBeforePreserving() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
|
||||||
|
.withDirty(true);
|
||||||
|
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||||
|
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||||
|
new WorktreeRequest("cb-578c-b", null));
|
||||||
|
|
||||||
|
sessions.release(s.paneId());
|
||||||
|
|
||||||
|
assertEquals(1, worktrees.snapshotCalls().size(),
|
||||||
|
"acceptance criterion 1: a dirty release snapshots the worktree exactly once");
|
||||||
|
FakeWorktrees.SnapshotCall call = worktrees.lastSnapshot();
|
||||||
|
assertEquals(s.worktree(), call.worktreePath());
|
||||||
|
assertEquals(s.branch(), call.branch());
|
||||||
|
assertTrue(worktrees.removeCalls().isEmpty(), "the dirty worktree is still preserved, not removed");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void releaseNotifiesTheListenerWithWorktreeBranchAndSnapshotRef() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
|
||||||
|
.withDirty(true);
|
||||||
|
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||||
|
java.util.List<SessionManager.ReleaseDetail> released = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||||
|
sessions.onRelease(released::add);
|
||||||
|
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||||
|
new WorktreeRequest("cb-578c-c", null));
|
||||||
|
|
||||||
|
sessions.release(s.paneId());
|
||||||
|
|
||||||
|
assertEquals(1, released.size());
|
||||||
|
SessionManager.ReleaseDetail detail = released.getFirst();
|
||||||
|
assertEquals(s.terminalId(), detail.terminalId());
|
||||||
|
assertEquals(s.worktree(), detail.worktreePath(),
|
||||||
|
"acceptance criterion 6: a failed ticket's detail must carry the worktree path");
|
||||||
|
assertEquals(s.branch(), detail.branch(),
|
||||||
|
"acceptance criterion 6: a failed ticket's detail must carry the branch");
|
||||||
|
assertEquals("wip1", detail.snapshotRef(),
|
||||||
|
"acceptance criterion 6: a failed ticket's detail must carry the snapshot ref");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void releaseNotifiesTheListenerWithTheAgentSessionId() {
|
||||||
|
// CB-584 (issue #65 criterion 5): a failed ticket's detail must also name the agent
|
||||||
|
// session, alongside worktree/branch/snapshot, so a lead can resume the conversation
|
||||||
|
// rather than only re-dispatch a fresh member onto the same files.
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
|
||||||
|
.withDirty(true);
|
||||||
|
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||||
|
java.util.List<SessionManager.ReleaseDetail> released = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||||
|
sessions.onRelease(released::add);
|
||||||
|
MemberSession s = sessions.acquire("ltms-local", MemberRole.DEV, null, "/caller/proj", null,
|
||||||
|
new WorktreeRequest("cb-584-e", null), "cb-584-session", null);
|
||||||
|
|
||||||
|
assertNotNull(s.agentSessionId(), "a named session must mint an agent session id to assert on");
|
||||||
|
sessions.release(s.paneId());
|
||||||
|
|
||||||
|
assertEquals(1, released.size());
|
||||||
|
SessionManager.ReleaseDetail detail = released.getFirst();
|
||||||
|
assertEquals(s.agentSessionId(), detail.agentSessionId());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aFailingSnapshotStillPreservesTheWorktreeStopsThePaneAndNotifies() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
|
||||||
|
.withDirty(true).failSnapshot("git commit-tree exited 128");
|
||||||
|
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||||
|
java.util.List<SessionManager.ReleaseDetail> released = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||||
|
sessions.onRelease(released::add);
|
||||||
|
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||||
|
new WorktreeRequest("cb-578c-d", null));
|
||||||
|
|
||||||
|
assertDoesNotThrow(() -> sessions.release(s.paneId()),
|
||||||
|
"acceptance criterion 5: a failing snapshot must not propagate out of release()");
|
||||||
|
|
||||||
|
assertTrue(herdr.called("pane.close"), "acceptance criterion 5: the pane must still stop");
|
||||||
|
assertTrue(worktrees.removeCalls().isEmpty(),
|
||||||
|
"acceptance criterion 5: the worktree must still be preserved when the snapshot fails");
|
||||||
|
assertEquals(1, released.size(), "acceptance criterion 5: the listener must still be notified");
|
||||||
|
assertNull(released.getFirst().snapshotRef(),
|
||||||
|
"a failed snapshot leaves no ref to report");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586, criteria 4 and 1. Once a worktree session is spawned the repo is known, so the
|
||||||
|
* operator census and the retention sweep delegate to that repo's {@code refs/wip/*}.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void wipRefsAndSweepDelegateToTheFleetRepoOnceKnown() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
|
||||||
|
.withWipRefs(new Worktrees.WipRefStats(3, 42L)).withPruneResult(2);
|
||||||
|
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||||
|
|
||||||
|
assertTrue(sessions.wipRefs().isEmpty(),
|
||||||
|
"no worktree spawned yet means no repo is known and nothing to report");
|
||||||
|
assertEquals(0, sessions.sweepWipRefs(TimeUnit.HOURS.toMillis(24)),
|
||||||
|
"no worktree spawned yet means the sweep is a no-op");
|
||||||
|
|
||||||
|
sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||||
|
new WorktreeRequest("cb-586", null));
|
||||||
|
|
||||||
|
Worktrees.WipRefStats stats = sessions.wipRefs().orElseThrow();
|
||||||
|
assertEquals(3, stats.count(), "the census comes from the fleet repo");
|
||||||
|
assertEquals(42L, stats.costBytes(), "the cost comes from the fleet repo");
|
||||||
|
assertEquals(2, sessions.sweepWipRefs(TimeUnit.HOURS.toMillis(24)),
|
||||||
|
"the sweep runs against the fleet repo");
|
||||||
|
|
||||||
|
List<FakeWorktrees.PruneCall> prunes = worktrees.pruneCalls();
|
||||||
|
assertEquals(1, prunes.size(), "the no-op short-circuits before reaching the seam, so only "
|
||||||
|
+ "the repo-known sweep issues a call");
|
||||||
|
assertEquals("/repo", prunes.getFirst().repoRoot(), "the sweep targets the fleet repo");
|
||||||
|
assertEquals(TimeUnit.HOURS.toMillis(24), prunes.getFirst().minAgeMillis(),
|
||||||
|
"the caller's age floor is passed through");
|
||||||
|
}
|
||||||
|
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
<?xml version="1.0" encoding="UTF-8"?>
|
<?xml version="1.0" encoding="UTF-8"?>
|
||||||
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
|
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
|
||||||
<!--
|
<!--
|
||||||
CB-504 — launchd agent for bridged (macOS).
|
CB-504 / CB-594 — launchd agent for bridged (macOS).
|
||||||
|
|
||||||
This is the real supervision target today: the dogfooded daemon runs on macOS, where there is
|
This is the real supervision target today: the dogfooded daemon runs on macOS, where there is
|
||||||
no systemd. A systemd unit ships alongside (deploy/bridged.service) for the Linux gateways
|
no systemd. A systemd unit ships alongside (deploy/bridged.service) for the Linux gateways
|
||||||
@@ -9,14 +9,33 @@
|
|||||||
|
|
||||||
Install:
|
Install:
|
||||||
cp deploy/dev.ltms.bridged.plist ~/Library/LaunchAgents/
|
cp deploy/dev.ltms.bridged.plist ~/Library/LaunchAgents/
|
||||||
# edit the paths + JAVA_HOME below to match this host, then:
|
|
||||||
launchctl load -w ~/Library/LaunchAgents/dev.ltms.bridged.plist
|
launchctl load -w ~/Library/LaunchAgents/dev.ltms.bridged.plist
|
||||||
launchctl list | grep bridged
|
launchctl list | grep bridged
|
||||||
|
|
||||||
|
The paths below are already filled in for this host (resolved 2026-08-16 from
|
||||||
|
`/usr/libexec/java_home`... except that reported the system Applet-plugin JVM, not the jenv-
|
||||||
|
managed JDK 25 actually used to build/run bridged, so JAVA_HOME here is the real one:
|
||||||
|
`JENV_VERSION=25.0.3 java -XshowSettings:properties -version 2>&1 | grep java.home`; `which mvn`;
|
||||||
|
`echo $HOME`). If this file is copied to a different host, re-resolve all three paths and check
|
||||||
|
no placeholder path is left behind; scripts/redeploy-bridged.sh's check mode does not (and
|
||||||
|
cannot) check this file for you.
|
||||||
|
|
||||||
|
CB-594 — launchd cannot run a login shell (see the PATH comment on EnvironmentVariables below,
|
||||||
|
and scripts/bridged-launchd-wrapper.sh for the fix): ProgramArguments below execs THAT wrapper,
|
||||||
|
not java directly, so WORKER_GITEA_TOKEN and AI_GATEWAY_TOKEN still get sourced from
|
||||||
|
${SHARED_ENV}/tools/secrets.sh even though launchd itself never sources anything.
|
||||||
|
|
||||||
Note on ordering: launchd has no "start after herdr" primitive for user agents, and neither
|
Note on ordering: launchd has no "start after herdr" primitive for user agents, and neither
|
||||||
does systemd in a way that survives a socket appearing late. bridged retries the herdr socket
|
does systemd in a way that survives a socket appearing late. bridged retries the herdr socket
|
||||||
on startup instead, so an agent that comes up before herdr converges rather than dying — that
|
on startup instead, so an agent that comes up before herdr converges rather than dying — that
|
||||||
retry is the actual fix; KeepAlive below is the backstop.
|
retry is the actual fix; KeepAlive below is the backstop.
|
||||||
|
|
||||||
|
CB-594 — KeepAlive vs. scripts/redeploy-bridged.sh: a bare SIGTERM makes this JVM exit 143 even
|
||||||
|
with its shutdown hook running to completion (measured, see the CB-594 report), which
|
||||||
|
SuccessfulExit:false below reads as a crash and races to restart the OLD jar. The redeploy
|
||||||
|
script now detects a loaded agent and uses `launchctl unload`/`load` instead of a raw kill, so
|
||||||
|
only one supervisor ever touches the process at a time — read that script's own output on a
|
||||||
|
redeploy for the confirmation.
|
||||||
-->
|
-->
|
||||||
<plist version="1.0">
|
<plist version="1.0">
|
||||||
<dict>
|
<dict>
|
||||||
@@ -25,22 +44,23 @@
|
|||||||
|
|
||||||
<key>ProgramArguments</key>
|
<key>ProgramArguments</key>
|
||||||
<array>
|
<array>
|
||||||
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home/bin/java</string>
|
<string>/Users/dai.ha/LTMS/claude-bridge/scripts/bridged-launchd-wrapper.sh</string>
|
||||||
|
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin/java</string>
|
||||||
<string>-jar</string>
|
<string>-jar</string>
|
||||||
<string>/Users/CHANGEME/src/claude-bridge/bridged/target/bridged.jar</string>
|
<string>/Users/dai.ha/LTMS/claude-bridge/bridged/target/bridged.jar</string>
|
||||||
<string>bridged.yaml</string>
|
<string>bridged.yaml</string>
|
||||||
</array>
|
</array>
|
||||||
|
|
||||||
<!-- Config path in ProgramArguments is relative, so the working directory must be the module. -->
|
<!-- Config path in ProgramArguments is relative, so the working directory must be the module. -->
|
||||||
<key>WorkingDirectory</key>
|
<key>WorkingDirectory</key>
|
||||||
<string>/Users/CHANGEME/src/claude-bridge/bridged</string>
|
<string>/Users/dai.ha/LTMS/claude-bridge/bridged</string>
|
||||||
|
|
||||||
<key>EnvironmentVariables</key>
|
<key>EnvironmentVariables</key>
|
||||||
<dict>
|
<dict>
|
||||||
<key>JAVA_HOME</key>
|
<key>JAVA_HOME</key>
|
||||||
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home</string>
|
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home</string>
|
||||||
<key>HERDR_SOCKET_PATH</key>
|
<key>HERDR_SOCKET_PATH</key>
|
||||||
<string>/Users/CHANGEME/.config/herdr/herdr.sock</string>
|
<string>/Users/dai.ha/.config/herdr/herdr.sock</string>
|
||||||
<!--
|
<!--
|
||||||
PATH matters more than it looks (CB-511): bridged propagates its own PATH to every worker
|
PATH matters more than it looks (CB-511): bridged propagates its own PATH to every worker
|
||||||
it spawns, so this line decides whether the fleet can run a build at all. launchd does NOT
|
it spawns, so this line decides whether the fleet can run a build at all. launchd does NOT
|
||||||
@@ -48,19 +68,39 @@
|
|||||||
bare /usr/bin:/bin and no JDK or Maven. Keep the toolchain entries first.
|
bare /usr/bin:/bin and no JDK or Maven. Keep the toolchain entries first.
|
||||||
-->
|
-->
|
||||||
<key>PATH</key>
|
<key>PATH</key>
|
||||||
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home/bin:/Users/CHANGEME/Tool/apache-maven-3.9.16/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin</string>
|
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin:/Users/dai.ha/Softwares/apache-maven/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin</string>
|
||||||
<!--
|
<!--
|
||||||
Worker/API tokens are NOT set here: this file is committed. Export them from a private
|
Worker/API tokens are NOT set here: this file is committed. CB-594 —
|
||||||
launchd override or a wrapper script. bridged reads the API token from the env var named
|
scripts/bridged-launchd-wrapper.sh (named in ProgramArguments above) is what supplies
|
||||||
by auth.tokenEnv (default BRIDGED_API_TOKEN) and only in auth.mode: token.
|
them, by execing a login shell that sources ${SHARED_ENV}/tools/secrets.sh before the
|
||||||
|
daemon itself starts. bridged also reads the API token from the env var named by
|
||||||
|
auth.tokenEnv (default BRIDGED_API_TOKEN) and only in auth.mode: token — the wrapper
|
||||||
|
covers that one too, since it is the same login shell.
|
||||||
-->
|
-->
|
||||||
</dict>
|
</dict>
|
||||||
|
|
||||||
<key>RunAtLoad</key>
|
<key>RunAtLoad</key>
|
||||||
<true/>
|
<true/>
|
||||||
|
|
||||||
<!-- Restart on crash, but not in a tight loop if the config is bad (bridged fails fast on a
|
<!--
|
||||||
non-loopback bind without token auth — that is a config error, not a transient one). -->
|
CB-600 — read this before assuming ThrottleInterval bounds anything. It paces restarts to at
|
||||||
|
most one per 10s; it does NOT cap how many times launchd retries. If bridged fails fast on
|
||||||
|
every start — a bad bridged.yaml, for example auth.mode: token with the token env var unset,
|
||||||
|
which throws in main() before the daemon ever binds a port — launchd restarts it forever,
|
||||||
|
once every 10s, until a human intervenes. LaunchAgents have no "give up after N attempts"
|
||||||
|
primitive, so this is not something a config change here can fix.
|
||||||
|
|
||||||
|
That loop stops only two ways: (1) `launchctl unload -w ~/Library/LaunchAgents/dev.ltms.bridged.plist`,
|
||||||
|
or (2) the underlying cause gets fixed, so the process starts successfully and stays up (no
|
||||||
|
more exits to restart). scripts/redeploy-bridged.sh does not add a third way — it does not
|
||||||
|
make bridged self-disable on a config error, on purpose: a fail-fast exit path that
|
||||||
|
sometimes decides "this is unrecoverable, stop trying" is one more thing that can misfire,
|
||||||
|
and a wrongly self-disabled daemon needs the exact same manual `launchctl load -w` recovery
|
||||||
|
this comment already names — so it buys nothing an operator watching for the crash loop
|
||||||
|
doesn't already have, at the cost of a new way to be silently down. Watch for it with
|
||||||
|
`launchctl list dev.ltms.bridged` (a high restart count) or by tailing bridged.out for the
|
||||||
|
same startup error repeating every ~10s.
|
||||||
|
-->
|
||||||
<key>KeepAlive</key>
|
<key>KeepAlive</key>
|
||||||
<dict>
|
<dict>
|
||||||
<key>SuccessfulExit</key>
|
<key>SuccessfulExit</key>
|
||||||
@@ -69,10 +109,17 @@
|
|||||||
<key>ThrottleInterval</key>
|
<key>ThrottleInterval</key>
|
||||||
<integer>10</integer>
|
<integer>10</integer>
|
||||||
|
|
||||||
|
<!--
|
||||||
|
CB-594 — same file scripts/redeploy-bridged.sh already tails ($BRIDGED/bridged.out), and both
|
||||||
|
streams point at it, not two separate log files: the script's fresh-line / ERROR-count checks
|
||||||
|
after a restart read this one path regardless of whether launchd or the script started the
|
||||||
|
process, and a stdout/stderr split would make half of what happens during a launchd-driven
|
||||||
|
restart invisible to it.
|
||||||
|
-->
|
||||||
<key>StandardOutPath</key>
|
<key>StandardOutPath</key>
|
||||||
<string>/Users/CHANGEME/src/claude-bridge/bridged/logs/bridged.out.log</string>
|
<string>/Users/dai.ha/LTMS/claude-bridge/bridged/bridged.out</string>
|
||||||
<key>StandardErrorPath</key>
|
<key>StandardErrorPath</key>
|
||||||
<string>/Users/CHANGEME/src/claude-bridge/bridged/logs/bridged.err.log</string>
|
<string>/Users/dai.ha/LTMS/claude-bridge/bridged/bridged.out</string>
|
||||||
|
|
||||||
<key>ProcessType</key>
|
<key>ProcessType</key>
|
||||||
<string>Background</string>
|
<string>Background</string>
|
||||||
|
|||||||
@@ -0,0 +1,467 @@
|
|||||||
|
# CB-591 — move the fleet onto the LLM and MCP gateway
|
||||||
|
|
||||||
|
**Status: DONE — the fleet is on the gateway as of 2026-08-15.** `local` runs on `/anthropic` and
|
||||||
|
`gx` on `/v1`, both at `weight: 100`; `local-direct` stays at `weight: 0` as the escape hatch. Getting
|
||||||
|
here took a revert and two upstream fixes — see §7.1, which is the useful part of this document. One
|
||||||
|
risk is **accepted rather than solved**: a stream cut by any mid-response timer arrives as HTTP 200
|
||||||
|
with no terminator, and our third-party members cannot detect it (§7.2).
|
||||||
|
· **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway)
|
||||||
|
· **Upstream issue:** [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31)
|
||||||
|
|
||||||
|
The gateway went live on 2026-08-15 and replaced Bifrost. This plan says what that means for a
|
||||||
|
**member definition** in `bridged.yaml`, because that is the part of this repo the change actually
|
||||||
|
touches.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. What changed upstream
|
||||||
|
|
||||||
|
One front door for every LLM and MCP client: `https://llm.ltms.dev`, one token per consumer.
|
||||||
|
|
||||||
|
| Surface | URL |
|
||||||
|
|---|---|
|
||||||
|
| OpenAI chat | `https://llm.ltms.dev/v1/chat/completions` |
|
||||||
|
| OpenAI models | `https://llm.ltms.dev/v1/models` |
|
||||||
|
| **Anthropic messages** | `https://llm.ltms.dev/anthropic/v1/messages` |
|
||||||
|
| MCP, all servers multiplexed | `https://llm.ltms.dev/mcp` |
|
||||||
|
|
||||||
|
Anything outside that list returns **404 before any token is checked**, on purpose — the gateway must
|
||||||
|
never become a blanket proxy.
|
||||||
|
|
||||||
|
The model backend is unchanged: GX10 vLLM at `10.10.10.26:8000` (`gx00.gw`), model name exactly
|
||||||
|
`deepseek-v4-flash`. The direct LAN path stays open on purpose as an escape hatch.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Where claude-bridge sits today
|
||||||
|
|
||||||
|
We do **not** use the gateway. The `local` profile talks straight to the vLLM:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
local:
|
||||||
|
kind: claude-code
|
||||||
|
baseUrl: http://gx00.gw:8000 # direct vLLM — no auth, LAN only
|
||||||
|
model: deepseek-v4-flash
|
||||||
|
configDir: /Users/dai.ha/.ccs/instances/gx10
|
||||||
|
```
|
||||||
|
|
||||||
|
Three facts about our side that decide the shape of this work:
|
||||||
|
|
||||||
|
1. **`baseUrl` becomes `ANTHROPIC_BASE_URL`** in the member's environment, and `tokenEnv` becomes
|
||||||
|
`ANTHROPIC_AUTH_TOKEN` (the value is read from a host env var and never stored in config).
|
||||||
|
`local` sets no `tokenEnv` today, because a direct vLLM needs no token.
|
||||||
|
2. **`SubscriptionGuard` refuses any host not on an allowlist**, and that allowlist is
|
||||||
|
`guard.offSubscriptionHosts: [gx00.gw]`. It is built once in `Bridged.java:93` and handed to the
|
||||||
|
launcher, so **it is a restart-required key**, not a hot one. Changing `baseUrl` without changing
|
||||||
|
this makes every `local` spawn throw.
|
||||||
|
3. **The wiki names us as a blocker.** Under *Not done yet*: retiring the shared `legacy` token is
|
||||||
|
blocked because "kb, brain, **claude-bridge** and the workstation still share it. Each needs its
|
||||||
|
own consumer first."
|
||||||
|
|
||||||
|
Context7 is mounted twice today, both times straight at `https://ct7.ltms.dev/mcp` — once in
|
||||||
|
`.mcp.json` (the primary) and once in `opencode.json` (the `sol` and `terra` members).
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
subgraph now["Today"]
|
||||||
|
M1["local member<br/>claude-code"] -->|"ANTHROPIC_BASE_URL"| V1["vLLM gx00.gw:8000<br/>no auth, LAN only"]
|
||||||
|
M2["sol / terra<br/>opencode"] --> CT1["ct7.ltms.dev/mcp"]
|
||||||
|
P1["primary"] --> CT1
|
||||||
|
end
|
||||||
|
subgraph after["Proposed"]
|
||||||
|
M3["local member"] -->|"ANTHROPIC_BASE_URL<br/>+ ANTHROPIC_AUTH_TOKEN"| G["llm.ltms.dev/anthropic<br/>consumer: claude-bridge"]
|
||||||
|
G --> V2["vLLM gx00.gw:8000"]
|
||||||
|
M4["local-direct<br/>weight 0, escape hatch"] --> V2
|
||||||
|
end
|
||||||
|
```
|
||||||
|
|
||||||
|
*The member definition is the only thing that moves. The model behind it does not.*
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. The member definition change
|
||||||
|
|
||||||
|
The gateway serves an Anthropic surface *and* an OpenAI surface, so **both member kinds can point at
|
||||||
|
it**. That is the main opportunity here, and it is bigger than the `local` profile alone.
|
||||||
|
|
||||||
|
### 3a. `local` — claude-code, on `/anthropic`
|
||||||
|
|
||||||
|
| Key | Today | After | Note |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `baseUrl` | `http://gx00.gw:8000` | `https://llm.ltms.dev/anthropic` | see the schema warning below |
|
||||||
|
| `tokenEnv` | *(unset)* | `AI_GATEWAY_TOKEN` | new consumer token, `llmk-claude-bridge-<32 hex>` |
|
||||||
|
| `model` | `deepseek-v4-flash` | unchanged | must stay **exact**; a regex match returns an empty `/v1/models` while completions keep working |
|
||||||
|
| `guard.offSubscriptionHosts` | `[gx00.gw]` | `[gx00.gw, llm.ltms.dev]` | **restart required** |
|
||||||
|
|
||||||
|
### 3b. A new opencode profile on `/v1` — no code needed
|
||||||
|
|
||||||
|
`OpenCodeLauncher` already supports a pinned OpenAI-compatible endpoint (CB-508). Given `baseUrl` it
|
||||||
|
writes a custom provider block into the worker's opencode config:
|
||||||
|
|
||||||
|
- `baseUrl` → `options.baseURL`. `openAiBaseUrl` appends `/v1` to a bare host, and takes a URL that
|
||||||
|
already has a path **as-is** — so `https://llm.ltms.dev/v1` works unchanged.
|
||||||
|
- `tokenEnv` → `options.apiKey` (falls back to a placeholder when unset, since a local vLLM ignores it).
|
||||||
|
- `model:` **must** be `<provider>/<model>` when `baseUrl` is set — a bare name is rejected loudly
|
||||||
|
rather than silently falling back to opencode's default gateway.
|
||||||
|
|
||||||
|
So the profile is pure config:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
gx:
|
||||||
|
kind: opencode
|
||||||
|
baseUrl: https://llm.ltms.dev/v1
|
||||||
|
tokenEnv: AI_GATEWAY_TOKEN
|
||||||
|
model: gx/deepseek-v4-flash # provider id is ours to choose; the half after / is the model
|
||||||
|
argv: ["opencode"]
|
||||||
|
mcpUrl: http://127.0.0.1:8765/mcp
|
||||||
|
gitTokenEnv: WORKER_GITEA_TOKEN
|
||||||
|
weight: 100 # same tier as `local` — free
|
||||||
|
maxLoad: 2
|
||||||
|
# deliberately NO credentialId — this is our own box, not the shared OpenAI account
|
||||||
|
```
|
||||||
|
|
||||||
|
**Why this matters more than it looks.** Today every opencode member is `sol` or `terra`, and those
|
||||||
|
are two models on **one** OpenAI account sharing `credentialId: openai-shared` — so an exhaustion on
|
||||||
|
either locks out both, and half the fleet's opencode capacity dies at once. A gateway-backed opencode
|
||||||
|
profile is free, is not on that credential, and therefore is not in that quarantine pair. It removes
|
||||||
|
a single point of failure rather than just adding capacity.
|
||||||
|
|
||||||
|
**Note the asymmetry, it is deliberate:** `SubscriptionGuard` does not apply to opencode at all — the
|
||||||
|
guard exists to stop a *Claude* worker borrowing the operator's subscription, and opencode reads its
|
||||||
|
own provider credentials. So 3b needs **no allowlist change**; only 3a does.
|
||||||
|
|
||||||
|
**Both still need a restart, for a different reason.** `tokenEnv` is resolved by
|
||||||
|
`HerdrPeerLauncher.resolveEnv` → `env.apply(name)`, which reads the **daemon's own process
|
||||||
|
environment**. The running `bridged` inherited its environment when it started, so a variable added to
|
||||||
|
`secrets.sh` afterwards is simply not there — the launcher would inject an empty token and the
|
||||||
|
gateway would answer 401. This is the same failure as trap 1 in `scripts/redeploy-bridged.sh`
|
||||||
|
(`WORKER_GITEA_TOKEN`), and it has the same fix: **restart from a login shell**, and use
|
||||||
|
`scripts/redeploy-bridged.sh --check` to confirm the name resolves before restarting anything.
|
||||||
|
|
||||||
|
### 3c. What this does to `ccs`
|
||||||
|
|
||||||
|
Once a profile carries `baseUrl`, `tokenEnv` and `model` itself, the ccs instance stops being what
|
||||||
|
routes a member. Be precise about what is left, though: `configDir` still supplies **folder trust**
|
||||||
|
and `settings.json`, and dropping it is what produced the trust dialog and the wrong-model error
|
||||||
|
recorded in `bridged.yaml`. So ccs goes from *deciding where the tokens go* to *holding client-side
|
||||||
|
state*. Less load-bearing, not removable.
|
||||||
|
|
||||||
|
### Why `/anthropic` and never `/v1/chat/completions`
|
||||||
|
|
||||||
|
The gateway declares its Anthropic backend as `schema.name: Anthropic`, which means **no
|
||||||
|
translation** — streaming, tool use and thinking blocks pass through exactly as they do against vLLM
|
||||||
|
directly.
|
||||||
|
|
||||||
|
Declared as `OpenAI`, Envoy's translator looks for a `thinking_blocks` field that our vLLM does not
|
||||||
|
send (it sends `reasoning_content`), and **every thinking delta disappears silently**. Claude Code
|
||||||
|
speaks the Anthropic protocol, so `/anthropic` is both correct and the only safe choice.
|
||||||
|
|
||||||
|
This is the exact failure shape this repo keeps hitting: it compiles, it answers, it looks healthy,
|
||||||
|
and a capability is quietly off. Treat it as a `silent-default` risk, not a config preference.
|
||||||
|
|
||||||
|
**Open question for 3b — ANSWERED, 2026-08-15.** The worry was that the OpenAI surface might drop
|
||||||
|
reasoning the way the wiki documents for a mis-declared Anthropic backend. It does not. Checked at
|
||||||
|
the API before any profile was switched:
|
||||||
|
|
||||||
|
| surface | request | result |
|
||||||
|
|---|---|---|
|
||||||
|
| `/anthropic/v1/messages` | `deepseek-v4-flash`, 64 tokens | 200, response carries a real `"type":"thinking"` block |
|
||||||
|
| `/v1/chat/completions` | same | 200, message carries a populated `reasoning_content` (and a `reasoning` field) |
|
||||||
|
| `/v1/models` | — | 200, exactly `["deepseek-v4-flash"]` — the exact-name trap is clear |
|
||||||
|
| `/v1/models`, **no token** | — | **401** — Caddy is gating, as designed |
|
||||||
|
|
||||||
|
So reasoning survives on **both** surfaces, and the `/anthropic` choice for `local` is about protocol
|
||||||
|
correctness rather than a repair for a known loss. The last row matters on its own: the wiki warns
|
||||||
|
the gateway's own `SecurityPolicy` fails open, so it is worth knowing the proxy in front really does
|
||||||
|
refuse an unauthenticated request here.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Decisions
|
||||||
|
|
||||||
|
### D1 — switch, but keep the direct path as an explicit profile · **recommended**
|
||||||
|
|
||||||
|
Switching buys four things we do not have:
|
||||||
|
|
||||||
|
- **Free opencode capacity, off the shared credential.** The largest single win. See §3b — it retires
|
||||||
|
a real single point of failure, not just a cost line.
|
||||||
|
- **Per-consumer usage figures.** The cockpit counts requests per consumer. That is the first real
|
||||||
|
measurement of what the fleet consumes, and it feeds [CB-589](https://git.ltms.dev/lms/claude-bridge/issues/74) Gap 2 directly.
|
||||||
|
- **Our own revocable token.** One consumer to revoke if a worker ever leaks it, instead of a shared
|
||||||
|
`legacy` token used by four systems.
|
||||||
|
- **It works off-LAN.** `gx00.gw` resolves on the LAN only.
|
||||||
|
|
||||||
|
The cost is honest and worth stating: we add a TLS edge, an auth proxy and a gateway to the path of
|
||||||
|
every member spawn. The wiki keeps the direct route open precisely because "if the gateway breaks,
|
||||||
|
nothing that matters is blocked."
|
||||||
|
|
||||||
|
So keep it. Add a second profile `local-direct` pointing at `http://gx00.gw:8000` with **`weight: 0`**
|
||||||
|
— never auto-selected, still spawnable with an explicit `bridge_spawn{profile: "local-direct"}`.
|
||||||
|
That is exactly what CB-554 made `weight: 0` mean, and it turns the escape hatch into something the
|
||||||
|
lead can actually reach during an incident.
|
||||||
|
|
||||||
|
### D2 — do members also mount the gateway's `/mcp`? · **OPEN, operator's call**
|
||||||
|
|
||||||
|
Not a detail. `CLAUDE.md` states in two places that a member mounts **only** the bridge MCP, and a
|
||||||
|
worker's honesty rule leans on it ("never claim the result of a check you had no way to run").
|
||||||
|
|
||||||
|
- **Keep bridge-only.** The invariant stays true and simple. Workers stay cheap and narrow.
|
||||||
|
- **Add the gateway MCP.** Implementers get context7 documentation lookups, which is genuinely useful
|
||||||
|
for library work. But `mcpUrl` in `BridgedConfig.Profile` is a **single `String`**, so a
|
||||||
|
claude-code member can mount exactly one MCP — this needs a code change, not a config edit.
|
||||||
|
|
||||||
|
Note the invariant is **already inaccurate**: `opencode.json` gives `sol` and `terra` both context7
|
||||||
|
and gitea. So the choice is really "make the rule true" or "make the rule match reality". Either is
|
||||||
|
defensible; picking one is not mine to do.
|
||||||
|
|
||||||
|
### D3 — token scope
|
||||||
|
|
||||||
|
One consumer, `claude-bridge`, its token in `${SHARED_ENV}/tools/secrets.sh` as `AI_GATEWAY_TOKEN`,
|
||||||
|
referenced by name only. Never the literal value in `bridged.yaml` — `tokenEnv` exists for this.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Units of work
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TB
|
||||||
|
U1["U1 · consumer token<br/>issue via cockpit, add to secrets.sh"]
|
||||||
|
U2["U2 · profile + guard<br/>bridged.yaml, restart"]
|
||||||
|
U3["U3 · verify live<br/>spawn, prove thinking survives"]
|
||||||
|
U4["U4 · context7 via gateway<br/>.mcp.json + opencode.json"]
|
||||||
|
U5["U5 · docs<br/>CLAUDE.md, wiki 11-Features"]
|
||||||
|
U1 --> U2 --> U3
|
||||||
|
U4 --> U5
|
||||||
|
U3 --> U5
|
||||||
|
```
|
||||||
|
|
||||||
|
| # | Scope | Who | Why |
|
||||||
|
|---|---|---|---|
|
||||||
|
| U1 | Issue the `claude-bridge` consumer at `auth.ltms.dev`; store as `AI_GATEWAY_TOKEN` | **operator** | touches secrets and a host we do not own |
|
||||||
|
| U2a | New `gx` opencode profile on `/v1` — pure config, no guard change | **lead** | `bridged.yaml` is gitignored, so a worker cannot see or edit it |
|
||||||
|
| U2b | `local` → `/anthropic`; add `local-direct` weight 0; add `llm.ltms.dev` to the guard allowlist | **lead** | same |
|
||||||
|
| U2c | One restart from a **login shell**, after U2a and U2b | **lead** | picks up `AI_GATEWAY_TOKEN` into the daemon env *and* the guard allowlist, in one stop |
|
||||||
|
| U3 | Live spawn on both new profiles; confirm reasoning survives on each surface | **lead** | needs real spawns and the running daemon |
|
||||||
|
| U4 | Point `.mcp.json` and `opencode.json` context7 at the gateway `/mcp`; rename pinned tools | delegatable | tracked files, self-contained |
|
||||||
|
| U5 | Fix the "members mount only the bridge" claim; add a `wiki/11-Features.md` entry | delegatable | writing, clear criteria |
|
||||||
|
|
||||||
|
U1 blocks U2a, U2b and U3. U4 and U5 do not depend on it.
|
||||||
|
|
||||||
|
**Write U2a and U2b, then restart once (U2c), then verify `gx` before `local`.** Since both profiles
|
||||||
|
need the same restart there is no reason to do two, but there is still a reason to *verify* in order:
|
||||||
|
`gx` exercises the token and the gateway with no guard involved, so if it fails the cause is upstream.
|
||||||
|
`local` adds the guard allowlist on top, so a failure there points at our config instead. Testing them
|
||||||
|
in that order separates the two causes instead of confusing them.
|
||||||
|
|
||||||
|
> **U1 status, 2026-08-15:** the operator issued the consumer and exported it as `AI_GATEWAY_TOKEN`
|
||||||
|
> (one key for every agent and MCP client behind `llm.ltms.dev`). Confirmed: it resolves in a login
|
||||||
|
> shell, is 48 characters and carries the documented `llmk-` prefix. The value was never printed.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Traps carried over from the wiki
|
||||||
|
|
||||||
|
Each of these cost someone real debugging time upstream. They apply to us.
|
||||||
|
|
||||||
|
1. **Rotating a token restarts the auth proxy, which drops in-flight streaming responses.** For us
|
||||||
|
that means rotating `AI_GATEWAY_TOKEN` kills every live member mid-turn, and an async ticket's
|
||||||
|
report goes with it. This is the same rule as a daemon redeploy: **drain the fleet first**
|
||||||
|
(`bridge_list` → `bridge_poll` anything wanted → `bridge_stop`), then rotate.
|
||||||
|
2. **The gateway's own `SecurityPolicy` fails open.** Standalone `aigw run` accepts it and silently
|
||||||
|
ignores it — an unauthenticated request returned **200**. Auth is the Caddy proxy in front, and
|
||||||
|
nothing else. Never reason as if the gateway authenticates.
|
||||||
|
3. **Exact model name.** A regex match routes fine but returns an **empty** `/v1/models` list while
|
||||||
|
completions keep working. A wrong name returns a bare 404 that reads exactly like a dead gateway.
|
||||||
|
4. **MCP tool names changed prefix separator.** Bifrost used one dash (`ct7-resolve-library-id`); the
|
||||||
|
gateway uses **two underscores** (`ct7__resolve-library-id`). Relevant only if U4 is done.
|
||||||
|
5. **`/v1/models` 404 vs empty list are different faults.** 404 means no route loaded at all; empty
|
||||||
|
means the model match is a regex. Do not conflate them when diagnosing.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Verification — what would prove this works
|
||||||
|
|
||||||
|
Merging config is not proving it. The checks, in order:
|
||||||
|
|
||||||
|
1. `bridge_spawn{profile: "gx"}` succeeds and the member completes a real turn ending in
|
||||||
|
`bridge_reply`. This is the first proof of the token, the URL and the model name, and it risks
|
||||||
|
nothing the fleet depends on.
|
||||||
|
2. `bridge_spawn{profile: "local"}` succeeds. If the guard allowlist was missed, this **throws** — a
|
||||||
|
loud, self-correcting failure, which is the good kind. If the restart was missed, it also throws,
|
||||||
|
for the same reason.
|
||||||
|
3. A `local` member completes a turn. That exercises streaming through two TLS edges, the auth proxy
|
||||||
|
and the gateway.
|
||||||
|
4. **Reasoning survives, checked separately on each surface.** For `local` on `/anthropic` this is
|
||||||
|
the check that catches the `/v1` versus `/anthropic` mistake, and it is the only one that does —
|
||||||
|
nothing else distinguishes a working passthrough from a translator quietly dropping thinking
|
||||||
|
deltas. For `gx` on `/v1`, this answers the open question in §3 rather than assuming it.
|
||||||
|
5. The cockpit at `auth.ltms.dev` shows requests counted against the `claude-bridge` consumer, not
|
||||||
|
`legacy`. That is the whole point of taking our own token.
|
||||||
|
6. `bridge_spawn{profile: "local-direct"}` still works, so the escape hatch is real rather than
|
||||||
|
theoretical.
|
||||||
|
7. `bridge_list` shows `gx` carrying no `credentialId`, so a `sol`/`terra` exhaustion cannot
|
||||||
|
quarantine it. This is the single-point-of-failure claim in §3b, checked rather than asserted.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7.1 What the live run actually found — 2026-08-15
|
||||||
|
|
||||||
|
U1–U2c were done, the daemon restarted onto them, and both new profiles were spawned for real. The
|
||||||
|
migration was then **reverted**. This section is the result, so none of it has to be re-derived.
|
||||||
|
|
||||||
|
### The blocker
|
||||||
|
|
||||||
|
`llm.ltms.dev` answers **HTTP 413 Request Entity Too Large** above **32 KiB (32768 bytes)**, on both
|
||||||
|
surfaces:
|
||||||
|
|
||||||
|
```
|
||||||
|
/v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200
|
||||||
|
/v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413
|
||||||
|
```
|
||||||
|
|
||||||
|
32 KiB is far below one real agent turn.
|
||||||
|
|
||||||
|
**Root cause — confirmed by the systems/vms side, 2026-08-15.** My guess that it was a Caddy
|
||||||
|
`request_body max_size` was **wrong**. It is Envoy, inside `aigw` on `llm.vm`. Envoy Gateway defaults
|
||||||
|
a listener's `per_connection_buffer_limit_bytes` to **32768**, and the AI Gateway buffers the *whole*
|
||||||
|
request body before it can route on the model name — so that default is not a network tuning knob
|
||||||
|
here, it is a hard ceiling on prompt size. Read out of the live Envoy `config_dump`:
|
||||||
|
|
||||||
|
```
|
||||||
|
listener default/llm/http per_connection_buffer_limit_bytes: 32768
|
||||||
|
```
|
||||||
|
|
||||||
|
Nobody chose 32 KiB; it was inherited from the default. Both TLS edges are innocent: the same
|
||||||
|
boundary reproduces on the LAN path and the internet path, and both 413s carry an `x-llm-consumer`
|
||||||
|
header their auth proxy sets only *after* authenticating — so the body cleared both edges and the
|
||||||
|
auth. Directly on `llm.vm`, `aigw` 413s at 39 KB while the vLLM backend accepts the same 39 KB and
|
||||||
|
answers 200.
|
||||||
|
|
||||||
|
**Do not plan around 32 KiB.** The intended ceiling is far higher. Their fix — a `ClientTrafficPolicy`
|
||||||
|
setting `bufferLimit: 8Mi` — is written but **not deployed** as of this note, pending their operator's
|
||||||
|
approval. I have not re-tested and will not until they confirm, so as not to measure a half-changed
|
||||||
|
system. Fixed in **systems/vms**, not here.
|
||||||
|
|
||||||
|
### The part worth remembering
|
||||||
|
|
||||||
|
Two members were spawned at the same moment with the same message:
|
||||||
|
|
||||||
|
| | `local` (claude-code, `/anthropic`) | `gx` (opencode, `/v1`) |
|
||||||
|
|---|---|---|
|
||||||
|
| READY → BUSY | 19:07:26 | 19:07:45 |
|
||||||
|
| BUSY → DONE | **19:08:51 (66s)** | **never — 10+ min, ticket FAILED** |
|
||||||
|
|
||||||
|
**`local` passed.** It passed only because the probe was three trivial questions in a fresh session,
|
||||||
|
so the request fit under 32 KiB. The profile looked healthy and was a landmine set to fire on the
|
||||||
|
first turn that reads a file.
|
||||||
|
|
||||||
|
So §7's checklist was not wrong, it was **too easy**. Any future run of it must use a task that reads
|
||||||
|
a real file. A liveness probe proves the token and the URL; it does not prove the path.
|
||||||
|
|
||||||
|
`gx` did not fail loudly either. Reproduced outside the bridge by running `opencode` by hand with the
|
||||||
|
launcher's own generated config:
|
||||||
|
|
||||||
|
```
|
||||||
|
Error: Request Entity Too Large
|
||||||
|
...compacts context, retries...
|
||||||
|
Error: Request Entity Too Large
|
||||||
|
```
|
||||||
|
|
||||||
|
opencode **catches the 413, compacts, and retries — indefinitely**. A member that fails loudly costs
|
||||||
|
one turn; this one costs the whole task and is indistinguishable from a slow worker.
|
||||||
|
|
||||||
|
> **Diagnosing a stuck opencode member.** Do not read its pane. The launcher writes its config to a
|
||||||
|
> temp dir and passes it as `OPENCODE_CONFIG` — find it with
|
||||||
|
> `ls -dt /var/folders/*/*/T/bridged-opencode-* | head -1`, check the provider block and the key's
|
||||||
|
> length and prefix (never its value), then reproduce with `opencode run --auto -m <provider>/<model>`
|
||||||
|
> using the same `OPENCODE_CONFIG`. That is what turned "it hangs" into a one-line error.
|
||||||
|
|
||||||
|
### What checked out, and needs no re-testing
|
||||||
|
|
||||||
|
- Token accepted on both surfaces. **Unauthenticated → 401**, so the Caddy proxy really does gate —
|
||||||
|
the wiki's "SecurityPolicy fails open" warning is about the gateway itself, not the edge.
|
||||||
|
- `/v1/models` returns exactly `["deepseek-v4-flash"]`, so trap 3 is clear.
|
||||||
|
- **Reasoning survives both surfaces** — see §3b above.
|
||||||
|
- The launcher's generated opencode provider block is correct, carrying a real 48-character `llmk-`
|
||||||
|
key rather than the `bridged-local-noauth` placeholder.
|
||||||
|
- `SubscriptionGuard` accepted `llm.ltms.dev` after the allowlist edit and the restart: `local`
|
||||||
|
spawned without throwing, which is the check that catches a missed restart.
|
||||||
|
|
||||||
|
### Resolution — both ceilings fixed, migration completed
|
||||||
|
|
||||||
|
systems/vms fixed both, and each was re-checked from this side rather than taken on trust:
|
||||||
|
|
||||||
|
| ceiling | was | now | our own check |
|
||||||
|
|---|---|---|---|
|
||||||
|
| listener buffer | 32 KiB | 32 Mi | 1.2 MB body → **200** (was 413) |
|
||||||
|
| LLM route timeout | 60s | 86400s | the request that truncated: **101s, `message_stop` present, 4000/4000** |
|
||||||
|
|
||||||
|
The timeout moved in two steps on 2026-08-15: 60s → 1800s, then 1800s → **86400s (24 hours)** after
|
||||||
|
the truncation risk below was discussed. They tried `request: 0s` first, which removes the
|
||||||
|
total-duration timer completely. It works, but on an `AIGatewayRoute` the **idle timeout is derived
|
||||||
|
from the request timeout**, so `0s` also removed any bound on a stalled connection. 86400s keeps a
|
||||||
|
reaper for dead connections while putting the truncation timer out of practical reach.
|
||||||
|
|
||||||
|
Neither was deliberate. The 32 KiB was Envoy Gateway's default `per_connection_buffer_limit_bytes`;
|
||||||
|
the 60s was Envoy AI Gateway's own documented default. The 60s bounded **generation** as well as
|
||||||
|
prompt size — a tiny prompt with a long answer returned 504 at 60.05s.
|
||||||
|
|
||||||
|
Two configuration facts worth keeping, from their bisection:
|
||||||
|
|
||||||
|
- **`ClientTrafficPolicy` is honoured in standalone `aigw run`; `BackendTrafficPolicy` is NOT.** A
|
||||||
|
`BackendTrafficPolicy` setting `requestTimeout` is accepted, logs nothing, and leaves the routes
|
||||||
|
unchanged (upstream `envoyproxy/gateway#9513`). What works is `timeouts: {request: …}` on each
|
||||||
|
`AIGatewayRoute` rule. Nothing from the outside distinguishes the two — the same silent-default
|
||||||
|
shape as their `SecurityPolicy` caveat.
|
||||||
|
- In that stack, "the config was accepted" proves nothing. Read the live `config_dump`.
|
||||||
|
|
||||||
|
## 7.2 The risk we accepted, and why we could not remove it
|
||||||
|
|
||||||
|
Raising the timeout made the failure **rare, not impossible**, and the residual failure is silent.
|
||||||
|
|
||||||
|
On a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked encoding *cleanly* instead of
|
||||||
|
resetting the connection, so the client receives what looks like a complete transfer
|
||||||
|
(`envoyproxy/envoy#17186` — acknowledged as a bug in 2021, closed by a stale bot, never fixed). The
|
||||||
|
December 2025 fix `envoyproxy/envoy#42269` changes locally-originated resets from `NO_ERROR` to
|
||||||
|
`INTERNAL_ERROR`, but it is **HTTP/2 only** and SSE clients here speak HTTP/1.1.
|
||||||
|
|
||||||
|
Measured on our side while the timeout was still 60s:
|
||||||
|
|
||||||
|
```
|
||||||
|
HTTP 200 61.07s 141992 bytes
|
||||||
|
message_stop 0 message_delta 0 error events 0
|
||||||
|
emitted 2473 of 4000, ending on a WELL-FORMED SSE frame
|
||||||
|
```
|
||||||
|
|
||||||
|
A syntactically valid stream that simply stops. Any timer firing mid-stream — route timeout, idle
|
||||||
|
timeout, `max_stream_duration` — fails this same way.
|
||||||
|
|
||||||
|
**The recommended defence does not transfer to us.** The right fix is to treat a stream with no
|
||||||
|
`message_stop` / `[DONE]` / `finish_reason` as failed. We cannot: our members are Claude Code and
|
||||||
|
opencode, third-party clients whose SSE parsing we do not own, and there is no seam to insert the
|
||||||
|
check. Whether either detects a missing terminator is unverified — and opencode's handling of the 413
|
||||||
|
(swallow, compact, retry forever, never surface an error) does not suggest it is strict.
|
||||||
|
|
||||||
|
So the honest statement of our position:
|
||||||
|
|
||||||
|
> Gateway traffic is acceptable at 86400s because a single request would have to run for 24 hours to
|
||||||
|
> trip the bug — **not** because we could detect it if it did.
|
||||||
|
|
||||||
|
At 86400s our **own** limit binds first, which is the ordering we want. `MessageService.ASYNC_TIMEOUT_MS`
|
||||||
|
caps a turn at 30 minutes, so a runaway request ends as a clean `FAILED` ticket that we raised, rather
|
||||||
|
than as a silently truncated `200` that we cannot see. While the gateway sat at 1800s the two numbers
|
||||||
|
were equal and did not nest, so a gateway-side stall could have been misread as a bug in our own ticket
|
||||||
|
handling. That ambiguity is now gone.
|
||||||
|
|
||||||
|
**If a member ever returns a confident but truncated answer, suspect this before anything in our own
|
||||||
|
code.** That is the whole reason this section exists.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Related
|
||||||
|
|
||||||
|
- [CB-589 / #74](https://git.ltms.dev/lms/claude-bridge/issues/74) — cost-first placement and a
|
||||||
|
gateway that reports live capacity. The per-consumer figures this migration unlocks are the first
|
||||||
|
input that ticket actually needs.
|
||||||
|
- `docs/CB-500-Multi-Tier-Coordination.md` §11 — the distributed-sandbox topology this gateway is
|
||||||
|
part of.
|
||||||
+36
-1
@@ -1,7 +1,9 @@
|
|||||||
# M4 - Fleet health, recovery, routing, and capacity
|
# M4 - Fleet health, recovery, routing, and capacity
|
||||||
|
|
||||||
**Status:** Design accepted on 2026-08-15. CB-573 part 1 has shipped the classification model and
|
**Status:** Design accepted on 2026-08-15. CB-573 part 1 has shipped the classification model and
|
||||||
the `bridge_list` capacity view; the remaining M4 units are not yet shipped.
|
the `bridge_list` capacity view; the remaining M4 units are not yet shipped. See
|
||||||
|
[Unit 2 — what has landed so far](#unit-2---what-has-landed-so-far) before planning Unit 2 work:
|
||||||
|
some of its criteria were met by separate CB tickets, and one of them contradicts the unit text.
|
||||||
**Scope:** Fleet evidence, safe mechanical repair, lead routing, capacity reporting, and optional
|
**Scope:** Fleet evidence, safe mechanical repair, lead routing, capacity reporting, and optional
|
||||||
human notification.
|
human notification.
|
||||||
**Grounded in:** `health/FleetHealth`, `health/PaneBudget`, `inject/StatusPoller`,
|
**Grounded in:** `health/FleetHealth`, `health/PaneBudget`, `inject/StatusPoller`,
|
||||||
@@ -799,6 +801,39 @@ Acceptance criteria:
|
|||||||
21. Tests cover both real traces, all repair refusals, clipping, explicit-reply and next-turn races,
|
21. Tests cover both real traces, all repair refusals, clipping, explicit-reply and next-turn races,
|
||||||
restart without capture, concurrent send and release, and preserved discovery after restart.
|
restart without capture, concurrent send and release, and preserved discovery after restart.
|
||||||
|
|
||||||
|
#### Unit 2 - what has landed so far
|
||||||
|
|
||||||
|
Checked against `main` at `e09cac6` on 2026-08-15. Unit 2 was written as one block, but parts of it
|
||||||
|
have since been built by separate CB tickets. Read this before planning the rest, or that work gets
|
||||||
|
done twice.
|
||||||
|
|
||||||
|
The check was a symbol survey of `bridged/src/main/java` plus the merge history. It tells you whether
|
||||||
|
the machinery exists at all. It is **not** a line-by-line audit of whether each criterion is fully
|
||||||
|
met, and I did not run one.
|
||||||
|
|
||||||
|
| Criterion | Marker searched for | Found in main source | Reading |
|
||||||
|
|---|---|---|---|
|
||||||
|
| 1 | `TurnToken` | 8 files | **Done** — unit 2a, merged as `fec284e`. Criterion 1 was corrected first; see the note under it. |
|
||||||
|
| 2-5, 9 | `REPAIRED` | 0 files | Not started. The whole guarded-repair path is absent. |
|
||||||
|
| 6, 7, 10 | `reconcileLostBoundary` | 0 files | Not started. |
|
||||||
|
| 12 | CB-568 failure operation | via CB-580 | **Partial.** CB-580 (`0af902e`) routes `GONE` and `NEVER_READY` into the one idempotent target-wide failure. I did not check that release and abnormal stop go through the same call. |
|
||||||
|
| 14 | `DELEGATION_ORPHANED` | 3 files | **Partial.** The health state exists. The teardown-invariant check that creates it, and the retry rule, do not. |
|
||||||
|
| 15 | `SPAWN_ROLLBACK` | 0 files | **Contradicted — see below.** |
|
||||||
|
| 16 | — | — | Partial at best. CB-576 made release preserve a dirty worktree; whether explicit stop is state-aware is not checked. |
|
||||||
|
| 17, 18 | `preservedWorktrees` | 0 files | Not started. No manifest, and no lead-only `bridge_list` field. |
|
||||||
|
| 19 | `WORK_PRODUCT_AT_RISK` | 0 files | Not started. |
|
||||||
|
|
||||||
|
**Criterion 15 no longer matches the code, and the code is right.** It says "normal `COMPLETED`
|
||||||
|
remove worktrees". Since CB-576 (`500bfa2`) that is false on purpose: a `COMPLETED` release now
|
||||||
|
preserves the worktree when it still holds uncommitted work, because deleting it destroys work
|
||||||
|
nobody can get back. CB-576 was filed after exactly that loss. CB-581 goes further — if the
|
||||||
|
dirty-check itself fails, the worktree is preserved rather than removed, since "we could not tell"
|
||||||
|
must not be treated as "it is clean".
|
||||||
|
|
||||||
|
So criterion 15 should be rewritten as: `SPAWN_ROLLBACK` and a `COMPLETED` release with a **clean**
|
||||||
|
worktree remove it; abnormal causes, shutdown, a dirty worktree, and a failed dirty-check all
|
||||||
|
preserve it. `SPAWN_ROLLBACK` itself does not exist yet.
|
||||||
|
|
||||||
### Unit 3 - Typed inbox and member routing
|
### Unit 3 - Typed inbox and member routing
|
||||||
|
|
||||||
Scope: semantic record, AMQP migration, both adapters, member routing, polling, and member health in
|
Scope: semantic record, AMQP migration, both adapters, member routing, polling, and member health in
|
||||||
|
|||||||
+1
-1
@@ -26,7 +26,7 @@
|
|||||||
],
|
],
|
||||||
"enabled": true,
|
"enabled": true,
|
||||||
"environment": {
|
"environment": {
|
||||||
"GITEA_ACCESS_TOKEN": "{env:GITEA_ACCESS_TOKEN}",
|
"GITEA_ACCESS_TOKEN": "{env:WORKER_GITEA_TOKEN}",
|
||||||
"GITEA_HOST": "{env:GITEA_HOST}"
|
"GITEA_HOST": "{env:GITEA_HOST}"
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
Executable
+32
@@ -0,0 +1,32 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
#
|
||||||
|
# CB-594 — the only reason this file exists: launchd does not run a login shell.
|
||||||
|
#
|
||||||
|
# WORKER_GITEA_TOKEN and AI_GATEWAY_TOKEN live in ${SHARED_ENV}/tools/secrets.sh, sourced only by a
|
||||||
|
# LOGIN shell (.zprofile/.zshrc etc). launchd execs a job's ProgramArguments directly — no shell, no
|
||||||
|
# profile, nothing sourced (the plist's own PATH comment documents the same gap one variable over).
|
||||||
|
# A daemon started that way boots fine and looks healthy; the failure is invisible until a worker
|
||||||
|
# tries to open a PR (WORKER_GITEA_TOKEN empty) or a gateway profile gets a 401 (AI_GATEWAY_TOKEN
|
||||||
|
# empty) — hours later, with nothing tying the two together (CB-591, CLAUDE.md "Redeploying the
|
||||||
|
# daemon"). Bridged now also logs which required secret names resolved at startup (see
|
||||||
|
# Bridged.reportRequiredSecrets), but that log line can only tell the truth if the tokens had a
|
||||||
|
# chance to be sourced in the first place — which is this script's entire job.
|
||||||
|
#
|
||||||
|
# So: launchd execs THIS script instead of java directly. This script execs a login shell
|
||||||
|
# ('zsh -l'), which sources secrets.sh, and that shell execs the real command in its place — one
|
||||||
|
# process throughout (exec, not a subshell fork), so launchd's PID tracking, KeepAlive, and
|
||||||
|
# StandardOut/ErrorPath all still see the one process they expect.
|
||||||
|
#
|
||||||
|
# The plist passes the full command as THIS script's own arguments, e.g.:
|
||||||
|
# ProgramArguments = [ .../bridged-launchd-wrapper.sh, /path/to/java, -jar, /path/to/bridged.jar,
|
||||||
|
# bridged.yaml ]
|
||||||
|
# so the wrapper stays generic and the actual command lives in exactly one place (the plist), not
|
||||||
|
# duplicated here.
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
if [ "$#" -eq 0 ]; then
|
||||||
|
echo "bridged-launchd-wrapper.sh: no command given — check the plist's ProgramArguments" >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
exec /bin/zsh -lc 'exec "$@"' -- "$@"
|
||||||
Executable
+135
@@ -0,0 +1,135 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
#
|
||||||
|
# CB-596 step 1: measure which credentials a live member actually holds.
|
||||||
|
#
|
||||||
|
# The claim under test is that a member's herdr pane starts a LOGIN shell, that shell sources
|
||||||
|
# ${SHARED_ENV}/tools/secrets.sh, and so a member inherits every name that file exports — while
|
||||||
|
# CB-592 blocks exactly one of them (GITEA_ACCESS_TOKEN). That is an inference from the code, not a
|
||||||
|
# measurement, and issue #82 says plainly: do not build a fix on the inference. This is the
|
||||||
|
# measurement.
|
||||||
|
#
|
||||||
|
# WHY THIS IS A SCRIPT AND NOT A COMMAND SOMEONE TYPES
|
||||||
|
#
|
||||||
|
# Enumerating credential names inside a member is exactly the action that should need the operator's
|
||||||
|
# explicit approval, and the command classifier refuses it. That refusal is correct. This script is
|
||||||
|
# the seam: it is one auditable file the operator can read once, top to bottom, and then run — rather
|
||||||
|
# than approving an ad-hoc shell pipeline whose behaviour they have to take on trust.
|
||||||
|
#
|
||||||
|
# WHAT IT WILL NOT DO
|
||||||
|
#
|
||||||
|
# * It never prints a credential value, and never any prefix or suffix of one. Not one character.
|
||||||
|
# Issue #82's criterion 1 asked for a 6-character prefix; this prints a truncated SHA-256 instead.
|
||||||
|
# A prefix of a short secret is most of the secret, and it would end up pasted into a ticket. The
|
||||||
|
# hash answers every question the prefix was for — is it set, is it the same value as over there,
|
||||||
|
# is it the CB-592 sentinel — and answers none of the ones it should not.
|
||||||
|
# * It never writes anywhere, never contacts the network, and never touches secrets.sh, which is
|
||||||
|
# the operator's file.
|
||||||
|
#
|
||||||
|
# HOW TO RUN IT
|
||||||
|
#
|
||||||
|
# 1. As the operator, in a member's pane (a spawned worker's terminal):
|
||||||
|
# bash scripts/probe-member-credentials.sh
|
||||||
|
# 2. For the comparison row, in your OWN shell — a lead, not a member:
|
||||||
|
# bash scripts/probe-member-credentials.sh --allow-outside-member
|
||||||
|
#
|
||||||
|
# The two outputs side by side are the finding: any name whose hash matches between them is a
|
||||||
|
# credential the member holds in full.
|
||||||
|
#
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
# The names ${SHARED_ENV}/tools/secrets.sh exports, recorded on 2026-08-16 (issue #82). Names only —
|
||||||
|
# this list contains no values and never should. If secrets.sh gains a name, this list goes stale and
|
||||||
|
# the probe silently stops asking about it; that staleness is itself part of what #82's criterion 4
|
||||||
|
# has to solve, so it is called out in the summary rather than hidden.
|
||||||
|
NAMES=(
|
||||||
|
AI_GATEWAY_TOKEN BESZEL_ADMIN_EMAIL BESZEL_ADMIN_PASSWORD
|
||||||
|
BESZEL_HUB_URL BESZEL_KEY BESZEL_UNIVERSAL_TOKEN
|
||||||
|
BRAIN_MCP_TOKEN CF_ACCOUNT_ID CF_API_TOKEN
|
||||||
|
CF_USER_TOKEN CONFLUENCE_API_TOKEN CONFLUENCE_USERNAME
|
||||||
|
CONTEXT7_TOKEN GITEA_HOST GITLAB_OAUTH_CLIENT_SECRET
|
||||||
|
GITLAB_PERSONAL_ACCESS_TOKEN GRAFANA_ADMIN_PASSWORD GRAFANA_ADMIN_USER
|
||||||
|
HASS_TOKEN HW_PASSWORD HW_USER
|
||||||
|
LTMS_API_KEY MEMORY_MCP_TOKEN METRICS_PUSH_TOKEN
|
||||||
|
OPENCODE_AUTOMODE_MODEL TELEGRAM_BOT_TOKEN TELEGRAM_CHAT_ID
|
||||||
|
TS_API_KEY TS_AUTHKEY WORKER_GITEA_TOKEN
|
||||||
|
GITEA_ACCESS_TOKEN
|
||||||
|
)
|
||||||
|
|
||||||
|
allow_outside=0
|
||||||
|
for arg in "$@"; do
|
||||||
|
case "$arg" in
|
||||||
|
--allow-outside-member) allow_outside=1 ;;
|
||||||
|
-h|--help) sed -n '2,40p' "$0"; exit 0 ;;
|
||||||
|
*) echo "unknown argument: $arg" >&2; exit 2 ;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
|
||||||
|
if [ "${BRIDGED_MEMBER:-}" != "1" ] && [ "$allow_outside" -eq 0 ]; then
|
||||||
|
cat >&2 <<'EOF'
|
||||||
|
refusing to run: BRIDGED_MEMBER is not 1, so this is not a member's shell.
|
||||||
|
|
||||||
|
The finding this probe exists for is what a MEMBER holds. Run it in a spawned worker's pane. If you
|
||||||
|
meant to take the comparison reading from your own shell, pass --allow-outside-member and the output
|
||||||
|
will be labelled as such.
|
||||||
|
EOF
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Prefer sha256sum (Linux), fall back to shasum (macOS). If neither exists, report presence and
|
||||||
|
# length only — degraded, but never a value.
|
||||||
|
hasher=""
|
||||||
|
if command -v sha256sum >/dev/null 2>&1; then
|
||||||
|
hasher="sha256sum"
|
||||||
|
elif command -v shasum >/dev/null 2>&1; then
|
||||||
|
hasher="shasum -a 256"
|
||||||
|
fi
|
||||||
|
|
||||||
|
digest() { # value -> first 12 hex chars of its sha256, or "-" when no hasher is available
|
||||||
|
[ -z "$hasher" ] && { printf '%s' "-"; return; }
|
||||||
|
printf '%s' "$1" | $hasher | cut -c1-12
|
||||||
|
}
|
||||||
|
|
||||||
|
if [ "${BRIDGED_MEMBER:-}" = "1" ]; then
|
||||||
|
where="MEMBER (BRIDGED_MEMBER=1)"
|
||||||
|
else
|
||||||
|
where="NOT a member — comparison reading only"
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo "CB-596 credential probe"
|
||||||
|
echo "reading from : $where"
|
||||||
|
echo "shell : ${SHELL:-unknown}"
|
||||||
|
echo "hash : ${hasher:-none available — lengths only}"
|
||||||
|
# Only printed so the two readings can be told apart when they are pasted side by side.
|
||||||
|
echo "host : $(hostname 2>/dev/null || echo unknown)"
|
||||||
|
echo
|
||||||
|
printf '%-30s %-7s %6s %s\n' "NAME" "STATE" "LEN" "SHA256-12"
|
||||||
|
printf '%-30s %-7s %6s %s\n' "------------------------------" "-------" "------" "------------"
|
||||||
|
|
||||||
|
set_count=0
|
||||||
|
for name in "${NAMES[@]}"; do
|
||||||
|
value="${!name:-}"
|
||||||
|
if [ -z "$value" ]; then
|
||||||
|
printf '%-30s %-7s %6s %s\n' "$name" "unset" "-" "-"
|
||||||
|
else
|
||||||
|
set_count=$((set_count + 1))
|
||||||
|
printf '%-30s %-7s %6s %s\n' "$name" "SET" "${#value}" "$(digest "$value")"
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
|
||||||
|
echo
|
||||||
|
echo "$set_count of ${#NAMES[@]} names are set in this shell."
|
||||||
|
echo
|
||||||
|
cat <<'EOF'
|
||||||
|
How to read this:
|
||||||
|
|
||||||
|
* Take the MEMBER reading and the comparison reading, and line them up. A name whose SHA256-12
|
||||||
|
matches on both sides is a credential the member holds in full. That is the finding.
|
||||||
|
* GITEA_ACCESS_TOKEN is the control. CB-592 replaces it with a blocked sentinel, so its hash
|
||||||
|
should DIFFER between the two readings. If it matches, CB-592 is not working and that is the
|
||||||
|
most urgent thing on this page.
|
||||||
|
* AI_GATEWAY_TOKEN matching is expected and correct, not a leak: bridged.yaml names it in
|
||||||
|
`tokenEnv:` for the local and gx profiles, so a member reaching the gateway is by design.
|
||||||
|
* A name that is set here but is NOT in the list above will not appear at all. The list was
|
||||||
|
recorded on 2026-08-16 and does not update itself. Anything added to secrets.sh since then is
|
||||||
|
invisible to this probe — which is the same gap issue #82 criterion 4 asks to close properly.
|
||||||
|
EOF
|
||||||
Executable
+361
@@ -0,0 +1,361 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
#
|
||||||
|
# Rebuild and restart the bridged daemon.
|
||||||
|
#
|
||||||
|
# A merge is not a deployment: the running daemon holds the jar it was started with, so code merged
|
||||||
|
# to main does nothing until this runs. See CLAUDE.md -> "Redeploying the daemon".
|
||||||
|
#
|
||||||
|
# This script exists to turn six remembered traps into one auditable command:
|
||||||
|
#
|
||||||
|
# 1. A piped `mvn` hides BUILD FAILURE behind a zero exit, so the build here is never piped.
|
||||||
|
# 2. The daemon must start from a LOGIN shell, or the tokens it hands to members are empty:
|
||||||
|
# WORKER_GITEA_TOKEN (workers cannot open a PR) and AI_GATEWAY_TOKEN (401 at llm.ltms.dev).
|
||||||
|
# Both are read from the DAEMON's own environment at spawn time, so a value added to
|
||||||
|
# secrets.sh after startup is absent. Nothing logs this here, so the script checks and says
|
||||||
|
# so — and since CB-594, bridged's own startup log says so too, by env var name.
|
||||||
|
# 3. An old daemon that never actually died looks identical from the outside, so the script waits
|
||||||
|
# for the process to exit and for the port to free before it starts a new one.
|
||||||
|
# 4. "It started" is not "it works": the script polls /healthz until it answers, and reports the
|
||||||
|
# herdr protocol number, because healthz can be green while every spawn fails on a protocol
|
||||||
|
# mismatch.
|
||||||
|
# 5. Restarting under live members drops their tickets, so the script refuses unless you confirm
|
||||||
|
# the fleet is drained.
|
||||||
|
# 6. CB-594 — the launchd agent (deploy/dev.ltms.bridged.plist), if installed and loaded, is a
|
||||||
|
# SECOND supervisor: its KeepAlive.SuccessfulExit=false restarts the daemon on any nonzero
|
||||||
|
# exit, and a bare SIGTERM makes this JVM exit 143 even with its shutdown hook running to
|
||||||
|
# completion (measured — see the CB-594 report). A plain `kill` here would race launchd's own
|
||||||
|
# restart of the OLD jar. So this script detects whether the agent is loaded and, only then,
|
||||||
|
# swaps `kill` + manual `nohup` for `launchctl unload`/`load` — the one supervisor in control
|
||||||
|
# at any moment is whichever one you asked to act, never both.
|
||||||
|
#
|
||||||
|
# Usage:
|
||||||
|
# scripts/redeploy-bridged.sh # build, confirm, restart, verify
|
||||||
|
# scripts/redeploy-bridged.sh --yes # skip the drain confirmation (fleet already checked)
|
||||||
|
# scripts/redeploy-bridged.sh --no-build # restart the jar already on disk
|
||||||
|
# scripts/redeploy-bridged.sh --check # report state and exit; changes nothing
|
||||||
|
#
|
||||||
|
# Exits non-zero on any failure. A failed build never stops the running daemon.
|
||||||
|
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||||
|
BRIDGED="$REPO/bridged"
|
||||||
|
JAR="$BRIDGED/target/bridged.jar"
|
||||||
|
OUT="$BRIDGED/bridged.out"
|
||||||
|
# Matches BOTH the absolute form and the relative `java -jar target/bridged.jar` a hand-start
|
||||||
|
# produces from inside bridged/. Anchoring on the absolute path alone was a real bug: the daemon
|
||||||
|
# restarted correctly and the script still reported "no process appeared", because it launched with
|
||||||
|
# a relative path and then looked for an absolute one.
|
||||||
|
PATTERN='target/bridged.jar'
|
||||||
|
HEALTH='http://127.0.0.1:8765/healthz'
|
||||||
|
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
|
||||||
|
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start
|
||||||
|
|
||||||
|
# CB-594: the launchd agent this script must not fight with (see trap 6 above).
|
||||||
|
LAUNCHD_LABEL='dev.ltms.bridged'
|
||||||
|
LAUNCHD_PLIST="$HOME/Library/LaunchAgents/$LAUNCHD_LABEL.plist"
|
||||||
|
|
||||||
|
DO_BUILD=1; ASSUME_YES=0; CHECK_ONLY=0
|
||||||
|
for arg in "$@"; do
|
||||||
|
case "$arg" in
|
||||||
|
--yes|-y) ASSUME_YES=1 ;;
|
||||||
|
--no-build) DO_BUILD=0 ;;
|
||||||
|
--check) CHECK_ONLY=1 ;;
|
||||||
|
-h|--help) sed -n '3,37p' "${BASH_SOURCE[0]}"; exit 0 ;;
|
||||||
|
*) echo "unknown option: $arg (try --help)" >&2; exit 2 ;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
|
||||||
|
say() { printf '\n\033[1m== %s\033[0m\n' "$*"; }
|
||||||
|
ok() { printf ' ok %s\n' "$*"; }
|
||||||
|
warn() { printf ' WARN %s\n' "$*"; }
|
||||||
|
die() { printf '\n FAIL %s\n\n' "$*" >&2; exit 1; }
|
||||||
|
|
||||||
|
jar_id() { [ -f "$JAR" ] && shasum -a 256 "$JAR" | cut -c1-12 || echo "absent"; }
|
||||||
|
running_pid() { pgrep -f "$PATTERN" || true; }
|
||||||
|
# `launchctl list <label>` exits 0 iff the label is loaded (registered with launchd) — true whether
|
||||||
|
# or not it is currently running, which is exactly "supervision is active" for our purposes. Read-
|
||||||
|
# only: neither helper below changes anything, so both are also safe under --check.
|
||||||
|
launchd_installed() { [ -f "$LAUNCHD_PLIST" ]; }
|
||||||
|
launchd_loaded() { launchctl list "$LAUNCHD_LABEL" >/dev/null 2>&1; }
|
||||||
|
|
||||||
|
# CB-600: the script computes its own log path from where it sits on disk (REPO, above); the
|
||||||
|
# plist hard-codes an absolute StandardOutPath. Nothing forced the two to agree — if this script
|
||||||
|
# were ever run from a checkout other than the one the loaded plist names, launchd would start and
|
||||||
|
# log the daemon correctly, while every check below (the fresh "bridged listening" line, the
|
||||||
|
# ERROR-count scan) would read a different, empty or stale file and the script would report a
|
||||||
|
# clean restart while the daemon crash-loops. Pure and side-effect-free besides `die`/`ok` — reads
|
||||||
|
# the two paths, resolves them, compares — so it never touches launchd or the daemon and can be
|
||||||
|
# exercised by sourcing this script (see the SOURCED guard below) without installing the agent.
|
||||||
|
check_log_path_matches_plist() {
|
||||||
|
local script_out="$1" plist_path="$2"
|
||||||
|
local plist_out resolved_out resolved_plist_out
|
||||||
|
# Checked by exit status, not by emptiness: on a missing file/key PlistBuddy exits nonzero but
|
||||||
|
# still writes a message ("File Doesn't Exist, Will Create: ...") that command substitution
|
||||||
|
# would happily capture as if it were the real value — testing only `-z` missed that case.
|
||||||
|
if ! plist_out="$(/usr/libexec/PlistBuddy -c 'Print :StandardOutPath' "$plist_path" 2>/dev/null)" \
|
||||||
|
|| [ -z "$plist_out" ]; then
|
||||||
|
die "launchd agent is loaded but PlistBuddy could not read StandardOutPath from
|
||||||
|
$plist_path
|
||||||
|
— cannot verify the daemon logs where this script is about to look. Fix the plist before
|
||||||
|
redeploying supervised."
|
||||||
|
fi
|
||||||
|
resolved_out="$(cd "$(dirname "$script_out")" 2>/dev/null && pwd -P)/$(basename "$script_out")" || true
|
||||||
|
resolved_plist_out="$(cd "$(dirname "$plist_out")" 2>/dev/null && pwd -P)/$(basename "$plist_out")" || true
|
||||||
|
if [ -z "$resolved_out" ] || [ -z "$resolved_plist_out" ] || [ "$resolved_out" != "$resolved_plist_out" ]; then
|
||||||
|
die "log path mismatch — this script reads
|
||||||
|
$script_out (resolved: ${resolved_out:-<directory does not exist>})
|
||||||
|
but the loaded plist's StandardOutPath is
|
||||||
|
$plist_out (resolved: ${resolved_plist_out:-<directory does not exist>})
|
||||||
|
Under supervision the daemon writes to the PLIST's path, not necessarily this script's — every
|
||||||
|
post-restart check below (the fresh 'bridged listening' line, the ERROR-count scan) would read
|
||||||
|
the wrong file and could report a clean restart while the daemon crash-loops. Fix the mismatch
|
||||||
|
(move this checkout to match the plist, or edit the plist's StandardOutPath/StandardErrorPath)
|
||||||
|
before redeploying supervised."
|
||||||
|
fi
|
||||||
|
ok "log path check: script and plist agree ($resolved_out)"
|
||||||
|
}
|
||||||
|
|
||||||
|
# CB-600: sourceable for testing. When this file is SOURCED (not executed) it stops here — nothing
|
||||||
|
# below runs — so a test harness can `source` it to call check_log_path_matches_plist (or the
|
||||||
|
# other pure helpers above) against a throwaway plist fixture without ever reaching the mutating
|
||||||
|
# flow (build/stop/start) or touching the real daemon or launchd. On a normal `./redeploy-bridged.sh`
|
||||||
|
# invocation `(return 0 2>/dev/null)` fails (return is illegal at top level of an executed script),
|
||||||
|
# so this whole block is a no-op and every line below still runs exactly as before.
|
||||||
|
if (return 0 2>/dev/null); then
|
||||||
|
return 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------- report state
|
||||||
|
|
||||||
|
say "current state"
|
||||||
|
OLD_PID="$(running_pid)"
|
||||||
|
if [ -n "$OLD_PID" ]; then
|
||||||
|
ok "daemon running, pid $OLD_PID"
|
||||||
|
else
|
||||||
|
warn "no daemon running — this will be a cold start"
|
||||||
|
fi
|
||||||
|
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
|
||||||
|
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
|
||||||
|
|
||||||
|
# CB-594: supervision state. Installed and loaded are different facts — a copied-but-never-loaded
|
||||||
|
# plist supervises nothing, and a loaded label with no file backing it (rare, but possible after an
|
||||||
|
# edited/moved plist) is still what launchd will act on.
|
||||||
|
if launchd_installed; then
|
||||||
|
ok "launchd agent installed: $LAUNCHD_PLIST"
|
||||||
|
else
|
||||||
|
warn "launchd agent NOT installed (no supervision — a crash will not restart the daemon)."
|
||||||
|
fi
|
||||||
|
SUPERVISED=0
|
||||||
|
if launchd_loaded; then
|
||||||
|
SUPERVISED=1
|
||||||
|
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
|
||||||
|
# CB-600: fail loudly here, before ANY other check runs, if this script and the loaded plist
|
||||||
|
# would read different log files — every check after this point is worthless otherwise.
|
||||||
|
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
|
||||||
|
else
|
||||||
|
warn "launchd agent not loaded — this script is the only thing that will restart the daemon."
|
||||||
|
fi
|
||||||
|
|
||||||
|
# The trap with no log line. Checked in a LOGIN shell, because that is how the daemon is started
|
||||||
|
# below. Never prints the value — only whether it resolved.
|
||||||
|
if zsh -lc '[ -n "${WORKER_GITEA_TOKEN:-}" ]' 2>/dev/null; then
|
||||||
|
ok "WORKER_GITEA_TOKEN resolves in a login shell"
|
||||||
|
else
|
||||||
|
warn "WORKER_GITEA_TOKEN is EMPTY in a login shell."
|
||||||
|
warn "The daemon will start fine and workers will silently fail to open PRs."
|
||||||
|
warn "Fix \${SHARED_ENV}/tools/secrets.sh before relying on worker checkpoints."
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Same trap, second variable (CB-591). A profile's `tokenEnv:` is resolved from the DAEMON's own
|
||||||
|
# process environment by HerdrPeerLauncher.resolveEnv, so a token added to secrets.sh after the
|
||||||
|
# daemon started is simply absent. The launcher then injects an empty token and llm.ltms.dev answers
|
||||||
|
# 401 — long after the restart, and with nothing tying the two together.
|
||||||
|
if zsh -lc '[ -n "${AI_GATEWAY_TOKEN:-}" ]' 2>/dev/null; then
|
||||||
|
ok "AI_GATEWAY_TOKEN resolves in a login shell"
|
||||||
|
else
|
||||||
|
warn "AI_GATEWAY_TOKEN is EMPTY in a login shell."
|
||||||
|
warn "Any profile whose tokenEnv is AI_GATEWAY_TOKEN will get an empty token and 401 at the gateway."
|
||||||
|
warn "This only matters once a profile points at llm.ltms.dev — harmless before that."
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [ "$CHECK_ONLY" = 1 ]; then
|
||||||
|
say "--check: nothing changed"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------- build
|
||||||
|
# Deliberately before the stop: a failed build must never leave the fleet down.
|
||||||
|
|
||||||
|
if [ "$DO_BUILD" = 1 ]; then
|
||||||
|
say "build"
|
||||||
|
BUILD_LOG="$(mktemp -t bridged-build)"
|
||||||
|
echo " log: $BUILD_LOG"
|
||||||
|
if ! mvn -f "$BRIDGED/pom.xml" clean install > "$BUILD_LOG" 2>&1; then
|
||||||
|
grep -E 'ERROR|BUILD FAILURE|Tests run:.*Failures: [1-9]|Tests run:.*Errors: [1-9]' "$BUILD_LOG" \
|
||||||
|
| head -20 || true
|
||||||
|
die "build failed — the running daemon was NOT touched. Full log: $BUILD_LOG"
|
||||||
|
fi
|
||||||
|
grep -E '^\[INFO\] Tests run:.*Failures' "$BUILD_LOG" | tail -1 | sed 's/^\[INFO\] / /' || true
|
||||||
|
ok "BUILD SUCCESS"
|
||||||
|
ok "jar now: $(jar_id)"
|
||||||
|
else
|
||||||
|
say "build skipped (--no-build)"
|
||||||
|
fi
|
||||||
|
|
||||||
|
[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build"
|
||||||
|
|
||||||
|
# ----------------------------------------------------------------- drain gate
|
||||||
|
|
||||||
|
if [ -n "$OLD_PID" ] && [ "$ASSUME_YES" = 0 ]; then
|
||||||
|
say "drain check"
|
||||||
|
echo " A restart drops every in-flight ticket and rendezvous. A member's report"
|
||||||
|
echo " is NOT recoverable once its ticket is gone."
|
||||||
|
echo
|
||||||
|
echo " Confirm with bridge_list that no members are live, and bridge_poll anything"
|
||||||
|
echo " you still want, BEFORE continuing."
|
||||||
|
echo
|
||||||
|
read -r -p " Fleet drained? type yes to restart: " reply
|
||||||
|
[ "$reply" = "yes" ] || die "aborted — nothing changed"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------ stop
|
||||||
|
#
|
||||||
|
# CB-594: when SUPERVISED, launchd owns the stop — never a raw `kill` here. A bare SIGTERM makes
|
||||||
|
# this JVM exit 143 even with its shutdown hook running to completion (verified separately: a
|
||||||
|
# throwaway Java process with an equivalent shutdown hook, sent SIGTERM from a login shell that
|
||||||
|
# could `wait` on it directly, reported exit code 143 every time — never 0). launchd's
|
||||||
|
# KeepAlive.SuccessfulExit=false treats any nonzero exit as a crash and restarts the OLD jar,
|
||||||
|
# which would race this script's own restart of the NEW one. `launchctl unload` avoids that race
|
||||||
|
# by deregistering the job first, so no KeepAlive is left armed when the process actually stops.
|
||||||
|
|
||||||
|
if [ -n "$OLD_PID" ]; then
|
||||||
|
say "stop"
|
||||||
|
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)" # verify a FRESH line appears later
|
||||||
|
if [ "$SUPERVISED" = 1 ]; then
|
||||||
|
echo " supervision is ON: using 'launchctl unload' (not kill) so launchd's own KeepAlive"
|
||||||
|
echo " cannot restart the OLD jar out from under this script — see the CB-594 comment above."
|
||||||
|
launchctl unload -w "$LAUNCHD_PLIST" \
|
||||||
|
|| die "launchctl unload failed — the daemon may still be under supervision; investigate before retrying"
|
||||||
|
else
|
||||||
|
kill "$OLD_PID"
|
||||||
|
fi
|
||||||
|
for _ in $(seq "$STOP_WAIT"); do
|
||||||
|
[ -z "$(running_pid)" ] && break
|
||||||
|
sleep 1
|
||||||
|
done
|
||||||
|
if [ -n "$(running_pid)" ]; then
|
||||||
|
die "pid $OLD_PID still alive after ${STOP_WAIT}s. Not escalating to kill -9 automatically:
|
||||||
|
the shutdown hook releases sessions and worktrees in order, and killing it hard can
|
||||||
|
leave worktrees and panes behind. Investigate, then kill -9 by hand if you accept that."
|
||||||
|
fi
|
||||||
|
ok "pid $OLD_PID exited"
|
||||||
|
elif [ "$SUPERVISED" = 1 ]; then
|
||||||
|
# Loaded but not currently running (e.g. throttled after a crash loop). Unload it anyway so the
|
||||||
|
# start step below does a clean load, never a load stacked on an already-loaded label.
|
||||||
|
say "stop"
|
||||||
|
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
|
||||||
|
launchctl unload -w "$LAUNCHD_PLIST" 2>/dev/null || true
|
||||||
|
ok "launchd agent unloaded (was already not running)"
|
||||||
|
else
|
||||||
|
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------ start
|
||||||
|
# Unsupervised: login shell (zsh -l) is what puts the secrets on the daemon's environment, and cwd
|
||||||
|
# must be bridged/ because the daemon resolves bridged.yaml, logs/ and target/ relative to it.
|
||||||
|
# Supervised: launchd does both — deploy/dev.ltms.bridged.plist points ProgramArguments at
|
||||||
|
# scripts/bridged-launchd-wrapper.sh (CB-594), which is what execs the login shell in launchd's
|
||||||
|
# place, and WorkingDirectory in the plist already pins bridged/.
|
||||||
|
|
||||||
|
say "start"
|
||||||
|
if [ "$SUPERVISED" = 1 ]; then
|
||||||
|
echo " supervision is ON: using 'launchctl load' so launchd starts and keeps supervising this"
|
||||||
|
echo " process, instead of a manual nohup that launchd would know nothing about."
|
||||||
|
# CB-600: 'launchctl unload -w' above already persisted Disabled=true for this label. A load -w
|
||||||
|
# that succeeds clears it; a load -w that FAILS leaves the agent both stopped and disabled — worse
|
||||||
|
# than before this script ran, because a later reboot or login will not bring it back either. One
|
||||||
|
# retry covers a transient race (e.g. launchd not yet fully done deregistering); if it still fails,
|
||||||
|
# die with the exact recovery command rather than a bare "failed".
|
||||||
|
if ! launchctl load -w "$LAUNCHD_PLIST" 2>/dev/null; then
|
||||||
|
warn "launchctl load failed on the first attempt — retrying once after a short pause"
|
||||||
|
sleep 2
|
||||||
|
launchctl load -w "$LAUNCHD_PLIST" || die "launchctl load failed twice.
|
||||||
|
The agent is now STOPPED and DISABLED — it will NOT come back on its own, not even after a
|
||||||
|
reboot or login, because 'launchctl unload -w' above persisted Disabled=true and load -w
|
||||||
|
never got the chance to clear it. Recover with:
|
||||||
|
launchctl load -w \"$LAUNCHD_PLIST\"
|
||||||
|
If that still fails, check 'launchctl list $LAUNCHD_LABEL', validate the plist with
|
||||||
|
'plutil -lint \"$LAUNCHD_PLIST\"', and check $OUT before assuming a retry will succeed."
|
||||||
|
fi
|
||||||
|
else
|
||||||
|
# Absolute jar path so `ps` names which checkout is running.
|
||||||
|
( cd "$BRIDGED" && zsh -lc "nohup java -jar '$JAR' >> bridged.out 2>&1 &" )
|
||||||
|
fi
|
||||||
|
|
||||||
|
for _ in $(seq 10); do
|
||||||
|
NEW_PID="$(running_pid)"
|
||||||
|
[ -n "$NEW_PID" ] && break
|
||||||
|
sleep 1
|
||||||
|
done
|
||||||
|
[ -n "${NEW_PID:-}" ] || die "no process appeared. Last lines of $OUT:
|
||||||
|
$(tail -20 "$OUT" 2>/dev/null)"
|
||||||
|
[ "$NEW_PID" != "${OLD_PID:-}" ] || die "pid unchanged ($NEW_PID) — the old daemon never died"
|
||||||
|
ok "started, pid $NEW_PID"
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------ verify
|
||||||
|
|
||||||
|
say "verify"
|
||||||
|
|
||||||
|
HEALTH_BODY=""
|
||||||
|
for _ in $(seq "$HEALTH_WAIT"); do
|
||||||
|
if HEALTH_BODY="$(curl -fsS --max-time 2 "$HEALTH" 2>/dev/null)"; then break; fi
|
||||||
|
HEALTH_BODY=""
|
||||||
|
sleep 1
|
||||||
|
done
|
||||||
|
|
||||||
|
if [ -z "$HEALTH_BODY" ]; then
|
||||||
|
# 503 still means the daemon is up — it means herdr is unreachable. Say which.
|
||||||
|
CODE="$(curl -s -o /dev/null -w '%{http_code}' --max-time 2 "$HEALTH" 2>/dev/null || echo 000)"
|
||||||
|
if [ "$CODE" = "503" ]; then
|
||||||
|
warn "/healthz answers 503 degraded — the daemon is up but herdr is unreachable."
|
||||||
|
warn "Spawns will fail. Check herdr before delegating anything."
|
||||||
|
curl -s --max-time 2 "$HEALTH" 2>/dev/null | head -3 || true
|
||||||
|
else
|
||||||
|
die "/healthz never answered within ${HEALTH_WAIT}s (last code: $CODE). Last lines of $OUT:
|
||||||
|
$(tail -30 "$OUT" 2>/dev/null)"
|
||||||
|
fi
|
||||||
|
else
|
||||||
|
ok "/healthz 200 — $HEALTH_BODY"
|
||||||
|
warn "healthz green only proves herdr ANSWERS. If its protocol number changed, spawns can still"
|
||||||
|
warn "fail — prove a real spawn before trusting the fleet."
|
||||||
|
fi
|
||||||
|
|
||||||
|
# A fresh listening line, strictly after the restart mark. An old daemon that never died would
|
||||||
|
# otherwise let an old line pass for a new one.
|
||||||
|
if tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null | grep -q 'bridged listening'; then
|
||||||
|
ok "$(tail -n "+$((RESTART_MARK + 1))" "$OUT" | grep 'bridged listening' | tail -1)"
|
||||||
|
else
|
||||||
|
warn "no fresh 'bridged listening' line after the restart — check $OUT yourself"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Config keys the daemon accepted or deferred at boot. This is usually WHY you restarted.
|
||||||
|
say "config at boot"
|
||||||
|
tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null \
|
||||||
|
| grep -iE 'deferred|classification:|fleet health:|coverage' | tail -8 | sed 's/^/ /' \
|
||||||
|
|| echo " (nothing reported)"
|
||||||
|
|
||||||
|
# Errors since the restart, anchored to the marker so old noise cannot leak in.
|
||||||
|
ERRS="$(tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null | grep -cE ' (ERROR|SEVERE) ' || true)"
|
||||||
|
say "result"
|
||||||
|
ok "pid $NEW_PID, jar $(jar_id)"
|
||||||
|
if [ "${ERRS:-0}" -gt 0 ]; then
|
||||||
|
warn "$ERRS ERROR lines since restart:"
|
||||||
|
tail -n "+$((RESTART_MARK + 1))" "$OUT" | grep -E ' (ERROR|SEVERE) ' | tail -5 | sed 's/^/ /'
|
||||||
|
else
|
||||||
|
ok "no ERROR lines since restart"
|
||||||
|
fi
|
||||||
|
echo
|
||||||
|
echo " Next: call bridge_whoami and confirm it still answers 'primary'. A lead whose tab label"
|
||||||
|
echo " no longer matches fleet.leaders.*.tab is demoted to worker and refuses orchestration."
|
||||||
|
echo
|
||||||
Reference in New Issue
Block a user