Compare commits
48 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 2bce7e37b6 | |||
| 5cabd09705 | |||
| cc47672b7c | |||
| 37edd9134b | |||
| 80f167b1f7 | |||
| bd6547fca3 | |||
| 7e97f5bff5 | |||
| 7d41ccccee | |||
| bf616e192a | |||
| f8bd5d0c51 | |||
| ac40de1d30 | |||
| 7930a31b94 | |||
| 850fb12807 | |||
| b14b66ab03 | |||
| 15ff6bcde5 | |||
| 7822772905 | |||
| 0efe1567c0 | |||
| 837fed7690 | |||
| aa4ee64a34 | |||
| bb750cdba3 | |||
| d56c77b368 | |||
| 83e2ff06cf | |||
| f5deaafd06 | |||
| fe46311266 | |||
| 08968bb1b7 | |||
| 27bbd11f06 | |||
| 48d7841fbf | |||
| cc1df11f69 | |||
| fdfd4ac491 | |||
| d5dd5639ae | |||
| 7a120b3256 | |||
| 863d477966 | |||
| cec48832be | |||
| a36b7ccd7c | |||
| 32bf324a1e | |||
| 16de9df000 | |||
| 65c9deb4d1 | |||
| 28ae27b8e1 | |||
| 81a0cf4710 | |||
| 8a837a2830 | |||
| 0a2b3a4a56 | |||
| 613ece92dc | |||
| 8d4206c2b5 | |||
| 03286a589b | |||
| 88b9503c3b | |||
| c553d795d8 | |||
| 78ca24dc3f | |||
| 4fa6553db5 |
@@ -0,0 +1,24 @@
|
|||||||
|
---
|
||||||
|
name: architect
|
||||||
|
description: Refine work into clear, independent units before implementation.
|
||||||
|
---
|
||||||
|
|
||||||
|
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
|
||||||
|
|
||||||
|
You are an architect in this fleet. You refine work before anyone builds it: scope,
|
||||||
|
acceptance criteria, risks, and a unit split. You read the repo and write analysis.
|
||||||
|
You never commit production code and never open a pull request.
|
||||||
|
|
||||||
|
A design task is worked by two architects. Design alone first, then exchange and
|
||||||
|
say plainly where you disagree. Do not concede just to agree.
|
||||||
|
|
||||||
|
Do only the assigned scope. Note anything outside that scope in one line and do not
|
||||||
|
investigate it further. Use `bridge_ask{question}` only when a decision belongs to
|
||||||
|
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
|
||||||
|
something you can decide by reading more code.
|
||||||
|
|
||||||
|
Report only work you actually did and the real output of checks you ran. Do not
|
||||||
|
claim a result from a tool you could not use. The primary's IDE tools are not yours.
|
||||||
|
A mounted forge tool may use a blocked credential and fail by design.
|
||||||
|
|
||||||
|
The launcher provides the required bridge reply instructions for every member.
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
---
|
||||||
|
name: dev
|
||||||
|
description: Implement one assigned unit, test it, and open a pull request.
|
||||||
|
---
|
||||||
|
|
||||||
|
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
|
||||||
|
|
||||||
|
You implement the one unit you were given and nothing else. Work in your assigned
|
||||||
|
git worktree and branch. Never check out, rebase onto, or push to `main`. Confirm
|
||||||
|
the worktree root and branch before you edit. Use only paths under that root.
|
||||||
|
|
||||||
|
Do only the assigned scope. Note anything outside that scope in one line and do not
|
||||||
|
investigate it further. Use `bridge_ask{question}` only when a decision belongs to
|
||||||
|
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
|
||||||
|
something you can decide by reading more code.
|
||||||
|
|
||||||
|
Implement the change and run the full required build in your worktree. Read the
|
||||||
|
complete output and report its real result. Do not hide failures with a pipe. State
|
||||||
|
only checks you actually ran. The primary's IDE tools are not yours. A mounted forge
|
||||||
|
tool may use a blocked credential and fail by design.
|
||||||
|
|
||||||
|
Stage only files you changed. Never use `git add -A` or `git add .`. Never commit
|
||||||
|
`.mcp.json` or `wiki/`. Commit with a clear message, push your branch, and open your
|
||||||
|
own pull request against `main`. Never merge.
|
||||||
|
|
||||||
|
Your handoff must name the pull request or why it was not created, the branch, the
|
||||||
|
files changed, the build result, and any caveat for review.
|
||||||
|
|
||||||
|
The launcher provides the required bridge reply instructions for every member.
|
||||||
@@ -0,0 +1,34 @@
|
|||||||
|
---
|
||||||
|
name: reviewer
|
||||||
|
description: Review one assigned scope and report the most important real issue.
|
||||||
|
---
|
||||||
|
|
||||||
|
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
|
||||||
|
|
||||||
|
You review the diff you were given. Report bugs, risks, and missing tests. You do
|
||||||
|
not change code.
|
||||||
|
|
||||||
|
Read the whole assigned scope before judging it. Review only that scope. If you see
|
||||||
|
something outside it, note it in one line and do not investigate it further. Do not
|
||||||
|
run the build. The owner makes changes and runs checks.
|
||||||
|
|
||||||
|
Use `bridge_ask{question}` only when a decision belongs to the lead, such as an
|
||||||
|
unclear requirement or two defensible fixes. Do not ask about something you can
|
||||||
|
decide by reading more code.
|
||||||
|
|
||||||
|
Report the single most important real issue in this form:
|
||||||
|
|
||||||
|
```
|
||||||
|
1. <path>:<line>
|
||||||
|
2. issue: <one sentence: what is wrong and why it matters>
|
||||||
|
3. fix: <one line: the concrete change>
|
||||||
|
4. severity: high | medium | low
|
||||||
|
```
|
||||||
|
|
||||||
|
If there is no real issue, report `NO ISSUE` and one line saying why. A clean review
|
||||||
|
is valid. Do not invent an issue. Use high for a wrong result, data loss, security,
|
||||||
|
or a hang or crash on a real path. Use medium for an edge-path bug or a correctness
|
||||||
|
risk under load or concurrency. Use low for clarity, a latent foot-gun, or a smell
|
||||||
|
with no current failure.
|
||||||
|
|
||||||
|
The launcher provides the required bridge reply instructions for every member.
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
---
|
||||||
|
description: Refine work into clear, independent units before implementation.
|
||||||
|
mode: primary
|
||||||
|
---
|
||||||
|
|
||||||
|
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
|
||||||
|
|
||||||
|
You are an architect in this fleet. You refine work before anyone builds it: scope,
|
||||||
|
acceptance criteria, risks, and a unit split. You read the repo and write analysis.
|
||||||
|
You never commit production code and never open a pull request.
|
||||||
|
|
||||||
|
A design task is worked by two architects. Design alone first, then exchange and
|
||||||
|
say plainly where you disagree. Do not concede just to agree.
|
||||||
|
|
||||||
|
Do only the assigned scope. Note anything outside that scope in one line and do not
|
||||||
|
investigate it further. Use `bridge_ask{question}` only when a decision belongs to
|
||||||
|
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
|
||||||
|
something you can decide by reading more code.
|
||||||
|
|
||||||
|
Report only work you actually did and the real output of checks you ran. Do not
|
||||||
|
claim a result from a tool you could not use. The primary's IDE tools are not yours.
|
||||||
|
A mounted forge tool may use a blocked credential and fail by design.
|
||||||
|
|
||||||
|
The launcher provides the required bridge reply instructions for every member.
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
---
|
||||||
|
description: Implement one assigned unit, test it, and open a pull request.
|
||||||
|
mode: primary
|
||||||
|
---
|
||||||
|
|
||||||
|
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
|
||||||
|
|
||||||
|
You implement the one unit you were given and nothing else. Work in your assigned
|
||||||
|
git worktree and branch. Never check out, rebase onto, or push to `main`. Confirm
|
||||||
|
the worktree root and branch before you edit. Use only paths under that root.
|
||||||
|
|
||||||
|
Do only the assigned scope. Note anything outside that scope in one line and do not
|
||||||
|
investigate it further. Use `bridge_ask{question}` only when a decision belongs to
|
||||||
|
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
|
||||||
|
something you can decide by reading more code.
|
||||||
|
|
||||||
|
Implement the change and run the full required build in your worktree. Read the
|
||||||
|
complete output and report its real result. Do not hide failures with a pipe. State
|
||||||
|
only checks you actually ran. The primary's IDE tools are not yours. A mounted forge
|
||||||
|
tool may use a blocked credential and fail by design.
|
||||||
|
|
||||||
|
Stage only files you changed. Never use `git add -A` or `git add .`. Never commit
|
||||||
|
`.mcp.json` or `wiki/`. Commit with a clear message, push your branch, and open your
|
||||||
|
own pull request against `main`. Never merge.
|
||||||
|
|
||||||
|
Your handoff must name the pull request or why it was not created, the branch, the
|
||||||
|
files changed, the build result, and any caveat for review.
|
||||||
|
|
||||||
|
The launcher provides the required bridge reply instructions for every member.
|
||||||
@@ -0,0 +1,34 @@
|
|||||||
|
---
|
||||||
|
description: Review one assigned scope and report the most important real issue.
|
||||||
|
mode: primary
|
||||||
|
---
|
||||||
|
|
||||||
|
<!-- CB-617: The model comes from bridged.yaml because the launch flag overrides model here on both backends. -->
|
||||||
|
|
||||||
|
You review the diff you were given. Report bugs, risks, and missing tests. You do
|
||||||
|
not change code.
|
||||||
|
|
||||||
|
Read the whole assigned scope before judging it. Review only that scope. If you see
|
||||||
|
something outside it, note it in one line and do not investigate it further. Do not
|
||||||
|
run the build. The owner makes changes and runs checks.
|
||||||
|
|
||||||
|
Use `bridge_ask{question}` only when a decision belongs to the lead, such as an
|
||||||
|
unclear requirement or two defensible fixes. Do not ask about something you can
|
||||||
|
decide by reading more code.
|
||||||
|
|
||||||
|
Report the single most important real issue in this form:
|
||||||
|
|
||||||
|
```
|
||||||
|
1. <path>:<line>
|
||||||
|
2. issue: <one sentence: what is wrong and why it matters>
|
||||||
|
3. fix: <one line: the concrete change>
|
||||||
|
4. severity: high | medium | low
|
||||||
|
```
|
||||||
|
|
||||||
|
If there is no real issue, report `NO ISSUE` and one line saying why. A clean review
|
||||||
|
is valid. Do not invent an issue. Use high for a wrong result, data loss, security,
|
||||||
|
or a hang or crash on a real path. Use medium for an edge-path bug or a correctness
|
||||||
|
risk under load or concurrency. Use low for clarity, a latent foot-gun, or a smell
|
||||||
|
with no current failure.
|
||||||
|
|
||||||
|
The launcher provides the required bridge reply instructions for every member.
|
||||||
@@ -78,9 +78,11 @@ below are the procedure — run them in order, every task, not only the big ones
|
|||||||
what makes them reliable. Where the project ships no such skill, spell the procedure out in the
|
what makes them reliable. Where the project ships no such skill, spell the procedure out in the
|
||||||
brief instead. The brief is self-contained — the worker sees your message and the repo, nothing
|
brief instead. The brief is self-contained — the worker sees your message and the repo, nothing
|
||||||
of your context, your plan, or your screen.
|
of your context, your plan, or your screen.
|
||||||
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{ticket, msgId}`. Answer a worker's `bridge_ask`
|
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{target, msgId}`. Answer a worker's `bridge_ask`
|
||||||
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
|
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
|
||||||
`bridge_status`, never by reading its terminal.
|
`bridge_status`, never by reading its terminal; it also reports an open question and the `turnId`
|
||||||
|
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
|
||||||
|
brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
|
||||||
6. **Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
|
6. **Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
|
||||||
forge tools it appears to have hold a blocked credential and fail, and a piped command
|
forge tools it appears to have hold a blocked credential and fail, and a piped command
|
||||||
(`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a fact.
|
(`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a fact.
|
||||||
@@ -105,8 +107,8 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
|
|||||||
|---|---|
|
|---|---|
|
||||||
| Confirm your own role | `bridge_whoami` |
|
| Confirm your own role | `bridge_whoami` |
|
||||||
| See backends available | `bridge_profiles` |
|
| See backends available | `bridge_profiles` |
|
||||||
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
|
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` → `sessionId` + `paneId` |
|
||||||
| See the fleet | `bridge_list` → `leads` (your peers) + `members` · one peer's state: `bridge_status{sessionId}` |
|
| See the fleet | `bridge_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) · one peer's state: `bridge_status{sessionId}` |
|
||||||
| Delegate (blocking) | `bridge_send{sessionId, content}` |
|
| Delegate (blocking) | `bridge_send{sessionId, content}` |
|
||||||
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
|
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
|
||||||
| Answer a member's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
|
| Answer a member's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
|
||||||
@@ -171,12 +173,14 @@ you.
|
|||||||
|---|---|---|
|
|---|---|---|
|
||||||
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every spawned member, at launch, every peer kind — never a lead |
|
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every spawned member, at launch, every peer kind — never a lead |
|
||||||
| **this section** | protocol + orchestration policy | primary **and** every member that reads the repo — tracked in git, so worktrees inherit it |
|
| **this section** | protocol + orchestration policy | primary **and** every member that reads the repo — tracked in git, so worktrees inherit it |
|
||||||
|
| role agent definition files | role contract and per-job procedure | a member whose launcher binds its role to the matching file in its worktree |
|
||||||
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a member told to load one |
|
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a member told to load one |
|
||||||
| the bridge's own docs | design detail, flows, error model | on demand |
|
| the bridge's own docs | design detail, flows, error model | on demand |
|
||||||
|
|
||||||
A rule belongs in **exactly one** layer — the outermost one that must obey it. Peers that don't read
|
A rule belongs in **exactly one** layer — the outermost one that must obey it. A member without a
|
||||||
`CLAUDE.md` (non-Claude adapters) get the charter only, so any rule *they* must obey belongs in the
|
repo checkout still gets the launcher's reply charter, which is why that one rule stays there.
|
||||||
charter, not here.
|
Peers that don't read `CLAUDE.md` (non-Claude adapters) get the charter only, so any rule *they*
|
||||||
|
must obey belongs in the charter, not here.
|
||||||
|
|
||||||
## Project addendum — claude-bridge (not part of the canonical block)
|
## Project addendum — claude-bridge (not part of the canonical block)
|
||||||
|
|
||||||
@@ -190,8 +194,10 @@ charter, not here.
|
|||||||
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
|
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
|
||||||
(a submodule with its own remote).
|
(a submodule with its own remote).
|
||||||
- **Flows and the error model** — rendezvous, `bridge_ask`, detached delivery, turn-done fallback —
|
- **Flows and the error model** — rendezvous, `bridge_ask`, detached delivery, turn-done fallback —
|
||||||
are diagrammed in `docs/MCP-Contract.md` §6, kept out of this file because it loads into every
|
are diagrammed in `docs/MCP-Contract.md` **§6 only**. The rest of that page is a pre-build design
|
||||||
session's context.
|
doc whose tool names, parameter names and REST paths never caught up with the code, so do not use
|
||||||
|
it as the tool reference (CB-609). Section 6 is kept out of this file because this file loads into
|
||||||
|
every session's context.
|
||||||
|
|
||||||
### Redeploying the daemon — the lead may do this (primary only)
|
### Redeploying the daemon — the lead may do this (primary only)
|
||||||
|
|
||||||
|
|||||||
@@ -18,8 +18,9 @@ one unified Claude setup and the **sole communication gateway** (REST/SSE stays
|
|||||||
clients; any broker is `bridged`-internal, below the gateway).
|
clients; any broker is `bridged`-internal, below the gateway).
|
||||||
herdr owns the PTYs, multiplexing, persistence, and **agent-status events**; `bridged` owns
|
herdr owns the PTYs, multiplexing, persistence, and **agent-status events**; `bridged` owns
|
||||||
policy (subscription boundary, session lifecycle, status-gated delivery) and the client
|
policy (subscription boundary, session lifecycle, status-gated delivery) and the client
|
||||||
contract. The worker `claude` launches with `ANTHROPIC_BASE_URL=https://ollama.ltms.dev` + a
|
contract. A Claude member launches with `ANTHROPIC_BASE_URL` pointed at the gateway,
|
||||||
bearer token; the primary Opus stays env-clean and calls `bridged`'s MCP tools.
|
`https://llm.ltms.dev/anthropic`, plus a bearer token; the lead stays env-clean and calls
|
||||||
|
`bridged`'s MCP tools. See the wiki's **[13 User Guide](wiki/13-User-Guide.md)** to run it.
|
||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
flowchart LR
|
flowchart LR
|
||||||
@@ -31,7 +32,7 @@ flowchart LR
|
|||||||
end
|
end
|
||||||
HERDR["herdr<br/>panes · agent-status"]
|
HERDR["herdr<br/>panes · agent-status"]
|
||||||
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
|
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
|
||||||
M["ollama.ltms.dev<br/>(worker model)"]
|
M["llm.ltms.dev<br/>(the one gateway)"]
|
||||||
|
|
||||||
OPUS -->|"MCP bridge_send (blocks)"| SRV
|
OPUS -->|"MCP bridge_send (blocks)"| SRV
|
||||||
W -.->|"MCP bridge_reply"| SRV
|
W -.->|"MCP bridge_reply"| SRV
|
||||||
|
|||||||
@@ -88,15 +88,28 @@ bind:
|
|||||||
# backoffMs: 60000
|
# backoffMs: 60000
|
||||||
# quietNudgeCap: 3
|
# quietNudgeCap: 3
|
||||||
|
|
||||||
# Fleet health detection is dormant unless enabled. It reads one whole-fleet agent list per tick.
|
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
|
||||||
# It can run without a webhook; bridge_list then reports healthCoverage: detection-only.
|
# per tick.
|
||||||
|
# intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes
|
||||||
|
# Math.max(15, intervalSeconds), so a lower value is silently raised, not
|
||||||
|
# rejected.
|
||||||
|
# workingSuspectAfterSeconds, paneProbeIntervalSeconds → accepted and parsed, but NOT YET READ by
|
||||||
|
# anything — the dormant monitor only consumes intervalSeconds today (CB-573
|
||||||
|
# shipped ahead of the evidence publishers these two knobs are for). Setting
|
||||||
|
# them changes nothing right now, and no minimum is enforced on either, because
|
||||||
|
# nothing reads them to enforce one. They exist so a later build can start
|
||||||
|
# honouring them without another config-shape change.
|
||||||
|
# notifications.mode → "webhook" flips what bridge_list REPORTS (healthCoverage: "full" instead
|
||||||
|
# of "detection-only") — it does NOT make bridged send any webhook call; no
|
||||||
|
# delivery mechanism is implemented yet. Any other value, or omitting the
|
||||||
|
# block, reports "detection-only".
|
||||||
# health:
|
# health:
|
||||||
# enabled: true
|
# enabled: true
|
||||||
# intervalSeconds: 30 # minimum 15
|
# intervalSeconds: 30
|
||||||
# workingSuspectAfterSeconds: 600 # minimum 300
|
# workingSuspectAfterSeconds: 600
|
||||||
# paneProbeIntervalSeconds: 60 # minimum 60
|
# paneProbeIntervalSeconds: 60
|
||||||
# notifications:
|
# notifications:
|
||||||
# mode: disabled # disabled (default) or webhook
|
# mode: disabled
|
||||||
|
|
||||||
# herdr Unix socket. Omit to use the client default
|
# herdr Unix socket. Omit to use the client default
|
||||||
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
|
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
|
||||||
@@ -195,6 +208,29 @@ profiles:
|
|||||||
# `bridge_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile
|
# `bridge_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile
|
||||||
# is named directly. Negative is refused at config load — there is no sane meaning for it.
|
# is named directly. Negative is refused at config load — there is no sane meaning for it.
|
||||||
maxLoad: 2
|
maxLoad: 2
|
||||||
|
# subscription: true
|
||||||
|
# THE KNOB THAT DECIDES WHO PAYS (CB-539). Default false. When true, this profile's members
|
||||||
|
# run on the OPERATOR'S OWN Claude subscription instead of a metered endpoint — every spawn
|
||||||
|
# bills your plan and eats your usage limit. Off-subscription is the whole point of this
|
||||||
|
# daemon, so treat `true` as a deliberate exception, not a convenience.
|
||||||
|
#
|
||||||
|
# What changes when it is set (ClaudeCodeLauncher):
|
||||||
|
# - no ANTHROPIC_BASE_URL and no ANTHROPIC_AUTH_TOKEN are injected — the member inherits
|
||||||
|
# the operator's own Claude Code auth, which is exactly why it bills the plan;
|
||||||
|
# - SubscriptionGuard never vets it, because there is no baseUrl to vet;
|
||||||
|
# - no token is required, so `tokenEnv` is irrelevant here.
|
||||||
|
#
|
||||||
|
# MUTUALLY EXCLUSIVE with `baseUrl` — setting both is refused at config load (CB-542). On the
|
||||||
|
# subscription path no guard would vet the URL, so allowing both would be a way around the
|
||||||
|
# guard rather than a configuration.
|
||||||
|
#
|
||||||
|
# GOTCHA 1 — it is invisible to the startup secret check. `Bridged.reportRequiredSecrets`
|
||||||
|
# skips subscription profiles on purpose (they need no token), so a boot log that reports
|
||||||
|
# every secret as fine says nothing about these profiles.
|
||||||
|
#
|
||||||
|
# GOTCHA 2 — `maxLoad` is the ONLY throttle you have here. There is no metering, no budget
|
||||||
|
# and no refusal on cost; the cap on live members is the single thing standing between a
|
||||||
|
# fan-out and your monthly limit. Set it deliberately and keep it small.
|
||||||
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
|
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
|
||||||
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
|
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
|
||||||
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
|
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
|
||||||
@@ -262,6 +298,26 @@ profiles:
|
|||||||
# argv: ["opencode"]
|
# argv: ["opencode"]
|
||||||
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
|
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
|
||||||
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
|
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
|
||||||
|
#
|
||||||
|
# `weighted` IS NOT "cheapest first" — read this before you set weights (CB-589).
|
||||||
|
# It is smooth weighted round-robin: it spreads spawns across EVERY profile that has a free slot,
|
||||||
|
# in weight ratio. It has no idea which profile costs money. So with local:10 / paid:2 you do not
|
||||||
|
# get "use local, overflow to paid" — you get roughly one spawn in six going to the paid profile
|
||||||
|
# while the local box still has a free slot.
|
||||||
|
#
|
||||||
|
# There is a sharper second effect. The policy's running score map lives for the daemon's whole
|
||||||
|
# life. While a profile is at maxLoad it is filtered out and its score FREEZES, so the paid
|
||||||
|
# profiles keep accumulating against it. When the local slot frees up it returns with a stale
|
||||||
|
# score and can LOSE the next pick — a paid spawn while the free box sits idle.
|
||||||
|
#
|
||||||
|
# Until a real cost-first policy exists, the workaround is to make the ratio decisive rather than
|
||||||
|
# proportional: give the free profile a weight so large that it wins every pick it is eligible
|
||||||
|
# for, and paid profiles only ever take genuine overflow. On this host that is local weight 100
|
||||||
|
# against paid weights of ~1.
|
||||||
|
#
|
||||||
|
# The gotcha with that workaround: it expresses a PREFERENCE ORDER through a RATIO knob. Add a
|
||||||
|
# future profile at weight 150 and it silently outranks the free box, with nothing to warn you.
|
||||||
|
# Re-check the weights whenever you add a profile.
|
||||||
placement: weighted
|
placement: weighted
|
||||||
|
|
||||||
# How long a credential sits out after a BACKEND_EXHAUSTED classification (CB-578 stage B), in
|
# How long a credential sits out after a BACKEND_EXHAUSTED classification (CB-578 stage B), in
|
||||||
@@ -286,6 +342,11 @@ placement: weighted
|
|||||||
# / credentialId. Those are hot because the placement policy (and, for credentialId,
|
# / credentialId. Those are hot because the placement policy (and, for credentialId,
|
||||||
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
|
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
|
||||||
# not by itself enough to make a key hot.
|
# not by itself enough to make a key hot.
|
||||||
|
# EXCEPT `fleet.leaders`: Bridged.main reads it once at startup to build the lead tab
|
||||||
|
# scanner and launcher, and neither is rebuilt on reload. A changed/added/removed
|
||||||
|
# `fleet.leaders` entry is silently accepted — the reload reports "config reloaded"
|
||||||
|
# with nothing in the deferred list — but has NO effect until you restart. Treat it
|
||||||
|
# as deferred in practice, even though today's reload output does not say so.
|
||||||
# DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value
|
# DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value
|
||||||
# until you restart: `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`,
|
# until you restart: `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`,
|
||||||
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, `quarantineCooldownSeconds` (CB-578
|
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, `quarantineCooldownSeconds` (CB-578
|
||||||
@@ -368,6 +429,12 @@ fleet:
|
|||||||
# An auto-launched lead is NOT a member: it gets no worker reply charter, is never registered with
|
# An auto-launched lead is NOT a member: it gets no worker reply charter, is never registered with
|
||||||
# the session lifecycle (the idle reaper would kill your orchestrator), and stays on the
|
# the session lifecycle (the idle reaper would kill your orchestrator), and stays on the
|
||||||
# subscription — ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says.
|
# subscription — ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says.
|
||||||
|
#
|
||||||
|
# GET THE `tab:` VALUE RIGHT. A pane that does not match any configured `tab:` (a typo, a renamed
|
||||||
|
# tab, a pane no entry names at all) is not recognised as a lead — it resolves as an ordinary
|
||||||
|
# WORKER instead, silently, and every orchestration call it makes (spawn/stop/send/drain) is
|
||||||
|
# refused. There is no error at startup for this: an unmatched pane is simply not a lead. If your
|
||||||
|
# primary suddenly can't spawn or send, check this section first.
|
||||||
# leaders:
|
# leaders:
|
||||||
# opus-5.0:
|
# opus-5.0:
|
||||||
# profile: opus # omit to never create this lead, only recognise it
|
# profile: opus # omit to never create this lead, only recognise it
|
||||||
@@ -406,6 +473,91 @@ guard:
|
|||||||
- gx00.gw
|
- gx00.gw
|
||||||
- gx01.gw
|
- gx01.gw
|
||||||
|
|
||||||
|
# Member credential policy (CB-596, gitea issue #82). A herdr pane runs a LOGIN shell, and that
|
||||||
|
# shell re-sources the operator's own secret store — so a spawned member inherits every credential
|
||||||
|
# the operator's shell holds, not just the ones bridged means to give it. Measured on this host:
|
||||||
|
# 31 credential names, all set, with only ONE (GITEA_ACCESS_TOKEN) blocked before this — and that
|
||||||
|
# block was a single name hardcoded in HerdrPeerLauncher.java, not driven by this file. This block
|
||||||
|
# replaces that hardcoded shadow with a config-driven list of names.
|
||||||
|
#
|
||||||
|
# ROUND-2 CORRECTION, measured live: the pane-creation env overlay below (applied at tab.create /
|
||||||
|
# pane.split, BEFORE the pane's login shell runs) does NOT survive that login shell for any name
|
||||||
|
# secrets.sh actually exports — the shell re-exports it afterwards and overwrites the sentinel.
|
||||||
|
# Proof: GITEA_ACCESS_TOKEN comes back blocked only because secrets.sh itself carries a guarded
|
||||||
|
# export (`[ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...`) — that guard, not this
|
||||||
|
# file, is what wins. No other name in `known` below has a matching guard in secrets.sh yet (1
|
||||||
|
# guard measured against 33 export lines there). So today this block's overlay is REAL protection
|
||||||
|
# only for a name secrets.sh does not export, or a peer kind whose pane never runs a login shell —
|
||||||
|
# for everything secrets.sh exports and guards, the guard in secrets.sh (out of scope for this
|
||||||
|
# ticket) is what actually blocks it, not this list. An exec-time fix (winning after the login
|
||||||
|
# shell finishes, before the agent process starts) was attempted and found to have no seam in the
|
||||||
|
# current herdr protocol — AgentControl.start takes a fixed `kind` (herdr resolves the executable)
|
||||||
|
# plus trailing CLI args for that binary, not an arbitrary argv or an env map; only tab.create /
|
||||||
|
# pane.split accept `env`, and that is this same pane-creation overlay. See gitea #82 for the open
|
||||||
|
# design question this leaves.
|
||||||
|
#
|
||||||
|
# DENY-BY-DEFAULT, NOT A DENY-LIST. A deny-list (name the bad ones, let everything else through) is
|
||||||
|
# silently wrong the moment the operator's store gains a new secret — nothing would ever report it.
|
||||||
|
# Deny-by-default inverts that: `known` bounds the blast radius to names actually enumerated below,
|
||||||
|
# and EVERY one of them is blocked UNLESS it is also in `allow`. Omitting this block entirely (the
|
||||||
|
# shipped default) blocks NOTHING — unlike most optional blocks in this file, absence here is a real
|
||||||
|
# gap, not a safe "feature off". A name that is neither `known` nor `allow`-ed is not silently let
|
||||||
|
# through either: the daemon logs a WARN naming any credential-shaped env var it finds on neither
|
||||||
|
# list (never its value), so a secret added to the store later does not go unnoticed forever.
|
||||||
|
#
|
||||||
|
# policy → only "deny-by-default" exists today (an operator-authored deny-list was deliberately
|
||||||
|
# rejected — see above). An unrecognized value refuses to start, naming it.
|
||||||
|
# allow → credential names a member legitimately needs. Left OUT of the pane's env overlay
|
||||||
|
# entirely, so the value the pane's own (login) shell exports passes through untouched.
|
||||||
|
# known → every credential name the operator's store is known to export. Every name here NOT
|
||||||
|
# also in `allow` is overlaid with a non-secret sentinel value before the pane's login
|
||||||
|
# shell runs — real protection only for names the login shell does not itself re-export
|
||||||
|
# (see the ROUND-2 CORRECTION note above for the ones it does).
|
||||||
|
#
|
||||||
|
# HOT-RELOADABLE the same way `fleet:` is (CB-559): read fresh on every spawn, so editing this list
|
||||||
|
# and reloading config (or restarting) changes what the NEXT spawn inherits; already-running members
|
||||||
|
# are unaffected either way.
|
||||||
|
# memberCredentials:
|
||||||
|
# policy: deny-by-default
|
||||||
|
# allow:
|
||||||
|
# - AI_GATEWAY_TOKEN # named in a profile's tokenEnv (local/gx) — a member reaching the
|
||||||
|
# # gateway is by design, not a leak
|
||||||
|
# - WORKER_GITEA_TOKEN # the repo-scoped forge token a member needs to open its own PR (CB-302)
|
||||||
|
# - CONTEXT7_TOKEN # already decided as allowed by CB-593
|
||||||
|
# - GITEA_HOST # not a credential — a hostname, paired with the forge token above
|
||||||
|
# known:
|
||||||
|
# - AI_GATEWAY_TOKEN
|
||||||
|
# - BESZEL_ADMIN_EMAIL
|
||||||
|
# - BESZEL_ADMIN_PASSWORD
|
||||||
|
# - BESZEL_HUB_URL
|
||||||
|
# - BESZEL_KEY
|
||||||
|
# - BESZEL_UNIVERSAL_TOKEN
|
||||||
|
# - BRAIN_MCP_TOKEN
|
||||||
|
# - CF_ACCOUNT_ID
|
||||||
|
# - CF_API_TOKEN
|
||||||
|
# - CF_USER_TOKEN
|
||||||
|
# - CONFLUENCE_API_TOKEN
|
||||||
|
# - CONFLUENCE_USERNAME
|
||||||
|
# - CONTEXT7_TOKEN
|
||||||
|
# - GITEA_ACCESS_TOKEN
|
||||||
|
# - GITEA_HOST
|
||||||
|
# - GITLAB_OAUTH_CLIENT_SECRET
|
||||||
|
# - GITLAB_PERSONAL_ACCESS_TOKEN
|
||||||
|
# - GRAFANA_ADMIN_PASSWORD
|
||||||
|
# - GRAFANA_ADMIN_USER
|
||||||
|
# - HASS_TOKEN
|
||||||
|
# - HW_PASSWORD
|
||||||
|
# - HW_USER
|
||||||
|
# - LTMS_API_KEY
|
||||||
|
# - MEMORY_MCP_TOKEN
|
||||||
|
# - METRICS_PUSH_TOKEN
|
||||||
|
# - OPENCODE_AUTOMODE_MODEL
|
||||||
|
# - TELEGRAM_BOT_TOKEN
|
||||||
|
# - TELEGRAM_CHAT_ID
|
||||||
|
# - TS_API_KEY
|
||||||
|
# - TS_AUTHKEY
|
||||||
|
# - WORKER_GITEA_TOKEN
|
||||||
|
|
||||||
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
|
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
|
||||||
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
|
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
|
||||||
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
|
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
|
||||||
@@ -423,10 +575,15 @@ guard:
|
|||||||
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
|
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
|
||||||
# contextCap → force-release a session after this many delegated turns
|
# contextCap → force-release a session after this many delegated turns
|
||||||
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
|
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
|
||||||
|
# clearAfterTurn → whether a reusable worker discards its conversation context after every
|
||||||
|
# completed delegated turn (default false). Works for claude-code workers
|
||||||
|
# only — any other peer kind (e.g. opencode) logs "context reset is
|
||||||
|
# unsupported for peer kind …" once and the reset is a no-op.
|
||||||
# lifecycle:
|
# lifecycle:
|
||||||
# idleTtlSeconds: 300
|
# idleTtlSeconds: 300
|
||||||
# contextCap: 10
|
# contextCap: 10
|
||||||
# drainTimeoutSeconds: 5
|
# drainTimeoutSeconds: 5
|
||||||
|
# clearAfterTurn: false
|
||||||
|
|
||||||
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
|
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
|
||||||
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
|
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
|
||||||
@@ -434,10 +591,14 @@ guard:
|
|||||||
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
|
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
|
||||||
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
|
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
|
||||||
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
|
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
|
||||||
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
|
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
|
||||||
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
|
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
|
||||||
|
# prefetch → CB-527: consumer basicQos, capping how many unacked messages the inbox holds
|
||||||
|
# in-heap per owned target (the rest sits on the broker's durable queue instead of
|
||||||
|
# growing the JVM heap). Default 32 when omitted.
|
||||||
# broker:
|
# broker:
|
||||||
# uri: amqp://guest:guest@127.0.0.1:5672
|
# uri: amqp://guest:guest@127.0.0.1:5672
|
||||||
|
# prefetch: 32
|
||||||
|
|
||||||
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
|
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
|
||||||
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
|
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
|
||||||
|
|||||||
@@ -87,6 +87,11 @@ public final class Bridged {
|
|||||||
// else can fail on a silently-empty one. A daemon started without a login shell (launchd)
|
// else can fail on a silently-empty one. A daemon started without a login shell (launchd)
|
||||||
// boots fine either way — this is the only thing that says so out loud.
|
// boots fine either way — this is the only thing that says so out loud.
|
||||||
reportRequiredSecrets(cfg);
|
reportRequiredSecrets(cfg);
|
||||||
|
// CB-596: an absent (or empty) memberCredentials: block blocks NOTHING — no credential
|
||||||
|
// name is hardcoded any more to fall back on. Say so loudly, the same way a missing
|
||||||
|
// secret is reported above, so upgrading past this commit never silently drops CB-592's
|
||||||
|
// protection.
|
||||||
|
reportMemberCredentialsGap(cfg);
|
||||||
// CB-559: `cfg` stays the startup snapshot — every validation and every piece of one-time
|
// CB-559: `cfg` stays the startup snapshot — every validation and every piece of one-time
|
||||||
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
|
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
|
||||||
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
|
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
|
||||||
@@ -139,13 +144,15 @@ public final class Bridged {
|
|||||||
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
|
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
|
||||||
claudeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
|
claudeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
|
||||||
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
|
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
|
||||||
() -> config.get().fleet()));
|
() -> config.get().fleet(),
|
||||||
|
() -> config.get().memberCredentials()));
|
||||||
}
|
}
|
||||||
if (!opencodeProfiles.isEmpty()) {
|
if (!opencodeProfiles.isEmpty()) {
|
||||||
adapters.add(new OpenCodeLauncher(agents, spaces,
|
adapters.add(new OpenCodeLauncher(agents, spaces,
|
||||||
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
|
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
|
||||||
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
|
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
|
||||||
() -> config.get().fleet()));
|
() -> config.get().fleet(),
|
||||||
|
() -> config.get().memberCredentials()));
|
||||||
}
|
}
|
||||||
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
|
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
|
||||||
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
|
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
|
||||||
@@ -352,8 +359,9 @@ public final class Bridged {
|
|||||||
// connection, so keep the reference to close it in the ordered shutdown hook.
|
// connection, so keep the reference to close it in the ordered shutdown hook.
|
||||||
final ReplyInbox replyInbox;
|
final ReplyInbox replyInbox;
|
||||||
if (cfg.broker() != null && cfg.broker().isConfigured()) {
|
if (cfg.broker() != null && cfg.broker().isConfigured()) {
|
||||||
replyInbox = AmqpReplyInbox.open(cfg.broker().uri());
|
replyInbox = AmqpReplyInbox.open(cfg.broker().uri(), cfg.broker().prefetchOrDefault());
|
||||||
log.info("reply inbox: AMQP broker (durable) at {}", cfg.broker().uri());
|
log.info("reply inbox: AMQP broker (durable) at {} (prefetch={})",
|
||||||
|
cfg.broker().uri(), cfg.broker().prefetchOrDefault());
|
||||||
} else {
|
} else {
|
||||||
replyInbox = new InMemoryReplyInbox();
|
replyInbox = new InMemoryReplyInbox();
|
||||||
log.info("reply inbox: in-memory (soft-state)");
|
log.info("reply inbox: in-memory (soft-state)");
|
||||||
@@ -440,6 +448,11 @@ public final class Bridged {
|
|||||||
reason += "; worktree=" + detail.worktreePath() + " branch=" + detail.branch()
|
reason += "; worktree=" + detail.worktreePath() + " branch=" + detail.branch()
|
||||||
+ " snapshot=" + (detail.snapshotRef() != null ? detail.snapshotRef() : "none");
|
+ " snapshot=" + (detail.snapshotRef() != null ? detail.snapshotRef() : "none");
|
||||||
}
|
}
|
||||||
|
// CB-584 (issue #65 criterion 5): also name the agent session, so a lead can resume the
|
||||||
|
// member's conversation instead of only re-dispatching a fresh one onto the same files.
|
||||||
|
if (detail.agentSessionId() != null) {
|
||||||
|
reason += " agentSessionId=" + detail.agentSessionId();
|
||||||
|
}
|
||||||
messages.abandon(detail.terminalId(), reason);
|
messages.abandon(detail.terminalId(), reason);
|
||||||
replyInbox.release(detail.terminalId());
|
replyInbox.release(detail.terminalId());
|
||||||
primaryRegistry.forgetDelegation(detail.terminalId()); // CB-532: don't leak the lead binding
|
primaryRegistry.forgetDelegation(detail.terminalId()); // CB-532: don't leak the lead binding
|
||||||
@@ -554,7 +567,10 @@ public final class Bridged {
|
|||||||
* where set (opt-in). Derived from the config, not hard-coded, so a new profile is covered for
|
* where set (opt-in). Derived from the config, not hard-coded, so a new profile is covered for
|
||||||
* free. A var required by more than one profile is one entry naming every profile that needs
|
* free. A var required by more than one profile is one entry naming every profile that needs
|
||||||
* it. Deliberately excludes {@code auth.tokenEnv}: that one is already enforced loudly, by a
|
* it. Deliberately excludes {@code auth.tokenEnv}: that one is already enforced loudly, by a
|
||||||
* startup throw, a few lines above this method's call site.
|
* startup throw in {@code main()} — about 370 lines <em>below</em> this method's call site
|
||||||
|
* ({@link #reportRequiredSecrets(BridgedConfig)}), not a few lines above it. That throw only
|
||||||
|
* fires when {@code auth.mode: token} is configured; under the default loopback-trust mode it
|
||||||
|
* never runs, and {@code auth.tokenEnv} is simply not required.
|
||||||
*
|
*
|
||||||
* <p>Package-private and pure (no I/O, no logging) so the derivation is unit-testable without
|
* <p>Package-private and pure (no I/O, no logging) so the derivation is unit-testable without
|
||||||
* capturing log output; {@link #reportRequiredSecrets(BridgedConfig)} is the logging caller.
|
* capturing log output; {@link #reportRequiredSecrets(BridgedConfig)} is the logging caller.
|
||||||
@@ -604,6 +620,31 @@ public final class Bridged {
|
|||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: {@code known:} empty (block absent entirely, or present but empty) means {@link
|
||||||
|
* BridgedConfig.MemberCredentials#blockedSet()} is empty too — every member pane inherits the
|
||||||
|
* operator's whole secret store, unblocked, exactly the defect this ticket fixes. Unlike a
|
||||||
|
* missing token ({@link #reportRequiredSecrets}), there is no name to point at: the point is
|
||||||
|
* that the block itself is missing. Warn once at startup and say what to add; never refuse to
|
||||||
|
* start over it — see {@link #reportRequiredSecrets} for why a daemon that boots and says
|
||||||
|
* what is wrong beats one that will not boot at all.
|
||||||
|
*
|
||||||
|
* <p>Package-private so the test can capture the log directly, the same way {@link
|
||||||
|
* #requiredSecretEnvVars} is exposed for {@link #reportRequiredSecrets}'s own test.
|
||||||
|
*/
|
||||||
|
static void reportMemberCredentialsGap(BridgedConfig cfg) {
|
||||||
|
BridgedConfig.MemberCredentials creds = cfg.memberCredentials();
|
||||||
|
if (creds != null && !creds.known().isEmpty()) {
|
||||||
|
log.info("memberCredentials: {} known name(s), {} allowed — blocking {} on every spawn",
|
||||||
|
creds.known().size(), creds.allow().size(), creds.blockedSet().size());
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
log.warn("memberCredentials: absent or empty — the daemon will start anyway, and every "
|
||||||
|
+ "member pane inherits the operator's WHOLE secret store, unblocked (CB-592's "
|
||||||
|
+ "protection is lost). Add a memberCredentials: block (policy/allow/known) to "
|
||||||
|
+ "bridged.yaml — see bridged.example.yaml — and restart.");
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
|
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
|
||||||
*
|
*
|
||||||
|
|||||||
@@ -5,7 +5,9 @@ import com.fasterxml.jackson.core.JsonParser;
|
|||||||
import com.fasterxml.jackson.core.JsonToken;
|
import com.fasterxml.jackson.core.JsonToken;
|
||||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
import com.fasterxml.jackson.dataformat.yaml.YAMLFactory;
|
import com.fasterxml.jackson.dataformat.yaml.YAMLFactory;
|
||||||
|
import dev.ltms.bridged.msg.AmqpReplyInbox;
|
||||||
import dev.ltms.bridged.peer.MemberRole;
|
import dev.ltms.bridged.peer.MemberRole;
|
||||||
|
import dev.ltms.bridged.placement.PlacementPolicies;
|
||||||
import org.slf4j.Logger;
|
import org.slf4j.Logger;
|
||||||
import org.slf4j.LoggerFactory;
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
@@ -64,6 +66,9 @@ import java.util.Set;
|
|||||||
* {@code BackendQuarantine} built at startup, so it is DEFERRED: changing it
|
* {@code BackendQuarantine} built at startup, so it is DEFERRED: changing it
|
||||||
* needs a restart, and a quarantine already running keeps whatever cooldown was
|
* needs a restart, and a quarantine already running keeps whatever cooldown was
|
||||||
* live when it started.
|
* live when it started.
|
||||||
|
* @param memberCredentials deny-by-default policy (CB-596) for which of the operator's own host
|
||||||
|
* credentials a spawned member's pane inherits. {@code null} (the block
|
||||||
|
* omitted) blocks nothing — see {@link MemberCredentials}.
|
||||||
*/
|
*/
|
||||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||||
public record BridgedConfig(
|
public record BridgedConfig(
|
||||||
@@ -83,7 +88,19 @@ public record BridgedConfig(
|
|||||||
String placement,
|
String placement,
|
||||||
Auth auth,
|
Auth auth,
|
||||||
ConfigReload configReload,
|
ConfigReload configReload,
|
||||||
Integer quarantineCooldownSeconds) {
|
Integer quarantineCooldownSeconds,
|
||||||
|
MemberCredentials memberCredentials) {
|
||||||
|
|
||||||
|
/** Back-compat form before the CB-596 {@code memberCredentials:} block was added. */
|
||||||
|
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
|
||||||
|
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
|
||||||
|
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
|
||||||
|
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
|
||||||
|
ConfigReload configReload, Integer quarantineCooldownSeconds) {
|
||||||
|
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
|
||||||
|
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
|
||||||
|
configReload, quarantineCooldownSeconds, null);
|
||||||
|
}
|
||||||
|
|
||||||
/** Default cooldown (CB-578 stage B) when {@code quarantineCooldownSeconds} is absent/non-positive. */
|
/** Default cooldown (CB-578 stage B) when {@code quarantineCooldownSeconds} is absent/non-positive. */
|
||||||
public static final int DEFAULT_QUARANTINE_COOLDOWN_SECONDS = 1800;
|
public static final int DEFAULT_QUARANTINE_COOLDOWN_SECONDS = 1800;
|
||||||
@@ -532,16 +549,24 @@ public record BridgedConfig(
|
|||||||
* stays soft-state. Production default is LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1
|
* stays soft-state. Production default is LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1
|
||||||
* and is a URI-only swap.
|
* and is a URI-only swap.
|
||||||
*
|
*
|
||||||
* @param uri AMQP connection URI, e.g. {@code amqp://guest:guest@127.0.0.1:5672/}. Blank/{@code null}
|
* @param uri AMQP connection URI, e.g. {@code amqp://guest:guest@127.0.0.1:5672/}. Blank/
|
||||||
* ⇒ the broker block is treated as absent (in-memory adapter).
|
* {@code null} ⇒ the broker block is treated as absent (in-memory adapter).
|
||||||
|
* @param prefetch CB-527: the consumer's {@code basicQos} prefetch count, bounding how many
|
||||||
|
* unacked messages the AMQP inbox holds in-heap per owned target. {@code null}/
|
||||||
|
* non-positive ⇒ {@link AmqpReplyInbox#DEFAULT_PREFETCH}.
|
||||||
*/
|
*/
|
||||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||||
public record Broker(String uri) {
|
public record Broker(String uri, Integer prefetch) {
|
||||||
|
|
||||||
/** True when a usable broker URI is configured (an empty block does not enable AMQP). */
|
/** True when a usable broker URI is configured (an empty block does not enable AMQP). */
|
||||||
public boolean isConfigured() {
|
public boolean isConfigured() {
|
||||||
return uri != null && !uri.isBlank();
|
return uri != null && !uri.isBlank();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** The prefetch to use, defaulting to {@link AmqpReplyInbox#DEFAULT_PREFETCH} when unset. */
|
||||||
|
public int prefetchOrDefault() {
|
||||||
|
return (prefetch != null && prefetch > 0) ? prefetch : AmqpReplyInbox.DEFAULT_PREFETCH;
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -916,6 +941,58 @@ public record BridgedConfig(
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: which of the operator's own host credentials a spawned member's pane may inherit.
|
||||||
|
*
|
||||||
|
* <p>A herdr pane runs a login shell that re-sources the operator's own secret store, so a
|
||||||
|
* member inherits every credential the operator's shell holds — measured at 31 names on this
|
||||||
|
* host, of which only one ({@code GITEA_ACCESS_TOKEN}) used to be blocked, and that block was a
|
||||||
|
* single name hardcoded in {@link HerdrPeerLauncher} rather than driven by config (gitea issue
|
||||||
|
* #82). This record replaces that hardcoded shadow with a config-driven one.
|
||||||
|
*
|
||||||
|
* <p><b>deny-by-default, not a deny-list.</b> A deny-list (block these specific names, let
|
||||||
|
* everything else through) is silently wrong the moment a new secret is added to the operator's
|
||||||
|
* store — nothing would ever report it. Deny-by-default inverts that: {@link #known} bounds the
|
||||||
|
* blast radius to names the operator has actually enumerated, and every one of them is blocked
|
||||||
|
* UNLESS it is also in {@link #allow}. A name that shows up in neither list is not silently
|
||||||
|
* allowed — see {@code HerdrPeerLauncher}'s gap detector, which logs it.
|
||||||
|
*
|
||||||
|
* @param policy how the block is computed. Only {@link #POLICY_DENY_BY_DEFAULT} is understood
|
||||||
|
* today; {@code null}/blank defaults to it. An operator's own deny-list is
|
||||||
|
* deliberately not supported — see above.
|
||||||
|
* @param allow credential names a member legitimately needs (e.g. the gateway token it reaches
|
||||||
|
* the LLM through, the repo-scoped forge token it opens its own PR with). Every
|
||||||
|
* name here is left unmentioned in the pane's env overlay, so the value the pane's
|
||||||
|
* own (login) shell exports passes through untouched.
|
||||||
|
* @param known every credential name the operator's store is known to export. Every name here
|
||||||
|
* that is NOT also in {@link #allow} is overlaid with a non-secret sentinel value,
|
||||||
|
* shadowing whatever the pane's login shell would otherwise export for it.
|
||||||
|
*/
|
||||||
|
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||||
|
public record MemberCredentials(String policy, List<String> allow, List<String> known) {
|
||||||
|
|
||||||
|
/** The only policy this build understands: block every {@code known} name not in {@code allow}. */
|
||||||
|
public static final String POLICY_DENY_BY_DEFAULT = "deny-by-default";
|
||||||
|
|
||||||
|
public MemberCredentials {
|
||||||
|
policy = (policy == null || policy.isBlank()) ? POLICY_DENY_BY_DEFAULT : policy.toLowerCase();
|
||||||
|
allow = allow == null ? List.of() : List.copyOf(allow);
|
||||||
|
known = known == null ? List.of() : List.copyOf(known);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** {@link #allow} as a set, for membership checks. */
|
||||||
|
public Set<String> allowSet() {
|
||||||
|
return Set.copyOf(allow);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** {@link #known} minus {@link #allow} — the names a spawn must shadow. */
|
||||||
|
public Set<String> blockedSet() {
|
||||||
|
Set<String> blocked = new java.util.LinkedHashSet<>(known);
|
||||||
|
blocked.removeAll(allowSet());
|
||||||
|
return blocked;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* The candidate profiles an unqualified spawn of {@code role} chooses between, in definition
|
* The candidate profiles an unqualified spawn of {@code role} chooses between, in definition
|
||||||
* order (CB-557).
|
* order (CB-557).
|
||||||
@@ -958,11 +1035,16 @@ public record BridgedConfig(
|
|||||||
/**
|
/**
|
||||||
* Top-level keys this version understands. Used only to warn about the rest — see
|
* Top-level keys this version understands. Used only to warn about the rest — see
|
||||||
* {@link #warnUnknownTopLevelKeys}. Keep in step with the record components.
|
* {@link #warnUnknownTopLevelKeys}. Keep in step with the record components.
|
||||||
|
*
|
||||||
|
* <p>Package-private (not {@code private}) so a test can assert every key here is documented in
|
||||||
|
* {@code bridged.example.yaml} — the only committed description of the config schema, since
|
||||||
|
* {@code bridged.yaml} itself is gitignored.
|
||||||
*/
|
*/
|
||||||
private static final Set<String> KNOWN_TOP_LEVEL_KEYS = Set.of(
|
static final Set<String> KNOWN_TOP_LEVEL_KEYS = Set.of(
|
||||||
"bind", "herdrSocket", "profiles", "guard", "worktreeRoot",
|
"bind", "herdrSocket", "profiles", "guard", "worktreeRoot",
|
||||||
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
|
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
|
||||||
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds");
|
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds",
|
||||||
|
"memberCredentials");
|
||||||
|
|
||||||
/** Load and validate config from {@code path}. */
|
/** Load and validate config from {@code path}. */
|
||||||
public static BridgedConfig load(Path path) {
|
public static BridgedConfig load(Path path) {
|
||||||
@@ -973,7 +1055,15 @@ public record BridgedConfig(
|
|||||||
warnUnknownTopLevelKeys(yaml, path);
|
warnUnknownTopLevelKeys(yaml, path);
|
||||||
rejectDuplicateMemberSlots(yaml);
|
rejectDuplicateMemberSlots(yaml);
|
||||||
rejectNegativeMaxLoad(yaml);
|
rejectNegativeMaxLoad(yaml);
|
||||||
|
rejectUnknownKind(yaml);
|
||||||
|
rejectUnknownAuthMode(yaml);
|
||||||
|
rejectUnknownPlacement(yaml);
|
||||||
|
rejectUnknownMemberCredentialsPolicy(yaml);
|
||||||
BridgedConfig cfg = YAML.readValue(yaml, BridgedConfig.class);
|
BridgedConfig cfg = YAML.readValue(yaml, BridgedConfig.class);
|
||||||
|
// CB-606: validated here, eagerly, using PlacementPolicies.fromName as the single source
|
||||||
|
// of truth — not lazily at first spawn (see CompositePeerLauncher's placementPolicy
|
||||||
|
// Supplier), where a bad name would still start a daemon that looks healthy.
|
||||||
|
rejectUnknownPlacementPolicy(cfg.placement());
|
||||||
return cfg.withDefaults();
|
return cfg.withDefaults();
|
||||||
} catch (IOException e) {
|
} catch (IOException e) {
|
||||||
throw new UncheckedIOException("cannot read bridged config at " + path, e);
|
throw new UncheckedIOException("cannot read bridged config at " + path, e);
|
||||||
@@ -1267,6 +1357,201 @@ public record BridgedConfig(
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** The peer kinds this build has an adapter for — {@link Profile#kind()}'s only valid values. */
|
||||||
|
private static final Set<String> KNOWN_KINDS = Set.of(Profile.KIND_CLAUDE_CODE, Profile.KIND_OPENCODE);
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reject a profile whose {@code kind:} is not one of {@link #KNOWN_KINDS} (CB-604), naming the
|
||||||
|
* profile, the value it set, and the accepted set.
|
||||||
|
*
|
||||||
|
* <p>{@link Profile}'s compact constructor only lower-cases {@code kind} and compares it against
|
||||||
|
* {@code KIND_OPENCODE} — anything else, including a typo like {@code opencod}, silently falls
|
||||||
|
* into the claude-code bucket ({@link dev.ltms.bridged.member.CompositePeerLauncher} routes by
|
||||||
|
* exact adapter claim, not by membership in a known set). With {@code argv:} also unset, the argv
|
||||||
|
* default special-cases only the exact string {@code "claude-code"}, so the launch command falls
|
||||||
|
* back to {@code List.of(kind)} — the daemon then tries to run a program literally named after the
|
||||||
|
* typo. {@code CompositePeerLauncher}'s constructor already treats a profile claimed by two
|
||||||
|
* adapters as fatal (CB-402); an unrecognized kind is the same class of adapter-routing mistake
|
||||||
|
* and gets the same treatment here, at config load, rather than surfacing later as a failed spawn.
|
||||||
|
*
|
||||||
|
* @param yaml the raw config text
|
||||||
|
* @throws IllegalStateException when any profile's {@code kind} is a non-blank value not in
|
||||||
|
* {@link #KNOWN_KINDS} (case-insensitive)
|
||||||
|
*/
|
||||||
|
static void rejectUnknownKind(String yaml) {
|
||||||
|
Map<?, ?> raw;
|
||||||
|
try {
|
||||||
|
raw = YAML.readValue(yaml, Map.class);
|
||||||
|
} catch (IOException | IllegalArgumentException e) {
|
||||||
|
return; // a malformed file is reported by the real parse, not here
|
||||||
|
}
|
||||||
|
if (raw == null || !(raw.get("profiles") instanceof Map<?, ?> profiles)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
List<String> bad = profiles.entrySet().stream()
|
||||||
|
.filter(e -> e.getValue() instanceof Map<?, ?> p
|
||||||
|
&& p.get("kind") instanceof String k && !k.isBlank()
|
||||||
|
&& !KNOWN_KINDS.contains(k.toLowerCase()))
|
||||||
|
.map(e -> String.valueOf(e.getKey()) + "=" + ((Map<?, ?>) e.getValue()).get("kind"))
|
||||||
|
.sorted()
|
||||||
|
.toList();
|
||||||
|
if (!bad.isEmpty()) {
|
||||||
|
throw new IllegalStateException("refusing to start: profile(s) [" + String.join(", ", bad)
|
||||||
|
+ "] set an unrecognized kind — accepted values are "
|
||||||
|
+ String.join(", ", KNOWN_KINDS.stream().sorted().toList())
|
||||||
|
+ " (case-insensitive); an unrecognized kind would otherwise fall back to the"
|
||||||
|
+ " claude-code adapter and try to launch a program named after the typo.");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** The auth modes this build understands — {@link Auth#mode()}'s only valid values. */
|
||||||
|
private static final Set<String> KNOWN_AUTH_MODES = Set.of(Auth.MODE_LOOPBACK_TRUST, Auth.MODE_TOKEN);
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reject an {@code auth.mode} that is not one of {@link #KNOWN_AUTH_MODES} (CB-606), naming the
|
||||||
|
* value and the accepted set.
|
||||||
|
*
|
||||||
|
* <p>{@link Auth}'s compact constructor only lower-cases {@code mode}, and
|
||||||
|
* {@link Auth#tokenMode()} only compares the result against {@code MODE_TOKEN} — anything else,
|
||||||
|
* including a typo like {@code toekn}, silently behaves as {@code loopback-trust}. That fallback
|
||||||
|
* is otherwise checked only by {@link #validateAuthExposure()}, and only when the bind is
|
||||||
|
* non-loopback: on a loopback bind (the common case) the typo is invisible end to end — the
|
||||||
|
* daemon starts cleanly and authenticates nobody while the operator believes {@code token} mode
|
||||||
|
* is active. Refuse it here, unconditionally, at config load, rather than let it hide behind the
|
||||||
|
* bind check.
|
||||||
|
*
|
||||||
|
* @param yaml the raw config text
|
||||||
|
* @throws IllegalStateException when {@code auth.mode} is a non-blank value not in
|
||||||
|
* {@link #KNOWN_AUTH_MODES} (case-insensitive)
|
||||||
|
*/
|
||||||
|
static void rejectUnknownAuthMode(String yaml) {
|
||||||
|
Map<?, ?> raw;
|
||||||
|
try {
|
||||||
|
raw = YAML.readValue(yaml, Map.class);
|
||||||
|
} catch (IOException | IllegalArgumentException e) {
|
||||||
|
return; // a malformed file is reported by the real parse, not here
|
||||||
|
}
|
||||||
|
if (raw == null || !(raw.get("auth") instanceof Map<?, ?> auth)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (!(auth.get("mode") instanceof String mode) || mode.isBlank()
|
||||||
|
|| KNOWN_AUTH_MODES.contains(mode.toLowerCase())) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
throw new IllegalStateException("refusing to start: auth.mode=" + mode
|
||||||
|
+ " is not recognized — accepted values are "
|
||||||
|
+ String.join(", ", KNOWN_AUTH_MODES.stream().sorted().toList())
|
||||||
|
+ " (case-insensitive); an unrecognized mode would otherwise silently fall back to"
|
||||||
|
+ " loopback-trust, which authenticates nobody.");
|
||||||
|
}
|
||||||
|
|
||||||
|
/** The per-profile placements this build understands — {@link Profile#placement()}'s only valid values. */
|
||||||
|
private static final Set<String> KNOWN_PLACEMENTS = Set.of("tab", "pane");
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reject a profile whose {@code placement:} is not one of {@link #KNOWN_PLACEMENTS} (CB-606),
|
||||||
|
* naming the profile, the value it set, and the accepted set.
|
||||||
|
*
|
||||||
|
* <p>{@link Profile}'s compact constructor only lower-cases {@code placement}, and
|
||||||
|
* {@link Profile#tabPlacement()} only compares the result against {@code "tab"} — anything else,
|
||||||
|
* including a typo like {@code tabb}, silently falls back to the legacy pane placement with no
|
||||||
|
* signal anywhere.
|
||||||
|
*
|
||||||
|
* @param yaml the raw config text
|
||||||
|
* @throws IllegalStateException when any profile's {@code placement} is a non-blank value not in
|
||||||
|
* {@link #KNOWN_PLACEMENTS} (case-insensitive)
|
||||||
|
*/
|
||||||
|
static void rejectUnknownPlacement(String yaml) {
|
||||||
|
Map<?, ?> raw;
|
||||||
|
try {
|
||||||
|
raw = YAML.readValue(yaml, Map.class);
|
||||||
|
} catch (IOException | IllegalArgumentException e) {
|
||||||
|
return; // a malformed file is reported by the real parse, not here
|
||||||
|
}
|
||||||
|
if (raw == null || !(raw.get("profiles") instanceof Map<?, ?> profiles)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
List<String> bad = profiles.entrySet().stream()
|
||||||
|
.filter(e -> e.getValue() instanceof Map<?, ?> p
|
||||||
|
&& p.get("placement") instanceof String pl && !pl.isBlank()
|
||||||
|
&& !KNOWN_PLACEMENTS.contains(pl.toLowerCase()))
|
||||||
|
.map(e -> String.valueOf(e.getKey()) + "=" + ((Map<?, ?>) e.getValue()).get("placement"))
|
||||||
|
.sorted()
|
||||||
|
.toList();
|
||||||
|
if (!bad.isEmpty()) {
|
||||||
|
throw new IllegalStateException("refusing to start: profile(s) [" + String.join(", ", bad)
|
||||||
|
+ "] set an unrecognized placement — accepted values are "
|
||||||
|
+ String.join(", ", KNOWN_PLACEMENTS.stream().sorted().toList())
|
||||||
|
+ " (case-insensitive); an unrecognized placement would otherwise fall back to"
|
||||||
|
+ " legacy pane placement with no signal anywhere.");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** The member-credential policies this build understands — {@link MemberCredentials#policy()}'s only valid value. */
|
||||||
|
private static final Set<String> KNOWN_MEMBER_CREDENTIALS_POLICIES =
|
||||||
|
Set.of(MemberCredentials.POLICY_DENY_BY_DEFAULT);
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reject a {@code memberCredentials.policy} that is not {@link #KNOWN_MEMBER_CREDENTIALS_POLICIES}
|
||||||
|
* (CB-596), naming the value and the accepted set.
|
||||||
|
*
|
||||||
|
* <p>{@link MemberCredentials}'s compact constructor only lower-cases {@code policy} and defaults
|
||||||
|
* a blank one to {@link MemberCredentials#POLICY_DENY_BY_DEFAULT} — nothing rejects an actual
|
||||||
|
* typo like {@code deny-by-defualt}. There is only one policy today, so such a typo would
|
||||||
|
* currently behave identically to the real value by accident; the day a second policy exists
|
||||||
|
* that accident becomes a silent behavior change. Refuse it now, at config load, following the
|
||||||
|
* same pattern as {@link #rejectUnknownAuthMode} and {@link #rejectUnknownPlacement}.
|
||||||
|
*
|
||||||
|
* @param yaml the raw config text
|
||||||
|
* @throws IllegalStateException when {@code memberCredentials.policy} is a non-blank value not in
|
||||||
|
* {@link #KNOWN_MEMBER_CREDENTIALS_POLICIES} (case-insensitive)
|
||||||
|
*/
|
||||||
|
static void rejectUnknownMemberCredentialsPolicy(String yaml) {
|
||||||
|
Map<?, ?> raw;
|
||||||
|
try {
|
||||||
|
raw = YAML.readValue(yaml, Map.class);
|
||||||
|
} catch (IOException | IllegalArgumentException e) {
|
||||||
|
return; // a malformed file is reported by the real parse, not here
|
||||||
|
}
|
||||||
|
if (raw == null || !(raw.get("memberCredentials") instanceof Map<?, ?> mc)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (!(mc.get("policy") instanceof String policy) || policy.isBlank()
|
||||||
|
|| KNOWN_MEMBER_CREDENTIALS_POLICIES.contains(policy.toLowerCase())) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
throw new IllegalStateException("refusing to start: memberCredentials.policy=" + policy
|
||||||
|
+ " is not recognized — accepted values are "
|
||||||
|
+ String.join(", ", KNOWN_MEMBER_CREDENTIALS_POLICIES.stream().sorted().toList())
|
||||||
|
+ " (case-insensitive).");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reject a top-level {@code placement:} policy name {@link PlacementPolicies#fromName} does not
|
||||||
|
* recognize (CB-606), at config load rather than lazily at first spawn.
|
||||||
|
*
|
||||||
|
* <p>{@code CompositePeerLauncher} only calls {@link PlacementPolicies#fromName} per spawn,
|
||||||
|
* through a {@code Supplier} that re-reads live config (CB-559, so a hot-reloaded placement
|
||||||
|
* policy takes effect without a restart) — so a bad name still starts a daemon that looks
|
||||||
|
* healthy and fails only the first time something spawns without naming a profile. Every other
|
||||||
|
* field this class validates fails here, at load; this one gets the same treatment, calling
|
||||||
|
* {@link PlacementPolicies#fromName} itself as the single source of truth for what is valid
|
||||||
|
* rather than duplicating its accepted set.
|
||||||
|
*
|
||||||
|
* @param placement the raw, possibly null/blank {@code placement} value as parsed (before
|
||||||
|
* {@link #withDefaults()} runs); {@code fromName} itself treats null/blank as
|
||||||
|
* {@code fixed}, so this call changes no default
|
||||||
|
* @throws IllegalStateException when {@code placement} is a name {@link PlacementPolicies} does
|
||||||
|
* not recognize
|
||||||
|
*/
|
||||||
|
private static void rejectUnknownPlacementPolicy(String placement) {
|
||||||
|
try {
|
||||||
|
PlacementPolicies.fromName(placement);
|
||||||
|
} catch (IllegalArgumentException e) {
|
||||||
|
throw new IllegalStateException("refusing to start: " + e.getMessage(), e);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
static List<String> unknownTopLevelKeys(String yaml) {
|
static List<String> unknownTopLevelKeys(String yaml) {
|
||||||
Map<?, ?> raw;
|
Map<?, ?> raw;
|
||||||
try {
|
try {
|
||||||
@@ -1311,9 +1596,17 @@ public record BridgedConfig(
|
|||||||
// config that never mentions it should still get a sane cooldown rather than a null one.
|
// config that never mentions it should still get a sane cooldown rather than a null one.
|
||||||
Integer quarantineCooldown = (quarantineCooldownSeconds != null && quarantineCooldownSeconds > 0)
|
Integer quarantineCooldown = (quarantineCooldownSeconds != null && quarantineCooldownSeconds > 0)
|
||||||
? quarantineCooldownSeconds : DEFAULT_QUARANTINE_COOLDOWN_SECONDS;
|
? quarantineCooldownSeconds : DEFAULT_QUARANTINE_COOLDOWN_SECONDS;
|
||||||
|
// memberCredentials IS defaulted, like guard/lifecycle/auth above, so no reader ever sees a
|
||||||
|
// null. CB-596: an empty MemberCredentials (empty known, empty allow) blocks NOTHING — unlike
|
||||||
|
// guard/lifecycle, an absent block is not a safe "feature off" default here, it is a gap. It
|
||||||
|
// is deliberately not pre-populated with a Java-side name list (that would just reintroduce
|
||||||
|
// the hardcoded-list defect this record replaces); the block must be configured in
|
||||||
|
// bridged.yaml to protect anything. See bridged.example.yaml's memberCredentials: comment.
|
||||||
|
MemberCredentials mc = memberCredentials != null ? memberCredentials
|
||||||
|
: new MemberCredentials(null, List.of(), List.of());
|
||||||
return new BridgedConfig(b, herdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
|
return new BridgedConfig(b, herdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
|
||||||
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
|
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
|
||||||
quarantineCooldown);
|
quarantineCooldown, mc);
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
|
|||||||
@@ -15,6 +15,7 @@ import dev.ltms.bridged.msg.MessageService;
|
|||||||
import dev.ltms.bridged.msg.Rendezvous;
|
import dev.ltms.bridged.msg.Rendezvous;
|
||||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||||
import dev.ltms.bridged.placement.BackendQuarantine;
|
import dev.ltms.bridged.placement.BackendQuarantine;
|
||||||
|
import dev.ltms.bridged.placement.PlacementException;
|
||||||
import dev.ltms.bridged.session.SessionManager;
|
import dev.ltms.bridged.session.SessionManager;
|
||||||
import dev.ltms.bridged.session.MemberSession;
|
import dev.ltms.bridged.session.MemberSession;
|
||||||
import dev.ltms.bridged.session.WorktreeRequest;
|
import dev.ltms.bridged.session.WorktreeRequest;
|
||||||
@@ -576,13 +577,25 @@ public final class BridgeMcp {
|
|||||||
return text("acknowledged " + msgId);
|
return text("acknowledged " + msgId);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** {@code bridge_status}: the live lifecycle status of a worker session. */
|
/**
|
||||||
|
* {@code bridge_status}: the live lifecycle status of a worker session, plus — when the worker
|
||||||
|
* is paused mid-turn in an async {@code bridge_ask} (CB-582) — the open question and how to
|
||||||
|
* answer it, so a lead on its normal poll cadence does not need the ticket to notice.
|
||||||
|
*/
|
||||||
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
|
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
|
||||||
if (isBlank(sessionId)) {
|
if (isBlank(sessionId)) {
|
||||||
return error("sessionId is required");
|
return error("sessionId is required");
|
||||||
}
|
}
|
||||||
try {
|
try {
|
||||||
return text(messages.status(sessionId).name().toLowerCase());
|
String base = messages.status(sessionId).name().toLowerCase();
|
||||||
|
MessageService.PendingAsk ask = messages.pendingAsk(sessionId);
|
||||||
|
if (ask == null) {
|
||||||
|
return text(base);
|
||||||
|
}
|
||||||
|
return text(base + "\n\n[question — worker is waiting for your answer]\n" + ask.question()
|
||||||
|
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + ask.turnId()
|
||||||
|
+ "\" and content set to your answer; the worker resumes the same turn."
|
||||||
|
+ " (ticket " + ask.ticket() + ")");
|
||||||
} catch (HerdrException e) {
|
} catch (HerdrException e) {
|
||||||
return error("herdr error for session " + sessionId + ": " + e.getMessage());
|
return error("herdr error for session " + sessionId + ": " + e.getMessage());
|
||||||
}
|
}
|
||||||
@@ -692,6 +705,10 @@ public final class BridgeMcp {
|
|||||||
return text(json(memberView(member)));
|
return text(json(memberView(member)));
|
||||||
} catch (GuardException e) {
|
} catch (GuardException e) {
|
||||||
return error("subscription boundary: " + e.getMessage());
|
return error("subscription boundary: " + e.getMessage());
|
||||||
|
} catch (PlacementException e) {
|
||||||
|
// CB-599: no candidate had capacity (maxLoad, quarantine, or all-exhausted) — distinct
|
||||||
|
// from "profile does not exist" below.
|
||||||
|
return error("no capacity: " + e.getMessage());
|
||||||
} catch (IllegalArgumentException e) {
|
} catch (IllegalArgumentException e) {
|
||||||
return error(e.getMessage()); // unknown / no-default profile, or a refused resumeSessionId
|
return error(e.getMessage()); // unknown / no-default profile, or a refused resumeSessionId
|
||||||
} catch (PeerUnreachableException e) {
|
} catch (PeerUnreachableException e) {
|
||||||
|
|||||||
@@ -9,6 +9,10 @@ import dev.ltms.bridged.peer.Capability;
|
|||||||
import org.slf4j.Logger;
|
import org.slf4j.Logger;
|
||||||
import org.slf4j.LoggerFactory;
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
|
import java.io.IOException;
|
||||||
|
import java.io.UncheckedIOException;
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
import java.util.EnumSet;
|
import java.util.EnumSet;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
@@ -83,6 +87,21 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
|||||||
fleet);
|
fleet);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Production constructor, plus the CB-596 {@code memberCredentials} policy supplier.
|
||||||
|
*/
|
||||||
|
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||||
|
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||||
|
Function<String, String> env,
|
||||||
|
long spawnReadyTimeoutMs, long spawnReadyPollMs,
|
||||||
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
|
||||||
|
this(agents, spaces, guard, profiles, defaultProfile, env,
|
||||||
|
spawnReadyTimeoutMs,
|
||||||
|
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
|
||||||
|
fleet, memberCredentials);
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
|
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
|
||||||
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
|
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
|
||||||
@@ -125,6 +144,39 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
|||||||
this.guard = guard;
|
this.guard = guard;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Full testability constructor, plus the CB-596 {@code memberCredentials} policy supplier.
|
||||||
|
*/
|
||||||
|
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||||
|
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||||
|
Function<String, String> env,
|
||||||
|
long spawnReadyTimeoutMs,
|
||||||
|
LongSupplier nowMillis, Runnable sleeper,
|
||||||
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
|
||||||
|
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
|
||||||
|
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials);
|
||||||
|
this.guard = guard;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Full testability constructor, plus an injectable host-env-names source for the CB-596
|
||||||
|
* criterion-4 gap detector. Test seam only — every production call site leaves this at the
|
||||||
|
* default (the real {@code System.getenv()} key set) via the constructor above.
|
||||||
|
*/
|
||||||
|
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||||
|
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||||
|
Function<String, String> env,
|
||||||
|
long spawnReadyTimeoutMs,
|
||||||
|
LongSupplier nowMillis, Runnable sleeper,
|
||||||
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
Supplier<BridgedConfig.MemberCredentials> memberCredentials,
|
||||||
|
Supplier<Set<String>> hostEnvNames) {
|
||||||
|
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
|
||||||
|
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials, hostEnvNames);
|
||||||
|
this.guard = guard;
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* {@inheritDoc}
|
* {@inheritDoc}
|
||||||
*
|
*
|
||||||
@@ -182,7 +234,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
|||||||
// model flag so --model keeps outranking the operator's own argv.
|
// model flag so --model keeps outranking the operator's own argv.
|
||||||
// mutableArgv: argvWithBridge may hand back the profile's own (immutable) List.of when it
|
// mutableArgv: argvWithBridge may hand back the profile's own (immutable) List.of when it
|
||||||
// has neither MCP nor a charter — session flags must be added into a list we own.
|
// has neither MCP nor a charter — session flags must be added into a list we own.
|
||||||
List<String> argv = mutableArgv(argvWithBridge(cfg, spec.charter()));
|
List<String> argv = mutableArgv(argvWithBridge(cfg, spec));
|
||||||
String agentSessionId = applySessionIdentity(argv, spec.sessionName(), spec.resumeSessionId());
|
String agentSessionId = applySessionIdentity(argv, spec.sessionName(), spec.resumeSessionId());
|
||||||
return new Launch(workerEnv, argvWithModel(argv, cfg), agentSessionId);
|
return new Launch(workerEnv, argvWithModel(argv, cfg), agentSessionId);
|
||||||
}
|
}
|
||||||
@@ -217,13 +269,31 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* The launch argv, plus an inline {@code --mcp-config} when {@code worker.mcpUrl} is set and
|
* The launch argv, plus an inline {@code --mcp-config} when {@code worker.mcpUrl} is set, the
|
||||||
* {@code --append-system-prompt} when the base composed a charter. Neither touches the profile's
|
* CB-617 charter flags, and {@code --agent <role>} when the role has an agent-definition file
|
||||||
* config; both are pure command-line flags. This inline-flag mount is Claude Code specific —
|
* under the worker's cwd. Neither touches the profile's config; all are pure command-line flags.
|
||||||
* other adapters mount MCP and instructions their own way.
|
* This inline-flag mount is Claude Code specific — other adapters mount MCP and instructions
|
||||||
|
* their own way.
|
||||||
|
*
|
||||||
|
* <p>CB-617: the role charter is operator-authored and often multi-line, so it can never be a
|
||||||
|
* single inline argv element — herdr refuses to shell-encode a multi-line argument
|
||||||
|
* ({@code invalid_agent_argument}). It is written to a temp file instead and mounted with
|
||||||
|
* {@code --append-system-prompt-file}, which this host confirms Claude Code accepts for a
|
||||||
|
* multi-line file.
|
||||||
|
*
|
||||||
|
* <p>CB-618: Claude Code refuses to start when BOTH {@code --append-system-prompt} and
|
||||||
|
* {@code --append-system-prompt-file} are on the command line ("Cannot use both ... Please use
|
||||||
|
* only one"), so the two charters can never travel on separate flags. When both are present they
|
||||||
|
* are concatenated into the one file, role charter first and reply charter last — last is where
|
||||||
|
* the reply rule must sit, because it is the rule that must survive. When only the reply charter
|
||||||
|
* is present it keeps its proven inline {@code --append-system-prompt} delivery, which is also
|
||||||
|
* the only form that reaches a member with no repo checkout.
|
||||||
*/
|
*/
|
||||||
private List<String> argvWithBridge(BridgedConfig.Profile cfg, String charter) {
|
private List<String> argvWithBridge(BridgedConfig.Profile cfg, LaunchSpec spec) {
|
||||||
if (!cfg.hasMcp() && charter == null) {
|
String roleCharter = nonBlank(spec.roleCharter());
|
||||||
|
String replyCharter = nonBlank(spec.replyCharter());
|
||||||
|
Path agentFile = agentDefinitionFile(spec.cwd(), spec.role(), ".claude", "agents");
|
||||||
|
if (!cfg.hasMcp() && roleCharter == null && replyCharter == null && agentFile == null) {
|
||||||
return cfg.argv();
|
return cfg.argv();
|
||||||
}
|
}
|
||||||
List<String> argv = mutableArgv(cfg.argv());
|
List<String> argv = mutableArgv(cfg.argv());
|
||||||
@@ -233,13 +303,45 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
|||||||
argv.add("--mcp-config");
|
argv.add("--mcp-config");
|
||||||
argv.add(mcpJson);
|
argv.add(mcpJson);
|
||||||
}
|
}
|
||||||
if (charter != null) {
|
if (roleCharter != null) {
|
||||||
|
String combined = replyCharter == null ? roleCharter : roleCharter + "\n\n" + replyCharter;
|
||||||
|
argv.add("--append-system-prompt-file");
|
||||||
|
argv.add(writeCharterFile(combined).toString());
|
||||||
|
} else if (replyCharter != null) {
|
||||||
argv.add("--append-system-prompt");
|
argv.add("--append-system-prompt");
|
||||||
argv.add(charter);
|
argv.add(replyCharter);
|
||||||
|
}
|
||||||
|
if (agentFile != null) {
|
||||||
|
argv.add("--agent");
|
||||||
|
argv.add(spec.role().wireName());
|
||||||
}
|
}
|
||||||
return argv;
|
return argv;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** {@code s}, or {@code null} when {@code s} is null/blank — the charter-presence test used above. */
|
||||||
|
private static String nonBlank(String s) {
|
||||||
|
return (s == null || s.isBlank()) ? null : s;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Write the role charter to a fresh temp file so it can be mounted with
|
||||||
|
* {@code --append-system-prompt-file} instead of riding inline in argv (CB-617). Best-effort
|
||||||
|
* cleaned via {@code deleteOnExit} — the same disposable-worker-config cleanup
|
||||||
|
* {@link OpenCodeLauncher#writeConfig} already uses for its charter file, since the process that
|
||||||
|
* reads this file (the spawned peer) outlives this JVM call and there is no spawn-scoped teardown
|
||||||
|
* hook to delete it synchronously.
|
||||||
|
*/
|
||||||
|
private static Path writeCharterFile(String charterText) {
|
||||||
|
try {
|
||||||
|
Path file = Files.createTempFile("bridged-role-charter-", ".md");
|
||||||
|
Files.writeString(file, charterText);
|
||||||
|
file.toFile().deleteOnExit();
|
||||||
|
return file;
|
||||||
|
} catch (IOException e) {
|
||||||
|
throw new UncheckedIOException("cannot write role charter temp file", e);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Pin the model on the command line as well as in {@code ANTHROPIC_MODEL} (CB-533).
|
* Pin the model on the command line as well as in {@code ANTHROPIC_MODEL} (CB-533).
|
||||||
*
|
*
|
||||||
|
|||||||
@@ -17,9 +17,12 @@ import dev.ltms.bridged.peer.SpawnRequest;
|
|||||||
import org.slf4j.Logger;
|
import org.slf4j.Logger;
|
||||||
import org.slf4j.LoggerFactory;
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
import java.security.SecureRandom;
|
import java.security.SecureRandom;
|
||||||
import java.util.ArrayList;
|
import java.util.ArrayList;
|
||||||
import java.util.Collection;
|
import java.util.Collection;
|
||||||
|
import java.util.HashSet;
|
||||||
import java.util.LinkedHashMap;
|
import java.util.LinkedHashMap;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
@@ -90,6 +93,27 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
*/
|
*/
|
||||||
private final Supplier<BridgedConfig.Fleet> fleet;
|
private final Supplier<BridgedConfig.Fleet> fleet;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: the live {@code memberCredentials:} policy, read once per spawn (same hot-reload shape
|
||||||
|
* as {@link #fleet}). {@code null} — either the supplier itself, or what it returns — means no
|
||||||
|
* policy is configured and {@link #applyMemberCredentialPolicy} shadows nothing.
|
||||||
|
*/
|
||||||
|
private final Supplier<BridgedConfig.MemberCredentials> memberCredentials;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Enumerates the daemon's own process environment variable NAMES ONLY, never values — the CB-596
|
||||||
|
* criterion-4 gap detector's data source (see {@link #logCredentialGap}). Injectable for tests;
|
||||||
|
* production always resolves to the real {@code System.getenv()} key set.
|
||||||
|
*
|
||||||
|
* <p>Deliberately the daemon's own environment, not the spawned pane's: nothing in the herdr
|
||||||
|
* client surface lets the daemon read back an arbitrary command's output from a pane before the
|
||||||
|
* peer starts in it, so there is no channel to inspect the pane's environment directly. The
|
||||||
|
* daemon's own process is started the same way (a login shell sourcing the same secret store —
|
||||||
|
* see CB-592's investigation of {@code secrets.sh}), so on a single-host deployment its env
|
||||||
|
* mirrors what the pane's login shell is about to export.
|
||||||
|
*/
|
||||||
|
private final Supplier<Set<String>> hostEnvNames;
|
||||||
|
|
||||||
/** The final instruction always requires a bridge reply when the bridge MCP is mounted. */
|
/** The final instruction always requires a bridge reply when the bridge MCP is mounted. */
|
||||||
protected static final String REPLY_CHARTER =
|
protected static final String REPLY_CHARTER =
|
||||||
"You are a spawned member in the claude-bridge fleet. Every message you receive arrives "
|
"You are a spawned member in the claude-bridge fleet. Every message you receive arrives "
|
||||||
@@ -166,6 +190,41 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
long spawnReadyTimeoutMs,
|
long spawnReadyTimeoutMs,
|
||||||
LongSupplier nowMillis, Runnable sleeper,
|
LongSupplier nowMillis, Runnable sleeper,
|
||||||
Supplier<BridgedConfig.Fleet> fleet) {
|
Supplier<BridgedConfig.Fleet> fleet) {
|
||||||
|
this(namePrefix, agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
|
||||||
|
nowMillis, sleeper, fleet, null);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* As above, plus the live {@code memberCredentials} policy (CB-596).
|
||||||
|
*
|
||||||
|
* @param memberCredentials live member-credential policy, read once per spawn; {@code null} ⇒
|
||||||
|
* no policy configured, so a spawn shadows nothing. A separate
|
||||||
|
* constructor rather than a new parameter on the one above, so every
|
||||||
|
* existing call site keeps the pre-CB-596 default without an edit.
|
||||||
|
*/
|
||||||
|
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
|
||||||
|
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||||
|
Function<String, String> env,
|
||||||
|
long spawnReadyTimeoutMs,
|
||||||
|
LongSupplier nowMillis, Runnable sleeper,
|
||||||
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
|
||||||
|
this(namePrefix, agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
|
||||||
|
nowMillis, sleeper, fleet, memberCredentials, null);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* As above, plus an injectable {@link #hostEnvNames} source for the CB-596 gap detector. Test
|
||||||
|
* seam only — every production call site leaves this {@code null} and gets the real host env.
|
||||||
|
*/
|
||||||
|
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
|
||||||
|
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||||
|
Function<String, String> env,
|
||||||
|
long spawnReadyTimeoutMs,
|
||||||
|
LongSupplier nowMillis, Runnable sleeper,
|
||||||
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
Supplier<BridgedConfig.MemberCredentials> memberCredentials,
|
||||||
|
Supplier<Set<String>> hostEnvNames) {
|
||||||
this.fleet = fleet;
|
this.fleet = fleet;
|
||||||
this.namePrefix = namePrefix;
|
this.namePrefix = namePrefix;
|
||||||
this.agents = agents;
|
this.agents = agents;
|
||||||
@@ -176,6 +235,8 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
this.spawnReadyTimeoutMs = spawnReadyTimeoutMs;
|
this.spawnReadyTimeoutMs = spawnReadyTimeoutMs;
|
||||||
this.nowMillis = nowMillis;
|
this.nowMillis = nowMillis;
|
||||||
this.sleeper = sleeper;
|
this.sleeper = sleeper;
|
||||||
|
this.memberCredentials = memberCredentials;
|
||||||
|
this.hostEnvNames = hostEnvNames != null ? hostEnvNames : () -> System.getenv().keySet();
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- adapter seams -------------------------------------------------------------------------
|
// --- adapter seams -------------------------------------------------------------------------
|
||||||
@@ -220,8 +281,35 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/** All per-spawn values adapters may need, including the base-composed effective charter. */
|
/**
|
||||||
protected record LaunchSpec(String sessionName, String resumeSessionId, MemberRole role, String charter) {
|
* All per-spawn values adapters may need.
|
||||||
|
*
|
||||||
|
* <p>{@code charter} is the base-composed effective charter (role charter, then the reply
|
||||||
|
* charter, joined by a blank line) — kept for an adapter that mounts both as one blob (opencode
|
||||||
|
* writes it to a single file) and as the exact input {@link CharterReceipt#compose} fingerprints.
|
||||||
|
* {@code roleCharter} and {@code replyCharter} are the same text split back into its two parts
|
||||||
|
* (CB-617), for an adapter that must deliver them differently: the role charter is
|
||||||
|
* operator-authored and often multi-line, so it cannot travel as an inline argv element (herdr
|
||||||
|
* refuses to shell-encode a multi-line argument); the reply charter is always one line and is
|
||||||
|
* demonstrated to encode, so it may still go inline. {@code cwd} is the spawn's resolved working
|
||||||
|
* directory (CB-112), needed to look up a role's agent-definition file before the peer starts.
|
||||||
|
*/
|
||||||
|
protected record LaunchSpec(String sessionName, String resumeSessionId, MemberRole role, String charter,
|
||||||
|
String roleCharter, String replyCharter, String cwd) {
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The role's agent-definition file under {@code <cwd>/<dir1>/<dir2>/<role>.md}, or {@code null}
|
||||||
|
* when absent or inapplicable (no role, no cwd, or the file does not exist) — CB-617. A member
|
||||||
|
* whose role has no such file must still spawn, so this is a lookup, never a requirement: the
|
||||||
|
* caller passes {@code --agent <role>} only when the return value is non-null.
|
||||||
|
*/
|
||||||
|
protected static Path agentDefinitionFile(String cwd, MemberRole role, String dir1, String dir2) {
|
||||||
|
if (role == null || cwd == null || cwd.isBlank()) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
Path candidate = Path.of(cwd, dir1, dir2, role.wireName() + ".md");
|
||||||
|
return Files.isRegularFile(candidate) ? candidate : null;
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- profile surface -----------------------------------------------------------------------
|
// --- profile surface -----------------------------------------------------------------------
|
||||||
@@ -328,9 +416,10 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
// start has no bridge_spawn result and no roster row, so the failure log below is the only
|
// start has no bridge_spawn result and no roster row, so the failure log below is the only
|
||||||
// surface the byte count can appear on. The charter text itself is never logged.
|
// surface the byte count can appear on. The charter text itself is never logged.
|
||||||
CharterReceipt receipt = CharterReceipt.compose(role, cfg.profile(), roleCharter, charter);
|
CharterReceipt receipt = CharterReceipt.compose(role, cfg.profile(), roleCharter, charter);
|
||||||
|
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
|
||||||
try {
|
try {
|
||||||
Launch launch = buildLaunch(cfg, new LaunchSpec(sessionName, resumeSessionId, role, charter));
|
Launch launch = buildLaunch(cfg, new LaunchSpec(sessionName, resumeSessionId, role, charter,
|
||||||
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
|
roleCharter, replyCharter, cwd));
|
||||||
Agent agent = cfg.tabPlacement()
|
Agent agent = cfg.tabPlacement()
|
||||||
? spawnInTab(cfg, launch.env(), launch.argv(), cwd, role, liveFleet)
|
? spawnInTab(cfg, launch.env(), launch.argv(), cwd, role, liveFleet)
|
||||||
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd, charter);
|
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd, charter);
|
||||||
@@ -767,12 +856,13 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* CB-592: overlay value that shadows the admin {@code GITEA_ACCESS_TOKEN} a herdr pane
|
* CB-596: overlay value that shadows any host credential a herdr pane otherwise inherits from
|
||||||
* otherwise inherits from herdr's own login-shell process environment (gitea issue #77).
|
* herdr's own login-shell process environment (gitea issue #82, superseding CB-592's single
|
||||||
* herdr spawns a pane from its <em>own</em> process environment and layers our map on top —
|
* hardcoded {@code GITEA_ACCESS_TOKEN} name — see {@link #applyMemberCredentialPolicy}). herdr
|
||||||
|
* spawns a pane from its <em>own</em> process environment and layers our map on top —
|
||||||
* {@link dev.ltms.bridged.herdr.WorkspaceControl#createTab} and {@code #splitPane} send only
|
* {@link dev.ltms.bridged.herdr.WorkspaceControl#createTab} and {@code #splitPane} send only
|
||||||
* the keys we put in that map, so any key we never mention passes straight through from
|
* the keys we put in that map, so any key we never mention passes straight through from
|
||||||
* herdr's own shell, admin token included.
|
* herdr's own shell, admin credentials included.
|
||||||
*
|
*
|
||||||
* <p>Deliberately a non-blank sentinel, not {@code ""}. Whether an empty-string overlay value
|
* <p>Deliberately a non-blank sentinel, not {@code ""}. Whether an empty-string overlay value
|
||||||
* overrides an inherited variable or is skipped as blank could not be settled by reading this
|
* overrides an inherited variable or is skipped as blank could not be settled by reading this
|
||||||
@@ -782,21 +872,21 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
* javadoc), and that is only demonstrated for a non-blank value, so this reuses the same,
|
* javadoc), and that is only demonstrated for a non-blank value, so this reuses the same,
|
||||||
* proven-reliable shape rather than the unverified one.
|
* proven-reliable shape rather than the unverified one.
|
||||||
*
|
*
|
||||||
* <p><b>MEASURED ON A LIVE PANE, 2026-08-15: this sentinel alone does NOT hold.</b> The overlay
|
* <p><b>MEASURED ON A LIVE PANE, 2026-08-15 (CB-592): this sentinel alone does NOT hold.</b> The
|
||||||
* itself works — {@code GITEA_TOKEN} is injected here, is exported by no shell file, and does
|
* overlay itself works — {@code GITEA_TOKEN} is injected here, is exported by no shell file, and
|
||||||
* reach the pane. The sentinel loses one step later. A herdr pane runs a <em>login</em> shell,
|
* does reach the pane. The sentinel loses one step later. A herdr pane runs a <em>login</em>
|
||||||
* {@code ~/.zprofile} sources {@code ${SHARED_ENV}/tools/secrets.sh}, and that file does a plain
|
* shell, {@code ~/.zprofile} sources {@code ${SHARED_ENV}/tools/secrets.sh}, and that file does a
|
||||||
* unconditional {@code export GITEA_ACCESS_TOKEN=...}. A login shell overwrites a value already
|
* plain unconditional {@code export GITEA_ACCESS_TOKEN=...}. A login shell overwrites a value
|
||||||
* in the environment, so the real admin token is put back over this sentinel before the member
|
* already in the environment, so the real admin token is put back over this sentinel before the
|
||||||
* process ever starts. That defeat applies to <em>every</em> name {@code secrets.sh} exports,
|
* member process ever starts. That defeat applies to <em>every</em> name {@code secrets.sh}
|
||||||
* and no launcher-side overlay can win against it.
|
* exports, and no launcher-side overlay can win against it.
|
||||||
*
|
*
|
||||||
* <p>So this constant is not the control on its own — {@link #MEMBER_MARKER} is the other half.
|
* <p>So this constant is not the control on its own — {@link #MEMBER_MARKER} is the other half.
|
||||||
* Keeping the sentinel is still worth it: it is correct for any peer kind whose pane does not
|
* Keeping the sentinel is still worth it: it is correct for any peer kind whose pane does not
|
||||||
* start a login shell, and it makes the intent explicit at the one place every adapter passes.
|
* start a login shell, and it makes the intent explicit at the one place every adapter passes.
|
||||||
*/
|
*/
|
||||||
private static final String BLOCKED_GITEA_ACCESS_TOKEN =
|
private static final String BLOCKED_CREDENTIAL_SENTINEL =
|
||||||
"blocked-by-bridged-cb592-see-gitea-issue-77";
|
"blocked-by-bridged-cb596-see-gitea-issue-82";
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* CB-592: marks a pane as a bridged member so a shell startup file can decline to export
|
* CB-592: marks a pane as a bridged member so a shell startup file can decline to export
|
||||||
@@ -835,11 +925,11 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
|
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
|
||||||
* {@code baseUrl} and nothing else.
|
* {@code baseUrl} and nothing else.
|
||||||
*
|
*
|
||||||
* <p>The CB-592 shadow and marker are put in <em>last</em>, after the profile's own
|
* <p>The CB-596 credential shadow and the CB-592 marker are put in <em>last</em>, after the
|
||||||
* {@code env:}, so no profile — present or future — can restore the admin token, or hide that
|
* profile's own {@code env:}, so no profile — present or future — can restore a blocked
|
||||||
* the pane is a member, by naming either in config. This is the one place both are applied:
|
* credential, or hide that the pane is a member, by naming either in config. This is the one
|
||||||
* every {@code buildLaunch} in every adapter calls this first, so a new profile, and a peer
|
* place both are applied: every {@code buildLaunch} in every adapter calls this first, so a new
|
||||||
* kind not yet written, gets them for free.
|
* profile, and a peer kind not yet written, gets them for free.
|
||||||
*/
|
*/
|
||||||
protected Map<String, String> baseEnv(BridgedConfig.Profile cfg) {
|
protected Map<String, String> baseEnv(BridgedConfig.Profile cfg) {
|
||||||
Map<String, String> workerEnv = new LinkedHashMap<>();
|
Map<String, String> workerEnv = new LinkedHashMap<>();
|
||||||
@@ -850,11 +940,70 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
|||||||
if (cfg != null && cfg.env() != null) {
|
if (cfg != null && cfg.env() != null) {
|
||||||
workerEnv.putAll(cfg.env());
|
workerEnv.putAll(cfg.env());
|
||||||
}
|
}
|
||||||
workerEnv.put("GITEA_ACCESS_TOKEN", BLOCKED_GITEA_ACCESS_TOKEN);
|
applyMemberCredentialPolicy(workerEnv);
|
||||||
workerEnv.put(MEMBER_MARKER, "1");
|
workerEnv.put(MEMBER_MARKER, "1");
|
||||||
return workerEnv;
|
return workerEnv;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: shadow every configured {@code memberCredentials.known} name that is not also
|
||||||
|
* {@code allow}-ed, replacing CB-592's single hardcoded {@code GITEA_ACCESS_TOKEN} name (gitea
|
||||||
|
* issue #82). An allow-listed name is deliberately left unmentioned here — see {@link
|
||||||
|
* #BLOCKED_CREDENTIAL_SENTINEL}'s javadoc for why an overlay entry is the only way to shadow an
|
||||||
|
* inherited value, which is exactly why an allowed name must get NO entry: any entry at all,
|
||||||
|
* blank or not, risks overriding the real value the pane needs.
|
||||||
|
*
|
||||||
|
* <p>No {@code memberCredentials} configured — the supplier is {@code null}, or it resolves to
|
||||||
|
* one whose {@code known} list is empty — shadows nothing. This is a real, config-driven gap
|
||||||
|
* (see {@link BridgedConfig.MemberCredentials}'s javadoc), not a safe default: deny-by-default
|
||||||
|
* only defends names the operator has actually enumerated in {@code known}.
|
||||||
|
*/
|
||||||
|
private void applyMemberCredentialPolicy(Map<String, String> workerEnv) {
|
||||||
|
BridgedConfig.MemberCredentials creds = memberCredentials == null ? null : memberCredentials.get();
|
||||||
|
if (creds == null) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
for (String name : creds.blockedSet()) {
|
||||||
|
workerEnv.put(name, BLOCKED_CREDENTIAL_SENTINEL);
|
||||||
|
}
|
||||||
|
logCredentialGap(creds);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Credential-shaped env var name heuristic for {@link #logCredentialGap} — case-insensitive. */
|
||||||
|
private static final Pattern CREDENTIAL_SHAPED_NAME =
|
||||||
|
Pattern.compile("(?i).*(TOKEN|SECRET|_KEY|APIKEY|PASSWORD|CREDENTIAL|AUTH).*");
|
||||||
|
|
||||||
|
/** Guards {@link #logCredentialGap} to one WARN per launcher instance, not one per spawn. */
|
||||||
|
private final AtomicBoolean credentialGapLogged = new AtomicBoolean();
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596 criterion 4: a credential-shaped host env var name on neither {@code known} nor
|
||||||
|
* {@code allow} is not silently allowed — it is reported. {@link #hostEnvNames} enumerates the
|
||||||
|
* daemon's own environment (see that field's javadoc for why the daemon's env is read rather
|
||||||
|
* than the spawned pane's, which the daemon has no channel to inspect at spawn time); this logs
|
||||||
|
* every such NAME, at WARN, at most once per launcher instance — never a value, a prefix of a
|
||||||
|
* value, or a hash of a value, so the log itself cannot leak anything.
|
||||||
|
*/
|
||||||
|
private void logCredentialGap(BridgedConfig.MemberCredentials creds) {
|
||||||
|
Set<String> covered = new HashSet<>(creds.known());
|
||||||
|
covered.addAll(creds.allow());
|
||||||
|
List<String> gap = hostEnvNames.get().stream()
|
||||||
|
.filter(name -> CREDENTIAL_SHAPED_NAME.matcher(name).matches())
|
||||||
|
.filter(name -> !covered.contains(name))
|
||||||
|
.sorted()
|
||||||
|
.toList();
|
||||||
|
if (gap.isEmpty()) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (credentialGapLogged.compareAndSet(false, true)) {
|
||||||
|
log.warn("memberCredentials gap: {} credential-shaped env var name(s) are on neither "
|
||||||
|
+ "known: nor allow: — every member pane inherits them UNBLOCKED — {}. "
|
||||||
|
+ "Add each to memberCredentials.known (blocked by default) or .allow "
|
||||||
|
+ "(if a member legitimately needs it).",
|
||||||
|
gap.size(), gap);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/** Defensive copy of {@code argv} plus room to append launch flags. */
|
/** Defensive copy of {@code argv} plus room to append launch flags. */
|
||||||
protected static List<String> mutableArgv(List<String> argv) {
|
protected static List<String> mutableArgv(List<String> argv) {
|
||||||
return new ArrayList<>(argv);
|
return new ArrayList<>(argv);
|
||||||
|
|||||||
@@ -104,6 +104,20 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
|||||||
defaultConfigRoot(), defaultDiscoveryRoot(), fleet);
|
defaultConfigRoot(), defaultDiscoveryRoot(), fleet);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Production constructor, plus the CB-596 {@code memberCredentials} policy supplier.
|
||||||
|
*/
|
||||||
|
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
|
||||||
|
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||||
|
Function<String, String> env,
|
||||||
|
long spawnReadyTimeoutMs, long spawnReadyPollMs,
|
||||||
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
|
||||||
|
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
|
||||||
|
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
|
||||||
|
defaultConfigRoot(), defaultDiscoveryRoot(), fleet, memberCredentials);
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply a
|
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply a
|
||||||
* fake clock ({@code nowMillis}), poll-loop wait ({@code sleeper}), and a temp {@code configRoot}
|
* fake clock ({@code nowMillis}), poll-loop wait ({@code sleeper}), and a temp {@code configRoot}
|
||||||
@@ -151,6 +165,23 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
|||||||
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
|
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Full testability constructor, plus the CB-596 {@code memberCredentials} policy supplier.
|
||||||
|
*/
|
||||||
|
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
|
||||||
|
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||||
|
Function<String, String> env,
|
||||||
|
long spawnReadyTimeoutMs,
|
||||||
|
LongSupplier nowMillis, Runnable sleeper,
|
||||||
|
Path configRoot, Path discoveryRoot,
|
||||||
|
Supplier<BridgedConfig.Fleet> fleet,
|
||||||
|
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
|
||||||
|
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
|
||||||
|
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials);
|
||||||
|
this.configRoot = configRoot;
|
||||||
|
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
|
||||||
|
}
|
||||||
|
|
||||||
private static Path defaultConfigRoot() {
|
private static Path defaultConfigRoot() {
|
||||||
return Path.of(System.getProperty("java.io.tmpdir"));
|
return Path.of(System.getProperty("java.io.tmpdir"));
|
||||||
}
|
}
|
||||||
@@ -177,8 +208,24 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
|||||||
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg, spec.charter()).toString());
|
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg, spec.charter()).toString());
|
||||||
}
|
}
|
||||||
applyGitToken(workerEnv, cfg);
|
applyGitToken(workerEnv, cfg);
|
||||||
return new Launch(workerEnv,
|
List<String> argv = argvWithResume(argvWithModel(argvWithAuto(cfg), cfg), spec.resumeSessionId());
|
||||||
argvWithResume(argvWithModel(argvWithAuto(cfg), cfg), spec.resumeSessionId()));
|
return new Launch(workerEnv, argvWithAgent(argv, spec));
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The launch argv plus, when the role has an agent-definition file under the worker's cwd,
|
||||||
|
* opencode's {@code --agent <role>} flag (CB-617). A role with no such file gets nothing added —
|
||||||
|
* the member must still spawn.
|
||||||
|
*/
|
||||||
|
private List<String> argvWithAgent(List<String> argv, LaunchSpec spec) {
|
||||||
|
Path agentFile = agentDefinitionFile(spec.cwd(), spec.role(), ".opencode", "agent");
|
||||||
|
if (agentFile == null) {
|
||||||
|
return argv;
|
||||||
|
}
|
||||||
|
List<String> withAgent = mutableArgv(argv);
|
||||||
|
withAgent.add("--agent");
|
||||||
|
withAgent.add(spec.role().wireName());
|
||||||
|
return withAgent;
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
|
|||||||
@@ -7,6 +7,7 @@ import com.rabbitmq.client.ConnectionFactory;
|
|||||||
import com.rabbitmq.client.DeliverCallback;
|
import com.rabbitmq.client.DeliverCallback;
|
||||||
import com.rabbitmq.client.Recoverable;
|
import com.rabbitmq.client.Recoverable;
|
||||||
import com.rabbitmq.client.RecoveryListener;
|
import com.rabbitmq.client.RecoveryListener;
|
||||||
|
import com.rabbitmq.client.Return;
|
||||||
import org.slf4j.Logger;
|
import org.slf4j.Logger;
|
||||||
import org.slf4j.LoggerFactory;
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
@@ -14,7 +15,13 @@ import java.io.IOException;
|
|||||||
import java.nio.charset.StandardCharsets;
|
import java.nio.charset.StandardCharsets;
|
||||||
import java.util.LinkedHashMap;
|
import java.util.LinkedHashMap;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
|
import java.util.NavigableMap;
|
||||||
|
import java.util.concurrent.CompletableFuture;
|
||||||
import java.util.concurrent.ConcurrentHashMap;
|
import java.util.concurrent.ConcurrentHashMap;
|
||||||
|
import java.util.concurrent.ConcurrentSkipListMap;
|
||||||
|
import java.util.concurrent.ExecutionException;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.TimeoutException;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
|
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
|
||||||
@@ -29,6 +36,27 @@ import java.util.concurrent.ConcurrentHashMap;
|
|||||||
* leaves them on the broker — it redelivers on reconnect. That is the durability the in-memory
|
* leaves them on the broker — it redelivers on reconnect. That is the durability the in-memory
|
||||||
* adapter cannot give, with the port contract preserved.
|
* adapter cannot give, with the port contract preserved.
|
||||||
*
|
*
|
||||||
|
* <p><strong>Prefetch bounds the held backlog (CB-527).</strong> The consumer channel calls
|
||||||
|
* {@code basicQos} with a configurable prefetch count ({@link #DEFAULT_PREFETCH} unless the caller
|
||||||
|
* passes another value to {@link #open(String, int)}) before starting any consumer. Without a bound,
|
||||||
|
* the broker pushes its entire queue into {@link #held} the instant a target is {@link #own owned},
|
||||||
|
* so an undrained primary grows the JVM heap without limit and any queue-level control
|
||||||
|
* ({@code x-max-length}, per-message TTL) never fires because the queue never actually holds a
|
||||||
|
* backlog. Prefetch keeps the backlog where it is visible — on the broker — until the owner drains it.
|
||||||
|
*
|
||||||
|
* <p><strong>Publishes require a confirmed, routable delivery (CB-528).</strong> {@link #publish}
|
||||||
|
* runs on a channel separate from the consume/ack channel ({@link #channel}), so a slow or blocked
|
||||||
|
* publish confirm can never hold {@link #channelLock} and stall an ack — the ack path never waits on
|
||||||
|
* a publish confirm. That publish channel is in publisher-confirm mode and every publish sets the
|
||||||
|
* {@code mandatory} flag, so an unroutable publish (queue not declared, e.g. the owner never called
|
||||||
|
* {@link #own}) is returned by the broker instead of silently dropped. The broker sends the
|
||||||
|
* <em>return</em> for an unroutable message before the <em>confirm</em> that covers it — the ack/nack
|
||||||
|
* callback checks the returned-set at confirm time rather than assuming an ack means routed — so
|
||||||
|
* "confirmed" here means "durably queued", not merely "accepted by the broker". A returned or nacked
|
||||||
|
* (or un-confirmed within the timeout) publish surfaces as an {@link IllegalStateException} on the
|
||||||
|
* caller's thread; the caller — {@link MessageService#reply} — must not report success for a
|
||||||
|
* black-holed reply.
|
||||||
|
*
|
||||||
* <p><strong>Ownership is explicit.</strong> {@link #own} declares the queue and starts the consumer;
|
* <p><strong>Ownership is explicit.</strong> {@link #own} declares the queue and starts the consumer;
|
||||||
* {@link #release} cancels it. {@link #publish} sends to the queue but does <em>not</em> imply ownership
|
* {@link #release} cancels it. {@link #publish} sends to the queue but does <em>not</em> imply ownership
|
||||||
* and does not attach a consumer. This split is required by CB-308 federation, where one gateway may
|
* and does not attach a consumer. This split is required by CB-308 federation, where one gateway may
|
||||||
@@ -54,6 +82,12 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
|||||||
private static final String QUEUE_PREFIX = "agent.";
|
private static final String QUEUE_PREFIX = "agent.";
|
||||||
private static final String QUEUE_SUFFIX = ".inbox";
|
private static final String QUEUE_SUFFIX = ".inbox";
|
||||||
|
|
||||||
|
/** CB-527: the prefetch used when a caller does not pass an explicit value to {@link #open(String, int)}. */
|
||||||
|
public static final int DEFAULT_PREFETCH = 32;
|
||||||
|
|
||||||
|
/** How long {@link #publish} waits for its publisher confirm before failing the call (CB-528). */
|
||||||
|
private static final long CONFIRM_TIMEOUT_MS = 10_000L;
|
||||||
|
|
||||||
private final Connection connection;
|
private final Connection connection;
|
||||||
private final Channel channel;
|
private final Channel channel;
|
||||||
/** All channel operations (publish/declare/ack/cancel) serialize on this — a Channel is not thread-safe. */
|
/** All channel operations (publish/declare/ack/cancel) serialize on this — a Channel is not thread-safe. */
|
||||||
@@ -63,39 +97,87 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
|||||||
/** Targets whose queue is declared and consumer is running, mapped to their broker consumer tag. */
|
/** Targets whose queue is declared and consumer is running, mapped to their broker consumer tag. */
|
||||||
private final ConcurrentHashMap<String, String> consumerTags = new ConcurrentHashMap<>();
|
private final ConcurrentHashMap<String, String> consumerTags = new ConcurrentHashMap<>();
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-528: a dedicated channel for {@link #publish}, kept separate from {@link #channel} (consume
|
||||||
|
* + ack) so a publish confirm round trip never blocks under {@link #channelLock} and stalls an ack.
|
||||||
|
*/
|
||||||
|
private final Channel publishChannel;
|
||||||
|
private final Object publishChannelLock = new Object();
|
||||||
|
/** In-flight publishes awaiting their confirm, keyed by the publish channel's sequence number. */
|
||||||
|
private final ConcurrentSkipListMap<Long, Pending> pendingBySeq = new ConcurrentSkipListMap<>();
|
||||||
|
/**
|
||||||
|
* The same in-flight publishes, keyed by {@code msgId} — a broker {@code Return} carries no delivery
|
||||||
|
* tag. Assumes {@code msgId} is unique per in-flight publish: a second {@link #publish} for a
|
||||||
|
* {@code msgId} still awaiting its confirm would overwrite this entry and misdirect
|
||||||
|
* {@link #onReturn}'s lookup. Not reachable today — {@code MessageService.reply} generates a fresh
|
||||||
|
* {@code UUID} per call — so no guard is added for it.
|
||||||
|
*/
|
||||||
|
private final ConcurrentHashMap<String, Pending> pendingByMsgId = new ConcurrentHashMap<>();
|
||||||
|
|
||||||
/** A message pulled off the broker but not yet acked: its delivery-tag plus the port payload. */
|
/** A message pulled off the broker but not yet acked: its delivery-tag plus the port payload. */
|
||||||
private record Held(long deliveryTag, InboxMessage message) {}
|
private record Held(long deliveryTag, InboxMessage message) {}
|
||||||
|
|
||||||
/** Connect to {@code uri} (e.g. {@code amqp://guest:guest@127.0.0.1:5672/}) and open the inbox. */
|
/** A publish awaiting its confirm; {@link #returned} records whether the broker already returned it. */
|
||||||
|
private static final class Pending {
|
||||||
|
final String msgId;
|
||||||
|
final CompletableFuture<Void> confirmed = new CompletableFuture<>();
|
||||||
|
volatile boolean returned;
|
||||||
|
|
||||||
|
Pending(String msgId) {
|
||||||
|
this.msgId = msgId;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Connect to {@code uri} (e.g. {@code amqp://guest:guest@127.0.0.1:5672/}) with {@link #DEFAULT_PREFETCH}. */
|
||||||
public static AmqpReplyInbox open(String uri) {
|
public static AmqpReplyInbox open(String uri) {
|
||||||
|
return open(uri, DEFAULT_PREFETCH);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** As {@link #open(String)}, with an explicit consumer prefetch (CB-527: caps the held backlog per target). */
|
||||||
|
public static AmqpReplyInbox open(String uri, int prefetch) {
|
||||||
try {
|
try {
|
||||||
ConnectionFactory factory = new ConnectionFactory();
|
ConnectionFactory factory = new ConnectionFactory();
|
||||||
factory.setUri(uri);
|
factory.setUri(uri);
|
||||||
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
|
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
|
||||||
factory.setAutomaticRecoveryEnabled(true);
|
factory.setAutomaticRecoveryEnabled(true);
|
||||||
factory.setTopologyRecoveryEnabled(true);
|
factory.setTopologyRecoveryEnabled(true);
|
||||||
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"));
|
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"), prefetch);
|
||||||
} catch (Exception e) {
|
} catch (Exception e) {
|
||||||
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
|
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Wrap an already-open connection (injection seam for the contract test). */
|
/** Wrap an already-open connection with {@link #DEFAULT_PREFETCH} (injection seam for the contract test). */
|
||||||
AmqpReplyInbox(Connection connection) {
|
AmqpReplyInbox(Connection connection) {
|
||||||
|
this(connection, DEFAULT_PREFETCH);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** As above, with an explicit prefetch (injection seam for the contract test). */
|
||||||
|
AmqpReplyInbox(Connection connection, int prefetch) {
|
||||||
this.connection = connection;
|
this.connection = connection;
|
||||||
try {
|
try {
|
||||||
this.channel = connection.createChannel();
|
this.channel = connection.createChannel();
|
||||||
|
// CB-527: bound the held backlog per owned target — must be set before any own()/basicConsume.
|
||||||
|
this.channel.basicQos(prefetch);
|
||||||
|
this.publishChannel = connection.createChannel();
|
||||||
|
this.publishChannel.confirmSelect();
|
||||||
|
this.publishChannel.addReturnListener(this::onReturn);
|
||||||
|
this.publishChannel.addConfirmListener(this::onAck, this::onNack);
|
||||||
} catch (IOException e) {
|
} catch (IOException e) {
|
||||||
throw new IllegalStateException("cannot open AMQP channel", e);
|
throw new IllegalStateException("cannot open AMQP channel", e);
|
||||||
}
|
}
|
||||||
// On automatic recovery the broker redelivers unacked messages with FRESH delivery-tags; the
|
// On automatic recovery the broker redelivers unacked messages with FRESH delivery-tags; the
|
||||||
// tags we were holding are now stale. Drop the held snapshot so the re-attached consumer
|
// tags we were holding are now stale. Drop the held snapshot so the re-attached consumer
|
||||||
// repopulates it with valid tags (dedup by msgId still prevents any double-queue).
|
// repopulates it with valid tags (dedup by msgId still prevents any double-queue). Any publish
|
||||||
|
// confirm still in flight when the connection dropped is equally stale — its sequence number
|
||||||
|
// meant nothing on the old channel and means nothing on the recovered one, so fail it now
|
||||||
|
// rather than let it silently ride out CONFIRM_TIMEOUT_MS.
|
||||||
if (connection instanceof Recoverable recoverable) {
|
if (connection instanceof Recoverable recoverable) {
|
||||||
recoverable.addRecoveryListener(new RecoveryListener() {
|
recoverable.addRecoveryListener(new RecoveryListener() {
|
||||||
@Override
|
@Override
|
||||||
public void handleRecovery(Recoverable recoverable) {
|
public void handleRecovery(Recoverable recoverable) {
|
||||||
held.clear();
|
held.clear();
|
||||||
|
failPendingPublishesOnRecovery();
|
||||||
log.info("AMQP connection recovered; cleared held replies for fresh redelivery");
|
log.info("AMQP connection recovered; cleared held replies for fresh redelivery");
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -141,6 +223,12 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Publish {@code content} and block until the broker's publisher confirm for it lands (CB-528).
|
||||||
|
* Throws {@link IllegalStateException} if the message is returned as unroutable, nacked, or not
|
||||||
|
* confirmed within {@link #CONFIRM_TIMEOUT_MS} — the caller must treat that as a failed publish,
|
||||||
|
* not a lost-and-forgotten one.
|
||||||
|
*/
|
||||||
@Override
|
@Override
|
||||||
public void publish(String target, String msgId, String content) {
|
public void publish(String target, String msgId, String content) {
|
||||||
AMQP.BasicProperties props = new AMQP.BasicProperties.Builder()
|
AMQP.BasicProperties props = new AMQP.BasicProperties.Builder()
|
||||||
@@ -148,12 +236,35 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
|||||||
.deliveryMode(2) // persistent — survives a broker restart
|
.deliveryMode(2) // persistent — survives a broker restart
|
||||||
.contentType("text/plain")
|
.contentType("text/plain")
|
||||||
.build();
|
.build();
|
||||||
try {
|
Pending pending = new Pending(msgId);
|
||||||
synchronized (channelLock) {
|
long seq;
|
||||||
channel.basicPublish("", queueName(target), props, content.getBytes(StandardCharsets.UTF_8));
|
synchronized (publishChannelLock) {
|
||||||
|
seq = publishChannel.getNextPublishSeqNo();
|
||||||
|
pendingBySeq.put(seq, pending);
|
||||||
|
pendingByMsgId.put(msgId, pending);
|
||||||
|
try {
|
||||||
|
publishChannel.basicPublish("", queueName(target), true, props,
|
||||||
|
content.getBytes(StandardCharsets.UTF_8));
|
||||||
|
} catch (IOException e) {
|
||||||
|
pendingBySeq.remove(seq, pending);
|
||||||
|
pendingByMsgId.remove(msgId, pending);
|
||||||
|
throw new IllegalStateException("cannot publish reply to " + queueName(target), e);
|
||||||
}
|
}
|
||||||
} catch (IOException e) {
|
}
|
||||||
throw new IllegalStateException("cannot publish reply to " + queueName(target), e);
|
try {
|
||||||
|
pending.confirmed.get(CONFIRM_TIMEOUT_MS, TimeUnit.MILLISECONDS);
|
||||||
|
} catch (ExecutionException e) {
|
||||||
|
Throwable cause = e.getCause();
|
||||||
|
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
|
||||||
|
} catch (TimeoutException e) {
|
||||||
|
throw new IllegalStateException("publish confirm for reply " + msgId + " to " + queueName(target)
|
||||||
|
+ " timed out after " + CONFIRM_TIMEOUT_MS + "ms — broker may be unreachable or overloaded", e);
|
||||||
|
} catch (InterruptedException e) {
|
||||||
|
Thread.currentThread().interrupt();
|
||||||
|
throw new IllegalStateException("interrupted awaiting publish confirm for " + msgId, e);
|
||||||
|
} finally {
|
||||||
|
pendingBySeq.remove(seq, pending);
|
||||||
|
pendingByMsgId.remove(msgId, pending);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -222,17 +333,115 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
|||||||
};
|
};
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** Broker return for an unroutable {@code mandatory} publish — arrives BEFORE its confirm (CB-528). */
|
||||||
|
private void onReturn(Return r) {
|
||||||
|
String msgId = r.getProperties() == null ? null : r.getProperties().getMessageId();
|
||||||
|
Pending pending = msgId == null ? null : pendingByMsgId.get(msgId);
|
||||||
|
if (pending != null) {
|
||||||
|
pending.returned = true;
|
||||||
|
} else {
|
||||||
|
log.warn("AMQP return for reply {} (routingKey={}, {} {}) with no matching in-flight publish"
|
||||||
|
+ " — already resolved by a prior confirm", msgId, r.getRoutingKey(), r.getReplyCode(),
|
||||||
|
r.getReplyText());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private void onAck(long seq, boolean multiple) {
|
||||||
|
resolveConfirm(seq, multiple, true);
|
||||||
|
}
|
||||||
|
|
||||||
|
private void onNack(long seq, boolean multiple) {
|
||||||
|
resolveConfirm(seq, multiple, false);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Resolve every pending publish covered by this confirm (a single seq, or — {@code multiple} —
|
||||||
|
* every seq up to and including it). Checks {@link Pending#returned} at confirm time: since the
|
||||||
|
* broker's return for an unroutable message always precedes its confirm, an ack that arrives after
|
||||||
|
* a return means "confirmed but never routed", not "durably queued".
|
||||||
|
*/
|
||||||
|
private void resolveConfirm(long seq, boolean multiple, boolean ack) {
|
||||||
|
NavigableMap<Long, Pending> covered = multiple
|
||||||
|
? pendingBySeq.headMap(seq, true)
|
||||||
|
: pendingBySeq.subMap(seq, true, seq, true);
|
||||||
|
for (var it = covered.entrySet().iterator(); it.hasNext(); ) {
|
||||||
|
Pending pending = it.next().getValue();
|
||||||
|
it.remove();
|
||||||
|
pendingByMsgId.remove(pending.msgId, pending);
|
||||||
|
if (ack && !pending.returned) {
|
||||||
|
pending.confirmed.complete(null);
|
||||||
|
} else if (ack) {
|
||||||
|
pending.confirmed.completeExceptionally(new IllegalStateException(
|
||||||
|
"reply " + pending.msgId + " was returned as unroutable (queue not declared/owned)"));
|
||||||
|
} else {
|
||||||
|
pending.confirmed.completeExceptionally(new IllegalStateException(
|
||||||
|
"broker nacked publish of reply " + pending.msgId));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Fail every publish still awaiting its confirm — their sequence numbers are stale after recovery.
|
||||||
|
* Guarded by {@link #publishChannelLock}, the same lock {@link #publish} holds while it takes its
|
||||||
|
* sequence number and registers its {@link Pending}: without it, a {@link #publish} that starts
|
||||||
|
* after the connection has already recovered (so it publishes — and will be confirmed — on the
|
||||||
|
* <em>new</em> channel) can register between this sweep's iteration and its clear, and this sweep
|
||||||
|
* then fails a publish that actually succeeded. {@link #publish} only holds the lock for the
|
||||||
|
* seq/map-put/{@code basicPublish} — it awaits the confirm outside it — so this sweep can only ever
|
||||||
|
* wait for an in-flight {@code basicPublish} call to return, never for a broker round trip. No
|
||||||
|
* deadlock.
|
||||||
|
*
|
||||||
|
* <p>Package-private (rather than {@code private}) only so the unit test can drive it directly
|
||||||
|
* against a concurrent {@link #publish} without a live broker reconnect.
|
||||||
|
*/
|
||||||
|
void failPendingPublishesOnRecovery() {
|
||||||
|
synchronized (publishChannelLock) {
|
||||||
|
for (var it = pendingBySeq.entrySet().iterator(); it.hasNext(); ) {
|
||||||
|
Pending pending = it.next().getValue();
|
||||||
|
it.remove();
|
||||||
|
pendingByMsgId.remove(pending.msgId, pending);
|
||||||
|
pending.confirmed.completeExceptionally(new IllegalStateException(
|
||||||
|
"AMQP connection recovered mid-publish; confirm status of reply " + pending.msgId
|
||||||
|
+ " is unknown"));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Fail every publish still awaiting its confirm with a clear, immediate error instead of leaving it
|
||||||
|
* to time out after {@link #CONFIRM_TIMEOUT_MS} once the channels are closed underneath it. Guarded
|
||||||
|
* by {@link #publishChannelLock} for the same reason as {@link #failPendingPublishesOnRecovery}.
|
||||||
|
*/
|
||||||
|
private void failPendingPublishesOnClose() {
|
||||||
|
synchronized (publishChannelLock) {
|
||||||
|
for (var it = pendingBySeq.entrySet().iterator(); it.hasNext(); ) {
|
||||||
|
Pending pending = it.next().getValue();
|
||||||
|
it.remove();
|
||||||
|
pendingByMsgId.remove(pending.msgId, pending);
|
||||||
|
pending.confirmed.completeExceptionally(new IllegalStateException(
|
||||||
|
"AMQP reply inbox closed while publish of reply " + pending.msgId
|
||||||
|
+ " was still awaiting its confirm"));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
private static String queueName(String target) {
|
private static String queueName(String target) {
|
||||||
return QUEUE_PREFIX + target + QUEUE_SUFFIX;
|
return QUEUE_PREFIX + target + QUEUE_SUFFIX;
|
||||||
}
|
}
|
||||||
|
|
||||||
@Override
|
@Override
|
||||||
public void close() {
|
public void close() {
|
||||||
|
failPendingPublishesOnClose();
|
||||||
try {
|
try {
|
||||||
channel.close();
|
channel.close();
|
||||||
} catch (Exception e) {
|
} catch (Exception e) {
|
||||||
log.debug("AMQP channel close: {}", e.toString());
|
log.debug("AMQP channel close: {}", e.toString());
|
||||||
}
|
}
|
||||||
|
try {
|
||||||
|
publishChannel.close();
|
||||||
|
} catch (Exception e) {
|
||||||
|
log.debug("AMQP publish channel close: {}", e.toString());
|
||||||
|
}
|
||||||
try {
|
try {
|
||||||
connection.close();
|
connection.close();
|
||||||
} catch (Exception e) {
|
} catch (Exception e) {
|
||||||
|
|||||||
@@ -164,18 +164,30 @@ public final class MessageService {
|
|||||||
|
|
||||||
/** An in-flight or finished async delegation, keyed by its ticket. */
|
/** An in-flight or finished async delegation, keyed by its ticket. */
|
||||||
private static final class Task {
|
private static final class Task {
|
||||||
|
private final String ticket;
|
||||||
private final String target;
|
private final String target;
|
||||||
private final CompletableFuture<Reply> future = new CompletableFuture<>();
|
private final CompletableFuture<Reply> future = new CompletableFuture<>();
|
||||||
private final long createdNanos;
|
private final long createdNanos;
|
||||||
private volatile Reply question;
|
private volatile Reply question;
|
||||||
private volatile String turnId;
|
private volatile String turnId;
|
||||||
|
|
||||||
private Task(String target, long createdNanos) {
|
private Task(String ticket, String target, long createdNanos) {
|
||||||
|
this.ticket = ticket;
|
||||||
this.target = target;
|
this.target = target;
|
||||||
this.createdNanos = createdNanos;
|
this.createdNanos = createdNanos;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* A worker session's currently-open {@code bridge_ask} question, surfaced so {@code bridge_status}
|
||||||
|
* can show it without the caller needing the ticket first (CB-582). Only covers async
|
||||||
|
* (fire-and-poll) delegations, which track the question on their {@link Task}; a blocking
|
||||||
|
* ({@code wait:true}) send already hands the question straight back to its own caller, so there is
|
||||||
|
* nothing hidden left for {@code bridge_status} to surface in that case.
|
||||||
|
*/
|
||||||
|
public record PendingAsk(String ticket, String question, String turnId) {
|
||||||
|
}
|
||||||
|
|
||||||
private final AgentControl agents;
|
private final AgentControl agents;
|
||||||
private final Injector injector;
|
private final Injector injector;
|
||||||
private final Rendezvous rendezvous;
|
private final Rendezvous rendezvous;
|
||||||
@@ -200,9 +212,11 @@ public final class MessageService {
|
|||||||
* Create with an explicit {@link ReplyInbox} and optional {@link ReplyPushLoop}.
|
* Create with an explicit {@link ReplyInbox} and optional {@link ReplyPushLoop}.
|
||||||
*
|
*
|
||||||
* @param pushLoop nullable — when non-null, the push loop is notified on the no-waiter reply
|
* @param pushLoop nullable — when non-null, the push loop is notified on the no-waiter reply
|
||||||
* branch ({@link #reply}) so it can nudge the primary to drain the inbox, and
|
* branch ({@link #reply}) so it can nudge the primary to drain the inbox,
|
||||||
* (CB-588) whenever an async ticket started by {@link #sendAsync} reaches a
|
* (CB-588) whenever an async ticket started by {@link #sendAsync} reaches a
|
||||||
* terminal phase, and whenever {@link #poll} hands a terminal ticket to its caller
|
* terminal phase, whenever {@link #poll} hands a terminal ticket to its caller,
|
||||||
|
* and (CB-582) whenever an async ticket's worker pauses mid-turn in
|
||||||
|
* {@code bridge_ask} or that pause ends (answered or lapsed)
|
||||||
*/
|
*/
|
||||||
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
|
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
|
||||||
ReplyInbox inbox, ReplyPushLoop pushLoop) {
|
ReplyInbox inbox, ReplyPushLoop pushLoop) {
|
||||||
@@ -495,6 +509,15 @@ public final class MessageService {
|
|||||||
rendezvous.closeAsk(ticket.turnId());
|
rendezvous.closeAsk(ticket.turnId());
|
||||||
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
|
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
|
||||||
}
|
}
|
||||||
|
// CB-582: the question just became visible via bridge_poll (Phase.ASKING) for an async
|
||||||
|
// (wait:false) delegation — nudge the lead's own pane the same way a terminal ticket does
|
||||||
|
// (CB-588), since the lead's normal poll cadence is minutes away and the reverse-rendezvous
|
||||||
|
// window (~55s, see BridgeMcp/BridgedApp) is far shorter. A blocking (wait:true) send has
|
||||||
|
// no Task and gets the question directly in its own reply, so task == null there — nothing
|
||||||
|
// to nudge.
|
||||||
|
if (task != null && pushLoop != null) {
|
||||||
|
pushLoop.onQuestionOpened(task.ticket, workerSession, ticket.turnId(), question);
|
||||||
|
}
|
||||||
}
|
}
|
||||||
try {
|
try {
|
||||||
String answer = ticket.answer().get(timeoutMillis, TimeUnit.MILLISECONDS);
|
String answer = ticket.answer().get(timeoutMillis, TimeUnit.MILLISECONDS);
|
||||||
@@ -513,6 +536,14 @@ public final class MessageService {
|
|||||||
// Only the fresh owner tears down the shared turn; a duplicate must leave it open.
|
// Only the fresh owner tears down the shared turn; a duplicate must leave it open.
|
||||||
if (ticket.fresh()) {
|
if (ticket.fresh()) {
|
||||||
rendezvous.closeAsk(ticket.turnId());
|
rendezvous.closeAsk(ticket.turnId());
|
||||||
|
// CB-582: tear the push loop's copy down at the same point, not only on the three
|
||||||
|
// paths that call clearAsyncQuestion. The answer future can complete exceptionally
|
||||||
|
// (ExecutionException) or the thread be interrupted, and both leave this method by
|
||||||
|
// throwing — the question would stay pending forever, keep being named in nudges
|
||||||
|
// until its own cap, and never be removed from the map. Already-closed is a no-op.
|
||||||
|
if (pushLoop != null) {
|
||||||
|
pushLoop.questionClosed(ticket.turnId());
|
||||||
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -589,7 +620,7 @@ public final class MessageService {
|
|||||||
*/
|
*/
|
||||||
public String sendAsync(String target, String content, Runnable onAccepted) {
|
public String sendAsync(String target, String content, Runnable onAccepted) {
|
||||||
String ticket = "task-" + ticketSeq.incrementAndGet();
|
String ticket = "task-" + ticketSeq.incrementAndGet();
|
||||||
Task task = new Task(target, nowNanos.getAsLong());
|
Task task = new Task(ticket, target, nowNanos.getAsLong());
|
||||||
tasks.put(ticket, task);
|
tasks.put(ticket, task);
|
||||||
if (pushLoop != null) {
|
if (pushLoop != null) {
|
||||||
// CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a
|
// CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a
|
||||||
@@ -715,6 +746,12 @@ public final class MessageService {
|
|||||||
|
|
||||||
/** Clear an answered or lapsed question, but only when it matches the ticket's current turn. */
|
/** Clear an answered or lapsed question, but only when it matches the ticket's current turn. */
|
||||||
private void clearAsyncQuestion(String turnId, boolean forgetTurn) {
|
private void clearAsyncQuestion(String turnId, boolean forgetTurn) {
|
||||||
|
// CB-582: tell the push loop first — like ticketCollected, a removal for a turnId it never
|
||||||
|
// nudged about (or already dropped) is a harmless no-op, so this is safe to call unconditionally
|
||||||
|
// rather than threading the guard below through it.
|
||||||
|
if (pushLoop != null) {
|
||||||
|
pushLoop.questionClosed(turnId);
|
||||||
|
}
|
||||||
Task task = asyncTasksByTurn.get(turnId);
|
Task task = asyncTasksByTurn.get(turnId);
|
||||||
if (task != null && turnId.equals(task.turnId)) {
|
if (task != null && turnId.equals(task.turnId)) {
|
||||||
task.question = null;
|
task.question = null;
|
||||||
@@ -746,6 +783,23 @@ public final class MessageService {
|
|||||||
return asyncTasksByTurn.values().stream().anyMatch(task -> target.equals(task.target));
|
return asyncTasksByTurn.values().stream().anyMatch(task -> target.equals(task.target));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The question {@code workerSession} is currently paused on via {@code bridge_ask}, if any
|
||||||
|
* (CB-582) — {@code bridge_status} uses this to show a pending question without the caller
|
||||||
|
* needing the ticket. {@code null} when the session has no open async question (including a
|
||||||
|
* session mid a <em>blocking</em> {@code bridge_ask}, which has no {@link Task} to look up — see
|
||||||
|
* {@link PendingAsk}).
|
||||||
|
*/
|
||||||
|
public PendingAsk pendingAsk(String workerSession) {
|
||||||
|
for (Task task : tasks.values()) {
|
||||||
|
Reply q = task.question;
|
||||||
|
if (q != null && workerSession.equals(task.target)) {
|
||||||
|
return new PendingAsk(task.ticket, q.text(), q.turnId());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
|
||||||
/** Release the async executor. */
|
/** Release the async executor. */
|
||||||
public void close() {
|
public void close() {
|
||||||
asyncExecutor.shutdown();
|
asyncExecutor.shutdown();
|
||||||
|
|||||||
@@ -8,6 +8,8 @@ import dev.ltms.bridged.metrics.Metrics;
|
|||||||
import org.slf4j.Logger;
|
import org.slf4j.Logger;
|
||||||
import org.slf4j.LoggerFactory;
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
|
import java.util.ArrayList;
|
||||||
|
import java.util.HashSet;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Set;
|
import java.util.Set;
|
||||||
import java.util.concurrent.ConcurrentHashMap;
|
import java.util.concurrent.ConcurrentHashMap;
|
||||||
@@ -16,37 +18,56 @@ import java.util.concurrent.TimeUnit;
|
|||||||
import java.util.stream.Collectors;
|
import java.util.stream.Collectors;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Mechanism (b) of CB-307: a dedicated, status-gated push loop that nudges the primary's own
|
* A status-gated push loop that nudges a lead's own herdr pane when it has uncollected work
|
||||||
* herdr pane when a worker reply lands with no live {@code bridge_send} to resolve it.
|
* waiting: a worker reply queued with no live {@code bridge_send} to resolve it (CB-307), an
|
||||||
|
* async delegation ticket ({@code bridge_send(wait:false)}) that reached a terminal phase
|
||||||
|
* (CB-588), or an async ticket's worker pausing mid-turn in {@code bridge_ask} to await an answer
|
||||||
|
* (CB-582).
|
||||||
*
|
*
|
||||||
* <p>The loop is triggered by {@link #onReplyQueued(String)} (called from
|
* <p><strong>CB-590: one schedule per lead.</strong> All three kinds of work are triggered
|
||||||
* {@link MessageService#reply} after the durable inbox publish). It checks four conditions
|
* through their own entry point — {@link #onReplyQueued(String)},
|
||||||
* at each tick via {@link #decide(String, int)}, then either injects a drain nudge,
|
* {@link #onTicketTerminal(String, String, boolean)}, and
|
||||||
* waits for the primary to become injectable, or stops reminding.
|
* {@link #onQuestionOpened(String, String, String, String)} — but each resolves the lead that
|
||||||
|
* should be nudged and coalesces onto a single per-lead reminder schedule, tracked in
|
||||||
|
* {@link #activeLeads}. Earlier this was two independent schedules (one keyed by worker target
|
||||||
|
* for replies, one keyed by lead for tickets) that could both decide to inject into the same pane
|
||||||
|
* in the same window — a race, not routine behaviour, but the expensive kind: it interrupts the
|
||||||
|
* lead's live turn twice. Collapsing to one schedule per lead makes that structurally impossible:
|
||||||
|
* at most one scheduled tick chain is ever live for a given lead (guarded by {@link #activeLeads}'
|
||||||
|
* compare-and-set), so at most one {@code agents.send} to that lead's pane is ever in flight.
|
||||||
*
|
*
|
||||||
* <p>Bounded: at most {@link #maxReminders} nudges per target, with a configurable backoff
|
* <p>Each tick examines <em>everything</em> pending for that lead — reply targets whose inbox
|
||||||
* between them. The reply is never lost — the durable inbox is the backstop.
|
* still holds an unacked message ({@link #pendingReplies}), tickets not yet collected
|
||||||
*
|
* ({@link #pendingTickets}), and open questions not yet answered or lapsed
|
||||||
* <p><strong>CB-588 ticket nudges.</strong> {@link #onTicketTerminal(String, String, boolean)} is a
|
* ({@link #pendingQuestions}) — and sends at most one combined nudge per tick
|
||||||
* second, independent entry point for an async delegation ticket ({@code bridge_send(wait:false)})
|
* ({@link #injectNudge(String, int, int, int)}). Work that arrives while the lead is busy is
|
||||||
* reaching a terminal phase. That path takes the rendezvous fast path in {@link MessageService#reply}
|
* never lost: it is re-read fresh on every tick until the lead is injectable or its own reminder
|
||||||
* and never reaches {@link #onReplyQueued}, so without this the lead's own charter — prefer
|
* cap ({@link #maxReminders}) is reached — each source spends from its own budget, so one source
|
||||||
* {@code wait:false} for anything non-trivial — was exactly the mode this loop failed to cover. It
|
* exhausting its cap does not stop nudges about the others (post-CB-590 regression fix; see
|
||||||
* reuses the same status gating, bounded/backoff reminders, and metrics, keyed by the nudge-receiving
|
* {@link #decide}) — whichever the durable inbox / pending set doesn't already answer via
|
||||||
* lead terminal rather than the worker target so several tickets finishing together coalesce into one
|
* {@code STOP}.
|
||||||
* nudge. The two entry points do not interact: {@link #onReplyQueued} / {@link #decide} / their nudge
|
|
||||||
* text and bound are unchanged.
|
|
||||||
*/
|
*/
|
||||||
public final class ReplyPushLoop {
|
public final class ReplyPushLoop {
|
||||||
|
|
||||||
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
|
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
|
||||||
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
|
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
|
||||||
/** CB-588: singular form, one uncollected ticket. */
|
/** Coalesced form, several uncollected replies for the same lead. */
|
||||||
|
static final String REPLIES_NUDGE_FORMAT =
|
||||||
|
"%d workers returned replies — run bridge_poll(target=...) for each to collect them: %s";
|
||||||
|
/** Singular form, one uncollected ticket. */
|
||||||
static final String TICKET_NUDGE_FORMAT =
|
static final String TICKET_NUDGE_FORMAT =
|
||||||
"Ticket %s finished%s — run bridge_poll(ticket=%s) to collect it";
|
"Ticket %s finished%s — run bridge_poll(ticket=%s) to collect it";
|
||||||
/** CB-588: coalesced form, several uncollected tickets for the same lead. */
|
/** Coalesced form, several uncollected tickets for the same lead. */
|
||||||
static final String TICKETS_NUDGE_FORMAT =
|
static final String TICKETS_NUDGE_FORMAT =
|
||||||
"%d tickets finished%s — run bridge_poll(ticket=...) for each to collect them: %s";
|
"%d tickets finished%s — run bridge_poll(ticket=...) for each to collect them: %s";
|
||||||
|
/** Singular form, one worker paused mid-turn in bridge_ask (CB-582) — names the answer call directly. */
|
||||||
|
static final String QUESTION_NUDGE_FORMAT =
|
||||||
|
"Worker %s asked a question (ticket %s) — answer it with bridge_send(turnId=\"%s\", "
|
||||||
|
+ "content=...) to resume its turn:\n%s";
|
||||||
|
/** Coalesced form, several open questions for the same lead. */
|
||||||
|
static final String QUESTIONS_NUDGE_FORMAT =
|
||||||
|
"%d workers are paused on a question — run bridge_poll(ticket=...) for each, then answer "
|
||||||
|
+ "with bridge_send(turnId=..., content=...): %s";
|
||||||
|
|
||||||
private final PrimaryRegistry primaryRegistry;
|
private final PrimaryRegistry primaryRegistry;
|
||||||
private final AgentControl agents;
|
private final AgentControl agents;
|
||||||
@@ -56,11 +77,23 @@ public final class ReplyPushLoop {
|
|||||||
private final long backoffMs;
|
private final long backoffMs;
|
||||||
private final Metrics metrics; // CB-512: nullable — no registry in unit tests
|
private final Metrics metrics; // CB-512: nullable — no registry in unit tests
|
||||||
|
|
||||||
/** Track targets that have an active schedule. */
|
/**
|
||||||
private final ConcurrentHashMap<String, Boolean> activeTargets = new ConcurrentHashMap<>();
|
* Worker targets with a reply queued, keyed by target. Each entry carries its own nudge
|
||||||
/** CB-588: tickets that have gone terminal but not yet been polled, keyed by ticket. */
|
* count (CB-598) rather than sharing one counter per lead per source: a target's count only
|
||||||
|
* ever reflects nudges that actually named that target, so a target that joins while the
|
||||||
|
* schedule is already deep into another target's reminders still reads as fresh.
|
||||||
|
*/
|
||||||
|
private final ConcurrentHashMap<String, ReplyEntry> pendingReplies = new ConcurrentHashMap<>();
|
||||||
|
/** Tickets that have gone terminal but not yet been polled, keyed by ticket. */
|
||||||
private final ConcurrentHashMap<String, PendingTicket> pendingTickets = new ConcurrentHashMap<>();
|
private final ConcurrentHashMap<String, PendingTicket> pendingTickets = new ConcurrentHashMap<>();
|
||||||
/** CB-588: leads with an active ticket-reminder schedule. */
|
/**
|
||||||
|
* Open {@code bridge_ask} questions not yet answered or lapsed, keyed by {@code turnId}
|
||||||
|
* (CB-582). A question's own nudge count is tracked the same per-item way as
|
||||||
|
* {@link #pendingTickets} (CB-598): a fresh question keeps its source eligible regardless of
|
||||||
|
* how depleted an older, still-open question's count is.
|
||||||
|
*/
|
||||||
|
private final ConcurrentHashMap<String, PendingQuestion> pendingQuestions = new ConcurrentHashMap<>();
|
||||||
|
/** CB-590: leads with an active combined reminder schedule (replies and/or tickets and/or questions). */
|
||||||
private final ConcurrentHashMap<String, Boolean> activeLeads = new ConcurrentHashMap<>();
|
private final ConcurrentHashMap<String, Boolean> activeLeads = new ConcurrentHashMap<>();
|
||||||
|
|
||||||
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
||||||
@@ -89,178 +122,169 @@ public final class ReplyPushLoop {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- decision logic (package-private for unit-testing) -------------------------------------
|
// --- pending-work lookups (package-private for unit-testing) -------------------------------
|
||||||
|
|
||||||
/** The action the loop should take for a target at the given reminder count. */
|
/** The action the loop should take for a lead at the given reminder count. */
|
||||||
enum Action { INJECT, WAIT_BUSY, STOP }
|
enum Action { INJECT, WAIT_BUSY, STOP }
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Pure decision function: examine the current state and return what the loop should do.
|
* Reply targets still pending for {@code lead} — registered via {@link #onReplyQueued} and
|
||||||
*
|
* whose inbox still holds an unacked message. A target whose inbox has since drained (acked,
|
||||||
* @param target the worker session (target terminal id)
|
* or collected via a live {@code bridge_send} rendezvous instead) is dropped from
|
||||||
* @param reminderCount how many nudges have been sent so far for this target
|
* {@link #pendingReplies} here rather than lingering forever; there is no explicit "reply
|
||||||
* @return the action the caller should take
|
* collected" callback the way {@link #ticketCollected} exists for tickets, so the inbox itself
|
||||||
|
* is the only signal.
|
||||||
*/
|
*/
|
||||||
Action decide(String target, int reminderCount) {
|
private Set<String> pendingReplyTargetsFor(String lead) {
|
||||||
// CB-532: the destination is per-delegation — the lead that sent this worker its work, not
|
Set<String> result = new HashSet<>();
|
||||||
// "the primary". With two leads orchestrating one fleet the singular question has no right
|
for (var entry : pendingReplies.entrySet()) {
|
||||||
// answer, and answering it anyway interrupted whichever lead happened to call bridge_send
|
String target = entry.getKey();
|
||||||
// first with results it never asked for.
|
ReplyEntry owning = entry.getValue();
|
||||||
var nudgeTarget = primaryRegistry.nudgeTargetFor(target);
|
if (!lead.equals(owning.lead())) continue;
|
||||||
if (nudgeTarget.isEmpty()) {
|
if (inbox.peek(target).isEmpty()) {
|
||||||
log.debug("push: no lead is known to be waiting on {}, stopping reminder", target);
|
pendingReplies.remove(target, owning);
|
||||||
return Action.STOP;
|
continue;
|
||||||
}
|
|
||||||
if (inbox.peek(target).isEmpty()) {
|
|
||||||
log.debug("push: inbox empty for {}, stopping reminder", target);
|
|
||||||
return Action.STOP;
|
|
||||||
}
|
|
||||||
if (reminderCount >= maxReminders) {
|
|
||||||
log.debug("push: reminder cap ({}) reached for {}, stopping", maxReminders, target);
|
|
||||||
countNudge("exhausted");
|
|
||||||
return Action.STOP;
|
|
||||||
}
|
|
||||||
String leadTerminal = nudgeTarget.get();
|
|
||||||
AgentStatus status;
|
|
||||||
try {
|
|
||||||
status = agents.status(leadTerminal);
|
|
||||||
} catch (RuntimeException e) {
|
|
||||||
log.debug("push: status check failed for lead {}, will retry", leadTerminal, e);
|
|
||||||
return Action.WAIT_BUSY;
|
|
||||||
}
|
|
||||||
if (status.injectable()) {
|
|
||||||
return Action.INJECT;
|
|
||||||
}
|
|
||||||
log.debug("push: lead {} is {} (not injectable), waiting", leadTerminal, status);
|
|
||||||
return Action.WAIT_BUSY;
|
|
||||||
}
|
|
||||||
|
|
||||||
// --- public entrypoint ---------------------------------------------------------------------
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Called when a reply is queued for {@code target}. Idempotent per target: a second call while
|
|
||||||
* a schedule is active is a no-op. The schedule nudges the primary, then schedules a follow-up
|
|
||||||
* check (reminder on backoff, or re-check on WAIT_BUSY), until the inbox is empty or the cap
|
|
||||||
* is reached.
|
|
||||||
*/
|
|
||||||
public void onReplyQueued(String target) {
|
|
||||||
if (activeTargets.putIfAbsent(target, Boolean.TRUE) != null) {
|
|
||||||
log.debug("push: already active for {}, ignoring duplicate trigger", target);
|
|
||||||
return; // already scheduled
|
|
||||||
}
|
|
||||||
log.debug("push: starting reminder loop for {}", target);
|
|
||||||
scheduleNext(target, 0);
|
|
||||||
}
|
|
||||||
|
|
||||||
/** Execute one loop tick — called on the scheduler thread. */
|
|
||||||
private void tick(String target, int reminderCount) {
|
|
||||||
var action = decide(target, reminderCount);
|
|
||||||
switch (action) {
|
|
||||||
case INJECT -> {
|
|
||||||
injectNudge(target, reminderCount);
|
|
||||||
scheduleNext(target, reminderCount + 1);
|
|
||||||
}
|
|
||||||
// Re-check after the configured backoff; the primary may become injectable soon.
|
|
||||||
case WAIT_BUSY -> scheduleNext(target, reminderCount);
|
|
||||||
case STOP -> {
|
|
||||||
activeTargets.remove(target);
|
|
||||||
log.debug("push: reminder loop ended for {}", target);
|
|
||||||
}
|
}
|
||||||
|
result.add(target);
|
||||||
}
|
}
|
||||||
|
return result;
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Send the nudge and log the event. */
|
/** A pending reply target: which lead to nudge, and how many nudges have named it so far. */
|
||||||
private void injectNudge(String target, int reminderCount) {
|
private record ReplyEntry(String lead, int nudgeCount) {
|
||||||
// Re-read rather than threading it down from decide(): the delegating lead can change
|
|
||||||
// between the decision and the injection, and the nudge should follow the current one.
|
|
||||||
var lead = primaryRegistry.nudgeTargetFor(target);
|
|
||||||
if (lead.isEmpty()) {
|
|
||||||
log.debug("push: lead for {} disappeared before the nudge could be sent", target);
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
String leadTerminal = lead.get();
|
|
||||||
String nudge = NUDGE_FORMAT.formatted(target, target);
|
|
||||||
try {
|
|
||||||
agents.send(leadTerminal, nudge);
|
|
||||||
log.debug("push: nudge {}/{} sent to lead {} for target {}",
|
|
||||||
reminderCount + 1, maxReminders, leadTerminal, target);
|
|
||||||
countNudge("delivered");
|
|
||||||
} catch (RuntimeException e) {
|
|
||||||
log.warn("push: failed to nudge lead {} for target {} (reminder {}/{}): {}",
|
|
||||||
leadTerminal, target, reminderCount + 1, maxReminders, e.toString());
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
/** Schedule the next tick on the scheduler thread pool. */
|
|
||||||
private void scheduleNext(String target, int nextReminderCount) {
|
|
||||||
scheduler.schedule(() -> tick(target, nextReminderCount), backoffMs, TimeUnit.MILLISECONDS);
|
|
||||||
}
|
|
||||||
|
|
||||||
// --- CB-588: async ticket terminal nudges ---------------------------------------------------
|
|
||||||
|
|
||||||
/** A ticket awaiting collection: which lead to nudge, and whether it ended in failure. */
|
|
||||||
private record PendingTicket(String ticket, String lead, boolean failed) {
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Called when an async delegation ticket ({@code bridge_send(wait:false)}, CB-107) reaches a
|
* A ticket awaiting collection: which lead to nudge, whether it ended in failure, and how
|
||||||
* terminal phase — DONE or a failure. Unlike {@link #onReplyQueued}, which nudges about the
|
* many nudges have named it so far (CB-598 — tracked per ticket, not per lead per source).
|
||||||
* durable-inbox no-waiter path, this covers the path {@code MessageService.reply} takes when a
|
|
||||||
* fire-and-poll send's own rendezvous waiter resolves the reply directly: that path returns
|
|
||||||
* before {@link #onReplyQueued} is ever called, so without this entry point a ticket finishing
|
|
||||||
* that way never nudged anyone (CB-588 / gitea #72).
|
|
||||||
*
|
|
||||||
* <p>Idempotent per lead: several tickets going terminal for the same lead while its schedule is
|
|
||||||
* already active coalesce onto that schedule's next tick rather than firing a nudge each.
|
|
||||||
*
|
|
||||||
* @param ticket the ticket to nudge about
|
|
||||||
* @param target the worker session the ticket was sent to — resolves which lead delegated it
|
|
||||||
* @param failed whether the ticket ended in a failure phase rather than {@code DONE}
|
|
||||||
*/
|
*/
|
||||||
public void onTicketTerminal(String ticket, String target, boolean failed) {
|
private record PendingTicket(String ticket, String lead, boolean failed, int nudgeCount) {
|
||||||
var lead = primaryRegistry.nudgeTargetFor(target);
|
|
||||||
if (lead.isEmpty()) {
|
|
||||||
log.debug("push: no lead is known to be waiting on ticket {} (target {}), skipping nudge",
|
|
||||||
ticket, target);
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
pendingTickets.put(ticket, new PendingTicket(ticket, lead.get(), failed));
|
|
||||||
if (activeLeads.putIfAbsent(lead.get(), Boolean.TRUE) != null) {
|
|
||||||
log.debug("push: ticket reminder loop already active for lead {}, {} coalesced in",
|
|
||||||
lead.get(), ticket);
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
log.debug("push: starting ticket reminder loop for lead {}", lead.get());
|
|
||||||
scheduleTicketTick(lead.get(), 0);
|
|
||||||
}
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Called when a ticket's terminal state has been collected via {@code bridge_poll}. Removes it
|
|
||||||
* from the pending set so a scheduled tick — and any nudge it sends — never names a ticket the
|
|
||||||
* lead already has (CB-588 acceptance #5). A ticket that was never pending (unknown ticket, or
|
|
||||||
* one nudged with no push loop configured) is a no-op.
|
|
||||||
*/
|
|
||||||
public void ticketCollected(String ticket) {
|
|
||||||
pendingTickets.remove(ticket);
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Tickets still pending for {@code lead}, snapshotted fresh for one tick. */
|
/** Tickets still pending for {@code lead}, snapshotted fresh for one tick. */
|
||||||
private List<PendingTicket> pendingFor(String lead) {
|
private List<PendingTicket> pendingTicketsFor(String lead) {
|
||||||
return pendingTickets.values().stream().filter(t -> lead.equals(t.lead())).toList();
|
return pendingTickets.values().stream().filter(t -> lead.equals(t.lead())).toList();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** Ticket ids still pending for {@code lead} — a plain snapshot for race comparison. */
|
||||||
|
private Set<String> pendingTicketIdsFor(String lead) {
|
||||||
|
return pendingTicketsFor(lead).stream().map(PendingTicket::ticket)
|
||||||
|
.collect(Collectors.toUnmodifiableSet());
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Pure decision function for ticket nudges, mirroring {@link #decide(String, int)} but keyed by
|
* An open question awaiting the lead's answer: which ticket it belongs to, which worker asked,
|
||||||
* the nudge-receiving lead terminal rather than the worker session — several tickets from
|
* which lead to nudge, the question text, and how many nudges have named it so far (CB-598 —
|
||||||
* different workers delegated by the same lead coalesce onto it.
|
* tracked per question, not per lead per source).
|
||||||
*/
|
*/
|
||||||
Action decideTickets(String lead, int reminderCount) {
|
private record PendingQuestion(String turnId, String ticket, String target, String lead,
|
||||||
if (pendingFor(lead).isEmpty()) {
|
String question, int nudgeCount) {
|
||||||
log.debug("push: nothing pending for lead {}, stopping ticket reminder", lead);
|
}
|
||||||
|
|
||||||
|
/** Questions still open for {@code lead}, snapshotted fresh for one tick. */
|
||||||
|
private List<PendingQuestion> pendingQuestionsFor(String lead) {
|
||||||
|
return pendingQuestions.values().stream().filter(q -> lead.equals(q.lead())).toList();
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Question turnIds still open for {@code lead} — a plain snapshot for race comparison. */
|
||||||
|
private Set<String> pendingQuestionTurnIdsFor(String lead) {
|
||||||
|
return pendingQuestionsFor(lead).stream().map(PendingQuestion::turnId)
|
||||||
|
.collect(Collectors.toUnmodifiableSet());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The reply-source reminder count {@link #decide} should see for {@code lead} on this tick:
|
||||||
|
* the <em>minimum</em> nudge count among the reply targets currently pending for it (CB-598).
|
||||||
|
*
|
||||||
|
* <p>Before this, the count passed to {@code decide} was a single counter carried forward
|
||||||
|
* across scheduled ticks ({@code scheduleNext(lead, count + 1, ...)}), incremented whenever
|
||||||
|
* the source had <em>any</em> pending work — not tied to which target that work was. A target
|
||||||
|
* that joined while an older target's count was already near the cap inherited that count on
|
||||||
|
* its very next tick, even though no nudge had ever named it. Taking the minimum over what is
|
||||||
|
* actually pending now means a fresh target (count 0) keeps the source eligible regardless of
|
||||||
|
* how many times an older, still-undrained target has already been nudged; that older target
|
||||||
|
* keeps riding along in the combined nudge text without spending any more of its own budget
|
||||||
|
* (see {@link #bumpNudgeCounts}). Returns 0 when nothing is pending — {@link #decide} never
|
||||||
|
* consults the count in that case, since {@code hasReplyWork} is false.
|
||||||
|
*/
|
||||||
|
private int minReplyNudgeCountFor(String lead) {
|
||||||
|
int min = Integer.MAX_VALUE;
|
||||||
|
for (String target : pendingReplyTargetsFor(lead)) {
|
||||||
|
ReplyEntry entry = pendingReplies.get(target);
|
||||||
|
if (entry != null) {
|
||||||
|
min = Math.min(min, entry.nudgeCount());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return min == Integer.MAX_VALUE ? 0 : min;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** As {@link #minReplyNudgeCountFor}, for the ticket source. */
|
||||||
|
private int minTicketNudgeCountFor(String lead) {
|
||||||
|
int min = Integer.MAX_VALUE;
|
||||||
|
for (PendingTicket ticket : pendingTicketsFor(lead)) {
|
||||||
|
min = Math.min(min, ticket.nudgeCount());
|
||||||
|
}
|
||||||
|
return min == Integer.MAX_VALUE ? 0 : min;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** As {@link #minReplyNudgeCountFor}, for the question source (CB-582). */
|
||||||
|
private int minQuestionNudgeCountFor(String lead) {
|
||||||
|
int min = Integer.MAX_VALUE;
|
||||||
|
for (PendingQuestion q : pendingQuestionsFor(lead)) {
|
||||||
|
min = Math.min(min, q.nudgeCount());
|
||||||
|
}
|
||||||
|
return min == Integer.MAX_VALUE ? 0 : min;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Pure decision function: examine everything pending for {@code lead} — reply targets and
|
||||||
|
* tickets alike — and return what the loop should do.
|
||||||
|
*
|
||||||
|
* <p><strong>CB-590-fix: one schedule, two budgets.</strong> The single per-lead schedule
|
||||||
|
* (CB-590) still ticks once for both sources, but each source is capped independently —
|
||||||
|
* {@code replyReminderCount} against a reply target still pending, {@code ticketReminderCount}
|
||||||
|
* against a ticket still pending. A busy reply stream that exhausts its own cap must not stop
|
||||||
|
* the loop from nudging about a ticket that still has budget left, and vice versa: either
|
||||||
|
* source being eligible (has pending work AND is under its own cap) is enough for
|
||||||
|
* {@link Action#INJECT}. Only when neither source has eligible work does the loop
|
||||||
|
* {@link Action#STOP}.
|
||||||
|
*
|
||||||
|
* <p><strong>CB-598: the counts are per-item, not per-tick.</strong> {@link #tick} no longer
|
||||||
|
* carries these counts forward across scheduled calls — it recomputes them fresh every tick via
|
||||||
|
* {@link #minReplyNudgeCountFor} / {@link #minTicketNudgeCountFor}, so this function itself did
|
||||||
|
* not need to change; only what its caller feeds it did.
|
||||||
|
*
|
||||||
|
* @param lead the lead terminal to nudge
|
||||||
|
* @param replyReminderCount the lowest nudge count among reply targets pending for this lead
|
||||||
|
* @param ticketReminderCount the lowest nudge count among tickets pending for this lead
|
||||||
|
* @return the action the caller should take
|
||||||
|
*/
|
||||||
|
Action decide(String lead, int replyReminderCount, int ticketReminderCount) {
|
||||||
|
return decide(lead, replyReminderCount, ticketReminderCount, minQuestionNudgeCountFor(lead));
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* As {@link #decide(String, int, int)}, with the question source (CB-582) folded in on the
|
||||||
|
* same footing as replies and tickets: its own eligibility (has open questions AND under its
|
||||||
|
* own {@link #maxReminders} budget) is enough on its own to {@link Action#INJECT}, exactly like
|
||||||
|
* the other two.
|
||||||
|
*
|
||||||
|
* @param questionReminderCount the lowest nudge count among questions open for this lead
|
||||||
|
*/
|
||||||
|
Action decide(String lead, int replyReminderCount, int ticketReminderCount, int questionReminderCount) {
|
||||||
|
boolean hasReplyWork = !pendingReplyTargetsFor(lead).isEmpty();
|
||||||
|
boolean hasTicketWork = !pendingTicketIdsFor(lead).isEmpty();
|
||||||
|
boolean hasQuestionWork = !pendingQuestionTurnIdsFor(lead).isEmpty();
|
||||||
|
if (!hasReplyWork && !hasTicketWork && !hasQuestionWork) {
|
||||||
|
log.debug("push: nothing pending for lead {}, stopping reminder", lead);
|
||||||
return Action.STOP;
|
return Action.STOP;
|
||||||
}
|
}
|
||||||
if (reminderCount >= maxReminders) {
|
boolean replyEligible = hasReplyWork && replyReminderCount < maxReminders;
|
||||||
log.debug("push: ticket reminder cap ({}) reached for lead {}, stopping", maxReminders, lead);
|
boolean ticketEligible = hasTicketWork && ticketReminderCount < maxReminders;
|
||||||
|
boolean questionEligible = hasQuestionWork && questionReminderCount < maxReminders;
|
||||||
|
if (!replyEligible && !ticketEligible && !questionEligible) {
|
||||||
|
log.debug("push: reminder cap ({}) reached for lead {} on every source with pending work, stopping",
|
||||||
|
maxReminders, lead);
|
||||||
countNudge("exhausted");
|
countNudge("exhausted");
|
||||||
return Action.STOP;
|
return Action.STOP;
|
||||||
}
|
}
|
||||||
@@ -278,95 +302,285 @@ public final class ReplyPushLoop {
|
|||||||
return Action.WAIT_BUSY;
|
return Action.WAIT_BUSY;
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Execute one ticket-loop tick — called on the scheduler thread. */
|
// --- public entrypoints ----------------------------------------------------------------------
|
||||||
private void ticketTick(String lead, int reminderCount) {
|
|
||||||
Set<String> pendingBefore = pendingIdsFor(lead);
|
|
||||||
var action = decideTickets(lead, reminderCount);
|
|
||||||
switch (action) {
|
|
||||||
case INJECT -> {
|
|
||||||
injectTicketNudge(lead, reminderCount);
|
|
||||||
scheduleTicketTick(lead, reminderCount + 1);
|
|
||||||
}
|
|
||||||
case WAIT_BUSY -> scheduleTicketTick(lead, reminderCount);
|
|
||||||
case STOP -> stopOrRestartTicketLoop(lead, pendingBefore);
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
/** Ticket IDs pending for {@code lead} right now, as a plain snapshot for race comparison. */
|
/**
|
||||||
private Set<String> pendingIdsFor(String lead) {
|
* Called when a reply is queued for {@code target}. Resolves the lead delegating to
|
||||||
return pendingFor(lead).stream().map(PendingTicket::ticket).collect(Collectors.toUnmodifiableSet());
|
* {@code target} (CB-532) and coalesces onto that lead's single reminder schedule — starting
|
||||||
|
* one if none is active, joining an already-active one otherwise. A no-op if no lead is known
|
||||||
|
* to be waiting on {@code target}: there is nobody to nudge yet, and the durable inbox is the
|
||||||
|
* backstop until a lead is recorded.
|
||||||
|
*/
|
||||||
|
public void onReplyQueued(String target) {
|
||||||
|
var lead = primaryRegistry.nudgeTargetFor(target);
|
||||||
|
if (lead.isEmpty()) {
|
||||||
|
log.debug("push: no lead is known to be waiting on {}, skipping reminder", target);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
pendingReplies.compute(target, (t, existing) ->
|
||||||
|
new ReplyEntry(lead.get(), existing == null ? 0 : existing.nudgeCount()));
|
||||||
|
startOrCoalesce(lead.get());
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Release {@code lead}'s active-schedule slot, then restart it only if a ticket landed that
|
* Called when an async delegation ticket ({@code bridge_send(wait:false)}, CB-107) reaches a
|
||||||
* {@code pendingBefore} — the snapshot taken just before this tick's decision — did not already
|
* terminal phase — DONE or a failure. Unlike {@link #onReplyQueued}, which nudges about the
|
||||||
* account for. {@code onTicketTerminal} reads {@code activeLeads} to decide whether to coalesce
|
* durable-inbox no-waiter path, this covers the path {@code MessageService.reply} takes when a
|
||||||
* onto an existing schedule or start one, so a ticket that lands between {@code decideTickets}
|
* fire-and-poll send's own rendezvous waiter resolves the reply directly: that path returns
|
||||||
* returning {@link Action#STOP} and this removal running sees the (soon-to-be-stale) slot as
|
* before {@link #onReplyQueued} is ever called, so without this entry point a ticket finishing
|
||||||
* occupied, coalesces onto a schedule that is about to die, and gets no nudge scheduled at all —
|
* that way never nudged anyone (CB-588 / gitea #72).
|
||||||
* a lost nudge, the exact failure CB-588 exists to remove (found in review, gitea PR #73).
|
|
||||||
*
|
*
|
||||||
* <p>Restarting on ANY non-empty {@code pendingFor(lead)} would be wrong: when STOP is reached
|
* <p>Resolves the delegating lead the same way {@link #onReplyQueued} does and coalesces onto
|
||||||
* because the reminder cap was hit rather than the backlog draining, the same never-collected
|
* the same per-lead schedule (CB-590) — several tickets, or a ticket and a reply, finishing
|
||||||
* ticket is expected to still be sitting there — that is the cap doing its job — and restarting
|
* for the same lead while its schedule is already active all ride the existing schedule's next
|
||||||
* would nudge about it forever, defeating the bound (the original CB-307 bounded-reminder
|
* tick rather than firing a nudge each.
|
||||||
* guarantee, carried into CB-588 by acceptance criterion #7 — this exact regression showed up as
|
*
|
||||||
* two existing tests failing once a naive "any pending ticket restarts" version of this fix went
|
* @param ticket the ticket to nudge about
|
||||||
* in: {@code successfulTicketNudgeIncrementsDelivered} and {@code ticketNudgesSendUpToCapThenStop}).
|
* @param target the worker session the ticket was sent to — resolves which lead delegated it
|
||||||
* Diffing the current pending set against {@code pendingBefore} tells the two cases apart: a
|
* @param failed whether the ticket ended in a failure phase rather than {@code DONE}
|
||||||
* ticket present before this tick's decision is stale backlog, not a race; only a ticket absent
|
*/
|
||||||
* from {@code pendingBefore} can only have arrived during the decision-to-release window, which is
|
public void onTicketTerminal(String ticket, String target, boolean failed) {
|
||||||
* exactly the race this method closes.
|
var lead = primaryRegistry.nudgeTargetFor(target);
|
||||||
|
if (lead.isEmpty()) {
|
||||||
|
log.debug("push: no lead is known to be waiting on ticket {} (target {}), skipping nudge",
|
||||||
|
ticket, target);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
pendingTickets.compute(ticket, (id, existing) ->
|
||||||
|
new PendingTicket(ticket, lead.get(), failed, existing == null ? 0 : existing.nudgeCount()));
|
||||||
|
startOrCoalesce(lead.get());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Called when a ticket's terminal state has been collected via {@code bridge_poll}. Removes it
|
||||||
|
* from the pending set so a scheduled tick — and any nudge it sends — never names a ticket the
|
||||||
|
* lead already has. A ticket that was never pending (unknown ticket, or one nudged with no push
|
||||||
|
* loop configured) is a no-op.
|
||||||
|
*/
|
||||||
|
public void ticketCollected(String ticket) {
|
||||||
|
pendingTickets.remove(ticket);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Called when an async ticket's worker pauses mid-turn in {@code bridge_ask} (CB-582): the
|
||||||
|
* question is now visible via {@code bridge_poll} (Phase.ASKING), but the reverse-rendezvous
|
||||||
|
* window it opened with (~55s default, see {@code BridgeMcp}/{@code BridgedApp}) is far shorter
|
||||||
|
* than a lead's normal minutes-long poll cadence — exactly the gap this closes. Resolves the
|
||||||
|
* delegating lead the same way {@link #onTicketTerminal} does and coalesces onto the same
|
||||||
|
* per-lead schedule (CB-590).
|
||||||
|
*
|
||||||
|
* @param ticket the async ticket the question belongs to (for {@code bridge_poll})
|
||||||
|
* @param target the worker session that asked
|
||||||
|
* @param turnId correlation id the lead answers with ({@code bridge_send turnId=...})
|
||||||
|
* @param question the question text
|
||||||
|
*/
|
||||||
|
public void onQuestionOpened(String ticket, String target, String turnId, String question) {
|
||||||
|
var lead = primaryRegistry.nudgeTargetFor(target);
|
||||||
|
if (lead.isEmpty()) {
|
||||||
|
log.debug("push: no lead is known to be waiting on {}'s question (turnId {}), skipping nudge",
|
||||||
|
target, turnId);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
pendingQuestions.put(turnId, new PendingQuestion(turnId, ticket, target, lead.get(), question, 0));
|
||||||
|
startOrCoalesce(lead.get());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Called when a worker's {@code bridge_ask} resolves — answered or lapsed unanswered — so a
|
||||||
|
* scheduled tick never nudges about a question the lead already handled. A {@code turnId} that
|
||||||
|
* was never pending (never nudged, or already closed) is a no-op.
|
||||||
|
*/
|
||||||
|
public void questionClosed(String turnId) {
|
||||||
|
pendingQuestions.remove(turnId);
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- the schedule ----------------------------------------------------------------------------
|
||||||
|
|
||||||
|
/** Start a reminder schedule for {@code lead}, or join the one already running. */
|
||||||
|
private void startOrCoalesce(String lead) {
|
||||||
|
if (activeLeads.putIfAbsent(lead, Boolean.TRUE) != null) {
|
||||||
|
log.debug("push: reminder loop already active for lead {}, work coalesced in", lead);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
log.debug("push: starting reminder loop for lead {}", lead);
|
||||||
|
scheduleNext(lead);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Execute one loop tick — called on the scheduler thread (or directly by a test; package-private
|
||||||
|
* for the same reason as {@link #stopOrRestart}).
|
||||||
|
*
|
||||||
|
* <p><strong>CB-598.</strong> The reminder counts fed into {@link #decide} are recomputed fresh
|
||||||
|
* every tick from what is actually pending right now ({@link #minReplyNudgeCountFor} /
|
||||||
|
* {@link #minTicketNudgeCountFor}), rather than carried forward as running counters across
|
||||||
|
* scheduled calls. A counter carried forward has no memory of which item it was counting for:
|
||||||
|
* a target or ticket that joined mid-backoff — after the previous tick fired but before this one
|
||||||
|
* did — is already sitting in {@code repliesBefore} / {@code ticketsBefore} below by the time this
|
||||||
|
* tick takes its snapshot, indistinguishable at that point from backlog the cap is meant to
|
||||||
|
* silence. Recomputing from the per-item counts fixes that: a newly-joined item's own count is
|
||||||
|
* still 0, so it keeps its source eligible regardless of how depleted an older, still-undrained
|
||||||
|
* item's count is.
|
||||||
|
*/
|
||||||
|
void tick(String lead) {
|
||||||
|
Set<String> repliesBefore = pendingReplyTargetsFor(lead);
|
||||||
|
Set<String> ticketsBefore = pendingTicketIdsFor(lead);
|
||||||
|
Set<String> questionsBefore = pendingQuestionTurnIdsFor(lead);
|
||||||
|
int replyReminderCount = minReplyNudgeCountFor(lead);
|
||||||
|
int ticketReminderCount = minTicketNudgeCountFor(lead);
|
||||||
|
int questionReminderCount = minQuestionNudgeCountFor(lead);
|
||||||
|
var action = decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount);
|
||||||
|
switch (action) {
|
||||||
|
case INJECT -> {
|
||||||
|
injectNudge(lead, replyReminderCount, ticketReminderCount, questionReminderCount);
|
||||||
|
scheduleNext(lead);
|
||||||
|
}
|
||||||
|
// Re-check after the configured backoff; the lead may become injectable soon.
|
||||||
|
case WAIT_BUSY -> scheduleNext(lead);
|
||||||
|
case STOP -> stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Release {@code lead}'s active-schedule slot, then restart it only if work landed that
|
||||||
|
* {@code repliesBefore} / {@code ticketsBefore} — the snapshots taken just before this tick's
|
||||||
|
* decision — did not already account for. {@link #onReplyQueued} / {@link #onTicketTerminal}
|
||||||
|
* read {@link #activeLeads} to decide whether to coalesce onto an existing schedule or start
|
||||||
|
* one, so work that lands between {@link #decide} returning {@link Action#STOP} and this
|
||||||
|
* removal running sees the (soon-to-be-stale) slot as occupied, coalesces onto a schedule that
|
||||||
|
* is about to die, and gets no nudge scheduled at all — a lost nudge, exactly what CB-588 (and
|
||||||
|
* now CB-590) exist to remove (originally found in review, gitea PR #73, for the ticket-only
|
||||||
|
* loop; carried forward here for the unified one).
|
||||||
|
*
|
||||||
|
* <p>Restarting on ANY non-empty pending set would be wrong: when STOP is reached because the
|
||||||
|
* reminder cap was hit rather than the backlog draining, the same never-collected work is
|
||||||
|
* expected to still be sitting there — that is the cap doing its job — and restarting would
|
||||||
|
* nudge about it forever, defeating the bound. Diffing the current pending sets against the
|
||||||
|
* "before" snapshots tells the two cases apart: an item present before this tick's decision is
|
||||||
|
* stale backlog, not a race; only an item absent from the "before" snapshot can only have
|
||||||
|
* arrived during the decision-to-release window, which is exactly the race this method closes.
|
||||||
*
|
*
|
||||||
* <p>Package-private so a test can drive the interleaving directly rather than trying to force a
|
* <p>Package-private so a test can drive the interleaving directly rather than trying to force a
|
||||||
* genuine thread race: pass the exact {@code pendingBefore} snapshot a race requires (or does
|
* genuine thread race.
|
||||||
* not) and call this to prove the recheck responds correctly either way.
|
|
||||||
*
|
*
|
||||||
* <p>Terminates rather than spinning: this method restarts the schedule at most once per call, and
|
* <p>Terminates rather than spinning: this method restarts the schedule at most once per call,
|
||||||
* a fresh {@link #onTicketTerminal} racing the recheck below still terminates in one of two ways —
|
* and a fresh {@link #onReplyQueued} / {@link #onTicketTerminal} racing the recheck below still
|
||||||
* either it observes the slot already vacated (by the {@code activeLeads.remove} above, which
|
* terminates in one of two ways — either it observes the slot already vacated (by the
|
||||||
* happens-before this recheck in program order) and claims it itself, or it lands first and this
|
* {@code activeLeads.remove} above, which happens-before this recheck in program order) and
|
||||||
* recheck then observes its ticket in {@code pendingTickets} and reclaims the slot instead. Exactly
|
* claims it itself, or it lands first and this recheck then observes its work in
|
||||||
* one side always wins; neither can miss the other, so this never loops on its own account.
|
* {@link #pendingReplies} / {@link #pendingTickets} and reclaims the slot instead. Exactly one
|
||||||
|
* side always wins; neither can miss the other, so this never loops on its own account.
|
||||||
*/
|
*/
|
||||||
void stopOrRestartTicketLoop(String lead, Set<String> pendingBefore) {
|
void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore) {
|
||||||
activeLeads.remove(lead);
|
stopOrRestart(lead, repliesBefore, ticketsBefore, pendingQuestionTurnIdsFor(lead));
|
||||||
boolean ticketRacedIn = pendingFor(lead).stream().anyMatch(t -> !pendingBefore.contains(t.ticket()));
|
|
||||||
if (ticketRacedIn && activeLeads.putIfAbsent(lead, Boolean.TRUE) == null) {
|
|
||||||
log.debug("push: a ticket for lead {} raced the reminder loop's stop — restarting", lead);
|
|
||||||
scheduleTicketTick(lead, 0);
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
log.debug("push: ticket reminder loop ended for lead {}", lead);
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Send the coalesced ticket nudge and log the event. */
|
/**
|
||||||
private void injectTicketNudge(String lead, int reminderCount) {
|
* As {@link #stopOrRestart(String, Set, Set)}, with the question source's (CB-582) own "before"
|
||||||
// Re-read rather than threading it down from decideTickets(): a ticket can be collected (or
|
* snapshot folded into the same race check: a question that raced in during the
|
||||||
// another can arrive) between the decision and the injection.
|
* decision-to-release window reclaims the schedule slot exactly like a raced-in reply or ticket.
|
||||||
List<PendingTicket> pending = pendingFor(lead);
|
*/
|
||||||
if (pending.isEmpty()) {
|
void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore,
|
||||||
log.debug("push: pending tickets for lead {} drained before the nudge could be sent", lead);
|
Set<String> questionsBefore) {
|
||||||
|
activeLeads.remove(lead);
|
||||||
|
boolean racedIn = pendingReplyTargetsFor(lead).stream().anyMatch(t -> !repliesBefore.contains(t))
|
||||||
|
|| pendingTicketIdsFor(lead).stream().anyMatch(t -> !ticketsBefore.contains(t))
|
||||||
|
|| pendingQuestionTurnIdsFor(lead).stream().anyMatch(t -> !questionsBefore.contains(t));
|
||||||
|
if (racedIn && activeLeads.putIfAbsent(lead, Boolean.TRUE) == null) {
|
||||||
|
log.debug("push: new work for lead {} raced the reminder loop's stop — restarting", lead);
|
||||||
|
scheduleNext(lead);
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
String nudge = formatTicketsNudge(pending);
|
log.debug("push: reminder loop ended for lead {}", lead);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Send one combined nudge covering everything currently pending for {@code lead}. */
|
||||||
|
private void injectNudge(String lead, int replyReminderCount, int ticketReminderCount,
|
||||||
|
int questionReminderCount) {
|
||||||
|
// Re-read rather than threading it down from decide(): a reply can drain, a ticket be
|
||||||
|
// collected, or a question be answered (or another arrive), between the decision and the
|
||||||
|
// injection.
|
||||||
|
Set<String> replyTargets = pendingReplyTargetsFor(lead);
|
||||||
|
List<PendingTicket> tickets = pendingTicketsFor(lead);
|
||||||
|
List<PendingQuestion> questions = pendingQuestionsFor(lead);
|
||||||
|
if (replyTargets.isEmpty() && tickets.isEmpty() && questions.isEmpty()) {
|
||||||
|
log.debug("push: pending work for lead {} drained before the nudge could be sent", lead);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
String nudge = formatNudge(replyTargets, tickets, questions);
|
||||||
try {
|
try {
|
||||||
agents.send(lead, nudge);
|
agents.send(lead, nudge);
|
||||||
log.debug("push: ticket nudge {}/{} sent to lead {} for {} ticket(s)",
|
log.debug("push: nudge sent to lead {} (reply {}/{}, ticket {}/{}, question {}/{}; "
|
||||||
reminderCount + 1, maxReminders, lead, pending.size());
|
+ "{} reply target(s), {} ticket(s), {} question(s))",
|
||||||
|
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
|
||||||
|
questionReminderCount + 1, maxReminders,
|
||||||
|
replyTargets.size(), tickets.size(), questions.size());
|
||||||
countNudge("delivered");
|
countNudge("delivered");
|
||||||
} catch (RuntimeException e) {
|
} catch (RuntimeException e) {
|
||||||
log.warn("push: failed to nudge lead {} for {} ticket(s) (reminder {}/{}): {}",
|
log.warn("push: failed to nudge lead {} (reply {}/{}, ticket {}/{}, question {}/{}): {}",
|
||||||
lead, pending.size(), reminderCount + 1, maxReminders, e.toString());
|
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
|
||||||
|
questionReminderCount + 1, maxReminders, e.toString());
|
||||||
|
}
|
||||||
|
// Bump every item actually named in this nudge, not just whatever the shared source-level
|
||||||
|
// eligibility used to gate (CB-598) — each item's own count is what the next tick's
|
||||||
|
// minReplyNudgeCountFor / minTicketNudgeCountFor / minQuestionNudgeCountFor will read. An
|
||||||
|
// item already at or over the cap keeps riding along in the text (still pending, still
|
||||||
|
// named) but its extra bumps here are inert: decide() already treats it as ineligible once
|
||||||
|
// its count reaches maxReminders.
|
||||||
|
bumpNudgeCounts(replyTargets, tickets, questions);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Record that every one of these items was just named in a sent (or attempted) nudge. */
|
||||||
|
private void bumpNudgeCounts(Set<String> replyTargets, List<PendingTicket> tickets,
|
||||||
|
List<PendingQuestion> questions) {
|
||||||
|
for (String target : replyTargets) {
|
||||||
|
pendingReplies.computeIfPresent(target, (t, e) -> new ReplyEntry(e.lead(), e.nudgeCount() + 1));
|
||||||
|
}
|
||||||
|
for (PendingTicket ticket : tickets) {
|
||||||
|
pendingTickets.computeIfPresent(ticket.ticket(),
|
||||||
|
(id, e) -> new PendingTicket(e.ticket(), e.lead(), e.failed(), e.nudgeCount() + 1));
|
||||||
|
}
|
||||||
|
for (PendingQuestion question : questions) {
|
||||||
|
pendingQuestions.computeIfPresent(question.turnId(), (id, e) ->
|
||||||
|
new PendingQuestion(e.turnId(), e.ticket(), e.target(), e.lead(), e.question(),
|
||||||
|
e.nudgeCount() + 1));
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Schedule the next ticket-loop tick on the scheduler thread pool. */
|
/** Schedule the next tick on the scheduler thread pool. */
|
||||||
private void scheduleTicketTick(String lead, int nextReminderCount) {
|
private void scheduleNext(String lead) {
|
||||||
scheduler.schedule(() -> ticketTick(lead, nextReminderCount), backoffMs, TimeUnit.MILLISECONDS);
|
scheduler.schedule(() -> tick(lead),
|
||||||
|
backoffMs, TimeUnit.MILLISECONDS);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Render one or several pending tickets as a single nudge line. */
|
// --- nudge formatting ------------------------------------------------------------------------
|
||||||
|
|
||||||
|
/** Render everything pending for one lead as a single nudge line. */
|
||||||
|
private static String formatNudge(Set<String> replyTargets, List<PendingTicket> tickets,
|
||||||
|
List<PendingQuestion> questions) {
|
||||||
|
List<String> parts = new ArrayList<>();
|
||||||
|
if (!replyTargets.isEmpty()) {
|
||||||
|
parts.add(formatRepliesNudge(replyTargets));
|
||||||
|
}
|
||||||
|
if (!tickets.isEmpty()) {
|
||||||
|
parts.add(formatTicketsNudge(tickets));
|
||||||
|
}
|
||||||
|
if (!questions.isEmpty()) {
|
||||||
|
parts.add(formatQuestionsNudge(questions));
|
||||||
|
}
|
||||||
|
return String.join(" | ", parts);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Render one or several pending reply targets. */
|
||||||
|
private static String formatRepliesNudge(Set<String> targets) {
|
||||||
|
if (targets.size() == 1) {
|
||||||
|
String target = targets.iterator().next();
|
||||||
|
return NUDGE_FORMAT.formatted(target, target);
|
||||||
|
}
|
||||||
|
String ids = String.join(", ", targets);
|
||||||
|
return REPLIES_NUDGE_FORMAT.formatted(targets.size(), ids);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Render one or several pending tickets. */
|
||||||
private static String formatTicketsNudge(List<PendingTicket> pending) {
|
private static String formatTicketsNudge(List<PendingTicket> pending) {
|
||||||
if (pending.size() == 1) {
|
if (pending.size() == 1) {
|
||||||
PendingTicket t = pending.get(0);
|
PendingTicket t = pending.get(0);
|
||||||
@@ -380,27 +594,39 @@ public final class ReplyPushLoop {
|
|||||||
return TICKETS_NUDGE_FORMAT.formatted(pending.size(), failedNote, ids);
|
return TICKETS_NUDGE_FORMAT.formatted(pending.size(), failedNote, ids);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** Render one or several open questions (CB-582). */
|
||||||
|
private static String formatQuestionsNudge(List<PendingQuestion> pending) {
|
||||||
|
if (pending.size() == 1) {
|
||||||
|
PendingQuestion q = pending.get(0);
|
||||||
|
return QUESTION_NUDGE_FORMAT.formatted(q.target(), q.ticket(), q.turnId(), q.question());
|
||||||
|
}
|
||||||
|
String ids = pending.stream()
|
||||||
|
.map(q -> q.ticket() + " (turnId=" + q.turnId() + ")")
|
||||||
|
.collect(Collectors.joining(", "));
|
||||||
|
return QUESTIONS_NUDGE_FORMAT.formatted(pending.size(), ids);
|
||||||
|
}
|
||||||
|
|
||||||
// --- lifecycle -----------------------------------------------------------------------------
|
// --- lifecycle -----------------------------------------------------------------------------
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Whether any reminder loop is currently active for some target (CB-551). The idle-lead heartbeat
|
* Whether any reminder loop is currently active for some lead (CB-551). The idle-lead heartbeat
|
||||||
* uses this to stand aside: while the push loop is actively nudging the lead, a concurrent
|
* uses this to stand aside: while the push loop is actively nudging a lead, a concurrent
|
||||||
* heartbeat injection would start a second competing turn in the same pane — racing loops multiply
|
* heartbeat injection would start a second competing turn in the same pane — racing loops
|
||||||
* turns and context burn. "Active" means a schedule exists in {@link #activeTargets} or
|
* multiply turns and context burn. "Active" means a schedule exists in {@link #activeLeads},
|
||||||
* {@link #activeLeads} (CB-588 ticket nudges are a second source of pane injections the heartbeat
|
* which now covers reply-queued (CB-307), ticket-terminal (CB-588), and question-open (CB-582)
|
||||||
* must equally stand aside for); the sets are bounded by what has been triggered, not by any
|
* work (CB-590) — bounded by what has been triggered, not by any persistent state.
|
||||||
* persistent state.
|
|
||||||
*/
|
*/
|
||||||
public boolean isActive() {
|
public boolean isActive() {
|
||||||
return !activeTargets.isEmpty() || !activeLeads.isEmpty();
|
return !activeLeads.isEmpty();
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Shut down the scheduler. Outstanding reminders are cancelled. */
|
/** Shut down the scheduler. Outstanding reminders are cancelled. */
|
||||||
public void stop() {
|
public void stop() {
|
||||||
scheduler.shutdownNow();
|
scheduler.shutdownNow();
|
||||||
activeTargets.clear();
|
|
||||||
activeLeads.clear();
|
activeLeads.clear();
|
||||||
|
pendingReplies.clear();
|
||||||
pendingTickets.clear();
|
pendingTickets.clear();
|
||||||
|
pendingQuestions.clear();
|
||||||
}
|
}
|
||||||
|
|
||||||
/** @see #stop() */
|
/** @see #stop() */
|
||||||
|
|||||||
@@ -13,6 +13,7 @@ import dev.ltms.bridged.herdr.HerdrClient;
|
|||||||
import dev.ltms.bridged.herdr.HerdrException;
|
import dev.ltms.bridged.herdr.HerdrException;
|
||||||
import dev.ltms.bridged.inject.MemberPresence;
|
import dev.ltms.bridged.inject.MemberPresence;
|
||||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||||
|
import dev.ltms.bridged.placement.PlacementException;
|
||||||
import dev.ltms.bridged.msg.MessageService;
|
import dev.ltms.bridged.msg.MessageService;
|
||||||
import dev.ltms.bridged.session.SessionManager;
|
import dev.ltms.bridged.session.SessionManager;
|
||||||
import dev.ltms.bridged.peer.MemberRole;
|
import dev.ltms.bridged.peer.MemberRole;
|
||||||
@@ -234,7 +235,14 @@ public final class BridgedApp {
|
|||||||
List<Map<String, Object>> out = sessions.roster().stream()
|
List<Map<String, Object>> out = sessions.roster().stream()
|
||||||
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
|
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
|
||||||
.toList();
|
.toList();
|
||||||
ctx.status(200).json(Map.of("workers", out));
|
Map<String, Object> body = new LinkedHashMap<>();
|
||||||
|
body.put("workers", out);
|
||||||
|
// CB-586: operator visibility for the refs/wip snapshot store without shelling into the
|
||||||
|
// repo — how many snapshot refs exist and roughly what they cost. Present only once a
|
||||||
|
// worktree session has established the repo, so a never-snapshotted fleet reports nothing.
|
||||||
|
sessions.wipRefs().ifPresent(st -> body.put("wipRefs",
|
||||||
|
Map.of("count", st.count(), "costBytes", st.costBytes())));
|
||||||
|
ctx.status(200).json(body);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** The configured worker profiles and which one a no-argument spawn uses. */
|
/** The configured worker profiles and which one a no-argument spawn uses. */
|
||||||
@@ -292,6 +300,11 @@ public final class BridgedApp {
|
|||||||
ctx.status(201).json(view(member));
|
ctx.status(201).json(view(member));
|
||||||
} catch (GuardException e) {
|
} catch (GuardException e) {
|
||||||
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
|
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
|
||||||
|
} catch (PlacementException e) {
|
||||||
|
// CB-599: no candidate had capacity (maxLoad, quarantine, or all-exhausted) — a benign,
|
||||||
|
// likely-transient refusal, distinct from "profile does not exist" below. 503: the
|
||||||
|
// request was valid and will likely succeed later.
|
||||||
|
ctx.status(503).json(Map.of("error", "no_capacity", "detail", e.getMessage()));
|
||||||
} catch (IllegalArgumentException e) {
|
} catch (IllegalArgumentException e) {
|
||||||
ctx.status(400).json(Map.of("error", "unknown_profile", "detail", e.getMessage()));
|
ctx.status(400).json(Map.of("error", "unknown_profile", "detail", e.getMessage()));
|
||||||
} catch (PeerUnreachableException e) {
|
} catch (PeerUnreachableException e) {
|
||||||
@@ -503,10 +516,20 @@ public final class BridgedApp {
|
|||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
try {
|
try {
|
||||||
ctx.status(200).json(Map.of(
|
Map<String, Object> body = new LinkedHashMap<>();
|
||||||
"sessionId", id,
|
body.put("sessionId", id);
|
||||||
"status", messages.status(id).name().toLowerCase(),
|
body.put("status", messages.status(id).name().toLowerCase());
|
||||||
"ready", presence.isPresent(id)));
|
body.put("ready", presence.isPresent(id));
|
||||||
|
// CB-582: a worker paused mid-turn in an async bridge_ask is otherwise invisible to a
|
||||||
|
// status poll — surface the open question and how to answer it, same as bridge_poll's
|
||||||
|
// Phase.ASKING view.
|
||||||
|
MessageService.PendingAsk ask = messages.pendingAsk(id);
|
||||||
|
if (ask != null) {
|
||||||
|
body.put("question", ask.question());
|
||||||
|
body.put("turnId", ask.turnId());
|
||||||
|
body.put("ticket", ask.ticket());
|
||||||
|
}
|
||||||
|
ctx.status(200).json(body);
|
||||||
} catch (HerdrException e) {
|
} catch (HerdrException e) {
|
||||||
herdrError(ctx, e);
|
herdrError(ctx, e);
|
||||||
}
|
}
|
||||||
@@ -532,6 +555,11 @@ public final class BridgedApp {
|
|||||||
if (v.detail() != null) {
|
if (v.detail() != null) {
|
||||||
body.put("detail", v.detail());
|
body.put("detail", v.detail());
|
||||||
}
|
}
|
||||||
|
// CB-582: Phase.ASKING carries the question in v.reply() (handled above) and its answer-
|
||||||
|
// correlation id here — a REST caller polling this ticket otherwise has no way to answer it.
|
||||||
|
if (v.turnId() != null) {
|
||||||
|
body.put("turnId", v.turnId());
|
||||||
|
}
|
||||||
ctx.status(200).json(body);
|
ctx.status(200).json(body);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -12,9 +12,12 @@ import java.nio.file.Files;
|
|||||||
import java.nio.file.Path;
|
import java.nio.file.Path;
|
||||||
import java.nio.file.StandardCopyOption;
|
import java.nio.file.StandardCopyOption;
|
||||||
import java.security.SecureRandom;
|
import java.security.SecureRandom;
|
||||||
|
import java.util.ArrayList;
|
||||||
|
import java.util.HashSet;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
import java.util.Optional;
|
import java.util.Optional;
|
||||||
|
import java.util.Set;
|
||||||
import java.util.concurrent.TimeUnit;
|
import java.util.concurrent.TimeUnit;
|
||||||
import java.util.concurrent.atomic.AtomicLong;
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
import java.util.stream.Collectors;
|
import java.util.stream.Collectors;
|
||||||
@@ -299,6 +302,129 @@ public final class GitWorktrees implements Worktrees {
|
|||||||
return index;
|
return index;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* One {@code refs/wip/<branch>} snapshot ref as read by {@link #listWipRefs}: its full ref name,
|
||||||
|
* the snapshot commit's sha, and that commit's committer time in unix millis (the age of the
|
||||||
|
* snapshot — a snapshot is written once and never rewritten, so the commit date is the ref's).
|
||||||
|
*/
|
||||||
|
private record WipRef(String refName, String sha, long committerMillis) {
|
||||||
|
String branch() {
|
||||||
|
return refName.substring("refs/wip/".length());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public WipRefStats wipRefs(String repoRoot) {
|
||||||
|
List<WipRef> refs = listWipRefs(repoRoot);
|
||||||
|
long costBytes = 0;
|
||||||
|
for (WipRef ref : refs) {
|
||||||
|
costBytes += treeSize(repoRoot, ref.sha());
|
||||||
|
}
|
||||||
|
return new WipRefStats(refs.size(), costBytes);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
|
||||||
|
// The rule is documented on Worktrees#pruneWipRefs: delete only a snapshot whose tree
|
||||||
|
// content is already reachable from main AND that is older than minAgeMillis. Reachability
|
||||||
|
// is the floor that keeps a worker's last copy; the age floor keeps a just-written snapshot
|
||||||
|
// from being swept while a lead may still be looking at it.
|
||||||
|
List<WipRef> refs = listWipRefs(repoRoot);
|
||||||
|
if (refs.isEmpty()) {
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
long nowMillis = System.currentTimeMillis();
|
||||||
|
// Resolve what main carries once per sweep, not once per ref.
|
||||||
|
Set<String> mainObjects = reachableObjectsFromMain(repoRoot);
|
||||||
|
int deleted = 0;
|
||||||
|
for (WipRef ref : refs) {
|
||||||
|
long ageMillis = nowMillis - ref.committerMillis();
|
||||||
|
if (ageMillis <= minAgeMillis) {
|
||||||
|
continue; // too recent — never swept, even if it looks recoverable (CB-586)
|
||||||
|
}
|
||||||
|
String tree = exec("git", "-C", repoRoot, "rev-parse", ref.sha() + "^{tree}").trim();
|
||||||
|
if (!mainObjects.contains(tree)) {
|
||||||
|
// Last copy of the snapshot's content — the worker's work exists nowhere else.
|
||||||
|
// Never delete automatically (CB-586 criterion 2).
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
exec("git", "-C", repoRoot, "update-ref", "-d", ref.refName());
|
||||||
|
deleted++;
|
||||||
|
log.info("pruned snapshot ref refs/wip/{} commit={} (age {}h): its tree is already "
|
||||||
|
+ "reachable from main, so the work is preserved; recover from reflog via "
|
||||||
|
+ "git update-ref refs/wip/{} {}",
|
||||||
|
ref.branch(), ref.sha(), TimeUnit.MILLISECONDS.toHours(ageMillis),
|
||||||
|
ref.branch(), ref.sha());
|
||||||
|
}
|
||||||
|
return deleted;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Every {@code refs/wip/*} ref (see {@link WipRef}). The committer date is read as a unix
|
||||||
|
* count of seconds and converted to millis. {@code %00} (NUL) separates the fields because a
|
||||||
|
* branch name may contain spaces.
|
||||||
|
*/
|
||||||
|
private List<WipRef> listWipRefs(String repoRoot) {
|
||||||
|
String out = exec("git", "-C", repoRoot, "for-each-ref",
|
||||||
|
"--format=%(refname)%00%(objectname)%00%(committerdate:unix)", "refs/wip/");
|
||||||
|
List<WipRef> refs = new ArrayList<>();
|
||||||
|
for (String line : out.split("\\R")) {
|
||||||
|
if (line.isBlank()) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
String[] parts = line.split("\u0000", -1);
|
||||||
|
if (parts.length == 3 && !parts[1].isBlank()) {
|
||||||
|
refs.add(new WipRef(parts[0], parts[1], Long.parseLong(parts[2]) * 1000L));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return refs;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The set of object shas reachable from {@code main}, or an empty set when {@code main} cannot
|
||||||
|
* be resolved. An empty set is the safe direction: the retention sweep then concludes nothing
|
||||||
|
* is recoverable, so it deletes nothing — a repo with no {@code main} must never cause a
|
||||||
|
* worker's last copy of a snapshot to be dropped on a reachability misreading.
|
||||||
|
*/
|
||||||
|
private Set<String> reachableObjectsFromMain(String repoRoot) {
|
||||||
|
if (exitCode("git", "-C", repoRoot, "rev-parse", "--verify", "main") != 0) {
|
||||||
|
log.debug("refs/wip retention: no 'main' ref in {} — treating nothing as reachable", repoRoot);
|
||||||
|
return Set.of();
|
||||||
|
}
|
||||||
|
String out = exec("git", "-C", repoRoot, "rev-list", "--objects", "main");
|
||||||
|
Set<String> objects = new HashSet<>();
|
||||||
|
for (String line : out.split("\\R")) {
|
||||||
|
if (line.isBlank()) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
int sp = line.indexOf(' ');
|
||||||
|
objects.add(sp < 0 ? line : line.substring(0, sp));
|
||||||
|
}
|
||||||
|
return objects;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Approximate cost of a snapshot: the sum of every blob's size in its committed tree. */
|
||||||
|
private long treeSize(String repoRoot, String sha) {
|
||||||
|
String out = exec("git", "-C", repoRoot, "ls-tree", "-r", "-l", sha);
|
||||||
|
long total = 0;
|
||||||
|
for (String line : out.split("\\R")) {
|
||||||
|
if (line.isBlank()) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
// ls-tree -l row: "<mode> <type> <object> <size>\t<path>"; the size is only numeric for
|
||||||
|
// blobs (trees read "-"), so gate on the type token and take the 4th whitespace field.
|
||||||
|
String[] parts = line.split("\\s+");
|
||||||
|
if (parts.length >= 4 && "blob".equals(parts[1])) {
|
||||||
|
try {
|
||||||
|
total += Long.parseLong(parts[3]);
|
||||||
|
} catch (NumberFormatException ignored) {
|
||||||
|
// a '-' size (or any anomaly) contributes nothing to the rough figure
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return total;
|
||||||
|
}
|
||||||
|
|
||||||
/** Resolve the directory that will hold per-session worktree checkouts. */
|
/** Resolve the directory that will hold per-session worktree checkouts. */
|
||||||
private Path resolveRoot(String repoRoot) {
|
private Path resolveRoot(String repoRoot) {
|
||||||
if (configuredRoot != null && !configuredRoot.isBlank()) {
|
if (configuredRoot != null && !configuredRoot.isBlank()) {
|
||||||
|
|||||||
@@ -53,6 +53,14 @@ public final class SessionManager implements TurnListener {
|
|||||||
private final int contextCap;
|
private final int contextCap;
|
||||||
private final boolean clearAfterTurn;
|
private final boolean clearAfterTurn;
|
||||||
private volatile MemberLifecycle memberLifecycle = MemberLifecycle.NONE;
|
private volatile MemberLifecycle memberLifecycle = MemberLifecycle.NONE;
|
||||||
|
/**
|
||||||
|
* CB-586: the repo root the fleet actually works in, remembered the first time a worktree
|
||||||
|
* session is spawned (worktrees are checkouts of it). {@code refs/wip/*} live there, and this
|
||||||
|
* single cached value is what the snapshot retention sweep and the operator-visible census run
|
||||||
|
* against. The daemon is bridged into one project at a time, so "the first worktree's repo" is
|
||||||
|
* the repo; {@code null} until any worktree is spawned, meaning nothing to sweep or measure.
|
||||||
|
*/
|
||||||
|
private volatile String fleetRepoRoot;
|
||||||
|
|
||||||
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
|
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
|
||||||
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
|
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||||
@@ -293,8 +301,10 @@ public final class SessionManager implements TurnListener {
|
|||||||
// blocked caller fails fast with a real reason instead of sitting on a rendezvous
|
// blocked caller fails fast with a real reason instead of sitting on a rendezvous
|
||||||
// nothing will ever resolve. CB-578 stage C: carry the worktree/branch/snapshot ref
|
// nothing will ever resolve. CB-578 stage C: carry the worktree/branch/snapshot ref
|
||||||
// too, so a failed ticket's detail can point a lead at the same tree to re-dispatch.
|
// too, so a failed ticket's detail can point a lead at the same tree to re-dispatch.
|
||||||
|
// CB-584 (issue #65 criterion 5): carry agentSessionId alongside them, so a lead can
|
||||||
|
// also resume the member's conversation, not just re-dispatch onto its files.
|
||||||
notifyReleased(new ReleaseDetail(removed.terminalId(), removed.worktree(),
|
notifyReleased(new ReleaseDetail(removed.terminalId(), removed.worktree(),
|
||||||
removed.branch(), snapshotRef));
|
removed.branch(), snapshotRef, removed.agentSessionId()));
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
// CB-581: the pane must always stop, even if the dirty check above threw. A session removed
|
// CB-581: the pane must always stop, even if the dirty check above threw. A session removed
|
||||||
@@ -348,7 +358,8 @@ public final class SessionManager implements TurnListener {
|
|||||||
* are {@code null} for a shared-tree session; {@code snapshotRef} is {@code null} unless this
|
* are {@code null} for a shared-tree session; {@code snapshotRef} is {@code null} unless this
|
||||||
* release snapshotted a dirty worktree into {@code refs/wip/<branch>}.
|
* release snapshotted a dirty worktree into {@code refs/wip/<branch>}.
|
||||||
*/
|
*/
|
||||||
public record ReleaseDetail(String terminalId, String worktreePath, String branch, String snapshotRef) {
|
public record ReleaseDetail(String terminalId, String worktreePath, String branch, String snapshotRef,
|
||||||
|
String agentSessionId) {
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -450,6 +461,11 @@ public final class SessionManager implements TurnListener {
|
|||||||
// The non-worktree path always used this chain; only this branch was missed.
|
// The non-worktree path always used this chain; only this branch was missed.
|
||||||
String repoRoot = worktrees.repoRoot(
|
String repoRoot = worktrees.repoRoot(
|
||||||
launcher.effectiveCwd(new SpawnRequest(preResolvedProfile, requestedCwd, callerCwd)));
|
launcher.effectiveCwd(new SpawnRequest(preResolvedProfile, requestedCwd, callerCwd)));
|
||||||
|
if (fleetRepoRoot == null) {
|
||||||
|
// CB-586: remember the repo whose worktrees the fleet spawns — its refs/wip/* are the
|
||||||
|
// snapshot store the retention sweep and the operator census operate on.
|
||||||
|
fleetRepoRoot = repoRoot;
|
||||||
|
}
|
||||||
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
|
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
|
||||||
String path = null;
|
String path = null;
|
||||||
PeerHandle handle;
|
PeerHandle handle;
|
||||||
@@ -740,6 +756,26 @@ public final class SessionManager implements TurnListener {
|
|||||||
return registry.size();
|
return registry.size();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586: the operator-visible census of {@code refs/wip/*} in the repo the fleet works in —
|
||||||
|
* how many snapshot refs exist and roughly what they cost. Empty (no repo known) until at
|
||||||
|
* least one worktree session has been spawned, exactly so a fleet that has never snapshotted
|
||||||
|
* anything surfaces nothing new, as it did before CB-586.
|
||||||
|
*/
|
||||||
|
public Optional<Worktrees.WipRefStats> wipRefs() {
|
||||||
|
String repo = fleetRepoRoot;
|
||||||
|
return repo == null ? Optional.empty() : Optional.of(worktrees.wipRefs(repo));
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586: run the snapshot retention sweep in the fleet's repo (a no-op until a worktree has
|
||||||
|
* been spawned, which establishes the repo). Returns how many {@code refs/wip/*} it deleted.
|
||||||
|
*/
|
||||||
|
public int sweepWipRefs(long minAgeMillis) {
|
||||||
|
String repo = fleetRepoRoot;
|
||||||
|
return repo == null ? 0 : worktrees.pruneWipRefs(repo, minAgeMillis);
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* The registered session owning {@code terminalId}, or {@code null} if none does.
|
* The registered session owning {@code terminalId}, or {@code null} if none does.
|
||||||
*
|
*
|
||||||
|
|||||||
@@ -14,12 +14,29 @@ public final class SessionReaper {
|
|||||||
|
|
||||||
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
|
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
|
||||||
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
|
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
|
||||||
|
/** CB-586: the refs/wip age floor — never sweep a snapshot younger than 24h (the CB-586 rule). */
|
||||||
|
private static final long WIP_MIN_AGE_MILLIS = TimeUnit.HOURS.toMillis(24);
|
||||||
|
/**
|
||||||
|
* CB-586: how often the retention sweep runs. Given the 24h age floor, running it every few
|
||||||
|
* hours means a ref is dropped within hours of becoming eligible, never within minutes.
|
||||||
|
*/
|
||||||
|
private static final long WIP_SWEEP_INTERVAL_NANOS = TimeUnit.HOURS.toNanos(6);
|
||||||
|
|
||||||
private final SessionManager sessions;
|
private final SessionManager sessions;
|
||||||
private final long idleTtlNanos;
|
private final long idleTtlNanos;
|
||||||
private final long intervalMillis;
|
private final long intervalMillis;
|
||||||
private volatile boolean running;
|
private volatile boolean running;
|
||||||
private Thread thread;
|
private Thread thread;
|
||||||
|
/**
|
||||||
|
* When the retention sweep last ran, and whether it ever has. The flag is not a convenience:
|
||||||
|
* a "never yet" sentinel value cannot be compared by subtraction. {@code Long.MIN_VALUE} was
|
||||||
|
* the obvious choice and it silently overflows — {@code System.nanoTime()} is positive on this
|
||||||
|
* platform, so {@code now - Long.MIN_VALUE} wraps to a large negative number, the interval gate
|
||||||
|
* reads it as "swept moments ago", and it returns before ever assigning the field. The sweep
|
||||||
|
* then never runs at all, for the life of the process, with nothing in the log to say so.
|
||||||
|
*/
|
||||||
|
private volatile boolean sweptOnce;
|
||||||
|
private volatile long lastWipSweepNanos;
|
||||||
|
|
||||||
/** Construct a reaper with the default 5-second polling interval. */
|
/** Construct a reaper with the default 5-second polling interval. */
|
||||||
public SessionReaper(SessionManager sessions, long idleTtlSeconds) {
|
public SessionReaper(SessionManager sessions, long idleTtlSeconds) {
|
||||||
@@ -49,10 +66,39 @@ public final class SessionReaper {
|
|||||||
} catch (RuntimeException e) {
|
} catch (RuntimeException e) {
|
||||||
log.warn("session reaper iteration failed; continuing", e);
|
log.warn("session reaper iteration failed; continuing", e);
|
||||||
}
|
}
|
||||||
|
maybeSweepWipRefs();
|
||||||
sleep();
|
sleep();
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586: run the refs/wip retention sweep on a slow cadence (hours, not the per-iteration
|
||||||
|
* millisecond loop). Best-effort — a failure must never take the idle-reap loop down with it.
|
||||||
|
*/
|
||||||
|
private void maybeSweepWipRefs() {
|
||||||
|
long now = System.nanoTime();
|
||||||
|
// The first pass always sweeps: a restart is a fine moment to sweep, the 24h age floor
|
||||||
|
// makes it safe, and it means the feature is observable right after a redeploy instead of
|
||||||
|
// six hours later. Only after that does the interval gate apply, and by then both operands
|
||||||
|
// come from nanoTime, so the subtraction is well-defined.
|
||||||
|
if (sweptOnce && now - lastWipSweepNanos < WIP_SWEEP_INTERVAL_NANOS) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
try {
|
||||||
|
int deleted = sessions.sweepWipRefs(WIP_MIN_AGE_MILLIS);
|
||||||
|
if (deleted > 0) {
|
||||||
|
log.info("refs/wip retention sweep deleted {} snapshot ref(s) older than 24h whose "
|
||||||
|
+ "content was already reachable from main", deleted);
|
||||||
|
}
|
||||||
|
} catch (RuntimeException e) {
|
||||||
|
log.warn("refs/wip retention sweep failed; continuing", e);
|
||||||
|
}
|
||||||
|
// Set even when the sweep threw, so a broken repo is retried on the slow cadence rather
|
||||||
|
// than hammering git on every 5-second iteration.
|
||||||
|
lastWipSweepNanos = now;
|
||||||
|
sweptOnce = true;
|
||||||
|
}
|
||||||
|
|
||||||
private void sleep() {
|
private void sleep() {
|
||||||
try {
|
try {
|
||||||
Thread.sleep(intervalMillis);
|
Thread.sleep(intervalMillis);
|
||||||
|
|||||||
@@ -54,4 +54,48 @@ public interface Worktrees {
|
|||||||
* tolerance — a worktree that is gone holds nothing to snapshot)
|
* tolerance — a worktree that is gone holds nothing to snapshot)
|
||||||
*/
|
*/
|
||||||
Optional<String> snapshot(String worktreePath, String branch, String message);
|
Optional<String> snapshot(String worktreePath, String branch, String message);
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586: how many {@code refs/wip/*} snapshot refs exist in {@code repoRoot} and roughly what
|
||||||
|
* they cost. This is the operator-visible surface for the snapshot growth CB-578 stage C left
|
||||||
|
* behind — counts of refs alone hide that each one pins a whole tree for {@code git gc}.
|
||||||
|
*
|
||||||
|
* @param repoRoot the repository to scan
|
||||||
|
* @return count of snapshot refs, and {@code costBytes} = the approximate total working-tree
|
||||||
|
* size of every snapshot's committed content (summed per ref, so shared objects are
|
||||||
|
* counted once per ref that carries them)
|
||||||
|
*/
|
||||||
|
WipRefStats wipRefs(String repoRoot);
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586: run the {@code refs/wip/*} retention sweep and return how many refs it deleted.
|
||||||
|
*
|
||||||
|
* <p>The retention rule is <em>reachability plus an age floor</em>. A snapshot ref is deleted
|
||||||
|
* only when <strong>both</strong> hold:
|
||||||
|
* <ol>
|
||||||
|
* <li>its commit's <em>tree content</em> is already reachable from {@code main} — the work
|
||||||
|
* the snapshot preserved has been recovered, so dropping the ref loses nothing; and</li>
|
||||||
|
* <li>the ref is older than {@code minAgeMillis} — a very recent snapshot is never swept
|
||||||
|
* while a lead may still be looking at it.</li>
|
||||||
|
* </ol>
|
||||||
|
*
|
||||||
|
* <p>Reachability is the safety property. A snapshot exists precisely because the work was not
|
||||||
|
* committed anywhere else, so a snapshot whose content is <em>not</em> reachable from
|
||||||
|
* {@code main} is the <strong>last copy</strong> of a worker's work and must never be deleted
|
||||||
|
* automatically — that is the failure CB-576 and CB-578 stage C were built to stop. Age alone
|
||||||
|
* must never drive a deletion, because age-based sweeping is exactly how the last copy gets
|
||||||
|
* destroyed. (Both numbers and the rule are CB-586's decision; this method only implements it.)
|
||||||
|
*
|
||||||
|
* <p>Every deletion logs the ref name and the commit sha, so an operator who finds they lost
|
||||||
|
* the wrong thing can still recover it from git's reflog.
|
||||||
|
*
|
||||||
|
* @param repoRoot the repository whose {@code refs/wip/*} to sweep
|
||||||
|
* @param minAgeMillis the age floor; a ref younger than this is never touched
|
||||||
|
* @return the number of snapshot refs deleted
|
||||||
|
*/
|
||||||
|
int pruneWipRefs(String repoRoot, long minAgeMillis);
|
||||||
|
|
||||||
|
/** CB-586: the operator-visible census of {@code refs/wip/*} in one repository. */
|
||||||
|
record WipRefStats(int count, long costBytes) {
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,108 @@
|
|||||||
|
package dev.ltms.bridged;
|
||||||
|
|
||||||
|
import ch.qos.logback.classic.Logger;
|
||||||
|
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||||
|
import ch.qos.logback.core.read.ListAppender;
|
||||||
|
import dev.ltms.bridged.config.BridgedConfig;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: an absent (or empty) {@code memberCredentials:} block blocks nothing — no name is
|
||||||
|
* hardcoded any more to fall back on, so the daemon must say so out loud at startup rather than
|
||||||
|
* silently dropping CB-592's protection. Mirrors {@link RequiredSecretEnvVarsTest}'s pattern for
|
||||||
|
* the CB-594 startup-secrets report, capturing the real log via a {@link ListAppender}.
|
||||||
|
*/
|
||||||
|
class MemberCredentialsGapReportTest {
|
||||||
|
|
||||||
|
private static BridgedConfig load(Path dir, String yaml) throws Exception {
|
||||||
|
Path f = dir.resolve("bridged.yaml");
|
||||||
|
Files.writeString(f, yaml);
|
||||||
|
return BridgedConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static ListAppender<ILoggingEvent> attach() {
|
||||||
|
Logger logger = (Logger) LoggerFactory.getLogger(Bridged.class);
|
||||||
|
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||||
|
appender.start();
|
||||||
|
logger.addAppender(appender);
|
||||||
|
return appender;
|
||||||
|
}
|
||||||
|
|
||||||
|
private static void detach(ListAppender<ILoggingEvent> appender) {
|
||||||
|
((Logger) LoggerFactory.getLogger(Bridged.class)).detachAppender(appender);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void anAbsentBlockWarnsThatEveryMemberInheritsTheWholeStore(@TempDir Path dir) throws Exception {
|
||||||
|
BridgedConfig cfg = load(dir, "bind:\n host: 127.0.0.1\n port: 8765\n");
|
||||||
|
|
||||||
|
ListAppender<ILoggingEvent> appender = attach();
|
||||||
|
try {
|
||||||
|
Bridged.reportMemberCredentialsGap(cfg);
|
||||||
|
} finally {
|
||||||
|
detach(appender);
|
||||||
|
}
|
||||||
|
|
||||||
|
assertTrue(appender.list.stream().anyMatch(e ->
|
||||||
|
e.getLevel() == ch.qos.logback.classic.Level.WARN
|
||||||
|
&& e.getFormattedMessage().contains("memberCredentials")
|
||||||
|
&& e.getFormattedMessage().contains("WHOLE secret store")),
|
||||||
|
"an absent block must WARN that protection is lost, not stay silent");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void anEmptyKnownListWarnsTheSameAsAbsent(@TempDir Path dir) throws Exception {
|
||||||
|
BridgedConfig cfg = load(dir, """
|
||||||
|
memberCredentials:
|
||||||
|
policy: deny-by-default
|
||||||
|
""");
|
||||||
|
|
||||||
|
ListAppender<ILoggingEvent> appender = attach();
|
||||||
|
try {
|
||||||
|
Bridged.reportMemberCredentialsGap(cfg);
|
||||||
|
} finally {
|
||||||
|
detach(appender);
|
||||||
|
}
|
||||||
|
|
||||||
|
assertTrue(appender.list.stream().anyMatch(e ->
|
||||||
|
e.getLevel() == ch.qos.logback.classic.Level.WARN
|
||||||
|
&& e.getFormattedMessage().contains("memberCredentials")),
|
||||||
|
"policy: with no known: names still blocks nothing and must warn the same way");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aPopulatedKnownListLogsInfoNotWarn(@TempDir Path dir) throws Exception {
|
||||||
|
BridgedConfig cfg = load(dir, """
|
||||||
|
memberCredentials:
|
||||||
|
policy: deny-by-default
|
||||||
|
allow: [AI_GATEWAY_TOKEN]
|
||||||
|
known: [AI_GATEWAY_TOKEN, GITEA_ACCESS_TOKEN]
|
||||||
|
""");
|
||||||
|
|
||||||
|
// logback-test.xml pins dev.ltms.bridged to WARN (see its own comment); raise it here so
|
||||||
|
// the INFO line this test asserts on actually reaches the appender, and restore after.
|
||||||
|
Logger logger = (Logger) LoggerFactory.getLogger(Bridged.class);
|
||||||
|
ch.qos.logback.classic.Level original = logger.getLevel();
|
||||||
|
logger.setLevel(ch.qos.logback.classic.Level.INFO);
|
||||||
|
ListAppender<ILoggingEvent> appender = attach();
|
||||||
|
try {
|
||||||
|
Bridged.reportMemberCredentialsGap(cfg);
|
||||||
|
} finally {
|
||||||
|
detach(appender);
|
||||||
|
logger.setLevel(original);
|
||||||
|
}
|
||||||
|
|
||||||
|
assertFalse(appender.list.stream().anyMatch(e -> e.getLevel() == ch.qos.logback.classic.Level.WARN),
|
||||||
|
"a configured, non-empty known: list must not warn — the block is doing its job");
|
||||||
|
assertTrue(appender.list.stream().anyMatch(e -> e.getFormattedMessage().contains("blocking 1")),
|
||||||
|
"the INFO line should say how many names are actually blocked (known minus allow)");
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -10,6 +10,7 @@ import java.nio.file.Path;
|
|||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
import java.util.Set;
|
import java.util.Set;
|
||||||
|
import java.util.regex.Pattern;
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.*;
|
import static org.junit.jupiter.api.Assertions.*;
|
||||||
|
|
||||||
@@ -1039,6 +1040,44 @@ class BridgedConfigTest {
|
|||||||
"an opencode worker with no argv defaults to the opencode binary, never claude");
|
"an opencode worker with no argv defaults to the opencode binary, never claude");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-604: an unrecognized {@code kind:} used to silently fall into the claude-code bucket — not
|
||||||
|
* matching {@code "opencode"} was the only check. With {@code argv:} also unset that meant the
|
||||||
|
* daemon tried to launch a program literally named after the typo.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void unknownKindIsRefusedAtLoadNamingTheValueAndTheAcceptedSet(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("kind-typo.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
profiles:
|
||||||
|
gemini:
|
||||||
|
kind: opencod
|
||||||
|
model: google/gemini-2.5-pro
|
||||||
|
""");
|
||||||
|
|
||||||
|
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
|
||||||
|
assertTrue(e.getMessage().contains("gemini"), "error names the profile: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("opencod"), "error names the bad value: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("claude-code") && e.getMessage().contains("opencode"),
|
||||||
|
"error names the accepted set: " + e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void blankKindStillDefaultsToClaudeCode(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("kind-blank.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
profiles:
|
||||||
|
gx10:
|
||||||
|
baseUrl: http://gx10.gw:8000
|
||||||
|
kind: ""
|
||||||
|
argv: ["ccs", "gx10"]
|
||||||
|
""");
|
||||||
|
|
||||||
|
BridgedConfig cfg = BridgedConfig.load(f);
|
||||||
|
assertEquals(BridgedConfig.Profile.KIND_CLAUDE_CODE, cfg.profiles().get("gx10").kind(),
|
||||||
|
"a blank kind: is documented to behave exactly like an absent one");
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void authDefaultsToLoopbackTrustSoExistingConfigsBehaveAsBefore(@TempDir Path dir) throws Exception {
|
void authDefaultsToLoopbackTrustSoExistingConfigsBehaveAsBefore(@TempDir Path dir) throws Exception {
|
||||||
Path f = dir.resolve("no-auth-block.yaml");
|
Path f = dir.resolve("no-auth-block.yaml");
|
||||||
@@ -1157,6 +1196,10 @@ class BridgedConfigTest {
|
|||||||
idleAfterSeconds: 600
|
idleAfterSeconds: 600
|
||||||
backoffMs: 45000
|
backoffMs: 45000
|
||||||
quietNudgeCap: 5
|
quietNudgeCap: 5
|
||||||
|
memberCredentials:
|
||||||
|
policy: deny-by-default
|
||||||
|
allow: [AI_GATEWAY_TOKEN]
|
||||||
|
known: [AI_GATEWAY_TOKEN, GITEA_ACCESS_TOKEN]
|
||||||
""");
|
""");
|
||||||
|
|
||||||
BridgedConfig cfg = BridgedConfig.load(f);
|
BridgedConfig cfg = BridgedConfig.load(f);
|
||||||
@@ -1187,6 +1230,84 @@ class BridgedConfigTest {
|
|||||||
assertEquals(600, cfg.leadHeartbeat().idleAfterSeconds(), "leadHeartbeat binds at the top level");
|
assertEquals(600, cfg.leadHeartbeat().idleAfterSeconds(), "leadHeartbeat binds at the top level");
|
||||||
assertEquals(45_000L, cfg.leadHeartbeat().backoffMs());
|
assertEquals(45_000L, cfg.leadHeartbeat().backoffMs());
|
||||||
assertEquals(5, cfg.leadHeartbeat().quietNudgeCap());
|
assertEquals(5, cfg.leadHeartbeat().quietNudgeCap());
|
||||||
|
|
||||||
|
assertEquals(BridgedConfig.MemberCredentials.POLICY_DENY_BY_DEFAULT, cfg.memberCredentials().policy(),
|
||||||
|
"memberCredentials binds at the top level");
|
||||||
|
assertEquals(Set.of("AI_GATEWAY_TOKEN"), cfg.memberCredentials().allowSet());
|
||||||
|
assertEquals(Set.of("GITEA_ACCESS_TOKEN"), cfg.memberCredentials().blockedSet());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* A top-level key {@code BridgedConfig} reads but that appears nowhere in
|
||||||
|
* {@code bridged.example.yaml} — live or commented — is invisible drift: {@code bridged.yaml}
|
||||||
|
* is gitignored, so the example is the ONLY committed description of the config schema, and
|
||||||
|
* neither {@link #shippedExampleConfigParses} (example → code: does the example still parse)
|
||||||
|
* nor {@link #everyOptionalKnobDocumentedInTheExampleBinds} (a hand-maintained list of keys
|
||||||
|
* that must bind) can catch a brand-new key nobody added to either.
|
||||||
|
*
|
||||||
|
* <p>This test compares the OTHER direction: every key in {@link BridgedConfig#KNOWN_TOP_LEVEL_KEYS}
|
||||||
|
* (the parser's own accepted set, which backs the unknown-key WARN) must appear as a top-level
|
||||||
|
* key in the example text, live or commented-out — see {@link #topLevelKeyDocumented}.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void everyKnownTopLevelKeyIsDocumentedInTheExample() throws Exception {
|
||||||
|
Path example = Path.of("bridged.example.yaml");
|
||||||
|
assertTrue(Files.exists(example), "bridged.example.yaml must ship next to the pom");
|
||||||
|
String text = Files.readString(example);
|
||||||
|
|
||||||
|
List<String> undocumented = BridgedConfig.KNOWN_TOP_LEVEL_KEYS.stream()
|
||||||
|
.filter(key -> !topLevelKeyDocumented(text, key))
|
||||||
|
.sorted()
|
||||||
|
.toList();
|
||||||
|
|
||||||
|
assertTrue(undocumented.isEmpty(), () -> "key(s) " + undocumented
|
||||||
|
+ " are read by BridgedConfig but appear nowhere in bridged.example.yaml — "
|
||||||
|
+ "document each one there, commented out if optional. bridged.yaml is "
|
||||||
|
+ "gitignored, so this file is the only committed description of the config "
|
||||||
|
+ "schema an operator or a worker can see.");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Most of {@code bridged.example.yaml} is deliberately commented out — optional sections are
|
||||||
|
* documented as commented blocks so the shipped file stays a working minimal config. A key
|
||||||
|
* documented ONLY as a comment must still count as documented; parsing the file as YAML and
|
||||||
|
* reading its live key set (as an earlier attempt at this guard did) gets this wrong, because
|
||||||
|
* every commented section then looks entirely absent.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void commentedOnlyTopLevelKeyCountsAsDocumented() {
|
||||||
|
String yaml = """
|
||||||
|
bind:
|
||||||
|
port: 8765
|
||||||
|
# broker:
|
||||||
|
# uri: amqp://guest:guest@127.0.0.1:5672
|
||||||
|
""";
|
||||||
|
assertTrue(topLevelKeyDocumented(yaml, "broker"),
|
||||||
|
"a key documented only inside a commented-out block must still count as documented");
|
||||||
|
}
|
||||||
|
|
||||||
|
/** A key that appears in neither a live nor a commented top-level line must NOT count. */
|
||||||
|
@Test
|
||||||
|
void absentTopLevelKeyIsNotDocumented() {
|
||||||
|
String yaml = """
|
||||||
|
bind:
|
||||||
|
port: 8765
|
||||||
|
""";
|
||||||
|
assertFalse(topLevelKeyDocumented(yaml, "broker"),
|
||||||
|
"a key mentioned nowhere in the example must not be reported as documented");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* True when {@code key} appears as a top-level YAML key in {@code yaml} — either live
|
||||||
|
* ({@code key:} at column 0) or commented out ({@code # key:}, also at column 0, with only
|
||||||
|
* whitespace between the {@code #} and the key). Anchoring on column 0 is what keeps this a
|
||||||
|
* top-level check: an indented occurrence (a nested field, or prose inside a comment that
|
||||||
|
* happens to end in a colon) never matches, because {@code ^} requires the key's own first
|
||||||
|
* character — or the sole leading {@code #} — to sit at the very start of the line.
|
||||||
|
*/
|
||||||
|
private static boolean topLevelKeyDocumented(String yaml, String key) {
|
||||||
|
Pattern p = Pattern.compile("(?m)^(?:#\\s*)?" + Pattern.quote(key) + ":");
|
||||||
|
return p.matcher(yaml).find();
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -1295,6 +1416,158 @@ class BridgedConfigTest {
|
|||||||
assertTrue(e.getMessage().contains("maxLoad"), "error names the key: " + e.getMessage());
|
assertTrue(e.getMessage().contains("maxLoad"), "error names the key: " + e.getMessage());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-606: an unrecognized {@code auth.mode} used to silently fall back to
|
||||||
|
* {@code loopback-trust} — {@link BridgedConfig.Auth#tokenMode()} only checked equality
|
||||||
|
* against {@code "token"}. On a loopback bind {@link BridgedConfig#validateAuthExposure()}
|
||||||
|
* never runs (it only fires for a non-loopback bind), so the typo was completely invisible:
|
||||||
|
* the daemon started cleanly and authenticated nobody while the operator believed token mode
|
||||||
|
* was active.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void unknownAuthModeIsRefusedAtLoadNamingTheValueAndTheAcceptedSet(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("auth-mode-typo.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
auth:
|
||||||
|
mode: toekn
|
||||||
|
""");
|
||||||
|
|
||||||
|
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
|
||||||
|
assertTrue(e.getMessage().contains("toekn"), "error names the bad value: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("loopback-trust") && e.getMessage().contains("token"),
|
||||||
|
"error names the accepted set: " + e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-606: an unrecognized per-profile {@code placement:} used to silently fall back to legacy
|
||||||
|
* pane placement — {@link BridgedConfig.Profile#tabPlacement()} only checked equality against
|
||||||
|
* {@code "tab"}.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void unknownProfilePlacementIsRefusedAtLoadNamingTheProfileAndTheAcceptedSet(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("placement-typo.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
profiles:
|
||||||
|
gx10:
|
||||||
|
baseUrl: http://gx10.gw:8000
|
||||||
|
placement: tabb
|
||||||
|
""");
|
||||||
|
|
||||||
|
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
|
||||||
|
assertTrue(e.getMessage().contains("gx10"), "error names the profile: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("tabb"), "error names the bad value: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("tab") && e.getMessage().contains("pane"),
|
||||||
|
"error names the accepted set: " + e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: {@code memberCredentials.policy} is validated the same way {@code auth.mode} and
|
||||||
|
* per-profile {@code placement} are (CB-606's pattern) — a typo must not silently behave as the
|
||||||
|
* one real policy, because the day a second policy exists that silent fallback becomes a real
|
||||||
|
* behavior change instead of a happy accident.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void unknownMemberCredentialsPolicyIsRefusedAtLoadNamingTheValueAndTheAcceptedSet(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("member-credentials-policy-typo.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
memberCredentials:
|
||||||
|
policy: deny-by-defualt
|
||||||
|
known:
|
||||||
|
- GITEA_ACCESS_TOKEN
|
||||||
|
""");
|
||||||
|
|
||||||
|
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
|
||||||
|
assertTrue(e.getMessage().contains("deny-by-defualt"), "error names the bad value: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("deny-by-default"), "error names the accepted set: " + e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
/** {@code allow}/{@code known} bind and {@link BridgedConfig.MemberCredentials#blockedSet()} is known minus allow. */
|
||||||
|
@Test
|
||||||
|
void memberCredentialsBindsAllowAndKnownAndComputesBlockedSet(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("member-credentials.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
port: 8080
|
||||||
|
memberCredentials:
|
||||||
|
policy: deny-by-default
|
||||||
|
allow:
|
||||||
|
- AI_GATEWAY_TOKEN
|
||||||
|
- WORKER_GITEA_TOKEN
|
||||||
|
known:
|
||||||
|
- AI_GATEWAY_TOKEN
|
||||||
|
- WORKER_GITEA_TOKEN
|
||||||
|
- GITEA_ACCESS_TOKEN
|
||||||
|
- GITLAB_PERSONAL_ACCESS_TOKEN
|
||||||
|
""");
|
||||||
|
|
||||||
|
BridgedConfig.MemberCredentials mc = BridgedConfig.load(f).memberCredentials();
|
||||||
|
assertEquals(BridgedConfig.MemberCredentials.POLICY_DENY_BY_DEFAULT, mc.policy());
|
||||||
|
assertEquals(Set.of("AI_GATEWAY_TOKEN", "WORKER_GITEA_TOKEN"), mc.allowSet());
|
||||||
|
assertEquals(Set.of("GITEA_ACCESS_TOKEN", "GITLAB_PERSONAL_ACCESS_TOKEN"), mc.blockedSet(),
|
||||||
|
"blockedSet is known minus allow");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596: omitting {@code memberCredentials:} entirely must NOT crash a reader that assumes a
|
||||||
|
* non-null block (the same "fill in nested defaults" contract every other structural field
|
||||||
|
* gets — see {@link BridgedConfig#withDefaults()}), but it also must not pretend anything is
|
||||||
|
* blocked: an empty {@code known} list blocks nothing, and that is a real gap the operator must
|
||||||
|
* close by configuring this block, not a safe default.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void absentMemberCredentialsDefaultsToAnEmptyNonNullBlock(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("member-credentials-absent.yaml");
|
||||||
|
Files.writeString(f, "bind:\n port: 8080\n");
|
||||||
|
|
||||||
|
BridgedConfig.MemberCredentials mc = BridgedConfig.load(f).memberCredentials();
|
||||||
|
assertNotNull(mc, "withDefaults() must never leave this null");
|
||||||
|
assertTrue(mc.known().isEmpty(), "no known list configured — nothing is blocked");
|
||||||
|
assertTrue(mc.allowSet().isEmpty());
|
||||||
|
assertTrue(mc.blockedSet().isEmpty());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void absentProfilePlacementDefaultsToTab(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("placement-absent.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
profiles:
|
||||||
|
gx10:
|
||||||
|
baseUrl: http://gx10.gw:8000
|
||||||
|
""");
|
||||||
|
|
||||||
|
BridgedConfig cfg = BridgedConfig.load(f);
|
||||||
|
BridgedConfig.Profile w = cfg.profiles().get("gx10");
|
||||||
|
assertTrue(w.tabPlacement(), "an absent placement must keep defaulting to tab");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-606: the top-level {@code placement:} policy name WAS validated, but only lazily, by
|
||||||
|
* {@code PlacementPolicies.fromName} through {@code CompositePeerLauncher}'s per-spawn
|
||||||
|
* {@code Supplier} — so a bad name still started a daemon that looked healthy and failed only
|
||||||
|
* the first time something spawned without naming a profile. This must now fail at load.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void unknownTopLevelPlacementPolicyIsRefusedAtLoadNotLazilyAtFirstSpawn(@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("placement-policy-typo.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
port: 8080
|
||||||
|
placement: weightd
|
||||||
|
""");
|
||||||
|
|
||||||
|
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
|
||||||
|
assertTrue(e.getMessage().contains("weightd"), "error names the bad value: " + e.getMessage());
|
||||||
|
assertTrue(e.getMessage().contains("fixed") && e.getMessage().contains("round-robin")
|
||||||
|
&& e.getMessage().contains("weighted"),
|
||||||
|
"error names the accepted set: " + e.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void subscriptionFlagBindsAndDefaultsFalse(@TempDir Path dir) throws Exception {
|
void subscriptionFlagBindsAndDefaultsFalse(@TempDir Path dir) throws Exception {
|
||||||
Path f = dir.resolve("subscription.yaml");
|
Path f = dir.resolve("subscription.yaml");
|
||||||
|
|||||||
@@ -7,6 +7,7 @@ import java.util.ArrayList;
|
|||||||
import java.util.LinkedHashMap;
|
import java.util.LinkedHashMap;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.CopyOnWriteArrayList;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Recording fake {@link HerdrClient} for unit/acceptance tests. Returns canned frames
|
* Recording fake {@link HerdrClient} for unit/acceptance tests. Returns canned frames
|
||||||
@@ -22,7 +23,13 @@ public final class FakeHerdr implements HerdrClient {
|
|||||||
public static final long WORKER_PID = 4242;
|
public static final long WORKER_PID = 4242;
|
||||||
|
|
||||||
private final ObjectMapper mapper = new ObjectMapper();
|
private final ObjectMapper mapper = new ObjectMapper();
|
||||||
public final List<Call> calls = new ArrayList<>();
|
/**
|
||||||
|
* Thread-safe on purpose. Background loops — {@link dev.ltms.bridged.msg.ReplyPushLoop} and the
|
||||||
|
* lead heartbeat — call this fake from their own scheduler threads while a test polls
|
||||||
|
* {@link #called} from the test thread. A plain {@code ArrayList} threw
|
||||||
|
* {@code ConcurrentModificationException} out of {@code called()} when a nudge landed mid-stream.
|
||||||
|
*/
|
||||||
|
public final List<Call> calls = new CopyOnWriteArrayList<>();
|
||||||
private boolean healthy = true;
|
private boolean healthy = true;
|
||||||
private final List<String> extraWorkspaces = new ArrayList<>();
|
private final List<String> extraWorkspaces = new ArrayList<>();
|
||||||
private final List<String> extraAgents = new ArrayList<>();
|
private final List<String> extraAgents = new ArrayList<>();
|
||||||
|
|||||||
@@ -16,12 +16,15 @@ import dev.ltms.bridged.inject.MemberPresence;
|
|||||||
import dev.ltms.bridged.session.MemberSession;
|
import dev.ltms.bridged.session.MemberSession;
|
||||||
import dev.ltms.bridged.session.WorktreeRequest;
|
import dev.ltms.bridged.session.WorktreeRequest;
|
||||||
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
||||||
|
import dev.ltms.bridged.member.CompositePeerLauncher;
|
||||||
import dev.ltms.bridged.placement.BackendQuarantine;
|
import dev.ltms.bridged.placement.BackendQuarantine;
|
||||||
|
import dev.ltms.bridged.placement.PlacementPolicies;
|
||||||
import io.modelcontextprotocol.spec.McpSchema;
|
import io.modelcontextprotocol.spec.McpSchema;
|
||||||
import dev.ltms.bridged.msg.InMemoryReplyInbox;
|
import dev.ltms.bridged.msg.InMemoryReplyInbox;
|
||||||
import org.junit.jupiter.api.BeforeEach;
|
import org.junit.jupiter.api.BeforeEach;
|
||||||
import org.junit.jupiter.api.Test;
|
import org.junit.jupiter.api.Test;
|
||||||
|
|
||||||
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
import java.util.Set;
|
import java.util.Set;
|
||||||
import java.util.concurrent.CompletableFuture;
|
import java.util.concurrent.CompletableFuture;
|
||||||
@@ -423,6 +426,35 @@ class BridgeMcpTest {
|
|||||||
assertTrue(textOf(res).contains("unknown worker profile"), textOf(res));
|
assertTrue(textOf(res).contains("unknown worker profile"), textOf(res));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-599: a profile at its {@code maxLoad} cap must surface a readable reason on the MCP
|
||||||
|
* surface too, not merely flip {@code isError} with an opaque or absent message.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void spawnAtMaxLoadSurfacesTheCapacityReason() {
|
||||||
|
FakeHerdr h = new FakeHerdr();
|
||||||
|
BridgedConfig.Profile wcfg = new BridgedConfig.Profile(
|
||||||
|
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
|
||||||
|
"tab", "bridged-workers", "worker: {profile} #{n}", null,
|
||||||
|
null, null, null, null, null, null, null, 0, null, null, null);
|
||||||
|
Map<String, BridgedConfig.Profile> profiles = Map.of(wcfg.profile(), wcfg);
|
||||||
|
ClaudeCodeLauncher delegate = new ClaudeCodeLauncher(
|
||||||
|
new AgentControl(h), new WorkspaceControl(h), new SubscriptionGuard(Set.of("gx00.gw")),
|
||||||
|
profiles, wcfg.profile(), k -> "BRIDGED_WORKER_TOKEN".equals(k) ? "tok" : null);
|
||||||
|
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||||
|
List.of(delegate), wcfg.profile(), profiles, PlacementPolicies.fixed(), _ -> 0);
|
||||||
|
SessionManager sm = new SessionManager(composite);
|
||||||
|
|
||||||
|
McpSchema.CallToolResult res = BridgeMcp.spawn(sm, "ltms-local");
|
||||||
|
|
||||||
|
assertTrue(res.isError());
|
||||||
|
String text = textOf(res);
|
||||||
|
assertTrue(text.contains("no capacity"), "surfaces a capacity reason, not a bare error: " + text);
|
||||||
|
assertTrue(text.contains("ltms-local"), "names the profile: " + text);
|
||||||
|
assertTrue(text.contains("maxLoad"), "explains the refusal: " + text);
|
||||||
|
assertFalse(h.called("agent.start"), "at cap, the spawn is refused before any herdr call");
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void spawnPassesTheRequestedCwdToTheWorker() {
|
void spawnPassesTheRequestedCwdToTheWorker() {
|
||||||
FakeHerdr h = new FakeHerdr();
|
FakeHerdr h = new FakeHerdr();
|
||||||
@@ -741,6 +773,52 @@ class BridgeMcpTest {
|
|||||||
assertEquals("blocked", textOf(res));
|
assertEquals("blocked", textOf(res));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-582: a lead polling {@code bridge_status} on its normal cadence — not {@code bridge_poll}
|
||||||
|
* — must also see a worker's open async {@code bridge_ask} question, since the reverse-rendezvous
|
||||||
|
* window it opened with is far shorter than that cadence.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void statusReportsAnOpenQuestionWhenTheWorkerIsMidAsk() throws Exception {
|
||||||
|
String ticket = messages.sendAsync(T, "task that asks");
|
||||||
|
long deadline = System.currentTimeMillis() + 3000;
|
||||||
|
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||||
|
Thread.sleep(5);
|
||||||
|
}
|
||||||
|
assertTrue(rendezvous.isWaiting(T), "sendAsync should have opened its rendezvous waiter");
|
||||||
|
|
||||||
|
CompletableFuture<MessageService.AskResult> ask =
|
||||||
|
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||||
|
|
||||||
|
MessageService.TaskView asking;
|
||||||
|
deadline = System.currentTimeMillis() + 3000;
|
||||||
|
do {
|
||||||
|
asking = messages.poll(ticket);
|
||||||
|
Thread.sleep(5);
|
||||||
|
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
|
||||||
|
assertEquals(MessageService.Phase.ASKING, asking.phase());
|
||||||
|
|
||||||
|
McpSchema.CallToolResult res = BridgeMcp.status(messages, T);
|
||||||
|
assertNotEquals(Boolean.TRUE, res.isError());
|
||||||
|
String out = textOf(res);
|
||||||
|
assertTrue(out.startsWith("idle"), "the live status must still lead the text: " + out);
|
||||||
|
assertTrue(out.contains("which config file?"), "the question text must be shown: " + out);
|
||||||
|
assertTrue(out.contains("turnId=\"" + asking.turnId() + "\""), "the turnId must be shown: " + out);
|
||||||
|
assertTrue(out.contains("(ticket " + ticket + ")"), "the ticket must be shown: " + out);
|
||||||
|
|
||||||
|
// Clean up the still-open ask so the background thread does not linger past the test.
|
||||||
|
String turnId = asking.turnId();
|
||||||
|
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||||
|
() -> messages.answer(turnId, "config.yaml", 5000));
|
||||||
|
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||||
|
deadline = System.currentTimeMillis() + 3000;
|
||||||
|
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||||
|
Thread.sleep(5);
|
||||||
|
}
|
||||||
|
assertTrue(rendezvous.resolve(T, "done"));
|
||||||
|
answer.get(5, TimeUnit.SECONDS);
|
||||||
|
}
|
||||||
|
|
||||||
// --- bridge_whoami: the caller's own identity, so an agent never has to guess its role -------
|
// --- bridge_whoami: the caller's own identity, so an agent never has to guess its role -------
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
|
|||||||
@@ -1,5 +1,8 @@
|
|||||||
package dev.ltms.bridged.member;
|
package dev.ltms.bridged.member;
|
||||||
|
|
||||||
|
import ch.qos.logback.classic.Logger;
|
||||||
|
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||||
|
import ch.qos.logback.core.read.ListAppender;
|
||||||
import dev.ltms.bridged.config.BridgedConfig;
|
import dev.ltms.bridged.config.BridgedConfig;
|
||||||
import dev.ltms.bridged.guard.GuardException;
|
import dev.ltms.bridged.guard.GuardException;
|
||||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||||
@@ -12,7 +15,11 @@ import dev.ltms.bridged.peer.PeerHandle;
|
|||||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||||
import dev.ltms.bridged.peer.SpawnRequest;
|
import dev.ltms.bridged.peer.SpawnRequest;
|
||||||
import org.junit.jupiter.api.Test;
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
import java.util.Set;
|
import java.util.Set;
|
||||||
@@ -102,10 +109,24 @@ class ClaudeCodeLauncherTest {
|
|||||||
"the operator's own args are preserved, in order, ahead of the model flag");
|
"the operator's own args are preserved, in order, ahead of the model flag");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- CB-617/CB-618: both charters travel in ONE file; the reply charter alone stays inline ---
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-617: herdr refuses to shell-encode a multi-line inline argv argument
|
||||||
|
* ({@code invalid_agent_argument}) — the exact failure this reproduced on profile {@code opus}.
|
||||||
|
* The role charter is operator-authored and often multi-line, so it must never appear as an argv
|
||||||
|
* element.
|
||||||
|
*
|
||||||
|
* <p>CB-618: and Claude Code itself refuses to start when both {@code --append-system-prompt} and
|
||||||
|
* {@code --append-system-prompt-file} are given ("Cannot use both ... Please use only one"), so
|
||||||
|
* the reply charter cannot ride inline alongside a role charter either. Both go in the one file,
|
||||||
|
* reply charter last. This drives the real launcher entry point ({@code spawn}), the same path a
|
||||||
|
* live spawn takes — not the argv builder in isolation.
|
||||||
|
*/
|
||||||
@Test
|
@Test
|
||||||
void appendsTheBaseComposedRoleAndReplyCharter() {
|
void bothChartersTravelInOneFileAndNeverOnBothFlags() {
|
||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
String roleCharter = "You review changes.";
|
String roleCharter = "You review changes.\nLine two.\nLine three.";
|
||||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||||
"sonnet", "http://gx00.gw:8000", null, null, "BRIDGED_WORKER_TOKEN",
|
"sonnet", "http://gx00.gw:8000", null, null, "BRIDGED_WORKER_TOKEN",
|
||||||
List.of("claude"), "tab", "bridged-workers", "w #{n}",
|
List.of("claude"), "tab", "bridged-workers", "w #{n}",
|
||||||
@@ -117,11 +138,20 @@ class ClaudeCodeLauncherTest {
|
|||||||
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.REVIEWER));
|
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.REVIEWER));
|
||||||
|
|
||||||
List<String> args = spawnedArgs(herdr);
|
List<String> args = spawnedArgs(herdr);
|
||||||
int flag = args.indexOf("--append-system-prompt");
|
assertTrue(args.stream().noneMatch(a -> a.contains("\n")),
|
||||||
assertEquals(1, args.stream().filter("--append-system-prompt"::equals).count(),
|
"no argv element may be multi-line — herdr cannot shell-encode one: " + args);
|
||||||
"the composed charter is passed once");
|
|
||||||
assertEquals(roleCharter + "\n\n" + HerdrPeerLauncher.REPLY_CHARTER, args.get(flag + 1),
|
int fileFlag = args.indexOf("--append-system-prompt-file");
|
||||||
"the role charter comes first and the reply rule comes last");
|
assertTrue(fileFlag >= 0, "the role charter is mounted via --append-system-prompt-file: " + args);
|
||||||
|
assertFalse(args.contains("--append-system-prompt"),
|
||||||
|
"CB-618: Claude Code refuses to start with both flags — the reply charter must not "
|
||||||
|
+ "ride inline beside a role charter: " + args);
|
||||||
|
assertDoesNotThrow(() -> {
|
||||||
|
String written = Files.readString(Path.of(args.get(fileFlag + 1)));
|
||||||
|
assertTrue(written.startsWith(roleCharter), "the file opens with the role charter: " + written);
|
||||||
|
assertTrue(written.endsWith(HerdrPeerLauncher.REPLY_CHARTER),
|
||||||
|
"the reply charter is last — it is the rule that must survive: " + written);
|
||||||
|
}, "the --append-system-prompt-file path must be a readable file");
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -131,11 +161,13 @@ class ClaudeCodeLauncherTest {
|
|||||||
|
|
||||||
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.DEV));
|
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.DEV));
|
||||||
|
|
||||||
assertFalse(spawnedArgs(herdr).contains("--append-system-prompt"));
|
List<String> args = spawnedArgs(herdr);
|
||||||
|
assertFalse(args.contains("--append-system-prompt"));
|
||||||
|
assertFalse(args.contains("--append-system-prompt-file"));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void profileWithoutMcpStillGetsItsRoleCharter() {
|
void profileWithoutMcpStillGetsItsRoleCharterAsAFileNotInline() {
|
||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
String roleCharter = "You design changes.";
|
String roleCharter = "You design changes.";
|
||||||
ClaudeCodeLauncher svc = labelService(herdr, () -> fleet(Map.of("architect", roleCharter), null));
|
ClaudeCodeLauncher svc = labelService(herdr, () -> fleet(Map.of("architect", roleCharter), null));
|
||||||
@@ -143,12 +175,52 @@ class ClaudeCodeLauncherTest {
|
|||||||
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.ARCHITECT));
|
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.ARCHITECT));
|
||||||
|
|
||||||
List<String> args = spawnedArgs(herdr);
|
List<String> args = spawnedArgs(herdr);
|
||||||
int flag = args.indexOf("--append-system-prompt");
|
int flag = args.indexOf("--append-system-prompt-file");
|
||||||
assertTrue(flag >= 0, "a role charter does not need an MCP mount");
|
assertTrue(flag >= 0, "a role charter does not need an MCP mount");
|
||||||
assertEquals(roleCharter, args.get(flag + 1));
|
assertDoesNotThrow(() -> assertEquals(roleCharter,
|
||||||
|
Files.readString(Path.of(args.get(flag + 1)))),
|
||||||
|
"the file holds the role charter");
|
||||||
|
assertFalse(args.contains("--append-system-prompt"), "no reply charter without an MCP mount");
|
||||||
assertFalse(args.contains("--mcp-config"));
|
assertFalse(args.contains("--mcp-config"));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- CB-617: --agent <role> when the role has an agent-definition file --------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void agentFlagIsPassedWhenTheRoleAgentDefinitionFileExists(@TempDir Path cwd) throws Exception {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Path agentsDir = Files.createDirectories(cwd.resolve(".claude/agents"));
|
||||||
|
Files.writeString(agentsDir.resolve("reviewer.md"), "---\nname: reviewer\n---\nBe a reviewer.");
|
||||||
|
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||||
|
"sonnet", "http://gx00.gw:8000", null, null, "BRIDGED_WORKER_TOKEN",
|
||||||
|
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
|
||||||
|
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||||
|
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||||
|
|
||||||
|
svc.spawn(new SpawnRequest("sonnet", cwd.toString(), null, null, null, MemberRole.REVIEWER));
|
||||||
|
|
||||||
|
List<String> args = spawnedArgs(herdr);
|
||||||
|
int flag = args.indexOf("--agent");
|
||||||
|
assertTrue(flag >= 0, "--agent is passed when the role's agent file exists: " + args);
|
||||||
|
assertEquals("reviewer", args.get(flag + 1));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void noAgentFlagWhenTheRoleAgentDefinitionFileIsAbsent(@TempDir Path cwd) {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||||
|
"sonnet", "http://gx00.gw:8000", null, null, "BRIDGED_WORKER_TOKEN",
|
||||||
|
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
|
||||||
|
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||||
|
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||||
|
|
||||||
|
PeerHandle handle = svc.spawn(new SpawnRequest("sonnet", cwd.toString(), null, null, null, MemberRole.REVIEWER));
|
||||||
|
|
||||||
|
assertNotNull(handle, "the member still spawns with no agent-definition file");
|
||||||
|
assertFalse(spawnedArgs(herdr).contains("--agent"),
|
||||||
|
"no --agent flag when the role has no agent-definition file");
|
||||||
|
}
|
||||||
|
|
||||||
private ClaudeCodeLauncher multiProfile(FakeHerdr herdr) {
|
private ClaudeCodeLauncher multiProfile(FakeHerdr herdr) {
|
||||||
BridgedConfig.Profile gx10 = new BridgedConfig.Profile("gx10", "http://gx10.gw:8000", "coder",
|
BridgedConfig.Profile gx10 = new BridgedConfig.Profile("gx10", "http://gx10.gw:8000", "coder",
|
||||||
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
|
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
|
||||||
@@ -676,39 +748,90 @@ class ClaudeCodeLauncherTest {
|
|||||||
"the guard-checked baseUrl must win over any env: entry, or the boundary is bypassable");
|
"the guard-checked baseUrl must win over any env: entry, or the boundary is bypassable");
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- CB-592: the admin GITEA_ACCESS_TOKEN never reaches a member -----------------------------
|
// --- CB-596: config-driven member-credential policy (replaces CB-592's hardcoded single name) --
|
||||||
|
|
||||||
/**
|
/** A representative {@code memberCredentials} — 4 allowed, 4 blocked, matching the real ticket shape. */
|
||||||
* herdr's env map is an overlay onto its own (login-shell) process environment, so a worker
|
private static final BridgedConfig.MemberCredentials TEST_MEMBER_CREDENTIALS = new BridgedConfig.MemberCredentials(
|
||||||
* inherits whatever the daemon's shell carries — including the admin GITEA_ACCESS_TOKEN — for
|
null,
|
||||||
* every key baseEnv does not explicitly shadow. This pins that the launcher DOES send an
|
List.of("AI_GATEWAY_TOKEN", "WORKER_GITEA_TOKEN", "CONTEXT7_TOKEN", "GITEA_HOST"),
|
||||||
* explicit (non-blank) GITEA_ACCESS_TOKEN to herdr on every spawn, whatever the profile is, so
|
List.of("AI_GATEWAY_TOKEN", "WORKER_GITEA_TOKEN", "CONTEXT7_TOKEN", "GITEA_HOST",
|
||||||
* a future baseEnv refactor cannot silently drop it and reopen the leak. Asserted against what
|
"GITEA_ACCESS_TOKEN", "GITLAB_PERSONAL_ACCESS_TOKEN", "TS_AUTHKEY", "HASS_TOKEN"));
|
||||||
* tab.create's params actually carry, not an internal map built in the test (gitea #77).
|
|
||||||
*
|
|
||||||
* <p>Scope, measured on a live pane 2026-08-15: this pins what the launcher SENDS, and that is
|
|
||||||
* all it can pin. It does not prove the value survives, and it does not: the pane runs a login
|
|
||||||
* shell, ~/.zprofile sources secrets.sh, and its unconditional `export GITEA_ACCESS_TOKEN=...`
|
|
||||||
* puts the real token back over this sentinel. Closing that needs the operator to guard the
|
|
||||||
* export on BRIDGED_MEMBER — see everySpawnMarksThePaneAsAMember below.
|
|
||||||
*/
|
|
||||||
@Test
|
|
||||||
void everySpawnShadowsTheAdminGiteaAccessToken() {
|
|
||||||
FakeHerdr herdr = new FakeHerdr();
|
|
||||||
service(herdr, List.of("claude"), null).spawn();
|
|
||||||
|
|
||||||
String shadowed = startEnv(herdr).get("GITEA_ACCESS_TOKEN");
|
private ClaudeCodeLauncher serviceWithCredentials(FakeHerdr herdr, BridgedConfig.MemberCredentials creds) {
|
||||||
assertNotNull(shadowed, "GITEA_ACCESS_TOKEN must be explicitly overlaid, not left unmentioned");
|
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||||
assertFalse(shadowed.isBlank(), "a blank overlay value's override behaviour is unverified — must be non-blank");
|
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||||
|
List.of("claude"), "tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
|
||||||
|
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||||
|
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||||
|
_ -> null, 0, System::currentTimeMillis, () -> {}, null, () -> creds);
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* No profile — present or future — may restore the admin token by naming it in {@code env:}.
|
* herdr's env map is an overlay onto its own (login-shell) process environment, so a worker
|
||||||
|
* inherits whatever the daemon's shell carries — including the operator's own credentials — for
|
||||||
|
* every key {@code baseEnv} does not explicitly shadow. This pins that every {@code known} name
|
||||||
|
* NOT also {@code allow}-ed gets an explicit (non-blank) sentinel overlay, whatever the profile
|
||||||
|
* is. Asserted against what tab.create's params actually carry, not an internal map built in the
|
||||||
|
* test (gitea #82).
|
||||||
|
*
|
||||||
|
* <p>Scope, measured on a live pane 2026-08-15 (CB-592): this pins what the launcher SENDS, and
|
||||||
|
* that is all a unit test can pin. It does not prove the value survives, and for names the
|
||||||
|
* operator's secrets.sh also exports it does not: the pane runs a login shell that puts the real
|
||||||
|
* value back over this sentinel unless the export is guarded on BRIDGED_MEMBER — see
|
||||||
|
* everySpawnMarksThePaneAsAMember below.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void everyKnownNameNotAllowedIsShadowedWithTheSentinel() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
serviceWithCredentials(herdr, TEST_MEMBER_CREDENTIALS).spawn();
|
||||||
|
|
||||||
|
Map<String, String> env = startEnv(herdr);
|
||||||
|
for (String blocked : List.of("GITEA_ACCESS_TOKEN", "GITLAB_PERSONAL_ACCESS_TOKEN", "TS_AUTHKEY", "HASS_TOKEN")) {
|
||||||
|
String shadowed = env.get(blocked);
|
||||||
|
assertNotNull(shadowed, blocked + " must be explicitly overlaid, not left unmentioned");
|
||||||
|
assertFalse(shadowed.isBlank(), blocked + "'s overlay value must be non-blank");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* An {@code allow}-ed name must get NO overlay entry at all — any entry, blank or not, risks
|
||||||
|
* overriding the real value the pane needs, and the whole point of {@code allow} is that the
|
||||||
|
* pane's own inherited value passes through untouched.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void everyAllowedNameGetsNoOverlayEntrySoTheRealValuePassesThrough() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
serviceWithCredentials(herdr, TEST_MEMBER_CREDENTIALS).spawn();
|
||||||
|
|
||||||
|
Map<String, String> env = startEnv(herdr);
|
||||||
|
for (String allowed : List.of("AI_GATEWAY_TOKEN", "WORKER_GITEA_TOKEN", "CONTEXT7_TOKEN", "GITEA_HOST")) {
|
||||||
|
assertFalse(env.containsKey(allowed),
|
||||||
|
allowed + " is allow-listed — the launcher must not mention it at all");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* No {@code memberCredentials} configured (the pre-CB-596 constructor overloads still used
|
||||||
|
* throughout this file, and the shape a fresh {@code bridged.yaml} with no memberCredentials:
|
||||||
|
* block resolves to) blocks NOTHING. This documents the transitional gap rather than hiding it —
|
||||||
|
* see {@code HerdrPeerLauncher#applyMemberCredentialPolicy}'s javadoc.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void noMemberCredentialsConfiguredBlocksNothing() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
service(herdr, List.of("claude"), null).spawn();
|
||||||
|
|
||||||
|
assertNull(startEnv(herdr).get("GITEA_ACCESS_TOKEN"),
|
||||||
|
"with no memberCredentials configured, nothing is shadowed — config must supply the policy");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* No profile — present or future — may restore a blocked name by naming it in {@code env:}.
|
||||||
* The shadow is applied after the profile's own env in {@link HerdrPeerLauncher#baseEnv}
|
* The shadow is applied after the profile's own env in {@link HerdrPeerLauncher#baseEnv}
|
||||||
* precisely so this can never happen; this test pins that ordering.
|
* precisely so this can never happen; this test pins that ordering.
|
||||||
*/
|
*/
|
||||||
@Test
|
@Test
|
||||||
void aProfileEnvEntryCannotRestoreTheAdminGiteaAccessToken() {
|
void aProfileEnvEntryCannotRestoreABlockedName() {
|
||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||||
@@ -716,10 +839,48 @@ class ClaudeCodeLauncherTest {
|
|||||||
null, Map.of("GITEA_ACCESS_TOKEN", "admin-secret-from-profile-config"), null, null);
|
null, Map.of("GITEA_ACCESS_TOKEN", "admin-secret-from-profile-config"), null, null);
|
||||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||||
_ -> null).spawn();
|
_ -> null, 0, System::currentTimeMillis, () -> {}, null, () -> TEST_MEMBER_CREDENTIALS)
|
||||||
|
.spawn(cfg.profile(), null, null);
|
||||||
|
|
||||||
assertNotEquals("admin-secret-from-profile-config", startEnv(herdr).get("GITEA_ACCESS_TOKEN"),
|
assertNotEquals("admin-secret-from-profile-config", startEnv(herdr).get("GITEA_ACCESS_TOKEN"),
|
||||||
"a profile's own env: must not be able to smuggle the admin token back in");
|
"a profile's own env: must not be able to smuggle a blocked name back in");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-596 criterion 4: a credential-shaped host env var name on neither {@code known} nor
|
||||||
|
* {@code allow} is not silently allowed — it must be reported (never its value). This pins the
|
||||||
|
* WARN naming the gap, using an injected host-env-names source rather than the real
|
||||||
|
* {@code System.getenv()} so the test is deterministic.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void aCredentialShapedNameOnNeitherListIsLoggedAsAGap() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||||
|
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||||
|
List.of("claude"), "tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
|
||||||
|
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||||
|
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||||
|
_ -> null, 0, System::currentTimeMillis, () -> {}, null, () -> TEST_MEMBER_CREDENTIALS,
|
||||||
|
() -> Set.of("PATH", "HOME", "AI_GATEWAY_TOKEN", "A_BRAND_NEW_SECRET_TOKEN"));
|
||||||
|
|
||||||
|
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
|
||||||
|
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||||
|
appender.start();
|
||||||
|
logger.addAppender(appender);
|
||||||
|
try {
|
||||||
|
svc.spawn();
|
||||||
|
} finally {
|
||||||
|
logger.detachAppender(appender);
|
||||||
|
}
|
||||||
|
|
||||||
|
assertTrue(appender.list.stream().anyMatch(e ->
|
||||||
|
e.getFormattedMessage().contains("memberCredentials gap")
|
||||||
|
&& e.getFormattedMessage().contains("A_BRAND_NEW_SECRET_TOKEN")),
|
||||||
|
"the gap must name the unrecognized credential-shaped var, never a value");
|
||||||
|
assertFalse(appender.list.stream().anyMatch(e -> e.getFormattedMessage().contains("PATH")),
|
||||||
|
"PATH/HOME are not credential-shaped and must not be reported as a gap");
|
||||||
|
assertFalse(appender.list.stream().anyMatch(e -> e.getFormattedMessage().contains("AI_GATEWAY_TOKEN")),
|
||||||
|
"a name already on allow: is covered, not a gap");
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
|
|||||||
@@ -8,6 +8,7 @@ import dev.ltms.bridged.herdr.FakeHerdr;
|
|||||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||||
import dev.ltms.bridged.peer.Capability;
|
import dev.ltms.bridged.peer.Capability;
|
||||||
import dev.ltms.bridged.peer.CharterReceipt;
|
import dev.ltms.bridged.peer.CharterReceipt;
|
||||||
|
import dev.ltms.bridged.peer.MemberRole;
|
||||||
import dev.ltms.bridged.peer.PeerHandle;
|
import dev.ltms.bridged.peer.PeerHandle;
|
||||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||||
import dev.ltms.bridged.peer.SpawnRequest;
|
import dev.ltms.bridged.peer.SpawnRequest;
|
||||||
@@ -52,6 +53,14 @@ class OpenCodeLauncherTest {
|
|||||||
0, System::currentTimeMillis, () -> { }, configRoot, configRoot, fleet);
|
0, System::currentTimeMillis, () -> { }, configRoot, configRoot, fleet);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private static OpenCodeLauncher serviceWithCredentials(FakeHerdr herdr, Path configRoot,
|
||||||
|
BridgedConfig.Profile cfg,
|
||||||
|
BridgedConfig.MemberCredentials creds) {
|
||||||
|
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||||
|
Map.of(cfg.profile(), cfg), cfg.profile(), k -> "GITEA_ACCESS_TOKEN".equals(k) ? "tok" : null,
|
||||||
|
0, System::currentTimeMillis, () -> { }, configRoot, configRoot, null, () -> creds);
|
||||||
|
}
|
||||||
|
|
||||||
@SuppressWarnings("unchecked")
|
@SuppressWarnings("unchecked")
|
||||||
private static Map<String, Object> lastStart(FakeHerdr herdr) {
|
private static Map<String, Object> lastStart(FakeHerdr herdr) {
|
||||||
return (Map<String, Object>) herdr.lastCall("agent.start").params();
|
return (Map<String, Object>) herdr.lastCall("agent.start").params();
|
||||||
@@ -200,6 +209,35 @@ class OpenCodeLauncherTest {
|
|||||||
"--auto is unconditional: a model-less worker still must never block on approval");
|
"--auto is unconditional: a model-less worker still must never block on approval");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- CB-617: --agent <role> when the role has an agent-definition file --------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void agentFlagIsPassedWhenTheRoleAgentDefinitionFileExists(@TempDir Path root) throws Exception {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Path agentsDir = Files.createDirectories(root.resolve(".opencode/agent"));
|
||||||
|
Files.writeString(agentsDir.resolve("dev.md"), "You are a dev.");
|
||||||
|
OpenCodeLauncher svc = service(herdr, root, opencodeCfg(null, null, null));
|
||||||
|
|
||||||
|
svc.spawn(new SpawnRequest(null, root.toString(), null, null, null, MemberRole.DEV));
|
||||||
|
|
||||||
|
List<String> args = startArgs(herdr);
|
||||||
|
int flag = args.indexOf("--agent");
|
||||||
|
assertTrue(flag >= 0, "--agent is passed when the role's agent file exists: " + args);
|
||||||
|
assertEquals("dev", args.get(flag + 1));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void noAgentFlagWhenTheRoleAgentDefinitionFileIsAbsent(@TempDir Path root) {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
OpenCodeLauncher svc = service(herdr, root, opencodeCfg(null, null, null));
|
||||||
|
|
||||||
|
PeerHandle handle = svc.spawn(new SpawnRequest(null, root.toString(), null, null, null, MemberRole.DEV));
|
||||||
|
|
||||||
|
assertNotNull(handle, "the member still spawns with no agent-definition file");
|
||||||
|
assertFalse(startArgs(herdr).contains("--agent"),
|
||||||
|
"no --agent flag when the role has no agent-definition file");
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void injectsForgeTokenWhenProfileGrantsIt(@TempDir Path root) {
|
void injectsForgeTokenWhenProfileGrantsIt(@TempDir Path root) {
|
||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
@@ -209,18 +247,33 @@ class OpenCodeLauncherTest {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* CB-592: the shadow lives in {@link HerdrPeerLauncher#baseEnv}, shared by every adapter — this
|
* CB-596: the config-driven shadow lives in {@link HerdrPeerLauncher#baseEnv}, shared by every
|
||||||
* pins that the opencode path gets it too, not just Claude's. See the matching test in
|
* adapter — this pins that the opencode path gets it too, not just Claude's. See the matching
|
||||||
* {@code ClaudeCodeLauncherTest} for the full rationale (gitea #77).
|
* tests in {@code ClaudeCodeLauncherTest} for the full rationale (gitea #82, superseding CB-592's
|
||||||
|
* single hardcoded name).
|
||||||
*/
|
*/
|
||||||
@Test
|
@Test
|
||||||
void everySpawnShadowsTheAdminGiteaAccessToken(@TempDir Path root) {
|
void aKnownNameNotAllowedIsShadowedWithTheSentinel(@TempDir Path root) {
|
||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
service(herdr, root, opencodeCfg(null, null, null)).spawn();
|
BridgedConfig.MemberCredentials creds = new BridgedConfig.MemberCredentials(
|
||||||
|
null, List.of("AI_GATEWAY_TOKEN"), List.of("AI_GATEWAY_TOKEN", "GITEA_ACCESS_TOKEN"));
|
||||||
|
serviceWithCredentials(herdr, root, opencodeCfg(null, null, null), creds).spawn();
|
||||||
|
|
||||||
String shadowed = startEnv(herdr).get("GITEA_ACCESS_TOKEN");
|
String shadowed = startEnv(herdr).get("GITEA_ACCESS_TOKEN");
|
||||||
assertNotNull(shadowed, "GITEA_ACCESS_TOKEN must be explicitly overlaid, not left unmentioned");
|
assertNotNull(shadowed, "GITEA_ACCESS_TOKEN must be explicitly overlaid, not left unmentioned");
|
||||||
assertFalse(shadowed.isBlank(), "a blank overlay value's override behaviour is unverified — must be non-blank");
|
assertFalse(shadowed.isBlank(), "a blank overlay value's override behaviour is unverified — must be non-blank");
|
||||||
|
assertFalse(startEnv(herdr).containsKey("AI_GATEWAY_TOKEN"),
|
||||||
|
"an allow-listed name must get no overlay entry at all");
|
||||||
|
}
|
||||||
|
|
||||||
|
/** No {@code memberCredentials} configured (the pre-CB-596 constructor overloads) blocks nothing. */
|
||||||
|
@Test
|
||||||
|
void noMemberCredentialsConfiguredBlocksNothing(@TempDir Path root) {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
service(herdr, root, opencodeCfg(null, null, null)).spawn();
|
||||||
|
|
||||||
|
assertNull(startEnv(herdr).get("GITEA_ACCESS_TOKEN"),
|
||||||
|
"with no memberCredentials configured, nothing is shadowed — config must supply the policy");
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
|
|||||||
@@ -16,6 +16,7 @@ import java.util.List;
|
|||||||
import java.util.concurrent.TimeUnit;
|
import java.util.concurrent.TimeUnit;
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -166,6 +167,88 @@ class AmqpReplyInboxContractTest {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void prefetchBoundsTheHeldBacklog() throws Exception {
|
||||||
|
String target = "worker-prefetch-" + System.nanoTime();
|
||||||
|
int prefetch = 4;
|
||||||
|
int published = 10;
|
||||||
|
try (AmqpReplyInbox inbox = new AmqpReplyInbox(newConnection(), prefetch);
|
||||||
|
Connection inspect = newConnection()) {
|
||||||
|
inbox.own(target);
|
||||||
|
for (int i = 0; i < published; i++) {
|
||||||
|
inbox.publish(target, "m" + i, "payload " + i);
|
||||||
|
}
|
||||||
|
awaitHeldAtLeast(inbox, target, prefetch);
|
||||||
|
|
||||||
|
long depth;
|
||||||
|
try (Channel ch = inspect.createChannel()) {
|
||||||
|
depth = ch.queueDeclarePassive(queueName(target)).getMessageCount();
|
||||||
|
}
|
||||||
|
assertTrue(depth >= published - prefetch,
|
||||||
|
"broker should still hold at least " + (published - prefetch)
|
||||||
|
+ " undelivered messages behind a prefetch of " + prefetch + ", saw " + depth);
|
||||||
|
|
||||||
|
drainUntilEmpty(inbox, target, inspect);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void unroutablePublishReportsFailureNotSilentSuccess() throws Exception {
|
||||||
|
String target = "worker-unroutable-" + System.nanoTime();
|
||||||
|
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
|
||||||
|
// Deliberately never own(target): the queue is never declared, so the default-exchange
|
||||||
|
// route to agent.<target>.inbox does not exist and the broker must return the publish.
|
||||||
|
IllegalStateException ex = assertThrows(IllegalStateException.class,
|
||||||
|
() -> inbox.publish(target, "m1", "nobody home"));
|
||||||
|
assertTrue(ex.getMessage() != null && ex.getMessage().toLowerCase().contains("unroutable"),
|
||||||
|
"expected an unroutable-publish failure, got: " + ex.getMessage());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void confirmedPublishDeliversNormally() throws Exception {
|
||||||
|
String target = "worker-confirm-" + System.nanoTime();
|
||||||
|
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
|
||||||
|
inbox.own(target);
|
||||||
|
inbox.publish(target, "m1", "confirmed delivery"); // must return normally: routed and confirmed
|
||||||
|
|
||||||
|
List<ReplyInbox.InboxMessage> got = awaitPeek(inbox, target);
|
||||||
|
assertEquals(1, got.size());
|
||||||
|
assertEquals("confirmed delivery", got.getFirst().content());
|
||||||
|
inbox.ack(target, "m1");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Poll peek until at least {@code n} replies for {@code target} are held, or ~10s elapse. */
|
||||||
|
@SuppressWarnings("BusyWait")
|
||||||
|
private static void awaitHeldAtLeast(AmqpReplyInbox inbox, String target, int n) throws InterruptedException {
|
||||||
|
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(10);
|
||||||
|
while (inbox.peek(target).size() < n && System.nanoTime() < deadline) {
|
||||||
|
Thread.sleep(50);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Repeatedly ack whatever is currently held (freeing prefetch slots for the next delivery) until
|
||||||
|
* both the local snapshot and the broker's own queue depth are empty, or ~10s elapse.
|
||||||
|
*/
|
||||||
|
@SuppressWarnings("BusyWait")
|
||||||
|
private static void drainUntilEmpty(AmqpReplyInbox inbox, String target, Connection inspect) throws Exception {
|
||||||
|
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(10);
|
||||||
|
while (System.nanoTime() < deadline) {
|
||||||
|
for (ReplyInbox.InboxMessage msg : inbox.peek(target)) {
|
||||||
|
inbox.ack(target, msg.msgId());
|
||||||
|
}
|
||||||
|
try (Channel ch = inspect.createChannel()) {
|
||||||
|
if (ch.queueDeclarePassive(queueName(target)).getMessageCount() == 0 && inbox.peek(target).isEmpty()) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
Thread.sleep(50);
|
||||||
|
}
|
||||||
|
throw new AssertionError("did not drain " + target + " to empty within the deadline");
|
||||||
|
}
|
||||||
|
|
||||||
/** Poll peek (broker delivery is async) until a reply for {@code target} appears or ~10s elapse. */
|
/** Poll peek (broker delivery is async) until a reply for {@code target} appears or ~10s elapse. */
|
||||||
@SuppressWarnings("BusyWait") // deliberate poll for async broker delivery, bounded by the deadline
|
@SuppressWarnings("BusyWait") // deliberate poll for async broker delivery, bounded by the deadline
|
||||||
private static List<ReplyInbox.InboxMessage> awaitPeek(AmqpReplyInbox inbox, String target)
|
private static List<ReplyInbox.InboxMessage> awaitPeek(AmqpReplyInbox inbox, String target)
|
||||||
|
|||||||
@@ -0,0 +1,288 @@
|
|||||||
|
package dev.ltms.bridged.msg;
|
||||||
|
|
||||||
|
import com.rabbitmq.client.AMQP;
|
||||||
|
import com.rabbitmq.client.Channel;
|
||||||
|
import com.rabbitmq.client.ConfirmCallback;
|
||||||
|
import com.rabbitmq.client.Connection;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.Timeout;
|
||||||
|
|
||||||
|
import java.lang.reflect.InvocationHandler;
|
||||||
|
import java.lang.reflect.Proxy;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.concurrent.CopyOnWriteArrayList;
|
||||||
|
import java.util.concurrent.CountDownLatch;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.atomic.AtomicInteger;
|
||||||
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
import java.util.concurrent.atomic.AtomicReference;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-528 follow-up: {@link AmqpReplyInbox#failPendingPublishesOnRecovery()} must not fail a publish
|
||||||
|
* that registers concurrently with (and was not yet visible when) the recovery sweep began — that
|
||||||
|
* would report a publish that actually succeeded as failed, and {@code MessageService.reply} retries
|
||||||
|
* with a fresh {@code msgId}, so the reply is delivered twice. A live broker reconnect cannot be
|
||||||
|
* forced reliably, so this drives {@link AmqpReplyInbox#failPendingPublishesOnRecovery} and a real
|
||||||
|
* {@link AmqpReplyInbox#publish} against each other directly, against fake AMQP channels built with
|
||||||
|
* {@link Proxy} (no mocking library is on the classpath).
|
||||||
|
*
|
||||||
|
* <p>Also covers Finding 2 (CB-528 follow-up): {@link AmqpReplyInbox#close()} must fail an in-flight
|
||||||
|
* publish promptly instead of leaving it to idle out the 10s confirm timeout.
|
||||||
|
*/
|
||||||
|
class AmqpReplyInboxRecoveryRaceTest {
|
||||||
|
|
||||||
|
/** Large enough that thousands of entries are still unprocessed by the time the very first one
|
||||||
|
* is observed as failed (see {@code sweepIsHoldingTheLock} below) — that gap is what makes the
|
||||||
|
* head start deterministic instead of a coin flip. 100,000 gave the same guarantee but made the
|
||||||
|
* test far more expensive than the guarantee needs; the ordering no longer depends on a timing
|
||||||
|
* window sized to the full backlog; just to the tail of it. */
|
||||||
|
private static final int STALE_PUBLISHES = 2_000;
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@Timeout(30)
|
||||||
|
void recoverySweepDoesNotFailAPublishThatRegistersWhileItIsRunning() throws Exception {
|
||||||
|
AtomicLong seqCounter = new AtomicLong();
|
||||||
|
List<Long> seqOrder = new CopyOnWriteArrayList<>();
|
||||||
|
List<String> msgIdOrder = new CopyOnWriteArrayList<>();
|
||||||
|
AtomicReference<ConfirmCallback> ackCallback = new AtomicReference<>();
|
||||||
|
AtomicReference<ConfirmCallback> nackCallback = new AtomicReference<>();
|
||||||
|
|
||||||
|
Channel publishChannel = fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback);
|
||||||
|
Channel consumeChannel = fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback);
|
||||||
|
Connection connection = fakeConnection(consumeChannel, publishChannel);
|
||||||
|
|
||||||
|
AmqpReplyInbox inbox = new AmqpReplyInbox(connection, AmqpReplyInbox.DEFAULT_PREFETCH);
|
||||||
|
|
||||||
|
// STALE_PUBLISHES in-flight publishes that never get confirmed — they sit in pendingBySeq /
|
||||||
|
// pendingByMsgId exactly like publishes whose confirm never arrived before a connection drop.
|
||||||
|
// Virtual threads make this many concurrent blocking publish() calls cheap.
|
||||||
|
CountDownLatch staleStarted = new CountDownLatch(STALE_PUBLISHES);
|
||||||
|
// Counted down by the FIRST stale publish thread to observe its own failure. That can only
|
||||||
|
// happen from inside failPendingPublishesOnRecovery() — nothing else in this test ever
|
||||||
|
// completes a stale Pending exceptionally (no nack/return is simulated for any "stale-*"
|
||||||
|
// msgId) — so seeing it fire is direct, observable proof the sweep is inside its loop, not a
|
||||||
|
// timing guess. It replaces the old fixed Thread.sleep(5) head start.
|
||||||
|
CountDownLatch sweepIsHoldingTheLock = new CountDownLatch(1);
|
||||||
|
for (int i = 0; i < STALE_PUBLISHES; i++) {
|
||||||
|
String msgId = "stale-" + i;
|
||||||
|
Thread.ofVirtual().start(() -> {
|
||||||
|
staleStarted.countDown();
|
||||||
|
try {
|
||||||
|
inbox.publish("worker-stale", msgId, "x");
|
||||||
|
} catch (IllegalStateException expected) {
|
||||||
|
sweepIsHoldingTheLock.countDown();
|
||||||
|
}
|
||||||
|
});
|
||||||
|
}
|
||||||
|
staleStarted.await();
|
||||||
|
// Let the registrations (the synchronized put into pendingBySeq/pendingByMsgId) actually land
|
||||||
|
// for all of them before the sweep starts, so the sweep begins with a large, real backlog.
|
||||||
|
Thread.sleep(300);
|
||||||
|
|
||||||
|
AtomicReference<Throwable> sweepError = new AtomicReference<>();
|
||||||
|
Thread sweepThread = new Thread(() -> {
|
||||||
|
try {
|
||||||
|
inbox.failPendingPublishesOnRecovery();
|
||||||
|
} catch (Throwable t) {
|
||||||
|
sweepError.set(t);
|
||||||
|
}
|
||||||
|
}, "recovery-sweep");
|
||||||
|
sweepThread.start();
|
||||||
|
|
||||||
|
// Deterministic head start: block until the sweep has actually failed one of the stale
|
||||||
|
// publishes. failPendingPublishesOnRecovery() (once guarded, as it is on main) holds
|
||||||
|
// publishChannelLock for its ENTIRE loop, not just per entry — so this failure proves the
|
||||||
|
// sweep is, at this instant, still holding that lock. With STALE_PUBLISHES this large, the
|
||||||
|
// remaining ~1,999 entries give an enormous margin between "first failure observed" and "sweep
|
||||||
|
// releases the lock": there is no window left for "fresh" to slip in before the sweep starts,
|
||||||
|
// or to win the lock ahead of it — see the case-2 note below. This also means Case 1 (the sweep
|
||||||
|
// is already inside its loop, holding the lock, when "fresh" tries to register) is now
|
||||||
|
// guaranteed by construction rather than merely likely under a fixed sleep.
|
||||||
|
assertTrue(sweepIsHoldingTheLock.await(20, TimeUnit.SECONDS),
|
||||||
|
"the sweep never failed a single stale publish — it may not have started");
|
||||||
|
|
||||||
|
// Case 2 ("fresh" wins publishChannelLock before the sweep even starts, so it genuinely
|
||||||
|
// published on the stale channel and the sweep correctly fails it) is impossible by
|
||||||
|
// construction in this test: freshThread.start() below is reached only after
|
||||||
|
// sweepIsHoldingTheLock has counted down, which can only happen once
|
||||||
|
// failPendingPublishesOnRecovery() is already running and has already failed a stale entry.
|
||||||
|
// There is no code path that lets "fresh" start before the sweep starts. That case is real
|
||||||
|
// and correct production behaviour (see AmqpReplyInbox#failPendingPublishesOnRecovery's
|
||||||
|
// javadoc), it is just not reachable from this deterministic ordering, so it does not need a
|
||||||
|
// separate assertion here.
|
||||||
|
|
||||||
|
// This is the exact interleaving CB-528's follow-up describes: "the still-running recovery
|
||||||
|
// sweep" racing a publish that registers while it is mid-flight.
|
||||||
|
AtomicReference<Throwable> publishError = new AtomicReference<>();
|
||||||
|
Thread freshThread = new Thread(() -> {
|
||||||
|
try {
|
||||||
|
inbox.publish("worker-fresh", "fresh", "hello");
|
||||||
|
} catch (Throwable t) {
|
||||||
|
publishError.set(t);
|
||||||
|
}
|
||||||
|
}, "fresh-publish");
|
||||||
|
freshThread.start();
|
||||||
|
|
||||||
|
sweepThread.join(20_000);
|
||||||
|
assertNull(sweepError.get(), "sweep threw: " + sweepError.get());
|
||||||
|
|
||||||
|
// Simulate the broker's real confirm for "fresh" now that the sweep is done, so a correct
|
||||||
|
// implementation's publish() returns normally instead of idling out CONFIRM_TIMEOUT_MS. Poll
|
||||||
|
// for the registration rather than checking once: freshThread may still be contending for
|
||||||
|
// publishChannelLock (behind the sweep's own hold on it, and possibly other stale threads
|
||||||
|
// still unwinding) even though the sweep itself has already finished.
|
||||||
|
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(9);
|
||||||
|
int idx = -1;
|
||||||
|
while (idx < 0 && System.nanoTime() < deadline) {
|
||||||
|
idx = msgIdOrder.indexOf("fresh");
|
||||||
|
if (idx < 0) {
|
||||||
|
Thread.sleep(20);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (idx >= 0 && ackCallback.get() != null) {
|
||||||
|
ackCallback.get().handle(seqOrder.get(idx), false);
|
||||||
|
}
|
||||||
|
freshThread.join(15_000);
|
||||||
|
|
||||||
|
assertNull(publishError.get(),
|
||||||
|
"a publish that registered while the recovery sweep was running must not be failed by "
|
||||||
|
+ "it, but got: " + publishError.get());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@Timeout(15)
|
||||||
|
void closeFailsInFlightPublishPromptlyInsteadOfWaitingOutTheConfirmTimeout() throws Exception {
|
||||||
|
AtomicLong seqCounter = new AtomicLong();
|
||||||
|
List<Long> seqOrder = new CopyOnWriteArrayList<>();
|
||||||
|
List<String> msgIdOrder = new CopyOnWriteArrayList<>();
|
||||||
|
AtomicReference<ConfirmCallback> ackCallback = new AtomicReference<>();
|
||||||
|
AtomicReference<ConfirmCallback> nackCallback = new AtomicReference<>();
|
||||||
|
|
||||||
|
Channel publishChannel = fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback);
|
||||||
|
Channel consumeChannel = fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback);
|
||||||
|
Connection connection = fakeConnection(consumeChannel, publishChannel);
|
||||||
|
|
||||||
|
AmqpReplyInbox inbox = new AmqpReplyInbox(connection, AmqpReplyInbox.DEFAULT_PREFETCH);
|
||||||
|
|
||||||
|
CountDownLatch publishReturned = new CountDownLatch(1);
|
||||||
|
AtomicReference<Throwable> publishError = new AtomicReference<>();
|
||||||
|
AtomicLong elapsedMillis = new AtomicLong();
|
||||||
|
Thread publishThread = new Thread(() -> {
|
||||||
|
long start = System.nanoTime();
|
||||||
|
try {
|
||||||
|
inbox.publish("worker-close", "never-confirmed", "x");
|
||||||
|
} catch (Throwable t) {
|
||||||
|
publishError.set(t);
|
||||||
|
} finally {
|
||||||
|
elapsedMillis.set((System.nanoTime() - start) / 1_000_000);
|
||||||
|
publishReturned.countDown();
|
||||||
|
}
|
||||||
|
}, "publish-during-close");
|
||||||
|
publishThread.start();
|
||||||
|
|
||||||
|
Thread.sleep(200); // let publish() register before close() runs
|
||||||
|
inbox.close();
|
||||||
|
|
||||||
|
assertTrue(publishReturned.await(5, TimeUnit.SECONDS), "publish() did not return after close()");
|
||||||
|
assertNotNull(publishError.get(), "a publish in flight when close() runs must fail, not hang");
|
||||||
|
assertTrue(publishError.get().getMessage() != null
|
||||||
|
&& publishError.get().getMessage().toLowerCase().contains("closed"),
|
||||||
|
"expected a clear closed-inbox message, got: " + publishError.get());
|
||||||
|
assertTrue(elapsedMillis.get() < 5_000,
|
||||||
|
"close() should fail the in-flight publish promptly, not wait out the confirm timeout — took "
|
||||||
|
+ elapsedMillis.get() + "ms");
|
||||||
|
}
|
||||||
|
|
||||||
|
/** A {@link Proxy}-backed {@link Channel}: only the calls {@link AmqpReplyInbox} actually makes
|
||||||
|
* are meaningfully implemented; everything else returns a harmless default. */
|
||||||
|
private static Channel fakeChannel(AtomicLong seqCounter, List<Long> seqOrder, List<String> msgIdOrder,
|
||||||
|
AtomicReference<ConfirmCallback> ackCallback,
|
||||||
|
AtomicReference<ConfirmCallback> nackCallback) {
|
||||||
|
InvocationHandler handler = (proxy, method, args) -> {
|
||||||
|
String name = method.getName();
|
||||||
|
if (name.equals("getNextPublishSeqNo")) {
|
||||||
|
long value = seqCounter.incrementAndGet();
|
||||||
|
seqOrder.add(value);
|
||||||
|
return value;
|
||||||
|
}
|
||||||
|
if (name.equals("basicPublish")) {
|
||||||
|
AMQP.BasicProperties props = (AMQP.BasicProperties) args[3];
|
||||||
|
msgIdOrder.add(props.getMessageId());
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
if (name.equals("addConfirmListener")) {
|
||||||
|
ackCallback.set((ConfirmCallback) args[0]);
|
||||||
|
nackCallback.set((ConfirmCallback) args[1]);
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
if (name.equals("equals")) {
|
||||||
|
return proxy == args[0];
|
||||||
|
}
|
||||||
|
if (name.equals("hashCode")) {
|
||||||
|
return System.identityHashCode(proxy);
|
||||||
|
}
|
||||||
|
if (name.equals("toString")) {
|
||||||
|
return "FakeChannel";
|
||||||
|
}
|
||||||
|
return defaultValue(method.getReturnType());
|
||||||
|
};
|
||||||
|
return (Channel) Proxy.newProxyInstance(AmqpReplyInboxRecoveryRaceTest.class.getClassLoader(),
|
||||||
|
new Class<?>[] {Channel.class}, handler);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** A {@link Proxy}-backed {@link Connection} handing out {@code first} then {@code second} from
|
||||||
|
* successive {@code createChannel()} calls, matching {@link AmqpReplyInbox}'s constructor. */
|
||||||
|
private static Connection fakeConnection(Channel first, Channel second) {
|
||||||
|
AtomicInteger calls = new AtomicInteger();
|
||||||
|
InvocationHandler handler = (proxy, method, args) -> {
|
||||||
|
String name = method.getName();
|
||||||
|
if (name.equals("createChannel") && (args == null || args.length == 0)) {
|
||||||
|
return calls.getAndIncrement() == 0 ? first : second;
|
||||||
|
}
|
||||||
|
if (name.equals("equals")) {
|
||||||
|
return proxy == args[0];
|
||||||
|
}
|
||||||
|
if (name.equals("hashCode")) {
|
||||||
|
return System.identityHashCode(proxy);
|
||||||
|
}
|
||||||
|
if (name.equals("toString")) {
|
||||||
|
return "FakeConnection";
|
||||||
|
}
|
||||||
|
return defaultValue(method.getReturnType());
|
||||||
|
};
|
||||||
|
return (Connection) Proxy.newProxyInstance(AmqpReplyInboxRecoveryRaceTest.class.getClassLoader(),
|
||||||
|
new Class<?>[] {Connection.class}, handler);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static Object defaultValue(Class<?> type) {
|
||||||
|
if (!type.isPrimitive() || type == void.class) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
if (type == boolean.class) {
|
||||||
|
return Boolean.FALSE;
|
||||||
|
}
|
||||||
|
if (type == long.class) {
|
||||||
|
return 0L;
|
||||||
|
}
|
||||||
|
if (type == short.class) {
|
||||||
|
return (short) 0;
|
||||||
|
}
|
||||||
|
if (type == byte.class) {
|
||||||
|
return (byte) 0;
|
||||||
|
}
|
||||||
|
if (type == char.class) {
|
||||||
|
return (char) 0;
|
||||||
|
}
|
||||||
|
if (type == double.class) {
|
||||||
|
return 0.0d;
|
||||||
|
}
|
||||||
|
if (type == float.class) {
|
||||||
|
return 0.0f;
|
||||||
|
}
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -714,6 +714,47 @@ class MessageServiceTest {
|
|||||||
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
|
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- CB-582: bridge_status pendingAsk() ------------------------------------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void pendingAskReturnsNullWhenNoQuestionIsOpen() throws Exception {
|
||||||
|
assertNull(messages.pendingAsk(T), "no async ticket at all -> no pending ask");
|
||||||
|
|
||||||
|
String ticket = messages.sendAsync(T, "long task");
|
||||||
|
awaitWaiting();
|
||||||
|
assertNull(messages.pendingAsk(T), "a plain pending delegation is not a question");
|
||||||
|
|
||||||
|
injectDelivery();
|
||||||
|
assertTrue(rendezvous.resolve(T, "done"));
|
||||||
|
awaitTicketPhase(ticket, MessageService.Phase.DONE);
|
||||||
|
assertNull(messages.pendingAsk(T), "a finished ticket carries no open question either");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void pendingAskReturnsTheOpenQuestionForAnAsyncTicket() throws Exception {
|
||||||
|
String ticket = messages.sendAsync(T, "task that asks");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
|
||||||
|
CompletableFuture<MessageService.AskResult> ask =
|
||||||
|
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||||
|
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
|
||||||
|
|
||||||
|
MessageService.PendingAsk pending = messages.pendingAsk(T);
|
||||||
|
assertNotNull(pending, "bridge_status should see the open question");
|
||||||
|
assertEquals(ticket, pending.ticket());
|
||||||
|
assertEquals("which config file?", pending.question());
|
||||||
|
assertEquals(asking.turnId(), pending.turnId());
|
||||||
|
|
||||||
|
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||||
|
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
|
||||||
|
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||||
|
assertNull(messages.pendingAsk(T), "an answered question is no longer pending");
|
||||||
|
awaitWaiting();
|
||||||
|
assertTrue(rendezvous.resolve(T, "done"));
|
||||||
|
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
|
||||||
|
}
|
||||||
|
|
||||||
// --- CB-588: async ticket terminal nudges ---------------------------------------------------
|
// --- CB-588: async ticket terminal nudges ---------------------------------------------------
|
||||||
//
|
//
|
||||||
// MessageService.reply's rendezvous fast path is exactly what an async ticket always takes
|
// MessageService.reply's rendezvous fast path is exactly what an async ticket always takes
|
||||||
@@ -833,6 +874,110 @@ class MessageServiceTest {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- CB-582: bridge_ask question-open nudges --------------------------------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void anAsyncTicketThatPausesOnAQuestionNudgesTheLeadWithNoPriorPollCall() throws Exception {
|
||||||
|
try (var wiring = wireWithPushLoop(5, 50)) {
|
||||||
|
String ticket = wiring.service().sendAsync(T, "task that asks");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
|
||||||
|
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
|
||||||
|
() -> wiring.service().ask(T, "which config file?", 5000));
|
||||||
|
MessageService.TaskView asking = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.ASKING);
|
||||||
|
|
||||||
|
awaitNudge(wiring.leadHerdr());
|
||||||
|
String nudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
|
||||||
|
assertTrue(nudge.contains(ticket), "the nudge should name the ticket: " + nudge);
|
||||||
|
assertTrue(nudge.contains(asking.turnId()), "the nudge should name the turnId: " + nudge);
|
||||||
|
assertTrue(nudge.contains("bridge_send(turnId="),
|
||||||
|
"the nudge should name the exact answer call: " + nudge);
|
||||||
|
assertTrue(nudge.contains("which config file?"), "the nudge should include the question: " + nudge);
|
||||||
|
|
||||||
|
// Clean up the still-open ask so the background thread does not linger past the test.
|
||||||
|
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||||
|
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
|
||||||
|
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||||
|
awaitWaiting();
|
||||||
|
assertTrue(rendezvous.resolve(T, "done"));
|
||||||
|
answer.get(5, TimeUnit.SECONDS);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void answeringAQuestionStopsFurtherNudgesAboutIt() throws Exception {
|
||||||
|
try (var wiring = wireWithPushLoop(5, 50)) {
|
||||||
|
String ticket = wiring.service().sendAsync(T, "task that asks");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
|
||||||
|
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
|
||||||
|
() -> wiring.service().ask(T, "which config file?", 5000));
|
||||||
|
MessageService.TaskView asking = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.ASKING);
|
||||||
|
|
||||||
|
awaitNudge(wiring.leadHerdr());
|
||||||
|
long callsBeforeAnswer = wiring.leadHerdr().calls.stream()
|
||||||
|
.filter(c -> c.method().equals("agent.prompt")).count();
|
||||||
|
|
||||||
|
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||||
|
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
|
||||||
|
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||||
|
awaitWaiting();
|
||||||
|
assertTrue(rendezvous.resolve(T, "done"));
|
||||||
|
answer.get(5, TimeUnit.SECONDS);
|
||||||
|
|
||||||
|
// Let several more ticks (and the ticket's own now-legitimate terminal nudge) fire —
|
||||||
|
// none of them may still name the question's turnId, which is closed.
|
||||||
|
Thread.sleep(300);
|
||||||
|
boolean anyNamesClosedQuestion = wiring.leadHerdr().calls.stream()
|
||||||
|
.filter(c -> c.method().equals("agent.prompt"))
|
||||||
|
.skip(callsBeforeAnswer)
|
||||||
|
.anyMatch(c -> c.params().toString().contains(asking.turnId()));
|
||||||
|
assertFalse(anyNamesClosedQuestion,
|
||||||
|
"no nudge sent after the answer may still name the now-closed turnId " + asking.turnId());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void anAskThatLeavesByThrowingStillClosesItsQuestion() throws Exception {
|
||||||
|
// CB-582 follow-up. ask() calls clearAsyncQuestion on three paths — no-waiter, timed out,
|
||||||
|
// and (from answer()) answered — but it can also leave by *throwing*: an interrupt while
|
||||||
|
// blocked on the answer, or an ExecutionException from the answer future. Those paths run
|
||||||
|
// only the finally block, so before the fix the push loop kept the question pending for
|
||||||
|
// good: named in every nudge until its own cap, then never removed from the map at all.
|
||||||
|
PrimaryRegistry registry = new PrimaryRegistry(null);
|
||||||
|
registry.recordDelegation(T, LEAD);
|
||||||
|
var scheduler = java.util.concurrent.Executors.newSingleThreadScheduledExecutor();
|
||||||
|
// A backoff far longer than the test: the schedule is started but no tick ever fires, so
|
||||||
|
// decide() is read directly and nothing here depends on timing.
|
||||||
|
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, new AgentControl(new FakeHerdr()), inbox,
|
||||||
|
scheduler, 5, 60_000);
|
||||||
|
MessageService service = new MessageService(agents, injector, rendezvous, inbox, pushLoop);
|
||||||
|
try {
|
||||||
|
String ticket = service.sendAsync(T, "task that asks");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
|
||||||
|
Thread asker = new Thread(() -> assertThrows(IllegalStateException.class,
|
||||||
|
() -> service.ask(T, "which config file?", 30_000)));
|
||||||
|
asker.start();
|
||||||
|
awaitTicketPhaseOn(service, ticket, MessageService.Phase.ASKING);
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, pushLoop.decide(LEAD, 0, 0, 0),
|
||||||
|
"the open question should be the one thing keeping this lead's schedule alive");
|
||||||
|
|
||||||
|
asker.interrupt();
|
||||||
|
asker.join(5000);
|
||||||
|
assertFalse(asker.isAlive(), "the interrupted ask should have left ask() by throwing");
|
||||||
|
|
||||||
|
assertEquals(ReplyPushLoop.Action.STOP, pushLoop.decide(LEAD, 0, 0, 0),
|
||||||
|
"an ask that threw must still close its question, or the loop nudges about it for good");
|
||||||
|
} finally {
|
||||||
|
service.close();
|
||||||
|
scheduler.shutdownNow();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void aFleetWithNoPushLoopConfiguredBehavesExactlyAsToday() throws Exception {
|
void aFleetWithNoPushLoopConfiguredBehavesExactlyAsToday() throws Exception {
|
||||||
// `messages` (the shared field) uses the no-pushLoop constructor — poll() must not throw,
|
// `messages` (the shared field) uses the no-pushLoop constructor — poll() must not throw,
|
||||||
|
|||||||
@@ -20,6 +20,7 @@ import java.util.concurrent.CountDownLatch;
|
|||||||
import java.util.concurrent.Executors;
|
import java.util.concurrent.Executors;
|
||||||
import java.util.concurrent.ScheduledExecutorService;
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
import java.util.concurrent.TimeUnit;
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.atomic.AtomicInteger;
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.*;
|
import static org.junit.jupiter.api.Assertions.*;
|
||||||
|
|
||||||
@@ -27,6 +28,12 @@ import static org.junit.jupiter.api.Assertions.*;
|
|||||||
* Unit tests for {@link ReplyPushLoop}: decision logic, nudge injection, idempotency,
|
* Unit tests for {@link ReplyPushLoop}: decision logic, nudge injection, idempotency,
|
||||||
* bounded reminders, and stop conditions.
|
* bounded reminders, and stop conditions.
|
||||||
*
|
*
|
||||||
|
* <p>CB-590 collapsed the CB-307 reply-nudge schedule and the CB-588 ticket-nudge schedule into
|
||||||
|
* one schedule per lead ({@link ReplyPushLoop#decide}), so most tests below register pending work
|
||||||
|
* through the public entry points ({@code onReplyQueued} / {@code onTicketTerminal}) before
|
||||||
|
* exercising {@code decide} directly, mirroring how the two entry points now share one decision
|
||||||
|
* function keyed by the lead terminal rather than by worker target.
|
||||||
|
*
|
||||||
* <p>Uses a {@link RecordingHerdrClient} that synchronizes access to its call list so the
|
* <p>Uses a {@link RecordingHerdrClient} that synchronizes access to its call list so the
|
||||||
* scheduler thread and test thread never have memory ordering issues. The {@code decide()}
|
* scheduler thread and test thread never have memory ordering issues. The {@code decide()}
|
||||||
* tests use a simple client with no concurrency concern.
|
* tests use a simple client with no concurrency concern.
|
||||||
@@ -35,6 +42,7 @@ class ReplyPushLoopTest {
|
|||||||
|
|
||||||
private static final String PRIMARY = "term_primary";
|
private static final String PRIMARY = "term_primary";
|
||||||
private static final String WORKER = "term_worker";
|
private static final String WORKER = "term_worker";
|
||||||
|
private static final String WORKER2 = "term_worker2";
|
||||||
private static final ObjectMapper MAPPER = new ObjectMapper();
|
private static final ObjectMapper MAPPER = new ObjectMapper();
|
||||||
|
|
||||||
private PrimaryRegistry registry;
|
private PrimaryRegistry registry;
|
||||||
@@ -58,38 +66,50 @@ class ReplyPushLoopTest {
|
|||||||
// --- decide() logic ------------------------------------------------------------------------
|
// --- decide() logic ------------------------------------------------------------------------
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideWithoutPrimaryIsStop() {
|
void onReplyQueuedWithNoKnownLeadNeverStartsASchedule() throws Exception {
|
||||||
agents = agentWithStatus("idle");
|
var rec = recordingClient();
|
||||||
var loop = new ReplyPushLoop(
|
agents = new AgentControl(rec);
|
||||||
new PrimaryRegistry(null), agents, inbox, scheduler, 5, 100);
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(WORKER, 0));
|
var loop = new ReplyPushLoop(new PrimaryRegistry(null), agents, inbox, scheduler, 5, 50);
|
||||||
|
|
||||||
|
loop.onReplyQueued(WORKER); // no lead known -> never registered, never scheduled
|
||||||
|
|
||||||
|
assertFalse(loop.isActive(), "no lead known means nothing to nudge yet");
|
||||||
|
Thread.sleep(150);
|
||||||
|
assertEquals(0, rec.sendCount(), "must not nudge when no lead is known to be waiting");
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideWithEmptyInboxIsStop() {
|
void decideWithNothingPendingIsStop() {
|
||||||
agents = agentWithStatus("idle");
|
agents = agentWithStatus("idle");
|
||||||
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
|
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(PRIMARY, 0, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideAtCapIsStop() {
|
void decideAtCapIsStop() {
|
||||||
agents = agentWithStatus("idle");
|
agents = agentWithStatus("idle");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100).decide(WORKER, 2));
|
var loop = loop(2, 100);
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 2, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideUnderCapWithInjectablePrimaryIsInject() {
|
void decideUnderCapWithInjectablePrimaryIsInject() {
|
||||||
agents = agentWithStatus("idle");
|
agents = agentWithStatus("idle");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
|
var loop = loop();
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideUnderCapWithBlockedPrimaryIsInject() {
|
void decideUnderCapWithBlockedPrimaryIsInject() {
|
||||||
agents = agentWithStatus("blocked");
|
agents = agentWithStatus("blocked");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
|
var loop = loop();
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0),
|
||||||
"BLOCKED is injectable");
|
"BLOCKED is injectable");
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -97,7 +117,9 @@ class ReplyPushLoopTest {
|
|||||||
void decideUnderCapWithDonePrimaryIsInject() {
|
void decideUnderCapWithDonePrimaryIsInject() {
|
||||||
agents = agentWithStatus("done");
|
agents = agentWithStatus("done");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
|
var loop = loop();
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0),
|
||||||
"DONE is injectable");
|
"DONE is injectable");
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -105,23 +127,29 @@ class ReplyPushLoopTest {
|
|||||||
void decideUnderCapWithBusyPrimaryIsWaitBusy() {
|
void decideUnderCapWithBusyPrimaryIsWaitBusy() {
|
||||||
agents = agentWithStatus("working");
|
agents = agentWithStatus("working");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
|
var loop = loop();
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decide(PRIMARY, 0, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideUnderCapWithUnknownPrimaryIsWaitBusy() {
|
void decideUnderCapWithUnknownPrimaryIsWaitBusy() {
|
||||||
agents = agentWithStatus("unknown");
|
agents = agentWithStatus("unknown");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
|
var loop = loop();
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decide(PRIMARY, 0, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideStopsAfterInboxIsEmptied() {
|
void decideStopsAfterInboxIsEmptied() {
|
||||||
agents = agentWithStatus("idle");
|
agents = agentWithStatus("idle");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
|
var loop = loop();
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0));
|
||||||
inbox.ack(WORKER, "m1");
|
inbox.ack(WORKER, "m1");
|
||||||
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- onReplyQueued integration -------------------------------------------------------------
|
// --- onReplyQueued integration -------------------------------------------------------------
|
||||||
@@ -201,12 +229,19 @@ class ReplyPushLoopTest {
|
|||||||
assertTrue(nudge.contains("bridge_poll(target=term_worker)"));
|
assertTrue(nudge.contains("bridge_poll(target=term_worker)"));
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- CB-588: async ticket terminal nudges — decideTickets() logic --------------------------
|
@Test
|
||||||
|
void repliesNudgeFormatIsCorrect() {
|
||||||
|
String multi = ReplyPushLoop.REPLIES_NUDGE_FORMAT.formatted(2, "term_worker1, term_worker2");
|
||||||
|
assertTrue(multi.contains("2 workers"));
|
||||||
|
assertTrue(multi.contains("bridge_poll(target=...)"));
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- CB-588: async ticket terminal nudges — decide() logic on tickets -----------------------
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void decideTicketsWithNothingPendingIsStop() {
|
void decideTicketsWithNothingPendingIsStop() {
|
||||||
agents = agentWithStatus("idle");
|
agents = agentWithStatus("idle");
|
||||||
assertEquals(ReplyPushLoop.Action.STOP, loop().decideTickets(PRIMARY, 0));
|
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(PRIMARY, 0, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -214,7 +249,7 @@ class ReplyPushLoopTest {
|
|||||||
agents = agentWithStatus("idle");
|
agents = agentWithStatus("idle");
|
||||||
var loop = loop(2, 100_000);
|
var loop = loop(2, 100_000);
|
||||||
loop.onTicketTerminal("task-1", WORKER, false);
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
assertEquals(ReplyPushLoop.Action.STOP, loop.decideTickets(PRIMARY, 2));
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 2));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -222,7 +257,7 @@ class ReplyPushLoopTest {
|
|||||||
agents = agentWithStatus("idle");
|
agents = agentWithStatus("idle");
|
||||||
var loop = loop(5, 100_000);
|
var loop = loop(5, 100_000);
|
||||||
loop.onTicketTerminal("task-1", WORKER, false);
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
assertEquals(ReplyPushLoop.Action.INJECT, loop.decideTickets(PRIMARY, 0));
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -230,7 +265,7 @@ class ReplyPushLoopTest {
|
|||||||
agents = agentWithStatus("working");
|
agents = agentWithStatus("working");
|
||||||
var loop = loop(5, 100_000);
|
var loop = loop(5, 100_000);
|
||||||
loop.onTicketTerminal("task-1", WORKER, false);
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decideTickets(PRIMARY, 0));
|
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decide(PRIMARY, 0, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -238,7 +273,7 @@ class ReplyPushLoopTest {
|
|||||||
agents = agentWithStatus("idle");
|
agents = agentWithStatus("idle");
|
||||||
var loop = new ReplyPushLoop(new PrimaryRegistry(null), agents, inbox, scheduler, 5, 100_000);
|
var loop = new ReplyPushLoop(new PrimaryRegistry(null), agents, inbox, scheduler, 5, 100_000);
|
||||||
loop.onTicketTerminal("task-1", WORKER, false); // no lead known -> never registered as pending
|
loop.onTicketTerminal("task-1", WORKER, false); // no lead known -> never registered as pending
|
||||||
assertEquals(ReplyPushLoop.Action.STOP, loop.decideTickets(PRIMARY, 0));
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 0));
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- CB-588: onTicketTerminal integration ---------------------------------------------------
|
// --- CB-588: onTicketTerminal integration ---------------------------------------------------
|
||||||
@@ -255,7 +290,7 @@ class ReplyPushLoopTest {
|
|||||||
String nudge = rec.sentParams().getFirst().getValue().toString();
|
String nudge = rec.sentParams().getFirst().getValue().toString();
|
||||||
assertTrue(nudge.contains("task-1"), "nudge should name the ticket");
|
assertTrue(nudge.contains("task-1"), "nudge should name the ticket");
|
||||||
assertTrue(nudge.contains("bridge_poll(ticket="), "nudge should name the exact ticket-poll call");
|
assertTrue(nudge.contains("bridge_poll(ticket="), "nudge should name the exact ticket-poll call");
|
||||||
assertFalse(nudge.contains("bridge_poll(target="), "a ticket nudge must not tell the lead to run the target-poll call");
|
assertFalse(nudge.contains("bridge_poll(target="), "a ticket-only nudge must not tell the lead to run the target-poll call");
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -346,11 +381,11 @@ class ReplyPushLoopTest {
|
|||||||
void aTicketStillPendingWhenTheLoopStopsIsNotStranded() {
|
void aTicketStillPendingWhenTheLoopStopsIsNotStranded() {
|
||||||
// Regression for the race a reviewer found in gitea PR #73: onTicketTerminal's
|
// Regression for the race a reviewer found in gitea PR #73: onTicketTerminal's
|
||||||
// activeLeads.putIfAbsent can see the lead's slot as still occupied a moment before
|
// activeLeads.putIfAbsent can see the lead's slot as still occupied a moment before
|
||||||
// decideTickets' STOP releases it, so the ticket coalesces onto a schedule that is about to
|
// decide's STOP releases it, so the ticket coalesces onto a schedule that is about to
|
||||||
// die and nothing ever nudges about it. Forcing that exact thread interleaving is not
|
// die and nothing ever nudges about it. Forcing that exact thread interleaving is not
|
||||||
// reliable, so this drives stopOrRestartTicketLoop — the STOP path's own release-and-recheck —
|
// reliable, so this drives stopOrRestart — the STOP path's own release-and-recheck —
|
||||||
// directly, arranging the state it must not lose a ticket in: a ticket pending for the lead
|
// directly, arranging the state it must not lose a ticket in: a ticket pending for the lead
|
||||||
// that was NOT part of the pre-decision snapshot (pendingBefore=empty), standing in for one
|
// that was NOT part of the pre-decision snapshot (ticketsBefore=empty), standing in for one
|
||||||
// that races in during the decision-to-release window.
|
// that races in during the decision-to-release window.
|
||||||
var rec = recordingClient();
|
var rec = recordingClient();
|
||||||
agents = new AgentControl(rec);
|
agents = new AgentControl(rec);
|
||||||
@@ -358,9 +393,9 @@ class ReplyPushLoopTest {
|
|||||||
|
|
||||||
loop.onTicketTerminal("task-1", WORKER, false); // pendingTickets={task-1}; activeLeads={PRIMARY}
|
loop.onTicketTerminal("task-1", WORKER, false); // pendingTickets={task-1}; activeLeads={PRIMARY}
|
||||||
|
|
||||||
// Stand in for the scheduler thread reaching decideTickets==STOP for this lead — with nothing
|
// Stand in for the scheduler thread reaching decide==STOP for this lead — with nothing
|
||||||
// pending at decide time — while task-1 races in before the release below runs.
|
// pending at decide time — while task-1 races in before the release below runs.
|
||||||
loop.stopOrRestartTicketLoop(PRIMARY, Set.of());
|
loop.stopOrRestart(PRIMARY, Set.of(), Set.of());
|
||||||
|
|
||||||
assertTrue(loop.isActive(), "a ticket that raced the loop's stop must reclaim the schedule "
|
assertTrue(loop.isActive(), "a ticket that raced the loop's stop must reclaim the schedule "
|
||||||
+ "slot, not be stranded with no schedule left to ever nudge about it");
|
+ "slot, not be stranded with no schedule left to ever nudge about it");
|
||||||
@@ -368,23 +403,67 @@ class ReplyPushLoopTest {
|
|||||||
|
|
||||||
@Test
|
@Test
|
||||||
void aStaleUncollectedTicketAtCapDoesNotRestartTheLoop() {
|
void aStaleUncollectedTicketAtCapDoesNotRestartTheLoop() {
|
||||||
// The other direction of the same fix: restarting on ANY non-empty pendingFor(lead) would be
|
// The other direction of the same fix: restarting on ANY non-empty pending set would be
|
||||||
// wrong. When STOP is reached because the reminder cap was hit, the same never-collected
|
// wrong. When STOP is reached because the reminder cap was hit, the same never-collected
|
||||||
// ticket is expected to still be there — that is the cap doing its job (acceptance criterion
|
// ticket is expected to still be there — that is the cap doing its job (acceptance criterion
|
||||||
// #7: nudges stay bounded). task-1 here was already accounted for at decide time (it is in
|
// #5: no spin / nudges stay bounded). task-1 here was already accounted for at decide time
|
||||||
// pendingBefore), so it must not restart the loop just because it is still sitting there.
|
// (it is in ticketsBefore), so it must not restart the loop just because it is still sitting
|
||||||
|
// there.
|
||||||
var rec = recordingClient();
|
var rec = recordingClient();
|
||||||
agents = new AgentControl(rec);
|
agents = new AgentControl(rec);
|
||||||
ReplyPushLoop loop = loop(1, 100_000);
|
ReplyPushLoop loop = loop(1, 100_000);
|
||||||
|
|
||||||
loop.onTicketTerminal("task-1", WORKER, false); // pendingTickets={task-1}; activeLeads={PRIMARY}
|
loop.onTicketTerminal("task-1", WORKER, false); // pendingTickets={task-1}; activeLeads={PRIMARY}
|
||||||
|
|
||||||
loop.stopOrRestartTicketLoop(PRIMARY, Set.of("task-1"));
|
loop.stopOrRestart(PRIMARY, Set.of(), Set.of("task-1"));
|
||||||
|
|
||||||
assertFalse(loop.isActive(), "a stale ticket already accounted for at decide time must not "
|
assertFalse(loop.isActive(), "a stale ticket already accounted for at decide time must not "
|
||||||
+ "restart the loop — that would defeat the reminder cap");
|
+ "restart the loop — that would defeat the reminder cap");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- mirror of the two stopOrRestart tests above, for the reply arm ------------------------
|
||||||
|
//
|
||||||
|
// Both tests above only ever passed Set.of() for repliesBefore, so racedIn's reply branch
|
||||||
|
// (`pendingReplyTargetsFor(lead).stream().anyMatch(t -> !repliesBefore.contains(t))`) was
|
||||||
|
// never exercised by anything other than an always-empty snapshot. The reviewer flagged this:
|
||||||
|
// racedIn is symmetric in the code, and only half of it was pinned by a test.
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aReplyStillPendingWhenTheLoopStopsIsNotStranded() {
|
||||||
|
// Mirrors aTicketStillPendingWhenTheLoopStopsIsNotStranded: a reply target that raced in
|
||||||
|
// during the decision-to-release window (absent from the "before" snapshot) must reclaim
|
||||||
|
// the schedule slot rather than being stranded with no schedule left to nudge about it.
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
ReplyPushLoop loop = loop(1, 100_000); // long backoff — no natural tick fires during this test
|
||||||
|
|
||||||
|
loop.onReplyQueued(WORKER); // pendingReplies={term_worker}; activeLeads={PRIMARY}
|
||||||
|
|
||||||
|
loop.stopOrRestart(PRIMARY, Set.of(), Set.of());
|
||||||
|
|
||||||
|
assertTrue(loop.isActive(), "a reply that raced the loop's stop must reclaim the schedule "
|
||||||
|
+ "slot, not be stranded with no schedule left to ever nudge about it");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aStaleUncollectedReplyAtCapDoesNotRestartTheLoop() {
|
||||||
|
// Mirrors aStaleUncollectedTicketAtCapDoesNotRestartTheLoop: a reply target already
|
||||||
|
// accounted for at decide time (present in repliesBefore) must not restart the loop —
|
||||||
|
// that is the reminder cap doing its job, not a race.
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
ReplyPushLoop loop = loop(1, 100_000);
|
||||||
|
|
||||||
|
loop.onReplyQueued(WORKER); // pendingReplies={term_worker}; activeLeads={PRIMARY}
|
||||||
|
|
||||||
|
loop.stopOrRestart(PRIMARY, Set.of(WORKER), Set.of());
|
||||||
|
|
||||||
|
assertFalse(loop.isActive(), "a stale reply target already accounted for at decide time "
|
||||||
|
+ "must not restart the loop — that would defeat the reminder cap");
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void ticketNudgeFormatIsCorrect() {
|
void ticketNudgeFormatIsCorrect() {
|
||||||
String single = ReplyPushLoop.TICKET_NUDGE_FORMAT.formatted("task-1", "", "task-1");
|
String single = ReplyPushLoop.TICKET_NUDGE_FORMAT.formatted("task-1", "", "task-1");
|
||||||
@@ -402,7 +481,306 @@ class ReplyPushLoopTest {
|
|||||||
void inboxNudgeStillUsesTheOriginalTargetPollCall() {
|
void inboxNudgeStillUsesTheOriginalTargetPollCall() {
|
||||||
String nudge = ReplyPushLoop.NUDGE_FORMAT.formatted(WORKER, WORKER);
|
String nudge = ReplyPushLoop.NUDGE_FORMAT.formatted(WORKER, WORKER);
|
||||||
assertTrue(nudge.contains("bridge_poll(target=" + WORKER + ")"),
|
assertTrue(nudge.contains("bridge_poll(target=" + WORKER + ")"),
|
||||||
"CB-588 must not change the CB-307 inbox nudge's call shape");
|
"CB-588/CB-590 must not change the CB-307 inbox nudge's call shape");
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- CB-590: one schedule per lead — no overlap, no lost nudges -----------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void replyAndTicketForTheSameLeadCoalesceIntoOneSendNeverOverlapping() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
var loop = loop(1, 300); // backoff wide enough that both entry points land before the first tick
|
||||||
|
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one combined nudge should have been sent");
|
||||||
|
Thread.sleep(300);
|
||||||
|
assertEquals(1, rec.sendCount(),
|
||||||
|
"a reply and a ticket for the same lead must coalesce onto ONE schedule — "
|
||||||
|
+ "two nudge injections into the same lead pane must never overlap");
|
||||||
|
String nudge = rec.sentParams().getFirst().getValue().toString();
|
||||||
|
assertTrue(nudge.contains("bridge_poll(target=" + WORKER + ")"),
|
||||||
|
"the combined nudge must still mention the reply: " + nudge);
|
||||||
|
assertTrue(nudge.contains("task-1"), "the combined nudge must still mention the ticket: " + nudge);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aReplyQueuedWhileTheLeadIsBusyIsNotLostWhenATicketArrivesToo() throws Exception {
|
||||||
|
// The lead is busy for its first two status checks, then becomes injectable. A reply is
|
||||||
|
// queued while busy; a ticket for the same lead arrives before the lead frees up. Neither
|
||||||
|
// may be dropped — deferred is fine, lost is not (acceptance criterion #2).
|
||||||
|
var rec = new BusyThenIdleHerdrClient(2);
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
var loop = loop(1, 50); // cap=1: WAIT_BUSY doesn't count against it, so exactly one send once injectable
|
||||||
|
|
||||||
|
loop.onReplyQueued(WORKER); // schedule starts, first tick(s) WAIT_BUSY
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false); // coalesces onto the same waiting schedule
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
|
||||||
|
"once the lead becomes injectable, the deferred work must still be nudged");
|
||||||
|
Thread.sleep(200);
|
||||||
|
assertEquals(1, rec.sendCount(), "exactly one nudge once injectable — reply and ticket coalesced");
|
||||||
|
String nudge = rec.sentParams().getFirst().getValue().toString();
|
||||||
|
assertTrue(nudge.contains("bridge_poll(target=" + WORKER + ")"), "the reply must not be dropped: " + nudge);
|
||||||
|
assertTrue(nudge.contains("task-1"), "the ticket must not be dropped: " + nudge);
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- CB-590 follow-up: per-source reminder budgets — the regression this round exists for ---
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void oneExhaustedSourceDoesNotBlockANudgeForTheOtherSource() {
|
||||||
|
// The live trace this ticket was filed from: an undrained reply target got nudged up to
|
||||||
|
// its cap (5 reminders), then a ticket for the SAME lead went terminal shortly before the
|
||||||
|
// next scheduled tick — so it coalesced onto the still-active schedule (arriving BEFORE
|
||||||
|
// that tick's "before" snapshot, not during the decision-to-release race stopOrRestart
|
||||||
|
// guards). With CB-590's single shared reminder counter, that tick's decide() saw
|
||||||
|
// reminderCount already at the cap and returned STOP regardless of the ticket, and because
|
||||||
|
// the ticket was already present in that tick's "before" snapshot, stopOrRestart's
|
||||||
|
// racedIn check (proven correct on its own above) did not save it either — it is not a
|
||||||
|
// race, it looks like ordinary stale backlog. The ticket was then stranded: pending
|
||||||
|
// forever with no live schedule, never named in any nudge.
|
||||||
|
//
|
||||||
|
// Fixed by giving each source its own counter. Here the reply source is AT its cap (2/2)
|
||||||
|
// and the ticket source has NEVER been nudged (0/2) — decide() must still return INJECT,
|
||||||
|
// because the ticket is still eligible on its own budget.
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
var loop = loop(2, 100_000); // huge backoff — this test drives decide()/isActive() directly
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 2, 0),
|
||||||
|
"the reply source is exhausted (2/2), but the ticket source has never been "
|
||||||
|
+ "nudged (0/2) — the lead must still be injected so the ticket is not "
|
||||||
|
+ "lost, exactly the CB-590 follow-up regression");
|
||||||
|
|
||||||
|
// isActive() (criterion #5): a real tick that takes the INJECT branch above never calls
|
||||||
|
// stopOrRestart, so the schedule started by onTicketTerminal above stays live — the
|
||||||
|
// ticket is not left stranded with isActive()==false while it is still pending.
|
||||||
|
assertTrue(loop.isActive(), "the schedule must stay active while the ticket source still "
|
||||||
|
+ "has budget left, even though the reply source sharing it is exhausted");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void bothSourcesExhaustedIsStillStop() {
|
||||||
|
// The flip side: per-source budgets must not turn into unbounded nudging. When BOTH
|
||||||
|
// sources are at their cap, decide() must still STOP — a per-source budget is still a
|
||||||
|
// budget.
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
var loop = loop(2, 100_000);
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
|
||||||
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 2, 2),
|
||||||
|
"both the reply and the ticket source are at their own cap — must still stop");
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- CB-598: work arriving during a backoff must not read as stale backlog -------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aTargetArrivingDuringTheBackoffGetsNudgedDespiteAnAlreadyCappedSibling() {
|
||||||
|
// The bug: reminder counts used to be a single counter per lead per source, carried
|
||||||
|
// forward across scheduled ticks (scheduleNext(lead, count + 1, ...)) rather than tracked
|
||||||
|
// per pending item. WORKER gets nudged once here, which — with cap=1 — exhausts the
|
||||||
|
// shared reply-source counter for this lead. WORKER2 then queues a reply for the SAME
|
||||||
|
// lead "during the backoff": while the schedule from WORKER's tick is still active, before
|
||||||
|
// the next tick's own start-of-tick snapshot runs. At that next tick, the OLD code passed
|
||||||
|
// the already-exhausted shared counter into decide() regardless of WORKER2 never having
|
||||||
|
// been named in any nudge, and — because WORKER2 was already present in that tick's
|
||||||
|
// "before" snapshot — stopOrRestart's race check (proven correct on its own elsewhere in
|
||||||
|
// this file) does not save it either: it looks like ordinary stale backlog, not a race.
|
||||||
|
// WORKER2 was then stranded forever with no live schedule and no nudge ever naming it.
|
||||||
|
//
|
||||||
|
// tick() is driven directly (package-private, same reasoning as stopOrRestart being
|
||||||
|
// directly testable) so the exact interleaving is deterministic instead of racing the
|
||||||
|
// scheduler thread over a real ~15s backoff.
|
||||||
|
//
|
||||||
|
// Before the fix, this test fails on the second assertEquals: rec.sendCount() stays at 1
|
||||||
|
// (decide() returns STOP on the second tick(), so injectNudge is never called a second
|
||||||
|
// time) and the "must still get one" assertion never even runs.
|
||||||
|
int cap = 1;
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
inbox.own(WORKER2);
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
var loop = loop(cap, 100_000); // huge backoff — nothing fires on its own; we drive tick()
|
||||||
|
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
loop.tick(PRIMARY); // first tick: nudges WORKER alone; WORKER's own count reaches the cap
|
||||||
|
assertEquals(1, rec.sendCount(), "the first tick should nudge about WORKER");
|
||||||
|
|
||||||
|
// WORKER2 "arrives during the backoff": queued for the same lead while the schedule from
|
||||||
|
// the tick above is still active (activeLeads still holds PRIMARY), before the next tick
|
||||||
|
// (simulated below) takes its own start-of-tick snapshot.
|
||||||
|
inbox.publish(WORKER2, "m2", "hello2");
|
||||||
|
loop.onReplyQueued(WORKER2);
|
||||||
|
|
||||||
|
loop.tick(PRIMARY); // the tick that would fire once that backoff elapsed
|
||||||
|
|
||||||
|
assertEquals(2, rec.sendCount(),
|
||||||
|
"WORKER2 was never named in any nudge yet and must still get one, even though "
|
||||||
|
+ "WORKER's own reminder count is already at the cap");
|
||||||
|
String secondNudge = rec.sentParams().get(1).getValue().toString();
|
||||||
|
assertTrue(secondNudge.contains(WORKER2), "the never-named target must be named: " + secondNudge);
|
||||||
|
|
||||||
|
// Criterion #3: isActive() must reflect that this lead still had a live nudge to give —
|
||||||
|
// the second tick took the INJECT branch, so the schedule stayed live rather than being
|
||||||
|
// torn down under WORKER2.
|
||||||
|
assertTrue(loop.isActive(), "the schedule must stay active after nudging the fresh target");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aTicketArrivingDuringTheBackoffGetsNudgedDespiteAnAlreadyCappedSibling() {
|
||||||
|
// Mirrors the reply-side test above for the ticket source.
|
||||||
|
int cap = 1;
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
var loop = loop(cap, 100_000);
|
||||||
|
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
loop.tick(PRIMARY); // first tick: nudges task-1 alone; its count reaches the cap
|
||||||
|
assertEquals(1, rec.sendCount(), "the first tick should nudge about task-1");
|
||||||
|
|
||||||
|
loop.onTicketTerminal("task-2", WORKER, false); // arrives during the backoff, same lead
|
||||||
|
loop.tick(PRIMARY);
|
||||||
|
|
||||||
|
assertEquals(2, rec.sendCount(),
|
||||||
|
"task-2 was never named in any nudge yet and must still get one, even though "
|
||||||
|
+ "task-1's reminder count is already at the cap");
|
||||||
|
String secondNudge = rec.sentParams().get(1).getValue().toString();
|
||||||
|
assertTrue(secondNudge.contains("task-2"), "the never-named ticket must be named: " + secondNudge);
|
||||||
|
assertTrue(loop.isActive(), "the schedule must stay active after nudging the fresh ticket");
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- CB-582: bridge_ask question-open nudges -------------------------------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void onQuestionOpenedWithNoKnownLeadNeverStartsASchedule() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
var loop = new ReplyPushLoop(new PrimaryRegistry(null), agents, inbox, scheduler, 5, 50);
|
||||||
|
|
||||||
|
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
|
||||||
|
|
||||||
|
assertFalse(loop.isActive(), "no lead known means nothing to nudge yet");
|
||||||
|
Thread.sleep(150);
|
||||||
|
assertEquals(0, rec.sendCount(), "must not nudge when no lead is known to be waiting");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void decideQuestionsWithNothingPendingIsStop() {
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(PRIMARY, 0, 0, 0));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void decideQuestionsAtCapIsStop() {
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
var loop = loop(2, 100_000);
|
||||||
|
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
|
||||||
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 0, 2));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void decideQuestionsUnderCapWithInjectableLeadIsInject() {
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
var loop = loop(5, 100_000);
|
||||||
|
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0, 0));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void onQuestionOpenedCausesExactlyOneNudgeNamingTheTicketAndTurnId() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
|
||||||
|
loop(1, 50).onQuestionOpened("task-1", WORKER, "term_worker#1", "which config file?");
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one question nudge should have been sent");
|
||||||
|
assertEquals(1, rec.sendCount());
|
||||||
|
String nudge = rec.sentParams().getFirst().getValue().toString();
|
||||||
|
assertTrue(nudge.contains("task-1"), "nudge should name the ticket: " + nudge);
|
||||||
|
assertTrue(nudge.contains("term_worker#1"), "nudge should name the turnId: " + nudge);
|
||||||
|
assertTrue(nudge.contains("bridge_send(turnId="), "nudge should name the exact answer call: " + nudge);
|
||||||
|
assertTrue(nudge.contains("which config file?"), "nudge should include the question text: " + nudge);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void questionClosedPreventsFurtherNudging() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
var loop = loop(1, 100);
|
||||||
|
|
||||||
|
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
|
||||||
|
loop.questionClosed("term_worker#1"); // answered/lapsed before the first tick fired
|
||||||
|
|
||||||
|
Thread.sleep(300); // let the scheduled tick run
|
||||||
|
assertEquals(0, rec.sendCount(), "an already-closed question must never be nudged");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void questionNudgesSendUpToCapThenStop() throws Exception {
|
||||||
|
int cap = 2;
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
rec.sendLatch = new CountDownLatch(cap);
|
||||||
|
|
||||||
|
loop(cap, 50).onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(5, TimeUnit.SECONDS), cap + " question nudges should have fired");
|
||||||
|
Thread.sleep(300);
|
||||||
|
assertEquals(cap, rec.sendCount(), "exactly " + cap + " question nudges (cap=" + cap + ")");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void questionAndTicketForTheSameLeadCoalesceIntoOneSend() throws Exception {
|
||||||
|
var rec = recordingClient();
|
||||||
|
agents = new AgentControl(rec);
|
||||||
|
var loop = loop(1, 300); // backoff wide enough that both entry points land before the first tick
|
||||||
|
|
||||||
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
loop.onQuestionOpened("task-2", WORKER, "term_worker#1", "which config?");
|
||||||
|
|
||||||
|
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one combined nudge should have been sent");
|
||||||
|
Thread.sleep(300);
|
||||||
|
assertEquals(1, rec.sendCount(),
|
||||||
|
"a ticket and a question for the same lead must coalesce onto ONE schedule");
|
||||||
|
String nudge = rec.sentParams().getFirst().getValue().toString();
|
||||||
|
assertTrue(nudge.contains("task-1"), "the combined nudge must still mention the ticket: " + nudge);
|
||||||
|
assertTrue(nudge.contains("term_worker#1"), "the combined nudge must still mention the question: " + nudge);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void oneExhaustedQuestionSourceDoesNotBlockANudgeForTheOtherSources() {
|
||||||
|
// Mirrors oneExhaustedSourceDoesNotBlockANudgeForTheOtherSource for the question source:
|
||||||
|
// the question source is at its cap (2/2), but the ticket source has never been nudged
|
||||||
|
// (0/2) — decide() must still INJECT so the ticket is not stranded.
|
||||||
|
agents = agentWithStatus("idle");
|
||||||
|
var loop = loop(2, 100_000);
|
||||||
|
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
|
||||||
|
loop.onTicketTerminal("task-2", WORKER, false);
|
||||||
|
|
||||||
|
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0, 2),
|
||||||
|
"the question source is exhausted (2/2), but the ticket source has never been "
|
||||||
|
+ "nudged (0/2) — the lead must still be injected so the ticket is not lost");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void questionNudgeFormatIsCorrect() {
|
||||||
|
String single = ReplyPushLoop.QUESTION_NUDGE_FORMAT.formatted(
|
||||||
|
WORKER, "task-1", "term_worker#1", "which config?");
|
||||||
|
assertTrue(single.contains("Worker term_worker"));
|
||||||
|
assertTrue(single.contains("bridge_send(turnId=\"term_worker#1\""));
|
||||||
|
assertTrue(single.contains("which config?"));
|
||||||
|
|
||||||
|
String multi = ReplyPushLoop.QUESTIONS_NUDGE_FORMAT.formatted(2, "task-1 (turnId=t1), task-2 (turnId=t2)");
|
||||||
|
assertTrue(multi.contains("2 workers"));
|
||||||
|
assertTrue(multi.contains("bridge_poll(ticket=...)"));
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- metrics (CB-512) ----------------------------------------------------------------------
|
// --- metrics (CB-512) ----------------------------------------------------------------------
|
||||||
@@ -430,8 +808,10 @@ class ReplyPushLoopTest {
|
|||||||
agents = agentWithStatus("idle");
|
agents = agentWithStatus("idle");
|
||||||
inbox.publish(WORKER, "m1", "hello");
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
Metrics metrics = new Metrics();
|
Metrics metrics = new Metrics();
|
||||||
|
var loop = loop(2, 100, metrics);
|
||||||
|
loop.onReplyQueued(WORKER);
|
||||||
|
|
||||||
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100, metrics).decide(WORKER, 2));
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 2, 0));
|
||||||
|
|
||||||
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
|
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
|
||||||
"hitting the reminder cap must count as exhausted");
|
"hitting the reminder cap must count as exhausted");
|
||||||
@@ -459,7 +839,7 @@ class ReplyPushLoopTest {
|
|||||||
var loop = loop(2, 100_000, metrics);
|
var loop = loop(2, 100_000, metrics);
|
||||||
loop.onTicketTerminal("task-1", WORKER, false);
|
loop.onTicketTerminal("task-1", WORKER, false);
|
||||||
|
|
||||||
assertEquals(ReplyPushLoop.Action.STOP, loop.decideTickets(PRIMARY, 2));
|
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 2));
|
||||||
|
|
||||||
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
|
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
|
||||||
"hitting the ticket reminder cap must count as exhausted");
|
"hitting the ticket reminder cap must count as exhausted");
|
||||||
@@ -547,4 +927,49 @@ class ReplyPushLoopTest {
|
|||||||
private static RecordingHerdrClient recordingClient() {
|
private static RecordingHerdrClient recordingClient() {
|
||||||
return new RecordingHerdrClient();
|
return new RecordingHerdrClient();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Thread-safe fake that reports {@code working} (not injectable) for its first
|
||||||
|
* {@code busyChecks} status calls, then {@code idle} forever after — used to prove work queued
|
||||||
|
* while the lead is busy is deferred, not dropped, once it becomes injectable.
|
||||||
|
*/
|
||||||
|
private static final class BusyThenIdleHerdrClient implements HerdrClient {
|
||||||
|
private final List<Map.Entry<String, Object>> calls =
|
||||||
|
Collections.synchronizedList(new ArrayList<>());
|
||||||
|
private final AtomicInteger statusChecks = new AtomicInteger();
|
||||||
|
private final int busyChecks;
|
||||||
|
volatile CountDownLatch sendLatch = new CountDownLatch(1);
|
||||||
|
|
||||||
|
BusyThenIdleHerdrClient(int busyChecks) {
|
||||||
|
this.busyChecks = busyChecks;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public JsonNode call(String method, Object params) {
|
||||||
|
if ("agent.get".equals(method)) {
|
||||||
|
String status = statusChecks.getAndIncrement() < busyChecks ? "working" : "idle";
|
||||||
|
return MAPPER.createObjectNode()
|
||||||
|
.set("agent", MAPPER.createObjectNode()
|
||||||
|
.put("terminal_id", PRIMARY)
|
||||||
|
.put("agent_status", status));
|
||||||
|
}
|
||||||
|
if ("agent.prompt".equals(method)) {
|
||||||
|
calls.add(Map.entry(method, params));
|
||||||
|
sendLatch.countDown();
|
||||||
|
}
|
||||||
|
return MAPPER.createObjectNode();
|
||||||
|
}
|
||||||
|
|
||||||
|
long sendCount() {
|
||||||
|
return calls.size();
|
||||||
|
}
|
||||||
|
|
||||||
|
List<Map.Entry<String, Object>> sentParams() {
|
||||||
|
return List.copyOf(calls);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
}
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -18,6 +18,8 @@ import dev.ltms.bridged.session.GitWorktrees;
|
|||||||
import dev.ltms.bridged.session.SessionManager;
|
import dev.ltms.bridged.session.SessionManager;
|
||||||
import dev.ltms.bridged.session.Worktrees;
|
import dev.ltms.bridged.session.Worktrees;
|
||||||
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
||||||
|
import dev.ltms.bridged.member.CompositePeerLauncher;
|
||||||
|
import dev.ltms.bridged.placement.PlacementPolicies;
|
||||||
import io.javalin.Javalin;
|
import io.javalin.Javalin;
|
||||||
import org.junit.jupiter.api.AfterEach;
|
import org.junit.jupiter.api.AfterEach;
|
||||||
import org.junit.jupiter.api.Test;
|
import org.junit.jupiter.api.Test;
|
||||||
@@ -240,6 +242,43 @@ class BridgedAppTest {
|
|||||||
assertFalse(herdr.called("agent.start"), "an unknown profile must not spawn anything");
|
assertFalse(herdr.called("agent.start"), "an unknown profile must not spawn anything");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-599: a profile at its {@code maxLoad} cap must not surface as a bare 500 — the caller
|
||||||
|
* needs a structured, readable reason, distinct from "unknown_profile".
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void spawnAtMaxLoadIs503WithTheCapacityReasonNotABare500() throws Exception {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
BridgedConfig.Profile wcfg = new BridgedConfig.Profile(
|
||||||
|
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
|
||||||
|
"tab", "bridged-workers", "worker: {profile} #{n}", null,
|
||||||
|
null, null, null, null, null, null, null, 0, null, null, null);
|
||||||
|
Map<String, BridgedConfig.Profile> profiles = Map.of(wcfg.profile(), wcfg);
|
||||||
|
ClaudeCodeLauncher delegate = new ClaudeCodeLauncher(
|
||||||
|
new AgentControl(herdr), new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
|
||||||
|
profiles, wcfg.profile(), k -> "BRIDGED_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
|
||||||
|
CompositePeerLauncher workers = new CompositePeerLauncher(
|
||||||
|
List.of(delegate), wcfg.profile(), profiles, PlacementPolicies.fixed(), _ -> 0);
|
||||||
|
SessionManager sessions = new SessionManager(workers, new GitWorktrees());
|
||||||
|
this.presence = sessions.asPresence();
|
||||||
|
Injector injector = new Injector(new AgentControl(herdr));
|
||||||
|
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||||
|
sessions.onAcquire(inbox::own);
|
||||||
|
MessageService messages = new MessageService(new AgentControl(herdr), injector, new Rendezvous(), inbox);
|
||||||
|
app = new BridgedApp(herdr, workers, sessions, messages, this.presence, null).build().start("127.0.0.1", 0);
|
||||||
|
int port = app.port();
|
||||||
|
|
||||||
|
HttpResponse<String> res = req(port, "POST", "/members?profile=ltms-local");
|
||||||
|
|
||||||
|
assertEquals(503, res.statusCode(), res.body());
|
||||||
|
JsonNode body = mapper.readTree(res.body());
|
||||||
|
assertEquals("no_capacity", body.get("error").asText());
|
||||||
|
String detail = body.get("detail").asText();
|
||||||
|
assertTrue(detail.contains("ltms-local"), "detail names the profile: " + detail);
|
||||||
|
assertTrue(detail.contains("maxLoad"), "detail explains the refusal: " + detail);
|
||||||
|
assertFalse(herdr.called("agent.start"), "at cap, the spawn is refused before any herdr call");
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void spawnWorkerReusesExistingWorkerSpace() throws Exception {
|
void spawnWorkerReusesExistingWorkerSpace() throws Exception {
|
||||||
// A space labelled "bridged-workers" already exists → no second workspace.create.
|
// A space labelled "bridged-workers" already exists → no second workspace.create.
|
||||||
@@ -445,6 +484,67 @@ class BridgedAppTest {
|
|||||||
assertTrue(mapper.readTree(req(port, "GET", "/sessions/term_a/status").body()).get("ready").asBoolean());
|
assertTrue(mapper.readTree(req(port, "GET", "/sessions/term_a/status").body()).get("ready").asBoolean());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-582: a lead polling {@code GET /sessions/{id}/status} on its normal cadence — not the
|
||||||
|
* ticket-scoped {@code /tasks/{ticket}} — must also see a worker's open async {@code bridge_ask}
|
||||||
|
* question, since the reverse-rendezvous window it opened with is far shorter than that cadence.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void sessionStatusReportsAnOpenQuestionWhenTheWorkerIsMidAsk() throws Exception {
|
||||||
|
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // poller delivers the injection
|
||||||
|
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||||
|
|
||||||
|
HttpResponse<String> accepted = postMessage(port, "{\"content\":\"do it\",\"wait\":false}");
|
||||||
|
assertEquals(202, accepted.statusCode());
|
||||||
|
String ticket = mapper.readTree(accepted.body()).get("ticket").asText();
|
||||||
|
|
||||||
|
Thread.sleep(200); // let the background async send open its rendezvous waiter
|
||||||
|
|
||||||
|
var ask = java.util.concurrent.CompletableFuture.supplyAsync(() -> {
|
||||||
|
try {
|
||||||
|
return postJson(port, "/sessions/term_a/ask",
|
||||||
|
"{\"question\":\"which config file?\",\"timeoutMs\":5000}");
|
||||||
|
} catch (Exception e) {
|
||||||
|
throw new RuntimeException(e);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
JsonNode task;
|
||||||
|
long deadline = System.currentTimeMillis() + 3000;
|
||||||
|
do {
|
||||||
|
task = mapper.readTree(req(port, "GET", "/tasks/" + ticket).body());
|
||||||
|
if ("asking".equals(task.path("phase").asText())) break;
|
||||||
|
//noinspection BusyWait
|
||||||
|
Thread.sleep(10);
|
||||||
|
} while (System.currentTimeMillis() < deadline);
|
||||||
|
assertEquals("asking", task.get("phase").asText());
|
||||||
|
String turnId = task.get("turnId").asText();
|
||||||
|
|
||||||
|
HttpResponse<String> status = req(port, "GET", "/sessions/term_a/status");
|
||||||
|
assertEquals(200, status.statusCode());
|
||||||
|
JsonNode body = mapper.readTree(status.body());
|
||||||
|
assertEquals("idle", body.get("status").asText(), "the live status must still be reported");
|
||||||
|
assertEquals("which config file?", body.get("question").asText());
|
||||||
|
assertEquals(turnId, body.get("turnId").asText());
|
||||||
|
assertEquals(ticket, body.get("ticket").asText());
|
||||||
|
|
||||||
|
// Answer it — via the same /message route bridge_send uses, keyed by turnId — so the
|
||||||
|
// background ask thread does not linger past the test.
|
||||||
|
var answer = java.util.concurrent.CompletableFuture.supplyAsync(() -> {
|
||||||
|
try {
|
||||||
|
return postJson(port, "/sessions/term_a/message",
|
||||||
|
"{\"content\":\"config.yaml\",\"turnId\":\"" + turnId + "\",\"timeoutMs\":4000}");
|
||||||
|
} catch (Exception e) {
|
||||||
|
throw new RuntimeException(e);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
HttpResponse<String> askResponse = ask.get(6, java.util.concurrent.TimeUnit.SECONDS);
|
||||||
|
assertEquals(200, askResponse.statusCode());
|
||||||
|
assertEquals("config.yaml", mapper.readTree(askResponse.body()).get("answer").asText());
|
||||||
|
postJson(port, "/sessions/term_a/reply", "{\"content\":\"done\"}");
|
||||||
|
answer.get(6, java.util.concurrent.TimeUnit.SECONDS);
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void stopWorkerInPanePlacementClosesOnlyThePane() throws Exception {
|
void stopWorkerInPanePlacementClosesOnlyThePane() throws Exception {
|
||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
|||||||
@@ -27,11 +27,15 @@ public final class FakeWorktrees implements Worktrees {
|
|||||||
public record SnapshotCall(String worktreePath, String branch, String message) {
|
public record SnapshotCall(String worktreePath, String branch, String message) {
|
||||||
}
|
}
|
||||||
|
|
||||||
|
public record PruneCall(String repoRoot, long minAgeMillis) {
|
||||||
|
}
|
||||||
|
|
||||||
private final List<AddCall> addCalls = new CopyOnWriteArrayList<>();
|
private final List<AddCall> addCalls = new CopyOnWriteArrayList<>();
|
||||||
private final List<RemoveCall> removeCalls = new CopyOnWriteArrayList<>();
|
private final List<RemoveCall> removeCalls = new CopyOnWriteArrayList<>();
|
||||||
private final List<OverlayCall> overlayCalls = new CopyOnWriteArrayList<>();
|
private final List<OverlayCall> overlayCalls = new CopyOnWriteArrayList<>();
|
||||||
private final List<RepoRootCall> repoRootCalls = new CopyOnWriteArrayList<>();
|
private final List<RepoRootCall> repoRootCalls = new CopyOnWriteArrayList<>();
|
||||||
private final List<SnapshotCall> snapshotCalls = new CopyOnWriteArrayList<>();
|
private final List<SnapshotCall> snapshotCalls = new CopyOnWriteArrayList<>();
|
||||||
|
private final List<PruneCall> pruneCalls = new CopyOnWriteArrayList<>();
|
||||||
private final Set<String> existingPaths = ConcurrentHashMap.newKeySet();
|
private final Set<String> existingPaths = ConcurrentHashMap.newKeySet();
|
||||||
private final Set<String> trackedPaths = ConcurrentHashMap.newKeySet();
|
private final Set<String> trackedPaths = ConcurrentHashMap.newKeySet();
|
||||||
private final AtomicLong snapshotSeq = new AtomicLong();
|
private final AtomicLong snapshotSeq = new AtomicLong();
|
||||||
@@ -40,6 +44,8 @@ public final class FakeWorktrees implements Worktrees {
|
|||||||
private volatile boolean dirty = false;
|
private volatile boolean dirty = false;
|
||||||
private volatile String repoRoot = "/repo";
|
private volatile String repoRoot = "/repo";
|
||||||
private volatile String prefix = "/worktrees";
|
private volatile String prefix = "/worktrees";
|
||||||
|
private volatile WipRefStats wipRefs = new WipRefStats(0, 0L);
|
||||||
|
private volatile int pruneResult = 0;
|
||||||
|
|
||||||
public FakeWorktrees withRepoRoot(String root) {
|
public FakeWorktrees withRepoRoot(String root) {
|
||||||
this.repoRoot = root;
|
this.repoRoot = root;
|
||||||
@@ -82,6 +88,18 @@ public final class FakeWorktrees implements Worktrees {
|
|||||||
return this;
|
return this;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** Configure the value returned by {@link #wipRefs}. */
|
||||||
|
public FakeWorktrees withWipRefs(WipRefStats stats) {
|
||||||
|
this.wipRefs = stats;
|
||||||
|
return this;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Configure the value returned by {@link #pruneWipRefs}. */
|
||||||
|
public FakeWorktrees withPruneResult(int deleted) {
|
||||||
|
this.pruneResult = deleted;
|
||||||
|
return this;
|
||||||
|
}
|
||||||
|
|
||||||
@Override
|
@Override
|
||||||
public String add(String repoRoot, String branch, String baseRef) {
|
public String add(String repoRoot, String branch, String baseRef) {
|
||||||
addCalls.add(new AddCall(repoRoot, branch, baseRef));
|
addCalls.add(new AddCall(repoRoot, branch, baseRef));
|
||||||
@@ -135,6 +153,17 @@ public final class FakeWorktrees implements Worktrees {
|
|||||||
return Optional.of("wip" + snapshotSeq.incrementAndGet());
|
return Optional.of("wip" + snapshotSeq.incrementAndGet());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public WipRefStats wipRefs(String repoRoot) {
|
||||||
|
return wipRefs;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
|
||||||
|
pruneCalls.add(new PruneCall(repoRoot, minAgeMillis));
|
||||||
|
return pruneResult;
|
||||||
|
}
|
||||||
|
|
||||||
public List<AddCall> addCalls() {
|
public List<AddCall> addCalls() {
|
||||||
return List.copyOf(addCalls);
|
return List.copyOf(addCalls);
|
||||||
}
|
}
|
||||||
@@ -167,6 +196,10 @@ public final class FakeWorktrees implements Worktrees {
|
|||||||
return List.copyOf(snapshotCalls);
|
return List.copyOf(snapshotCalls);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
public List<PruneCall> pruneCalls() {
|
||||||
|
return List.copyOf(pruneCalls);
|
||||||
|
}
|
||||||
|
|
||||||
public SnapshotCall lastSnapshot() {
|
public SnapshotCall lastSnapshot() {
|
||||||
return snapshotCalls.isEmpty() ? null : snapshotCalls.getLast();
|
return snapshotCalls.isEmpty() ? null : snapshotCalls.getLast();
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -3,6 +3,7 @@ package dev.ltms.bridged.session;
|
|||||||
import org.junit.jupiter.api.Test;
|
import org.junit.jupiter.api.Test;
|
||||||
import org.junit.jupiter.api.io.TempDir;
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.charset.StandardCharsets;
|
||||||
import java.nio.file.Files;
|
import java.nio.file.Files;
|
||||||
import java.nio.file.Path;
|
import java.nio.file.Path;
|
||||||
import java.util.HashSet;
|
import java.util.HashSet;
|
||||||
@@ -138,6 +139,53 @@ class GitWorktreesTest {
|
|||||||
return out;
|
return out;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** Write {@code content} as a blob into the object database; returns its sha. */
|
||||||
|
private static String blobOf(Path cwd, String content) throws Exception {
|
||||||
|
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "hash-object", "-w", "--stdin")
|
||||||
|
.redirectErrorStream(true).start();
|
||||||
|
p.getOutputStream().write(content.getBytes(StandardCharsets.UTF_8));
|
||||||
|
p.getOutputStream().close();
|
||||||
|
String out = new String(p.getInputStream().readAllBytes()).trim();
|
||||||
|
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git hash-object timed out");
|
||||||
|
assertEquals(0, p.exitValue(), "git hash-object failed:\n" + out);
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Build a single-file tree object from {@code blob}; returns the tree's sha. */
|
||||||
|
private static String treeOf(Path cwd, String path, String blob) throws Exception {
|
||||||
|
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "mktree")
|
||||||
|
.redirectErrorStream(true).start();
|
||||||
|
p.getOutputStream().write(("100644 blob " + blob + "\t" + path + "\n").getBytes(StandardCharsets.UTF_8));
|
||||||
|
p.getOutputStream().close();
|
||||||
|
String out = new String(p.getInputStream().readAllBytes()).trim();
|
||||||
|
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git mktree timed out");
|
||||||
|
assertEquals(0, p.exitValue(), "git mktree failed:\n" + out);
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** {@code git commit-tree} rooted at {@code tree} with a chosen committer date; returns the sha. */
|
||||||
|
private static String commitTree(Path cwd, String tree, String parent, String committerDate,
|
||||||
|
String message) throws Exception {
|
||||||
|
ProcessBuilder pb = new ProcessBuilder("git", "-C", cwd.toString(), "commit-tree",
|
||||||
|
tree, "-p", parent, "-m", message);
|
||||||
|
pb.environment().put("GIT_COMMITTER_DATE", committerDate);
|
||||||
|
Process p = pb.redirectErrorStream(true).start();
|
||||||
|
String out = new String(p.getInputStream().readAllBytes()).trim();
|
||||||
|
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git commit-tree timed out");
|
||||||
|
assertEquals(0, p.exitValue(), "git commit-tree failed:\n" + out);
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** {@code git update-ref <ref> <sha>} — create the snapshot ref directly. */
|
||||||
|
private static void updateRef(Path cwd, String ref, String sha) throws Exception {
|
||||||
|
git(cwd, "update-ref", ref, sha);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** True when {@code ref} exists in the repo (for-each-ref on a missing ref is empty, not an error). */
|
||||||
|
private static boolean refExists(Path cwd, String ref) throws Exception {
|
||||||
|
return !forEachRef(cwd, ref).trim().isEmpty();
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* The heart of CB-525: a provisioned worktree must not inherit the primary's MCP servers. Without
|
* The heart of CB-525: a provisioned worktree must not inherit the primary's MCP servers. Without
|
||||||
* the isolation step the checked-out {@code .mcp.json} carries them in, and a worker navigating
|
* the isolation step the checked-out {@code .mcp.json} carries them in, and a worker navigating
|
||||||
@@ -474,4 +522,104 @@ class GitWorktreesTest {
|
|||||||
assertThrows(WorktreeException.class, () -> gitWorktrees.snapshot(wt, branch, "test snapshot"),
|
assertThrows(WorktreeException.class, () -> gitWorktrees.snapshot(wt, branch, "test snapshot"),
|
||||||
"an unresolvable real index must fail loudly, not silently snapshot from an empty index");
|
"an unresolvable real index must fail loudly, not silently snapshot from an empty index");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586, criterion 2. A snapshot whose content is NOT reachable from {@code main} is the last
|
||||||
|
* copy of a worker's work, and must never be deleted automatically — even when it is old and
|
||||||
|
* even when the caller passes a zero age floor. Uses the real snapshot path on a dirty worktree,
|
||||||
|
* so the unreachable tree is exactly the shape CB-576/CB-578 stage C exist to protect.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void anUnreachableSnapshotIsNeverPruned(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = initRepo(tmp.resolve("repo"));
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
String branch = "cb-586-unreachable";
|
||||||
|
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
|
||||||
|
Files.writeString(Path.of(wt).resolve("worker-draft.txt"), "work that exists nowhere else\n");
|
||||||
|
|
||||||
|
Optional<String> ref = gitWorktrees.snapshot(wt, branch, "snapshot with unreachable content");
|
||||||
|
assertTrue(ref.isPresent());
|
||||||
|
|
||||||
|
// Age floor 0 makes age a non-issue: only reachability can save it — and it must.
|
||||||
|
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), 0),
|
||||||
|
"the unreachable snapshot is the last copy and must not be pruned");
|
||||||
|
assertTrue(refExists(repo, "refs/wip/" + branch),
|
||||||
|
"an unreachable snapshot must survive the sweep");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586, criterion 1 (the reachable half). A snapshot whose tree content IS already reachable
|
||||||
|
* from {@code main} and which is older than the age floor is pure duplication — the work is
|
||||||
|
* recovered — so it must be pruned.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void aReachableSnapshotOlderThanTheFloorIsPruned(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = initRepo(tmp.resolve("repo"));
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
|
||||||
|
// A snapshot whose tree is exactly main's current tree: fully reachable from main.
|
||||||
|
String mainTree = revParse(repo, "main^{tree}");
|
||||||
|
String old = commitTree(repo, mainTree, revParse(repo, "HEAD"), "2020-01-01T00:00:00", "snapshot");
|
||||||
|
updateRef(repo, "refs/wip/recovered", old);
|
||||||
|
|
||||||
|
assertEquals(1, gitWorktrees.pruneWipRefs(repo.toString(), TimeUnit.HOURS.toMillis(24)),
|
||||||
|
"an old, main-reachable snapshot must be pruned");
|
||||||
|
assertFalse(refExists(repo, "refs/wip/recovered"),
|
||||||
|
"the reachable snapshot's ref must be gone after the sweep");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586, the age floor. A snapshot whose content IS reachable from {@code main} but which is
|
||||||
|
* younger than the age floor must not be swept — a lead may still be looking at it.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void aReachableButRecentSnapshotIsNotPruned(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = initRepo(tmp.resolve("repo"));
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
|
||||||
|
// Reachable from main, but committed "now" — a fresh snapshot. The 24h floor must protect it.
|
||||||
|
String mainTree = revParse(repo, "main^{tree}");
|
||||||
|
String fresh = commitTree(repo, mainTree, revParse(repo, "HEAD"),
|
||||||
|
"2038-01-01T00:00:00", "snapshot just taken");
|
||||||
|
updateRef(repo, "refs/wip/fresh", fresh);
|
||||||
|
|
||||||
|
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), TimeUnit.HOURS.toMillis(24)),
|
||||||
|
"a recent snapshot must be kept even when reachable");
|
||||||
|
assertTrue(refExists(repo, "refs/wip/fresh"),
|
||||||
|
"the recent reachable snapshot must survive the sweep");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586, criterion 5. A fleet that has never snapshotted anything has no {@code refs/wip/*},
|
||||||
|
* so a sweep is a no-op and the census reports none — identical to before CB-586 existed.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void aFleetWithNoSnapshotsPrunesNothingAndReportsNothing(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = initRepo(tmp.resolve("repo"));
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
|
||||||
|
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), 0),
|
||||||
|
"no snapshot refs means nothing to prune");
|
||||||
|
Worktrees.WipRefStats stats = gitWorktrees.wipRefs(repo.toString());
|
||||||
|
assertEquals(0, stats.count(), "a never-snapshotted fleet has zero refs/wip refs");
|
||||||
|
assertEquals(0L, stats.costBytes(), "a never-snapshotted fleet costs zero bytes");
|
||||||
|
}
|
||||||
|
|
||||||
|
/** CB-586, criterion 4: the census reports how many refs exist and roughly what they cost. */
|
||||||
|
@Test
|
||||||
|
void wipRefsReportsCountAndCost(@TempDir Path tmp) throws Exception {
|
||||||
|
Path repo = initRepo(tmp.resolve("repo"));
|
||||||
|
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||||
|
|
||||||
|
String blob = blobOf(repo, "a recoverable snapshot's worth of content");
|
||||||
|
String tree = treeOf(repo, "snapshot.txt", blob);
|
||||||
|
updateRef(repo, "refs/wip/one", commitTree(repo, tree, revParse(repo, "HEAD"),
|
||||||
|
"2020-01-01T00:00:00", "snapshot"));
|
||||||
|
updateRef(repo, "refs/wip/two", commitTree(repo, tree, revParse(repo, "HEAD"),
|
||||||
|
"2020-01-02T00:00:00", "snapshot"));
|
||||||
|
|
||||||
|
Worktrees.WipRefStats stats = gitWorktrees.wipRefs(repo.toString());
|
||||||
|
assertEquals(2, stats.count(), "two snapshot refs are reported");
|
||||||
|
assertTrue(stats.costBytes() > 0, "the cost of the snapshots is a positive byte count");
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -138,6 +138,16 @@ class SessionManagerTest {
|
|||||||
return java.util.Optional.of("wip" + snapshotSeq.incrementAndGet());
|
return java.util.Optional.of("wip" + snapshotSeq.incrementAndGet());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public WipRefStats wipRefs(String repoRoot) {
|
||||||
|
return new WipRefStats(0, 0L);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
|
||||||
List<String> removeCalls() {
|
List<String> removeCalls() {
|
||||||
return List.copyOf(removeCalls);
|
return List.copyOf(removeCalls);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -11,10 +11,12 @@ import org.junit.jupiter.api.Test;
|
|||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
import java.util.Set;
|
import java.util.Set;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
import java.util.concurrent.atomic.AtomicLong;
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
|
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
|
||||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -90,6 +92,59 @@ class SessionReaperTest {
|
|||||||
return ticks.get() >= n;
|
return ticks.get() >= n;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586: the retention sweep must actually reach the seam when the loop runs.
|
||||||
|
*
|
||||||
|
* <p>The sweep's own tests call {@code SessionManager.sweepWipRefs} directly, which walks
|
||||||
|
* around the reaper's interval gate — and the gate is where it broke. {@code lastWipSweepNanos}
|
||||||
|
* started at {@code Long.MIN_VALUE}, so {@code now - lastWipSweepNanos} overflowed to a large
|
||||||
|
* negative number, the gate read that as "swept moments ago", and it returned <em>before</em>
|
||||||
|
* the assignment that would have fixed the field. The sweep never ran once, for the life of
|
||||||
|
* the process, and every direct-call test still passed.
|
||||||
|
*
|
||||||
|
* <p>So this asserts through the loop: spawn a worktree session (which is what tells the
|
||||||
|
* manager which repo holds {@code refs/wip/*}), start the reaper, and require a real prune call
|
||||||
|
* to arrive at the fake.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void theLoopActuallyRunsTheWipRetentionSweep() throws InterruptedException {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||||
|
SessionManager sessions = new SessionManager(launcher(herdr), worktrees);
|
||||||
|
// Without a worktree session the repo is unknown and the sweep is a legitimate no-op, so
|
||||||
|
// this step is what makes the assertion below meaningful rather than vacuous.
|
||||||
|
sessions.acquire("ltms-local", null, "/caller/proj", null, new WorktreeRequest("cb-586", null));
|
||||||
|
assertTrue(worktrees.pruneCalls().isEmpty(), "nothing has swept before the reaper starts");
|
||||||
|
|
||||||
|
SessionReaper reaper = new SessionReaper(sessions, IDLE_TTL_SECONDS, SHORT_INTERVAL_MILLIS);
|
||||||
|
reaper.start();
|
||||||
|
try {
|
||||||
|
long deadline = System.currentTimeMillis() + 3000;
|
||||||
|
while (worktrees.pruneCalls().isEmpty() && System.currentTimeMillis() < deadline) {
|
||||||
|
Thread.sleep(10);
|
||||||
|
}
|
||||||
|
} finally {
|
||||||
|
reaper.stop();
|
||||||
|
}
|
||||||
|
|
||||||
|
assertFalse(worktrees.pruneCalls().isEmpty(),
|
||||||
|
"the reaper loop must run the refs/wip retention sweep; it never reached the seam");
|
||||||
|
assertEquals("/repo", worktrees.pruneCalls().getFirst().repoRoot(),
|
||||||
|
"the sweep must target the repo the fleet's worktrees came from");
|
||||||
|
assertEquals(TimeUnit.HOURS.toMillis(24), worktrees.pruneCalls().getFirst().minAgeMillis(),
|
||||||
|
"the 24h age floor is the safety rule and must reach the seam intact");
|
||||||
|
}
|
||||||
|
|
||||||
|
/** As {@link #launcher()}, on a caller-supplied herdr so the test can inspect it. */
|
||||||
|
private static ClaudeCodeLauncher launcher(FakeHerdr herdr) {
|
||||||
|
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||||
|
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||||
|
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||||
|
"worker: {profile} #{n}", null, null, null);
|
||||||
|
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||||
|
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void stopIsIdempotentAndSafeBeforeStart() {
|
void stopIsIdempotentAndSafeBeforeStart() {
|
||||||
SessionReaper reaper = reaper();
|
SessionReaper reaper = reaper();
|
||||||
|
|||||||
@@ -422,6 +422,28 @@ class WorktreeSessionManagerTest {
|
|||||||
"acceptance criterion 6: a failed ticket's detail must carry the snapshot ref");
|
"acceptance criterion 6: a failed ticket's detail must carry the snapshot ref");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void releaseNotifiesTheListenerWithTheAgentSessionId() {
|
||||||
|
// CB-584 (issue #65 criterion 5): a failed ticket's detail must also name the agent
|
||||||
|
// session, alongside worktree/branch/snapshot, so a lead can resume the conversation
|
||||||
|
// rather than only re-dispatch a fresh member onto the same files.
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
|
||||||
|
.withDirty(true);
|
||||||
|
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||||
|
java.util.List<SessionManager.ReleaseDetail> released = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||||
|
sessions.onRelease(released::add);
|
||||||
|
MemberSession s = sessions.acquire("ltms-local", MemberRole.DEV, null, "/caller/proj", null,
|
||||||
|
new WorktreeRequest("cb-584-e", null), "cb-584-session", null);
|
||||||
|
|
||||||
|
assertNotNull(s.agentSessionId(), "a named session must mint an agent session id to assert on");
|
||||||
|
sessions.release(s.paneId());
|
||||||
|
|
||||||
|
assertEquals(1, released.size());
|
||||||
|
SessionManager.ReleaseDetail detail = released.getFirst();
|
||||||
|
assertEquals(s.agentSessionId(), detail.agentSessionId());
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void aFailingSnapshotStillPreservesTheWorktreeStopsThePaneAndNotifies() {
|
void aFailingSnapshotStillPreservesTheWorktreeStopsThePaneAndNotifies() {
|
||||||
FakeHerdr herdr = new FakeHerdr();
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
@@ -444,4 +466,37 @@ class WorktreeSessionManagerTest {
|
|||||||
"a failed snapshot leaves no ref to report");
|
"a failed snapshot leaves no ref to report");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-586, criteria 4 and 1. Once a worktree session is spawned the repo is known, so the
|
||||||
|
* operator census and the retention sweep delegate to that repo's {@code refs/wip/*}.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void wipRefsAndSweepDelegateToTheFleetRepoOnceKnown() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
|
||||||
|
.withWipRefs(new Worktrees.WipRefStats(3, 42L)).withPruneResult(2);
|
||||||
|
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||||
|
|
||||||
|
assertTrue(sessions.wipRefs().isEmpty(),
|
||||||
|
"no worktree spawned yet means no repo is known and nothing to report");
|
||||||
|
assertEquals(0, sessions.sweepWipRefs(TimeUnit.HOURS.toMillis(24)),
|
||||||
|
"no worktree spawned yet means the sweep is a no-op");
|
||||||
|
|
||||||
|
sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||||
|
new WorktreeRequest("cb-586", null));
|
||||||
|
|
||||||
|
Worktrees.WipRefStats stats = sessions.wipRefs().orElseThrow();
|
||||||
|
assertEquals(3, stats.count(), "the census comes from the fleet repo");
|
||||||
|
assertEquals(42L, stats.costBytes(), "the cost comes from the fleet repo");
|
||||||
|
assertEquals(2, sessions.sweepWipRefs(TimeUnit.HOURS.toMillis(24)),
|
||||||
|
"the sweep runs against the fleet repo");
|
||||||
|
|
||||||
|
List<FakeWorktrees.PruneCall> prunes = worktrees.pruneCalls();
|
||||||
|
assertEquals(1, prunes.size(), "the no-op short-circuits before reaching the seam, so only "
|
||||||
|
+ "the repo-known sweep issues a call");
|
||||||
|
assertEquals("/repo", prunes.getFirst().repoRoot(), "the sweep targets the fleet repo");
|
||||||
|
assertEquals(TimeUnit.HOURS.toMillis(24), prunes.getFirst().minAgeMillis(),
|
||||||
|
"the caller's age floor is passed through");
|
||||||
|
}
|
||||||
|
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -82,8 +82,25 @@
|
|||||||
<key>RunAtLoad</key>
|
<key>RunAtLoad</key>
|
||||||
<true/>
|
<true/>
|
||||||
|
|
||||||
<!-- Restart on crash, but not in a tight loop if the config is bad (bridged fails fast on a
|
<!--
|
||||||
non-loopback bind without token auth — that is a config error, not a transient one). -->
|
CB-600 — read this before assuming ThrottleInterval bounds anything. It paces restarts to at
|
||||||
|
most one per 10s; it does NOT cap how many times launchd retries. If bridged fails fast on
|
||||||
|
every start — a bad bridged.yaml, for example auth.mode: token with the token env var unset,
|
||||||
|
which throws in main() before the daemon ever binds a port — launchd restarts it forever,
|
||||||
|
once every 10s, until a human intervenes. LaunchAgents have no "give up after N attempts"
|
||||||
|
primitive, so this is not something a config change here can fix.
|
||||||
|
|
||||||
|
That loop stops only two ways: (1) `launchctl unload -w ~/Library/LaunchAgents/dev.ltms.bridged.plist`,
|
||||||
|
or (2) the underlying cause gets fixed, so the process starts successfully and stays up (no
|
||||||
|
more exits to restart). scripts/redeploy-bridged.sh does not add a third way — it does not
|
||||||
|
make bridged self-disable on a config error, on purpose: a fail-fast exit path that
|
||||||
|
sometimes decides "this is unrecoverable, stop trying" is one more thing that can misfire,
|
||||||
|
and a wrongly self-disabled daemon needs the exact same manual `launchctl load -w` recovery
|
||||||
|
this comment already names — so it buys nothing an operator watching for the crash loop
|
||||||
|
doesn't already have, at the cost of a new way to be silently down. Watch for it with
|
||||||
|
`launchctl list dev.ltms.bridged` (a high restart count) or by tailing bridged.out for the
|
||||||
|
same startup error repeating every ~10s.
|
||||||
|
-->
|
||||||
<key>KeepAlive</key>
|
<key>KeepAlive</key>
|
||||||
<dict>
|
<dict>
|
||||||
<key>SuccessfulExit</key>
|
<key>SuccessfulExit</key>
|
||||||
|
|||||||
+19
-4
@@ -1,9 +1,24 @@
|
|||||||
# MCP Contract — `bridged`'s unified gateway
|
# MCP Contract — `bridged`'s unified gateway
|
||||||
|
|
||||||
> **Status:** 🟡 Design (2026-07-14). Greenfield — no MCP code exists yet; the pom carries
|
> **Status: 🔴 HISTORICAL DESIGN — do NOT use as the tool reference.** Written 2026-07-14, before
|
||||||
> only Javalin/Jackson. This page defines the tool surface that CB-104 and its followers
|
> any MCP code existed. The system shipped and this page never caught up, so **its tool names,
|
||||||
> implement. It supersedes nothing; it fills the "MCP server face" left open by the
|
> parameter names and REST paths are wrong today**. Audited 2026-08-17; the specific drift:
|
||||||
> [Architecture](1-Architecture) page.
|
>
|
||||||
|
> - **Tools it names that do not exist:** `bridge_read`, `bridge_cancel`.
|
||||||
|
> - **Shipped tools it omits:** `bridge_poll`, `bridge_ack`, `bridge_profiles`, `bridge_whoami`.
|
||||||
|
> - **Parameter names are wrong nearly everywhere** — it says `message`/`target`/`timeout_seconds`/
|
||||||
|
> `block` where the code takes `content`/`sessionId`/`timeoutMs`/`wait`; `text` where
|
||||||
|
> `bridge_reply` takes `content`; `target` where `bridge_stop` takes `paneId`.
|
||||||
|
> - **REST paths are wrong:** it says `POST /workers` and `DELETE /workers/{paneId}`; the daemon
|
||||||
|
> serves `POST /members` and `DELETE /members/{paneId}`.
|
||||||
|
>
|
||||||
|
> **The authoritative tool surface is the live MCP schema** (each tool's own description and
|
||||||
|
> parameters, as mounted), with the intent→tool table in `CLAUDE.md` as the short form. Both were
|
||||||
|
> checked against `mcp/BridgeMcp.java` on 2026-08-17 and are accurate.
|
||||||
|
>
|
||||||
|
> What is still worth reading here is **§6 — the flows and the error model** (rendezvous,
|
||||||
|
> `bridge_ask`, detached delivery, the turn-done fallback). The shapes it describes are the ones
|
||||||
|
> that shipped; only the names around them drifted. Rewriting this page is tracked as **CB-609**.
|
||||||
|
|
||||||
`bridged` is the **sole communication gateway** for every Claude session in the bridge. Both
|
`bridged` is the **sole communication gateway** for every Claude session in the bridge. Both
|
||||||
the **primary** (Opus, on subscription) and every **worker** (off-subscription Claude Code)
|
the **primary** (Opus, on subscription) and every **worker** (off-subscription Claude Code)
|
||||||
|
|||||||
+2
-2
@@ -17,7 +17,7 @@ who the workers are, how the lead picks one, and how it runs many at once.
|
|||||||
- **Workers** — a herd of `claude` panes in herdr, each an addressable `bridged` session
|
- **Workers** — a herd of `claude` panes in herdr, each an addressable `bridged` session
|
||||||
with its **own model/env**:
|
with its **own model/env**:
|
||||||
- **Claude workers** (clean env, e.g. Sonnet) — reasoning-heavy or high-accuracy subtasks.
|
- **Claude workers** (clean env, e.g. Sonnet) — reasoning-heavy or high-accuracy subtasks.
|
||||||
- **Local workers** (`ANTHROPIC_BASE_URL=https://ollama.ltms.dev`) — bulk, cheap, or
|
- **Local workers** (`ANTHROPIC_BASE_URL=https://llm.ltms.dev/anthropic`) — bulk, cheap, or
|
||||||
embarrassingly parallel subtasks.
|
embarrassingly parallel subtasks.
|
||||||
|
|
||||||
Every worker is still a *real Claude Code process* (inherits `CLAUDE.md`, hooks, skills,
|
Every worker is still a *real Claude Code process* (inherits `CLAUDE.md`, hooks, skills,
|
||||||
@@ -35,7 +35,7 @@ flowchart TB
|
|||||||
WL1["w-local-1<br/>ANTHROPIC_BASE_URL set"]
|
WL1["w-local-1<br/>ANTHROPIC_BASE_URL set"]
|
||||||
WL2["w-local-2<br/>ANTHROPIC_BASE_URL set"]
|
WL2["w-local-2<br/>ANTHROPIC_BASE_URL set"]
|
||||||
ANT["api.anthropic.com<br/>(Pro/Max)"]
|
ANT["api.anthropic.com<br/>(Pro/Max)"]
|
||||||
OLL["ollama.ltms.dev<br/>(local model)"]
|
OLL["llm.ltms.dev<br/>(gateway to the local model)"]
|
||||||
|
|
||||||
LEAD -->|"blocking POST /message (target role)"| BD
|
LEAD -->|"blocking POST /message (target role)"| BD
|
||||||
BD -->|"Unix socket · send_text · events.subscribe"| HERDR
|
BD -->|"Unix socket · send_text · events.subscribe"| HERDR
|
||||||
|
|||||||
+1
-1
@@ -4,7 +4,7 @@
|
|||||||
"CLAUDE.md"
|
"CLAUDE.md"
|
||||||
],
|
],
|
||||||
"mcp": {
|
"mcp": {
|
||||||
"bridged": {
|
"fleetd": {
|
||||||
"type": "remote",
|
"type": "remote",
|
||||||
"url": "http://127.0.0.1:8765/mcp",
|
"url": "http://127.0.0.1:8765/mcp",
|
||||||
"enabled": true
|
"enabled": true
|
||||||
|
|||||||
Executable
+135
@@ -0,0 +1,135 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
#
|
||||||
|
# CB-596 step 1: measure which credentials a live member actually holds.
|
||||||
|
#
|
||||||
|
# The claim under test is that a member's herdr pane starts a LOGIN shell, that shell sources
|
||||||
|
# ${SHARED_ENV}/tools/secrets.sh, and so a member inherits every name that file exports — while
|
||||||
|
# CB-592 blocks exactly one of them (GITEA_ACCESS_TOKEN). That is an inference from the code, not a
|
||||||
|
# measurement, and issue #82 says plainly: do not build a fix on the inference. This is the
|
||||||
|
# measurement.
|
||||||
|
#
|
||||||
|
# WHY THIS IS A SCRIPT AND NOT A COMMAND SOMEONE TYPES
|
||||||
|
#
|
||||||
|
# Enumerating credential names inside a member is exactly the action that should need the operator's
|
||||||
|
# explicit approval, and the command classifier refuses it. That refusal is correct. This script is
|
||||||
|
# the seam: it is one auditable file the operator can read once, top to bottom, and then run — rather
|
||||||
|
# than approving an ad-hoc shell pipeline whose behaviour they have to take on trust.
|
||||||
|
#
|
||||||
|
# WHAT IT WILL NOT DO
|
||||||
|
#
|
||||||
|
# * It never prints a credential value, and never any prefix or suffix of one. Not one character.
|
||||||
|
# Issue #82's criterion 1 asked for a 6-character prefix; this prints a truncated SHA-256 instead.
|
||||||
|
# A prefix of a short secret is most of the secret, and it would end up pasted into a ticket. The
|
||||||
|
# hash answers every question the prefix was for — is it set, is it the same value as over there,
|
||||||
|
# is it the CB-592 sentinel — and answers none of the ones it should not.
|
||||||
|
# * It never writes anywhere, never contacts the network, and never touches secrets.sh, which is
|
||||||
|
# the operator's file.
|
||||||
|
#
|
||||||
|
# HOW TO RUN IT
|
||||||
|
#
|
||||||
|
# 1. As the operator, in a member's pane (a spawned worker's terminal):
|
||||||
|
# bash scripts/probe-member-credentials.sh
|
||||||
|
# 2. For the comparison row, in your OWN shell — a lead, not a member:
|
||||||
|
# bash scripts/probe-member-credentials.sh --allow-outside-member
|
||||||
|
#
|
||||||
|
# The two outputs side by side are the finding: any name whose hash matches between them is a
|
||||||
|
# credential the member holds in full.
|
||||||
|
#
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
# The names ${SHARED_ENV}/tools/secrets.sh exports, recorded on 2026-08-16 (issue #82). Names only —
|
||||||
|
# this list contains no values and never should. If secrets.sh gains a name, this list goes stale and
|
||||||
|
# the probe silently stops asking about it; that staleness is itself part of what #82's criterion 4
|
||||||
|
# has to solve, so it is called out in the summary rather than hidden.
|
||||||
|
NAMES=(
|
||||||
|
AI_GATEWAY_TOKEN BESZEL_ADMIN_EMAIL BESZEL_ADMIN_PASSWORD
|
||||||
|
BESZEL_HUB_URL BESZEL_KEY BESZEL_UNIVERSAL_TOKEN
|
||||||
|
BRAIN_MCP_TOKEN CF_ACCOUNT_ID CF_API_TOKEN
|
||||||
|
CF_USER_TOKEN CONFLUENCE_API_TOKEN CONFLUENCE_USERNAME
|
||||||
|
CONTEXT7_TOKEN GITEA_HOST GITLAB_OAUTH_CLIENT_SECRET
|
||||||
|
GITLAB_PERSONAL_ACCESS_TOKEN GRAFANA_ADMIN_PASSWORD GRAFANA_ADMIN_USER
|
||||||
|
HASS_TOKEN HW_PASSWORD HW_USER
|
||||||
|
LTMS_API_KEY MEMORY_MCP_TOKEN METRICS_PUSH_TOKEN
|
||||||
|
OPENCODE_AUTOMODE_MODEL TELEGRAM_BOT_TOKEN TELEGRAM_CHAT_ID
|
||||||
|
TS_API_KEY TS_AUTHKEY WORKER_GITEA_TOKEN
|
||||||
|
GITEA_ACCESS_TOKEN
|
||||||
|
)
|
||||||
|
|
||||||
|
allow_outside=0
|
||||||
|
for arg in "$@"; do
|
||||||
|
case "$arg" in
|
||||||
|
--allow-outside-member) allow_outside=1 ;;
|
||||||
|
-h|--help) sed -n '2,40p' "$0"; exit 0 ;;
|
||||||
|
*) echo "unknown argument: $arg" >&2; exit 2 ;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
|
||||||
|
if [ "${BRIDGED_MEMBER:-}" != "1" ] && [ "$allow_outside" -eq 0 ]; then
|
||||||
|
cat >&2 <<'EOF'
|
||||||
|
refusing to run: BRIDGED_MEMBER is not 1, so this is not a member's shell.
|
||||||
|
|
||||||
|
The finding this probe exists for is what a MEMBER holds. Run it in a spawned worker's pane. If you
|
||||||
|
meant to take the comparison reading from your own shell, pass --allow-outside-member and the output
|
||||||
|
will be labelled as such.
|
||||||
|
EOF
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Prefer sha256sum (Linux), fall back to shasum (macOS). If neither exists, report presence and
|
||||||
|
# length only — degraded, but never a value.
|
||||||
|
hasher=""
|
||||||
|
if command -v sha256sum >/dev/null 2>&1; then
|
||||||
|
hasher="sha256sum"
|
||||||
|
elif command -v shasum >/dev/null 2>&1; then
|
||||||
|
hasher="shasum -a 256"
|
||||||
|
fi
|
||||||
|
|
||||||
|
digest() { # value -> first 12 hex chars of its sha256, or "-" when no hasher is available
|
||||||
|
[ -z "$hasher" ] && { printf '%s' "-"; return; }
|
||||||
|
printf '%s' "$1" | $hasher | cut -c1-12
|
||||||
|
}
|
||||||
|
|
||||||
|
if [ "${BRIDGED_MEMBER:-}" = "1" ]; then
|
||||||
|
where="MEMBER (BRIDGED_MEMBER=1)"
|
||||||
|
else
|
||||||
|
where="NOT a member — comparison reading only"
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo "CB-596 credential probe"
|
||||||
|
echo "reading from : $where"
|
||||||
|
echo "shell : ${SHELL:-unknown}"
|
||||||
|
echo "hash : ${hasher:-none available — lengths only}"
|
||||||
|
# Only printed so the two readings can be told apart when they are pasted side by side.
|
||||||
|
echo "host : $(hostname 2>/dev/null || echo unknown)"
|
||||||
|
echo
|
||||||
|
printf '%-30s %-7s %6s %s\n' "NAME" "STATE" "LEN" "SHA256-12"
|
||||||
|
printf '%-30s %-7s %6s %s\n' "------------------------------" "-------" "------" "------------"
|
||||||
|
|
||||||
|
set_count=0
|
||||||
|
for name in "${NAMES[@]}"; do
|
||||||
|
value="${!name:-}"
|
||||||
|
if [ -z "$value" ]; then
|
||||||
|
printf '%-30s %-7s %6s %s\n' "$name" "unset" "-" "-"
|
||||||
|
else
|
||||||
|
set_count=$((set_count + 1))
|
||||||
|
printf '%-30s %-7s %6s %s\n' "$name" "SET" "${#value}" "$(digest "$value")"
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
|
||||||
|
echo
|
||||||
|
echo "$set_count of ${#NAMES[@]} names are set in this shell."
|
||||||
|
echo
|
||||||
|
cat <<'EOF'
|
||||||
|
How to read this:
|
||||||
|
|
||||||
|
* Take the MEMBER reading and the comparison reading, and line them up. A name whose SHA256-12
|
||||||
|
matches on both sides is a credential the member holds in full. That is the finding.
|
||||||
|
* GITEA_ACCESS_TOKEN is the control. CB-592 replaces it with a blocked sentinel, so its hash
|
||||||
|
should DIFFER between the two readings. If it matches, CB-592 is not working and that is the
|
||||||
|
most urgent thing on this page.
|
||||||
|
* AI_GATEWAY_TOKEN matching is expected and correct, not a leak: bridged.yaml names it in
|
||||||
|
`tokenEnv:` for the local and gx profiles, so a member reaching the gateway is by design.
|
||||||
|
* A name that is set here but is NOT in the list above will not appear at all. The list was
|
||||||
|
recorded on 2026-08-16 and does not update itself. Anything added to secrets.sh since then is
|
||||||
|
invisible to this probe — which is the same gap issue #82 criterion 4 asks to close properly.
|
||||||
|
EOF
|
||||||
@@ -79,6 +79,53 @@ running_pid() { pgrep -f "$PATTERN" || true; }
|
|||||||
launchd_installed() { [ -f "$LAUNCHD_PLIST" ]; }
|
launchd_installed() { [ -f "$LAUNCHD_PLIST" ]; }
|
||||||
launchd_loaded() { launchctl list "$LAUNCHD_LABEL" >/dev/null 2>&1; }
|
launchd_loaded() { launchctl list "$LAUNCHD_LABEL" >/dev/null 2>&1; }
|
||||||
|
|
||||||
|
# CB-600: the script computes its own log path from where it sits on disk (REPO, above); the
|
||||||
|
# plist hard-codes an absolute StandardOutPath. Nothing forced the two to agree — if this script
|
||||||
|
# were ever run from a checkout other than the one the loaded plist names, launchd would start and
|
||||||
|
# log the daemon correctly, while every check below (the fresh "bridged listening" line, the
|
||||||
|
# ERROR-count scan) would read a different, empty or stale file and the script would report a
|
||||||
|
# clean restart while the daemon crash-loops. Pure and side-effect-free besides `die`/`ok` — reads
|
||||||
|
# the two paths, resolves them, compares — so it never touches launchd or the daemon and can be
|
||||||
|
# exercised by sourcing this script (see the SOURCED guard below) without installing the agent.
|
||||||
|
check_log_path_matches_plist() {
|
||||||
|
local script_out="$1" plist_path="$2"
|
||||||
|
local plist_out resolved_out resolved_plist_out
|
||||||
|
# Checked by exit status, not by emptiness: on a missing file/key PlistBuddy exits nonzero but
|
||||||
|
# still writes a message ("File Doesn't Exist, Will Create: ...") that command substitution
|
||||||
|
# would happily capture as if it were the real value — testing only `-z` missed that case.
|
||||||
|
if ! plist_out="$(/usr/libexec/PlistBuddy -c 'Print :StandardOutPath' "$plist_path" 2>/dev/null)" \
|
||||||
|
|| [ -z "$plist_out" ]; then
|
||||||
|
die "launchd agent is loaded but PlistBuddy could not read StandardOutPath from
|
||||||
|
$plist_path
|
||||||
|
— cannot verify the daemon logs where this script is about to look. Fix the plist before
|
||||||
|
redeploying supervised."
|
||||||
|
fi
|
||||||
|
resolved_out="$(cd "$(dirname "$script_out")" 2>/dev/null && pwd -P)/$(basename "$script_out")" || true
|
||||||
|
resolved_plist_out="$(cd "$(dirname "$plist_out")" 2>/dev/null && pwd -P)/$(basename "$plist_out")" || true
|
||||||
|
if [ -z "$resolved_out" ] || [ -z "$resolved_plist_out" ] || [ "$resolved_out" != "$resolved_plist_out" ]; then
|
||||||
|
die "log path mismatch — this script reads
|
||||||
|
$script_out (resolved: ${resolved_out:-<directory does not exist>})
|
||||||
|
but the loaded plist's StandardOutPath is
|
||||||
|
$plist_out (resolved: ${resolved_plist_out:-<directory does not exist>})
|
||||||
|
Under supervision the daemon writes to the PLIST's path, not necessarily this script's — every
|
||||||
|
post-restart check below (the fresh 'bridged listening' line, the ERROR-count scan) would read
|
||||||
|
the wrong file and could report a clean restart while the daemon crash-loops. Fix the mismatch
|
||||||
|
(move this checkout to match the plist, or edit the plist's StandardOutPath/StandardErrorPath)
|
||||||
|
before redeploying supervised."
|
||||||
|
fi
|
||||||
|
ok "log path check: script and plist agree ($resolved_out)"
|
||||||
|
}
|
||||||
|
|
||||||
|
# CB-600: sourceable for testing. When this file is SOURCED (not executed) it stops here — nothing
|
||||||
|
# below runs — so a test harness can `source` it to call check_log_path_matches_plist (or the
|
||||||
|
# other pure helpers above) against a throwaway plist fixture without ever reaching the mutating
|
||||||
|
# flow (build/stop/start) or touching the real daemon or launchd. On a normal `./redeploy-bridged.sh`
|
||||||
|
# invocation `(return 0 2>/dev/null)` fails (return is illegal at top level of an executed script),
|
||||||
|
# so this whole block is a no-op and every line below still runs exactly as before.
|
||||||
|
if (return 0 2>/dev/null); then
|
||||||
|
return 0
|
||||||
|
fi
|
||||||
|
|
||||||
# ---------------------------------------------------------------- report state
|
# ---------------------------------------------------------------- report state
|
||||||
|
|
||||||
say "current state"
|
say "current state"
|
||||||
@@ -103,6 +150,9 @@ SUPERVISED=0
|
|||||||
if launchd_loaded; then
|
if launchd_loaded; then
|
||||||
SUPERVISED=1
|
SUPERVISED=1
|
||||||
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
|
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
|
||||||
|
# CB-600: fail loudly here, before ANY other check runs, if this script and the loaded plist
|
||||||
|
# would read different log files — every check after this point is worthless otherwise.
|
||||||
|
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
|
||||||
else
|
else
|
||||||
warn "launchd agent not loaded — this script is the only thing that will restart the daemon."
|
warn "launchd agent not loaded — this script is the only thing that will restart the daemon."
|
||||||
fi
|
fi
|
||||||
@@ -222,7 +272,22 @@ say "start"
|
|||||||
if [ "$SUPERVISED" = 1 ]; then
|
if [ "$SUPERVISED" = 1 ]; then
|
||||||
echo " supervision is ON: using 'launchctl load' so launchd starts and keeps supervising this"
|
echo " supervision is ON: using 'launchctl load' so launchd starts and keeps supervising this"
|
||||||
echo " process, instead of a manual nohup that launchd would know nothing about."
|
echo " process, instead of a manual nohup that launchd would know nothing about."
|
||||||
launchctl load -w "$LAUNCHD_PLIST" || die "launchctl load failed"
|
# CB-600: 'launchctl unload -w' above already persisted Disabled=true for this label. A load -w
|
||||||
|
# that succeeds clears it; a load -w that FAILS leaves the agent both stopped and disabled — worse
|
||||||
|
# than before this script ran, because a later reboot or login will not bring it back either. One
|
||||||
|
# retry covers a transient race (e.g. launchd not yet fully done deregistering); if it still fails,
|
||||||
|
# die with the exact recovery command rather than a bare "failed".
|
||||||
|
if ! launchctl load -w "$LAUNCHD_PLIST" 2>/dev/null; then
|
||||||
|
warn "launchctl load failed on the first attempt — retrying once after a short pause"
|
||||||
|
sleep 2
|
||||||
|
launchctl load -w "$LAUNCHD_PLIST" || die "launchctl load failed twice.
|
||||||
|
The agent is now STOPPED and DISABLED — it will NOT come back on its own, not even after a
|
||||||
|
reboot or login, because 'launchctl unload -w' above persisted Disabled=true and load -w
|
||||||
|
never got the chance to clear it. Recover with:
|
||||||
|
launchctl load -w \"$LAUNCHD_PLIST\"
|
||||||
|
If that still fails, check 'launchctl list $LAUNCHD_LABEL', validate the plist with
|
||||||
|
'plutil -lint \"$LAUNCHD_PLIST\"', and check $OUT before assuming a retry will succeed."
|
||||||
|
fi
|
||||||
else
|
else
|
||||||
# Absolute jar path so `ps` names which checkout is running.
|
# Absolute jar path so `ps` names which checkout is running.
|
||||||
( cd "$BRIDGED" && zsh -lc "nohup java -jar '$JAR' >> bridged.out 2>&1 &" )
|
( cd "$BRIDGED" && zsh -lc "nohup java -jar '$JAR' >> bridged.out 2>&1 &" )
|
||||||
|
|||||||
+1
-1
Submodule wiki updated: 7c50cce52e...aa750de78e
Reference in New Issue
Block a user