Compare commits
122 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| a0505dc614 | |||
| d2f30f1654 | |||
| 73137f198f | |||
| 4ffe49f3bb | |||
| aabecce901 | |||
| 6754b4edbc | |||
| 787ae0ed7a | |||
| e3050efe8b | |||
| 11998cd626 | |||
| 8e5394f63f | |||
| 428a12af62 | |||
| a332dfdb2c | |||
| d0f4ae057b | |||
| 7b3beaa209 | |||
| b3b2bf3da6 | |||
| 8bb2aa0be4 | |||
| 4ce3149bfd | |||
| 1fc9e85bf1 | |||
| 38544d467c | |||
| 809b7d9b20 | |||
| 8cf7215d56 | |||
| 1ef93e57cc | |||
| cf0c9b9316 | |||
| 337dbd491e | |||
| 7cf6075b79 | |||
| fbdcd709c9 | |||
| 8d3f10d291 | |||
| 38f4fd64ee | |||
| 3fc39b981d | |||
| 70328ca0f8 | |||
| efab9b8c49 | |||
| 2374de28e4 | |||
| dd18bd1f38 | |||
| efeffb4ab7 | |||
| ed4f4b08ad | |||
| 28a1f3d6f5 | |||
| b12d70716b | |||
| 0f5985b419 | |||
| 6e06058b07 | |||
| e5f4fb81ab | |||
| d28ab0968b | |||
| 9a64d42599 | |||
| f4e0ca41e6 | |||
| 29a2f97c25 | |||
| d2db8c7dc9 | |||
| 95311c6e8e | |||
| 10ab58e4fc | |||
| 0ba597e394 | |||
| c468953963 | |||
| 25d53e6ef7 | |||
| a9a37af957 | |||
| 7d497aa423 | |||
| b205bcc2aa | |||
| 11eccc3a1b | |||
| 4939d40362 | |||
| e7a7711a4e | |||
| 1fd2cfa716 | |||
| 3e8e314656 | |||
| 20fc42b572 | |||
| f82073717a | |||
| 16b52fac6c | |||
| 13b6ae2628 | |||
| 7e838ba8b9 | |||
| c3e3554bde | |||
| b5bc5d4ab5 | |||
| d83972ede2 | |||
| 02c6909546 | |||
| b92a669ddc | |||
| bb29b001e4 | |||
| f0ff25221e | |||
| 780cb342ad | |||
| 5c563f02c8 | |||
| ad593c9bb9 | |||
| 736fd9cf4b | |||
| 05244a82b3 | |||
| ef4996a01e | |||
| edbd8d816a | |||
| 9425a9b696 | |||
| d0688c8a60 | |||
| 2c467c2553 | |||
| c6430d8edd | |||
| bfee23acc3 | |||
| 7dec74f1b4 | |||
| 482598e2a6 | |||
| 5f5d16fbd4 | |||
| 7df7985a16 | |||
| 724b35b46e | |||
| 2eb2d6112e | |||
| 7c458e8bf2 | |||
| 133f03e428 | |||
| 804279175d | |||
| 9dea289975 | |||
| cb4a6869b9 | |||
| 28b45d97e5 | |||
| 1a397e962e | |||
| 209e1231ea | |||
| 5d8b9d365c | |||
| a6aeda39e7 | |||
| 31b3c24caa | |||
| d105da978d | |||
| 5051a06443 | |||
| a134eccc57 | |||
| 6f275227d2 | |||
| 283ccf8423 | |||
| b96fba4a03 | |||
| 6d97d210b4 | |||
| 37b23cd704 | |||
| 367facf6a6 | |||
| 7f9a9c09f9 | |||
| e854957247 | |||
| b4b7cf5155 | |||
| cbb35ad947 | |||
| 656588f597 | |||
| e33377b2ca | |||
| 4b4a8688c2 | |||
| 7e48d4b86c | |||
| 41cc785534 | |||
| 60fa86a107 | |||
| f288cee2bb | |||
| a5d6ce1a37 | |||
| a52ca35d34 | |||
| 9417de1123 |
@@ -59,7 +59,7 @@ as `matches HEAD`, `drift`, or `unknown`; do not turn an unclear timestamp into
|
||||
Report the process identifier (PID) and uptime too:
|
||||
|
||||
```bash
|
||||
PIDS="$(pgrep -f 'target/fleetd.jar' || true)"
|
||||
PIDS="$(pgrep -f 'fleetd.jar' || true)"
|
||||
if [ -z "$PIDS" ]; then
|
||||
printf '%s\n' 'fleetd: not running'
|
||||
else
|
||||
|
||||
@@ -159,11 +159,41 @@ fails.
|
||||
|
||||
`{action: "cancel", token}` drops a pending request without rolling.
|
||||
|
||||
**A change is coming: the roll will restart the process instead of sending `/clear` (fleetd #726
|
||||
unit 2, written 2026-10-04).**
|
||||
|
||||
Today a roll types `/clear` into your pane. Your `claude` process keeps running, so a newer CLI on
|
||||
disk is never loaded. Unit 2 replaces that: the daemon ends the old pane, launches a fresh one,
|
||||
waits for the new terminal to be recognised as a lead, and only then sends the bootstrap text.
|
||||
|
||||
**Everything below about `/clear` is accurate while the old jar is running.** Unit 2 was not merged
|
||||
when this note was written, and a merge is not a deployment.
|
||||
|
||||
**How to tell which one is live: read your own tool list.** If `fleet_handover`'s description says
|
||||
it will "clear your pane", the daemon is serving the old behaviour. If it names a restart, the new
|
||||
behaviour is live. The description comes from the running daemon, so it cannot disagree with the
|
||||
code that is actually loaded.
|
||||
|
||||
Two things change for you once it is live. The `status` outcomes are different: three new failures
|
||||
replace the `/clear` ones. And the "never observed as WORKING after 8 consecutive IDLE/DONE polls"
|
||||
warning described below can no longer appear, because that wait is deleted — so if you still see
|
||||
it, the old jar is running. `TURN_NEVER_SETTLED` does not change, and still means nothing was
|
||||
touched.
|
||||
|
||||
**Delete this note and rewrite the `/clear` paragraphs once the new jar is live.**
|
||||
|
||||
**Things that will surprise you:**
|
||||
|
||||
- **`accepted` does not mean your pane has been cleared.** It means every gate passed and the roll
|
||||
is scheduled to run once your current turn ends. Say your goodbye in the same turn — you will not
|
||||
get another one.
|
||||
- **If you are still running after that turn, the roll did not happen.** A roll that works clears
|
||||
you, so surviving your own goodbye is itself the signal that it refused. Check with
|
||||
`fleet_handover{action: "status", token}`, using the token you confirmed. `TURN_NEVER_SETTLED`
|
||||
means your turn ran past `leadRollover.turnSettleSeconds` and **no `/clear` was ever sent**: your
|
||||
context is intact and nothing was lost. Open a fresh request and retry. Never assume the roll
|
||||
succeeded because `confirm` answered `accepted` — by the time it refuses, there is no caller left
|
||||
to tell, so this check is the only thing that closes that gap.
|
||||
- **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved
|
||||
from your connection, so you can only ever roll yourself.
|
||||
- **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you
|
||||
@@ -193,13 +223,18 @@ fails.
|
||||
|
||||
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
|
||||
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
|
||||
- **The bootstrap prompt works end to end. Measured 2026-09-22.** This used to say the fix was
|
||||
unproven (fleetd #489) and told you to expect a failure. That is no longer true. The daemon log
|
||||
now holds four `lead-rollover: rolled` lines, and three of them ran on 2026-09-22 at 10:01:43,
|
||||
10:38:28 and 11:15:47. Each one cleared the old lead and started a fresh session against the
|
||||
handover file, with the configured `bootstrapText` arriving as its first message. No context was
|
||||
lost. The old `Unknown command: /clearFresh` failure from 2026-09-12 does not appear in the log
|
||||
at all. Re-measure both numbers with:
|
||||
- **The bootstrap prompt works end to end. Measured 2026-09-22, re-measured 2026-10-04.** This used
|
||||
to say the fix was unproven (fleetd #489) and told you to expect a failure. That is no longer
|
||||
true. On 2026-09-22 the daemon log held four `lead-rollover: rolled` lines. On 2026-10-04 it holds
|
||||
**20**, against a control of 86 `lead-rollover:` lines. Each roll cleared the old lead and started
|
||||
a fresh session against the handover file, with the configured `bootstrapText` arriving as its
|
||||
first message. No context was lost. The old `Unknown command: /clearFresh` failure from 2026-09-12
|
||||
does not appear in the log at all.
|
||||
|
||||
19 of the 20 carry an `elapsedMs`: median 16507 ms, maximum 48261 ms, and two above 45000 ms. That
|
||||
figure times the **whole** roll, and the wait for your own turn to end dominates it, so do not
|
||||
read it as the cost of the clear. Expect a roll to take tens of seconds, and do not treat a slow
|
||||
one as a failed one. Re-measure all of these with:
|
||||
|
||||
```bash
|
||||
grep -c "lead-rollover: rolled" fleetd/fleetd.out # successful rolls
|
||||
|
||||
@@ -22,13 +22,22 @@ scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already chec
|
||||
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
|
||||
```
|
||||
|
||||
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
|
||||
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
|
||||
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
|
||||
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
|
||||
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
|
||||
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
|
||||
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
|
||||
`--no-build` skips the build and restarts whatever jar is at `fleetd/run/fleetd.jar` — the runtime
|
||||
path, not Maven's output path. Use it only when you just built and nothing changed since. It gives
|
||||
up the protection in the next paragraph: no build runs, so a stale or missing jar is not caught
|
||||
early. The script still checks the file is there and dies with `no jar at … — run without
|
||||
--no-build` if it is not, but it cannot tell you the jar is old.
|
||||
|
||||
The daemon runs from `fleetd/run/fleetd.jar`, not from `fleetd/target/fleetd.jar` where Maven
|
||||
writes its output (fleetd #664). That split is what makes a bare `mvn install`/`mvn clean` in the
|
||||
main clone harmless now: neither can reach the file the running daemon holds open, because that
|
||||
file no longer lives under `target/` at all. Verify a merge by building in a throwaway git
|
||||
worktree anyway — a build still produces nothing the fleet runs until this script's own `mv` of
|
||||
`target/fleetd.jar` onto `run/fleetd.jar`, performed only after the old daemon is confirmed gone.
|
||||
Let only `scripts/redeploy-fleetd.sh` touch `fleetd/run/fleetd.jar`. Run `--check` first: it prints
|
||||
the BUILT jar (`target/fleetd.jar`) and the RUNNING jar (`run/fleetd.jar`) as two separately
|
||||
labelled hash-and-mtime facts, so a mismatch between them — a build sitting unswapped, or a stale
|
||||
runtime jar — is visible before you decide anything.
|
||||
|
||||
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
|
||||
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
|
||||
|
||||
@@ -27,23 +27,34 @@ through its `fleet_*` tools. No session addresses a peer, a broker, or the netwo
|
||||
**Every role reads this file.** A member runs in a git worktree of this same repo, so it inherits
|
||||
this `CLAUDE.md` verbatim, and every rule below is role-conditional.
|
||||
|
||||
**Call `fleet_whoami`.** It returns `primary`, `worker`, or `architect`, resolved by the daemon from
|
||||
your connection — unforgeable, and the same resolution its authorization gate uses. A worker also
|
||||
carries its `sessionId`, `profile`, `worktree` and `branch`; an architect carries the slot name it
|
||||
was bound to. Don't infer what you can ask.
|
||||
**Call `fleet_whoami`.** It returns `primary`, `worker`, `architect`, `collaborator`, or `observer`,
|
||||
resolved by the daemon from your connection — unforgeable, and the same resolution its authorization
|
||||
gate uses. A worker also carries its `sessionId`, `profile`, `worktree` and `branch`; an architect
|
||||
carries the slot name it was bound to; a collaborator carries its registry name and its own
|
||||
`sessionId`, and **no `leader` key** — a collaborator is a named peer, not a primary. An **observer**
|
||||
carries only its own `sessionId`: a pane the daemon could not place as any of the above, authorized
|
||||
to `READ`/`METRICS` and to `REPLY`/`ASK` on its own pane and nothing more — never `SEND`, never a
|
||||
ticket. Don't infer what you can ask.
|
||||
|
||||
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
|
||||
fires: the reply charter in your system prompt (*"You are a spawned member in the
|
||||
claude-bridge fleet"*) ⇒ **spawned member**; fleet tools prefixed `mcp__fleet__*` ⇒ **spawned
|
||||
member** (the launcher fixes that mount name; a primary's mount is named by whoever wrote its
|
||||
`.mcp.json`, so it varies — and a member spawned before CB-632 still says `mcp__bridge__*`); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
|
||||
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect —
|
||||
only `fleet_whoami` does. **Still unsure ⇒ act as a worker**, the most restricted member role. The
|
||||
two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate —
|
||||
loud and self-correcting — while a member acting as the primary ends its turn with no `fleet_reply`,
|
||||
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect,
|
||||
or a worker from an **observer** — an observer is just as unspawned as a collaborator and carries
|
||||
none of these signals either, so only `fleet_whoami` tells the two apart. **And none of them fires
|
||||
for a collaborator at all**: every signal in the ladder detects a *spawned* member, while a
|
||||
collaborator is a tab a person opened by hand, so it has no charter, no fixed mount name and a
|
||||
normal environment. A collaborator — or an observer — that cannot call `fleet_whoami` therefore falls
|
||||
to the line below and acts as a worker. That is the safe direction — it under-privileges, and the
|
||||
refusals are loud — but it means a collaborator or an observer has no way to learn what it is except
|
||||
by asking. **Still unsure ⇒ act as a worker**, the most restricted member role this ladder can name.
|
||||
The two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate
|
||||
— loud and self-correcting — while a member acting as the primary ends its turn with no `fleet_reply`,
|
||||
and the sender silently receives nothing. Fail toward the recoverable error.
|
||||
|
||||
### Invariants — both roles, no exceptions
|
||||
### Invariants — every role, no exceptions
|
||||
|
||||
1. **Never set, export, or forward `ANTHROPIC_BASE_URL`** (or `ANTHROPIC_AUTH_TOKEN`). The primary
|
||||
stays on subscription; only the bridge puts a member off it, at spawn. Mounting the bridge must
|
||||
@@ -51,9 +62,10 @@ and the sender silently receives nothing. Fail toward the recoverable error.
|
||||
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
|
||||
side cannot see your screen. An answer that isn't in a `fleet_*` call is silently discarded.
|
||||
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
|
||||
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead or architect**;
|
||||
reply/ask are only-as-itself — any peer may answer for its own pane, and for no other. A call
|
||||
outside your role is refused, not queued.
|
||||
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead, architect, or
|
||||
collaborator** — and a collaborator may send only to a lead or another collaborator, never to a
|
||||
spawned member's terminal; reply/ask are only-as-itself — any peer may answer for its own pane,
|
||||
and for no other. A call outside your role is refused, not queued.
|
||||
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
|
||||
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
|
||||
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
|
||||
@@ -93,7 +105,11 @@ below are the procedure — run them in order, every task, not only the big ones
|
||||
5. **Collect** — `fleet_poll{ticket}` → `fleet_ack{target, msgId}`. Answer a worker's `fleet_ask`
|
||||
with `fleet_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
|
||||
`fleet_status`, never by reading its terminal; it also reports an open question and the `turnId`
|
||||
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
|
||||
that answers it — but **only to the caller that created that delegation**, so a question raised
|
||||
under an architect's brief is invisible to you, and seeing none does not mean there is none.
|
||||
**Only that same creator can answer it.** A `turnId` you came by any other way is refused, so an
|
||||
architect's worker waits for that architect and not for you.
|
||||
**A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
|
||||
brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
|
||||
**A correction cannot reach a busy member.** A `fleet_send` to a working member is *accepted* and
|
||||
returns a ticket, and is then never delivered — measured here three times in one session, and the
|
||||
@@ -158,9 +174,10 @@ you decide.
|
||||
| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) + `loopHealth` (`RUNNING`, `STALLED`, or `STOPPED` for `statusPoller` and `sessionReaper`) · one peer's state: `fleet_status{sessionId}` |
|
||||
| Delegate (blocking) | `fleet_send{sessionId, content}` |
|
||||
| Delegate (long task) | `fleet_send{sessionId, content, wait:false}` → ticket → `fleet_poll{ticket}` |
|
||||
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
|
||||
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId`, and only the caller that created that delegation |
|
||||
| Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
|
||||
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
|
||||
| Message a **collaborator** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` reports a `collaborators` array, and each row carries that peer's `name` and the `sessionId` you send to. It is visible to you, to an architect and to another collaborator, never to a worker. Coordination only, **never** a task |
|
||||
| Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}` if they are on this host. **Not** `fleet_reply`: it has no peer route and the publish is refused |
|
||||
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
|
||||
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
|
||||
@@ -234,6 +251,27 @@ simply complies has thrown away the reason there are two of you.
|
||||
7. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
|
||||
project marks as not-yours-to-commit.
|
||||
|
||||
### Collaborator — a named peer, not a member
|
||||
|
||||
`fleet_whoami` answered `collaborator`, so your pane's tab matches a `fleet.collaborators.<name>.tab`
|
||||
entry. You are **not** a member: nothing delegates to you, you have no brief, no worktree and no
|
||||
ticket, and **you owe no `fleet_reply`** — the turn contract above is for a session a lead spawned,
|
||||
and it does not apply to you. Read it only to understand what the members around you are doing.
|
||||
|
||||
What you may do: observe the fleet (`fleet_list`, `fleet_profiles`, `fleet_whoami`), and send to a
|
||||
lead or to another collaborator. What you may not: spawn, stop or drain anything, roll a lead's
|
||||
session, answer a member's `fleet_ask`, poll a ticket, or send to a spawned member's terminal. Each
|
||||
of those is refused at the gate, not queued.
|
||||
|
||||
Two limits worth knowing before you hit them. **You cannot reach a worker** — not even to help one —
|
||||
because a worker belongs to the lead that spawned it, and routing around that would make you a
|
||||
second orchestrator with no plan. Send to the lead instead. And **you cannot read a ticket**, so you
|
||||
cannot collect a delegation's reply: `fleet_poll` refuses you at the role gate, and a ticket also
|
||||
records the terminal that created it, so even a leaked id reads nothing.
|
||||
|
||||
Being named buys you a channel, not authority. Your `fleet_send` to a lead is coordination between
|
||||
peers: the lead owes you no obedience, and you owe it none.
|
||||
|
||||
### Where each rule lives (don't duplicate — extend the right layer)
|
||||
|
||||
| Layer | Scope | Reaches |
|
||||
@@ -301,10 +339,22 @@ must obey belongs in the charter, not here.
|
||||
(#362). **Read `plugin/` before designing anything about onboarding a project.** Two limits are
|
||||
structural, not bugs: a plugin cannot carry the role agent files, because
|
||||
`ClaudeCodeLauncher.java:371` requires `<cwd>/.claude/agents/<role>.md` in the member's own
|
||||
worktree; and a plugin cannot deliver anything to members at all, because
|
||||
`ClaudeCodeLauncher.java:285` exports `CLAUDE_CONFIG_DIR` and every Claude profile here sets it,
|
||||
so a member never reads the operator's plugin store. **The plugin is the lead-side surface;
|
||||
member-facing assets travel in the worktree.**
|
||||
worktree; and a plugin reaches a member only through `CLAUDE_CONFIG_DIR`, which
|
||||
`ClaudeCodeLauncher.java:286` exports with `putIfPresent` — so only for a profile that sets
|
||||
`configDir`. Every `claude-code` profile does set one (the four without are `opencode`, which
|
||||
never reads that variable). **But measured 2026-10-04: two of them point at
|
||||
`~/.ccs/instances/ltms`, which is the operator's own `CLAUDE_CONFIG_DIR` on this host.** So for an
|
||||
`opus` or `sonnet` member, "a member never reads the operator's plugin store" is false — it reads
|
||||
the same store, because that store is the one its `configDir` names. It stays true for `local` and
|
||||
`local-direct`, which point at `~/.ccs/instances/gx10`. `ClaudeCodeLauncher`'s own javadoc names
|
||||
the related hazard: that file is rewritten on every spawn, so for those two profiles fleetd and the
|
||||
operator's live session write the same `.claude.json`, and its compare-and-swap "narrows the
|
||||
lost-update window, it does not close it". Re-measure which profiles share the operator's dir with
|
||||
`awk '/^profiles:/{i=1;next} /^[a-z]/{i=0} i&&/^ [a-z-]+:$/{p=$1} i&&/configDir:/{print p,$2}'
|
||||
fleetd/fleetd.yaml` against `echo $CLAUDE_CONFIG_DIR`; delete this note once no profile names the
|
||||
operator's dir. **Treat the plugin as the lead-side surface and put member-facing assets in the
|
||||
worktree** — that conclusion holds either way, because a worktree asset does not depend on which
|
||||
config dir a member reads.
|
||||
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
|
||||
(a submodule with its own remote).
|
||||
- **A provisioned worktree neutralizes `.mcp.json`, `opencode.json` and `.autoenv`** — the repo's
|
||||
@@ -323,6 +373,13 @@ must obey belongs in the charter, not here.
|
||||
reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if
|
||||
that page names a `fleet_*` tool the server does not register. The flows are kept out of this
|
||||
file because this file loads into every session's context.
|
||||
- **The daemon runs from `fleetd/run/fleetd.jar`, not `fleetd/target/fleetd.jar`** (fleetd #664).
|
||||
Maven's own output still lands at `fleetd/target/fleetd.jar` — that part of the build is
|
||||
unchanged — but the running daemon never has that file open, so a bare `mvn install`/`mvn clean`
|
||||
in the main clone no longer corrupts anything a live process is reading. Verify merges in a
|
||||
throwaway git worktree anyway: a build in the main clone still ships nothing until
|
||||
`scripts/redeploy-fleetd.sh` moves it into place with its own atomic `mv`, performed only after
|
||||
the old daemon is confirmed gone. Let only that script touch `fleetd/run/fleetd.jar`.
|
||||
|
||||
### Redeploying the daemon — the lead may do this (primary only)
|
||||
|
||||
@@ -352,7 +409,7 @@ Before you call any work done, check the row that matches what you touched:
|
||||
| `ConnectionIdentity` / how a caller is resolved | the `fleet_whoami` paragraph and the fallback ladder |
|
||||
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__fleet__*`), and the layering table's top row |
|
||||
| the injector / status gating | invariant 4 |
|
||||
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
|
||||
| worktree provisioning or the parity overlay | the "every role reads this file" premise — it rests on the worker's worktree being a checkout of this repo |
|
||||
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
|
||||
| a new peer kind (non-Claude adapter) | what that peer can read — anything it must obey belongs in its charter, not in the block |
|
||||
| **anything an operator can use, configure, or observe** — an MCP tool, a `fleetd.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
|
||||
|
||||
@@ -47,7 +47,7 @@
|
||||
<string>/Users/dai.ha/LTMS/claude-bridge/scripts/fleetd-launchd-wrapper.sh</string>
|
||||
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin/java</string>
|
||||
<string>-jar</string>
|
||||
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/target/fleetd.jar</string>
|
||||
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/run/fleetd.jar</string>
|
||||
<string>fleetd.yaml</string>
|
||||
</array>
|
||||
|
||||
|
||||
@@ -50,7 +50,7 @@ WorkingDirectory=%h/LTMS/fleetd/fleetd
|
||||
# and looks healthy, and the failure appears hours later as a member that cannot open a pull
|
||||
# request. exec keeps it one process, so systemd tracks the right PID.
|
||||
# This also avoids a SECOND copy of the secrets in a systemd drop-in: one source of truth.
|
||||
ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"
|
||||
ExecStart=/bin/zsh -lc "exec java -jar run/fleetd.jar fleetd.yaml"
|
||||
|
||||
# PrivateTmp MUST stay false -- see herdr.service. fleetd creates the member ZDOTDIR scrub dir and
|
||||
# the opencode config dir under java.io.tmpdir, and the member pane (a herdr child, a different
|
||||
|
||||
@@ -2,6 +2,10 @@
|
||||
target/
|
||||
dependency-reduced-pom.xml
|
||||
|
||||
# The daemon's runtime jar (fleetd #664). scripts/redeploy-fleetd.sh moves the built jar here
|
||||
# with a same-filesystem rename; this is never Maven's output path and never belongs in git.
|
||||
run/
|
||||
|
||||
# Local runtime config (copy from fleetd.example.yaml). Both names are ignored: fleetd.yaml is
|
||||
# the current name, and bridged.yaml is the legacy name Fleetd still falls back to.
|
||||
fleetd.yaml
|
||||
|
||||
+48
-25
@@ -62,11 +62,12 @@ bind:
|
||||
#
|
||||
# Leads are configured under `fleet.leaders:` — see THE FLEET further down.
|
||||
#
|
||||
# Two things stop the tab-name convention from becoming a way to claim leadership: the configured
|
||||
# member spaces are excluded from the scan, so nothing fleetd places can land in a matching tab;
|
||||
# and startup REFUSES a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel`
|
||||
# override, also matches — so the two namespaces cannot overlap by accident. The label is a NAME,
|
||||
# never a capability: what a pane may do is decided by the role the daemon resolves for it.
|
||||
# Two things stop the tab-name convention from becoming a way to claim leadership: startup REFUSES
|
||||
# a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel` override, also
|
||||
# matches, so the two namespaces cannot overlap by accident; and the CallerResolver asks the live
|
||||
# spawned-member roster BEFORE any tab map, so a live member is never mistaken for a lead no matter
|
||||
# what its tab says. The label is a NAME, never a capability: what a pane may do is decided by the
|
||||
# role the daemon resolves for it.
|
||||
|
||||
# CB-551: IDLE-LEAD HEARTBEAT — nudge the single lead back to work when it has been continuously
|
||||
# idle (no open fleet_send driving it) past the quiet period. The fleet is one lead + architects +
|
||||
@@ -98,17 +99,17 @@ bind:
|
||||
# contextHighNudge: false
|
||||
|
||||
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
|
||||
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
|
||||
# own pane and bootstrap a fresh session against that file.
|
||||
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to end its
|
||||
# own pane, launch a fresh one, and bootstrap that fresh session against the handover file.
|
||||
#
|
||||
# Opt-in on purpose — it clears the lead's own pane on request, so upgrading the daemon must never
|
||||
# acquire that ability for you. Absent block = feature off, and nothing is constructed at all. Even
|
||||
# once present, nothing but an explicit confirm() call — one that passes every check — can ever
|
||||
# cause a /clear: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
|
||||
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so
|
||||
# it cannot clear the pane inline (that pane is still WORKING); instead it schedules a one-shot
|
||||
# Opt-in on purpose — it tears down the lead's own pane on request, so upgrading the daemon must
|
||||
# never acquire that ability for you. Absent block = feature off, and nothing is constructed at all.
|
||||
# Even once present, nothing but an explicit confirm() call — one that passes every check — can ever
|
||||
# tear a pane down: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
|
||||
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so it
|
||||
# cannot act on the pane inline (that pane is still WORKING); instead it schedules a one-shot
|
||||
# continuation that waits for the SAME confirm() call's turn to end, then does the actual work. See
|
||||
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order (fleetd #480 correction).
|
||||
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order.
|
||||
#
|
||||
# handoverPath: REQUIRED when this block is present — where the handover file a fresh lead session
|
||||
# reads must live. No default (an operator-specific path); a present block with no
|
||||
@@ -117,25 +118,30 @@ bind:
|
||||
# working directory when that lead has none configured) — never against whatever
|
||||
# directory the daemon process happens to have been started in. An absolute path is
|
||||
# used unchanged. Prefer an absolute path if the daemon and the lead's pane might not
|
||||
# share a working directory (fleetd #480 follow-up).
|
||||
# share a working directory.
|
||||
# requireOperatorConfirm: true # default true — confirm() refuses unless the caller also passes
|
||||
# # operatorConfirmed: true
|
||||
# maxDocAgeSeconds: 3600 # default 3600 — refuse a handover file older than this
|
||||
# turnSettleSeconds: 20 # default 20 — how long the deferred roll waits for the CALLING
|
||||
# # lead's own turn to end (its pane to report injectable again)
|
||||
# # before sending /clear at all. If this elapses, /clear is NEVER
|
||||
# # sent — a lead that never goes idle is still doing real work.
|
||||
# clearSettleSeconds: 20 # default 20 — how long to wait for the pane to become injectable
|
||||
# # again AFTER /clear before giving up (never sends bootstrapText
|
||||
# # if this elapses). A separate, second wait from turnSettleSeconds.
|
||||
# # before tearing the old pane down at all. If this elapses, nothing
|
||||
# # is torn down — a lead that never goes idle is still doing real
|
||||
# # work.
|
||||
# relaunchReadySeconds: 45 # default 45 — bounds two later waits, after the old pane is gone
|
||||
# # and a fresh one has been launched: first, for the fresh pane to
|
||||
# # reach a real turn boundary (never sends bootstrapText if THIS one
|
||||
# # elapses); second, for the new terminal to be recognised as this
|
||||
# # lead (bootstrapText is sent either way once the first wait
|
||||
# # passes). A separate, later pair of waits from turnSettleSeconds.
|
||||
# bootstrapText: "..." # default names the RESOLVED (absolute) handoverPath — sent to
|
||||
# # the lead once its pane settles after /clear
|
||||
# # the freshly relaunched lead's pane once it reaches a real turn
|
||||
# # boundary
|
||||
# leadRollover:
|
||||
# handoverPath: /path/to/handover.md
|
||||
# requireOperatorConfirm: true
|
||||
# maxDocAgeSeconds: 3600
|
||||
# turnSettleSeconds: 20
|
||||
# clearSettleSeconds: 20
|
||||
# relaunchReadySeconds: 45
|
||||
# bootstrapText: "Fresh lead session: read the handover file and carry on."
|
||||
|
||||
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
|
||||
@@ -659,9 +665,10 @@ fleet:
|
||||
# tabPrefix: "lead:" # only used to guard against a worker tabLabel colliding with
|
||||
# # this convention at startup; plays no part in matching a lead
|
||||
# scanIntervalSeconds: 10 # rescan cadence, and the worst case before a new tab is seen
|
||||
# workspace: leads # where a launched lead's tab is created (default "leads").
|
||||
# # MUST NOT be a member workspace — those are excluded from the
|
||||
# # scan, so a lead placed in one is never found again.
|
||||
# workspace: leads # where a launched lead's tab is created (default "fleet",
|
||||
# # the same shared space the members use). Sharing that space
|
||||
# # with members is the normal shipped shape: the scanner tells
|
||||
# # a lead from a member by the exact tab label, not by workspace.
|
||||
# cwd: /path/to/repo # the launched lead's working directory (default: fleetd's own)
|
||||
# kind: claude # descriptive; reported by fleet_whoami
|
||||
# gpt-sol-5.6:
|
||||
@@ -669,6 +676,22 @@ fleet:
|
||||
# kind: opencode
|
||||
# model: openai/gpt-5.6-terra
|
||||
|
||||
# A collaborator tab, keyed by name (fleetd #669). A pane whose tab matches resolves as the
|
||||
# COLLABORATOR role: `fleet_whoami` answers `collaborator`, and the session may observe the fleet
|
||||
# (`fleet_list`, `fleet_profiles`, `fleet_whoami`) and `fleet_send` to a lead or to another
|
||||
# collaborator. It may NOT spawn, stop or drain anything, roll a lead's session, answer a
|
||||
# member's `fleet_ask`, poll a ticket, send across hosts, or send to a spawned member's terminal.
|
||||
# Each of those is refused at the gate rather than queued.
|
||||
# Recognise-only, like a profile-less `leaders:` entry above: there is no `profile:`, no
|
||||
# `instances:` and no `kind:`. `tab:` is REQUIRED and is the only field identity depends on,
|
||||
# matched case-insensitively — the same GET-THE-VALUE-RIGHT warning above the `leaders:` block
|
||||
# applies here too.
|
||||
# A lead cannot discover a collaborator yet (#703): `fleet_list` has no `collaborators` key, so
|
||||
# the collaborator must speak first, or pass on the `sessionId` its own `fleet_whoami` reports.
|
||||
# collaborators:
|
||||
# reviewer-alex:
|
||||
# tab: "collab: alex"
|
||||
|
||||
# architects:
|
||||
# architect-1:
|
||||
# profile: opus # a strong model, on the operator's subscription
|
||||
|
||||
+4
-2
@@ -189,8 +189,10 @@
|
||||
|
||||
<build>
|
||||
<!-- CB-634: the cutover renamed the module dir (bridged/ -> fleetd/), the jar, and the
|
||||
launchd plist together. The installed plist names fleetd/target/fleetd.jar and
|
||||
KeepAlive is armed, so this name, the plist, and the wrapper must move as one. -->
|
||||
launchd plist together. fleetd #664: the installed plist and the systemd unit now name
|
||||
fleetd/run/fleetd.jar, not this plugin's own output path — see
|
||||
scripts/redeploy-fleetd.sh for the mv that gets a build from here to there. KeepAlive is
|
||||
armed, so this name, the plist, and the wrapper must still move as one. -->
|
||||
<finalName>fleetd</finalName>
|
||||
<plugins>
|
||||
<plugin>
|
||||
|
||||
@@ -217,25 +217,27 @@ public final class Fleetd {
|
||||
|
||||
/**
|
||||
* The {@link Injector}'s readiness gate (CB-534): a target is deliverable if it is a spawned
|
||||
* member whose agent has connected the bridge MCP, <em>or</em> a lead.
|
||||
* member whose agent has connected the bridge MCP, a lead, <em>or</em> a collaborator.
|
||||
*
|
||||
* <p>The gate exists for one reason — to hold a delivery out of a <em>spawned</em> member's boot
|
||||
* window, where herdr already reports {@code idle} but the TUI would drop an injected paste. That
|
||||
* hazard is a property of spawning. A lead is never spawned: the operator started it and named it
|
||||
* (or labelled its tab) only once it was up, so there is no boot window to guard.
|
||||
* (or labelled its tab) only once it was up, so there is no boot window to guard. A collaborator
|
||||
* is the same way — a person's own tab, matched to a configured name, never spawned.
|
||||
*
|
||||
* <p>A lead is also never enrolled in {@link MemberPresence} — {@code FleetMcp} marks presence
|
||||
* for every spawned member (worker and architect), deliberately, since that map doubles as the
|
||||
* member roster's availability signal and a lead counted there would show up as an available
|
||||
* member. So without the second disjunct a lead is permanently un-deliverable: every
|
||||
* lead→lead send sat on the gate for {@code READINESS_GRACE_POLLS} (~60s) and then failed
|
||||
* having never been typed into the pane.
|
||||
* <p>Neither a lead nor a collaborator is ever enrolled in {@link MemberPresence} — {@code
|
||||
* FleetMcp} marks presence only for a worker, an architect, or the unconfigured-pane floor,
|
||||
* never for a lead or a collaborator. So without the second and third disjuncts a lead or
|
||||
* collaborator is permanently un-deliverable: every send to one sat on the gate for
|
||||
* {@code READINESS_GRACE_POLLS} (~60s) and then failed having never been typed into the pane.
|
||||
*
|
||||
* <p>The lead set is read through the supplier on each call rather than snapshotted, so a lead
|
||||
* discovered by {@code leadScan} after startup becomes deliverable without a restart.
|
||||
* <p>Both sets are read through their supplier on each call rather than snapshotted, so a lead or
|
||||
* collaborator discovered by {@code leadScan} after startup becomes deliverable without a restart.
|
||||
*/
|
||||
static Predicate<String> deliverableTo(MemberPresence presence, Supplier<Map<String, String>> leads) {
|
||||
return target -> presence.isPresent(target) || leads.get().containsKey(target);
|
||||
static Predicate<String> deliverableTo(MemberPresence presence, Supplier<Map<String, String>> leads,
|
||||
Supplier<Map<String, String>> collaborators) {
|
||||
return target -> presence.isPresent(target) || leads.get().containsKey(target)
|
||||
|| collaborators.get().containsKey(target);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -939,21 +941,27 @@ public final class Fleetd {
|
||||
* cfg.leadHeartbeat()}
|
||||
* @param leadAgents the {@link AgentControl} instance that reaches the LEAD's pane (not
|
||||
* {@code memberAgents}), normally {@code router.leadAgents()}
|
||||
* @param leadSpaces the {@link WorkspaceControl} instance that reaches the LEAD's
|
||||
* workspace, normally {@code router.leadSpaces()} — used to tear down
|
||||
* a rolled lead's old pane and confirm it is gone
|
||||
* @param launcher starts the fresh lead a roll relaunches once the old one is gone
|
||||
* @param config the live {@link ConfigRef}, captured only inside the returned
|
||||
* supplier and the workspace lookup — never dereferenced here
|
||||
* supplier and the two lookups below — never dereferenced here
|
||||
* @param liveLeadTerminals terminal id → lead NAME for every CURRENTLY recognised lead, normally
|
||||
* the same {@code leads} supplier {@code main} already builds for
|
||||
* {@code HerdrRouter}/{@link #leadSeatLookup} — never a value snapshot
|
||||
* @return a constructed {@link LeadRollover}, or {@code null} when {@code leadRollover:} is
|
||||
* absent from the startup config
|
||||
*/
|
||||
static LeadRollover leadRollover(FleetConfig cfg, AgentControl leadAgents, ConfigRef config,
|
||||
static LeadRollover leadRollover(FleetConfig cfg, AgentControl leadAgents,
|
||||
WorkspaceControl leadSpaces, LeadLauncher launcher, ConfigRef config,
|
||||
Supplier<Map<String, String>> liveLeadTerminals) {
|
||||
if (cfg.leadRollover() == null) {
|
||||
return null;
|
||||
}
|
||||
Function<String, String> leadNameForTerminal = terminal -> liveLeadTerminals.get().get(terminal);
|
||||
Function<String, String> leadWorkspace = terminal -> {
|
||||
String leadName = liveLeadTerminals.get().get(terminal);
|
||||
String leadName = leadNameForTerminal.apply(terminal);
|
||||
if (leadName == null) {
|
||||
return null;
|
||||
}
|
||||
@@ -961,7 +969,8 @@ public final class Fleetd {
|
||||
FleetConfig.Leader leader = fleet == null ? null : fleet.leaders().get(leadName);
|
||||
return leader == null ? null : leader.cwd();
|
||||
};
|
||||
return new LeadRollover(leadAgents, () -> config.get().leadRollover(), leadWorkspace);
|
||||
return new LeadRollover(leadAgents, leadSpaces, launcher, () -> config.get().leadRollover(),
|
||||
leadWorkspace, leadNameForTerminal, liveLeadTerminals);
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -43,6 +43,7 @@ import dev.ltms.fleet.msg.MessageService;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import dev.ltms.fleet.msg.ReplyPushLoop;
|
||||
import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
@@ -141,8 +142,12 @@ final class FleetdAssembly {
|
||||
? ports.connectHerdr(Path.of(cfg.memberHerdrSocket()))
|
||||
: herdr;
|
||||
AtomicReference<Supplier<Map<String, String>>> leadsRef = new AtomicReference<>(Map::of);
|
||||
// fleetd #669 Unit E: a collaborator's pane is opened by a person, exactly like a lead's,
|
||||
// so its terminal must also route to the lead herdr daemon rather than the member one.
|
||||
AtomicReference<Supplier<Map<String, String>>> collaboratorTerminalsRef = new AtomicReference<>(Map::of);
|
||||
HerdrRouter router = new HerdrRouter(herdr, memberHerdr,
|
||||
target -> leadsRef.get().get().containsKey(target));
|
||||
target -> leadsRef.get().get().containsKey(target)
|
||||
|| collaboratorTerminalsRef.get().get().containsKey(target));
|
||||
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
|
||||
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
|
||||
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
|
||||
@@ -250,32 +255,57 @@ final class FleetdAssembly {
|
||||
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
|
||||
}
|
||||
// CB-531/CB-579: discover leads by the tab labels the operator writes, one scanner per
|
||||
// configured lead's own exact `tab:` label.
|
||||
// configured lead's own exact `tab:` label. fleetd #669: the same scan also recognises a
|
||||
// configured collaborator's tab, so one herdr pass answers both.
|
||||
final Supplier<Map<String, String>> leads;
|
||||
final Supplier<Map<String, String>> collaboratorTerminals;
|
||||
var leaders = cfg.fleet().leaders();
|
||||
if (!leaders.isEmpty()) {
|
||||
var collaboratorsConfig = cfg.fleet().collaborators();
|
||||
if (!leaders.isEmpty() || !collaboratorsConfig.isEmpty()) {
|
||||
Map<String, String> tabToName = new LinkedHashMap<>();
|
||||
leaders.forEach((name, leader) -> {
|
||||
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
|
||||
tabToName.put(leader.tab(), name);
|
||||
}
|
||||
});
|
||||
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
|
||||
Map<String, String> collaboratorTabToName = new LinkedHashMap<>();
|
||||
collaboratorsConfig.forEach((name, collaborator) -> {
|
||||
if (collaborator != null && collaborator.tab() != null && !collaborator.tab().isBlank()) {
|
||||
collaboratorTabToName.put(collaborator.tab(), name);
|
||||
}
|
||||
});
|
||||
// A collaborator-only fleet configures no `leaders:` entry to read a scan interval from
|
||||
// — FleetConfig.Collaborator carries no scanIntervalSeconds of its own. Falling back to
|
||||
// FleetConfig.Leader's own compact-constructor default keeps a collaborator-only
|
||||
// deployment on the same rescan cadence as the default lead cadence, instead of
|
||||
// inventing a second number for the same kind of scan.
|
||||
int scanIntervalSeconds = leaders.isEmpty()
|
||||
? 10
|
||||
: leaders.values().iterator().next().scanIntervalSeconds();
|
||||
// This must use the lead daemon: scanning member tabs would demote the lead to a worker.
|
||||
leads = new LeadTabScanner(herdr, tabToName, Set.of(),
|
||||
LeadTabScanner scanner = new LeadTabScanner(herdr, tabToName, collaboratorTabToName, Set.of(),
|
||||
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), ports.nanoClock());
|
||||
log.info("lead scan: tabs {} host a lead (rescan every {}s, shared fleet space)",
|
||||
tabToName.keySet(), scanIntervalSeconds);
|
||||
leads = scanner;
|
||||
collaboratorTerminals = scanner::collaborators;
|
||||
log.info("lead/collaborator scan: tabs {} host a lead, tabs {} host a collaborator "
|
||||
+ "(rescan every {}s, shared fleet space)",
|
||||
tabToName.keySet(), collaboratorTabToName.keySet(), scanIntervalSeconds);
|
||||
} else {
|
||||
leads = () -> leadTerminals;
|
||||
collaboratorTerminals = Map::of;
|
||||
}
|
||||
leadsRef.set(leads);
|
||||
collaboratorTerminalsRef.set(collaboratorTerminals);
|
||||
|
||||
// Constructed unconditionally — it is cheap and side-effect free — so a LeadRollover built
|
||||
// below can relaunch a lead even on a boot where herdr was down for the ensureLeads() call.
|
||||
LeadLauncher leadLauncher = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg);
|
||||
|
||||
// CB-558: start any declared lead that is not already running. After the scanner is built,
|
||||
// and only when herdr answered — the launcher's whole safety property is that it can count
|
||||
// live leads first, and must never guess and risk a second orchestrator.
|
||||
if (herdrUp && !leaders.isEmpty()) {
|
||||
int launched = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg).ensureLeads();
|
||||
int launched = leadLauncher.ensureLeads();
|
||||
if (launched > 0) {
|
||||
log.info("lead auto-launch: {} lead(s) started", launched);
|
||||
}
|
||||
@@ -341,7 +371,7 @@ final class FleetdAssembly {
|
||||
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
|
||||
MemberPresence presence = sessions.asPresence();
|
||||
TurnListener turnListener = Fleetd.turnListener(completion, sessions);
|
||||
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads);
|
||||
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads, collaboratorTerminals);
|
||||
// fleetd #556: registration is wired directly to `completion`, not folded into the
|
||||
// `turnListener` fan-out above — so it survives `sessions.onDelivered` (or any future
|
||||
// listener) throwing, regardless of call order.
|
||||
@@ -414,7 +444,8 @@ final class FleetdAssembly {
|
||||
heartbeatScheduler.shutdownNow();
|
||||
}
|
||||
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed.
|
||||
LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), config, leads);
|
||||
LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), router.leadSpaces(),
|
||||
leadLauncher, config, leads);
|
||||
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
|
||||
pushLoop, metrics);
|
||||
|
||||
@@ -452,6 +483,11 @@ final class FleetdAssembly {
|
||||
ConnectionIdentity identity = new ConnectionIdentity(
|
||||
new PaneLocator(herdr, memberHerdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
|
||||
|
||||
// fleetd #669 Unit D: a live spawned member resolves as its own role, whatever a tab map
|
||||
// says about the same terminal. fleetd #702: SessionManager.spawnedMemberRole also answers
|
||||
// for a pane mid-teardown, not only one still in the registry — see its javadoc.
|
||||
Function<String, MemberRole> spawnedMemberRole = sessions::spawnedMemberRole;
|
||||
|
||||
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
|
||||
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
|
||||
final CallerResolver callers;
|
||||
@@ -461,11 +497,13 @@ final class FleetdAssembly {
|
||||
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
|
||||
+ " is unset or empty — export it before starting fleetd");
|
||||
}
|
||||
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
|
||||
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members,
|
||||
spawnedMemberRole, collaboratorTerminals);
|
||||
log.info("auth: token mode (bearer required for non-worker callers, env {})",
|
||||
cfg.auth().tokenEnv());
|
||||
} else {
|
||||
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
|
||||
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members,
|
||||
spawnedMemberRole, collaboratorTerminals);
|
||||
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
|
||||
}
|
||||
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
package dev.ltms.fleet.auth;
|
||||
|
||||
import java.util.function.Predicate;
|
||||
|
||||
/**
|
||||
* The authorization table (CB-505), stated once and enforced on both entry paths.
|
||||
*
|
||||
@@ -20,16 +22,22 @@ public final class Authz {
|
||||
SPAWN,
|
||||
/** Tear a worker peer down. */
|
||||
STOP,
|
||||
/** Deliver a turn to a session (or answer a worker's question). */
|
||||
/** Deliver a turn to a local session, addressed by {@code sessionId}. */
|
||||
SEND,
|
||||
/** Resolve a worker's blocked question and resume its turn, addressed by {@code turnId}. */
|
||||
ANSWER,
|
||||
/** Address a peer lead on another daemon over the coordination broker, by {@code coordId}. */
|
||||
COORD_SEND,
|
||||
/** A worker's terminal reply for its own turn. */
|
||||
REPLY,
|
||||
/** A worker's mid-turn question to the primary. */
|
||||
ASK,
|
||||
/** Collect held replies from a session's inbox. */
|
||||
DRAIN,
|
||||
/** Read-only observation: status, roster, profiles, task polling. */
|
||||
/** Read-only roster, profile, and identity observation: no ticket, task, or turn state. */
|
||||
READ,
|
||||
/** Poll a ticket, or read a session's status. */
|
||||
TASK_READ,
|
||||
/**
|
||||
* Read (never ack) this daemon's own held lead-to-lead coordination mail (fleetd #421).
|
||||
*
|
||||
@@ -54,43 +62,97 @@ public final class Authz {
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether {@code caller} may perform {@code action} against {@code targetSession}.
|
||||
* The fail-closed classifier: answers no for every target, so a collaborator's {@code SEND}
|
||||
* is refused unless a caller supplies a real one. {@code CallerResolver#knownLeadOrCollaborator()}
|
||||
* is the real one, read from the same lead and collaborator maps {@code CallerResolver#resolve}
|
||||
* consults, so a target that classifier calls known is one {@code resolve} would actually
|
||||
* resolve as a lead or collaborator.
|
||||
*/
|
||||
public static final Predicate<String> NO_KNOWN_LEAD_OR_COLLABORATOR = target -> false;
|
||||
|
||||
/**
|
||||
* Convenience form for a caller with no classifier to supply. Fails closed: a collaborator's
|
||||
* {@code SEND} is refused, as if no terminal were a configured lead or collaborator — the
|
||||
* same decision {@link #NO_KNOWN_LEAD_OR_COLLABORATOR} gives explicitly. Every other action's
|
||||
* result is identical to the four-argument form's, since none of them consult the classifier.
|
||||
*
|
||||
* @param targetSession the session id in the request path; only consulted for the worker-scoped
|
||||
* actions ({@code REPLY}, {@code ASK}), ignored otherwise, may be
|
||||
* {@code null}
|
||||
* <p>Its default classifier denies every collaborator, so a caller enforcing authorization
|
||||
* must use the four-argument form instead.
|
||||
*/
|
||||
public static boolean permits(Principal caller, Action action, String targetSession) {
|
||||
return permits(caller, action, targetSession, NO_KNOWN_LEAD_OR_COLLABORATOR);
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether {@code caller} may perform {@code action} against {@code targetSession}.
|
||||
*
|
||||
* @param targetSession the session id in the request path; only consulted for the
|
||||
* worker-scoped actions ({@code REPLY}, {@code ASK}) and for a
|
||||
* collaborator's {@code SEND}, ignored otherwise, may be
|
||||
* {@code null}
|
||||
* @param knownLeadOrCollaborator whether a terminal is a configured lead or collaborator —
|
||||
* consulted only for a collaborator's {@code SEND}, to confine
|
||||
* it to another named peer and never a spawned member's
|
||||
* terminal
|
||||
*/
|
||||
public static boolean permits(Principal caller, Action action, String targetSession,
|
||||
Predicate<String> knownLeadOrCollaborator) {
|
||||
if (caller == null || caller.isAnonymous()) {
|
||||
return false; // authenticated as nothing ⇒ authorized for nothing
|
||||
}
|
||||
return switch (action) {
|
||||
// Fleet lifecycle is the primary's alone — spawn, stop, drain. An architect
|
||||
// deliberately does NOT get these (CB-548), so it cannot tear down or stand up workers
|
||||
// even though it coordinates them; and a worker driving any of these would be a worker
|
||||
// escalating into the orchestrator role.
|
||||
// Fleet lifecycle is the primary's alone — spawn, stop, drain. An architect and a
|
||||
// collaborator deliberately do NOT get these, so neither can tear down or stand up
|
||||
// workers even though one of them coordinates them; and a worker driving any of these
|
||||
// would be a worker escalating into the orchestrator role.
|
||||
case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary();
|
||||
|
||||
// Delivering a turn is open to the primary and the architect: an architect delegates
|
||||
// to workers (that is the role's point) but still has no lifecycle rights. A worker is
|
||||
// excluded — sending would be it escalating.
|
||||
case SEND -> caller.isPrimary() || caller.isArchitect();
|
||||
// Delivering a turn to a local session is open to the primary, the architect, and a
|
||||
// collaborator whose target is itself a configured lead or collaborator: the architect
|
||||
// delegates to workers (that is the role's point); a collaborator may reach only
|
||||
// another named peer, never a spawned member's terminal. A worker is excluded —
|
||||
// sending would be it escalating.
|
||||
case SEND -> caller.isPrimary() || caller.isArchitect()
|
||||
|| (caller.isCollaborator() && knownLeadOrCollaborator.test(targetSession));
|
||||
|
||||
// Resolving a worker's blocked question is part of delegating to it, open to the same
|
||||
// two roles that may stand up that delegation in the first place. Not a collaborator:
|
||||
// resuming another session's turn is lifecycle-adjacent, not peer messaging.
|
||||
case ANSWER -> caller.isPrimary() || caller.isArchitect();
|
||||
|
||||
// Leaves the daemon over the coordination broker rather than addressing a local
|
||||
// session, open to the same two roles as ANSWER. Not a collaborator: it is a
|
||||
// local-tab peer with no cross-host route.
|
||||
case COORD_SEND -> caller.isPrimary() || caller.isArchitect();
|
||||
|
||||
// The load-bearing rule: a caller acts only as the pane it occupies. CB-532 widened who
|
||||
// that can be — a lead answering another lead is replying for its OWN terminal, which
|
||||
// this already permits — while the rule itself is unchanged, and is what stops anyone
|
||||
// forging a reply for a rendezvous someone else is waiting on. An architect's own pane
|
||||
// passes through the same check, so it can answer a funnel that delegated to it. An
|
||||
// unnamed primary (token/loopback, no pane) owns nothing and is still excluded.
|
||||
// forging a reply for a rendezvous someone else is waiting on. An architect's or a
|
||||
// collaborator's own pane passes through the same check, so each can answer a funnel
|
||||
// that delegated to it. An unnamed primary (token/loopback, no pane) owns nothing and
|
||||
// is still excluded.
|
||||
case REPLY, ASK -> caller.ownsSession(targetSession);
|
||||
|
||||
// Observation is open to every authenticated role: a worker legitimately polls its own
|
||||
// status, and the roster carries no secrets.
|
||||
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
|
||||
// READ is roster, profile, and identity observation — fleet_list, fleet_profiles, and
|
||||
// fleet_whoami — and carries no secrets: no ticket reply, no pending question, and no
|
||||
// other session's turn state. Those live under TASK_READ. METRICS is the separate
|
||||
// Prometheus scrape. Both are open to every authenticated role, including a
|
||||
// collaborator and the unconfigured-pane floor.
|
||||
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect()
|
||||
|| caller.isCollaborator() || caller.isObserver();
|
||||
|
||||
// Ticket polling and session status, open to every role READ is open to except a
|
||||
// collaborator or an observer. MessageService compares a ticket's creator to the
|
||||
// caller on every read as well, so dropping this gate would not expose another
|
||||
// session's reply — it would move the refusal later and widen what a caller that
|
||||
// never orchestrates can probe.
|
||||
case TASK_READ -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
|
||||
|
||||
// fleetd #421: reading held lead-to-lead mail is the primary's alone. An architect
|
||||
// holds READ today (CB-548), so "not primary" must mean not-architect here too — this
|
||||
// is coordination between leads, not observation of the roster.
|
||||
// is coordination between leads, not observation of the roster. The same reasoning
|
||||
// excludes a collaborator.
|
||||
case COORD_READ -> caller.isPrimary();
|
||||
};
|
||||
}
|
||||
|
||||
@@ -7,6 +7,7 @@ import java.nio.charset.StandardCharsets;
|
||||
import java.security.MessageDigest;
|
||||
import java.util.Map;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.Predicate;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
@@ -19,6 +20,13 @@ import java.util.function.Supplier;
|
||||
*
|
||||
* <p><strong>Resolution order</strong> — connection identity first, token second, nothing third:
|
||||
* <ol>
|
||||
* <li>A loopback peer PID that maps to a pane this gateway itself spawned ⇒ that member's own
|
||||
* role: {@link Role#WORKER} for a dev, hunter, or reviewer; {@link Role#ARCHITECT} for an
|
||||
* architect, but only while the live slot role still confirms it (fleetd #424 — a slot
|
||||
* revoked from config demotes an already-bound session on its very next request, so the
|
||||
* roster's own role is never granted on its word alone). No tab map is consulted — a live
|
||||
* spawned member's identity comes from the registry that spawned it, never from a label a
|
||||
* pane could also carry.</li>
|
||||
* <li>A loopback peer PID that maps to a pane named by {@code leaders:}, by the legacy
|
||||
* {@code primary.terminal} pin, or by an operator-labelled lead tab (CB-307, CB-530, CB-531)
|
||||
* ⇒ {@link Role#PRIMARY}, carrying that lead's
|
||||
@@ -28,9 +36,11 @@ import java.util.function.Supplier;
|
||||
* so two leads can work as peers rather than one being demoted.</li>
|
||||
* <li>A loopback peer PID that maps to a pane bound to a CB-548 architect slot ⇒
|
||||
* {@link Role#ARCHITECT}, carrying the slot name. Just unforgeable as a worker's, and
|
||||
* resolved from the <em>live</em> terminal→slot binding (never a request argument), before
|
||||
* the generic worker fallback.</li>
|
||||
* <li>A loopback peer PID that maps to any other herdr pane ⇒ {@link Role#WORKER}. This is
|
||||
* resolved from the <em>live</em> terminal→slot binding (never a request argument). This is
|
||||
* the case the previous step does not catch: a binding with no live spawned-member session.</li>
|
||||
* <li>A loopback peer PID that maps to an operator-labelled collaborator tab ⇒
|
||||
* {@link Role#COLLABORATOR}, carrying that collaborator's name.</li>
|
||||
* <li>A loopback peer PID that maps to any other herdr pane ⇒ {@link Role#OBSERVER}. This is
|
||||
* unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured
|
||||
* regardless of auth mode, so enabling auth never breaks the fleet.</li>
|
||||
* <li>Otherwise, under {@code token} mode, a valid bearer token ⇒ {@link Role#PRIMARY}.</li>
|
||||
@@ -64,6 +74,20 @@ public final class CallerResolver {
|
||||
private final Supplier<Map<String, String>> architectTerminals;
|
||||
private final Function<String, MemberRole> memberSlotRoles;
|
||||
private final Function<String, String> memberSlotNames;
|
||||
/**
|
||||
* terminal_id → the role of the live spawned member occupying it, or {@code null} for a
|
||||
* terminal no spawned member occupies. Consulted first, ahead of every tab map: a live
|
||||
* spawned member's identity is its own, whatever a tab map says about the same terminal.
|
||||
* A function rather than the roster itself, so a resolve on the hot path never scans a list —
|
||||
* the lookup strategy is the caller's to choose.
|
||||
*/
|
||||
private final Function<String, MemberRole> spawnedMemberRole;
|
||||
/**
|
||||
* terminal_id → collaborator name; empty when none are configured. A supplier for the same
|
||||
* reason as {@link #leadTerminals}: a collaborator tab recognised after construction (the tab
|
||||
* scan discovering a newly-labelled tab) takes effect without a restart.
|
||||
*/
|
||||
private final Supplier<Map<String, String>> collaboratorTerminals;
|
||||
|
||||
/** Loopback-trust resolver: no token required, historical behaviour. Test-only. */
|
||||
CallerResolver(ConnectionIdentity identity) {
|
||||
@@ -124,17 +148,42 @@ public final class CallerResolver {
|
||||
/**
|
||||
* Live registry form that can confirm a bound slot is an architect slot.
|
||||
*
|
||||
* <p>This is the only public construction path. It keeps terminal bindings and slot roles in
|
||||
* the same {@link MemberRegistry}, so a configured architect can resolve as an architect.
|
||||
* <p>It keeps terminal bindings and slot roles in the same {@link MemberRegistry}, so a
|
||||
* configured architect can resolve as an architect. No spawned-member roster or collaborator
|
||||
* registry is consulted — equivalent to {@link #withLeadsAndMembers(ConnectionIdentity,
|
||||
* boolean, String, Supplier, MemberRegistry, Function, Supplier)} with both absent. Kept for
|
||||
* every caller that has neither to offer, so adding them did not churn every construction site.
|
||||
*/
|
||||
public static CallerResolver withLeadsAndMembers(ConnectionIdentity identity,
|
||||
boolean tokenMode, String token,
|
||||
Supplier<Map<String, String>> leadTerminals,
|
||||
MemberRegistry members) {
|
||||
return withLeadsAndMembers(identity, tokenMode, token, leadTerminals, members, null, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Live registry form that also resolves a live spawned member to its own role, and a
|
||||
* configured collaborator tab to {@link Role#COLLABORATOR}.
|
||||
*
|
||||
* <p>This is the only public construction path that exercises the full resolution order.
|
||||
*
|
||||
* @param spawnedMemberRole terminal_id → the role of the live spawned member occupying
|
||||
* it, or {@code null} for a terminal no spawned member occupies.
|
||||
* {@code null} here means no roster is consulted at all (every
|
||||
* terminal falls through to the tab maps), not that none matches.
|
||||
* @param collaboratorTerminals terminal_id → collaborator name, live like {@code leadTerminals}
|
||||
*/
|
||||
public static CallerResolver withLeadsAndMembers(ConnectionIdentity identity,
|
||||
boolean tokenMode, String token,
|
||||
Supplier<Map<String, String>> leadTerminals,
|
||||
MemberRegistry members,
|
||||
Function<String, MemberRole> spawnedMemberRole,
|
||||
Supplier<Map<String, String>> collaboratorTerminals) {
|
||||
return new CallerResolver(identity, tokenMode, token, leadTerminals,
|
||||
members == null ? null : members::snapshot,
|
||||
members == null ? null : members::roleForSlot,
|
||||
members == null ? null : members::nameForSlot);
|
||||
members == null ? null : members::nameForSlot,
|
||||
spawnedMemberRole, collaboratorTerminals);
|
||||
}
|
||||
|
||||
private static Supplier<Map<String, String>> fixed(Map<String, String> leadTerminals) {
|
||||
@@ -160,6 +209,17 @@ public final class CallerResolver {
|
||||
Supplier<Map<String, String>> architectTerminals,
|
||||
Function<String, MemberRole> memberSlotRoles,
|
||||
Function<String, String> memberSlotNames) {
|
||||
this(identity, tokenMode, token, leadTerminals, architectTerminals, memberSlotRoles,
|
||||
memberSlotNames, null, null);
|
||||
}
|
||||
|
||||
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
|
||||
Supplier<Map<String, String>> leadTerminals,
|
||||
Supplier<Map<String, String>> architectTerminals,
|
||||
Function<String, MemberRole> memberSlotRoles,
|
||||
Function<String, String> memberSlotNames,
|
||||
Function<String, MemberRole> spawnedMemberRole,
|
||||
Supplier<Map<String, String>> collaboratorTerminals) {
|
||||
if (tokenMode && (token == null || token.isBlank())) {
|
||||
throw new IllegalArgumentException(
|
||||
"auth.mode=token requires a non-empty token; check that the env var named by "
|
||||
@@ -172,6 +232,8 @@ public final class CallerResolver {
|
||||
this.architectTerminals = architectTerminals == null ? Map::of : architectTerminals;
|
||||
this.memberSlotRoles = memberSlotRoles == null ? _ -> null : memberSlotRoles;
|
||||
this.memberSlotNames = memberSlotNames == null ? Function.identity() : memberSlotNames;
|
||||
this.spawnedMemberRole = spawnedMemberRole == null ? _ -> null : spawnedMemberRole;
|
||||
this.collaboratorTerminals = collaboratorTerminals == null ? Map::of : collaboratorTerminals;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -188,14 +250,24 @@ public final class CallerResolver {
|
||||
}
|
||||
|
||||
/**
|
||||
* The currently-recognised architect slots, {@code terminal_id → slot name} (CB-548).
|
||||
* The currently-recognised collaborator tabs, {@code terminal_id → name}.
|
||||
*
|
||||
* <p>Read from the same supplier {@link #resolve} consults, so a slot that is <em>listed</em>
|
||||
* here but would not <em>resolve</em> (or the reverse) cannot drift apart. Live for the same
|
||||
* reason as {@link #leads()}.
|
||||
* <p>Read from the same supplier {@link #resolve} consults, for the reason given in
|
||||
* {@link #leads()}. Live for the same reason as {@link #leads()}.
|
||||
*/
|
||||
public Map<String, String> members() {
|
||||
return architectTerminals.get();
|
||||
public Map<String, String> collaborators() {
|
||||
return collaboratorTerminals.get();
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether {@code target} names a terminal this resolver would resolve as a lead or a
|
||||
* collaborator — the classifier a collaborator's {@code SEND} is checked against, read from the
|
||||
* exact maps {@link #resolve} consults so a target that would resolve as a lead or collaborator
|
||||
* is never the one a collaborator is refused to reach, or the reverse.
|
||||
*/
|
||||
public Predicate<String> knownLeadOrCollaborator() {
|
||||
return target -> leadTerminals.get().containsKey(target)
|
||||
|| collaboratorTerminals.get().containsKey(target);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -208,6 +280,25 @@ public final class CallerResolver {
|
||||
public Principal resolve(String remoteAddr, int remotePort, String authorizationHeader) {
|
||||
ConnectionIdentity.Caller c = identity.resolve(remoteAddr, remotePort);
|
||||
if (c.terminal() != null) {
|
||||
MemberRole spawnedRole = spawnedMemberRole.apply(c.terminal());
|
||||
if (spawnedRole != null) {
|
||||
// A live spawned member occupies this pane. Its identity is its own, whatever a tab
|
||||
// map says about the same terminal — checked before every tab map, consulting none
|
||||
// of them, so a tab label can never override a roster entry for the same terminal.
|
||||
if (spawnedRole == MemberRole.ARCHITECT) {
|
||||
// The roster only answers THAT this pane is a live spawned member; config still
|
||||
// decides WHAT that member's slot grants (fleetd #424). A slot revoked after the
|
||||
// bind must still demote this session on its very next request, so the roster's
|
||||
// own ARCHITECT role is confirmed against the live slot role, exactly as the
|
||||
// architect-slot step below confirms a binding with no live member session.
|
||||
String slot = architectTerminals.get().get(c.terminal());
|
||||
if (slot != null && memberSlotRoles.apply(slot) == MemberRole.ARCHITECT) {
|
||||
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
|
||||
}
|
||||
return Principal.worker(c.terminal(), c.pid());
|
||||
}
|
||||
return Principal.worker(c.terminal(), c.pid());
|
||||
}
|
||||
String lead = leadTerminals.get().get(c.terminal());
|
||||
if (lead != null) {
|
||||
// The config names this pane as a lead's own. The pane mapping is exactly as
|
||||
@@ -221,11 +312,18 @@ public final class CallerResolver {
|
||||
// The config/live binding names this pane as an architect slot's own. Same
|
||||
// unforgeable pane mapping; the live binding, never a request argument, decides.
|
||||
// Check the slot role too: this defence in depth prevents a bad lifecycle bind from
|
||||
// escalating a dev, hunter or reviewer into an architect. Checked before
|
||||
// the worker fallback.
|
||||
// escalating a dev, hunter or reviewer into an architect. This is the case the
|
||||
// spawned-member step above does not catch: a binding with no live member session.
|
||||
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
|
||||
}
|
||||
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
|
||||
String collaborator = collaboratorTerminals.get().get(c.terminal());
|
||||
if (collaborator != null) {
|
||||
// An operator-labelled collaborator tab, confirmed live by the same scan that
|
||||
// confirms a lead tab. Checked last among the tab maps so a pane also matching one
|
||||
// of the above keeps that stronger role.
|
||||
return Principal.collaborator(collaborator, c.terminal(), c.pid());
|
||||
}
|
||||
return Principal.observer(c.terminal(), c.pid()); // unforgeable; never token-gated
|
||||
}
|
||||
|
||||
if (tokenMode) {
|
||||
|
||||
@@ -11,7 +11,9 @@ package dev.ltms.fleet.auth;
|
||||
* @param pid the connecting process id, or {@code -1} when not resolvable (audit context)
|
||||
* @param name for a lead resolved from the CB-530 {@code leaders:} registry, which lead it is;
|
||||
* for an architect resolved from the CB-548 {@code architects:} registry, which
|
||||
* slot it occupies; {@code null} for every other caller, including an unnamed primary
|
||||
* slot it occupies; for a collaborator resolved from the {@code collaborators:}
|
||||
* registry, which collaborator it is; {@code null} for every other caller,
|
||||
* including an unnamed primary
|
||||
*/
|
||||
public record Principal(Role role, String terminal, long pid, String name) {
|
||||
|
||||
@@ -73,6 +75,26 @@ public record Principal(Role role, String terminal, long pid, String name) {
|
||||
return new Principal(Role.ARCHITECT, terminal, pid, slotName);
|
||||
}
|
||||
|
||||
/**
|
||||
* A collaborator: a human-opened tab recognised by its exact label in the
|
||||
* {@code collaborators:} registry.
|
||||
*
|
||||
* <p>Carries {@link Role#COLLABORATOR}. {@code name} is reporting only — it lets
|
||||
* {@code fleet_whoami} say which collaborator is asking. Identity is the {@code terminal}:
|
||||
* like a worker's it comes from the connection, so {@code ownsSession} works exactly as it
|
||||
* does for a worker — a collaborator acts as its own pane and no other.
|
||||
*/
|
||||
public static Principal collaborator(String name, String terminal, long pid) {
|
||||
return new Principal(Role.COLLABORATOR, terminal, pid, name);
|
||||
}
|
||||
|
||||
/**
|
||||
* The unconfigured-pane floor: a loopback caller whose pane matched no other role.
|
||||
*/
|
||||
public static Principal observer(String terminal, long pid) {
|
||||
return new Principal(Role.OBSERVER, terminal, pid);
|
||||
}
|
||||
|
||||
public boolean isPrimary() {
|
||||
return role == Role.PRIMARY;
|
||||
}
|
||||
@@ -81,15 +103,23 @@ public record Principal(Role role, String terminal, long pid, String name) {
|
||||
return role == Role.ARCHITECT;
|
||||
}
|
||||
|
||||
public boolean isCollaborator() {
|
||||
return role == Role.COLLABORATOR;
|
||||
}
|
||||
|
||||
public boolean isWorker() {
|
||||
return role == Role.WORKER;
|
||||
}
|
||||
|
||||
public boolean isObserver() {
|
||||
return role == Role.OBSERVER;
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether this caller is a spawned member with its own pane.
|
||||
*
|
||||
* <p>Both workers and architects are spawned members. A lead is excluded because recording it
|
||||
* as present would count it as an available member in the roster.
|
||||
* <p>Both workers and architects are spawned members. A lead is not: it is a peer the
|
||||
* operator started and named, never a pane this daemon spawned.
|
||||
*/
|
||||
public boolean isSpawnedMember() {
|
||||
return role == Role.WORKER || role == Role.ARCHITECT;
|
||||
@@ -120,6 +150,8 @@ public record Principal(Role role, String terminal, long pid, String name) {
|
||||
return switch (role) {
|
||||
case WORKER -> "worker:" + terminal;
|
||||
case ARCHITECT -> "architect:" + name;
|
||||
case COLLABORATOR -> "collaborator:" + name;
|
||||
case OBSERVER -> "observer:" + terminal;
|
||||
case PRIMARY -> name == null ? "primary" : "leader:" + name;
|
||||
case ANONYMOUS -> "anonymous";
|
||||
};
|
||||
|
||||
@@ -34,6 +34,28 @@ public enum Role {
|
||||
*/
|
||||
ARCHITECT,
|
||||
|
||||
/**
|
||||
* A config-declared, human-opened tab recognised by its exact label (the {@code
|
||||
* fleet.collaborators.<name>.tab} registry). Never spawned — identity comes from the
|
||||
* connection, never a request argument, exactly like {@link #WORKER} and {@link #ARCHITECT}.
|
||||
* May {@code SEND} only to a configured lead or collaborator, {@code REPLY}/{@code ASK} only
|
||||
* as its own pane, and {@code READ}/{@code METRICS}; may not {@code SPAWN}/{@code STOP}/
|
||||
* {@code DRAIN}/{@code HANDOVER}, poll a ticket ({@code TASK_READ}), or reach the
|
||||
* coordination broker ({@code COORD_SEND}/{@code COORD_READ}).
|
||||
*/
|
||||
COLLABORATOR,
|
||||
|
||||
/**
|
||||
* A loopback pane that resolved to none of the roles above: not a live spawned member, not a
|
||||
* configured lead, not a bound architect slot, not a configured collaborator tab. Unforgeable
|
||||
* like a worker's — derived from the connection's pane, never from a request argument, and
|
||||
* honoured regardless of auth mode. May {@code READ} and {@code METRICS}, and {@code REPLY}/
|
||||
* {@code ASK} only as its own pane; may not {@code SPAWN}/{@code STOP}/{@code DRAIN}/
|
||||
* {@code HANDOVER}, {@code SEND}, poll a ticket ({@code TASK_READ}), or reach the coordination
|
||||
* broker ({@code COORD_SEND}/{@code COORD_READ}).
|
||||
*/
|
||||
OBSERVER,
|
||||
|
||||
/** Authenticated as nothing. Authorized for nothing but {@code /healthz}. */
|
||||
ANONYMOUS
|
||||
}
|
||||
|
||||
@@ -67,7 +67,7 @@ import java.util.function.Supplier;
|
||||
* {@code models:} above: {@code dev.ltms.fleet.lead.LeadRollover} holds a
|
||||
* {@code Supplier<FleetConfig.LeadRollover>} (the same {@code () -> config.get().x()} shape)
|
||||
* and reads {@code handoverPath}/{@code requireOperatorConfirm}/{@code maxDocAgeSeconds}/
|
||||
* {@code turnSettleSeconds}/{@code clearSettleSeconds}/{@code bootstrapText} fresh on every
|
||||
* {@code turnSettleSeconds}/{@code relaunchReadySeconds}/{@code bootstrapText} fresh on every
|
||||
* {@code open()}/{@code confirm()} call (and on the deferred post-{@code confirm()}
|
||||
* continuation fleetd #480's correction added — see {@code LeadRollover}'s class doc) rather
|
||||
* than capturing them into fields at construction — unlike its closest
|
||||
@@ -615,7 +615,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
// may bind to, AND what a slot already bound still grants) through its own instance of that
|
||||
// same supplier shape — see MemberRegistry.live and its class doc for the binding rule:
|
||||
// removing a slot revokes ARCHITECT on the bound pane's very next request, and only the slot
|
||||
// OCCUPANCY survives, so the demoted session keeps its slot key until it unbinds. Only
|
||||
// OCCUPANCY survives, so the demoted session keeps its slot key until it unbinds.
|
||||
// fleet.leaders is frozen (Fleetd.java:281 reads cfg.fleet().leaders() off the startup
|
||||
// snapshot to build both the LeadTabScanner's tab-label-to-name map, wired into
|
||||
// CallerResolver.withLeadsAndMembers at Fleetd.java:620/624, and — when herdr answered —
|
||||
@@ -635,6 +635,11 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
+ "live through that same supplier for placement AND through a separate supplier "
|
||||
+ "on MemberRegistry for spawn-time identity — both already applied");
|
||||
}
|
||||
if (!Objects.equals(collaboratorsOf(old), collaboratorsOf(fresh))) {
|
||||
changed.add("fleet: fleet.collaborators (each collaborator's tab) is read once at "
|
||||
+ "startup to build the LeadTabScanner's identity map, which is not rebuilt on "
|
||||
+ "reload, so a collaborator added, removed, or given a new tab: needs a restart");
|
||||
}
|
||||
// Kept in step with SPLIT_KEYS the same way changedColdKeys is kept in step with COLD_KEYS —
|
||||
// every message here must be traceable to one of the split keys the class doc documents.
|
||||
// NOTE what this does NOT prove, per the javadoc above: it does not catch a SPLIT_KEYS
|
||||
@@ -650,6 +655,11 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
return cfg.fleet() == null ? Map.of() : cfg.fleet().leaders();
|
||||
}
|
||||
|
||||
/** {@code cfg.fleet().collaborators()}, defensively, in case a caller hands in a non-defaulted config. */
|
||||
private static Map<String, FleetConfig.Collaborator> collaboratorsOf(FleetConfig cfg) {
|
||||
return cfg.fleet() == null ? Map.of() : cfg.fleet().collaborators();
|
||||
}
|
||||
|
||||
/**
|
||||
* {@link FleetConfig.Profile} record components deliberately left out of
|
||||
* {@link #sameLaunchSettings} because they are read <em>live</em>, not baked in at spawn — see
|
||||
|
||||
@@ -1129,11 +1129,8 @@ public record FleetConfig(
|
||||
* {@code tab} can never be discovered, launched or not
|
||||
* @param instances how many of this lead should be live (default 1). The daemon
|
||||
* launches only the shortfall, so a restart adopts rather than doubles
|
||||
* @param tabPrefix no longer used to find a lead's tab — {@code tab} is matched
|
||||
* exactly. Its only remaining job is the startup collision guard
|
||||
* ({@link #validateLeadTabPrefixes()}), which still uses it to refuse
|
||||
* a worker {@code tabLabel} template that could be misread as a lead.
|
||||
* Default {@code "lead:"}
|
||||
* @param tabPrefix lead-tab naming convention checked against member labels. Lead
|
||||
* identity uses {@code tab}. Default {@code "lead:"}
|
||||
* @param scanIntervalSeconds how long a tab scan is cached before herdr is asked again; also the
|
||||
* worst case before a newly-labelled tab is recognised. Default 10
|
||||
* @param kind which agent runs there ({@code claude}, {@code opencode}, …)
|
||||
@@ -1179,6 +1176,25 @@ public record FleetConfig(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* A tab fleetd recognises as a collaborator, keyed by name (fleetd #669).
|
||||
*
|
||||
* <p>Recognise-only: there is no {@code profile}, no {@code instances} and no {@code kind}.
|
||||
* Nothing here ever launches a pane.
|
||||
*
|
||||
* <p>{@code tabPrefix} is absent. Identity is matched on the exact {@code tab} alone.
|
||||
*
|
||||
* @param tab the exact tab label hosting this collaborator, matched case-insensitively; the
|
||||
* only field identity depends on. Required — an entry with no {@code tab} can
|
||||
* never be discovered.
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record Collaborator(String tab) {
|
||||
public Collaborator {
|
||||
tab = (tab == null || tab.isBlank()) ? null : tab.strip();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* One entry of a {@code fleet:} role pool — a role paired with the backend it runs on.
|
||||
*
|
||||
@@ -1216,15 +1232,18 @@ public record FleetConfig(
|
||||
* is exactly compatible with that. The pool is also what replaced {@code defaultProfile:} — an
|
||||
* unqualified spawn names a role, and the role's pool supplies the candidates.
|
||||
*
|
||||
* @param leaders panes that orchestrate rather than are orchestrated, keyed by lead name
|
||||
* @param architects profiles the {@code architect} role may run on
|
||||
* @param developers profiles the {@code dev} role may run on
|
||||
* @param hunters profiles the {@code hunter} role may run on
|
||||
* @param reviewers profiles the {@code reviewer} role may run on
|
||||
* @param charters optional launch-charter text keyed by singular role wire name
|
||||
* @param tabLabel template for a member tab's label; {@code {role}}, {@code {profile}},
|
||||
* {@code {model}} and {@code {n}} (a per role+profile counter) are
|
||||
* substituted. Default {@link #DEFAULT_TAB_LABEL}
|
||||
* @param leaders panes that orchestrate rather than are orchestrated, keyed by lead name
|
||||
* @param architects profiles the {@code architect} role may run on
|
||||
* @param developers profiles the {@code dev} role may run on
|
||||
* @param hunters profiles the {@code hunter} role may run on
|
||||
* @param reviewers profiles the {@code reviewer} role may run on
|
||||
* @param charters optional launch-charter text keyed by singular role wire name
|
||||
* @param tabLabel template for a member tab's label; {@code {role}}, {@code {profile}},
|
||||
* {@code {model}} and {@code {n}} (a per role+profile counter) are
|
||||
* substituted. Default {@link #DEFAULT_TAB_LABEL}
|
||||
* @param collaborators tabs fleetd recognises as collaborators (fleetd #669), keyed by name.
|
||||
* Recognise-only, exactly like a {@code profile}-less {@link Leader}:
|
||||
* nothing here is ever auto-launched.
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record Fleet(Map<String, Leader> leaders,
|
||||
@@ -1233,13 +1252,11 @@ public record FleetConfig(
|
||||
Map<String, Slot> hunters,
|
||||
Map<String, Slot> reviewers,
|
||||
Map<String, String> charters,
|
||||
String tabLabel) {
|
||||
String tabLabel,
|
||||
Map<String, Collaborator> collaborators) {
|
||||
|
||||
/**
|
||||
* Role first, so the tab bar reads as the fleet and so the label shares a namespace with a
|
||||
* lead's {@code tabPrefix}. Because {@code {role}} comes from a closed enum, a generated
|
||||
* member label can never begin with {@code "lead:"} — the clash that
|
||||
* {@link #validateLeadTabPrefixes()} used to have to check for is unrepresentable here.
|
||||
* Role first, so the tab bar identifies the member's fleet role.
|
||||
*/
|
||||
public static final String DEFAULT_TAB_LABEL = "{role}: {profile} #{n}";
|
||||
|
||||
@@ -1251,26 +1268,30 @@ public record FleetConfig(
|
||||
reviewers = unmodifiableOrEmpty(reviewers);
|
||||
charters = unmodifiableOrEmpty(charters);
|
||||
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? DEFAULT_TAB_LABEL : tabLabel;
|
||||
collaborators = unmodifiableOrEmpty(collaborators);
|
||||
}
|
||||
|
||||
/**
|
||||
* A fleet with no configured launch charters — the shape every deployment had before
|
||||
* CB-566, and what most tests want.
|
||||
* A fleet with no configured launch charters and no collaborators — the shape every
|
||||
* deployment had before CB-566, and what most tests want.
|
||||
*
|
||||
* <p>Kept deliberately, even though an overload that drops a new field is normally the
|
||||
* shape to avoid. It is safe here because nothing <em>reads</em> a charter through a
|
||||
* constructor: the launcher reads {@code fleet.charters()} from the live config. Jackson
|
||||
* binds the canonical constructor, so this one cannot swallow an operator's YAML.
|
||||
* {@code collaborators} is dropped the same way and for the same reason: no caller of
|
||||
* this overload has ever needed to set it, so it defaults to empty here exactly as the
|
||||
* canonical constructor would default an absent YAML key.
|
||||
*/
|
||||
public Fleet(Map<String, Leader> leaders, Map<String, Slot> architects,
|
||||
Map<String, Slot> developers, Map<String, Slot> reviewers,
|
||||
Map<String, String> charters, String tabLabel) {
|
||||
this(leaders, architects, developers, null, reviewers, charters, tabLabel);
|
||||
this(leaders, architects, developers, null, reviewers, charters, tabLabel, null);
|
||||
}
|
||||
|
||||
public Fleet(Map<String, Leader> leaders, Map<String, Slot> architects,
|
||||
Map<String, Slot> developers, Map<String, Slot> reviewers, String tabLabel) {
|
||||
this(leaders, architects, developers, null, reviewers, null, tabLabel);
|
||||
this(leaders, architects, developers, null, reviewers, null, tabLabel, null);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1395,15 +1416,16 @@ public record FleetConfig(
|
||||
* before anything exists to call — the same fact already true of adding a brand-new
|
||||
* {@code profiles:} entry.
|
||||
*
|
||||
* <p><strong>{@code turnSettleSeconds} (fleetd #480 correction):</strong> {@code confirm()} is
|
||||
* called FROM the calling lead's own turn, so its pane is still {@code WORKING} the instant
|
||||
* {@code confirm()} validates every gate and schedules the roll. {@code
|
||||
* dev.ltms.fleet.lead.LeadRollover}'s deferred continuation waits up to this many seconds for
|
||||
* that SAME pane to report an injectable state again — i.e. for the calling turn to actually
|
||||
* end — before it sends {@code /clear} at all. If that wait times out, no {@code /clear} is
|
||||
* ever sent: a lead that never goes idle is still doing real work, and clearing it would
|
||||
* destroy live context. This is a separate wait from {@code clearSettleSeconds} below, which
|
||||
* bounds the SECOND wait, for the pane to re-settle AFTER {@code /clear} has already gone out.
|
||||
* <p><strong>{@code turnSettleSeconds}:</strong> {@code confirm()} is called FROM the calling
|
||||
* lead's own turn, so its pane is still {@code WORKING} the instant {@code confirm()} validates
|
||||
* every gate and schedules the roll. {@code dev.ltms.fleet.lead.LeadRollover}'s deferred
|
||||
* continuation waits up to this many seconds for that SAME pane to report {@code IDLE} or
|
||||
* {@code DONE} — i.e. for the calling turn to actually end — before it ends the old pane's
|
||||
* process at all. {@code BLOCKED} does not count: that is a live turn merely paused, not one
|
||||
* that has finished. If that wait times out, the old pane is never touched: a lead that never
|
||||
* goes idle is still doing real work, and the roll ends that pane's whole process — there is no
|
||||
* way back from this once it runs, so this wait is the only thing standing between "still
|
||||
* working" and "gone".
|
||||
*
|
||||
* @param handoverPath required when this block is present — where the handover file a fresh
|
||||
* lead session reads must live. There is no sane non-null default for an
|
||||
@@ -1422,37 +1444,51 @@ public record FleetConfig(
|
||||
* @param maxDocAgeSeconds default 3600 — refuse a handover file whose modified time is older
|
||||
* than this many seconds, so a stale leftover from an earlier rollover
|
||||
* attempt can never be mistaken for a fresh one.
|
||||
* @param turnSettleSeconds default 20 — bound on how long the deferred roll waits for the
|
||||
* CALLING lead's own turn to end (its pane to report injectable again)
|
||||
* before sending {@code /clear} at all. See the paragraph above.
|
||||
* @param clearSettleSeconds default 20 — bound on how long to wait for the lead's pane to
|
||||
* report an injectable state again after {@code /clear} before giving up. A
|
||||
* roll that times out here never sends {@code bootstrapText}.
|
||||
* @param turnSettleSeconds default 300 — bound on how long the deferred roll waits for the
|
||||
* CALLING lead's own turn to end (its pane to report {@code IDLE} or
|
||||
* {@code DONE}) before ending that pane's process at all. See the
|
||||
* paragraph above.
|
||||
* @param relaunchReadySeconds default 45 — bound on EACH of two separate waits that run after
|
||||
* the old lead's pane has been torn down and a fresh one launched: first,
|
||||
* for the fresh pane itself to reach a real turn boundary ({@code IDLE} or
|
||||
* {@code DONE}, never merely {@code BLOCKED}) — the safety gate, since
|
||||
* typing into a pane that has not finished booting loses the keystrokes;
|
||||
* second, for the fresh terminal to show up as a recognised lead, which is
|
||||
* bookkeeping rather than a safety gate, so a timeout on this second wait
|
||||
* does not withhold {@code bootstrapText} — it is sent once the pane is
|
||||
* ready regardless. Recognition comes from the same periodically-refreshed
|
||||
* scan {@code LeadTabScanner} already keeps ({@code scanIntervalSeconds},
|
||||
* 10s live), so a budget has to clear more than one scan interval to leave
|
||||
* any real margin for the CLI's own boot time; 20 was rejected for exactly
|
||||
* that reason — at a 10s scan interval it only buys two scans. 45 buys
|
||||
* roughly four. Only a timeout on the FIRST wait (the pane never becomes
|
||||
* ready) withholds {@code bootstrapText}.
|
||||
* @param bootstrapText default a sentence naming the RESOLVED handover path — sent to the
|
||||
* lead's pane once it settles after {@code /clear}, telling the fresh
|
||||
* session where to read the handover and carry on. Left {@code null} here
|
||||
* when the operator configures none: the default sentence cannot be built
|
||||
* at construction time because it must name the path AFTER {@code
|
||||
* dev.ltms.fleet.lead.LeadRollover#open} has resolved a relative {@code
|
||||
* handoverPath} against the calling lead's workspace, which this record has
|
||||
* no way to know — see {@link #bootstrapTextFor(String)}.
|
||||
* fresh lead's pane once it reaches a real turn boundary after relaunch,
|
||||
* telling the fresh session where to read the handover and carry on. Left
|
||||
* {@code null} here when the operator configures none: the default sentence
|
||||
* cannot be built at construction time because it must name the path AFTER
|
||||
* {@code dev.ltms.fleet.lead.LeadRollover#open} has resolved a relative
|
||||
* {@code handoverPath} against the calling lead's workspace, which this
|
||||
* record has no way to know — see {@link #bootstrapTextFor(String)}.
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record LeadRollover(String handoverPath, Boolean requireOperatorConfirm,
|
||||
Integer maxDocAgeSeconds, Integer turnSettleSeconds,
|
||||
Integer clearSettleSeconds, String bootstrapText) {
|
||||
Integer relaunchReadySeconds, String bootstrapText) {
|
||||
public LeadRollover {
|
||||
requireOperatorConfirm = requireOperatorConfirm == null || requireOperatorConfirm;
|
||||
maxDocAgeSeconds = (maxDocAgeSeconds == null || maxDocAgeSeconds <= 0) ? 3600 : maxDocAgeSeconds;
|
||||
turnSettleSeconds = (turnSettleSeconds == null || turnSettleSeconds <= 0) ? 20 : turnSettleSeconds;
|
||||
clearSettleSeconds = (clearSettleSeconds == null || clearSettleSeconds <= 0) ? 20 : clearSettleSeconds;
|
||||
turnSettleSeconds = (turnSettleSeconds == null || turnSettleSeconds <= 0) ? 300 : turnSettleSeconds;
|
||||
relaunchReadySeconds = (relaunchReadySeconds == null || relaunchReadySeconds <= 0)
|
||||
? 45 : relaunchReadySeconds;
|
||||
bootstrapText = (bootstrapText == null || bootstrapText.isBlank()) ? null : bootstrapText;
|
||||
}
|
||||
|
||||
/**
|
||||
* The text actually sent to the lead's pane once it settles after {@code /clear}: the
|
||||
* operator's configured {@link #bootstrapText} when one is set, otherwise the default
|
||||
* sentence built from {@code resolvedHandoverPath}.
|
||||
* The text actually sent to the fresh lead's pane once it reaches a real turn boundary
|
||||
* after relaunch: the operator's configured {@link #bootstrapText} when one is set,
|
||||
* otherwise the default sentence built from {@code resolvedHandoverPath}.
|
||||
*
|
||||
* @param resolvedHandoverPath the ABSOLUTE path {@code dev.ltms.fleet.lead.LeadRollover
|
||||
* #open} already resolved — never the raw configured {@link
|
||||
@@ -1895,6 +1931,7 @@ public record FleetConfig(
|
||||
rejectNegativeMaxLoad(yaml);
|
||||
rejectAutoCompactWindowOutOfRange(yaml);
|
||||
warnConflictingAutoCompactWindows(yaml);
|
||||
warnRetiredClearSettleSecondsKey(yaml);
|
||||
rejectMalformedProfilePatterns(yaml);
|
||||
rejectUnknownKind(yaml);
|
||||
rejectUnknownAuthMode(yaml);
|
||||
@@ -1913,7 +1950,7 @@ public record FleetConfig(
|
||||
|
||||
/** The {@code fleet:} child blocks whose direct children are slot names. */
|
||||
private static final Set<String> FLEET_POOL_KEYS =
|
||||
Set.of("leaders", "architects", "developers", "hunters", "reviewers");
|
||||
Set.of("leaders", "architects", "developers", "hunters", "reviewers", "collaborators");
|
||||
|
||||
/**
|
||||
* Reject a {@code fleet:} role pool whose slot names repeat (CB-548, re-homed by CB-557).
|
||||
@@ -1923,7 +1960,7 @@ public record FleetConfig(
|
||||
* daemon would never know. Jackson's YAML parser does not fail on duplicate mapping keys by
|
||||
* default, so duplicates are caught here, at parse time, before the map is built.
|
||||
*
|
||||
* <p>Only the five pools <em>directly under the top-level {@code fleet:}</em> are considered,
|
||||
* <p>Only the six pools <em>directly under the top-level {@code fleet:}</em> are considered,
|
||||
* and only their direct child keys (the slot names). A nested field elsewhere, even one also
|
||||
* named {@code developers:}, is ignored, so parsing of the rest of the config is unaffected.
|
||||
*
|
||||
@@ -2318,6 +2355,33 @@ public record FleetConfig(
|
||||
names, String.join(", ", detail));
|
||||
}
|
||||
|
||||
/**
|
||||
* Warn when a {@code leadRollover:} block still sets the retired {@code clearSettleSeconds}
|
||||
* key. {@link LeadRollover} carries {@code @JsonIgnoreProperties(ignoreUnknown = true)} and no
|
||||
* longer declares that component, so Jackson drops it with no signal of its own — this raw-YAML
|
||||
* check is the only place an operator's now-inert setting is reported at all; by the time a
|
||||
* {@link LeadRollover} instance exists to run a validator against, the key is already gone.
|
||||
*
|
||||
* @param yaml the raw config text
|
||||
*/
|
||||
static void warnRetiredClearSettleSecondsKey(String yaml) {
|
||||
Map<?, ?> raw;
|
||||
try {
|
||||
raw = YAML.readValue(yaml, Map.class);
|
||||
} catch (IOException | IllegalArgumentException e) {
|
||||
return;
|
||||
}
|
||||
if (raw == null || !(raw.get("leadRollover") instanceof Map<?, ?> leadRollover)) {
|
||||
return;
|
||||
}
|
||||
if (leadRollover.containsKey("clearSettleSeconds")) {
|
||||
log.warn("leadRollover.clearSettleSeconds is retired and no longer read. Set "
|
||||
+ "leadRollover.relaunchReadySeconds instead: it bounds how long to wait, after "
|
||||
+ "a lead is relaunched, for its pane to become ready and then for it to be "
|
||||
+ "recognised as a lead. Remove clearSettleSeconds from fleetd.yaml.");
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Reject a profile whose {@code errorPattern} (fleetd #201 Unit 5) or {@code exhaustedPattern}
|
||||
* (CB-578 stage A) is not a valid Java regex, naming the profile, the key, and the parser's own
|
||||
@@ -2682,30 +2746,21 @@ public record FleetConfig(
|
||||
}
|
||||
|
||||
/**
|
||||
* Reject a lead-scan convention that a worker tab would also satisfy (CB-531).
|
||||
* Reject a member tab-label template that could render as a configured lead or collaborator
|
||||
* tab or match a lead-tab naming convention, and reject two {@code fleet.leaders} or
|
||||
* {@code fleet.collaborators} entries — across either registry — that share one exact tab.
|
||||
*
|
||||
* <p>The scan reads a tab label and concludes "a lead lives here". fleetd also <em>writes</em>
|
||||
* tab labels — every member gets one rendered into its tab. Choose a lead {@code tabPrefix} that
|
||||
* a member template matches and the daemon starts labelling its own members as leads, promoting
|
||||
* the entire fleet to {@link dev.ltms.fleet.auth.Role#PRIMARY} with no message and no diff.
|
||||
* The member-space exclusion in {@link dev.ltms.fleet.herdr.LeadTabScanner} already blocks the
|
||||
* realistic path, but defence that depends on one workspace label holding is not defence enough
|
||||
* for a privilege boundary.
|
||||
* <p>{@code fleet.collaborators} has no {@code tabPrefix}: identity is matched on the exact
|
||||
* {@code tab} alone, so only the exact-render check applies there, not the prefix check.
|
||||
*
|
||||
* <p>CB-557 shrank this check rather than removing it. The default template is
|
||||
* {@code "{role}: {profile} #{n}"} and {@code {role}} comes from a closed enum, so a
|
||||
* <em>generated</em> label can no longer collide by construction. What remains checkable is what
|
||||
* an operator still writes by hand: the {@code fleet.tabLabel} template and any per-profile
|
||||
* {@code tabLabel} override.
|
||||
*
|
||||
* <p>Fatal rather than a warning, unlike {@link #warnUnknownTopLevelKeys}: an unknown key means
|
||||
* a feature does nothing, while this means a feature does the opposite of what it says.
|
||||
*
|
||||
* @throws IllegalStateException when the fleet template or any profile's {@code tabLabel}
|
||||
* override starts with a configured lead prefix
|
||||
* @throws IllegalStateException when the fleet template or a profile {@code tabLabel} override
|
||||
* can render as a configured lead or collaborator tab or match a
|
||||
* lead-tab prefix, or when two entries — of either registry, or
|
||||
* one of each — carry the same exact {@code tab}
|
||||
* (case-insensitively)
|
||||
*/
|
||||
public void validateLeadTabPrefixes() {
|
||||
if (fleet == null || fleet.leaders().isEmpty()) {
|
||||
if (fleet == null) {
|
||||
return;
|
||||
}
|
||||
List<String> bad = new ArrayList<>();
|
||||
@@ -2713,28 +2768,193 @@ public record FleetConfig(
|
||||
if (leader == null) {
|
||||
return;
|
||||
}
|
||||
String tab = leader.tab();
|
||||
String prefix = leader.tabPrefix();
|
||||
// The fleet-wide template is checked once per prefix: it labels every member that has no
|
||||
// override, so one bad template promotes the entire fleet, not one profile.
|
||||
if (startsWithIgnoreCase(fleet.tabLabel(), prefix)) {
|
||||
if (templateCanRenderAs(fleet.tabLabel(), tab)) {
|
||||
bad.add("fleet.tabLabel=\"" + fleet.tabLabel() + "\" can render as the tab of "
|
||||
+ "lead '" + leadName + "' (\"" + tab + "\")");
|
||||
} else if (startsWithIgnoreCase(fleet.tabLabel(), prefix)) {
|
||||
bad.add("fleet.tabLabel=\"" + fleet.tabLabel() + "\" starts with the tabPrefix of "
|
||||
+ "lead '" + leadName + "' (\"" + prefix + "\")");
|
||||
}
|
||||
profiles().entrySet().stream()
|
||||
.filter(e -> startsWithIgnoreCase(e.getValue().tabLabel(), prefix))
|
||||
.map(Map.Entry::getKey)
|
||||
.sorted()
|
||||
.forEach(p -> bad.add("profile '" + p + "' overrides tabLabel with \""
|
||||
+ profiles().get(p).tabLabel() + "\", which starts with the tabPrefix of "
|
||||
+ "lead '" + leadName + "' (\"" + prefix + "\")"));
|
||||
.forEach(p -> {
|
||||
String label = profiles().get(p).tabLabel();
|
||||
if (templateCanRenderAs(label, tab)) {
|
||||
bad.add("profile '" + p + "' overrides tabLabel with \"" + label
|
||||
+ "\", which can render as the tab of lead '" + leadName
|
||||
+ "' (\"" + tab + "\")");
|
||||
} else if (startsWithIgnoreCase(label, prefix)) {
|
||||
bad.add("profile '" + p + "' overrides tabLabel with \"" + label
|
||||
+ "\", which starts with the tabPrefix of lead '" + leadName
|
||||
+ "' (\"" + prefix + "\")");
|
||||
}
|
||||
});
|
||||
});
|
||||
fleet.collaborators().forEach((collabName, collaborator) -> {
|
||||
if (collaborator == null) {
|
||||
return;
|
||||
}
|
||||
String tab = collaborator.tab();
|
||||
if (templateCanRenderAs(fleet.tabLabel(), tab)) {
|
||||
bad.add("fleet.tabLabel=\"" + fleet.tabLabel() + "\" can render as the tab of "
|
||||
+ "collaborator '" + collabName + "' (\"" + tab + "\")");
|
||||
}
|
||||
profiles().entrySet().stream()
|
||||
.map(Map.Entry::getKey)
|
||||
.sorted()
|
||||
.forEach(p -> {
|
||||
String label = profiles().get(p).tabLabel();
|
||||
if (templateCanRenderAs(label, tab)) {
|
||||
bad.add("profile '" + p + "' overrides tabLabel with \"" + label
|
||||
+ "\", which can render as the tab of collaborator '"
|
||||
+ collabName + "' (\"" + tab + "\")");
|
||||
}
|
||||
});
|
||||
});
|
||||
if (!bad.isEmpty()) {
|
||||
throw new IllegalStateException("refusing to start: " + String.join("; ", bad)
|
||||
+ ". A member labelled that way, while its pane carries no entry in the "
|
||||
+ "spawned-member roster, is read back as a lead or collaborator and granted "
|
||||
+ "that identity's authority. Change one of the two so member tabs cannot be "
|
||||
+ "confused with a lead's or collaborator's tab.");
|
||||
}
|
||||
|
||||
List<String> collisions = new ArrayList<>();
|
||||
List<String> leadNames = fleet.leaders().keySet().stream().sorted().toList();
|
||||
for (int i = 0; i < leadNames.size(); i++) {
|
||||
String nameA = leadNames.get(i);
|
||||
Leader a = fleet.leaders().get(nameA);
|
||||
if (a == null || a.tab() == null || a.tab().isBlank()) {
|
||||
continue;
|
||||
}
|
||||
for (int j = i + 1; j < leadNames.size(); j++) {
|
||||
String nameB = leadNames.get(j);
|
||||
Leader b = fleet.leaders().get(nameB);
|
||||
if (b == null || b.tab() == null || b.tab().isBlank()) {
|
||||
continue;
|
||||
}
|
||||
if (a.tab().equalsIgnoreCase(b.tab())) {
|
||||
collisions.add("lead '" + nameA + "' and lead '" + nameB + "' both use tab \""
|
||||
+ a.tab() + "\"");
|
||||
}
|
||||
}
|
||||
}
|
||||
List<String> collabNames = fleet.collaborators().keySet().stream().sorted().toList();
|
||||
for (int i = 0; i < collabNames.size(); i++) {
|
||||
String nameA = collabNames.get(i);
|
||||
Collaborator a = fleet.collaborators().get(nameA);
|
||||
if (a == null || a.tab() == null || a.tab().isBlank()) {
|
||||
continue;
|
||||
}
|
||||
for (int j = i + 1; j < collabNames.size(); j++) {
|
||||
String nameB = collabNames.get(j);
|
||||
Collaborator b = fleet.collaborators().get(nameB);
|
||||
if (b == null || b.tab() == null || b.tab().isBlank()) {
|
||||
continue;
|
||||
}
|
||||
if (a.tab().equalsIgnoreCase(b.tab())) {
|
||||
collisions.add("collaborator '" + nameA + "' and collaborator '" + nameB
|
||||
+ "' both use tab \"" + a.tab() + "\"");
|
||||
}
|
||||
}
|
||||
}
|
||||
for (String leadName : leadNames) {
|
||||
Leader lead = fleet.leaders().get(leadName);
|
||||
if (lead == null || lead.tab() == null || lead.tab().isBlank()) {
|
||||
continue;
|
||||
}
|
||||
for (String collabName : collabNames) {
|
||||
Collaborator collaborator = fleet.collaborators().get(collabName);
|
||||
if (collaborator == null || collaborator.tab() == null
|
||||
|| collaborator.tab().isBlank()) {
|
||||
continue;
|
||||
}
|
||||
if (lead.tab().equalsIgnoreCase(collaborator.tab())) {
|
||||
collisions.add("lead '" + leadName + "' and collaborator '" + collabName
|
||||
+ "' both use tab \"" + lead.tab() + "\"");
|
||||
}
|
||||
}
|
||||
}
|
||||
if (collisions.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
throw new IllegalStateException("refusing to start: " + String.join("; ", collisions)
|
||||
+ ". Tab identity is matched exactly, so only one of two entries sharing a tab can "
|
||||
+ "ever be found — the other is silently unreachable. Give each lead and "
|
||||
+ "collaborator its own exact tab.");
|
||||
}
|
||||
|
||||
private static boolean templateCanRenderAs(String template, String tab) {
|
||||
if (template == null || template.isBlank() || tab == null || tab.isBlank()) {
|
||||
return false;
|
||||
}
|
||||
var placeholders = Pattern.compile("\\{(?:role|profile|model|n)}").matcher(template);
|
||||
StringBuilder expression = new StringBuilder("^");
|
||||
int literalStart = 0;
|
||||
while (placeholders.find()) {
|
||||
expression.append(Pattern.quote(template.substring(literalStart, placeholders.start())));
|
||||
expression.append(".*");
|
||||
literalStart = placeholders.end();
|
||||
}
|
||||
expression.append(Pattern.quote(template.substring(literalStart))).append("$");
|
||||
return Pattern.compile(expression.toString(), Pattern.CASE_INSENSITIVE).matcher(tab).matches();
|
||||
}
|
||||
|
||||
/** Case-insensitive prefix test that tolerates a null or blank label. */
|
||||
private static boolean startsWithIgnoreCase(String label, String prefix) {
|
||||
if (label == null || prefix == null || prefix.isBlank()) {
|
||||
return false;
|
||||
}
|
||||
String stripped = label.strip();
|
||||
return stripped.regionMatches(true, 0, prefix, 0, prefix.length());
|
||||
}
|
||||
|
||||
/**
|
||||
* Reject a profile that places its members by {@code "pane"} while any {@code fleet.leaders}
|
||||
* or {@code fleet.collaborators} entry names a {@code tab}. A pane-placed member lands inside
|
||||
* the focused tab rather than its own, so it can land inside a lead's or collaborator's own
|
||||
* labelled tab. {@link dev.ltms.fleet.herdr.LeadTabScanner} identifies a lead or collaborator
|
||||
* purely by that tab's label — it does not exclude the member space — so a member that ends up
|
||||
* there, while its pane carries no entry in the spawned-member roster, is read back as that
|
||||
* lead or collaborator and granted that identity's authority.
|
||||
*
|
||||
* <p>Only an entry with a non-blank {@code tab} is in scope: one with no {@code tab} feeds
|
||||
* nothing into {@link dev.ltms.fleet.herdr.LeadTabScanner}, so it creates no hazard here.
|
||||
*
|
||||
* @throws IllegalStateException when any {@code profiles:} entry is pane-placed while any
|
||||
* {@code fleet.leaders} or {@code fleet.collaborators} entry
|
||||
* names a non-blank {@code tab}
|
||||
*/
|
||||
public void validatePanePlacementAgainstLeadTabs() {
|
||||
if (fleet == null) {
|
||||
return;
|
||||
}
|
||||
boolean anyLeaderHasTab = fleet.leaders().values().stream()
|
||||
.anyMatch(leader -> leader != null && leader.tab() != null && !leader.tab().isBlank());
|
||||
boolean anyCollaboratorHasTab = fleet.collaborators().values().stream()
|
||||
.anyMatch(c -> c != null && c.tab() != null && !c.tab().isBlank());
|
||||
if (!anyLeaderHasTab && !anyCollaboratorHasTab) {
|
||||
return;
|
||||
}
|
||||
List<String> bad = new ArrayList<>();
|
||||
profiles().entrySet().stream()
|
||||
.filter(e -> !e.getValue().tabPlacement())
|
||||
.map(Map.Entry::getKey)
|
||||
.sorted()
|
||||
.forEach(bad::add);
|
||||
if (bad.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
throw new IllegalStateException("refusing to start: " + String.join("; ", bad)
|
||||
+ ". Every member labelled that way would be read back as a lead and granted "
|
||||
+ "spawn/stop/send on the whole fleet. Change one of the two so member tabs and "
|
||||
+ "lead tabs cannot be confused.");
|
||||
throw new IllegalStateException("refusing to start: profile(s) " + bad
|
||||
+ " use placement: pane while fleet.leaders or fleet.collaborators names a tab. A "
|
||||
+ "pane-placed member can land inside that labelled tab, and while its pane "
|
||||
+ "carries no entry in the spawned-member roster, it is read back as the lead or "
|
||||
+ "collaborator and granted that identity's authority. Set placement: tab for "
|
||||
+ "each named profile, or remove the tab from every fleet.leaders and "
|
||||
+ "fleet.collaborators entry.");
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -2757,15 +2977,6 @@ public record FleetConfig(
|
||||
}
|
||||
}
|
||||
|
||||
/** Case-insensitive prefix test that tolerates a null/blank label. */
|
||||
private static boolean startsWithIgnoreCase(String label, String prefix) {
|
||||
if (label == null || prefix == null || prefix.isBlank()) {
|
||||
return false;
|
||||
}
|
||||
String stripped = label.strip();
|
||||
return stripped.regionMatches(true, 0, prefix, 0, prefix.length());
|
||||
}
|
||||
|
||||
/**
|
||||
* Reject a subscription profile whose {@code env:} block tries to reseat the Anthropic binding
|
||||
* (CB-542).
|
||||
@@ -2850,8 +3061,14 @@ public record FleetConfig(
|
||||
* so duplicates are unrepresentable by construction once loaded — and {@link #load(Path)}
|
||||
* already rejects a duplicated slot name at parse time, before the map collapses.
|
||||
*
|
||||
* @throws IllegalStateException when a slot names no profile or an unknown one, or when a lead
|
||||
* can be neither found nor created, naming the offending entry
|
||||
* <p>Also rejects a {@code fleet.collaborators} entry with no (or a blank) {@code tab}. A
|
||||
* {@code profile}-less lead is still useful recognise-only — {@code tab} is the only field
|
||||
* that matters to it either way. A collaborator carries no other field at all, so a blank
|
||||
* {@code tab} leaves nothing for the entry to mean.
|
||||
*
|
||||
* @throws IllegalStateException when a slot names no profile or an unknown one, when a lead
|
||||
* can be neither found nor created, or when a collaborator names
|
||||
* no tab, naming the offending entry
|
||||
*/
|
||||
public void validateMembers() {
|
||||
if (fleet == null) {
|
||||
@@ -2889,6 +3106,16 @@ public record FleetConfig(
|
||||
+ "auto-launched, labelled) purely by its tab, so every entry must name one.");
|
||||
}
|
||||
});
|
||||
fleet.collaborators().forEach((name, collaborator) -> {
|
||||
if (collaborator == null) {
|
||||
return;
|
||||
}
|
||||
if (collaborator.tab() == null || collaborator.tab().isBlank()) {
|
||||
bad.add("fleet.collaborators." + name + " has no tab: — a collaborator is "
|
||||
+ "recognised purely by its tab, and carries no other field, so every "
|
||||
+ "entry must name one.");
|
||||
}
|
||||
});
|
||||
if (!bad.isEmpty()) {
|
||||
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
|
||||
}
|
||||
@@ -2937,11 +3164,11 @@ public record FleetConfig(
|
||||
* Runs every validator this class declares — found by reflection, not by name.
|
||||
*
|
||||
* <p>fleetd ticket "central allow-list of usable models", follow-up: mutation testing found
|
||||
* that although each of the six validators above was well pinned on its own, nothing proved
|
||||
* that although each validator above was well pinned on its own, nothing proved
|
||||
* either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()}) still
|
||||
* invoked it — deleting a call site left the full suite green. The fix is not a seventh test
|
||||
* per caller; a hand-maintained list of six names here would have the exact same defect its
|
||||
* own javadoc would warn against: the seventh validator someone adds next month has no reason
|
||||
* invoked it — deleting a call site left the full suite green. The fix is not one more test
|
||||
* per caller; a hand-maintained list of names here would have the exact same defect its
|
||||
* own javadoc would warn against: the next validator someone adds has no reason
|
||||
* to be added to it. So this method does not name any validator. It sweeps {@link
|
||||
* #getClass()}'s own public, no-argument, {@code void} methods whose name starts with {@code
|
||||
* "validate"} (excluding itself) and invokes every one it finds, via {@link
|
||||
@@ -2950,7 +3177,7 @@ public record FleetConfig(
|
||||
* which it silently never runs.
|
||||
*
|
||||
* <p>{@code Fleetd.main} and {@link ConfigRef#reload()} each call this one method instead of
|
||||
* the six individually — see the comments at those two call sites for why
|
||||
* each validator individually — see the comments at those two call sites for why
|
||||
* each must run it.
|
||||
*
|
||||
* <p>Methods run in a fixed (alphabetical) order, so a config with more than one violation
|
||||
@@ -2967,9 +3194,9 @@ public record FleetConfig(
|
||||
/**
|
||||
* The reflective sweep behind {@link #validateAll()}, kept as its own method — taking any
|
||||
* {@code target}, not just {@code this} — so a test can prove the MECHANISM is generic (it
|
||||
* would sweep a seventh {@code validateXxx()} method added to any class, not just something
|
||||
* special-cased to today's six on {@link FleetConfig}) without needing to add a real, unwanted
|
||||
* seventh validator to this class just to exercise that claim. See {@code
|
||||
* would sweep any new {@code validateXxx()} method added to any class, not just something
|
||||
* special-cased to the set {@link FleetConfig} declares today) without needing to add a real,
|
||||
* unwanted extra validator to this class just to exercise that claim. See {@code
|
||||
* FleetConfigValidateAllTest} for that proof.
|
||||
*
|
||||
* @param target an object whose public, no-argument, {@code void} methods named {@code
|
||||
|
||||
@@ -11,12 +11,18 @@ public final class HerdrRouter implements AutoCloseable {
|
||||
private final AgentControl memberAgents;
|
||||
private final WorkspaceControl leadSpaces;
|
||||
private final WorkspaceControl memberSpaces;
|
||||
private final Predicate<String> isLead;
|
||||
private final Predicate<String> routeToLead;
|
||||
|
||||
public HerdrRouter(HerdrClient lead, HerdrClient member, Predicate<String> isLead) {
|
||||
/**
|
||||
* @param routeToLead true for a terminal whose pane lives in the lead herdr daemon — a lead's
|
||||
* own pane or a configured collaborator's, both opened by a person at a
|
||||
* terminal rather than spawned, so both are found in the lead daemon rather
|
||||
* than the member one
|
||||
*/
|
||||
public HerdrRouter(HerdrClient lead, HerdrClient member, Predicate<String> routeToLead) {
|
||||
this.lead = Objects.requireNonNull(lead, "lead");
|
||||
this.member = member != null ? member : lead;
|
||||
this.isLead = Objects.requireNonNull(isLead, "isLead");
|
||||
this.routeToLead = Objects.requireNonNull(routeToLead, "routeToLead");
|
||||
leadAgents = new AgentControl(this.lead);
|
||||
memberAgents = this.member == this.lead ? leadAgents : new AgentControl(this.member);
|
||||
leadSpaces = new WorkspaceControl(this.lead);
|
||||
@@ -27,7 +33,7 @@ public final class HerdrRouter implements AutoCloseable {
|
||||
public WorkspaceControl leadSpaces() { return leadSpaces; }
|
||||
public AgentControl memberAgents() { return memberAgents; }
|
||||
public WorkspaceControl memberSpaces() { return memberSpaces; }
|
||||
public AgentControl agentsFor(String targetId) { return isLead.test(targetId) ? leadAgents : memberAgents; }
|
||||
public AgentControl agentsFor(String targetId) { return routeToLead.test(targetId) ? leadAgents : memberAgents; }
|
||||
|
||||
HerdrClient leadClient() { return lead; }
|
||||
HerdrClient memberClient() { return member; }
|
||||
|
||||
@@ -37,8 +37,11 @@ import java.util.function.Supplier;
|
||||
* <p><strong>Direction of trust.</strong> The label names the lead; it never <em>grants</em>
|
||||
* anything a pane could take for itself. Three properties keep that honest:
|
||||
* <ol>
|
||||
* <li>Worker spaces are excluded wholesale ({@code excludedWorkspaceLabels}), so a worker cannot
|
||||
* become a lead by being placed — as a split, say — inside a matching tab.</li>
|
||||
* <li>{@code excludedWorkspaceLabels} can filter a workspace out of the scan, but this class does
|
||||
* not by itself stop a worker from landing inside a matching tab — a caller may pass an empty
|
||||
* set, and the daemon does. The guard against that is {@code
|
||||
* FleetConfig.validatePanePlacementAgainstLeadTabs}: it refuses, at startup, any profile that
|
||||
* places members by pane while a lead names a tab.</li>
|
||||
* <li>A worker cannot rename a tab: {@code tab.rename} is reachable only through
|
||||
* {@link WorkspaceControl}, which no {@code fleet_*} tool exposes. The label is writable by
|
||||
* the human at the terminal and by nobody the bridge is defending against.</li>
|
||||
@@ -93,13 +96,19 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(LeadTabScanner.class);
|
||||
|
||||
/** What a matched tab names: a lead or a collaborator. */
|
||||
private enum Kind { LEAD, COLLABORATOR }
|
||||
|
||||
/** One matched tab's name and what it names. */
|
||||
private record Entry(String name, Kind kind) {}
|
||||
|
||||
private final HerdrClient herdr;
|
||||
private final Map<String, String> tabToName;
|
||||
private final Map<String, Entry> tabToEntry;
|
||||
private final Set<String> excludedWorkspaceLabels;
|
||||
private final long ttlNanos;
|
||||
private final LongSupplier clock;
|
||||
|
||||
private Map<String, String> cached = Map.of();
|
||||
private Map<String, Entry> cached = Map.of();
|
||||
private long scannedAtNanos;
|
||||
private boolean everScanned;
|
||||
|
||||
@@ -123,26 +132,53 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
|
||||
*/
|
||||
public LeadTabScanner(HerdrClient herdr, Map<String, String> tabToName,
|
||||
Set<String> excludedWorkspaceLabels, long ttlNanos, LongSupplier clock) {
|
||||
this(herdr, tabToName, Map.of(), excludedWorkspaceLabels, ttlNanos, clock);
|
||||
}
|
||||
|
||||
/**
|
||||
* As {@link #LeadTabScanner(HerdrClient, Map, Set, long, LongSupplier)}, additionally scanning
|
||||
* for configured collaborator tabs in the same pass.
|
||||
*
|
||||
* @param collaboratorTabToName every configured collaborator's exact tab label → its name
|
||||
* ({@code fleet.collaborators.<name>.tab}), matched the same way as
|
||||
* {@code tabToName}
|
||||
*/
|
||||
public LeadTabScanner(HerdrClient herdr, Map<String, String> tabToName,
|
||||
Map<String, String> collaboratorTabToName,
|
||||
Set<String> excludedWorkspaceLabels, long ttlNanos, LongSupplier clock) {
|
||||
this.herdr = herdr;
|
||||
this.tabToName = normalize(tabToName);
|
||||
this.tabToEntry = buildTabIndex(tabToName, collaboratorTabToName);
|
||||
this.excludedWorkspaceLabels = excludedWorkspaceLabels == null
|
||||
? Set.of() : Set.copyOf(excludedWorkspaceLabels);
|
||||
this.ttlNanos = ttlNanos;
|
||||
this.clock = clock;
|
||||
}
|
||||
|
||||
/** Keys stripped and lower-cased once, so every lookup is a plain map hit. */
|
||||
private static Map<String, String> normalize(Map<String, String> tabToName) {
|
||||
if (tabToName == null || tabToName.isEmpty()) {
|
||||
return Map.of();
|
||||
/**
|
||||
* Keys stripped and lower-cased once, so every lookup is a plain map hit. Leads and
|
||||
* collaborators merge into a single index, so {@link #scan()} matches both kinds in one pass
|
||||
* over the tab list; a label naming both a lead and a collaborator takes the lead entry —
|
||||
* leads are put last, so a colliding key's lead entry is the one that overwrites — since a lead
|
||||
* can already do everything a collaborator can. Config validation already refuses a lead and a
|
||||
* collaborator sharing one exact tab, so this ordering is defence in depth, not the control.
|
||||
*/
|
||||
private static Map<String, Entry> buildTabIndex(Map<String, String> tabToName,
|
||||
Map<String, String> collaboratorTabToName) {
|
||||
Map<String, Entry> out = new LinkedHashMap<>();
|
||||
putNormalized(out, collaboratorTabToName, Kind.COLLABORATOR);
|
||||
putNormalized(out, tabToName, Kind.LEAD);
|
||||
return Collections.unmodifiableMap(out);
|
||||
}
|
||||
|
||||
private static void putNormalized(Map<String, Entry> out, Map<String, String> tabToName, Kind kind) {
|
||||
if (tabToName == null) {
|
||||
return;
|
||||
}
|
||||
Map<String, String> out = new LinkedHashMap<>();
|
||||
tabToName.forEach((tab, name) -> {
|
||||
if (tab != null && !tab.isBlank() && name != null && !name.isBlank()) {
|
||||
out.put(tab.strip().toLowerCase(Locale.ROOT), name);
|
||||
out.put(tab.strip().toLowerCase(Locale.ROOT), new Entry(name, kind));
|
||||
}
|
||||
});
|
||||
return Collections.unmodifiableMap(out);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -153,6 +189,29 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
|
||||
*/
|
||||
@Override
|
||||
public synchronized Map<String, String> get() {
|
||||
return byKind(refresh(), Kind.LEAD);
|
||||
}
|
||||
|
||||
/**
|
||||
* The current {@code terminal_id → collaborator name} map, sharing the same scan and cache as
|
||||
* {@link #get()} — both kinds are matched in one pass, so this never costs a second herdr call.
|
||||
*/
|
||||
public synchronized Map<String, String> collaborators() {
|
||||
return byKind(refresh(), Kind.COLLABORATOR);
|
||||
}
|
||||
|
||||
private static Map<String, String> byKind(Map<String, Entry> entries, Kind kind) {
|
||||
Map<String, String> out = new LinkedHashMap<>();
|
||||
entries.forEach((terminal, entry) -> {
|
||||
if (entry.kind() == kind) {
|
||||
out.put(terminal, entry.name());
|
||||
}
|
||||
});
|
||||
return Collections.unmodifiableMap(out);
|
||||
}
|
||||
|
||||
/** Rescans if the cache has expired, otherwise returns the cached answer. */
|
||||
private Map<String, Entry> refresh() {
|
||||
long now = clock.getAsLong();
|
||||
if (everScanned && now - scannedAtNanos < ttlNanos) {
|
||||
return cached;
|
||||
@@ -162,21 +221,21 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
|
||||
scannedAtNanos = now;
|
||||
everScanned = true;
|
||||
try {
|
||||
Map<String, String> fresh = scan();
|
||||
Map<String, Entry> fresh = scan();
|
||||
if (!fresh.equals(cached)) {
|
||||
log.info("lead panes: {}", fresh);
|
||||
log.info("lead/collaborator panes: {}", fresh);
|
||||
}
|
||||
cached = fresh;
|
||||
} catch (HerdrException e) {
|
||||
log.warn("lead-tab scan failed, keeping the {} lead(s) already known: {}",
|
||||
log.warn("lead-tab scan failed, keeping the {} entr(y/ies) already known: {}",
|
||||
cached.size(), e.getMessage());
|
||||
}
|
||||
return cached;
|
||||
}
|
||||
|
||||
/** One full pass: labelled tabs → live agents in them → those panes' terminals. */
|
||||
private Map<String, String> scan() {
|
||||
Map<String, String> nameByTab = new LinkedHashMap<>();
|
||||
private Map<String, Entry> scan() {
|
||||
Map<String, Entry> entryByTab = new LinkedHashMap<>();
|
||||
for (JsonNode w : herdr.call("workspace.list").path("workspaces")) {
|
||||
Workspace ws = Workspace.from(w);
|
||||
if (ws.workspaceId() == null || excludedWorkspaceLabels.contains(ws.label())) {
|
||||
@@ -184,21 +243,22 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
|
||||
}
|
||||
for (JsonNode t : herdr.call("tab.list", Map.of("workspace_id", ws.workspaceId())).path("tabs")) {
|
||||
Tab tab = Tab.from(t);
|
||||
String name = leadNameOf(tab.label());
|
||||
if (name != null && tab.tabId() != null) {
|
||||
nameByTab.put(tab.tabId(), name);
|
||||
Entry entry = entryOf(tab.label());
|
||||
if (entry != null && tab.tabId() != null) {
|
||||
entryByTab.put(tab.tabId(), entry);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if (nameByTab.isEmpty()) {
|
||||
if (entryByTab.isEmpty()) {
|
||||
gracedTerminals = Set.of();
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
// fleetd #359: a labelled tab is only a lead when herdr also reports a running agent in
|
||||
// it — the same liveness signal LeadLauncher.countLeads trusts for the identical purpose.
|
||||
// Without this, a tab left behind by a session that has since died reads as live forever.
|
||||
// fleetd #359: a labelled tab is only a lead (or collaborator) when herdr also reports a
|
||||
// running agent in it — the same liveness signal LeadLauncher.countLeads trusts for the
|
||||
// identical purpose. Without this, a tab left behind by a session that has since died reads
|
||||
// as live forever.
|
||||
Set<String> tabsWithAgent = new HashSet<>();
|
||||
for (JsonNode a : herdr.call("agent.list").path("agents")) {
|
||||
String tabId = a.path("tab_id").asText(null);
|
||||
@@ -207,18 +267,18 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
|
||||
}
|
||||
}
|
||||
|
||||
Map<String, String> byTerminal = new LinkedHashMap<>();
|
||||
Map<String, Entry> byTerminal = new LinkedHashMap<>();
|
||||
Set<String> stillGraced = new HashSet<>();
|
||||
// One pane.list for every tab: panes carry tab_id, so the join is local.
|
||||
for (JsonNode p : herdr.call("pane.list", Map.of()).path("panes")) {
|
||||
String tabId = p.path("tab_id").asText(null);
|
||||
String name = nameByTab.get(tabId);
|
||||
Entry entry = entryByTab.get(tabId);
|
||||
String terminal = p.path("terminal_id").asText(null);
|
||||
if (name == null || terminal == null || terminal.isBlank()) {
|
||||
if (entry == null || terminal == null || terminal.isBlank()) {
|
||||
continue;
|
||||
}
|
||||
if (tabsWithAgent.contains(tabId)) {
|
||||
byTerminal.put(terminal, name);
|
||||
byTerminal.put(terminal, entry);
|
||||
continue;
|
||||
}
|
||||
// No agent reported for this tab, but its tab/pane are still here — this is the
|
||||
@@ -227,7 +287,7 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
|
||||
// reported as live; a terminal we never reported live gets none, so the original #359
|
||||
// fix (a genuinely dead tab is never reported) is unaffected for the common case.
|
||||
if (cached.containsKey(terminal) && !gracedTerminals.contains(terminal)) {
|
||||
byTerminal.put(terminal, name);
|
||||
byTerminal.put(terminal, entry);
|
||||
stillGraced.add(terminal);
|
||||
}
|
||||
}
|
||||
@@ -236,18 +296,19 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
|
||||
}
|
||||
|
||||
/**
|
||||
* The lead name a tab label declares, or {@code null} if it names none of the configured leads.
|
||||
* The entry a tab label declares, or {@code null} if it names neither a configured lead nor a
|
||||
* configured collaborator.
|
||||
*
|
||||
* <p>Exact match (case-insensitive, ends stripped) against {@link #tabToName} — no prefix
|
||||
* <p>Exact match (case-insensitive, ends stripped) against {@link #tabToEntry} — no prefix
|
||||
* stripping, so an operator's {@code "lead: something-else"} tab is never mistaken for a
|
||||
* configured lead just because it shares a prefix. The match strips a trailing
|
||||
* {@link PendingCloseMarker} first, so a tab {@code LeadLauncher} has flagged as maybe-dead but
|
||||
* not yet closed keeps resolving normally while that reconcile is pending.
|
||||
*/
|
||||
private String leadNameOf(String label) {
|
||||
private Entry entryOf(String label) {
|
||||
if (label == null) {
|
||||
return null;
|
||||
}
|
||||
return tabToName.get(PendingCloseMarker.strip(label).toLowerCase(Locale.ROOT));
|
||||
return tabToEntry.get(PendingCloseMarker.strip(label).toLowerCase(Locale.ROOT));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,146 @@
|
||||
package dev.ltms.fleet.herdr;
|
||||
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.util.List;
|
||||
import java.util.function.IntFunction;
|
||||
|
||||
/**
|
||||
* The three protections every {@code agent.start} caller needs against herdr's pane-typed launch
|
||||
* surface (fleetd #220, #727): a byte-limit check on the assembled command line, a bounded retry
|
||||
* on {@code agent_pane_busy} (the target pane's shell has not reached its prompt yet), and a
|
||||
* bounded retry on {@code agent_name_taken} with a fresh name each attempt. One implementation —
|
||||
* every caller of {@code agent.start}, lead or member, goes through this seam rather than carrying
|
||||
* its own copy.
|
||||
*/
|
||||
public final class ResilientAgentLaunch {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(ResilientAgentLaunch.class);
|
||||
|
||||
private ResilientAgentLaunch() {
|
||||
}
|
||||
|
||||
/**
|
||||
* The pty line buffer herdr types a launch command into: BSD/macOS {@code MAX_CANON}. Not a
|
||||
* fleetd choice and not configurable — see {@link #checkFits}.
|
||||
*/
|
||||
public static final int PANE_COMMAND_BYTE_LIMIT = 1024;
|
||||
|
||||
/** Per-argument allowance for the separating space and a shell quote pair fleetd cannot see. */
|
||||
private static final int QUOTING_OVERHEAD_PER_ARG = 3;
|
||||
|
||||
/** herdr rejects a duplicate agent {@code name}; a caller retries a bumped name this many times. */
|
||||
public static final int NAME_RETRIES = 8;
|
||||
|
||||
/**
|
||||
* Retries for {@code agent.start} against a seed pane whose shell has not reached its prompt
|
||||
* yet — {@code tab.create}/{@code pane.split} return as soon as the pane exists, and herdr
|
||||
* refuses to start an agent in a pane that is not "an available shell" ({@code agent_pane_busy}).
|
||||
*/
|
||||
public static final int SHELL_READY_RETRIES = 20;
|
||||
|
||||
/** Raised by {@link #checkFits} when the assembled command cannot fit the pane line. */
|
||||
public static final class TooLargeException extends RuntimeException {
|
||||
public TooLargeException(String message) {
|
||||
super(message);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Verify the assembled launch line fits the pane herdr types it into. herdr does not exec the
|
||||
* launch command — it TYPES it into the pane as one line, and a pty line buffer holds only
|
||||
* {@value #PANE_COMMAND_BYTE_LIMIT} bytes. Everything past that byte is dropped with no error
|
||||
* anywhere: herdr answers "agent started", the backend exits on the mangled argument it was
|
||||
* handed, the pane closes, and the only symptom is a readiness timeout with no reason. That is
|
||||
* how fleetd #214 broke every claude-code spawn — one 50-byte flag pushed a 978-byte command to
|
||||
* 1028, and the tail that got cut was {@code --autocompact 250000}.
|
||||
*
|
||||
* <p>So measure it here and refuse, loudly and immediately, rather than start something that
|
||||
* cannot work. The estimate is deliberately conservative: fleetd cannot see herdr's quoting, so
|
||||
* every argument is charged its own bytes plus a separator and a quote pair. An over-estimate
|
||||
* costs a clear error at a length that was already unsafe; an under-estimate would let the
|
||||
* silent truncation back in.
|
||||
*
|
||||
* @param label names the launch in the refusal message (a profile name)
|
||||
* @param argv the full argv, including the executable at index 0
|
||||
* @throws TooLargeException naming the limit, the estimate, and the longest argument
|
||||
*/
|
||||
public static void checkFits(String label, List<String> argv) {
|
||||
int bytes = 0;
|
||||
String longest = null;
|
||||
int longestBytes = 0;
|
||||
for (String arg : argv) {
|
||||
int argBytes = arg == null ? 0 : arg.getBytes(StandardCharsets.UTF_8).length;
|
||||
bytes += argBytes + QUOTING_OVERHEAD_PER_ARG;
|
||||
if (argBytes > longestBytes) {
|
||||
longestBytes = argBytes;
|
||||
longest = arg;
|
||||
}
|
||||
}
|
||||
if (bytes <= PANE_COMMAND_BYTE_LIMIT) {
|
||||
return;
|
||||
}
|
||||
String culprit = longest == null ? "<none>"
|
||||
: longest.substring(0, Math.min(longest.length(), 60)) + (longest.length() > 60 ? "…" : "");
|
||||
throw new TooLargeException(
|
||||
"launch command for " + label + " is about " + bytes + " bytes, over the "
|
||||
+ PANE_COMMAND_BYTE_LIMIT + "-byte limit of the pane line herdr types it into. "
|
||||
+ "The pty would drop the tail silently and the backend would exit on a mangled "
|
||||
+ "argument. Longest argument is " + longestBytes + " bytes: " + culprit
|
||||
+ " — move it off the command line (a file flag) or shorten it.");
|
||||
}
|
||||
|
||||
/**
|
||||
* Start an agent into {@code paneId}, retrying {@code agent_pane_busy} up to {@code retries}
|
||||
* times with {@code sleeper} run between attempts.
|
||||
*
|
||||
* @throws HerdrException the last {@code agent_pane_busy} failure once {@code retries} is
|
||||
* spent, or immediately for any other herdr failure
|
||||
*/
|
||||
public static Agent startAwaitingShellPrompt(AgentControl agents, String name, String kind,
|
||||
List<String> args, String paneId,
|
||||
int retries, Runnable sleeper) {
|
||||
HerdrException busy = null;
|
||||
for (int attempt = 0; attempt < retries; attempt++) {
|
||||
try {
|
||||
return agents.start(name, kind, args, paneId);
|
||||
} catch (HerdrException e) {
|
||||
if (!"agent_pane_busy".equals(e.code())) throw e;
|
||||
log.debug("pane {} not at its shell prompt yet, retrying agent.start", paneId);
|
||||
busy = e;
|
||||
sleeper.run();
|
||||
}
|
||||
}
|
||||
throw busy;
|
||||
}
|
||||
|
||||
/**
|
||||
* Start an agent under a freshly generated name each attempt, retrying {@code agent_name_taken}
|
||||
* up to {@code nameRetries} times — herdr refuses a duplicate {@code name} outright, so a stale
|
||||
* registry entry (a crashed session, a name the registry has not yet released) must not block a
|
||||
* legitimate relaunch. Each attempt also carries its own {@link #startAwaitingShellPrompt} retry.
|
||||
*
|
||||
* @param nameForAttempt called once per attempt (0-based) to produce that attempt's name
|
||||
* @throws HerdrException the last {@code agent_name_taken} failure once {@code nameRetries} is
|
||||
* spent, or immediately for any other herdr failure
|
||||
*/
|
||||
public static Agent startUniquelyNamed(AgentControl agents, String kind, List<String> args,
|
||||
String paneId, IntFunction<String> nameForAttempt,
|
||||
int nameRetries, int shellReadyRetries, Runnable sleeper) {
|
||||
HerdrException last = null;
|
||||
for (int attempt = 0; attempt < nameRetries; attempt++) {
|
||||
String name = nameForAttempt.apply(attempt);
|
||||
try {
|
||||
return startAwaitingShellPrompt(agents, name, kind, args, paneId,
|
||||
shellReadyRetries, sleeper);
|
||||
} catch (HerdrException e) {
|
||||
if (!"agent_name_taken".equals(e.code())) throw e;
|
||||
log.debug("agent name '{}' taken, retrying", name);
|
||||
last = e;
|
||||
}
|
||||
}
|
||||
throw last;
|
||||
}
|
||||
}
|
||||
@@ -172,12 +172,6 @@ public final class LeadContextGauge {
|
||||
* {@code "claude"} (including {@code null}, meaning undetected) reports
|
||||
* {@link State#UNKNOWN} — this reader only understands Claude Code's own
|
||||
* transcript format
|
||||
*/
|
||||
public Reading read(String configDir, String sessionId, String agentType) {
|
||||
return read(configDir, sessionId, agentType, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* @param effectiveWindowTokens the caller's resolved effective auto-compact window for this
|
||||
* lead's own profile, or {@code null} when it cannot be resolved.
|
||||
* HIGH fires at {@link #HIGH_THRESHOLD_FRACTION} of this value;
|
||||
|
||||
@@ -5,6 +5,7 @@ import dev.ltms.fleet.herdr.Agent;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.HerdrException;
|
||||
import dev.ltms.fleet.herdr.PendingCloseMarker;
|
||||
import dev.ltms.fleet.herdr.ResilientAgentLaunch;
|
||||
import dev.ltms.fleet.herdr.Tab;
|
||||
import dev.ltms.fleet.herdr.Workspace;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
@@ -12,6 +13,7 @@ import dev.ltms.fleet.launch.ClaudeCodeArguments;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.security.SecureRandom;
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.LinkedHashSet;
|
||||
@@ -19,6 +21,7 @@ import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
|
||||
/**
|
||||
@@ -80,9 +83,19 @@ public final class LeadLauncher {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(LeadLauncher.class);
|
||||
|
||||
/** Attempts {@link #relaunch(String)} makes before giving up and returning {@code null}. */
|
||||
static final int RELAUNCH_ATTEMPTS = 3;
|
||||
|
||||
private final AgentControl agents;
|
||||
private final WorkspaceControl spaces;
|
||||
private final FleetConfig cfg;
|
||||
private final Runnable sleeper;
|
||||
|
||||
// Per-process token mixed into each lead agent name so a fresh daemon process (seq back at 0)
|
||||
// cannot collide with a same-name lead that outlived a restart — the same scheme
|
||||
// HerdrPeerLauncher uses for members (fleetd #727).
|
||||
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
|
||||
private final AtomicLong nameSeq = new AtomicLong();
|
||||
|
||||
/**
|
||||
* @param agents herdr agent control (start, list)
|
||||
@@ -90,9 +103,27 @@ public final class LeadLauncher {
|
||||
* @param cfg the loaded config — {@code fleet.leaders}, {@code profiles} and each lead's tab
|
||||
*/
|
||||
public LeadLauncher(AgentControl agents, WorkspaceControl spaces, FleetConfig cfg) {
|
||||
this(agents, spaces, cfg, () -> sleepUninterruptibly(300));
|
||||
}
|
||||
|
||||
/**
|
||||
* Test seam: as above, plus an injectable {@code sleeper} for the {@code agent_pane_busy}
|
||||
* retry (fleetd #727), so a test can prove the retry budget without a real sleep.
|
||||
*/
|
||||
LeadLauncher(AgentControl agents, WorkspaceControl spaces, FleetConfig cfg, Runnable sleeper) {
|
||||
this.agents = agents;
|
||||
this.spaces = spaces;
|
||||
this.cfg = cfg;
|
||||
this.sleeper = sleeper;
|
||||
}
|
||||
|
||||
/** Uninterruptible sleep — the production {@link #sleeper} between {@code agent_pane_busy} retries. */
|
||||
private static void sleepUninterruptibly(long ms) {
|
||||
try {
|
||||
Thread.sleep(ms);
|
||||
} catch (InterruptedException e) {
|
||||
Thread.currentThread().interrupt();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -173,24 +204,14 @@ public final class LeadLauncher {
|
||||
log.info("lead '{}': {} live, {} wanted — nothing to start", name, running, wanted);
|
||||
continue;
|
||||
}
|
||||
if (!lead.isCreatable()) {
|
||||
// A lead with a `tab:` but no `profile:` is recognise-only by design: the operator
|
||||
// opens it by hand. Say so once rather than looking like a silent failure.
|
||||
log.info("lead '{}' is not live, and names no profile — it can be recognised but not "
|
||||
+ "launched. Add `profile:` under fleet.leaders.{} to have fleetd start it.",
|
||||
name, name);
|
||||
continue;
|
||||
}
|
||||
|
||||
FleetConfig.Profile profile = cfg.profiles().get(lead.profile());
|
||||
if (profile == null) {
|
||||
log.warn("lead '{}' names profile '{}', which is not configured — not launching",
|
||||
name, lead.profile());
|
||||
ResolvedLead resolved = resolveLaunchable(name);
|
||||
if (resolved == null) {
|
||||
continue;
|
||||
}
|
||||
|
||||
for (int i = running; i < wanted; i++) {
|
||||
if (launch(name, lead, profile)) {
|
||||
if (launch(name, resolved.lead(), resolved.profile()) != null) {
|
||||
started++;
|
||||
}
|
||||
}
|
||||
@@ -198,11 +219,87 @@ public final class LeadLauncher {
|
||||
return started;
|
||||
}
|
||||
|
||||
/** A declared lead paired with the profile it launches on — {@link #resolveLaunchable}'s result. */
|
||||
private record ResolvedLead(FleetConfig.Leader lead, FleetConfig.Profile profile) {
|
||||
}
|
||||
|
||||
/**
|
||||
* The declared {@code Leader} and its {@code Profile} for {@code name}, read from the config
|
||||
* snapshot this launcher was constructed with.
|
||||
*
|
||||
* @return the resolved pair, or {@code null} (having logged) if {@code name} is not declared
|
||||
* under {@code fleet.leaders}, that lead names no {@code profile:} (a {@code tab:}-only,
|
||||
* recognise-only lead), or its {@code profile:} is not configured. Shared by
|
||||
* {@link #ensureLeads()} and {@link #relaunch(String)} so the three refusals and their
|
||||
* wording live in one place.
|
||||
*/
|
||||
private ResolvedLead resolveLaunchable(String name) {
|
||||
FleetConfig.Leader lead = cfg.fleet().leaders().get(name);
|
||||
if (lead == null) {
|
||||
log.warn("lead '{}' is not declared under fleet.leaders — not launching", name);
|
||||
return null;
|
||||
}
|
||||
if (!lead.isCreatable()) {
|
||||
// A lead with a `tab:` but no `profile:` is recognise-only by design: the operator
|
||||
// opens it by hand. Say so once rather than looking like a silent failure.
|
||||
log.info("lead '{}' names no profile — it can be recognised but not launched. Add "
|
||||
+ "`profile:` under fleet.leaders.{} to have fleetd start it.", name, name);
|
||||
return null;
|
||||
}
|
||||
|
||||
FleetConfig.Profile profile = cfg.profiles().get(lead.profile());
|
||||
if (profile == null) {
|
||||
log.warn("lead '{}' names profile '{}', which is not configured — not launching",
|
||||
name, lead.profile());
|
||||
return null;
|
||||
}
|
||||
return new ResolvedLead(lead, profile);
|
||||
}
|
||||
|
||||
/**
|
||||
* Start the named lead from the config snapshot this launcher was constructed with — not a
|
||||
* live read, so a lead's {@code profile:} or {@code tab:} edited in config needs a daemon
|
||||
* restart to take effect here — outside of {@link #ensureLeads()}'s {@code instances}
|
||||
* bookkeeping.
|
||||
*
|
||||
* @return the started {@link Agent}, or {@code null} if {@code name} is not declared under
|
||||
* {@code fleet.leaders}, that lead names no {@code profile:} (a {@code tab:}-only,
|
||||
* recognise-only lead), its {@code profile:} is not configured, or every attempt up to
|
||||
* {@link #RELAUNCH_ATTEMPTS} failed to start it. Never throws.
|
||||
*
|
||||
* <p>Does not count how many instances of this lead are already live. {@link #ensureLeads()}'s
|
||||
* count exists to avoid starting a second orchestrator; the caller of this method has already
|
||||
* decided to replace the lead and owns that decision.
|
||||
*
|
||||
* <p>Retries the whole launch attempt — not only the {@code agent_name_taken}/
|
||||
* {@code agent_pane_busy} cases {@link ResilientAgentLaunch} already retries inside one
|
||||
* {@code agents.start} call — up to {@link #RELAUNCH_ATTEMPTS} times, sleeping via the
|
||||
* injected sleeper between attempts, and returns the agent from the first attempt that
|
||||
* succeeds.
|
||||
*/
|
||||
public Agent relaunch(String name) {
|
||||
ResolvedLead resolved = resolveLaunchable(name);
|
||||
if (resolved == null) {
|
||||
return null;
|
||||
}
|
||||
|
||||
for (int attempt = 1; attempt <= RELAUNCH_ATTEMPTS; attempt++) {
|
||||
Agent started = launch(name, resolved.lead(), resolved.profile());
|
||||
if (started != null) {
|
||||
return started;
|
||||
}
|
||||
if (attempt < RELAUNCH_ATTEMPTS) {
|
||||
sleeper.run();
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/**
|
||||
* How many live leads exist per configured name, and which of that name's labelled tabs are
|
||||
* <em>not</em> live: a running agent in a tab labelled with that lead's exact {@code tab}
|
||||
* (CB-579). Member workspaces are excluded, exactly as the scanner excludes them: a member must
|
||||
* not be counted as a lead because it happens to sit in a matching tab.
|
||||
* (CB-579). A member sitting in the same shared workspace is not counted as a lead because its
|
||||
* tab carries a different label, not because any workspace is excluded from this count.
|
||||
*
|
||||
* <p>There used to be a second path here — a running agent on the terminal a
|
||||
* {@code fleet.leaders.<name>.terminal} pin named, for a lead opened and pinned by hand. That
|
||||
@@ -306,8 +403,17 @@ public final class LeadLauncher {
|
||||
return null;
|
||||
}
|
||||
|
||||
/** Start one lead. Returns false (having logged) rather than throwing on any failure. */
|
||||
private boolean launch(String name, FleetConfig.Leader lead, FleetConfig.Profile profile) {
|
||||
/**
|
||||
* Start one lead. Returns null (having logged) rather than throwing on any failure.
|
||||
*
|
||||
* <p>Goes through the same {@link ResilientAgentLaunch} seam every member spawn uses
|
||||
* (fleetd #727): the assembled argv is refused outright if it cannot fit the pane line herdr
|
||||
* types it into, a stale {@code agent_name_taken} (a crashed session's name the registry has
|
||||
* not yet released) is retried under a fresh per-attempt name rather than refusing the whole
|
||||
* relaunch, and a seed pane whose shell has not reached its prompt yet ({@code
|
||||
* agent_pane_busy}) is retried rather than failing on the first miss.
|
||||
*/
|
||||
private Agent launch(String name, FleetConfig.Leader lead, FleetConfig.Profile profile) {
|
||||
String label = lead.tabLabel();
|
||||
String cwd = (lead.cwd() == null || lead.cwd().isBlank())
|
||||
? System.getProperty("user.dir") : lead.cwd();
|
||||
@@ -323,8 +429,12 @@ public final class LeadLauncher {
|
||||
// Same shape as the member launchers: herdr resolves the executable from `kind`, so
|
||||
// argv[0] (the configured launcher, e.g. `ccs`) is dropped and only the rest is passed.
|
||||
List<String> argv = leadArgv(profile);
|
||||
Agent started = agents.start("lead-" + name, herdrKind(profile),
|
||||
argv.isEmpty() ? argv : argv.subList(1, argv.size()), tab.rootPaneId());
|
||||
ResilientAgentLaunch.checkFits(profile.profile(), argv);
|
||||
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
|
||||
Agent started = ResilientAgentLaunch.startUniquelyNamed(agents, herdrKind(profile), args,
|
||||
tab.rootPaneId(),
|
||||
attempt -> "lead-" + name + "-" + nameNonce + "-" + nameSeq.incrementAndGet(),
|
||||
ResilientAgentLaunch.NAME_RETRIES, ResilientAgentLaunch.SHELL_READY_RETRIES, sleeper);
|
||||
|
||||
// Label AFTER the start succeeds. A label written before would survive a failed start
|
||||
// and then read back as a live lead on the next boot, which is the exact staleness the
|
||||
@@ -334,7 +444,7 @@ public final class LeadLauncher {
|
||||
log.info("lead '{}' launched: profile={} tab={} pane={} terminal={} label='{}' cwd={}",
|
||||
name, profile.profile(), tab.tab().tabId(), started.paneId(),
|
||||
started.terminalId(), label, cwd);
|
||||
return true;
|
||||
return started;
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("lead '{}' failed to launch on profile '{}': {}",
|
||||
name, profile.profile(), e.getMessage());
|
||||
@@ -346,7 +456,7 @@ public final class LeadLauncher {
|
||||
tab.tab().tabId(), cleanup.getMessage());
|
||||
}
|
||||
}
|
||||
return false;
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -1,8 +1,11 @@
|
||||
package dev.ltms.fleet.lead;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.herdr.Agent;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.AgentStatus;
|
||||
import dev.ltms.fleet.herdr.HerdrException;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
@@ -24,60 +27,69 @@ import java.util.function.Supplier;
|
||||
/**
|
||||
* fleetd #480: replace a lead session that has decided it is ready to be rolled over, without an
|
||||
* operator doing it by hand. A lead writes a handover file, calls {@link #open}, and then — once
|
||||
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation
|
||||
* clears the lead's own pane and bootstraps a fresh session against that file.
|
||||
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation ends
|
||||
* the lead's own pane, launches a fresh one, and bootstraps that fresh session against the file.
|
||||
*
|
||||
* <p>This is the executor behind the {@code fleet_handover} MCP tool ({@code
|
||||
* dev.ltms.fleet.mcp.FleetMcp#handover}), which drives {@link #open}, {@link #confirm}, {@link
|
||||
* #cancel}, and {@link #status} from a tool call — wired in fleetd #480 Unit C. <strong>An earlier
|
||||
* version of this paragraph said nothing called this class at all; that stopped being true once
|
||||
* that unit landed, and this correction exists so the javadoc does not go on claiming it.</strong>
|
||||
* #cancel}, and {@link #status} from a tool call.
|
||||
*
|
||||
* <p><strong>{@code confirm()} cannot roll inline — a fleetd #480 correction.</strong> The first
|
||||
* version of this class called {@code agents.send(lead, "/clear")} directly from inside {@code
|
||||
* confirm()}, then polled for the pane to become injectable again. That is wrong, because {@code
|
||||
* confirm()} is called BY the lead, FROM the lead's own turn: the lead's pane is {@code WORKING}
|
||||
* for the whole duration of that call and cannot possibly report injectable until {@code confirm()}
|
||||
* itself returns. The poll always timed out — but only after the {@code /clear} had already been
|
||||
* sent and queued in the pane, where it fired the instant the turn ended anyway. The result was the
|
||||
* worst outcome this feature can produce: a silently destroyed lead context with no fresh session
|
||||
* ever started, and a refusal return value that claimed nothing had happened.
|
||||
*
|
||||
* <p>The fix: {@link #confirm} validates every gate, then does no I/O against the lead's own pane
|
||||
* at all — it only records that the request is approved and hands a one-shot continuation to
|
||||
* {@code continuationRunner} before returning. That continuation is what actually touches the pane,
|
||||
* once the calling turn has ended, in this order:
|
||||
* <p><strong>{@code confirm()} cannot roll inline.</strong> {@code confirm()} is called BY the
|
||||
* lead, FROM the lead's own turn: the lead's pane is {@code WORKING} for the whole duration of that
|
||||
* call and cannot possibly report a real turn boundary until {@code confirm()} itself returns. So
|
||||
* {@link #confirm} validates every gate, then does no I/O against the lead's own pane at all — it
|
||||
* only records that the request is approved and hands a one-shot continuation to {@code
|
||||
* continuationRunner} before returning. That continuation is what actually touches the pane, once
|
||||
* the calling turn has ended, in this order:
|
||||
* <ol>
|
||||
* <li>wait for the lead's own pane to report a real turn boundary — {@code IDLE} or {@code
|
||||
* DONE}, never merely {@code BLOCKED} — i.e. wait for the very {@code confirm()} call that
|
||||
* approved this roll to finish its turn — bounded by {@code turnSettleSeconds}. <strong>If
|
||||
* this never happens, nothing else in this list runs: no {@code /clear} is ever sent.</strong>
|
||||
* A lead that never goes idle is a lead still doing real work, and clearing it would throw
|
||||
* away live context — exactly the failure this correction exists to prevent.</li>
|
||||
* <li>{@code agents.send(lead, "/clear")}</li>
|
||||
* <li>wait for {@code /clear} to be picked up and settle, bounded by {@code clearSettleSeconds}
|
||||
* (fleetd #489: no longer a plain re-check of the same boundary — {@code /clear} starts no
|
||||
* turn of its own, so this instead nudges the submit keystroke while no pickup has been seen,
|
||||
* then waits for a real {@code WORKING} → {@code IDLE}/{@code DONE} boundary once one has;
|
||||
* see {@link #waitForClearPickupAndSettle})</li>
|
||||
* <li>{@code agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()))}</li>
|
||||
* this never happens, nothing else in this list runs: the old pane is never touched.</strong>
|
||||
* A lead that never goes idle is a lead still doing real work, and tearing it down would throw
|
||||
* away live context.</li>
|
||||
* <li>capture the old pane id (and, through it, the old tab) from {@link AgentControl#get}, with
|
||||
* a bounded retry — the terminal-to-pane lookup it goes through can itself report a genuinely
|
||||
* live agent as not found (see {@code AgentControl#agentCall}'s own re-resolve-once
|
||||
* behaviour), and one false negative here must not abort an otherwise-healthy roll. Neither id
|
||||
* is ever re-resolved from the terminal again after this — once the pane below is closed there
|
||||
* is nothing left to resolve it from.</li>
|
||||
* <li>resolve the lead's configured name from its terminal, for the relaunch step below.</li>
|
||||
* <li>end the old session: close the pane (an already-gone pane counts as success; any other
|
||||
* failure propagates), then close its tab only when the pane was that tab's sole occupant —
|
||||
* the same pane-then-tab teardown {@code HerdrPeerLauncher#stop} uses for a member.</li>
|
||||
* <li>confirm the old pane is actually gone by polling {@link
|
||||
* dev.ltms.fleet.herdr.WorkspaceControl#locatePane} for a {@code null} result — never {@link
|
||||
* AgentControl#status}, and never the live-lead terminal map, each of which answers a
|
||||
* different question. <strong>If the old pane is never confirmed gone, no relaunch is
|
||||
* attempted</strong> — see {@link RollState#OLD_PANE_NEVER_DIED}.</li>
|
||||
* <li>launch a fresh lead with {@code LeadLauncher#relaunch}. <strong>If every attempt fails,
|
||||
* {@code bootstrapText} is never sent</strong> — see {@link RollState#RELAUNCH_FAILED}.</li>
|
||||
* <li>wait for the fresh pane to reach a real turn boundary ({@code IDLE} or {@code DONE},
|
||||
* never merely {@code BLOCKED}), bounded by {@code relaunchReadySeconds}. This is the
|
||||
* safety gate: typing into a pane that has not actually finished booting loses the
|
||||
* keystrokes. <strong>If the pane never becomes ready, {@code bootstrapText} is never
|
||||
* sent</strong> — see {@link RollState#RELAUNCH_NEVER_READY}.</li>
|
||||
* <li>wait for the fresh terminal to be recognised as a live lead — present in the live-lead
|
||||
* terminal map — bounded by {@code relaunchReadySeconds}. This is bookkeeping, not a
|
||||
* safety gate: {@code bootstrapText} is sent either way once the pane is ready, whether or
|
||||
* not this wait itself times out — see {@link RollState#RELAUNCH_NOT_RECOGNISED}.</li>
|
||||
* <li>{@code agents.send(newTerminal, cfg.bootstrapTextFor(p.handoverPath()))} — sent to the
|
||||
* FRESH terminal, never the one that was just torn down.</li>
|
||||
* </ol>
|
||||
* A {@link #confirm} that returns {@link RollDecision#approved()} therefore means <em>"every gate
|
||||
* passed and the roll is scheduled"</em>, never <em>"the pane has been cleared"</em> — the pane may
|
||||
* still be mid-turn, possibly for a long time, when the caller gets that answer back.
|
||||
* passed and the roll is scheduled"</em>, never <em>"the lead has already been replaced"</em> — the
|
||||
* old pane may still be mid-turn, possibly for a long time, when the caller gets that answer back.
|
||||
*
|
||||
* <p><strong>The safety invariant survives this change, restated precisely.</strong> The ticket
|
||||
* that first defined this class required "no timer, no scheduler, no background thread" so that
|
||||
* nothing but an explicit {@link #confirm} call could ever cause a {@code /clear}. That invariant
|
||||
* is about INITIATIVE, not about synchronicity, and this correction keeps it: {@code
|
||||
* continuationRunner} launches a single-shot task that exists only because one specific,
|
||||
* <p><strong>The safety invariant.</strong> "No timer, no scheduler, no background thread" means
|
||||
* that nothing but an explicit {@link #confirm} call can ever tear a lead's pane down.
|
||||
* {@code continuationRunner} launches a single-shot task that exists only because one specific,
|
||||
* already-approved {@link #confirm} call created it — it is not recurring, it is not started at
|
||||
* construction time or on any schedule, and no two invocations of it ever share state. A recurring
|
||||
* heartbeat or timer that could decide on its own initiative to roll a pane is still, and will
|
||||
* always be, absent from this class. <strong>Nothing but an explicit {@link #confirm} call that
|
||||
* passes every gate can ever cause a {@code /clear} — that call may simply finish its own work
|
||||
* slightly later than the method return, as a continuation of the same approved request, rather
|
||||
* than entirely inside the method body.</strong>
|
||||
* heartbeat or timer that could decide on its own initiative to roll a pane is absent from this
|
||||
* class. <strong>Nothing but an explicit {@link #confirm} call that passes every gate can ever tear
|
||||
* a pane down — that call may simply finish its own work slightly later than the method return, as
|
||||
* a continuation of the same approved request, rather than entirely inside the method body.</strong>
|
||||
*
|
||||
* <p><strong>Identity is resolved by the caller, never looked up here — a second fleetd #480
|
||||
* correction.</strong> The first version resolved the pane to clear via {@code
|
||||
@@ -115,19 +127,17 @@ public final class LeadRollover {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(LeadRollover.class);
|
||||
|
||||
/** Poll interval while waiting for the lead's pane to settle after {@code /clear}. */
|
||||
static final long SETTLE_POLL_MS = 250;
|
||||
/** Poll interval shared by every bounded wait in this class. */
|
||||
static final long POLL_INTERVAL_MS = 250;
|
||||
|
||||
/**
|
||||
* How many consecutive not-yet-picked-up polls {@link #waitForClearPickupAndSettle} allows
|
||||
* before releasing rather than wedging the roll — the same constant and the same
|
||||
* release-not-wedge choice {@link dev.ltms.fleet.inject.Injector} already makes for its own
|
||||
* post-turn {@code /clear} housekeeping (fleetd #306). <strong>This bounds the number of
|
||||
* consecutive polls, not the number of nudges:</strong> the first {@code PICKUP_GRACE_POLLS - 1}
|
||||
* of those polls each send a nudge, and the {@code PICKUP_GRACE_POLLS}th releases instead of
|
||||
* nudging again — so 8 polls produce 7 nudges, not 8.
|
||||
* How long {@link #waitUntilPaneGone} polls {@link WorkspaceControl#locatePane} before giving
|
||||
* up on ever seeing the old pane disappear. Not configurable: once {@link #endOldSession} has
|
||||
* closed the pane (and, usually, its tab), herdr dropping the pane from its own bookkeeping is
|
||||
* expected to show up within one or two polls, not on an operator-tunable timescale the way a
|
||||
* CLI boot is.
|
||||
*/
|
||||
static final int PICKUP_GRACE_POLLS = 8;
|
||||
static final int PANE_DEATH_TIMEOUT_SECONDS = 10;
|
||||
|
||||
/**
|
||||
* One request opened by {@link #open}, pending its {@link #confirm} (or {@link #cancel}).
|
||||
@@ -164,15 +174,21 @@ public final class LeadRollover {
|
||||
* The handover file's modified time is not after {@link #open}'s request timestamp, or is
|
||||
* older than {@code maxDocAgeSeconds}.
|
||||
*/
|
||||
HANDOVER_STALE
|
||||
HANDOVER_STALE,
|
||||
/**
|
||||
* This lead terminal already has a roll running: an earlier {@link #confirm} call claimed
|
||||
* it and that roll's continuation has not released it yet. {@code detail} names the lead
|
||||
* terminal and the token that holds the claim.
|
||||
*/
|
||||
ROLL_ALREADY_RUNNING
|
||||
}
|
||||
|
||||
/**
|
||||
* The outcome of a {@link #confirm} call. {@link #approved()} means every gate passed and the
|
||||
* roll has been handed to a one-shot continuation — <strong>not</strong> that the pane has been
|
||||
* cleared; the continuation may still be waiting for the calling turn to end when this returns.
|
||||
* Whether the deferred roll itself later goes on to clear the pane, refuse for never going
|
||||
* idle, or refuse for never re-settling after {@code /clear} is logged only (see this class's
|
||||
* roll has been handed to a one-shot continuation — <strong>not</strong> that the lead has
|
||||
* already been replaced; the continuation may still be waiting for the calling turn to end when
|
||||
* this returns. Whether the deferred roll itself later goes on to tear the old pane down and
|
||||
* relaunch the lead, or refuses at any of its own steps, is logged only (see this class's
|
||||
* javadoc) — there is deliberately no synchronous caller left by that point to hand a result to.
|
||||
*/
|
||||
public record RollDecision(boolean accepted, RefusalReason reason, String detail) {
|
||||
@@ -208,7 +224,7 @@ public final class LeadRollover {
|
||||
|
||||
/**
|
||||
* What is known about one token, right now — the answer {@link #status} gives. Distinguishes
|
||||
* three terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
|
||||
* five terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
|
||||
* that has been approved but has not finished yet, and two answers for a token that names no
|
||||
* active work at all: still pending confirmation, or nothing known about this token at all.
|
||||
*/
|
||||
@@ -231,41 +247,64 @@ public final class LeadRollover {
|
||||
* #status} could wrongly answer {@link #UNKNOWN} ("nothing was ever requested") for a roll
|
||||
* that is, in fact, actively running. This is not sticky: the deferred continuation
|
||||
* overwrites this same entry with a terminal state ({@link #ROLLED}, {@link
|
||||
* #TURN_NEVER_SETTLED}, {@link #CLEAR_NEVER_SETTLED}, or {@link #FAILED}) once it finishes
|
||||
* — including by throwing, which fleetd #615's catch in {@link #runRollover} now turns into
|
||||
* {@link #FAILED} instead of leaving this entry stuck forever.
|
||||
* #TURN_NEVER_SETTLED}, {@link #OLD_PANE_NEVER_DIED}, {@link #RELAUNCH_FAILED}, {@link
|
||||
* #RELAUNCH_NOT_RECOGNISED}, or {@link #FAILED}) once it finishes — including by throwing,
|
||||
* which {@link #runRollover}'s catch turns into {@link #FAILED} instead of leaving this
|
||||
* entry stuck forever.
|
||||
*/
|
||||
IN_PROGRESS,
|
||||
/**
|
||||
* {@link #confirm} was approved and the deferred continuation completed the entire roll:
|
||||
* the calling lead's turn settled, {@code /clear} was sent and settled, and {@code
|
||||
* bootstrapText} was sent.
|
||||
* {@link #confirm} was approved and the deferred continuation completed the entire roll: the
|
||||
* calling lead's turn settled, the old pane was torn down and confirmed gone, a fresh lead
|
||||
* was launched and recognised, and {@code bootstrapText} was sent to it.
|
||||
*/
|
||||
ROLLED,
|
||||
/**
|
||||
* {@link #confirm} was approved, but the calling lead's own turn never reached a boundary
|
||||
* (IDLE or DONE) within {@code turnSettleSeconds} — no {@code /clear} was ever sent, at
|
||||
* all. This is the branch the fleetd #480 correction exists to make safe, and the one this
|
||||
* status exists to make VISIBLE: before this, a lead that hit this case had no way to find
|
||||
* out, and would carry on believing it was about to be replaced. See this class's javadoc.
|
||||
* (IDLE or DONE) within {@code turnSettleSeconds} — the old pane was never touched at all.
|
||||
* This is the state that makes a lead's own stuck turn VISIBLE: without it, a lead that hit
|
||||
* this case would have no way to find out, and would carry on believing it was about to be
|
||||
* replaced. See this class's javadoc.
|
||||
*/
|
||||
TURN_NEVER_SETTLED,
|
||||
/**
|
||||
* {@link #confirm} was approved and {@code /clear} was sent, but the pane never re-settled
|
||||
* within {@code clearSettleSeconds} — {@code bootstrapText} was never sent.
|
||||
* {@link #confirm} was approved and the calling lead's turn settled, the old pane was closed
|
||||
* (and its tab, if it was the sole occupant), but {@link
|
||||
* dev.ltms.fleet.herdr.WorkspaceControl#locatePane} kept reporting it as still present for
|
||||
* the whole pane-death timeout. No relaunch was ever attempted, and {@code bootstrapText}
|
||||
* was never sent.
|
||||
*/
|
||||
CLEAR_NEVER_SETTLED,
|
||||
OLD_PANE_NEVER_DIED,
|
||||
/**
|
||||
* fleetd #615: the deferred continuation threw a {@link RuntimeException} — most likely a
|
||||
* {@link dev.ltms.fleet.herdr.HerdrException} out of one of the two unwrapped {@code
|
||||
* agents.send} calls in {@link #runRollover} — and the continuation thread died with it.
|
||||
* Before this state existed, that throw left {@link #outcomes} holding {@link #IN_PROGRESS}
|
||||
* forever, because the production {@code continuationRunner} is a bare virtual thread with
|
||||
* no uncaught-exception handler and nothing downstream of the throw ever ran to write a
|
||||
* terminal outcome. {@code detail} names the exception, so a reader has something to act on
|
||||
* — the same diagnostic style as {@link #TURN_NEVER_SETTLED} and {@link
|
||||
* #CLEAR_NEVER_SETTLED}. The roll is dead at this point and does not retry itself; a stuck
|
||||
* lead must {@link #open} a fresh request.
|
||||
* The old pane was confirmed gone, but {@code LeadLauncher#relaunch} returned {@code null}
|
||||
* — every launch attempt failed. {@code bootstrapText} was never sent, and no fresh terminal
|
||||
* exists for this roll to have recognised.
|
||||
*/
|
||||
RELAUNCH_FAILED,
|
||||
/**
|
||||
* A fresh lead was launched, but its pane never reached a real turn boundary ({@code IDLE}
|
||||
* or {@code DONE}, never merely {@code BLOCKED}) within {@code relaunchReadySeconds} — the
|
||||
* CLI never finished booting, or it stayed paused on a startup prompt. {@code bootstrapText}
|
||||
* was never sent: typing into a pane that is not actually ready to accept input loses the
|
||||
* keystrokes.
|
||||
*/
|
||||
RELAUNCH_NEVER_READY,
|
||||
/**
|
||||
* A fresh lead was launched and its pane reached a real turn boundary, so {@code
|
||||
* bootstrapText} WAS sent to it, but the terminal was never recognised as a live lead —
|
||||
* present in the live-lead terminal map — within {@code relaunchReadySeconds}. The session
|
||||
* itself is alive and bootstrapped; only the daemon's own bookkeeping has not caught up, and
|
||||
* an operator should check why the tab was not recognised.
|
||||
*/
|
||||
RELAUNCH_NOT_RECOGNISED,
|
||||
/**
|
||||
* The deferred continuation threw a {@link RuntimeException} and the continuation thread
|
||||
* died with it. Without this state, that throw would leave {@link #outcomes} holding {@link
|
||||
* #IN_PROGRESS} forever, because the production {@code continuationRunner} is a bare virtual
|
||||
* thread with no uncaught-exception handler and nothing downstream of the throw ever runs to
|
||||
* write a terminal outcome. {@code detail} names the exception, so a reader has something to
|
||||
* act on. The roll is dead at this point and does not retry itself; a stuck lead must
|
||||
* {@link #open} a fresh request.
|
||||
*/
|
||||
FAILED,
|
||||
/**
|
||||
@@ -286,6 +325,10 @@ public final class LeadRollover {
|
||||
public record RollStatus(RollState state, String detail) {}
|
||||
|
||||
private final AgentControl agents;
|
||||
/** Workspace/tab/pane control — used to tear down the old pane and confirm it is gone. */
|
||||
private final WorkspaceControl spaces;
|
||||
/** Starts the fresh lead that replaces the one this roll tears down. */
|
||||
private final LeadLauncher launcher;
|
||||
private final Supplier<FleetConfig.LeadRollover> configSupplier;
|
||||
/**
|
||||
* Terminal id → that lead's configured workspace directory (their {@code
|
||||
@@ -295,8 +338,21 @@ public final class LeadRollover {
|
||||
* daemon-cwd bug this parameter exists to fix.
|
||||
*/
|
||||
private final Function<String, String> leadWorkspace;
|
||||
/**
|
||||
* Terminal id → that lead's configured name under {@code fleet.leaders}, or {@code null} when
|
||||
* the terminal names no currently-recognised lead. The deferred continuation calls this, on the
|
||||
* OLD terminal, before tearing it down, so it knows which lead to pass to {@link
|
||||
* LeadLauncher#relaunch}.
|
||||
*/
|
||||
private final Function<String, String> leadNameForTerminal;
|
||||
/**
|
||||
* The daemon's current terminal id → lead name map, read fresh on every poll. The deferred
|
||||
* continuation polls this for the FRESH terminal {@link LeadLauncher#relaunch} returns, to
|
||||
* learn when that terminal has been recognised as a live lead — see this class's javadoc.
|
||||
*/
|
||||
private final Supplier<Map<String, String>> liveLeadTerminals;
|
||||
private final LongSupplier nowMillis;
|
||||
private final Runnable settleSleeper;
|
||||
private final Runnable pollSleeper;
|
||||
/**
|
||||
* Launches the post-{@code confirm()} continuation. Production uses a single unstarted virtual
|
||||
* thread per confirmed request — see this class's javadoc for why that is a single-shot task,
|
||||
@@ -305,6 +361,15 @@ public final class LeadRollover {
|
||||
*/
|
||||
private final Consumer<Runnable> continuationRunner;
|
||||
private final Map<String, PendingRollover> pending = new ConcurrentHashMap<>();
|
||||
/**
|
||||
* Lead terminal → the token of the roll currently holding that terminal exclusive, for
|
||||
* {@link #confirm}'s single-flight claim. {@link #confirm} claims an entry here with an
|
||||
* atomic put-if-absent once every other gate has passed, refusing with {@link
|
||||
* RefusalReason#ROLL_ALREADY_RUNNING} when a claim is already held; {@link #runRollover}
|
||||
* releases it in a {@code finally}, on both the success and the thrown-exception path. A
|
||||
* terminal absent from this map has no roll currently in flight for it.
|
||||
*/
|
||||
private final Map<String, String> rollingByTerminal = new ConcurrentHashMap<>();
|
||||
/**
|
||||
* Finished tokens → what actually happened, for {@link #status}. Bounded by {@link
|
||||
* #OUTCOME_HISTORY_CAP}, oldest evicted first ({@code removeEldestEntry} on an insertion-order
|
||||
@@ -323,29 +388,40 @@ public final class LeadRollover {
|
||||
}
|
||||
});
|
||||
|
||||
/** Production constructor — wall clock, real sleep between settle polls, a real virtual thread. */
|
||||
public LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
|
||||
Function<String, String> leadWorkspace) {
|
||||
this(agents, configSupplier, leadWorkspace, System::currentTimeMillis,
|
||||
() -> sleepUninterruptibly(SETTLE_POLL_MS),
|
||||
/** Production constructor — wall clock, real sleep between polls, a real virtual thread. */
|
||||
public LeadRollover(AgentControl agents, WorkspaceControl spaces, LeadLauncher launcher,
|
||||
Supplier<FleetConfig.LeadRollover> configSupplier,
|
||||
Function<String, String> leadWorkspace,
|
||||
Function<String, String> leadNameForTerminal,
|
||||
Supplier<Map<String, String>> liveLeadTerminals) {
|
||||
this(agents, spaces, launcher, configSupplier, leadWorkspace, leadNameForTerminal,
|
||||
liveLeadTerminals, System::currentTimeMillis,
|
||||
() -> sleepUninterruptibly(POLL_INTERVAL_MS),
|
||||
r -> Thread.ofVirtual().name("lead-rollover-continuation-").start(r));
|
||||
}
|
||||
|
||||
/**
|
||||
* Full constructor — an injectable wall-clock supplier, settle-poll sleeper, and continuation
|
||||
* runner, for tests. {@code nowMillis} MUST be a wall-clock source (e.g. {@code
|
||||
* Full constructor — an injectable wall-clock supplier, poll sleeper, and continuation runner,
|
||||
* for tests. {@code nowMillis} MUST be a wall-clock source (e.g. {@code
|
||||
* System.currentTimeMillis()}), never {@code System.nanoTime()}: the freshness check compares
|
||||
* against a file's modified time, which only a wall clock is comparable to, and {@code
|
||||
* nanoTime} freezes while the host sleeps (fleetd #386).
|
||||
* nanoTime} freezes while the host sleeps.
|
||||
*/
|
||||
LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
|
||||
Function<String, String> leadWorkspace, LongSupplier nowMillis,
|
||||
Runnable settleSleeper, Consumer<Runnable> continuationRunner) {
|
||||
LeadRollover(AgentControl agents, WorkspaceControl spaces, LeadLauncher launcher,
|
||||
Supplier<FleetConfig.LeadRollover> configSupplier,
|
||||
Function<String, String> leadWorkspace,
|
||||
Function<String, String> leadNameForTerminal,
|
||||
Supplier<Map<String, String>> liveLeadTerminals,
|
||||
LongSupplier nowMillis, Runnable pollSleeper, Consumer<Runnable> continuationRunner) {
|
||||
this.agents = agents;
|
||||
this.spaces = spaces;
|
||||
this.launcher = launcher;
|
||||
this.configSupplier = configSupplier;
|
||||
this.leadWorkspace = leadWorkspace;
|
||||
this.leadNameForTerminal = leadNameForTerminal;
|
||||
this.liveLeadTerminals = liveLeadTerminals;
|
||||
this.nowMillis = nowMillis;
|
||||
this.settleSleeper = settleSleeper;
|
||||
this.pollSleeper = pollSleeper;
|
||||
this.continuationRunner = continuationRunner;
|
||||
}
|
||||
|
||||
@@ -487,6 +563,16 @@ public final class LeadRollover {
|
||||
return docCheck;
|
||||
}
|
||||
|
||||
// Single-flight claim: atomic put-if-absent, taken only after every other gate has
|
||||
// passed, so a refused confirm() never takes it. A non-null previous value means a
|
||||
// different, still-running roll already holds this lead terminal.
|
||||
String holder = rollingByTerminal.putIfAbsent(p.leadTerminal(), token);
|
||||
if (holder != null) {
|
||||
return RollDecision.refused(RefusalReason.ROLL_ALREADY_RUNNING,
|
||||
"lead terminal " + p.leadTerminal() + " already has a roll running under token "
|
||||
+ holder);
|
||||
}
|
||||
|
||||
// Record IN_PROGRESS BEFORE removing from `pending` — see RollState#IN_PROGRESS and
|
||||
// OUTCOME_HISTORY_CAP's javadoc. This ordering means `token` is written into `outcomes`
|
||||
// while it is STILL present in `pending`; status() checks `outcomes` first (see that
|
||||
@@ -494,34 +580,48 @@ public final class LeadRollover {
|
||||
// remove-then-put ordering would leave in which the token is in neither map.
|
||||
outcomes.put(token, new RollStatus(RollState.IN_PROGRESS,
|
||||
"confirm() approved this roll and handed it to the deferred continuation; it has "
|
||||
+ "not finished yet — still waiting for the calling turn to settle, for "
|
||||
+ "/clear to be sent and settle, or for bootstrapText to be sent"));
|
||||
+ "not finished yet — still waiting for the calling turn to settle, for the "
|
||||
+ "old pane to be torn down and confirmed gone, for the fresh lead to be "
|
||||
+ "recognised, or for bootstrapText to be sent"));
|
||||
pending.remove(token);
|
||||
log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends",
|
||||
token, callerTerminal);
|
||||
continuationRunner.accept(() -> runRollover(p, cfg));
|
||||
try {
|
||||
continuationRunner.accept(() -> runRollover(p, cfg));
|
||||
} catch (RuntimeException e) {
|
||||
// continuationRunner can reject the hand-off itself (e.g. a bounded executor's
|
||||
// RejectedExecutionException) before runRollover ever starts, so runRollover's own
|
||||
// finally — the only other place that releases rollingByTerminal — never runs either.
|
||||
// Release the claim here and overwrite the IN_PROGRESS entry with a terminal outcome,
|
||||
// or this lead terminal could never be rolled again and status() would report
|
||||
// IN_PROGRESS forever for a roll that in fact never started.
|
||||
log.warn("lead-rollover: continuationRunner rejected token={} lead={}: {} — the roll "
|
||||
+ "never started; releasing its claim and reporting it as FAILED",
|
||||
token, callerTerminal, e.toString(), e);
|
||||
rollingByTerminal.remove(p.leadTerminal(), token);
|
||||
outcomes.put(token, new RollStatus(RollState.FAILED,
|
||||
"continuationRunner rejected this roll before it ever started: " + e.toString()
|
||||
+ " — the roll never ran; open() a fresh rollover request"));
|
||||
}
|
||||
return RollDecision.approved();
|
||||
}
|
||||
|
||||
/**
|
||||
* The single-shot continuation {@link #confirm} hands to {@code continuationRunner}. Runs
|
||||
* entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the
|
||||
* four-step order. There is no result to return to by this point, so every outcome is logged
|
||||
* only.
|
||||
* full order. There is no result to return to by this point, so every outcome is logged only.
|
||||
*
|
||||
* <p><strong>fleetd #615 — the whole body is wrapped in one {@code try}.</strong> The two {@code
|
||||
* agents.send} calls below are not wrapped individually: {@code send} → {@code agentCall} →
|
||||
* {@code herdr.call} can throw an unchecked {@link dev.ltms.fleet.herdr.HerdrException} (see
|
||||
* {@code AgentControl.java}), and the production {@code continuationRunner} is a bare virtual
|
||||
* thread with no uncaught-exception handler (see this class's public constructor). Before this
|
||||
* fix, either throw killed the continuation thread silently, leaving the {@link
|
||||
* RollState#IN_PROGRESS} entry {@link #confirm} wrote at hand-off stuck forever — {@link
|
||||
* #status} had no way to tell a dead roll from one still genuinely running. The {@code catch}
|
||||
* below is scoped to the method body rather than to each {@code send} call individually, so it
|
||||
* also covers anything else added to this continuation later, not just today's two call sites —
|
||||
* the same reasoning that put the write-a-terminal-outcome step at each of this method's other
|
||||
* exits (see the {@link RollState#TURN_NEVER_SETTLED} and {@link RollState#CLEAR_NEVER_SETTLED}
|
||||
* branches below) rather than inside the helpers that detect them.</p>
|
||||
* <p><strong>The whole body is wrapped in one {@code try}.</strong> Several calls below —
|
||||
* {@code agents.get}, {@code agents.close}, {@code agents.send} — can throw an unchecked {@link
|
||||
* dev.ltms.fleet.herdr.HerdrException} (see {@code AgentControl.java}), and the production
|
||||
* {@code continuationRunner} is a bare virtual thread with no uncaught-exception handler (see
|
||||
* this class's public constructor). An uncaught throw would kill the continuation thread
|
||||
* silently, leaving the {@link RollState#IN_PROGRESS} entry {@link #confirm} wrote at hand-off
|
||||
* stuck forever — {@link #status} would have no way to tell a dead roll from one still
|
||||
* genuinely running. The {@code catch} below is scoped to the method body rather than to each
|
||||
* call individually, so it also covers every call in this continuation, not a fixed list of
|
||||
* call sites — the same reasoning that put the write-a-terminal-outcome step at each of this
|
||||
* method's other exits rather than inside the helpers that detect them.</p>
|
||||
*
|
||||
* <p>Only {@link RuntimeException} is caught, matching the local convention {@link
|
||||
* #waitUntilAtTurnBoundary} already set around its own {@code agents.status} call — not the
|
||||
@@ -539,6 +639,12 @@ public final class LeadRollover {
|
||||
"the roll's continuation threw " + e.toString() + " — the roll is dead and will "
|
||||
+ "not retry itself; check the daemon log for the stack trace, then open() "
|
||||
+ "a fresh rollover request"));
|
||||
} finally {
|
||||
// Release the single-flight claim on both the normal return and the thrown-exception
|
||||
// path above — a release only on success would leave this lead terminal unrollable
|
||||
// forever after one failure. The conditional two-argument remove only clears the
|
||||
// entry this roll itself holds, never a different roll's claim on the same terminal.
|
||||
rollingByTerminal.remove(p.leadTerminal(), p.token());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -546,53 +652,237 @@ public final class LeadRollover {
|
||||
private void runRolloverUnguarded(PendingRollover p, FleetConfig.LeadRollover cfg) {
|
||||
String lead = p.leadTerminal();
|
||||
long rollStartMillis = nowMillis.getAsLong();
|
||||
|
||||
TurnSettleResult turnResult = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds());
|
||||
if (!turnResult.settled()) {
|
||||
// fleetd #494 follow-up: this line had the SAME defect as the /clear-timeout line below
|
||||
// — cfg.turnSettleSeconds() is the CONFIGURED budget, not how long this wait actually
|
||||
// ran. Print the measured elapsed time alongside it, labelled, exactly like the /clear
|
||||
// path already does.
|
||||
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
|
||||
+ "confirm() — refusing to send /clear at all; the calling lead's own "
|
||||
+ "turn is still live and clearing it now would destroy live context "
|
||||
+ "confirm() — the old pane is never touched; the calling lead's own "
|
||||
+ "turn is still live and tearing it down now would destroy live context "
|
||||
+ "(token={}, configured={}s elapsed={}ms)",
|
||||
lead, p.token(), cfg.turnSettleSeconds(), turnResult.elapsedMillis());
|
||||
outcomes.put(p.token(), new RollStatus(RollState.TURN_NEVER_SETTLED,
|
||||
"the calling lead's own turn never reached a boundary (IDLE or DONE) within "
|
||||
+ "turnSettleSeconds=" + cfg.turnSettleSeconds() + "s (measured elapsed="
|
||||
+ turnResult.elapsedMillis() + "ms) — no /clear was ever sent. If this "
|
||||
+ "keeps happening, raise turnSettleSeconds in fleetd.yaml"));
|
||||
+ turnResult.elapsedMillis() + "ms) — the old pane was never touched. If "
|
||||
+ "this keeps happening, raise turnSettleSeconds in fleetd.yaml"));
|
||||
return;
|
||||
}
|
||||
|
||||
// This deliberately bypasses Injector, exactly like ClaudeCodeLauncher#clearContext:
|
||||
// /clear is housekeeping, not a delegated turn, and routing it through Injector wedges the
|
||||
// pane forever (see this class's javadoc).
|
||||
agents.send(lead, "/clear");
|
||||
ClearSettleResult clearResult = waitForClearPickupAndSettle(lead, cfg.clearSettleSeconds());
|
||||
if (!clearResult.settled()) {
|
||||
// fleetd #494: cfg.clearSettleSeconds() is the CONFIGURED budget, not how long the wait
|
||||
// actually ran — an operator reading only that number wrongly believes it is a measured
|
||||
// duration. Print the measured elapsed time and nudge count alongside it, each labelled,
|
||||
// so the two can be compared at a glance.
|
||||
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
|
||||
+ "/clear — NOT sending bootstrapText (token={}, configured={}s "
|
||||
+ "elapsed={}ms nudges={})",
|
||||
lead, p.token(), cfg.clearSettleSeconds(), clearResult.elapsedMillis(),
|
||||
clearResult.nudges());
|
||||
outcomes.put(p.token(), new RollStatus(RollState.CLEAR_NEVER_SETTLED,
|
||||
"/clear was sent, but the pane never re-settled within clearSettleSeconds="
|
||||
+ cfg.clearSettleSeconds() + "s (measured elapsed=" + clearResult.elapsedMillis()
|
||||
+ "ms, nudges=" + clearResult.nudges() + ") — bootstrapText was never sent"));
|
||||
// Captured once, here, and never re-resolved from `lead` again below: once the pane is
|
||||
// closed there is nothing left for a terminal lookup to find.
|
||||
Agent oldAgent = captureAgentWithRetry(lead);
|
||||
String oldPaneId = oldAgent.paneId();
|
||||
String leadName = leadNameForTerminal.apply(lead);
|
||||
|
||||
endOldSession(oldPaneId);
|
||||
DeathResult deathResult = waitUntilPaneGone(oldPaneId);
|
||||
if (!deathResult.gone()) {
|
||||
log.warn("lead-rollover: old pane {} for lead {} was never confirmed gone after being "
|
||||
+ "closed — not attempting a relaunch (token={}, timeout={}s "
|
||||
+ "elapsed={}ms)",
|
||||
oldPaneId, lead, p.token(), PANE_DEATH_TIMEOUT_SECONDS, deathResult.elapsedMillis());
|
||||
outcomes.put(p.token(), new RollStatus(RollState.OLD_PANE_NEVER_DIED,
|
||||
"the old pane was closed, but locatePane kept reporting it as still present "
|
||||
+ "after a pane-death timeout=" + PANE_DEATH_TIMEOUT_SECONDS
|
||||
+ "s (measured elapsed=" + deathResult.elapsedMillis() + "ms) — no "
|
||||
+ "relaunch was attempted"));
|
||||
return;
|
||||
}
|
||||
agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()));
|
||||
|
||||
Agent newAgent = launcher.relaunch(leadName);
|
||||
if (newAgent == null) {
|
||||
log.warn("lead-rollover: relaunch of lead '{}' (old terminal {}) failed every attempt "
|
||||
+ "— bootstrapText was never sent (token={})", leadName, lead, p.token());
|
||||
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_FAILED,
|
||||
"lead '" + leadName + "' could not be relaunched — every attempt failed; "
|
||||
+ "bootstrapText was never sent"));
|
||||
return;
|
||||
}
|
||||
|
||||
ReadinessResult readinessResult = waitUntilPaneReady(newAgent.terminalId(),
|
||||
cfg.relaunchReadySeconds());
|
||||
if (!readinessResult.ready()) {
|
||||
log.warn("lead-rollover: fresh pane for lead '{}' (terminal {}) never reached a real "
|
||||
+ "turn boundary — bootstrapText was never sent (token={}, configured={}s "
|
||||
+ "elapsed={}ms)",
|
||||
leadName, newAgent.terminalId(), p.token(), cfg.relaunchReadySeconds(),
|
||||
readinessResult.elapsedMillis());
|
||||
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_NEVER_READY,
|
||||
"fresh terminal " + newAgent.terminalId() + " never reached a real turn "
|
||||
+ "boundary (IDLE or DONE) within relaunchReadySeconds="
|
||||
+ cfg.relaunchReadySeconds() + "s (measured elapsed="
|
||||
+ readinessResult.elapsedMillis() + "ms) — bootstrapText was never "
|
||||
+ "sent"));
|
||||
return;
|
||||
}
|
||||
|
||||
IdentityResult identityResult = waitUntilRecognisedAsLead(newAgent.terminalId(),
|
||||
cfg.relaunchReadySeconds());
|
||||
agents.send(newAgent.terminalId(), cfg.bootstrapTextFor(p.handoverPath()));
|
||||
if (!identityResult.ready()) {
|
||||
log.warn("lead-rollover: fresh terminal {} for lead '{}' is alive and bootstrapped, but "
|
||||
+ "was never recognised as a live lead — an operator should check why "
|
||||
+ "the tab was not recognised (token={}, configured={}s elapsed={}ms)",
|
||||
newAgent.terminalId(), leadName, p.token(), cfg.relaunchReadySeconds(),
|
||||
identityResult.elapsedMillis());
|
||||
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_NOT_RECOGNISED,
|
||||
"bootstrapText was sent to fresh terminal " + newAgent.terminalId() + ", but "
|
||||
+ "it was never recognised as a live lead within relaunchReadySeconds="
|
||||
+ cfg.relaunchReadySeconds() + "s (measured elapsed="
|
||||
+ identityResult.elapsedMillis() + "ms) — check why the tab was not "
|
||||
+ "recognised"));
|
||||
return;
|
||||
}
|
||||
|
||||
long rollElapsedMillis = nowMillis.getAsLong() - rollStartMillis;
|
||||
log.info("lead-rollover: rolled token={} lead={} elapsedMs={}", p.token(), lead, rollElapsedMillis);
|
||||
log.info("lead-rollover: rolled token={} oldLead={} newTerminal={} elapsedMs={}",
|
||||
p.token(), lead, newAgent.terminalId(), rollElapsedMillis);
|
||||
outcomes.put(p.token(), new RollStatus(RollState.ROLLED,
|
||||
"rolled successfully in " + rollElapsedMillis + "ms"));
|
||||
"rolled successfully in " + rollElapsedMillis + "ms; new terminal="
|
||||
+ newAgent.terminalId()));
|
||||
}
|
||||
|
||||
/** Attempts {@link #captureAgentWithRetry} makes before letting the failure propagate. */
|
||||
static final int CAPTURE_RETRIES = 3;
|
||||
|
||||
/**
|
||||
* {@link AgentControl#get} for {@code lead}, retried up to {@link #CAPTURE_RETRIES} times. The
|
||||
* terminal-to-pane lookup it goes through can report a genuinely live agent as not found (see
|
||||
* {@code AgentControl#agentCall}'s own re-resolve-once behaviour), and one such false negative
|
||||
* must not abort an otherwise-healthy roll. The result is captured once by the caller and never
|
||||
* looked up again — see this class's javadoc.
|
||||
*
|
||||
* @throws RuntimeException the last failure, if every attempt fails — {@link #runRollover}'s
|
||||
* catch turns that into {@link RollState#FAILED}
|
||||
*/
|
||||
private Agent captureAgentWithRetry(String lead) {
|
||||
RuntimeException last = null;
|
||||
for (int attempt = 1; attempt <= CAPTURE_RETRIES; attempt++) {
|
||||
try {
|
||||
return agents.get(lead);
|
||||
} catch (RuntimeException e) {
|
||||
last = e;
|
||||
log.debug("lead-rollover: agents.get({}) failed on attempt {}/{}: {}",
|
||||
lead, attempt, CAPTURE_RETRIES, e.toString());
|
||||
if (attempt < CAPTURE_RETRIES) {
|
||||
pollSleeper.run();
|
||||
}
|
||||
}
|
||||
}
|
||||
throw last;
|
||||
}
|
||||
|
||||
/**
|
||||
* End the old lead's session: close its pane, then close its tab only when the pane was that
|
||||
* tab's sole occupant — the same pane-then-tab teardown {@code HerdrPeerLauncher#stop} uses for
|
||||
* a member. An already-gone pane counts as success; any other {@code agents.close} failure
|
||||
* propagates, so a genuinely failed teardown is never reported as done. A failing
|
||||
* {@code spaces.closeTab} never propagates — by the time it runs the pane is already closed, so
|
||||
* it is cosmetic tidying, not a real teardown failure.
|
||||
*/
|
||||
private void endOldSession(String paneId) {
|
||||
WorkspaceControl.PaneLocation loc = spaces.locatePane(paneId);
|
||||
try {
|
||||
agents.close(paneId);
|
||||
} catch (HerdrException e) {
|
||||
if (!isAlreadyGone(e)) {
|
||||
throw e;
|
||||
}
|
||||
log.debug("lead-rollover: pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
|
||||
}
|
||||
if (loc != null && loc.tabPaneCount() == 1) {
|
||||
try {
|
||||
spaces.closeTab(loc.tabId());
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("lead-rollover: tab.close({}) failed — the pane is already torn down, so "
|
||||
+ "continuing; the tab may need manual cleanup: {}", loc.tabId(), e.getMessage());
|
||||
}
|
||||
} else if (loc != null) {
|
||||
log.debug("lead-rollover: not closing tab {} — it holds {} panes (not a dedicated lead "
|
||||
+ "tab)", loc.tabId(), loc.tabPaneCount());
|
||||
}
|
||||
}
|
||||
|
||||
/** True when a herdr error means the target is already gone (safe to treat as done). */
|
||||
private static boolean isAlreadyGone(HerdrException e) {
|
||||
return e.code() != null && e.code().endsWith("_not_found");
|
||||
}
|
||||
|
||||
/**
|
||||
* Poll {@link WorkspaceControl#locatePane} for {@code paneId} until it reports {@code null}
|
||||
* (the pane is gone) or {@link #PANE_DEATH_TIMEOUT_SECONDS} elapses. Deliberately never calls
|
||||
* {@link AgentControl#status} and never reads the live-lead terminal map — both answer a
|
||||
* different question (whether an AGENT is live, not whether this PANE still exists) and
|
||||
* {@code locatePane} alone catches a {@link HerdrException} from the underlying {@code
|
||||
* pane.get} and turns it into {@code null} — see this class's javadoc.
|
||||
*/
|
||||
private DeathResult waitUntilPaneGone(String paneId) {
|
||||
long startMillis = nowMillis.getAsLong();
|
||||
long deadline = startMillis + TimeUnit.SECONDS.toMillis(PANE_DEATH_TIMEOUT_SECONDS);
|
||||
while (nowMillis.getAsLong() < deadline) {
|
||||
if (spaces.locatePane(paneId) == null) {
|
||||
return new DeathResult(true, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
pollSleeper.run();
|
||||
}
|
||||
return new DeathResult(false, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
|
||||
/** The measured outcome of {@link #waitUntilPaneGone}. */
|
||||
private record DeathResult(boolean gone, long elapsedMillis) {}
|
||||
|
||||
/**
|
||||
* Poll until {@code newTerminal}'s own pane reaches a real turn boundary ({@link
|
||||
* AgentStatus#IDLE} or {@link AgentStatus#DONE}, never merely {@link AgentStatus#BLOCKED}) —
|
||||
* the same exclusion {@link #waitUntilAtTurnBoundary} applies to the calling lead's own turn,
|
||||
* applied here to the fresh one, so {@code bootstrapText} is never typed into a pane that has
|
||||
* not actually finished booting — or {@code readySeconds} elapses. A failed status read
|
||||
* degrades to "not yet ready" and is retried on the next poll.
|
||||
*/
|
||||
private ReadinessResult waitUntilPaneReady(String newTerminal, int readySeconds) {
|
||||
long startMillis = nowMillis.getAsLong();
|
||||
long deadline = startMillis + TimeUnit.SECONDS.toMillis(readySeconds);
|
||||
while (nowMillis.getAsLong() < deadline) {
|
||||
AgentStatus status;
|
||||
try {
|
||||
status = agents.status(newTerminal);
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("lead-rollover: status check failed while waiting for {} to be ready: {}",
|
||||
newTerminal, e.toString());
|
||||
status = null;
|
||||
}
|
||||
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
|
||||
return new ReadinessResult(true, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
pollSleeper.run();
|
||||
}
|
||||
return new ReadinessResult(false, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
|
||||
/** The measured outcome of {@link #waitUntilPaneReady}. */
|
||||
private record ReadinessResult(boolean ready, long elapsedMillis) {}
|
||||
|
||||
/**
|
||||
* Poll until {@code newTerminal} is present in {@link #liveLeadTerminals} or {@code
|
||||
* readySeconds} elapses. This is bookkeeping, not a safety gate: the pane's own readiness (see
|
||||
* {@link #waitUntilPaneReady}) is what decides whether {@code bootstrapText} is safe to send —
|
||||
* a timeout here only means the daemon's own lead-discovery scan has not caught up yet.
|
||||
*/
|
||||
private IdentityResult waitUntilRecognisedAsLead(String newTerminal, int readySeconds) {
|
||||
long startMillis = nowMillis.getAsLong();
|
||||
long deadline = startMillis + TimeUnit.SECONDS.toMillis(readySeconds);
|
||||
while (nowMillis.getAsLong() < deadline) {
|
||||
if (liveLeadTerminals.get().containsKey(newTerminal)) {
|
||||
return new IdentityResult(true, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
pollSleeper.run();
|
||||
}
|
||||
return new IdentityResult(false, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
|
||||
/** The measured outcome of {@link #waitUntilRecognisedAsLead}. */
|
||||
private record IdentityResult(boolean ready, long elapsedMillis) {}
|
||||
|
||||
/** Drop a pending request without rolling. @return whether a pending request existed for {@code token} */
|
||||
public boolean cancel(String token) {
|
||||
return pending.remove(token) != null;
|
||||
@@ -682,15 +972,12 @@ public final class LeadRollover {
|
||||
|
||||
/**
|
||||
* Poll {@link AgentControl#status} until {@code target} reports a real turn boundary — {@link
|
||||
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used once by
|
||||
* {@link #runRollover}, to wait for the CALLING turn's own pane to settle before {@code /clear}
|
||||
* is ever sent at all — the {@code turnSettleSeconds} gate that makes this correction safe. The
|
||||
* SECOND wait, after {@code /clear}, is {@link #waitForClearPickupAndSettle} instead (fleetd
|
||||
* #489) — a plain boundary check is not enough there, because {@code /clear} starts no turn of
|
||||
* its own, so this method would (wrongly) report "settled" on its very first poll whether or not
|
||||
* {@code /clear} was actually picked up. A failed status read degrades to "not yet settled" and
|
||||
* is retried on the next poll, the same posture {@code LeadHeartbeatLoop} and {@code
|
||||
* HerdrPeerLauncher}'s readiness gate already take toward an unreadable status.
|
||||
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used by
|
||||
* {@link #runRollover} to wait for the CALLING turn's own pane to settle before the old pane is
|
||||
* touched at all — the {@code turnSettleSeconds} gate that makes tearing it down safe. A failed
|
||||
* status read degrades to "not yet settled" and is retried on the next poll, the same posture
|
||||
* {@code LeadHeartbeatLoop} and {@code HerdrPeerLauncher}'s readiness gate already take toward
|
||||
* an unreadable status.
|
||||
*
|
||||
* <p><strong>Deliberately not {@link AgentStatus#injectable()}.</strong> {@code injectable()}
|
||||
* answers the {@code Injector}'s question — "may I deliver a message without stepping on a live
|
||||
@@ -698,16 +985,16 @@ public final class LeadRollover {
|
||||
* an approval prompt is safe to queue a message behind. This class asks a stricter question —
|
||||
* "has the turn actually ended" — and {@code BLOCKED} answers no: it is a live turn that is
|
||||
* merely paused, not one that has finished. Reusing {@code injectable()} here would let this
|
||||
* wait fire {@code /clear} while the lead's own {@code confirm()}-calling turn is still live and
|
||||
* paused on a prompt — exactly the live-context-destroying failure the {@code turnSettleSeconds}
|
||||
* gate exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link
|
||||
* #waitForClearPickupAndSettle} keeps the same exclusion of {@code BLOCKED}, for the same
|
||||
* reason, on the second wait.)
|
||||
* wait tear the old pane down while the lead's own {@code confirm()}-calling turn is still live
|
||||
* and paused on a prompt — exactly the live-context-destroying failure {@code turnSettleSeconds}
|
||||
* exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link
|
||||
* #waitUntilPaneReady} applies the same exclusion of {@code BLOCKED} to the fresh lead's own
|
||||
* turn.)
|
||||
*
|
||||
* @return a {@link TurnSettleResult} whose {@code settled()} is {@code true} once a real
|
||||
* boundary was observed, {@code false} if {@code settleSeconds} elapses first.
|
||||
* {@code elapsedMillis()} is a MEASURED value from the injected {@link #nowMillis}
|
||||
* clock, never the configured {@code settleSeconds} budget (fleetd #494 follow-up).
|
||||
* clock, never the configured {@code settleSeconds} budget.
|
||||
*/
|
||||
private TurnSettleResult waitUntilAtTurnBoundary(String target, int settleSeconds) {
|
||||
long startMillis = nowMillis.getAsLong();
|
||||
@@ -724,142 +1011,11 @@ public final class LeadRollover {
|
||||
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
|
||||
return new TurnSettleResult(true, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
settleSleeper.run();
|
||||
pollSleeper.run();
|
||||
}
|
||||
return new TurnSettleResult(false, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
|
||||
/**
|
||||
* The measured outcome of {@link #waitUntilAtTurnBoundary} — fleetd #494 follow-up. The sibling
|
||||
* of {@link ClearSettleResult} for the FIRST wait, which never nudges, so it carries no nudge
|
||||
* count.
|
||||
*/
|
||||
/** The measured outcome of {@link #waitUntilAtTurnBoundary}. */
|
||||
private record TurnSettleResult(boolean settled, long elapsedMillis) {}
|
||||
|
||||
/**
|
||||
* The SECOND wait in {@link #runRollover} — after {@code /clear} has been sent, waits for it to
|
||||
* settle, bounded by {@code settleSeconds}. <strong>fleetd #489 — the paste-race fix.</strong>
|
||||
* {@code /clear} does not start a real turn of its own, so a pane with no submit race simply
|
||||
* stays {@link AgentStatus#IDLE} the whole time: {@link #waitUntilAtTurnBoundary} would (wrongly)
|
||||
* call that "settled" on its very first poll, whether or not the {@code /clear} Enter actually
|
||||
* landed. That was Fault 1, measured live on 2026-09-12 — the second gate was a no-op, so a
|
||||
* {@code bootstrapText} send followed immediately, racing Fault 2: {@link AgentControl#submit}'s
|
||||
* own javadoc already records that the submit accompanying a delivery "can race the paste —
|
||||
* especially right as the worker's TUI becomes interactive — leaving the text unsubmitted"
|
||||
* (CB-113). Because {@code runRollover} deliberately bypasses {@code Injector} for {@code
|
||||
* /clear} (see this class's javadoc), it inherited none of {@code Injector}'s nudging — so the
|
||||
* lost {@code /clear} Enter sat in the input box and {@code bootstrapText} was typed right after
|
||||
* it, landing as one concatenated line.
|
||||
*
|
||||
* <p>This method copies the pickup-nudge pattern {@link dev.ltms.fleet.inject.Injector} already
|
||||
* ships for exactly this, on its own post-turn {@code /clear} housekeeping (fleetd #306; see
|
||||
* {@code Injector.java:288-340} and {@code Injector.java:437-442}):
|
||||
* <ul>
|
||||
* <li>an {@link AgentStatus#WORKING} sample means {@code /clear} was picked up as a real
|
||||
* turn;</li>
|
||||
* <li>until that happens, each poll that still reports {@link AgentStatus#IDLE} or {@link
|
||||
* AgentStatus#DONE} re-sends the submit keystroke ({@link AgentControl#submit}) to nudge
|
||||
* the raced Enter — for the first {@code PICKUP_GRACE_POLLS - 1} of {@link
|
||||
* #PICKUP_GRACE_POLLS} consecutive such polls (i.e. {@code PICKUP_GRACE_POLLS - 1}
|
||||
* nudges: 7, not 8, given {@code PICKUP_GRACE_POLLS = 8}). A second Enter on an empty
|
||||
* Claude Code prompt is a no-op, so repeating it is safe;</li>
|
||||
* <li>the {@code PICKUP_GRACE_POLLS}th consecutive such poll, with {@code WORKING} still never
|
||||
* observed, releases rather than wedges the roll instead of nudging again — the same
|
||||
* choice {@code Injector} makes — and returns {@code settled() == true} anyway, logged at
|
||||
* {@code warn} with the measured elapsed time (fleetd #494) so an operator can see which
|
||||
* path ran and how long it actually took;</li>
|
||||
* <li>once {@code WORKING} has been observed, nudging stops and this instead waits for a real
|
||||
* {@code working → IDLE/DONE} completion boundary before returning {@code true}.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p><strong>{@link AgentStatus#BLOCKED} is deliberately excluded from both the nudge and the
|
||||
* boundary check</strong> — the same reasoning as {@link #waitUntilAtTurnBoundary}'s own
|
||||
* javadoc: a paused live turn is not a settled one, and re-sending Enter into an open approval
|
||||
* prompt could wrongly answer it. A {@code BLOCKED} sample (or an unreadable/{@link
|
||||
* AgentStatus#UNKNOWN} one) simply keeps this polling, with no nudge and no release, until either
|
||||
* a real boundary is reached or {@code settleSeconds} runs out.
|
||||
*
|
||||
* <p>{@link AgentControl#submit} can itself throw; a {@link RuntimeException} from it is
|
||||
* swallowed and logged at {@code debug}, exactly like {@code Injector.java:437-442} — a failed
|
||||
* nudge must not abort the roll.
|
||||
*
|
||||
* @return a {@link ClearSettleResult} whose {@code settled()} is {@code true} once {@code
|
||||
* /clear} has settled, or once the nudge budget was exhausted with no pickup ever
|
||||
* observed (released rather than wedged); {@code false} if {@code settleSeconds} elapses
|
||||
* first — the caller must NOT send {@code bootstrapText} in that case, exactly as before
|
||||
* this fix. {@code elapsedMillis()} and {@code nudges()} are MEASURED values (from the
|
||||
* injected {@link #nowMillis} clock and an actual nudge count), never the configured
|
||||
* {@code settleSeconds} budget (fleetd #494).
|
||||
*/
|
||||
private ClearSettleResult waitForClearPickupAndSettle(String target, int settleSeconds) {
|
||||
long startMillis = nowMillis.getAsLong();
|
||||
long deadline = startMillis + TimeUnit.SECONDS.toMillis(settleSeconds);
|
||||
boolean pickedUp = false; // a WORKING sample has been observed since /clear was sent
|
||||
int idlePollsAwaitingPickup = 0;
|
||||
int nudges = 0;
|
||||
while (nowMillis.getAsLong() < deadline) {
|
||||
AgentStatus status;
|
||||
try {
|
||||
status = agents.status(target);
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("lead-rollover: status check failed while waiting for {} to settle after "
|
||||
+ "/clear: {}", target, e.toString());
|
||||
status = null;
|
||||
}
|
||||
if (status == AgentStatus.WORKING) {
|
||||
pickedUp = true;
|
||||
} else if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
|
||||
if (pickedUp) {
|
||||
// a real WORKING -> IDLE/DONE completion boundary
|
||||
return new ClearSettleResult(true, nowMillis.getAsLong() - startMillis, nudges);
|
||||
}
|
||||
if (++idlePollsAwaitingPickup >= PICKUP_GRACE_POLLS) {
|
||||
long elapsedMillis = nowMillis.getAsLong() - startMillis;
|
||||
// fleetd #494: this release trades a possibly-unsubmitted /clear for progress
|
||||
// instead of wedging the roll — that trade is deliberate and stays. But it is
|
||||
// also exactly the case that reported false success in the real incident (the
|
||||
// whole roll "succeeded" after 438ms of a 20s budget), so raise it to WARN and
|
||||
// print the MEASURED elapsed time next to the target pane, not just the count.
|
||||
//
|
||||
// fleetd #494 follow-up (2nd pass): BOTH numbers in this line must come from
|
||||
// the loop's own counters, never from the PICKUP_GRACE_POLLS constant.
|
||||
// `idlePollsAwaitingPickup` and `nudges` each have exactly one write site in
|
||||
// this loop, on the same branch, so on this branch they cannot differ from
|
||||
// PICKUP_GRACE_POLLS / PICKUP_GRACE_POLLS - 1 today — no test can prove the
|
||||
// difference on this line, and printing the counters does not change that.
|
||||
// What it does buy: one source of truth instead of two, so a later change to
|
||||
// the loop (an early return, a second increment site, a different exit
|
||||
// condition) cannot leave this message reporting a number the loop no longer
|
||||
// produces. The place where `nudges` genuinely varies with the run — and is
|
||||
// covered by a test that can tell it apart from a constant — is the
|
||||
// /clear-timeout warn in runRollover, which prints clearResult.nudges().
|
||||
log.warn("lead-rollover: /clear on {} was never observed as WORKING after {} "
|
||||
+ "consecutive IDLE/DONE polls ({} of those were nudged) — "
|
||||
+ "releasing rather than wedging the roll (elapsed={}ms)",
|
||||
target, idlePollsAwaitingPickup, nudges, elapsedMillis);
|
||||
return new ClearSettleResult(true, elapsedMillis, nudges);
|
||||
}
|
||||
try {
|
||||
agents.submit(target); // nudge a raced Enter (CB-113) so /clear actually submits
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("lead-rollover: resubmit to {} failed (will retry next poll): {}",
|
||||
target, e.getMessage());
|
||||
} finally {
|
||||
nudges++; // an attempted nudge, whether or not the submit call itself threw
|
||||
}
|
||||
}
|
||||
// AgentStatus.BLOCKED or UNKNOWN (or an unreadable status, above): neither a pickup
|
||||
// signal nor a boundary — keep polling without nudging or releasing.
|
||||
settleSleeper.run();
|
||||
}
|
||||
return new ClearSettleResult(false, nowMillis.getAsLong() - startMillis, nudges);
|
||||
}
|
||||
|
||||
/**
|
||||
* The measured outcome of {@link #waitForClearPickupAndSettle} — fleetd #494. Carries the
|
||||
* MEASURED elapsed time (from the injected {@link #nowMillis} clock) and nudge count alongside
|
||||
* the settle/timeout decision, so callers can log them instead of the configured budget, which
|
||||
* is not how long the wait actually ran.
|
||||
*/
|
||||
private record ClearSettleResult(boolean settled, long elapsedMillis, int nudges) {}
|
||||
}
|
||||
|
||||
@@ -114,6 +114,13 @@ public final class FleetMcp {
|
||||
* dev.ltms.fleet.herdr.PaneLocator} — that a real assembly wired up. See {@link #identity()}.
|
||||
*/
|
||||
private final ConnectionIdentity identity;
|
||||
/**
|
||||
* Kept as a field (rather than only captured by the {@code contextExtractor} closure) so
|
||||
* {@link #denyFor} can read {@link CallerResolver#knownLeadOrCollaborator()} — the classifier a
|
||||
* collaborator's {@code SEND} is checked against, built from the same lead and collaborator
|
||||
* maps {@link #identity}-based resolution reads.
|
||||
*/
|
||||
private final CallerResolver callers;
|
||||
private final Metrics metrics; // CB-502: null → auth failures not counted
|
||||
private final CapacitySource capacity;
|
||||
private final HealthCoverageSource healthCoverage;
|
||||
@@ -327,7 +334,7 @@ public final class FleetMcp {
|
||||
* caller explicitly saying so — never by omitting a {@link CallerResolver} the way the old
|
||||
* {@code callers == null} idiom allowed. {@code callers} itself is required either way: even
|
||||
* under {@link #UNENFORCED}, the one real {@link CallerResolver} still resolves every caller's
|
||||
* {@link Principal} (so {@code markSpawnedMemberPresent}/{@code recordPrimarySingleton} see a
|
||||
* {@link Principal} (so {@code markTrackedCallerPresent}/{@code recordPrimarySingleton} see a
|
||||
* real identity), and {@link #denyFor} is the only thing that changes.
|
||||
*/
|
||||
public enum AuthorizationMode { ENFORCED, UNENFORCED }
|
||||
@@ -409,6 +416,7 @@ public final class FleetMcp {
|
||||
this.authorizationEnforced = Objects.requireNonNull(authorizationMode, "authorizationMode")
|
||||
== AuthorizationMode.ENFORCED;
|
||||
this.identity = identity;
|
||||
this.callers = callers;
|
||||
this.leadChannel = leadChannel;
|
||||
this.peers = peers == null ? List.of() : List.copyOf(peers);
|
||||
this.capacity = capacity;
|
||||
@@ -434,10 +442,10 @@ public final class FleetMcp {
|
||||
// fall back to here — AuthorizationMode governs enforcement, not identity.
|
||||
Principal p = callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
|
||||
req.getHeader("Authorization"));
|
||||
// CB-532: guard on the ROLE, not on the terminal being null. This excludes a
|
||||
// lead, which carries its pane too, while including every spawned member role.
|
||||
// Enrolling a lead would count it as an available member in the roster.
|
||||
markSpawnedMemberPresent(p, presence);
|
||||
// Guards on the ROLE, not on the terminal being null — this marks presence for
|
||||
// a worker, an architect, or the unconfigured-pane floor, and excludes a lead
|
||||
// or a collaborator even though each carries its own pane too.
|
||||
markTrackedCallerPresent(p, presence);
|
||||
return McpTransportContext.create(Map.of(
|
||||
CALLER_TERMINAL, orEmpty(p.terminal()),
|
||||
CALLER_PID, Long.toString(p.pid()),
|
||||
@@ -472,7 +480,7 @@ public final class FleetMcp {
|
||||
// Answering a worker's fleet_ask (CB-205): resolve its blocked question and
|
||||
// block for the worker's reply as it resumes the same turn. This is the same
|
||||
// delegation, so ownership is left untouched (CB-548) — never re-recorded.
|
||||
return answer(messages, turnId, content, timeoutMs(a));
|
||||
return answer(messages, turnId, content, timeoutMs(a), caller);
|
||||
}
|
||||
// CB-548: delegator ownership (which lead's reply nudge this worker routes to,
|
||||
// CB-532) is recorded only once the send is ACCEPTED — MessageService has won the
|
||||
@@ -482,8 +490,8 @@ public final class FleetMcp {
|
||||
Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, caller);
|
||||
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
|
||||
return Boolean.FALSE.equals(a.get("wait"))
|
||||
? sendAsync(messages, target, content, onAccepted, workers.profiles())
|
||||
: send(messages, target, content, timeoutMs(a), onAccepted, workers.profiles());
|
||||
? sendAsync(messages, target, content, onAccepted, workers.profiles(), caller)
|
||||
: send(messages, target, content, timeoutMs(a), onAccepted, workers.profiles(), caller);
|
||||
};
|
||||
// fleet_reply's identity is the CONNECTION, never an argument — so the authz check
|
||||
// is "is this caller a worker at all", and it can only ever reply as itself.
|
||||
@@ -506,7 +514,7 @@ public final class FleetMcp {
|
||||
(exchange, req) -> {
|
||||
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_status", req.arguments()), null);
|
||||
if (denied != null) return denied;
|
||||
return status(messages, str(req.arguments(), "sessionId"));
|
||||
return status(messages, str(req.arguments(), "sessionId"), callerTerminal(exchange));
|
||||
};
|
||||
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> pollHandler =
|
||||
(exchange, req) -> {
|
||||
@@ -516,7 +524,7 @@ public final class FleetMcp {
|
||||
// The action depends on the ARGUMENTS, not on the tool name -- see pollAction.
|
||||
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_poll", a), target);
|
||||
if (denied != null) return denied;
|
||||
return poll(messages, leadChannel, str(a, "ticket"), target, coordId);
|
||||
return poll(messages, leadChannel, str(a, "ticket"), target, coordId, callerTerminal(exchange));
|
||||
};
|
||||
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
|
||||
// Acking removes a reply from the inbox, so it is a drain, not a read.
|
||||
@@ -552,8 +560,10 @@ public final class FleetMcp {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine, outage,
|
||||
leadSeats, leadContextGauge, leadConfigDirs, callers.leads(),
|
||||
callerTerminal(exchange),
|
||||
callers.collaborators(), collaboratorsVisibleTo(principal(exchange)),
|
||||
new CoordinationSource(leadChannel, peers),
|
||||
coordinatorVisibleTo(principal(exchange)));
|
||||
coordinatorVisibleTo(principal(exchange)),
|
||||
leadsVisibleTo(principal(exchange)), membersVisibleTo(principal(exchange)));
|
||||
};
|
||||
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> stopHandler =
|
||||
(exchange, req) -> {
|
||||
@@ -689,8 +699,8 @@ public final class FleetMcp {
|
||||
if (!authorizationEnforced) {
|
||||
return null; // AuthorizationMode.UNENFORCED: authorization not enforced (fleetd #518)
|
||||
}
|
||||
if (Authz.permits(caller, action, target)) {
|
||||
if (action != Authz.Action.READ) {
|
||||
if (Authz.permits(caller, action, target, callers.knownLeadOrCollaborator())) {
|
||||
if (action != Authz.Action.READ && action != Authz.Action.TASK_READ) {
|
||||
AuditLog.allowed(caller, action, target); // reads would drown the trail
|
||||
}
|
||||
return null;
|
||||
@@ -735,7 +745,53 @@ public final class FleetMcp {
|
||||
return caller.isPrimary();
|
||||
}
|
||||
|
||||
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
|
||||
/**
|
||||
* fleetd #703: who may see {@code fleet_list}'s {@code collaborators} array — a roster of the
|
||||
* human-opened tabs this daemon recognises as named peers. Visible to exactly the roles that
|
||||
* may {@link Authz.Action#SEND} to a named peer ({@code Authz.java}'s {@code SEND} case):
|
||||
* the primary, an architect, and a collaborator (reaching another collaborator or a lead). A
|
||||
* worker holds {@code READ} but can never {@code SEND} to a collaborator, so listing them to a
|
||||
* worker would expose which human tabs exist on the host with no use to that caller. Split out
|
||||
* for the same reason as {@link #coordinatorVisibleTo}: the decision must be unit-testable
|
||||
* without fabricating an SDK {@code McpSyncServerExchange}, and the handler must call this
|
||||
* named predicate rather than inlining the check.
|
||||
*/
|
||||
static boolean collaboratorsVisibleTo(Principal caller) {
|
||||
return caller.isPrimary() || caller.isArchitect() || caller.isCollaborator();
|
||||
}
|
||||
|
||||
/**
|
||||
* Who may see {@code fleet_list}'s {@code leads} array — exactly the roles that may
|
||||
* {@link Authz.Action#SEND} to a lead: the primary, an architect, and a collaborator. A
|
||||
* collaborator's own {@code fleet_whoami} carries no lead address, and {@code leads} is the
|
||||
* only place this tool gives one, so a collaborator needs this array to use the send it
|
||||
* already holds. A worker can never {@code SEND} at all, so it still sees neither this array
|
||||
* nor {@code members}; a worker's own facts come from {@code fleet_whoami} instead. Split
|
||||
* out for the same reason as {@link #coordinatorVisibleTo} and
|
||||
* {@link #collaboratorsVisibleTo}: the decision must be unit-testable without fabricating an
|
||||
* SDK {@code McpSyncServerExchange}, and the handler must call this named predicate rather
|
||||
* than inlining the check.
|
||||
*/
|
||||
static boolean leadsVisibleTo(Principal caller) {
|
||||
return caller.isPrimary() || caller.isArchitect() || caller.isCollaborator();
|
||||
}
|
||||
|
||||
/**
|
||||
* Who may see {@code fleet_list}'s {@code members} array — the primary and an architect,
|
||||
* which may {@link Authz.Action#SEND} to a member. A worker can never {@code SEND} at all,
|
||||
* and a collaborator may {@code SEND} only to a lead or another collaborator, never to a
|
||||
* spawned member, so neither sees this array even though {@link #leadsVisibleTo} grants a
|
||||
* collaborator the sibling one.
|
||||
*/
|
||||
static boolean membersVisibleTo(Principal caller) {
|
||||
return caller.isPrimary() || caller.isArchitect();
|
||||
}
|
||||
|
||||
/**
|
||||
* The terminal of the caller on this call's connection, or {@code null} when that caller carries
|
||||
* no terminal, which is only the unnamed primary. A named lead, an architect and a worker each
|
||||
* carry one.
|
||||
*/
|
||||
private static String callerTerminal(McpSyncServerExchange exchange) {
|
||||
Object v = exchange.transportContext().get(CALLER_TERMINAL);
|
||||
String s = v == null ? null : v.toString();
|
||||
@@ -777,9 +833,14 @@ public final class FleetMcp {
|
||||
return identity;
|
||||
}
|
||||
|
||||
/** Mark a connected spawned member available for the injector readiness gate. */
|
||||
static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) {
|
||||
if (caller.isSpawnedMember()) {
|
||||
/**
|
||||
* Mark a caller present for the injector readiness gate, when its deliverability depends on
|
||||
* proving a live MCP contact: a worker, an architect, or the unconfigured-pane floor. A lead
|
||||
* or a collaborator is excluded — each is already deliverable through its own named-registry
|
||||
* entry.
|
||||
*/
|
||||
static void markTrackedCallerPresent(Principal caller, MemberPresence presence) {
|
||||
if (caller.isSpawnedMember() || caller.isObserver()) {
|
||||
presence.markPresent(caller.terminal());
|
||||
}
|
||||
}
|
||||
@@ -835,8 +896,8 @@ public final class FleetMcp {
|
||||
* fleetd #612 B3 — as {@link #quarantineSource()}, {@code public} for the same cross-package
|
||||
* reason, for the real {@link LeadRollover} (or {@code null}) this daemon was assembled with.
|
||||
* {@code FleetdLeadRolloverAssemblyTest} drives {@code open}/{@code confirm} on this exact
|
||||
* instance and waits for the real continuation to send {@code /clear} and {@code bootstrapText}
|
||||
* through the real {@code router.leadAgents()}.
|
||||
* instance and waits for the real continuation to end the old pane, relaunch a fresh one, and
|
||||
* send {@code bootstrapText} through the real {@code router.leadAgents()}.
|
||||
*/
|
||||
public LeadRollover leadRollover() {
|
||||
return leadRollover;
|
||||
@@ -852,7 +913,8 @@ public final class FleetMcp {
|
||||
* a BUSY interloper never claims a turn it did not win. {@code null} disables recording.
|
||||
*/
|
||||
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content,
|
||||
Long timeoutMs, Runnable onAccepted, Set<String> profiles) {
|
||||
Long timeoutMs, Runnable onAccepted, Set<String> profiles,
|
||||
String callerTerminal) {
|
||||
if (isBlank(sessionId) || isBlank(content)) {
|
||||
return error("sessionId and content are required");
|
||||
}
|
||||
@@ -862,7 +924,7 @@ public final class FleetMcp {
|
||||
}
|
||||
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
|
||||
try {
|
||||
return formatReply(messages.send(sessionId, content, timeout, onAccepted), timeout);
|
||||
return formatReply(messages.send(sessionId, content, timeout, onAccepted, callerTerminal), timeout);
|
||||
} catch (HerdrException e) {
|
||||
return error("herdr error contacting session " + sessionId + ": " + e.getMessage());
|
||||
}
|
||||
@@ -872,13 +934,15 @@ public final class FleetMcp {
|
||||
* {@code fleet_send} carrying a {@code turnId}: the primary's answer to a worker's
|
||||
* {@code fleet_ask} (CB-205). Resolves the worker's blocked question and blocks for its reply as
|
||||
* it resumes the same turn — surfaced to the primary identically to a normal send.
|
||||
* {@code callerTerminal} must match the turn's recorded owner or this is refused.
|
||||
*/
|
||||
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs) {
|
||||
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs,
|
||||
String callerTerminal) {
|
||||
if (isBlank(turnId) || isBlank(content)) {
|
||||
return error("turnId and content are required to answer a worker's question");
|
||||
}
|
||||
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
|
||||
return formatReply(messages.answer(turnId, content, timeout), timeout);
|
||||
return formatReply(messages.answer(turnId, content, timeout, callerTerminal), timeout);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -927,6 +991,8 @@ public final class FleetMcp {
|
||||
+ "\" and content set to your answer; the worker resumes the same turn.");
|
||||
case STALE_TURN -> error("that question is no longer open — it timed out or was already "
|
||||
+ "answered (turnId stale)");
|
||||
case NOT_TURN_OWNER -> error("this turn belongs to a different delegation — only the caller "
|
||||
+ "that opened it may answer it");
|
||||
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> text("[no reply within " + timeout + "ms — worker "
|
||||
+ r.outcome().name().toLowerCase().replace("timed_out_", "") + "; retry or poll status]");
|
||||
// fleetd #571: delivery is unknown here — agent.prompt pastes and submits in one call,
|
||||
@@ -948,6 +1014,17 @@ public final class FleetMcp {
|
||||
*/
|
||||
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content,
|
||||
Runnable onAccepted, Set<String> profiles) {
|
||||
return sendAsync(messages, sessionId, content, onAccepted, profiles, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, recording {@code creatorTerminal} as this ticket's owner (the caller's own
|
||||
* terminal, resolved from the connection) so a later {@code fleet_poll{ticket}} only hands the
|
||||
* result back to that same caller — see {@link MessageService#poll(String, String)}.
|
||||
*/
|
||||
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content,
|
||||
Runnable onAccepted, Set<String> profiles,
|
||||
String creatorTerminal) {
|
||||
if (isBlank(sessionId) || isBlank(content)) {
|
||||
return error("sessionId and content are required");
|
||||
}
|
||||
@@ -955,7 +1032,7 @@ public final class FleetMcp {
|
||||
if (targetError != null) {
|
||||
return targetError;
|
||||
}
|
||||
String ticket = messages.sendAsync(sessionId, content, onAccepted);
|
||||
String ticket = messages.sendAsync(sessionId, content, onAccepted, creatorTerminal);
|
||||
return text("accepted — task delegated. Poll fleet_poll with ticket=" + ticket);
|
||||
}
|
||||
|
||||
@@ -1035,31 +1112,21 @@ public final class FleetMcp {
|
||||
}
|
||||
|
||||
/**
|
||||
* Which authorization action a {@code fleet_poll} call needs, decided by its arguments
|
||||
* (fleetd #272, widened by fleetd #421).
|
||||
* Which authorization action a {@code fleet_poll} call needs, decided by its arguments.
|
||||
*
|
||||
* <p>{@code fleet_poll} is now <strong>three operations behind one tool name</strong>. With
|
||||
* <p>{@code fleet_poll} is <strong>three operations behind one tool name</strong>. With
|
||||
* {@code ticket} it observes an async delegation and changes nothing, which is a {@link
|
||||
* Authz.Action#READ}. With {@code target} it calls {@link MessageService#drainReplies} on that
|
||||
* session -- the replies are removed from the inbox and a second call returns nothing -- so it
|
||||
* is a {@link Authz.Action#DRAIN}, the same gate {@code fleet_ack} already uses for removing a
|
||||
* single message, and the same one the REST path uses at {@code FleetApp.drainReplies}. With
|
||||
* {@code coordId} it reads (never acks) this daemon's own held lead-to-lead mail, which is a
|
||||
* {@link Authz.Action#COORD_READ} -- <strong>not</strong> {@code READ}, even though nothing is
|
||||
* consumed: {@code READ}'s grant is open to every authenticated role on the premise that the
|
||||
* roster carries no secrets, and a lead-to-lead body is not the roster. Mapping a non-destructive
|
||||
* peer-mail read to {@code READ} would let any worker read every peer lead's mail in full.
|
||||
*
|
||||
* <p>Before this method existed (fleetd #272) the handler passed a constant {@code READ} for
|
||||
* both of the original branches. {@code READ} is open to every authenticated role, so any
|
||||
* worker could read a peer's id out of {@code fleet_list} and destroy the replies that peer had
|
||||
* queued for the primary. The gate failed open, and it did so because the required action is a
|
||||
* function of the arguments while the handler chose it before looking at them.
|
||||
* Authz.Action#TASK_READ}. With {@code target} it calls {@link MessageService#drainReplies} on
|
||||
* that session -- the replies are removed from the inbox and a second call returns nothing --
|
||||
* so it is a {@link Authz.Action#DRAIN}, the same gate {@code fleet_ack} already uses for
|
||||
* removing a single message, and the same one the REST path uses at
|
||||
* {@code FleetApp.drainReplies}. With {@code coordId} it reads (never acks) this daemon's own
|
||||
* held lead-to-lead mail, which is a {@link Authz.Action#COORD_READ} -- <strong>not</strong>
|
||||
* {@code TASK_READ} or {@code READ}: a lead-to-lead body is a different inbox from either, and
|
||||
* folding it into either would let any worker or architect read every peer lead's mail in full.
|
||||
*
|
||||
* <p>The choice lives in this method, and not inline in the handler, so that a test can assert
|
||||
* the mapping the handler actually uses. {@code FleetMcpAuthzTest} already checked every
|
||||
* {@link Authz.Action} against every {@link Role} and passed throughout -- it tested the policy
|
||||
* table, which was correct, while the defect was in which action the caller handed it.
|
||||
* the mapping the handler actually uses.
|
||||
*
|
||||
* <p>Checked first, and exclusively of {@code target}: a call naming {@code coordId} is reading
|
||||
* a different inbox entirely (this daemon's own lead channel, never a worker's), so it takes
|
||||
@@ -1072,7 +1139,30 @@ public final class FleetMcp {
|
||||
if (!isBlank(coordId)) {
|
||||
return Authz.Action.COORD_READ;
|
||||
}
|
||||
return isBlank(target) ? Authz.Action.READ : Authz.Action.DRAIN;
|
||||
return isBlank(target) ? Authz.Action.TASK_READ : Authz.Action.DRAIN;
|
||||
}
|
||||
|
||||
/**
|
||||
* Which authorization action a {@code fleet_send} call needs, decided by its arguments.
|
||||
*
|
||||
* <p>{@code fleet_send} is three call shapes behind one tool name, mirroring {@link
|
||||
* #pollAction}. With {@code coordId} it addresses a peer lead on another daemon over the
|
||||
* coordination broker, which is {@link Authz.Action#COORD_SEND}. With {@code turnId} it
|
||||
* resolves a worker's blocked {@code fleet_ask} and resumes that turn, which is {@link
|
||||
* Authz.Action#ANSWER}. Otherwise it delivers to a local session by {@code sessionId}, which is
|
||||
* the plain {@link Authz.Action#SEND}.
|
||||
*
|
||||
* <p>Checked in the same order the handler branches: {@code coordId} first and exclusively of
|
||||
* {@code turnId}, matching {@link #sendToLead}'s own mutual-exclusion check.
|
||||
*
|
||||
* @param coordId the {@code coordId} argument of the call, or {@code null}/blank when absent
|
||||
* @param turnId the {@code turnId} argument of the call, or {@code null}/blank when absent
|
||||
*/
|
||||
static Authz.Action sendAction(String coordId, String turnId) {
|
||||
if (!isBlank(coordId)) {
|
||||
return Authz.Action.COORD_SEND;
|
||||
}
|
||||
return isBlank(turnId) ? Authz.Action.SEND : Authz.Action.ANSWER;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1097,10 +1187,11 @@ public final class FleetMcp {
|
||||
*/
|
||||
private static Authz.Action authzAction(FleetTool tool, Map<String, Object> arguments) {
|
||||
return switch (tool) {
|
||||
case SEND -> Authz.Action.SEND;
|
||||
case SEND -> sendAction(str(arguments, "coordId"), str(arguments, "turnId"));
|
||||
case REPLY -> Authz.Action.REPLY;
|
||||
case ASK -> Authz.Action.ASK;
|
||||
case STATUS, LIST, PROFILES, WHOAMI -> Authz.Action.READ;
|
||||
case STATUS -> Authz.Action.TASK_READ;
|
||||
case LIST, PROFILES, WHOAMI -> Authz.Action.READ;
|
||||
case POLL -> pollAction(str(arguments, "target"), str(arguments, "coordId"));
|
||||
case ACK -> Authz.Action.DRAIN;
|
||||
case SPAWN -> Authz.Action.SPAWN;
|
||||
@@ -1120,9 +1211,24 @@ public final class FleetMcp {
|
||||
* exclusively of {@code ticket}/{@code target} — see {@link #pollAction}'s javadoc for why this
|
||||
* is a different inbox (this daemon's own {@link LeadChannel}) that authorizes differently
|
||||
* ({@link Authz.Action#COORD_READ}, primary-only) from either of the original two branches.
|
||||
*
|
||||
* <p>Does not check who owns {@code ticket} — see the overload that takes {@code
|
||||
* callerTerminal} for that. Callers that do not resolve a caller terminal (tests, or a surface
|
||||
* with no connection-based identity) use this one.
|
||||
*/
|
||||
static McpSchema.CallToolResult poll(MessageService messages, LeadChannel leadChannel, String ticket,
|
||||
String target, String coordId) {
|
||||
return poll(messages, leadChannel, ticket, target, coordId, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, refusing a ticket lookup whose caller's terminal differs from the terminal that
|
||||
* created it — see {@link MessageService#poll(String, String)}. {@code callerTerminal} is the
|
||||
* CALLING session's terminal id, resolved by the MCP layer from the connection, never a
|
||||
* client-supplied value.
|
||||
*/
|
||||
static McpSchema.CallToolResult poll(MessageService messages, LeadChannel leadChannel, String ticket,
|
||||
String target, String coordId, String callerTerminal) {
|
||||
if (!isBlank(coordId)) {
|
||||
return pollHeldPeerMail(leadChannel, coordId);
|
||||
}
|
||||
@@ -1136,7 +1242,7 @@ public final class FleetMcp {
|
||||
if (isBlank(ticket)) {
|
||||
return error("ticket (or target) is required");
|
||||
}
|
||||
MessageService.TaskView v = messages.poll(ticket);
|
||||
MessageService.TaskView v = messages.poll(ticket, callerTerminal);
|
||||
if (v == null) {
|
||||
return error("unknown ticket: " + ticket + " (never issued, or expired)");
|
||||
}
|
||||
@@ -1243,16 +1349,20 @@ public final class FleetMcp {
|
||||
|
||||
/**
|
||||
* {@code fleet_status}: the live lifecycle status of a worker session, plus — when the worker
|
||||
* is paused mid-turn in an async {@code fleet_ask} (CB-582) — the open question and how to
|
||||
* answer it, so a lead on its normal poll cadence does not need the ticket to notice.
|
||||
* is paused mid-turn in an async {@code fleet_ask} — the open question and how to answer it, so
|
||||
* a lead on its normal poll cadence does not need the ticket to notice. The question, its
|
||||
* {@code turnId} and its ticket id are shown only to the caller whose terminal created that
|
||||
* delegation, or to a caller with no terminal at all (the unnamed primary); any other caller
|
||||
* still sees the base status. {@code callerTerminal} is the CALLING session's terminal id,
|
||||
* resolved by the MCP layer from the connection, never a client-supplied value.
|
||||
*/
|
||||
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
|
||||
static McpSchema.CallToolResult status(MessageService messages, String sessionId, String callerTerminal) {
|
||||
if (isBlank(sessionId)) {
|
||||
return error("sessionId is required");
|
||||
}
|
||||
try {
|
||||
String base = messages.status(sessionId).name().toLowerCase();
|
||||
MessageService.PendingAsk ask = messages.pendingAsk(sessionId);
|
||||
MessageService.PendingAsk ask = messages.pendingAsk(sessionId, callerTerminal);
|
||||
if (ask == null) {
|
||||
return text(base);
|
||||
}
|
||||
@@ -1295,6 +1405,26 @@ public final class FleetMcp {
|
||||
}
|
||||
return text(json(m));
|
||||
}
|
||||
if (caller.isCollaborator()) {
|
||||
// A collaborator's name is its slot in the collaborators: registry; sessionId is its
|
||||
// pane so a peer knows where to reach it. No leader key: a collaborator is not a
|
||||
// primary for authorization, unlike a lead.
|
||||
if (caller.name() != null) {
|
||||
m.put("collaborator", caller.name());
|
||||
}
|
||||
if (caller.terminal() != null) {
|
||||
m.put("sessionId", caller.terminal());
|
||||
}
|
||||
return text(json(m));
|
||||
}
|
||||
if (caller.isObserver()) {
|
||||
// No architect slot, collaborator name, or lead name to report — only the pane itself,
|
||||
// so a peer that already knows this terminal can still address it.
|
||||
if (caller.terminal() != null) {
|
||||
m.put("sessionId", caller.terminal());
|
||||
}
|
||||
return text(json(m));
|
||||
}
|
||||
if (!caller.isWorker()) {
|
||||
// CB-530: which lead, once more than one pane is configured as one. `role` deliberately
|
||||
// still reads "primary" — the fallback ladder in CLAUDE.md keys on it, and a lead IS a
|
||||
@@ -1375,7 +1505,7 @@ public final class FleetMcp {
|
||||
}
|
||||
if (isBlank(callerTerminal)) {
|
||||
// An unnamed primary (token/loopback path, no resolved pane) has nowhere for the
|
||||
// eventual /clear + bootstrap to land — LeadRollover#open would throw
|
||||
// eventual relaunch + bootstrap to land — LeadRollover#open would throw
|
||||
// IllegalArgumentException for the same reason; refuse cleanly here instead.
|
||||
return error("fleet_handover requires a named lead pane (a resolved connection terminal) "
|
||||
+ "to open a rollover request against — an unnamed primary has none");
|
||||
@@ -1727,7 +1857,8 @@ public final class FleetMcp {
|
||||
Map<String, String> leads, String selfTerm,
|
||||
CoordinationSource coordination) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine,
|
||||
OutageSource.none(), LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
|
||||
OutageSource.none(), LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
|
||||
Map.of(), false, coordination, false, true, true);
|
||||
}
|
||||
|
||||
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
|
||||
@@ -1736,7 +1867,8 @@ public final class FleetMcp {
|
||||
QuarantineSource quarantine, OutageSource outage,
|
||||
Map<String, String> leads, String selfTerm) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
|
||||
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, CoordinationSource.none(), false);
|
||||
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
|
||||
Map.of(), false, CoordinationSource.none(), false, true, true);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1753,7 +1885,8 @@ public final class FleetMcp {
|
||||
QuarantineSource quarantine, Map<String, String> leads, String selfTerm,
|
||||
CoordinationSource coordination) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, OutageSource.none(),
|
||||
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
|
||||
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
|
||||
Map.of(), false, coordination, false, true, true);
|
||||
}
|
||||
|
||||
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
|
||||
@@ -1762,7 +1895,8 @@ public final class FleetMcp {
|
||||
QuarantineSource quarantine, OutageSource outage,
|
||||
Map<String, String> leads, String selfTerm, CoordinationSource coordination) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
|
||||
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
|
||||
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
|
||||
Map.of(), false, coordination, false, true, true);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1784,7 +1918,8 @@ public final class FleetMcp {
|
||||
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
|
||||
CoordinationSource coordination) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
|
||||
leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
|
||||
leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
|
||||
Map.of(), false, coordination, false, true, true);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1801,14 +1936,23 @@ public final class FleetMcp {
|
||||
* {@code false} (fleetd #463: a forgotten argument fails closed, not
|
||||
* open), so a test that wants the {@code coordinator} row must pass
|
||||
* an explicit {@code true}
|
||||
* @param leadsVisible whether this caller may see the {@code leads} array (see
|
||||
* {@link #leadsVisibleTo}) — unlike {@code callerIsPrimary}, this is
|
||||
* also {@code true} for an architect or a collaborator, so it cannot
|
||||
* be derived from {@code callerIsPrimary} alone
|
||||
* @param membersVisible whether this caller may see the {@code members} array (see
|
||||
* {@link #membersVisibleTo}); also {@code true} for an architect, but
|
||||
* unlike {@code leadsVisible}, never for a collaborator
|
||||
*/
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, OutageSource outage,
|
||||
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
|
||||
CoordinationSource coordination, boolean callerIsPrimary) {
|
||||
CoordinationSource coordination, boolean callerIsPrimary,
|
||||
boolean leadsVisible, boolean membersVisible) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine,
|
||||
outage, leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, callerIsPrimary);
|
||||
outage, leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
|
||||
Map.of(), false, coordination, callerIsPrimary, leadsVisible, membersVisible);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1818,6 +1962,17 @@ public final class FleetMcp {
|
||||
* matter); the one caller that matters for caching, {@code fleet_list}'s MCP handler, passes
|
||||
* its own single long-lived instance instead (see {@code FleetMcp}'s {@code leadContextGauge}
|
||||
* field).
|
||||
*
|
||||
* @param collaborators terminal_id → collaborator name, live from the resolver
|
||||
* ({@link dev.ltms.fleet.auth.CallerResolver#collaborators()})
|
||||
* @param collaboratorsVisible whether this caller may see the {@code collaborators} array (see
|
||||
* {@link #collaboratorsVisibleTo}); every wrapper overload above
|
||||
* passes {@code false}, so a test that wants the row must call this
|
||||
* overload with an explicit {@code true}
|
||||
* @param leadsVisible whether this caller may see the {@code leads} array (see
|
||||
* {@link #leadsVisibleTo})
|
||||
* @param membersVisible whether this caller may see the {@code members} array (see
|
||||
* {@link #membersVisibleTo})
|
||||
*/
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
@@ -1826,32 +1981,56 @@ public final class FleetMcp {
|
||||
LeadSeatSource leadSeats, LeadContextGauge contextGauge,
|
||||
LeadConfigDirSource leadConfigDirs,
|
||||
Map<String, String> leads, String selfTerm,
|
||||
CoordinationSource coordination, boolean callerIsPrimary) {
|
||||
Map<String, String> collaborators, boolean collaboratorsVisible,
|
||||
CoordinationSource coordination, boolean callerIsPrimary,
|
||||
boolean leadsVisible, boolean membersVisible) {
|
||||
try {
|
||||
Map<String, Agent> live = workers.list().stream()
|
||||
.map(Agent.class::cast)
|
||||
.filter(a -> a.terminalId() != null)
|
||||
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
|
||||
List<Map<String, Object>> leadRows = leads.entrySet().stream()
|
||||
.sorted(Map.Entry.comparingByValue())
|
||||
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm, contextGauge,
|
||||
leadConfigDirs))
|
||||
.toList();
|
||||
// Neither row's assembly (leadView/memberCapacityView probing herdr for live status)
|
||||
// runs unless at least one of them needs the live-agent lookup backing it.
|
||||
Map<String, Agent> live = (leadsVisible || membersVisible)
|
||||
? workers.list().stream()
|
||||
.map(Agent.class::cast)
|
||||
.filter(a -> a.terminalId() != null)
|
||||
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b))
|
||||
: Map.of();
|
||||
// fleetd #209: this is the caller-driven fleet_list read that actually reports
|
||||
// agentSessionId (via memberCapacityView -> SessionManager.rosterView), so it uses the
|
||||
// resolving roster; the heartbeat/health/metrics timers stay on the plain sessions.roster().
|
||||
List<MemberSession> roster = sessions.rosterResolved();
|
||||
List<Map<String, Object>> out = roster.stream()
|
||||
.map(s -> memberCapacityView(s, live.get(s.terminalId()), messages, capacity.clock().getAsLong()))
|
||||
.toList();
|
||||
Set<String> profiles = new java.util.TreeSet<>(capacity.configuredProfiles().get());
|
||||
roster.stream().map(MemberSession::profile).forEach(profiles::add);
|
||||
Map<String, Object> result = new LinkedHashMap<>();
|
||||
result.put("leads", leadRows); result.put("members", out);
|
||||
// READ is permission to enter this tool, not permission to receive every field it can
|
||||
// build -- gate BEFORE assembling each row, so the key is absent rather than
|
||||
// present-and-empty; a caller without either row gets its own facts from fleet_whoami.
|
||||
if (leadsVisible) {
|
||||
List<Map<String, Object>> leadRows = leads.entrySet().stream()
|
||||
.sorted(Map.Entry.comparingByValue())
|
||||
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm, contextGauge,
|
||||
leadConfigDirs))
|
||||
.toList();
|
||||
result.put("leads", leadRows);
|
||||
}
|
||||
if (membersVisible) {
|
||||
List<Map<String, Object>> out = roster.stream()
|
||||
.map(s -> memberCapacityView(s, live.get(s.terminalId()), messages, capacity.clock().getAsLong()))
|
||||
.toList();
|
||||
result.put("members", out);
|
||||
}
|
||||
result.put("healthCoverage", healthCoverage.value().get());
|
||||
result.put("loopHealth", Map.of(
|
||||
"statusPoller", loopHealth.statusPoller().get().name(),
|
||||
"sessionReaper", loopHealth.sessionReaper().get().name()));
|
||||
// fleetd #703: a collaborator tab is a person's own tab, so the row is assembled and
|
||||
// included only for the roles that may SEND to a named peer -- gate BEFORE assembling
|
||||
// it, same reason as the coordinator row just below: the key must be absent for a
|
||||
// worker, never present-and-empty.
|
||||
if (collaboratorsVisible) {
|
||||
result.put("collaborators", collaborators.entrySet().stream()
|
||||
.sorted(Map.Entry.comparingByValue())
|
||||
.map(e -> collaboratorRow(e.getKey(), e.getValue()))
|
||||
.toList());
|
||||
}
|
||||
// fleetd #439: coordinator/coordinatorView is lead-to-lead coordination state and must
|
||||
// never reach a worker or an architect -- gate BEFORE assembling it, not after, so the
|
||||
// key is absent rather than present-and-empty.
|
||||
@@ -2163,6 +2342,20 @@ public final class FleetMcp {
|
||||
return m;
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #703: one row of the {@code collaborators} array — a named peer's registry name and
|
||||
* the herdr {@code terminal_id} a {@code fleet_send} must target to reach it. No status, no
|
||||
* context gauge, no profile: a collaborator is never spawned and carries no profile, so those
|
||||
* fields have no meaning for it, and the scan behind {@code terminal} only reports a tab it
|
||||
* actually found in herdr, so a tab nobody has open does not appear here at all.
|
||||
*/
|
||||
private static Map<String, Object> collaboratorRow(String terminal, String name) {
|
||||
Map<String, Object> m = new LinkedHashMap<>();
|
||||
m.put("name", name);
|
||||
m.put("sessionId", terminal);
|
||||
return m;
|
||||
}
|
||||
|
||||
/**
|
||||
* Reads the lead context gauge for one lead. {@code configDir} is this lead's configured
|
||||
* {@code CLAUDE_CONFIG_DIR} override (see {@link LeadConfigDirSource}), derived from
|
||||
@@ -2357,7 +2550,14 @@ public final class FleetMcp {
|
||||
|
||||
private static McpSchema.Tool listTool() {
|
||||
return tool(FleetTool.LIST.wireName(),
|
||||
"List the whole fleet the bridge tracks, in two parts. 'leads' are your PEERS — other "
|
||||
"List the whole fleet the bridge tracks, in two parts. 'members' is visible to "
|
||||
+ "the primary and an architect only. 'leads' is visible to those two AND a "
|
||||
+ "collaborator — exactly the roles that may fleet_send to a lead, so a "
|
||||
+ "collaborator can learn a lead's sessionId before using the send it already "
|
||||
+ "holds. A worker holds READ to call this tool at all, but gets neither "
|
||||
+ "array, never an empty one; a worker reads its own session, "
|
||||
+ "profile, state, worktree, branch and owner from fleet_whoami instead. "
|
||||
+ "'leads' are your PEERS — other "
|
||||
+ "orchestrators, each with its sessionId (the address to fleet_send to), "
|
||||
+ "name, live status, and 'self': true on your own row; this is how you "
|
||||
+ "discover a peer lead without being told its address. 'members' are the "
|
||||
@@ -2388,7 +2588,12 @@ public final class FleetMcp {
|
||||
+ "msgId/from/preview, never the full body), and one row per coordinator.peers "
|
||||
+ "coord-id ('peers': coordId/reachable, plus pending/consumers when reachable) — "
|
||||
+ "this is peer DISCOVERY for cross-host leads, distinct from the local 'leads' "
|
||||
+ "array above. It is omitted entirely when no coordinator is configured.",
|
||||
+ "array above. It is omitted entirely when no coordinator is configured. A "
|
||||
+ "'collaborators' array, visible only to the primary, an architect, and a "
|
||||
+ "collaborator (never a worker), reports every other named peer tab this daemon "
|
||||
+ "recognises: each row is 'name' (its registry name) and 'sessionId' (the "
|
||||
+ "terminal id a fleet_send targets to reach it). Absent entirely for a worker, "
|
||||
+ "whatever is configured.",
|
||||
objectSchema(Map.of(), List.of()));
|
||||
}
|
||||
|
||||
@@ -2438,17 +2643,20 @@ public final class FleetMcp {
|
||||
private static McpSchema.Tool handoverTool() {
|
||||
return tool(FleetTool.HANDOVER.wireName(),
|
||||
"Replace your OWN lead session once its context is full: write a handover file, "
|
||||
+ "then use this to have fleetd clear your pane and bootstrap a fresh lead "
|
||||
+ "session against it. Four actions: 'open' (requests a token and the "
|
||||
+ "handoverPath you must write the handover file to before confirming), "
|
||||
+ "'confirm' (validates every gate and — only if every one passes — schedules "
|
||||
+ "the roll; it does NOT itself clear the pane, the roll runs once this call's "
|
||||
+ "own turn ends), 'cancel' (drops a pending request without rolling), and "
|
||||
+ "'status' (read-only: what happened to a token after 'confirm' — still "
|
||||
+ "running (approved but not finished yet), the roll completed, the calling "
|
||||
+ "turn never settled within turnSettleSeconds so no /clear was ever sent, or "
|
||||
+ "/clear itself never settled so bootstrapText was never sent; never "
|
||||
+ "schedules, cancels or retries anything). Primary-only. "
|
||||
+ "then use this to have fleetd end your pane's process and relaunch a fresh "
|
||||
+ "lead session bootstrapped against it. Four actions: 'open' (requests a "
|
||||
+ "token and the handoverPath you must write the handover file to before "
|
||||
+ "confirming), 'confirm' (validates every gate and — only if every one "
|
||||
+ "passes — schedules the roll; it does NOT itself end your pane, the roll "
|
||||
+ "runs once this call's own turn ends), 'cancel' (drops a pending request "
|
||||
+ "without rolling), and 'status' (read-only: what happened to a token after "
|
||||
+ "'confirm' — still running (approved but not finished yet), the roll "
|
||||
+ "completed, the calling turn never settled within turnSettleSeconds so "
|
||||
+ "nothing was touched, the old pane never confirmed dead so no relaunch was "
|
||||
+ "attempted, the relaunch itself failed, the fresh pane never became ready "
|
||||
+ "so bootstrapText was never sent, or the fresh pane became ready and was "
|
||||
+ "bootstrapped but was never recognised as a live lead; never schedules, "
|
||||
+ "cancels or retries anything). Primary-only. "
|
||||
+ "There is deliberately no terminal/session/leadTerminal parameter: the pane "
|
||||
+ "to roll is always resolved from YOUR OWN connection, never a value you "
|
||||
+ "pass, so you can only ever roll yourself — never another lead. Requires "
|
||||
|
||||
@@ -6,6 +6,7 @@ import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.AgentStatus;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.HerdrException;
|
||||
import dev.ltms.fleet.herdr.ResilientAgentLaunch;
|
||||
import dev.ltms.fleet.herdr.Tab;
|
||||
import dev.ltms.fleet.herdr.Workspace;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
@@ -67,16 +68,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(HerdrPeerLauncher.class);
|
||||
|
||||
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
|
||||
private static final int NAME_RETRIES = 8;
|
||||
|
||||
/**
|
||||
* Retries for {@code agent.start} against a seed pane whose shell has not reached its prompt
|
||||
* yet — {@code tab.create}/{@code pane.split} return as soon as the pane exists, and herdr
|
||||
* refuses to start an agent in a pane that is not "an available shell" ({@code agent_pane_busy}).
|
||||
*/
|
||||
private static final int SHELL_READY_RETRIES = 20;
|
||||
|
||||
private final String namePrefix; // label prefix: naming + reap scheme
|
||||
private final AgentControl agents;
|
||||
private final WorkspaceControl spaces;
|
||||
@@ -766,90 +757,26 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
|
||||
// argv[0] — the configured executable — is dropped and only the extra args are passed.
|
||||
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
|
||||
checkPaneCommandFits(cfg, argv);
|
||||
HerdrException last = null;
|
||||
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
|
||||
long seq = nameSeq.incrementAndGet();
|
||||
String name = namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + seq;
|
||||
try {
|
||||
return new Started(startAwaitingShellPrompt(name, args, paneId), seq);
|
||||
} catch (HerdrException e) {
|
||||
if (!"agent_name_taken".equals(e.code())) throw e;
|
||||
log.debug("peer name '{}' taken, retrying", name);
|
||||
last = e;
|
||||
}
|
||||
try {
|
||||
ResilientAgentLaunch.checkFits(cfg.profile(), argv);
|
||||
} catch (ResilientAgentLaunch.TooLargeException e) {
|
||||
throw new PeerUnreachableException(e.getMessage());
|
||||
}
|
||||
throw last;
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #220: herdr does not exec the launch command — it TYPES it into the pane as one line,
|
||||
* and a pty line buffer holds only {@value #PANE_COMMAND_BYTE_LIMIT} bytes (BSD/macOS {@code
|
||||
* MAX_CANON}). Everything past that byte is dropped. Nothing reports it: herdr answers "agent
|
||||
* started", the backend exits on the mangled argument it was handed, the pane closes, and the
|
||||
* only symptom is {@link #waitUntilInjectableOrThrow} timing out 20 seconds later with no
|
||||
* reason. That is exactly how #214 broke every claude-code spawn — one 50-byte flag pushed a
|
||||
* 978-byte command to 1028, and the tail that got cut was {@code --autocompact 250000}.
|
||||
*
|
||||
* <p>So measure it here and refuse, loudly and immediately, rather than spawn something that
|
||||
* cannot work. The estimate is deliberately conservative: fleetd cannot see herdr's quoting, so
|
||||
* every argument is charged its own bytes plus a separator and a quote pair. An over-estimate
|
||||
* costs a clear error at a length that was already unsafe; an under-estimate would let the
|
||||
* silent truncation back in.
|
||||
*
|
||||
* @throws PeerUnreachableException when the command cannot fit — the same failure the spawn
|
||||
* would have hit anyway, named at the point it is still
|
||||
* explainable
|
||||
*/
|
||||
private void checkPaneCommandFits(FleetConfig.Profile cfg, List<String> argv) {
|
||||
int bytes = 0;
|
||||
String longest = null;
|
||||
int longestBytes = 0;
|
||||
for (String arg : argv) {
|
||||
int argBytes = arg == null ? 0 : arg.getBytes(java.nio.charset.StandardCharsets.UTF_8).length;
|
||||
bytes += argBytes + QUOTING_OVERHEAD_PER_ARG;
|
||||
if (argBytes > longestBytes) {
|
||||
longestBytes = argBytes;
|
||||
longest = arg;
|
||||
}
|
||||
}
|
||||
if (bytes <= PANE_COMMAND_BYTE_LIMIT) {
|
||||
return;
|
||||
}
|
||||
String culprit = longest == null ? "<none>"
|
||||
: longest.substring(0, Math.min(longest.length(), 60)) + (longest.length() > 60 ? "…" : "");
|
||||
throw new PeerUnreachableException(
|
||||
"launch command for profile " + cfg.profile() + " is about " + bytes + " bytes, over the "
|
||||
+ PANE_COMMAND_BYTE_LIMIT + "-byte limit of the pane line herdr types it into. "
|
||||
+ "The pty would drop the tail silently and the backend would exit on a mangled "
|
||||
+ "argument. Longest argument is " + longestBytes + " bytes: " + culprit
|
||||
+ " — move it off the command line (a file flag) or shorten it.");
|
||||
long[] lastSeq = {0};
|
||||
Agent agent = ResilientAgentLaunch.startUniquelyNamed(agents, namePrefix, args, paneId,
|
||||
attempt -> {
|
||||
lastSeq[0] = nameSeq.incrementAndGet();
|
||||
return namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + lastSeq[0];
|
||||
},
|
||||
ResilientAgentLaunch.NAME_RETRIES, ResilientAgentLaunch.SHELL_READY_RETRIES, sleeper);
|
||||
return new Started(agent, lastSeq[0]);
|
||||
}
|
||||
|
||||
/**
|
||||
* The pty line buffer herdr types a launch command into: BSD/macOS {@code MAX_CANON}. Not a
|
||||
* fleetd choice and not configurable — see {@link #checkPaneCommandFits}.
|
||||
* fleetd choice and not configurable — see {@link ResilientAgentLaunch#checkFits}.
|
||||
*/
|
||||
static final int PANE_COMMAND_BYTE_LIMIT = 1024;
|
||||
|
||||
/** Per-argument allowance for the separating space and a shell quote pair fleetd cannot see. */
|
||||
private static final int QUOTING_OVERHEAD_PER_ARG = 3;
|
||||
|
||||
/** Start the agent into {@code paneId}, waiting out the seed shell's boot with the sleeper. */
|
||||
private Agent startAwaitingShellPrompt(String name, List<String> args, String paneId) {
|
||||
HerdrException busy = null;
|
||||
for (int attempt = 0; attempt < SHELL_READY_RETRIES; attempt++) {
|
||||
try {
|
||||
return agents.start(name, namePrefix, args, paneId);
|
||||
} catch (HerdrException e) {
|
||||
if (!"agent_pane_busy".equals(e.code())) throw e;
|
||||
log.debug("pane {} not at its shell prompt yet, retrying agent.start", paneId);
|
||||
busy = e;
|
||||
sleeper.run();
|
||||
}
|
||||
}
|
||||
throw busy;
|
||||
}
|
||||
static final int PANE_COMMAND_BYTE_LIMIT = ResilientAgentLaunch.PANE_COMMAND_BYTE_LIMIT;
|
||||
|
||||
// --- discovery + reap ----------------------------------------------------------------------
|
||||
|
||||
|
||||
@@ -87,7 +87,7 @@ public final class MessageService {
|
||||
/**
|
||||
* The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the
|
||||
* question and {@code turnId} correlates the answer. Not terminal — the primary answers with
|
||||
* {@link #answer(String, String, long)} and the turn resumes.
|
||||
* {@link #answer(String, String, long, String)} and the turn resumes.
|
||||
*/
|
||||
QUESTION,
|
||||
/** Timed out after the message was delivered — the worker is still working. */
|
||||
@@ -120,10 +120,16 @@ public final class MessageService {
|
||||
/** Another send to this session was in flight for the whole window. */
|
||||
BUSY,
|
||||
/**
|
||||
* An answer ({@link #answer(String, String, long)}) referenced a {@code turnId} that is no
|
||||
* longer open — the worker's {@code fleet_ask} already timed out or was answered.
|
||||
* An answer ({@link #answer(String, String, long, String)}) referenced a {@code turnId}
|
||||
* that is no longer open — the worker's {@code fleet_ask} already timed out or was answered.
|
||||
*/
|
||||
STALE_TURN
|
||||
STALE_TURN,
|
||||
/**
|
||||
* An answer ({@link #answer(String, String, long, String)}) named a {@code turnId} that is
|
||||
* still open, but the answering caller is not the caller whose accepted delegation opened
|
||||
* it. Distinct from {@link #STALE_TURN} so a refusal is never reported as a lapsed turn.
|
||||
*/
|
||||
NOT_TURN_OWNER
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -133,7 +139,7 @@ public final class MessageService {
|
||||
* {@link Outcome#COMPLETED_UNREPLIED}), or the question for {@link Outcome#QUESTION},
|
||||
* else {@code null}
|
||||
* @param turnId correlation id for a {@link Outcome#QUESTION} (answered via
|
||||
* {@link #answer(String, String, long)}), else {@code null}
|
||||
* {@link #answer(String, String, long, String)}), else {@code null}
|
||||
*/
|
||||
public record Reply(Outcome outcome, String text, String turnId) {
|
||||
/** A reply with no correlation id (the common terminal outcomes). */
|
||||
@@ -275,10 +281,19 @@ public final class MessageService {
|
||||
* {@link #abandon}) can never match again regardless of this flag's value.
|
||||
*/
|
||||
private volatile boolean askTimedOut;
|
||||
/**
|
||||
* The terminal of the caller whose {@code fleet_send{wait:false}} created this ticket, or
|
||||
* {@code null} when that caller had no terminal (the unnamed primary) or the ticket was
|
||||
* created through an overload that does not record one. {@link #poll(String, String)}
|
||||
* compares a polling caller's own terminal against this field before handing back the
|
||||
* ticket's state.
|
||||
*/
|
||||
private final String creatorTerminal;
|
||||
|
||||
private Task(String ticket, String target, LongSupplier nowNanos) {
|
||||
private Task(String ticket, String target, LongSupplier nowNanos, String creatorTerminal) {
|
||||
this.ticket = ticket;
|
||||
this.target = target;
|
||||
this.creatorTerminal = creatorTerminal;
|
||||
this.createdNanos = nowNanos.getAsLong();
|
||||
future.whenComplete((reply, ex) -> completedNanos = nowNanos.getAsLong());
|
||||
}
|
||||
@@ -340,6 +355,14 @@ public final class MessageService {
|
||||
*/
|
||||
private final ConcurrentHashMap<String, Boolean> queuedDeliveries = new ConcurrentHashMap<>();
|
||||
private final AtomicLong ticketSeq = new AtomicLong();
|
||||
/**
|
||||
* Minted once per {@code MessageService} instance and folded into every ticket id (see
|
||||
* {@link #sendAsync(String, String, Runnable, String)}). {@link #ticketSeq} alone restarts at
|
||||
* zero for every instance, so without this a ticket id can be reused across instances and
|
||||
* resolve to an unrelated {@link Task} with no error; this nonce makes that impossible, because
|
||||
* an id minted by one instance can never match the id space of another.
|
||||
*/
|
||||
private final String ticketBootNonce = UUID.randomUUID().toString().substring(0, 6);
|
||||
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
|
||||
Thread.ofVirtual().name("bridge-async-", 0).factory());
|
||||
|
||||
@@ -656,7 +679,7 @@ public final class MessageService {
|
||||
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, TIMED_OUT_UNCONFIRMED, BUSY -> "timeout";
|
||||
case WORKER_FAILED -> "failed";
|
||||
case BACKEND_EXHAUSTED -> "backend_exhausted";
|
||||
case STALE_TURN, QUESTION -> null; // not a completed delegation
|
||||
case STALE_TURN, QUESTION, NOT_TURN_OWNER -> null; // not a completed delegation
|
||||
};
|
||||
}
|
||||
|
||||
@@ -915,14 +938,17 @@ public final class MessageService {
|
||||
|
||||
/**
|
||||
* Deliver {@code content} to {@code target} (a herdr {@code terminal_id}) and block until the
|
||||
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
|
||||
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses. {@code callerTerminal}
|
||||
* is the terminal of the caller making this call — {@code null} for the unnamed primary — and is
|
||||
* recorded as the turn's owner, the only caller {@link #answer(String, String, long, String)} will
|
||||
* later accept an answer from if the worker pauses mid-turn to ask.
|
||||
*/
|
||||
public Reply send(String target, String content, long timeoutMillis) {
|
||||
return send(target, content, timeoutMillis, null);
|
||||
public Reply send(String target, String content, long timeoutMillis, String callerTerminal) {
|
||||
return send(target, content, timeoutMillis, null, callerTerminal);
|
||||
}
|
||||
|
||||
/**
|
||||
* As {@link #send(String, String, long)}, but with an accepted-delivery hook.
|
||||
* As {@link #send(String, String, long, String)}, but with an accepted-delivery hook.
|
||||
*
|
||||
* <p>{@code onAccepted} is invoked exactly once, once this send has won {@code target}'s send
|
||||
* lock and so become the <em>accepted target turn</em> — it runs <em>before</em> delivery is
|
||||
@@ -933,12 +959,13 @@ public final class MessageService {
|
||||
* acceptance means a concurrent sender that times out {@code BUSY} can never steal ownership it
|
||||
* never earned. {@code null} disables the hook.
|
||||
*/
|
||||
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted) {
|
||||
return send(target, content, timeoutMillis, onAccepted, null);
|
||||
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, String callerTerminal) {
|
||||
return send(target, content, timeoutMillis, onAccepted, null, callerTerminal);
|
||||
}
|
||||
|
||||
/** Run a send, optionally stopping an async task that teardown already failed before acceptance. */
|
||||
private Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, Task task) {
|
||||
private Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, Task task,
|
||||
String callerTerminal) {
|
||||
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
|
||||
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
|
||||
|
||||
@@ -958,7 +985,7 @@ public final class MessageService {
|
||||
// open race). Opening first also means a throwing onAccepted (fired before enqueue) or an
|
||||
// enqueue failure is safely closed by the finally below: nothing is left queued, and the
|
||||
// failed send leaves no stale waiter behind.
|
||||
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
|
||||
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target, Rendezvous.Owner.of(callerTerminal));
|
||||
// CB-640: this send now owns target's delivery, so any earlier stranded-reply or
|
||||
// still-queued fact no longer describes the live state — clear both rather than let
|
||||
// them outlive the send that supersedes them.
|
||||
@@ -1137,19 +1164,30 @@ public final class MessageService {
|
||||
* mid-turn (already picked up), so the answer flows back through its own open {@code fleet_ask}
|
||||
* call, not a new status-gated delivery. The forward waiter is opened <em>before</em> the worker
|
||||
* is unblocked so a reply that lands the instant it resumes is not lost.
|
||||
*
|
||||
* <p>{@code callerTerminal} is the terminal of the caller making this call — {@code null} for
|
||||
* the unnamed primary. It is checked against the turn's recorded owner (the caller whose
|
||||
* accepted delegation opened it, see {@link #send(String, String, long, String)} and
|
||||
* {@link #sendAsync(String, String, Runnable, String)}) before anything else runs: a mismatch,
|
||||
* including a turn with no owner on record at all, returns {@link Outcome#NOT_TURN_OWNER}
|
||||
* without touching the rendezvous, the session lock, or any async task bookkeeping.
|
||||
*/
|
||||
public Reply answer(String turnId, String content, long timeoutMillis) {
|
||||
public Reply answer(String turnId, String content, long timeoutMillis, String callerTerminal) {
|
||||
String workerSession = rendezvous.askSession(turnId);
|
||||
if (workerSession == null) {
|
||||
return new Reply(Outcome.STALE_TURN, null); // the ask lapsed (timed out or already answered)
|
||||
}
|
||||
Rendezvous.Owner owner = rendezvous.askOwner(turnId);
|
||||
if (!Rendezvous.Owner.permits(owner, callerTerminal)) {
|
||||
return new Reply(Outcome.NOT_TURN_OWNER, null);
|
||||
}
|
||||
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
|
||||
ReentrantLock lock = sessionLocks.computeIfAbsent(workerSession, _ -> new ReentrantLock());
|
||||
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
|
||||
return new Reply(Outcome.BUSY, null);
|
||||
}
|
||||
try {
|
||||
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession);
|
||||
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession, owner);
|
||||
// fleetd #575: this try used to open below, AFTER the Task lookup/registration and the
|
||||
// STALE_TURN early return that follows it — so that return was covered only by a
|
||||
// hand-rolled copy of the finally's own cleanup pair, not the finally itself. Widening the
|
||||
@@ -1279,8 +1317,20 @@ public final class MessageService {
|
||||
* @return the ticket to poll for the eventual result
|
||||
*/
|
||||
public String sendAsync(String target, String content, Runnable onAccepted) {
|
||||
String ticket = "task-" + ticketSeq.incrementAndGet();
|
||||
Task task = new Task(ticket, target, nowNanos);
|
||||
return sendAsync(target, content, onAccepted, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* As {@link #sendAsync(String, String, Runnable)}, recording {@code creatorTerminal} as this
|
||||
* ticket's owner. {@link #poll(String, String)} refuses a later caller whose own terminal
|
||||
* differs from this one; {@code null} records no owner (a caller with no terminal — the
|
||||
* unnamed primary — is always allowed to poll the result regardless).
|
||||
*
|
||||
* @return the ticket to poll for the eventual result
|
||||
*/
|
||||
public String sendAsync(String target, String content, Runnable onAccepted, String creatorTerminal) {
|
||||
String ticket = "task-" + ticketBootNonce + "-" + ticketSeq.incrementAndGet();
|
||||
Task task = new Task(ticket, target, nowNanos, creatorTerminal);
|
||||
tasks.put(ticket, task);
|
||||
if (pushLoop != null) {
|
||||
// CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a
|
||||
@@ -1311,7 +1361,7 @@ public final class MessageService {
|
||||
}
|
||||
asyncExecutor.submit(() -> {
|
||||
try {
|
||||
Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task);
|
||||
Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task, creatorTerminal);
|
||||
if (result.outcome() == Outcome.QUESTION) {
|
||||
// Keep the accepted owner until answer() finishes it. markAsyncQuestion may run
|
||||
// just after resolveQuestion wakes this thread.
|
||||
@@ -1343,15 +1393,31 @@ public final class MessageService {
|
||||
}
|
||||
|
||||
/**
|
||||
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket;
|
||||
* otherwise a {@link Phase#PENDING} view (with the live worker status as detail), a
|
||||
* {@link Phase#DONE} view carrying the reply, or a {@link Phase#FAILED} view with the reason.
|
||||
* As {@link #poll(String, String)}, with no caller terminal — the ticket's ownership is never
|
||||
* checked, so this overload must only be used where the caller's identity is otherwise
|
||||
* irrelevant (a test, or a surface that does not resolve a caller terminal at all).
|
||||
*/
|
||||
public TaskView poll(String ticket) {
|
||||
return poll(ticket, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket.
|
||||
* Refuses a {@code callerTerminal} that differs from the terminal that created the ticket (see
|
||||
* {@link #sendAsync(String, String, Runnable, String)}) with a {@link Phase#FAILED} view that
|
||||
* carries no reply text — a caller with no terminal (the unnamed primary) is never refused.
|
||||
* Otherwise returns a {@link Phase#PENDING} view (with the live worker status as detail), a
|
||||
* {@link Phase#DONE} view carrying the reply, or a {@link Phase#FAILED} view with the reason.
|
||||
*/
|
||||
public TaskView poll(String ticket, String callerTerminal) {
|
||||
Task task = tasks.get(ticket);
|
||||
if (task == null) {
|
||||
return null;
|
||||
}
|
||||
if (!ownsTicket(task, callerTerminal)) {
|
||||
return new TaskView(ticket, Phase.FAILED, null, null,
|
||||
"forbidden: this ticket was created by a different session", null);
|
||||
}
|
||||
CompletableFuture<Reply> f = task.future;
|
||||
if (!f.isDone()) {
|
||||
Reply question = task.question;
|
||||
@@ -1387,6 +1453,17 @@ public final class MessageService {
|
||||
return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether {@code callerTerminal} may read {@code task}'s state. A caller with no terminal
|
||||
* always may — that is the unnamed primary, resolved by token or loopback trust, which never
|
||||
* carries a herdr pane and must keep reading every ticket. Otherwise the caller's terminal must
|
||||
* equal the terminal recorded on the task; a task with no recorded terminal matches no
|
||||
* terminal-bearing caller.
|
||||
*/
|
||||
private static boolean ownsTicket(Task task, String callerTerminal) {
|
||||
return callerTerminal == null || callerTerminal.equals(task.creatorTerminal);
|
||||
}
|
||||
|
||||
/**
|
||||
* Test seam only — carries no production behaviour, and nothing in this class calls it;
|
||||
* {@link #pruneTerminalTickets} still reads {@link Task#completedNanos} directly.
|
||||
@@ -1706,16 +1783,17 @@ public final class MessageService {
|
||||
}
|
||||
|
||||
/**
|
||||
* The question {@code workerSession} is currently paused on via {@code fleet_ask}, if any
|
||||
* (CB-582) — {@code fleet_status} uses this to show a pending question without the caller
|
||||
* needing the ticket. {@code null} when the session has no open async question (including a
|
||||
* session mid a <em>blocking</em> {@code fleet_ask}, which has no {@link Task} to look up — see
|
||||
* {@link PendingAsk}).
|
||||
* The question {@code workerSession} is currently paused on via {@code fleet_ask}, if any —
|
||||
* {@code fleet_status} uses this to show a pending question without the caller needing the
|
||||
* ticket. {@code null} when the session has no open async question (including a session mid a
|
||||
* <em>blocking</em> {@code fleet_ask}, which has no {@link Task} to look up — see
|
||||
* {@link PendingAsk}), or when {@code callerTerminal} does not own the task the question
|
||||
* belongs to (see {@link #ownsTicket(Task, String)}).
|
||||
*/
|
||||
public PendingAsk pendingAsk(String workerSession) {
|
||||
public PendingAsk pendingAsk(String workerSession, String callerTerminal) {
|
||||
for (Task task : tasks.values()) {
|
||||
Reply q = task.question;
|
||||
if (q != null && workerSession.equals(task.target)) {
|
||||
if (q != null && workerSession.equals(task.target) && ownsTicket(task, callerTerminal)) {
|
||||
return new PendingAsk(task.ticket, q.text(), q.turnId());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
package dev.ltms.fleet.msg;
|
||||
|
||||
import java.util.UUID;
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
@@ -60,7 +61,7 @@ public final class Rendezvous {
|
||||
}
|
||||
|
||||
/** A worker's open mid-turn question: the worker session it belongs to and the answer future. */
|
||||
private record AskWaiter(String session, CompletableFuture<String> answer) {
|
||||
private record AskWaiter(String session, CompletableFuture<String> answer, Owner owner) {
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -70,11 +71,45 @@ public final class Rendezvous {
|
||||
public record AskTicket(String turnId, CompletableFuture<String> answer, boolean fresh) {
|
||||
}
|
||||
|
||||
private final ConcurrentHashMap<String, CompletableFuture<Resolution>> waiters = new ConcurrentHashMap<>();
|
||||
/**
|
||||
* The caller whose accepted delegation opened a turn — the only caller allowed to answer it.
|
||||
* A {@code null} terminal means the unnamed primary, an authenticated caller with no pane.
|
||||
*/
|
||||
public record Owner(String terminal) {
|
||||
public static final Owner UNNAMED_PRIMARY = new Owner(null);
|
||||
|
||||
public static Owner of(String terminal) {
|
||||
return terminal == null ? UNNAMED_PRIMARY : new Owner(terminal);
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether {@code callerTerminal} matches {@code owner}. A {@code null} owner means no
|
||||
* owner was ever recorded, and that state matches no caller, not even one whose own
|
||||
* terminal is {@code null} — "no record" and "recorded as the unnamed primary" are
|
||||
* different states.
|
||||
*/
|
||||
public static boolean permits(Owner owner, String callerTerminal) {
|
||||
return owner != null && java.util.Objects.equals(owner.terminal(), callerTerminal);
|
||||
}
|
||||
}
|
||||
|
||||
/** A registered forward waiter together with the owner its delegation was opened under. */
|
||||
private record ForwardWaiter(Owner owner, CompletableFuture<Resolution> future) {
|
||||
}
|
||||
|
||||
private final ConcurrentHashMap<String, ForwardWaiter> waiters = new ConcurrentHashMap<>();
|
||||
|
||||
/** Reverse rendezvous (CB-205): worker questions awaiting the primary's answer, keyed by {@code turnId}. */
|
||||
private final ConcurrentHashMap<String, AskWaiter> asks = new ConcurrentHashMap<>();
|
||||
private final AtomicLong askSeq = new AtomicLong();
|
||||
/**
|
||||
* Minted once per {@code Rendezvous} instance and folded into every {@code turnId} (see
|
||||
* {@link #openAsk(String)}). {@link #askSeq} alone restarts at zero for every instance, so
|
||||
* without this a {@code turnId} minted by one instance could be minted again by another and
|
||||
* resolve to an unrelated ask with no error; this nonce makes that impossible, because an id
|
||||
* minted by one instance can never match the id space of another.
|
||||
*/
|
||||
private final String askBootNonce = UUID.randomUUID().toString().substring(0, 6);
|
||||
/** Per-session index of the currently-open ask, so duplicate fleet_ask calls coalesce onto one turn. */
|
||||
private final ConcurrentHashMap<String, String> openAsksBySession = new ConcurrentHashMap<>();
|
||||
|
||||
@@ -89,8 +124,17 @@ public final class Rendezvous {
|
||||
* code a double open is impossible; this is a tripwire for the day that no longer holds.
|
||||
*/
|
||||
public CompletableFuture<Resolution> open(String session) {
|
||||
return open(session, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Same as {@link #open(String)}, additionally recording {@code owner} as the caller whose
|
||||
* delegation opened this waiter. A {@code null} owner records no owner at all — the
|
||||
* fail-closed default {@link Owner#permits} refuses to everyone.
|
||||
*/
|
||||
public CompletableFuture<Resolution> open(String session, Owner owner) {
|
||||
CompletableFuture<Resolution> waiter = new CompletableFuture<>();
|
||||
CompletableFuture<Resolution> existing = waiters.putIfAbsent(session, waiter);
|
||||
ForwardWaiter existing = waiters.putIfAbsent(session, new ForwardWaiter(owner, waiter));
|
||||
if (existing != null) {
|
||||
throw new IllegalStateException(
|
||||
"rendezvous double-open for session " + session + " — a waiter is already registered");
|
||||
@@ -105,7 +149,13 @@ public final class Rendezvous {
|
||||
* successful {@code open} after a finished turn requires this close to have happened first).
|
||||
*/
|
||||
public void close(String session, CompletableFuture<Resolution> waiter) {
|
||||
waiters.remove(session, waiter);
|
||||
waiters.computeIfPresent(session, (s, w) -> w.future() == waiter ? null : w);
|
||||
}
|
||||
|
||||
/** The owner recorded for {@code session}'s open waiter, or {@code null} if none is open. */
|
||||
public Owner ownerOf(String session) {
|
||||
ForwardWaiter w = waiters.get(session);
|
||||
return w == null ? null : w.owner();
|
||||
}
|
||||
|
||||
/** Whether a send is currently awaiting a resolution for {@code session}. */
|
||||
@@ -119,7 +169,8 @@ public final class Rendezvous {
|
||||
* send (see the CB-116 note above) rather than whichever send happens to be waiting when they fire.
|
||||
*/
|
||||
public CompletableFuture<Resolution> currentWaiter(String session) {
|
||||
return waiters.get(session);
|
||||
ForwardWaiter w = waiters.get(session);
|
||||
return w == null ? null : w.future();
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -145,9 +196,9 @@ public final class Rendezvous {
|
||||
while (true) {
|
||||
AskWaiter[] minted = { null };
|
||||
String turnId = openAsksBySession.computeIfAbsent(session, _ -> {
|
||||
String newTurnId = session + "#" + askSeq.incrementAndGet();
|
||||
String newTurnId = session + "#" + askBootNonce + "-" + askSeq.incrementAndGet();
|
||||
CompletableFuture<String> answer = new CompletableFuture<>();
|
||||
AskWaiter waiter = new AskWaiter(session, answer);
|
||||
AskWaiter waiter = new AskWaiter(session, answer, ownerOf(session));
|
||||
asks.put(newTurnId, waiter);
|
||||
minted[0] = waiter;
|
||||
return newTurnId;
|
||||
@@ -183,6 +234,16 @@ public final class Rendezvous {
|
||||
return w == null ? null : w.session();
|
||||
}
|
||||
|
||||
/**
|
||||
* The owner recorded for {@code turnId} when its ask turn was freshly opened — the caller
|
||||
* whose delegation {@link #answerAsk} must match. {@code null} if {@code turnId} is unknown or
|
||||
* lapsed, or if the ask opened with no forward waiter owner on record.
|
||||
*/
|
||||
public Owner askOwner(String turnId) {
|
||||
AskWaiter w = asks.get(turnId);
|
||||
return w == null ? null : w.owner();
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve a worker's blocked {@code fleet_ask} with the primary's {@code answer}, unblocking it
|
||||
* to resume its turn.
|
||||
@@ -245,7 +306,7 @@ public final class Rendezvous {
|
||||
}
|
||||
|
||||
private boolean complete(String session, Resolution resolution) {
|
||||
CompletableFuture<Resolution> waiter = waiters.get(session);
|
||||
return waiter != null && waiter.complete(resolution);
|
||||
ForwardWaiter waiter = waiters.get(session);
|
||||
return waiter != null && waiter.future().complete(resolution);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -49,22 +49,52 @@ import java.util.stream.Collectors;
|
||||
*/
|
||||
public final class FleetApp {
|
||||
|
||||
/** The authorization action the matching route handler hands to {@link #allow}. */
|
||||
/**
|
||||
* The authorization action the matching route handler hands to {@link #allow}, for a route
|
||||
* whose action does not depend on the request body.
|
||||
*/
|
||||
static Authz.Action routeAction(String route) {
|
||||
return routeAction(route, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, plus the one route whose action depends on the body: {@code POST
|
||||
* /sessions/{id}/message} carries a {@code turnId} (the answer-a-blocked-worker shape) or not
|
||||
* (a plain delivery), mirroring {@code FleetMcp#sendAction}'s split of the same two call
|
||||
* shapes over MCP. {@code turnId} is ignored by every other route.
|
||||
*
|
||||
* @param turnId the request body's {@code turnId}, or {@code null}/blank when absent or not
|
||||
* applicable to this route
|
||||
*/
|
||||
static Authz.Action routeAction(String route, String turnId) {
|
||||
return switch (route) {
|
||||
case "GET /metrics" -> Authz.Action.METRICS;
|
||||
case "POST /members" -> Authz.Action.SPAWN;
|
||||
case "DELETE /members/{paneId}" -> Authz.Action.STOP;
|
||||
case "POST /sessions/{id}/message" -> Authz.Action.SEND;
|
||||
case "POST /sessions/{id}/message" -> turnId == null || turnId.isBlank()
|
||||
? Authz.Action.SEND : Authz.Action.ANSWER;
|
||||
case "POST /sessions/{id}/reply" -> Authz.Action.REPLY;
|
||||
case "GET /sessions/{id}/replies" -> Authz.Action.DRAIN;
|
||||
case "POST /sessions/{id}/ask" -> Authz.Action.ASK;
|
||||
case "GET /sessions", "GET /agents", "GET /members", "GET /profiles",
|
||||
"GET /member-credentials", "GET /sessions/{id}/status", "GET /tasks/{ticket}" -> Authz.Action.READ;
|
||||
"GET /member-credentials" -> Authz.Action.READ;
|
||||
case "GET /sessions/{id}/status", "GET /tasks/{ticket}" -> Authz.Action.TASK_READ;
|
||||
default -> throw new IllegalArgumentException("route has no authorization gate: " + route);
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* The second gate for {@code POST /sessions/{id}/message}: checked only when {@code turnId}
|
||||
* is present and non-blank, against {@link Authz.Action#ANSWER}. A request with no {@code
|
||||
* turnId} passes this gate unconditionally, without consulting {@code permit} at all, having
|
||||
* already cleared the coarse {@link Authz.Action#SEND} grant checked ahead of it.
|
||||
*
|
||||
* @param permit reports whether the caller holds the named grant
|
||||
*/
|
||||
static boolean answerGatePasses(String turnId, Predicate<Authz.Action> permit) {
|
||||
return turnId == null || turnId.isBlank() || permit.test(Authz.Action.ANSWER);
|
||||
}
|
||||
|
||||
/** Default blocking window for a message; kept under typical HTTP idle timeouts. */
|
||||
private static final long DEFAULT_MESSAGE_TIMEOUT_MS = 25_000;
|
||||
private static final long MAX_MESSAGE_TIMEOUT_MS = 120_000;
|
||||
@@ -236,6 +266,21 @@ public final class FleetApp {
|
||||
return app;
|
||||
}
|
||||
|
||||
/**
|
||||
* The authorization decision behind {@link #allow}, taking the caller directly rather than
|
||||
* pulling it from a servlet {@link Context} — unit-testable without fabricating a live
|
||||
* request, the same reason {@code FleetMcp#denyFor} is split from {@code FleetMcp#deny}.
|
||||
*
|
||||
* @param knownLeadOrCollaborator the classifier a collaborator's {@code SEND} is checked
|
||||
* against; pass {@link #auth}'s own {@code
|
||||
* knownLeadOrCollaborator()} to exercise the real production
|
||||
* gate, as {@link #allow} does
|
||||
*/
|
||||
static boolean permitsFor(Principal caller, Authz.Action action, String target,
|
||||
Predicate<String> knownLeadOrCollaborator) {
|
||||
return Authz.permits(caller, action, target, knownLeadOrCollaborator);
|
||||
}
|
||||
|
||||
/**
|
||||
* Gate a handler on the CB-505 authorization table. Returns {@code true} when the request may
|
||||
* proceed; otherwise writes the error response and returns {@code false}.
|
||||
@@ -249,8 +294,9 @@ public final class FleetApp {
|
||||
return true; // legacy: authorization not enforced
|
||||
}
|
||||
Principal caller = ctx.attribute(CALLER);
|
||||
if (Authz.permits(caller, action, target)) {
|
||||
if (action != Authz.Action.READ && action != Authz.Action.METRICS) {
|
||||
if (permitsFor(caller, action, target, auth.knownLeadOrCollaborator())) {
|
||||
if (action != Authz.Action.READ && action != Authz.Action.METRICS
|
||||
&& action != Authz.Action.TASK_READ) {
|
||||
AuditLog.allowed(caller, action, target); // reads would drown the trail
|
||||
}
|
||||
return true;
|
||||
@@ -603,26 +649,39 @@ public final class FleetApp {
|
||||
* status-gated injector and block until the worker returns a structured {@code fleet_reply}.
|
||||
* Times out with a typed 202 (working / queued / busy) rather than an error — the message may
|
||||
* still land.
|
||||
*
|
||||
* <p>Two call shapes share this route, exactly as {@code fleet_send} does over MCP (see
|
||||
* {@code FleetMcp#sendAction}): a plain delivery to {@code id}, and -- when the body carries
|
||||
* {@code turnId} -- resolving a worker's blocked question. The coarse {@link
|
||||
* Authz.Action#SEND} grant is checked first, before the body is read at all; only once that
|
||||
* passes is the body parsed, and a present {@code turnId} is then checked again against
|
||||
* {@link Authz.Action#ANSWER}. A body that fails to parse is rejected with 400 and reaches
|
||||
* neither {@code messages.answer} nor {@code messages.send}.
|
||||
*/
|
||||
private void sendMessage(Context ctx) {
|
||||
String id = ctx.pathParam("id");
|
||||
if (!allow(ctx, routeAction("POST /sessions/{id}/message"), id)) {
|
||||
return;
|
||||
}
|
||||
String content;
|
||||
String turnId;
|
||||
long timeout;
|
||||
boolean wait;
|
||||
Principal caller = ctx.attribute(CALLER);
|
||||
String callerTerminal = caller == null ? null : caller.terminal();
|
||||
JsonNode body;
|
||||
try {
|
||||
JsonNode body = mapper.readTree(ctx.body());
|
||||
content = body.path("content").asText("");
|
||||
turnId = body.path("turnId").asText(null);
|
||||
timeout = body.path("timeoutMs").asLong(DEFAULT_MESSAGE_TIMEOUT_MS);
|
||||
wait = body.path("wait").asBoolean(true); // default: block for the reply (CB-104)
|
||||
body = mapper.readTree(ctx.body());
|
||||
} catch (Exception e) {
|
||||
body = null;
|
||||
}
|
||||
if (body == null) {
|
||||
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
|
||||
return;
|
||||
}
|
||||
String turnId = body.path("turnId").asText(null);
|
||||
if (!answerGatePasses(turnId, action -> allow(ctx, action, id))) {
|
||||
return;
|
||||
}
|
||||
String content = body.path("content").asText("");
|
||||
long timeout = body.path("timeoutMs").asLong(DEFAULT_MESSAGE_TIMEOUT_MS);
|
||||
boolean wait = body.path("wait").asBoolean(true); // default: block for the reply (CB-104)
|
||||
if (content.isBlank()) {
|
||||
ctx.status(400).json(Map.of("error", "bad_request", "detail", "content is required"));
|
||||
return;
|
||||
@@ -631,19 +690,19 @@ public final class FleetApp {
|
||||
|
||||
// Answering a worker's fleet_ask (CB-205): always blocks, and derives the worker from turnId.
|
||||
if (turnId != null && !turnId.isBlank()) {
|
||||
writeReply(ctx, id, messages.answer(turnId, content, timeout), timeout);
|
||||
writeReply(ctx, id, messages.answer(turnId, content, timeout, callerTerminal), timeout);
|
||||
return;
|
||||
}
|
||||
|
||||
if (!wait) {
|
||||
// Fire-and-poll (CB-107): return a ticket immediately; the caller polls GET /tasks/{ticket}.
|
||||
String ticket = messages.sendAsync(id, content);
|
||||
String ticket = messages.sendAsync(id, content, null, callerTerminal);
|
||||
ctx.status(202).json(Map.of("sessionId", id, "ticket", ticket, "status", "accepted"));
|
||||
return;
|
||||
}
|
||||
|
||||
try {
|
||||
writeReply(ctx, id, messages.send(id, content, timeout), timeout);
|
||||
writeReply(ctx, id, messages.send(id, content, timeout, callerTerminal), timeout);
|
||||
} catch (HerdrException e) {
|
||||
herdrError(ctx, e);
|
||||
}
|
||||
@@ -662,6 +721,10 @@ public final class FleetApp {
|
||||
case STALE_TURN -> ctx.status(409).json(Map.of(
|
||||
"sessionId", id, "error", "stale_turn",
|
||||
"detail", "that question is no longer open (timed out or already answered)"));
|
||||
case NOT_TURN_OWNER -> ctx.status(403).json(Map.of(
|
||||
"sessionId", id, "error", "not_turn_owner",
|
||||
"detail", "this turn belongs to a different delegation — only the caller that "
|
||||
+ "opened it may answer it"));
|
||||
case REPLIED, COMPLETED_UNREPLIED -> {
|
||||
// replySource distinguishes a structured fleet_reply from the CB-106 completion
|
||||
// fallback (a scrape of the worker's transcript when it finished without replying).
|
||||
@@ -676,9 +739,10 @@ public final class FleetApp {
|
||||
// silent fall-through. That is exactly the bug this ticket exists to fix:
|
||||
// `default -> "done"` used to sit here and would have told a REST caller the
|
||||
// delegation completed for TIMED_OUT_UNCONFIRMED, the one outcome where delivery
|
||||
// is unknown. REPLIED, COMPLETED_UNREPLIED, QUESTION and STALE_TURN can never
|
||||
// actually reach this inner switch — the outer switch above always dispatches
|
||||
// them first — but they still need an arm to keep this switch exhaustive.
|
||||
// is unknown. REPLIED, COMPLETED_UNREPLIED, QUESTION, STALE_TURN and
|
||||
// NOT_TURN_OWNER can never actually reach this inner switch — the outer switch
|
||||
// above always dispatches them first — but they still need an arm to keep this
|
||||
// switch exhaustive.
|
||||
"status", switch (reply.outcome()) {
|
||||
case TIMED_OUT_WORKING -> "working";
|
||||
case TIMED_OUT_QUEUED -> "queued";
|
||||
@@ -688,7 +752,7 @@ public final class FleetApp {
|
||||
case BUSY -> "busy";
|
||||
case WORKER_FAILED -> "failed";
|
||||
case BACKEND_EXHAUSTED -> "backend_exhausted";
|
||||
case REPLIED, COMPLETED_UNREPLIED, QUESTION, STALE_TURN -> "done"; // unreachable
|
||||
case REPLIED, COMPLETED_UNREPLIED, QUESTION, STALE_TURN, NOT_TURN_OWNER -> "done"; // unreachable
|
||||
},
|
||||
"detail", reply.outcome() == MessageService.Outcome.TIMED_OUT_UNCONFIRMED
|
||||
? "no reply within " + timeout + "ms; delivery is unconfirmed — the "
|
||||
@@ -817,10 +881,12 @@ public final class FleetApp {
|
||||
body.put("sessionId", id);
|
||||
body.put("status", messages.status(id).name().toLowerCase());
|
||||
body.put("ready", deliverable.test(id));
|
||||
// CB-582: a worker paused mid-turn in an async fleet_ask is otherwise invisible to a
|
||||
// status poll — surface the open question and how to answer it, same as fleet_poll's
|
||||
// Phase.ASKING view.
|
||||
MessageService.PendingAsk ask = messages.pendingAsk(id);
|
||||
// A worker paused mid-turn in an async fleet_ask is otherwise invisible to a status
|
||||
// poll — surface the open question and how to answer it, same as fleet_poll's
|
||||
// Phase.ASKING view, but only to the caller whose terminal created that delegation, or
|
||||
// to a caller with no terminal at all (the unnamed primary).
|
||||
Principal caller = ctx.attribute(CALLER);
|
||||
MessageService.PendingAsk ask = messages.pendingAsk(id, caller == null ? null : caller.terminal());
|
||||
if (ask != null) {
|
||||
body.put("question", ask.question());
|
||||
body.put("turnId", ask.turnId());
|
||||
@@ -837,7 +903,8 @@ public final class FleetApp {
|
||||
if (!allow(ctx, routeAction("GET /tasks/{ticket}"), null)) {
|
||||
return;
|
||||
}
|
||||
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"));
|
||||
Principal caller = ctx.attribute(CALLER);
|
||||
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"), caller == null ? null : caller.terminal());
|
||||
if (v == null) {
|
||||
ctx.status(404).json(Map.of("error", "unknown_ticket", "detail", "no such task (or it has expired)"));
|
||||
return;
|
||||
|
||||
@@ -26,6 +26,7 @@ import java.util.concurrent.atomic.AtomicBoolean;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.Consumer;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
* Authoritative in-daemon registry of the worker sessions this {@code fleetd} process spawned.
|
||||
@@ -57,6 +58,18 @@ public final class SessionManager implements TurnListener {
|
||||
* Populated on every spawn path, removed on {@link #release}.
|
||||
*/
|
||||
private final ConcurrentHashMap<String /*paneId*/, PeerHandle> handles = new ConcurrentHashMap<>();
|
||||
/**
|
||||
* fleetd #702: a pane mid-teardown, keyed by paneId, held from just before its registry entry
|
||||
* is removed until {@link #releaseRemoved} finishes. {@link #spawnedMemberRole} consults this
|
||||
* alongside the registry, so a caller resolving the pane's terminal during that window still
|
||||
* sees a live member and never falls through to a tab map.
|
||||
*
|
||||
* <p>Depth-counted rather than a plain set: two threads can be tearing down the same pane at
|
||||
* once (the CAS in {@link #releaseIfCurrent} exists for exactly that race), and with a set the
|
||||
* loser's {@code finally} would unmark the pane while the winner is still mid-teardown,
|
||||
* reopening the window this exists to close.
|
||||
*/
|
||||
private final ConcurrentHashMap<String /*paneId*/, Releasing> releasing = new ConcurrentHashMap<>();
|
||||
private final MemberPresence presence;
|
||||
private final SecureRandom nonceRandom = new SecureRandom();
|
||||
private final AtomicLong nonceSeq = new AtomicLong();
|
||||
@@ -242,6 +255,10 @@ public final class SessionManager implements TurnListener {
|
||||
handle.id(), handle.terminalId(), resolvedProfile, actualRole, cwd, ownerTerminal, now, now, 0,
|
||||
MemberSession.State.SPAWNING, null, null, handle.charterReceipt(), handle.agentSessionId());
|
||||
registry.put(handle.id(), session);
|
||||
// A presence contact that already arrived for this terminal found no registry
|
||||
// entry to transition and gave up silently. Retry it now that one exists; remove
|
||||
// this call and such a session stays in SPAWNING even though it is present.
|
||||
reconcilePresence(handle.terminalId());
|
||||
handles.put(handle.id(), handle);
|
||||
log.debug("acquired session id={} terminal={} profile={} owner={}",
|
||||
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
|
||||
@@ -303,9 +320,11 @@ public final class SessionManager implements TurnListener {
|
||||
* is a logged path an operator can reclaim, the cost of a deleted one is unrecoverable work.
|
||||
*/
|
||||
private MemberSession release(String paneId, ReleaseCause cause) {
|
||||
MemberSession removed = registry.remove(paneId);
|
||||
releaseRemoved(paneId, removed, handles.remove(paneId), cause);
|
||||
return removed;
|
||||
return releaseWindow(paneId, registry.get(paneId), () -> {
|
||||
MemberSession removed = registry.remove(paneId);
|
||||
releaseRemoved(paneId, removed, handles.remove(paneId), cause);
|
||||
return removed;
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -314,16 +333,85 @@ public final class SessionManager implements TurnListener {
|
||||
* DONE record from stopping a worker that delivery has made BUSY.
|
||||
*/
|
||||
private boolean releaseIfCurrent(MemberSession expected, ReleaseCause cause) {
|
||||
if (!registry.remove(expected.paneId(), expected)) {
|
||||
// A lifecycle transition replaced the record between the caller's check and this remove.
|
||||
// Log it: this race is by definition unobservable otherwise, and a reaper that silently
|
||||
// declines to reap is the hardest kind of behaviour to diagnose after the fact.
|
||||
log.debug("skipping reap of pane={}: its registry record changed after the idle check "
|
||||
+ "(most likely a delivery made it BUSY)", expected.paneId());
|
||||
return false;
|
||||
return releaseWindow(expected.paneId(), expected, () -> {
|
||||
if (!registry.remove(expected.paneId(), expected)) {
|
||||
// A lifecycle transition replaced the record between the caller's check and this
|
||||
// remove. Log it: this race is by definition unobservable otherwise, and a reaper
|
||||
// that silently declines to reap is the hardest kind of behaviour to diagnose
|
||||
// after the fact.
|
||||
log.debug("skipping reap of pane={}: its registry record changed after the idle "
|
||||
+ "check (most likely a delivery made it BUSY)", expected.paneId());
|
||||
return false;
|
||||
}
|
||||
releaseRemoved(expected.paneId(), expected, handles.remove(expected.paneId()), cause);
|
||||
return true;
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #702: mark {@code paneId} as mid-teardown — using {@code known}'s terminal/role when
|
||||
* it is available — for the whole of {@code teardown}, which removes the registry entry and
|
||||
* then runs {@link #releaseRemoved}. Shared by both registry-removal sites ({@link #release}'s
|
||||
* unconditional remove and {@link #releaseIfCurrent}'s CAS remove) so neither can leave the
|
||||
* other's window unmarked.
|
||||
*
|
||||
* <p>The mark is written before {@code teardown} runs — so it covers the removal itself, not
|
||||
* only what comes after it — and cleared in a {@code finally}, so an unchecked throw out of
|
||||
* {@code teardown} (including one from {@link PeerLauncher#stop}, which declares nothing) can
|
||||
* never leave the pane marked for the rest of the daemon's life.
|
||||
*/
|
||||
private <T> T releaseWindow(String paneId, MemberSession known, Supplier<T> teardown) {
|
||||
releasing.compute(paneId, (_, prior) -> Releasing.enter(prior, known));
|
||||
try {
|
||||
return teardown.get();
|
||||
} finally {
|
||||
releasing.compute(paneId, (_, prior) -> prior == null ? null : prior.leave());
|
||||
}
|
||||
releaseRemoved(expected.paneId(), expected, handles.remove(expected.paneId()), cause);
|
||||
return true;
|
||||
}
|
||||
|
||||
/**
|
||||
* Depth count plus the terminal/role a mid-teardown pane belongs to, for
|
||||
* {@link #spawnedMemberRole}. The terminal/role come from whichever call into
|
||||
* {@link #releaseWindow} first knew them: a call that finds the registry entry already gone
|
||||
* passes a {@code null} session, and must not blank out what the first call recorded.
|
||||
*/
|
||||
record Releasing(int depth, String terminalId, MemberRole role) {
|
||||
static Releasing enter(Releasing prior, MemberSession known) {
|
||||
int depth = (prior == null ? 0 : prior.depth()) + 1;
|
||||
String terminalId = known != null ? known.terminalId() : prior == null ? null : prior.terminalId();
|
||||
MemberRole role = known != null ? known.role() : prior == null ? null : prior.role();
|
||||
return new Releasing(depth, terminalId, role);
|
||||
}
|
||||
|
||||
Releasing leave() {
|
||||
return depth <= 1 ? null : new Releasing(depth - 1, terminalId, role);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The role of the live spawned member occupying {@code terminal} — whether it is currently in
|
||||
* the registry, or mid-teardown between {@link #release} removing its registry entry and
|
||||
* {@link #releaseRemoved} actually stopping its pane (fleetd #702). {@code null} for a terminal
|
||||
* that is neither: this method is the one reader a caller resolver consults before any tab
|
||||
* map, so a live or releasing member's identity never falls back to a tab label.
|
||||
*
|
||||
* <p>Checks the registry directly via {@link #findByTerminal} rather than {@link #roster()},
|
||||
* so this hot-path lookup (consulted on every resolve) never pays for a list copy or a stream.
|
||||
*/
|
||||
public MemberRole spawnedMemberRole(String terminal) {
|
||||
MemberSession session = findByTerminal(terminal);
|
||||
if (session != null) {
|
||||
return session.role();
|
||||
}
|
||||
if (terminal == null) {
|
||||
return null;
|
||||
}
|
||||
for (Releasing r : releasing.values()) {
|
||||
if (terminal.equals(r.terminalId())) {
|
||||
return r.role();
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private void releaseRemoved(String paneId, MemberSession removed, PeerHandle removedHandle,
|
||||
@@ -382,6 +470,12 @@ public final class SessionManager implements TurnListener {
|
||||
MemberSession resolved = resolveAgentSessionId(removed, removedHandle);
|
||||
notifyReleased(new ReleaseDetail(resolved.terminalId(), resolved.worktree(),
|
||||
resolved.branch(), snapshotRef, resolved.agentSessionId()));
|
||||
String terminal = removed.terminalId();
|
||||
if (terminal != null && !terminal.isBlank()) {
|
||||
// Without this, a terminal stays marked present after its pane is gone, so a
|
||||
// later send to the same id would read as deliverable instead of refused.
|
||||
presence.forget(terminal);
|
||||
}
|
||||
}
|
||||
}
|
||||
// CB-581: the pane must always stop, even if the dirty check above threw. A session removed
|
||||
@@ -725,6 +819,10 @@ public final class SessionManager implements TurnListener {
|
||||
handle.charterReceipt(),
|
||||
handle.agentSessionId());
|
||||
registry.put(handle.id(), session);
|
||||
// A presence contact that already arrived for this terminal found no registry entry to
|
||||
// transition and gave up silently. Retry it now that one exists; remove this call and
|
||||
// such a session stays in SPAWNING even though it is present.
|
||||
reconcilePresence(handle.terminalId());
|
||||
handles.put(handle.id(), handle);
|
||||
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
|
||||
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
|
||||
@@ -899,6 +997,20 @@ public final class SessionManager implements TurnListener {
|
||||
transitionByTerminal(terminalId, MemberSession.State.SPAWNING, MemberSession.State.READY);
|
||||
}
|
||||
|
||||
/**
|
||||
* Completes a newly registered session's {@code SPAWNING -> READY} transition when {@code
|
||||
* terminalId} was already marked present before this ran. A terminal never marked present is
|
||||
* left in {@code SPAWNING}; it reaches {@code READY} normally through {@link #onReady} once
|
||||
* its own contact arrives. Callers must run this only once the session's registry entry is
|
||||
* already visible — {@link #onReady}'s transition matches against that entry, and reconciling
|
||||
* before the entry exists finds nothing to transition.
|
||||
*/
|
||||
private void reconcilePresence(String terminalId) {
|
||||
if (terminalId != null && !terminalId.isBlank() && presence.isPresent(terminalId)) {
|
||||
onReady(terminalId);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Lifecycle hook: a message was delivered into the worker — it is now busy on a turn.
|
||||
* The turn count is bumped and the activity timestamp is refreshed. A {@code DONE} session
|
||||
|
||||
@@ -14,10 +14,12 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* CB-534: the injector's readiness gate must open for a lead as well as for a present worker.
|
||||
* fleetd #669 follow-up: the same gate must also open for a collaborator, which — like a lead —
|
||||
* is never enrolled in {@link MemberPresence} and never discovered by the lead scan.
|
||||
*
|
||||
* <p>The bug these cover was silent and slow: a lead was never marked present (only workers are), so
|
||||
* every lead→lead delivery sat on the gate for the full readiness grace and failed ~60s later without
|
||||
* a keystroke ever reaching the pane.
|
||||
* <p>The bug these cover was silent and slow: a lead (and later a collaborator) was never marked
|
||||
* present (only workers are) and never counted as a lead, so every send to one sat on the gate for
|
||||
* the full readiness grace and failed ~60s later without a keystroke ever reaching the pane.
|
||||
*/
|
||||
class FleetDeliverabilityTest {
|
||||
|
||||
@@ -25,19 +27,25 @@ class FleetDeliverabilityTest {
|
||||
return () -> m;
|
||||
}
|
||||
|
||||
private static Supplier<Map<String, String>> collaborators(Map<String, String> m) {
|
||||
return () -> m;
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a worker that has connected its MCP is deliverable")
|
||||
void presentWorkerIsDeliverable() {
|
||||
MemberPresence presence = new MemberPresence();
|
||||
presence.markPresent("term_worker");
|
||||
|
||||
assertTrue(Fleetd.deliverableTo(presence, leads(Map.of())).test("term_worker"));
|
||||
assertTrue(Fleetd.deliverableTo(presence, leads(Map.of()), collaborators(Map.of()))
|
||||
.test("term_worker"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a worker still in its boot window is held back")
|
||||
void absentWorkerIsNotDeliverable() {
|
||||
assertFalse(Fleetd.deliverableTo(new MemberPresence(), leads(Map.of())).test("term_booting"));
|
||||
assertFalse(Fleetd.deliverableTo(new MemberPresence(), leads(Map.of()), collaborators(Map.of()))
|
||||
.test("term_booting"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -45,19 +53,31 @@ class FleetDeliverabilityTest {
|
||||
void leadIsDeliverableWithoutPresence() {
|
||||
MemberPresence presence = new MemberPresence();
|
||||
Predicate<String> deliverable =
|
||||
Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")));
|
||||
Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")), collaborators(Map.of()));
|
||||
|
||||
assertFalse(presence.isPresent("term_lead"), "a lead is never enrolled in worker presence");
|
||||
assertTrue(deliverable.test("term_lead"), "…and must be deliverable anyway");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unknown terminal is deliverable to neither")
|
||||
@DisplayName("a collaborator is deliverable without ever being marked present or scanned as a lead")
|
||||
void collaboratorIsDeliverableWithoutPresenceOrLeadStatus() {
|
||||
MemberPresence presence = new MemberPresence();
|
||||
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads(Map.of()),
|
||||
collaborators(Map.of("term_collab", "kevin")));
|
||||
|
||||
assertFalse(presence.isPresent("term_collab"), "a collaborator is never enrolled in worker presence");
|
||||
assertTrue(deliverable.test("term_collab"), "…and must be deliverable anyway");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unknown terminal is deliverable to none of presence, leads, or collaborators")
|
||||
void strangerIsNotDeliverable() {
|
||||
MemberPresence presence = new MemberPresence();
|
||||
presence.markPresent("term_worker");
|
||||
|
||||
assertFalse(Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")))
|
||||
assertFalse(Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")),
|
||||
collaborators(Map.of("term_collab", "kevin")))
|
||||
.test("term_stranger"));
|
||||
}
|
||||
|
||||
@@ -65,22 +85,47 @@ class FleetDeliverabilityTest {
|
||||
@DisplayName("a lead discovered after startup becomes deliverable with no restart")
|
||||
void leadSetIsReadThroughOnEveryCall() {
|
||||
Map<String, String> discovered = new HashMap<>();
|
||||
Predicate<String> deliverable = Fleetd.deliverableTo(new MemberPresence(), leads(discovered));
|
||||
Predicate<String> deliverable =
|
||||
Fleetd.deliverableTo(new MemberPresence(), leads(discovered), collaborators(Map.of()));
|
||||
|
||||
assertFalse(deliverable.test("term_late"));
|
||||
discovered.put("term_late", "gpt-sol-5.6"); // leadScan picks up a newly labelled tab
|
||||
assertTrue(deliverable.test("term_late"), "the supplier must be re-read, not snapshotted");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a collaborator discovered after startup becomes deliverable with no restart")
|
||||
void collaboratorSetIsReadThroughOnEveryCall() {
|
||||
Map<String, String> discovered = new HashMap<>();
|
||||
Predicate<String> deliverable =
|
||||
Fleetd.deliverableTo(new MemberPresence(), leads(Map.of()), collaborators(discovered));
|
||||
|
||||
assertFalse(deliverable.test("term_late_collab"));
|
||||
discovered.put("term_late_collab", "kevin"); // the same tab scan picks up a newly labelled collaborator tab
|
||||
assertTrue(deliverable.test("term_late_collab"), "the supplier must be re-read, not snapshotted");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("forgetting a torn-down worker does not strip a lead of its deliverability")
|
||||
void forgetDoesNotDisarmALead() {
|
||||
MemberPresence presence = new MemberPresence();
|
||||
Predicate<String> deliverable =
|
||||
Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")));
|
||||
Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")), collaborators(Map.of()));
|
||||
|
||||
presence.forget("term_lead"); // the injector's cleanup path runs against every target
|
||||
|
||||
assertTrue(deliverable.test("term_lead"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("forgetting a torn-down worker does not strip a collaborator of its deliverability")
|
||||
void forgetDoesNotDisarmACollaborator() {
|
||||
MemberPresence presence = new MemberPresence();
|
||||
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads(Map.of()),
|
||||
collaborators(Map.of("term_collab", "kevin")));
|
||||
|
||||
presence.forget("term_collab"); // the injector's cleanup path runs against every target
|
||||
|
||||
assertTrue(deliverable.test("term_collab"));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,166 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.auth.Authz;
|
||||
import dev.ltms.fleet.auth.Principal;
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Method;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* Asserts that the {@link FleetMcp} built by {@link FleetdAssembly#assembleAndStart} applies the
|
||||
* authorization table: a worker is refused {@code SPAWN}, and the primary is allowed it.
|
||||
*/
|
||||
class FleetdAssemblyAuthorizationModeTest {
|
||||
|
||||
private static final class TestResourcePorts implements ResourcePorts {
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> new ReplyInbox() {
|
||||
@Override public void own(String target) { }
|
||||
@Override public void release(String target) { }
|
||||
@Override public void publish(String target, String msgId, String content) { }
|
||||
@Override public List<InboxMessage> peek(String target) { return List.of(); }
|
||||
@Override public boolean ack(String target, String msgId) { return false; }
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Binding a real port would clash with any daemon already listening on it.
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private TestResourcePorts ports;
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
if (ports != null && ports.shutdownHook != null) {
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
health:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
""");
|
||||
return FleetConfig.load(file);
|
||||
}
|
||||
|
||||
private FleetMcp assemble(Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ports = new TestResourcePorts();
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
|
||||
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
|
||||
return runtime.mcp();
|
||||
}
|
||||
|
||||
/**
|
||||
* Invokes {@code FleetMcp#denyFor}, which is package-private to {@code dev.ltms.fleet.mcp}
|
||||
* while this test is in {@code dev.ltms.fleet}. Nothing here catches a missing method: if
|
||||
* {@code denyFor} is renamed or removed, {@link NoSuchMethodException} propagates and the
|
||||
* test fails.
|
||||
*/
|
||||
private static McpSchema.CallToolResult denyFor(FleetMcp mcp, Principal caller, Authz.Action action,
|
||||
String target) throws Exception {
|
||||
Method m = FleetMcp.class.getDeclaredMethod("denyFor", Principal.class, Authz.Action.class, String.class);
|
||||
m.setAccessible(true);
|
||||
return (McpSchema.CallToolResult) m.invoke(mcp, caller, action, target);
|
||||
}
|
||||
|
||||
@Test
|
||||
void productionBootPathRefusesAnUnauthorizedCallerThroughTheAssembledFleetMcp(@TempDir Path dir)
|
||||
throws Exception {
|
||||
FleetMcp mcp = assemble(dir);
|
||||
|
||||
McpSchema.CallToolResult deniedForWorker = denyFor(mcp, Principal.worker("term_a", 200),
|
||||
Authz.Action.SPAWN, "term_a");
|
||||
assertNotNull(deniedForWorker,
|
||||
"a worker must not be able to fleet_spawn through the assembled FleetMcp");
|
||||
assertTrue(deniedForWorker.isError(), "a refusal is returned as an MCP tool error");
|
||||
|
||||
McpSchema.CallToolResult allowedForPrimary = denyFor(mcp, Principal.primary(100),
|
||||
Authz.Action.SPAWN, "term_a");
|
||||
assertNull(allowedForPrimary,
|
||||
"control: the primary must still be allowed to fleet_spawn — otherwise the worker "
|
||||
+ "refusal above would pass even with the gate wired backwards");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,165 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertSame;
|
||||
|
||||
/**
|
||||
* fleetd #669 Unit E. Reaches the real {@link dev.ltms.fleet.herdr.HerdrRouter} that {@link
|
||||
* FleetdAssembly#assembleAndStart} builds and wires — not a copy built for this test — and proves
|
||||
* that a configured collaborator's terminal routes to the LEAD herdr daemon.
|
||||
*
|
||||
* <p>Two distinct {@link FakeHerdr} instances are required, the same pattern {@code
|
||||
* FleetdAssemblyConnectionIdentityTest} and {@code FleetdLeadRolloverAssemblyTest} already use:
|
||||
* with one client shared between {@code herdrSocket} and {@code memberHerdrSocket},
|
||||
* {@code HerdrRouter} folds {@code leadAgents} and {@code memberAgents} into the same instance
|
||||
* (see its constructor), and {@code agentsFor} would return that one object regardless of whether
|
||||
* the collaborator map was ever consulted — invisible to a mutation of the predicate this ticket
|
||||
* fixes. This test's two sockets resolve to two different fakes, so the assertion only passes when
|
||||
* the collaborator's terminal is actually recognised and routed to the lead one.
|
||||
*/
|
||||
class FleetdAssemblyCollaboratorHerdrRoutingTest {
|
||||
|
||||
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
|
||||
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
|
||||
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
HerdrClient client = herdrsBySocket.get(socketPath);
|
||||
if (client == null) {
|
||||
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
|
||||
}
|
||||
return client;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> new ReplyInbox() {
|
||||
@Override public void own(String target) { }
|
||||
@Override public void release(String target) { }
|
||||
@Override public void publish(String target, String msgId, String content) { }
|
||||
@Override public List<InboxMessage> peek(String target) { return List.of(); }
|
||||
@Override public boolean ack(String target, String msgId) { return false; }
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Do not bind a real port in this assembly test.
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: "%s"
|
||||
memberHerdrSocket: "%s"
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
fleet:
|
||||
collaborators:
|
||||
reviewer-alex:
|
||||
tab: "collab: alex"
|
||||
profiles:
|
||||
sonnet:
|
||||
subscription: true
|
||||
argv: ["ccs", "sonnet"]
|
||||
""".formatted(LEAD_SOCKET, MEMBER_SOCKET));
|
||||
return FleetConfig.load(file);
|
||||
}
|
||||
|
||||
@Test
|
||||
void assembledRouterRoutesACollaboratorTerminalToTheLeadDaemon(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
|
||||
// The fixed FakeHerdr fixture already ties terminal "term_a" to a live agent on tab
|
||||
// "w2:t7" (pane "w2:p7") — seeding only the tab LABEL to match the configured collaborator
|
||||
// is enough to make LeadTabScanner resolve "term_a" as that collaborator. Seeded on the
|
||||
// LEAD fake only: a collaborator's pane lives in the lead daemon, exactly like a lead's.
|
||||
FakeHerdr lead = new FakeHerdr().withTab("w2", "w2:t7", "collab: alex");
|
||||
FakeHerdr member = new FakeHerdr();
|
||||
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
|
||||
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
|
||||
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
|
||||
try {
|
||||
assertSame(runtime.router().leadAgents(), runtime.router().agentsFor("term_a"),
|
||||
"a configured collaborator's terminal must route to the LEAD daemon — "
|
||||
+ "FleetdAssembly must wire the collaborator map into the router's "
|
||||
+ "predicate, not just LeadTabScanner.get()");
|
||||
assertSame(runtime.router().memberAgents(), runtime.router().agentsFor("term_shell"),
|
||||
"control: a terminal naming neither a lead nor a collaborator (term_shell, on "
|
||||
+ "the unlabelled tab w2:t8) must still route to the member daemon");
|
||||
} finally {
|
||||
assertNotNull(ports.shutdownHook, "control: assembly must capture its shutdown hook");
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,200 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.LeadTabScanner;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadCoordLoop;
|
||||
import dev.ltms.fleet.msg.LeadMessage;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Field;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertInstanceOf;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #670 — pins the {@code excludedWorkspaceLabels} argument {@link FleetdAssembly}'s
|
||||
* production boot path passes to {@link LeadTabScanner} ({@code Set.of()}).
|
||||
*
|
||||
* <p>{@code LeadTabScannerTest} already covers this constructor parameter, but it builds its own
|
||||
* {@link LeadTabScanner} with its own set, so it tests the seam and proves nothing about the
|
||||
* producer. This test instead reaches the exact object {@link FleetdAssembly#assembleAndStart}
|
||||
* builds: a {@code fleet.leaders:} block makes the assembly construct a real
|
||||
* {@link LeadTabScanner} for its local {@code leads} supplier, and a {@code coordinator:} block
|
||||
* makes it hand that same supplier instance to {@link LeadCoordLoop} (fleetd #637), which stores
|
||||
* it as a field. Reflection recovers it from there, and then from the scanner itself, so the
|
||||
* assertion is against the real production argument rather than a copy built for this test.
|
||||
*/
|
||||
class FleetdAssemblyLeadTabScannerExclusionTest {
|
||||
|
||||
private static final class FakeLeadChannel implements LeadChannelHandle {
|
||||
@Override
|
||||
public void publish(String toCoordId, LeadMessage message) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<LeadMessage> peek() {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void ack(String msgId) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public String selfCoordId() {
|
||||
return "test-lead";
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean heldDurable() {
|
||||
return true;
|
||||
}
|
||||
|
||||
@Override
|
||||
public MailboxState inspect(String coordId) {
|
||||
return MailboxState.unknown(coordId);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static final class TestResourcePorts implements ResourcePorts {
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> new ReplyInbox() {
|
||||
@Override public void own(String target) { }
|
||||
@Override public void release(String target) { }
|
||||
@Override public void publish(String target, String msgId, String content) { }
|
||||
@Override public List<InboxMessage> peek(String target) { return List.of(); }
|
||||
@Override public boolean ack(String target, String msgId) { return false; }
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> new FakeLeadChannel();
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Do not bind a real port in this assembly test.
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
coordinator:
|
||||
uri: "amqp://fake-lead-broker/vh"
|
||||
selfId: "test-lead"
|
||||
fleet:
|
||||
leaders:
|
||||
primary:
|
||||
tab: "lead: primary"
|
||||
profile: sonnet
|
||||
profiles:
|
||||
sonnet:
|
||||
subscription: true
|
||||
argv: ["ccs", "sonnet"]
|
||||
""");
|
||||
return FleetConfig.load(file);
|
||||
}
|
||||
|
||||
@Test
|
||||
void productionBootPathPassesNoExcludedWorkspaceLabels(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
TestResourcePorts ports = new TestResourcePorts();
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
|
||||
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
|
||||
try {
|
||||
LeadCoordLoop coordLoop = runtime.leadCoordLoop();
|
||||
assertNotNull(coordLoop, "control: a configured coordinator: block must build LeadCoordLoop");
|
||||
|
||||
Field leadsField = LeadCoordLoop.class.getDeclaredField("leads");
|
||||
leadsField.setAccessible(true);
|
||||
@SuppressWarnings("unchecked")
|
||||
Supplier<Map<String, String>> leads = (Supplier<Map<String, String>>) leadsField.get(coordLoop);
|
||||
|
||||
assertInstanceOf(LeadTabScanner.class, leads,
|
||||
"control: a non-empty fleet.leaders: block must make FleetdAssembly build a real "
|
||||
+ "LeadTabScanner for its `leads` supplier, not the Map::of fallback — "
|
||||
+ "otherwise this test would pass for the wrong reason");
|
||||
|
||||
Field excludedField = LeadTabScanner.class.getDeclaredField("excludedWorkspaceLabels");
|
||||
excludedField.setAccessible(true);
|
||||
Set<?> excluded = (Set<?>) excludedField.get(leads);
|
||||
|
||||
assertTrue(excluded.isEmpty(),
|
||||
"FleetdAssembly must pass an empty excludedWorkspaceLabels to "
|
||||
+ "LeadTabScanner — scanning member tabs would demote the lead to a worker");
|
||||
} finally {
|
||||
assertNotNull(ports.shutdownHook, "control: assembly must capture its shutdown hook");
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,190 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.inject.StatusPoller;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadCoordLoop;
|
||||
import dev.ltms.fleet.msg.LeadMessage;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import dev.ltms.fleet.msg.ReplyPushLoop;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Field;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
|
||||
/**
|
||||
* Asserts that the assembled loops use the production reminder, coordination, and delivery timing
|
||||
* defaults when no {@code primary:} block configures the reply-push values.
|
||||
*/
|
||||
class FleetdAssemblyTimingDefaultsTest {
|
||||
|
||||
private static final class FakeLeadChannel implements LeadChannelHandle {
|
||||
@Override
|
||||
public void publish(String toCoordId, LeadMessage message) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<LeadMessage> peek() {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void ack(String msgId) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public String selfCoordId() {
|
||||
return "test-lead";
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean heldDurable() {
|
||||
return true;
|
||||
}
|
||||
|
||||
@Override
|
||||
public MailboxState inspect(String coordId) {
|
||||
return MailboxState.unknown(coordId);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static final class TestResourcePorts implements ResourcePorts {
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> new ReplyInbox() {
|
||||
@Override public void own(String target) { }
|
||||
@Override public void release(String target) { }
|
||||
@Override public void publish(String target, String msgId, String content) { }
|
||||
@Override public List<InboxMessage> peek(String target) { return List.of(); }
|
||||
@Override public boolean ack(String target, String msgId) { return false; }
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> new FakeLeadChannel();
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private TestResourcePorts ports;
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
if (ports != null && ports.shutdownHook != null) {
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
coordinator:
|
||||
uri: "amqp://fake-lead-broker/vh"
|
||||
selfId: "test-lead"
|
||||
""");
|
||||
return FleetConfig.load(file);
|
||||
}
|
||||
|
||||
private FleetdRuntime assemble(Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ports = new TestResourcePorts();
|
||||
return FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
|
||||
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
|
||||
}
|
||||
|
||||
private static long longField(Object target, String name) throws Exception {
|
||||
Field field = target.getClass().getDeclaredField(name);
|
||||
field.setAccessible(true);
|
||||
return field.getLong(target);
|
||||
}
|
||||
|
||||
@Test
|
||||
void productionBootPathUsesTheExpectedLoopTimingDefaults(@TempDir Path dir) throws Exception {
|
||||
FleetdRuntime runtime = assemble(dir);
|
||||
|
||||
ReplyPushLoop pushLoop = runtime.pushLoop();
|
||||
assertEquals(5, longField(pushLoop, "maxReminders"),
|
||||
"without primary:, ReplyPushLoop must stop after five reminder attempts");
|
||||
assertEquals(15_000L, longField(pushLoop, "backoffMs"),
|
||||
"without primary:, ReplyPushLoop must wait fifteen seconds before the next reminder");
|
||||
|
||||
LeadCoordLoop leadCoordLoop = runtime.leadCoordLoop();
|
||||
assertNotNull(leadCoordLoop, "control: coordinator: must build LeadCoordLoop");
|
||||
assertEquals(3_000L, longField(leadCoordLoop, "intervalMs"),
|
||||
"LeadCoordLoop must poll for peer-lead mail every three seconds");
|
||||
|
||||
StatusPoller poller = runtime.poller();
|
||||
assertEquals(Injector.POLL_INTERVAL_MILLIS, longField(poller, "intervalMillis"),
|
||||
"StatusPoller must use Injector's delivery poll interval");
|
||||
}
|
||||
}
|
||||
@@ -6,6 +6,8 @@ import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.lead.LeadLauncher;
|
||||
import dev.ltms.fleet.lead.LeadRollover;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
@@ -34,7 +36,7 @@ import static org.junit.jupiter.api.Assertions.fail;
|
||||
* fleetd #612 Unit A) with three methods: {@code unrelatedAnchorStillPresent} (a scaffold anchor,
|
||||
* not an independent claim — needs no replacement of its own), {@code
|
||||
* mainStillCallsTheLeadRolloverFactory} (the call-site pin replaced by {@link
|
||||
* #assembledLeadRolloverRunsTheRealClearAndBootstrapSequence}), and {@code
|
||||
* #assembledLeadRolloverEndsTheOldPaneThroughTheRealHerdrRouter}), and {@code
|
||||
* factoryGatesOnConfigPresence} (the absent-config claim replaced by {@link
|
||||
* #absentLeadRolloverConfigMeansNoRolloverIsBuilt} — a claim this ticket found was NOT actually
|
||||
* covered behaviourally anywhere else: {@code LeadRolloverTest}'s only related assertion is
|
||||
@@ -50,8 +52,8 @@ import static org.junit.jupiter.api.Assertions.fail;
|
||||
* invisible to this test, even though the two are genuinely different daemons in production. This
|
||||
* version configures two distinct sockets and two distinct {@link FakeHerdr} instances (the same
|
||||
* pattern {@code FleetdAssemblyConnectionIdentityTest}, fleetd #612 B2, already uses to separate
|
||||
* lead from member) and asserts the roll's {@code /clear}/bootstrap sends land on the LEAD fake
|
||||
* and never on the MEMBER one.
|
||||
* lead from member) and asserts the roll's {@code pane.close} call lands on the LEAD fake and
|
||||
* never on the MEMBER one.
|
||||
*/
|
||||
class FleetdLeadRolloverAssemblyTest {
|
||||
|
||||
@@ -177,11 +179,47 @@ class FleetdLeadRolloverAssemblyTest {
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
/**
|
||||
* Unlike {@link #writeConfig}, this names a {@code profile:} for the lead and declares it
|
||||
* under {@code profiles:}, so {@code LeadLauncher#relaunch} can actually start a fresh agent
|
||||
* instead of refusing with "names no profile". {@code relaunchReadySeconds} is cut to 2s so
|
||||
* the recognition wait (expected to time out — see the test) does not cost real test seconds.
|
||||
*/
|
||||
private static FleetConfig writeConfigWithRelaunchableLead(Path dir, Path leadCwd) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: "%s"
|
||||
memberHerdrSocket: "%s"
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
cwd: "%s"
|
||||
profile: opus
|
||||
profiles:
|
||||
opus:
|
||||
subscription: true
|
||||
argv: ["ccs", "opus"]
|
||||
leadRollover:
|
||||
handoverPath: handover.md
|
||||
requireOperatorConfirm: false
|
||||
relaunchReadySeconds: 2
|
||||
""".formatted(LEAD_SOCKET, MEMBER_SOCKET, leadCwd.toString()));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled LeadRollover runs the full open/confirm/continuation "
|
||||
+ "sequence — /clear, then bootstrapText — through the real herdr router")
|
||||
void assembledLeadRolloverRunsTheRealClearAndBootstrapSequence(@TempDir Path dir) throws Exception {
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled LeadRollover runs the open/confirm/continuation "
|
||||
+ "sequence through the real herdr router — ending the old pane, then giving up once it "
|
||||
+ "never reports gone")
|
||||
void assembledLeadRolloverEndsTheOldPaneThroughTheRealHerdrRouter(@TempDir Path dir) throws Exception {
|
||||
Path leadCwd = dir.resolve("lead-workspace");
|
||||
Files.createDirectories(leadCwd);
|
||||
FleetConfig cfg = writeConfig(dir, leadCwd);
|
||||
@@ -204,6 +242,11 @@ class FleetdLeadRolloverAssemblyTest {
|
||||
+ "Fleetd.leadRollover(...) call site — a mutation to `LeadRollover leadRollover = "
|
||||
+ "null;` at that call site can never pass this");
|
||||
|
||||
// assembleAndStart's own boot work (the orphan-worker reap) makes a real call on the
|
||||
// member daemon before the roll ever starts. Clear it here so the assertion below measures
|
||||
// only what the roll itself does, not what daemon startup does.
|
||||
member.calls.clear();
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open("term_a", "fleetd #612 B3 test");
|
||||
String expectedHandoverPath = leadCwd.resolve("handover.md").normalize().toString();
|
||||
assertEquals(expectedHandoverPath, pending.handoverPath());
|
||||
@@ -223,40 +266,109 @@ class FleetdLeadRolloverAssemblyTest {
|
||||
// constructor), so this polls the real FleetMcp.leadRollover() instance's status(token)
|
||||
// until the real continuation finishes.
|
||||
LeadRollover.RollStatus status = pollUntilTerminal(rollover, pending.token());
|
||||
assertEquals(LeadRollover.RollState.ROLLED, status.state(),
|
||||
"the full happy path must complete: FakeHerdr's default agent status is 'idle', so "
|
||||
+ "the turn-boundary wait settles immediately and the post-/clear wait "
|
||||
+ "releases via its pickup-grace path — detail: " + status.detail());
|
||||
|
||||
// Prove the real herdr router actually sent BOTH messages, in order, to the real LEAD
|
||||
// pane — this is the one thing a source-text pin on the call site could never show.
|
||||
// FakeHerdr's pane.get is a fixed canned response that never reports a pane as gone, so the
|
||||
// real router's death poll runs out its whole budget and the roll stops here — proving the
|
||||
// real teardown call landed on the real LEAD pane without ever reaching a relaunch or a send.
|
||||
assertEquals(LeadRollover.RollState.OLD_PANE_NEVER_DIED, status.state(),
|
||||
"the old pane never reports gone against this fake, so the roll must stop with "
|
||||
+ "OLD_PANE_NEVER_DIED rather than ever relaunching or sending anything — "
|
||||
+ "detail: " + status.detail());
|
||||
|
||||
// Prove the real herdr router actually closed the real LEAD pane — this is the one thing a
|
||||
// source-text pin on the call site could never show.
|
||||
boolean closedOldPane = lead.calls.stream()
|
||||
.anyMatch(c -> c.method().equals("pane.close")
|
||||
&& c.params() instanceof Map<?, ?> m && "w2:p7".equals(m.get("pane_id")));
|
||||
assertTrue(closedOldPane, "endOldSession must close the real old pane (w2:p7) through the "
|
||||
+ "real LEAD herdr client, got calls: " + lead.calls);
|
||||
|
||||
// No agent.prompt is ever sent on this path: the roll stops at the pane-death wait, strictly
|
||||
// before the relaunch and the final send step.
|
||||
List<FakeHerdr.Call> prompts = lead.calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt"))
|
||||
.toList();
|
||||
assertTrue(prompts.size() >= 2, "expected at least a /clear send and a bootstrapText send "
|
||||
+ "on the LEAD daemon, got " + prompts.size() + " agent.prompt calls: " + prompts);
|
||||
assertEquals("/clear", ((Map<String, Object>) prompts.get(0).params()).get("text"),
|
||||
"the first send must be the literal /clear housekeeping command");
|
||||
Object secondText = ((Map<String, Object>) prompts.get(1).params()).get("text");
|
||||
assertTrue(secondText instanceof String && ((String) secondText).contains(expectedHandoverPath),
|
||||
"the second send must be the default bootstrapText naming the resolved handover "
|
||||
+ "path, got: " + secondText);
|
||||
assertTrue(prompts.isEmpty(), "a roll that stops at OLD_PANE_NEVER_DIED must never reach the "
|
||||
+ "send step, got agent.prompt call(s) on the LEAD daemon: " + prompts);
|
||||
|
||||
// fleetd #612 B3 correction: prove the roll never touches the MEMBER daemon. A mutation
|
||||
// swapping router.leadAgents() for router.memberAgents() at the real call site would move
|
||||
// both sends above onto `member` instead, which this assertion catches — the thing the
|
||||
// single-fake version of this test could never see, because both wrapped the same client.
|
||||
// the pane.close call above onto `member` instead, which this assertion catches — the thing
|
||||
// the single-fake version of this test could never see, because both wrapped the same client.
|
||||
assertTrue(member.calls.isEmpty(), "the roll must be wired to the LEAD daemon only — got "
|
||||
+ member.calls.size() + " call(s) recorded on the MEMBER daemon since the roll began: "
|
||||
+ member.calls);
|
||||
}
|
||||
|
||||
/**
|
||||
* Exercises the relaunch site {@link #assembledLeadRolloverEndsTheOldPaneThroughTheRealHerdrRouter}
|
||||
* never reaches: with the old pane confirmed gone, the roll relaunches a fresh lead, and
|
||||
* {@code bootstrapText} must reach it even though recognition times out (FakeHerdr's
|
||||
* {@code tab.list} is a fixed canned response that never reflects the relaunch's own
|
||||
* {@code tab.rename}, so the fresh terminal is never recognised as a live lead). Same
|
||||
* dual-socket shape as the sibling test: two distinct {@link FakeHerdr} instances, so a
|
||||
* {@code bootstrapText} send wired to the wrong daemon is visible.
|
||||
*/
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] bootstrapText reaches the fresh LEAD terminal even when recognition "
|
||||
+ "times out, and the MEMBER daemon never sees it")
|
||||
void bootstrapTextReachesTheFreshLeadTerminalEvenWhenRecognitionTimesOut(@TempDir Path dir) throws Exception {
|
||||
Path leadCwd = dir.resolve("lead-workspace");
|
||||
Files.createDirectories(leadCwd);
|
||||
FleetConfig cfg = writeConfigWithRelaunchableLead(dir, leadCwd);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
FakeHerdr lead = new FakeHerdr();
|
||||
lead.withTab("w2", "w2:t7", "lead: opus");
|
||||
// Lets the old pane (w2:p7) report gone once pane.close actually reaches it, so the roll
|
||||
// proceeds to relaunch instead of stopping at OLD_PANE_NEVER_DIED.
|
||||
lead.paneGoneAfterClose("w2:p7");
|
||||
FakeHerdr member = new FakeHerdr();
|
||||
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
|
||||
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
|
||||
LeadRollover rollover = runtime.mcp().leadRollover();
|
||||
assertNotNull(rollover, "leadRollover: is present in this test's config, so a real "
|
||||
+ "LeadRollover must have been built");
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open("term_a", "bootstrapText relaunch test");
|
||||
Thread.sleep(50);
|
||||
Files.writeString(Path.of(pending.handoverPath()), "handover content for bootstrapText test");
|
||||
|
||||
LeadRollover.RollDecision decision = rollover.confirm("term_a", pending.token(), true);
|
||||
assertTrue(decision.accepted(), "confirm() must approve — got: " + decision);
|
||||
|
||||
LeadRollover.RollStatus status = pollUntilTerminal(rollover, pending.token());
|
||||
assertEquals(LeadRollover.RollState.RELAUNCH_NOT_RECOGNISED, status.state(),
|
||||
"the fresh terminal is never recognised against this fake's static tab.list, so the "
|
||||
+ "roll must reach RELAUNCH_NOT_RECOGNISED — not an earlier failure state and "
|
||||
+ "not ROLLED — detail: " + status.detail());
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
List<FakeHerdr.Call> leadPrompts = lead.calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt"))
|
||||
.toList();
|
||||
assertEquals(1, leadPrompts.size(), "exactly one bootstrapText send is expected, on the LEAD "
|
||||
+ "daemon, once recognition gives up — got: " + leadPrompts);
|
||||
Object text = ((Map<String, Object>) leadPrompts.get(0).params()).get("text");
|
||||
assertTrue(text instanceof String && ((String) text).contains("handover.md"),
|
||||
"the send must be bootstrapText naming the resolved handover path, got: " + text);
|
||||
|
||||
// Scoped to agent.prompt specifically, not every MEMBER call: the orphan-worker reap also
|
||||
// talks to the MEMBER daemon once, unconditionally, at daemon boot — unrelated to this roll.
|
||||
List<FakeHerdr.Call> memberPrompts = member.calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt"))
|
||||
.toList();
|
||||
assertTrue(memberPrompts.isEmpty(), "the roll must be wired to the LEAD daemon only — got "
|
||||
+ memberPrompts.size() + " agent.prompt call(s) on the MEMBER daemon instead: "
|
||||
+ memberPrompts);
|
||||
assertTrue(memberPrompts.isEmpty(), "bootstrapText must never be sent to the MEMBER daemon, "
|
||||
+ "got: " + memberPrompts);
|
||||
}
|
||||
|
||||
private static LeadRollover.RollStatus pollUntilTerminal(LeadRollover rollover, String token)
|
||||
throws InterruptedException {
|
||||
long deadline = System.nanoTime() + java.util.concurrent.TimeUnit.SECONDS.toNanos(10);
|
||||
long deadline = System.nanoTime() + java.util.concurrent.TimeUnit.SECONDS.toNanos(15);
|
||||
while (System.nanoTime() < deadline) {
|
||||
LeadRollover.RollStatus status = rollover.status(token);
|
||||
if (status.state() != LeadRollover.RollState.PENDING
|
||||
@@ -283,8 +395,10 @@ class FleetdLeadRolloverAssemblyTest {
|
||||
""");
|
||||
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
|
||||
AgentControl agents = new AgentControl(new FakeHerdr());
|
||||
WorkspaceControl spaces = new WorkspaceControl(new FakeHerdr());
|
||||
LeadLauncher launcher = new LeadLauncher(agents, spaces, config.get());
|
||||
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, config, Map::of);
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, spaces, launcher, config, Map::of);
|
||||
|
||||
assertNull(rollover, "leadRollover: is absent from this config, so the factory's opt-in "
|
||||
+ "gate (`if (cfg.leadRollover() == null) return null;`) must fire and no "
|
||||
|
||||
@@ -4,6 +4,8 @@ import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.lead.LeadLauncher;
|
||||
import dev.ltms.fleet.lead.LeadRollover;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
@@ -54,6 +56,11 @@ class FleetdLeadRolloverWorkspaceLookupTest {
|
||||
return new AgentControl(new FakeHerdr());
|
||||
}
|
||||
|
||||
/** None of this class's tests reach the deferred continuation, so a plain fake is enough. */
|
||||
private static LeadLauncher fakeLauncher(FleetConfig cfg) {
|
||||
return new LeadLauncher(fakeAgents(), new WorkspaceControl(new FakeHerdr()), cfg);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] Fleetd.leadRollover(...) resolves a relative handoverPath against "
|
||||
+ "the CALLING lead's configured cwd, not the daemon's own working directory")
|
||||
@@ -74,7 +81,8 @@ class FleetdLeadRolloverWorkspaceLookupTest {
|
||||
""".formatted(leadCwd.toString()));
|
||||
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
|
||||
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
|
||||
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config,
|
||||
() -> Map.of("term_opus", "opus"));
|
||||
assertNotNull(rollover, "leadRollover: is present in the loaded config, so the factory "
|
||||
+ "must construct an object");
|
||||
@@ -105,7 +113,8 @@ class FleetdLeadRolloverWorkspaceLookupTest {
|
||||
|
||||
// No lead has been discovered yet — exactly the real shape of a lead the live tab scan
|
||||
// has not yet scanned, or one with no fleet.leaders entry at all.
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config, Map::of);
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
|
||||
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config, Map::of);
|
||||
assertNotNull(rollover);
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open("term_unknown", "test");
|
||||
@@ -144,7 +153,8 @@ class FleetdLeadRolloverWorkspaceLookupTest {
|
||||
// below — exactly the natural mistake to make, since leads are discovered by a live tab
|
||||
// scan that runs AFTER this factory is constructed at startup.
|
||||
Map<String, String> liveLeadTerminals = new HashMap<>();
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
|
||||
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config,
|
||||
() -> liveLeadTerminals);
|
||||
assertNotNull(rollover);
|
||||
|
||||
|
||||
@@ -14,6 +14,8 @@ class AuthzTest {
|
||||
private static final Principal ANON = Principal.anonymous();
|
||||
private static final Principal ARCH_DESIGN = Principal.architect("lead-designer", "term_design", 400);
|
||||
private static final Principal ARCH_OTHER = Principal.architect("reviewer", "term_review", 500);
|
||||
private static final Principal COLLABORATOR = Principal.collaborator("ops", "term_collab", 600);
|
||||
private static final Principal OBSERVER = Principal.observer("term_observer", 700);
|
||||
|
||||
@Test
|
||||
void anonymousIsAuthorizedForNothing() {
|
||||
@@ -29,6 +31,22 @@ class AuthzTest {
|
||||
assertTrue(Authz.isUnauthenticated(null));
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code MessageService.answer}'s turn-ownership check treats a caller with no terminal as
|
||||
* matching a turn recorded for the unnamed primary. That rule only stays safe because an
|
||||
* unauthenticated caller — whose terminal is also {@code null} — never reaches {@code answer}
|
||||
* at all: {@link #anonymousIsAuthorizedForNothing} already covers every action including
|
||||
* {@code ANSWER}, but this test names the exact coupling so a future change to either side
|
||||
* cannot drift without turning this test red.
|
||||
*/
|
||||
@Test
|
||||
void anAnonymousCallerIsRefusedAnswerSoItCanNeverBeMistakenForTheUnnamedPrimary() {
|
||||
assertFalse(Authz.permits(ANON, ANSWER, null),
|
||||
"an anonymous caller, whose terminal is also null, must never reach answer() — the "
|
||||
+ "turn-ownership check's null-terminal match for the unnamed primary owner "
|
||||
+ "relies on this gate refusing it first");
|
||||
}
|
||||
|
||||
@Test
|
||||
void orchestrationBelongsToThePrimaryAlone() {
|
||||
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, SEND, DRAIN}) {
|
||||
@@ -38,6 +56,37 @@ class AuthzTest {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code fleet_send} is three call shapes behind one action name until {@code
|
||||
* FleetMcp#sendAction} picks one: a plain local {@link Authz.Action#SEND}, the {@code coordId}
|
||||
* route ({@link Authz.Action#COORD_SEND}), and the {@code turnId} answer form ({@link
|
||||
* Authz.Action#ANSWER}). All three carry the same grant as the undivided action did — a worker
|
||||
* is excluded from every one, exactly as it was excluded from the one combined action before.
|
||||
*/
|
||||
@Test
|
||||
void theThreeSendShapesCarryTheSameGrantAsTheOldUndividedAction() {
|
||||
for (Authz.Action a : new Authz.Action[]{SEND, COORD_SEND, ANSWER}) {
|
||||
assertTrue(Authz.permits(PRIMARY, a, "term_a"), "the primary may " + a);
|
||||
assertTrue(Authz.permits(ARCH_DESIGN, a, "term_a"), "an architect may " + a);
|
||||
assertFalse(Authz.permits(WORKER_A, a, "term_a"),
|
||||
"a worker performing " + a + " would be escalating into the orchestrator role");
|
||||
assertFalse(Authz.permits(ANON, a, "term_a"));
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code fleet_poll{ticket}} and {@code fleet_status} are {@link Authz.Action#TASK_READ}, split
|
||||
* out of the roster-only {@link Authz.Action#READ} (fleetd #678). The grant is unchanged from
|
||||
* what the undivided {@code READ} action gave every one of these callers.
|
||||
*/
|
||||
@Test
|
||||
void taskReadCarriesTheSameGrantReadDidBeforeTheSplit() {
|
||||
assertTrue(Authz.permits(PRIMARY, TASK_READ, null));
|
||||
assertTrue(Authz.permits(WORKER_A, TASK_READ, null));
|
||||
assertTrue(Authz.permits(ARCH_DESIGN, TASK_READ, null));
|
||||
assertFalse(Authz.permits(ANON, TASK_READ, null));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aWorkerMayReplyAndAskOnlyAsItself() {
|
||||
assertTrue(Authz.permits(WORKER_A, REPLY, "term_a"));
|
||||
@@ -135,4 +184,132 @@ class AuthzTest {
|
||||
assertFalse(Authz.isUnauthenticated(WORKER_A));
|
||||
assertFalse(Authz.isUnauthenticated(PRIMARY));
|
||||
}
|
||||
|
||||
// ── the collaborator matrix ─────────────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* {@code SEND} for a collaborator is the one grant that is conditional rather than fixed:
|
||||
* flipping only the classifier's answer for the target flips only this outcome.
|
||||
*/
|
||||
@Test
|
||||
void aCollaboratorMaySendOnlyWhenTheClassifierAcceptsTheTarget() {
|
||||
assertTrue(Authz.permits(COLLABORATOR, SEND, "term_lead", target -> true),
|
||||
"the classifier accepting the target must grant SEND");
|
||||
assertFalse(Authz.permits(COLLABORATOR, SEND, "term_lead", target -> false),
|
||||
"the classifier refusing the target must deny SEND");
|
||||
assertFalse(Authz.permits(COLLABORATOR, SEND, "term_lead"),
|
||||
"the real production classifier recognises no terminal yet, so SEND is refused today");
|
||||
}
|
||||
|
||||
/**
|
||||
* Control for the test above: every other action's result for a collaborator does not move
|
||||
* when the classifier does. Only {@code SEND} is wired to it.
|
||||
*/
|
||||
@Test
|
||||
void theClassifierMovesOnlySendForACollaborator() {
|
||||
for (Authz.Action a : Authz.Action.values()) {
|
||||
if (a == SEND) {
|
||||
continue;
|
||||
}
|
||||
assertEquals(
|
||||
Authz.permits(COLLABORATOR, a, "term_lead"),
|
||||
Authz.permits(COLLABORATOR, a, "term_lead", target -> true),
|
||||
a + " must not depend on the classifier at all");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void aCollaboratorMayReadAndScrapeMetrics() {
|
||||
assertTrue(Authz.permits(COLLABORATOR, READ, null));
|
||||
assertTrue(Authz.permits(COLLABORATOR, METRICS, null));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aCollaboratorMayReplyAndAskOnlyAsItsOwnPane() {
|
||||
assertTrue(Authz.permits(COLLABORATOR, REPLY, "term_collab"),
|
||||
"its own pane is its own");
|
||||
assertTrue(Authz.permits(COLLABORATOR, ASK, "term_collab"));
|
||||
|
||||
assertFalse(Authz.permits(COLLABORATOR, REPLY, "term_design"),
|
||||
"a collaborator must not reply on another pane");
|
||||
assertFalse(Authz.permits(COLLABORATOR, REPLY, null),
|
||||
"an absent target must not pass the own-session rule");
|
||||
}
|
||||
|
||||
/**
|
||||
* Every action denied to a collaborator, asserted denied even when the classifier would
|
||||
* accept any target — proving none of these is actually gated on the classifier at all.
|
||||
*/
|
||||
@Test
|
||||
void aCollaboratorIsDeniedLifecycleCoordinationAndTicketPolling() {
|
||||
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, DRAIN, HANDOVER, ANSWER, COORD_SEND,
|
||||
COORD_READ, TASK_READ}) {
|
||||
assertFalse(Authz.permits(COLLABORATOR, a, "term_lead", target -> true),
|
||||
"a collaborator must not " + a + " even when the classifier accepts every target");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void aCollaboratorIsNotCountedAsPrimaryWorkerOrArchitect() {
|
||||
assertFalse(COLLABORATOR.isPrimary());
|
||||
assertFalse(COLLABORATOR.isWorker());
|
||||
assertFalse(COLLABORATOR.isArchitect());
|
||||
assertTrue(COLLABORATOR.isCollaborator());
|
||||
}
|
||||
|
||||
/**
|
||||
* A collaborator is never spawned, so it must not be enrolled in the presence map as an
|
||||
* available member. Control: both a worker and an architect — which ARE spawned — still are.
|
||||
*/
|
||||
@Test
|
||||
void isSpawnedMemberIsFalseForACollaboratorButTrueForAWorkerAndAnArchitect() {
|
||||
assertFalse(COLLABORATOR.isSpawnedMember());
|
||||
assertTrue(WORKER_A.isSpawnedMember());
|
||||
assertTrue(ARCH_DESIGN.isSpawnedMember());
|
||||
}
|
||||
|
||||
// ── the observer matrix ─────────────────────────────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
void anObserverMayReadAndScrapeMetrics() {
|
||||
assertTrue(Authz.permits(OBSERVER, READ, null));
|
||||
assertTrue(Authz.permits(OBSERVER, METRICS, null));
|
||||
}
|
||||
|
||||
@Test
|
||||
void anObserverMayReplyAndAskOnlyAsItsOwnPane() {
|
||||
assertTrue(Authz.permits(OBSERVER, REPLY, "term_observer"), "its own pane is its own");
|
||||
assertTrue(Authz.permits(OBSERVER, ASK, "term_observer"));
|
||||
|
||||
assertFalse(Authz.permits(OBSERVER, REPLY, "term_design"),
|
||||
"an observer must not reply on another pane");
|
||||
assertFalse(Authz.permits(OBSERVER, REPLY, null),
|
||||
"an absent target must not pass the own-session rule");
|
||||
}
|
||||
|
||||
/**
|
||||
* Every action beyond READ/METRICS/REPLY/ASK, asserted denied for an observer — including
|
||||
* {@code TASK_READ}, which is the entire point of this role: an unconfigured pane must not be
|
||||
* able to poll a ticket or read another session's status.
|
||||
*/
|
||||
@Test
|
||||
void anObserverIsDeniedEverythingBeyondReadMetricsReplyAndAsk() {
|
||||
for (Authz.Action a : Authz.Action.values()) {
|
||||
if (a == READ || a == METRICS || a == REPLY || a == ASK) {
|
||||
continue;
|
||||
}
|
||||
assertFalse(Authz.permits(OBSERVER, a, "term_observer", target -> true),
|
||||
"an observer must not " + a + " even when the classifier accepts every target");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void anObserverIsNotCountedAsAnyOtherRole() {
|
||||
assertFalse(OBSERVER.isPrimary());
|
||||
assertFalse(OBSERVER.isWorker());
|
||||
assertFalse(OBSERVER.isArchitect());
|
||||
assertFalse(OBSERVER.isCollaborator());
|
||||
assertFalse(OBSERVER.isSpawnedMember());
|
||||
assertTrue(OBSERVER.isObserver());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,13 +1,25 @@
|
||||
package dev.ltms.fleet.auth;
|
||||
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.PaneLocator;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.mcp.ConnectionIdentity;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import dev.ltms.fleet.session.WorktreeRequest;
|
||||
import dev.ltms.fleet.session.Worktrees;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Optional;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
@@ -45,17 +57,23 @@ class CallerResolverTest {
|
||||
return members;
|
||||
}
|
||||
|
||||
/**
|
||||
* With no roster wired up at all (the simple constructor), a loopback pane that owns a herdr
|
||||
* pane but is not recognised as a live spawned member lands on the {@link Role#OBSERVER} floor
|
||||
* — unforgeable and never token-gated, exactly like a worker's own identity, because it comes
|
||||
* from the same connection-derived pane mapping.
|
||||
*/
|
||||
@Test
|
||||
void aLoopbackWorkerPaneResolvesToWorkerRegardlessOfAuthMode() {
|
||||
void aLoopbackPaneWithNoLiveRosterResolvesToObserverRegardlessOfAuthMode() {
|
||||
Principal underTrust = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
|
||||
Principal underToken = new CallerResolver(workerIdentity(), true, "s3cret")
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.WORKER, underTrust.role());
|
||||
assertEquals(Role.OBSERVER, underTrust.role());
|
||||
assertEquals("term_a", underTrust.terminal());
|
||||
assertEquals(Role.WORKER, underToken.role(),
|
||||
"worker identity is unforgeable and must never be token-gated — otherwise enabling "
|
||||
+ "auth would lock the whole fleet out of fleet_reply");
|
||||
assertEquals(Role.OBSERVER, underToken.role(),
|
||||
"the floor is unforgeable and must never be token-gated — otherwise enabling auth "
|
||||
+ "would lock every unconfigured pane out of even READ");
|
||||
assertEquals("term_a", underToken.terminal());
|
||||
}
|
||||
|
||||
@@ -79,20 +97,20 @@ class CallerResolverTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void otherPanesRemainWorkersWhenAPinIsSet() {
|
||||
void otherPanesRemainAtTheFloorWhenAPinIsSet() {
|
||||
Principal p = CallerResolver.pinnedTo(workerIdentity(), false, null, "term_someone_else")
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.WORKER, p.role());
|
||||
assertEquals(Role.OBSERVER, p.role());
|
||||
assertEquals("term_a", p.terminal());
|
||||
}
|
||||
|
||||
/** The pin is optional config, so an absent or whitespace one must change nothing at all. */
|
||||
@Test
|
||||
void aBlankPinLeavesWorkerResolutionUntouched() {
|
||||
assertEquals(Role.WORKER,
|
||||
void aBlankPinLeavesFloorResolutionUntouched() {
|
||||
assertEquals(Role.OBSERVER,
|
||||
CallerResolver.pinnedTo(workerIdentity(), false, null, " ").resolve("127.0.0.1", 42, null).role());
|
||||
assertEquals(Role.WORKER,
|
||||
assertEquals(Role.OBSERVER,
|
||||
CallerResolver.pinnedTo(workerIdentity(), false, null, null).resolve("127.0.0.1", 42, null).role());
|
||||
}
|
||||
|
||||
@@ -213,13 +231,13 @@ class CallerResolverTest {
|
||||
* mid-scan teardown into a refusal — the real match is still found and resolves as a worker.
|
||||
*/
|
||||
@Test
|
||||
void aHerdrErrorOnANonOwningPaneStillResolvesTheRealWorker() {
|
||||
void aHerdrErrorOnANonOwningPaneStillResolvesTheRealPane() {
|
||||
FakeHerdr vanishedElsewhere = new FakeHerdr().processInfoFailsForPane("w2:p9", "pane_not_found");
|
||||
ConnectionIdentity id = new ConnectionIdentity(new PaneLocator(vanishedElsewhere), _ -> FakeHerdr.WORKER_PID);
|
||||
|
||||
Principal p = new CallerResolver(id).resolve("127.0.0.1", 55555, null);
|
||||
|
||||
assertEquals(Role.WORKER, p.role());
|
||||
assertEquals(Role.OBSERVER, p.role());
|
||||
assertEquals("term_a", p.terminal());
|
||||
}
|
||||
|
||||
@@ -263,12 +281,12 @@ class CallerResolverTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void aPaneAbsentFromTheRegistryIsStillAWorker() {
|
||||
void aPaneAbsentFromTheRegistryFallsToTheObserverFloor() {
|
||||
Principal p = new CallerResolver(workerIdentity(), false, null,
|
||||
Map.of("term_elsewhere", "gpt-sol-5.6"))
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.WORKER, p.role());
|
||||
assertEquals(Role.OBSERVER, p.role());
|
||||
assertEquals("term_a", p.terminal());
|
||||
assertNull(p.name());
|
||||
}
|
||||
@@ -294,16 +312,16 @@ class CallerResolverTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void anEmptyRegistryLeavesEveryPaneAWorker() {
|
||||
void anEmptyRegistryLeavesEveryPaneAtTheObserverFloor() {
|
||||
Map<String, String> noLeads = null;
|
||||
assertEquals(Role.WORKER,
|
||||
assertEquals(Role.OBSERVER,
|
||||
new CallerResolver(workerIdentity(), false, null, Map.of())
|
||||
.resolve("127.0.0.1", 42, null).role());
|
||||
assertEquals(Role.WORKER,
|
||||
assertEquals(Role.OBSERVER,
|
||||
new CallerResolver(workerIdentity(), false, null, noLeads)
|
||||
.resolve("127.0.0.1", 42, null).role());
|
||||
// CB-531: and the same for the live-registry form, whose supplier may also be absent.
|
||||
assertEquals(Role.WORKER,
|
||||
// And the same for the live-registry form, whose supplier may also be absent.
|
||||
assertEquals(Role.OBSERVER,
|
||||
CallerResolver.withLeads(workerIdentity(), false, null, null)
|
||||
.resolve("127.0.0.1", 42, null).role());
|
||||
}
|
||||
@@ -376,7 +394,7 @@ class CallerResolverTest {
|
||||
Map<String, String> live = new java.util.HashMap<>();
|
||||
CallerResolver r = CallerResolver.withLeads(workerIdentity(), false, null, () -> live);
|
||||
|
||||
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
|
||||
assertEquals(Role.OBSERVER, r.resolve("127.0.0.1", 42, null).role());
|
||||
|
||||
live.put("term_a", "gpt-sol-5.6"); // the scanner sees a newly-labelled tab
|
||||
|
||||
@@ -393,7 +411,7 @@ class CallerResolverTest {
|
||||
|
||||
mutable.put("term_a", "sneaky");
|
||||
|
||||
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
|
||||
assertEquals(Role.OBSERVER, r.resolve("127.0.0.1", 42, null).role());
|
||||
}
|
||||
|
||||
// ── CB-548: architect slots ─────────────────────────────────────────────────────────────────
|
||||
@@ -423,13 +441,13 @@ class CallerResolverTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void anUnboundPaneStillResolvesAsAWorker() {
|
||||
void anUnboundPaneResolvesToTheObserverFloor() {
|
||||
MemberRegistry members = new MemberRegistry(new FleetConfig.Fleet(Map.of(),
|
||||
Map.of("lead-designer", new FleetConfig.Slot("sonnet")), Map.of(), Map.of(), null));
|
||||
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
|
||||
Map::of, members).resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.WORKER, p.role());
|
||||
assertEquals(Role.OBSERVER, p.role());
|
||||
assertNull(p.name());
|
||||
}
|
||||
|
||||
@@ -455,12 +473,11 @@ class CallerResolverTest {
|
||||
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
|
||||
Map::of, members);
|
||||
|
||||
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
|
||||
assertEquals(Role.OBSERVER, r.resolve("127.0.0.1", 42, null).role());
|
||||
|
||||
assertTrue(members.bind("architect:lead-designer", "term_a")); // the later lifecycle binds the slot
|
||||
|
||||
assertEquals(Role.ARCHITECT, r.resolve("127.0.0.1", 42, null).role());
|
||||
assertEquals("architect:lead-designer", r.members().get("term_a"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -483,12 +500,13 @@ class CallerResolverTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void aBoundNonArchitectSlotStillResolvesAsAWorker() {
|
||||
void aBoundNonArchitectSlotResolvesToTheObserverFloorNotArchitect() {
|
||||
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
|
||||
Map::of, boundMembers("dev:builder", MemberRole.DEV))
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.WORKER, p.role(), "a dev binding must never grant architect rights");
|
||||
assertEquals(Role.OBSERVER, p.role(), "a dev binding must never grant architect rights, "
|
||||
+ "and this construction path wires no roster to recognise it as the live dev it is");
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -499,17 +517,17 @@ class CallerResolverTest {
|
||||
assertThrows(IllegalArgumentException.class, () -> new CallerResolver(id, true, " "));
|
||||
}
|
||||
@Test
|
||||
void aWorkerOnAnyLoopbackSourceAddressIsStillAWorkerNotThePrimary() {
|
||||
// fleetd #305: the escalation. ConnectionIdentity used to accept only 127.0.0.1, so a
|
||||
// worker connecting from 127.0.0.2 resolved to no terminal, and this resolver's own
|
||||
// (wider) loopback check then made it the PRIMARY — granting spawn, stop, send and drain.
|
||||
// Measured on the Linux fleet host: binding a source of 127.0.0.2 succeeds there, so the
|
||||
// path is real and not theoretical.
|
||||
void aPaneOnAnyLoopbackSourceAddressIsStillAtTheFloorNotThePrimary() {
|
||||
// fleetd #305: the escalation this guards against. ConnectionIdentity used to accept only
|
||||
// 127.0.0.1, so a pane connecting from 127.0.0.2 resolved to no terminal, and this
|
||||
// resolver's own (wider) loopback check then made it the PRIMARY — granting spawn, stop,
|
||||
// send and drain. Measured on the Linux fleet host: binding a source of 127.0.0.2 succeeds
|
||||
// there, so the path is real and not theoretical.
|
||||
CallerResolver r = new CallerResolver(workerIdentity(), false, null);
|
||||
for (String src : new String[]{"127.0.0.1", "127.0.0.2", "127.1.2.3", "::ffff:127.0.0.2"}) {
|
||||
Principal p = r.resolve(src, 55555, null);
|
||||
assertEquals(Role.WORKER, p.role(), "a worker must stay a worker from source " + src);
|
||||
assertEquals("term_a", p.terminal(), "worker terminal from source " + src);
|
||||
assertEquals(Role.OBSERVER, p.role(), "the pane must stay off PRIMARY from source " + src);
|
||||
assertEquals("term_a", p.terminal(), "pane terminal from source " + src);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -523,4 +541,375 @@ class CallerResolverTest {
|
||||
}
|
||||
}
|
||||
|
||||
// ── fleetd #669 Unit D: a live spawned member outranks every tab map ────────────────────────
|
||||
|
||||
/**
|
||||
* Criterion 1: a terminal present in BOTH the spawned-member roster AND the lead tab map
|
||||
* resolves as its member role, not as a lead — the roster is checked first, consulting no tab
|
||||
* map at all when it matches.
|
||||
*/
|
||||
@Test
|
||||
void aSpawnedMemberWinsOverALeadTabForTheSamePane() {
|
||||
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
|
||||
() -> Map.of("term_a", "opus-5.0"), new MemberRegistry(null),
|
||||
t -> "term_a".equals(t) ? MemberRole.DEV : null, Map::of)
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.WORKER, p.role(),
|
||||
"a live spawned member's own identity must win over a tab map naming the same pane a lead");
|
||||
assertEquals("term_a", p.terminal());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #702: between {@link SessionManager#release} removing the registry entry and the
|
||||
* pane actually stopping, a resolve for that terminal must still see the live member and
|
||||
* never fall through to a tab map — wired through the real {@link SessionManager}, not a
|
||||
* hand-rolled stand-in for its {@code spawnedMemberRole}.
|
||||
*
|
||||
* <p>Reuses the ticket's own test idea: an injected {@code hasUncommitted} resolves the
|
||||
* releasing terminal from inside {@code release}'s window — a real call landing inside the
|
||||
* window, so no sleep and no race.
|
||||
*
|
||||
* <p>The property asserted is "no tab map is consulted", never "the same role is returned".
|
||||
* The mandatory control is the lead tab map: it names this exact terminal, and the same
|
||||
* resolve taken <em>outside</em> the window (before release runs) must still return the lead
|
||||
* role — without that control, the in-window assertion would also pass on an empty map and
|
||||
* prove nothing.
|
||||
*/
|
||||
@Test
|
||||
void aPaneMidTeardownResolvesAsItsOwnRoleConsultingNoTabMap() {
|
||||
FakeHerdr herdr = new FakeHerdr().pinNextStarts(1, "term_a", "w2:p7");
|
||||
FleetConfig.Profile cfg = new FleetConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
|
||||
AtomicReference<Principal> duringWindow = new AtomicReference<>();
|
||||
AtomicReference<CallerResolver> resolverRef = new AtomicReference<>();
|
||||
Worktrees worktrees = new Worktrees() {
|
||||
@Override
|
||||
public String add(String repoRoot, String branch, String baseRef) {
|
||||
return "/wt/" + branch.replace('/', '_');
|
||||
}
|
||||
|
||||
@Override
|
||||
public void remove(String repoRoot, String worktreePath) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void deleteBranch(String repoRoot, String branch) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean hasUncommitted(String worktreePath) {
|
||||
// Runs from INSIDE release()'s git-status shell-out: the registry entry is
|
||||
// already gone, but the pane has not stopped yet.
|
||||
duringWindow.set(resolverRef.get().resolve("127.0.0.1", 42, null));
|
||||
return false;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public String repoRoot(String cwd) {
|
||||
return "/repo";
|
||||
}
|
||||
|
||||
@Override
|
||||
public Optional<String> snapshot(String worktreePath, String branch, String message) {
|
||||
return Optional.empty();
|
||||
}
|
||||
|
||||
@Override
|
||||
public WipRefStats wipRefs(String repoRoot) {
|
||||
return new WipRefStats(0, 0L);
|
||||
}
|
||||
|
||||
@Override
|
||||
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
|
||||
return 0;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void shareWithGroup(String repoRoot, String worktreePath) {
|
||||
}
|
||||
};
|
||||
SessionManager sessions = new SessionManager(workers, worktrees);
|
||||
|
||||
// The lead tab map names "term_a" before anything is ever spawned onto it — the control
|
||||
// this test needs. Built up front so the SAME resolver answers every resolve() call below.
|
||||
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> FakeHerdr.WORKER_PID);
|
||||
CallerResolver resolver = CallerResolver.withLeadsAndMembers(identity, false, null,
|
||||
() -> Map.of("term_a", "the-lead"), null, sessions::spawnedMemberRole, Map::of);
|
||||
resolverRef.set(resolver);
|
||||
|
||||
Principal before = resolver.resolve("127.0.0.1", 42, null);
|
||||
assertEquals(Role.PRIMARY, before.role(),
|
||||
"control: with no live or releasing member on this terminal, the lead tab map must "
|
||||
+ "win — this is what proves the in-window assertion below is not passing on "
|
||||
+ "an empty map");
|
||||
assertEquals("the-lead", before.name());
|
||||
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("fleetd-702", null));
|
||||
assertEquals("term_a", s.terminalId(), "sanity: the spawn resolved to the pinned pane");
|
||||
|
||||
sessions.release(s.paneId());
|
||||
|
||||
assertNotNull(duringWindow.get(), "the dirty check must have run and captured a resolve");
|
||||
assertEquals(Role.WORKER, duringWindow.get().role(),
|
||||
"inside the window the pane must resolve as its own live-member role, consulting no "
|
||||
+ "tab map — a lead tab naming the same terminal must not win");
|
||||
assertEquals("term_a", duringWindow.get().terminal());
|
||||
}
|
||||
|
||||
/**
|
||||
* {@link SessionManager#releaseRemoved} unbinds the architect slot before the git-status
|
||||
* shell-out that opens the teardown window, so a resolve landing inside that window must see
|
||||
* the slot already unbound and resolve {@link Role#WORKER} — never {@link Role#ARCHITECT},
|
||||
* and never by checking role equality against the live session, which would hold even if the
|
||||
* unbind ran too late.
|
||||
*
|
||||
* <p>The control is the same resolve taken outside the window, while the slot is still bound,
|
||||
* which must return {@link Role#ARCHITECT} — without it this test would also pass against a
|
||||
* slot that was never bound, and prove nothing.
|
||||
*/
|
||||
@Test
|
||||
void aReleasingArchitectIsDemotedToWorkerInsideTheTeardownWindow() {
|
||||
FakeHerdr herdr = new FakeHerdr().pinNextStarts(1, "term_a", "w2:p7");
|
||||
FleetConfig.Profile cfg = new FleetConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
|
||||
AtomicReference<Principal> duringWindow = new AtomicReference<>();
|
||||
AtomicReference<CallerResolver> resolverRef = new AtomicReference<>();
|
||||
Worktrees worktrees = new Worktrees() {
|
||||
@Override
|
||||
public String add(String repoRoot, String branch, String baseRef) {
|
||||
return "/wt/" + branch.replace('/', '_');
|
||||
}
|
||||
|
||||
@Override
|
||||
public void remove(String repoRoot, String worktreePath) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void deleteBranch(String repoRoot, String branch) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean hasUncommitted(String worktreePath) {
|
||||
// Runs from INSIDE release()'s git-status shell-out: the architect slot is already
|
||||
// unbound by this point, but the pane has not stopped yet.
|
||||
duringWindow.set(resolverRef.get().resolve("127.0.0.1", 42, null));
|
||||
return false;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public String repoRoot(String cwd) {
|
||||
return "/repo";
|
||||
}
|
||||
|
||||
@Override
|
||||
public Optional<String> snapshot(String worktreePath, String branch, String message) {
|
||||
return Optional.empty();
|
||||
}
|
||||
|
||||
@Override
|
||||
public WipRefStats wipRefs(String repoRoot) {
|
||||
return new WipRefStats(0, 0L);
|
||||
}
|
||||
|
||||
@Override
|
||||
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
|
||||
return 0;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void shareWithGroup(String repoRoot, String worktreePath) {
|
||||
}
|
||||
};
|
||||
SessionManager sessions = new SessionManager(workers, worktrees);
|
||||
|
||||
MemberRegistry members = new MemberRegistry(new FleetConfig.Fleet(Map.of(),
|
||||
Map.of("lead-designer", new FleetConfig.Slot("ltms-local")), Map.of(), Map.of(), null));
|
||||
sessions.setMemberLifecycle(members);
|
||||
|
||||
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> FakeHerdr.WORKER_PID);
|
||||
CallerResolver resolver = CallerResolver.withLeadsAndMembers(identity, false, null,
|
||||
Map::of, members, sessions::spawnedMemberRole, Map::of);
|
||||
resolverRef.set(resolver);
|
||||
|
||||
MemberSession s = sessions.acquire("ltms-local", MemberRole.ARCHITECT, null, "/caller/proj", null,
|
||||
new WorktreeRequest("fleetd-702d", null));
|
||||
assertEquals("term_a", s.terminalId(), "sanity: the spawn resolved to the pinned pane");
|
||||
assertEquals(MemberRole.ARCHITECT, s.role(), "sanity: the slot bind succeeded");
|
||||
|
||||
Principal before = resolver.resolve("127.0.0.1", 42, null);
|
||||
assertEquals(Role.ARCHITECT, before.role(),
|
||||
"control: with the slot still bound, the pane must resolve as an architect — this "
|
||||
+ "is what proves the in-window assertion below is not passing against a "
|
||||
+ "slot that was never bound");
|
||||
assertEquals("lead-designer", before.name());
|
||||
|
||||
sessions.release(s.paneId());
|
||||
|
||||
assertNotNull(duringWindow.get(), "the dirty check must have run and captured a resolve");
|
||||
assertEquals(Role.WORKER, duringWindow.get().role(),
|
||||
"inside the window the architect slot is already unbound, so the result must be "
|
||||
+ "WORKER — asserting role-equality with the live session here would tempt "
|
||||
+ "moving the unbind earlier or later, which would be wrong either way");
|
||||
assertEquals("term_a", duringWindow.get().terminal());
|
||||
}
|
||||
|
||||
/** A spawned architect in the roster resolves ARCHITECT, carrying its bound slot's name. */
|
||||
@Test
|
||||
void aSpawnedArchitectInTheRosterResolvesArchitectWithItsSlotName() {
|
||||
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
|
||||
boundMembers("architect:lead-designer", MemberRole.ARCHITECT),
|
||||
t -> "term_a".equals(t) ? MemberRole.ARCHITECT : null, Map::of)
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.ARCHITECT, p.role());
|
||||
assertEquals("lead-designer", p.name());
|
||||
assertEquals("term_a", p.terminal());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #424 regression: the roster only answers THAT a pane is a live spawned member; config
|
||||
* still decides WHAT that member's slot grants. A slot revoked after the bind must still demote
|
||||
* the session on its very next request, exactly as it would for a pane with no roster entry at
|
||||
* all — the roster's own ARCHITECT role must never be granted on its word alone.
|
||||
*
|
||||
* <p>{@code bind} refuses an unconfigured slot, so the revoked state can only be reached by
|
||||
* binding while the slot is configured and then swapping the config out from under it, the way
|
||||
* a live reload does.
|
||||
*/
|
||||
@Test
|
||||
void aRevokedArchitectSlotDemotesALiveSpawnedArchitectToWorker() {
|
||||
FleetConfig.Fleet configured = new FleetConfig.Fleet(Map.of(),
|
||||
Map.of("lead-designer", new FleetConfig.Slot("sonnet")), Map.of(), Map.of(), null);
|
||||
java.util.concurrent.atomic.AtomicReference<FleetConfig.Fleet> live =
|
||||
new java.util.concurrent.atomic.AtomicReference<>(configured);
|
||||
MemberRegistry members = MemberRegistry.live(live::get);
|
||||
assertTrue(members.bind("architect:lead-designer", "term_a"));
|
||||
|
||||
live.set(new FleetConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(), null)); // slot revoked
|
||||
|
||||
// Setup controls: the slot is really gone from config, but the occupancy is still there —
|
||||
// otherwise this test would pass for the wrong reason.
|
||||
assertNull(members.roleForSlot("architect:lead-designer"), "setup control: the slot must be gone from config");
|
||||
assertEquals("architect:lead-designer", members.snapshot().get("term_a"),
|
||||
"setup control: the binding itself must still be there");
|
||||
|
||||
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
|
||||
members, t -> "term_a".equals(t) ? MemberRole.ARCHITECT : null, Map::of)
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.WORKER, p.role(),
|
||||
"a revoked slot must demote a live spawned architect on its very next request");
|
||||
}
|
||||
|
||||
/** Criterion 3: a configured collaborator tab that is not a spawned member resolves COLLABORATOR. */
|
||||
@Test
|
||||
void aConfiguredCollaboratorTabResolvesToCollaboratorCarryingItsName() {
|
||||
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
|
||||
new MemberRegistry(null), t -> null, () -> Map.of("term_a", "ops"))
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.COLLABORATOR, p.role());
|
||||
assertEquals("ops", p.name());
|
||||
assertEquals("term_a", p.terminal());
|
||||
}
|
||||
|
||||
/** Regression: an empty collaborator registry leaves every pane at the unconfigured-pane floor. */
|
||||
@Test
|
||||
void anEmptyCollaboratorRegistryLeavesEveryPaneAtTheObserverFloor() {
|
||||
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
|
||||
new MemberRegistry(null), t -> null, Map::of)
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.OBSERVER, p.role());
|
||||
assertNull(p.name());
|
||||
}
|
||||
|
||||
@Test
|
||||
void describeNamesTheCollaborator() {
|
||||
assertEquals("collaborator:ops", Principal.collaborator("ops", "term_a", 1).describe());
|
||||
}
|
||||
|
||||
// ── fleetd #669 Unit D: knownLeadOrCollaborator() reads the same maps resolve() does ───────────
|
||||
|
||||
@Test
|
||||
void knownLeadOrCollaboratorIsTrueForALeadTerminal() {
|
||||
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
|
||||
() -> Map.of("term_lead", "opus-5.0"), new MemberRegistry(null), t -> null, Map::of);
|
||||
|
||||
assertTrue(r.knownLeadOrCollaborator().test("term_lead"));
|
||||
assertFalse(r.knownLeadOrCollaborator().test("term_other"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void knownLeadOrCollaboratorIsTrueForACollaboratorTerminal() {
|
||||
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
|
||||
new MemberRegistry(null), t -> null, () -> Map.of("term_collab", "ops"));
|
||||
|
||||
assertTrue(r.knownLeadOrCollaborator().test("term_collab"));
|
||||
assertFalse(r.knownLeadOrCollaborator().test("term_other"));
|
||||
}
|
||||
|
||||
// ── fleetd #705: narrowing the unconfigured-pane floor to OBSERVER ──────────────────────────
|
||||
|
||||
/**
|
||||
* The case this ticket exists for: a pane the resolver cannot place as a live spawned member,
|
||||
* a lead, a bound architect slot, or a configured collaborator must land on the narrow
|
||||
* {@link Role#OBSERVER} floor, never the {@link Role#WORKER} the old fallback granted.
|
||||
*
|
||||
* <p>The second assertion is the control the ticket requires: a terminal the roster DOES
|
||||
* recognise as a live spawned member must still resolve its own role. Without it, this test
|
||||
* would also pass if the fix accidentally turned every caller into an observer.
|
||||
*/
|
||||
@Test
|
||||
void anUnconfiguredPaneResolvesObserverButARegisteredMemberStillResolvesItsOwnRole() {
|
||||
Principal unconfigured = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
|
||||
assertEquals(Role.OBSERVER, unconfigured.role(),
|
||||
"a pane matching none of the configured or live-roster roles must fall to the "
|
||||
+ "floor, not WORKER");
|
||||
assertEquals("term_a", unconfigured.terminal());
|
||||
|
||||
Principal registered = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
|
||||
Map::of, new MemberRegistry(null),
|
||||
t -> "term_a".equals(t) ? MemberRole.DEV : null, Map::of)
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
assertEquals(Role.WORKER, registered.role(),
|
||||
"control: a live spawned member must keep resolving its own role, never the "
|
||||
+ "unconfigured-pane floor");
|
||||
}
|
||||
|
||||
@Test
|
||||
void describeNamesTheObserverByItsPane() {
|
||||
assertEquals("observer:term_a", Principal.observer("term_a", 1).describe());
|
||||
}
|
||||
|
||||
@Test
|
||||
void knownLeadOrCollaboratorIsFalseForASpawnedMembersTerminal() {
|
||||
// The exact scenario a collaborator's SEND must never reach: a live spawned member's own
|
||||
// terminal, which is neither a configured lead nor a configured collaborator.
|
||||
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
|
||||
new MemberRegistry(null), t -> "term_a".equals(t) ? MemberRole.DEV : null, Map::of);
|
||||
|
||||
assertFalse(r.knownLeadOrCollaborator().test("term_a"));
|
||||
}
|
||||
|
||||
}
|
||||
|
||||
@@ -295,9 +295,10 @@ class MemberRegistryLiveTest {
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
Principal after = resolver.resolve("127.0.0.1", 42, null);
|
||||
assertEquals(Role.WORKER, after.role(),
|
||||
"removing the slot from config must demote the bound session to worker on its "
|
||||
+ "NEXT request — this is the ticket's whole point");
|
||||
assertEquals(Role.OBSERVER, after.role(),
|
||||
"removing the slot from config must demote the bound session on its NEXT request — "
|
||||
+ "this harness wires no live roster for term_a, so the demotion lands on "
|
||||
+ "the unconfigured-pane floor");
|
||||
assertEquals("term_a", after.terminal(), "same pane, same terminal — only the role changed");
|
||||
}
|
||||
|
||||
|
||||
@@ -107,6 +107,8 @@ class ConfigRefTest {
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertTrue(out.deferred().isEmpty());
|
||||
assertTrue(out.split().stream().noneMatch(s -> s.contains("fleet.collaborators")),
|
||||
out.split().toString());
|
||||
assertEquals("new charter", ref.get().fleet().charterFor(
|
||||
dev.ltms.fleet.peer.MemberRole.ARCHITECT));
|
||||
}
|
||||
@@ -786,14 +788,100 @@ class ConfigRefTest {
|
||||
assertTrue(out.split().getFirst().startsWith("fleet:"), out.split().toString());
|
||||
assertTrue(out.split().getFirst().contains("restart"), out.split().toString());
|
||||
assertTrue(out.split().getFirst().contains("live"), out.split().toString());
|
||||
assertTrue(out.split().stream().noneMatch(s -> s.contains("fleet.collaborators")),
|
||||
out.split().toString());
|
||||
assertTrue(out.summary().contains("partially live"), out.summary());
|
||||
// The snapshot still carries the new value — the LeadTabScanner's identity map and
|
||||
// LeadLauncher's auto-launch are what wait for a restart; a reload rebuilds neither.
|
||||
assertEquals("lead: opus-b", ref.get().fleet().leaders().get("opus").tab());
|
||||
}
|
||||
|
||||
@Test
|
||||
void addingAFleetCollaboratorIsReportedAsSplit(@TempDir Path dir) throws Exception {
|
||||
assertCollaboratorChangeIsReported(dir, """
|
||||
fleet:
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "collaborator: alex"
|
||||
""", "add a collaborator");
|
||||
}
|
||||
|
||||
@Test
|
||||
void removingAFleetCollaboratorIsReportedAsSplit(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "collaborator: alex"
|
||||
"""));
|
||||
ConfigRef ref = refFor(f);
|
||||
|
||||
Files.writeString(f, yaml("fleet: {}\n"));
|
||||
assertCollaboratorSplit(ref.reload(), "remove a collaborator");
|
||||
}
|
||||
|
||||
@Test
|
||||
void changingAFleetCollaboratorTabIsReportedAsSplit(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "collaborator: alex-a"
|
||||
"""));
|
||||
ConfigRef ref = refFor(f);
|
||||
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "collaborator: alex-b"
|
||||
"""));
|
||||
assertCollaboratorSplit(ref.reload(), "change a collaborator tab");
|
||||
}
|
||||
|
||||
@Test
|
||||
void unchangedFleetCollaboratorsProduceNoSplitReport(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
String config = yaml("""
|
||||
fleet:
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "collaborator: alex"
|
||||
""");
|
||||
Files.writeString(f, config);
|
||||
ConfigRef ref = refFor(f);
|
||||
|
||||
Files.writeString(f, config);
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertTrue(out.split().isEmpty(), out.split().toString());
|
||||
}
|
||||
|
||||
private static void assertCollaboratorChangeIsReported(Path dir, String changed, String action)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml("fleet: {}\n"));
|
||||
ConfigRef ref = refFor(f);
|
||||
|
||||
Files.writeString(f, yaml(changed));
|
||||
assertCollaboratorSplit(ref.reload(), action);
|
||||
}
|
||||
|
||||
private static void assertCollaboratorSplit(ConfigRef.Outcome out, String action) {
|
||||
assertTrue(out.applied(), action);
|
||||
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
|
||||
assertEquals(1, out.split().size(), out.split().toString());
|
||||
String report = out.split().getFirst();
|
||||
assertTrue(report.startsWith("fleet: fleet.collaborators"), report);
|
||||
assertTrue(report.contains("read once at startup"), report);
|
||||
assertTrue(report.contains("restart"), report);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #333: {@code fleet.leaders} is the ONLY frozen part of {@code fleet:}. A reload that
|
||||
* fleetd #333: {@code fleet.leaders} is a frozen part of {@code fleet:}. A reload that
|
||||
* changes {@code tabLabel} (or charters, or a role pool) without touching {@code fleet.leaders}
|
||||
* must stay fully hot with nothing reported — proving {@link ConfigRef#changedSplitKeys}
|
||||
* compares {@code fleet.leaders} specifically rather than the whole {@code Fleet} record, which
|
||||
|
||||
@@ -676,12 +676,8 @@ class FleetConfigTest {
|
||||
assertEquals(5, hb.quietNudgeCap());
|
||||
}
|
||||
|
||||
/**
|
||||
* The hazard the guard exists for: fleetd writes worker tab labels and reads lead tab labels.
|
||||
* Overlap the two and every worker it spawns is read back as a lead.
|
||||
*/
|
||||
@Test
|
||||
void aLeadPrefixThatAProfileTabLabelOverrideAlsoMatchesRefusesToStart(@TempDir Path dir)
|
||||
void aProfileTabLabelOverrideMatchingALeadTabRefusesToStart(@TempDir Path dir)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("collide.yaml");
|
||||
Files.writeString(f, """
|
||||
@@ -689,32 +685,33 @@ class FleetConfigTest {
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
tabLabel: "lead: {profile} #{n}"
|
||||
tabLabel: "alpha"
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
tabPrefix: "lead:"
|
||||
tab: "alpha"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e =
|
||||
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
|
||||
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
|
||||
assertTrue(e.getMessage().contains("alpha"), "the message must name the offending label");
|
||||
}
|
||||
|
||||
/** A bad fleet-wide template promotes every member, not one profile — so it is checked too. */
|
||||
@Test
|
||||
void aFleetTabLabelThatMatchesALeadPrefixRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
void aFleetTabLabelTemplateThatCanRenderAsALeadTabRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("collide-template.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
pha: {}
|
||||
fleet:
|
||||
tabLabel: "lead: {role} {profile}"
|
||||
tabLabel: "al{profile}"
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
tab: "alpha"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
@@ -723,12 +720,28 @@ class FleetConfigTest {
|
||||
assertTrue(e.getMessage().contains("fleet.tabLabel"));
|
||||
}
|
||||
|
||||
/**
|
||||
* The point of making role the label's first field: {@code {role}} comes from a closed enum, so
|
||||
* a generated label cannot begin with {@code "lead:"} however the fleet is configured.
|
||||
*/
|
||||
@Test
|
||||
void theDefaultTabLabelCannotCollideWithTheDefaultLeadPrefix(@TempDir Path dir) throws Exception {
|
||||
void anExactFleetTabLabelCollisionRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("exact-tab-collision.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
tabLabel: "alpha"
|
||||
leaders:
|
||||
alpha:
|
||||
tab: "alpha"
|
||||
""");
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class,
|
||||
() -> FleetConfig.load(f).validateAll());
|
||||
assertTrue(e.getMessage().contains("fleet.tabLabel"),
|
||||
"the message must name the offending label");
|
||||
assertTrue(e.getMessage().contains("alpha"), "the message must name the colliding lead tab");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aFleetTabLabelTemplateThatCannotRenderAsALeadTabIsAllowed(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("ok.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
@@ -737,17 +750,13 @@ class FleetConfigTest {
|
||||
gx10:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
fleet:
|
||||
tabLabel: "worker-{profile}"
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
tab: "alpha"
|
||||
""");
|
||||
|
||||
assertDoesNotThrow(() -> FleetConfig.load(f).validateLeadTabPrefixes());
|
||||
for (MemberRole role : MemberRole.values()) {
|
||||
assertFalse(FleetConfig.Fleet.DEFAULT_TAB_LABEL
|
||||
.replace("{role}", role.wireName()).startsWith("lead:"),
|
||||
"no role renders a label that reads as a lead");
|
||||
}
|
||||
assertDoesNotThrow(() -> FleetConfig.load(f).validateAll());
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -765,6 +774,165 @@ class FleetConfigTest {
|
||||
"a label that collides with a convention nobody reads is not a problem");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #677: identity is matched on a lead's exact {@code tab} alone, so two leads sharing
|
||||
* one tab means only one of them is ever found — the guard must catch this independently of
|
||||
* the member-template checks above.
|
||||
*/
|
||||
@Test
|
||||
void twoLeadsSharingTheSameExactTabRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("shared-tab.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "shared tab"
|
||||
sonnet:
|
||||
tab: "shared tab"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e =
|
||||
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
|
||||
assertTrue(e.getMessage().contains("opus"), "the message must name one offending lead");
|
||||
assertTrue(e.getMessage().contains("sonnet"), "the message must name the other offending lead");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #693: the guard matches tabs case-insensitively, because
|
||||
* {@code LeadTabScanner} keys its tab map on a lowercased label — two tabs differing only in
|
||||
* case collide there too, and the guard must catch that independently of the exact-match case
|
||||
* above.
|
||||
*/
|
||||
@Test
|
||||
void twoLeadsSharingTheSameTabInDifferentCaseRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("shared-tab-case.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "Shared Tab"
|
||||
sonnet:
|
||||
tab: "shared tab"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e =
|
||||
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
|
||||
assertTrue(e.getMessage().contains("opus"), "the message must name one offending lead");
|
||||
assertTrue(e.getMessage().contains("sonnet"), "the message must name the other offending lead");
|
||||
}
|
||||
|
||||
/** Control for {@link #twoLeadsSharingTheSameExactTabRefusesToStart}: distinct tabs load cleanly. */
|
||||
@Test
|
||||
void twoLeadsWithDistinctExactTabsAreAllowed(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("distinct-tabs.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "opus tab"
|
||||
sonnet:
|
||||
tab: "sonnet tab"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
assertDoesNotThrow(cfg::validateLeadTabPrefixes);
|
||||
}
|
||||
|
||||
// ── validatePanePlacementAgainstLeadTabs ────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* The hazard this guard closes: a pane-placed member lands inside the focused tab rather than
|
||||
* its own, so it can land inside a lead's labelled tab and, while its pane carries no entry in
|
||||
* the spawned-member roster, be read back as that lead.
|
||||
*/
|
||||
@Test
|
||||
void aPanePlacedProfileWithALeadTabRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("pane-hazard.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
placement: pane
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class,
|
||||
cfg::validatePanePlacementAgainstLeadTabs);
|
||||
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aPanePlacedProfileWithNoLeadTabIsAllowed(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("pane-no-tab.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
placement: pane
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
profile: gx10
|
||||
""");
|
||||
|
||||
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
|
||||
"a leader with no tab feeds nothing into the scanner, so pane placement is safe");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aTabPlacedProfileWithALeadTabIsAllowed(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("tab-safe.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
placement: tab
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
""");
|
||||
|
||||
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
|
||||
"a member in its own tab cannot land inside a lead's tab");
|
||||
}
|
||||
|
||||
/** Proves the reflective sweep behind {@code validateAll} really reaches this validator. */
|
||||
@Test
|
||||
void validateAllAlsoRefusesPanePlacementAgainstALeadTab(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("pane-hazard-sweep.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
placement: pane
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAll);
|
||||
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
|
||||
}
|
||||
|
||||
// ── CB-530/CB-579: the leaders registry ─────────────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
@@ -833,6 +1001,274 @@ class FleetConfigTest {
|
||||
"no primary.terminal pin ⇒ nothing registered, even with fleet.leaders configured");
|
||||
}
|
||||
|
||||
// ── fleetd #669: the collaborators registry ────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* {@code Fleet} is {@code @JsonIgnoreProperties(ignoreUnknown = true)}, so a config naming
|
||||
* {@code fleet.collaborators.<name>.tab} loads with no exception whether or not the key is
|
||||
* ever read into the object model. Asserting only "no exception" would pass both before and
|
||||
* after the real fix, so this asserts the parsed value is actually reachable from the loaded
|
||||
* {@code FleetConfig} — the one thing a vacuous "no exception" test cannot tell apart.
|
||||
*/
|
||||
@Test
|
||||
void collaboratorsBlockIsActuallyParsedNotSilentlyDropped(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("collaborators.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
collaborators:
|
||||
reviewer-alex:
|
||||
tab: "collab: alex"
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertEquals("collab: alex", cfg.fleet().collaborators().get("reviewer-alex").tab());
|
||||
}
|
||||
|
||||
/** A collaborator carries no field other than {@code tab}, so a blank one is meaningless. */
|
||||
@Test
|
||||
void aCollaboratorWithNoTabRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("useless-collaborator.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
collaborators:
|
||||
ghost: {}
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateMembers);
|
||||
assertTrue(e.getMessage().contains("ghost"), "the message must name the useless entry");
|
||||
assertTrue(e.getMessage().contains("tab:"), "the message must say what is missing");
|
||||
}
|
||||
|
||||
/** Control for {@link #aCollaboratorWithNoTabRefusesToStart}: a named tab loads cleanly. */
|
||||
@Test
|
||||
void aCollaboratorWithATabIsAllowed(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("named-collaborator.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
collaborators:
|
||||
reviewer-alex:
|
||||
tab: "collab: alex"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
assertDoesNotThrow(cfg::validateMembers);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #669: the fleet-wide {@code tabLabel} template can render as a collaborator tab, the
|
||||
* same hazard {@link #aFleetTabLabelTemplateThatCanRenderAsALeadTabRefusesToStart} covers on
|
||||
* the lead side. Drives the fleet-wide branch directly, with no profile override involved.
|
||||
*/
|
||||
@Test
|
||||
void aFleetTabLabelTemplateThatCanRenderAsACollaboratorTabRefusesToStart(@TempDir Path dir)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("collide-template-collaborator.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
tabLabel: "al{profile}"
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "alpha"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e =
|
||||
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
|
||||
assertTrue(e.getMessage().contains("fleet.tabLabel"));
|
||||
assertTrue(e.getMessage().contains("alex"), "the message must name the offending collaborator");
|
||||
}
|
||||
|
||||
/**
|
||||
* A member tabLabel that can render as a configured collaborator tab is the same hazard as the
|
||||
* lead case above — while its pane carries no entry in the spawned-member roster, a member
|
||||
* labelled that way is read back as the collaborator.
|
||||
*/
|
||||
@Test
|
||||
void aProfileTabLabelOverrideMatchingACollaboratorTabRefusesToStart(@TempDir Path dir)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("collide-collaborator.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
tabLabel: "collab-tab"
|
||||
fleet:
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "collab-tab"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e =
|
||||
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
|
||||
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
|
||||
assertTrue(e.getMessage().contains("collab-tab"), "the message must name the offending label");
|
||||
}
|
||||
|
||||
/** Control: a profile tabLabel that cannot render as the collaborator tab is allowed. */
|
||||
@Test
|
||||
void aProfileTabLabelThatCannotRenderAsACollaboratorTabIsAllowed(@TempDir Path dir)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("ok-collaborator.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
tabLabel: "worker-{profile}"
|
||||
fleet:
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "collab-tab"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
assertDoesNotThrow(cfg::validateLeadTabPrefixes);
|
||||
}
|
||||
|
||||
/** fleetd #669: identity is matched on a collaborator's exact tab, so two sharing one are unreachable. */
|
||||
@Test
|
||||
void twoCollaboratorsSharingTheSameExactTabRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("shared-collaborator-tab.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "shared tab"
|
||||
sam:
|
||||
tab: "Shared Tab"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e =
|
||||
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
|
||||
assertTrue(e.getMessage().contains("alex"), "the message must name one offending collaborator");
|
||||
assertTrue(e.getMessage().contains("sam"), "the message must name the other offending collaborator");
|
||||
}
|
||||
|
||||
/** Control for {@link #twoCollaboratorsSharingTheSameExactTabRefusesToStart}: distinct tabs load cleanly. */
|
||||
@Test
|
||||
void twoCollaboratorsWithDistinctExactTabsAreAllowed(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("distinct-collaborator-tabs.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "alex tab"
|
||||
sam:
|
||||
tab: "sam tab"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
assertDoesNotThrow(cfg::validateLeadTabPrefixes);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #669: a collaborator tab equal to a lead tab crosses a privilege boundary — the worst
|
||||
* of the three new collisions, since only one of the two identities is ever found.
|
||||
*/
|
||||
@Test
|
||||
void aCollaboratorTabEqualToALeadTabRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("lead-collaborator-collision.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "shared tab"
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "Shared Tab"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e =
|
||||
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
|
||||
assertTrue(e.getMessage().contains("opus"), "the message must name the offending lead");
|
||||
assertTrue(e.getMessage().contains("alex"), "the message must name the offending collaborator");
|
||||
}
|
||||
|
||||
/** Control for {@link #aCollaboratorTabEqualToALeadTabRefusesToStart}: distinct tabs load cleanly. */
|
||||
@Test
|
||||
void aLeadAndACollaboratorWithDistinctTabsAreAllowed(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("lead-collaborator-ok.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead tab"
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "collab tab"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
assertDoesNotThrow(cfg::validateLeadTabPrefixes);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #669: a pane-placed member can land in a collaborator's labelled tab exactly as it
|
||||
* can land in a lead's — {@code validatePanePlacementAgainstLeadTabs()} must fire even when
|
||||
* {@code fleet.leaders} is empty, which is the early-return the brief flagged as the bug.
|
||||
*/
|
||||
@Test
|
||||
void aPanePlacedProfileWithACollaboratorTabRefusesToStartEvenWithNoLeaders(@TempDir Path dir)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("pane-hazard-collaborator.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
placement: pane
|
||||
fleet:
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "collab: alex"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class,
|
||||
cfg::validatePanePlacementAgainstLeadTabs);
|
||||
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
|
||||
}
|
||||
|
||||
/** Control: a pane-placed profile with no lead or collaborator tab configured is allowed. */
|
||||
@Test
|
||||
void aPanePlacedProfileWithNoLeaderOrCollaboratorTabIsAllowed(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("pane-no-tab-at-all.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
placement: pane
|
||||
fleet:
|
||||
collaborators:
|
||||
alex: {}
|
||||
""");
|
||||
|
||||
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
|
||||
"a collaborator with no tab feeds nothing into the scanner, so pane placement is safe");
|
||||
}
|
||||
|
||||
// ── CB-548: the architects registry ────────────────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
@@ -1105,6 +1541,60 @@ class FleetConfigTest {
|
||||
assertEquals(Set.of("sonnet"), cfg.fleet().pool(MemberRole.REVIEWER).keySet());
|
||||
}
|
||||
|
||||
/**
|
||||
* A duplicated name in {@code fleet.collaborators} is refused at parse time, like any other
|
||||
* {@code fleet:} pool. See {@link #duplicateSlotNamesInOnePoolAreRejectedAtParseTime}.
|
||||
*/
|
||||
@Test
|
||||
void duplicateCollaboratorNamesAreRejectedAtParseTime(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("collaborator-dup.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "collab: alex"
|
||||
alex:
|
||||
tab: "collab: alex, second"
|
||||
""");
|
||||
|
||||
IllegalStateException e =
|
||||
assertThrows(IllegalStateException.class, () -> FleetConfig.load(f));
|
||||
assertTrue(e.getMessage().contains("alex"),
|
||||
"the refusal names the duplicated entry, was: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("fleet.collaborators"),
|
||||
"the refusal names the pool the duplicate is in, was: " + e.getMessage());
|
||||
}
|
||||
|
||||
/**
|
||||
* Control for {@link #duplicateCollaboratorNamesAreRejectedAtParseTime}: the same name reused
|
||||
* across the collaborators registry and a member role pool is the role × profile matrix doing
|
||||
* its job in the other pool, not a mistake — only a repeat within one pool loses an entry.
|
||||
*/
|
||||
@Test
|
||||
void theSameNameInCollaboratorsAndAnotherPoolIsNotADuplicate(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("collaborator-cross-pool.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
fleet:
|
||||
architects:
|
||||
alex:
|
||||
profile: sonnet
|
||||
collaborators:
|
||||
alex:
|
||||
tab: "collab: alex"
|
||||
""");
|
||||
|
||||
FleetConfig cfg = assertDoesNotThrow(() -> FleetConfig.load(f));
|
||||
assertEquals(Set.of("alex"), cfg.fleet().pool(MemberRole.ARCHITECT).keySet());
|
||||
assertEquals(Set.of("alex"), cfg.fleet().collaborators().keySet());
|
||||
}
|
||||
|
||||
@Test
|
||||
void duplicateKeysOutsideTheFleetPoolsAreUnaffected(@TempDir Path dir) throws Exception {
|
||||
// The duplicate check is scoped to the fleet pools — a duplicate elsewhere is not this
|
||||
@@ -3089,4 +3579,41 @@ class FleetConfigTest {
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertTrue(cfg.models().offIds().isEmpty());
|
||||
}
|
||||
|
||||
// ── fleetd #651: leadRollover.turnSettleSeconds default resolution ─────────────────────────
|
||||
|
||||
@Test
|
||||
void turnSettleSecondsDefaultsTo300WhenUnset(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bare-rollover.yaml");
|
||||
Files.writeString(f, "bind:\n port: 8080\nleadRollover: {}\n");
|
||||
|
||||
FleetConfig.LeadRollover rollover = FleetConfig.load(f).leadRollover();
|
||||
assertNotNull(rollover);
|
||||
assertEquals(300, rollover.turnSettleSeconds());
|
||||
}
|
||||
|
||||
@Test
|
||||
void turnSettleSecondsUsesAnExplicitPositiveValue(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("rollover.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
leadRollover:
|
||||
turnSettleSeconds: 45
|
||||
""");
|
||||
|
||||
FleetConfig.LeadRollover rollover = FleetConfig.load(f).leadRollover();
|
||||
assertEquals(45, rollover.turnSettleSeconds());
|
||||
}
|
||||
|
||||
@Test
|
||||
void turnSettleSecondsFallsBackTo300WhenZeroOrNegative(@TempDir Path dir) throws Exception {
|
||||
Path zero = dir.resolve("zero.yaml");
|
||||
Files.writeString(zero, "bind:\n port: 8080\nleadRollover:\n turnSettleSeconds: 0\n");
|
||||
assertEquals(300, FleetConfig.load(zero).leadRollover().turnSettleSeconds());
|
||||
|
||||
Path negative = dir.resolve("negative.yaml");
|
||||
Files.writeString(negative, "bind:\n port: 8080\nleadRollover:\n turnSettleSeconds: -5\n");
|
||||
assertEquals(300, FleetConfig.load(negative).leadRollover().turnSettleSeconds());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -17,62 +17,21 @@ import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* The gap this class exists to close: mutation testing on the fleetd ticket "central allow-list
|
||||
* of usable models" found that although {@link FleetConfig#validateModels()}'s own logic was well
|
||||
* pinned, nothing proved either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()})
|
||||
* still invoked it — deleting the call site left the full suite green (1478/0/0/0). A follow-up
|
||||
* measurement (same technique — remove one call site, run the suite, not read the code) found the
|
||||
* SAME gap for all five of {@link FleetConfig}'s other validators at startup, and for four of the
|
||||
* six inside {@link ConfigRef#reload()}. This is a class of gap, not one line's mistake: every one
|
||||
* of those thirteen tests called the validator itself directly, never the real caller that was
|
||||
* supposed to.
|
||||
* Tests the reflective validator sweep and {@link FleetConfig#validateAll()} reachability.
|
||||
*
|
||||
* <p>The fix replaces the six individual {@code cfg.validateXxx()} calls at each of the two real
|
||||
* call sites with one {@link FleetConfig#validateAll()}, which reaches every validator by
|
||||
* reflection rather than by a hand-maintained list of names. A hand-maintained list of six names
|
||||
* would have exactly the defect it replaces: the seventh validator someone adds next month has no
|
||||
* reason to be added to it, and nothing would say so. This class proves TWO separate claims, and
|
||||
* keeps them separate on purpose:
|
||||
*
|
||||
* <ol>
|
||||
* <li>{@link #theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames()} and its neighbours
|
||||
* prove the reflective sweep itself ({@link FleetConfig#invokeAllValidators}) is a general
|
||||
* mechanism — it runs whatever public, no-arg, void {@code validateXxx()} methods a class
|
||||
* happens to declare today, including a class with more of them than {@link FleetConfig}
|
||||
* has right now. This is the proof that a future, real seventh validator on {@link
|
||||
* FleetConfig} would be swept automatically, without needing to add a real (unwanted)
|
||||
* seventh validator just to exercise the claim.</li>
|
||||
* <li>{@link #validateAllReachesEveryOneOfTodaysSixValidators()} proves {@link
|
||||
* FleetConfig#validateAll()} itself is wired to that same generic mechanism and genuinely
|
||||
* reaches each of today's six real validators — reusing the exact minimal failing
|
||||
* configurations {@code FleetConfigTest} already established for each one directly, so a
|
||||
* single call to {@code validateAll()} is shown to reproduce every one of those six
|
||||
* failures.</li>
|
||||
* </ol>
|
||||
*
|
||||
* <p>Together with the direct-{@code Fleetd.main}-invocation tests in {@code
|
||||
* FleetdStartupValidationTest} (which prove the real startup call site still calls {@code
|
||||
* validateAll()}) and the {@code ConfigRefTest} reload tests (which prove the same for {@link
|
||||
* ConfigRef#reload()}), removing {@code cfg.validateAll();} from either real call site now fails
|
||||
* a test in this module.
|
||||
*
|
||||
* <p><b>What is NOT pinned, measured rather than assumed.</b> Reverting {@link
|
||||
* FleetConfig#validateAll()} to a hardcoded list of today's six method calls leaves the whole
|
||||
* suite green (measured at review: 1491 tests, 0 failures). Nothing ties {@code validateAll()} to
|
||||
* the generic sweep — claim 1 proves {@link FleetConfig#invokeAllValidators} is generic, and claim
|
||||
* 2 proves {@code validateAll()} reaches today's six, and a hardcoded list satisfies both. So the
|
||||
* reflective sweep is a convenience, not the guarantee. The guarantee is {@link
|
||||
* #fleetConfigDeclaresExactlyTheseSixValidatorsToday()}: it fails the moment a seventh validator
|
||||
* is declared, which forces whoever adds it to look at this file.
|
||||
* <p>{@link #theSweepRunsEveryValidateMethodOnAnUnrelatedClass()} and its neighbours
|
||||
* prove that {@link FleetConfig#invokeAllValidators} runs each public, no-arg, void
|
||||
* {@code validateXxx()} method on its target. {@link #fleetConfigDeclaresExactlyTheseValidatorsToday()}
|
||||
* is the canary for the validator set. {@link #validateAllReachesEveryOneOfTodaysRealValidators()}
|
||||
* is the reachability check for that set.
|
||||
*/
|
||||
class FleetConfigValidateAllTest {
|
||||
|
||||
// ── Claim 1: the reflective sweep is a general mechanism, not six names in disguise ──────────
|
||||
// ── Claim 1: the reflective sweep is a general mechanism ─────────────────────────────────────
|
||||
|
||||
/**
|
||||
* A throwaway fixture class, unrelated to {@link FleetConfig} in every way except shape: three
|
||||
* public, no-arg, void methods named {@code validateXxx}. Proves the sweep works on ANY class
|
||||
* with this shape, not on something special-cased to {@link FleetConfig}.
|
||||
* Fixture with public, no-arg, void methods named {@code validateXxx}. It proves the sweep uses
|
||||
* the target's method shape rather than special handling for {@link FleetConfig}.
|
||||
*/
|
||||
static class ThreeValidators {
|
||||
final List<String> ran = new ArrayList<>();
|
||||
@@ -91,7 +50,7 @@ class FleetConfigValidateAllTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames() {
|
||||
void theSweepRunsEveryValidateMethodOnAnUnrelatedClass() {
|
||||
ThreeValidators target = new ThreeValidators();
|
||||
FleetConfig.invokeAllValidators(target);
|
||||
assertEquals(List.of("validateAlpha", "validateBeta", "validateGamma"), target.ran,
|
||||
@@ -101,12 +60,8 @@ class FleetConfigValidateAllTest {
|
||||
}
|
||||
|
||||
/**
|
||||
* The core of the "self-maintaining" requirement: the exact same class shape as {@link
|
||||
* ThreeValidators}, plus one more method — standing in for "a developer adds a validator next
|
||||
* month". Nothing about the sweep changes to pick it up; the new method is invoked purely
|
||||
* because it exists and matches the shape. This is what makes adding a seventh real validator
|
||||
* to {@link FleetConfig} safe without touching {@link FleetConfig#validateAll()} or either
|
||||
* call site — there is no "wire it in" step left to forget.
|
||||
* Fixture with an added valid method. It proves the sweep reaches a method because it matches
|
||||
* the validator shape.
|
||||
*/
|
||||
static class FourValidators {
|
||||
final List<String> ran = new ArrayList<>();
|
||||
@@ -134,7 +89,7 @@ class FleetConfigValidateAllTest {
|
||||
FleetConfig.invokeAllValidators(target);
|
||||
assertEquals(List.of("validateAlpha", "validateBeta", "validateDelta", "validateGamma"),
|
||||
sorted(target.ran),
|
||||
"the fourth method must be reached automatically — proving a class can grow the "
|
||||
"the added method must be reached automatically — proving a class can grow the "
|
||||
+ "set of things it validates with no change to the sweep itself");
|
||||
}
|
||||
|
||||
@@ -209,18 +164,24 @@ class FleetConfigValidateAllTest {
|
||||
+ "name) must all be skipped");
|
||||
}
|
||||
|
||||
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches all six today ──
|
||||
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches every validator ──
|
||||
|
||||
/**
|
||||
* Reflectively enumerates {@link FleetConfig}'s own public, no-arg, void {@code validateXxx()}
|
||||
* methods (excluding {@code validateAll} itself) — the exact same filter {@link
|
||||
* FleetConfig#invokeAllValidators} applies. This is not the mechanism proof (that is claim 1,
|
||||
* above, on an unrelated class) — it is a visible denominator: today there are six, named
|
||||
* here, so a reader adding a seventh sees this assertion name the new count rather than a
|
||||
* silent pass at the old one.
|
||||
* above, on an unrelated class) — it is a visible denominator: today there are eight, named in
|
||||
* the {@code Set.of} below, so a reader adding or removing one sees this assertion name the new
|
||||
* count rather than a silent pass at the old one. The count lives only in that set, not in this
|
||||
* method's name, so the two cannot drift apart.
|
||||
*
|
||||
* <p>This assertion alone proves only that the validator exists with the right shape — it
|
||||
* cannot prove {@code validateAll()} actually reaches it. Only {@link
|
||||
* #validateAllReachesEveryOneOfTodaysRealValidators()} proves reachability, which is why this
|
||||
* method's failure message sends the reader there too.
|
||||
*/
|
||||
@Test
|
||||
void fleetConfigDeclaresExactlyTheseSixValidatorsToday() {
|
||||
void fleetConfigDeclaresExactlyTheseValidatorsToday() {
|
||||
Set<String> names = new TreeSet<>();
|
||||
for (Method m : FleetConfig.class.getMethods()) {
|
||||
if (java.lang.reflect.Modifier.isPublic(m.getModifiers())
|
||||
@@ -233,12 +194,18 @@ class FleetConfigValidateAllTest {
|
||||
}
|
||||
assertEquals(new TreeSet<>(Set.of("validateAuthExposure", "validateLeadTabPrefixes",
|
||||
"validateSubscriptionProfiles", "validateCharters", "validateMembers",
|
||||
"validateModels", "validateLeadRollover")), names,
|
||||
"FleetConfig's public validate*() methods changed. Do TWO things, in this "
|
||||
"validateModels", "validateLeadRollover", "validatePanePlacementAgainstLeadTabs")),
|
||||
names,
|
||||
"FleetConfig's public validate*() methods changed. Do THREE things, in this "
|
||||
+ "order. First confirm validateAll() still delegates to "
|
||||
+ "invokeAllValidators(this) — a hardcoded list there passes every other "
|
||||
+ "test in this class, so this assertion is the only place that will ever "
|
||||
+ "make you check. Only then update the expected set to match.");
|
||||
+ "make you check. Second, update the expected set below to match. Third, "
|
||||
+ "add or remove a case for that validator in "
|
||||
+ "validateAllReachesEveryOneOfTodaysRealValidators() below — this "
|
||||
+ "assertion proves only that the validator exists with the right shape, "
|
||||
+ "never that validateAll() reaches it; that enumeration is the test that "
|
||||
+ "does.");
|
||||
}
|
||||
|
||||
/** A minimal, otherwise-valid file — same shape FleetConfigTest and ConfigRefTest use. */
|
||||
@@ -262,14 +229,21 @@ class FleetConfigValidateAllTest {
|
||||
}
|
||||
|
||||
/**
|
||||
* The heart of claim 2: for each of today's six real validators, a minimal file that fails
|
||||
* ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
|
||||
* directly — must also fail through {@link FleetConfig#validateAll()}. If a future edit to
|
||||
* {@code validateAll()} silently dropped one validator from the sweep (e.g. a typo'd name
|
||||
* filter), exactly one of these six would start passing when it must not.
|
||||
* The heart of claim 2: for every one of today's real validators, a minimal file that
|
||||
* fails ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
|
||||
* directly, or a dedicated minimal fixture where no other test drives that validator through
|
||||
* {@code validateAll()} — must also fail through {@link FleetConfig#validateAll()}. If a
|
||||
* future edit to {@code validateAll()} silently dropped one of these from the sweep
|
||||
* (e.g. a typo'd name filter), exactly one of them would start passing when it must not.
|
||||
*
|
||||
* <p>This is the single place that proves {@code validateAll()} reaches a given validator.
|
||||
* Adding or removing a validator on {@link FleetConfig} must add or remove a case here, not
|
||||
* only an updated name in {@link #fleetConfigDeclaresExactlyTheseValidatorsToday()}'s expected
|
||||
* set — that assertion proves the validator's shape, never that {@code validateAll()} reaches
|
||||
* it.
|
||||
*/
|
||||
@Test
|
||||
void validateAllReachesEveryOneOfTodaysSixValidators(@TempDir Path dir) throws Exception {
|
||||
void validateAllReachesEveryOneOfTodaysRealValidators(@TempDir Path dir) throws Exception {
|
||||
// validateAuthExposure: a non-loopback bind without token mode.
|
||||
assertValidateAllRefuses(dir, "auth-exposure.yaml", """
|
||||
bind:
|
||||
@@ -277,16 +251,16 @@ class FleetConfigValidateAllTest {
|
||||
port: 8765
|
||||
""", "auth.mode: token");
|
||||
|
||||
// validateLeadTabPrefixes: a fleet-wide tabLabel that starts with a lead's own tabPrefix.
|
||||
// validateLeadTabPrefixes: a fleet-wide tabLabel that equals a lead tab.
|
||||
assertValidateAllRefuses(dir, "lead-tab-prefixes.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
fleet:
|
||||
tabLabel: "lead: {role} {profile}"
|
||||
tabLabel: "alpha"
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
tab: "alpha"
|
||||
""", "fleet.tabLabel");
|
||||
|
||||
// validateSubscriptionProfiles: subscription: true with env: reseating ANTHROPIC_BASE_URL.
|
||||
@@ -342,6 +316,29 @@ class FleetConfigValidateAllTest {
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
""", "rogue");
|
||||
|
||||
// validatePanePlacementAgainstLeadTabs: a pane-placed profile while a lead names a tab.
|
||||
assertValidateAllRefuses(dir, "pane-placement.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
gx10:
|
||||
placement: pane
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
""", "gx10");
|
||||
|
||||
// validateLeadRollover: a leadRollover: block present with no handoverPath.
|
||||
assertValidateAllRefuses(dir, "lead-rollover.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
leadRollover:
|
||||
requireOperatorConfirm: false
|
||||
""", "handoverPath");
|
||||
}
|
||||
|
||||
private static void assertValidateAllRefuses(Path dir, String fileName, String yaml,
|
||||
|
||||
+1
-1
@@ -103,7 +103,7 @@ class FleetConfigWithDefaultsPreservesEveryComponentTest {
|
||||
// comment there), same as broker/primary/leadHeartbeat/... above — a real, non-null value
|
||||
// here proves it, rather than leaving it null and proving nothing.
|
||||
v.put("leadRollover", new FleetConfig.LeadRollover(
|
||||
"/handover/guard.md", true, 3600, 20, 20, "read the handover file"));
|
||||
"/handover/guard.md", true, 3600, 20, 45, "read the handover file"));
|
||||
assertNamesMatchComponents(v);
|
||||
return v;
|
||||
}
|
||||
|
||||
@@ -7,6 +7,7 @@ import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.CopyOnWriteArrayList;
|
||||
|
||||
@@ -59,6 +60,10 @@ public final class FakeHerdr implements HerdrClient {
|
||||
private Runnable onAgentStart; // fires the instant agent.start is called — see onAgentStart(Runnable)
|
||||
private volatile int agentGetOkCalls = Integer.MAX_VALUE; // how many agent.get calls succeed first
|
||||
private volatile String agentGetFailCode = null; // error code every agent.get call after that reports
|
||||
/** pane ids that {@link #paneGoneAfterClose} has opted into reporting gone — see that method. */
|
||||
private final Set<String> paneGoneAfterCloseIds = ConcurrentHashMap.newKeySet();
|
||||
/** pane ids a {@code pane.close} call has actually reached, for {@link #paneGoneAfterCloseIds}. */
|
||||
private final Set<String> closedPaneIds = ConcurrentHashMap.newKeySet();
|
||||
|
||||
public FakeHerdr healthy(boolean h) {
|
||||
this.healthy = h;
|
||||
@@ -220,6 +225,19 @@ public final class FakeHerdr implements HerdrClient {
|
||||
return this;
|
||||
}
|
||||
|
||||
/**
|
||||
* Make {@code pane.get(paneId)} report the pane gone (a {@code pane_not_found} {@link
|
||||
* HerdrException}, exactly as {@link WorkspaceControl#locatePane} expects to see once a pane
|
||||
* has really disappeared) once a {@code pane.close} call for that same {@code paneId} has
|
||||
* actually reached this fake. Every other pane, and this pane before its own close, keeps
|
||||
* reporting the default canned {@code pane.get} response — opt-in, by pane id, so no existing
|
||||
* test's {@code pane.get} behaviour changes.
|
||||
*/
|
||||
public FakeHerdr paneGoneAfterClose(String paneId) {
|
||||
paneGoneAfterCloseIds.add(paneId);
|
||||
return this;
|
||||
}
|
||||
|
||||
/**
|
||||
* Run {@code hook} synchronously the instant an {@code agent.start} call reaches this fake —
|
||||
* i.e. the instant the peer PROCESS would start against a real herdr daemon. A test uses this
|
||||
@@ -422,9 +440,18 @@ public final class FakeHerdr implements HerdrClient {
|
||||
}
|
||||
yield mapper.readTree("{\"type\":\"ok\"}");
|
||||
}
|
||||
case "pane.get" -> mapper.readTree("""
|
||||
case "pane.get" -> {
|
||||
Object paneIdParam = params instanceof Map<?, ?> m ? m.get("pane_id") : null;
|
||||
String paneIdKey = paneIdParam == null ? null : String.valueOf(paneIdParam);
|
||||
if (paneIdKey != null && paneGoneAfterCloseIds.contains(paneIdKey)
|
||||
&& closedPaneIds.contains(paneIdKey)) {
|
||||
throw new HerdrException("herdr error [pane_not_found]: pane.get failed",
|
||||
"pane_not_found", null);
|
||||
}
|
||||
yield mapper.readTree("""
|
||||
{"type":"pane_info","pane":{"pane_id":"w9:pW","workspace_id":"w9",
|
||||
"tab_id":"w9:t2","agent_status":"idle"}}""");
|
||||
}
|
||||
case "pane.list" -> noPanes
|
||||
? mapper.readTree("{\"type\":\"pane_list\",\"panes\":[]}")
|
||||
: mapper.readTree("""
|
||||
@@ -458,6 +485,9 @@ public final class FakeHerdr implements HerdrClient {
|
||||
throw new HerdrException("herdr error [" + code + "]: pane.close failed",
|
||||
code, null);
|
||||
}
|
||||
if (paneIdParam != null) {
|
||||
closedPaneIds.add(String.valueOf(paneIdParam));
|
||||
}
|
||||
yield mapper.readTree("{\"type\":\"ok\"}");
|
||||
}
|
||||
default -> throw new HerdrException("fake has no canned response for " + method);
|
||||
|
||||
@@ -2,6 +2,8 @@ package dev.ltms.fleet.herdr;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertSame;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotSame;
|
||||
|
||||
@@ -28,4 +30,44 @@ class HerdrRouterTest {
|
||||
assertSame(router.leadAgents(), router.agentsFor("lead"));
|
||||
assertSame(router.memberAgents(), router.agentsFor("member"));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #669 Unit E. The predicate shape here is exactly what {@code FleetdAssembly} builds:
|
||||
* true when the terminal is a known lead OR a known collaborator. Two distinct clients are
|
||||
* required — with one shared client {@code agentsFor} would return the same object regardless
|
||||
* of the predicate's answer, and this assertion would pass whether or not the collaborator map
|
||||
* was ever consulted.
|
||||
*/
|
||||
@Test
|
||||
void collaboratorTerminalRoutesToTheLeadDaemon() {
|
||||
FakeHerdr lead = new FakeHerdr();
|
||||
FakeHerdr member = new FakeHerdr();
|
||||
Map<String, String> leads = Map.of("term_lead", "primary");
|
||||
Map<String, String> collaborators = Map.of("term_collab", "reviewer-alex");
|
||||
HerdrRouter router = new HerdrRouter(lead, member,
|
||||
id -> leads.containsKey(id) || collaborators.containsKey(id));
|
||||
|
||||
assertSame(router.leadAgents(), router.agentsFor("term_collab"),
|
||||
"a configured collaborator's terminal must route to the LEAD daemon, not the member "
|
||||
+ "one — its pane is opened by a person, exactly like a lead's");
|
||||
}
|
||||
|
||||
/**
|
||||
* Companion to {@link #collaboratorTerminalRoutesToTheLeadDaemon}: a terminal that is neither a
|
||||
* known lead nor a known collaborator must still route to the member daemon. Without this, a
|
||||
* predicate of {@code _ -> true} would also pass the test above.
|
||||
*/
|
||||
@Test
|
||||
void terminalInNeitherMapStillRoutesToTheMemberDaemon() {
|
||||
FakeHerdr lead = new FakeHerdr();
|
||||
FakeHerdr member = new FakeHerdr();
|
||||
Map<String, String> leads = Map.of("term_lead", "primary");
|
||||
Map<String, String> collaborators = Map.of("term_collab", "reviewer-alex");
|
||||
HerdrRouter router = new HerdrRouter(lead, member,
|
||||
id -> leads.containsKey(id) || collaborators.containsKey(id));
|
||||
|
||||
assertSame(router.memberAgents(), router.agentsFor("term_worker"),
|
||||
"a terminal absent from both maps must stay on the member daemon — the fix widens "
|
||||
+ "the predicate, it does not make it unconditionally true");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -163,6 +163,13 @@ class LeadTabScannerTest {
|
||||
return new LeadTabScanner(herdr, tabToName, Set.of("fleetd-workers"), TTL, clock::get);
|
||||
}
|
||||
|
||||
private LeadTabScanner scannerWithCollaborators(TopologyHerdr herdr, Map<String, String> tabToName,
|
||||
Map<String, String> collaboratorTabToName,
|
||||
AtomicLong clock) {
|
||||
return new LeadTabScanner(herdr, tabToName, collaboratorTabToName,
|
||||
Set.of("fleetd-workers"), TTL, clock::get);
|
||||
}
|
||||
|
||||
@Test
|
||||
void everyConfiguredTabBecomesALeadNamedByItsEntry() {
|
||||
Map<String, String> leads = scanner(twoLeads(), twoLeadsConfigured(), new AtomicLong()).get();
|
||||
@@ -198,9 +205,12 @@ class LeadTabScannerTest {
|
||||
}
|
||||
|
||||
/**
|
||||
* The guard that matters: fleetd labels its own worker tabs, so if a worker space were scanned
|
||||
* a naming accident would promote the fleet. The exclusion is by workspace, not by hoping the
|
||||
* worker template never collides.
|
||||
* Covers the {@code excludedWorkspaceLabels} parameter: a tab in an excluded workspace is never
|
||||
* matched, whatever its label. Production always constructs this class with an empty set (CB-558,
|
||||
* {@code FleetdAssembly}), so this parameter plays no part in the live guard against a worker
|
||||
* tab being mistaken for a lead — that guard is {@code CallerResolver} asking the live
|
||||
* spawned-member roster before any tab map. This test exists because the parameter still exists
|
||||
* and is worth covering on its own terms.
|
||||
*/
|
||||
@Test
|
||||
void aTabInAWorkerSpaceIsNeverALeadEvenWhenItsLabelMatches() {
|
||||
@@ -212,6 +222,29 @@ class LeadTabScannerTest {
|
||||
assertFalse(scanner(herdr, tabToName, new AtomicLong()).get().containsKey("term_impostor"));
|
||||
}
|
||||
|
||||
/**
|
||||
* The shipped shape (CB-558, {@code FleetdAssembly}): production always constructs this class
|
||||
* with an empty {@code excludedWorkspaceLabels}, and a lead's {@code workspace:} default is the
|
||||
* same shared {@code "fleet"} space the members use. A lead tab is still found when it sits in
|
||||
* the exact same workspace as a member-labelled tab — the scanner tells them apart by the exact
|
||||
* tab label, not by which workspace either one is in.
|
||||
*/
|
||||
@Test
|
||||
void aLeadIsDiscoveredWhenItsWorkspaceIsTheSameAsTheMemberWorkspace() {
|
||||
TopologyHerdr herdr = new TopologyHerdr()
|
||||
.workspace("w1", "fleet")
|
||||
.tab("w1:t1", "w1", "lead: opus-5.0")
|
||||
.tab("w1:t2", "w1", "worker: gx10 #1")
|
||||
.pane("w1:p1", "w1:t1", "term_opus")
|
||||
.pane("w1:p2", "w1:t2", "term_worker");
|
||||
LeadTabScanner s = new LeadTabScanner(herdr, Map.of("lead: opus-5.0", "opus-5.0"),
|
||||
Set.of(), TTL, new AtomicLong()::get);
|
||||
|
||||
assertEquals("opus-5.0", s.get().get("term_opus"),
|
||||
"a lead sharing the members' workspace is still discovered — the label, not the "
|
||||
+ "workspace, is what matches it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aLabelWithNoConfiguredEntryIsIgnored() {
|
||||
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
|
||||
@@ -232,8 +265,9 @@ class LeadTabScannerTest {
|
||||
|
||||
@Test
|
||||
void everyPaneInALeadTabResolvesAsThatLead() {
|
||||
// A human may split their own lead tab. Both panes are theirs, so both are that lead —
|
||||
// nothing fleetd placed can land here (see the worker-space test above).
|
||||
// A human may split their own lead tab. Both panes are theirs, so both are that lead.
|
||||
// A pane-placed member landing here instead is refused at startup by
|
||||
// FleetConfig.validatePanePlacementAgainstLeadTabs, not by this scanner.
|
||||
TopologyHerdr herdr = twoLeads().pane("w1:p1b", "w1:t1", "term_opus_split");
|
||||
|
||||
assertEquals("opus-5.0",
|
||||
@@ -463,4 +497,75 @@ class LeadTabScannerTest {
|
||||
|
||||
assertEquals(afterFirst, herdr.calls, "the failure path must be rate-limited too");
|
||||
}
|
||||
|
||||
// ── fleetd #669 Unit D: a collaborator tab is matched the same way as a lead tab, one pass ─────
|
||||
|
||||
@Test
|
||||
void aConfiguredCollaboratorTabIsReportedByCollaboratorsNotByGet() {
|
||||
TopologyHerdr herdr = new TopologyHerdr()
|
||||
.workspace("w1", "main")
|
||||
.tab("w1:t1", "w1", "collab: ops")
|
||||
.pane("w1:p1", "w1:t1", "term_ops");
|
||||
LeadTabScanner s = scannerWithCollaborators(herdr, Map.of(), Map.of("collab: ops", "ops"),
|
||||
new AtomicLong());
|
||||
|
||||
assertEquals(Map.of("term_ops", "ops"), s.collaborators(),
|
||||
"a collaborator tab is matched exactly like a lead tab");
|
||||
assertEquals(Map.of(), s.get(), "a collaborator tab must never also appear as a lead");
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 4, scanner level: a collaborator tab that is labelled but runs no agent is not
|
||||
* reported — the same #359 liveness cross-check a lead tab gets.
|
||||
*/
|
||||
@Test
|
||||
void aDeadCollaboratorTabIsNotReported() {
|
||||
TopologyHerdr herdr = new TopologyHerdr()
|
||||
.workspace("w1", "main")
|
||||
.tab("w1:t1", "w1", "collab: ops")
|
||||
.pane("w1:p1", "w1:t1", "term_ops")
|
||||
.deadAgent("w1:t1");
|
||||
LeadTabScanner s = scannerWithCollaborators(herdr, Map.of(), Map.of("collab: ops", "ops"),
|
||||
new AtomicLong());
|
||||
|
||||
assertFalse(s.collaborators().containsKey("term_ops"),
|
||||
"a dead collaborator tab must never resolve as a live collaborator");
|
||||
}
|
||||
|
||||
/**
|
||||
* Both kinds are matched in a single pass over the same tab list — not a second scanner, not a
|
||||
* second scan. Proven by herdr call count: scanning one lead tab and one collaborator tab in the
|
||||
* same instance costs exactly as many calls as scanning two lead tabs in {@link #twoLeads()}.
|
||||
*/
|
||||
@Test
|
||||
void leadsAndCollaboratorsAreMatchedInOnePassOverTheSameScan() {
|
||||
TopologyHerdr oneOfEach = new TopologyHerdr()
|
||||
.workspace("w1", "main")
|
||||
.tab("w1:t1", "w1", "lead: opus-5.0")
|
||||
.tab("w1:t2", "w1", "collab: ops")
|
||||
.pane("w1:p1", "w1:t1", "term_opus")
|
||||
.pane("w1:p2", "w1:t2", "term_ops");
|
||||
LeadTabScanner s = scannerWithCollaborators(oneOfEach, Map.of("lead: opus-5.0", "opus-5.0"),
|
||||
Map.of("collab: ops", "ops"), new AtomicLong());
|
||||
|
||||
assertEquals(Map.of("term_opus", "opus-5.0"), s.get());
|
||||
assertEquals(Map.of("term_ops", "ops"), s.collaborators());
|
||||
|
||||
TopologyHerdr twoLeadsBaseline = twoLeads();
|
||||
scanner(twoLeadsBaseline, twoLeadsConfigured(), new AtomicLong()).get();
|
||||
|
||||
assertEquals(twoLeadsBaseline.calls, oneOfEach.calls,
|
||||
"one lead tab + one collaborator tab must cost exactly as many herdr calls as two "
|
||||
+ "lead tabs — proof this is one pass, not a second scan");
|
||||
}
|
||||
|
||||
/** Regression: with no collaborators configured, every existing lead-only behaviour is unchanged. */
|
||||
@Test
|
||||
void anEmptyCollaboratorMapLeavesCollaboratorsEmptyAndGetUnaffected() {
|
||||
LeadTabScanner s = scannerWithCollaborators(twoLeads(), twoLeadsConfigured(), Map.of(),
|
||||
new AtomicLong());
|
||||
|
||||
assertEquals(Map.of(), s.collaborators());
|
||||
assertEquals(Map.of("term_opus", "opus-5.0", "term_gpt", "gpt-sol-5.6"), s.get());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -81,12 +81,6 @@ class LeadContextGaugeHighThresholdTest {
|
||||
assertEquals(LeadContextGauge.State.HIGH,
|
||||
atGauge.read(atConfigDir, SESSION_ID, "claude", null).state(),
|
||||
"the fixed default must still be 200,000 when no window is resolvable");
|
||||
|
||||
String legacyConfigDir = writeTranscript(tmp.resolve("legacy"), SESSION_ID, 200_000);
|
||||
LeadContextGauge legacyGauge = new LeadContextGauge();
|
||||
assertEquals(LeadContextGauge.State.HIGH,
|
||||
legacyGauge.read(legacyConfigDir, SESSION_ID, "claude").state(),
|
||||
"the 3-arg read() (no window argument at all) must behave exactly like passing a null window");
|
||||
}
|
||||
|
||||
@Test
|
||||
|
||||
@@ -85,7 +85,7 @@ class LeadContextGaugeTest {
|
||||
usageLine(40_000, 5_000, 3_000)); // last record: 48,000
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
|
||||
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude");
|
||||
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude", null);
|
||||
assertEquals(48_000L, first.tokens(), "must total input+cache_read+cache_creation of the LAST usage record");
|
||||
assertEquals(LeadContextGauge.State.OK, first.state());
|
||||
|
||||
@@ -93,7 +93,7 @@ class LeadContextGaugeTest {
|
||||
// must change with it, not stay pinned to the first fixture's total.
|
||||
String otherSession = "22222222-2222-2222-2222-222222222222";
|
||||
writeTranscript(tmp, otherSession, usageLine(100_000, 50_000, 50_000)); // last record: 200,000
|
||||
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude");
|
||||
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude", null);
|
||||
assertEquals(200_000L, second.tokens());
|
||||
assertTrue(second.tokens() != first.tokens(), "changing N in the fixture must change the reported number");
|
||||
}
|
||||
@@ -111,12 +111,12 @@ class LeadContextGaugeTest {
|
||||
compactionLine(),
|
||||
usageLine(3_000, 0, 0));
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude");
|
||||
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude", null);
|
||||
assertEquals(2, twoCompactions.compactions());
|
||||
|
||||
String sessionZeroCompactions = "44444444-4444-4444-4444-444444444444";
|
||||
writeTranscript(tmp, sessionZeroCompactions, usageLine(3_000, 0, 0));
|
||||
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude");
|
||||
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude", null);
|
||||
assertEquals(0, zeroCompactions.compactions(), "changing K in the fixture must change the reported count");
|
||||
}
|
||||
|
||||
@@ -126,7 +126,7 @@ class LeadContextGaugeTest {
|
||||
@DisplayName("a missing transcript file reports UNKNOWN with no token number")
|
||||
void missingFileIsUnknown(@TempDir Path tmp) {
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
|
||||
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude", null);
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
|
||||
assertNull(reading.tokens());
|
||||
}
|
||||
@@ -148,7 +148,7 @@ class LeadContextGaugeTest {
|
||||
assumeFalse(Files.isReadable(file),
|
||||
"runs as root (CI container): the read bit does not stop root, so this case cannot be set up here");
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
|
||||
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude", null);
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
|
||||
assertNull(reading.tokens());
|
||||
} finally {
|
||||
@@ -171,7 +171,7 @@ class LeadContextGaugeTest {
|
||||
Files.writeString(file, lastCompleteLine + "\n" + tornLine, StandardCharsets.UTF_8);
|
||||
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
|
||||
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude", null);
|
||||
assertEquals(LeadContextGauge.State.OK, reading.state(),
|
||||
"a torn final line must not turn a good earlier reading into UNKNOWN");
|
||||
assertEquals(6_000L, reading.tokens(),
|
||||
@@ -187,7 +187,7 @@ class LeadContextGaugeTest {
|
||||
"{this is not json at all",
|
||||
"neither is this{{{");
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
|
||||
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude", null);
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
|
||||
"every line unparseable is the real format-change signal and must still report UNKNOWN");
|
||||
assertNull(reading.tokens());
|
||||
@@ -224,12 +224,12 @@ class LeadContextGaugeTest {
|
||||
AtomicLong now = new AtomicLong(0);
|
||||
LeadContextGauge gauge = new LeadContextGauge(now::get, 5_000);
|
||||
|
||||
gauge.read(configDir, SESSION_ID, "claude");
|
||||
gauge.read(configDir, SESSION_ID, "claude"); // still inside the TTL window
|
||||
gauge.read(configDir, SESSION_ID, "claude", null);
|
||||
gauge.read(configDir, SESSION_ID, "claude", null); // still inside the TTL window
|
||||
assertEquals(1, gauge.diskReadCount(), "two reads inside the TTL must touch disk once");
|
||||
|
||||
now.set(6_000); // past the TTL
|
||||
gauge.read(configDir, SESSION_ID, "claude");
|
||||
gauge.read(configDir, SESSION_ID, "claude", null);
|
||||
assertEquals(2, gauge.diskReadCount(), "a read past the TTL must touch disk again");
|
||||
}
|
||||
|
||||
@@ -241,8 +241,8 @@ class LeadContextGaugeTest {
|
||||
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode").state());
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null).state());
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude").state());
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode", null).state());
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null, null).state());
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude", null).state());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,10 +1,16 @@
|
||||
package dev.ltms.fleet.lead;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.ResilientAgentLaunch;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
@@ -64,6 +70,21 @@ class LeadLauncherTest {
|
||||
return new LeadLauncher(new AgentControl(herdr), new WorkspaceControl(herdr), cfg);
|
||||
}
|
||||
|
||||
/** As {@link #launcher}, plus a fast no-op sleeper so a busy-retry test never real-sleeps. */
|
||||
private static LeadLauncher fastLauncher(FakeHerdr herdr, FleetConfig cfg) {
|
||||
return new LeadLauncher(new AgentControl(herdr), new WorkspaceControl(herdr), cfg, () -> { });
|
||||
}
|
||||
|
||||
/** {@link #opusProfile()} with one argv element long enough to overflow the pane line limit. */
|
||||
private static FleetConfig.Profile hugeArgvProfile() {
|
||||
return new FleetConfig.Profile(
|
||||
"opus", null, "claude-opus-5", null, "FLEETD_WORKER_TOKEN",
|
||||
List.of("ccs", "x".repeat(1500)), "tab", "fleet", null,
|
||||
"http://127.0.0.1:8765/mcp", null, null,
|
||||
null, null, null,
|
||||
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true, null);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static List<String> startedArgs(FakeHerdr herdr) {
|
||||
return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("args");
|
||||
@@ -86,7 +107,11 @@ class LeadLauncherTest {
|
||||
|
||||
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
|
||||
assertTrue(herdr.called("agent.start"), "a lead must actually be started");
|
||||
assertEquals("lead-opus", startedName(herdr));
|
||||
// fleetd #727: the name carries a per-process nonce and a per-start sequence number — the
|
||||
// same unique-naming scheme HerdrPeerLauncher uses for members — rather than the fixed
|
||||
// "lead-opus" a stale registry entry could block a legitimate relaunch under.
|
||||
assertTrue(startedName(herdr).matches("lead-opus-[0-9a-f]{6}-\\d+"),
|
||||
"name is lead-<name>-<nonce>-<seq>: " + startedName(herdr));
|
||||
}
|
||||
|
||||
/** The tab is labelled with the configured `tab:` so the scanner finds the lead on the next resolve. */
|
||||
@@ -436,4 +461,247 @@ class LeadLauncherTest {
|
||||
assertEquals(0, launcher(herdr, cfg).ensureLeads());
|
||||
assertTrue(herdr.calls.isEmpty(), "nothing declared ⇒ nothing scanned");
|
||||
}
|
||||
|
||||
// ── fleetd #727: the same three launch protections every member spawn gets ───────────────────
|
||||
|
||||
/**
|
||||
* A freshly created pane may not have redrawn its prompt yet, so herdr answers
|
||||
* {@code agent_pane_busy}. The lead launch must wait it out rather than fail on the first miss
|
||||
* — exactly the retry {@code HerdrPeerLauncher} already gives every member.
|
||||
*/
|
||||
@Test
|
||||
void aLeadLaunchRetriesWhileTheSeedShellBoots() {
|
||||
FakeHerdr herdr = new FakeHerdr().agentPaneBusyTimes(2);
|
||||
|
||||
assertEquals(1, fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
|
||||
"the lead must still start once the shell is ready");
|
||||
assertEquals(3, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
|
||||
"two busy rejections, then the successful start");
|
||||
}
|
||||
|
||||
/**
|
||||
* The busy retry is bounded, not an infinite poll. If the pane never becomes ready the launch
|
||||
* must eventually give up and log the failure, not hang the daemon's reconcile loop forever —
|
||||
* proven here by a budget that would still be busy on attempt
|
||||
* {@value ResilientAgentLaunch#SHELL_READY_RETRIES} and a call count that stops exactly there.
|
||||
*/
|
||||
@Test
|
||||
void aLeadLaunchGivesUpAfterTheBoundedBusyBudgetRatherThanLoopingForever() {
|
||||
FakeHerdr herdr = new FakeHerdr().agentPaneBusyTimes(999);
|
||||
|
||||
assertEquals(0, fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
|
||||
"a pane that never becomes ready must not be reported as a started lead");
|
||||
assertEquals(ResilientAgentLaunch.SHELL_READY_RETRIES,
|
||||
herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
|
||||
"the retry budget is bounded: it stops after exactly SHELL_READY_RETRIES attempts");
|
||||
}
|
||||
|
||||
/**
|
||||
* herdr refuses a duplicate agent {@code name} outright. A stale registry entry — a crashed
|
||||
* lead session, or a name the registry has not yet released — must not permanently block a
|
||||
* legitimate relaunch, so each retry attempt carries a fresh per-start name.
|
||||
*/
|
||||
@Test
|
||||
@SuppressWarnings("unchecked")
|
||||
void aLeadLaunchRetriesUnderAFreshNameWhenTheOldNameIsStillTaken() {
|
||||
FakeHerdr herdr = new FakeHerdr().agentNameTakenTimes(2);
|
||||
|
||||
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
|
||||
"the lead must still start once a free name is found");
|
||||
List<String> names = herdr.calls.stream()
|
||||
.filter(c -> c.method().equals("agent.start"))
|
||||
.map(c -> ((Map<String, Object>) c.params()).get("name").toString())
|
||||
.toList();
|
||||
assertEquals(3, names.size(), "2 rejected + 1 success");
|
||||
assertEquals(3, Set.copyOf(names).size(), "each attempt must use a distinct name");
|
||||
}
|
||||
|
||||
/**
|
||||
* The name-collision retry is bounded too. If the name is taken on every attempt, the launch
|
||||
* must give up rather than keep minting new names forever — the per-start naming scheme means
|
||||
* every failed attempt was refused outright by herdr (no process exists under a name herdr
|
||||
* refused), so a bounded, exhausted retry never leaves a second process running: nothing is
|
||||
* running at all.
|
||||
*/
|
||||
@Test
|
||||
void aLeadLaunchNeverEndsUpWithASecondProcessWhenTheNameStaysTaken() {
|
||||
FakeHerdr herdr = new FakeHerdr().agentNameTakenTimes(999);
|
||||
|
||||
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
|
||||
"a name that is never free must not be reported as a started lead");
|
||||
assertEquals(ResilientAgentLaunch.NAME_RETRIES,
|
||||
herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
|
||||
"the retry budget is bounded: it stops after exactly NAME_RETRIES attempts, "
|
||||
+ "never racing a duplicate into existence");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #220/#727: herdr types the launch command into the pane as one line, and a pty line
|
||||
* buffer holds only 1024 bytes — past that the tail is dropped with no error at all, and the
|
||||
* backend exits on a mangled argument. The lead launch must refuse an over-long command outright
|
||||
* rather than let it be typed and silently truncated.
|
||||
*/
|
||||
@Test
|
||||
void anOverlongLeadArgvIsRefusedRatherThanTypedAndTruncated() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(LeadLauncher.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
|
||||
int started;
|
||||
try {
|
||||
started = launcher(herdr, configWith(lead("opus", "lead: opus", 1), hugeArgvProfile()))
|
||||
.ensureLeads();
|
||||
} finally {
|
||||
logger.detachAppender(appender);
|
||||
}
|
||||
|
||||
assertEquals(0, started, "an over-long command must never be reported as a started lead");
|
||||
assertFalse(herdr.called("agent.start"),
|
||||
"nothing may be started — a truncated command is worse than no spawn");
|
||||
String warn = appender.list.stream()
|
||||
.filter(e -> e.getLevel().equals(Level.WARN))
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.filter(m -> m.contains("failed to launch"))
|
||||
.findFirst()
|
||||
.orElseThrow(() -> new AssertionError("expected a WARN naming the launch failure: "
|
||||
+ appender.list));
|
||||
assertTrue(warn.contains("1024"), "names the limit: " + warn);
|
||||
assertTrue(warn.contains("opus"), "names the profile: " + warn);
|
||||
assertTrue(warn.contains("x".repeat(60)), "names the culprit argument: " + warn);
|
||||
}
|
||||
|
||||
// ── fleetd #726 unit 1: the single-lead relaunch seam ─────────────────────────────────────
|
||||
|
||||
/**
|
||||
* The returned agent's {@code terminalId()}/{@code paneId()} are the ones the fake
|
||||
* {@code AgentControl} actually started — not a coincidental field left over from the caller.
|
||||
* {@code paneId()} echoes the exact {@code pane_id} the launch's own {@code agent.start} call
|
||||
* carried (protocol 19: the agent starts into the pane it is asked to), and {@code
|
||||
* terminalId()} is herdr's own generated id, which the fake always shapes as {@code
|
||||
* term_new_<n>}.
|
||||
*/
|
||||
@Test
|
||||
void relaunchReturnsTheStartedAgent() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
dev.ltms.fleet.herdr.Agent started =
|
||||
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
|
||||
|
||||
assertNotNull(started, "a launchable, configured lead must start");
|
||||
Object startedPaneIdParam = ((Map<?, ?>) herdr.lastCall("agent.start").params()).get("pane_id");
|
||||
assertEquals(startedPaneIdParam, started.paneId(),
|
||||
"paneId() must be the pane the agent.start call actually targeted");
|
||||
assertTrue(started.terminalId() != null && started.terminalId().startsWith("term_new_"),
|
||||
"terminalId() must be herdr's own generated id: " + started.terminalId());
|
||||
}
|
||||
|
||||
/** The new tab is labelled with the lead's configured {@code tab:}, and AFTER the start. */
|
||||
@Test
|
||||
void relaunchLabelsTheNewTabAfterStarting() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
dev.ltms.fleet.herdr.Agent started =
|
||||
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
|
||||
|
||||
assertNotNull(started);
|
||||
assertEquals("lead: opus", ((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("label"));
|
||||
|
||||
int startIndex = indexOfLastCall(herdr, "agent.start");
|
||||
int renameIndex = indexOfLastCall(herdr, "tab.rename");
|
||||
assertTrue(renameIndex > startIndex,
|
||||
"the tab must be renamed AFTER the start succeeds, not before: start=" + startIndex
|
||||
+ " rename=" + renameIndex);
|
||||
}
|
||||
|
||||
private static int indexOfLastCall(FakeHerdr herdr, String method) {
|
||||
int idx = -1;
|
||||
List<FakeHerdr.Call> calls = herdr.calls;
|
||||
for (int i = 0; i < calls.size(); i++) {
|
||||
if (calls.get(i).method().equals(method)) {
|
||||
idx = i;
|
||||
}
|
||||
}
|
||||
return idx;
|
||||
}
|
||||
|
||||
@Test
|
||||
void relaunchOfAnUnknownLeadNameReturnsNullAndStartsNothing() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
dev.ltms.fleet.herdr.Agent started =
|
||||
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("not-declared");
|
||||
|
||||
assertNull(started);
|
||||
assertFalse(herdr.called("agent.start"));
|
||||
assertFalse(herdr.called("workspace.create"));
|
||||
assertFalse(herdr.called("tab.create"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void relaunchOfARecogniseOnlyLeadReturnsNullAndStartsNothing() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
dev.ltms.fleet.herdr.Agent started =
|
||||
launcher(herdr, configWith(lead(null, "lead: dead", 1))).relaunch("opus");
|
||||
|
||||
assertNull(started);
|
||||
assertFalse(herdr.called("agent.start"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void relaunchWithAnUnconfiguredProfileReturnsNullAndStartsNothing() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
dev.ltms.fleet.herdr.Agent started =
|
||||
launcher(herdr, configWith(lead("nope", "lead: opus", 1))).relaunch("opus");
|
||||
|
||||
assertNull(started);
|
||||
assertFalse(herdr.called("agent.start"));
|
||||
}
|
||||
|
||||
/**
|
||||
* The outer retry {@link LeadLauncher#relaunch(String)} owns, separate from {@code
|
||||
* ResilientAgentLaunch}'s internal {@code agent_name_taken} retry: a failed attempt must not
|
||||
* be the end of the whole relaunch. Each of the first two attempts exhausts {@code
|
||||
* ResilientAgentLaunch.NAME_RETRIES} name attempts (every one of them rejected), so each
|
||||
* attempt's own tab is created and then closed; the third attempt's first name is free.
|
||||
*/
|
||||
@Test
|
||||
void relaunchRetriesTheWholeAttemptAndSucceedsOnTheThird() {
|
||||
FakeHerdr herdr = new FakeHerdr()
|
||||
.agentNameTakenTimes(2 * ResilientAgentLaunch.NAME_RETRIES);
|
||||
|
||||
dev.ltms.fleet.herdr.Agent started =
|
||||
fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
|
||||
|
||||
assertNotNull(started, "the third attempt's first name is free — it must succeed");
|
||||
assertEquals(3, herdr.calls.stream().filter(c -> c.method().equals("tab.create")).count(),
|
||||
"one tab per attempt: three attempts");
|
||||
assertEquals(2, herdr.calls.stream().filter(c -> c.method().equals("tab.close")).count(),
|
||||
"the two failed attempts' tabs must be closed");
|
||||
}
|
||||
|
||||
/**
|
||||
* Every attempt fails outright (a herdr error {@code ResilientAgentLaunch} does not retry at
|
||||
* all) — {@link LeadLauncher#relaunch(String)} must give up after exactly {@code
|
||||
* RELAUNCH_ATTEMPTS} and must not leak any of the tabs it created along the way.
|
||||
*/
|
||||
@Test
|
||||
void relaunchGivesUpAfterExactlyRelaunchAttemptsAndLeaksNoTab() {
|
||||
FakeHerdr herdr = new FakeHerdr().agentStartFailsWith("some_other_error");
|
||||
|
||||
dev.ltms.fleet.herdr.Agent started =
|
||||
fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
|
||||
|
||||
assertNull(started, "every attempt failed — relaunch must give up, not hang or guess");
|
||||
assertEquals(LeadLauncher.RELAUNCH_ATTEMPTS,
|
||||
herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
|
||||
"exactly RELAUNCH_ATTEMPTS attempts, no more, no fewer");
|
||||
long tabsCreated = herdr.calls.stream().filter(c -> c.method().equals("tab.create")).count();
|
||||
long tabsClosed = herdr.calls.stream().filter(c -> c.method().equals("tab.close")).count();
|
||||
assertEquals(LeadLauncher.RELAUNCH_ATTEMPTS, tabsCreated);
|
||||
assertEquals(tabsCreated, tabsClosed, "every tab this method created must be closed — no leaks");
|
||||
}
|
||||
}
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -63,6 +63,17 @@ class FleetMcpAuthzTest {
|
||||
|
||||
/** A fully wired FleetMcp on fakes — constructing it is itself part of what is under test. */
|
||||
private FleetMcp mcp(boolean enforce) {
|
||||
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
|
||||
return mcp(enforce, CallerResolver.withLeadsAndMembers(identity, false, null,
|
||||
Map::of, new MemberRegistry(null)));
|
||||
}
|
||||
|
||||
/**
|
||||
* As {@link #mcp(boolean)}, with an explicit {@link CallerResolver} — so a test can wire known
|
||||
* leads/collaborators and drive {@code denyFor}'s real {@code knownLeadOrCollaborator()}
|
||||
* classifier instead of the default empty one.
|
||||
*/
|
||||
private FleetMcp mcp(boolean enforce, CallerResolver callers) {
|
||||
FleetConfig.Profile cfg = new FleetConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
|
||||
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
@@ -83,9 +94,7 @@ class FleetMcpAuthzTest {
|
||||
// to be omitted to reach "legacy" is now always real, and AuthorizationMode is the
|
||||
// separate, explicit choice that governs enforcement.
|
||||
mcp = new FleetMcp(messages, workers, sessions, identity, sessions.asPresence(),
|
||||
new PrimaryRegistry(null),
|
||||
CallerResolver.withLeadsAndMembers(identity, false, null,
|
||||
Map::of, new MemberRegistry(null)),
|
||||
new PrimaryRegistry(null), callers,
|
||||
enforce ? FleetMcp.AuthorizationMode.ENFORCED : FleetMcp.AuthorizationMode.UNENFORCED,
|
||||
metrics, FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), null, FleetMcp.OutageSource.none(),
|
||||
@@ -97,6 +106,7 @@ class FleetMcpAuthzTest {
|
||||
private static final Principal WORKER_A = Principal.worker("term_a", 200);
|
||||
private static final Principal ANON = Principal.anonymous();
|
||||
private static final Principal ARCH_DESIGN = Principal.architect("lead-designer", "term_design", 400);
|
||||
private static final Principal COLLABORATOR = Principal.collaborator("ops", "term_collab", 600);
|
||||
|
||||
// --- the table, enforced on THIS path too ---------------------------------------------------
|
||||
|
||||
@@ -156,6 +166,45 @@ class FleetMcpAuthzTest {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #669 Unit A: {@code SEND} is split into three actions ({@link Authz.Action#SEND},
|
||||
* {@link Authz.Action#COORD_SEND}, {@link Authz.Action#ANSWER}), each carrying the same grant
|
||||
* the one undivided action gave. An architect holds all three, exactly as it held the one.
|
||||
*/
|
||||
@Test
|
||||
void anArchitectMayUseAllThreeSendShapesOverMcp() {
|
||||
FleetMcp m = mcp(true);
|
||||
for (Authz.Action a : new Authz.Action[]{Authz.Action.SEND, Authz.Action.COORD_SEND,
|
||||
Authz.Action.ANSWER}) {
|
||||
assertNull(m.denyFor(ARCH_DESIGN, a, "term_a"),
|
||||
a + " carries the same grant the undivided SEND action gave an architect");
|
||||
}
|
||||
}
|
||||
|
||||
/** The other half of the same split: a worker is excluded from all three, as it was from one. */
|
||||
@Test
|
||||
void aWorkerMayNotUseAnySendShapeOverMcp() {
|
||||
FleetMcp m = mcp(true);
|
||||
for (Authz.Action a : new Authz.Action[]{Authz.Action.SEND, Authz.Action.COORD_SEND,
|
||||
Authz.Action.ANSWER}) {
|
||||
McpSchema.CallToolResult denied = m.denyFor(WORKER_A, a, "term_a");
|
||||
assertNotNull(denied, a + " must stay refused to a worker");
|
||||
assertTrue(denied.isError(), "a refusal is returned as an MCP tool error");
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #669 Unit A / #678: {@code TASK_READ} (ticket polling, session status) is split out of
|
||||
* the roster-only {@code READ}, carrying forward the grant the undivided action gave. A worker
|
||||
* still has both — it never gained or lost anything by the split.
|
||||
*/
|
||||
@Test
|
||||
void aWorkerKeepsBothReadActionsAfterTheSplit() {
|
||||
FleetMcp m = mcp(true);
|
||||
assertNull(m.denyFor(WORKER_A, Authz.Action.READ, null));
|
||||
assertNull(m.denyFor(WORKER_A, Authz.Action.TASK_READ, null));
|
||||
}
|
||||
|
||||
@Test
|
||||
void anArchitectMayReplyAndAskOnlyAsItsOwnPaneOverMcp() {
|
||||
FleetMcp m = mcp(true);
|
||||
@@ -187,6 +236,39 @@ class FleetMcpAuthzTest {
|
||||
"the caller IS authenticated — it is just not the right role");
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code denyFor} passes the real production classifier, not a test-supplied one — no
|
||||
* terminal is recognised as a configured lead or collaborator, so a collaborator's SEND is
|
||||
* refused over MCP.
|
||||
*/
|
||||
@Test
|
||||
void aCollaboratorMayNotSendOverMcpWithTheRealProductionClassifier() {
|
||||
FleetMcp m = mcp(true);
|
||||
McpSchema.CallToolResult denied = m.denyFor(COLLABORATOR, Authz.Action.SEND, "term_lead");
|
||||
assertNotNull(denied, "no terminal is recognised as a lead or collaborator yet");
|
||||
assertTrue(denied.isError());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #669 Unit D: wires a real {@link CallerResolver} with a known lead and a known
|
||||
* collaborator tab, and leaves a spawned member's own terminal recognised by neither map — so a
|
||||
* collaborator's SEND reaches both named peers and is refused for the spawned member's terminal,
|
||||
* over MCP's {@code denyFor}.
|
||||
*/
|
||||
@Test
|
||||
void aCollaboratorMaySendToAKnownLeadOrCollaboratorButNotToASpawnedMembersTerminal() {
|
||||
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
|
||||
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
|
||||
() -> Map.of("term_lead_known", "lead-x"), new MemberRegistry(null),
|
||||
t -> null, () -> Map.of("term_collab_known", "ops2"));
|
||||
FleetMcp m = mcp(true, callers);
|
||||
|
||||
assertNull(m.denyFor(COLLABORATOR, Authz.Action.SEND, "term_lead_known"));
|
||||
assertNull(m.denyFor(COLLABORATOR, Authz.Action.SEND, "term_collab_known"));
|
||||
assertNotNull(m.denyFor(COLLABORATOR, Authz.Action.SEND, "term_a"),
|
||||
"a spawned member's own terminal must stay unreachable, even once the classifier is real");
|
||||
}
|
||||
|
||||
@Test
|
||||
void theLegacyConstructorLeavesTheGateOpen() {
|
||||
// The 22 pre-existing FleetMcpTest cases rely on no authorization being enforced.
|
||||
@@ -273,10 +355,10 @@ class FleetMcpAuthzTest {
|
||||
+ "the anchors have drifted, this test is not testing what it claims to");
|
||||
|
||||
Pattern trailingArg = Pattern.compile(
|
||||
"listFleet\\([^;]*?,\\s*(coordinatorVisibleTo\\(principal\\(exchange\\)\\)|true|false)\\s*\\)\\s*;",
|
||||
"listFleet\\([^;]*?,\\s*(coordinatorVisibleTo\\(principal\\(exchange\\)\\)|true|false)\\s*[,)]",
|
||||
Pattern.DOTALL);
|
||||
Matcher m = trailingArg.matcher(handlerBlock);
|
||||
assertTrue(m.find(), "could not locate listFleet(...)'s trailing boolean argument in the "
|
||||
assertTrue(m.find(), "could not locate listFleet(...)'s coordinator boolean argument in the "
|
||||
+ "listHandler block -- the call shape changed, update this test's anchor: " + handlerBlock);
|
||||
String trailing = m.group(1);
|
||||
assertEquals("coordinatorVisibleTo(principal(exchange))", trailing,
|
||||
@@ -284,6 +366,197 @@ class FleetMcpAuthzTest {
|
||||
+ "calling, not pass a literal boolean -- found: " + trailing);
|
||||
}
|
||||
|
||||
// --- fleetd #703: who may see fleet_list's collaborators array -------------------------------
|
||||
|
||||
/**
|
||||
* fleetd #703: {@link FleetMcp#collaboratorsVisibleTo} is the whole policy decision for
|
||||
* {@code fleet_list}'s {@code collaborators} array. Visible to exactly the roles that may
|
||||
* {@code SEND} to a named peer -- the primary, an architect, and a collaborator itself -- never
|
||||
* a worker, which holds {@code READ} but can never {@code SEND} to a collaborator, and never an
|
||||
* anonymous caller.
|
||||
*/
|
||||
@Test
|
||||
void onlyPrimaryArchitectAndCollaboratorMaySeeTheCollaboratorsArray() {
|
||||
assertTrue(FleetMcp.collaboratorsVisibleTo(PRIMARY), "the primary must see the collaborators array");
|
||||
assertTrue(FleetMcp.collaboratorsVisibleTo(ARCH_DESIGN), "an architect must see the collaborators array");
|
||||
assertTrue(FleetMcp.collaboratorsVisibleTo(COLLABORATOR), "a collaborator must see its own peer roster");
|
||||
assertFalse(FleetMcp.collaboratorsVisibleTo(WORKER_A),
|
||||
"a worker holds READ but can never SEND to a collaborator, so it must not see the array");
|
||||
assertFalse(FleetMcp.collaboratorsVisibleTo(ANON), "authenticated as nothing must not see it either");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #703, same reasoning as {@link #theFleetListHandlerActuallyConsultsCoordinatorVisibleTo}:
|
||||
* the predicate above can be perfectly correct while the one production call site never asks it.
|
||||
* This reads {@code FleetMcp.java}'s own source and asserts the {@code fleet_list} handler's
|
||||
* {@code listFleet(...)} call both threads {@code callers.collaborators()} into the payload and
|
||||
* asks {@code collaboratorsVisibleTo(principal(exchange))} for the visibility flag, rather than a
|
||||
* literal boolean or an empty map.
|
||||
*/
|
||||
@Test
|
||||
void theFleetListHandlerActuallyConsultsCollaboratorsVisibleTo() throws Exception {
|
||||
String source = Files.readString(MCP_SOURCE);
|
||||
|
||||
int start = source.indexOf("listHandler =");
|
||||
assertTrue(start >= 0, "could not find the fleet_list handler (listHandler) in " + MCP_SOURCE
|
||||
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
|
||||
int end = source.indexOf("stopHandler =", start);
|
||||
assertTrue(end > start, "could not find the handler declared after listHandler to bound the scrape");
|
||||
String handlerBlock = source.substring(start, end);
|
||||
|
||||
// CONTROL: the block we scraped really does contain a call to listFleet(...) -- if this
|
||||
// fails, the anchors above moved and the assertions below would otherwise pass on nothing.
|
||||
assertTrue(handlerBlock.contains("listFleet("),
|
||||
"control failed: the scraped listHandler block contains no listFleet( call at all -- "
|
||||
+ "the anchors have drifted, this test is not testing what it claims to");
|
||||
|
||||
assertTrue(handlerBlock.contains("callers.collaborators()"),
|
||||
"the fleet_list handler must thread callers.collaborators() into listFleet(...), not an "
|
||||
+ "empty or literal map -- block: " + handlerBlock);
|
||||
assertTrue(handlerBlock.contains("collaboratorsVisibleTo(principal(exchange))"),
|
||||
"the fleet_list handler must ask collaboratorsVisibleTo(principal(exchange)) who is "
|
||||
+ "calling, not pass a literal boolean -- block: " + handlerBlock);
|
||||
}
|
||||
|
||||
// --- who may see fleet_list's leads and members arrays ---------------------------------------
|
||||
|
||||
/**
|
||||
* {@link FleetMcp#leadsVisibleTo} is the whole policy decision for {@code fleet_list}'s
|
||||
* {@code leads} array: visible to exactly the roles that may {@code SEND} to a lead -- the
|
||||
* primary, an architect, and a collaborator -- never a worker, which holds {@code READ} but
|
||||
* can never {@code SEND} at all, and never an anonymous caller.
|
||||
*/
|
||||
@Test
|
||||
void primaryArchitectAndCollaboratorMaySeeTheLeadsArray() {
|
||||
assertTrue(FleetMcp.leadsVisibleTo(PRIMARY), "the primary must see the leads array");
|
||||
assertTrue(FleetMcp.leadsVisibleTo(ARCH_DESIGN), "an architect must see the leads array");
|
||||
assertTrue(FleetMcp.leadsVisibleTo(COLLABORATOR),
|
||||
"a collaborator may SEND to a lead, so it must see the leads array to learn where");
|
||||
assertFalse(FleetMcp.leadsVisibleTo(WORKER_A),
|
||||
"a worker holds READ but can never SEND, so it must not see the leads array");
|
||||
assertFalse(FleetMcp.leadsVisibleTo(ANON), "authenticated as nothing must not see it either");
|
||||
}
|
||||
|
||||
/**
|
||||
* As {@link #primaryArchitectAndCollaboratorMaySeeTheLeadsArray}, for the {@code members}
|
||||
* array -- but a collaborator may {@code SEND} only to a lead or another collaborator, never
|
||||
* to a spawned member, so it must not see this one.
|
||||
*/
|
||||
@Test
|
||||
void onlyPrimaryAndArchitectMaySeeTheMembersArray() {
|
||||
assertTrue(FleetMcp.membersVisibleTo(PRIMARY), "the primary must see the members array");
|
||||
assertTrue(FleetMcp.membersVisibleTo(ARCH_DESIGN), "an architect must see the members array");
|
||||
assertFalse(FleetMcp.membersVisibleTo(WORKER_A),
|
||||
"a worker holds READ but can never SEND, so it must not see the members array");
|
||||
assertFalse(FleetMcp.membersVisibleTo(COLLABORATOR),
|
||||
"a collaborator may SEND to a lead, never to a spawned member, so it must not see the members array");
|
||||
assertFalse(FleetMcp.membersVisibleTo(ANON), "authenticated as nothing must not see it either");
|
||||
}
|
||||
|
||||
/**
|
||||
* Same reasoning as {@link #theFleetListHandlerActuallyConsultsCoordinatorVisibleTo}: the
|
||||
* predicate above can be perfectly correct while the one production call site never asks it.
|
||||
* This reads {@code FleetMcp.java}'s own source and asserts the {@code fleet_list} handler's
|
||||
* {@code listFleet(...)} call asks {@code leadsVisibleTo(principal(exchange))} for the leads
|
||||
* visibility flag, rather than a literal boolean.
|
||||
*/
|
||||
@Test
|
||||
void theFleetListHandlerActuallyConsultsLeadsVisibleTo() throws Exception {
|
||||
String source = Files.readString(MCP_SOURCE);
|
||||
|
||||
int start = source.indexOf("listHandler =");
|
||||
assertTrue(start >= 0, "could not find the fleet_list handler (listHandler) in " + MCP_SOURCE
|
||||
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
|
||||
int end = source.indexOf("stopHandler =", start);
|
||||
assertTrue(end > start, "could not find the handler declared after listHandler to bound the scrape");
|
||||
String handlerBlock = source.substring(start, end);
|
||||
|
||||
// CONTROL: the block we scraped really does contain a call to listFleet(...) -- if this
|
||||
// fails, the anchors above moved and the assertion below would otherwise pass on nothing.
|
||||
assertTrue(handlerBlock.contains("listFleet("),
|
||||
"control failed: the scraped listHandler block contains no listFleet( call at all -- "
|
||||
+ "the anchors have drifted, this test is not testing what it claims to");
|
||||
|
||||
assertTrue(handlerBlock.contains("leadsVisibleTo(principal(exchange))"),
|
||||
"the fleet_list handler must ask leadsVisibleTo(principal(exchange)) who is calling, "
|
||||
+ "not pass a literal boolean -- block: " + handlerBlock);
|
||||
}
|
||||
|
||||
/** As {@link #theFleetListHandlerActuallyConsultsLeadsVisibleTo}, for {@code membersVisibleTo}. */
|
||||
@Test
|
||||
void theFleetListHandlerActuallyConsultsMembersVisibleTo() throws Exception {
|
||||
String source = Files.readString(MCP_SOURCE);
|
||||
|
||||
int start = source.indexOf("listHandler =");
|
||||
assertTrue(start >= 0, "could not find the fleet_list handler (listHandler) in " + MCP_SOURCE
|
||||
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
|
||||
int end = source.indexOf("stopHandler =", start);
|
||||
assertTrue(end > start, "could not find the handler declared after listHandler to bound the scrape");
|
||||
String handlerBlock = source.substring(start, end);
|
||||
|
||||
assertTrue(handlerBlock.contains("listFleet("),
|
||||
"control failed: the scraped listHandler block contains no listFleet( call at all -- "
|
||||
+ "the anchors have drifted, this test is not testing what it claims to");
|
||||
|
||||
assertTrue(handlerBlock.contains("membersVisibleTo(principal(exchange))"),
|
||||
"the fleet_list handler must ask membersVisibleTo(principal(exchange)) who is calling, "
|
||||
+ "not pass a literal boolean -- block: " + handlerBlock);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code fleet_poll{ticket}} must thread the calling connection's own terminal into
|
||||
* {@link MessageService#poll(String, String)}, so a worker cannot read a ticket a different
|
||||
* session created.
|
||||
*/
|
||||
@Test
|
||||
void theFleetPollHandlerActuallyThreadsCallerTerminalIntoPoll() throws Exception {
|
||||
String source = Files.readString(MCP_SOURCE);
|
||||
|
||||
int start = source.indexOf("pollHandler =");
|
||||
assertTrue(start >= 0, "could not find the fleet_poll handler (pollHandler) in " + MCP_SOURCE
|
||||
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
|
||||
int end = source.indexOf("ackHandler =", start);
|
||||
assertTrue(end > start, "could not find the handler declared after pollHandler to bound the scrape");
|
||||
String handlerBlock = source.substring(start, end);
|
||||
|
||||
// CONTROL: the block we scraped really does call poll(...) -- if this fails, the anchors
|
||||
// above moved and the assertions below would otherwise pass on nothing.
|
||||
assertTrue(handlerBlock.contains("poll(messages,"),
|
||||
"control failed: the scraped pollHandler block contains no poll(messages, ...) call "
|
||||
+ "at all -- the anchors have drifted, this test is not testing what it claims to");
|
||||
|
||||
assertTrue(handlerBlock.contains("callerTerminal(exchange)"),
|
||||
"the fleet_poll handler must thread callerTerminal(exchange) into poll(...), not omit "
|
||||
+ "it or pass a literal null -- block: " + handlerBlock);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code fleet_status} must thread the calling connection's own terminal into
|
||||
* {@link FleetMcp#status(MessageService, String, String)}, so a caller that did not create a
|
||||
* worker's open delegation cannot read its pending question through the status handler either.
|
||||
*/
|
||||
@Test
|
||||
void theFleetStatusHandlerActuallyThreadsCallerTerminalIntoStatus() throws Exception {
|
||||
String source = Files.readString(MCP_SOURCE);
|
||||
|
||||
int start = source.indexOf("statusHandler =");
|
||||
assertTrue(start >= 0, "could not find the fleet_status handler (statusHandler) in " + MCP_SOURCE
|
||||
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
|
||||
int end = source.indexOf("pollHandler =", start);
|
||||
assertTrue(end > start, "could not find the handler declared after statusHandler to bound the scrape");
|
||||
String handlerBlock = source.substring(start, end);
|
||||
|
||||
// CONTROL: the block we scraped really does call status(...) -- if this fails, the anchors
|
||||
// above moved and the assertions below would otherwise pass on nothing.
|
||||
assertTrue(handlerBlock.contains("status(messages,"),
|
||||
"control failed: the scraped statusHandler block contains no status(messages, ...) "
|
||||
+ "call at all -- the anchors have drifted, this test is not testing what it claims to");
|
||||
|
||||
assertTrue(handlerBlock.contains("callerTerminal(exchange)"),
|
||||
"the fleet_status handler must thread callerTerminal(exchange) into status(...), not "
|
||||
+ "omit it or pass a literal null -- block: " + handlerBlock);
|
||||
}
|
||||
|
||||
// --- which action each tool hands the gate (fleetd #272) ------------------------------------
|
||||
|
||||
/**
|
||||
@@ -297,35 +570,53 @@ class FleetMcpAuthzTest {
|
||||
* the whole time the defect was live -- the table was right, the action fed to it was wrong.
|
||||
*/
|
||||
@Test
|
||||
void pollingByTargetIsADrainAndPollingByTicketIsARead() {
|
||||
void pollingByTargetIsADrainAndPollingByTicketIsATaskRead() {
|
||||
assertEquals(Authz.Action.DRAIN, FleetMcp.pollAction("term_b", null),
|
||||
"poll by target removes the replies — that is a drain, not an observation");
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null, null),
|
||||
"poll by ticket changes nothing");
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(" ", null),
|
||||
assertEquals(Authz.Action.TASK_READ, FleetMcp.pollAction(null, null),
|
||||
"poll by ticket changes nothing, but is not the roster-only READ action");
|
||||
assertEquals(Authz.Action.TASK_READ, FleetMcp.pollAction(" ", null),
|
||||
"a blank target is an absent target");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #421: a coordId branch is a THIRD operation behind fleet_poll's one name, and it must
|
||||
* map to {@link Authz.Action#COORD_READ} — never {@link Authz.Action#READ}, even though this
|
||||
* branch also consumes nothing. READ's grant is open to every authenticated role on the premise
|
||||
* that the roster carries no secrets; a lead-to-lead body is not the roster, so folding this
|
||||
* branch into READ would let any worker read every peer lead's mail in full. coordId also takes
|
||||
* priority over target when both happen to be present — it addresses a different inbox entirely.
|
||||
* map to {@link Authz.Action#COORD_READ} — never {@link Authz.Action#READ} or {@link
|
||||
* Authz.Action#TASK_READ}, even though this branch also consumes nothing. A lead-to-lead body
|
||||
* is not the roster and not a ticket/status read, so folding this branch into either would let
|
||||
* any worker or architect read every peer lead's mail in full. coordId also takes priority over
|
||||
* target when both happen to be present — it addresses a different inbox entirely.
|
||||
*/
|
||||
@Test
|
||||
void pollingByCoordIdIsACoordReadNeverAPlainRead() {
|
||||
void pollingByCoordIdIsACoordReadNeverAPlainOrTaskRead() {
|
||||
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction(null, "mac-opus"),
|
||||
"reading held peer mail must not be mapped to the everyone-readable READ action");
|
||||
"reading held peer mail must not be mapped to a widely-readable action");
|
||||
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction(" ", "mac-opus"),
|
||||
"a blank target must not fall through to READ/DRAIN when coordId is present");
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null, " "),
|
||||
assertEquals(Authz.Action.TASK_READ, FleetMcp.pollAction(null, " "),
|
||||
"a blank coordId is an absent coordId, same as target/ticket");
|
||||
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction("term_b", "mac-opus"),
|
||||
"coordId takes priority over target — this is a different inbox, not a drain");
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code fleet_send} is three call shapes behind one tool name, exactly as {@code fleet_poll}
|
||||
* is (fleetd #669 Unit A). {@link FleetMcp#sendAction} picks the action from the arguments, not
|
||||
* the handler, for the same reason {@link FleetMcp#pollAction} does: a test can assert the
|
||||
* mapping the handler actually uses.
|
||||
*/
|
||||
@Test
|
||||
void sendMapsToThreeDifferentActionsByItsArguments() {
|
||||
assertEquals(Authz.Action.SEND, FleetMcp.sendAction(null, null),
|
||||
"a plain delivery, with neither coordId nor turnId, is a local SEND");
|
||||
assertEquals(Authz.Action.COORD_SEND, FleetMcp.sendAction("mac-opus", null),
|
||||
"coordId addresses a peer lead over the coordination broker");
|
||||
assertEquals(Authz.Action.ANSWER, FleetMcp.sendAction(null, "turn-1"),
|
||||
"turnId resolves a worker's blocked question");
|
||||
assertEquals(Authz.Action.COORD_SEND, FleetMcp.sendAction("mac-opus", "turn-1"),
|
||||
"coordId takes priority over turnId, mirroring sendToLead's own mutual-exclusion check");
|
||||
}
|
||||
|
||||
@Test
|
||||
void everyRegisteredToolHasItsHandlerActionPinned() {
|
||||
// fleetd #469: this used to scrape FleetMcp.java's tool("…") calls for the registered set —
|
||||
@@ -344,16 +635,22 @@ class FleetMcpAuthzTest {
|
||||
() -> tool + " is registered but has no pinned authorization action"));
|
||||
|
||||
assertEquals(Authz.Action.SEND, FleetMcp.toolAction("fleet_send", Map.of()));
|
||||
assertEquals(Authz.Action.SEND,
|
||||
FleetMcp.toolAction("fleet_send", Map.of("sessionId", "term_a", "content", "hi")));
|
||||
assertEquals(Authz.Action.COORD_SEND,
|
||||
FleetMcp.toolAction("fleet_send", Map.of("coordId", "mac-opus", "content", "hi")));
|
||||
assertEquals(Authz.Action.ANSWER,
|
||||
FleetMcp.toolAction("fleet_send", Map.of("turnId", "turn-1", "content", "hi")));
|
||||
assertEquals(Authz.Action.REPLY, FleetMcp.toolAction("fleet_reply", Map.of()));
|
||||
assertEquals(Authz.Action.ASK, FleetMcp.toolAction("fleet_ask", Map.of()));
|
||||
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_status", Map.of()));
|
||||
assertEquals(Authz.Action.TASK_READ, FleetMcp.toolAction("fleet_status", Map.of()));
|
||||
assertEquals(Authz.Action.DRAIN, FleetMcp.toolAction("fleet_ack", Map.of()));
|
||||
assertEquals(Authz.Action.SPAWN, FleetMcp.toolAction("fleet_spawn", Map.of()));
|
||||
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_list", Map.of()));
|
||||
assertEquals(Authz.Action.STOP, FleetMcp.toolAction("fleet_stop", Map.of()));
|
||||
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_profiles", Map.of()));
|
||||
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_whoami", Map.of()));
|
||||
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_poll", Map.of("ticket", "task")));
|
||||
assertEquals(Authz.Action.TASK_READ, FleetMcp.toolAction("fleet_poll", Map.of("ticket", "task")));
|
||||
assertEquals(Authz.Action.DRAIN, FleetMcp.toolAction("fleet_poll", Map.of("target", "term_b")));
|
||||
assertEquals(Authz.Action.COORD_READ,
|
||||
FleetMcp.toolAction("fleet_poll", Map.of("coordId", "mac-opus")));
|
||||
|
||||
@@ -11,6 +11,7 @@ import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.PaneLocator;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.lead.LeadLauncher;
|
||||
import dev.ltms.fleet.lead.LeadRollover;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.msg.InMemoryReplyInbox;
|
||||
@@ -24,6 +25,8 @@ import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.io.UncheckedIOException;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
@@ -36,14 +39,14 @@ import static org.junit.jupiter.api.Assertions.*;
|
||||
* fleetd #480 Unit C — the {@code fleet_handover} MCP tool, the surface that finally calls
|
||||
* {@link LeadRollover#open}/{@link LeadRollover#confirm}/{@link LeadRollover#cancel}.
|
||||
*
|
||||
* <p>Uses {@link LeadRollover}'s PUBLIC constructor (real wall clock, real 250ms settle poll, a
|
||||
* real virtual-thread continuation runner) rather than its package-private test constructor —
|
||||
* this test lives in {@code dev.ltms.fleet.mcp}, not {@code dev.ltms.fleet.lead}, and does not
|
||||
* need to control the post-{@code confirm()} continuation's timing: it only asserts the
|
||||
* SYNCHRONOUS return value of {@code open}/{@code confirm}/{@code cancel}, which is exactly what
|
||||
* {@code FleetMcp.handover} forwards to the client. {@code turnSettleSeconds}/{@code
|
||||
* clearSettleSeconds} are kept at 1s so a confirmed request's background continuation (which this
|
||||
* class does not wait on or assert against) gives up quickly rather than polling for 20s on a
|
||||
* <p>Uses {@link LeadRollover}'s PUBLIC constructor (real wall clock, real 250ms poll, a real
|
||||
* virtual-thread continuation runner) rather than its package-private test constructor — this
|
||||
* test lives in {@code dev.ltms.fleet.mcp}, not {@code dev.ltms.fleet.lead}, and does not need to
|
||||
* control the post-{@code confirm()} continuation's timing: it only asserts the SYNCHRONOUS return
|
||||
* value of {@code open}/{@code confirm}/{@code cancel}, which is exactly what {@code
|
||||
* FleetMcp.handover} forwards to the client. {@code turnSettleSeconds}/{@code
|
||||
* relaunchReadySeconds} are kept at 1s so a confirmed request's background continuation (which
|
||||
* this class does not wait on or assert against) gives up quickly rather than polling for 20s on a
|
||||
* daemon virtual thread.
|
||||
*/
|
||||
class FleetMcpHandoverTest {
|
||||
@@ -67,10 +70,26 @@ class FleetMcpHandoverTest {
|
||||
return new FleetConfig.LeadRollover(handoverPath, false, 3600, 1, 1, "read the handover file");
|
||||
}
|
||||
|
||||
private static FleetConfig minimalFleetConfig() {
|
||||
try {
|
||||
Path yaml = Files.createTempFile("fleet-mcp-handover-test", ".yaml");
|
||||
Files.writeString(yaml, "bind:\n port: 8080\n");
|
||||
return FleetConfig.load(yaml);
|
||||
} catch (IOException e) {
|
||||
throw new UncheckedIOException(e);
|
||||
}
|
||||
}
|
||||
|
||||
private LeadRollover newRollover(String handoverPath) {
|
||||
// Every handoverPath this test class uses comes from tmp.resolve(...), which is already
|
||||
// absolute, so the workspace lookup is never actually consulted — a no-op lookup is enough.
|
||||
return new LeadRollover(agents, () -> cfg(handoverPath), _ -> null);
|
||||
// None of this class's tests reach the recognition-wait or the relaunch call, so the
|
||||
// launcher's own functional correctness is irrelevant here — any constructed instance,
|
||||
// backed by the same fake herdr, is enough.
|
||||
WorkspaceControl spaces = new WorkspaceControl(herdr);
|
||||
LeadLauncher launcher = new LeadLauncher(agents, spaces, minimalFleetConfig());
|
||||
return new LeadRollover(agents, spaces, launcher, () -> cfg(handoverPath),
|
||||
_ -> null, _ -> null, Map::of);
|
||||
}
|
||||
|
||||
/** A fully wired FleetMcp on fakes (mirrors FleetMcpAuthzTest's helper), plus a leadRollover. */
|
||||
|
||||
@@ -121,6 +121,7 @@ class FleetMcpLeadContextGaugeWiringTest {
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
|
||||
FleetMcp.LeadSeatSource.none(), contextGauge, leadConfigDirs,
|
||||
Map.of(LEAD_TERMINAL, LEAD_NAME), LEAD_TERMINAL, FleetMcp.CoordinationSource.none(), false);
|
||||
Map.of(LEAD_TERMINAL, LEAD_NAME), LEAD_TERMINAL, Map.of(), false,
|
||||
FleetMcp.CoordinationSource.none(), false, true, true);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -10,6 +10,7 @@ import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.AgentStatus;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.inject.LoopWatchdog;
|
||||
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||
import dev.ltms.fleet.herdr.PaneLocator;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
@@ -76,7 +77,7 @@ class FleetMcpTest {
|
||||
|
||||
private void assertSendRoundTrips(String target, Set<String> profiles) throws Exception {
|
||||
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
|
||||
() -> FleetMcp.send(messages, target, "hi", 4000L, null, profiles));
|
||||
() -> FleetMcp.send(messages, target, "hi", 4000L, null, profiles, null));
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting(target) && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
@@ -102,7 +103,7 @@ class FleetMcpTest {
|
||||
void sendThenReplyRoundTrips() throws Exception {
|
||||
// fleet_send blocks; fleet_reply resolves it with the worker's structured answer.
|
||||
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
|
||||
() -> FleetMcp.send(messages, "term_a", "review this", 4000L, null, Set.of()));
|
||||
() -> FleetMcp.send(messages, "term_a", "review this", 4000L, null, Set.of(), null));
|
||||
|
||||
// Wait until the send has opened its waiter so the reply resolves it (CB-307: reply now
|
||||
// queues in the inbox if no waiter is open, which would break the round-trip).
|
||||
@@ -180,7 +181,7 @@ class FleetMcpTest {
|
||||
String turnId = afterTurnId.substring(0, afterTurnId.indexOf('"'));
|
||||
|
||||
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
|
||||
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L));
|
||||
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, null));
|
||||
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
|
||||
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
@@ -278,7 +279,7 @@ class FleetMcpTest {
|
||||
assertEquals("late reply", messages.drainReplies("term_a").getFirst().content());
|
||||
|
||||
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
|
||||
() -> FleetMcp.answer(messages, firstTurnId, "config.yaml", 5000L));
|
||||
() -> FleetMcp.answer(messages, firstTurnId, "config.yaml", 5000L, null));
|
||||
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
|
||||
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
@@ -287,6 +288,89 @@ class FleetMcpTest {
|
||||
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
|
||||
}
|
||||
|
||||
// --- fleetd #715: fleet_send{turnId} is gated on the caller that owns the turn -------------
|
||||
|
||||
/**
|
||||
* A blocking {@code fleet_send} from one caller opens the turn; a {@code fleet_send{turnId}}
|
||||
* from a different caller is refused as an error, and the real owner's answer still succeeds.
|
||||
*/
|
||||
@Test
|
||||
void aDifferentCallersMcpAnswerIsRefusedForABlockingSendButTheRealOwnerSucceeds() throws Exception {
|
||||
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
|
||||
() -> FleetMcp.send(messages, T, "do X", 5000L, null, Set.of(), "term_owner"));
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting(T));
|
||||
|
||||
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
|
||||
() -> FleetMcp.ask(messages, T, "which config?", 5000L));
|
||||
McpSchema.CallToolResult question = send.get(5, TimeUnit.SECONDS);
|
||||
assertTrue(textOf(question).contains("[question]"), textOf(question));
|
||||
String questionText = textOf(question);
|
||||
String afterTurnId = questionText.substring(questionText.indexOf("turnId=\"") + "turnId=\"".length());
|
||||
String turnId = afterTurnId.substring(0, afterTurnId.indexOf('"'));
|
||||
|
||||
McpSchema.CallToolResult hijacked = FleetMcp.answer(messages, turnId, "evil.yaml", 500L, "term_attacker");
|
||||
assertTrue(hijacked.isError(), "a caller that did not open this turn must get an error, not an answer");
|
||||
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
|
||||
|
||||
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
|
||||
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, "term_owner"));
|
||||
assertEquals("config.yaml", textOf(ask.get(5, TimeUnit.SECONDS)));
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
FleetMcp.reply(messages, T, Role.WORKER, "done");
|
||||
assertEquals("done", textOf(answer.get(5, TimeUnit.SECONDS)));
|
||||
}
|
||||
|
||||
/**
|
||||
* Same hijack and control through the fire-and-poll ({@code sendAsync}) path: the owner comes
|
||||
* from the ticket's recorded creator terminal, not from a caller threaded through a live call.
|
||||
*/
|
||||
@Test
|
||||
void aDifferentCallersMcpAnswerIsRefusedForAnAsyncSendButTheRealOwnerSucceeds() throws Exception {
|
||||
McpSchema.CallToolResult accepted =
|
||||
FleetMcp.sendAsync(messages, T, "do it", null, Set.of(), "term_owner");
|
||||
String ticket = textOf(accepted).substring(textOf(accepted).indexOf("ticket=") + "ticket=".length()).trim();
|
||||
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting(T));
|
||||
|
||||
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
|
||||
() -> FleetMcp.ask(messages, T, "which config?", 5000L));
|
||||
MessageService.TaskView asking = messages.poll(ticket);
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
asking = messages.poll(ticket);
|
||||
}
|
||||
assertEquals(MessageService.Phase.ASKING, asking.phase());
|
||||
String turnId = asking.turnId();
|
||||
|
||||
McpSchema.CallToolResult hijacked = FleetMcp.answer(messages, turnId, "evil.yaml", 500L, "term_attacker");
|
||||
assertTrue(hijacked.isError(), "a caller that did not create this delegation must get an error");
|
||||
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
|
||||
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase(),
|
||||
"a refused answer must not advance the async ticket's phase");
|
||||
|
||||
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
|
||||
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, "term_owner"));
|
||||
assertEquals("config.yaml", textOf(ask.get(5, TimeUnit.SECONDS)));
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
FleetMcp.reply(messages, T, Role.WORKER, "done");
|
||||
assertEquals("done", textOf(answer.get(5, TimeUnit.SECONDS)));
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollUnknownTicketIsAnError() {
|
||||
McpSchema.CallToolResult res = FleetMcp.poll(messages, "task-999", null);
|
||||
@@ -296,7 +380,7 @@ class FleetMcpTest {
|
||||
|
||||
@Test
|
||||
void sendTimesOutWithAWorkingNote() {
|
||||
McpSchema.CallToolResult res = FleetMcp.send(messages, "term_a", "hi", 120L, null, Set.of());
|
||||
McpSchema.CallToolResult res = FleetMcp.send(messages, "term_a", "hi", 120L, null, Set.of(), null);
|
||||
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
|
||||
assertTrue(textOf(res).contains("no reply"), "got: " + textOf(res));
|
||||
}
|
||||
@@ -312,7 +396,7 @@ class FleetMcpTest {
|
||||
void sendTimesOutWithAnUnconfirmedNoteNotARetryInvitation() throws Exception {
|
||||
herdr.agentSendFailsWith("send_failed");
|
||||
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
|
||||
() -> FleetMcp.send(messages, T, "hi", 150L, null, Set.of()));
|
||||
() -> FleetMcp.send(messages, T, "hi", 150L, null, Set.of(), null));
|
||||
long deadline = System.currentTimeMillis() + 2000;
|
||||
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
@@ -332,13 +416,13 @@ class FleetMcpTest {
|
||||
|
||||
@Test
|
||||
void sendRejectsMissingArgs() {
|
||||
assertTrue(FleetMcp.send(messages, null, "hi", null, null, Set.of()).isError());
|
||||
assertTrue(FleetMcp.send(messages, "term_a", " ", null, null, Set.of()).isError());
|
||||
assertTrue(FleetMcp.send(messages, null, "hi", null, null, Set.of(), null).isError());
|
||||
assertTrue(FleetMcp.send(messages, "term_a", " ", null, null, Set.of(), null).isError());
|
||||
}
|
||||
|
||||
@Test
|
||||
void sendRejectsAConfiguredProfileNameBeforeAcceptingIt() {
|
||||
McpSchema.CallToolResult blocking = FleetMcp.send(messages, "sol", "hi", 100L, null, Set.of("sol"));
|
||||
McpSchema.CallToolResult blocking = FleetMcp.send(messages, "sol", "hi", 100L, null, Set.of("sol"), null);
|
||||
McpSchema.CallToolResult async = FleetMcp.sendAsync(messages, "sol", "hi", null, Set.of("sol"));
|
||||
|
||||
assertTrue(blocking.isError());
|
||||
@@ -357,7 +441,7 @@ class FleetMcpTest {
|
||||
assertSendRoundTrips("term_live_member", profiles);
|
||||
|
||||
// A herdr-owned pane outside the bridge roster cannot be classified at accept time.
|
||||
McpSchema.CallToolResult result = FleetMcp.send(messages, "external-pane", "hi", 10L, null, profiles);
|
||||
McpSchema.CallToolResult result = FleetMcp.send(messages, "external-pane", "hi", 10L, null, profiles, null);
|
||||
assertFalse(result.isError(), "an unclassified target must not be rejected at acceptance time");
|
||||
}
|
||||
|
||||
@@ -449,7 +533,7 @@ class FleetMcpTest {
|
||||
void askThenAnswerRoundTrips() throws Exception {
|
||||
// The primary delegates and blocks; wait until its waiter is open before the worker asks.
|
||||
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
|
||||
() -> FleetMcp.send(messages, "term_a", "do X", 5000L, null, Set.of()));
|
||||
() -> FleetMcp.send(messages, "term_a", "do X", 5000L, null, Set.of(), null));
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
@@ -471,7 +555,7 @@ class FleetMcpTest {
|
||||
|
||||
// The primary answers via fleet_send(turnId); this blocks again for the worker's reply.
|
||||
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
|
||||
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L));
|
||||
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, null));
|
||||
|
||||
// The worker's ask returns the answer — it resumes the same turn.
|
||||
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
|
||||
@@ -497,7 +581,7 @@ class FleetMcpTest {
|
||||
|
||||
@Test
|
||||
void answerToAStaleTurnIsAnError() {
|
||||
McpSchema.CallToolResult res = FleetMcp.answer(messages, "term_a#999", "too late", 500L);
|
||||
McpSchema.CallToolResult res = FleetMcp.answer(messages, "term_a#999", "too late", 500L, null);
|
||||
assertTrue(res.isError());
|
||||
assertTrue(textOf(res).contains("no longer open"), textOf(res));
|
||||
}
|
||||
@@ -691,7 +775,7 @@ class FleetMcpTest {
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
|
||||
|
||||
String out = textOf(res);
|
||||
// fleetd #361: reports both which coord-id a peer must use to reach ME, and this daemon's
|
||||
@@ -720,7 +804,7 @@ class FleetMcpTest {
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"mailbox\":{\"status\":\"unknown\"}"), out);
|
||||
@@ -779,19 +863,82 @@ class FleetMcpTest {
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
|
||||
boolean callerIsPrimary = Principal.worker("term_a", 1).isPrimary();
|
||||
Principal worker = Principal.worker("term_a", 1);
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), callerIsPrimary);
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), worker.isPrimary(),
|
||||
FleetMcp.leadsVisibleTo(worker), FleetMcp.membersVisibleTo(worker));
|
||||
|
||||
String out = textOf(res);
|
||||
assertFalse(out.contains("\"coordinator\""), "a worker must never see the coordinator key at all: " + out);
|
||||
assertFalse(out.contains("mac-opus"), "no fragment of the coordinator row may leak either: " + out);
|
||||
assertTrue(out.contains("\"leads\""), "the rest of the result must still be present: " + out);
|
||||
assertTrue(out.contains("\"members\""), out);
|
||||
assertFalse(out.contains("\"leads\""), "a worker must never see the leads key at all: " + out);
|
||||
assertFalse(out.contains("\"members\""), "a worker must never see the members key at all: " + out);
|
||||
assertTrue(out.contains("\"healthCoverage\""), "the rest of the result must still be present: " + out);
|
||||
}
|
||||
|
||||
/**
|
||||
* A worker's result must carry neither the {@code leads} nor the {@code members} key, and no
|
||||
* fragment of either row leaks even though both are fully populated for this call -- a key
|
||||
* check alone would pass on an implementation that still built the rows and only renamed or
|
||||
* nested them.
|
||||
*/
|
||||
@Test
|
||||
void listLeaksNoLeadOrMemberRowFragmentToAWorker() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_owner",
|
||||
new WorktreeRequest("cb-304", null));
|
||||
Principal worker = Principal.worker("term_a", 1);
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of("term_lead_x", "mac-opus"), "", FleetMcp.CoordinationSource.none(), worker.isPrimary(),
|
||||
FleetMcp.leadsVisibleTo(worker), FleetMcp.membersVisibleTo(worker));
|
||||
|
||||
String out = textOf(res);
|
||||
assertFalse(out.contains("\"leads\""), out);
|
||||
assertFalse(out.contains("\"members\""), out);
|
||||
assertFalse(out.contains(s.terminalId()), "no member row fragment may leak to a worker: " + out);
|
||||
assertFalse(out.contains("term_owner"), "no member owner fragment may leak to a worker: " + out);
|
||||
assertFalse(out.contains("mac-opus"), "no lead row fragment may leak to a worker: " + out);
|
||||
assertTrue(out.contains("\"healthCoverage\""), out);
|
||||
}
|
||||
|
||||
/**
|
||||
* A collaborator may {@code SEND} to a lead, so it must see the {@code leads} array -- the
|
||||
* only place {@code fleet_whoami} does not already give it a lead's address. It may never
|
||||
* {@code SEND} to a spawned member, so the {@code members} key must stay absent for it, with
|
||||
* no fragment of a populated member row leaking either.
|
||||
*/
|
||||
@Test
|
||||
void listShowsLeadsButNotMembersToACollaborator() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_owner",
|
||||
new WorktreeRequest("cb-304", null));
|
||||
Principal collaborator = Principal.collaborator("ops", "term_collab", 600);
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of("term_lead_x", "mac-opus"), "", FleetMcp.CoordinationSource.none(), collaborator.isPrimary(),
|
||||
FleetMcp.leadsVisibleTo(collaborator), FleetMcp.membersVisibleTo(collaborator));
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"leads\""), "a collaborator must see the leads array: " + out);
|
||||
assertTrue(out.contains("mac-opus"), "a collaborator must see the lead's name/address: " + out);
|
||||
assertFalse(out.contains("\"members\""), "a collaborator must never see the members key: " + out);
|
||||
assertFalse(out.contains(s.terminalId()), "no member row fragment may leak to a collaborator: " + out);
|
||||
assertFalse(out.contains("term_owner"), "no member owner fragment may leak to a collaborator: " + out);
|
||||
assertTrue(out.contains("\"healthCoverage\""), out);
|
||||
}
|
||||
|
||||
@@ -807,16 +954,19 @@ class FleetMcpTest {
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
|
||||
boolean callerIsPrimary = Principal.architect("lead-designer", "term_design", 400).isPrimary();
|
||||
Principal architect = Principal.architect("lead-designer", "term_design", 400);
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), callerIsPrimary);
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), architect.isPrimary(),
|
||||
FleetMcp.leadsVisibleTo(architect), FleetMcp.membersVisibleTo(architect));
|
||||
|
||||
String out = textOf(res);
|
||||
assertFalse(out.contains("\"coordinator\""), "an architect must never see the coordinator key either: " + out);
|
||||
assertTrue(out.contains("\"leads\""), "an architect must still see the leads array: " + out);
|
||||
assertTrue(out.contains("\"members\""), "an architect must still see the members array: " + out);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -841,7 +991,7 @@ class FleetMcpTest {
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true));
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true));
|
||||
|
||||
assertTrue(gatedAsPrimary.contains("\"coordinator\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"selfId\":\"mac-opus\""), gatedAsPrimary);
|
||||
@@ -850,6 +1000,8 @@ class FleetMcpTest {
|
||||
assertTrue(gatedAsPrimary.contains("\"heldDurable\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"held\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"peers\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"leads\""), "the primary must still see the leads array: " + gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"members\""), "the primary must still see the members array: " + gatedAsPrimary);
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -868,7 +1020,7 @@ class FleetMcpTest {
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"msgId\":\"m1\""), out);
|
||||
@@ -879,6 +1031,83 @@ class FleetMcpTest {
|
||||
assertTrue(out.contains("x".repeat(80) + "…"), "expected an 80-char preview with an ellipsis: " + out);
|
||||
}
|
||||
|
||||
// --- fleetd #703: fleet_list's collaborators array -------------------------------------------
|
||||
|
||||
/** Calls the canonical {@code listFleet} overload directly, so a test can set the collaborators
|
||||
* payload and its visibility independently of a real {@code Principal} / MCP exchange. */
|
||||
private static McpSchema.CallToolResult listFleetWithCollaborators(FakeHerdr h,
|
||||
Map<String, String> collaborators, boolean collaboratorsVisible) {
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
return FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
|
||||
FleetMcp.LeadSeatSource.none(), new LeadContextGauge(), FleetMcp.LeadConfigDirSource.none(),
|
||||
Map.of(), "", collaborators, collaboratorsVisible,
|
||||
FleetMcp.CoordinationSource.none(), false, true, true);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #703 acceptance A: a visible caller with one configured collaborator gets a
|
||||
* {@code collaborators} row whose key ({@code CallerResolver.collaborators()}'s
|
||||
* {@code terminal_id -> name} entry) lands as that row's {@code sessionId}, and {@code leads}/
|
||||
* {@code members} are unaffected by the new key.
|
||||
*/
|
||||
@Test
|
||||
void listReportsACollaboratorsRowKeyedByTheCollaboratorsTerminalId() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
String out = textOf(listFleetWithCollaborators(h, Map.of("term_collab", "ops"), true));
|
||||
|
||||
assertTrue(out.contains("\"collaborators\":["), out);
|
||||
assertTrue(out.contains("\"name\":\"ops\""), out);
|
||||
assertTrue(out.contains("\"sessionId\":\"term_collab\""), out);
|
||||
assertTrue(out.contains("\"leads\":[]"), out);
|
||||
assertTrue(out.contains("\"members\":[]"), out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #703 acceptance A, control half: with no collaborator configured, the key is absent
|
||||
* (caller cannot see it) or an empty array (caller can), and {@code leads}/{@code members} are
|
||||
* unchanged either way.
|
||||
*/
|
||||
@Test
|
||||
void listOmitsOrEmptiesCollaboratorsWhenNoneAreConfigured() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
|
||||
String visibleButEmpty = textOf(listFleetWithCollaborators(h, Map.of(), true));
|
||||
assertTrue(visibleButEmpty.contains("\"collaborators\":[]"), visibleButEmpty);
|
||||
assertTrue(visibleButEmpty.contains("\"leads\":[]"), visibleButEmpty);
|
||||
assertTrue(visibleButEmpty.contains("\"members\":[]"), visibleButEmpty);
|
||||
|
||||
String notVisible = textOf(listFleetWithCollaborators(h, Map.of(), false));
|
||||
assertFalse(notVisible.contains("\"collaborators\""), notVisible);
|
||||
assertTrue(notVisible.contains("\"leads\":[]"), notVisible);
|
||||
assertTrue(notVisible.contains("\"members\":[]"), notVisible);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #703 acceptance B: a worker must not see the {@code collaborators} array at all, while
|
||||
* an architect -- a role that also holds READ, same as a worker -- does see it. Both halves are
|
||||
* asserted: a test that only checked the worker-hidden half would pass even if the feature were
|
||||
* never wired up for anyone.
|
||||
*/
|
||||
@Test
|
||||
void listOmitsCollaboratorsForAWorkerAndIncludesThemForAnArchitect() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
Map<String, String> collaborators = Map.of("term_collab", "ops");
|
||||
|
||||
String asWorker = textOf(listFleetWithCollaborators(h, collaborators,
|
||||
FleetMcp.collaboratorsVisibleTo(Principal.worker("term_w", 1))));
|
||||
assertFalse(asWorker.contains("\"collaborators\""),
|
||||
"a worker must not see the collaborators array: " + asWorker);
|
||||
|
||||
String asArchitect = textOf(listFleetWithCollaborators(h, collaborators,
|
||||
FleetMcp.collaboratorsVisibleTo(Principal.architect("design", "term_arch", 2))));
|
||||
assertTrue(asArchitect.contains("\"collaborators\":["),
|
||||
"an architect must see the collaborators array: " + asArchitect);
|
||||
assertTrue(asArchitect.contains("\"sessionId\":\"term_collab\""), asArchitect);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #421: {@code mailbox.pending} counts only broker-ready messages, so a blocked lead's
|
||||
* normal, healthy state is {@code "pending": 0} next to a non-empty {@code held[]} — which
|
||||
@@ -900,7 +1129,7 @@ class FleetMcpTest {
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"pending\":0"), out);
|
||||
@@ -929,7 +1158,7 @@ class FleetMcpTest {
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"heldDurable\":false"),
|
||||
@@ -998,7 +1227,7 @@ class FleetMcpTest {
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "",
|
||||
new FleetMcp.CoordinationSource(channel, List.of("fleet01-lead", "fleet02-lead", "fleet03-lead")), true);
|
||||
new FleetMcp.CoordinationSource(channel, List.of("fleet01-lead", "fleet02-lead", "fleet03-lead")), true, true, true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"coordId\":\"fleet01-lead\",\"status\":\"exists\",\"pending\":2,\"consumers\":1"), out);
|
||||
@@ -1691,8 +1920,8 @@ class FleetMcpTest {
|
||||
Principal architect = Principal.architect("lead-designer", "term_design", 400);
|
||||
MemberPresence presence = new MemberPresence();
|
||||
|
||||
FleetMcp.markSpawnedMemberPresent(worker, presence);
|
||||
FleetMcp.markSpawnedMemberPresent(architect, presence);
|
||||
FleetMcp.markTrackedCallerPresent(worker, presence);
|
||||
FleetMcp.markTrackedCallerPresent(architect, presence);
|
||||
|
||||
assertTrue(presence.isPresent("term_worker"));
|
||||
assertTrue(presence.isPresent("term_design"));
|
||||
@@ -1703,25 +1932,40 @@ class FleetMcpTest {
|
||||
Principal lead = Principal.leader("opus", "term_lead", 100);
|
||||
MemberPresence presence = new MemberPresence();
|
||||
|
||||
FleetMcp.markSpawnedMemberPresent(lead, presence);
|
||||
FleetMcp.markSpawnedMemberPresent(Principal.anonymous(), presence);
|
||||
FleetMcp.markTrackedCallerPresent(lead, presence);
|
||||
FleetMcp.markTrackedCallerPresent(Principal.anonymous(), presence);
|
||||
|
||||
assertFalse(presence.isPresent("term_lead"));
|
||||
}
|
||||
|
||||
/**
|
||||
* The item whose absence would be silent: an observer's own MCP contact must still mark
|
||||
* presence, or a pane resolving to the unconfigured-pane floor would sit on the injector
|
||||
* readiness gate forever once something addresses it.
|
||||
*/
|
||||
@Test
|
||||
void anObserverContactMarksPresence() {
|
||||
Principal observer = Principal.observer("term_observer", 800);
|
||||
MemberPresence presence = new MemberPresence();
|
||||
|
||||
FleetMcp.markTrackedCallerPresent(observer, presence);
|
||||
|
||||
assertTrue(presence.isPresent("term_observer"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void statusReportsLiveAgentStatus() {
|
||||
FakeHerdr blocked = new FakeHerdr().agentStatus("blocked");
|
||||
AgentControl blockedAgents = new AgentControl(blocked);
|
||||
McpSchema.CallToolResult res = FleetMcp.status(
|
||||
new MessageService(blockedAgents, new Injector(blockedAgents), rendezvous), "term_a");
|
||||
new MessageService(blockedAgents, new Injector(blockedAgents), rendezvous), "term_a", null);
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
assertEquals("blocked", textOf(res));
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-582: a lead polling {@code fleet_status} on its normal cadence — not {@code fleet_poll}
|
||||
* — must also see a worker's open async {@code fleet_ask} question, since the reverse-rendezvous
|
||||
* A lead polling {@code fleet_status} on its normal cadence — not {@code fleet_poll} — must
|
||||
* also see a worker's open async {@code fleet_ask} question, since the reverse-rendezvous
|
||||
* window it opened with is far shorter than that cadence.
|
||||
*/
|
||||
@Test
|
||||
@@ -1744,7 +1988,7 @@ class FleetMcpTest {
|
||||
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
|
||||
assertEquals(MessageService.Phase.ASKING, asking.phase());
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.status(messages, T);
|
||||
McpSchema.CallToolResult res = FleetMcp.status(messages, T, null);
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
assertTrue(out.startsWith("idle"), "the live status must still lead the text: " + out);
|
||||
@@ -1755,7 +1999,64 @@ class FleetMcpTest {
|
||||
// Clean up the still-open ask so the background thread does not linger past the test.
|
||||
String turnId = asking.turnId();
|
||||
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||
() -> messages.answer(turnId, "config.yaml", 5000));
|
||||
() -> messages.answer(turnId, "config.yaml", 5000, null));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
answer.get(5, TimeUnit.SECONDS);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code fleet_status}'s pending-ask block (the question, its {@code turnId} and its ticket)
|
||||
* is shown only to the caller whose terminal created the delegation, or to a caller with no
|
||||
* terminal at all (the unnamed primary) — a different terminal-bearing caller still sees the
|
||||
* base status line, but none of the pending-ask fields.
|
||||
*/
|
||||
@Test
|
||||
void statusGatesThePendingAskFieldsByTheDelegationsCreatorTerminal() throws Exception {
|
||||
String ticket = messages.sendAsync(T, "task that asks", null, "term_creator");
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting(T), "sendAsync should have opened its rendezvous waiter");
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||
|
||||
MessageService.TaskView asking;
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
do {
|
||||
asking = messages.poll(ticket);
|
||||
Thread.sleep(5);
|
||||
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
|
||||
assertEquals(MessageService.Phase.ASKING, asking.phase());
|
||||
|
||||
String other = textOf(FleetMcp.status(messages, T, "term_other"));
|
||||
assertTrue(other.startsWith("idle"), "the base status must still be shown: " + other);
|
||||
assertFalse(other.contains("which config file?"),
|
||||
"a non-creating caller must not see the question text: " + other);
|
||||
assertFalse(other.contains(asking.turnId()),
|
||||
"a non-creating caller must not see the turnId: " + other);
|
||||
assertFalse(other.contains(ticket),
|
||||
"a non-creating caller must not see the ticket: " + other);
|
||||
|
||||
String creator = textOf(FleetMcp.status(messages, T, "term_creator"));
|
||||
assertTrue(creator.contains("which config file?"), "the creator must see the question: " + creator);
|
||||
assertTrue(creator.contains(asking.turnId()), "the creator must see the turnId: " + creator);
|
||||
assertTrue(creator.contains(ticket), "the creator must see the ticket: " + creator);
|
||||
|
||||
String unnamed = textOf(FleetMcp.status(messages, T, null));
|
||||
assertTrue(unnamed.contains("which config file?"),
|
||||
"a caller with no terminal (the unnamed primary) must see the question: " + unnamed);
|
||||
|
||||
// Clean up the still-open ask so the background thread does not linger past the test.
|
||||
String turnId = asking.turnId();
|
||||
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||
() -> messages.answer(turnId, "config.yaml", 5000, "term_creator"));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||
@@ -1837,6 +2138,53 @@ class FleetMcpTest {
|
||||
assertTrue(out.contains("\"sessionId\":\"term_design\""), out);
|
||||
}
|
||||
|
||||
/**
|
||||
* A collaborator reports its own role and name, never the {@code leader} key a lead gets —
|
||||
* {@code role} already reads {@code "collaborator"}, so a {@code leader} key alongside it
|
||||
* would be self-contradicting. Control: the same call shape fed a named lead must still carry
|
||||
* {@code leader}, so this is not passing because the key stopped being emitted for everyone.
|
||||
*/
|
||||
@Test
|
||||
void whoamiReportsACollaboratorWithNoLeaderKeyButALeadStillGetsOne() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
McpSchema.CallToolResult collabRes = FleetMcp.whoami(
|
||||
Principal.collaborator("ops", "term_collab", 700), sessions);
|
||||
assertNotEquals(Boolean.TRUE, collabRes.isError());
|
||||
String collabOut = textOf(collabRes);
|
||||
assertTrue(collabOut.contains("\"role\":\"collaborator\""), collabOut);
|
||||
assertTrue(collabOut.contains("\"collaborator\":\"ops\""), collabOut);
|
||||
assertTrue(collabOut.contains("\"sessionId\":\"term_collab\""), collabOut);
|
||||
assertFalse(collabOut.contains("leader"), collabOut);
|
||||
|
||||
McpSchema.CallToolResult leadRes = FleetMcp.whoami(
|
||||
Principal.leader("opus", "term_lead", 100), sessions);
|
||||
String leadOut = textOf(leadRes);
|
||||
assertTrue(leadOut.contains("\"leader\":\"opus\""), leadOut);
|
||||
}
|
||||
|
||||
/**
|
||||
* An observer reports its own role and pane, never a {@code leader} key. Without an explicit
|
||||
* branch it would reach the lead branch by elimination and look right only because the
|
||||
* {@code leader} key is guarded on a non-null name — this pins the branch rather than the
|
||||
* accident.
|
||||
*/
|
||||
@Test
|
||||
void whoamiReportsAnObserverNotALead() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.whoami(
|
||||
Principal.observer("term_observer", 900), sessions);
|
||||
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"role\":\"observer\""), out);
|
||||
assertTrue(out.contains("\"sessionId\":\"term_observer\""), out);
|
||||
assertFalse(out.contains("leader"), out);
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-548: an architect SEND delegates as its own pane (recording the per-target delegation) but
|
||||
* must NEVER become the legacy singleton "primary" fallback — the per-target map does not cure
|
||||
|
||||
@@ -0,0 +1,280 @@
|
||||
package dev.ltms.fleet.msg;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
import java.util.stream.Stream;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* Pins that no file under {@code src/main/java} calls the fail-open
|
||||
* {@link MessageService#poll(String)} overload. That overload skips the ownership check in
|
||||
* {@code MessageService}'s {@code ownsTicket} entirely, so a caller of it can read any session's
|
||||
* ticket. Every production caller must go through {@link MessageService#poll(String, String)}
|
||||
* and pass a {@code callerTerminal} explicitly, even when it is {@code null}.
|
||||
*
|
||||
* <p>This reads each file's own source text rather than reflecting on compiled bytecode, because
|
||||
* the risk is a future one-word edit at a call site, not a missing overload.
|
||||
*
|
||||
* <p>The scan below finds a violation by its receiver, {@code messages.poll(}, rather than the
|
||||
* bare method name, so it does not mistake {@link java.util.Queue#poll()} for a violation. That
|
||||
* anchor only covers a {@code MessageService} reached through a variable or field named
|
||||
* {@code messages}, so {@link #everyMessageServiceDeclarationIsNamedMessages} pins the naming
|
||||
* convention the anchor depends on: a declaration under any other name would be invisible to the
|
||||
* scan above, and must turn this second check red instead of passing silently.
|
||||
*/
|
||||
class MessageServicePollUsageTest {
|
||||
|
||||
private static final Path PRODUCTION_SOURCE = Path.of("src/main/java");
|
||||
|
||||
@Test
|
||||
void noProductionFileCallsTheSingleArgumentPollOverload() throws IOException {
|
||||
List<String> violations = new ArrayList<>();
|
||||
List<String> twoArgSites = new ArrayList<>();
|
||||
int filesScanned = scanForPollCalls(PRODUCTION_SOURCE, violations, twoArgSites);
|
||||
|
||||
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "no
|
||||
// violations found" having looked at nothing.
|
||||
assertTrue(filesScanned > 0, "control failed: the scan under " + PRODUCTION_SOURCE
|
||||
+ " visited zero .java files -- the path is wrong, so the absence of violations "
|
||||
+ "below proves nothing");
|
||||
|
||||
assertTrue(violations.isEmpty(), "found a call to the fail-open MessageService.poll(String) "
|
||||
+ "overload, which skips the ownership check entirely -- pass a callerTerminal "
|
||||
+ "explicitly (even if null) through poll(String, String) instead: " + violations);
|
||||
|
||||
// CONTROL: the arity parser actually finds the two genuine two-argument call sites (the
|
||||
// MCP handler in FleetMcp and the REST handler in FleetApp). If this drops, the parser
|
||||
// itself is broken, not the production code -- a broken parser (or a scan root that
|
||||
// reaches no real source) must fail loudly here rather than pass vacuously above.
|
||||
assertEquals(2, twoArgSites.size(), "control failed: expected exactly the two known "
|
||||
+ "two-argument messages.poll(...) call sites, found: " + twoArgSites);
|
||||
assertTrue(twoArgSites.stream().anyMatch(s -> s.contains("FleetMcp.java")),
|
||||
"control failed: did not find the FleetMcp.java messages.poll(ticket, callerTerminal) "
|
||||
+ "site among: " + twoArgSites);
|
||||
assertTrue(twoArgSites.stream().anyMatch(s -> s.contains("FleetApp.java")),
|
||||
"control failed: did not find the FleetApp.java messages.poll(...) site among: "
|
||||
+ twoArgSites);
|
||||
}
|
||||
|
||||
/**
|
||||
* The {@code messages.poll(} anchor above only sees a {@code MessageService} reached through
|
||||
* a variable, field, or parameter named {@code messages}. This asserts that every such
|
||||
* declaration under {@code src/main/java} uses that name, so a differently named declaration
|
||||
* -- invisible to the scan above -- fails loudly here instead of letting that scan pass on a
|
||||
* call site it never looked at.
|
||||
*/
|
||||
@Test
|
||||
void everyMessageServiceDeclarationIsNamedMessages() throws IOException {
|
||||
List<String> names = new ArrayList<>();
|
||||
int filesScanned = scanForDeclarationNames(PRODUCTION_SOURCE, names);
|
||||
|
||||
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "every
|
||||
// declaration is named messages" having looked at nothing.
|
||||
assertTrue(filesScanned > 0, "control failed: the scan under " + PRODUCTION_SOURCE
|
||||
+ " visited zero .java files -- the path is wrong, so the result below proves nothing");
|
||||
|
||||
// CONTROL: the declaration pattern actually finds real declarations. Zero means the
|
||||
// pattern is broken, not that every MessageService variable, field, or parameter vanished.
|
||||
assertTrue(names.size() > 0, "control failed: found zero MessageService declarations under "
|
||||
+ PRODUCTION_SOURCE + " -- the declaration pattern is broken, update it before "
|
||||
+ "trusting the naming check below");
|
||||
|
||||
List<String> other = names.stream().filter(n -> !n.equals("messages")).distinct().toList();
|
||||
assertTrue(other.isEmpty(), "found a MessageService declaration not named \"messages\": "
|
||||
+ other + " -- the messages.poll( scan above only looks for that name, so a call "
|
||||
+ "through a differently named variable or field is invisible to it; either rename "
|
||||
+ "the declaration or widen that scan's anchor to cover it");
|
||||
}
|
||||
|
||||
private static int scanForPollCalls(Path root, List<String> violations, List<String> twoArgSites)
|
||||
throws IOException {
|
||||
List<Path> files = javaFiles(root);
|
||||
for (Path file : files) {
|
||||
scanFileForPollCalls(file, violations, twoArgSites);
|
||||
}
|
||||
return files.size();
|
||||
}
|
||||
|
||||
private static int scanForDeclarationNames(Path root, List<String> names) throws IOException {
|
||||
List<Path> files = javaFiles(root);
|
||||
Pattern declaration = Pattern.compile("MessageService\\s+([A-Za-z_][A-Za-z0-9_]*)");
|
||||
for (Path file : files) {
|
||||
if (file.getFileName().toString().equals("MessageService.java")) {
|
||||
continue; // the type's own declaration, not a caller holding a reference to it
|
||||
}
|
||||
String stripped = stripComments(Files.readString(file));
|
||||
Matcher m = declaration.matcher(stripped);
|
||||
while (m.find()) {
|
||||
int j = m.end();
|
||||
while (j < stripped.length() && Character.isWhitespace(stripped.charAt(j))) j++;
|
||||
if (j < stripped.length() && stripped.charAt(j) == '(') {
|
||||
continue; // a method named like the convention, e.g. "MessageService messages()"
|
||||
}
|
||||
names.add(m.group(1));
|
||||
}
|
||||
}
|
||||
return files.size();
|
||||
}
|
||||
|
||||
private static List<Path> javaFiles(Path root) throws IOException {
|
||||
try (Stream<Path> paths = Files.walk(root)) {
|
||||
return paths.filter(p -> p.toString().endsWith(".java")).toList();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Replaces {@code //} and {@code /* *}{@code /} comment text with nothing, leaving code,
|
||||
* string/char literals and line breaks untouched -- so a comment that merely mentions
|
||||
* {@code MessageService} in prose can never be read as a declaration.
|
||||
*/
|
||||
private static String stripComments(String source) {
|
||||
StringBuilder out = new StringBuilder(source.length());
|
||||
boolean inString = false;
|
||||
boolean inChar = false;
|
||||
int i = 0;
|
||||
while (i < source.length()) {
|
||||
char c = source.charAt(i);
|
||||
if (inString) {
|
||||
out.append(c);
|
||||
if (c == '\\' && i + 1 < source.length()) { out.append(source.charAt(i + 1)); i += 2; continue; }
|
||||
if (c == '"') inString = false;
|
||||
i++;
|
||||
continue;
|
||||
}
|
||||
if (inChar) {
|
||||
out.append(c);
|
||||
if (c == '\\' && i + 1 < source.length()) { out.append(source.charAt(i + 1)); i += 2; continue; }
|
||||
if (c == '\'') inChar = false;
|
||||
i++;
|
||||
continue;
|
||||
}
|
||||
if (c == '"') { inString = true; out.append(c); i++; continue; }
|
||||
if (c == '\'') { inChar = true; out.append(c); i++; continue; }
|
||||
if (c == '/' && i + 1 < source.length() && source.charAt(i + 1) == '/') {
|
||||
while (i < source.length() && source.charAt(i) != '\n') i++;
|
||||
continue; // leaves the newline itself for the next iteration to append
|
||||
}
|
||||
if (c == '/' && i + 1 < source.length() && source.charAt(i + 1) == '*') {
|
||||
i += 2;
|
||||
while (i < source.length() && !(source.charAt(i) == '*' && i + 1 < source.length()
|
||||
&& source.charAt(i + 1) == '/')) {
|
||||
if (source.charAt(i) == '\n') out.append('\n');
|
||||
i++;
|
||||
}
|
||||
i += 2;
|
||||
continue;
|
||||
}
|
||||
out.append(c);
|
||||
i++;
|
||||
}
|
||||
return out.toString();
|
||||
}
|
||||
|
||||
private static void scanFileForPollCalls(Path file, List<String> violations, List<String> twoArgSites)
|
||||
throws IOException {
|
||||
String source = Files.readString(file);
|
||||
String needle = "messages.poll(";
|
||||
int from = 0;
|
||||
int idx;
|
||||
while ((idx = source.indexOf(needle, from)) >= 0) {
|
||||
int argsStart = idx + needle.length();
|
||||
String args = extractBalancedArgs(source, argsStart, file, idx);
|
||||
int closeParenIndex = argsStart + args.length();
|
||||
from = closeParenIndex + 1;
|
||||
|
||||
if (args.isBlank()) {
|
||||
continue; // MessageService has no zero-argument poll() -- nothing to classify
|
||||
}
|
||||
String site = file + ":" + lineOf(source, idx);
|
||||
if (topLevelCommaCount(args) == 0) {
|
||||
violations.add(site + " -- messages.poll(" + args.trim() + ")");
|
||||
} else {
|
||||
twoArgSites.add(site);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The text between {@code messages.poll(} and its matching close paren: balanced over nested
|
||||
* calls, and never split by a paren or comma sitting inside a string or char literal.
|
||||
*/
|
||||
private static String extractBalancedArgs(String source, int start, Path file, int callIndex) {
|
||||
int depth = 1;
|
||||
boolean inString = false;
|
||||
boolean inChar = false;
|
||||
int i = start;
|
||||
while (i < source.length()) {
|
||||
char c = source.charAt(i);
|
||||
if (inString) {
|
||||
if (c == '\\') { i += 2; continue; }
|
||||
if (c == '"') inString = false;
|
||||
} else if (inChar) {
|
||||
if (c == '\\') { i += 2; continue; }
|
||||
if (c == '\'') inChar = false;
|
||||
} else if (c == '"') {
|
||||
inString = true;
|
||||
} else if (c == '\'') {
|
||||
inChar = true;
|
||||
} else if (c == '(') {
|
||||
depth++;
|
||||
} else if (c == ')') {
|
||||
depth--;
|
||||
if (depth == 0) return source.substring(start, i);
|
||||
}
|
||||
i++;
|
||||
}
|
||||
throw new IllegalStateException(
|
||||
"unbalanced parentheses scanning " + file + ":" + lineOf(source, callIndex));
|
||||
}
|
||||
|
||||
/**
|
||||
* Commas at paren/bracket/brace depth zero, skipping string and char literals -- the argument
|
||||
* separators a human reader would see, not every comma character in the text.
|
||||
*/
|
||||
private static int topLevelCommaCount(String args) {
|
||||
int depth = 0;
|
||||
int commas = 0;
|
||||
boolean inString = false;
|
||||
boolean inChar = false;
|
||||
int i = 0;
|
||||
while (i < args.length()) {
|
||||
char c = args.charAt(i);
|
||||
if (inString) {
|
||||
if (c == '\\') { i += 2; continue; }
|
||||
if (c == '"') inString = false;
|
||||
} else if (inChar) {
|
||||
if (c == '\\') { i += 2; continue; }
|
||||
if (c == '\'') inChar = false;
|
||||
} else if (c == '"') {
|
||||
inString = true;
|
||||
} else if (c == '\'') {
|
||||
inChar = true;
|
||||
} else if (c == '(' || c == '[' || c == '{') {
|
||||
depth++;
|
||||
} else if (c == ')' || c == ']' || c == '}') {
|
||||
depth--;
|
||||
} else if (c == ',' && depth == 0) {
|
||||
commas++;
|
||||
}
|
||||
i++;
|
||||
}
|
||||
return commas;
|
||||
}
|
||||
|
||||
private static int lineOf(String source, int index) {
|
||||
int line = 1;
|
||||
for (int i = 0; i < index; i++) {
|
||||
if (source.charAt(i) == '\n') line++;
|
||||
}
|
||||
return line;
|
||||
}
|
||||
}
|
||||
@@ -60,13 +60,25 @@ class MessageServiceTest {
|
||||
inbox.own(T);
|
||||
}
|
||||
|
||||
/**
|
||||
* A send budget large enough that a test's own setup — {@link #awaitWaiting()} plus whatever
|
||||
* status transitions it drives afterward — can never compete with it for the same clock. A test
|
||||
* that needs {@code send.get(...)}'s own window to be the only timing bound it depends on uses
|
||||
* {@link #sendAsync(String, long)} with this value instead of the default 5000 ms.
|
||||
*/
|
||||
private static final long GENEROUS_SEND_BUDGET_MILLIS = 30_000;
|
||||
|
||||
/** Run {@code send} on a background thread; the current thread drives the worker's turn. */
|
||||
private CompletableFuture<MessageService.Reply> sendAsync() {
|
||||
return sendAsync("do the task");
|
||||
}
|
||||
|
||||
private CompletableFuture<MessageService.Reply> sendAsync(String content) {
|
||||
return CompletableFuture.supplyAsync(() -> messages.send(T, content, 5000));
|
||||
return sendAsync(content, 5000);
|
||||
}
|
||||
|
||||
private CompletableFuture<MessageService.Reply> sendAsync(String content, long timeoutMillis) {
|
||||
return CompletableFuture.supplyAsync(() -> messages.send(T, content, timeoutMillis, null));
|
||||
}
|
||||
|
||||
private void awaitWaiting() throws InterruptedException {
|
||||
@@ -80,7 +92,7 @@ class MessageServiceTest {
|
||||
|
||||
@Test
|
||||
void completionFallbackResolvesATurnThatNeverCalledFleetReply() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync("do the task", GENEROUS_SEND_BUDGET_MILLIS);
|
||||
awaitWaiting();
|
||||
|
||||
herdr.readText("$ prompt"); // pre-turn pane: no answer yet (baseline reference)
|
||||
@@ -96,6 +108,32 @@ class MessageServiceTest {
|
||||
assertTrue(reply.completed(), "a scraped completion still counts as completed");
|
||||
}
|
||||
|
||||
/**
|
||||
* Pins {@link #GENEROUS_SEND_BUDGET_MILLIS} as the budget {@link
|
||||
* #completionFallbackResolvesATurnThatNeverCalledFleetReply} depends on. A 5500 ms delay between
|
||||
* {@link #awaitWaiting()} and the status transitions that drive completion stands in for a loaded
|
||||
* machine's setup overhead — comfortably past the 5000 ms budget this send no longer uses, and
|
||||
* still well inside this method's own 30 000 ms budget. The only clock this test depends on is
|
||||
* {@code send.get}'s own 10 s window.
|
||||
*/
|
||||
@Test
|
||||
void completionFallbackSurvivesASlowHarnessBecauseItsSendBudgetIsNotTheBindingClock() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync("do the task", GENEROUS_SEND_BUDGET_MILLIS);
|
||||
awaitWaiting();
|
||||
|
||||
Thread.sleep(5500);
|
||||
|
||||
herdr.readText("$ prompt");
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
injector.onStatus(T, AgentStatus.WORKING);
|
||||
herdr.readText("BUILD GREEN: 391 files");
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
MessageService.Reply reply = send.get(10, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, reply.outcome(),
|
||||
"a slow harness must not be mistaken for a timed-out delivery");
|
||||
}
|
||||
|
||||
@Test
|
||||
void completionFallbackReplacesAnEchoedInjectedBriefWithNoReportOutcome() throws Exception {
|
||||
String brief = "Implement the requested change. ".repeat(20);
|
||||
@@ -150,7 +188,7 @@ class MessageServiceTest {
|
||||
localRendezvous, new InMemoryReplyInbox());
|
||||
String brief = "Implement the requested change. ".repeat(20);
|
||||
CompletableFuture<MessageService.Reply> send =
|
||||
CompletableFuture.supplyAsync(() -> localMessages.send(T, brief, 5000));
|
||||
CompletableFuture.supplyAsync(() -> localMessages.send(T, brief, 5000, (String) null));
|
||||
long deadline = System.currentTimeMillis() + 2000;
|
||||
while (!localRendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
@@ -303,7 +341,7 @@ class MessageServiceTest {
|
||||
|
||||
// The primary answers via fleet_send(turnId); this blocks again for the worker's reply.
|
||||
CompletableFuture<MessageService.Reply> answer =
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
|
||||
|
||||
// The worker's ask returns the answer — it resumes the same turn.
|
||||
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
|
||||
@@ -339,7 +377,7 @@ class MessageServiceTest {
|
||||
|
||||
// The primary answers that one turnId; both asks unblock with the same answer.
|
||||
CompletableFuture<MessageService.Reply> answer =
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
|
||||
|
||||
MessageService.AskResult a1 = ask1.get(5, TimeUnit.SECONDS);
|
||||
MessageService.AskResult a2 = ask2.get(5, TimeUnit.SECONDS);
|
||||
@@ -434,7 +472,7 @@ class MessageServiceTest {
|
||||
"the duplicate's own timeout elapses first");
|
||||
|
||||
CompletableFuture<MessageService.Reply> answered =
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "fleetd.yaml", 500));
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "fleetd.yaml", 500, null));
|
||||
|
||||
MessageService.AskResult a = fresh.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome(),
|
||||
@@ -469,7 +507,7 @@ class MessageServiceTest {
|
||||
|
||||
CompletableFuture<MessageService.Reply> lateAnswer = new CompletableFuture<>();
|
||||
messages.setAskTimeoutRaceHookForTest(() ->
|
||||
lateAnswer.complete(messages.answer(turnId, "too late", 500)));
|
||||
lateAnswer.complete(messages.answer(turnId, "too late", 500, null)));
|
||||
try {
|
||||
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.AskOutcome.TIMED_OUT, a.outcome());
|
||||
@@ -501,7 +539,7 @@ class MessageServiceTest {
|
||||
|
||||
@Test
|
||||
void answeringAnUnknownTurnIsStale() {
|
||||
MessageService.Reply r = messages.answer(T + "#999", "too late", 500);
|
||||
MessageService.Reply r = messages.answer(T + "#999", "too late", 500, null);
|
||||
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
|
||||
"an answer to a turn that never existed (or already lapsed) is stale, not a hang");
|
||||
}
|
||||
@@ -534,7 +572,7 @@ class MessageServiceTest {
|
||||
|
||||
messages.setAnswerAskLapseRaceHookForTest(() -> rendezvous.answerAsk(turnId, "raced in first"));
|
||||
try {
|
||||
MessageService.Reply r = messages.answer(turnId, "too late", 500);
|
||||
MessageService.Reply r = messages.answer(turnId, "too late", 500, null);
|
||||
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
|
||||
"an ask already answered by the race must be seen as lapsed, not double-delivered");
|
||||
} finally {
|
||||
@@ -558,7 +596,7 @@ class MessageServiceTest {
|
||||
void sendTimesOutBeforeDeliveryIsQueuedNotWorking() {
|
||||
// Nothing ever delivers the message and nothing resolves the send, so the reply future
|
||||
// times out with delivery still incomplete — the message is still queued for the worker.
|
||||
MessageService.Reply r = messages.send(T, "never delivered", 50);
|
||||
MessageService.Reply r = messages.send(T, "never delivered", 50, null);
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome(),
|
||||
"an undelivered send that times out is still queued, not working");
|
||||
assertNull(r.text());
|
||||
@@ -567,7 +605,7 @@ class MessageServiceTest {
|
||||
@Test
|
||||
void sendTimesOutAfterDeliveryIsStillWorking() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send =
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 300));
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 300, null));
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver — the delivered future now completes
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker starts but never replies
|
||||
@@ -592,7 +630,7 @@ class MessageServiceTest {
|
||||
void sendTimeoutUsesCancellationDeliveredWhenPickupWinsTheRace() {
|
||||
messages.setTimeoutCancellationRaceHookForTest(() -> injector.onStatus(T, AgentStatus.IDLE));
|
||||
try {
|
||||
MessageService.Reply reply = messages.send(T, "race delivery", 50);
|
||||
MessageService.Reply reply = messages.send(T, "race delivery", 50, null);
|
||||
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, reply.outcome(),
|
||||
"cancel reporting DELIVERED means the worker received the timed-out message");
|
||||
@@ -614,7 +652,7 @@ class MessageServiceTest {
|
||||
void sendTimesOutWithAttemptedDeliveryReportsUnconfirmedNotQueued() throws Exception {
|
||||
herdr.agentSendFailsWith("send_failed");
|
||||
CompletableFuture<MessageService.Reply> send =
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "brief", 150));
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "brief", 150, null));
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE); // triggers the failing delivery attempt → ATTEMPTED
|
||||
|
||||
@@ -641,7 +679,7 @@ class MessageServiceTest {
|
||||
|
||||
// The primary answers, unblocking the worker; but the worker never sends the follow-up
|
||||
// fleet_reply, so the answering send rides out its short window as still-working.
|
||||
MessageService.Reply answer = messages.answer(q.turnId(), "config.yaml", 200);
|
||||
MessageService.Reply answer = messages.answer(q.turnId(), "config.yaml", 200, null);
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answer.outcome(),
|
||||
"an answered worker that never replies times out as still working");
|
||||
|
||||
@@ -677,7 +715,7 @@ class MessageServiceTest {
|
||||
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
|
||||
|
||||
CompletableFuture<MessageService.Reply> answer =
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
|
||||
ask.get(5, TimeUnit.SECONDS); // worker resumed with the answer
|
||||
|
||||
awaitWaiting(); // the answering call has (re)opened its own forward waiter
|
||||
@@ -685,7 +723,7 @@ class MessageServiceTest {
|
||||
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.REPLIED, done.outcome());
|
||||
|
||||
MessageService.Reply probe = messages.send(T, "probe after normal reply", 300);
|
||||
MessageService.Reply probe = messages.send(T, "probe after normal reply", 300, null);
|
||||
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
|
||||
"the session lock must be released after a normal REPLIED answer(), or this bounded "
|
||||
+ "follow-up send would come back BUSY instead of timing out on its own work");
|
||||
@@ -718,13 +756,13 @@ class MessageServiceTest {
|
||||
// The primary answers, unblocking the worker; the worker never sends its follow-up
|
||||
// fleet_reply, so the answering call rides out its short window as still-working.
|
||||
CompletableFuture<MessageService.Reply> answer =
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 200));
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 200, null));
|
||||
MessageService.Reply answered = answer.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answered.outcome(),
|
||||
"an answered worker that never replies times out as still working");
|
||||
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
|
||||
|
||||
MessageService.Reply probe = messages.send(T, "probe after timed-out-working", 300);
|
||||
MessageService.Reply probe = messages.send(T, "probe after timed-out-working", 300, null);
|
||||
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
|
||||
"the session lock must be released after a TIMED_OUT_WORKING answer(), or this bounded "
|
||||
+ "follow-up send would come back BUSY instead of timing out on its own work");
|
||||
@@ -753,7 +791,7 @@ class MessageServiceTest {
|
||||
new java.util.concurrent.atomic.AtomicReference<>();
|
||||
Thread answerer = new Thread(() -> {
|
||||
try {
|
||||
messages.answer(turnId, "config.yaml", 5000);
|
||||
messages.answer(turnId, "config.yaml", 5000, null);
|
||||
caught.set(new AssertionError("expected answer() to throw"));
|
||||
} catch (Throwable t) {
|
||||
caught.set(t);
|
||||
@@ -775,7 +813,7 @@ class MessageServiceTest {
|
||||
|
||||
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
|
||||
|
||||
MessageService.Reply probe = messages.send(T, "probe after exceptional failure", 300);
|
||||
MessageService.Reply probe = messages.send(T, "probe after exceptional failure", 300, null);
|
||||
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
|
||||
"the session lock must be released when answer()'s reply future fails exceptionally, "
|
||||
+ "or this bounded follow-up send would come back BUSY instead of timing out on its own work");
|
||||
@@ -802,7 +840,7 @@ class MessageServiceTest {
|
||||
new java.util.concurrent.atomic.AtomicReference<>();
|
||||
Thread answerer = new Thread(() -> {
|
||||
try {
|
||||
messages.answer(turnId, "config.yaml", 5000);
|
||||
messages.answer(turnId, "config.yaml", 5000, null);
|
||||
caught.set(new AssertionError("expected answer() to throw"));
|
||||
} catch (Throwable t) {
|
||||
caught.set(t);
|
||||
@@ -821,17 +859,121 @@ class MessageServiceTest {
|
||||
|
||||
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
|
||||
|
||||
MessageService.Reply probe = messages.send(T, "probe after interruption", 300);
|
||||
MessageService.Reply probe = messages.send(T, "probe after interruption", 300, null);
|
||||
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
|
||||
"the session lock must be released when answer()'s wait is interrupted, or this bounded "
|
||||
+ "follow-up send would come back BUSY instead of timing out on its own work");
|
||||
}
|
||||
|
||||
/**
|
||||
* Drive an async send on {@code T} to a resolved reply, polling as {@code owner} until the
|
||||
* ticket reports {@link MessageService.Phase#DONE} (or the 2s deadline runs out). {@code owner}
|
||||
* must be a terminal this ticket's creator check actually accepts, or this loops to the
|
||||
* deadline and returns a non-{@code DONE} view.
|
||||
*/
|
||||
private MessageService.TaskView driveAsyncTicketToDone(String ticket, String owner) throws InterruptedException {
|
||||
MessageService.TaskView view = null;
|
||||
long deadline = System.currentTimeMillis() + 2000;
|
||||
while (view == null || view.phase() != MessageService.Phase.DONE) {
|
||||
if (System.currentTimeMillis() >= deadline) break;
|
||||
view = messages.poll(ticket, owner);
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(5);
|
||||
}
|
||||
return view;
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollReturnsNullForAnUnknownTicket() {
|
||||
assertNull(messages.poll("task-999999"), "a ticket that was never minted is unknown");
|
||||
}
|
||||
|
||||
// --- fleetd #719: a per-boot nonce keeps one instance's ticket ids out of another's space ---
|
||||
|
||||
/** A second, fully independent instance — its own agents/injector/rendezvous/inbox, not shared. */
|
||||
private MessageService newIndependentInstance() {
|
||||
FakeHerdr otherHerdr = new FakeHerdr().readText("BUILD GREEN: 391 files");
|
||||
AgentControl otherAgents = new AgentControl(otherHerdr);
|
||||
Injector otherInjector = new Injector(otherAgents);
|
||||
return new MessageService(otherAgents, otherInjector, new Rendezvous(), new InMemoryReplyInbox());
|
||||
}
|
||||
|
||||
@Test
|
||||
void twoInstancesMintDisjointTicketIds() {
|
||||
MessageService other = newIndependentInstance();
|
||||
String ticketFromThis = messages.sendAsync(T, "task on first instance", null, null);
|
||||
String ticketFromOther = other.sendAsync(T, "task on second instance", null, null);
|
||||
assertNotEquals(ticketFromThis, ticketFromOther,
|
||||
"each instance mints its own id space, so even a first ticket from each must differ");
|
||||
}
|
||||
|
||||
@Test
|
||||
void foreignInstanceTicketDoesNotResolve() {
|
||||
MessageService other = newIndependentInstance();
|
||||
String ticket = messages.sendAsync(T, "task on first instance", null, null);
|
||||
// `other` must reach the same sequence number, or this test passes against an empty map
|
||||
// instead of against a colliding id.
|
||||
other.sendAsync(T, "task on second instance", null, null);
|
||||
|
||||
// control: the id resolves in the instance that minted it, so a null below cannot be
|
||||
// explained by broken plumbing — only by the ticket being foreign to `other`.
|
||||
assertNotNull(messages.poll(ticket), "the minting instance must still resolve its own ticket");
|
||||
assertNull(other.poll(ticket), "a ticket minted by a different instance must not resolve here");
|
||||
}
|
||||
|
||||
// --- fleetd #705: a ticket's creator terminal gates who may poll it -------------------------
|
||||
|
||||
@Test
|
||||
void pollByAnotherTerminalIsRefused() throws Exception {
|
||||
String ticket = messages.sendAsync(T, "long task", null, "term_a");
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker works
|
||||
assertTrue(rendezvous.resolve(T, "secret async result"), "a reply resolves the async send");
|
||||
|
||||
MessageService.TaskView owner = driveAsyncTicketToDone(ticket, "term_a");
|
||||
assertNotNull(owner, "the creator must still be able to read its own ticket");
|
||||
assertEquals(MessageService.Phase.DONE, owner.phase());
|
||||
|
||||
MessageService.TaskView refused = messages.poll(ticket, "term_b");
|
||||
assertNotNull(refused, "a different terminal gets a refusal, not silence");
|
||||
assertNotEquals(MessageService.Phase.DONE, refused.phase(),
|
||||
"a different terminal must never see the ticket as DONE");
|
||||
assertNull(refused.reply(), "a refusal must never carry the reply text");
|
||||
assertFalse(String.valueOf(refused).contains("secret async result"),
|
||||
"the reply text must not appear anywhere in the refused view");
|
||||
}
|
||||
|
||||
@Test
|
||||
void unnamedPrimaryStillReadsAnyTicket() throws Exception {
|
||||
String ticket = messages.sendAsync(T, "long task", null, "term_lead");
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker works
|
||||
assertTrue(rendezvous.resolve(T, "primary-visible result"), "a reply resolves the async send");
|
||||
|
||||
// callerTerminal == null is the unnamed primary (resolved by token or loopback trust, with
|
||||
// no herdr pane) — it must read a ticket a terminal-bearing lead created.
|
||||
MessageService.TaskView view = driveAsyncTicketToDone(ticket, null);
|
||||
assertNotNull(view, "the unnamed primary must be able to read any ticket");
|
||||
assertEquals(MessageService.Phase.DONE, view.phase());
|
||||
assertEquals("primary-visible result", view.reply());
|
||||
}
|
||||
|
||||
@Test
|
||||
void creatorReadsItsOwnTicket() throws Exception {
|
||||
String ticket = messages.sendAsync(T, "long task", null, "term_creator");
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE); // deliver
|
||||
injector.onStatus(T, AgentStatus.WORKING); // worker works
|
||||
assertTrue(rendezvous.resolve(T, "own result"), "a reply resolves the async send");
|
||||
|
||||
MessageService.TaskView view = driveAsyncTicketToDone(ticket, "term_creator");
|
||||
assertNotNull(view, "the ticket's own creator must be able to read it");
|
||||
assertEquals(MessageService.Phase.DONE, view.phase());
|
||||
assertEquals("own result", view.reply());
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollReportsACompletedTicket() throws Exception {
|
||||
String ticket = messages.sendAsync(T, "long task");
|
||||
@@ -858,11 +1000,11 @@ class MessageServiceTest {
|
||||
@Test
|
||||
void concurrentSendToSameSessionWhileFirstHoldsItIsBusy() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> first =
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "first", 5000));
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "first", 5000, null));
|
||||
awaitWaiting(); // the first send now holds the session lock, blocked on its reply
|
||||
|
||||
// A second send to the SAME session cannot take the lock within its short window.
|
||||
MessageService.Reply busy = messages.send(T, "second", 100);
|
||||
MessageService.Reply busy = messages.send(T, "second", 100, null);
|
||||
assertEquals(MessageService.Outcome.BUSY, busy.outcome(),
|
||||
"a second send while another holds the session is busy, not a hang");
|
||||
assertNull(busy.text());
|
||||
@@ -892,13 +1034,13 @@ class MessageServiceTest {
|
||||
PrimaryRegistry reg = new PrimaryRegistry(null);
|
||||
// L accepts a delegation to W: the send wins the lock and queues delivery → L is recorded.
|
||||
CompletableFuture<MessageService.Reply> first = CompletableFuture.supplyAsync(
|
||||
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L)));
|
||||
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L), null));
|
||||
awaitWaiting();
|
||||
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
|
||||
"an accepted send owns the delegation");
|
||||
|
||||
// A attempts W while L holds it → BUSY (lock never taken) → its hook never fires.
|
||||
MessageService.Reply busy = messages.send(T, "second", 100, () -> reg.recordDelegation(T, LEAD_A));
|
||||
MessageService.Reply busy = messages.send(T, "second", 100, () -> reg.recordDelegation(T, LEAD_A), null);
|
||||
assertEquals(MessageService.Outcome.BUSY, busy.outcome());
|
||||
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
|
||||
"a BUSY send must not steal the delegator ownership it never earned");
|
||||
@@ -919,7 +1061,7 @@ class MessageServiceTest {
|
||||
void anAcceptedSendAfterThePriorOwnerFinishesBecomesTheNewOwner() throws Exception {
|
||||
PrimaryRegistry reg = new PrimaryRegistry(null);
|
||||
CompletableFuture<MessageService.Reply> first = CompletableFuture.supplyAsync(
|
||||
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L)));
|
||||
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L), null));
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
injector.onStatus(T, AgentStatus.WORKING);
|
||||
@@ -930,7 +1072,7 @@ class MessageServiceTest {
|
||||
|
||||
// L finished; A's later accepted send takes the delegation over.
|
||||
CompletableFuture<MessageService.Reply> second = CompletableFuture.supplyAsync(
|
||||
() -> messages.send(T, "second", 5000, () -> reg.recordDelegation(T, LEAD_A)));
|
||||
() -> messages.send(T, "second", 5000, () -> reg.recordDelegation(T, LEAD_A), null));
|
||||
awaitWaiting();
|
||||
assertEquals(LEAD_A, reg.nudgeTargetFor(T).orElseThrow(),
|
||||
"an accepted send after the owner finished becomes the new delegator");
|
||||
@@ -951,7 +1093,7 @@ class MessageServiceTest {
|
||||
void answeringAnAskDoesNotRewriteDelegatorOwnership() throws Exception {
|
||||
PrimaryRegistry reg = new PrimaryRegistry(null);
|
||||
CompletableFuture<MessageService.Reply> send = CompletableFuture.supplyAsync(
|
||||
() -> messages.send(T, "do X", 5000, () -> reg.recordDelegation(T, LEAD_L)));
|
||||
() -> messages.send(T, "do X", 5000, () -> reg.recordDelegation(T, LEAD_L), null));
|
||||
awaitWaiting();
|
||||
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(), "L owns the delegation");
|
||||
|
||||
@@ -966,7 +1108,7 @@ class MessageServiceTest {
|
||||
|
||||
// L answers the ask on the same turn; the answer path must not touch ownership.
|
||||
CompletableFuture<MessageService.Reply> answer =
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting(); // the answering send reopened its forward waiter
|
||||
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
|
||||
@@ -986,7 +1128,7 @@ class MessageServiceTest {
|
||||
void aThrowingAcceptedHookLeavesNoStaleWaiterOrQueuedOrphan() {
|
||||
assertThrows(IllegalStateException.class,
|
||||
() -> messages.send(T, "doomed", 500,
|
||||
() -> { throw new IllegalStateException("ownership hook failed"); }),
|
||||
() -> { throw new IllegalStateException("ownership hook failed"); }, null),
|
||||
"a throwing ownership hook fails the send loudly");
|
||||
|
||||
assertFalse(rendezvous.isWaiting(T), "the failed send must not leave a stale rendezvous waiter");
|
||||
@@ -1114,7 +1256,7 @@ class MessageServiceTest {
|
||||
@Test
|
||||
void abandonFailsASendThatIsStillWaitingOnAReleasedSession() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send =
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000, null));
|
||||
awaitWaiting();
|
||||
|
||||
assertTrue(messages.abandon(T, "session released"), "a live waiter is abandoned");
|
||||
@@ -1134,7 +1276,7 @@ class MessageServiceTest {
|
||||
@Test
|
||||
void abandonDoesNotOverwriteAnAlreadyResolvedSend() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send =
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000, null));
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "the real answer"));
|
||||
|
||||
@@ -1265,7 +1407,7 @@ class MessageServiceTest {
|
||||
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase());
|
||||
|
||||
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
|
||||
() -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
@@ -1305,7 +1447,7 @@ class MessageServiceTest {
|
||||
// same turnId must see it as lapsed rather than resolving a question nobody is waiting on.
|
||||
assertEquals(MessageService.AskOutcome.TIMED_OUT, ask.get(5, TimeUnit.SECONDS).outcome());
|
||||
assertEquals(MessageService.Outcome.STALE_TURN,
|
||||
messages.answer(asking.turnId(), "config.yaml", 200).outcome());
|
||||
messages.answer(asking.turnId(), "config.yaml", 200, null).outcome());
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -1342,7 +1484,7 @@ class MessageServiceTest {
|
||||
|
||||
// The primary answers, but its own bounded wait for the worker's resumed turn is short and
|
||||
// expires before the worker (still genuinely working) gets back to it.
|
||||
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150);
|
||||
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150, null);
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome(),
|
||||
"the primary's own bounded wait gives up before the worker finishes resuming");
|
||||
@@ -1371,7 +1513,7 @@ class MessageServiceTest {
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
|
||||
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
|
||||
|
||||
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150);
|
||||
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150, null);
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome());
|
||||
|
||||
@@ -1425,7 +1567,7 @@ class MessageServiceTest {
|
||||
messages.setFinishAsyncTaskRaceHookForTest(() -> messages.forgetTurnForTest(turnId));
|
||||
|
||||
CompletableFuture<MessageService.Reply> answer =
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000));
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000, null));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting(); // answer() opened its own forward waiter for the resumed worker turn
|
||||
|
||||
@@ -1531,7 +1673,7 @@ class MessageServiceTest {
|
||||
String turnId = asking.turnId();
|
||||
|
||||
CompletableFuture<MessageService.Reply> answer =
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000));
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000, null));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer(),
|
||||
"the worker's own ask() call must have already unblocked with the primary's answer "
|
||||
+ "before we force the race below");
|
||||
@@ -1578,7 +1720,7 @@ class MessageServiceTest {
|
||||
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
|
||||
String turnId = asking.turnId();
|
||||
|
||||
MessageService.Reply answerReply = messages.answer(turnId, "config.yaml", 150);
|
||||
MessageService.Reply answerReply = messages.answer(turnId, "config.yaml", 150, null);
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome(),
|
||||
"the primary's own bounded wait must give up first, leaving turnId stamped with no "
|
||||
@@ -1644,7 +1786,7 @@ class MessageServiceTest {
|
||||
|
||||
MessageService.TaskView asking = messages.poll(first);
|
||||
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
|
||||
() -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
@@ -1678,7 +1820,7 @@ class MessageServiceTest {
|
||||
|
||||
// The primary answers it — answer() resumes the turn and blocks for what comes next.
|
||||
CompletableFuture<MessageService.Reply> answer1 = CompletableFuture.supplyAsync(
|
||||
() -> messages.answer(asking1.turnId(), "a1", 5000));
|
||||
() -> messages.answer(asking1.turnId(), "a1", 5000, null));
|
||||
assertEquals("a1", ask1.get(5, TimeUnit.SECONDS).answer());
|
||||
|
||||
// Still in the SAME resumed turn — before replying — the worker asks again.
|
||||
@@ -1699,7 +1841,7 @@ class MessageServiceTest {
|
||||
|
||||
// The primary answers the second question; the worker finally sends its real fleet_reply.
|
||||
CompletableFuture<MessageService.Reply> answer2 = CompletableFuture.supplyAsync(
|
||||
() -> messages.answer(turnId2, "a2", 5000));
|
||||
() -> messages.answer(turnId2, "a2", 5000, null));
|
||||
assertEquals("a2", ask2.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
@@ -1766,20 +1908,20 @@ class MessageServiceTest {
|
||||
}
|
||||
}
|
||||
|
||||
// --- CB-582: fleet_status pendingAsk() ------------------------------------------------------
|
||||
// --- fleet_status pendingAsk() -----------------------------------------------------------
|
||||
|
||||
@Test
|
||||
void pendingAskReturnsNullWhenNoQuestionIsOpen() throws Exception {
|
||||
assertNull(messages.pendingAsk(T), "no async ticket at all -> no pending ask");
|
||||
assertNull(messages.pendingAsk(T, null), "no async ticket at all -> no pending ask");
|
||||
|
||||
String ticket = messages.sendAsync(T, "long task");
|
||||
awaitWaiting();
|
||||
assertNull(messages.pendingAsk(T), "a plain pending delegation is not a question");
|
||||
assertNull(messages.pendingAsk(T, null), "a plain pending delegation is not a question");
|
||||
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
awaitTicketPhase(ticket, MessageService.Phase.DONE);
|
||||
assertNull(messages.pendingAsk(T), "a finished ticket carries no open question either");
|
||||
assertNull(messages.pendingAsk(T, null), "a finished ticket carries no open question either");
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -1792,21 +1934,151 @@ class MessageServiceTest {
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
|
||||
|
||||
MessageService.PendingAsk pending = messages.pendingAsk(T);
|
||||
MessageService.PendingAsk pending = messages.pendingAsk(T, null);
|
||||
assertNotNull(pending, "fleet_status should see the open question");
|
||||
assertEquals(ticket, pending.ticket());
|
||||
assertEquals("which config file?", pending.question());
|
||||
assertEquals(asking.turnId(), pending.turnId());
|
||||
|
||||
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
|
||||
() -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
assertNull(messages.pendingAsk(T), "an answered question is no longer pending");
|
||||
assertNull(messages.pendingAsk(T, null), "an answered question is no longer pending");
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
|
||||
}
|
||||
|
||||
/**
|
||||
* A caller's own terminal must match the terminal that created the delegation to see its
|
||||
* pending question; a different terminal-bearing caller sees nothing, and a caller with no
|
||||
* terminal at all (the unnamed primary) always sees it.
|
||||
*/
|
||||
@Test
|
||||
void pendingAskGatesTheQuestionByTheDelegationsCreatorTerminal() throws Exception {
|
||||
String ticket = messages.sendAsync(T, "task that asks", null, "term_creator");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
|
||||
|
||||
assertNull(messages.pendingAsk(T, "term_other"),
|
||||
"a caller whose terminal did not create the delegation must not see the question");
|
||||
|
||||
MessageService.PendingAsk own = messages.pendingAsk(T, "term_creator");
|
||||
assertNotNull(own, "the creating caller must see its own open question");
|
||||
assertEquals("which config file?", own.question());
|
||||
|
||||
MessageService.PendingAsk unnamed = messages.pendingAsk(T, null);
|
||||
assertNotNull(unnamed, "a caller with no terminal (the unnamed primary) must always see the question");
|
||||
assertEquals("which config file?", unnamed.question());
|
||||
|
||||
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||
() -> messages.answer(asking.turnId(), "config.yaml", 5000, "term_creator"));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
|
||||
}
|
||||
|
||||
/**
|
||||
* A task created with no recorded creator terminal (a short {@code sendAsync} overload) must
|
||||
* not hand its open question to any caller that does have a terminal — only a caller with no
|
||||
* terminal at all may still see it.
|
||||
*/
|
||||
@Test
|
||||
void pendingAskDeniesATerminalBearingCallerWhenTheTaskRecordsNoCreator() throws Exception {
|
||||
String ticket = messages.sendAsync(T, "task that asks"); // no creatorTerminal recorded
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
|
||||
|
||||
assertNull(messages.pendingAsk(T, "term_someone"),
|
||||
"a terminal-bearing caller must not see a question whose task records no creator");
|
||||
assertNotNull(messages.pendingAsk(T, null),
|
||||
"the unnamed primary must still see it even with no recorded creator");
|
||||
|
||||
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||
() -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
|
||||
}
|
||||
|
||||
// --- fleetd #715: answer() is gated on the caller that owns the turn -----------------------
|
||||
|
||||
/**
|
||||
* A blocking {@code fleet_send} from one caller opens the turn; a different caller's answer is
|
||||
* refused with no side effect on the rendezvous or the worker's blocked {@code fleet_ask} —
|
||||
* only the real owner can answer it.
|
||||
*/
|
||||
@Test
|
||||
void aDifferentCallersAnswerIsRefusedForABlockingSendDelegationButTheRealOwnerSucceeds() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = CompletableFuture.supplyAsync(
|
||||
() -> messages.send(T, "do X", 5000, "term_owner"));
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
|
||||
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
|
||||
assertFalse(rendezvous.isWaiting(T), "the forward waiter closes once the question surfaces");
|
||||
|
||||
// The hijack: a different caller answers the SAME turnId.
|
||||
MessageService.Reply hijacked = messages.answer(q.turnId(), "evil.yaml", 500, "term_attacker");
|
||||
assertEquals(MessageService.Outcome.NOT_TURN_OWNER, hijacked.outcome(),
|
||||
"a caller that did not open this turn must be refused, not served");
|
||||
assertFalse(rendezvous.isWaiting(T), "a refused answer must not open a forward waiter");
|
||||
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
|
||||
assertEquals(T, rendezvous.askSession(q.turnId()), "a refused answer must leave the ask turn open");
|
||||
|
||||
// The control: the real owner answers the same turnId and the turn resumes normally.
|
||||
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||
() -> messages.answer(q.turnId(), "config.yaml", 5000, "term_owner"));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
|
||||
}
|
||||
|
||||
/**
|
||||
* Same hijack and control as the blocking case, but the delegation is opened through the
|
||||
* fire-and-poll path ({@code sendAsync}) — the owner comes from the ticket's recorded creator
|
||||
* terminal, not a caller argument threaded through a live blocking call.
|
||||
*/
|
||||
@Test
|
||||
void aDifferentCallersAnswerIsRefusedForAnAsyncSendDelegationButTheRealOwnerSucceeds() throws Exception {
|
||||
String ticket = messages.sendAsync(T, "task that asks", null, "term_owner");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
|
||||
|
||||
MessageService.Reply hijacked = messages.answer(asking.turnId(), "evil.yaml", 500, "term_attacker");
|
||||
assertEquals(MessageService.Outcome.NOT_TURN_OWNER, hijacked.outcome(),
|
||||
"a caller that did not create this delegation must be refused, not served");
|
||||
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
|
||||
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase(),
|
||||
"a refused answer must not advance the async ticket's phase");
|
||||
|
||||
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||
() -> messages.answer(asking.turnId(), "config.yaml", 5000, "term_owner"));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
|
||||
assertEquals("done", awaitTicketPhase(ticket, MessageService.Phase.DONE).reply());
|
||||
}
|
||||
|
||||
// --- CB-588: async ticket terminal nudges ---------------------------------------------------
|
||||
//
|
||||
// MessageService.reply's rendezvous fast path is exactly what an async ticket always takes
|
||||
@@ -2063,7 +2335,7 @@ class MessageServiceTest {
|
||||
|
||||
// Clean up the still-open ask so the background thread does not linger past the test.
|
||||
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
|
||||
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000, null));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
@@ -2089,7 +2361,7 @@ class MessageServiceTest {
|
||||
.filter(c -> c.method().equals("agent.prompt")).count();
|
||||
|
||||
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
|
||||
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000, null));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting();
|
||||
java.util.concurrent.CountDownLatch terminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||
@@ -2482,7 +2754,7 @@ class MessageServiceTest {
|
||||
void hasQueuedDeliveryIsTrueAfterAnUndeliveredSendTimesOut() {
|
||||
// Nothing ever delivers the message (never goes IDLE/BLOCKED), so the send times out with
|
||||
// TIMED_OUT_QUEUED — same setup as sendTimesOutBeforeDeliveryIsQueuedNotWorking above.
|
||||
MessageService.Reply r = messages.send(T, "never delivered", 50);
|
||||
MessageService.Reply r = messages.send(T, "never delivered", 50, null);
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome());
|
||||
|
||||
assertTrue(messages.hasQueuedDelivery(T),
|
||||
@@ -2494,7 +2766,7 @@ class MessageServiceTest {
|
||||
|
||||
@Test
|
||||
void hasQueuedDeliveryClearsOnceTheTargetsNextDeliveryIsAccepted() throws Exception {
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50).outcome());
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50, null).outcome());
|
||||
assertTrue(messages.hasQueuedDelivery(T));
|
||||
|
||||
// A fresh send accepts delivery (opens its own waiter) — the stale queued fact is cleared.
|
||||
@@ -2511,7 +2783,7 @@ class MessageServiceTest {
|
||||
|
||||
@Test
|
||||
void hasQueuedDeliveryClearsOnAbandon() {
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50).outcome());
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50, null).outcome());
|
||||
assertTrue(messages.hasQueuedDelivery(T));
|
||||
|
||||
messages.abandon(T, "session released");
|
||||
|
||||
@@ -0,0 +1,187 @@
|
||||
package dev.ltms.fleet.msg;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.stream.Stream;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* Pins that no file under {@code src/main/java} calls the fail-open
|
||||
* {@link Rendezvous#open(String)} overload. That overload opens a forward waiter with no
|
||||
* recorded {@link Rendezvous.Owner}, so the turn it opens can never be answered by anyone,
|
||||
* not even the unnamed primary. Every production caller must go through
|
||||
* {@link Rendezvous#open(String, Rendezvous.Owner)} and record an explicit owner.
|
||||
*
|
||||
* <p>This reads each file's own source text rather than reflecting on compiled bytecode, because
|
||||
* the risk is a future one-word edit at a call site, not a missing overload.
|
||||
*
|
||||
* <p>The scan below finds a violation by its receiver, {@code rendezvous.open(}, classified by
|
||||
* argument count. The one test method here points that exact scanner at a file known to hold
|
||||
* many real one-argument calls before it ever looks at production, so a scanner that stops
|
||||
* matching fails loudly on the known-positive case instead of leaving a clean production result
|
||||
* looking like evidence it never produced.
|
||||
*/
|
||||
class RendezvousOpenUsageTest {
|
||||
|
||||
private static final Path PRODUCTION_SOURCE = Path.of("src/main/java");
|
||||
private static final Path KNOWN_TEST_CALLER =
|
||||
Path.of("src/test/java/dev/ltms/fleet/inject/CompletionResolverTest.java");
|
||||
|
||||
/**
|
||||
* A pattern that cannot find a known one-argument {@code rendezvous.open(} call would also
|
||||
* find none in production -- not because production is clean, but because the pattern does
|
||||
* not match the text it is supposed to catch. That failure mode is exactly what let a
|
||||
* {@code \b}-based regex read "no callers" under {@code git grep -E} when 81 real ones
|
||||
* existed: {@code git grep} does not treat {@code \b} as a word boundary, so the pattern
|
||||
* silently matched nothing anywhere, clean code and real calls alike. This test runs the
|
||||
* known-positive check first, with the same scanning method the production check then
|
||||
* depends on, so that mistake fails loudly here instead of reading as a clean result.
|
||||
*/
|
||||
@Test
|
||||
void noProductionFileCallsTheSingleArgumentOpenOverload() throws IOException {
|
||||
List<String> knownCalls = new ArrayList<>();
|
||||
int knownFilesScanned = scanForOneArgOpenCalls(KNOWN_TEST_CALLER, knownCalls);
|
||||
|
||||
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "found
|
||||
// nothing" having looked at nothing.
|
||||
assertTrue(knownFilesScanned > 0, "control failed: the scan under " + KNOWN_TEST_CALLER
|
||||
+ " visited zero .java files -- the path is wrong, so neither result below proves "
|
||||
+ "anything");
|
||||
|
||||
// CONTROL: the scanner actually finds real one-argument rendezvous.open( calls when
|
||||
// pointed at a file known to hold many. If this is not satisfied, the matching logic
|
||||
// itself is broken, and the production result below is the scanner failing silently,
|
||||
// not production code actually being clean.
|
||||
assertTrue(knownCalls.size() >= 60, "control failed: the scanner found only "
|
||||
+ knownCalls.size() + " one-argument rendezvous.open( call(s) in " + KNOWN_TEST_CALLER
|
||||
+ ", which is known to hold many -- the matching logic itself is broken: " + knownCalls);
|
||||
|
||||
List<String> violations = new ArrayList<>();
|
||||
int filesScanned = scanForOneArgOpenCalls(PRODUCTION_SOURCE, violations);
|
||||
|
||||
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "no
|
||||
// violations found" having looked at nothing.
|
||||
assertTrue(filesScanned > 0, "control failed: the scan under " + PRODUCTION_SOURCE
|
||||
+ " visited zero .java files -- the path is wrong, so the absence of violations "
|
||||
+ "below proves nothing");
|
||||
|
||||
assertTrue(violations.isEmpty(), "found a call to the fail-open Rendezvous.open(String) "
|
||||
+ "overload, which opens a forward waiter with no recorded owner -- record an "
|
||||
+ "explicit Rendezvous.Owner through open(String, Owner) instead: " + violations);
|
||||
}
|
||||
|
||||
private static int scanForOneArgOpenCalls(Path root, List<String> sites) throws IOException {
|
||||
List<Path> files = javaFiles(root);
|
||||
for (Path file : files) {
|
||||
scanFileForOpenCalls(file, sites);
|
||||
}
|
||||
return files.size();
|
||||
}
|
||||
|
||||
private static List<Path> javaFiles(Path root) throws IOException {
|
||||
try (Stream<Path> paths = Files.walk(root)) {
|
||||
return paths.filter(p -> p.toString().endsWith(".java")).toList();
|
||||
}
|
||||
}
|
||||
|
||||
private static void scanFileForOpenCalls(Path file, List<String> sites) throws IOException {
|
||||
String source = Files.readString(file);
|
||||
String needle = "rendezvous.open(";
|
||||
int from = 0;
|
||||
int idx;
|
||||
while ((idx = source.indexOf(needle, from)) >= 0) {
|
||||
int argsStart = idx + needle.length();
|
||||
String args = extractBalancedArgs(source, argsStart, file, idx);
|
||||
int closeParenIndex = argsStart + args.length();
|
||||
from = closeParenIndex + 1;
|
||||
|
||||
if (args.isBlank()) {
|
||||
continue; // Rendezvous has no zero-argument open() -- a javadoc "rendezvous.open()" mention, not a call
|
||||
}
|
||||
if (topLevelCommaCount(args) == 0) {
|
||||
sites.add(file + ":" + lineOf(source, idx) + " -- rendezvous.open(" + args.trim() + ")");
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The text between {@code rendezvous.open(} and its matching close paren: balanced over
|
||||
* nested calls, and never split by a paren or comma sitting inside a string or char literal.
|
||||
*/
|
||||
private static String extractBalancedArgs(String source, int start, Path file, int callIndex) {
|
||||
int depth = 1;
|
||||
boolean inString = false;
|
||||
boolean inChar = false;
|
||||
int i = start;
|
||||
while (i < source.length()) {
|
||||
char c = source.charAt(i);
|
||||
if (inString) {
|
||||
if (c == '\\') { i += 2; continue; }
|
||||
if (c == '"') inString = false;
|
||||
} else if (inChar) {
|
||||
if (c == '\\') { i += 2; continue; }
|
||||
if (c == '\'') inChar = false;
|
||||
} else if (c == '"') {
|
||||
inString = true;
|
||||
} else if (c == '\'') {
|
||||
inChar = true;
|
||||
} else if (c == '(') {
|
||||
depth++;
|
||||
} else if (c == ')') {
|
||||
depth--;
|
||||
if (depth == 0) return source.substring(start, i);
|
||||
}
|
||||
i++;
|
||||
}
|
||||
throw new IllegalStateException(
|
||||
"unbalanced parentheses scanning " + file + ":" + lineOf(source, callIndex));
|
||||
}
|
||||
|
||||
/**
|
||||
* Commas at paren/bracket/brace depth zero, skipping string and char literals -- the argument
|
||||
* separators a human reader would see, not every comma character in the text.
|
||||
*/
|
||||
private static int topLevelCommaCount(String args) {
|
||||
int depth = 0;
|
||||
int commas = 0;
|
||||
boolean inString = false;
|
||||
boolean inChar = false;
|
||||
int i = 0;
|
||||
while (i < args.length()) {
|
||||
char c = args.charAt(i);
|
||||
if (inString) {
|
||||
if (c == '\\') { i += 2; continue; }
|
||||
if (c == '"') inString = false;
|
||||
} else if (inChar) {
|
||||
if (c == '\\') { i += 2; continue; }
|
||||
if (c == '\'') inChar = false;
|
||||
} else if (c == '"') {
|
||||
inString = true;
|
||||
} else if (c == '\'') {
|
||||
inChar = true;
|
||||
} else if (c == '(' || c == '[' || c == '{') {
|
||||
depth++;
|
||||
} else if (c == ')' || c == ']' || c == '}') {
|
||||
depth--;
|
||||
} else if (c == ',' && depth == 0) {
|
||||
commas++;
|
||||
}
|
||||
i++;
|
||||
}
|
||||
return commas;
|
||||
}
|
||||
|
||||
private static int lineOf(String source, int index) {
|
||||
int line = 1;
|
||||
for (int i = 0; i < index; i++) {
|
||||
if (source.charAt(i) == '\n') line++;
|
||||
}
|
||||
return line;
|
||||
}
|
||||
}
|
||||
@@ -132,4 +132,127 @@ class RendezvousTest {
|
||||
assertNull(rendezvous.askSession(t.turnId()), "a closed ask is forgotten");
|
||||
assertFalse(rendezvous.answerAsk(t.turnId(), "late"), "a closed ask can no longer be answered");
|
||||
}
|
||||
|
||||
// ── fleetd #715: turn ownership ────────────────────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
void openWithNoOwnerRecordsNoOwnerAndOpenWithAnOwnerRecordsIt() {
|
||||
rendezvous.open(W);
|
||||
assertNull(rendezvous.ownerOf(W), "the no-owner overload records no owner at all");
|
||||
rendezvous.close(W, rendezvous.currentWaiter(W));
|
||||
|
||||
rendezvous.open(W, Rendezvous.Owner.of("term_lead"));
|
||||
assertEquals(Rendezvous.Owner.of("term_lead"), rendezvous.ownerOf(W));
|
||||
}
|
||||
|
||||
@Test
|
||||
void ownerOfIsNullWhenNoWaiterIsOpen() {
|
||||
assertNull(rendezvous.ownerOf(W), "no waiter open means no owner to report");
|
||||
}
|
||||
|
||||
@Test
|
||||
void openAskStampsTheForwardWaitersOwnerOntoTheFreshTurnOnly() {
|
||||
rendezvous.open(W, Rendezvous.Owner.of("term_lead"));
|
||||
|
||||
Rendezvous.AskTicket fresh = rendezvous.openAsk(W);
|
||||
assertTrue(fresh.fresh());
|
||||
assertEquals(Rendezvous.Owner.of("term_lead"), rendezvous.askOwner(fresh.turnId()),
|
||||
"a freshly-opened ask copies the forward waiter's current owner");
|
||||
|
||||
Rendezvous.AskTicket coalesced = rendezvous.openAsk(W);
|
||||
assertFalse(coalesced.fresh());
|
||||
assertEquals(fresh.turnId(), coalesced.turnId());
|
||||
assertEquals(Rendezvous.Owner.of("term_lead"), rendezvous.askOwner(coalesced.turnId()),
|
||||
"a coalesced duplicate ask rides the fresh owner's turn, unchanged");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aSecondAskAfterTheFirstClosesStampsWhateverOwnerIsOpenAtThatLaterMoment() {
|
||||
rendezvous.open(W, Rendezvous.Owner.of("term_lead"));
|
||||
Rendezvous.AskTicket first = rendezvous.openAsk(W);
|
||||
rendezvous.closeAsk(first.turnId());
|
||||
|
||||
// The forward waiter is reopened under a different owner before the second ask — mirrors
|
||||
// answer() reopening with the owner it already checked, which can differ turn to turn.
|
||||
rendezvous.close(W, rendezvous.currentWaiter(W));
|
||||
rendezvous.open(W, Rendezvous.Owner.of("term_other"));
|
||||
|
||||
Rendezvous.AskTicket second = rendezvous.openAsk(W);
|
||||
assertTrue(second.fresh());
|
||||
assertEquals(Rendezvous.Owner.of("term_other"), rendezvous.askOwner(second.turnId()),
|
||||
"a freshly-opened ask always copies whatever owner is open right now, not a stale one");
|
||||
}
|
||||
|
||||
@Test
|
||||
void askOwnerIsNullForAnUnknownOrLapsedTurn() {
|
||||
assertNull(rendezvous.askOwner("no-such#1"));
|
||||
Rendezvous.AskTicket t = rendezvous.openAsk(W);
|
||||
rendezvous.closeAsk(t.turnId());
|
||||
assertNull(rendezvous.askOwner(t.turnId()), "a closed ask no longer reports an owner");
|
||||
}
|
||||
|
||||
// ── fleetd #729: per-boot nonce guards turnId against cross-instance reuse ────────────────
|
||||
|
||||
@Test
|
||||
void twoInstancesMintDisjointTurnIds() {
|
||||
Rendezvous other = new Rendezvous();
|
||||
Rendezvous.AskTicket fromThis = rendezvous.openAsk(W);
|
||||
Rendezvous.AskTicket fromOther = other.openAsk(W);
|
||||
assertNotEquals(fromThis.turnId(), fromOther.turnId(),
|
||||
"each instance mints its own id space, so even a first ask from each must differ");
|
||||
}
|
||||
|
||||
@Test
|
||||
void foreignInstanceTurnIdDoesNotResolve() {
|
||||
Rendezvous other = new Rendezvous();
|
||||
|
||||
// `other` must reach the same sequence number as `rendezvous` (two asks each, the first
|
||||
// closed so the second mints fresh), or this test passes against an empty map instead of
|
||||
// against a colliding id.
|
||||
Rendezvous.AskTicket firstFromThis = rendezvous.openAsk(W);
|
||||
rendezvous.closeAsk(firstFromThis.turnId());
|
||||
Rendezvous.AskTicket secondFromThis = rendezvous.openAsk(W);
|
||||
|
||||
Rendezvous.AskTicket firstFromOther = other.openAsk(W);
|
||||
other.closeAsk(firstFromOther.turnId());
|
||||
other.openAsk(W);
|
||||
|
||||
// control: the id resolves in the instance that minted it, so a false below cannot be
|
||||
// explained by broken plumbing — only by the turnId being foreign to `other`.
|
||||
assertTrue(rendezvous.answerAsk(secondFromThis.turnId(), "answer from this instance"),
|
||||
"the minting instance must still resolve its own turnId");
|
||||
assertFalse(other.answerAsk(secondFromThis.turnId(), "answer from other instance"),
|
||||
"a turnId minted by a different instance must not resolve here");
|
||||
}
|
||||
|
||||
@Test
|
||||
void openAskStillCoalescesDuplicatesAndStillMintsDistinctIdsPerAsk() {
|
||||
Rendezvous.AskTicket t1 = rendezvous.openAsk(W);
|
||||
Rendezvous.AskTicket t2 = rendezvous.openAsk(W);
|
||||
assertEquals(t1.turnId(), t2.turnId(),
|
||||
"a second openAsk while one is open still coalesces onto the same turn");
|
||||
assertFalse(t2.fresh(), "the coalesced ask is still reported as not fresh");
|
||||
|
||||
rendezvous.closeAsk(t1.turnId());
|
||||
Rendezvous.AskTicket t3 = rendezvous.openAsk(W);
|
||||
assertNotEquals(t1.turnId(), t3.turnId(), "two asks from the same session still get different turnIds");
|
||||
}
|
||||
|
||||
@Test
|
||||
void ownerPermitsIsFailClosedOnARecordAndThreeStatesAreDistinct() {
|
||||
assertFalse(Rendezvous.Owner.permits(null, null),
|
||||
"no owner on record refuses even a caller with no terminal");
|
||||
assertFalse(Rendezvous.Owner.permits(null, "term_a"),
|
||||
"no owner on record refuses a terminal-bearing caller too");
|
||||
assertTrue(Rendezvous.Owner.permits(Rendezvous.Owner.UNNAMED_PRIMARY, null),
|
||||
"the unnamed primary owner matches a caller with no terminal");
|
||||
assertFalse(Rendezvous.Owner.permits(Rendezvous.Owner.UNNAMED_PRIMARY, "term_a"),
|
||||
"the unnamed primary owner does not match a terminal-bearing caller");
|
||||
assertTrue(Rendezvous.Owner.permits(Rendezvous.Owner.of("term_a"), "term_a"),
|
||||
"a named owner matches the same terminal");
|
||||
assertFalse(Rendezvous.Owner.permits(Rendezvous.Owner.of("term_a"), "term_b"),
|
||||
"a named owner refuses a different terminal");
|
||||
assertFalse(Rendezvous.Owner.permits(Rendezvous.Owner.of("term_a"), null),
|
||||
"a named owner refuses the unnamed primary");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,8 +1,14 @@
|
||||
package dev.ltms.fleet.rest;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.fleet.auth.CallerResolver;
|
||||
import dev.ltms.fleet.auth.Authz;
|
||||
import dev.ltms.fleet.auth.MemberRegistry;
|
||||
import dev.ltms.fleet.auth.Principal;
|
||||
import dev.ltms.fleet.auth.Role;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
@@ -15,9 +21,11 @@ import dev.ltms.fleet.metrics.FleetMetrics;
|
||||
import dev.ltms.fleet.metrics.Metrics;
|
||||
import dev.ltms.fleet.msg.MessageService;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.session.FakeWorktrees;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.testing.CapturedLog;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
@@ -28,10 +36,15 @@ import java.net.http.HttpRequest;
|
||||
import java.net.http.HttpResponse;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
import java.util.Locale;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.function.Predicate;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
import java.util.stream.Collectors;
|
||||
@@ -122,12 +135,465 @@ class FleetAppAuthTest {
|
||||
assertEquals(Authz.Action.DRAIN, FleetApp.routeAction("GET /sessions/{id}/replies"));
|
||||
assertEquals(Authz.Action.ASK, FleetApp.routeAction("POST /sessions/{id}/ask"));
|
||||
for (String route : Set.of("GET /sessions", "GET /agents", "GET /members", "GET /profiles",
|
||||
"GET /member-credentials", "GET /sessions/{id}/status", "GET /tasks/{ticket}")) {
|
||||
"GET /member-credentials")) {
|
||||
assertEquals(Authz.Action.READ, FleetApp.routeAction(route), route);
|
||||
}
|
||||
for (String route : Set.of("GET /sessions/{id}/status", "GET /tasks/{ticket}")) {
|
||||
assertEquals(Authz.Action.TASK_READ, FleetApp.routeAction(route), route);
|
||||
}
|
||||
assertThrows(IllegalArgumentException.class, () -> FleetApp.routeAction("GET /healthz"));
|
||||
}
|
||||
|
||||
// --- GET /tasks/{ticket} must not be the no-check overload ----------------------------------
|
||||
|
||||
/**
|
||||
* {@code GET /tasks/{ticket}} must resolve its caller the same way {@code allow(...)} does
|
||||
* and thread that terminal into {@link MessageService#poll(String, String)}, not the
|
||||
* no-check overload that ignores who is asking.
|
||||
*/
|
||||
@Test
|
||||
void theTaskStatusRouteActuallyThreadsTheCallersTerminalIntoPoll() throws Exception {
|
||||
String source = Files.readString(REST_SOURCE);
|
||||
|
||||
int start = source.indexOf("private void taskStatus(Context ctx) {");
|
||||
assertTrue(start >= 0, "could not find taskStatus in " + REST_SOURCE
|
||||
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
|
||||
int end = source.indexOf("private static void herdrError(Context ctx, HerdrException e) {", start);
|
||||
assertTrue(end > start, "could not find the method declared after taskStatus to bound the scrape");
|
||||
String handlerBlock = source.substring(start, end);
|
||||
|
||||
// CONTROL: the block we scraped really does call messages.poll(...) -- if this fails, the
|
||||
// anchors above moved and the assertions below would otherwise pass on nothing.
|
||||
assertTrue(handlerBlock.contains("messages.poll("),
|
||||
"control failed: the scraped taskStatus block contains no messages.poll( call at all "
|
||||
+ "-- the anchors have drifted, this test is not testing what it claims to");
|
||||
|
||||
assertTrue(handlerBlock.contains("caller.terminal()"),
|
||||
"the taskStatus route must thread the resolved caller's terminal into messages.poll(...), "
|
||||
+ "not the no-check overload -- block: " + handlerBlock);
|
||||
assertTrue(handlerBlock.contains("ctx.attribute(CALLER)"),
|
||||
"the taskStatus route must resolve its caller the same way allow(...) does, not via a "
|
||||
+ "second, separate resolution path -- block: " + handlerBlock);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code GET /sessions/{id}/status} must resolve its caller the same way {@code allow(...)}
|
||||
* does and thread that terminal into {@link MessageService#pendingAsk(String, String)}, not
|
||||
* the no-check overload that ignores who is asking.
|
||||
*/
|
||||
@Test
|
||||
void theSessionStatusRouteActuallyThreadsTheCallersTerminalIntoPendingAsk() throws Exception {
|
||||
String source = Files.readString(REST_SOURCE);
|
||||
|
||||
int start = source.indexOf("private void sessionStatus(Context ctx) {");
|
||||
assertTrue(start >= 0, "could not find sessionStatus in " + REST_SOURCE
|
||||
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
|
||||
int end = source.indexOf("private void taskStatus(Context ctx) {", start);
|
||||
assertTrue(end > start, "could not find the method declared after sessionStatus to bound the scrape");
|
||||
String handlerBlock = source.substring(start, end);
|
||||
|
||||
// CONTROL: the block we scraped really does call messages.pendingAsk(...) -- if this fails,
|
||||
// the anchors above moved and the assertions below would otherwise pass on nothing.
|
||||
assertTrue(handlerBlock.contains("messages.pendingAsk("),
|
||||
"control failed: the scraped sessionStatus block contains no messages.pendingAsk( "
|
||||
+ "call at all -- the anchors have drifted, this test is not testing what it claims to");
|
||||
|
||||
assertTrue(handlerBlock.contains("caller.terminal()"),
|
||||
"the sessionStatus route must thread the resolved caller's terminal into "
|
||||
+ "messages.pendingAsk(...), not the no-check overload -- block: " + handlerBlock);
|
||||
assertTrue(handlerBlock.contains("ctx.attribute(CALLER)"),
|
||||
"the sessionStatus route must resolve its caller the same way allow(...) does, not via "
|
||||
+ "a second, separate resolution path -- block: " + handlerBlock);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code GET /tasks/{ticket}} refuses a worker whose terminal did not create the ticket,
|
||||
* while the creating worker and the unnamed primary both still read it. The ticket is minted
|
||||
* directly on the shared {@link MessageService}, the same way {@code MessageServiceTest}
|
||||
* drives {@link MessageService#poll(String, String)}, so this exercises only the REST poll
|
||||
* route's own handling of the ownership already recorded on the ticket.
|
||||
*/
|
||||
@Test
|
||||
void restPollRefusesADifferentWorkerButAllowsTheCreatorAndTheUnnamedPrimary() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
Injector injector = new Injector(agents);
|
||||
MessageService messages = new MessageService(agents, injector, new Rendezvous());
|
||||
|
||||
Javalin creatorApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID); // -> term_a
|
||||
Javalin otherWorkerApp = startOnSharedService(messages, herdr, 9001L); // -> term_shell
|
||||
Javalin primaryApp = startOnSharedService(messages, herdr, 999_999L); // no pane -> primary
|
||||
try {
|
||||
String ticket = messages.sendAsync("term_a", "long task", null, "term_a");
|
||||
|
||||
HttpResponse<String> refused = send(otherWorkerApp.port(), "GET", "/tasks/" + ticket, null, null);
|
||||
assertEquals(200, refused.statusCode());
|
||||
assertTrue(refused.body().contains("forbidden"),
|
||||
"a different worker's terminal must be refused, not shown the ticket: " + refused.body());
|
||||
assertFalse(refused.body().contains("\"reply\""),
|
||||
"a refusal must never carry reply text: " + refused.body());
|
||||
|
||||
HttpResponse<String> own = send(creatorApp.port(), "GET", "/tasks/" + ticket, null, null);
|
||||
assertEquals(200, own.statusCode());
|
||||
assertFalse(own.body().contains("forbidden"),
|
||||
"the creating worker must read its own ticket: " + own.body());
|
||||
|
||||
HttpResponse<String> primary = send(primaryApp.port(), "GET", "/tasks/" + ticket, null, null);
|
||||
assertEquals(200, primary.statusCode());
|
||||
assertFalse(primary.body().contains("forbidden"),
|
||||
"the unnamed primary must read any ticket: " + primary.body());
|
||||
} finally {
|
||||
creatorApp.stop();
|
||||
otherWorkerApp.stop();
|
||||
primaryApp.stop();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code POST /sessions/{id}/message} with {@code wait:false} must record the creating
|
||||
* caller's own terminal on the ticket it returns, so that caller can still poll its own
|
||||
* ticket over REST, while a different terminal is refused.
|
||||
*/
|
||||
@Test
|
||||
void restSendAsyncRecordsTheCreatingCallersTerminalSoItCanStillPollItsOwnTicket() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
Injector injector = new Injector(agents);
|
||||
MessageService messages = new MessageService(agents, injector, new Rendezvous());
|
||||
|
||||
Javalin leadApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID, Map.of("term_a", "lead-x"));
|
||||
Javalin otherWorkerApp = startOnSharedService(messages, herdr, 9001L); // -> term_shell
|
||||
try {
|
||||
ObjectMapper mapper = new ObjectMapper();
|
||||
HttpResponse<String> created = send(leadApp.port(), "POST", "/sessions/term_a/message",
|
||||
"{\"content\":\"long task\",\"wait\":false}", null);
|
||||
assertEquals(202, created.statusCode(), created.body());
|
||||
String ticket = mapper.readTree(created.body()).path("ticket").asText(null);
|
||||
assertNotNull(ticket, "the accepted response carried no ticket: " + created.body());
|
||||
|
||||
HttpResponse<String> own = send(leadApp.port(), "GET", "/tasks/" + ticket, null, null);
|
||||
assertEquals(200, own.statusCode());
|
||||
assertFalse(own.body().contains("forbidden"),
|
||||
"the session that created the ticket over REST must be able to poll it: " + own.body());
|
||||
|
||||
HttpResponse<String> refused = send(otherWorkerApp.port(), "GET", "/tasks/" + ticket, null, null);
|
||||
assertEquals(200, refused.statusCode());
|
||||
assertTrue(refused.body().contains("forbidden: this ticket was created by a different session"),
|
||||
"a different terminal must still be refused with the ownership detail, not some "
|
||||
+ "other rejection: " + refused.body());
|
||||
} finally {
|
||||
leadApp.stop();
|
||||
otherWorkerApp.stop();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code GET /sessions/{id}/status} shows a worker's pending {@code fleet_ask} question, its
|
||||
* {@code turnId} and its ticket only to the caller whose terminal created that delegation, or
|
||||
* to a caller with no terminal at all (the unnamed primary) — a different terminal-bearing
|
||||
* caller still sees the base status line, but none of the pending-ask fields.
|
||||
*/
|
||||
@Test
|
||||
void restStatusGatesThePendingAskFieldsByTheDelegationsCreatorTerminal() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
Injector injector = new Injector(agents);
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
MessageService messages = new MessageService(agents, injector, rendezvous);
|
||||
|
||||
Javalin creatorApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID); // -> term_a
|
||||
Javalin otherWorkerApp = startOnSharedService(messages, herdr, 9001L); // -> term_shell
|
||||
Javalin primaryApp = startOnSharedService(messages, herdr, 999_999L); // no pane -> primary
|
||||
try {
|
||||
ObjectMapper mapper = new ObjectMapper();
|
||||
String ticket = messages.sendAsync("term_target", "task that asks", null, "term_a");
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_target"), "sendAsync should have opened its rendezvous waiter");
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
|
||||
() -> messages.ask("term_target", "which config file?", 5000));
|
||||
|
||||
MessageService.TaskView asking;
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
do {
|
||||
asking = messages.poll(ticket, null);
|
||||
Thread.sleep(5);
|
||||
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
|
||||
assertEquals(MessageService.Phase.ASKING, asking.phase());
|
||||
String turnId = asking.turnId();
|
||||
|
||||
JsonNode other = mapper.readTree(
|
||||
send(otherWorkerApp.port(), "GET", "/sessions/term_target/status", null, null).body());
|
||||
assertEquals("idle", other.get("status").asText(), "the base status must still be shown");
|
||||
assertFalse(other.has("question"), "a non-creating caller must not see the question: " + other);
|
||||
assertFalse(other.has("turnId"), "a non-creating caller must not see the turnId: " + other);
|
||||
assertFalse(other.has("ticket"), "a non-creating caller must not see the ticket: " + other);
|
||||
|
||||
JsonNode own = mapper.readTree(
|
||||
send(creatorApp.port(), "GET", "/sessions/term_target/status", null, null).body());
|
||||
assertEquals("which config file?", own.get("question").asText(), "the creator must see the question");
|
||||
assertEquals(turnId, own.get("turnId").asText(), "the creator must see the turnId");
|
||||
assertEquals(ticket, own.get("ticket").asText(), "the creator must see the ticket");
|
||||
|
||||
JsonNode primary = mapper.readTree(
|
||||
send(primaryApp.port(), "GET", "/sessions/term_target/status", null, null).body());
|
||||
assertEquals("which config file?", primary.get("question").asText(),
|
||||
"a caller with no terminal (the unnamed primary) must see the question");
|
||||
|
||||
// Clean up the still-open ask so the background thread does not linger past the test.
|
||||
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||
() -> messages.answer(turnId, "config.yaml", 5000, "term_a"));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.resolve("term_target", "done"));
|
||||
answer.get(5, TimeUnit.SECONDS);
|
||||
} finally {
|
||||
creatorApp.stop();
|
||||
otherWorkerApp.stop();
|
||||
primaryApp.stop();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code POST /sessions/{id}/message} with a {@code turnId} refuses a caller whose terminal
|
||||
* did not open the turn, even when that caller otherwise holds ANSWER rights, and leaves the
|
||||
* turn open for the real owner to resolve. Covers the REST adapter's ANSWER gate for an
|
||||
* async-send ({@code wait:false}) delegation.
|
||||
*/
|
||||
@Test
|
||||
void restAnswerIsRefusedForADifferentCallerOnAnAsyncSendDelegationButTheRealOwnerSucceeds() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
Injector injector = new Injector(agents);
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
MessageService messages = new MessageService(agents, injector, rendezvous);
|
||||
|
||||
Javalin ownerApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID, Map.of("term_a", "lead-owner"));
|
||||
Javalin attackerApp = startOnSharedService(messages, herdr, 9001L, Map.of("term_shell", "lead-attacker"));
|
||||
try {
|
||||
ObjectMapper mapper = new ObjectMapper();
|
||||
HttpResponse<String> created = send(ownerApp.port(), "POST", "/sessions/term_target/message",
|
||||
"{\"content\":\"long task\",\"wait\":false}", null);
|
||||
assertEquals(202, created.statusCode(), created.body());
|
||||
String ticket = mapper.readTree(created.body()).path("ticket").asText(null);
|
||||
assertNotNull(ticket, "the accepted response carried no ticket: " + created.body());
|
||||
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_target"), "the async send should have opened its rendezvous waiter");
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
|
||||
() -> messages.ask("term_target", "which config file?", 5000));
|
||||
|
||||
MessageService.TaskView asking;
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
do {
|
||||
asking = messages.poll(ticket, null);
|
||||
Thread.sleep(5);
|
||||
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
|
||||
assertEquals(MessageService.Phase.ASKING, asking.phase());
|
||||
String turnId = asking.turnId();
|
||||
|
||||
HttpResponse<String> hijacked = send(attackerApp.port(), "POST", "/sessions/term_target/message",
|
||||
"{\"content\":\"hijack\",\"turnId\":\"" + turnId + "\"}", null);
|
||||
assertEquals(403, hijacked.statusCode(), hijacked.body());
|
||||
assertTrue(hijacked.body().contains("not_turn_owner"),
|
||||
"a different caller's answer must be refused as not_turn_owner: " + hijacked.body());
|
||||
assertEquals("term_target", rendezvous.askSession(turnId),
|
||||
"a refused answer must leave the ask turn open");
|
||||
|
||||
CompletableFuture<HttpResponse<String>> owned = CompletableFuture.supplyAsync(() -> {
|
||||
try {
|
||||
return send(ownerApp.port(), "POST", "/sessions/term_target/message",
|
||||
"{\"content\":\"config.yaml\",\"turnId\":\"" + turnId + "\"}", null);
|
||||
} catch (Exception e) {
|
||||
throw new RuntimeException(e);
|
||||
}
|
||||
});
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.resolve("term_target", "done"));
|
||||
HttpResponse<String> ownedResponse = owned.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(200, ownedResponse.statusCode(), ownedResponse.body());
|
||||
assertEquals("done", mapper.readTree(ownedResponse.body()).path("reply").asText(null),
|
||||
"the real owner's answer must resolve the worker's turn");
|
||||
} finally {
|
||||
ownerApp.stop();
|
||||
attackerApp.stop();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, but the delegation is a blocking send ({@code wait:true}) instead of a ticket —
|
||||
* the owning caller's own HTTP request is the one that surfaces the worker's question and
|
||||
* later carries the real answer. Covers the REST adapter's ANSWER gate for a blocking-send
|
||||
* delegation.
|
||||
*/
|
||||
@Test
|
||||
void restAnswerIsRefusedForADifferentCallerOnABlockingSendDelegationButTheRealOwnerSucceeds() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
Injector injector = new Injector(agents);
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
MessageService messages = new MessageService(agents, injector, rendezvous);
|
||||
|
||||
Javalin ownerApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID, Map.of("term_a", "lead-owner"));
|
||||
Javalin attackerApp = startOnSharedService(messages, herdr, 9001L, Map.of("term_shell", "lead-attacker"));
|
||||
try {
|
||||
ObjectMapper mapper = new ObjectMapper();
|
||||
CompletableFuture<HttpResponse<String>> blocking = CompletableFuture.supplyAsync(() -> {
|
||||
try {
|
||||
return send(ownerApp.port(), "POST", "/sessions/term_target/message",
|
||||
"{\"content\":\"long task\",\"wait\":true,\"timeoutMs\":5000}", null);
|
||||
} catch (Exception e) {
|
||||
throw new RuntimeException(e);
|
||||
}
|
||||
});
|
||||
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_target"), "the blocking send should have opened its rendezvous waiter");
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
|
||||
() -> messages.ask("term_target", "which config file?", 5000));
|
||||
|
||||
HttpResponse<String> questionResponse = blocking.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(202, questionResponse.statusCode(), questionResponse.body());
|
||||
JsonNode question = mapper.readTree(questionResponse.body());
|
||||
assertEquals("question", question.get("status").asText());
|
||||
String turnId = question.get("turnId").asText();
|
||||
|
||||
HttpResponse<String> hijacked = send(attackerApp.port(), "POST", "/sessions/term_target/message",
|
||||
"{\"content\":\"hijack\",\"turnId\":\"" + turnId + "\"}", null);
|
||||
assertEquals(403, hijacked.statusCode(), hijacked.body());
|
||||
assertTrue(hijacked.body().contains("not_turn_owner"),
|
||||
"a different caller's answer must be refused as not_turn_owner: " + hijacked.body());
|
||||
assertEquals("term_target", rendezvous.askSession(turnId),
|
||||
"a refused answer must leave the ask turn open");
|
||||
|
||||
CompletableFuture<HttpResponse<String>> owned = CompletableFuture.supplyAsync(() -> {
|
||||
try {
|
||||
return send(ownerApp.port(), "POST", "/sessions/term_target/message",
|
||||
"{\"content\":\"config.yaml\",\"turnId\":\"" + turnId + "\"}", null);
|
||||
} catch (Exception e) {
|
||||
throw new RuntimeException(e);
|
||||
}
|
||||
});
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.resolve("term_target", "done"));
|
||||
HttpResponse<String> ownedResponse = owned.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(200, ownedResponse.statusCode(), ownedResponse.body());
|
||||
assertEquals("done", mapper.readTree(ownedResponse.body()).path("reply").asText(null),
|
||||
"the real owner's answer must resolve the worker's turn");
|
||||
} finally {
|
||||
ownerApp.stop();
|
||||
attackerApp.stop();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* As {@link #start}, but shares {@code messages} and {@code herdr} across several app
|
||||
* instances bound to different pids, each returned as its own started {@link Javalin} rather
|
||||
* than through the shared {@code app} field, so several differently-resolved callers can
|
||||
* poll the same ticket.
|
||||
*/
|
||||
private Javalin startOnSharedService(MessageService messages, FakeHerdr herdr, long pid) {
|
||||
return startOnSharedService(messages, herdr, pid, Map.of());
|
||||
}
|
||||
|
||||
/**
|
||||
* As {@link #startOnSharedService(MessageService, FakeHerdr, long)}, but {@code leadTerminals}
|
||||
* resolves the given pid's terminal to a named lead (a caller with SEND permission) instead of
|
||||
* a plain worker, for a test that needs a terminal-bearing caller able to create a ticket.
|
||||
*
|
||||
* <p>Every connecting pane not already claimed by {@code leadTerminals} is wired into the live
|
||||
* roster as a spawned worker, so a caller's resolved role matches what its own test expects:
|
||||
* a {@link Role#WORKER}, never the unconfigured-pane {@link Role#OBSERVER} floor a roster-less
|
||||
* resolver would otherwise fall to.
|
||||
*/
|
||||
private Javalin startOnSharedService(MessageService messages, FakeHerdr herdr, long pid,
|
||||
Map<String, String> leadTerminals) {
|
||||
FleetConfig.Profile wcfg = new FleetConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
|
||||
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
|
||||
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
|
||||
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
|
||||
k -> "FLEETD_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
|
||||
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
|
||||
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
|
||||
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
|
||||
() -> leadTerminals, new MemberRegistry(null),
|
||||
t -> leadTerminals.containsKey(t) ? null : MemberRole.DEV, Map::of);
|
||||
Metrics appMetrics = FleetMetrics.create(sessions, new dev.ltms.fleet.msg.InMemoryReplyInbox());
|
||||
|
||||
return new FleetApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
|
||||
callers, appMetrics).build().start("127.0.0.1", 0);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #669 Unit A: {@code POST /sessions/{id}/message} is two call shapes behind one route,
|
||||
* mirroring {@code fleet_send}'s MCP-side split into {@link Authz.Action#SEND} and {@link
|
||||
* Authz.Action#ANSWER} ({@code FleetMcp#sendAction}). The route never carries a {@code coordId}
|
||||
* shape — that peer-lead route is MCP-only — so only these two apply here.
|
||||
*/
|
||||
@Test
|
||||
void theMessageRouteIsASendWithNoTurnIdAndAnAnswerWithOne() {
|
||||
assertEquals(Authz.Action.SEND, FleetApp.routeAction("POST /sessions/{id}/message", null));
|
||||
assertEquals(Authz.Action.SEND, FleetApp.routeAction("POST /sessions/{id}/message", " "));
|
||||
assertEquals(Authz.Action.ANSWER, FleetApp.routeAction("POST /sessions/{id}/message", "turn-1"));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #689: {@code answerGatePasses} is the second, conditional gate behind {@code
|
||||
* sendMessage}'s coarse {@link Authz.Action#SEND} check. With the {@code ANSWER} grant denied,
|
||||
* a {@code turnId}-bearing request is refused while a plain one still passes — and the denied
|
||||
* permit is queried only for the {@code turnId} case, never for the plain one, which is what
|
||||
* proves this is a genuinely separate, conditional check rather than the {@code SEND} check
|
||||
* renamed or an unconditional call whose result is ignored. Flipping only the {@code ANSWER}
|
||||
* grant to allowed then flips only the {@code turnId} shape's outcome.
|
||||
*/
|
||||
@Test
|
||||
void answerGatePassesOnlyWhenTurnIdAbsentOrAnswerGranted() {
|
||||
List<Authz.Action> queried = new ArrayList<>();
|
||||
Predicate<Authz.Action> denyAnswer = action -> {
|
||||
queried.add(action);
|
||||
return false;
|
||||
};
|
||||
|
||||
assertFalse(FleetApp.answerGatePasses("turn-1", denyAnswer),
|
||||
"ANSWER denied ⇒ the turnId shape is refused");
|
||||
assertEquals(List.of(Authz.Action.ANSWER), queried,
|
||||
"the ANSWER grant, specifically, must be the one consulted");
|
||||
|
||||
queried.clear();
|
||||
assertTrue(FleetApp.answerGatePasses(null, denyAnswer),
|
||||
"no turnId ⇒ the plain shape passes even though ANSWER is denied");
|
||||
assertTrue(FleetApp.answerGatePasses(" ", denyAnswer), "a blank turnId is treated as absent");
|
||||
assertEquals(List.of(), queried, "the plain shape must never consult the permit at all");
|
||||
|
||||
assertTrue(FleetApp.answerGatePasses("turn-1", action -> true),
|
||||
"flipping only the ANSWER grant to allowed flips only the turnId shape's outcome");
|
||||
}
|
||||
|
||||
private static Set<String> routesTheServerRegisters() {
|
||||
try {
|
||||
String source = Files.readString(REST_SOURCE).lines()
|
||||
@@ -147,6 +613,74 @@ class FleetAppAuthTest {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code permitsFor} is the exact decision {@link FleetApp#allow} makes, passing the real
|
||||
* production classifier rather than a test-supplied one — built from an empty {@link
|
||||
* CallerResolver}, so no terminal is recognised as a configured lead or collaborator and a
|
||||
* collaborator's SEND is refused through the REST gate.
|
||||
*/
|
||||
@Test
|
||||
void aCollaboratorMayNotSendOverRestWithTheRealProductionClassifier() {
|
||||
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(new FakeHerdr()), _ -> 700L);
|
||||
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
|
||||
Map::of, new MemberRegistry(null));
|
||||
Principal collaborator = Principal.collaborator("ops", "term_collab", 700);
|
||||
assertFalse(FleetApp.permitsFor(collaborator, Authz.Action.SEND, "term_lead",
|
||||
callers.knownLeadOrCollaborator()),
|
||||
"no terminal is recognised as a lead or collaborator yet");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #669 Unit D: wires a real {@link CallerResolver} with a known lead and a known
|
||||
* collaborator tab, and a spawned member's own terminal recognised by neither map. A
|
||||
* collaborator's SEND reaches the known lead and the known collaborator, and is refused for the
|
||||
* spawned member's terminal — over the REST route, not just the unit-level classifier, so a
|
||||
* test covering only MCP cannot leave this route open.
|
||||
*/
|
||||
@Test
|
||||
void aCollaboratorMaySendToAKnownLeadOrCollaboratorButNotToASpawnedMembersTerminalOverRest() throws Exception {
|
||||
int port = startWithRealClassifier(FakeHerdr.WORKER_PID,
|
||||
Map.of("term_lead_known", "lead-x"), Map.of("term_a", "ops2"));
|
||||
|
||||
HttpResponse<String> toLead = send(port, "POST", "/sessions/term_lead_known/message",
|
||||
"{\"content\":\"hi\",\"wait\":false}", null);
|
||||
assertEquals(202, toLead.statusCode(), toLead.body());
|
||||
|
||||
HttpResponse<String> toSpawnedMembersTerminal = send(port, "POST", "/sessions/term_worker/message",
|
||||
"{\"content\":\"hi\",\"wait\":false}", null);
|
||||
assertEquals(403, toSpawnedMembersTerminal.statusCode(), toSpawnedMembersTerminal.body());
|
||||
}
|
||||
|
||||
/**
|
||||
* As {@link #start}, but with explicit lead/collaborator maps and no spawned-member roster, so
|
||||
* a test can wire the real {@link CallerResolver#knownLeadOrCollaborator()} classifier instead
|
||||
* of the default empty one.
|
||||
*/
|
||||
private int startWithRealClassifier(long pid, Map<String, String> leadTerminals,
|
||||
Map<String, String> collaboratorTerminals) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile wcfg = new FleetConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
|
||||
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
|
||||
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
|
||||
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
|
||||
k -> "FLEETD_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
|
||||
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
|
||||
Injector injector = new Injector(agents);
|
||||
MessageService messages = new MessageService(agents, injector, new Rendezvous());
|
||||
|
||||
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
|
||||
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
|
||||
() -> leadTerminals, new MemberRegistry(null), t -> null, () -> collaboratorTerminals);
|
||||
metrics = FleetMetrics.create(sessions, new dev.ltms.fleet.msg.InMemoryReplyInbox());
|
||||
|
||||
app = new FleetApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
|
||||
callers, metrics).build().start("127.0.0.1", 0);
|
||||
return app.port();
|
||||
}
|
||||
|
||||
// --- loopback-trust: the caller is the primary -------------------------------------------
|
||||
|
||||
@Test
|
||||
@@ -192,6 +726,107 @@ class FleetAppAuthTest {
|
||||
"draining an inbox is the primary's collection step");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #669 Unit A: the {@code turnId} shape of {@code POST /sessions/{id}/message} maps to
|
||||
* {@link Authz.Action#ANSWER}, not the plain {@link Authz.Action#SEND} the test above drives —
|
||||
* a worker must stay refused on this shape too, exactly as it was refused on the one undivided
|
||||
* action before the split.
|
||||
*/
|
||||
@Test
|
||||
void aWorkerMayNotAnswerAnotherSessionsBlockedQuestionOverRest() throws Exception {
|
||||
int port = start(FakeHerdr.WORKER_PID, false, null);
|
||||
|
||||
assertEquals(403, send(port, "POST", "/sessions/term_b/message",
|
||||
"{\"turnId\":\"turn-1\",\"content\":\"hi\"}", null).statusCode(),
|
||||
"resolving another session's blocked question would be a worker escalating too");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #689: a caller refused the coarse {@link Authz.Action#SEND} grant is refused on
|
||||
* {@code SEND} specifically, even on the {@code turnId}-bearing shape that otherwise raises
|
||||
* the check to {@link Authz.Action#ANSWER} — proving {@code turnId} was never read from the
|
||||
* body before the refusal (reading it would have changed which action is named in the 403).
|
||||
* The same caller refused with no body at all gets the identical detail, which could not hold
|
||||
* if the decision depended on anything read from the body. Control: a caller who IS granted
|
||||
* reaches past the gate and the body is used normally.
|
||||
*/
|
||||
@Test
|
||||
void aDeniedCallerIsRefusedOnSendEvenWithATurnIdBodyAndNeverReadsTheBody() throws Exception {
|
||||
int workerPort = start(FakeHerdr.WORKER_PID, false, null); // denied: not primary/architect
|
||||
|
||||
HttpResponse<String> withTurnId = send(workerPort, "POST", "/sessions/term_b/message",
|
||||
"{\"turnId\":\"turn-1\",\"content\":\"hi\"}", null);
|
||||
assertEquals(403, withTurnId.statusCode());
|
||||
assertTrue(withTurnId.body().contains("may not SEND"),
|
||||
"the SEND check must be the one that fired, not ANSWER — ANSWER would only be "
|
||||
+ "reachable by having already read turnId out of the body");
|
||||
|
||||
HttpResponse<String> noBody = send(workerPort, "POST", "/sessions/term_b/message", null, null);
|
||||
assertEquals(403, noBody.statusCode());
|
||||
assertTrue(noBody.body().contains("may not SEND"),
|
||||
"refused identically with no body at all — the refusal cannot depend on body content");
|
||||
|
||||
// Control: a primary IS granted SEND, so the same turnId body is read and acted on —
|
||||
// reaching messages.answer, which reports this unknown turnId as a stale one.
|
||||
int primaryPort = start(999_999, false, null);
|
||||
HttpResponse<String> granted = send(primaryPort, "POST", "/sessions/term_b/message",
|
||||
"{\"turnId\":\"turn-1\",\"content\":\"hi\"}", null);
|
||||
assertEquals(409, granted.statusCode());
|
||||
assertTrue(granted.body().contains("stale_turn"), "a granted caller's body IS read and acted on");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #689 (ticket comment 18353): the only place {@code sendMessage}'s call to {@code
|
||||
* answerGatePasses} is observable is the audit trail — {@code allow()} logs an {@code
|
||||
* "allowed"} entry for every granted action except {@code READ}/{@code METRICS}/{@code
|
||||
* TASK_READ}, and {@code ANSWER} is none of those. A granted {@code turnId} request must
|
||||
* therefore log both a {@code SEND} and an {@code ANSWER} entry; a granted plain request must
|
||||
* log {@code SEND} alone. A unit test of the extracted helper pins the helper; this pins the
|
||||
* call site — deleting the {@code answerGatePasses} call from {@code sendMessage} leaves the
|
||||
* helper's own test green but turns this one red.
|
||||
*/
|
||||
@Test
|
||||
void aGrantedTurnIdRequestAuditsBothSendAndAnswerButAPlainRequestAuditsSendAlone() throws Exception {
|
||||
int port = start(999_999, false, null); // primary: granted both SEND and ANSWER
|
||||
ObjectMapper mapper = new ObjectMapper();
|
||||
|
||||
try (CapturedLog audit = CapturedLog.at("audit", Level.INFO)) {
|
||||
send(port, "POST", "/sessions/term_b/message",
|
||||
"{\"turnId\":\"turn-1\",\"content\":\"hi\"}", null);
|
||||
|
||||
List<String> allowed = allowedActions(audit, mapper);
|
||||
assertTrue(allowed.contains("SEND"),
|
||||
"a turnId request must still clear the coarse SEND grant first");
|
||||
assertTrue(allowed.contains("ANSWER"),
|
||||
"a turnId request must ALSO clear the ANSWER grant — this is the call site itself");
|
||||
}
|
||||
|
||||
try (CapturedLog audit = CapturedLog.at("audit", Level.INFO)) {
|
||||
send(port, "POST", "/sessions/term_b/message",
|
||||
"{\"content\":\"hi\",\"timeoutMs\":50}", null);
|
||||
|
||||
List<String> allowed = allowedActions(audit, mapper);
|
||||
assertEquals(List.of("SEND"), allowed,
|
||||
"a plain request must log SEND and nothing else — ANSWER is conditional on "
|
||||
+ "turnId, not something every request happens to log");
|
||||
}
|
||||
}
|
||||
|
||||
private static List<String> allowedActions(CapturedLog audit, ObjectMapper mapper) {
|
||||
return audit.events().stream()
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.map(line -> {
|
||||
try {
|
||||
return mapper.readTree(line);
|
||||
} catch (Exception e) {
|
||||
throw new AssertionError("audit line is not valid JSON: " + line, e);
|
||||
}
|
||||
})
|
||||
.filter(n -> "allowed".equals(n.path("outcome").asText()))
|
||||
.map(n -> n.path("action").asText())
|
||||
.toList();
|
||||
}
|
||||
|
||||
// --- token mode ---------------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
|
||||
@@ -675,6 +675,26 @@ class FleetAppTest {
|
||||
assertEquals(400, postMessage(port, "{}").statusCode());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #689: a body that fails to parse is rejected with 400 before {@code turnId} is ever
|
||||
* read from it, so it reaches neither {@code messages.answer} (which needs a {@code turnId})
|
||||
* nor {@code messages.send} — confirmed here for {@code send} by the fake agent's idle status,
|
||||
* which would otherwise make an immediate {@code agent.prompt} delivery observable.
|
||||
*/
|
||||
@Test
|
||||
void malformedBodyReturns400AndNeverReachesSendOrAnswer() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // idle ⇒ send would deliver right away if reached
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
HttpResponse<String> res = postMessage(port, "not json at all");
|
||||
assertEquals(400, res.statusCode());
|
||||
JsonNode err = mapper.readTree(res.body());
|
||||
assertEquals("bad_request", err.get("error").asText());
|
||||
assertEquals("body must be JSON", err.get("detail").asText());
|
||||
assertFalse(herdr.called("agent.prompt"),
|
||||
"a malformed body must never reach messages.send's delivery");
|
||||
}
|
||||
|
||||
@Test
|
||||
void sessionStatusReportsLiveAgentStatus() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr().agentStatus("blocked");
|
||||
|
||||
@@ -0,0 +1,92 @@
|
||||
package dev.ltms.fleet.session;
|
||||
|
||||
import dev.ltms.fleet.inject.MemberPresence;
|
||||
import dev.ltms.fleet.peer.Capability;
|
||||
import dev.ltms.fleet.peer.PeerHandle;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Set;
|
||||
|
||||
/**
|
||||
* {@link PeerLauncher} decorator that marks presence for a spawned terminal before returning its
|
||||
* handle to the caller — the contact-then-register ordering fleetd #722 covers, where the
|
||||
* terminal's MCP contact lands before {@link SessionManager#acquire} runs its own
|
||||
* {@code registry.put}. The presence view is set after construction, via {@link #presence},
|
||||
* because it is owned by the {@link SessionManager} this launcher is passed into.
|
||||
*/
|
||||
final class PresenceRacingLauncher implements PeerLauncher {
|
||||
|
||||
private final PeerLauncher delegate;
|
||||
volatile MemberPresence presence;
|
||||
|
||||
PresenceRacingLauncher(PeerLauncher delegate) {
|
||||
this.delegate = delegate;
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req) {
|
||||
PeerHandle handle = delegate.spawn(req);
|
||||
presence.markPresent(handle.terminalId());
|
||||
return handle;
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
PeerHandle handle = delegate.spawn(req, decision);
|
||||
presence.markPresent(handle.terminalId());
|
||||
return handle;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<Capability> capabilities() {
|
||||
return delegate.capabilities();
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<Capability> capabilitiesFor(String profileName) {
|
||||
return delegate.capabilitiesFor(profileName);
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return delegate.profiles();
|
||||
}
|
||||
|
||||
@Override
|
||||
public String defaultProfile() {
|
||||
return delegate.defaultProfile();
|
||||
}
|
||||
|
||||
@Override
|
||||
public String effectiveCwd(SpawnRequest req) {
|
||||
return delegate.effectiveCwd(req);
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<String> parityOverlay(String profileName) {
|
||||
return delegate.parityOverlay(profileName);
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<?> list() {
|
||||
return delegate.list();
|
||||
}
|
||||
|
||||
@Override
|
||||
public int reapOrphanWorkers() {
|
||||
return delegate.reapOrphanWorkers();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void stop(String id) {
|
||||
delegate.stop(id);
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean clearContext(String id) {
|
||||
return delegate.clearContext(id);
|
||||
}
|
||||
}
|
||||
@@ -422,6 +422,53 @@ class SessionManagerTest {
|
||||
"turn completion moves BUSY → DONE");
|
||||
}
|
||||
|
||||
// --- fleetd #722: registration and presence must reach READY whichever lands first --------
|
||||
|
||||
@Test
|
||||
void registerThenContactReachesReadyForPlainSpawn() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
|
||||
sessions.asPresence().markPresent(session.terminalId());
|
||||
|
||||
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"a presence contact that arrives after registration reaches READY");
|
||||
}
|
||||
|
||||
@Test
|
||||
void contactThenRegisterStillReachesReadyForPlainSpawn() {
|
||||
// The racing launcher marks presence for the spawned terminal from inside spawn() —
|
||||
// before SessionManager.acquire's own registry.put runs — modeling an MCP contact that
|
||||
// lands in that window.
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile cfg = new FleetConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
PresenceRacingLauncher race = new PresenceRacingLauncher(workers);
|
||||
SessionManager sessions = new SessionManager(race);
|
||||
race.presence = sessions.asPresence();
|
||||
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
|
||||
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"a presence contact that lands before registry.put must still reach READY");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aTerminalNeverMarkedPresentStaysSpawningAfterRegistration() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
|
||||
assertEquals(MemberSession.State.SPAWNING, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"registration alone must not advance a terminal that was never marked present");
|
||||
}
|
||||
|
||||
@Test
|
||||
void releaseTearsDownWorkerAndRemovesFromRosterAndIsIdempotent() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
@@ -1405,6 +1452,105 @@ class SessionManagerTest {
|
||||
+ "dirty check threw");
|
||||
}
|
||||
|
||||
// --- fleetd #736: a release must forget the member's presence entry, not just its registry
|
||||
// row ---------------------------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
void releaseByPaneIdForgetsThePresenceEntry() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
assertTrue(sessions.asPresence().isPresent(terminal), "present before the release");
|
||||
|
||||
sessions.release(session.paneId());
|
||||
|
||||
assertFalse(sessions.asPresence().isPresent(terminal),
|
||||
"release must forget the terminal's presence, not just remove its registry row");
|
||||
}
|
||||
|
||||
@Test
|
||||
void reapIdleForgetsThePresenceEntryToo() {
|
||||
long[] clock = {0};
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
assertTrue(sessions.asPresence().isPresent(terminal), "present before the reap");
|
||||
|
||||
clock[0] = 11;
|
||||
assertEquals(1, sessions.reapIdle(10), "READY session past TTL is reaped");
|
||||
|
||||
assertFalse(sessions.asPresence().isPresent(terminal),
|
||||
"the idle-reap release path (releaseIfCurrent) goes through the same teardown "
|
||||
+ "funnel as an explicit release, so it must forget presence too");
|
||||
}
|
||||
|
||||
@Test
|
||||
void shutdownDrainAlsoForgetsThePresenceEntry() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
assertTrue(sessions.asPresence().isPresent(terminal), "present before the drain");
|
||||
|
||||
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
|
||||
|
||||
assertFalse(sessions.asPresence().isPresent(terminal),
|
||||
"a shutdown drain still ends the member's process, so presence must be cleared "
|
||||
+ "exactly as it is for any other release cause");
|
||||
}
|
||||
|
||||
@Test
|
||||
void releaseOfAnUnknownPaneIdDoesNotThrow() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
|
||||
assertDoesNotThrow(() -> sessions.release("no-such-pane"),
|
||||
"releasing a pane id that was never registered must be a no-op, not a throw");
|
||||
}
|
||||
|
||||
@Test
|
||||
void releaseStillForgetsPresenceWhenDirtyCheckThrows() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
RecordingWorktrees worktrees = new RecordingWorktrees();
|
||||
SessionManager sessions = sessionManager(herdr, worktrees);
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("fleetd-736", null));
|
||||
String terminal = s.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
worktrees.failHasUncommittedWith(new WorktreeException("git status exited 128"));
|
||||
|
||||
assertDoesNotThrow(() -> sessions.release(s.paneId()),
|
||||
"a throwing dirty check must not abort the release");
|
||||
|
||||
assertFalse(sessions.asPresence().isPresent(terminal),
|
||||
"presence must be forgotten even when the dirty check throws, which pins the "
|
||||
+ "forget call to the finally block that runs no matter what happened above");
|
||||
}
|
||||
|
||||
@Test
|
||||
void releaseLeavesADifferentStillLiveMembersPresenceUntouched() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
MemberSession released = sessions.acquire("ltms-local", null, "/caller/a", "ownerA");
|
||||
MemberSession stillLive = sessions.acquire("ltms-local", null, "/caller/b", "ownerB");
|
||||
sessions.asPresence().markPresent(released.terminalId());
|
||||
sessions.asPresence().markPresent(stillLive.terminalId());
|
||||
assertTrue(sessions.asPresence().isPresent(stillLive.terminalId()),
|
||||
"present before the release of the other member");
|
||||
|
||||
sessions.release(released.paneId());
|
||||
|
||||
assertFalse(sessions.asPresence().isPresent(released.terminalId()),
|
||||
"the released terminal is forgotten");
|
||||
assertTrue(sessions.asPresence().isPresent(stillLive.terminalId()),
|
||||
"a still-live member's presence must survive an unrelated release");
|
||||
}
|
||||
|
||||
// --- fleetd #316: the dirty check must be re-taken after the worker is stopped, not trusted
|
||||
// stale from before it ------------------------------------------------------------------------
|
||||
|
||||
@@ -2436,4 +2582,160 @@ class SessionManagerTest {
|
||||
return "threw:" + e.getClass().getName() + ":" + e.getMessage();
|
||||
}
|
||||
}
|
||||
|
||||
// --- fleetd #702: spawnedMemberRole must still answer for a pane mid-teardown ----------------
|
||||
|
||||
/**
|
||||
* A {@link Worktrees} test double whose {@code hasUncommitted} runs an injected hook before
|
||||
* answering. This is what lets a test resolve a releasing pane's terminal from inside the
|
||||
* window {@link SessionManager#release} opens between removing the registry entry and
|
||||
* actually stopping the pane — the hook runs synchronously on the release call's own thread,
|
||||
* at the exact point {@code release} shells out to {@code git status}, so there is no sleep
|
||||
* and no race to land in it.
|
||||
*/
|
||||
private static final class HookedWorktrees implements Worktrees {
|
||||
private Runnable hook;
|
||||
private boolean dirty = false;
|
||||
|
||||
HookedWorktrees onHasUncommitted(Runnable hook) {
|
||||
this.hook = hook;
|
||||
return this;
|
||||
}
|
||||
|
||||
@Override
|
||||
public String add(String repoRoot, String branch, String baseRef) {
|
||||
return "/wt/" + branch.replace('/', '_');
|
||||
}
|
||||
|
||||
@Override
|
||||
public void remove(String repoRoot, String worktreePath) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void deleteBranch(String repoRoot, String branch) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean hasUncommitted(String worktreePath) {
|
||||
if (hook != null) {
|
||||
hook.run();
|
||||
}
|
||||
return dirty;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public String repoRoot(String cwd) {
|
||||
return "/repo";
|
||||
}
|
||||
|
||||
@Override
|
||||
public java.util.Optional<String> snapshot(String worktreePath, String branch, String message) {
|
||||
return java.util.Optional.empty();
|
||||
}
|
||||
|
||||
@Override
|
||||
public WipRefStats wipRefs(String repoRoot) {
|
||||
return new WipRefStats(0, 0L);
|
||||
}
|
||||
|
||||
@Override
|
||||
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
|
||||
return 0;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void shareWithGroup(String repoRoot, String worktreePath) {
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnedMemberRoleResolvesTheLiveRegisteredRole() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller", null);
|
||||
|
||||
assertEquals(s.role(), sessions.spawnedMemberRole(s.terminalId()),
|
||||
"a registered session resolves to its own role");
|
||||
assertNull(sessions.spawnedMemberRole("term_unknown"),
|
||||
"a terminal with no session at all resolves to null");
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnedMemberRoleIsNullOnceReleaseFullyCompletes() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, new HookedWorktrees());
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("fleetd-702a", null));
|
||||
|
||||
sessions.release(s.paneId());
|
||||
|
||||
assertNull(sessions.spawnedMemberRole(s.terminalId()),
|
||||
"once release has fully finished, the terminal is neither registered nor releasing");
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnedMemberRoleStillAnswersBetweenTheRegistryRemovalAndThePaneStop() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
java.util.concurrent.atomic.AtomicReference<MemberRole> duringWindow = new java.util.concurrent.atomic.AtomicReference<>();
|
||||
HookedWorktrees worktrees = new HookedWorktrees();
|
||||
SessionManager sessions = sessionManager(herdr, worktrees);
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("fleetd-702b", null));
|
||||
worktrees.onHasUncommitted(() -> {
|
||||
// This runs from INSIDE release()'s git-status shell-out: the registry entry is
|
||||
// already gone, but the pane has not stopped yet — a real call landing inside the
|
||||
// exact window fleetd #702 reports, so no sleep and no race is needed to reach it.
|
||||
assertTrue(sessions.get(s.paneId()).isEmpty(),
|
||||
"sanity: the registry entry is already gone at this point");
|
||||
duringWindow.set(sessions.spawnedMemberRole(s.terminalId()));
|
||||
});
|
||||
|
||||
sessions.release(s.paneId());
|
||||
|
||||
assertEquals(s.role(), duringWindow.get(),
|
||||
"spawnedMemberRole must still answer the live role while the pane is mid-teardown, "
|
||||
+ "not only while the session is still in the registry");
|
||||
assertNull(sessions.spawnedMemberRole(s.terminalId()),
|
||||
"and once release has fully finished, the window is closed too");
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code Releasing.enter} and {@code Releasing.leave} are called directly here, never through
|
||||
* {@link SessionManager#release} or {@link SessionManager#spawnedMemberRole}, so a test naming
|
||||
* one of them exercises only that one — a regression in the other can never hide behind it.
|
||||
*/
|
||||
@Test
|
||||
void releasingLeaveStepsDownADepthGreaterThanOneInsteadOfRemovingIt() {
|
||||
SessionManager.Releasing depthTwo = new SessionManager.Releasing(2, "term_a", MemberRole.DEV);
|
||||
|
||||
SessionManager.Releasing afterLeave = depthTwo.leave();
|
||||
|
||||
assertNotNull(afterLeave,
|
||||
"depth 2 means another release of the SAME pane is still mid-teardown; leave() must "
|
||||
+ "step the depth down, never remove the marker outright — removing it here "
|
||||
+ "is what a plain Set would do, and would reopen the window the still-in-"
|
||||
+ "flight release is relying on staying closed");
|
||||
assertEquals(1, afterLeave.depth());
|
||||
assertEquals("term_a", afterLeave.terminalId());
|
||||
assertEquals(MemberRole.DEV, afterLeave.role());
|
||||
}
|
||||
|
||||
@Test
|
||||
void releasingEnterPreservesThePriorTerminalWhenTheOverlappingCallHasNoSessionOfItsOwn() {
|
||||
SessionManager.Releasing prior = new SessionManager.Releasing(1, "term_a", MemberRole.DEV);
|
||||
|
||||
SessionManager.Releasing afterEnter = SessionManager.Releasing.enter(prior, null);
|
||||
|
||||
assertEquals(2, afterEnter.depth(), "depth still increments whether or not this enter knows its session");
|
||||
assertEquals("term_a", afterEnter.terminalId(),
|
||||
"an overlapping release that finds the registry entry already gone has no session of "
|
||||
+ "its own to pass as known, and must not blank out the terminal the first "
|
||||
+ "call already recorded — that terminal is what spawnedMemberRole matches "
|
||||
+ "against");
|
||||
assertEquals(MemberRole.DEV, afterEnter.role());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -159,6 +159,40 @@ class WorktreeSessionManagerTest {
|
||||
assertEquals(expectedPath, s.cwd(), "session cwd is the worktree path");
|
||||
}
|
||||
|
||||
// --- fleetd #722: registration and presence must reach READY whichever lands first --------
|
||||
|
||||
@Test
|
||||
void registerThenContactReachesReadyForWorktreeSpawn() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
|
||||
new WorktreeRequest("cb-722", null));
|
||||
sessions.asPresence().markPresent(session.terminalId());
|
||||
|
||||
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"a presence contact that arrives after worktree registration reaches READY");
|
||||
}
|
||||
|
||||
@Test
|
||||
void contactThenRegisterStillReachesReadyForWorktreeSpawn() {
|
||||
// The racing launcher marks presence for the spawned terminal from inside spawn() —
|
||||
// before SessionManager.acquireWithWorktree's own registry.put runs — modeling an MCP
|
||||
// contact that lands in that window.
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
PresenceRacingLauncher race = new PresenceRacingLauncher(workerService(herdr));
|
||||
SessionManager sessions = new SessionManager(race, worktrees);
|
||||
race.presence = sessions.asPresence();
|
||||
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
|
||||
new WorktreeRequest("cb-722", null));
|
||||
|
||||
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"a presence contact that lands before worktree registration must still reach READY");
|
||||
}
|
||||
|
||||
@Test
|
||||
void worktreeArchitectAcquireAlsoBindsItsSlot() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
+22
-5
@@ -213,6 +213,17 @@ map_masked_lines() {
|
||||
done < "$file"
|
||||
}
|
||||
|
||||
# Masks every `scheme://user:pass@host` userinfo on one line of text, replacing just that
|
||||
# userinfo with `<redacted>` and leaving the rest of the line untouched, byte for byte. The
|
||||
# pattern stops at the first `/`, whitespace, or `@` reached after `://` — a URI's userinfo
|
||||
# component cannot contain any of those three characters — so a URI with no userinfo, followed
|
||||
# later on the same line by an unrelated `@`, never matches. The `g` flag matters: a line can
|
||||
# carry more than one URI. Shared by every caller that prints a line which may hold a
|
||||
# credentialed URI, so the bound lives in exactly one place.
|
||||
mask_url_userinfo() {
|
||||
printf '%s\n' "$1" | sed -E 's#://[^@/[:space:]]*@#://<redacted>@#g'
|
||||
}
|
||||
|
||||
redact() {
|
||||
local old_file="$1" new_file="$2"
|
||||
local line prefix content indent lead key old_line=0 new_line=0 in_hunk=0
|
||||
@@ -270,7 +281,7 @@ redact() {
|
||||
continue
|
||||
fi
|
||||
fi
|
||||
printf '%s\n' "$line" | sed -E 's#://[^@]*@#://<redacted>@#g'
|
||||
mask_url_userinfo "$line"
|
||||
done
|
||||
[ "$saved_nocasematch" = 1 ] || shopt -u nocasematch
|
||||
}
|
||||
@@ -580,6 +591,12 @@ install_candidate() {
|
||||
|
||||
# -------------------------------------------------------------------------------- the report path
|
||||
#
|
||||
# Masks basic-auth userinfo (scheme://user:pass@host) in a daemon verdict line before it reaches
|
||||
# the terminal.
|
||||
mask_verdict_userinfo() {
|
||||
mask_url_userinfo "$1"
|
||||
}
|
||||
|
||||
# Prints the literal command the operator (or a test) can run to restore the backup by hand — the
|
||||
# absolute path to THIS script plus the overrides actually in force, so it works from any cwd.
|
||||
restore_command_line() {
|
||||
@@ -599,8 +616,8 @@ restore_and_confirm() {
|
||||
ok "restored from $backup"
|
||||
if wait_for_verdict "$LOG" "$mark2" "$WAIT_SECONDS"; then
|
||||
case "$VERDICT_KIND" in
|
||||
refused) warn "the RESTORE was also refused by the daemon: $VERDICT_LINE" ;;
|
||||
*) ok "restore confirmed: $VERDICT_LINE" ;;
|
||||
refused) warn "the RESTORE was also refused by the daemon: $(mask_verdict_userinfo "$VERDICT_LINE")" ;;
|
||||
*) ok "restore confirmed: $(mask_verdict_userinfo "$VERDICT_LINE")" ;;
|
||||
esac
|
||||
else
|
||||
warn "the restore is on disk, but no confirming verdict line appeared within ${WAIT_SECONDS}s"
|
||||
@@ -616,7 +633,7 @@ report_outcome() {
|
||||
|
||||
say "waiting for the daemon's verdict (up to ${WAIT_SECONDS}s)"
|
||||
if wait_for_verdict "$LOG" "$mark" "$WAIT_SECONDS"; then
|
||||
kind="$VERDICT_KIND"; line="$VERDICT_LINE"
|
||||
kind="$VERDICT_KIND"; line="$(mask_verdict_userinfo "$VERDICT_LINE")"
|
||||
else
|
||||
kind="none"
|
||||
fi
|
||||
@@ -673,7 +690,7 @@ check_mode() {
|
||||
local verdict
|
||||
verdict="$(last_verdict_line "$LOG")"
|
||||
if [ -n "$verdict" ]; then
|
||||
ok "last verdict in log: $verdict"
|
||||
ok "last verdict in log: $(mask_verdict_userinfo "$verdict")"
|
||||
else
|
||||
warn "no reload verdict line found in $LOG"
|
||||
fi
|
||||
|
||||
+131
-66
@@ -71,20 +71,22 @@ set -euo pipefail
|
||||
|
||||
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
MODULE="$REPO/fleetd"
|
||||
JAR="$MODULE/target/fleetd.jar"
|
||||
# fleetd #493: never build into the path a running process holds. The build writes here first
|
||||
# (Maven's shade plugin has finalName=fleetd, so `clean install` still lands its output at
|
||||
# target/fleetd.jar — that part is unchanged and out of this script's control), but this script
|
||||
# now moves it out to JAR_STAGED immediately, and only swaps it back to JAR (a plain `mv`, so a
|
||||
# rename, never a byte-by-byte overwrite) after the OLD daemon has been confirmed exited. See
|
||||
# stage_built_jar/swap_staged_jar below.
|
||||
JAR_STAGED="$MODULE/target/fleetd-new.jar"
|
||||
# fleetd #664: the runtime path and Maven's output path are no longer the same file. Maven's
|
||||
# shade plugin (finalName=fleetd) always lands a fresh build at target/fleetd.jar — that is
|
||||
# Maven's own output directory and this script does not change it — but the daemon is launched
|
||||
# from $JAR instead, outside target/ entirely. That split is the whole fix: neither `mvn install`
|
||||
# nor `mvn clean` can ever reach the file a running daemon holds open, because that file no
|
||||
# longer lives under target/ at all. See swap_if_built/swap_staged_jar below for the one `mv`
|
||||
# that moves a build from one path to the other, and only after the old daemon is confirmed gone.
|
||||
BUILD_JAR="$MODULE/target/fleetd.jar"
|
||||
JAR="$MODULE/run/fleetd.jar"
|
||||
OUT="$MODULE/fleetd.out"
|
||||
# Matches BOTH the absolute form and the relative `java -jar target/fleetd.jar` a hand-start
|
||||
# produces from inside fleetd/. Anchoring on the absolute path alone was a real bug: the daemon
|
||||
# restarted correctly and the script still reported "no process appeared", because it launched with
|
||||
# a relative path and then looked for an absolute one.
|
||||
PATTERN='target/fleetd.jar'
|
||||
# Matches a fleetd daemon's command line wherever its jar sits — absolute or relative, under
|
||||
# run/, under target/, or anywhere else a build or a hand-start might point it. Detecting a
|
||||
# daemon this script did not start, including one running from a jar outside $JAR's own
|
||||
# directory, is this pattern's whole job; running_pid()'s `comm = java` allowlist below is what
|
||||
# keeps that breadth from counting a shell that merely types the pattern as literal text.
|
||||
PATTERN='fleetd.jar'
|
||||
HEALTH='http://127.0.0.1:8765/healthz'
|
||||
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
|
||||
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start — fleetd #603: also the pid-
|
||||
@@ -169,10 +171,10 @@ hash256() {
|
||||
fi
|
||||
}
|
||||
|
||||
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the STAGED
|
||||
# jar right after a build (before it has been swapped in) without ever changing what a bare
|
||||
# `jar_id` (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line
|
||||
# both call it with no args on purpose, so neither can ever be fooled by a leftover staged file.
|
||||
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the BUILT
|
||||
# jar at $BUILD_JAR (before it has been swapped in) without ever changing what a bare `jar_id`
|
||||
# (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line both call
|
||||
# it with no args on purpose, so neither can ever be fooled by a leftover build output.
|
||||
# fleetd #550 — THREE distinct answers now, not two: `[ -f "$f" ]` already separates "the jar is
|
||||
# not there" (-> "absent") from "the jar is there"; for the second case, hash256 itself separates
|
||||
# "hashed it" (a 12-char hex string) from "could not hash it" (-> "unhashable", when no hasher is
|
||||
@@ -180,15 +182,35 @@ hash256() {
|
||||
# was the whole defect this ticket fixes.
|
||||
jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
|
||||
|
||||
# fleetd #664 — under the old layout $JAR and the build output were the same file, so "jar on
|
||||
# disk" was one fact. Now they are two: $BUILD_JAR (target/fleetd.jar, whatever Maven last wrote,
|
||||
# by this script or by a bare `mvn install` run by hand) and $JAR (run/fleetd.jar, whatever the
|
||||
# daemon actually has open). Printing one label for both was the trap this ticket exists to close
|
||||
# — during the incident it would have shown the NEW jar's hash while the JVM ran the OLD one.
|
||||
# Pure (reads jar_id/date, never mutates), so a test can call it directly without reaching the
|
||||
# main flow — the same shape swap_if_built/drain_gate_refusal already use.
|
||||
report_jar_state() {
|
||||
local built_hash running_hash built_mtime running_mtime
|
||||
built_hash="$(jar_id "$BUILD_JAR")"
|
||||
running_hash="$(jar_id "$JAR")"
|
||||
built_mtime="$([ -f "$BUILD_JAR" ] && date -r "$BUILD_JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none')"
|
||||
running_mtime="$([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none')"
|
||||
ok "built jar (target/fleetd.jar): $built_hash ($built_mtime)"
|
||||
ok "running jar (run/fleetd.jar): $running_hash ($running_mtime)"
|
||||
if [ "$built_hash" != "absent" ] && [ "$running_hash" != "absent" ] && [ "$built_hash" != "$running_hash" ]; then
|
||||
warn "built jar and running jar differ — target/fleetd.jar was rebuilt since the running daemon last started and is not yet live"
|
||||
fi
|
||||
}
|
||||
|
||||
# fleetd #593 — `pgrep -f "$PATTERN"` matches ANY process whose full command line CONTAINS the
|
||||
# pattern text, and that is not the same thing as "is the daemon". A shell that merely embeds the
|
||||
# pattern as literal text — a human typing this exact investigation by hand, an ssh-shaped
|
||||
# `sh -c '...; ...'`, a pipeline, or any other non-exec'ing shell that never replaced itself with
|
||||
# the pattern-holding command — still shows up in that match, and it is the INSTRUMENT, not the
|
||||
# daemon. Measured live on this Mac: `sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 30' &`
|
||||
# daemon. Measured live on this Mac: `sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 30' &`
|
||||
# leaves a real `sh` process alive (it forks for the `sleep`, it does not exec into it) whose own
|
||||
# `ps -o args` is `sh -c echo "target/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
|
||||
# matches that line right alongside the real `java -jar target/fleetd.jar` process. `pgrep -c`
|
||||
# `ps -o args` is `sh -c echo "run/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
|
||||
# matches that line right alongside the real `java -jar run/fleetd.jar` process. `pgrep -c`
|
||||
# (an in-one-call count) does not exist on BSD/macOS at all, so this cannot be fixed by switching
|
||||
# pgrep flags — it has to filter what pgrep already found, after the fact, in a way that still
|
||||
# runs on BSD.
|
||||
@@ -210,7 +232,7 @@ jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
|
||||
# launched as `java -jar ...` — a native image, a renamed launcher — `running_pid()` silently
|
||||
# returns nothing and `assert_single_daemon` stops noticing a second daemon at all. For a guard,
|
||||
# that false-negative direction is the worse one to be wrong in. This is not a new assumption,
|
||||
# though: `PATTERN='target/fleetd.jar'` two lines up already assumes the daemon is a jar, which
|
||||
# though: `PATTERN='fleetd.jar'` two lines up already assumes the daemon is a jar, which
|
||||
# is only ever run by `java`. If that launch method changes, `PATTERN` stops matching anything
|
||||
# before this allowlist would ever get the chance to be wrong — the allowlist rides on the same
|
||||
# assumption that is already load-bearing, it does not add a new one. Whoever changes the launch
|
||||
@@ -229,16 +251,12 @@ running_pid() {
|
||||
printf '%s' "$out"
|
||||
}
|
||||
|
||||
# fleetd #493 — three small, independently testable pieces of "never build into the path a
|
||||
# running process holds":
|
||||
# fleetd #493/#664 — the independently testable pieces of "never build into the path a running
|
||||
# process holds":
|
||||
#
|
||||
# stage_built_jar moves the jar Maven just produced OUT of the live path and onto the staging
|
||||
# path, immediately after a successful build. Dies (leaving the OLD daemon
|
||||
# untouched — this runs before the stop step) if Maven reported success but
|
||||
# left no jar behind, or if the move itself fails.
|
||||
# require_no_build_jar the --no-build path never builds or stages anything: it must find a
|
||||
# jar already sitting at the live path from an earlier successful run, and
|
||||
# die with the same truthful message this script has always used if not.
|
||||
# require_no_build_jar the --no-build path never builds anything: it must find a jar already
|
||||
# sitting at the live path ($JAR, under run/) from an earlier successful run,
|
||||
# and die with the same truthful message this script has always used if not.
|
||||
# wait_for_daemon_exit polls running_pid() for up to $1 seconds and reports whether the OLD
|
||||
# daemon actually exited — extracted to its own function so the main flow
|
||||
# can be relied on to call swap_staged_jar only AFTER this returns success,
|
||||
@@ -248,14 +266,10 @@ running_pid() {
|
||||
# so this is never a write into a path a running process holds — by the time
|
||||
# it runs, nothing holds that path anymore. If it fails, the caller must not
|
||||
# start a new daemon: die() below already refuses that by exiting the script.
|
||||
stage_built_jar() {
|
||||
[ -f "$JAR" ] || die "build succeeded but produced no jar at $JAR — cannot stage it for restart.
|
||||
The running daemon was NOT touched."
|
||||
mv -f "$JAR" "$JAR_STAGED" \
|
||||
|| die "could not move the freshly built jar from $JAR to the staging path $JAR_STAGED.
|
||||
The running daemon was NOT touched."
|
||||
}
|
||||
|
||||
# fleetd #664: the "staged" jar swap_if_built passes in is now $BUILD_JAR
|
||||
# itself (target/fleetd.jar, Maven's own output) — a build no longer needs to
|
||||
# be moved off the live path right after compiling, because target/ was never
|
||||
# the live path to begin with.
|
||||
require_no_build_jar() {
|
||||
[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build"
|
||||
}
|
||||
@@ -282,8 +296,7 @@ swap_staged_jar() {
|
||||
#
|
||||
# The defect: the swap step used to be guarded inline by `if [ "$DO_BUILD" = 1 ]` in the main flow.
|
||||
# Changing that to `if false` left the suite green and the swap never ran, so a redeploy reported
|
||||
# every step succeeding while the daemon started on no jar at all (stage_built_jar has already moved
|
||||
# the freshly built one to $JAR_STAGED by then) or on a stale one.
|
||||
# every step succeeding while the daemon started on no jar at all or on a stale one.
|
||||
# test_swap_ordered_after_wait_and_before_start could not catch it: it reads this script's own text
|
||||
# and compares line positions, and a same-line edit moves no line.
|
||||
#
|
||||
@@ -312,7 +325,12 @@ swap_if_built() {
|
||||
local do_build="$1"
|
||||
should_swap "$do_build" || return 0
|
||||
say "swap"
|
||||
swap_staged_jar "$JAR_STAGED" "$JAR"
|
||||
# fleetd #664: $JAR now lives under run/, a directory target/ never created. mkdir -p here,
|
||||
# not inside swap_staged_jar itself — that function's own contract is tested on a missing
|
||||
# parent directory (a failing mv), and widening it to auto-create one would change what that
|
||||
# test proves.
|
||||
mkdir -p "$(dirname "$JAR")"
|
||||
swap_staged_jar "$BUILD_JAR" "$JAR"
|
||||
ok "jar in place: $(jar_id)"
|
||||
}
|
||||
|
||||
@@ -610,6 +628,45 @@ check_log_path_matches_plist() {
|
||||
ok "log path check: script and plist agree ($resolved_out)"
|
||||
}
|
||||
|
||||
# Reads the launchd plist's ProgramArguments for the argument that follows "-jar", resolves it
|
||||
# alongside $jar_path, and dies when the two differ. Call it only when the agent is loaded; it
|
||||
# never touches launchd or the daemon itself.
|
||||
check_jar_path_matches_plist() {
|
||||
local jar_path="$1" plist_path="$2"
|
||||
local plist_args plist_jar resolved_jar resolved_plist_jar
|
||||
if ! plist_args="$(/usr/libexec/PlistBuddy -c 'Print :ProgramArguments' "$plist_path" 2>/dev/null)"; then
|
||||
die "launchd agent is loaded but PlistBuddy could not read ProgramArguments from
|
||||
$plist_path
|
||||
— cannot verify which jar the supervised daemon launches. Fix the plist before
|
||||
redeploying supervised."
|
||||
fi
|
||||
plist_jar="$(printf '%s\n' "$plist_args" | awk '
|
||||
{ gsub(/^[ \t]+|[ \t]+$/, "") }
|
||||
prev == "-jar" { print; exit }
|
||||
{ prev = $0 }
|
||||
')"
|
||||
if [ -z "$plist_jar" ]; then
|
||||
die "launchd agent is loaded but its ProgramArguments at
|
||||
$plist_path
|
||||
do not contain a '-jar <path>' pair — cannot verify which jar the supervised daemon
|
||||
launches. Fix the plist before redeploying supervised."
|
||||
fi
|
||||
resolved_jar="$(cd "$(dirname "$jar_path")" 2>/dev/null && pwd -P)/$(basename "$jar_path")" || true
|
||||
resolved_plist_jar="$(cd "$(dirname "$plist_jar")" 2>/dev/null && pwd -P)/$(basename "$plist_jar")" || true
|
||||
if [ -z "$resolved_jar" ] || [ -z "$resolved_plist_jar" ] || [ "$resolved_jar" != "$resolved_plist_jar" ]; then
|
||||
die "jar path mismatch — this script deploys to
|
||||
$jar_path (resolved: ${resolved_jar:-<directory does not exist>})
|
||||
but the loaded plist's ProgramArguments names
|
||||
$plist_jar (resolved: ${resolved_plist_jar:-<directory does not exist>})
|
||||
The swap renames the built jar into place, so the old path stops existing after a redeploy;
|
||||
a launchd-initiated start from this plist (a reboot, or KeepAlive after a crash) would then
|
||||
run java against a missing file. Reinstall the plist at
|
||||
$plist_path
|
||||
so its ProgramArguments names $jar_path before redeploying supervised."
|
||||
fi
|
||||
ok "jar path check: script and plist agree ($resolved_jar)"
|
||||
}
|
||||
|
||||
# fleetd #552: the post-restart fresh-log capture, pulled out of the main flow so it is testable by
|
||||
# sourcing (the same reason systemd_installed/systemd_loaded above guard their OWN mktemp inline
|
||||
# instead of leaving it bare) even though its only caller sits below the SOURCED guard. By the time
|
||||
@@ -826,16 +883,25 @@ report_shutdown_drain() {
|
||||
# --no-build, staged jar present -> ALSO "nothing changed", deliberately: --no-build itself builds
|
||||
# and stages nothing (see require_no_build_jar above), so a staged jar found here is a leftover
|
||||
# from an earlier, unrelated run. THIS run truly changed nothing, and the next DO_BUILD=1 run
|
||||
# wipes that leftover before it builds (`rm -f "$JAR_STAGED"` in the build section above) — so
|
||||
# there is nothing here for the operator to lose track of.
|
||||
# wipes that leftover for free — `mvn clean` deletes all of target/, $BUILD_JAR included,
|
||||
# before the build even starts — so there is nothing here for the operator to lose track of.
|
||||
# --no-build, staged jar absent -> "nothing changed"
|
||||
drain_gate_refusal() {
|
||||
local do_build="$1" staged_path="$2"
|
||||
if [ "$do_build" = 1 ] && [ -f "$staged_path" ]; then
|
||||
# fleetd #664: under the old layout a build emptied the live path ($JAR) immediately, so
|
||||
# --no-build's own check ("no jar at $JAR") was the thing that refused a rerun here. That is
|
||||
# no longer true: a build never touches $JAR at all now, so $JAR still holds whatever was
|
||||
# already running before this gate fired (reaching this message at all requires OLD_PID to
|
||||
# have been set, which means a daemon was running from $JAR already) — a --no-build rerun
|
||||
# would NOT refuse, it would just restart that same old jar and silently throw away the one
|
||||
# sitting at %s.
|
||||
printf 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at
|
||||
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart —
|
||||
the freshly built jar is no longer at the live path that --no-build requires — or
|
||||
remove %s by hand if you want to discard this build.' "$staged_path" "$JAR" "$staged_path"
|
||||
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart — a
|
||||
--no-build rerun would NOT refuse here: %s already exists from before this run, so it
|
||||
would restart the daemon on that OLD jar and silently discard the one you just built — or
|
||||
remove %s by hand if you want to discard this build instead.' \
|
||||
"$staged_path" "$JAR" "$JAR" "$staged_path"
|
||||
else
|
||||
printf 'aborted — nothing changed'
|
||||
fi
|
||||
@@ -934,6 +1000,7 @@ report_supervisor_state() {
|
||||
SUPERVISED=1
|
||||
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
|
||||
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
|
||||
check_jar_path_matches_plist "$JAR" "$LAUNCHD_PLIST"
|
||||
;;
|
||||
systemd)
|
||||
SUPERVISED=1
|
||||
@@ -1170,7 +1237,7 @@ if [ -n "$OLD_PID" ]; then
|
||||
else
|
||||
warn "no daemon running — this will be a cold start"
|
||||
fi
|
||||
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
|
||||
report_jar_state
|
||||
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
|
||||
|
||||
# CB-594 / fleetd #492: supervision state. Installed and loaded are different facts — a
|
||||
@@ -1252,9 +1319,6 @@ stop_if_check_only "$CHECK_ONLY"
|
||||
|
||||
if [ "$DO_BUILD" = 1 ]; then
|
||||
say "build"
|
||||
# fleetd #493: wipe a leftover staged jar from a previous failed/interrupted run BEFORE doing
|
||||
# anything else, so that run's leftovers can never be mistaken for this run's output.
|
||||
rm -f "$JAR_STAGED"
|
||||
BUILD_LOG="$(mktemp -t fleetd-build.XXXXXX)"
|
||||
echo " log: $BUILD_LOG"
|
||||
if ! mvn -f "$MODULE/pom.xml" clean install > "$BUILD_LOG" 2>&1; then
|
||||
@@ -1264,16 +1328,16 @@ if [ "$DO_BUILD" = 1 ]; then
|
||||
fi
|
||||
grep -E '^\[INFO\] Tests run:.*Failures' "$BUILD_LOG" | tail -1 | sed 's/^\[INFO\] / /' || true
|
||||
ok "BUILD SUCCESS"
|
||||
# fleetd #493: move the freshly built jar off the live path immediately — the running (OLD)
|
||||
# daemon, if any, is still up at this point (build always runs before stop). From here until the
|
||||
# swap step below (after the OLD daemon is confirmed gone), $JAR_STAGED is the only artefact this
|
||||
# script treats as "the new jar" — $JAR itself is not touched again until the swap.
|
||||
stage_built_jar
|
||||
ok "jar now: $(jar_id "$JAR_STAGED")"
|
||||
# fleetd #664: nothing to stage — $BUILD_JAR (target/fleetd.jar) is Maven's own output path and
|
||||
# was never the live path, so the running (OLD) daemon, if any, was never at risk from this build
|
||||
# at all. From here until the swap step below (after the OLD daemon is confirmed gone),
|
||||
# $BUILD_JAR is the artefact this script treats as "the new jar" — $JAR itself is not touched
|
||||
# again until the swap.
|
||||
ok "jar now: $(jar_id "$BUILD_JAR")"
|
||||
else
|
||||
say "build skipped (--no-build)"
|
||||
# fleetd #493: --no-build never builds or stages anything — it restarts whatever jar is already
|
||||
# sitting at the live path from an earlier successful run. Same check, same message as before.
|
||||
# fleetd #493: --no-build never builds anything — it restarts whatever jar is already sitting at
|
||||
# the live path from an earlier successful run. Same check, same message as before.
|
||||
require_no_build_jar
|
||||
fi
|
||||
|
||||
@@ -1282,8 +1346,8 @@ fi
|
||||
# fleetd #555: run_drain_gate above is drain_gate_required + the prompt + drain_confirmed, called
|
||||
# unconditionally — it returns immediately when the gate is not required, and composes/dies through
|
||||
# refuse_drain_gate itself when the reply does not confirm. See #493/#517/#528 for why "nothing
|
||||
# changed" would be a lie once a build has staged a jar.
|
||||
run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"
|
||||
# changed" would be a lie once a build has produced a jar not yet swapped in.
|
||||
run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"
|
||||
|
||||
# ------------------------------------------------------------------ stop
|
||||
#
|
||||
@@ -1334,21 +1398,22 @@ fi
|
||||
|
||||
# ------------------------------------------------------------------ swap
|
||||
#
|
||||
# fleetd #493: every branch above has now either confirmed the OLD daemon actually exited
|
||||
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only
|
||||
# NOW is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
|
||||
# launchd/systemd's ExecStart) will read from — this mv is the one and only write to $JAR anywhere
|
||||
# fleetd #493/#664: every branch above has now either confirmed the OLD daemon actually exited
|
||||
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only NOW
|
||||
# is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
|
||||
# launchd/systemd's ExecStart) will read from — a rename from $BUILD_JAR (target/) to $JAR (run/),
|
||||
# both under $MODULE and so on one filesystem. This mv is the one and only write to $JAR anywhere
|
||||
# in this script's mutating flow. If it fails, do not start: die() below exits before "start" runs.
|
||||
swap_if_built "$DO_BUILD"
|
||||
|
||||
# ------------------------------------------------------------------ start
|
||||
# Unsupervised: login shell (zsh -l) is what puts the secrets on the daemon's environment, and cwd
|
||||
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and target/ relative to it.
|
||||
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and run/ relative to it.
|
||||
# Supervised (launchd): launchd does both — deploy/dev.ltms.fleetd.plist points ProgramArguments at
|
||||
# scripts/fleetd-launchd-wrapper.sh (CB-594), which is what execs the login shell in launchd's
|
||||
# place, and WorkingDirectory in the plist already pins fleetd/.
|
||||
# Supervised (systemd --user): the unit does both too — measured on the second host, ExecStart is
|
||||
# `/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` (a login shell, same reason as
|
||||
# `/bin/zsh -lc "exec java -jar run/fleetd.jar fleetd.yaml"` (a login shell, same reason as
|
||||
# above) and WorkingDirectory is already pinned to fleetd/.
|
||||
|
||||
say "start"
|
||||
|
||||
@@ -233,6 +233,66 @@ test_redaction_holds() {
|
||||
assert_contains "weight" "$RUN_OUTPUT" "a diff must have been demonstrably printed at all"
|
||||
}
|
||||
|
||||
# redact()'s key-name filter only inspects the KEY, so a diff line whose key does not match
|
||||
# TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY still reaches the final userinfo
|
||||
# sed even when its VALUE holds a credentialed URI. "note" is not a sensitive key name, so this
|
||||
# line must fall all the way through to that sed, not the earlier whole-value branch. The
|
||||
# trailing prose on both sides of the userinfo is a positive control: it proves the line reached
|
||||
# the userinfo sed (which touches only the userinfo) rather than the earlier branch (which would
|
||||
# have replaced the whole value with a bare "<redacted>" and dropped the prose).
|
||||
test_diff_line_userinfo_is_masked_with_positive_control() {
|
||||
local dir
|
||||
dir="$(new_fixture)"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.note=see amqp://alice:wonderland@rabbit.local:5672/vhost for details'
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 0 "$RUN_RC" "diff-userinfo-case reload exit code"
|
||||
assert_not_contains "alice:wonderland" "$RUN_OUTPUT" "the userinfo must never reach the output"
|
||||
assert_contains "amqp://<redacted>@rabbit.local:5672/vhost" "$RUN_OUTPUT" \
|
||||
"the userinfo must be MASKED, not deleted — the rest of the value must survive"
|
||||
assert_contains "note:" "$RUN_OUTPUT" "the key name must still reach the output"
|
||||
assert_contains "see " "$RUN_OUTPUT" "prose BEFORE the userinfo must still reach the output"
|
||||
assert_contains "for details" "$RUN_OUTPUT" "prose AFTER the userinfo must still reach the output"
|
||||
}
|
||||
|
||||
# A diff line can hold a URL with no userinfo, followed later on the same line by an unrelated @
|
||||
# (free text in a string value, for example an email address). The line must pass through the
|
||||
# userinfo sed byte for byte: the match must stop at the end of the URL and must not treat the
|
||||
# later @ as a second userinfo delimiter.
|
||||
test_diff_line_uri_without_userinfo_survives_a_later_at_sign() {
|
||||
local dir
|
||||
dir="$(new_fixture)"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.note2=see https://docs.local/guide and mail ops@example.com'
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 0 "$RUN_RC" "diff-no-userinfo-with-later-at-sign reload exit code"
|
||||
assert_contains "note2: see https://docs.local/guide and mail ops@example.com" "$RUN_OUTPUT" \
|
||||
"a URL with no userinfo plus a later @ on the same line must pass through byte for byte"
|
||||
}
|
||||
|
||||
# Two credentialed URIs on one diff line must both be masked — the g flag matters.
|
||||
test_diff_line_masks_multiple_userinfo_with_g_flag() {
|
||||
local dir
|
||||
dir="$(new_fixture)"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.note3=amqp://u1:p1@host1/vhost1 and amqp://u2:p2@host2/vhost2'
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 0 "$RUN_RC" "diff-two-userinfo-on-one-line reload exit code"
|
||||
assert_not_contains "u1:p1" "$RUN_OUTPUT" "the first userinfo must never reach the output"
|
||||
assert_not_contains "u2:p2" "$RUN_OUTPUT" "the second userinfo must never reach the output"
|
||||
assert_contains "amqp://<redacted>@host1/vhost1" "$RUN_OUTPUT" "the first URI must be masked"
|
||||
assert_contains "amqp://<redacted>@host2/vhost2" "$RUN_OUTPUT" "the second URI must be masked"
|
||||
}
|
||||
|
||||
# ------------------------------------------------------- acceptance criterion 9: forgotten value
|
||||
# `--set .a.b=` is a plausible typo (the value simply forgotten), and it must be refused outright
|
||||
# rather than silently nulling the field — a null numeric field falls back to its default, which
|
||||
@@ -629,6 +689,65 @@ test_refusal_shape_from_parse_failure_wording_is_recognised() {
|
||||
assert_equals 4 "$RUN_RC" "the parse-failure refusal shape must also exit 4, not be read as silence"
|
||||
}
|
||||
|
||||
# A verdict line carrying a credentialed URI has its userinfo masked, with a positive control
|
||||
# proving the rest of the line still reaches the output unchanged.
|
||||
test_verdict_userinfo_is_masked_with_positive_control() {
|
||||
local dir
|
||||
dir="$(new_fixture)"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
|
||||
sleep 1
|
||||
printf 'config reload from %s refused, keeping the running config: refusing to start: malformed pattern — profiles.local.errorPattern ("amqp://user:hunter2@host/vhost"): Unclosed character class near index 8\n' \
|
||||
"$dir/fleetd.yaml" >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 4 "$RUN_RC" "refusal-with-userinfo exit code"
|
||||
assert_not_contains "user:hunter2" "$RUN_OUTPUT" "the userinfo must never reach the output"
|
||||
assert_contains "amqp://<redacted>@host/vhost" "$RUN_OUTPUT" \
|
||||
"the userinfo must be MASKED, not deleted — the rest of the quoted value must survive"
|
||||
# Positive control: the diagnostic prose on both sides of the userinfo must still reach the
|
||||
# output. Without this, a mutant that drops the whole verdict line would pass identically.
|
||||
assert_contains "malformed pattern" "$RUN_OUTPUT" "prose BEFORE the userinfo must still reach the output"
|
||||
assert_contains "Unclosed character class near index 8" "$RUN_OUTPUT" \
|
||||
"prose AFTER the userinfo must still reach the output"
|
||||
}
|
||||
|
||||
# An ordinary refusal line quotes the offending pattern, not a credential, and must survive byte
|
||||
# for byte: the rewrite is scoped to userinfo only, and the quoted pattern is the detail an
|
||||
# operator needs to fix the refusal.
|
||||
test_ordinary_refusal_line_passes_through_unchanged() {
|
||||
local dir real_line
|
||||
dir="$(new_fixture)"
|
||||
real_line="config reload from $dir/fleetd.yaml refused, keeping the running config: refusing to start: malformed pattern — profiles.local.errorPattern (\"[unclosed\"): Unclosed character class near index 8"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
|
||||
sleep 1
|
||||
printf '%s\n' "$real_line" >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 4 "$RUN_RC" "ordinary refusal exit code"
|
||||
assert_contains "$real_line" "$RUN_OUTPUT" \
|
||||
"an ordinary refusal with no userinfo must pass through byte for byte, unchanged"
|
||||
}
|
||||
|
||||
# A verdict line can hold a URI with NO userinfo and a later, unrelated @ further on in the same
|
||||
# line (an email address in diagnostic prose, for example). The rewrite must stop at the end of
|
||||
# the URI and must not treat the later @ as a second userinfo delimiter.
|
||||
test_uri_without_userinfo_survives_a_later_at_sign() {
|
||||
local dir real_line
|
||||
dir="$(new_fixture)"
|
||||
real_line="config reload from $dir/fleetd.yaml refused, keeping the running config: broker.uri amqp://broker.local/vhost unreachable, contact ops@example.com"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
|
||||
sleep 1
|
||||
printf '%s\n' "$real_line" >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 4 "$RUN_RC" "no-userinfo-with-later-at-sign exit code"
|
||||
assert_contains "$real_line" "$RUN_OUTPUT" \
|
||||
"a URI with no userinfo plus a later @ in the same line must pass through byte for byte"
|
||||
}
|
||||
|
||||
# --set runs yq over the whole candidate. It warns when that changes more lines than the requested
|
||||
# pairs, but a simple file with only the intended changed line must stay quiet.
|
||||
new_fixture_reformat_sensitive() {
|
||||
@@ -694,6 +813,12 @@ echo "== acceptance criterion 6: the marker works =="
|
||||
test_marker_skips_lines_before_it
|
||||
echo "== acceptance criterion 7 (+13: redaction is proven to have run) =="
|
||||
test_redaction_holds
|
||||
echo "== fleetd #692: a diff line's userinfo is masked, rest of the value survives =="
|
||||
test_diff_line_userinfo_is_masked_with_positive_control
|
||||
echo "== fleetd #692: a diff line's URI with no userinfo survives a later @ in the line =="
|
||||
test_diff_line_uri_without_userinfo_survives_a_later_at_sign
|
||||
echo "== fleetd #692: two userinfo URIs on one diff line are both masked =="
|
||||
test_diff_line_masks_multiple_userinfo_with_g_flag
|
||||
echo "== acceptance criterion 9: a forgotten value refuses and installs nothing =="
|
||||
test_forgotten_value_refuses_and_installs_nothing
|
||||
echo "== acceptance criterion 10: an explicit clear writes a bare null =="
|
||||
@@ -720,6 +845,12 @@ echo "== extra: --check is read-only and always exits 0 =="
|
||||
test_check_is_read_only_and_exits_zero
|
||||
echo "== extra: the parse-failure refusal shape is also recognised =="
|
||||
test_refusal_shape_from_parse_failure_wording_is_recognised
|
||||
echo "== verdict-redaction criteria 2+3: verdict userinfo is masked, rest of line survives =="
|
||||
test_verdict_userinfo_is_masked_with_positive_control
|
||||
echo "== verdict-redaction criterion 4: an ordinary refusal passes through unchanged =="
|
||||
test_ordinary_refusal_line_passes_through_unchanged
|
||||
echo "== fleetd #638: a URI with no userinfo survives a later @ in the same line =="
|
||||
test_uri_without_userinfo_survives_a_later_at_sign
|
||||
echo "== acceptance criterion 17: --set warns about yq formatting churn =="
|
||||
test_set_warns_when_yq_reformats_extra_lines
|
||||
echo "== acceptance criterion 18: --set stays quiet without formatting churn =="
|
||||
|
||||
+213
-63
@@ -447,7 +447,7 @@ test_assert_single_daemon_rejects_two_pids() {
|
||||
test_running_pid_excludes_self_matching_wrapper_shell() {
|
||||
local before after wrapper_pid
|
||||
before="$(running_pid)"
|
||||
sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' &
|
||||
sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 20' &
|
||||
wrapper_pid=$!
|
||||
sleep 0.3
|
||||
after="$(running_pid)"
|
||||
@@ -475,11 +475,26 @@ test_running_pid_excludes_self_matching_wrapper_shell() {
|
||||
# `comm` from the actually-executed binary's own path, not from `exec -a`'s argv[0] override (BSD
|
||||
# ties `comm` to argv[0], which is what makes this technique work here) — so on Linux this
|
||||
# specific fixture might report `comm=sh`, not `comm=java`, even though the REAL daemon (a literal
|
||||
# `java -jar target/fleetd.jar` process, never fabricated) is unaffected either way. I could not
|
||||
# `java -jar run/fleetd.jar` process, never fabricated) is unaffected either way. I could not
|
||||
# verify this fixture's behavior on Linux, so test_running_pid_counts_a_pid_whose_comm_is_java
|
||||
# below backstops the same claim (the allowlist admits a pid whose comm is `java`) with a stubbed
|
||||
# `ps`, which is identical bash on every platform and carries no such platform question.
|
||||
test_running_pid_finds_a_real_java_named_second_process() {
|
||||
local before after standin_pid
|
||||
before="$(running_pid)"
|
||||
( exec -a java sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 20' ) &
|
||||
standin_pid=$!
|
||||
sleep 0.3
|
||||
after="$(running_pid)"
|
||||
kill "$standin_pid" 2>/dev/null || true
|
||||
wait "$standin_pid" 2>/dev/null || true
|
||||
printf '%s\n' "$after" | grep -qxF "$standin_pid" \
|
||||
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) whose own argv holds the pattern: before=[$before] after=[$after]"
|
||||
}
|
||||
|
||||
# PATTERN matches a fleetd jar in either build layout, not only the run/ one: a process whose
|
||||
# argv names a jar under target/ must be found too, the same way the run/ case above is.
|
||||
test_running_pid_finds_a_real_java_named_process_from_target_dir() {
|
||||
local before after standin_pid
|
||||
before="$(running_pid)"
|
||||
( exec -a java sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' ) &
|
||||
@@ -489,7 +504,24 @@ test_running_pid_finds_a_real_java_named_second_process() {
|
||||
kill "$standin_pid" 2>/dev/null || true
|
||||
wait "$standin_pid" 2>/dev/null || true
|
||||
printf '%s\n' "$after" | grep -qxF "$standin_pid" \
|
||||
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) whose own argv holds the pattern: before=[$before] after=[$after]"
|
||||
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) naming a jar under target/: before=[$before] after=[$after]"
|
||||
}
|
||||
|
||||
# Broadening PATTERN to match both build layouts must not also broaden it into matching a
|
||||
# non-exec'ing shell that merely holds the target/ text as a literal argument, the same
|
||||
# self-matching shape test_running_pid_excludes_self_matching_wrapper_shell above excludes for
|
||||
# the run/ text.
|
||||
test_running_pid_excludes_self_matching_wrapper_shell_naming_target_dir() {
|
||||
local before after wrapper_pid
|
||||
before="$(running_pid)"
|
||||
sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' &
|
||||
wrapper_pid=$!
|
||||
sleep 0.3
|
||||
after="$(running_pid)"
|
||||
kill "$wrapper_pid" 2>/dev/null || true
|
||||
wait "$wrapper_pid" 2>/dev/null || true
|
||||
[ "$after" = "$before" ] \
|
||||
|| fail "running_pid() counted a self-matching wrapper shell (pid $wrapper_pid, holding 'target/fleetd.jar' as literal text in its own argv, not the daemon): before=[$before] after=[$after]"
|
||||
}
|
||||
|
||||
# fleetd #593 CORRECTION 1, hole 2 — the round-1 filter denied known shell names (sh/bash/zsh/
|
||||
@@ -557,29 +589,30 @@ test_die_message_does_not_recommend_bare_pgrep_as_remediation() {
|
||||
|| fail "assert_single_daemon's die message does not say in words that a pattern can match the caller (fleetd #593)"
|
||||
}
|
||||
|
||||
# fleetd #511 — jar_id()'s no-argument default was unpinned by any test: nothing proved it reports
|
||||
# $JAR (the live path) rather than $JAR_STAGED. Both halves matter, so this pins both: the bare call
|
||||
# must hash the live jar, and an explicit path argument must hash THAT file, not fall back to $JAR.
|
||||
# Two files with different content, so a default pointed at the wrong one reports the wrong hash
|
||||
# rather than accidentally matching.
|
||||
# fleetd #511/#664 — jar_id()'s no-argument default was unpinned by any test: nothing proved it
|
||||
# reports $JAR (the live path) rather than whatever explicit path a caller passes it (e.g.
|
||||
# $BUILD_JAR). Both halves matter, so this pins both: the bare call must hash the live jar, and an
|
||||
# explicit path argument must hash THAT file, not fall back to $JAR. Two files with different
|
||||
# content, so a default pointed at the wrong one reports the wrong hash rather than accidentally
|
||||
# matching.
|
||||
test_jar_id_defaults_to_live_and_reports_explicit_path() {
|
||||
local dir saved_jar="$JAR" saved_staged="$JAR_STAGED"
|
||||
local live_hash staged_hash default_result explicit_result
|
||||
local dir saved_jar="$JAR"
|
||||
local live_hash other_hash default_result explicit_result other_path
|
||||
dir="$TMP/jar-id"; mkdir -p "$dir"
|
||||
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
|
||||
JAR="$dir/fleetd.jar"; other_path="$dir/other.jar"
|
||||
printf 'live jar bytes' > "$JAR"
|
||||
printf 'staged jar bytes, not the same content' > "$JAR_STAGED"
|
||||
printf 'other jar bytes, not the same content' > "$other_path"
|
||||
# fleetd #550: this reference hash must be computed the same portable way jar_id() itself now
|
||||
# computes one — a bare, unguarded call to the macOS-only hasher here was exactly the item-2
|
||||
# defect, dying with "command not found" on any Linux runner that has no such hasher at all.
|
||||
live_hash="$(hash256 "$JAR")"
|
||||
staged_hash="$(hash256 "$JAR_STAGED")"
|
||||
other_hash="$(hash256 "$other_path")"
|
||||
default_result="$(jar_id)"
|
||||
explicit_result="$(jar_id "$JAR_STAGED")"
|
||||
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
|
||||
[ "$live_hash" != "$staged_hash" ] || fail "test fixture error: live and staged jars hashed the same"
|
||||
explicit_result="$(jar_id "$other_path")"
|
||||
JAR="$saved_jar"
|
||||
[ "$live_hash" != "$other_hash" ] || fail "test fixture error: live and other jars hashed the same"
|
||||
assert_equals "$live_hash" "$default_result" "jar_id with no arguments must report the hash of \$JAR"
|
||||
assert_equals "$staged_hash" "$explicit_result" "jar_id \"\$JAR_STAGED\" must report the hash of the staged jar, not fall back to \$JAR"
|
||||
assert_equals "$other_hash" "$explicit_result" "jar_id with an explicit path must report the hash of that path, not fall back to \$JAR"
|
||||
}
|
||||
|
||||
# fleetd #550 — closes a gap the test above leaves open. That test's own reference hash is now ALSO
|
||||
@@ -651,34 +684,77 @@ test_jar_id_reports_unhashable_when_no_hasher_on_path() {
|
||||
assert_equals "unhashable" "$explicit_result" "jar_id (explicit path) with no hasher on PATH must report the same third state"
|
||||
}
|
||||
|
||||
# fleetd #493 — never build into the path a running process holds. stage_built_jar/swap_staged_jar
|
||||
# are exercised directly against real files on disk (not stubs), because the whole point is file
|
||||
# fleetd #664 — report_jar_state is the --check fix: under the old layout $JAR and the build
|
||||
# output were the same file, so a single "jar on disk" fact covered both. Now they can disagree,
|
||||
# and this is the function that is supposed to show that. Agreeing case: two files with IDENTICAL
|
||||
# content must print both labels and never warn.
|
||||
test_report_jar_state_agrees_when_hashes_match() {
|
||||
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
|
||||
dir="$TMP/report-jar-agree"; mkdir -p "$dir"
|
||||
BUILD_JAR="$dir/target-fleetd.jar"; JAR="$dir/run-fleetd.jar"
|
||||
printf 'identical jar bytes' > "$BUILD_JAR"
|
||||
printf 'identical jar bytes' > "$JAR"
|
||||
output="$(report_jar_state)"
|
||||
BUILD_JAR="$saved_build"; JAR="$saved_jar"
|
||||
printf '%s' "$output" | grep -qF 'built jar' \
|
||||
|| fail "report_jar_state did not label the built jar"
|
||||
printf '%s' "$output" | grep -qF 'running jar' \
|
||||
|| fail "report_jar_state did not label the running jar"
|
||||
if printf '%s' "$output" | grep -qF 'differ'; then
|
||||
fail "report_jar_state warned about a mismatch when both jars have identical content"
|
||||
fi
|
||||
}
|
||||
|
||||
# The disagreeing case: this is the whole point of the ticket — a built jar that is NOT the
|
||||
# running jar must be visibly flagged, not silently printed as two unremarkable facts.
|
||||
test_report_jar_state_warns_when_hashes_differ() {
|
||||
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
|
||||
dir="$TMP/report-jar-differ"; mkdir -p "$dir"
|
||||
BUILD_JAR="$dir/target-fleetd.jar"; JAR="$dir/run-fleetd.jar"
|
||||
printf 'freshly built jar bytes' > "$BUILD_JAR"
|
||||
printf 'older running jar bytes' > "$JAR"
|
||||
output="$(report_jar_state)"
|
||||
BUILD_JAR="$saved_build"; JAR="$saved_jar"
|
||||
printf '%s' "$output" | grep -qF 'differ' \
|
||||
|| fail "report_jar_state did not warn when the built jar and running jar disagree"
|
||||
}
|
||||
|
||||
# Neither file existing (a fresh checkout, never built or deployed) must report two "absent"
|
||||
# facts and never a false mismatch warning — "absent" vs "absent" is agreement, not a diff.
|
||||
test_report_jar_state_both_absent_is_not_a_mismatch() {
|
||||
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
|
||||
dir="$TMP/report-jar-absent"; mkdir -p "$dir"
|
||||
BUILD_JAR="$dir/no-such-target.jar"; JAR="$dir/no-such-run.jar"
|
||||
output="$(report_jar_state)"
|
||||
BUILD_JAR="$saved_build"; JAR="$saved_jar"
|
||||
printf '%s' "$output" | grep -qF 'absent' \
|
||||
|| fail "report_jar_state did not report absent for a missing built/running jar"
|
||||
if printf '%s' "$output" | grep -qF 'differ'; then
|
||||
fail "report_jar_state warned about a mismatch when both jars are simply absent"
|
||||
fi
|
||||
}
|
||||
|
||||
# Reads JAR and BUILD_JAR exactly as the script sources them, with nothing here assigning
|
||||
# either first. JAR must resolve outside $MODULE/target/, and JAR must differ from BUILD_JAR:
|
||||
# the daemon's live path and Maven's own build output are never the same file.
|
||||
test_jar_and_build_jar_are_sourced_outside_target_and_differ() {
|
||||
source "$ROOT/scripts/redeploy-fleetd.sh"
|
||||
case "$JAR" in
|
||||
"$MODULE"/target/*)
|
||||
fail "\$JAR must not live under \$MODULE/target/ — got $JAR" ;;
|
||||
esac
|
||||
[ "$JAR" != "$BUILD_JAR" ] \
|
||||
|| fail "\$JAR and \$BUILD_JAR must not be the same path — got $JAR"
|
||||
}
|
||||
|
||||
# fleetd #493/#664 — never build into the path a running process holds. swap_staged_jar is
|
||||
# exercised directly against real files on disk (not stubs), because the whole point is file
|
||||
# behavior (does the content move, does the source disappear, does a failure leave both sides
|
||||
# intact) that a stubbed function cannot prove.
|
||||
test_stage_built_jar_moves_off_live_path() {
|
||||
local dir jar staged saved_jar="$JAR" saved_staged="$JAR_STAGED"
|
||||
dir="$TMP/stage-ok"; mkdir -p "$dir"
|
||||
jar="$dir/fleetd.jar"; staged="$dir/fleetd-new.jar"
|
||||
printf 'built jar bytes' > "$jar"
|
||||
JAR="$jar"; JAR_STAGED="$staged"
|
||||
stage_built_jar || fail "stage_built_jar rejected a real build output"
|
||||
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
|
||||
[ ! -f "$jar" ] || fail "stage_built_jar left the jar behind at the live path $jar"
|
||||
[ -f "$staged" ] || fail "stage_built_jar did not create the staged jar at $staged"
|
||||
grep -qF 'built jar bytes' "$staged" || fail "staged jar does not carry the built content"
|
||||
}
|
||||
|
||||
test_stage_built_jar_dies_when_build_produced_nothing() {
|
||||
local dir output rc=0 saved_jar="$JAR" saved_staged="$JAR_STAGED"
|
||||
dir="$TMP/stage-missing"; mkdir -p "$dir"
|
||||
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
|
||||
output="$(stage_built_jar 2>&1)" || rc=$?
|
||||
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
|
||||
[ "$rc" -ne 0 ] || fail "stage_built_jar accepted a missing build output"
|
||||
printf '%s' "$output" | grep -qF "$dir/fleetd.jar" \
|
||||
|| fail "refusal message does not name the missing jar path"
|
||||
}
|
||||
|
||||
# intact) that a stubbed function cannot prove. stage_built_jar no longer exists: under the #664
|
||||
# layout $BUILD_JAR (target/fleetd.jar) was never the live path, so a build has nothing to be
|
||||
# staged OUT of — swap_staged_jar is called directly against $BUILD_JAR/$JAR (see
|
||||
# test_swap_ordered_after_wait_and_before_start and the report_jar_state tests below for the rest
|
||||
# of that seam).
|
||||
test_swap_staged_jar_moves_staged_onto_live() {
|
||||
local dir staged live
|
||||
dir="$TMP/swap-ok"; mkdir -p "$dir"
|
||||
@@ -989,24 +1065,26 @@ test_swap_ordered_after_wait_and_before_start() {
|
||||
|| fail "swap_if_built (line $swap_line) is not before the start section (line $start_line)"
|
||||
}
|
||||
|
||||
# fleetd #511: the drain-gate abort message (fired when a build has staged a jar but the operator
|
||||
# declines the drain confirmation) used to tell the operator to "Rerun (with or without --no-build)"
|
||||
# to finish the restart. That is wrong — by the time this message can fire, stage_built_jar has
|
||||
# already moved the jar off $JAR, so a rerun WITH --no-build hits require_no_build_jar's own refusal
|
||||
# ("no jar at $JAR — run without --no-build"). Like test_swap_ordered_after_wait_and_before_start
|
||||
# above, this code path is never reached by sourcing (the SOURCED guard stops before the main flow),
|
||||
# so the only way to pin its exact wording is to read the source.
|
||||
# fleetd #511: the drain-gate abort message (fired when a build has produced a jar but the
|
||||
# operator declines the drain confirmation) used to tell the operator to "Rerun (with or without
|
||||
# --no-build)" to finish the restart. That is wrong. fleetd #664 changed WHY it is wrong: under
|
||||
# the old layout a build emptied the live path immediately, so --no-build's own check refused a
|
||||
# rerun for you; now a build never touches the live path at all, so --no-build would NOT refuse —
|
||||
# it would quietly restart the daemon on the OLD jar and throw away the one just built. Like
|
||||
# test_swap_ordered_after_wait_and_before_start above, this code path is never reached by sourcing
|
||||
# (the SOURCED guard stops before the main flow), so the only way to pin its exact wording is to
|
||||
# read the source.
|
||||
test_drain_gate_abort_message_says_no_no_build() {
|
||||
local src="$ROOT/scripts/redeploy-fleetd.sh" msg
|
||||
msg="$(grep -A3 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
|
||||
msg="$(grep -A6 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
|
||||
[ -n "$msg" ] || fail "could not find the drain-gate staged-jar abort message in redeploy-fleetd.sh"
|
||||
if printf '%s' "$msg" | grep -qF 'with or without --no-build'; then
|
||||
fail "abort message still claims a rerun WITH --no-build can finish the restart"
|
||||
fi
|
||||
printf '%s' "$msg" | grep -qF 'WITHOUT --no-build' \
|
||||
|| fail "abort message does not tell the operator to rerun without --no-build"
|
||||
printf '%s' "$msg" | grep -qF 'no longer at the live path' \
|
||||
|| fail "abort message does not say why --no-build cannot finish the restart"
|
||||
printf '%s' "$msg" | grep -qF 'silently discard' \
|
||||
|| fail "abort message does not say that --no-build would silently discard the build just made"
|
||||
}
|
||||
|
||||
# fleetd #517 — the drain-gate abort branch itself. Before this, the only test of this message was
|
||||
@@ -1149,19 +1227,22 @@ test_refuse_drain_gate_no_build_staged_absent() {
|
||||
#
|
||||
# fleetd #555 — the main flow's own call site moved: it used to read
|
||||
# `refuse_drain_gate "$DO_BUILD" "$JAR_STAGED"` directly; it now reads
|
||||
# `run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"`, and run_drain_gate (tested
|
||||
# directly below by test_run_drain_gate_*) is what calls refuse_drain_gate with its own local names.
|
||||
# This grep now pins THAT call site — the thing that would go missing if a future edit deleted the
|
||||
# main flow's call to run_drain_gate altogether, the same residual gap #521/#528 already accepted for
|
||||
# swap_if_built/refuse_drain_gate (sourcing stops before the main flow runs, so no test in this file
|
||||
# can do better than reading the source for this one specific gap).
|
||||
# `run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"` (fleetd #664 renamed the
|
||||
# fourth argument from $JAR_STAGED to $BUILD_JAR — same role, the not-yet-swapped-in jar — when
|
||||
# that path stopped being a separate staging file and became target/fleetd.jar itself), and
|
||||
# run_drain_gate (tested directly below by test_run_drain_gate_*) is what calls refuse_drain_gate
|
||||
# with its own local names. This grep now pins THAT call site — the thing that would go missing if
|
||||
# a future edit deleted the main flow's call to run_drain_gate altogether, the same residual gap
|
||||
# #521/#528 already accepted for swap_if_built/refuse_drain_gate (sourcing stops before the main
|
||||
# flow runs, so no test in this file can do better than reading the source for this one specific
|
||||
# gap).
|
||||
#
|
||||
# The grep ends `|| true`: this file runs under `set -euo pipefail`, so an ABSENT needle would fail
|
||||
# the assignment and `set -e` would kill the whole suite before the `[ -n ... ] || fail` guard below
|
||||
# ever ran — the exact dead-check shape fleetd #528 also flags as a sweep finding (see the PR body).
|
||||
test_run_drain_gate_call_site_present() {
|
||||
local src="$ROOT/scripts/redeploy-fleetd.sh" call_line
|
||||
call_line="$(grep -Fn 'run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"' "$src" | head -1 | cut -d: -f1 || true)"
|
||||
call_line="$(grep -Fn 'run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"' "$src" | head -1 | cut -d: -f1 || true)"
|
||||
[ -n "$call_line" ] \
|
||||
|| fail "could not find the main flow's run_drain_gate call site in redeploy-fleetd.sh"
|
||||
}
|
||||
@@ -1269,12 +1350,71 @@ test_run_drain_gate_declined_reply_refuses() {
|
||||
source "$ROOT/scripts/redeploy-fleetd.sh"
|
||||
}
|
||||
|
||||
# Writes a launchd-plist fixture naming jar_path as the ProgramArguments entry after "-jar", so
|
||||
# check_jar_path_matches_plist has something real to read back.
|
||||
write_launchd_plist_fixture() {
|
||||
local path="$1" jar_path="$2"
|
||||
cat > "$path" <<PLIST
|
||||
<?xml version="1.0" encoding="UTF-8"?>
|
||||
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
|
||||
<plist version="1.0">
|
||||
<dict>
|
||||
<key>Label</key>
|
||||
<string>test.fixture</string>
|
||||
<key>ProgramArguments</key>
|
||||
<array>
|
||||
<string>/usr/bin/java</string>
|
||||
<string>-jar</string>
|
||||
<string>$jar_path</string>
|
||||
<string>fleetd.yaml</string>
|
||||
</array>
|
||||
</dict>
|
||||
</plist>
|
||||
PLIST
|
||||
}
|
||||
|
||||
# Agreeing case: a plist whose ProgramArguments names the same jar, resolved, must proceed and
|
||||
# say so, never die.
|
||||
test_check_jar_path_matches_plist_agrees_ok() {
|
||||
local dir jar plist output rc=0
|
||||
dir="$TMP/jar-path-agree"; mkdir -p "$dir/run"
|
||||
jar="$dir/run/fleetd.jar"
|
||||
plist="$dir/agree.plist"
|
||||
write_launchd_plist_fixture "$plist" "$jar"
|
||||
output="$(check_jar_path_matches_plist "$jar" "$plist" 2>&1)" || rc=$?
|
||||
[ "$rc" -eq 0 ] \
|
||||
|| fail "check_jar_path_matches_plist must succeed when the plist names the same jar: $output"
|
||||
printf '%s' "$output" | grep -qF 'jar path check' \
|
||||
|| fail "check_jar_path_matches_plist did not print the agreement line: $output"
|
||||
}
|
||||
|
||||
# The disagreeing case, and the positive control this check exists for: a plist naming a
|
||||
# different jar must die, naming both paths.
|
||||
test_check_jar_path_matches_plist_mismatch_dies() {
|
||||
local dir jar other_jar plist output rc=0
|
||||
dir="$TMP/jar-path-mismatch"; mkdir -p "$dir/run" "$dir/target"
|
||||
jar="$dir/run/fleetd.jar"
|
||||
other_jar="$dir/target/fleetd.jar"
|
||||
plist="$dir/mismatch.plist"
|
||||
write_launchd_plist_fixture "$plist" "$other_jar"
|
||||
output="$(check_jar_path_matches_plist "$jar" "$plist" 2>&1)" || rc=$?
|
||||
[ "$rc" -ne 0 ] \
|
||||
|| fail "check_jar_path_matches_plist must die when the plist names a different jar"
|
||||
printf '%s' "$output" | grep -qF "$jar" \
|
||||
|| fail "die message does not name this script's jar path: $output"
|
||||
printf '%s' "$output" | grep -qF "$other_jar" \
|
||||
|| fail "die message does not name the plist's jar path: $output"
|
||||
}
|
||||
|
||||
# fleetd #555 item 4 — the report-state dispatch on $SUPERVISOR_KIND. Inverting this used to report
|
||||
# the wrong supervisor and, for the launchd arm specifically, skip check_log_path_matches_plist.
|
||||
CHECK_LOG_PATH_CALLED=0
|
||||
CHECK_JAR_PATH_CALLED=0
|
||||
stub_check_log_path_recorder() {
|
||||
CHECK_LOG_PATH_CALLED=0
|
||||
CHECK_JAR_PATH_CALLED=0
|
||||
check_log_path_matches_plist() { CHECK_LOG_PATH_CALLED=1; }
|
||||
check_jar_path_matches_plist() { CHECK_JAR_PATH_CALLED=1; }
|
||||
}
|
||||
|
||||
test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path() {
|
||||
@@ -1285,6 +1425,8 @@ test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path() {
|
||||
[ "$SUPERVISED" = 1 ] || fail "report_supervisor_state launchd must set SUPERVISED=1"
|
||||
[ "$CHECK_LOG_PATH_CALLED" = 1 ] \
|
||||
|| fail "report_supervisor_state launchd must call check_log_path_matches_plist"
|
||||
[ "$CHECK_JAR_PATH_CALLED" = 1 ] \
|
||||
|| fail "report_supervisor_state launchd must call check_jar_path_matches_plist"
|
||||
source "$ROOT/scripts/redeploy-fleetd.sh"
|
||||
}
|
||||
|
||||
@@ -1296,6 +1438,8 @@ test_report_supervisor_state_systemd_sets_supervised_without_log_path_check() {
|
||||
[ "$SUPERVISED" = 1 ] || fail "report_supervisor_state systemd must set SUPERVISED=1"
|
||||
[ "$CHECK_LOG_PATH_CALLED" = 0 ] \
|
||||
|| fail "report_supervisor_state systemd must NOT call check_log_path_matches_plist"
|
||||
[ "$CHECK_JAR_PATH_CALLED" = 0 ] \
|
||||
|| fail "report_supervisor_state systemd must NOT call check_jar_path_matches_plist"
|
||||
source "$ROOT/scripts/redeploy-fleetd.sh"
|
||||
}
|
||||
|
||||
@@ -2153,6 +2297,8 @@ test_assert_single_daemon_accepts_one_pid
|
||||
test_assert_single_daemon_rejects_two_pids
|
||||
test_running_pid_excludes_self_matching_wrapper_shell
|
||||
test_running_pid_finds_a_real_java_named_second_process
|
||||
test_running_pid_finds_a_real_java_named_process_from_target_dir
|
||||
test_running_pid_excludes_self_matching_wrapper_shell_naming_target_dir
|
||||
test_running_pid_drops_a_pid_whose_comm_is_not_java
|
||||
test_running_pid_drops_a_pid_that_exited_before_the_comm_lookup
|
||||
test_running_pid_counts_a_pid_whose_comm_is_java
|
||||
@@ -2161,8 +2307,10 @@ test_jar_id_defaults_to_live_and_reports_explicit_path
|
||||
test_hash256_computes_a_real_sha256
|
||||
test_jar_id_reports_absent_for_missing_file
|
||||
test_jar_id_reports_unhashable_when_no_hasher_on_path
|
||||
test_stage_built_jar_moves_off_live_path
|
||||
test_stage_built_jar_dies_when_build_produced_nothing
|
||||
test_report_jar_state_agrees_when_hashes_match
|
||||
test_report_jar_state_warns_when_hashes_differ
|
||||
test_report_jar_state_both_absent_is_not_a_mismatch
|
||||
test_jar_and_build_jar_are_sourced_outside_target_and_differ
|
||||
test_swap_staged_jar_moves_staged_onto_live
|
||||
test_swap_staged_jar_dies_without_staged_file
|
||||
test_swap_staged_jar_dies_when_mv_fails
|
||||
@@ -2204,6 +2352,8 @@ test_drain_confirmed_false_on_anything_else
|
||||
test_run_drain_gate_skips_prompt_when_not_required
|
||||
test_run_drain_gate_confirmed_reply_does_not_refuse
|
||||
test_run_drain_gate_declined_reply_refuses
|
||||
test_check_jar_path_matches_plist_agrees_ok
|
||||
test_check_jar_path_matches_plist_mismatch_dies
|
||||
test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path
|
||||
test_report_supervisor_state_systemd_sets_supervised_without_log_path_check
|
||||
test_report_supervisor_state_none_leaves_supervised_zero
|
||||
|
||||
Reference in New Issue
Block a user