Compare commits

..

1 Commits

Author SHA1 Message Date
Dai Ha fe465cc425 fleetd #637: scale LeadContextGauge's HIGH threshold with the effective auto-compact window
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Failing after 2m2s
HIGH_THRESHOLD_TOKENS was a hardcoded 200_000, while the event it warns about
(auto-compaction) is configured per profile via autoCompactWindow and can legally go
as low as 100_000 — making HIGH unreachable before a compaction on such a profile.

LeadContextGauge.read now takes an optional effective window and fires HIGH at 2/3 of
it, falling back to the fixed 200_000 when no window is resolvable (unresolved callers,
including the pre-existing 3-arg read(), keep today's behaviour exactly).

FleetConfig.Profile.effectiveAutoCompactWindow() resolves that window the way a
launched Claude Code session actually reads it: env.CLAUDE_CODE_AUTO_COMPACT_WINDOW
wins over the autoCompactWindow launch flag when both are set.

Wired into both real consumers: fleet_list's context row (FleetMcp.LeadConfigDirSource,
widened with a back-compat constructor so no unrelated call site changes) and the lead
heartbeat's context-high nudge (Fleetd.leadContextLookup/leadContextSource, widened the
same way).
2026-10-03 15:54:33 +02:00
75 changed files with 1929 additions and 9754 deletions
+1 -1
View File
@@ -59,7 +59,7 @@ as `matches HEAD`, `drift`, or `unknown`; do not turn an unclear timestamp into
Report the process identifier (PID) and uptime too:
```bash
PIDS="$(pgrep -f 'fleetd.jar' || true)"
PIDS="$(pgrep -f 'target/fleetd.jar' || true)"
if [ -z "$PIDS" ]; then
printf '%s\n' 'fleetd: not running'
else
+7 -42
View File
@@ -159,41 +159,11 @@ fails.
`{action: "cancel", token}` drops a pending request without rolling.
**A change is coming: the roll will restart the process instead of sending `/clear` (fleetd #726
unit 2, written 2026-10-04).**
Today a roll types `/clear` into your pane. Your `claude` process keeps running, so a newer CLI on
disk is never loaded. Unit 2 replaces that: the daemon ends the old pane, launches a fresh one,
waits for the new terminal to be recognised as a lead, and only then sends the bootstrap text.
**Everything below about `/clear` is accurate while the old jar is running.** Unit 2 was not merged
when this note was written, and a merge is not a deployment.
**How to tell which one is live: read your own tool list.** If `fleet_handover`'s description says
it will "clear your pane", the daemon is serving the old behaviour. If it names a restart, the new
behaviour is live. The description comes from the running daemon, so it cannot disagree with the
code that is actually loaded.
Two things change for you once it is live. The `status` outcomes are different: three new failures
replace the `/clear` ones. And the "never observed as WORKING after 8 consecutive IDLE/DONE polls"
warning described below can no longer appear, because that wait is deleted — so if you still see
it, the old jar is running. `TURN_NEVER_SETTLED` does not change, and still means nothing was
touched.
**Delete this note and rewrite the `/clear` paragraphs once the new jar is live.**
**Things that will surprise you:**
- **`accepted` does not mean your pane has been cleared.** It means every gate passed and the roll
is scheduled to run once your current turn ends. Say your goodbye in the same turn — you will not
get another one.
- **If you are still running after that turn, the roll did not happen.** A roll that works clears
you, so surviving your own goodbye is itself the signal that it refused. Check with
`fleet_handover{action: "status", token}`, using the token you confirmed. `TURN_NEVER_SETTLED`
means your turn ran past `leadRollover.turnSettleSeconds` and **no `/clear` was ever sent**: your
context is intact and nothing was lost. Open a fresh request and retry. Never assume the roll
succeeded because `confirm` answered `accepted` — by the time it refuses, there is no caller left
to tell, so this check is the only thing that closes that gap.
- **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved
from your connection, so you can only ever roll yourself.
- **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you
@@ -223,18 +193,13 @@ touched.
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
- **The bootstrap prompt works end to end. Measured 2026-09-22, re-measured 2026-10-04.** This used
to say the fix was unproven (fleetd #489) and told you to expect a failure. That is no longer
true. On 2026-09-22 the daemon log held four `lead-rollover: rolled` lines. On 2026-10-04 it holds
**20**, against a control of 86 `lead-rollover:` lines. Each roll cleared the old lead and started
a fresh session against the handover file, with the configured `bootstrapText` arriving as its
first message. No context was lost. The old `Unknown command: /clearFresh` failure from 2026-09-12
does not appear in the log at all.
19 of the 20 carry an `elapsedMs`: median 16507 ms, maximum 48261 ms, and two above 45000 ms. That
figure times the **whole** roll, and the wait for your own turn to end dominates it, so do not
read it as the cost of the clear. Expect a roll to take tens of seconds, and do not treat a slow
one as a failed one. Re-measure all of these with:
- **The bootstrap prompt works end to end. Measured 2026-09-22.** This used to say the fix was
unproven (fleetd #489) and told you to expect a failure. That is no longer true. The daemon log
now holds four `lead-rollover: rolled` lines, and three of them ran on 2026-09-22 at 10:01:43,
10:38:28 and 11:15:47. Each one cleared the old lead and started a fresh session against the
handover file, with the configured `bootstrapText` arriving as its first message. No context was
lost. The old `Unknown command: /clearFresh` failure from 2026-09-12 does not appear in the log
at all. Re-measure both numbers with:
```bash
grep -c "lead-rollover: rolled" fleetd/fleetd.out # successful rolls
+7 -16
View File
@@ -22,22 +22,13 @@ scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already chec
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
```
`--no-build` skips the build and restarts whatever jar is at `fleetd/run/fleetd.jar` — the runtime
path, not Maven's output path. Use it only when you just built and nothing changed since. It gives
up the protection in the next paragraph: no build runs, so a stale or missing jar is not caught
early. The script still checks the file is there and dies with `no jar at … — run without
--no-build` if it is not, but it cannot tell you the jar is old.
The daemon runs from `fleetd/run/fleetd.jar`, not from `fleetd/target/fleetd.jar` where Maven
writes its output (fleetd #664). That split is what makes a bare `mvn install`/`mvn clean` in the
main clone harmless now: neither can reach the file the running daemon holds open, because that
file no longer lives under `target/` at all. Verify a merge by building in a throwaway git
worktree anyway — a build still produces nothing the fleet runs until this script's own `mv` of
`target/fleetd.jar` onto `run/fleetd.jar`, performed only after the old daemon is confirmed gone.
Let only `scripts/redeploy-fleetd.sh` touch `fleetd/run/fleetd.jar`. Run `--check` first: it prints
the BUILT jar (`target/fleetd.jar`) and the RUNNING jar (`run/fleetd.jar`) as two separately
labelled hash-and-mtime facts, so a mismatch between them — a build sitting unswapped, or a stale
runtime jar — is visible before you decide anything.
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
+19 -76
View File
@@ -27,34 +27,23 @@ through its `fleet_*` tools. No session addresses a peer, a broker, or the netwo
**Every role reads this file.** A member runs in a git worktree of this same repo, so it inherits
this `CLAUDE.md` verbatim, and every rule below is role-conditional.
**Call `fleet_whoami`.** It returns `primary`, `worker`, `architect`, `collaborator`, or `observer`,
resolved by the daemon from your connection — unforgeable, and the same resolution its authorization
gate uses. A worker also carries its `sessionId`, `profile`, `worktree` and `branch`; an architect
carries the slot name it was bound to; a collaborator carries its registry name and its own
`sessionId`, and **no `leader` key** — a collaborator is a named peer, not a primary. An **observer**
carries only its own `sessionId`: a pane the daemon could not place as any of the above, authorized
to `READ`/`METRICS` and to `REPLY`/`ASK` on its own pane and nothing more — never `SEND`, never a
ticket. Don't infer what you can ask.
**Call `fleet_whoami`.** It returns `primary`, `worker`, or `architect`, resolved by the daemon from
your connection — unforgeable, and the same resolution its authorization gate uses. A worker also
carries its `sessionId`, `profile`, `worktree` and `branch`; an architect carries the slot name it
was bound to. Don't infer what you can ask.
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
fires: the reply charter in your system prompt (*"You are a spawned member in the
claude-bridge fleet"*) ⇒ **spawned member**; fleet tools prefixed `mcp__fleet__*` ⇒ **spawned
member** (the launcher fixes that mount name; a primary's mount is named by whoever wrote its
`.mcp.json`, so it varies — and a member spawned before CB-632 still says `mcp__bridge__*`); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect,
or a worker from an **observer** — an observer is just as unspawned as a collaborator and carries
none of these signals either, so only `fleet_whoami` tells the two apart. **And none of them fires
for a collaborator at all**: every signal in the ladder detects a *spawned* member, while a
collaborator is a tab a person opened by hand, so it has no charter, no fixed mount name and a
normal environment. A collaborator — or an observer — that cannot call `fleet_whoami` therefore falls
to the line below and acts as a worker. That is the safe direction — it under-privileges, and the
refusals are loud — but it means a collaborator or an observer has no way to learn what it is except
by asking. **Still unsure ⇒ act as a worker**, the most restricted member role this ladder can name.
The two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate
— loud and self-correcting — while a member acting as the primary ends its turn with no `fleet_reply`,
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect —
only `fleet_whoami` does. **Still unsure ⇒ act as a worker**, the most restricted member role. The
two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate —
loud and self-correcting — while a member acting as the primary ends its turn with no `fleet_reply`,
and the sender silently receives nothing. Fail toward the recoverable error.
### Invariants — every role, no exceptions
### Invariants — both roles, no exceptions
1. **Never set, export, or forward `ANTHROPIC_BASE_URL`** (or `ANTHROPIC_AUTH_TOKEN`). The primary
stays on subscription; only the bridge puts a member off it, at spawn. Mounting the bridge must
@@ -62,10 +51,9 @@ and the sender silently receives nothing. Fail toward the recoverable error.
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `fleet_*` call is silently discarded.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead, architect, or
collaborator** — and a collaborator may send only to a lead or another collaborator, never to a
spawned member's terminal; reply/ask are only-as-itself — any peer may answer for its own pane,
and for no other. A call outside your role is refused, not queued.
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead or architect**;
reply/ask are only-as-itself — any peer may answer for its own pane, and for no other. A call
outside your role is refused, not queued.
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
@@ -105,11 +93,7 @@ below are the procedure — run them in order, every task, not only the big ones
5. **Collect** — `fleet_poll{ticket}` → `fleet_ack{target, msgId}`. Answer a worker's `fleet_ask`
with `fleet_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`fleet_status`, never by reading its terminal; it also reports an open question and the `turnId`
that answers it — but **only to the caller that created that delegation**, so a question raised
under an architect's brief is invisible to you, and seeing none does not mean there is none.
**Only that same creator can answer it.** A `turnId` you came by any other way is refused, so an
architect's worker waits for that architect and not for you.
**A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
**A correction cannot reach a busy member.** A `fleet_send` to a working member is *accepted* and
returns a ticket, and is then never delivered — measured here three times in one session, and the
@@ -174,10 +158,9 @@ you decide.
| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) + `loopHealth` (`RUNNING`, `STALLED`, or `STOPPED` for `statusPoller` and `sessionReaper`) · one peer's state: `fleet_status{sessionId}` |
| Delegate (blocking) | `fleet_send{sessionId, content}` |
| Delegate (long task) | `fleet_send{sessionId, content, wait:false}` → ticket → `fleet_poll{ticket}` |
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId`, and only the caller that created that delegation |
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
| Message a **collaborator** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` reports a `collaborators` array, and each row carries that peer's `name` and the `sessionId` you send to. It is visible to you, to an architect and to another collaborator, never to a worker. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}` if they are on this host. **Not** `fleet_reply`: it has no peer route and the publish is refused |
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
@@ -251,27 +234,6 @@ simply complies has thrown away the reason there are two of you.
7. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
project marks as not-yours-to-commit.
### Collaborator — a named peer, not a member
`fleet_whoami` answered `collaborator`, so your pane's tab matches a `fleet.collaborators.<name>.tab`
entry. You are **not** a member: nothing delegates to you, you have no brief, no worktree and no
ticket, and **you owe no `fleet_reply`** — the turn contract above is for a session a lead spawned,
and it does not apply to you. Read it only to understand what the members around you are doing.
What you may do: observe the fleet (`fleet_list`, `fleet_profiles`, `fleet_whoami`), and send to a
lead or to another collaborator. What you may not: spawn, stop or drain anything, roll a lead's
session, answer a member's `fleet_ask`, poll a ticket, or send to a spawned member's terminal. Each
of those is refused at the gate, not queued.
Two limits worth knowing before you hit them. **You cannot reach a worker** — not even to help one —
because a worker belongs to the lead that spawned it, and routing around that would make you a
second orchestrator with no plan. Send to the lead instead. And **you cannot read a ticket**, so you
cannot collect a delegation's reply: `fleet_poll` refuses you at the role gate, and a ticket also
records the terminal that created it, so even a leaked id reads nothing.
Being named buys you a channel, not authority. Your `fleet_send` to a lead is coordination between
peers: the lead owes you no obedience, and you owe it none.
### Where each rule lives (don't duplicate — extend the right layer)
| Layer | Scope | Reaches |
@@ -339,22 +301,10 @@ must obey belongs in the charter, not here.
(#362). **Read `plugin/` before designing anything about onboarding a project.** Two limits are
structural, not bugs: a plugin cannot carry the role agent files, because
`ClaudeCodeLauncher.java:371` requires `<cwd>/.claude/agents/<role>.md` in the member's own
worktree; and a plugin reaches a member only through `CLAUDE_CONFIG_DIR`, which
`ClaudeCodeLauncher.java:286` exports with `putIfPresent` — so only for a profile that sets
`configDir`. Every `claude-code` profile does set one (the four without are `opencode`, which
never reads that variable). **But measured 2026-10-04: two of them point at
`~/.ccs/instances/ltms`, which is the operator's own `CLAUDE_CONFIG_DIR` on this host.** So for an
`opus` or `sonnet` member, "a member never reads the operator's plugin store" is false — it reads
the same store, because that store is the one its `configDir` names. It stays true for `local` and
`local-direct`, which point at `~/.ccs/instances/gx10`. `ClaudeCodeLauncher`'s own javadoc names
the related hazard: that file is rewritten on every spawn, so for those two profiles fleetd and the
operator's live session write the same `.claude.json`, and its compare-and-swap "narrows the
lost-update window, it does not close it". Re-measure which profiles share the operator's dir with
`awk '/^profiles:/{i=1;next} /^[a-z]/{i=0} i&&/^ [a-z-]+:$/{p=$1} i&&/configDir:/{print p,$2}'
fleetd/fleetd.yaml` against `echo $CLAUDE_CONFIG_DIR`; delete this note once no profile names the
operator's dir. **Treat the plugin as the lead-side surface and put member-facing assets in the
worktree** — that conclusion holds either way, because a worktree asset does not depend on which
config dir a member reads.
worktree; and a plugin cannot deliver anything to members at all, because
`ClaudeCodeLauncher.java:285` exports `CLAUDE_CONFIG_DIR` and every Claude profile here sets it,
so a member never reads the operator's plugin store. **The plugin is the lead-side surface;
member-facing assets travel in the worktree.**
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **A provisioned worktree neutralizes `.mcp.json`, `opencode.json` and `.autoenv`** — the repo's
@@ -373,13 +323,6 @@ must obey belongs in the charter, not here.
reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if
that page names a `fleet_*` tool the server does not register. The flows are kept out of this
file because this file loads into every session's context.
- **The daemon runs from `fleetd/run/fleetd.jar`, not `fleetd/target/fleetd.jar`** (fleetd #664).
Maven's own output still lands at `fleetd/target/fleetd.jar` — that part of the build is
unchanged — but the running daemon never has that file open, so a bare `mvn install`/`mvn clean`
in the main clone no longer corrupts anything a live process is reading. Verify merges in a
throwaway git worktree anyway: a build in the main clone still ships nothing until
`scripts/redeploy-fleetd.sh` moves it into place with its own atomic `mv`, performed only after
the old daemon is confirmed gone. Let only that script touch `fleetd/run/fleetd.jar`.
### Redeploying the daemon — the lead may do this (primary only)
@@ -409,7 +352,7 @@ Before you call any work done, check the row that matches what you touched:
| `ConnectionIdentity` / how a caller is resolved | the `fleet_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__fleet__*`), and the layering table's top row |
| the injector / status gating | invariant 4 |
| worktree provisioning or the parity overlay | the "every role reads this file" premise — it rests on the worker's worktree being a checkout of this repo |
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
| a new peer kind (non-Claude adapter) | what that peer can read — anything it must obey belongs in its charter, not in the block |
| **anything an operator can use, configure, or observe** — an MCP tool, a `fleetd.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
+1 -1
View File
@@ -47,7 +47,7 @@
<string>/Users/dai.ha/LTMS/claude-bridge/scripts/fleetd-launchd-wrapper.sh</string>
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin/java</string>
<string>-jar</string>
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/run/fleetd.jar</string>
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/target/fleetd.jar</string>
<string>fleetd.yaml</string>
</array>
+1 -1
View File
@@ -50,7 +50,7 @@ WorkingDirectory=%h/LTMS/fleetd/fleetd
# and looks healthy, and the failure appears hours later as a member that cannot open a pull
# request. exec keeps it one process, so systemd tracks the right PID.
# This also avoids a SECOND copy of the secrets in a systemd drop-in: one source of truth.
ExecStart=/bin/zsh -lc "exec java -jar run/fleetd.jar fleetd.yaml"
ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"
# PrivateTmp MUST stay false -- see herdr.service. fleetd creates the member ZDOTDIR scrub dir and
# the opencode config dir under java.io.tmpdir, and the member pane (a herdr child, a different
-4
View File
@@ -2,10 +2,6 @@
target/
dependency-reduced-pom.xml
# The daemon's runtime jar (fleetd #664). scripts/redeploy-fleetd.sh moves the built jar here
# with a same-filesystem rename; this is never Maven's output path and never belongs in git.
run/
# Local runtime config (copy from fleetd.example.yaml). Both names are ignored: fleetd.yaml is
# the current name, and bridged.yaml is the legacy name Fleetd still falls back to.
fleetd.yaml
+25 -48
View File
@@ -62,12 +62,11 @@ bind:
#
# Leads are configured under `fleet.leaders:` — see THE FLEET further down.
#
# Two things stop the tab-name convention from becoming a way to claim leadership: startup REFUSES
# a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel` override, also
# matches, so the two namespaces cannot overlap by accident; and the CallerResolver asks the live
# spawned-member roster BEFORE any tab map, so a live member is never mistaken for a lead no matter
# what its tab says. The label is a NAME, never a capability: what a pane may do is decided by the
# role the daemon resolves for it.
# Two things stop the tab-name convention from becoming a way to claim leadership: the configured
# member spaces are excluded from the scan, so nothing fleetd places can land in a matching tab;
# and startup REFUSES a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel`
# override, also matches — so the two namespaces cannot overlap by accident. The label is a NAME,
# never a capability: what a pane may do is decided by the role the daemon resolves for it.
# CB-551: IDLE-LEAD HEARTBEAT — nudge the single lead back to work when it has been continuously
# idle (no open fleet_send driving it) past the quiet period. The fleet is one lead + architects +
@@ -99,17 +98,17 @@ bind:
# contextHighNudge: false
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to end its
# own pane, launch a fresh one, and bootstrap that fresh session against the handover file.
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
# own pane and bootstrap a fresh session against that file.
#
# Opt-in on purpose — it tears down the lead's own pane on request, so upgrading the daemon must
# never acquire that ability for you. Absent block = feature off, and nothing is constructed at all.
# Even once present, nothing but an explicit confirm() call — one that passes every check — can ever
# tear a pane down: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so it
# cannot act on the pane inline (that pane is still WORKING); instead it schedules a one-shot
# Opt-in on purpose — it clears the lead's own pane on request, so upgrading the daemon must never
# acquire that ability for you. Absent block = feature off, and nothing is constructed at all. Even
# once present, nothing but an explicit confirm() call — one that passes every check — can ever
# cause a /clear: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so
# it cannot clear the pane inline (that pane is still WORKING); instead it schedules a one-shot
# continuation that waits for the SAME confirm() call's turn to end, then does the actual work. See
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order.
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order (fleetd #480 correction).
#
# handoverPath: REQUIRED when this block is present — where the handover file a fresh lead session
# reads must live. No default (an operator-specific path); a present block with no
@@ -118,30 +117,25 @@ bind:
# working directory when that lead has none configured) — never against whatever
# directory the daemon process happens to have been started in. An absolute path is
# used unchanged. Prefer an absolute path if the daemon and the lead's pane might not
# share a working directory.
# share a working directory (fleetd #480 follow-up).
# requireOperatorConfirm: true # default true — confirm() refuses unless the caller also passes
# # operatorConfirmed: true
# maxDocAgeSeconds: 3600 # default 3600 — refuse a handover file older than this
# turnSettleSeconds: 20 # default 20 — how long the deferred roll waits for the CALLING
# # lead's own turn to end (its pane to report injectable again)
# # before tearing the old pane down at all. If this elapses, nothing
# # is torn down — a lead that never goes idle is still doing real
# # work.
# relaunchReadySeconds: 45 # default 45 — bounds two later waits, after the old pane is gone
# # and a fresh one has been launched: first, for the fresh pane to
# # reach a real turn boundary (never sends bootstrapText if THIS one
# # elapses); second, for the new terminal to be recognised as this
# # lead (bootstrapText is sent either way once the first wait
# # passes). A separate, later pair of waits from turnSettleSeconds.
# # before sending /clear at all. If this elapses, /clear is NEVER
# # sent — a lead that never goes idle is still doing real work.
# clearSettleSeconds: 20 # default 20 — how long to wait for the pane to become injectable
# # again AFTER /clear before giving up (never sends bootstrapText
# # if this elapses). A separate, second wait from turnSettleSeconds.
# bootstrapText: "..." # default names the RESOLVED (absolute) handoverPath — sent to
# # the freshly relaunched lead's pane once it reaches a real turn
# # boundary
# # the lead once its pane settles after /clear
# leadRollover:
# handoverPath: /path/to/handover.md
# requireOperatorConfirm: true
# maxDocAgeSeconds: 3600
# turnSettleSeconds: 20
# relaunchReadySeconds: 45
# clearSettleSeconds: 20
# bootstrapText: "Fresh lead session: read the handover file and carry on."
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
@@ -665,10 +659,9 @@ fleet:
# tabPrefix: "lead:" # only used to guard against a worker tabLabel colliding with
# # this convention at startup; plays no part in matching a lead
# scanIntervalSeconds: 10 # rescan cadence, and the worst case before a new tab is seen
# workspace: leads # where a launched lead's tab is created (default "fleet",
# # the same shared space the members use). Sharing that space
# # with members is the normal shipped shape: the scanner tells
# # a lead from a member by the exact tab label, not by workspace.
# workspace: leads # where a launched lead's tab is created (default "leads").
# # MUST NOT be a member workspace — those are excluded from the
# # scan, so a lead placed in one is never found again.
# cwd: /path/to/repo # the launched lead's working directory (default: fleetd's own)
# kind: claude # descriptive; reported by fleet_whoami
# gpt-sol-5.6:
@@ -676,22 +669,6 @@ fleet:
# kind: opencode
# model: openai/gpt-5.6-terra
# A collaborator tab, keyed by name (fleetd #669). A pane whose tab matches resolves as the
# COLLABORATOR role: `fleet_whoami` answers `collaborator`, and the session may observe the fleet
# (`fleet_list`, `fleet_profiles`, `fleet_whoami`) and `fleet_send` to a lead or to another
# collaborator. It may NOT spawn, stop or drain anything, roll a lead's session, answer a
# member's `fleet_ask`, poll a ticket, send across hosts, or send to a spawned member's terminal.
# Each of those is refused at the gate rather than queued.
# Recognise-only, like a profile-less `leaders:` entry above: there is no `profile:`, no
# `instances:` and no `kind:`. `tab:` is REQUIRED and is the only field identity depends on,
# matched case-insensitively — the same GET-THE-VALUE-RIGHT warning above the `leaders:` block
# applies here too.
# A lead cannot discover a collaborator yet (#703): `fleet_list` has no `collaborators` key, so
# the collaborator must speak first, or pass on the `sessionId` its own `fleet_whoami` reports.
# collaborators:
# reviewer-alex:
# tab: "collab: alex"
# architects:
# architect-1:
# profile: opus # a strong model, on the operator's subscription
+2 -4
View File
@@ -189,10 +189,8 @@
<build>
<!-- CB-634: the cutover renamed the module dir (bridged/ -> fleetd/), the jar, and the
launchd plist together. fleetd #664: the installed plist and the systemd unit now name
fleetd/run/fleetd.jar, not this plugin's own output path — see
scripts/redeploy-fleetd.sh for the mv that gets a build from here to there. KeepAlive is
armed, so this name, the plist, and the wrapper must still move as one. -->
launchd plist together. The installed plist names fleetd/target/fleetd.jar and
KeepAlive is armed, so this name, the plist, and the wrapper must move as one. -->
<finalName>fleetd</finalName>
<plugins>
<plugin>
+34 -31
View File
@@ -217,27 +217,25 @@ public final class Fleetd {
/**
* The {@link Injector}'s readiness gate (CB-534): a target is deliverable if it is a spawned
* member whose agent has connected the bridge MCP, a lead, <em>or</em> a collaborator.
* member whose agent has connected the bridge MCP, <em>or</em> a lead.
*
* <p>The gate exists for one reason — to hold a delivery out of a <em>spawned</em> member's boot
* window, where herdr already reports {@code idle} but the TUI would drop an injected paste. That
* hazard is a property of spawning. A lead is never spawned: the operator started it and named it
* (or labelled its tab) only once it was up, so there is no boot window to guard. A collaborator
* is the same way — a person's own tab, matched to a configured name, never spawned.
* (or labelled its tab) only once it was up, so there is no boot window to guard.
*
* <p>Neither a lead nor a collaborator is ever enrolled in {@link MemberPresence} — {@code
* FleetMcp} marks presence only for a worker, an architect, or the unconfigured-pane floor,
* never for a lead or a collaborator. So without the second and third disjuncts a lead or
* collaborator is permanently un-deliverable: every send to one sat on the gate for
* {@code READINESS_GRACE_POLLS} (~60s) and then failed having never been typed into the pane.
* <p>A lead is also never enrolled in {@link MemberPresence} — {@code FleetMcp} marks presence
* for every spawned member (worker and architect), deliberately, since that map doubles as the
* member roster's availability signal and a lead counted there would show up as an available
* member. So without the second disjunct a lead is permanently un-deliverable: every
* lead→lead send sat on the gate for {@code READINESS_GRACE_POLLS} (~60s) and then failed
* having never been typed into the pane.
*
* <p>Both sets are read through their supplier on each call rather than snapshotted, so a lead or
* collaborator discovered by {@code leadScan} after startup becomes deliverable without a restart.
* <p>The lead set is read through the supplier on each call rather than snapshotted, so a lead
* discovered by {@code leadScan} after startup becomes deliverable without a restart.
*/
static Predicate<String> deliverableTo(MemberPresence presence, Supplier<Map<String, String>> leads,
Supplier<Map<String, String>> collaborators) {
return target -> presence.isPresent(target) || leads.get().containsKey(target)
|| collaborators.get().containsKey(target);
static Predicate<String> deliverableTo(MemberPresence presence, Supplier<Map<String, String>> leads) {
return target -> presence.isPresent(target) || leads.get().containsKey(target);
}
/**
@@ -941,27 +939,21 @@ public final class Fleetd {
* cfg.leadHeartbeat()}
* @param leadAgents the {@link AgentControl} instance that reaches the LEAD's pane (not
* {@code memberAgents}), normally {@code router.leadAgents()}
* @param leadSpaces the {@link WorkspaceControl} instance that reaches the LEAD's
* workspace, normally {@code router.leadSpaces()} — used to tear down
* a rolled lead's old pane and confirm it is gone
* @param launcher starts the fresh lead a roll relaunches once the old one is gone
* @param config the live {@link ConfigRef}, captured only inside the returned
* supplier and the two lookups below — never dereferenced here
* supplier and the workspace lookup — never dereferenced here
* @param liveLeadTerminals terminal id → lead NAME for every CURRENTLY recognised lead, normally
* the same {@code leads} supplier {@code main} already builds for
* {@code HerdrRouter}/{@link #leadSeatLookup} — never a value snapshot
* @return a constructed {@link LeadRollover}, or {@code null} when {@code leadRollover:} is
* absent from the startup config
*/
static LeadRollover leadRollover(FleetConfig cfg, AgentControl leadAgents,
WorkspaceControl leadSpaces, LeadLauncher launcher, ConfigRef config,
static LeadRollover leadRollover(FleetConfig cfg, AgentControl leadAgents, ConfigRef config,
Supplier<Map<String, String>> liveLeadTerminals) {
if (cfg.leadRollover() == null) {
return null;
}
Function<String, String> leadNameForTerminal = terminal -> liveLeadTerminals.get().get(terminal);
Function<String, String> leadWorkspace = terminal -> {
String leadName = leadNameForTerminal.apply(terminal);
String leadName = liveLeadTerminals.get().get(terminal);
if (leadName == null) {
return null;
}
@@ -969,8 +961,7 @@ public final class Fleetd {
FleetConfig.Leader leader = fleet == null ? null : fleet.leaders().get(leadName);
return leader == null ? null : leader.cwd();
};
return new LeadRollover(leadAgents, leadSpaces, launcher, () -> config.get().leadRollover(),
leadWorkspace, leadNameForTerminal, liveLeadTerminals);
return new LeadRollover(leadAgents, () -> config.get().leadRollover(), leadWorkspace);
}
/**
@@ -1132,12 +1123,17 @@ public final class Fleetd {
* @param liveLeadTerminals terminal id → lead name for every CURRENTLY recognised lead
* @param configDirForLeadName lead name → {@code configDir}, normally {@link
* #leadConfigDirLookup}'s return
* @param windowForLeadName lead name → that lead's profile's effective auto-compact window,
* normally {@link #leadContextWindowLookup}'s return, and passed
* through to {@link LeadContextGauge#read} so the heartbeat's own
* HIGH reading scales with that lead's real window, or {@code null}
* when it cannot be resolved — either way {@link LeadContextGauge}
* falls back to its own fixed HIGH threshold
*/
static Function<String, LeadContextGauge.Reading> leadContextLookup(LeadContextGauge gauge, AgentControl agents,
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName) {
return leadContextLookup(gauge, agents, liveLeadTerminals, configDirForLeadName, _ -> null);
}
/**
* As above, additionally resolving each lead's effective auto-compact window (normally {@link
* #leadContextWindowLookup}'s return) and passing it through to {@link LeadContextGauge#read},
* so the heartbeat's own HIGH reading scales with that lead's real window instead of always the
* gauge's fixed fallback.
*/
static Function<String, LeadContextGauge.Reading> leadContextLookup(LeadContextGauge gauge, AgentControl agents,
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName,
@@ -1166,6 +1162,13 @@ public final class Fleetd {
* factory {@code main} calls, rather than only a lookup nothing in {@code main} is proven to use
* (see {@code FleetdLeadConfigDirSourceWiringTest}'s javadoc for the measured gap this shape closes).
*/
static LeadHeartbeatLoop.LeadContextSource leadContextSource(LeadContextGauge gauge, AgentControl agents,
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName) {
return new LeadHeartbeatLoop.LeadContextSource(
leadContextLookup(gauge, agents, liveLeadTerminals, configDirForLeadName));
}
/** As above, additionally threading the effective-window lookup through. */
static LeadHeartbeatLoop.LeadContextSource leadContextSource(LeadContextGauge gauge, AgentControl agents,
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName,
Function<String, Long> windowForLeadName) {
@@ -43,7 +43,6 @@ import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.msg.ReplyPushLoop;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
@@ -142,12 +141,8 @@ final class FleetdAssembly {
? ports.connectHerdr(Path.of(cfg.memberHerdrSocket()))
: herdr;
AtomicReference<Supplier<Map<String, String>>> leadsRef = new AtomicReference<>(Map::of);
// fleetd #669 Unit E: a collaborator's pane is opened by a person, exactly like a lead's,
// so its terminal must also route to the lead herdr daemon rather than the member one.
AtomicReference<Supplier<Map<String, String>>> collaboratorTerminalsRef = new AtomicReference<>(Map::of);
HerdrRouter router = new HerdrRouter(herdr, memberHerdr,
target -> leadsRef.get().get().containsKey(target)
|| collaboratorTerminalsRef.get().get().containsKey(target));
target -> leadsRef.get().get().containsKey(target));
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
@@ -255,57 +250,32 @@ final class FleetdAssembly {
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
}
// CB-531/CB-579: discover leads by the tab labels the operator writes, one scanner per
// configured lead's own exact `tab:` label. fleetd #669: the same scan also recognises a
// configured collaborator's tab, so one herdr pass answers both.
// configured lead's own exact `tab:` label.
final Supplier<Map<String, String>> leads;
final Supplier<Map<String, String>> collaboratorTerminals;
var leaders = cfg.fleet().leaders();
var collaboratorsConfig = cfg.fleet().collaborators();
if (!leaders.isEmpty() || !collaboratorsConfig.isEmpty()) {
if (!leaders.isEmpty()) {
Map<String, String> tabToName = new LinkedHashMap<>();
leaders.forEach((name, leader) -> {
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
tabToName.put(leader.tab(), name);
}
});
Map<String, String> collaboratorTabToName = new LinkedHashMap<>();
collaboratorsConfig.forEach((name, collaborator) -> {
if (collaborator != null && collaborator.tab() != null && !collaborator.tab().isBlank()) {
collaboratorTabToName.put(collaborator.tab(), name);
}
});
// A collaborator-only fleet configures no `leaders:` entry to read a scan interval from
// — FleetConfig.Collaborator carries no scanIntervalSeconds of its own. Falling back to
// FleetConfig.Leader's own compact-constructor default keeps a collaborator-only
// deployment on the same rescan cadence as the default lead cadence, instead of
// inventing a second number for the same kind of scan.
int scanIntervalSeconds = leaders.isEmpty()
? 10
: leaders.values().iterator().next().scanIntervalSeconds();
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
// This must use the lead daemon: scanning member tabs would demote the lead to a worker.
LeadTabScanner scanner = new LeadTabScanner(herdr, tabToName, collaboratorTabToName, Set.of(),
leads = new LeadTabScanner(herdr, tabToName, Set.of(),
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), ports.nanoClock());
leads = scanner;
collaboratorTerminals = scanner::collaborators;
log.info("lead/collaborator scan: tabs {} host a lead, tabs {} host a collaborator "
+ "(rescan every {}s, shared fleet space)",
tabToName.keySet(), collaboratorTabToName.keySet(), scanIntervalSeconds);
log.info("lead scan: tabs {} host a lead (rescan every {}s, shared fleet space)",
tabToName.keySet(), scanIntervalSeconds);
} else {
leads = () -> leadTerminals;
collaboratorTerminals = Map::of;
}
leadsRef.set(leads);
collaboratorTerminalsRef.set(collaboratorTerminals);
// Constructed unconditionally — it is cheap and side-effect free — so a LeadRollover built
// below can relaunch a lead even on a boot where herdr was down for the ensureLeads() call.
LeadLauncher leadLauncher = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg);
// CB-558: start any declared lead that is not already running. After the scanner is built,
// and only when herdr answered — the launcher's whole safety property is that it can count
// live leads first, and must never guess and risk a second orchestrator.
if (herdrUp && !leaders.isEmpty()) {
int launched = leadLauncher.ensureLeads();
int launched = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg).ensureLeads();
if (launched > 0) {
log.info("lead auto-launch: {} lead(s) started", launched);
}
@@ -371,7 +341,7 @@ final class FleetdAssembly {
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
MemberPresence presence = sessions.asPresence();
TurnListener turnListener = Fleetd.turnListener(completion, sessions);
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads, collaboratorTerminals);
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads);
// fleetd #556: registration is wired directly to `completion`, not folded into the
// `turnListener` fan-out above — so it survives `sessions.onDelivered` (or any future
// listener) throwing, regardless of call order.
@@ -444,8 +414,7 @@ final class FleetdAssembly {
heartbeatScheduler.shutdownNow();
}
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed.
LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), router.leadSpaces(),
leadLauncher, config, leads);
LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), config, leads);
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
pushLoop, metrics);
@@ -483,11 +452,6 @@ final class FleetdAssembly {
ConnectionIdentity identity = new ConnectionIdentity(
new PaneLocator(herdr, memberHerdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// fleetd #669 Unit D: a live spawned member resolves as its own role, whatever a tab map
// says about the same terminal. fleetd #702: SessionManager.spawnedMemberRole also answers
// for a pane mid-teardown, not only one still in the registry — see its javadoc.
Function<String, MemberRole> spawnedMemberRole = sessions::spawnedMemberRole;
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
final CallerResolver callers;
@@ -497,13 +461,11 @@ final class FleetdAssembly {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting fleetd");
}
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members,
spawnedMemberRole, collaboratorTerminals);
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members,
spawnedMemberRole, collaboratorTerminals);
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
@@ -1,7 +1,5 @@
package dev.ltms.fleet.auth;
import java.util.function.Predicate;
/**
* The authorization table (CB-505), stated once and enforced on both entry paths.
*
@@ -22,22 +20,16 @@ public final class Authz {
SPAWN,
/** Tear a worker peer down. */
STOP,
/** Deliver a turn to a local session, addressed by {@code sessionId}. */
/** Deliver a turn to a session (or answer a worker's question). */
SEND,
/** Resolve a worker's blocked question and resume its turn, addressed by {@code turnId}. */
ANSWER,
/** Address a peer lead on another daemon over the coordination broker, by {@code coordId}. */
COORD_SEND,
/** A worker's terminal reply for its own turn. */
REPLY,
/** A worker's mid-turn question to the primary. */
ASK,
/** Collect held replies from a session's inbox. */
DRAIN,
/** Read-only roster, profile, and identity observation: no ticket, task, or turn state. */
/** Read-only observation: status, roster, profiles, task polling. */
READ,
/** Poll a ticket, or read a session's status. */
TASK_READ,
/**
* Read (never ack) this daemon's own held lead-to-lead coordination mail (fleetd #421).
*
@@ -61,98 +53,44 @@ public final class Authz {
HANDOVER
}
/**
* The fail-closed classifier: answers no for every target, so a collaborator's {@code SEND}
* is refused unless a caller supplies a real one. {@code CallerResolver#knownLeadOrCollaborator()}
* is the real one, read from the same lead and collaborator maps {@code CallerResolver#resolve}
* consults, so a target that classifier calls known is one {@code resolve} would actually
* resolve as a lead or collaborator.
*/
public static final Predicate<String> NO_KNOWN_LEAD_OR_COLLABORATOR = target -> false;
/**
* Convenience form for a caller with no classifier to supply. Fails closed: a collaborator's
* {@code SEND} is refused, as if no terminal were a configured lead or collaborator — the
* same decision {@link #NO_KNOWN_LEAD_OR_COLLABORATOR} gives explicitly. Every other action's
* result is identical to the four-argument form's, since none of them consult the classifier.
*
* <p>Its default classifier denies every collaborator, so a caller enforcing authorization
* must use the four-argument form instead.
*/
public static boolean permits(Principal caller, Action action, String targetSession) {
return permits(caller, action, targetSession, NO_KNOWN_LEAD_OR_COLLABORATOR);
}
/**
* Whether {@code caller} may perform {@code action} against {@code targetSession}.
*
* @param targetSession the session id in the request path; only consulted for the
* worker-scoped actions ({@code REPLY}, {@code ASK}) and for a
* collaborator's {@code SEND}, ignored otherwise, may be
* {@code null}
* @param knownLeadOrCollaborator whether a terminal is a configured lead or collaborator —
* consulted only for a collaborator's {@code SEND}, to confine
* it to another named peer and never a spawned member's
* terminal
* @param targetSession the session id in the request path; only consulted for the worker-scoped
* actions ({@code REPLY}, {@code ASK}), ignored otherwise, may be
* {@code null}
*/
public static boolean permits(Principal caller, Action action, String targetSession,
Predicate<String> knownLeadOrCollaborator) {
public static boolean permits(Principal caller, Action action, String targetSession) {
if (caller == null || caller.isAnonymous()) {
return false; // authenticated as nothing ⇒ authorized for nothing
}
return switch (action) {
// Fleet lifecycle is the primary's alone — spawn, stop, drain. An architect and a
// collaborator deliberately do NOT get these, so neither can tear down or stand up
// workers even though one of them coordinates them; and a worker driving any of these
// would be a worker escalating into the orchestrator role.
// Fleet lifecycle is the primary's alone — spawn, stop, drain. An architect
// deliberately does NOT get these (CB-548), so it cannot tear down or stand up workers
// even though it coordinates them; and a worker driving any of these would be a worker
// escalating into the orchestrator role.
case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary();
// Delivering a turn to a local session is open to the primary, the architect, and a
// collaborator whose target is itself a configured lead or collaborator: the architect
// delegates to workers (that is the role's point); a collaborator may reach only
// another named peer, never a spawned member's terminal. A worker is excluded —
// sending would be it escalating.
case SEND -> caller.isPrimary() || caller.isArchitect()
|| (caller.isCollaborator() && knownLeadOrCollaborator.test(targetSession));
// Resolving a worker's blocked question is part of delegating to it, open to the same
// two roles that may stand up that delegation in the first place. Not a collaborator:
// resuming another session's turn is lifecycle-adjacent, not peer messaging.
case ANSWER -> caller.isPrimary() || caller.isArchitect();
// Leaves the daemon over the coordination broker rather than addressing a local
// session, open to the same two roles as ANSWER. Not a collaborator: it is a
// local-tab peer with no cross-host route.
case COORD_SEND -> caller.isPrimary() || caller.isArchitect();
// Delivering a turn is open to the primary and the architect: an architect delegates
// to workers (that is the role's point) but still has no lifecycle rights. A worker is
// excluded — sending would be it escalating.
case SEND -> caller.isPrimary() || caller.isArchitect();
// The load-bearing rule: a caller acts only as the pane it occupies. CB-532 widened who
// that can be — a lead answering another lead is replying for its OWN terminal, which
// this already permits — while the rule itself is unchanged, and is what stops anyone
// forging a reply for a rendezvous someone else is waiting on. An architect's or a
// collaborator's own pane passes through the same check, so each can answer a funnel
// that delegated to it. An unnamed primary (token/loopback, no pane) owns nothing and
// is still excluded.
// forging a reply for a rendezvous someone else is waiting on. An architect's own pane
// passes through the same check, so it can answer a funnel that delegated to it. An
// unnamed primary (token/loopback, no pane) owns nothing and is still excluded.
case REPLY, ASK -> caller.ownsSession(targetSession);
// READ is roster, profile, and identity observation — fleet_list, fleet_profiles, and
// fleet_whoami — and carries no secrets: no ticket reply, no pending question, and no
// other session's turn state. Those live under TASK_READ. METRICS is the separate
// Prometheus scrape. Both are open to every authenticated role, including a
// collaborator and the unconfigured-pane floor.
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect()
|| caller.isCollaborator() || caller.isObserver();
// Ticket polling and session status, open to every role READ is open to except a
// collaborator or an observer. MessageService compares a ticket's creator to the
// caller on every read as well, so dropping this gate would not expose another
// session's reply — it would move the refusal later and widen what a caller that
// never orchestrates can probe.
case TASK_READ -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
// Observation is open to every authenticated role: a worker legitimately polls its own
// status, and the roster carries no secrets.
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
// fleetd #421: reading held lead-to-lead mail is the primary's alone. An architect
// holds READ today (CB-548), so "not primary" must mean not-architect here too — this
// is coordination between leads, not observation of the roster. The same reasoning
// excludes a collaborator.
// is coordination between leads, not observation of the roster.
case COORD_READ -> caller.isPrimary();
};
}
@@ -7,7 +7,6 @@ import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
import java.util.Map;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.function.Supplier;
/**
@@ -20,13 +19,6 @@ import java.util.function.Supplier;
*
* <p><strong>Resolution order</strong> — connection identity first, token second, nothing third:
* <ol>
* <li>A loopback peer PID that maps to a pane this gateway itself spawned ⇒ that member's own
* role: {@link Role#WORKER} for a dev, hunter, or reviewer; {@link Role#ARCHITECT} for an
* architect, but only while the live slot role still confirms it (fleetd #424 — a slot
* revoked from config demotes an already-bound session on its very next request, so the
* roster's own role is never granted on its word alone). No tab map is consulted — a live
* spawned member's identity comes from the registry that spawned it, never from a label a
* pane could also carry.</li>
* <li>A loopback peer PID that maps to a pane named by {@code leaders:}, by the legacy
* {@code primary.terminal} pin, or by an operator-labelled lead tab (CB-307, CB-530, CB-531)
* ⇒ {@link Role#PRIMARY}, carrying that lead's
@@ -36,11 +28,9 @@ import java.util.function.Supplier;
* so two leads can work as peers rather than one being demoted.</li>
* <li>A loopback peer PID that maps to a pane bound to a CB-548 architect slot ⇒
* {@link Role#ARCHITECT}, carrying the slot name. Just unforgeable as a worker's, and
* resolved from the <em>live</em> terminal→slot binding (never a request argument). This is
* the case the previous step does not catch: a binding with no live spawned-member session.</li>
* <li>A loopback peer PID that maps to an operator-labelled collaborator tab ⇒
* {@link Role#COLLABORATOR}, carrying that collaborator's name.</li>
* <li>A loopback peer PID that maps to any other herdr pane ⇒ {@link Role#OBSERVER}. This is
* resolved from the <em>live</em> terminal→slot binding (never a request argument), before
* the generic worker fallback.</li>
* <li>A loopback peer PID that maps to any other herdr pane ⇒ {@link Role#WORKER}. This is
* unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured
* regardless of auth mode, so enabling auth never breaks the fleet.</li>
* <li>Otherwise, under {@code token} mode, a valid bearer token ⇒ {@link Role#PRIMARY}.</li>
@@ -74,20 +64,6 @@ public final class CallerResolver {
private final Supplier<Map<String, String>> architectTerminals;
private final Function<String, MemberRole> memberSlotRoles;
private final Function<String, String> memberSlotNames;
/**
* terminal_id → the role of the live spawned member occupying it, or {@code null} for a
* terminal no spawned member occupies. Consulted first, ahead of every tab map: a live
* spawned member's identity is its own, whatever a tab map says about the same terminal.
* A function rather than the roster itself, so a resolve on the hot path never scans a list —
* the lookup strategy is the caller's to choose.
*/
private final Function<String, MemberRole> spawnedMemberRole;
/**
* terminal_id → collaborator name; empty when none are configured. A supplier for the same
* reason as {@link #leadTerminals}: a collaborator tab recognised after construction (the tab
* scan discovering a newly-labelled tab) takes effect without a restart.
*/
private final Supplier<Map<String, String>> collaboratorTerminals;
/** Loopback-trust resolver: no token required, historical behaviour. Test-only. */
CallerResolver(ConnectionIdentity identity) {
@@ -148,42 +124,17 @@ public final class CallerResolver {
/**
* Live registry form that can confirm a bound slot is an architect slot.
*
* <p>It keeps terminal bindings and slot roles in the same {@link MemberRegistry}, so a
* configured architect can resolve as an architect. No spawned-member roster or collaborator
* registry is consulted — equivalent to {@link #withLeadsAndMembers(ConnectionIdentity,
* boolean, String, Supplier, MemberRegistry, Function, Supplier)} with both absent. Kept for
* every caller that has neither to offer, so adding them did not churn every construction site.
* <p>This is the only public construction path. It keeps terminal bindings and slot roles in
* the same {@link MemberRegistry}, so a configured architect can resolve as an architect.
*/
public static CallerResolver withLeadsAndMembers(ConnectionIdentity identity,
boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
MemberRegistry members) {
return withLeadsAndMembers(identity, tokenMode, token, leadTerminals, members, null, null);
}
/**
* Live registry form that also resolves a live spawned member to its own role, and a
* configured collaborator tab to {@link Role#COLLABORATOR}.
*
* <p>This is the only public construction path that exercises the full resolution order.
*
* @param spawnedMemberRole terminal_id → the role of the live spawned member occupying
* it, or {@code null} for a terminal no spawned member occupies.
* {@code null} here means no roster is consulted at all (every
* terminal falls through to the tab maps), not that none matches.
* @param collaboratorTerminals terminal_id → collaborator name, live like {@code leadTerminals}
*/
public static CallerResolver withLeadsAndMembers(ConnectionIdentity identity,
boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
MemberRegistry members,
Function<String, MemberRole> spawnedMemberRole,
Supplier<Map<String, String>> collaboratorTerminals) {
return new CallerResolver(identity, tokenMode, token, leadTerminals,
members == null ? null : members::snapshot,
members == null ? null : members::roleForSlot,
members == null ? null : members::nameForSlot,
spawnedMemberRole, collaboratorTerminals);
members == null ? null : members::nameForSlot);
}
private static Supplier<Map<String, String>> fixed(Map<String, String> leadTerminals) {
@@ -209,17 +160,6 @@ public final class CallerResolver {
Supplier<Map<String, String>> architectTerminals,
Function<String, MemberRole> memberSlotRoles,
Function<String, String> memberSlotNames) {
this(identity, tokenMode, token, leadTerminals, architectTerminals, memberSlotRoles,
memberSlotNames, null, null);
}
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals,
Function<String, MemberRole> memberSlotRoles,
Function<String, String> memberSlotNames,
Function<String, MemberRole> spawnedMemberRole,
Supplier<Map<String, String>> collaboratorTerminals) {
if (tokenMode && (token == null || token.isBlank())) {
throw new IllegalArgumentException(
"auth.mode=token requires a non-empty token; check that the env var named by "
@@ -232,8 +172,6 @@ public final class CallerResolver {
this.architectTerminals = architectTerminals == null ? Map::of : architectTerminals;
this.memberSlotRoles = memberSlotRoles == null ? _ -> null : memberSlotRoles;
this.memberSlotNames = memberSlotNames == null ? Function.identity() : memberSlotNames;
this.spawnedMemberRole = spawnedMemberRole == null ? _ -> null : spawnedMemberRole;
this.collaboratorTerminals = collaboratorTerminals == null ? Map::of : collaboratorTerminals;
}
/**
@@ -250,24 +188,14 @@ public final class CallerResolver {
}
/**
* The currently-recognised collaborator tabs, {@code terminal_id → name}.
* The currently-recognised architect slots, {@code terminal_id → slot name} (CB-548).
*
* <p>Read from the same supplier {@link #resolve} consults, for the reason given in
* {@link #leads()}. Live for the same reason as {@link #leads()}.
* <p>Read from the same supplier {@link #resolve} consults, so a slot that is <em>listed</em>
* here but would not <em>resolve</em> (or the reverse) cannot drift apart. Live for the same
* reason as {@link #leads()}.
*/
public Map<String, String> collaborators() {
return collaboratorTerminals.get();
}
/**
* Whether {@code target} names a terminal this resolver would resolve as a lead or a
* collaborator — the classifier a collaborator's {@code SEND} is checked against, read from the
* exact maps {@link #resolve} consults so a target that would resolve as a lead or collaborator
* is never the one a collaborator is refused to reach, or the reverse.
*/
public Predicate<String> knownLeadOrCollaborator() {
return target -> leadTerminals.get().containsKey(target)
|| collaboratorTerminals.get().containsKey(target);
public Map<String, String> members() {
return architectTerminals.get();
}
/**
@@ -280,25 +208,6 @@ public final class CallerResolver {
public Principal resolve(String remoteAddr, int remotePort, String authorizationHeader) {
ConnectionIdentity.Caller c = identity.resolve(remoteAddr, remotePort);
if (c.terminal() != null) {
MemberRole spawnedRole = spawnedMemberRole.apply(c.terminal());
if (spawnedRole != null) {
// A live spawned member occupies this pane. Its identity is its own, whatever a tab
// map says about the same terminal — checked before every tab map, consulting none
// of them, so a tab label can never override a roster entry for the same terminal.
if (spawnedRole == MemberRole.ARCHITECT) {
// The roster only answers THAT this pane is a live spawned member; config still
// decides WHAT that member's slot grants (fleetd #424). A slot revoked after the
// bind must still demote this session on its very next request, so the roster's
// own ARCHITECT role is confirmed against the live slot role, exactly as the
// architect-slot step below confirms a binding with no live member session.
String slot = architectTerminals.get().get(c.terminal());
if (slot != null && memberSlotRoles.apply(slot) == MemberRole.ARCHITECT) {
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
}
return Principal.worker(c.terminal(), c.pid());
}
return Principal.worker(c.terminal(), c.pid());
}
String lead = leadTerminals.get().get(c.terminal());
if (lead != null) {
// The config names this pane as a lead's own. The pane mapping is exactly as
@@ -312,18 +221,11 @@ public final class CallerResolver {
// The config/live binding names this pane as an architect slot's own. Same
// unforgeable pane mapping; the live binding, never a request argument, decides.
// Check the slot role too: this defence in depth prevents a bad lifecycle bind from
// escalating a dev, hunter or reviewer into an architect. This is the case the
// spawned-member step above does not catch: a binding with no live member session.
// escalating a dev, hunter or reviewer into an architect. Checked before
// the worker fallback.
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
}
String collaborator = collaboratorTerminals.get().get(c.terminal());
if (collaborator != null) {
// An operator-labelled collaborator tab, confirmed live by the same scan that
// confirms a lead tab. Checked last among the tab maps so a pane also matching one
// of the above keeps that stronger role.
return Principal.collaborator(collaborator, c.terminal(), c.pid());
}
return Principal.observer(c.terminal(), c.pid()); // unforgeable; never token-gated
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
}
if (tokenMode) {
@@ -11,9 +11,7 @@ package dev.ltms.fleet.auth;
* @param pid the connecting process id, or {@code -1} when not resolvable (audit context)
* @param name for a lead resolved from the CB-530 {@code leaders:} registry, which lead it is;
* for an architect resolved from the CB-548 {@code architects:} registry, which
* slot it occupies; for a collaborator resolved from the {@code collaborators:}
* registry, which collaborator it is; {@code null} for every other caller,
* including an unnamed primary
* slot it occupies; {@code null} for every other caller, including an unnamed primary
*/
public record Principal(Role role, String terminal, long pid, String name) {
@@ -75,26 +73,6 @@ public record Principal(Role role, String terminal, long pid, String name) {
return new Principal(Role.ARCHITECT, terminal, pid, slotName);
}
/**
* A collaborator: a human-opened tab recognised by its exact label in the
* {@code collaborators:} registry.
*
* <p>Carries {@link Role#COLLABORATOR}. {@code name} is reporting only — it lets
* {@code fleet_whoami} say which collaborator is asking. Identity is the {@code terminal}:
* like a worker's it comes from the connection, so {@code ownsSession} works exactly as it
* does for a worker — a collaborator acts as its own pane and no other.
*/
public static Principal collaborator(String name, String terminal, long pid) {
return new Principal(Role.COLLABORATOR, terminal, pid, name);
}
/**
* The unconfigured-pane floor: a loopback caller whose pane matched no other role.
*/
public static Principal observer(String terminal, long pid) {
return new Principal(Role.OBSERVER, terminal, pid);
}
public boolean isPrimary() {
return role == Role.PRIMARY;
}
@@ -103,23 +81,15 @@ public record Principal(Role role, String terminal, long pid, String name) {
return role == Role.ARCHITECT;
}
public boolean isCollaborator() {
return role == Role.COLLABORATOR;
}
public boolean isWorker() {
return role == Role.WORKER;
}
public boolean isObserver() {
return role == Role.OBSERVER;
}
/**
* Whether this caller is a spawned member with its own pane.
*
* <p>Both workers and architects are spawned members. A lead is not: it is a peer the
* operator started and named, never a pane this daemon spawned.
* <p>Both workers and architects are spawned members. A lead is excluded because recording it
* as present would count it as an available member in the roster.
*/
public boolean isSpawnedMember() {
return role == Role.WORKER || role == Role.ARCHITECT;
@@ -145,32 +115,11 @@ public record Principal(Role role, String terminal, long pid, String name) {
return terminal != null && terminal.equals(sessionId);
}
/**
* Stable identity used to own tickets and open turns. The unnamed primary has no owner key so
* it can use the message layer's primary-wide ticket access rule.
*/
public String ownerKey() {
return switch (role) {
case PRIMARY -> name == null ? null : prefixed("leader", name);
case WORKER -> prefixed("worker", terminal);
case ARCHITECT -> prefixed("architect", terminal);
case COLLABORATOR -> prefixed("collaborator", name);
case OBSERVER -> prefixed("observer", terminal);
case ANONYMOUS -> "anonymous";
};
}
private static String prefixed(String role, String identity) {
return role + ":" + identity;
}
/** Short, non-sensitive description for audit lines and error details. */
public String describe() {
return switch (role) {
case WORKER -> "worker:" + terminal;
case ARCHITECT -> "architect:" + name;
case COLLABORATOR -> "collaborator:" + name;
case OBSERVER -> "observer:" + terminal;
case PRIMARY -> name == null ? "primary" : "leader:" + name;
case ANONYMOUS -> "anonymous";
};
@@ -34,28 +34,6 @@ public enum Role {
*/
ARCHITECT,
/**
* A config-declared, human-opened tab recognised by its exact label (the {@code
* fleet.collaborators.<name>.tab} registry). Never spawned — identity comes from the
* connection, never a request argument, exactly like {@link #WORKER} and {@link #ARCHITECT}.
* May {@code SEND} only to a configured lead or collaborator, {@code REPLY}/{@code ASK} only
* as its own pane, and {@code READ}/{@code METRICS}; may not {@code SPAWN}/{@code STOP}/
* {@code DRAIN}/{@code HANDOVER}, poll a ticket ({@code TASK_READ}), or reach the
* coordination broker ({@code COORD_SEND}/{@code COORD_READ}).
*/
COLLABORATOR,
/**
* A loopback pane that resolved to none of the roles above: not a live spawned member, not a
* configured lead, not a bound architect slot, not a configured collaborator tab. Unforgeable
* like a worker's — derived from the connection's pane, never from a request argument, and
* honoured regardless of auth mode. May {@code READ} and {@code METRICS}, and {@code REPLY}/
* {@code ASK} only as its own pane; may not {@code SPAWN}/{@code STOP}/{@code DRAIN}/
* {@code HANDOVER}, {@code SEND}, poll a ticket ({@code TASK_READ}), or reach the coordination
* broker ({@code COORD_SEND}/{@code COORD_READ}).
*/
OBSERVER,
/** Authenticated as nothing. Authorized for nothing but {@code /healthz}. */
ANONYMOUS
}
@@ -67,7 +67,7 @@ import java.util.function.Supplier;
* {@code models:} above: {@code dev.ltms.fleet.lead.LeadRollover} holds a
* {@code Supplier<FleetConfig.LeadRollover>} (the same {@code () -> config.get().x()} shape)
* and reads {@code handoverPath}/{@code requireOperatorConfirm}/{@code maxDocAgeSeconds}/
* {@code turnSettleSeconds}/{@code relaunchReadySeconds}/{@code bootstrapText} fresh on every
* {@code turnSettleSeconds}/{@code clearSettleSeconds}/{@code bootstrapText} fresh on every
* {@code open()}/{@code confirm()} call (and on the deferred post-{@code confirm()}
* continuation fleetd #480's correction added — see {@code LeadRollover}'s class doc) rather
* than capturing them into fields at construction — unlike its closest
@@ -615,7 +615,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
// may bind to, AND what a slot already bound still grants) through its own instance of that
// same supplier shape — see MemberRegistry.live and its class doc for the binding rule:
// removing a slot revokes ARCHITECT on the bound pane's very next request, and only the slot
// OCCUPANCY survives, so the demoted session keeps its slot key until it unbinds.
// OCCUPANCY survives, so the demoted session keeps its slot key until it unbinds. Only
// fleet.leaders is frozen (Fleetd.java:281 reads cfg.fleet().leaders() off the startup
// snapshot to build both the LeadTabScanner's tab-label-to-name map, wired into
// CallerResolver.withLeadsAndMembers at Fleetd.java:620/624, and — when herdr answered —
@@ -635,11 +635,6 @@ public final class ConfigRef implements Supplier<FleetConfig> {
+ "live through that same supplier for placement AND through a separate supplier "
+ "on MemberRegistry for spawn-time identity — both already applied");
}
if (!Objects.equals(collaboratorsOf(old), collaboratorsOf(fresh))) {
changed.add("fleet: fleet.collaborators (each collaborator's tab) is read once at "
+ "startup to build the LeadTabScanner's identity map, which is not rebuilt on "
+ "reload, so a collaborator added, removed, or given a new tab: needs a restart");
}
// Kept in step with SPLIT_KEYS the same way changedColdKeys is kept in step with COLD_KEYS —
// every message here must be traceable to one of the split keys the class doc documents.
// NOTE what this does NOT prove, per the javadoc above: it does not catch a SPLIT_KEYS
@@ -655,11 +650,6 @@ public final class ConfigRef implements Supplier<FleetConfig> {
return cfg.fleet() == null ? Map.of() : cfg.fleet().leaders();
}
/** {@code cfg.fleet().collaborators()}, defensively, in case a caller hands in a non-defaulted config. */
private static Map<String, FleetConfig.Collaborator> collaboratorsOf(FleetConfig cfg) {
return cfg.fleet() == null ? Map.of() : cfg.fleet().collaborators();
}
/**
* {@link FleetConfig.Profile} record components deliberately left out of
* {@link #sameLaunchSettings} because they are read <em>live</em>, not baked in at spawn — see
@@ -1129,8 +1129,11 @@ public record FleetConfig(
* {@code tab} can never be discovered, launched or not
* @param instances how many of this lead should be live (default 1). The daemon
* launches only the shortfall, so a restart adopts rather than doubles
* @param tabPrefix lead-tab naming convention checked against member labels. Lead
* identity uses {@code tab}. Default {@code "lead:"}
* @param tabPrefix no longer used to find a lead's tab — {@code tab} is matched
* exactly. Its only remaining job is the startup collision guard
* ({@link #validateLeadTabPrefixes()}), which still uses it to refuse
* a worker {@code tabLabel} template that could be misread as a lead.
* Default {@code "lead:"}
* @param scanIntervalSeconds how long a tab scan is cached before herdr is asked again; also the
* worst case before a newly-labelled tab is recognised. Default 10
* @param kind which agent runs there ({@code claude}, {@code opencode}, …)
@@ -1176,25 +1179,6 @@ public record FleetConfig(
}
}
/**
* A tab fleetd recognises as a collaborator, keyed by name (fleetd #669).
*
* <p>Recognise-only: there is no {@code profile}, no {@code instances} and no {@code kind}.
* Nothing here ever launches a pane.
*
* <p>{@code tabPrefix} is absent. Identity is matched on the exact {@code tab} alone.
*
* @param tab the exact tab label hosting this collaborator, matched case-insensitively; the
* only field identity depends on. Required — an entry with no {@code tab} can
* never be discovered.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Collaborator(String tab) {
public Collaborator {
tab = (tab == null || tab.isBlank()) ? null : tab.strip();
}
}
/**
* One entry of a {@code fleet:} role pool — a role paired with the backend it runs on.
*
@@ -1232,18 +1216,15 @@ public record FleetConfig(
* is exactly compatible with that. The pool is also what replaced {@code defaultProfile:} — an
* unqualified spawn names a role, and the role's pool supplies the candidates.
*
* @param leaders panes that orchestrate rather than are orchestrated, keyed by lead name
* @param architects profiles the {@code architect} role may run on
* @param developers profiles the {@code dev} role may run on
* @param hunters profiles the {@code hunter} role may run on
* @param reviewers profiles the {@code reviewer} role may run on
* @param charters optional launch-charter text keyed by singular role wire name
* @param tabLabel template for a member tab's label; {@code {role}}, {@code {profile}},
* {@code {model}} and {@code {n}} (a per role+profile counter) are
* substituted. Default {@link #DEFAULT_TAB_LABEL}
* @param collaborators tabs fleetd recognises as collaborators (fleetd #669), keyed by name.
* Recognise-only, exactly like a {@code profile}-less {@link Leader}:
* nothing here is ever auto-launched.
* @param leaders panes that orchestrate rather than are orchestrated, keyed by lead name
* @param architects profiles the {@code architect} role may run on
* @param developers profiles the {@code dev} role may run on
* @param hunters profiles the {@code hunter} role may run on
* @param reviewers profiles the {@code reviewer} role may run on
* @param charters optional launch-charter text keyed by singular role wire name
* @param tabLabel template for a member tab's label; {@code {role}}, {@code {profile}},
* {@code {model}} and {@code {n}} (a per role+profile counter) are
* substituted. Default {@link #DEFAULT_TAB_LABEL}
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Fleet(Map<String, Leader> leaders,
@@ -1252,11 +1233,13 @@ public record FleetConfig(
Map<String, Slot> hunters,
Map<String, Slot> reviewers,
Map<String, String> charters,
String tabLabel,
Map<String, Collaborator> collaborators) {
String tabLabel) {
/**
* Role first, so the tab bar identifies the member's fleet role.
* Role first, so the tab bar reads as the fleet and so the label shares a namespace with a
* lead's {@code tabPrefix}. Because {@code {role}} comes from a closed enum, a generated
* member label can never begin with {@code "lead:"} — the clash that
* {@link #validateLeadTabPrefixes()} used to have to check for is unrepresentable here.
*/
public static final String DEFAULT_TAB_LABEL = "{role}: {profile} #{n}";
@@ -1268,30 +1251,26 @@ public record FleetConfig(
reviewers = unmodifiableOrEmpty(reviewers);
charters = unmodifiableOrEmpty(charters);
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? DEFAULT_TAB_LABEL : tabLabel;
collaborators = unmodifiableOrEmpty(collaborators);
}
/**
* A fleet with no configured launch charters and no collaborators — the shape every
* deployment had before CB-566, and what most tests want.
* A fleet with no configured launch charters — the shape every deployment had before
* CB-566, and what most tests want.
*
* <p>Kept deliberately, even though an overload that drops a new field is normally the
* shape to avoid. It is safe here because nothing <em>reads</em> a charter through a
* constructor: the launcher reads {@code fleet.charters()} from the live config. Jackson
* binds the canonical constructor, so this one cannot swallow an operator's YAML.
* {@code collaborators} is dropped the same way and for the same reason: no caller of
* this overload has ever needed to set it, so it defaults to empty here exactly as the
* canonical constructor would default an absent YAML key.
*/
public Fleet(Map<String, Leader> leaders, Map<String, Slot> architects,
Map<String, Slot> developers, Map<String, Slot> reviewers,
Map<String, String> charters, String tabLabel) {
this(leaders, architects, developers, null, reviewers, charters, tabLabel, null);
this(leaders, architects, developers, null, reviewers, charters, tabLabel);
}
public Fleet(Map<String, Leader> leaders, Map<String, Slot> architects,
Map<String, Slot> developers, Map<String, Slot> reviewers, String tabLabel) {
this(leaders, architects, developers, null, reviewers, null, tabLabel, null);
this(leaders, architects, developers, null, reviewers, null, tabLabel);
}
/**
@@ -1416,16 +1395,15 @@ public record FleetConfig(
* before anything exists to call — the same fact already true of adding a brand-new
* {@code profiles:} entry.
*
* <p><strong>{@code turnSettleSeconds}:</strong> {@code confirm()} is called FROM the calling
* lead's own turn, so its pane is still {@code WORKING} the instant {@code confirm()} validates
* every gate and schedules the roll. {@code dev.ltms.fleet.lead.LeadRollover}'s deferred
* continuation waits up to this many seconds for that SAME pane to report {@code IDLE} or
* {@code DONE} — i.e. for the calling turn to actually end — before it ends the old pane's
* process at all. {@code BLOCKED} does not count: that is a live turn merely paused, not one
* that has finished. If that wait times out, the old pane is never touched: a lead that never
* goes idle is still doing real work, and the roll ends that pane's whole process — there is no
* way back from this once it runs, so this wait is the only thing standing between "still
* working" and "gone".
* <p><strong>{@code turnSettleSeconds} (fleetd #480 correction):</strong> {@code confirm()} is
* called FROM the calling lead's own turn, so its pane is still {@code WORKING} the instant
* {@code confirm()} validates every gate and schedules the roll. {@code
* dev.ltms.fleet.lead.LeadRollover}'s deferred continuation waits up to this many seconds for
* that SAME pane to report an injectable state again — i.e. for the calling turn to actually
* end — before it sends {@code /clear} at all. If that wait times out, no {@code /clear} is
* ever sent: a lead that never goes idle is still doing real work, and clearing it would
* destroy live context. This is a separate wait from {@code clearSettleSeconds} below, which
* bounds the SECOND wait, for the pane to re-settle AFTER {@code /clear} has already gone out.
*
* @param handoverPath required when this block is present — where the handover file a fresh
* lead session reads must live. There is no sane non-null default for an
@@ -1444,51 +1422,37 @@ public record FleetConfig(
* @param maxDocAgeSeconds default 3600 — refuse a handover file whose modified time is older
* than this many seconds, so a stale leftover from an earlier rollover
* attempt can never be mistaken for a fresh one.
* @param turnSettleSeconds default 300 — bound on how long the deferred roll waits for the
* CALLING lead's own turn to end (its pane to report {@code IDLE} or
* {@code DONE}) before ending that pane's process at all. See the
* paragraph above.
* @param relaunchReadySeconds default 45 — bound on EACH of two separate waits that run after
* the old lead's pane has been torn down and a fresh one launched: first,
* for the fresh pane itself to reach a real turn boundary ({@code IDLE} or
* {@code DONE}, never merely {@code BLOCKED}) — the safety gate, since
* typing into a pane that has not finished booting loses the keystrokes;
* second, for the fresh terminal to show up as a recognised lead, which is
* bookkeeping rather than a safety gate, so a timeout on this second wait
* does not withhold {@code bootstrapText} — it is sent once the pane is
* ready regardless. Recognition comes from the same periodically-refreshed
* scan {@code LeadTabScanner} already keeps ({@code scanIntervalSeconds},
* 10s live), so a budget has to clear more than one scan interval to leave
* any real margin for the CLI's own boot time; 20 was rejected for exactly
* that reason — at a 10s scan interval it only buys two scans. 45 buys
* roughly four. Only a timeout on the FIRST wait (the pane never becomes
* ready) withholds {@code bootstrapText}.
* @param turnSettleSeconds default 20 — bound on how long the deferred roll waits for the
* CALLING lead's own turn to end (its pane to report injectable again)
* before sending {@code /clear} at all. See the paragraph above.
* @param clearSettleSeconds default 20 — bound on how long to wait for the lead's pane to
* report an injectable state again after {@code /clear} before giving up. A
* roll that times out here never sends {@code bootstrapText}.
* @param bootstrapText default a sentence naming the RESOLVED handover path — sent to the
* fresh lead's pane once it reaches a real turn boundary after relaunch,
* telling the fresh session where to read the handover and carry on. Left
* {@code null} here when the operator configures none: the default sentence
* cannot be built at construction time because it must name the path AFTER
* {@code dev.ltms.fleet.lead.LeadRollover#open} has resolved a relative
* {@code handoverPath} against the calling lead's workspace, which this
* record has no way to know — see {@link #bootstrapTextFor(String)}.
* lead's pane once it settles after {@code /clear}, telling the fresh
* session where to read the handover and carry on. Left {@code null} here
* when the operator configures none: the default sentence cannot be built
* at construction time because it must name the path AFTER {@code
* dev.ltms.fleet.lead.LeadRollover#open} has resolved a relative {@code
* handoverPath} against the calling lead's workspace, which this record has
* no way to know — see {@link #bootstrapTextFor(String)}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record LeadRollover(String handoverPath, Boolean requireOperatorConfirm,
Integer maxDocAgeSeconds, Integer turnSettleSeconds,
Integer relaunchReadySeconds, String bootstrapText) {
Integer clearSettleSeconds, String bootstrapText) {
public LeadRollover {
requireOperatorConfirm = requireOperatorConfirm == null || requireOperatorConfirm;
maxDocAgeSeconds = (maxDocAgeSeconds == null || maxDocAgeSeconds <= 0) ? 3600 : maxDocAgeSeconds;
turnSettleSeconds = (turnSettleSeconds == null || turnSettleSeconds <= 0) ? 300 : turnSettleSeconds;
relaunchReadySeconds = (relaunchReadySeconds == null || relaunchReadySeconds <= 0)
? 45 : relaunchReadySeconds;
turnSettleSeconds = (turnSettleSeconds == null || turnSettleSeconds <= 0) ? 20 : turnSettleSeconds;
clearSettleSeconds = (clearSettleSeconds == null || clearSettleSeconds <= 0) ? 20 : clearSettleSeconds;
bootstrapText = (bootstrapText == null || bootstrapText.isBlank()) ? null : bootstrapText;
}
/**
* The text actually sent to the fresh lead's pane once it reaches a real turn boundary
* after relaunch: the operator's configured {@link #bootstrapText} when one is set,
* otherwise the default sentence built from {@code resolvedHandoverPath}.
* The text actually sent to the lead's pane once it settles after {@code /clear}: the
* operator's configured {@link #bootstrapText} when one is set, otherwise the default
* sentence built from {@code resolvedHandoverPath}.
*
* @param resolvedHandoverPath the ABSOLUTE path {@code dev.ltms.fleet.lead.LeadRollover
* #open} already resolved — never the raw configured {@link
@@ -1931,7 +1895,6 @@ public record FleetConfig(
rejectNegativeMaxLoad(yaml);
rejectAutoCompactWindowOutOfRange(yaml);
warnConflictingAutoCompactWindows(yaml);
warnRetiredClearSettleSecondsKey(yaml);
rejectMalformedProfilePatterns(yaml);
rejectUnknownKind(yaml);
rejectUnknownAuthMode(yaml);
@@ -1950,7 +1913,7 @@ public record FleetConfig(
/** The {@code fleet:} child blocks whose direct children are slot names. */
private static final Set<String> FLEET_POOL_KEYS =
Set.of("leaders", "architects", "developers", "hunters", "reviewers", "collaborators");
Set.of("leaders", "architects", "developers", "hunters", "reviewers");
/**
* Reject a {@code fleet:} role pool whose slot names repeat (CB-548, re-homed by CB-557).
@@ -1960,7 +1923,7 @@ public record FleetConfig(
* daemon would never know. Jackson's YAML parser does not fail on duplicate mapping keys by
* default, so duplicates are caught here, at parse time, before the map is built.
*
* <p>Only the six pools <em>directly under the top-level {@code fleet:}</em> are considered,
* <p>Only the five pools <em>directly under the top-level {@code fleet:}</em> are considered,
* and only their direct child keys (the slot names). A nested field elsewhere, even one also
* named {@code developers:}, is ignored, so parsing of the rest of the config is unaffected.
*
@@ -2355,33 +2318,6 @@ public record FleetConfig(
names, String.join(", ", detail));
}
/**
* Warn when a {@code leadRollover:} block still sets the retired {@code clearSettleSeconds}
* key. {@link LeadRollover} carries {@code @JsonIgnoreProperties(ignoreUnknown = true)} and no
* longer declares that component, so Jackson drops it with no signal of its own — this raw-YAML
* check is the only place an operator's now-inert setting is reported at all; by the time a
* {@link LeadRollover} instance exists to run a validator against, the key is already gone.
*
* @param yaml the raw config text
*/
static void warnRetiredClearSettleSecondsKey(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return;
}
if (raw == null || !(raw.get("leadRollover") instanceof Map<?, ?> leadRollover)) {
return;
}
if (leadRollover.containsKey("clearSettleSeconds")) {
log.warn("leadRollover.clearSettleSeconds is retired and no longer read. Set "
+ "leadRollover.relaunchReadySeconds instead: it bounds how long to wait, after "
+ "a lead is relaunched, for its pane to become ready and then for it to be "
+ "recognised as a lead. Remove clearSettleSeconds from fleetd.yaml.");
}
}
/**
* Reject a profile whose {@code errorPattern} (fleetd #201 Unit 5) or {@code exhaustedPattern}
* (CB-578 stage A) is not a valid Java regex, naming the profile, the key, and the parser's own
@@ -2746,21 +2682,30 @@ public record FleetConfig(
}
/**
* Reject a member tab-label template that could render as a configured lead or collaborator
* tab or match a lead-tab naming convention, and reject two {@code fleet.leaders} or
* {@code fleet.collaborators} entries — across either registry — that share one exact tab.
* Reject a lead-scan convention that a worker tab would also satisfy (CB-531).
*
* <p>{@code fleet.collaborators} has no {@code tabPrefix}: identity is matched on the exact
* {@code tab} alone, so only the exact-render check applies there, not the prefix check.
* <p>The scan reads a tab label and concludes "a lead lives here". fleetd also <em>writes</em>
* tab labels — every member gets one rendered into its tab. Choose a lead {@code tabPrefix} that
* a member template matches and the daemon starts labelling its own members as leads, promoting
* the entire fleet to {@link dev.ltms.fleet.auth.Role#PRIMARY} with no message and no diff.
* The member-space exclusion in {@link dev.ltms.fleet.herdr.LeadTabScanner} already blocks the
* realistic path, but defence that depends on one workspace label holding is not defence enough
* for a privilege boundary.
*
* @throws IllegalStateException when the fleet template or a profile {@code tabLabel} override
* can render as a configured lead or collaborator tab or match a
* lead-tab prefix, or when two entries — of either registry, or
* one of each — carry the same exact {@code tab}
* (case-insensitively)
* <p>CB-557 shrank this check rather than removing it. The default template is
* {@code "{role}: {profile} #{n}"} and {@code {role}} comes from a closed enum, so a
* <em>generated</em> label can no longer collide by construction. What remains checkable is what
* an operator still writes by hand: the {@code fleet.tabLabel} template and any per-profile
* {@code tabLabel} override.
*
* <p>Fatal rather than a warning, unlike {@link #warnUnknownTopLevelKeys}: an unknown key means
* a feature does nothing, while this means a feature does the opposite of what it says.
*
* @throws IllegalStateException when the fleet template or any profile's {@code tabLabel}
* override starts with a configured lead prefix
*/
public void validateLeadTabPrefixes() {
if (fleet == null) {
if (fleet == null || fleet.leaders().isEmpty()) {
return;
}
List<String> bad = new ArrayList<>();
@@ -2768,193 +2713,28 @@ public record FleetConfig(
if (leader == null) {
return;
}
String tab = leader.tab();
String prefix = leader.tabPrefix();
if (templateCanRenderAs(fleet.tabLabel(), tab)) {
bad.add("fleet.tabLabel=\"" + fleet.tabLabel() + "\" can render as the tab of "
+ "lead '" + leadName + "' (\"" + tab + "\")");
} else if (startsWithIgnoreCase(fleet.tabLabel(), prefix)) {
// The fleet-wide template is checked once per prefix: it labels every member that has no
// override, so one bad template promotes the entire fleet, not one profile.
if (startsWithIgnoreCase(fleet.tabLabel(), prefix)) {
bad.add("fleet.tabLabel=\"" + fleet.tabLabel() + "\" starts with the tabPrefix of "
+ "lead '" + leadName + "' (\"" + prefix + "\")");
}
profiles().entrySet().stream()
.filter(e -> startsWithIgnoreCase(e.getValue().tabLabel(), prefix))
.map(Map.Entry::getKey)
.sorted()
.forEach(p -> {
String label = profiles().get(p).tabLabel();
if (templateCanRenderAs(label, tab)) {
bad.add("profile '" + p + "' overrides tabLabel with \"" + label
+ "\", which can render as the tab of lead '" + leadName
+ "' (\"" + tab + "\")");
} else if (startsWithIgnoreCase(label, prefix)) {
bad.add("profile '" + p + "' overrides tabLabel with \"" + label
+ "\", which starts with the tabPrefix of lead '" + leadName
+ "' (\"" + prefix + "\")");
}
});
.forEach(p -> bad.add("profile '" + p + "' overrides tabLabel with \""
+ profiles().get(p).tabLabel() + "\", which starts with the tabPrefix of "
+ "lead '" + leadName + "' (\"" + prefix + "\")"));
});
fleet.collaborators().forEach((collabName, collaborator) -> {
if (collaborator == null) {
return;
}
String tab = collaborator.tab();
if (templateCanRenderAs(fleet.tabLabel(), tab)) {
bad.add("fleet.tabLabel=\"" + fleet.tabLabel() + "\" can render as the tab of "
+ "collaborator '" + collabName + "' (\"" + tab + "\")");
}
profiles().entrySet().stream()
.map(Map.Entry::getKey)
.sorted()
.forEach(p -> {
String label = profiles().get(p).tabLabel();
if (templateCanRenderAs(label, tab)) {
bad.add("profile '" + p + "' overrides tabLabel with \"" + label
+ "\", which can render as the tab of collaborator '"
+ collabName + "' (\"" + tab + "\")");
}
});
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join("; ", bad)
+ ". A member labelled that way, while its pane carries no entry in the "
+ "spawned-member roster, is read back as a lead or collaborator and granted "
+ "that identity's authority. Change one of the two so member tabs cannot be "
+ "confused with a lead's or collaborator's tab.");
}
List<String> collisions = new ArrayList<>();
List<String> leadNames = fleet.leaders().keySet().stream().sorted().toList();
for (int i = 0; i < leadNames.size(); i++) {
String nameA = leadNames.get(i);
Leader a = fleet.leaders().get(nameA);
if (a == null || a.tab() == null || a.tab().isBlank()) {
continue;
}
for (int j = i + 1; j < leadNames.size(); j++) {
String nameB = leadNames.get(j);
Leader b = fleet.leaders().get(nameB);
if (b == null || b.tab() == null || b.tab().isBlank()) {
continue;
}
if (a.tab().equalsIgnoreCase(b.tab())) {
collisions.add("lead '" + nameA + "' and lead '" + nameB + "' both use tab \""
+ a.tab() + "\"");
}
}
}
List<String> collabNames = fleet.collaborators().keySet().stream().sorted().toList();
for (int i = 0; i < collabNames.size(); i++) {
String nameA = collabNames.get(i);
Collaborator a = fleet.collaborators().get(nameA);
if (a == null || a.tab() == null || a.tab().isBlank()) {
continue;
}
for (int j = i + 1; j < collabNames.size(); j++) {
String nameB = collabNames.get(j);
Collaborator b = fleet.collaborators().get(nameB);
if (b == null || b.tab() == null || b.tab().isBlank()) {
continue;
}
if (a.tab().equalsIgnoreCase(b.tab())) {
collisions.add("collaborator '" + nameA + "' and collaborator '" + nameB
+ "' both use tab \"" + a.tab() + "\"");
}
}
}
for (String leadName : leadNames) {
Leader lead = fleet.leaders().get(leadName);
if (lead == null || lead.tab() == null || lead.tab().isBlank()) {
continue;
}
for (String collabName : collabNames) {
Collaborator collaborator = fleet.collaborators().get(collabName);
if (collaborator == null || collaborator.tab() == null
|| collaborator.tab().isBlank()) {
continue;
}
if (lead.tab().equalsIgnoreCase(collaborator.tab())) {
collisions.add("lead '" + leadName + "' and collaborator '" + collabName
+ "' both use tab \"" + lead.tab() + "\"");
}
}
}
if (collisions.isEmpty()) {
return;
}
throw new IllegalStateException("refusing to start: " + String.join("; ", collisions)
+ ". Tab identity is matched exactly, so only one of two entries sharing a tab can "
+ "ever be found — the other is silently unreachable. Give each lead and "
+ "collaborator its own exact tab.");
}
private static boolean templateCanRenderAs(String template, String tab) {
if (template == null || template.isBlank() || tab == null || tab.isBlank()) {
return false;
}
var placeholders = Pattern.compile("\\{(?:role|profile|model|n)}").matcher(template);
StringBuilder expression = new StringBuilder("^");
int literalStart = 0;
while (placeholders.find()) {
expression.append(Pattern.quote(template.substring(literalStart, placeholders.start())));
expression.append(".*");
literalStart = placeholders.end();
}
expression.append(Pattern.quote(template.substring(literalStart))).append("$");
return Pattern.compile(expression.toString(), Pattern.CASE_INSENSITIVE).matcher(tab).matches();
}
/** Case-insensitive prefix test that tolerates a null or blank label. */
private static boolean startsWithIgnoreCase(String label, String prefix) {
if (label == null || prefix == null || prefix.isBlank()) {
return false;
}
String stripped = label.strip();
return stripped.regionMatches(true, 0, prefix, 0, prefix.length());
}
/**
* Reject a profile that places its members by {@code "pane"} while any {@code fleet.leaders}
* or {@code fleet.collaborators} entry names a {@code tab}. A pane-placed member lands inside
* the focused tab rather than its own, so it can land inside a lead's or collaborator's own
* labelled tab. {@link dev.ltms.fleet.herdr.LeadTabScanner} identifies a lead or collaborator
* purely by that tab's label — it does not exclude the member space — so a member that ends up
* there, while its pane carries no entry in the spawned-member roster, is read back as that
* lead or collaborator and granted that identity's authority.
*
* <p>Only an entry with a non-blank {@code tab} is in scope: one with no {@code tab} feeds
* nothing into {@link dev.ltms.fleet.herdr.LeadTabScanner}, so it creates no hazard here.
*
* @throws IllegalStateException when any {@code profiles:} entry is pane-placed while any
* {@code fleet.leaders} or {@code fleet.collaborators} entry
* names a non-blank {@code tab}
*/
public void validatePanePlacementAgainstLeadTabs() {
if (fleet == null) {
return;
}
boolean anyLeaderHasTab = fleet.leaders().values().stream()
.anyMatch(leader -> leader != null && leader.tab() != null && !leader.tab().isBlank());
boolean anyCollaboratorHasTab = fleet.collaborators().values().stream()
.anyMatch(c -> c != null && c.tab() != null && !c.tab().isBlank());
if (!anyLeaderHasTab && !anyCollaboratorHasTab) {
return;
}
List<String> bad = new ArrayList<>();
profiles().entrySet().stream()
.filter(e -> !e.getValue().tabPlacement())
.map(Map.Entry::getKey)
.sorted()
.forEach(bad::add);
if (bad.isEmpty()) {
return;
}
throw new IllegalStateException("refusing to start: profile(s) " + bad
+ " use placement: pane while fleet.leaders or fleet.collaborators names a tab. A "
+ "pane-placed member can land inside that labelled tab, and while its pane "
+ "carries no entry in the spawned-member roster, it is read back as the lead or "
+ "collaborator and granted that identity's authority. Set placement: tab for "
+ "each named profile, or remove the tab from every fleet.leaders and "
+ "fleet.collaborators entry.");
throw new IllegalStateException("refusing to start: " + String.join("; ", bad)
+ ". Every member labelled that way would be read back as a lead and granted "
+ "spawn/stop/send on the whole fleet. Change one of the two so member tabs and "
+ "lead tabs cannot be confused.");
}
/**
@@ -2977,6 +2757,15 @@ public record FleetConfig(
}
}
/** Case-insensitive prefix test that tolerates a null/blank label. */
private static boolean startsWithIgnoreCase(String label, String prefix) {
if (label == null || prefix == null || prefix.isBlank()) {
return false;
}
String stripped = label.strip();
return stripped.regionMatches(true, 0, prefix, 0, prefix.length());
}
/**
* Reject a subscription profile whose {@code env:} block tries to reseat the Anthropic binding
* (CB-542).
@@ -3061,14 +2850,8 @@ public record FleetConfig(
* so duplicates are unrepresentable by construction once loaded — and {@link #load(Path)}
* already rejects a duplicated slot name at parse time, before the map collapses.
*
* <p>Also rejects a {@code fleet.collaborators} entry with no (or a blank) {@code tab}. A
* {@code profile}-less lead is still useful recognise-only — {@code tab} is the only field
* that matters to it either way. A collaborator carries no other field at all, so a blank
* {@code tab} leaves nothing for the entry to mean.
*
* @throws IllegalStateException when a slot names no profile or an unknown one, when a lead
* can be neither found nor created, or when a collaborator names
* no tab, naming the offending entry
* @throws IllegalStateException when a slot names no profile or an unknown one, or when a lead
* can be neither found nor created, naming the offending entry
*/
public void validateMembers() {
if (fleet == null) {
@@ -3106,16 +2889,6 @@ public record FleetConfig(
+ "auto-launched, labelled) purely by its tab, so every entry must name one.");
}
});
fleet.collaborators().forEach((name, collaborator) -> {
if (collaborator == null) {
return;
}
if (collaborator.tab() == null || collaborator.tab().isBlank()) {
bad.add("fleet.collaborators." + name + " has no tab: — a collaborator is "
+ "recognised purely by its tab, and carries no other field, so every "
+ "entry must name one.");
}
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
}
@@ -3164,11 +2937,11 @@ public record FleetConfig(
* Runs every validator this class declares — found by reflection, not by name.
*
* <p>fleetd ticket "central allow-list of usable models", follow-up: mutation testing found
* that although each validator above was well pinned on its own, nothing proved
* that although each of the six validators above was well pinned on its own, nothing proved
* either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()}) still
* invoked it — deleting a call site left the full suite green. The fix is not one more test
* per caller; a hand-maintained list of names here would have the exact same defect its
* own javadoc would warn against: the next validator someone adds has no reason
* invoked it — deleting a call site left the full suite green. The fix is not a seventh test
* per caller; a hand-maintained list of six names here would have the exact same defect its
* own javadoc would warn against: the seventh validator someone adds next month has no reason
* to be added to it. So this method does not name any validator. It sweeps {@link
* #getClass()}'s own public, no-argument, {@code void} methods whose name starts with {@code
* "validate"} (excluding itself) and invokes every one it finds, via {@link
@@ -3177,7 +2950,7 @@ public record FleetConfig(
* which it silently never runs.
*
* <p>{@code Fleetd.main} and {@link ConfigRef#reload()} each call this one method instead of
* each validator individually — see the comments at those two call sites for why
* the six individually — see the comments at those two call sites for why
* each must run it.
*
* <p>Methods run in a fixed (alphabetical) order, so a config with more than one violation
@@ -3194,9 +2967,9 @@ public record FleetConfig(
/**
* The reflective sweep behind {@link #validateAll()}, kept as its own method — taking any
* {@code target}, not just {@code this} — so a test can prove the MECHANISM is generic (it
* would sweep any new {@code validateXxx()} method added to any class, not just something
* special-cased to the set {@link FleetConfig} declares today) without needing to add a real,
* unwanted extra validator to this class just to exercise that claim. See {@code
* would sweep a seventh {@code validateXxx()} method added to any class, not just something
* special-cased to today's six on {@link FleetConfig}) without needing to add a real, unwanted
* seventh validator to this class just to exercise that claim. See {@code
* FleetConfigValidateAllTest} for that proof.
*
* @param target an object whose public, no-argument, {@code void} methods named {@code
@@ -11,18 +11,12 @@ public final class HerdrRouter implements AutoCloseable {
private final AgentControl memberAgents;
private final WorkspaceControl leadSpaces;
private final WorkspaceControl memberSpaces;
private final Predicate<String> routeToLead;
private final Predicate<String> isLead;
/**
* @param routeToLead true for a terminal whose pane lives in the lead herdr daemon — a lead's
* own pane or a configured collaborator's, both opened by a person at a
* terminal rather than spawned, so both are found in the lead daemon rather
* than the member one
*/
public HerdrRouter(HerdrClient lead, HerdrClient member, Predicate<String> routeToLead) {
public HerdrRouter(HerdrClient lead, HerdrClient member, Predicate<String> isLead) {
this.lead = Objects.requireNonNull(lead, "lead");
this.member = member != null ? member : lead;
this.routeToLead = Objects.requireNonNull(routeToLead, "routeToLead");
this.isLead = Objects.requireNonNull(isLead, "isLead");
leadAgents = new AgentControl(this.lead);
memberAgents = this.member == this.lead ? leadAgents : new AgentControl(this.member);
leadSpaces = new WorkspaceControl(this.lead);
@@ -33,7 +27,7 @@ public final class HerdrRouter implements AutoCloseable {
public WorkspaceControl leadSpaces() { return leadSpaces; }
public AgentControl memberAgents() { return memberAgents; }
public WorkspaceControl memberSpaces() { return memberSpaces; }
public AgentControl agentsFor(String targetId) { return routeToLead.test(targetId) ? leadAgents : memberAgents; }
public AgentControl agentsFor(String targetId) { return isLead.test(targetId) ? leadAgents : memberAgents; }
HerdrClient leadClient() { return lead; }
HerdrClient memberClient() { return member; }
@@ -37,11 +37,8 @@ import java.util.function.Supplier;
* <p><strong>Direction of trust.</strong> The label names the lead; it never <em>grants</em>
* anything a pane could take for itself. Three properties keep that honest:
* <ol>
* <li>{@code excludedWorkspaceLabels} can filter a workspace out of the scan, but this class does
* not by itself stop a worker from landing inside a matching tab — a caller may pass an empty
* set, and the daemon does. The guard against that is {@code
* FleetConfig.validatePanePlacementAgainstLeadTabs}: it refuses, at startup, any profile that
* places members by pane while a lead names a tab.</li>
* <li>Worker spaces are excluded wholesale ({@code excludedWorkspaceLabels}), so a worker cannot
* become a lead by being placed — as a split, say — inside a matching tab.</li>
* <li>A worker cannot rename a tab: {@code tab.rename} is reachable only through
* {@link WorkspaceControl}, which no {@code fleet_*} tool exposes. The label is writable by
* the human at the terminal and by nobody the bridge is defending against.</li>
@@ -96,19 +93,13 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
private static final Logger log = LoggerFactory.getLogger(LeadTabScanner.class);
/** What a matched tab names: a lead or a collaborator. */
private enum Kind { LEAD, COLLABORATOR }
/** One matched tab's name and what it names. */
private record Entry(String name, Kind kind) {}
private final HerdrClient herdr;
private final Map<String, Entry> tabToEntry;
private final Map<String, String> tabToName;
private final Set<String> excludedWorkspaceLabels;
private final long ttlNanos;
private final LongSupplier clock;
private Map<String, Entry> cached = Map.of();
private Map<String, String> cached = Map.of();
private long scannedAtNanos;
private boolean everScanned;
@@ -132,53 +123,26 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
*/
public LeadTabScanner(HerdrClient herdr, Map<String, String> tabToName,
Set<String> excludedWorkspaceLabels, long ttlNanos, LongSupplier clock) {
this(herdr, tabToName, Map.of(), excludedWorkspaceLabels, ttlNanos, clock);
}
/**
* As {@link #LeadTabScanner(HerdrClient, Map, Set, long, LongSupplier)}, additionally scanning
* for configured collaborator tabs in the same pass.
*
* @param collaboratorTabToName every configured collaborator's exact tab label → its name
* ({@code fleet.collaborators.<name>.tab}), matched the same way as
* {@code tabToName}
*/
public LeadTabScanner(HerdrClient herdr, Map<String, String> tabToName,
Map<String, String> collaboratorTabToName,
Set<String> excludedWorkspaceLabels, long ttlNanos, LongSupplier clock) {
this.herdr = herdr;
this.tabToEntry = buildTabIndex(tabToName, collaboratorTabToName);
this.tabToName = normalize(tabToName);
this.excludedWorkspaceLabels = excludedWorkspaceLabels == null
? Set.of() : Set.copyOf(excludedWorkspaceLabels);
this.ttlNanos = ttlNanos;
this.clock = clock;
}
/**
* Keys stripped and lower-cased once, so every lookup is a plain map hit. Leads and
* collaborators merge into a single index, so {@link #scan()} matches both kinds in one pass
* over the tab list; a label naming both a lead and a collaborator takes the lead entry —
* leads are put last, so a colliding key's lead entry is the one that overwrites — since a lead
* can already do everything a collaborator can. Config validation already refuses a lead and a
* collaborator sharing one exact tab, so this ordering is defence in depth, not the control.
*/
private static Map<String, Entry> buildTabIndex(Map<String, String> tabToName,
Map<String, String> collaboratorTabToName) {
Map<String, Entry> out = new LinkedHashMap<>();
putNormalized(out, collaboratorTabToName, Kind.COLLABORATOR);
putNormalized(out, tabToName, Kind.LEAD);
return Collections.unmodifiableMap(out);
}
private static void putNormalized(Map<String, Entry> out, Map<String, String> tabToName, Kind kind) {
if (tabToName == null) {
return;
/** Keys stripped and lower-cased once, so every lookup is a plain map hit. */
private static Map<String, String> normalize(Map<String, String> tabToName) {
if (tabToName == null || tabToName.isEmpty()) {
return Map.of();
}
Map<String, String> out = new LinkedHashMap<>();
tabToName.forEach((tab, name) -> {
if (tab != null && !tab.isBlank() && name != null && !name.isBlank()) {
out.put(tab.strip().toLowerCase(Locale.ROOT), new Entry(name, kind));
out.put(tab.strip().toLowerCase(Locale.ROOT), name);
}
});
return Collections.unmodifiableMap(out);
}
/**
@@ -189,29 +153,6 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
*/
@Override
public synchronized Map<String, String> get() {
return byKind(refresh(), Kind.LEAD);
}
/**
* The current {@code terminal_id → collaborator name} map, sharing the same scan and cache as
* {@link #get()} — both kinds are matched in one pass, so this never costs a second herdr call.
*/
public synchronized Map<String, String> collaborators() {
return byKind(refresh(), Kind.COLLABORATOR);
}
private static Map<String, String> byKind(Map<String, Entry> entries, Kind kind) {
Map<String, String> out = new LinkedHashMap<>();
entries.forEach((terminal, entry) -> {
if (entry.kind() == kind) {
out.put(terminal, entry.name());
}
});
return Collections.unmodifiableMap(out);
}
/** Rescans if the cache has expired, otherwise returns the cached answer. */
private Map<String, Entry> refresh() {
long now = clock.getAsLong();
if (everScanned && now - scannedAtNanos < ttlNanos) {
return cached;
@@ -221,21 +162,21 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
scannedAtNanos = now;
everScanned = true;
try {
Map<String, Entry> fresh = scan();
Map<String, String> fresh = scan();
if (!fresh.equals(cached)) {
log.info("lead/collaborator panes: {}", fresh);
log.info("lead panes: {}", fresh);
}
cached = fresh;
} catch (HerdrException e) {
log.warn("lead-tab scan failed, keeping the {} entr(y/ies) already known: {}",
log.warn("lead-tab scan failed, keeping the {} lead(s) already known: {}",
cached.size(), e.getMessage());
}
return cached;
}
/** One full pass: labelled tabs → live agents in them → those panes' terminals. */
private Map<String, Entry> scan() {
Map<String, Entry> entryByTab = new LinkedHashMap<>();
private Map<String, String> scan() {
Map<String, String> nameByTab = new LinkedHashMap<>();
for (JsonNode w : herdr.call("workspace.list").path("workspaces")) {
Workspace ws = Workspace.from(w);
if (ws.workspaceId() == null || excludedWorkspaceLabels.contains(ws.label())) {
@@ -243,22 +184,21 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
}
for (JsonNode t : herdr.call("tab.list", Map.of("workspace_id", ws.workspaceId())).path("tabs")) {
Tab tab = Tab.from(t);
Entry entry = entryOf(tab.label());
if (entry != null && tab.tabId() != null) {
entryByTab.put(tab.tabId(), entry);
String name = leadNameOf(tab.label());
if (name != null && tab.tabId() != null) {
nameByTab.put(tab.tabId(), name);
}
}
}
if (entryByTab.isEmpty()) {
if (nameByTab.isEmpty()) {
gracedTerminals = Set.of();
return Map.of();
}
// fleetd #359: a labelled tab is only a lead (or collaborator) when herdr also reports a
// running agent in it — the same liveness signal LeadLauncher.countLeads trusts for the
// identical purpose. Without this, a tab left behind by a session that has since died reads
// as live forever.
// fleetd #359: a labelled tab is only a lead when herdr also reports a running agent in
// it — the same liveness signal LeadLauncher.countLeads trusts for the identical purpose.
// Without this, a tab left behind by a session that has since died reads as live forever.
Set<String> tabsWithAgent = new HashSet<>();
for (JsonNode a : herdr.call("agent.list").path("agents")) {
String tabId = a.path("tab_id").asText(null);
@@ -267,18 +207,18 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
}
}
Map<String, Entry> byTerminal = new LinkedHashMap<>();
Map<String, String> byTerminal = new LinkedHashMap<>();
Set<String> stillGraced = new HashSet<>();
// One pane.list for every tab: panes carry tab_id, so the join is local.
for (JsonNode p : herdr.call("pane.list", Map.of()).path("panes")) {
String tabId = p.path("tab_id").asText(null);
Entry entry = entryByTab.get(tabId);
String name = nameByTab.get(tabId);
String terminal = p.path("terminal_id").asText(null);
if (entry == null || terminal == null || terminal.isBlank()) {
if (name == null || terminal == null || terminal.isBlank()) {
continue;
}
if (tabsWithAgent.contains(tabId)) {
byTerminal.put(terminal, entry);
byTerminal.put(terminal, name);
continue;
}
// No agent reported for this tab, but its tab/pane are still here — this is the
@@ -287,7 +227,7 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
// reported as live; a terminal we never reported live gets none, so the original #359
// fix (a genuinely dead tab is never reported) is unaffected for the common case.
if (cached.containsKey(terminal) && !gracedTerminals.contains(terminal)) {
byTerminal.put(terminal, entry);
byTerminal.put(terminal, name);
stillGraced.add(terminal);
}
}
@@ -296,19 +236,18 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
}
/**
* The entry a tab label declares, or {@code null} if it names neither a configured lead nor a
* configured collaborator.
* The lead name a tab label declares, or {@code null} if it names none of the configured leads.
*
* <p>Exact match (case-insensitive, ends stripped) against {@link #tabToEntry} — no prefix
* <p>Exact match (case-insensitive, ends stripped) against {@link #tabToName} — no prefix
* stripping, so an operator's {@code "lead: something-else"} tab is never mistaken for a
* configured lead just because it shares a prefix. The match strips a trailing
* {@link PendingCloseMarker} first, so a tab {@code LeadLauncher} has flagged as maybe-dead but
* not yet closed keeps resolving normally while that reconcile is pending.
*/
private Entry entryOf(String label) {
private String leadNameOf(String label) {
if (label == null) {
return null;
}
return tabToEntry.get(PendingCloseMarker.strip(label).toLowerCase(Locale.ROOT));
return tabToName.get(PendingCloseMarker.strip(label).toLowerCase(Locale.ROOT));
}
}
@@ -1,146 +0,0 @@
package dev.ltms.fleet.herdr;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.charset.StandardCharsets;
import java.util.List;
import java.util.function.IntFunction;
/**
* The three protections every {@code agent.start} caller needs against herdr's pane-typed launch
* surface (fleetd #220, #727): a byte-limit check on the assembled command line, a bounded retry
* on {@code agent_pane_busy} (the target pane's shell has not reached its prompt yet), and a
* bounded retry on {@code agent_name_taken} with a fresh name each attempt. One implementation —
* every caller of {@code agent.start}, lead or member, goes through this seam rather than carrying
* its own copy.
*/
public final class ResilientAgentLaunch {
private static final Logger log = LoggerFactory.getLogger(ResilientAgentLaunch.class);
private ResilientAgentLaunch() {
}
/**
* The pty line buffer herdr types a launch command into: BSD/macOS {@code MAX_CANON}. Not a
* fleetd choice and not configurable — see {@link #checkFits}.
*/
public static final int PANE_COMMAND_BYTE_LIMIT = 1024;
/** Per-argument allowance for the separating space and a shell quote pair fleetd cannot see. */
private static final int QUOTING_OVERHEAD_PER_ARG = 3;
/** herdr rejects a duplicate agent {@code name}; a caller retries a bumped name this many times. */
public static final int NAME_RETRIES = 8;
/**
* Retries for {@code agent.start} against a seed pane whose shell has not reached its prompt
* yet — {@code tab.create}/{@code pane.split} return as soon as the pane exists, and herdr
* refuses to start an agent in a pane that is not "an available shell" ({@code agent_pane_busy}).
*/
public static final int SHELL_READY_RETRIES = 20;
/** Raised by {@link #checkFits} when the assembled command cannot fit the pane line. */
public static final class TooLargeException extends RuntimeException {
public TooLargeException(String message) {
super(message);
}
}
/**
* Verify the assembled launch line fits the pane herdr types it into. herdr does not exec the
* launch command — it TYPES it into the pane as one line, and a pty line buffer holds only
* {@value #PANE_COMMAND_BYTE_LIMIT} bytes. Everything past that byte is dropped with no error
* anywhere: herdr answers "agent started", the backend exits on the mangled argument it was
* handed, the pane closes, and the only symptom is a readiness timeout with no reason. That is
* how fleetd #214 broke every claude-code spawn — one 50-byte flag pushed a 978-byte command to
* 1028, and the tail that got cut was {@code --autocompact 250000}.
*
* <p>So measure it here and refuse, loudly and immediately, rather than start something that
* cannot work. The estimate is deliberately conservative: fleetd cannot see herdr's quoting, so
* every argument is charged its own bytes plus a separator and a quote pair. An over-estimate
* costs a clear error at a length that was already unsafe; an under-estimate would let the
* silent truncation back in.
*
* @param label names the launch in the refusal message (a profile name)
* @param argv the full argv, including the executable at index 0
* @throws TooLargeException naming the limit, the estimate, and the longest argument
*/
public static void checkFits(String label, List<String> argv) {
int bytes = 0;
String longest = null;
int longestBytes = 0;
for (String arg : argv) {
int argBytes = arg == null ? 0 : arg.getBytes(StandardCharsets.UTF_8).length;
bytes += argBytes + QUOTING_OVERHEAD_PER_ARG;
if (argBytes > longestBytes) {
longestBytes = argBytes;
longest = arg;
}
}
if (bytes <= PANE_COMMAND_BYTE_LIMIT) {
return;
}
String culprit = longest == null ? "<none>"
: longest.substring(0, Math.min(longest.length(), 60)) + (longest.length() > 60 ? "…" : "");
throw new TooLargeException(
"launch command for " + label + " is about " + bytes + " bytes, over the "
+ PANE_COMMAND_BYTE_LIMIT + "-byte limit of the pane line herdr types it into. "
+ "The pty would drop the tail silently and the backend would exit on a mangled "
+ "argument. Longest argument is " + longestBytes + " bytes: " + culprit
+ " — move it off the command line (a file flag) or shorten it.");
}
/**
* Start an agent into {@code paneId}, retrying {@code agent_pane_busy} up to {@code retries}
* times with {@code sleeper} run between attempts.
*
* @throws HerdrException the last {@code agent_pane_busy} failure once {@code retries} is
* spent, or immediately for any other herdr failure
*/
public static Agent startAwaitingShellPrompt(AgentControl agents, String name, String kind,
List<String> args, String paneId,
int retries, Runnable sleeper) {
HerdrException busy = null;
for (int attempt = 0; attempt < retries; attempt++) {
try {
return agents.start(name, kind, args, paneId);
} catch (HerdrException e) {
if (!"agent_pane_busy".equals(e.code())) throw e;
log.debug("pane {} not at its shell prompt yet, retrying agent.start", paneId);
busy = e;
sleeper.run();
}
}
throw busy;
}
/**
* Start an agent under a freshly generated name each attempt, retrying {@code agent_name_taken}
* up to {@code nameRetries} times — herdr refuses a duplicate {@code name} outright, so a stale
* registry entry (a crashed session, a name the registry has not yet released) must not block a
* legitimate relaunch. Each attempt also carries its own {@link #startAwaitingShellPrompt} retry.
*
* @param nameForAttempt called once per attempt (0-based) to produce that attempt's name
* @throws HerdrException the last {@code agent_name_taken} failure once {@code nameRetries} is
* spent, or immediately for any other herdr failure
*/
public static Agent startUniquelyNamed(AgentControl agents, String kind, List<String> args,
String paneId, IntFunction<String> nameForAttempt,
int nameRetries, int shellReadyRetries, Runnable sleeper) {
HerdrException last = null;
for (int attempt = 0; attempt < nameRetries; attempt++) {
String name = nameForAttempt.apply(attempt);
try {
return startAwaitingShellPrompt(agents, name, kind, args, paneId,
shellReadyRetries, sleeper);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("agent name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
}
@@ -68,10 +68,8 @@ import java.util.function.LongSupplier;
* (never the whole 52 MB a long-lived transcript reaches on the host this was measured on), and
* {@link #DEFAULT_CACHE_TTL_MILLIS} bounds how often that bounded read actually happens — a burst
* of {@code fleet_list} calls inside one TTL window reads the file once. One instance's cache is
* keyed by {@code (configDir, sessionId, highThreshold)}, so it is safe to share across every lead
* a single {@code fleet_list} call reports on, and a call that resolves a different effective
* window for the same lead never reads back a state computed against the other window's
* threshold.
* keyed by {@code (configDir, sessionId)}, so it is safe to share across every lead a single
* {@code fleet_list} call reports on.
*/
public final class LeadContextGauge {
@@ -172,6 +170,12 @@ public final class LeadContextGauge {
* {@code "claude"} (including {@code null}, meaning undetected) reports
* {@link State#UNKNOWN} — this reader only understands Claude Code's own
* transcript format
*/
public Reading read(String configDir, String sessionId, String agentType) {
return read(configDir, sessionId, agentType, null);
}
/**
* @param effectiveWindowTokens the caller's resolved effective auto-compact window for this
* lead's own profile, or {@code null} when it cannot be resolved.
* HIGH fires at {@link #HIGH_THRESHOLD_FRACTION} of this value;
@@ -188,14 +192,13 @@ public final class LeadContextGauge {
String base = (configDir == null || configDir.isBlank())
? System.getProperty("user.home") + "/.claude"
: configDir;
long highThreshold = highThreshold(effectiveWindowTokens);
String cacheKey = base + '\u0000' + sessionId + '\u0000' + highThreshold;
String cacheKey = base + '\u0000' + sessionId;
long now = clock.getAsLong();
CacheEntry cached = cache.get(cacheKey);
if (cached != null && now - cached.readAtMillis() < ttlMillis) {
return cached.reading();
}
Reading fresh = readUncached(base, sessionId, highThreshold);
Reading fresh = readUncached(base, sessionId, highThreshold(effectiveWindowTokens));
cache.put(cacheKey, new CacheEntry(fresh, now));
return fresh;
}
@@ -5,7 +5,6 @@ import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.PendingCloseMarker;
import dev.ltms.fleet.herdr.ResilientAgentLaunch;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl;
@@ -13,7 +12,6 @@ import dev.ltms.fleet.launch.ClaudeCodeArguments;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
@@ -21,7 +19,6 @@ import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.atomic.AtomicLong;
import dev.ltms.fleet.peer.PeerLauncher;
/**
@@ -83,19 +80,9 @@ public final class LeadLauncher {
private static final Logger log = LoggerFactory.getLogger(LeadLauncher.class);
/** Attempts {@link #relaunch(String)} makes before giving up and returning {@code null}. */
static final int RELAUNCH_ATTEMPTS = 3;
private final AgentControl agents;
private final WorkspaceControl spaces;
private final FleetConfig cfg;
private final Runnable sleeper;
// Per-process token mixed into each lead agent name so a fresh daemon process (seq back at 0)
// cannot collide with a same-name lead that outlived a restart — the same scheme
// HerdrPeerLauncher uses for members (fleetd #727).
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
private final AtomicLong nameSeq = new AtomicLong();
/**
* @param agents herdr agent control (start, list)
@@ -103,27 +90,9 @@ public final class LeadLauncher {
* @param cfg the loaded config — {@code fleet.leaders}, {@code profiles} and each lead's tab
*/
public LeadLauncher(AgentControl agents, WorkspaceControl spaces, FleetConfig cfg) {
this(agents, spaces, cfg, () -> sleepUninterruptibly(300));
}
/**
* Test seam: as above, plus an injectable {@code sleeper} for the {@code agent_pane_busy}
* retry (fleetd #727), so a test can prove the retry budget without a real sleep.
*/
LeadLauncher(AgentControl agents, WorkspaceControl spaces, FleetConfig cfg, Runnable sleeper) {
this.agents = agents;
this.spaces = spaces;
this.cfg = cfg;
this.sleeper = sleeper;
}
/** Uninterruptible sleep — the production {@link #sleeper} between {@code agent_pane_busy} retries. */
private static void sleepUninterruptibly(long ms) {
try {
Thread.sleep(ms);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}
/**
@@ -204,14 +173,24 @@ public final class LeadLauncher {
log.info("lead '{}': {} live, {} wanted — nothing to start", name, running, wanted);
continue;
}
if (!lead.isCreatable()) {
// A lead with a `tab:` but no `profile:` is recognise-only by design: the operator
// opens it by hand. Say so once rather than looking like a silent failure.
log.info("lead '{}' is not live, and names no profile — it can be recognised but not "
+ "launched. Add `profile:` under fleet.leaders.{} to have fleetd start it.",
name, name);
continue;
}
ResolvedLead resolved = resolveLaunchable(name);
if (resolved == null) {
FleetConfig.Profile profile = cfg.profiles().get(lead.profile());
if (profile == null) {
log.warn("lead '{}' names profile '{}', which is not configured — not launching",
name, lead.profile());
continue;
}
for (int i = running; i < wanted; i++) {
if (launch(name, resolved.lead(), resolved.profile()) != null) {
if (launch(name, lead, profile)) {
started++;
}
}
@@ -219,87 +198,11 @@ public final class LeadLauncher {
return started;
}
/** A declared lead paired with the profile it launches on — {@link #resolveLaunchable}'s result. */
private record ResolvedLead(FleetConfig.Leader lead, FleetConfig.Profile profile) {
}
/**
* The declared {@code Leader} and its {@code Profile} for {@code name}, read from the config
* snapshot this launcher was constructed with.
*
* @return the resolved pair, or {@code null} (having logged) if {@code name} is not declared
* under {@code fleet.leaders}, that lead names no {@code profile:} (a {@code tab:}-only,
* recognise-only lead), or its {@code profile:} is not configured. Shared by
* {@link #ensureLeads()} and {@link #relaunch(String)} so the three refusals and their
* wording live in one place.
*/
private ResolvedLead resolveLaunchable(String name) {
FleetConfig.Leader lead = cfg.fleet().leaders().get(name);
if (lead == null) {
log.warn("lead '{}' is not declared under fleet.leaders — not launching", name);
return null;
}
if (!lead.isCreatable()) {
// A lead with a `tab:` but no `profile:` is recognise-only by design: the operator
// opens it by hand. Say so once rather than looking like a silent failure.
log.info("lead '{}' names no profile — it can be recognised but not launched. Add "
+ "`profile:` under fleet.leaders.{} to have fleetd start it.", name, name);
return null;
}
FleetConfig.Profile profile = cfg.profiles().get(lead.profile());
if (profile == null) {
log.warn("lead '{}' names profile '{}', which is not configured — not launching",
name, lead.profile());
return null;
}
return new ResolvedLead(lead, profile);
}
/**
* Start the named lead from the config snapshot this launcher was constructed with — not a
* live read, so a lead's {@code profile:} or {@code tab:} edited in config needs a daemon
* restart to take effect here — outside of {@link #ensureLeads()}'s {@code instances}
* bookkeeping.
*
* @return the started {@link Agent}, or {@code null} if {@code name} is not declared under
* {@code fleet.leaders}, that lead names no {@code profile:} (a {@code tab:}-only,
* recognise-only lead), its {@code profile:} is not configured, or every attempt up to
* {@link #RELAUNCH_ATTEMPTS} failed to start it. Never throws.
*
* <p>Does not count how many instances of this lead are already live. {@link #ensureLeads()}'s
* count exists to avoid starting a second orchestrator; the caller of this method has already
* decided to replace the lead and owns that decision.
*
* <p>Retries the whole launch attempt — not only the {@code agent_name_taken}/
* {@code agent_pane_busy} cases {@link ResilientAgentLaunch} already retries inside one
* {@code agents.start} call — up to {@link #RELAUNCH_ATTEMPTS} times, sleeping via the
* injected sleeper between attempts, and returns the agent from the first attempt that
* succeeds.
*/
public Agent relaunch(String name) {
ResolvedLead resolved = resolveLaunchable(name);
if (resolved == null) {
return null;
}
for (int attempt = 1; attempt <= RELAUNCH_ATTEMPTS; attempt++) {
Agent started = launch(name, resolved.lead(), resolved.profile());
if (started != null) {
return started;
}
if (attempt < RELAUNCH_ATTEMPTS) {
sleeper.run();
}
}
return null;
}
/**
* How many live leads exist per configured name, and which of that name's labelled tabs are
* <em>not</em> live: a running agent in a tab labelled with that lead's exact {@code tab}
* (CB-579). A member sitting in the same shared workspace is not counted as a lead because its
* tab carries a different label, not because any workspace is excluded from this count.
* (CB-579). Member workspaces are excluded, exactly as the scanner excludes them: a member must
* not be counted as a lead because it happens to sit in a matching tab.
*
* <p>There used to be a second path here — a running agent on the terminal a
* {@code fleet.leaders.<name>.terminal} pin named, for a lead opened and pinned by hand. That
@@ -403,17 +306,8 @@ public final class LeadLauncher {
return null;
}
/**
* Start one lead. Returns null (having logged) rather than throwing on any failure.
*
* <p>Goes through the same {@link ResilientAgentLaunch} seam every member spawn uses
* (fleetd #727): the assembled argv is refused outright if it cannot fit the pane line herdr
* types it into, a stale {@code agent_name_taken} (a crashed session's name the registry has
* not yet released) is retried under a fresh per-attempt name rather than refusing the whole
* relaunch, and a seed pane whose shell has not reached its prompt yet ({@code
* agent_pane_busy}) is retried rather than failing on the first miss.
*/
private Agent launch(String name, FleetConfig.Leader lead, FleetConfig.Profile profile) {
/** Start one lead. Returns false (having logged) rather than throwing on any failure. */
private boolean launch(String name, FleetConfig.Leader lead, FleetConfig.Profile profile) {
String label = lead.tabLabel();
String cwd = (lead.cwd() == null || lead.cwd().isBlank())
? System.getProperty("user.dir") : lead.cwd();
@@ -429,12 +323,8 @@ public final class LeadLauncher {
// Same shape as the member launchers: herdr resolves the executable from `kind`, so
// argv[0] (the configured launcher, e.g. `ccs`) is dropped and only the rest is passed.
List<String> argv = leadArgv(profile);
ResilientAgentLaunch.checkFits(profile.profile(), argv);
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
Agent started = ResilientAgentLaunch.startUniquelyNamed(agents, herdrKind(profile), args,
tab.rootPaneId(),
attempt -> "lead-" + name + "-" + nameNonce + "-" + nameSeq.incrementAndGet(),
ResilientAgentLaunch.NAME_RETRIES, ResilientAgentLaunch.SHELL_READY_RETRIES, sleeper);
Agent started = agents.start("lead-" + name, herdrKind(profile),
argv.isEmpty() ? argv : argv.subList(1, argv.size()), tab.rootPaneId());
// Label AFTER the start succeeds. A label written before would survive a failed start
// and then read back as a live lead on the next boot, which is the exact staleness the
@@ -444,7 +334,7 @@ public final class LeadLauncher {
log.info("lead '{}' launched: profile={} tab={} pane={} terminal={} label='{}' cwd={}",
name, profile.profile(), tab.tab().tabId(), started.paneId(),
started.terminalId(), label, cwd);
return started;
return true;
} catch (RuntimeException e) {
log.warn("lead '{}' failed to launch on profile '{}': {}",
name, profile.profile(), e.getMessage());
@@ -456,7 +346,7 @@ public final class LeadLauncher {
tab.tab().tabId(), cleanup.getMessage());
}
}
return null;
return false;
}
}
@@ -1,11 +1,8 @@
package dev.ltms.fleet.lead;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.WorkspaceControl;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -27,69 +24,60 @@ import java.util.function.Supplier;
/**
* fleetd #480: replace a lead session that has decided it is ready to be rolled over, without an
* operator doing it by hand. A lead writes a handover file, calls {@link #open}, and then — once
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation ends
* the lead's own pane, launches a fresh one, and bootstraps that fresh session against the file.
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation
* clears the lead's own pane and bootstraps a fresh session against that file.
*
* <p>This is the executor behind the {@code fleet_handover} MCP tool ({@code
* dev.ltms.fleet.mcp.FleetMcp#handover}), which drives {@link #open}, {@link #confirm}, {@link
* #cancel}, and {@link #status} from a tool call.
* #cancel}, and {@link #status} from a tool call — wired in fleetd #480 Unit C. <strong>An earlier
* version of this paragraph said nothing called this class at all; that stopped being true once
* that unit landed, and this correction exists so the javadoc does not go on claiming it.</strong>
*
* <p><strong>{@code confirm()} cannot roll inline.</strong> {@code confirm()} is called BY the
* lead, FROM the lead's own turn: the lead's pane is {@code WORKING} for the whole duration of that
* call and cannot possibly report a real turn boundary until {@code confirm()} itself returns. So
* {@link #confirm} validates every gate, then does no I/O against the lead's own pane at all — it
* only records that the request is approved and hands a one-shot continuation to {@code
* continuationRunner} before returning. That continuation is what actually touches the pane, once
* the calling turn has ended, in this order:
* <p><strong>{@code confirm()} cannot roll inline — a fleetd #480 correction.</strong> The first
* version of this class called {@code agents.send(lead, "/clear")} directly from inside {@code
* confirm()}, then polled for the pane to become injectable again. That is wrong, because {@code
* confirm()} is called BY the lead, FROM the lead's own turn: the lead's pane is {@code WORKING}
* for the whole duration of that call and cannot possibly report injectable until {@code confirm()}
* itself returns. The poll always timed out — but only after the {@code /clear} had already been
* sent and queued in the pane, where it fired the instant the turn ended anyway. The result was the
* worst outcome this feature can produce: a silently destroyed lead context with no fresh session
* ever started, and a refusal return value that claimed nothing had happened.
*
* <p>The fix: {@link #confirm} validates every gate, then does no I/O against the lead's own pane
* at all — it only records that the request is approved and hands a one-shot continuation to
* {@code continuationRunner} before returning. That continuation is what actually touches the pane,
* once the calling turn has ended, in this order:
* <ol>
* <li>wait for the lead's own pane to report a real turn boundary — {@code IDLE} or {@code
* DONE}, never merely {@code BLOCKED} — i.e. wait for the very {@code confirm()} call that
* approved this roll to finish its turn — bounded by {@code turnSettleSeconds}. <strong>If
* this never happens, nothing else in this list runs: the old pane is never touched.</strong>
* A lead that never goes idle is a lead still doing real work, and tearing it down would throw
* away live context.</li>
* <li>capture the old pane id (and, through it, the old tab) from {@link AgentControl#get}, with
* a bounded retry — the terminal-to-pane lookup it goes through can itself report a genuinely
* live agent as not found (see {@code AgentControl#agentCall}'s own re-resolve-once
* behaviour), and one false negative here must not abort an otherwise-healthy roll. Neither id
* is ever re-resolved from the terminal again after this — once the pane below is closed there
* is nothing left to resolve it from.</li>
* <li>resolve the lead's configured name from its terminal, for the relaunch step below.</li>
* <li>end the old session: close the pane (an already-gone pane counts as success; any other
* failure propagates), then close its tab only when the pane was that tab's sole occupant —
* the same pane-then-tab teardown {@code HerdrPeerLauncher#stop} uses for a member.</li>
* <li>confirm the old pane is actually gone by polling {@link
* dev.ltms.fleet.herdr.WorkspaceControl#locatePane} for a {@code null} result — never {@link
* AgentControl#status}, and never the live-lead terminal map, each of which answers a
* different question. <strong>If the old pane is never confirmed gone, no relaunch is
* attempted</strong> — see {@link RollState#OLD_PANE_NEVER_DIED}.</li>
* <li>launch a fresh lead with {@code LeadLauncher#relaunch}. <strong>If every attempt fails,
* {@code bootstrapText} is never sent</strong> — see {@link RollState#RELAUNCH_FAILED}.</li>
* <li>wait for the fresh pane to reach a real turn boundary ({@code IDLE} or {@code DONE},
* never merely {@code BLOCKED}), bounded by {@code relaunchReadySeconds}. This is the
* safety gate: typing into a pane that has not actually finished booting loses the
* keystrokes. <strong>If the pane never becomes ready, {@code bootstrapText} is never
* sent</strong> — see {@link RollState#RELAUNCH_NEVER_READY}.</li>
* <li>wait for the fresh terminal to be recognised as a live lead — present in the live-lead
* terminal map — bounded by {@code relaunchReadySeconds}. This is bookkeeping, not a
* safety gate: {@code bootstrapText} is sent either way once the pane is ready, whether or
* not this wait itself times out — see {@link RollState#RELAUNCH_NOT_RECOGNISED}.</li>
* <li>{@code agents.send(newTerminal, cfg.bootstrapTextFor(p.handoverPath()))} — sent to the
* FRESH terminal, never the one that was just torn down.</li>
* this never happens, nothing else in this list runs: no {@code /clear} is ever sent.</strong>
* A lead that never goes idle is a lead still doing real work, and clearing it would throw
* away live context — exactly the failure this correction exists to prevent.</li>
* <li>{@code agents.send(lead, "/clear")}</li>
* <li>wait for {@code /clear} to be picked up and settle, bounded by {@code clearSettleSeconds}
* (fleetd #489: no longer a plain re-check of the same boundary — {@code /clear} starts no
* turn of its own, so this instead nudges the submit keystroke while no pickup has been seen,
* then waits for a real {@code WORKING} → {@code IDLE}/{@code DONE} boundary once one has;
* see {@link #waitForClearPickupAndSettle})</li>
* <li>{@code agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()))}</li>
* </ol>
* A {@link #confirm} that returns {@link RollDecision#approved()} therefore means <em>"every gate
* passed and the roll is scheduled"</em>, never <em>"the lead has already been replaced"</em> — the
* old pane may still be mid-turn, possibly for a long time, when the caller gets that answer back.
* passed and the roll is scheduled"</em>, never <em>"the pane has been cleared"</em> — the pane may
* still be mid-turn, possibly for a long time, when the caller gets that answer back.
*
* <p><strong>The safety invariant.</strong> "No timer, no scheduler, no background thread" means
* that nothing but an explicit {@link #confirm} call can ever tear a lead's pane down.
* {@code continuationRunner} launches a single-shot task that exists only because one specific,
* <p><strong>The safety invariant survives this change, restated precisely.</strong> The ticket
* that first defined this class required "no timer, no scheduler, no background thread" so that
* nothing but an explicit {@link #confirm} call could ever cause a {@code /clear}. That invariant
* is about INITIATIVE, not about synchronicity, and this correction keeps it: {@code
* continuationRunner} launches a single-shot task that exists only because one specific,
* already-approved {@link #confirm} call created it — it is not recurring, it is not started at
* construction time or on any schedule, and no two invocations of it ever share state. A recurring
* heartbeat or timer that could decide on its own initiative to roll a pane is absent from this
* class. <strong>Nothing but an explicit {@link #confirm} call that passes every gate can ever tear
* a pane down — that call may simply finish its own work slightly later than the method return, as
* a continuation of the same approved request, rather than entirely inside the method body.</strong>
* heartbeat or timer that could decide on its own initiative to roll a pane is still, and will
* always be, absent from this class. <strong>Nothing but an explicit {@link #confirm} call that
* passes every gate can ever cause a {@code /clear} — that call may simply finish its own work
* slightly later than the method return, as a continuation of the same approved request, rather
* than entirely inside the method body.</strong>
*
* <p><strong>Identity is resolved by the caller, never looked up here — a second fleetd #480
* correction.</strong> The first version resolved the pane to clear via {@code
@@ -127,17 +115,19 @@ public final class LeadRollover {
private static final Logger log = LoggerFactory.getLogger(LeadRollover.class);
/** Poll interval shared by every bounded wait in this class. */
static final long POLL_INTERVAL_MS = 250;
/** Poll interval while waiting for the lead's pane to settle after {@code /clear}. */
static final long SETTLE_POLL_MS = 250;
/**
* How long {@link #waitUntilPaneGone} polls {@link WorkspaceControl#locatePane} before giving
* up on ever seeing the old pane disappear. Not configurable: once {@link #endOldSession} has
* closed the pane (and, usually, its tab), herdr dropping the pane from its own bookkeeping is
* expected to show up within one or two polls, not on an operator-tunable timescale the way a
* CLI boot is.
* How many consecutive not-yet-picked-up polls {@link #waitForClearPickupAndSettle} allows
* before releasing rather than wedging the roll — the same constant and the same
* release-not-wedge choice {@link dev.ltms.fleet.inject.Injector} already makes for its own
* post-turn {@code /clear} housekeeping (fleetd #306). <strong>This bounds the number of
* consecutive polls, not the number of nudges:</strong> the first {@code PICKUP_GRACE_POLLS - 1}
* of those polls each send a nudge, and the {@code PICKUP_GRACE_POLLS}th releases instead of
* nudging again — so 8 polls produce 7 nudges, not 8.
*/
static final int PANE_DEATH_TIMEOUT_SECONDS = 10;
static final int PICKUP_GRACE_POLLS = 8;
/**
* One request opened by {@link #open}, pending its {@link #confirm} (or {@link #cancel}).
@@ -174,21 +164,15 @@ public final class LeadRollover {
* The handover file's modified time is not after {@link #open}'s request timestamp, or is
* older than {@code maxDocAgeSeconds}.
*/
HANDOVER_STALE,
/**
* This lead terminal already has a roll running: an earlier {@link #confirm} call claimed
* it and that roll's continuation has not released it yet. {@code detail} names the lead
* terminal and the token that holds the claim.
*/
ROLL_ALREADY_RUNNING
HANDOVER_STALE
}
/**
* The outcome of a {@link #confirm} call. {@link #approved()} means every gate passed and the
* roll has been handed to a one-shot continuation — <strong>not</strong> that the lead has
* already been replaced; the continuation may still be waiting for the calling turn to end when
* this returns. Whether the deferred roll itself later goes on to tear the old pane down and
* relaunch the lead, or refuses at any of its own steps, is logged only (see this class's
* roll has been handed to a one-shot continuation — <strong>not</strong> that the pane has been
* cleared; the continuation may still be waiting for the calling turn to end when this returns.
* Whether the deferred roll itself later goes on to clear the pane, refuse for never going
* idle, or refuse for never re-settling after {@code /clear} is logged only (see this class's
* javadoc) — there is deliberately no synchronous caller left by that point to hand a result to.
*/
public record RollDecision(boolean accepted, RefusalReason reason, String detail) {
@@ -224,7 +208,7 @@ public final class LeadRollover {
/**
* What is known about one token, right now — the answer {@link #status} gives. Distinguishes
* five terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
* three terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
* that has been approved but has not finished yet, and two answers for a token that names no
* active work at all: still pending confirmation, or nothing known about this token at all.
*/
@@ -247,64 +231,41 @@ public final class LeadRollover {
* #status} could wrongly answer {@link #UNKNOWN} ("nothing was ever requested") for a roll
* that is, in fact, actively running. This is not sticky: the deferred continuation
* overwrites this same entry with a terminal state ({@link #ROLLED}, {@link
* #TURN_NEVER_SETTLED}, {@link #OLD_PANE_NEVER_DIED}, {@link #RELAUNCH_FAILED}, {@link
* #RELAUNCH_NEVER_READY}, {@link #RELAUNCH_NOT_RECOGNISED}, or {@link #FAILED}) once it
* finishes — including by throwing, which {@link #runRollover}'s catch turns into {@link
* #FAILED} instead of leaving this entry stuck forever.
* #TURN_NEVER_SETTLED}, {@link #CLEAR_NEVER_SETTLED}, or {@link #FAILED}) once it finishes
* — including by throwing, which fleetd #615's catch in {@link #runRollover} now turns into
* {@link #FAILED} instead of leaving this entry stuck forever.
*/
IN_PROGRESS,
/**
* {@link #confirm} was approved and the deferred continuation completed the entire roll: the
* calling lead's turn settled, the old pane was torn down and confirmed gone, a fresh lead
* was launched and recognised, and {@code bootstrapText} was sent to it.
* {@link #confirm} was approved and the deferred continuation completed the entire roll:
* the calling lead's turn settled, {@code /clear} was sent and settled, and {@code
* bootstrapText} was sent.
*/
ROLLED,
/**
* {@link #confirm} was approved, but the calling lead's own turn never reached a boundary
* (IDLE or DONE) within {@code turnSettleSeconds} — the old pane was never touched at all.
* This is the state that makes a lead's own stuck turn VISIBLE: without it, a lead that hit
* this case would have no way to find out, and would carry on believing it was about to be
* replaced. See this class's javadoc.
* (IDLE or DONE) within {@code turnSettleSeconds} — no {@code /clear} was ever sent, at
* all. This is the branch the fleetd #480 correction exists to make safe, and the one this
* status exists to make VISIBLE: before this, a lead that hit this case had no way to find
* out, and would carry on believing it was about to be replaced. See this class's javadoc.
*/
TURN_NEVER_SETTLED,
/**
* {@link #confirm} was approved and the calling lead's turn settled, the old pane was closed
* (and its tab, if it was the sole occupant), but {@link
* dev.ltms.fleet.herdr.WorkspaceControl#locatePane} kept reporting it as still present for
* the whole pane-death timeout. No relaunch was ever attempted, and {@code bootstrapText}
* was never sent.
* {@link #confirm} was approved and {@code /clear} was sent, but the pane never re-settled
* within {@code clearSettleSeconds} — {@code bootstrapText} was never sent.
*/
OLD_PANE_NEVER_DIED,
CLEAR_NEVER_SETTLED,
/**
* The old pane was confirmed gone, but {@code LeadLauncher#relaunch} returned {@code null}
* — every launch attempt failed. {@code bootstrapText} was never sent, and no fresh terminal
* exists for this roll to have recognised.
*/
RELAUNCH_FAILED,
/**
* A fresh lead was launched, but its pane never reached a real turn boundary ({@code IDLE}
* or {@code DONE}, never merely {@code BLOCKED}) within {@code relaunchReadySeconds} — the
* CLI never finished booting, or it stayed paused on a startup prompt. {@code bootstrapText}
* was never sent: typing into a pane that is not actually ready to accept input loses the
* keystrokes.
*/
RELAUNCH_NEVER_READY,
/**
* A fresh lead was launched and its pane reached a real turn boundary, so {@code
* bootstrapText} WAS sent to it, but the terminal was never recognised as a live lead —
* present in the live-lead terminal map — within {@code relaunchReadySeconds}. The session
* itself is alive and bootstrapped; only the daemon's own bookkeeping has not caught up, and
* an operator should check why the tab was not recognised.
*/
RELAUNCH_NOT_RECOGNISED,
/**
* The deferred continuation threw a {@link RuntimeException} and the continuation thread
* died with it. Without this state, that throw would leave {@link #outcomes} holding {@link
* #IN_PROGRESS} forever, because the production {@code continuationRunner} is a bare virtual
* thread with no uncaught-exception handler and nothing downstream of the throw ever runs to
* write a terminal outcome. {@code detail} names the exception, so a reader has something to
* act on. The roll is dead at this point and does not retry itself; a stuck lead must
* {@link #open} a fresh request.
* fleetd #615: the deferred continuation threw a {@link RuntimeException} — most likely a
* {@link dev.ltms.fleet.herdr.HerdrException} out of one of the two unwrapped {@code
* agents.send} calls in {@link #runRollover} — and the continuation thread died with it.
* Before this state existed, that throw left {@link #outcomes} holding {@link #IN_PROGRESS}
* forever, because the production {@code continuationRunner} is a bare virtual thread with
* no uncaught-exception handler and nothing downstream of the throw ever ran to write a
* terminal outcome. {@code detail} names the exception, so a reader has something to act on
* — the same diagnostic style as {@link #TURN_NEVER_SETTLED} and {@link
* #CLEAR_NEVER_SETTLED}. The roll is dead at this point and does not retry itself; a stuck
* lead must {@link #open} a fresh request.
*/
FAILED,
/**
@@ -325,10 +286,6 @@ public final class LeadRollover {
public record RollStatus(RollState state, String detail) {}
private final AgentControl agents;
/** Workspace/tab/pane control — used to tear down the old pane and confirm it is gone. */
private final WorkspaceControl spaces;
/** Starts the fresh lead that replaces the one this roll tears down. */
private final LeadLauncher launcher;
private final Supplier<FleetConfig.LeadRollover> configSupplier;
/**
* Terminal id → that lead's configured workspace directory (their {@code
@@ -338,21 +295,8 @@ public final class LeadRollover {
* daemon-cwd bug this parameter exists to fix.
*/
private final Function<String, String> leadWorkspace;
/**
* Terminal id → that lead's configured name under {@code fleet.leaders}, or {@code null} when
* the terminal names no currently-recognised lead. The deferred continuation calls this, on the
* OLD terminal, before tearing it down, so it knows which lead to pass to {@link
* LeadLauncher#relaunch}.
*/
private final Function<String, String> leadNameForTerminal;
/**
* The daemon's current terminal id → lead name map, read fresh on every poll. The deferred
* continuation polls this for the FRESH terminal {@link LeadLauncher#relaunch} returns, to
* learn when that terminal has been recognised as a live lead — see this class's javadoc.
*/
private final Supplier<Map<String, String>> liveLeadTerminals;
private final LongSupplier nowMillis;
private final Runnable pollSleeper;
private final Runnable settleSleeper;
/**
* Launches the post-{@code confirm()} continuation. Production uses a single unstarted virtual
* thread per confirmed request — see this class's javadoc for why that is a single-shot task,
@@ -361,15 +305,6 @@ public final class LeadRollover {
*/
private final Consumer<Runnable> continuationRunner;
private final Map<String, PendingRollover> pending = new ConcurrentHashMap<>();
/**
* Lead terminal → the token of the roll currently holding that terminal exclusive, for
* {@link #confirm}'s single-flight claim. {@link #confirm} claims an entry here with an
* atomic put-if-absent once every other gate has passed, refusing with {@link
* RefusalReason#ROLL_ALREADY_RUNNING} when a claim is already held; {@link #runRollover}
* releases it in a {@code finally}, on both the success and the thrown-exception path. A
* terminal absent from this map has no roll currently in flight for it.
*/
private final Map<String, String> rollingByTerminal = new ConcurrentHashMap<>();
/**
* Finished tokens → what actually happened, for {@link #status}. Bounded by {@link
* #OUTCOME_HISTORY_CAP}, oldest evicted first ({@code removeEldestEntry} on an insertion-order
@@ -388,40 +323,29 @@ public final class LeadRollover {
}
});
/** Production constructor — wall clock, real sleep between polls, a real virtual thread. */
public LeadRollover(AgentControl agents, WorkspaceControl spaces, LeadLauncher launcher,
Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace,
Function<String, String> leadNameForTerminal,
Supplier<Map<String, String>> liveLeadTerminals) {
this(agents, spaces, launcher, configSupplier, leadWorkspace, leadNameForTerminal,
liveLeadTerminals, System::currentTimeMillis,
() -> sleepUninterruptibly(POLL_INTERVAL_MS),
/** Production constructor — wall clock, real sleep between settle polls, a real virtual thread. */
public LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace) {
this(agents, configSupplier, leadWorkspace, System::currentTimeMillis,
() -> sleepUninterruptibly(SETTLE_POLL_MS),
r -> Thread.ofVirtual().name("lead-rollover-continuation-").start(r));
}
/**
* Full constructor — an injectable wall-clock supplier, poll sleeper, and continuation runner,
* for tests. {@code nowMillis} MUST be a wall-clock source (e.g. {@code
* Full constructor — an injectable wall-clock supplier, settle-poll sleeper, and continuation
* runner, for tests. {@code nowMillis} MUST be a wall-clock source (e.g. {@code
* System.currentTimeMillis()}), never {@code System.nanoTime()}: the freshness check compares
* against a file's modified time, which only a wall clock is comparable to, and {@code
* nanoTime} freezes while the host sleeps.
* nanoTime} freezes while the host sleeps (fleetd #386).
*/
LeadRollover(AgentControl agents, WorkspaceControl spaces, LeadLauncher launcher,
Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace,
Function<String, String> leadNameForTerminal,
Supplier<Map<String, String>> liveLeadTerminals,
LongSupplier nowMillis, Runnable pollSleeper, Consumer<Runnable> continuationRunner) {
LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace, LongSupplier nowMillis,
Runnable settleSleeper, Consumer<Runnable> continuationRunner) {
this.agents = agents;
this.spaces = spaces;
this.launcher = launcher;
this.configSupplier = configSupplier;
this.leadWorkspace = leadWorkspace;
this.leadNameForTerminal = leadNameForTerminal;
this.liveLeadTerminals = liveLeadTerminals;
this.nowMillis = nowMillis;
this.pollSleeper = pollSleeper;
this.settleSleeper = settleSleeper;
this.continuationRunner = continuationRunner;
}
@@ -563,16 +487,6 @@ public final class LeadRollover {
return docCheck;
}
// Single-flight claim: atomic put-if-absent, taken only after every other gate has
// passed, so a refused confirm() never takes it. A non-null previous value means a
// different, still-running roll already holds this lead terminal.
String holder = rollingByTerminal.putIfAbsent(p.leadTerminal(), token);
if (holder != null) {
return RollDecision.refused(RefusalReason.ROLL_ALREADY_RUNNING,
"lead terminal " + p.leadTerminal() + " already has a roll running under token "
+ holder);
}
// Record IN_PROGRESS BEFORE removing from `pending` — see RollState#IN_PROGRESS and
// OUTCOME_HISTORY_CAP's javadoc. This ordering means `token` is written into `outcomes`
// while it is STILL present in `pending`; status() checks `outcomes` first (see that
@@ -580,48 +494,34 @@ public final class LeadRollover {
// remove-then-put ordering would leave in which the token is in neither map.
outcomes.put(token, new RollStatus(RollState.IN_PROGRESS,
"confirm() approved this roll and handed it to the deferred continuation; it has "
+ "not finished yet — still waiting for the calling turn to settle, for the "
+ "old pane to be torn down and confirmed gone, for the fresh lead to be "
+ "recognised, or for bootstrapText to be sent"));
+ "not finished yet — still waiting for the calling turn to settle, for "
+ "/clear to be sent and settle, or for bootstrapText to be sent"));
pending.remove(token);
log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends",
token, callerTerminal);
try {
continuationRunner.accept(() -> runRollover(p, cfg));
} catch (RuntimeException e) {
// continuationRunner can reject the hand-off itself (e.g. a bounded executor's
// RejectedExecutionException) before runRollover ever starts, so runRollover's own
// finally — the only other place that releases rollingByTerminal — never runs either.
// Release the claim here and overwrite the IN_PROGRESS entry with a terminal outcome,
// or this lead terminal could never be rolled again and status() would report
// IN_PROGRESS forever for a roll that in fact never started.
log.warn("lead-rollover: continuationRunner rejected token={} lead={}: {} — the roll "
+ "never started; releasing its claim and reporting it as FAILED",
token, callerTerminal, e.toString(), e);
rollingByTerminal.remove(p.leadTerminal(), token);
outcomes.put(token, new RollStatus(RollState.FAILED,
"continuationRunner rejected this roll before it ever started: " + e.toString()
+ " — the roll never ran; open() a fresh rollover request"));
}
continuationRunner.accept(() -> runRollover(p, cfg));
return RollDecision.approved();
}
/**
* The single-shot continuation {@link #confirm} hands to {@code continuationRunner}. Runs
* entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the
* full order. There is no result to return to by this point, so every outcome is logged only.
* four-step order. There is no result to return to by this point, so every outcome is logged
* only.
*
* <p><strong>The whole body is wrapped in one {@code try}.</strong> Several calls below —
* {@code agents.get}, {@code agents.close}, {@code agents.send} — can throw an unchecked {@link
* dev.ltms.fleet.herdr.HerdrException} (see {@code AgentControl.java}), and the production
* {@code continuationRunner} is a bare virtual thread with no uncaught-exception handler (see
* this class's public constructor). An uncaught throw would kill the continuation thread
* silently, leaving the {@link RollState#IN_PROGRESS} entry {@link #confirm} wrote at hand-off
* stuck forever — {@link #status} would have no way to tell a dead roll from one still
* genuinely running. The {@code catch} below is scoped to the method body rather than to each
* call individually, so it also covers every call in this continuation, not a fixed list of
* call sites — the same reasoning that put the write-a-terminal-outcome step at each of this
* method's other exits rather than inside the helpers that detect them.</p>
* <p><strong>fleetd #615 — the whole body is wrapped in one {@code try}.</strong> The two {@code
* agents.send} calls below are not wrapped individually: {@code send} → {@code agentCall} →
* {@code herdr.call} can throw an unchecked {@link dev.ltms.fleet.herdr.HerdrException} (see
* {@code AgentControl.java}), and the production {@code continuationRunner} is a bare virtual
* thread with no uncaught-exception handler (see this class's public constructor). Before this
* fix, either throw killed the continuation thread silently, leaving the {@link
* RollState#IN_PROGRESS} entry {@link #confirm} wrote at hand-off stuck forever — {@link
* #status} had no way to tell a dead roll from one still genuinely running. The {@code catch}
* below is scoped to the method body rather than to each {@code send} call individually, so it
* also covers anything else added to this continuation later, not just today's two call sites —
* the same reasoning that put the write-a-terminal-outcome step at each of this method's other
* exits (see the {@link RollState#TURN_NEVER_SETTLED} and {@link RollState#CLEAR_NEVER_SETTLED}
* branches below) rather than inside the helpers that detect them.</p>
*
* <p>Only {@link RuntimeException} is caught, matching the local convention {@link
* #waitUntilAtTurnBoundary} already set around its own {@code agents.status} call — not the
@@ -639,12 +539,6 @@ public final class LeadRollover {
"the roll's continuation threw " + e.toString() + " — the roll is dead and will "
+ "not retry itself; check the daemon log for the stack trace, then open() "
+ "a fresh rollover request"));
} finally {
// Release the single-flight claim on both the normal return and the thrown-exception
// path above — a release only on success would leave this lead terminal unrollable
// forever after one failure. The conditional two-argument remove only clears the
// entry this roll itself holds, never a different roll's claim on the same terminal.
rollingByTerminal.remove(p.leadTerminal(), p.token());
}
}
@@ -652,237 +546,53 @@ public final class LeadRollover {
private void runRolloverUnguarded(PendingRollover p, FleetConfig.LeadRollover cfg) {
String lead = p.leadTerminal();
long rollStartMillis = nowMillis.getAsLong();
TurnSettleResult turnResult = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds());
if (!turnResult.settled()) {
// fleetd #494 follow-up: this line had the SAME defect as the /clear-timeout line below
// — cfg.turnSettleSeconds() is the CONFIGURED budget, not how long this wait actually
// ran. Print the measured elapsed time alongside it, labelled, exactly like the /clear
// path already does.
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
+ "confirm() — the old pane is never touched; the calling lead's own "
+ "turn is still live and tearing it down now would destroy live context "
+ "confirm() — refusing to send /clear at all; the calling lead's own "
+ "turn is still live and clearing it now would destroy live context "
+ "(token={}, configured={}s elapsed={}ms)",
lead, p.token(), cfg.turnSettleSeconds(), turnResult.elapsedMillis());
outcomes.put(p.token(), new RollStatus(RollState.TURN_NEVER_SETTLED,
"the calling lead's own turn never reached a boundary (IDLE or DONE) within "
+ "turnSettleSeconds=" + cfg.turnSettleSeconds() + "s (measured elapsed="
+ turnResult.elapsedMillis() + "ms) — the old pane was never touched. If "
+ "this keeps happening, raise turnSettleSeconds in fleetd.yaml"));
+ turnResult.elapsedMillis() + "ms) — no /clear was ever sent. If this "
+ "keeps happening, raise turnSettleSeconds in fleetd.yaml"));
return;
}
// Captured once, here, and never re-resolved from `lead` again below: once the pane is
// closed there is nothing left for a terminal lookup to find.
Agent oldAgent = captureAgentWithRetry(lead);
String oldPaneId = oldAgent.paneId();
String leadName = leadNameForTerminal.apply(lead);
endOldSession(oldPaneId);
DeathResult deathResult = waitUntilPaneGone(oldPaneId);
if (!deathResult.gone()) {
log.warn("lead-rollover: old pane {} for lead {} was never confirmed gone after being "
+ "closed — not attempting a relaunch (token={}, timeout={}s "
+ "elapsed={}ms)",
oldPaneId, lead, p.token(), PANE_DEATH_TIMEOUT_SECONDS, deathResult.elapsedMillis());
outcomes.put(p.token(), new RollStatus(RollState.OLD_PANE_NEVER_DIED,
"the old pane was closed, but locatePane kept reporting it as still present "
+ "after a pane-death timeout=" + PANE_DEATH_TIMEOUT_SECONDS
+ "s (measured elapsed=" + deathResult.elapsedMillis() + "ms) — no "
+ "relaunch was attempted"));
// This deliberately bypasses Injector, exactly like ClaudeCodeLauncher#clearContext:
// /clear is housekeeping, not a delegated turn, and routing it through Injector wedges the
// pane forever (see this class's javadoc).
agents.send(lead, "/clear");
ClearSettleResult clearResult = waitForClearPickupAndSettle(lead, cfg.clearSettleSeconds());
if (!clearResult.settled()) {
// fleetd #494: cfg.clearSettleSeconds() is the CONFIGURED budget, not how long the wait
// actually ran — an operator reading only that number wrongly believes it is a measured
// duration. Print the measured elapsed time and nudge count alongside it, each labelled,
// so the two can be compared at a glance.
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
+ "/clear — NOT sending bootstrapText (token={}, configured={}s "
+ "elapsed={}ms nudges={})",
lead, p.token(), cfg.clearSettleSeconds(), clearResult.elapsedMillis(),
clearResult.nudges());
outcomes.put(p.token(), new RollStatus(RollState.CLEAR_NEVER_SETTLED,
"/clear was sent, but the pane never re-settled within clearSettleSeconds="
+ cfg.clearSettleSeconds() + "s (measured elapsed=" + clearResult.elapsedMillis()
+ "ms, nudges=" + clearResult.nudges() + ") — bootstrapText was never sent"));
return;
}
Agent newAgent = launcher.relaunch(leadName);
if (newAgent == null) {
log.warn("lead-rollover: relaunch of lead '{}' (old terminal {}) failed every attempt "
+ "— bootstrapText was never sent (token={})", leadName, lead, p.token());
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_FAILED,
"lead '" + leadName + "' could not be relaunched — every attempt failed; "
+ "bootstrapText was never sent"));
return;
}
ReadinessResult readinessResult = waitUntilPaneReady(newAgent.terminalId(),
cfg.relaunchReadySeconds());
if (!readinessResult.ready()) {
log.warn("lead-rollover: fresh pane for lead '{}' (terminal {}) never reached a real "
+ "turn boundary — bootstrapText was never sent (token={}, configured={}s "
+ "elapsed={}ms)",
leadName, newAgent.terminalId(), p.token(), cfg.relaunchReadySeconds(),
readinessResult.elapsedMillis());
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_NEVER_READY,
"fresh terminal " + newAgent.terminalId() + " never reached a real turn "
+ "boundary (IDLE or DONE) within relaunchReadySeconds="
+ cfg.relaunchReadySeconds() + "s (measured elapsed="
+ readinessResult.elapsedMillis() + "ms) — bootstrapText was never "
+ "sent"));
return;
}
IdentityResult identityResult = waitUntilRecognisedAsLead(newAgent.terminalId(),
cfg.relaunchReadySeconds());
agents.send(newAgent.terminalId(), cfg.bootstrapTextFor(p.handoverPath()));
if (!identityResult.ready()) {
log.warn("lead-rollover: fresh terminal {} for lead '{}' is alive and bootstrapped, but "
+ "was never recognised as a live lead — an operator should check why "
+ "the tab was not recognised (token={}, configured={}s elapsed={}ms)",
newAgent.terminalId(), leadName, p.token(), cfg.relaunchReadySeconds(),
identityResult.elapsedMillis());
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_NOT_RECOGNISED,
"bootstrapText was sent to fresh terminal " + newAgent.terminalId() + ", but "
+ "it was never recognised as a live lead within relaunchReadySeconds="
+ cfg.relaunchReadySeconds() + "s (measured elapsed="
+ identityResult.elapsedMillis() + "ms) — check why the tab was not "
+ "recognised"));
return;
}
agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()));
long rollElapsedMillis = nowMillis.getAsLong() - rollStartMillis;
log.info("lead-rollover: rolled token={} oldLead={} newTerminal={} elapsedMs={}",
p.token(), lead, newAgent.terminalId(), rollElapsedMillis);
log.info("lead-rollover: rolled token={} lead={} elapsedMs={}", p.token(), lead, rollElapsedMillis);
outcomes.put(p.token(), new RollStatus(RollState.ROLLED,
"rolled successfully in " + rollElapsedMillis + "ms; new terminal="
+ newAgent.terminalId()));
"rolled successfully in " + rollElapsedMillis + "ms"));
}
/** Attempts {@link #captureAgentWithRetry} makes before letting the failure propagate. */
static final int CAPTURE_RETRIES = 3;
/**
* {@link AgentControl#get} for {@code lead}, retried up to {@link #CAPTURE_RETRIES} times. The
* terminal-to-pane lookup it goes through can report a genuinely live agent as not found (see
* {@code AgentControl#agentCall}'s own re-resolve-once behaviour), and one such false negative
* must not abort an otherwise-healthy roll. The result is captured once by the caller and never
* looked up again — see this class's javadoc.
*
* @throws RuntimeException the last failure, if every attempt fails — {@link #runRollover}'s
* catch turns that into {@link RollState#FAILED}
*/
private Agent captureAgentWithRetry(String lead) {
RuntimeException last = null;
for (int attempt = 1; attempt <= CAPTURE_RETRIES; attempt++) {
try {
return agents.get(lead);
} catch (RuntimeException e) {
last = e;
log.debug("lead-rollover: agents.get({}) failed on attempt {}/{}: {}",
lead, attempt, CAPTURE_RETRIES, e.toString());
if (attempt < CAPTURE_RETRIES) {
pollSleeper.run();
}
}
}
throw last;
}
/**
* End the old lead's session: close its pane, then close its tab only when the pane was that
* tab's sole occupant — the same pane-then-tab teardown {@code HerdrPeerLauncher#stop} uses for
* a member. An already-gone pane counts as success; any other {@code agents.close} failure
* propagates, so a genuinely failed teardown is never reported as done. A failing
* {@code spaces.closeTab} never propagates — by the time it runs the pane is already closed, so
* it is cosmetic tidying, not a real teardown failure.
*/
private void endOldSession(String paneId) {
WorkspaceControl.PaneLocation loc = spaces.locatePane(paneId);
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) {
throw e;
}
log.debug("lead-rollover: pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
try {
spaces.closeTab(loc.tabId());
} catch (RuntimeException e) {
log.warn("lead-rollover: tab.close({}) failed — the pane is already torn down, so "
+ "continuing; the tab may need manual cleanup: {}", loc.tabId(), e.getMessage());
}
} else if (loc != null) {
log.debug("lead-rollover: not closing tab {} — it holds {} panes (not a dedicated lead "
+ "tab)", loc.tabId(), loc.tabPaneCount());
}
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
/**
* Poll {@link WorkspaceControl#locatePane} for {@code paneId} until it reports {@code null}
* (the pane is gone) or {@link #PANE_DEATH_TIMEOUT_SECONDS} elapses. Deliberately never calls
* {@link AgentControl#status} and never reads the live-lead terminal map — both answer a
* different question (whether an AGENT is live, not whether this PANE still exists) and
* {@code locatePane} alone catches a {@link HerdrException} from the underlying {@code
* pane.get} and turns it into {@code null} — see this class's javadoc.
*/
private DeathResult waitUntilPaneGone(String paneId) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(PANE_DEATH_TIMEOUT_SECONDS);
while (nowMillis.getAsLong() < deadline) {
if (spaces.locatePane(paneId) == null) {
return new DeathResult(true, nowMillis.getAsLong() - startMillis);
}
pollSleeper.run();
}
return new DeathResult(false, nowMillis.getAsLong() - startMillis);
}
/** The measured outcome of {@link #waitUntilPaneGone}. */
private record DeathResult(boolean gone, long elapsedMillis) {}
/**
* Poll until {@code newTerminal}'s own pane reaches a real turn boundary ({@link
* AgentStatus#IDLE} or {@link AgentStatus#DONE}, never merely {@link AgentStatus#BLOCKED}) —
* the same exclusion {@link #waitUntilAtTurnBoundary} applies to the calling lead's own turn,
* applied here to the fresh one, so {@code bootstrapText} is never typed into a pane that has
* not actually finished booting — or {@code readySeconds} elapses. A failed status read
* degrades to "not yet ready" and is retried on the next poll.
*/
private ReadinessResult waitUntilPaneReady(String newTerminal, int readySeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(readySeconds);
while (nowMillis.getAsLong() < deadline) {
AgentStatus status;
try {
status = agents.status(newTerminal);
} catch (RuntimeException e) {
log.debug("lead-rollover: status check failed while waiting for {} to be ready: {}",
newTerminal, e.toString());
status = null;
}
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
return new ReadinessResult(true, nowMillis.getAsLong() - startMillis);
}
pollSleeper.run();
}
return new ReadinessResult(false, nowMillis.getAsLong() - startMillis);
}
/** The measured outcome of {@link #waitUntilPaneReady}. */
private record ReadinessResult(boolean ready, long elapsedMillis) {}
/**
* Poll until {@code newTerminal} is present in {@link #liveLeadTerminals} or {@code
* readySeconds} elapses. This is bookkeeping, not a safety gate: the pane's own readiness (see
* {@link #waitUntilPaneReady}) is what decides whether {@code bootstrapText} is safe to send —
* a timeout here only means the daemon's own lead-discovery scan has not caught up yet.
*/
private IdentityResult waitUntilRecognisedAsLead(String newTerminal, int readySeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(readySeconds);
while (nowMillis.getAsLong() < deadline) {
if (liveLeadTerminals.get().containsKey(newTerminal)) {
return new IdentityResult(true, nowMillis.getAsLong() - startMillis);
}
pollSleeper.run();
}
return new IdentityResult(false, nowMillis.getAsLong() - startMillis);
}
/** The measured outcome of {@link #waitUntilRecognisedAsLead}. */
private record IdentityResult(boolean ready, long elapsedMillis) {}
/** Drop a pending request without rolling. @return whether a pending request existed for {@code token} */
public boolean cancel(String token) {
return pending.remove(token) != null;
@@ -972,12 +682,15 @@ public final class LeadRollover {
/**
* Poll {@link AgentControl#status} until {@code target} reports a real turn boundary — {@link
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used by
* {@link #runRollover} to wait for the CALLING turn's own pane to settle before the old pane is
* touched at all — the {@code turnSettleSeconds} gate that makes tearing it down safe. A failed
* status read degrades to "not yet settled" and is retried on the next poll, the same posture
* {@code LeadHeartbeatLoop} and {@code HerdrPeerLauncher}'s readiness gate already take toward
* an unreadable status.
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used once by
* {@link #runRollover}, to wait for the CALLING turn's own pane to settle before {@code /clear}
* is ever sent at all — the {@code turnSettleSeconds} gate that makes this correction safe. The
* SECOND wait, after {@code /clear}, is {@link #waitForClearPickupAndSettle} instead (fleetd
* #489) — a plain boundary check is not enough there, because {@code /clear} starts no turn of
* its own, so this method would (wrongly) report "settled" on its very first poll whether or not
* {@code /clear} was actually picked up. A failed status read degrades to "not yet settled" and
* is retried on the next poll, the same posture {@code LeadHeartbeatLoop} and {@code
* HerdrPeerLauncher}'s readiness gate already take toward an unreadable status.
*
* <p><strong>Deliberately not {@link AgentStatus#injectable()}.</strong> {@code injectable()}
* answers the {@code Injector}'s question — "may I deliver a message without stepping on a live
@@ -985,16 +698,16 @@ public final class LeadRollover {
* an approval prompt is safe to queue a message behind. This class asks a stricter question —
* "has the turn actually ended" — and {@code BLOCKED} answers no: it is a live turn that is
* merely paused, not one that has finished. Reusing {@code injectable()} here would let this
* wait tear the old pane down while the lead's own {@code confirm()}-calling turn is still live
* and paused on a prompt — exactly the live-context-destroying failure {@code turnSettleSeconds}
* exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link
* #waitUntilPaneReady} applies the same exclusion of {@code BLOCKED} to the fresh lead's own
* turn.)
* wait fire {@code /clear} while the lead's own {@code confirm()}-calling turn is still live and
* paused on a prompt — exactly the live-context-destroying failure the {@code turnSettleSeconds}
* gate exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link
* #waitForClearPickupAndSettle} keeps the same exclusion of {@code BLOCKED}, for the same
* reason, on the second wait.)
*
* @return a {@link TurnSettleResult} whose {@code settled()} is {@code true} once a real
* boundary was observed, {@code false} if {@code settleSeconds} elapses first.
* {@code elapsedMillis()} is a MEASURED value from the injected {@link #nowMillis}
* clock, never the configured {@code settleSeconds} budget.
* clock, never the configured {@code settleSeconds} budget (fleetd #494 follow-up).
*/
private TurnSettleResult waitUntilAtTurnBoundary(String target, int settleSeconds) {
long startMillis = nowMillis.getAsLong();
@@ -1011,11 +724,142 @@ public final class LeadRollover {
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
return new TurnSettleResult(true, nowMillis.getAsLong() - startMillis);
}
pollSleeper.run();
settleSleeper.run();
}
return new TurnSettleResult(false, nowMillis.getAsLong() - startMillis);
}
/** The measured outcome of {@link #waitUntilAtTurnBoundary}. */
/**
* The measured outcome of {@link #waitUntilAtTurnBoundary} — fleetd #494 follow-up. The sibling
* of {@link ClearSettleResult} for the FIRST wait, which never nudges, so it carries no nudge
* count.
*/
private record TurnSettleResult(boolean settled, long elapsedMillis) {}
/**
* The SECOND wait in {@link #runRollover} — after {@code /clear} has been sent, waits for it to
* settle, bounded by {@code settleSeconds}. <strong>fleetd #489 — the paste-race fix.</strong>
* {@code /clear} does not start a real turn of its own, so a pane with no submit race simply
* stays {@link AgentStatus#IDLE} the whole time: {@link #waitUntilAtTurnBoundary} would (wrongly)
* call that "settled" on its very first poll, whether or not the {@code /clear} Enter actually
* landed. That was Fault 1, measured live on 2026-09-12 — the second gate was a no-op, so a
* {@code bootstrapText} send followed immediately, racing Fault 2: {@link AgentControl#submit}'s
* own javadoc already records that the submit accompanying a delivery "can race the paste —
* especially right as the worker's TUI becomes interactive — leaving the text unsubmitted"
* (CB-113). Because {@code runRollover} deliberately bypasses {@code Injector} for {@code
* /clear} (see this class's javadoc), it inherited none of {@code Injector}'s nudging — so the
* lost {@code /clear} Enter sat in the input box and {@code bootstrapText} was typed right after
* it, landing as one concatenated line.
*
* <p>This method copies the pickup-nudge pattern {@link dev.ltms.fleet.inject.Injector} already
* ships for exactly this, on its own post-turn {@code /clear} housekeeping (fleetd #306; see
* {@code Injector.java:288-340} and {@code Injector.java:437-442}):
* <ul>
* <li>an {@link AgentStatus#WORKING} sample means {@code /clear} was picked up as a real
* turn;</li>
* <li>until that happens, each poll that still reports {@link AgentStatus#IDLE} or {@link
* AgentStatus#DONE} re-sends the submit keystroke ({@link AgentControl#submit}) to nudge
* the raced Enter — for the first {@code PICKUP_GRACE_POLLS - 1} of {@link
* #PICKUP_GRACE_POLLS} consecutive such polls (i.e. {@code PICKUP_GRACE_POLLS - 1}
* nudges: 7, not 8, given {@code PICKUP_GRACE_POLLS = 8}). A second Enter on an empty
* Claude Code prompt is a no-op, so repeating it is safe;</li>
* <li>the {@code PICKUP_GRACE_POLLS}th consecutive such poll, with {@code WORKING} still never
* observed, releases rather than wedges the roll instead of nudging again — the same
* choice {@code Injector} makes — and returns {@code settled() == true} anyway, logged at
* {@code warn} with the measured elapsed time (fleetd #494) so an operator can see which
* path ran and how long it actually took;</li>
* <li>once {@code WORKING} has been observed, nudging stops and this instead waits for a real
* {@code working → IDLE/DONE} completion boundary before returning {@code true}.</li>
* </ul>
*
* <p><strong>{@link AgentStatus#BLOCKED} is deliberately excluded from both the nudge and the
* boundary check</strong> — the same reasoning as {@link #waitUntilAtTurnBoundary}'s own
* javadoc: a paused live turn is not a settled one, and re-sending Enter into an open approval
* prompt could wrongly answer it. A {@code BLOCKED} sample (or an unreadable/{@link
* AgentStatus#UNKNOWN} one) simply keeps this polling, with no nudge and no release, until either
* a real boundary is reached or {@code settleSeconds} runs out.
*
* <p>{@link AgentControl#submit} can itself throw; a {@link RuntimeException} from it is
* swallowed and logged at {@code debug}, exactly like {@code Injector.java:437-442} — a failed
* nudge must not abort the roll.
*
* @return a {@link ClearSettleResult} whose {@code settled()} is {@code true} once {@code
* /clear} has settled, or once the nudge budget was exhausted with no pickup ever
* observed (released rather than wedged); {@code false} if {@code settleSeconds} elapses
* first — the caller must NOT send {@code bootstrapText} in that case, exactly as before
* this fix. {@code elapsedMillis()} and {@code nudges()} are MEASURED values (from the
* injected {@link #nowMillis} clock and an actual nudge count), never the configured
* {@code settleSeconds} budget (fleetd #494).
*/
private ClearSettleResult waitForClearPickupAndSettle(String target, int settleSeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(settleSeconds);
boolean pickedUp = false; // a WORKING sample has been observed since /clear was sent
int idlePollsAwaitingPickup = 0;
int nudges = 0;
while (nowMillis.getAsLong() < deadline) {
AgentStatus status;
try {
status = agents.status(target);
} catch (RuntimeException e) {
log.debug("lead-rollover: status check failed while waiting for {} to settle after "
+ "/clear: {}", target, e.toString());
status = null;
}
if (status == AgentStatus.WORKING) {
pickedUp = true;
} else if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
if (pickedUp) {
// a real WORKING -> IDLE/DONE completion boundary
return new ClearSettleResult(true, nowMillis.getAsLong() - startMillis, nudges);
}
if (++idlePollsAwaitingPickup >= PICKUP_GRACE_POLLS) {
long elapsedMillis = nowMillis.getAsLong() - startMillis;
// fleetd #494: this release trades a possibly-unsubmitted /clear for progress
// instead of wedging the roll — that trade is deliberate and stays. But it is
// also exactly the case that reported false success in the real incident (the
// whole roll "succeeded" after 438ms of a 20s budget), so raise it to WARN and
// print the MEASURED elapsed time next to the target pane, not just the count.
//
// fleetd #494 follow-up (2nd pass): BOTH numbers in this line must come from
// the loop's own counters, never from the PICKUP_GRACE_POLLS constant.
// `idlePollsAwaitingPickup` and `nudges` each have exactly one write site in
// this loop, on the same branch, so on this branch they cannot differ from
// PICKUP_GRACE_POLLS / PICKUP_GRACE_POLLS - 1 today — no test can prove the
// difference on this line, and printing the counters does not change that.
// What it does buy: one source of truth instead of two, so a later change to
// the loop (an early return, a second increment site, a different exit
// condition) cannot leave this message reporting a number the loop no longer
// produces. The place where `nudges` genuinely varies with the run — and is
// covered by a test that can tell it apart from a constant — is the
// /clear-timeout warn in runRollover, which prints clearResult.nudges().
log.warn("lead-rollover: /clear on {} was never observed as WORKING after {} "
+ "consecutive IDLE/DONE polls ({} of those were nudged) — "
+ "releasing rather than wedging the roll (elapsed={}ms)",
target, idlePollsAwaitingPickup, nudges, elapsedMillis);
return new ClearSettleResult(true, elapsedMillis, nudges);
}
try {
agents.submit(target); // nudge a raced Enter (CB-113) so /clear actually submits
} catch (RuntimeException e) {
log.debug("lead-rollover: resubmit to {} failed (will retry next poll): {}",
target, e.getMessage());
} finally {
nudges++; // an attempted nudge, whether or not the submit call itself threw
}
}
// AgentStatus.BLOCKED or UNKNOWN (or an unreadable status, above): neither a pickup
// signal nor a boundary — keep polling without nudging or releasing.
settleSleeper.run();
}
return new ClearSettleResult(false, nowMillis.getAsLong() - startMillis, nudges);
}
/**
* The measured outcome of {@link #waitForClearPickupAndSettle} — fleetd #494. Carries the
* MEASURED elapsed time (from the injected {@link #nowMillis} clock) and nudge count alongside
* the settle/timeout decision, so callers can log them instead of the configured budget, which
* is not how long the wait actually ran.
*/
private record ClearSettleResult(boolean settled, long elapsedMillis, int nudges) {}
}
@@ -114,13 +114,6 @@ public final class FleetMcp {
* dev.ltms.fleet.herdr.PaneLocator} — that a real assembly wired up. See {@link #identity()}.
*/
private final ConnectionIdentity identity;
/**
* Kept as a field (rather than only captured by the {@code contextExtractor} closure) so
* {@link #denyFor} can read {@link CallerResolver#knownLeadOrCollaborator()} — the classifier a
* collaborator's {@code SEND} is checked against, built from the same lead and collaborator
* maps {@link #identity}-based resolution reads.
*/
private final CallerResolver callers;
private final Metrics metrics; // CB-502: null → auth failures not counted
private final CapacitySource capacity;
private final HealthCoverageSource healthCoverage;
@@ -300,6 +293,14 @@ public final class FleetMcp {
public record LeadConfigDirSource(Function<String, String> configDirFor, Function<String, Long> windowFor) {
/** Inert source — every lead reads {@link LeadContextGauge}'s built-in defaults. */
public static LeadConfigDirSource none() { return new LeadConfigDirSource(_ -> null, _ -> null); }
/**
* Constructor that resolves no window — every lead this source answers for keeps
* {@link LeadContextGauge}'s fixed HIGH threshold.
*/
public LeadConfigDirSource(Function<String, String> configDirFor) {
this(configDirFor, _ -> null);
}
}
/**
@@ -334,7 +335,7 @@ public final class FleetMcp {
* caller explicitly saying so — never by omitting a {@link CallerResolver} the way the old
* {@code callers == null} idiom allowed. {@code callers} itself is required either way: even
* under {@link #UNENFORCED}, the one real {@link CallerResolver} still resolves every caller's
* {@link Principal} (so {@code markTrackedCallerPresent}/{@code recordPrimarySingleton} see a
* {@link Principal} (so {@code markSpawnedMemberPresent}/{@code recordPrimarySingleton} see a
* real identity), and {@link #denyFor} is the only thing that changes.
*/
public enum AuthorizationMode { ENFORCED, UNENFORCED }
@@ -416,7 +417,6 @@ public final class FleetMcp {
this.authorizationEnforced = Objects.requireNonNull(authorizationMode, "authorizationMode")
== AuthorizationMode.ENFORCED;
this.identity = identity;
this.callers = callers;
this.leadChannel = leadChannel;
this.peers = peers == null ? List.of() : List.copyOf(peers);
this.capacity = capacity;
@@ -442,10 +442,10 @@ public final class FleetMcp {
// fall back to here — AuthorizationMode governs enforcement, not identity.
Principal p = callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
req.getHeader("Authorization"));
// Guards on the ROLE, not on the terminal being null — this marks presence for
// a worker, an architect, or the unconfigured-pane floor, and excludes a lead
// or a collaborator even though each carries its own pane too.
markTrackedCallerPresent(p, presence);
// CB-532: guard on the ROLE, not on the terminal being null. This excludes a
// lead, which carries its pane too, while including every spawned member role.
// Enrolling a lead would count it as an available member in the roster.
markSpawnedMemberPresent(p, presence);
return McpTransportContext.create(Map.of(
CALLER_TERMINAL, orEmpty(p.terminal()),
CALLER_PID, Long.toString(p.pid()),
@@ -459,14 +459,12 @@ public final class FleetMcp {
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_send", req.arguments()),
str(req.arguments(), "sessionId"));
if (denied != null) return denied;
Principal caller = principal(exchange);
String callerTerminal = caller.terminal();
String callerOwner = caller.ownerKey();
String caller = callerTerminal(exchange);
// CB-548: only a PRIMARY caller may claim the legacy singleton "primary" fallback.
// An architect delegates as its own pane but must never become the fallback that
// no-delegation inbox nudges target as if it were the primary (the per-target
// delegation map does not cure the singleton).
recordPrimarySingleton(primaryRegistry, callerTerminal, caller);
recordPrimarySingleton(primaryRegistry, caller, principal(exchange));
Map<String, Object> a = req.arguments();
String target = str(a, "sessionId");
String content = str(a, "content");
@@ -482,18 +480,18 @@ public final class FleetMcp {
// Answering a worker's fleet_ask (CB-205): resolve its blocked question and
// block for the worker's reply as it resumes the same turn. This is the same
// delegation, so ownership is left untouched (CB-548) — never re-recorded.
return answer(messages, turnId, content, timeoutMs(a), callerOwner);
return answer(messages, turnId, content, timeoutMs(a));
}
// CB-548: delegator ownership (which lead's reply nudge this worker routes to,
// CB-532) is recorded only once the send is ACCEPTED — MessageService has won the
// session lock and queued delivery — via the accepted-delivery callback, never at
// request time. A concurrent sender that times out BUSY therefore cannot steal a
// live turn's reply routing without ever owning the turn.
Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, callerTerminal);
Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, caller);
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
return Boolean.FALSE.equals(a.get("wait"))
? sendAsync(messages, target, content, onAccepted, workers.profiles(), caller)
: send(messages, target, content, timeoutMs(a), onAccepted, workers.profiles(), callerOwner);
? sendAsync(messages, target, content, onAccepted, workers.profiles())
: send(messages, target, content, timeoutMs(a), onAccepted, workers.profiles());
};
// fleet_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself.
@@ -516,7 +514,7 @@ public final class FleetMcp {
(exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_status", req.arguments()), null);
if (denied != null) return denied;
return status(messages, str(req.arguments(), "sessionId"), principal(exchange).ownerKey());
return status(messages, str(req.arguments(), "sessionId"));
};
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> pollHandler =
(exchange, req) -> {
@@ -526,8 +524,7 @@ public final class FleetMcp {
// The action depends on the ARGUMENTS, not on the tool name -- see pollAction.
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_poll", a), target);
if (denied != null) return denied;
return poll(messages, leadChannel, str(a, "ticket"), target, coordId,
principal(exchange).ownerKey());
return poll(messages, leadChannel, str(a, "ticket"), target, coordId);
};
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
// Acking removes a reply from the inbox, so it is a drain, not a read.
@@ -563,10 +560,8 @@ public final class FleetMcp {
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine, outage,
leadSeats, leadContextGauge, leadConfigDirs, callers.leads(),
callerTerminal(exchange),
callers.collaborators(), collaboratorsVisibleTo(principal(exchange)),
new CoordinationSource(leadChannel, peers),
coordinatorVisibleTo(principal(exchange)),
leadsVisibleTo(principal(exchange)), membersVisibleTo(principal(exchange)));
coordinatorVisibleTo(principal(exchange)));
};
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> stopHandler =
(exchange, req) -> {
@@ -702,8 +697,8 @@ public final class FleetMcp {
if (!authorizationEnforced) {
return null; // AuthorizationMode.UNENFORCED: authorization not enforced (fleetd #518)
}
if (Authz.permits(caller, action, target, callers.knownLeadOrCollaborator())) {
if (action != Authz.Action.READ && action != Authz.Action.TASK_READ) {
if (Authz.permits(caller, action, target)) {
if (action != Authz.Action.READ) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
}
return null;
@@ -748,53 +743,7 @@ public final class FleetMcp {
return caller.isPrimary();
}
/**
* fleetd #703: who may see {@code fleet_list}'s {@code collaborators} array — a roster of the
* human-opened tabs this daemon recognises as named peers. Visible to exactly the roles that
* may {@link Authz.Action#SEND} to a named peer ({@code Authz.java}'s {@code SEND} case):
* the primary, an architect, and a collaborator (reaching another collaborator or a lead). A
* worker holds {@code READ} but can never {@code SEND} to a collaborator, so listing them to a
* worker would expose which human tabs exist on the host with no use to that caller. Split out
* for the same reason as {@link #coordinatorVisibleTo}: the decision must be unit-testable
* without fabricating an SDK {@code McpSyncServerExchange}, and the handler must call this
* named predicate rather than inlining the check.
*/
static boolean collaboratorsVisibleTo(Principal caller) {
return caller.isPrimary() || caller.isArchitect() || caller.isCollaborator();
}
/**
* Who may see {@code fleet_list}'s {@code leads} array — exactly the roles that may
* {@link Authz.Action#SEND} to a lead: the primary, an architect, and a collaborator. A
* collaborator's own {@code fleet_whoami} carries no lead address, and {@code leads} is the
* only place this tool gives one, so a collaborator needs this array to use the send it
* already holds. A worker can never {@code SEND} at all, so it still sees neither this array
* nor {@code members}; a worker's own facts come from {@code fleet_whoami} instead. Split
* out for the same reason as {@link #coordinatorVisibleTo} and
* {@link #collaboratorsVisibleTo}: the decision must be unit-testable without fabricating an
* SDK {@code McpSyncServerExchange}, and the handler must call this named predicate rather
* than inlining the check.
*/
static boolean leadsVisibleTo(Principal caller) {
return caller.isPrimary() || caller.isArchitect() || caller.isCollaborator();
}
/**
* Who may see {@code fleet_list}'s {@code members} array — the primary and an architect,
* which may {@link Authz.Action#SEND} to a member. A worker can never {@code SEND} at all,
* and a collaborator may {@code SEND} only to a lead or another collaborator, never to a
* spawned member, so neither sees this array even though {@link #leadsVisibleTo} grants a
* collaborator the sibling one.
*/
static boolean membersVisibleTo(Principal caller) {
return caller.isPrimary() || caller.isArchitect();
}
/**
* The terminal of the caller on this call's connection, or {@code null} when that caller carries
* no terminal, which is only the unnamed primary. A named lead, an architect and a worker each
* carry one.
*/
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
private static String callerTerminal(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_TERMINAL);
String s = v == null ? null : v.toString();
@@ -836,14 +785,9 @@ public final class FleetMcp {
return identity;
}
/**
* Mark a caller present for the injector readiness gate, when its deliverability depends on
* proving a live MCP contact: a worker, an architect, or the unconfigured-pane floor. A lead
* or a collaborator is excluded — each is already deliverable through its own named-registry
* entry.
*/
static void markTrackedCallerPresent(Principal caller, MemberPresence presence) {
if (caller.isSpawnedMember() || caller.isObserver()) {
/** Mark a connected spawned member available for the injector readiness gate. */
static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) {
if (caller.isSpawnedMember()) {
presence.markPresent(caller.terminal());
}
}
@@ -899,8 +843,8 @@ public final class FleetMcp {
* fleetd #612 B3 — as {@link #quarantineSource()}, {@code public} for the same cross-package
* reason, for the real {@link LeadRollover} (or {@code null}) this daemon was assembled with.
* {@code FleetdLeadRolloverAssemblyTest} drives {@code open}/{@code confirm} on this exact
* instance and waits for the real continuation to end the old pane, relaunch a fresh one, and
* send {@code bootstrapText} through the real {@code router.leadAgents()}.
* instance and waits for the real continuation to send {@code /clear} and {@code bootstrapText}
* through the real {@code router.leadAgents()}.
*/
public LeadRollover leadRollover() {
return leadRollover;
@@ -916,8 +860,7 @@ public final class FleetMcp {
* a BUSY interloper never claims a turn it did not win. {@code null} disables recording.
*/
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content,
Long timeoutMs, Runnable onAccepted, Set<String> profiles,
String callerOwner) {
Long timeoutMs, Runnable onAccepted, Set<String> profiles) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
@@ -927,7 +870,7 @@ public final class FleetMcp {
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
try {
return formatReply(messages.send(sessionId, content, timeout, onAccepted, callerOwner), timeout);
return formatReply(messages.send(sessionId, content, timeout, onAccepted), timeout);
} catch (HerdrException e) {
return error("herdr error contacting session " + sessionId + ": " + e.getMessage());
}
@@ -937,15 +880,13 @@ public final class FleetMcp {
* {@code fleet_send} carrying a {@code turnId}: the primary's answer to a worker's
* {@code fleet_ask} (CB-205). Resolves the worker's blocked question and blocks for its reply as
* it resumes the same turn — surfaced to the primary identically to a normal send.
* {@code callerOwner} must match the turn's recorded owner or this is refused.
*/
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs,
String callerOwner) {
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs) {
if (isBlank(turnId) || isBlank(content)) {
return error("turnId and content are required to answer a worker's question");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
return formatReply(messages.answer(turnId, content, timeout, callerOwner), timeout);
return formatReply(messages.answer(turnId, content, timeout), timeout);
}
/**
@@ -994,8 +935,6 @@ public final class FleetMcp {
+ "\" and content set to your answer; the worker resumes the same turn.");
case STALE_TURN -> error("that question is no longer open — it timed out or was already "
+ "answered (turnId stale)");
case NOT_TURN_OWNER -> error("this turn belongs to a different delegation — only the caller "
+ "that opened it may answer it");
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> text("[no reply within " + timeout + "ms — worker "
+ r.outcome().name().toLowerCase().replace("timed_out_", "") + "; retry or poll status]");
// fleetd #571: delivery is unknown here — agent.prompt pastes and submits in one call,
@@ -1017,16 +956,6 @@ public final class FleetMcp {
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content,
Runnable onAccepted, Set<String> profiles) {
return sendAsync(messages, sessionId, content, onAccepted, profiles, null);
}
/**
* As above, recording {@code creator}'s owner key so a later {@code fleet_poll{ticket}} only
* hands the result back to the same caller — see {@link MessageService#poll(String, String)}.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content,
Runnable onAccepted, Set<String> profiles,
Principal creator) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
@@ -1034,7 +963,7 @@ public final class FleetMcp {
if (targetError != null) {
return targetError;
}
String ticket = messages.sendAsync(sessionId, content, onAccepted, creator);
String ticket = messages.sendAsync(sessionId, content, onAccepted);
return text("accepted — task delegated. Poll fleet_poll with ticket=" + ticket);
}
@@ -1114,21 +1043,31 @@ public final class FleetMcp {
}
/**
* Which authorization action a {@code fleet_poll} call needs, decided by its arguments.
* Which authorization action a {@code fleet_poll} call needs, decided by its arguments
* (fleetd #272, widened by fleetd #421).
*
* <p>{@code fleet_poll} is <strong>three operations behind one tool name</strong>. With
* <p>{@code fleet_poll} is now <strong>three operations behind one tool name</strong>. With
* {@code ticket} it observes an async delegation and changes nothing, which is a {@link
* Authz.Action#TASK_READ}. With {@code target} it calls {@link MessageService#drainReplies} on
* that session -- the replies are removed from the inbox and a second call returns nothing --
* so it is a {@link Authz.Action#DRAIN}, the same gate {@code fleet_ack} already uses for
* removing a single message, and the same one the REST path uses at
* {@code FleetApp.drainReplies}. With {@code coordId} it reads (never acks) this daemon's own
* held lead-to-lead mail, which is a {@link Authz.Action#COORD_READ} -- <strong>not</strong>
* {@code TASK_READ} or {@code READ}: a lead-to-lead body is a different inbox from either, and
* folding it into either would let any worker or architect read every peer lead's mail in full.
* Authz.Action#READ}. With {@code target} it calls {@link MessageService#drainReplies} on that
* session -- the replies are removed from the inbox and a second call returns nothing -- so it
* is a {@link Authz.Action#DRAIN}, the same gate {@code fleet_ack} already uses for removing a
* single message, and the same one the REST path uses at {@code FleetApp.drainReplies}. With
* {@code coordId} it reads (never acks) this daemon's own held lead-to-lead mail, which is a
* {@link Authz.Action#COORD_READ} -- <strong>not</strong> {@code READ}, even though nothing is
* consumed: {@code READ}'s grant is open to every authenticated role on the premise that the
* roster carries no secrets, and a lead-to-lead body is not the roster. Mapping a non-destructive
* peer-mail read to {@code READ} would let any worker read every peer lead's mail in full.
*
* <p>Before this method existed (fleetd #272) the handler passed a constant {@code READ} for
* both of the original branches. {@code READ} is open to every authenticated role, so any
* worker could read a peer's id out of {@code fleet_list} and destroy the replies that peer had
* queued for the primary. The gate failed open, and it did so because the required action is a
* function of the arguments while the handler chose it before looking at them.
*
* <p>The choice lives in this method, and not inline in the handler, so that a test can assert
* the mapping the handler actually uses.
* the mapping the handler actually uses. {@code FleetMcpAuthzTest} already checked every
* {@link Authz.Action} against every {@link Role} and passed throughout -- it tested the policy
* table, which was correct, while the defect was in which action the caller handed it.
*
* <p>Checked first, and exclusively of {@code target}: a call naming {@code coordId} is reading
* a different inbox entirely (this daemon's own lead channel, never a worker's), so it takes
@@ -1141,30 +1080,7 @@ public final class FleetMcp {
if (!isBlank(coordId)) {
return Authz.Action.COORD_READ;
}
return isBlank(target) ? Authz.Action.TASK_READ : Authz.Action.DRAIN;
}
/**
* Which authorization action a {@code fleet_send} call needs, decided by its arguments.
*
* <p>{@code fleet_send} is three call shapes behind one tool name, mirroring {@link
* #pollAction}. With {@code coordId} it addresses a peer lead on another daemon over the
* coordination broker, which is {@link Authz.Action#COORD_SEND}. With {@code turnId} it
* resolves a worker's blocked {@code fleet_ask} and resumes that turn, which is {@link
* Authz.Action#ANSWER}. Otherwise it delivers to a local session by {@code sessionId}, which is
* the plain {@link Authz.Action#SEND}.
*
* <p>Checked in the same order the handler branches: {@code coordId} first and exclusively of
* {@code turnId}, matching {@link #sendToLead}'s own mutual-exclusion check.
*
* @param coordId the {@code coordId} argument of the call, or {@code null}/blank when absent
* @param turnId the {@code turnId} argument of the call, or {@code null}/blank when absent
*/
static Authz.Action sendAction(String coordId, String turnId) {
if (!isBlank(coordId)) {
return Authz.Action.COORD_SEND;
}
return isBlank(turnId) ? Authz.Action.SEND : Authz.Action.ANSWER;
return isBlank(target) ? Authz.Action.READ : Authz.Action.DRAIN;
}
/**
@@ -1189,11 +1105,10 @@ public final class FleetMcp {
*/
private static Authz.Action authzAction(FleetTool tool, Map<String, Object> arguments) {
return switch (tool) {
case SEND -> sendAction(str(arguments, "coordId"), str(arguments, "turnId"));
case SEND -> Authz.Action.SEND;
case REPLY -> Authz.Action.REPLY;
case ASK -> Authz.Action.ASK;
case STATUS -> Authz.Action.TASK_READ;
case LIST, PROFILES, WHOAMI -> Authz.Action.READ;
case STATUS, LIST, PROFILES, WHOAMI -> Authz.Action.READ;
case POLL -> pollAction(str(arguments, "target"), str(arguments, "coordId"));
case ACK -> Authz.Action.DRAIN;
case SPAWN -> Authz.Action.SPAWN;
@@ -1213,23 +1128,9 @@ public final class FleetMcp {
* exclusively of {@code ticket}/{@code target} — see {@link #pollAction}'s javadoc for why this
* is a different inbox (this daemon's own {@link LeadChannel}) that authorizes differently
* ({@link Authz.Action#COORD_READ}, primary-only) from either of the original two branches.
*
* <p>Does not check who owns {@code ticket} — see the overload that takes {@code
* callerTerminal} for that. Callers that do not resolve a caller terminal (tests, or a surface
* with no connection-based identity) use this one.
*/
static McpSchema.CallToolResult poll(MessageService messages, LeadChannel leadChannel, String ticket,
String target, String coordId) {
return poll(messages, leadChannel, ticket, target, coordId, null);
}
/**
* As above, refusing a ticket lookup whose caller owner key differs from the key that created it
* — see {@link MessageService#poll(String, String)}. {@code callerOwner} comes from the calling
* connection's resolved principal, never a client-supplied value.
*/
static McpSchema.CallToolResult poll(MessageService messages, LeadChannel leadChannel, String ticket,
String target, String coordId, String callerOwner) {
if (!isBlank(coordId)) {
return pollHeldPeerMail(leadChannel, coordId);
}
@@ -1243,7 +1144,7 @@ public final class FleetMcp {
if (isBlank(ticket)) {
return error("ticket (or target) is required");
}
MessageService.TaskView v = messages.poll(ticket, callerOwner);
MessageService.TaskView v = messages.poll(ticket);
if (v == null) {
return error("unknown ticket: " + ticket + " (never issued, or expired)");
}
@@ -1350,19 +1251,16 @@ public final class FleetMcp {
/**
* {@code fleet_status}: the live lifecycle status of a worker session, plus — when the worker
* is paused mid-turn in an async {@code fleet_ask} — the open question and how to answer it, so
* a lead on its normal poll cadence does not need the ticket to notice. The question, its
* {@code turnId} and its ticket id are shown only to the caller whose owner key created that
* delegation, or to the unnamed primary; any other caller still sees the base status.
* {@code callerOwner} comes from the calling connection's resolved principal.
* is paused mid-turn in an async {@code fleet_ask} (CB-582) — the open question and how to
* answer it, so a lead on its normal poll cadence does not need the ticket to notice.
*/
static McpSchema.CallToolResult status(MessageService messages, String sessionId, String callerOwner) {
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
if (isBlank(sessionId)) {
return error("sessionId is required");
}
try {
String base = messages.status(sessionId).name().toLowerCase();
MessageService.PendingAsk ask = messages.pendingAsk(sessionId, callerOwner);
MessageService.PendingAsk ask = messages.pendingAsk(sessionId);
if (ask == null) {
return text(base);
}
@@ -1405,26 +1303,6 @@ public final class FleetMcp {
}
return text(json(m));
}
if (caller.isCollaborator()) {
// A collaborator's name is its slot in the collaborators: registry; sessionId is its
// pane so a peer knows where to reach it. No leader key: a collaborator is not a
// primary for authorization, unlike a lead.
if (caller.name() != null) {
m.put("collaborator", caller.name());
}
if (caller.terminal() != null) {
m.put("sessionId", caller.terminal());
}
return text(json(m));
}
if (caller.isObserver()) {
// No architect slot, collaborator name, or lead name to report — only the pane itself,
// so a peer that already knows this terminal can still address it.
if (caller.terminal() != null) {
m.put("sessionId", caller.terminal());
}
return text(json(m));
}
if (!caller.isWorker()) {
// CB-530: which lead, once more than one pane is configured as one. `role` deliberately
// still reads "primary" — the fallback ladder in CLAUDE.md keys on it, and a lead IS a
@@ -1505,7 +1383,7 @@ public final class FleetMcp {
}
if (isBlank(callerTerminal)) {
// An unnamed primary (token/loopback path, no resolved pane) has nowhere for the
// eventual relaunch + bootstrap to land — LeadRollover#open would throw
// eventual /clear + bootstrap to land — LeadRollover#open would throw
// IllegalArgumentException for the same reason; refuse cleanly here instead.
return error("fleet_handover requires a named lead pane (a resolved connection terminal) "
+ "to open a rollover request against — an unnamed primary has none");
@@ -1857,8 +1735,7 @@ public final class FleetMcp {
Map<String, String> leads, String selfTerm,
CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine,
OutageSource.none(), LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, false, true, true);
OutageSource.none(), LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
}
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
@@ -1867,8 +1744,7 @@ public final class FleetMcp {
QuarantineSource quarantine, OutageSource outage,
Map<String, String> leads, String selfTerm) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, CoordinationSource.none(), false, true, true);
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, CoordinationSource.none(), false);
}
/**
@@ -1885,8 +1761,7 @@ public final class FleetMcp {
QuarantineSource quarantine, Map<String, String> leads, String selfTerm,
CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, OutageSource.none(),
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, false, true, true);
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
}
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
@@ -1895,8 +1770,7 @@ public final class FleetMcp {
QuarantineSource quarantine, OutageSource outage,
Map<String, String> leads, String selfTerm, CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, false, true, true);
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
}
/**
@@ -1918,8 +1792,7 @@ public final class FleetMcp {
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, false, true, true);
leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
}
/**
@@ -1936,23 +1809,14 @@ public final class FleetMcp {
* {@code false} (fleetd #463: a forgotten argument fails closed, not
* open), so a test that wants the {@code coordinator} row must pass
* an explicit {@code true}
* @param leadsVisible whether this caller may see the {@code leads} array (see
* {@link #leadsVisibleTo}) — unlike {@code callerIsPrimary}, this is
* also {@code true} for an architect or a collaborator, so it cannot
* be derived from {@code callerIsPrimary} alone
* @param membersVisible whether this caller may see the {@code members} array (see
* {@link #membersVisibleTo}); also {@code true} for an architect, but
* unlike {@code leadsVisible}, never for a collaborator
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, OutageSource outage,
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
CoordinationSource coordination, boolean callerIsPrimary,
boolean leadsVisible, boolean membersVisible) {
CoordinationSource coordination, boolean callerIsPrimary) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine,
outage, leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, callerIsPrimary, leadsVisible, membersVisible);
outage, leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, callerIsPrimary);
}
/**
@@ -1962,17 +1826,6 @@ public final class FleetMcp {
* matter); the one caller that matters for caching, {@code fleet_list}'s MCP handler, passes
* its own single long-lived instance instead (see {@code FleetMcp}'s {@code leadContextGauge}
* field).
*
* @param collaborators terminal_id → collaborator name, live from the resolver
* ({@link dev.ltms.fleet.auth.CallerResolver#collaborators()})
* @param collaboratorsVisible whether this caller may see the {@code collaborators} array (see
* {@link #collaboratorsVisibleTo}); every wrapper overload above
* passes {@code false}, so a test that wants the row must call this
* overload with an explicit {@code true}
* @param leadsVisible whether this caller may see the {@code leads} array (see
* {@link #leadsVisibleTo})
* @param membersVisible whether this caller may see the {@code members} array (see
* {@link #membersVisibleTo})
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
@@ -1981,56 +1834,32 @@ public final class FleetMcp {
LeadSeatSource leadSeats, LeadContextGauge contextGauge,
LeadConfigDirSource leadConfigDirs,
Map<String, String> leads, String selfTerm,
Map<String, String> collaborators, boolean collaboratorsVisible,
CoordinationSource coordination, boolean callerIsPrimary,
boolean leadsVisible, boolean membersVisible) {
CoordinationSource coordination, boolean callerIsPrimary) {
try {
// Neither row's assembly (leadView/memberCapacityView probing herdr for live status)
// runs unless at least one of them needs the live-agent lookup backing it.
Map<String, Agent> live = (leadsVisible || membersVisible)
? workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b))
: Map.of();
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> leadRows = leads.entrySet().stream()
.sorted(Map.Entry.comparingByValue())
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm, contextGauge,
leadConfigDirs))
.toList();
// fleetd #209: this is the caller-driven fleet_list read that actually reports
// agentSessionId (via memberCapacityView -> SessionManager.rosterView), so it uses the
// resolving roster; the heartbeat/health/metrics timers stay on the plain sessions.roster().
List<MemberSession> roster = sessions.rosterResolved();
List<Map<String, Object>> out = roster.stream()
.map(s -> memberCapacityView(s, live.get(s.terminalId()), messages, capacity.clock().getAsLong()))
.toList();
Set<String> profiles = new java.util.TreeSet<>(capacity.configuredProfiles().get());
roster.stream().map(MemberSession::profile).forEach(profiles::add);
Map<String, Object> result = new LinkedHashMap<>();
// READ is permission to enter this tool, not permission to receive every field it can
// build -- gate BEFORE assembling each row, so the key is absent rather than
// present-and-empty; a caller without either row gets its own facts from fleet_whoami.
if (leadsVisible) {
List<Map<String, Object>> leadRows = leads.entrySet().stream()
.sorted(Map.Entry.comparingByValue())
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm, contextGauge,
leadConfigDirs))
.toList();
result.put("leads", leadRows);
}
if (membersVisible) {
List<Map<String, Object>> out = roster.stream()
.map(s -> memberCapacityView(s, live.get(s.terminalId()), messages, capacity.clock().getAsLong()))
.toList();
result.put("members", out);
}
result.put("leads", leadRows); result.put("members", out);
result.put("healthCoverage", healthCoverage.value().get());
result.put("loopHealth", Map.of(
"statusPoller", loopHealth.statusPoller().get().name(),
"sessionReaper", loopHealth.sessionReaper().get().name()));
// fleetd #703: a collaborator tab is a person's own tab, so the row is assembled and
// included only for the roles that may SEND to a named peer -- gate BEFORE assembling
// it, same reason as the coordinator row just below: the key must be absent for a
// worker, never present-and-empty.
if (collaboratorsVisible) {
result.put("collaborators", collaborators.entrySet().stream()
.sorted(Map.Entry.comparingByValue())
.map(e -> collaboratorRow(e.getKey(), e.getValue()))
.toList());
}
// fleetd #439: coordinator/coordinatorView is lead-to-lead coordination state and must
// never reach a worker or an architect -- gate BEFORE assembling it, not after, so the
// key is absent rather than present-and-empty.
@@ -2342,20 +2171,6 @@ public final class FleetMcp {
return m;
}
/**
* fleetd #703: one row of the {@code collaborators} array — a named peer's registry name and
* the herdr {@code terminal_id} a {@code fleet_send} must target to reach it. No status, no
* context gauge, no profile: a collaborator is never spawned and carries no profile, so those
* fields have no meaning for it, and the scan behind {@code terminal} only reports a tab it
* actually found in herdr, so a tab nobody has open does not appear here at all.
*/
private static Map<String, Object> collaboratorRow(String terminal, String name) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("name", name);
m.put("sessionId", terminal);
return m;
}
/**
* Reads the lead context gauge for one lead. {@code configDir} is this lead's configured
* {@code CLAUDE_CONFIG_DIR} override (see {@link LeadConfigDirSource}), derived from
@@ -2550,14 +2365,7 @@ public final class FleetMcp {
private static McpSchema.Tool listTool() {
return tool(FleetTool.LIST.wireName(),
"List the whole fleet the bridge tracks, in two parts. 'members' is visible to "
+ "the primary and an architect only. 'leads' is visible to those two AND a "
+ "collaborator — exactly the roles that may fleet_send to a lead, so a "
+ "collaborator can learn a lead's sessionId before using the send it already "
+ "holds. A worker holds READ to call this tool at all, but gets neither "
+ "array, never an empty one; a worker reads its own session, "
+ "profile, state, worktree, branch and owner from fleet_whoami instead. "
+ "'leads' are your PEERS — other "
"List the whole fleet the bridge tracks, in two parts. 'leads' are your PEERS — other "
+ "orchestrators, each with its sessionId (the address to fleet_send to), "
+ "name, live status, and 'self': true on your own row; this is how you "
+ "discover a peer lead without being told its address. 'members' are the "
@@ -2588,12 +2396,7 @@ public final class FleetMcp {
+ "msgId/from/preview, never the full body), and one row per coordinator.peers "
+ "coord-id ('peers': coordId/reachable, plus pending/consumers when reachable) — "
+ "this is peer DISCOVERY for cross-host leads, distinct from the local 'leads' "
+ "array above. It is omitted entirely when no coordinator is configured. A "
+ "'collaborators' array, visible only to the primary, an architect, and a "
+ "collaborator (never a worker), reports every other named peer tab this daemon "
+ "recognises: each row is 'name' (its registry name) and 'sessionId' (the "
+ "terminal id a fleet_send targets to reach it). Absent entirely for a worker, "
+ "whatever is configured.",
+ "array above. It is omitted entirely when no coordinator is configured.",
objectSchema(Map.of(), List.of()));
}
@@ -2643,20 +2446,17 @@ public final class FleetMcp {
private static McpSchema.Tool handoverTool() {
return tool(FleetTool.HANDOVER.wireName(),
"Replace your OWN lead session once its context is full: write a handover file, "
+ "then use this to have fleetd end your pane's process and relaunch a fresh "
+ "lead session bootstrapped against it. Four actions: 'open' (requests a "
+ "token and the handoverPath you must write the handover file to before "
+ "confirming), 'confirm' (validates every gate and — only if every one "
+ "passes — schedules the roll; it does NOT itself end your pane, the roll "
+ "runs once this call's own turn ends), 'cancel' (drops a pending request "
+ "without rolling), and 'status' (read-only: what happened to a token after "
+ "'confirm' — still running (approved but not finished yet), the roll "
+ "completed, the calling turn never settled within turnSettleSeconds so "
+ "nothing was touched, the old pane never confirmed dead so no relaunch was "
+ "attempted, the relaunch itself failed, the fresh pane never became ready "
+ "so bootstrapText was never sent, or the fresh pane became ready and was "
+ "bootstrapped but was never recognised as a live lead; never schedules, "
+ "cancels or retries anything). Primary-only. "
+ "then use this to have fleetd clear your pane and bootstrap a fresh lead "
+ "session against it. Four actions: 'open' (requests a token and the "
+ "handoverPath you must write the handover file to before confirming), "
+ "'confirm' (validates every gate and — only if every one passes — schedules "
+ "the roll; it does NOT itself clear the pane, the roll runs once this call's "
+ "own turn ends), 'cancel' (drops a pending request without rolling), and "
+ "'status' (read-only: what happened to a token after 'confirm' — still "
+ "running (approved but not finished yet), the roll completed, the calling "
+ "turn never settled within turnSettleSeconds so no /clear was ever sent, or "
+ "/clear itself never settled so bootstrapText was never sent; never "
+ "schedules, cancels or retries anything). Primary-only. "
+ "There is deliberately no terminal/session/leadTerminal parameter: the pane "
+ "to roll is always resolved from YOUR OWN connection, never a value you "
+ "pass, so you can only ever roll yourself — never another lead. Requires "
@@ -6,7 +6,6 @@ import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.ResilientAgentLaunch;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl;
@@ -68,6 +67,16 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(HerdrPeerLauncher.class);
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
/**
* Retries for {@code agent.start} against a seed pane whose shell has not reached its prompt
* yet — {@code tab.create}/{@code pane.split} return as soon as the pane exists, and herdr
* refuses to start an agent in a pane that is not "an available shell" ({@code agent_pane_busy}).
*/
private static final int SHELL_READY_RETRIES = 20;
private final String namePrefix; // label prefix: naming + reap scheme
private final AgentControl agents;
private final WorkspaceControl spaces;
@@ -757,26 +766,90 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
// argv[0] — the configured executable — is dropped and only the extra args are passed.
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
try {
ResilientAgentLaunch.checkFits(cfg.profile(), argv);
} catch (ResilientAgentLaunch.TooLargeException e) {
throw new PeerUnreachableException(e.getMessage());
checkPaneCommandFits(cfg, argv);
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(startAwaitingShellPrompt(name, args, paneId), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("peer name '{}' taken, retrying", name);
last = e;
}
}
long[] lastSeq = {0};
Agent agent = ResilientAgentLaunch.startUniquelyNamed(agents, namePrefix, args, paneId,
attempt -> {
lastSeq[0] = nameSeq.incrementAndGet();
return namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + lastSeq[0];
},
ResilientAgentLaunch.NAME_RETRIES, ResilientAgentLaunch.SHELL_READY_RETRIES, sleeper);
return new Started(agent, lastSeq[0]);
throw last;
}
/**
* fleetd #220: herdr does not exec the launch command — it TYPES it into the pane as one line,
* and a pty line buffer holds only {@value #PANE_COMMAND_BYTE_LIMIT} bytes (BSD/macOS {@code
* MAX_CANON}). Everything past that byte is dropped. Nothing reports it: herdr answers "agent
* started", the backend exits on the mangled argument it was handed, the pane closes, and the
* only symptom is {@link #waitUntilInjectableOrThrow} timing out 20 seconds later with no
* reason. That is exactly how #214 broke every claude-code spawn — one 50-byte flag pushed a
* 978-byte command to 1028, and the tail that got cut was {@code --autocompact 250000}.
*
* <p>So measure it here and refuse, loudly and immediately, rather than spawn something that
* cannot work. The estimate is deliberately conservative: fleetd cannot see herdr's quoting, so
* every argument is charged its own bytes plus a separator and a quote pair. An over-estimate
* costs a clear error at a length that was already unsafe; an under-estimate would let the
* silent truncation back in.
*
* @throws PeerUnreachableException when the command cannot fit — the same failure the spawn
* would have hit anyway, named at the point it is still
* explainable
*/
private void checkPaneCommandFits(FleetConfig.Profile cfg, List<String> argv) {
int bytes = 0;
String longest = null;
int longestBytes = 0;
for (String arg : argv) {
int argBytes = arg == null ? 0 : arg.getBytes(java.nio.charset.StandardCharsets.UTF_8).length;
bytes += argBytes + QUOTING_OVERHEAD_PER_ARG;
if (argBytes > longestBytes) {
longestBytes = argBytes;
longest = arg;
}
}
if (bytes <= PANE_COMMAND_BYTE_LIMIT) {
return;
}
String culprit = longest == null ? "<none>"
: longest.substring(0, Math.min(longest.length(), 60)) + (longest.length() > 60 ? "…" : "");
throw new PeerUnreachableException(
"launch command for profile " + cfg.profile() + " is about " + bytes + " bytes, over the "
+ PANE_COMMAND_BYTE_LIMIT + "-byte limit of the pane line herdr types it into. "
+ "The pty would drop the tail silently and the backend would exit on a mangled "
+ "argument. Longest argument is " + longestBytes + " bytes: " + culprit
+ " — move it off the command line (a file flag) or shorten it.");
}
/**
* The pty line buffer herdr types a launch command into: BSD/macOS {@code MAX_CANON}. Not a
* fleetd choice and not configurable — see {@link ResilientAgentLaunch#checkFits}.
* fleetd choice and not configurable — see {@link #checkPaneCommandFits}.
*/
static final int PANE_COMMAND_BYTE_LIMIT = ResilientAgentLaunch.PANE_COMMAND_BYTE_LIMIT;
static final int PANE_COMMAND_BYTE_LIMIT = 1024;
/** Per-argument allowance for the separating space and a shell quote pair fleetd cannot see. */
private static final int QUOTING_OVERHEAD_PER_ARG = 3;
/** Start the agent into {@code paneId}, waiting out the seed shell's boot with the sleeper. */
private Agent startAwaitingShellPrompt(String name, List<String> args, String paneId) {
HerdrException busy = null;
for (int attempt = 0; attempt < SHELL_READY_RETRIES; attempt++) {
try {
return agents.start(name, namePrefix, args, paneId);
} catch (HerdrException e) {
if (!"agent_pane_busy".equals(e.code())) throw e;
log.debug("pane {} not at its shell prompt yet, retrying agent.start", paneId);
busy = e;
sleeper.run();
}
}
throw busy;
}
// --- discovery + reap ----------------------------------------------------------------------
@@ -1,6 +1,5 @@
package dev.ltms.fleet.msg;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrRouter;
@@ -88,7 +87,7 @@ public final class MessageService {
/**
* The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the
* question and {@code turnId} correlates the answer. Not terminal — the primary answers with
* {@link #answer(String, String, long, String)} and the turn resumes.
* {@link #answer(String, String, long)} and the turn resumes.
*/
QUESTION,
/** Timed out after the message was delivered — the worker is still working. */
@@ -121,16 +120,10 @@ public final class MessageService {
/** Another send to this session was in flight for the whole window. */
BUSY,
/**
* An answer ({@link #answer(String, String, long, String)}) referenced a {@code turnId}
* that is no longer open — the worker's {@code fleet_ask} already timed out or was answered.
* An answer ({@link #answer(String, String, long)}) referenced a {@code turnId} that is no
* longer open — the worker's {@code fleet_ask} already timed out or was answered.
*/
STALE_TURN,
/**
* An answer ({@link #answer(String, String, long, String)}) named a {@code turnId} that is
* still open, but the answering caller is not the caller whose accepted delegation opened
* it. Distinct from {@link #STALE_TURN} so a refusal is never reported as a lapsed turn.
*/
NOT_TURN_OWNER
STALE_TURN
}
/**
@@ -140,7 +133,7 @@ public final class MessageService {
* {@link Outcome#COMPLETED_UNREPLIED}), or the question for {@link Outcome#QUESTION},
* else {@code null}
* @param turnId correlation id for a {@link Outcome#QUESTION} (answered via
* {@link #answer(String, String, long, String)}), else {@code null}
* {@link #answer(String, String, long)}), else {@code null}
*/
public record Reply(Outcome outcome, String text, String turnId) {
/** A reply with no correlation id (the common terminal outcomes). */
@@ -282,16 +275,10 @@ public final class MessageService {
* {@link #abandon}) can never match again regardless of this flag's value.
*/
private volatile boolean askTimedOut;
/**
* The owner key of the caller whose {@code fleet_send{wait:false}} created this ticket, or
* {@code null} for the unnamed primary and overloads that do not record a caller.
*/
private final String creatorOwner;
private Task(String ticket, String target, LongSupplier nowNanos, String creatorOwner) {
private Task(String ticket, String target, LongSupplier nowNanos) {
this.ticket = ticket;
this.target = target;
this.creatorOwner = creatorOwner;
this.createdNanos = nowNanos.getAsLong();
future.whenComplete((reply, ex) -> completedNanos = nowNanos.getAsLong());
}
@@ -353,14 +340,6 @@ public final class MessageService {
*/
private final ConcurrentHashMap<String, Boolean> queuedDeliveries = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
/**
* Minted once per {@code MessageService} instance and folded into every ticket id (see
* {@link #sendAsync(String, String, Runnable, Principal)}). {@link #ticketSeq} alone restarts at
* zero for every instance, so without this a ticket id can be reused across instances and
* resolve to an unrelated {@link Task} with no error; this nonce makes that impossible, because
* an id minted by one instance can never match the id space of another.
*/
private final String ticketBootNonce = UUID.randomUUID().toString().substring(0, 6);
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
Thread.ofVirtual().name("bridge-async-", 0).factory());
@@ -677,7 +656,7 @@ public final class MessageService {
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, TIMED_OUT_UNCONFIRMED, BUSY -> "timeout";
case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted";
case STALE_TURN, QUESTION, NOT_TURN_OWNER -> null; // not a completed delegation
case STALE_TURN, QUESTION -> null; // not a completed delegation
};
}
@@ -936,17 +915,14 @@ public final class MessageService {
/**
* Deliver {@code content} to {@code target} (a herdr {@code terminal_id}) and block until the
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses. {@code callerOwner}
* identifies the caller making this call and is recorded as the turn's owner. It is the only
* caller {@link #answer(String, String, long, String)} will
* later accept an answer from if the worker pauses mid-turn to ask.
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
*/
public Reply send(String target, String content, long timeoutMillis, String callerOwner) {
return send(target, content, timeoutMillis, null, callerOwner);
public Reply send(String target, String content, long timeoutMillis) {
return send(target, content, timeoutMillis, null);
}
/**
* As {@link #send(String, String, long, String)}, but with an accepted-delivery hook.
* As {@link #send(String, String, long)}, but with an accepted-delivery hook.
*
* <p>{@code onAccepted} is invoked exactly once, once this send has won {@code target}'s send
* lock and so become the <em>accepted target turn</em> — it runs <em>before</em> delivery is
@@ -957,13 +933,12 @@ public final class MessageService {
* acceptance means a concurrent sender that times out {@code BUSY} can never steal ownership it
* never earned. {@code null} disables the hook.
*/
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, String callerOwner) {
return send(target, content, timeoutMillis, onAccepted, null, callerOwner);
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted) {
return send(target, content, timeoutMillis, onAccepted, null);
}
/** Run a send, optionally stopping an async task that teardown already failed before acceptance. */
private Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, Task task,
String callerOwner) {
private Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, Task task) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
@@ -983,7 +958,7 @@ public final class MessageService {
// open race). Opening first also means a throwing onAccepted (fired before enqueue) or an
// enqueue failure is safely closed by the finally below: nothing is left queued, and the
// failed send leaves no stale waiter behind.
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target, Rendezvous.Owner.of(callerOwner));
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
// CB-640: this send now owns target's delivery, so any earlier stranded-reply or
// still-queued fact no longer describes the live state — clear both rather than let
// them outlive the send that supersedes them.
@@ -1162,30 +1137,19 @@ public final class MessageService {
* mid-turn (already picked up), so the answer flows back through its own open {@code fleet_ask}
* call, not a new status-gated delivery. The forward waiter is opened <em>before</em> the worker
* is unblocked so a reply that lands the instant it resumes is not lost.
*
* <p>{@code callerOwner} identifies the caller making this call. It is checked against the
* turn's recorded owner (the caller whose
* accepted delegation opened it, see {@link #send(String, String, long, String)} and
* {@link #sendAsync(String, String, Runnable, Principal)}) before anything else runs: a mismatch,
* including a turn with no owner on record at all, returns {@link Outcome#NOT_TURN_OWNER}
* without touching the rendezvous, the session lock, or any async task bookkeeping.
*/
public Reply answer(String turnId, String content, long timeoutMillis, String callerOwner) {
public Reply answer(String turnId, String content, long timeoutMillis) {
String workerSession = rendezvous.askSession(turnId);
if (workerSession == null) {
return new Reply(Outcome.STALE_TURN, null); // the ask lapsed (timed out or already answered)
}
Rendezvous.Owner owner = rendezvous.askOwner(turnId);
if (!Rendezvous.Owner.permits(owner, callerOwner)) {
return new Reply(Outcome.NOT_TURN_OWNER, null);
}
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(workerSession, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null);
}
try {
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession, owner);
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession);
// fleetd #575: this try used to open below, AFTER the Task lookup/registration and the
// STALE_TURN early return that follows it — so that return was covered only by a
// hand-rolled copy of the finally's own cleanup pair, not the finally itself. Widening the
@@ -1315,20 +1279,8 @@ public final class MessageService {
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content, Runnable onAccepted) {
return sendAsync(target, content, onAccepted, null);
}
/**
* As {@link #sendAsync(String, String, Runnable)}, recording {@code creator}'s owner key as this
* ticket's owner. The key is derived here from the resolved principal so callers cannot pass a
* terminal address where an owner identity is required.
*
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content, Runnable onAccepted, Principal creator) {
String ticket = "task-" + ticketBootNonce + "-" + ticketSeq.incrementAndGet();
String creatorOwner = creator == null ? null : creator.ownerKey();
Task task = new Task(ticket, target, nowNanos, creatorOwner);
String ticket = "task-" + ticketSeq.incrementAndGet();
Task task = new Task(ticket, target, nowNanos);
tasks.put(ticket, task);
if (pushLoop != null) {
// CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a
@@ -1359,7 +1311,7 @@ public final class MessageService {
}
asyncExecutor.submit(() -> {
try {
Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task, creatorOwner);
Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task);
if (result.outcome() == Outcome.QUESTION) {
// Keep the accepted owner until answer() finishes it. markAsyncQuestion may run
// just after resolveQuestion wakes this thread.
@@ -1391,31 +1343,15 @@ public final class MessageService {
}
/**
* As {@link #poll(String, String)}, with no caller owner key — the unnamed primary's ticket rule
* checked, so this overload must only be used where the caller's identity is otherwise
* irrelevant.
*/
public TaskView poll(String ticket) {
return poll(ticket, null);
}
/**
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket.
* Refuses a {@code callerOwner} that differs from the owner that created the ticket (see
* {@link #sendAsync(String, String, Runnable, Principal)}) with a {@link Phase#FAILED} view that
* carries no reply text. The unnamed primary has a {@code null} owner key and is never refused.
* Otherwise returns a {@link Phase#PENDING} view (with the live worker status as detail), a
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket;
* otherwise a {@link Phase#PENDING} view (with the live worker status as detail), a
* {@link Phase#DONE} view carrying the reply, or a {@link Phase#FAILED} view with the reason.
*/
public TaskView poll(String ticket, String callerOwner) {
public TaskView poll(String ticket) {
Task task = tasks.get(ticket);
if (task == null) {
return null;
}
if (!ownsTicket(task, callerOwner)) {
return new TaskView(ticket, Phase.FAILED, null, null,
"forbidden: this ticket was created by a different session", null);
}
CompletableFuture<Reply> f = task.future;
if (!f.isDone()) {
Reply question = task.question;
@@ -1451,16 +1387,6 @@ public final class MessageService {
return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
}
/**
* Whether {@code callerOwner} may read {@code task}'s state. A {@code null} caller key is the
* unnamed primary and may read every ticket. Other callers must match the task's owner key. This
* differs from {@link Rendezvous.Owner#permits}: a missing rendezvous owner is not an authenticated
* unnamed primary, so that gate refuses every caller when no owner was recorded.
*/
private static boolean ownsTicket(Task task, String callerOwner) {
return callerOwner == null || callerOwner.equals(task.creatorOwner);
}
/**
* Test seam only — carries no production behaviour, and nothing in this class calls it;
* {@link #pruneTerminalTickets} still reads {@link Task#completedNanos} directly.
@@ -1780,17 +1706,16 @@ public final class MessageService {
}
/**
* The question {@code workerSession} is currently paused on via {@code fleet_ask}, if any —
* {@code fleet_status} uses this to show a pending question without the caller needing the
* ticket. {@code null} when the session has no open async question (including a session mid a
* <em>blocking</em> {@code fleet_ask}, which has no {@link Task} to look up — see
* {@link PendingAsk}), or when {@code callerOwner} does not own the task the question
* belongs to (see {@link #ownsTicket(Task, String)}).
* The question {@code workerSession} is currently paused on via {@code fleet_ask}, if any
* (CB-582) — {@code fleet_status} uses this to show a pending question without the caller
* needing the ticket. {@code null} when the session has no open async question (including a
* session mid a <em>blocking</em> {@code fleet_ask}, which has no {@link Task} to look up — see
* {@link PendingAsk}).
*/
public PendingAsk pendingAsk(String workerSession, String callerOwner) {
public PendingAsk pendingAsk(String workerSession) {
for (Task task : tasks.values()) {
Reply q = task.question;
if (q != null && workerSession.equals(task.target) && ownsTicket(task, callerOwner)) {
if (q != null && workerSession.equals(task.target)) {
return new PendingAsk(task.ticket, q.text(), q.turnId());
}
}
@@ -1,6 +1,5 @@
package dev.ltms.fleet.msg;
import java.util.UUID;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicLong;
@@ -61,7 +60,7 @@ public final class Rendezvous {
}
/** A worker's open mid-turn question: the worker session it belongs to and the answer future. */
private record AskWaiter(String session, CompletableFuture<String> answer, Owner owner) {
private record AskWaiter(String session, CompletableFuture<String> answer) {
}
/**
@@ -71,44 +70,11 @@ public final class Rendezvous {
public record AskTicket(String turnId, CompletableFuture<String> answer, boolean fresh) {
}
/**
* The caller whose accepted delegation opened a turn — the only caller allowed to answer it.
* A {@code null} owner key means the unnamed primary.
*/
public record Owner(String ownerKey) {
public static final Owner UNNAMED_PRIMARY = new Owner(null);
public static Owner of(String ownerKey) {
return ownerKey == null ? UNNAMED_PRIMARY : new Owner(ownerKey);
}
/**
* Whether {@code callerOwner} matches {@code owner}. A {@code null} owner means no owner was
* recorded, so it matches no caller. {@link #UNNAMED_PRIMARY} records the unnamed primary
* with an owner object whose key is {@code null}.
*/
public static boolean permits(Owner owner, String callerOwner) {
return owner != null && java.util.Objects.equals(owner.ownerKey(), callerOwner);
}
}
/** A registered forward waiter together with the owner its delegation was opened under. */
private record ForwardWaiter(Owner owner, CompletableFuture<Resolution> future) {
}
private final ConcurrentHashMap<String, ForwardWaiter> waiters = new ConcurrentHashMap<>();
private final ConcurrentHashMap<String, CompletableFuture<Resolution>> waiters = new ConcurrentHashMap<>();
/** Reverse rendezvous (CB-205): worker questions awaiting the primary's answer, keyed by {@code turnId}. */
private final ConcurrentHashMap<String, AskWaiter> asks = new ConcurrentHashMap<>();
private final AtomicLong askSeq = new AtomicLong();
/**
* Minted once per {@code Rendezvous} instance and folded into every {@code turnId} (see
* {@link #openAsk(String)}). {@link #askSeq} alone restarts at zero for every instance, so
* without this a {@code turnId} minted by one instance could be minted again by another and
* resolve to an unrelated ask with no error; this nonce makes that impossible, because an id
* minted by one instance can never match the id space of another.
*/
private final String askBootNonce = UUID.randomUUID().toString().substring(0, 6);
/** Per-session index of the currently-open ask, so duplicate fleet_ask calls coalesce onto one turn. */
private final ConcurrentHashMap<String, String> openAsksBySession = new ConcurrentHashMap<>();
@@ -123,17 +89,8 @@ public final class Rendezvous {
* code a double open is impossible; this is a tripwire for the day that no longer holds.
*/
public CompletableFuture<Resolution> open(String session) {
return open(session, null);
}
/**
* Same as {@link #open(String)}, additionally recording {@code owner} as the caller whose
* delegation opened this waiter. A {@code null} owner records no owner at all — the
* fail-closed default {@link Owner#permits} refuses to everyone.
*/
public CompletableFuture<Resolution> open(String session, Owner owner) {
CompletableFuture<Resolution> waiter = new CompletableFuture<>();
ForwardWaiter existing = waiters.putIfAbsent(session, new ForwardWaiter(owner, waiter));
CompletableFuture<Resolution> existing = waiters.putIfAbsent(session, waiter);
if (existing != null) {
throw new IllegalStateException(
"rendezvous double-open for session " + session + " — a waiter is already registered");
@@ -148,13 +105,7 @@ public final class Rendezvous {
* successful {@code open} after a finished turn requires this close to have happened first).
*/
public void close(String session, CompletableFuture<Resolution> waiter) {
waiters.computeIfPresent(session, (s, w) -> w.future() == waiter ? null : w);
}
/** The owner recorded for {@code session}'s open waiter, or {@code null} if none is open. */
public Owner ownerOf(String session) {
ForwardWaiter w = waiters.get(session);
return w == null ? null : w.owner();
waiters.remove(session, waiter);
}
/** Whether a send is currently awaiting a resolution for {@code session}. */
@@ -168,8 +119,7 @@ public final class Rendezvous {
* send (see the CB-116 note above) rather than whichever send happens to be waiting when they fire.
*/
public CompletableFuture<Resolution> currentWaiter(String session) {
ForwardWaiter w = waiters.get(session);
return w == null ? null : w.future();
return waiters.get(session);
}
/**
@@ -195,9 +145,9 @@ public final class Rendezvous {
while (true) {
AskWaiter[] minted = { null };
String turnId = openAsksBySession.computeIfAbsent(session, _ -> {
String newTurnId = session + "#" + askBootNonce + "-" + askSeq.incrementAndGet();
String newTurnId = session + "#" + askSeq.incrementAndGet();
CompletableFuture<String> answer = new CompletableFuture<>();
AskWaiter waiter = new AskWaiter(session, answer, ownerOf(session));
AskWaiter waiter = new AskWaiter(session, answer);
asks.put(newTurnId, waiter);
minted[0] = waiter;
return newTurnId;
@@ -233,16 +183,6 @@ public final class Rendezvous {
return w == null ? null : w.session();
}
/**
* The owner recorded for {@code turnId} when its ask turn was freshly opened — the caller
* whose delegation {@link #answerAsk} must match. {@code null} if {@code turnId} is unknown or
* lapsed, or if the ask opened with no forward waiter owner on record.
*/
public Owner askOwner(String turnId) {
AskWaiter w = asks.get(turnId);
return w == null ? null : w.owner();
}
/**
* Resolve a worker's blocked {@code fleet_ask} with the primary's {@code answer}, unblocking it
* to resume its turn.
@@ -305,7 +245,7 @@ public final class Rendezvous {
}
private boolean complete(String session, Resolution resolution) {
ForwardWaiter waiter = waiters.get(session);
return waiter != null && waiter.future().complete(resolution);
CompletableFuture<Resolution> waiter = waiters.get(session);
return waiter != null && waiter.complete(resolution);
}
}
@@ -49,52 +49,22 @@ import java.util.stream.Collectors;
*/
public final class FleetApp {
/**
* The authorization action the matching route handler hands to {@link #allow}, for a route
* whose action does not depend on the request body.
*/
/** The authorization action the matching route handler hands to {@link #allow}. */
static Authz.Action routeAction(String route) {
return routeAction(route, null);
}
/**
* As above, plus the one route whose action depends on the body: {@code POST
* /sessions/{id}/message} carries a {@code turnId} (the answer-a-blocked-worker shape) or not
* (a plain delivery), mirroring {@code FleetMcp#sendAction}'s split of the same two call
* shapes over MCP. {@code turnId} is ignored by every other route.
*
* @param turnId the request body's {@code turnId}, or {@code null}/blank when absent or not
* applicable to this route
*/
static Authz.Action routeAction(String route, String turnId) {
return switch (route) {
case "GET /metrics" -> Authz.Action.METRICS;
case "POST /members" -> Authz.Action.SPAWN;
case "DELETE /members/{paneId}" -> Authz.Action.STOP;
case "POST /sessions/{id}/message" -> turnId == null || turnId.isBlank()
? Authz.Action.SEND : Authz.Action.ANSWER;
case "POST /sessions/{id}/message" -> Authz.Action.SEND;
case "POST /sessions/{id}/reply" -> Authz.Action.REPLY;
case "GET /sessions/{id}/replies" -> Authz.Action.DRAIN;
case "POST /sessions/{id}/ask" -> Authz.Action.ASK;
case "GET /sessions", "GET /agents", "GET /members", "GET /profiles",
"GET /member-credentials" -> Authz.Action.READ;
case "GET /sessions/{id}/status", "GET /tasks/{ticket}" -> Authz.Action.TASK_READ;
"GET /member-credentials", "GET /sessions/{id}/status", "GET /tasks/{ticket}" -> Authz.Action.READ;
default -> throw new IllegalArgumentException("route has no authorization gate: " + route);
};
}
/**
* The second gate for {@code POST /sessions/{id}/message}: checked only when {@code turnId}
* is present and non-blank, against {@link Authz.Action#ANSWER}. A request with no {@code
* turnId} passes this gate unconditionally, without consulting {@code permit} at all, having
* already cleared the coarse {@link Authz.Action#SEND} grant checked ahead of it.
*
* @param permit reports whether the caller holds the named grant
*/
static boolean answerGatePasses(String turnId, Predicate<Authz.Action> permit) {
return turnId == null || turnId.isBlank() || permit.test(Authz.Action.ANSWER);
}
/** Default blocking window for a message; kept under typical HTTP idle timeouts. */
private static final long DEFAULT_MESSAGE_TIMEOUT_MS = 25_000;
private static final long MAX_MESSAGE_TIMEOUT_MS = 120_000;
@@ -266,21 +236,6 @@ public final class FleetApp {
return app;
}
/**
* The authorization decision behind {@link #allow}, taking the caller directly rather than
* pulling it from a servlet {@link Context} — unit-testable without fabricating a live
* request, the same reason {@code FleetMcp#denyFor} is split from {@code FleetMcp#deny}.
*
* @param knownLeadOrCollaborator the classifier a collaborator's {@code SEND} is checked
* against; pass {@link #auth}'s own {@code
* knownLeadOrCollaborator()} to exercise the real production
* gate, as {@link #allow} does
*/
static boolean permitsFor(Principal caller, Authz.Action action, String target,
Predicate<String> knownLeadOrCollaborator) {
return Authz.permits(caller, action, target, knownLeadOrCollaborator);
}
/**
* Gate a handler on the CB-505 authorization table. Returns {@code true} when the request may
* proceed; otherwise writes the error response and returns {@code false}.
@@ -294,9 +249,8 @@ public final class FleetApp {
return true; // legacy: authorization not enforced
}
Principal caller = ctx.attribute(CALLER);
if (permitsFor(caller, action, target, auth.knownLeadOrCollaborator())) {
if (action != Authz.Action.READ && action != Authz.Action.METRICS
&& action != Authz.Action.TASK_READ) {
if (Authz.permits(caller, action, target)) {
if (action != Authz.Action.READ && action != Authz.Action.METRICS) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
}
return true;
@@ -649,39 +603,26 @@ public final class FleetApp {
* status-gated injector and block until the worker returns a structured {@code fleet_reply}.
* Times out with a typed 202 (working / queued / busy) rather than an error — the message may
* still land.
*
* <p>Two call shapes share this route, exactly as {@code fleet_send} does over MCP (see
* {@code FleetMcp#sendAction}): a plain delivery to {@code id}, and -- when the body carries
* {@code turnId} -- resolving a worker's blocked question. The coarse {@link
* Authz.Action#SEND} grant is checked first, before the body is read at all; only once that
* passes is the body parsed, and a present {@code turnId} is then checked again against
* {@link Authz.Action#ANSWER}. A body that fails to parse is rejected with 400 and reaches
* neither {@code messages.answer} nor {@code messages.send}.
*/
private void sendMessage(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, routeAction("POST /sessions/{id}/message"), id)) {
return;
}
Principal caller = ctx.attribute(CALLER);
String callerOwner = caller == null ? null : caller.ownerKey();
JsonNode body;
String content;
String turnId;
long timeout;
boolean wait;
try {
body = mapper.readTree(ctx.body());
JsonNode body = mapper.readTree(ctx.body());
content = body.path("content").asText("");
turnId = body.path("turnId").asText(null);
timeout = body.path("timeoutMs").asLong(DEFAULT_MESSAGE_TIMEOUT_MS);
wait = body.path("wait").asBoolean(true); // default: block for the reply (CB-104)
} catch (Exception e) {
body = null;
}
if (body == null) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
String turnId = body.path("turnId").asText(null);
if (!answerGatePasses(turnId, action -> allow(ctx, action, id))) {
return;
}
String content = body.path("content").asText("");
long timeout = body.path("timeoutMs").asLong(DEFAULT_MESSAGE_TIMEOUT_MS);
boolean wait = body.path("wait").asBoolean(true); // default: block for the reply (CB-104)
if (content.isBlank()) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "content is required"));
return;
@@ -690,19 +631,19 @@ public final class FleetApp {
// Answering a worker's fleet_ask (CB-205): always blocks, and derives the worker from turnId.
if (turnId != null && !turnId.isBlank()) {
writeReply(ctx, id, messages.answer(turnId, content, timeout, callerOwner), timeout);
writeReply(ctx, id, messages.answer(turnId, content, timeout), timeout);
return;
}
if (!wait) {
// Fire-and-poll (CB-107): return a ticket immediately; the caller polls GET /tasks/{ticket}.
String ticket = messages.sendAsync(id, content, null, caller);
String ticket = messages.sendAsync(id, content);
ctx.status(202).json(Map.of("sessionId", id, "ticket", ticket, "status", "accepted"));
return;
}
try {
writeReply(ctx, id, messages.send(id, content, timeout, callerOwner), timeout);
writeReply(ctx, id, messages.send(id, content, timeout), timeout);
} catch (HerdrException e) {
herdrError(ctx, e);
}
@@ -721,10 +662,6 @@ public final class FleetApp {
case STALE_TURN -> ctx.status(409).json(Map.of(
"sessionId", id, "error", "stale_turn",
"detail", "that question is no longer open (timed out or already answered)"));
case NOT_TURN_OWNER -> ctx.status(403).json(Map.of(
"sessionId", id, "error", "not_turn_owner",
"detail", "this turn belongs to a different delegation — only the caller that "
+ "opened it may answer it"));
case REPLIED, COMPLETED_UNREPLIED -> {
// replySource distinguishes a structured fleet_reply from the CB-106 completion
// fallback (a scrape of the worker's transcript when it finished without replying).
@@ -739,10 +676,9 @@ public final class FleetApp {
// silent fall-through. That is exactly the bug this ticket exists to fix:
// `default -> "done"` used to sit here and would have told a REST caller the
// delegation completed for TIMED_OUT_UNCONFIRMED, the one outcome where delivery
// is unknown. REPLIED, COMPLETED_UNREPLIED, QUESTION, STALE_TURN and
// NOT_TURN_OWNER can never actually reach this inner switch — the outer switch
// above always dispatches them first — but they still need an arm to keep this
// switch exhaustive.
// is unknown. REPLIED, COMPLETED_UNREPLIED, QUESTION and STALE_TURN can never
// actually reach this inner switch — the outer switch above always dispatches
// them first — but they still need an arm to keep this switch exhaustive.
"status", switch (reply.outcome()) {
case TIMED_OUT_WORKING -> "working";
case TIMED_OUT_QUEUED -> "queued";
@@ -752,7 +688,7 @@ public final class FleetApp {
case BUSY -> "busy";
case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted";
case REPLIED, COMPLETED_UNREPLIED, QUESTION, STALE_TURN, NOT_TURN_OWNER -> "done"; // unreachable
case REPLIED, COMPLETED_UNREPLIED, QUESTION, STALE_TURN -> "done"; // unreachable
},
"detail", reply.outcome() == MessageService.Outcome.TIMED_OUT_UNCONFIRMED
? "no reply within " + timeout + "ms; delivery is unconfirmed — the "
@@ -881,12 +817,10 @@ public final class FleetApp {
body.put("sessionId", id);
body.put("status", messages.status(id).name().toLowerCase());
body.put("ready", deliverable.test(id));
// A worker paused mid-turn in an async fleet_ask is otherwise invisible to a status
// poll — surface the open question and how to answer it, same as fleet_poll's
// Phase.ASKING view, but only to the caller whose owner key created that delegation, or
// to the unnamed primary.
Principal caller = ctx.attribute(CALLER);
MessageService.PendingAsk ask = messages.pendingAsk(id, caller == null ? null : caller.ownerKey());
// CB-582: a worker paused mid-turn in an async fleet_ask is otherwise invisible to a
// status poll — surface the open question and how to answer it, same as fleet_poll's
// Phase.ASKING view.
MessageService.PendingAsk ask = messages.pendingAsk(id);
if (ask != null) {
body.put("question", ask.question());
body.put("turnId", ask.turnId());
@@ -903,8 +837,7 @@ public final class FleetApp {
if (!allow(ctx, routeAction("GET /tasks/{ticket}"), null)) {
return;
}
Principal caller = ctx.attribute(CALLER);
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"), caller == null ? null : caller.ownerKey());
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"));
if (v == null) {
ctx.status(404).json(Map.of("error", "unknown_ticket", "detail", "no such task (or it has expired)"));
return;
@@ -26,7 +26,6 @@ import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Consumer;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* Authoritative in-daemon registry of the worker sessions this {@code fleetd} process spawned.
@@ -58,18 +57,6 @@ public final class SessionManager implements TurnListener {
* Populated on every spawn path, removed on {@link #release}.
*/
private final ConcurrentHashMap<String /*paneId*/, PeerHandle> handles = new ConcurrentHashMap<>();
/**
* fleetd #702: a pane mid-teardown, keyed by paneId, held from just before its registry entry
* is removed until {@link #releaseRemoved} finishes. {@link #spawnedMemberRole} consults this
* alongside the registry, so a caller resolving the pane's terminal during that window still
* sees a live member and never falls through to a tab map.
*
* <p>Depth-counted rather than a plain set: two threads can be tearing down the same pane at
* once (the CAS in {@link #releaseIfCurrent} exists for exactly that race), and with a set the
* loser's {@code finally} would unmark the pane while the winner is still mid-teardown,
* reopening the window this exists to close.
*/
private final ConcurrentHashMap<String /*paneId*/, Releasing> releasing = new ConcurrentHashMap<>();
private final MemberPresence presence;
private final SecureRandom nonceRandom = new SecureRandom();
private final AtomicLong nonceSeq = new AtomicLong();
@@ -255,10 +242,6 @@ public final class SessionManager implements TurnListener {
handle.id(), handle.terminalId(), resolvedProfile, actualRole, cwd, ownerTerminal, now, now, 0,
MemberSession.State.SPAWNING, null, null, handle.charterReceipt(), handle.agentSessionId());
registry.put(handle.id(), session);
// A presence contact that already arrived for this terminal found no registry
// entry to transition and gave up silently. Retry it now that one exists; remove
// this call and such a session stays in SPAWNING even though it is present.
reconcilePresence(handle.terminalId());
handles.put(handle.id(), handle);
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
@@ -320,11 +303,9 @@ public final class SessionManager implements TurnListener {
* is a logged path an operator can reclaim, the cost of a deleted one is unrecoverable work.
*/
private MemberSession release(String paneId, ReleaseCause cause) {
return releaseWindow(paneId, registry.get(paneId), () -> {
MemberSession removed = registry.remove(paneId);
releaseRemoved(paneId, removed, handles.remove(paneId), cause);
return removed;
});
MemberSession removed = registry.remove(paneId);
releaseRemoved(paneId, removed, handles.remove(paneId), cause);
return removed;
}
/**
@@ -333,85 +314,16 @@ public final class SessionManager implements TurnListener {
* DONE record from stopping a worker that delivery has made BUSY.
*/
private boolean releaseIfCurrent(MemberSession expected, ReleaseCause cause) {
return releaseWindow(expected.paneId(), expected, () -> {
if (!registry.remove(expected.paneId(), expected)) {
// A lifecycle transition replaced the record between the caller's check and this
// remove. Log it: this race is by definition unobservable otherwise, and a reaper
// that silently declines to reap is the hardest kind of behaviour to diagnose
// after the fact.
log.debug("skipping reap of pane={}: its registry record changed after the idle "
+ "check (most likely a delivery made it BUSY)", expected.paneId());
return false;
}
releaseRemoved(expected.paneId(), expected, handles.remove(expected.paneId()), cause);
return true;
});
}
/**
* fleetd #702: mark {@code paneId} as mid-teardown — using {@code known}'s terminal/role when
* it is available — for the whole of {@code teardown}, which removes the registry entry and
* then runs {@link #releaseRemoved}. Shared by both registry-removal sites ({@link #release}'s
* unconditional remove and {@link #releaseIfCurrent}'s CAS remove) so neither can leave the
* other's window unmarked.
*
* <p>The mark is written before {@code teardown} runs — so it covers the removal itself, not
* only what comes after it — and cleared in a {@code finally}, so an unchecked throw out of
* {@code teardown} (including one from {@link PeerLauncher#stop}, which declares nothing) can
* never leave the pane marked for the rest of the daemon's life.
*/
private <T> T releaseWindow(String paneId, MemberSession known, Supplier<T> teardown) {
releasing.compute(paneId, (_, prior) -> Releasing.enter(prior, known));
try {
return teardown.get();
} finally {
releasing.compute(paneId, (_, prior) -> prior == null ? null : prior.leave());
if (!registry.remove(expected.paneId(), expected)) {
// A lifecycle transition replaced the record between the caller's check and this remove.
// Log it: this race is by definition unobservable otherwise, and a reaper that silently
// declines to reap is the hardest kind of behaviour to diagnose after the fact.
log.debug("skipping reap of pane={}: its registry record changed after the idle check "
+ "(most likely a delivery made it BUSY)", expected.paneId());
return false;
}
}
/**
* Depth count plus the terminal/role a mid-teardown pane belongs to, for
* {@link #spawnedMemberRole}. The terminal/role come from whichever call into
* {@link #releaseWindow} first knew them: a call that finds the registry entry already gone
* passes a {@code null} session, and must not blank out what the first call recorded.
*/
record Releasing(int depth, String terminalId, MemberRole role) {
static Releasing enter(Releasing prior, MemberSession known) {
int depth = (prior == null ? 0 : prior.depth()) + 1;
String terminalId = known != null ? known.terminalId() : prior == null ? null : prior.terminalId();
MemberRole role = known != null ? known.role() : prior == null ? null : prior.role();
return new Releasing(depth, terminalId, role);
}
Releasing leave() {
return depth <= 1 ? null : new Releasing(depth - 1, terminalId, role);
}
}
/**
* The role of the live spawned member occupying {@code terminal} — whether it is currently in
* the registry, or mid-teardown between {@link #release} removing its registry entry and
* {@link #releaseRemoved} actually stopping its pane (fleetd #702). {@code null} for a terminal
* that is neither: this method is the one reader a caller resolver consults before any tab
* map, so a live or releasing member's identity never falls back to a tab label.
*
* <p>Checks the registry directly via {@link #findByTerminal} rather than {@link #roster()},
* so this hot-path lookup (consulted on every resolve) never pays for a list copy or a stream.
*/
public MemberRole spawnedMemberRole(String terminal) {
MemberSession session = findByTerminal(terminal);
if (session != null) {
return session.role();
}
if (terminal == null) {
return null;
}
for (Releasing r : releasing.values()) {
if (terminal.equals(r.terminalId())) {
return r.role();
}
}
return null;
releaseRemoved(expected.paneId(), expected, handles.remove(expected.paneId()), cause);
return true;
}
private void releaseRemoved(String paneId, MemberSession removed, PeerHandle removedHandle,
@@ -470,12 +382,6 @@ public final class SessionManager implements TurnListener {
MemberSession resolved = resolveAgentSessionId(removed, removedHandle);
notifyReleased(new ReleaseDetail(resolved.terminalId(), resolved.worktree(),
resolved.branch(), snapshotRef, resolved.agentSessionId()));
String terminal = removed.terminalId();
if (terminal != null && !terminal.isBlank()) {
// Without this, a terminal stays marked present after its pane is gone, so a
// later send to the same id would read as deliverable instead of refused.
presence.forget(terminal);
}
}
}
// CB-581: the pane must always stop, even if the dirty check above threw. A session removed
@@ -819,10 +725,6 @@ public final class SessionManager implements TurnListener {
handle.charterReceipt(),
handle.agentSessionId());
registry.put(handle.id(), session);
// A presence contact that already arrived for this terminal found no registry entry to
// transition and gave up silently. Retry it now that one exists; remove this call and
// such a session stays in SPAWNING even though it is present.
reconcilePresence(handle.terminalId());
handles.put(handle.id(), handle);
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
@@ -997,20 +899,6 @@ public final class SessionManager implements TurnListener {
transitionByTerminal(terminalId, MemberSession.State.SPAWNING, MemberSession.State.READY);
}
/**
* Completes a newly registered session's {@code SPAWNING -> READY} transition when {@code
* terminalId} was already marked present before this ran. A terminal never marked present is
* left in {@code SPAWNING}; it reaches {@code READY} normally through {@link #onReady} once
* its own contact arrives. Callers must run this only once the session's registry entry is
* already visible — {@link #onReady}'s transition matches against that entry, and reconciling
* before the entry exists finds nothing to transition.
*/
private void reconcilePresence(String terminalId) {
if (terminalId != null && !terminalId.isBlank() && presence.isPresent(terminalId)) {
onReady(terminalId);
}
}
/**
* Lifecycle hook: a message was delivered into the worker — it is now busy on a turn.
* The turn count is bumped and the activity timestamp is refreshed. A {@code DONE} session
@@ -14,12 +14,10 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-534: the injector's readiness gate must open for a lead as well as for a present worker.
* fleetd #669 follow-up: the same gate must also open for a collaborator, which — like a lead —
* is never enrolled in {@link MemberPresence} and never discovered by the lead scan.
*
* <p>The bug these cover was silent and slow: a lead (and later a collaborator) was never marked
* present (only workers are) and never counted as a lead, so every send to one sat on the gate for
* the full readiness grace and failed ~60s later without a keystroke ever reaching the pane.
* <p>The bug these cover was silent and slow: a lead was never marked present (only workers are), so
* every lead→lead delivery sat on the gate for the full readiness grace and failed ~60s later without
* a keystroke ever reaching the pane.
*/
class FleetDeliverabilityTest {
@@ -27,25 +25,19 @@ class FleetDeliverabilityTest {
return () -> m;
}
private static Supplier<Map<String, String>> collaborators(Map<String, String> m) {
return () -> m;
}
@Test
@DisplayName("a worker that has connected its MCP is deliverable")
void presentWorkerIsDeliverable() {
MemberPresence presence = new MemberPresence();
presence.markPresent("term_worker");
assertTrue(Fleetd.deliverableTo(presence, leads(Map.of()), collaborators(Map.of()))
.test("term_worker"));
assertTrue(Fleetd.deliverableTo(presence, leads(Map.of())).test("term_worker"));
}
@Test
@DisplayName("a worker still in its boot window is held back")
void absentWorkerIsNotDeliverable() {
assertFalse(Fleetd.deliverableTo(new MemberPresence(), leads(Map.of()), collaborators(Map.of()))
.test("term_booting"));
assertFalse(Fleetd.deliverableTo(new MemberPresence(), leads(Map.of())).test("term_booting"));
}
@Test
@@ -53,31 +45,19 @@ class FleetDeliverabilityTest {
void leadIsDeliverableWithoutPresence() {
MemberPresence presence = new MemberPresence();
Predicate<String> deliverable =
Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")), collaborators(Map.of()));
Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")));
assertFalse(presence.isPresent("term_lead"), "a lead is never enrolled in worker presence");
assertTrue(deliverable.test("term_lead"), "…and must be deliverable anyway");
}
@Test
@DisplayName("a collaborator is deliverable without ever being marked present or scanned as a lead")
void collaboratorIsDeliverableWithoutPresenceOrLeadStatus() {
MemberPresence presence = new MemberPresence();
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads(Map.of()),
collaborators(Map.of("term_collab", "kevin")));
assertFalse(presence.isPresent("term_collab"), "a collaborator is never enrolled in worker presence");
assertTrue(deliverable.test("term_collab"), "…and must be deliverable anyway");
}
@Test
@DisplayName("an unknown terminal is deliverable to none of presence, leads, or collaborators")
@DisplayName("an unknown terminal is deliverable to neither")
void strangerIsNotDeliverable() {
MemberPresence presence = new MemberPresence();
presence.markPresent("term_worker");
assertFalse(Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")),
collaborators(Map.of("term_collab", "kevin")))
assertFalse(Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")))
.test("term_stranger"));
}
@@ -85,47 +65,22 @@ class FleetDeliverabilityTest {
@DisplayName("a lead discovered after startup becomes deliverable with no restart")
void leadSetIsReadThroughOnEveryCall() {
Map<String, String> discovered = new HashMap<>();
Predicate<String> deliverable =
Fleetd.deliverableTo(new MemberPresence(), leads(discovered), collaborators(Map.of()));
Predicate<String> deliverable = Fleetd.deliverableTo(new MemberPresence(), leads(discovered));
assertFalse(deliverable.test("term_late"));
discovered.put("term_late", "gpt-sol-5.6"); // leadScan picks up a newly labelled tab
assertTrue(deliverable.test("term_late"), "the supplier must be re-read, not snapshotted");
}
@Test
@DisplayName("a collaborator discovered after startup becomes deliverable with no restart")
void collaboratorSetIsReadThroughOnEveryCall() {
Map<String, String> discovered = new HashMap<>();
Predicate<String> deliverable =
Fleetd.deliverableTo(new MemberPresence(), leads(Map.of()), collaborators(discovered));
assertFalse(deliverable.test("term_late_collab"));
discovered.put("term_late_collab", "kevin"); // the same tab scan picks up a newly labelled collaborator tab
assertTrue(deliverable.test("term_late_collab"), "the supplier must be re-read, not snapshotted");
}
@Test
@DisplayName("forgetting a torn-down worker does not strip a lead of its deliverability")
void forgetDoesNotDisarmALead() {
MemberPresence presence = new MemberPresence();
Predicate<String> deliverable =
Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")), collaborators(Map.of()));
Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")));
presence.forget("term_lead"); // the injector's cleanup path runs against every target
assertTrue(deliverable.test("term_lead"));
}
@Test
@DisplayName("forgetting a torn-down worker does not strip a collaborator of its deliverability")
void forgetDoesNotDisarmACollaborator() {
MemberPresence presence = new MemberPresence();
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads(Map.of()),
collaborators(Map.of("term_collab", "kevin")));
presence.forget("term_collab"); // the injector's cleanup path runs against every target
assertTrue(deliverable.test("term_collab"));
}
}
@@ -1,166 +0,0 @@
package dev.ltms.fleet;
import dev.ltms.fleet.auth.Authz;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Method;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Asserts that the {@link FleetMcp} built by {@link FleetdAssembly#assembleAndStart} applies the
* authorization table: a worker is refused {@code SPAWN}, and the primary is allowed it.
*/
class FleetdAssemblyAuthorizationModeTest {
private static final class TestResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new ReplyInbox() {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) { }
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public boolean ack(String target, String msgId) { return false; }
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Binding a real port would clash with any daemon already listening on it.
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
};
}
}
private TestResourcePorts ports;
@AfterEach
void tearDown() {
if (ports != null && ports.shutdownHook != null) {
ports.shutdownHook.run();
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
health:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
""");
return FleetConfig.load(file);
}
private FleetMcp assemble(Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ports = new TestResourcePorts();
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
return runtime.mcp();
}
/**
* Invokes {@code FleetMcp#denyFor}, which is package-private to {@code dev.ltms.fleet.mcp}
* while this test is in {@code dev.ltms.fleet}. Nothing here catches a missing method: if
* {@code denyFor} is renamed or removed, {@link NoSuchMethodException} propagates and the
* test fails.
*/
private static McpSchema.CallToolResult denyFor(FleetMcp mcp, Principal caller, Authz.Action action,
String target) throws Exception {
Method m = FleetMcp.class.getDeclaredMethod("denyFor", Principal.class, Authz.Action.class, String.class);
m.setAccessible(true);
return (McpSchema.CallToolResult) m.invoke(mcp, caller, action, target);
}
@Test
void productionBootPathRefusesAnUnauthorizedCallerThroughTheAssembledFleetMcp(@TempDir Path dir)
throws Exception {
FleetMcp mcp = assemble(dir);
McpSchema.CallToolResult deniedForWorker = denyFor(mcp, Principal.worker("term_a", 200),
Authz.Action.SPAWN, "term_a");
assertNotNull(deniedForWorker,
"a worker must not be able to fleet_spawn through the assembled FleetMcp");
assertTrue(deniedForWorker.isError(), "a refusal is returned as an MCP tool error");
McpSchema.CallToolResult allowedForPrimary = denyFor(mcp, Principal.primary(100),
Authz.Action.SPAWN, "term_a");
assertNull(allowedForPrimary,
"control: the primary must still be allowed to fleet_spawn — otherwise the worker "
+ "refusal above would pass even with the gate wired backwards");
}
}
@@ -1,165 +0,0 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertSame;
/**
* fleetd #669 Unit E. Reaches the real {@link dev.ltms.fleet.herdr.HerdrRouter} that {@link
* FleetdAssembly#assembleAndStart} builds and wires — not a copy built for this test — and proves
* that a configured collaborator's terminal routes to the LEAD herdr daemon.
*
* <p>Two distinct {@link FakeHerdr} instances are required, the same pattern {@code
* FleetdAssemblyConnectionIdentityTest} and {@code FleetdLeadRolloverAssemblyTest} already use:
* with one client shared between {@code herdrSocket} and {@code memberHerdrSocket},
* {@code HerdrRouter} folds {@code leadAgents} and {@code memberAgents} into the same instance
* (see its constructor), and {@code agentsFor} would return that one object regardless of whether
* the collaborator map was ever consulted — invisible to a mutation of the predicate this ticket
* fixes. This test's two sockets resolve to two different fakes, so the assertion only passes when
* the collaborator's terminal is actually recognised and routed to the lead one.
*/
class FleetdAssemblyCollaboratorHerdrRoutingTest {
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
private static final class RecordingResourcePorts implements ResourcePorts {
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
HerdrClient client = herdrsBySocket.get(socketPath);
if (client == null) {
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
}
return client;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new ReplyInbox() {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) { }
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public boolean ack(String target, String msgId) { return false; }
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Do not bind a real port in this assembly test.
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
};
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: "%s"
memberHerdrSocket: "%s"
idleSleepGuard:
enabled: false
fleet:
collaborators:
reviewer-alex:
tab: "collab: alex"
profiles:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
""".formatted(LEAD_SOCKET, MEMBER_SOCKET));
return FleetConfig.load(file);
}
@Test
void assembledRouterRoutesACollaboratorTerminalToTheLeadDaemon(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
RecordingResourcePorts ports = new RecordingResourcePorts();
// The fixed FakeHerdr fixture already ties terminal "term_a" to a live agent on tab
// "w2:t7" (pane "w2:p7") — seeding only the tab LABEL to match the configured collaborator
// is enough to make LeadTabScanner resolve "term_a" as that collaborator. Seeded on the
// LEAD fake only: a collaborator's pane lives in the lead daemon, exactly like a lead's.
FakeHerdr lead = new FakeHerdr().withTab("w2", "w2:t7", "collab: alex");
FakeHerdr member = new FakeHerdr();
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
try {
assertSame(runtime.router().leadAgents(), runtime.router().agentsFor("term_a"),
"a configured collaborator's terminal must route to the LEAD daemon — "
+ "FleetdAssembly must wire the collaborator map into the router's "
+ "predicate, not just LeadTabScanner.get()");
assertSame(runtime.router().memberAgents(), runtime.router().agentsFor("term_shell"),
"control: a terminal naming neither a lead nor a collaborator (term_shell, on "
+ "the unlabelled tab w2:t8) must still route to the member daemon");
} finally {
assertNotNull(ports.shutdownHook, "control: assembly must capture its shutdown hook");
ports.shutdownHook.run();
}
}
}
@@ -58,31 +58,20 @@ import static org.junit.jupiter.api.Assertions.assertEquals;
* {@code HttpClient} — no accessor needed for this half.
*
* <p><strong>{@code GET /sessions} could not be driven the same way</strong>, so this class does
* not pin the merge half of the deleted test's javadoc. This class configures no {@code auth:}
* block, so it runs under the default {@code loopback-trust} mode ({@code FleetConfig}). Under
* that mode, {@code /sessions} requires {@code Authz.Action.READ}, which — through the REAL
* assembly's real {@code CallerResolver}/{@code ConnectionIdentity} (built with a hardcoded
* {@code new LsofPeerPidLookup()}) — needs {@code Caller.resolved()}, i.e. a real positive pid
* from {@code lsof}. {@code LsofPeerPidLookup} excludes its own pid (see its javadoc), and a
* JUnit test's HTTP client and the daemon under test share one JVM pid, so the resolved pid is
* always {@code -1} and every such request is refused as {@code ANONYMOUS} (fleetd #317's
* fail-closed rule) before the route handler — and its {@code memberHerdr} merge — is ever
* reached. Verified directly: driving {@code GET /sessions} here returns {@code 401
* unauthenticated}, not the merged body. {@code FleetAppTwoDaemonTest} avoids this because it
* builds {@code FleetApp} with {@code callers: null}, which is not what the real assembly
* passes. The {@code /healthz} pin below is what this class relies on for CB-185's {@code
* FleetApp} half; {@code FleetAppTwoDaemonTest} remains the full behavioural proof that
* {@code FleetApp} itself merges {@code /sessions} correctly once handed two clients.
*
* <p><strong>This refusal is {@code loopback-trust}-specific, not a property of {@code
* CallerResolver} in general.</strong> Under {@code auth.mode: token}, {@code
* CallerResolver#resolve} returns before ever consulting {@code Caller.resolved()} or {@code
* Caller.scanComplete()}: a request carrying a valid bearer token in its {@code Authorization}
* header resolves to {@code Role#PRIMARY} with no pid lookup at all, so the same-JVM-pid
* exclusion above never comes into play. {@code FleetdQuarantineOutageDualWindowAssemblyTest}
* and {@code FleetdListReportingSourcesAssemblyTest} both drive {@code Authz.Action.READ} this
* way, over a real {@code McpSyncClient}/{@code HttpClient} against a real {@code
* FleetdAssembly#assembleAndStart}, and both get the real response rather than a refusal.
* not pin the merge half of the deleted test's javadoc. {@code /sessions} requires
* {@code Authz.Action.READ}, which — through the REAL assembly's real {@code
* CallerResolver}/{@code ConnectionIdentity} (built with a hardcoded {@code
* new LsofPeerPidLookup()}) — needs {@code Caller.resolved()}, i.e. a real positive pid from
* {@code lsof}. {@code LsofPeerPidLookup} excludes its own pid (see its javadoc), and a JUnit
* test's HTTP client and the daemon under test share one JVM pid, so the resolved pid is always
* {@code -1} and every such request is refused as {@code ANONYMOUS} (fleetd #317's fail-closed
* rule) before the route handler — and its {@code memberHerdr} merge — is ever reached. Verified
* directly: driving {@code GET /sessions} here returns {@code 401 unauthenticated}, not the
* merged body. {@code FleetAppTwoDaemonTest} avoids this because it builds {@code FleetApp} with
* {@code callers: null}, which is not what the real assembly passes. The {@code /healthz} pin
* below is what this class relies on for CB-185's {@code FleetApp} half; {@code
* FleetAppTwoDaemonTest} remains the full behavioural proof that {@code FleetApp} itself merges
* {@code /sessions} correctly once handed two clients.
*
* <p><strong>fleetd #629 follow-up.</strong> The fix below (see {@link TwoHerdrResourcePorts})
* makes {@link #healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp}'s fake {@code
@@ -1,200 +0,0 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.LeadTabScanner;
import dev.ltms.fleet.msg.LeadChannelHandle;
import dev.ltms.fleet.msg.LeadCoordLoop;
import dev.ltms.fleet.msg.LeadMessage;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Field;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertInstanceOf;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #670 — pins the {@code excludedWorkspaceLabels} argument {@link FleetdAssembly}'s
* production boot path passes to {@link LeadTabScanner} ({@code Set.of()}).
*
* <p>{@code LeadTabScannerTest} already covers this constructor parameter, but it builds its own
* {@link LeadTabScanner} with its own set, so it tests the seam and proves nothing about the
* producer. This test instead reaches the exact object {@link FleetdAssembly#assembleAndStart}
* builds: a {@code fleet.leaders:} block makes the assembly construct a real
* {@link LeadTabScanner} for its local {@code leads} supplier, and a {@code coordinator:} block
* makes it hand that same supplier instance to {@link LeadCoordLoop} (fleetd #637), which stores
* it as a field. Reflection recovers it from there, and then from the scanner itself, so the
* assertion is against the real production argument rather than a copy built for this test.
*/
class FleetdAssemblyLeadTabScannerExclusionTest {
private static final class FakeLeadChannel implements LeadChannelHandle {
@Override
public void publish(String toCoordId, LeadMessage message) {
}
@Override
public List<LeadMessage> peek() {
return List.of();
}
@Override
public void ack(String msgId) {
}
@Override
public String selfCoordId() {
return "test-lead";
}
@Override
public boolean heldDurable() {
return true;
}
@Override
public MailboxState inspect(String coordId) {
return MailboxState.unknown(coordId);
}
@Override
public void close() {
}
}
private static final class TestResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new ReplyInbox() {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) { }
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public boolean ack(String target, String msgId) { return false; }
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> new FakeLeadChannel();
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Do not bind a real port in this assembly test.
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
};
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
coordinator:
uri: "amqp://fake-lead-broker/vh"
selfId: "test-lead"
fleet:
leaders:
primary:
tab: "lead: primary"
profile: sonnet
profiles:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
""");
return FleetConfig.load(file);
}
@Test
void productionBootPathPassesNoExcludedWorkspaceLabels(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
TestResourcePorts ports = new TestResourcePorts();
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
try {
LeadCoordLoop coordLoop = runtime.leadCoordLoop();
assertNotNull(coordLoop, "control: a configured coordinator: block must build LeadCoordLoop");
Field leadsField = LeadCoordLoop.class.getDeclaredField("leads");
leadsField.setAccessible(true);
@SuppressWarnings("unchecked")
Supplier<Map<String, String>> leads = (Supplier<Map<String, String>>) leadsField.get(coordLoop);
assertInstanceOf(LeadTabScanner.class, leads,
"control: a non-empty fleet.leaders: block must make FleetdAssembly build a real "
+ "LeadTabScanner for its `leads` supplier, not the Map::of fallback — "
+ "otherwise this test would pass for the wrong reason");
Field excludedField = LeadTabScanner.class.getDeclaredField("excludedWorkspaceLabels");
excludedField.setAccessible(true);
Set<?> excluded = (Set<?>) excludedField.get(leads);
assertTrue(excluded.isEmpty(),
"FleetdAssembly must pass an empty excludedWorkspaceLabels to "
+ "LeadTabScanner — scanning member tabs would demote the lead to a worker");
} finally {
assertNotNull(ports.shutdownHook, "control: assembly must capture its shutdown hook");
ports.shutdownHook.run();
}
}
}
@@ -1,190 +0,0 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.inject.StatusPoller;
import dev.ltms.fleet.msg.LeadChannelHandle;
import dev.ltms.fleet.msg.LeadCoordLoop;
import dev.ltms.fleet.msg.LeadMessage;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.msg.ReplyPushLoop;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Field;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotNull;
/**
* Asserts that the assembled loops use the production reminder, coordination, and delivery timing
* defaults when no {@code primary:} block configures the reply-push values.
*/
class FleetdAssemblyTimingDefaultsTest {
private static final class FakeLeadChannel implements LeadChannelHandle {
@Override
public void publish(String toCoordId, LeadMessage message) {
}
@Override
public List<LeadMessage> peek() {
return List.of();
}
@Override
public void ack(String msgId) {
}
@Override
public String selfCoordId() {
return "test-lead";
}
@Override
public boolean heldDurable() {
return true;
}
@Override
public MailboxState inspect(String coordId) {
return MailboxState.unknown(coordId);
}
@Override
public void close() {
}
}
private static final class TestResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new ReplyInbox() {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) { }
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public boolean ack(String target, String msgId) { return false; }
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> new FakeLeadChannel();
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
};
}
}
private TestResourcePorts ports;
@AfterEach
void tearDown() {
if (ports != null && ports.shutdownHook != null) {
ports.shutdownHook.run();
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
coordinator:
uri: "amqp://fake-lead-broker/vh"
selfId: "test-lead"
""");
return FleetConfig.load(file);
}
private FleetdRuntime assemble(Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ports = new TestResourcePorts();
return FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
}
private static long longField(Object target, String name) throws Exception {
Field field = target.getClass().getDeclaredField(name);
field.setAccessible(true);
return field.getLong(target);
}
@Test
void productionBootPathUsesTheExpectedLoopTimingDefaults(@TempDir Path dir) throws Exception {
FleetdRuntime runtime = assemble(dir);
ReplyPushLoop pushLoop = runtime.pushLoop();
assertEquals(5, longField(pushLoop, "maxReminders"),
"without primary:, ReplyPushLoop must stop after five reminder attempts");
assertEquals(15_000L, longField(pushLoop, "backoffMs"),
"without primary:, ReplyPushLoop must wait fifteen seconds before the next reminder");
LeadCoordLoop leadCoordLoop = runtime.leadCoordLoop();
assertNotNull(leadCoordLoop, "control: coordinator: must build LeadCoordLoop");
assertEquals(3_000L, longField(leadCoordLoop, "intervalMs"),
"LeadCoordLoop must poll for peer-lead mail every three seconds");
StatusPoller poller = runtime.poller();
assertEquals(Injector.POLL_INTERVAL_MILLIS, longField(poller, "intervalMillis"),
"StatusPoller must use Injector's delivery poll interval");
}
}
@@ -24,8 +24,8 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} case — because
* neither of those tests constructs its {@code ConfigRef} through {@code main}; both build their own
* instance directly, wired with the check by hand. That silent regression is exactly the shape
* {@link FleetdBackendQuarantineAssemblyTest}, {@link FleetdLeadSeatAssemblyTest} and {@link
* FleetdCompletionResolverAssemblyTest} already guard against for their own constructor arguments —
* {@link FleetdBackendQuarantineWiringTest}, {@link FleetdLeadSeatWiringTest} and {@link
* FleetdCompletionResolverWiringTest} already guard against for their own constructor arguments —
* this class is the same class of gap for fleetd #474's {@code extraValidation} argument, following
* their approach.
*
@@ -2,67 +2,16 @@ package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* {@code AgentControl} caches {@code paneByTerminal}, so {@code HerdrRouter} must be its only
* production factory — a second instance means a second cache; the same reasoning applies to
* {@code WorkspaceControl}. {@code HerdrRouter}'s constructor is the one place both are built.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* HerdrRouter} and never runs {@code FleetdAssembly.assembleAndStart} — a green result proves only
* that neither watched file's text contains {@code new AgentControl(} or {@code new
* WorkspaceControl(}. It does not prove the instances {@code HerdrRouter} does build are the ones
* actually wired through the rest of the daemon, and it does not cover a bypass written into a
* production file other than the two this test reads.
*/
class FleetdHerdrControlConstructionTest {
private static String source(String relativePath) throws Exception {
return Files.readString(Path.of(relativePath));
}
@Test
@DisplayName("[SOURCE TEXT] Fleetd.java never constructs AgentControl or WorkspaceControl directly")
void fleetdDelegatesStatefulControlsToTheRouter() throws Exception {
String source = source("src/main/java/dev/ltms/fleet/Fleetd.java");
// A broken read (wrong working directory, wrong path, a file that came back empty) would
// make the assertFalse checks below pass vacuously — a "clean" negative check that actually
// checked nothing. Guard against that first, with an anchor that has nothing to do with
// this mutation, so a bad read fails loudly here instead of silently proving nothing below.
assertTrue(source.contains("public final class Fleetd"),
"the read of Fleetd.java did not come back containing its own class declaration — "
+ "the assertFalse checks below would pass vacuously on a broken read; fix the "
+ "read before trusting this test.");
assertFalse(source.contains("new AgentControl("),
"Fleetd.java must not construct AgentControl directly — HerdrRouter is its only "
+ "production factory");
assertFalse(source.contains("new WorkspaceControl("),
"Fleetd.java must not construct WorkspaceControl directly — HerdrRouter is its only "
+ "production factory");
}
@Test
@DisplayName("[SOURCE TEXT] FleetdAssembly.java never constructs AgentControl or WorkspaceControl directly")
void fleetdAssemblyDelegatesStatefulControlsToTheRouter() throws Exception {
String source = source("src/main/java/dev/ltms/fleet/FleetdAssembly.java");
assertTrue(source.contains("final class FleetdAssembly"),
"the read of FleetdAssembly.java did not come back containing its own class "
+ "declaration — the assertFalse checks below would pass vacuously on a broken "
+ "read; fix the read before trusting this test.");
assertFalse(source.contains("new AgentControl("),
"FleetdAssembly.java must not construct AgentControl directly — HerdrRouter is its "
+ "only production factory");
assertFalse(source.contains("new WorkspaceControl("),
"FleetdAssembly.java must not construct WorkspaceControl directly — HerdrRouter is "
+ "its only production factory");
// AgentControl caches paneByTerminal, so the router must be its only production factory.
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
assertFalse(source.contains("new AgentControl("));
assertFalse(source.contains("new WorkspaceControl("));
}
}
@@ -1,63 +0,0 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.mcp.FleetMcp;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
/**
* Pins {@link Fleetd#leadConfigDirSource}'s own wiring of the window lookup into the returned
* {@link FleetMcp.LeadConfigDirSource}, not only the detached {@link Fleetd#leadContextWindowLookup}
* factory it delegates to. Calls the producer directly, with real {@link FleetConfig.Profile}/
* {@link FleetConfig.Leader} fixtures, and asserts on {@code windowFor()} — the companion of
* {@link FleetdLeadConfigDirSourceWiringTest}, which pins the same factory's {@code configDirFor()}.
*/
class FleetdLeadConfigDirSourceWindowWiringTest {
private static FleetConfig.Profile profileWithWindow(String name, Integer autoCompactWindow) {
return new FleetConfig.Profile(name, null, "claude-sonnet-5", null, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, null, true, null, null, null, null, null, autoCompactWindow, null);
}
private static FleetConfig.Leader leadOnProfile(String profile) {
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
}
@Test
@DisplayName("the returned source resolves the lead's REAL configured effective window, not a hardcoded null")
void resolvesTheRealConfiguredWindow() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithWindow("opus", 250_000));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
assertEquals(250_000L, source.windowFor().apply("primary"),
"windowFor must delegate to the real leadContextWindowLookup, not a stub that always "
+ "returns null");
}
@Test
@DisplayName("a lead on a profile with no window configured still resolves to null, not a crash")
void leadWithNoWindowConfiguredResolvesToNull() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithWindow("opus", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
assertNull(source.windowFor().apply("primary"));
}
@Test
@DisplayName("an unrecognised lead name resolves to null, not a thrown exception")
void unrecognisedLeadNameResolvesToNull() {
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(Map::of, Map.of());
assertNull(source.windowFor().apply("ghost-lead"));
}
}
@@ -99,7 +99,7 @@ class FleetdLeadContextLookupTest {
@DisplayName("an unrecognised terminal resolves to UNKNOWN, not a thrown exception")
void unrecognisedTerminalResolvesToUnknown() {
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), throwingAgentControl(), Map::of, name -> null, name -> null);
new LeadContextGauge(), throwingAgentControl(), Map::of, name -> null);
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply("ghost-terminal"));
@@ -112,7 +112,7 @@ class FleetdLeadContextLookupTest {
void agentsGetThrowingDegradesToUnknown() {
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), throwingAgentControl(), () -> liveLeadTerminals, name -> null, name -> null);
new LeadContextGauge(), throwingAgentControl(), () -> liveLeadTerminals, name -> null);
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply(LEAD_TERMINAL));
@@ -128,7 +128,7 @@ class FleetdLeadContextLookupTest {
AgentControl agents = agentControlStub(SESSION_ID, "claude", "idle");
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), agents, () -> liveLeadTerminals,
name -> LEAD_NAME.equals(name) ? tmp.toString() : null, name -> null);
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
@@ -146,7 +146,7 @@ class FleetdLeadContextLookupTest {
AgentControl agents = agentControlStub(SESSION_ID, "opencode", "idle");
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), agents, () -> liveLeadTerminals,
name -> LEAD_NAME.equals(name) ? tmp.toString() : null, name -> null);
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
@@ -1,223 +0,0 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.lang.reflect.Field;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #659: {@code FleetdAssembly.java} wires {@code Fleetd.leadContextSource}'s window-lookup
* argument with {@code Fleetd.leadContextWindowLookup(() -> config.get().profiles(), leaders)} —
* but nothing called the real assembled {@link LeadHeartbeatLoop} far enough to prove that
* argument is the one the live heartbeat reads through. Measured: swapping that one call-site
* argument for {@code _ -> null} compiles with 0 errors and leaves the full suite green.
*
* <p>This test drives the REAL {@link LeadHeartbeatLoop} the real {@link
* FleetdAssembly#assembleAndStart} builds, reached through {@link FleetdRuntime#heartbeat()}, and
* reads its private {@code contextSource} field via reflection — the loop exposes no public
* accessor for it, the same reason {@link FleetdLeadConfigDirSourceAssemblyTest} reflects on
* {@code FleetMcp.leadConfigDirs}. The configured profile's {@code autoCompactWindow: 100000}
* resolves a HIGH threshold of {@code 66666} ({@link LeadContextGauge}'s {@code 2/3} fraction) —
* far below the fixed {@code 200000} fallback a lost window argument would silently revert to.
* {@code 90000} live tokens sits between the two: HIGH under the real window, OK under the
* fallback — a property the fallback can never produce by accident.
*/
class FleetdLeadContextSourceWindowAssemblyTest {
private static final String LEAD_NAME = "opus";
private static final String LEAD_TAB = "lead: opus";
private static final String LEAD_PROFILE = "sonnet";
/** {@code FakeHerdr}'s own default {@code agent.list} entry: terminal {@code term_a}, session {@code sess-1111}. */
private static final String LEAD_TERMINAL = "term_a";
private static final String LEAD_SESSION_ID = "sess-1111";
private static final class RecordingResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> replyInbox;
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
}
@Override
public void startHttp(Javalin app, String host, int port) {
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
@Override
public void own(String target) {
}
@Override
public void release(String target) {
}
@Override
public void publish(String target, String msgId, String content) {
}
@Override
public List<InboxMessage> peek(String target) {
return List.of();
}
@Override
public boolean ack(String target, String msgId) {
return false;
}
@Override
public void close() {
}
}
private static String usageLine(long tokens) {
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
}
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} carrying one usage record. */
private static void writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
Files.createDirectories(projectDir);
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
}
private static FleetConfig writeConfig(Path dir, String configDir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
leadHeartbeat:
idleAfterSeconds: 600
backoffMs: 15000
quietNudgeCap: 5
fleet:
leaders:
%s:
tab: "%s"
profile: %s
profiles:
%s:
subscription: true
argv: ["ccs", "sonnet"]
configDir: "%s"
autoCompactWindow: 100000
""".formatted(LEAD_NAME, LEAD_TAB, LEAD_PROFILE, LEAD_PROFILE, configDir));
return FleetConfig.load(f);
}
@SuppressWarnings("unchecked")
private static LeadHeartbeatLoop.LeadContextSource contextSourceOf(LeadHeartbeatLoop heartbeat) throws Exception {
Field field = LeadHeartbeatLoop.class.getDeclaredField("contextSource");
field.setAccessible(true);
return (LeadHeartbeatLoop.LeadContextSource) field.get(heartbeat);
}
@Test
@DisplayName("[BEHAVIOURAL] the real assembled heartbeat loop resolves HIGH against the lead's "
+ "REAL configured window, not the fixed 200000 fallback a lost window argument reverts to")
void assembledHeartbeatContextSourceResolvesTheRealConfiguredWindow(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir, dir.toString());
writeTranscript(dir, LEAD_SESSION_ID, 90_000);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
RecordingResourcePorts ports = new RecordingResourcePorts();
// Label FakeHerdr's own default pane's tab (term_a / w2:p7 / w2:t7, already carrying a live
// agent on session sess-1111) to match fleet.leaders.opus.tab exactly, so LeadTabScanner
// recognises it as the live "opus" lead without a second auto-launched pane.
ports.herdr.withTab("w2", "w2:t7", LEAD_TAB).agentSessionId(LEAD_SESSION_ID);
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
try {
LeadHeartbeatLoop.LeadContextSource source = contextSourceOf(runtime.heartbeat());
LeadContextGauge.Reading reading = source.readingFor().apply(LEAD_TERMINAL);
assertEquals(LeadContextGauge.State.HIGH, reading.state(),
"profiles." + LEAD_PROFILE + ".autoCompactWindow: 100000 resolves a HIGH threshold "
+ "of 66666 tokens — 90000 live tokens must read HIGH against it. Mutating "
+ "FleetdAssembly's window-lookup argument to `_ -> null` falls back to the "
+ "fixed 200000 threshold, under which 90000 reads OK instead: " + reading);
} finally {
// Surefire runs the whole suite in one JVM fork (fleetd/pom.xml sets no forkCount /
// reuseForks), so the scheduler/loops this assembly starts must be torn down here, on the
// failure path too — hence try/finally rather than a bare statement at the end.
runtime.close();
}
}
}
@@ -87,8 +87,7 @@ class FleetdLeadContextSourceWiringTest {
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
agentControlStub(), () -> liveLeadTerminals,
name -> LEAD_NAME.equals(name) ? tmp.toString() : null, name -> null);
agentControlStub(), () -> liveLeadTerminals, name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
LeadContextGauge.Reading reading = source.readingFor().apply(LEAD_TERMINAL);
@@ -102,7 +101,7 @@ class FleetdLeadContextSourceWiringTest {
@DisplayName("an unrecognised lead terminal resolves to UNKNOWN, not a thrown exception")
void unrecognisedTerminalResolvesToUnknown() {
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
agentControlStub(), Map::of, name -> null, name -> null);
agentControlStub(), Map::of, name -> null);
assertEquals(LeadContextGauge.State.UNKNOWN, source.readingFor().apply("ghost-terminal").state());
}
@@ -6,8 +6,6 @@ import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
@@ -36,7 +34,7 @@ import static org.junit.jupiter.api.Assertions.fail;
* fleetd #612 Unit A) with three methods: {@code unrelatedAnchorStillPresent} (a scaffold anchor,
* not an independent claim — needs no replacement of its own), {@code
* mainStillCallsTheLeadRolloverFactory} (the call-site pin replaced by {@link
* #assembledLeadRolloverEndsTheOldPaneThroughTheRealHerdrRouter}), and {@code
* #assembledLeadRolloverRunsTheRealClearAndBootstrapSequence}), and {@code
* factoryGatesOnConfigPresence} (the absent-config claim replaced by {@link
* #absentLeadRolloverConfigMeansNoRolloverIsBuilt} — a claim this ticket found was NOT actually
* covered behaviourally anywhere else: {@code LeadRolloverTest}'s only related assertion is
@@ -52,8 +50,8 @@ import static org.junit.jupiter.api.Assertions.fail;
* invisible to this test, even though the two are genuinely different daemons in production. This
* version configures two distinct sockets and two distinct {@link FakeHerdr} instances (the same
* pattern {@code FleetdAssemblyConnectionIdentityTest}, fleetd #612 B2, already uses to separate
* lead from member) and asserts the roll's {@code pane.close} call lands on the LEAD fake and
* never on the MEMBER one.
* lead from member) and asserts the roll's {@code /clear}/bootstrap sends land on the LEAD fake
* and never on the MEMBER one.
*/
class FleetdLeadRolloverAssemblyTest {
@@ -179,47 +177,11 @@ class FleetdLeadRolloverAssemblyTest {
return FleetConfig.load(f);
}
/**
* Unlike {@link #writeConfig}, this names a {@code profile:} for the lead and declares it
* under {@code profiles:}, so {@code LeadLauncher#relaunch} can actually start a fresh agent
* instead of refusing with "names no profile". {@code relaunchReadySeconds} is cut to 2s so
* the recognition wait (expected to time out — see the test) does not cost real test seconds.
*/
private static FleetConfig writeConfigWithRelaunchableLead(Path dir, Path leadCwd) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: "%s"
memberHerdrSocket: "%s"
idleSleepGuard:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
fleet:
leaders:
opus:
tab: "lead: opus"
cwd: "%s"
profile: opus
profiles:
opus:
subscription: true
argv: ["ccs", "opus"]
leadRollover:
handoverPath: handover.md
requireOperatorConfirm: false
relaunchReadySeconds: 2
""".formatted(LEAD_SOCKET, MEMBER_SOCKET, leadCwd.toString()));
return FleetConfig.load(f);
}
@SuppressWarnings("unchecked")
@Test
@DisplayName("[BEHAVIOURAL] the real assembled LeadRollover runs the open/confirm/continuation "
+ "sequence through the real herdr router — ending the old pane, then giving up once it "
+ "never reports gone")
void assembledLeadRolloverEndsTheOldPaneThroughTheRealHerdrRouter(@TempDir Path dir) throws Exception {
@DisplayName("[BEHAVIOURAL] the real assembled LeadRollover runs the full open/confirm/continuation "
+ "sequence — /clear, then bootstrapText — through the real herdr router")
void assembledLeadRolloverRunsTheRealClearAndBootstrapSequence(@TempDir Path dir) throws Exception {
Path leadCwd = dir.resolve("lead-workspace");
Files.createDirectories(leadCwd);
FleetConfig cfg = writeConfig(dir, leadCwd);
@@ -242,11 +204,6 @@ class FleetdLeadRolloverAssemblyTest {
+ "Fleetd.leadRollover(...) call site — a mutation to `LeadRollover leadRollover = "
+ "null;` at that call site can never pass this");
// assembleAndStart's own boot work (the orphan-worker reap) makes a real call on the
// member daemon before the roll ever starts. Clear it here so the assertion below measures
// only what the roll itself does, not what daemon startup does.
member.calls.clear();
LeadRollover.PendingRollover pending = rollover.open("term_a", "fleetd #612 B3 test");
String expectedHandoverPath = leadCwd.resolve("handover.md").normalize().toString();
assertEquals(expectedHandoverPath, pending.handoverPath());
@@ -266,109 +223,40 @@ class FleetdLeadRolloverAssemblyTest {
// constructor), so this polls the real FleetMcp.leadRollover() instance's status(token)
// until the real continuation finishes.
LeadRollover.RollStatus status = pollUntilTerminal(rollover, pending.token());
assertEquals(LeadRollover.RollState.ROLLED, status.state(),
"the full happy path must complete: FakeHerdr's default agent status is 'idle', so "
+ "the turn-boundary wait settles immediately and the post-/clear wait "
+ "releases via its pickup-grace path — detail: " + status.detail());
// FakeHerdr's pane.get is a fixed canned response that never reports a pane as gone, so the
// real router's death poll runs out its whole budget and the roll stops here — proving the
// real teardown call landed on the real LEAD pane without ever reaching a relaunch or a send.
assertEquals(LeadRollover.RollState.OLD_PANE_NEVER_DIED, status.state(),
"the old pane never reports gone against this fake, so the roll must stop with "
+ "OLD_PANE_NEVER_DIED rather than ever relaunching or sending anything — "
+ "detail: " + status.detail());
// Prove the real herdr router actually closed the real LEAD pane — this is the one thing a
// source-text pin on the call site could never show.
boolean closedOldPane = lead.calls.stream()
.anyMatch(c -> c.method().equals("pane.close")
&& c.params() instanceof Map<?, ?> m && "w2:p7".equals(m.get("pane_id")));
assertTrue(closedOldPane, "endOldSession must close the real old pane (w2:p7) through the "
+ "real LEAD herdr client, got calls: " + lead.calls);
// No agent.prompt is ever sent on this path: the roll stops at the pane-death wait, strictly
// before the relaunch and the final send step.
// Prove the real herdr router actually sent BOTH messages, in order, to the real LEAD
// pane — this is the one thing a source-text pin on the call site could never show.
List<FakeHerdr.Call> prompts = lead.calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.toList();
assertTrue(prompts.isEmpty(), "a roll that stops at OLD_PANE_NEVER_DIED must never reach the "
+ "send step, got agent.prompt call(s) on the LEAD daemon: " + prompts);
assertTrue(prompts.size() >= 2, "expected at least a /clear send and a bootstrapText send "
+ "on the LEAD daemon, got " + prompts.size() + " agent.prompt calls: " + prompts);
assertEquals("/clear", ((Map<String, Object>) prompts.get(0).params()).get("text"),
"the first send must be the literal /clear housekeeping command");
Object secondText = ((Map<String, Object>) prompts.get(1).params()).get("text");
assertTrue(secondText instanceof String && ((String) secondText).contains(expectedHandoverPath),
"the second send must be the default bootstrapText naming the resolved handover "
+ "path, got: " + secondText);
// fleetd #612 B3 correction: prove the roll never touches the MEMBER daemon. A mutation
// swapping router.leadAgents() for router.memberAgents() at the real call site would move
// the pane.close call above onto `member` instead, which this assertion catches — the thing
// the single-fake version of this test could never see, because both wrapped the same client.
assertTrue(member.calls.isEmpty(), "the roll must be wired to the LEAD daemon only — got "
+ member.calls.size() + " call(s) recorded on the MEMBER daemon since the roll began: "
+ member.calls);
}
/**
* Exercises the relaunch site {@link #assembledLeadRolloverEndsTheOldPaneThroughTheRealHerdrRouter}
* never reaches: with the old pane confirmed gone, the roll relaunches a fresh lead, and
* {@code bootstrapText} must reach it even though recognition times out (FakeHerdr's
* {@code tab.list} is a fixed canned response that never reflects the relaunch's own
* {@code tab.rename}, so the fresh terminal is never recognised as a live lead). Same
* dual-socket shape as the sibling test: two distinct {@link FakeHerdr} instances, so a
* {@code bootstrapText} send wired to the wrong daemon is visible.
*/
@Test
@DisplayName("[BEHAVIOURAL] bootstrapText reaches the fresh LEAD terminal even when recognition "
+ "times out, and the MEMBER daemon never sees it")
void bootstrapTextReachesTheFreshLeadTerminalEvenWhenRecognitionTimesOut(@TempDir Path dir) throws Exception {
Path leadCwd = dir.resolve("lead-workspace");
Files.createDirectories(leadCwd);
FleetConfig cfg = writeConfigWithRelaunchableLead(dir, leadCwd);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
RecordingResourcePorts ports = new RecordingResourcePorts();
FakeHerdr lead = new FakeHerdr();
lead.withTab("w2", "w2:t7", "lead: opus");
// Lets the old pane (w2:p7) report gone once pane.close actually reaches it, so the roll
// proceeds to relaunch instead of stopping at OLD_PANE_NEVER_DIED.
lead.paneGoneAfterClose("w2:p7");
FakeHerdr member = new FakeHerdr();
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
LeadRollover rollover = runtime.mcp().leadRollover();
assertNotNull(rollover, "leadRollover: is present in this test's config, so a real "
+ "LeadRollover must have been built");
LeadRollover.PendingRollover pending = rollover.open("term_a", "bootstrapText relaunch test");
Thread.sleep(50);
Files.writeString(Path.of(pending.handoverPath()), "handover content for bootstrapText test");
LeadRollover.RollDecision decision = rollover.confirm("term_a", pending.token(), true);
assertTrue(decision.accepted(), "confirm() must approve — got: " + decision);
LeadRollover.RollStatus status = pollUntilTerminal(rollover, pending.token());
assertEquals(LeadRollover.RollState.RELAUNCH_NOT_RECOGNISED, status.state(),
"the fresh terminal is never recognised against this fake's static tab.list, so the "
+ "roll must reach RELAUNCH_NOT_RECOGNISED — not an earlier failure state and "
+ "not ROLLED — detail: " + status.detail());
@SuppressWarnings("unchecked")
List<FakeHerdr.Call> leadPrompts = lead.calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.toList();
assertEquals(1, leadPrompts.size(), "exactly one bootstrapText send is expected, on the LEAD "
+ "daemon, once recognition gives up — got: " + leadPrompts);
Object text = ((Map<String, Object>) leadPrompts.get(0).params()).get("text");
assertTrue(text instanceof String && ((String) text).contains("handover.md"),
"the send must be bootstrapText naming the resolved handover path, got: " + text);
// Scoped to agent.prompt specifically, not every MEMBER call: the orphan-worker reap also
// talks to the MEMBER daemon once, unconditionally, at daemon boot — unrelated to this roll.
// both sends above onto `member` instead, which this assertion catches — the thing the
// single-fake version of this test could never see, because both wrapped the same client.
List<FakeHerdr.Call> memberPrompts = member.calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.toList();
assertTrue(memberPrompts.isEmpty(), "bootstrapText must never be sent to the MEMBER daemon, "
+ "got: " + memberPrompts);
assertTrue(memberPrompts.isEmpty(), "the roll must be wired to the LEAD daemon only — got "
+ memberPrompts.size() + " agent.prompt call(s) on the MEMBER daemon instead: "
+ memberPrompts);
}
private static LeadRollover.RollStatus pollUntilTerminal(LeadRollover rollover, String token)
throws InterruptedException {
long deadline = System.nanoTime() + java.util.concurrent.TimeUnit.SECONDS.toNanos(15);
long deadline = System.nanoTime() + java.util.concurrent.TimeUnit.SECONDS.toNanos(10);
while (System.nanoTime() < deadline) {
LeadRollover.RollStatus status = rollover.status(token);
if (status.state() != LeadRollover.RollState.PENDING
@@ -395,10 +283,8 @@ class FleetdLeadRolloverAssemblyTest {
""");
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
AgentControl agents = new AgentControl(new FakeHerdr());
WorkspaceControl spaces = new WorkspaceControl(new FakeHerdr());
LeadLauncher launcher = new LeadLauncher(agents, spaces, config.get());
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, spaces, launcher, config, Map::of);
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, config, Map::of);
assertNull(rollover, "leadRollover: is absent from this config, so the factory's opt-in "
+ "gate (`if (cfg.leadRollover() == null) return null;`) must fire and no "
@@ -4,8 +4,6 @@ import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.lead.LeadRollover;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
@@ -56,11 +54,6 @@ class FleetdLeadRolloverWorkspaceLookupTest {
return new AgentControl(new FakeHerdr());
}
/** None of this class's tests reach the deferred continuation, so a plain fake is enough. */
private static LeadLauncher fakeLauncher(FleetConfig cfg) {
return new LeadLauncher(fakeAgents(), new WorkspaceControl(new FakeHerdr()), cfg);
}
@Test
@DisplayName("[BEHAVIOURAL] Fleetd.leadRollover(...) resolves a relative handoverPath against "
+ "the CALLING lead's configured cwd, not the daemon's own working directory")
@@ -81,8 +74,7 @@ class FleetdLeadRolloverWorkspaceLookupTest {
""".formatted(leadCwd.toString()));
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config,
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
() -> Map.of("term_opus", "opus"));
assertNotNull(rollover, "leadRollover: is present in the loaded config, so the factory "
+ "must construct an object");
@@ -113,8 +105,7 @@ class FleetdLeadRolloverWorkspaceLookupTest {
// No lead has been discovered yet — exactly the real shape of a lead the live tab scan
// has not yet scanned, or one with no fleet.leaders entry at all.
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config, Map::of);
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config, Map::of);
assertNotNull(rollover);
LeadRollover.PendingRollover pending = rollover.open("term_unknown", "test");
@@ -153,8 +144,7 @@ class FleetdLeadRolloverWorkspaceLookupTest {
// below — exactly the natural mistake to make, since leads are discovered by a live tab
// scan that runs AFTER this factory is constructed at startup.
Map<String, String> liveLeadTerminals = new HashMap<>();
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config,
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
() -> liveLeadTerminals);
assertNotNull(rollover);
@@ -14,8 +14,6 @@ class AuthzTest {
private static final Principal ANON = Principal.anonymous();
private static final Principal ARCH_DESIGN = Principal.architect("lead-designer", "term_design", 400);
private static final Principal ARCH_OTHER = Principal.architect("reviewer", "term_review", 500);
private static final Principal COLLABORATOR = Principal.collaborator("ops", "term_collab", 600);
private static final Principal OBSERVER = Principal.observer("term_observer", 700);
@Test
void anonymousIsAuthorizedForNothing() {
@@ -31,22 +29,6 @@ class AuthzTest {
assertTrue(Authz.isUnauthenticated(null));
}
/**
* {@code MessageService.answer}'s turn-ownership check treats a caller with no terminal as
* matching a turn recorded for the unnamed primary. That rule only stays safe because an
* unauthenticated caller — whose terminal is also {@code null} — never reaches {@code answer}
* at all: {@link #anonymousIsAuthorizedForNothing} already covers every action including
* {@code ANSWER}, but this test names the exact coupling so a future change to either side
* cannot drift without turning this test red.
*/
@Test
void anAnonymousCallerIsRefusedAnswerSoItCanNeverBeMistakenForTheUnnamedPrimary() {
assertFalse(Authz.permits(ANON, ANSWER, null),
"an anonymous caller, whose terminal is also null, must never reach answer() — the "
+ "turn-ownership check's null-terminal match for the unnamed primary owner "
+ "relies on this gate refusing it first");
}
@Test
void orchestrationBelongsToThePrimaryAlone() {
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, SEND, DRAIN}) {
@@ -56,37 +38,6 @@ class AuthzTest {
}
}
/**
* {@code fleet_send} is three call shapes behind one action name until {@code
* FleetMcp#sendAction} picks one: a plain local {@link Authz.Action#SEND}, the {@code coordId}
* route ({@link Authz.Action#COORD_SEND}), and the {@code turnId} answer form ({@link
* Authz.Action#ANSWER}). All three carry the same grant as the undivided action did — a worker
* is excluded from every one, exactly as it was excluded from the one combined action before.
*/
@Test
void theThreeSendShapesCarryTheSameGrantAsTheOldUndividedAction() {
for (Authz.Action a : new Authz.Action[]{SEND, COORD_SEND, ANSWER}) {
assertTrue(Authz.permits(PRIMARY, a, "term_a"), "the primary may " + a);
assertTrue(Authz.permits(ARCH_DESIGN, a, "term_a"), "an architect may " + a);
assertFalse(Authz.permits(WORKER_A, a, "term_a"),
"a worker performing " + a + " would be escalating into the orchestrator role");
assertFalse(Authz.permits(ANON, a, "term_a"));
}
}
/**
* {@code fleet_poll{ticket}} and {@code fleet_status} are {@link Authz.Action#TASK_READ}, split
* out of the roster-only {@link Authz.Action#READ} (fleetd #678). The grant is unchanged from
* what the undivided {@code READ} action gave every one of these callers.
*/
@Test
void taskReadCarriesTheSameGrantReadDidBeforeTheSplit() {
assertTrue(Authz.permits(PRIMARY, TASK_READ, null));
assertTrue(Authz.permits(WORKER_A, TASK_READ, null));
assertTrue(Authz.permits(ARCH_DESIGN, TASK_READ, null));
assertFalse(Authz.permits(ANON, TASK_READ, null));
}
@Test
void aWorkerMayReplyAndAskOnlyAsItself() {
assertTrue(Authz.permits(WORKER_A, REPLY, "term_a"));
@@ -184,132 +135,4 @@ class AuthzTest {
assertFalse(Authz.isUnauthenticated(WORKER_A));
assertFalse(Authz.isUnauthenticated(PRIMARY));
}
// ── the collaborator matrix ─────────────────────────────────────────────────────────────────
/**
* {@code SEND} for a collaborator is the one grant that is conditional rather than fixed:
* flipping only the classifier's answer for the target flips only this outcome.
*/
@Test
void aCollaboratorMaySendOnlyWhenTheClassifierAcceptsTheTarget() {
assertTrue(Authz.permits(COLLABORATOR, SEND, "term_lead", target -> true),
"the classifier accepting the target must grant SEND");
assertFalse(Authz.permits(COLLABORATOR, SEND, "term_lead", target -> false),
"the classifier refusing the target must deny SEND");
assertFalse(Authz.permits(COLLABORATOR, SEND, "term_lead"),
"the real production classifier recognises no terminal yet, so SEND is refused today");
}
/**
* Control for the test above: every other action's result for a collaborator does not move
* when the classifier does. Only {@code SEND} is wired to it.
*/
@Test
void theClassifierMovesOnlySendForACollaborator() {
for (Authz.Action a : Authz.Action.values()) {
if (a == SEND) {
continue;
}
assertEquals(
Authz.permits(COLLABORATOR, a, "term_lead"),
Authz.permits(COLLABORATOR, a, "term_lead", target -> true),
a + " must not depend on the classifier at all");
}
}
@Test
void aCollaboratorMayReadAndScrapeMetrics() {
assertTrue(Authz.permits(COLLABORATOR, READ, null));
assertTrue(Authz.permits(COLLABORATOR, METRICS, null));
}
@Test
void aCollaboratorMayReplyAndAskOnlyAsItsOwnPane() {
assertTrue(Authz.permits(COLLABORATOR, REPLY, "term_collab"),
"its own pane is its own");
assertTrue(Authz.permits(COLLABORATOR, ASK, "term_collab"));
assertFalse(Authz.permits(COLLABORATOR, REPLY, "term_design"),
"a collaborator must not reply on another pane");
assertFalse(Authz.permits(COLLABORATOR, REPLY, null),
"an absent target must not pass the own-session rule");
}
/**
* Every action denied to a collaborator, asserted denied even when the classifier would
* accept any target — proving none of these is actually gated on the classifier at all.
*/
@Test
void aCollaboratorIsDeniedLifecycleCoordinationAndTicketPolling() {
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, DRAIN, HANDOVER, ANSWER, COORD_SEND,
COORD_READ, TASK_READ}) {
assertFalse(Authz.permits(COLLABORATOR, a, "term_lead", target -> true),
"a collaborator must not " + a + " even when the classifier accepts every target");
}
}
@Test
void aCollaboratorIsNotCountedAsPrimaryWorkerOrArchitect() {
assertFalse(COLLABORATOR.isPrimary());
assertFalse(COLLABORATOR.isWorker());
assertFalse(COLLABORATOR.isArchitect());
assertTrue(COLLABORATOR.isCollaborator());
}
/**
* A collaborator is never spawned, so it must not be enrolled in the presence map as an
* available member. Control: both a worker and an architect — which ARE spawned — still are.
*/
@Test
void isSpawnedMemberIsFalseForACollaboratorButTrueForAWorkerAndAnArchitect() {
assertFalse(COLLABORATOR.isSpawnedMember());
assertTrue(WORKER_A.isSpawnedMember());
assertTrue(ARCH_DESIGN.isSpawnedMember());
}
// ── the observer matrix ─────────────────────────────────────────────────────────────────────
@Test
void anObserverMayReadAndScrapeMetrics() {
assertTrue(Authz.permits(OBSERVER, READ, null));
assertTrue(Authz.permits(OBSERVER, METRICS, null));
}
@Test
void anObserverMayReplyAndAskOnlyAsItsOwnPane() {
assertTrue(Authz.permits(OBSERVER, REPLY, "term_observer"), "its own pane is its own");
assertTrue(Authz.permits(OBSERVER, ASK, "term_observer"));
assertFalse(Authz.permits(OBSERVER, REPLY, "term_design"),
"an observer must not reply on another pane");
assertFalse(Authz.permits(OBSERVER, REPLY, null),
"an absent target must not pass the own-session rule");
}
/**
* Every action beyond READ/METRICS/REPLY/ASK, asserted denied for an observer — including
* {@code TASK_READ}, which is the entire point of this role: an unconfigured pane must not be
* able to poll a ticket or read another session's status.
*/
@Test
void anObserverIsDeniedEverythingBeyondReadMetricsReplyAndAsk() {
for (Authz.Action a : Authz.Action.values()) {
if (a == READ || a == METRICS || a == REPLY || a == ASK) {
continue;
}
assertFalse(Authz.permits(OBSERVER, a, "term_observer", target -> true),
"an observer must not " + a + " even when the classifier accepts every target");
}
}
@Test
void anObserverIsNotCountedAsAnyOtherRole() {
assertFalse(OBSERVER.isPrimary());
assertFalse(OBSERVER.isWorker());
assertFalse(OBSERVER.isArchitect());
assertFalse(OBSERVER.isCollaborator());
assertFalse(OBSERVER.isSpawnedMember());
assertTrue(OBSERVER.isObserver());
}
}
@@ -1,25 +1,13 @@
package dev.ltms.fleet.auth;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.session.WorktreeRequest;
import dev.ltms.fleet.session.Worktrees;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
import static org.junit.jupiter.api.Assertions.*;
@@ -57,23 +45,17 @@ class CallerResolverTest {
return members;
}
/**
* With no roster wired up at all (the simple constructor), a loopback pane that owns a herdr
* pane but is not recognised as a live spawned member lands on the {@link Role#OBSERVER} floor
* — unforgeable and never token-gated, exactly like a worker's own identity, because it comes
* from the same connection-derived pane mapping.
*/
@Test
void aLoopbackPaneWithNoLiveRosterResolvesToObserverRegardlessOfAuthMode() {
void aLoopbackWorkerPaneResolvesToWorkerRegardlessOfAuthMode() {
Principal underTrust = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
Principal underToken = new CallerResolver(workerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.OBSERVER, underTrust.role());
assertEquals(Role.WORKER, underTrust.role());
assertEquals("term_a", underTrust.terminal());
assertEquals(Role.OBSERVER, underToken.role(),
"the floor is unforgeable and must never be token-gated — otherwise enabling auth "
+ "would lock every unconfigured pane out of even READ");
assertEquals(Role.WORKER, underToken.role(),
"worker identity is unforgeable and must never be token-gated — otherwise enabling "
+ "auth would lock the whole fleet out of fleet_reply");
assertEquals("term_a", underToken.terminal());
}
@@ -97,20 +79,20 @@ class CallerResolverTest {
}
@Test
void otherPanesRemainAtTheFloorWhenAPinIsSet() {
void otherPanesRemainWorkersWhenAPinIsSet() {
Principal p = CallerResolver.pinnedTo(workerIdentity(), false, null, "term_someone_else")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.OBSERVER, p.role());
assertEquals(Role.WORKER, p.role());
assertEquals("term_a", p.terminal());
}
/** The pin is optional config, so an absent or whitespace one must change nothing at all. */
@Test
void aBlankPinLeavesFloorResolutionUntouched() {
assertEquals(Role.OBSERVER,
void aBlankPinLeavesWorkerResolutionUntouched() {
assertEquals(Role.WORKER,
CallerResolver.pinnedTo(workerIdentity(), false, null, " ").resolve("127.0.0.1", 42, null).role());
assertEquals(Role.OBSERVER,
assertEquals(Role.WORKER,
CallerResolver.pinnedTo(workerIdentity(), false, null, null).resolve("127.0.0.1", 42, null).role());
}
@@ -231,13 +213,13 @@ class CallerResolverTest {
* mid-scan teardown into a refusal — the real match is still found and resolves as a worker.
*/
@Test
void aHerdrErrorOnANonOwningPaneStillResolvesTheRealPane() {
void aHerdrErrorOnANonOwningPaneStillResolvesTheRealWorker() {
FakeHerdr vanishedElsewhere = new FakeHerdr().processInfoFailsForPane("w2:p9", "pane_not_found");
ConnectionIdentity id = new ConnectionIdentity(new PaneLocator(vanishedElsewhere), _ -> FakeHerdr.WORKER_PID);
Principal p = new CallerResolver(id).resolve("127.0.0.1", 55555, null);
assertEquals(Role.OBSERVER, p.role());
assertEquals(Role.WORKER, p.role());
assertEquals("term_a", p.terminal());
}
@@ -281,12 +263,12 @@ class CallerResolverTest {
}
@Test
void aPaneAbsentFromTheRegistryFallsToTheObserverFloor() {
void aPaneAbsentFromTheRegistryIsStillAWorker() {
Principal p = new CallerResolver(workerIdentity(), false, null,
Map.of("term_elsewhere", "gpt-sol-5.6"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.OBSERVER, p.role());
assertEquals(Role.WORKER, p.role());
assertEquals("term_a", p.terminal());
assertNull(p.name());
}
@@ -312,16 +294,16 @@ class CallerResolverTest {
}
@Test
void anEmptyRegistryLeavesEveryPaneAtTheObserverFloor() {
void anEmptyRegistryLeavesEveryPaneAWorker() {
Map<String, String> noLeads = null;
assertEquals(Role.OBSERVER,
assertEquals(Role.WORKER,
new CallerResolver(workerIdentity(), false, null, Map.of())
.resolve("127.0.0.1", 42, null).role());
assertEquals(Role.OBSERVER,
assertEquals(Role.WORKER,
new CallerResolver(workerIdentity(), false, null, noLeads)
.resolve("127.0.0.1", 42, null).role());
// And the same for the live-registry form, whose supplier may also be absent.
assertEquals(Role.OBSERVER,
// CB-531: and the same for the live-registry form, whose supplier may also be absent.
assertEquals(Role.WORKER,
CallerResolver.withLeads(workerIdentity(), false, null, null)
.resolve("127.0.0.1", 42, null).role());
}
@@ -394,7 +376,7 @@ class CallerResolverTest {
Map<String, String> live = new java.util.HashMap<>();
CallerResolver r = CallerResolver.withLeads(workerIdentity(), false, null, () -> live);
assertEquals(Role.OBSERVER, r.resolve("127.0.0.1", 42, null).role());
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
live.put("term_a", "gpt-sol-5.6"); // the scanner sees a newly-labelled tab
@@ -411,7 +393,7 @@ class CallerResolverTest {
mutable.put("term_a", "sneaky");
assertEquals(Role.OBSERVER, r.resolve("127.0.0.1", 42, null).role());
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
}
// ── CB-548: architect slots ─────────────────────────────────────────────────────────────────
@@ -441,13 +423,13 @@ class CallerResolverTest {
}
@Test
void anUnboundPaneResolvesToTheObserverFloor() {
void anUnboundPaneStillResolvesAsAWorker() {
MemberRegistry members = new MemberRegistry(new FleetConfig.Fleet(Map.of(),
Map.of("lead-designer", new FleetConfig.Slot("sonnet")), Map.of(), Map.of(), null));
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, members).resolve("127.0.0.1", 42, null);
assertEquals(Role.OBSERVER, p.role());
assertEquals(Role.WORKER, p.role());
assertNull(p.name());
}
@@ -473,11 +455,12 @@ class CallerResolverTest {
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, members);
assertEquals(Role.OBSERVER, r.resolve("127.0.0.1", 42, null).role());
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
assertTrue(members.bind("architect:lead-designer", "term_a")); // the later lifecycle binds the slot
assertEquals(Role.ARCHITECT, r.resolve("127.0.0.1", 42, null).role());
assertEquals("architect:lead-designer", r.members().get("term_a"));
}
@Test
@@ -500,13 +483,12 @@ class CallerResolverTest {
}
@Test
void aBoundNonArchitectSlotResolvesToTheObserverFloorNotArchitect() {
void aBoundNonArchitectSlotStillResolvesAsAWorker() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, boundMembers("dev:builder", MemberRole.DEV))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.OBSERVER, p.role(), "a dev binding must never grant architect rights, "
+ "and this construction path wires no roster to recognise it as the live dev it is");
assertEquals(Role.WORKER, p.role(), "a dev binding must never grant architect rights");
}
@Test
@@ -517,17 +499,17 @@ class CallerResolverTest {
assertThrows(IllegalArgumentException.class, () -> new CallerResolver(id, true, " "));
}
@Test
void aPaneOnAnyLoopbackSourceAddressIsStillAtTheFloorNotThePrimary() {
// fleetd #305: the escalation this guards against. ConnectionIdentity used to accept only
// 127.0.0.1, so a pane connecting from 127.0.0.2 resolved to no terminal, and this
// resolver's own (wider) loopback check then made it the PRIMARY — granting spawn, stop,
// send and drain. Measured on the Linux fleet host: binding a source of 127.0.0.2 succeeds
// there, so the path is real and not theoretical.
void aWorkerOnAnyLoopbackSourceAddressIsStillAWorkerNotThePrimary() {
// fleetd #305: the escalation. ConnectionIdentity used to accept only 127.0.0.1, so a
// worker connecting from 127.0.0.2 resolved to no terminal, and this resolver's own
// (wider) loopback check then made it the PRIMARY — granting spawn, stop, send and drain.
// Measured on the Linux fleet host: binding a source of 127.0.0.2 succeeds there, so the
// path is real and not theoretical.
CallerResolver r = new CallerResolver(workerIdentity(), false, null);
for (String src : new String[]{"127.0.0.1", "127.0.0.2", "127.1.2.3", "::ffff:127.0.0.2"}) {
Principal p = r.resolve(src, 55555, null);
assertEquals(Role.OBSERVER, p.role(), "the pane must stay off PRIMARY from source " + src);
assertEquals("term_a", p.terminal(), "pane terminal from source " + src);
assertEquals(Role.WORKER, p.role(), "a worker must stay a worker from source " + src);
assertEquals("term_a", p.terminal(), "worker terminal from source " + src);
}
}
@@ -541,375 +523,4 @@ class CallerResolverTest {
}
}
// ── fleetd #669 Unit D: a live spawned member outranks every tab map ────────────────────────
/**
* Criterion 1: a terminal present in BOTH the spawned-member roster AND the lead tab map
* resolves as its member role, not as a lead — the roster is checked first, consulting no tab
* map at all when it matches.
*/
@Test
void aSpawnedMemberWinsOverALeadTabForTheSamePane() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
() -> Map.of("term_a", "opus-5.0"), new MemberRegistry(null),
t -> "term_a".equals(t) ? MemberRole.DEV : null, Map::of)
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role(),
"a live spawned member's own identity must win over a tab map naming the same pane a lead");
assertEquals("term_a", p.terminal());
}
/**
* fleetd #702: between {@link SessionManager#release} removing the registry entry and the
* pane actually stopping, a resolve for that terminal must still see the live member and
* never fall through to a tab map — wired through the real {@link SessionManager}, not a
* hand-rolled stand-in for its {@code spawnedMemberRole}.
*
* <p>Reuses the ticket's own test idea: an injected {@code hasUncommitted} resolves the
* releasing terminal from inside {@code release}'s window — a real call landing inside the
* window, so no sleep and no race.
*
* <p>The property asserted is "no tab map is consulted", never "the same role is returned".
* The mandatory control is the lead tab map: it names this exact terminal, and the same
* resolve taken <em>outside</em> the window (before release runs) must still return the lead
* role — without that control, the in-window assertion would also pass on an empty map and
* prove nothing.
*/
@Test
void aPaneMidTeardownResolvesAsItsOwnRoleConsultingNoTabMap() {
FakeHerdr herdr = new FakeHerdr().pinNextStarts(1, "term_a", "w2:p7");
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
AtomicReference<Principal> duringWindow = new AtomicReference<>();
AtomicReference<CallerResolver> resolverRef = new AtomicReference<>();
Worktrees worktrees = new Worktrees() {
@Override
public String add(String repoRoot, String branch, String baseRef) {
return "/wt/" + branch.replace('/', '_');
}
@Override
public void remove(String repoRoot, String worktreePath) {
}
@Override
public void deleteBranch(String repoRoot, String branch) {
}
@Override
public boolean hasUncommitted(String worktreePath) {
// Runs from INSIDE release()'s git-status shell-out: the registry entry is
// already gone, but the pane has not stopped yet.
duringWindow.set(resolverRef.get().resolve("127.0.0.1", 42, null));
return false;
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
}
@Override
public String repoRoot(String cwd) {
return "/repo";
}
@Override
public Optional<String> snapshot(String worktreePath, String branch, String message) {
return Optional.empty();
}
@Override
public WipRefStats wipRefs(String repoRoot) {
return new WipRefStats(0, 0L);
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
return 0;
}
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
}
};
SessionManager sessions = new SessionManager(workers, worktrees);
// The lead tab map names "term_a" before anything is ever spawned onto it — the control
// this test needs. Built up front so the SAME resolver answers every resolve() call below.
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> FakeHerdr.WORKER_PID);
CallerResolver resolver = CallerResolver.withLeadsAndMembers(identity, false, null,
() -> Map.of("term_a", "the-lead"), null, sessions::spawnedMemberRole, Map::of);
resolverRef.set(resolver);
Principal before = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, before.role(),
"control: with no live or releasing member on this terminal, the lead tab map must "
+ "win — this is what proves the in-window assertion below is not passing on "
+ "an empty map");
assertEquals("the-lead", before.name());
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("fleetd-702", null));
assertEquals("term_a", s.terminalId(), "sanity: the spawn resolved to the pinned pane");
sessions.release(s.paneId());
assertNotNull(duringWindow.get(), "the dirty check must have run and captured a resolve");
assertEquals(Role.WORKER, duringWindow.get().role(),
"inside the window the pane must resolve as its own live-member role, consulting no "
+ "tab map — a lead tab naming the same terminal must not win");
assertEquals("term_a", duringWindow.get().terminal());
}
/**
* {@link SessionManager#releaseRemoved} unbinds the architect slot before the git-status
* shell-out that opens the teardown window, so a resolve landing inside that window must see
* the slot already unbound and resolve {@link Role#WORKER} — never {@link Role#ARCHITECT},
* and never by checking role equality against the live session, which would hold even if the
* unbind ran too late.
*
* <p>The control is the same resolve taken outside the window, while the slot is still bound,
* which must return {@link Role#ARCHITECT} — without it this test would also pass against a
* slot that was never bound, and prove nothing.
*/
@Test
void aReleasingArchitectIsDemotedToWorkerInsideTheTeardownWindow() {
FakeHerdr herdr = new FakeHerdr().pinNextStarts(1, "term_a", "w2:p7");
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
AtomicReference<Principal> duringWindow = new AtomicReference<>();
AtomicReference<CallerResolver> resolverRef = new AtomicReference<>();
Worktrees worktrees = new Worktrees() {
@Override
public String add(String repoRoot, String branch, String baseRef) {
return "/wt/" + branch.replace('/', '_');
}
@Override
public void remove(String repoRoot, String worktreePath) {
}
@Override
public void deleteBranch(String repoRoot, String branch) {
}
@Override
public boolean hasUncommitted(String worktreePath) {
// Runs from INSIDE release()'s git-status shell-out: the architect slot is already
// unbound by this point, but the pane has not stopped yet.
duringWindow.set(resolverRef.get().resolve("127.0.0.1", 42, null));
return false;
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
}
@Override
public String repoRoot(String cwd) {
return "/repo";
}
@Override
public Optional<String> snapshot(String worktreePath, String branch, String message) {
return Optional.empty();
}
@Override
public WipRefStats wipRefs(String repoRoot) {
return new WipRefStats(0, 0L);
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
return 0;
}
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
}
};
SessionManager sessions = new SessionManager(workers, worktrees);
MemberRegistry members = new MemberRegistry(new FleetConfig.Fleet(Map.of(),
Map.of("lead-designer", new FleetConfig.Slot("ltms-local")), Map.of(), Map.of(), null));
sessions.setMemberLifecycle(members);
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> FakeHerdr.WORKER_PID);
CallerResolver resolver = CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, members, sessions::spawnedMemberRole, Map::of);
resolverRef.set(resolver);
MemberSession s = sessions.acquire("ltms-local", MemberRole.ARCHITECT, null, "/caller/proj", null,
new WorktreeRequest("fleetd-702d", null));
assertEquals("term_a", s.terminalId(), "sanity: the spawn resolved to the pinned pane");
assertEquals(MemberRole.ARCHITECT, s.role(), "sanity: the slot bind succeeded");
Principal before = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, before.role(),
"control: with the slot still bound, the pane must resolve as an architect — this "
+ "is what proves the in-window assertion below is not passing against a "
+ "slot that was never bound");
assertEquals("lead-designer", before.name());
sessions.release(s.paneId());
assertNotNull(duringWindow.get(), "the dirty check must have run and captured a resolve");
assertEquals(Role.WORKER, duringWindow.get().role(),
"inside the window the architect slot is already unbound, so the result must be "
+ "WORKER — asserting role-equality with the live session here would tempt "
+ "moving the unbind earlier or later, which would be wrong either way");
assertEquals("term_a", duringWindow.get().terminal());
}
/** A spawned architect in the roster resolves ARCHITECT, carrying its bound slot's name. */
@Test
void aSpawnedArchitectInTheRosterResolvesArchitectWithItsSlotName() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
boundMembers("architect:lead-designer", MemberRole.ARCHITECT),
t -> "term_a".equals(t) ? MemberRole.ARCHITECT : null, Map::of)
.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, p.role());
assertEquals("lead-designer", p.name());
assertEquals("term_a", p.terminal());
}
/**
* fleetd #424 regression: the roster only answers THAT a pane is a live spawned member; config
* still decides WHAT that member's slot grants. A slot revoked after the bind must still demote
* the session on its very next request, exactly as it would for a pane with no roster entry at
* all — the roster's own ARCHITECT role must never be granted on its word alone.
*
* <p>{@code bind} refuses an unconfigured slot, so the revoked state can only be reached by
* binding while the slot is configured and then swapping the config out from under it, the way
* a live reload does.
*/
@Test
void aRevokedArchitectSlotDemotesALiveSpawnedArchitectToWorker() {
FleetConfig.Fleet configured = new FleetConfig.Fleet(Map.of(),
Map.of("lead-designer", new FleetConfig.Slot("sonnet")), Map.of(), Map.of(), null);
java.util.concurrent.atomic.AtomicReference<FleetConfig.Fleet> live =
new java.util.concurrent.atomic.AtomicReference<>(configured);
MemberRegistry members = MemberRegistry.live(live::get);
assertTrue(members.bind("architect:lead-designer", "term_a"));
live.set(new FleetConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(), null)); // slot revoked
// Setup controls: the slot is really gone from config, but the occupancy is still there —
// otherwise this test would pass for the wrong reason.
assertNull(members.roleForSlot("architect:lead-designer"), "setup control: the slot must be gone from config");
assertEquals("architect:lead-designer", members.snapshot().get("term_a"),
"setup control: the binding itself must still be there");
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
members, t -> "term_a".equals(t) ? MemberRole.ARCHITECT : null, Map::of)
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role(),
"a revoked slot must demote a live spawned architect on its very next request");
}
/** Criterion 3: a configured collaborator tab that is not a spawned member resolves COLLABORATOR. */
@Test
void aConfiguredCollaboratorTabResolvesToCollaboratorCarryingItsName() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
new MemberRegistry(null), t -> null, () -> Map.of("term_a", "ops"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.COLLABORATOR, p.role());
assertEquals("ops", p.name());
assertEquals("term_a", p.terminal());
}
/** Regression: an empty collaborator registry leaves every pane at the unconfigured-pane floor. */
@Test
void anEmptyCollaboratorRegistryLeavesEveryPaneAtTheObserverFloor() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
new MemberRegistry(null), t -> null, Map::of)
.resolve("127.0.0.1", 42, null);
assertEquals(Role.OBSERVER, p.role());
assertNull(p.name());
}
@Test
void describeNamesTheCollaborator() {
assertEquals("collaborator:ops", Principal.collaborator("ops", "term_a", 1).describe());
}
// ── fleetd #669 Unit D: knownLeadOrCollaborator() reads the same maps resolve() does ───────────
@Test
void knownLeadOrCollaboratorIsTrueForALeadTerminal() {
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
() -> Map.of("term_lead", "opus-5.0"), new MemberRegistry(null), t -> null, Map::of);
assertTrue(r.knownLeadOrCollaborator().test("term_lead"));
assertFalse(r.knownLeadOrCollaborator().test("term_other"));
}
@Test
void knownLeadOrCollaboratorIsTrueForACollaboratorTerminal() {
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
new MemberRegistry(null), t -> null, () -> Map.of("term_collab", "ops"));
assertTrue(r.knownLeadOrCollaborator().test("term_collab"));
assertFalse(r.knownLeadOrCollaborator().test("term_other"));
}
// ── fleetd #705: narrowing the unconfigured-pane floor to OBSERVER ──────────────────────────
/**
* The case this ticket exists for: a pane the resolver cannot place as a live spawned member,
* a lead, a bound architect slot, or a configured collaborator must land on the narrow
* {@link Role#OBSERVER} floor, never the {@link Role#WORKER} the old fallback granted.
*
* <p>The second assertion is the control the ticket requires: a terminal the roster DOES
* recognise as a live spawned member must still resolve its own role. Without it, this test
* would also pass if the fix accidentally turned every caller into an observer.
*/
@Test
void anUnconfiguredPaneResolvesObserverButARegisteredMemberStillResolvesItsOwnRole() {
Principal unconfigured = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
assertEquals(Role.OBSERVER, unconfigured.role(),
"a pane matching none of the configured or live-roster roles must fall to the "
+ "floor, not WORKER");
assertEquals("term_a", unconfigured.terminal());
Principal registered = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, new MemberRegistry(null),
t -> "term_a".equals(t) ? MemberRole.DEV : null, Map::of)
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, registered.role(),
"control: a live spawned member must keep resolving its own role, never the "
+ "unconfigured-pane floor");
}
@Test
void describeNamesTheObserverByItsPane() {
assertEquals("observer:term_a", Principal.observer("term_a", 1).describe());
}
@Test
void knownLeadOrCollaboratorIsFalseForASpawnedMembersTerminal() {
// The exact scenario a collaborator's SEND must never reach: a live spawned member's own
// terminal, which is neither a configured lead nor a configured collaborator.
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
new MemberRegistry(null), t -> "term_a".equals(t) ? MemberRole.DEV : null, Map::of);
assertFalse(r.knownLeadOrCollaborator().test("term_a"));
}
}
@@ -295,10 +295,9 @@ class MemberRegistryLiveTest {
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
Principal after = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.OBSERVER, after.role(),
"removing the slot from config must demote the bound session on its NEXT request — "
+ "this harness wires no live roster for term_a, so the demotion lands on "
+ "the unconfigured-pane floor");
assertEquals(Role.WORKER, after.role(),
"removing the slot from config must demote the bound session to worker on its "
+ "NEXT request — this is the ticket's whole point");
assertEquals("term_a", after.terminal(), "same pane, same terminal — only the role changed");
}
@@ -1,31 +0,0 @@
package dev.ltms.fleet.auth;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
class PrincipalTest {
@Test
void ownerKeyCoversEveryRole() {
assertEquals("leader:opus", Principal.leader("opus", "term_lead", 1).ownerKey());
assertNull(Principal.primary(2).ownerKey());
assertEquals("worker:term_worker", Principal.worker("term_worker", 3).ownerKey());
assertEquals("architect:term_arch", Principal.architect("opus", "term_arch", 4).ownerKey());
assertEquals("collaborator:ops", Principal.collaborator("ops", "term_collab", 5).ownerKey());
assertEquals("observer:term_observer", Principal.observer("term_observer", 6).ownerKey());
assertEquals("anonymous", Principal.anonymous().ownerKey());
}
@Test
void rolePrefixesKeepLeadAndArchitectKeysDistinct() {
String lead = Principal.leader("opus", "term_lead", 1).ownerKey();
String architect = Principal.architect("design", "opus", 2).ownerKey();
assertEquals("leader:opus", lead);
assertEquals("architect:opus", architect);
assertNotEquals(lead, architect);
}
}
@@ -107,8 +107,6 @@ class ConfigRefTest {
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty());
assertTrue(out.split().stream().noneMatch(s -> s.contains("fleet.collaborators")),
out.split().toString());
assertEquals("new charter", ref.get().fleet().charterFor(
dev.ltms.fleet.peer.MemberRole.ARCHITECT));
}
@@ -788,100 +786,14 @@ class ConfigRefTest {
assertTrue(out.split().getFirst().startsWith("fleet:"), out.split().toString());
assertTrue(out.split().getFirst().contains("restart"), out.split().toString());
assertTrue(out.split().getFirst().contains("live"), out.split().toString());
assertTrue(out.split().stream().noneMatch(s -> s.contains("fleet.collaborators")),
out.split().toString());
assertTrue(out.summary().contains("partially live"), out.summary());
// The snapshot still carries the new value — the LeadTabScanner's identity map and
// LeadLauncher's auto-launch are what wait for a restart; a reload rebuilds neither.
assertEquals("lead: opus-b", ref.get().fleet().leaders().get("opus").tab());
}
@Test
void addingAFleetCollaboratorIsReportedAsSplit(@TempDir Path dir) throws Exception {
assertCollaboratorChangeIsReported(dir, """
fleet:
collaborators:
alex:
tab: "collaborator: alex"
""", "add a collaborator");
}
@Test
void removingAFleetCollaboratorIsReportedAsSplit(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
fleet:
collaborators:
alex:
tab: "collaborator: alex"
"""));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("fleet: {}\n"));
assertCollaboratorSplit(ref.reload(), "remove a collaborator");
}
@Test
void changingAFleetCollaboratorTabIsReportedAsSplit(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
fleet:
collaborators:
alex:
tab: "collaborator: alex-a"
"""));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("""
fleet:
collaborators:
alex:
tab: "collaborator: alex-b"
"""));
assertCollaboratorSplit(ref.reload(), "change a collaborator tab");
}
@Test
void unchangedFleetCollaboratorsProduceNoSplitReport(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
String config = yaml("""
fleet:
collaborators:
alex:
tab: "collaborator: alex"
""");
Files.writeString(f, config);
ConfigRef ref = refFor(f);
Files.writeString(f, config);
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.split().isEmpty(), out.split().toString());
}
private static void assertCollaboratorChangeIsReported(Path dir, String changed, String action)
throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("fleet: {}\n"));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml(changed));
assertCollaboratorSplit(ref.reload(), action);
}
private static void assertCollaboratorSplit(ConfigRef.Outcome out, String action) {
assertTrue(out.applied(), action);
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
assertEquals(1, out.split().size(), out.split().toString());
String report = out.split().getFirst();
assertTrue(report.startsWith("fleet: fleet.collaborators"), report);
assertTrue(report.contains("read once at startup"), report);
assertTrue(report.contains("restart"), report);
}
/**
* fleetd #333: {@code fleet.leaders} is a frozen part of {@code fleet:}. A reload that
* fleetd #333: {@code fleet.leaders} is the ONLY frozen part of {@code fleet:}. A reload that
* changes {@code tabLabel} (or charters, or a role pool) without touching {@code fleet.leaders}
* must stay fully hot with nothing reported — proving {@link ConfigRef#changedSplitKeys}
* compares {@code fleet.leaders} specifically rather than the whole {@code Fleet} record, which
@@ -676,8 +676,12 @@ class FleetConfigTest {
assertEquals(5, hb.quietNudgeCap());
}
/**
* The hazard the guard exists for: fleetd writes worker tab labels and reads lead tab labels.
* Overlap the two and every worker it spawns is read back as a lead.
*/
@Test
void aProfileTabLabelOverrideMatchingALeadTabRefusesToStart(@TempDir Path dir)
void aLeadPrefixThatAProfileTabLabelOverrideAlsoMatchesRefusesToStart(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("collide.yaml");
Files.writeString(f, """
@@ -685,33 +689,32 @@ class FleetConfigTest {
port: 8080
profiles:
gx10:
tabLabel: "alpha"
tabLabel: "lead: {profile} #{n}"
fleet:
leaders:
opus:
tab: "alpha"
tab: "lead: opus"
tabPrefix: "lead:"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
assertTrue(e.getMessage().contains("alpha"), "the message must name the offending label");
}
/** A bad fleet-wide template promotes every member, not one profile — so it is checked too. */
@Test
void aFleetTabLabelTemplateThatCanRenderAsALeadTabRefusesToStart(@TempDir Path dir) throws Exception {
void aFleetTabLabelThatMatchesALeadPrefixRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("collide-template.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
pha: {}
fleet:
tabLabel: "al{profile}"
tabLabel: "lead: {role} {profile}"
leaders:
opus:
tab: "alpha"
tab: "lead: opus"
""");
FleetConfig cfg = FleetConfig.load(f);
@@ -720,28 +723,12 @@ class FleetConfigTest {
assertTrue(e.getMessage().contains("fleet.tabLabel"));
}
/**
* The point of making role the label's first field: {@code {role}} comes from a closed enum, so
* a generated label cannot begin with {@code "lead:"} however the fleet is configured.
*/
@Test
void anExactFleetTabLabelCollisionRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("exact-tab-collision.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
tabLabel: "alpha"
leaders:
alpha:
tab: "alpha"
""");
IllegalStateException e = assertThrows(IllegalStateException.class,
() -> FleetConfig.load(f).validateAll());
assertTrue(e.getMessage().contains("fleet.tabLabel"),
"the message must name the offending label");
assertTrue(e.getMessage().contains("alpha"), "the message must name the colliding lead tab");
}
@Test
void aFleetTabLabelTemplateThatCannotRenderAsALeadTabIsAllowed(@TempDir Path dir) throws Exception {
void theDefaultTabLabelCannotCollideWithTheDefaultLeadPrefix(@TempDir Path dir) throws Exception {
Path f = dir.resolve("ok.yaml");
Files.writeString(f, """
bind:
@@ -750,13 +737,17 @@ class FleetConfigTest {
gx10:
baseUrl: http://gx00.gw:8000
fleet:
tabLabel: "worker-{profile}"
leaders:
opus:
tab: "alpha"
tab: "lead: opus"
""");
assertDoesNotThrow(() -> FleetConfig.load(f).validateAll());
assertDoesNotThrow(() -> FleetConfig.load(f).validateLeadTabPrefixes());
for (MemberRole role : MemberRole.values()) {
assertFalse(FleetConfig.Fleet.DEFAULT_TAB_LABEL
.replace("{role}", role.wireName()).startsWith("lead:"),
"no role renders a label that reads as a lead");
}
}
@Test
@@ -774,165 +765,6 @@ class FleetConfigTest {
"a label that collides with a convention nobody reads is not a problem");
}
/**
* fleetd #677: identity is matched on a lead's exact {@code tab} alone, so two leads sharing
* one tab means only one of them is ever found — the guard must catch this independently of
* the member-template checks above.
*/
@Test
void twoLeadsSharingTheSameExactTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("shared-tab.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "shared tab"
sonnet:
tab: "shared tab"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("opus"), "the message must name one offending lead");
assertTrue(e.getMessage().contains("sonnet"), "the message must name the other offending lead");
}
/**
* fleetd #693: the guard matches tabs case-insensitively, because
* {@code LeadTabScanner} keys its tab map on a lowercased label — two tabs differing only in
* case collide there too, and the guard must catch that independently of the exact-match case
* above.
*/
@Test
void twoLeadsSharingTheSameTabInDifferentCaseRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("shared-tab-case.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "Shared Tab"
sonnet:
tab: "shared tab"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("opus"), "the message must name one offending lead");
assertTrue(e.getMessage().contains("sonnet"), "the message must name the other offending lead");
}
/** Control for {@link #twoLeadsSharingTheSameExactTabRefusesToStart}: distinct tabs load cleanly. */
@Test
void twoLeadsWithDistinctExactTabsAreAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("distinct-tabs.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "opus tab"
sonnet:
tab: "sonnet tab"
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateLeadTabPrefixes);
}
// ── validatePanePlacementAgainstLeadTabs ────────────────────────────────────────────────────
/**
* The hazard this guard closes: a pane-placed member lands inside the focused tab rather than
* its own, so it can land inside a lead's labelled tab and, while its pane carries no entry in
* the spawned-member roster, be read back as that lead.
*/
@Test
void aPanePlacedProfileWithALeadTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-hazard.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
tab: "lead: opus"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class,
cfg::validatePanePlacementAgainstLeadTabs);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
}
@Test
void aPanePlacedProfileWithNoLeadTabIsAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-no-tab.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
profile: gx10
""");
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
"a leader with no tab feeds nothing into the scanner, so pane placement is safe");
}
@Test
void aTabPlacedProfileWithALeadTabIsAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("tab-safe.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: tab
fleet:
leaders:
opus:
tab: "lead: opus"
""");
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
"a member in its own tab cannot land inside a lead's tab");
}
/** Proves the reflective sweep behind {@code validateAll} really reaches this validator. */
@Test
void validateAllAlsoRefusesPanePlacementAgainstALeadTab(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-hazard-sweep.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
tab: "lead: opus"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAll);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
}
// ── CB-530/CB-579: the leaders registry ─────────────────────────────────────────────────────
@Test
@@ -1001,274 +833,6 @@ class FleetConfigTest {
"no primary.terminal pin ⇒ nothing registered, even with fleet.leaders configured");
}
// ── fleetd #669: the collaborators registry ────────────────────────────────────────────────
/**
* {@code Fleet} is {@code @JsonIgnoreProperties(ignoreUnknown = true)}, so a config naming
* {@code fleet.collaborators.<name>.tab} loads with no exception whether or not the key is
* ever read into the object model. Asserting only "no exception" would pass both before and
* after the real fix, so this asserts the parsed value is actually reachable from the loaded
* {@code FleetConfig} — the one thing a vacuous "no exception" test cannot tell apart.
*/
@Test
void collaboratorsBlockIsActuallyParsedNotSilentlyDropped(@TempDir Path dir) throws Exception {
Path f = dir.resolve("collaborators.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
collaborators:
reviewer-alex:
tab: "collab: alex"
""");
FleetConfig cfg = FleetConfig.load(f);
assertEquals("collab: alex", cfg.fleet().collaborators().get("reviewer-alex").tab());
}
/** A collaborator carries no field other than {@code tab}, so a blank one is meaningless. */
@Test
void aCollaboratorWithNoTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("useless-collaborator.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
collaborators:
ghost: {}
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateMembers);
assertTrue(e.getMessage().contains("ghost"), "the message must name the useless entry");
assertTrue(e.getMessage().contains("tab:"), "the message must say what is missing");
}
/** Control for {@link #aCollaboratorWithNoTabRefusesToStart}: a named tab loads cleanly. */
@Test
void aCollaboratorWithATabIsAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("named-collaborator.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
collaborators:
reviewer-alex:
tab: "collab: alex"
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateMembers);
}
/**
* fleetd #669: the fleet-wide {@code tabLabel} template can render as a collaborator tab, the
* same hazard {@link #aFleetTabLabelTemplateThatCanRenderAsALeadTabRefusesToStart} covers on
* the lead side. Drives the fleet-wide branch directly, with no profile override involved.
*/
@Test
void aFleetTabLabelTemplateThatCanRenderAsACollaboratorTabRefusesToStart(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("collide-template-collaborator.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
tabLabel: "al{profile}"
collaborators:
alex:
tab: "alpha"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("fleet.tabLabel"));
assertTrue(e.getMessage().contains("alex"), "the message must name the offending collaborator");
}
/**
* A member tabLabel that can render as a configured collaborator tab is the same hazard as the
* lead case above — while its pane carries no entry in the spawned-member roster, a member
* labelled that way is read back as the collaborator.
*/
@Test
void aProfileTabLabelOverrideMatchingACollaboratorTabRefusesToStart(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("collide-collaborator.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
tabLabel: "collab-tab"
fleet:
collaborators:
alex:
tab: "collab-tab"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
assertTrue(e.getMessage().contains("collab-tab"), "the message must name the offending label");
}
/** Control: a profile tabLabel that cannot render as the collaborator tab is allowed. */
@Test
void aProfileTabLabelThatCannotRenderAsACollaboratorTabIsAllowed(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("ok-collaborator.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
tabLabel: "worker-{profile}"
fleet:
collaborators:
alex:
tab: "collab-tab"
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateLeadTabPrefixes);
}
/** fleetd #669: identity is matched on a collaborator's exact tab, so two sharing one are unreachable. */
@Test
void twoCollaboratorsSharingTheSameExactTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("shared-collaborator-tab.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
collaborators:
alex:
tab: "shared tab"
sam:
tab: "Shared Tab"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("alex"), "the message must name one offending collaborator");
assertTrue(e.getMessage().contains("sam"), "the message must name the other offending collaborator");
}
/** Control for {@link #twoCollaboratorsSharingTheSameExactTabRefusesToStart}: distinct tabs load cleanly. */
@Test
void twoCollaboratorsWithDistinctExactTabsAreAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("distinct-collaborator-tabs.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
collaborators:
alex:
tab: "alex tab"
sam:
tab: "sam tab"
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateLeadTabPrefixes);
}
/**
* fleetd #669: a collaborator tab equal to a lead tab crosses a privilege boundary — the worst
* of the three new collisions, since only one of the two identities is ever found.
*/
@Test
void aCollaboratorTabEqualToALeadTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("lead-collaborator-collision.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "shared tab"
collaborators:
alex:
tab: "Shared Tab"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("opus"), "the message must name the offending lead");
assertTrue(e.getMessage().contains("alex"), "the message must name the offending collaborator");
}
/** Control for {@link #aCollaboratorTabEqualToALeadTabRefusesToStart}: distinct tabs load cleanly. */
@Test
void aLeadAndACollaboratorWithDistinctTabsAreAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("lead-collaborator-ok.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "lead tab"
collaborators:
alex:
tab: "collab tab"
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateLeadTabPrefixes);
}
/**
* fleetd #669: a pane-placed member can land in a collaborator's labelled tab exactly as it
* can land in a lead's — {@code validatePanePlacementAgainstLeadTabs()} must fire even when
* {@code fleet.leaders} is empty, which is the early-return the brief flagged as the bug.
*/
@Test
void aPanePlacedProfileWithACollaboratorTabRefusesToStartEvenWithNoLeaders(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("pane-hazard-collaborator.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
collaborators:
alex:
tab: "collab: alex"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class,
cfg::validatePanePlacementAgainstLeadTabs);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
}
/** Control: a pane-placed profile with no lead or collaborator tab configured is allowed. */
@Test
void aPanePlacedProfileWithNoLeaderOrCollaboratorTabIsAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-no-tab-at-all.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
collaborators:
alex: {}
""");
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
"a collaborator with no tab feeds nothing into the scanner, so pane placement is safe");
}
// ── CB-548: the architects registry ────────────────────────────────────────────────────────
@Test
@@ -1541,60 +1105,6 @@ class FleetConfigTest {
assertEquals(Set.of("sonnet"), cfg.fleet().pool(MemberRole.REVIEWER).keySet());
}
/**
* A duplicated name in {@code fleet.collaborators} is refused at parse time, like any other
* {@code fleet:} pool. See {@link #duplicateSlotNamesInOnePoolAreRejectedAtParseTime}.
*/
@Test
void duplicateCollaboratorNamesAreRejectedAtParseTime(@TempDir Path dir) throws Exception {
Path f = dir.resolve("collaborator-dup.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
collaborators:
alex:
tab: "collab: alex"
alex:
tab: "collab: alex, second"
""");
IllegalStateException e =
assertThrows(IllegalStateException.class, () -> FleetConfig.load(f));
assertTrue(e.getMessage().contains("alex"),
"the refusal names the duplicated entry, was: " + e.getMessage());
assertTrue(e.getMessage().contains("fleet.collaborators"),
"the refusal names the pool the duplicate is in, was: " + e.getMessage());
}
/**
* Control for {@link #duplicateCollaboratorNamesAreRejectedAtParseTime}: the same name reused
* across the collaborators registry and a member role pool is the role × profile matrix doing
* its job in the other pool, not a mistake — only a repeat within one pool loses an entry.
*/
@Test
void theSameNameInCollaboratorsAndAnotherPoolIsNotADuplicate(@TempDir Path dir) throws Exception {
Path f = dir.resolve("collaborator-cross-pool.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
fleet:
architects:
alex:
profile: sonnet
collaborators:
alex:
tab: "collab: alex"
""");
FleetConfig cfg = assertDoesNotThrow(() -> FleetConfig.load(f));
assertEquals(Set.of("alex"), cfg.fleet().pool(MemberRole.ARCHITECT).keySet());
assertEquals(Set.of("alex"), cfg.fleet().collaborators().keySet());
}
@Test
void duplicateKeysOutsideTheFleetPoolsAreUnaffected(@TempDir Path dir) throws Exception {
// The duplicate check is scoped to the fleet pools — a duplicate elsewhere is not this
@@ -3579,41 +3089,4 @@ class FleetConfigTest {
FleetConfig cfg = FleetConfig.load(f);
assertTrue(cfg.models().offIds().isEmpty());
}
// ── fleetd #651: leadRollover.turnSettleSeconds default resolution ─────────────────────────
@Test
void turnSettleSecondsDefaultsTo300WhenUnset(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bare-rollover.yaml");
Files.writeString(f, "bind:\n port: 8080\nleadRollover: {}\n");
FleetConfig.LeadRollover rollover = FleetConfig.load(f).leadRollover();
assertNotNull(rollover);
assertEquals(300, rollover.turnSettleSeconds());
}
@Test
void turnSettleSecondsUsesAnExplicitPositiveValue(@TempDir Path dir) throws Exception {
Path f = dir.resolve("rollover.yaml");
Files.writeString(f, """
bind:
port: 8080
leadRollover:
turnSettleSeconds: 45
""");
FleetConfig.LeadRollover rollover = FleetConfig.load(f).leadRollover();
assertEquals(45, rollover.turnSettleSeconds());
}
@Test
void turnSettleSecondsFallsBackTo300WhenZeroOrNegative(@TempDir Path dir) throws Exception {
Path zero = dir.resolve("zero.yaml");
Files.writeString(zero, "bind:\n port: 8080\nleadRollover:\n turnSettleSeconds: 0\n");
assertEquals(300, FleetConfig.load(zero).leadRollover().turnSettleSeconds());
Path negative = dir.resolve("negative.yaml");
Files.writeString(negative, "bind:\n port: 8080\nleadRollover:\n turnSettleSeconds: -5\n");
assertEquals(300, FleetConfig.load(negative).leadRollover().turnSettleSeconds());
}
}
@@ -17,21 +17,62 @@ import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Tests the reflective validator sweep and {@link FleetConfig#validateAll()} reachability.
* The gap this class exists to close: mutation testing on the fleetd ticket "central allow-list
* of usable models" found that although {@link FleetConfig#validateModels()}'s own logic was well
* pinned, nothing proved either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()})
* still invoked it — deleting the call site left the full suite green (1478/0/0/0). A follow-up
* measurement (same technique — remove one call site, run the suite, not read the code) found the
* SAME gap for all five of {@link FleetConfig}'s other validators at startup, and for four of the
* six inside {@link ConfigRef#reload()}. This is a class of gap, not one line's mistake: every one
* of those thirteen tests called the validator itself directly, never the real caller that was
* supposed to.
*
* <p>{@link #theSweepRunsEveryValidateMethodOnAnUnrelatedClass()} and its neighbours
* prove that {@link FleetConfig#invokeAllValidators} runs each public, no-arg, void
* {@code validateXxx()} method on its target. {@link #fleetConfigDeclaresExactlyTheseValidatorsToday()}
* is the canary for the validator set. {@link #validateAllReachesEveryOneOfTodaysRealValidators()}
* is the reachability check for that set.
* <p>The fix replaces the six individual {@code cfg.validateXxx()} calls at each of the two real
* call sites with one {@link FleetConfig#validateAll()}, which reaches every validator by
* reflection rather than by a hand-maintained list of names. A hand-maintained list of six names
* would have exactly the defect it replaces: the seventh validator someone adds next month has no
* reason to be added to it, and nothing would say so. This class proves TWO separate claims, and
* keeps them separate on purpose:
*
* <ol>
* <li>{@link #theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames()} and its neighbours
* prove the reflective sweep itself ({@link FleetConfig#invokeAllValidators}) is a general
* mechanism — it runs whatever public, no-arg, void {@code validateXxx()} methods a class
* happens to declare today, including a class with more of them than {@link FleetConfig}
* has right now. This is the proof that a future, real seventh validator on {@link
* FleetConfig} would be swept automatically, without needing to add a real (unwanted)
* seventh validator just to exercise the claim.</li>
* <li>{@link #validateAllReachesEveryOneOfTodaysSixValidators()} proves {@link
* FleetConfig#validateAll()} itself is wired to that same generic mechanism and genuinely
* reaches each of today's six real validators — reusing the exact minimal failing
* configurations {@code FleetConfigTest} already established for each one directly, so a
* single call to {@code validateAll()} is shown to reproduce every one of those six
* failures.</li>
* </ol>
*
* <p>Together with the direct-{@code Fleetd.main}-invocation tests in {@code
* FleetdStartupValidationTest} (which prove the real startup call site still calls {@code
* validateAll()}) and the {@code ConfigRefTest} reload tests (which prove the same for {@link
* ConfigRef#reload()}), removing {@code cfg.validateAll();} from either real call site now fails
* a test in this module.
*
* <p><b>What is NOT pinned, measured rather than assumed.</b> Reverting {@link
* FleetConfig#validateAll()} to a hardcoded list of today's six method calls leaves the whole
* suite green (measured at review: 1491 tests, 0 failures). Nothing ties {@code validateAll()} to
* the generic sweep — claim 1 proves {@link FleetConfig#invokeAllValidators} is generic, and claim
* 2 proves {@code validateAll()} reaches today's six, and a hardcoded list satisfies both. So the
* reflective sweep is a convenience, not the guarantee. The guarantee is {@link
* #fleetConfigDeclaresExactlyTheseSixValidatorsToday()}: it fails the moment a seventh validator
* is declared, which forces whoever adds it to look at this file.
*/
class FleetConfigValidateAllTest {
// ── Claim 1: the reflective sweep is a general mechanism ─────────────────────────────────────
// ── Claim 1: the reflective sweep is a general mechanism, not six names in disguise ──────────
/**
* Fixture with public, no-arg, void methods named {@code validateXxx}. It proves the sweep uses
* the target's method shape rather than special handling for {@link FleetConfig}.
* A throwaway fixture class, unrelated to {@link FleetConfig} in every way except shape: three
* public, no-arg, void methods named {@code validateXxx}. Proves the sweep works on ANY class
* with this shape, not on something special-cased to {@link FleetConfig}.
*/
static class ThreeValidators {
final List<String> ran = new ArrayList<>();
@@ -50,7 +91,7 @@ class FleetConfigValidateAllTest {
}
@Test
void theSweepRunsEveryValidateMethodOnAnUnrelatedClass() {
void theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames() {
ThreeValidators target = new ThreeValidators();
FleetConfig.invokeAllValidators(target);
assertEquals(List.of("validateAlpha", "validateBeta", "validateGamma"), target.ran,
@@ -60,8 +101,12 @@ class FleetConfigValidateAllTest {
}
/**
* Fixture with an added valid method. It proves the sweep reaches a method because it matches
* the validator shape.
* The core of the "self-maintaining" requirement: the exact same class shape as {@link
* ThreeValidators}, plus one more method — standing in for "a developer adds a validator next
* month". Nothing about the sweep changes to pick it up; the new method is invoked purely
* because it exists and matches the shape. This is what makes adding a seventh real validator
* to {@link FleetConfig} safe without touching {@link FleetConfig#validateAll()} or either
* call site — there is no "wire it in" step left to forget.
*/
static class FourValidators {
final List<String> ran = new ArrayList<>();
@@ -89,7 +134,7 @@ class FleetConfigValidateAllTest {
FleetConfig.invokeAllValidators(target);
assertEquals(List.of("validateAlpha", "validateBeta", "validateDelta", "validateGamma"),
sorted(target.ran),
"the added method must be reached automatically — proving a class can grow the "
"the fourth method must be reached automatically — proving a class can grow the "
+ "set of things it validates with no change to the sweep itself");
}
@@ -164,24 +209,18 @@ class FleetConfigValidateAllTest {
+ "name) must all be skipped");
}
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches every validator ──
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches all six today ──
/**
* Reflectively enumerates {@link FleetConfig}'s own public, no-arg, void {@code validateXxx()}
* methods (excluding {@code validateAll} itself) — the exact same filter {@link
* FleetConfig#invokeAllValidators} applies. This is not the mechanism proof (that is claim 1,
* above, on an unrelated class) — it is a visible denominator: today there are eight, named in
* the {@code Set.of} below, so a reader adding or removing one sees this assertion name the new
* count rather than a silent pass at the old one. The count lives only in that set, not in this
* method's name, so the two cannot drift apart.
*
* <p>This assertion alone proves only that the validator exists with the right shape — it
* cannot prove {@code validateAll()} actually reaches it. Only {@link
* #validateAllReachesEveryOneOfTodaysRealValidators()} proves reachability, which is why this
* method's failure message sends the reader there too.
* above, on an unrelated class) — it is a visible denominator: today there are six, named
* here, so a reader adding a seventh sees this assertion name the new count rather than a
* silent pass at the old one.
*/
@Test
void fleetConfigDeclaresExactlyTheseValidatorsToday() {
void fleetConfigDeclaresExactlyTheseSixValidatorsToday() {
Set<String> names = new TreeSet<>();
for (Method m : FleetConfig.class.getMethods()) {
if (java.lang.reflect.Modifier.isPublic(m.getModifiers())
@@ -194,18 +233,12 @@ class FleetConfigValidateAllTest {
}
assertEquals(new TreeSet<>(Set.of("validateAuthExposure", "validateLeadTabPrefixes",
"validateSubscriptionProfiles", "validateCharters", "validateMembers",
"validateModels", "validateLeadRollover", "validatePanePlacementAgainstLeadTabs")),
names,
"FleetConfig's public validate*() methods changed. Do THREE things, in this "
"validateModels", "validateLeadRollover")), names,
"FleetConfig's public validate*() methods changed. Do TWO things, in this "
+ "order. First confirm validateAll() still delegates to "
+ "invokeAllValidators(this) — a hardcoded list there passes every other "
+ "test in this class, so this assertion is the only place that will ever "
+ "make you check. Second, update the expected set below to match. Third, "
+ "add or remove a case for that validator in "
+ "validateAllReachesEveryOneOfTodaysRealValidators() below — this "
+ "assertion proves only that the validator exists with the right shape, "
+ "never that validateAll() reaches it; that enumeration is the test that "
+ "does.");
+ "make you check. Only then update the expected set to match.");
}
/** A minimal, otherwise-valid file — same shape FleetConfigTest and ConfigRefTest use. */
@@ -229,21 +262,14 @@ class FleetConfigValidateAllTest {
}
/**
* The heart of claim 2: for every one of today's real validators, a minimal file that
* fails ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
* directly, or a dedicated minimal fixture where no other test drives that validator through
* {@code validateAll()} — must also fail through {@link FleetConfig#validateAll()}. If a
* future edit to {@code validateAll()} silently dropped one of these from the sweep
* (e.g. a typo'd name filter), exactly one of them would start passing when it must not.
*
* <p>This is the single place that proves {@code validateAll()} reaches a given validator.
* Adding or removing a validator on {@link FleetConfig} must add or remove a case here, not
* only an updated name in {@link #fleetConfigDeclaresExactlyTheseValidatorsToday()}'s expected
* set — that assertion proves the validator's shape, never that {@code validateAll()} reaches
* it.
* The heart of claim 2: for each of today's six real validators, a minimal file that fails
* ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
* directly — must also fail through {@link FleetConfig#validateAll()}. If a future edit to
* {@code validateAll()} silently dropped one validator from the sweep (e.g. a typo'd name
* filter), exactly one of these six would start passing when it must not.
*/
@Test
void validateAllReachesEveryOneOfTodaysRealValidators(@TempDir Path dir) throws Exception {
void validateAllReachesEveryOneOfTodaysSixValidators(@TempDir Path dir) throws Exception {
// validateAuthExposure: a non-loopback bind without token mode.
assertValidateAllRefuses(dir, "auth-exposure.yaml", """
bind:
@@ -251,16 +277,16 @@ class FleetConfigValidateAllTest {
port: 8765
""", "auth.mode: token");
// validateLeadTabPrefixes: a fleet-wide tabLabel that equals a lead tab.
// validateLeadTabPrefixes: a fleet-wide tabLabel that starts with a lead's own tabPrefix.
assertValidateAllRefuses(dir, "lead-tab-prefixes.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
tabLabel: "alpha"
tabLabel: "lead: {role} {profile}"
leaders:
opus:
tab: "alpha"
tab: "lead: opus"
""", "fleet.tabLabel");
// validateSubscriptionProfiles: subscription: true with env: reseating ANTHROPIC_BASE_URL.
@@ -316,29 +342,6 @@ class FleetConfigValidateAllTest {
allow:
- model: claude-sonnet-5
""", "rogue");
// validatePanePlacementAgainstLeadTabs: a pane-placed profile while a lead names a tab.
assertValidateAllRefuses(dir, "pane-placement.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
tab: "lead: opus"
""", "gx10");
// validateLeadRollover: a leadRollover: block present with no handoverPath.
assertValidateAllRefuses(dir, "lead-rollover.yaml", """
bind:
host: 127.0.0.1
port: 8765
leadRollover:
requireOperatorConfirm: false
""", "handoverPath");
}
private static void assertValidateAllRefuses(Path dir, String fileName, String yaml,
@@ -103,7 +103,7 @@ class FleetConfigWithDefaultsPreservesEveryComponentTest {
// comment there), same as broker/primary/leadHeartbeat/... above — a real, non-null value
// here proves it, rather than leaving it null and proving nothing.
v.put("leadRollover", new FleetConfig.LeadRollover(
"/handover/guard.md", true, 3600, 20, 45, "read the handover file"));
"/handover/guard.md", true, 3600, 20, 20, "read the handover file"));
assertNamesMatchComponents(v);
return v;
}
@@ -7,7 +7,6 @@ import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.CopyOnWriteArrayList;
@@ -52,7 +51,6 @@ public final class FakeHerdr implements HerdrClient {
private boolean noPanes = false;
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String agentType = "claude"; // detected agent kind on agent.get; null = undetected
private volatile String agentSessionId = null; // agent_session.value on agent.get; null = omitted
private volatile String readText = "worker transcript tail"; // canned agent.read output
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
private String pinnedStartTerminal;
@@ -60,10 +58,6 @@ public final class FakeHerdr implements HerdrClient {
private Runnable onAgentStart; // fires the instant agent.start is called — see onAgentStart(Runnable)
private volatile int agentGetOkCalls = Integer.MAX_VALUE; // how many agent.get calls succeed first
private volatile String agentGetFailCode = null; // error code every agent.get call after that reports
/** pane ids that {@link #paneGoneAfterClose} has opted into reporting gone — see that method. */
private final Set<String> paneGoneAfterCloseIds = ConcurrentHashMap.newKeySet();
/** pane ids a {@code pane.close} call has actually reached, for {@link #paneGoneAfterCloseIds}. */
private final Set<String> closedPaneIds = ConcurrentHashMap.newKeySet();
public FakeHerdr healthy(boolean h) {
this.healthy = h;
@@ -179,16 +173,6 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Set the {@code agent_session.value} that {@code agent.get} reports for {@code term_a} — the
* default omits the field entirely (herdr not yet having resolved one), matching the real
* daemon's own "not resolved yet" shape.
*/
public FakeHerdr agentSessionId(String sessionId) {
this.agentSessionId = sessionId;
return this;
}
/**
* Make {@code agent.get} succeed normally for its first {@code okCalls} invocations, then fail
* every call after that with {@code code} — fleetd #176 fix 1's "backend exited mid-wait"
@@ -225,19 +209,6 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Make {@code pane.get(paneId)} report the pane gone (a {@code pane_not_found} {@link
* HerdrException}, exactly as {@link WorkspaceControl#locatePane} expects to see once a pane
* has really disappeared) once a {@code pane.close} call for that same {@code paneId} has
* actually reached this fake. Every other pane, and this pane before its own close, keeps
* reporting the default canned {@code pane.get} response — opt-in, by pane id, so no existing
* test's {@code pane.get} behaviour changes.
*/
public FakeHerdr paneGoneAfterClose(String paneId) {
paneGoneAfterCloseIds.add(paneId);
return this;
}
/**
* Run {@code hook} synchronously the instant an {@code agent.start} call reaches this fake —
* i.e. the instant the peer PROCESS would start against a real herdr daemon. A test uses this
@@ -340,12 +311,10 @@ public final class FakeHerdr implements HerdrClient {
}
}
String agentField = agentType == null ? "null" : "\"" + agentType + "\"";
String sessionField = agentSessionId == null ? ""
: ",\"agent_session\":{\"kind\":\"id\",\"value\":\"" + agentSessionId + "\"}";
yield mapper.readTree(("""
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":%s,
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"%s}}""")
.formatted(agentField, agentStatus, sessionField));
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
.formatted(agentField, agentStatus));
}
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
@@ -440,18 +409,9 @@ public final class FakeHerdr implements HerdrClient {
}
yield mapper.readTree("{\"type\":\"ok\"}");
}
case "pane.get" -> {
Object paneIdParam = params instanceof Map<?, ?> m ? m.get("pane_id") : null;
String paneIdKey = paneIdParam == null ? null : String.valueOf(paneIdParam);
if (paneIdKey != null && paneGoneAfterCloseIds.contains(paneIdKey)
&& closedPaneIds.contains(paneIdKey)) {
throw new HerdrException("herdr error [pane_not_found]: pane.get failed",
"pane_not_found", null);
}
yield mapper.readTree("""
case "pane.get" -> mapper.readTree("""
{"type":"pane_info","pane":{"pane_id":"w9:pW","workspace_id":"w9",
"tab_id":"w9:t2","agent_status":"idle"}}""");
}
case "pane.list" -> noPanes
? mapper.readTree("{\"type\":\"pane_list\",\"panes\":[]}")
: mapper.readTree("""
@@ -485,9 +445,6 @@ public final class FakeHerdr implements HerdrClient {
throw new HerdrException("herdr error [" + code + "]: pane.close failed",
code, null);
}
if (paneIdParam != null) {
closedPaneIds.add(String.valueOf(paneIdParam));
}
yield mapper.readTree("{\"type\":\"ok\"}");
}
default -> throw new HerdrException("fake has no canned response for " + method);
@@ -2,8 +2,6 @@ package dev.ltms.fleet.herdr;
import org.junit.jupiter.api.Test;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertNotSame;
@@ -30,44 +28,4 @@ class HerdrRouterTest {
assertSame(router.leadAgents(), router.agentsFor("lead"));
assertSame(router.memberAgents(), router.agentsFor("member"));
}
/**
* fleetd #669 Unit E. The predicate shape here is exactly what {@code FleetdAssembly} builds:
* true when the terminal is a known lead OR a known collaborator. Two distinct clients are
* required — with one shared client {@code agentsFor} would return the same object regardless
* of the predicate's answer, and this assertion would pass whether or not the collaborator map
* was ever consulted.
*/
@Test
void collaboratorTerminalRoutesToTheLeadDaemon() {
FakeHerdr lead = new FakeHerdr();
FakeHerdr member = new FakeHerdr();
Map<String, String> leads = Map.of("term_lead", "primary");
Map<String, String> collaborators = Map.of("term_collab", "reviewer-alex");
HerdrRouter router = new HerdrRouter(lead, member,
id -> leads.containsKey(id) || collaborators.containsKey(id));
assertSame(router.leadAgents(), router.agentsFor("term_collab"),
"a configured collaborator's terminal must route to the LEAD daemon, not the member "
+ "one — its pane is opened by a person, exactly like a lead's");
}
/**
* Companion to {@link #collaboratorTerminalRoutesToTheLeadDaemon}: a terminal that is neither a
* known lead nor a known collaborator must still route to the member daemon. Without this, a
* predicate of {@code _ -> true} would also pass the test above.
*/
@Test
void terminalInNeitherMapStillRoutesToTheMemberDaemon() {
FakeHerdr lead = new FakeHerdr();
FakeHerdr member = new FakeHerdr();
Map<String, String> leads = Map.of("term_lead", "primary");
Map<String, String> collaborators = Map.of("term_collab", "reviewer-alex");
HerdrRouter router = new HerdrRouter(lead, member,
id -> leads.containsKey(id) || collaborators.containsKey(id));
assertSame(router.memberAgents(), router.agentsFor("term_worker"),
"a terminal absent from both maps must stay on the member daemon — the fix widens "
+ "the predicate, it does not make it unconditionally true");
}
}
@@ -163,13 +163,6 @@ class LeadTabScannerTest {
return new LeadTabScanner(herdr, tabToName, Set.of("fleetd-workers"), TTL, clock::get);
}
private LeadTabScanner scannerWithCollaborators(TopologyHerdr herdr, Map<String, String> tabToName,
Map<String, String> collaboratorTabToName,
AtomicLong clock) {
return new LeadTabScanner(herdr, tabToName, collaboratorTabToName,
Set.of("fleetd-workers"), TTL, clock::get);
}
@Test
void everyConfiguredTabBecomesALeadNamedByItsEntry() {
Map<String, String> leads = scanner(twoLeads(), twoLeadsConfigured(), new AtomicLong()).get();
@@ -205,12 +198,9 @@ class LeadTabScannerTest {
}
/**
* Covers the {@code excludedWorkspaceLabels} parameter: a tab in an excluded workspace is never
* matched, whatever its label. Production always constructs this class with an empty set (CB-558,
* {@code FleetdAssembly}), so this parameter plays no part in the live guard against a worker
* tab being mistaken for a lead — that guard is {@code CallerResolver} asking the live
* spawned-member roster before any tab map. This test exists because the parameter still exists
* and is worth covering on its own terms.
* The guard that matters: fleetd labels its own worker tabs, so if a worker space were scanned
* a naming accident would promote the fleet. The exclusion is by workspace, not by hoping the
* worker template never collides.
*/
@Test
void aTabInAWorkerSpaceIsNeverALeadEvenWhenItsLabelMatches() {
@@ -222,29 +212,6 @@ class LeadTabScannerTest {
assertFalse(scanner(herdr, tabToName, new AtomicLong()).get().containsKey("term_impostor"));
}
/**
* The shipped shape (CB-558, {@code FleetdAssembly}): production always constructs this class
* with an empty {@code excludedWorkspaceLabels}, and a lead's {@code workspace:} default is the
* same shared {@code "fleet"} space the members use. A lead tab is still found when it sits in
* the exact same workspace as a member-labelled tab — the scanner tells them apart by the exact
* tab label, not by which workspace either one is in.
*/
@Test
void aLeadIsDiscoveredWhenItsWorkspaceIsTheSameAsTheMemberWorkspace() {
TopologyHerdr herdr = new TopologyHerdr()
.workspace("w1", "fleet")
.tab("w1:t1", "w1", "lead: opus-5.0")
.tab("w1:t2", "w1", "worker: gx10 #1")
.pane("w1:p1", "w1:t1", "term_opus")
.pane("w1:p2", "w1:t2", "term_worker");
LeadTabScanner s = new LeadTabScanner(herdr, Map.of("lead: opus-5.0", "opus-5.0"),
Set.of(), TTL, new AtomicLong()::get);
assertEquals("opus-5.0", s.get().get("term_opus"),
"a lead sharing the members' workspace is still discovered — the label, not the "
+ "workspace, is what matches it");
}
@Test
void aLabelWithNoConfiguredEntryIsIgnored() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
@@ -265,9 +232,8 @@ class LeadTabScannerTest {
@Test
void everyPaneInALeadTabResolvesAsThatLead() {
// A human may split their own lead tab. Both panes are theirs, so both are that lead.
// A pane-placed member landing here instead is refused at startup by
// FleetConfig.validatePanePlacementAgainstLeadTabs, not by this scanner.
// A human may split their own lead tab. Both panes are theirs, so both are that lead —
// nothing fleetd placed can land here (see the worker-space test above).
TopologyHerdr herdr = twoLeads().pane("w1:p1b", "w1:t1", "term_opus_split");
assertEquals("opus-5.0",
@@ -497,75 +463,4 @@ class LeadTabScannerTest {
assertEquals(afterFirst, herdr.calls, "the failure path must be rate-limited too");
}
// ── fleetd #669 Unit D: a collaborator tab is matched the same way as a lead tab, one pass ─────
@Test
void aConfiguredCollaboratorTabIsReportedByCollaboratorsNotByGet() {
TopologyHerdr herdr = new TopologyHerdr()
.workspace("w1", "main")
.tab("w1:t1", "w1", "collab: ops")
.pane("w1:p1", "w1:t1", "term_ops");
LeadTabScanner s = scannerWithCollaborators(herdr, Map.of(), Map.of("collab: ops", "ops"),
new AtomicLong());
assertEquals(Map.of("term_ops", "ops"), s.collaborators(),
"a collaborator tab is matched exactly like a lead tab");
assertEquals(Map.of(), s.get(), "a collaborator tab must never also appear as a lead");
}
/**
* Criterion 4, scanner level: a collaborator tab that is labelled but runs no agent is not
* reported — the same #359 liveness cross-check a lead tab gets.
*/
@Test
void aDeadCollaboratorTabIsNotReported() {
TopologyHerdr herdr = new TopologyHerdr()
.workspace("w1", "main")
.tab("w1:t1", "w1", "collab: ops")
.pane("w1:p1", "w1:t1", "term_ops")
.deadAgent("w1:t1");
LeadTabScanner s = scannerWithCollaborators(herdr, Map.of(), Map.of("collab: ops", "ops"),
new AtomicLong());
assertFalse(s.collaborators().containsKey("term_ops"),
"a dead collaborator tab must never resolve as a live collaborator");
}
/**
* Both kinds are matched in a single pass over the same tab list — not a second scanner, not a
* second scan. Proven by herdr call count: scanning one lead tab and one collaborator tab in the
* same instance costs exactly as many calls as scanning two lead tabs in {@link #twoLeads()}.
*/
@Test
void leadsAndCollaboratorsAreMatchedInOnePassOverTheSameScan() {
TopologyHerdr oneOfEach = new TopologyHerdr()
.workspace("w1", "main")
.tab("w1:t1", "w1", "lead: opus-5.0")
.tab("w1:t2", "w1", "collab: ops")
.pane("w1:p1", "w1:t1", "term_opus")
.pane("w1:p2", "w1:t2", "term_ops");
LeadTabScanner s = scannerWithCollaborators(oneOfEach, Map.of("lead: opus-5.0", "opus-5.0"),
Map.of("collab: ops", "ops"), new AtomicLong());
assertEquals(Map.of("term_opus", "opus-5.0"), s.get());
assertEquals(Map.of("term_ops", "ops"), s.collaborators());
TopologyHerdr twoLeadsBaseline = twoLeads();
scanner(twoLeadsBaseline, twoLeadsConfigured(), new AtomicLong()).get();
assertEquals(twoLeadsBaseline.calls, oneOfEach.calls,
"one lead tab + one collaborator tab must cost exactly as many herdr calls as two "
+ "lead tabs — proof this is one pass, not a second scan");
}
/** Regression: with no collaborators configured, every existing lead-only behaviour is unchanged. */
@Test
void anEmptyCollaboratorMapLeavesCollaboratorsEmptyAndGetUnaffected() {
LeadTabScanner s = scannerWithCollaborators(twoLeads(), twoLeadsConfigured(), Map.of(),
new AtomicLong());
assertEquals(Map.of(), s.collaborators());
assertEquals(Map.of("term_opus", "opus-5.0", "term_gpt", "gpt-sol-5.6"), s.get());
}
}
@@ -81,24 +81,11 @@ class LeadContextGaugeHighThresholdTest {
assertEquals(LeadContextGauge.State.HIGH,
atGauge.read(atConfigDir, SESSION_ID, "claude", null).state(),
"the fixed default must still be 200,000 when no window is resolvable");
}
@Test
@DisplayName("a second read with a different window, within the TTL, reports against its own window, not the first call's cached state")
void aSecondReadWithADifferentWindowWithinTheTtlReportsAgainstItsOwnWindow(@TempDir Path tmp) throws IOException {
String configDir = writeTranscript(tmp, SESSION_ID, 90_000);
long[] now = {0L};
LeadContextGauge gauge = new LeadContextGauge(() -> now[0], 5_000);
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude", 100_000L);
assertEquals(LeadContextGauge.State.HIGH, first.state(),
"90,000 tokens against a 100,000 window is HIGH");
now[0] += 1_000; // stays inside the 5,000ms TTL — the cache key must still vary with the window
LeadContextGauge.Reading second = gauge.read(configDir, SESSION_ID, "claude", 1_000_000L);
assertEquals(LeadContextGauge.State.OK, second.state(),
"90,000 tokens against a 1,000,000 window must report OK regardless of the previous call's "
+ "window, even while that call's cache entry is still within its TTL");
String legacyConfigDir = writeTranscript(tmp.resolve("legacy"), SESSION_ID, 200_000);
LeadContextGauge legacyGauge = new LeadContextGauge();
assertEquals(LeadContextGauge.State.HIGH,
legacyGauge.read(legacyConfigDir, SESSION_ID, "claude").state(),
"the 3-arg read() (no window argument at all) must behave exactly like passing a null window");
}
}
@@ -85,7 +85,7 @@ class LeadContextGaugeTest {
usageLine(40_000, 5_000, 3_000)); // last record: 48,000
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude", null);
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude");
assertEquals(48_000L, first.tokens(), "must total input+cache_read+cache_creation of the LAST usage record");
assertEquals(LeadContextGauge.State.OK, first.state());
@@ -93,7 +93,7 @@ class LeadContextGaugeTest {
// must change with it, not stay pinned to the first fixture's total.
String otherSession = "22222222-2222-2222-2222-222222222222";
writeTranscript(tmp, otherSession, usageLine(100_000, 50_000, 50_000)); // last record: 200,000
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude", null);
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude");
assertEquals(200_000L, second.tokens());
assertTrue(second.tokens() != first.tokens(), "changing N in the fixture must change the reported number");
}
@@ -111,12 +111,12 @@ class LeadContextGaugeTest {
compactionLine(),
usageLine(3_000, 0, 0));
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude", null);
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude");
assertEquals(2, twoCompactions.compactions());
String sessionZeroCompactions = "44444444-4444-4444-4444-444444444444";
writeTranscript(tmp, sessionZeroCompactions, usageLine(3_000, 0, 0));
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude", null);
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude");
assertEquals(0, zeroCompactions.compactions(), "changing K in the fixture must change the reported count");
}
@@ -126,7 +126,7 @@ class LeadContextGaugeTest {
@DisplayName("a missing transcript file reports UNKNOWN with no token number")
void missingFileIsUnknown(@TempDir Path tmp) {
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude", null);
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
assertNull(reading.tokens());
}
@@ -148,7 +148,7 @@ class LeadContextGaugeTest {
assumeFalse(Files.isReadable(file),
"runs as root (CI container): the read bit does not stop root, so this case cannot be set up here");
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude", null);
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
assertNull(reading.tokens());
} finally {
@@ -171,7 +171,7 @@ class LeadContextGaugeTest {
Files.writeString(file, lastCompleteLine + "\n" + tornLine, StandardCharsets.UTF_8);
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude", null);
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.OK, reading.state(),
"a torn final line must not turn a good earlier reading into UNKNOWN");
assertEquals(6_000L, reading.tokens(),
@@ -187,7 +187,7 @@ class LeadContextGaugeTest {
"{this is not json at all",
"neither is this{{{");
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude", null);
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
"every line unparseable is the real format-change signal and must still report UNKNOWN");
assertNull(reading.tokens());
@@ -224,12 +224,12 @@ class LeadContextGaugeTest {
AtomicLong now = new AtomicLong(0);
LeadContextGauge gauge = new LeadContextGauge(now::get, 5_000);
gauge.read(configDir, SESSION_ID, "claude", null);
gauge.read(configDir, SESSION_ID, "claude", null); // still inside the TTL window
gauge.read(configDir, SESSION_ID, "claude");
gauge.read(configDir, SESSION_ID, "claude"); // still inside the TTL window
assertEquals(1, gauge.diskReadCount(), "two reads inside the TTL must touch disk once");
now.set(6_000); // past the TTL
gauge.read(configDir, SESSION_ID, "claude", null);
gauge.read(configDir, SESSION_ID, "claude");
assertEquals(2, gauge.diskReadCount(), "a read past the TTL must touch disk again");
}
@@ -241,8 +241,8 @@ class LeadContextGaugeTest {
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
LeadContextGauge gauge = new LeadContextGauge();
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode", null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null, null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude", null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode").state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude").state());
}
}
@@ -1,16 +1,10 @@
package dev.ltms.fleet.lead;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.ResilientAgentLaunch;
import dev.ltms.fleet.herdr.WorkspaceControl;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.LinkedHashMap;
import java.util.List;
@@ -70,21 +64,6 @@ class LeadLauncherTest {
return new LeadLauncher(new AgentControl(herdr), new WorkspaceControl(herdr), cfg);
}
/** As {@link #launcher}, plus a fast no-op sleeper so a busy-retry test never real-sleeps. */
private static LeadLauncher fastLauncher(FakeHerdr herdr, FleetConfig cfg) {
return new LeadLauncher(new AgentControl(herdr), new WorkspaceControl(herdr), cfg, () -> { });
}
/** {@link #opusProfile()} with one argv element long enough to overflow the pane line limit. */
private static FleetConfig.Profile hugeArgvProfile() {
return new FleetConfig.Profile(
"opus", null, "claude-opus-5", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "x".repeat(1500)), "tab", "fleet", null,
"http://127.0.0.1:8765/mcp", null, null,
null, null, null,
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true, null);
}
@SuppressWarnings("unchecked")
private static List<String> startedArgs(FakeHerdr herdr) {
return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("args");
@@ -107,11 +86,7 @@ class LeadLauncherTest {
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertTrue(herdr.called("agent.start"), "a lead must actually be started");
// fleetd #727: the name carries a per-process nonce and a per-start sequence number — the
// same unique-naming scheme HerdrPeerLauncher uses for members — rather than the fixed
// "lead-opus" a stale registry entry could block a legitimate relaunch under.
assertTrue(startedName(herdr).matches("lead-opus-[0-9a-f]{6}-\\d+"),
"name is lead-<name>-<nonce>-<seq>: " + startedName(herdr));
assertEquals("lead-opus", startedName(herdr));
}
/** The tab is labelled with the configured `tab:` so the scanner finds the lead on the next resolve. */
@@ -461,247 +436,4 @@ class LeadLauncherTest {
assertEquals(0, launcher(herdr, cfg).ensureLeads());
assertTrue(herdr.calls.isEmpty(), "nothing declared ⇒ nothing scanned");
}
// ── fleetd #727: the same three launch protections every member spawn gets ───────────────────
/**
* A freshly created pane may not have redrawn its prompt yet, so herdr answers
* {@code agent_pane_busy}. The lead launch must wait it out rather than fail on the first miss
* — exactly the retry {@code HerdrPeerLauncher} already gives every member.
*/
@Test
void aLeadLaunchRetriesWhileTheSeedShellBoots() {
FakeHerdr herdr = new FakeHerdr().agentPaneBusyTimes(2);
assertEquals(1, fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"the lead must still start once the shell is ready");
assertEquals(3, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"two busy rejections, then the successful start");
}
/**
* The busy retry is bounded, not an infinite poll. If the pane never becomes ready the launch
* must eventually give up and log the failure, not hang the daemon's reconcile loop forever —
* proven here by a budget that would still be busy on attempt
* {@value ResilientAgentLaunch#SHELL_READY_RETRIES} and a call count that stops exactly there.
*/
@Test
void aLeadLaunchGivesUpAfterTheBoundedBusyBudgetRatherThanLoopingForever() {
FakeHerdr herdr = new FakeHerdr().agentPaneBusyTimes(999);
assertEquals(0, fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"a pane that never becomes ready must not be reported as a started lead");
assertEquals(ResilientAgentLaunch.SHELL_READY_RETRIES,
herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"the retry budget is bounded: it stops after exactly SHELL_READY_RETRIES attempts");
}
/**
* herdr refuses a duplicate agent {@code name} outright. A stale registry entry — a crashed
* lead session, or a name the registry has not yet released — must not permanently block a
* legitimate relaunch, so each retry attempt carries a fresh per-start name.
*/
@Test
@SuppressWarnings("unchecked")
void aLeadLaunchRetriesUnderAFreshNameWhenTheOldNameIsStillTaken() {
FakeHerdr herdr = new FakeHerdr().agentNameTakenTimes(2);
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"the lead must still start once a free name is found");
List<String> names = herdr.calls.stream()
.filter(c -> c.method().equals("agent.start"))
.map(c -> ((Map<String, Object>) c.params()).get("name").toString())
.toList();
assertEquals(3, names.size(), "2 rejected + 1 success");
assertEquals(3, Set.copyOf(names).size(), "each attempt must use a distinct name");
}
/**
* The name-collision retry is bounded too. If the name is taken on every attempt, the launch
* must give up rather than keep minting new names forever — the per-start naming scheme means
* every failed attempt was refused outright by herdr (no process exists under a name herdr
* refused), so a bounded, exhausted retry never leaves a second process running: nothing is
* running at all.
*/
@Test
void aLeadLaunchNeverEndsUpWithASecondProcessWhenTheNameStaysTaken() {
FakeHerdr herdr = new FakeHerdr().agentNameTakenTimes(999);
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"a name that is never free must not be reported as a started lead");
assertEquals(ResilientAgentLaunch.NAME_RETRIES,
herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"the retry budget is bounded: it stops after exactly NAME_RETRIES attempts, "
+ "never racing a duplicate into existence");
}
/**
* fleetd #220/#727: herdr types the launch command into the pane as one line, and a pty line
* buffer holds only 1024 bytes — past that the tail is dropped with no error at all, and the
* backend exits on a mangled argument. The lead launch must refuse an over-long command outright
* rather than let it be typed and silently truncated.
*/
@Test
void anOverlongLeadArgvIsRefusedRatherThanTypedAndTruncated() {
FakeHerdr herdr = new FakeHerdr();
Logger logger = (Logger) LoggerFactory.getLogger(LeadLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
int started;
try {
started = launcher(herdr, configWith(lead("opus", "lead: opus", 1), hugeArgvProfile()))
.ensureLeads();
} finally {
logger.detachAppender(appender);
}
assertEquals(0, started, "an over-long command must never be reported as a started lead");
assertFalse(herdr.called("agent.start"),
"nothing may be started — a truncated command is worse than no spawn");
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains("failed to launch"))
.findFirst()
.orElseThrow(() -> new AssertionError("expected a WARN naming the launch failure: "
+ appender.list));
assertTrue(warn.contains("1024"), "names the limit: " + warn);
assertTrue(warn.contains("opus"), "names the profile: " + warn);
assertTrue(warn.contains("x".repeat(60)), "names the culprit argument: " + warn);
}
// ── fleetd #726 unit 1: the single-lead relaunch seam ─────────────────────────────────────
/**
* The returned agent's {@code terminalId()}/{@code paneId()} are the ones the fake
* {@code AgentControl} actually started — not a coincidental field left over from the caller.
* {@code paneId()} echoes the exact {@code pane_id} the launch's own {@code agent.start} call
* carried (protocol 19: the agent starts into the pane it is asked to), and {@code
* terminalId()} is herdr's own generated id, which the fake always shapes as {@code
* term_new_<n>}.
*/
@Test
void relaunchReturnsTheStartedAgent() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
assertNotNull(started, "a launchable, configured lead must start");
Object startedPaneIdParam = ((Map<?, ?>) herdr.lastCall("agent.start").params()).get("pane_id");
assertEquals(startedPaneIdParam, started.paneId(),
"paneId() must be the pane the agent.start call actually targeted");
assertTrue(started.terminalId() != null && started.terminalId().startsWith("term_new_"),
"terminalId() must be herdr's own generated id: " + started.terminalId());
}
/** The new tab is labelled with the lead's configured {@code tab:}, and AFTER the start. */
@Test
void relaunchLabelsTheNewTabAfterStarting() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
assertNotNull(started);
assertEquals("lead: opus", ((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("label"));
int startIndex = indexOfLastCall(herdr, "agent.start");
int renameIndex = indexOfLastCall(herdr, "tab.rename");
assertTrue(renameIndex > startIndex,
"the tab must be renamed AFTER the start succeeds, not before: start=" + startIndex
+ " rename=" + renameIndex);
}
private static int indexOfLastCall(FakeHerdr herdr, String method) {
int idx = -1;
List<FakeHerdr.Call> calls = herdr.calls;
for (int i = 0; i < calls.size(); i++) {
if (calls.get(i).method().equals(method)) {
idx = i;
}
}
return idx;
}
@Test
void relaunchOfAnUnknownLeadNameReturnsNullAndStartsNothing() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("not-declared");
assertNull(started);
assertFalse(herdr.called("agent.start"));
assertFalse(herdr.called("workspace.create"));
assertFalse(herdr.called("tab.create"));
}
@Test
void relaunchOfARecogniseOnlyLeadReturnsNullAndStartsNothing() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead(null, "lead: dead", 1))).relaunch("opus");
assertNull(started);
assertFalse(herdr.called("agent.start"));
}
@Test
void relaunchWithAnUnconfiguredProfileReturnsNullAndStartsNothing() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead("nope", "lead: opus", 1))).relaunch("opus");
assertNull(started);
assertFalse(herdr.called("agent.start"));
}
/**
* The outer retry {@link LeadLauncher#relaunch(String)} owns, separate from {@code
* ResilientAgentLaunch}'s internal {@code agent_name_taken} retry: a failed attempt must not
* be the end of the whole relaunch. Each of the first two attempts exhausts {@code
* ResilientAgentLaunch.NAME_RETRIES} name attempts (every one of them rejected), so each
* attempt's own tab is created and then closed; the third attempt's first name is free.
*/
@Test
void relaunchRetriesTheWholeAttemptAndSucceedsOnTheThird() {
FakeHerdr herdr = new FakeHerdr()
.agentNameTakenTimes(2 * ResilientAgentLaunch.NAME_RETRIES);
dev.ltms.fleet.herdr.Agent started =
fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
assertNotNull(started, "the third attempt's first name is free — it must succeed");
assertEquals(3, herdr.calls.stream().filter(c -> c.method().equals("tab.create")).count(),
"one tab per attempt: three attempts");
assertEquals(2, herdr.calls.stream().filter(c -> c.method().equals("tab.close")).count(),
"the two failed attempts' tabs must be closed");
}
/**
* Every attempt fails outright (a herdr error {@code ResilientAgentLaunch} does not retry at
* all) — {@link LeadLauncher#relaunch(String)} must give up after exactly {@code
* RELAUNCH_ATTEMPTS} and must not leak any of the tabs it created along the way.
*/
@Test
void relaunchGivesUpAfterExactlyRelaunchAttemptsAndLeaksNoTab() {
FakeHerdr herdr = new FakeHerdr().agentStartFailsWith("some_other_error");
dev.ltms.fleet.herdr.Agent started =
fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
assertNull(started, "every attempt failed — relaunch must give up, not hang or guess");
assertEquals(LeadLauncher.RELAUNCH_ATTEMPTS,
herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"exactly RELAUNCH_ATTEMPTS attempts, no more, no fewer");
long tabsCreated = herdr.calls.stream().filter(c -> c.method().equals("tab.create")).count();
long tabsClosed = herdr.calls.stream().filter(c -> c.method().equals("tab.close")).count();
assertEquals(LeadLauncher.RELAUNCH_ATTEMPTS, tabsCreated);
assertEquals(tabsCreated, tabsClosed, "every tab this method created must be closed — no leaks");
}
}
File diff suppressed because it is too large Load Diff
@@ -63,17 +63,6 @@ class FleetMcpAuthzTest {
/** A fully wired FleetMcp on fakes — constructing it is itself part of what is under test. */
private FleetMcp mcp(boolean enforce) {
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
return mcp(enforce, CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, new MemberRegistry(null)));
}
/**
* As {@link #mcp(boolean)}, with an explicit {@link CallerResolver} — so a test can wire known
* leads/collaborators and drive {@code denyFor}'s real {@code knownLeadOrCollaborator()}
* classifier instead of the default empty one.
*/
private FleetMcp mcp(boolean enforce, CallerResolver callers) {
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
@@ -94,7 +83,9 @@ class FleetMcpAuthzTest {
// to be omitted to reach "legacy" is now always real, and AuthorizationMode is the
// separate, explicit choice that governs enforcement.
mcp = new FleetMcp(messages, workers, sessions, identity, sessions.asPresence(),
new PrimaryRegistry(null), callers,
new PrimaryRegistry(null),
CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, new MemberRegistry(null)),
enforce ? FleetMcp.AuthorizationMode.ENFORCED : FleetMcp.AuthorizationMode.UNENFORCED,
metrics, FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), null, FleetMcp.OutageSource.none(),
@@ -106,7 +97,6 @@ class FleetMcpAuthzTest {
private static final Principal WORKER_A = Principal.worker("term_a", 200);
private static final Principal ANON = Principal.anonymous();
private static final Principal ARCH_DESIGN = Principal.architect("lead-designer", "term_design", 400);
private static final Principal COLLABORATOR = Principal.collaborator("ops", "term_collab", 600);
// --- the table, enforced on THIS path too ---------------------------------------------------
@@ -166,45 +156,6 @@ class FleetMcpAuthzTest {
}
}
/**
* fleetd #669 Unit A: {@code SEND} is split into three actions ({@link Authz.Action#SEND},
* {@link Authz.Action#COORD_SEND}, {@link Authz.Action#ANSWER}), each carrying the same grant
* the one undivided action gave. An architect holds all three, exactly as it held the one.
*/
@Test
void anArchitectMayUseAllThreeSendShapesOverMcp() {
FleetMcp m = mcp(true);
for (Authz.Action a : new Authz.Action[]{Authz.Action.SEND, Authz.Action.COORD_SEND,
Authz.Action.ANSWER}) {
assertNull(m.denyFor(ARCH_DESIGN, a, "term_a"),
a + " carries the same grant the undivided SEND action gave an architect");
}
}
/** The other half of the same split: a worker is excluded from all three, as it was from one. */
@Test
void aWorkerMayNotUseAnySendShapeOverMcp() {
FleetMcp m = mcp(true);
for (Authz.Action a : new Authz.Action[]{Authz.Action.SEND, Authz.Action.COORD_SEND,
Authz.Action.ANSWER}) {
McpSchema.CallToolResult denied = m.denyFor(WORKER_A, a, "term_a");
assertNotNull(denied, a + " must stay refused to a worker");
assertTrue(denied.isError(), "a refusal is returned as an MCP tool error");
}
}
/**
* fleetd #669 Unit A / #678: {@code TASK_READ} (ticket polling, session status) is split out of
* the roster-only {@code READ}, carrying forward the grant the undivided action gave. A worker
* still has both — it never gained or lost anything by the split.
*/
@Test
void aWorkerKeepsBothReadActionsAfterTheSplit() {
FleetMcp m = mcp(true);
assertNull(m.denyFor(WORKER_A, Authz.Action.READ, null));
assertNull(m.denyFor(WORKER_A, Authz.Action.TASK_READ, null));
}
@Test
void anArchitectMayReplyAndAskOnlyAsItsOwnPaneOverMcp() {
FleetMcp m = mcp(true);
@@ -236,39 +187,6 @@ class FleetMcpAuthzTest {
"the caller IS authenticated — it is just not the right role");
}
/**
* {@code denyFor} passes the real production classifier, not a test-supplied one — no
* terminal is recognised as a configured lead or collaborator, so a collaborator's SEND is
* refused over MCP.
*/
@Test
void aCollaboratorMayNotSendOverMcpWithTheRealProductionClassifier() {
FleetMcp m = mcp(true);
McpSchema.CallToolResult denied = m.denyFor(COLLABORATOR, Authz.Action.SEND, "term_lead");
assertNotNull(denied, "no terminal is recognised as a lead or collaborator yet");
assertTrue(denied.isError());
}
/**
* fleetd #669 Unit D: wires a real {@link CallerResolver} with a known lead and a known
* collaborator tab, and leaves a spawned member's own terminal recognised by neither map — so a
* collaborator's SEND reaches both named peers and is refused for the spawned member's terminal,
* over MCP's {@code denyFor}.
*/
@Test
void aCollaboratorMaySendToAKnownLeadOrCollaboratorButNotToASpawnedMembersTerminal() {
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
() -> Map.of("term_lead_known", "lead-x"), new MemberRegistry(null),
t -> null, () -> Map.of("term_collab_known", "ops2"));
FleetMcp m = mcp(true, callers);
assertNull(m.denyFor(COLLABORATOR, Authz.Action.SEND, "term_lead_known"));
assertNull(m.denyFor(COLLABORATOR, Authz.Action.SEND, "term_collab_known"));
assertNotNull(m.denyFor(COLLABORATOR, Authz.Action.SEND, "term_a"),
"a spawned member's own terminal must stay unreachable, even once the classifier is real");
}
@Test
void theLegacyConstructorLeavesTheGateOpen() {
// The 22 pre-existing FleetMcpTest cases rely on no authorization being enforced.
@@ -355,10 +273,10 @@ class FleetMcpAuthzTest {
+ "the anchors have drifted, this test is not testing what it claims to");
Pattern trailingArg = Pattern.compile(
"listFleet\\([^;]*?,\\s*(coordinatorVisibleTo\\(principal\\(exchange\\)\\)|true|false)\\s*[,)]",
"listFleet\\([^;]*?,\\s*(coordinatorVisibleTo\\(principal\\(exchange\\)\\)|true|false)\\s*\\)\\s*;",
Pattern.DOTALL);
Matcher m = trailingArg.matcher(handlerBlock);
assertTrue(m.find(), "could not locate listFleet(...)'s coordinator boolean argument in the "
assertTrue(m.find(), "could not locate listFleet(...)'s trailing boolean argument in the "
+ "listHandler block -- the call shape changed, update this test's anchor: " + handlerBlock);
String trailing = m.group(1);
assertEquals("coordinatorVisibleTo(principal(exchange))", trailing,
@@ -366,197 +284,6 @@ class FleetMcpAuthzTest {
+ "calling, not pass a literal boolean -- found: " + trailing);
}
// --- fleetd #703: who may see fleet_list's collaborators array -------------------------------
/**
* fleetd #703: {@link FleetMcp#collaboratorsVisibleTo} is the whole policy decision for
* {@code fleet_list}'s {@code collaborators} array. Visible to exactly the roles that may
* {@code SEND} to a named peer -- the primary, an architect, and a collaborator itself -- never
* a worker, which holds {@code READ} but can never {@code SEND} to a collaborator, and never an
* anonymous caller.
*/
@Test
void onlyPrimaryArchitectAndCollaboratorMaySeeTheCollaboratorsArray() {
assertTrue(FleetMcp.collaboratorsVisibleTo(PRIMARY), "the primary must see the collaborators array");
assertTrue(FleetMcp.collaboratorsVisibleTo(ARCH_DESIGN), "an architect must see the collaborators array");
assertTrue(FleetMcp.collaboratorsVisibleTo(COLLABORATOR), "a collaborator must see its own peer roster");
assertFalse(FleetMcp.collaboratorsVisibleTo(WORKER_A),
"a worker holds READ but can never SEND to a collaborator, so it must not see the array");
assertFalse(FleetMcp.collaboratorsVisibleTo(ANON), "authenticated as nothing must not see it either");
}
/**
* fleetd #703, same reasoning as {@link #theFleetListHandlerActuallyConsultsCoordinatorVisibleTo}:
* the predicate above can be perfectly correct while the one production call site never asks it.
* This reads {@code FleetMcp.java}'s own source and asserts the {@code fleet_list} handler's
* {@code listFleet(...)} call both threads {@code callers.collaborators()} into the payload and
* asks {@code collaboratorsVisibleTo(principal(exchange))} for the visibility flag, rather than a
* literal boolean or an empty map.
*/
@Test
void theFleetListHandlerActuallyConsultsCollaboratorsVisibleTo() throws Exception {
String source = Files.readString(MCP_SOURCE);
int start = source.indexOf("listHandler =");
assertTrue(start >= 0, "could not find the fleet_list handler (listHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("stopHandler =", start);
assertTrue(end > start, "could not find the handler declared after listHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does contain a call to listFleet(...) -- if this
// fails, the anchors above moved and the assertions below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("listFleet("),
"control failed: the scraped listHandler block contains no listFleet( call at all -- "
+ "the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("callers.collaborators()"),
"the fleet_list handler must thread callers.collaborators() into listFleet(...), not an "
+ "empty or literal map -- block: " + handlerBlock);
assertTrue(handlerBlock.contains("collaboratorsVisibleTo(principal(exchange))"),
"the fleet_list handler must ask collaboratorsVisibleTo(principal(exchange)) who is "
+ "calling, not pass a literal boolean -- block: " + handlerBlock);
}
// --- who may see fleet_list's leads and members arrays ---------------------------------------
/**
* {@link FleetMcp#leadsVisibleTo} is the whole policy decision for {@code fleet_list}'s
* {@code leads} array: visible to exactly the roles that may {@code SEND} to a lead -- the
* primary, an architect, and a collaborator -- never a worker, which holds {@code READ} but
* can never {@code SEND} at all, and never an anonymous caller.
*/
@Test
void primaryArchitectAndCollaboratorMaySeeTheLeadsArray() {
assertTrue(FleetMcp.leadsVisibleTo(PRIMARY), "the primary must see the leads array");
assertTrue(FleetMcp.leadsVisibleTo(ARCH_DESIGN), "an architect must see the leads array");
assertTrue(FleetMcp.leadsVisibleTo(COLLABORATOR),
"a collaborator may SEND to a lead, so it must see the leads array to learn where");
assertFalse(FleetMcp.leadsVisibleTo(WORKER_A),
"a worker holds READ but can never SEND, so it must not see the leads array");
assertFalse(FleetMcp.leadsVisibleTo(ANON), "authenticated as nothing must not see it either");
}
/**
* As {@link #primaryArchitectAndCollaboratorMaySeeTheLeadsArray}, for the {@code members}
* array -- but a collaborator may {@code SEND} only to a lead or another collaborator, never
* to a spawned member, so it must not see this one.
*/
@Test
void onlyPrimaryAndArchitectMaySeeTheMembersArray() {
assertTrue(FleetMcp.membersVisibleTo(PRIMARY), "the primary must see the members array");
assertTrue(FleetMcp.membersVisibleTo(ARCH_DESIGN), "an architect must see the members array");
assertFalse(FleetMcp.membersVisibleTo(WORKER_A),
"a worker holds READ but can never SEND, so it must not see the members array");
assertFalse(FleetMcp.membersVisibleTo(COLLABORATOR),
"a collaborator may SEND to a lead, never to a spawned member, so it must not see the members array");
assertFalse(FleetMcp.membersVisibleTo(ANON), "authenticated as nothing must not see it either");
}
/**
* Same reasoning as {@link #theFleetListHandlerActuallyConsultsCoordinatorVisibleTo}: the
* predicate above can be perfectly correct while the one production call site never asks it.
* This reads {@code FleetMcp.java}'s own source and asserts the {@code fleet_list} handler's
* {@code listFleet(...)} call asks {@code leadsVisibleTo(principal(exchange))} for the leads
* visibility flag, rather than a literal boolean.
*/
@Test
void theFleetListHandlerActuallyConsultsLeadsVisibleTo() throws Exception {
String source = Files.readString(MCP_SOURCE);
int start = source.indexOf("listHandler =");
assertTrue(start >= 0, "could not find the fleet_list handler (listHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("stopHandler =", start);
assertTrue(end > start, "could not find the handler declared after listHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does contain a call to listFleet(...) -- if this
// fails, the anchors above moved and the assertion below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("listFleet("),
"control failed: the scraped listHandler block contains no listFleet( call at all -- "
+ "the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("leadsVisibleTo(principal(exchange))"),
"the fleet_list handler must ask leadsVisibleTo(principal(exchange)) who is calling, "
+ "not pass a literal boolean -- block: " + handlerBlock);
}
/** As {@link #theFleetListHandlerActuallyConsultsLeadsVisibleTo}, for {@code membersVisibleTo}. */
@Test
void theFleetListHandlerActuallyConsultsMembersVisibleTo() throws Exception {
String source = Files.readString(MCP_SOURCE);
int start = source.indexOf("listHandler =");
assertTrue(start >= 0, "could not find the fleet_list handler (listHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("stopHandler =", start);
assertTrue(end > start, "could not find the handler declared after listHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
assertTrue(handlerBlock.contains("listFleet("),
"control failed: the scraped listHandler block contains no listFleet( call at all -- "
+ "the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("membersVisibleTo(principal(exchange))"),
"the fleet_list handler must ask membersVisibleTo(principal(exchange)) who is calling, "
+ "not pass a literal boolean -- block: " + handlerBlock);
}
/**
* {@code fleet_poll{ticket}} must thread the calling connection's owner key into
* {@link MessageService#poll(String, String)}, so a worker cannot read a ticket a different
* session created.
*/
@Test
void theFleetPollHandlerActuallyThreadsCallerOwnerIntoPoll() throws Exception {
String source = Files.readString(MCP_SOURCE);
int start = source.indexOf("pollHandler =");
assertTrue(start >= 0, "could not find the fleet_poll handler (pollHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("ackHandler =", start);
assertTrue(end > start, "could not find the handler declared after pollHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does call poll(...) -- if this fails, the anchors
// above moved and the assertions below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("poll(messages,"),
"control failed: the scraped pollHandler block contains no poll(messages, ...) call "
+ "at all -- the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("principal(exchange).ownerKey()"),
"the fleet_poll handler must thread principal(exchange).ownerKey() into poll(...), not omit "
+ "it or pass a literal null -- block: " + handlerBlock);
}
/**
* {@code fleet_status} must thread the calling connection's owner key into
* {@link FleetMcp#status(MessageService, String, String)}, so a caller that did not create a
* worker's open delegation cannot read its pending question through the status handler either.
*/
@Test
void theFleetStatusHandlerActuallyThreadsCallerOwnerIntoStatus() throws Exception {
String source = Files.readString(MCP_SOURCE);
int start = source.indexOf("statusHandler =");
assertTrue(start >= 0, "could not find the fleet_status handler (statusHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("pollHandler =", start);
assertTrue(end > start, "could not find the handler declared after statusHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does call status(...) -- if this fails, the anchors
// above moved and the assertions below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("status(messages,"),
"control failed: the scraped statusHandler block contains no status(messages, ...) "
+ "call at all -- the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("principal(exchange).ownerKey()"),
"the fleet_status handler must thread principal(exchange).ownerKey() into status(...), not "
+ "omit it or pass a literal null -- block: " + handlerBlock);
}
// --- which action each tool hands the gate (fleetd #272) ------------------------------------
/**
@@ -570,53 +297,35 @@ class FleetMcpAuthzTest {
* the whole time the defect was live -- the table was right, the action fed to it was wrong.
*/
@Test
void pollingByTargetIsADrainAndPollingByTicketIsATaskRead() {
void pollingByTargetIsADrainAndPollingByTicketIsARead() {
assertEquals(Authz.Action.DRAIN, FleetMcp.pollAction("term_b", null),
"poll by target removes the replies — that is a drain, not an observation");
assertEquals(Authz.Action.TASK_READ, FleetMcp.pollAction(null, null),
"poll by ticket changes nothing, but is not the roster-only READ action");
assertEquals(Authz.Action.TASK_READ, FleetMcp.pollAction(" ", null),
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null, null),
"poll by ticket changes nothing");
assertEquals(Authz.Action.READ, FleetMcp.pollAction(" ", null),
"a blank target is an absent target");
}
/**
* fleetd #421: a coordId branch is a THIRD operation behind fleet_poll's one name, and it must
* map to {@link Authz.Action#COORD_READ} — never {@link Authz.Action#READ} or {@link
* Authz.Action#TASK_READ}, even though this branch also consumes nothing. A lead-to-lead body
* is not the roster and not a ticket/status read, so folding this branch into either would let
* any worker or architect read every peer lead's mail in full. coordId also takes priority over
* target when both happen to be present — it addresses a different inbox entirely.
* map to {@link Authz.Action#COORD_READ} — never {@link Authz.Action#READ}, even though this
* branch also consumes nothing. READ's grant is open to every authenticated role on the premise
* that the roster carries no secrets; a lead-to-lead body is not the roster, so folding this
* branch into READ would let any worker read every peer lead's mail in full. coordId also takes
* priority over target when both happen to be present — it addresses a different inbox entirely.
*/
@Test
void pollingByCoordIdIsACoordReadNeverAPlainOrTaskRead() {
void pollingByCoordIdIsACoordReadNeverAPlainRead() {
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction(null, "mac-opus"),
"reading held peer mail must not be mapped to a widely-readable action");
"reading held peer mail must not be mapped to the everyone-readable READ action");
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction(" ", "mac-opus"),
"a blank target must not fall through to READ/DRAIN when coordId is present");
assertEquals(Authz.Action.TASK_READ, FleetMcp.pollAction(null, " "),
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null, " "),
"a blank coordId is an absent coordId, same as target/ticket");
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction("term_b", "mac-opus"),
"coordId takes priority over target — this is a different inbox, not a drain");
}
/**
* {@code fleet_send} is three call shapes behind one tool name, exactly as {@code fleet_poll}
* is (fleetd #669 Unit A). {@link FleetMcp#sendAction} picks the action from the arguments, not
* the handler, for the same reason {@link FleetMcp#pollAction} does: a test can assert the
* mapping the handler actually uses.
*/
@Test
void sendMapsToThreeDifferentActionsByItsArguments() {
assertEquals(Authz.Action.SEND, FleetMcp.sendAction(null, null),
"a plain delivery, with neither coordId nor turnId, is a local SEND");
assertEquals(Authz.Action.COORD_SEND, FleetMcp.sendAction("mac-opus", null),
"coordId addresses a peer lead over the coordination broker");
assertEquals(Authz.Action.ANSWER, FleetMcp.sendAction(null, "turn-1"),
"turnId resolves a worker's blocked question");
assertEquals(Authz.Action.COORD_SEND, FleetMcp.sendAction("mac-opus", "turn-1"),
"coordId takes priority over turnId, mirroring sendToLead's own mutual-exclusion check");
}
@Test
void everyRegisteredToolHasItsHandlerActionPinned() {
// fleetd #469: this used to scrape FleetMcp.java's tool("…") calls for the registered set —
@@ -635,22 +344,16 @@ class FleetMcpAuthzTest {
() -> tool + " is registered but has no pinned authorization action"));
assertEquals(Authz.Action.SEND, FleetMcp.toolAction("fleet_send", Map.of()));
assertEquals(Authz.Action.SEND,
FleetMcp.toolAction("fleet_send", Map.of("sessionId", "term_a", "content", "hi")));
assertEquals(Authz.Action.COORD_SEND,
FleetMcp.toolAction("fleet_send", Map.of("coordId", "mac-opus", "content", "hi")));
assertEquals(Authz.Action.ANSWER,
FleetMcp.toolAction("fleet_send", Map.of("turnId", "turn-1", "content", "hi")));
assertEquals(Authz.Action.REPLY, FleetMcp.toolAction("fleet_reply", Map.of()));
assertEquals(Authz.Action.ASK, FleetMcp.toolAction("fleet_ask", Map.of()));
assertEquals(Authz.Action.TASK_READ, FleetMcp.toolAction("fleet_status", Map.of()));
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_status", Map.of()));
assertEquals(Authz.Action.DRAIN, FleetMcp.toolAction("fleet_ack", Map.of()));
assertEquals(Authz.Action.SPAWN, FleetMcp.toolAction("fleet_spawn", Map.of()));
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_list", Map.of()));
assertEquals(Authz.Action.STOP, FleetMcp.toolAction("fleet_stop", Map.of()));
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_profiles", Map.of()));
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_whoami", Map.of()));
assertEquals(Authz.Action.TASK_READ, FleetMcp.toolAction("fleet_poll", Map.of("ticket", "task")));
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_poll", Map.of("ticket", "task")));
assertEquals(Authz.Action.DRAIN, FleetMcp.toolAction("fleet_poll", Map.of("target", "term_b")));
assertEquals(Authz.Action.COORD_READ,
FleetMcp.toolAction("fleet_poll", Map.of("coordId", "mac-opus")));
@@ -11,7 +11,6 @@ import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.msg.InMemoryReplyInbox;
@@ -25,8 +24,6 @@ import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
@@ -39,14 +36,14 @@ import static org.junit.jupiter.api.Assertions.*;
* fleetd #480 Unit C — the {@code fleet_handover} MCP tool, the surface that finally calls
* {@link LeadRollover#open}/{@link LeadRollover#confirm}/{@link LeadRollover#cancel}.
*
* <p>Uses {@link LeadRollover}'s PUBLIC constructor (real wall clock, real 250ms poll, a real
* virtual-thread continuation runner) rather than its package-private test constructor — this
* test lives in {@code dev.ltms.fleet.mcp}, not {@code dev.ltms.fleet.lead}, and does not need to
* control the post-{@code confirm()} continuation's timing: it only asserts the SYNCHRONOUS return
* value of {@code open}/{@code confirm}/{@code cancel}, which is exactly what {@code
* FleetMcp.handover} forwards to the client. {@code turnSettleSeconds}/{@code
* relaunchReadySeconds} are kept at 1s so a confirmed request's background continuation (which
* this class does not wait on or assert against) gives up quickly rather than polling for 20s on a
* <p>Uses {@link LeadRollover}'s PUBLIC constructor (real wall clock, real 250ms settle poll, a
* real virtual-thread continuation runner) rather than its package-private test constructor —
* this test lives in {@code dev.ltms.fleet.mcp}, not {@code dev.ltms.fleet.lead}, and does not
* need to control the post-{@code confirm()} continuation's timing: it only asserts the
* SYNCHRONOUS return value of {@code open}/{@code confirm}/{@code cancel}, which is exactly what
* {@code FleetMcp.handover} forwards to the client. {@code turnSettleSeconds}/{@code
* clearSettleSeconds} are kept at 1s so a confirmed request's background continuation (which this
* class does not wait on or assert against) gives up quickly rather than polling for 20s on a
* daemon virtual thread.
*/
class FleetMcpHandoverTest {
@@ -70,26 +67,10 @@ class FleetMcpHandoverTest {
return new FleetConfig.LeadRollover(handoverPath, false, 3600, 1, 1, "read the handover file");
}
private static FleetConfig minimalFleetConfig() {
try {
Path yaml = Files.createTempFile("fleet-mcp-handover-test", ".yaml");
Files.writeString(yaml, "bind:\n port: 8080\n");
return FleetConfig.load(yaml);
} catch (IOException e) {
throw new UncheckedIOException(e);
}
}
private LeadRollover newRollover(String handoverPath) {
// Every handoverPath this test class uses comes from tmp.resolve(...), which is already
// absolute, so the workspace lookup is never actually consulted — a no-op lookup is enough.
// None of this class's tests reach the recognition-wait or the relaunch call, so the
// launcher's own functional correctness is irrelevant here — any constructed instance,
// backed by the same fake herdr, is enough.
WorkspaceControl spaces = new WorkspaceControl(herdr);
LeadLauncher launcher = new LeadLauncher(agents, spaces, minimalFleetConfig());
return new LeadRollover(agents, spaces, launcher, () -> cfg(handoverPath),
_ -> null, _ -> null, Map::of);
return new LeadRollover(agents, () -> cfg(handoverPath), _ -> null);
}
/** A fully wired FleetMcp on fakes (mirrors FleetMcpAuthzTest's helper), plus a leadRollover. */
@@ -86,7 +86,7 @@ class FleetMcpLeadContextGaugeWiringTest {
LeadContextGauge contextGauge = new LeadContextGauge();
AtomicReference<String> configuredDir = new AtomicReference<>(dirA.toString());
FleetMcp.LeadConfigDirSource source = new FleetMcp.LeadConfigDirSource(name ->
LEAD_NAME.equals(name) ? configuredDir.get() : null, _ -> null);
LEAD_NAME.equals(name) ? configuredDir.get() : null);
String firstRead = textOf(listFleet(herdr, sessions, contextGauge, source));
assertTrue(firstRead.contains("\"tokens\":11000"),
@@ -121,7 +121,6 @@ class FleetMcpLeadContextGaugeWiringTest {
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
FleetMcp.LeadSeatSource.none(), contextGauge, leadConfigDirs,
Map.of(LEAD_TERMINAL, LEAD_NAME), LEAD_TERMINAL, Map.of(), false,
FleetMcp.CoordinationSource.none(), false, true, true);
Map.of(LEAD_TERMINAL, LEAD_NAME), LEAD_TERMINAL, FleetMcp.CoordinationSource.none(), false);
}
}
@@ -10,7 +10,6 @@ import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.inject.LoopWatchdog;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.Injector;
@@ -77,7 +76,7 @@ class FleetMcpTest {
private void assertSendRoundTrips(String target, Set<String> profiles) throws Exception {
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, target, "hi", 4000L, null, profiles, null));
() -> FleetMcp.send(messages, target, "hi", 4000L, null, profiles));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(target) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
@@ -103,7 +102,7 @@ class FleetMcpTest {
void sendThenReplyRoundTrips() throws Exception {
// fleet_send blocks; fleet_reply resolves it with the worker's structured answer.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, "term_a", "review this", 4000L, null, Set.of(), null));
() -> FleetMcp.send(messages, "term_a", "review this", 4000L, null, Set.of()));
// Wait until the send has opened its waiter so the reply resolves it (CB-307: reply now
// queues in the inbox if no waiter is open, which would break the round-trip).
@@ -181,7 +180,7 @@ class FleetMcpTest {
String turnId = afterTurnId.substring(0, afterTurnId.indexOf('"'));
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, null));
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L));
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
deadline = System.currentTimeMillis() + 3000;
@@ -279,7 +278,7 @@ class FleetMcpTest {
assertEquals("late reply", messages.drainReplies("term_a").getFirst().content());
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, firstTurnId, "config.yaml", 5000L, null));
() -> FleetMcp.answer(messages, firstTurnId, "config.yaml", 5000L));
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
@@ -288,91 +287,6 @@ class FleetMcpTest {
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
}
// --- fleetd #715: fleet_send{turnId} is gated on the caller that owns the turn -------------
/**
* A blocking {@code fleet_send} from one caller opens the turn; a {@code fleet_send{turnId}}
* from a different caller is refused as an error, and the real owner's answer still succeeds.
*/
@Test
void aDifferentCallersMcpAnswerIsRefusedForABlockingSendButTheRealOwnerSucceeds() throws Exception {
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, T, "do X", 5000L, null, Set.of(), "term_owner"));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T));
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
() -> FleetMcp.ask(messages, T, "which config?", 5000L));
McpSchema.CallToolResult question = send.get(5, TimeUnit.SECONDS);
assertTrue(textOf(question).contains("[question]"), textOf(question));
String questionText = textOf(question);
String afterTurnId = questionText.substring(questionText.indexOf("turnId=\"") + "turnId=\"".length());
String turnId = afterTurnId.substring(0, afterTurnId.indexOf('"'));
McpSchema.CallToolResult hijacked = FleetMcp.answer(messages, turnId, "evil.yaml", 500L, "term_attacker");
assertTrue(hijacked.isError(), "a caller that did not open this turn must get an error, not an answer");
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, "term_owner"));
assertEquals("config.yaml", textOf(ask.get(5, TimeUnit.SECONDS)));
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
FleetMcp.reply(messages, T, Role.WORKER, "done");
assertEquals("done", textOf(answer.get(5, TimeUnit.SECONDS)));
}
/**
* Same hijack and control through the fire-and-poll ({@code sendAsync}) path: the owner comes
* from the resolved principal, not from a caller argument threaded through a live call.
*/
@Test
void aDifferentCallersMcpAnswerIsRefusedForAnAsyncSendButTheRealOwnerSucceeds() throws Exception {
Principal owner = Principal.worker("term_owner", 1);
McpSchema.CallToolResult accepted =
FleetMcp.sendAsync(messages, T, "do it", null, Set.of(), owner);
String ticket = textOf(accepted).substring(textOf(accepted).indexOf("ticket=") + "ticket=".length()).trim();
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T));
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
() -> FleetMcp.ask(messages, T, "which config?", 5000L));
MessageService.TaskView asking = messages.poll(ticket);
deadline = System.currentTimeMillis() + 3000;
while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
asking = messages.poll(ticket);
}
assertEquals(MessageService.Phase.ASKING, asking.phase());
String turnId = asking.turnId();
McpSchema.CallToolResult hijacked = FleetMcp.answer(messages, turnId, "evil.yaml", 500L,
"worker:term_attacker");
assertTrue(hijacked.isError(), "a caller that did not create this delegation must get an error");
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase(),
"a refused answer must not advance the async ticket's phase");
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, owner.ownerKey()));
assertEquals("config.yaml", textOf(ask.get(5, TimeUnit.SECONDS)));
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
FleetMcp.reply(messages, T, Role.WORKER, "done");
assertEquals("done", textOf(answer.get(5, TimeUnit.SECONDS)));
}
@Test
void pollUnknownTicketIsAnError() {
McpSchema.CallToolResult res = FleetMcp.poll(messages, "task-999", null);
@@ -382,7 +296,7 @@ class FleetMcpTest {
@Test
void sendTimesOutWithAWorkingNote() {
McpSchema.CallToolResult res = FleetMcp.send(messages, "term_a", "hi", 120L, null, Set.of(), null);
McpSchema.CallToolResult res = FleetMcp.send(messages, "term_a", "hi", 120L, null, Set.of());
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
assertTrue(textOf(res).contains("no reply"), "got: " + textOf(res));
}
@@ -398,7 +312,7 @@ class FleetMcpTest {
void sendTimesOutWithAnUnconfirmedNoteNotARetryInvitation() throws Exception {
herdr.agentSendFailsWith("send_failed");
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, T, "hi", 150L, null, Set.of(), null));
() -> FleetMcp.send(messages, T, "hi", 150L, null, Set.of()));
long deadline = System.currentTimeMillis() + 2000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
@@ -418,13 +332,13 @@ class FleetMcpTest {
@Test
void sendRejectsMissingArgs() {
assertTrue(FleetMcp.send(messages, null, "hi", null, null, Set.of(), null).isError());
assertTrue(FleetMcp.send(messages, "term_a", " ", null, null, Set.of(), null).isError());
assertTrue(FleetMcp.send(messages, null, "hi", null, null, Set.of()).isError());
assertTrue(FleetMcp.send(messages, "term_a", " ", null, null, Set.of()).isError());
}
@Test
void sendRejectsAConfiguredProfileNameBeforeAcceptingIt() {
McpSchema.CallToolResult blocking = FleetMcp.send(messages, "sol", "hi", 100L, null, Set.of("sol"), null);
McpSchema.CallToolResult blocking = FleetMcp.send(messages, "sol", "hi", 100L, null, Set.of("sol"));
McpSchema.CallToolResult async = FleetMcp.sendAsync(messages, "sol", "hi", null, Set.of("sol"));
assertTrue(blocking.isError());
@@ -443,7 +357,7 @@ class FleetMcpTest {
assertSendRoundTrips("term_live_member", profiles);
// A herdr-owned pane outside the bridge roster cannot be classified at accept time.
McpSchema.CallToolResult result = FleetMcp.send(messages, "external-pane", "hi", 10L, null, profiles, null);
McpSchema.CallToolResult result = FleetMcp.send(messages, "external-pane", "hi", 10L, null, profiles);
assertFalse(result.isError(), "an unclassified target must not be rejected at acceptance time");
}
@@ -535,7 +449,7 @@ class FleetMcpTest {
void askThenAnswerRoundTrips() throws Exception {
// The primary delegates and blocks; wait until its waiter is open before the worker asks.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, "term_a", "do X", 5000L, null, Set.of(), null));
() -> FleetMcp.send(messages, "term_a", "do X", 5000L, null, Set.of()));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
@@ -557,7 +471,7 @@ class FleetMcpTest {
// The primary answers via fleet_send(turnId); this blocks again for the worker's reply.
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, null));
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L));
// The worker's ask returns the answer — it resumes the same turn.
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
@@ -583,7 +497,7 @@ class FleetMcpTest {
@Test
void answerToAStaleTurnIsAnError() {
McpSchema.CallToolResult res = FleetMcp.answer(messages, "term_a#999", "too late", 500L, null);
McpSchema.CallToolResult res = FleetMcp.answer(messages, "term_a#999", "too late", 500L);
assertTrue(res.isError());
assertTrue(textOf(res).contains("no longer open"), textOf(res));
}
@@ -777,7 +691,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
String out = textOf(res);
// fleetd #361: reports both which coord-id a peer must use to reach ME, and this daemon's
@@ -806,7 +720,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
String out = textOf(res);
assertTrue(out.contains("\"mailbox\":{\"status\":\"unknown\"}"), out);
@@ -865,82 +779,19 @@ class FleetMcpTest {
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
Principal worker = Principal.worker("term_a", 1);
boolean callerIsPrimary = Principal.worker("term_a", 1).isPrimary();
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), worker.isPrimary(),
FleetMcp.leadsVisibleTo(worker), FleetMcp.membersVisibleTo(worker));
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), callerIsPrimary);
String out = textOf(res);
assertFalse(out.contains("\"coordinator\""), "a worker must never see the coordinator key at all: " + out);
assertFalse(out.contains("mac-opus"), "no fragment of the coordinator row may leak either: " + out);
assertFalse(out.contains("\"leads\""), "a worker must never see the leads key at all: " + out);
assertFalse(out.contains("\"members\""), "a worker must never see the members key at all: " + out);
assertTrue(out.contains("\"healthCoverage\""), "the rest of the result must still be present: " + out);
}
/**
* A worker's result must carry neither the {@code leads} nor the {@code members} key, and no
* fragment of either row leaks even though both are fully populated for this call -- a key
* check alone would pass on an implementation that still built the rows and only renamed or
* nested them.
*/
@Test
void listLeaksNoLeadOrMemberRowFragmentToAWorker() {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_owner",
new WorktreeRequest("cb-304", null));
Principal worker = Principal.worker("term_a", 1);
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of("term_lead_x", "mac-opus"), "", FleetMcp.CoordinationSource.none(), worker.isPrimary(),
FleetMcp.leadsVisibleTo(worker), FleetMcp.membersVisibleTo(worker));
String out = textOf(res);
assertFalse(out.contains("\"leads\""), out);
assertFalse(out.contains("\"members\""), out);
assertFalse(out.contains(s.terminalId()), "no member row fragment may leak to a worker: " + out);
assertFalse(out.contains("term_owner"), "no member owner fragment may leak to a worker: " + out);
assertFalse(out.contains("mac-opus"), "no lead row fragment may leak to a worker: " + out);
assertTrue(out.contains("\"healthCoverage\""), out);
}
/**
* A collaborator may {@code SEND} to a lead, so it must see the {@code leads} array -- the
* only place {@code fleet_whoami} does not already give it a lead's address. It may never
* {@code SEND} to a spawned member, so the {@code members} key must stay absent for it, with
* no fragment of a populated member row leaking either.
*/
@Test
void listShowsLeadsButNotMembersToACollaborator() {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_owner",
new WorktreeRequest("cb-304", null));
Principal collaborator = Principal.collaborator("ops", "term_collab", 600);
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of("term_lead_x", "mac-opus"), "", FleetMcp.CoordinationSource.none(), collaborator.isPrimary(),
FleetMcp.leadsVisibleTo(collaborator), FleetMcp.membersVisibleTo(collaborator));
String out = textOf(res);
assertTrue(out.contains("\"leads\""), "a collaborator must see the leads array: " + out);
assertTrue(out.contains("mac-opus"), "a collaborator must see the lead's name/address: " + out);
assertFalse(out.contains("\"members\""), "a collaborator must never see the members key: " + out);
assertFalse(out.contains(s.terminalId()), "no member row fragment may leak to a collaborator: " + out);
assertFalse(out.contains("term_owner"), "no member owner fragment may leak to a collaborator: " + out);
assertTrue(out.contains("\"leads\""), "the rest of the result must still be present: " + out);
assertTrue(out.contains("\"members\""), out);
assertTrue(out.contains("\"healthCoverage\""), out);
}
@@ -956,19 +807,16 @@ class FleetMcpTest {
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
Principal architect = Principal.architect("lead-designer", "term_design", 400);
boolean callerIsPrimary = Principal.architect("lead-designer", "term_design", 400).isPrimary();
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), architect.isPrimary(),
FleetMcp.leadsVisibleTo(architect), FleetMcp.membersVisibleTo(architect));
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), callerIsPrimary);
String out = textOf(res);
assertFalse(out.contains("\"coordinator\""), "an architect must never see the coordinator key either: " + out);
assertTrue(out.contains("\"leads\""), "an architect must still see the leads array: " + out);
assertTrue(out.contains("\"members\""), "an architect must still see the members array: " + out);
}
/**
@@ -993,7 +841,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true));
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true));
assertTrue(gatedAsPrimary.contains("\"coordinator\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"selfId\":\"mac-opus\""), gatedAsPrimary);
@@ -1002,8 +850,6 @@ class FleetMcpTest {
assertTrue(gatedAsPrimary.contains("\"heldDurable\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"held\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"peers\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"leads\""), "the primary must still see the leads array: " + gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"members\""), "the primary must still see the members array: " + gatedAsPrimary);
}
@Test
@@ -1022,7 +868,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
String out = textOf(res);
assertTrue(out.contains("\"msgId\":\"m1\""), out);
@@ -1033,83 +879,6 @@ class FleetMcpTest {
assertTrue(out.contains("x".repeat(80) + "…"), "expected an 80-char preview with an ellipsis: " + out);
}
// --- fleetd #703: fleet_list's collaborators array -------------------------------------------
/** Calls the canonical {@code listFleet} overload directly, so a test can set the collaborators
* payload and its visibility independently of a real {@code Principal} / MCP exchange. */
private static McpSchema.CallToolResult listFleetWithCollaborators(FakeHerdr h,
Map<String, String> collaborators, boolean collaboratorsVisible) {
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
return FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
FleetMcp.LeadSeatSource.none(), new LeadContextGauge(), FleetMcp.LeadConfigDirSource.none(),
Map.of(), "", collaborators, collaboratorsVisible,
FleetMcp.CoordinationSource.none(), false, true, true);
}
/**
* fleetd #703 acceptance A: a visible caller with one configured collaborator gets a
* {@code collaborators} row whose key ({@code CallerResolver.collaborators()}'s
* {@code terminal_id -> name} entry) lands as that row's {@code sessionId}, and {@code leads}/
* {@code members} are unaffected by the new key.
*/
@Test
void listReportsACollaboratorsRowKeyedByTheCollaboratorsTerminalId() {
FakeHerdr h = new FakeHerdr();
String out = textOf(listFleetWithCollaborators(h, Map.of("term_collab", "ops"), true));
assertTrue(out.contains("\"collaborators\":["), out);
assertTrue(out.contains("\"name\":\"ops\""), out);
assertTrue(out.contains("\"sessionId\":\"term_collab\""), out);
assertTrue(out.contains("\"leads\":[]"), out);
assertTrue(out.contains("\"members\":[]"), out);
}
/**
* fleetd #703 acceptance A, control half: with no collaborator configured, the key is absent
* (caller cannot see it) or an empty array (caller can), and {@code leads}/{@code members} are
* unchanged either way.
*/
@Test
void listOmitsOrEmptiesCollaboratorsWhenNoneAreConfigured() {
FakeHerdr h = new FakeHerdr();
String visibleButEmpty = textOf(listFleetWithCollaborators(h, Map.of(), true));
assertTrue(visibleButEmpty.contains("\"collaborators\":[]"), visibleButEmpty);
assertTrue(visibleButEmpty.contains("\"leads\":[]"), visibleButEmpty);
assertTrue(visibleButEmpty.contains("\"members\":[]"), visibleButEmpty);
String notVisible = textOf(listFleetWithCollaborators(h, Map.of(), false));
assertFalse(notVisible.contains("\"collaborators\""), notVisible);
assertTrue(notVisible.contains("\"leads\":[]"), notVisible);
assertTrue(notVisible.contains("\"members\":[]"), notVisible);
}
/**
* fleetd #703 acceptance B: a worker must not see the {@code collaborators} array at all, while
* an architect -- a role that also holds READ, same as a worker -- does see it. Both halves are
* asserted: a test that only checked the worker-hidden half would pass even if the feature were
* never wired up for anyone.
*/
@Test
void listOmitsCollaboratorsForAWorkerAndIncludesThemForAnArchitect() {
FakeHerdr h = new FakeHerdr();
Map<String, String> collaborators = Map.of("term_collab", "ops");
String asWorker = textOf(listFleetWithCollaborators(h, collaborators,
FleetMcp.collaboratorsVisibleTo(Principal.worker("term_w", 1))));
assertFalse(asWorker.contains("\"collaborators\""),
"a worker must not see the collaborators array: " + asWorker);
String asArchitect = textOf(listFleetWithCollaborators(h, collaborators,
FleetMcp.collaboratorsVisibleTo(Principal.architect("design", "term_arch", 2))));
assertTrue(asArchitect.contains("\"collaborators\":["),
"an architect must see the collaborators array: " + asArchitect);
assertTrue(asArchitect.contains("\"sessionId\":\"term_collab\""), asArchitect);
}
/**
* fleetd #421: {@code mailbox.pending} counts only broker-ready messages, so a blocked lead's
* normal, healthy state is {@code "pending": 0} next to a non-empty {@code held[]} — which
@@ -1131,7 +900,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
String out = textOf(res);
assertTrue(out.contains("\"pending\":0"), out);
@@ -1160,7 +929,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
String out = textOf(res);
assertTrue(out.contains("\"heldDurable\":false"),
@@ -1229,7 +998,7 @@ class FleetMcpTest {
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "",
new FleetMcp.CoordinationSource(channel, List.of("fleet01-lead", "fleet02-lead", "fleet03-lead")), true, true, true);
new FleetMcp.CoordinationSource(channel, List.of("fleet01-lead", "fleet02-lead", "fleet03-lead")), true);
String out = textOf(res);
assertTrue(out.contains("\"coordId\":\"fleet01-lead\",\"status\":\"exists\",\"pending\":2,\"consumers\":1"), out);
@@ -1922,8 +1691,8 @@ class FleetMcpTest {
Principal architect = Principal.architect("lead-designer", "term_design", 400);
MemberPresence presence = new MemberPresence();
FleetMcp.markTrackedCallerPresent(worker, presence);
FleetMcp.markTrackedCallerPresent(architect, presence);
FleetMcp.markSpawnedMemberPresent(worker, presence);
FleetMcp.markSpawnedMemberPresent(architect, presence);
assertTrue(presence.isPresent("term_worker"));
assertTrue(presence.isPresent("term_design"));
@@ -1934,40 +1703,25 @@ class FleetMcpTest {
Principal lead = Principal.leader("opus", "term_lead", 100);
MemberPresence presence = new MemberPresence();
FleetMcp.markTrackedCallerPresent(lead, presence);
FleetMcp.markTrackedCallerPresent(Principal.anonymous(), presence);
FleetMcp.markSpawnedMemberPresent(lead, presence);
FleetMcp.markSpawnedMemberPresent(Principal.anonymous(), presence);
assertFalse(presence.isPresent("term_lead"));
}
/**
* The item whose absence would be silent: an observer's own MCP contact must still mark
* presence, or a pane resolving to the unconfigured-pane floor would sit on the injector
* readiness gate forever once something addresses it.
*/
@Test
void anObserverContactMarksPresence() {
Principal observer = Principal.observer("term_observer", 800);
MemberPresence presence = new MemberPresence();
FleetMcp.markTrackedCallerPresent(observer, presence);
assertTrue(presence.isPresent("term_observer"));
}
@Test
void statusReportsLiveAgentStatus() {
FakeHerdr blocked = new FakeHerdr().agentStatus("blocked");
AgentControl blockedAgents = new AgentControl(blocked);
McpSchema.CallToolResult res = FleetMcp.status(
new MessageService(blockedAgents, new Injector(blockedAgents), rendezvous), "term_a", null);
new MessageService(blockedAgents, new Injector(blockedAgents), rendezvous), "term_a");
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("blocked", textOf(res));
}
/**
* A lead polling {@code fleet_status} on its normal cadence — not {@code fleet_poll} — must
* also see a worker's open async {@code fleet_ask} question, since the reverse-rendezvous
* CB-582: a lead polling {@code fleet_status} on its normal cadence — not {@code fleet_poll}
* — must also see a worker's open async {@code fleet_ask} question, since the reverse-rendezvous
* window it opened with is far shorter than that cadence.
*/
@Test
@@ -1990,7 +1744,7 @@ class FleetMcpTest {
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
assertEquals(MessageService.Phase.ASKING, asking.phase());
McpSchema.CallToolResult res = FleetMcp.status(messages, T, null);
McpSchema.CallToolResult res = FleetMcp.status(messages, T);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.startsWith("idle"), "the live status must still lead the text: " + out);
@@ -2001,66 +1755,7 @@ class FleetMcpTest {
// Clean up the still-open ask so the background thread does not linger past the test.
String turnId = asking.turnId();
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(turnId, "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.resolve(T, "done"));
answer.get(5, TimeUnit.SECONDS);
}
/**
* {@code fleet_status}'s pending-ask block (the question, its {@code turnId} and its ticket)
* is shown only to the caller whose owner key created the delegation, or to the unnamed primary.
* A different caller still sees the base status line, but none of the pending-ask fields.
*/
@Test
void statusGatesThePendingAskFieldsByTheDelegationsCreatorOwner() throws Exception {
Principal creator = Principal.worker("term_creator", 1);
String ticket = messages.sendAsync(T, "task that asks", null, creator);
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T), "sendAsync should have opened its rendezvous waiter");
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking;
deadline = System.currentTimeMillis() + 3000;
do {
asking = messages.poll(ticket);
Thread.sleep(5);
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
assertEquals(MessageService.Phase.ASKING, asking.phase());
String other = textOf(FleetMcp.status(messages, T, "worker:term_other"));
assertTrue(other.startsWith("idle"), "the base status must still be shown: " + other);
assertFalse(other.contains("which config file?"),
"a non-creating caller must not see the question text: " + other);
assertFalse(other.contains(asking.turnId()),
"a non-creating caller must not see the turnId: " + other);
assertFalse(other.contains(ticket),
"a non-creating caller must not see the ticket: " + other);
String creatorStatus = textOf(FleetMcp.status(messages, T, creator.ownerKey()));
assertTrue(creatorStatus.contains("which config file?"),
"the creator must see the question: " + creatorStatus);
assertTrue(creatorStatus.contains(asking.turnId()),
"the creator must see the turnId: " + creatorStatus);
assertTrue(creatorStatus.contains(ticket), "the creator must see the ticket: " + creatorStatus);
String unnamed = textOf(FleetMcp.status(messages, T, null));
assertTrue(unnamed.contains("which config file?"),
"a caller with no terminal (the unnamed primary) must see the question: " + unnamed);
// Clean up the still-open ask so the background thread does not linger past the test.
String turnId = asking.turnId();
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(turnId, "config.yaml", 5000, creator.ownerKey()));
() -> messages.answer(turnId, "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
@@ -2142,53 +1837,6 @@ class FleetMcpTest {
assertTrue(out.contains("\"sessionId\":\"term_design\""), out);
}
/**
* A collaborator reports its own role and name, never the {@code leader} key a lead gets —
* {@code role} already reads {@code "collaborator"}, so a {@code leader} key alongside it
* would be self-contradicting. Control: the same call shape fed a named lead must still carry
* {@code leader}, so this is not passing because the key stopped being emitted for everyone.
*/
@Test
void whoamiReportsACollaboratorWithNoLeaderKeyButALeadStillGetsOne() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
McpSchema.CallToolResult collabRes = FleetMcp.whoami(
Principal.collaborator("ops", "term_collab", 700), sessions);
assertNotEquals(Boolean.TRUE, collabRes.isError());
String collabOut = textOf(collabRes);
assertTrue(collabOut.contains("\"role\":\"collaborator\""), collabOut);
assertTrue(collabOut.contains("\"collaborator\":\"ops\""), collabOut);
assertTrue(collabOut.contains("\"sessionId\":\"term_collab\""), collabOut);
assertFalse(collabOut.contains("leader"), collabOut);
McpSchema.CallToolResult leadRes = FleetMcp.whoami(
Principal.leader("opus", "term_lead", 100), sessions);
String leadOut = textOf(leadRes);
assertTrue(leadOut.contains("\"leader\":\"opus\""), leadOut);
}
/**
* An observer reports its own role and pane, never a {@code leader} key. Without an explicit
* branch it would reach the lead branch by elimination and look right only because the
* {@code leader} key is guarded on a non-null name — this pins the branch rather than the
* accident.
*/
@Test
void whoamiReportsAnObserverNotALead() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
McpSchema.CallToolResult res = FleetMcp.whoami(
Principal.observer("term_observer", 900), sessions);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"observer\""), out);
assertTrue(out.contains("\"sessionId\":\"term_observer\""), out);
assertFalse(out.contains("leader"), out);
}
/**
* CB-548: an architect SEND delegates as its own pane (recording the per-target delegation) but
* must NEVER become the legacy singleton "primary" fallback — the per-target map does not cure
@@ -1,280 +0,0 @@
package dev.ltms.fleet.msg;
import org.junit.jupiter.api.Test;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.List;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
import java.util.stream.Stream;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Pins that no file under {@code src/main/java} calls the fail-open
* {@link MessageService#poll(String)} overload. That overload skips the ownership check in
* {@code MessageService}'s {@code ownsTicket} entirely, so a caller of it can read any session's
* ticket. Every production caller must go through {@link MessageService#poll(String, String)}
* and pass a {@code callerOwner} explicitly, even when it is {@code null}.
*
* <p>This reads each file's own source text rather than reflecting on compiled bytecode, because
* the risk is a future one-word edit at a call site, not a missing overload.
*
* <p>The scan below finds a violation by its receiver, {@code messages.poll(}, rather than the
* bare method name, so it does not mistake {@link java.util.Queue#poll()} for a violation. That
* anchor only covers a {@code MessageService} reached through a variable or field named
* {@code messages}, so {@link #everyMessageServiceDeclarationIsNamedMessages} pins the naming
* convention the anchor depends on: a declaration under any other name would be invisible to the
* scan above, and must turn this second check red instead of passing silently.
*/
class MessageServicePollUsageTest {
private static final Path PRODUCTION_SOURCE = Path.of("src/main/java");
@Test
void noProductionFileCallsTheSingleArgumentPollOverload() throws IOException {
List<String> violations = new ArrayList<>();
List<String> twoArgSites = new ArrayList<>();
int filesScanned = scanForPollCalls(PRODUCTION_SOURCE, violations, twoArgSites);
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "no
// violations found" having looked at nothing.
assertTrue(filesScanned > 0, "control failed: the scan under " + PRODUCTION_SOURCE
+ " visited zero .java files -- the path is wrong, so the absence of violations "
+ "below proves nothing");
assertTrue(violations.isEmpty(), "found a call to the fail-open MessageService.poll(String) "
+ "overload, which skips the ownership check entirely -- pass a callerOwner "
+ "explicitly (even if null) through poll(String, String) instead: " + violations);
// CONTROL: the arity parser actually finds the two genuine two-argument call sites (the
// MCP handler in FleetMcp and the REST handler in FleetApp). If this drops, the parser
// itself is broken, not the production code -- a broken parser (or a scan root that
// reaches no real source) must fail loudly here rather than pass vacuously above.
assertEquals(2, twoArgSites.size(), "control failed: expected exactly the two known "
+ "two-argument messages.poll(...) call sites, found: " + twoArgSites);
assertTrue(twoArgSites.stream().anyMatch(s -> s.contains("FleetMcp.java")),
"control failed: did not find the FleetMcp.java messages.poll(ticket, callerOwner) "
+ "site among: " + twoArgSites);
assertTrue(twoArgSites.stream().anyMatch(s -> s.contains("FleetApp.java")),
"control failed: did not find the FleetApp.java messages.poll(...) site among: "
+ twoArgSites);
}
/**
* The {@code messages.poll(} anchor above only sees a {@code MessageService} reached through
* a variable, field, or parameter named {@code messages}. This asserts that every such
* declaration under {@code src/main/java} uses that name, so a differently named declaration
* -- invisible to the scan above -- fails loudly here instead of letting that scan pass on a
* call site it never looked at.
*/
@Test
void everyMessageServiceDeclarationIsNamedMessages() throws IOException {
List<String> names = new ArrayList<>();
int filesScanned = scanForDeclarationNames(PRODUCTION_SOURCE, names);
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "every
// declaration is named messages" having looked at nothing.
assertTrue(filesScanned > 0, "control failed: the scan under " + PRODUCTION_SOURCE
+ " visited zero .java files -- the path is wrong, so the result below proves nothing");
// CONTROL: the declaration pattern actually finds real declarations. Zero means the
// pattern is broken, not that every MessageService variable, field, or parameter vanished.
assertTrue(names.size() > 0, "control failed: found zero MessageService declarations under "
+ PRODUCTION_SOURCE + " -- the declaration pattern is broken, update it before "
+ "trusting the naming check below");
List<String> other = names.stream().filter(n -> !n.equals("messages")).distinct().toList();
assertTrue(other.isEmpty(), "found a MessageService declaration not named \"messages\": "
+ other + " -- the messages.poll( scan above only looks for that name, so a call "
+ "through a differently named variable or field is invisible to it; either rename "
+ "the declaration or widen that scan's anchor to cover it");
}
private static int scanForPollCalls(Path root, List<String> violations, List<String> twoArgSites)
throws IOException {
List<Path> files = javaFiles(root);
for (Path file : files) {
scanFileForPollCalls(file, violations, twoArgSites);
}
return files.size();
}
private static int scanForDeclarationNames(Path root, List<String> names) throws IOException {
List<Path> files = javaFiles(root);
Pattern declaration = Pattern.compile("MessageService\\s+([A-Za-z_][A-Za-z0-9_]*)");
for (Path file : files) {
if (file.getFileName().toString().equals("MessageService.java")) {
continue; // the type's own declaration, not a caller holding a reference to it
}
String stripped = stripComments(Files.readString(file));
Matcher m = declaration.matcher(stripped);
while (m.find()) {
int j = m.end();
while (j < stripped.length() && Character.isWhitespace(stripped.charAt(j))) j++;
if (j < stripped.length() && stripped.charAt(j) == '(') {
continue; // a method named like the convention, e.g. "MessageService messages()"
}
names.add(m.group(1));
}
}
return files.size();
}
private static List<Path> javaFiles(Path root) throws IOException {
try (Stream<Path> paths = Files.walk(root)) {
return paths.filter(p -> p.toString().endsWith(".java")).toList();
}
}
/**
* Replaces {@code //} and {@code /* *}{@code /} comment text with nothing, leaving code,
* string/char literals and line breaks untouched -- so a comment that merely mentions
* {@code MessageService} in prose can never be read as a declaration.
*/
private static String stripComments(String source) {
StringBuilder out = new StringBuilder(source.length());
boolean inString = false;
boolean inChar = false;
int i = 0;
while (i < source.length()) {
char c = source.charAt(i);
if (inString) {
out.append(c);
if (c == '\\' && i + 1 < source.length()) { out.append(source.charAt(i + 1)); i += 2; continue; }
if (c == '"') inString = false;
i++;
continue;
}
if (inChar) {
out.append(c);
if (c == '\\' && i + 1 < source.length()) { out.append(source.charAt(i + 1)); i += 2; continue; }
if (c == '\'') inChar = false;
i++;
continue;
}
if (c == '"') { inString = true; out.append(c); i++; continue; }
if (c == '\'') { inChar = true; out.append(c); i++; continue; }
if (c == '/' && i + 1 < source.length() && source.charAt(i + 1) == '/') {
while (i < source.length() && source.charAt(i) != '\n') i++;
continue; // leaves the newline itself for the next iteration to append
}
if (c == '/' && i + 1 < source.length() && source.charAt(i + 1) == '*') {
i += 2;
while (i < source.length() && !(source.charAt(i) == '*' && i + 1 < source.length()
&& source.charAt(i + 1) == '/')) {
if (source.charAt(i) == '\n') out.append('\n');
i++;
}
i += 2;
continue;
}
out.append(c);
i++;
}
return out.toString();
}
private static void scanFileForPollCalls(Path file, List<String> violations, List<String> twoArgSites)
throws IOException {
String source = Files.readString(file);
String needle = "messages.poll(";
int from = 0;
int idx;
while ((idx = source.indexOf(needle, from)) >= 0) {
int argsStart = idx + needle.length();
String args = extractBalancedArgs(source, argsStart, file, idx);
int closeParenIndex = argsStart + args.length();
from = closeParenIndex + 1;
if (args.isBlank()) {
continue; // MessageService has no zero-argument poll() -- nothing to classify
}
String site = file + ":" + lineOf(source, idx);
if (topLevelCommaCount(args) == 0) {
violations.add(site + " -- messages.poll(" + args.trim() + ")");
} else {
twoArgSites.add(site);
}
}
}
/**
* The text between {@code messages.poll(} and its matching close paren: balanced over nested
* calls, and never split by a paren or comma sitting inside a string or char literal.
*/
private static String extractBalancedArgs(String source, int start, Path file, int callIndex) {
int depth = 1;
boolean inString = false;
boolean inChar = false;
int i = start;
while (i < source.length()) {
char c = source.charAt(i);
if (inString) {
if (c == '\\') { i += 2; continue; }
if (c == '"') inString = false;
} else if (inChar) {
if (c == '\\') { i += 2; continue; }
if (c == '\'') inChar = false;
} else if (c == '"') {
inString = true;
} else if (c == '\'') {
inChar = true;
} else if (c == '(') {
depth++;
} else if (c == ')') {
depth--;
if (depth == 0) return source.substring(start, i);
}
i++;
}
throw new IllegalStateException(
"unbalanced parentheses scanning " + file + ":" + lineOf(source, callIndex));
}
/**
* Commas at paren/bracket/brace depth zero, skipping string and char literals -- the argument
* separators a human reader would see, not every comma character in the text.
*/
private static int topLevelCommaCount(String args) {
int depth = 0;
int commas = 0;
boolean inString = false;
boolean inChar = false;
int i = 0;
while (i < args.length()) {
char c = args.charAt(i);
if (inString) {
if (c == '\\') { i += 2; continue; }
if (c == '"') inString = false;
} else if (inChar) {
if (c == '\\') { i += 2; continue; }
if (c == '\'') inChar = false;
} else if (c == '"') {
inString = true;
} else if (c == '\'') {
inChar = true;
} else if (c == '(' || c == '[' || c == '{') {
depth++;
} else if (c == ')' || c == ']' || c == '}') {
depth--;
} else if (c == ',' && depth == 0) {
commas++;
}
i++;
}
return commas;
}
private static int lineOf(String source, int index) {
int line = 1;
for (int i = 0; i < index; i++) {
if (source.charAt(i) == '\n') line++;
}
return line;
}
}
@@ -4,7 +4,6 @@ import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.FakeHerdr;
@@ -61,25 +60,13 @@ class MessageServiceTest {
inbox.own(T);
}
/**
* A send budget large enough that a test's own setup — {@link #awaitWaiting()} plus whatever
* status transitions it drives afterward — can never compete with it for the same clock. A test
* that needs {@code send.get(...)}'s own window to be the only timing bound it depends on uses
* {@link #sendAsync(String, long)} with this value instead of the default 5000 ms.
*/
private static final long GENEROUS_SEND_BUDGET_MILLIS = 30_000;
/** Run {@code send} on a background thread; the current thread drives the worker's turn. */
private CompletableFuture<MessageService.Reply> sendAsync() {
return sendAsync("do the task");
}
private CompletableFuture<MessageService.Reply> sendAsync(String content) {
return sendAsync(content, 5000);
}
private CompletableFuture<MessageService.Reply> sendAsync(String content, long timeoutMillis) {
return CompletableFuture.supplyAsync(() -> messages.send(T, content, timeoutMillis, null));
return CompletableFuture.supplyAsync(() -> messages.send(T, content, 5000));
}
private void awaitWaiting() throws InterruptedException {
@@ -93,7 +80,7 @@ class MessageServiceTest {
@Test
void completionFallbackResolvesATurnThatNeverCalledFleetReply() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync("do the task", GENEROUS_SEND_BUDGET_MILLIS);
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
herdr.readText("$ prompt"); // pre-turn pane: no answer yet (baseline reference)
@@ -109,32 +96,6 @@ class MessageServiceTest {
assertTrue(reply.completed(), "a scraped completion still counts as completed");
}
/**
* Pins {@link #GENEROUS_SEND_BUDGET_MILLIS} as the budget {@link
* #completionFallbackResolvesATurnThatNeverCalledFleetReply} depends on. A 5500 ms delay between
* {@link #awaitWaiting()} and the status transitions that drive completion stands in for a loaded
* machine's setup overhead — comfortably past the 5000 ms budget this send no longer uses, and
* still well inside this method's own 30 000 ms budget. The only clock this test depends on is
* {@code send.get}'s own 10 s window.
*/
@Test
void completionFallbackSurvivesASlowHarnessBecauseItsSendBudgetIsNotTheBindingClock() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync("do the task", GENEROUS_SEND_BUDGET_MILLIS);
awaitWaiting();
Thread.sleep(5500);
herdr.readText("$ prompt");
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
herdr.readText("BUILD GREEN: 391 files");
injector.onStatus(T, AgentStatus.IDLE);
MessageService.Reply reply = send.get(10, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, reply.outcome(),
"a slow harness must not be mistaken for a timed-out delivery");
}
@Test
void completionFallbackReplacesAnEchoedInjectedBriefWithNoReportOutcome() throws Exception {
String brief = "Implement the requested change. ".repeat(20);
@@ -189,7 +150,7 @@ class MessageServiceTest {
localRendezvous, new InMemoryReplyInbox());
String brief = "Implement the requested change. ".repeat(20);
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> localMessages.send(T, brief, 5000, (String) null));
CompletableFuture.supplyAsync(() -> localMessages.send(T, brief, 5000));
long deadline = System.currentTimeMillis() + 2000;
while (!localRendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
@@ -342,7 +303,7 @@ class MessageServiceTest {
// The primary answers via fleet_send(turnId); this blocks again for the worker's reply.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
// The worker's ask returns the answer — it resumes the same turn.
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
@@ -378,7 +339,7 @@ class MessageServiceTest {
// The primary answers that one turnId; both asks unblock with the same answer.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
MessageService.AskResult a1 = ask1.get(5, TimeUnit.SECONDS);
MessageService.AskResult a2 = ask2.get(5, TimeUnit.SECONDS);
@@ -473,7 +434,7 @@ class MessageServiceTest {
"the duplicate's own timeout elapses first");
CompletableFuture<MessageService.Reply> answered =
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "fleetd.yaml", 500, null));
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "fleetd.yaml", 500));
MessageService.AskResult a = fresh.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome(),
@@ -508,7 +469,7 @@ class MessageServiceTest {
CompletableFuture<MessageService.Reply> lateAnswer = new CompletableFuture<>();
messages.setAskTimeoutRaceHookForTest(() ->
lateAnswer.complete(messages.answer(turnId, "too late", 500, null)));
lateAnswer.complete(messages.answer(turnId, "too late", 500)));
try {
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.TIMED_OUT, a.outcome());
@@ -540,7 +501,7 @@ class MessageServiceTest {
@Test
void answeringAnUnknownTurnIsStale() {
MessageService.Reply r = messages.answer(T + "#999", "too late", 500, null);
MessageService.Reply r = messages.answer(T + "#999", "too late", 500);
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
"an answer to a turn that never existed (or already lapsed) is stale, not a hang");
}
@@ -573,7 +534,7 @@ class MessageServiceTest {
messages.setAnswerAskLapseRaceHookForTest(() -> rendezvous.answerAsk(turnId, "raced in first"));
try {
MessageService.Reply r = messages.answer(turnId, "too late", 500, null);
MessageService.Reply r = messages.answer(turnId, "too late", 500);
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
"an ask already answered by the race must be seen as lapsed, not double-delivered");
} finally {
@@ -597,7 +558,7 @@ class MessageServiceTest {
void sendTimesOutBeforeDeliveryIsQueuedNotWorking() {
// Nothing ever delivers the message and nothing resolves the send, so the reply future
// times out with delivery still incomplete — the message is still queued for the worker.
MessageService.Reply r = messages.send(T, "never delivered", 50, null);
MessageService.Reply r = messages.send(T, "never delivered", 50);
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome(),
"an undelivered send that times out is still queued, not working");
assertNull(r.text());
@@ -606,7 +567,7 @@ class MessageServiceTest {
@Test
void sendTimesOutAfterDeliveryIsStillWorking() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 300, null));
CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 300));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver — the delivered future now completes
injector.onStatus(T, AgentStatus.WORKING); // worker starts but never replies
@@ -631,7 +592,7 @@ class MessageServiceTest {
void sendTimeoutUsesCancellationDeliveredWhenPickupWinsTheRace() {
messages.setTimeoutCancellationRaceHookForTest(() -> injector.onStatus(T, AgentStatus.IDLE));
try {
MessageService.Reply reply = messages.send(T, "race delivery", 50, null);
MessageService.Reply reply = messages.send(T, "race delivery", 50);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, reply.outcome(),
"cancel reporting DELIVERED means the worker received the timed-out message");
@@ -653,7 +614,7 @@ class MessageServiceTest {
void sendTimesOutWithAttemptedDeliveryReportsUnconfirmedNotQueued() throws Exception {
herdr.agentSendFailsWith("send_failed");
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "brief", 150, null));
CompletableFuture.supplyAsync(() -> messages.send(T, "brief", 150));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // triggers the failing delivery attempt → ATTEMPTED
@@ -680,7 +641,7 @@ class MessageServiceTest {
// The primary answers, unblocking the worker; but the worker never sends the follow-up
// fleet_reply, so the answering send rides out its short window as still-working.
MessageService.Reply answer = messages.answer(q.turnId(), "config.yaml", 200, null);
MessageService.Reply answer = messages.answer(q.turnId(), "config.yaml", 200);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answer.outcome(),
"an answered worker that never replies times out as still working");
@@ -716,7 +677,7 @@ class MessageServiceTest {
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
ask.get(5, TimeUnit.SECONDS); // worker resumed with the answer
awaitWaiting(); // the answering call has (re)opened its own forward waiter
@@ -724,7 +685,7 @@ class MessageServiceTest {
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, done.outcome());
MessageService.Reply probe = messages.send(T, "probe after normal reply", 300, null);
MessageService.Reply probe = messages.send(T, "probe after normal reply", 300);
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
"the session lock must be released after a normal REPLIED answer(), or this bounded "
+ "follow-up send would come back BUSY instead of timing out on its own work");
@@ -757,13 +718,13 @@ class MessageServiceTest {
// The primary answers, unblocking the worker; the worker never sends its follow-up
// fleet_reply, so the answering call rides out its short window as still-working.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 200, null));
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 200));
MessageService.Reply answered = answer.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answered.outcome(),
"an answered worker that never replies times out as still working");
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
MessageService.Reply probe = messages.send(T, "probe after timed-out-working", 300, null);
MessageService.Reply probe = messages.send(T, "probe after timed-out-working", 300);
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
"the session lock must be released after a TIMED_OUT_WORKING answer(), or this bounded "
+ "follow-up send would come back BUSY instead of timing out on its own work");
@@ -792,7 +753,7 @@ class MessageServiceTest {
new java.util.concurrent.atomic.AtomicReference<>();
Thread answerer = new Thread(() -> {
try {
messages.answer(turnId, "config.yaml", 5000, null);
messages.answer(turnId, "config.yaml", 5000);
caught.set(new AssertionError("expected answer() to throw"));
} catch (Throwable t) {
caught.set(t);
@@ -814,7 +775,7 @@ class MessageServiceTest {
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
MessageService.Reply probe = messages.send(T, "probe after exceptional failure", 300, null);
MessageService.Reply probe = messages.send(T, "probe after exceptional failure", 300);
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
"the session lock must be released when answer()'s reply future fails exceptionally, "
+ "or this bounded follow-up send would come back BUSY instead of timing out on its own work");
@@ -841,7 +802,7 @@ class MessageServiceTest {
new java.util.concurrent.atomic.AtomicReference<>();
Thread answerer = new Thread(() -> {
try {
messages.answer(turnId, "config.yaml", 5000, null);
messages.answer(turnId, "config.yaml", 5000);
caught.set(new AssertionError("expected answer() to throw"));
} catch (Throwable t) {
caught.set(t);
@@ -860,147 +821,17 @@ class MessageServiceTest {
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
MessageService.Reply probe = messages.send(T, "probe after interruption", 300, null);
MessageService.Reply probe = messages.send(T, "probe after interruption", 300);
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
"the session lock must be released when answer()'s wait is interrupted, or this bounded "
+ "follow-up send would come back BUSY instead of timing out on its own work");
}
/**
* Drive an async send on {@code T} to a resolved reply, polling as {@code owner} until the
* ticket reports {@link MessageService.Phase#DONE} (or the 2s deadline runs out). {@code owner}
* must be a terminal this ticket's creator check actually accepts, or this loops to the
* deadline and returns a non-{@code DONE} view.
*/
private MessageService.TaskView driveAsyncTicketToDone(String ticket, String owner) throws InterruptedException {
MessageService.TaskView view = null;
long deadline = System.currentTimeMillis() + 2000;
while (view == null || view.phase() != MessageService.Phase.DONE) {
if (System.currentTimeMillis() >= deadline) break;
view = messages.poll(ticket, owner);
//noinspection BusyWait
Thread.sleep(5);
}
return view;
}
@Test
void pollReturnsNullForAnUnknownTicket() {
assertNull(messages.poll("task-999999"), "a ticket that was never minted is unknown");
}
// --- fleetd #719: a per-boot nonce keeps one instance's ticket ids out of another's space ---
/** A second, fully independent instance — its own agents/injector/rendezvous/inbox, not shared. */
private MessageService newIndependentInstance() {
FakeHerdr otherHerdr = new FakeHerdr().readText("BUILD GREEN: 391 files");
AgentControl otherAgents = new AgentControl(otherHerdr);
Injector otherInjector = new Injector(otherAgents);
return new MessageService(otherAgents, otherInjector, new Rendezvous(), new InMemoryReplyInbox());
}
@Test
void twoInstancesMintDisjointTicketIds() {
MessageService other = newIndependentInstance();
String ticketFromThis = messages.sendAsync(T, "task on first instance", null, null);
String ticketFromOther = other.sendAsync(T, "task on second instance", null, null);
assertNotEquals(ticketFromThis, ticketFromOther,
"each instance mints its own id space, so even a first ticket from each must differ");
}
@Test
void foreignInstanceTicketDoesNotResolve() {
MessageService other = newIndependentInstance();
String ticket = messages.sendAsync(T, "task on first instance", null, null);
// `other` must reach the same sequence number, or this test passes against an empty map
// instead of against a colliding id.
other.sendAsync(T, "task on second instance", null, null);
// control: the id resolves in the instance that minted it, so a null below cannot be
// explained by broken plumbing — only by the ticket being foreign to `other`.
assertNotNull(messages.poll(ticket), "the minting instance must still resolve its own ticket");
assertNull(other.poll(ticket), "a ticket minted by a different instance must not resolve here");
}
// --- ticket ownership -----------------------------------------------------------------------
@Test
void leadBIsRefusedFromLeadAsTicket() throws Exception {
Principal leadA = Principal.leader("opus", "term_a", 1);
Principal leadB = Principal.leader("sol", "term_b", 2);
String ticket = messages.sendAsync(T, "long task", null, leadA);
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker works
assertTrue(rendezvous.resolve(T, "secret async result"), "a reply resolves the async send");
MessageService.TaskView owner = driveAsyncTicketToDone(ticket, leadA.ownerKey());
assertNotNull(owner, "the creator must still be able to read its own ticket");
assertEquals(MessageService.Phase.DONE, owner.phase());
MessageService.TaskView refused = messages.poll(ticket, leadB.ownerKey());
assertNotNull(refused, "a different lead gets a refusal, not silence");
assertNotEquals(MessageService.Phase.DONE, refused.phase(),
"a different lead must never see the ticket as DONE");
assertNull(refused.reply(), "a refusal must never carry the reply text");
assertFalse(String.valueOf(refused).contains("secret async result"),
"the reply text must not appear anywhere in the refused view");
}
@Test
void unnamedPrimaryStillReadsAnyTicket() throws Exception {
String ticket = messages.sendAsync(T, "long task", null,
Principal.leader("opus", "term_lead", 1));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker works
assertTrue(rendezvous.resolve(T, "primary-visible result"), "a reply resolves the async send");
MessageService.TaskView view = driveAsyncTicketToDone(ticket, null);
assertNotNull(view, "the unnamed primary must be able to read any ticket");
assertEquals(MessageService.Phase.DONE, view.phase());
assertEquals("primary-visible result", view.reply());
}
@Test
void namedLeadCanPollItsTicketAfterItsTerminalChanges() throws Exception {
Principal oldLead = Principal.leader("opus", "term_OLD", 1);
Principal newLead = Principal.leader("opus", "term_NEW", 2);
assertNotEquals(oldLead.terminal(), newLead.terminal(), "the test requires different terminals");
String ticket = messages.sendAsync(T, "long task", null, oldLead);
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker works
assertTrue(rendezvous.resolve(T, "own result"), "a reply resolves the async send");
MessageService.TaskView view = driveAsyncTicketToDone(ticket, newLead.ownerKey());
assertNotNull(view, "the same named lead must read the ticket from its new terminal");
assertEquals(MessageService.Phase.DONE, view.phase());
assertEquals("own result", view.reply());
}
@Test
void anonymousOwnerKeyIsRefusedByTheTicketGateItself() {
Principal lead = Principal.leader("opus", "term_lead", 1);
String ticket = messages.sendAsync(T, "long task", null, lead);
MessageService.TaskView refused = messages.poll(ticket, Principal.anonymous().ownerKey());
assertNotNull(refused);
assertEquals(MessageService.Phase.FAILED, refused.phase());
assertEquals("forbidden: this ticket was created by a different session", refused.detail());
}
@Test
void architectOwnershipUsesTerminalRatherThanSlot() {
Principal oldArchitect = Principal.architect("opus", "term_OLD", 1);
Principal newArchitect = Principal.architect("opus", "term_NEW", 2);
String ticket = messages.sendAsync(T, "long task", null, oldArchitect);
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket, oldArchitect.ownerKey()).phase());
assertEquals(MessageService.Phase.FAILED, messages.poll(ticket, newArchitect.ownerKey()).phase());
}
@Test
void pollReportsACompletedTicket() throws Exception {
String ticket = messages.sendAsync(T, "long task");
@@ -1027,11 +858,11 @@ class MessageServiceTest {
@Test
void concurrentSendToSameSessionWhileFirstHoldsItIsBusy() throws Exception {
CompletableFuture<MessageService.Reply> first =
CompletableFuture.supplyAsync(() -> messages.send(T, "first", 5000, null));
CompletableFuture.supplyAsync(() -> messages.send(T, "first", 5000));
awaitWaiting(); // the first send now holds the session lock, blocked on its reply
// A second send to the SAME session cannot take the lock within its short window.
MessageService.Reply busy = messages.send(T, "second", 100, null);
MessageService.Reply busy = messages.send(T, "second", 100);
assertEquals(MessageService.Outcome.BUSY, busy.outcome(),
"a second send while another holds the session is busy, not a hang");
assertNull(busy.text());
@@ -1061,13 +892,13 @@ class MessageServiceTest {
PrimaryRegistry reg = new PrimaryRegistry(null);
// L accepts a delegation to W: the send wins the lock and queues delivery → L is recorded.
CompletableFuture<MessageService.Reply> first = CompletableFuture.supplyAsync(
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L), null));
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L)));
awaitWaiting();
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
"an accepted send owns the delegation");
// A attempts W while L holds it → BUSY (lock never taken) → its hook never fires.
MessageService.Reply busy = messages.send(T, "second", 100, () -> reg.recordDelegation(T, LEAD_A), null);
MessageService.Reply busy = messages.send(T, "second", 100, () -> reg.recordDelegation(T, LEAD_A));
assertEquals(MessageService.Outcome.BUSY, busy.outcome());
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
"a BUSY send must not steal the delegator ownership it never earned");
@@ -1088,7 +919,7 @@ class MessageServiceTest {
void anAcceptedSendAfterThePriorOwnerFinishesBecomesTheNewOwner() throws Exception {
PrimaryRegistry reg = new PrimaryRegistry(null);
CompletableFuture<MessageService.Reply> first = CompletableFuture.supplyAsync(
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L), null));
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L)));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
@@ -1099,7 +930,7 @@ class MessageServiceTest {
// L finished; A's later accepted send takes the delegation over.
CompletableFuture<MessageService.Reply> second = CompletableFuture.supplyAsync(
() -> messages.send(T, "second", 5000, () -> reg.recordDelegation(T, LEAD_A), null));
() -> messages.send(T, "second", 5000, () -> reg.recordDelegation(T, LEAD_A)));
awaitWaiting();
assertEquals(LEAD_A, reg.nudgeTargetFor(T).orElseThrow(),
"an accepted send after the owner finished becomes the new delegator");
@@ -1120,7 +951,7 @@ class MessageServiceTest {
void answeringAnAskDoesNotRewriteDelegatorOwnership() throws Exception {
PrimaryRegistry reg = new PrimaryRegistry(null);
CompletableFuture<MessageService.Reply> send = CompletableFuture.supplyAsync(
() -> messages.send(T, "do X", 5000, () -> reg.recordDelegation(T, LEAD_L), null));
() -> messages.send(T, "do X", 5000, () -> reg.recordDelegation(T, LEAD_L)));
awaitWaiting();
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(), "L owns the delegation");
@@ -1135,7 +966,7 @@ class MessageServiceTest {
// L answers the ask on the same turn; the answer path must not touch ownership.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting(); // the answering send reopened its forward waiter
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
@@ -1155,7 +986,7 @@ class MessageServiceTest {
void aThrowingAcceptedHookLeavesNoStaleWaiterOrQueuedOrphan() {
assertThrows(IllegalStateException.class,
() -> messages.send(T, "doomed", 500,
() -> { throw new IllegalStateException("ownership hook failed"); }, null),
() -> { throw new IllegalStateException("ownership hook failed"); }),
"a throwing ownership hook fails the send loudly");
assertFalse(rendezvous.isWaiting(T), "the failed send must not leave a stale rendezvous waiter");
@@ -1283,7 +1114,7 @@ class MessageServiceTest {
@Test
void abandonFailsASendThatIsStillWaitingOnAReleasedSession() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000, null));
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
awaitWaiting();
assertTrue(messages.abandon(T, "session released"), "a live waiter is abandoned");
@@ -1303,7 +1134,7 @@ class MessageServiceTest {
@Test
void abandonDoesNotOverwriteAnAlreadyResolvedSend() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000, null));
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
awaitWaiting();
assertTrue(rendezvous.resolve(T, "the real answer"));
@@ -1434,7 +1265,7 @@ class MessageServiceTest {
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase());
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
@@ -1474,7 +1305,7 @@ class MessageServiceTest {
// same turnId must see it as lapsed rather than resolving a question nobody is waiting on.
assertEquals(MessageService.AskOutcome.TIMED_OUT, ask.get(5, TimeUnit.SECONDS).outcome());
assertEquals(MessageService.Outcome.STALE_TURN,
messages.answer(asking.turnId(), "config.yaml", 200, null).outcome());
messages.answer(asking.turnId(), "config.yaml", 200).outcome());
}
@Test
@@ -1511,7 +1342,7 @@ class MessageServiceTest {
// The primary answers, but its own bounded wait for the worker's resumed turn is short and
// expires before the worker (still genuinely working) gets back to it.
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150, null);
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150);
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome(),
"the primary's own bounded wait gives up before the worker finishes resuming");
@@ -1540,7 +1371,7 @@ class MessageServiceTest {
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150, null);
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150);
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome());
@@ -1594,7 +1425,7 @@ class MessageServiceTest {
messages.setFinishAsyncTaskRaceHookForTest(() -> messages.forgetTurnForTest(turnId));
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000, null));
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting(); // answer() opened its own forward waiter for the resumed worker turn
@@ -1700,7 +1531,7 @@ class MessageServiceTest {
String turnId = asking.turnId();
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000, null));
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer(),
"the worker's own ask() call must have already unblocked with the primary's answer "
+ "before we force the race below");
@@ -1747,7 +1578,7 @@ class MessageServiceTest {
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
String turnId = asking.turnId();
MessageService.Reply answerReply = messages.answer(turnId, "config.yaml", 150, null);
MessageService.Reply answerReply = messages.answer(turnId, "config.yaml", 150);
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome(),
"the primary's own bounded wait must give up first, leaving turnId stamped with no "
@@ -1813,7 +1644,7 @@ class MessageServiceTest {
MessageService.TaskView asking = messages.poll(first);
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
@@ -1847,7 +1678,7 @@ class MessageServiceTest {
// The primary answers it — answer() resumes the turn and blocks for what comes next.
CompletableFuture<MessageService.Reply> answer1 = CompletableFuture.supplyAsync(
() -> messages.answer(asking1.turnId(), "a1", 5000, null));
() -> messages.answer(asking1.turnId(), "a1", 5000));
assertEquals("a1", ask1.get(5, TimeUnit.SECONDS).answer());
// Still in the SAME resumed turn — before replying — the worker asks again.
@@ -1868,7 +1699,7 @@ class MessageServiceTest {
// The primary answers the second question; the worker finally sends its real fleet_reply.
CompletableFuture<MessageService.Reply> answer2 = CompletableFuture.supplyAsync(
() -> messages.answer(turnId2, "a2", 5000, null));
() -> messages.answer(turnId2, "a2", 5000));
assertEquals("a2", ask2.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
@@ -1935,20 +1766,20 @@ class MessageServiceTest {
}
}
// --- fleet_status pendingAsk() -----------------------------------------------------------
// --- CB-582: fleet_status pendingAsk() ------------------------------------------------------
@Test
void pendingAskReturnsNullWhenNoQuestionIsOpen() throws Exception {
assertNull(messages.pendingAsk(T, null), "no async ticket at all -> no pending ask");
assertNull(messages.pendingAsk(T), "no async ticket at all -> no pending ask");
String ticket = messages.sendAsync(T, "long task");
awaitWaiting();
assertNull(messages.pendingAsk(T, null), "a plain pending delegation is not a question");
assertNull(messages.pendingAsk(T), "a plain pending delegation is not a question");
injectDelivery();
assertTrue(rendezvous.resolve(T, "done"));
awaitTicketPhase(ticket, MessageService.Phase.DONE);
assertNull(messages.pendingAsk(T, null), "a finished ticket carries no open question either");
assertNull(messages.pendingAsk(T), "a finished ticket carries no open question either");
}
@Test
@@ -1961,183 +1792,21 @@ class MessageServiceTest {
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
MessageService.PendingAsk pending = messages.pendingAsk(T, null);
MessageService.PendingAsk pending = messages.pendingAsk(T);
assertNotNull(pending, "fleet_status should see the open question");
assertEquals(ticket, pending.ticket());
assertEquals("which config file?", pending.question());
assertEquals(asking.turnId(), pending.turnId());
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertNull(messages.pendingAsk(T, null), "an answered question is no longer pending");
assertNull(messages.pendingAsk(T), "an answered question is no longer pending");
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
/**
* A caller's owner key must match the key that created the delegation to see its pending
* question. The unnamed primary always sees it.
*/
@Test
void pendingAskGatesTheQuestionByTheDelegationsCreatorOwner() throws Exception {
Principal creator = Principal.worker("term_creator", 1);
String ticket = messages.sendAsync(T, "task that asks", null, creator);
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
assertNull(messages.pendingAsk(T, "worker:term_other"),
"a caller whose key did not create the delegation must not see the question");
MessageService.PendingAsk own = messages.pendingAsk(T, creator.ownerKey());
assertNotNull(own, "the creating caller must see its own open question");
assertEquals("which config file?", own.question());
MessageService.PendingAsk unnamed = messages.pendingAsk(T, null);
assertNotNull(unnamed, "a caller with no terminal (the unnamed primary) must always see the question");
assertEquals("which config file?", unnamed.question());
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, creator.ownerKey()));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
/**
* A task created with no recorded owner (a short {@code sendAsync} overload) must not hand its
* open question to a caller with an owner key. Only the unnamed primary may still see it.
*/
@Test
void pendingAskDeniesATerminalBearingCallerWhenTheTaskRecordsNoCreator() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
assertNull(messages.pendingAsk(T, "worker:term_someone"),
"a caller with an owner key must not see a question whose task records no creator");
assertNotNull(messages.pendingAsk(T, null),
"the unnamed primary must still see it even with no recorded creator");
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
// --- fleetd #715: answer() is gated on the caller that owns the turn -----------------------
/**
* A blocking {@code fleet_send} from one caller opens the turn; a different caller's answer is
* refused with no side effect on the rendezvous or the worker's blocked {@code fleet_ask} —
* only the real owner can answer it.
*/
@Test
void aDifferentCallersAnswerIsRefusedForABlockingSendDelegationButTheRealOwnerSucceeds() throws Exception {
CompletableFuture<MessageService.Reply> send = CompletableFuture.supplyAsync(
() -> messages.send(T, "do X", 5000, "term_owner"));
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertFalse(rendezvous.isWaiting(T), "the forward waiter closes once the question surfaces");
// The hijack: a different caller answers the SAME turnId.
MessageService.Reply hijacked = messages.answer(q.turnId(), "evil.yaml", 500, "term_attacker");
assertEquals(MessageService.Outcome.NOT_TURN_OWNER, hijacked.outcome(),
"a caller that did not open this turn must be refused, not served");
assertFalse(rendezvous.isWaiting(T), "a refused answer must not open a forward waiter");
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
assertEquals(T, rendezvous.askSession(q.turnId()), "a refused answer must leave the ask turn open");
// The control: the real owner answers the same turnId and the turn resumes normally.
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(q.turnId(), "config.yaml", 5000, "term_owner"));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
/**
* Same hijack and control as the blocking case, but the delegation is opened through the
* fire-and-poll path ({@code sendAsync}) — the owner comes from the ticket's recorded creator
* terminal, not a caller argument threaded through a live blocking call.
*/
@Test
void aDifferentCallersAnswerIsRefusedForAnAsyncSendDelegationButTheRealOwnerSucceeds() throws Exception {
Principal owner = Principal.worker("term_owner", 1);
String ticket = messages.sendAsync(T, "task that asks", null, owner);
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
MessageService.Reply hijacked = messages.answer(asking.turnId(), "evil.yaml", 500,
"worker:term_attacker");
assertEquals(MessageService.Outcome.NOT_TURN_OWNER, hijacked.outcome(),
"a caller that did not create this delegation must be refused, not served");
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase(),
"a refused answer must not advance the async ticket's phase");
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, owner.ownerKey()));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
assertEquals("done", awaitTicketPhase(ticket, MessageService.Phase.DONE).reply());
}
@Test
void namedLeadCanSeeAndAnswerAnAskAfterItsTerminalChangesWhileLeadBIsRefused() throws Exception {
Principal oldLead = Principal.leader("opus", "term_OLD", 1);
Principal newLead = Principal.leader("opus", "term_NEW", 2);
Principal leadB = Principal.leader("sol", "term_SOL", 3);
assertNotEquals(oldLead.terminal(), newLead.terminal(), "the test requires different terminals");
String ticket = messages.sendAsync(T, "task that asks", null, oldLead);
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
assertNotNull(messages.pendingAsk(T, newLead.ownerKey()),
"the same lead at its new terminal must see the pending ask");
assertNull(messages.pendingAsk(T, leadB.ownerKey()),
"another lead must not see the pending ask");
assertEquals(MessageService.Outcome.NOT_TURN_OWNER,
messages.answer(asking.turnId(), "evil.yaml", 500, leadB.ownerKey()).outcome());
assertFalse(ask.isDone(), "another lead must not resolve the worker's ask");
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, newLead.ownerKey()));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
assertEquals("done", awaitTicketPhase(ticket, MessageService.Phase.DONE).reply());
}
// --- CB-588: async ticket terminal nudges ---------------------------------------------------
//
// MessageService.reply's rendezvous fast path is exactly what an async ticket always takes
@@ -2394,7 +2063,7 @@ class MessageServiceTest {
// Clean up the still-open ask so the background thread does not linger past the test.
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000, null));
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
@@ -2420,7 +2089,7 @@ class MessageServiceTest {
.filter(c -> c.method().equals("agent.prompt")).count();
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000, null));
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
java.util.concurrent.CountDownLatch terminalReached = new java.util.concurrent.CountDownLatch(1);
@@ -2813,7 +2482,7 @@ class MessageServiceTest {
void hasQueuedDeliveryIsTrueAfterAnUndeliveredSendTimesOut() {
// Nothing ever delivers the message (never goes IDLE/BLOCKED), so the send times out with
// TIMED_OUT_QUEUED — same setup as sendTimesOutBeforeDeliveryIsQueuedNotWorking above.
MessageService.Reply r = messages.send(T, "never delivered", 50, null);
MessageService.Reply r = messages.send(T, "never delivered", 50);
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome());
assertTrue(messages.hasQueuedDelivery(T),
@@ -2825,7 +2494,7 @@ class MessageServiceTest {
@Test
void hasQueuedDeliveryClearsOnceTheTargetsNextDeliveryIsAccepted() throws Exception {
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50, null).outcome());
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50).outcome());
assertTrue(messages.hasQueuedDelivery(T));
// A fresh send accepts delivery (opens its own waiter) — the stale queued fact is cleared.
@@ -2842,7 +2511,7 @@ class MessageServiceTest {
@Test
void hasQueuedDeliveryClearsOnAbandon() {
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50, null).outcome());
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50).outcome());
assertTrue(messages.hasQueuedDelivery(T));
messages.abandon(T, "session released");
@@ -1,187 +0,0 @@
package dev.ltms.fleet.msg;
import org.junit.jupiter.api.Test;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.List;
import java.util.stream.Stream;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Pins that no file under {@code src/main/java} calls the fail-open
* {@link Rendezvous#open(String)} overload. That overload opens a forward waiter with no
* recorded {@link Rendezvous.Owner}, so the turn it opens can never be answered by anyone,
* not even the unnamed primary. Every production caller must go through
* {@link Rendezvous#open(String, Rendezvous.Owner)} and record an explicit owner.
*
* <p>This reads each file's own source text rather than reflecting on compiled bytecode, because
* the risk is a future one-word edit at a call site, not a missing overload.
*
* <p>The scan below finds a violation by its receiver, {@code rendezvous.open(}, classified by
* argument count. The one test method here points that exact scanner at a file known to hold
* many real one-argument calls before it ever looks at production, so a scanner that stops
* matching fails loudly on the known-positive case instead of leaving a clean production result
* looking like evidence it never produced.
*/
class RendezvousOpenUsageTest {
private static final Path PRODUCTION_SOURCE = Path.of("src/main/java");
private static final Path KNOWN_TEST_CALLER =
Path.of("src/test/java/dev/ltms/fleet/inject/CompletionResolverTest.java");
/**
* A pattern that cannot find a known one-argument {@code rendezvous.open(} call would also
* find none in production -- not because production is clean, but because the pattern does
* not match the text it is supposed to catch. That failure mode is exactly what let a
* {@code \b}-based regex read "no callers" under {@code git grep -E} when 81 real ones
* existed: {@code git grep} does not treat {@code \b} as a word boundary, so the pattern
* silently matched nothing anywhere, clean code and real calls alike. This test runs the
* known-positive check first, with the same scanning method the production check then
* depends on, so that mistake fails loudly here instead of reading as a clean result.
*/
@Test
void noProductionFileCallsTheSingleArgumentOpenOverload() throws IOException {
List<String> knownCalls = new ArrayList<>();
int knownFilesScanned = scanForOneArgOpenCalls(KNOWN_TEST_CALLER, knownCalls);
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "found
// nothing" having looked at nothing.
assertTrue(knownFilesScanned > 0, "control failed: the scan under " + KNOWN_TEST_CALLER
+ " visited zero .java files -- the path is wrong, so neither result below proves "
+ "anything");
// CONTROL: the scanner actually finds real one-argument rendezvous.open( calls when
// pointed at a file known to hold many. If this is not satisfied, the matching logic
// itself is broken, and the production result below is the scanner failing silently,
// not production code actually being clean.
assertTrue(knownCalls.size() >= 60, "control failed: the scanner found only "
+ knownCalls.size() + " one-argument rendezvous.open( call(s) in " + KNOWN_TEST_CALLER
+ ", which is known to hold many -- the matching logic itself is broken: " + knownCalls);
List<String> violations = new ArrayList<>();
int filesScanned = scanForOneArgOpenCalls(PRODUCTION_SOURCE, violations);
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "no
// violations found" having looked at nothing.
assertTrue(filesScanned > 0, "control failed: the scan under " + PRODUCTION_SOURCE
+ " visited zero .java files -- the path is wrong, so the absence of violations "
+ "below proves nothing");
assertTrue(violations.isEmpty(), "found a call to the fail-open Rendezvous.open(String) "
+ "overload, which opens a forward waiter with no recorded owner -- record an "
+ "explicit Rendezvous.Owner through open(String, Owner) instead: " + violations);
}
private static int scanForOneArgOpenCalls(Path root, List<String> sites) throws IOException {
List<Path> files = javaFiles(root);
for (Path file : files) {
scanFileForOpenCalls(file, sites);
}
return files.size();
}
private static List<Path> javaFiles(Path root) throws IOException {
try (Stream<Path> paths = Files.walk(root)) {
return paths.filter(p -> p.toString().endsWith(".java")).toList();
}
}
private static void scanFileForOpenCalls(Path file, List<String> sites) throws IOException {
String source = Files.readString(file);
String needle = "rendezvous.open(";
int from = 0;
int idx;
while ((idx = source.indexOf(needle, from)) >= 0) {
int argsStart = idx + needle.length();
String args = extractBalancedArgs(source, argsStart, file, idx);
int closeParenIndex = argsStart + args.length();
from = closeParenIndex + 1;
if (args.isBlank()) {
continue; // Rendezvous has no zero-argument open() -- a javadoc "rendezvous.open()" mention, not a call
}
if (topLevelCommaCount(args) == 0) {
sites.add(file + ":" + lineOf(source, idx) + " -- rendezvous.open(" + args.trim() + ")");
}
}
}
/**
* The text between {@code rendezvous.open(} and its matching close paren: balanced over
* nested calls, and never split by a paren or comma sitting inside a string or char literal.
*/
private static String extractBalancedArgs(String source, int start, Path file, int callIndex) {
int depth = 1;
boolean inString = false;
boolean inChar = false;
int i = start;
while (i < source.length()) {
char c = source.charAt(i);
if (inString) {
if (c == '\\') { i += 2; continue; }
if (c == '"') inString = false;
} else if (inChar) {
if (c == '\\') { i += 2; continue; }
if (c == '\'') inChar = false;
} else if (c == '"') {
inString = true;
} else if (c == '\'') {
inChar = true;
} else if (c == '(') {
depth++;
} else if (c == ')') {
depth--;
if (depth == 0) return source.substring(start, i);
}
i++;
}
throw new IllegalStateException(
"unbalanced parentheses scanning " + file + ":" + lineOf(source, callIndex));
}
/**
* Commas at paren/bracket/brace depth zero, skipping string and char literals -- the argument
* separators a human reader would see, not every comma character in the text.
*/
private static int topLevelCommaCount(String args) {
int depth = 0;
int commas = 0;
boolean inString = false;
boolean inChar = false;
int i = 0;
while (i < args.length()) {
char c = args.charAt(i);
if (inString) {
if (c == '\\') { i += 2; continue; }
if (c == '"') inString = false;
} else if (inChar) {
if (c == '\\') { i += 2; continue; }
if (c == '\'') inChar = false;
} else if (c == '"') {
inString = true;
} else if (c == '\'') {
inChar = true;
} else if (c == '(' || c == '[' || c == '{') {
depth++;
} else if (c == ')' || c == ']' || c == '}') {
depth--;
} else if (c == ',' && depth == 0) {
commas++;
}
i++;
}
return commas;
}
private static int lineOf(String source, int index) {
int line = 1;
for (int i = 0; i < index; i++) {
if (source.charAt(i) == '\n') line++;
}
return line;
}
}
@@ -132,127 +132,4 @@ class RendezvousTest {
assertNull(rendezvous.askSession(t.turnId()), "a closed ask is forgotten");
assertFalse(rendezvous.answerAsk(t.turnId(), "late"), "a closed ask can no longer be answered");
}
// ── fleetd #715: turn ownership ────────────────────────────────────────────────────────────
@Test
void openWithNoOwnerRecordsNoOwnerAndOpenWithAnOwnerRecordsIt() {
rendezvous.open(W);
assertNull(rendezvous.ownerOf(W), "the no-owner overload records no owner at all");
rendezvous.close(W, rendezvous.currentWaiter(W));
rendezvous.open(W, Rendezvous.Owner.of("term_lead"));
assertEquals(Rendezvous.Owner.of("term_lead"), rendezvous.ownerOf(W));
}
@Test
void ownerOfIsNullWhenNoWaiterIsOpen() {
assertNull(rendezvous.ownerOf(W), "no waiter open means no owner to report");
}
@Test
void openAskStampsTheForwardWaitersOwnerOntoTheFreshTurnOnly() {
rendezvous.open(W, Rendezvous.Owner.of("term_lead"));
Rendezvous.AskTicket fresh = rendezvous.openAsk(W);
assertTrue(fresh.fresh());
assertEquals(Rendezvous.Owner.of("term_lead"), rendezvous.askOwner(fresh.turnId()),
"a freshly-opened ask copies the forward waiter's current owner");
Rendezvous.AskTicket coalesced = rendezvous.openAsk(W);
assertFalse(coalesced.fresh());
assertEquals(fresh.turnId(), coalesced.turnId());
assertEquals(Rendezvous.Owner.of("term_lead"), rendezvous.askOwner(coalesced.turnId()),
"a coalesced duplicate ask rides the fresh owner's turn, unchanged");
}
@Test
void aSecondAskAfterTheFirstClosesStampsWhateverOwnerIsOpenAtThatLaterMoment() {
rendezvous.open(W, Rendezvous.Owner.of("term_lead"));
Rendezvous.AskTicket first = rendezvous.openAsk(W);
rendezvous.closeAsk(first.turnId());
// The forward waiter is reopened under a different owner before the second ask — mirrors
// answer() reopening with the owner it already checked, which can differ turn to turn.
rendezvous.close(W, rendezvous.currentWaiter(W));
rendezvous.open(W, Rendezvous.Owner.of("term_other"));
Rendezvous.AskTicket second = rendezvous.openAsk(W);
assertTrue(second.fresh());
assertEquals(Rendezvous.Owner.of("term_other"), rendezvous.askOwner(second.turnId()),
"a freshly-opened ask always copies whatever owner is open right now, not a stale one");
}
@Test
void askOwnerIsNullForAnUnknownOrLapsedTurn() {
assertNull(rendezvous.askOwner("no-such#1"));
Rendezvous.AskTicket t = rendezvous.openAsk(W);
rendezvous.closeAsk(t.turnId());
assertNull(rendezvous.askOwner(t.turnId()), "a closed ask no longer reports an owner");
}
// ── fleetd #729: per-boot nonce guards turnId against cross-instance reuse ────────────────
@Test
void twoInstancesMintDisjointTurnIds() {
Rendezvous other = new Rendezvous();
Rendezvous.AskTicket fromThis = rendezvous.openAsk(W);
Rendezvous.AskTicket fromOther = other.openAsk(W);
assertNotEquals(fromThis.turnId(), fromOther.turnId(),
"each instance mints its own id space, so even a first ask from each must differ");
}
@Test
void foreignInstanceTurnIdDoesNotResolve() {
Rendezvous other = new Rendezvous();
// `other` must reach the same sequence number as `rendezvous` (two asks each, the first
// closed so the second mints fresh), or this test passes against an empty map instead of
// against a colliding id.
Rendezvous.AskTicket firstFromThis = rendezvous.openAsk(W);
rendezvous.closeAsk(firstFromThis.turnId());
Rendezvous.AskTicket secondFromThis = rendezvous.openAsk(W);
Rendezvous.AskTicket firstFromOther = other.openAsk(W);
other.closeAsk(firstFromOther.turnId());
other.openAsk(W);
// control: the id resolves in the instance that minted it, so a false below cannot be
// explained by broken plumbing — only by the turnId being foreign to `other`.
assertTrue(rendezvous.answerAsk(secondFromThis.turnId(), "answer from this instance"),
"the minting instance must still resolve its own turnId");
assertFalse(other.answerAsk(secondFromThis.turnId(), "answer from other instance"),
"a turnId minted by a different instance must not resolve here");
}
@Test
void openAskStillCoalescesDuplicatesAndStillMintsDistinctIdsPerAsk() {
Rendezvous.AskTicket t1 = rendezvous.openAsk(W);
Rendezvous.AskTicket t2 = rendezvous.openAsk(W);
assertEquals(t1.turnId(), t2.turnId(),
"a second openAsk while one is open still coalesces onto the same turn");
assertFalse(t2.fresh(), "the coalesced ask is still reported as not fresh");
rendezvous.closeAsk(t1.turnId());
Rendezvous.AskTicket t3 = rendezvous.openAsk(W);
assertNotEquals(t1.turnId(), t3.turnId(), "two asks from the same session still get different turnIds");
}
@Test
void ownerPermitsIsFailClosedOnARecordAndThreeStatesAreDistinct() {
assertFalse(Rendezvous.Owner.permits(null, null),
"no owner on record refuses even the unnamed primary");
assertFalse(Rendezvous.Owner.permits(null, "worker:term_a"),
"no owner on record refuses a caller with an owner key too");
assertTrue(Rendezvous.Owner.permits(Rendezvous.Owner.UNNAMED_PRIMARY, null),
"the recorded unnamed primary matches a caller with a null owner key");
assertFalse(Rendezvous.Owner.permits(Rendezvous.Owner.UNNAMED_PRIMARY, "worker:term_a"),
"the recorded unnamed primary does not match another owner key");
assertTrue(Rendezvous.Owner.permits(Rendezvous.Owner.of("worker:term_a"), "worker:term_a"),
"an owner matches the same key");
assertFalse(Rendezvous.Owner.permits(Rendezvous.Owner.of("worker:term_a"), "worker:term_b"),
"an owner refuses a different key");
assertFalse(Rendezvous.Owner.permits(Rendezvous.Owner.of("worker:term_a"), null),
"a named owner refuses the unnamed primary");
}
}
@@ -1,14 +1,8 @@
package dev.ltms.fleet.rest;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.spi.ILoggingEvent;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.auth.Authz;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.auth.Role;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
@@ -21,11 +15,9 @@ import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.session.FakeWorktrees;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.testing.CapturedLog;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
@@ -36,15 +28,10 @@ import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Locale;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;
import java.util.function.Predicate;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
import java.util.stream.Collectors;
@@ -135,467 +122,12 @@ class FleetAppAuthTest {
assertEquals(Authz.Action.DRAIN, FleetApp.routeAction("GET /sessions/{id}/replies"));
assertEquals(Authz.Action.ASK, FleetApp.routeAction("POST /sessions/{id}/ask"));
for (String route : Set.of("GET /sessions", "GET /agents", "GET /members", "GET /profiles",
"GET /member-credentials")) {
"GET /member-credentials", "GET /sessions/{id}/status", "GET /tasks/{ticket}")) {
assertEquals(Authz.Action.READ, FleetApp.routeAction(route), route);
}
for (String route : Set.of("GET /sessions/{id}/status", "GET /tasks/{ticket}")) {
assertEquals(Authz.Action.TASK_READ, FleetApp.routeAction(route), route);
}
assertThrows(IllegalArgumentException.class, () -> FleetApp.routeAction("GET /healthz"));
}
// --- GET /tasks/{ticket} must not be the no-check overload ----------------------------------
/**
* {@code GET /tasks/{ticket}} must resolve its caller the same way {@code allow(...)} does
* and thread that owner key into {@link MessageService#poll(String, String)}, not the
* no-check overload that ignores who is asking.
*/
@Test
void theTaskStatusRouteActuallyThreadsTheCallersOwnerKeyIntoPoll() throws Exception {
String source = Files.readString(REST_SOURCE);
int start = source.indexOf("private void taskStatus(Context ctx) {");
assertTrue(start >= 0, "could not find taskStatus in " + REST_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("private static void herdrError(Context ctx, HerdrException e) {", start);
assertTrue(end > start, "could not find the method declared after taskStatus to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does call messages.poll(...) -- if this fails, the
// anchors above moved and the assertions below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("messages.poll("),
"control failed: the scraped taskStatus block contains no messages.poll( call at all "
+ "-- the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("caller.ownerKey()"),
"the taskStatus route must thread the resolved caller's owner key into messages.poll(...), "
+ "not the no-check overload -- block: " + handlerBlock);
assertTrue(handlerBlock.contains("ctx.attribute(CALLER)"),
"the taskStatus route must resolve its caller the same way allow(...) does, not via a "
+ "second, separate resolution path -- block: " + handlerBlock);
}
/**
* {@code GET /sessions/{id}/status} must resolve its caller the same way {@code allow(...)}
* does and thread that owner key into {@link MessageService#pendingAsk(String, String)}, not
* the no-check overload that ignores who is asking.
*/
@Test
void theSessionStatusRouteActuallyThreadsTheCallersOwnerKeyIntoPendingAsk() throws Exception {
String source = Files.readString(REST_SOURCE);
int start = source.indexOf("private void sessionStatus(Context ctx) {");
assertTrue(start >= 0, "could not find sessionStatus in " + REST_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("private void taskStatus(Context ctx) {", start);
assertTrue(end > start, "could not find the method declared after sessionStatus to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does call messages.pendingAsk(...) -- if this fails,
// the anchors above moved and the assertions below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("messages.pendingAsk("),
"control failed: the scraped sessionStatus block contains no messages.pendingAsk( "
+ "call at all -- the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("caller.ownerKey()"),
"the sessionStatus route must thread the resolved caller's owner key into "
+ "messages.pendingAsk(...), not the no-check overload -- block: " + handlerBlock);
assertTrue(handlerBlock.contains("ctx.attribute(CALLER)"),
"the sessionStatus route must resolve its caller the same way allow(...) does, not via "
+ "a second, separate resolution path -- block: " + handlerBlock);
}
/**
* {@code GET /tasks/{ticket}} refuses a worker whose terminal did not create the ticket,
* while the creating worker and the unnamed primary both still read it. The ticket is minted
* directly on the shared {@link MessageService}, the same way {@code MessageServiceTest}
* drives {@link MessageService#poll(String, String)}, so this exercises only the REST poll
* route's own handling of the ownership already recorded on the ticket.
*/
@Test
void restPollRefusesADifferentWorkerButAllowsTheCreatorAndTheUnnamedPrimary() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
MessageService messages = new MessageService(agents, injector, new Rendezvous());
Javalin creatorApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID); // -> term_a
Javalin otherWorkerApp = startOnSharedService(messages, herdr, 9001L); // -> term_shell
Javalin primaryApp = startOnSharedService(messages, herdr, 999_999L); // no pane -> primary
try {
String ticket = messages.sendAsync("term_a", "long task", null,
Principal.worker("term_a", FakeHerdr.WORKER_PID));
HttpResponse<String> refused = send(otherWorkerApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, refused.statusCode());
assertTrue(refused.body().contains("forbidden"),
"a different worker's terminal must be refused, not shown the ticket: " + refused.body());
assertFalse(refused.body().contains("\"reply\""),
"a refusal must never carry reply text: " + refused.body());
HttpResponse<String> own = send(creatorApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, own.statusCode());
assertFalse(own.body().contains("forbidden"),
"the creating worker must read its own ticket: " + own.body());
HttpResponse<String> primary = send(primaryApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, primary.statusCode());
assertFalse(primary.body().contains("forbidden"),
"the unnamed primary must read any ticket: " + primary.body());
} finally {
creatorApp.stop();
otherWorkerApp.stop();
primaryApp.stop();
}
}
/**
* {@code POST /sessions/{id}/message} with {@code wait:false} must record the creating
* caller's own terminal on the ticket it returns, so that caller can still poll its own
* ticket over REST, while a different terminal is refused.
*/
@Test
void restSendAsyncRecordsTheCreatingCallersTerminalSoItCanStillPollItsOwnTicket() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
MessageService messages = new MessageService(agents, injector, new Rendezvous());
Javalin leadApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID, Map.of("term_a", "lead-x"));
Javalin otherWorkerApp = startOnSharedService(messages, herdr, 9001L); // -> term_shell
try {
ObjectMapper mapper = new ObjectMapper();
HttpResponse<String> created = send(leadApp.port(), "POST", "/sessions/term_a/message",
"{\"content\":\"long task\",\"wait\":false}", null);
assertEquals(202, created.statusCode(), created.body());
String ticket = mapper.readTree(created.body()).path("ticket").asText(null);
assertNotNull(ticket, "the accepted response carried no ticket: " + created.body());
HttpResponse<String> own = send(leadApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, own.statusCode());
assertFalse(own.body().contains("forbidden"),
"the session that created the ticket over REST must be able to poll it: " + own.body());
HttpResponse<String> refused = send(otherWorkerApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, refused.statusCode());
assertTrue(refused.body().contains("forbidden: this ticket was created by a different session"),
"a different terminal must still be refused with the ownership detail, not some "
+ "other rejection: " + refused.body());
} finally {
leadApp.stop();
otherWorkerApp.stop();
}
}
/**
* {@code GET /sessions/{id}/status} shows a worker's pending {@code fleet_ask} question, its
* {@code turnId} and its ticket only to the caller whose owner key created that delegation, or
* to the unnamed primary. A different caller still sees the base status line, but none of the
* pending-ask fields.
*/
@Test
void restStatusGatesThePendingAskFieldsByTheDelegationsCreatorOwner() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
Rendezvous rendezvous = new Rendezvous();
MessageService messages = new MessageService(agents, injector, rendezvous);
Javalin creatorApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID); // -> term_a
Javalin otherWorkerApp = startOnSharedService(messages, herdr, 9001L); // -> term_shell
Javalin primaryApp = startOnSharedService(messages, herdr, 999_999L); // no pane -> primary
try {
ObjectMapper mapper = new ObjectMapper();
String ticket = messages.sendAsync("term_target", "task that asks", null,
Principal.worker("term_a", FakeHerdr.WORKER_PID));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_target"), "sendAsync should have opened its rendezvous waiter");
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
() -> messages.ask("term_target", "which config file?", 5000));
MessageService.TaskView asking;
deadline = System.currentTimeMillis() + 3000;
do {
asking = messages.poll(ticket, null);
Thread.sleep(5);
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
assertEquals(MessageService.Phase.ASKING, asking.phase());
String turnId = asking.turnId();
JsonNode other = mapper.readTree(
send(otherWorkerApp.port(), "GET", "/sessions/term_target/status", null, null).body());
assertEquals("idle", other.get("status").asText(), "the base status must still be shown");
assertFalse(other.has("question"), "a non-creating caller must not see the question: " + other);
assertFalse(other.has("turnId"), "a non-creating caller must not see the turnId: " + other);
assertFalse(other.has("ticket"), "a non-creating caller must not see the ticket: " + other);
JsonNode own = mapper.readTree(
send(creatorApp.port(), "GET", "/sessions/term_target/status", null, null).body());
assertEquals("which config file?", own.get("question").asText(), "the creator must see the question");
assertEquals(turnId, own.get("turnId").asText(), "the creator must see the turnId");
assertEquals(ticket, own.get("ticket").asText(), "the creator must see the ticket");
JsonNode primary = mapper.readTree(
send(primaryApp.port(), "GET", "/sessions/term_target/status", null, null).body());
assertEquals("which config file?", primary.get("question").asText(),
"a caller with no terminal (the unnamed primary) must see the question");
// Clean up the still-open ask so the background thread does not linger past the test.
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(turnId, "config.yaml", 5000, "worker:term_a"));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.resolve("term_target", "done"));
answer.get(5, TimeUnit.SECONDS);
} finally {
creatorApp.stop();
otherWorkerApp.stop();
primaryApp.stop();
}
}
/**
* {@code POST /sessions/{id}/message} with a {@code turnId} refuses a caller whose terminal
* did not open the turn, even when that caller otherwise holds ANSWER rights, and leaves the
* turn open for the real owner to resolve. Covers the REST adapter's ANSWER gate for an
* async-send ({@code wait:false}) delegation.
*/
@Test
void restAnswerIsRefusedForADifferentCallerOnAnAsyncSendDelegationButTheRealOwnerSucceeds() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
Rendezvous rendezvous = new Rendezvous();
MessageService messages = new MessageService(agents, injector, rendezvous);
Javalin ownerApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID, Map.of("term_a", "lead-owner"));
Javalin attackerApp = startOnSharedService(messages, herdr, 9001L, Map.of("term_shell", "lead-attacker"));
try {
ObjectMapper mapper = new ObjectMapper();
HttpResponse<String> created = send(ownerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"long task\",\"wait\":false}", null);
assertEquals(202, created.statusCode(), created.body());
String ticket = mapper.readTree(created.body()).path("ticket").asText(null);
assertNotNull(ticket, "the accepted response carried no ticket: " + created.body());
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_target"), "the async send should have opened its rendezvous waiter");
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
() -> messages.ask("term_target", "which config file?", 5000));
MessageService.TaskView asking;
deadline = System.currentTimeMillis() + 3000;
do {
asking = messages.poll(ticket, null);
Thread.sleep(5);
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
assertEquals(MessageService.Phase.ASKING, asking.phase());
String turnId = asking.turnId();
HttpResponse<String> hijacked = send(attackerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"hijack\",\"turnId\":\"" + turnId + "\"}", null);
assertEquals(403, hijacked.statusCode(), hijacked.body());
assertTrue(hijacked.body().contains("not_turn_owner"),
"a different caller's answer must be refused as not_turn_owner: " + hijacked.body());
assertEquals("term_target", rendezvous.askSession(turnId),
"a refused answer must leave the ask turn open");
CompletableFuture<HttpResponse<String>> owned = CompletableFuture.supplyAsync(() -> {
try {
return send(ownerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"config.yaml\",\"turnId\":\"" + turnId + "\"}", null);
} catch (Exception e) {
throw new RuntimeException(e);
}
});
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.resolve("term_target", "done"));
HttpResponse<String> ownedResponse = owned.get(5, TimeUnit.SECONDS);
assertEquals(200, ownedResponse.statusCode(), ownedResponse.body());
assertEquals("done", mapper.readTree(ownedResponse.body()).path("reply").asText(null),
"the real owner's answer must resolve the worker's turn");
} finally {
ownerApp.stop();
attackerApp.stop();
}
}
/**
* As above, but the delegation is a blocking send ({@code wait:true}) instead of a ticket —
* the owning caller's own HTTP request is the one that surfaces the worker's question and
* later carries the real answer. Covers the REST adapter's ANSWER gate for a blocking-send
* delegation.
*/
@Test
void restAnswerIsRefusedForADifferentCallerOnABlockingSendDelegationButTheRealOwnerSucceeds() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
Rendezvous rendezvous = new Rendezvous();
MessageService messages = new MessageService(agents, injector, rendezvous);
Javalin ownerApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID, Map.of("term_a", "lead-owner"));
Javalin attackerApp = startOnSharedService(messages, herdr, 9001L, Map.of("term_shell", "lead-attacker"));
try {
ObjectMapper mapper = new ObjectMapper();
CompletableFuture<HttpResponse<String>> blocking = CompletableFuture.supplyAsync(() -> {
try {
return send(ownerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"long task\",\"wait\":true,\"timeoutMs\":5000}", null);
} catch (Exception e) {
throw new RuntimeException(e);
}
});
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_target"), "the blocking send should have opened its rendezvous waiter");
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
() -> messages.ask("term_target", "which config file?", 5000));
HttpResponse<String> questionResponse = blocking.get(5, TimeUnit.SECONDS);
assertEquals(202, questionResponse.statusCode(), questionResponse.body());
JsonNode question = mapper.readTree(questionResponse.body());
assertEquals("question", question.get("status").asText());
String turnId = question.get("turnId").asText();
HttpResponse<String> hijacked = send(attackerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"hijack\",\"turnId\":\"" + turnId + "\"}", null);
assertEquals(403, hijacked.statusCode(), hijacked.body());
assertTrue(hijacked.body().contains("not_turn_owner"),
"a different caller's answer must be refused as not_turn_owner: " + hijacked.body());
assertEquals("term_target", rendezvous.askSession(turnId),
"a refused answer must leave the ask turn open");
CompletableFuture<HttpResponse<String>> owned = CompletableFuture.supplyAsync(() -> {
try {
return send(ownerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"config.yaml\",\"turnId\":\"" + turnId + "\"}", null);
} catch (Exception e) {
throw new RuntimeException(e);
}
});
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.resolve("term_target", "done"));
HttpResponse<String> ownedResponse = owned.get(5, TimeUnit.SECONDS);
assertEquals(200, ownedResponse.statusCode(), ownedResponse.body());
assertEquals("done", mapper.readTree(ownedResponse.body()).path("reply").asText(null),
"the real owner's answer must resolve the worker's turn");
} finally {
ownerApp.stop();
attackerApp.stop();
}
}
/**
* As {@link #start}, but shares {@code messages} and {@code herdr} across several app
* instances bound to different pids, each returned as its own started {@link Javalin} rather
* than through the shared {@code app} field, so several differently-resolved callers can
* poll the same ticket.
*/
private Javalin startOnSharedService(MessageService messages, FakeHerdr herdr, long pid) {
return startOnSharedService(messages, herdr, pid, Map.of());
}
/**
* As {@link #startOnSharedService(MessageService, FakeHerdr, long)}, but {@code leadTerminals}
* resolves the given pid's terminal to a named lead (a caller with SEND permission) instead of
* a plain worker, for a test that needs a terminal-bearing caller able to create a ticket.
*
* <p>Every connecting pane not already claimed by {@code leadTerminals} is wired into the live
* roster as a spawned worker, so a caller's resolved role matches what its own test expects:
* a {@link Role#WORKER}, never the unconfigured-pane {@link Role#OBSERVER} floor a roster-less
* resolver would otherwise fall to.
*/
private Javalin startOnSharedService(MessageService messages, FakeHerdr herdr, long pid,
Map<String, String> leadTerminals) {
FleetConfig.Profile wcfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
AgentControl agents = new AgentControl(herdr);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
k -> "FLEETD_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
() -> leadTerminals, new MemberRegistry(null),
t -> leadTerminals.containsKey(t) ? null : MemberRole.DEV, Map::of);
Metrics appMetrics = FleetMetrics.create(sessions, new dev.ltms.fleet.msg.InMemoryReplyInbox());
return new FleetApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
callers, appMetrics).build().start("127.0.0.1", 0);
}
/**
* fleetd #669 Unit A: {@code POST /sessions/{id}/message} is two call shapes behind one route,
* mirroring {@code fleet_send}'s MCP-side split into {@link Authz.Action#SEND} and {@link
* Authz.Action#ANSWER} ({@code FleetMcp#sendAction}). The route never carries a {@code coordId}
* shape — that peer-lead route is MCP-only — so only these two apply here.
*/
@Test
void theMessageRouteIsASendWithNoTurnIdAndAnAnswerWithOne() {
assertEquals(Authz.Action.SEND, FleetApp.routeAction("POST /sessions/{id}/message", null));
assertEquals(Authz.Action.SEND, FleetApp.routeAction("POST /sessions/{id}/message", " "));
assertEquals(Authz.Action.ANSWER, FleetApp.routeAction("POST /sessions/{id}/message", "turn-1"));
}
/**
* fleetd #689: {@code answerGatePasses} is the second, conditional gate behind {@code
* sendMessage}'s coarse {@link Authz.Action#SEND} check. With the {@code ANSWER} grant denied,
* a {@code turnId}-bearing request is refused while a plain one still passes — and the denied
* permit is queried only for the {@code turnId} case, never for the plain one, which is what
* proves this is a genuinely separate, conditional check rather than the {@code SEND} check
* renamed or an unconditional call whose result is ignored. Flipping only the {@code ANSWER}
* grant to allowed then flips only the {@code turnId} shape's outcome.
*/
@Test
void answerGatePassesOnlyWhenTurnIdAbsentOrAnswerGranted() {
List<Authz.Action> queried = new ArrayList<>();
Predicate<Authz.Action> denyAnswer = action -> {
queried.add(action);
return false;
};
assertFalse(FleetApp.answerGatePasses("turn-1", denyAnswer),
"ANSWER denied ⇒ the turnId shape is refused");
assertEquals(List.of(Authz.Action.ANSWER), queried,
"the ANSWER grant, specifically, must be the one consulted");
queried.clear();
assertTrue(FleetApp.answerGatePasses(null, denyAnswer),
"no turnId ⇒ the plain shape passes even though ANSWER is denied");
assertTrue(FleetApp.answerGatePasses(" ", denyAnswer), "a blank turnId is treated as absent");
assertEquals(List.of(), queried, "the plain shape must never consult the permit at all");
assertTrue(FleetApp.answerGatePasses("turn-1", action -> true),
"flipping only the ANSWER grant to allowed flips only the turnId shape's outcome");
}
private static Set<String> routesTheServerRegisters() {
try {
String source = Files.readString(REST_SOURCE).lines()
@@ -615,74 +147,6 @@ class FleetAppAuthTest {
}
}
/**
* {@code permitsFor} is the exact decision {@link FleetApp#allow} makes, passing the real
* production classifier rather than a test-supplied one — built from an empty {@link
* CallerResolver}, so no terminal is recognised as a configured lead or collaborator and a
* collaborator's SEND is refused through the REST gate.
*/
@Test
void aCollaboratorMayNotSendOverRestWithTheRealProductionClassifier() {
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(new FakeHerdr()), _ -> 700L);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, new MemberRegistry(null));
Principal collaborator = Principal.collaborator("ops", "term_collab", 700);
assertFalse(FleetApp.permitsFor(collaborator, Authz.Action.SEND, "term_lead",
callers.knownLeadOrCollaborator()),
"no terminal is recognised as a lead or collaborator yet");
}
/**
* fleetd #669 Unit D: wires a real {@link CallerResolver} with a known lead and a known
* collaborator tab, and a spawned member's own terminal recognised by neither map. A
* collaborator's SEND reaches the known lead and the known collaborator, and is refused for the
* spawned member's terminal — over the REST route, not just the unit-level classifier, so a
* test covering only MCP cannot leave this route open.
*/
@Test
void aCollaboratorMaySendToAKnownLeadOrCollaboratorButNotToASpawnedMembersTerminalOverRest() throws Exception {
int port = startWithRealClassifier(FakeHerdr.WORKER_PID,
Map.of("term_lead_known", "lead-x"), Map.of("term_a", "ops2"));
HttpResponse<String> toLead = send(port, "POST", "/sessions/term_lead_known/message",
"{\"content\":\"hi\",\"wait\":false}", null);
assertEquals(202, toLead.statusCode(), toLead.body());
HttpResponse<String> toSpawnedMembersTerminal = send(port, "POST", "/sessions/term_worker/message",
"{\"content\":\"hi\",\"wait\":false}", null);
assertEquals(403, toSpawnedMembersTerminal.statusCode(), toSpawnedMembersTerminal.body());
}
/**
* As {@link #start}, but with explicit lead/collaborator maps and no spawned-member roster, so
* a test can wire the real {@link CallerResolver#knownLeadOrCollaborator()} classifier instead
* of the default empty one.
*/
private int startWithRealClassifier(long pid, Map<String, String> leadTerminals,
Map<String, String> collaboratorTerminals) {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile wcfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
AgentControl agents = new AgentControl(herdr);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
k -> "FLEETD_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
Injector injector = new Injector(agents);
MessageService messages = new MessageService(agents, injector, new Rendezvous());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
() -> leadTerminals, new MemberRegistry(null), t -> null, () -> collaboratorTerminals);
metrics = FleetMetrics.create(sessions, new dev.ltms.fleet.msg.InMemoryReplyInbox());
app = new FleetApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
callers, metrics).build().start("127.0.0.1", 0);
return app.port();
}
// --- loopback-trust: the caller is the primary -------------------------------------------
@Test
@@ -728,107 +192,6 @@ class FleetAppAuthTest {
"draining an inbox is the primary's collection step");
}
/**
* fleetd #669 Unit A: the {@code turnId} shape of {@code POST /sessions/{id}/message} maps to
* {@link Authz.Action#ANSWER}, not the plain {@link Authz.Action#SEND} the test above drives —
* a worker must stay refused on this shape too, exactly as it was refused on the one undivided
* action before the split.
*/
@Test
void aWorkerMayNotAnswerAnotherSessionsBlockedQuestionOverRest() throws Exception {
int port = start(FakeHerdr.WORKER_PID, false, null);
assertEquals(403, send(port, "POST", "/sessions/term_b/message",
"{\"turnId\":\"turn-1\",\"content\":\"hi\"}", null).statusCode(),
"resolving another session's blocked question would be a worker escalating too");
}
/**
* fleetd #689: a caller refused the coarse {@link Authz.Action#SEND} grant is refused on
* {@code SEND} specifically, even on the {@code turnId}-bearing shape that otherwise raises
* the check to {@link Authz.Action#ANSWER} — proving {@code turnId} was never read from the
* body before the refusal (reading it would have changed which action is named in the 403).
* The same caller refused with no body at all gets the identical detail, which could not hold
* if the decision depended on anything read from the body. Control: a caller who IS granted
* reaches past the gate and the body is used normally.
*/
@Test
void aDeniedCallerIsRefusedOnSendEvenWithATurnIdBodyAndNeverReadsTheBody() throws Exception {
int workerPort = start(FakeHerdr.WORKER_PID, false, null); // denied: not primary/architect
HttpResponse<String> withTurnId = send(workerPort, "POST", "/sessions/term_b/message",
"{\"turnId\":\"turn-1\",\"content\":\"hi\"}", null);
assertEquals(403, withTurnId.statusCode());
assertTrue(withTurnId.body().contains("may not SEND"),
"the SEND check must be the one that fired, not ANSWER — ANSWER would only be "
+ "reachable by having already read turnId out of the body");
HttpResponse<String> noBody = send(workerPort, "POST", "/sessions/term_b/message", null, null);
assertEquals(403, noBody.statusCode());
assertTrue(noBody.body().contains("may not SEND"),
"refused identically with no body at all — the refusal cannot depend on body content");
// Control: a primary IS granted SEND, so the same turnId body is read and acted on —
// reaching messages.answer, which reports this unknown turnId as a stale one.
int primaryPort = start(999_999, false, null);
HttpResponse<String> granted = send(primaryPort, "POST", "/sessions/term_b/message",
"{\"turnId\":\"turn-1\",\"content\":\"hi\"}", null);
assertEquals(409, granted.statusCode());
assertTrue(granted.body().contains("stale_turn"), "a granted caller's body IS read and acted on");
}
/**
* fleetd #689 (ticket comment 18353): the only place {@code sendMessage}'s call to {@code
* answerGatePasses} is observable is the audit trail — {@code allow()} logs an {@code
* "allowed"} entry for every granted action except {@code READ}/{@code METRICS}/{@code
* TASK_READ}, and {@code ANSWER} is none of those. A granted {@code turnId} request must
* therefore log both a {@code SEND} and an {@code ANSWER} entry; a granted plain request must
* log {@code SEND} alone. A unit test of the extracted helper pins the helper; this pins the
* call site — deleting the {@code answerGatePasses} call from {@code sendMessage} leaves the
* helper's own test green but turns this one red.
*/
@Test
void aGrantedTurnIdRequestAuditsBothSendAndAnswerButAPlainRequestAuditsSendAlone() throws Exception {
int port = start(999_999, false, null); // primary: granted both SEND and ANSWER
ObjectMapper mapper = new ObjectMapper();
try (CapturedLog audit = CapturedLog.at("audit", Level.INFO)) {
send(port, "POST", "/sessions/term_b/message",
"{\"turnId\":\"turn-1\",\"content\":\"hi\"}", null);
List<String> allowed = allowedActions(audit, mapper);
assertTrue(allowed.contains("SEND"),
"a turnId request must still clear the coarse SEND grant first");
assertTrue(allowed.contains("ANSWER"),
"a turnId request must ALSO clear the ANSWER grant — this is the call site itself");
}
try (CapturedLog audit = CapturedLog.at("audit", Level.INFO)) {
send(port, "POST", "/sessions/term_b/message",
"{\"content\":\"hi\",\"timeoutMs\":50}", null);
List<String> allowed = allowedActions(audit, mapper);
assertEquals(List.of("SEND"), allowed,
"a plain request must log SEND and nothing else — ANSWER is conditional on "
+ "turnId, not something every request happens to log");
}
}
private static List<String> allowedActions(CapturedLog audit, ObjectMapper mapper) {
return audit.events().stream()
.map(ILoggingEvent::getFormattedMessage)
.map(line -> {
try {
return mapper.readTree(line);
} catch (Exception e) {
throw new AssertionError("audit line is not valid JSON: " + line, e);
}
})
.filter(n -> "allowed".equals(n.path("outcome").asText()))
.map(n -> n.path("action").asText())
.toList();
}
// --- token mode ---------------------------------------------------------------------------
@Test
@@ -675,26 +675,6 @@ class FleetAppTest {
assertEquals(400, postMessage(port, "{}").statusCode());
}
/**
* fleetd #689: a body that fails to parse is rejected with 400 before {@code turnId} is ever
* read from it, so it reaches neither {@code messages.answer} (which needs a {@code turnId})
* nor {@code messages.send} — confirmed here for {@code send} by the fake agent's idle status,
* which would otherwise make an immediate {@code agent.prompt} delivery observable.
*/
@Test
void malformedBodyReturns400AndNeverReachesSendOrAnswer() throws Exception {
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // idle ⇒ send would deliver right away if reached
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = postMessage(port, "not json at all");
assertEquals(400, res.statusCode());
JsonNode err = mapper.readTree(res.body());
assertEquals("bad_request", err.get("error").asText());
assertEquals("body must be JSON", err.get("detail").asText());
assertFalse(herdr.called("agent.prompt"),
"a malformed body must never reach messages.send's delivery");
}
@Test
void sessionStatusReportsLiveAgentStatus() throws Exception {
FakeHerdr herdr = new FakeHerdr().agentStatus("blocked");
@@ -1,92 +0,0 @@
package dev.ltms.fleet.session;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.PlacementDecision;
import java.util.List;
import java.util.Set;
/**
* {@link PeerLauncher} decorator that marks presence for a spawned terminal before returning its
* handle to the caller — the contact-then-register ordering fleetd #722 covers, where the
* terminal's MCP contact lands before {@link SessionManager#acquire} runs its own
* {@code registry.put}. The presence view is set after construction, via {@link #presence},
* because it is owned by the {@link SessionManager} this launcher is passed into.
*/
final class PresenceRacingLauncher implements PeerLauncher {
private final PeerLauncher delegate;
volatile MemberPresence presence;
PresenceRacingLauncher(PeerLauncher delegate) {
this.delegate = delegate;
}
@Override
public PeerHandle spawn(SpawnRequest req) {
PeerHandle handle = delegate.spawn(req);
presence.markPresent(handle.terminalId());
return handle;
}
@Override
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
PeerHandle handle = delegate.spawn(req, decision);
presence.markPresent(handle.terminalId());
return handle;
}
@Override
public Set<Capability> capabilities() {
return delegate.capabilities();
}
@Override
public Set<Capability> capabilitiesFor(String profileName) {
return delegate.capabilitiesFor(profileName);
}
@Override
public Set<String> profiles() {
return delegate.profiles();
}
@Override
public String defaultProfile() {
return delegate.defaultProfile();
}
@Override
public String effectiveCwd(SpawnRequest req) {
return delegate.effectiveCwd(req);
}
@Override
public List<String> parityOverlay(String profileName) {
return delegate.parityOverlay(profileName);
}
@Override
public List<?> list() {
return delegate.list();
}
@Override
public int reapOrphanWorkers() {
return delegate.reapOrphanWorkers();
}
@Override
public void stop(String id) {
delegate.stop(id);
}
@Override
public boolean clearContext(String id) {
return delegate.clearContext(id);
}
}
@@ -422,53 +422,6 @@ class SessionManagerTest {
"turn completion moves BUSY → DONE");
}
// --- fleetd #722: registration and presence must reach READY whichever lands first --------
@Test
void registerThenContactReachesReadyForPlainSpawn() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"a presence contact that arrives after registration reaches READY");
}
@Test
void contactThenRegisterStillReachesReadyForPlainSpawn() {
// The racing launcher marks presence for the spawned terminal from inside spawn() —
// before SessionManager.acquire's own registry.put runs — modeling an MCP contact that
// lands in that window.
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
PresenceRacingLauncher race = new PresenceRacingLauncher(workers);
SessionManager sessions = new SessionManager(race);
race.presence = sessions.asPresence();
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"a presence contact that lands before registry.put must still reach READY");
}
@Test
void aTerminalNeverMarkedPresentStaysSpawningAfterRegistration() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
assertEquals(MemberSession.State.SPAWNING, sessions.get(session.paneId()).orElseThrow().state(),
"registration alone must not advance a terminal that was never marked present");
}
@Test
void releaseTearsDownWorkerAndRemovesFromRosterAndIsIdempotent() {
FakeHerdr herdr = new FakeHerdr();
@@ -1452,105 +1405,6 @@ class SessionManagerTest {
+ "dirty check threw");
}
// --- fleetd #736: a release must forget the member's presence entry, not just its registry
// row ---------------------------------------------------------------------------------------
@Test
void releaseByPaneIdForgetsThePresenceEntry() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
assertTrue(sessions.asPresence().isPresent(terminal), "present before the release");
sessions.release(session.paneId());
assertFalse(sessions.asPresence().isPresent(terminal),
"release must forget the terminal's presence, not just remove its registry row");
}
@Test
void reapIdleForgetsThePresenceEntryToo() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
assertTrue(sessions.asPresence().isPresent(terminal), "present before the reap");
clock[0] = 11;
assertEquals(1, sessions.reapIdle(10), "READY session past TTL is reaped");
assertFalse(sessions.asPresence().isPresent(terminal),
"the idle-reap release path (releaseIfCurrent) goes through the same teardown "
+ "funnel as an explicit release, so it must forget presence too");
}
@Test
void shutdownDrainAlsoForgetsThePresenceEntry() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
assertTrue(sessions.asPresence().isPresent(terminal), "present before the drain");
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
assertFalse(sessions.asPresence().isPresent(terminal),
"a shutdown drain still ends the member's process, so presence must be cleared "
+ "exactly as it is for any other release cause");
}
@Test
void releaseOfAnUnknownPaneIdDoesNotThrow() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
assertDoesNotThrow(() -> sessions.release("no-such-pane"),
"releasing a pane id that was never registered must be a no-op, not a throw");
}
@Test
void releaseStillForgetsPresenceWhenDirtyCheckThrows() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("fleetd-736", null));
String terminal = s.terminalId();
sessions.asPresence().markPresent(terminal);
worktrees.failHasUncommittedWith(new WorktreeException("git status exited 128"));
assertDoesNotThrow(() -> sessions.release(s.paneId()),
"a throwing dirty check must not abort the release");
assertFalse(sessions.asPresence().isPresent(terminal),
"presence must be forgotten even when the dirty check throws, which pins the "
+ "forget call to the finally block that runs no matter what happened above");
}
@Test
void releaseLeavesADifferentStillLiveMembersPresenceUntouched() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession released = sessions.acquire("ltms-local", null, "/caller/a", "ownerA");
MemberSession stillLive = sessions.acquire("ltms-local", null, "/caller/b", "ownerB");
sessions.asPresence().markPresent(released.terminalId());
sessions.asPresence().markPresent(stillLive.terminalId());
assertTrue(sessions.asPresence().isPresent(stillLive.terminalId()),
"present before the release of the other member");
sessions.release(released.paneId());
assertFalse(sessions.asPresence().isPresent(released.terminalId()),
"the released terminal is forgotten");
assertTrue(sessions.asPresence().isPresent(stillLive.terminalId()),
"a still-live member's presence must survive an unrelated release");
}
// --- fleetd #316: the dirty check must be re-taken after the worker is stopped, not trusted
// stale from before it ------------------------------------------------------------------------
@@ -2582,160 +2436,4 @@ class SessionManagerTest {
return "threw:" + e.getClass().getName() + ":" + e.getMessage();
}
}
// --- fleetd #702: spawnedMemberRole must still answer for a pane mid-teardown ----------------
/**
* A {@link Worktrees} test double whose {@code hasUncommitted} runs an injected hook before
* answering. This is what lets a test resolve a releasing pane's terminal from inside the
* window {@link SessionManager#release} opens between removing the registry entry and
* actually stopping the pane — the hook runs synchronously on the release call's own thread,
* at the exact point {@code release} shells out to {@code git status}, so there is no sleep
* and no race to land in it.
*/
private static final class HookedWorktrees implements Worktrees {
private Runnable hook;
private boolean dirty = false;
HookedWorktrees onHasUncommitted(Runnable hook) {
this.hook = hook;
return this;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
return "/wt/" + branch.replace('/', '_');
}
@Override
public void remove(String repoRoot, String worktreePath) {
}
@Override
public void deleteBranch(String repoRoot, String branch) {
}
@Override
public boolean hasUncommitted(String worktreePath) {
if (hook != null) {
hook.run();
}
return dirty;
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
}
@Override
public String repoRoot(String cwd) {
return "/repo";
}
@Override
public java.util.Optional<String> snapshot(String worktreePath, String branch, String message) {
return java.util.Optional.empty();
}
@Override
public WipRefStats wipRefs(String repoRoot) {
return new WipRefStats(0, 0L);
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
return 0;
}
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
}
}
@Test
void spawnedMemberRoleResolvesTheLiveRegisteredRole() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession s = sessions.acquire("ltms-local", null, "/caller", null);
assertEquals(s.role(), sessions.spawnedMemberRole(s.terminalId()),
"a registered session resolves to its own role");
assertNull(sessions.spawnedMemberRole("term_unknown"),
"a terminal with no session at all resolves to null");
}
@Test
void spawnedMemberRoleIsNullOnceReleaseFullyCompletes() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, new HookedWorktrees());
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("fleetd-702a", null));
sessions.release(s.paneId());
assertNull(sessions.spawnedMemberRole(s.terminalId()),
"once release has fully finished, the terminal is neither registered nor releasing");
}
@Test
void spawnedMemberRoleStillAnswersBetweenTheRegistryRemovalAndThePaneStop() {
FakeHerdr herdr = new FakeHerdr();
java.util.concurrent.atomic.AtomicReference<MemberRole> duringWindow = new java.util.concurrent.atomic.AtomicReference<>();
HookedWorktrees worktrees = new HookedWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("fleetd-702b", null));
worktrees.onHasUncommitted(() -> {
// This runs from INSIDE release()'s git-status shell-out: the registry entry is
// already gone, but the pane has not stopped yet — a real call landing inside the
// exact window fleetd #702 reports, so no sleep and no race is needed to reach it.
assertTrue(sessions.get(s.paneId()).isEmpty(),
"sanity: the registry entry is already gone at this point");
duringWindow.set(sessions.spawnedMemberRole(s.terminalId()));
});
sessions.release(s.paneId());
assertEquals(s.role(), duringWindow.get(),
"spawnedMemberRole must still answer the live role while the pane is mid-teardown, "
+ "not only while the session is still in the registry");
assertNull(sessions.spawnedMemberRole(s.terminalId()),
"and once release has fully finished, the window is closed too");
}
/**
* {@code Releasing.enter} and {@code Releasing.leave} are called directly here, never through
* {@link SessionManager#release} or {@link SessionManager#spawnedMemberRole}, so a test naming
* one of them exercises only that one — a regression in the other can never hide behind it.
*/
@Test
void releasingLeaveStepsDownADepthGreaterThanOneInsteadOfRemovingIt() {
SessionManager.Releasing depthTwo = new SessionManager.Releasing(2, "term_a", MemberRole.DEV);
SessionManager.Releasing afterLeave = depthTwo.leave();
assertNotNull(afterLeave,
"depth 2 means another release of the SAME pane is still mid-teardown; leave() must "
+ "step the depth down, never remove the marker outright — removing it here "
+ "is what a plain Set would do, and would reopen the window the still-in-"
+ "flight release is relying on staying closed");
assertEquals(1, afterLeave.depth());
assertEquals("term_a", afterLeave.terminalId());
assertEquals(MemberRole.DEV, afterLeave.role());
}
@Test
void releasingEnterPreservesThePriorTerminalWhenTheOverlappingCallHasNoSessionOfItsOwn() {
SessionManager.Releasing prior = new SessionManager.Releasing(1, "term_a", MemberRole.DEV);
SessionManager.Releasing afterEnter = SessionManager.Releasing.enter(prior, null);
assertEquals(2, afterEnter.depth(), "depth still increments whether or not this enter knows its session");
assertEquals("term_a", afterEnter.terminalId(),
"an overlapping release that finds the registry entry already gone has no session of "
+ "its own to pass as known, and must not blank out the terminal the first "
+ "call already recorded — that terminal is what spawnedMemberRole matches "
+ "against");
assertEquals(MemberRole.DEV, afterEnter.role());
}
}
@@ -159,40 +159,6 @@ class WorktreeSessionManagerTest {
assertEquals(expectedPath, s.cwd(), "session cwd is the worktree path");
}
// --- fleetd #722: registration and presence must reach READY whichever lands first --------
@Test
void registerThenContactReachesReadyForWorktreeSpawn() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
MemberSession session = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-722", null));
sessions.asPresence().markPresent(session.terminalId());
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"a presence contact that arrives after worktree registration reaches READY");
}
@Test
void contactThenRegisterStillReachesReadyForWorktreeSpawn() {
// The racing launcher marks presence for the spawned terminal from inside spawn() —
// before SessionManager.acquireWithWorktree's own registry.put runs — modeling an MCP
// contact that lands in that window.
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
PresenceRacingLauncher race = new PresenceRacingLauncher(workerService(herdr));
SessionManager sessions = new SessionManager(race, worktrees);
race.presence = sessions.asPresence();
MemberSession session = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-722", null));
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"a presence contact that lands before worktree registration must still reach READY");
}
@Test
void worktreeArchitectAcquireAlsoBindsItsSlot() {
FakeHerdr herdr = new FakeHerdr();
+26 -160
View File
@@ -2,10 +2,6 @@
#
# The one auditable way to edit the live fleetd.yaml.
#
# `--set` uses yq and rewrites the whole YAML document in yq's output style. Use `--from` for a
# candidate whose comment alignment or other formatting carries meaning: it copies that file
# verbatim while keeping this script's backup, parse check, atomic install, and verdict read-back.
#
# fleetd ticket #635 — why this exists at all: fleetd.yaml is gitignored and holds the live
# fleet's settings. A bad raw edit reaches a daemon that is already serving, so a direct `Edit`
# on it is refused by policy. This script is the allow-listed alternative, and it is not just
@@ -50,11 +46,6 @@
# scripts/config-edit.sh --dry-run --set <yq-path>=<value>
# scripts/config-edit.sh --restore
#
# `--set` rewrites the whole file in yq's output style, not only the requested keys. The script
# warns before installation when the candidate changes more lines than its number of --set pairs.
# Use `--from <candidate.yaml>` when comment alignment or other formatting is meaningful: --from
# copies the candidate verbatim, with no yq round-trip.
#
# `--set .a.b=` (an empty value — a forgotten typo) is REFUSED, not accepted as "clear the
# field": a null value falls back to its default rather than erroring, which is silent, not
# safe. To clear a key on purpose, write a literal null: `--set .a.b=null`. Every other value
@@ -163,10 +154,18 @@ done
# story: a YAML block scalar (`|`, `|-`, `>`, `>-`, ...) puts the VALUE on the lines that follow
# the key, each indented deeper than it. The key-name match above only ever sees the key line
# itself, so those continuation lines used to flow straight through unredacted while the key line
# right above them printed a reassuring "<redacted>". The redactor maps masked continuation lines
# from each complete file before it reads the diff. It then masks a printed line when that file
# line is inside a masked key's value. This covers block-scalar bodies even when the key line is
# outside the printed hunk, and it keeps blank lines inside the value masked.
# right above them printed a reassuring "<redacted>" — an incomplete redactor that looks complete
# is worse than one that visibly does nothing, because it stops a reviewer from looking further.
# The fix is structural, not another name to match: once a key line is masked, every following
# line indented STRICTLY DEEPER than that key is masked too, by indentation alone, until the
# indentation returns to the key's own level or shallower. This needs no knowledge of the key's
# name, so it covers a block scalar under any masked key — but ONLY while that key's own line is
# itself inside the hunk being printed. `diff -u` prints just three lines of context, so a block
# scalar's body often reaches this function with its key line left out; there is then nothing to
# anchor to, `masked` is never set, and the body prints in full. A blank line inside a block
# scalar loses the anchor the same way, because a blank diff line measures as indent 0. Both are
# measured and filed as fleetd #639 — do not read this paragraph as a guarantee that a masked
# key's value can never be printed.
#
# `redact` is always fed `diff -u` output, and every line of a unified diff starts with exactly
# one of ' ', '+', '-' (the three body markers; '@'/'-'/'+' for the three header-line kinds too).
@@ -175,113 +174,37 @@ done
# column shallower than it really is, and either wrongly escapes a continuation mask or wrongly
# ends one early. Tabs are out of scope: YAML forbids them for indentation, and this is a bounded
# fix, not a YAML parser.
map_masked_lines() {
local file="$1" side="$2" line content indent lead key line_number=0
local masked=0 masked_indent=0
case "$side" in
old) OLD_MASKED_LINES=() ;;
new) NEW_MASKED_LINES=() ;;
*) die "internal error: unknown redaction map side $side" ;;
esac
while IFS= read -r line || [ -n "$line" ]; do
line_number=$((line_number + 1))
content="$line"
indent=0
while [ "${content:$indent:1}" = " " ]; do indent=$((indent + 1)); done
if [ "$masked" = 1 ]; then
if [ -z "${content// /}" ] || [ "$indent" -gt "$masked_indent" ]; then
case "$side" in
old) OLD_MASKED_LINES[$line_number]=1 ;;
new) NEW_MASKED_LINES[$line_number]=1 ;;
esac
continue
fi
masked=0
fi
if [[ "$content" =~ ^([[:space:]]*)([A-Za-z0-9_.-]+:) ]]; then
lead="${BASH_REMATCH[1]}"
key="${BASH_REMATCH[2]}"
if [[ "$key" =~ (TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY) ]]; then
masked=1
masked_indent="$indent"
fi
fi
done < "$file"
}
# Masks every `scheme://user:pass@host` userinfo on one line of text, replacing just that
# userinfo with `<redacted>` and leaving the rest of the line untouched, byte for byte. The
# pattern stops at the first `/`, whitespace, or `@` reached after `://` — a URI's userinfo
# component cannot contain any of those three characters — so a URI with no userinfo, followed
# later on the same line by an unrelated `@`, never matches. The `g` flag matters: a line can
# carry more than one URI. Shared by every caller that prints a line which may hold a
# credentialed URI, so the bound lives in exactly one place.
mask_url_userinfo() {
printf '%s\n' "$1" | sed -E 's#://[^@/[:space:]]*@#://<redacted>@#g'
}
redact() {
local old_file="$1" new_file="$2"
local line prefix content indent lead key old_line=0 new_line=0 in_hunk=0
local old_masked new_masked saved_nocasematch=0
local line prefix content indent lead key
local masked=0 masked_indent=0 saved_nocasematch=0
shopt -q nocasematch && saved_nocasematch=1
shopt -s nocasematch
map_masked_lines "$old_file" old
map_masked_lines "$new_file" new
while IFS= read -r line || [ -n "$line" ]; do
if [[ "$line" =~ ^@@\ -([0-9]+)(,([0-9]+))?\ \+([0-9]+)(,([0-9]+))?\ @@ ]]; then
old_line="${BASH_REMATCH[1]}"
new_line="${BASH_REMATCH[4]}"
in_hunk=1
printf '%s\n' "$line"
continue
fi
sed -E 's#://[^@]*@#://<redacted>@#g' | while IFS= read -r line || [ -n "$line" ]; do
case "$line" in
[\ +-]*) prefix="${line:0:1}"; content="${line:1}" ;;
*) prefix=""; content="$line" ;;
esac
old_masked=0
new_masked=0
if [ "$in_hunk" = 1 ]; then
case "$prefix" in
' ')
[ "${OLD_MASKED_LINES[$old_line]:-}" = 1 ] && old_masked=1
[ "${NEW_MASKED_LINES[$new_line]:-}" = 1 ] && new_masked=1
old_line=$((old_line + 1)); new_line=$((new_line + 1)) ;;
-)
[ "${OLD_MASKED_LINES[$old_line]:-}" = 1 ] && old_masked=1
old_line=$((old_line + 1)) ;;
+)
[ "${NEW_MASKED_LINES[$new_line]:-}" = 1 ] && new_masked=1
new_line=$((new_line + 1)) ;;
esac
fi
indent=0
while [ "${content:$indent:1}" = " " ]; do indent=$((indent + 1)); done
if [ "$old_masked" = 1 ] || [ "$new_masked" = 1 ]; then
if [ "$masked" = 1 ] && [ "$indent" -gt "$masked_indent" ]; then
printf '%s%*s<redacted>\n' "$prefix" "$indent" ""
continue
fi
masked=0
if [[ "$content" =~ ^([[:space:]]*)([A-Za-z0-9_.-]+:) ]]; then
lead="${BASH_REMATCH[1]}"
key="${BASH_REMATCH[2]}"
if [[ "$key" =~ (TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY) ]]; then
printf '%s%s%s <redacted>\n' "$prefix" "$lead" "$key"
masked=1
masked_indent="$indent"
continue
fi
fi
mask_url_userinfo "$line"
printf '%s\n' "$line"
done
[ "$saved_nocasematch" = 1 ] || shopt -u nocasematch
}
@@ -540,50 +463,6 @@ parse_check() {
yq eval '.' "$1" >/dev/null 2>&1
}
# Count logical changed lines in a unified diff. A replacement counts once, while an added or
# deleted line also counts once. One changed `--set` value normally produces one changed line.
changed_line_count() {
local before="$1" after="$2" line count=0 old_count=0 new_count=0
while IFS= read -r line || [ -n "$line" ]; do
case "$line" in
---\ *|+++\ *|@@\ *)
if [ "$old_count" -gt "$new_count" ]; then
count=$((count + old_count))
else
count=$((count + new_count))
fi
old_count=0
new_count=0
;;
-*) old_count=$((old_count + 1)) ;;
+*) new_count=$((new_count + 1)) ;;
*)
if [ "$old_count" -gt "$new_count" ]; then
count=$((count + old_count))
else
count=$((count + new_count))
fi
old_count=0
new_count=0
;;
esac
done < <(diff -u "$before" "$after" || true)
if [ "$old_count" -gt "$new_count" ]; then
count=$((count + old_count))
else
count=$((count + new_count))
fi
printf '%s' "$count"
}
warn_set_reformat() {
local before="$1" after="$2" changed
changed="$(changed_line_count "$before" "$after")"
if [ "$changed" -gt "${#SETS[@]}" ]; then
warn "--set changed $changed candidate lines for ${#SETS[@]} pair(s); yq reformatted the whole file. Use --from for meaningful comment alignment or formatting."
fi
}
install_candidate() {
local cand="$1" live="$2"
mv -f "$cand" "$live"
@@ -591,12 +470,6 @@ install_candidate() {
# -------------------------------------------------------------------------------- the report path
#
# Masks basic-auth userinfo (scheme://user:pass@host) in a daemon verdict line before it reaches
# the terminal.
mask_verdict_userinfo() {
mask_url_userinfo "$1"
}
# Prints the literal command the operator (or a test) can run to restore the backup by hand — the
# absolute path to THIS script plus the overrides actually in force, so it works from any cwd.
restore_command_line() {
@@ -616,8 +489,8 @@ restore_and_confirm() {
ok "restored from $backup"
if wait_for_verdict "$LOG" "$mark2" "$WAIT_SECONDS"; then
case "$VERDICT_KIND" in
refused) warn "the RESTORE was also refused by the daemon: $(mask_verdict_userinfo "$VERDICT_LINE")" ;;
*) ok "restore confirmed: $(mask_verdict_userinfo "$VERDICT_LINE")" ;;
refused) warn "the RESTORE was also refused by the daemon: $VERDICT_LINE" ;;
*) ok "restore confirmed: $VERDICT_LINE" ;;
esac
else
warn "the restore is on disk, but no confirming verdict line appeared within ${WAIT_SECONDS}s"
@@ -633,7 +506,7 @@ report_outcome() {
say "waiting for the daemon's verdict (up to ${WAIT_SECONDS}s)"
if wait_for_verdict "$LOG" "$mark" "$WAIT_SECONDS"; then
kind="$VERDICT_KIND"; line="$(mask_verdict_userinfo "$VERDICT_LINE")"
kind="$VERDICT_KIND"; line="$VERDICT_LINE"
else
kind="none"
fi
@@ -690,7 +563,7 @@ check_mode() {
local verdict
verdict="$(last_verdict_line "$LOG")"
if [ -n "$verdict" ]; then
ok "last verdict in log: $(mask_verdict_userinfo "$verdict")"
ok "last verdict in log: $verdict"
else
warn "no reload verdict line found in $LOG"
fi
@@ -751,14 +624,10 @@ run_edit() {
fi
ok "candidate parses"
if [ "$MODE" = "set" ]; then
warn_set_reformat "$backup" "$cand"
fi
apply_mode "$cand" "$orig_mode"
say "change (redacted)"
diff -u "$backup" "$cand" | redact "$backup" "$cand" || true
diff -u "$backup" "$cand" | redact || true
say "install"
install_candidate "$cand" "$CONFIG" \
@@ -786,11 +655,8 @@ dry_run_diff() {
rm -f "$cand"; CAND=""
die "candidate does not parse as valid YAML — this was a --dry-run, nothing would have been installed either"
fi
if [ "$MODE" = "set" ]; then
warn_set_reformat "$CONFIG" "$cand"
fi
say "dry run — diff (redacted), nothing installed"
diff -u "$CONFIG" "$cand" | redact "$CONFIG" "$cand" || true
diff -u "$CONFIG" "$cand" | redact || true
rm -f "$cand"; CAND=""
return 0
}
+66 -131
View File
@@ -71,22 +71,20 @@ set -euo pipefail
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
MODULE="$REPO/fleetd"
# fleetd #664: the runtime path and Maven's output path are no longer the same file. Maven's
# shade plugin (finalName=fleetd) always lands a fresh build at target/fleetd.jar — that is
# Maven's own output directory and this script does not change it — but the daemon is launched
# from $JAR instead, outside target/ entirely. That split is the whole fix: neither `mvn install`
# nor `mvn clean` can ever reach the file a running daemon holds open, because that file no
# longer lives under target/ at all. See swap_if_built/swap_staged_jar below for the one `mv`
# that moves a build from one path to the other, and only after the old daemon is confirmed gone.
BUILD_JAR="$MODULE/target/fleetd.jar"
JAR="$MODULE/run/fleetd.jar"
JAR="$MODULE/target/fleetd.jar"
# fleetd #493: never build into the path a running process holds. The build writes here first
# (Maven's shade plugin has finalName=fleetd, so `clean install` still lands its output at
# target/fleetd.jar — that part is unchanged and out of this script's control), but this script
# now moves it out to JAR_STAGED immediately, and only swaps it back to JAR (a plain `mv`, so a
# rename, never a byte-by-byte overwrite) after the OLD daemon has been confirmed exited. See
# stage_built_jar/swap_staged_jar below.
JAR_STAGED="$MODULE/target/fleetd-new.jar"
OUT="$MODULE/fleetd.out"
# Matches a fleetd daemon's command line wherever its jar sits — absolute or relative, under
# run/, under target/, or anywhere else a build or a hand-start might point it. Detecting a
# daemon this script did not start, including one running from a jar outside $JAR's own
# directory, is this pattern's whole job; running_pid()'s `comm = java` allowlist below is what
# keeps that breadth from counting a shell that merely types the pattern as literal text.
PATTERN='fleetd.jar'
# Matches BOTH the absolute form and the relative `java -jar target/fleetd.jar` a hand-start
# produces from inside fleetd/. Anchoring on the absolute path alone was a real bug: the daemon
# restarted correctly and the script still reported "no process appeared", because it launched with
# a relative path and then looked for an absolute one.
PATTERN='target/fleetd.jar'
HEALTH='http://127.0.0.1:8765/healthz'
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start — fleetd #603: also the pid-
@@ -171,10 +169,10 @@ hash256() {
fi
}
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the BUILT
# jar at $BUILD_JAR (before it has been swapped in) without ever changing what a bare `jar_id`
# (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line both call
# it with no args on purpose, so neither can ever be fooled by a leftover build output.
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the STAGED
# jar right after a build (before it has been swapped in) without ever changing what a bare
# `jar_id` (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line
# both call it with no args on purpose, so neither can ever be fooled by a leftover staged file.
# fleetd #550 — THREE distinct answers now, not two: `[ -f "$f" ]` already separates "the jar is
# not there" (-> "absent") from "the jar is there"; for the second case, hash256 itself separates
# "hashed it" (a 12-char hex string) from "could not hash it" (-> "unhashable", when no hasher is
@@ -182,35 +180,15 @@ hash256() {
# was the whole defect this ticket fixes.
jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
# fleetd #664 — under the old layout $JAR and the build output were the same file, so "jar on
# disk" was one fact. Now they are two: $BUILD_JAR (target/fleetd.jar, whatever Maven last wrote,
# by this script or by a bare `mvn install` run by hand) and $JAR (run/fleetd.jar, whatever the
# daemon actually has open). Printing one label for both was the trap this ticket exists to close
# — during the incident it would have shown the NEW jar's hash while the JVM ran the OLD one.
# Pure (reads jar_id/date, never mutates), so a test can call it directly without reaching the
# main flow — the same shape swap_if_built/drain_gate_refusal already use.
report_jar_state() {
local built_hash running_hash built_mtime running_mtime
built_hash="$(jar_id "$BUILD_JAR")"
running_hash="$(jar_id "$JAR")"
built_mtime="$([ -f "$BUILD_JAR" ] && date -r "$BUILD_JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none')"
running_mtime="$([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none')"
ok "built jar (target/fleetd.jar): $built_hash ($built_mtime)"
ok "running jar (run/fleetd.jar): $running_hash ($running_mtime)"
if [ "$built_hash" != "absent" ] && [ "$running_hash" != "absent" ] && [ "$built_hash" != "$running_hash" ]; then
warn "built jar and running jar differ — target/fleetd.jar was rebuilt since the running daemon last started and is not yet live"
fi
}
# fleetd #593 — `pgrep -f "$PATTERN"` matches ANY process whose full command line CONTAINS the
# pattern text, and that is not the same thing as "is the daemon". A shell that merely embeds the
# pattern as literal text — a human typing this exact investigation by hand, an ssh-shaped
# `sh -c '...; ...'`, a pipeline, or any other non-exec'ing shell that never replaced itself with
# the pattern-holding command — still shows up in that match, and it is the INSTRUMENT, not the
# daemon. Measured live on this Mac: `sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 30' &`
# daemon. Measured live on this Mac: `sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 30' &`
# leaves a real `sh` process alive (it forks for the `sleep`, it does not exec into it) whose own
# `ps -o args` is `sh -c echo "run/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
# matches that line right alongside the real `java -jar run/fleetd.jar` process. `pgrep -c`
# `ps -o args` is `sh -c echo "target/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
# matches that line right alongside the real `java -jar target/fleetd.jar` process. `pgrep -c`
# (an in-one-call count) does not exist on BSD/macOS at all, so this cannot be fixed by switching
# pgrep flags — it has to filter what pgrep already found, after the fact, in a way that still
# runs on BSD.
@@ -232,7 +210,7 @@ report_jar_state() {
# launched as `java -jar ...` — a native image, a renamed launcher — `running_pid()` silently
# returns nothing and `assert_single_daemon` stops noticing a second daemon at all. For a guard,
# that false-negative direction is the worse one to be wrong in. This is not a new assumption,
# though: `PATTERN='fleetd.jar'` two lines up already assumes the daemon is a jar, which
# though: `PATTERN='target/fleetd.jar'` two lines up already assumes the daemon is a jar, which
# is only ever run by `java`. If that launch method changes, `PATTERN` stops matching anything
# before this allowlist would ever get the chance to be wrong — the allowlist rides on the same
# assumption that is already load-bearing, it does not add a new one. Whoever changes the launch
@@ -251,12 +229,16 @@ running_pid() {
printf '%s' "$out"
}
# fleetd #493/#664 — the independently testable pieces of "never build into the path a running
# process holds":
# fleetd #493 — three small, independently testable pieces of "never build into the path a
# running process holds":
#
# require_no_build_jar the --no-build path never builds anything: it must find a jar already
# sitting at the live path ($JAR, under run/) from an earlier successful run,
# and die with the same truthful message this script has always used if not.
# stage_built_jar moves the jar Maven just produced OUT of the live path and onto the staging
# path, immediately after a successful build. Dies (leaving the OLD daemon
# untouched — this runs before the stop step) if Maven reported success but
# left no jar behind, or if the move itself fails.
# require_no_build_jar the --no-build path never builds or stages anything: it must find a
# jar already sitting at the live path from an earlier successful run, and
# die with the same truthful message this script has always used if not.
# wait_for_daemon_exit polls running_pid() for up to $1 seconds and reports whether the OLD
# daemon actually exited — extracted to its own function so the main flow
# can be relied on to call swap_staged_jar only AFTER this returns success,
@@ -266,10 +248,14 @@ running_pid() {
# so this is never a write into a path a running process holds — by the time
# it runs, nothing holds that path anymore. If it fails, the caller must not
# start a new daemon: die() below already refuses that by exiting the script.
# fleetd #664: the "staged" jar swap_if_built passes in is now $BUILD_JAR
# itself (target/fleetd.jar, Maven's own output) — a build no longer needs to
# be moved off the live path right after compiling, because target/ was never
# the live path to begin with.
stage_built_jar() {
[ -f "$JAR" ] || die "build succeeded but produced no jar at $JAR — cannot stage it for restart.
The running daemon was NOT touched."
mv -f "$JAR" "$JAR_STAGED" \
|| die "could not move the freshly built jar from $JAR to the staging path $JAR_STAGED.
The running daemon was NOT touched."
}
require_no_build_jar() {
[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build"
}
@@ -296,7 +282,8 @@ swap_staged_jar() {
#
# The defect: the swap step used to be guarded inline by `if [ "$DO_BUILD" = 1 ]` in the main flow.
# Changing that to `if false` left the suite green and the swap never ran, so a redeploy reported
# every step succeeding while the daemon started on no jar at all or on a stale one.
# every step succeeding while the daemon started on no jar at all (stage_built_jar has already moved
# the freshly built one to $JAR_STAGED by then) or on a stale one.
# test_swap_ordered_after_wait_and_before_start could not catch it: it reads this script's own text
# and compares line positions, and a same-line edit moves no line.
#
@@ -325,12 +312,7 @@ swap_if_built() {
local do_build="$1"
should_swap "$do_build" || return 0
say "swap"
# fleetd #664: $JAR now lives under run/, a directory target/ never created. mkdir -p here,
# not inside swap_staged_jar itself — that function's own contract is tested on a missing
# parent directory (a failing mv), and widening it to auto-create one would change what that
# test proves.
mkdir -p "$(dirname "$JAR")"
swap_staged_jar "$BUILD_JAR" "$JAR"
swap_staged_jar "$JAR_STAGED" "$JAR"
ok "jar in place: $(jar_id)"
}
@@ -628,45 +610,6 @@ check_log_path_matches_plist() {
ok "log path check: script and plist agree ($resolved_out)"
}
# Reads the launchd plist's ProgramArguments for the argument that follows "-jar", resolves it
# alongside $jar_path, and dies when the two differ. Call it only when the agent is loaded; it
# never touches launchd or the daemon itself.
check_jar_path_matches_plist() {
local jar_path="$1" plist_path="$2"
local plist_args plist_jar resolved_jar resolved_plist_jar
if ! plist_args="$(/usr/libexec/PlistBuddy -c 'Print :ProgramArguments' "$plist_path" 2>/dev/null)"; then
die "launchd agent is loaded but PlistBuddy could not read ProgramArguments from
$plist_path
— cannot verify which jar the supervised daemon launches. Fix the plist before
redeploying supervised."
fi
plist_jar="$(printf '%s\n' "$plist_args" | awk '
{ gsub(/^[ \t]+|[ \t]+$/, "") }
prev == "-jar" { print; exit }
{ prev = $0 }
')"
if [ -z "$plist_jar" ]; then
die "launchd agent is loaded but its ProgramArguments at
$plist_path
do not contain a '-jar <path>' pair — cannot verify which jar the supervised daemon
launches. Fix the plist before redeploying supervised."
fi
resolved_jar="$(cd "$(dirname "$jar_path")" 2>/dev/null && pwd -P)/$(basename "$jar_path")" || true
resolved_plist_jar="$(cd "$(dirname "$plist_jar")" 2>/dev/null && pwd -P)/$(basename "$plist_jar")" || true
if [ -z "$resolved_jar" ] || [ -z "$resolved_plist_jar" ] || [ "$resolved_jar" != "$resolved_plist_jar" ]; then
die "jar path mismatch — this script deploys to
$jar_path (resolved: ${resolved_jar:-<directory does not exist>})
but the loaded plist's ProgramArguments names
$plist_jar (resolved: ${resolved_plist_jar:-<directory does not exist>})
The swap renames the built jar into place, so the old path stops existing after a redeploy;
a launchd-initiated start from this plist (a reboot, or KeepAlive after a crash) would then
run java against a missing file. Reinstall the plist at
$plist_path
so its ProgramArguments names $jar_path before redeploying supervised."
fi
ok "jar path check: script and plist agree ($resolved_jar)"
}
# fleetd #552: the post-restart fresh-log capture, pulled out of the main flow so it is testable by
# sourcing (the same reason systemd_installed/systemd_loaded above guard their OWN mktemp inline
# instead of leaving it bare) even though its only caller sits below the SOURCED guard. By the time
@@ -883,25 +826,16 @@ report_shutdown_drain() {
# --no-build, staged jar present -> ALSO "nothing changed", deliberately: --no-build itself builds
# and stages nothing (see require_no_build_jar above), so a staged jar found here is a leftover
# from an earlier, unrelated run. THIS run truly changed nothing, and the next DO_BUILD=1 run
# wipes that leftover for free — `mvn clean` deletes all of target/, $BUILD_JAR included,
# before the build even starts — so there is nothing here for the operator to lose track of.
# wipes that leftover before it builds (`rm -f "$JAR_STAGED"` in the build section above) — so
# there is nothing here for the operator to lose track of.
# --no-build, staged jar absent -> "nothing changed"
drain_gate_refusal() {
local do_build="$1" staged_path="$2"
if [ "$do_build" = 1 ] && [ -f "$staged_path" ]; then
# fleetd #664: under the old layout a build emptied the live path ($JAR) immediately, so
# --no-build's own check ("no jar at $JAR") was the thing that refused a rerun here. That is
# no longer true: a build never touches $JAR at all now, so $JAR still holds whatever was
# already running before this gate fired (reaching this message at all requires OLD_PID to
# have been set, which means a daemon was running from $JAR already) — a --no-build rerun
# would NOT refuse, it would just restart that same old jar and silently throw away the one
# sitting at %s.
printf 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart — a
--no-build rerun would NOT refuse here: %s already exists from before this run, so it
would restart the daemon on that OLD jar and silently discard the one you just built — or
remove %s by hand if you want to discard this build instead.' \
"$staged_path" "$JAR" "$JAR" "$staged_path"
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart —
the freshly built jar is no longer at the live path that --no-build requires — or
remove %s by hand if you want to discard this build.' "$staged_path" "$JAR" "$staged_path"
else
printf 'aborted — nothing changed'
fi
@@ -1000,7 +934,6 @@ report_supervisor_state() {
SUPERVISED=1
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
check_jar_path_matches_plist "$JAR" "$LAUNCHD_PLIST"
;;
systemd)
SUPERVISED=1
@@ -1237,7 +1170,7 @@ if [ -n "$OLD_PID" ]; then
else
warn "no daemon running — this will be a cold start"
fi
report_jar_state
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
# CB-594 / fleetd #492: supervision state. Installed and loaded are different facts — a
@@ -1319,6 +1252,9 @@ stop_if_check_only "$CHECK_ONLY"
if [ "$DO_BUILD" = 1 ]; then
say "build"
# fleetd #493: wipe a leftover staged jar from a previous failed/interrupted run BEFORE doing
# anything else, so that run's leftovers can never be mistaken for this run's output.
rm -f "$JAR_STAGED"
BUILD_LOG="$(mktemp -t fleetd-build.XXXXXX)"
echo " log: $BUILD_LOG"
if ! mvn -f "$MODULE/pom.xml" clean install > "$BUILD_LOG" 2>&1; then
@@ -1328,16 +1264,16 @@ if [ "$DO_BUILD" = 1 ]; then
fi
grep -E '^\[INFO\] Tests run:.*Failures' "$BUILD_LOG" | tail -1 | sed 's/^\[INFO\] / /' || true
ok "BUILD SUCCESS"
# fleetd #664: nothing to stage — $BUILD_JAR (target/fleetd.jar) is Maven's own output path and
# was never the live path, so the running (OLD) daemon, if any, was never at risk from this build
# at all. From here until the swap step below (after the OLD daemon is confirmed gone),
# $BUILD_JAR is the artefact this script treats as "the new jar" — $JAR itself is not touched
# again until the swap.
ok "jar now: $(jar_id "$BUILD_JAR")"
# fleetd #493: move the freshly built jar off the live path immediately — the running (OLD)
# daemon, if any, is still up at this point (build always runs before stop). From here until the
# swap step below (after the OLD daemon is confirmed gone), $JAR_STAGED is the only artefact this
# script treats as "the new jar" — $JAR itself is not touched again until the swap.
stage_built_jar
ok "jar now: $(jar_id "$JAR_STAGED")"
else
say "build skipped (--no-build)"
# fleetd #493: --no-build never builds anything — it restarts whatever jar is already sitting at
# the live path from an earlier successful run. Same check, same message as before.
# fleetd #493: --no-build never builds or stages anything — it restarts whatever jar is already
# sitting at the live path from an earlier successful run. Same check, same message as before.
require_no_build_jar
fi
@@ -1346,8 +1282,8 @@ fi
# fleetd #555: run_drain_gate above is drain_gate_required + the prompt + drain_confirmed, called
# unconditionally — it returns immediately when the gate is not required, and composes/dies through
# refuse_drain_gate itself when the reply does not confirm. See #493/#517/#528 for why "nothing
# changed" would be a lie once a build has produced a jar not yet swapped in.
run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"
# changed" would be a lie once a build has staged a jar.
run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"
# ------------------------------------------------------------------ stop
#
@@ -1398,22 +1334,21 @@ fi
# ------------------------------------------------------------------ swap
#
# fleetd #493/#664: every branch above has now either confirmed the OLD daemon actually exited
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only NOW
# is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
# launchd/systemd's ExecStart) will read from — a rename from $BUILD_JAR (target/) to $JAR (run/),
# both under $MODULE and so on one filesystem. This mv is the one and only write to $JAR anywhere
# fleetd #493: every branch above has now either confirmed the OLD daemon actually exited
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only
# NOW is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
# launchd/systemd's ExecStart) will read from — this mv is the one and only write to $JAR anywhere
# in this script's mutating flow. If it fails, do not start: die() below exits before "start" runs.
swap_if_built "$DO_BUILD"
# ------------------------------------------------------------------ start
# Unsupervised: login shell (zsh -l) is what puts the secrets on the daemon's environment, and cwd
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and run/ relative to it.
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and target/ relative to it.
# Supervised (launchd): launchd does both — deploy/dev.ltms.fleetd.plist points ProgramArguments at
# scripts/fleetd-launchd-wrapper.sh (CB-594), which is what execs the login shell in launchd's
# place, and WorkingDirectory in the plist already pins fleetd/.
# Supervised (systemd --user): the unit does both too — measured on the second host, ExecStart is
# `/bin/zsh -lc "exec java -jar run/fleetd.jar fleetd.yaml"` (a login shell, same reason as
# `/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` (a login shell, same reason as
# above) and WorkingDirectory is already pinned to fleetd/.
say "start"
-274
View File
@@ -233,66 +233,6 @@ test_redaction_holds() {
assert_contains "weight" "$RUN_OUTPUT" "a diff must have been demonstrably printed at all"
}
# redact()'s key-name filter only inspects the KEY, so a diff line whose key does not match
# TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY still reaches the final userinfo
# sed even when its VALUE holds a credentialed URI. "note" is not a sensitive key name, so this
# line must fall all the way through to that sed, not the earlier whole-value branch. The
# trailing prose on both sides of the userinfo is a positive control: it proves the line reached
# the userinfo sed (which touches only the userinfo) rather than the earlier branch (which would
# have replaced the whole value with a bare "<redacted>" and dropped the prose).
test_diff_line_userinfo_is_masked_with_positive_control() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.note=see amqp://alice:wonderland@rabbit.local:5672/vhost for details'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "diff-userinfo-case reload exit code"
assert_not_contains "alice:wonderland" "$RUN_OUTPUT" "the userinfo must never reach the output"
assert_contains "amqp://<redacted>@rabbit.local:5672/vhost" "$RUN_OUTPUT" \
"the userinfo must be MASKED, not deleted — the rest of the value must survive"
assert_contains "note:" "$RUN_OUTPUT" "the key name must still reach the output"
assert_contains "see " "$RUN_OUTPUT" "prose BEFORE the userinfo must still reach the output"
assert_contains "for details" "$RUN_OUTPUT" "prose AFTER the userinfo must still reach the output"
}
# A diff line can hold a URL with no userinfo, followed later on the same line by an unrelated @
# (free text in a string value, for example an email address). The line must pass through the
# userinfo sed byte for byte: the match must stop at the end of the URL and must not treat the
# later @ as a second userinfo delimiter.
test_diff_line_uri_without_userinfo_survives_a_later_at_sign() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.note2=see https://docs.local/guide and mail ops@example.com'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "diff-no-userinfo-with-later-at-sign reload exit code"
assert_contains "note2: see https://docs.local/guide and mail ops@example.com" "$RUN_OUTPUT" \
"a URL with no userinfo plus a later @ on the same line must pass through byte for byte"
}
# Two credentialed URIs on one diff line must both be masked — the g flag matters.
test_diff_line_masks_multiple_userinfo_with_g_flag() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.note3=amqp://u1:p1@host1/vhost1 and amqp://u2:p2@host2/vhost2'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "diff-two-userinfo-on-one-line reload exit code"
assert_not_contains "u1:p1" "$RUN_OUTPUT" "the first userinfo must never reach the output"
assert_not_contains "u2:p2" "$RUN_OUTPUT" "the second userinfo must never reach the output"
assert_contains "amqp://<redacted>@host1/vhost1" "$RUN_OUTPUT" "the first URI must be masked"
assert_contains "amqp://<redacted>@host2/vhost2" "$RUN_OUTPUT" "the second URI must be masked"
}
# ------------------------------------------------------- acceptance criterion 9: forgotten value
# `--set .a.b=` is a plausible typo (the value simply forgotten), and it must be refused outright
# rather than silently nulling the field — a null numeric field falls back to its default, which
@@ -534,90 +474,6 @@ test_passphrase_key_is_redacted() {
assert_not_contains "FAKELEAK-PASSPHRASE" "$RUN_OUTPUT" "passphrase case: the passphrase VALUE must never leak"
}
# ------------------- acceptance criterion 19: the key line falls outside the printed hunk
# fleetd #656 — criteria 15a/15b both put the edit right next to the key line, so the key line is
# always inside diff -u's default 3-line context. Neither covers the actual case #639 fixed: an
# 8-line block-scalar body with only its SIXTH line changed, so the printed hunk (3 lines of
# context on each side of the change) covers body lines 3-8 and never includes the "token:" key
# line at all. The old, line-by-line redact() only ever masks after it has SEEN the key line go
# past; with the key line outside the hunk it never sets its mask, and the whole body — the
# changed line included — passes through raw. The control key sits right after the body, inside
# the same hunk, so the positive control below proves the fix is not simply printing nothing.
new_fixture_hunk_without_key_line() {
local dir
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
cat > "$dir/fleetd.yaml" <<'YAML'
bind:
host: 127.0.0.1
port: 19999
broker:
uri: amqp://user:hunter2@host/vhost
auth:
token: |
SECRET-LINE-1
SECRET-LINE-2
SECRET-LINE-3
SECRET-LINE-4
SECRET-LINE-5
SECRET-LINE-6
SECRET-LINE-7
SECRET-LINE-8
control: CTRL-MUST-APPEAR
profiles:
sonnet:
weight: 3
maxLoad: 5
YAML
: > "$dir/fleetd.out"
printf '%s' "$dir"
}
test_key_line_outside_hunk_is_still_redacted() {
local dir
dir="$(new_fixture_hunk_without_key_line)"
sed 's/SECRET-LINE-6$/SECRET-LINE-6-CHANGED/' "$dir/fleetd.yaml" > "$dir/candidate.yaml"
start_run "$dir" 5 --from "$dir/candidate.yaml"
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "hunk-without-key-line case reload exit code"
# Positive control FIRST: without this, a diff that printed nothing at all would pass the
# negative assertion right below identically to a correctly redacted one.
assert_contains "CTRL-MUST-APPEAR" "$RUN_OUTPUT" "hunk-without-key-line case: the non-secret control line must still print unmasked"
assert_not_contains "SECRET-LINE-6-CHANGED" "$RUN_OUTPUT" "hunk-without-key-line case: the changed body line must never leak, even with the key line outside the printed hunk"
}
# ------------------------------- acceptance criterion 20: a blank line inside the value
# fleetd #656 — the old, line-by-line redact() reset its mask on any line whose indentation was
# not STRICTLY greater than the key's, and a wholly blank line has indentation 0, so it reset the
# mask exactly like the "control:" line that legitimately ends the block scalar. Everything after
# the blank line then printed raw. The current fix tracks masked lines by FILE line number instead
# of by indentation seen so far, so a blank line inside the value stays masked.
new_fixture_blank_line_in_value() {
local dir
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
printf 'bind:\n host: 127.0.0.1\n port: 19999\nbroker:\n uri: amqp://user:hunter2@host/vhost\nauth:\n token: |\n LEAK-BEFORE-BLANK\n\n LEAK-AFTER-BLANK\n control: CTRL-MUST-APPEAR\nprofiles:\n sonnet:\n weight: 3\n maxLoad: 5\n' > "$dir/fleetd.yaml"
: > "$dir/fleetd.out"
printf '%s' "$dir"
}
test_blank_line_inside_value_is_still_redacted() {
local dir
dir="$(new_fixture_blank_line_in_value)"
sed 's/LEAK-AFTER-BLANK$/LEAK-AFTER-BLANK-CHANGED/' "$dir/fleetd.yaml" > "$dir/candidate.yaml"
start_run "$dir" 5 --from "$dir/candidate.yaml"
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "blank-line-in-value case reload exit code"
assert_contains "CTRL-MUST-APPEAR" "$RUN_OUTPUT" "blank-line-in-value case: the non-secret control line must still print unmasked"
assert_not_contains "LEAK-AFTER-BLANK-CHANGED" "$RUN_OUTPUT" "blank-line-in-value case: the line after the blank must never leak"
}
# ----------------------------------- acceptance criterion 16: a failing --set must not echo value
# fleetd #635 follow-up (ticket comment 17673, defect 8) — apply_set_pairs used to echo the FULL
# "$kv" (path=value, exactly as typed) in its yq-failure messages, so a broken --set with a
@@ -689,116 +545,6 @@ test_refusal_shape_from_parse_failure_wording_is_recognised() {
assert_equals 4 "$RUN_RC" "the parse-failure refusal shape must also exit 4, not be read as silence"
}
# A verdict line carrying a credentialed URI has its userinfo masked, with a positive control
# proving the rest of the line still reaches the output unchanged.
test_verdict_userinfo_is_masked_with_positive_control() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
sleep 1
printf 'config reload from %s refused, keeping the running config: refusing to start: malformed pattern — profiles.local.errorPattern ("amqp://user:hunter2@host/vhost"): Unclosed character class near index 8\n' \
"$dir/fleetd.yaml" >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 4 "$RUN_RC" "refusal-with-userinfo exit code"
assert_not_contains "user:hunter2" "$RUN_OUTPUT" "the userinfo must never reach the output"
assert_contains "amqp://<redacted>@host/vhost" "$RUN_OUTPUT" \
"the userinfo must be MASKED, not deleted — the rest of the quoted value must survive"
# Positive control: the diagnostic prose on both sides of the userinfo must still reach the
# output. Without this, a mutant that drops the whole verdict line would pass identically.
assert_contains "malformed pattern" "$RUN_OUTPUT" "prose BEFORE the userinfo must still reach the output"
assert_contains "Unclosed character class near index 8" "$RUN_OUTPUT" \
"prose AFTER the userinfo must still reach the output"
}
# An ordinary refusal line quotes the offending pattern, not a credential, and must survive byte
# for byte: the rewrite is scoped to userinfo only, and the quoted pattern is the detail an
# operator needs to fix the refusal.
test_ordinary_refusal_line_passes_through_unchanged() {
local dir real_line
dir="$(new_fixture)"
real_line="config reload from $dir/fleetd.yaml refused, keeping the running config: refusing to start: malformed pattern — profiles.local.errorPattern (\"[unclosed\"): Unclosed character class near index 8"
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
sleep 1
printf '%s\n' "$real_line" >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 4 "$RUN_RC" "ordinary refusal exit code"
assert_contains "$real_line" "$RUN_OUTPUT" \
"an ordinary refusal with no userinfo must pass through byte for byte, unchanged"
}
# A verdict line can hold a URI with NO userinfo and a later, unrelated @ further on in the same
# line (an email address in diagnostic prose, for example). The rewrite must stop at the end of
# the URI and must not treat the later @ as a second userinfo delimiter.
test_uri_without_userinfo_survives_a_later_at_sign() {
local dir real_line
dir="$(new_fixture)"
real_line="config reload from $dir/fleetd.yaml refused, keeping the running config: broker.uri amqp://broker.local/vhost unreachable, contact ops@example.com"
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
sleep 1
printf '%s\n' "$real_line" >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 4 "$RUN_RC" "no-userinfo-with-later-at-sign exit code"
assert_contains "$real_line" "$RUN_OUTPUT" \
"a URI with no userinfo plus a later @ in the same line must pass through byte for byte"
}
# --set runs yq over the whole candidate. It warns when that changes more lines than the requested
# pairs, but a simple file with only the intended changed line must stay quiet.
new_fixture_reformat_sensitive() {
local dir
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
cat > "$dir/fleetd.yaml" <<'YAML'
# A section comment that documents the next block.
bind:
host: 127.0.0.1 # Keep this aligned with the port note.
port: 19999 # A fixture port.
# These comments use their placement as documentation.
profiles:
sonnet:
weight: 3
bootstrapText: >-
First line.
Second line.
YAML
: > "$dir/fleetd.out"
printf '%s' "$dir"
}
test_set_warns_when_yq_reformats_extra_lines() {
local dir
dir="$(new_fixture_reformat_sensitive)"
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "reformat warning case reload exit code"
assert_contains "yq reformatted the whole file" "$RUN_OUTPUT" \
"a --set that changes extra candidate lines must warn before installation"
}
test_set_stays_quiet_without_formatting_churn() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "no-reformat warning case reload exit code"
assert_not_contains "yq reformatted the whole file" "$RUN_OUTPUT" \
"a --set that changes only its requested candidate line must not warn"
}
echo "== acceptance criterion 1: refusal restores byte for byte =="
test_refusal_restores_byte_for_byte
echo "== acceptance criterion 2: clean reload keeps the edit =="
@@ -813,12 +559,6 @@ echo "== acceptance criterion 6: the marker works =="
test_marker_skips_lines_before_it
echo "== acceptance criterion 7 (+13: redaction is proven to have run) =="
test_redaction_holds
echo "== fleetd #692: a diff line's userinfo is masked, rest of the value survives =="
test_diff_line_userinfo_is_masked_with_positive_control
echo "== fleetd #692: a diff line's URI with no userinfo survives a later @ in the line =="
test_diff_line_uri_without_userinfo_survives_a_later_at_sign
echo "== fleetd #692: two userinfo URIs on one diff line are both masked =="
test_diff_line_masks_multiple_userinfo_with_g_flag
echo "== acceptance criterion 9: a forgotten value refuses and installs nothing =="
test_forgotten_value_refuses_and_installs_nothing
echo "== acceptance criterion 10: an explicit clear writes a bare null =="
@@ -833,10 +573,6 @@ echo "== acceptance criterion 15a: a block scalar's continuation lines are redac
test_block_scalar_continuation_is_redacted
echo "== acceptance criterion 15b: a passphrase key is also recognised =="
test_passphrase_key_is_redacted
echo "== acceptance criterion 19: the key line falls outside the printed hunk =="
test_key_line_outside_hunk_is_still_redacted
echo "== acceptance criterion 20: a blank line inside the value =="
test_blank_line_inside_value_is_still_redacted
echo "== acceptance criterion 16: a failing --set must not echo its value =="
test_failing_set_does_not_echo_its_value
echo "== extra: dry-run never installs, and redacts =="
@@ -845,15 +581,5 @@ echo "== extra: --check is read-only and always exits 0 =="
test_check_is_read_only_and_exits_zero
echo "== extra: the parse-failure refusal shape is also recognised =="
test_refusal_shape_from_parse_failure_wording_is_recognised
echo "== verdict-redaction criteria 2+3: verdict userinfo is masked, rest of line survives =="
test_verdict_userinfo_is_masked_with_positive_control
echo "== verdict-redaction criterion 4: an ordinary refusal passes through unchanged =="
test_ordinary_refusal_line_passes_through_unchanged
echo "== fleetd #638: a URI with no userinfo survives a later @ in the same line =="
test_uri_without_userinfo_survives_a_later_at_sign
echo "== acceptance criterion 17: --set warns about yq formatting churn =="
test_set_warns_when_yq_reformats_extra_lines
echo "== acceptance criterion 18: --set stays quiet without formatting churn =="
test_set_stays_quiet_without_formatting_churn
printf 'PASS: config-edit acceptance criteria\n'
+63 -213
View File
@@ -447,7 +447,7 @@ test_assert_single_daemon_rejects_two_pids() {
test_running_pid_excludes_self_matching_wrapper_shell() {
local before after wrapper_pid
before="$(running_pid)"
sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 20' &
sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' &
wrapper_pid=$!
sleep 0.3
after="$(running_pid)"
@@ -475,26 +475,11 @@ test_running_pid_excludes_self_matching_wrapper_shell() {
# `comm` from the actually-executed binary's own path, not from `exec -a`'s argv[0] override (BSD
# ties `comm` to argv[0], which is what makes this technique work here) — so on Linux this
# specific fixture might report `comm=sh`, not `comm=java`, even though the REAL daemon (a literal
# `java -jar run/fleetd.jar` process, never fabricated) is unaffected either way. I could not
# `java -jar target/fleetd.jar` process, never fabricated) is unaffected either way. I could not
# verify this fixture's behavior on Linux, so test_running_pid_counts_a_pid_whose_comm_is_java
# below backstops the same claim (the allowlist admits a pid whose comm is `java`) with a stubbed
# `ps`, which is identical bash on every platform and carries no such platform question.
test_running_pid_finds_a_real_java_named_second_process() {
local before after standin_pid
before="$(running_pid)"
( exec -a java sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 20' ) &
standin_pid=$!
sleep 0.3
after="$(running_pid)"
kill "$standin_pid" 2>/dev/null || true
wait "$standin_pid" 2>/dev/null || true
printf '%s\n' "$after" | grep -qxF "$standin_pid" \
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) whose own argv holds the pattern: before=[$before] after=[$after]"
}
# PATTERN matches a fleetd jar in either build layout, not only the run/ one: a process whose
# argv names a jar under target/ must be found too, the same way the run/ case above is.
test_running_pid_finds_a_real_java_named_process_from_target_dir() {
local before after standin_pid
before="$(running_pid)"
( exec -a java sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' ) &
@@ -504,24 +489,7 @@ test_running_pid_finds_a_real_java_named_process_from_target_dir() {
kill "$standin_pid" 2>/dev/null || true
wait "$standin_pid" 2>/dev/null || true
printf '%s\n' "$after" | grep -qxF "$standin_pid" \
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) naming a jar under target/: before=[$before] after=[$after]"
}
# Broadening PATTERN to match both build layouts must not also broaden it into matching a
# non-exec'ing shell that merely holds the target/ text as a literal argument, the same
# self-matching shape test_running_pid_excludes_self_matching_wrapper_shell above excludes for
# the run/ text.
test_running_pid_excludes_self_matching_wrapper_shell_naming_target_dir() {
local before after wrapper_pid
before="$(running_pid)"
sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' &
wrapper_pid=$!
sleep 0.3
after="$(running_pid)"
kill "$wrapper_pid" 2>/dev/null || true
wait "$wrapper_pid" 2>/dev/null || true
[ "$after" = "$before" ] \
|| fail "running_pid() counted a self-matching wrapper shell (pid $wrapper_pid, holding 'target/fleetd.jar' as literal text in its own argv, not the daemon): before=[$before] after=[$after]"
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) whose own argv holds the pattern: before=[$before] after=[$after]"
}
# fleetd #593 CORRECTION 1, hole 2 — the round-1 filter denied known shell names (sh/bash/zsh/
@@ -589,30 +557,29 @@ test_die_message_does_not_recommend_bare_pgrep_as_remediation() {
|| fail "assert_single_daemon's die message does not say in words that a pattern can match the caller (fleetd #593)"
}
# fleetd #511/#664 — jar_id()'s no-argument default was unpinned by any test: nothing proved it
# reports $JAR (the live path) rather than whatever explicit path a caller passes it (e.g.
# $BUILD_JAR). Both halves matter, so this pins both: the bare call must hash the live jar, and an
# explicit path argument must hash THAT file, not fall back to $JAR. Two files with different
# content, so a default pointed at the wrong one reports the wrong hash rather than accidentally
# matching.
# fleetd #511 — jar_id()'s no-argument default was unpinned by any test: nothing proved it reports
# $JAR (the live path) rather than $JAR_STAGED. Both halves matter, so this pins both: the bare call
# must hash the live jar, and an explicit path argument must hash THAT file, not fall back to $JAR.
# Two files with different content, so a default pointed at the wrong one reports the wrong hash
# rather than accidentally matching.
test_jar_id_defaults_to_live_and_reports_explicit_path() {
local dir saved_jar="$JAR"
local live_hash other_hash default_result explicit_result other_path
local dir saved_jar="$JAR" saved_staged="$JAR_STAGED"
local live_hash staged_hash default_result explicit_result
dir="$TMP/jar-id"; mkdir -p "$dir"
JAR="$dir/fleetd.jar"; other_path="$dir/other.jar"
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
printf 'live jar bytes' > "$JAR"
printf 'other jar bytes, not the same content' > "$other_path"
printf 'staged jar bytes, not the same content' > "$JAR_STAGED"
# fleetd #550: this reference hash must be computed the same portable way jar_id() itself now
# computes one — a bare, unguarded call to the macOS-only hasher here was exactly the item-2
# defect, dying with "command not found" on any Linux runner that has no such hasher at all.
live_hash="$(hash256 "$JAR")"
other_hash="$(hash256 "$other_path")"
staged_hash="$(hash256 "$JAR_STAGED")"
default_result="$(jar_id)"
explicit_result="$(jar_id "$other_path")"
JAR="$saved_jar"
[ "$live_hash" != "$other_hash" ] || fail "test fixture error: live and other jars hashed the same"
explicit_result="$(jar_id "$JAR_STAGED")"
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ "$live_hash" != "$staged_hash" ] || fail "test fixture error: live and staged jars hashed the same"
assert_equals "$live_hash" "$default_result" "jar_id with no arguments must report the hash of \$JAR"
assert_equals "$other_hash" "$explicit_result" "jar_id with an explicit path must report the hash of that path, not fall back to \$JAR"
assert_equals "$staged_hash" "$explicit_result" "jar_id \"\$JAR_STAGED\" must report the hash of the staged jar, not fall back to \$JAR"
}
# fleetd #550 — closes a gap the test above leaves open. That test's own reference hash is now ALSO
@@ -684,77 +651,34 @@ test_jar_id_reports_unhashable_when_no_hasher_on_path() {
assert_equals "unhashable" "$explicit_result" "jar_id (explicit path) with no hasher on PATH must report the same third state"
}
# fleetd #664 — report_jar_state is the --check fix: under the old layout $JAR and the build
# output were the same file, so a single "jar on disk" fact covered both. Now they can disagree,
# and this is the function that is supposed to show that. Agreeing case: two files with IDENTICAL
# content must print both labels and never warn.
test_report_jar_state_agrees_when_hashes_match() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-agree"; mkdir -p "$dir"
BUILD_JAR="$dir/target-fleetd.jar"; JAR="$dir/run-fleetd.jar"
printf 'identical jar bytes' > "$BUILD_JAR"
printf 'identical jar bytes' > "$JAR"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'built jar' \
|| fail "report_jar_state did not label the built jar"
printf '%s' "$output" | grep -qF 'running jar' \
|| fail "report_jar_state did not label the running jar"
if printf '%s' "$output" | grep -qF 'differ'; then
fail "report_jar_state warned about a mismatch when both jars have identical content"
fi
}
# The disagreeing case: this is the whole point of the ticket — a built jar that is NOT the
# running jar must be visibly flagged, not silently printed as two unremarkable facts.
test_report_jar_state_warns_when_hashes_differ() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-differ"; mkdir -p "$dir"
BUILD_JAR="$dir/target-fleetd.jar"; JAR="$dir/run-fleetd.jar"
printf 'freshly built jar bytes' > "$BUILD_JAR"
printf 'older running jar bytes' > "$JAR"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'differ' \
|| fail "report_jar_state did not warn when the built jar and running jar disagree"
}
# Neither file existing (a fresh checkout, never built or deployed) must report two "absent"
# facts and never a false mismatch warning — "absent" vs "absent" is agreement, not a diff.
test_report_jar_state_both_absent_is_not_a_mismatch() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-absent"; mkdir -p "$dir"
BUILD_JAR="$dir/no-such-target.jar"; JAR="$dir/no-such-run.jar"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'absent' \
|| fail "report_jar_state did not report absent for a missing built/running jar"
if printf '%s' "$output" | grep -qF 'differ'; then
fail "report_jar_state warned about a mismatch when both jars are simply absent"
fi
}
# Reads JAR and BUILD_JAR exactly as the script sources them, with nothing here assigning
# either first. JAR must resolve outside $MODULE/target/, and JAR must differ from BUILD_JAR:
# the daemon's live path and Maven's own build output are never the same file.
test_jar_and_build_jar_are_sourced_outside_target_and_differ() {
source "$ROOT/scripts/redeploy-fleetd.sh"
case "$JAR" in
"$MODULE"/target/*)
fail "\$JAR must not live under \$MODULE/target/ — got $JAR" ;;
esac
[ "$JAR" != "$BUILD_JAR" ] \
|| fail "\$JAR and \$BUILD_JAR must not be the same path — got $JAR"
}
# fleetd #493/#664 — never build into the path a running process holds. swap_staged_jar is
# exercised directly against real files on disk (not stubs), because the whole point is file
# fleetd #493 — never build into the path a running process holds. stage_built_jar/swap_staged_jar
# are exercised directly against real files on disk (not stubs), because the whole point is file
# behavior (does the content move, does the source disappear, does a failure leave both sides
# intact) that a stubbed function cannot prove. stage_built_jar no longer exists: under the #664
# layout $BUILD_JAR (target/fleetd.jar) was never the live path, so a build has nothing to be
# staged OUT of — swap_staged_jar is called directly against $BUILD_JAR/$JAR (see
# test_swap_ordered_after_wait_and_before_start and the report_jar_state tests below for the rest
# of that seam).
# intact) that a stubbed function cannot prove.
test_stage_built_jar_moves_off_live_path() {
local dir jar staged saved_jar="$JAR" saved_staged="$JAR_STAGED"
dir="$TMP/stage-ok"; mkdir -p "$dir"
jar="$dir/fleetd.jar"; staged="$dir/fleetd-new.jar"
printf 'built jar bytes' > "$jar"
JAR="$jar"; JAR_STAGED="$staged"
stage_built_jar || fail "stage_built_jar rejected a real build output"
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ ! -f "$jar" ] || fail "stage_built_jar left the jar behind at the live path $jar"
[ -f "$staged" ] || fail "stage_built_jar did not create the staged jar at $staged"
grep -qF 'built jar bytes' "$staged" || fail "staged jar does not carry the built content"
}
test_stage_built_jar_dies_when_build_produced_nothing() {
local dir output rc=0 saved_jar="$JAR" saved_staged="$JAR_STAGED"
dir="$TMP/stage-missing"; mkdir -p "$dir"
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
output="$(stage_built_jar 2>&1)" || rc=$?
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ "$rc" -ne 0 ] || fail "stage_built_jar accepted a missing build output"
printf '%s' "$output" | grep -qF "$dir/fleetd.jar" \
|| fail "refusal message does not name the missing jar path"
}
test_swap_staged_jar_moves_staged_onto_live() {
local dir staged live
dir="$TMP/swap-ok"; mkdir -p "$dir"
@@ -1065,26 +989,24 @@ test_swap_ordered_after_wait_and_before_start() {
|| fail "swap_if_built (line $swap_line) is not before the start section (line $start_line)"
}
# fleetd #511: the drain-gate abort message (fired when a build has produced a jar but the
# operator declines the drain confirmation) used to tell the operator to "Rerun (with or without
# --no-build)" to finish the restart. That is wrong. fleetd #664 changed WHY it is wrong: under
# the old layout a build emptied the live path immediately, so --no-build's own check refused a
# rerun for you; now a build never touches the live path at all, so --no-build would NOT refuse —
# it would quietly restart the daemon on the OLD jar and throw away the one just built. Like
# test_swap_ordered_after_wait_and_before_start above, this code path is never reached by sourcing
# (the SOURCED guard stops before the main flow), so the only way to pin its exact wording is to
# read the source.
# fleetd #511: the drain-gate abort message (fired when a build has staged a jar but the operator
# declines the drain confirmation) used to tell the operator to "Rerun (with or without --no-build)"
# to finish the restart. That is wrong — by the time this message can fire, stage_built_jar has
# already moved the jar off $JAR, so a rerun WITH --no-build hits require_no_build_jar's own refusal
# ("no jar at $JAR — run without --no-build"). Like test_swap_ordered_after_wait_and_before_start
# above, this code path is never reached by sourcing (the SOURCED guard stops before the main flow),
# so the only way to pin its exact wording is to read the source.
test_drain_gate_abort_message_says_no_no_build() {
local src="$ROOT/scripts/redeploy-fleetd.sh" msg
msg="$(grep -A6 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
msg="$(grep -A3 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
[ -n "$msg" ] || fail "could not find the drain-gate staged-jar abort message in redeploy-fleetd.sh"
if printf '%s' "$msg" | grep -qF 'with or without --no-build'; then
fail "abort message still claims a rerun WITH --no-build can finish the restart"
fi
printf '%s' "$msg" | grep -qF 'WITHOUT --no-build' \
|| fail "abort message does not tell the operator to rerun without --no-build"
printf '%s' "$msg" | grep -qF 'silently discard' \
|| fail "abort message does not say that --no-build would silently discard the build just made"
printf '%s' "$msg" | grep -qF 'no longer at the live path' \
|| fail "abort message does not say why --no-build cannot finish the restart"
}
# fleetd #517 — the drain-gate abort branch itself. Before this, the only test of this message was
@@ -1227,22 +1149,19 @@ test_refuse_drain_gate_no_build_staged_absent() {
#
# fleetd #555 — the main flow's own call site moved: it used to read
# `refuse_drain_gate "$DO_BUILD" "$JAR_STAGED"` directly; it now reads
# `run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"` (fleetd #664 renamed the
# fourth argument from $JAR_STAGED to $BUILD_JAR — same role, the not-yet-swapped-in jar — when
# that path stopped being a separate staging file and became target/fleetd.jar itself), and
# run_drain_gate (tested directly below by test_run_drain_gate_*) is what calls refuse_drain_gate
# with its own local names. This grep now pins THAT call site — the thing that would go missing if
# a future edit deleted the main flow's call to run_drain_gate altogether, the same residual gap
# #521/#528 already accepted for swap_if_built/refuse_drain_gate (sourcing stops before the main
# flow runs, so no test in this file can do better than reading the source for this one specific
# gap).
# `run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"`, and run_drain_gate (tested
# directly below by test_run_drain_gate_*) is what calls refuse_drain_gate with its own local names.
# This grep now pins THAT call site — the thing that would go missing if a future edit deleted the
# main flow's call to run_drain_gate altogether, the same residual gap #521/#528 already accepted for
# swap_if_built/refuse_drain_gate (sourcing stops before the main flow runs, so no test in this file
# can do better than reading the source for this one specific gap).
#
# The grep ends `|| true`: this file runs under `set -euo pipefail`, so an ABSENT needle would fail
# the assignment and `set -e` would kill the whole suite before the `[ -n ... ] || fail` guard below
# ever ran — the exact dead-check shape fleetd #528 also flags as a sweep finding (see the PR body).
test_run_drain_gate_call_site_present() {
local src="$ROOT/scripts/redeploy-fleetd.sh" call_line
call_line="$(grep -Fn 'run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"' "$src" | head -1 | cut -d: -f1 || true)"
call_line="$(grep -Fn 'run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"' "$src" | head -1 | cut -d: -f1 || true)"
[ -n "$call_line" ] \
|| fail "could not find the main flow's run_drain_gate call site in redeploy-fleetd.sh"
}
@@ -1350,71 +1269,12 @@ test_run_drain_gate_declined_reply_refuses() {
source "$ROOT/scripts/redeploy-fleetd.sh"
}
# Writes a launchd-plist fixture naming jar_path as the ProgramArguments entry after "-jar", so
# check_jar_path_matches_plist has something real to read back.
write_launchd_plist_fixture() {
local path="$1" jar_path="$2"
cat > "$path" <<PLIST
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>test.fixture</string>
<key>ProgramArguments</key>
<array>
<string>/usr/bin/java</string>
<string>-jar</string>
<string>$jar_path</string>
<string>fleetd.yaml</string>
</array>
</dict>
</plist>
PLIST
}
# Agreeing case: a plist whose ProgramArguments names the same jar, resolved, must proceed and
# say so, never die.
test_check_jar_path_matches_plist_agrees_ok() {
local dir jar plist output rc=0
dir="$TMP/jar-path-agree"; mkdir -p "$dir/run"
jar="$dir/run/fleetd.jar"
plist="$dir/agree.plist"
write_launchd_plist_fixture "$plist" "$jar"
output="$(check_jar_path_matches_plist "$jar" "$plist" 2>&1)" || rc=$?
[ "$rc" -eq 0 ] \
|| fail "check_jar_path_matches_plist must succeed when the plist names the same jar: $output"
printf '%s' "$output" | grep -qF 'jar path check' \
|| fail "check_jar_path_matches_plist did not print the agreement line: $output"
}
# The disagreeing case, and the positive control this check exists for: a plist naming a
# different jar must die, naming both paths.
test_check_jar_path_matches_plist_mismatch_dies() {
local dir jar other_jar plist output rc=0
dir="$TMP/jar-path-mismatch"; mkdir -p "$dir/run" "$dir/target"
jar="$dir/run/fleetd.jar"
other_jar="$dir/target/fleetd.jar"
plist="$dir/mismatch.plist"
write_launchd_plist_fixture "$plist" "$other_jar"
output="$(check_jar_path_matches_plist "$jar" "$plist" 2>&1)" || rc=$?
[ "$rc" -ne 0 ] \
|| fail "check_jar_path_matches_plist must die when the plist names a different jar"
printf '%s' "$output" | grep -qF "$jar" \
|| fail "die message does not name this script's jar path: $output"
printf '%s' "$output" | grep -qF "$other_jar" \
|| fail "die message does not name the plist's jar path: $output"
}
# fleetd #555 item 4 — the report-state dispatch on $SUPERVISOR_KIND. Inverting this used to report
# the wrong supervisor and, for the launchd arm specifically, skip check_log_path_matches_plist.
CHECK_LOG_PATH_CALLED=0
CHECK_JAR_PATH_CALLED=0
stub_check_log_path_recorder() {
CHECK_LOG_PATH_CALLED=0
CHECK_JAR_PATH_CALLED=0
check_log_path_matches_plist() { CHECK_LOG_PATH_CALLED=1; }
check_jar_path_matches_plist() { CHECK_JAR_PATH_CALLED=1; }
}
test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path() {
@@ -1425,8 +1285,6 @@ test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path() {
[ "$SUPERVISED" = 1 ] || fail "report_supervisor_state launchd must set SUPERVISED=1"
[ "$CHECK_LOG_PATH_CALLED" = 1 ] \
|| fail "report_supervisor_state launchd must call check_log_path_matches_plist"
[ "$CHECK_JAR_PATH_CALLED" = 1 ] \
|| fail "report_supervisor_state launchd must call check_jar_path_matches_plist"
source "$ROOT/scripts/redeploy-fleetd.sh"
}
@@ -1438,8 +1296,6 @@ test_report_supervisor_state_systemd_sets_supervised_without_log_path_check() {
[ "$SUPERVISED" = 1 ] || fail "report_supervisor_state systemd must set SUPERVISED=1"
[ "$CHECK_LOG_PATH_CALLED" = 0 ] \
|| fail "report_supervisor_state systemd must NOT call check_log_path_matches_plist"
[ "$CHECK_JAR_PATH_CALLED" = 0 ] \
|| fail "report_supervisor_state systemd must NOT call check_jar_path_matches_plist"
source "$ROOT/scripts/redeploy-fleetd.sh"
}
@@ -2297,8 +2153,6 @@ test_assert_single_daemon_accepts_one_pid
test_assert_single_daemon_rejects_two_pids
test_running_pid_excludes_self_matching_wrapper_shell
test_running_pid_finds_a_real_java_named_second_process
test_running_pid_finds_a_real_java_named_process_from_target_dir
test_running_pid_excludes_self_matching_wrapper_shell_naming_target_dir
test_running_pid_drops_a_pid_whose_comm_is_not_java
test_running_pid_drops_a_pid_that_exited_before_the_comm_lookup
test_running_pid_counts_a_pid_whose_comm_is_java
@@ -2307,10 +2161,8 @@ test_jar_id_defaults_to_live_and_reports_explicit_path
test_hash256_computes_a_real_sha256
test_jar_id_reports_absent_for_missing_file
test_jar_id_reports_unhashable_when_no_hasher_on_path
test_report_jar_state_agrees_when_hashes_match
test_report_jar_state_warns_when_hashes_differ
test_report_jar_state_both_absent_is_not_a_mismatch
test_jar_and_build_jar_are_sourced_outside_target_and_differ
test_stage_built_jar_moves_off_live_path
test_stage_built_jar_dies_when_build_produced_nothing
test_swap_staged_jar_moves_staged_onto_live
test_swap_staged_jar_dies_without_staged_file
test_swap_staged_jar_dies_when_mv_fails
@@ -2352,8 +2204,6 @@ test_drain_confirmed_false_on_anything_else
test_run_drain_gate_skips_prompt_when_not_required
test_run_drain_gate_confirmed_reply_does_not_refuse
test_run_drain_gate_declined_reply_refuses
test_check_jar_path_matches_plist_agrees_ok
test_check_jar_path_matches_plist_mismatch_dies
test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path
test_report_supervisor_state_systemd_sets_supervised_without_log_path_check
test_report_supervisor_state_none_leaves_supervised_zero