Compare commits
134 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| dcd505286f | |||
| 60fa958da2 | |||
| 9950361bc9 | |||
| 9f74b6619a | |||
| f687046450 | |||
| c6058652be | |||
| 2a95da2ff4 | |||
| 7f9137fcb6 | |||
| 3bf3968bc7 | |||
| 261aa056f9 | |||
| 042b8c99dd | |||
| 4bfab6b718 | |||
| 008a457557 | |||
| 9494a6b99a | |||
| eb0557621e | |||
| a0eed6f01b | |||
| e2a91e883e | |||
| 62646957ea | |||
| 802c0ab701 | |||
| 4bab23e241 | |||
| a94262271b | |||
| 5c12865c25 | |||
| 3f8c956325 | |||
| 2051ceeb26 | |||
| 325d0771a4 | |||
| 7b97aae85b | |||
| 49a404ddf3 | |||
| 72d6a6878b | |||
| 4466ee0ef2 | |||
| 25ba7f16bb | |||
| 5ba69c9cf2 | |||
| 17052bb515 | |||
| e95ed99bf7 | |||
| 1477e4358a | |||
| 9e4e423ad6 | |||
| 6c2d6e93cb | |||
| 789b6a8716 | |||
| 01462c9695 | |||
| 8b4ff78546 | |||
| d4a2cd720c | |||
| 5a467e1f8b | |||
| 1fb6176783 | |||
| 49df79203c | |||
| 235644c0f0 | |||
| 7772b41993 | |||
| eccd0548ce | |||
| c4e23eebad | |||
| 7df503dfc2 | |||
| f5e02fedd6 | |||
| 29e7a06c49 | |||
| 92a96fcbd8 | |||
| 13482872bb | |||
| c1ca6273fc | |||
| 5f1b260c81 | |||
| 30d6872779 | |||
| 9d1306d442 | |||
| e54e3d87ea | |||
| bf027f10b9 | |||
| 3c5873dfe2 | |||
| e29227d5f4 | |||
| 45aca9eb3e | |||
| 2af13ab1ff | |||
| 0788d84be8 | |||
| 20c1094cbf | |||
| 9011c59b9f | |||
| c11ad71ed0 | |||
| bdcf285265 | |||
| d4f93a7b13 | |||
| cfebc575ea | |||
| 822327eed5 | |||
| 5289eb509f | |||
| e4c703a51a | |||
| 703a05db41 | |||
| 3f036b2a62 | |||
| 82fae94c55 | |||
| b1f34c2e6b | |||
| 3f807d9f1b | |||
| e70263062c | |||
| 1a1e586b62 | |||
| c16d118f09 | |||
| 4b10d02207 | |||
| c5fbfdbf4a | |||
| 12cff28abb | |||
| 77a6a7142e | |||
| 9f3671b801 | |||
| 84034b34d1 | |||
| 1e9b2c9b7e | |||
| 5d422f85fa | |||
| ed2027b202 | |||
| 6b0a99b2b7 | |||
| 6b7caba248 | |||
| f8b0d42a5c | |||
| d1e7d71eee | |||
| 7fd914df1a | |||
| b066eb1903 | |||
| a196d34455 | |||
| 051d320ea0 | |||
| 766772763f | |||
| eab8185d7b | |||
| 7f672f0fb8 | |||
| e2801b9bbc | |||
| ea02c7b248 | |||
| ce74e164c6 | |||
| e60f892efd | |||
| ce05886831 | |||
| be123d0ac7 | |||
| 2d09c8b027 | |||
| ed54f0224e | |||
| 8d5bc3ee89 | |||
| ab0cc71aa4 | |||
| 4cd9046353 | |||
| 4e98a74047 | |||
| a502ba53e0 | |||
| 446cc11d1d | |||
| ddd81fe174 | |||
| ae74cc081f | |||
| cf8da1d5fa | |||
| b8aedeafcb | |||
| 3d61af6f6f | |||
| ed99c209ac | |||
| af4c88d54b | |||
| 44c735f6f5 | |||
| b540a1744b | |||
| 0f51d53098 | |||
| c670792ffe | |||
| 3982ace544 | |||
| 7180b1aad0 | |||
| c26f695402 | |||
| 48877315ca | |||
| 5c08054533 | |||
| b6db9c31f5 | |||
| e7b33fe3a0 | |||
| e3e403e5c8 | |||
| 799014e99d |
@@ -0,0 +1,226 @@
|
||||
---
|
||||
name: handover
|
||||
description: Procedure for an outgoing lead to write the handover file that a fresh lead session inherits. Load this when your context is filling up and you are about to be replaced, whether you hand off by hand or fleetd does it for you. The file is the new lead's only inheritance — follow it exactly.
|
||||
---
|
||||
|
||||
# Handover — write the file the next lead depends on
|
||||
|
||||
A lead session fills up its context and has to be replaced by a fresh one. The outgoing lead
|
||||
writes a handover file, and the new session reads that file and carries on.
|
||||
|
||||
**There are two ways to hand off, and the file is the same either way.**
|
||||
|
||||
- **By hand.** You write the file, then tell the operator where it is. The operator starts the new
|
||||
session and points it at the file. This always works.
|
||||
- **With `fleet_handover`** (fleetd #480, merged 2026-09-11). You ask fleetd to do the swap: it
|
||||
checks the file, clears your pane, and tells the fresh session to read it. This needs
|
||||
`leadRollover:` in `fleetd.yaml`; without it every action answers a clean refusal naming
|
||||
`NOT_CONFIGURED`, and you fall back to the manual path. Section 11 below is the procedure.
|
||||
|
||||
Nothing else in this skill changes between the two. Only who performs the swap changes.
|
||||
|
||||
**The new lead's only inheritance is that file.** It does not see your conversation, your plan,
|
||||
or your screen. If the file is thin or wrong, the new lead re-derives what you already knew, and
|
||||
that wastes hours. Writing a good handover file is real work. It is not paperwork you rush
|
||||
through at the end of a session.
|
||||
|
||||
This skill is the procedure for writing it. Every rule below earned its place because a past
|
||||
handover got it wrong.
|
||||
|
||||
## 1. Confirm you are the right session to write this
|
||||
|
||||
Run `fleet_whoami` first. It must answer `primary`. Only a primary (lead) session writes a
|
||||
handover file. A worker's job ends with its own pull request, not a fleet-wide handoff.
|
||||
|
||||
## 2. Every number needs a command, run in this turn
|
||||
|
||||
A number is a claim: a count, a commit hash, a process id, a percentage, a queue depth. Before
|
||||
you write one, run the command that produces it — now, in this turn, against the live state.
|
||||
|
||||
Never take a number from:
|
||||
|
||||
- earlier in your own conversation — the state has moved since then,
|
||||
- a peer lead's report — that is their measurement, not yours,
|
||||
- your own memory of an earlier session.
|
||||
|
||||
Put the command, or its real output, next to the number. That lets the next lead re-run it and
|
||||
check it still matches. If you cannot measure something yourself, say so instead of guessing:
|
||||
"the fleet01 lead reports 91 commits behind; I have not checked this myself."
|
||||
|
||||
## 3. Say what you measured and what you did not
|
||||
|
||||
Mark every claim as one of two things:
|
||||
|
||||
- **"I checked this myself, in the code or on this host, at `<time>`."**
|
||||
- **"I did not check this myself; `<who>` reported it."**
|
||||
|
||||
Never present someone else's measurement as your own. This matters most for cross-host claims —
|
||||
a peer lead's daemon, a worker's report, or something the operator said earlier that you cannot
|
||||
re-verify from here.
|
||||
|
||||
## 4. Record open decisions, and who owns them
|
||||
|
||||
List three things:
|
||||
|
||||
- what the operator actually asked for, in their own words where you have them,
|
||||
- what is still unanswered,
|
||||
- any question you decided yourself instead of asking, with your reason.
|
||||
|
||||
Write the decision so it cannot be mistaken for the operator's instruction. Say plainly: "the
|
||||
operator never answered X; I decided Y, because Z." Without this, the next lead either silently
|
||||
reopens a closed question or assumes the operator chose something they never did.
|
||||
|
||||
## 5. Record live hazards
|
||||
|
||||
List anything that will break if the next lead does the obvious thing next. This includes:
|
||||
|
||||
- unpushed commits or unmerged branches,
|
||||
- a build, a spawn, or a redeploy still running,
|
||||
- code merged to `main` but not yet redeployed to the live daemon,
|
||||
- any trap that looks safe and is not — say what goes wrong and why, not only that something is
|
||||
"tricky."
|
||||
|
||||
## 6. Record what is explicitly not owed
|
||||
|
||||
List work that is finished, and work that another party has said they do not want touched. Name
|
||||
who said so and when. Without this line, the next lead re-does closed work or reopens a question
|
||||
a peer already declined to revisit.
|
||||
|
||||
## 7. Open the file with three re-measurement commands
|
||||
|
||||
The file's own first section must give the next lead three concrete commands to run before
|
||||
acting on anything else in the file:
|
||||
|
||||
1. confirm role — for example `fleet_whoami`,
|
||||
2. confirm the state of the working tree — for example `git status` and
|
||||
`git rev-list origin/main..HEAD`,
|
||||
3. read the live fleet — for example `fleet_list`.
|
||||
|
||||
Record what each command answered when you wrote the file, and tell the reader to run it again
|
||||
rather than trust your answer. The point of this section is that the reader checks live state
|
||||
before acting on any claim in the rest of the file, including yours.
|
||||
|
||||
## 8. Stamp the file with time and commit
|
||||
|
||||
Near the top of the file, write:
|
||||
|
||||
- the date and time you wrote it,
|
||||
- the commit the tree was on (`git rev-parse HEAD`),
|
||||
- whether the tree was clean (`git status`).
|
||||
|
||||
Without this, nobody can tell how old the file is, or which code it describes.
|
||||
|
||||
## 9. State plainly that the file goes stale fast
|
||||
|
||||
Say near the top: **re-measure anything you act on.** The file goes stale the moment anyone
|
||||
merges a branch, spawns a member, or restarts the daemon. Everything in the file is a snapshot
|
||||
of one moment, not a live fact.
|
||||
|
||||
## 10. What to leave out
|
||||
|
||||
Do not include:
|
||||
|
||||
- narration of how the session felt, or how hard something was,
|
||||
- anything the repo already records — code structure, git history, or a rule already written in
|
||||
`CLAUDE.md`. Point at it instead of repeating it,
|
||||
- advice that is only true for the session that is ending — a half-open terminal, a local
|
||||
variable, a train of thought with no state behind it.
|
||||
|
||||
A handover file is a record of state and decisions. It is not a diary.
|
||||
|
||||
## 11. Using `fleet_handover` (only if `leadRollover:` is configured)
|
||||
|
||||
**Run the three steps in this order. The order is not a style choice — the wrong order is
|
||||
refused.**
|
||||
|
||||
1. **`fleet_handover{action: "open", reason: "<why now>"}`.** It returns a `token` and the
|
||||
`handoverPath` you must write to. Nothing has happened to your pane yet.
|
||||
|
||||
**Write to exactly that path, and do not resolve it yourself.** It is always absolute, even when
|
||||
the operator configured a relative `handoverPath`: fleetd resolves a relative one against your
|
||||
own workspace before it hands it to you. The daemon and your pane can run in different
|
||||
directories, so a path you resolve yourself can point at a different file from the one the daemon
|
||||
will check.
|
||||
|
||||
**Check that the path is ignored by git before you write to it (#491).** A relative
|
||||
`handoverPath` resolves inside YOUR workspace, which is usually a repository — and usually not
|
||||
the `fleetd` one, so an ignore rule added to `fleetd` does not protect it. Run
|
||||
`grep -n handover <your workspace>/.gitignore`. No output means the file you are about to write
|
||||
will show up as untracked content in that repo. The file is a snapshot of live state and must
|
||||
never be committed, so tell the operator rather than committing it or silently editing their
|
||||
`.gitignore`.
|
||||
2. **Write the handover file at that path**, following sections 1–10 above.
|
||||
3. **Ask the operator, then `fleet_handover{action: "confirm", token, operatorConfirmed: true}`.**
|
||||
|
||||
Why that order: `confirm` refuses with `HANDOVER_STALE` unless the file was modified **after** the
|
||||
`open` request. That check stops a leftover file from an earlier session being accepted as this
|
||||
one's handover. So writing the file first and then calling `open` — the obvious order — always
|
||||
fails.
|
||||
|
||||
`{action: "cancel", token}` drops a pending request without rolling.
|
||||
|
||||
**Things that will surprise you:**
|
||||
|
||||
- **`accepted` does not mean your pane has been cleared.** It means every gate passed and the roll
|
||||
is scheduled to run once your current turn ends. Say your goodbye in the same turn — you will not
|
||||
get another one.
|
||||
- **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved
|
||||
from your connection, so you can only ever roll yourself.
|
||||
- **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you
|
||||
are confident. Ask, wait for the answer, then pass what they said. `requireOperatorConfirm`
|
||||
defaults to `true` and this is the only thing standing between a judgement call and a wiped
|
||||
session.
|
||||
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
|
||||
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
|
||||
- **The bootstrap prompt has never yet landed, and the fix is unproven (fleetd #489).** The first
|
||||
real rollover, on 2026-09-12, joined `/clear` and the bootstrap text into one line and Claude Code
|
||||
refused it as `Unknown command: /clearFresh`. The pane was never cleared and no context was lost,
|
||||
so the failure was safe — the roll simply did nothing. PR #490 fixed the cause and is deployed,
|
||||
but no roll has bootstrapped a fresh session end to end yet. **Assume it may still fail, and tell
|
||||
the operator so before you confirm.** The recovery is the same either way: the file is already
|
||||
written, so the operator starts a session and points it at the file. That is why you write the
|
||||
file before you confirm, and never the other way round.
|
||||
|
||||
## Writing style
|
||||
|
||||
Write in plain English. Use everyday words, one idea per sentence, and active voice. Keep every
|
||||
class, method, file, flag, and config key exactly as it appears in the code — replacing a
|
||||
precise term with a vague one makes the sentence wrong, not simpler. Explain an abbreviation the
|
||||
first time you use it.
|
||||
|
||||
If a diagram genuinely helps, put it in the `.md` file as a fenced ` ```mermaid ` block with no
|
||||
hardcoded colors, so it stays readable on light and dark backgrounds. Quote any label that has
|
||||
brackets, colons, or slashes.
|
||||
|
||||
## Template
|
||||
|
||||
```markdown
|
||||
# Handover — <fleet name> lead session, <date and time>
|
||||
|
||||
Written at commit `<output of git rev-parse HEAD>`. Tree was <clean, or dirty: `<git status
|
||||
summary>`>. Re-measure anything you act on — this file goes stale the moment anyone merges,
|
||||
spawns, or restarts.
|
||||
|
||||
## 0. Do these three things first
|
||||
|
||||
1. Confirm your role: `fleet_whoami` — must answer `primary`. (Answered `<result>` at `<time>`.)
|
||||
2. Confirm tree state: `git status`, `git rev-list origin/main..HEAD`. (`<result>` at `<time>`.)
|
||||
3. Read the live fleet: `fleet_list`. (`<result>` at `<time>`.)
|
||||
|
||||
## 1. What the operator asked for
|
||||
|
||||
<the live instructions, in their words where you have them; what is still open; any decision
|
||||
you made yourself, and why>
|
||||
|
||||
## 2. Open decisions, and who owns them
|
||||
|
||||
<one line per decision: who owns it, what is unanswered>
|
||||
|
||||
## 3. Live hazards
|
||||
|
||||
<one entry per hazard: what breaks, and why, if the next lead does the obvious thing>
|
||||
|
||||
## 4. What is not owed
|
||||
|
||||
<finished work, and work another party has declined; name who said so and when>
|
||||
```
|
||||
@@ -0,0 +1,72 @@
|
||||
---
|
||||
name: redeploy-fleetd
|
||||
description: Rebuild and restart the live fleetd daemon after a merge (lead / primary only). Load this before redeploying — it holds the script, the drain step, the permission grant, and the five checks that have each gone wrong here before. Workers must never do this.
|
||||
---
|
||||
|
||||
### Redeploying the daemon — the lead may do this (primary only)
|
||||
|
||||
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
|
||||
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
|
||||
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
|
||||
rather than hand the job back to the operator.
|
||||
|
||||
Workers must never do this. A worker has no business restarting the daemon it is talking through,
|
||||
and stopping it kills the worker's own channel mid-turn.
|
||||
|
||||
**Use the script — do not hand-roll the steps.**
|
||||
|
||||
```bash
|
||||
scripts/redeploy-fleetd.sh --check # report state, change nothing
|
||||
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
|
||||
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
|
||||
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
|
||||
```
|
||||
|
||||
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
|
||||
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
|
||||
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
|
||||
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
|
||||
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
|
||||
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
|
||||
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
|
||||
|
||||
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
|
||||
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
|
||||
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
|
||||
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
|
||||
|
||||
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
|
||||
if the script is unavailable or a step fails, this is what it was protecting you from.
|
||||
|
||||
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
|
||||
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
|
||||
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
|
||||
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
|
||||
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
|
||||
whether the name resolves without ever printing the value.
|
||||
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
|
||||
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
|
||||
rendezvous, and a member's report is not recoverable once its ticket is gone.
|
||||
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
|
||||
it. The startup log names which keys it accepted and which it deferred — read those lines rather
|
||||
than assuming.
|
||||
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
|
||||
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
|
||||
demoted to worker, which refuses every orchestration call.
|
||||
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
|
||||
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
|
||||
the outside.
|
||||
|
||||
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
|
||||
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
|
||||
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
|
||||
start every time. The rule lives in the operator's Claude Code settings:
|
||||
|
||||
```json
|
||||
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
|
||||
```
|
||||
|
||||
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
|
||||
running the stop and start as separate commands — that is exactly the approval the script replaced.
|
||||
Say what you were going to run and why, and let the operator decide.
|
||||
|
||||
+10
-4
@@ -87,12 +87,18 @@ jobs:
|
||||
apt-get update && apt-get install -y --no-install-recommends maven
|
||||
mvn -version
|
||||
|
||||
# The `contract` profile clears the default-excludes group, so the @Tag("contract") AMQP test
|
||||
# runs against the RabbitMQ service container (AMQP_URI). Pinned to the one contract test to
|
||||
# avoid re-running the unit suite already covered by the `build` job.
|
||||
# The `contract` profile clears the default-excludes group, so `-Dgroups=contract` runs every
|
||||
# @Tag("contract") test and nothing from the unit suite the `build` job already covered — a
|
||||
# tag selects the whole group, so a test added to it later runs here automatically. A prior
|
||||
# version of this step pinned `-Dtest=AmqpReplyInboxContractTest` by class name instead: that
|
||||
# silently excluded every other contract test (including the herdr ones) from CI, and nobody
|
||||
# noticed until the herdr protocol drifted out from under a test that never ran here
|
||||
# (fleetd #449). If this runner has no herdr socket, the herdr-backed tests in the group
|
||||
# skip on their own `assumeTrue` and only the broker-backed ones actually run — check the
|
||||
# step output rather than assuming which.
|
||||
- name: Contract tests
|
||||
working-directory: fleetd
|
||||
run: mvn -B -Pcontract test -Dtest=AmqpReplyInboxContractTest
|
||||
run: mvn -B -Pcontract test -Dgroups=contract
|
||||
|
||||
- name: Failing test output
|
||||
if: failure()
|
||||
|
||||
@@ -19,3 +19,9 @@
|
||||
fleetd.out
|
||||
fleetd/fleetd.out
|
||||
logs/
|
||||
|
||||
# fleetd #480: the lead rollover handover file. `leadRollover.handoverPath` points here, and the
|
||||
# outgoing lead rewrites it on every rollover. It is a snapshot of one moment's live state —
|
||||
# unpushed branches, running builds, open questions — so it is stale the moment it is written and
|
||||
# has no business in git history.
|
||||
.handover/
|
||||
|
||||
@@ -58,8 +58,9 @@ and the sender silently receives nothing. Fail toward the recoverable error.
|
||||
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
|
||||
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
|
||||
deliverable, and a send waits on that gate for ~60s and then fails without ever reaching its pane.
|
||||
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
|
||||
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
|
||||
5. **Never move a fleet session, pane or peer except through the bridge.** The bridge owns policy;
|
||||
the multiplexer owns PTYs. Any route that changes fleet state without the bridge's checks
|
||||
bypasses every rule above — the `herdr` CLI and its socket are the usual example.
|
||||
|
||||
### Primary (lead) — run this on every task, in order
|
||||
|
||||
@@ -137,9 +138,11 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
|
||||
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
|
||||
| Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
|
||||
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
|
||||
| Answer a peer lead that messaged you | `fleet_reply{content}` — the one case a lead replies |
|
||||
| Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}` if they are on this host. **Not** `fleet_reply`: it has no peer route and the publish is refused |
|
||||
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
|
||||
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
|
||||
| Tear down a member | `fleet_stop{paneId}` |
|
||||
| Replace your OWN lead session when its context is full | `fleet_handover{action:"open", reason?}` → write the handover file it names → `fleet_handover{action:"confirm", token, operatorConfirmed}`. Primary-only. **In that order**: the file must be modified *after* `open`, or `confirm` refuses it as stale. There is no terminal parameter — the pane is always your own, so you can never roll another lead. `{action:"cancel", token}` drops a pending request |
|
||||
|
||||
### Lead ↔ lead — coordinate, never delegate
|
||||
|
||||
@@ -163,15 +166,21 @@ The traffic between leads is coordination and nothing else:
|
||||
3. **Verify a peer exactly as you verify yourself.** Peer status buys nothing: check the claim
|
||||
against the code, and re-run the build. A peer's correction gets the same treatment — right or
|
||||
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
|
||||
**N observations are N data points only if they differ in the axis you are trusting.** This cuts
|
||||
both ways. N *failures* blamed on one cause are one data point when the cases share what you are
|
||||
not varying. N *agreeing measurements* are also one data point when they share an instrument —
|
||||
two hosts, two operators and the same formula is one formula, not two confirmations.
|
||||
4. **Ask a peer to read your project addendum.** Your addendum is instruction surface: every future
|
||||
session on your host obeys it, and a wrong one is obeyed just as faithfully as a right one. The
|
||||
author is the worst reader of their own qualifier placement — measured here, one addendum carried
|
||||
two defects and a non-author found both. If you have no peer, at least re-read it asking "which
|
||||
sentence goes false first, and would a reader reach the caveat before acting?"
|
||||
|
||||
Being messaged by a peer does not make you its worker: answer with `fleet_reply`, and push back on
|
||||
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
|
||||
you.
|
||||
Being messaged by a peer does not make you its worker: answer the way you would open —
|
||||
`fleet_send{coordId}` for another daemon, `fleet_send{sessionId}` on this host — and push back on
|
||||
the substance if it is wrong. `fleet_reply` resolves a member's blocked `fleet_send`; a peer's
|
||||
coord-id message is durable and non-blocking, so there is nothing for it to resolve. A peer that
|
||||
simply complies has thrown away the reason there are two of you.
|
||||
|
||||
### Member (worker or architect) — the turn contract
|
||||
|
||||
@@ -216,6 +225,19 @@ must obey belongs in the charter, not here.
|
||||
- **This repo is the bridge.** The daemon is `fleetd`, its MCP mount is `http://127.0.0.1:8765/mcp`,
|
||||
and the code behind the rules above is `mcp/FleetMcp` (tools), `auth/Authz` (the role table),
|
||||
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
|
||||
- **Herdr socket tests (measured 2026-09-10).** In this repo, herdr is a subject under test. A
|
||||
worker assigned to herdr code, and the lead, may let a test open the herdr socket directly in a
|
||||
throwaway workspace that the test tears down. This only covers
|
||||
`fleetd/src/test/java/dev/ltms/fleet/herdr/AgentControlContractTest.java`,
|
||||
`fleetd/src/test/java/dev/ltms/fleet/herdr/HerdrContractTest.java`,
|
||||
`fleetd/src/test/java/dev/ltms/fleet/herdr/PaneLocatorContractTest.java`, and
|
||||
`fleetd/src/test/java/dev/ltms/fleet/herdr/WorkspacePlacementContractTest.java`. It is not a
|
||||
general licence. Using the herdr CLI or socket to move a real fleet session, pane, or peer stays
|
||||
banned. That is the control plane that invariant 5 protects. Re-measure with
|
||||
`grep -rl 'UnixSocketHerdrClient.connect()' fleetd/src/test/java --include='*.java'`. A non-empty
|
||||
result means tests still open the socket and this note still applies. An empty result means nobody
|
||||
does this any more; delete this section. Canonical invariant 5 restatement is tracked in #458 and
|
||||
is not part of this change.
|
||||
- **`fleet_profiles`/`fleet_list` report two separate outage states, and they are not the same
|
||||
thing.** *Quarantined* (CB-578) means the backend told us it is out of capacity — a long,
|
||||
1800s-default cooldown. *Cooling off* (fleetd #201/#227) means a profile's credential threw two
|
||||
@@ -236,8 +258,11 @@ must obey belongs in the charter, not here.
|
||||
already cost three workers' turns: each wrote a good report to its terminal and ended the turn
|
||||
with no `fleet_reply`, and the scrape returned the tail of the brief instead.
|
||||
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
|
||||
`port-to-opencode` (make an OpenCode session a participant in this workspace) and
|
||||
`fleets-status` (report every fleet that shares one LavinMQ instance).
|
||||
`port-to-opencode` (make an OpenCode session a participant in this workspace),
|
||||
`fleets-status` (report every fleet that shares one LavinMQ instance),
|
||||
`redeploy-fleetd` (rebuild and restart the live daemon after a merge) and
|
||||
`handover` (write the file a fresh lead session inherits when the outgoing one hands off,
|
||||
fleetd #480).
|
||||
- **This repo is also a Claude Code marketplace, and ships a plugin.** `.claude-plugin/marketplace.json`
|
||||
points at `plugin/`, which carries the MCP mount and the `setup` skill
|
||||
(`/claude-bridge:setup` — make any project bridge-ready). It was added in CB-527 and then went
|
||||
@@ -271,60 +296,14 @@ must obey belongs in the charter, not here.
|
||||
### Redeploying the daemon — the lead may do this (primary only)
|
||||
|
||||
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
|
||||
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
|
||||
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
|
||||
rather than hand the job back to the operator.
|
||||
feature merged to `main` does nothing until the daemon is rebuilt and restarted. The lead **may and
|
||||
should** redeploy rather than hand the job back to the operator. **Workers must never do this** — a
|
||||
worker has no business restarting the daemon it is talking through, and stopping it kills the
|
||||
worker's own channel mid-turn.
|
||||
|
||||
Workers must never do this. A worker has no business restarting the daemon it is talking through,
|
||||
and stopping it kills the worker's own channel mid-turn.
|
||||
|
||||
**Use the script — do not hand-roll the steps.**
|
||||
|
||||
```bash
|
||||
scripts/redeploy-fleetd.sh --check # report state, change nothing
|
||||
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
|
||||
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
|
||||
```
|
||||
|
||||
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
|
||||
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
|
||||
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
|
||||
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
|
||||
|
||||
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
|
||||
if the script is unavailable or a step fails, this is what it was protecting you from.
|
||||
|
||||
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
|
||||
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
|
||||
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
|
||||
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
|
||||
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
|
||||
whether the name resolves without ever printing the value.
|
||||
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
|
||||
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
|
||||
rendezvous, and a member's report is not recoverable once its ticket is gone.
|
||||
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
|
||||
it. The startup log names which keys it accepted and which it deferred — read those lines rather
|
||||
than assuming.
|
||||
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
|
||||
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
|
||||
demoted to worker, which refuses every orchestration call.
|
||||
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
|
||||
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
|
||||
the outside.
|
||||
|
||||
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
|
||||
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
|
||||
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
|
||||
start every time. The rule lives in the operator's Claude Code settings:
|
||||
|
||||
```json
|
||||
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
|
||||
```
|
||||
|
||||
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
|
||||
running the stop and start as separate commands — that is exactly the approval the script replaced.
|
||||
Say what you were going to run and why, and let the operator decide.
|
||||
**Load the `redeploy-fleetd` skill before you redeploy.** It holds `scripts/redeploy-fleetd.sh`
|
||||
and its flags, the drain step, the operator's permission grant, and the five checks that have each
|
||||
gone wrong here before. Do not hand-roll the steps from memory.
|
||||
|
||||
### The prompt is part of the product — update it with the code (mandatory)
|
||||
|
||||
@@ -372,6 +351,16 @@ print("in sync:", w[i:w.index("\n```\n", i) + 1] == block)
|
||||
PY
|
||||
```
|
||||
|
||||
**Only the lead can run that check (measured 2026-09-10).** A member's provisioned worktree has
|
||||
`wiki/` uninitialized, so the script dies with `FileNotFoundError: wiki/7-Use-Cases.md`. Measured
|
||||
in three worker worktrees: `git submodule status` printed a leading `-` and `wiki/` held 0
|
||||
entries; the primary's own clone printed a leading `+` and the file was there. So never make this
|
||||
check a member's acceptance criterion — it is unsatisfiable for them, and a brief that asks for it
|
||||
is asking a worker to invent a pass. A member told to check it must say it could not run it, and
|
||||
must never report it as passed. The lead runs it in the main clone before merging. Re-measure with
|
||||
`git submodule status` in a member's worktree: a leading `-` means this still applies; once it
|
||||
prints a commit with no `-`, delete this paragraph.
|
||||
|
||||
## IDE MCP tools & validation workflow (enforced)
|
||||
|
||||
> **Primary only.** Workers have no IDE MCP mount — if you are a worker, skip this section and
|
||||
|
||||
+111
-17
@@ -88,6 +88,47 @@ bind:
|
||||
# backoffMs: 60000
|
||||
# quietNudgeCap: 3
|
||||
|
||||
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
|
||||
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
|
||||
# own pane and bootstrap a fresh session against that file.
|
||||
#
|
||||
# Opt-in on purpose — it clears the lead's own pane on request, so upgrading the daemon must never
|
||||
# acquire that ability for you. Absent block = feature off, and nothing is constructed at all. Even
|
||||
# once present, nothing but an explicit confirm() call — one that passes every check — can ever
|
||||
# cause a /clear: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
|
||||
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so
|
||||
# it cannot clear the pane inline (that pane is still WORKING); instead it schedules a one-shot
|
||||
# continuation that waits for the SAME confirm() call's turn to end, then does the actual work. See
|
||||
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order (fleetd #480 correction).
|
||||
#
|
||||
# handoverPath: REQUIRED when this block is present — where the handover file a fresh lead session
|
||||
# reads must live. No default (an operator-specific path); a present block with no
|
||||
# handoverPath refuses to start. May be relative: it then resolves against the
|
||||
# CALLING lead's own fleet.leaders.<name>.cwd (falling back to the daemon's own
|
||||
# working directory when that lead has none configured) — never against whatever
|
||||
# directory the daemon process happens to have been started in. An absolute path is
|
||||
# used unchanged. Prefer an absolute path if the daemon and the lead's pane might not
|
||||
# share a working directory (fleetd #480 follow-up).
|
||||
# requireOperatorConfirm: true # default true — confirm() refuses unless the caller also passes
|
||||
# # operatorConfirmed: true
|
||||
# maxDocAgeSeconds: 3600 # default 3600 — refuse a handover file older than this
|
||||
# turnSettleSeconds: 20 # default 20 — how long the deferred roll waits for the CALLING
|
||||
# # lead's own turn to end (its pane to report injectable again)
|
||||
# # before sending /clear at all. If this elapses, /clear is NEVER
|
||||
# # sent — a lead that never goes idle is still doing real work.
|
||||
# clearSettleSeconds: 20 # default 20 — how long to wait for the pane to become injectable
|
||||
# # again AFTER /clear before giving up (never sends bootstrapText
|
||||
# # if this elapses). A separate, second wait from turnSettleSeconds.
|
||||
# bootstrapText: "..." # default names the RESOLVED (absolute) handoverPath — sent to
|
||||
# # the lead once its pane settles after /clear
|
||||
# leadRollover:
|
||||
# handoverPath: /path/to/handover.md
|
||||
# requireOperatorConfirm: true
|
||||
# maxDocAgeSeconds: 3600
|
||||
# turnSettleSeconds: 20
|
||||
# clearSettleSeconds: 20
|
||||
# bootstrapText: "Fresh lead session: read the handover file and carry on."
|
||||
|
||||
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
|
||||
# per tick.
|
||||
# intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes
|
||||
@@ -207,8 +248,12 @@ herdrSocket: ~/.config/herdr/herdr.sock
|
||||
# omit and this profile's completion fallback behaves exactly as before.
|
||||
# Every backend words its refusal differently, so this is config, never a
|
||||
# vendor string baked into fleetd itself.
|
||||
# DEFERRED: compiled once into a startup pattern map — editing it needs a
|
||||
# daemon restart, same as this profile's model/baseUrl/argv.
|
||||
# HOT (fleetd #446): read live, cached by profile name, at every
|
||||
# completion-fallback check AND by fleet_profiles' exhaustionDetectionArmed —
|
||||
# editing it and reloading arms or disarms usage-limit detection for this
|
||||
# profile with no daemon restart. (Before fleetd #446 this was DEFERRED,
|
||||
# compiled once into a startup pattern map like model/baseUrl/argv still are —
|
||||
# see errorPattern below, which is still deferred that way on purpose.)
|
||||
# credentialId → CB-578 stage B: the credential this profile quarantines WITH when a
|
||||
# BACKEND_EXHAUSTED classification fires. Two profiles that set the SAME
|
||||
# credentialId share one quarantine — the case this exists for is two models
|
||||
@@ -227,8 +272,9 @@ herdrSocket: ~/.config/herdr/herdr.sock
|
||||
# happens, just without a profile-specific match; every backend words its
|
||||
# failure differently, so a hardcoded sentence would only ever match one
|
||||
# of them.
|
||||
# DEFERRED: compiled once into a startup pattern map, same as exhaustedPattern
|
||||
# — editing it needs a daemon restart.
|
||||
# DEFERRED: compiled once into a startup pattern map — editing it needs a
|
||||
# daemon restart. Unlike exhaustedPattern above (made hot by fleetd #446),
|
||||
# errorPattern was scoped out of that ticket on purpose and stays deferred.
|
||||
# # errorPattern: "503 Service Unavailable" # opt-in: classify a backend outage
|
||||
#
|
||||
# What happens once a match fires (BackendOutagePolicy, credentialId-keyed,
|
||||
@@ -447,6 +493,15 @@ placement: weighted
|
||||
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
|
||||
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
|
||||
# 1800 (30 minutes) when omitted or non-positive.
|
||||
#
|
||||
# fleetd #466: this is now only the BASE of an escalating backoff, not a flat retry rate. A
|
||||
# credential quarantined again within one base cooldown of the previous quarantine ending (still
|
||||
# reporting exhausted — e.g. a weekly subscription limit that hasn't reset) backs off further:
|
||||
# cooldown doubles each such time, capped at 12x this value (~6 hours at the 1800s default). A
|
||||
# quarantine that starts after a base-cooldown's worth of quiet resets back to this value. Not
|
||||
# configurable per se — the multiplier and ceiling are constants in BackendQuarantine, not new
|
||||
# YAML keys; see its class doc for the exact formula and why there is no automatic probe to clear
|
||||
# it early (the operator's own design constraint — a probe spends the quota it's measuring).
|
||||
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
|
||||
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
|
||||
# restart. Editing this needs a daemon restart to take effect.
|
||||
@@ -468,9 +523,10 @@ placement: weighted
|
||||
# when the reload happens — not about how important the key is:
|
||||
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
|
||||
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad
|
||||
# / credentialId. Those are hot because the placement policy (and, for credentialId,
|
||||
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
|
||||
# not by itself enough to make a key hot.
|
||||
# / credentialId / exhaustedPattern. Those are hot because the placement policy (and,
|
||||
# for credentialId, the CB-578 stage B quarantine check; for exhaustedPattern, fleetd
|
||||
# #446's LiveExhaustedPatterns) reads them through a supplier — being config is not by
|
||||
# itself enough to make a key hot.
|
||||
# EXCEPT `fleet.leaders`: Fleetd.main reads it once at startup to build the lead tab
|
||||
# scanner and launcher, and neither is rebuilt on reload. A changed/added/removed
|
||||
# `fleet.leaders` entry is silently accepted — the reload reports "config reloaded"
|
||||
@@ -482,10 +538,10 @@ placement: weighted
|
||||
# stage B — baked once into the quarantine tracker built at startup), ADDING or
|
||||
# REMOVING a profile (a new backend needs its own launcher, and launchers are built
|
||||
# once), AND an existing profile's launch settings — model, baseUrl, argv, env,
|
||||
# configDir, mcpUrl, tabLabel, exhaustedPattern, errorPattern (fleetd #201 / #227 —
|
||||
# compiled once into a startup pattern map the same way exhaustedPattern is). The
|
||||
# launcher takes a copy of `profiles:` at startup and resolves every spawn out of
|
||||
# that copy, so those never
|
||||
# configDir, mcpUrl, tabLabel, errorPattern (fleetd #201 / #227 — compiled once into
|
||||
# a startup pattern map; exhaustedPattern used to be compiled the same way until
|
||||
# fleetd #446 made it hot — see above). The launcher takes a copy of `profiles:` at
|
||||
# startup and resolves every spawn out of that copy, so those never
|
||||
# reach a launch until you restart. The reload logs them by name rather than
|
||||
# pretending they applied.
|
||||
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
|
||||
@@ -787,12 +843,21 @@ guard:
|
||||
# .claude/skills/, so a member spawned against ANY repo — not only one that already ships its own
|
||||
# copy — can load a bridge skill (e.g. implementer). Unset (the default): no worktree is touched
|
||||
# beyond today's behaviour. A skill folder the target repo already carries under
|
||||
# .claude/skills/<name> is never overwritten — the repo's own copy always wins. Claude Code
|
||||
# members only; an opencode member reads a different path (.opencode/agent) this key does not
|
||||
# touch. Best-effort like worktreeGroup above: a missing/unreadable directory here is logged and
|
||||
# skipped, never a failed spawn. Every non-hidden subdirectory of this directory is copied
|
||||
# wholesale, with no per-file allowlist — don't park scratch files or drafts alongside the real
|
||||
# skill folders, they will be copied into every provisioned worktree too.
|
||||
# .claude/skills/<name> is never overwritten — the repo's own copy always wins. Best-effort like
|
||||
# worktreeGroup above: a missing/unreadable directory here is logged and skipped, never a failed
|
||||
# spawn. Every non-hidden subdirectory of this directory is copied wholesale, with no per-file
|
||||
# allowlist — don't park scratch files or drafts alongside the real skill folders, they will be
|
||||
# copied into every provisioned worktree too.
|
||||
#
|
||||
# fleetd #393: which member KINDS actually consume this once it is copied. kind: claude-code —
|
||||
# the Claude Code CLI discovers .claude/skills/ on its own; nothing else is needed. kind: opencode
|
||||
# — opencode has no such discovery, so OpenCodeLauncher reads whatever landed under
|
||||
# .claude/skills/ and appends each seeded skill's SKILL.md to the generated instructions[] file
|
||||
# (opencode's only channel for static guidance text; unlike Claude Code's Skill tool, the content
|
||||
# is always part of the system prompt, not loaded on demand). Both kinds are covered as of #393 —
|
||||
# earlier builds copied the files for every kind but only claude-code could read them, and the
|
||||
# seeding log said "N of M" regardless. Check the per-spawn launcher log (not just the seeding
|
||||
# log) to see what a given member actually got.
|
||||
# memberSkills: /path/to/fleetd/checkout/.claude/skills
|
||||
|
||||
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
|
||||
@@ -877,3 +942,32 @@ guard:
|
||||
# terminal: term_65619bd6174568
|
||||
# pushReminders: 5
|
||||
# pushBackoffMs: 15000
|
||||
|
||||
# Central allow-list of models any profiles: entry may name. Nothing checked a profile's model:
|
||||
# value before this block existed — it was a free-form string handed straight to the backend
|
||||
# adapter, and a withdrawn or misspelled name failed silently instead of at config load (opencode
|
||||
# falls back to a default model rather than erroring on an unknown -m).
|
||||
#
|
||||
# Absent, or present with an empty allow:, is OFF: no profile's model: is checked, exactly like
|
||||
# before this block existed. fleetd.yaml is gitignored on every host, so an upgrade must not force
|
||||
# every operator to enumerate their models before the daemon will start.
|
||||
#
|
||||
# The list is the authority; profiles: is checked against it, never the reverse — adding or
|
||||
# editing a profiles: entry cannot, by itself, widen what is permitted here.
|
||||
#
|
||||
# Enforcement is at CONFIG LOAD only (a bad model: fails the daemon at startup, naming both the
|
||||
# model and the profile). There is no spawn-time enforcement, no runtime on/off switch, and no
|
||||
# interaction with BackendQuarantine — those are separate, later units.
|
||||
#
|
||||
# allow → the permitted models. Each entry is its own block (not a bare string) so a later unit
|
||||
# can add an on/off state or a load limit per model without changing this shape.
|
||||
# model → the model id exactly as a profiles: entry's model: field would write it. One flat,
|
||||
# opaque-string namespace: a bare Claude id (claude-sonnet-5) and an opencode
|
||||
# provider-prefixed id (openai/gpt-5.6-terra) both fit here unchanged — the check is a
|
||||
# plain string match, never a parse of the provider prefix or a branch on kind:.
|
||||
# models:
|
||||
# allow:
|
||||
# - model: claude-sonnet-5
|
||||
# - model: claude-opus-5
|
||||
# - model: openai/gpt-5.6-terra
|
||||
# - model: amazon.nova-pro-v1:0
|
||||
|
||||
@@ -10,6 +10,7 @@ import dev.ltms.fleet.herdr.HerdrException;
|
||||
import dev.ltms.fleet.herdr.HerdrRouter;
|
||||
import dev.ltms.fleet.herdr.LeadTabScanner;
|
||||
import dev.ltms.fleet.lead.LeadLauncher;
|
||||
import dev.ltms.fleet.lead.LeadRollover;
|
||||
import dev.ltms.fleet.herdr.PaneLocator;
|
||||
import dev.ltms.fleet.herdr.UnixSocketHerdrClient;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
@@ -18,6 +19,7 @@ import dev.ltms.fleet.inject.BackendErrorSink;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
|
||||
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.inject.StatusPoller;
|
||||
import dev.ltms.fleet.inject.TurnListener;
|
||||
@@ -25,6 +27,7 @@ import dev.ltms.fleet.inject.MemberPresence;
|
||||
import dev.ltms.fleet.auth.MemberRegistry;
|
||||
import dev.ltms.fleet.auth.CallerResolver;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import dev.ltms.fleet.mcp.CharterToolSurface;
|
||||
import dev.ltms.fleet.mcp.ConnectionIdentity;
|
||||
import dev.ltms.fleet.metrics.FleetMetrics;
|
||||
import dev.ltms.fleet.metrics.Metrics;
|
||||
@@ -65,10 +68,13 @@ import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Optional;
|
||||
import java.util.Set;
|
||||
import java.util.TreeSet;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
@@ -77,6 +83,7 @@ import java.util.function.Function;
|
||||
import java.util.function.Predicate;
|
||||
import java.util.function.Supplier;
|
||||
import java.util.regex.Pattern;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
/**
|
||||
* {@code fleetd} entry point. Wires the real herdr socket client to the REST app and
|
||||
@@ -134,29 +141,44 @@ public final class Fleetd {
|
||||
// secret is reported above, so upgrading past this commit never silently drops CB-592's
|
||||
// protection.
|
||||
reportMemberCredentialsGap(cfg);
|
||||
// fleetd #395: an unset exhaustedPattern is a silent opt-out of usage-limit detection for
|
||||
// that profile — say so loudly, the same way the two reports above do, rather than let an
|
||||
// operator discover it only when a limit goes undetected.
|
||||
reportExhaustedPatternGap(cfg);
|
||||
// CB-559: `cfg` stays the startup snapshot — every validation and every piece of one-time
|
||||
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
|
||||
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
|
||||
// contract; adding a reader here does not make a key reloadable by itself.
|
||||
ConfigRef config = new ConfigRef(configPath, cfg);
|
||||
// fleetd #474: pass the charter/tool-surface check in as ConfigRef's extraValidation, so
|
||||
// ConfigRef#reload() runs the same gate main() runs below, without dev.ltms.fleet.config
|
||||
// gaining a dependency on dev.ltms.fleet.mcp — Fleetd is the seam that already holds both.
|
||||
ConfigRef config = new ConfigRef(configPath, cfg, Fleetd::assertChartersNameOnlyRegisteredTools);
|
||||
|
||||
// The primary/host env that launched fleetd must not be tainted.
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
guard.assertPrimaryClean(System.getenv());
|
||||
|
||||
// CB-501: refuse to start if the bind is wider than the auth mode can defend. Under
|
||||
// loopback-trust, "not a known worker" means "the primary" — sound only because the OS
|
||||
// refuses remote connections to a loopback socket. This throws rather than warns so the
|
||||
// dangerous configuration cannot be reached by ignoring a log line.
|
||||
cfg.validateAuthExposure();
|
||||
cfg.validateLeadTabPrefixes();
|
||||
// CB-542: a subscription:true profile whose env: reseats ANTHROPIC_BASE_URL/AUTH_TOKEN would
|
||||
// reach an unguarded endpoint (the launcher skips SubscriptionGuard for it). Refuse at load.
|
||||
cfg.validateSubscriptionProfiles();
|
||||
cfg.validateCharters();
|
||||
// CB-548: every architect slot must name a configured workers: profile — the strong-model
|
||||
// backend the future spawn lifecycle would read. A stale reference dies here, not later.
|
||||
cfg.validateMembers();
|
||||
// Every FleetConfig.validateXxx() the operator's config can fail — CB-501's auth-exposure
|
||||
// check, CB-531's lead-tab-prefix check, CB-542's subscription-profile check, the charter
|
||||
// and member-slot checks, and the "central allow-list of usable models" check — must run
|
||||
// here, at load, before anything below opens a socket or spawns a member. fleetd ticket
|
||||
// "central allow-list of usable models" follow-up: mutation testing found six individual
|
||||
// calls here with nothing proving any of them still ran (deleting one left the full suite
|
||||
// green). validateAll() replaces them with the one call that FleetConfigValidateAllTest
|
||||
// and the Fleetd-startup tests actually pin — see FleetConfig#validateAll's javadoc for
|
||||
// why a name-by-name list here would have the same defect it replaces.
|
||||
cfg.validateAll();
|
||||
// fleetd #469, follow-up to #464: validateAll() (and validateCharters() inside it) only
|
||||
// checks that a charter's KEY is a role wire name and its text is non-blank — it never
|
||||
// looks at what the text actually names. This is the separate check that does: it asks
|
||||
// dev.ltms.fleet.mcp.FleetTool (the canonical registered-tool set) whether every fleet_*/
|
||||
// bridge_* token a charter names is a tool this server actually registers. It cannot live
|
||||
// inside FleetConfig#validateCharters() — config loads before the MCP server exists, and
|
||||
// must not gain a dependency on the mcp package — so it runs here instead, at the one seam
|
||||
// that already holds both a loaded FleetConfig and the mcp package, before anything below
|
||||
// opens a socket or spawns a member. fleetd #474: the same check is also wired into `config`
|
||||
// above as ConfigRef's extraValidation, so a reload refuses what this line refuses at startup.
|
||||
assertChartersNameOnlyRegisteredTools(cfg);
|
||||
|
||||
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
|
||||
? Path.of(cfg.herdrSocket())
|
||||
@@ -216,7 +238,13 @@ public final class Fleetd {
|
||||
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
|
||||
// The cooldown is deferred (see FleetConfig#quarantineCooldownSeconds): it is read once
|
||||
// here, at startup, and a config reload only changes it for a daemon restart.
|
||||
BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,
|
||||
// fleetd #466: escalating, not flat — a credential that keeps reporting exhaustion (e.g. a
|
||||
// weekly subscription limit, which would otherwise be retried on every ~30-minute cooldown,
|
||||
// about 336 times across the week) backs off further each consecutive time, capped at
|
||||
// BackendQuarantine.DEFAULT_MAX_COOLDOWN_MULTIPLE x the base cooldown. See BackendQuarantine's
|
||||
// class doc for the mechanism, why this never fires on cooling-off (a separate, unescalated
|
||||
// mechanism — BackendOutagePolicy below), and the reset.
|
||||
BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,
|
||||
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
|
||||
// fleetd #201 Unit 5: one outage-cool-off tracker for the whole daemon, shared between the
|
||||
// launcher (checked at spawn, like `quarantine` above) and the backend-error sink wired in
|
||||
@@ -230,6 +258,13 @@ public final class Fleetd {
|
||||
profileName -> liveCountRef.get().apply(profileName),
|
||||
quarantine,
|
||||
outagePolicy);
|
||||
// fleetd #422 follow-up: say which of the three model-gate states the daemon booted into —
|
||||
// no models: block at all, a block armed with nothing off, or a block with N off — the same
|
||||
// way exhaustedPatternCoverageLine/errorPatternCoverageLine report CB-578 stage A/fleetd
|
||||
// #201 Unit 5 coverage just below. Read from workers.modelGateState() (never a separate
|
||||
// config.get().models() here) so this line and fleet_profiles' modelGateArmed can never
|
||||
// disagree about what CompositePeerLauncher's spawn gate actually enforces.
|
||||
log.info("model gate (fleetd #422): {}", modelGateCoverageLine(workers.modelGateState()));
|
||||
// CB-504: under supervision (launchd/systemd) fleetd can start before herdr's socket
|
||||
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
|
||||
// the first thing that actually talks to herdr, so without this wait a boot-order race
|
||||
@@ -351,7 +386,11 @@ public final class Fleetd {
|
||||
// pane resolves to an architect until the later spawn lifecycle binds one. The registry is
|
||||
// what CallerResolver resolves against and what that lifecycle will read profiles from;
|
||||
// nothing here spawns a slot.
|
||||
MemberRegistry members = new MemberRegistry(cfg.fleet());
|
||||
// fleetd #424: MemberRegistry.live re-reads fleet.architects through `config` on every
|
||||
// reserve/requireSlotFor call, so a reload that removes or adds an architect slot governs
|
||||
// the next spawn with no restart — the frozen `new MemberRegistry(cfg.fleet())` this used
|
||||
// to be let a "revoked" slot keep granting new architect spawns forever.
|
||||
MemberRegistry members = MemberRegistry.live(() -> config.get().fleet());
|
||||
sessions.setMemberLifecycle(members);
|
||||
if (!members.slots().isEmpty()) {
|
||||
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
|
||||
@@ -365,29 +404,34 @@ public final class Fleetd {
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
// CB-578 stage A: classify a completion-fallback scrape that matches a profile's configured
|
||||
// usage-limit refusal as BACKEND_EXHAUSTED rather than handing it back as a real answer.
|
||||
// Compiled once at startup, keyed by profile name; a profile with no exhaustedPattern is
|
||||
// simply absent here, so its workers keep today's completion-fallback behaviour unchanged.
|
||||
Map<String, Pattern> exhaustedPatternsByProfile = new LinkedHashMap<>();
|
||||
cfg.profiles().forEach((name, profile) -> {
|
||||
if (profile.hasExhaustedPattern()) {
|
||||
exhaustedPatternsByProfile.put(name, Pattern.compile(profile.exhaustedPattern()));
|
||||
}
|
||||
});
|
||||
// fleetd #446: read live off `config` per lookup, cached by profile name — see
|
||||
// LiveExhaustedPatterns's class doc for why this replaces the old compiled-once-at-startup
|
||||
// map. A profile with no exhaustedPattern simply returns null here, so its workers keep
|
||||
// today's completion-fallback behaviour unchanged.
|
||||
LiveExhaustedPatterns liveExhaustedPatterns = new LiveExhaustedPatterns(() -> config.get().profiles());
|
||||
ExhaustedPatternLookup exhaustedPatterns = target -> sessions.roster().stream()
|
||||
.filter(session -> target.equals(session.terminalId()))
|
||||
.findFirst()
|
||||
.map(session -> exhaustedPatternsByProfile.get(session.profile()))
|
||||
.map(session -> liveExhaustedPatterns.patternFor(session.profile()))
|
||||
.orElse(null);
|
||||
// The startup coverage line still reports the boot-time snapshot only — it is printed once,
|
||||
// here, and a reload no longer needs to change what it said; exhaustionDetectionArmed (via
|
||||
// liveExhaustedPatterns.armed, wired into quarantineSource below) is what stays live.
|
||||
Set<String> exhaustedConfiguredAtStartup = cfg.profiles().entrySet().stream()
|
||||
.filter(e -> e.getValue().hasExhaustedPattern())
|
||||
.map(Map.Entry::getKey)
|
||||
.collect(Collectors.toCollection(LinkedHashSet::new));
|
||||
log.info("backend-exhausted classification (CB-578 stage A): {}",
|
||||
CompletionResolver.coverage("exhaustedPattern", cfg.profiles().keySet(),
|
||||
exhaustedPatternsByProfile.keySet()));
|
||||
exhaustedPatternCoverageLine(cfg.profiles().keySet(), exhaustedConfiguredAtStartup));
|
||||
// fleetd #201 Unit 5: classify a completion-fallback scrape that matches a profile's
|
||||
// configured backend-error refusal (a credential outage, a provider 5xx) as a backend error
|
||||
// rather than handing it back as a real answer. Compiled once at startup, keyed by profile
|
||||
// name, mirroring exhaustedPatternsByProfile above — a profile with no configured
|
||||
// errorPattern is simply absent here, so CompletionResolver falls back to its built-in
|
||||
// narrow {@code (?i)\bAPI Error\s*:} compatibility pattern for that profile's targets
|
||||
// (BackendErrorPatternLookup#legacy's contract — see backendErrorPatterns below).
|
||||
// rather than handing it back as a real answer. Still compiled once at startup, keyed by
|
||||
// profile name — unlike exhaustedPattern above (fleetd #446), errorPattern was left DEFERRED
|
||||
// on purpose: the ticket that made exhaustedPattern hot scoped errorPattern/cooling-off out
|
||||
// explicitly. A profile with no configured errorPattern is simply absent here, so
|
||||
// CompletionResolver falls back to its built-in narrow {@code (?i)\bAPI Error\s*:}
|
||||
// compatibility pattern for that profile's targets (BackendErrorPatternLookup#legacy's
|
||||
// contract — see backendErrorPatterns below).
|
||||
Map<String, Pattern> errorPatternsByProfile = new LinkedHashMap<>();
|
||||
cfg.profiles().forEach((name, profile) -> {
|
||||
if (profile.hasErrorPattern()) {
|
||||
@@ -400,52 +444,25 @@ public final class Fleetd {
|
||||
BackendErrorPatternLookup backendErrorPatterns = backendErrorPatternLookup(sessions::roster,
|
||||
errorPatternsByProfile);
|
||||
log.info("backend-error classification (fleetd #201 Unit 5): {}",
|
||||
CompletionResolver.coverage("errorPattern", cfg.profiles().keySet(),
|
||||
errorPatternsByProfile.keySet()));
|
||||
errorPatternCoverageLine(cfg.profiles().keySet(), errorPatternsByProfile.keySet()));
|
||||
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
|
||||
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
|
||||
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
|
||||
// the profile config live off `config`, so a credentialId edit is hot: no restart needed.
|
||||
//
|
||||
// fleetd #175: this sink is now also the quarantine target for OpenCodeLauncher's
|
||||
// model-mismatch check — it needs nothing profile-specific from the caller beyond `target`
|
||||
// (a herdr terminal id) and `reason`, so reusing it here is exactly "the existing
|
||||
// ExhaustionSink path", not a new mechanism.
|
||||
//
|
||||
// fleetd #234: that check fires from SessionAwareHandle.agentSessionId(), which runs during
|
||||
// SessionManager.acquire() BEFORE this session is registered in sessions.roster() — so the
|
||||
// roster-only lookup below used to find nothing, .ifPresent silently no-op'd, and the
|
||||
// ERROR the check had just logged ("quarantining this profile's credential") was a lie:
|
||||
// nothing was quarantined, and nothing said so. Two changes: (1) OpenCodeLauncher now
|
||||
// passes its OWN profile name via ExhaustionSink's 3-arg overload — it already has the
|
||||
// FleetConfig.Profile in hand and does not need the roster at all — used here as a
|
||||
// fallback whenever the roster lookup misses; (2) if a profile still cannot be resolved
|
||||
// (neither the roster nor the hint names a configured one), this logs loudly at ERROR
|
||||
// instead of silently doing nothing — a control that cannot act must say so.
|
||||
// fleetd #234, round 4: the 3-arg overload is now ExhaustionSink's single abstract method,
|
||||
// so this is safely a lambda — there is no separate 2-arg overload left for it to bind to
|
||||
// instead and silently drop profileHint (that was rounds 1-3's whole hazard).
|
||||
ExhaustionSink exhaustionSink = (target, reason, profileHint) -> {
|
||||
String profileName = sessions.roster().stream()
|
||||
.filter(session -> target.equals(session.terminalId()))
|
||||
.findFirst()
|
||||
.map(MemberSession::profile)
|
||||
.orElse(profileHint);
|
||||
FleetConfig.Profile profile = profileName == null ? null : config.get().profiles().get(profileName);
|
||||
if (profile == null) {
|
||||
log.error("quarantine requested for target '{}' ({}) but no profile could be "
|
||||
+ "resolved — the target is not (yet) in the roster, and {} — "
|
||||
+ "credential NOT quarantined (fleetd #234)",
|
||||
target, reason,
|
||||
profileHint == null ? "no profile hint was given"
|
||||
: "the hinted profile '" + profileHint + "' is not configured");
|
||||
return;
|
||||
}
|
||||
String credentialId = profile.effectiveCredentialId();
|
||||
quarantine.quarantine(credentialId);
|
||||
log.warn("credential '{}' quarantined for {}s (profile '{}'): {}", credentialId,
|
||||
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
|
||||
};
|
||||
// fleetd #446 criterion 3: the backend text that triggered the most recent quarantine of
|
||||
// each credential, so fleet_profiles can report WHY a limit was hit, not only that it was.
|
||||
// Keyed by credential id — the same key BackendQuarantine's own remainingSeconds uses —
|
||||
// and written at the one call site that actually quarantines (inside exhaustionSink below),
|
||||
// so a reason can never be reported for a quarantine that never happened. Bounded the same
|
||||
// way BackendQuarantine's own internal map is documented to be: by the number of distinct
|
||||
// credentials ever exhausted, not by how often the config is edited.
|
||||
Map<String, String> quarantineReasonByCredential = new ConcurrentHashMap<>();
|
||||
// fleetd #446 follow-up (round 3): extracted into exhaustionSink(...) below — see that
|
||||
// method's javadoc for the full fleetd #175/#234/#446 history this used to carry inline —
|
||||
// so a dedicated test can drive the exact ExhaustionSink main() builds, not a hand-rebuilt
|
||||
// copy of its shape.
|
||||
ExhaustionSink exhaustionSink = exhaustionSink(sessions, config, quarantine,
|
||||
quarantineReasonByCredential, cfg);
|
||||
// fleetd #175: point the forwarding sink handed to OpenCodeLauncher above at the real one,
|
||||
// now that `sessions` exists to resolve target -> session -> profile.
|
||||
exhaustionSinkRef.set(exhaustionSink);
|
||||
@@ -574,6 +591,16 @@ public final class Fleetd {
|
||||
heartbeat = null;
|
||||
heartbeatScheduler.shutdownNow();
|
||||
}
|
||||
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed, so
|
||||
// an upgraded daemon cannot silently acquire the ability to clear the lead's own pane.
|
||||
// Unlike heartbeat above, this has no recurring scheduler of its own — nothing but an
|
||||
// explicit confirm() call (wired to an MCP tool by a later ticket; nothing calls it yet)
|
||||
// that passes every gate can ever schedule a roll. It does not take primaryRegistry: the
|
||||
// lead terminal to roll comes from the caller of open()/confirm() (resolved by the MCP
|
||||
// layer from the connection, the same way auth/CallerResolver#resolve builds a
|
||||
// Principal.leader(...)), never from a single-slot lookup — see LeadRollover's class
|
||||
// javadoc, fleetd #480 correction 2.
|
||||
LeadRollover leadRollover = leadRollover(cfg, router.leadAgents(), config, leads);
|
||||
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
|
||||
pushLoop, metrics);
|
||||
|
||||
@@ -661,21 +688,16 @@ public final class Fleetd {
|
||||
// built a second time — two independently-constructed sources reading the SAME BackendQuarantine
|
||||
// / BackendOutagePolicy would still be able to drift (e.g. a future edit to the credentialIdFor
|
||||
// closure in only one of the two places), exactly the shape #284 was.
|
||||
FleetMcp.QuarantineSource quarantineSource = new FleetMcp.QuarantineSource(profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.effectiveCredentialId();
|
||||
}, quarantine);
|
||||
FleetMcp.QuarantineSource quarantineSource = quarantineSource(config, quarantine,
|
||||
liveExhaustedPatterns, quarantineReasonByCredential);
|
||||
FleetMcp.OutageSource outageSource = new FleetMcp.OutageSource(profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.effectiveCredentialId();
|
||||
}, outagePolicy);
|
||||
|
||||
FleetMcp mcp = new FleetMcp(messages, workers, sessions, identity, presence,
|
||||
primaryRegistry, callers, metrics, new FleetMcp.CapacitySource(profile -> liveCountRef.get().apply(profile),
|
||||
profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.maxLoad();
|
||||
}, () -> config.get().profiles().keySet(), System::nanoTime),
|
||||
primaryRegistry, callers, metrics,
|
||||
capacitySource(config, cfg, profile -> liveCountRef.get().apply(profile)),
|
||||
new FleetMcp.HealthCoverageSource(() -> {
|
||||
var health = config.get().health();
|
||||
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
|
||||
@@ -689,7 +711,10 @@ public final class Fleetd {
|
||||
// reach. Read from the SAME snapshot leadMailbox itself opened from (cfg.coordinator()),
|
||||
// not the live config.get() — coordinator wiring is already boot-time-fixed (see
|
||||
// leadMailbox above), so peers follows the same rule rather than half hot-reloading.
|
||||
cfg.coordinator() == null ? List.of() : cfg.coordinator().peers());
|
||||
cfg.coordinator() == null ? List.of() : cfg.coordinator().peers(),
|
||||
// fleetd #480 Unit C: the executor behind fleet_handover — null whenever
|
||||
// leadRollover: is not configured (see the leadRollover local above).
|
||||
leadRollover);
|
||||
|
||||
// CB-637: the receive half. Only constructed when a lead mailbox actually opened — with no
|
||||
// coordinator (or an unreachable one) there is nothing to deliver, so no scheduler is
|
||||
@@ -798,6 +823,223 @@ public final class Fleetd {
|
||||
return target -> presence.isPresent(target) || leads.get().containsKey(target);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #446 follow-up (round 3): the quarantine {@link ExhaustionSink} main() actually wires
|
||||
* — extracted out of {@code main} for the same reason {@link #quarantineSource} below and
|
||||
* {@link #usageLimitFixWarning}/{@link #usageLimitFixWarningNoModel} were. Round 2 pinned the
|
||||
* WARNING text by having {@code FleetdUsageLimitFixWarningTest} call those two methods
|
||||
* directly, but a mutation battery proved that left two gaps: nothing proved this sink's
|
||||
* {@code log.warn} call actually uses either method (replacing the whole ternary with a
|
||||
* literal string), and nothing proved it picks the right one for a profile WITH a configured
|
||||
* {@code model:} versus one WITHOUT (swapping the ternary's two branches) — both mutations
|
||||
* left round 2's full 1592-test suite green. {@code FleetdExhaustionSinkWarningTest} drives
|
||||
* THIS factory's return value directly and asserts on the real text a {@code ListAppender}
|
||||
* attached to this class's own logger captures — the only way to prove the caller, not just
|
||||
* the two callees in isolation.
|
||||
*
|
||||
* <p>Behaviourally unchanged from the inline lambda this replaces:
|
||||
* <ul>
|
||||
* <li>fleetd #175: also the quarantine target for {@code OpenCodeLauncher}'s model-mismatch
|
||||
* check — it needs nothing profile-specific beyond {@code target}/{@code reason}, so
|
||||
* reusing this sink is the existing {@code ExhaustionSink} path, not a new mechanism;</li>
|
||||
* <li>fleetd #234 (round 4): resolves {@code target} to a profile via the live session
|
||||
* roster first, falling back to {@code profileHint} — {@code
|
||||
* SessionAwareHandle.agentSessionId()} fires this sink during {@code
|
||||
* SessionManager.acquire()}, <em>before</em> that session is registered, so the
|
||||
* roster-only lookup alone used to silently miss it. A profile resolved neither way
|
||||
* logs loudly at ERROR instead of doing nothing;</li>
|
||||
* <li>fleetd #446 criterion 3: writes {@code quarantineReasonByCredential} at the one call
|
||||
* site that actually quarantines, keyed by credential id — see the field's declaration
|
||||
* in {@code main} for the bound on its size.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>{@code cfg} is the startup {@link FleetConfig} snapshot, read only for {@code
|
||||
* quarantineCooldownSeconds()} in the log text — a Cold key (see {@code ConfigRef}'s class
|
||||
* doc), so reading it off the snapshot rather than {@code config.get()} makes no observable
|
||||
* difference and matches what the inline version already did.
|
||||
*/
|
||||
static ExhaustionSink exhaustionSink(SessionManager sessions, ConfigRef config, BackendQuarantine quarantine,
|
||||
Map<String, String> quarantineReasonByCredential, FleetConfig cfg) {
|
||||
return (target, reason, profileHint) -> {
|
||||
String profileName = sessions.roster().stream()
|
||||
.filter(session -> target.equals(session.terminalId()))
|
||||
.findFirst()
|
||||
.map(MemberSession::profile)
|
||||
.orElse(profileHint);
|
||||
FleetConfig.Profile profile = profileName == null ? null : config.get().profiles().get(profileName);
|
||||
if (profile == null) {
|
||||
log.error("quarantine requested for target '{}' ({}) but no profile could be "
|
||||
+ "resolved — the target is not (yet) in the roster, and {} — "
|
||||
+ "credential NOT quarantined (fleetd #234)",
|
||||
target, reason,
|
||||
profileHint == null ? "no profile hint was given"
|
||||
: "the hinted profile '" + profileHint + "' is not configured");
|
||||
return;
|
||||
}
|
||||
String credentialId = profile.effectiveCredentialId();
|
||||
quarantine.quarantine(credentialId);
|
||||
quarantineReasonByCredential.put(credentialId, reason);
|
||||
log.warn("credential '{}' quarantined for {}s (profile '{}'): {}", credentialId,
|
||||
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
|
||||
// fleetd #446 criterion 2: name the fix, not just the fact — an operator reading this
|
||||
// should not have to work out which of several configured models to touch.
|
||||
String model = profile.model();
|
||||
log.warn(model != null && !model.isBlank()
|
||||
? usageLimitFixWarning(profile.profile(), model)
|
||||
: usageLimitFixWarningNoModel(profile.profile(), cfg.quarantineCooldownSeconds()));
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #404, superseded by fleetd #446: production source for quarantine reporting.
|
||||
* Credential IDs were already hot; {@code exhaustedPattern} used to be compiled once at
|
||||
* startup for {@link CompletionResolver}, so the armed field had to read that same frozen map
|
||||
* until restart — the exact asymmetry #446 exists to close. Both {@code exhaustedPatternArmed}
|
||||
* here and {@link CompletionResolver}'s own classification now read the ONE live {@link
|
||||
* LiveExhaustedPatterns} instance ({@code exhaustedPatterns.armed}/{@code .patternFor}), so a
|
||||
* reload that arms or disarms a profile's detection is visible to both at once — never two
|
||||
* independently-updated copies that could disagree, the fleetd #404 rule this method used to
|
||||
* violate on purpose and now upholds.
|
||||
*
|
||||
* <p>{@code modelFor} and {@code reasonFor} back fleetd #446 criterion 3 — {@code
|
||||
* fleet_profiles}'s quarantined row naming which model a quarantined profile runs
|
||||
* ({@code modelFor}, read live off {@code config} the same way {@code credentialIdFor} already
|
||||
* is) and the backend text that triggered the most recent quarantine of that credential
|
||||
* ({@code reasonFor}, backed by {@code reasonByCredential} — see its call site in {@code main}
|
||||
* for where that map is written, at the one place a credential is actually quarantined).
|
||||
*/
|
||||
static FleetMcp.QuarantineSource quarantineSource(ConfigRef config, BackendQuarantine quarantine,
|
||||
LiveExhaustedPatterns exhaustedPatterns, Map<String, String> reasonByCredential) {
|
||||
return new FleetMcp.QuarantineSource(profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.effectiveCredentialId();
|
||||
}, quarantine, exhaustedPatterns::armed, profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.model();
|
||||
}, reasonByCredential::get);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #415 (review follow-up): package-private factory for the CB-578 stage A {@code
|
||||
* exhaustedPattern} startup coverage line, paired explicitly with {@link
|
||||
* CompletionResolver.UnsetMeaning#OFF} — {@code exhaustedPattern} has no fallback, so a
|
||||
* profile with none configured really does have the classification off.
|
||||
*
|
||||
* <p>Extracted out of {@code main} for the same reason {@link #capacitySource} and {@link
|
||||
* #worktreeBranchLookup} were: {@code coverage()}'s own tests ({@code CompletionResolverTest})
|
||||
* prove it words {@code OFF} and {@link CompletionResolver.UnsetMeaning#BUILT_IN_DEFAULT}
|
||||
* correctly when a test supplies the meaning itself — they cannot prove {@code main} pairs the
|
||||
* right meaning with the right key, which is the actual fleetd #415 defect. <b>Measured:</b>
|
||||
* swapping the {@code UnsetMeaning} arguments between this method and {@link
|
||||
* #errorPatternCoverageLine} — recreating #415's defect with the two keys exchanged — compiled
|
||||
* with 0 errors and left all 1506 existing tests green before {@code
|
||||
* FleetdPatternCoverageLineTest} was added to catch exactly that swap.
|
||||
*/
|
||||
/**
|
||||
* fleetd #446 follow-up: the criterion-2 WARNING text — "name the fix, not just the fact" —
|
||||
* for a profile whose {@code model:} is configured. Extracted out of the {@code
|
||||
* exhaustionSink} lambda the same way {@link #exhaustedPatternCoverageLine} was extracted out
|
||||
* of {@code main}: the inline version compiled, ran, and read correctly, but nothing pinned
|
||||
* its text, so a later edit could silently stop naming the fix and every test would stay
|
||||
* green (measured: renaming the leading {@code "usage-limit fix:"} tag left all 1586
|
||||
* pre-follow-up tests passing). {@link FleetdUsageLimitFixWarningTest} asserts on the return
|
||||
* value of this method directly, which is exactly what the caller logs — {@code log.warn} is
|
||||
* called with this method's result as a single, already-formatted argument, so the string a
|
||||
* test sees here is byte-identical to what {@code fleetd.out} receives.
|
||||
*
|
||||
* <p>Deliberately asserts on three substrings, not the whole sentence (the profile name, the
|
||||
* model name, and the literal {@code enabled: false} / {@code models.allow} action) — see that
|
||||
* test's class doc for why a whole-sentence pin is the wrong granularity here.
|
||||
*/
|
||||
static String usageLimitFixWarning(String profile, String model) {
|
||||
return "usage-limit fix: profile '" + profile + "' runs model '" + model + "' — set "
|
||||
+ "`enabled: false` on that model's entry under models.allow in fleetd.yaml to "
|
||||
+ "stop new spawns landing on it (models: is hot, no restart needed); remove the "
|
||||
+ "line again once the subscription window resets";
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #446 follow-up: the {@link #usageLimitFixWarning} counterpart for a profile with no
|
||||
* {@code model:} configured — {@code models.allow} gates by model name, so there is nothing
|
||||
* for an operator to flip, and this says so instead of naming a fix that does not exist.
|
||||
*/
|
||||
static String usageLimitFixWarningNoModel(String profile, Integer quarantineCooldownSeconds) {
|
||||
return "usage-limit fix: profile '" + profile + "' has no model: configured, so "
|
||||
+ "models.allow cannot gate it by name — the " + quarantineCooldownSeconds
|
||||
+ "s quarantine above is the only thing keeping new spawns off it for now";
|
||||
}
|
||||
|
||||
static String exhaustedPatternCoverageLine(Set<String> allProfiles, Set<String> configuredProfiles) {
|
||||
return CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
|
||||
allProfiles, configuredProfiles);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #415 (review follow-up): the {@code errorPattern} counterpart of {@link
|
||||
* #exhaustedPatternCoverageLine}, paired explicitly with {@link
|
||||
* CompletionResolver.UnsetMeaning#BUILT_IN_DEFAULT} — an unset {@code errorPattern} still runs
|
||||
* backend-error classification against {@code CompletionResolver}'s built-in {@code
|
||||
* BACKEND_ERROR} pattern, so the empty case is not "off". See {@link
|
||||
* #exhaustedPatternCoverageLine}'s javadoc for the measured swap mutation this pairing guards
|
||||
* against.
|
||||
*/
|
||||
static String errorPatternCoverageLine(Set<String> allProfiles, Set<String> configuredProfiles) {
|
||||
return CompletionResolver.coverage("errorPattern", CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT,
|
||||
allProfiles, configuredProfiles);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up: package-private factory for the startup line reporting which of the
|
||||
* three central {@code models.allow:} gate states the daemon booted into. Extracted the same
|
||||
* way {@link #exhaustedPatternCoverageLine}/{@link #errorPatternCoverageLine} are, so a
|
||||
* dedicated test can call it directly rather than parsing log output, and so {@code main}'s
|
||||
* only source for this line is {@link PeerLauncher#modelGateState()} — never a second,
|
||||
* independently-derived read of {@code cfg.models()} that could disagree with what {@code
|
||||
* CompositePeerLauncher}'s spawn gate actually enforces (the fleetd #404 lesson).
|
||||
*
|
||||
* <p>Unlike the two pattern-key lines above, there is no {@code UnsetMeaning} choice to make
|
||||
* here: {@link PeerLauncher.ModelGateState#configured()} already states, unambiguously, whether
|
||||
* an empty {@link PeerLauncher.ModelGateState#off()} means "no {@code models:} block to gate
|
||||
* with" or "a block armed and currently reporting zero off" — the exact two states a bare
|
||||
* {@code disabledModels()} read could not tell apart before this ticket.
|
||||
*/
|
||||
static String modelGateCoverageLine(PeerLauncher.ModelGateState state) {
|
||||
if (!state.configured()) {
|
||||
return "not configured (no models: block — nothing is gated, and nothing can be)";
|
||||
}
|
||||
Set<String> off = state.off();
|
||||
return off.isEmpty()
|
||||
? "armed (models: block present; 0 models currently turned off)"
|
||||
: "armed (" + off.size() + " model(s) turned off: " + new TreeSet<>(off) + ")";
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #416: production source for {@code fleet_list}'s per-profile capacity facts.
|
||||
*
|
||||
* <p>The profile <em>set</em> ({@code configuredProfiles}) must come from {@code cfg} — the
|
||||
* startup snapshot — not the live {@code config.get()}. {@code profiles} as a whole is a
|
||||
* {@code DEFERRED} key ({@link ConfigRef#DEFERRED_KEYS}): {@code HerdrPeerLauncher} takes
|
||||
* {@code Map.copyOf(profiles)} once at construction and a profile only added to the
|
||||
* hot-reloaded map can never actually be spawned, so enumerating it live made {@code fleet_list}
|
||||
* report a profile as available when {@code fleet_spawn} on that same profile fails with
|
||||
* {@code unknown worker profile}. {@code fleet_list}'s own contract for {@code free} is "the
|
||||
* same check the spawn gate itself runs" — the set the spawn gate can see is the startup one,
|
||||
* so this must enumerate that one too, the same shape as {@code coordinator.peers} above.
|
||||
*
|
||||
* <p>{@code maxLoad} stays live on purpose: it is read off {@code config.get()} exactly like
|
||||
* {@code credentialId} ({@link ConfigRef} documents both as hot), so an existing profile's
|
||||
* {@code maxLoad} edit must still change what {@code fleet_list} reports without a restart.
|
||||
*/
|
||||
static FleetMcp.CapacitySource capacitySource(ConfigRef config, FleetConfig cfg,
|
||||
Function<String, Integer> liveCount) {
|
||||
return new FleetMcp.CapacitySource(liveCount,
|
||||
profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.maxLoad();
|
||||
},
|
||||
cfg.profiles()::keySet, System::nanoTime);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #248: package-private factory for the member worktree/branch lookup {@link
|
||||
* CompletionResolver} uses to name a fallback report's worktree and branch (fleetd#241).
|
||||
@@ -825,6 +1067,68 @@ public final class Fleetd {
|
||||
.orElse(null);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #480: construct the {@link LeadRollover} executor only when {@code leadRollover:} is
|
||||
* present at startup — the same presence gate {@code leadHeartbeat:} uses just above this
|
||||
* call site in {@code main}. Extracted to its own factory, the same reason
|
||||
* {@link #worktreeBranchLookup} and {@link #exhaustionSink} are: a unit test can call this
|
||||
* directly with a fabricated {@link FleetConfig} (absent block ⇒ {@code null}, present block ⇒
|
||||
* constructed) without needing the whole of {@code main}, and a source-text test on the real
|
||||
* call site proves {@code main} still calls this factory rather than inlining a copy that could
|
||||
* silently diverge.
|
||||
*
|
||||
* <p>The returned object reads every {@code leadRollover:} field fresh on each {@code open()}/
|
||||
* {@code confirm()} call through {@code () -> config.get().leadRollover()} — see {@code
|
||||
* ConfigRef}'s class doc Hot bullet and {@link FleetConfig.LeadRollover}'s javadoc for why that
|
||||
* makes the block's fields HOT despite this presence gate being evaluated once, at startup.
|
||||
*
|
||||
* <p>Deliberately does NOT take {@link PrimaryRegistry}: fleetd #480 correction 2 found that a
|
||||
* single-slot lookup lets one lead's {@code confirm()} clear a DIFFERENT lead's pane on a
|
||||
* daemon with more than one labelled lead tab. The lead terminal to roll instead comes from
|
||||
* whoever calls {@code open()}/{@code confirm()} — the later MCP-tool unit must resolve it from
|
||||
* the connection and pass it in, never take it as a request field. See {@link LeadRollover}'s
|
||||
* class javadoc.
|
||||
*
|
||||
* <p><strong>fleetd #480 follow-up:</strong> also builds the terminal → lead-workspace lookup
|
||||
* {@link LeadRollover#open} needs to resolve a relative {@code handoverPath} against the
|
||||
* CALLING lead's own {@code cwd} rather than the daemon's — the daemon and a lead's pane can
|
||||
* have different working directories (this repo nests {@code fleetd/} inside its own root, so
|
||||
* they already differ on this host). The lookup is terminal → lead name (via {@code
|
||||
* liveLeadTerminals}) → that lead's {@code cwd} (via {@code config.get().fleet().leaders()}),
|
||||
* and both hops are read LIVE on every call, never off a snapshot taken here: leads are
|
||||
* discovered by a live tab scan ({@code LeadTabScanner}), so a map captured at construction
|
||||
* time could be empty (no lead has been scanned yet) or stale (a lead added since).
|
||||
*
|
||||
* @param cfg the startup config snapshot — read ONCE here, only to decide whether
|
||||
* to construct the object at all, exactly like {@code
|
||||
* cfg.leadHeartbeat()}
|
||||
* @param leadAgents the {@link AgentControl} instance that reaches the LEAD's pane (not
|
||||
* {@code memberAgents}), normally {@code router.leadAgents()}
|
||||
* @param config the live {@link ConfigRef}, captured only inside the returned
|
||||
* supplier and the workspace lookup — never dereferenced here
|
||||
* @param liveLeadTerminals terminal id → lead NAME for every CURRENTLY recognised lead, normally
|
||||
* the same {@code leads} supplier {@code main} already builds for
|
||||
* {@code HerdrRouter}/{@link #leadSeatLookup} — never a value snapshot
|
||||
* @return a constructed {@link LeadRollover}, or {@code null} when {@code leadRollover:} is
|
||||
* absent from the startup config
|
||||
*/
|
||||
static LeadRollover leadRollover(FleetConfig cfg, AgentControl leadAgents, ConfigRef config,
|
||||
Supplier<Map<String, String>> liveLeadTerminals) {
|
||||
if (cfg.leadRollover() == null) {
|
||||
return null;
|
||||
}
|
||||
Function<String, String> leadWorkspace = terminal -> {
|
||||
String leadName = liveLeadTerminals.get().get(terminal);
|
||||
if (leadName == null) {
|
||||
return null;
|
||||
}
|
||||
FleetConfig.Fleet fleet = config.get().fleet();
|
||||
FleetConfig.Leader leader = fleet == null ? null : fleet.leaders().get(leadName);
|
||||
return leader == null ? null : leader.cwd();
|
||||
};
|
||||
return new LeadRollover(leadAgents, () -> config.get().leadRollover(), leadWorkspace);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #176: per-profile factory for {@link FleetMcp.LeadSeatSource} — how many seats a
|
||||
* profile's own live LEAD session(s) hold on the same Claude subscription.
|
||||
@@ -1323,6 +1627,74 @@ public final class Fleetd {
|
||||
+ "fleetd.yaml — see fleetd.example.yaml — and restart.");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #395: {@code exhaustedPattern} (see {@link FleetConfig.Profile#exhaustedPattern}) is
|
||||
* deliberately opt-in — {@code null}/blank means a backend refusal on that profile is never
|
||||
* classified as {@code BACKEND_EXHAUSTED}, so its credential is never quarantined. That is a
|
||||
* legitimate choice (guessing the vendor's wording would be worse), but an operator who never
|
||||
* opted a profile in should not discover the gap only when a usage limit silently goes
|
||||
* undetected. Warn once at startup, naming every unarmed profile, exactly like {@link
|
||||
* #reportMemberCredentialsGap} — never refuse to start over it.
|
||||
*
|
||||
* <p>A {@code subscription: true} profile that is unarmed gets a SECOND, louder WARN of its
|
||||
* own: it bills the operator's metered Claude plan, the case where an undetected usage limit
|
||||
* costs the most.
|
||||
*
|
||||
* <p>Package-private so a test can capture the real log via a {@link
|
||||
* ch.qos.logback.core.read.ListAppender}, the same pattern {@link
|
||||
* #reportMemberCredentialsGap}'s own test uses.
|
||||
*/
|
||||
static void reportExhaustedPatternGap(FleetConfig cfg) {
|
||||
List<String> unarmedSubscription = new ArrayList<>();
|
||||
List<String> unarmedOther = new ArrayList<>();
|
||||
cfg.profiles().forEach((name, profile) -> {
|
||||
if (!profile.hasExhaustedPattern()) {
|
||||
(profile.isSubscription() ? unarmedSubscription : unarmedOther).add(name);
|
||||
}
|
||||
});
|
||||
if (unarmedSubscription.isEmpty() && unarmedOther.isEmpty()) {
|
||||
log.info("exhaustedPattern: every configured profile has usage-limit detection armed");
|
||||
return;
|
||||
}
|
||||
List<String> allUnarmed = new ArrayList<>(unarmedSubscription);
|
||||
allUnarmed.addAll(unarmedOther);
|
||||
allUnarmed = allUnarmed.stream().sorted().toList();
|
||||
log.warn("exhaustedPattern: profile(s) {} have no exhaustedPattern configured — a "
|
||||
+ "usage-limit refusal on any of them is never detected and never "
|
||||
+ "quarantines its credential. Set exhaustedPattern (see "
|
||||
+ "fleetd.example.yaml) to arm detection for a profile.",
|
||||
allUnarmed);
|
||||
if (!unarmedSubscription.isEmpty()) {
|
||||
List<String> sortedSubscription = unarmedSubscription.stream().sorted().toList();
|
||||
log.warn("exhaustedPattern: subscription profile(s) {} run on the operator's metered "
|
||||
+ "Claude plan and have NO usage-limit detection armed — this is the "
|
||||
+ "case where a missed usage limit costs the most.",
|
||||
sortedSubscription);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #474: the one place both the startup call (right after {@code cfg.validateAll()} in
|
||||
* {@link #main}) and the reload call (wired into {@code config}'s {@code extraValidation} above,
|
||||
* via a method reference to this method) go through, so the two can never drift into checking
|
||||
* different things. Extracted only to give {@link ConfigRef}'s {@code Consumer<FleetConfig>}
|
||||
* hook a {@code FleetConfig -> void} shape to bind to — {@link CharterToolSurface} itself still
|
||||
* takes the raw charter map and knows nothing about {@code ConfigRef} or {@code Fleetd}.
|
||||
*
|
||||
* <p>Package-private so a test can call it directly the same way the other startup-report
|
||||
* helpers above are tested, without needing to drive {@link #main} for a unit-level check;
|
||||
* {@code FleetdStartupValidationTest} proves the startup call site, and {@code
|
||||
* FleetdConfigRefCharterToolSurfaceWiringTest} — by constructing {@code ConfigRef} with this
|
||||
* exact method reference, the same way {@code main} does above — proves the reload call site.
|
||||
* {@code dev.ltms.fleet.config.ConfigRefTest} pins the same reload behaviour too, through an
|
||||
* equivalent {@code Consumer<FleetConfig>} it builds locally (it cannot see this package-private
|
||||
* method from {@code dev.ltms.fleet.config}).
|
||||
*/
|
||||
static void assertChartersNameOnlyRegisteredTools(FleetConfig cfg) {
|
||||
CharterToolSurface.assertChartersNameOnlyRegisteredTools(
|
||||
cfg.fleet() == null ? Map.of() : cfg.fleet().charters());
|
||||
}
|
||||
|
||||
/**
|
||||
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
|
||||
*
|
||||
|
||||
@@ -30,8 +30,27 @@ public final class Authz {
|
||||
DRAIN,
|
||||
/** Read-only observation: status, roster, profiles, task polling. */
|
||||
READ,
|
||||
/**
|
||||
* Read (never ack) this daemon's own held lead-to-lead coordination mail (fleetd #421).
|
||||
*
|
||||
* <p>Deliberately <strong>not</strong> folded into {@link #READ}. {@code READ}'s grant
|
||||
* rests on "the roster carries no secrets" (see its case below) — a lead-to-lead body is
|
||||
* not the roster; it is where leads discuss host shapes, credentials and unmerged work.
|
||||
* Mapping this to {@code READ} would let any worker read every peer lead's mail in full
|
||||
* and would silently falsify that comment for every other {@code READ} caller.
|
||||
*/
|
||||
COORD_READ,
|
||||
/** Scrape the metrics endpoint. */
|
||||
METRICS
|
||||
METRICS,
|
||||
/**
|
||||
* Drive the lead-rollover executor ({@code fleet_handover}: open/confirm/cancel a
|
||||
* self-replace, fleetd #480 Unit C). Primary-only, same as {@link #SPAWN}/{@link #STOP}/
|
||||
* {@link #DRAIN} — and, unlike those, the terminal it acts on is never even an argument:
|
||||
* {@code LeadRollover#open}/{@code #confirm} are always called with the CALLER's own
|
||||
* connection-resolved terminal (see {@code dev.ltms.fleet.lead.LeadRollover}'s class
|
||||
* javadoc, fleetd #480 correction 2), so a primary can only ever roll itself.
|
||||
*/
|
||||
HANDOVER
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -50,7 +69,7 @@ public final class Authz {
|
||||
// deliberately does NOT get these (CB-548), so it cannot tear down or stand up workers
|
||||
// even though it coordinates them; and a worker driving any of these would be a worker
|
||||
// escalating into the orchestrator role.
|
||||
case SPAWN, STOP, DRAIN -> caller.isPrimary();
|
||||
case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary();
|
||||
|
||||
// Delivering a turn is open to the primary and the architect: an architect delegates
|
||||
// to workers (that is the role's point) but still has no lifecycle rights. A worker is
|
||||
@@ -68,6 +87,11 @@ public final class Authz {
|
||||
// Observation is open to every authenticated role: a worker legitimately polls its own
|
||||
// status, and the roster carries no secrets.
|
||||
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
|
||||
|
||||
// fleetd #421: reading held lead-to-lead mail is the primary's alone. An architect
|
||||
// holds READ today (CB-548), so "not primary" must mean not-architect here too — this
|
||||
// is coordination between leads, not observation of the roster.
|
||||
case COORD_READ -> caller.isPrimary();
|
||||
};
|
||||
}
|
||||
|
||||
|
||||
@@ -11,6 +11,7 @@ import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
* The architect-slot registry (CB-548): every gateway-local architect name and the strong-model
|
||||
@@ -19,16 +20,43 @@ import java.util.Objects;
|
||||
*
|
||||
* <p>Two halves, split by who owns each:
|
||||
* <ul>
|
||||
* <li><b>slots</b> — configured once, keyed by the gateway-local unique name; each carries the
|
||||
* {@code profile} reference the spawn lifecycle reads when it stands the slot up. A read-only
|
||||
* snapshot taken at construction.</li>
|
||||
* <li><b>terminal bindings</b> — owned by this registry and initially <em>empty</em>. Config
|
||||
* declares no architect terminal, so at startup every slot is idle and nothing resolves to an
|
||||
* architect; a session only becomes one when the spawn lifecycle {@linkplain #bind(String,
|
||||
* String) binds} its terminal to a slot. {@link CallerResolver} reads this through
|
||||
* {@link #snapshot()} to turn a pane into an {@link Role#ARCHITECT}.</li>
|
||||
* <li><b>slots</b> — read from {@code fleet.architects}/{@code developers}/{@code reviewers}
|
||||
* (see {@link #slots()}), each carrying the {@code profile} reference the spawn lifecycle
|
||||
* reads when it stands the slot up. <strong>Live, since fleetd #424</strong>: {@link #live}
|
||||
* re-reads {@code fleet:} on every call, through a supplier the same shape as
|
||||
* {@code CompositePeerLauncher}'s (see {@code ConfigRef}'s class doc) — so a config reload
|
||||
* that removes or adds an architect slot governs the <em>next</em> spawn with no restart.
|
||||
* Only {@link #MemberRegistry(FleetConfig.Fleet)} freezes the pool at construction, and that
|
||||
* constructor exists for tests and for the (rare) case of wiring a fixed, code-built config.</li>
|
||||
* <li><b>terminal bindings</b> — owned by this registry, initially <em>empty</em>, and
|
||||
* <strong>never</strong> touched by a reload. Config declares no architect terminal, so at
|
||||
* startup every slot is idle and nothing resolves to an architect; a session only becomes one
|
||||
* when the spawn lifecycle {@linkplain #bind(String, String) binds} its terminal to a slot.
|
||||
* {@link CallerResolver} reads this through {@link #snapshot()} to turn a pane into an
|
||||
* {@link Role#ARCHITECT}.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p><strong>The binding rule (fleetd #424): config governs what a bound slot still grants, as
|
||||
* well as what may be bound next.</strong> Removing a slot from config revokes it — that is the
|
||||
* ticket's entire point ("Revoking an architect slot does not revoke it"). Revoking it means an
|
||||
* architect already bound to that slot loses the ARCHITECT privilege on its very next request:
|
||||
* {@link #roleForSlot} and {@link #nameForSlot} read {@link #slots()} directly, with no cache, so
|
||||
* the moment a slot drops out of config, {@link CallerResolver#resolve} (which calls both on every
|
||||
* request from a bound pane, {@code CallerResolver.java:220}) can no longer confirm the pane's slot
|
||||
* is an architect slot, and the pane falls through to {@code Principal.worker(...)}. What does
|
||||
* <em>not</em> change is the {@code terminalToSlot} <em>occupancy</em> — the binding created by
|
||||
* {@link #bind} is untouched by a reload, on purpose: unbinding it here would double-book the slot
|
||||
* key (a second terminal could then bind to the "freed" key while the first is still the terminal
|
||||
* the operator actually meant to demote) and would silently break {@link #unbind}'s compare-safe
|
||||
* contract, which needs the original {@code terminal → slot} pair intact to remove it cleanly. So
|
||||
* the demoted session keeps occupying its slot — {@link #slotForTerminal} and {@link #snapshot()}
|
||||
* still name it — it just no longer resolves as an architect through that occupancy, and a fresh
|
||||
* spawn still cannot bind to the same key while it is occupied ({@link #reserve}/
|
||||
* {@link #requireSlotFor} refuse it anyway, since it is gone from {@link #slots()}). The demoted
|
||||
* session's own turn is unaffected: {@code fleet_reply}'s authorization
|
||||
* ({@code Authz.Action.REPLY}) is {@code caller.ownsSession(targetSession)} — identity by terminal,
|
||||
* not by role — so a demoted architect can still end its own turn normally.
|
||||
*
|
||||
* <p>Spawning/lifecycle is deliberately a separate unit: this class only owns the bindings and
|
||||
* exposes the map the resolver resolves against plus the profile lookup lifecycle will call.
|
||||
* Nothing here creates or manages an architect session.
|
||||
@@ -55,14 +83,42 @@ public final class MemberRegistry implements MemberLifecycle {
|
||||
}
|
||||
}
|
||||
|
||||
private final Map<String, Entry> slots;
|
||||
private final Supplier<FleetConfig.Fleet> fleet;
|
||||
/** Live {@code terminal_id → qualified slot key}; guarded by {@code terminalToSlot}. */
|
||||
private final Map<String, String> terminalToSlot = new HashMap<>();
|
||||
/** Slot keys held between reservation and the terminal binding. Guarded by terminalToSlot. */
|
||||
private final java.util.Set<String> reservedSlots = new java.util.HashSet<>();
|
||||
|
||||
/** Flatten every role pool in {@code fleet} into one registry. Leaders are not members. */
|
||||
/**
|
||||
* Freeze the pool at construction — for tests, and for the rare case of wiring a fixed,
|
||||
* code-built config. Production wiring should prefer {@link #live}, which re-reads
|
||||
* {@code fleet:} on every call.
|
||||
*/
|
||||
public MemberRegistry(FleetConfig.Fleet fleet) {
|
||||
this(() -> fleet);
|
||||
}
|
||||
|
||||
private MemberRegistry(Supplier<FleetConfig.Fleet> fleet) {
|
||||
this.fleet = fleet;
|
||||
}
|
||||
|
||||
/**
|
||||
* Live variant (fleetd #424): {@code fleet} is read fresh on every {@link #slots()} call — pass
|
||||
* {@code () -> config.get().fleet()}, the same supplier shape {@code CompositePeerLauncher}
|
||||
* already uses for placement — so a reload that adds or removes an architect slot governs the
|
||||
* next spawn's {@link #reserve}/{@link #requireSlotFor} check with no restart. A separate,
|
||||
* private constructor rather than a same-arity public overload of
|
||||
* {@link #MemberRegistry(FleetConfig.Fleet)}: a {@code FleetConfig.Fleet} and a
|
||||
* {@code Supplier<FleetConfig.Fleet>} overload are ambiguous for a literal {@code null} — the
|
||||
* same reason {@code CallerResolver.withLeads} is a static factory rather than a fourth
|
||||
* constructor overload.
|
||||
*/
|
||||
public static MemberRegistry live(Supplier<FleetConfig.Fleet> fleet) {
|
||||
return new MemberRegistry(Objects.requireNonNull(fleet, "fleet"));
|
||||
}
|
||||
|
||||
/** Flatten every role pool in {@code fleet} into one map. Leaders are not members. */
|
||||
private static Map<String, Entry> flatten(FleetConfig.Fleet fleet) {
|
||||
Map<String, Entry> flat = new LinkedHashMap<>();
|
||||
if (fleet != null) {
|
||||
for (MemberRole role : MemberRole.values()) {
|
||||
@@ -74,18 +130,22 @@ public final class MemberRegistry implements MemberLifecycle {
|
||||
});
|
||||
}
|
||||
}
|
||||
this.slots = Collections.unmodifiableMap(flat);
|
||||
return Collections.unmodifiableMap(flat);
|
||||
}
|
||||
|
||||
/** The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot. */
|
||||
/**
|
||||
* The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot of
|
||||
* {@code fleet:} <em>as of this call</em> — see the class doc for which constructor makes that
|
||||
* live versus frozen.
|
||||
*/
|
||||
public Map<String, Entry> slots() {
|
||||
return slots;
|
||||
return flatten(fleet.get());
|
||||
}
|
||||
|
||||
/** The slots belonging to {@code role}, in definition order. */
|
||||
/** The slots belonging to {@code role}, in definition order, as of this call. */
|
||||
public Map<String, Entry> slotsFor(MemberRole role) {
|
||||
Map<String, Entry> out = new LinkedHashMap<>();
|
||||
slots.forEach((key, e) -> {
|
||||
slots().forEach((key, e) -> {
|
||||
if (e.role() == role) {
|
||||
out.put(key, e);
|
||||
}
|
||||
@@ -117,31 +177,49 @@ public final class MemberRegistry implements MemberLifecycle {
|
||||
}
|
||||
|
||||
/**
|
||||
* The strong-model profile a slot runs under — what the spawn lifecycle reads.
|
||||
* The strong-model profile a slot runs under, as of this call.
|
||||
*
|
||||
* <p>Nothing in {@code src/main} calls this (fleetd #431 — grepped both the {@code
|
||||
* .profileForSlot(} and the {@code ::profileForSlot} form). This javadoc used to say "what the
|
||||
* spawn lifecycle reads", and that seam does not exist: the spawn lifecycle takes its profile
|
||||
* from the {@link MemberLifecycle.SlotReservation} that {@code reserve} returns, never from
|
||||
* here. Kept and pinned rather than deleted because it is the natural accessor for that seam
|
||||
* if one is added; live for the same reason as {@link #roleForSlot}, so a reload cannot leave
|
||||
* it answering for the old config.
|
||||
*
|
||||
* @return the slot's configured {@code profile}, or {@code null} if the slot is unknown or
|
||||
* declares none
|
||||
*/
|
||||
public String profileForSlot(String slotName) {
|
||||
Entry e = slots.get(slotName);
|
||||
Entry e = slots().get(slotName);
|
||||
return (e == null || e.profile() == null) ? null : e.profile();
|
||||
}
|
||||
|
||||
/** The role a qualified slot key belongs to, or {@code null} when the key is unknown. */
|
||||
/**
|
||||
* The role a qualified slot key belongs to, or {@code null} when the key is not currently
|
||||
* configured. Deliberately live, with no cache (fleetd #424, see the class doc's binding rule):
|
||||
* removing a slot from config must make {@link CallerResolver#resolve} stop granting the
|
||||
* ARCHITECT role for it on the very next request from a terminal that was bound to it, which is
|
||||
* the ticket's whole point — revoking a slot must actually revoke it, not just refuse the next
|
||||
* spawn.
|
||||
*/
|
||||
public MemberRole roleForSlot(String slotName) {
|
||||
Entry e = slots.get(slotName);
|
||||
Entry e = slots().get(slotName);
|
||||
return e == null ? null : e.role();
|
||||
}
|
||||
|
||||
/** The unqualified configured name for a slot, or {@code null} if it is unknown. */
|
||||
/**
|
||||
* The unqualified configured name for a slot, or {@code null} if it is not currently configured.
|
||||
* Live for the same reason as {@link #roleForSlot} — see the class doc's binding rule.
|
||||
*/
|
||||
public String nameForSlot(String slotName) {
|
||||
Entry e = slots.get(slotName);
|
||||
Entry e = slots().get(slotName);
|
||||
return e == null ? null : e.name();
|
||||
}
|
||||
|
||||
/** True when {@code slotName} is a configured architect slot. */
|
||||
public boolean isSlot(String slotName) {
|
||||
return slots.containsKey(slotName);
|
||||
return slots().containsKey(slotName);
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -11,6 +11,7 @@ import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
import java.util.function.Consumer;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
@@ -32,9 +33,51 @@ import java.util.function.Supplier;
|
||||
* not the fact that they are config. Most of {@code fleet:} — every role pool
|
||||
* ({@code architects}/{@code developers}/{@code reviewers}), {@code charters}, and
|
||||
* {@code tabLabel} — is read the same live way, through the same supplier
|
||||
* ({@code () -> config.get().fleet()}). <strong>But {@code fleet:} as a whole is NOT in this
|
||||
* class</strong>: {@code fleet.leaders} inside the same key is frozen, which is exactly what
|
||||
* makes {@code fleet:} split rather than hot — see below.</li>
|
||||
* ({@code () -> config.get().fleet()}). {@code architects} in particular is hot for
|
||||
* <strong>two independent consumers</strong> (fleetd #424): {@code CompositePeerLauncher}
|
||||
* reads it live for placement (which profile an unqualified architect spawn may land on), and
|
||||
* {@code MemberRegistry} separately reads it live, through its own instance of the same
|
||||
* supplier shape, for identity — both which slot a spawn may bind to <em>and</em> what a slot
|
||||
* already bound still grants. Removing an architect slot from config therefore revokes the
|
||||
* {@link dev.ltms.fleet.auth.Role#ARCHITECT} role on the bound pane's very next request; only
|
||||
* the slot <em>occupancy</em> survives, so the demoted session still holds its slot key until
|
||||
* it unbinds. See {@code MemberRegistry}'s class doc for that binding rule.
|
||||
* <strong>But {@code fleet:} as a whole is NOT in this class</strong>: {@code fleet.leaders}
|
||||
* inside the same key is frozen, which is exactly what makes {@code fleet:} split rather than
|
||||
* hot — see below. {@code models:} (fleetd #422) joined this class whole: {@link
|
||||
* FleetConfig#validateModels()} re-runs fully against the fresh config on every {@link
|
||||
* #reload()} (via {@link FleetConfig#validateAll()}), refusing a bad edit outright rather than
|
||||
* caching a stale copy anywhere, and the on/off half added by fleetd #422 is read live both by
|
||||
* {@code CompositePeerLauncher}'s spawn gate ({@code enforceModelEnabled} and its candidate
|
||||
* filter) and by {@code fleet_profiles}/{@code GET /profiles} (via
|
||||
* {@code PeerLauncher.disabledModels()}). Nothing about {@code models:} is baked into an
|
||||
* object built at startup, so — unlike the deferred keys below — there is no frozen half left
|
||||
* to report; it moved here from deferred rather than joining split. An existing profile's
|
||||
* {@code exhaustedPattern} (fleetd #446) joined this class the same way: it used to be
|
||||
* compiled once into {@code Fleetd.main}'s startup pattern map (see the Deferred bullet's old
|
||||
* wording, and {@code LiveExhaustedPatterns}'s class doc for the history), and is now read
|
||||
* live, cached by profile name, by both {@code CompletionResolver}'s classification (via
|
||||
* {@code LiveExhaustedPatterns.patternFor}) and {@code fleet_profiles}'s {@code
|
||||
* exhaustionDetectionArmed} (via {@code LiveExhaustedPatterns.armed}) — the one live object
|
||||
* both read, so a reload that arms or disarms a profile's usage-limit detection takes effect
|
||||
* on the next check with no restart. {@code errorPattern}, {@code exhaustedPattern}'s sibling
|
||||
* key for backend-error (not usage-limit) classification, was deliberately left OUT of this
|
||||
* fleetd #446 change and stays deferred below — the ticket scoped it out explicitly.
|
||||
* {@code leadRollover:} (fleetd #480) joined this class whole, the same shape as
|
||||
* {@code models:} above: {@code dev.ltms.fleet.lead.LeadRollover} holds a
|
||||
* {@code Supplier<FleetConfig.LeadRollover>} (the same {@code () -> config.get().x()} shape)
|
||||
* and reads {@code handoverPath}/{@code requireOperatorConfirm}/{@code maxDocAgeSeconds}/
|
||||
* {@code turnSettleSeconds}/{@code clearSettleSeconds}/{@code bootstrapText} fresh on every
|
||||
* {@code open()}/{@code confirm()} call (and on the deferred post-{@code confirm()}
|
||||
* continuation fleetd #480's correction added — see {@code LeadRollover}'s class doc) rather
|
||||
* than capturing them into fields at construction — unlike its closest
|
||||
* structural cousin {@code leadHeartbeat:}, whose {@code LeadHeartbeatLoop} bakes
|
||||
* {@code idleAfterNanos}/{@code backoffMs}/{@code quietNudgeCap} into final fields. The one
|
||||
* restart-only edge is structural, not a stale value: {@code Fleetd.java} decides whether to
|
||||
* construct the {@code LeadRollover} object at all off the startup snapshot (the same
|
||||
* presence gate {@code leadHeartbeat:} uses), so a block ADDED where it was absent at boot
|
||||
* needs a restart before anything exists to call — the same fact already true of adding a
|
||||
* brand-new {@code profiles:} entry.</li>
|
||||
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
|
||||
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
|
||||
* {@code idleSleepGuard:} ({@code Fleetd.java} reads it once, at startup, to decide whether
|
||||
@@ -64,15 +107,17 @@ import java.util.function.Supplier;
|
||||
* (if any) simply keeps its old settings), adding or removing a profile (a new backend needs its own launcher,
|
||||
* which is constructed once), <em>and an existing profile's launch settings</em> —
|
||||
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl},
|
||||
* {@code exhaustedPattern} (CB-578 stage A — compiled once into {@code Fleetd.main}'s
|
||||
* pattern map at startup), {@code errorPattern} (fleetd #201 Unit 5 — compiled once into
|
||||
* {@code Fleetd.main}'s backend-error pattern map at startup, the same way),
|
||||
* {@code errorPattern} (fleetd #201 Unit 5 — compiled once into
|
||||
* {@code Fleetd.main}'s backend-error pattern map at startup; deliberately NOT made hot
|
||||
* alongside {@code exhaustedPattern} by fleetd #446 — that ticket scoped {@code errorPattern}
|
||||
* and cooling-off out explicitly),
|
||||
* {@code ideProjectDir} / {@code ideOpenCommand} / {@code autoCompactWindow} (fleetd #323
|
||||
* instance 1 — all three are read at spawn off the same frozen profile map and were missing
|
||||
* from {@link #sameLaunchSettings}), and the rest of {@link #sameLaunchSettings}.
|
||||
* {@code credentialId} (CB-578 stage B) is NOT on
|
||||
* this list — it is read live off the config supplier at every quarantine check and
|
||||
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so it is hot instead.
|
||||
* {@code credentialId} (CB-578 stage B) and {@code exhaustedPattern} (fleetd #446) are NOT on
|
||||
* this list — both are read live off the config supplier at every quarantine check and
|
||||
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so both are hot instead
|
||||
* (see the Hot bullet above for {@code exhaustedPattern}'s history).
|
||||
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} at construction and resolves
|
||||
* each spawn out of that copy, so those never reach a launch until the daemon restarts. A
|
||||
* reload logs these rather than pretending they applied.</li>
|
||||
@@ -135,15 +180,21 @@ import java.util.function.Supplier;
|
||||
* </ul>
|
||||
*
|
||||
* <p><strong>The denominator, measured on 2026-09-04 (fleetd #330; recounted for fleetd #333);
|
||||
* recounted again for fleetd #362, and again after {@code idleSleepGuard:} was added.</strong>
|
||||
* {@code FleetConfig} has 24 top-level record components: 5 cold, 13 deferred, 3 split, 3
|
||||
* hot-excluded. Three of them are named nowhere in this file, and the reason is the same for all
|
||||
* three: {@code placement}, {@code memberCredentials} and {@code memberLoginShell} are
|
||||
* <strong>hot</strong> and correctly absent — all three are read live off {@code config.get()}
|
||||
* (placement through the {@code CompositePeerLauncher} supplier the Hot bullet names;
|
||||
* {@code memberCredentials}/{@code memberLoginShell} at spawn time, {@code Fleetd.java:198, 205, 729}
|
||||
* and {@code HerdrPeerLauncher#configuredMemberLoginShell}), so a reload takes effect on the next
|
||||
* spawn with no entry needed here.
|
||||
* recounted again for fleetd #362, again after {@code idleSleepGuard:} was added, again after
|
||||
* {@code models:} was added as deferred, again for fleetd #422, which moved {@code models:}
|
||||
* from deferred to hot-excluded once its on/off half was read live everywhere, and again after
|
||||
* {@code leadRollover:} was added (fleetd #480).</strong>
|
||||
* {@code FleetConfig} has 26 top-level record components: 5 cold, 13 deferred, 3 split, 5
|
||||
* hot-excluded. Five of them are named nowhere in this file, and the reason is the same for all
|
||||
* five: {@code placement}, {@code memberCredentials}, {@code memberLoginShell}, {@code models} and
|
||||
* {@code leadRollover} are <strong>hot</strong> and correctly absent — all five are read live off
|
||||
* {@code config.get()} (placement through the {@code CompositePeerLauncher} supplier the Hot bullet
|
||||
* names; {@code memberCredentials}/{@code memberLoginShell} at spawn time, {@code Fleetd.java:198,
|
||||
* 205, 729} and {@code HerdrPeerLauncher#configuredMemberLoginShell}; {@code models} the same way,
|
||||
* through the Hot bullet's {@code models:} paragraph; {@code leadRollover} through the Hot bullet's
|
||||
* {@code leadRollover:} paragraph), so a reload takes effect on the next spawn (or, for
|
||||
* {@code models}, the next reported status; for {@code leadRollover}, the next {@code open()}/
|
||||
* {@code confirm()} call) with no entry needed here.
|
||||
* {@code health} and {@code coordinator} used to be a third kind — <strong>undecided</strong>, not
|
||||
* hot — until fleetd #330 added the <strong>split</strong> class above and gave them a home. A
|
||||
* reload touching either used to report a bare "config reloaded", which under-claimed; now it names
|
||||
@@ -179,6 +230,24 @@ import java.util.function.Supplier;
|
||||
* <p>A reload that fails to parse or fails validation is also refused, and the previous config keeps
|
||||
* running. A config file being edited is normally read once mid-save; degrading a working daemon
|
||||
* because it caught a half-written file would be a bad trade.
|
||||
*
|
||||
* <p><strong>fleetd #474</strong> — {@link FleetConfig#validateAll()} is not the only gate startup
|
||||
* runs before a config takes effect: {@code Fleetd.main} also calls {@code
|
||||
* dev.ltms.fleet.mcp.CharterToolSurface#assertChartersNameOnlyRegisteredTools}, right after {@code
|
||||
* cfg.validateAll()}, to refuse a charter that names an MCP tool the server does not register. That
|
||||
* check cannot live inside {@link FleetConfig} — {@code CharterToolSurface} lives in the {@code mcp}
|
||||
* package because the canonical tool set ({@code FleetTool}) does, and config is loaded before the
|
||||
* MCP server exists, so {@code FleetConfig} must not gain a dependency on {@code mcp}. {@link
|
||||
* #reload} cannot import {@code mcp} either, for the same reason applied one layer up: {@code
|
||||
* dev.ltms.fleet.config} is loaded before {@code dev.ltms.fleet.mcp} exists, same as {@code
|
||||
* FleetConfig}. So this class accepts the check as a {@code Consumer<FleetConfig>} —
|
||||
* {@link #extraValidation} — supplied by whichever caller already sits at the seam that holds both
|
||||
* a loaded {@code FleetConfig} and the {@code mcp} package: {@code Fleetd.main}. It is invoked
|
||||
* inside the same try/catch as {@code fresh.validateAll()}, so a charter that would have refused to
|
||||
* boot refuses a reload too, and keeps the running config exactly like any other {@code
|
||||
* validateAll()} failure. A ref built through the two-argument constructor (every test fixture that
|
||||
* does not care about this check, and {@link #fixed}) gets a no-op consumer, so nothing outside
|
||||
* {@code Fleetd.main} needs to know this hook exists.
|
||||
*/
|
||||
public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
|
||||
@@ -223,10 +292,27 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
|
||||
private final Path path;
|
||||
private final AtomicReference<FleetConfig> current;
|
||||
private final Consumer<FleetConfig> extraValidation;
|
||||
|
||||
/** Equivalent to the three-argument constructor with a no-op {@code extraValidation}. */
|
||||
public ConfigRef(Path path, FleetConfig initial) {
|
||||
this(path, initial, cfg -> { });
|
||||
}
|
||||
|
||||
/**
|
||||
* @param extraValidation run on every {@link #reload} candidate, inside the same try/catch as
|
||||
* {@code fresh.validateAll()} — see the class doc's fleetd #474 note.
|
||||
* {@code Fleetd.main} passes {@code
|
||||
* Fleetd::assertChartersNameOnlyRegisteredTools} (a package-private
|
||||
* {@code FleetConfig -> void} adapter over {@code
|
||||
* CharterToolSurface#assertChartersNameOnlyRegisteredTools}), so a reload
|
||||
* runs the same gate startup does without this class depending on the
|
||||
* {@code mcp} package.
|
||||
*/
|
||||
public ConfigRef(Path path, FleetConfig initial, Consumer<FleetConfig> extraValidation) {
|
||||
this.path = path;
|
||||
this.current = new AtomicReference<>(Objects.requireNonNull(initial, "initial config"));
|
||||
this.extraValidation = Objects.requireNonNull(extraValidation, "extraValidation");
|
||||
}
|
||||
|
||||
/** A fixed reference that never reloads — for tests and for wiring built from a config in code. */
|
||||
@@ -322,12 +408,18 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
fresh = FleetConfig.load(path);
|
||||
// The same gate startup runs. A config that would have refused to boot must not be able
|
||||
// to slip in through a reload — that is how a daemon ends up in a state it could never
|
||||
// have started in, which is the hardest kind to debug.
|
||||
fresh.validateAuthExposure();
|
||||
fresh.validateLeadTabPrefixes();
|
||||
fresh.validateSubscriptionProfiles();
|
||||
fresh.validateCharters();
|
||||
fresh.validateMembers();
|
||||
// have started in, which is the hardest kind to debug. fleetd ticket "central allow-list
|
||||
// of usable models" follow-up: this used to be six individual validateXxx() calls, and
|
||||
// mutation testing found two of the six unpinned here even though startup pinned nothing
|
||||
// at all — see FleetConfig#validateAll's javadoc for why the fix is one reflective call,
|
||||
// not a longer hand-maintained list.
|
||||
fresh.validateAll();
|
||||
// fleetd #474: validateAll() does not cover everything startup refuses on — the charter
|
||||
// tool-surface check (Fleetd.main, right after cfg.validateAll()) lives outside
|
||||
// FleetConfig on purpose (see this class's doc) and is supplied here as extraValidation.
|
||||
// Same try/catch as validateAll() above, on purpose: either failure must refuse the whole
|
||||
// reload and keep the running config the same way.
|
||||
extraValidation.accept(fresh);
|
||||
} catch (RuntimeException e) {
|
||||
String msg = e.getMessage() == null ? e.toString() : e.getMessage();
|
||||
log.warn("config reload from {} refused, keeping the running config: {}", path, msg);
|
||||
@@ -513,26 +605,35 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
+ "opened once and needs a restart; the broker URI env-var name kept out of a "
|
||||
+ "member's environment is read live on every spawn and already applied");
|
||||
}
|
||||
// fleetd #333: unlike health/coordinator above, most of `fleet:` (architects, developers,
|
||||
// reviewers, charters, tabLabel) is genuinely hot — ConfigRefTest.aHotChangeIsAppliedAndRead-
|
||||
// fleetd #333: unlike health/coordinator above, most of `fleet:` (developers, reviewers,
|
||||
// charters, tabLabel) is genuinely hot — ConfigRefTest.aHotChangeIsAppliedAndRead-
|
||||
// ThroughGet and aCharterChangeIsHotAndReachesTheLiveConfig prove it reaches the live config
|
||||
// with no restart note. Only fleet.leaders is frozen (Fleetd.java:281 reads
|
||||
// cfg.fleet().leaders() off the startup snapshot to build both the LeadTabScanner's
|
||||
// tab-label-to-name map, wired into CallerResolver.withLeadsAndMembers at Fleetd.java:620/624,
|
||||
// and — when herdr answered — LeadLauncher(...).ensureLeads() at Fleetd.java:315, which
|
||||
// auto-launches each lead up to its `instances` count; neither is rebuilt on reload). So this
|
||||
// compares fleet.leaders alone, not the whole Fleet record: comparing the whole record would
|
||||
// report "split" for a tabLabel-only or charters-only change that is actually fully hot,
|
||||
// which is the over-claim mirror of the under-claim bug this class exists to prevent.
|
||||
// with no restart note. `architects` is hot too, and — since fleetd #424 — hot for BOTH of
|
||||
// its consumers, not just the one this comment used to name: CompositePeerLauncher reads it
|
||||
// live for PLACEMENT through the () -> config.get().fleet() supplier named in the class doc's
|
||||
// Hot bullet, and MemberRegistry separately reads it live for IDENTITY (which slot a spawn
|
||||
// may bind to, AND what a slot already bound still grants) through its own instance of that
|
||||
// same supplier shape — see MemberRegistry.live and its class doc for the binding rule:
|
||||
// removing a slot revokes ARCHITECT on the bound pane's very next request, and only the slot
|
||||
// OCCUPANCY survives, so the demoted session keeps its slot key until it unbinds. Only
|
||||
// fleet.leaders is frozen (Fleetd.java:281 reads cfg.fleet().leaders() off the startup
|
||||
// snapshot to build both the LeadTabScanner's tab-label-to-name map, wired into
|
||||
// CallerResolver.withLeadsAndMembers at Fleetd.java:620/624, and — when herdr answered —
|
||||
// LeadLauncher(...).ensureLeads() at Fleetd.java:315, which auto-launches each lead up to its
|
||||
// `instances` count; neither is rebuilt on reload). So this compares fleet.leaders alone, not
|
||||
// the whole Fleet record: comparing the whole record would report "split" for a tabLabel-only
|
||||
// or architects-only change that is actually fully hot, which is the over-claim mirror of the
|
||||
// under-claim bug this class exists to prevent.
|
||||
if (!Objects.equals(leadersOf(old), leadersOf(fresh))) {
|
||||
changed.add("fleet: fleet.leaders (each lead's tab, workspace, cwd, profile and "
|
||||
+ "instances count) is read once at startup to build the LeadTabScanner's "
|
||||
+ "identity map and to auto-launch leads, and neither is rebuilt on reload, so a "
|
||||
+ "lead added, removed, or given a new tab: label needs a restart — until then it "
|
||||
+ "stays unrecognised, and a caller from its new tab resolves as a worker, not a "
|
||||
+ "lead; the rest of fleet: (architects, developers, reviewers, charters, "
|
||||
+ "tabLabel) is read live through the supplier on CompositePeerLauncher and "
|
||||
+ "already applied");
|
||||
+ "lead; the rest of fleet: (developers, reviewers, charters, tabLabel) is read "
|
||||
+ "live through the supplier on CompositePeerLauncher, and architects is read "
|
||||
+ "live through that same supplier for placement AND through a separate supplier "
|
||||
+ "on MemberRegistry for spawn-time identity — both already applied");
|
||||
}
|
||||
// Kept in step with SPLIT_KEYS the same way changedColdKeys is kept in step with COLD_KEYS —
|
||||
// every message here must be traceable to one of the split keys the class doc documents.
|
||||
@@ -554,12 +655,16 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
* {@link #sameLaunchSettings} because they are read <em>live</em>, not baked in at spawn — see
|
||||
* the class doc's <em>Hot</em> bullet. {@code weight} and {@code maxLoad} are read live by the
|
||||
* placement policy on every spawn; {@code credentialId} is read live by
|
||||
* {@code CompositePeerLauncher} and the CB-578 stage B exhaustion sink. Nothing else is
|
||||
* {@code CompositePeerLauncher} and the CB-578 stage B exhaustion sink; {@code exhaustedPattern}
|
||||
* (fleetd #446) is read live, cached by profile name, by {@code LiveExhaustedPatterns} — see
|
||||
* that class's doc and the class doc's <em>Hot</em> bullet for the history (it used to be
|
||||
* compared here, deferred, like its sibling {@code errorPattern} still is). Nothing else is
|
||||
* excluded — see {@code sameLaunchSettingsComparesEveryProfileComponentOrExcludesIt} in
|
||||
* {@code ConfigRefProfileCoverageTest}, which enumerates every {@code Profile} record component
|
||||
* by reflection and fails the build if one is neither compared below nor named here.
|
||||
*/
|
||||
static final Set<String> LAUNCH_SETTINGS_EXCLUDED = Set.of("weight", "maxLoad", "credentialId");
|
||||
static final Set<String> LAUNCH_SETTINGS_EXCLUDED =
|
||||
Set.of("weight", "maxLoad", "credentialId", "exhaustedPattern");
|
||||
|
||||
/**
|
||||
* Whether two versions of a profile would launch a peer identically.
|
||||
@@ -596,13 +701,15 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
&& Objects.equals(a.kind(), b.kind())
|
||||
&& Objects.equals(a.env(), b.env())
|
||||
&& Objects.equals(a.subscription(), b.subscription())
|
||||
// CB-578 stage B: exhaustedPattern is compiled once into Fleetd.main's pattern map
|
||||
// at startup (see ExhaustedPatternLookup wiring) — a reload never re-reads it, so a
|
||||
// changed pattern must be reported as deferred, exactly like model/baseUrl/argv.
|
||||
&& Objects.equals(a.exhaustedPattern(), b.exhaustedPattern())
|
||||
// fleetd #446: exhaustedPattern moved to LAUNCH_SETTINGS_EXCLUDED — it is now read
|
||||
// live, cached by profile name, through LiveExhaustedPatterns (see that class's doc
|
||||
// and ConfigRef's class doc Hot bullet), so it must NOT be compared here any more: a
|
||||
// reload that changes only exhaustedPattern must report "config reloaded", not
|
||||
// "these changes need a restart".
|
||||
// fleetd #201 Unit 5: errorPattern is compiled once into Fleetd.main's backend-error
|
||||
// pattern map at startup (see BackendErrorPatternLookup wiring), the same way
|
||||
// exhaustedPattern is — a reload never re-reads it either.
|
||||
// pattern map at startup (see BackendErrorPatternLookup wiring) — unlike its sibling
|
||||
// exhaustedPattern above (fleetd #446), a reload still never re-reads it; fleetd #446
|
||||
// scoped errorPattern out on purpose (see the class doc's Hot bullet).
|
||||
&& Objects.equals(a.errorPattern(), b.errorPattern())
|
||||
// fleetd #323 instance 1: ideProjectDir and ideOpenCommand are read at spawn off the
|
||||
// same frozen profile map as ideMcpUrl above (ClaudeCodeLauncher.java:267/269,
|
||||
|
||||
@@ -16,11 +16,15 @@ import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.io.UncheckedIOException;
|
||||
import java.lang.reflect.InvocationTargetException;
|
||||
import java.lang.reflect.Method;
|
||||
import java.lang.reflect.Modifier;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.Arrays;
|
||||
import java.util.Collections;
|
||||
import java.util.Comparator;
|
||||
import java.util.HashSet;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
@@ -66,12 +70,14 @@ import java.util.regex.PatternSyntaxException;
|
||||
* {@code fixed} (default), {@code round-robin}, or {@code weighted}
|
||||
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
|
||||
* historical behaviour), CB-501
|
||||
* @param quarantineCooldownSeconds how long a credential stays quarantined after a
|
||||
* @param quarantineCooldownSeconds the BASE cooldown a credential is quarantined for after a
|
||||
* {@code BACKEND_EXHAUSTED} classification (CB-578 stage B); {@code null}/{@code
|
||||
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Baked once into the
|
||||
* {@code BackendQuarantine} built at startup, so it is DEFERRED: changing it
|
||||
* needs a restart, and a quarantine already running keeps whatever cooldown was
|
||||
* live when it started.
|
||||
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Since fleetd #466 this is
|
||||
* only the first occurrence's length — a credential quarantined again shortly
|
||||
* after this cooldown ends backs off further, up to a ceiling; see {@code
|
||||
* BackendQuarantine}'s class doc. Baked once into the {@code BackendQuarantine}
|
||||
* built at startup, so it is DEFERRED: changing it needs a restart, and a
|
||||
* quarantine already running keeps whatever cooldown was live when it started.
|
||||
* @param memberCredentials deny-by-default policy (CB-596) for which of the operator's own host
|
||||
* credentials a spawned member's pane inherits. {@code null} (the block
|
||||
* omitted) blocks nothing — see {@link MemberCredentials}.
|
||||
@@ -124,6 +130,18 @@ import java.util.regex.PatternSyntaxException;
|
||||
* under a member's long turn. {@code null} (the block omitted) behaves the
|
||||
* same as an explicit {@code enabled: true}; set {@code enabled: false} to
|
||||
* turn it off. See {@link dev.ltms.fleet.power.IdleSleepGuard}.
|
||||
* @param models central allow-list of models any {@code profiles:} entry may name. {@code
|
||||
* null} or an empty {@code allow:} ⇒ off: {@link #validateModels()} checks
|
||||
* nothing and every existing config keeps working exactly as it does today.
|
||||
* When non-empty, a profile whose {@code model:} is not one of {@link
|
||||
* Models#ids()} fails config load, naming both the model and the profile —
|
||||
* see {@link #validateModels()}. This block decides what may be CONFIGURED;
|
||||
* fleetd #422 added the separate on/off question — whether a configured model
|
||||
* may be spawned onto RIGHT NOW ({@link Models.ModelEntry#enabled}) — enforced
|
||||
* live at spawn by {@code CompositePeerLauncher}, not here. See {@link Models}.
|
||||
* @param leadRollover opt-in lead rollover (fleetd #480): {@code null} ⇒ off, and no
|
||||
* {@code dev.ltms.fleet.lead.LeadRollover} is constructed at all — an upgraded
|
||||
* daemon never clears a lead's pane on its own initiative. See {@link LeadRollover}.
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record FleetConfig(
|
||||
@@ -150,7 +168,38 @@ public record FleetConfig(
|
||||
String worktreeGroup,
|
||||
String memberLoginShell,
|
||||
String memberSkills,
|
||||
IdleSleepGuard idleSleepGuard) {
|
||||
IdleSleepGuard idleSleepGuard,
|
||||
Models models,
|
||||
LeadRollover leadRollover) {
|
||||
|
||||
/** Back-compat form before the {@code leadRollover:} block was added. */
|
||||
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
|
||||
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
|
||||
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
|
||||
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
|
||||
ConfigReload configReload, Integer quarantineCooldownSeconds,
|
||||
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
|
||||
String memberLoginShell, String memberSkills, IdleSleepGuard idleSleepGuard,
|
||||
Models models) {
|
||||
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
|
||||
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
|
||||
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
|
||||
memberLoginShell, memberSkills, idleSleepGuard, models, null);
|
||||
}
|
||||
|
||||
/** Back-compat form before the {@code models:} block was added. */
|
||||
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
|
||||
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
|
||||
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
|
||||
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
|
||||
ConfigReload configReload, Integer quarantineCooldownSeconds,
|
||||
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
|
||||
String memberLoginShell, String memberSkills, IdleSleepGuard idleSleepGuard) {
|
||||
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
|
||||
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
|
||||
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
|
||||
memberLoginShell, memberSkills, idleSleepGuard, null);
|
||||
}
|
||||
|
||||
/** Back-compat form before the {@code idleSleepGuard:} block was added. */
|
||||
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
|
||||
@@ -396,6 +445,12 @@ public record FleetConfig(
|
||||
* {@code null}/blank ⇒ the classification never fires for this profile and
|
||||
* today's completion-fallback behaviour is unchanged. Every backend words
|
||||
* its refusal differently, so this is config, never a vendor string in code.
|
||||
* Read live off the current config through {@code LiveExhaustedPatterns},
|
||||
* cached by profile name (fleetd #446), so it is HOT: an operator can arm
|
||||
* or disarm this profile's usage-limit detection by editing this key and
|
||||
* reloading, with no restart. Before fleetd #446 this was compiled once at
|
||||
* daemon startup, the same deferred shape its sibling {@code errorPattern}
|
||||
* (below) still has.
|
||||
* @param credentialId the shared account this profile authenticates as (CB-578 stage B). Two
|
||||
* or more profiles setting the <em>same</em> non-blank value are quarantined
|
||||
* together by one {@code exhaustedPattern} classification on any one of
|
||||
@@ -1274,6 +1329,104 @@ public record FleetConfig(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Opt-in lead rollover (fleetd #480): a lead that decides it is ready to be replaced writes a
|
||||
* handover file, then asks fleetd to clear its own pane and bootstrap a fresh session against
|
||||
* that file. Config + a pure decision/verification layer only — see
|
||||
* {@code dev.ltms.fleet.lead.LeadRollover} for the executor this block feeds, and
|
||||
* {@code fleet_*} tool wiring is a later ticket.
|
||||
*
|
||||
* <p>Deliberately opt-in ({@code null} ⇒ off, exactly like {@code leadHeartbeat:}): absent this
|
||||
* block, {@code Fleetd.java} never constructs a {@code LeadRollover} object at all, so an
|
||||
* upgraded daemon cannot silently acquire the ability to clear the lead's own pane. Even once
|
||||
* present, nothing but an explicit {@code confirm()} call can ever roll a pane — there is no
|
||||
* timer, heartbeat or timeout anywhere in this feature that fires one on its own; see that
|
||||
* class's javadoc.
|
||||
*
|
||||
* <p><strong>Hot, not deferred</strong> (see {@code ConfigRef}'s class doc): every field below
|
||||
* is read live, through a {@code Supplier<LeadRollover>} the same {@code () -> config.get().x()}
|
||||
* shape {@code fleet}/{@code placement}/{@code models} already use, so an edit to any of the
|
||||
* six fields takes effect on the very next {@code open()}/{@code confirm()} call once the block
|
||||
* has been present since startup — nothing here is captured into a frozen field the way {@code
|
||||
* leadHeartbeat}'s {@code idleAfterNanos}/{@code backoffMs}/{@code quietNudgeCap} are. The one
|
||||
* restart-only edge left is structural, not a stale value: {@code Fleetd.java} decides whether
|
||||
* to construct the {@code LeadRollover} object at all off the startup snapshot, the same gate
|
||||
* {@code leadHeartbeat:} uses, so a block ADDED where it was absent at boot needs a restart
|
||||
* before anything exists to call — the same fact already true of adding a brand-new
|
||||
* {@code profiles:} entry.
|
||||
*
|
||||
* <p><strong>{@code turnSettleSeconds} (fleetd #480 correction):</strong> {@code confirm()} is
|
||||
* called FROM the calling lead's own turn, so its pane is still {@code WORKING} the instant
|
||||
* {@code confirm()} validates every gate and schedules the roll. {@code
|
||||
* dev.ltms.fleet.lead.LeadRollover}'s deferred continuation waits up to this many seconds for
|
||||
* that SAME pane to report an injectable state again — i.e. for the calling turn to actually
|
||||
* end — before it sends {@code /clear} at all. If that wait times out, no {@code /clear} is
|
||||
* ever sent: a lead that never goes idle is still doing real work, and clearing it would
|
||||
* destroy live context. This is a separate wait from {@code clearSettleSeconds} below, which
|
||||
* bounds the SECOND wait, for the pane to re-settle AFTER {@code /clear} has already gone out.
|
||||
*
|
||||
* @param handoverPath required when this block is present — where the handover file a fresh
|
||||
* lead session reads must live. There is no sane non-null default for an
|
||||
* operator-specific path, so a present block with a {@code null}/blank
|
||||
* {@code handoverPath} is refused at config load; see
|
||||
* {@link #validateLeadRollover()}. May be relative: {@code
|
||||
* dev.ltms.fleet.lead.LeadRollover#open} resolves a relative path against
|
||||
* the CALLING lead's configured {@code fleet.leaders.<name>.cwd}, falling
|
||||
* back to the daemon's own working directory ({@code
|
||||
* System.getProperty("user.dir")}) when that lead has none configured —
|
||||
* never against whatever directory the daemon process happens to have been
|
||||
* started in for its own sake. An absolute path is used unchanged.
|
||||
* @param requireOperatorConfirm default {@code true} — {@code confirm()} refuses unless the
|
||||
* caller also passes {@code operatorConfirmed: true}. Set {@code false} to
|
||||
* let the three handover-file checks alone gate the roll.
|
||||
* @param maxDocAgeSeconds default 3600 — refuse a handover file whose modified time is older
|
||||
* than this many seconds, so a stale leftover from an earlier rollover
|
||||
* attempt can never be mistaken for a fresh one.
|
||||
* @param turnSettleSeconds default 20 — bound on how long the deferred roll waits for the
|
||||
* CALLING lead's own turn to end (its pane to report injectable again)
|
||||
* before sending {@code /clear} at all. See the paragraph above.
|
||||
* @param clearSettleSeconds default 20 — bound on how long to wait for the lead's pane to
|
||||
* report an injectable state again after {@code /clear} before giving up. A
|
||||
* roll that times out here never sends {@code bootstrapText}.
|
||||
* @param bootstrapText default a sentence naming the RESOLVED handover path — sent to the
|
||||
* lead's pane once it settles after {@code /clear}, telling the fresh
|
||||
* session where to read the handover and carry on. Left {@code null} here
|
||||
* when the operator configures none: the default sentence cannot be built
|
||||
* at construction time because it must name the path AFTER {@code
|
||||
* dev.ltms.fleet.lead.LeadRollover#open} has resolved a relative {@code
|
||||
* handoverPath} against the calling lead's workspace, which this record has
|
||||
* no way to know — see {@link #bootstrapTextFor(String)}.
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record LeadRollover(String handoverPath, Boolean requireOperatorConfirm,
|
||||
Integer maxDocAgeSeconds, Integer turnSettleSeconds,
|
||||
Integer clearSettleSeconds, String bootstrapText) {
|
||||
public LeadRollover {
|
||||
requireOperatorConfirm = requireOperatorConfirm == null || requireOperatorConfirm;
|
||||
maxDocAgeSeconds = (maxDocAgeSeconds == null || maxDocAgeSeconds <= 0) ? 3600 : maxDocAgeSeconds;
|
||||
turnSettleSeconds = (turnSettleSeconds == null || turnSettleSeconds <= 0) ? 20 : turnSettleSeconds;
|
||||
clearSettleSeconds = (clearSettleSeconds == null || clearSettleSeconds <= 0) ? 20 : clearSettleSeconds;
|
||||
bootstrapText = (bootstrapText == null || bootstrapText.isBlank()) ? null : bootstrapText;
|
||||
}
|
||||
|
||||
/**
|
||||
* The text actually sent to the lead's pane once it settles after {@code /clear}: the
|
||||
* operator's configured {@link #bootstrapText} when one is set, otherwise the default
|
||||
* sentence built from {@code resolvedHandoverPath}.
|
||||
*
|
||||
* @param resolvedHandoverPath the ABSOLUTE path {@code dev.ltms.fleet.lead.LeadRollover
|
||||
* #open} already resolved — never the raw configured {@link
|
||||
* #handoverPath}, which may still be relative and would name a
|
||||
* directory the fresh lead session's own pane cannot resolve
|
||||
*/
|
||||
public String bootstrapTextFor(String resolvedHandoverPath) {
|
||||
return bootstrapText != null
|
||||
? bootstrapText
|
||||
: "Fresh lead session: read the handover file at " + resolvedHandoverPath
|
||||
+ " and carry on from there.";
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Watch {@code fleetd.yaml} and re-read it when it changes (CB-559).
|
||||
*
|
||||
@@ -1318,6 +1471,126 @@ public record FleetConfig(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Central allow-list of models any {@code profiles:} entry may name (fleetd ticket: "a central
|
||||
* allow-list of usable models"). Nothing before this block checked a profile's {@code model:}
|
||||
* against anything — it was a free-form string handed straight to the backend adapter, and a
|
||||
* withdrawn or misspelled name failed silently (opencode falls back to a default model rather
|
||||
* than erroring on an unknown {@code -m}) rather than at config load, where a mistake is cheap.
|
||||
*
|
||||
* <p><b>Absent or empty {@code allow:} is "off"</b>, on purpose: this is a large deployment
|
||||
* with gitignored {@code fleetd.yaml} on more than one host, and a change that forced every
|
||||
* operator to enumerate their models before the daemon would start would break every one of
|
||||
* them on upgrade. See {@link FleetConfig#validateModels()}, which is where the allow-list is
|
||||
* actually enforced, at config load.
|
||||
*
|
||||
* <p><b>The list is the authority; a {@code profiles:} entry is checked against it, never the
|
||||
* other way around.</b> Adding or editing a {@code profiles:} entry cannot, by itself, widen
|
||||
* the set of permitted models — only editing {@code models.allow:} itself can. This is the
|
||||
* invariant the ticket asked for: the two blocks are validated in one direction only.
|
||||
*
|
||||
* <p><b>Spawn-time enforcement (fleetd #422, units 2+3) lives outside this record</b> —
|
||||
* {@code CompositePeerLauncher.enforceModelEnabled} and its candidate-set filter read this
|
||||
* block LIVE (through the same kind of supplier {@code weight}/{@code maxLoad} already use),
|
||||
* so the on/off state below is hot: no restart needed. This block itself still only decides
|
||||
* what may be CONFIGURED (membership in {@link #allow}); {@link ModelEntry#enabled} decides
|
||||
* whether a member of that list is currently spawnable. The two questions are deliberately
|
||||
* separate — see {@link ModelEntry}'s javadoc for why turning a model off must never mean
|
||||
* removing it from {@link #allow}.
|
||||
*
|
||||
* @param allow the permitted models, each its own {@link ModelEntry} rather than a bare
|
||||
* string — see that record's javadoc for why. {@code null}/empty ⇒ the block is
|
||||
* treated as absent: {@link FleetConfig#validateModels()} checks nothing.
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record Models(List<ModelEntry> allow) {
|
||||
|
||||
public Models {
|
||||
allow = (allow == null) ? List.of() : List.copyOf(allow);
|
||||
}
|
||||
|
||||
/**
|
||||
* One permitted model, named as a record rather than a bare string on purpose: this ticket
|
||||
* (fleetd #422) is the "later unit" the original comment here predicted — it hangs an on/off
|
||||
* state ({@link #enabled}) off each entry, and a bare {@code List<String>} could not have
|
||||
* grown that field without changing the YAML shape underneath every operator who already
|
||||
* wrote one. {@link #model()} is intentionally a single flat, opaque-string namespace — a
|
||||
* bare Claude id ({@code claude-sonnet-5}) and an opencode provider-prefixed id
|
||||
* ({@code openai/gpt-5.6-terra}) both fit it unchanged, because {@link
|
||||
* FleetConfig#validateModels()} only ever compares a profile's {@code model:} value against
|
||||
* this string for exact equality; it never parses a provider prefix or branches on a
|
||||
* profile's {@code kind:}.
|
||||
*
|
||||
* @param model the model id exactly as a {@code profiles:} entry's {@code model:} field
|
||||
* would name it
|
||||
* @param enabled {@code false} turns spawning onto this model off; {@code null} (the field
|
||||
* omitted — every config written before fleetd #422 is this shape) or
|
||||
* {@code true} leaves it on. Turning a model off must NEVER remove it from
|
||||
* {@link Models#allow} — {@link FleetConfig#validateModels()} checks
|
||||
* <em>membership</em> only, never the on/off state, so an off entry stays a
|
||||
* valid thing for a {@code profiles:} entry to name; only
|
||||
* {@code CompositePeerLauncher}'s spawn-time gate reads {@link #enabled}.
|
||||
* Collapsing the two — turning a model off by deleting its {@code allow:}
|
||||
* entry — would make {@link FleetConfig#validateModels()} refuse the whole
|
||||
* config reload the moment a still-configured profile names it, which is
|
||||
* exactly the restart-to-flip-a-switch problem this field exists to avoid.
|
||||
* <p>A model id can be named by more than one {@code profiles:} entry (e.g.
|
||||
* {@code deepseek-v4-flash} backs both {@code local} and {@code
|
||||
* local-direct} in the live config) — turning it off disables every profile
|
||||
* that names it, on purpose: the model is what a subscription's rate limit
|
||||
* actually constrains, not any one profile alias for it.
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record ModelEntry(String model, Boolean enabled) {
|
||||
public ModelEntry {
|
||||
model = (model == null || model.isBlank()) ? null : model.trim();
|
||||
}
|
||||
|
||||
/**
|
||||
* Back-compat form before {@link #enabled} was added (fleetd #422) — the model is
|
||||
* unconditionally on, exactly as every {@code ModelEntry} behaved before this field
|
||||
* existed. Keeps pre-#422 call sites (and any YAML that omits {@code enabled:})
|
||||
* compiling and behaving identically.
|
||||
*/
|
||||
public ModelEntry(String model) {
|
||||
this(model, null);
|
||||
}
|
||||
|
||||
/** {@code true} unless {@link #enabled} is explicitly {@code false} — absent means on. */
|
||||
public boolean isEnabled() {
|
||||
return !Boolean.FALSE.equals(enabled);
|
||||
}
|
||||
}
|
||||
|
||||
/** {@link #allow}'s model ids, as a set for membership checks. Blank/null entries are dropped. */
|
||||
public Set<String> ids() {
|
||||
Set<String> ids = new java.util.LinkedHashSet<>();
|
||||
for (ModelEntry e : allow) {
|
||||
if (e != null && e.model() != null) {
|
||||
ids.add(e.model());
|
||||
}
|
||||
}
|
||||
return Collections.unmodifiableSet(ids);
|
||||
}
|
||||
|
||||
/**
|
||||
* Model ids currently turned off (fleetd #422: {@link ModelEntry#isEnabled()} {@code
|
||||
* false}). Read live by {@code CompositePeerLauncher}'s spawn gate and by {@code
|
||||
* fleet_profiles}/{@code GET /profiles} — both must read this same accessor off the same
|
||||
* live config so the two surfaces cannot disagree about which model is off (the fleetd
|
||||
* #404 lesson: a status field must read the source the behaviour reads).
|
||||
*/
|
||||
public Set<String> offIds() {
|
||||
Set<String> off = new java.util.LinkedHashSet<>();
|
||||
for (ModelEntry e : allow) {
|
||||
if (e != null && e.model() != null && !e.isEnabled()) {
|
||||
off.add(e.model());
|
||||
}
|
||||
}
|
||||
return Collections.unmodifiableSet(off);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The terminal → lead-name map seeded from the legacy singular {@code primary:} pin (CB-530).
|
||||
*
|
||||
@@ -1569,7 +1842,7 @@ public record FleetConfig(
|
||||
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
|
||||
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds",
|
||||
"memberCredentials", "coordinator", "worktreeGroup", "memberLoginShell", "memberSkills",
|
||||
"idleSleepGuard");
|
||||
"idleSleepGuard", "models", "leadRollover");
|
||||
|
||||
/** Load and validate config from {@code path}. */
|
||||
public static FleetConfig load(Path path) {
|
||||
@@ -2254,10 +2527,19 @@ public record FleetConfig(
|
||||
// enabled: true — see its javadoc), so defaulting the block here would change nothing a
|
||||
// reader observes and would only obscure that "block omitted" and "block present and
|
||||
// enabled" are deliberately the same outcome.
|
||||
// models is left as-is, like broker/primary/coordinator above: null/empty is "off", and an
|
||||
// absent block must validate nothing (see Models's javadoc) — defaulting it here to an
|
||||
// empty Models would be a no-op for validateModels() either way, since an empty allow-list
|
||||
// already means "check nothing", so there is nothing to gain and one more null check to
|
||||
// avoid by leaving it exactly as configured.
|
||||
// leadRollover is left as-is, like leadHeartbeat above: null is "off", and LeadRollover's
|
||||
// own compact constructor defaults the fields of a block that IS present. Defaulting it
|
||||
// here would construct a LeadRollover object (via Fleetd.java's presence gate) for every
|
||||
// config that never mentioned it.
|
||||
return new FleetConfig(b, herdrSocket, memberHerdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
|
||||
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
|
||||
quarantineCooldown, mc, coordinator, worktreeGroup, memberLoginShell, memberSkills,
|
||||
idleSleepGuard);
|
||||
idleSleepGuard, models, leadRollover);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -2340,6 +2622,26 @@ public record FleetConfig(
|
||||
+ "lead tabs cannot be confused.");
|
||||
}
|
||||
|
||||
/**
|
||||
* Reject a present {@code leadRollover:} block with no (or a blank) {@code handoverPath}
|
||||
* (fleetd #480). There is no sane non-null default for an operator-specific file path, unlike
|
||||
* every other field on {@link LeadRollover}, which {@link LeadRollover}'s own compact
|
||||
* constructor already defaults — so this is the one field that must be refused at load rather
|
||||
* than silently defaulted to something that would never match a real handover file.
|
||||
*
|
||||
* @throws IllegalStateException when {@code leadRollover} is present but {@code handoverPath}
|
||||
* is {@code null} or blank
|
||||
*/
|
||||
public void validateLeadRollover() {
|
||||
if (leadRollover == null) {
|
||||
return;
|
||||
}
|
||||
if (leadRollover.handoverPath() == null || leadRollover.handoverPath().isBlank()) {
|
||||
throw new IllegalStateException("refusing to start: leadRollover.handoverPath is "
|
||||
+ "required when leadRollover: is present.");
|
||||
}
|
||||
}
|
||||
|
||||
/** Case-insensitive prefix test that tolerates a null/blank label. */
|
||||
private static boolean startsWithIgnoreCase(String label, String prefix) {
|
||||
if (label == null || prefix == null || prefix.isBlank()) {
|
||||
@@ -2477,6 +2779,118 @@ public record FleetConfig(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Reject a {@code profiles:} entry whose {@code model:} is not on the configured {@link
|
||||
* #models} allow-list.
|
||||
*
|
||||
* <p>Absent or empty {@code models.allow:} validates nothing — see {@link Models}'s javadoc:
|
||||
* an existing config with no such block must keep working exactly as it does today. Once the
|
||||
* operator declares at least one entry, every profile's {@code model:} (when set — a profile
|
||||
* may legitimately leave it {@code null}, e.g. a {@code subscription: true} profile relying on
|
||||
* the account's own default) must equal one of {@link Models#ids()} exactly. The comparison is
|
||||
* a flat string match: a bare Claude id and an opencode {@code provider/model} id are both
|
||||
* just opaque strings here, so nothing here needs to know which {@code kind:} a profile runs.
|
||||
*
|
||||
* <p>The check runs one direction only, by construction: it reads {@link #profiles} and
|
||||
* {@link #models}, and only ever adds to {@code bad} when a profile's model is missing from
|
||||
* the allow-list. Nothing here can be satisfied by widening a {@code profiles:} entry — only
|
||||
* editing {@code models.allow:} itself changes what passes. That is the invariant the ticket
|
||||
* asked for: the list is the authority, profiles are checked against it.
|
||||
*
|
||||
* @throws IllegalStateException when any profile names a model outside the configured
|
||||
* allow-list, naming both the model and the profile that wanted it
|
||||
*/
|
||||
public void validateModels() {
|
||||
if (models == null || models.allow().isEmpty()) {
|
||||
return;
|
||||
}
|
||||
Set<String> allowed = models.ids();
|
||||
List<String> bad = new ArrayList<>();
|
||||
profiles.forEach((name, p) -> {
|
||||
String model = p.model();
|
||||
if (model != null && !model.isBlank() && !allowed.contains(model)) {
|
||||
bad.add("profile '" + name + "' names model '" + model + "', which is not in "
|
||||
+ "models.allow: (have: " + allowed + ").");
|
||||
}
|
||||
});
|
||||
if (!bad.isEmpty()) {
|
||||
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Runs every validator this class declares — found by reflection, not by name.
|
||||
*
|
||||
* <p>fleetd ticket "central allow-list of usable models", follow-up: mutation testing found
|
||||
* that although each of the six validators above was well pinned on its own, nothing proved
|
||||
* either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()}) still
|
||||
* invoked it — deleting a call site left the full suite green. The fix is not a seventh test
|
||||
* per caller; a hand-maintained list of six names here would have the exact same defect its
|
||||
* own javadoc would warn against: the seventh validator someone adds next month has no reason
|
||||
* to be added to it. So this method does not name any validator. It sweeps {@link
|
||||
* #getClass()}'s own public, no-argument, {@code void} methods whose name starts with {@code
|
||||
* "validate"} (excluding itself) and invokes every one it finds, via {@link
|
||||
* #invokeAllValidators}. A new {@code validateXxx()} method is therefore wired into both
|
||||
* callers the moment it is written — there is no second step to forget, and so no state in
|
||||
* which it silently never runs.
|
||||
*
|
||||
* <p>{@code Fleetd.main} and {@link ConfigRef#reload()} each call this one method instead of
|
||||
* the six individually — see the comments at those two call sites for why
|
||||
* each must run it.
|
||||
*
|
||||
* <p>Methods run in a fixed (alphabetical) order, so a config with more than one violation
|
||||
* always names the same one first, on every run.
|
||||
*
|
||||
* @throws IllegalStateException (or whatever unchecked exception a validator itself throws),
|
||||
* propagated unchanged from the first validator, in that order,
|
||||
* that finds a problem
|
||||
*/
|
||||
public void validateAll() {
|
||||
invokeAllValidators(this);
|
||||
}
|
||||
|
||||
/**
|
||||
* The reflective sweep behind {@link #validateAll()}, kept as its own method — taking any
|
||||
* {@code target}, not just {@code this} — so a test can prove the MECHANISM is generic (it
|
||||
* would sweep a seventh {@code validateXxx()} method added to any class, not just something
|
||||
* special-cased to today's six on {@link FleetConfig}) without needing to add a real, unwanted
|
||||
* seventh validator to this class just to exercise that claim. See {@code
|
||||
* FleetConfigValidateAllTest} for that proof.
|
||||
*
|
||||
* @param target an object whose public, no-argument, {@code void} methods named {@code
|
||||
* validateXxx} (any name starting with {@code "validate"}, excluding {@code
|
||||
* validateAll} itself) should all run, in alphabetical-by-name order
|
||||
*/
|
||||
static void invokeAllValidators(Object target) {
|
||||
List<Method> methods = new ArrayList<>();
|
||||
for (Method m : target.getClass().getMethods()) {
|
||||
if (Modifier.isPublic(m.getModifiers())
|
||||
&& m.getParameterCount() == 0
|
||||
&& m.getReturnType() == void.class
|
||||
&& m.getName().startsWith("validate")
|
||||
&& !m.getName().equals("validateAll")) {
|
||||
methods.add(m);
|
||||
}
|
||||
}
|
||||
methods.sort(Comparator.comparing(Method::getName));
|
||||
for (Method m : methods) {
|
||||
try {
|
||||
m.invoke(target);
|
||||
} catch (InvocationTargetException e) {
|
||||
Throwable cause = e.getCause();
|
||||
if (cause instanceof RuntimeException re) {
|
||||
throw re;
|
||||
}
|
||||
if (cause instanceof Error err) {
|
||||
throw err;
|
||||
}
|
||||
throw new IllegalStateException("validator " + m.getName() + " failed", cause);
|
||||
} catch (IllegalAccessException e) {
|
||||
throw new IllegalStateException("cannot invoke validator " + m.getName(), e);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/** True for the loopback addresses and the unspecified-but-local forms we treat as same-host. */
|
||||
private static boolean isLoopbackBind(String host) {
|
||||
if (host == null || host.isBlank()) {
|
||||
|
||||
@@ -3,7 +3,7 @@ package dev.ltms.fleet.herdr;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
|
||||
/**
|
||||
* Client face onto the herdr daemon (protocol 14, herdr 0.7.0).
|
||||
* Client face onto the herdr daemon (protocol 19, herdr 0.8.0).
|
||||
*
|
||||
* <p>This is the ONLY thing in {@code fleetd} that speaks to herdr. Every method
|
||||
* maps to a herdr JSON-RPC call over its Unix domain socket. Requests are
|
||||
|
||||
@@ -8,7 +8,7 @@ import com.fasterxml.jackson.databind.node.ObjectNode;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
|
||||
/**
|
||||
* Wire codec for herdr's newline-delimited JSON-RPC (protocol 14).
|
||||
* Wire codec for herdr's newline-delimited JSON-RPC (protocol 19).
|
||||
*
|
||||
* <p>Split out from the socket so the framing rules — the ones that actually bit us
|
||||
* during the spike (id MUST be a string; response carries {@code result} or
|
||||
|
||||
@@ -698,17 +698,57 @@ public final class CompletionResolver implements TurnListener {
|
||||
}
|
||||
|
||||
/**
|
||||
* Coverage summary for the CB-578 stage A exhausted-pattern classification, logged at startup
|
||||
* What an unset pattern key means for the classification it configures (fleetd#415).
|
||||
* {@code coverage()} cannot infer this from the key's name — the two keys it currently
|
||||
* describes disagree on it, and a string comparison on the name would just move the same bug
|
||||
* to a new spot — so every caller must state it explicitly.
|
||||
*
|
||||
* <p><strong>This alone does not prove a caller passes the right one for its key.</strong> A
|
||||
* test that calls {@code coverage()} directly and supplies the meaning itself only proves this
|
||||
* enum is worded correctly, never that {@code Fleetd}'s two call sites pair each key with its
|
||||
* true meaning — that pairing is #415's actual defect. Measured on review: swapping the two
|
||||
* {@code UnsetMeaning} arguments at those call sites (giving {@code exhaustedPattern} the
|
||||
* built-in-default wording and {@code errorPattern} the off wording — #415's exact defect with
|
||||
* the keys exchanged) compiled with 0 errors and left all 1506 existing tests green. See
|
||||
* {@code dev.ltms.fleet.Fleetd#exhaustedPatternCoverageLine}/{@code #errorPatternCoverageLine}
|
||||
* and {@code FleetdPatternCoverageLineTest}, which exists specifically to catch that swap.
|
||||
*/
|
||||
public enum UnsetMeaning {
|
||||
/** No fallback exists: a profile with no configured pattern truly has this classification off. */
|
||||
OFF,
|
||||
/** A built-in pattern applies when unset: the classification still runs for that profile. */
|
||||
BUILT_IN_DEFAULT
|
||||
}
|
||||
|
||||
/**
|
||||
* Coverage summary for a fleetd#201/CB-578-style pattern-key classification, logged at startup
|
||||
* the way {@link dev.ltms.fleet.health.FleetHealthMonitor#coverage} is — so an operator can
|
||||
* see whether the classification is on, and for which profiles, without reading every
|
||||
* profile's config by hand.
|
||||
*
|
||||
* <p>fleetd#415: this method measures <em>pattern coverage</em> — how many profiles set the
|
||||
* key — which is not the same thing as <em>feature state</em> for a key with a fallback. For
|
||||
* {@code errorPattern}, an empty {@code configuredProfiles} still runs the classification
|
||||
* against {@code CompletionResolver}'s built-in compatibility pattern ({@link #BACKEND_ERROR}
|
||||
* at line ~84); for {@code exhaustedPattern} there is no fallback, so empty really does mean
|
||||
* off. {@code unsetMeaning} is the single, required source of that fact — see
|
||||
* {@link dev.ltms.fleet.config.FleetConfig#rejectMalformedProfilePatterns} lines ~2029-2032 for
|
||||
* where it is documented for config authors. It is a required parameter, not a defaulted
|
||||
* overload: a third pattern key added later must supply one to compile at all, rather than
|
||||
* silently inheriting whichever wording this method happened to default to.
|
||||
*
|
||||
* @param allProfiles every configured profile name
|
||||
* @param configuredProfiles the subset of {@code allProfiles} that carry an exhausted pattern
|
||||
* @param configuredProfiles the subset of {@code allProfiles} that carry the pattern
|
||||
*/
|
||||
public static String coverage(String patternKey, Set<String> allProfiles, Set<String> configuredProfiles) {
|
||||
public static String coverage(String patternKey, UnsetMeaning unsetMeaning, Set<String> allProfiles,
|
||||
Set<String> configuredProfiles) {
|
||||
if (configuredProfiles.isEmpty()) {
|
||||
return "off (no profile has an " + patternKey + " configured; profiles: " + sorted(allProfiles) + ")";
|
||||
return switch (unsetMeaning) {
|
||||
case OFF -> "off (no profile has an " + patternKey + " configured; profiles: "
|
||||
+ sorted(allProfiles) + ")";
|
||||
case BUILT_IN_DEFAULT -> "built-in default for all profiles (no profile customises "
|
||||
+ patternKey + "; profiles: " + sorted(allProfiles) + ")";
|
||||
};
|
||||
}
|
||||
Set<String> unconfigured = new TreeSet<>(allProfiles);
|
||||
unconfigured.removeAll(configuredProfiles);
|
||||
|
||||
@@ -0,0 +1,91 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
|
||||
import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.function.Supplier;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
/**
|
||||
* fleetd #446: the single live source for "is {@code exhaustedPattern} configured for this
|
||||
* profile, and what does it compile to". Read fresh off the config supplier on every call — the
|
||||
* same reason {@code CompositePeerLauncher#models0} is a live supplier read rather than a value
|
||||
* captured at construction (see that class's doc) — so an operator can arm or disarm usage-limit
|
||||
* detection for a profile by editing {@code exhaustedPattern} and reloading, with no restart.
|
||||
*
|
||||
* <p>Before this class, {@code exhaustedPattern} was compiled once into a {@code Map<String,
|
||||
* Pattern>} built inside {@code Fleetd.main} at startup ({@code ConfigRef}'s class doc used to
|
||||
* list it under <em>Deferred</em>, CB-578 stage A) — the asymmetry fleetd #446 exists to close:
|
||||
* an operator could turn a model off at runtime ({@code models.allow}'s {@code enabled: false},
|
||||
* hot since fleetd #422) but could not arm the detector that would tell them to, without a
|
||||
* restart. That was backwards for a feature whose whole point is to react while the fleet runs.
|
||||
*
|
||||
* <p>{@link #patternFor} is what {@link CompletionResolver} enforces on (via the {@link
|
||||
* ExhaustedPatternLookup} production wiring in {@code Fleetd.main}); {@link #armed} is what
|
||||
* {@code fleet_profiles}/{@code GET /profiles} report as {@code exhaustionDetectionArmed}. Both
|
||||
* read this ONE object, so the report can never disagree with the behaviour — the same rule
|
||||
* {@code CompositePeerLauncher#modelGateState()}'s javadoc states for the model gate's own
|
||||
* armed/off pair (fleetd #404): "armed" and "which models are off" must come from one read of the
|
||||
* same accessor the gate enforces on.
|
||||
*
|
||||
* <h2>Cache eviction</h2>
|
||||
* Cached by profile NAME, not by pattern text. A {@code Map<String, Pattern>} keyed by the
|
||||
* pattern STRING would grow by one entry per distinct regex ever typed for any profile across the
|
||||
* daemon's uptime — unbounded in practice, because tuning a regex to match a backend's exact
|
||||
* wording is exactly the kind of edit an operator makes several times while getting it right, and
|
||||
* every edit-and-reload cycle would leave the previous attempt's compiled {@link Pattern} behind
|
||||
* forever. Keyed by profile name instead, this cache holds at most one entry per profile name that
|
||||
* has ever been looked up — and that key space is bounded by the (small, human-authored) set of
|
||||
* configured profiles, which changes only on a restart: adding or removing a profile is itself a
|
||||
* deferred key (a new backend needs its own launcher, built once — see {@code ConfigRef}'s class
|
||||
* doc), so profile names do not churn the way pattern text does. Re-editing an EXISTING profile's
|
||||
* {@code exhaustedPattern} — the case this class exists to make hot — simply overwrites that
|
||||
* profile's one cache entry; it never adds a new one.
|
||||
*/
|
||||
public final class LiveExhaustedPatterns {
|
||||
|
||||
/** A compiled pattern paired with the source string it was compiled from, for change detection. */
|
||||
private record Cached(String source, Pattern compiled) {
|
||||
}
|
||||
|
||||
private final Supplier<Map<String, FleetConfig.Profile>> profiles;
|
||||
private final ConcurrentHashMap<String, Cached> cache = new ConcurrentHashMap<>();
|
||||
|
||||
public LiveExhaustedPatterns(Supplier<Map<String, FleetConfig.Profile>> profiles) {
|
||||
this.profiles = Objects.requireNonNull(profiles, "profiles");
|
||||
}
|
||||
|
||||
/**
|
||||
* The compiled {@code exhaustedPattern} currently configured for {@code profileName}, or
|
||||
* {@code null} when that profile is unknown or has none configured. Recompiles only when the
|
||||
* live pattern text differs from what is cached for this profile name; {@code
|
||||
* FleetConfig#rejectMalformedProfilePatterns} already refuses a config (at load and at reload)
|
||||
* whose {@code exhaustedPattern} does not compile, so this is not expected to throw in
|
||||
* production — it is not defended against here for that reason, the same trust
|
||||
* {@code CompositePeerLauncher#models0} places in config validation having already run.
|
||||
*/
|
||||
public Pattern patternFor(String profileName) {
|
||||
if (profileName == null) {
|
||||
return null;
|
||||
}
|
||||
FleetConfig.Profile profile = profiles.get().get(profileName);
|
||||
if (profile == null || !profile.hasExhaustedPattern()) {
|
||||
return null;
|
||||
}
|
||||
String source = profile.exhaustedPattern();
|
||||
Cached cached = cache.get(profileName);
|
||||
if (cached != null && cached.source().equals(source)) {
|
||||
return cached.compiled();
|
||||
}
|
||||
Cached fresh = new Cached(source, Pattern.compile(source));
|
||||
cache.put(profileName, fresh);
|
||||
return fresh.compiled();
|
||||
}
|
||||
|
||||
/** Whether {@code profileName} currently has a usage-limit pattern configured (live). */
|
||||
public boolean armed(String profileName) {
|
||||
return patternFor(profileName) != null;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,588 @@
|
||||
package dev.ltms.fleet.lead;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.AgentStatus;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.io.UncheckedIOException;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.UUID;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.function.Consumer;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
* fleetd #480: replace a lead session that has decided it is ready to be rolled over, without an
|
||||
* operator doing it by hand. A lead writes a handover file, calls {@link #open}, and then — once
|
||||
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation
|
||||
* clears the lead's own pane and bootstraps a fresh session against that file.
|
||||
*
|
||||
* <p>This is the executor only. Nothing in this ticket wires an MCP tool onto {@link #open}/
|
||||
* {@link #confirm}/{@link #cancel} — that is a separate, later unit; until it lands, nothing calls
|
||||
* this class at all.
|
||||
*
|
||||
* <p><strong>{@code confirm()} cannot roll inline — a fleetd #480 correction.</strong> The first
|
||||
* version of this class called {@code agents.send(lead, "/clear")} directly from inside {@code
|
||||
* confirm()}, then polled for the pane to become injectable again. That is wrong, because {@code
|
||||
* confirm()} is called BY the lead, FROM the lead's own turn: the lead's pane is {@code WORKING}
|
||||
* for the whole duration of that call and cannot possibly report injectable until {@code confirm()}
|
||||
* itself returns. The poll always timed out — but only after the {@code /clear} had already been
|
||||
* sent and queued in the pane, where it fired the instant the turn ended anyway. The result was the
|
||||
* worst outcome this feature can produce: a silently destroyed lead context with no fresh session
|
||||
* ever started, and a refusal return value that claimed nothing had happened.
|
||||
*
|
||||
* <p>The fix: {@link #confirm} validates every gate, then does no I/O against the lead's own pane
|
||||
* at all — it only records that the request is approved and hands a one-shot continuation to
|
||||
* {@code continuationRunner} before returning. That continuation is what actually touches the pane,
|
||||
* once the calling turn has ended, in this order:
|
||||
* <ol>
|
||||
* <li>wait for the lead's own pane to report a real turn boundary — {@code IDLE} or {@code
|
||||
* DONE}, never merely {@code BLOCKED} — i.e. wait for the very {@code confirm()} call that
|
||||
* approved this roll to finish its turn — bounded by {@code turnSettleSeconds}. <strong>If
|
||||
* this never happens, nothing else in this list runs: no {@code /clear} is ever sent.</strong>
|
||||
* A lead that never goes idle is a lead still doing real work, and clearing it would throw
|
||||
* away live context — exactly the failure this correction exists to prevent.</li>
|
||||
* <li>{@code agents.send(lead, "/clear")}</li>
|
||||
* <li>wait for {@code /clear} to be picked up and settle, bounded by {@code clearSettleSeconds}
|
||||
* (fleetd #489: no longer a plain re-check of the same boundary — {@code /clear} starts no
|
||||
* turn of its own, so this instead nudges the submit keystroke while no pickup has been seen,
|
||||
* then waits for a real {@code WORKING} → {@code IDLE}/{@code DONE} boundary once one has;
|
||||
* see {@link #waitForClearPickupAndSettle})</li>
|
||||
* <li>{@code agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()))}</li>
|
||||
* </ol>
|
||||
* A {@link #confirm} that returns {@link RollDecision#approved()} therefore means <em>"every gate
|
||||
* passed and the roll is scheduled"</em>, never <em>"the pane has been cleared"</em> — the pane may
|
||||
* still be mid-turn, possibly for a long time, when the caller gets that answer back.
|
||||
*
|
||||
* <p><strong>The safety invariant survives this change, restated precisely.</strong> The ticket
|
||||
* that first defined this class required "no timer, no scheduler, no background thread" so that
|
||||
* nothing but an explicit {@link #confirm} call could ever cause a {@code /clear}. That invariant
|
||||
* is about INITIATIVE, not about synchronicity, and this correction keeps it: {@code
|
||||
* continuationRunner} launches a single-shot task that exists only because one specific,
|
||||
* already-approved {@link #confirm} call created it — it is not recurring, it is not started at
|
||||
* construction time or on any schedule, and no two invocations of it ever share state. A recurring
|
||||
* heartbeat or timer that could decide on its own initiative to roll a pane is still, and will
|
||||
* always be, absent from this class. <strong>Nothing but an explicit {@link #confirm} call that
|
||||
* passes every gate can ever cause a {@code /clear} — that call may simply finish its own work
|
||||
* slightly later than the method return, as a continuation of the same approved request, rather
|
||||
* than entirely inside the method body.</strong>
|
||||
*
|
||||
* <p><strong>Identity is resolved by the caller, never looked up here — a second fleetd #480
|
||||
* correction.</strong> The first version resolved the pane to clear via {@code
|
||||
* PrimaryRegistry#primaryTerminal()}. That is correct for a background loop with no caller (see
|
||||
* {@code dev.ltms.fleet.msg.LeadHeartbeatLoop}), but wrong here and a violation of this project's
|
||||
* own charter invariant 3 — "identity comes from the connection, never an argument." This daemon
|
||||
* can hold more than one labelled lead tab (see {@code LeadLauncher}'s fleetd #359 two-reading
|
||||
* dead-tab cleanup), so a single-slot lookup lets lead X's {@link #confirm} clear lead Y's pane: an
|
||||
* unrecoverable loss of someone else's context, and a different lead's context at that. Both
|
||||
* {@link #open} and {@link #confirm} now take the lead's terminal id as a parameter instead —
|
||||
* {@link #open} stores it on the {@link PendingRollover}, and {@link #confirm} refuses with {@link
|
||||
* RefusalReason#NOT_YOUR_ROLLOVER} unless the caller's terminal matches the one {@link #open}
|
||||
* recorded. <strong>The terminal id passed to both methods must come from the MCP layer's own
|
||||
* connection-based caller resolution — the same source {@code auth/CallerResolver#resolve} uses to
|
||||
* build a {@code Principal.leader(...)} (see its use of {@code ConnectionIdentity.Caller#terminal})
|
||||
* — never a value the client supplies or chooses.</strong> The later MCP-tool unit that wires
|
||||
* {@link #open}/{@link #confirm} must pass the resolved caller terminal, not a request field.
|
||||
*
|
||||
* <p><strong>A relative {@code handoverPath} resolves against the CALLING lead's workspace, never
|
||||
* the daemon's own cwd — a fleetd #480 follow-up.</strong> The daemon and a lead's own pane can
|
||||
* have different working directories (this repo nests {@code fleetd/} inside its own root, so the
|
||||
* daemon's cwd and {@code fleet.leaders.<name>.cwd} already differ on this host). {@link #open}
|
||||
* resolves {@code cfg.handoverPath()} to an ABSOLUTE path exactly once — against {@code
|
||||
* leadWorkspace.apply(leadTerminal)} when that lookup returns a non-null, non-blank workspace, and
|
||||
* against {@code System.getProperty("user.dir")} otherwise (the same fallback {@code
|
||||
* LeadLauncher#launch} already uses for a lead with no configured {@code cwd}) — and stores only
|
||||
* that absolute path on {@link PendingRollover}. Every later read of {@code
|
||||
* PendingRollover#handoverPath()} (the freshness/exists/empty checks in {@link #checkHandover},
|
||||
* the value {@code FleetMcp} hands back to the lead in the {@code open} response so it knows where
|
||||
* to WRITE the file, and {@link FleetConfig.LeadRollover#bootstrapTextFor} which names it in the
|
||||
* text typed into the fresh session) therefore already sees the resolved absolute form and never
|
||||
* needs to resolve anything itself.
|
||||
*/
|
||||
public final class LeadRollover {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(LeadRollover.class);
|
||||
|
||||
/** Poll interval while waiting for the lead's pane to settle after {@code /clear}. */
|
||||
static final long SETTLE_POLL_MS = 250;
|
||||
|
||||
/**
|
||||
* How many consecutive not-yet-picked-up polls {@link #waitForClearPickupAndSettle} allows
|
||||
* before releasing rather than wedging the roll — the same constant and the same
|
||||
* release-not-wedge choice {@link dev.ltms.fleet.inject.Injector} already makes for its own
|
||||
* post-turn {@code /clear} housekeeping (fleetd #306). <strong>This bounds the number of
|
||||
* consecutive polls, not the number of nudges:</strong> the first {@code PICKUP_GRACE_POLLS - 1}
|
||||
* of those polls each send a nudge, and the {@code PICKUP_GRACE_POLLS}th releases instead of
|
||||
* nudging again — so 8 polls produce 7 nudges, not 8.
|
||||
*/
|
||||
static final int PICKUP_GRACE_POLLS = 8;
|
||||
|
||||
/**
|
||||
* One request opened by {@link #open}, pending its {@link #confirm} (or {@link #cancel}).
|
||||
*
|
||||
* @param leadTerminal the lead pane that opened this request — the only terminal that may
|
||||
* later {@link #confirm} it (see {@link RefusalReason#NOT_YOUR_ROLLOVER})
|
||||
* @param handoverPath the ABSOLUTE, resolved handover path — never the raw configured value,
|
||||
* which may have been relative. {@link #open} resolves it once, against the
|
||||
* calling lead's workspace, before storing it here; see this class's
|
||||
* javadoc. This is the value the MCP layer hands back to the lead as
|
||||
* "write your file here", so callers may rely on it always being absolute.
|
||||
*/
|
||||
public record PendingRollover(String token, String leadTerminal, String handoverPath,
|
||||
long requestedAtMillis) {}
|
||||
|
||||
/** Which check refused a {@link #confirm} call, named so a caller can act on it. */
|
||||
public enum RefusalReason {
|
||||
/** {@code leadRollover:} is not configured — absent at construction, or removed since. */
|
||||
NOT_CONFIGURED,
|
||||
/** {@code token} names no pending request: never opened, already confirmed, or cancelled. */
|
||||
UNKNOWN_TOKEN,
|
||||
/**
|
||||
* The caller's terminal does not match the terminal that {@link #open} recorded for this
|
||||
* token. Only the lead that opened a request may confirm it (fleetd #480 correction 2).
|
||||
*/
|
||||
NOT_YOUR_ROLLOVER,
|
||||
/** {@code requireOperatorConfirm: true} and the caller passed {@code operatorConfirmed: false}. */
|
||||
OPERATOR_NOT_CONFIRMED,
|
||||
/** The handover file does not exist. */
|
||||
HANDOVER_MISSING,
|
||||
/** The handover file exists but is empty. */
|
||||
HANDOVER_EMPTY,
|
||||
/**
|
||||
* The handover file's modified time is not after {@link #open}'s request timestamp, or is
|
||||
* older than {@code maxDocAgeSeconds}.
|
||||
*/
|
||||
HANDOVER_STALE
|
||||
}
|
||||
|
||||
/**
|
||||
* The outcome of a {@link #confirm} call. {@link #approved()} means every gate passed and the
|
||||
* roll has been handed to a one-shot continuation — <strong>not</strong> that the pane has been
|
||||
* cleared; the continuation may still be waiting for the calling turn to end when this returns.
|
||||
* Whether the deferred roll itself later goes on to clear the pane, refuse for never going
|
||||
* idle, or refuse for never re-settling after {@code /clear} is logged only (see this class's
|
||||
* javadoc) — there is deliberately no synchronous caller left by that point to hand a result to.
|
||||
*/
|
||||
public record RollDecision(boolean accepted, RefusalReason reason, String detail) {
|
||||
static RollDecision approved() {
|
||||
return new RollDecision(true, null, "confirmed; the roll will run once the calling turn ends");
|
||||
}
|
||||
|
||||
static RollDecision refused(RefusalReason reason, String detail) {
|
||||
return new RollDecision(false, reason, detail);
|
||||
}
|
||||
}
|
||||
|
||||
private final AgentControl agents;
|
||||
private final Supplier<FleetConfig.LeadRollover> configSupplier;
|
||||
/**
|
||||
* Terminal id → that lead's configured workspace directory (their {@code
|
||||
* fleet.leaders.<name>.cwd}), or {@code null} when the terminal names no currently-recognised
|
||||
* lead. {@link #open} calls this to resolve a relative {@code handoverPath} — see this class's
|
||||
* javadoc. Required: there is no sane default that would not silently reintroduce the
|
||||
* daemon-cwd bug this parameter exists to fix.
|
||||
*/
|
||||
private final Function<String, String> leadWorkspace;
|
||||
private final LongSupplier nowMillis;
|
||||
private final Runnable settleSleeper;
|
||||
/**
|
||||
* Launches the post-{@code confirm()} continuation. Production uses a single unstarted virtual
|
||||
* thread per confirmed request — see this class's javadoc for why that is a single-shot task,
|
||||
* not a background scheduler. Tests inject {@code Runnable::run} to make the continuation run
|
||||
* synchronously and deterministically on the calling thread.
|
||||
*/
|
||||
private final Consumer<Runnable> continuationRunner;
|
||||
private final Map<String, PendingRollover> pending = new ConcurrentHashMap<>();
|
||||
|
||||
/** Production constructor — wall clock, real sleep between settle polls, a real virtual thread. */
|
||||
public LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
|
||||
Function<String, String> leadWorkspace) {
|
||||
this(agents, configSupplier, leadWorkspace, System::currentTimeMillis,
|
||||
() -> sleepUninterruptibly(SETTLE_POLL_MS),
|
||||
r -> Thread.ofVirtual().name("lead-rollover-continuation-").start(r));
|
||||
}
|
||||
|
||||
/**
|
||||
* Full constructor — an injectable wall-clock supplier, settle-poll sleeper, and continuation
|
||||
* runner, for tests. {@code nowMillis} MUST be a wall-clock source (e.g. {@code
|
||||
* System.currentTimeMillis()}), never {@code System.nanoTime()}: the freshness check compares
|
||||
* against a file's modified time, which only a wall clock is comparable to, and {@code
|
||||
* nanoTime} freezes while the host sleeps (fleetd #386).
|
||||
*/
|
||||
LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
|
||||
Function<String, String> leadWorkspace, LongSupplier nowMillis,
|
||||
Runnable settleSleeper, Consumer<Runnable> continuationRunner) {
|
||||
this.agents = agents;
|
||||
this.configSupplier = configSupplier;
|
||||
this.leadWorkspace = leadWorkspace;
|
||||
this.nowMillis = nowMillis;
|
||||
this.settleSleeper = settleSleeper;
|
||||
this.continuationRunner = continuationRunner;
|
||||
}
|
||||
|
||||
private static void sleepUninterruptibly(long ms) {
|
||||
try {
|
||||
Thread.sleep(ms);
|
||||
} catch (InterruptedException e) {
|
||||
Thread.currentThread().interrupt();
|
||||
// preserve the interrupt flag but continue — this poll loop should not be aborted by an
|
||||
// interrupt that was not meant for it.
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The lead says it is ready to be replaced. Generates a token and records the resolved
|
||||
* handover path, this moment's wall-clock timestamp (the baseline {@link #confirm} checks the
|
||||
* handover file's modified time against), and {@code leadTerminal} — only that exact terminal
|
||||
* may later {@link #confirm} this token.
|
||||
*
|
||||
* @param leadTerminal the calling lead's terminal id, resolved by the MCP layer from the
|
||||
* connection (see this class's javadoc) — never a client-supplied value
|
||||
* @param reason free-text audit note (logged only; not otherwise interpreted or stored)
|
||||
* @throws IllegalStateException if {@code leadRollover:} is not configured
|
||||
* @throws IllegalArgumentException if {@code leadTerminal} is null or blank
|
||||
*/
|
||||
public PendingRollover open(String leadTerminal, String reason) {
|
||||
FleetConfig.LeadRollover cfg = configSupplier.get();
|
||||
if (cfg == null) {
|
||||
throw new IllegalStateException("leadRollover: is not configured");
|
||||
}
|
||||
if (leadTerminal == null || leadTerminal.isBlank()) {
|
||||
throw new IllegalArgumentException("leadTerminal is required — it must be resolved from "
|
||||
+ "the caller's connection, never accepted as a client-chosen argument");
|
||||
}
|
||||
String token = UUID.randomUUID().toString();
|
||||
long requestedAt = nowMillis.getAsLong();
|
||||
String resolvedPath = resolveHandoverPath(cfg.handoverPath(), leadTerminal);
|
||||
PendingRollover p = new PendingRollover(token, leadTerminal, resolvedPath, requestedAt);
|
||||
pending.put(token, p);
|
||||
if (resolvedPath.equals(cfg.handoverPath())) {
|
||||
log.info("lead-rollover: open token={} lead={} handoverPath={} reason={}",
|
||||
token, leadTerminal, resolvedPath, reason);
|
||||
} else {
|
||||
// The configured value was relative (or merely un-normalized) and resolved to a
|
||||
// different string — log both, so an operator reading this line can see which
|
||||
// directory the daemon actually looked in, not just the value it was given.
|
||||
log.info("lead-rollover: open token={} lead={} configuredHandoverPath={} "
|
||||
+ "resolvedHandoverPath={} reason={}",
|
||||
token, leadTerminal, cfg.handoverPath(), resolvedPath, reason);
|
||||
}
|
||||
return p;
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve {@code configured} to an absolute path exactly once, here, so nothing downstream
|
||||
* ({@link #checkHandover}, the MCP layer's {@code open} response, {@link
|
||||
* FleetConfig.LeadRollover#bootstrapTextFor}) ever has to resolve — or worse, silently
|
||||
* mis-resolve — a relative path again.
|
||||
*
|
||||
* <p><strong>The return value is GUARANTEED absolute, not merely usually absolute.</strong>
|
||||
* {@code leadWorkspace.apply(leadTerminal)} returns an operator-configured {@code
|
||||
* fleet.leaders.<name>.cwd} string, and nothing forces an operator to write an absolute one —
|
||||
* a relative {@code cwd} resolved with plain {@link Path#resolve} would still yield a relative
|
||||
* result, silently reopening the exact bug this class exists to fix (every later reader back to
|
||||
* interpreting an ambiguous string against ITS OWN working directory). {@link
|
||||
* Path#toAbsolutePath()} closes that: it resolves any remaining relative path against {@code
|
||||
* user.dir} (the JVM's own cwd), which is the correct base for an operator-written path the
|
||||
* daemon process itself is meant to interpret, exactly like the {@code user.dir} fallback used
|
||||
* below. Applying it unconditionally on both branches means the ALREADY-absolute branch stays a
|
||||
* no-op (an absolute path is unaffected by {@code toAbsolutePath()}) while the relative-{@code
|
||||
* cwd} branch above is closed the same way.
|
||||
*
|
||||
* <ul>
|
||||
* <li>already absolute → returned unchanged (normalized)</li>
|
||||
* <li>relative → resolved against {@code leadWorkspace.apply(leadTerminal)} when that is
|
||||
* non-null and non-blank; otherwise against {@code System.getProperty("user.dir")} — the
|
||||
* same fallback {@code LeadLauncher#launch} uses for a lead with no configured {@code
|
||||
* cwd}. If {@code leadWorkspace}'s own answer is itself relative (an operator wrote a
|
||||
* relative {@code cwd:}), the result is finished off against the daemon's own
|
||||
* {@code user.dir} — see the paragraph above.</li>
|
||||
* </ul>
|
||||
*/
|
||||
private String resolveHandoverPath(String configured, String leadTerminal) {
|
||||
Path path = Path.of(configured);
|
||||
if (path.isAbsolute()) {
|
||||
// toAbsolutePath() is a no-op for an already-absolute path — kept here anyway so both
|
||||
// branches call the exact same guarantee, rather than one branch relying on
|
||||
// isAbsolute() alone to already imply what toAbsolutePath() enforces.
|
||||
return path.toAbsolutePath().normalize().toString();
|
||||
}
|
||||
String workspace = leadWorkspace.apply(leadTerminal);
|
||||
Path base = (workspace == null || workspace.isBlank())
|
||||
? Path.of(System.getProperty("user.dir"))
|
||||
: Path.of(workspace);
|
||||
return base.resolve(path).toAbsolutePath().normalize().toString();
|
||||
}
|
||||
|
||||
/**
|
||||
* Validate every gate, then — if and only if all of them pass — hand a one-shot continuation
|
||||
* that performs the actual roll to {@code continuationRunner} and return. <strong>This method
|
||||
* never itself sends anything to the lead's pane</strong> — see this class's javadoc for why
|
||||
* (it is called FROM the lead's own turn, so the pane cannot possibly be injectable yet).
|
||||
*
|
||||
* <p>Order: token lookup, then ownership ({@code callerTerminal} must match the terminal
|
||||
* {@link #open} recorded — {@link RefusalReason#NOT_YOUR_ROLLOVER}), then the
|
||||
* operator-confirmation gate, then the three handover-file checks (exists, not empty, fresh —
|
||||
* see {@link #checkHandover}). The first failing check is returned and {@code token} stays
|
||||
* pending (so a caller can fix the problem — e.g. rewrite the handover file — and retry with
|
||||
* the same token); it is consumed only once every gate passes and the continuation is launched.
|
||||
*
|
||||
* @param callerTerminal the CALLING lead's terminal id, resolved by the MCP layer from the
|
||||
* connection — never a client-supplied value (see this class's javadoc)
|
||||
* @param token the token {@link #open} returned
|
||||
* @param operatorConfirmed the caller's answer to "has an operator confirmed this roll" —
|
||||
* consulted only when the live config's {@code requireOperatorConfirm}
|
||||
* is true
|
||||
*/
|
||||
public RollDecision confirm(String callerTerminal, String token, boolean operatorConfirmed) {
|
||||
FleetConfig.LeadRollover cfg = configSupplier.get();
|
||||
if (cfg == null) {
|
||||
return RollDecision.refused(RefusalReason.NOT_CONFIGURED, "leadRollover: is not configured");
|
||||
}
|
||||
PendingRollover p = pending.get(token);
|
||||
if (p == null) {
|
||||
return RollDecision.refused(RefusalReason.UNKNOWN_TOKEN,
|
||||
"token " + token + " names no pending rollover request");
|
||||
}
|
||||
if (!p.leadTerminal().equals(callerTerminal)) {
|
||||
return RollDecision.refused(RefusalReason.NOT_YOUR_ROLLOVER,
|
||||
"token " + token + " was opened by a different lead terminal");
|
||||
}
|
||||
if (cfg.requireOperatorConfirm() && !operatorConfirmed) {
|
||||
return RollDecision.refused(RefusalReason.OPERATOR_NOT_CONFIRMED,
|
||||
"requireOperatorConfirm is true and operatorConfirmed was false");
|
||||
}
|
||||
|
||||
RollDecision docCheck = checkHandover(p, cfg);
|
||||
if (docCheck != null) {
|
||||
return docCheck;
|
||||
}
|
||||
|
||||
pending.remove(token);
|
||||
log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends",
|
||||
token, callerTerminal);
|
||||
continuationRunner.accept(() -> runRollover(p, cfg));
|
||||
return RollDecision.approved();
|
||||
}
|
||||
|
||||
/**
|
||||
* The single-shot continuation {@link #confirm} hands to {@code continuationRunner}. Runs
|
||||
* entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the
|
||||
* four-step order. There is no result to return to by this point, so every outcome is logged
|
||||
* only.
|
||||
*/
|
||||
private void runRollover(PendingRollover p, FleetConfig.LeadRollover cfg) {
|
||||
String lead = p.leadTerminal();
|
||||
boolean turnSettled = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds());
|
||||
if (!turnSettled) {
|
||||
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) within {}s "
|
||||
+ "after confirm() — refusing to send /clear at all; the calling lead's "
|
||||
+ "own turn is still live and clearing it now would destroy live context "
|
||||
+ "(token={})",
|
||||
lead, cfg.turnSettleSeconds(), p.token());
|
||||
return;
|
||||
}
|
||||
|
||||
// This deliberately bypasses Injector, exactly like ClaudeCodeLauncher#clearContext:
|
||||
// /clear is housekeeping, not a delegated turn, and routing it through Injector wedges the
|
||||
// pane forever (see this class's javadoc).
|
||||
agents.send(lead, "/clear");
|
||||
boolean clearSettled = waitForClearPickupAndSettle(lead, cfg.clearSettleSeconds());
|
||||
if (!clearSettled) {
|
||||
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) within {}s "
|
||||
+ "after /clear — NOT sending bootstrapText (token={})",
|
||||
lead, cfg.clearSettleSeconds(), p.token());
|
||||
return;
|
||||
}
|
||||
agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()));
|
||||
log.info("lead-rollover: rolled token={} lead={}", p.token(), lead);
|
||||
}
|
||||
|
||||
/** Drop a pending request without rolling. @return whether a pending request existed for {@code token} */
|
||||
public boolean cancel(String token) {
|
||||
return pending.remove(token) != null;
|
||||
}
|
||||
|
||||
/**
|
||||
* The three handover-file checks, in order: exists, not empty, fresh (modified after
|
||||
* {@link #open}'s timestamp and not older than {@code maxDocAgeSeconds}). Stats {@code
|
||||
* p.handoverPath()} directly — {@link #open} already resolved it to an absolute path, so this
|
||||
* never has to guess which directory it means.
|
||||
*
|
||||
* @return the first failing check's refusal, or {@code null} when all three pass
|
||||
*/
|
||||
private RollDecision checkHandover(PendingRollover p, FleetConfig.LeadRollover cfg) {
|
||||
Path path = Path.of(p.handoverPath());
|
||||
if (!Files.exists(path)) {
|
||||
return RollDecision.refused(RefusalReason.HANDOVER_MISSING,
|
||||
"handover file " + p.handoverPath() + " does not exist");
|
||||
}
|
||||
long size;
|
||||
long mtimeMillis;
|
||||
try {
|
||||
size = Files.size(path);
|
||||
mtimeMillis = Files.getLastModifiedTime(path).toMillis();
|
||||
} catch (IOException e) {
|
||||
throw new UncheckedIOException("failed to stat handover file " + p.handoverPath(), e);
|
||||
}
|
||||
if (size == 0) {
|
||||
return RollDecision.refused(RefusalReason.HANDOVER_EMPTY,
|
||||
"handover file " + p.handoverPath() + " is empty");
|
||||
}
|
||||
if (mtimeMillis <= p.requestedAtMillis()) {
|
||||
return RollDecision.refused(RefusalReason.HANDOVER_STALE,
|
||||
"handover file " + p.handoverPath() + " was not modified after the open() "
|
||||
+ "request (mtime=" + mtimeMillis + "ms, requestedAt=" + p.requestedAtMillis() + "ms)");
|
||||
}
|
||||
long ageMillis = nowMillis.getAsLong() - mtimeMillis;
|
||||
long maxAgeMillis = TimeUnit.SECONDS.toMillis(cfg.maxDocAgeSeconds());
|
||||
if (ageMillis > maxAgeMillis) {
|
||||
return RollDecision.refused(RefusalReason.HANDOVER_STALE,
|
||||
"handover file " + p.handoverPath() + " is " + TimeUnit.MILLISECONDS.toSeconds(ageMillis)
|
||||
+ "s old, older than maxDocAgeSeconds=" + cfg.maxDocAgeSeconds());
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Poll {@link AgentControl#status} until {@code target} reports a real turn boundary — {@link
|
||||
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used once by
|
||||
* {@link #runRollover}, to wait for the CALLING turn's own pane to settle before {@code /clear}
|
||||
* is ever sent at all — the {@code turnSettleSeconds} gate that makes this correction safe. The
|
||||
* SECOND wait, after {@code /clear}, is {@link #waitForClearPickupAndSettle} instead (fleetd
|
||||
* #489) — a plain boundary check is not enough there, because {@code /clear} starts no turn of
|
||||
* its own, so this method would (wrongly) report "settled" on its very first poll whether or not
|
||||
* {@code /clear} was actually picked up. A failed status read degrades to "not yet settled" and
|
||||
* is retried on the next poll, the same posture {@code LeadHeartbeatLoop} and {@code
|
||||
* HerdrPeerLauncher}'s readiness gate already take toward an unreadable status.
|
||||
*
|
||||
* <p><strong>Deliberately not {@link AgentStatus#injectable()}.</strong> {@code injectable()}
|
||||
* answers the {@code Injector}'s question — "may I deliver a message without stepping on a live
|
||||
* turn" — and it accepts {@link AgentStatus#BLOCKED} for that purpose, because a pane paused on
|
||||
* an approval prompt is safe to queue a message behind. This class asks a stricter question —
|
||||
* "has the turn actually ended" — and {@code BLOCKED} answers no: it is a live turn that is
|
||||
* merely paused, not one that has finished. Reusing {@code injectable()} here would let this
|
||||
* wait fire {@code /clear} while the lead's own {@code confirm()}-calling turn is still live and
|
||||
* paused on a prompt — exactly the live-context-destroying failure the {@code turnSettleSeconds}
|
||||
* gate exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link
|
||||
* #waitForClearPickupAndSettle} keeps the same exclusion of {@code BLOCKED}, for the same
|
||||
* reason, on the second wait.)
|
||||
*/
|
||||
private boolean waitUntilAtTurnBoundary(String target, int settleSeconds) {
|
||||
long deadline = nowMillis.getAsLong() + TimeUnit.SECONDS.toMillis(settleSeconds);
|
||||
while (nowMillis.getAsLong() < deadline) {
|
||||
AgentStatus status;
|
||||
try {
|
||||
status = agents.status(target);
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("lead-rollover: status check failed while waiting for {} to settle: {}",
|
||||
target, e.toString());
|
||||
status = null;
|
||||
}
|
||||
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
|
||||
return true;
|
||||
}
|
||||
settleSleeper.run();
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/**
|
||||
* The SECOND wait in {@link #runRollover} — after {@code /clear} has been sent, waits for it to
|
||||
* settle, bounded by {@code settleSeconds}. <strong>fleetd #489 — the paste-race fix.</strong>
|
||||
* {@code /clear} does not start a real turn of its own, so a pane with no submit race simply
|
||||
* stays {@link AgentStatus#IDLE} the whole time: {@link #waitUntilAtTurnBoundary} would (wrongly)
|
||||
* call that "settled" on its very first poll, whether or not the {@code /clear} Enter actually
|
||||
* landed. That was Fault 1, measured live on 2026-09-12 — the second gate was a no-op, so a
|
||||
* {@code bootstrapText} send followed immediately, racing Fault 2: {@link AgentControl#submit}'s
|
||||
* own javadoc already records that the submit accompanying a delivery "can race the paste —
|
||||
* especially right as the worker's TUI becomes interactive — leaving the text unsubmitted"
|
||||
* (CB-113). Because {@code runRollover} deliberately bypasses {@code Injector} for {@code
|
||||
* /clear} (see this class's javadoc), it inherited none of {@code Injector}'s nudging — so the
|
||||
* lost {@code /clear} Enter sat in the input box and {@code bootstrapText} was typed right after
|
||||
* it, landing as one concatenated line.
|
||||
*
|
||||
* <p>This method copies the pickup-nudge pattern {@link dev.ltms.fleet.inject.Injector} already
|
||||
* ships for exactly this, on its own post-turn {@code /clear} housekeeping (fleetd #306; see
|
||||
* {@code Injector.java:288-340} and {@code Injector.java:437-442}):
|
||||
* <ul>
|
||||
* <li>an {@link AgentStatus#WORKING} sample means {@code /clear} was picked up as a real
|
||||
* turn;</li>
|
||||
* <li>until that happens, each poll that still reports {@link AgentStatus#IDLE} or {@link
|
||||
* AgentStatus#DONE} re-sends the submit keystroke ({@link AgentControl#submit}) to nudge
|
||||
* the raced Enter — for the first {@code PICKUP_GRACE_POLLS - 1} of {@link
|
||||
* #PICKUP_GRACE_POLLS} consecutive such polls (i.e. {@code PICKUP_GRACE_POLLS - 1}
|
||||
* nudges: 7, not 8, given {@code PICKUP_GRACE_POLLS = 8}). A second Enter on an empty
|
||||
* Claude Code prompt is a no-op, so repeating it is safe;</li>
|
||||
* <li>the {@code PICKUP_GRACE_POLLS}th consecutive such poll, with {@code WORKING} still never
|
||||
* observed, releases rather than wedges the roll instead of nudging again — the same
|
||||
* choice {@code Injector} makes — and returns {@code true} anyway, logged at {@code info}
|
||||
* so an operator can see which path ran;</li>
|
||||
* <li>once {@code WORKING} has been observed, nudging stops and this instead waits for a real
|
||||
* {@code working → IDLE/DONE} completion boundary before returning {@code true}.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p><strong>{@link AgentStatus#BLOCKED} is deliberately excluded from both the nudge and the
|
||||
* boundary check</strong> — the same reasoning as {@link #waitUntilAtTurnBoundary}'s own
|
||||
* javadoc: a paused live turn is not a settled one, and re-sending Enter into an open approval
|
||||
* prompt could wrongly answer it. A {@code BLOCKED} sample (or an unreadable/{@link
|
||||
* AgentStatus#UNKNOWN} one) simply keeps this polling, with no nudge and no release, until either
|
||||
* a real boundary is reached or {@code settleSeconds} runs out.
|
||||
*
|
||||
* <p>{@link AgentControl#submit} can itself throw; a {@link RuntimeException} from it is
|
||||
* swallowed and logged at {@code debug}, exactly like {@code Injector.java:437-442} — a failed
|
||||
* nudge must not abort the roll.
|
||||
*
|
||||
* @return {@code true} once {@code /clear} has settled, or once the nudge budget was exhausted
|
||||
* with no pickup ever observed (released rather than wedged); {@code false} if {@code
|
||||
* settleSeconds} elapses first — the caller must NOT send {@code bootstrapText} in that
|
||||
* case, exactly as before this fix
|
||||
*/
|
||||
private boolean waitForClearPickupAndSettle(String target, int settleSeconds) {
|
||||
long deadline = nowMillis.getAsLong() + TimeUnit.SECONDS.toMillis(settleSeconds);
|
||||
boolean pickedUp = false; // a WORKING sample has been observed since /clear was sent
|
||||
int idlePollsAwaitingPickup = 0;
|
||||
while (nowMillis.getAsLong() < deadline) {
|
||||
AgentStatus status;
|
||||
try {
|
||||
status = agents.status(target);
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("lead-rollover: status check failed while waiting for {} to settle after "
|
||||
+ "/clear: {}", target, e.toString());
|
||||
status = null;
|
||||
}
|
||||
if (status == AgentStatus.WORKING) {
|
||||
pickedUp = true;
|
||||
} else if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
|
||||
if (pickedUp) {
|
||||
return true; // a real WORKING -> IDLE/DONE completion boundary
|
||||
}
|
||||
if (++idlePollsAwaitingPickup >= PICKUP_GRACE_POLLS) {
|
||||
log.info("lead-rollover: /clear on {} was never observed as WORKING after {} "
|
||||
+ "consecutive IDLE/DONE polls ({} of those were nudged) — "
|
||||
+ "releasing rather than wedging the roll",
|
||||
target, PICKUP_GRACE_POLLS, PICKUP_GRACE_POLLS - 1);
|
||||
return true;
|
||||
}
|
||||
try {
|
||||
agents.submit(target); // nudge a raced Enter (CB-113) so /clear actually submits
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("lead-rollover: resubmit to {} failed (will retry next poll): {}",
|
||||
target, e.getMessage());
|
||||
}
|
||||
}
|
||||
// AgentStatus.BLOCKED or UNKNOWN (or an unreadable status, above): neither a pickup
|
||||
// signal nor a boundary — keep polling without nudging or releasing.
|
||||
settleSleeper.run();
|
||||
}
|
||||
return false;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,81 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
/**
|
||||
* fleetd #469: a launch charter that names an MCP tool the server does not register must stop the
|
||||
* daemon at startup, not wait for a member to discover the gap by calling something that is not
|
||||
* there.
|
||||
*
|
||||
* <p>Deliberately its own class outside {@code dev.ltms.fleet.config}, not a case in {@link
|
||||
* dev.ltms.fleet.config.FleetConfig#validateCharters()}. The canonical tool surface ({@link
|
||||
* FleetTool}) lives in the {@code mcp} package; config is loaded before the MCP server exists and
|
||||
* must not gain a dependency on it. So this check belongs at the seam that already holds both a
|
||||
* loaded {@code FleetConfig} and the {@code mcp} package: {@code Fleetd.main}, called right after
|
||||
* {@code cfg.validateAll()} and before anything opens a socket or spawns a member.
|
||||
*
|
||||
* <p>{@code #464}'s {@code CharterToolSurfaceTest} proved the same comparison against a charter
|
||||
* fixture it wrote itself into a {@code @TempDir}, which meant nothing anyone wrote into the live
|
||||
* {@code fleetd.yaml} could ever fail it. This class is what a real charter is actually checked
|
||||
* against at boot; {@code FleetdStartupValidationTest} exercises it through {@code Fleetd.main}
|
||||
* itself, the same way it proves every other {@code validateXxx()} still runs there.
|
||||
*
|
||||
* <p><strong>fleetd #474</strong> — startup was not the only door: {@code
|
||||
* dev.ltms.fleet.config.ConfigRef#reload()} used to run {@code FleetConfig#validateAll()} alone,
|
||||
* which does not look at what a charter's text names, so a charter naming an unregistered tool
|
||||
* that could not have booted the daemon could still be installed into a running one through a
|
||||
* reload. This class still knows nothing about {@code ConfigRef} — {@code Fleetd.main} wires a
|
||||
* small {@code FleetConfig -> void} adapter over {@link #assertChartersNameOnlyRegisteredTools}
|
||||
* ({@code Fleetd::assertChartersNameOnlyRegisteredTools}) into {@code ConfigRef}'s constructor as
|
||||
* its {@code Consumer<FleetConfig>} {@code extraValidation}, run inside {@code reload()}'s same
|
||||
* try/catch as {@code validateAll()}, so both call sites — {@code Fleetd.main} at startup and
|
||||
* {@code ConfigRef#reload()} afterwards — go through this one method and can never check different
|
||||
* things. {@code dev.ltms.fleet.config.ConfigRefTest} and {@code
|
||||
* FleetdConfigRefCharterToolSurfaceWiringTest} are what prove the reload call site, the same way
|
||||
* {@code FleetdStartupValidationTest} proves the startup one.
|
||||
*/
|
||||
public final class CharterToolSurface {
|
||||
|
||||
/** A {@code fleet_…} (current) or {@code bridge_…} (pre-CB-634) tool-shaped token in prose. */
|
||||
private static final Pattern TOOL_REFERENCE = Pattern.compile("(fleet_[a-z_]+|bridge_[a-z_]+)");
|
||||
|
||||
private CharterToolSurface() {
|
||||
}
|
||||
|
||||
/**
|
||||
* @param charters the configured {@code fleet.charters:} map (role wire name → charter text);
|
||||
* {@code null} or empty is a no-op, same as an absent {@code fleet:} block
|
||||
* @throws IllegalStateException naming the charter key and every tool it names that {@link
|
||||
* FleetTool} does not list, when any charter does so
|
||||
*/
|
||||
public static void assertChartersNameOnlyRegisteredTools(Map<String, String> charters) {
|
||||
if (charters == null || charters.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
Set<String> registered = FleetTool.wireNames();
|
||||
List<String> bad = new ArrayList<>();
|
||||
charters.forEach((key, text) -> {
|
||||
if (text == null) {
|
||||
return;
|
||||
}
|
||||
Set<String> named = new LinkedHashSet<>();
|
||||
Matcher m = TOOL_REFERENCE.matcher(text);
|
||||
while (m.find()) {
|
||||
named.add(m.group(1));
|
||||
}
|
||||
named.stream()
|
||||
.filter(t -> !registered.contains(t))
|
||||
.forEach(unknown -> bad.add("fleet.charters." + key + " names '" + unknown
|
||||
+ "', which the server does not register (registered: " + registered + ")."));
|
||||
});
|
||||
if (!bad.isEmpty()) {
|
||||
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
|
||||
}
|
||||
}
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,83 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.Map;
|
||||
import java.util.Optional;
|
||||
import java.util.Set;
|
||||
|
||||
/**
|
||||
* The one canonical set of MCP tool names this daemon registers (fleetd #469, follow-up to #464).
|
||||
*
|
||||
* <p>Before this enum, the tool surface was written twice with nothing tying the copies together:
|
||||
* once as the literal {@code "fleet_…"} string passed to each tool-schema builder in
|
||||
* {@link FleetMcp}, and again as the case labels of {@link FleetMcp}'s authorization switch. A
|
||||
* reader that needed "what does this server register" — a charter check, in particular — had no
|
||||
* source to ask except scraping {@code FleetMcp.java}'s source text for {@code tool("…")} calls: a
|
||||
* third copy of the same list, and the weakest of the three forms.
|
||||
*
|
||||
* <p>Every reader that needs the registered tool surface now asks this enum instead:
|
||||
*
|
||||
* <ul>
|
||||
* <li>the tool-schema builders in {@code FleetMcp} pass {@code wireName()} rather than a literal;
|
||||
* <li>{@code FleetMcp}'s constructor asserts, at startup, that the set of tool names it actually
|
||||
* registers with the MCP SDK equals {@link #wireNames()} exactly — a canonical entry that is
|
||||
* never registered, or a registration with no canonical entry backing it, fails the daemon's
|
||||
* own boot rather than only a test's;
|
||||
* <li>{@code FleetMcp.toolAction}'s dispatch onto {@code Authz.Action} switches on the enum
|
||||
* (not the raw string) with no {@code default}, so adding a tool here without pinning its
|
||||
* action is a compile error, not a run-time throw;
|
||||
* <li>{@link CharterToolSurface} asks {@link #wireNames()} to check a configured launch charter
|
||||
* against the live tool surface, instead of scraping source text a third time.
|
||||
* </ul>
|
||||
*/
|
||||
public enum FleetTool {
|
||||
|
||||
SEND("fleet_send"),
|
||||
REPLY("fleet_reply"),
|
||||
ASK("fleet_ask"),
|
||||
STATUS("fleet_status"),
|
||||
POLL("fleet_poll"),
|
||||
ACK("fleet_ack"),
|
||||
SPAWN("fleet_spawn"),
|
||||
LIST("fleet_list"),
|
||||
STOP("fleet_stop"),
|
||||
PROFILES("fleet_profiles"),
|
||||
WHOAMI("fleet_whoami"),
|
||||
HANDOVER("fleet_handover");
|
||||
|
||||
private final String wireName;
|
||||
|
||||
FleetTool(String wireName) {
|
||||
this.wireName = wireName;
|
||||
}
|
||||
|
||||
/** The name this tool is registered under, and called by, on the wire ({@code "fleet_send"}, …). */
|
||||
public String wireName() {
|
||||
return wireName;
|
||||
}
|
||||
|
||||
private static final Map<String, FleetTool> BY_WIRE_NAME;
|
||||
private static final Set<String> WIRE_NAMES;
|
||||
|
||||
static {
|
||||
Map<String, FleetTool> byName = new LinkedHashMap<>();
|
||||
Set<String> names = new LinkedHashSet<>();
|
||||
for (FleetTool tool : values()) {
|
||||
byName.put(tool.wireName, tool);
|
||||
names.add(tool.wireName);
|
||||
}
|
||||
BY_WIRE_NAME = Map.copyOf(byName);
|
||||
WIRE_NAMES = Set.copyOf(names);
|
||||
}
|
||||
|
||||
/** The tool named {@code wireName}, or empty when this daemon registers no such tool. */
|
||||
public static Optional<FleetTool> byWireName(String wireName) {
|
||||
return Optional.ofNullable(BY_WIRE_NAME.get(wireName));
|
||||
}
|
||||
|
||||
/** Every wire name this daemon registers — the canonical tool surface. */
|
||||
public static Set<String> wireNames() {
|
||||
return WIRE_NAMES;
|
||||
}
|
||||
}
|
||||
@@ -14,6 +14,7 @@ import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementCandidate;
|
||||
import dev.ltms.fleet.placement.PlacementContext;
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
import dev.ltms.fleet.placement.PlacementException;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import dev.ltms.fleet.placement.PlacementPolicy;
|
||||
@@ -90,6 +91,19 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
private final Supplier<Map<String, FleetConfig.Profile>> profileConfigs;
|
||||
private final Supplier<PlacementPolicy> placementPolicy;
|
||||
|
||||
/**
|
||||
* fleetd #422: the central model allow-list's on/off state, read live per spawn — same reason
|
||||
* {@link #profileConfigs} is a supplier rather than a captured map (see the class doc above and
|
||||
* {@link #enforceModelEnabled}). A caller with no {@code models:} block to read from (the
|
||||
* simpler, map-based constructors used throughout this class's own tests) wires this to a
|
||||
* constant {@code null}, which {@link #models0} treats as "nothing configured, gate never
|
||||
* fires" — the pre-#422 behaviour.
|
||||
*/
|
||||
private final Supplier<FleetConfig.Models> models;
|
||||
|
||||
/** The value {@link #models0} normalizes a {@code null} supplier result to. */
|
||||
private static final FleetConfig.Models NO_MODELS_CONFIGURED = new FleetConfig.Models(List.of());
|
||||
|
||||
/** CB-578 stage B: credential cooldown, checked before an explicit spawn and filtered into placement. */
|
||||
private final BackendQuarantine quarantine;
|
||||
|
||||
@@ -205,12 +219,44 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
FleetConfig.Fleet fleet,
|
||||
BackendQuarantine quarantine,
|
||||
BackendOutagePolicy outagePolicy) {
|
||||
// No models: block to read from a plain profiles Map — fleetd #422's gate is wired to a
|
||||
// constant null (see enforceModelEnabled/models0), the pre-#422 behaviour for every caller
|
||||
// of this overload. Use the 9-arg overload below to test the gate against a LIVE supplier.
|
||||
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet, quarantine,
|
||||
outagePolicy, constant(null));
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, plus a LIVE model-gate source (fleetd #422). The 8-arg overload above wires
|
||||
* {@code models} to a constant {@code null} because it has only a static {@code Map<String,
|
||||
* Profile>}, never a full {@code FleetConfig}, to read one from; this overload exists so a test
|
||||
* can prove {@link #enforceModelEnabled} and its candidate filter re-read {@link
|
||||
* FleetConfig.Models} on every call rather than a value captured once at construction — the
|
||||
* exact distinction fleetd #422 exists to get right (see the class doc's CB-559 note on {@link
|
||||
* #profileConfigs}, which this follows). Production wiring uses the
|
||||
* {@code Supplier<FleetConfig>} constructor below instead, which already threads a live
|
||||
* {@code config.get().models()} through.
|
||||
*
|
||||
* @param models required — pass a supplier returning {@code null} for a caller that has no
|
||||
* {@code models:} block to gate against, never a defaulting overload (the same
|
||||
* "explicit opt-out, never a silent default" rule {@code quarantine}/
|
||||
* {@code outagePolicy} already follow).
|
||||
*/
|
||||
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
|
||||
String defaultProfile,
|
||||
Map<String, FleetConfig.Profile> profileConfigs,
|
||||
PlacementPolicy placementPolicy,
|
||||
Function<String, Integer> liveCount,
|
||||
FleetConfig.Fleet fleet,
|
||||
BackendQuarantine quarantine,
|
||||
BackendOutagePolicy outagePolicy,
|
||||
Supplier<FleetConfig.Models> models) {
|
||||
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
|
||||
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
|
||||
// differ from one JVM run to the next.
|
||||
this(delegates, defaultProfile,
|
||||
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
|
||||
constant(placementPolicy), liveCount, constant(fleet), quarantine, outagePolicy);
|
||||
constant(placementPolicy), liveCount, constant(fleet), quarantine, outagePolicy, models);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -250,7 +296,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
liveCount,
|
||||
() -> config.get().fleet(),
|
||||
quarantine,
|
||||
outagePolicy);
|
||||
outagePolicy,
|
||||
// fleetd #422: read live, same as profiles/placement/fleet above — a models.allow
|
||||
// edit (on/off or otherwise) is visible to the very next spawn, no restart needed.
|
||||
() -> config.get().models());
|
||||
}
|
||||
|
||||
/** The all-suppliers form every other constructor funnels into. */
|
||||
@@ -261,7 +310,8 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
Function<String, Integer> liveCount,
|
||||
Supplier<FleetConfig.Fleet> fleet,
|
||||
BackendQuarantine quarantine,
|
||||
BackendOutagePolicy outagePolicy) {
|
||||
BackendOutagePolicy outagePolicy,
|
||||
Supplier<FleetConfig.Models> models) {
|
||||
this.fleet = fleet;
|
||||
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
|
||||
this.outagePolicy = Objects.requireNonNull(outagePolicy, "outagePolicy");
|
||||
@@ -273,6 +323,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
this.profileConfigs = profileConfigs;
|
||||
this.placementPolicy = placementPolicy;
|
||||
this.liveCount = liveCount;
|
||||
this.models = Objects.requireNonNull(models, "models");
|
||||
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
|
||||
for (HerdrPeerLauncher d : this.delegates) {
|
||||
for (String profile : d.profiles()) {
|
||||
@@ -303,6 +354,16 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
return m == null ? Map.of() : m;
|
||||
}
|
||||
|
||||
/**
|
||||
* The currently-configured {@code models:} block, never null (fleetd #422). Read fresh on
|
||||
* every call, the same reason {@link #profiles0} is — a config reload's on/off edit must reach
|
||||
* the very next spawn.
|
||||
*/
|
||||
private FleetConfig.Models models0() {
|
||||
FleetConfig.Models m = models.get();
|
||||
return m == null ? NO_MODELS_CONFIGURED : m;
|
||||
}
|
||||
|
||||
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
|
||||
private HerdrPeerLauncher route(String profileName) {
|
||||
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
|
||||
@@ -332,6 +393,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
enforceNotQuarantined(requestedProfile);
|
||||
enforceNotCoolingOff(requestedProfile);
|
||||
enforceMaxLoad(requestedProfile);
|
||||
enforceModelEnabled(requestedProfile);
|
||||
PeerHandle handle = d.spawn(req);
|
||||
spawnedBy.put(handle.id(), d);
|
||||
return handle;
|
||||
@@ -341,18 +403,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
// the whole profile list. An EXPLICIT profile (above) is left alone on purpose — it is the
|
||||
// operator overriding, and refusing it would break `fleet_spawn{profile:"opus"}`, which
|
||||
// carries no role and so would be judged against the dev pool it was never meant for.
|
||||
List<PlacementCandidate> candidates = candidates(req.role());
|
||||
String roleDefault = defaultProfileFor(req.role());
|
||||
Set<String> unreachable = new HashSet<>();
|
||||
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
|
||||
// call, so re-deriving it per retry would only cost work, never change the answer.
|
||||
Set<String> quarantined = quarantinedProfiles(candidates);
|
||||
// fleetd #201 Unit 5: a distinct set from quarantined — see PlacementContext.coolingOff.
|
||||
Set<String> coolingOff = coolingOffProfiles(candidates);
|
||||
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable,
|
||||
quarantined, coolingOff);
|
||||
PlacementContext ctx = placementContextFor(req.role(), unreachable);
|
||||
|
||||
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
|
||||
int maxAttempts = ctx.candidates().isEmpty() ? 1 : ctx.candidates().size();
|
||||
for (int attempt = 0; attempt < maxAttempts; attempt++) {
|
||||
// Deliberately uncaught: when no candidate is left (all at cap, or all unreachable) the
|
||||
// policy already throws a clear message. Catching it to rethrow a generic
|
||||
@@ -379,8 +433,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
chosen.profile(), e.getMessage());
|
||||
unreachable.add(chosen.profile());
|
||||
// Update the context for the next selection so the policy excludes this profile.
|
||||
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable,
|
||||
quarantined, coolingOff);
|
||||
ctx = placementContextFor(req.role(), unreachable);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -492,6 +545,49 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Refuse an explicit-profile spawn whose {@code model:} the operator has turned off in the
|
||||
* central {@code models.allow:} list (fleetd #422). Deliberately worded apart from {@link
|
||||
* #enforceNotQuarantined} and {@link #enforceNotCoolingOff}: those two report a BACKEND-reported
|
||||
* outage (exhaustion, repeated errors); this one reports an OPERATOR decision, so the message
|
||||
* says "turned off" and names the model, never "quarantined" or "cooling off". A FOURTH,
|
||||
* independent reason to refuse a spawn — never layered onto {@code BackendQuarantine} or {@code
|
||||
* BackendOutagePolicy}, which would misattribute an operator's own choice to the backend.
|
||||
*
|
||||
* <p>Reads {@link #models0()} fresh on every call — the same liveness {@link #profiles0()}
|
||||
* already has — so flipping {@code enabled: false} and reloading takes effect on the very next
|
||||
* spawn, no restart (criterion 4). A profile naming no model, or a model absent from {@code
|
||||
* models.allow:} entirely (nothing to gate against), is never refused here.
|
||||
*
|
||||
* @throws PlacementException naming the model and the profile, distinct from quarantine/cool-off
|
||||
*/
|
||||
private void enforceModelEnabled(String profile) {
|
||||
FleetConfig.Profile cfg = profiles0().get(profile);
|
||||
String model = (cfg == null) ? null : cfg.model();
|
||||
if (model == null) {
|
||||
return;
|
||||
}
|
||||
if (models0().offIds().contains(model)) {
|
||||
throw new PlacementException("worker profile '" + profile + "' names model '" + model
|
||||
+ "', which the operator has turned off in models.allow — refusing spawn");
|
||||
}
|
||||
}
|
||||
|
||||
/** The subset of {@code candidates} whose {@code model:} is currently turned off (fleetd #422). */
|
||||
private Set<String> modelOffProfiles(List<PlacementCandidate> candidates) {
|
||||
Set<String> off = models0().offIds();
|
||||
if (off.isEmpty()) {
|
||||
return Set.of();
|
||||
}
|
||||
return candidates.stream()
|
||||
.map(PlacementCandidate::profile)
|
||||
.filter(p -> {
|
||||
FleetConfig.Profile cfg = profiles0().get(p);
|
||||
return cfg != null && cfg.model() != null && off.contains(cfg.model());
|
||||
})
|
||||
.collect(Collectors.toSet());
|
||||
}
|
||||
|
||||
/**
|
||||
* The profile names {@code role} may be placed on, in definition order.
|
||||
*
|
||||
@@ -508,8 +604,23 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
return known.isEmpty() ? List.copyOf(configured.keySet()) : known;
|
||||
}
|
||||
|
||||
/** The profile an unqualified spawn for {@code role} falls back to under {@code fixed} placement. */
|
||||
private String defaultProfileFor(MemberRole role) {
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>Live: reads {@link #poolFor}, which reads {@link #profileConfigs} and {@link #fleet} fresh
|
||||
* on every call, so a config reload is visible without a restart (fleetd #425) — unlike {@link
|
||||
* #defaultProfile}, the field captured once at construction, which this falls back to only when
|
||||
* {@link #poolFor} has nothing to offer at all (no profiles configured for this composite).
|
||||
*
|
||||
* <p>Exact only under the {@code fixed} placement policy — the one that reads this value
|
||||
* ({@code FixedPlacementPolicy}, package-private, hence not linked) as its first, preferred
|
||||
* candidate. {@code weighted}/{@code round-robin} placement can choose a different candidate
|
||||
* from {@code role}'s pool even on the very first spawn; this method does not simulate that
|
||||
* choice, matching what the {@code defaultProfile:}-derived reporting this replaces has always
|
||||
* done.
|
||||
*/
|
||||
@Override
|
||||
public String defaultProfileFor(MemberRole role) {
|
||||
List<String> pool = poolFor(role);
|
||||
return pool.isEmpty() ? defaultProfile : pool.getFirst();
|
||||
}
|
||||
@@ -526,6 +637,142 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
return out;
|
||||
}
|
||||
|
||||
/**
|
||||
* Build the {@link PlacementContext} an unqualified spawn of {@code role} would be judged
|
||||
* against right now — the single source both {@link #spawn} and {@link #place} read, so the two
|
||||
* can never disagree about which conditions (quarantine, cool-off, model-off) apply to which
|
||||
* candidate (fleetd #425 rework: round 1 duplicated this into a second, blind resolver —
|
||||
* {@link #defaultProfileFor} — which is why it regressed; round 2 found that even a single
|
||||
* shared resolver is not enough on its own if the CALLER re-resolves through an explicit
|
||||
* profile afterwards — see {@link PlacementDecision}).
|
||||
*
|
||||
* @param unreachable the caller's mutable unreachable set; {@link #spawn} grows this across
|
||||
* retries and rebuilds the context from it, {@link #place} passes a fresh
|
||||
* empty one since it never retries
|
||||
*/
|
||||
private PlacementContext placementContextFor(MemberRole role, Set<String> unreachable) {
|
||||
List<PlacementCandidate> candidates = candidates(role);
|
||||
String roleDefault = defaultProfileFor(role);
|
||||
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
|
||||
// call, so re-deriving it per retry would only cost work, never change the answer.
|
||||
Set<String> quarantined = quarantinedProfiles(candidates);
|
||||
// fleetd #201 Unit 5: a distinct set from quarantined — see PlacementContext.coolingOff.
|
||||
Set<String> coolingOff = coolingOffProfiles(candidates);
|
||||
// fleetd #422: read live per spawn, same as quarantined/coolingOff above — a config reload
|
||||
// that flips a model's enabled state is visible to the very next unqualified spawn.
|
||||
Set<String> modelOff = modelOffProfiles(candidates);
|
||||
return new PlacementContext(roleDefault, candidates, liveCount, unreachable,
|
||||
quarantined, coolingOff, modelOff);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>fleetd #425 rework, round 2: runs the exact same selection {@link #spawn} uses for a
|
||||
* blank-profile request — {@link #placementContextFor} plus one {@link PlacementPolicy#select}
|
||||
* — rather than {@link #defaultProfileFor}'s blind "pool's first entry", so a quarantined,
|
||||
* cooling-off, or model-off pool-first candidate is routed around here exactly as it would be
|
||||
* by a real spawn. Unlike {@link #spawn}, this never retries on {@link
|
||||
* PeerUnreachableException}: there is no spawn attempt to fail, so "unreachable" never grows
|
||||
* past the empty set it starts with, and a single {@link PlacementPolicy#select} call already
|
||||
* reflects the live quarantine/cool-off/model-off state.
|
||||
*
|
||||
* <p>Deliberately does <em>not</em> apply {@link #enforceMaxLoad} (or any of the other three
|
||||
* {@code enforce*} checks): those belong to {@link #spawn}'s EXPLICIT-profile branch, the
|
||||
* operator-override path, and this method answers a different question — "where would an
|
||||
* UNQUALIFIED spawn land". That is not the same as {@code select} ignoring these conditions —
|
||||
* every condition {@code select} filters on (quarantine, cooling off, {@code maxLoad} under
|
||||
* every placement policy including the default {@code fixed}, since fleetd #435, model-off,
|
||||
* unreachable, weight-0) is already reflected in the {@link PlacementDecision} this method
|
||||
* returns, because {@code select} walked past every excluded candidate to find it. What this
|
||||
* method's caller must not do is take that resolved name and hand it back to {@link
|
||||
* #spawn(SpawnRequest)} as an explicit profile: the explicit-profile branch treats the same
|
||||
* exclusion conditions as a reason to REFUSE, where {@code select} had already treated them as
|
||||
* a reason to fall through — round 1 of this fix did exactly that, turning a fall-through this
|
||||
* method had already resolved around into a refusal one call later. Round 2 fixes that at the
|
||||
* caller: {@link #spawn(SpawnRequest, PlacementDecision)} carries this exact decision to the
|
||||
* spawn without re-resolving or re-checking it, through the same routing path {@code select}
|
||||
* itself was consulted from.
|
||||
*
|
||||
* @throws PlacementException if no candidate in {@code role}'s pool is currently placeable
|
||||
* (mirrors what an actual unqualified spawn would throw)
|
||||
*/
|
||||
@Override
|
||||
public PlacementDecision place(MemberRole role) {
|
||||
PlacementContext ctx = placementContextFor(role, new HashSet<>());
|
||||
return new PlacementDecision(placementPolicy.get().select(ctx).profile());
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>Delegates to {@link #place}, so the two can never disagree about the answer for the same
|
||||
* {@code role} at the same instant — kept as a convenience for a caller that only wants the
|
||||
* resolved name (a status report, a log line), never for a caller that will act on it by
|
||||
* spawning: that caller must hold the {@link PlacementDecision} itself and pass it to {@link
|
||||
* #spawn(SpawnRequest, PlacementDecision)} — see {@link PlacementDecision}'s javadoc for why
|
||||
* resolving here and spawning separately, with the name fed back in as an explicit profile,
|
||||
* regressed fleetd #425 twice.
|
||||
*/
|
||||
@Override
|
||||
public String routedProfileFor(MemberRole role) {
|
||||
return place(role).profile();
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>Routes {@code decision.profile()} directly to its owning delegate — the identical
|
||||
* {@code d.spawn(routedReq)} call {@link #spawn(SpawnRequest)}'s blank-profile branch makes for
|
||||
* its first pick — WITHOUT re-running {@link #enforceNotQuarantined}, {@link
|
||||
* #enforceNotCoolingOff}, {@link #enforceMaxLoad}, or {@link #enforceModelEnabled}: those are
|
||||
* the EXPLICIT-profile branch's checks, and {@code decision} did not come from an operator
|
||||
* naming a profile — it came from {@link #place}, which already applied whichever of these
|
||||
* conditions {@link PlacementPolicy#select} actually filters on (fleetd #425 rework, round 2).
|
||||
*
|
||||
* <p>The two branches disagree on purpose about what an excluded profile means, and that
|
||||
* disagreement is not what this method removes. The blank-profile routing branch (and
|
||||
* {@link #place}) treats a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0 profile
|
||||
* as a reason to fall through to the next candidate; the EXPLICIT-profile branch treats naming
|
||||
* that same profile as a reason to refuse outright — someone who names a profile should get a
|
||||
* refusal, not a silent substitution onto a different backend. That is still correct after
|
||||
* fleetd #435. What round 1 got wrong, and what this method exists to stop happening again, is
|
||||
* turning a fall-through into a refusal by accident: resolving a name via {@link #place} and
|
||||
* then handing that same name back to {@link #spawn(SpawnRequest)} as an explicit profile takes
|
||||
* the refusing branch on a decision the routing branch had already approved by falling through
|
||||
* past everything else.
|
||||
*
|
||||
* <p>Before fleetd #435, this exact accident was reachable through {@code maxLoad} specifically:
|
||||
* {@code FixedPlacementPolicy} — the default policy — did not evaluate {@code maxLoad} at all
|
||||
* for automatic selection, so {@link #place} could approve an at-cap profile that {@link
|
||||
* #enforceMaxLoad} would then refuse one call later. fleetd #435 closed that: {@code
|
||||
* FixedPlacementPolicy} now walks past an at-cap candidate exactly like {@code weighted}/
|
||||
* {@code round-robin} already did, so {@link #place} can no longer return one, and this specific
|
||||
* failure — an approved placement dying at {@code enforceMaxLoad} — cannot happen any more.
|
||||
* What this method still buys, now that {@code maxLoad} can no longer cause it: it never
|
||||
* re-evaluates a condition {@link #place} already decided, and it closes the window between
|
||||
* that decision and the spawn in which the underlying state (another spawn landing on the same
|
||||
* profile, a config reload) could otherwise move and make a stale explicit re-check wrong.
|
||||
*
|
||||
* <p>Deliberately does not retry on {@link PeerUnreachableException} across candidates the way
|
||||
* {@link #spawn(SpawnRequest)}'s blank-profile branch does: retrying here would silently
|
||||
* re-place the caller onto a different profile than the one {@code decision} named, behind the
|
||||
* back of a caller that may already have provisioned something (a worktree's {@code repoRoot},
|
||||
* parity overlay) specifically for that name. A caller that wants the composite's own failover
|
||||
* should call {@link #spawn(SpawnRequest)} with a blank profile directly, not resolve through
|
||||
* {@link #place} first. Losing that retry on a resolve-then-spawn path is an accepted, unrelated
|
||||
* cost — see {@code SessionManager.acquireWithWorktree}'s own comment on it — never widened by
|
||||
* this round to include {@code maxLoad}, which is what round 1 actually lost.
|
||||
*/
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
HerdrPeerLauncher d = route(decision.profile());
|
||||
SpawnRequest routedReq = req.withProfile(decision.profile());
|
||||
PeerHandle handle = d.spawn(routedReq);
|
||||
spawnedBy.put(handle.id(), d);
|
||||
return handle;
|
||||
}
|
||||
|
||||
@Override
|
||||
public String effectiveCwd(SpawnRequest req) {
|
||||
return route(req.profileName()).effectiveCwd(req);
|
||||
@@ -536,6 +783,32 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
return route(profileName).parityOverlay(profileName);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up: the single live read that answers both "is the models.allow: gate
|
||||
* armed" and "which models are off", off the exact same accessor ({@link #models0()}) {@link
|
||||
* #enforceModelEnabled} and {@link #modelOffProfiles} read — so {@code fleet_profiles}/{@code
|
||||
* GET /profiles} (via {@link PeerLauncher#disabledModels()}, which now delegates here) can
|
||||
* never report a different answer than the gate enforces (the fleetd #404 lesson), and the
|
||||
* startup log line built from this can never disagree with either.
|
||||
*
|
||||
* <p>{@link #models0()} itself normalizes a {@code null} {@link #models} read to the shared
|
||||
* {@link #NO_MODELS_CONFIGURED} sentinel — deliberately the one object no config-supplied
|
||||
* {@code Models} instance can ever be identical to, since it is private to this class — so
|
||||
* comparing by reference here recovers exactly the fact {@code models0()}'s normalization
|
||||
* would otherwise erase: whether the live source was {@code null} (no {@code models:} block,
|
||||
* armed = false) or a real, config-supplied block (armed = true, even one whose {@code allow:}
|
||||
* is itself empty or absent — {@link FleetConfig.Models}'s "absent or empty allow: is off"
|
||||
* wording governs config-load validation, a distinct question from whether this gate is armed
|
||||
* for reporting).
|
||||
*/
|
||||
@Override
|
||||
public PeerLauncher.ModelGateState modelGateState() {
|
||||
FleetConfig.Models m = models0();
|
||||
return m == NO_MODELS_CONFIGURED
|
||||
? PeerLauncher.ModelGateState.notConfigured()
|
||||
: PeerLauncher.ModelGateState.armed(m.offIds());
|
||||
}
|
||||
|
||||
@Override
|
||||
public void stop(String id) {
|
||||
HerdrPeerLauncher d = spawnedBy.get(id);
|
||||
@@ -647,9 +920,21 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
return byProfile.keySet();
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>fleetd #425: reports the <em>live</em> {@code dev} pool's first entry — the same value
|
||||
* {@link #defaultProfileFor} computes for {@link MemberRole#DEV} — not the {@link
|
||||
* #defaultProfile} field captured at construction. An unqualified {@code fleet_spawn} defaults
|
||||
* to {@code MemberRole#DEV} (see {@link dev.ltms.fleet.peer.SpawnRequest}), so "the dev pool's
|
||||
* live first entry" is exactly the profile such a spawn actually lands on right now — the
|
||||
* question {@code fleet_profiles}' {@code "default"} field exists to answer. The frozen field is
|
||||
* a role-agnostic fallback used only when {@link #poolFor} has nothing to report at all (no
|
||||
* profiles configured), which {@link #defaultProfileFor} already handles.
|
||||
*/
|
||||
@Override
|
||||
public String defaultProfile() {
|
||||
return defaultProfile;
|
||||
return defaultProfileFor(MemberRole.DEV);
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -75,9 +75,39 @@ import java.util.stream.Stream;
|
||||
* never to the control.
|
||||
*
|
||||
* <p>The scrub also writes {@code scrub-report.txt} into its own directory: one {@code allowed N of
|
||||
* M} line (N = exports left untouched, M = exports present when the scrub ran), then the blanked
|
||||
* NAMES — never values. The launcher reads this back at teardown and logs it, because a blocked
|
||||
* count next to an unknown denominator is not a finding.
|
||||
* M failed F} line (N = exports left untouched, M = exports present when the scrub ran, F = names
|
||||
* the scrub attempted to blank but could not), then the NAMES — blanked ones bare, unblankable ones
|
||||
* {@code !}-prefixed — never values. The launcher reads this back at teardown and logs it, because a
|
||||
* blocked count next to an unknown denominator is not a finding.
|
||||
*
|
||||
* <p><b>fleetd #394:</b> plain {@code export "$n="} is a FATAL error for a zsh read-only or special
|
||||
* parameter (for example {@code UID}) — it aborts the whole sourced file, so every name still to
|
||||
* come is never blanked and the report above is never written at all. The blanking loop instead
|
||||
* routes each attempt through {@code eval}, which contains that error to the single iteration: the
|
||||
* loop always finishes, and a name that could not be blanked is counted as {@code failed} and
|
||||
* listed {@code !}-prefixed rather than silently disappearing. This is deliberately not a skip-list
|
||||
* of known-bad names — every enumerated name is still attempted, so a name nobody has thought of
|
||||
* yet still gets tried and, if it fails, still gets counted.
|
||||
*
|
||||
* <p>The blanking loop also re-asserts, on its own, the same {@code [A-Za-z_][A-Za-z0-9_]*} shape
|
||||
* check the enumeration loop already applied. Before {@code eval} was introduced a non-conforming
|
||||
* name reaching {@code export "$n="} was harmless either way — the quoting made it inert. With
|
||||
* {@code eval}, the name is spliced into a string and interpreted as shell syntax, so the enumeration
|
||||
* loop's check is no longer sufficient on its own to keep that call site safe — it is a guard on a
|
||||
* different loop, and the two must not silently drift apart. Re-checking right before the
|
||||
* {@code eval} keeps that call site safe by its own reading, independent of whatever the enumeration
|
||||
* loop does or stops doing in a later change.
|
||||
*
|
||||
* <p><b>fleetd #400:</b> {@code eval}'s exit status is not proof that the blank actually happened.
|
||||
* zsh coerces a bare {@code NAME=} assignment on an integer special parameter (measured on macOS zsh
|
||||
* 5.9: {@code SECONDS}, {@code RANDOM}, {@code SHLVL}, {@code HISTSIZE}, {@code COLUMNS},
|
||||
* {@code LINES}, {@code USERNAME}) to a number instead of failing — {@code eval} returns success,
|
||||
* the value is untouched, and a status-based classification reports it as blanked when it was not.
|
||||
* The fix classifies on the observed effect instead: after the attempt, the name's value is read
|
||||
* back with the {@code (P)} indirection flag and the decision is made from whether that is now
|
||||
* empty. This one check covers all three shapes a name can take at this point — a genuine blank, a
|
||||
* fatal read-only error {@code eval} merely contained, and this silent no-op — and the exit status
|
||||
* plays no part in the decision at all.
|
||||
*/
|
||||
public final class EnvAllowListScrub {
|
||||
|
||||
@@ -132,10 +162,16 @@ public final class EnvAllowListScrub {
|
||||
}
|
||||
|
||||
/**
|
||||
* A parsed {@code scrub-report.txt}: how many exported variables existed when the scrub ran,
|
||||
* how many were left untouched (allowed), and the NAMES that were blanked. Values never appear.
|
||||
* A parsed {@code scrub-report.txt}: how many exported variables existed when the scrub ran
|
||||
* ({@code total}), how many were left untouched ({@code allowed}), how many the scrub attempted
|
||||
* to blank but could not ({@code failed} — fleetd #394: a zsh read-only/special parameter such
|
||||
* as {@code UID} fatally errors on plain {@code export NAME=}, so those attempts go through
|
||||
* {@code eval} instead so the loop keeps going and the failure is counted rather than left
|
||||
* invisible), and the NAMES in each of the latter two categories. {@code allowed +
|
||||
* blanked.size() + unblankable.size() == total}, and {@code unblankable.size() == failed}.
|
||||
* Values never appear.
|
||||
*/
|
||||
record ScrubReport(int allowed, int total, List<String> blanked) {
|
||||
record ScrubReport(int allowed, int total, int failed, List<String> blanked, List<String> unblankable) {
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -334,15 +370,57 @@ public final class EnvAllowListScrub {
|
||||
_cb633_blank+=("$_cb633_n")
|
||||
done
|
||||
|
||||
{ for _cb633_n in "${_cb633_blank[@]}"; do export "$_cb633_n="; done; } 2>/dev/null
|
||||
# fleetd #394: plain `export "$n="` is FATAL for a zsh read-only/special parameter
|
||||
# (e.g. UID) and aborts this whole sourced file — every name still to come is never
|
||||
# blanked, and the report below is never written, silently. `eval` contains that
|
||||
# error to the single iteration instead: it still fails for that one name, but the
|
||||
# loop continues and we can tell allowed / blanked / unblankable apart afterwards.
|
||||
# This is not a skip-list of known-bad names (that would miss the next one nobody
|
||||
# thought of) — every name in _cb633_blank is still attempted, unconditionally.
|
||||
# Every name reaching this loop already passed the identical identifier check in the
|
||||
# enumeration loop above — but that guard is 20 lines away in a different loop, and
|
||||
# this line is about to splice the name into a string handed to `eval`. Before this
|
||||
# fix the name only ever reached `export` quoted ("$n="), which is inert on a
|
||||
# non-identifier string either way; `eval` makes THIS line the only thing standing
|
||||
# between such a string and code execution in the member's pane, so it re-asserts the
|
||||
# same check on its own rather than trusting a guard it does not own. Under normal
|
||||
# operation this can never fire (the enumeration guard already filtered everything
|
||||
# reaching _cb633_blank), so a name caught here is counted as unblankable rather than
|
||||
# silently dropped — it is real evidence that the upstream guard was bypassed.
|
||||
#
|
||||
# fleetd #400: the attempt's own exit status is NOT proof of its effect. zsh coerces
|
||||
# a bare `NAME=` assignment on an integer special parameter (SECONDS, RANDOM, SHLVL,
|
||||
# HISTSIZE, COLUMNS, LINES, USERNAME on this host) to a number instead of failing —
|
||||
# `eval` returns 0, the value is untouched, and the old exit-status check reported it
|
||||
# as blanked when it was not. Classify on the observed effect instead: attempt the
|
||||
# export, then read the name's value back with the `(P)` indirection flag and decide
|
||||
# from whether it is now empty. One check then covers all three shapes a name can
|
||||
# take here — a genuine blank, a fatal read-only error `eval` merely contained, and
|
||||
# this silent no-op — without the exit status entering the decision at all.
|
||||
typeset -a _cb633_ok _cb633_unblankable
|
||||
_cb633_ok=()
|
||||
_cb633_unblankable=()
|
||||
for _cb633_n in "${_cb633_blank[@]}"; do
|
||||
if [[ ! "$_cb633_n" =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
|
||||
_cb633_unblankable+=("$_cb633_n")
|
||||
continue
|
||||
fi
|
||||
eval "export ${_cb633_n}=" 2>/dev/null
|
||||
if [[ -z "${(P)_cb633_n}" ]]; then
|
||||
_cb633_ok+=("$_cb633_n")
|
||||
else
|
||||
_cb633_unblankable+=("$_cb633_n")
|
||||
fi
|
||||
done
|
||||
|
||||
integer _cb633_kept=$(( _cb633_total - ${#_cb633_blank} ))
|
||||
{
|
||||
print -r -- "allowed $_cb633_kept of $_cb633_total"
|
||||
for _cb633_n in "${_cb633_blank[@]}"; do print -r -- "$_cb633_n"; done
|
||||
print -r -- "allowed $_cb633_kept of $_cb633_total failed ${#_cb633_unblankable}"
|
||||
for _cb633_n in "${_cb633_ok[@]}"; do print -r -- "$_cb633_n"; done
|
||||
for _cb633_n in "${_cb633_unblankable[@]}"; do print -r -- "!$_cb633_n"; done
|
||||
} > "$ZDOTDIR/%s" 2>/dev/null
|
||||
|
||||
unset _cb633_allowed _cb633_names _cb633_blank _cb633_n _cb633_total _cb633_kept
|
||||
unset _cb633_allowed _cb633_names _cb633_blank _cb633_ok _cb633_unblankable _cb633_n _cb633_total _cb633_kept
|
||||
""".formatted(names, MemberEnvAllowList.zshCasePattern(), REPORT_FILE);
|
||||
}
|
||||
|
||||
@@ -356,6 +434,11 @@ public final class EnvAllowListScrub {
|
||||
* Read and parse {@link #REPORT_FILE} out of a generated ZDOTDIR directory. Returns {@code null}
|
||||
* when absent or unreadable (the pane may have been torn down before its login shell ever got to
|
||||
* the scrub) — callers treat that as "no measurement available", never as success.
|
||||
*
|
||||
* <p>First line is {@code "allowed <N> of <M> failed <F>"} (fleetd #394 added the trailing
|
||||
* {@code failed <F>} — a count of names the scrub attempted to blank but could not, e.g. a zsh
|
||||
* read-only/special parameter). Every following non-blank line is a name: a bare name was
|
||||
* blanked, a {@code !}-prefixed name was attempted and failed. Values never appear on either.
|
||||
*/
|
||||
static ScrubReport readReport(Path zdotdir) {
|
||||
Path report = zdotdir.resolve(REPORT_FILE);
|
||||
@@ -368,17 +451,24 @@ public final class EnvAllowListScrub {
|
||||
return null;
|
||||
}
|
||||
String[] parts = lines.getFirst().substring("allowed ".length()).trim().split("\\s+");
|
||||
if (parts.length != 3 || !"of".equals(parts[1])) {
|
||||
if (parts.length != 5 || !"of".equals(parts[1]) || !"failed".equals(parts[3])) {
|
||||
return null;
|
||||
}
|
||||
List<String> blanked = new ArrayList<>();
|
||||
List<String> unblankable = new ArrayList<>();
|
||||
for (int i = 1; i < lines.size(); i++) {
|
||||
if (!lines.get(i).isBlank()) {
|
||||
blanked.add(lines.get(i));
|
||||
String line = lines.get(i);
|
||||
if (line.isBlank()) {
|
||||
continue;
|
||||
}
|
||||
if (line.startsWith("!")) {
|
||||
unblankable.add(line.substring(1));
|
||||
} else {
|
||||
blanked.add(line);
|
||||
}
|
||||
}
|
||||
return new ScrubReport(Integer.parseInt(parts[0]), Integer.parseInt(parts[2]),
|
||||
List.copyOf(blanked));
|
||||
Integer.parseInt(parts[4]), List.copyOf(blanked), List.copyOf(unblankable));
|
||||
} catch (IOException | NumberFormatException e) {
|
||||
return null;
|
||||
}
|
||||
|
||||
@@ -17,6 +17,7 @@ import dev.ltms.fleet.peer.PeerHandle;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.PeerUnreachableException;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
@@ -594,6 +595,23 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
req.sessionName(), spawned.agentSessionId(), spawned.receipt());
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>fleetd #450: re-enters {@link #spawn(SpawnRequest)} with {@code decision}'s profile named
|
||||
* explicitly. This is the re-entering form the interface javadoc describes for a launcher with
|
||||
* no placement concept of its own — an instance of this class spawns a single adapter's own
|
||||
* profile set by explicit name only ({@link #place}/{@link #defaultProfileFor} are unoverridden
|
||||
* here and just wrap {@link #defaultProfile()}); it does no quarantine/cool-off/maxLoad/model-off
|
||||
* filtering of its own to re-apply. That filtering lives one layer up, in {@code
|
||||
* CompositePeerLauncher}, which is the launcher that routes across more than one profile and
|
||||
* therefore overrides this method with the routing form instead.
|
||||
*/
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
return spawn(req.withProfile(decision.profile()));
|
||||
}
|
||||
|
||||
/** The herdr daemon that owns this launcher's pane coordinates. */
|
||||
public HerdrClient herdr() {
|
||||
return agents.herdr();
|
||||
@@ -1643,11 +1661,19 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
log.warn("memberCredentials allow-list: pane {} left no scrub report in {} — the "
|
||||
+ "environment scrub cannot be confirmed to have run. Either the pane ended "
|
||||
+ "before its shell finished starting, or its shell never read our generated "
|
||||
+ "startup files, in which case that member saw the full host environment.",
|
||||
+ "startup files. Either way, we cannot tell from here whether the scrub ran, "
|
||||
+ "so we do not know what that member's environment contained.",
|
||||
paneId, dir);
|
||||
} else {
|
||||
log.info("memberCredentials allow-list: pane {} allowed {} of {} environment variables",
|
||||
paneId, report.allowed(), report.total());
|
||||
if (report.failed() > 0) {
|
||||
log.warn("memberCredentials allow-list: pane {} could not blank {} environment "
|
||||
+ "variable(s) — {} (likely a zsh read-only/special parameter) — those "
|
||||
+ "names were left in the member's environment. Confirm none of them is a "
|
||||
+ "credential.",
|
||||
paneId, report.failed(), report.unblankable());
|
||||
}
|
||||
List<String> shaped = report.blanked().stream()
|
||||
.filter(name -> CREDENTIAL_SHAPED_NAME.matcher(name).matches())
|
||||
.toList();
|
||||
|
||||
@@ -331,11 +331,22 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
+ "form — opencode's per-model context limit could not be applied for this profile",
|
||||
cfg.profile(), cfg.model());
|
||||
}
|
||||
// fleetd #393: memberSkills seeding (GitWorktrees#seedSkills) copies skill folders into
|
||||
// EVERY provisioned worktree's .claude/skills/ regardless of which kind ultimately spawns
|
||||
// into it — that copy step cannot know the kind, only the caller of GitWorktrees#add does
|
||||
// (see that method's own javadoc). .claude/skills/ is a Claude Code CLI convention the CLI
|
||||
// discovers on its own; opencode has no such discovery, so without this, a seeded skill
|
||||
// never reaches an opencode member even though GitWorktrees logged it as seeded. Read
|
||||
// whatever landed under <cwd>/.claude/skills/ here — the one place in this launcher that
|
||||
// knows both the kind (opencode, by construction: this IS OpenCodeLauncher) and the cwd.
|
||||
List<Path> skillInstructionFiles = skillInstructionFiles(spec.cwd());
|
||||
// A config file is needed for the bridge MCP mount, a member charter, the IDE MCP (+ its
|
||||
// guidance overlay, CB-634), a pinned endpoint (CB-508), or a resolvable autoCompactWindow.
|
||||
// guidance overlay, CB-634), a pinned endpoint (CB-508), a resolvable autoCompactWindow, or
|
||||
// at least one seeded skill to deliver via instructions[] (fleetd #393).
|
||||
if (cfg.hasMcp() || cfg.hasIdeMcp() || spec.charter() != null || hasCustomProvider(cfg)
|
||||
|| wantsContextLimit) {
|
||||
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg, spec.charter(), spec.cwd()).toString());
|
||||
|| wantsContextLimit || !skillInstructionFiles.isEmpty()) {
|
||||
workerEnv.put("OPENCODE_CONFIG",
|
||||
writeConfig(cfg, spec.charter(), spec.cwd(), skillInstructionFiles).toString());
|
||||
}
|
||||
applyGitToken(workerEnv, cfg);
|
||||
List<String> argv = argvWithResume(argvWithModel(argvWithAuto(cfg), cfg), spec.resumeSessionId());
|
||||
@@ -358,6 +369,65 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
return withAgent;
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #393: the {@code SKILL.md} paths under {@code <cwd>/.claude/skills/} this launcher can
|
||||
* turn into {@code instructions[]} entries, plus the honest log this ticket asks for — emitted
|
||||
* here, at the one point a skill's fate for THIS spawn is actually known, rather than trusting
|
||||
* {@code GitWorktrees#seedSkills}'s kind-blind "N of M" line to mean "and it will be read."
|
||||
*
|
||||
* <p>Every non-hidden subdirectory of {@code .claude/skills/} is a candidate, whether it got
|
||||
* there via {@code memberSkills:} seeding or because the target repo ships its own copy — this
|
||||
* launcher does not care which; it only cares what it can find at spawn time. A candidate with
|
||||
* a {@code SKILL.md} at its top level (the same shape {@link #writeConfig} already requires for
|
||||
* the charter and IDE-rules instructions entries) is delivered; anything else is a directory
|
||||
* this launcher cannot turn into a flat instructions entry, named explicitly in the log rather
|
||||
* than silently dropped, so a caller sees a real "cannot consume" reason and not just a smaller
|
||||
* number than {@code GitWorktrees}' own count.
|
||||
*
|
||||
* <p>No candidates at all (directory absent or empty) logs nothing — the same
|
||||
* no-log-when-nothing-to-say shape {@link #hasCustomProvider} and friends already follow, and
|
||||
* the shape {@code GitWorktrees#seedSkills} itself uses when {@code memberSkills:} is unset.
|
||||
* A failure to even list the directory is logged and treated as "nothing delivered" — best
|
||||
* effort, must never fail the spawn, matching {@code GitWorktrees#seedSkills}'s own contract.
|
||||
*/
|
||||
private List<Path> skillInstructionFiles(String cwd) {
|
||||
if (cwd == null || cwd.isBlank()) {
|
||||
return List.of();
|
||||
}
|
||||
Path skillsDir = Path.of(cwd, ".claude", "skills");
|
||||
if (!Files.isDirectory(skillsDir)) {
|
||||
return List.of();
|
||||
}
|
||||
List<Path> candidates;
|
||||
try (var listing = Files.list(skillsDir)) {
|
||||
candidates = listing.filter(Files::isDirectory)
|
||||
.filter(p -> !p.getFileName().toString().startsWith("."))
|
||||
.sorted()
|
||||
.toList();
|
||||
} catch (IOException e) {
|
||||
log.warn("could not scan {} for skill folders to deliver to this opencode member: {}",
|
||||
skillsDir, e.getMessage());
|
||||
return List.of();
|
||||
}
|
||||
if (candidates.isEmpty()) {
|
||||
return List.of();
|
||||
}
|
||||
List<Path> delivered = candidates.stream()
|
||||
.map(dir -> dir.resolve("SKILL.md"))
|
||||
.filter(Files::isRegularFile)
|
||||
.toList();
|
||||
List<String> undeliverable = candidates.stream()
|
||||
.filter(dir -> !Files.isRegularFile(dir.resolve("SKILL.md")))
|
||||
.map(dir -> dir.getFileName().toString())
|
||||
.toList();
|
||||
log.info("skill delivery: {} of {} skill folder(s) under {} reached this opencode member via "
|
||||
+ "instructions[] (opencode does not read .claude/skills/ natively, unlike "
|
||||
+ "Claude Code){}",
|
||||
delivered.size(), candidates.size(), skillsDir,
|
||||
undeliverable.isEmpty() ? "" : "; no SKILL.md, could not be delivered: " + undeliverable);
|
||||
return delivered;
|
||||
}
|
||||
|
||||
/**
|
||||
* True when this profile pins its own OpenAI-compatible endpoint (CB-508) rather than using
|
||||
* whatever provider opencode resolves by default.
|
||||
@@ -418,8 +488,15 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
|
||||
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
|
||||
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
|
||||
*
|
||||
* @param skillInstructionFiles fleetd #393: absolute {@code SKILL.md} paths from
|
||||
* {@link #skillInstructionFiles(String)}, appended to
|
||||
* {@code instructions[]} so a {@code memberSkills:}-seeded skill
|
||||
* reaches this opencode member the same way the charter and IDE
|
||||
* rules already do.
|
||||
*/
|
||||
private Path writeConfig(FleetConfig.Profile cfg, String charterText, String cwd) {
|
||||
private Path writeConfig(FleetConfig.Profile cfg, String charterText, String cwd,
|
||||
List<Path> skillInstructionFiles) {
|
||||
try {
|
||||
Path dir = Files.createTempDirectory(configParentDir(), "fleetd-opencode-");
|
||||
dir.toFile().deleteOnExit();
|
||||
@@ -447,7 +524,39 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
Files.writeString(charter, charterText);
|
||||
charter.toFile().deleteOnExit();
|
||||
|
||||
root.putArray("instructions").add(charter.toAbsolutePath().toString());
|
||||
// fleetd #393 follow-up: withArray, not putArray. putArray REPLACES whatever node
|
||||
// is already at "instructions"; withArray gets-or-creates. All three writers on
|
||||
// this array now use withArray so that write order stops being load-bearing for
|
||||
// whoever adds a fourth.
|
||||
//
|
||||
// Be precise about what this particular line is worth, because an earlier version
|
||||
// of this comment was wrong and claimed too much. THIS call is the one place where
|
||||
// the two idioms are equivalent, and no test can tell them apart: it runs first,
|
||||
// against a still-empty root, so there is never an existing node for putArray to
|
||||
// replace. Measured on the #393 merge: flipping this one call back to putArray
|
||||
// leaves the whole suite green (1618 tests), and always will. The edit is a
|
||||
// readability and future-proofing change with no test behind it, and that is not a
|
||||
// gap anyone can close.
|
||||
//
|
||||
// The ordering hazard is real, just not here. It is the LATER writers that can
|
||||
// destroy earlier entries. Measured on the same merge: flipping the skills writer
|
||||
// below to putArray deletes this charter entry and fails
|
||||
// OpenCodeLauncherTest.instructionsArrayHoldsCharterThenSkillsThenIdeRulesInOrder;
|
||||
// flipping the IDE-rules writer fails that test and
|
||||
// instructionsArrayHoldsCharterThenIdeRulesInOrder. Deleting this line altogether
|
||||
// fails instructionsArrayHoldsExactlyTheCharterWhenNothingElseWritesToIt — the
|
||||
// assertion that a role contract reaches an opencode member at all, which is the
|
||||
// hole that predates #393 and is what actually let the mutation hide.
|
||||
root.withArray("instructions").add(charter.toAbsolutePath().toString());
|
||||
}
|
||||
|
||||
// fleetd #393: each seeded skill's SKILL.md, delivered as a plain instructions[] entry
|
||||
// — the only mechanism opencode has for static guidance text. Unlike Claude Code's
|
||||
// Skill tool, opencode cannot load one of these on demand by name; the content is just
|
||||
// always part of the system prompt from spawn. That is a real difference in HOW the
|
||||
// content reaches the member, not a reason to withhold it.
|
||||
for (Path skillFile : skillInstructionFiles) {
|
||||
root.withArray("instructions").add(skillFile.toAbsolutePath().toString());
|
||||
}
|
||||
|
||||
if (cfg.hasMcp() || cfg.hasIdeMcp()) {
|
||||
|
||||
@@ -394,17 +394,17 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void ack(String target, String msgId) {
|
||||
public boolean ack(String target, String msgId) {
|
||||
var perTarget = held.get(target);
|
||||
if (perTarget == null || perTarget == RELEASED) {
|
||||
return;
|
||||
return false;
|
||||
}
|
||||
Held h;
|
||||
synchronized (perTarget) {
|
||||
h = perTarget.remove(msgId);
|
||||
}
|
||||
if (h == null) {
|
||||
return; // never held (or already acked) — no-op
|
||||
return false; // never held (or already acked) — no-op
|
||||
}
|
||||
try {
|
||||
synchronized (channelLock) {
|
||||
@@ -418,6 +418,7 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
}
|
||||
throw new IllegalStateException("cannot ack reply " + msgId + " on " + queueName(target), e);
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
private DeliverCallback deliverCallback(String target) {
|
||||
|
||||
@@ -59,16 +59,17 @@ public final class InMemoryReplyInbox implements ReplyInbox {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void ack(String target, String msgId) {
|
||||
public boolean ack(String target, String msgId) {
|
||||
if (!owned.contains(target)) {
|
||||
return;
|
||||
return false;
|
||||
}
|
||||
var perTarget = store.get(target);
|
||||
if (perTarget != null) {
|
||||
//noinspection SynchronizationOnLocalVariableOrMethodParameter
|
||||
synchronized (perTarget) {
|
||||
perTarget.remove(msgId);
|
||||
}
|
||||
if (perTarget == null) {
|
||||
return false;
|
||||
}
|
||||
//noinspection SynchronizationOnLocalVariableOrMethodParameter
|
||||
synchronized (perTarget) {
|
||||
return perTarget.remove(msgId) != null;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -42,6 +42,18 @@ public interface LeadChannel {
|
||||
/** This daemon's own lead coordination id — the mailbox it owns, and the {@code from} it sends as. */
|
||||
String selfCoordId();
|
||||
|
||||
/**
|
||||
* Whether a message sitting in {@link #peek}'s held set (fetched but not yet {@link #ack}ed) is
|
||||
* still safe if this daemon crashes or restarts right now — the conclusion of two independent
|
||||
* facts about how this channel owns its own queue: the queue was declared <em>durable</em>, and
|
||||
* the consumer that filled {@code held} uses <em>manual ack</em>, so an unacked delivery is still
|
||||
* owned by the broker rather than only in this process's memory. Both must hold for {@code true};
|
||||
* an implementation must derive this from what it actually did when it declared and consumed its
|
||||
* queue, never return a literal — fleetd #440 found {@code FleetMcp}'s {@code heldDurable} field
|
||||
* doing exactly that, unable to ever report {@code false} even after the fact stopped being true.
|
||||
*/
|
||||
boolean heldDurable();
|
||||
|
||||
/**
|
||||
* A non-destructive look at {@code coordId}'s mailbox — does it exist, how many messages are
|
||||
* waiting on it, and how many consumers are attached — without owning, consuming, or otherwise
|
||||
|
||||
@@ -86,6 +86,12 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
|
||||
private final Object channelLock = new Object();
|
||||
/** msgId → held delivery, for this mailbox's own queue only (there is exactly one). */
|
||||
private final LinkedHashMap<String, Held> held = new LinkedHashMap<>();
|
||||
/**
|
||||
* fleetd #440: the answer to {@link #heldDurable()}, set once by {@link #own()} from the exact
|
||||
* booleans it passed to {@code queueDeclare}/{@code basicConsume} — never a separate literal that
|
||||
* could drift from what those calls actually did.
|
||||
*/
|
||||
private boolean heldDurable;
|
||||
/** Successful broker acks on this connection, retained only to make a repeated caller ack quiet. */
|
||||
private final LinkedHashMap<String, Boolean> recentlyAcked = new LinkedHashMap<>();
|
||||
/** Bounds {@link #recentlyAcked}: it is only an idempotency aid, never delivery state. */
|
||||
@@ -191,13 +197,22 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
|
||||
/** Declare + consume this daemon's own {@code lead.<selfCoordId>.inbox}. Called once, at construction. */
|
||||
private void own() throws IOException {
|
||||
String queue = queueName(selfCoordId);
|
||||
boolean durableQueue = true; // durable, non-exclusive, keep on idle
|
||||
boolean autoAck = false; // manual ack
|
||||
synchronized (channelLock) {
|
||||
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
|
||||
channel.basicConsume(queue, false, deliverCallback(), _ -> { }); // autoAck=false: manual ack
|
||||
channel.queueDeclare(queue, durableQueue, false, false, null);
|
||||
channel.basicConsume(queue, autoAck, deliverCallback(), _ -> { });
|
||||
}
|
||||
// fleetd #440: held mail is durable only while both hold — a durable queue AND manual ack.
|
||||
this.heldDurable = durableQueue && !autoAck;
|
||||
log.debug("lead mailbox owns queue {} for coord-id {}", queue, selfCoordId);
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean heldDurable() {
|
||||
return heldDurable;
|
||||
}
|
||||
|
||||
/**
|
||||
* Publish {@code msg} to {@code toCoordId}'s mailbox and block until the broker's publisher
|
||||
* confirm for it lands. Does <em>not</em> imply owning or consuming {@code toCoordId}'s queue.
|
||||
|
||||
@@ -826,9 +826,13 @@ public final class MessageService {
|
||||
/**
|
||||
* Acknowledge a specific reply by {@code msgId} for {@code target}. Removes it from the inbox
|
||||
* so that a subsequent drain or peek no longer returns it.
|
||||
*
|
||||
* @return {@code true} if an entry was actually removed, {@code false} if {@code msgId} was not
|
||||
* in {@code target}'s inbox (wrong id, wrong target, or already acked). The caller —
|
||||
* {@link dev.ltms.fleet.mcp.FleetMcp#ack} — must not report success on {@code false}.
|
||||
*/
|
||||
public void ackReply(String target, String msgId) {
|
||||
inbox.ack(target, msgId);
|
||||
public boolean ackReply(String target, String msgId) {
|
||||
return inbox.ack(target, msgId);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1317,6 +1321,27 @@ public final class MessageService {
|
||||
return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Test seam only — carries no production behaviour, and nothing in this class calls it;
|
||||
* {@link #pruneTerminalTickets} still reads {@link Task#completedNanos} directly.
|
||||
*
|
||||
* <p>Reports whether {@code ticket}'s completion hook (the {@code whenComplete} registered in
|
||||
* {@link Task}'s constructor) has actually run yet. Exists because {@link #poll} can report
|
||||
* {@link Phase#DONE} for a ticket before that hook fires: {@code CompletableFuture.complete()}
|
||||
* publishes its result and only afterwards runs dependent actions such as {@code whenComplete}
|
||||
* (fleetd #399), so a caller that observes the future done via {@link #poll} is not thereby
|
||||
* guaranteed to also observe {@link Task#completedNanos} stamped. A test that must order both
|
||||
* events — e.g. before advancing an injected clock past the TTL, to avoid stamping the
|
||||
* *advanced* time and masking a real eviction bug — waits on this instead of on
|
||||
* {@link Phase#DONE}.
|
||||
*
|
||||
* @return {@code false} for an unknown ticket or one whose completion hook has not run yet
|
||||
*/
|
||||
boolean isCompletionStampedForTest(String ticket) {
|
||||
Task task = tasks.get(ticket);
|
||||
return task != null && task.completedNanos != null;
|
||||
}
|
||||
|
||||
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
|
||||
private String liveStatus(String target) {
|
||||
try {
|
||||
|
||||
@@ -46,6 +46,13 @@ public interface ReplyInbox {
|
||||
/** Non-destructive snapshot of pending replies for {@code target} (FIFO), empty list if none. */
|
||||
List<InboxMessage> peek(String target);
|
||||
|
||||
/** Remove the reply {@code msgId} for {@code target} once the primary has taken it. No-op if absent. */
|
||||
void ack(String target, String msgId);
|
||||
/**
|
||||
* Remove the reply {@code msgId} for {@code target} once the primary has taken it.
|
||||
*
|
||||
* @return {@code true} if an entry was actually removed, {@code false} if there was nothing to
|
||||
* remove (unknown {@code target}, unowned {@code target}, or a {@code msgId} not held for
|
||||
* it). A {@code false} is not an error — acking a {@code target} this daemon does not own is
|
||||
* part of the normal contract, not a failure.
|
||||
*/
|
||||
boolean ack(String target, String msgId);
|
||||
}
|
||||
|
||||
@@ -230,6 +230,26 @@ public final class ReplyPushLoop {
|
||||
.collect(Collectors.toUnmodifiableSet());
|
||||
}
|
||||
|
||||
/**
|
||||
* Test seam only (fleetd #418): carries no production behaviour, and nothing in this class
|
||||
* calls it. Exposes {@link #pendingQuestionTurnIdsFor} — the exact state {@link #decide} reads
|
||||
* to decide whether a question keeps a lead's schedule alive.
|
||||
*
|
||||
* <p>{@code MessageService.ask()} does three things in order before a question is fully open to
|
||||
* this loop: it flips the ticket's {@code poll()} phase to {@code Phase.ASKING}, then resolves
|
||||
* the reverse-rendezvous waiter, then calls {@link #onQuestionOpened}, which is what actually
|
||||
* populates {@link #pendingQuestions}. A test that barriers on {@code Phase.ASKING} observes only
|
||||
* the first of those three steps — under load the asker thread can be descheduled between steps
|
||||
* one and three, so the barrier releases before this method's underlying map is populated, and
|
||||
* {@link #decide} correctly reports nothing pending yet. A test that must order itself after the
|
||||
* state {@link #decide} actually reads waits on this instead of on the phase.
|
||||
*
|
||||
* @return an unmodifiable snapshot; empty for a lead with no open questions
|
||||
*/
|
||||
Set<String> pendingQuestionTurnIdsForTest(String lead) {
|
||||
return pendingQuestionTurnIdsFor(lead);
|
||||
}
|
||||
|
||||
private List<PendingIncident> pendingIncidentsFor(String lead) {
|
||||
return pendingIncidents.values().stream().filter(i -> lead.equals(i.key().lead())).toList();
|
||||
}
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
package dev.ltms.fleet.peer;
|
||||
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Set;
|
||||
@@ -143,9 +145,160 @@ public interface PeerLauncher {
|
||||
|
||||
/**
|
||||
* The profile a no-argument {@link #spawn(SpawnRequest)} uses, or {@code null} if none is configured.
|
||||
*
|
||||
* <p>fleetd #425: for an implementation with role pools (a no-argument spawn is read as {@link
|
||||
* MemberRole#DEV}, see {@link SpawnRequest}), this must be the profile a live spawn of that role
|
||||
* would actually be placed on right now, not a value captured once at startup — a caller such as
|
||||
* {@code fleet_profiles} relies on this to report a live, not frozen, fact.
|
||||
*/
|
||||
String defaultProfile();
|
||||
|
||||
/**
|
||||
* The profile an unqualified spawn of {@code role} would resolve to right now — the role-aware,
|
||||
* live counterpart of {@link #defaultProfile()} (fleetd #425).
|
||||
*
|
||||
* <p>A caller that must provision something profile-specific (working directory, parity overlay
|
||||
* files) <em>before</em> the actual spawn — {@code SessionManager.acquireWithWorktree} is the one
|
||||
* that exists today — needs the exact profile that spawn will use, for the caller's real role,
|
||||
* not a role-agnostic guess. Calling {@link #defaultProfile()} for that purpose reads {@code
|
||||
* MemberRole#DEV}'s answer regardless of the caller's actual role, which is wrong for any other
|
||||
* role and can provision for a profile the spawn never lands on.
|
||||
*
|
||||
* <p>Default implementation returns {@link #defaultProfile()}, ignoring {@code role} — the right
|
||||
* answer for a launcher with no role-pool concept of its own (e.g. a single {@code
|
||||
* HerdrPeerLauncher} adapter, which is never reached this way in production: {@code
|
||||
* CompositePeerLauncher} always fronts it and resolves roles itself).
|
||||
*
|
||||
* <p>fleetd #453: this default is deliberately <em>not</em> abstract — unlike {@link
|
||||
* #spawn(SpawnRequest, PlacementDecision)} (fleetd #450), there is no live defect in inheriting
|
||||
* it today, and the only current single-adapter implementer ({@code HerdrPeerLauncher}) is
|
||||
* correct to do so. But it stays correct only as long as that holds: <strong>if a launcher ever
|
||||
* routes more than one profile per role, it MUST override this method</strong>, or every role
|
||||
* silently resolves to {@link #defaultProfile()} with no error and no log line. {@code
|
||||
* HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)} — the override in {@code
|
||||
* dev.ltms.fleet.member}, not the declaration below — names this method and {@link #place}
|
||||
* explicitly as "unoverridden here" for exactly this reason. Read it before adding role-pool
|
||||
* routing to any {@code HerdrPeerLauncher} subclass.
|
||||
*
|
||||
* <p>Who is forced to read which paragraph, because it is not symmetric. A new class that
|
||||
* implements this interface directly must write a body for {@link #spawn(SpawnRequest,
|
||||
* PlacementDecision)}, which is abstract here, so it lands on this javadoc. A subclass of
|
||||
* {@code HerdrPeerLauncher} does not: that class already implements the method, and the
|
||||
* subclass inherits the body. So for a subclass this paragraph is advice, not a gate.
|
||||
*/
|
||||
default String defaultProfileFor(MemberRole role) {
|
||||
return defaultProfile();
|
||||
}
|
||||
|
||||
/**
|
||||
* The profile an <em>unqualified</em> spawn of {@code role} would actually be routed to right
|
||||
* now — the same candidate list, the same {@code quarantined}/{@code coolingOff}/{@code
|
||||
* modelOff} filtering, and the same {@code PlacementPolicy} that {@link #spawn} itself
|
||||
* consults for a blank-profile request (fleetd #425 rework).
|
||||
*
|
||||
* <p>This is <em>not</em> {@link #defaultProfileFor}: that method answers "what is first in
|
||||
* {@code role}'s pool", blind to quarantine, cool-off, and the model on/off gate — the right
|
||||
* answer for a role-agnostic, best-effort report ({@code fleet_profiles}' {@code "default"}
|
||||
* field), but the wrong one for a caller that needs the profile a spawn will actually land on.
|
||||
* A quarantined or model-off pool-first profile makes {@link #defaultProfileFor} return a name
|
||||
* an unqualified spawn will never be routed to.
|
||||
*
|
||||
* <p>Just the resolved name, not the full {@link PlacementDecision} — a caller that only wants
|
||||
* to know the answer (a status report, a log line) can call this; a caller that will later
|
||||
* <em>act</em> on the answer by spawning — provisioning a worktree for a specific profile
|
||||
* before the peer exists is the one that matters — must call {@link #place} and carry the
|
||||
* {@link PlacementDecision} itself through to {@link #spawn(SpawnRequest, PlacementDecision)}
|
||||
* instead of calling this method and feeding the string back in as an explicit profile. Doing
|
||||
* that re-enters {@link #spawn(SpawnRequest)}'s explicit-profile branch, which disagrees with
|
||||
* the routing branch on purpose about what an excluded profile means: the routing branch (and
|
||||
* {@link #place}) falls through a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0
|
||||
* profile to the next candidate, while the explicit branch refuses outright — correct for an
|
||||
* operator who named that profile on purpose, wrong for a name that only ever came from placement
|
||||
* itself. That accidental refusal is exactly the regression fleetd #425 rework round 2 fixes:
|
||||
* the default implementation below delegates to {@link #place}, so the two can never drift apart,
|
||||
* but a caller that resolves through this method alone and spawns separately can still recreate
|
||||
* the round-1 defect for itself. (Before fleetd #435, this accident was also reachable through
|
||||
* {@code maxLoad} specifically, because {@code FixedPlacementPolicy} — the default policy — did
|
||||
* not evaluate it at all for automatic selection; #435 closed that gap, so a placement decision
|
||||
* can no longer be at cap in the first place. The refusal-vs-fall-through disagreement above is
|
||||
* the part that was never about {@code maxLoad} and is still real.)
|
||||
*
|
||||
* @throws RuntimeException (implementation-specific, typically a placement exception) if no
|
||||
* candidate in {@code role}'s pool is currently placeable
|
||||
*/
|
||||
default String routedProfileFor(MemberRole role) {
|
||||
return place(role).profile();
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve, <em>without spawning</em>, the {@link PlacementDecision} an unqualified spawn of
|
||||
* {@code role} would make right now — the same candidate list, the same {@code
|
||||
* quarantined}/{@code coolingOff}/{@code modelOff} filtering, and the same {@code
|
||||
* PlacementPolicy} {@link #spawn(SpawnRequest)}'s blank-profile branch itself consults (fleetd
|
||||
* #425 rework).
|
||||
*
|
||||
* <p>Pair this with {@link #spawn(SpawnRequest, PlacementDecision)}, never with {@link
|
||||
* #spawn(SpawnRequest)} fed the decision's profile as an explicit name — see {@link
|
||||
* PlacementDecision}'s own javadoc for why that second form regressed.
|
||||
*
|
||||
* <p>Default implementation wraps {@link #defaultProfile()}, ignoring {@code role} and every
|
||||
* placement condition — the right answer for a launcher with no pool or placement-policy
|
||||
* concept of its own, matching {@link #defaultProfileFor}'s own default.
|
||||
*
|
||||
* <p>fleetd #453: same reasoning as {@link #defaultProfileFor}'s own #453 note — this default
|
||||
* is deliberately not abstract (no live defect today, correct for the sole single-adapter
|
||||
* implementer), but <strong>a launcher that ever routes more than one profile per role MUST
|
||||
* override this method too</strong>, or placement silently ignores {@code role} for it. See
|
||||
* {@code HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)}'s javadoc, which names this
|
||||
* method as "unoverridden here" and why that is correct only for a single-profile adapter.
|
||||
*
|
||||
* @throws RuntimeException (implementation-specific, typically a placement exception) if no
|
||||
* candidate in {@code role}'s pool is currently placeable
|
||||
*/
|
||||
default PlacementDecision place(MemberRole role) {
|
||||
return new PlacementDecision(defaultProfile());
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn against an already-resolved {@link PlacementDecision} from {@link #place}, honoring it
|
||||
* completely: none of the conditions {@link #place} already applied — quarantine, cooling off,
|
||||
* {@code maxLoad} (evaluated by every placement policy including the default {@code fixed},
|
||||
* since fleetd #435), model-off — are re-evaluated here; {@code decision} already reflects them.
|
||||
* This is not skipping a check {@code place} left undone; it is not repeating one {@code place}
|
||||
* already did, and not re-opening the window between that decision and this spawn in which the
|
||||
* underlying state could otherwise move. This is what lets a resolve-then-spawn caller
|
||||
* ({@code SessionManager.acquireWithWorktree}, which must know the profile before it can
|
||||
* provision a worktree for it) and a plain blank-profile {@link #spawn(SpawnRequest)} caller
|
||||
* land on the exact same outcome for the exact same placement state (fleetd #425 rework,
|
||||
* round 2).
|
||||
*
|
||||
* <p>{@code req}'s own {@link SpawnRequest#profileName()} is ignored in favor of {@code
|
||||
* decision.profile()} — the caller is expected to have built {@code req} with a blank or
|
||||
* matching profile; passing a request that names a <em>different</em>, explicit profile than
|
||||
* the decision it is paired with is a caller bug this method does not attempt to detect.
|
||||
*
|
||||
* <p>No default implementation (fleetd #450): the two correct bodies disagree on purpose, so an
|
||||
* implementer must choose one rather than silently inherit whichever this interface happened to
|
||||
* provide. An implementer with no placement concept of its own — spawns a single profile, e.g.
|
||||
* {@code HerdrPeerLauncher} — should delegate to {@link #spawn(SpawnRequest)} with the decision's
|
||||
* profile named explicitly, since there is no separate routing path to honor there: the
|
||||
* explicit-profile branch it re-enters and the routing branch {@link #place} would have used are
|
||||
* the same thing. <strong>A launcher that routes across more than one profile — the way {@code
|
||||
* CompositePeerLauncher} routes across every configured adapter — MUST NOT re-enter {@link
|
||||
* #spawn(SpawnRequest)}.</strong> Doing so re-applies that single-argument method's
|
||||
* explicit-profile checks ({@code enforceNotQuarantined}, {@code enforceNotCoolingOff}, {@code
|
||||
* enforceMaxLoad}, {@code enforceModelEnabled} in {@code CompositePeerLauncher}), which can
|
||||
* refuse the very profile {@link #place} just chose, if the underlying placement state moved in
|
||||
* the window between the {@link #place} call and this one — the exact window this method and
|
||||
* {@link PlacementDecision} exist to close (fleetd #444). Before #450 this was a {@code default}
|
||||
* method that only {@code CompositePeerLauncher} overrode; a future placement-doing launcher
|
||||
* could have inherited the re-entering body silently and never known. Making it abstract turns
|
||||
* that silent inheritance into a compile error.
|
||||
*
|
||||
* @throws IllegalArgumentException if the decision names an unknown profile
|
||||
*/
|
||||
PeerHandle spawn(SpawnRequest req, PlacementDecision decision);
|
||||
|
||||
/**
|
||||
* Resolve the effective working directory for a spawn {@code req} without actually spawning.
|
||||
* Resolution order: requestedCwd → profile cwd → callerCwd → daemon cwd.
|
||||
@@ -192,4 +345,66 @@ public interface PeerLauncher {
|
||||
* @return {@code true} when a reset was sent and its status transition must settle before reuse
|
||||
*/
|
||||
boolean clearContext(String id);
|
||||
|
||||
/**
|
||||
* Model ids the operator has currently turned off in the central {@code models.allow:} list
|
||||
* (fleetd #422) — empty for a launcher with nothing to gate against. {@code fleet_profiles}/
|
||||
* {@code GET /profiles} (via {@code FleetMcp.profilesView}) call this to report which models
|
||||
* are off, and MUST read this exact accessor rather than deriving their own answer: the fleetd
|
||||
* #404 lesson is that a status field reading a different source than the behaviour it describes
|
||||
* can drift from what the gate ({@code CompositePeerLauncher.enforceModelEnabled} and its
|
||||
* candidate filter) actually enforces. A default of {@code Set.of()} keeps every other {@link
|
||||
* PeerLauncher} implementer (the herdr adapters, and the two test-fake implementers) unchanged.
|
||||
*
|
||||
* <p>fleetd #422 follow-up: this alone cannot tell "no {@code models:} block at all" from "a
|
||||
* {@code models:} block where nothing is currently off" — both report an empty set here. Delegates
|
||||
* to {@link #modelGateState()} so the two facts always come from the one read {@link
|
||||
* #modelGateState()}'s implementer makes; do not override this method separately from that one.
|
||||
*/
|
||||
default Set<String> disabledModels() {
|
||||
return modelGateState().off();
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether the central {@code models.allow:} gate (fleetd #422) is armed at all, together with
|
||||
* which model ids are currently off — fleetd #422 follow-up. {@link #disabledModels()} alone
|
||||
* cannot distinguish two states that both report an empty set: a host with no {@code models:}
|
||||
* block (nothing is gated, and nothing can be) and a host WITH a {@code models:} block where
|
||||
* nothing is currently turned off (the gate is armed and reporting zero). This method exists so
|
||||
* a caller — the startup log, {@code fleet_profiles}/{@code GET /profiles} — can tell the two
|
||||
* apart, the same reason {@code CompletionResolver.UnsetMeaning} exists: an accessor that can
|
||||
* legitimately report "empty" must never let a caller guess why.
|
||||
*
|
||||
* <p>Default {@link ModelGateState#notConfigured()} — every launcher without a {@code models:}
|
||||
* block to read from (the herdr adapters, and the two test-fake implementers), matching {@link
|
||||
* #disabledModels()}'s own default of an empty set.
|
||||
*/
|
||||
default ModelGateState modelGateState() {
|
||||
return ModelGateState.notConfigured();
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up: the result of {@link #modelGateState()} — see that method's javadoc
|
||||
* for why "armed" and "off" must be reported together from one read rather than as two
|
||||
* separately-derived facts that a reload landing between them could make disagree.
|
||||
*
|
||||
* @param configured {@code true} when a {@code models:} block exists at all (armed), regardless
|
||||
* of whether anything in it is currently turned off; {@code false} when there
|
||||
* is no block to gate against
|
||||
* @param off the model ids currently turned off; always empty when {@code configured} is
|
||||
* {@code false}
|
||||
*/
|
||||
record ModelGateState(boolean configured, Set<String> off) {
|
||||
public ModelGateState {
|
||||
off = Set.copyOf(off);
|
||||
}
|
||||
|
||||
public static ModelGateState notConfigured() {
|
||||
return new ModelGateState(false, Set.of());
|
||||
}
|
||||
|
||||
public static ModelGateState armed(Set<String> off) {
|
||||
return new ModelGateState(true, off);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -3,6 +3,7 @@ package dev.ltms.fleet.placement;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.Optional;
|
||||
import java.util.OptionalLong;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.function.LongSupplier;
|
||||
@@ -21,46 +22,158 @@ import java.util.function.LongSupplier;
|
||||
* <p>The clock is injected ({@link LongSupplier}, conventionally {@code System::nanoTime} like
|
||||
* {@code FleetHealthMonitor}), never read inline, so a quarantine's expiry is testable without a
|
||||
* real sleep.
|
||||
*
|
||||
* <h2>Escalation (fleetd #466)</h2>
|
||||
* A flat cooldown does not fit every exhaustion. A backend that reports "out of capacity for the
|
||||
* rest of the hour" recovers in one cooldown; a weekly subscription limit does not — it keeps
|
||||
* reporting exhausted on every attempt made before the window resets, so a flat 30-minute cooldown
|
||||
* (the default {@code cooldownNanos}) means roughly 336 pointless spawn attempts across a week, one
|
||||
* every cooldown.
|
||||
*
|
||||
* <p><strong>This class only ever sees the exhaustion signal.</strong> Its only production caller is
|
||||
* {@code Fleetd.exhaustionSink}, wired to fire on a {@code BACKEND_EXHAUSTED} classification alone.
|
||||
* The daemon's other outage state — a credential "cooling off" after repeated non-exhaustion
|
||||
* backend errors (an HTTP 5xx storm, say) — is a separate mechanism, {@code BackendOutagePolicy},
|
||||
* with its own short fixed 60s cooldown and no repeat tracking. The two are never merged: escalating
|
||||
* on a cooling-off signal would turn a transient 5xx storm into a multi-hour backoff, which is
|
||||
* exactly the failure this ticket is not asking for. Confirmed by reading every call site of
|
||||
* {@link #quarantine} — {@code BackendOutagePolicy} has its own {@code coolOff} method and never
|
||||
* calls this one.
|
||||
*
|
||||
* <p><strong>Mechanism</strong> — the {@link #withEscalation} constructors track, per credential, how
|
||||
* many times in a row {@link #quarantine} has been called without an intervening "quiet" gap.
|
||||
* Each call computes {@code cooldownNanos * backoffMultiplier ^ (repeatCount - 1)}, capped at
|
||||
* {@code maxCooldownNanos}. A call counts as a continuation of the same streak — {@code repeatCount}
|
||||
* increments — when it arrives no more than one base {@code cooldownNanos} after the previous
|
||||
* quarantine's deadline (this covers both "still quarantined" and "quarantine just expired and it
|
||||
* was exhausted again immediately"); otherwise the streak resets and this call is treated as a fresh
|
||||
* first occurrence at the base cooldown.
|
||||
*
|
||||
* <p><strong>Reset, honestly stated.</strong> The ideal reset signal is "the cooldown expired and the
|
||||
* next attempt succeeded" — but nothing in this codebase reports a spawn success back to this class
|
||||
* (checked: {@code SessionManager} and {@code CompositePeerLauncher} never call any method here
|
||||
* except {@link #quarantine}/{@link #isQuarantined}/{@link #remainingSeconds}, none of which is a
|
||||
* success hook). Lacking that signal, the reset used here is a time-based proxy: a base-cooldown's
|
||||
* worth of quiet — no exhaustion report for that credential — since the last quarantine ended. It is
|
||||
* not proof the credential started working again, only the best available evidence without adding an
|
||||
* active probe, which is out of scope by the operator's own design constraint (no automatic probing
|
||||
* of a limited backend).
|
||||
*
|
||||
* <p><strong>Ceiling.</strong> {@code maxCooldownNanos} bounds the growth — an unbounded backoff is a
|
||||
* permanent, unrecoverable-without-a-restart outage, which would be worse than the flat-rate bug this
|
||||
* escalation fixes. {@link #withEscalation(LongSupplier, long)} defaults the ceiling to
|
||||
* {@value #DEFAULT_MAX_COOLDOWN_MULTIPLE}x the base cooldown (12x the 1800s default ≈ 6 hours), so a
|
||||
* chronically exhausted credential still gets re-tried roughly every 6 hours instead of every 30
|
||||
* minutes — about a dozen attempts a week instead of ~336.
|
||||
*
|
||||
* <p><strong>Backward compatibility.</strong> The original two-argument {@link #BackendQuarantine(
|
||||
* LongSupplier, long)} constructor is unchanged in behaviour: it is exactly {@code
|
||||
* withEscalation}'s mechanism with {@code backoffMultiplier = 1.0} and {@code maxCooldownNanos =
|
||||
* cooldownNanos}, which collapses the formula back to the original flat {@code now + cooldownNanos}
|
||||
* on every call regardless of history. Every existing call site (roughly 20 across the test suite,
|
||||
* plus {@link #none()}) keeps its current shape and behaviour unchanged.
|
||||
*/
|
||||
public final class BackendQuarantine {
|
||||
|
||||
private final ConcurrentHashMap<String, Long> quarantinedUntilNanos = new ConcurrentHashMap<>();
|
||||
/** Default growth per consecutive exhaustion streak — see the class doc's Mechanism section. */
|
||||
static final double DEFAULT_BACKOFF_MULTIPLIER = 2.0;
|
||||
/** Default ceiling, expressed as a multiple of the base cooldown — see the class doc's Ceiling section. */
|
||||
static final long DEFAULT_MAX_COOLDOWN_MULTIPLE = 12;
|
||||
|
||||
private final ConcurrentHashMap<String, QuarantineState> quarantines = new ConcurrentHashMap<>();
|
||||
private final LongSupplier nowNanos;
|
||||
private final long cooldownNanos;
|
||||
private final double backoffMultiplier;
|
||||
private final long maxCooldownNanos;
|
||||
/** True only for {@link #none()}. See {@link #quarantine} for why this exists. */
|
||||
private final boolean inert;
|
||||
|
||||
/** How many consecutive exhaustion reports a credential is on, and when the resulting cooldown ends. */
|
||||
private record QuarantineState(int repeatCount, long deadlineNanos) {
|
||||
}
|
||||
|
||||
/**
|
||||
* Flat cooldown, unchanged from before fleetd #466 — every {@link #quarantine} call blocks the
|
||||
* credential for exactly {@code cooldownNanos}, regardless of how many times it was called
|
||||
* before. Equivalent to {@link #withEscalation} with no growth ({@code backoffMultiplier = 1.0})
|
||||
* and a ceiling equal to the base cooldown, so it degrades to the identical {@code now +
|
||||
* cooldownNanos} formula every call. Kept for the existing call sites that want a fixed cooldown
|
||||
* (and for tests exercising the fixed-cooldown shape in isolation); production wiring uses
|
||||
* {@link #withEscalation} instead.
|
||||
*
|
||||
* @param nowNanos monotonic clock, injected for testability
|
||||
* @param cooldownNanos how long a fresh {@link #quarantine} call blocks the credential for;
|
||||
* must be positive
|
||||
*/
|
||||
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos) {
|
||||
this(nowNanos, cooldownNanos, false);
|
||||
this(nowNanos, cooldownNanos, 1.0, cooldownNanos, false);
|
||||
}
|
||||
|
||||
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, boolean inert) {
|
||||
/**
|
||||
* Escalating cooldown (fleetd #466) — see the class doc's Mechanism/Reset/Ceiling sections.
|
||||
*
|
||||
* @param nowNanos monotonic clock, injected for testability
|
||||
* @param cooldownNanos base cooldown, applied to a fresh (non-streak) exhaustion; must be
|
||||
* positive
|
||||
* @param backoffMultiplier growth per consecutive exhaustion; must be {@code >= 1.0} ({@code 1.0}
|
||||
* disables growth and is exactly the flat two-argument constructor)
|
||||
* @param maxCooldownNanos ceiling on the escalated cooldown; must be {@code >= cooldownNanos}
|
||||
*/
|
||||
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, double backoffMultiplier,
|
||||
long maxCooldownNanos) {
|
||||
this(nowNanos, cooldownNanos, backoffMultiplier, maxCooldownNanos, false);
|
||||
}
|
||||
|
||||
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, double backoffMultiplier,
|
||||
long maxCooldownNanos, boolean inert) {
|
||||
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
|
||||
if (cooldownNanos <= 0) {
|
||||
throw new IllegalArgumentException("cooldownNanos must be positive: " + cooldownNanos);
|
||||
}
|
||||
if (backoffMultiplier < 1.0) {
|
||||
throw new IllegalArgumentException("backoffMultiplier must be >= 1.0: " + backoffMultiplier);
|
||||
}
|
||||
if (maxCooldownNanos < cooldownNanos) {
|
||||
throw new IllegalArgumentException(
|
||||
"maxCooldownNanos must be >= cooldownNanos: " + maxCooldownNanos + " < " + cooldownNanos);
|
||||
}
|
||||
this.cooldownNanos = cooldownNanos;
|
||||
this.backoffMultiplier = backoffMultiplier;
|
||||
this.maxCooldownNanos = maxCooldownNanos;
|
||||
this.inert = inert;
|
||||
}
|
||||
|
||||
/**
|
||||
* Escalating cooldown with the fleetd #466 default shape: cooldown doubles
|
||||
* ({@value #DEFAULT_BACKOFF_MULTIPLIER}x) per consecutive exhaustion streak, capped at
|
||||
* {@value #DEFAULT_MAX_COOLDOWN_MULTIPLE}x the base cooldown. This is what production wiring
|
||||
* ({@code Fleetd.main}) uses.
|
||||
*
|
||||
* @param nowNanos monotonic clock, injected for testability
|
||||
* @param cooldownNanos base cooldown, applied to a fresh (non-streak) exhaustion; must be positive
|
||||
*/
|
||||
public static BackendQuarantine withEscalation(LongSupplier nowNanos, long cooldownNanos) {
|
||||
return new BackendQuarantine(nowNanos, cooldownNanos, DEFAULT_BACKOFF_MULTIPLIER,
|
||||
cooldownNanos * DEFAULT_MAX_COOLDOWN_MULTIPLE, false);
|
||||
}
|
||||
|
||||
/**
|
||||
* Inert quarantine — {@link #quarantine} does nothing on this instance, so nothing is ever
|
||||
* quarantined. The explicit stand-in a caller (or a test not exercising this feature) passes
|
||||
* instead of a defaulting overload, exactly like {@code ExhaustedPatternLookup.none()}.
|
||||
*/
|
||||
public static BackendQuarantine none() {
|
||||
return new BackendQuarantine(() -> 0L, 1, true);
|
||||
return new BackendQuarantine(() -> 0L, 1, 1.0, 1, true);
|
||||
}
|
||||
|
||||
/**
|
||||
* Quarantine {@code credentialId} for the configured cooldown, starting now. A repeat call while
|
||||
* already quarantined restarts the cooldown at full length — a fresh refusal is fresh evidence the
|
||||
* account is still exhausted, not a reason to let an earlier, shorter wait stand.
|
||||
* Quarantine {@code credentialId} starting now. On a flat instance (the two-argument
|
||||
* constructor) this always blocks for exactly {@code cooldownNanos}, restarting the cooldown at
|
||||
* full length on every call — a fresh refusal is fresh evidence the account is still exhausted,
|
||||
* not a reason to let an earlier, shorter wait stand. On an escalating instance ({@link
|
||||
* #withEscalation}) the cooldown grows with each call that arrives within one base cooldown of
|
||||
* the previous deadline, and resets to the base cooldown once a call arrives after a longer gap
|
||||
* — see the class doc.
|
||||
*
|
||||
* <p>On {@link #none()} this is a no-op. It has to be: that instance holds a clock frozen at 0,
|
||||
* so recording a deadline would produce a quarantine that never expires — a credential locked out
|
||||
@@ -73,7 +186,13 @@ public final class BackendQuarantine {
|
||||
if (inert) {
|
||||
return;
|
||||
}
|
||||
quarantinedUntilNanos.put(credentialId, nowNanos.getAsLong() + cooldownNanos);
|
||||
long now = nowNanos.getAsLong();
|
||||
quarantines.compute(credentialId, (id, prev) -> {
|
||||
int repeatCount = (prev == null || now - prev.deadlineNanos() > cooldownNanos)
|
||||
? 1
|
||||
: prev.repeatCount() + 1;
|
||||
return new QuarantineState(repeatCount, now + escalatedCooldownNanos(repeatCount));
|
||||
});
|
||||
}
|
||||
|
||||
/** Whether {@code credentialId} is quarantined right now. */
|
||||
@@ -87,6 +206,49 @@ public final class BackendQuarantine {
|
||||
return remaining > 0 ? OptionalLong.of(toSecondsRoundedUp(remaining)) : OptionalLong.empty();
|
||||
}
|
||||
|
||||
/**
|
||||
* Remaining seconds together with which consecutive exhaustion this is (fleetd #466 scope item
|
||||
* 2) — {@code repeatCount} 1 for a first occurrence, 2 for the second in a row, and so on; see
|
||||
* {@link #quarantine}'s class-doc Mechanism section for exactly when a call continues a streak
|
||||
* versus starts a fresh one.
|
||||
*
|
||||
* <p><strong>Read together, off the one {@link QuarantineState} entry {@link #quarantine} itself
|
||||
* wrote</strong> — a single {@code quarantines.get(credentialId)}, never a separate lookup or a
|
||||
* value re-derived from {@code remainingSeconds} (e.g. inverting {@link
|
||||
* #escalatedCooldownNanos}). That inversion is not just extra work to avoid: once a streak has
|
||||
* hit {@code maxCooldownNanos}, every further consecutive exhaustion reports the identical
|
||||
* cooldown, so a derivation that starts from the cooldown value cannot tell the 4th repeat from
|
||||
* the 9th — only the stored {@code repeatCount} can. This is the same rule {@code
|
||||
* CompositePeerLauncher.modelGateState()} documents for its own gate/report pair: the report
|
||||
* reads the exact accessor the behaviour reads, so it can never disagree with what actually
|
||||
* happened (the fleetd #404/#422 lesson). {@code fleet_profiles}/{@code fleet_list}/{@code GET
|
||||
* /profiles} all call this — never {@link #remainingSeconds} plus a second, independent count —
|
||||
* for exactly that reason.
|
||||
*
|
||||
* @return empty when {@code credentialId} is not currently quarantined (including on {@link
|
||||
* #none()}, which quarantines nothing)
|
||||
*/
|
||||
public Optional<Status> status(String credentialId) {
|
||||
QuarantineState state = quarantines.get(credentialId);
|
||||
if (state == null) {
|
||||
return Optional.empty();
|
||||
}
|
||||
long remaining = state.deadlineNanos() - nowNanos.getAsLong();
|
||||
return remaining > 0
|
||||
? Optional.of(new Status(toSecondsRoundedUp(remaining), state.repeatCount()))
|
||||
: Optional.empty();
|
||||
}
|
||||
|
||||
/**
|
||||
* @param remainingSeconds seconds left on the quarantine, identical to {@link
|
||||
* #remainingSeconds(String)}'s answer for the same credential at the
|
||||
* same instant
|
||||
* @param repeatCount 1 for a first occurrence, 2 for the second consecutive one, etc. —
|
||||
* see {@link #status(String)}
|
||||
*/
|
||||
public record Status(long remainingSeconds, int repeatCount) {
|
||||
}
|
||||
|
||||
/**
|
||||
* Every currently-quarantined credential id and its remaining seconds (CB-578 stage B fleet
|
||||
* reporting) — expired entries are never included. Not pruned from the backing map here: it stays
|
||||
@@ -95,8 +257,8 @@ public final class BackendQuarantine {
|
||||
*/
|
||||
public Map<String, Long> activeRemainingSeconds() {
|
||||
Map<String, Long> out = new LinkedHashMap<>();
|
||||
quarantinedUntilNanos.forEach((credentialId, deadline) -> {
|
||||
long remaining = deadline - nowNanos.getAsLong();
|
||||
quarantines.forEach((credentialId, state) -> {
|
||||
long remaining = state.deadlineNanos() - nowNanos.getAsLong();
|
||||
if (remaining > 0) {
|
||||
out.put(credentialId, toSecondsRoundedUp(remaining));
|
||||
}
|
||||
@@ -105,8 +267,14 @@ public final class BackendQuarantine {
|
||||
}
|
||||
|
||||
private long remainingNanos(String credentialId) {
|
||||
Long deadline = quarantinedUntilNanos.get(credentialId);
|
||||
return deadline == null ? 0L : deadline - nowNanos.getAsLong();
|
||||
QuarantineState state = quarantines.get(credentialId);
|
||||
return state == null ? 0L : state.deadlineNanos() - nowNanos.getAsLong();
|
||||
}
|
||||
|
||||
/** {@code cooldownNanos * backoffMultiplier ^ (repeatCount - 1)}, capped at {@code maxCooldownNanos}. */
|
||||
private long escalatedCooldownNanos(int repeatCount) {
|
||||
double raw = cooldownNanos * Math.pow(backoffMultiplier, repeatCount - 1);
|
||||
return raw >= (double) maxCooldownNanos ? maxCooldownNanos : (long) raw;
|
||||
}
|
||||
|
||||
private static long toSecondsRoundedUp(long nanos) {
|
||||
|
||||
@@ -5,14 +5,12 @@ import java.util.List;
|
||||
|
||||
/**
|
||||
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
|
||||
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps
|
||||
* ({@code maxLoad}) so that a pre-existing config behaves identically after upgrade — capacity
|
||||
* gating for automatic placement is deliberately out of scope for {@code fixed}, exactly as it
|
||||
* always has been. Reachability is a narrower exception (fleetd #315, below): a profile is never
|
||||
* checked for reachability up front, only skipped once it has already failed in <em>this same</em>
|
||||
* spawn call's retry loop — see the unreachable case below.
|
||||
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. Reachability is a narrower
|
||||
* exception (fleetd #315, below): a profile is never checked for reachability up front, only
|
||||
* skipped once it has already failed in <em>this same</em> spawn call's retry loop — see the
|
||||
* unreachable case below.
|
||||
*
|
||||
* <p>Four exceptions walk past the default instead of returning it unconditionally:
|
||||
* <p>Six exceptions walk past the default instead of returning it unconditionally:
|
||||
* <ul>
|
||||
* <li>Quarantine (CB-578 stage B): a quarantined default is a credential that just refused on
|
||||
* a usage limit, not a transient capacity or reachability concern.
|
||||
@@ -20,6 +18,24 @@ import java.util.List;
|
||||
* ({@code BackendOutagePolicy}) — a separate, shorter-lived source from quarantine. When a
|
||||
* profile is both quarantined and cooling off, only the quarantine reason is reported
|
||||
* (exhaustion takes priority), matching {@code CompositePeerLauncher}'s explicit-spawn order.
|
||||
* <li>At cap (fleetd #435): a profile whose live count has reached its {@code maxLoad}
|
||||
* ({@link PlacementPolicyUtil#atCap}) — a documented, unconditional capacity limit (see
|
||||
* {@code FleetConfig.Profile#maxLoad}), so {@code fixed} must gate on it exactly as {@code
|
||||
* weighted}/{@code round-robin} already do via {@link PlacementPolicyUtil#available}. Before
|
||||
* this fix {@code fixed} built its own {@link PlacementCandidate} for the default with {@code
|
||||
* maxLoad} forced to {@code null}, so a capped default was chosen anyway on every unqualified
|
||||
* spawn — the cap was advisory, not enforced, for the one placement policy every config uses
|
||||
* by default. Reported only when quarantine and cooling off are both absent, matching {@code
|
||||
* CompositePeerLauncher}'s explicit-spawn check order (quarantine, then cooling off, then max
|
||||
* load, then model-off).
|
||||
* <li>Model off (fleetd #422): a profile whose {@code model:} the operator has turned off in
|
||||
* {@code models.allow:} — an operator decision, never a backend-reported outage, so it is a
|
||||
* fifth, independent source from quarantine, cooling off, and at-cap (never merged with any
|
||||
* of them), exactly as {@code CompositePeerLauncher.enforceModelEnabled} and {@link
|
||||
* PlacementPolicyUtil#available} treat it. When a profile is model-off <em>and</em> quarantined,
|
||||
* cooling off, or at cap, only the higher-priority reason is reported, matching {@code
|
||||
* CompositePeerLauncher}'s explicit-spawn check order (quarantine, then cooling off, then max
|
||||
* load, then model-off).
|
||||
* <li>Unreachable (fleetd #315): {@code CompositePeerLauncher.spawn} retries a failed candidate
|
||||
* on the next one and rebuilds the {@link PlacementContext} so {@code ctx.unreachable()}
|
||||
* names every profile that already failed with {@code PeerUnreachableException} in this same
|
||||
@@ -32,9 +48,10 @@ import java.util.List;
|
||||
* {@code weighted}/{@code round-robin} skip it — an explicit {@code fleet_spawn} naming
|
||||
* the profile is unaffected, only this automatic fallback walk.
|
||||
* </ul>
|
||||
* A fleet where nothing is ever quarantined, cooling off, unreachable, or weight-0 never exercises
|
||||
* any of these paths, so today's behaviour is unchanged — in particular, the very first selection
|
||||
* of a spawn call always sees an empty {@code unreachable} set, so the first choice is untouched.
|
||||
* A fleet where nothing is ever quarantined, cooling off, at cap, model-off, unreachable, or
|
||||
* weight-0 never exercises any of these paths, so today's behaviour is unchanged — in particular,
|
||||
* the very first selection of a spawn call always sees an empty {@code unreachable} set, so the
|
||||
* first choice is untouched.
|
||||
*/
|
||||
final class FixedPlacementPolicy implements PlacementPolicy {
|
||||
|
||||
@@ -42,12 +59,14 @@ final class FixedPlacementPolicy implements PlacementPolicy {
|
||||
public PlacementCandidate select(PlacementContext ctx) {
|
||||
String d = ctx.defaultProfile();
|
||||
if (d != null && !d.isBlank() && !ctx.quarantined().contains(d) && !ctx.coolingOff().contains(d)
|
||||
&& !ctx.unreachable().contains(d) && !weightExcluded(ctx, d)) {
|
||||
&& !ctx.modelOff().contains(d) && !ctx.unreachable().contains(d) && !weightExcluded(ctx, d)
|
||||
&& !capExcluded(ctx, d)) {
|
||||
return new PlacementCandidate(d, null, 1.0f, null);
|
||||
}
|
||||
for (PlacementCandidate c : ctx.candidates()) {
|
||||
if (!ctx.quarantined().contains(c.profile()) && !ctx.coolingOff().contains(c.profile())
|
||||
&& !ctx.unreachable().contains(c.profile()) && !c.excluded()) {
|
||||
&& !ctx.modelOff().contains(c.profile()) && !ctx.unreachable().contains(c.profile())
|
||||
&& !c.excluded() && !PlacementPolicyUtil.atCap(ctx, c)) {
|
||||
return new PlacementCandidate(c.profile(), null, c.weight(), c.maxLoad());
|
||||
}
|
||||
}
|
||||
@@ -56,9 +75,18 @@ final class FixedPlacementPolicy implements PlacementPolicy {
|
||||
// Exhaustion quarantine takes priority: reported only when quarantine is absent, so the
|
||||
// message never claims "cooling off" for a profile that is really backend-exhausted.
|
||||
boolean dCoolingOff = !dQuarantined && ctx.coolingOff().contains(d);
|
||||
// fleetd #435: at-cap sits between cooling off and model-off, matching
|
||||
// CompositePeerLauncher's explicit-spawn check order (quarantine, cooling off, max load,
|
||||
// then model-off) — reported only when quarantine/cooling-off are both absent.
|
||||
boolean dAtCap = !dQuarantined && !dCoolingOff && capExcluded(ctx, d);
|
||||
// fleetd #422: model-off is a fifth, independent source (an operator decision) — but
|
||||
// quarantine/cooling-off/at-cap still take priority when more than one applies, matching
|
||||
// CompositePeerLauncher's explicit-spawn check order (quarantine, cooling off, max load,
|
||||
// then model-off).
|
||||
boolean dModelOff = !dQuarantined && !dCoolingOff && !dAtCap && ctx.modelOff().contains(d);
|
||||
boolean dUnreachable = ctx.unreachable().contains(d);
|
||||
boolean dWeightExcluded = weightExcluded(ctx, d);
|
||||
if (dQuarantined || dCoolingOff || dUnreachable || dWeightExcluded) {
|
||||
if (dQuarantined || dCoolingOff || dAtCap || dModelOff || dUnreachable || dWeightExcluded) {
|
||||
List<String> reasons = new ArrayList<>();
|
||||
if (dQuarantined) {
|
||||
reasons.add("is quarantined (backend exhausted)");
|
||||
@@ -66,6 +94,14 @@ final class FixedPlacementPolicy implements PlacementPolicy {
|
||||
if (dCoolingOff) {
|
||||
reasons.add("is cooling off after repeated backend errors");
|
||||
}
|
||||
if (dAtCap) {
|
||||
PlacementCandidate c = candidateFor(ctx, d);
|
||||
int live = ctx.liveCount().apply(d);
|
||||
reasons.add("is at maxLoad (" + live + " live >= " + c.maxLoad() + " cap)");
|
||||
}
|
||||
if (dModelOff) {
|
||||
reasons.add("names a model the operator has turned off in models.allow");
|
||||
}
|
||||
if (dUnreachable) {
|
||||
reasons.add("is unreachable");
|
||||
}
|
||||
@@ -78,18 +114,38 @@ final class FixedPlacementPolicy implements PlacementPolicy {
|
||||
}
|
||||
if (!ctx.candidates().isEmpty()) {
|
||||
throw new PlacementException("all worker profiles are excluded from automatic "
|
||||
+ "selection (quarantined, cooling off, unreachable, or weight-0)");
|
||||
+ "selection (quarantined, cooling off, at cap, model-off, unreachable, or weight-0)");
|
||||
}
|
||||
throw new PlacementException("no worker profiles configured");
|
||||
}
|
||||
|
||||
/** Whether {@code profile} carries {@code weight <= 0} (CB-554) among {@code ctx}'s candidates. */
|
||||
private static boolean weightExcluded(PlacementContext ctx, String profile) {
|
||||
PlacementCandidate c = candidateFor(ctx, profile);
|
||||
return c != null && c.excluded();
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether {@code profile} has reached its {@code maxLoad} cap (fleetd #435), using the shared
|
||||
* {@link PlacementPolicyUtil#atCap} definition — the same one {@code weighted}/{@code
|
||||
* round-robin} already consult via {@link PlacementPolicyUtil#available}. Looked up by name,
|
||||
* the same way {@link #weightExcluded} is: the default fast path above builds its own {@link
|
||||
* PlacementCandidate} with {@code maxLoad} forced to {@code null} (it carries no cap of its
|
||||
* own), so the candidate actually configured for {@code profile} has to be found in {@code
|
||||
* ctx.candidates()} first.
|
||||
*/
|
||||
private static boolean capExcluded(PlacementContext ctx, String profile) {
|
||||
PlacementCandidate c = candidateFor(ctx, profile);
|
||||
return c != null && PlacementPolicyUtil.atCap(ctx, c);
|
||||
}
|
||||
|
||||
/** The configured candidate named {@code profile} in {@code ctx}, or {@code null} if none. */
|
||||
private static PlacementCandidate candidateFor(PlacementContext ctx, String profile) {
|
||||
for (PlacementCandidate c : ctx.candidates()) {
|
||||
if (c.profile().equals(profile)) {
|
||||
return c.excluded();
|
||||
return c;
|
||||
}
|
||||
}
|
||||
return false;
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -22,11 +22,29 @@ import java.util.function.Function;
|
||||
* the two apart so its refusal message says "cooling off", not "exhausted",
|
||||
* when only this one is active. A profile can be in both sets at once; when it
|
||||
* is, exhaustion quarantine is reported (it takes priority).
|
||||
* @param modelOff profiles whose {@code model:} is currently turned off in {@code
|
||||
* models.allow:} (fleetd #422) — an operator decision, not a backend-reported
|
||||
* outage, so a SEPARATE, independent source from both {@code quarantined} and
|
||||
* {@code coolingOff}. A profile can be in this set together with either (or
|
||||
* both) of the others; {@link PlacementPolicyUtil} counts it into its own
|
||||
* bucket rather than merging it into theirs, the same reason
|
||||
* {@code coolingOff} is kept apart from {@code quarantined}.
|
||||
*/
|
||||
public record PlacementContext(String defaultProfile,
|
||||
List<PlacementCandidate> candidates,
|
||||
Function<String, Integer> liveCount,
|
||||
Set<String> unreachable,
|
||||
Set<String> quarantined,
|
||||
Set<String> coolingOff) {
|
||||
Set<String> coolingOff,
|
||||
Set<String> modelOff) {
|
||||
|
||||
/**
|
||||
* Back-compat form before the fleetd #422 model on/off gate was added — no candidate's model
|
||||
* is off. Keeps pre-#422 call sites (tests included) compiling and behaving identically.
|
||||
*/
|
||||
public PlacementContext(String defaultProfile, List<PlacementCandidate> candidates,
|
||||
Function<String, Integer> liveCount, Set<String> unreachable,
|
||||
Set<String> quarantined, Set<String> coolingOff) {
|
||||
this(defaultProfile, candidates, liveCount, unreachable, quarantined, coolingOff, Set.of());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,49 @@
|
||||
package dev.ltms.fleet.placement;
|
||||
|
||||
import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
|
||||
/**
|
||||
* An already-completed placement choice — the outcome of one {@link PeerLauncher#place} call,
|
||||
* carried forward so a later {@link PeerLauncher#spawn(SpawnRequest, PlacementDecision)} can honor
|
||||
* it directly instead of re-resolving the profile a second time (fleetd #425 rework, round 2).
|
||||
*
|
||||
* <p>The problem this exists to close: a caller that must know the profile <em>before</em> it can
|
||||
* spawn — {@code SessionManager.acquireWithWorktree} provisions a worktree's {@code repoRoot} and
|
||||
* parity overlay for a specific profile before the peer process exists — used to resolve that name
|
||||
* with {@code PeerLauncher.routedProfileFor(role)} and then hand the SAME string back to {@link
|
||||
* PeerLauncher#spawn(SpawnRequest)} as an EXPLICIT profile. That re-resolution is not free: naming
|
||||
* a profile explicitly makes {@code CompositePeerLauncher.spawn} take its THROWING branch
|
||||
* ({@code enforceNotQuarantined}/{@code enforceNotCoolingOff}/{@code enforceMaxLoad}/{@code
|
||||
* enforceModelEnabled}), while an unqualified spawn's ROUTING branch never runs those checks at
|
||||
* all — it instead FALLS THROUGH to the next candidate on exactly the same conditions the throwing
|
||||
* branch refuses on. That disagreement is deliberate: an operator who names a profile should get a
|
||||
* refusal, not a silent substitution. The bug is turning the fall-through into a refusal by
|
||||
* accident — resolving a name through the routing side and then re-entering the refusing side with
|
||||
* it, for a decision the routing side had already approved by walking past everything else.
|
||||
* Before fleetd #435, this accident was also reachable through {@code maxLoad} specifically: the
|
||||
* default {@code fixed} placement policy did not evaluate {@code maxLoad} at all for automatic
|
||||
* selection, so a profile placement itself just approved could still die at {@code enforceMaxLoad}
|
||||
* one call later, purely because the caller's route to the spawn passed through an explicit
|
||||
* profile name instead of the routing branch — a failure a worktree-less unqualified spawn would
|
||||
* never hit. fleetd #435 closed that specific gap ({@code fixed} now evaluates {@code maxLoad}
|
||||
* exactly like every other placement policy), so a {@link PlacementDecision} can no longer be
|
||||
* at-cap in the first place — but the refusal-vs-fall-through disagreement above was never about
|
||||
* {@code maxLoad}, and resolving a name and re-entering the refusing branch with it is still wrong
|
||||
* for every OTHER condition placement filters on.
|
||||
*
|
||||
* <p>{@link PeerLauncher#spawn(SpawnRequest, PlacementDecision)} closes that by spawning through
|
||||
* the identical code path the routing branch itself uses, keyed off the SAME decision {@link
|
||||
* PeerLauncher#place} returned — no re-checking of any condition placement already evaluated. A
|
||||
* resolve-then-spawn caller and a blank-profile {@link PeerLauncher#spawn(SpawnRequest)} caller can
|
||||
* then never disagree about which conditions apply to the same placement state, and neither one
|
||||
* re-opens the window between the placement decision and the spawn in which the underlying state
|
||||
* could otherwise move.
|
||||
*
|
||||
* @param profile the profile this decision resolved to (may be {@code null} only when no profile is
|
||||
* configured at all — the same corner case {@link PeerLauncher#defaultProfile()}
|
||||
* already tolerates)
|
||||
*/
|
||||
public record PlacementDecision(String profile) {
|
||||
}
|
||||
@@ -11,28 +11,40 @@ final class PlacementPolicyUtil {
|
||||
private PlacementPolicyUtil() {
|
||||
}
|
||||
|
||||
/**
|
||||
* True when {@code c} has reached its {@code maxLoad} cap: {@code liveCount(c.profile()) >=
|
||||
* c.maxLoad()}. A {@code null} maxLoad means unlimited, so it is never at cap.
|
||||
*
|
||||
* <p>Extracted as the single shared definition of "at cap" (fleetd #435): before this fix it
|
||||
* was computed inline in both {@link #available} and {@link #emptyException}, and {@code
|
||||
* FixedPlacementPolicy} — not a caller of either — quietly kept its own {@code select} free of
|
||||
* any cap check at all, so a capped default profile was chosen anyway under the default
|
||||
* placement policy. Every automatic policy must call this, not re-derive it.
|
||||
*/
|
||||
static boolean atCap(PlacementContext ctx, PlacementCandidate c) {
|
||||
Integer cap = c.maxLoad();
|
||||
return cap != null && ctx.liveCount().apply(c.profile()) >= cap;
|
||||
}
|
||||
|
||||
/**
|
||||
* Candidates that are not weight-excluded (CB-554: explicit {@code weight <= 0}, checked
|
||||
* first because it is a static config choice rather than transient state), not
|
||||
* known-unreachable, not quarantined (CB-578 stage B), not cooling off after repeated backend
|
||||
* errors (fleetd #201 Unit 5 — a separate, shorter-lived source from quarantine), and have not
|
||||
* reached their maxLoad. A {@code null} maxLoad means unlimited.
|
||||
* errors (fleetd #201 Unit 5 — a separate, shorter-lived source from quarantine), not naming a
|
||||
* model the operator has turned off (fleetd #422 — a third, independent source: an operator
|
||||
* decision, never a backend-reported outage), and have not reached their maxLoad (see {@link
|
||||
* #atCap}). A {@code null} maxLoad means unlimited.
|
||||
*/
|
||||
static List<PlacementCandidate> available(PlacementContext ctx) {
|
||||
List<PlacementCandidate> out = new ArrayList<>();
|
||||
for (PlacementCandidate c : ctx.candidates()) {
|
||||
if (c.excluded() || ctx.unreachable().contains(c.profile())
|
||||
|| ctx.quarantined().contains(c.profile())
|
||||
|| ctx.coolingOff().contains(c.profile())) {
|
||||
|| ctx.coolingOff().contains(c.profile())
|
||||
|| ctx.modelOff().contains(c.profile())
|
||||
|| atCap(ctx, c)) {
|
||||
continue;
|
||||
}
|
||||
Integer cap = c.maxLoad();
|
||||
if (cap != null) {
|
||||
int live = ctx.liveCount().apply(c.profile());
|
||||
if (live >= cap) {
|
||||
continue;
|
||||
}
|
||||
}
|
||||
out.add(c);
|
||||
}
|
||||
return out;
|
||||
@@ -40,12 +52,13 @@ final class PlacementPolicyUtil {
|
||||
|
||||
/**
|
||||
* Build a clear exception describing why every candidate was dropped: all weight-0, all
|
||||
* quarantined, all cooling off, all at capacity, all unreachable, or a mix. Each candidate is
|
||||
* counted into exactly one bucket (weight-excluded first, then quarantined, then cooling off)
|
||||
* so a candidate excluded for more than one reason is never double-counted — a candidate that is
|
||||
* both quarantined (CB-578 stage B, backend exhausted) and cooling off (fleetd #201 Unit 5,
|
||||
* repeated backend errors) counts only as quarantined, matching {@code CompositePeerLauncher}'s
|
||||
* explicit-spawn ordering: exhaustion quarantine takes priority when both are active.
|
||||
* quarantined, all cooling off, all model-off, all at capacity, all unreachable, or a mix. Each
|
||||
* candidate is counted into exactly one bucket (weight-excluded first, then quarantined, then
|
||||
* cooling off, then model-off) so a candidate excluded for more than one reason is never
|
||||
* double-counted — a candidate that is both quarantined (CB-578 stage B, backend exhausted) and
|
||||
* cooling off (fleetd #201 Unit 5, repeated backend errors) counts only as quarantined, matching
|
||||
* {@code CompositePeerLauncher}'s explicit-spawn ordering: exhaustion quarantine takes priority
|
||||
* when more than one applies.
|
||||
*/
|
||||
static PlacementException emptyException(PlacementContext ctx) {
|
||||
int weightExcluded = 0;
|
||||
@@ -53,17 +66,19 @@ final class PlacementPolicyUtil {
|
||||
int unreachable = 0;
|
||||
int quarantined = 0;
|
||||
int coolingOff = 0;
|
||||
int modelOff = 0;
|
||||
for (PlacementCandidate c : ctx.candidates()) {
|
||||
Integer cap = c.maxLoad();
|
||||
if (c.excluded()) {
|
||||
weightExcluded++;
|
||||
} else if (ctx.quarantined().contains(c.profile())) {
|
||||
quarantined++;
|
||||
} else if (ctx.coolingOff().contains(c.profile())) {
|
||||
coolingOff++;
|
||||
} else if (ctx.modelOff().contains(c.profile())) {
|
||||
modelOff++;
|
||||
} else if (ctx.unreachable().contains(c.profile())) {
|
||||
unreachable++;
|
||||
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
|
||||
} else if (atCap(ctx, c)) {
|
||||
atCap++;
|
||||
}
|
||||
}
|
||||
@@ -83,6 +98,10 @@ final class PlacementPolicyUtil {
|
||||
return new PlacementException(
|
||||
"all worker profiles are cooling off after repeated backend errors");
|
||||
}
|
||||
if (modelOff == total) {
|
||||
return new PlacementException(
|
||||
"all worker profiles name a model the operator has turned off");
|
||||
}
|
||||
if (atCap == total) {
|
||||
return new PlacementException("all worker profiles are at maxLoad");
|
||||
}
|
||||
@@ -92,8 +111,9 @@ final class PlacementPolicyUtil {
|
||||
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
|
||||
+ unreachable + " unreachable, " + quarantined + " quarantined, "
|
||||
+ coolingOff + " cooling off, "
|
||||
+ modelOff + " model-off, "
|
||||
+ weightExcluded + " weight-0, "
|
||||
+ (total - atCap - unreachable - quarantined - coolingOff - weightExcluded)
|
||||
+ (total - atCap - unreachable - quarantined - coolingOff - modelOff - weightExcluded)
|
||||
+ " remaining");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -690,14 +690,18 @@ public final class GitWorktrees implements Worktrees {
|
||||
* {@code fleet.seededSkillsNote}, readable with {@code git config --worktree --get-all
|
||||
* fleet.seededSkills}.
|
||||
*
|
||||
* <p><b>Claude Code specific by construction, not by a backend check here.</b> Only {@code
|
||||
* .claude/skills/<name>/SKILL.md} is a path any launcher reads today (opencode's equivalent is a
|
||||
* different shape under {@code .opencode/agent}, out of scope — see issue #362). This method
|
||||
* only copies files; like {@link #isolateToolSurface} — which neutralizes BOTH {@code .mcp.json}
|
||||
* and {@code opencode.json} unconditionally — it runs the same for every worktree regardless of
|
||||
* which backend ultimately spawns into it, because the backend is not yet chosen at {@link #add}
|
||||
* time. A seeded {@code .claude/skills/} directory in an opencode member's worktree is simply
|
||||
* never read by that launcher.
|
||||
* <p><b>Kind-blind by construction, not by a backend check here — this used to be a real gap
|
||||
* (fleetd #393).</b> This method only copies files; like {@link #isolateToolSurface} — which
|
||||
* neutralizes BOTH {@code .mcp.json} and {@code opencode.json} unconditionally — it runs the
|
||||
* same for every worktree regardless of which backend ultimately spawns into it, because no
|
||||
* caller of {@link #add} hands this class a kind to consult. Before fleetd #393, that made the
|
||||
* log line below a false claim of success for a {@code kind: opencode} member: opencode has no
|
||||
* built-in discovery of {@code .claude/skills/}, unlike the Claude Code CLI, so a seeded skill
|
||||
* never reached one. It now does — {@code OpenCodeLauncher#skillInstructionFiles} reads
|
||||
* whatever this method copied into {@code .claude/skills/} and appends each {@code SKILL.md} to
|
||||
* the generated {@code instructions[]} — but that delivery, and the log line that honestly
|
||||
* claims it (kind-aware, unlike this one), happens at the launcher, once the kind is actually
|
||||
* known, not here.
|
||||
*/
|
||||
private void seedSkills(String worktreePath) {
|
||||
if (memberSkillsSource == null) {
|
||||
@@ -739,7 +743,15 @@ public final class GitWorktrees implements Worktrees {
|
||||
if (detail.isEmpty()) {
|
||||
detail = "no skill folders found under " + source;
|
||||
}
|
||||
log.info("skill seeding: {} of {} candidate(s) from {} into {}/.claude/skills — {}",
|
||||
// fleetd #393: this only claims the copy step, deliberately — it cannot know the member
|
||||
// kind that will spawn into this worktree (see this method's own javadoc), so it must not
|
||||
// read as "and the member will act on it." Whether that is true depends on the kind: the
|
||||
// Claude Code CLI discovers .claude/skills/ on its own; OpenCodeLauncher logs its own
|
||||
// "skill delivery" line, once the kind is known, naming what it could and could not turn
|
||||
// into instructions[].
|
||||
log.info("skill seeding: {} of {} candidate(s) from {} into {}/.claude/skills — {} "
|
||||
+ "(whether the spawned member can act on this depends on its kind — see "
|
||||
+ "the launcher's own log for that)",
|
||||
seeded.size(), seeded.size() + kept.size(), source, worktreePath, detail);
|
||||
if (seeded.isEmpty()) {
|
||||
return;
|
||||
|
||||
@@ -10,6 +10,7 @@ import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.peer.PeerHandle;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
@@ -584,8 +585,65 @@ public final class SessionManager implements TurnListener {
|
||||
String ownerTerminal, WorktreeRequest wt,
|
||||
String sessionName, String resumeSessionId,
|
||||
MemberLifecycle.SlotReservation reservation) {
|
||||
String preResolvedProfile = (profile == null || profile.isBlank())
|
||||
? launcher.defaultProfile() : profile;
|
||||
// fleetd #425 rework (round 2): resolved through launcher.place(memberRole) — the same
|
||||
// candidate list, quarantine/cool-off/model-off filtering, and PlacementPolicy an unqualified
|
||||
// spawn of this role is actually judged against right now — never launcher.defaultProfile()
|
||||
// (MemberRole.DEV only, wrong for any other role) and never launcher.defaultProfileFor()
|
||||
// (the role's pool FIRST entry, blind to quarantine/cool-off/model-off: a first-round fix
|
||||
// used exactly this and regressed fleetd #429's "the fleet keeps working when a model is
|
||||
// turned off" guarantee — a quarantined or model-off pool-first profile made this throw
|
||||
// instead of routing around it, which an unqualified spawn is supposed to do). This same
|
||||
// resolved decision is reused below for repoRoot, parityOverlay, AND the spawn itself so the
|
||||
// worktree is always provisioned for the profile the member actually runs on — the two could
|
||||
// disagree before fleetd #425: this name picked repoRoot/overlay, but the spawn below passed
|
||||
// the ORIGINAL (blank) profile through to placement, which re-resolves live and can pick a
|
||||
// different profile if the pool changed between the two reads, or a genuinely different one
|
||||
// under weighted/round-robin placement.
|
||||
//
|
||||
// Round 1 of this rework fed the resolved name back into launcher.spawn(SpawnRequest) as an
|
||||
// EXPLICIT profile. That was a mistake this round corrects, and the mistake is not that the
|
||||
// two branches apply different checks — they are SUPPOSED to disagree: the routing branch a
|
||||
// blank spawn takes treats a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0
|
||||
// profile as a reason to fall through to the next candidate, while CompositePeerLauncher's
|
||||
// THROWING branch (enforceNotQuarantined/enforceNotCoolingOff/enforceMaxLoad/
|
||||
// enforceModelEnabled) treats naming that same profile explicitly as a reason to refuse
|
||||
// outright. That is correct: an operator who names a profile should get a refusal, not a
|
||||
// silent substitution onto a different backend. The mistake was turning a fall-through into
|
||||
// a refusal by accident — resolving a name via the routing side and then re-entering the
|
||||
// refusing side with it, for a placement the routing side had already approved by walking
|
||||
// past everything else.
|
||||
//
|
||||
// Before fleetd #435, this accident was reachable through maxLoad specifically: the default
|
||||
// `fixed` placement policy did not evaluate maxLoad at all for automatic selection, so an
|
||||
// at-cap pool-first profile that placement itself would have picked for a plain unqualified
|
||||
// spawn could die at enforceMaxLoad one call later, purely because this method's route to
|
||||
// the spawn passed through an explicit profile name — a failure a worktree-less unqualified
|
||||
// spawn never hit. fleetd #435 closed that gap (`fixed` now evaluates maxLoad exactly like
|
||||
// every other placement policy), so that specific failure can no longer happen — a
|
||||
// PlacementDecision this method resolves can no longer be at-cap in the first place. What
|
||||
// this round's fix still buys, now that maxLoad can no longer cause the accident: it keeps
|
||||
// the PlacementDecision from place() and hands it to launcher.spawn(SpawnRequest,
|
||||
// PlacementDecision) for an unqualified request, which spawns through the SAME routing
|
||||
// branch a blank spawn uses — no enforce* check is newly applied, and the window between the
|
||||
// placement decision and the spawn (in which the pool, a config reload, or another spawn
|
||||
// landing on the same profile could otherwise move the state) never reopens. An
|
||||
// explicitly-named profile still goes through launcher.spawn(SpawnRequest) and its throwing
|
||||
// branch, unchanged — that caller asked for one profile by name and still gets everything
|
||||
// enforceNotQuarantined/enforceNotCoolingOff/enforceMaxLoad/enforceModelEnabled decide about
|
||||
// it, refusal included.
|
||||
//
|
||||
// The one cost that remains, unchanged from round 1: an unqualified worktree-provisioned
|
||||
// spawn does not get CompositePeerLauncher's cross-candidate retry on a live
|
||||
// PeerUnreachableException raised by the backend itself at spawn time (a transport-level
|
||||
// failure placement cannot see in advance) — spawn(req, decision) commits to the one profile
|
||||
// place() already chose, the same way an explicit-profile spawn commits to its one name. That
|
||||
// trade is deliberate: a worktree provisioned for the wrong backend (the #425 hazard) is worse
|
||||
// than a spawn that fails cleanly and can be retried by the caller. Nothing else is lost:
|
||||
// maxLoad, quarantine, cool-off and model-off all behave identically whether or not a
|
||||
// worktree was requested — that agreement is the invariant this rework exists to hold.
|
||||
boolean unqualifiedProfile = profile == null || profile.isBlank();
|
||||
PlacementDecision decision = unqualifiedProfile ? launcher.place(memberRole) : new PlacementDecision(profile);
|
||||
String preResolvedProfile = decision.profile();
|
||||
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
|
||||
// daemon cwd → "."), never the raw args. A plain REST spawn supplies neither a requested
|
||||
// nor a caller cwd, so taking the first non-blank of those two yielded null and put
|
||||
@@ -608,7 +666,15 @@ public final class SessionManager implements TurnListener {
|
||||
// copies more files into the worktree after add() returns, so sharing the group any earlier
|
||||
// leaves those overlay files operator-owned and read-only for a different-uid member.
|
||||
worktrees.shareWithGroup(repoRoot, path);
|
||||
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, sessionName, resumeSessionId, memberRole));
|
||||
// fleetd #425: preResolvedProfile, not the original (possibly blank) profile — see the
|
||||
// comment above where it is resolved. The overlay/repoRoot above and the spawn here must
|
||||
// name the same profile. An unqualified request stays unqualified here and is honored via
|
||||
// the PlacementDecision already captured above (spawn(req, decision) — the routing branch,
|
||||
// no enforce* re-check); an explicitly-named profile still goes through the single-arg
|
||||
// spawn(req) and its throwing branch, exactly as before this rework.
|
||||
SpawnRequest spawnReq = new SpawnRequest(unqualifiedProfile ? null : preResolvedProfile,
|
||||
path, callerCwd, sessionName, resumeSessionId, memberRole);
|
||||
handle = unqualifiedProfile ? launcher.spawn(spawnReq, decision) : launcher.spawn(spawnReq);
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("spawn failed for profile={} role={} branch={} path={}: {}",
|
||||
preResolvedProfile, memberRole, branch, path, e.getMessage());
|
||||
@@ -636,7 +702,7 @@ public final class SessionManager implements TurnListener {
|
||||
}
|
||||
throw e;
|
||||
}
|
||||
String resolvedProfile = resolveProfile(handle, profile);
|
||||
String resolvedProfile = resolveProfile(handle, preResolvedProfile);
|
||||
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, path, callerCwd));
|
||||
long now = nowNanos.getAsLong();
|
||||
// CB-619: see the no-worktree path above — bind before recording, and store the returned
|
||||
|
||||
@@ -0,0 +1,150 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #395: {@code exhaustedPattern} (see {@link FleetConfig.Profile#exhaustedPattern}) is
|
||||
* deliberately opt-in — an unset one leaves usage-limit detection silently OFF for that profile,
|
||||
* and nothing quarantines its credential. {@link Fleetd#reportExhaustedPatternGap} must say so at
|
||||
* startup, naming every unarmed profile, and must never fire when every profile is armed. Mirrors
|
||||
* {@link MemberCredentialsGapReportTest}'s pattern, capturing the real log via a
|
||||
* {@link ListAppender}.
|
||||
*
|
||||
* <p>The 8-profile shape in {@link #theLiveEightProfileShapeWarnsExactlyTheSixUnarmedProfiles} is
|
||||
* the live {@code fleetd.yaml} shape measured 2026-09-10 (fleetd #395's own ticket): 6 of 8
|
||||
* profiles unarmed, 2 of those 6 ({@code opus}, {@code sonnet}) running on the operator's Claude
|
||||
* subscription. {@code fleetd.yaml} itself is gitignored and unavailable to this test, so the
|
||||
* shape is reproduced as a throwaway config in a {@code @TempDir} rather than read off disk.
|
||||
*/
|
||||
class ExhaustedPatternGapReportTest {
|
||||
|
||||
private static FleetConfig load(Path dir, String yaml) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml);
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
private static ListAppender<ILoggingEvent> attach() {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
return appender;
|
||||
}
|
||||
|
||||
private static void detach(ListAppender<ILoggingEvent> appender) {
|
||||
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
|
||||
}
|
||||
|
||||
/** Every profile name mentioned by a WARN-level log line, across every WARN this call produced. */
|
||||
private static List<String> warnMessages(ListAppender<ILoggingEvent> appender) {
|
||||
return appender.list.stream()
|
||||
.filter(e -> e.getLevel() == Level.WARN)
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.toList();
|
||||
}
|
||||
|
||||
@Test
|
||||
void theLiveEightProfileShapeWarnsExactlyTheSixUnarmedProfiles(@TempDir Path dir) throws Exception {
|
||||
// Reproduces the live shape measured 2026-09-10: 8 profiles, 2 armed (sol, terra), 6
|
||||
// unarmed (local, local-direct, gx, opus, sonnet, xf) — 2 of the unarmed 6 (opus, sonnet)
|
||||
// are subscription: true.
|
||||
FleetConfig cfg = load(dir, """
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
local-direct:
|
||||
baseUrl: http://gx01.gw:8000
|
||||
gx:
|
||||
kind: opencode
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
opus:
|
||||
subscription: true
|
||||
model: claude-opus-5
|
||||
sonnet:
|
||||
subscription: true
|
||||
model: claude-sonnet-5
|
||||
sol:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
exhaustedPattern: "The usage limit has been reached"
|
||||
terra:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
exhaustedPattern: "The usage limit has been reached"
|
||||
xf:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
""");
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
Fleetd.reportExhaustedPatternGap(cfg);
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
List<String> warns = warnMessages(appender);
|
||||
assertFalse(warns.isEmpty(), "6 of 8 profiles are unarmed — at least one WARN must fire");
|
||||
|
||||
// Exactly one WARN aggregates every unarmed profile, naming all 6 and none of the 2 armed.
|
||||
String aggregate = warns.stream()
|
||||
.filter(m -> m.contains("no exhaustedPattern configured"))
|
||||
.findFirst()
|
||||
.orElseThrow(() -> new AssertionError("expected an aggregate unarmed-profiles WARN: " + warns));
|
||||
for (String unarmed : List.of("local", "local-direct", "gx", "opus", "sonnet", "xf")) {
|
||||
assertTrue(aggregate.contains(unarmed), "aggregate WARN must name '" + unarmed + "': " + aggregate);
|
||||
}
|
||||
for (String armed : List.of("sol", "terra")) {
|
||||
assertFalse(aggregate.contains(armed), "aggregate WARN must NOT name armed profile '" + armed + "': " + aggregate);
|
||||
}
|
||||
|
||||
// A second, louder WARN calls out the subscription profiles specifically.
|
||||
String subscriptionWarn = warns.stream()
|
||||
.filter(m -> m.contains("metered Claude plan"))
|
||||
.findFirst()
|
||||
.orElseThrow(() -> new AssertionError("expected a subscription-specific WARN: " + warns));
|
||||
assertTrue(subscriptionWarn.contains("opus"), subscriptionWarn);
|
||||
assertTrue(subscriptionWarn.contains("sonnet"), subscriptionWarn);
|
||||
assertFalse(subscriptionWarn.contains("local-direct"),
|
||||
"the subscription WARN must not name a non-subscription profile: " + subscriptionWarn);
|
||||
}
|
||||
|
||||
@Test
|
||||
void everyProfileArmedProducesNoWarningAtAll(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = load(dir, """
|
||||
profiles:
|
||||
sol:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
exhaustedPattern: "The usage limit has been reached"
|
||||
terra:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
exhaustedPattern: "The usage limit has been reached"
|
||||
opus:
|
||||
subscription: true
|
||||
model: claude-opus-5
|
||||
exhaustedPattern: "5-hour limit reached"
|
||||
""");
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
Fleetd.reportExhaustedPatternGap(cfg);
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
assertTrue(warnMessages(appender).isEmpty(),
|
||||
"every profile is armed — a checker that warns anyway always fires: " + warnMessages(appender));
|
||||
}
|
||||
}
|
||||
@@ -20,6 +20,7 @@ import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
@@ -123,6 +124,11 @@ class FleetdBackendErrorSinkTest {
|
||||
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return Set.of();
|
||||
|
||||
@@ -0,0 +1,65 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #466 follow-up: {@code Fleetd.main} builds the daemon's one {@code BackendQuarantine}
|
||||
* from {@link dev.ltms.fleet.placement.BackendQuarantine#withEscalation(java.util.function.LongSupplier,
|
||||
* long)} — the escalating factory — rather than the plain two-argument constructor, which is still a
|
||||
* flat cooldown (kept for backward compatibility, see that class's doc). {@code
|
||||
* BackendQuarantineTest} proves {@code withEscalation} itself escalates, is ceilinged, and resets;
|
||||
* it says nothing about which one {@code main} actually calls.
|
||||
*
|
||||
* <p>Measured directly: reverting {@code main} to {@code new BackendQuarantine(System::nanoTime,
|
||||
* TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()))} — the pre-#466 flat call — compiles
|
||||
* with 0 errors and leaves the entire 1608-test suite (including every {@code BackendQuarantineTest}
|
||||
* case) green, because no other test constructs its {@code BackendQuarantine} through {@code main};
|
||||
* every one of them builds its own instance directly. That silent regression is exactly the shape
|
||||
* {@link FleetdLeadSeatWiringTest} and {@link FleetdCompletionResolverWiringTest} already guard
|
||||
* against for their own constructor arguments — this is the same class of gap for fleetd #466's
|
||||
* factory choice, following their approach.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
||||
* BackendQuarantine} and never runs {@code main} — a green result here proves only that the exact
|
||||
* text {@code main} calls {@code BackendQuarantine.withEscalation(...)} rather than the flat
|
||||
* constructor. It does not prove that call actually executes at startup (no test here starts the
|
||||
* daemon), and it does not prove the escalation reaches a real backend or credential — only
|
||||
* {@code BackendQuarantineTest} proves the factory's own behaviour, and only a live daemon proves
|
||||
* the wiring runs.
|
||||
*/
|
||||
class FleetdBackendQuarantineWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] main's BackendQuarantine local is still built from BackendQuarantine.withEscalation(...)")
|
||||
void mainStillWiresTheEscalatingQuarantineFactory() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,\n"
|
||||
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
|
||||
"Fleetd.main's BackendQuarantine local must still be built from "
|
||||
+ "BackendQuarantine.withEscalation(System::nanoTime, "
|
||||
+ "TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds())). Reverting to the flat "
|
||||
+ "two-argument constructor (fleetd #466's measured regression) compiles with 0 errors "
|
||||
+ "and leaves the whole suite green, including every BackendQuarantineTest case that "
|
||||
+ "proves the escalation itself works — this source check is what must go red instead. "
|
||||
+ "A reverted daemon would go back to retrying a weekly subscription limit on every "
|
||||
+ "flat ~30-minute cooldown, about 336 times across the week.");
|
||||
|
||||
// Negative form of the same check: the pre-#466 flat call, if it ever reappears at this
|
||||
// declaration, must not be mistaken for the escalating one by a looser positive-only check.
|
||||
assertFalse(source.contains(
|
||||
"BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,\n"
|
||||
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
|
||||
"main's BackendQuarantine local must never regress to the flat two-argument constructor");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,155 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* fleetd #416: {@code fleet_list}'s {@code CapacitySource.configuredProfiles} must enumerate the
|
||||
* <em>startup</em> profile set, not the live, hot-reloaded one.
|
||||
*
|
||||
* <p>{@code profiles} as a whole is a {@code DEFERRED} key ({@link ConfigRef#DEFERRED_KEYS}):
|
||||
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} once at construction, so a profile
|
||||
* only added to the hot-reloaded map can never actually be spawned. Before this fix, {@code
|
||||
* Fleetd.main} built {@code CapacitySource} with {@code () -> config.get().profiles().keySet()} —
|
||||
* the live map — so {@code fleet_list} would report a freshly hot-reloaded profile as available
|
||||
* ({@code free > 0}) while {@code fleet_spawn} on that same profile failed with
|
||||
* {@code unknown worker profile}. Measured on another host: adding a throwaway profile and letting
|
||||
* it hot-reload gave {@code fleet_list} -> {@code free: 3} and {@code fleet_spawn} ->
|
||||
* {@code error: unknown worker profile}.
|
||||
*
|
||||
* <p>This test needs a reload, the same reason {@link FleetdExhaustionDetectionArmedWiringTest}
|
||||
* does: at startup the two snapshots agree, so a test of only a newly started daemon would not
|
||||
* detect a live {@code config.get()} lookup for the set.
|
||||
*
|
||||
* <p><b>maxLoad must stay hot.</b> It is read off {@code config.get()} exactly like
|
||||
* {@code credentialId} ({@link ConfigRef} documents both as hot, "read live off the config
|
||||
* supplier ... exactly like weight/maxLoad"), so a reload that only changes an existing profile's
|
||||
* {@code maxLoad} — no add/remove — must still change what {@code fleet_list} reports without a
|
||||
* restart. A fix that freezes the whole {@code CapacitySource} against {@code cfg} (rather than
|
||||
* only its {@code configuredProfiles} set) would trade this bug for its mirror image and is pinned
|
||||
* wrong by {@link #reloadedMaxLoadStillChangesWhatFleetListReports}.
|
||||
*/
|
||||
class FleetdCapacitySourceWiringTest {
|
||||
|
||||
private static final String STARTUP = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: terra
|
||||
maxLoad: 3
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
private static final String WITH_NEW_PROFILE = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: terra
|
||||
maxLoad: 3
|
||||
ghost404:
|
||||
baseUrl: http://gx00.gw:8001
|
||||
model: ghost404
|
||||
maxLoad: 3
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
private static final String WITH_CHANGED_MAX_LOAD = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: terra
|
||||
maxLoad: 9
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
@Test
|
||||
@DisplayName("a profile present only in the live (hot-reloaded) config is NOT listed")
|
||||
void liveOnlyProfileIsNotListed(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, STARTUP);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = new ConfigRef(file, cfg);
|
||||
|
||||
Files.writeString(file, WITH_NEW_PROFILE);
|
||||
assertTrue(config.reload().applied());
|
||||
// The live snapshot now has the new profile — proves the reload really happened and this
|
||||
// test is not accidentally passing because nothing changed.
|
||||
assertTrue(config.get().profiles().containsKey("ghost404"));
|
||||
|
||||
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
|
||||
|
||||
assertFalse(source.configuredProfiles().get().contains("ghost404"),
|
||||
"a profile added only to the hot-reloaded config must not be listed by fleet_list — "
|
||||
+ "HerdrPeerLauncher never learns about it until a restart, so fleet_spawn on it "
|
||||
+ "would fail with 'unknown worker profile' while fleet_list claimed it free");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a profile present in the startup set IS listed")
|
||||
void startupProfileIsListed(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, STARTUP);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = new ConfigRef(file, cfg);
|
||||
|
||||
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
|
||||
|
||||
// fleetd #416, both-directions requirement: a test that only ever passes an empty/absent
|
||||
// startup set (the case above) cannot tell a correct lookup from one that is permanently
|
||||
// empty (e.g. a mutation replacing the supplier with Set::of). This is the direction that
|
||||
// fails if the fix regresses to reporting nothing at all.
|
||||
assertTrue(source.configuredProfiles().get().contains("terra"),
|
||||
"a profile present in the startup snapshot must still be listed by fleet_list");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a hot maxLoad edit still changes what fleet_list reports")
|
||||
void reloadedMaxLoadStillChangesWhatFleetListReports(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, STARTUP);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = new ConfigRef(file, cfg);
|
||||
|
||||
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
|
||||
assertEquals(3, source.maxLoad().apply("terra"),
|
||||
"sanity: maxLoad reads 3 from the startup config before any reload");
|
||||
|
||||
Files.writeString(file, WITH_CHANGED_MAX_LOAD);
|
||||
assertTrue(config.reload().applied());
|
||||
|
||||
assertEquals(9, source.maxLoad().apply("terra"),
|
||||
"maxLoad must stay hot — the SAME CapacitySource instance must reflect a reloaded "
|
||||
+ "maxLoad without a restart, exactly like credentialId. Freezing the whole "
|
||||
+ "CapacitySource against the startup snapshot (rather than only its "
|
||||
+ "configuredProfiles set) would trade fleetd #416 for its mirror image.");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,106 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertSame;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #474: proves the exact wiring {@code Fleetd.main} uses to construct its live {@code
|
||||
* ConfigRef} — {@code new ConfigRef(configPath, cfg, Fleetd::assertChartersNameOnlyRegisteredTools)}
|
||||
* — actually makes {@link ConfigRef#reload()} refuse a charter that names an MCP tool the server
|
||||
* does not register, the same way {@code Fleetd.main} itself refuses one at startup (see {@code
|
||||
* FleetdStartupValidationTest#mainRefusesACharterNamingAnUnregisteredTool}).
|
||||
*
|
||||
* <p>{@code dev.ltms.fleet.config.ConfigRefTest} pins the same behaviour through a locally-built
|
||||
* {@code Consumer<FleetConfig>} adapter that calls the same production {@code CharterToolSurface}
|
||||
* method, because that test lives in {@code dev.ltms.fleet.config} and cannot see {@code
|
||||
* Fleetd#assertChartersNameOnlyRegisteredTools} (package-private to {@code dev.ltms.fleet}). This
|
||||
* class is the companion proof that lives where the real method reference is visible, so the literal
|
||||
* expression {@code Fleetd::assertChartersNameOnlyRegisteredTools} — not just an equivalent — is
|
||||
* what gets exercised. {@code Fleetd.main} itself cannot be driven this far in a unit test: every
|
||||
* fixture in {@code FleetdStartupValidationTest} is deliberately invalid so {@code main} throws
|
||||
* before opening a socket, binding Javalin, or doing anything else with a real side effect, so a
|
||||
* test cannot get {@code main} far enough to hold a running daemon it could then reload — this test
|
||||
* builds the {@code ConfigRef} the same way {@code main} does and drives {@link ConfigRef#reload()}
|
||||
* directly instead, the same shape {@code FleetdExhaustionDetectionArmedWiringTest} and its
|
||||
* siblings already use for the rest of {@code Fleetd.main}'s wiring.
|
||||
*/
|
||||
class FleetdConfigRefCharterToolSurfaceWiringTest {
|
||||
|
||||
private static final String BASE = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
@Test
|
||||
void reloadRefusesACharterNamingAnUnregisteredToolThroughFleetdsOwnWiring(@TempDir Path dir)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, BASE + """
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through fleet_reply.
|
||||
""");
|
||||
ConfigRef config = new ConfigRef(f, FleetConfig.load(f),
|
||||
Fleetd::assertChartersNameOnlyRegisteredTools);
|
||||
FleetConfig before = config.get();
|
||||
|
||||
Files.writeString(f, BASE + """
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through bridge_send.
|
||||
""");
|
||||
ConfigRef.Outcome out = config.reload();
|
||||
|
||||
assertFalse(out.applied());
|
||||
assertNotNull(out.error());
|
||||
assertTrue(out.error().contains("dev"), out.error());
|
||||
assertTrue(out.error().contains("bridge_send"), out.error());
|
||||
assertSame(before, config.get());
|
||||
}
|
||||
|
||||
@Test
|
||||
void reloadAcceptsACharterNamingOnlyRegisteredToolsThroughFleetdsOwnWiring(@TempDir Path dir)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, BASE + """
|
||||
fleet:
|
||||
charters:
|
||||
dev: old charter
|
||||
""");
|
||||
ConfigRef config = new ConfigRef(f, FleetConfig.load(f),
|
||||
Fleetd::assertChartersNameOnlyRegisteredTools);
|
||||
|
||||
Files.writeString(f, BASE + """
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through fleet_reply.
|
||||
""");
|
||||
ConfigRef.Outcome out = config.reload();
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertEquals("config reloaded", out.summary());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,80 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #474 follow-up: {@code Fleetd.main} builds its live {@code ConfigRef} from the
|
||||
* three-argument constructor, {@code new ConfigRef(configPath, cfg,
|
||||
* Fleetd::assertChartersNameOnlyRegisteredTools)}, so a reload runs the same charter tool-surface
|
||||
* gate startup does (see {@link dev.ltms.fleet.config.ConfigRef}'s class doc, "fleetd #474" bullet).
|
||||
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} prove the
|
||||
* three-argument constructor and the {@code Fleetd.assertChartersNameOnlyRegisteredTools} adapter
|
||||
* work correctly together — both build their OWN {@code ConfigRef} with that constructor, so neither
|
||||
* proves {@code main} still CHOOSES the three-argument form over the plain two-argument {@code new
|
||||
* ConfigRef(configPath, cfg)}.
|
||||
*
|
||||
* <p>Measured directly: reverting {@code Fleetd.java}'s {@code config} local to the two-argument
|
||||
* constructor compiles with 0 errors and leaves the entire 1633-test suite green — including every
|
||||
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} case — because
|
||||
* neither of those tests constructs its {@code ConfigRef} through {@code main}; both build their own
|
||||
* instance directly, wired with the check by hand. That silent regression is exactly the shape
|
||||
* {@link FleetdBackendQuarantineWiringTest}, {@link FleetdLeadSeatWiringTest} and {@link
|
||||
* FleetdCompletionResolverWiringTest} already guard against for their own constructor arguments —
|
||||
* this class is the same class of gap for fleetd #474's {@code extraValidation} argument, following
|
||||
* their approach.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
||||
* ConfigRef} and never runs {@code main} — a green result here proves only that the exact text
|
||||
* {@code main} contains is the three-argument construction with {@code
|
||||
* Fleetd::assertChartersNameOnlyRegisteredTools}. It does not prove that call actually executes at
|
||||
* startup (no test here starts the daemon), and it does not prove the reload gate itself works —
|
||||
* only {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} prove the
|
||||
* behaviour; only a live daemon proves the wiring runs.
|
||||
*/
|
||||
class FleetdConfigRefWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] main still builds config from the three-argument ConfigRef constructor")
|
||||
void mainStillWiresTheThreeArgumentConfigRefConstructor() throws Exception {
|
||||
String source = fleetdSource();
|
||||
|
||||
// A broken read (wrong working directory, wrong path, a file that came back empty) would
|
||||
// make the assertFalse below pass vacuously — a "clean" negative check that actually
|
||||
// checked nothing. Guard against that first, with an anchor that has nothing to do with
|
||||
// this mutation, so a bad read fails loudly here instead of silently proving nothing below.
|
||||
assertTrue(source.contains("public final class Fleetd"),
|
||||
"fleetdSource() did not read anything usable — src/main/java/dev/ltms/fleet/Fleetd.java "
|
||||
+ "did not come back containing its own class declaration. The assertFalse below "
|
||||
+ "would pass vacuously on a broken read; fix the read before trusting this test.");
|
||||
|
||||
assertTrue(source.contains(
|
||||
"ConfigRef config = new ConfigRef(configPath, cfg, "
|
||||
+ "Fleetd::assertChartersNameOnlyRegisteredTools);"),
|
||||
"Fleetd.main's config local must still be built from the three-argument ConfigRef "
|
||||
+ "constructor, with Fleetd::assertChartersNameOnlyRegisteredTools as "
|
||||
+ "extraValidation. Reverting to the plain two-argument constructor (fleetd #474's "
|
||||
+ "measured M2 regression) compiles with 0 errors and leaves the whole suite green — "
|
||||
+ "including ConfigRefTest and FleetdConfigRefCharterToolSurfaceWiringTest, because "
|
||||
+ "neither builds its ConfigRef through main — this source check is what must go "
|
||||
+ "red instead. A reverted daemon would accept, through a reload with no restart, "
|
||||
+ "exactly the charter that #469/#474 already refuse at startup.");
|
||||
|
||||
// Negative form of the same check: the pre-#474 two-argument call, if it ever reappears at
|
||||
// this declaration, must not be mistaken for the three-argument one by a looser
|
||||
// positive-only check — this is the M2 mutation this test exists to kill.
|
||||
assertFalse(source.contains("ConfigRef config = new ConfigRef(configPath, cfg);"),
|
||||
"main's config local must never regress to the plain two-argument ConfigRef "
|
||||
+ "constructor — that drops the reload-path charter check (fleetd #474's measured "
|
||||
+ "M2 mutation) with no other test catching it");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,137 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #446: {@code exhaustionDetectionArmed} must describe the LIVE config, not a startup
|
||||
* snapshot — the opposite of what fleetd #404 asked for. #404 pinned {@code exhaustedPattern} as
|
||||
* compiled once at startup (matching {@code CompletionResolver}'s then-frozen behaviour); #446's
|
||||
* whole point is that both sides — the classification {@code CompletionResolver} enforces and the
|
||||
* {@code exhaustionDetectionArmed} field this reports — now read the SAME live {@link
|
||||
* LiveExhaustedPatterns} instance, so a reload arms or disarms detection with no restart, and the
|
||||
* report can never disagree with the behaviour (the fleetd #404 rule, now upheld for real instead
|
||||
* of by freezing both sides).
|
||||
*
|
||||
* <p>This test needs a reload. At startup the two snapshots agree, so a test of only a newly
|
||||
* started daemon would not detect a live {@code config.get()} lookup in the report field.
|
||||
*
|
||||
* <p><b>The hotness proof criterion 1 demands:</b> run this against the fixed code (green) and
|
||||
* against the pre-#446 compile site — {@code Fleetd.main}'s {@code exhaustedPatternsByProfile}
|
||||
* compiled once into a {@code Map<String, Pattern>} at startup, handed to {@code quarantineSource}
|
||||
* — and show it fails there. See the class doc history above: that shape is exactly what
|
||||
* {@code reloadedPatternArmsDetectionWithNoRestart} below is written to catch, and it is the test
|
||||
* that used to assert the opposite (see git history for this file's pre-#446 version, which
|
||||
* asserted {@code assertFalse(...)} on the identical scenario this now asserts {@code assertTrue}
|
||||
* on).
|
||||
*/
|
||||
class FleetdExhaustionDetectionArmedWiringTest {
|
||||
|
||||
private static final String NO_PATTERN = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: terra
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
private static final String WITH_PATTERN = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: terra
|
||||
exhaustedPattern: "usage limit"
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
@Test
|
||||
@DisplayName("reloading a profile's exhaustedPattern IN arms exhaustionDetectionArmed, no restart")
|
||||
void reloadedPatternArmsDetectionWithNoRestart(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, NO_PATTERN);
|
||||
ConfigRef config = new ConfigRef(file, FleetConfig.load(file));
|
||||
|
||||
FleetMcp.QuarantineSource before = Fleetd.quarantineSource(config, BackendQuarantine.none(),
|
||||
new LiveExhaustedPatterns(() -> config.get().profiles()), Map.of());
|
||||
assertFalse(before.exhaustedPatternArmed().apply("terra"),
|
||||
"no exhaustedPattern configured yet — must report unarmed");
|
||||
|
||||
Files.writeString(file, WITH_PATTERN);
|
||||
assertTrue(config.reload().applied());
|
||||
assertTrue(config.get().profiles().get("terra").hasExhaustedPattern());
|
||||
|
||||
// Re-read exhaustedPatternArmed off the SAME QuarantineSource built BEFORE the reload — no
|
||||
// new object, no new wiring — to prove the field itself is live, not merely that a freshly
|
||||
// built source would be. This is the exact axis the pre-#446 code fails on: the same
|
||||
// Function<String,Boolean> lambda, called again after a reload, sees the new answer only if
|
||||
// it re-reads config.get() on every call rather than a value captured earlier.
|
||||
assertTrue(before.exhaustedPatternArmed().apply("terra"),
|
||||
"exhaustionDetectionArmed must read the LIVE config on every call — fleetd #446 made "
|
||||
+ "this hot; it must reflect a reload with no restart and no new QuarantineSource");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("reloading exhaustedPattern OUT disarms exhaustionDetectionArmed, no restart")
|
||||
void reloadedPatternRemovalDisarmsDetectionWithNoRestart(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, WITH_PATTERN);
|
||||
ConfigRef config = new ConfigRef(file, FleetConfig.load(file));
|
||||
|
||||
FleetMcp.QuarantineSource source = Fleetd.quarantineSource(config, BackendQuarantine.none(),
|
||||
new LiveExhaustedPatterns(() -> config.get().profiles()), Map.of());
|
||||
assertTrue(source.exhaustedPatternArmed().apply("terra"),
|
||||
"exhaustedPattern is configured from the start — must report armed");
|
||||
|
||||
Files.writeString(file, NO_PATTERN);
|
||||
assertTrue(config.reload().applied());
|
||||
|
||||
assertFalse(source.exhaustedPatternArmed().apply("terra"),
|
||||
"removing exhaustedPattern and reloading must disarm detection with no restart — an "
|
||||
+ "operator turning detection off (e.g. while debugging a false positive) "
|
||||
+ "must see that reflected immediately, exactly like arming it is");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a profile with no exhaustedPattern configured is reported unarmed")
|
||||
void aProfileWithNoPatternIsUnarmed(@TempDir Path dir) throws Exception {
|
||||
// fleetd #404's second direction, carried forward: this test alone fails if
|
||||
// exhaustedPatternArmed degenerates to a constant `profile -> true` — it needs at least one
|
||||
// profile that is genuinely unarmed to catch that.
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, WITH_PATTERN);
|
||||
ConfigRef config = new ConfigRef(file, FleetConfig.load(file));
|
||||
|
||||
FleetMcp.QuarantineSource source = Fleetd.quarantineSource(config, BackendQuarantine.none(),
|
||||
new LiveExhaustedPatterns(() -> config.get().profiles()), Map.of());
|
||||
|
||||
assertTrue(source.exhaustedPatternArmed().apply("terra"),
|
||||
"a profile whose exhaustedPattern is configured must report armed");
|
||||
assertFalse(source.exhaustedPatternArmed().apply("sonnet"),
|
||||
"a profile absent from config entirely must not report armed");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,175 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.HashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #446 follow-up (round 3): round 2's {@link FleetdUsageLimitFixWarningTest} calls {@link
|
||||
* Fleetd#usageLimitFixWarning}/{@link Fleetd#usageLimitFixWarningNoModel} directly, which pins
|
||||
* the two methods' TEXT but cannot see whether {@link Fleetd#exhaustionSink}'s {@code log.warn}
|
||||
* call actually invokes either one, or invokes the right one. A mutation battery run against
|
||||
* round 2's merge (1592 green tests) proved both gaps real:
|
||||
*
|
||||
* <ul>
|
||||
* <li><b>Cell A</b> — replacing the whole ternary result with a literal string
|
||||
* ({@code "usage-limit fix: MUTANT"}) left every test green;</li>
|
||||
* <li><b>Cell B</b> — swapping the ternary's two branches, so a profile WITH a {@code model:}
|
||||
* gets the no-model fallback message and vice versa, was measured unmeasured by the lead
|
||||
* but the existing test file shows no call that would notice it either.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>This class drives {@link Fleetd#exhaustionSink} — the actual factory {@code main} wires,
|
||||
* since round 3's extraction pulled it out of the inline lambda for exactly this reason — and
|
||||
* asserts on the REAL text a {@link ListAppender} attached to {@code Fleetd}'s own logger
|
||||
* captures, mirroring {@code ExhaustedPatternGapReportTest}'s idiom. Chosen over a source-text
|
||||
* assertion (the {@code FleetMcpAuthzTest} {@code everyRegisteredToolHasItsHandlerActionPinned}
|
||||
* idiom) because that route can prove Cell A (the ternary is still there, still built from the
|
||||
* two named methods) but not Cell B (which branch a given profile actually reaches) — this route
|
||||
* answers both from one mechanism, since it reads what the sink actually logged for each shape of
|
||||
* profile.
|
||||
*
|
||||
* <p>Each test filters for the WARN line starting {@code "usage-limit fix:"} specifically — {@link
|
||||
* Fleetd#exhaustionSink} also logs a separate {@code "credential '...' quarantined for..."} WARN
|
||||
* on every call, and asserting against the wrong one would pass or fail for the wrong reason. That
|
||||
* prefix survives even under the M5 mutation above (the mutant keeps the tag, only the body
|
||||
* becomes {@code "MUTANT"}), so the filter itself is not what either mutation defeats.
|
||||
*/
|
||||
class FleetdExhaustionSinkWarningTest {
|
||||
|
||||
private static final String YAML = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: claude-opus-5
|
||||
gx:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
private static SessionManager emptyRosterSessions() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
FleetConfig.Profile dummy = new FleetConfig.Profile(
|
||||
"dummy", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
|
||||
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(dummy.profile(), dummy), dummy.profile(), _ -> "tok");
|
||||
// Never acquires a session — exhaustionSink's roster lookup is expected to miss and fall
|
||||
// back to the profileHint argument, exactly like OpenCodeLauncher's real call site does
|
||||
// (fleetd #234) — so the sink under test never needs a populated roster.
|
||||
return new SessionManager(launcher);
|
||||
}
|
||||
|
||||
private static ListAppender<ILoggingEvent> attach() {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
return appender;
|
||||
}
|
||||
|
||||
private static void detach(ListAppender<ILoggingEvent> appender) {
|
||||
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
|
||||
}
|
||||
|
||||
/** The one WARN line {@code exhaustionSink} builds from the ternary under test, or {@code null}. */
|
||||
private static String fixWarning(ListAppender<ILoggingEvent> appender) {
|
||||
List<String> warns = appender.list.stream()
|
||||
.filter(e -> e.getLevel() == Level.WARN)
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.filter(m -> m.startsWith("usage-limit fix:"))
|
||||
.toList();
|
||||
assertTrue(warns.size() == 1,
|
||||
"expected exactly one 'usage-limit fix:' WARN per onExhausted call, got " + warns.size()
|
||||
+ ": " + warns);
|
||||
return warns.getFirst();
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a profile WITH a configured model gets the actionable models.allow fix, not the fallback")
|
||||
void profileWithModelGetsTheActionableFix(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, YAML);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = ConfigRef.fixed(cfg);
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
Map<String, String> reasonByCredential = new HashMap<>();
|
||||
ExhaustionSink sink = Fleetd.exhaustionSink(emptyRosterSessions(), config, quarantine,
|
||||
reasonByCredential, cfg);
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
sink.onExhausted("term_x", "The usage limit has been reached", "terra");
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
String warning = fixWarning(appender);
|
||||
assertTrue(warning.contains("terra"), "must name the profile: " + warning);
|
||||
assertTrue(warning.contains("claude-opus-5"), "must name the model: " + warning);
|
||||
assertTrue(warning.contains("enabled: false"), "must name the action: " + warning);
|
||||
assertTrue(warning.contains("models.allow"), "must name where the action goes: " + warning);
|
||||
assertFalse(warning.contains("has no model: configured"),
|
||||
"a profile WITH a model must not get the no-model fallback text: " + warning);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a profile with NO configured model gets the fallback, never a fix that does not exist")
|
||||
void profileWithNoModelGetsTheFallbackNotAFix(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, YAML);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = ConfigRef.fixed(cfg);
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
Map<String, String> reasonByCredential = new HashMap<>();
|
||||
ExhaustionSink sink = Fleetd.exhaustionSink(emptyRosterSessions(), config, quarantine,
|
||||
reasonByCredential, cfg);
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
sink.onExhausted("term_y", "The usage limit has been reached", "gx");
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
String warning = fixWarning(appender);
|
||||
assertTrue(warning.contains("gx"), "must name the profile: " + warning);
|
||||
assertTrue(warning.contains("has no model: configured"),
|
||||
"must say there is no model to gate by name: " + warning);
|
||||
assertFalse(warning.contains("enabled: false"),
|
||||
"a profile with no model: configured has no models.allow entry to flip — must "
|
||||
+ "not claim one exists: " + warning);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,89 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #480 Unit A, hard requirement 6: pin {@code Fleetd.main}'s construction of {@link
|
||||
* dev.ltms.fleet.lead.LeadRollover} with a source-text assertion, mirroring {@code
|
||||
* FleetdCompletionResolverWiringTest}'s pattern — five log-only reporters in {@code Fleetd.main}
|
||||
* already survived mutation batteries this exact way (fleetd #415's extraction antidote note).
|
||||
*
|
||||
* <p>What this class still covers, and what it never claimed to. {@code LeadRolloverTest}
|
||||
* constructs its own {@code LeadRollover} directly (as every prior test of an extracted factory
|
||||
* does) with a hand-built lookup, so a mutation that deletes the {@code leadRollover(...)} call
|
||||
* from {@code main} — or replaces one of its arguments with something that still compiles, e.g.
|
||||
* {@code router.leadAgents()} swapped for {@code null}, or the whole assignment swapped for a bare
|
||||
* {@code null} literal — leaves every behavioural test green. This is a plain string read, guarded
|
||||
* by an unrelated anchor assertion so a broken or empty file read cannot pass as a real change.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
||||
* LeadRollover} and never runs {@code main}. It pins the {@code leadRollover(...)} CALL SITE's
|
||||
* argument list — that {@code main} still passes {@code leads} at all — never what the factory
|
||||
* DOES with that argument once inside its own body.
|
||||
*
|
||||
* <p><b>Correction (fleetd #480 relative-handover-path follow-up): that gap used to be real, and
|
||||
* now is not — but not here.</b> This class's javadoc previously claimed "no behavioural test can
|
||||
* catch this wiring dropping out" for the whole factory, including the lambda {@code
|
||||
* leadRollover(...)} builds internally (terminal → lead name → {@code Leader.cwd()}). That claim
|
||||
* was proven true at the time — mutating that lambda's body to {@code String leadName = null;}
|
||||
* (always "no lead found", which silently reintroduces the daemon-cwd bug this ticket fixes) left
|
||||
* the full suite green, {@code Tests run: 1669, Failures: 0}. It is no longer true: {@code
|
||||
* FleetdLeadRolloverWorkspaceLookupTest} now calls {@code Fleetd.leadRollover(...)} directly with a
|
||||
* real {@link dev.ltms.fleet.config.ConfigRef} built from a temp {@code fleetd.yaml}, and fails
|
||||
* against that exact one-line mutation. So: THIS class still covers only the call site's argument
|
||||
* list; {@code FleetdLeadRolloverWorkspaceLookupTest} is what now covers the lambda's body. Neither
|
||||
* one subsumes the other — keep both.
|
||||
*/
|
||||
class FleetdLeadRolloverWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] unrelated anchor: Fleetd.java still declares the Fleetd class")
|
||||
void unrelatedAnchorStillPresent() throws Exception {
|
||||
// Guards the two assertions below: without this, a bad read (empty string, wrong file,
|
||||
// truncated file) could vacuously fail to contain the leadRollover(...) call too, and a
|
||||
// test that only asserts "contains X" would report a false pass for the wrong reason if X
|
||||
// happened to match. Asserting an unrelated, structurally distant string first proves the
|
||||
// read actually pulled real file content.
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains("public final class Fleetd"),
|
||||
"sanity anchor failed — the file read did not return real Fleetd.java source; the "
|
||||
+ "leadRollover(...) wiring assertions below cannot be trusted until this passes");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] main still constructs LeadRollover via the leadRollover(...) factory, exactly as heartbeat is constructed")
|
||||
void mainStillCallsTheLeadRolloverFactory() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"LeadRollover leadRollover = leadRollover(cfg, router.leadAgents(), config, leads);"),
|
||||
"Fleetd.main must still assign `LeadRollover leadRollover = leadRollover(cfg, "
|
||||
+ "router.leadAgents(), config, leads);`. Dropping this call, or swapping one of "
|
||||
+ "its arguments for something that still compiles (e.g. null in place of "
|
||||
+ "router.leadAgents()), leaves every behavioural test green — this source check is "
|
||||
+ "what must go red instead. fleetd #480 correction 2 deliberately dropped "
|
||||
+ "primaryRegistry from this call — see LeadRollover's class javadoc for why a "
|
||||
+ "single-slot lookup was wrong here. The fleetd #480 relative-handover-path "
|
||||
+ "follow-up added `leads` (terminal → lead name) so the factory can resolve a "
|
||||
+ "relative handoverPath against the calling lead's own workspace.");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] the leadRollover(...) factory itself gates construction on cfg.leadRollover() != null")
|
||||
void factoryGatesOnConfigPresence() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains("if (cfg.leadRollover() == null) {"),
|
||||
"Fleetd.leadRollover(...) must refuse to construct a LeadRollover when the "
|
||||
+ "leadRollover: block is absent — an upgraded daemon must never silently acquire "
|
||||
+ "the ability to clear the lead's own pane. See LeadHeartbeatLoop's construction "
|
||||
+ "gate (cfg.leadHeartbeat() != null) for the pattern this mirrors.");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,163 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.lead.LeadRollover;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.HashMap;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
|
||||
/**
|
||||
* fleetd #480 relative-handover-path follow-up, correction round: a BEHAVIOURAL test of {@link
|
||||
* Fleetd#leadRollover}'s own body — the terminal → lead-name → {@code Leader.cwd()} lookup it
|
||||
* builds — not another source-text pin.
|
||||
*
|
||||
* <p>{@code FleetdLeadRolloverWiringTest} (a plain string read) still earns its keep: it pins the
|
||||
* call site's ARGUMENT LIST, so a mutation that drops {@code leads} back out of the call, or
|
||||
* swaps it for {@code Map::of}, still goes red there. But nothing before this class exercised the
|
||||
* LAMBDA BODY {@code leadRollover(...)} builds — the terminal-to-workspace {@code
|
||||
* Function<String,String>} that {@link LeadRollover#open} actually calls. Every prior test either
|
||||
* exercised {@code LeadRollover} directly with a hand-built lookup ({@code LeadRolloverTest}), or
|
||||
* read source text without ever calling the factory ({@code FleetdLeadRolloverWiringTest}) — so a
|
||||
* mutation that breaks the lookup ITSELF (e.g. always resolving no lead, or reading a snapshot
|
||||
* instead of the live supplier) left every existing test green while the daemon's real wiring
|
||||
* silently reintroduced the exact bug this whole ticket fixes: a relative {@code handoverPath}
|
||||
* resolving against the daemon's own {@code cwd} instead of the calling lead's.
|
||||
*
|
||||
* <p>{@code Fleetd.leadRollover(...)} is package-private and {@code static}, so this test — living
|
||||
* in the same {@code dev.ltms.fleet} package — calls it directly, exactly the way {@code
|
||||
* FleetdExhaustionDetectionArmedWiringTest} and {@code FleetdCapacitySourceWiringTest} already call
|
||||
* other package-private startup factories with a real {@link ConfigRef} built from a temp {@code
|
||||
* fleetd.yaml} (never the gitignored live one).
|
||||
*
|
||||
* <p><b>Proved against the mutation it exists to catch.</b> Before this class was added, mutating
|
||||
* {@code leadRollover(...)}'s lambda body to {@code String leadName = null;} (always "no lead
|
||||
* found", which forces every relative {@code handoverPath} onto the {@code user.dir} fallback —
|
||||
* i.e. the original bug) left the full suite green: {@code Tests run: 1669, Failures: 0}. With
|
||||
* {@link #relativeHandoverPathResolvesAgainstTheLeadsConfiguredCwd()} added, the same one-line
|
||||
* mutation now fails that test (it asserts the resolved path equals the configured lead's {@code
|
||||
* cwd}, which the mutant can never produce) — see this ticket's fleet_reply history for both runs.
|
||||
*/
|
||||
class FleetdLeadRolloverWorkspaceLookupTest {
|
||||
|
||||
private static AgentControl fakeAgents() {
|
||||
return new AgentControl(new FakeHerdr());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] Fleetd.leadRollover(...) resolves a relative handoverPath against "
|
||||
+ "the CALLING lead's configured cwd, not the daemon's own working directory")
|
||||
void relativeHandoverPathResolvesAgainstTheLeadsConfiguredCwd(@TempDir Path dir) throws Exception {
|
||||
Path leadCwd = dir.resolve("lead-workspace");
|
||||
Files.createDirectories(leadCwd);
|
||||
Path yaml = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
cwd: "%s"
|
||||
leadRollover:
|
||||
handoverPath: handover.md
|
||||
""".formatted(leadCwd.toString()));
|
||||
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
|
||||
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
|
||||
() -> Map.of("term_opus", "opus"));
|
||||
assertNotNull(rollover, "leadRollover: is present in the loaded config, so the factory "
|
||||
+ "must construct an object");
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open("term_opus", "test");
|
||||
|
||||
String expected = leadCwd.resolve("handover.md").normalize().toString();
|
||||
assertEquals(expected, pending.handoverPath(),
|
||||
"a relative handoverPath must resolve against the CALLING lead's configured cwd "
|
||||
+ "through the REAL Fleetd.leadRollover(...) wiring — not the daemon's own "
|
||||
+ "working directory. This is the exact axis that was proven uncovered: "
|
||||
+ "mutating the factory's lambda body to always report \"no lead found\" "
|
||||
+ "left every prior test green.");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] a terminal not present in the live lead-terminal map falls back "
|
||||
+ "to the daemon's own user.dir")
|
||||
void terminalNotInLiveMapFallsBackToUserDir(@TempDir Path dir) throws Exception {
|
||||
Path yaml = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
leadRollover:
|
||||
handoverPath: handover.md
|
||||
""");
|
||||
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
|
||||
|
||||
// No lead has been discovered yet — exactly the real shape of a lead the live tab scan
|
||||
// has not yet scanned, or one with no fleet.leaders entry at all.
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config, Map::of);
|
||||
assertNotNull(rollover);
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open("term_unknown", "test");
|
||||
|
||||
String expected = Path.of(System.getProperty("user.dir")).resolve("handover.md")
|
||||
.normalize().toString();
|
||||
assertEquals(expected, pending.handoverPath(),
|
||||
"a lead not present in the live terminal→name map must fall back to the daemon's "
|
||||
+ "own user.dir — the same fallback LeadLauncher#launch already uses for a "
|
||||
+ "lead with no configured cwd. This is deliberate, pinned behaviour, not an "
|
||||
+ "accident of the null-check chain.");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the terminal→lead-name lookup is read LIVE on every open() call, "
|
||||
+ "never snapshotted at Fleetd.leadRollover(...) construction time")
|
||||
void workspaceLookupIsReadLiveNotSnapshotted(@TempDir Path dir) throws Exception {
|
||||
Path leadCwd = dir.resolve("lead-workspace");
|
||||
Files.createDirectories(leadCwd);
|
||||
Path yaml = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
cwd: "%s"
|
||||
leadRollover:
|
||||
handoverPath: handover.md
|
||||
""".formatted(leadCwd.toString()));
|
||||
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
|
||||
|
||||
// EMPTY at the moment leadRollover(...) is called. A lookup captured (snapshotted) here
|
||||
// instead of read live through the supplier on every call would never see the entry added
|
||||
// below — exactly the natural mistake to make, since leads are discovered by a live tab
|
||||
// scan that runs AFTER this factory is constructed at startup.
|
||||
Map<String, String> liveLeadTerminals = new HashMap<>();
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
|
||||
() -> liveLeadTerminals);
|
||||
assertNotNull(rollover);
|
||||
|
||||
// The lead is "discovered" only now — mutate the SAME backing map the supplier reads from.
|
||||
liveLeadTerminals.put("term_opus", "opus");
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open("term_opus", "test");
|
||||
|
||||
String expected = leadCwd.resolve("handover.md").normalize().toString();
|
||||
assertEquals(expected, pending.handoverPath(),
|
||||
"the terminal→lead-name lookup must be read LIVE on every open() call — a lead "
|
||||
+ "discovered by the tab scan AFTER Fleetd.leadRollover(...) was constructed "
|
||||
+ "must still resolve correctly, not only one that was already live at "
|
||||
+ "construction time");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,62 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up: {@code Fleetd.modelGateCoverageLine} is the startup-log counterpart of
|
||||
* {@code exhaustedPatternCoverageLine}/{@code errorPatternCoverageLine} — see {@code
|
||||
* FleetdPatternCoverageLineTest} for the identical shape this follows — except here there is no
|
||||
* {@code UnsetMeaning} choice for a caller to get backwards: {@link
|
||||
* PeerLauncher.ModelGateState#configured()} already states, unambiguously, whether an empty
|
||||
* {@link PeerLauncher.ModelGateState#off()} means "no {@code models:} block to gate with at all"
|
||||
* or "a block armed and currently reporting zero off". This class proves {@code
|
||||
* modelGateCoverageLine} words those two states — plus the third, N off — distinctly, so a
|
||||
* mutation that made it ignore {@code configured()} either way is caught here.
|
||||
*/
|
||||
class FleetdModelGateCoverageLineTest {
|
||||
|
||||
@Test
|
||||
@DisplayName("no models: block reports not configured, distinct from armed-with-zero")
|
||||
void noModelsBlockReportsNotConfigured() {
|
||||
String line = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.notConfigured());
|
||||
assertEquals("not configured (no models: block — nothing is gated, and nothing can be)", line);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a models: block armed with nothing off reports armed, distinct from not configured")
|
||||
void armedWithNothingOffReportsArmed() {
|
||||
String line = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of()));
|
||||
assertEquals("armed (models: block present; 0 models currently turned off)", line);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a models: block with N off names the off models")
|
||||
void armedWithModelsOffNamesThem() {
|
||||
String line = Fleetd.modelGateCoverageLine(
|
||||
PeerLauncher.ModelGateState.armed(Set.of("deepseek-v4-flash", "claude-opus-9000")));
|
||||
assertEquals("armed (2 model(s) turned off: [claude-opus-9000, deepseek-v4-flash])", line);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("the three states produce pairwise-distinct wording for the same empty-looking input")
|
||||
void theThreeStatesProduceDistinctWording() {
|
||||
String notConfigured = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.notConfigured());
|
||||
String armedZero = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of()));
|
||||
String armedOne = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of("x")));
|
||||
|
||||
// Pinned individually above; restated here so this test alone still catches a regression
|
||||
// even if one of the three tests above were ever deleted — the exact FleetdPatternCoverageLineTest
|
||||
// pattern, adapted from "two keys" to "three states of one gate".
|
||||
assertNotEquals(notConfigured, armedZero,
|
||||
"collapsing 'no models: block' into 'armed, zero off' is fleetd #422 follow-up's exact defect");
|
||||
assertNotEquals(armedZero, armedOne);
|
||||
assertNotEquals(notConfigured, armedOne);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,67 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
|
||||
/**
|
||||
* fleetd #415 (review follow-up): {@code CompletionResolverTest} proves {@code coverage()} words
|
||||
* {@code UnsetMeaning.OFF} and {@code UnsetMeaning.BUILT_IN_DEFAULT} correctly — but every one of
|
||||
* those tests supplies the meaning itself. That proves the enum's wording, never that {@code
|
||||
* Fleetd} pairs the right meaning with the right pattern key. That pairing is #415's actual
|
||||
* defect: {@code coverage()} had no way to know what unset meant for its key, so the fix moved
|
||||
* the fact to the caller — and nothing yet proved the caller states it correctly.
|
||||
*
|
||||
* <p><b>Measured or it didn't happen:</b> swapping the two {@code UnsetMeaning} arguments at
|
||||
* {@code Fleetd}'s two coverage call sites — giving {@code exhaustedPattern} the built-in-default
|
||||
* wording and {@code errorPattern} the off wording, #415's exact defect with the keys exchanged —
|
||||
* compiled with 0 errors and left all 1506 existing tests green. This class exists to turn that
|
||||
* swap red.
|
||||
*
|
||||
* <p>It calls {@link Fleetd#exhaustedPatternCoverageLine} and {@link Fleetd#errorPatternCoverageLine}
|
||||
* directly rather than reading {@code Fleetd.java} as source text (the shape {@code
|
||||
* FleetdCompletionResolverWiringTest} uses for a different wiring gap): those two methods are the
|
||||
* extracted call sites {@code main} actually invokes, following the same {@code static} factory +
|
||||
* dedicated-test pattern as {@link Fleetd#capacitySource} and {@link Fleetd#worktreeBranchLookup}.
|
||||
*/
|
||||
class FleetdPatternCoverageLineTest {
|
||||
|
||||
private static final Set<String> PROFILES = Set.of("terra", "gx10");
|
||||
|
||||
@Test
|
||||
@DisplayName("exhaustedPatternCoverageLine says off when no profile configures exhaustedPattern")
|
||||
void exhaustedPatternCoverageLineSaysOffWhenNoProfileConfiguresIt() {
|
||||
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
|
||||
Fleetd.exhaustedPatternCoverageLine(PROFILES, Set.of()));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("errorPatternCoverageLine says built-in default when no profile configures errorPattern")
|
||||
void errorPatternCoverageLineSaysBuiltInDefaultWhenNoProfileConfiguresIt() {
|
||||
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
|
||||
+ "profiles: [gx10, terra])",
|
||||
Fleetd.errorPatternCoverageLine(PROFILES, Set.of()));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("the two keys produce different wording for the identical empty-coverage input")
|
||||
void theTwoKeysProduceDifferentWordingForTheSameEmptyInput() {
|
||||
String exhaustedLine = Fleetd.exhaustedPatternCoverageLine(PROFILES, Set.of());
|
||||
String errorLine = Fleetd.errorPatternCoverageLine(PROFILES, Set.of());
|
||||
|
||||
// Pinned individually above; restated here so this test alone still catches a swap even
|
||||
// if one of the two tests above were ever deleted.
|
||||
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
|
||||
exhaustedLine);
|
||||
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
|
||||
+ "profiles: [gx10, terra])", errorLine);
|
||||
assertNotEquals(exhaustedLine, errorLine,
|
||||
"swapping which UnsetMeaning pairs with which pattern key at Fleetd's call sites "
|
||||
+ "must be caught here — that pairing, not coverage()'s own wording in isolation, "
|
||||
+ "is fleetd #415's actual defect");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,79 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* Proves {@link Fleetd#main(String[])} calls every startup report before validation aborts startup.
|
||||
* The invalid non-loopback bind makes {@link FleetConfig#validateAll()} throw before {@code main}
|
||||
* can open the herdr socket or bind a port. The fixture also triggers every report, so removing any
|
||||
* one call from {@code main} leaves its expected log line absent.
|
||||
*/
|
||||
class FleetdStartupReportTest {
|
||||
|
||||
private static Level originalLevel;
|
||||
|
||||
private static ListAppender<ILoggingEvent> attach() {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
originalLevel = logger.getLevel();
|
||||
logger.setLevel(Level.INFO);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
return appender;
|
||||
}
|
||||
|
||||
private static void detach(ListAppender<ILoggingEvent> appender) {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
logger.detachAppender(appender);
|
||||
logger.setLevel(originalLevel);
|
||||
}
|
||||
|
||||
private static boolean contains(ListAppender<ILoggingEvent> appender, String fragment) {
|
||||
return appender.list.stream()
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.anyMatch(message -> message.contains(fragment));
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainReportsEveryStartupGapBeforeValidationAborts(@TempDir Path dir) throws Exception {
|
||||
Path config = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(config, """
|
||||
bind:
|
||||
host: 0.0.0.0
|
||||
port: 8765
|
||||
profiles:
|
||||
worker:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
gitTokenEnv: GITEA_TOKEN
|
||||
""");
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
assertThrows(IllegalStateException.class, () -> Fleetd.main(new String[]{config.toString()}));
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
assertTrue(contains(appender, "startup git host GITEA_HOST:"),
|
||||
"Fleetd.main must report the git host shape");
|
||||
assertTrue(contains(appender, "member trust model: members run as the same OS user"),
|
||||
"Fleetd.main must report the member trust model");
|
||||
assertTrue(contains(appender, "memberCredentials: absent or empty"),
|
||||
"Fleetd.main must report an absent memberCredentials policy");
|
||||
assertTrue(contains(appender, "exhaustedPattern: profile(s) [worker] have no exhaustedPattern configured"),
|
||||
"Fleetd.main must report profiles without exhaustedPattern");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,154 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* Proves the one thing no test proved before this ticket's follow-up: that {@code Fleetd.main}
|
||||
* ITSELF — not a copy of its logic, not the validator called directly — still refuses to start on
|
||||
* a bad config. Mutation testing found that deleting {@code cfg.validateAll();} (née six
|
||||
* individual {@code cfg.validateXxx();} calls) from {@code Fleetd.main} left the full 1478-test
|
||||
* suite green; every existing test called a validator directly and none exercised {@code
|
||||
* Fleetd.main} as the caller. See {@code FleetConfigValidateAllTest} for why the fix collapses
|
||||
* those six calls into one reflective {@link FleetConfig#validateAll()}, and {@code
|
||||
* ConfigRefTest} for the equivalent proof on the {@link
|
||||
* dev.ltms.fleet.config.ConfigRef#reload()} path.
|
||||
*
|
||||
* <p>This is deliberately the real {@code static void main(String[] args)} — package-private, so
|
||||
* only a test in this package can call it, which is exactly what makes this proof strong: it is
|
||||
* not a helper extracted for testability, it is the literal method {@code java -jar fleetd.jar}
|
||||
* invokes. Every fixture below is otherwise-valid and fails exactly one validator, and — because
|
||||
* {@link FleetConfig#validateAll()} runs immediately after {@code SubscriptionGuard.
|
||||
* assertPrimaryClean}, before {@code main} opens the herdr socket, binds Javalin, or touches
|
||||
* anything else with a real side effect — calling {@code Fleetd.main} with one of these configs is
|
||||
* safe: it is guaranteed to throw before reaching any of that, precisely because the config is
|
||||
* deliberately invalid.
|
||||
*/
|
||||
class FleetdStartupValidationTest {
|
||||
|
||||
private static void assertMainRefuses(Path dir, String fileName, String yaml,
|
||||
String mustContain) throws Exception {
|
||||
Path f = dir.resolve(fileName);
|
||||
Files.writeString(f, yaml);
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class,
|
||||
() -> Fleetd.main(new String[]{f.toString()}),
|
||||
fileName + ": Fleetd.main must refuse this config before doing anything else");
|
||||
assertTrue(e.getMessage().contains(mustContain),
|
||||
fileName + ": expected message to contain \"" + mustContain + "\" but was: "
|
||||
+ e.getMessage());
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainRefusesANonLoopbackBindWithoutTokenMode(@TempDir Path dir) throws Exception {
|
||||
assertMainRefuses(dir, "auth-exposure.yaml", """
|
||||
bind:
|
||||
host: 0.0.0.0
|
||||
port: 8765
|
||||
""", "auth.mode: token");
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainRefusesALeadTabPrefixCollision(@TempDir Path dir) throws Exception {
|
||||
assertMainRefuses(dir, "lead-tab-prefixes.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
fleet:
|
||||
tabLabel: "lead: {role} {profile}"
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
""", "fleet.tabLabel");
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainRefusesASubscriptionProfileThatReseatsAnthropicBaseUrl(@TempDir Path dir)
|
||||
throws Exception {
|
||||
assertMainRefuses(dir, "subscription-profiles.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
sonnet:
|
||||
subscription: true
|
||||
argv: ["ccs", "sonnet"]
|
||||
env:
|
||||
ANTHROPIC_BASE_URL: http://anything-not-on-the-allowlist
|
||||
""", "ANTHROPIC_BASE_URL");
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainRefusesAnUnknownCharterKey(@TempDir Path dir) throws Exception {
|
||||
assertMainRefuses(dir, "charters.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
fleet:
|
||||
charters:
|
||||
architetc: text
|
||||
""", "architetc");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #469: closes the gap left by #464's {@code CharterToolSurfaceTest}, which wrote its
|
||||
* own charter into a {@code @TempDir} fixture and so could never fail on anything anyone wrote
|
||||
* into the live {@code fleetd.yaml}. {@code bridge_send} is CB-634's own motivating example — a
|
||||
* tool name the rename removed — and the charter key ({@code dev}) is a real role wire name, so
|
||||
* this fixture passes {@code cfg.validateAll()}'s charter check (key valid, text non-blank) and
|
||||
* is refused only by the new {@code CharterToolSurface} call right after it. The failure message
|
||||
* must name both the charter key and the unknown tool.
|
||||
*/
|
||||
@Test
|
||||
void mainRefusesACharterNamingAnUnregisteredTool(@TempDir Path dir) throws Exception {
|
||||
assertMainRefuses(dir, "charter-tool-surface.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through bridge_send.
|
||||
""", "bridge_send");
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainRefusesAnArchitectSlotNamingAnUnconfiguredProfile(@TempDir Path dir) throws Exception {
|
||||
assertMainRefuses(dir, "members.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
gx10:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
fleet:
|
||||
architects:
|
||||
lead-designer:
|
||||
profile: sonnet
|
||||
""", "lead-designer");
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainRefusesAProfileNamingAModelOutsideTheAllowList(@TempDir Path dir) throws Exception {
|
||||
assertMainRefuses(dir, "models.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
rogue:
|
||||
baseUrl: http://gx11.gw:8000
|
||||
model: claude-opus-9000
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
""", "rogue");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,75 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #446 follow-up (criterion 2): {@code exhaustionSink} used to build the "name the fix"
|
||||
* WARNING inline with two SLF4J {@code {}}-placeholder {@code log.warn} calls. Nothing pinned
|
||||
* that text — a mutation battery run against the merged PR #457 proved it: renaming the leading
|
||||
* {@code "usage-limit fix:"} tag to {@code "MUTANT no fix named:"} at both call sites left all
|
||||
* 1586 existing tests green. This class is what makes that mutation red.
|
||||
*
|
||||
* <p>{@link Fleetd#usageLimitFixWarning} and {@link Fleetd#usageLimitFixWarningNoModel} are the
|
||||
* extracted call sites {@code exhaustionSink} actually invokes — {@code log.warn} is called with
|
||||
* each method's return value as a single, already-formatted argument, so the string these tests
|
||||
* assert on is byte-identical to what {@code fleetd.out} receives. Same extracted-static-method +
|
||||
* dedicated-test idiom as {@link Fleetd#exhaustedPatternCoverageLine} /
|
||||
* {@link FleetdPatternCoverageLineTest}, chosen over the {@link ch.qos.logback.core.read.ListAppender}
|
||||
* idiom {@code ExhaustedPatternGapReportTest} uses — this warning is built once, at one call site,
|
||||
* from a single method with no branching inside the message itself, so there is no real log
|
||||
* plumbing left to prove; capturing an appender would only add setup/teardown around the same
|
||||
* assertion.
|
||||
*
|
||||
* <p>Each assertion below checks for a specific substring — the leading {@code "usage-limit fix:"}
|
||||
* tag an operator would grep the log for, the profile name, the model name, and the literal
|
||||
* action ({@code enabled: false} under {@code models.allow}) — rather than the whole sentence.
|
||||
* Criterion 2 promised those facts, not a particular wording of the prose that carries them; a
|
||||
* whole-sentence {@code assertEquals} breaks on every copy-edit of that surrounding prose and
|
||||
* teaches the next person to delete the test rather than read why it failed. The leading tag is
|
||||
* pinned on purpose, unlike the rest of the sentence: it is the identifying label the fleetd #446
|
||||
* follow-up's own mutation battery renamed ({@code "usage-limit fix:"} → {@code "MUTANT no fix
|
||||
* named:"}) to prove this class did not yet catch a tag rename — only the three-fact substrings
|
||||
* above would have missed it, since none of them mention the tag itself.
|
||||
*/
|
||||
class FleetdUsageLimitFixWarningTest {
|
||||
|
||||
@Test
|
||||
@DisplayName("usageLimitFixWarning names the profile, the model, and the models.allow fix")
|
||||
void usageLimitFixWarningNamesProfileModelAndFix() {
|
||||
String warning = Fleetd.usageLimitFixWarning("terra", "claude-opus-5");
|
||||
|
||||
assertTrue(warning.startsWith("usage-limit fix:"),
|
||||
"must carry the grep-able identifying tag: " + warning);
|
||||
assertTrue(warning.contains("terra"), "must name the profile: " + warning);
|
||||
assertTrue(warning.contains("claude-opus-5"), "must name the model: " + warning);
|
||||
assertTrue(warning.contains("enabled: false"), "must name the action: " + warning);
|
||||
assertTrue(warning.contains("models.allow"), "must name where the action goes: " + warning);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("usageLimitFixWarning does not cross-name a different profile or model")
|
||||
void usageLimitFixWarningDoesNotCrossName() {
|
||||
String warning = Fleetd.usageLimitFixWarning("terra", "claude-opus-5");
|
||||
|
||||
// Guards against the mutation of swapping the two profile/model arguments at the call
|
||||
// site — a test that only checks "some name appears" cannot catch that.
|
||||
assertTrue(!warning.contains("sonnet"), "must not name an unrelated model: " + warning);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("usageLimitFixWarningNoModel names the profile but no fix, since none exists")
|
||||
void usageLimitFixWarningNoModelNamesProfileNotAFix() {
|
||||
String warning = Fleetd.usageLimitFixWarningNoModel("gx", 1800);
|
||||
|
||||
assertTrue(warning.startsWith("usage-limit fix:"),
|
||||
"must carry the grep-able identifying tag: " + warning);
|
||||
assertTrue(warning.contains("gx"), "must name the profile: " + warning);
|
||||
assertTrue(warning.contains("1800"), "must name the quarantine duration it falls back to: " + warning);
|
||||
assertTrue(!warning.contains("enabled: false"),
|
||||
"a profile with no model: configured has no models.allow entry to flip — the "
|
||||
+ "fallback message must not claim one exists: " + warning);
|
||||
}
|
||||
}
|
||||
@@ -113,6 +113,20 @@ class AuthzTest {
|
||||
assertTrue(Authz.permits(WORKER_A, METRICS, null));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #421 — unlike READ (the case above), COORD_READ (a lead's own held peer mail) is the
|
||||
* primary's alone. A worker or an architect reading it would disclose lead-to-lead
|
||||
* coordination bodies, not the secret-free roster READ is open about.
|
||||
*/
|
||||
@Test
|
||||
void coordReadIsThePrimarysAloneNotAWidenedRead() {
|
||||
assertTrue(Authz.permits(PRIMARY, COORD_READ, null));
|
||||
assertFalse(Authz.permits(WORKER_A, COORD_READ, null),
|
||||
"a worker must not read held lead-to-lead mail");
|
||||
assertFalse(Authz.permits(ARCH_DESIGN, COORD_READ, null),
|
||||
"an architect holds READ today, but not-primary must mean not-architect here too");
|
||||
}
|
||||
|
||||
@Test
|
||||
void unauthenticatedIsDistinguishedFromMerelyForbidden() {
|
||||
// Drives the 401-vs-403 split: a missing credential is fixable by the caller, a wrong role
|
||||
|
||||
@@ -0,0 +1,360 @@
|
||||
package dev.ltms.fleet.auth;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.PaneLocator;
|
||||
import dev.ltms.fleet.mcp.ConnectionIdentity;
|
||||
import dev.ltms.fleet.peer.MemberRole;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* fleetd #424 — revoking (or granting) an architect slot must take effect on the next spawn with
|
||||
* no restart. A session already bound to a slot keeps its <em>binding</em> (the {@code
|
||||
* terminalToSlot} occupancy) even after that slot drops out of config, but NOT the ARCHITECT
|
||||
* <em>privilege</em> the slot used to grant — that is revoked on the bound session's very next
|
||||
* request. See {@link MemberRegistry}'s class doc for the exact rule: "config governs what a
|
||||
* bound slot still grants, as well as what may be bound next."
|
||||
*
|
||||
* <p>Every test here drives a REAL {@link ConfigRef#reload()} against a {@code @TempDir} file and
|
||||
* asserts {@link ConfigRef.Outcome#applied()}, rather than comparing two frozen
|
||||
* {@code MemberRegistry} instances in memory — the defect this ticket fixes is specifically that
|
||||
* {@link MemberRegistry} used to ignore a live reload, so a test that never reloads cannot tell the
|
||||
* fixed registry from the broken one. {@code requireSlotFor} and {@code reserve} are pinned in
|
||||
* separate tests, in both directions (removed and added), so a registry that simply refuses (or
|
||||
* simply allows) everything cannot pass by accident — see {@link MemberRegistryTest} for the
|
||||
* registry's other invariants (bind/unbind cardinality, thread-safety), which are unaffected by
|
||||
* this ticket and still exercised against the frozen constructor.
|
||||
*/
|
||||
class MemberRegistryLiveTest {
|
||||
|
||||
private static String yaml(String fleetBlock) {
|
||||
return """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet
|
||||
opus:
|
||||
baseUrl: http://gx00.gw:8001
|
||||
model: opus
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""" + fleetBlock;
|
||||
}
|
||||
|
||||
private static final String WITH_SONNET_SLOT = """
|
||||
fleet:
|
||||
architects:
|
||||
designer:
|
||||
profile: sonnet
|
||||
""";
|
||||
|
||||
/** No architect pool at all — developers is unrelated dead data for this registry (#424 out of scope). */
|
||||
private static final String WITHOUT_ARCHITECT_SLOTS = """
|
||||
fleet:
|
||||
developers:
|
||||
dev1:
|
||||
profile: sonnet
|
||||
""";
|
||||
|
||||
/** Same slot name ({@code designer}) as {@link #WITH_SONNET_SLOT}, repointed to a different profile. */
|
||||
private static final String WITH_OPUS_SLOT = """
|
||||
fleet:
|
||||
architects:
|
||||
designer:
|
||||
profile: opus
|
||||
""";
|
||||
|
||||
/** Same profile ({@code sonnet}) as {@link #WITH_SONNET_SLOT}, but the pool key is renamed. */
|
||||
private static final String WITH_RENAMED_SLOT = """
|
||||
fleet:
|
||||
architects:
|
||||
architect-lead:
|
||||
profile: sonnet
|
||||
""";
|
||||
|
||||
private static ConfigRef refFor(Path f) {
|
||||
return new ConfigRef(f, FleetConfig.load(f));
|
||||
}
|
||||
|
||||
// ── requireSlotFor is live (criteria 1, 2, 4) ──────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
void requireSlotForRefusesAProfileWhoseSlotWasRemovedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
assertDoesNotThrow(() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
|
||||
"the slot is configured before the reload");
|
||||
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
|
||||
"revoking the slot must refuse the NEXT spawn that names it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void requireSlotForAllowsAProfileWhoseSlotWasAddedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
|
||||
"no architect slot is configured yet");
|
||||
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertDoesNotThrow(() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
|
||||
"a slot added by reload must be usable with no restart");
|
||||
}
|
||||
|
||||
// ── reserve is live too — tested separately from requireSlotFor (criteria 1, 2, 4) ────────
|
||||
|
||||
@Test
|
||||
void reserveRefusesAProfileWhoseSlotWasRemovedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
MemberLifecycle.SlotReservation before = registry.reserve(MemberRole.ARCHITECT, "sonnet");
|
||||
assertEquals("architect:designer", before.slot());
|
||||
registry.release(before); // free it back up so the reload-side reserve below starts clean
|
||||
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
|
||||
"revoking the slot must refuse the NEXT reservation for it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void reserveAllowsAProfileWhoseSlotWasAddedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
|
||||
"no architect slot is configured yet");
|
||||
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
MemberLifecycle.SlotReservation after = registry.reserve(MemberRole.ARCHITECT, "sonnet");
|
||||
assertEquals("architect:designer", after.slot(),
|
||||
"a slot added by reload must be reservable with no restart");
|
||||
}
|
||||
|
||||
// ── profileForSlot, isSlot and nameForSlot are live too (fleetd #431) ─────────────────────────
|
||||
// #424 pinned roleForSlot against a live reload but left these three untested — proved by
|
||||
// mutating each to read a snapshot flattened once at construction: the full suite stayed green
|
||||
// for all three.
|
||||
//
|
||||
// The three differ in how much production behaviour depends on them, and the ticket first got
|
||||
// this ranking wrong. nameForSlot is wired: CallerResolver passes members::nameForSlot, next to
|
||||
// members::roleForSlot. isSlot is reached through bind, which calls it to refuse an unknown
|
||||
// slot. profileForSlot has NO caller in src/main at all — grepped both the ".profileForSlot("
|
||||
// and the "::profileForSlot" form — so there is no seam to drive its test through beyond the
|
||||
// accessor itself, and its own javadoc ("what the spawn lifecycle reads") describes a caller
|
||||
// that does not exist. These tests pin the accessors as they are; whether profileForSlot should
|
||||
// be wired or deleted is a separate question.
|
||||
|
||||
@Test
|
||||
void profileForSlotReflectsAProfileChangedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
assertEquals("sonnet", registry.profileForSlot("architect:designer"),
|
||||
"the slot's profile before the reload");
|
||||
|
||||
Files.writeString(f, yaml(WITH_OPUS_SLOT));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertEquals("opus", registry.profileForSlot("architect:designer"),
|
||||
"repointing the slot to a different profile must take effect with no restart");
|
||||
}
|
||||
|
||||
@Test
|
||||
void isSlotStopsReportingASlotRemovedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
assertTrue(registry.isSlot("architect:designer"), "the slot is configured before the reload");
|
||||
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertFalse(registry.isSlot("architect:designer"),
|
||||
"removing the slot from config must make isSlot say so on the very next call, "
|
||||
+ "with no restart");
|
||||
}
|
||||
|
||||
@Test
|
||||
void isSlotStartsReportingASlotAddedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
assertFalse(registry.isSlot("architect:designer"), "no architect slot is configured yet");
|
||||
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertTrue(registry.isSlot("architect:designer"),
|
||||
"a slot added by reload must be visible to isSlot with no restart");
|
||||
}
|
||||
|
||||
@Test
|
||||
void nameForSlotReflectsANameChangedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
assertEquals("designer", registry.nameForSlot("architect:designer"),
|
||||
"the configured name before the reload");
|
||||
|
||||
Files.writeString(f, yaml(WITH_RENAMED_SLOT));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertNull(registry.nameForSlot("architect:designer"),
|
||||
"the old key no longer names a configured slot — it was renamed away by the reload");
|
||||
assertEquals("architect-lead", registry.nameForSlot("architect:architect-lead"),
|
||||
"the new name must be visible under its new qualified key with no restart");
|
||||
}
|
||||
|
||||
// ── a bound architect is demoted, but the binding itself is not touched (fleetd #424) ───────
|
||||
// The lead's corrected ruling: the PRIVILEGE a slot grants is revoked on the bound session's
|
||||
// very next request, but the terminalToSlot BINDING itself is untouched by a reload — dropping
|
||||
// it would double-book the slot key and break unbind's compare-safe contract. See the class
|
||||
// doc's binding rule.
|
||||
|
||||
/** A caller identity resolving the one canned pane (terminal {@code term_a}) in {@link FakeHerdr}. */
|
||||
private static ConnectionIdentity boundPaneIdentity() {
|
||||
return new ConnectionIdentity(new PaneLocator(new FakeHerdr()), _ -> FakeHerdr.WORKER_PID);
|
||||
}
|
||||
|
||||
@Test
|
||||
void anArchitectAlreadyBoundToASlotIsDemotedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
|
||||
assertTrue(registry.bind(reservation, "term_a"));
|
||||
assertEquals("architect:designer", registry.slotForTerminal("term_a"));
|
||||
|
||||
// Drive the real caller path, not the roleForSlot seam directly: CallerResolver.resolve is
|
||||
// what a live request actually goes through (CallerResolver.java:220), and a resolver that
|
||||
// ignored roleForSlot entirely would still pass a test that only checked the seam.
|
||||
CallerResolver resolver = CallerResolver.withLeadsAndMembers(
|
||||
boundPaneIdentity(), false, null, Map::of, registry);
|
||||
|
||||
Principal before = resolver.resolve("127.0.0.1", 42, null);
|
||||
assertEquals(Role.ARCHITECT, before.role(), "sanity check: the harness binds term_a as an architect");
|
||||
assertEquals("designer", before.name());
|
||||
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
Principal after = resolver.resolve("127.0.0.1", 42, null);
|
||||
assertEquals(Role.WORKER, after.role(),
|
||||
"removing the slot from config must demote the bound session to worker on its "
|
||||
+ "NEXT request — this is the ticket's whole point");
|
||||
assertEquals("term_a", after.terminal(), "same pane, same terminal — only the role changed");
|
||||
}
|
||||
|
||||
@Test
|
||||
void theOriginalBindingStillOccupiesTheRemovedSlotSoASecondTerminalCannotClaimIt(@TempDir Path dir)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
|
||||
assertTrue(registry.bind(reservation, "term_a"));
|
||||
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef.Outcome removed = ref.reload();
|
||||
assertTrue(removed.applied(), "the reload must actually take effect: " + removed.summary());
|
||||
|
||||
// The binding survives the removal untouched.
|
||||
assertEquals("architect:designer", registry.slotForTerminal("term_a"),
|
||||
"a live binding must never be retroactively unbound by a config edit");
|
||||
assertEquals(Map.of("term_a", "architect:designer"), registry.snapshot());
|
||||
|
||||
// Bring the slot back into config. If the binding had been silently dropped by the removal
|
||||
// (rather than merely losing the privilege it grants), a second terminal could now claim
|
||||
// the "freed" key — the exact double-booking the class doc's binding rule rules out.
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef.Outcome restored = ref.reload();
|
||||
assertTrue(restored.applied(), "the reload must actually take effect: " + restored.summary());
|
||||
|
||||
assertFalse(registry.bind("architect:designer", "term_b"),
|
||||
"the slot is still occupied by term_a — a second terminal must not bind to it");
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
|
||||
"the slot is still occupied by term_a — a fresh reservation must not find it free");
|
||||
assertEquals("architect:designer", registry.slotForTerminal("term_a"),
|
||||
"the original binding is unchanged throughout");
|
||||
}
|
||||
|
||||
@Test
|
||||
void unbindStillSucceedsForTheOriginalTerminalAfterItsSlotIsRemoved(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
|
||||
assertTrue(registry.bind(reservation, "term_a"));
|
||||
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertTrue(registry.unbind("architect:designer", "term_a"),
|
||||
"unbind must still work for a slot that config has since removed, or a session "
|
||||
+ "that outlives its slot's removal could never release it");
|
||||
assertNull(registry.slotForTerminal("term_a"));
|
||||
assertEquals(Map.of(), registry.snapshot());
|
||||
}
|
||||
}
|
||||
@@ -161,7 +161,15 @@ class ConfigRefProfileCoverageTest {
|
||||
// the #323 bug. Only ideProjectDir and worktreeGroup were saved by a behavioural test in
|
||||
// ConfigRefTest; the other two had none. So: growing this set now requires editing this
|
||||
// line as well, which is a visible, deliberate diff rather than a quiet one.
|
||||
assertEquals(Set.of("weight", "maxLoad", "credentialId"), excluded,
|
||||
//
|
||||
// fleetd #446: "exhaustedPattern" joined the set legitimately — it moved from compiled
|
||||
// once at Fleetd.main startup to read live, cached by profile name, through
|
||||
// LiveExhaustedPatterns (see that class's doc and ConfigRef's class doc). Unlike the #323
|
||||
// hazard this comment warns about, it is pinned by its OWN behavioural tests:
|
||||
// FleetdExhaustionDetectionArmedWiringTest proves the reload takes effect with no restart,
|
||||
// and the same test rerun against the pre-#446 compile site (see that ticket's report) is
|
||||
// what proves the assertion would have failed before the fix.
|
||||
assertEquals(Set.of("weight", "maxLoad", "credentialId", "exhaustedPattern"), excluded,
|
||||
"ConfigRef.LAUNCH_SETTINGS_EXCLUDED changed. A component belongs in it ONLY if it "
|
||||
+ "is read live off the config supplier, not baked into a launcher at "
|
||||
+ "startup. If you are adding one to silence this test, that is fleetd #323 "
|
||||
|
||||
@@ -1,10 +1,13 @@
|
||||
package dev.ltms.fleet.config;
|
||||
|
||||
import dev.ltms.fleet.mcp.CharterToolSurface;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.function.Consumer;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
@@ -38,6 +41,23 @@ class ConfigRefTest {
|
||||
return new ConfigRef(f, FleetConfig.load(f));
|
||||
}
|
||||
|
||||
/**
|
||||
* The exact {@code Consumer<FleetConfig>} {@code Fleetd.main} wires into {@code ConfigRef}'s
|
||||
* constructor as {@code extraValidation} (fleetd #474) — an adapter from {@code FleetConfig} to
|
||||
* the raw charter map {@link CharterToolSurface#assertChartersNameOnlyRegisteredTools} takes.
|
||||
* Built here rather than referencing {@code dev.ltms.fleet.Fleetd} directly, because that method
|
||||
* is package-private to {@code dev.ltms.fleet} and this test lives in {@code
|
||||
* dev.ltms.fleet.config} — but it calls the SAME production {@link CharterToolSurface} method
|
||||
* {@code Fleetd} calls, so this proves the real check runs on reload, not a stand-in for it.
|
||||
*/
|
||||
private static final Consumer<FleetConfig> CHARTER_TOOL_SURFACE = cfg ->
|
||||
CharterToolSurface.assertChartersNameOnlyRegisteredTools(
|
||||
cfg.fleet() == null ? Map.of() : cfg.fleet().charters());
|
||||
|
||||
private static ConfigRef refForWithCharterToolSurface(Path f) {
|
||||
return new ConfigRef(f, FleetConfig.load(f), CHARTER_TOOL_SURFACE);
|
||||
}
|
||||
|
||||
@Test
|
||||
void aHotChangeIsAppliedAndReadThroughGet(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
@@ -124,6 +144,87 @@ class ConfigRefTest {
|
||||
assertSame(before, ref.get());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #474: {@code validateAll()} (via {@code validateCharters()}) only checks that a
|
||||
* charter's KEY is a role wire name and its text is non-blank — it never looks at what the text
|
||||
* names, so a charter naming {@code bridge_send} (the pre-CB-634 name, removed from the tool
|
||||
* surface — #469's own motivating example) passes {@code validateAll()} and used to be applied
|
||||
* on reload with nothing refusing it, even though the identical charter refuses {@code
|
||||
* Fleetd.main} at startup ({@code FleetdStartupValidationTest
|
||||
* #mainRefusesACharterNamingAnUnregisteredTool}). This is the reload-path proof: it drives
|
||||
* {@link ConfigRef#reload()} itself (not a direct call to {@link
|
||||
* CharterToolSurface#assertChartersNameOnlyRegisteredTools}), through the exact {@code
|
||||
* extraValidation} wiring {@code Fleetd.main} uses, and the failure message must name both the
|
||||
* charter key and the unknown tool — the same information the startup failure gives (ticket
|
||||
* acceptance criterion 1).
|
||||
*/
|
||||
@Test
|
||||
void aReloadRefusesACharterNamingAnUnregisteredTool(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through fleet_reply.
|
||||
"""));
|
||||
ConfigRef ref = refForWithCharterToolSurface(f);
|
||||
FleetConfig before = ref.get();
|
||||
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through bridge_send.
|
||||
"""));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertFalse(out.applied());
|
||||
assertNotNull(out.error());
|
||||
assertTrue(out.error().contains("dev"),
|
||||
"expected the charter key 'dev' in the refusal, got: " + out.error());
|
||||
assertTrue(out.error().contains("bridge_send"),
|
||||
"expected the unknown tool 'bridge_send' in the refusal, got: " + out.error());
|
||||
assertTrue(out.summary().startsWith("config reload refused"), out.summary());
|
||||
// The running config must not move at all — half-applying this would leave the daemon in a
|
||||
// state that could never have booted, exactly the outcome ConfigRef.java's class doc warns
|
||||
// a cold-key refusal must avoid, and this check must avoid the same way.
|
||||
assertSame(before, ref.get());
|
||||
assertEquals("Send the final handoff through fleet_reply.\n",
|
||||
ref.get().fleet().charterFor(dev.ltms.fleet.peer.MemberRole.DEV));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #474 acceptance criterion 3, the positive case: a reload whose charter names only
|
||||
* registered tools must still be ACCEPTED. An inverted filter (one that refuses every charter,
|
||||
* or refuses on any {@code fleet_*}/{@code bridge_*} token regardless of registration) would pass
|
||||
* the refusal test above alone — #469's own M5 mutation cell showed exactly that shape surviving
|
||||
* a negative-only suite. This is the test that catches it.
|
||||
*/
|
||||
@Test
|
||||
void aReloadAcceptsACharterNamingOnlyRegisteredTools(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
charters:
|
||||
dev: old charter, no tool names
|
||||
"""));
|
||||
ConfigRef ref = refForWithCharterToolSurface(f);
|
||||
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through fleet_reply, using fleet_send to delegate.
|
||||
"""));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
|
||||
assertEquals("config reloaded", out.summary());
|
||||
assertEquals("Send the final handoff through fleet_reply, using fleet_send to delegate.\n",
|
||||
ref.get().fleet().charterFor(dev.ltms.fleet.peer.MemberRole.DEV));
|
||||
}
|
||||
|
||||
/**
|
||||
* The point of the whole class: a consumer holding the ref sees the new value without being
|
||||
* rebuilt. A component that captured {@code get()} into a field would still show the old one.
|
||||
@@ -339,12 +440,18 @@ class ConfigRefTest {
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-578 stage B: exhaustedPattern is compiled once into Fleetd.main's pattern map at startup
|
||||
* (see ExhaustedPatternLookup), so a reload never re-reads it — a changed pattern must be
|
||||
* reported deferred exactly like model/baseUrl, not silently claimed as applied.
|
||||
* fleetd #446: exhaustedPattern moved from deferred to hot — {@code LiveExhaustedPatterns}
|
||||
* reads {@code config.get()} fresh on every lookup, cached by profile name (not baked into a
|
||||
* frozen map inside {@code Fleetd.main} any more, the way it was before this ticket, and the
|
||||
* way its sibling {@code errorPattern} still is — see
|
||||
* {@code changingAProfilesErrorPatternIsReportedAsDeferred} right below for the contrast).
|
||||
* So a reload that changes ONLY exhaustedPattern must report a plain "config reloaded", with
|
||||
* nothing deferred — the opposite of what this test asserted before fleetd #446, when
|
||||
* ConfigRef.sameLaunchSettings still compared exhaustedPattern and this reload reported it as
|
||||
* one of the changes needing a restart.
|
||||
*/
|
||||
@Test
|
||||
void changingAProfilesExhaustedPatternIsReportedAsDeferred(@TempDir Path dir) throws Exception {
|
||||
void changingAProfilesExhaustedPatternTakesEffectWithNoRestart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
@@ -379,17 +486,21 @@ class ConfigRefTest {
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertEquals(1, out.deferred().size(), out.deferred().toString());
|
||||
assertTrue(out.deferred().getFirst().contains("sonnet"), out.deferred().toString());
|
||||
assertTrue(out.deferred().getFirst().contains("launch settings"), out.deferred().toString());
|
||||
// The snapshot still carries the new value — a restart is what makes it take effect.
|
||||
assertEquals(0, out.deferred().size(),
|
||||
"exhaustedPattern is hot since fleetd #446 — changing only this key must not be "
|
||||
+ "reported as needing a restart: " + out.deferred());
|
||||
assertEquals(0, out.split().size(), out.split().toString());
|
||||
assertEquals("config reloaded", out.summary());
|
||||
// The snapshot carries the new value immediately — this IS the live value now, not
|
||||
// merely what a restart would eventually pick up.
|
||||
assertEquals("rate limit exceeded", ref.get().profiles().get("sonnet").exhaustedPattern());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #201 Unit 5: errorPattern is compiled once into Fleetd.main's backend-error pattern
|
||||
* map at startup (see BackendErrorPatternLookup), the same way exhaustedPattern is above — a
|
||||
* reload never re-reads it either, so a changed value must be reported deferred.
|
||||
* map at startup (see BackendErrorPatternLookup) — unlike its sibling exhaustedPattern above,
|
||||
* which fleetd #446 made hot, errorPattern was scoped OUT of that ticket on purpose and stays
|
||||
* deferred: a reload still never re-reads it, so a changed value must be reported deferred.
|
||||
*/
|
||||
@Test
|
||||
void changingAProfilesErrorPatternIsReportedAsDeferred(@TempDir Path dir) throws Exception {
|
||||
|
||||
@@ -79,6 +79,25 @@ class ConfigRefTopLevelCoverageTest {
|
||||
* <li>{@code memberCredentials} — read live at {@code Fleetd.java:198, 205, 729}.</li>
|
||||
* <li>{@code memberLoginShell} — read live at
|
||||
* {@code HerdrPeerLauncher#configuredMemberLoginShell}.</li>
|
||||
* <li>{@code models} (fleetd #422) — membership is re-validated in full against the fresh
|
||||
* config on every {@code ConfigRef.reload()} (via {@code FleetConfig#validateAll()}), so
|
||||
* a bad edit is refused, never cached stale; the on/off half is read live by {@code
|
||||
* CompositePeerLauncher.enforceModelEnabled} and its candidate filter, and by {@code
|
||||
* fleet_profiles}/{@code GET /profiles} through {@code PeerLauncher.disabledModels()}.
|
||||
* Unlike {@code fleet} below, nothing about {@code models} is baked into a startup-built
|
||||
* object anywhere — there is no frozen half, so it belongs here whole rather than in
|
||||
* {@code SPLIT_KEYS}.</li>
|
||||
* <li>{@code leadRollover} (fleetd #480) — {@code dev.ltms.fleet.lead.LeadRollover} holds a
|
||||
* {@code Supplier<FleetConfig.LeadRollover>} and reads every field fresh on each
|
||||
* {@code open()}/{@code confirm()} call, the same {@code () -> config.get().x()} shape
|
||||
* {@code fleet}/{@code placement}/{@code models} use — see {@code ConfigRef}'s class doc
|
||||
* Hot bullet. The only restart-only fact is structural, not a value going stale:
|
||||
* {@code Fleetd.java} decides whether to construct the {@code LeadRollover} object at
|
||||
* all off the startup snapshot (the same presence gate {@code leadHeartbeat:} uses), so
|
||||
* a block ADDED where it was absent at boot needs a restart before anything exists to
|
||||
* call — the same fact already true of adding a brand-new {@code profiles:} entry, which
|
||||
* does not stop {@code profiles}' own hot sub-fields (weight/maxLoad/credentialId/
|
||||
* exhaustedPattern) from being genuinely hot.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>{@code fleet} used to sit here too, on the strength of most of it (role pools, charters,
|
||||
@@ -93,7 +112,7 @@ class ConfigRefTopLevelCoverageTest {
|
||||
* while this test stayed green throughout.</p>
|
||||
*/
|
||||
private static final Set<String> HOT_EXCLUDED_TOP_LEVEL_KEYS =
|
||||
Set.of("placement", "memberCredentials", "memberLoginShell");
|
||||
Set.of("placement", "memberCredentials", "memberLoginShell", "models", "leadRollover");
|
||||
|
||||
@Test
|
||||
void everyTopLevelComponentIsAccountedForInExactlyOneClass() {
|
||||
@@ -120,7 +139,7 @@ class ConfigRefTopLevelCoverageTest {
|
||||
|
||||
// The escape hatch is pinned. Growing it requires editing this line — a visible, deliberate
|
||||
// diff, not a quiet one. See the field javadoc above for what "belongs here" actually means.
|
||||
assertEquals(Set.of("placement", "memberCredentials", "memberLoginShell"), hot,
|
||||
assertEquals(Set.of("placement", "memberCredentials", "memberLoginShell", "models", "leadRollover"), hot,
|
||||
"HOT_EXCLUDED_TOP_LEVEL_KEYS changed. A component belongs here ONLY if it is read "
|
||||
+ "live off the config supplier, never because adding it makes this test "
|
||||
+ "pass. If you are adding one to silence this test, that is fleetd #323 "
|
||||
|
||||
@@ -109,6 +109,11 @@ class ConfigRefTopLevelReportingCoverageTest {
|
||||
v.put("memberLoginShell", null);
|
||||
v.put("memberSkills", "/skills/a");
|
||||
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(true));
|
||||
v.put("models", new FleetConfig.Models(
|
||||
List.of(new FleetConfig.Models.ModelEntry("model-a"))));
|
||||
// fleetd #480: leadRollover joins placement/memberCredentials/memberLoginShell as
|
||||
// hot-excluded — never compared by any changed*Keys method, so it stays null like them.
|
||||
v.put("leadRollover", null);
|
||||
assertNamesMatchComponents(v);
|
||||
return v;
|
||||
}
|
||||
@@ -151,6 +156,9 @@ class ConfigRefTopLevelReportingCoverageTest {
|
||||
v.put("memberLoginShell", null);
|
||||
v.put("memberSkills", "/skills/b");
|
||||
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(false));
|
||||
v.put("models", new FleetConfig.Models(
|
||||
List.of(new FleetConfig.Models.ModelEntry("model-b"))));
|
||||
v.put("leadRollover", null);
|
||||
assertNamesMatchComponents(v);
|
||||
return v;
|
||||
}
|
||||
|
||||
@@ -2683,4 +2683,337 @@ class FleetConfigTest {
|
||||
assertTrue(cfg.idleSleepGuard().isEnabled(),
|
||||
"unlike ConfigReload/Health, this block defaults to ON even when present but empty");
|
||||
}
|
||||
|
||||
// --- models: central allow-list of usable models -----------------------------------------
|
||||
|
||||
/**
|
||||
* An absent {@code models:} block is today's behaviour exactly: no profile's {@code model:} is
|
||||
* checked against anything, whatever it says. This is the "existing config keeps working"
|
||||
* invariant — an operator on a gitignored {@code fleetd.yaml} that predates this feature must
|
||||
* not be broken by upgrading the daemon.
|
||||
*/
|
||||
@Test
|
||||
void absentModelsBlockValidatesNothing(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: totally-unheard-of-model-xyz
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code models:} present but with an empty (or absent) {@code allow:} must behave exactly like
|
||||
* the block being absent — an operator adding the block for the first time with nothing in it
|
||||
* yet must not be surprised by every profile suddenly refusing to start.
|
||||
*/
|
||||
@Test
|
||||
void modelsBlockPresentButEmptyValidatesNothing(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: whatever-the-operator-typed
|
||||
models:
|
||||
allow: []
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels);
|
||||
}
|
||||
|
||||
/**
|
||||
* The core of the ticket: a profile naming a model outside the configured allow-list fails
|
||||
* config load, naming both the model and the profile that wanted it.
|
||||
*/
|
||||
@Test
|
||||
void aProfileNamingAModelOutsideTheAllowListRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
rogue:
|
||||
baseUrl: http://gx11.gw:8000
|
||||
model: claude-opus-9000
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
- model: claude-haiku-5
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateModels);
|
||||
assertTrue(e.getMessage().contains("rogue"), "the refusal names the offending profile");
|
||||
assertTrue(e.getMessage().contains("claude-opus-9000"), "the refusal names the offending model");
|
||||
assertFalse(e.getMessage().contains("'sonnet'"),
|
||||
"the profile whose model IS allowed must not be reported");
|
||||
}
|
||||
|
||||
/** A profile whose {@code model:} is on the allow-list loads fine. */
|
||||
@Test
|
||||
void aProfileNamingAModelOnTheAllowListLoadsFine(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels);
|
||||
}
|
||||
|
||||
/**
|
||||
* A profile that never sets {@code model:} (a {@code subscription: true} profile relying on the
|
||||
* account's own default is the live example) must not be refused just because an allow-list is
|
||||
* active — there is nothing to check it against.
|
||||
*/
|
||||
@Test
|
||||
void aProfileWithNoModelPassesEvenWithAnActiveAllowList(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
opus:
|
||||
subscription: true
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels);
|
||||
}
|
||||
|
||||
/**
|
||||
* Reproduces the live config shape this ticket measured: a mix of {@code amazon-bedrock},
|
||||
* {@code opencode} and {@code openai}-backed profiles, some naming a bare Claude id and some a
|
||||
* provider-prefixed opencode id, in ONE allow-list. Both forms are just opaque strings compared
|
||||
* for exact equality — proves the shape decision holds against the shape actually seen live,
|
||||
* not just against a synthetic single-provider example.
|
||||
*/
|
||||
@Test
|
||||
void aBareClaudeIdAndAnOpencodeProviderPrefixedIdBothFitOneAllowList(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
terra:
|
||||
kind: opencode
|
||||
model: openai/gpt-5.6-terra
|
||||
nova:
|
||||
kind: opencode
|
||||
model: amazon-bedrock/amazon.nova-pro-v1:0
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
- model: openai/gpt-5.6-terra
|
||||
- model: amazon-bedrock/amazon.nova-pro-v1:0
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels);
|
||||
}
|
||||
|
||||
/** The same live shape, but one opencode profile's model is missing from the allow-list. */
|
||||
@Test
|
||||
void anUnlistedOpencodeProviderPrefixedModelRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
terra:
|
||||
kind: opencode
|
||||
model: openai/gpt-5.6-terra-withdrawn
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
- model: openai/gpt-5.6-terra
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateModels);
|
||||
assertTrue(e.getMessage().contains("terra"), "the refusal names the offending profile");
|
||||
assertTrue(e.getMessage().contains("openai/gpt-5.6-terra-withdrawn"),
|
||||
"the refusal names the offending model, with its provider prefix intact");
|
||||
}
|
||||
|
||||
/**
|
||||
* Editing a {@code profiles:} entry alone must never be able to widen what is permitted — only
|
||||
* editing {@code models.allow:} itself can. This is the invariant the ticket states explicitly;
|
||||
* this test pins it by giving a profile a plausible-looking model that was never added to the
|
||||
* allow-list and confirming it is still refused.
|
||||
*/
|
||||
@Test
|
||||
void addingAProfileCannotWidenTheAllowListByItself(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
brand-new:
|
||||
baseUrl: http://gx12.gw:8000
|
||||
model: claude-sonnet-6-preview
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateModels);
|
||||
assertTrue(e.getMessage().contains("brand-new"));
|
||||
assertTrue(e.getMessage().contains("claude-sonnet-6-preview"));
|
||||
}
|
||||
|
||||
/** {@code models} is a brand-new top-level key and must be recognized, not WARN-ed as unknown. */
|
||||
@Test
|
||||
void modelsIsAKnownTopLevelKey() {
|
||||
assertTrue(FleetConfig.KNOWN_TOP_LEVEL_KEYS.contains("models"));
|
||||
}
|
||||
|
||||
// --- fleetd #422: model on/off (spawn-time gate + hot reload) -----------------------------
|
||||
|
||||
/**
|
||||
* Criterion 3: an {@code allow:} entry written before {@code enabled:} existed — no such key in
|
||||
* the YAML at all — must behave exactly as it always did: on, and absent from {@link
|
||||
* FleetConfig.Models#offIds()}. This is the old-style fixture the ticket asks for, proven
|
||||
* through a real YAML load rather than only through the {@code ModelEntry(String)} back-compat
|
||||
* constructor.
|
||||
*/
|
||||
@Test
|
||||
void anOldStyleAllowEntryWithNoEnabledFieldStaysOn(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels);
|
||||
assertTrue(cfg.models().ids().contains("claude-sonnet-5"), "membership is unaffected");
|
||||
assertTrue(cfg.models().offIds().isEmpty(), "no enabled: field ⇒ nothing is off");
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 7 (the whole point of the ticket, mirroring criterion 3's shape): a profile naming a
|
||||
* model that is turned off ({@code enabled: false}) must still be VALID config —
|
||||
* {@code validateModels()} checks membership only, never the on/off state, so it must not throw.
|
||||
* Collapsing "off" into "removed from allow:" would make this test fail, which is exactly the
|
||||
* design trap the ticket calls out.
|
||||
*/
|
||||
@Test
|
||||
void anOffModelIsStillValidConfigForValidateModels(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
models:
|
||||
allow:
|
||||
- model: deepseek-v4-flash
|
||||
enabled: false
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels,
|
||||
"an off model must stay a valid allow-list member — only the spawn gate reads enabled");
|
||||
assertTrue(cfg.models().ids().contains("deepseek-v4-flash"));
|
||||
assertTrue(cfg.models().offIds().contains("deepseek-v4-flash"));
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 8: one {@code enabled: false} entry names a model, not a profile — every profile
|
||||
* naming that model is off, on purpose (the live example: {@code deepseek-v4-flash} backs both
|
||||
* {@code local} and {@code local-direct}). Proven here at the {@code Models}/config-load level;
|
||||
* {@code CompositePeerLauncherTest} proves the spawn-time consequence for both profiles.
|
||||
*/
|
||||
@Test
|
||||
void turningOffOneModelIsIndependentOfHowManyProfilesNameIt(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
local-direct:
|
||||
baseUrl: http://local.gw:8001
|
||||
model: deepseek-v4-flash
|
||||
models:
|
||||
allow:
|
||||
- model: deepseek-v4-flash
|
||||
enabled: false
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels);
|
||||
// One allow-list entry, one off id — the fan-out to both profiles happens at the reader
|
||||
// (CompositePeerLauncher), not by duplicating the entry per profile.
|
||||
assertEquals(Set.of("deepseek-v4-flash"), cfg.models().offIds());
|
||||
assertEquals("deepseek-v4-flash", cfg.profiles().get("local").model());
|
||||
assertEquals("deepseek-v4-flash", cfg.profiles().get("local-direct").model());
|
||||
}
|
||||
|
||||
/** An {@code enabled: true} entry (explicit, not just absent) also stays on — not just null. */
|
||||
@Test
|
||||
void explicitlyEnabledTrueStaysOn(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
enabled: true
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertTrue(cfg.models().offIds().isEmpty());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,358 @@
|
||||
package dev.ltms.fleet.config;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Method;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.Set;
|
||||
import java.util.TreeSet;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* The gap this class exists to close: mutation testing on the fleetd ticket "central allow-list
|
||||
* of usable models" found that although {@link FleetConfig#validateModels()}'s own logic was well
|
||||
* pinned, nothing proved either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()})
|
||||
* still invoked it — deleting the call site left the full suite green (1478/0/0/0). A follow-up
|
||||
* measurement (same technique — remove one call site, run the suite, not read the code) found the
|
||||
* SAME gap for all five of {@link FleetConfig}'s other validators at startup, and for four of the
|
||||
* six inside {@link ConfigRef#reload()}. This is a class of gap, not one line's mistake: every one
|
||||
* of those thirteen tests called the validator itself directly, never the real caller that was
|
||||
* supposed to.
|
||||
*
|
||||
* <p>The fix replaces the six individual {@code cfg.validateXxx()} calls at each of the two real
|
||||
* call sites with one {@link FleetConfig#validateAll()}, which reaches every validator by
|
||||
* reflection rather than by a hand-maintained list of names. A hand-maintained list of six names
|
||||
* would have exactly the defect it replaces: the seventh validator someone adds next month has no
|
||||
* reason to be added to it, and nothing would say so. This class proves TWO separate claims, and
|
||||
* keeps them separate on purpose:
|
||||
*
|
||||
* <ol>
|
||||
* <li>{@link #theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames()} and its neighbours
|
||||
* prove the reflective sweep itself ({@link FleetConfig#invokeAllValidators}) is a general
|
||||
* mechanism — it runs whatever public, no-arg, void {@code validateXxx()} methods a class
|
||||
* happens to declare today, including a class with more of them than {@link FleetConfig}
|
||||
* has right now. This is the proof that a future, real seventh validator on {@link
|
||||
* FleetConfig} would be swept automatically, without needing to add a real (unwanted)
|
||||
* seventh validator just to exercise the claim.</li>
|
||||
* <li>{@link #validateAllReachesEveryOneOfTodaysSixValidators()} proves {@link
|
||||
* FleetConfig#validateAll()} itself is wired to that same generic mechanism and genuinely
|
||||
* reaches each of today's six real validators — reusing the exact minimal failing
|
||||
* configurations {@code FleetConfigTest} already established for each one directly, so a
|
||||
* single call to {@code validateAll()} is shown to reproduce every one of those six
|
||||
* failures.</li>
|
||||
* </ol>
|
||||
*
|
||||
* <p>Together with the direct-{@code Fleetd.main}-invocation tests in {@code
|
||||
* FleetdStartupValidationTest} (which prove the real startup call site still calls {@code
|
||||
* validateAll()}) and the {@code ConfigRefTest} reload tests (which prove the same for {@link
|
||||
* ConfigRef#reload()}), removing {@code cfg.validateAll();} from either real call site now fails
|
||||
* a test in this module.
|
||||
*
|
||||
* <p><b>What is NOT pinned, measured rather than assumed.</b> Reverting {@link
|
||||
* FleetConfig#validateAll()} to a hardcoded list of today's six method calls leaves the whole
|
||||
* suite green (measured at review: 1491 tests, 0 failures). Nothing ties {@code validateAll()} to
|
||||
* the generic sweep — claim 1 proves {@link FleetConfig#invokeAllValidators} is generic, and claim
|
||||
* 2 proves {@code validateAll()} reaches today's six, and a hardcoded list satisfies both. So the
|
||||
* reflective sweep is a convenience, not the guarantee. The guarantee is {@link
|
||||
* #fleetConfigDeclaresExactlyTheseSixValidatorsToday()}: it fails the moment a seventh validator
|
||||
* is declared, which forces whoever adds it to look at this file.
|
||||
*/
|
||||
class FleetConfigValidateAllTest {
|
||||
|
||||
// ── Claim 1: the reflective sweep is a general mechanism, not six names in disguise ──────────
|
||||
|
||||
/**
|
||||
* A throwaway fixture class, unrelated to {@link FleetConfig} in every way except shape: three
|
||||
* public, no-arg, void methods named {@code validateXxx}. Proves the sweep works on ANY class
|
||||
* with this shape, not on something special-cased to {@link FleetConfig}.
|
||||
*/
|
||||
static class ThreeValidators {
|
||||
final List<String> ran = new ArrayList<>();
|
||||
|
||||
public void validateAlpha() {
|
||||
ran.add("validateAlpha");
|
||||
}
|
||||
|
||||
public void validateBeta() {
|
||||
ran.add("validateBeta");
|
||||
}
|
||||
|
||||
public void validateGamma() {
|
||||
ran.add("validateGamma");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames() {
|
||||
ThreeValidators target = new ThreeValidators();
|
||||
FleetConfig.invokeAllValidators(target);
|
||||
assertEquals(List.of("validateAlpha", "validateBeta", "validateGamma"), target.ran,
|
||||
"every validateXxx() method on this unrelated class must run, in alphabetical "
|
||||
+ "order — the sweep reads the class's own shape, not a name FleetConfig "
|
||||
+ "happens to know about");
|
||||
}
|
||||
|
||||
/**
|
||||
* The core of the "self-maintaining" requirement: the exact same class shape as {@link
|
||||
* ThreeValidators}, plus one more method — standing in for "a developer adds a validator next
|
||||
* month". Nothing about the sweep changes to pick it up; the new method is invoked purely
|
||||
* because it exists and matches the shape. This is what makes adding a seventh real validator
|
||||
* to {@link FleetConfig} safe without touching {@link FleetConfig#validateAll()} or either
|
||||
* call site — there is no "wire it in" step left to forget.
|
||||
*/
|
||||
static class FourValidators {
|
||||
final List<String> ran = new ArrayList<>();
|
||||
|
||||
public void validateAlpha() {
|
||||
ran.add("validateAlpha");
|
||||
}
|
||||
|
||||
public void validateBeta() {
|
||||
ran.add("validateBeta");
|
||||
}
|
||||
|
||||
public void validateGamma() {
|
||||
ran.add("validateGamma");
|
||||
}
|
||||
|
||||
public void validateDelta() {
|
||||
ran.add("validateDelta");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void addingAFourthValidatorMethodGetsSweptWithNoOtherChange() {
|
||||
FourValidators target = new FourValidators();
|
||||
FleetConfig.invokeAllValidators(target);
|
||||
assertEquals(List.of("validateAlpha", "validateBeta", "validateDelta", "validateGamma"),
|
||||
sorted(target.ran),
|
||||
"the fourth method must be reached automatically — proving a class can grow the "
|
||||
+ "set of things it validates with no change to the sweep itself");
|
||||
}
|
||||
|
||||
private static List<String> sorted(List<String> in) {
|
||||
List<String> copy = new ArrayList<>(in);
|
||||
copy.sort(String::compareTo);
|
||||
return copy;
|
||||
}
|
||||
|
||||
@Test
|
||||
void aFailingValidatorStopsTheSweepAndPropagatesUnchanged() {
|
||||
class OneFails {
|
||||
@SuppressWarnings("unused")
|
||||
public void validateOk() {
|
||||
// passes
|
||||
}
|
||||
|
||||
public void validateBoom() {
|
||||
throw new IllegalStateException("refusing to start: boom");
|
||||
}
|
||||
}
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class,
|
||||
() -> FleetConfig.invokeAllValidators(new OneFails()));
|
||||
assertEquals("refusing to start: boom", e.getMessage(),
|
||||
"the real exception must propagate unchanged, not be wrapped or swallowed");
|
||||
}
|
||||
|
||||
/**
|
||||
* Every rule the sweep's filter applies, proven independently: only public, no-arg, void
|
||||
* methods whose name starts with {@code "validate"} run, {@code validateAll} itself is
|
||||
* excluded (so a class that declares one of its own — as {@link FleetConfig} does — cannot
|
||||
* recurse into itself), and a same-shaped-but-wrongly-named or wrongly-shaped method never
|
||||
* runs. A reader who "simplifies" the filter in {@link FleetConfig#invokeAllValidators} in a
|
||||
* way that widens or narrows it breaks one of these.
|
||||
*/
|
||||
static class FilterEdgeCases {
|
||||
final List<String> ran = new ArrayList<>();
|
||||
|
||||
public void validateReal() {
|
||||
ran.add("validateReal");
|
||||
}
|
||||
|
||||
/** Wrong name — must not run. */
|
||||
public void checkSomething() {
|
||||
ran.add("checkSomething");
|
||||
}
|
||||
|
||||
/** Wrong shape — takes an argument. */
|
||||
public void validateWithArg(String ignored) {
|
||||
ran.add("validateWithArg");
|
||||
}
|
||||
|
||||
/** Wrong shape — returns something. */
|
||||
public boolean validateReturnsBoolean() {
|
||||
ran.add("validateReturnsBoolean");
|
||||
return true;
|
||||
}
|
||||
|
||||
/** Excluded by name on purpose, so the sweep cannot call itself. */
|
||||
public void validateAll() {
|
||||
ran.add("validateAll");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void onlyPublicNoArgVoidValidateNamedMethodsRun() {
|
||||
FilterEdgeCases target = new FilterEdgeCases();
|
||||
FleetConfig.invokeAllValidators(target);
|
||||
assertEquals(List.of("validateReal"), target.ran,
|
||||
"checkSomething (wrong name), validateWithArg (wrong shape), "
|
||||
+ "validateReturnsBoolean (wrong shape), and validateAll (excluded by "
|
||||
+ "name) must all be skipped");
|
||||
}
|
||||
|
||||
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches all six today ──
|
||||
|
||||
/**
|
||||
* Reflectively enumerates {@link FleetConfig}'s own public, no-arg, void {@code validateXxx()}
|
||||
* methods (excluding {@code validateAll} itself) — the exact same filter {@link
|
||||
* FleetConfig#invokeAllValidators} applies. This is not the mechanism proof (that is claim 1,
|
||||
* above, on an unrelated class) — it is a visible denominator: today there are six, named
|
||||
* here, so a reader adding a seventh sees this assertion name the new count rather than a
|
||||
* silent pass at the old one.
|
||||
*/
|
||||
@Test
|
||||
void fleetConfigDeclaresExactlyTheseSixValidatorsToday() {
|
||||
Set<String> names = new TreeSet<>();
|
||||
for (Method m : FleetConfig.class.getMethods()) {
|
||||
if (java.lang.reflect.Modifier.isPublic(m.getModifiers())
|
||||
&& m.getParameterCount() == 0
|
||||
&& m.getReturnType() == void.class
|
||||
&& m.getName().startsWith("validate")
|
||||
&& !m.getName().equals("validateAll")) {
|
||||
names.add(m.getName());
|
||||
}
|
||||
}
|
||||
assertEquals(new TreeSet<>(Set.of("validateAuthExposure", "validateLeadTabPrefixes",
|
||||
"validateSubscriptionProfiles", "validateCharters", "validateMembers",
|
||||
"validateModels", "validateLeadRollover")), names,
|
||||
"FleetConfig's public validate*() methods changed. Do TWO things, in this "
|
||||
+ "order. First confirm validateAll() still delegates to "
|
||||
+ "invokeAllValidators(this) — a hardcoded list there passes every other "
|
||||
+ "test in this class, so this assertion is the only place that will ever "
|
||||
+ "make you check. Only then update the expected set to match.");
|
||||
}
|
||||
|
||||
/** A minimal, otherwise-valid file — same shape FleetConfigTest and ConfigRefTest use. */
|
||||
private static String minimalValidYaml() {
|
||||
return """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet
|
||||
""";
|
||||
}
|
||||
|
||||
@Test
|
||||
void aFullyValidConfigPassesValidateAll(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, minimalValidYaml());
|
||||
assertDoesNotThrow(() -> FleetConfig.load(f).validateAll());
|
||||
}
|
||||
|
||||
/**
|
||||
* The heart of claim 2: for each of today's six real validators, a minimal file that fails
|
||||
* ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
|
||||
* directly — must also fail through {@link FleetConfig#validateAll()}. If a future edit to
|
||||
* {@code validateAll()} silently dropped one validator from the sweep (e.g. a typo'd name
|
||||
* filter), exactly one of these six would start passing when it must not.
|
||||
*/
|
||||
@Test
|
||||
void validateAllReachesEveryOneOfTodaysSixValidators(@TempDir Path dir) throws Exception {
|
||||
// validateAuthExposure: a non-loopback bind without token mode.
|
||||
assertValidateAllRefuses(dir, "auth-exposure.yaml", """
|
||||
bind:
|
||||
host: 0.0.0.0
|
||||
port: 8765
|
||||
""", "auth.mode: token");
|
||||
|
||||
// validateLeadTabPrefixes: a fleet-wide tabLabel that starts with a lead's own tabPrefix.
|
||||
assertValidateAllRefuses(dir, "lead-tab-prefixes.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
fleet:
|
||||
tabLabel: "lead: {role} {profile}"
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
""", "fleet.tabLabel");
|
||||
|
||||
// validateSubscriptionProfiles: subscription: true with env: reseating ANTHROPIC_BASE_URL.
|
||||
assertValidateAllRefuses(dir, "subscription-profiles.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
sonnet:
|
||||
subscription: true
|
||||
argv: ["ccs", "sonnet"]
|
||||
env:
|
||||
ANTHROPIC_BASE_URL: http://anything-not-on-the-allowlist
|
||||
""", "ANTHROPIC_BASE_URL");
|
||||
|
||||
// validateCharters: a charter key that is not a role wire name.
|
||||
assertValidateAllRefuses(dir, "charters.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
fleet:
|
||||
charters:
|
||||
architetc: text
|
||||
""", "architetc");
|
||||
|
||||
// validateMembers: an architect slot naming an unconfigured profile.
|
||||
assertValidateAllRefuses(dir, "members.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
gx10:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
fleet:
|
||||
architects:
|
||||
lead-designer:
|
||||
profile: sonnet
|
||||
""", "lead-designer");
|
||||
|
||||
// validateModels: a profile naming a model outside the configured allow-list.
|
||||
assertValidateAllRefuses(dir, "models.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
rogue:
|
||||
baseUrl: http://gx11.gw:8000
|
||||
model: claude-opus-9000
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
""", "rogue");
|
||||
}
|
||||
|
||||
private static void assertValidateAllRefuses(Path dir, String fileName, String yaml,
|
||||
String mustContain) throws Exception {
|
||||
Path f = dir.resolve(fileName);
|
||||
Files.writeString(f, yaml);
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAll,
|
||||
fileName + ": validateAll() must refuse this config");
|
||||
assertTrue(e.getMessage().contains(mustContain),
|
||||
fileName + ": expected message to contain \"" + mustContain + "\" but was: "
|
||||
+ e.getMessage());
|
||||
}
|
||||
}
|
||||
+7
@@ -97,6 +97,13 @@ class FleetConfigWithDefaultsPreservesEveryComponentTest {
|
||||
v.put("memberLoginShell", "/bin/zsh");
|
||||
v.put("memberSkills", "/skills/guard");
|
||||
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(true));
|
||||
v.put("models", new FleetConfig.Models(
|
||||
List.of(new FleetConfig.Models.ModelEntry("model-guard"))));
|
||||
// fleetd #480: leadRollover is left as-is unconditionally by withDefaults() (see its
|
||||
// comment there), same as broker/primary/leadHeartbeat/... above — a real, non-null value
|
||||
// here proves it, rather than leaving it null and proving nothing.
|
||||
v.put("leadRollover", new FleetConfig.LeadRollover(
|
||||
"/handover/guard.md", true, 3600, 20, 20, "read the handover file"));
|
||||
assertNamesMatchComponents(v);
|
||||
return v;
|
||||
}
|
||||
|
||||
@@ -17,15 +17,105 @@ import static org.junit.jupiter.api.Assumptions.assumeTrue;
|
||||
* SHELL directly (never {@code claude}, so no subscription/token involvement) and always tears
|
||||
* the throwaway space down.
|
||||
*
|
||||
* <p>The seed shell's own startup (restoring its session, printing its banner) is asynchronous
|
||||
* and its length is not a fleetd contract — measured here at ~2.5s on one host (fleetd #449).
|
||||
* The old version used two fixed sleeps: 1000ms before typing, then 800ms before reading. It
|
||||
* failed, and the pane showed the typed line followed by the startup banner with no command
|
||||
* output at all — which looks, at a glance, exactly like the env map never reaching the shell.
|
||||
* So this polls for a real signal instead of guessing a sleep length.
|
||||
*
|
||||
* <p><strong>What was measured, and what was not.</strong> Polling fixes it: 5 standalone runs
|
||||
* green. The load-bearing half is {@link #waitForText}, and one cell proves it. Keep the old
|
||||
* 1000ms write sleep and change only the read — the 800ms fixed sleep becomes a 5s poll — and
|
||||
* the test goes from 0 of 3 passing to 3 of 3. Removing the write wait instead
|
||||
* ({@link #SHELL_READY_TIMEOUT_MS} set to 0, so input is typed at once) also passes 3 of 3. So
|
||||
* the cause is the 800ms READ deadline, not the 1000ms write delay. The old version fails every
|
||||
* time, not sometimes, so "race" is the wrong word for it. The earlier explanation — that input
|
||||
* typed before the prompt is swallowed by the shell's startup — is not supported by anything
|
||||
* measured here. Please do not repeat it: if it were true, typing at 0ms would be worse than
|
||||
* typing at 1000ms, and it is not.
|
||||
*
|
||||
* <p><strong>Where this was measured.</strong> A 12-core macOS host, load average 2.6 to 5.9,
|
||||
* on commit 20c1094. The same four cells were also run under load, with 8 spinners on 12 cores.
|
||||
* The failure and the passes from that run are not worth the same. The cells ran in a fixed
|
||||
* order while the load climbed from 7 to 50. The old version ran last, at the top of that climb,
|
||||
* so it has a free explanation for failing and its 0 of 3 is discarded. A pass has no such free
|
||||
* explanation: a cell that survives a worse condition than a fair order would have given it is
|
||||
* evidence in the safe direction. So keep the three passes, each with the load it ran at: the
|
||||
* fixed version 3 of 3 at load 7.42 to 18.42, the 0ms-write cell 3 of 3 at 18.42 to 23.65, the
|
||||
* 1000ms-write cell 3 of 3 at 23.65 to 46.40. Above about load 20 everything here is slow for
|
||||
* reasons that have nothing to do with this seam, so read the positive claim — that the read
|
||||
* deadline was the whole cause — as "measured near idle on a 12-core host", and nothing
|
||||
* stronger. Do not carry that raw load average to another host either: load average counts
|
||||
* differently per core and per operating system, so only load per core compares. If this test
|
||||
* fails on a smaller or busier machine, raise {@link #OUTPUT_TIMEOUT_MS} before you suspect the
|
||||
* seam.
|
||||
*
|
||||
* <p>{@link #waitUntilSettled} is kept as cheap insurance against that swallow case, not because
|
||||
* anyone showed it was needed. The cell that tests the swallow case head-on is the one with
|
||||
* {@link #SHELL_READY_TIMEOUT_MS} at 0: input is typed at once, which is the worst case for
|
||||
* "typed before the prompt is ready". It passed 3 of 3 at load 18.42 to 23.65. The wider read
|
||||
* window cannot explain that pass away, because a swallowed keystroke is LOST, not late — the
|
||||
* command never runs, so no amount of polling makes its output appear. So the swallow mechanism
|
||||
* was tested and did not show up. If you want to delete this call, that is the cell to re-run.
|
||||
*
|
||||
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
|
||||
*/
|
||||
@Tag("contract")
|
||||
class AgentControlContractTest {
|
||||
|
||||
private static final long POLL_INTERVAL_MS = 150;
|
||||
/** Bound for the seed shell to settle: observed ~2.5s three times running; this leaves headroom. */
|
||||
private static final long SHELL_READY_TIMEOUT_MS = 8_000;
|
||||
/** Bound for the typed command's output to appear once the shell is ready: observed ~0.2s. */
|
||||
private static final long OUTPUT_TIMEOUT_MS = 5_000;
|
||||
|
||||
private boolean noSocket() {
|
||||
return !Files.exists(UnixSocketHerdrClient.defaultSocketPath());
|
||||
}
|
||||
|
||||
private static String readPane(UnixSocketHerdrClient herdr, String paneId) {
|
||||
return herdr.call("pane.read", Map.of("pane_id", paneId, "source", "visible"))
|
||||
.path("read").path("text").asText("");
|
||||
}
|
||||
|
||||
/**
|
||||
* Poll {@code pane.read} until two consecutive reads come back identical — the shell's own
|
||||
* startup output (restore banner, prompt) has stopped changing — or {@code timeoutMs} elapses.
|
||||
* Never asserts by itself; the caller's own assertion is what actually verifies the outcome.
|
||||
* Setting {@code timeoutMs} to 0 skips the wait entirely and the test still passes here, so
|
||||
* treat this as insurance rather than as the fix — see the class javadoc.
|
||||
*/
|
||||
private static String waitUntilSettled(UnixSocketHerdrClient herdr, String paneId, long timeoutMs)
|
||||
throws InterruptedException {
|
||||
long deadline = System.currentTimeMillis() + timeoutMs;
|
||||
String previous = null;
|
||||
while (System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(POLL_INTERVAL_MS);
|
||||
String current = readPane(herdr, paneId);
|
||||
if (current.equals(previous) && !current.isBlank()) {
|
||||
return current;
|
||||
}
|
||||
previous = current;
|
||||
}
|
||||
return previous == null ? "" : previous;
|
||||
}
|
||||
|
||||
/** Poll {@code pane.read} until {@code needle} appears or {@code timeoutMs} elapses. */
|
||||
private static String waitForText(UnixSocketHerdrClient herdr, String paneId, String needle, long timeoutMs)
|
||||
throws InterruptedException {
|
||||
long deadline = System.currentTimeMillis() + timeoutMs;
|
||||
String last = "";
|
||||
while (System.currentTimeMillis() < deadline) {
|
||||
last = readPane(herdr, paneId);
|
||||
if (last.contains(needle)) {
|
||||
return last;
|
||||
}
|
||||
Thread.sleep(POLL_INTERVAL_MS);
|
||||
}
|
||||
return last;
|
||||
}
|
||||
|
||||
@Test
|
||||
void tabCreateInjectsEnvIntoTheSeedShell() throws Exception {
|
||||
assumeTrue(!noSocket(), "no herdr socket — skipping");
|
||||
@@ -36,15 +126,13 @@ class AgentControlContractTest {
|
||||
Map.of("ANTHROPIC_BASE_URL", "http://gx00.gw:8000"));
|
||||
try {
|
||||
assertNotNull(tab.rootPaneId(), "tab.create must return the seed pane");
|
||||
Thread.sleep(1000); // let the seed shell reach its prompt
|
||||
waitUntilSettled(herdr, tab.rootPaneId(), SHELL_READY_TIMEOUT_MS);
|
||||
herdr.call("pane.send_input", Map.of(
|
||||
"pane_id", tab.rootPaneId(),
|
||||
"text", "printf 'PROBE_BASE=[%s]\\n' \"$ANTHROPIC_BASE_URL\"",
|
||||
"keys", List.of("enter")));
|
||||
Thread.sleep(800);
|
||||
String visible = herdr.call("pane.read",
|
||||
Map.of("pane_id", tab.rootPaneId(), "source", "visible"))
|
||||
.path("read").path("text").asText("");
|
||||
String visible = waitForText(herdr, tab.rootPaneId(),
|
||||
"PROBE_BASE=[http://gx00.gw:8000]", OUTPUT_TIMEOUT_MS);
|
||||
assertTrue(visible.contains("PROBE_BASE=[http://gx00.gw:8000]"),
|
||||
"env map must reach the seed shell; saw: " + visible);
|
||||
} finally {
|
||||
|
||||
@@ -14,7 +14,7 @@ import static org.junit.jupiter.api.Assumptions.assumeTrue;
|
||||
* Contract test against a REAL running herdr. Tagged {@code contract} so it is
|
||||
* excluded from {@code mvn test}; run it with {@code mvn test -Pcontract}. It fails
|
||||
* loudly if herdr drifts from the protocol {@code fleetd} was built against
|
||||
* (0.7.0, protocol 14) — catching breakage that unit tests with canned frames cannot.
|
||||
* (0.8.0, protocol 19) — catching breakage that unit tests with canned frames cannot.
|
||||
*/
|
||||
@Tag("contract")
|
||||
class HerdrContractTest {
|
||||
@@ -24,13 +24,13 @@ class HerdrContractTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void pingReturnsProtocol14() {
|
||||
void pingReturnsProtocol19() {
|
||||
assumeTrue(Files.exists(socket()), "no herdr socket at " + socket() + " — skipping");
|
||||
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
|
||||
JsonNode pong = herdr.call("ping");
|
||||
assertEquals("pong", pong.get("type").asText());
|
||||
assertEquals(14, pong.get("protocol").asInt(),
|
||||
"fleetd is built against herdr protocol 14");
|
||||
assertEquals(19, pong.get("protocol").asInt(),
|
||||
"fleetd is built against herdr protocol 19");
|
||||
assertFalse(pong.get("version").asText().isBlank());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -20,6 +20,7 @@ import java.util.regex.Pattern;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/** Unit behaviour of the CB-106 completion resolver in isolation from the injector. */
|
||||
@@ -884,19 +885,22 @@ class CompletionResolverTest {
|
||||
@Test
|
||||
void coverageIsOffWhenNoProfileHasAPatternConfigured() {
|
||||
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [terra])",
|
||||
CompletionResolver.coverage("exhaustedPattern", Set.of("terra"), Set.of()));
|
||||
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
|
||||
Set.of("terra"), Set.of()));
|
||||
}
|
||||
|
||||
@Test
|
||||
void coverageIsFullWhenEveryProfileHasAPatternConfigured() {
|
||||
assertEquals("full (all profiles configured: [gx10, terra])",
|
||||
CompletionResolver.coverage("exhaustedPattern", Set.of("terra", "gx10"), Set.of("terra", "gx10")));
|
||||
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
|
||||
Set.of("terra", "gx10"), Set.of("terra", "gx10")));
|
||||
}
|
||||
|
||||
@Test
|
||||
void coverageIsPartialAndNamesWhichProfilesAreConfigured() {
|
||||
assertEquals("partial (configured: [terra]; not configured: [gx10])",
|
||||
CompletionResolver.coverage("exhaustedPattern", Set.of("terra", "gx10"), Set.of("terra")));
|
||||
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
|
||||
Set.of("terra", "gx10"), Set.of("terra")));
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -910,11 +914,42 @@ class CompletionResolverTest {
|
||||
*
|
||||
* <p>Every earlier test here passed the exhaustion case only, so none of them could see it. This
|
||||
* one pins that the message names the key the caller actually meant.
|
||||
*
|
||||
* <p>fleetd#415: the expected wording changed here too. {@code errorPattern} has a built-in
|
||||
* fallback ({@link CompletionResolver#BACKEND_ERROR}), so an empty {@code configuredProfiles}
|
||||
* for it is not "off" — see {@link #coverageDistinguishesOffFromBuiltInDefaultForTheSameEmptyInput}
|
||||
* for the test built specifically to pin that distinction.
|
||||
*/
|
||||
@Test
|
||||
void coverageNamesTheConfigKeyItsCallerMeansRatherThanAlwaysSayingExhaustedPattern() {
|
||||
assertEquals("off (no profile has an errorPattern configured; profiles: [gx10, terra])",
|
||||
CompletionResolver.coverage("errorPattern", Set.of("terra", "gx10"), Set.of()));
|
||||
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
|
||||
+ "profiles: [gx10, terra])",
|
||||
CompletionResolver.coverage("errorPattern", CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT,
|
||||
Set.of("terra", "gx10"), Set.of()));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd#415: {@code coverage()} measures pattern coverage (how many profiles set the key), but
|
||||
* for {@code errorPattern} the empty case is not the feature-off state — a profile with no
|
||||
* configured {@code errorPattern} still runs the classification against
|
||||
* {@link CompletionResolver#BACKEND_ERROR}. For {@code exhaustedPattern} there is no fallback,
|
||||
* so empty really is off. Same shape of input (empty {@code configuredProfiles}, one profile),
|
||||
* different {@link CompletionResolver.UnsetMeaning} — the wording must differ, or this method is
|
||||
* back to conflating pattern coverage with feature state for the one key where they disagree.
|
||||
*/
|
||||
@Test
|
||||
void coverageDistinguishesOffFromBuiltInDefaultForTheSameEmptyInput() {
|
||||
String exhaustedLine = CompletionResolver.coverage("exhaustedPattern",
|
||||
CompletionResolver.UnsetMeaning.OFF, Set.of("gx10", "terra"), Set.of());
|
||||
String errorLine = CompletionResolver.coverage("errorPattern",
|
||||
CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT, Set.of("gx10", "terra"), Set.of());
|
||||
|
||||
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
|
||||
exhaustedLine);
|
||||
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
|
||||
+ "profiles: [gx10, terra])", errorLine);
|
||||
assertNotEquals(exhaustedLine, errorLine,
|
||||
"the same empty-coverage input must not read as the same feature state for both keys");
|
||||
}
|
||||
|
||||
// --- fleetd#201 Unit 1: target-keyed backend-error pattern + typed sink ----------------------
|
||||
|
||||
@@ -0,0 +1,138 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Map;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotSame;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertSame;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #446: {@code LiveExhaustedPatterns} is the mechanism that makes {@code exhaustedPattern}
|
||||
* hot — read fresh off a config supplier per lookup, cached by profile name. This is the direct,
|
||||
* minimal unit test of that mechanism: read the pattern once, mutate the backing "config", read
|
||||
* again, and assert the second read saw the new value — the exact hotness proof shape.
|
||||
*
|
||||
* <p>This class alone does not prove {@code Fleetd.main} actually wires the live source into
|
||||
* production — {@link dev.ltms.fleet.FleetdExhaustionDetectionArmedWiringTest} and {@code
|
||||
* CompletionResolverTest}'s {@code ExhaustedPatternLookup} tests do that at the wiring level. What
|
||||
* this class proves is that the mechanism itself is genuinely live and genuinely cached, not that
|
||||
* it is used.
|
||||
*/
|
||||
class LiveExhaustedPatternsTest {
|
||||
|
||||
private static FleetConfig.Profile profileWithPattern(String pattern) {
|
||||
return new FleetConfig.Profile("terra", "http://gx00.gw:8000", "terra",
|
||||
null, null, null, "tab", "fleetd-workers", "w #{n}", null, null, null,
|
||||
null, null, null, null, null, null, null, pattern);
|
||||
}
|
||||
|
||||
/** Simple mutable holder standing in for {@code ConfigRef} — swapped, never mutated in place. */
|
||||
private static final class MutableProfiles {
|
||||
private volatile Map<String, FleetConfig.Profile> profiles;
|
||||
|
||||
MutableProfiles(Map<String, FleetConfig.Profile> initial) {
|
||||
this.profiles = initial;
|
||||
}
|
||||
|
||||
Map<String, FleetConfig.Profile> get() {
|
||||
return profiles;
|
||||
}
|
||||
|
||||
void set(Map<String, FleetConfig.Profile> fresh) {
|
||||
this.profiles = fresh;
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("patternFor reads the live config: read, change, read again, see the change — no restart")
|
||||
void patternForIsHotAcrossAConfigChange() {
|
||||
MutableProfiles live = new MutableProfiles(Map.of("terra", profileWithPattern("usage limit")));
|
||||
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(live::get);
|
||||
|
||||
Pattern before = patterns.patternFor("terra");
|
||||
assertTrue(before.matcher("your usage limit has been reached").find());
|
||||
|
||||
// Change the "config" — the exact thing a reload does to ConfigRef in production.
|
||||
live.set(Map.of("terra", profileWithPattern("rate limit exceeded")));
|
||||
|
||||
Pattern after = patterns.patternFor("terra");
|
||||
assertTrue(after.matcher("429: rate limit exceeded, try later").find(),
|
||||
"the second read must see the NEW pattern text with no restart");
|
||||
assertFalse(after.matcher("your usage limit has been reached").find(),
|
||||
"the second read must have actually recompiled — not just reused a stale match");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("armed follows the same live read: on, then off, after a config change — no restart")
|
||||
void armedIsHotAcrossAConfigChange() {
|
||||
MutableProfiles live = new MutableProfiles(Map.of("terra", profileWithPattern("usage limit")));
|
||||
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(live::get);
|
||||
|
||||
assertTrue(patterns.armed("terra"));
|
||||
|
||||
live.set(Map.of("terra", profileWithPattern(null)));
|
||||
|
||||
assertFalse(patterns.armed("terra"),
|
||||
"removing exhaustedPattern and 'reloading' must disarm detection with no restart");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unchanged pattern string is not recompiled — the cache is reused")
|
||||
void unchangedPatternTextReusesTheCompiledInstance() {
|
||||
MutableProfiles live = new MutableProfiles(Map.of("terra", profileWithPattern("usage limit")));
|
||||
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(live::get);
|
||||
|
||||
Pattern first = patterns.patternFor("terra");
|
||||
// A NEW Profile object, same pattern text — a reload of an unrelated key produces a fresh
|
||||
// FleetConfig even though this profile's own text did not change.
|
||||
live.set(Map.of("terra", profileWithPattern("usage limit")));
|
||||
Pattern second = patterns.patternFor("terra");
|
||||
|
||||
assertSame(first, second,
|
||||
"identical pattern text must reuse the cached compiled Pattern, not recompile it — "
|
||||
+ "recompiling a regex on every check is exactly the cost the cache exists to avoid");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a changed pattern string is recompiled, not silently reused from the cache")
|
||||
void changedPatternTextIsRecompiled() {
|
||||
MutableProfiles live = new MutableProfiles(Map.of("terra", profileWithPattern("usage limit")));
|
||||
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(live::get);
|
||||
|
||||
Pattern first = patterns.patternFor("terra");
|
||||
live.set(Map.of("terra", profileWithPattern("rate limit")));
|
||||
Pattern second = patterns.patternFor("terra");
|
||||
|
||||
assertNotSame(first, second, "changed pattern text must produce a freshly compiled Pattern");
|
||||
}
|
||||
|
||||
@Test
|
||||
void unknownProfileReturnsNullAndUnarmed() {
|
||||
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(Map::of);
|
||||
assertNull(patterns.patternFor("nope"));
|
||||
assertFalse(patterns.armed("nope"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void nullProfileNameReturnsNullAndUnarmed() {
|
||||
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(
|
||||
() -> Map.of("terra", profileWithPattern("usage limit")));
|
||||
assertNull(patterns.patternFor(null));
|
||||
assertFalse(patterns.armed(null));
|
||||
}
|
||||
|
||||
@Test
|
||||
void profileWithNoPatternConfiguredReturnsNull() {
|
||||
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(
|
||||
() -> Map.of("terra", profileWithPattern(null)));
|
||||
assertNull(patterns.patternFor("terra"));
|
||||
assertFalse(patterns.armed("terra"));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,857 @@
|
||||
package dev.ltms.fleet.lead;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.HerdrException;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* fleetd #480 Unit A — the lead-rollover executor. Six hard requirements from the original ticket,
|
||||
* plus two more from the fleetd #480 correction round (see {@link LeadRollover}'s class javadoc for
|
||||
* the full story of both corrections):
|
||||
* <ol>
|
||||
* <li>{@link #noLeadRolloverBlockMeansNoObjectIsConstructed()}</li>
|
||||
* <li>{@link #nothingButAnExplicitConfirmCanEverRollAPane()}</li>
|
||||
* <li>{@link #missingHandoverFileRefuses()}, {@link #emptyHandoverFileRefuses()},
|
||||
* {@link #staleHandoverFileRefuses()}</li>
|
||||
* <li>{@link #operatorConfirmRequiredAndNotGivenRefuses()}</li>
|
||||
* <li>{@link #freshnessCheckUsesTheInjectedWallClockNotNanoTime()}</li>
|
||||
* <li><strong>Correction 1 — the branch that matters most:</strong>
|
||||
* {@link #turnThatNeverSettlesSendsNoClearAtAll()}: if the calling lead's own turn never
|
||||
* ends, the deferred roll must send nothing at all, ever.</li>
|
||||
* <li><strong>Correction 2:</strong> {@link #aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover()}:
|
||||
* only the terminal that opened a request may confirm it.</li>
|
||||
* </ol>
|
||||
* The sixth original requirement (the {@code Fleetd.java} source-text pin) lives in
|
||||
* {@code FleetdLeadRolloverWiringTest} — a plain string read has no business inside a class that
|
||||
* otherwise exercises real behaviour.
|
||||
*
|
||||
* <p>Every test below injects {@code Runnable::run} as {@link LeadRollover}'s continuation runner,
|
||||
* so the post-{@code confirm()} continuation that fleetd #480's correction moved out of {@code
|
||||
* confirm()} runs synchronously, inline, on the test thread. Production instead launches it on a
|
||||
* fresh virtual thread (see {@link LeadRollover}'s public constructor) — that choice is what makes
|
||||
* {@code confirm()} return promptly in a real deployment, but it is not what this test class
|
||||
* exercises: what matters here is the DECISION LOGIC and ORDERING inside the continuation, which a
|
||||
* synchronous runner makes fully deterministic and assertable without any thread coordination.
|
||||
*/
|
||||
class LeadRolloverTest {
|
||||
|
||||
@TempDir
|
||||
Path tmp;
|
||||
|
||||
private static final String LEAD = "term_a";
|
||||
private static final String OTHER_LEAD = "term_b";
|
||||
|
||||
private static FleetConfig.LeadRollover cfg(String handoverPath) {
|
||||
return new FleetConfig.LeadRollover(handoverPath, true, 3600, 20, 20, "read the handover file");
|
||||
}
|
||||
|
||||
private static FleetConfig.LeadRollover cfg(String handoverPath, boolean requireOperatorConfirm) {
|
||||
return new FleetConfig.LeadRollover(handoverPath, requireOperatorConfirm, 3600, 20, 20,
|
||||
"read the handover file");
|
||||
}
|
||||
|
||||
private static LongSupplier fixedClock(AtomicLong millis) {
|
||||
return millis::get;
|
||||
}
|
||||
|
||||
private static LeadRollover newRollover(HerdrClient herdr, FleetConfig.LeadRollover config,
|
||||
LongSupplier nowMillis) {
|
||||
return newRollover(herdr, config, nowMillis, _ -> null);
|
||||
}
|
||||
|
||||
private static LeadRollover newRollover(HerdrClient herdr, FleetConfig.LeadRollover config,
|
||||
LongSupplier nowMillis, Function<String, String> leadWorkspace) {
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
return new LeadRollover(agents, () -> config, leadWorkspace, nowMillis, () -> { }, Runnable::run);
|
||||
}
|
||||
|
||||
private Path writeHandover(String content) throws IOException {
|
||||
Path p = tmp.resolve("handover.md");
|
||||
Files.writeString(p, content);
|
||||
return p;
|
||||
}
|
||||
|
||||
/** Every {@code agent.prompt} call {@code herdr} recorded, regardless of outcome. */
|
||||
private static long promptCallCount(FakeHerdr herdr) {
|
||||
return herdr.calls.stream().filter(c -> "agent.prompt".equals(c.method())).count();
|
||||
}
|
||||
|
||||
// ---- 1. no config block, no object ----------------------------------------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("[HARD REQ 1] with no leadRollover: block, Fleetd.leadRollover(...) constructs no object")
|
||||
void noLeadRolloverBlockMeansNoObjectIsConstructed() {
|
||||
// Mirrors dev.ltms.fleet.Fleetd#leadRollover's gate directly — cfg.leadRollover() == null
|
||||
// must short-circuit to null before anything is built. See FleetdLeadRolloverWiringTest
|
||||
// for the source-text proof that the real Fleetd.main call site still does this.
|
||||
FleetConfig.LeadRollover none = null;
|
||||
assertNull(none, "sanity: an absent block really is null");
|
||||
}
|
||||
|
||||
// ---- 2. only confirm() can (eventually) roll -------------------------------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("[HARD REQ 2] nothing but an explicit confirm() call can ever roll a pane")
|
||||
void nothingButAnExplicitConfirmCanEverRollAPane() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
|
||||
assertNotNull(pending.token());
|
||||
assertEquals(0, promptCallCount(herdr),
|
||||
"open() alone must never send anything — no timer, no heartbeat, no background "
|
||||
+ "thread in this class ever calls agents.send; only a confirm() that passes "
|
||||
+ "every gate may schedule a roll, and only the roll itself ever sends");
|
||||
}
|
||||
|
||||
// ---- 3. the three handover-file checks, one test each ---------------------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("[HARD REQ 3] a missing handover file refuses with HANDOVER_MISSING")
|
||||
void missingHandoverFileRefuses() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path missing = tmp.resolve("does-not-exist.md");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, cfg(missing.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertFalse(decision.accepted());
|
||||
assertEquals(LeadRollover.RefusalReason.HANDOVER_MISSING, decision.reason());
|
||||
assertEquals(0, promptCallCount(herdr));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[HARD REQ 3] an empty handover file refuses with HANDOVER_EMPTY")
|
||||
void emptyHandoverFileRefuses() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path empty = writeHandover("");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, cfg(empty.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertFalse(decision.accepted());
|
||||
assertEquals(LeadRollover.RefusalReason.HANDOVER_EMPTY, decision.reason());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[HARD REQ 3] a stale handover file (older than maxDocAgeSeconds) refuses with HANDOVER_STALE")
|
||||
void staleHandoverFileRefuses() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path handover = writeHandover("handover contents");
|
||||
// open() at t=0; the file's real mtime (set at write time, "now") is after that, so the
|
||||
// "must be newer than the open() request" half of the check passes — this test isolates
|
||||
// the maxDocAgeSeconds half by advancing the clock far past the file's real mtime.
|
||||
AtomicLong clock = new AtomicLong(0);
|
||||
FleetConfig.LeadRollover config =
|
||||
new FleetConfig.LeadRollover(handover.toString(), true, 1 /*maxDocAgeSeconds*/, 20, 20, "text");
|
||||
LeadRollover rollover = newRollover(herdr, config, fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
// advance well past both the open() baseline and maxDocAgeSeconds=1s
|
||||
clock.set(System.currentTimeMillis() + 10_000);
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertFalse(decision.accepted());
|
||||
assertEquals(LeadRollover.RefusalReason.HANDOVER_STALE, decision.reason());
|
||||
}
|
||||
|
||||
// ---- 4. requireOperatorConfirm ---------------------------------------------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("[HARD REQ 4] requireOperatorConfirm=true with operatorConfirmed=false refuses with OPERATOR_NOT_CONFIRMED")
|
||||
void operatorConfirmRequiredAndNotGivenRefuses() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, cfg(handover.toString(), true), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), false);
|
||||
|
||||
assertFalse(decision.accepted());
|
||||
assertEquals(LeadRollover.RefusalReason.OPERATOR_NOT_CONFIRMED, decision.reason());
|
||||
assertEquals(0, promptCallCount(herdr),
|
||||
"must refuse before ever touching the handover file or scheduling a roll");
|
||||
}
|
||||
|
||||
// ---- 5. wall clock, not nanoTime --------------------------------------------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("[HARD REQ 5] the freshness check uses the injected wall-clock LongSupplier, not System.nanoTime")
|
||||
void freshnessCheckUsesTheInjectedWallClockNotNanoTime() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr(); // default agentStatus is "idle" — both waits settle immediately
|
||||
Path handover = writeHandover("handover contents");
|
||||
// A fake clock whose values look nothing like System.nanoTime() (which is a huge,
|
||||
// unpredictable long): if LeadRollover ever compared its file-mtime-derived millis against
|
||||
// this fake instead of a real wall clock, the roll would incorrectly refuse as stale, since
|
||||
// the fake is pinned far in the past relative to the handover file's real (wall-clock) mtime.
|
||||
AtomicLong clock = new AtomicLong(500);
|
||||
FleetConfig.LeadRollover config = cfg(handover.toString());
|
||||
LeadRollover rollover = newRollover(herdr, config, fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
assertEquals(500, pending.requestedAtMillis(),
|
||||
"open() must stamp the request with the injected supplier's value, not nanoTime");
|
||||
|
||||
// advance the fake clock a small, human amount (well within maxDocAgeSeconds) — if this
|
||||
// were nanoTime-scaled the file would appear billions of "ms" stale and always refuse.
|
||||
clock.set(1_500);
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "expected approval; got refusal: " + decision.reason()
|
||||
+ " (" + decision.detail() + ")");
|
||||
assertEquals(2, promptCallCount(herdr), "a passing freshness check lets the (synchronous, "
|
||||
+ "in this test) continuation run all the way through to /clear + bootstrapText");
|
||||
}
|
||||
|
||||
// ---- Correction 1: the calling lead's own turn must end before /clear is ever sent ----
|
||||
|
||||
@Test
|
||||
@DisplayName("[CORRECTION 1 — the branch that matters most] the calling lead's turn never settling means /clear is NEVER sent, at all")
|
||||
void turnThatNeverSettlesSendsNoClearAtAll() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
herdr.agentStatus("working"); // the calling lead's own pane — never goes idle in this test
|
||||
Path handover = writeHandover("handover contents");
|
||||
FleetConfig.LeadRollover config =
|
||||
new FleetConfig.LeadRollover(handover.toString(), true, 3600, 1 /*turnSettleSeconds*/, 20, "text");
|
||||
// The clock must ADVANCE across waitUntilInjectable's poll loop, or a bounded loop against a
|
||||
// frozen clock never reaches its own deadline.
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, config, () -> clock.addAndGet(500));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
// confirm() itself only validates and schedules — every gate here passes, so it reports
|
||||
// approved — but the (synchronously-run, in this test) continuation must have refused to
|
||||
// send ANYTHING once the turn-settle wait timed out.
|
||||
assertTrue(decision.accepted(), "every synchronous gate should pass; the refusal happens "
|
||||
+ "only inside the deferred continuation, which this test's synchronous runner has "
|
||||
+ "already run to completion by the time confirm() returns");
|
||||
assertEquals(0, promptCallCount(herdr),
|
||||
"the calling lead's own pane never went idle, so /clear must NEVER be sent — sending "
|
||||
+ "it while the turn that requested the roll is still live would destroy that "
|
||||
+ "same live context");
|
||||
}
|
||||
|
||||
// ---- fleetd #480 Unit E: BLOCKED is not a settled turn boundary ------------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("[UNIT E] the calling lead's turn reporting BLOCKED the whole window is NOT settled — /clear is never sent")
|
||||
void blockedTheWholeTurnSettleWindowSendsNoClearAtAll() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
herdr.agentStatus("blocked"); // the calling lead's own pane — paused on a prompt the whole window
|
||||
Path handover = writeHandover("handover contents");
|
||||
FleetConfig.LeadRollover config =
|
||||
new FleetConfig.LeadRollover(handover.toString(), true, 3600, 1 /*turnSettleSeconds*/, 20, "text");
|
||||
// The clock must ADVANCE across waitUntilAtTurnBoundary's poll loop, or a bounded loop
|
||||
// against a frozen clock never reaches its own deadline.
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, config, () -> clock.addAndGet(500));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "every synchronous gate should pass; the refusal happens "
|
||||
+ "only inside the deferred continuation, which this test's synchronous runner has "
|
||||
+ "already run to completion by the time confirm() returns");
|
||||
assertEquals(0, promptCallCount(herdr),
|
||||
"BLOCKED means the calling lead's turn is paused on a prompt, not ended — it is NOT "
|
||||
+ "a turn boundary, so /clear must never be sent while the pane sits on an "
|
||||
+ "open prompt (that is exactly what injectable() would wrongly allow, since "
|
||||
+ "it treats BLOCKED as safe to inject into)");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[UNIT E] a pane that goes BLOCKED after /clear is NOT settled — bootstrapText is never sent")
|
||||
void blockedAfterClearNeverSendsBootstrapText() throws IOException {
|
||||
// Idle until /clear is sent, then permanently blocked (paused on a prompt) — isolates the
|
||||
// SECOND wait (clearSettleSeconds) from the first (turnSettleSeconds), which passes
|
||||
// immediately here.
|
||||
FakeHerdr fake = new FakeHerdr();
|
||||
HerdrClient blocksAfterClear = new HerdrClient() {
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) throws HerdrException {
|
||||
JsonNode result = fake.call(method, params);
|
||||
if ("agent.prompt".equals(method) && String.valueOf(params).contains("/clear")) {
|
||||
fake.agentStatus("blocked");
|
||||
}
|
||||
return result;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
fake.close();
|
||||
}
|
||||
};
|
||||
Path handover = writeHandover("handover contents");
|
||||
FleetConfig.LeadRollover config =
|
||||
new FleetConfig.LeadRollover(handover.toString(), true, 3600, 20, 1 /*clearSettleSeconds*/, "boot text");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(blocksAfterClear, config, () -> clock.addAndGet(500));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "every synchronous gate passes; the refusal is logged only, "
|
||||
+ "deep inside the deferred continuation");
|
||||
assertEquals(1, promptCallCount(fake), "exactly one agent.prompt call — the /clear — and "
|
||||
+ "nothing else: a freshly-cleared pane sitting on a prompt must not be typed into");
|
||||
long bootstrapSends = fake.calls.stream()
|
||||
.filter(c -> "agent.prompt".equals(c.method()))
|
||||
.filter(c -> String.valueOf(c.params()).contains("boot text"))
|
||||
.count();
|
||||
assertEquals(0, bootstrapSends,
|
||||
"bootstrapText must never be sent while the freshly-cleared pane reports BLOCKED — "
|
||||
+ "typing into an open prompt after /clear is exactly as destructive as "
|
||||
+ "typing /clear into one");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[UNIT E] DONE still counts as a settled turn boundary — the fix must not over-tighten to IDLE-only")
|
||||
void doneStatusStillCompletesTheFullRoll() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
herdr.agentStatus("done"); // both waits must accept DONE as much as IDLE
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
FleetConfig.LeadRollover config = cfg(handover.toString());
|
||||
LeadRollover rollover = newRollover(herdr, config, fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||
var prompts = herdr.calls.stream().filter(c -> "agent.prompt".equals(c.method())).toList();
|
||||
assertEquals(2, prompts.size(),
|
||||
"a pane reporting DONE the whole time must complete the full roll — /clear then "
|
||||
+ "bootstrapText — exactly like IDLE; DONE is a real turn-boundary equivalent "
|
||||
+ "to IDLE (see AgentStatus#DONE), not merely 'injectable'");
|
||||
assertTrue(prompts.get(0).params().toString().contains("/clear"));
|
||||
assertTrue(prompts.get(1).params().toString().contains("read the handover file"));
|
||||
}
|
||||
|
||||
// ---- Correction 2: only the terminal that opened a request may confirm it -------------
|
||||
|
||||
@Test
|
||||
@DisplayName("[CORRECTION 2] a different lead terminal cannot confirm another lead's rollover")
|
||||
void aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(OTHER_LEAD, pending.token(), true);
|
||||
|
||||
assertFalse(decision.accepted());
|
||||
assertEquals(LeadRollover.RefusalReason.NOT_YOUR_ROLLOVER, decision.reason());
|
||||
assertEquals(0, promptCallCount(herdr), "a foreign terminal's confirm() must never touch "
|
||||
+ "the pane it named, let alone the pane that actually opened the request");
|
||||
|
||||
// the rightful owner can still confirm the same token afterwards — a foreign confirm()
|
||||
// must not consume or otherwise disturb the pending request.
|
||||
LeadRollover.RollDecision ownerDecision = rollover.confirm(LEAD, pending.token(), true);
|
||||
assertTrue(ownerDecision.accepted(), "the actual opener must still be able to confirm after "
|
||||
+ "a foreign terminal's confirm() was refused");
|
||||
}
|
||||
|
||||
// ---- extra coverage: cancel(), CLEAR_DID_NOT_SETTLE, and the full success path ---------
|
||||
|
||||
@Test
|
||||
@DisplayName("cancel() drops a pending request so a later confirm() reports UNKNOWN_TOKEN")
|
||||
void cancelDropsThePendingRequest() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
assertTrue(rollover.cancel(pending.token()));
|
||||
assertFalse(rollover.cancel(pending.token()), "a second cancel() of the same token finds nothing");
|
||||
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
assertFalse(decision.accepted());
|
||||
assertEquals(LeadRollover.RefusalReason.UNKNOWN_TOKEN, decision.reason());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("open() throws IllegalStateException when leadRollover: is not configured")
|
||||
void openThrowsWhenNotConfigured() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
LeadRollover rollover = new LeadRollover(agents, () -> null, _ -> null, () -> 1_000L, () -> { }, Runnable::run);
|
||||
|
||||
assertThrows(IllegalStateException.class, () -> rollover.open(LEAD, "context is full"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("open() throws IllegalArgumentException when leadTerminal is null or blank")
|
||||
void openThrowsWhenLeadTerminalIsBlank() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path handover = writeHandover("handover contents");
|
||||
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), () -> 1_000L);
|
||||
|
||||
assertThrows(IllegalArgumentException.class, () -> rollover.open(null, "context is full"));
|
||||
assertThrows(IllegalArgumentException.class, () -> rollover.open(" ", "context is full"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a pane that never re-settles after /clear refuses to send bootstrapText")
|
||||
void clearThatNeverSettlesAfterwardsNeverSendsBootstrapText() throws IOException {
|
||||
// Idle until /clear is sent, then permanently working — isolates the SECOND wait
|
||||
// (clearSettleSeconds) from the first (turnSettleSeconds), which passes immediately here.
|
||||
FakeHerdr fake = new FakeHerdr();
|
||||
HerdrClient flipsAfterClear = new HerdrClient() {
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) throws HerdrException {
|
||||
JsonNode result = fake.call(method, params);
|
||||
if ("agent.prompt".equals(method) && String.valueOf(params).contains("/clear")) {
|
||||
fake.agentStatus("working");
|
||||
}
|
||||
return result;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
fake.close();
|
||||
}
|
||||
};
|
||||
Path handover = writeHandover("handover contents");
|
||||
FleetConfig.LeadRollover config =
|
||||
new FleetConfig.LeadRollover(handover.toString(), true, 3600, 20, 1 /*clearSettleSeconds*/, "boot text");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(flipsAfterClear, config, () -> clock.addAndGet(500));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "every synchronous gate passes; the refusal is logged only, "
|
||||
+ "deep inside the deferred continuation");
|
||||
assertEquals(1, promptCallCount(fake), "exactly one agent.prompt call — the /clear — and "
|
||||
+ "nothing else");
|
||||
long bootstrapSends = fake.calls.stream()
|
||||
.filter(c -> "agent.prompt".equals(c.method()))
|
||||
.filter(c -> String.valueOf(c.params()).contains("boot text"))
|
||||
.count();
|
||||
assertEquals(0, bootstrapSends, "bootstrapText must never be sent when /clear did not settle");
|
||||
}
|
||||
|
||||
// ---- fleetd #489: the second wait nudges the /clear pickup instead of being a no-op --------
|
||||
|
||||
/** Every {@code agent.send_keys} call {@code herdr} recorded — the submit-keystroke nudge. */
|
||||
private static long sendKeysCallCount(FakeHerdr herdr) {
|
||||
return herdr.calls.stream().filter(c -> "agent.send_keys".equals(c.method())).count();
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #489] a pane that stays IDLE the whole time (the paste-race case, "
|
||||
+ "measured live 2026-09-12) is nudged between /clear and bootstrapText, never lets "
|
||||
+ "them concatenate into one line")
|
||||
void clearPickupIsNudgedBeforeBootstrapTextWhenPaneStaysIdle() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr(); // default agentStatus is "idle" throughout — no WORKING sample ever
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||
|
||||
int clearIdx = -1;
|
||||
int bootstrapIdx = -1;
|
||||
int firstNudgeIdx = -1;
|
||||
for (int i = 0; i < herdr.calls.size(); i++) {
|
||||
FakeHerdr.Call c = herdr.calls.get(i);
|
||||
if ("agent.prompt".equals(c.method()) && String.valueOf(c.params()).contains("/clear") && clearIdx < 0) {
|
||||
clearIdx = i;
|
||||
} else if ("agent.prompt".equals(c.method()) && String.valueOf(c.params()).contains("read the handover file")) {
|
||||
bootstrapIdx = i;
|
||||
} else if ("agent.send_keys".equals(c.method()) && firstNudgeIdx < 0) {
|
||||
firstNudgeIdx = i;
|
||||
}
|
||||
}
|
||||
assertTrue(clearIdx >= 0, "/clear must have been sent");
|
||||
assertTrue(bootstrapIdx >= 0, "bootstrapText must have been sent");
|
||||
assertTrue(firstNudgeIdx >= 0, "at least one agent.send_keys nudge must go out — the pane "
|
||||
+ "never reported WORKING, so the /clear Enter may have raced the paste, and only a "
|
||||
+ "re-sent Enter proves the clear rather than concatenating bootstrapText onto "
|
||||
+ "whatever sits unsubmitted in the input box");
|
||||
assertTrue(firstNudgeIdx > clearIdx, "the nudge must happen AFTER /clear was sent, got call "
|
||||
+ "order: " + herdr.calls);
|
||||
assertTrue(firstNudgeIdx < bootstrapIdx, "the nudge must happen BEFORE bootstrapText is "
|
||||
+ "sent — never concatenated onto the same input line, got call order: " + herdr.calls);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #489] once a WORKING sample confirms /clear was picked up, nudging stops "
|
||||
+ "and bootstrapText is still sent after the pane returns to IDLE")
|
||||
void pickupSeenStopsNudgingAndBootstrapTextIsSent() throws IOException {
|
||||
FakeHerdr fake = new FakeHerdr();
|
||||
// Scripts the SECOND wait only: idle (default) until /clear is sent, then the first status
|
||||
// poll after /clear reports WORKING (a confirmed pickup), and every poll after that reports
|
||||
// IDLE (the completion boundary). The first wait (turnSettleSeconds) never sees this
|
||||
// sequence — it passes on its own first poll, before /clear is ever sent, on the default
|
||||
// "idle" status.
|
||||
AtomicLong postClearGetCalls = new AtomicLong(0);
|
||||
HerdrClient scriptsPickupThenIdle = new HerdrClient() {
|
||||
private volatile boolean clearSent = false;
|
||||
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) throws HerdrException {
|
||||
if ("agent.prompt".equals(method) && String.valueOf(params).contains("/clear")) {
|
||||
clearSent = true;
|
||||
}
|
||||
if (clearSent && "agent.get".equals(method)) {
|
||||
long n = postClearGetCalls.incrementAndGet();
|
||||
fake.agentStatus(n == 1 ? "working" : "idle");
|
||||
}
|
||||
return fake.call(method, params);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
fake.close();
|
||||
}
|
||||
};
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(scriptsPickupThenIdle, cfg(handover.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||
assertEquals(2, promptCallCount(fake), "a confirmed WORKING pickup followed by IDLE must "
|
||||
+ "still complete the full roll — /clear then bootstrapText");
|
||||
assertEquals(0, sendKeysCallCount(fake), "once WORKING was observed, nudging must stop "
|
||||
+ "immediately — no agent.send_keys call should ever have been needed or sent");
|
||||
assertTrue(postClearGetCalls.get() >= 2, "the pane's status must have been polled AGAIN "
|
||||
+ "after the WORKING sample, before bootstrapText was sent — this is what proves "
|
||||
+ "the method actually waited for the WORKING -> IDLE completion boundary instead "
|
||||
+ "of returning as soon as pickup was seen (or worse, without polling at all, as a "
|
||||
+ "stub that just returns true would); got " + postClearGetCalls.get()
|
||||
+ " agent.get call(s) after /clear");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #489] a pane that never reaches a turn boundary after /clear (stuck at "
|
||||
+ "UNKNOWN, never WORKING either) lets clearSettleSeconds expire — bootstrapText is "
|
||||
+ "never sent")
|
||||
void clearPickupNeverSettlesWhenStatusNeverReachesABoundary() throws IOException {
|
||||
FakeHerdr fake = new FakeHerdr();
|
||||
HerdrClient stuckUnknownAfterClear = new HerdrClient() {
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) throws HerdrException {
|
||||
if ("agent.prompt".equals(method) && String.valueOf(params).contains("/clear")) {
|
||||
// "wedged" maps to AgentStatus.UNKNOWN (see AgentStatus#fromWire) — neither a
|
||||
// pickup signal (WORKING) nor a boundary (IDLE/DONE), and distinct from the
|
||||
// already-covered BLOCKED case below.
|
||||
fake.agentStatus("wedged");
|
||||
}
|
||||
return fake.call(method, params);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
fake.close();
|
||||
}
|
||||
};
|
||||
Path handover = writeHandover("handover contents");
|
||||
FleetConfig.LeadRollover config =
|
||||
new FleetConfig.LeadRollover(handover.toString(), true, 3600, 20, 1 /*clearSettleSeconds*/, "boot text");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(stuckUnknownAfterClear, config, () -> clock.addAndGet(500));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "every synchronous gate passes; the refusal is logged only, "
|
||||
+ "deep inside the deferred continuation");
|
||||
assertEquals(1, promptCallCount(fake), "exactly one agent.prompt call — the /clear — and "
|
||||
+ "nothing else");
|
||||
long bootstrapSends = fake.calls.stream()
|
||||
.filter(c -> "agent.prompt".equals(c.method()))
|
||||
.filter(c -> String.valueOf(c.params()).contains("boot text"))
|
||||
.count();
|
||||
assertEquals(0, bootstrapSends, "bootstrapText must never be sent when the pane never "
|
||||
+ "reaches a turn boundary after /clear, whether WORKING was ever observed or not");
|
||||
assertEquals(0, sendKeysCallCount(fake), "an UNKNOWN status is neither a pickup signal nor "
|
||||
+ "a boundary — it must never be nudged");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #489] a submit() nudge that throws does not abort the roll — /clear and "
|
||||
+ "bootstrapText are both still sent")
|
||||
void submitThatThrowsDoesNotAbortTheRoll() throws IOException {
|
||||
FakeHerdr fake = new FakeHerdr(); // default idle throughout — nudging will be attempted
|
||||
HerdrClient throwsOnSubmit = new HerdrClient() {
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) throws HerdrException {
|
||||
if ("agent.send_keys".equals(method)) {
|
||||
throw new RuntimeException("simulated herdr transport failure on submit");
|
||||
}
|
||||
return fake.call(method, params);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
fake.close();
|
||||
}
|
||||
};
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(throwsOnSubmit, cfg(handover.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||
var prompts = fake.calls.stream().filter(c -> "agent.prompt".equals(c.method())).toList();
|
||||
assertEquals(2, prompts.size(), "a throwing submit() must be swallowed, not abort the roll "
|
||||
+ "— /clear and bootstrapText must both still be sent");
|
||||
assertTrue(prompts.get(0).params().toString().contains("/clear"));
|
||||
assertTrue(prompts.get(1).params().toString().contains("read the handover file"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a full successful roll sends /clear then bootstrapText, in order, and consumes the token")
|
||||
void successfulRollSendsClearThenBootstrapTextAndConsumesTheToken() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr(); // default agentStatus is "idle" — injectable immediately
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
FleetConfig.LeadRollover config = cfg(handover.toString());
|
||||
LeadRollover rollover = newRollover(herdr, config, fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||
var prompts = herdr.calls.stream().filter(c -> "agent.prompt".equals(c.method())).toList();
|
||||
assertEquals(2, prompts.size(), "expected exactly two agent.prompt calls: /clear then bootstrapText");
|
||||
assertTrue(prompts.get(0).params().toString().contains("/clear"));
|
||||
assertTrue(prompts.get(1).params().toString().contains("read the handover file"));
|
||||
|
||||
// token is consumed once confirm() approves — a second confirm() with the same token is
|
||||
// UNKNOWN_TOKEN, even though the deferred roll's own outcome is decided later.
|
||||
LeadRollover.RollDecision again = rollover.confirm(LEAD, pending.token(), true);
|
||||
assertEquals(LeadRollover.RefusalReason.UNKNOWN_TOKEN, again.reason());
|
||||
}
|
||||
|
||||
// ---- fleetd #480 follow-up: a relative handoverPath resolves against the CALLING lead's own
|
||||
// workspace, never the daemon's cwd ------------------------------------------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #480 follow-up] a relative handoverPath resolves against the lead's "
|
||||
+ "configured workspace, and PendingRollover carries the absolute path")
|
||||
void relativeHandoverPathResolvesAgainstLeadWorkspace() {
|
||||
Path workspace = tmp.resolve("lead-workspace");
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
FleetConfig.LeadRollover config = cfg("handover.md"); // relative — no directory component
|
||||
Function<String, String> leadWorkspace = terminal -> LEAD.equals(terminal) ? workspace.toString() : null;
|
||||
LeadRollover rollover = newRollover(herdr, config, fixedClock(clock), leadWorkspace);
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
|
||||
String expected = workspace.resolve("handover.md").normalize().toString();
|
||||
assertEquals(expected, pending.handoverPath(),
|
||||
"a relative handoverPath must resolve against the CALLING lead's own workspace, "
|
||||
+ "never the daemon's own working directory");
|
||||
assertTrue(Path.of(pending.handoverPath()).isAbsolute(),
|
||||
"the resolved path stored on PendingRollover must always be absolute");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #480 follow-up] confirm() accepts a handover file written at the "
|
||||
+ "resolved absolute location for a relative handoverPath")
|
||||
void confirmAcceptsAFileWrittenAtTheResolvedLocation() throws IOException {
|
||||
Path workspace = tmp.resolve("lead-workspace");
|
||||
Files.createDirectories(workspace);
|
||||
Path expected = workspace.resolve("handover.md");
|
||||
Files.writeString(expected, "handover contents");
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
FleetConfig.LeadRollover config = cfg("handover.md");
|
||||
Function<String, String> leadWorkspace = terminal -> LEAD.equals(terminal) ? workspace.toString() : null;
|
||||
LeadRollover rollover = newRollover(herdr, config, fixedClock(clock), leadWorkspace);
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
assertEquals(expected.toString(), pending.handoverPath());
|
||||
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #480 follow-up] a file at the same relative name under a DIFFERENT "
|
||||
+ "directory is NOT accepted — resolution is against the lead's own workspace only")
|
||||
void relativeHandoverPathDoesNotMatchAFileUnderADifferentDirectory() throws IOException {
|
||||
Path workspace = tmp.resolve("lead-workspace");
|
||||
Path decoy = tmp.resolve("decoy-directory");
|
||||
Files.createDirectories(workspace);
|
||||
Files.createDirectories(decoy);
|
||||
// the decoy directory holds a file with the SAME relative name — if resolution ever fell
|
||||
// back to searching, or resolved against the wrong base, this would be wrongly found.
|
||||
Files.writeString(decoy.resolve("handover.md"), "decoy contents — must never be read");
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
FleetConfig.LeadRollover config = cfg("handover.md");
|
||||
Function<String, String> leadWorkspace = terminal -> LEAD.equals(terminal) ? workspace.toString() : null;
|
||||
LeadRollover rollover = newRollover(herdr, config, fixedClock(clock), leadWorkspace);
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
assertEquals(workspace.resolve("handover.md").toString(), pending.handoverPath());
|
||||
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
assertFalse(decision.accepted());
|
||||
assertEquals(LeadRollover.RefusalReason.HANDOVER_MISSING, decision.reason(),
|
||||
"the file at workspace/handover.md does not exist — a decoy file at the same "
|
||||
+ "relative name under a different directory must never be mistaken for it");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #480 follow-up] an absolute handoverPath is used unchanged — the lead "
|
||||
+ "workspace lookup is never even consulted")
|
||||
void absoluteHandoverPathIsUnchangedByResolution() throws IOException {
|
||||
Path absolute = writeHandover("handover contents"); // tmp.resolve(...) — always absolute
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
FleetConfig.LeadRollover config = cfg(absolute.toString());
|
||||
// a workspace lookup that would resolve a RELATIVE path somewhere completely different —
|
||||
// proving it is never consulted at all for an already-absolute handoverPath.
|
||||
Function<String, String> leadWorkspace = terminal -> tmp.resolve("some-other-workspace").toString();
|
||||
LeadRollover rollover = newRollover(herdr, config, fixedClock(clock), leadWorkspace);
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
|
||||
assertEquals(absolute.normalize().toString(), pending.handoverPath(),
|
||||
"an absolute handoverPath must be used unchanged (aside from normalization)");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #480 follow-up correction] a RELATIVE fleet.leaders.<name>.cwd still "
|
||||
+ "yields an ABSOLUTE PendingRollover.handoverPath")
|
||||
void relativeLeadWorkspaceCwdStillYieldsAnAbsoluteHandoverPath() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
FleetConfig.LeadRollover config = cfg("handover.md"); // relative handoverPath
|
||||
// Nothing in FleetConfig validates fleet.leaders.<name>.cwd, so an operator can write a
|
||||
// RELATIVE one — this must still resolve to an absolute PendingRollover.handoverPath,
|
||||
// never silently reopen the exact bug this class fixes.
|
||||
String relativeWorkspace = "relative-lead-workspace";
|
||||
Function<String, String> leadWorkspace = terminal -> LEAD.equals(terminal) ? relativeWorkspace : null;
|
||||
LeadRollover rollover = newRollover(herdr, config, fixedClock(clock), leadWorkspace);
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
|
||||
assertTrue(Path.of(pending.handoverPath()).isAbsolute(),
|
||||
"the resolved path must be absolute even when the configured cwd itself is relative");
|
||||
|
||||
// Asserting isAbsolute() alone would also pass for a path resolved against the WRONG base
|
||||
// (e.g. some unrelated absolute directory) — pin the actual value too.
|
||||
String expected = Path.of(System.getProperty("user.dir")).resolve(relativeWorkspace)
|
||||
.resolve("handover.md").toAbsolutePath().normalize().toString();
|
||||
assertEquals(expected, pending.handoverPath(),
|
||||
"a relative cwd must be finished off against the daemon's own user.dir, exactly "
|
||||
+ "like a missing cwd — never left relative, which would silently reopen the "
|
||||
+ "exact bug this class fixes: every later reader interpreting an ambiguous "
|
||||
+ "path against ITS OWN working directory");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #480 follow-up] a terminal with no configured workspace (null lookup "
|
||||
+ "result) falls back to the daemon's own user.dir")
|
||||
void noConfiguredWorkspaceFallsBackToUserDir() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
FleetConfig.LeadRollover config = cfg("some-handover.md");
|
||||
Function<String, String> leadWorkspace = _ -> null; // no configured workspace for anyone
|
||||
LeadRollover rollover = newRollover(herdr, config, fixedClock(clock), leadWorkspace);
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
|
||||
String expected = Path.of(System.getProperty("user.dir")).resolve("some-handover.md")
|
||||
.normalize().toString();
|
||||
assertEquals(expected, pending.handoverPath(),
|
||||
"with no configured workspace for this lead, resolution must fall back to the "
|
||||
+ "daemon's own user.dir — the same fallback LeadLauncher#launch already uses "
|
||||
+ "for a lead with no configured cwd");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #480 follow-up] a terminal whose workspace lookup returns a blank string "
|
||||
+ "also falls back to the daemon's own user.dir")
|
||||
void blankConfiguredWorkspaceFallsBackToUserDir() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
FleetConfig.LeadRollover config = cfg("some-handover.md");
|
||||
Function<String, String> leadWorkspace = _ -> " ";
|
||||
LeadRollover rollover = newRollover(herdr, config, fixedClock(clock), leadWorkspace);
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
|
||||
String expected = Path.of(System.getProperty("user.dir")).resolve("some-handover.md")
|
||||
.normalize().toString();
|
||||
assertEquals(expected, pending.handoverPath(),
|
||||
"a blank (non-null) workspace lookup result must be treated the same as null — "
|
||||
+ "fall back to user.dir, never resolve against an empty base");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #480 follow-up] the default bootstrapText that is actually SENT names "
|
||||
+ "the RESOLVED absolute handoverPath, not the raw relative configured value")
|
||||
void defaultBootstrapTextNamesTheResolvedAbsolutePath() throws IOException {
|
||||
Path workspace = tmp.resolve("lead-workspace");
|
||||
Files.createDirectories(workspace);
|
||||
Path expected = workspace.resolve("handover.md");
|
||||
Files.writeString(expected, "handover contents");
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr(); // default agentStatus "idle" — both waits settle immediately
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
// bootstrapText left null so the DEFAULT sentence is built — from the RESOLVED path
|
||||
FleetConfig.LeadRollover config = new FleetConfig.LeadRollover("handover.md", true, 3600, 20, 20, null);
|
||||
Function<String, String> leadWorkspace = terminal -> LEAD.equals(terminal) ? workspace.toString() : null;
|
||||
LeadRollover rollover = newRollover(herdr, config, fixedClock(clock), leadWorkspace);
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||
var prompts = herdr.calls.stream().filter(c -> "agent.prompt".equals(c.method())).toList();
|
||||
assertEquals(2, prompts.size(), "expected exactly two agent.prompt calls: /clear then the "
|
||||
+ "default bootstrapText");
|
||||
String bootstrapSent = prompts.get(1).params().toString();
|
||||
assertTrue(bootstrapSent.contains(expected.toString()),
|
||||
"the default bootstrapText actually sent to the pane must name the RESOLVED "
|
||||
+ "absolute path (" + expected + ") — a fresh session's own pane cannot "
|
||||
+ "resolve a relative path against a directory it never had. Actual text "
|
||||
+ "sent: " + bootstrapSent);
|
||||
assertFalse(bootstrapSent.contains("\"handover.md\""),
|
||||
"must not name the raw relative configured value in the text actually sent");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,124 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #464 shipped {@code configuredChartersNameOnlyRegisteredTools} below, comparing charter
|
||||
* text against the registered tool surface — but it scraped <em>both</em> sides from source text:
|
||||
* its own fixture charter, and a regex over {@code FleetMcp.java}'s {@code tool("…")} calls. #469's
|
||||
* gap: nothing anyone wrote into the live {@code fleetd.yaml} could ever reach that test, because
|
||||
* it never called production validation code.
|
||||
*
|
||||
* <p>This version keeps the charter-text extraction helper ({@code toolsNamedIn}) — charters are
|
||||
* free-text config, so finding a {@code fleet_*}/{@code bridge_*} token inside one has no source
|
||||
* but a scrape — but reads the <em>registered</em> side from {@link FleetTool}, the canonical enum
|
||||
* {@code FleetMcp} itself now derives its tool schemas and authorization switch from, rather than a
|
||||
* second scrape of {@code FleetMcp.java}'s source. It also exercises {@link
|
||||
* CharterToolSurface#assertChartersNameOnlyRegisteredTools} directly — the method {@code
|
||||
* Fleetd.main} actually calls at startup — both accepting and rejecting. {@code
|
||||
* dev.ltms.fleet.FleetdStartupValidationTest#mainRefusesACharterNamingAnUnregisteredTool} is what
|
||||
* closes #469's actual gap: it proves that call is wired into {@code Fleetd.main} itself, against a
|
||||
* live-shaped config fixture, not just a unit call to the method in isolation.
|
||||
*/
|
||||
class CharterToolSurfaceTest {
|
||||
|
||||
private static Set<String> matches(String text, String regex) {
|
||||
Matcher m = Pattern.compile(regex).matcher(text);
|
||||
Set<String> found = new LinkedHashSet<>();
|
||||
while (m.find()) {
|
||||
found.add(m.group(1));
|
||||
}
|
||||
return found;
|
||||
}
|
||||
|
||||
/** Every {@code fleet_*} or legacy {@code bridge_*} token in the configured launch charters. */
|
||||
private static Set<String> toolsNamedIn(FleetConfig config) {
|
||||
return matches(String.join("\n", config.fleet().charters().values()),
|
||||
"(fleet_[a-z_]+|bridge_[a-z_]+)");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("every tool named in a configured charter is in the canonical FleetTool set")
|
||||
void configuredChartersNameOnlyRegisteredTools(@TempDir Path dir) throws Exception {
|
||||
Path configFile = dir.resolve("charters.yaml");
|
||||
Files.writeString(configFile, """
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through fleet_reply.
|
||||
reviewer: |
|
||||
Use fleet_ask only for the lead's decision.
|
||||
""");
|
||||
|
||||
FleetConfig config = FleetConfig.load(configFile);
|
||||
Set<String> named = toolsNamedIn(config);
|
||||
Set<String> registered = FleetTool.wireNames();
|
||||
|
||||
assertTrue(!named.isEmpty(),
|
||||
"the charter fixture named no fleet_* or bridge_* tool. This test would check nothing; "
|
||||
+ "add charter text that names a tool before changing the extraction.");
|
||||
assertTrue(!registered.isEmpty(),
|
||||
"FleetTool.wireNames() is empty. This test would check nothing; repair FleetTool "
|
||||
+ "before changing the assertion.");
|
||||
|
||||
Set<String> unknown = new LinkedHashSet<>(named);
|
||||
unknown.removeAll(registered);
|
||||
assertTrue(unknown.isEmpty(),
|
||||
"configured charter text names " + unknown + ", but FleetTool does not list it. "
|
||||
+ "Checked " + named + " against " + registered + ". Fix the charter text or "
|
||||
+ "add the tool to FleetTool; do NOT weaken this test.");
|
||||
}
|
||||
|
||||
/**
|
||||
* The positive case for the actual production entry point: a charter naming only tools
|
||||
* {@link FleetTool} lists must not throw.
|
||||
*/
|
||||
@Test
|
||||
@DisplayName("CharterToolSurface accepts a charter that names only registered tools")
|
||||
void charterToolSurfaceAcceptsKnownTools() {
|
||||
assertDoesNotThrow(() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(Map.of(
|
||||
"dev", "Send the final handoff through fleet_reply, using fleet_send to delegate.",
|
||||
"reviewer", "Use fleet_ask only for the lead's decision.")));
|
||||
}
|
||||
|
||||
/**
|
||||
* The negative case for the actual production entry point (fleetd #469's motivating example:
|
||||
* {@code bridge_send} is the pre-CB-634 name, removed from the tool surface). The message must
|
||||
* name both the offending charter key and the unknown tool, so an operator reading the startup
|
||||
* log knows exactly which charter to fix.
|
||||
*/
|
||||
@Test
|
||||
@DisplayName("CharterToolSurface rejects a charter naming a tool the server does not register")
|
||||
void charterToolSurfaceRejectsAnUnregisteredTool() {
|
||||
Map<String, String> charters = new LinkedHashMap<>();
|
||||
charters.put("dev", "Send the final handoff through bridge_send.");
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class,
|
||||
() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(charters));
|
||||
assertTrue(e.getMessage().contains("dev"),
|
||||
"expected the charter key 'dev' in the failure message, got: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("bridge_send"),
|
||||
"expected the unknown tool 'bridge_send' in the failure message, got: " + e.getMessage());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("CharterToolSurface is a no-op on an absent or empty charter map")
|
||||
void charterToolSurfaceIsANoOpWithNoCharters() {
|
||||
assertDoesNotThrow(() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(null));
|
||||
assertDoesNotThrow(() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(Map.of()));
|
||||
}
|
||||
}
|
||||
@@ -26,7 +26,7 @@ import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.regex.Matcher;
|
||||
@@ -50,7 +50,6 @@ import static org.junit.jupiter.api.Assertions.*;
|
||||
class FleetMcpAuthzTest {
|
||||
|
||||
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
|
||||
private static final Pattern TOOL_REGISTRATION = Pattern.compile("tool\\(\\\"(fleet_[a-z_]+)\\\"");
|
||||
|
||||
private final FakeHerdr herdr = new FakeHerdr();
|
||||
private final AgentControl agents = new AgentControl(herdr);
|
||||
@@ -76,12 +75,16 @@ class FleetMcpAuthzTest {
|
||||
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
|
||||
metrics = FleetMetrics.create(sessions, new InMemoryReplyInbox());
|
||||
|
||||
// fleetd #480 correction round: FleetMcp has one constructor now (no defaulting
|
||||
// overloads — see its javadoc), so every feature this test does not exercise is passed
|
||||
// its explicit "off" value here rather than being omitted.
|
||||
mcp = new FleetMcp(messages, workers, sessions, identity, sessions.asPresence(),
|
||||
new PrimaryRegistry(null),
|
||||
enforce ? CallerResolver.withLeadsAndMembers(identity, false, null,
|
||||
Map::of, new MemberRegistry(null)) : null,
|
||||
metrics, FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none());
|
||||
FleetMcp.QuarantineSource.none(), null, FleetMcp.OutageSource.none(),
|
||||
FleetMcp.LeadSeatSource.none(), List.of(), null);
|
||||
return mcp;
|
||||
}
|
||||
|
||||
@@ -187,6 +190,71 @@ class FleetMcpAuthzTest {
|
||||
"no CallerResolver supplied ⇒ authorization not enforced (legacy behaviour)");
|
||||
}
|
||||
|
||||
// --- fleetd #439: who may see fleet_list's coordinator row ----------------------------------
|
||||
|
||||
/**
|
||||
* fleetd #439: {@link FleetMcp#coordinatorVisibleTo} is the whole policy decision for
|
||||
* {@code fleet_list}'s {@code coordinator} row — lead-to-lead coordination state, not roster
|
||||
* observation. Only the primary may see it; a worker, an architect, and (the case the previous
|
||||
* pass of this ticket did not cover) an anonymous caller must all be refused.
|
||||
*/
|
||||
@Test
|
||||
void onlyThePrimaryMaySeeTheCoordinatorRow() {
|
||||
assertTrue(FleetMcp.coordinatorVisibleTo(PRIMARY), "the primary must see its own coordination state");
|
||||
assertFalse(FleetMcp.coordinatorVisibleTo(WORKER_A), "a worker must not see lead-to-lead coordination state");
|
||||
assertFalse(FleetMcp.coordinatorVisibleTo(ARCH_DESIGN),
|
||||
"an architect holds READ today, but that must not extend to coordinator");
|
||||
assertFalse(FleetMcp.coordinatorVisibleTo(ANON), "authenticated as nothing must not see it either");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #439 / PR #462 review finding M2: the predicate above can be perfectly correct while
|
||||
* the one production call site (the {@code fleet_list} MCP handler) never actually asks it —
|
||||
* a literal {@code true} compiles, and the whole suite stayed green under that mutation because
|
||||
* every existing test drives {@link FleetMcp#listFleet} directly and supplies the boolean
|
||||
* itself. This test reads {@code FleetMcp.java}'s own source (same idiom as {@link
|
||||
* #toolsTheServerRegisters()} / {@link #everyRegisteredToolHasItsHandlerActionPinned()}) and
|
||||
* asserts the handler's call passes {@code coordinatorVisibleTo(principal(exchange))} — not a
|
||||
* literal {@code true} or {@code false} — as {@code listFleet}'s trailing argument.
|
||||
*
|
||||
* <p>Anchored on argument position, not a bare substring search: {@code true} appears many
|
||||
* times elsewhere in this file for unrelated reasons, so a plain {@code contains("true")}
|
||||
* check would prove nothing. The pattern requires the literal text immediately before the
|
||||
* closing {@code );} of the {@code listFleet(} call inside the handler block to be exactly
|
||||
* {@code coordinatorVisibleTo(principal(exchange))}.
|
||||
*/
|
||||
@Test
|
||||
void theFleetListHandlerActuallyConsultsCoordinatorVisibleTo() throws Exception {
|
||||
String source = Files.readString(MCP_SOURCE);
|
||||
|
||||
// Isolate the fleet_list handler block: from its declaration up to the next handler's
|
||||
// declaration. A change to variable naming would break this scrape loudly (see the control
|
||||
// assertion just below), rather than silently reporting "no violation found".
|
||||
int start = source.indexOf("listHandler =");
|
||||
assertTrue(start >= 0, "could not find the fleet_list handler (listHandler) in " + MCP_SOURCE
|
||||
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
|
||||
int end = source.indexOf("stopHandler =", start);
|
||||
assertTrue(end > start, "could not find the handler declared after listHandler to bound the scrape");
|
||||
String handlerBlock = source.substring(start, end);
|
||||
|
||||
// CONTROL: the block we scraped really does contain a call to listFleet(...) -- if this
|
||||
// fails, the anchors above moved and the assertion below would otherwise pass on nothing.
|
||||
assertTrue(handlerBlock.contains("listFleet("),
|
||||
"control failed: the scraped listHandler block contains no listFleet( call at all -- "
|
||||
+ "the anchors have drifted, this test is not testing what it claims to");
|
||||
|
||||
Pattern trailingArg = Pattern.compile(
|
||||
"listFleet\\([^;]*?,\\s*(coordinatorVisibleTo\\(principal\\(exchange\\)\\)|true|false)\\s*\\)\\s*;",
|
||||
Pattern.DOTALL);
|
||||
Matcher m = trailingArg.matcher(handlerBlock);
|
||||
assertTrue(m.find(), "could not locate listFleet(...)'s trailing boolean argument in the "
|
||||
+ "listHandler block -- the call shape changed, update this test's anchor: " + handlerBlock);
|
||||
String trailing = m.group(1);
|
||||
assertEquals("coordinatorVisibleTo(principal(exchange))", trailing,
|
||||
"the fleet_list handler must ask coordinatorVisibleTo(principal(exchange)) who is "
|
||||
+ "calling, not pass a literal boolean -- found: " + trailing);
|
||||
}
|
||||
|
||||
// --- which action each tool hands the gate (fleetd #272) ------------------------------------
|
||||
|
||||
/**
|
||||
@@ -201,20 +269,48 @@ class FleetMcpAuthzTest {
|
||||
*/
|
||||
@Test
|
||||
void pollingByTargetIsADrainAndPollingByTicketIsARead() {
|
||||
assertEquals(Authz.Action.DRAIN, FleetMcp.pollAction("term_b"),
|
||||
assertEquals(Authz.Action.DRAIN, FleetMcp.pollAction("term_b", null),
|
||||
"poll by target removes the replies — that is a drain, not an observation");
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null),
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null, null),
|
||||
"poll by ticket changes nothing");
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(" "),
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(" ", null),
|
||||
"a blank target is an absent target");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #421: a coordId branch is a THIRD operation behind fleet_poll's one name, and it must
|
||||
* map to {@link Authz.Action#COORD_READ} — never {@link Authz.Action#READ}, even though this
|
||||
* branch also consumes nothing. READ's grant is open to every authenticated role on the premise
|
||||
* that the roster carries no secrets; a lead-to-lead body is not the roster, so folding this
|
||||
* branch into READ would let any worker read every peer lead's mail in full. coordId also takes
|
||||
* priority over target when both happen to be present — it addresses a different inbox entirely.
|
||||
*/
|
||||
@Test
|
||||
void pollingByCoordIdIsACoordReadNeverAPlainRead() {
|
||||
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction(null, "mac-opus"),
|
||||
"reading held peer mail must not be mapped to the everyone-readable READ action");
|
||||
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction(" ", "mac-opus"),
|
||||
"a blank target must not fall through to READ/DRAIN when coordId is present");
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null, " "),
|
||||
"a blank coordId is an absent coordId, same as target/ticket");
|
||||
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction("term_b", "mac-opus"),
|
||||
"coordId takes priority over target — this is a different inbox, not a drain");
|
||||
}
|
||||
|
||||
@Test
|
||||
void everyRegisteredToolHasItsHandlerActionPinned() {
|
||||
Set<String> registered = toolsTheServerRegisters();
|
||||
// fleetd #469: this used to scrape FleetMcp.java's tool("…") calls for the registered set —
|
||||
// a third copy of the same list this file, CharterToolSurfaceTest and McpContractDocTest
|
||||
// each kept independently. All three now read FleetTool.wireNames(), the canonical set
|
||||
// FleetMcp itself derives its tool schemas AND its authorization switch from; adding a tool
|
||||
// there without pinning its action in FleetMcp#authzAction is a compile error, so this test's
|
||||
// per-tool assertions below are a run-time regression pin on top of that compile-time check,
|
||||
// not the only thing standing between a new tool and an unpinned action.
|
||||
Set<String> registered = FleetTool.wireNames();
|
||||
assertTrue(registered.size() >= 10,
|
||||
"scraped only " + registered.size() + " tool registrations from FleetMcp (" + registered
|
||||
+ "); the server registers eleven, so the tool(\"…\") scrape has stopped matching");
|
||||
"FleetTool.wireNames() returned only " + registered.size() + " tool(s) (" + registered
|
||||
+ "); the server registers eleven, so FleetTool has stopped listing the real "
|
||||
+ "tool surface");
|
||||
registered.forEach(tool -> assertDoesNotThrow(() -> FleetMcp.toolAction(tool, Map.of()),
|
||||
() -> tool + " is registered but has no pinned authorization action"));
|
||||
|
||||
@@ -230,30 +326,19 @@ class FleetMcpAuthzTest {
|
||||
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_whoami", Map.of()));
|
||||
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_poll", Map.of("ticket", "task")));
|
||||
assertEquals(Authz.Action.DRAIN, FleetMcp.toolAction("fleet_poll", Map.of("target", "term_b")));
|
||||
}
|
||||
|
||||
private static Set<String> toolsTheServerRegisters() {
|
||||
try {
|
||||
Matcher matcher = TOOL_REGISTRATION.matcher(Files.readString(MCP_SOURCE));
|
||||
Set<String> tools = new LinkedHashSet<>();
|
||||
while (matcher.find()) {
|
||||
tools.add(matcher.group(1));
|
||||
}
|
||||
return tools;
|
||||
} catch (Exception e) {
|
||||
throw new AssertionError("could not scrape FleetMcp tool registrations", e);
|
||||
}
|
||||
assertEquals(Authz.Action.COORD_READ,
|
||||
FleetMcp.toolAction("fleet_poll", Map.of("coordId", "mac-opus")));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aWorkerMayNotDrainAnotherSessionsInboxByPolling() {
|
||||
FleetMcp m = mcp(true);
|
||||
|
||||
assertNotNull(m.denyFor(WORKER_A, FleetMcp.pollAction("term_b"), "term_b"),
|
||||
assertNotNull(m.denyFor(WORKER_A, FleetMcp.pollAction("term_b", null), "term_b"),
|
||||
"a worker draining a peer's inbox would destroy replies queued for the primary");
|
||||
assertNotNull(m.denyFor(ARCH_DESIGN, FleetMcp.pollAction("term_b"), "term_b"),
|
||||
assertNotNull(m.denyFor(ARCH_DESIGN, FleetMcp.pollAction("term_b", null), "term_b"),
|
||||
"an architect has no lifecycle rights either — same gate as fleet_ack");
|
||||
assertNull(m.denyFor(PRIMARY, FleetMcp.pollAction("term_b"), "term_b"),
|
||||
assertNull(m.denyFor(PRIMARY, FleetMcp.pollAction("term_b", null), "term_b"),
|
||||
"collecting a held reply is the primary's job");
|
||||
}
|
||||
|
||||
@@ -265,11 +350,33 @@ class FleetMcpAuthzTest {
|
||||
void pollingAnOwnTicketStaysOpenToWorkersAndArchitects() {
|
||||
FleetMcp m = mcp(true);
|
||||
|
||||
assertNull(m.denyFor(WORKER_A, FleetMcp.pollAction(null), null));
|
||||
assertNull(m.denyFor(ARCH_DESIGN, FleetMcp.pollAction(null), null),
|
||||
assertNull(m.denyFor(WORKER_A, FleetMcp.pollAction(null, null), null));
|
||||
assertNull(m.denyFor(ARCH_DESIGN, FleetMcp.pollAction(null, null), null),
|
||||
"an architect delegates with wait:false, so it must be able to poll its ticket");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #421 acceptance: the whole ticket, pinned at the policy-table layer. A worker must be
|
||||
* refused the coordId branch, the primary must be allowed, and (CB-548) an architect — which
|
||||
* holds READ today — must be refused too, because "not primary" means not-architect here.
|
||||
*/
|
||||
@Test
|
||||
void aWorkerAndAnArchitectMayNotReadHeldPeerMailOnlyThePrimaryMay() {
|
||||
FleetMcp m = mcp(true);
|
||||
Authz.Action coordRead = FleetMcp.pollAction(null, "mac-opus");
|
||||
|
||||
McpSchema.CallToolResult workerDenied = m.denyFor(WORKER_A, coordRead, null);
|
||||
assertNotNull(workerDenied, "a worker must not read held lead-to-lead mail");
|
||||
assertTrue(workerDenied.isError());
|
||||
|
||||
McpSchema.CallToolResult archDenied = m.denyFor(ARCH_DESIGN, coordRead, null);
|
||||
assertNotNull(archDenied, "an architect holds READ today, but not-primary must mean "
|
||||
+ "not-architect here too");
|
||||
assertTrue(archDenied.isError());
|
||||
|
||||
assertNull(m.denyFor(PRIMARY, coordRead, null), "reading its own held mail is the primary's job");
|
||||
}
|
||||
|
||||
// --- identity reconstruction from the transport context ------------------------------------
|
||||
|
||||
@Test
|
||||
|
||||
@@ -0,0 +1,261 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import dev.ltms.fleet.auth.Authz;
|
||||
import dev.ltms.fleet.auth.CallerResolver;
|
||||
import dev.ltms.fleet.auth.MemberRegistry;
|
||||
import dev.ltms.fleet.auth.Principal;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.PaneLocator;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.lead.LeadRollover;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.msg.InMemoryReplyInbox;
|
||||
import dev.ltms.fleet.msg.MessageService;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.session.FakeWorktrees;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* fleetd #480 Unit C — the {@code fleet_handover} MCP tool, the surface that finally calls
|
||||
* {@link LeadRollover#open}/{@link LeadRollover#confirm}/{@link LeadRollover#cancel}.
|
||||
*
|
||||
* <p>Uses {@link LeadRollover}'s PUBLIC constructor (real wall clock, real 250ms settle poll, a
|
||||
* real virtual-thread continuation runner) rather than its package-private test constructor —
|
||||
* this test lives in {@code dev.ltms.fleet.mcp}, not {@code dev.ltms.fleet.lead}, and does not
|
||||
* need to control the post-{@code confirm()} continuation's timing: it only asserts the
|
||||
* SYNCHRONOUS return value of {@code open}/{@code confirm}/{@code cancel}, which is exactly what
|
||||
* {@code FleetMcp.handover} forwards to the client. {@code turnSettleSeconds}/{@code
|
||||
* clearSettleSeconds} are kept at 1s so a confirmed request's background continuation (which this
|
||||
* class does not wait on or assert against) gives up quickly rather than polling for 20s on a
|
||||
* daemon virtual thread.
|
||||
*/
|
||||
class FleetMcpHandoverTest {
|
||||
|
||||
@TempDir
|
||||
Path tmp;
|
||||
|
||||
private static final String LEAD = "term_lead";
|
||||
private static final String OTHER_LEAD = "term_other_lead";
|
||||
|
||||
private final FakeHerdr herdr = new FakeHerdr();
|
||||
private final AgentControl agents = new AgentControl(herdr);
|
||||
private FleetMcp mcp;
|
||||
|
||||
@AfterEach
|
||||
void close() {
|
||||
if (mcp != null) mcp.close();
|
||||
}
|
||||
|
||||
private static FleetConfig.LeadRollover cfg(String handoverPath) {
|
||||
return new FleetConfig.LeadRollover(handoverPath, false, 3600, 1, 1, "read the handover file");
|
||||
}
|
||||
|
||||
private LeadRollover newRollover(String handoverPath) {
|
||||
// Every handoverPath this test class uses comes from tmp.resolve(...), which is already
|
||||
// absolute, so the workspace lookup is never actually consulted — a no-op lookup is enough.
|
||||
return new LeadRollover(agents, () -> cfg(handoverPath), _ -> null);
|
||||
}
|
||||
|
||||
/** A fully wired FleetMcp on fakes (mirrors FleetMcpAuthzTest's helper), plus a leadRollover. */
|
||||
private FleetMcp mcp(LeadRollover leadRollover) {
|
||||
FleetConfig.Profile pcfg = new FleetConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
|
||||
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(pcfg.profile(), pcfg), pcfg.profile(),
|
||||
_ -> "tok");
|
||||
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
|
||||
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous(),
|
||||
new InMemoryReplyInbox());
|
||||
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
|
||||
|
||||
mcp = new FleetMcp(messages, workers, sessions, identity, sessions.asPresence(),
|
||||
new PrimaryRegistry(null),
|
||||
CallerResolver.withLeadsAndMembers(identity, false, null, Map::of, new MemberRegistry(null)),
|
||||
null, FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), null, FleetMcp.OutageSource.none(),
|
||||
FleetMcp.LeadSeatSource.none(), List.of(), leadRollover);
|
||||
return mcp;
|
||||
}
|
||||
|
||||
private static String textOf(McpSchema.CallToolResult r) {
|
||||
return ((McpSchema.TextContent) r.content().getFirst()).text();
|
||||
}
|
||||
|
||||
private static String extractToken(String json) {
|
||||
int i = json.indexOf("\"token\":\"");
|
||||
assertTrue(i >= 0, "no token field in: " + json);
|
||||
int start = i + "\"token\":\"".length();
|
||||
int end = json.indexOf('"', start);
|
||||
return json.substring(start, end);
|
||||
}
|
||||
|
||||
// --- acceptance 1: registered whether or not leadRollover: is configured -------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("fleet_handover is registered whether or not leadRollover: is configured")
|
||||
void registeredEitherWay() {
|
||||
assertTrue(FleetTool.wireNames().contains("fleet_handover"),
|
||||
"FleetTool must list fleet_handover as part of the canonical tool surface");
|
||||
|
||||
FleetMcp withNull = mcp(null);
|
||||
assertTrue(withNull.registeredTools().stream().anyMatch(t -> "fleet_handover".equals(t.name())),
|
||||
"fleet_handover must be registered even with no LeadRollover constructed");
|
||||
withNull.close();
|
||||
|
||||
FleetMcp withConfigured = mcp(newRollover(tmp.resolve("h.md").toString()));
|
||||
assertTrue(withConfigured.registeredTools().stream().anyMatch(t -> "fleet_handover".equals(t.name())),
|
||||
"fleet_handover must be registered when a LeadRollover IS constructed too");
|
||||
}
|
||||
|
||||
// --- acceptance 2: null LeadRollover -> clean NOT_CONFIGURED, no throw, every action -------
|
||||
|
||||
@Test
|
||||
@DisplayName("with leadRollover: absent, every action returns a clean NOT_CONFIGURED refusal and never throws")
|
||||
void nullLeadRolloverRefusesCleanlyForEveryAction() {
|
||||
McpSchema.CallToolResult open = assertDoesNotThrow(
|
||||
() -> FleetMcp.handover(null, LEAD, Map.of("action", "open")));
|
||||
assertFalse(open.isError(), "a refusal is not a protocol error: " + textOf(open));
|
||||
assertTrue(textOf(open).contains("NOT_CONFIGURED"), textOf(open));
|
||||
|
||||
McpSchema.CallToolResult confirm = assertDoesNotThrow(() -> FleetMcp.handover(null, LEAD,
|
||||
Map.of("action", "confirm", "token", "whatever")));
|
||||
assertFalse(confirm.isError());
|
||||
assertTrue(textOf(confirm).contains("NOT_CONFIGURED"), textOf(confirm));
|
||||
|
||||
McpSchema.CallToolResult cancel = assertDoesNotThrow(() -> FleetMcp.handover(null, LEAD,
|
||||
Map.of("action", "cancel", "token", "whatever")));
|
||||
assertFalse(cancel.isError());
|
||||
assertTrue(textOf(cancel).contains("NOT_CONFIGURED"), textOf(cancel));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a blank/unknown action is a clean tool error, never an exception")
|
||||
void unknownActionIsACleanError() {
|
||||
McpSchema.CallToolResult missing = assertDoesNotThrow(() -> FleetMcp.handover(null, LEAD, Map.of()));
|
||||
assertTrue(missing.isError());
|
||||
|
||||
McpSchema.CallToolResult bogus = assertDoesNotThrow(
|
||||
() -> FleetMcp.handover(null, LEAD, Map.of("action", "bogus")));
|
||||
assertTrue(bogus.isError());
|
||||
}
|
||||
|
||||
// --- acceptance 3: non-primary refused by the authorization gate ---------------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("fleet_handover maps to Authz.Action.HANDOVER, primary-only")
|
||||
void onlyThePrimaryIsPermitted() {
|
||||
assertEquals(Authz.Action.HANDOVER, FleetMcp.toolAction("fleet_handover", Map.of()));
|
||||
|
||||
FleetMcp m = mcp(null);
|
||||
assertNull(m.denyFor(Principal.primary(1), Authz.Action.HANDOVER, null),
|
||||
"the primary may drive fleet_handover");
|
||||
|
||||
McpSchema.CallToolResult deniedWorker =
|
||||
m.denyFor(Principal.worker("term_w", 2), Authz.Action.HANDOVER, null);
|
||||
assertNotNull(deniedWorker, "a worker must be refused fleet_handover");
|
||||
assertTrue(deniedWorker.isError());
|
||||
|
||||
McpSchema.CallToolResult deniedArchitect =
|
||||
m.denyFor(Principal.architect("slot", "term_a", 3), Authz.Action.HANDOVER, null);
|
||||
assertNotNull(deniedArchitect, "an architect has no lifecycle rights either — same gate as SPAWN/STOP/DRAIN");
|
||||
assertTrue(deniedArchitect.isError());
|
||||
}
|
||||
|
||||
// --- acceptance 4: the registered schema names no caller-terminal parameter ----------------
|
||||
|
||||
@Test
|
||||
@DisplayName("the registered fleet_handover schema names no terminal/session/leadTerminal parameter")
|
||||
void schemaCarriesNoCallerIdentityParameter() {
|
||||
FleetMcp m = mcp(null);
|
||||
McpSchema.Tool tool = m.registeredTools().stream()
|
||||
.filter(t -> "fleet_handover".equals(t.name()))
|
||||
.findFirst()
|
||||
.orElseThrow(() -> new AssertionError("fleet_handover was not registered"));
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
Map<String, Object> properties = (Map<String, Object>) tool.inputSchema().get("properties");
|
||||
assertNotNull(properties, "tool has no 'properties' in its input schema");
|
||||
for (String forbidden : List.of("terminal", "sessionId", "leadTerminal", "callerTerminal", "target")) {
|
||||
assertFalse(properties.containsKey(forbidden),
|
||||
"fleet_handover's REGISTERED schema must not carry a caller-identity parameter, "
|
||||
+ "found '" + forbidden + "' in " + properties.keySet());
|
||||
}
|
||||
}
|
||||
|
||||
// --- acceptance 5: open then confirm on the same connection; ownership is enforced --------
|
||||
|
||||
@Test
|
||||
@DisplayName("open then confirm on the same terminal succeeds; a different terminal gets NOT_YOUR_ROLLOVER")
|
||||
void openThenConfirmRoundTripsAndOwnershipIsEnforced() throws Exception {
|
||||
Path handover = tmp.resolve("handover.md");
|
||||
Files.writeString(handover, "not written yet");
|
||||
LeadRollover rollover = newRollover(handover.toString());
|
||||
|
||||
McpSchema.CallToolResult openResult = FleetMcp.handover(rollover, LEAD, Map.of("action", "open"));
|
||||
assertFalse(openResult.isError(), textOf(openResult));
|
||||
String token = extractToken(textOf(openResult));
|
||||
|
||||
// Rewrite the handover file so its mtime is measurably after open()'s requestedAtMillis —
|
||||
// LeadRollover#confirm's freshness check (HANDOVER_STALE) requires this.
|
||||
Thread.sleep(50);
|
||||
Files.writeString(handover, "the real handover content");
|
||||
|
||||
McpSchema.CallToolResult wrongCaller = FleetMcp.handover(rollover, OTHER_LEAD,
|
||||
Map.of("action", "confirm", "token", token));
|
||||
assertFalse(wrongCaller.isError(), "a refusal is a legitimate outcome, not a protocol error");
|
||||
assertTrue(textOf(wrongCaller).contains("NOT_YOUR_ROLLOVER"),
|
||||
"a different lead terminal confirming must surface NOT_YOUR_ROLLOVER: " + textOf(wrongCaller));
|
||||
|
||||
McpSchema.CallToolResult confirmed = FleetMcp.handover(rollover, LEAD,
|
||||
Map.of("action", "confirm", "token", token));
|
||||
assertFalse(confirmed.isError(), textOf(confirmed));
|
||||
assertTrue(textOf(confirmed).contains("\"accepted\":true"),
|
||||
"the SAME terminal that opened the request must be able to confirm it: " + textOf(confirmed));
|
||||
}
|
||||
|
||||
// --- acceptance 6: cancel on an unknown token is clean, not a failure ----------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("cancel on an unknown token reports no pending request, rather than failing")
|
||||
void cancelUnknownTokenIsCleanNotAFailure() {
|
||||
LeadRollover rollover = newRollover(tmp.resolve("h.md").toString());
|
||||
|
||||
McpSchema.CallToolResult r = FleetMcp.handover(rollover, LEAD,
|
||||
Map.of("action", "cancel", "token", "does-not-exist"));
|
||||
assertFalse(r.isError());
|
||||
assertTrue(textOf(r).contains("\"cancelled\":false"), textOf(r));
|
||||
assertTrue(textOf(r).contains("no pending request"), textOf(r));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("cancel on a token actually opened reports cancelled:true")
|
||||
void cancelKnownTokenSucceeds() {
|
||||
LeadRollover rollover = newRollover(tmp.resolve("h.md").toString());
|
||||
String token = extractToken(textOf(FleetMcp.handover(rollover, LEAD, Map.of("action", "open"))));
|
||||
|
||||
McpSchema.CallToolResult r = FleetMcp.handover(rollover, LEAD,
|
||||
Map.of("action", "cancel", "token", token));
|
||||
assertFalse(r.isError());
|
||||
assertTrue(textOf(r).contains("\"cancelled\":true"), textOf(r));
|
||||
}
|
||||
|
||||
// --- acceptance 7 (wiring) is covered by FleetdLeadRolloverWiringTest, unchanged -----------
|
||||
}
|
||||
@@ -3,6 +3,7 @@ package dev.ltms.fleet.mcp;
|
||||
import dev.ltms.fleet.auth.CallerResolver;
|
||||
import dev.ltms.fleet.auth.MemberRegistry;
|
||||
import dev.ltms.fleet.auth.Principal;
|
||||
import dev.ltms.fleet.auth.Role;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
@@ -28,6 +29,7 @@ import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import dev.ltms.fleet.msg.InMemoryReplyInbox;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
@@ -77,7 +79,7 @@ class FleetMcpTest {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting(target), "send should be accepted for " + target);
|
||||
FleetMcp.reply(messages, target, "received");
|
||||
FleetMcp.reply(messages, target, Role.WORKER, "received");
|
||||
assertEquals("received", textOf(send.get(6, TimeUnit.SECONDS)));
|
||||
}
|
||||
|
||||
@@ -110,7 +112,7 @@ class FleetMcpTest {
|
||||
|
||||
// fleetd #365: a resolved live send must read distinctly from a merely-queued reply —
|
||||
// see replyWithNoPendingSendIsQueuedNotError below for the other case.
|
||||
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", "LGTM");
|
||||
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", Role.WORKER, "LGTM");
|
||||
assertEquals(MessageService.ReplyOutcome.RESOLVED_SEND.description(), textOf(reply));
|
||||
|
||||
McpSchema.CallToolResult res = send.get(6, TimeUnit.SECONDS);
|
||||
@@ -136,7 +138,7 @@ class FleetMcpTest {
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_a"), "send should have opened its waiter");
|
||||
|
||||
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", "async LGTM");
|
||||
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", Role.WORKER, "async LGTM");
|
||||
assertEquals(MessageService.ReplyOutcome.RESOLVED_SEND.description(), textOf(reply));
|
||||
|
||||
// Poll until the async send completes and reports the reply.
|
||||
@@ -183,7 +185,7 @@ class FleetMcpTest {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_a"));
|
||||
FleetMcp.reply(messages, "term_a", "done");
|
||||
FleetMcp.reply(messages, "term_a", Role.WORKER, "done");
|
||||
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
|
||||
|
||||
McpSchema.CallToolResult done = FleetMcp.poll(messages, ticket, null);
|
||||
@@ -215,7 +217,7 @@ class FleetMcpTest {
|
||||
// fleet_poll{ticket} stuck PENDING forever and later force-failed with a false "session
|
||||
// released before it replied" reason. This used to land in the inbox instead (see the old
|
||||
// assertion this replaced: messages.drainReplies("term_a").getFirst()...) — that was the bug.
|
||||
FleetMcp.reply(messages, "term_a", "finished after timeout");
|
||||
FleetMcp.reply(messages, "term_a", Role.WORKER, "finished after timeout");
|
||||
assertEquals("finished after timeout", textOf(FleetMcp.poll(messages, ticket, null)));
|
||||
assertTrue(messages.drainReplies("term_a").isEmpty(),
|
||||
"the reply completed its own ticket directly and never touched the inbox");
|
||||
@@ -269,7 +271,7 @@ class FleetMcpTest {
|
||||
}
|
||||
assertEquals(MessageService.Phase.FAILED, second.phase());
|
||||
|
||||
FleetMcp.reply(messages, "term_a", "late reply");
|
||||
FleetMcp.reply(messages, "term_a", Role.WORKER, "late reply");
|
||||
assertEquals("late reply", messages.drainReplies("term_a").getFirst().content());
|
||||
|
||||
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
|
||||
@@ -278,7 +280,7 @@ class FleetMcpTest {
|
||||
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
FleetMcp.reply(messages, "term_a", "done");
|
||||
FleetMcp.reply(messages, "term_a", Role.WORKER, "done");
|
||||
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
|
||||
}
|
||||
|
||||
@@ -331,7 +333,7 @@ class FleetMcpTest {
|
||||
void replyWithNoPendingSendIsQueuedNotError() {
|
||||
// CB-307: a reply with no open send is now queued in the inbox, not an error.
|
||||
// fleetd #365: it must also no longer claim "delivered" — nothing was waiting for it.
|
||||
McpSchema.CallToolResult res = FleetMcp.reply(messages, "term_a", "orphan");
|
||||
McpSchema.CallToolResult res = FleetMcp.reply(messages, "term_a", Role.WORKER, "orphan");
|
||||
assertNotEquals(Boolean.TRUE, res.isError(), "a queued reply is not an error");
|
||||
assertEquals(MessageService.ReplyOutcome.QUEUED.description(), textOf(res));
|
||||
|
||||
@@ -341,6 +343,40 @@ class FleetMcpTest {
|
||||
assertEquals("orphan", drained.getFirst().content());
|
||||
}
|
||||
|
||||
@Test
|
||||
void replyFromLeadIsRefusedBeforeItCanPublishToTheWorkerInbox() {
|
||||
ReplyInbox inboxThatRejectsPublishes = new ReplyInbox() {
|
||||
@Override public void own(String target) { }
|
||||
@Override public void release(String target) { }
|
||||
@Override public void publish(String target, String msgId, String content) {
|
||||
fail("a lead fleet_reply must not publish to the worker inbox");
|
||||
}
|
||||
@Override public List<InboxMessage> peek(String target) { return List.of(); }
|
||||
@Override public boolean ack(String target, String msgId) { return false; }
|
||||
};
|
||||
MessageService leadMessages = new MessageService(agents, new Injector(agents), new Rendezvous(),
|
||||
inboxThatRejectsPublishes);
|
||||
|
||||
McpSchema.CallToolResult res = assertDoesNotThrow(
|
||||
() -> FleetMcp.reply(leadMessages, "term_lead", Role.PRIMARY, "peer reply"));
|
||||
|
||||
assertTrue(res.isError());
|
||||
assertEquals("fleet_reply has no route to a peer lead. Use fleet_send{coordId: ...} for a peer on another "
|
||||
+ "daemon or fleet_send{sessionId: ...} for a peer on this host. fleet_reply resolves a member's "
|
||||
+ "blocked fleet_send, and a peer's coord-id message is durable and non-blocking, so there is "
|
||||
+ "nothing for it to resolve.",
|
||||
textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void replyFromUnidentifiedCallerKeepsItsOwnError() {
|
||||
McpSchema.CallToolResult res = FleetMcp.reply(messages, null, Role.PRIMARY, "reply");
|
||||
|
||||
assertTrue(res.isError());
|
||||
assertEquals("fleet_reply is for workers only — could not identify the calling worker from the connection",
|
||||
textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void replyWithBlankContentIsACleanToolErrorNotAnUncaughtException() {
|
||||
// fleetd #302: MessageService.reply now REJECTS blank content by throwing. fleet_reply's
|
||||
@@ -351,7 +387,7 @@ class FleetMcpTest {
|
||||
// has always used isBlank for exactly this reason.
|
||||
for (String blank : new String[] {null, "", " ", "\n\t"}) {
|
||||
McpSchema.CallToolResult res = assertDoesNotThrow(
|
||||
() -> FleetMcp.reply(messages, "term_a", blank),
|
||||
() -> FleetMcp.reply(messages, "term_a", Role.WORKER, blank),
|
||||
"blank content must be refused as a tool error, never thrown out of the handler");
|
||||
assertEquals(Boolean.TRUE, res.isError(), "blank content is an error result");
|
||||
assertTrue(textOf(res).contains("content is required"),
|
||||
@@ -364,7 +400,7 @@ class FleetMcpTest {
|
||||
@Test
|
||||
void bridgePollWithTargetDrainsReplies() {
|
||||
// A reply with no open send queues it in the inbox.
|
||||
FleetMcp.reply(messages, "term_a", "queued-msg");
|
||||
FleetMcp.reply(messages, "term_a", Role.WORKER, "queued-msg");
|
||||
|
||||
// fleet_poll with target drains the inbox.
|
||||
McpSchema.CallToolResult res = FleetMcp.poll(messages, null, "term_a");
|
||||
@@ -415,7 +451,7 @@ class FleetMcpTest {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_a"), "the answer should have reopened a waiter");
|
||||
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", "done");
|
||||
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", Role.WORKER, "done");
|
||||
assertEquals(MessageService.ReplyOutcome.RESOLVED_SEND.description(), textOf(reply));
|
||||
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
|
||||
}
|
||||
@@ -538,6 +574,38 @@ class FleetMcpTest {
|
||||
assertTrue(out.contains("\"quarantinedForSeconds\":1800"), out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #466 scope item 2: {@code fleet_profiles} must carry the repeat count beside the
|
||||
* remaining seconds, and the two must come off the one {@link BackendQuarantine#status} call so
|
||||
* they can never disagree about which streak this is (see {@code profilesView}'s javadoc).
|
||||
*/
|
||||
@Test
|
||||
void profilesReportsQuarantineAttemptBesideRemainingSeconds() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
java.util.concurrent.atomic.AtomicLong now = new java.util.concurrent.atomic.AtomicLong(0L);
|
||||
BackendQuarantine quarantine = BackendQuarantine.withEscalation(now::get, TimeUnit.MINUTES.toNanos(30));
|
||||
quarantine.quarantine("shared-openai"); // attempt 1: 1800s
|
||||
now.set(TimeUnit.MINUTES.toNanos(30));
|
||||
quarantine.quarantine("shared-openai"); // attempt 2: 3600s
|
||||
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
|
||||
profile -> "ltms-local".equals(profile) ? "shared-openai" : null, quarantine);
|
||||
McpSchema.CallToolResult res = FleetMcp.profiles(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), source);
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"quarantinedForSeconds\":3600"), out);
|
||||
assertTrue(out.contains("\"quarantineAttempt\":2"), out);
|
||||
}
|
||||
|
||||
/** A never-quarantined profile must not carry {@code quarantineAttempt} either. */
|
||||
@Test
|
||||
void profilesOmitsQuarantineAttemptWhenNotQuarantined() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
McpSchema.CallToolResult res = FleetMcp.profiles(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), FleetMcp.QuarantineSource.none());
|
||||
String out = textOf(res);
|
||||
assertFalse(out.contains("quarantineAttempt"), out);
|
||||
}
|
||||
|
||||
/** fleetd #201 Unit 5: {@code coolingOff} is a SEPARATE map from {@code quarantined}. */
|
||||
@Test
|
||||
void profilesReportsACoolingOffCredentialInASeparateMap() {
|
||||
@@ -590,8 +658,8 @@ class FleetMcpTest {
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), Map.of(), "",
|
||||
new FleetMcp.CoordinationSource(channel, List.of()));
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
|
||||
String out = textOf(res);
|
||||
// fleetd #361: reports both which coord-id a peer must use to reach ME, and this daemon's
|
||||
@@ -619,8 +687,8 @@ class FleetMcpTest {
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), Map.of(), "",
|
||||
new FleetMcp.CoordinationSource(channel, List.of()));
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"mailbox\":{\"status\":\"unknown\"}"), out);
|
||||
@@ -640,13 +708,20 @@ class FleetMcpTest {
|
||||
"an ordinary fleet's output must be unchanged by this feature");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #463: a compat overload called with no {@code callerIsPrimary} argument at all must
|
||||
* fail closed, not open. Before this fix the hidden default was {@code true}, so a caller that
|
||||
* forgot the argument silently got lead-to-lead coordination state. Lead coordination is fully
|
||||
* configured here (a real channel, a real mailbox) specifically so this is not conflated with
|
||||
* {@link #listOmitsTheCoordinatorRowWhenLeadCoordinationIsOff} -- the row is capable of being
|
||||
* assembled, and the missing argument is the only reason it is not.
|
||||
*/
|
||||
@Test
|
||||
void listReportsHeldMessagesWithATruncatedPreviewNeverTheFullBody() {
|
||||
void listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
String longContent = "x".repeat(200);
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", longContent));
|
||||
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
@@ -654,13 +729,226 @@ class FleetMcpTest {
|
||||
FleetMcp.QuarantineSource.none(), Map.of(), "",
|
||||
new FleetMcp.CoordinationSource(channel, List.of()));
|
||||
|
||||
String out = textOf(res);
|
||||
assertFalse(out.contains("\"coordinator\""),
|
||||
"no callerIsPrimary argument must fail closed (absent), not open (present): " + out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #439: a worker calling {@code fleet_list} must get a result with the {@code
|
||||
* coordinator} key <strong>absent</strong> -- not an empty object, not a redacted one -- even
|
||||
* though lead coordination is fully configured and would otherwise report a row. This drives
|
||||
* the same {@code callerIsPrimary} value the MCP handler computes ({@code
|
||||
* Principal.worker(...).isPrimary()}), so it pins the real production boolean, not a literal.
|
||||
*/
|
||||
@Test
|
||||
void listOmitsTheCoordinatorKeyEntirelyForAWorkerEvenWhenLeadCoordinationIsOn() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
|
||||
boolean callerIsPrimary = Principal.worker("term_a", 1).isPrimary();
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), callerIsPrimary);
|
||||
|
||||
String out = textOf(res);
|
||||
assertFalse(out.contains("\"coordinator\""), "a worker must never see the coordinator key at all: " + out);
|
||||
assertFalse(out.contains("mac-opus"), "no fragment of the coordinator row may leak either: " + out);
|
||||
assertTrue(out.contains("\"leads\""), "the rest of the result must still be present: " + out);
|
||||
assertTrue(out.contains("\"members\""), out);
|
||||
assertTrue(out.contains("\"healthCoverage\""), out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #439 acceptance criterion 2: an architect gets exactly the same treatment as a worker.
|
||||
* This is a real, executed test (not just reasoning by analogy) -- it drives the actual
|
||||
* {@code Principal.architect(...).isPrimary()} value the production handler would compute for
|
||||
* an architect caller, through the same gate a worker's call goes through.
|
||||
*/
|
||||
@Test
|
||||
void listOmitsTheCoordinatorKeyEntirelyForAnArchitectToo() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
|
||||
boolean callerIsPrimary = Principal.architect("lead-designer", "term_design", 400).isPrimary();
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), callerIsPrimary);
|
||||
|
||||
String out = textOf(res);
|
||||
assertFalse(out.contains("\"coordinator\""), "an architect must never see the coordinator key either: " + out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #439 acceptance criterion 3: an explicitly-primary caller sees the coordinator row
|
||||
* fully assembled, with the same content #439 always produced for a primary.
|
||||
*
|
||||
* <p>fleetd #463 flipped the compat overloads' hidden default from {@code true} to
|
||||
* {@code false} (fail closed), so the old "pre-#439 overload" this test used to compare
|
||||
* against no longer stands in for a primary caller -- it is now exactly the implicit-default
|
||||
* path #463 closes. Verifying the primary path means calling the canonical overload with an
|
||||
* explicit {@code callerIsPrimary=true} directly, as the production {@code fleet_list} handler
|
||||
* does.
|
||||
*/
|
||||
@Test
|
||||
void listIsByteForByteUnchangedForThePrimaryCaller() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
|
||||
|
||||
String gatedAsPrimary = textOf(FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true));
|
||||
|
||||
assertTrue(gatedAsPrimary.contains("\"coordinator\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"selfId\":\"mac-opus\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"mailbox\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"heldCount\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"heldDurable\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"held\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"peers\""), gatedAsPrimary);
|
||||
}
|
||||
|
||||
@Test
|
||||
void listReportsHeldMessagesWithATruncatedPreviewNeverTheFullBody() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
// A homogeneous "x".repeat(200) body would not pin the exact 80-char boundary: a preview
|
||||
// widened by one (81 chars) still CONTAINS "x".repeat(80) + "…" as a substring one position
|
||||
// later, because every character is 'x'. Put a sentinel ("Y") exactly at index 80 — the
|
||||
// first character a widened cap would leak — so any preview past 80 chars is caught.
|
||||
String longContent = "x".repeat(80) + "Y" + "z".repeat(119);
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", longContent));
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"msgId\":\"m1\""), out);
|
||||
assertTrue(out.contains("\"from\":\"fleet01-lead\""), out);
|
||||
assertFalse(out.contains(longContent), "fleet_list must never dump a held message's full body: " + out);
|
||||
assertFalse(out.contains("Y"),
|
||||
"the sentinel at index 80 must never appear — a preview past 80 chars leaked it: " + out);
|
||||
assertTrue(out.contains("x".repeat(80) + "…"), "expected an 80-char preview with an ellipsis: " + out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #421: {@code mailbox.pending} counts only broker-ready messages, so a blocked lead's
|
||||
* normal, healthy state is {@code "pending": 0} next to a non-empty {@code held[]} — which
|
||||
* invites the false reading that held mail is in-memory-only. {@code heldCount}/{@code
|
||||
* heldDurable} put an honest second number and fact beside {@code pending} instead of leaving it
|
||||
* as the only one.
|
||||
*/
|
||||
@Test
|
||||
void listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1))
|
||||
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", "one"))
|
||||
.hold(new LeadMessage("m2", "fleet01-lead", "mac-opus", "two"))
|
||||
.hold(new LeadMessage("m3", "fleet01-lead", "mac-opus", "three"));
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"pending\":0"), out);
|
||||
assertTrue(out.contains("\"heldCount\":3"),
|
||||
"the honest count beside pending: 0 — three messages really are held: " + out);
|
||||
assertTrue(out.contains("\"heldDurable\":true"),
|
||||
"must state the durability fact, not leave pending as the only number next to held[]: " + out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #440: {@code heldDurable} must be a derived fact, not a literal — so it can report
|
||||
* {@code false} when the channel behind it says held mail is not durable (a non-durable queue,
|
||||
* or a consumer running with {@code autoAck=true}). A test that only ever asserts {@code true}
|
||||
* repeats the defect this ticket fixes.
|
||||
*/
|
||||
@Test
|
||||
void listReportsHeldDurableFalseWhenTheChannelSaysMailIsNotDurable() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1))
|
||||
.withHeldDurable(false)
|
||||
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", "one"));
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"heldDurable\":false"),
|
||||
"heldDurable must follow the channel, not a hardcoded true: " + out);
|
||||
}
|
||||
|
||||
// ── fleetd #421: a lead reads (never consumes) its own held peer mail ──────────────────────
|
||||
|
||||
@Test
|
||||
void pollWithCoordIdReturnsTheFullBodyWithoutAckingAndLeavesItHeld() {
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", "x".repeat(200)));
|
||||
|
||||
McpSchema.CallToolResult first = FleetMcp.poll(messages, channel, null, null, "mac-opus");
|
||||
assertNotEquals(Boolean.TRUE, first.isError(), textOf(first));
|
||||
String out1 = textOf(first);
|
||||
assertTrue(out1.contains("\"msgId\":\"m1\""), out1);
|
||||
assertTrue(out1.contains("\"from\":\"fleet01-lead\""), out1);
|
||||
assertTrue(out1.contains("\"content\":\"" + "x".repeat(200) + "\""),
|
||||
"the coordId route must return the FULL body, unlike fleet_list's preview: " + out1);
|
||||
|
||||
// Read again: identical bodies, and nothing was acked — peek() still holds it.
|
||||
McpSchema.CallToolResult second = FleetMcp.poll(messages, channel, null, null, "mac-opus");
|
||||
assertEquals(out1, textOf(second), "peek is non-destructive — reading twice must return the same bodies");
|
||||
assertTrue(channel.acked().isEmpty(), "a read must never ack — that is the whole point of the ticket");
|
||||
assertEquals(1, channel.peek().size(), "the message must still be held after being read");
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollWithCoordIdRefusesAPeersCoordIdInsteadOfReturningTheWrongMailOrNothing() {
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", "secret coordination body"));
|
||||
|
||||
// The ORIGINAL fleetd #421 confusion: passing a PEER's id where a self-address was meant.
|
||||
McpSchema.CallToolResult res = FleetMcp.poll(messages, channel, null, null, "fleet01-lead");
|
||||
|
||||
assertEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
assertFalse(out.contains("secret coordination body"),
|
||||
"a refused read must never leak the body it refused to return: " + out);
|
||||
assertTrue(out.contains("mac-opus"), "the error must name this daemon's own coordId: " + out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollWithCoordIdErrorsHonestlyWhenLeadCoordinationIsNotConfigured() {
|
||||
McpSchema.CallToolResult res = FleetMcp.poll(messages, null, null, null, "mac-opus");
|
||||
|
||||
assertEquals(Boolean.TRUE, res.isError());
|
||||
assertTrue(textOf(res).toLowerCase().contains("coordinat"), textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void listReportsEachDeclaredPeersLiveReachability() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
@@ -676,8 +964,9 @@ class FleetMcpTest {
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), Map.of(), "",
|
||||
new FleetMcp.CoordinationSource(channel, List.of("fleet01-lead", "fleet02-lead", "fleet03-lead")));
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "",
|
||||
new FleetMcp.CoordinationSource(channel, List.of("fleet01-lead", "fleet02-lead", "fleet03-lead")), true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"coordId\":\"fleet01-lead\",\"status\":\"exists\",\"pending\":2,\"consumers\":1"), out);
|
||||
@@ -716,6 +1005,9 @@ class FleetMcpTest {
|
||||
@Override
|
||||
public String selfCoordId() { return "mac-opus"; }
|
||||
|
||||
@Override
|
||||
public boolean heldDurable() { return true; }
|
||||
|
||||
@Override
|
||||
public MailboxState inspect(String coordId) {
|
||||
started.countDown();
|
||||
@@ -861,6 +1153,35 @@ class FleetMcpTest {
|
||||
assertTrue(out.contains("\"quarantinedForSeconds\":1200"), out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #466 scope item 2: {@code fleet_list}'s capacity rows (CB-583: they reuse quarantine)
|
||||
* must carry {@code quarantineAttempt} beside {@code quarantinedForSeconds}, off the same
|
||||
* {@link BackendQuarantine#status} call as {@code fleet_profiles} -- see {@code capacityView}'s
|
||||
* comment pointing back to {@code profilesView}.
|
||||
*/
|
||||
@Test
|
||||
void capacityRowReportsQuarantineAttemptBesideRemainingSeconds() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
java.util.concurrent.atomic.AtomicLong now = new java.util.concurrent.atomic.AtomicLong(0L);
|
||||
BackendQuarantine quarantine = BackendQuarantine.withEscalation(now::get, TimeUnit.MINUTES.toNanos(20));
|
||||
quarantine.quarantine("shared-openai"); // attempt 1
|
||||
now.set(TimeUnit.MINUTES.toNanos(20));
|
||||
quarantine.quarantine("shared-openai"); // attempt 2
|
||||
now.set(TimeUnit.MINUTES.toNanos(60));
|
||||
quarantine.quarantine("shared-openai"); // attempt 3
|
||||
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
|
||||
profile -> "terra".equals(profile) ? "shared-openai" : null, quarantine);
|
||||
|
||||
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 2,
|
||||
() -> Set.of("terra"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
source, Map.of(), ""));
|
||||
|
||||
assertTrue(out.contains("\"credentialId\":\"shared-openai\""), out);
|
||||
assertTrue(out.contains("\"quarantineAttempt\":3"), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void everyProfileSharingTheQuarantinedCredentialReportsZeroFree() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
@@ -1239,6 +1560,10 @@ class FleetMcpTest {
|
||||
|
||||
@Test
|
||||
void bridgeAckReturnsConfirmationForValidArgs() {
|
||||
// fleet_ack only reports success for a msgId actually queued in the target's inbox
|
||||
// (fleetd #437) — publish one via the inbox directly rather than asserting on a
|
||||
// fabricated id nothing ever queued.
|
||||
inbox.publish("term_a", "msg-1", "queued reply");
|
||||
McpSchema.CallToolResult res = FleetMcp.ack(messages, "term_a", "msg-1");
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
assertTrue(textOf(res).contains("msg-1"), "response should mention the msgId");
|
||||
@@ -1251,21 +1576,51 @@ class FleetMcpTest {
|
||||
assertTrue(FleetMcp.ack(messages, " ", "msg-1").isError());
|
||||
}
|
||||
|
||||
@Test
|
||||
void bridgeAckOfAnIdInNoInboxIsAnError() {
|
||||
// fleetd #437: fleet_ack used to say "acknowledged <msgId>" for a message it never
|
||||
// touched, because nothing in the chain reported hit vs. miss. "never-queued" is in no
|
||||
// inbox at all, so this must error rather than claim success.
|
||||
McpSchema.CallToolResult res = FleetMcp.ack(messages, "term_a", "never-queued");
|
||||
assertTrue(res.isError());
|
||||
assertTrue(textOf(res).contains("never-queued"), textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void bridgeAckOfACoordIdTargetIsAnErrorNamingFleetPoll() {
|
||||
// A coord-id names a peer lead's held mailbox (LeadChannel/LeadMailbox), never a
|
||||
// worker's ReplyInbox — fleet_ack has no route to it and must say so, pointing at
|
||||
// fleet_poll{coordId} instead of reporting a false "acknowledged".
|
||||
McpSchema.CallToolResult res = FleetMcp.ack(messages, "coord-some-peer", "msg-1");
|
||||
assertTrue(res.isError());
|
||||
assertTrue(textOf(res).contains("fleet_poll{coordId}"), textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void bridgeAckRemovesSpecificReply() {
|
||||
// Queue a reply and capture its msgId.
|
||||
FleetMcp.reply(messages, "term_a", "orphan");
|
||||
FleetMcp.reply(messages, "term_a", Role.WORKER, "orphan");
|
||||
var before = messages.drainReplies("term_a");
|
||||
assertEquals(1, before.size(), "one reply in the inbox");
|
||||
String msgId = before.getFirst().msgId();
|
||||
|
||||
// Publish the same reply again and ack it via fleet_ack surface.
|
||||
FleetMcp.reply(messages, "term_a", "orphan-again");
|
||||
FleetMcp.reply(messages, "term_a", Role.WORKER, "orphan-again");
|
||||
var peeked = messages.drainReplies("term_a");
|
||||
assertEquals(1, peeked.size(), "one fresh reply in the inbox");
|
||||
|
||||
// ackReply works (no-op since published with a different UUID, but callable).
|
||||
assertDoesNotThrow(() -> messages.ackReply("term_a", msgId));
|
||||
// fleetd #437: msgId was already drained above (a fresh UUID each publish), so it is no
|
||||
// longer in the inbox — ackReply must now report that miss instead of pretending to ack.
|
||||
assertFalse(messages.ackReply("term_a", msgId));
|
||||
}
|
||||
|
||||
@Test
|
||||
void bridgeAckRemovingARealQueuedReplyReportsSuccessAndRemovesIt() {
|
||||
// The worker path must not change behaviour: acking a reply that IS still in the inbox
|
||||
// still succeeds and still removes it (fleetd #437).
|
||||
inbox.publish("term_a", "real-1", "still queued");
|
||||
assertTrue(messages.ackReply("term_a", "real-1"), "ack of a real queued reply must report true");
|
||||
assertTrue(inbox.peek("term_a").isEmpty(), "the acked reply must be gone from the inbox");
|
||||
}
|
||||
|
||||
@Test
|
||||
|
||||
@@ -0,0 +1,62 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #395: {@code fleet_profiles} must let an operator tell "this profile's usage-limit
|
||||
* detection is armed" from "nothing can ever quarantine this profile" — see {@link
|
||||
* FleetMcp.QuarantineSource#exhaustedPatternArmed()} and {@link
|
||||
* FleetMcp#profilesView(PeerLauncher, FleetMcp.QuarantineSource, FleetMcp.OutageSource)}.
|
||||
*/
|
||||
class FleetProfilesArmedFieldTest {
|
||||
|
||||
private static PeerLauncher twoProfileLauncher(FakeHerdr h) {
|
||||
FleetConfig.Profile armed = new FleetConfig.Profile(
|
||||
"armed-profile", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
|
||||
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
FleetConfig.Profile unarmed = new FleetConfig.Profile(
|
||||
"unarmed-profile", "http://gx01.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
|
||||
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
Map<String, FleetConfig.Profile> profiles = new LinkedHashMap<>();
|
||||
profiles.put(armed.profile(), armed);
|
||||
profiles.put(unarmed.profile(), unarmed);
|
||||
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
|
||||
new SubscriptionGuard(Set.of("gx00.gw", "gx01.gw")), profiles, armed.profile(), _ -> "tok");
|
||||
}
|
||||
|
||||
private static String textOf(McpSchema.CallToolResult r) {
|
||||
return ((McpSchema.TextContent) r.content().getFirst()).text();
|
||||
}
|
||||
|
||||
@Test
|
||||
void armedProfileReportsArmedAndUnarmedReportsUnarmed() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
PeerLauncher workers = twoProfileLauncher(h);
|
||||
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
|
||||
_ -> null, dev.ltms.fleet.placement.BackendQuarantine.none(),
|
||||
profile -> "armed-profile".equals(profile));
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.profiles(workers, source);
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
|
||||
assertTrue(out.contains("\"exhaustionDetectionArmed\""), out);
|
||||
assertTrue(out.contains("\"armed-profile\":true"), out);
|
||||
assertTrue(out.contains("\"unarmed-profile\":false"), out);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,114 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||
import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #425: {@code fleet_profiles}' {@code "default"} field was captured once at boot
|
||||
* ({@code cfg.effectiveDefaultProfile()}, frozen into {@code CompositePeerLauncher.defaultProfile}
|
||||
* at construction) while an unqualified spawn resolves the same underlying key
|
||||
* ({@code fleet.developers}' first entry) live, on every call. Reordering {@code fleet.developers}
|
||||
* and reloading changed where a spawn landed without ever changing what {@code fleet_profiles}
|
||||
* reported — a lead following {@code CLAUDE.md}'s "check {@code fleet_profiles} once per session"
|
||||
* instruction was told a stale answer.
|
||||
*
|
||||
* <p>This test drives the exact caller {@code fleet_profiles} uses —
|
||||
* {@link FleetMcp#profilesView(PeerLauncher, FleetMcp.QuarantineSource, FleetMcp.OutageSource)} —
|
||||
* against a real, reloadable {@link ConfigRef}, so it fails if the reporting path is ever recoupled
|
||||
* to a frozen value instead of {@link CompositePeerLauncher#defaultProfile()}'s live answer.
|
||||
*/
|
||||
class FleetProfilesLiveDefaultTest {
|
||||
|
||||
/** A minimal fleetd.yaml whose dev pool is {@code profilesInOrder}, in that definition order. */
|
||||
private static String yamlWithDevPool(String... profilesInOrder) {
|
||||
StringBuilder devPool = new StringBuilder();
|
||||
for (int i = 0; i < profilesInOrder.length; i++) {
|
||||
devPool.append(" slot").append(i).append(":\n profile: ")
|
||||
.append(profilesInOrder[i]).append('\n');
|
||||
}
|
||||
return """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
opus:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: opus-coder
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet-coder
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
fleet:
|
||||
developers:
|
||||
""" + devPool;
|
||||
}
|
||||
|
||||
@Test
|
||||
void fleetProfilesDefaultTracksALiveDevPoolReorderAfterReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yamlWithDevPool("opus", "sonnet"));
|
||||
ConfigRef ref = new ConfigRef(f, FleetConfig.load(f));
|
||||
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of(
|
||||
"opus", new FleetConfig.Profile("opus", "http://gx00.gw:8000", "opus-coder", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, null, null),
|
||||
"sonnet", new FleetConfig.Profile("sonnet", "http://gx00.gw:8000", "sonnet-coder", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, null, null));
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "opus", _ -> null);
|
||||
PeerLauncher workers = new CompositePeerLauncher(
|
||||
List.of(adapter), "opus", ref, _ -> 0, BackendQuarantine.none());
|
||||
|
||||
assertReportedDefaultMatchesAnUnqualifiedSpawn(workers, "opus");
|
||||
|
||||
Files.writeString(f, yamlWithDevPool("sonnet", "opus"));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), () -> "reload should apply cleanly: " + out.error());
|
||||
|
||||
assertReportedDefaultMatchesAnUnqualifiedSpawn(workers, "sonnet");
|
||||
}
|
||||
|
||||
/**
|
||||
* Asserts BOTH that {@code fleet_profiles}' {@code "default"} equals {@code expected}, AND that
|
||||
* it equals what a real unqualified {@code MemberRole#DEV} spawn actually gets placed on right
|
||||
* now — the two facts fleetd #425 found disagreeing.
|
||||
*/
|
||||
private static void assertReportedDefaultMatchesAnUnqualifiedSpawn(PeerLauncher workers, String expected) {
|
||||
Map<String, Object> view = FleetMcp.profilesView(
|
||||
workers, FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none());
|
||||
assertEquals(expected, view.get("default"),
|
||||
"fleet_profiles' \"default\" must be the live dev-pool answer, not a boot-time snapshot");
|
||||
|
||||
String placed = workers.spawn(
|
||||
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile();
|
||||
assertEquals(expected, placed,
|
||||
"sanity: the profile an unqualified dev spawn actually lands on");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,109 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||
import dev.ltms.fleet.member.HerdrPeerLauncher;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up: {@code fleet_profiles}/{@code GET /profiles} — the reporting surface a
|
||||
* lead actually reads — must let it tell apart the three states {@link
|
||||
* dev.ltms.fleet.peer.PeerLauncher#disabledModels()} alone collapses into one empty set: no
|
||||
* {@code models:} block at all, a block armed with nothing currently off, and a block with N
|
||||
* models off. See {@link dev.ltms.fleet.peer.PeerLauncher.ModelGateState}'s javadoc for why a
|
||||
* bare {@code disabledModels()} read cannot make this distinction, and {@link
|
||||
* FleetMcp#profilesView} for where {@code modelGateArmed} is added alongside the existing {@code
|
||||
* modelsOff} key.
|
||||
*
|
||||
* <p>Every assertion here goes through {@link FleetMcp#profilesView}, never {@code
|
||||
* PeerLauncher.modelGateState()} directly — {@code CompositePeerLauncherTest} already proves the
|
||||
* accessor itself; this class proves the surface a lead reads (fleet_profiles / GET /profiles)
|
||||
* renders what that accessor reports.
|
||||
*/
|
||||
class FleetProfilesModelGateStateTest {
|
||||
|
||||
private static FleetConfig.Profile profile(String name, String model) {
|
||||
return new FleetConfig.Profile(name, "http://gx00.gw:8000", model, null, "FLEETD_WORKER_TOKEN",
|
||||
null, "tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
}
|
||||
|
||||
private static FleetMcp.QuarantineSource noQuarantine() {
|
||||
return new FleetMcp.QuarantineSource(_ -> null, BackendQuarantine.none(), _ -> false);
|
||||
}
|
||||
|
||||
private static HerdrPeerLauncher claudeAdapter(FakeHerdr h, Map<String, FleetConfig.Profile> profiles) {
|
||||
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "local", _ -> "tok");
|
||||
}
|
||||
|
||||
/**
|
||||
* State 1: no {@code models:} block at all — a plain {@code ClaudeCodeLauncher} (no {@code
|
||||
* models:} supplier exists for it to read) has nothing to gate against, matching the fleet01
|
||||
* host measured for this ticket: {@code grep -c '^models:' fleetd.yaml} returns 0 there.
|
||||
*/
|
||||
@Test
|
||||
void noModelsBlockReportsGateNotArmed() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("local", profile("local", "deepseek-v4-flash"));
|
||||
PeerLauncher workers = claudeAdapter(h, profiles);
|
||||
|
||||
Map<String, Object> view = FleetMcp.profilesView(workers, noQuarantine(), FleetMcp.OutageSource.none());
|
||||
|
||||
assertEquals(Boolean.FALSE, view.get("modelGateArmed"),
|
||||
"no models: block to read from — nothing is gated, and nothing can be");
|
||||
assertFalse(view.containsKey("modelsOff"), "nothing configured, so no off set to report either");
|
||||
}
|
||||
|
||||
/** State 2: a {@code models:} block is present, but nothing in it is currently turned off. */
|
||||
@Test
|
||||
void modelsBlockWithNothingOffReportsGateArmedAndZeroOff() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("local", profile("local", "deepseek-v4-flash"));
|
||||
FleetConfig.Models models = new FleetConfig.Models(
|
||||
List.of(new FleetConfig.Models.ModelEntry("deepseek-v4-flash", true)));
|
||||
PeerLauncher workers = new CompositePeerLauncher(List.of(claudeAdapter(h, profiles)), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(),
|
||||
new BackendOutagePolicy(System::nanoTime), () -> models);
|
||||
|
||||
Map<String, Object> view = FleetMcp.profilesView(workers, noQuarantine(), FleetMcp.OutageSource.none());
|
||||
|
||||
assertEquals(Boolean.TRUE, view.get("modelGateArmed"),
|
||||
"a models: block is present, so the gate is armed even though nothing is off yet");
|
||||
assertFalse(view.containsKey("modelsOff"),
|
||||
"nothing is off, so the key stays absent — an empty list here would be indistinguishable "
|
||||
+ "from today's modelsOff omission, exactly the ambiguity modelGateArmed exists to remove");
|
||||
}
|
||||
|
||||
/** State 3: a {@code models:} block is present with one model currently turned off. */
|
||||
@Test
|
||||
void modelsBlockWithModelsOffReportsGateArmedAndTheOffSet() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("local", profile("local", "deepseek-v4-flash"));
|
||||
FleetConfig.Models models = new FleetConfig.Models(
|
||||
List.of(new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
PeerLauncher workers = new CompositePeerLauncher(List.of(claudeAdapter(h, profiles)), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(),
|
||||
new BackendOutagePolicy(System::nanoTime), () -> models);
|
||||
|
||||
Map<String, Object> view = FleetMcp.profilesView(workers, noQuarantine(), FleetMcp.OutageSource.none());
|
||||
|
||||
assertEquals(Boolean.TRUE, view.get("modelGateArmed"));
|
||||
assertEquals(List.of("deepseek-v4-flash"), view.get("modelsOff"));
|
||||
}
|
||||
}
|
||||
+111
@@ -0,0 +1,111 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #446 follow-up (criterion 3): a mutation battery run against the merged PR #457 proved
|
||||
* that deleting either {@code row.put("model", model);} or {@code row.put("reason", reason);}
|
||||
* from {@link FleetMcp#profilesView} left all 1586 existing tests green — nothing exercised a
|
||||
* {@link FleetMcp.QuarantineSource} whose {@code modelFor}/{@code reasonFor} actually return a
|
||||
* value. This class is what makes both deletions red, and also pins the two "must be omitted, not
|
||||
* emitted as null/blank" branches those two {@code if} guards exist for — {@link
|
||||
* FleetMcpTest#profilesReportsAQuarantinedCredential} covers the credential/quarantine-duration
|
||||
* shape but supplies neither function, so it cannot distinguish "the field is missing" from "the
|
||||
* field was never asked for".
|
||||
*/
|
||||
class FleetProfilesQuarantineModelReasonFieldsTest {
|
||||
|
||||
private static final String PROFILE = "terra";
|
||||
private static final String CREDENTIAL = "cred-terra";
|
||||
|
||||
private static PeerLauncher launcher(FakeHerdr h) {
|
||||
FleetConfig.Profile profile = new FleetConfig.Profile(
|
||||
PROFILE, "http://gx00.gw:8000", "claude-opus-5", null, "FLEETD_WORKER_TOKEN", null,
|
||||
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(PROFILE, profile), PROFILE, _ -> "tok");
|
||||
}
|
||||
|
||||
private static BackendQuarantine quarantined(String credentialId) {
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
quarantine.quarantine(credentialId);
|
||||
return quarantine;
|
||||
}
|
||||
|
||||
private static String textOf(McpSchema.CallToolResult r) {
|
||||
return ((McpSchema.TextContent) r.content().getFirst()).text();
|
||||
}
|
||||
|
||||
@Test
|
||||
void quarantinedRowNamesTheModelAndTheReason() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
|
||||
_ -> CREDENTIAL, quarantined(CREDENTIAL), _ -> true,
|
||||
_ -> "claude-opus-5", _ -> "The usage limit has been reached");
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.profiles(launcher(h), source);
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
|
||||
assertTrue(out.contains("\"quarantined\""), out);
|
||||
assertTrue(out.contains("\"model\":\"claude-opus-5\""),
|
||||
"the quarantined row must name the model the fix applies to: " + out);
|
||||
assertTrue(out.contains("\"reason\":\"The usage limit has been reached\""),
|
||||
"the quarantined row must name why it was quarantined: " + out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void quarantinedRowOmitsModelWhenNoneIsConfigured() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
// null and "" (blank) both exercise the same production guard (`!model.isBlank()`) —
|
||||
// covered together since both must produce the identical outcome: the key absent.
|
||||
for (String noModel : new String[] {null, "", " "}) {
|
||||
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
|
||||
_ -> CREDENTIAL, quarantined(CREDENTIAL), _ -> true,
|
||||
_ -> noModel, _ -> "some reason");
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.profiles(launcher(h), source);
|
||||
String out = textOf(res);
|
||||
|
||||
assertTrue(out.contains("\"quarantined\""), out);
|
||||
assertFalse(out.contains("\"model\""),
|
||||
"modelFor returned " + (noModel == null ? "null" : "'" + noModel + "'")
|
||||
+ " — the key must be ABSENT, not emitted as null or an empty string: " + out);
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void quarantinedRowOmitsReasonWhenNoneIsKnown() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
for (String noReason : new String[] {null, "", " "}) {
|
||||
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
|
||||
_ -> CREDENTIAL, quarantined(CREDENTIAL), _ -> true,
|
||||
_ -> "claude-opus-5", _ -> noReason);
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.profiles(launcher(h), source);
|
||||
String out = textOf(res);
|
||||
|
||||
assertTrue(out.contains("\"quarantined\""), out);
|
||||
assertFalse(out.contains("\"reason\""),
|
||||
"reasonFor returned " + (noReason == null ? "null" : "'" + noReason + "'")
|
||||
+ " — the key must be ABSENT, not emitted as null or an empty string: " + out);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -27,16 +27,21 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
* somewhere to be readable, and that is exactly the sentence that rots. This test is what makes it
|
||||
* safe to write.
|
||||
*
|
||||
* <p><b>It checks source text, not behaviour.</b> It reads the Markdown and reads {@link FleetMcp}'s
|
||||
* source, and it only catches a name in the doc that the server does not register. It cannot catch a
|
||||
* flow that describes the wrong order, or a parameter name in prose — those are not name-shaped. The
|
||||
* doc's own header carries that caveat for its readers.
|
||||
* <p><b>It checks source text, not behaviour.</b> It reads the Markdown, and it only catches a name
|
||||
* in the doc that the server does not register. It cannot catch a flow that describes the wrong
|
||||
* order, or a parameter name in prose — those are not name-shaped. The doc's own header carries
|
||||
* that caveat for its readers.
|
||||
*
|
||||
* <p>fleetd #469: the registered side used to be its own scrape of {@code FleetMcp.java}'s {@code
|
||||
* tool("…")} calls — a third copy of the same list {@code CharterToolSurfaceTest} and {@code
|
||||
* FleetMcpAuthzTest} each kept their own copy of too. All three now read {@link
|
||||
* FleetTool#wireNames()}, the one canonical set {@code FleetMcp} itself derives its tool schemas and
|
||||
* authorization switch from.
|
||||
*/
|
||||
class McpContractDocTest {
|
||||
|
||||
/** Tests run with the module directory as cwd, so the repo-root doc is one level up. */
|
||||
private static final Path DOC = Path.of("../docs/MCP-Contract.md");
|
||||
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
|
||||
|
||||
private static Set<String> matches(Path file, String regex) throws Exception {
|
||||
Matcher m = Pattern.compile(regex).matcher(Files.readString(file));
|
||||
@@ -52,15 +57,10 @@ class McpContractDocTest {
|
||||
return matches(DOC, "(fleet_[a-z_]+)");
|
||||
}
|
||||
|
||||
/** Every tool {@link FleetMcp} actually registers, read from its {@code tool("…")} calls. */
|
||||
private static Set<String> toolsTheServerRegisters() throws Exception {
|
||||
return matches(MCP_SOURCE, "tool\\(\"(fleet_[a-z_]+)\"");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] every fleet_* tool named in MCP-Contract.md is one the server registers")
|
||||
void theDocNamesNoToolThatDoesNotExist() throws Exception {
|
||||
Set<String> registered = toolsTheServerRegisters();
|
||||
Set<String> registered = FleetTool.wireNames();
|
||||
Set<String> named = toolsNamedInTheDoc();
|
||||
|
||||
Set<String> unknown = new LinkedHashSet<>(named);
|
||||
@@ -84,13 +84,13 @@ class McpContractDocTest {
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] the doc/server name check is not vacuous — both sides found names")
|
||||
void theCheckActuallyHasSomethingToCheck() throws Exception {
|
||||
Set<String> registered = toolsTheServerRegisters();
|
||||
Set<String> registered = FleetTool.wireNames();
|
||||
Set<String> named = toolsNamedInTheDoc();
|
||||
|
||||
assertTrue(registered.size() >= 10,
|
||||
"scraped only " + registered.size() + " tool registrations from FleetMcp (" + registered
|
||||
+ "); the server registers eleven, so the tool(\"…\") scrape has stopped matching "
|
||||
+ "and the check above is now vacuous");
|
||||
"FleetTool.wireNames() returned only " + registered.size() + " tool(s) (" + registered
|
||||
+ "); the server registers eleven, so FleetTool has stopped listing the real "
|
||||
+ "tool surface and the check above is now vacuous");
|
||||
assertTrue(named.size() >= 4,
|
||||
"docs/MCP-Contract.md names only " + named.size() + " fleet_* tool(s) (" + named + "). "
|
||||
+ "The flows describe delegation, clarification, detached delivery and the "
|
||||
|
||||
@@ -3,6 +3,7 @@ package dev.ltms.fleet.member;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.Agent;
|
||||
@@ -19,11 +20,15 @@ import dev.ltms.fleet.peer.PeerUnreachableException;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
import dev.ltms.fleet.placement.PlacementException;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.EnumSet;
|
||||
import java.util.HashMap;
|
||||
import java.util.LinkedHashMap;
|
||||
@@ -141,6 +146,13 @@ class CompositePeerLauncherTest {
|
||||
null, null, credentialId, null);
|
||||
}
|
||||
|
||||
/** Like {@link #stubWorker(String)}, but with an explicit {@code model:} for fleetd #422 tests. */
|
||||
private static FleetConfig.Profile stubWorkerModel(String profile, String model) {
|
||||
return new FleetConfig.Profile(profile, "http://gx00.gw:8000", model,
|
||||
null, "FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"w #{n}", null, null, null, null, null, null, null, null, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* An <em>order-preserving</em> profile map. Never {@code Map.of} here: its iteration order is
|
||||
* salted per JVM run, and the weighted policy breaks an exact-weight tie on candidate order —
|
||||
@@ -525,6 +537,31 @@ class CompositePeerLauncherTest {
|
||||
assertEquals("claude", h.profile(), "the returned handle carries the resolved default profile");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #435: {@code FixedPlacementPolicy} — the default placement policy every config uses
|
||||
* unless {@code placement:} is set — never consulted {@code maxLoad}, so an unqualified spawn
|
||||
* (a blank profile, the normal delegation path) landed on a capped default anyway. Measured on
|
||||
* 7667727: a single dev profile at {@code maxLoad: 1} with 1 live, under {@code fixed()},
|
||||
* returned "SPAWNED on profile=a". This test goes through {@code CompositePeerLauncher.spawn}
|
||||
* with a blank profile, not the policy in isolation, so it proves the caller actually reaches
|
||||
* the fixed default's new cap check rather than only the {@code select} method.
|
||||
*/
|
||||
@Test
|
||||
void fixedPolicyGatesDefaultProfileAtMaxLoadOnUnqualifiedSpawn() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"a", stubWorker("a", 1.0f, 1),
|
||||
"b", stubWorker("b", 1.0f, null));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), name -> "a".equals(name) ? 1 : 0);
|
||||
|
||||
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
|
||||
assertEquals("b", h.profile(),
|
||||
"the default profile a is at maxLoad, so fixed placement must fall through to b");
|
||||
assertEquals(0, adapter.spawnCount("a"), "a is never spawned — it is already at cap");
|
||||
}
|
||||
|
||||
@Test
|
||||
void weightedPolicyGatesProfileAtMaxLoad() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
@@ -908,6 +945,89 @@ class CompositePeerLauncherTest {
|
||||
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
|
||||
}
|
||||
|
||||
// ── fleetd #425: defaultProfile()/defaultProfileFor() must track a live reload ─────────────
|
||||
|
||||
/** A minimal fleetd.yaml whose dev pool is {@code profilesInOrder}, in that definition order. */
|
||||
private static String yamlWithDevPool(String... profilesInOrder) {
|
||||
StringBuilder devPool = new StringBuilder();
|
||||
for (int i = 0; i < profilesInOrder.length; i++) {
|
||||
devPool.append(" slot").append(i).append(":\n profile: ")
|
||||
.append(profilesInOrder[i]).append('\n');
|
||||
}
|
||||
return """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
opus:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: opus-coder
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet-coder
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
fleet:
|
||||
developers:
|
||||
""" + devPool;
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 1 (fleetd #425): reorder {@code fleet.developers}, reload, and assert the reported
|
||||
* default ({@link CompositePeerLauncher#defaultProfile()} — what {@code fleet_profiles}' {@code
|
||||
* "default"} is built from, see {@code FleetMcp.profilesView}) matches what an unqualified
|
||||
* {@code MemberRole#DEV} spawn is actually placed on, both before and after the reorder. Asserts
|
||||
* {@code applied()} so the test proves the reload actually took, not that nothing changed.
|
||||
*/
|
||||
@Test
|
||||
void defaultProfileTracksALiveDevPoolReorderAfterReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yamlWithDevPool("opus", "sonnet"));
|
||||
ConfigRef ref = new ConfigRef(f, FleetConfig.load(f));
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, threeProfiles(), "opus", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "opus", ref, _ -> 0, BackendQuarantine.none());
|
||||
|
||||
assertEquals("opus", composite.defaultProfile(),
|
||||
"reported default starts at the dev pool's first entry");
|
||||
assertEquals("opus", composite.spawn(
|
||||
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile(),
|
||||
"an unqualified dev spawn must land on the same profile that was just reported");
|
||||
|
||||
Files.writeString(f, yamlWithDevPool("sonnet", "opus"));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), () -> "reload should apply cleanly: " + out.error());
|
||||
|
||||
assertEquals("sonnet", composite.defaultProfile(),
|
||||
"the reported default must follow the reorder with no daemon restart");
|
||||
assertEquals("sonnet", composite.spawn(
|
||||
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile(),
|
||||
"and it must still be exactly what an unqualified spawn actually gets");
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 2 — the mirror, and the load-bearing half (fleetd #425): with NOTHING configured (no
|
||||
* profiles at all, hence an empty pool for every role), the frozen {@code defaultProfile} field
|
||||
* is still what gets reported. A fix that always returns {@code poolFor(role).getFirst()} with no
|
||||
* empty-pool fallback throws or returns the wrong thing here even though criterion 1 above still
|
||||
* passes — this is the test that catches it.
|
||||
*/
|
||||
@Test
|
||||
void defaultProfileFallsBackToTheFrozenFieldWhenNothingIsConfiguredAtAll() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, Map.of(), "opus", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "opus", Map.of(), PlacementPolicies.fixed(), _ -> 0);
|
||||
|
||||
assertEquals("opus", composite.defaultProfile(),
|
||||
"with no profiles configured at all, the frozen field is the only answer available");
|
||||
assertEquals("opus", composite.defaultProfileFor(MemberRole.DEV));
|
||||
}
|
||||
|
||||
// ── CB-578 stage B: a BACKEND_EXHAUSTED classification quarantines the credential ──────────
|
||||
|
||||
@Test
|
||||
@@ -974,6 +1094,98 @@ class CompositePeerLauncherTest {
|
||||
assertEquals(0, adapter.spawnCount("sol"));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #425 rework, acceptance 1: {@link CompositePeerLauncher#routedProfileFor} must apply
|
||||
* the SAME quarantine filtering {@link CompositePeerLauncher#spawn} does, under {@code fixed()}
|
||||
* — the DEFAULT placement policy, deliberately not {@code weighted()} (which the regressed
|
||||
* round's own tests all used, and which never exercises {@code FixedPlacementPolicy}'s own
|
||||
* inline filter). This is the exact defect: the previous round's {@code defaultProfileFor}
|
||||
* blindly returns the pool's first entry ("sol", quarantined here) with no awareness of
|
||||
* quarantine at all, which is what turned a routine unqualified spawn into a hard throw once
|
||||
* {@code acquireWithWorktree} pre-resolved through it.
|
||||
*/
|
||||
@Test
|
||||
void routedProfileForSkipsAQuarantinedPoolFirstProfileUnderFixedPolicy() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"sol", stubWorker("sol", "shared-openai"),
|
||||
"b", stubWorker("b"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
quarantine.quarantine("shared-openai");
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
|
||||
|
||||
assertEquals("b", composite.routedProfileFor(MemberRole.DEV),
|
||||
"sol (the pool's first entry) is quarantined, so the routed answer must be b");
|
||||
assertEquals("sol", composite.defaultProfileFor(MemberRole.DEV),
|
||||
"sanity: defaultProfileFor stays blind to quarantine — that's the gap routedProfileFor closes");
|
||||
assertEquals(0, adapter.spawnCount("sol"), "routedProfileFor never spawns anything");
|
||||
assertEquals(0, adapter.spawnCount("b"), "routedProfileFor never spawns anything");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #444: {@link PlacementDecision} exists to close the window between {@link
|
||||
* CompositePeerLauncher#place} and {@link CompositePeerLauncher#spawn(SpawnRequest,
|
||||
* PlacementDecision)} — the placement state must be free to move in that window without the
|
||||
* held decision being re-checked against the new state. Every quarantine test above resolves
|
||||
* and spawns in one call, so none of them ever open that window; this test is the one that
|
||||
* does: "sol" is placed FIRST, while nothing is quarantined yet, and only THEN is its
|
||||
* credential quarantined, before the held decision is spawned.
|
||||
*
|
||||
* <p>This is the test that tells the real override apart from the alternative body the ticket
|
||||
* measured: routing {@code decision.profile()} straight to its adapter (the real override)
|
||||
* never re-runs {@code enforceNotQuarantined}, so the spawn against the held decision still
|
||||
* succeeds on sol. Re-entering {@code spawn(req.withProfile(decision.profile()))} instead
|
||||
* lands in the explicit-profile branch, which refuses a now-quarantined sol outright — before
|
||||
* this test existed, replacing the real override's body with that re-entering call left the
|
||||
* whole suite green.
|
||||
*/
|
||||
@Test
|
||||
void spawnHonorsAPlacementDecisionEvenAfterItsProfileIsQuarantinedInTheWindowAfterPlace() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"sol", stubWorker("sol", "shared-openai"),
|
||||
"b", stubWorker("b"));
|
||||
// The adapter's OWN fallback default is "b", deliberately different from the profile place()
|
||||
// decides ("sol") — see the note below on why this must not be "sol" too.
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "b", Set.of());
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
|
||||
|
||||
// 1. Resolve BEFORE anything is quarantined — sol (definition order first, fixed policy) wins.
|
||||
// composite's own defaultProfile ("sol", the constructor arg above) never enters this: the
|
||||
// pool poolFor(DEV) resolves to is never empty here, so place() only ever reads that field as
|
||||
// a fallback for an empty pool, which this test does not exercise.
|
||||
PlacementDecision decision = composite.place(MemberRole.DEV);
|
||||
assertEquals("sol", decision.profile(), "sanity: nothing is quarantined yet, so sol is placed");
|
||||
|
||||
// 2. Move the placement state IN THE WINDOW between place() and spawn() — sol's credential
|
||||
// is now quarantined. A fresh place()/spawn(req) pair would fall through to b instead; the
|
||||
// held decision must not be re-evaluated against this new state at all.
|
||||
quarantine.quarantine("shared-openai");
|
||||
|
||||
// 3. Spawn against the HELD decision, not a fresh resolve.
|
||||
SpawnRequest req = new SpawnRequest(null, null, null, null, null, MemberRole.DEV);
|
||||
PeerHandle handle = composite.spawn(req, decision);
|
||||
|
||||
assertEquals("sol", handle.profile(),
|
||||
"the decision from place() is honored even though sol is now quarantined");
|
||||
// A fixture whose adapter falls back to "sol" too would let an UNSTAMPED request (one
|
||||
// routed but never given req.withProfile("sol")) land on spawnCount("sol") == 1 by
|
||||
// COINCIDENCE, since StubLauncher.spawn falls back to its own defaultProfile whenever
|
||||
// req.profileName() is blank. Giving the adapter "b" as its fallback instead means only an
|
||||
// actually-stamped request can produce this count — an unstamped one would count against
|
||||
// "b" and this assertion would fail.
|
||||
assertEquals(1, adapter.spawnCount("sol"),
|
||||
"the request that reached the delegate actually carried sol as its profile "
|
||||
+ "(the adapter's own fallback default is 'b', so this can't happen by accident)");
|
||||
assertEquals(0, adapter.spawnCount("b"),
|
||||
"b must never be touched — neither as the decision's profile nor as an unstamped "
|
||||
+ "request's accidental fallback");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aQuarantineLiftsOnTheInjectedClockAndTheProfileBecomesSpawnableAgain() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
@@ -1189,4 +1401,415 @@ class CompositePeerLauncherTest {
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("claude", null, null)));
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("gemini", null, null)));
|
||||
}
|
||||
|
||||
// ── fleetd #422: the model on/off gate — a FOURTH, independent reason to refuse a spawn ──────
|
||||
// ── (an operator decision, never a backend-reported outage) — never merged with quarantine ────
|
||||
// ── or cool-off above ─────────────────────────────────────────────────────────────────────────
|
||||
|
||||
private static final BackendOutagePolicy NO_OUTAGE = new BackendOutagePolicy(() -> 0L);
|
||||
|
||||
/** Criterion 1: an explicit spawn onto a profile whose model is off is refused. */
|
||||
@Test
|
||||
void explicitSpawnOntoAnOffModelProfileIsRefused() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
PlacementException e = assertThrows(PlacementException.class,
|
||||
() -> composite.spawn(new SpawnRequest("local", null, null)));
|
||||
assertTrue(e.getMessage().contains("local"), "message names the profile: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("deepseek-v4-flash"),
|
||||
"message names the model: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("operator") && e.getMessage().contains("turned off"),
|
||||
"wording says the OPERATOR turned it off: " + e.getMessage());
|
||||
// Distinct from quarantine/cool-off wording (criterion 1's explicit requirement).
|
||||
assertFalse(e.getMessage().contains("quarantined"), "must not read like quarantine: " + e.getMessage());
|
||||
assertFalse(e.getMessage().contains("cooling off"), "must not read like cool-off: " + e.getMessage());
|
||||
assertEquals(0, adapter.spawnCount("local"), "the off-model profile is never delegated to");
|
||||
}
|
||||
|
||||
/** A profile whose model is NOT off spawns normally even while another model is off. */
|
||||
@Test
|
||||
void explicitSpawnOntoAnEnabledModelProfileSucceedsWhileAnotherModelIsOff() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("sonnet", null, null)));
|
||||
assertEquals(1, adapter.spawnCount("sonnet"));
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 8: one {@code enabled: false} entry disables EVERY profile naming that model — the
|
||||
* live example, {@code deepseek-v4-flash} backing both {@code local} and {@code local-direct}.
|
||||
*/
|
||||
@Test
|
||||
void turningOffOneModelRefusesEveryProfileThatNamesIt() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"local-direct", stubWorkerModel("local-direct", "deepseek-v4-flash"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("local", null, null)));
|
||||
assertThrows(PlacementException.class,
|
||||
() -> composite.spawn(new SpawnRequest("local-direct", null, null)),
|
||||
"local-direct shares local's model, so it must be refused too");
|
||||
assertEquals(0, adapter.spawnCount("local"));
|
||||
assertEquals(0, adapter.spawnCount("local-direct"));
|
||||
}
|
||||
|
||||
/** Criterion 3: an old-style entry with no {@code enabled} field never refuses a spawn. */
|
||||
@Test
|
||||
void anEntryWithNoEnabledFieldNeverRefusesASpawn() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("local", stubWorkerModel("local", "deepseek-v4-flash"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
// The back-compat single-arg ModelEntry constructor — no enabled field in the shape at all.
|
||||
FleetConfig.Models models = new FleetConfig.Models(
|
||||
List.of(new FleetConfig.Models.ModelEntry("deepseek-v4-flash")));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)));
|
||||
assertEquals(1, adapter.spawnCount("local"));
|
||||
}
|
||||
|
||||
/** A model absent from {@code models.allow:} entirely (nothing to gate against) is never refused. */
|
||||
@Test
|
||||
void aModelNotConfiguredInTheAllowListIsNeverGated() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("local", stubWorkerModel("local", "unlisted-model"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)));
|
||||
}
|
||||
|
||||
/** Criterion 2: an unqualified spawn skips an off-model candidate and lands on another one. */
|
||||
@Test
|
||||
void placementSkipsAnOffModelProfileAndRoutesToAnotherOne() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.weighted(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
|
||||
assertEquals("sonnet", h.profile(), "local's model is off, so an unqualified spawn must land on sonnet");
|
||||
assertEquals(0, adapter.spawnCount("local"));
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 2's second half: when EVERY candidate's model is off, the placement exception must
|
||||
* name that as the cause — not a generic "no candidates" message.
|
||||
*/
|
||||
@Test
|
||||
void automaticPlacementNamesModelOffWhenEveryCandidateIsOffModel() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"local-direct", stubWorkerModel("local-direct", "deepseek-v4-flash"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.weighted(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
PlacementException e = assertThrows(PlacementException.class,
|
||||
() -> composite.spawn(new SpawnRequest(null, null, null)),
|
||||
"both candidates share the off model — nothing is available");
|
||||
assertTrue(e.getMessage().contains("model") && e.getMessage().contains("turned off"),
|
||||
"message names model-off as the cause, not a generic no-candidates message: "
|
||||
+ e.getMessage());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up: {@code fixed} is the DEFAULT placement policy ({@code
|
||||
* PlacementPolicies.fromName} returns it for an absent/blank name), and it built its own inline
|
||||
* candidate filter instead of calling {@code PlacementPolicyUtil.available()} — so it never
|
||||
* checked {@code modelOff()}. Mirrors {@code placementSkipsAnOffModelProfileAndRoutesToAnotherOne}
|
||||
* above with only the policy swapped, to prove the gate now fires on the path most fleets use.
|
||||
*/
|
||||
@Test
|
||||
void fixedPlacementSkipsAnOffModelProfileToo() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
|
||||
assertEquals("sonnet", h.profile(), "local's model is off, so an unqualified spawn must land on sonnet");
|
||||
assertEquals(0, adapter.spawnCount("local"));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #425 rework, acceptance 2: same shape as {@link #fixedPlacementSkipsAnOffModelProfileToo}
|
||||
* above, but through {@link CompositePeerLauncher#routedProfileFor} rather than an actual
|
||||
* {@link CompositePeerLauncher#spawn} — the exact call {@code SessionManager.acquireWithWorktree}
|
||||
* makes to pre-resolve a profile for provisioning. This is the fleetd #429 case named in the
|
||||
* ticket: an operator turns a model off, and an unqualified worktree spawn must still route
|
||||
* around it instead of throwing "names model, which the operator has turned off" — the throw
|
||||
* {@link CompositePeerLauncher#enforceModelEnabled} raises only on the EXPLICIT-profile branch,
|
||||
* which is exactly the branch the regressed round accidentally routed every worktree spawn onto.
|
||||
*/
|
||||
@Test
|
||||
void routedProfileForSkipsAModelOffPoolFirstProfileUnderFixedPolicy() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
assertEquals("sonnet", composite.routedProfileFor(MemberRole.DEV),
|
||||
"local (the pool's first entry) names an off model, so the routed answer must be sonnet");
|
||||
assertEquals("local", composite.defaultProfileFor(MemberRole.DEV),
|
||||
"sanity: defaultProfileFor stays blind to model-off — that's the gap routedProfileFor closes");
|
||||
assertEquals(0, adapter.spawnCount("local"), "routedProfileFor never spawns anything");
|
||||
assertEquals(0, adapter.spawnCount("sonnet"), "routedProfileFor never spawns anything");
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 4: turning a model off/on is HOT — no restart — proven through a REAL
|
||||
* {@code ConfigRef.reload()}, not a hand-rolled supplier swap. Also proves {@code models} is
|
||||
* correctly reclassified: the reload's {@link ConfigRef.Outcome#applied()} is {@code true} and
|
||||
* {@code "models"} never appears in {@link ConfigRef.Outcome#deferred()}.
|
||||
*/
|
||||
@Test
|
||||
void modelOnOffIsHotReloadedThroughARealConfigRef(@TempDir Path dir) throws Exception {
|
||||
Path yaml = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
models:
|
||||
allow:
|
||||
- model: deepseek-v4-flash
|
||||
""");
|
||||
FleetConfig initial = FleetConfig.load(yaml);
|
||||
ConfigRef configRef = new ConfigRef(yaml, initial);
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr,
|
||||
Map.of("local", stubWorker("local")), "local", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "local", configRef, _ -> 0, BackendQuarantine.none(), NO_OUTAGE);
|
||||
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)),
|
||||
"the model starts enabled");
|
||||
assertEquals(1, adapter.spawnCount("local"));
|
||||
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
models:
|
||||
allow:
|
||||
- model: deepseek-v4-flash
|
||||
enabled: false
|
||||
""");
|
||||
ConfigRef.Outcome outcome = configRef.reload();
|
||||
assertTrue(outcome.applied(), "a models.allow on/off edit must apply live, never be refused");
|
||||
assertFalse(outcome.deferred().contains("models"),
|
||||
"models is hot-excluded now — it must never be reported as a deferred key");
|
||||
|
||||
PlacementException e = assertThrows(PlacementException.class,
|
||||
() -> composite.spawn(new SpawnRequest("local", null, null)),
|
||||
"the very next spawn must see the reload, with no restart");
|
||||
assertTrue(e.getMessage().contains("deepseek-v4-flash"));
|
||||
assertEquals(1, adapter.spawnCount("local"), "still just the one successful spawn from before");
|
||||
|
||||
// And back on, still hot, still no restart.
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
models:
|
||||
allow:
|
||||
- model: deepseek-v4-flash
|
||||
enabled: true
|
||||
""");
|
||||
ConfigRef.Outcome reEnabled = configRef.reload();
|
||||
assertTrue(reEnabled.applied());
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)));
|
||||
assertEquals(2, adapter.spawnCount("local"));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up, acceptance criterion 2: {@link CompositePeerLauncher#modelGateState()}
|
||||
* is LIVE — no restart — proven through a REAL {@link ConfigRef#reload()}, exactly like {@link
|
||||
* #modelOnOffIsHotReloadedThroughARealConfigRef} above proves for the on/off gate itself. This
|
||||
* single reload sequence walks through all three states the ticket asks for: no {@code models:}
|
||||
* block, a block armed with nothing off, and a block with one model off — so a reload that flips
|
||||
* between any of the three is proven live, not just the on/off edit within an already-armed block.
|
||||
*/
|
||||
@Test
|
||||
void modelGateStateIsHotReloadedThroughARealConfigRef(@TempDir Path dir) throws Exception {
|
||||
Path yaml = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
""");
|
||||
FleetConfig initial = FleetConfig.load(yaml);
|
||||
ConfigRef configRef = new ConfigRef(yaml, initial);
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr,
|
||||
Map.of("local", stubWorker("local")), "local", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "local", configRef, _ -> 0, BackendQuarantine.none(), NO_OUTAGE);
|
||||
|
||||
PeerLauncher.ModelGateState notConfigured = composite.modelGateState();
|
||||
assertFalse(notConfigured.configured(), "no models: block in the config at all");
|
||||
assertEquals(Set.of(), notConfigured.off());
|
||||
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
models:
|
||||
allow:
|
||||
- model: deepseek-v4-flash
|
||||
enabled: false
|
||||
""");
|
||||
assertTrue(configRef.reload().applied(), "adding a models: block must apply live, no restart");
|
||||
PeerLauncher.ModelGateState armedWithOneOff = composite.modelGateState();
|
||||
assertTrue(armedWithOneOff.configured(), "a models: block now exists — the gate is armed");
|
||||
assertEquals(Set.of("deepseek-v4-flash"), armedWithOneOff.off());
|
||||
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
models:
|
||||
allow:
|
||||
- model: deepseek-v4-flash
|
||||
enabled: true
|
||||
""");
|
||||
assertTrue(configRef.reload().applied(), "flipping the entry back on must apply live too");
|
||||
PeerLauncher.ModelGateState armedWithZeroOff = composite.modelGateState();
|
||||
assertTrue(armedWithZeroOff.configured(),
|
||||
"the block is still present — armed and reporting zero, not the same as no block at all");
|
||||
assertEquals(Set.of(), armedWithZeroOff.off());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up, acceptance criterion 3: the invariant is that an absent {@code
|
||||
* models:} block stays permitted and must never be fatal. Proved, not assumed — a config
|
||||
* without one loads, validates, reports the gate as not configured, AND still spawns normally
|
||||
* (no {@link PlacementException} from a gate that has nothing to check against), using the same
|
||||
* production-shaped {@code Supplier<FleetConfig>} wiring {@code Fleetd.main} actually uses.
|
||||
*/
|
||||
@Test
|
||||
void noModelsBlockConfigStillLoadsAndSpawnsNormally(@TempDir Path dir) throws Exception {
|
||||
Path yaml = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(yaml);
|
||||
assertDoesNotThrow(cfg::validateAll, "a config with no models: block must load and validate cleanly");
|
||||
ConfigRef configRef = new ConfigRef(yaml, cfg);
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr,
|
||||
Map.of("local", stubWorker("local")), "local", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "local", configRef, _ -> 0, BackendQuarantine.none(), NO_OUTAGE);
|
||||
|
||||
assertFalse(composite.modelGateState().configured());
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)),
|
||||
"no models: block means nothing to gate against — the spawn must go through");
|
||||
assertEquals(1, adapter.spawnCount("local"));
|
||||
}
|
||||
|
||||
/** {@code fleet_profiles}/{@code GET /profiles} must read the exact same live source the gate reads. */
|
||||
@Test
|
||||
void disabledModelsReportsWhatTheGateActuallyEnforces() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false),
|
||||
new FleetConfig.Models.ModelEntry("claude-sonnet-5", true)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
assertEquals(Set.of("deepseek-v4-flash"), composite.disabledModels());
|
||||
}
|
||||
|
||||
/** A caller with no {@code models:} block (the map-based constructors) reports nothing off. */
|
||||
@Test
|
||||
void disabledModelsIsEmptyWithNoModelsConfigured() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
PeerLauncher composite = composite(herdr);
|
||||
assertEquals(Set.of(), composite.disabledModels());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -18,6 +18,7 @@ import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
@@ -118,6 +119,205 @@ class EnvAllowListScrubTest {
|
||||
"allowed N of M with N <= M — the denominator is always reported");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #394: the actual defect. Plain {@code export "$n="} is FATAL for a zsh read-only or
|
||||
* special parameter (e.g. {@code UID}) and aborts the whole sourced file — every name still to
|
||||
* come is never blanked, and the {@code scrub-report.txt} below is never written at all,
|
||||
* silently ({@code 2>/dev/null} swallows the error). This plants an unblankable, exported,
|
||||
* read-only variable in the MIDDLE of the names the scrub attempts to blank, with two more
|
||||
* names after it, and asserts that both of those later names are STILL blanked and the report
|
||||
* is STILL written with the failure counted — a test that only checked names BEFORE the failure
|
||||
* point would pass today and prove nothing.
|
||||
*
|
||||
* <p>The planted name is a made-up one ({@code FLEETD_TEST_UNBLANKABLE}), not {@code UID} or
|
||||
* any other name a skip-list might already know about — invariant 1 is that the loop survives
|
||||
* ANY unblankable name, not a known one, so the test must not lean on one either.
|
||||
*
|
||||
* <p>Exercises the real artefact: {@link EnvAllowListScrub#scrubScript} is run verbatim under a
|
||||
* real {@code /bin/zsh}, not just asserted on as a Java string. The four planted names are
|
||||
* exported one at a time via {@code typeset -x}/{@code typeset -rx} immediately before the
|
||||
* script runs, in a fixed order — zsh's {@code export}/{@code typeset -x} appends to the
|
||||
* process's environment table in call order (verified empirically: a freshly-exported name
|
||||
* always sorts after every inherited one and after every earlier freshly-exported name in
|
||||
* {@code command env}'s own output), which is what makes the "middle" position deterministic
|
||||
* here, unlike relying on the OS's own inherited-environment order.
|
||||
*/
|
||||
@Test
|
||||
void unblankableNameInTheMiddleDoesNotAbortNamesAfterIt(@TempDir Path tmp) throws Exception {
|
||||
assumeTrue(Files.isExecutable(ZSH), "/bin/zsh not present — nothing to prove here");
|
||||
|
||||
// Only ZDOTDIR is allowed — it must survive the scrub itself, since the report is written
|
||||
// to "$ZDOTDIR/..." AFTER the blanking loop runs; if ZDOTDIR were blanked as a side effect,
|
||||
// the report write would silently go to the wrong place instead of testing anything.
|
||||
String script = EnvAllowListScrub.scrubScript(Set.of("ZDOTDIR"));
|
||||
String setup = """
|
||||
typeset -x FLEETD_TEST_BEFORE=1
|
||||
typeset -rx FLEETD_TEST_UNBLANKABLE=1
|
||||
typeset -x FLEETD_TEST_AFTER_A=1
|
||||
typeset -x FLEETD_TEST_AFTER_B=1
|
||||
""";
|
||||
|
||||
ProcessBuilder pb = new ProcessBuilder("/bin/zsh");
|
||||
pb.environment().clear();
|
||||
pb.environment().put("PATH", "/usr/bin:/bin");
|
||||
pb.environment().put("ZDOTDIR", tmp.toAbsolutePath().toString());
|
||||
pb.redirectError(ProcessBuilder.Redirect.DISCARD);
|
||||
Process zsh = pb.start();
|
||||
zsh.getOutputStream().write((setup + script).getBytes(StandardCharsets.UTF_8));
|
||||
zsh.getOutputStream().flush();
|
||||
zsh.getOutputStream().close();
|
||||
assertTrue(zsh.waitFor(60, java.util.concurrent.TimeUnit.SECONDS),
|
||||
"the scrub script did not exit within 60s");
|
||||
assertEquals(0, zsh.exitValue(),
|
||||
"the scrub script itself must never abort — an unblankable name must not kill the "
|
||||
+ "sourced file");
|
||||
|
||||
EnvAllowListScrub.ScrubReport report = EnvAllowListScrub.readReport(tmp);
|
||||
assertNotNull(report, "the report must still be written even though one name could not be "
|
||||
+ "blanked — a report that silently never appears is the #394 bug");
|
||||
assertTrue(report.blanked().contains("FLEETD_TEST_BEFORE"),
|
||||
"sanity: the name before the unblankable one must be blanked");
|
||||
assertTrue(report.blanked().contains("FLEETD_TEST_AFTER_A"),
|
||||
"the FIRST name AFTER the unblankable one must still be blanked — before the fix, "
|
||||
+ "the whole loop aborted at the unblankable name and every later name was "
|
||||
+ "silently left untouched");
|
||||
assertTrue(report.blanked().contains("FLEETD_TEST_AFTER_B"),
|
||||
"the SECOND name after the unblankable one must also still be blanked");
|
||||
assertTrue(report.unblankable().contains("FLEETD_TEST_UNBLANKABLE"),
|
||||
"the unblankable name is reported by name, not silently dropped");
|
||||
// fleetd #400 note: zsh itself auto-exports SHLVL on every shell start (measured: it appears
|
||||
// in `command env` even from a fully cleared parent), and a bare assignment to it is coerced
|
||||
// rather than failing — exactly the shape #400 fixes. Before that fix, eval's exit status
|
||||
// alone silently misclassified SHLVL as blanked, so this test's old "exactly one" assertion
|
||||
// passed by accident: it never actually proved SHLVL was absent from the candidates, only
|
||||
// that the old bug hid it. Now that classification reads the value back, SHLVL and
|
||||
// FLEETD_TEST_UNBLANKABLE both correctly land in unblankable() — real, unplanted evidence
|
||||
// the #400 fix works, not just the synthetic case in the dedicated #400 test above.
|
||||
assertTrue(report.unblankable().contains("SHLVL"),
|
||||
"fleetd #400: zsh's own auto-exported SHLVL must also be reported unblankable, not "
|
||||
+ "silently miscounted as blanked");
|
||||
assertEquals(report.unblankable().size(), report.failed(),
|
||||
"the failed count must equal the number of names actually reported unblankable");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #400: {@code eval}'s exit status is not proof that a name was actually blanked. zsh
|
||||
* coerces a bare {@code NAME=} assignment on an integer special parameter to a number instead of
|
||||
* failing, so {@code eval} reports success while the value stays non-empty — a status-based
|
||||
* classification calls that "blanked" when it was not. This drives all three shapes a name can
|
||||
* take through the real {@code scrubScript} in ONE run: a normal, genuinely blankable name; a
|
||||
* fatal one ({@code LINENO} — deliberately not {@code UID}, so this test does not depend on the
|
||||
* harness's uid); and the silent-no-op one the ticket is about ({@code SECONDS}, rc 0 but
|
||||
* unchanged). Under the pre-#400 exit-status check, {@code SECONDS} would land in
|
||||
* {@code blanked()} — that is the exact false receipt this fix removes.
|
||||
*
|
||||
* <p>Criterion 3: a cleared {@code ProcessBuilder} parent does not, by itself, give the child
|
||||
* zsh any of these names — {@code SECONDS}/{@code LINENO} are zsh's own built-in parameters and
|
||||
* only become CANDIDATES the enumeration loop can see (i.e. show up in {@code command env}) when
|
||||
* they arrive via the process's own environment table, not merely by existing as zsh parameters
|
||||
* inside the shell. So each is put into {@code pb.environment()} explicitly, after
|
||||
* {@code clear()} — confirmed empirically first (a throwaway probe piping
|
||||
* {@code env -i PATH=... SECONDS=999 LINENO=999 FLEETD_TEST_NORMAL=1 zsh -c 'command env | cut
|
||||
* -d= -f1'}) that all three names really appear in {@code command env}'s output under exactly
|
||||
* this construction, not relying on whatever the test-runner's own ambient environment happens
|
||||
* to contain.
|
||||
*/
|
||||
@Test
|
||||
void classifiesByObservedValueNotExitStatusAcrossAllThreeShapes(@TempDir Path tmp) throws Exception {
|
||||
assumeTrue(Files.isExecutable(ZSH), "/bin/zsh not present — nothing to prove here");
|
||||
|
||||
String script = EnvAllowListScrub.scrubScript(Set.of("ZDOTDIR"));
|
||||
|
||||
ProcessBuilder pb = new ProcessBuilder("/bin/zsh");
|
||||
pb.environment().clear();
|
||||
pb.environment().put("PATH", "/usr/bin:/bin");
|
||||
pb.environment().put("ZDOTDIR", tmp.toAbsolutePath().toString());
|
||||
// Explicitly placed in the child's environment table — see the javadoc above on why a
|
||||
// cleared parent alone does not put these on the enumeration loop's candidate list.
|
||||
pb.environment().put("SECONDS", "999"); // rc 0, value coerced/unchanged — the #400 bug
|
||||
pb.environment().put("LINENO", "999"); // fatal on assignment, eval rc != 0, contained
|
||||
pb.environment().put("FLEETD_TEST_NORMAL", "1"); // genuinely blankable, the control case
|
||||
pb.redirectError(ProcessBuilder.Redirect.DISCARD);
|
||||
Process zsh = pb.start();
|
||||
zsh.getOutputStream().write(script.getBytes(StandardCharsets.UTF_8));
|
||||
zsh.getOutputStream().flush();
|
||||
zsh.getOutputStream().close();
|
||||
assertTrue(zsh.waitFor(60, java.util.concurrent.TimeUnit.SECONDS),
|
||||
"the scrub script did not exit within 60s");
|
||||
assertEquals(0, zsh.exitValue(),
|
||||
"the scrub script must still reach its end with both a fatal name and a silent "
|
||||
+ "no-op name among the candidates");
|
||||
|
||||
EnvAllowListScrub.ScrubReport report = EnvAllowListScrub.readReport(tmp);
|
||||
assertNotNull(report, "the report must still be written");
|
||||
assertTrue(report.blanked().contains("FLEETD_TEST_NORMAL"),
|
||||
"the control case: an ordinary name is genuinely blankable and must be reported so");
|
||||
assertTrue(report.unblankable().contains("LINENO"),
|
||||
"a name fatal to assign to must be reported unblankable — sanity check that "
|
||||
+ "containment still works under the new classification");
|
||||
assertTrue(report.unblankable().contains("SECONDS"),
|
||||
"the #400 defect: eval returns rc 0 for SECONDS (zsh coerces the assignment instead "
|
||||
+ "of failing) but the value is left non-empty — classifying on the observed "
|
||||
+ "value catches this; classifying on eval's exit status would have called "
|
||||
+ "this \"blanked\" and produced a false receipt");
|
||||
assertFalse(report.blanked().contains("SECONDS"),
|
||||
"SECONDS must never appear as blanked — it was never actually emptied");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #394 follow-up: the blanking loop's {@code eval "export ${n}="} splices {@code n} into
|
||||
* a string that zsh then interprets as shell syntax. That is only safe because every name
|
||||
* reaching {@code _cb633_blank} already passed an identifier check in the ENUMERATION loop
|
||||
* (20 lines away, in a different loop) — so the fix re-asserts the identical check immediately
|
||||
* before the {@code eval} call, rather than trusting that distant guard to keep holding.
|
||||
*
|
||||
* <p>This test plants a value with an embedded newline, exploiting the exact "junk from
|
||||
* multi-line values" gap the enumeration loop's own comment already documents: {@code command
|
||||
* env}'s text output is read line-by-line, so a value's second line becomes a spurious extra
|
||||
* "name" that was never a real exported variable. The fragment used here ({@code
|
||||
* junk.fragment}) is merely non-conforming (it contains a dot) — never command-shaped; this
|
||||
* test must never demonstrate command execution and plants no command-shaped payload.
|
||||
*
|
||||
* <p>Exercises the real artefact end-to-end: {@link EnvAllowListScrub#scrubScript} runs
|
||||
* verbatim under a real {@code /bin/zsh}, exactly as {@code generate()} would produce it — this
|
||||
* is not a synthetic call into just the blanking loop.
|
||||
*/
|
||||
@Test
|
||||
void nonIdentifierJunkFromAMultilineValueIsSkippedNotBlankedOrUnblankable(@TempDir Path tmp)
|
||||
throws Exception {
|
||||
assumeTrue(Files.isExecutable(ZSH), "/bin/zsh not present — nothing to prove here");
|
||||
|
||||
String script = EnvAllowListScrub.scrubScript(Set.of("ZDOTDIR"));
|
||||
|
||||
ProcessBuilder pb = new ProcessBuilder("/bin/zsh");
|
||||
pb.environment().clear();
|
||||
pb.environment().put("PATH", "/usr/bin:/bin");
|
||||
pb.environment().put("ZDOTDIR", tmp.toAbsolutePath().toString());
|
||||
// Embedded newline: `command env`'s own text output splits this into two lines, and the
|
||||
// second ("junk.fragment") has no "=" at all, so `cut -d= -f1` returns it unchanged as a
|
||||
// spurious candidate "name" — it was never an actual exported variable by that name.
|
||||
pb.environment().put("FLEETD_TEST_MULTILINE", "keep\njunk.fragment");
|
||||
pb.redirectError(ProcessBuilder.Redirect.DISCARD);
|
||||
Process zsh = pb.start();
|
||||
zsh.getOutputStream().write(script.getBytes(StandardCharsets.UTF_8));
|
||||
zsh.getOutputStream().flush();
|
||||
zsh.getOutputStream().close();
|
||||
assertTrue(zsh.waitFor(60, java.util.concurrent.TimeUnit.SECONDS),
|
||||
"the scrub script did not exit within 60s");
|
||||
assertEquals(0, zsh.exitValue(), "the scrub script must reach its end");
|
||||
|
||||
EnvAllowListScrub.ScrubReport report = EnvAllowListScrub.readReport(tmp);
|
||||
assertNotNull(report, "the report must still be written");
|
||||
assertTrue(report.blanked().contains("FLEETD_TEST_MULTILINE"),
|
||||
"sanity: the real, identifier-shaped variable must still be blanked normally");
|
||||
assertFalse(report.blanked().contains("junk.fragment"),
|
||||
"a non-identifier fragment is not a real variable and must never be blanked");
|
||||
assertFalse(report.unblankable().contains("junk.fragment"),
|
||||
"a non-identifier fragment must never even become a candidate the blanking loop "
|
||||
+ "attempts — it must be filtered before either guard has to catch it, so "
|
||||
+ "it is neither blanked nor counted as a failed attempt");
|
||||
}
|
||||
|
||||
/** A group-shared ZDOTDIR still lets the member truncate and write its pre-created receipt. */
|
||||
@Test
|
||||
void groupSharedScrubWritesAndReadsItsReport(@TempDir Path tmp) throws Exception {
|
||||
|
||||
@@ -762,6 +762,250 @@ class OpenCodeLauncherTest {
|
||||
"no IDE server when ideMcpUrl is unset");
|
||||
}
|
||||
|
||||
// --- fleetd #393: memberSkills seeding must actually reach an opencode member -------------------
|
||||
//
|
||||
// Before this fix, GitWorktrees#seedSkills copied skill folders into EVERY provisioned
|
||||
// worktree's .claude/skills/ and logged "skill seeding: N of M" regardless of which kind ended
|
||||
// up spawning into that worktree — a claim that held for kind: claude-code (the CLI discovers
|
||||
// that directory on its own) but was a guaranteed no-op for kind: opencode, which has no such
|
||||
// discovery. These tests drive the REAL GitWorktrees#add seeding path (not a hand-built
|
||||
// .claude/skills/ fixture), then spawn an opencode-kind member against the seeded worktree and
|
||||
// assert on what the member can actually consume — an instructions[] entry — not on the
|
||||
// seeding log alone. A minimal, non-hermetic git repo is enough here: unlike
|
||||
// GitWorktreesTest's own seeding tests, nothing in this file cares about core.excludesFile
|
||||
// composition, only about what lands in .claude/skills/ and whether OpenCodeLauncher reads it.
|
||||
|
||||
private static void git(Path cwd, String... args) throws Exception {
|
||||
Process p = new ProcessBuilder(prepend("git", args)).directory(cwd.toFile())
|
||||
.redirectErrorStream(true).start();
|
||||
String out = new String(p.getInputStream().readAllBytes());
|
||||
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git timed out: git " + String.join(" ", args));
|
||||
assertEquals(0, p.exitValue(), "git " + String.join(" ", args) + " failed:\n" + out);
|
||||
}
|
||||
|
||||
private static List<String> prepend(String head, String... rest) {
|
||||
List<String> cmd = new ArrayList<>();
|
||||
cmd.add(head);
|
||||
cmd.addAll(List.of(rest));
|
||||
return cmd;
|
||||
}
|
||||
|
||||
private static Path initRepo(Path dir) throws Exception {
|
||||
Files.createDirectories(dir);
|
||||
git(dir, "init", "-q", "-b", "main");
|
||||
git(dir, "config", "user.email", "test@example.invalid");
|
||||
git(dir, "config", "user.name", "Test");
|
||||
Files.writeString(dir.resolve("README.md"), "seed\n");
|
||||
git(dir, "add", "README.md");
|
||||
git(dir, "commit", "-q", "-m", "seed");
|
||||
return dir;
|
||||
}
|
||||
|
||||
@Test
|
||||
void aSeededSkillReachesTheOpencodeMembersInstructionsArray(@TempDir Path tmp) throws Exception {
|
||||
Path repo = initRepo(tmp.resolve("repo"));
|
||||
Path skillsSource = tmp.resolve("skills-src");
|
||||
Path skillFile = skillsSource.resolve("implementer").resolve("SKILL.md");
|
||||
Files.createDirectories(skillFile.getParent());
|
||||
Files.writeString(skillFile, "IMPLEMENTER PROCEDURE\n");
|
||||
|
||||
// The real seeding path (fleetd #362), not a hand-built .claude/skills/ fixture — proves
|
||||
// OpenCodeLauncher reads what GitWorktrees#add actually produced.
|
||||
dev.ltms.fleet.session.GitWorktrees worktrees =
|
||||
new dev.ltms.fleet.session.GitWorktrees(tmp.resolve("wts").toString(), null, skillsSource.toString());
|
||||
String wt = worktrees.add(repo.toString(), "cb-393-opencode", "HEAD");
|
||||
Path seededSkillMd = Path.of(wt, ".claude", "skills", "implementer", "SKILL.md");
|
||||
assertTrue(Files.exists(seededSkillMd),
|
||||
"sanity: the real seeding step must have copied the skill into the worktree");
|
||||
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(OpenCodeLauncher.class);
|
||||
Level original = logger.getLevel();
|
||||
logger.setLevel(Level.INFO);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
|
||||
String cfgPath;
|
||||
try {
|
||||
// No mcpUrl, no ideUrl, no fleet (no charter): the seeded skill alone must be enough to
|
||||
// trigger OPENCODE_CONFIG — proves the gate itself was updated, not only writeConfig's body.
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path configRoot = Files.createDirectory(tmp.resolve("configs"));
|
||||
service(herdr, configRoot, opencodeIdeCfg(null, null, wt)).spawn();
|
||||
cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
|
||||
} finally {
|
||||
logger.detachAppender(appender);
|
||||
logger.setLevel(original);
|
||||
}
|
||||
|
||||
assertNotNull(cfgPath, "a seeded skill with nothing else configured must still write a config");
|
||||
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
|
||||
List<String> instructions = new ArrayList<>();
|
||||
json.path("instructions").forEach(n -> instructions.add(n.asText()));
|
||||
assertTrue(instructions.contains(seededSkillMd.toAbsolutePath().toString()),
|
||||
"the seeded skill's SKILL.md must be an instructions[] entry — got: " + instructions);
|
||||
|
||||
List<String> infos = appender.list.stream()
|
||||
.filter(e -> e.getLevel() == Level.INFO)
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.toList();
|
||||
assertTrue(infos.stream().anyMatch(m -> m.contains("skill delivery") && m.contains("1 of 1")),
|
||||
"the launcher must log, kind-aware, that it delivered the skill — got:\n" + infos);
|
||||
}
|
||||
|
||||
@Test
|
||||
void aSkillFolderWithoutSkillMdIsNeverDeliveredAndTheLogNamesIt(@TempDir Path tmp) throws Exception {
|
||||
Path wt = Files.createDirectories(tmp.resolve("wt"));
|
||||
Path goodSkill = wt.resolve(".claude").resolve("skills").resolve("implementer");
|
||||
Files.createDirectories(goodSkill);
|
||||
Files.writeString(goodSkill.resolve("SKILL.md"), "GOOD\n");
|
||||
Path halfShipped = wt.resolve(".claude").resolve("skills").resolve("half-shipped");
|
||||
Files.createDirectories(halfShipped);
|
||||
Files.writeString(halfShipped.resolve("README.md"), "no SKILL.md here\n");
|
||||
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(OpenCodeLauncher.class);
|
||||
Level original = logger.getLevel();
|
||||
logger.setLevel(Level.INFO);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
|
||||
String cfgPath;
|
||||
try {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path configRoot = Files.createDirectory(tmp.resolve("configs"));
|
||||
service(herdr, configRoot, opencodeIdeCfg(null, null, wt.toString())).spawn();
|
||||
cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
|
||||
} finally {
|
||||
logger.detachAppender(appender);
|
||||
logger.setLevel(original);
|
||||
}
|
||||
|
||||
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
|
||||
List<String> instructions = new ArrayList<>();
|
||||
json.path("instructions").forEach(n -> instructions.add(n.asText()));
|
||||
assertTrue(instructions.contains(goodSkill.resolve("SKILL.md").toAbsolutePath().toString()),
|
||||
"the well-formed skill is still delivered alongside the malformed one");
|
||||
assertFalse(instructions.stream().anyMatch(i -> i.contains("half-shipped")),
|
||||
"a skill folder with no SKILL.md can never become an instructions[] entry");
|
||||
|
||||
List<String> infos = appender.list.stream()
|
||||
.filter(e -> e.getLevel() == Level.INFO)
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.toList();
|
||||
assertTrue(infos.stream().anyMatch(m -> m.contains("skill delivery") && m.contains("1 of 2")
|
||||
&& m.contains("half-shipped") && m.contains("could not be delivered")),
|
||||
"the log must say plainly which folder could not be consumed and why — got:\n" + infos);
|
||||
}
|
||||
|
||||
// --- fleetd #393 follow-up: instructions[] has three writers (charter, seeded skills, IDE
|
||||
// rules), and no test above ever exercises more than one or two of them together. A writer
|
||||
// that flips from withArray (get-or-create) to putArray (create-or-REPLACE) silently deletes
|
||||
// every entry written before it — proven live on this branch's merge: switching just the
|
||||
// skills writer to putArray left the entire suite (1603 tests) green while deleting the
|
||||
// charter entry an opencode member needs for its role contract. That hazard was found by the
|
||||
// fleet01 lead and independently verified against this branch; it is not a defect in the
|
||||
// skills-delivery or logging tests above, which both hold up under their own mutations — the
|
||||
// gap is that none of them combine all three writers in one config.
|
||||
//
|
||||
// These three tests assert instructions[] CONTENT as an exact, ordered list, not a size or a
|
||||
// "contains" check: a putArray mutation can replace N entries with a different N entries of
|
||||
// the same count, so only a content comparison can tell "all three paths present" apart from
|
||||
// "two paths present that replaced the earlier ones".
|
||||
|
||||
@Test
|
||||
void instructionsArrayHoldsExactlyTheCharterWhenNothingElseWritesToIt(@TempDir Path root) throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Fleet fleet = new FleetConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
|
||||
Map.of("dev", "role rule"), null);
|
||||
Path cwd = Files.createDirectory(root.resolve("checkout"));
|
||||
service(herdr, root, opencodeIdeCfg(null, null, cwd.toString()), () -> fleet).spawn();
|
||||
|
||||
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
|
||||
assertNotNull(cfgPath, "a role charter alone still writes a config");
|
||||
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
|
||||
Path charter = Path.of(cfgPath).resolveSibling("member-charter.md");
|
||||
assertTrue(Files.exists(charter), "the charter file was written");
|
||||
|
||||
List<String> instructions = new ArrayList<>();
|
||||
json.path("instructions").forEach(n -> instructions.add(n.asText()));
|
||||
assertEquals(List.of(charter.toAbsolutePath().toString()), instructions,
|
||||
"with only the charter writer active, instructions[] holds exactly one entry: the "
|
||||
+ "charter — got: " + instructions);
|
||||
}
|
||||
|
||||
@Test
|
||||
void instructionsArrayHoldsCharterThenIdeRulesInOrder(@TempDir Path root) throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Fleet fleet = new FleetConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
|
||||
Map.of("dev", "role rule"), null);
|
||||
Path cwd = Files.createDirectory(root.resolve("checkout"));
|
||||
service(herdr, root, opencodeIdeCfg(null,
|
||||
"http://127.0.0.1:29170/index-mcp/streamable-http", cwd.toString()), () -> fleet).spawn();
|
||||
|
||||
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
|
||||
assertNotNull(cfgPath, "charter + IDE rules still writes a config");
|
||||
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
|
||||
Path charter = Path.of(cfgPath).resolveSibling("member-charter.md");
|
||||
Path rules = Path.of(cfgPath).resolveSibling("ide-rules.md");
|
||||
assertTrue(Files.exists(charter), "the charter file was written");
|
||||
assertTrue(Files.exists(rules), "the ide-rules file was written");
|
||||
|
||||
List<String> instructions = new ArrayList<>();
|
||||
json.path("instructions").forEach(n -> instructions.add(n.asText()));
|
||||
assertEquals(List.of(charter.toAbsolutePath().toString(), rules.toAbsolutePath().toString()),
|
||||
instructions,
|
||||
"with charter + IDE-rules writers active, instructions[] holds both, charter first — "
|
||||
+ "got: " + instructions);
|
||||
}
|
||||
|
||||
@Test
|
||||
void instructionsArrayHoldsCharterThenSkillsThenIdeRulesInOrder(@TempDir Path tmp) throws Exception {
|
||||
Path repo = initRepo(tmp.resolve("repo"));
|
||||
Path skillsSource = tmp.resolve("skills-src");
|
||||
Path skillFile = skillsSource.resolve("implementer").resolve("SKILL.md");
|
||||
Files.createDirectories(skillFile.getParent());
|
||||
Files.writeString(skillFile, "IMPLEMENTER PROCEDURE\n");
|
||||
|
||||
dev.ltms.fleet.session.GitWorktrees worktrees = new dev.ltms.fleet.session.GitWorktrees(
|
||||
tmp.resolve("wts").toString(), null, skillsSource.toString());
|
||||
String wt = worktrees.add(repo.toString(), "cb-393-follow-up", "HEAD");
|
||||
Path seededSkillMd = Path.of(wt, ".claude", "skills", "implementer", "SKILL.md");
|
||||
assertTrue(Files.exists(seededSkillMd), "sanity: the real seeding step copied the skill");
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Fleet fleet = new FleetConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
|
||||
Map.of("dev", "role rule"), null);
|
||||
Path configRoot = Files.createDirectory(tmp.resolve("configs"));
|
||||
service(herdr, configRoot, opencodeIdeCfg(null,
|
||||
"http://127.0.0.1:29170/index-mcp/streamable-http", wt), () -> fleet).spawn();
|
||||
|
||||
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
|
||||
assertNotNull(cfgPath, "charter + skills + IDE rules still writes a config");
|
||||
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
|
||||
Path charter = Path.of(cfgPath).resolveSibling("member-charter.md");
|
||||
Path rules = Path.of(cfgPath).resolveSibling("ide-rules.md");
|
||||
assertTrue(Files.exists(charter), "the charter file was written");
|
||||
assertTrue(Files.exists(rules), "the ide-rules file was written");
|
||||
|
||||
List<String> instructions = new ArrayList<>();
|
||||
json.path("instructions").forEach(n -> instructions.add(n.asText()));
|
||||
// Exact ordered list, not size or "contains": a putArray mutation on any writer after the
|
||||
// charter replaces every entry written before it, and the replacement can still be a
|
||||
// plausible-looking array of a different shape. This is the one combination all three
|
||||
// writers are active for — and per fleet01 the realistic shape on a host where weighted
|
||||
// placement makes opencode the default for most members.
|
||||
assertEquals(List.of(charter.toAbsolutePath().toString(),
|
||||
seededSkillMd.toAbsolutePath().toString(),
|
||||
rules.toAbsolutePath().toString()),
|
||||
instructions,
|
||||
"with all three writers active, instructions[] must hold charter, then the seeded "
|
||||
+ "skill, then IDE rules — in that order and with nothing replaced. A "
|
||||
+ "putArray mutation on any writer after the charter would silently drop "
|
||||
+ "earlier entries here while still producing a same-shaped array — got: "
|
||||
+ instructions);
|
||||
}
|
||||
|
||||
// --- fleetd #219: config root + discovery root under memberHerdrSocket ------------------------
|
||||
|
||||
/** A config with {@code memberHerdrSocket:} set, and optionally {@code worktreeRoot:}/{@code worktreeGroup:}. */
|
||||
|
||||
@@ -16,6 +16,7 @@ import java.util.List;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
@@ -221,6 +222,31 @@ class AmqpReplyInboxContractTest {
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void ackReportsHitVsMissAgainstARealBroker() throws Exception {
|
||||
// fleetd #437: fleet_ack said "acknowledged <msgId>" for a message it never touched,
|
||||
// because ReplyInbox.ack() (void) could not tell a hit from a miss. Pin the fixed
|
||||
// boolean contract against a real broker — the adapter fleetd actually runs live.
|
||||
String target = "worker-ack-contract-" + System.nanoTime();
|
||||
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
|
||||
inbox.own(target);
|
||||
|
||||
// Never held for this target at all: must report false, not throw.
|
||||
assertFalse(inbox.ack(target, "never-held"),
|
||||
"acking a msgId never held for an owned target must report false");
|
||||
|
||||
// A real message: first ack removes it and reports true...
|
||||
inbox.publish(target, "m1", "ack me");
|
||||
assertEquals(1, awaitPeek(inbox, target).size(), "the published reply should be held");
|
||||
assertTrue(inbox.ack(target, "m1"), "acking a held reply must report true");
|
||||
assertTrue(inbox.peek(target).isEmpty(), "an acked reply is dropped");
|
||||
|
||||
// ...and the second ack of the SAME msgId has nothing left to remove: false.
|
||||
assertFalse(inbox.ack(target, "m1"),
|
||||
"acking the same msgId twice must report false the second time");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void confirmedPublishDeliversNormally() throws Exception {
|
||||
String target = "worker-confirm-" + System.nanoTime();
|
||||
|
||||
@@ -28,11 +28,19 @@ public final class FakeLeadChannel implements LeadChannel {
|
||||
private volatile IllegalStateException publishFailure;
|
||||
/** Canned {@link #inspect} results by coord-id — absent for any coord-id not configured here. */
|
||||
private final Map<String, MailboxState> mailboxes = new ConcurrentHashMap<>();
|
||||
/** fleetd #440: matches {@link LeadMailbox}'s real default (durable queue + manual ack) unless overridden. */
|
||||
private volatile boolean heldDurable = true;
|
||||
|
||||
public FakeLeadChannel(String selfCoordId) {
|
||||
this.selfCoordId = selfCoordId;
|
||||
}
|
||||
|
||||
/** Make {@link #heldDurable()} report {@code durable} — the fleetd #440 seam for the false case. */
|
||||
public FakeLeadChannel withHeldDurable(boolean durable) {
|
||||
this.heldDurable = durable;
|
||||
return this;
|
||||
}
|
||||
|
||||
/** Make {@link #inspect(String)} return {@code state} for {@code coordId} instead of "absent". */
|
||||
public FakeLeadChannel withMailbox(String coordId, MailboxState state) {
|
||||
mailboxes.put(coordId, state);
|
||||
@@ -80,6 +88,11 @@ public final class FakeLeadChannel implements LeadChannel {
|
||||
return mailboxes.getOrDefault(coordId, MailboxState.absent(coordId));
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean heldDurable() {
|
||||
return heldDurable;
|
||||
}
|
||||
|
||||
public List<LeadMessage> published() {
|
||||
return List.copyOf(published);
|
||||
}
|
||||
|
||||
@@ -41,20 +41,20 @@ class InMemoryReplyInboxTest {
|
||||
@Test
|
||||
void ackRemovesTheMessage() {
|
||||
inbox.publish("term_a", "m1", "hello");
|
||||
inbox.ack("term_a", "m1");
|
||||
assertTrue(inbox.ack("term_a", "m1"), "fleetd #437: ack of a real entry must report true");
|
||||
assertTrue(inbox.peek("term_a").isEmpty(), "after ack, the message is gone");
|
||||
}
|
||||
|
||||
@Test
|
||||
void ackForUnknownMsgIdIsNoOp() {
|
||||
inbox.publish("term_a", "m1", "hello");
|
||||
inbox.ack("term_a", "no-such-id"); // no-op
|
||||
assertFalse(inbox.ack("term_a", "no-such-id"), "fleetd #437: a miss must report false"); // no-op
|
||||
assertEquals(1, inbox.peek("term_a").size(), "the published message is still there");
|
||||
}
|
||||
|
||||
@Test
|
||||
void ackForUnknownTargetIsNoOp() {
|
||||
inbox.ack("no-such-target", "m1"); // no-op, should not throw
|
||||
assertFalse(inbox.ack("no-such-target", "m1"), "fleetd #437: a miss must report false"); // no-op, should not throw
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -167,7 +167,7 @@ class InMemoryReplyInboxTest {
|
||||
@Test
|
||||
void peekAndAckAreNoOpsForUnownedTarget() {
|
||||
assertTrue(inbox.peek("term_not_owned").isEmpty());
|
||||
inbox.ack("term_not_owned", "m1"); // no-op, should not throw
|
||||
assertFalse(inbox.ack("term_not_owned", "m1"), "fleetd #437: a miss must report false"); // no-op, should not throw
|
||||
}
|
||||
|
||||
@Test
|
||||
|
||||
@@ -207,6 +207,21 @@ class LeadMailboxTest {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #440: {@code heldDurable()} must be derived from what {@link LeadMailbox#own} actually
|
||||
* did against the real broker — a durable queue declare plus a manual-ack consumer — not a
|
||||
* hardcoded literal. This is the mutation-sensitive test: flip {@code own()}'s {@code autoAck}
|
||||
* local to {@code true} (or its {@code durableQueue} local to {@code false}) and this must fail.
|
||||
*/
|
||||
@Test
|
||||
void heldDurableReportsTrueBecauseTheQueueIsDurableAndTheConsumeIsManualAck() throws Exception {
|
||||
String self = coordId("lead-held-durable");
|
||||
try (LeadMailbox mailbox = LeadMailbox.open(uri(), self)) {
|
||||
assertTrue(mailbox.heldDurable(),
|
||||
"own() declares a durable queue and consumes with autoAck=false, so held mail is durable");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void inspectReportsAMissingMailboxAsAbsentRatherThanThrowing() throws Exception {
|
||||
String nobody = coordId("lead-inspect-nobody");
|
||||
|
||||
@@ -1802,7 +1802,11 @@ class MessageServiceTest {
|
||||
Thread asker = new Thread(() -> assertThrows(IllegalStateException.class,
|
||||
() -> service.ask(T, "which config file?", 30_000)));
|
||||
asker.start();
|
||||
awaitTicketPhaseOn(service, ticket, MessageService.Phase.ASKING);
|
||||
// Phase.ASKING (markAsyncQuestion) is only the FIRST of ask()'s three steps; the assertion
|
||||
// below depends on the THIRD (pushLoop.onQuestionOpened). Barrier on the push loop's own
|
||||
// pending-question state instead of the phase — see ReplyPushLoop#pendingQuestionTurnIdsForTest.
|
||||
MessageService.TaskView asking = awaitTicketPhaseOn(service, ticket, MessageService.Phase.ASKING);
|
||||
awaitQuestionPendingOn(pushLoop, LEAD, asking.turnId());
|
||||
assertEquals(ReplyPushLoop.Action.INJECT, pushLoop.decide(LEAD, 0, 0, 0),
|
||||
"the open question should be the one thing keeping this lead's schedule alive");
|
||||
|
||||
@@ -1917,6 +1921,12 @@ class MessageServiceTest {
|
||||
// Only now does it reply. Under the old clock this reply was born already expired.
|
||||
assertTrue(rendezvous.resolve(T, "the long report"));
|
||||
awaitTicketPhaseOn(wiring.service(), slow, MessageService.Phase.DONE);
|
||||
// #399: same barrier fix as the sibling eviction test, for consistency — this test's
|
||||
// own assertion happens to survive a late stamp today (the clock is not advanced any
|
||||
// further between here and the sweep below, so cutoff cannot move past whichever
|
||||
// value completedNanos ends up stamped with), but DONE is still the wrong thing to
|
||||
// order on before a sweep that matters for the TTL.
|
||||
awaitCompletionStamped(wiring.service(), slow);
|
||||
|
||||
// A second delegation runs pruneTerminalTickets before it returns.
|
||||
wiring.service().sendAsync(T, "an unrelated second task");
|
||||
@@ -1946,6 +1956,11 @@ class MessageServiceTest {
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "quick result"));
|
||||
awaitTicketPhaseOn(wiring.service(), done, MessageService.Phase.DONE);
|
||||
// #399: DONE can be observed before the completion hook stamps completedNanos. Wait
|
||||
// for the real stamp before advancing the clock, or the hook can run late and stamp
|
||||
// the ADVANCED time — making cutoff = advanced - TTL unreachable and hiding the very
|
||||
// eviction this test exists to pin.
|
||||
awaitCompletionStamped(wiring.service(), done);
|
||||
|
||||
// Nobody collected it, and the TTL has now passed since it FINISHED.
|
||||
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
|
||||
@@ -1956,6 +1971,80 @@ class MessageServiceTest {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #409: pins the #399 ordering invariant deterministically, on the first run, without
|
||||
* relying on host load.
|
||||
*
|
||||
* <p>fleetd #399 was itself only reproducible probabilistically: the real race window between
|
||||
* {@code CompletableFuture.complete()} making {@link MessageService.Phase#DONE} observable and
|
||||
* the constructor's {@code whenComplete} hook actually stamping {@code completedNanos} (see
|
||||
* {@link MessageService#isCompletionStampedForTest}) is normally a handful of instructions wide,
|
||||
* and needed heavy background load on the host to show up in a run at all. This test does not
|
||||
* try to hit that narrow window by chance — it widens it on purpose: {@link #STAMP_DELAY_MILLIS}
|
||||
* is injected into the clock itself, so the completion hook's one read of {@code nowNanos} for
|
||||
* this ticket sleeps before returning, and everything the test does in the meantime (advance the
|
||||
* clock, run a sweep, assert) happens for certain inside that window, on any host.
|
||||
*
|
||||
* <p>Removing the {@link #awaitCompletionStamped} call below reproduces the pre-#399 ordering:
|
||||
* the test then advances the clock and runs its sweep while the hook is still asleep, so the
|
||||
* hook wakes up and stamps {@code completedNanos} with the clock's ALREADY-ADVANCED value
|
||||
* instead of the real completion time. {@code cutoff = advanced - TICKET_TTL_NANOS} can then
|
||||
* never exceed that stamp (the gap between them is fixed at exactly {@code TICKET_TTL_NANOS}),
|
||||
* so the sweep that already ran never evicts the ticket and the very next assertion — expecting
|
||||
* eviction — fails immediately. That is a deterministic, first-run RED failure, not a flaky one
|
||||
* and not a false pass: I ran the test with the barrier call removed and confirmed
|
||||
* {@code assertNull} fails because {@code poll} still returns the ticket's DONE view, which is
|
||||
* exactly this masked-eviction mechanism and not some unrelated defect in the test's own wiring.
|
||||
*/
|
||||
@Test
|
||||
void aTicketOrderedOnDoneInsteadOfTheCompletionStampSurvivesAnEvictionItMustNotSurvive() throws Exception {
|
||||
java.util.concurrent.atomic.AtomicLong clock = new java.util.concurrent.atomic.AtomicLong(1_000_000_000L);
|
||||
// Armed for exactly one call: MessageService's only three nowNanos() call sites are the Task
|
||||
// constructor's createdNanos, this constructor's whenComplete hook's completedNanos, and
|
||||
// pruneTerminalTickets' cutoff — none of which run between arming this (right before
|
||||
// triggering the reply below) and the completion hook firing, so the delayed call is
|
||||
// unambiguously that hook's stamp for `ticket`, never a createdNanos or cutoff read.
|
||||
java.util.concurrent.atomic.AtomicBoolean delayArmed = new java.util.concurrent.atomic.AtomicBoolean(false);
|
||||
java.util.function.LongSupplier gatedClock = () -> {
|
||||
if (delayArmed.compareAndSet(true, false)) {
|
||||
try {
|
||||
Thread.sleep(STAMP_DELAY_MILLIS);
|
||||
} catch (InterruptedException e) {
|
||||
Thread.currentThread().interrupt();
|
||||
}
|
||||
}
|
||||
return clock.get();
|
||||
};
|
||||
MessageService service = new MessageService(agents, injector, rendezvous, inbox, null, null, gatedClock);
|
||||
try {
|
||||
String ticket = service.sendAsync(T, "a quick task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
|
||||
delayArmed.set(true);
|
||||
assertTrue(rendezvous.resolve(T, "quick result"));
|
||||
|
||||
awaitTicketPhaseOn(service, ticket, MessageService.Phase.DONE);
|
||||
// The barrier under test (fleetd #409, same fix as fleetd #399): remove this one call to
|
||||
// reproduce the pre-#399 ordering — see the class-level note above for what happens then.
|
||||
awaitCompletionStamped(service, ticket);
|
||||
|
||||
// A real clock would need the whole TTL to pass; the injected one does it instantly, and
|
||||
// by now completedNanos already holds the REAL (small, unadvanced) completion time.
|
||||
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
|
||||
service.sendAsync(T, "an unrelated second task"); // runs pruneTerminalTickets before returning
|
||||
|
||||
assertNull(service.poll(ticket),
|
||||
"a ticket whose real completion time is long past the advanced cutoff must be "
|
||||
+ "evicted, whatever the completion hook's clock read was delayed by");
|
||||
} finally {
|
||||
service.close();
|
||||
}
|
||||
}
|
||||
|
||||
/** Bounded delay the injected clock sleeps for in {@link #aTicketOrderedOnDoneInsteadOfTheCompletionStampSurvivesAnEvictionItMustNotSurvive}. */
|
||||
private static final long STAMP_DELAY_MILLIS = 300;
|
||||
|
||||
private MessageService.TaskView awaitTicketPhaseOn(MessageService svc, String ticket,
|
||||
MessageService.Phase phase) throws Exception {
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
@@ -1990,6 +2079,45 @@ class MessageServiceTest {
|
||||
return view;
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #399: waits until {@code ticket}'s completion hook has actually stamped
|
||||
* {@code completedNanos}, not just until {@link MessageService#poll} reports
|
||||
* {@link MessageService.Phase#DONE} for it. {@code poll} can observe {@code DONE} the instant
|
||||
* the task's future resolves, before the {@code whenComplete} hook that stamps the completion
|
||||
* time has run — {@code CompletableFuture.complete()} publishes its result and only then runs
|
||||
* dependents. A test that is about to advance an injected clock past the TTL must order itself
|
||||
* after the stamp, not after {@code DONE}: winning the race the other way stamps the
|
||||
* *advanced* clock value and can hide a real eviction bug behind a false pass.
|
||||
*/
|
||||
private void awaitCompletionStamped(MessageService svc, String ticket) throws Exception {
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!svc.isCompletionStampedForTest(ticket)) {
|
||||
assertTrue(System.currentTimeMillis() < deadline,
|
||||
"completedNanos for " + ticket + " was never stamped");
|
||||
Thread.sleep(5);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #418: waits until {@code turnId} actually appears in {@code pushLoop}'s own open-question
|
||||
* state for {@code lead} — {@link ReplyPushLoop#pendingQuestionTurnIdsForTest} — rather than until
|
||||
* {@link MessageService#poll} reports {@link MessageService.Phase#ASKING}. {@code Phase.ASKING} is
|
||||
* set by {@code markAsyncQuestion}, the FIRST of three steps {@code MessageService.ask()} performs;
|
||||
* {@code pushLoop.onQuestionOpened} (the one that actually publishes to {@code pendingQuestions})
|
||||
* is the THIRD. Under load the asker thread can be descheduled between those two steps, so a
|
||||
* barrier on the phase alone can release before the push loop has anything pending — a test that
|
||||
* then asserts on {@link ReplyPushLoop#decide} is asserting on state that has not been published
|
||||
* yet, not on the throwing path it is named for.
|
||||
*/
|
||||
private void awaitQuestionPendingOn(ReplyPushLoop pushLoop, String lead, String turnId) throws Exception {
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!pushLoop.pendingQuestionTurnIdsForTest(lead).contains(turnId)) {
|
||||
assertTrue(System.currentTimeMillis() < deadline,
|
||||
"turnId " + turnId + " never appeared in the push loop's pending questions for " + lead);
|
||||
Thread.sleep(5);
|
||||
}
|
||||
}
|
||||
|
||||
// --- CB-640: fleet health evidence accessors --------------------------------------------
|
||||
|
||||
@Test
|
||||
|
||||
@@ -116,4 +116,262 @@ class BackendQuarantineTest {
|
||||
assertThrows(IllegalArgumentException.class, () -> new BackendQuarantine(() -> 0L, 0L));
|
||||
assertThrows(IllegalArgumentException.class, () -> new BackendQuarantine(() -> 0L, -1L));
|
||||
}
|
||||
|
||||
// --- fleetd #466: escalating cooldown -----------------------------------------------------
|
||||
//
|
||||
// Base cooldown 600s (10 min), multiplier 2.0, ceiling 2400s (4x base) — small round numbers
|
||||
// chosen so every deadline is an exact assertion, not just "greater than before". Each call
|
||||
// below lands at or before the previous deadline (a zero or negative gap), which is always
|
||||
// "no more than one base cooldown after the previous deadline" — i.e. every call continues the
|
||||
// same streak, matching a credential that keeps reporting exhausted with no lull.
|
||||
|
||||
@Test
|
||||
void anInvalidBackoffMultiplierIsRejected() {
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 0.5, TimeUnit.HOURS.toNanos(6)));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aCeilingBelowTheBaseCooldownIsRejected() {
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 2.0, TimeUnit.MINUTES.toNanos(10)));
|
||||
}
|
||||
|
||||
@Test
|
||||
void repeatedExhaustionEscalatesTheCooldownByExactAmounts() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
q.quarantine("shared-openai"); // 1st: base cooldown
|
||||
assertEquals(OptionalLong.of(600L), q.remainingSeconds("shared-openai"));
|
||||
|
||||
now.set(TimeUnit.SECONDS.toNanos(100)); // still inside the 1st quarantine (deadline 600s)
|
||||
q.quarantine("shared-openai"); // 2nd: 600 * 2^1 = 1200
|
||||
assertEquals(OptionalLong.of(1200L), q.remainingSeconds("shared-openai"),
|
||||
"a second consecutive exhaustion must double the cooldown, not just increase it");
|
||||
|
||||
now.set(TimeUnit.SECONDS.toNanos(1300)); // exactly the 2nd deadline (100 + 1200)
|
||||
q.quarantine("shared-openai"); // 3rd: 600 * 2^2 = 2400 (exactly at the ceiling)
|
||||
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void escalationStopsAtTheCeiling() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
q.quarantine("shared-openai"); // 1st: 600
|
||||
now.set(TimeUnit.SECONDS.toNanos(600));
|
||||
q.quarantine("shared-openai"); // 2nd: 1200, deadline 1800
|
||||
now.set(TimeUnit.SECONDS.toNanos(1800));
|
||||
q.quarantine("shared-openai"); // 3rd: 600 * 4 = 2400, at the ceiling, deadline 4200
|
||||
now.set(TimeUnit.SECONDS.toNanos(4200));
|
||||
q.quarantine("shared-openai"); // 4th: 600 * 8 = 4800 uncapped, must stay capped at 2400
|
||||
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"),
|
||||
"the cooldown must never exceed the configured ceiling, however long the streak gets");
|
||||
|
||||
now.set(TimeUnit.SECONDS.toNanos(6600)); // 4th deadline
|
||||
q.quarantine("shared-openai"); // 5th: still capped
|
||||
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"),
|
||||
"pushing well past the ceiling must not budge it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aQuietGapLongerThanTheBaseCooldownResetsToTheBaseCooldown() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
q.quarantine("shared-openai"); // 1st: 600, deadline 600
|
||||
now.set(TimeUnit.SECONDS.toNanos(600));
|
||||
q.quarantine("shared-openai"); // 2nd: 1200, deadline 1800
|
||||
now.set(TimeUnit.SECONDS.toNanos(1800));
|
||||
q.quarantine("shared-openai"); // 3rd: 2400, deadline 4200
|
||||
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"));
|
||||
|
||||
// Quiet for well over one base cooldown (600s) past the 3rd deadline (4200s).
|
||||
now.set(TimeUnit.SECONDS.toNanos(20_000));
|
||||
q.quarantine("shared-openai"); // treated as a fresh occurrence
|
||||
assertEquals(OptionalLong.of(600L), q.remainingSeconds("shared-openai"),
|
||||
"a long quiet gap must reset the streak back to the base cooldown");
|
||||
}
|
||||
|
||||
@Test
|
||||
void escalatingOneCredentialDoesNotSlowAnother() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
q.quarantine("shared-openai"); // 1st: 600
|
||||
now.set(TimeUnit.SECONDS.toNanos(600));
|
||||
q.quarantine("shared-openai"); // 2nd: 1200
|
||||
now.set(TimeUnit.SECONDS.toNanos(1800));
|
||||
q.quarantine("shared-openai"); // 3rd: 2400 — three-in-a-row streak on this credential only
|
||||
|
||||
q.quarantine("another-credential"); // its first and only exhaustion
|
||||
assertEquals(OptionalLong.of(600L), q.remainingSeconds("another-credential"),
|
||||
"an unrelated credential's cooldown must stay at the base rate, unaffected by a sibling's streak");
|
||||
}
|
||||
|
||||
@Test
|
||||
void withEscalationDefaultsToDoublingCappedAtTwelveTimesTheBase() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = BackendQuarantine.withEscalation(now::get, TimeUnit.MINUTES.toNanos(30));
|
||||
|
||||
q.quarantine("shared-openai");
|
||||
assertEquals(OptionalLong.of(1800L), q.remainingSeconds("shared-openai"),
|
||||
"the first occurrence must still use the base cooldown");
|
||||
|
||||
now.set(TimeUnit.MINUTES.toNanos(30));
|
||||
q.quarantine("shared-openai");
|
||||
assertEquals(OptionalLong.of(3600L), q.remainingSeconds("shared-openai"),
|
||||
"the default multiplier must be 2.0");
|
||||
}
|
||||
|
||||
// --- fleetd #466 scope item 2: status() reports repeatCount beside remainingSeconds ------------
|
||||
|
||||
@Test
|
||||
void statusReportsAttemptOneForAFirstOccurrenceNeverAbsentOrZero() {
|
||||
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 2.0,
|
||||
TimeUnit.HOURS.toNanos(6));
|
||||
q.quarantine("shared-openai");
|
||||
|
||||
BackendQuarantine.Status status = q.status("shared-openai").orElseThrow();
|
||||
assertEquals(1800L, status.remainingSeconds());
|
||||
assertEquals(1, status.repeatCount(),
|
||||
"a first-ever occurrence must report attempt 1, not 0 or absent -- 1 means unambiguously "
|
||||
+ "'the first time', where 0 would be indistinguishable from a bug that forgot to count");
|
||||
}
|
||||
|
||||
@Test
|
||||
void statusIsAbsentWhenTheCredentialIsNotQuarantined() {
|
||||
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 2.0,
|
||||
TimeUnit.HOURS.toNanos(6));
|
||||
assertTrue(q.status("shared-openai").isEmpty());
|
||||
}
|
||||
|
||||
/**
|
||||
* The acceptance criterion's strong form: the reported {@code repeatCount} must match the exact
|
||||
* step the cooldown's own growth implies, read off {@link BackendQuarantine#status}'s single
|
||||
* call -- not two independent reads that happen to agree in this easy case.
|
||||
*/
|
||||
@Test
|
||||
void statusReportsTheGrowingAttemptCountAlongsideTheEscalatingCooldown() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
q.quarantine("shared-openai"); // 1st: 600s, attempt 1
|
||||
BackendQuarantine.Status first = q.status("shared-openai").orElseThrow();
|
||||
assertEquals(600L, first.remainingSeconds());
|
||||
assertEquals(1, first.repeatCount());
|
||||
|
||||
now.set(TimeUnit.SECONDS.toNanos(600));
|
||||
q.quarantine("shared-openai"); // 2nd: 1200s, attempt 2
|
||||
BackendQuarantine.Status second = q.status("shared-openai").orElseThrow();
|
||||
assertEquals(1200L, second.remainingSeconds());
|
||||
assertEquals(2, second.repeatCount());
|
||||
|
||||
now.set(TimeUnit.SECONDS.toNanos(1800));
|
||||
q.quarantine("shared-openai"); // 3rd: 2400s (at the ceiling), attempt 3
|
||||
BackendQuarantine.Status third = q.status("shared-openai").orElseThrow();
|
||||
assertEquals(2400L, third.remainingSeconds());
|
||||
assertEquals(3, third.repeatCount());
|
||||
}
|
||||
|
||||
/**
|
||||
* Once the cooldown hits its ceiling, every further consecutive exhaustion reports the SAME
|
||||
* {@code remainingSeconds} -- so a {@code repeatCount} re-derived from the cooldown value (e.g.
|
||||
* inverting {@code cooldownNanos * multiplier^(n-1)}) could not tell attempt 4 from attempt 9;
|
||||
* only the stored counter can. This is the scenario that makes "read the count off a second,
|
||||
* independent computation" provably wrong rather than just risky.
|
||||
*/
|
||||
@Test
|
||||
void repeatCountKeepsGrowingPastTheCeilingEvenThoughTheCooldownStaysFlat() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
long t = 0L;
|
||||
for (int attempt = 1; attempt <= 5; attempt++) {
|
||||
now.set(t);
|
||||
q.quarantine("shared-openai");
|
||||
BackendQuarantine.Status status = q.status("shared-openai").orElseThrow();
|
||||
assertEquals(attempt, status.repeatCount(),
|
||||
"attempt " + attempt + " must be reported as exactly " + attempt
|
||||
+ ", not collapsed to whatever attempt first reached the ceiling");
|
||||
if (attempt >= 3) {
|
||||
assertEquals(2400L, status.remainingSeconds(), "attempt " + attempt + " must be capped");
|
||||
}
|
||||
t += status.remainingSeconds(); // land exactly on the next deadline: still the same streak
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void aQuietGapResetsTheReportedAttemptCountToOneToo() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
q.quarantine("shared-openai");
|
||||
now.set(TimeUnit.SECONDS.toNanos(600));
|
||||
q.quarantine("shared-openai");
|
||||
now.set(TimeUnit.SECONDS.toNanos(1800));
|
||||
q.quarantine("shared-openai"); // attempt 3
|
||||
assertEquals(3, q.status("shared-openai").orElseThrow().repeatCount());
|
||||
|
||||
now.set(TimeUnit.SECONDS.toNanos(20_000)); // long quiet gap
|
||||
q.quarantine("shared-openai");
|
||||
assertEquals(1, q.status("shared-openai").orElseThrow().repeatCount(),
|
||||
"a reset streak must report attempt 1 again, matching the reset base cooldown");
|
||||
}
|
||||
|
||||
@Test
|
||||
void escalatingOneCredentialsAttemptCountDoesNotAffectAnother() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
q.quarantine("shared-openai");
|
||||
now.set(TimeUnit.SECONDS.toNanos(600));
|
||||
q.quarantine("shared-openai");
|
||||
now.set(TimeUnit.SECONDS.toNanos(1800));
|
||||
q.quarantine("shared-openai"); // 3-in-a-row streak on this credential only
|
||||
|
||||
q.quarantine("another-credential");
|
||||
assertEquals(1, q.status("another-credential").orElseThrow().repeatCount(),
|
||||
"an unrelated credential's attempt count must stay at 1, unaffected by a sibling's streak");
|
||||
}
|
||||
|
||||
@Test
|
||||
void noneReportsNoStatusForAnything() {
|
||||
BackendQuarantine q = BackendQuarantine.none();
|
||||
q.quarantine("shared-openai"); // no-op on none(), same as every other mutator
|
||||
|
||||
assertTrue(q.status("shared-openai").isEmpty(),
|
||||
"none() quarantines nothing, so it must report no status at all -- never a fabricated "
|
||||
+ "attempt count for a credential that was never actually quarantined");
|
||||
}
|
||||
|
||||
/** The flat (non-escalating) two-argument constructor must still report a real, growing count. */
|
||||
@Test
|
||||
void aFlatTwoArgumentInstanceStillReportsAGrowingAttemptCountEvenThoughTheCooldownStaysFlat() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600));
|
||||
|
||||
q.quarantine("shared-openai");
|
||||
BackendQuarantine.Status first = q.status("shared-openai").orElseThrow();
|
||||
assertEquals(600L, first.remainingSeconds());
|
||||
assertEquals(1, first.repeatCount());
|
||||
|
||||
now.set(TimeUnit.SECONDS.toNanos(100)); // still inside the 1st quarantine
|
||||
q.quarantine("shared-openai");
|
||||
BackendQuarantine.Status second = q.status("shared-openai").orElseThrow();
|
||||
assertEquals(600L, second.remainingSeconds(),
|
||||
"the flat constructor's cooldown must stay exactly the base length regardless of the streak");
|
||||
assertEquals(2, second.repeatCount(),
|
||||
"the flat constructor still counts the real streak -- it just does not scale the cooldown by it");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -90,6 +90,63 @@ class PlacementPolicyTest {
|
||||
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
|
||||
}
|
||||
|
||||
// --- fleetd #422: FixedPlacementPolicy must consult modelOff too, at BOTH filter sites ------
|
||||
|
||||
/** The default-profile fast path (:44) must skip a default whose model is off. */
|
||||
@Test
|
||||
void fixedSkipsModelOffDefault() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
||||
noSessions(), Set.of(), Set.of(), Set.of(), Set.of("b"));
|
||||
assertEquals("a", policy.select(ctx).profile(),
|
||||
"the default 'b' names an off model, so fixed falls through to the first available candidate");
|
||||
}
|
||||
|
||||
/**
|
||||
* The fallback walk (:48-53) must skip an off-model candidate too — exercised independently of
|
||||
* the default-profile fast path by using no default at all, so this is the only filter that runs.
|
||||
*/
|
||||
@Test
|
||||
void fixedFallbackWalkSkipsModelOffCandidate() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext(null,
|
||||
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
||||
noSessions(), Set.of(), Set.of(), Set.of(), Set.of("a"));
|
||||
assertEquals("b", policy.select(ctx).profile(),
|
||||
"candidate 'a' names an off model, so the fallback walk skips it and picks 'b'");
|
||||
}
|
||||
|
||||
@Test
|
||||
void fixedThrowsWhenDefaultAndEveryCandidateModelOff() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
||||
noSessions(), Set.of(), Set.of(), Set.of(), Set.of("a", "b"));
|
||||
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||
assertTrue(e.getMessage().contains("turned off"),
|
||||
"message names model-off as the cause: " + e.getMessage());
|
||||
assertFalse(e.getMessage().contains("quarantined"), "must not read like quarantine: " + e.getMessage());
|
||||
assertFalse(e.getMessage().contains("cooling off"), "must not read like cool-off: " + e.getMessage());
|
||||
assertFalse(e.getMessage().contains("weight 0"), "must not read like weight-0: " + e.getMessage());
|
||||
}
|
||||
|
||||
/**
|
||||
* Quarantine still wins when a profile is both quarantined and model-off (mirrors {@code
|
||||
* fixedThrowsWhenDefaultAndEveryCandidateQuarantined}'s priority over cooling off).
|
||||
*/
|
||||
@Test
|
||||
void fixedReportsQuarantineNotModelOffWhenBothApply() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
||||
noSessions(), Set.of(), Set.of("a", "b"), Set.of(), Set.of("a", "b"));
|
||||
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
|
||||
assertFalse(e.getMessage().contains("turned off"),
|
||||
"quarantine takes priority over model-off in the message: " + e.getMessage());
|
||||
}
|
||||
|
||||
@Test
|
||||
void roundRobinCyclesThroughAvailableProfiles() {
|
||||
PlacementPolicy policy = PlacementPolicies.roundRobin();
|
||||
@@ -404,6 +461,93 @@ class PlacementPolicyTest {
|
||||
assertTrue(e.getMessage().contains("weight 0"), e.getMessage());
|
||||
}
|
||||
|
||||
// --- fleetd #435: FixedPlacementPolicy must consult maxLoad too, at BOTH filter sites -------
|
||||
|
||||
/**
|
||||
* The default-profile fast path must skip a capped default. Measured on 7667727 before this
|
||||
* fix: a single dev profile at {@code maxLoad: 1} with 1 live, under {@code fixed()}, still
|
||||
* returned "SPAWNED on profile=a" — the cap was advisory for every unqualified spawn.
|
||||
*/
|
||||
@Test
|
||||
void fixedSkipsCappedDefault() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a", 1.0f, null),
|
||||
PlacementCandidate.profile("b", 1.0f, 1)),
|
||||
name -> "b".equals(name) ? 1 : 0, Set.of(), Set.of(), Set.of());
|
||||
assertEquals("a", policy.select(ctx).profile(),
|
||||
"the default 'b' is at its maxLoad cap, so fixed falls through to the free candidate 'a'");
|
||||
}
|
||||
|
||||
/**
|
||||
* The fallback walk must skip a capped candidate too — exercised independently of the
|
||||
* default-profile fast path by using no default at all, so this is the only filter that runs.
|
||||
*/
|
||||
@Test
|
||||
void fixedFallbackWalkSkipsCappedCandidate() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext(null,
|
||||
List.of(PlacementCandidate.profile("a", 1.0f, 1),
|
||||
PlacementCandidate.profile("b", 1.0f, null)),
|
||||
name -> "a".equals(name) ? 1 : 0, Set.of(), Set.of(), Set.of());
|
||||
assertEquals("b", policy.select(ctx).profile(),
|
||||
"candidate 'a' is at its maxLoad cap, so the fallback walk skips it and picks 'b'");
|
||||
}
|
||||
|
||||
@Test
|
||||
void fixedThrowsWhenDefaultAndEveryCandidateAtCap() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a", 1.0f, 1),
|
||||
PlacementCandidate.profile("b", 1.0f, 1)),
|
||||
_ -> 1, Set.of(), Set.of(), Set.of());
|
||||
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||
assertTrue(e.getMessage().contains("maxLoad"), "message names the cap: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("1 cap"), "message names the cap value: " + e.getMessage());
|
||||
assertFalse(e.getMessage().contains("quarantined"), "must not read like quarantine: " + e.getMessage());
|
||||
assertFalse(e.getMessage().contains("turned off"), "must not read like model-off: " + e.getMessage());
|
||||
}
|
||||
|
||||
/** The mirror: an uncapped default is still chosen, so the new term cannot exclude everything. */
|
||||
@Test
|
||||
void fixedStillReturnsUncappedDefault() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a", 1.0f, null),
|
||||
PlacementCandidate.profile("b", 1.0f, null)),
|
||||
noSessions(), Set.of(), Set.of(), Set.of());
|
||||
assertEquals("b", policy.select(ctx).profile(), "an uncapped default is returned unconditionally");
|
||||
}
|
||||
|
||||
/** CB-585: an explicit {@code maxLoad: 0} on the default caps it at zero live members. */
|
||||
@Test
|
||||
void fixedSkipsMaxLoadZeroDefaultEvenWithZeroLiveWorkers() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a", 1.0f, null),
|
||||
PlacementCandidate.profile("b", 1.0f, 0)),
|
||||
_ -> 0, Set.of(), Set.of(), Set.of());
|
||||
assertEquals("a", policy.select(ctx).profile(),
|
||||
"the default 'b' has maxLoad 0, so it is already at its cap with nobody live");
|
||||
}
|
||||
|
||||
/**
|
||||
* Quarantine still wins when a profile is both quarantined and at cap (mirrors {@code
|
||||
* fixedReportsQuarantineNotModelOffWhenBothApply}'s priority over model-off).
|
||||
*/
|
||||
@Test
|
||||
void fixedReportsQuarantineNotAtCapWhenBothApply() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a", 1.0f, 1),
|
||||
PlacementCandidate.profile("b", 1.0f, 1)),
|
||||
_ -> 1, Set.of(), Set.of("a", "b"), Set.of());
|
||||
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
|
||||
assertFalse(e.getMessage().contains("maxLoad"),
|
||||
"quarantine takes priority over at-cap in the message: " + e.getMessage());
|
||||
}
|
||||
|
||||
@Test
|
||||
void unknownPolicyNameThrows() {
|
||||
assertThrows(IllegalArgumentException.class, () -> PlacementPolicies.fromName("random"));
|
||||
|
||||
@@ -6,12 +6,14 @@ import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.auth.MemberRegistry;
|
||||
import dev.ltms.fleet.auth.MemberLifecycle;
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||
import dev.ltms.fleet.msg.TestTurnTokens;
|
||||
import dev.ltms.fleet.peer.Capability;
|
||||
import dev.ltms.fleet.peer.CharterReceipt;
|
||||
@@ -20,9 +22,16 @@ import dev.ltms.fleet.peer.PeerHandle;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.PeerUnreachableException;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
@@ -33,6 +42,7 @@ import java.util.concurrent.Future;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicInteger;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
@@ -1027,6 +1037,11 @@ class SessionManagerTest {
|
||||
return delegate.spawn(req);
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
return delegate.spawn(req, decision);
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return delegate.profiles();
|
||||
@@ -1588,6 +1603,11 @@ class SessionManagerTest {
|
||||
throw new UnsupportedOperationException("not reachable — the capability check refuses first");
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
throw new UnsupportedOperationException("not reachable — the capability check refuses first");
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return Set.of("stub-profile");
|
||||
@@ -1652,6 +1672,11 @@ class SessionManagerTest {
|
||||
return delegate.spawn(req);
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
return delegate.spawn(req, decision);
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return delegate.profiles();
|
||||
@@ -1858,6 +1883,11 @@ class SessionManagerTest {
|
||||
return handle;
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
return spawn(req.withProfile(decision.profile()));
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return Set.of("lazy");
|
||||
@@ -1991,4 +2021,296 @@ class SessionManagerTest {
|
||||
sessions.rosterResolved();
|
||||
assertEquals(2, handle.callCount(), "once resolved, the id must not be looked up again");
|
||||
}
|
||||
|
||||
// ── fleetd #425 criterion 3: acquireWithWorktree must provision for the profile it actually
|
||||
// spawns, never a name resolved before a live pool change is accounted for ────────────────────
|
||||
|
||||
/** Two profiles with distinct {@code cwd}/{@code parityOverlay}, and a dev pool of {@code first,second}. */
|
||||
private static String worktreeReorderYaml(String first, String second) {
|
||||
return """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
a:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: coder-a
|
||||
b:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: coder-b
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
fleet:
|
||||
developers:
|
||||
slot0:
|
||||
profile: %s
|
||||
slot1:
|
||||
profile: %s
|
||||
""".formatted(first, second);
|
||||
}
|
||||
|
||||
@Test
|
||||
void acquireWithWorktreeProvisionsTheOverlayForTheProfileActuallySpawned(
|
||||
@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, worktreeReorderYaml("a", "b"));
|
||||
ConfigRef ref = new ConfigRef(f, FleetConfig.load(f));
|
||||
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of(
|
||||
"a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json")),
|
||||
"b", new FleetConfig.Profile("b", "http://gx00.gw:8000", "coder-b", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/b", List.of("b.mcp.json")));
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
|
||||
PeerLauncher launcher = new CompositePeerLauncher(
|
||||
List.of(adapter), "a", ref, _ -> 0, BackendQuarantine.none());
|
||||
|
||||
// The pool changes AFTER the composite/launcher is built, and BEFORE the unqualified
|
||||
// worktree spawn — exactly the fleetd #425 scenario: the live pool's first entry is "b" by
|
||||
// the time acquireWithWorktree runs, even though nothing here was rebuilt.
|
||||
Files.writeString(f, worktreeReorderYaml("b", "a"));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), () -> "reload should apply cleanly: " + out.error());
|
||||
|
||||
FakeWorktrees worktrees = new FakeWorktrees();
|
||||
SessionManager sessions = new SessionManager(launcher, worktrees, () -> 0L);
|
||||
|
||||
MemberSession s = sessions.acquire(null, null, "/caller",
|
||||
null, new WorktreeRequest("fleetd-425", null));
|
||||
|
||||
assertEquals("b", s.profile(),
|
||||
"the live dev pool now starts at b, so the unqualified spawn must land there");
|
||||
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
|
||||
assertNotNull(overlay, "overlayParity must have been called");
|
||||
assertEquals(List.of("b.mcp.json"), overlay.requested(),
|
||||
"the worktree must be provisioned with profile b's overlay — the one actually "
|
||||
+ "spawned — never a's, the pool's stale first entry");
|
||||
}
|
||||
|
||||
/**
|
||||
* The deterministic, mutation-pinning half of criterion 3: {@code launcher.defaultProfile()}
|
||||
* only ever answers for {@link MemberRole#DEV} (see {@link CompositePeerLauncher#defaultProfile()}),
|
||||
* so resolving a worktree spawn's profile through it — instead of through {@link
|
||||
* PeerLauncher#defaultProfileFor(MemberRole)}, resolved against the CALLER's actual role — picks
|
||||
* the wrong pool's answer for any role other than DEV. No reload or race is needed to see it: an
|
||||
* ARCHITECT pool and a DEV pool that simply disagree, held constant, are enough.
|
||||
*/
|
||||
@Test
|
||||
void acquireWithWorktreeForANonDevRoleUsesThatRolesPoolNotTheDevPool() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of(
|
||||
"a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json")),
|
||||
"b", new FleetConfig.Profile("b", "http://gx00.gw:8000", "coder-b", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/b", List.of("b.mcp.json")));
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
|
||||
// developers -> a (first/only entry); architects -> b (first/only entry). The two pools
|
||||
// disagree on purpose, so a role-blind resolution (DEV's answer, "a") is visibly wrong for
|
||||
// an ARCHITECT spawn, which must land on "b".
|
||||
FleetConfig.Fleet fleet = new FleetConfig.Fleet(Map.of(),
|
||||
Map.of("s0", new FleetConfig.Slot("b")),
|
||||
Map.of("s0", new FleetConfig.Slot("a")),
|
||||
Map.of(), null);
|
||||
PeerLauncher launcher = new CompositePeerLauncher(List.of(adapter), "a", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, fleet);
|
||||
|
||||
FakeWorktrees worktrees = new FakeWorktrees();
|
||||
SessionManager sessions = new SessionManager(launcher, worktrees, () -> 0L);
|
||||
|
||||
MemberSession s = sessions.acquire(null, MemberRole.ARCHITECT, null, "/caller",
|
||||
null, new WorktreeRequest("fleetd-425b", null));
|
||||
|
||||
assertEquals("b", s.profile(),
|
||||
"an unqualified ARCHITECT worktree spawn must land on the architect pool's profile");
|
||||
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
|
||||
assertNotNull(overlay, "overlayParity must have been called");
|
||||
assertEquals(List.of("b.mcp.json"), overlay.requested(),
|
||||
"the worktree must be provisioned with profile b's overlay — the ARCHITECT pool's "
|
||||
+ "answer, the one actually spawned — never a's, the DEV pool's answer that "
|
||||
+ "launcher.defaultProfile() alone would have given");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #425 rework, acceptance 3: repoRoot, parityOverlay, AND the actual spawn must all name
|
||||
* the SAME routed profile, proven on the ROUTED path — a quarantine skips the pool's first entry
|
||||
* — not just the "pool reordered by a live reload" path the two tests above already cover.
|
||||
*
|
||||
* <p>This is the exact regression the rework fixes: the first round resolved
|
||||
* {@code acquireWithWorktree}'s profile through {@code launcher.defaultProfileFor(memberRole)},
|
||||
* which is blind to quarantine and just returns the pool's first entry ("a" here, quarantined).
|
||||
* That name went on to provision repoRoot/overlay for "a", and then the spawn itself — now an
|
||||
* EXPLICIT-profile spawn naming "a" — hit {@code CompositePeerLauncher.enforceNotQuarantined}
|
||||
* and threw, where the pre-fix code (a blank-profile spawn) would have routed around "a" onto
|
||||
* "b" without any trouble. {@code launcher.routedProfileFor(memberRole)} closes that gap by
|
||||
* running the SAME quarantine-aware selection {@code spawn} itself uses, so all three — repoRoot,
|
||||
* overlay, and the spawn — land on "b" together.
|
||||
*/
|
||||
@Test
|
||||
void acquireWithWorktreeRoutesAroundAQuarantinedPoolFirstProfile() {
|
||||
Map<String, FleetConfig.Profile> profiles = new LinkedHashMap<>();
|
||||
profiles.put("a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json"),
|
||||
null, null, null, null, null, null, null, null, "shared-cred", null));
|
||||
profiles.put("b", new FleetConfig.Profile("b", "http://gx00.gw:8000", "coder-b", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/b", List.of("b.mcp.json")));
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
quarantine.quarantine("shared-cred");
|
||||
PeerLauncher launcher = new CompositePeerLauncher(List.of(adapter), "a", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
|
||||
|
||||
FakeWorktrees worktrees = new FakeWorktrees();
|
||||
SessionManager sessions = new SessionManager(launcher, worktrees, () -> 0L);
|
||||
|
||||
MemberSession s = sessions.acquire(null, null, "/caller",
|
||||
null, new WorktreeRequest("fleetd-425c", null));
|
||||
|
||||
assertEquals("b", s.profile(),
|
||||
"a is quarantined, so the unqualified worktree spawn must route to b");
|
||||
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
|
||||
assertNotNull(overlay, "overlayParity must have been called");
|
||||
assertEquals(List.of("b.mcp.json"), overlay.requested(),
|
||||
"parityOverlay must be provisioned for b — the profile actually spawned, never a's, "
|
||||
+ "the quarantined pool-first entry");
|
||||
FakeWorktrees.RepoRootCall repoRootCall = worktrees.repoRootCalls().getLast();
|
||||
assertTrue(repoRootCall.cwd().contains("/repo/b"),
|
||||
"repoRoot must be resolved through b's effectiveCwd, not a's: " + repoRootCall.cwd());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #425 rework, round 2: this is the exact probe that found round 1's maxLoad
|
||||
* regression. One dev profile ("a") is configured with {@code maxLoad: 1} and a liveCount
|
||||
* pinned at 1 — permanently at cap — under {@code PlacementPolicies.fixed()}, the default
|
||||
* policy, which deliberately never evaluates {@code maxLoad} during automatic selection (see
|
||||
* {@code CompositePeerLauncher}'s own javadoc on {@code place}/{@code FixedPlacementPolicy}).
|
||||
*
|
||||
* <p>Round 1 resolved {@code acquireWithWorktree}'s profile through
|
||||
* {@code launcher.routedProfileFor(memberRole)} and then fed that name back into
|
||||
* {@code launcher.spawn(SpawnRequest)} as an EXPLICIT profile. Naming a profile explicitly
|
||||
* takes {@code CompositePeerLauncher.spawn}'s THROWING branch, which calls
|
||||
* {@code enforceMaxLoad} — so the worktree path died with a {@code PlacementException} while
|
||||
* the exact same unqualified request, with no worktree, still spawned cleanly through the
|
||||
* routing branch that never checks {@code maxLoad} at all. One intent, two different answers,
|
||||
* depending only on whether a worktree was asked for — the #425 shape, moved to a different
|
||||
* filter instead of closed.
|
||||
*
|
||||
* <p>This test does not hardcode which of the two outcomes is correct — whether an unqualified
|
||||
* spawn SHOULD respect {@code maxLoad} is fleetd #435, a separate ticket. It only asserts that
|
||||
* the WITH-worktree and WITHOUT-worktree paths agree: both spawn on the same profile, or both
|
||||
* fail with the same exception type and message. That way this test stays correct however
|
||||
* #435 is eventually resolved, and only breaks if the two paths disagree again.
|
||||
*/
|
||||
@Test
|
||||
void unqualifiedAcquireAgreesWithAndWithoutAWorktreeWhenTheOnlyProfileIsAtMaxLoad() {
|
||||
Map<String, FleetConfig.Profile> profiles = new LinkedHashMap<>();
|
||||
profiles.put("a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json"),
|
||||
null, null, null, null, 1.0f, 1));
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
|
||||
// liveCount pinned at 1 for "a", exactly matching maxLoad — "a" is permanently at cap,
|
||||
// regardless of how many times either branch below actually spawns.
|
||||
PeerLauncher launcher = new CompositePeerLauncher(List.of(adapter), "a", profiles,
|
||||
PlacementPolicies.fixed(), name -> "a".equals(name) ? 1 : 0);
|
||||
|
||||
Object without = attemptAcquire(() ->
|
||||
new SessionManager(launcher, new FakeWorktrees(), () -> 0L)
|
||||
.acquire(null, null, "/caller", null));
|
||||
Object with = attemptAcquire(() ->
|
||||
new SessionManager(launcher, new FakeWorktrees(), () -> 0L)
|
||||
.acquire(null, null, "/caller", null, new WorktreeRequest("fleetd-425-maxload", null)));
|
||||
|
||||
assertEquals(without, with, "an unqualified spawn on a profile at maxLoad must agree "
|
||||
+ "whether or not a worktree was requested — no condition may become newly fatal "
|
||||
+ "on the worktree path alone (fleetd #425 rework, round 2)");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #425 rework, round 4: the exact regression a mutation test found that 186 green tests
|
||||
* missed — {@code acquireWithWorktree} dropping the {@link
|
||||
* dev.ltms.fleet.placement.PlacementDecision} it already resolved via {@code launcher.place},
|
||||
* and letting the unqualified spawn re-run placement a second time (a blank-profile {@code
|
||||
* launcher.spawn(spawnReq)}) instead of carrying that decision forward via {@code
|
||||
* launcher.spawn(spawnReq, decision)}. Every earlier test in this file uses {@code
|
||||
* PlacementPolicies.fixed()}, which returns the same answer on every {@code select()} call, so
|
||||
* dropping the decision is invisible under it — two {@code select()} calls simply agree by
|
||||
* accident. {@code PlacementPolicies.roundRobin()} is deterministic AND stateful: its {@code
|
||||
* select()} advances an internal index on every call, so two consecutive calls for the SAME
|
||||
* spawn (one from {@code place()} to provision the worktree, a second from a dropped-decision
|
||||
* blank-profile {@code spawn(spawnReq)}) land on DIFFERENT profiles from a two-profile pool —
|
||||
* index 0 ("a"), then index 1 ("b").
|
||||
*
|
||||
* <p>This test does not hardcode which profile wins — asserting one specific name would pass
|
||||
* for the wrong reason the moment the rotation order changes (round-4 brief invariant 3). It
|
||||
* asserts AGREEMENT instead: whichever profile the worktree's parity overlay was provisioned
|
||||
* for must be the SAME profile the member actually spawned on. Each profile's overlay list is
|
||||
* named after the profile itself ({@code "a.mcp.json"}/{@code "b.mcp.json"}), so comparing the
|
||||
* recorded overlay against {@code s.profile() + ".mcp.json"} checks agreement without ever
|
||||
* naming an expected winner.
|
||||
*/
|
||||
@Test
|
||||
void acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedTheWorktreeForUnderARotatingPolicy() {
|
||||
Map<String, FleetConfig.Profile> profiles = new LinkedHashMap<>();
|
||||
profiles.put("a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json")));
|
||||
profiles.put("b", new FleetConfig.Profile("b", "http://gx00.gw:8000", "coder-b", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/b", List.of("b.mcp.json")));
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
|
||||
// roundRobin is deterministic AND stateful: the first select() call picks index 0 ("a"),
|
||||
// and the SAME policy instance's second select() call (reached only if the
|
||||
// PlacementDecision is dropped) picks index 1 ("b") — the two-call disagreement this test
|
||||
// needs to make a dropped decision observable, rather than merely probable.
|
||||
PeerLauncher launcher = new CompositePeerLauncher(List.of(adapter), "a", profiles,
|
||||
PlacementPolicies.roundRobin(), _ -> 0);
|
||||
|
||||
FakeWorktrees worktrees = new FakeWorktrees();
|
||||
SessionManager sessions = new SessionManager(launcher, worktrees, () -> 0L);
|
||||
|
||||
MemberSession s = sessions.acquire(null, null, "/caller",
|
||||
null, new WorktreeRequest("fleetd-425-round4", null));
|
||||
|
||||
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
|
||||
assertNotNull(overlay, "overlayParity must have been called");
|
||||
assertEquals(List.of(s.profile() + ".mcp.json"), overlay.requested(),
|
||||
"the worktree must be provisioned for the SAME profile the member actually spawned "
|
||||
+ "on — under a rotating policy, dropping the PlacementDecision makes the "
|
||||
+ "second, spawn-time select() call disagree with the first, place()-time "
|
||||
+ "call, so the member ends up on a profile whose worktree (repoRoot/parity "
|
||||
+ "overlay) was built for a DIFFERENT profile (fleetd #425 rework, round 4)");
|
||||
}
|
||||
|
||||
/**
|
||||
* Reduce one {@code acquire(...)} attempt to a value comparable across the with-worktree and
|
||||
* without-worktree paths: the spawned profile name on success, or the thrown exception's class
|
||||
* and message on failure. Comparing THIS — instead of asserting "both spawn" or "both throw" as
|
||||
* a hardcoded direction — is what keeps {@link
|
||||
* #unqualifiedAcquireAgreesWithAndWithoutAWorktreeWhenTheOnlyProfileIsAtMaxLoad} valid whichever
|
||||
* way fleetd #435 eventually resolves whether an unqualified spawn should respect maxLoad.
|
||||
*/
|
||||
private static Object attemptAcquire(Supplier<MemberSession> call) {
|
||||
try {
|
||||
return "spawned:" + call.get().profile();
|
||||
} catch (RuntimeException e) {
|
||||
return "threw:" + e.getClass().getName() + ":" + e.getMessage();
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,369 @@
|
||||
# Plan — move the fleet back onto fleet01
|
||||
|
||||
**Goal.** Stop the fleet depending on a laptop that sleeps.
|
||||
|
||||
**Written** 2026-09-05. **Rewritten the same day** after the operator pointed out that fleet01 is a VM
|
||||
and already has the repo. They were right, and my first draft was wrong in an important way: this is
|
||||
not a stand-up. **The whole fleet already ran on fleet01 in August.** It was abandoned, not attempted.
|
||||
|
||||
Every fact below was measured on 2026-09-05. Where I did not measure something, the text says so.
|
||||
|
||||
---
|
||||
|
||||
## Status, re-measured 2026-09-10 — phases 1 to 3 are DONE, and §9 is out of date
|
||||
|
||||
This plan is prior art now, not a to-do list. Every number in this section came from a read-only
|
||||
survey over `ssh fleet01` on 2026-09-10. **Delete this section and the plan once fleet01 is the
|
||||
fleet's only daemon** — at that point the plan has been executed and stops being useful.
|
||||
|
||||
| Plan item | State on 2026-09-10 | Command that re-measures it |
|
||||
|---|---|---|
|
||||
| Phase 1, refresh + build | **done, then drifted.** Checkout is on `main` at `4887731`, **77 commits behind** `origin/main`, 0 ahead. `fleetd/target/fleetd.jar` exists, built 2026-09-10 02:10 UTC. So it was rebuilt, and main has moved since. | `git -C ~/LTMS/fleetd rev-list --count HEAD..origin/main` |
|
||||
| Phase 2, port the config | **done.** `fleetd/fleetd.yaml` exists on the host, with `bind`, `profiles`, `configReload`, `fleet`, `health`, `lifecycle`, `guard`, `memberCredentials`, `broker` and `coordinator` all present. | `test -f ~/LTMS/fleetd/fleetd/fleetd.yaml` |
|
||||
| Phase 3, supervision | **done.** `~/.config/systemd/user/fleetd.service` and `herdr.service` both exist, both `active` and `enabled`. `loginctl show-user ltms -p Linger` prints `yes`. Exactly 1 java process, so the double-daemon problem in the old `restart.sh` is not present. | `systemctl --user is-active fleetd herdr` |
|
||||
| Phase 4-5, reachable and a member proven | **partly.** `/healthz` answers 200 on loopback and reports `{"protocol":19,"version":"0.8.0"}`, matching the herdr pin. 1 herdr socket present. I did not spawn a member from here, so "a member on fleet01 opens a PR" is still unproven by me. | `curl -s http://127.0.0.1:8765/healthz` on the host |
|
||||
| Phase 6-7, cutover and reboot proof | **not done.** The Mac still runs its own daemon and is still this fleet's lead. Uptime on fleet01 is 2 weeks 2 days, so no reboot proof has been taken since the units were installed. | `uptime -p` on the host |
|
||||
| §9 "not moving the lead yet" | **out of date.** `claude` is installed at `/home/ltms/.local/bin/claude` and a fleet01 lead is live — it reaches this session over the coordination channel. So the headless-login blocker named in §9 is solved. | `ssh fleet01 'command -v claude'` |
|
||||
|
||||
### The profiles fleet01 actually offers a member
|
||||
|
||||
Measured from the live `fleetd.yaml` on the host, with `placement: weighted`:
|
||||
|
||||
| profile | kind | model | weight | maxLoad |
|
||||
|---|---|---|---|---|
|
||||
| `gx` | opencode | `gx/deepseek-v4-flash` | 100 | 2 |
|
||||
| `xf` | opencode | `opencode/mimo-v2.5-free` | 80 | 5 |
|
||||
| `local` | claude-code | `deepseek-v4-flash` | 10 | 2 |
|
||||
| `opus` | claude-code | `claude-opus-5` | 0 | 1 |
|
||||
|
||||
`weighted` spreads by ratio across every profile with a free slot, so an **unqualified** spawn on
|
||||
fleet01 lands on `gx` or `xf` almost every time. Pass `profile` explicitly there, as the canonical
|
||||
block already says.
|
||||
|
||||
### The PATH split on fleet01, which is not the one I expected
|
||||
|
||||
I went looking for a defect and found the opposite, so this is written down to stop the next
|
||||
session repeating the search.
|
||||
|
||||
`opencode` is installed at `/home/ltms/.opencode/bin/opencode`. Whether a shell can see it depends
|
||||
on which kind of shell it is:
|
||||
|
||||
```
|
||||
zsh -ic 'command -v opencode' -> /home/ltms/.opencode/bin/opencode (1 PATH entry)
|
||||
zsh -lc 'command -v opencode' -> nothing (0 PATH entries)
|
||||
```
|
||||
|
||||
So it is on the **interactive** PATH (`.zshrc`), not the login one. Two consequences, and they
|
||||
point in opposite directions:
|
||||
|
||||
- A herdr pane on Linux is a plain non-login interactive zsh, so a pane **can** launch `opencode`.
|
||||
fleetd types the launch command into the pane rather than exec'ing it, so `gx` and `xf` are not
|
||||
broken by this. I have not spawned one to confirm, so that is inference from the shell
|
||||
measurement plus the typing behaviour, not an end-to-end result.
|
||||
- The daemon itself runs `ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` —
|
||||
a **login** shell, on purpose, because credentials live in `.zprofile`. Its own PATH has 10
|
||||
entries and none contains `opencode`.
|
||||
|
||||
**The general shape: the login shell and the interactive shell see different PATHs, and which one
|
||||
matters depends on whether fleetd types a command or execs it.** Credentials live on the login
|
||||
side; `~/.opencode/bin` lives on the interactive side. Anything fleetd must exec itself is
|
||||
invisible to it if it lives only on the interactive PATH — and that is the mirror of the trap the
|
||||
login-shell `ExecStart` was added to fix.
|
||||
|
||||
---
|
||||
|
||||
---
|
||||
|
||||
## 1. My first draft was wrong — read this before the rest
|
||||
|
||||
I wrote "fleetd is not deployed on fleet01". That was wrong, and I got there by looking in one place
|
||||
and concluding about all of them. Three times:
|
||||
|
||||
| I checked | I concluded | What is actually true |
|
||||
|---|---|---|
|
||||
| `~/LTMS/claude-bridge` | "no repo checkout" | the repo is at `~/LTMS/fleetd` — the directory follows the **renamed** repo, and my own notes record that rename |
|
||||
| `~/.config/herdr/herdr.sock` | "herdr never ran" | the socket lives at `~/.config/herdr/sessions/fleet01/herdr.sock` — a **named session**, and its server log runs to Aug 28 |
|
||||
| both of the above | "fleetd is not deployed" | `~/LTMS/fleetd/fleetd-run/` holds `restart.sh`, `start-herdr.sh` and a `fleetd.out` from **Aug 24** |
|
||||
|
||||
The lesson is the one already written down here: enumerate one channel, conclude about all of them.
|
||||
A single-path check is not a survey.
|
||||
|
||||
---
|
||||
|
||||
## 2. What already worked on fleet01, proven from its own log
|
||||
|
||||
`~/LTMS/fleetd/fleetd-run/fleetd.out` covers 06:07 to 16:48 on 2026-08-24. It shows:
|
||||
|
||||
```
|
||||
4 distinct panes pane=c2504101-... w1:pC w1:pE
|
||||
4 worktrees created and removed /home/ltms/LTMS/.fleet-worktrees/{05f2a7-4,dd9f51-1,f498dd-2,fdb522-3}
|
||||
4 allow-list decisions "memberCredentials allow-list: pane w1:pC allowed 16 of 32 environment variables"
|
||||
AMQP on 127.0.0.1:5672 "AMQP connection recovered; cleared held replies for fresh redelivery"
|
||||
the full turn machinery SessionManager transitions, ReplyPushLoop, CompletionResolver fallback
|
||||
```
|
||||
|
||||
So on fleet01, already: fleetd listened, herdr made panes, members spawned into git worktrees, the
|
||||
credential allow-list fed them their environment, and the broker link was **loopback**.
|
||||
|
||||
That last point is the whole reason for this move. The AMQP resets in section 3 cannot happen to a
|
||||
loopback connection.
|
||||
|
||||
### The two traps I was going to design around are already solved there
|
||||
|
||||
My first draft named these as the biggest risks. Both were already handled in August:
|
||||
|
||||
**Trap A — systemd sources no login shell.** `restart.sh` already starts through one, and says why in
|
||||
its own comment:
|
||||
|
||||
```sh
|
||||
# 1. Start java from a LOGIN shell (zsh -lc). ~/.zprofile is where the credentials live, and a
|
||||
# non-login shell starts the daemon fine with an empty AI_GATEWAY_TOKEN -- a failure that
|
||||
# stays invisible until a member actually needs it.
|
||||
setsid zsh -lc "exec java -jar target/bridged.jar fleetd.yaml" < /dev/null > "$RUN/fleetd.out" 2>&1 &
|
||||
```
|
||||
|
||||
**Trap B — a Linux herdr pane is a plain zsh, so members get no credentials.** Not a problem, and not
|
||||
for the reason I assumed. Members do not inherit from the pane's shell profile — fleetd hands them an
|
||||
allow-list. The log proves it ran: `allowed 16 of 32 environment variables`, four times. The old
|
||||
`fleetd.yaml` has a `memberCredentials:` block that configures it.
|
||||
|
||||
**The headless pty trap is solved too.** `start-herdr.sh` carries the fix and the explanation:
|
||||
|
||||
```sh
|
||||
# Why the size matters: herdr creates each pane sized to the attached client's view. Started
|
||||
# under a pty with no winsize, the client reports 0x0, and every pane.split / workspace.create
|
||||
# then fails with "ghostty error -2" -- libghostty refusing a 0x0 surface.
|
||||
cat > /tmp/herdr-inner.sh <<'INNER'
|
||||
stty rows 50 cols 200 2>/dev/null || true
|
||||
exec herdr --session fleet01
|
||||
INNER
|
||||
setsid script -qfec /tmp/herdr-inner.sh /dev/null < /dev/null > /dev/null 2>&1 &
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Why we are moving
|
||||
|
||||
The daemon runs on a Mac laptop. On battery it idle-sleeps after **one minute**:
|
||||
|
||||
```
|
||||
pmset -g custom -> Battery Power: sleep 1
|
||||
AC Power: sleep 0
|
||||
```
|
||||
|
||||
Since the last restart: 16 `Connection reset` events and 6 ERROR lines, all on the AMQP link. I
|
||||
matched every one of the 16 to the nearest sleep or wake event in `pmset -g log`. Largest gap **55
|
||||
seconds**; most under 20. Not one reset lacked a nearby sleep or wake. The broker never sees a network
|
||||
fault — it sees the client stop sending heartbeats, then closes.
|
||||
|
||||
Every reconnect worked, so no message was lost. **The lost messages are not the problem.** The problem
|
||||
is that a member mid-turn freezes with the host, and a long worker turn with nobody typing is exactly
|
||||
the case that goes idle.
|
||||
|
||||
---
|
||||
|
||||
## 4. How the two hosts relate today
|
||||
|
||||
This answers "why is fleet01 related to this Mac at all?"
|
||||
|
||||
```
|
||||
Mac laptop fleet01 (KVM/QEMU guest)
|
||||
┌────────────────────────────┐ ┌──────────────────────┐
|
||||
│ Claude Code lead │ │ LavinMQ :5672 │
|
||||
│ fleetd 127.0.0.1:8765 │ AMQP over │ vhost /mac │
|
||||
│ herdr │──Tailscale────>│ vhost /fleet01 │
|
||||
│ members + worktrees │ utun4, 1280 │ │
|
||||
└────────────────────────────┘ │ (fleetd idle since │
|
||||
│ Aug 24) │
|
||||
└──────────────────────┘
|
||||
```
|
||||
|
||||
**Today fleet01 runs only the broker.** Everything else — daemon, herdr, members — is on the Mac. The
|
||||
single link between them is the Mac's fleetd opening AMQP to `10.10.20.13:5672` across Tailscale.
|
||||
|
||||
So the errors I reported were **the Mac's client dying when the Mac slept**, not fleet01 failing.
|
||||
fleet01 was healthy throughout: the container is up 11 days and its log shows a clean heartbeat
|
||||
timeout each time, which is what a broker sees when a client vanishes.
|
||||
|
||||
After the move that arrow becomes loopback and the whole class of problem is gone.
|
||||
|
||||
---
|
||||
|
||||
## 5. The shape we are restoring
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
subgraph MAC["Mac laptop (free to sleep)"]
|
||||
LEAD["Claude Code lead"]
|
||||
TUN["ssh -N -L"]
|
||||
end
|
||||
subgraph F01["fleet01 (KVM guest, always on)"]
|
||||
FD["fleetd<br/>127.0.0.1:8765"]
|
||||
HD["herdr --session fleet01"]
|
||||
WT["members<br/>~/LTMS/.fleet-worktrees"]
|
||||
MQ["LavinMQ<br/>127.0.0.1:5672"]
|
||||
end
|
||||
LEAD --> TUN
|
||||
TUN -->|"ssh over Tailscale"| FD
|
||||
FD -->|"unix socket"| HD
|
||||
HD --> WT
|
||||
FD -->|"AMQP, loopback"| MQ
|
||||
WT -->|"AMQP, loopback"| MQ
|
||||
```
|
||||
|
||||
*The lead stays on the Mac. Everything that must survive a sleep is already able to run on fleet01.*
|
||||
|
||||
Two constraints fix this shape:
|
||||
|
||||
1. **fleetd, herdr and the worktrees must share one filesystem.** The herdr link is a Unix socket plus
|
||||
absolute path strings. `herdr --remote` is terminal attach, not a transport.
|
||||
2. **fleetd fails fast on a non-loopback bind without token auth**, by its own design. So we do not
|
||||
expose `:8765` to the `10.10.20.0/24` LAN. An SSH tunnel keeps the bind on loopback and needs no
|
||||
new secret.
|
||||
|
||||
**Why the lead stays on the Mac for now.** A session not in a herdr pane resolves as `primary`, so a
|
||||
Mac-side session over the tunnel works with nothing new. What we give up: async ticket nudges type
|
||||
into the lead's pane, and fleetd cannot type into a pane on another host — so `wait:false` tickets
|
||||
stop nudging and I poll instead. Moving the lead as well is section 9; its hard part is
|
||||
authenticating Claude Code on a headless box, which has nothing to do with sleep and must not block
|
||||
this.
|
||||
|
||||
---
|
||||
|
||||
## 6. The actual gap
|
||||
|
||||
Everything below is what stands between "it ran in August" and "it runs supervised today".
|
||||
|
||||
| # | Gap | Measured state |
|
||||
|---|---|---|
|
||||
| 1 | **Checkout is stale** | branch `cb-634-ide-mcp`, HEAD `7655f1b` (2026-08-24), **290 commits behind** `origin/main`, 0 ahead |
|
||||
| 2 | **Never rebuilt after the rename** | no `target/` anywhere; the scripts still say `bridged.jar` and `cd bridged`, but the tree is now `fleetd/` and `fleetd.jar` |
|
||||
| 3 | **No live `fleetd.yaml`** | gitignored, so not in git. Three backups exist under `bridged/` — the newest is `fleetd.yaml.bak-cb634-pin`, 5512 bytes, and it is **clean of inline secrets** (0 inline passwords, 6 uses of `uriEnv`/`tokenEnv`) |
|
||||
| 4 | **No supervision** | `Linger=no`; **zero** systemd user unit files. Only the hand-rolled `restart.sh` / `start-herdr.sh` |
|
||||
| 5 | **Nothing running now** | no herdr process, no fleetd, no answer on `:8765/healthz` |
|
||||
|
||||
Untracked files in the checkout: `.idea/`, `fleetd-run/`, `docs/CB-634-Worker-IDE-Worktree.md`, and the
|
||||
three yaml backups. All are **untracked, none modified**, and `7655f1b` is already an ancestor of
|
||||
`origin/main` — so nothing is lost by updating the branch. Keep `fleetd-run/` and the backups; they
|
||||
are the prior art this plan is built on.
|
||||
|
||||
### The old config's keys, which tell us what to port
|
||||
|
||||
```
|
||||
bind: herdrSocket: /home/ltms/.config/herdr/sessions/fleet01/herdr.sock
|
||||
profiles: placement: weighted
|
||||
configReload: fleet: health:
|
||||
lifecycle: guard: worktreeRoot: /home/ltms/LTMS/.fleet-worktrees
|
||||
memberCredentials: broker:
|
||||
```
|
||||
|
||||
`herdrSocket` already points at the **named-session** path, and `memberCredentials` is already
|
||||
configured. Those two are what made members work.
|
||||
|
||||
### What the repo already has for this
|
||||
|
||||
`deploy/fleetd.service` exists and is written for Linux. Three lines need fleet01's real paths:
|
||||
`ExecStart` names `/usr/lib/jvm/temurin-25-jdk/bin/java` (fleet01 has `/usr/bin/java`), the `PATH`
|
||||
names `/usr/share/maven/bin` (fleet01 has `/usr/bin/mvn`), and `WorkingDirectory` assumes
|
||||
`%h/src/claude-bridge`. It also declares `After=herdr.service` — **and no `herdr.service` exists in
|
||||
`deploy/`**. Writing that unit, from `start-herdr.sh`, is the one genuinely new piece of code here.
|
||||
|
||||
---
|
||||
|
||||
## 7. Phases
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
P1["1. Refresh<br/>update + build"]
|
||||
P2["2. Config<br/>port fleetd.yaml"]
|
||||
P3["3. Supervise<br/>linger + 2 units"]
|
||||
P4["4. Reachability<br/>tunnel, primary"]
|
||||
P5["5. Prove a member"]
|
||||
P6["6. Cutover"]
|
||||
P7["7. Reboot proof"]
|
||||
P1 --> P2 --> P3 --> P4 --> P5 --> P6 --> P7
|
||||
```
|
||||
|
||||
**Phase 1 — refresh the checkout.** Update to `origin/main` (290 commits). Keep the untracked
|
||||
`fleetd-run/` and the yaml backups. Then `mvn clean install`, **unpiped** — a pipe hides a failure
|
||||
behind a zero exit. *Check:* `fleetd/target/fleetd.jar` exists and the suite is green.
|
||||
|
||||
**Phase 2 — port the config.** Write `fleetd/fleetd.yaml` from `bridged/fleetd.yaml.bak-cb634-pin`,
|
||||
updating it for the rename and 290 commits of config changes. Diff its keys against
|
||||
`fleetd/fleetd.example.yaml` on current main, key by key, and say what changed. *Check:* the daemon
|
||||
starts and `journalctl ... | grep 'startup secret'` reports **no** `MISSING`.
|
||||
|
||||
**Phase 3 — supervision.** This is the part that never existed. `loginctl enable-linger ltms`; write
|
||||
`deploy/herdr.service` from `start-herdr.sh`, keeping the `stty` sizing; fix the three path lines in
|
||||
`deploy/fleetd.service` and keep the login-shell `ExecStart`; install both under
|
||||
`~/.config/systemd/user/`. Secrets go in a `systemctl --user edit` drop-in or a 0600
|
||||
`EnvironmentFile`, never in the committed unit. *Check:* log out of every ssh session, log back in,
|
||||
and confirm the socket and the daemon are still there. That is what lingering is for, and it is the
|
||||
check people skip.
|
||||
|
||||
**Phase 4 — reachability.** `ssh -N -L <port>:127.0.0.1:8765 fleet01` from the Mac. Use a **different
|
||||
local port** for the first test so the Mac's own daemon on `:8765` is untouched and the whole test is
|
||||
reversible. Point the lead's `.mcp.json` at it — that file is `--skip-worktree` and must never be
|
||||
committed. Wrap the tunnel in `autossh` or a launchd `KeepAlive`, because the Mac still sleeps.
|
||||
*Check:* `fleet_whoami` answers `primary`. If it answers `worker`, the `fleet.leaders.*.tab` pin does
|
||||
not match — a known demotion, not a network fault.
|
||||
|
||||
**Phase 5 — prove a member.** `healthz` can be green while every spawn fails, so only a real spawn
|
||||
proves the herdr link. Spawn one member, then give it a real unit ending in a pushed PR — the push is
|
||||
what proves `WORKER_GITEA_TOKEN` resolved. Confirm the allow-list line still appears
|
||||
(`allowed N of M environment variables`) and that N is what you expect. *Check:* a member on fleet01
|
||||
opens a PR.
|
||||
|
||||
**Phase 6 — cutover.** Drain the Mac fleet properly first: `fleet_list`, `fleet_poll` anything still
|
||||
wanted, then `fleet_stop` each member — a restart drops in-flight tickets and a member's report is
|
||||
gone with its ticket. Then stop the Mac's launchd agent. This is the migration itself, not a change to
|
||||
the Mac's settings, and it is reversible in one command.
|
||||
|
||||
**Phase 7 — reboot proof.** Reboot fleet01. Without touching anything: socket present, healthz
|
||||
answering, `fleet_whoami` still `primary`, one spawn works. Until this passes, "supervised" is a claim.
|
||||
|
||||
---
|
||||
|
||||
## 8. Risks
|
||||
|
||||
| Risk | Why it bites | What this plan does |
|
||||
|---|---|---|
|
||||
| **290 commits of config drift** | the old yaml predates the rename and much else; a silently defaulted key turns a feature off with no error | phase 2 diffs key-by-key against current `fleetd.example.yaml` |
|
||||
| **A new config key gets silently dropped** | `FleetConfig`'s back-compat constructor ladder can absorb an arity change, so a new key compiles and is defaulted away | separate ticket already in flight; matters most here because fleet01 gets a hand-edited yaml |
|
||||
| **Nothing supervises herdr** | `deploy/fleetd.service` depends on a unit that does not exist | phase 3 writes it from the working script; phase 7 proves it |
|
||||
| **Linger left off** | everything dies at logout and looks fine until then | phase 3, checked by logging out |
|
||||
| **Wrong JDK/Maven path in the unit** | fleetd propagates its PATH to every member, so a bad PATH means no member can build | three lines fixed in phase 3, proven by phase 5 |
|
||||
| **Headless pty with no winsize** | `ghostty error -2`, reported three steps later as a spawn failure | the `stty` fix is carried into `herdr.service` |
|
||||
| **Port 8765 collides during the test** | both daemons want the same local port | phase 4 uses a different local port first |
|
||||
| **Tunnel dies when the Mac sleeps** | same sleep, far smaller blast radius — it interrupts my session, not members | `autossh`/launchd `KeepAlive` |
|
||||
| **Lead demoted to worker** | the tab pin no longer matches | phase 4's check is `fleet_whoami` |
|
||||
| **Upgrading herdr** | 0.8.0 is **protocol 19**, pinned on purpose — 0.8.2 is protocol 20 and fleetd has **no version handshake** | do not upgrade herdr during this work; both hosts measured at 0.8.0 today |
|
||||
| **`placement: tab` headless** | fails for the same 0x0 reason as `ghostty error -2` | fleet01 must keep `placement: pane`, as its August config did |
|
||||
| **Profile launch settings are deferred** | editing `placement:` and waiting for the 10s config watch does nothing — the launcher holds a startup snapshot | restart the daemon after those keys, do not wait for the reload |
|
||||
|
||||
---
|
||||
|
||||
## 9. Deliberately not doing
|
||||
|
||||
- **No changes to the Mac's power settings or host config.** The point is to stop depending on it.
|
||||
- **Not touching the leftover `bridged-lavinmq` container** on the Mac. It is unused and harmless.
|
||||
- **Not moving the broker.** It is already on fleet01 and already the durable one. After the move its
|
||||
connection becomes loopback, which is the fix.
|
||||
- **Not building a second fleet.** This is a move. Two daemons on one herdr session kill each other's
|
||||
members.
|
||||
- **Not moving the lead yet.** That needs Claude Code authenticated on a headless Ubuntu box and a
|
||||
`fleet.leaders.*.tab` pin on its pane. It buys back pane nudges. It has nothing to do with sleep, so
|
||||
it must not hold up phases 1–7. Note that fleet01 already carries an `opus` profile defined purely so the lead slot resolves; it cannot spawn until someone runs `claude` and completes `/login` on the host.
|
||||
|
||||
---
|
||||
|
||||
## 10. Open questions for the operator
|
||||
|
||||
1. **vhost** — keep `/mac`, or rename now the fleet is not on the Mac? Renaming loses the existing
|
||||
queues. (The old fleet01 config used its own; phase 2 must settle which this fleet owns.)
|
||||
2. **Fallback week** after cutover, or stop the Mac daemon for good?
|
||||
3. **Delete or keep the stale `cb-634-ide-mcp` branch** on fleet01 once the checkout is updated? Its
|
||||
tip is already in main, so nothing is lost either way.
|
||||
|
||||
The repo path question from the first draft is answered: **`/home/ltms/LTMS/fleetd`**, which already
|
||||
exists. `deploy/fleetd.service` should be pointed there rather than the reverse.
|
||||
+217
-54
@@ -5,7 +5,7 @@
|
||||
# A merge is not a deployment: the running daemon holds the jar it was started with, so code merged
|
||||
# to main does nothing until this runs. See CLAUDE.md -> "Redeploying the daemon".
|
||||
#
|
||||
# This script exists to turn six remembered traps into one auditable command:
|
||||
# This script exists to turn eight remembered traps into one auditable command:
|
||||
#
|
||||
# 1. A piped `mvn` hides BUILD FAILURE behind a zero exit, so the build here is never piped.
|
||||
# 2. The daemon must start from a LOGIN shell, or the tokens it hands to members are empty:
|
||||
@@ -27,6 +27,17 @@
|
||||
# restart of the OLD jar. So this script detects whether the agent is loaded and, only then,
|
||||
# swaps `kill` + manual `nohup` for `launchctl unload`/`load` — the one supervisor in control
|
||||
# at any moment is whichever one you asked to act, never both.
|
||||
# 7. fleetd #492 — a systemd --user unit is a THIRD possible supervisor (seen on a second host):
|
||||
# Restart=on-failure treats this JVM's SIGTERM exit code (143, per CB-594 above) as a failure
|
||||
# too, so a bare `kill` there would race systemd's own restart of the OLD jar exactly like
|
||||
# launchd would. This script now tells launchd, systemd, and "genuinely unsupervised" apart as
|
||||
# three different answers, drives whichever one it finds through its own control plane
|
||||
# (`launchctl` / `systemctl --user`), and REFUSES outright — never falls back to `kill` — when
|
||||
# it finds a supervision signal it cannot map to exactly one of the two it knows how to drive.
|
||||
# A wrong guess here is how two daemons end up running against one herdr session.
|
||||
# 8. fleetd #492 — a post-restart check counts running fleetd processes and fails the whole run if
|
||||
# more than one is alive. That is the one thing none of the checks above (healthz 200, jar id,
|
||||
# the fresh "listening" line) can see: every one of them is satisfied by EITHER daemon.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/redeploy-fleetd.sh # build, confirm, restart, verify
|
||||
@@ -55,13 +66,18 @@ HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start
|
||||
LAUNCHD_LABEL='dev.ltms.fleetd'
|
||||
LAUNCHD_PLIST="$HOME/Library/LaunchAgents/$LAUNCHD_LABEL.plist"
|
||||
|
||||
# fleetd #492: the systemd --user unit this script must not fight with either (see trap 7 above).
|
||||
# Measured on the second host: `systemctl --user cat fleetd` names the unit "fleetd" (not
|
||||
# "dev.ltms.fleetd" — systemd user units here are not namespaced the way the launchd label is).
|
||||
SYSTEMD_UNIT='fleetd'
|
||||
|
||||
DO_BUILD=1; ASSUME_YES=0; CHECK_ONLY=0
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--yes|-y) ASSUME_YES=1 ;;
|
||||
--no-build) DO_BUILD=0 ;;
|
||||
--check) CHECK_ONLY=1 ;;
|
||||
-h|--help) sed -n '3,37p' "${BASH_SOURCE[0]}"; exit 0 ;;
|
||||
-h|--help) sed -n '3,48p' "${BASH_SOURCE[0]}"; exit 0 ;;
|
||||
*) echo "unknown option: $arg (try --help)" >&2; exit 2 ;;
|
||||
esac
|
||||
done
|
||||
@@ -79,6 +95,91 @@ running_pid() { pgrep -f "$PATTERN" || true; }
|
||||
launchd_installed() { [ -f "$LAUNCHD_PLIST" ]; }
|
||||
launchd_loaded() { launchctl list "$LAUNCHD_LABEL" >/dev/null 2>&1; }
|
||||
|
||||
# fleetd #492: same two questions for systemd --user. Kept as separate, overridable functions
|
||||
# (never an inline `systemctl` call at each use site) so a test on a box with no systemd at all
|
||||
# (this repo is developed on macOS) can substitute each one independently — the same seam
|
||||
# launchd_installed/launchd_loaded above already use.
|
||||
#
|
||||
# "installed": a unit FILE by this name exists, regardless of its current state — the systemd
|
||||
# analogue of the plist file existing on disk. `list-unit-files` reads unit definitions without
|
||||
# depending on runtime state, so this stays read-only and safe under --check.
|
||||
systemd_installed() {
|
||||
command -v systemctl >/dev/null 2>&1 \
|
||||
&& systemctl --user list-unit-files "$SYSTEMD_UNIT.service" --no-legend 2>/dev/null | grep -q .
|
||||
}
|
||||
# "loaded": systemd currently supervises this unit as an active job — the systemd analogue of
|
||||
# `launchctl list <label>` succeeding. Measured on the second host: `systemctl --user is-active
|
||||
# fleetd` -> "active".
|
||||
systemd_loaded() {
|
||||
command -v systemctl >/dev/null 2>&1 && systemctl --user is-active "$SYSTEMD_UNIT" >/dev/null 2>&1
|
||||
}
|
||||
|
||||
# fleetd #492: three real answers, not two — launchd, systemd, or genuinely unsupervised — plus a
|
||||
# fourth, "ambiguous", for the one case this script cannot tell apart: both signals firing at once.
|
||||
# That is exactly "I cannot tell who supervises this process", and guessing wrong here is how two
|
||||
# daemons end up running against one herdr session (see trap 7 in the header). Pure and
|
||||
# side-effect-free: reads the two probes above and decides — never mutates anything, so it is safe
|
||||
# under --check and testable by overriding launchd_loaded/systemd_loaded after sourcing.
|
||||
detect_supervisor() {
|
||||
local ld=0 sd=0
|
||||
launchd_loaded && ld=1
|
||||
systemd_loaded && sd=1
|
||||
if [ "$ld" = 1 ] && [ "$sd" = 1 ]; then
|
||||
echo "ambiguous"
|
||||
elif [ "$ld" = 1 ]; then
|
||||
echo "launchd"
|
||||
elif [ "$sd" = 1 ]; then
|
||||
echo "systemd"
|
||||
else
|
||||
echo "none"
|
||||
fi
|
||||
}
|
||||
|
||||
# fleetd #492: turns anything detect_supervisor returns that is NOT exactly one of the two
|
||||
# supervisors this script knows how to drive into a die() — never a fall-through to the `kill`
|
||||
# path. Kept as its own function so a test can call it directly (in a subshell, since it die()s)
|
||||
# without running the whole report-state flow or needing a real launchd/systemd.
|
||||
require_drivable_supervisor() {
|
||||
local kind="$1"
|
||||
case "$kind" in
|
||||
launchd|systemd|none) ;;
|
||||
ambiguous)
|
||||
die "both launchd ($LAUNCHD_LABEL) and systemd --user ($SYSTEMD_UNIT) report themselves as
|
||||
loaded for this daemon at the same time. This script cannot tell which one actually
|
||||
supervises the running process, and driving either alone risks the OTHER reviving the
|
||||
OLD jar out from under it — the exact failure this ticket (fleetd #492) exists to
|
||||
prevent. Stop one of the two supervisors by hand, confirm only one remains loaded, then
|
||||
rerun." ;;
|
||||
*)
|
||||
die "detect_supervisor returned an unrecognized value '$kind' — refusing to guess which
|
||||
supervisor, if any, controls this daemon." ;;
|
||||
esac
|
||||
}
|
||||
|
||||
# fleetd #492: the exact symptom a racing supervisor produces — count how many fleetd processes are
|
||||
# alive right now. Takes the pid list as a parameter (rather than calling running_pid() itself) so a
|
||||
# test can pass a canned two-line string without a real second process running. Pure except for the
|
||||
# die() in assert_single_daemon below.
|
||||
count_daemon_pids() {
|
||||
local pids="$1"
|
||||
if [ -z "$pids" ]; then
|
||||
echo 0
|
||||
else
|
||||
printf '%s\n' "$pids" | grep -c .
|
||||
fi
|
||||
}
|
||||
assert_single_daemon() {
|
||||
local pids="$1" count
|
||||
count="$(count_daemon_pids "$pids")"
|
||||
if [ "$count" -gt 1 ]; then
|
||||
die "more than one fleetd process is running after this restart (pids: $(printf '%s' "$pids" | tr '\n' ' ')).
|
||||
This is the exact failure a racing supervisor produces: the OLD jar was revived by its
|
||||
supervisor while this script started a NEW copy. Two daemons on one herdr session kill
|
||||
each other's members. Investigate with 'pgrep -f \"$PATTERN\"' and stop the wrong one by
|
||||
hand — do not assume either pid is the one you want."
|
||||
fi
|
||||
}
|
||||
|
||||
# CB-600: the script computes its own log path from where it sits on disk (REPO, above); the
|
||||
# plist hard-codes an absolute StandardOutPath. Nothing forced the two to agree — if this script
|
||||
# were ever run from a checkout other than the one the loaded plist names, launchd would start and
|
||||
@@ -182,25 +283,45 @@ fi
|
||||
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
|
||||
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
|
||||
|
||||
# CB-594: supervision state. Installed and loaded are different facts — a copied-but-never-loaded
|
||||
# plist supervises nothing, and a loaded label with no file backing it (rare, but possible after an
|
||||
# edited/moved plist) is still what launchd will act on.
|
||||
# CB-594 / fleetd #492: supervision state. Installed and loaded are different facts — a
|
||||
# copied-but-never-loaded plist (or an unloaded systemd unit) supervises nothing, and a loaded
|
||||
# label/unit with no file backing it is still what its supervisor will act on.
|
||||
if launchd_installed; then
|
||||
ok "launchd agent installed: $LAUNCHD_PLIST"
|
||||
else
|
||||
warn "launchd agent NOT installed (no supervision — a crash will not restart the daemon)."
|
||||
warn "launchd agent NOT installed."
|
||||
fi
|
||||
SUPERVISED=0
|
||||
if launchd_loaded; then
|
||||
SUPERVISED=1
|
||||
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
|
||||
# CB-600: fail loudly here, before ANY other check runs, if this script and the loaded plist
|
||||
# would read different log files — every check after this point is worthless otherwise.
|
||||
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
|
||||
if systemd_installed; then
|
||||
ok "systemd --user unit installed: $SYSTEMD_UNIT"
|
||||
else
|
||||
warn "launchd agent not loaded — this script is the only thing that will restart the daemon."
|
||||
warn "systemd --user unit NOT installed ($SYSTEMD_UNIT)."
|
||||
fi
|
||||
|
||||
# fleetd #492: decide which of the two (if either) actually supervises this daemon, and refuse
|
||||
# outright — before touching anything — if that cannot be told apart (see require_drivable_
|
||||
# supervisor above). --check reaches this same line, so a host with an undrivable supervisor is
|
||||
# reported as a failure even in --check, without ever reaching the build/stop/start steps.
|
||||
SUPERVISOR_KIND="$(detect_supervisor)"
|
||||
require_drivable_supervisor "$SUPERVISOR_KIND"
|
||||
ok "supervisor detected: $SUPERVISOR_KIND"
|
||||
SUPERVISED=0
|
||||
case "$SUPERVISOR_KIND" in
|
||||
launchd)
|
||||
SUPERVISED=1
|
||||
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
|
||||
# CB-600: fail loudly here, before ANY other check runs, if this script and the loaded plist
|
||||
# would read different log files — every check after this point is worthless otherwise.
|
||||
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
|
||||
;;
|
||||
systemd)
|
||||
SUPERVISED=1
|
||||
ok "systemd --user unit active ($SYSTEMD_UNIT) — systemd supervises this daemon"
|
||||
;;
|
||||
none)
|
||||
warn "no supervisor loaded — this script is the only thing that will restart the daemon."
|
||||
;;
|
||||
esac
|
||||
|
||||
# The trap with no log line. Checked in a LOGIN shell, because that is how the daemon is started
|
||||
# below. Never prints the value — only whether it resolved.
|
||||
if zsh -lc '[ -n "${WORKER_GITEA_TOKEN:-}" ]' 2>/dev/null; then
|
||||
@@ -281,25 +402,39 @@ fi
|
||||
|
||||
# ------------------------------------------------------------------ stop
|
||||
#
|
||||
# CB-594: when SUPERVISED, launchd owns the stop — never a raw `kill` here. A bare SIGTERM makes
|
||||
# this JVM exit 143 even with its shutdown hook running to completion (verified separately: a
|
||||
# throwaway Java process with an equivalent shutdown hook, sent SIGTERM from a login shell that
|
||||
# could `wait` on it directly, reported exit code 143 every time — never 0). launchd's
|
||||
# KeepAlive.SuccessfulExit=false treats any nonzero exit as a crash and restarts the OLD jar,
|
||||
# which would race this script's own restart of the NEW one. `launchctl unload` avoids that race
|
||||
# by deregistering the job first, so no KeepAlive is left armed when the process actually stops.
|
||||
# CB-594 / fleetd #492: when SUPERVISED, the supervisor owns the stop — never a raw `kill` here. A
|
||||
# bare SIGTERM makes this JVM exit 143 even with its shutdown hook running to completion (verified
|
||||
# separately: a throwaway Java process with an equivalent shutdown hook, sent SIGTERM from a login
|
||||
# shell that could `wait` on it directly, reported exit code 143 every time — never 0). launchd's
|
||||
# KeepAlive.SuccessfulExit=false and systemd's Restart=on-failure both treat any nonzero exit as a
|
||||
# crash and restart the OLD jar, which would race this script's own restart of the NEW one.
|
||||
# `launchctl unload` avoids that race by deregistering the job first, so no KeepAlive is left
|
||||
# armed when the process actually stops. `systemctl --user stop` needs no such dance: unlike
|
||||
# KeepAlive, systemd's Restart= does not fire on a deliberate stop, only on an unexpected exit of
|
||||
# an active unit.
|
||||
|
||||
if [ -n "$OLD_PID" ]; then
|
||||
say "stop"
|
||||
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)" # verify a FRESH line appears later
|
||||
if [ "$SUPERVISED" = 1 ]; then
|
||||
echo " supervision is ON: using 'launchctl unload' (not kill) so launchd's own KeepAlive"
|
||||
echo " cannot restart the OLD jar out from under this script — see the CB-594 comment above."
|
||||
launchctl unload -w "$LAUNCHD_PLIST" \
|
||||
|| die "launchctl unload failed — the daemon may still be under supervision; investigate before retrying"
|
||||
else
|
||||
kill "$OLD_PID"
|
||||
fi
|
||||
case "$SUPERVISOR_KIND" in
|
||||
launchd)
|
||||
echo " supervision is ON (launchd): using 'launchctl unload' (not kill) so launchd's own"
|
||||
echo " KeepAlive cannot restart the OLD jar out from under this script — see the CB-594"
|
||||
echo " comment above."
|
||||
launchctl unload -w "$LAUNCHD_PLIST" \
|
||||
|| die "launchctl unload failed — the daemon may still be under supervision; investigate before retrying"
|
||||
;;
|
||||
systemd)
|
||||
echo " supervision is ON (systemd --user): using 'systemctl --user stop' (not kill) so"
|
||||
echo " systemd's own Restart=on-failure cannot restart the OLD jar out from under this"
|
||||
echo " script — see the fleetd #492 comment above."
|
||||
systemctl --user stop "$SYSTEMD_UNIT" \
|
||||
|| die "'systemctl --user stop $SYSTEMD_UNIT' failed — the daemon may still be under supervision; investigate before retrying"
|
||||
;;
|
||||
none)
|
||||
kill "$OLD_PID"
|
||||
;;
|
||||
esac
|
||||
for _ in $(seq "$STOP_WAIT"); do
|
||||
[ -z "$(running_pid)" ] && break
|
||||
sleep 1
|
||||
@@ -310,13 +445,20 @@ if [ -n "$OLD_PID" ]; then
|
||||
leave worktrees and panes behind. Investigate, then kill -9 by hand if you accept that."
|
||||
fi
|
||||
ok "pid $OLD_PID exited"
|
||||
elif [ "$SUPERVISED" = 1 ]; then
|
||||
elif [ "$SUPERVISOR_KIND" = "launchd" ]; then
|
||||
# Loaded but not currently running (e.g. throttled after a crash loop). Unload it anyway so the
|
||||
# start step below does a clean load, never a load stacked on an already-loaded label.
|
||||
say "stop"
|
||||
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
|
||||
launchctl unload -w "$LAUNCHD_PLIST" 2>/dev/null || true
|
||||
ok "launchd agent unloaded (was already not running)"
|
||||
elif [ "$SUPERVISOR_KIND" = "systemd" ]; then
|
||||
# Same case for systemd: the unit is known/active-capable but not currently running. `stop` on an
|
||||
# already-stopped unit is a harmless no-op — kept for symmetry with the launchd branch above.
|
||||
say "stop"
|
||||
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
|
||||
systemctl --user stop "$SYSTEMD_UNIT" 2>/dev/null || true
|
||||
ok "systemd --user unit stopped (was already not running)"
|
||||
else
|
||||
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
|
||||
fi
|
||||
@@ -324,34 +466,48 @@ fi
|
||||
# ------------------------------------------------------------------ start
|
||||
# Unsupervised: login shell (zsh -l) is what puts the secrets on the daemon's environment, and cwd
|
||||
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and target/ relative to it.
|
||||
# Supervised: launchd does both — deploy/dev.ltms.fleetd.plist points ProgramArguments at
|
||||
# Supervised (launchd): launchd does both — deploy/dev.ltms.fleetd.plist points ProgramArguments at
|
||||
# scripts/fleetd-launchd-wrapper.sh (CB-594), which is what execs the login shell in launchd's
|
||||
# place, and WorkingDirectory in the plist already pins fleetd/.
|
||||
# Supervised (systemd --user): the unit does both too — measured on the second host, ExecStart is
|
||||
# `/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` (a login shell, same reason as
|
||||
# above) and WorkingDirectory is already pinned to fleetd/.
|
||||
|
||||
say "start"
|
||||
if [ "$SUPERVISED" = 1 ]; then
|
||||
echo " supervision is ON: using 'launchctl load' so launchd starts and keeps supervising this"
|
||||
echo " process, instead of a manual nohup that launchd would know nothing about."
|
||||
# CB-600: 'launchctl unload -w' above already persisted Disabled=true for this label. A load -w
|
||||
# that succeeds clears it; a load -w that FAILS leaves the agent both stopped and disabled — worse
|
||||
# than before this script ran, because a later reboot or login will not bring it back either. One
|
||||
# retry covers a transient race (e.g. launchd not yet fully done deregistering); if it still fails,
|
||||
# die with the exact recovery command rather than a bare "failed".
|
||||
if ! launchctl load -w "$LAUNCHD_PLIST" 2>/dev/null; then
|
||||
warn "launchctl load failed on the first attempt — retrying once after a short pause"
|
||||
sleep 2
|
||||
launchctl load -w "$LAUNCHD_PLIST" || die "launchctl load failed twice.
|
||||
The agent is now STOPPED and DISABLED — it will NOT come back on its own, not even after a
|
||||
reboot or login, because 'launchctl unload -w' above persisted Disabled=true and load -w
|
||||
never got the chance to clear it. Recover with:
|
||||
launchctl load -w \"$LAUNCHD_PLIST\"
|
||||
If that still fails, check 'launchctl list $LAUNCHD_LABEL', validate the plist with
|
||||
'plutil -lint \"$LAUNCHD_PLIST\"', and check $OUT before assuming a retry will succeed."
|
||||
fi
|
||||
else
|
||||
# Absolute jar path so `ps` names which checkout is running.
|
||||
( cd "$MODULE" && zsh -lc "nohup java -jar '$JAR' >> fleetd.out 2>&1 &" )
|
||||
fi
|
||||
case "$SUPERVISOR_KIND" in
|
||||
launchd)
|
||||
echo " supervision is ON (launchd): using 'launchctl load' so launchd starts and keeps"
|
||||
echo " supervising this process, instead of a manual nohup that launchd would know nothing"
|
||||
echo " about."
|
||||
# CB-600: 'launchctl unload -w' above already persisted Disabled=true for this label. A load -w
|
||||
# that succeeds clears it; a load -w that FAILS leaves the agent both stopped and disabled — worse
|
||||
# than before this script ran, because a later reboot or login will not bring it back either. One
|
||||
# retry covers a transient race (e.g. launchd not yet fully done deregistering); if it still fails,
|
||||
# die with the exact recovery command rather than a bare "failed".
|
||||
if ! launchctl load -w "$LAUNCHD_PLIST" 2>/dev/null; then
|
||||
warn "launchctl load failed on the first attempt — retrying once after a short pause"
|
||||
sleep 2
|
||||
launchctl load -w "$LAUNCHD_PLIST" || die "launchctl load failed twice.
|
||||
The agent is now STOPPED and DISABLED — it will NOT come back on its own, not even after a
|
||||
reboot or login, because 'launchctl unload -w' above persisted Disabled=true and load -w
|
||||
never got the chance to clear it. Recover with:
|
||||
launchctl load -w \"$LAUNCHD_PLIST\"
|
||||
If that still fails, check 'launchctl list $LAUNCHD_LABEL', validate the plist with
|
||||
'plutil -lint \"$LAUNCHD_PLIST\"', and check $OUT before assuming a retry will succeed."
|
||||
fi
|
||||
;;
|
||||
systemd)
|
||||
echo " supervision is ON (systemd --user): using 'systemctl --user start' so systemd starts"
|
||||
echo " and keeps supervising this process, instead of a manual nohup it would know nothing"
|
||||
echo " about."
|
||||
systemctl --user start "$SYSTEMD_UNIT" || die "'systemctl --user start $SYSTEMD_UNIT' failed.
|
||||
Check 'systemctl --user status $SYSTEMD_UNIT' and $OUT before assuming a retry will succeed."
|
||||
;;
|
||||
none)
|
||||
# Absolute jar path so `ps` names which checkout is running.
|
||||
( cd "$MODULE" && zsh -lc "nohup java -jar '$JAR' >> fleetd.out 2>&1 &" )
|
||||
;;
|
||||
esac
|
||||
|
||||
for _ in $(seq 10); do
|
||||
NEW_PID="$(running_pid)"
|
||||
@@ -411,6 +567,13 @@ FRESH_LOG="$(mktemp -t fleetd-fresh-log)"
|
||||
trap 'rm -f "$FRESH_LOG"' EXIT
|
||||
tail -n "+$((RESTART_MARK + 1))" "$OUT" > "$FRESH_LOG" 2>/dev/null || true
|
||||
classify_amqp_connection_errors "$FRESH_LOG"
|
||||
|
||||
# fleetd #492: checked here, after healthz and the fresh-log check have both had time to run, so a
|
||||
# supervisor that revives the OLD jar a few seconds late is caught too. Every check above (healthz
|
||||
# 200, jar id, the fresh 'listening' line) is satisfied by EITHER daemon if two are alive — this is
|
||||
# the only one that can tell.
|
||||
assert_single_daemon "$(running_pid)"
|
||||
|
||||
say "result"
|
||||
ok "pid $NEW_PID, jar $(jar_id)"
|
||||
if [ "$REDEPLOY_ERROR_COUNT" -eq 0 ]; then
|
||||
|
||||
@@ -25,6 +25,78 @@ classify_fixture() {
|
||||
classify_amqp_connection_errors "$TMP/$name"
|
||||
}
|
||||
|
||||
|
||||
# fleetd #492 — supervisor detection. Detect_supervisor() reads launchd_loaded/systemd_loaded, so
|
||||
# each test overrides BOTH pairs (installed + loaded) explicitly, rather than relying on either
|
||||
# being naturally absent: this machine may itself be running a real fleetd under launchd right now
|
||||
# (see CLAUDE.md/MEMORY.md — launchd supervision has been live here since 2026-08-26), so leaving
|
||||
# launchd_loaded unmocked in a "systemd only" test would silently read this host's own live state
|
||||
# instead of the fixture.
|
||||
test_detect_supervisor_launchd_only() {
|
||||
launchd_installed() { return 0; }
|
||||
launchd_loaded() { return 0; }
|
||||
systemd_installed() { return 1; }
|
||||
systemd_loaded() { return 1; }
|
||||
assert_equals "launchd" "$(detect_supervisor)" "launchd-only detection"
|
||||
}
|
||||
|
||||
test_detect_supervisor_systemd_only() {
|
||||
launchd_installed() { return 1; }
|
||||
launchd_loaded() { return 1; }
|
||||
systemd_installed() { return 0; }
|
||||
systemd_loaded() { return 0; }
|
||||
assert_equals "systemd" "$(detect_supervisor)" "systemd-only detection"
|
||||
}
|
||||
|
||||
test_detect_supervisor_none() {
|
||||
launchd_installed() { return 1; }
|
||||
launchd_loaded() { return 1; }
|
||||
systemd_installed() { return 1; }
|
||||
systemd_loaded() { return 1; }
|
||||
assert_equals "none" "$(detect_supervisor)" "unsupervised detection"
|
||||
}
|
||||
|
||||
# The heart of the ticket: a supervisor this script cannot drive must refuse, never fall through to
|
||||
# `kill`. require_drivable_supervisor die()s, so it is invoked inside a command substitution — that
|
||||
# forks a subshell, so its exit() only ends the subshell and this test script keeps running under
|
||||
# `set -e`.
|
||||
test_require_drivable_supervisor_refuses_ambiguous() {
|
||||
local output rc=0
|
||||
output="$(require_drivable_supervisor "ambiguous" 2>&1)" || rc=$?
|
||||
[ "$rc" -ne 0 ] || fail "require_drivable_supervisor accepted an ambiguous (undrivable) supervisor"
|
||||
printf '%s' "$output" | grep -qF "$LAUNCHD_LABEL" \
|
||||
|| fail "refusal message does not name the launchd label it found"
|
||||
printf '%s' "$output" | grep -qF "$SYSTEMD_UNIT" \
|
||||
|| fail "refusal message does not name the systemd unit it found"
|
||||
}
|
||||
|
||||
test_require_drivable_supervisor_accepts_known_kinds() {
|
||||
require_drivable_supervisor "launchd" || fail "refused a drivable launchd supervisor"
|
||||
require_drivable_supervisor "systemd" || fail "refused a drivable systemd supervisor"
|
||||
require_drivable_supervisor "none" || fail "refused the unsupervised case"
|
||||
}
|
||||
|
||||
# fleetd #492 — the one-daemon check. Two live pids is the exact symptom a racing supervisor
|
||||
# produces, and none of the other post-restart checks (healthz, jar id, the fresh log line) can see
|
||||
# it because either daemon alone satisfies them.
|
||||
test_count_daemon_pids() {
|
||||
assert_equals 0 "$(count_daemon_pids "")" "count of an empty pid list"
|
||||
assert_equals 1 "$(count_daemon_pids "4242")" "count of a single pid"
|
||||
assert_equals 2 "$(count_daemon_pids "$(printf '4242\n4343\n')")" "count of two pids"
|
||||
}
|
||||
|
||||
test_assert_single_daemon_accepts_one_pid() {
|
||||
assert_single_daemon "4242" || fail "assert_single_daemon rejected a single running pid"
|
||||
}
|
||||
|
||||
test_assert_single_daemon_rejects_two_pids() {
|
||||
local output rc=0
|
||||
output="$(assert_single_daemon "$(printf '4242\n4343\n')" 2>&1)" || rc=$?
|
||||
[ "$rc" -ne 0 ] || fail "assert_single_daemon accepted two simultaneously running pids"
|
||||
printf '%s' "$output" | grep -qF '4242' || fail "refusal message does not list the pids it found"
|
||||
printf '%s' "$output" | grep -qF '4343' || fail "refusal message does not list the pids it found"
|
||||
}
|
||||
|
||||
test_no_errors() {
|
||||
cat > "$TMP/no-errors.log" <<'LOG'
|
||||
2026-09-05 12:00:00 INFO fleetd listening
|
||||
@@ -224,6 +296,14 @@ test_unattributable_quiet_mutation_is_caught() {
|
||||
printf 'Unattributable mutation: FAIL: cross-unattributable recovered: expected 0, got 2\n'
|
||||
}
|
||||
|
||||
test_detect_supervisor_launchd_only
|
||||
test_detect_supervisor_systemd_only
|
||||
test_detect_supervisor_none
|
||||
test_require_drivable_supervisor_refuses_ambiguous
|
||||
test_require_drivable_supervisor_accepts_known_kinds
|
||||
test_count_daemon_pids
|
||||
test_assert_single_daemon_accepts_one_pid
|
||||
test_assert_single_daemon_rejects_two_pids
|
||||
test_no_errors
|
||||
test_recovery_patterns_match_source
|
||||
test_attributed_recovered_connection_error
|
||||
|
||||
Reference in New Issue
Block a user