Compare commits
15 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 325d0771a4 | |||
| 7b97aae85b | |||
| 49a404ddf3 | |||
| 72d6a6878b | |||
| 4466ee0ef2 | |||
| 25ba7f16bb | |||
| 5ba69c9cf2 | |||
| 17052bb515 | |||
| e95ed99bf7 | |||
| 1477e4358a | |||
| 6c2d6e93cb | |||
| 789b6a8716 | |||
| 01462c9695 | |||
| 8b4ff78546 | |||
| 5a467e1f8b |
@@ -0,0 +1,163 @@
|
||||
---
|
||||
name: handover
|
||||
description: Procedure for an outgoing lead to write the handover file before a fresh lead session replaces it (fleetd #480). Load this when your context is full and fleetd is about to clear your pane. The file is the new lead's only inheritance — follow it exactly.
|
||||
---
|
||||
|
||||
# Handover — write the file the next lead depends on
|
||||
|
||||
fleetd ticket #480 lets a lead session hand off to a fresh one. The outgoing lead writes a
|
||||
handover file, fleetd checks it, clears the pane, and tells the new session to read that file
|
||||
and carry on.
|
||||
|
||||
**The new lead's only inheritance is that file.** It does not see your conversation, your plan,
|
||||
or your screen. If the file is thin or wrong, the new lead re-derives what you already knew, and
|
||||
that wastes hours. Writing a good handover file is real work. It is not paperwork you rush
|
||||
through at the end of a session.
|
||||
|
||||
This skill is the procedure for writing it. Every rule below earned its place because a past
|
||||
handover got it wrong.
|
||||
|
||||
## 1. Confirm you are the right session to write this
|
||||
|
||||
Run `fleet_whoami` first. It must answer `primary`. Only a primary (lead) session writes a
|
||||
handover file. A worker's job ends with its own pull request, not a fleet-wide handoff.
|
||||
|
||||
## 2. Every number needs a command, run in this turn
|
||||
|
||||
A number is a claim: a count, a commit hash, a process id, a percentage, a queue depth. Before
|
||||
you write one, run the command that produces it — now, in this turn, against the live state.
|
||||
|
||||
Never take a number from:
|
||||
|
||||
- earlier in your own conversation — the state has moved since then,
|
||||
- a peer lead's report — that is their measurement, not yours,
|
||||
- your own memory of an earlier session.
|
||||
|
||||
Put the command, or its real output, next to the number. That lets the next lead re-run it and
|
||||
check it still matches. If you cannot measure something yourself, say so instead of guessing:
|
||||
"the fleet01 lead reports 91 commits behind; I have not checked this myself."
|
||||
|
||||
## 3. Say what you measured and what you did not
|
||||
|
||||
Mark every claim as one of two things:
|
||||
|
||||
- **"I checked this myself, in the code or on this host, at `<time>`."**
|
||||
- **"I did not check this myself; `<who>` reported it."**
|
||||
|
||||
Never present someone else's measurement as your own. This matters most for cross-host claims —
|
||||
a peer lead's daemon, a worker's report, or something the operator said earlier that you cannot
|
||||
re-verify from here.
|
||||
|
||||
## 4. Record open decisions, and who owns them
|
||||
|
||||
List three things:
|
||||
|
||||
- what the operator actually asked for, in their own words where you have them,
|
||||
- what is still unanswered,
|
||||
- any question you decided yourself instead of asking, with your reason.
|
||||
|
||||
Write the decision so it cannot be mistaken for the operator's instruction. Say plainly: "the
|
||||
operator never answered X; I decided Y, because Z." Without this, the next lead either silently
|
||||
reopens a closed question or assumes the operator chose something they never did.
|
||||
|
||||
## 5. Record live hazards
|
||||
|
||||
List anything that will break if the next lead does the obvious thing next. This includes:
|
||||
|
||||
- unpushed commits or unmerged branches,
|
||||
- a build, a spawn, or a redeploy still running,
|
||||
- code merged to `main` but not yet redeployed to the live daemon,
|
||||
- any trap that looks safe and is not — say what goes wrong and why, not only that something is
|
||||
"tricky."
|
||||
|
||||
## 6. Record what is explicitly not owed
|
||||
|
||||
List work that is finished, and work that another party has said they do not want touched. Name
|
||||
who said so and when. Without this line, the next lead re-does closed work or reopens a question
|
||||
a peer already declined to revisit.
|
||||
|
||||
## 7. Open the file with three re-measurement commands
|
||||
|
||||
The file's own first section must give the next lead three concrete commands to run before
|
||||
acting on anything else in the file:
|
||||
|
||||
1. confirm role — for example `fleet_whoami`,
|
||||
2. confirm the state of the working tree — for example `git status` and
|
||||
`git rev-list origin/main..HEAD`,
|
||||
3. read the live fleet — for example `fleet_list`.
|
||||
|
||||
Record what each command answered when you wrote the file, and tell the reader to run it again
|
||||
rather than trust your answer. The point of this section is that the reader checks live state
|
||||
before acting on any claim in the rest of the file, including yours.
|
||||
|
||||
## 8. Stamp the file with time and commit
|
||||
|
||||
Near the top of the file, write:
|
||||
|
||||
- the date and time you wrote it,
|
||||
- the commit the tree was on (`git rev-parse HEAD`),
|
||||
- whether the tree was clean (`git status`).
|
||||
|
||||
Without this, nobody can tell how old the file is, or which code it describes.
|
||||
|
||||
## 9. State plainly that the file goes stale fast
|
||||
|
||||
Say near the top: **re-measure anything you act on.** The file goes stale the moment anyone
|
||||
merges a branch, spawns a member, or restarts the daemon. Everything in the file is a snapshot
|
||||
of one moment, not a live fact.
|
||||
|
||||
## 10. What to leave out
|
||||
|
||||
Do not include:
|
||||
|
||||
- narration of how the session felt, or how hard something was,
|
||||
- anything the repo already records — code structure, git history, or a rule already written in
|
||||
`CLAUDE.md`. Point at it instead of repeating it,
|
||||
- advice that is only true for the session that is ending — a half-open terminal, a local
|
||||
variable, a train of thought with no state behind it.
|
||||
|
||||
A handover file is a record of state and decisions. It is not a diary.
|
||||
|
||||
## Writing style
|
||||
|
||||
Write in plain English. Use everyday words, one idea per sentence, and active voice. Keep every
|
||||
class, method, file, flag, and config key exactly as it appears in the code — replacing a
|
||||
precise term with a vague one makes the sentence wrong, not simpler. Explain an abbreviation the
|
||||
first time you use it.
|
||||
|
||||
If a diagram genuinely helps, put it in the `.md` file as a fenced ` ```mermaid ` block with no
|
||||
hardcoded colors, so it stays readable on light and dark backgrounds. Quote any label that has
|
||||
brackets, colons, or slashes.
|
||||
|
||||
## Template
|
||||
|
||||
```markdown
|
||||
# Handover — <fleet name> lead session, <date and time>
|
||||
|
||||
Written at commit `<output of git rev-parse HEAD>`. Tree was <clean, or dirty: `<git status
|
||||
summary>`>. Re-measure anything you act on — this file goes stale the moment anyone merges,
|
||||
spawns, or restarts.
|
||||
|
||||
## 0. Do these three things first
|
||||
|
||||
1. Confirm your role: `fleet_whoami` — must answer `primary`. (Answered `<result>` at `<time>`.)
|
||||
2. Confirm tree state: `git status`, `git rev-list origin/main..HEAD`. (`<result>` at `<time>`.)
|
||||
3. Read the live fleet: `fleet_list`. (`<result>` at `<time>`.)
|
||||
|
||||
## 1. What the operator asked for
|
||||
|
||||
<the live instructions, in their words where you have them; what is still open; any decision
|
||||
you made yourself, and why>
|
||||
|
||||
## 2. Open decisions, and who owns them
|
||||
|
||||
<one line per decision: who owns it, what is unanswered>
|
||||
|
||||
## 3. Live hazards
|
||||
|
||||
<one entry per hazard: what breaks, and why, if the next lead does the obvious thing>
|
||||
|
||||
## 4. What is not owed
|
||||
|
||||
<finished work, and work another party has declined; name who said so and when>
|
||||
```
|
||||
@@ -0,0 +1,72 @@
|
||||
---
|
||||
name: redeploy-fleetd
|
||||
description: Rebuild and restart the live fleetd daemon after a merge (lead / primary only). Load this before redeploying — it holds the script, the drain step, the permission grant, and the five checks that have each gone wrong here before. Workers must never do this.
|
||||
---
|
||||
|
||||
### Redeploying the daemon — the lead may do this (primary only)
|
||||
|
||||
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
|
||||
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
|
||||
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
|
||||
rather than hand the job back to the operator.
|
||||
|
||||
Workers must never do this. A worker has no business restarting the daemon it is talking through,
|
||||
and stopping it kills the worker's own channel mid-turn.
|
||||
|
||||
**Use the script — do not hand-roll the steps.**
|
||||
|
||||
```bash
|
||||
scripts/redeploy-fleetd.sh --check # report state, change nothing
|
||||
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
|
||||
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
|
||||
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
|
||||
```
|
||||
|
||||
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
|
||||
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
|
||||
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
|
||||
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
|
||||
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
|
||||
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
|
||||
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
|
||||
|
||||
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
|
||||
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
|
||||
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
|
||||
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
|
||||
|
||||
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
|
||||
if the script is unavailable or a step fails, this is what it was protecting you from.
|
||||
|
||||
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
|
||||
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
|
||||
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
|
||||
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
|
||||
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
|
||||
whether the name resolves without ever printing the value.
|
||||
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
|
||||
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
|
||||
rendezvous, and a member's report is not recoverable once its ticket is gone.
|
||||
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
|
||||
it. The startup log names which keys it accepted and which it deferred — read those lines rather
|
||||
than assuming.
|
||||
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
|
||||
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
|
||||
demoted to worker, which refuses every orchestration call.
|
||||
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
|
||||
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
|
||||
the outside.
|
||||
|
||||
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
|
||||
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
|
||||
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
|
||||
start every time. The rule lives in the operator's Claude Code settings:
|
||||
|
||||
```json
|
||||
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
|
||||
```
|
||||
|
||||
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
|
||||
running the stop and start as separate commands — that is exactly the approval the script replaced.
|
||||
Say what you were going to run and why, and let the operator decide.
|
||||
|
||||
@@ -257,8 +257,11 @@ must obey belongs in the charter, not here.
|
||||
already cost three workers' turns: each wrote a good report to its terminal and ended the turn
|
||||
with no `fleet_reply`, and the scrape returned the tail of the brief instead.
|
||||
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
|
||||
`port-to-opencode` (make an OpenCode session a participant in this workspace) and
|
||||
`fleets-status` (report every fleet that shares one LavinMQ instance).
|
||||
`port-to-opencode` (make an OpenCode session a participant in this workspace),
|
||||
`fleets-status` (report every fleet that shares one LavinMQ instance),
|
||||
`redeploy-fleetd` (rebuild and restart the live daemon after a merge) and
|
||||
`handover` (write the file a fresh lead session inherits when the outgoing one hands off,
|
||||
fleetd #480).
|
||||
- **This repo is also a Claude Code marketplace, and ships a plugin.** `.claude-plugin/marketplace.json`
|
||||
points at `plugin/`, which carries the MCP mount and the `setup` skill
|
||||
(`/claude-bridge:setup` — make any project bridge-ready). It was added in CB-527 and then went
|
||||
@@ -292,60 +295,14 @@ must obey belongs in the charter, not here.
|
||||
### Redeploying the daemon — the lead may do this (primary only)
|
||||
|
||||
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
|
||||
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
|
||||
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
|
||||
rather than hand the job back to the operator.
|
||||
feature merged to `main` does nothing until the daemon is rebuilt and restarted. The lead **may and
|
||||
should** redeploy rather than hand the job back to the operator. **Workers must never do this** — a
|
||||
worker has no business restarting the daemon it is talking through, and stopping it kills the
|
||||
worker's own channel mid-turn.
|
||||
|
||||
Workers must never do this. A worker has no business restarting the daemon it is talking through,
|
||||
and stopping it kills the worker's own channel mid-turn.
|
||||
|
||||
**Use the script — do not hand-roll the steps.**
|
||||
|
||||
```bash
|
||||
scripts/redeploy-fleetd.sh --check # report state, change nothing
|
||||
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
|
||||
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
|
||||
```
|
||||
|
||||
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
|
||||
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
|
||||
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
|
||||
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
|
||||
|
||||
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
|
||||
if the script is unavailable or a step fails, this is what it was protecting you from.
|
||||
|
||||
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
|
||||
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
|
||||
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
|
||||
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
|
||||
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
|
||||
whether the name resolves without ever printing the value.
|
||||
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
|
||||
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
|
||||
rendezvous, and a member's report is not recoverable once its ticket is gone.
|
||||
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
|
||||
it. The startup log names which keys it accepted and which it deferred — read those lines rather
|
||||
than assuming.
|
||||
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
|
||||
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
|
||||
demoted to worker, which refuses every orchestration call.
|
||||
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
|
||||
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
|
||||
the outside.
|
||||
|
||||
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
|
||||
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
|
||||
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
|
||||
start every time. The rule lives in the operator's Claude Code settings:
|
||||
|
||||
```json
|
||||
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
|
||||
```
|
||||
|
||||
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
|
||||
running the stop and start as separate commands — that is exactly the approval the script replaced.
|
||||
Say what you were going to run and why, and let the operator decide.
|
||||
**Load the `redeploy-fleetd` skill before you redeploy.** It holds `scripts/redeploy-fleetd.sh`
|
||||
and its flags, the drain step, the operator's permission grant, and the five checks that have each
|
||||
gone wrong here before. Do not hand-roll the steps from memory.
|
||||
|
||||
### The prompt is part of the product — update it with the code (mandatory)
|
||||
|
||||
|
||||
@@ -452,6 +452,15 @@ placement: weighted
|
||||
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
|
||||
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
|
||||
# 1800 (30 minutes) when omitted or non-positive.
|
||||
#
|
||||
# fleetd #466: this is now only the BASE of an escalating backoff, not a flat retry rate. A
|
||||
# credential quarantined again within one base cooldown of the previous quarantine ending (still
|
||||
# reporting exhausted — e.g. a weekly subscription limit that hasn't reset) backs off further:
|
||||
# cooldown doubles each such time, capped at 12x this value (~6 hours at the 1800s default). A
|
||||
# quarantine that starts after a base-cooldown's worth of quiet resets back to this value. Not
|
||||
# configurable per se — the multiplier and ceiling are constants in BackendQuarantine, not new
|
||||
# YAML keys; see its class doc for the exact formula and why there is no automatic probe to clear
|
||||
# it early (the operator's own design constraint — a probe spends the quota it's measuring).
|
||||
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
|
||||
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
|
||||
# restart. Editing this needs a daemon restart to take effect.
|
||||
|
||||
@@ -26,6 +26,7 @@ import dev.ltms.fleet.inject.MemberPresence;
|
||||
import dev.ltms.fleet.auth.MemberRegistry;
|
||||
import dev.ltms.fleet.auth.CallerResolver;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import dev.ltms.fleet.mcp.CharterToolSurface;
|
||||
import dev.ltms.fleet.mcp.ConnectionIdentity;
|
||||
import dev.ltms.fleet.metrics.FleetMetrics;
|
||||
import dev.ltms.fleet.metrics.Metrics;
|
||||
@@ -147,7 +148,10 @@ public final class Fleetd {
|
||||
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
|
||||
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
|
||||
// contract; adding a reader here does not make a key reloadable by itself.
|
||||
ConfigRef config = new ConfigRef(configPath, cfg);
|
||||
// fleetd #474: pass the charter/tool-surface check in as ConfigRef's extraValidation, so
|
||||
// ConfigRef#reload() runs the same gate main() runs below, without dev.ltms.fleet.config
|
||||
// gaining a dependency on dev.ltms.fleet.mcp — Fleetd is the seam that already holds both.
|
||||
ConfigRef config = new ConfigRef(configPath, cfg, Fleetd::assertChartersNameOnlyRegisteredTools);
|
||||
|
||||
// The primary/host env that launched fleetd must not be tainted.
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
@@ -163,6 +167,17 @@ public final class Fleetd {
|
||||
// and the Fleetd-startup tests actually pin — see FleetConfig#validateAll's javadoc for
|
||||
// why a name-by-name list here would have the same defect it replaces.
|
||||
cfg.validateAll();
|
||||
// fleetd #469, follow-up to #464: validateAll() (and validateCharters() inside it) only
|
||||
// checks that a charter's KEY is a role wire name and its text is non-blank — it never
|
||||
// looks at what the text actually names. This is the separate check that does: it asks
|
||||
// dev.ltms.fleet.mcp.FleetTool (the canonical registered-tool set) whether every fleet_*/
|
||||
// bridge_* token a charter names is a tool this server actually registers. It cannot live
|
||||
// inside FleetConfig#validateCharters() — config loads before the MCP server exists, and
|
||||
// must not gain a dependency on the mcp package — so it runs here instead, at the one seam
|
||||
// that already holds both a loaded FleetConfig and the mcp package, before anything below
|
||||
// opens a socket or spawns a member. fleetd #474: the same check is also wired into `config`
|
||||
// above as ConfigRef's extraValidation, so a reload refuses what this line refuses at startup.
|
||||
assertChartersNameOnlyRegisteredTools(cfg);
|
||||
|
||||
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
|
||||
? Path.of(cfg.herdrSocket())
|
||||
@@ -222,7 +237,13 @@ public final class Fleetd {
|
||||
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
|
||||
// The cooldown is deferred (see FleetConfig#quarantineCooldownSeconds): it is read once
|
||||
// here, at startup, and a config reload only changes it for a daemon restart.
|
||||
BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,
|
||||
// fleetd #466: escalating, not flat — a credential that keeps reporting exhaustion (e.g. a
|
||||
// weekly subscription limit, which would otherwise be retried on every ~30-minute cooldown,
|
||||
// about 336 times across the week) backs off further each consecutive time, capped at
|
||||
// BackendQuarantine.DEFAULT_MAX_COOLDOWN_MULTIPLE x the base cooldown. See BackendQuarantine's
|
||||
// class doc for the mechanism, why this never fires on cooling-off (a separate, unescalated
|
||||
// mechanism — BackendOutagePolicy below), and the reset.
|
||||
BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,
|
||||
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
|
||||
// fleetd #201 Unit 5: one outage-cool-off tracker for the whole daemon, shared between the
|
||||
// launcher (checked at spawn, like `quarantine` above) and the backend-error sink wired in
|
||||
@@ -1576,6 +1597,28 @@ public final class Fleetd {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #474: the one place both the startup call (right after {@code cfg.validateAll()} in
|
||||
* {@link #main}) and the reload call (wired into {@code config}'s {@code extraValidation} above,
|
||||
* via a method reference to this method) go through, so the two can never drift into checking
|
||||
* different things. Extracted only to give {@link ConfigRef}'s {@code Consumer<FleetConfig>}
|
||||
* hook a {@code FleetConfig -> void} shape to bind to — {@link CharterToolSurface} itself still
|
||||
* takes the raw charter map and knows nothing about {@code ConfigRef} or {@code Fleetd}.
|
||||
*
|
||||
* <p>Package-private so a test can call it directly the same way the other startup-report
|
||||
* helpers above are tested, without needing to drive {@link #main} for a unit-level check;
|
||||
* {@code FleetdStartupValidationTest} proves the startup call site, and {@code
|
||||
* FleetdConfigRefCharterToolSurfaceWiringTest} — by constructing {@code ConfigRef} with this
|
||||
* exact method reference, the same way {@code main} does above — proves the reload call site.
|
||||
* {@code dev.ltms.fleet.config.ConfigRefTest} pins the same reload behaviour too, through an
|
||||
* equivalent {@code Consumer<FleetConfig>} it builds locally (it cannot see this package-private
|
||||
* method from {@code dev.ltms.fleet.config}).
|
||||
*/
|
||||
static void assertChartersNameOnlyRegisteredTools(FleetConfig cfg) {
|
||||
CharterToolSurface.assertChartersNameOnlyRegisteredTools(
|
||||
cfg.fleet() == null ? Map.of() : cfg.fleet().charters());
|
||||
}
|
||||
|
||||
/**
|
||||
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
|
||||
*
|
||||
|
||||
@@ -11,6 +11,7 @@ import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
import java.util.function.Consumer;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
@@ -211,6 +212,24 @@ import java.util.function.Supplier;
|
||||
* <p>A reload that fails to parse or fails validation is also refused, and the previous config keeps
|
||||
* running. A config file being edited is normally read once mid-save; degrading a working daemon
|
||||
* because it caught a half-written file would be a bad trade.
|
||||
*
|
||||
* <p><strong>fleetd #474</strong> — {@link FleetConfig#validateAll()} is not the only gate startup
|
||||
* runs before a config takes effect: {@code Fleetd.main} also calls {@code
|
||||
* dev.ltms.fleet.mcp.CharterToolSurface#assertChartersNameOnlyRegisteredTools}, right after {@code
|
||||
* cfg.validateAll()}, to refuse a charter that names an MCP tool the server does not register. That
|
||||
* check cannot live inside {@link FleetConfig} — {@code CharterToolSurface} lives in the {@code mcp}
|
||||
* package because the canonical tool set ({@code FleetTool}) does, and config is loaded before the
|
||||
* MCP server exists, so {@code FleetConfig} must not gain a dependency on {@code mcp}. {@link
|
||||
* #reload} cannot import {@code mcp} either, for the same reason applied one layer up: {@code
|
||||
* dev.ltms.fleet.config} is loaded before {@code dev.ltms.fleet.mcp} exists, same as {@code
|
||||
* FleetConfig}. So this class accepts the check as a {@code Consumer<FleetConfig>} —
|
||||
* {@link #extraValidation} — supplied by whichever caller already sits at the seam that holds both
|
||||
* a loaded {@code FleetConfig} and the {@code mcp} package: {@code Fleetd.main}. It is invoked
|
||||
* inside the same try/catch as {@code fresh.validateAll()}, so a charter that would have refused to
|
||||
* boot refuses a reload too, and keeps the running config exactly like any other {@code
|
||||
* validateAll()} failure. A ref built through the two-argument constructor (every test fixture that
|
||||
* does not care about this check, and {@link #fixed}) gets a no-op consumer, so nothing outside
|
||||
* {@code Fleetd.main} needs to know this hook exists.
|
||||
*/
|
||||
public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
|
||||
@@ -255,10 +274,27 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
|
||||
private final Path path;
|
||||
private final AtomicReference<FleetConfig> current;
|
||||
private final Consumer<FleetConfig> extraValidation;
|
||||
|
||||
/** Equivalent to the three-argument constructor with a no-op {@code extraValidation}. */
|
||||
public ConfigRef(Path path, FleetConfig initial) {
|
||||
this(path, initial, cfg -> { });
|
||||
}
|
||||
|
||||
/**
|
||||
* @param extraValidation run on every {@link #reload} candidate, inside the same try/catch as
|
||||
* {@code fresh.validateAll()} — see the class doc's fleetd #474 note.
|
||||
* {@code Fleetd.main} passes {@code
|
||||
* Fleetd::assertChartersNameOnlyRegisteredTools} (a package-private
|
||||
* {@code FleetConfig -> void} adapter over {@code
|
||||
* CharterToolSurface#assertChartersNameOnlyRegisteredTools}), so a reload
|
||||
* runs the same gate startup does without this class depending on the
|
||||
* {@code mcp} package.
|
||||
*/
|
||||
public ConfigRef(Path path, FleetConfig initial, Consumer<FleetConfig> extraValidation) {
|
||||
this.path = path;
|
||||
this.current = new AtomicReference<>(Objects.requireNonNull(initial, "initial config"));
|
||||
this.extraValidation = Objects.requireNonNull(extraValidation, "extraValidation");
|
||||
}
|
||||
|
||||
/** A fixed reference that never reloads — for tests and for wiring built from a config in code. */
|
||||
@@ -360,6 +396,12 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
// at all — see FleetConfig#validateAll's javadoc for why the fix is one reflective call,
|
||||
// not a longer hand-maintained list.
|
||||
fresh.validateAll();
|
||||
// fleetd #474: validateAll() does not cover everything startup refuses on — the charter
|
||||
// tool-surface check (Fleetd.main, right after cfg.validateAll()) lives outside
|
||||
// FleetConfig on purpose (see this class's doc) and is supplied here as extraValidation.
|
||||
// Same try/catch as validateAll() above, on purpose: either failure must refuse the whole
|
||||
// reload and keep the running config the same way.
|
||||
extraValidation.accept(fresh);
|
||||
} catch (RuntimeException e) {
|
||||
String msg = e.getMessage() == null ? e.toString() : e.getMessage();
|
||||
log.warn("config reload from {} refused, keeping the running config: {}", path, msg);
|
||||
|
||||
@@ -70,12 +70,14 @@ import java.util.regex.PatternSyntaxException;
|
||||
* {@code fixed} (default), {@code round-robin}, or {@code weighted}
|
||||
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
|
||||
* historical behaviour), CB-501
|
||||
* @param quarantineCooldownSeconds how long a credential stays quarantined after a
|
||||
* @param quarantineCooldownSeconds the BASE cooldown a credential is quarantined for after a
|
||||
* {@code BACKEND_EXHAUSTED} classification (CB-578 stage B); {@code null}/{@code
|
||||
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Baked once into the
|
||||
* {@code BackendQuarantine} built at startup, so it is DEFERRED: changing it
|
||||
* needs a restart, and a quarantine already running keeps whatever cooldown was
|
||||
* live when it started.
|
||||
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Since fleetd #466 this is
|
||||
* only the first occurrence's length — a credential quarantined again shortly
|
||||
* after this cooldown ends backs off further, up to a ceiling; see {@code
|
||||
* BackendQuarantine}'s class doc. Baked once into the {@code BackendQuarantine}
|
||||
* built at startup, so it is DEFERRED: changing it needs a restart, and a
|
||||
* quarantine already running keeps whatever cooldown was live when it started.
|
||||
* @param memberCredentials deny-by-default policy (CB-596) for which of the operator's own host
|
||||
* credentials a spawned member's pane inherits. {@code null} (the block
|
||||
* omitted) blocks nothing — see {@link MemberCredentials}.
|
||||
|
||||
@@ -0,0 +1,81 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
/**
|
||||
* fleetd #469: a launch charter that names an MCP tool the server does not register must stop the
|
||||
* daemon at startup, not wait for a member to discover the gap by calling something that is not
|
||||
* there.
|
||||
*
|
||||
* <p>Deliberately its own class outside {@code dev.ltms.fleet.config}, not a case in {@link
|
||||
* dev.ltms.fleet.config.FleetConfig#validateCharters()}. The canonical tool surface ({@link
|
||||
* FleetTool}) lives in the {@code mcp} package; config is loaded before the MCP server exists and
|
||||
* must not gain a dependency on it. So this check belongs at the seam that already holds both a
|
||||
* loaded {@code FleetConfig} and the {@code mcp} package: {@code Fleetd.main}, called right after
|
||||
* {@code cfg.validateAll()} and before anything opens a socket or spawns a member.
|
||||
*
|
||||
* <p>{@code #464}'s {@code CharterToolSurfaceTest} proved the same comparison against a charter
|
||||
* fixture it wrote itself into a {@code @TempDir}, which meant nothing anyone wrote into the live
|
||||
* {@code fleetd.yaml} could ever fail it. This class is what a real charter is actually checked
|
||||
* against at boot; {@code FleetdStartupValidationTest} exercises it through {@code Fleetd.main}
|
||||
* itself, the same way it proves every other {@code validateXxx()} still runs there.
|
||||
*
|
||||
* <p><strong>fleetd #474</strong> — startup was not the only door: {@code
|
||||
* dev.ltms.fleet.config.ConfigRef#reload()} used to run {@code FleetConfig#validateAll()} alone,
|
||||
* which does not look at what a charter's text names, so a charter naming an unregistered tool
|
||||
* that could not have booted the daemon could still be installed into a running one through a
|
||||
* reload. This class still knows nothing about {@code ConfigRef} — {@code Fleetd.main} wires a
|
||||
* small {@code FleetConfig -> void} adapter over {@link #assertChartersNameOnlyRegisteredTools}
|
||||
* ({@code Fleetd::assertChartersNameOnlyRegisteredTools}) into {@code ConfigRef}'s constructor as
|
||||
* its {@code Consumer<FleetConfig>} {@code extraValidation}, run inside {@code reload()}'s same
|
||||
* try/catch as {@code validateAll()}, so both call sites — {@code Fleetd.main} at startup and
|
||||
* {@code ConfigRef#reload()} afterwards — go through this one method and can never check different
|
||||
* things. {@code dev.ltms.fleet.config.ConfigRefTest} and {@code
|
||||
* FleetdConfigRefCharterToolSurfaceWiringTest} are what prove the reload call site, the same way
|
||||
* {@code FleetdStartupValidationTest} proves the startup one.
|
||||
*/
|
||||
public final class CharterToolSurface {
|
||||
|
||||
/** A {@code fleet_…} (current) or {@code bridge_…} (pre-CB-634) tool-shaped token in prose. */
|
||||
private static final Pattern TOOL_REFERENCE = Pattern.compile("(fleet_[a-z_]+|bridge_[a-z_]+)");
|
||||
|
||||
private CharterToolSurface() {
|
||||
}
|
||||
|
||||
/**
|
||||
* @param charters the configured {@code fleet.charters:} map (role wire name → charter text);
|
||||
* {@code null} or empty is a no-op, same as an absent {@code fleet:} block
|
||||
* @throws IllegalStateException naming the charter key and every tool it names that {@link
|
||||
* FleetTool} does not list, when any charter does so
|
||||
*/
|
||||
public static void assertChartersNameOnlyRegisteredTools(Map<String, String> charters) {
|
||||
if (charters == null || charters.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
Set<String> registered = FleetTool.wireNames();
|
||||
List<String> bad = new ArrayList<>();
|
||||
charters.forEach((key, text) -> {
|
||||
if (text == null) {
|
||||
return;
|
||||
}
|
||||
Set<String> named = new LinkedHashSet<>();
|
||||
Matcher m = TOOL_REFERENCE.matcher(text);
|
||||
while (m.find()) {
|
||||
named.add(m.group(1));
|
||||
}
|
||||
named.stream()
|
||||
.filter(t -> !registered.contains(t))
|
||||
.forEach(unknown -> bad.add("fleet.charters." + key + " names '" + unknown
|
||||
+ "', which the server does not register (registered: " + registered + ")."));
|
||||
});
|
||||
if (!bad.isEmpty()) {
|
||||
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -490,6 +490,21 @@ public final class FleetMcp {
|
||||
McpSchema.Tool fleetProfiles = profilesTool();
|
||||
McpSchema.Tool fleetWhoami = whoamiTool();
|
||||
|
||||
// fleetd #469: the tool schemas above are already named from FleetTool.wireName(), but
|
||||
// this is the check that a schema was not accidentally dropped, duplicated, or added
|
||||
// under a name FleetTool does not list. It runs once, at construction (startup), rather
|
||||
// than being left to CharterToolSurface or a test to discover later — a canonical entry
|
||||
// this server never registers, or a registration with no canonical entry backing it, is a
|
||||
// startup failure, not a silent gap.
|
||||
Set<String> registeredToolNames = Set.of(fleetSend.name(), fleetReply.name(), fleetAsk.name(),
|
||||
fleetStatus.name(), fleetPoll.name(), fleetAck.name(), fleetSpawn.name(),
|
||||
fleetList.name(), fleetStop.name(), fleetProfiles.name(), fleetWhoami.name());
|
||||
if (!registeredToolNames.equals(FleetTool.wireNames())) {
|
||||
throw new IllegalStateException("fleetd #469: registered MCP tools " + registeredToolNames
|
||||
+ " do not match the canonical tool set " + FleetTool.wireNames()
|
||||
+ " -- FleetTool is the single source of truth for what this server registers");
|
||||
}
|
||||
|
||||
this.server = McpServer.sync(transport)
|
||||
.serverInfo("fleet", "0.1.0")
|
||||
.capabilities(McpSchema.ServerCapabilities.builder().tools(true).build())
|
||||
@@ -897,20 +912,34 @@ public final class FleetMcp {
|
||||
|
||||
/**
|
||||
* The action a registered tool handler actually hands to the authorization gate.
|
||||
* Keeping this choice beside the registered-tool inventory makes a new tool fail the coverage
|
||||
* test until its action is pinned.
|
||||
*
|
||||
* <p>{@code toolName} is a raw string off the wire (an MCP call names its tool by string, and a
|
||||
* malformed or stale client can send anything), so resolving it against {@link FleetTool} first
|
||||
* — and throwing on a miss — is still a run-time check by necessity. What moved to compile time
|
||||
* is the second step: {@link #authzAction(FleetTool, Map)} switches on the resolved {@link
|
||||
* FleetTool} itself with no {@code default}, so a new {@link FleetTool} constant with no pinned
|
||||
* action fails {@code mvn compile}, not just {@code FleetMcpAuthzTest} at run time.
|
||||
*/
|
||||
static Authz.Action toolAction(String toolName, Map<String, Object> arguments) {
|
||||
return switch (toolName) {
|
||||
case "fleet_send" -> Authz.Action.SEND;
|
||||
case "fleet_reply" -> Authz.Action.REPLY;
|
||||
case "fleet_ask" -> Authz.Action.ASK;
|
||||
case "fleet_status", "fleet_list", "fleet_profiles", "fleet_whoami" -> Authz.Action.READ;
|
||||
case "fleet_poll" -> pollAction(str(arguments, "target"), str(arguments, "coordId"));
|
||||
case "fleet_ack" -> Authz.Action.DRAIN;
|
||||
case "fleet_spawn" -> Authz.Action.SPAWN;
|
||||
case "fleet_stop" -> Authz.Action.STOP;
|
||||
default -> throw new IllegalArgumentException("unregistered tool: " + toolName);
|
||||
FleetTool tool = FleetTool.byWireName(toolName)
|
||||
.orElseThrow(() -> new IllegalArgumentException("unregistered tool: " + toolName));
|
||||
return authzAction(tool, arguments);
|
||||
}
|
||||
|
||||
/**
|
||||
* Exhaustive over {@link FleetTool} on purpose — no {@code default}. Adding a tool to {@link
|
||||
* FleetTool} without adding its case here is a compile error (fleetd #469).
|
||||
*/
|
||||
private static Authz.Action authzAction(FleetTool tool, Map<String, Object> arguments) {
|
||||
return switch (tool) {
|
||||
case SEND -> Authz.Action.SEND;
|
||||
case REPLY -> Authz.Action.REPLY;
|
||||
case ASK -> Authz.Action.ASK;
|
||||
case STATUS, LIST, PROFILES, WHOAMI -> Authz.Action.READ;
|
||||
case POLL -> pollAction(str(arguments, "target"), str(arguments, "coordId"));
|
||||
case ACK -> Authz.Action.DRAIN;
|
||||
case SPAWN -> Authz.Action.SPAWN;
|
||||
case STOP -> Authz.Action.STOP;
|
||||
};
|
||||
}
|
||||
|
||||
@@ -1269,6 +1298,13 @@ public final class FleetMcp {
|
||||
* (the backend text that triggered the most recent quarantine of that credential, omitted when
|
||||
* none is known) — so a lead can see WHICH model to turn off and WHY, without reading the
|
||||
* daemon log.
|
||||
*
|
||||
* <p>fleetd #466 scope item 2: each {@code quarantined} row also names {@code
|
||||
* quarantineAttempt} — 1 for a first-time exhaustion, 2 for the second in a row, and so on — so
|
||||
* an operator sees "this is the 5th time" instead of inferring it from {@code
|
||||
* quarantinedForSeconds} alone. Read off {@link BackendQuarantine#status(String)}, the same
|
||||
* one-call accessor {@code quarantinedForSeconds} itself comes from here (see its doc) — never a
|
||||
* separately derived count.
|
||||
*/
|
||||
public static Map<String, Object> profilesView(PeerLauncher workers, QuarantineSource quarantine, OutageSource outage) {
|
||||
Map<String, Object> result = new LinkedHashMap<>();
|
||||
@@ -1281,10 +1317,14 @@ public final class FleetMcp {
|
||||
exhaustionDetectionArmed.put(profile, quarantine.exhaustedPatternArmed().apply(profile));
|
||||
String credentialId = quarantine.credentialIdFor().apply(profile);
|
||||
if (credentialId != null) {
|
||||
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
|
||||
// fleetd #466 scope item 2: quarantinedForSeconds and quarantineAttempt come off the
|
||||
// ONE BackendQuarantine#status(credentialId) call — never a second, independent read
|
||||
// for the attempt count — so the two can never disagree about which streak this is.
|
||||
quarantine.quarantine().status(credentialId).ifPresent(status -> {
|
||||
Map<String, Object> row = new LinkedHashMap<>();
|
||||
row.put("credentialId", credentialId);
|
||||
row.put("quarantinedForSeconds", remaining);
|
||||
row.put("quarantinedForSeconds", status.remainingSeconds());
|
||||
row.put("quarantineAttempt", status.repeatCount());
|
||||
String model = quarantine.modelFor().apply(profile);
|
||||
if (model != null && !model.isBlank()) {
|
||||
row.put("model", model);
|
||||
@@ -1724,10 +1764,13 @@ public final class FleetMcp {
|
||||
}
|
||||
String credentialId = quarantine.credentialIdFor().apply(profile);
|
||||
if (credentialId != null) {
|
||||
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
|
||||
// fleetd #466 scope item 2: same one-call status() read as profilesView above — see its
|
||||
// comment for why this must not become two separate lookups.
|
||||
quarantine.quarantine().status(credentialId).ifPresent(status -> {
|
||||
row.put("free", 0);
|
||||
row.put("credentialId", credentialId);
|
||||
row.put("quarantinedForSeconds", remaining);
|
||||
row.put("quarantinedForSeconds", status.remainingSeconds());
|
||||
row.put("quarantineAttempt", status.repeatCount());
|
||||
});
|
||||
}
|
||||
String outageCredentialId = outage.credentialIdFor().apply(profile);
|
||||
@@ -1805,7 +1848,7 @@ public final class FleetMcp {
|
||||
// --- tool schemas --------------------------------------------------------------------------
|
||||
|
||||
private static McpSchema.Tool sendTool() {
|
||||
return tool("fleet_send",
|
||||
return tool(FleetTool.SEND.wireName(),
|
||||
"Delegate a task to a worker session. By default blocks until the worker replies and "
|
||||
+ "returns its reply (or a 'still working / queued' note on timeout). Pass wait:false "
|
||||
+ "for a long task to return a ticket immediately, then poll it with fleet_poll. To "
|
||||
@@ -1833,7 +1876,7 @@ public final class FleetMcp {
|
||||
|
||||
private static McpSchema.Tool askTool() {
|
||||
// No target/session arg — the worker's identity is resolved from the connection.
|
||||
return tool("fleet_ask",
|
||||
return tool(FleetTool.ASK.wireName(),
|
||||
"Pause your current delegated turn to ask the primary a question, blocking until it "
|
||||
+ "answers — then resume the same turn with the answer. Use this when only the "
|
||||
+ "primary has a decision or detail you need to continue. You do not address the "
|
||||
@@ -1846,7 +1889,7 @@ public final class FleetMcp {
|
||||
}
|
||||
|
||||
private static McpSchema.Tool pollTool() {
|
||||
return tool("fleet_poll",
|
||||
return tool(FleetTool.POLL.wireName(),
|
||||
"Check an async delegation (a fleet_send with wait:false) by its ticket: "
|
||||
+ "pending, done (with the worker's reply), or failed. When target (a worker "
|
||||
+ "session id) is present instead of ticket, drain that worker's inbox of "
|
||||
@@ -1866,7 +1909,7 @@ public final class FleetMcp {
|
||||
}
|
||||
|
||||
private static McpSchema.Tool ackTool() {
|
||||
return tool("fleet_ack",
|
||||
return tool(FleetTool.ACK.wireName(),
|
||||
"Acknowledge (remove) a specific reply from a worker's inbox. Use when the primary "
|
||||
+ "has processed a reply and wants to confirm it, leaving other pending replies "
|
||||
+ "in the inbox for later drain.",
|
||||
@@ -1880,7 +1923,7 @@ public final class FleetMcp {
|
||||
}
|
||||
|
||||
private static McpSchema.Tool spawnTool() {
|
||||
return tool("fleet_spawn",
|
||||
return tool(FleetTool.SPAWN.wireName(),
|
||||
"Spawn a new off-subscription member session. A member has two independent attributes: "
|
||||
+ "role (what it is for) and profile (which backend it runs on). Pass role to pick "
|
||||
+ "the contract — 'dev' implements a unit and opens its own PR, 'reviewer' reviews a "
|
||||
@@ -1912,7 +1955,7 @@ public final class FleetMcp {
|
||||
}
|
||||
|
||||
private static McpSchema.Tool profilesTool() {
|
||||
return tool("fleet_profiles",
|
||||
return tool(FleetTool.PROFILES.wireName(),
|
||||
"List the configured worker profiles (backends) and which one fleet_spawn uses by "
|
||||
+ "default. A 'quarantined' map is present when a backend-exhausted refusal put "
|
||||
+ "a profile's credential on cooldown — fleet_spawn onto it is refused until "
|
||||
@@ -1930,7 +1973,7 @@ public final class FleetMcp {
|
||||
}
|
||||
|
||||
private static McpSchema.Tool listTool() {
|
||||
return tool("fleet_list",
|
||||
return tool(FleetTool.LIST.wireName(),
|
||||
"List the whole fleet the bridge tracks, in two parts. 'leads' are your PEERS — other "
|
||||
+ "orchestrators, each with its sessionId (the address to fleet_send to), "
|
||||
+ "name, live status, and 'self': true on your own row; this is how you "
|
||||
@@ -1966,7 +2009,7 @@ public final class FleetMcp {
|
||||
}
|
||||
|
||||
private static McpSchema.Tool stopTool() {
|
||||
return tool("fleet_stop",
|
||||
return tool(FleetTool.STOP.wireName(),
|
||||
"Tear down a worker session by its paneId (from fleet_spawn or fleet_list).",
|
||||
objectSchema(Map.of(
|
||||
"paneId", stringProp("The worker's paneId to stop")),
|
||||
@@ -1975,7 +2018,7 @@ public final class FleetMcp {
|
||||
|
||||
private static McpSchema.Tool replyTool() {
|
||||
// No session/target arg — the caller's identity is resolved from the connection.
|
||||
return tool("fleet_reply",
|
||||
return tool(FleetTool.REPLY.wireName(),
|
||||
"Return your structured answer for a message you were sent, resolving the sender's "
|
||||
+ "blocked fleet_send. A worker MUST end every delegated turn with exactly "
|
||||
+ "one of these. A lead uses it only to answer another lead that messaged "
|
||||
@@ -1986,7 +2029,7 @@ public final class FleetMcp {
|
||||
}
|
||||
|
||||
private static McpSchema.Tool statusTool() {
|
||||
return tool("fleet_status",
|
||||
return tool(FleetTool.STATUS.wireName(),
|
||||
"Get the live lifecycle status (idle/working/blocked/unknown) of a worker session.",
|
||||
objectSchema(Map.of(
|
||||
"sessionId", stringProp("The worker session id to query")),
|
||||
@@ -1994,7 +2037,7 @@ public final class FleetMcp {
|
||||
}
|
||||
|
||||
private static McpSchema.Tool whoamiTool() {
|
||||
return tool("fleet_whoami",
|
||||
return tool(FleetTool.WHOAMI.wireName(),
|
||||
"Report who YOU are on the bridge — your role is resolved from your connection "
|
||||
+ "(unforgeable), never from anything you claim. Returns role 'primary' (you "
|
||||
+ "orchestrate: spawn/send/stop; reply ONLY to answer a peer lead that "
|
||||
|
||||
@@ -0,0 +1,82 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.Map;
|
||||
import java.util.Optional;
|
||||
import java.util.Set;
|
||||
|
||||
/**
|
||||
* The one canonical set of MCP tool names this daemon registers (fleetd #469, follow-up to #464).
|
||||
*
|
||||
* <p>Before this enum, the tool surface was written twice with nothing tying the copies together:
|
||||
* once as the literal {@code "fleet_…"} string passed to each tool-schema builder in
|
||||
* {@link FleetMcp}, and again as the case labels of {@link FleetMcp}'s authorization switch. A
|
||||
* reader that needed "what does this server register" — a charter check, in particular — had no
|
||||
* source to ask except scraping {@code FleetMcp.java}'s source text for {@code tool("…")} calls: a
|
||||
* third copy of the same list, and the weakest of the three forms.
|
||||
*
|
||||
* <p>Every reader that needs the registered tool surface now asks this enum instead:
|
||||
*
|
||||
* <ul>
|
||||
* <li>the tool-schema builders in {@code FleetMcp} pass {@code wireName()} rather than a literal;
|
||||
* <li>{@code FleetMcp}'s constructor asserts, at startup, that the set of tool names it actually
|
||||
* registers with the MCP SDK equals {@link #wireNames()} exactly — a canonical entry that is
|
||||
* never registered, or a registration with no canonical entry backing it, fails the daemon's
|
||||
* own boot rather than only a test's;
|
||||
* <li>{@code FleetMcp.toolAction}'s dispatch onto {@code Authz.Action} switches on the enum
|
||||
* (not the raw string) with no {@code default}, so adding a tool here without pinning its
|
||||
* action is a compile error, not a run-time throw;
|
||||
* <li>{@link CharterToolSurface} asks {@link #wireNames()} to check a configured launch charter
|
||||
* against the live tool surface, instead of scraping source text a third time.
|
||||
* </ul>
|
||||
*/
|
||||
public enum FleetTool {
|
||||
|
||||
SEND("fleet_send"),
|
||||
REPLY("fleet_reply"),
|
||||
ASK("fleet_ask"),
|
||||
STATUS("fleet_status"),
|
||||
POLL("fleet_poll"),
|
||||
ACK("fleet_ack"),
|
||||
SPAWN("fleet_spawn"),
|
||||
LIST("fleet_list"),
|
||||
STOP("fleet_stop"),
|
||||
PROFILES("fleet_profiles"),
|
||||
WHOAMI("fleet_whoami");
|
||||
|
||||
private final String wireName;
|
||||
|
||||
FleetTool(String wireName) {
|
||||
this.wireName = wireName;
|
||||
}
|
||||
|
||||
/** The name this tool is registered under, and called by, on the wire ({@code "fleet_send"}, …). */
|
||||
public String wireName() {
|
||||
return wireName;
|
||||
}
|
||||
|
||||
private static final Map<String, FleetTool> BY_WIRE_NAME;
|
||||
private static final Set<String> WIRE_NAMES;
|
||||
|
||||
static {
|
||||
Map<String, FleetTool> byName = new LinkedHashMap<>();
|
||||
Set<String> names = new LinkedHashSet<>();
|
||||
for (FleetTool tool : values()) {
|
||||
byName.put(tool.wireName, tool);
|
||||
names.add(tool.wireName);
|
||||
}
|
||||
BY_WIRE_NAME = Map.copyOf(byName);
|
||||
WIRE_NAMES = Set.copyOf(names);
|
||||
}
|
||||
|
||||
/** The tool named {@code wireName}, or empty when this daemon registers no such tool. */
|
||||
public static Optional<FleetTool> byWireName(String wireName) {
|
||||
return Optional.ofNullable(BY_WIRE_NAME.get(wireName));
|
||||
}
|
||||
|
||||
/** Every wire name this daemon registers — the canonical tool surface. */
|
||||
public static Set<String> wireNames() {
|
||||
return WIRE_NAMES;
|
||||
}
|
||||
}
|
||||
@@ -525,17 +525,28 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
charter.toFile().deleteOnExit();
|
||||
|
||||
// fleetd #393 follow-up: withArray, not putArray. putArray REPLACES whatever node
|
||||
// is already at "instructions" — harmless only as long as this block runs first
|
||||
// against a still-empty root, which is an ordering constraint nothing declared or
|
||||
// tested. The skills writer just below, and the IDE-rules writer further down,
|
||||
// both already use withArray (get-or-create) for exactly this reason; this was the
|
||||
// one straggler. Proven load-bearing on the fleetd #393 merge: flipping this one
|
||||
// call back to putArray left the whole suite green while silently deleting the
|
||||
// charter entry whenever skills or IDE rules ran after it — an opencode member
|
||||
// would launch with no role contract at all, worse than the bug #393 fixed, and
|
||||
// nothing caught it. See OpenCodeLauncherTest's
|
||||
// instructionsArrayHoldsCharterThenIdeRulesInOrder and
|
||||
// instructionsArrayHoldsCharterThenSkillsThenIdeRulesInOrder.
|
||||
// is already at "instructions"; withArray gets-or-creates. All three writers on
|
||||
// this array now use withArray so that write order stops being load-bearing for
|
||||
// whoever adds a fourth.
|
||||
//
|
||||
// Be precise about what this particular line is worth, because an earlier version
|
||||
// of this comment was wrong and claimed too much. THIS call is the one place where
|
||||
// the two idioms are equivalent, and no test can tell them apart: it runs first,
|
||||
// against a still-empty root, so there is never an existing node for putArray to
|
||||
// replace. Measured on the #393 merge: flipping this one call back to putArray
|
||||
// leaves the whole suite green (1618 tests), and always will. The edit is a
|
||||
// readability and future-proofing change with no test behind it, and that is not a
|
||||
// gap anyone can close.
|
||||
//
|
||||
// The ordering hazard is real, just not here. It is the LATER writers that can
|
||||
// destroy earlier entries. Measured on the same merge: flipping the skills writer
|
||||
// below to putArray deletes this charter entry and fails
|
||||
// OpenCodeLauncherTest.instructionsArrayHoldsCharterThenSkillsThenIdeRulesInOrder;
|
||||
// flipping the IDE-rules writer fails that test and
|
||||
// instructionsArrayHoldsCharterThenIdeRulesInOrder. Deleting this line altogether
|
||||
// fails instructionsArrayHoldsExactlyTheCharterWhenNothingElseWritesToIt — the
|
||||
// assertion that a role contract reaches an opencode member at all, which is the
|
||||
// hole that predates #393 and is what actually let the mutation hide.
|
||||
root.withArray("instructions").add(charter.toAbsolutePath().toString());
|
||||
}
|
||||
|
||||
|
||||
@@ -3,6 +3,7 @@ package dev.ltms.fleet.placement;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.Optional;
|
||||
import java.util.OptionalLong;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.function.LongSupplier;
|
||||
@@ -21,46 +22,158 @@ import java.util.function.LongSupplier;
|
||||
* <p>The clock is injected ({@link LongSupplier}, conventionally {@code System::nanoTime} like
|
||||
* {@code FleetHealthMonitor}), never read inline, so a quarantine's expiry is testable without a
|
||||
* real sleep.
|
||||
*
|
||||
* <h2>Escalation (fleetd #466)</h2>
|
||||
* A flat cooldown does not fit every exhaustion. A backend that reports "out of capacity for the
|
||||
* rest of the hour" recovers in one cooldown; a weekly subscription limit does not — it keeps
|
||||
* reporting exhausted on every attempt made before the window resets, so a flat 30-minute cooldown
|
||||
* (the default {@code cooldownNanos}) means roughly 336 pointless spawn attempts across a week, one
|
||||
* every cooldown.
|
||||
*
|
||||
* <p><strong>This class only ever sees the exhaustion signal.</strong> Its only production caller is
|
||||
* {@code Fleetd.exhaustionSink}, wired to fire on a {@code BACKEND_EXHAUSTED} classification alone.
|
||||
* The daemon's other outage state — a credential "cooling off" after repeated non-exhaustion
|
||||
* backend errors (an HTTP 5xx storm, say) — is a separate mechanism, {@code BackendOutagePolicy},
|
||||
* with its own short fixed 60s cooldown and no repeat tracking. The two are never merged: escalating
|
||||
* on a cooling-off signal would turn a transient 5xx storm into a multi-hour backoff, which is
|
||||
* exactly the failure this ticket is not asking for. Confirmed by reading every call site of
|
||||
* {@link #quarantine} — {@code BackendOutagePolicy} has its own {@code coolOff} method and never
|
||||
* calls this one.
|
||||
*
|
||||
* <p><strong>Mechanism</strong> — the {@link #withEscalation} constructors track, per credential, how
|
||||
* many times in a row {@link #quarantine} has been called without an intervening "quiet" gap.
|
||||
* Each call computes {@code cooldownNanos * backoffMultiplier ^ (repeatCount - 1)}, capped at
|
||||
* {@code maxCooldownNanos}. A call counts as a continuation of the same streak — {@code repeatCount}
|
||||
* increments — when it arrives no more than one base {@code cooldownNanos} after the previous
|
||||
* quarantine's deadline (this covers both "still quarantined" and "quarantine just expired and it
|
||||
* was exhausted again immediately"); otherwise the streak resets and this call is treated as a fresh
|
||||
* first occurrence at the base cooldown.
|
||||
*
|
||||
* <p><strong>Reset, honestly stated.</strong> The ideal reset signal is "the cooldown expired and the
|
||||
* next attempt succeeded" — but nothing in this codebase reports a spawn success back to this class
|
||||
* (checked: {@code SessionManager} and {@code CompositePeerLauncher} never call any method here
|
||||
* except {@link #quarantine}/{@link #isQuarantined}/{@link #remainingSeconds}, none of which is a
|
||||
* success hook). Lacking that signal, the reset used here is a time-based proxy: a base-cooldown's
|
||||
* worth of quiet — no exhaustion report for that credential — since the last quarantine ended. It is
|
||||
* not proof the credential started working again, only the best available evidence without adding an
|
||||
* active probe, which is out of scope by the operator's own design constraint (no automatic probing
|
||||
* of a limited backend).
|
||||
*
|
||||
* <p><strong>Ceiling.</strong> {@code maxCooldownNanos} bounds the growth — an unbounded backoff is a
|
||||
* permanent, unrecoverable-without-a-restart outage, which would be worse than the flat-rate bug this
|
||||
* escalation fixes. {@link #withEscalation(LongSupplier, long)} defaults the ceiling to
|
||||
* {@value #DEFAULT_MAX_COOLDOWN_MULTIPLE}x the base cooldown (12x the 1800s default ≈ 6 hours), so a
|
||||
* chronically exhausted credential still gets re-tried roughly every 6 hours instead of every 30
|
||||
* minutes — about a dozen attempts a week instead of ~336.
|
||||
*
|
||||
* <p><strong>Backward compatibility.</strong> The original two-argument {@link #BackendQuarantine(
|
||||
* LongSupplier, long)} constructor is unchanged in behaviour: it is exactly {@code
|
||||
* withEscalation}'s mechanism with {@code backoffMultiplier = 1.0} and {@code maxCooldownNanos =
|
||||
* cooldownNanos}, which collapses the formula back to the original flat {@code now + cooldownNanos}
|
||||
* on every call regardless of history. Every existing call site (roughly 20 across the test suite,
|
||||
* plus {@link #none()}) keeps its current shape and behaviour unchanged.
|
||||
*/
|
||||
public final class BackendQuarantine {
|
||||
|
||||
private final ConcurrentHashMap<String, Long> quarantinedUntilNanos = new ConcurrentHashMap<>();
|
||||
/** Default growth per consecutive exhaustion streak — see the class doc's Mechanism section. */
|
||||
static final double DEFAULT_BACKOFF_MULTIPLIER = 2.0;
|
||||
/** Default ceiling, expressed as a multiple of the base cooldown — see the class doc's Ceiling section. */
|
||||
static final long DEFAULT_MAX_COOLDOWN_MULTIPLE = 12;
|
||||
|
||||
private final ConcurrentHashMap<String, QuarantineState> quarantines = new ConcurrentHashMap<>();
|
||||
private final LongSupplier nowNanos;
|
||||
private final long cooldownNanos;
|
||||
private final double backoffMultiplier;
|
||||
private final long maxCooldownNanos;
|
||||
/** True only for {@link #none()}. See {@link #quarantine} for why this exists. */
|
||||
private final boolean inert;
|
||||
|
||||
/** How many consecutive exhaustion reports a credential is on, and when the resulting cooldown ends. */
|
||||
private record QuarantineState(int repeatCount, long deadlineNanos) {
|
||||
}
|
||||
|
||||
/**
|
||||
* Flat cooldown, unchanged from before fleetd #466 — every {@link #quarantine} call blocks the
|
||||
* credential for exactly {@code cooldownNanos}, regardless of how many times it was called
|
||||
* before. Equivalent to {@link #withEscalation} with no growth ({@code backoffMultiplier = 1.0})
|
||||
* and a ceiling equal to the base cooldown, so it degrades to the identical {@code now +
|
||||
* cooldownNanos} formula every call. Kept for the existing call sites that want a fixed cooldown
|
||||
* (and for tests exercising the fixed-cooldown shape in isolation); production wiring uses
|
||||
* {@link #withEscalation} instead.
|
||||
*
|
||||
* @param nowNanos monotonic clock, injected for testability
|
||||
* @param cooldownNanos how long a fresh {@link #quarantine} call blocks the credential for;
|
||||
* must be positive
|
||||
*/
|
||||
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos) {
|
||||
this(nowNanos, cooldownNanos, false);
|
||||
this(nowNanos, cooldownNanos, 1.0, cooldownNanos, false);
|
||||
}
|
||||
|
||||
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, boolean inert) {
|
||||
/**
|
||||
* Escalating cooldown (fleetd #466) — see the class doc's Mechanism/Reset/Ceiling sections.
|
||||
*
|
||||
* @param nowNanos monotonic clock, injected for testability
|
||||
* @param cooldownNanos base cooldown, applied to a fresh (non-streak) exhaustion; must be
|
||||
* positive
|
||||
* @param backoffMultiplier growth per consecutive exhaustion; must be {@code >= 1.0} ({@code 1.0}
|
||||
* disables growth and is exactly the flat two-argument constructor)
|
||||
* @param maxCooldownNanos ceiling on the escalated cooldown; must be {@code >= cooldownNanos}
|
||||
*/
|
||||
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, double backoffMultiplier,
|
||||
long maxCooldownNanos) {
|
||||
this(nowNanos, cooldownNanos, backoffMultiplier, maxCooldownNanos, false);
|
||||
}
|
||||
|
||||
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, double backoffMultiplier,
|
||||
long maxCooldownNanos, boolean inert) {
|
||||
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
|
||||
if (cooldownNanos <= 0) {
|
||||
throw new IllegalArgumentException("cooldownNanos must be positive: " + cooldownNanos);
|
||||
}
|
||||
if (backoffMultiplier < 1.0) {
|
||||
throw new IllegalArgumentException("backoffMultiplier must be >= 1.0: " + backoffMultiplier);
|
||||
}
|
||||
if (maxCooldownNanos < cooldownNanos) {
|
||||
throw new IllegalArgumentException(
|
||||
"maxCooldownNanos must be >= cooldownNanos: " + maxCooldownNanos + " < " + cooldownNanos);
|
||||
}
|
||||
this.cooldownNanos = cooldownNanos;
|
||||
this.backoffMultiplier = backoffMultiplier;
|
||||
this.maxCooldownNanos = maxCooldownNanos;
|
||||
this.inert = inert;
|
||||
}
|
||||
|
||||
/**
|
||||
* Escalating cooldown with the fleetd #466 default shape: cooldown doubles
|
||||
* ({@value #DEFAULT_BACKOFF_MULTIPLIER}x) per consecutive exhaustion streak, capped at
|
||||
* {@value #DEFAULT_MAX_COOLDOWN_MULTIPLE}x the base cooldown. This is what production wiring
|
||||
* ({@code Fleetd.main}) uses.
|
||||
*
|
||||
* @param nowNanos monotonic clock, injected for testability
|
||||
* @param cooldownNanos base cooldown, applied to a fresh (non-streak) exhaustion; must be positive
|
||||
*/
|
||||
public static BackendQuarantine withEscalation(LongSupplier nowNanos, long cooldownNanos) {
|
||||
return new BackendQuarantine(nowNanos, cooldownNanos, DEFAULT_BACKOFF_MULTIPLIER,
|
||||
cooldownNanos * DEFAULT_MAX_COOLDOWN_MULTIPLE, false);
|
||||
}
|
||||
|
||||
/**
|
||||
* Inert quarantine — {@link #quarantine} does nothing on this instance, so nothing is ever
|
||||
* quarantined. The explicit stand-in a caller (or a test not exercising this feature) passes
|
||||
* instead of a defaulting overload, exactly like {@code ExhaustedPatternLookup.none()}.
|
||||
*/
|
||||
public static BackendQuarantine none() {
|
||||
return new BackendQuarantine(() -> 0L, 1, true);
|
||||
return new BackendQuarantine(() -> 0L, 1, 1.0, 1, true);
|
||||
}
|
||||
|
||||
/**
|
||||
* Quarantine {@code credentialId} for the configured cooldown, starting now. A repeat call while
|
||||
* already quarantined restarts the cooldown at full length — a fresh refusal is fresh evidence the
|
||||
* account is still exhausted, not a reason to let an earlier, shorter wait stand.
|
||||
* Quarantine {@code credentialId} starting now. On a flat instance (the two-argument
|
||||
* constructor) this always blocks for exactly {@code cooldownNanos}, restarting the cooldown at
|
||||
* full length on every call — a fresh refusal is fresh evidence the account is still exhausted,
|
||||
* not a reason to let an earlier, shorter wait stand. On an escalating instance ({@link
|
||||
* #withEscalation}) the cooldown grows with each call that arrives within one base cooldown of
|
||||
* the previous deadline, and resets to the base cooldown once a call arrives after a longer gap
|
||||
* — see the class doc.
|
||||
*
|
||||
* <p>On {@link #none()} this is a no-op. It has to be: that instance holds a clock frozen at 0,
|
||||
* so recording a deadline would produce a quarantine that never expires — a credential locked out
|
||||
@@ -73,7 +186,13 @@ public final class BackendQuarantine {
|
||||
if (inert) {
|
||||
return;
|
||||
}
|
||||
quarantinedUntilNanos.put(credentialId, nowNanos.getAsLong() + cooldownNanos);
|
||||
long now = nowNanos.getAsLong();
|
||||
quarantines.compute(credentialId, (id, prev) -> {
|
||||
int repeatCount = (prev == null || now - prev.deadlineNanos() > cooldownNanos)
|
||||
? 1
|
||||
: prev.repeatCount() + 1;
|
||||
return new QuarantineState(repeatCount, now + escalatedCooldownNanos(repeatCount));
|
||||
});
|
||||
}
|
||||
|
||||
/** Whether {@code credentialId} is quarantined right now. */
|
||||
@@ -87,6 +206,49 @@ public final class BackendQuarantine {
|
||||
return remaining > 0 ? OptionalLong.of(toSecondsRoundedUp(remaining)) : OptionalLong.empty();
|
||||
}
|
||||
|
||||
/**
|
||||
* Remaining seconds together with which consecutive exhaustion this is (fleetd #466 scope item
|
||||
* 2) — {@code repeatCount} 1 for a first occurrence, 2 for the second in a row, and so on; see
|
||||
* {@link #quarantine}'s class-doc Mechanism section for exactly when a call continues a streak
|
||||
* versus starts a fresh one.
|
||||
*
|
||||
* <p><strong>Read together, off the one {@link QuarantineState} entry {@link #quarantine} itself
|
||||
* wrote</strong> — a single {@code quarantines.get(credentialId)}, never a separate lookup or a
|
||||
* value re-derived from {@code remainingSeconds} (e.g. inverting {@link
|
||||
* #escalatedCooldownNanos}). That inversion is not just extra work to avoid: once a streak has
|
||||
* hit {@code maxCooldownNanos}, every further consecutive exhaustion reports the identical
|
||||
* cooldown, so a derivation that starts from the cooldown value cannot tell the 4th repeat from
|
||||
* the 9th — only the stored {@code repeatCount} can. This is the same rule {@code
|
||||
* CompositePeerLauncher.modelGateState()} documents for its own gate/report pair: the report
|
||||
* reads the exact accessor the behaviour reads, so it can never disagree with what actually
|
||||
* happened (the fleetd #404/#422 lesson). {@code fleet_profiles}/{@code fleet_list}/{@code GET
|
||||
* /profiles} all call this — never {@link #remainingSeconds} plus a second, independent count —
|
||||
* for exactly that reason.
|
||||
*
|
||||
* @return empty when {@code credentialId} is not currently quarantined (including on {@link
|
||||
* #none()}, which quarantines nothing)
|
||||
*/
|
||||
public Optional<Status> status(String credentialId) {
|
||||
QuarantineState state = quarantines.get(credentialId);
|
||||
if (state == null) {
|
||||
return Optional.empty();
|
||||
}
|
||||
long remaining = state.deadlineNanos() - nowNanos.getAsLong();
|
||||
return remaining > 0
|
||||
? Optional.of(new Status(toSecondsRoundedUp(remaining), state.repeatCount()))
|
||||
: Optional.empty();
|
||||
}
|
||||
|
||||
/**
|
||||
* @param remainingSeconds seconds left on the quarantine, identical to {@link
|
||||
* #remainingSeconds(String)}'s answer for the same credential at the
|
||||
* same instant
|
||||
* @param repeatCount 1 for a first occurrence, 2 for the second consecutive one, etc. —
|
||||
* see {@link #status(String)}
|
||||
*/
|
||||
public record Status(long remainingSeconds, int repeatCount) {
|
||||
}
|
||||
|
||||
/**
|
||||
* Every currently-quarantined credential id and its remaining seconds (CB-578 stage B fleet
|
||||
* reporting) — expired entries are never included. Not pruned from the backing map here: it stays
|
||||
@@ -95,8 +257,8 @@ public final class BackendQuarantine {
|
||||
*/
|
||||
public Map<String, Long> activeRemainingSeconds() {
|
||||
Map<String, Long> out = new LinkedHashMap<>();
|
||||
quarantinedUntilNanos.forEach((credentialId, deadline) -> {
|
||||
long remaining = deadline - nowNanos.getAsLong();
|
||||
quarantines.forEach((credentialId, state) -> {
|
||||
long remaining = state.deadlineNanos() - nowNanos.getAsLong();
|
||||
if (remaining > 0) {
|
||||
out.put(credentialId, toSecondsRoundedUp(remaining));
|
||||
}
|
||||
@@ -105,8 +267,14 @@ public final class BackendQuarantine {
|
||||
}
|
||||
|
||||
private long remainingNanos(String credentialId) {
|
||||
Long deadline = quarantinedUntilNanos.get(credentialId);
|
||||
return deadline == null ? 0L : deadline - nowNanos.getAsLong();
|
||||
QuarantineState state = quarantines.get(credentialId);
|
||||
return state == null ? 0L : state.deadlineNanos() - nowNanos.getAsLong();
|
||||
}
|
||||
|
||||
/** {@code cooldownNanos * backoffMultiplier ^ (repeatCount - 1)}, capped at {@code maxCooldownNanos}. */
|
||||
private long escalatedCooldownNanos(int repeatCount) {
|
||||
double raw = cooldownNanos * Math.pow(backoffMultiplier, repeatCount - 1);
|
||||
return raw >= (double) maxCooldownNanos ? maxCooldownNanos : (long) raw;
|
||||
}
|
||||
|
||||
private static long toSecondsRoundedUp(long nanos) {
|
||||
|
||||
@@ -0,0 +1,65 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #466 follow-up: {@code Fleetd.main} builds the daemon's one {@code BackendQuarantine}
|
||||
* from {@link dev.ltms.fleet.placement.BackendQuarantine#withEscalation(java.util.function.LongSupplier,
|
||||
* long)} — the escalating factory — rather than the plain two-argument constructor, which is still a
|
||||
* flat cooldown (kept for backward compatibility, see that class's doc). {@code
|
||||
* BackendQuarantineTest} proves {@code withEscalation} itself escalates, is ceilinged, and resets;
|
||||
* it says nothing about which one {@code main} actually calls.
|
||||
*
|
||||
* <p>Measured directly: reverting {@code main} to {@code new BackendQuarantine(System::nanoTime,
|
||||
* TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()))} — the pre-#466 flat call — compiles
|
||||
* with 0 errors and leaves the entire 1608-test suite (including every {@code BackendQuarantineTest}
|
||||
* case) green, because no other test constructs its {@code BackendQuarantine} through {@code main};
|
||||
* every one of them builds its own instance directly. That silent regression is exactly the shape
|
||||
* {@link FleetdLeadSeatWiringTest} and {@link FleetdCompletionResolverWiringTest} already guard
|
||||
* against for their own constructor arguments — this is the same class of gap for fleetd #466's
|
||||
* factory choice, following their approach.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
||||
* BackendQuarantine} and never runs {@code main} — a green result here proves only that the exact
|
||||
* text {@code main} calls {@code BackendQuarantine.withEscalation(...)} rather than the flat
|
||||
* constructor. It does not prove that call actually executes at startup (no test here starts the
|
||||
* daemon), and it does not prove the escalation reaches a real backend or credential — only
|
||||
* {@code BackendQuarantineTest} proves the factory's own behaviour, and only a live daemon proves
|
||||
* the wiring runs.
|
||||
*/
|
||||
class FleetdBackendQuarantineWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] main's BackendQuarantine local is still built from BackendQuarantine.withEscalation(...)")
|
||||
void mainStillWiresTheEscalatingQuarantineFactory() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,\n"
|
||||
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
|
||||
"Fleetd.main's BackendQuarantine local must still be built from "
|
||||
+ "BackendQuarantine.withEscalation(System::nanoTime, "
|
||||
+ "TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds())). Reverting to the flat "
|
||||
+ "two-argument constructor (fleetd #466's measured regression) compiles with 0 errors "
|
||||
+ "and leaves the whole suite green, including every BackendQuarantineTest case that "
|
||||
+ "proves the escalation itself works — this source check is what must go red instead. "
|
||||
+ "A reverted daemon would go back to retrying a weekly subscription limit on every "
|
||||
+ "flat ~30-minute cooldown, about 336 times across the week.");
|
||||
|
||||
// Negative form of the same check: the pre-#466 flat call, if it ever reappears at this
|
||||
// declaration, must not be mistaken for the escalating one by a looser positive-only check.
|
||||
assertFalse(source.contains(
|
||||
"BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,\n"
|
||||
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
|
||||
"main's BackendQuarantine local must never regress to the flat two-argument constructor");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,106 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertSame;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #474: proves the exact wiring {@code Fleetd.main} uses to construct its live {@code
|
||||
* ConfigRef} — {@code new ConfigRef(configPath, cfg, Fleetd::assertChartersNameOnlyRegisteredTools)}
|
||||
* — actually makes {@link ConfigRef#reload()} refuse a charter that names an MCP tool the server
|
||||
* does not register, the same way {@code Fleetd.main} itself refuses one at startup (see {@code
|
||||
* FleetdStartupValidationTest#mainRefusesACharterNamingAnUnregisteredTool}).
|
||||
*
|
||||
* <p>{@code dev.ltms.fleet.config.ConfigRefTest} pins the same behaviour through a locally-built
|
||||
* {@code Consumer<FleetConfig>} adapter that calls the same production {@code CharterToolSurface}
|
||||
* method, because that test lives in {@code dev.ltms.fleet.config} and cannot see {@code
|
||||
* Fleetd#assertChartersNameOnlyRegisteredTools} (package-private to {@code dev.ltms.fleet}). This
|
||||
* class is the companion proof that lives where the real method reference is visible, so the literal
|
||||
* expression {@code Fleetd::assertChartersNameOnlyRegisteredTools} — not just an equivalent — is
|
||||
* what gets exercised. {@code Fleetd.main} itself cannot be driven this far in a unit test: every
|
||||
* fixture in {@code FleetdStartupValidationTest} is deliberately invalid so {@code main} throws
|
||||
* before opening a socket, binding Javalin, or doing anything else with a real side effect, so a
|
||||
* test cannot get {@code main} far enough to hold a running daemon it could then reload — this test
|
||||
* builds the {@code ConfigRef} the same way {@code main} does and drives {@link ConfigRef#reload()}
|
||||
* directly instead, the same shape {@code FleetdExhaustionDetectionArmedWiringTest} and its
|
||||
* siblings already use for the rest of {@code Fleetd.main}'s wiring.
|
||||
*/
|
||||
class FleetdConfigRefCharterToolSurfaceWiringTest {
|
||||
|
||||
private static final String BASE = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
@Test
|
||||
void reloadRefusesACharterNamingAnUnregisteredToolThroughFleetdsOwnWiring(@TempDir Path dir)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, BASE + """
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through fleet_reply.
|
||||
""");
|
||||
ConfigRef config = new ConfigRef(f, FleetConfig.load(f),
|
||||
Fleetd::assertChartersNameOnlyRegisteredTools);
|
||||
FleetConfig before = config.get();
|
||||
|
||||
Files.writeString(f, BASE + """
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through bridge_send.
|
||||
""");
|
||||
ConfigRef.Outcome out = config.reload();
|
||||
|
||||
assertFalse(out.applied());
|
||||
assertNotNull(out.error());
|
||||
assertTrue(out.error().contains("dev"), out.error());
|
||||
assertTrue(out.error().contains("bridge_send"), out.error());
|
||||
assertSame(before, config.get());
|
||||
}
|
||||
|
||||
@Test
|
||||
void reloadAcceptsACharterNamingOnlyRegisteredToolsThroughFleetdsOwnWiring(@TempDir Path dir)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, BASE + """
|
||||
fleet:
|
||||
charters:
|
||||
dev: old charter
|
||||
""");
|
||||
ConfigRef config = new ConfigRef(f, FleetConfig.load(f),
|
||||
Fleetd::assertChartersNameOnlyRegisteredTools);
|
||||
|
||||
Files.writeString(f, BASE + """
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through fleet_reply.
|
||||
""");
|
||||
ConfigRef.Outcome out = config.reload();
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertEquals("config reloaded", out.summary());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,80 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #474 follow-up: {@code Fleetd.main} builds its live {@code ConfigRef} from the
|
||||
* three-argument constructor, {@code new ConfigRef(configPath, cfg,
|
||||
* Fleetd::assertChartersNameOnlyRegisteredTools)}, so a reload runs the same charter tool-surface
|
||||
* gate startup does (see {@link dev.ltms.fleet.config.ConfigRef}'s class doc, "fleetd #474" bullet).
|
||||
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} prove the
|
||||
* three-argument constructor and the {@code Fleetd.assertChartersNameOnlyRegisteredTools} adapter
|
||||
* work correctly together — both build their OWN {@code ConfigRef} with that constructor, so neither
|
||||
* proves {@code main} still CHOOSES the three-argument form over the plain two-argument {@code new
|
||||
* ConfigRef(configPath, cfg)}.
|
||||
*
|
||||
* <p>Measured directly: reverting {@code Fleetd.java}'s {@code config} local to the two-argument
|
||||
* constructor compiles with 0 errors and leaves the entire 1633-test suite green — including every
|
||||
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} case — because
|
||||
* neither of those tests constructs its {@code ConfigRef} through {@code main}; both build their own
|
||||
* instance directly, wired with the check by hand. That silent regression is exactly the shape
|
||||
* {@link FleetdBackendQuarantineWiringTest}, {@link FleetdLeadSeatWiringTest} and {@link
|
||||
* FleetdCompletionResolverWiringTest} already guard against for their own constructor arguments —
|
||||
* this class is the same class of gap for fleetd #474's {@code extraValidation} argument, following
|
||||
* their approach.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
||||
* ConfigRef} and never runs {@code main} — a green result here proves only that the exact text
|
||||
* {@code main} contains is the three-argument construction with {@code
|
||||
* Fleetd::assertChartersNameOnlyRegisteredTools}. It does not prove that call actually executes at
|
||||
* startup (no test here starts the daemon), and it does not prove the reload gate itself works —
|
||||
* only {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} prove the
|
||||
* behaviour; only a live daemon proves the wiring runs.
|
||||
*/
|
||||
class FleetdConfigRefWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] main still builds config from the three-argument ConfigRef constructor")
|
||||
void mainStillWiresTheThreeArgumentConfigRefConstructor() throws Exception {
|
||||
String source = fleetdSource();
|
||||
|
||||
// A broken read (wrong working directory, wrong path, a file that came back empty) would
|
||||
// make the assertFalse below pass vacuously — a "clean" negative check that actually
|
||||
// checked nothing. Guard against that first, with an anchor that has nothing to do with
|
||||
// this mutation, so a bad read fails loudly here instead of silently proving nothing below.
|
||||
assertTrue(source.contains("public final class Fleetd"),
|
||||
"fleetdSource() did not read anything usable — src/main/java/dev/ltms/fleet/Fleetd.java "
|
||||
+ "did not come back containing its own class declaration. The assertFalse below "
|
||||
+ "would pass vacuously on a broken read; fix the read before trusting this test.");
|
||||
|
||||
assertTrue(source.contains(
|
||||
"ConfigRef config = new ConfigRef(configPath, cfg, "
|
||||
+ "Fleetd::assertChartersNameOnlyRegisteredTools);"),
|
||||
"Fleetd.main's config local must still be built from the three-argument ConfigRef "
|
||||
+ "constructor, with Fleetd::assertChartersNameOnlyRegisteredTools as "
|
||||
+ "extraValidation. Reverting to the plain two-argument constructor (fleetd #474's "
|
||||
+ "measured M2 regression) compiles with 0 errors and leaves the whole suite green — "
|
||||
+ "including ConfigRefTest and FleetdConfigRefCharterToolSurfaceWiringTest, because "
|
||||
+ "neither builds its ConfigRef through main — this source check is what must go "
|
||||
+ "red instead. A reverted daemon would accept, through a reload with no restart, "
|
||||
+ "exactly the charter that #469/#474 already refuse at startup.");
|
||||
|
||||
// Negative form of the same check: the pre-#474 two-argument call, if it ever reappears at
|
||||
// this declaration, must not be mistaken for the three-argument one by a looser
|
||||
// positive-only check — this is the M2 mutation this test exists to kill.
|
||||
assertFalse(source.contains("ConfigRef config = new ConfigRef(configPath, cfg);"),
|
||||
"main's config local must never regress to the plain two-argument ConfigRef "
|
||||
+ "constructor — that drops the reload-path charter check (fleetd #474's measured "
|
||||
+ "M2 mutation) with no other test catching it");
|
||||
}
|
||||
}
|
||||
@@ -95,6 +95,28 @@ class FleetdStartupValidationTest {
|
||||
""", "architetc");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #469: closes the gap left by #464's {@code CharterToolSurfaceTest}, which wrote its
|
||||
* own charter into a {@code @TempDir} fixture and so could never fail on anything anyone wrote
|
||||
* into the live {@code fleetd.yaml}. {@code bridge_send} is CB-634's own motivating example — a
|
||||
* tool name the rename removed — and the charter key ({@code dev}) is a real role wire name, so
|
||||
* this fixture passes {@code cfg.validateAll()}'s charter check (key valid, text non-blank) and
|
||||
* is refused only by the new {@code CharterToolSurface} call right after it. The failure message
|
||||
* must name both the charter key and the unknown tool.
|
||||
*/
|
||||
@Test
|
||||
void mainRefusesACharterNamingAnUnregisteredTool(@TempDir Path dir) throws Exception {
|
||||
assertMainRefuses(dir, "charter-tool-surface.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through bridge_send.
|
||||
""", "bridge_send");
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainRefusesAnArchitectSlotNamingAnUnconfiguredProfile(@TempDir Path dir) throws Exception {
|
||||
assertMainRefuses(dir, "members.yaml", """
|
||||
|
||||
@@ -1,10 +1,13 @@
|
||||
package dev.ltms.fleet.config;
|
||||
|
||||
import dev.ltms.fleet.mcp.CharterToolSurface;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.function.Consumer;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
@@ -38,6 +41,23 @@ class ConfigRefTest {
|
||||
return new ConfigRef(f, FleetConfig.load(f));
|
||||
}
|
||||
|
||||
/**
|
||||
* The exact {@code Consumer<FleetConfig>} {@code Fleetd.main} wires into {@code ConfigRef}'s
|
||||
* constructor as {@code extraValidation} (fleetd #474) — an adapter from {@code FleetConfig} to
|
||||
* the raw charter map {@link CharterToolSurface#assertChartersNameOnlyRegisteredTools} takes.
|
||||
* Built here rather than referencing {@code dev.ltms.fleet.Fleetd} directly, because that method
|
||||
* is package-private to {@code dev.ltms.fleet} and this test lives in {@code
|
||||
* dev.ltms.fleet.config} — but it calls the SAME production {@link CharterToolSurface} method
|
||||
* {@code Fleetd} calls, so this proves the real check runs on reload, not a stand-in for it.
|
||||
*/
|
||||
private static final Consumer<FleetConfig> CHARTER_TOOL_SURFACE = cfg ->
|
||||
CharterToolSurface.assertChartersNameOnlyRegisteredTools(
|
||||
cfg.fleet() == null ? Map.of() : cfg.fleet().charters());
|
||||
|
||||
private static ConfigRef refForWithCharterToolSurface(Path f) {
|
||||
return new ConfigRef(f, FleetConfig.load(f), CHARTER_TOOL_SURFACE);
|
||||
}
|
||||
|
||||
@Test
|
||||
void aHotChangeIsAppliedAndReadThroughGet(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
@@ -124,6 +144,87 @@ class ConfigRefTest {
|
||||
assertSame(before, ref.get());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #474: {@code validateAll()} (via {@code validateCharters()}) only checks that a
|
||||
* charter's KEY is a role wire name and its text is non-blank — it never looks at what the text
|
||||
* names, so a charter naming {@code bridge_send} (the pre-CB-634 name, removed from the tool
|
||||
* surface — #469's own motivating example) passes {@code validateAll()} and used to be applied
|
||||
* on reload with nothing refusing it, even though the identical charter refuses {@code
|
||||
* Fleetd.main} at startup ({@code FleetdStartupValidationTest
|
||||
* #mainRefusesACharterNamingAnUnregisteredTool}). This is the reload-path proof: it drives
|
||||
* {@link ConfigRef#reload()} itself (not a direct call to {@link
|
||||
* CharterToolSurface#assertChartersNameOnlyRegisteredTools}), through the exact {@code
|
||||
* extraValidation} wiring {@code Fleetd.main} uses, and the failure message must name both the
|
||||
* charter key and the unknown tool — the same information the startup failure gives (ticket
|
||||
* acceptance criterion 1).
|
||||
*/
|
||||
@Test
|
||||
void aReloadRefusesACharterNamingAnUnregisteredTool(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through fleet_reply.
|
||||
"""));
|
||||
ConfigRef ref = refForWithCharterToolSurface(f);
|
||||
FleetConfig before = ref.get();
|
||||
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through bridge_send.
|
||||
"""));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertFalse(out.applied());
|
||||
assertNotNull(out.error());
|
||||
assertTrue(out.error().contains("dev"),
|
||||
"expected the charter key 'dev' in the refusal, got: " + out.error());
|
||||
assertTrue(out.error().contains("bridge_send"),
|
||||
"expected the unknown tool 'bridge_send' in the refusal, got: " + out.error());
|
||||
assertTrue(out.summary().startsWith("config reload refused"), out.summary());
|
||||
// The running config must not move at all — half-applying this would leave the daemon in a
|
||||
// state that could never have booted, exactly the outcome ConfigRef.java's class doc warns
|
||||
// a cold-key refusal must avoid, and this check must avoid the same way.
|
||||
assertSame(before, ref.get());
|
||||
assertEquals("Send the final handoff through fleet_reply.\n",
|
||||
ref.get().fleet().charterFor(dev.ltms.fleet.peer.MemberRole.DEV));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #474 acceptance criterion 3, the positive case: a reload whose charter names only
|
||||
* registered tools must still be ACCEPTED. An inverted filter (one that refuses every charter,
|
||||
* or refuses on any {@code fleet_*}/{@code bridge_*} token regardless of registration) would pass
|
||||
* the refusal test above alone — #469's own M5 mutation cell showed exactly that shape surviving
|
||||
* a negative-only suite. This is the test that catches it.
|
||||
*/
|
||||
@Test
|
||||
void aReloadAcceptsACharterNamingOnlyRegisteredTools(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
charters:
|
||||
dev: old charter, no tool names
|
||||
"""));
|
||||
ConfigRef ref = refForWithCharterToolSurface(f);
|
||||
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through fleet_reply, using fleet_send to delegate.
|
||||
"""));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
|
||||
assertEquals("config reloaded", out.summary());
|
||||
assertEquals("Send the final handoff through fleet_reply, using fleet_send to delegate.\n",
|
||||
ref.get().fleet().charterFor(dev.ltms.fleet.peer.MemberRole.DEV));
|
||||
}
|
||||
|
||||
/**
|
||||
* The point of the whole class: a consumer holding the ref sees the new value without being
|
||||
* rebuilt. A component that captured {@code get()} into a field would still show the old one.
|
||||
|
||||
@@ -3,7 +3,9 @@ package dev.ltms.fleet.mcp;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
@@ -11,13 +13,30 @@ import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/** fleetd #464: launch charters must not name MCP tools the server does not register. */
|
||||
/**
|
||||
* fleetd #464 shipped {@code configuredChartersNameOnlyRegisteredTools} below, comparing charter
|
||||
* text against the registered tool surface — but it scraped <em>both</em> sides from source text:
|
||||
* its own fixture charter, and a regex over {@code FleetMcp.java}'s {@code tool("…")} calls. #469's
|
||||
* gap: nothing anyone wrote into the live {@code fleetd.yaml} could ever reach that test, because
|
||||
* it never called production validation code.
|
||||
*
|
||||
* <p>This version keeps the charter-text extraction helper ({@code toolsNamedIn}) — charters are
|
||||
* free-text config, so finding a {@code fleet_*}/{@code bridge_*} token inside one has no source
|
||||
* but a scrape — but reads the <em>registered</em> side from {@link FleetTool}, the canonical enum
|
||||
* {@code FleetMcp} itself now derives its tool schemas and authorization switch from, rather than a
|
||||
* second scrape of {@code FleetMcp.java}'s source. It also exercises {@link
|
||||
* CharterToolSurface#assertChartersNameOnlyRegisteredTools} directly — the method {@code
|
||||
* Fleetd.main} actually calls at startup — both accepting and rejecting. {@code
|
||||
* dev.ltms.fleet.FleetdStartupValidationTest#mainRefusesACharterNamingAnUnregisteredTool} is what
|
||||
* closes #469's actual gap: it proves that call is wired into {@code Fleetd.main} itself, against a
|
||||
* live-shaped config fixture, not just a unit call to the method in isolation.
|
||||
*/
|
||||
class CharterToolSurfaceTest {
|
||||
|
||||
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
|
||||
|
||||
private static Set<String> matches(String text, String regex) {
|
||||
Matcher m = Pattern.compile(regex).matcher(text);
|
||||
Set<String> found = new LinkedHashSet<>();
|
||||
@@ -33,13 +52,8 @@ class CharterToolSurfaceTest {
|
||||
"(fleet_[a-z_]+|bridge_[a-z_]+)");
|
||||
}
|
||||
|
||||
/** Every tool {@link FleetMcp} registers, read from its {@code tool("…")} calls. */
|
||||
private static Set<String> toolsTheServerRegisters() throws Exception {
|
||||
return matches(Files.readString(MCP_SOURCE), "tool\\(\\\"(fleet_[a-z_]+)\\\"");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] every tool named in a configured charter is registered by the server")
|
||||
@DisplayName("every tool named in a configured charter is in the canonical FleetTool set")
|
||||
void configuredChartersNameOnlyRegisteredTools(@TempDir Path dir) throws Exception {
|
||||
Path configFile = dir.resolve("charters.yaml");
|
||||
Files.writeString(configFile, """
|
||||
@@ -53,20 +67,58 @@ class CharterToolSurfaceTest {
|
||||
|
||||
FleetConfig config = FleetConfig.load(configFile);
|
||||
Set<String> named = toolsNamedIn(config);
|
||||
Set<String> registered = toolsTheServerRegisters();
|
||||
Set<String> registered = FleetTool.wireNames();
|
||||
|
||||
assertTrue(!named.isEmpty(),
|
||||
"the charter fixture named no fleet_* or bridge_* tool. This test would check nothing; "
|
||||
+ "add charter text that names a tool before changing the extraction.");
|
||||
assertTrue(!registered.isEmpty(),
|
||||
"the FleetMcp registration scrape found no tools. This test would check nothing; "
|
||||
+ "repair the tool(\"…\") extraction before changing the assertion.");
|
||||
"FleetTool.wireNames() is empty. This test would check nothing; repair FleetTool "
|
||||
+ "before changing the assertion.");
|
||||
|
||||
Set<String> unknown = new LinkedHashSet<>(named);
|
||||
unknown.removeAll(registered);
|
||||
assertTrue(unknown.isEmpty(),
|
||||
"configured charter text names " + unknown + ", but FleetMcp does not register it. "
|
||||
"configured charter text names " + unknown + ", but FleetTool does not list it. "
|
||||
+ "Checked " + named + " against " + registered + ". Fix the charter text or "
|
||||
+ "register the tool; do NOT weaken this test.");
|
||||
+ "add the tool to FleetTool; do NOT weaken this test.");
|
||||
}
|
||||
|
||||
/**
|
||||
* The positive case for the actual production entry point: a charter naming only tools
|
||||
* {@link FleetTool} lists must not throw.
|
||||
*/
|
||||
@Test
|
||||
@DisplayName("CharterToolSurface accepts a charter that names only registered tools")
|
||||
void charterToolSurfaceAcceptsKnownTools() {
|
||||
assertDoesNotThrow(() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(Map.of(
|
||||
"dev", "Send the final handoff through fleet_reply, using fleet_send to delegate.",
|
||||
"reviewer", "Use fleet_ask only for the lead's decision.")));
|
||||
}
|
||||
|
||||
/**
|
||||
* The negative case for the actual production entry point (fleetd #469's motivating example:
|
||||
* {@code bridge_send} is the pre-CB-634 name, removed from the tool surface). The message must
|
||||
* name both the offending charter key and the unknown tool, so an operator reading the startup
|
||||
* log knows exactly which charter to fix.
|
||||
*/
|
||||
@Test
|
||||
@DisplayName("CharterToolSurface rejects a charter naming a tool the server does not register")
|
||||
void charterToolSurfaceRejectsAnUnregisteredTool() {
|
||||
Map<String, String> charters = new LinkedHashMap<>();
|
||||
charters.put("dev", "Send the final handoff through bridge_send.");
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class,
|
||||
() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(charters));
|
||||
assertTrue(e.getMessage().contains("dev"),
|
||||
"expected the charter key 'dev' in the failure message, got: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("bridge_send"),
|
||||
"expected the unknown tool 'bridge_send' in the failure message, got: " + e.getMessage());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("CharterToolSurface is a no-op on an absent or empty charter map")
|
||||
void charterToolSurfaceIsANoOpWithNoCharters() {
|
||||
assertDoesNotThrow(() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(null));
|
||||
assertDoesNotThrow(() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(Map.of()));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -26,7 +26,6 @@ import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.regex.Matcher;
|
||||
@@ -50,7 +49,6 @@ import static org.junit.jupiter.api.Assertions.*;
|
||||
class FleetMcpAuthzTest {
|
||||
|
||||
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
|
||||
private static final Pattern TOOL_REGISTRATION = Pattern.compile("tool\\(\\\"(fleet_[a-z_]+)\\\"");
|
||||
|
||||
private final FakeHerdr herdr = new FakeHerdr();
|
||||
private final AgentControl agents = new AgentControl(herdr);
|
||||
@@ -296,10 +294,18 @@ class FleetMcpAuthzTest {
|
||||
|
||||
@Test
|
||||
void everyRegisteredToolHasItsHandlerActionPinned() {
|
||||
Set<String> registered = toolsTheServerRegisters();
|
||||
// fleetd #469: this used to scrape FleetMcp.java's tool("…") calls for the registered set —
|
||||
// a third copy of the same list this file, CharterToolSurfaceTest and McpContractDocTest
|
||||
// each kept independently. All three now read FleetTool.wireNames(), the canonical set
|
||||
// FleetMcp itself derives its tool schemas AND its authorization switch from; adding a tool
|
||||
// there without pinning its action in FleetMcp#authzAction is a compile error, so this test's
|
||||
// per-tool assertions below are a run-time regression pin on top of that compile-time check,
|
||||
// not the only thing standing between a new tool and an unpinned action.
|
||||
Set<String> registered = FleetTool.wireNames();
|
||||
assertTrue(registered.size() >= 10,
|
||||
"scraped only " + registered.size() + " tool registrations from FleetMcp (" + registered
|
||||
+ "); the server registers eleven, so the tool(\"…\") scrape has stopped matching");
|
||||
"FleetTool.wireNames() returned only " + registered.size() + " tool(s) (" + registered
|
||||
+ "); the server registers eleven, so FleetTool has stopped listing the real "
|
||||
+ "tool surface");
|
||||
registered.forEach(tool -> assertDoesNotThrow(() -> FleetMcp.toolAction(tool, Map.of()),
|
||||
() -> tool + " is registered but has no pinned authorization action"));
|
||||
|
||||
@@ -319,19 +325,6 @@ class FleetMcpAuthzTest {
|
||||
FleetMcp.toolAction("fleet_poll", Map.of("coordId", "mac-opus")));
|
||||
}
|
||||
|
||||
private static Set<String> toolsTheServerRegisters() {
|
||||
try {
|
||||
Matcher matcher = TOOL_REGISTRATION.matcher(Files.readString(MCP_SOURCE));
|
||||
Set<String> tools = new LinkedHashSet<>();
|
||||
while (matcher.find()) {
|
||||
tools.add(matcher.group(1));
|
||||
}
|
||||
return tools;
|
||||
} catch (Exception e) {
|
||||
throw new AssertionError("could not scrape FleetMcp tool registrations", e);
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void aWorkerMayNotDrainAnotherSessionsInboxByPolling() {
|
||||
FleetMcp m = mcp(true);
|
||||
|
||||
@@ -574,6 +574,38 @@ class FleetMcpTest {
|
||||
assertTrue(out.contains("\"quarantinedForSeconds\":1800"), out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #466 scope item 2: {@code fleet_profiles} must carry the repeat count beside the
|
||||
* remaining seconds, and the two must come off the one {@link BackendQuarantine#status} call so
|
||||
* they can never disagree about which streak this is (see {@code profilesView}'s javadoc).
|
||||
*/
|
||||
@Test
|
||||
void profilesReportsQuarantineAttemptBesideRemainingSeconds() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
java.util.concurrent.atomic.AtomicLong now = new java.util.concurrent.atomic.AtomicLong(0L);
|
||||
BackendQuarantine quarantine = BackendQuarantine.withEscalation(now::get, TimeUnit.MINUTES.toNanos(30));
|
||||
quarantine.quarantine("shared-openai"); // attempt 1: 1800s
|
||||
now.set(TimeUnit.MINUTES.toNanos(30));
|
||||
quarantine.quarantine("shared-openai"); // attempt 2: 3600s
|
||||
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
|
||||
profile -> "ltms-local".equals(profile) ? "shared-openai" : null, quarantine);
|
||||
McpSchema.CallToolResult res = FleetMcp.profiles(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), source);
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"quarantinedForSeconds\":3600"), out);
|
||||
assertTrue(out.contains("\"quarantineAttempt\":2"), out);
|
||||
}
|
||||
|
||||
/** A never-quarantined profile must not carry {@code quarantineAttempt} either. */
|
||||
@Test
|
||||
void profilesOmitsQuarantineAttemptWhenNotQuarantined() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
McpSchema.CallToolResult res = FleetMcp.profiles(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), FleetMcp.QuarantineSource.none());
|
||||
String out = textOf(res);
|
||||
assertFalse(out.contains("quarantineAttempt"), out);
|
||||
}
|
||||
|
||||
/** fleetd #201 Unit 5: {@code coolingOff} is a SEPARATE map from {@code quarantined}. */
|
||||
@Test
|
||||
void profilesReportsACoolingOffCredentialInASeparateMap() {
|
||||
@@ -1121,6 +1153,35 @@ class FleetMcpTest {
|
||||
assertTrue(out.contains("\"quarantinedForSeconds\":1200"), out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #466 scope item 2: {@code fleet_list}'s capacity rows (CB-583: they reuse quarantine)
|
||||
* must carry {@code quarantineAttempt} beside {@code quarantinedForSeconds}, off the same
|
||||
* {@link BackendQuarantine#status} call as {@code fleet_profiles} -- see {@code capacityView}'s
|
||||
* comment pointing back to {@code profilesView}.
|
||||
*/
|
||||
@Test
|
||||
void capacityRowReportsQuarantineAttemptBesideRemainingSeconds() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
java.util.concurrent.atomic.AtomicLong now = new java.util.concurrent.atomic.AtomicLong(0L);
|
||||
BackendQuarantine quarantine = BackendQuarantine.withEscalation(now::get, TimeUnit.MINUTES.toNanos(20));
|
||||
quarantine.quarantine("shared-openai"); // attempt 1
|
||||
now.set(TimeUnit.MINUTES.toNanos(20));
|
||||
quarantine.quarantine("shared-openai"); // attempt 2
|
||||
now.set(TimeUnit.MINUTES.toNanos(60));
|
||||
quarantine.quarantine("shared-openai"); // attempt 3
|
||||
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
|
||||
profile -> "terra".equals(profile) ? "shared-openai" : null, quarantine);
|
||||
|
||||
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 2,
|
||||
() -> Set.of("terra"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
source, Map.of(), ""));
|
||||
|
||||
assertTrue(out.contains("\"credentialId\":\"shared-openai\""), out);
|
||||
assertTrue(out.contains("\"quarantineAttempt\":3"), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void everyProfileSharingTheQuarantinedCredentialReportsZeroFree() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
|
||||
@@ -27,16 +27,21 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
* somewhere to be readable, and that is exactly the sentence that rots. This test is what makes it
|
||||
* safe to write.
|
||||
*
|
||||
* <p><b>It checks source text, not behaviour.</b> It reads the Markdown and reads {@link FleetMcp}'s
|
||||
* source, and it only catches a name in the doc that the server does not register. It cannot catch a
|
||||
* flow that describes the wrong order, or a parameter name in prose — those are not name-shaped. The
|
||||
* doc's own header carries that caveat for its readers.
|
||||
* <p><b>It checks source text, not behaviour.</b> It reads the Markdown, and it only catches a name
|
||||
* in the doc that the server does not register. It cannot catch a flow that describes the wrong
|
||||
* order, or a parameter name in prose — those are not name-shaped. The doc's own header carries
|
||||
* that caveat for its readers.
|
||||
*
|
||||
* <p>fleetd #469: the registered side used to be its own scrape of {@code FleetMcp.java}'s {@code
|
||||
* tool("…")} calls — a third copy of the same list {@code CharterToolSurfaceTest} and {@code
|
||||
* FleetMcpAuthzTest} each kept their own copy of too. All three now read {@link
|
||||
* FleetTool#wireNames()}, the one canonical set {@code FleetMcp} itself derives its tool schemas and
|
||||
* authorization switch from.
|
||||
*/
|
||||
class McpContractDocTest {
|
||||
|
||||
/** Tests run with the module directory as cwd, so the repo-root doc is one level up. */
|
||||
private static final Path DOC = Path.of("../docs/MCP-Contract.md");
|
||||
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
|
||||
|
||||
private static Set<String> matches(Path file, String regex) throws Exception {
|
||||
Matcher m = Pattern.compile(regex).matcher(Files.readString(file));
|
||||
@@ -52,15 +57,10 @@ class McpContractDocTest {
|
||||
return matches(DOC, "(fleet_[a-z_]+)");
|
||||
}
|
||||
|
||||
/** Every tool {@link FleetMcp} actually registers, read from its {@code tool("…")} calls. */
|
||||
private static Set<String> toolsTheServerRegisters() throws Exception {
|
||||
return matches(MCP_SOURCE, "tool\\(\"(fleet_[a-z_]+)\"");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] every fleet_* tool named in MCP-Contract.md is one the server registers")
|
||||
void theDocNamesNoToolThatDoesNotExist() throws Exception {
|
||||
Set<String> registered = toolsTheServerRegisters();
|
||||
Set<String> registered = FleetTool.wireNames();
|
||||
Set<String> named = toolsNamedInTheDoc();
|
||||
|
||||
Set<String> unknown = new LinkedHashSet<>(named);
|
||||
@@ -84,13 +84,13 @@ class McpContractDocTest {
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] the doc/server name check is not vacuous — both sides found names")
|
||||
void theCheckActuallyHasSomethingToCheck() throws Exception {
|
||||
Set<String> registered = toolsTheServerRegisters();
|
||||
Set<String> registered = FleetTool.wireNames();
|
||||
Set<String> named = toolsNamedInTheDoc();
|
||||
|
||||
assertTrue(registered.size() >= 10,
|
||||
"scraped only " + registered.size() + " tool registrations from FleetMcp (" + registered
|
||||
+ "); the server registers eleven, so the tool(\"…\") scrape has stopped matching "
|
||||
+ "and the check above is now vacuous");
|
||||
"FleetTool.wireNames() returned only " + registered.size() + " tool(s) (" + registered
|
||||
+ "); the server registers eleven, so FleetTool has stopped listing the real "
|
||||
+ "tool surface and the check above is now vacuous");
|
||||
assertTrue(named.size() >= 4,
|
||||
"docs/MCP-Contract.md names only " + named.size() + " fleet_* tool(s) (" + named + "). "
|
||||
+ "The flows describe delegation, clarification, detached delivery and the "
|
||||
|
||||
@@ -116,4 +116,262 @@ class BackendQuarantineTest {
|
||||
assertThrows(IllegalArgumentException.class, () -> new BackendQuarantine(() -> 0L, 0L));
|
||||
assertThrows(IllegalArgumentException.class, () -> new BackendQuarantine(() -> 0L, -1L));
|
||||
}
|
||||
|
||||
// --- fleetd #466: escalating cooldown -----------------------------------------------------
|
||||
//
|
||||
// Base cooldown 600s (10 min), multiplier 2.0, ceiling 2400s (4x base) — small round numbers
|
||||
// chosen so every deadline is an exact assertion, not just "greater than before". Each call
|
||||
// below lands at or before the previous deadline (a zero or negative gap), which is always
|
||||
// "no more than one base cooldown after the previous deadline" — i.e. every call continues the
|
||||
// same streak, matching a credential that keeps reporting exhausted with no lull.
|
||||
|
||||
@Test
|
||||
void anInvalidBackoffMultiplierIsRejected() {
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 0.5, TimeUnit.HOURS.toNanos(6)));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aCeilingBelowTheBaseCooldownIsRejected() {
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 2.0, TimeUnit.MINUTES.toNanos(10)));
|
||||
}
|
||||
|
||||
@Test
|
||||
void repeatedExhaustionEscalatesTheCooldownByExactAmounts() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
q.quarantine("shared-openai"); // 1st: base cooldown
|
||||
assertEquals(OptionalLong.of(600L), q.remainingSeconds("shared-openai"));
|
||||
|
||||
now.set(TimeUnit.SECONDS.toNanos(100)); // still inside the 1st quarantine (deadline 600s)
|
||||
q.quarantine("shared-openai"); // 2nd: 600 * 2^1 = 1200
|
||||
assertEquals(OptionalLong.of(1200L), q.remainingSeconds("shared-openai"),
|
||||
"a second consecutive exhaustion must double the cooldown, not just increase it");
|
||||
|
||||
now.set(TimeUnit.SECONDS.toNanos(1300)); // exactly the 2nd deadline (100 + 1200)
|
||||
q.quarantine("shared-openai"); // 3rd: 600 * 2^2 = 2400 (exactly at the ceiling)
|
||||
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void escalationStopsAtTheCeiling() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
q.quarantine("shared-openai"); // 1st: 600
|
||||
now.set(TimeUnit.SECONDS.toNanos(600));
|
||||
q.quarantine("shared-openai"); // 2nd: 1200, deadline 1800
|
||||
now.set(TimeUnit.SECONDS.toNanos(1800));
|
||||
q.quarantine("shared-openai"); // 3rd: 600 * 4 = 2400, at the ceiling, deadline 4200
|
||||
now.set(TimeUnit.SECONDS.toNanos(4200));
|
||||
q.quarantine("shared-openai"); // 4th: 600 * 8 = 4800 uncapped, must stay capped at 2400
|
||||
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"),
|
||||
"the cooldown must never exceed the configured ceiling, however long the streak gets");
|
||||
|
||||
now.set(TimeUnit.SECONDS.toNanos(6600)); // 4th deadline
|
||||
q.quarantine("shared-openai"); // 5th: still capped
|
||||
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"),
|
||||
"pushing well past the ceiling must not budge it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aQuietGapLongerThanTheBaseCooldownResetsToTheBaseCooldown() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
q.quarantine("shared-openai"); // 1st: 600, deadline 600
|
||||
now.set(TimeUnit.SECONDS.toNanos(600));
|
||||
q.quarantine("shared-openai"); // 2nd: 1200, deadline 1800
|
||||
now.set(TimeUnit.SECONDS.toNanos(1800));
|
||||
q.quarantine("shared-openai"); // 3rd: 2400, deadline 4200
|
||||
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"));
|
||||
|
||||
// Quiet for well over one base cooldown (600s) past the 3rd deadline (4200s).
|
||||
now.set(TimeUnit.SECONDS.toNanos(20_000));
|
||||
q.quarantine("shared-openai"); // treated as a fresh occurrence
|
||||
assertEquals(OptionalLong.of(600L), q.remainingSeconds("shared-openai"),
|
||||
"a long quiet gap must reset the streak back to the base cooldown");
|
||||
}
|
||||
|
||||
@Test
|
||||
void escalatingOneCredentialDoesNotSlowAnother() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
q.quarantine("shared-openai"); // 1st: 600
|
||||
now.set(TimeUnit.SECONDS.toNanos(600));
|
||||
q.quarantine("shared-openai"); // 2nd: 1200
|
||||
now.set(TimeUnit.SECONDS.toNanos(1800));
|
||||
q.quarantine("shared-openai"); // 3rd: 2400 — three-in-a-row streak on this credential only
|
||||
|
||||
q.quarantine("another-credential"); // its first and only exhaustion
|
||||
assertEquals(OptionalLong.of(600L), q.remainingSeconds("another-credential"),
|
||||
"an unrelated credential's cooldown must stay at the base rate, unaffected by a sibling's streak");
|
||||
}
|
||||
|
||||
@Test
|
||||
void withEscalationDefaultsToDoublingCappedAtTwelveTimesTheBase() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = BackendQuarantine.withEscalation(now::get, TimeUnit.MINUTES.toNanos(30));
|
||||
|
||||
q.quarantine("shared-openai");
|
||||
assertEquals(OptionalLong.of(1800L), q.remainingSeconds("shared-openai"),
|
||||
"the first occurrence must still use the base cooldown");
|
||||
|
||||
now.set(TimeUnit.MINUTES.toNanos(30));
|
||||
q.quarantine("shared-openai");
|
||||
assertEquals(OptionalLong.of(3600L), q.remainingSeconds("shared-openai"),
|
||||
"the default multiplier must be 2.0");
|
||||
}
|
||||
|
||||
// --- fleetd #466 scope item 2: status() reports repeatCount beside remainingSeconds ------------
|
||||
|
||||
@Test
|
||||
void statusReportsAttemptOneForAFirstOccurrenceNeverAbsentOrZero() {
|
||||
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 2.0,
|
||||
TimeUnit.HOURS.toNanos(6));
|
||||
q.quarantine("shared-openai");
|
||||
|
||||
BackendQuarantine.Status status = q.status("shared-openai").orElseThrow();
|
||||
assertEquals(1800L, status.remainingSeconds());
|
||||
assertEquals(1, status.repeatCount(),
|
||||
"a first-ever occurrence must report attempt 1, not 0 or absent -- 1 means unambiguously "
|
||||
+ "'the first time', where 0 would be indistinguishable from a bug that forgot to count");
|
||||
}
|
||||
|
||||
@Test
|
||||
void statusIsAbsentWhenTheCredentialIsNotQuarantined() {
|
||||
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 2.0,
|
||||
TimeUnit.HOURS.toNanos(6));
|
||||
assertTrue(q.status("shared-openai").isEmpty());
|
||||
}
|
||||
|
||||
/**
|
||||
* The acceptance criterion's strong form: the reported {@code repeatCount} must match the exact
|
||||
* step the cooldown's own growth implies, read off {@link BackendQuarantine#status}'s single
|
||||
* call -- not two independent reads that happen to agree in this easy case.
|
||||
*/
|
||||
@Test
|
||||
void statusReportsTheGrowingAttemptCountAlongsideTheEscalatingCooldown() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
q.quarantine("shared-openai"); // 1st: 600s, attempt 1
|
||||
BackendQuarantine.Status first = q.status("shared-openai").orElseThrow();
|
||||
assertEquals(600L, first.remainingSeconds());
|
||||
assertEquals(1, first.repeatCount());
|
||||
|
||||
now.set(TimeUnit.SECONDS.toNanos(600));
|
||||
q.quarantine("shared-openai"); // 2nd: 1200s, attempt 2
|
||||
BackendQuarantine.Status second = q.status("shared-openai").orElseThrow();
|
||||
assertEquals(1200L, second.remainingSeconds());
|
||||
assertEquals(2, second.repeatCount());
|
||||
|
||||
now.set(TimeUnit.SECONDS.toNanos(1800));
|
||||
q.quarantine("shared-openai"); // 3rd: 2400s (at the ceiling), attempt 3
|
||||
BackendQuarantine.Status third = q.status("shared-openai").orElseThrow();
|
||||
assertEquals(2400L, third.remainingSeconds());
|
||||
assertEquals(3, third.repeatCount());
|
||||
}
|
||||
|
||||
/**
|
||||
* Once the cooldown hits its ceiling, every further consecutive exhaustion reports the SAME
|
||||
* {@code remainingSeconds} -- so a {@code repeatCount} re-derived from the cooldown value (e.g.
|
||||
* inverting {@code cooldownNanos * multiplier^(n-1)}) could not tell attempt 4 from attempt 9;
|
||||
* only the stored counter can. This is the scenario that makes "read the count off a second,
|
||||
* independent computation" provably wrong rather than just risky.
|
||||
*/
|
||||
@Test
|
||||
void repeatCountKeepsGrowingPastTheCeilingEvenThoughTheCooldownStaysFlat() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
long t = 0L;
|
||||
for (int attempt = 1; attempt <= 5; attempt++) {
|
||||
now.set(t);
|
||||
q.quarantine("shared-openai");
|
||||
BackendQuarantine.Status status = q.status("shared-openai").orElseThrow();
|
||||
assertEquals(attempt, status.repeatCount(),
|
||||
"attempt " + attempt + " must be reported as exactly " + attempt
|
||||
+ ", not collapsed to whatever attempt first reached the ceiling");
|
||||
if (attempt >= 3) {
|
||||
assertEquals(2400L, status.remainingSeconds(), "attempt " + attempt + " must be capped");
|
||||
}
|
||||
t += status.remainingSeconds(); // land exactly on the next deadline: still the same streak
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void aQuietGapResetsTheReportedAttemptCountToOneToo() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
q.quarantine("shared-openai");
|
||||
now.set(TimeUnit.SECONDS.toNanos(600));
|
||||
q.quarantine("shared-openai");
|
||||
now.set(TimeUnit.SECONDS.toNanos(1800));
|
||||
q.quarantine("shared-openai"); // attempt 3
|
||||
assertEquals(3, q.status("shared-openai").orElseThrow().repeatCount());
|
||||
|
||||
now.set(TimeUnit.SECONDS.toNanos(20_000)); // long quiet gap
|
||||
q.quarantine("shared-openai");
|
||||
assertEquals(1, q.status("shared-openai").orElseThrow().repeatCount(),
|
||||
"a reset streak must report attempt 1 again, matching the reset base cooldown");
|
||||
}
|
||||
|
||||
@Test
|
||||
void escalatingOneCredentialsAttemptCountDoesNotAffectAnother() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
|
||||
TimeUnit.SECONDS.toNanos(2400));
|
||||
|
||||
q.quarantine("shared-openai");
|
||||
now.set(TimeUnit.SECONDS.toNanos(600));
|
||||
q.quarantine("shared-openai");
|
||||
now.set(TimeUnit.SECONDS.toNanos(1800));
|
||||
q.quarantine("shared-openai"); // 3-in-a-row streak on this credential only
|
||||
|
||||
q.quarantine("another-credential");
|
||||
assertEquals(1, q.status("another-credential").orElseThrow().repeatCount(),
|
||||
"an unrelated credential's attempt count must stay at 1, unaffected by a sibling's streak");
|
||||
}
|
||||
|
||||
@Test
|
||||
void noneReportsNoStatusForAnything() {
|
||||
BackendQuarantine q = BackendQuarantine.none();
|
||||
q.quarantine("shared-openai"); // no-op on none(), same as every other mutator
|
||||
|
||||
assertTrue(q.status("shared-openai").isEmpty(),
|
||||
"none() quarantines nothing, so it must report no status at all -- never a fabricated "
|
||||
+ "attempt count for a credential that was never actually quarantined");
|
||||
}
|
||||
|
||||
/** The flat (non-escalating) two-argument constructor must still report a real, growing count. */
|
||||
@Test
|
||||
void aFlatTwoArgumentInstanceStillReportsAGrowingAttemptCountEvenThoughTheCooldownStaysFlat() {
|
||||
AtomicLong now = new AtomicLong(0L);
|
||||
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600));
|
||||
|
||||
q.quarantine("shared-openai");
|
||||
BackendQuarantine.Status first = q.status("shared-openai").orElseThrow();
|
||||
assertEquals(600L, first.remainingSeconds());
|
||||
assertEquals(1, first.repeatCount());
|
||||
|
||||
now.set(TimeUnit.SECONDS.toNanos(100)); // still inside the 1st quarantine
|
||||
q.quarantine("shared-openai");
|
||||
BackendQuarantine.Status second = q.status("shared-openai").orElseThrow();
|
||||
assertEquals(600L, second.remainingSeconds(),
|
||||
"the flat constructor's cooldown must stay exactly the base length regardless of the streak");
|
||||
assertEquals(2, second.repeatCount(),
|
||||
"the flat constructor still counts the real streak -- it just does not scale the cooldown by it");
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user