Compare commits

..

15 Commits

Author SHA1 Message Date
Dai Ha 325d0771a4 fleetd #480 Unit B: add handover skill for lead session handoff
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Successful in 1m54s
Adds .claude/skills/handover/SKILL.md, the procedure an outgoing lead
follows to write the handover file a fresh lead session inherits when
fleetd clears the pane. Registers the new skill in CLAUDE.md's
primary-side skills list; no other change to CLAUDE.md.
2026-09-11 06:13:49 +07:00
Dai Ha 7b97aae85b docs: move the redeploy procedure out of CLAUDE.md into a skill
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m34s
CLAUDE.md loads into every session. The daemon-redeploy procedure is
needed only when someone redeploys, so it paid for context it did not
use: 58 lines, about 918 est. tokens, every session.

The procedure now lives in .claude/skills/redeploy-fleetd/SKILL.md,
which loads only when invoked. The moved text is byte-identical to
what was removed, plus one new paragraph documenting --no-build
(scripts/redeploy-fleetd.sh:34,62,266) — the script and the wiki
already had that flag, CLAUDE.md never did.

CLAUDE.md keeps an 11-line pointer, because two rules must stay
resident: a merge is not a deployment, and workers must never
redeploy. A session learns it needs the procedure before it needs
the skill, then the next line names the skill.

Also updates the addendum's primary-side skill list, as this file's
own rule for .claude/skills/** changes requires.

CLAUDE.md: 34,442 -> 31,591 chars.
Canonical block untouched — the wiki sync check still prints True.
2026-09-11 05:48:06 +07:00
Dai Ha 49a404ddf3 Merge #474 follow-up: pin main's ConfigRef wiring against the surviving mutation
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m37s
My battery on the #474 merge found one survivor: reverting Fleetd.java:154
from the three-argument ConfigRef constructor to the plain two-argument
one turns the live reload gate off and leaves all 1633 tests green. Both
new #474 tests build their own ConfigRef with the method reference, so
neither reads what main chose.

This adds FleetdConfigRefWiringTest, following the three source-text
precedents already in the tree (FleetdBackendQuarantineWiringTest,
FleetdLeadSeatWiringTest, FleetdCompletionResolverWiringTest) rather
than the weaker sibling pattern that builds the wiring itself. No
production change.

The worker branched fresh off 4466ee0 rather than continuing its old
branch off the stranded 435e022 base. That was its own call and it was
the right one: one merge base, a clean 80-line diff, and no cherry-pick
needed this time.

Its vacuity guard is worth keeping in mind for the next source-text
test: it asserts the file it read contains 'public final class Fleetd'
before asserting anything about the mutation, with a message saying the
assertFalse below would pass vacuously on a broken read. An unrelated
anchor is the right choice there, because a guard that shares the
mutation's text cannot tell a bad read from a real change.
2026-09-10 20:50:13 +07:00
Dai Ha 72d6a6878b fleetd #474 follow-up: pin main's config wiring against M2
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m32s
Fleetd.main's own choice of the three-argument ConfigRef constructor
(with Fleetd::assertChartersNameOnlyRegisteredTools as extraValidation)
was unpinned. Reverting Fleetd.java:154 to the plain two-argument
constructor compiled with 0 errors and left the whole suite green,
because ConfigRefTest and FleetdConfigRefCharterToolSurfaceWiringTest
each build their own ConfigRef directly rather than through main.

Adds FleetdConfigRefWiringTest, a source-text check on Fleetd.java
following the FleetdBackendQuarantineWiringTest/FleetdLeadSeatWiringTest/
FleetdCompletionResolverWiringTest precedent: asserts the exact
three-argument construction is present, asserts the plain two-argument
form is absent, and guards against a vacuous pass on a broken/empty
source read by first asserting an unrelated anchor is present.
2026-09-10 20:48:05 +07:00
Dai Ha 4466ee0ef2 fleetd #474: ConfigRef.reload() runs the charter tool-surface gate too
CI / build (push) Successful in 1m37s
CI / contract (push) Successful in 1m44s
A charter naming an MCP tool the server does not register refused Fleetd.main
at startup but slipped through ConfigRef.reload(), because reload() only ran
FleetConfig.validateAll(), which never looks at what a charter's text names.

CharterToolSurface stays in the mcp package (config must not depend on it), so
ConfigRef now accepts the check as a Consumer<FleetConfig> extraValidation,
run inside reload()'s same try/catch as validateAll(). Fleetd.main wires a new
package-private adapter, Fleetd.assertChartersNameOnlyRegisteredTools, into
both the startup call site and ConfigRef's constructor, so the two call sites
can never check different things.

Tests: ConfigRefTest (reload refuses/accepts, via a locally-built equivalent
consumer since Fleetd's method is package-private to dev.ltms.fleet) and the
new FleetdConfigRefCharterToolSurfaceWiringTest (same proof through the exact
Fleetd::assertChartersNameOnlyRegisteredTools reference production uses).
Verified deleting the new extraValidation.accept(fresh) call site fails both
new "refuses" tests by name.

(cherry picked from commit 97e4c1d658)

Lead review. Cherry-picked, not merged: the worker branched from 435e022,
which is not an ancestor of main (I had reset and rewritten that commit's
message while the worker was already on it). Merging its branch would have
added a second merge commit for work already on main. Same content, no
duplicated history. Rule learned: once a worker is spawned, its base
commit is published.

My own mutation battery, five cells, each a full `mvn -B clean test` on
this commit's own tree:

  CONTROL 1, unmutated       1633 tests, 0 failures, BUILD SUCCESS
  M1 delete accept(fresh)    KILLED  2 by name
  M3 3-arg ctor stores no-op KILLED  2 by name
  M4 set(fresh) before valid KILLED  6
  M5 accept outside try      KILLED  2 errors
  M2 main uses the 2-arg ctor SURVIVED

M2 is a real gap and it is a test gap, not a defect: the production code
here is correct. Reverting Fleetd.java:154 to `new ConfigRef(configPath,
cfg)` turns the live reload gate off and leaves all 1633 tests green,
because both new tests build their own ConfigRef with the method
reference rather than reading what main wires.

That is the fourth Fleetd.main call site to survive a battery (#446 M5
and M7, #466 M1, now this). Extracting a check into a well-tested helper
moves the untested surface UP, into the line that chooses to call it. The
repo already has the answer in three source-text wiring tests
(FleetdBackendQuarantineWiringTest, FleetdLeadSeatWiringTest,
FleetdCompletionResolverWiringTest); the worker followed the other
sibling pattern, which builds the wiring itself and so cannot pin main.
Delegated as a follow-up, with the surviving mutation as its acceptance
criterion.

Also mine to correct: my M3 cell's annotation said the no-op store "must
be 2" and measured 1. The 2-arg constructor delegates rather than
assigning, so my pattern only ever matched the mutant. The cell still
stands on its other proof (requireNonNull left = 0). Third battery in a
row with a wrong CONTROL annotation of my own.
2026-09-10 20:40:49 +07:00
Dai Ha 25ba7f16bb Merge #473: fleet_profiles and fleet_list report which attempt a quarantine is on (fleetd #466 item 2)
CI / contract (push) Successful in 1m20s
CI / build (push) Successful in 1m24s
The escalating cooldown landed in 789b6a8, but the only thing an operator
could see was quarantinedForSeconds. A long number does not say whether
this is the first exhaustion or the fifth, and after escalation those look
the same from outside: 3600 seconds could be a big base cooldown or a
credential on its fourth strike.

Both reports now carry quarantineAttempt beside quarantinedForSeconds.

The part that matters is HOW, not that the field exists.
BackendQuarantine#status(credentialId) does ONE quarantines.get() and
returns both values from the same QuarantineState. It is not two accessors
that a caller pairs up. That is the pattern
CompositePeerLauncher.modelGateState() already documents: a report that
reads a different source from the behaviour it describes will eventually
disagree with it, and the disagreement is invisible because both halves
look right on their own.

This is reporting only. Escalation, the ceiling and the quiet-gap reset are
untouched. The flat two-argument constructor still reports a real growing
attempt count even though its cooldown stays flat - which is the honest
answer, since the streak is real whether or not the cooldown uses it.

The worker's report was lost: its fleet_reply never arrived and its inbox
was empty. Nothing was lost with it, because the brief required pushing
first and putting the report in the PR body. That is the second time that
practice has saved a turn.

Numbers and my own mutation battery are on the ticket, measured on this
merge commit.
2026-09-10 20:21:14 +07:00
Dai Ha 5ba69c9cf2 fleetd #393: correct a comment that claimed two tests pin a call they cannot
CI / contract (push) Successful in 47s
CI / build (push) Failing after 1m51s
The comment above the charter writer said flipping that call back to
putArray was "proven load-bearing" and pointed at two named tests. I
mutated exactly that, on the merge commit, and it SURVIVED at 1618 green -
so neither named test covers it, and neither can.

The reason is structural, not a missing test: that writer runs first
against an empty array, so putArray has nothing to replace and the two
idioms are equivalent there. No test can distinguish them. The worker
measured the same thing independently and said so in their report; the
comment was left over from the pre-fix state, where the mutation being
described was a DIFFERENT one (a later writer destroying the charter).

A comment that names tests which do not cover the line is worse than no
comment. The next person mutates the line, sees green, and concludes the
tests are broken.

The corrected version states what each cell actually measures: this line
is unpinnable and why, the later writers ARE pinned and by which test
names, and deleting this line entirely fails the charter-reaches-the-member
assertion - the hole that predates #393.
2026-09-10 20:17:21 +07:00
Dai Ha 17052bb515 Merge #471: seeded skills reach an opencode member, and instructions[] stops depending on write order (fleetd #393)
memberSkills: copied skill folders into every provisioned worktree's
.claude/skills/ and stopped there. Claude Code reads that directory
natively; opencode never does. So the feature was INERT for opencode
members rather than broken: the copy succeeded, the files were correct,
and nothing ever read them. No test failed because there was nothing to
fail - the feature worked at the only layer it implemented.

An opencode member's only channel for static guidance is the instructions[]
array in its generated config. OpenCodeLauncher now adds each seeded
skill's SKILL.md there. A folder with no SKILL.md is never delivered and
the log names it.

The second half is an ordering hazard fleet01 found by reading, and that I
then measured. Three writers append to instructions[]: the role charter,
the seeded skills, and the IDE rules. The charter used putArray
(CREATE-OR-REPLACE) while the other two used withArray (get-or-create).
That was safe only because the charter ran first against an empty array -
an undeclared constraint that nothing tested. Measured on the earlier
merge: making the skills writer use putArray left 1603 tests green while
silently deleting the charter entry, so an opencode member would launch
with no role contract at all. Worse than the bug being fixed, and
invisible.

All three writers now use withArray, and three tests pin the array's
CONTENTS (never its size - a size assertion passes when putArray swaps two
entries for two others) across the combinations that matter:
charter-only, charter+ide, charter+skills+ide.

WHY NOTHING CAUGHT IT, MEASURED RATHER THAN ASSUMED. fleet01 first said
the missing axis was the COMBINATION of writers, then revised that to a
stronger claim: that nothing asserted the charter reaches an opencode
member at all. I checked the second claim on this merge and it is FALSE.
Deleting the charter writer outright fails 6 tests, and 3 of those existed
before #393: writesRemoteMcpConfigAndCharterInstructionsWhenMcpUrlSet,
roleCharterWithoutMcpOrCustomProviderStillWritesAConfig, and
aPinnedEndpointAndTheFleetMcpCoexistInOneConfig. The charter reaching a
member was already pinned.

So their FIRST diagnosis was the right one. The surviving mutation did not
delete the charter write; it made a LATER writer replace the whole array.
Every pre-existing test had exactly one writer active, and with one writer
putArray and withArray are indistinguishable. The gap was never "is the
charter delivered" - it was "are two writers ever active at once", which is
the combination axis. Recording this because the stronger claim is the more
quotable one and it would have sent the next reader looking for a hole that
is not there.

One honest residue: the charter writer's own idiom cannot be pinned.
Flipping it back to putArray leaves the suite green, and always will,
because it runs first against an empty array where the two idioms are
equivalent. The edit removes an undeclared constraint for the next person
to add a writer; it is not a change any test can detect. The comment in
the source claimed two named tests cover it - that was wrong, and the
commit after this one corrects it rather than leaving a false claim beside
the code.

What this does NOT do: opencode has no equivalent of Claude Code's Skill
tool, so the content is static system-prompt text present from spawn, not
something a member can invoke by name. This closes the DELIVERY gap and
cannot close the ACTIVATION gap. That residue is opencode's design.

Numbers and my own mutation battery are on the ticket, measured on this
merge commit.
2026-09-10 20:17:21 +07:00
Dai Ha e95ed99bf7 fleetd #466 scope item 2: report the quarantine repeat count, not only the seconds
CI / build (pull_request) Successful in 1m31s
CI / contract (pull_request) Successful in 1m33s
Add BackendQuarantine#status(credentialId) -> Optional<Status>, a single
QuarantineState read that answers both remainingSeconds and repeatCount
together -- the same "one accessor" pattern CompositePeerLauncher.
modelGateState() already uses, so the two facts can never disagree.

fleet_profiles/REST GET /profiles and fleet_list's capacity rows
(FleetMcp.profilesView/capacityView) now call status() instead of
remainingSeconds() and add a "quarantineAttempt" field beside
"quarantinedForSeconds": 1 for a first occurrence, 2 for the second
in a row, and so on.

No change to the escalation, ceiling, or reset logic itself -- this
unit is reporting only.
2026-09-10 20:09:13 +07:00
Dai Ha 1477e4358a Merge #472: one canonical tool-name set, and charters are checked against it (fleetd #469)
CI / contract (push) Successful in 57s
CI / build (push) Successful in 2m0s
A role charter is free text in config that tells a member which tools to
call, and nothing checked that those tools exist. A charter naming
bridge_send - a name CB-634 removed - started the daemon cleanly, and the
member found out at run time by calling something that was not there.

#464 shipped a test for this, but it wrote its own charter into a @TempDir,
so nothing anyone put in the real config could fail it. That was a defect
in my acceptance criteria, not in that work.

Now: FleetTool is one enum of the 11 registered wire names, and every
reader goes through it.

- FleetMcp's schema builders pass FleetTool.X.wireName() instead of a
  literal.
- FleetMcp's constructor asserts at startup that what it registers with the
  SDK equals FleetTool.wireNames() exactly, in both directions.
- toolAction(String, Map) resolves arbitrary wire input against
  FleetTool.byWireName() and keeps its run-time throw, which is necessary -
  network input has no closed compile-time form. It then hands off to
  authzAction(FleetTool, Map), a switch over the enum with NO default, so a
  new tool is a compile error at that layer.
- CharterToolSurface lives in mcp, not config, and Fleetd.main calls it
  right after validateAll(). Config must not depend on the MCP server:
  config loads before the server exists.

My brief undercounted the problem and the worker corrected it. I said there
were two tool-name inventories plus a test fixture. There were FIVE: the
registrations, the authz switch, and three separate source-text scrapes of
FleetMcp.java in CharterToolSurfaceTest, FleetMcpAuthzTest and
McpContractDocTest - none of which the ticket mentioned. Fixing those three
was required, not scope creep: once the literals moved into FleetTool their
regexes matched zero names, so one would have failed on its vacuity guard
and the other two would have gone quietly vacuous. Inventories after: one.

Verified independently on origin/main before accepting the wider diff:
three test files did read FleetMcp.java as source text, with a control file
at zero to prove the search discriminated.

Build number and my own mutation results are on the ticket and the PR,
measured on this merge commit rather than on the branch.
2026-09-10 19:59:03 +07:00
Dai Ha 6c2d6e93cb fleetd #469: one canonical FleetTool set backs registration, authz and charter checks
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m37s
FleetConfig.validateCharters() only checked that a charter key is a role
wire name and its text is non-blank; #464's CharterToolSurfaceTest compared
charter text against the registered tool surface, but wrote its own charter
into a @TempDir fixture, so nothing anyone wrote into the live fleetd.yaml
could ever fail it.

Add FleetTool, an enum in dev.ltms.fleet.mcp holding the one canonical set
of registered tool wire names. FleetMcp's tool schemas now derive their
names from it, its constructor asserts at startup that what it actually
registers with the SDK equals FleetTool.wireNames() exactly, and its authz
dispatch (toolAction/authzAction) resolves the wire string against FleetTool
before switching on the enum itself with no default -- adding a tool without
pinning its Authz.Action is now a compile error, not just a test gap.

Add CharterToolSurface (mcp package, not config -- config loads before the
MCP server exists) and call it from Fleetd.main right after
cfg.validateAll(), so a charter naming a tool the server does not register
refuses the daemon's startup, naming both the charter key and the unknown
tool. FleetdStartupValidationTest proves this through Fleetd.main itself
against a live-shaped config fixture (bridge_send, CB-634's own removed
name).

CharterToolSurfaceTest, FleetMcpAuthzTest and McpContractDocTest each kept
an independent regex scrape of FleetMcp.java's source for the registered
side of their own comparison -- three more copies of the same list nothing
tied together. All three now read FleetTool.wireNames() instead.

Proved canonical by removal: deleting FleetTool.ACK while ackTool() still
referenced it broke mvn compile in two places (FleetMcp.java:940,:1898);
registering a schema under a literal not backed by FleetTool
("fleet_ack_v2") failed FleetMcp's new startup assertion in every test that
constructs it (12 errors, IllegalStateException at FleetMcp.<init>). Both
reverted before this commit.
2026-09-10 19:55:02 +07:00
Dai Ha 789b6a8716 Merge #470: the exhaustion quarantine escalates instead of retrying flat (fleetd #466)
CI / contract (push) Successful in 1m56s
CI / build (push) Successful in 2m4s
A flat 30-minute cooldown suits a backend that is out of capacity for the
hour. It does not suit a weekly subscription limit: that keeps reporting
exhausted until the window resets, so the daemon retried it roughly 336
times across a week and learned nothing each time.

BackendQuarantine now tracks, per credential, how many consecutive
exhaustion reports it has seen with no quiet gap between them, and doubles
the cooldown each time, capped at 12x the base (about 6 hours at the 1800s
default). That is about a dozen attempts a week instead of ~336.

Three things on the record because they are judgement calls, not facts:

- The reset is a TIME PROXY, not a success signal. Nothing in this codebase
  reports a spawn success back to this class, so "it started working again"
  cannot be observed here. A base cooldown of quiet is the best available
  evidence, and the class doc says that plainly instead of implying the
  stronger thing.
- No automatic probing. That was the operator's design constraint and the
  implementation respects it: the daemon warns and waits, it never pokes a
  limited backend to see whether the limit lifted.
- The multiplier (2.0) and ceiling (12x) are constants, not config surface,
  so no new ConfigRef hot/cold/deferred question arises.
  quarantineCooldownSeconds stays Deferred and is now the BASE of the
  backoff; fleetd.example.yaml and FleetConfig's javadoc say so.

Escalation fires on the exhaustion signal alone. Cooling-off
(BackendOutagePolicy, a flat 60s after repeated non-exhaustion errors) is a
separate mechanism and deliberately NOT escalated: doing so would turn a
transient 5xx storm into a multi-hour outage.

The old two-argument constructor is behaviourally unchanged - the same
formula with multiplier 1.0 and a ceiling equal to the base, which collapses
to the original flat "now + cooldown". Every existing call site keeps its
shape.

The second commit exists because my own mutation on the first merge found
the wiring unpinned: putting Fleetd.main back on the flat constructor left
all 1608 tests green, so the factory was pinned and the decision to use it
was not. FleetdBackendQuarantineWiringTest closes that, following the five
existing *WiringTest files rather than inventing an idiom. It checks source
TEXT, and its class doc says so: it does not prove the call executes, and it
cannot tell "wrong factory" apart from "renamed the anchor" - both fail the
same assertion. That limit is real and recorded rather than papered over.

Build number and my own mutation results are on the ticket and the PR,
measured on this merge commit rather than on the branch.
2026-09-10 19:51:00 +07:00
Dai Ha 01462c9695 fleetd #466 follow-up: pin main's choice of the escalating quarantine factory
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m25s
BackendQuarantineTest proves BackendQuarantine.withEscalation itself
escalates, is ceilinged, and resets. Nothing proved Fleetd.main still
calls it: reverting main to the flat two-argument constructor compiled
with 0 errors and left the full 1608-test suite green, because every
other test builds its own BackendQuarantine directly.

Adds FleetdBackendQuarantineWiringTest, a source-text assertion on
Fleetd.java (same idiom as FleetdLeadSeatWiringTest and
FleetdCompletionResolverWiringTest) pinning that main's declaration is
built from BackendQuarantine.withEscalation(...), not `new
BackendQuarantine(...)`. Measured: reverting to the flat constructor
fails it (expected true, was false); renaming the anchored local
variable while keeping behaviour identical also fails it loudly, not
silently. The class doc states plainly this checks source text only,
not that the call executes or that escalation reaches a real backend.
2026-09-10 19:47:59 +07:00
Dai Ha 8b4ff78546 Merge #470: the exhaustion quarantine escalates instead of retrying flat (fleetd #466)
A flat 30-minute cooldown suits a backend that is out of capacity for the
hour. It does not suit a weekly subscription limit: that keeps reporting
exhausted until the window resets, so the daemon retried it roughly 336
times across a week and learned nothing each time.

BackendQuarantine now tracks, per credential, how many consecutive
exhaustion reports it has seen with no quiet gap between them, and doubles
the cooldown each time, capped at 12x the base (about 6 hours at the 1800s
default). That is about a dozen attempts a week instead of ~336.

Three things I want on the record because they are judgement calls, not
facts:

- The reset is a TIME PROXY, not a success signal. Nothing in this
  codebase reports a spawn success back to this class, so "it started
  working again" cannot be observed here. A base cooldown of quiet is the
  best available evidence. The class doc says this plainly rather than
  implying the stronger thing.
- No automatic probing. That was the operator's design constraint and the
  implementation respects it: the daemon warns and waits, it never pokes a
  limited backend to see whether the limit lifted.
- The multiplier (2.0) and ceiling (12x) are constants, not config surface,
  so no new ConfigRef hot/cold/deferred classification question arises.
  quarantineCooldownSeconds stays Deferred and is now the BASE of the
  backoff; fleetd.example.yaml and FleetConfig's javadoc say so.

Escalation fires on the exhaustion signal alone. Cooling-off
(BackendOutagePolicy, a flat 60s after repeated non-exhaustion errors) is a
separate mechanism and is deliberately NOT escalated: doing so would turn a
transient 5xx storm into a multi-hour outage.

The old two-argument constructor is behaviourally unchanged - it is the
same formula with multiplier 1.0 and a ceiling equal to the base, which
collapses to the original flat "now + cooldown". Every existing call site
keeps its shape.

Build number and my own mutation results are reported in the PR and on the
ticket, measured on this merge commit rather than on the branch.
2026-09-10 19:40:15 +07:00
Dai Ha 5a467e1f8b fleetd #466: escalate BackendQuarantine's cooldown on repeated exhaustion
CI / build (pull_request) Successful in 1m34s
CI / contract (pull_request) Successful in 1m36s
A flat 30-minute quarantine retries a weekly subscription limit about
336 times before the window resets. BackendQuarantine now doubles the
cooldown on each consecutive exhaustion of the same credential (no
more than one base cooldown after the previous quarantine's deadline),
capped at 12x the base cooldown (~6h at the 1800s default), and resets
back to the base cooldown once a base-cooldown's worth of quiet has
passed with no further exhaustion.

The flat two-argument constructor is unchanged (equivalent to
multiplier 1.0 / ceiling == base), so all ~20 existing call sites keep
their current shape and behaviour. Production wiring (Fleetd.main)
switches to the new BackendQuarantine.withEscalation factory.

This only touches the exhaustion path (BackendQuarantine's one
production caller is Fleetd.exhaustionSink, fired on BACKEND_EXHAUSTED
alone) and never the separate, unescalated cooling-off mechanism
(BackendOutagePolicy, fixed 60s) that guards against a transient
backend-error storm.
2026-09-10 19:33:09 +07:00
22 changed files with 1570 additions and 159 deletions
+163
View File
@@ -0,0 +1,163 @@
---
name: handover
description: Procedure for an outgoing lead to write the handover file before a fresh lead session replaces it (fleetd #480). Load this when your context is full and fleetd is about to clear your pane. The file is the new lead's only inheritance — follow it exactly.
---
# Handover — write the file the next lead depends on
fleetd ticket #480 lets a lead session hand off to a fresh one. The outgoing lead writes a
handover file, fleetd checks it, clears the pane, and tells the new session to read that file
and carry on.
**The new lead's only inheritance is that file.** It does not see your conversation, your plan,
or your screen. If the file is thin or wrong, the new lead re-derives what you already knew, and
that wastes hours. Writing a good handover file is real work. It is not paperwork you rush
through at the end of a session.
This skill is the procedure for writing it. Every rule below earned its place because a past
handover got it wrong.
## 1. Confirm you are the right session to write this
Run `fleet_whoami` first. It must answer `primary`. Only a primary (lead) session writes a
handover file. A worker's job ends with its own pull request, not a fleet-wide handoff.
## 2. Every number needs a command, run in this turn
A number is a claim: a count, a commit hash, a process id, a percentage, a queue depth. Before
you write one, run the command that produces it — now, in this turn, against the live state.
Never take a number from:
- earlier in your own conversation — the state has moved since then,
- a peer lead's report — that is their measurement, not yours,
- your own memory of an earlier session.
Put the command, or its real output, next to the number. That lets the next lead re-run it and
check it still matches. If you cannot measure something yourself, say so instead of guessing:
"the fleet01 lead reports 91 commits behind; I have not checked this myself."
## 3. Say what you measured and what you did not
Mark every claim as one of two things:
- **"I checked this myself, in the code or on this host, at `<time>`."**
- **"I did not check this myself; `<who>` reported it."**
Never present someone else's measurement as your own. This matters most for cross-host claims —
a peer lead's daemon, a worker's report, or something the operator said earlier that you cannot
re-verify from here.
## 4. Record open decisions, and who owns them
List three things:
- what the operator actually asked for, in their own words where you have them,
- what is still unanswered,
- any question you decided yourself instead of asking, with your reason.
Write the decision so it cannot be mistaken for the operator's instruction. Say plainly: "the
operator never answered X; I decided Y, because Z." Without this, the next lead either silently
reopens a closed question or assumes the operator chose something they never did.
## 5. Record live hazards
List anything that will break if the next lead does the obvious thing next. This includes:
- unpushed commits or unmerged branches,
- a build, a spawn, or a redeploy still running,
- code merged to `main` but not yet redeployed to the live daemon,
- any trap that looks safe and is not — say what goes wrong and why, not only that something is
"tricky."
## 6. Record what is explicitly not owed
List work that is finished, and work that another party has said they do not want touched. Name
who said so and when. Without this line, the next lead re-does closed work or reopens a question
a peer already declined to revisit.
## 7. Open the file with three re-measurement commands
The file's own first section must give the next lead three concrete commands to run before
acting on anything else in the file:
1. confirm role — for example `fleet_whoami`,
2. confirm the state of the working tree — for example `git status` and
`git rev-list origin/main..HEAD`,
3. read the live fleet — for example `fleet_list`.
Record what each command answered when you wrote the file, and tell the reader to run it again
rather than trust your answer. The point of this section is that the reader checks live state
before acting on any claim in the rest of the file, including yours.
## 8. Stamp the file with time and commit
Near the top of the file, write:
- the date and time you wrote it,
- the commit the tree was on (`git rev-parse HEAD`),
- whether the tree was clean (`git status`).
Without this, nobody can tell how old the file is, or which code it describes.
## 9. State plainly that the file goes stale fast
Say near the top: **re-measure anything you act on.** The file goes stale the moment anyone
merges a branch, spawns a member, or restarts the daemon. Everything in the file is a snapshot
of one moment, not a live fact.
## 10. What to leave out
Do not include:
- narration of how the session felt, or how hard something was,
- anything the repo already records — code structure, git history, or a rule already written in
`CLAUDE.md`. Point at it instead of repeating it,
- advice that is only true for the session that is ending — a half-open terminal, a local
variable, a train of thought with no state behind it.
A handover file is a record of state and decisions. It is not a diary.
## Writing style
Write in plain English. Use everyday words, one idea per sentence, and active voice. Keep every
class, method, file, flag, and config key exactly as it appears in the code — replacing a
precise term with a vague one makes the sentence wrong, not simpler. Explain an abbreviation the
first time you use it.
If a diagram genuinely helps, put it in the `.md` file as a fenced ` ```mermaid ` block with no
hardcoded colors, so it stays readable on light and dark backgrounds. Quote any label that has
brackets, colons, or slashes.
## Template
```markdown
# Handover — <fleet name> lead session, <date and time>
Written at commit `<output of git rev-parse HEAD>`. Tree was <clean, or dirty: `<git status
summary>`>. Re-measure anything you act on — this file goes stale the moment anyone merges,
spawns, or restarts.
## 0. Do these three things first
1. Confirm your role: `fleet_whoami` — must answer `primary`. (Answered `<result>` at `<time>`.)
2. Confirm tree state: `git status`, `git rev-list origin/main..HEAD`. (`<result>` at `<time>`.)
3. Read the live fleet: `fleet_list`. (`<result>` at `<time>`.)
## 1. What the operator asked for
<the live instructions, in their words where you have them; what is still open; any decision
you made yourself, and why>
## 2. Open decisions, and who owns them
<one line per decision: who owns it, what is unanswered>
## 3. Live hazards
<one entry per hazard: what breaks, and why, if the next lead does the obvious thing>
## 4. What is not owed
<finished work, and work another party has declined; name who said so and when>
```
+72
View File
@@ -0,0 +1,72 @@
---
name: redeploy-fleetd
description: Rebuild and restart the live fleetd daemon after a merge (lead / primary only). Load this before redeploying — it holds the script, the drain step, the permission grant, and the five checks that have each gone wrong here before. Workers must never do this.
---
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-fleetd.sh --check # report state, change nothing
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
```
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
+12 -55
View File
@@ -257,8 +257,11 @@ must obey belongs in the charter, not here.
already cost three workers' turns: each wrote a good report to its terminal and ended the turn
with no `fleet_reply`, and the scrape returned the tail of the brief instead.
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
`port-to-opencode` (make an OpenCode session a participant in this workspace) and
`fleets-status` (report every fleet that shares one LavinMQ instance).
`port-to-opencode` (make an OpenCode session a participant in this workspace),
`fleets-status` (report every fleet that shares one LavinMQ instance),
`redeploy-fleetd` (rebuild and restart the live daemon after a merge) and
`handover` (write the file a fresh lead session inherits when the outgoing one hands off,
fleetd #480).
- **This repo is also a Claude Code marketplace, and ships a plugin.** `.claude-plugin/marketplace.json`
points at `plugin/`, which carries the MCP mount and the `setup` skill
(`/claude-bridge:setup` — make any project bridge-ready). It was added in CB-527 and then went
@@ -292,60 +295,14 @@ must obey belongs in the charter, not here.
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
feature merged to `main` does nothing until the daemon is rebuilt and restarted. The lead **may and
should** redeploy rather than hand the job back to the operator. **Workers must never do this** — a
worker has no business restarting the daemon it is talking through, and stopping it kills the
worker's own channel mid-turn.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-fleetd.sh --check # report state, change nothing
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
```
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
**Load the `redeploy-fleetd` skill before you redeploy.** It holds `scripts/redeploy-fleetd.sh`
and its flags, the drain step, the operator's permission grant, and the five checks that have each
gone wrong here before. Do not hand-roll the steps from memory.
### The prompt is part of the product — update it with the code (mandatory)
+9
View File
@@ -452,6 +452,15 @@ placement: weighted
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
# 1800 (30 minutes) when omitted or non-positive.
#
# fleetd #466: this is now only the BASE of an escalating backoff, not a flat retry rate. A
# credential quarantined again within one base cooldown of the previous quarantine ending (still
# reporting exhausted — e.g. a weekly subscription limit that hasn't reset) backs off further:
# cooldown doubles each such time, capped at 12x this value (~6 hours at the 1800s default). A
# quarantine that starts after a base-cooldown's worth of quiet resets back to this value. Not
# configurable per se — the multiplier and ceiling are constants in BackendQuarantine, not new
# YAML keys; see its class doc for the exact formula and why there is no automatic probe to clear
# it early (the operator's own design constraint — a probe spends the quota it's measuring).
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
# restart. Editing this needs a daemon restart to take effect.
@@ -26,6 +26,7 @@ import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.mcp.CharterToolSurface;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
@@ -147,7 +148,10 @@ public final class Fleetd {
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
// contract; adding a reader here does not make a key reloadable by itself.
ConfigRef config = new ConfigRef(configPath, cfg);
// fleetd #474: pass the charter/tool-surface check in as ConfigRef's extraValidation, so
// ConfigRef#reload() runs the same gate main() runs below, without dev.ltms.fleet.config
// gaining a dependency on dev.ltms.fleet.mcp — Fleetd is the seam that already holds both.
ConfigRef config = new ConfigRef(configPath, cfg, Fleetd::assertChartersNameOnlyRegisteredTools);
// The primary/host env that launched fleetd must not be tainted.
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
@@ -163,6 +167,17 @@ public final class Fleetd {
// and the Fleetd-startup tests actually pin — see FleetConfig#validateAll's javadoc for
// why a name-by-name list here would have the same defect it replaces.
cfg.validateAll();
// fleetd #469, follow-up to #464: validateAll() (and validateCharters() inside it) only
// checks that a charter's KEY is a role wire name and its text is non-blank — it never
// looks at what the text actually names. This is the separate check that does: it asks
// dev.ltms.fleet.mcp.FleetTool (the canonical registered-tool set) whether every fleet_*/
// bridge_* token a charter names is a tool this server actually registers. It cannot live
// inside FleetConfig#validateCharters() — config loads before the MCP server exists, and
// must not gain a dependency on the mcp package — so it runs here instead, at the one seam
// that already holds both a loaded FleetConfig and the mcp package, before anything below
// opens a socket or spawns a member. fleetd #474: the same check is also wired into `config`
// above as ConfigRef's extraValidation, so a reload refuses what this line refuses at startup.
assertChartersNameOnlyRegisteredTools(cfg);
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
@@ -222,7 +237,13 @@ public final class Fleetd {
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
// The cooldown is deferred (see FleetConfig#quarantineCooldownSeconds): it is read once
// here, at startup, and a config reload only changes it for a daemon restart.
BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,
// fleetd #466: escalating, not flat — a credential that keeps reporting exhaustion (e.g. a
// weekly subscription limit, which would otherwise be retried on every ~30-minute cooldown,
// about 336 times across the week) backs off further each consecutive time, capped at
// BackendQuarantine.DEFAULT_MAX_COOLDOWN_MULTIPLE x the base cooldown. See BackendQuarantine's
// class doc for the mechanism, why this never fires on cooling-off (a separate, unescalated
// mechanism — BackendOutagePolicy below), and the reset.
BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
// fleetd #201 Unit 5: one outage-cool-off tracker for the whole daemon, shared between the
// launcher (checked at spawn, like `quarantine` above) and the backend-error sink wired in
@@ -1576,6 +1597,28 @@ public final class Fleetd {
}
}
/**
* fleetd #474: the one place both the startup call (right after {@code cfg.validateAll()} in
* {@link #main}) and the reload call (wired into {@code config}'s {@code extraValidation} above,
* via a method reference to this method) go through, so the two can never drift into checking
* different things. Extracted only to give {@link ConfigRef}'s {@code Consumer<FleetConfig>}
* hook a {@code FleetConfig -> void} shape to bind to — {@link CharterToolSurface} itself still
* takes the raw charter map and knows nothing about {@code ConfigRef} or {@code Fleetd}.
*
* <p>Package-private so a test can call it directly the same way the other startup-report
* helpers above are tested, without needing to drive {@link #main} for a unit-level check;
* {@code FleetdStartupValidationTest} proves the startup call site, and {@code
* FleetdConfigRefCharterToolSurfaceWiringTest} — by constructing {@code ConfigRef} with this
* exact method reference, the same way {@code main} does above — proves the reload call site.
* {@code dev.ltms.fleet.config.ConfigRefTest} pins the same reload behaviour too, through an
* equivalent {@code Consumer<FleetConfig>} it builds locally (it cannot see this package-private
* method from {@code dev.ltms.fleet.config}).
*/
static void assertChartersNameOnlyRegisteredTools(FleetConfig cfg) {
CharterToolSurface.assertChartersNameOnlyRegisteredTools(
cfg.fleet() == null ? Map.of() : cfg.fleet().charters());
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
*
@@ -11,6 +11,7 @@ import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Consumer;
import java.util.function.Supplier;
/**
@@ -211,6 +212,24 @@ import java.util.function.Supplier;
* <p>A reload that fails to parse or fails validation is also refused, and the previous config keeps
* running. A config file being edited is normally read once mid-save; degrading a working daemon
* because it caught a half-written file would be a bad trade.
*
* <p><strong>fleetd #474</strong> — {@link FleetConfig#validateAll()} is not the only gate startup
* runs before a config takes effect: {@code Fleetd.main} also calls {@code
* dev.ltms.fleet.mcp.CharterToolSurface#assertChartersNameOnlyRegisteredTools}, right after {@code
* cfg.validateAll()}, to refuse a charter that names an MCP tool the server does not register. That
* check cannot live inside {@link FleetConfig} — {@code CharterToolSurface} lives in the {@code mcp}
* package because the canonical tool set ({@code FleetTool}) does, and config is loaded before the
* MCP server exists, so {@code FleetConfig} must not gain a dependency on {@code mcp}. {@link
* #reload} cannot import {@code mcp} either, for the same reason applied one layer up: {@code
* dev.ltms.fleet.config} is loaded before {@code dev.ltms.fleet.mcp} exists, same as {@code
* FleetConfig}. So this class accepts the check as a {@code Consumer<FleetConfig>} —
* {@link #extraValidation} — supplied by whichever caller already sits at the seam that holds both
* a loaded {@code FleetConfig} and the {@code mcp} package: {@code Fleetd.main}. It is invoked
* inside the same try/catch as {@code fresh.validateAll()}, so a charter that would have refused to
* boot refuses a reload too, and keeps the running config exactly like any other {@code
* validateAll()} failure. A ref built through the two-argument constructor (every test fixture that
* does not care about this check, and {@link #fixed}) gets a no-op consumer, so nothing outside
* {@code Fleetd.main} needs to know this hook exists.
*/
public final class ConfigRef implements Supplier<FleetConfig> {
@@ -255,10 +274,27 @@ public final class ConfigRef implements Supplier<FleetConfig> {
private final Path path;
private final AtomicReference<FleetConfig> current;
private final Consumer<FleetConfig> extraValidation;
/** Equivalent to the three-argument constructor with a no-op {@code extraValidation}. */
public ConfigRef(Path path, FleetConfig initial) {
this(path, initial, cfg -> { });
}
/**
* @param extraValidation run on every {@link #reload} candidate, inside the same try/catch as
* {@code fresh.validateAll()} — see the class doc's fleetd #474 note.
* {@code Fleetd.main} passes {@code
* Fleetd::assertChartersNameOnlyRegisteredTools} (a package-private
* {@code FleetConfig -> void} adapter over {@code
* CharterToolSurface#assertChartersNameOnlyRegisteredTools}), so a reload
* runs the same gate startup does without this class depending on the
* {@code mcp} package.
*/
public ConfigRef(Path path, FleetConfig initial, Consumer<FleetConfig> extraValidation) {
this.path = path;
this.current = new AtomicReference<>(Objects.requireNonNull(initial, "initial config"));
this.extraValidation = Objects.requireNonNull(extraValidation, "extraValidation");
}
/** A fixed reference that never reloads — for tests and for wiring built from a config in code. */
@@ -360,6 +396,12 @@ public final class ConfigRef implements Supplier<FleetConfig> {
// at all — see FleetConfig#validateAll's javadoc for why the fix is one reflective call,
// not a longer hand-maintained list.
fresh.validateAll();
// fleetd #474: validateAll() does not cover everything startup refuses on — the charter
// tool-surface check (Fleetd.main, right after cfg.validateAll()) lives outside
// FleetConfig on purpose (see this class's doc) and is supplied here as extraValidation.
// Same try/catch as validateAll() above, on purpose: either failure must refuse the whole
// reload and keep the running config the same way.
extraValidation.accept(fresh);
} catch (RuntimeException e) {
String msg = e.getMessage() == null ? e.toString() : e.getMessage();
log.warn("config reload from {} refused, keeping the running config: {}", path, msg);
@@ -70,12 +70,14 @@ import java.util.regex.PatternSyntaxException;
* {@code fixed} (default), {@code round-robin}, or {@code weighted}
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
* historical behaviour), CB-501
* @param quarantineCooldownSeconds how long a credential stays quarantined after a
* @param quarantineCooldownSeconds the BASE cooldown a credential is quarantined for after a
* {@code BACKEND_EXHAUSTED} classification (CB-578 stage B); {@code null}/{@code
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Baked once into the
* {@code BackendQuarantine} built at startup, so it is DEFERRED: changing it
* needs a restart, and a quarantine already running keeps whatever cooldown was
* live when it started.
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Since fleetd #466 this is
* only the first occurrence's length — a credential quarantined again shortly
* after this cooldown ends backs off further, up to a ceiling; see {@code
* BackendQuarantine}'s class doc. Baked once into the {@code BackendQuarantine}
* built at startup, so it is DEFERRED: changing it needs a restart, and a
* quarantine already running keeps whatever cooldown was live when it started.
* @param memberCredentials deny-by-default policy (CB-596) for which of the operator's own host
* credentials a spawned member's pane inherits. {@code null} (the block
* omitted) blocks nothing — see {@link MemberCredentials}.
@@ -0,0 +1,81 @@
package dev.ltms.fleet.mcp;
import java.util.ArrayList;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* fleetd #469: a launch charter that names an MCP tool the server does not register must stop the
* daemon at startup, not wait for a member to discover the gap by calling something that is not
* there.
*
* <p>Deliberately its own class outside {@code dev.ltms.fleet.config}, not a case in {@link
* dev.ltms.fleet.config.FleetConfig#validateCharters()}. The canonical tool surface ({@link
* FleetTool}) lives in the {@code mcp} package; config is loaded before the MCP server exists and
* must not gain a dependency on it. So this check belongs at the seam that already holds both a
* loaded {@code FleetConfig} and the {@code mcp} package: {@code Fleetd.main}, called right after
* {@code cfg.validateAll()} and before anything opens a socket or spawns a member.
*
* <p>{@code #464}'s {@code CharterToolSurfaceTest} proved the same comparison against a charter
* fixture it wrote itself into a {@code @TempDir}, which meant nothing anyone wrote into the live
* {@code fleetd.yaml} could ever fail it. This class is what a real charter is actually checked
* against at boot; {@code FleetdStartupValidationTest} exercises it through {@code Fleetd.main}
* itself, the same way it proves every other {@code validateXxx()} still runs there.
*
* <p><strong>fleetd #474</strong> — startup was not the only door: {@code
* dev.ltms.fleet.config.ConfigRef#reload()} used to run {@code FleetConfig#validateAll()} alone,
* which does not look at what a charter's text names, so a charter naming an unregistered tool
* that could not have booted the daemon could still be installed into a running one through a
* reload. This class still knows nothing about {@code ConfigRef} — {@code Fleetd.main} wires a
* small {@code FleetConfig -> void} adapter over {@link #assertChartersNameOnlyRegisteredTools}
* ({@code Fleetd::assertChartersNameOnlyRegisteredTools}) into {@code ConfigRef}'s constructor as
* its {@code Consumer<FleetConfig>} {@code extraValidation}, run inside {@code reload()}'s same
* try/catch as {@code validateAll()}, so both call sites — {@code Fleetd.main} at startup and
* {@code ConfigRef#reload()} afterwards — go through this one method and can never check different
* things. {@code dev.ltms.fleet.config.ConfigRefTest} and {@code
* FleetdConfigRefCharterToolSurfaceWiringTest} are what prove the reload call site, the same way
* {@code FleetdStartupValidationTest} proves the startup one.
*/
public final class CharterToolSurface {
/** A {@code fleet_…} (current) or {@code bridge_…} (pre-CB-634) tool-shaped token in prose. */
private static final Pattern TOOL_REFERENCE = Pattern.compile("(fleet_[a-z_]+|bridge_[a-z_]+)");
private CharterToolSurface() {
}
/**
* @param charters the configured {@code fleet.charters:} map (role wire name → charter text);
* {@code null} or empty is a no-op, same as an absent {@code fleet:} block
* @throws IllegalStateException naming the charter key and every tool it names that {@link
* FleetTool} does not list, when any charter does so
*/
public static void assertChartersNameOnlyRegisteredTools(Map<String, String> charters) {
if (charters == null || charters.isEmpty()) {
return;
}
Set<String> registered = FleetTool.wireNames();
List<String> bad = new ArrayList<>();
charters.forEach((key, text) -> {
if (text == null) {
return;
}
Set<String> named = new LinkedHashSet<>();
Matcher m = TOOL_REFERENCE.matcher(text);
while (m.find()) {
named.add(m.group(1));
}
named.stream()
.filter(t -> !registered.contains(t))
.forEach(unknown -> bad.add("fleet.charters." + key + " names '" + unknown
+ "', which the server does not register (registered: " + registered + ")."));
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
}
}
}
@@ -490,6 +490,21 @@ public final class FleetMcp {
McpSchema.Tool fleetProfiles = profilesTool();
McpSchema.Tool fleetWhoami = whoamiTool();
// fleetd #469: the tool schemas above are already named from FleetTool.wireName(), but
// this is the check that a schema was not accidentally dropped, duplicated, or added
// under a name FleetTool does not list. It runs once, at construction (startup), rather
// than being left to CharterToolSurface or a test to discover later — a canonical entry
// this server never registers, or a registration with no canonical entry backing it, is a
// startup failure, not a silent gap.
Set<String> registeredToolNames = Set.of(fleetSend.name(), fleetReply.name(), fleetAsk.name(),
fleetStatus.name(), fleetPoll.name(), fleetAck.name(), fleetSpawn.name(),
fleetList.name(), fleetStop.name(), fleetProfiles.name(), fleetWhoami.name());
if (!registeredToolNames.equals(FleetTool.wireNames())) {
throw new IllegalStateException("fleetd #469: registered MCP tools " + registeredToolNames
+ " do not match the canonical tool set " + FleetTool.wireNames()
+ " -- FleetTool is the single source of truth for what this server registers");
}
this.server = McpServer.sync(transport)
.serverInfo("fleet", "0.1.0")
.capabilities(McpSchema.ServerCapabilities.builder().tools(true).build())
@@ -897,20 +912,34 @@ public final class FleetMcp {
/**
* The action a registered tool handler actually hands to the authorization gate.
* Keeping this choice beside the registered-tool inventory makes a new tool fail the coverage
* test until its action is pinned.
*
* <p>{@code toolName} is a raw string off the wire (an MCP call names its tool by string, and a
* malformed or stale client can send anything), so resolving it against {@link FleetTool} first
* — and throwing on a miss — is still a run-time check by necessity. What moved to compile time
* is the second step: {@link #authzAction(FleetTool, Map)} switches on the resolved {@link
* FleetTool} itself with no {@code default}, so a new {@link FleetTool} constant with no pinned
* action fails {@code mvn compile}, not just {@code FleetMcpAuthzTest} at run time.
*/
static Authz.Action toolAction(String toolName, Map<String, Object> arguments) {
return switch (toolName) {
case "fleet_send" -> Authz.Action.SEND;
case "fleet_reply" -> Authz.Action.REPLY;
case "fleet_ask" -> Authz.Action.ASK;
case "fleet_status", "fleet_list", "fleet_profiles", "fleet_whoami" -> Authz.Action.READ;
case "fleet_poll" -> pollAction(str(arguments, "target"), str(arguments, "coordId"));
case "fleet_ack" -> Authz.Action.DRAIN;
case "fleet_spawn" -> Authz.Action.SPAWN;
case "fleet_stop" -> Authz.Action.STOP;
default -> throw new IllegalArgumentException("unregistered tool: " + toolName);
FleetTool tool = FleetTool.byWireName(toolName)
.orElseThrow(() -> new IllegalArgumentException("unregistered tool: " + toolName));
return authzAction(tool, arguments);
}
/**
* Exhaustive over {@link FleetTool} on purpose — no {@code default}. Adding a tool to {@link
* FleetTool} without adding its case here is a compile error (fleetd #469).
*/
private static Authz.Action authzAction(FleetTool tool, Map<String, Object> arguments) {
return switch (tool) {
case SEND -> Authz.Action.SEND;
case REPLY -> Authz.Action.REPLY;
case ASK -> Authz.Action.ASK;
case STATUS, LIST, PROFILES, WHOAMI -> Authz.Action.READ;
case POLL -> pollAction(str(arguments, "target"), str(arguments, "coordId"));
case ACK -> Authz.Action.DRAIN;
case SPAWN -> Authz.Action.SPAWN;
case STOP -> Authz.Action.STOP;
};
}
@@ -1269,6 +1298,13 @@ public final class FleetMcp {
* (the backend text that triggered the most recent quarantine of that credential, omitted when
* none is known) — so a lead can see WHICH model to turn off and WHY, without reading the
* daemon log.
*
* <p>fleetd #466 scope item 2: each {@code quarantined} row also names {@code
* quarantineAttempt} — 1 for a first-time exhaustion, 2 for the second in a row, and so on — so
* an operator sees "this is the 5th time" instead of inferring it from {@code
* quarantinedForSeconds} alone. Read off {@link BackendQuarantine#status(String)}, the same
* one-call accessor {@code quarantinedForSeconds} itself comes from here (see its doc) — never a
* separately derived count.
*/
public static Map<String, Object> profilesView(PeerLauncher workers, QuarantineSource quarantine, OutageSource outage) {
Map<String, Object> result = new LinkedHashMap<>();
@@ -1281,10 +1317,14 @@ public final class FleetMcp {
exhaustionDetectionArmed.put(profile, quarantine.exhaustedPatternArmed().apply(profile));
String credentialId = quarantine.credentialIdFor().apply(profile);
if (credentialId != null) {
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
// fleetd #466 scope item 2: quarantinedForSeconds and quarantineAttempt come off the
// ONE BackendQuarantine#status(credentialId) call — never a second, independent read
// for the attempt count — so the two can never disagree about which streak this is.
quarantine.quarantine().status(credentialId).ifPresent(status -> {
Map<String, Object> row = new LinkedHashMap<>();
row.put("credentialId", credentialId);
row.put("quarantinedForSeconds", remaining);
row.put("quarantinedForSeconds", status.remainingSeconds());
row.put("quarantineAttempt", status.repeatCount());
String model = quarantine.modelFor().apply(profile);
if (model != null && !model.isBlank()) {
row.put("model", model);
@@ -1724,10 +1764,13 @@ public final class FleetMcp {
}
String credentialId = quarantine.credentialIdFor().apply(profile);
if (credentialId != null) {
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
// fleetd #466 scope item 2: same one-call status() read as profilesView above — see its
// comment for why this must not become two separate lookups.
quarantine.quarantine().status(credentialId).ifPresent(status -> {
row.put("free", 0);
row.put("credentialId", credentialId);
row.put("quarantinedForSeconds", remaining);
row.put("quarantinedForSeconds", status.remainingSeconds());
row.put("quarantineAttempt", status.repeatCount());
});
}
String outageCredentialId = outage.credentialIdFor().apply(profile);
@@ -1805,7 +1848,7 @@ public final class FleetMcp {
// --- tool schemas --------------------------------------------------------------------------
private static McpSchema.Tool sendTool() {
return tool("fleet_send",
return tool(FleetTool.SEND.wireName(),
"Delegate a task to a worker session. By default blocks until the worker replies and "
+ "returns its reply (or a 'still working / queued' note on timeout). Pass wait:false "
+ "for a long task to return a ticket immediately, then poll it with fleet_poll. To "
@@ -1833,7 +1876,7 @@ public final class FleetMcp {
private static McpSchema.Tool askTool() {
// No target/session arg — the worker's identity is resolved from the connection.
return tool("fleet_ask",
return tool(FleetTool.ASK.wireName(),
"Pause your current delegated turn to ask the primary a question, blocking until it "
+ "answers — then resume the same turn with the answer. Use this when only the "
+ "primary has a decision or detail you need to continue. You do not address the "
@@ -1846,7 +1889,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool pollTool() {
return tool("fleet_poll",
return tool(FleetTool.POLL.wireName(),
"Check an async delegation (a fleet_send with wait:false) by its ticket: "
+ "pending, done (with the worker's reply), or failed. When target (a worker "
+ "session id) is present instead of ticket, drain that worker's inbox of "
@@ -1866,7 +1909,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool ackTool() {
return tool("fleet_ack",
return tool(FleetTool.ACK.wireName(),
"Acknowledge (remove) a specific reply from a worker's inbox. Use when the primary "
+ "has processed a reply and wants to confirm it, leaving other pending replies "
+ "in the inbox for later drain.",
@@ -1880,7 +1923,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool spawnTool() {
return tool("fleet_spawn",
return tool(FleetTool.SPAWN.wireName(),
"Spawn a new off-subscription member session. A member has two independent attributes: "
+ "role (what it is for) and profile (which backend it runs on). Pass role to pick "
+ "the contract — 'dev' implements a unit and opens its own PR, 'reviewer' reviews a "
@@ -1912,7 +1955,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool profilesTool() {
return tool("fleet_profiles",
return tool(FleetTool.PROFILES.wireName(),
"List the configured worker profiles (backends) and which one fleet_spawn uses by "
+ "default. A 'quarantined' map is present when a backend-exhausted refusal put "
+ "a profile's credential on cooldown — fleet_spawn onto it is refused until "
@@ -1930,7 +1973,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool listTool() {
return tool("fleet_list",
return tool(FleetTool.LIST.wireName(),
"List the whole fleet the bridge tracks, in two parts. 'leads' are your PEERS — other "
+ "orchestrators, each with its sessionId (the address to fleet_send to), "
+ "name, live status, and 'self': true on your own row; this is how you "
@@ -1966,7 +2009,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool stopTool() {
return tool("fleet_stop",
return tool(FleetTool.STOP.wireName(),
"Tear down a worker session by its paneId (from fleet_spawn or fleet_list).",
objectSchema(Map.of(
"paneId", stringProp("The worker's paneId to stop")),
@@ -1975,7 +2018,7 @@ public final class FleetMcp {
private static McpSchema.Tool replyTool() {
// No session/target arg — the caller's identity is resolved from the connection.
return tool("fleet_reply",
return tool(FleetTool.REPLY.wireName(),
"Return your structured answer for a message you were sent, resolving the sender's "
+ "blocked fleet_send. A worker MUST end every delegated turn with exactly "
+ "one of these. A lead uses it only to answer another lead that messaged "
@@ -1986,7 +2029,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool statusTool() {
return tool("fleet_status",
return tool(FleetTool.STATUS.wireName(),
"Get the live lifecycle status (idle/working/blocked/unknown) of a worker session.",
objectSchema(Map.of(
"sessionId", stringProp("The worker session id to query")),
@@ -1994,7 +2037,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool whoamiTool() {
return tool("fleet_whoami",
return tool(FleetTool.WHOAMI.wireName(),
"Report who YOU are on the bridge — your role is resolved from your connection "
+ "(unforgeable), never from anything you claim. Returns role 'primary' (you "
+ "orchestrate: spawn/send/stop; reply ONLY to answer a peer lead that "
@@ -0,0 +1,82 @@
package dev.ltms.fleet.mcp;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
/**
* The one canonical set of MCP tool names this daemon registers (fleetd #469, follow-up to #464).
*
* <p>Before this enum, the tool surface was written twice with nothing tying the copies together:
* once as the literal {@code "fleet_…"} string passed to each tool-schema builder in
* {@link FleetMcp}, and again as the case labels of {@link FleetMcp}'s authorization switch. A
* reader that needed "what does this server register" — a charter check, in particular — had no
* source to ask except scraping {@code FleetMcp.java}'s source text for {@code tool("…")} calls: a
* third copy of the same list, and the weakest of the three forms.
*
* <p>Every reader that needs the registered tool surface now asks this enum instead:
*
* <ul>
* <li>the tool-schema builders in {@code FleetMcp} pass {@code wireName()} rather than a literal;
* <li>{@code FleetMcp}'s constructor asserts, at startup, that the set of tool names it actually
* registers with the MCP SDK equals {@link #wireNames()} exactly — a canonical entry that is
* never registered, or a registration with no canonical entry backing it, fails the daemon's
* own boot rather than only a test's;
* <li>{@code FleetMcp.toolAction}'s dispatch onto {@code Authz.Action} switches on the enum
* (not the raw string) with no {@code default}, so adding a tool here without pinning its
* action is a compile error, not a run-time throw;
* <li>{@link CharterToolSurface} asks {@link #wireNames()} to check a configured launch charter
* against the live tool surface, instead of scraping source text a third time.
* </ul>
*/
public enum FleetTool {
SEND("fleet_send"),
REPLY("fleet_reply"),
ASK("fleet_ask"),
STATUS("fleet_status"),
POLL("fleet_poll"),
ACK("fleet_ack"),
SPAWN("fleet_spawn"),
LIST("fleet_list"),
STOP("fleet_stop"),
PROFILES("fleet_profiles"),
WHOAMI("fleet_whoami");
private final String wireName;
FleetTool(String wireName) {
this.wireName = wireName;
}
/** The name this tool is registered under, and called by, on the wire ({@code "fleet_send"}, …). */
public String wireName() {
return wireName;
}
private static final Map<String, FleetTool> BY_WIRE_NAME;
private static final Set<String> WIRE_NAMES;
static {
Map<String, FleetTool> byName = new LinkedHashMap<>();
Set<String> names = new LinkedHashSet<>();
for (FleetTool tool : values()) {
byName.put(tool.wireName, tool);
names.add(tool.wireName);
}
BY_WIRE_NAME = Map.copyOf(byName);
WIRE_NAMES = Set.copyOf(names);
}
/** The tool named {@code wireName}, or empty when this daemon registers no such tool. */
public static Optional<FleetTool> byWireName(String wireName) {
return Optional.ofNullable(BY_WIRE_NAME.get(wireName));
}
/** Every wire name this daemon registers — the canonical tool surface. */
public static Set<String> wireNames() {
return WIRE_NAMES;
}
}
@@ -525,17 +525,28 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
charter.toFile().deleteOnExit();
// fleetd #393 follow-up: withArray, not putArray. putArray REPLACES whatever node
// is already at "instructions" — harmless only as long as this block runs first
// against a still-empty root, which is an ordering constraint nothing declared or
// tested. The skills writer just below, and the IDE-rules writer further down,
// both already use withArray (get-or-create) for exactly this reason; this was the
// one straggler. Proven load-bearing on the fleetd #393 merge: flipping this one
// call back to putArray left the whole suite green while silently deleting the
// charter entry whenever skills or IDE rules ran after it — an opencode member
// would launch with no role contract at all, worse than the bug #393 fixed, and
// nothing caught it. See OpenCodeLauncherTest's
// instructionsArrayHoldsCharterThenIdeRulesInOrder and
// instructionsArrayHoldsCharterThenSkillsThenIdeRulesInOrder.
// is already at "instructions"; withArray gets-or-creates. All three writers on
// this array now use withArray so that write order stops being load-bearing for
// whoever adds a fourth.
//
// Be precise about what this particular line is worth, because an earlier version
// of this comment was wrong and claimed too much. THIS call is the one place where
// the two idioms are equivalent, and no test can tell them apart: it runs first,
// against a still-empty root, so there is never an existing node for putArray to
// replace. Measured on the #393 merge: flipping this one call back to putArray
// leaves the whole suite green (1618 tests), and always will. The edit is a
// readability and future-proofing change with no test behind it, and that is not a
// gap anyone can close.
//
// The ordering hazard is real, just not here. It is the LATER writers that can
// destroy earlier entries. Measured on the same merge: flipping the skills writer
// below to putArray deletes this charter entry and fails
// OpenCodeLauncherTest.instructionsArrayHoldsCharterThenSkillsThenIdeRulesInOrder;
// flipping the IDE-rules writer fails that test and
// instructionsArrayHoldsCharterThenIdeRulesInOrder. Deleting this line altogether
// fails instructionsArrayHoldsExactlyTheCharterWhenNothingElseWritesToIt — the
// assertion that a role contract reaches an opencode member at all, which is the
// hole that predates #393 and is what actually let the mutation hide.
root.withArray("instructions").add(charter.toAbsolutePath().toString());
}
@@ -3,6 +3,7 @@ package dev.ltms.fleet.placement;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Objects;
import java.util.Optional;
import java.util.OptionalLong;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.LongSupplier;
@@ -21,46 +22,158 @@ import java.util.function.LongSupplier;
* <p>The clock is injected ({@link LongSupplier}, conventionally {@code System::nanoTime} like
* {@code FleetHealthMonitor}), never read inline, so a quarantine's expiry is testable without a
* real sleep.
*
* <h2>Escalation (fleetd #466)</h2>
* A flat cooldown does not fit every exhaustion. A backend that reports "out of capacity for the
* rest of the hour" recovers in one cooldown; a weekly subscription limit does not — it keeps
* reporting exhausted on every attempt made before the window resets, so a flat 30-minute cooldown
* (the default {@code cooldownNanos}) means roughly 336 pointless spawn attempts across a week, one
* every cooldown.
*
* <p><strong>This class only ever sees the exhaustion signal.</strong> Its only production caller is
* {@code Fleetd.exhaustionSink}, wired to fire on a {@code BACKEND_EXHAUSTED} classification alone.
* The daemon's other outage state — a credential "cooling off" after repeated non-exhaustion
* backend errors (an HTTP 5xx storm, say) — is a separate mechanism, {@code BackendOutagePolicy},
* with its own short fixed 60s cooldown and no repeat tracking. The two are never merged: escalating
* on a cooling-off signal would turn a transient 5xx storm into a multi-hour backoff, which is
* exactly the failure this ticket is not asking for. Confirmed by reading every call site of
* {@link #quarantine} — {@code BackendOutagePolicy} has its own {@code coolOff} method and never
* calls this one.
*
* <p><strong>Mechanism</strong> — the {@link #withEscalation} constructors track, per credential, how
* many times in a row {@link #quarantine} has been called without an intervening "quiet" gap.
* Each call computes {@code cooldownNanos * backoffMultiplier ^ (repeatCount - 1)}, capped at
* {@code maxCooldownNanos}. A call counts as a continuation of the same streak — {@code repeatCount}
* increments — when it arrives no more than one base {@code cooldownNanos} after the previous
* quarantine's deadline (this covers both "still quarantined" and "quarantine just expired and it
* was exhausted again immediately"); otherwise the streak resets and this call is treated as a fresh
* first occurrence at the base cooldown.
*
* <p><strong>Reset, honestly stated.</strong> The ideal reset signal is "the cooldown expired and the
* next attempt succeeded" — but nothing in this codebase reports a spawn success back to this class
* (checked: {@code SessionManager} and {@code CompositePeerLauncher} never call any method here
* except {@link #quarantine}/{@link #isQuarantined}/{@link #remainingSeconds}, none of which is a
* success hook). Lacking that signal, the reset used here is a time-based proxy: a base-cooldown's
* worth of quiet — no exhaustion report for that credential — since the last quarantine ended. It is
* not proof the credential started working again, only the best available evidence without adding an
* active probe, which is out of scope by the operator's own design constraint (no automatic probing
* of a limited backend).
*
* <p><strong>Ceiling.</strong> {@code maxCooldownNanos} bounds the growth — an unbounded backoff is a
* permanent, unrecoverable-without-a-restart outage, which would be worse than the flat-rate bug this
* escalation fixes. {@link #withEscalation(LongSupplier, long)} defaults the ceiling to
* {@value #DEFAULT_MAX_COOLDOWN_MULTIPLE}x the base cooldown (12x the 1800s default ≈ 6 hours), so a
* chronically exhausted credential still gets re-tried roughly every 6 hours instead of every 30
* minutes — about a dozen attempts a week instead of ~336.
*
* <p><strong>Backward compatibility.</strong> The original two-argument {@link #BackendQuarantine(
* LongSupplier, long)} constructor is unchanged in behaviour: it is exactly {@code
* withEscalation}'s mechanism with {@code backoffMultiplier = 1.0} and {@code maxCooldownNanos =
* cooldownNanos}, which collapses the formula back to the original flat {@code now + cooldownNanos}
* on every call regardless of history. Every existing call site (roughly 20 across the test suite,
* plus {@link #none()}) keeps its current shape and behaviour unchanged.
*/
public final class BackendQuarantine {
private final ConcurrentHashMap<String, Long> quarantinedUntilNanos = new ConcurrentHashMap<>();
/** Default growth per consecutive exhaustion streak — see the class doc's Mechanism section. */
static final double DEFAULT_BACKOFF_MULTIPLIER = 2.0;
/** Default ceiling, expressed as a multiple of the base cooldown — see the class doc's Ceiling section. */
static final long DEFAULT_MAX_COOLDOWN_MULTIPLE = 12;
private final ConcurrentHashMap<String, QuarantineState> quarantines = new ConcurrentHashMap<>();
private final LongSupplier nowNanos;
private final long cooldownNanos;
private final double backoffMultiplier;
private final long maxCooldownNanos;
/** True only for {@link #none()}. See {@link #quarantine} for why this exists. */
private final boolean inert;
/** How many consecutive exhaustion reports a credential is on, and when the resulting cooldown ends. */
private record QuarantineState(int repeatCount, long deadlineNanos) {
}
/**
* Flat cooldown, unchanged from before fleetd #466 — every {@link #quarantine} call blocks the
* credential for exactly {@code cooldownNanos}, regardless of how many times it was called
* before. Equivalent to {@link #withEscalation} with no growth ({@code backoffMultiplier = 1.0})
* and a ceiling equal to the base cooldown, so it degrades to the identical {@code now +
* cooldownNanos} formula every call. Kept for the existing call sites that want a fixed cooldown
* (and for tests exercising the fixed-cooldown shape in isolation); production wiring uses
* {@link #withEscalation} instead.
*
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos how long a fresh {@link #quarantine} call blocks the credential for;
* must be positive
*/
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos) {
this(nowNanos, cooldownNanos, false);
this(nowNanos, cooldownNanos, 1.0, cooldownNanos, false);
}
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, boolean inert) {
/**
* Escalating cooldown (fleetd #466) — see the class doc's Mechanism/Reset/Ceiling sections.
*
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos base cooldown, applied to a fresh (non-streak) exhaustion; must be
* positive
* @param backoffMultiplier growth per consecutive exhaustion; must be {@code >= 1.0} ({@code 1.0}
* disables growth and is exactly the flat two-argument constructor)
* @param maxCooldownNanos ceiling on the escalated cooldown; must be {@code >= cooldownNanos}
*/
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, double backoffMultiplier,
long maxCooldownNanos) {
this(nowNanos, cooldownNanos, backoffMultiplier, maxCooldownNanos, false);
}
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, double backoffMultiplier,
long maxCooldownNanos, boolean inert) {
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
if (cooldownNanos <= 0) {
throw new IllegalArgumentException("cooldownNanos must be positive: " + cooldownNanos);
}
if (backoffMultiplier < 1.0) {
throw new IllegalArgumentException("backoffMultiplier must be >= 1.0: " + backoffMultiplier);
}
if (maxCooldownNanos < cooldownNanos) {
throw new IllegalArgumentException(
"maxCooldownNanos must be >= cooldownNanos: " + maxCooldownNanos + " < " + cooldownNanos);
}
this.cooldownNanos = cooldownNanos;
this.backoffMultiplier = backoffMultiplier;
this.maxCooldownNanos = maxCooldownNanos;
this.inert = inert;
}
/**
* Escalating cooldown with the fleetd #466 default shape: cooldown doubles
* ({@value #DEFAULT_BACKOFF_MULTIPLIER}x) per consecutive exhaustion streak, capped at
* {@value #DEFAULT_MAX_COOLDOWN_MULTIPLE}x the base cooldown. This is what production wiring
* ({@code Fleetd.main}) uses.
*
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos base cooldown, applied to a fresh (non-streak) exhaustion; must be positive
*/
public static BackendQuarantine withEscalation(LongSupplier nowNanos, long cooldownNanos) {
return new BackendQuarantine(nowNanos, cooldownNanos, DEFAULT_BACKOFF_MULTIPLIER,
cooldownNanos * DEFAULT_MAX_COOLDOWN_MULTIPLE, false);
}
/**
* Inert quarantine — {@link #quarantine} does nothing on this instance, so nothing is ever
* quarantined. The explicit stand-in a caller (or a test not exercising this feature) passes
* instead of a defaulting overload, exactly like {@code ExhaustedPatternLookup.none()}.
*/
public static BackendQuarantine none() {
return new BackendQuarantine(() -> 0L, 1, true);
return new BackendQuarantine(() -> 0L, 1, 1.0, 1, true);
}
/**
* Quarantine {@code credentialId} for the configured cooldown, starting now. A repeat call while
* already quarantined restarts the cooldown at full length — a fresh refusal is fresh evidence the
* account is still exhausted, not a reason to let an earlier, shorter wait stand.
* Quarantine {@code credentialId} starting now. On a flat instance (the two-argument
* constructor) this always blocks for exactly {@code cooldownNanos}, restarting the cooldown at
* full length on every call — a fresh refusal is fresh evidence the account is still exhausted,
* not a reason to let an earlier, shorter wait stand. On an escalating instance ({@link
* #withEscalation}) the cooldown grows with each call that arrives within one base cooldown of
* the previous deadline, and resets to the base cooldown once a call arrives after a longer gap
* — see the class doc.
*
* <p>On {@link #none()} this is a no-op. It has to be: that instance holds a clock frozen at 0,
* so recording a deadline would produce a quarantine that never expires — a credential locked out
@@ -73,7 +186,13 @@ public final class BackendQuarantine {
if (inert) {
return;
}
quarantinedUntilNanos.put(credentialId, nowNanos.getAsLong() + cooldownNanos);
long now = nowNanos.getAsLong();
quarantines.compute(credentialId, (id, prev) -> {
int repeatCount = (prev == null || now - prev.deadlineNanos() > cooldownNanos)
? 1
: prev.repeatCount() + 1;
return new QuarantineState(repeatCount, now + escalatedCooldownNanos(repeatCount));
});
}
/** Whether {@code credentialId} is quarantined right now. */
@@ -87,6 +206,49 @@ public final class BackendQuarantine {
return remaining > 0 ? OptionalLong.of(toSecondsRoundedUp(remaining)) : OptionalLong.empty();
}
/**
* Remaining seconds together with which consecutive exhaustion this is (fleetd #466 scope item
* 2) — {@code repeatCount} 1 for a first occurrence, 2 for the second in a row, and so on; see
* {@link #quarantine}'s class-doc Mechanism section for exactly when a call continues a streak
* versus starts a fresh one.
*
* <p><strong>Read together, off the one {@link QuarantineState} entry {@link #quarantine} itself
* wrote</strong> — a single {@code quarantines.get(credentialId)}, never a separate lookup or a
* value re-derived from {@code remainingSeconds} (e.g. inverting {@link
* #escalatedCooldownNanos}). That inversion is not just extra work to avoid: once a streak has
* hit {@code maxCooldownNanos}, every further consecutive exhaustion reports the identical
* cooldown, so a derivation that starts from the cooldown value cannot tell the 4th repeat from
* the 9th — only the stored {@code repeatCount} can. This is the same rule {@code
* CompositePeerLauncher.modelGateState()} documents for its own gate/report pair: the report
* reads the exact accessor the behaviour reads, so it can never disagree with what actually
* happened (the fleetd #404/#422 lesson). {@code fleet_profiles}/{@code fleet_list}/{@code GET
* /profiles} all call this — never {@link #remainingSeconds} plus a second, independent count —
* for exactly that reason.
*
* @return empty when {@code credentialId} is not currently quarantined (including on {@link
* #none()}, which quarantines nothing)
*/
public Optional<Status> status(String credentialId) {
QuarantineState state = quarantines.get(credentialId);
if (state == null) {
return Optional.empty();
}
long remaining = state.deadlineNanos() - nowNanos.getAsLong();
return remaining > 0
? Optional.of(new Status(toSecondsRoundedUp(remaining), state.repeatCount()))
: Optional.empty();
}
/**
* @param remainingSeconds seconds left on the quarantine, identical to {@link
* #remainingSeconds(String)}'s answer for the same credential at the
* same instant
* @param repeatCount 1 for a first occurrence, 2 for the second consecutive one, etc. —
* see {@link #status(String)}
*/
public record Status(long remainingSeconds, int repeatCount) {
}
/**
* Every currently-quarantined credential id and its remaining seconds (CB-578 stage B fleet
* reporting) — expired entries are never included. Not pruned from the backing map here: it stays
@@ -95,8 +257,8 @@ public final class BackendQuarantine {
*/
public Map<String, Long> activeRemainingSeconds() {
Map<String, Long> out = new LinkedHashMap<>();
quarantinedUntilNanos.forEach((credentialId, deadline) -> {
long remaining = deadline - nowNanos.getAsLong();
quarantines.forEach((credentialId, state) -> {
long remaining = state.deadlineNanos() - nowNanos.getAsLong();
if (remaining > 0) {
out.put(credentialId, toSecondsRoundedUp(remaining));
}
@@ -105,8 +267,14 @@ public final class BackendQuarantine {
}
private long remainingNanos(String credentialId) {
Long deadline = quarantinedUntilNanos.get(credentialId);
return deadline == null ? 0L : deadline - nowNanos.getAsLong();
QuarantineState state = quarantines.get(credentialId);
return state == null ? 0L : state.deadlineNanos() - nowNanos.getAsLong();
}
/** {@code cooldownNanos * backoffMultiplier ^ (repeatCount - 1)}, capped at {@code maxCooldownNanos}. */
private long escalatedCooldownNanos(int repeatCount) {
double raw = cooldownNanos * Math.pow(backoffMultiplier, repeatCount - 1);
return raw >= (double) maxCooldownNanos ? maxCooldownNanos : (long) raw;
}
private static long toSecondsRoundedUp(long nanos) {
@@ -0,0 +1,65 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #466 follow-up: {@code Fleetd.main} builds the daemon's one {@code BackendQuarantine}
* from {@link dev.ltms.fleet.placement.BackendQuarantine#withEscalation(java.util.function.LongSupplier,
* long)} — the escalating factory — rather than the plain two-argument constructor, which is still a
* flat cooldown (kept for backward compatibility, see that class's doc). {@code
* BackendQuarantineTest} proves {@code withEscalation} itself escalates, is ceilinged, and resets;
* it says nothing about which one {@code main} actually calls.
*
* <p>Measured directly: reverting {@code main} to {@code new BackendQuarantine(System::nanoTime,
* TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()))} — the pre-#466 flat call — compiles
* with 0 errors and leaves the entire 1608-test suite (including every {@code BackendQuarantineTest}
* case) green, because no other test constructs its {@code BackendQuarantine} through {@code main};
* every one of them builds its own instance directly. That silent regression is exactly the shape
* {@link FleetdLeadSeatWiringTest} and {@link FleetdCompletionResolverWiringTest} already guard
* against for their own constructor arguments — this is the same class of gap for fleetd #466's
* factory choice, following their approach.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* BackendQuarantine} and never runs {@code main} — a green result here proves only that the exact
* text {@code main} calls {@code BackendQuarantine.withEscalation(...)} rather than the flat
* constructor. It does not prove that call actually executes at startup (no test here starts the
* daemon), and it does not prove the escalation reaches a real backend or credential — only
* {@code BackendQuarantineTest} proves the factory's own behaviour, and only a live daemon proves
* the wiring runs.
*/
class FleetdBackendQuarantineWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] main's BackendQuarantine local is still built from BackendQuarantine.withEscalation(...)")
void mainStillWiresTheEscalatingQuarantineFactory() throws Exception {
String source = fleetdSource();
assertTrue(source.contains(
"BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,\n"
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
"Fleetd.main's BackendQuarantine local must still be built from "
+ "BackendQuarantine.withEscalation(System::nanoTime, "
+ "TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds())). Reverting to the flat "
+ "two-argument constructor (fleetd #466's measured regression) compiles with 0 errors "
+ "and leaves the whole suite green, including every BackendQuarantineTest case that "
+ "proves the escalation itself works — this source check is what must go red instead. "
+ "A reverted daemon would go back to retrying a weekly subscription limit on every "
+ "flat ~30-minute cooldown, about 336 times across the week.");
// Negative form of the same check: the pre-#466 flat call, if it ever reappears at this
// declaration, must not be mistaken for the escalating one by a looser positive-only check.
assertFalse(source.contains(
"BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,\n"
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
"main's BackendQuarantine local must never regress to the flat two-argument constructor");
}
}
@@ -0,0 +1,106 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #474: proves the exact wiring {@code Fleetd.main} uses to construct its live {@code
* ConfigRef} — {@code new ConfigRef(configPath, cfg, Fleetd::assertChartersNameOnlyRegisteredTools)}
* — actually makes {@link ConfigRef#reload()} refuse a charter that names an MCP tool the server
* does not register, the same way {@code Fleetd.main} itself refuses one at startup (see {@code
* FleetdStartupValidationTest#mainRefusesACharterNamingAnUnregisteredTool}).
*
* <p>{@code dev.ltms.fleet.config.ConfigRefTest} pins the same behaviour through a locally-built
* {@code Consumer<FleetConfig>} adapter that calls the same production {@code CharterToolSurface}
* method, because that test lives in {@code dev.ltms.fleet.config} and cannot see {@code
* Fleetd#assertChartersNameOnlyRegisteredTools} (package-private to {@code dev.ltms.fleet}). This
* class is the companion proof that lives where the real method reference is visible, so the literal
* expression {@code Fleetd::assertChartersNameOnlyRegisteredTools} — not just an equivalent — is
* what gets exercised. {@code Fleetd.main} itself cannot be driven this far in a unit test: every
* fixture in {@code FleetdStartupValidationTest} is deliberately invalid so {@code main} throws
* before opening a socket, binding Javalin, or doing anything else with a real side effect, so a
* test cannot get {@code main} far enough to hold a running daemon it could then reload — this test
* builds the {@code ConfigRef} the same way {@code main} does and drives {@link ConfigRef#reload()}
* directly instead, the same shape {@code FleetdExhaustionDetectionArmedWiringTest} and its
* siblings already use for the rest of {@code Fleetd.main}'s wiring.
*/
class FleetdConfigRefCharterToolSurfaceWiringTest {
private static final String BASE = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
guard:
offSubscriptionHosts:
- gx00.gw
""";
@Test
void reloadRefusesACharterNamingAnUnregisteredToolThroughFleetdsOwnWiring(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, BASE + """
fleet:
charters:
dev: |
Send the final handoff through fleet_reply.
""");
ConfigRef config = new ConfigRef(f, FleetConfig.load(f),
Fleetd::assertChartersNameOnlyRegisteredTools);
FleetConfig before = config.get();
Files.writeString(f, BASE + """
fleet:
charters:
dev: |
Send the final handoff through bridge_send.
""");
ConfigRef.Outcome out = config.reload();
assertFalse(out.applied());
assertNotNull(out.error());
assertTrue(out.error().contains("dev"), out.error());
assertTrue(out.error().contains("bridge_send"), out.error());
assertSame(before, config.get());
}
@Test
void reloadAcceptsACharterNamingOnlyRegisteredToolsThroughFleetdsOwnWiring(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, BASE + """
fleet:
charters:
dev: old charter
""");
ConfigRef config = new ConfigRef(f, FleetConfig.load(f),
Fleetd::assertChartersNameOnlyRegisteredTools);
Files.writeString(f, BASE + """
fleet:
charters:
dev: |
Send the final handoff through fleet_reply.
""");
ConfigRef.Outcome out = config.reload();
assertTrue(out.applied());
assertEquals("config reloaded", out.summary());
}
}
@@ -0,0 +1,80 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #474 follow-up: {@code Fleetd.main} builds its live {@code ConfigRef} from the
* three-argument constructor, {@code new ConfigRef(configPath, cfg,
* Fleetd::assertChartersNameOnlyRegisteredTools)}, so a reload runs the same charter tool-surface
* gate startup does (see {@link dev.ltms.fleet.config.ConfigRef}'s class doc, "fleetd #474" bullet).
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} prove the
* three-argument constructor and the {@code Fleetd.assertChartersNameOnlyRegisteredTools} adapter
* work correctly together — both build their OWN {@code ConfigRef} with that constructor, so neither
* proves {@code main} still CHOOSES the three-argument form over the plain two-argument {@code new
* ConfigRef(configPath, cfg)}.
*
* <p>Measured directly: reverting {@code Fleetd.java}'s {@code config} local to the two-argument
* constructor compiles with 0 errors and leaves the entire 1633-test suite green — including every
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} case — because
* neither of those tests constructs its {@code ConfigRef} through {@code main}; both build their own
* instance directly, wired with the check by hand. That silent regression is exactly the shape
* {@link FleetdBackendQuarantineWiringTest}, {@link FleetdLeadSeatWiringTest} and {@link
* FleetdCompletionResolverWiringTest} already guard against for their own constructor arguments —
* this class is the same class of gap for fleetd #474's {@code extraValidation} argument, following
* their approach.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* ConfigRef} and never runs {@code main} — a green result here proves only that the exact text
* {@code main} contains is the three-argument construction with {@code
* Fleetd::assertChartersNameOnlyRegisteredTools}. It does not prove that call actually executes at
* startup (no test here starts the daemon), and it does not prove the reload gate itself works —
* only {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} prove the
* behaviour; only a live daemon proves the wiring runs.
*/
class FleetdConfigRefWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] main still builds config from the three-argument ConfigRef constructor")
void mainStillWiresTheThreeArgumentConfigRefConstructor() throws Exception {
String source = fleetdSource();
// A broken read (wrong working directory, wrong path, a file that came back empty) would
// make the assertFalse below pass vacuously — a "clean" negative check that actually
// checked nothing. Guard against that first, with an anchor that has nothing to do with
// this mutation, so a bad read fails loudly here instead of silently proving nothing below.
assertTrue(source.contains("public final class Fleetd"),
"fleetdSource() did not read anything usable — src/main/java/dev/ltms/fleet/Fleetd.java "
+ "did not come back containing its own class declaration. The assertFalse below "
+ "would pass vacuously on a broken read; fix the read before trusting this test.");
assertTrue(source.contains(
"ConfigRef config = new ConfigRef(configPath, cfg, "
+ "Fleetd::assertChartersNameOnlyRegisteredTools);"),
"Fleetd.main's config local must still be built from the three-argument ConfigRef "
+ "constructor, with Fleetd::assertChartersNameOnlyRegisteredTools as "
+ "extraValidation. Reverting to the plain two-argument constructor (fleetd #474's "
+ "measured M2 regression) compiles with 0 errors and leaves the whole suite green — "
+ "including ConfigRefTest and FleetdConfigRefCharterToolSurfaceWiringTest, because "
+ "neither builds its ConfigRef through main — this source check is what must go "
+ "red instead. A reverted daemon would accept, through a reload with no restart, "
+ "exactly the charter that #469/#474 already refuse at startup.");
// Negative form of the same check: the pre-#474 two-argument call, if it ever reappears at
// this declaration, must not be mistaken for the three-argument one by a looser
// positive-only check — this is the M2 mutation this test exists to kill.
assertFalse(source.contains("ConfigRef config = new ConfigRef(configPath, cfg);"),
"main's config local must never regress to the plain two-argument ConfigRef "
+ "constructor — that drops the reload-path charter check (fleetd #474's measured "
+ "M2 mutation) with no other test catching it");
}
}
@@ -95,6 +95,28 @@ class FleetdStartupValidationTest {
""", "architetc");
}
/**
* fleetd #469: closes the gap left by #464's {@code CharterToolSurfaceTest}, which wrote its
* own charter into a {@code @TempDir} fixture and so could never fail on anything anyone wrote
* into the live {@code fleetd.yaml}. {@code bridge_send} is CB-634's own motivating example — a
* tool name the rename removed — and the charter key ({@code dev}) is a real role wire name, so
* this fixture passes {@code cfg.validateAll()}'s charter check (key valid, text non-blank) and
* is refused only by the new {@code CharterToolSurface} call right after it. The failure message
* must name both the charter key and the unknown tool.
*/
@Test
void mainRefusesACharterNamingAnUnregisteredTool(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "charter-tool-surface.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
charters:
dev: |
Send the final handoff through bridge_send.
""", "bridge_send");
}
@Test
void mainRefusesAnArchitectSlotNamingAnUnconfiguredProfile(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "members.yaml", """
@@ -1,10 +1,13 @@
package dev.ltms.fleet.config;
import dev.ltms.fleet.mcp.CharterToolSurface;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.function.Consumer;
import static org.junit.jupiter.api.Assertions.*;
@@ -38,6 +41,23 @@ class ConfigRefTest {
return new ConfigRef(f, FleetConfig.load(f));
}
/**
* The exact {@code Consumer<FleetConfig>} {@code Fleetd.main} wires into {@code ConfigRef}'s
* constructor as {@code extraValidation} (fleetd #474) — an adapter from {@code FleetConfig} to
* the raw charter map {@link CharterToolSurface#assertChartersNameOnlyRegisteredTools} takes.
* Built here rather than referencing {@code dev.ltms.fleet.Fleetd} directly, because that method
* is package-private to {@code dev.ltms.fleet} and this test lives in {@code
* dev.ltms.fleet.config} — but it calls the SAME production {@link CharterToolSurface} method
* {@code Fleetd} calls, so this proves the real check runs on reload, not a stand-in for it.
*/
private static final Consumer<FleetConfig> CHARTER_TOOL_SURFACE = cfg ->
CharterToolSurface.assertChartersNameOnlyRegisteredTools(
cfg.fleet() == null ? Map.of() : cfg.fleet().charters());
private static ConfigRef refForWithCharterToolSurface(Path f) {
return new ConfigRef(f, FleetConfig.load(f), CHARTER_TOOL_SURFACE);
}
@Test
void aHotChangeIsAppliedAndReadThroughGet(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
@@ -124,6 +144,87 @@ class ConfigRefTest {
assertSame(before, ref.get());
}
/**
* fleetd #474: {@code validateAll()} (via {@code validateCharters()}) only checks that a
* charter's KEY is a role wire name and its text is non-blank — it never looks at what the text
* names, so a charter naming {@code bridge_send} (the pre-CB-634 name, removed from the tool
* surface — #469's own motivating example) passes {@code validateAll()} and used to be applied
* on reload with nothing refusing it, even though the identical charter refuses {@code
* Fleetd.main} at startup ({@code FleetdStartupValidationTest
* #mainRefusesACharterNamingAnUnregisteredTool}). This is the reload-path proof: it drives
* {@link ConfigRef#reload()} itself (not a direct call to {@link
* CharterToolSurface#assertChartersNameOnlyRegisteredTools}), through the exact {@code
* extraValidation} wiring {@code Fleetd.main} uses, and the failure message must name both the
* charter key and the unknown tool — the same information the startup failure gives (ticket
* acceptance criterion 1).
*/
@Test
void aReloadRefusesACharterNamingAnUnregisteredTool(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
fleet:
charters:
dev: |
Send the final handoff through fleet_reply.
"""));
ConfigRef ref = refForWithCharterToolSurface(f);
FleetConfig before = ref.get();
Files.writeString(f, yaml("""
fleet:
charters:
dev: |
Send the final handoff through bridge_send.
"""));
ConfigRef.Outcome out = ref.reload();
assertFalse(out.applied());
assertNotNull(out.error());
assertTrue(out.error().contains("dev"),
"expected the charter key 'dev' in the refusal, got: " + out.error());
assertTrue(out.error().contains("bridge_send"),
"expected the unknown tool 'bridge_send' in the refusal, got: " + out.error());
assertTrue(out.summary().startsWith("config reload refused"), out.summary());
// The running config must not move at all — half-applying this would leave the daemon in a
// state that could never have booted, exactly the outcome ConfigRef.java's class doc warns
// a cold-key refusal must avoid, and this check must avoid the same way.
assertSame(before, ref.get());
assertEquals("Send the final handoff through fleet_reply.\n",
ref.get().fleet().charterFor(dev.ltms.fleet.peer.MemberRole.DEV));
}
/**
* fleetd #474 acceptance criterion 3, the positive case: a reload whose charter names only
* registered tools must still be ACCEPTED. An inverted filter (one that refuses every charter,
* or refuses on any {@code fleet_*}/{@code bridge_*} token regardless of registration) would pass
* the refusal test above alone — #469's own M5 mutation cell showed exactly that shape surviving
* a negative-only suite. This is the test that catches it.
*/
@Test
void aReloadAcceptsACharterNamingOnlyRegisteredTools(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
fleet:
charters:
dev: old charter, no tool names
"""));
ConfigRef ref = refForWithCharterToolSurface(f);
Files.writeString(f, yaml("""
fleet:
charters:
dev: |
Send the final handoff through fleet_reply, using fleet_send to delegate.
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
assertEquals("config reloaded", out.summary());
assertEquals("Send the final handoff through fleet_reply, using fleet_send to delegate.\n",
ref.get().fleet().charterFor(dev.ltms.fleet.peer.MemberRole.DEV));
}
/**
* The point of the whole class: a consumer holding the ref sees the new value without being
* rebuilt. A component that captured {@code get()} into a field would still show the old one.
@@ -3,7 +3,9 @@ package dev.ltms.fleet.mcp;
import dev.ltms.fleet.config.FleetConfig;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.Map;
import java.util.Set;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
@@ -11,13 +13,30 @@ import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** fleetd #464: launch charters must not name MCP tools the server does not register. */
/**
* fleetd #464 shipped {@code configuredChartersNameOnlyRegisteredTools} below, comparing charter
* text against the registered tool surface — but it scraped <em>both</em> sides from source text:
* its own fixture charter, and a regex over {@code FleetMcp.java}'s {@code tool("…")} calls. #469's
* gap: nothing anyone wrote into the live {@code fleetd.yaml} could ever reach that test, because
* it never called production validation code.
*
* <p>This version keeps the charter-text extraction helper ({@code toolsNamedIn}) — charters are
* free-text config, so finding a {@code fleet_*}/{@code bridge_*} token inside one has no source
* but a scrape — but reads the <em>registered</em> side from {@link FleetTool}, the canonical enum
* {@code FleetMcp} itself now derives its tool schemas and authorization switch from, rather than a
* second scrape of {@code FleetMcp.java}'s source. It also exercises {@link
* CharterToolSurface#assertChartersNameOnlyRegisteredTools} directly — the method {@code
* Fleetd.main} actually calls at startup — both accepting and rejecting. {@code
* dev.ltms.fleet.FleetdStartupValidationTest#mainRefusesACharterNamingAnUnregisteredTool} is what
* closes #469's actual gap: it proves that call is wired into {@code Fleetd.main} itself, against a
* live-shaped config fixture, not just a unit call to the method in isolation.
*/
class CharterToolSurfaceTest {
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
private static Set<String> matches(String text, String regex) {
Matcher m = Pattern.compile(regex).matcher(text);
Set<String> found = new LinkedHashSet<>();
@@ -33,13 +52,8 @@ class CharterToolSurfaceTest {
"(fleet_[a-z_]+|bridge_[a-z_]+)");
}
/** Every tool {@link FleetMcp} registers, read from its {@code tool("…")} calls. */
private static Set<String> toolsTheServerRegisters() throws Exception {
return matches(Files.readString(MCP_SOURCE), "tool\\(\\\"(fleet_[a-z_]+)\\\"");
}
@Test
@DisplayName("[SOURCE TEXT] every tool named in a configured charter is registered by the server")
@DisplayName("every tool named in a configured charter is in the canonical FleetTool set")
void configuredChartersNameOnlyRegisteredTools(@TempDir Path dir) throws Exception {
Path configFile = dir.resolve("charters.yaml");
Files.writeString(configFile, """
@@ -53,20 +67,58 @@ class CharterToolSurfaceTest {
FleetConfig config = FleetConfig.load(configFile);
Set<String> named = toolsNamedIn(config);
Set<String> registered = toolsTheServerRegisters();
Set<String> registered = FleetTool.wireNames();
assertTrue(!named.isEmpty(),
"the charter fixture named no fleet_* or bridge_* tool. This test would check nothing; "
+ "add charter text that names a tool before changing the extraction.");
assertTrue(!registered.isEmpty(),
"the FleetMcp registration scrape found no tools. This test would check nothing; "
+ "repair the tool(\"…\") extraction before changing the assertion.");
"FleetTool.wireNames() is empty. This test would check nothing; repair FleetTool "
+ "before changing the assertion.");
Set<String> unknown = new LinkedHashSet<>(named);
unknown.removeAll(registered);
assertTrue(unknown.isEmpty(),
"configured charter text names " + unknown + ", but FleetMcp does not register it. "
"configured charter text names " + unknown + ", but FleetTool does not list it. "
+ "Checked " + named + " against " + registered + ". Fix the charter text or "
+ "register the tool; do NOT weaken this test.");
+ "add the tool to FleetTool; do NOT weaken this test.");
}
/**
* The positive case for the actual production entry point: a charter naming only tools
* {@link FleetTool} lists must not throw.
*/
@Test
@DisplayName("CharterToolSurface accepts a charter that names only registered tools")
void charterToolSurfaceAcceptsKnownTools() {
assertDoesNotThrow(() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(Map.of(
"dev", "Send the final handoff through fleet_reply, using fleet_send to delegate.",
"reviewer", "Use fleet_ask only for the lead's decision.")));
}
/**
* The negative case for the actual production entry point (fleetd #469's motivating example:
* {@code bridge_send} is the pre-CB-634 name, removed from the tool surface). The message must
* name both the offending charter key and the unknown tool, so an operator reading the startup
* log knows exactly which charter to fix.
*/
@Test
@DisplayName("CharterToolSurface rejects a charter naming a tool the server does not register")
void charterToolSurfaceRejectsAnUnregisteredTool() {
Map<String, String> charters = new LinkedHashMap<>();
charters.put("dev", "Send the final handoff through bridge_send.");
IllegalStateException e = assertThrows(IllegalStateException.class,
() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(charters));
assertTrue(e.getMessage().contains("dev"),
"expected the charter key 'dev' in the failure message, got: " + e.getMessage());
assertTrue(e.getMessage().contains("bridge_send"),
"expected the unknown tool 'bridge_send' in the failure message, got: " + e.getMessage());
}
@Test
@DisplayName("CharterToolSurface is a no-op on an absent or empty charter map")
void charterToolSurfaceIsANoOpWithNoCharters() {
assertDoesNotThrow(() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(null));
assertDoesNotThrow(() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(Map.of()));
}
}
@@ -26,7 +26,6 @@ import org.junit.jupiter.api.Test;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashSet;
import java.util.Map;
import java.util.Set;
import java.util.regex.Matcher;
@@ -50,7 +49,6 @@ import static org.junit.jupiter.api.Assertions.*;
class FleetMcpAuthzTest {
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
private static final Pattern TOOL_REGISTRATION = Pattern.compile("tool\\(\\\"(fleet_[a-z_]+)\\\"");
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
@@ -296,10 +294,18 @@ class FleetMcpAuthzTest {
@Test
void everyRegisteredToolHasItsHandlerActionPinned() {
Set<String> registered = toolsTheServerRegisters();
// fleetd #469: this used to scrape FleetMcp.java's tool("…") calls for the registered set —
// a third copy of the same list this file, CharterToolSurfaceTest and McpContractDocTest
// each kept independently. All three now read FleetTool.wireNames(), the canonical set
// FleetMcp itself derives its tool schemas AND its authorization switch from; adding a tool
// there without pinning its action in FleetMcp#authzAction is a compile error, so this test's
// per-tool assertions below are a run-time regression pin on top of that compile-time check,
// not the only thing standing between a new tool and an unpinned action.
Set<String> registered = FleetTool.wireNames();
assertTrue(registered.size() >= 10,
"scraped only " + registered.size() + " tool registrations from FleetMcp (" + registered
+ "); the server registers eleven, so the tool(\"…\") scrape has stopped matching");
"FleetTool.wireNames() returned only " + registered.size() + " tool(s) (" + registered
+ "); the server registers eleven, so FleetTool has stopped listing the real "
+ "tool surface");
registered.forEach(tool -> assertDoesNotThrow(() -> FleetMcp.toolAction(tool, Map.of()),
() -> tool + " is registered but has no pinned authorization action"));
@@ -319,19 +325,6 @@ class FleetMcpAuthzTest {
FleetMcp.toolAction("fleet_poll", Map.of("coordId", "mac-opus")));
}
private static Set<String> toolsTheServerRegisters() {
try {
Matcher matcher = TOOL_REGISTRATION.matcher(Files.readString(MCP_SOURCE));
Set<String> tools = new LinkedHashSet<>();
while (matcher.find()) {
tools.add(matcher.group(1));
}
return tools;
} catch (Exception e) {
throw new AssertionError("could not scrape FleetMcp tool registrations", e);
}
}
@Test
void aWorkerMayNotDrainAnotherSessionsInboxByPolling() {
FleetMcp m = mcp(true);
@@ -574,6 +574,38 @@ class FleetMcpTest {
assertTrue(out.contains("\"quarantinedForSeconds\":1800"), out);
}
/**
* fleetd #466 scope item 2: {@code fleet_profiles} must carry the repeat count beside the
* remaining seconds, and the two must come off the one {@link BackendQuarantine#status} call so
* they can never disagree about which streak this is (see {@code profilesView}'s javadoc).
*/
@Test
void profilesReportsQuarantineAttemptBesideRemainingSeconds() {
FakeHerdr h = new FakeHerdr();
java.util.concurrent.atomic.AtomicLong now = new java.util.concurrent.atomic.AtomicLong(0L);
BackendQuarantine quarantine = BackendQuarantine.withEscalation(now::get, TimeUnit.MINUTES.toNanos(30));
quarantine.quarantine("shared-openai"); // attempt 1: 1800s
now.set(TimeUnit.MINUTES.toNanos(30));
quarantine.quarantine("shared-openai"); // attempt 2: 3600s
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
profile -> "ltms-local".equals(profile) ? "shared-openai" : null, quarantine);
McpSchema.CallToolResult res = FleetMcp.profiles(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), source);
String out = textOf(res);
assertTrue(out.contains("\"quarantinedForSeconds\":3600"), out);
assertTrue(out.contains("\"quarantineAttempt\":2"), out);
}
/** A never-quarantined profile must not carry {@code quarantineAttempt} either. */
@Test
void profilesOmitsQuarantineAttemptWhenNotQuarantined() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = FleetMcp.profiles(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), FleetMcp.QuarantineSource.none());
String out = textOf(res);
assertFalse(out.contains("quarantineAttempt"), out);
}
/** fleetd #201 Unit 5: {@code coolingOff} is a SEPARATE map from {@code quarantined}. */
@Test
void profilesReportsACoolingOffCredentialInASeparateMap() {
@@ -1121,6 +1153,35 @@ class FleetMcpTest {
assertTrue(out.contains("\"quarantinedForSeconds\":1200"), out);
}
/**
* fleetd #466 scope item 2: {@code fleet_list}'s capacity rows (CB-583: they reuse quarantine)
* must carry {@code quarantineAttempt} beside {@code quarantinedForSeconds}, off the same
* {@link BackendQuarantine#status} call as {@code fleet_profiles} -- see {@code capacityView}'s
* comment pointing back to {@code profilesView}.
*/
@Test
void capacityRowReportsQuarantineAttemptBesideRemainingSeconds() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
java.util.concurrent.atomic.AtomicLong now = new java.util.concurrent.atomic.AtomicLong(0L);
BackendQuarantine quarantine = BackendQuarantine.withEscalation(now::get, TimeUnit.MINUTES.toNanos(20));
quarantine.quarantine("shared-openai"); // attempt 1
now.set(TimeUnit.MINUTES.toNanos(20));
quarantine.quarantine("shared-openai"); // attempt 2
now.set(TimeUnit.MINUTES.toNanos(60));
quarantine.quarantine("shared-openai"); // attempt 3
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
profile -> "terra".equals(profile) ? "shared-openai" : null, quarantine);
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 2,
() -> Set.of("terra"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
source, Map.of(), ""));
assertTrue(out.contains("\"credentialId\":\"shared-openai\""), out);
assertTrue(out.contains("\"quarantineAttempt\":3"), out);
}
@Test
void everyProfileSharingTheQuarantinedCredentialReportsZeroFree() {
FakeHerdr h = new FakeHerdr();
@@ -27,16 +27,21 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* somewhere to be readable, and that is exactly the sentence that rots. This test is what makes it
* safe to write.
*
* <p><b>It checks source text, not behaviour.</b> It reads the Markdown and reads {@link FleetMcp}'s
* source, and it only catches a name in the doc that the server does not register. It cannot catch a
* flow that describes the wrong order, or a parameter name in prose — those are not name-shaped. The
* doc's own header carries that caveat for its readers.
* <p><b>It checks source text, not behaviour.</b> It reads the Markdown, and it only catches a name
* in the doc that the server does not register. It cannot catch a flow that describes the wrong
* order, or a parameter name in prose — those are not name-shaped. The doc's own header carries
* that caveat for its readers.
*
* <p>fleetd #469: the registered side used to be its own scrape of {@code FleetMcp.java}'s {@code
* tool("…")} calls — a third copy of the same list {@code CharterToolSurfaceTest} and {@code
* FleetMcpAuthzTest} each kept their own copy of too. All three now read {@link
* FleetTool#wireNames()}, the one canonical set {@code FleetMcp} itself derives its tool schemas and
* authorization switch from.
*/
class McpContractDocTest {
/** Tests run with the module directory as cwd, so the repo-root doc is one level up. */
private static final Path DOC = Path.of("../docs/MCP-Contract.md");
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
private static Set<String> matches(Path file, String regex) throws Exception {
Matcher m = Pattern.compile(regex).matcher(Files.readString(file));
@@ -52,15 +57,10 @@ class McpContractDocTest {
return matches(DOC, "(fleet_[a-z_]+)");
}
/** Every tool {@link FleetMcp} actually registers, read from its {@code tool("…")} calls. */
private static Set<String> toolsTheServerRegisters() throws Exception {
return matches(MCP_SOURCE, "tool\\(\"(fleet_[a-z_]+)\"");
}
@Test
@DisplayName("[SOURCE TEXT] every fleet_* tool named in MCP-Contract.md is one the server registers")
void theDocNamesNoToolThatDoesNotExist() throws Exception {
Set<String> registered = toolsTheServerRegisters();
Set<String> registered = FleetTool.wireNames();
Set<String> named = toolsNamedInTheDoc();
Set<String> unknown = new LinkedHashSet<>(named);
@@ -84,13 +84,13 @@ class McpContractDocTest {
@Test
@DisplayName("[SOURCE TEXT] the doc/server name check is not vacuous — both sides found names")
void theCheckActuallyHasSomethingToCheck() throws Exception {
Set<String> registered = toolsTheServerRegisters();
Set<String> registered = FleetTool.wireNames();
Set<String> named = toolsNamedInTheDoc();
assertTrue(registered.size() >= 10,
"scraped only " + registered.size() + " tool registrations from FleetMcp (" + registered
+ "); the server registers eleven, so the tool(\"…\") scrape has stopped matching "
+ "and the check above is now vacuous");
"FleetTool.wireNames() returned only " + registered.size() + " tool(s) (" + registered
+ "); the server registers eleven, so FleetTool has stopped listing the real "
+ "tool surface and the check above is now vacuous");
assertTrue(named.size() >= 4,
"docs/MCP-Contract.md names only " + named.size() + " fleet_* tool(s) (" + named + "). "
+ "The flows describe delegation, clarification, detached delivery and the "
@@ -116,4 +116,262 @@ class BackendQuarantineTest {
assertThrows(IllegalArgumentException.class, () -> new BackendQuarantine(() -> 0L, 0L));
assertThrows(IllegalArgumentException.class, () -> new BackendQuarantine(() -> 0L, -1L));
}
// --- fleetd #466: escalating cooldown -----------------------------------------------------
//
// Base cooldown 600s (10 min), multiplier 2.0, ceiling 2400s (4x base) — small round numbers
// chosen so every deadline is an exact assertion, not just "greater than before". Each call
// below lands at or before the previous deadline (a zero or negative gap), which is always
// "no more than one base cooldown after the previous deadline" — i.e. every call continues the
// same streak, matching a credential that keeps reporting exhausted with no lull.
@Test
void anInvalidBackoffMultiplierIsRejected() {
assertThrows(IllegalArgumentException.class,
() -> new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 0.5, TimeUnit.HOURS.toNanos(6)));
}
@Test
void aCeilingBelowTheBaseCooldownIsRejected() {
assertThrows(IllegalArgumentException.class,
() -> new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 2.0, TimeUnit.MINUTES.toNanos(10)));
}
@Test
void repeatedExhaustionEscalatesTheCooldownByExactAmounts() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
q.quarantine("shared-openai"); // 1st: base cooldown
assertEquals(OptionalLong.of(600L), q.remainingSeconds("shared-openai"));
now.set(TimeUnit.SECONDS.toNanos(100)); // still inside the 1st quarantine (deadline 600s)
q.quarantine("shared-openai"); // 2nd: 600 * 2^1 = 1200
assertEquals(OptionalLong.of(1200L), q.remainingSeconds("shared-openai"),
"a second consecutive exhaustion must double the cooldown, not just increase it");
now.set(TimeUnit.SECONDS.toNanos(1300)); // exactly the 2nd deadline (100 + 1200)
q.quarantine("shared-openai"); // 3rd: 600 * 2^2 = 2400 (exactly at the ceiling)
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"));
}
@Test
void escalationStopsAtTheCeiling() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
q.quarantine("shared-openai"); // 1st: 600
now.set(TimeUnit.SECONDS.toNanos(600));
q.quarantine("shared-openai"); // 2nd: 1200, deadline 1800
now.set(TimeUnit.SECONDS.toNanos(1800));
q.quarantine("shared-openai"); // 3rd: 600 * 4 = 2400, at the ceiling, deadline 4200
now.set(TimeUnit.SECONDS.toNanos(4200));
q.quarantine("shared-openai"); // 4th: 600 * 8 = 4800 uncapped, must stay capped at 2400
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"),
"the cooldown must never exceed the configured ceiling, however long the streak gets");
now.set(TimeUnit.SECONDS.toNanos(6600)); // 4th deadline
q.quarantine("shared-openai"); // 5th: still capped
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"),
"pushing well past the ceiling must not budge it");
}
@Test
void aQuietGapLongerThanTheBaseCooldownResetsToTheBaseCooldown() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
q.quarantine("shared-openai"); // 1st: 600, deadline 600
now.set(TimeUnit.SECONDS.toNanos(600));
q.quarantine("shared-openai"); // 2nd: 1200, deadline 1800
now.set(TimeUnit.SECONDS.toNanos(1800));
q.quarantine("shared-openai"); // 3rd: 2400, deadline 4200
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"));
// Quiet for well over one base cooldown (600s) past the 3rd deadline (4200s).
now.set(TimeUnit.SECONDS.toNanos(20_000));
q.quarantine("shared-openai"); // treated as a fresh occurrence
assertEquals(OptionalLong.of(600L), q.remainingSeconds("shared-openai"),
"a long quiet gap must reset the streak back to the base cooldown");
}
@Test
void escalatingOneCredentialDoesNotSlowAnother() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
q.quarantine("shared-openai"); // 1st: 600
now.set(TimeUnit.SECONDS.toNanos(600));
q.quarantine("shared-openai"); // 2nd: 1200
now.set(TimeUnit.SECONDS.toNanos(1800));
q.quarantine("shared-openai"); // 3rd: 2400 — three-in-a-row streak on this credential only
q.quarantine("another-credential"); // its first and only exhaustion
assertEquals(OptionalLong.of(600L), q.remainingSeconds("another-credential"),
"an unrelated credential's cooldown must stay at the base rate, unaffected by a sibling's streak");
}
@Test
void withEscalationDefaultsToDoublingCappedAtTwelveTimesTheBase() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = BackendQuarantine.withEscalation(now::get, TimeUnit.MINUTES.toNanos(30));
q.quarantine("shared-openai");
assertEquals(OptionalLong.of(1800L), q.remainingSeconds("shared-openai"),
"the first occurrence must still use the base cooldown");
now.set(TimeUnit.MINUTES.toNanos(30));
q.quarantine("shared-openai");
assertEquals(OptionalLong.of(3600L), q.remainingSeconds("shared-openai"),
"the default multiplier must be 2.0");
}
// --- fleetd #466 scope item 2: status() reports repeatCount beside remainingSeconds ------------
@Test
void statusReportsAttemptOneForAFirstOccurrenceNeverAbsentOrZero() {
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 2.0,
TimeUnit.HOURS.toNanos(6));
q.quarantine("shared-openai");
BackendQuarantine.Status status = q.status("shared-openai").orElseThrow();
assertEquals(1800L, status.remainingSeconds());
assertEquals(1, status.repeatCount(),
"a first-ever occurrence must report attempt 1, not 0 or absent -- 1 means unambiguously "
+ "'the first time', where 0 would be indistinguishable from a bug that forgot to count");
}
@Test
void statusIsAbsentWhenTheCredentialIsNotQuarantined() {
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 2.0,
TimeUnit.HOURS.toNanos(6));
assertTrue(q.status("shared-openai").isEmpty());
}
/**
* The acceptance criterion's strong form: the reported {@code repeatCount} must match the exact
* step the cooldown's own growth implies, read off {@link BackendQuarantine#status}'s single
* call -- not two independent reads that happen to agree in this easy case.
*/
@Test
void statusReportsTheGrowingAttemptCountAlongsideTheEscalatingCooldown() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
q.quarantine("shared-openai"); // 1st: 600s, attempt 1
BackendQuarantine.Status first = q.status("shared-openai").orElseThrow();
assertEquals(600L, first.remainingSeconds());
assertEquals(1, first.repeatCount());
now.set(TimeUnit.SECONDS.toNanos(600));
q.quarantine("shared-openai"); // 2nd: 1200s, attempt 2
BackendQuarantine.Status second = q.status("shared-openai").orElseThrow();
assertEquals(1200L, second.remainingSeconds());
assertEquals(2, second.repeatCount());
now.set(TimeUnit.SECONDS.toNanos(1800));
q.quarantine("shared-openai"); // 3rd: 2400s (at the ceiling), attempt 3
BackendQuarantine.Status third = q.status("shared-openai").orElseThrow();
assertEquals(2400L, third.remainingSeconds());
assertEquals(3, third.repeatCount());
}
/**
* Once the cooldown hits its ceiling, every further consecutive exhaustion reports the SAME
* {@code remainingSeconds} -- so a {@code repeatCount} re-derived from the cooldown value (e.g.
* inverting {@code cooldownNanos * multiplier^(n-1)}) could not tell attempt 4 from attempt 9;
* only the stored counter can. This is the scenario that makes "read the count off a second,
* independent computation" provably wrong rather than just risky.
*/
@Test
void repeatCountKeepsGrowingPastTheCeilingEvenThoughTheCooldownStaysFlat() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
long t = 0L;
for (int attempt = 1; attempt <= 5; attempt++) {
now.set(t);
q.quarantine("shared-openai");
BackendQuarantine.Status status = q.status("shared-openai").orElseThrow();
assertEquals(attempt, status.repeatCount(),
"attempt " + attempt + " must be reported as exactly " + attempt
+ ", not collapsed to whatever attempt first reached the ceiling");
if (attempt >= 3) {
assertEquals(2400L, status.remainingSeconds(), "attempt " + attempt + " must be capped");
}
t += status.remainingSeconds(); // land exactly on the next deadline: still the same streak
}
}
@Test
void aQuietGapResetsTheReportedAttemptCountToOneToo() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
q.quarantine("shared-openai");
now.set(TimeUnit.SECONDS.toNanos(600));
q.quarantine("shared-openai");
now.set(TimeUnit.SECONDS.toNanos(1800));
q.quarantine("shared-openai"); // attempt 3
assertEquals(3, q.status("shared-openai").orElseThrow().repeatCount());
now.set(TimeUnit.SECONDS.toNanos(20_000)); // long quiet gap
q.quarantine("shared-openai");
assertEquals(1, q.status("shared-openai").orElseThrow().repeatCount(),
"a reset streak must report attempt 1 again, matching the reset base cooldown");
}
@Test
void escalatingOneCredentialsAttemptCountDoesNotAffectAnother() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
q.quarantine("shared-openai");
now.set(TimeUnit.SECONDS.toNanos(600));
q.quarantine("shared-openai");
now.set(TimeUnit.SECONDS.toNanos(1800));
q.quarantine("shared-openai"); // 3-in-a-row streak on this credential only
q.quarantine("another-credential");
assertEquals(1, q.status("another-credential").orElseThrow().repeatCount(),
"an unrelated credential's attempt count must stay at 1, unaffected by a sibling's streak");
}
@Test
void noneReportsNoStatusForAnything() {
BackendQuarantine q = BackendQuarantine.none();
q.quarantine("shared-openai"); // no-op on none(), same as every other mutator
assertTrue(q.status("shared-openai").isEmpty(),
"none() quarantines nothing, so it must report no status at all -- never a fabricated "
+ "attempt count for a credential that was never actually quarantined");
}
/** The flat (non-escalating) two-argument constructor must still report a real, growing count. */
@Test
void aFlatTwoArgumentInstanceStillReportsAGrowingAttemptCountEvenThoughTheCooldownStaysFlat() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600));
q.quarantine("shared-openai");
BackendQuarantine.Status first = q.status("shared-openai").orElseThrow();
assertEquals(600L, first.remainingSeconds());
assertEquals(1, first.repeatCount());
now.set(TimeUnit.SECONDS.toNanos(100)); // still inside the 1st quarantine
q.quarantine("shared-openai");
BackendQuarantine.Status second = q.status("shared-openai").orElseThrow();
assertEquals(600L, second.remainingSeconds(),
"the flat constructor's cooldown must stay exactly the base length regardless of the streak");
assertEquals(2, second.repeatCount(),
"the flat constructor still counts the real streak -- it just does not scale the cooldown by it");
}
}