Compare commits
139 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 26f198675c | |||
| 72f46d7c0c | |||
| 640f4d5f23 | |||
| 6edeb70bc4 | |||
| bc49d87cb8 | |||
| 14b169c410 | |||
| 2b52324d9a | |||
| b6147a39f6 | |||
| dbfc34cb6d | |||
| a20cb96730 | |||
| cda1a6a917 | |||
| 608e4496be | |||
| fa61dc587c | |||
| 526e3b7459 | |||
| b6b006c651 | |||
| 63eec8a0da | |||
| cbe872b538 | |||
| 7d9a807243 | |||
| 8915e40c7d | |||
| 3f7bc3815e | |||
| 6cb31a10e4 | |||
| 8368a274a0 | |||
| 17127efb88 | |||
| 203f034528 | |||
| 9ee16f5b85 | |||
| 388ef5a3c3 | |||
| be6c45ff78 | |||
| e99cb70a8b | |||
| 987ccef4c7 | |||
| 076cc43f7b | |||
| bad47a8444 | |||
| 955b9ea013 | |||
| 89cb8ff79b | |||
| fa62e9906d | |||
| d7390ccd37 | |||
| aa517ae0ec | |||
| 60496831c2 | |||
| 9a992d0f70 | |||
| 3762aca307 | |||
| 4e27bde2d7 | |||
| b85d9b0e46 | |||
| 856dfc6318 | |||
| 81c1d8e91c | |||
| d345b14e43 | |||
| 9640deeffc | |||
| ad3d81941f | |||
| 8b986a52e0 | |||
| f336bcef39 | |||
| 3ed7bfca67 | |||
| f5c6a0e4fc | |||
| a7aee5b982 | |||
| bf895616a5 | |||
| b874afb0af | |||
| 6eb34a654f | |||
| 7084d99b89 | |||
| ae7845c375 | |||
| 61115f6f61 | |||
| 4b9ebda1b3 | |||
| 1e68d7ee39 | |||
| 6cccd458d4 | |||
| 42820fbe75 | |||
| d91ff886da | |||
| 386e760a5c | |||
| 2e349139e9 | |||
| a639969a9a | |||
| 17c3a69c57 | |||
| 49a5875586 | |||
| 634d33b50b | |||
| 1db79bcaa9 | |||
| 4507bc5a70 | |||
| d7239ed23b | |||
| dfeb9340b4 | |||
| 1513d4f260 | |||
| 4ca7d72303 | |||
| c0545d003d | |||
| 275ac0d251 | |||
| c1e06c9e12 | |||
| 204da67d66 | |||
| b091c51eee | |||
| 84d631b030 | |||
| 3f8c38fc54 | |||
| a4dbc8f8b7 | |||
| ed2fd6646a | |||
| 5441a2b321 | |||
| 384867dfa3 | |||
| f40c19ecf0 | |||
| 034e17bb32 | |||
| 68b428c484 | |||
| e20ccab1eb | |||
| d83821bbce | |||
| ba2f4d16f8 | |||
| db4c98ac60 | |||
| 8f80d267a0 | |||
| 738d34a609 | |||
| a46e4058ac | |||
| f4f5f3106e | |||
| 8d79d229ff | |||
| cb64bc8157 | |||
| ed28b51f12 | |||
| 4f9aba40e7 | |||
| dab697fae0 | |||
| bfac14108f | |||
| 7a3b2bb7ee | |||
| f188947750 | |||
| f606fccf7f | |||
| 735b6af976 | |||
| 8b4320ed24 | |||
| 351ee1ea6d | |||
| d4a51c6274 | |||
| 26f380a00b | |||
| b8182c96c2 | |||
| 3da44eed63 | |||
| 93a9ed3f83 | |||
| 0b032f5a1a | |||
| fad99c4c5e | |||
| 87871eaefb | |||
| a476a14f1c | |||
| 5eb4267a4a | |||
| 7611b69667 | |||
| cc302fe4af | |||
| 4a8a780274 | |||
| 343ce0f4c0 | |||
| cec3e191d4 | |||
| 1850a5f324 | |||
| e4eb3dbed4 | |||
| 57cd96f5e6 | |||
| 202e37e3b3 | |||
| 90253f832d | |||
| f1640f5dcc | |||
| c7903c1efe | |||
| 7d711942fe | |||
| bec87f987c | |||
| 7f8a8829f9 | |||
| af9589783e | |||
| 190436c9cf | |||
| 8ea5c2bb1f | |||
| 8335b12562 | |||
| d25c863118 | |||
| be07ed2033 |
@@ -0,0 +1,23 @@
|
||||
---
|
||||
name: hunter
|
||||
description: Sweep one assigned scope for defects and report ranked findings without changes.
|
||||
---
|
||||
|
||||
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
|
||||
|
||||
You sweep the assigned package or scope for real defects. Read the full assigned scope before you
|
||||
judge it. Report several ranked findings when the evidence supports them. Change nothing: do not
|
||||
edit code, commit, push, or open a pull request.
|
||||
|
||||
You may run the build or tests to check a finding. Read the complete output and report the real
|
||||
result. Do not hide failures with a pipe. State only checks you actually ran. The primary's IDE
|
||||
tools are not yours. A mounted forge tool may use a blocked credential and fail by design.
|
||||
|
||||
Do only the assigned scope. Note anything outside it in one line and do not investigate it further.
|
||||
Use `fleet_ask{question}` only when a decision belongs to the lead, such as an unclear requirement
|
||||
or two defensible fixes. Do not ask about something you can decide by reading more code.
|
||||
|
||||
Your handoff must name the files you read, each ranked finding or `NO FINDINGS`, the checks you ran,
|
||||
and any caveat for review.
|
||||
|
||||
The launcher provides the required bridge reply instructions for every member.
|
||||
@@ -167,19 +167,59 @@ fails.
|
||||
- **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved
|
||||
from your connection, so you can only ever roll yourself.
|
||||
- **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you
|
||||
are confident. Ask, wait for the answer, then pass what they said. `requireOperatorConfirm`
|
||||
defaults to `true` and this is the only thing standing between a judgement call and a wiped
|
||||
session.
|
||||
are confident. Ask, wait for the answer, then pass what they said.
|
||||
|
||||
- **Whether you must ask at all depends on `leadRollover.requireOperatorConfirm`. Check it; do not
|
||||
assume.** The default is `true` (`FleetConfig.java:1426`), and then `confirm` refuses unless you
|
||||
also pass `operatorConfirmed: true`. **This host set it to `false` on 2026-09-22**, on the
|
||||
operator's explicit grant, because they do not want to approve routine context rolls. Where it is
|
||||
`false`, the three handover-file checks are the whole gate: the file must exist, be fresher than
|
||||
`maxDocAgeSeconds`, and have been modified after the open request.
|
||||
|
||||
Read the live value rather than trusting this line:
|
||||
|
||||
```bash
|
||||
grep -A1 'requireOperatorConfirm' fleetd/fleetd.yaml
|
||||
```
|
||||
|
||||
No match means the key is unset, so the default `true` applies and you must ask. The key is
|
||||
**deferred, not hot** — it is read once at boot, so an edit does nothing until the daemon is
|
||||
redeployed.
|
||||
|
||||
**Until fleetd #621 merges, the nudge text will tell you to ask the operator even where the
|
||||
daemon no longer requires it.** `LeadHeartbeatLoop.contextNotice()` hardcodes "ask the operator"
|
||||
and takes no config, so it cannot know. Trust the config value over the nudge text. Once #621 is
|
||||
merged and deployed, the nudge matches the config and this warning can be deleted.
|
||||
|
||||
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
|
||||
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
|
||||
- **The bootstrap prompt has never yet landed, and the fix is unproven (fleetd #489).** The first
|
||||
real rollover, on 2026-09-12, joined `/clear` and the bootstrap text into one line and Claude Code
|
||||
refused it as `Unknown command: /clearFresh`. The pane was never cleared and no context was lost,
|
||||
so the failure was safe — the roll simply did nothing. PR #490 fixed the cause and is deployed,
|
||||
but no roll has bootstrapped a fresh session end to end yet. **Assume it may still fail, and tell
|
||||
the operator so before you confirm.** The recovery is the same either way: the file is already
|
||||
written, so the operator starts a session and points it at the file. That is why you write the
|
||||
file before you confirm, and never the other way round.
|
||||
- **The bootstrap prompt works end to end. Measured 2026-09-22.** This used to say the fix was
|
||||
unproven (fleetd #489) and told you to expect a failure. That is no longer true. The daemon log
|
||||
now holds four `lead-rollover: rolled` lines, and three of them ran on 2026-09-22 at 10:01:43,
|
||||
10:38:28 and 11:15:47. Each one cleared the old lead and started a fresh session against the
|
||||
handover file, with the configured `bootstrapText` arriving as its first message. No context was
|
||||
lost. The old `Unknown command: /clearFresh` failure from 2026-09-12 does not appear in the log
|
||||
at all. Re-measure both numbers with:
|
||||
|
||||
```bash
|
||||
grep -c "lead-rollover: rolled" fleetd/fleetd.out # successful rolls
|
||||
grep -c "lead-rollover:" fleetd/fleetd.out # positive control: must be larger
|
||||
grep -c "Unknown command" fleetd/fleetd.out # the old failure: expect 0
|
||||
```
|
||||
|
||||
Run the control line too. A broken pattern returns a clean `0` that reads exactly like good news.
|
||||
If the first number stops growing across rolls, or `Unknown command` returns anything above 0,
|
||||
the bootstrap has regressed and this paragraph is stale again.
|
||||
|
||||
**You still write the file before you confirm, and never the other way round.** That order is not
|
||||
about the bootstrap being unreliable. It is what the daemon checks: the handover file must have
|
||||
been modified *after* the open request, or `confirm` refuses it as stale.
|
||||
|
||||
- **One warning in the log is normal and is not a failure.** Every one of the three rolls above also
|
||||
logged `/clear on term_… was never observed as WORKING after 8 consecutive IDLE/DONE polls —
|
||||
releasing rather than wedging the roll`. The daemon could not see the pane go WORKING after
|
||||
`/clear`, so it released instead of hanging. The roll then succeeded anyway. That is the safe
|
||||
branch behaving correctly. Do not report it as a broken roll.
|
||||
|
||||
## Writing style
|
||||
|
||||
|
||||
@@ -39,6 +39,10 @@ jobs:
|
||||
# that runs does so against the fake UDS herdr and fake ccs/claude stubs.
|
||||
run: mvn -B clean install
|
||||
|
||||
- name: javadoc reference lint
|
||||
working-directory: fleetd
|
||||
run: mvn -B -DskipTests javadoc:javadoc -Ddoclint=reference
|
||||
|
||||
# Deliberately NOT actions/upload-artifact: this Gitea instance presents as GHES, and
|
||||
# @actions/artifact v2+ (i.e. upload-artifact@v4) refuses to run there —
|
||||
# "GHESNotSupportedError ... not currently supported on GHES", which red-Xes an otherwise
|
||||
@@ -54,6 +58,23 @@ jobs:
|
||||
done
|
||||
exit 0
|
||||
|
||||
# fleetd #550 — nothing ran scripts/test-redeploy-fleetd.sh in CI before this, on any platform,
|
||||
# so it had run only on macOS by hand and two Linux-only bugs (this issue's items 1 and 2)
|
||||
# survived undetected: shasum is a macOS-only tool (it ships with Perl; GNU coreutils, i.e. every
|
||||
# mainstream Linux distro including this runner's ubuntu-latest, does not have it and ships
|
||||
# sha256sum instead). The gate here is the step's own exit code, nothing else: a `run:` step in
|
||||
# Gitea/GitHub Actions already fails the job on a non-zero exit with no extra scripting needed,
|
||||
# so this deliberately does NOT grep the output for a `FAIL:` count. That is the #550 item-2
|
||||
# lesson one level up — a suite that dies before it runs a single test prints zero FAIL lines,
|
||||
# which is exactly what a clean pass also prints, so counting FAIL lines can never be the gate.
|
||||
shell-tests:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- name: redeploy-fleetd.sh shell suite
|
||||
run: bash scripts/test-redeploy-fleetd.sh
|
||||
|
||||
# CB-521 — actually run the AMQP contract test in CI, against a REAL broker. The broker is a
|
||||
# RabbitMQ SERVICE CONTAINER, not Testcontainers-with-Docker: the runner image has no Docker, so
|
||||
# AmqpReplyInboxContractTest reads AMQP_URI (set below to the service's network alias) and binds
|
||||
|
||||
@@ -95,9 +95,21 @@ below are the procedure — run them in order, every task, not only the big ones
|
||||
`fleet_status`, never by reading its terminal; it also reports an open question and the `turnId`
|
||||
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
|
||||
brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
|
||||
**A correction cannot reach a busy member.** A `fleet_send` to a working member is *accepted* and
|
||||
returns a ticket, and is then never delivered — measured here three times in one session, and the
|
||||
member was released still executing a brief that had been retracted twice. The receipt is true and
|
||||
it is a fact about the *mailbox*; what you needed was a fact about the *pane*. **A push delivery
|
||||
needs the recipient free at send time; a pull channel needs only that they look before acting.** So
|
||||
put every correction on the **ticket**, which they can read whenever they look, and send as the
|
||||
notification. That obliges you, not them: **all corrections go to the ticket, and the brief is
|
||||
write-once.** The member cannot check which source is newer — it just always prefers the ticket —
|
||||
so the day you revise a brief in place instead of commenting, it obeys your rule and does the wrong
|
||||
thing. A *first* brief for a unit not yet running is not a correction, and may be the whole spec.
|
||||
6. **Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
|
||||
forge tools it appears to have hold a blocked credential and fail, and a piped command
|
||||
(`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a fact.
|
||||
forge MCP server it appears to have holds a blocked credential and fails every call, and a piped
|
||||
command (`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a
|
||||
fact. Its injected repo-scoped `GITEA_TOKEN` is a different credential and does work, so a worker
|
||||
reporting that it opened its own PR is reporting something it really can do.
|
||||
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
|
||||
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
|
||||
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
|
||||
@@ -127,12 +139,23 @@ prefer `wait:false` + `fleet_poll` for anything non-trivial: a blocking `fleet_s
|
||||
**Delegating does not delegate responsibility.** Workers open PRs; you are the gate. Never delegate
|
||||
the merge — and merging on a reviewer's word is delegating it by proxy.
|
||||
|
||||
**When a decision blocks you, consult architects — not the operator.** Spawn one or more architect
|
||||
members, give them the question and the evidence you have, and act on what they agree. They are
|
||||
authorized to settle it, not only to advise. Architects first form independent positions, then
|
||||
compare them. If they still disagree after that comparison, they return both positions and their
|
||||
checked evidence; the lead decides. Go to the operator only for an action the fleet has no
|
||||
authority to take, such as spending money, granting access, or making a promise to someone else.
|
||||
**Then write the decision on the ticket.** Taking the operator out of the loop also removes the
|
||||
signal they used to get, because that signal was the block itself — work stopped, so they found
|
||||
out. A ticket comment replaces it, and it reaches them whether or not they are at a terminal when
|
||||
you decide.
|
||||
|
||||
| Intent | Tool |
|
||||
|---|---|
|
||||
| Confirm your own role | `fleet_whoami` |
|
||||
| See backends available | `fleet_profiles` |
|
||||
| Start a member | `fleet_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` → `sessionId` + `paneId` |
|
||||
| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) · one peer's state: `fleet_status{sessionId}` |
|
||||
| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) + `loopHealth` (`RUNNING`, `STALLED`, or `STOPPED` for `statusPoller` and `sessionReaper`) · one peer's state: `fleet_status{sessionId}` |
|
||||
| Delegate (blocking) | `fleet_send{sessionId, content}` |
|
||||
| Delegate (long task) | `fleet_send{sessionId, content, wait:false}` → ticket → `fleet_poll{ticket}` |
|
||||
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
|
||||
@@ -189,20 +212,26 @@ simply complies has thrown away the reason there are two of you.
|
||||
3. **`fleet_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
|
||||
defensible fixes, "bug or intended?"). It blocks and you resume the *same* turn with the answer.
|
||||
Don't ask what you could decide yourself.
|
||||
4. **End the turn with exactly one `fleet_reply{content}`**, carrying your complete answer. This is
|
||||
4. **Re-read the ticket before you act on anything you were told earlier**, and again before you
|
||||
commit. A message reaches you only while you are free to receive it; the ticket is there whenever
|
||||
you look. **If a ticket comment contradicts your brief, the ticket comment is newer and it wins.**
|
||||
5. **End the turn with exactly one `fleet_reply{content}`**, carrying your complete answer. This is
|
||||
the whole handoff. No `fleet_reply` ⇒ the sender gets nothing and the exchange stalls.
|
||||
Do **not** lean on the completion fallback to carry your answer for you: when you end a turn
|
||||
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
|
||||
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
|
||||
lead with its end cut off.
|
||||
5. **Report honestly.** State only what you actually ran and its real output, including failures,
|
||||
6. **Report honestly.** State only what you actually ran and its real output, including failures,
|
||||
and never claim the result of a check you had no way to run. **Measure your own tools; do not
|
||||
assume them.** What you mount depends on your backend: an opencode member gets the bridge and
|
||||
nothing else, while a Claude Code member also inherits the operator's user-scope MCP servers,
|
||||
which the bridge never chose for you. Two rules follow. The primary's IDE tooling is still not
|
||||
yours, whatever you see. And **a mounted tool is not a working tool** — the forge server you may
|
||||
find there holds a deliberately blocked credential and fails every call, by design.
|
||||
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
|
||||
yours, whatever you see. And **a mounted tool is not a working tool** — the forge MCP server you
|
||||
may find there holds a deliberately blocked credential and fails every call, by design. That is
|
||||
not your only forge route, and the two must not be confused: the repo-scoped `GITEA_TOKEN` the
|
||||
daemon injects into your environment does work, and using it to open your own PR is part of the
|
||||
job. A blocked MCP tool is never a reason to skip that step.
|
||||
7. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
|
||||
project marks as not-yours-to-commit.
|
||||
|
||||
### Where each rule lives (don't duplicate — extend the right layer)
|
||||
@@ -257,6 +286,8 @@ must obey belongs in the charter, not here.
|
||||
it for a multi-finding sweep hands the worker two contradictory output contracts. That has
|
||||
already cost three workers' turns: each wrote a good report to its terminal and ended the turn
|
||||
with no `fleet_reply`, and the scrape returned the tail of the brief instead.
|
||||
Spawn `implementer` with role `dev`, `reviewer` with role `reviewer`, and `hunter` with role
|
||||
`hunter`.
|
||||
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
|
||||
`port-to-opencode` (make an OpenCode session a participant in this workspace),
|
||||
`fleets-status` (report every fleet that shares one LavinMQ instance),
|
||||
|
||||
@@ -5,6 +5,21 @@
|
||||
here took a revert and two upstream fixes — see §7.1, which is the useful part of this document. One
|
||||
risk is **accepted rather than solved**: a stream cut by any mid-response timer arrives as HTTP 200
|
||||
with no terminator, and our third-party members cannot detect it (§7.2).
|
||||
|
||||
> **Superseded in part — 2026-09-13.** Two claims on this page are no longer true of the live fleet.
|
||||
> I measured both on this host today.
|
||||
>
|
||||
> 1. **The model is named `acoder` now, not `deepseek-v4-flash`.** `acoder` is a stable alias, and
|
||||
> the model behind it changed on 2026-08-28: it is Qwen3.8-27B, not DeepSeek. The old name is
|
||||
> still served, so nothing broke — the gateway answers it and reports `"model": "acoder"` in the
|
||||
> reply, which is how you can see for yourself that it is an alias. `fleetd.yaml` moved to
|
||||
> `acoder` on 2026-09-13. Do not guess behaviour from the name; ask the gateway's own manifest,
|
||||
> `GET https://llm.ltms.dev/v1/deployment`, and read its `generation` field.
|
||||
> 2. **`local` sits at `weight: 0`, not 100.** Only `gx` is auto-selected today.
|
||||
>
|
||||
> §2 and §3 below are the plan as written in August. They are the record of the migration, so they
|
||||
> stay as they are. If this note stops matching `fleetd.yaml`, re-measure and rewrite the note.
|
||||
|
||||
· **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway)
|
||||
· **Upstream issue:** [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31)
|
||||
|
||||
@@ -380,7 +395,10 @@ one turn; this one costs the whole task and is indistinguishable from a slow wor
|
||||
|
||||
- Token accepted on both surfaces. **Unauthenticated → 401**, so the Caddy proxy really does gate —
|
||||
the wiki's "SecurityPolicy fails open" warning is about the gateway itself, not the edge.
|
||||
- `/v1/models` returns exactly `["deepseek-v4-flash"]`, so trap 3 is clear.
|
||||
- `/v1/models` returned exactly `["deepseek-v4-flash"]` **on 2026-08-15**, so trap 3 was clear
|
||||
then. It returns 6 ids now — `acoder`, `qwen3.8-27b-nvfp4`, `deepseek-v4-flash` and three
|
||||
embedding names — measured on this host 2026-09-13. The exact-name rule still holds; the
|
||||
one-item list does not.
|
||||
- **Reasoning survives both surfaces** — see §3b above.
|
||||
- The launcher's generated opencode provider block is correct, carrying a real 48-character `llmk-`
|
||||
key rather than the `fleetd-local-noauth` placeholder.
|
||||
|
||||
+27
-12
@@ -77,16 +77,25 @@ bind:
|
||||
# a lead turn nobody asked for), so upgrading the daemon must never switch it on for you. Absent
|
||||
# block = feature off, exactly as before.
|
||||
#
|
||||
# Three knobs, each with a default that errs on the side of not burning context:
|
||||
# Four knobs, each with a default that errs on the side of not burning context:
|
||||
# idleAfterSeconds: 300 # how long the lead must stay idle before the FIRST nudge (default 300 —
|
||||
# # absorbs normal post-turn pauses; re-prompting every pause burns context)
|
||||
# backoffMs: 60000 # re-check cadence / spacing between nudges past the quiet period (default 60000)
|
||||
# quietNudgeCap: 3 # cap on consecutive nudges that find NOTHING pending, then it stops
|
||||
# # until real state appears (default 3 — never nag an empty fleet forever)
|
||||
# contextHighNudge: false # fleetd #609 — when true, an idle lead whose OWN Claude Code context
|
||||
# # reads HIGH (see LeadContextGauge; fleet_list's context row) gets a text
|
||||
# # notice appended to its nudge telling it to consider fleet_handover. Text
|
||||
# # only — it never rolls a pane by itself, and only the operator can approve
|
||||
# # a roll. Fires once per HIGH stretch (a later OK reading re-arms it), and
|
||||
# # never spends the quietNudgeCap budget. Default false/absent = off, same
|
||||
# # as every other knob here — an upgraded daemon must not start telling
|
||||
# # leads to hand over on its own.
|
||||
# leadHeartbeat:
|
||||
# idleAfterSeconds: 300
|
||||
# backoffMs: 60000
|
||||
# quietNudgeCap: 3
|
||||
# contextHighNudge: false
|
||||
|
||||
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
|
||||
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
|
||||
@@ -200,13 +209,13 @@ herdrSocket: ~/.config/herdr/herdr.sock
|
||||
# e.g. `env DISPLAY=:10.0 idea {dir}`. Best-effort: a failure is logged, never
|
||||
# fails the spawn. Omit to open the member's module by hand. There is no close
|
||||
# half yet — an opened module stays open until the operator closes it.
|
||||
# autoCompactWindow → opt-in, default off. A bounded token window that forces a spawned member to
|
||||
# compact its context instead of running on the backend's own default and dying
|
||||
# mid-turn (losing its fleet_reply — the whole point of the turn — with it).
|
||||
# autoCompactWindow → opt-in, default off. A bounded token window that forces a launched Claude Code
|
||||
# lead or member to compact its context instead of running on the backend's own
|
||||
# default. A member that runs out of context can die mid-turn and lose its fleet_reply.
|
||||
# Validated at config load to [100000, 1000000] — the band Claude Code's own
|
||||
# --autocompact flag accepts.
|
||||
# CROSS-BACKEND SEMANTICS DIFFER: on claude-code this is a launch-time
|
||||
# `--autocompact <tokens>` flag — the member compacts AT this window. opencode
|
||||
# `--autocompact <tokens>` flag — the Claude Code session compacts AT this window. opencode
|
||||
# has no equivalent flag (it only forces `compaction.auto: true`, unconditionally,
|
||||
# already), so this is instead applied as the model's `limit.context` in the
|
||||
# generated opencode.json — the member compacts WITHIN this window, not exactly
|
||||
@@ -406,7 +415,7 @@ profiles:
|
||||
# ideMcpUrl: http://127.0.0.1:29170/index-mcp/streamable-http # opt-in (CB-634): IDE code intelligence, pinned to the worktree
|
||||
# ideProjectDir: fleetd # CB-634: module dir the IDE opens + the overlay pins (this repo's pom is in fleetd/)
|
||||
# ideOpenCommand: env DISPLAY=:10.0 idea {dir} # CB-634 auto-open: opens {dir} in the IDE at spawn; omit to open by hand
|
||||
# autoCompactWindow: 250000 # opt-in: bound member context; claude-code compacts AT this, opencode within it (model limit.context)
|
||||
# autoCompactWindow: 250000 # opt-in: bound Claude Code lead/member context; claude-code compacts AT this, opencode within it (model limit.context)
|
||||
gx11: # a second backend, so `placement: weighted` has a choice
|
||||
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
|
||||
placement: tab
|
||||
@@ -560,7 +569,7 @@ placement: weighted
|
||||
# older keys: `leaders:`, `members:`, `leadScan:` and `defaultProfile:`.
|
||||
#
|
||||
# A member is anything a lead spawns, and every member has two INDEPENDENT attributes:
|
||||
# role — which contract: architect, dev or reviewer. It picks the launch charter, the role
|
||||
# role — which contract: architect, dev, hunter or reviewer. It picks the launch charter, the role
|
||||
# file, the playbook skill and the authz row.
|
||||
# profile — which backend: one of the `profiles:` keys above (model, CLI adapter, cost).
|
||||
# They vary on their own. A reviewer may run on the same profile as the dev whose diff it reads,
|
||||
@@ -572,13 +581,13 @@ placement: weighted
|
||||
#
|
||||
# Each pool lists the profiles that role MAY run on — these are pools, not identities. That is also
|
||||
# what replaced `defaultProfile:`: an unqualified spawn names a role, and that role's pool supplies
|
||||
# the candidates, in definition order. A dev and a reviewer staying anonymous is exactly compatible
|
||||
# with being listed here; the entry key just names the entry.
|
||||
# the candidates, in definition order. A dev, hunter and reviewer staying anonymous is exactly
|
||||
# compatible with being listed here; the entry key just names the entry.
|
||||
fleet:
|
||||
# Optional launch-charter text, keyed only by the singular role wire names: architect, dev,
|
||||
# reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put secrets
|
||||
# here: a later launch step writes this text to a world-readable temp file, and ${ENV} interpolation
|
||||
# is deliberately not supported.
|
||||
# hunter, reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put
|
||||
# secrets here: a later launch step writes this text to a world-readable temp file, and ${ENV}
|
||||
# interpolation is deliberately not supported.
|
||||
charters:
|
||||
architect: |-
|
||||
You are an architect in this fleet. You refine work before anyone builds it:
|
||||
@@ -589,6 +598,9 @@ fleet:
|
||||
dev: |-
|
||||
You implement the one unit you were given, and nothing else. You test it,
|
||||
commit it, and open your own pull request. You never merge.
|
||||
hunter: |-
|
||||
You sweep the assigned scope for real defects. You may run the build or tests
|
||||
to check a finding. You change nothing, and report several ranked findings.
|
||||
reviewer: |-
|
||||
You review the diff you were given. You report bugs, risks and missing tests.
|
||||
You do not change code.
|
||||
@@ -666,6 +678,9 @@ fleet:
|
||||
developers:
|
||||
gx10:
|
||||
profile: gx10
|
||||
# hunters:
|
||||
# gx10:
|
||||
# profile: gx10 # a hunt may run checks, but never changes code
|
||||
# reviewers:
|
||||
# gx10:
|
||||
# profile: gx10 # the same backend may serve two roles; that is the point
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
|
||||
/**
|
||||
* fleetd #612 Unit A: everything {@code Fleetd.main} has ready once config is loaded, reported
|
||||
* and validated — the exact point {@code main} used to keep going straight into socket and broker
|
||||
* work. Building this (and handing it to {@link FleetdAssembly#assembleAndStart}) is the seam a
|
||||
* test now has to drive the real boot composition without being {@code main} itself.
|
||||
*
|
||||
* @param cfg the boot-time {@link FleetConfig} snapshot every one-time wiring decision reads —
|
||||
* see {@code Fleetd.main}'s own comment on why this must never be swapped for a live
|
||||
* reference once loaded
|
||||
* @param config the live {@link ConfigRef} the hot-reloadable paths read per use
|
||||
* @param guard the same {@link SubscriptionGuard} {@code Fleetd.main} already used to assert the
|
||||
* launching environment is clean, reused rather than rebuilt so the assembled
|
||||
* launchers see the identical instance {@code main} already validated with
|
||||
*/
|
||||
record AssemblyInputs(FleetConfig cfg, ConfigRef config, SubscriptionGuard guard) {
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,536 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.auth.CallerResolver;
|
||||
import dev.ltms.fleet.auth.MemberRegistry;
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.ConfigWatcher;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.health.FleetHealthMonitor;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.HerdrRouter;
|
||||
import dev.ltms.fleet.herdr.LeadTabScanner;
|
||||
import dev.ltms.fleet.herdr.PaneLocator;
|
||||
import dev.ltms.fleet.herdr.UnixSocketHerdrClient;
|
||||
import dev.ltms.fleet.inject.BackendErrorPatternLookup;
|
||||
import dev.ltms.fleet.inject.BackendErrorSink;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
|
||||
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
|
||||
import dev.ltms.fleet.inject.MemberPresence;
|
||||
import dev.ltms.fleet.inject.StatusPoller;
|
||||
import dev.ltms.fleet.inject.TurnListener;
|
||||
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||
import dev.ltms.fleet.lead.LeadLauncher;
|
||||
import dev.ltms.fleet.lead.LeadRollover;
|
||||
import dev.ltms.fleet.mcp.ConnectionIdentity;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import dev.ltms.fleet.mcp.LsofPeerPidLookup;
|
||||
import dev.ltms.fleet.mcp.LsofProcessCwdLookup;
|
||||
import dev.ltms.fleet.mcp.PrimaryRegistry;
|
||||
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||
import dev.ltms.fleet.member.HerdrPeerLauncher;
|
||||
import dev.ltms.fleet.member.MemberCredentialPolicyView;
|
||||
import dev.ltms.fleet.metrics.FleetMetrics;
|
||||
import dev.ltms.fleet.metrics.Metrics;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadCoordLoop;
|
||||
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
|
||||
import dev.ltms.fleet.msg.MessageService;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import dev.ltms.fleet.msg.ReplyPushLoop;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.power.CaffeinateSleepAssertionMechanism;
|
||||
import dev.ltms.fleet.power.IdleSleepGuard;
|
||||
import dev.ltms.fleet.rest.FleetApp;
|
||||
import dev.ltms.fleet.session.GitWorktrees;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import dev.ltms.fleet.session.SessionReaper;
|
||||
import io.javalin.Javalin;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.Predicate;
|
||||
import java.util.function.Supplier;
|
||||
import java.util.regex.Pattern;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
/**
|
||||
* fleetd #612 Unit A: the real boot assembly, extracted out of {@code Fleetd.main} so a test can
|
||||
* drive it directly. {@link #assembleAndStart} is <em>the same statements {@code main} used to run
|
||||
* inline</em>, in the same order, against a real {@link ResourcePorts} in production and a fake one
|
||||
* in a test — see {@code FleetdAssemblyLifecycleTest}. {@code Fleetd.main} keeps config loading and
|
||||
* {@code cfg.validateAll()}; everything from immediately after that call onward moved here —
|
||||
* including, since fleetd #612 A-gaps (gap 2), the two post-validation reports ({@code
|
||||
* reportRoleFallbackGaps}, {@code assertChartersNameOnlyRegisteredTools}) that Unit A originally
|
||||
* left behind in {@code main}. Those two calls do no I/O themselves, but leaving them outside this
|
||||
* method meant deleting either one compiled clean and left the whole suite green — nothing drove
|
||||
* {@code main} itself, so nothing could notice. They run first here, in the same relative order,
|
||||
* before the herdr socket or anything else that touches the outside world.
|
||||
*
|
||||
* <p><strong>Construction and start order is preserved exactly, on purpose.</strong> This is not
|
||||
* rebuilt into "construct everything, then start everything" — that would change boot timing. The
|
||||
* order recorded before any code moved (see the ticket and {@code FleetdAssemblyLifecycleTest}):
|
||||
* {@code SessionReaper} starts first (if {@code lifecycle.idleTtlSeconds} is configured), then
|
||||
* {@link StatusPoller}, then the optional {@link LeadHeartbeatLoop} and {@link FleetHealthMonitor},
|
||||
* then the optional {@link LeadCoordLoop} and {@link ConfigWatcher}, and the HTTP server starts
|
||||
* last of all. The close order (see {@link FleetdRuntime#close()}) is the mirror the original
|
||||
* shutdown hook always used.
|
||||
*
|
||||
* <p><strong>One statement could not move without reordering startup.</strong> {@code Fleetd.main}
|
||||
* registered its shutdown hook <em>before</em> building the Javalin {@code FleetApp} — the hook
|
||||
* itself never touched {@code app} (it still doesn't; see {@link FleetdRuntime#close()}), but the
|
||||
* hook needs a {@link FleetdRuntime} to close over, and the ticket asks that runtime to also own
|
||||
* {@code FleetApp}. Building the runtime before the app exists and mutating it afterward (via
|
||||
* {@link FleetdRuntime#attachApp}) preserves the exact original order — hook registered, then app
|
||||
* built, then HTTP started — without moving the app's construction earlier or the hook's
|
||||
* registration later. That is the one seam this ticket did not get to pin any other way.
|
||||
*
|
||||
* <p><strong>No inert variant.</strong> Deliberately, there is no overload of this method that
|
||||
* accepts a smaller/optional {@link ResourcePorts} or defaults one internally. A future edit that
|
||||
* wants to skip {@code FleetdAssembly} entirely and build its own graph is still possible — no
|
||||
* static analysis stops that — but it cannot do so by quietly swapping this call for an inert
|
||||
* substitute that still compiles, because none exists.
|
||||
*/
|
||||
final class FleetdAssembly {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(Fleetd.class);
|
||||
|
||||
/** CB-637: how often the lead coordination loop looks for peer messages — see {@code Fleetd}'s own constant. */
|
||||
private static final long LEAD_COORD_INTERVAL_MS = 3_000L;
|
||||
|
||||
private FleetdAssembly() {
|
||||
}
|
||||
|
||||
static FleetdRuntime assembleAndStart(AssemblyInputs inputs, ResourcePorts ports) {
|
||||
FleetConfig cfg = inputs.cfg();
|
||||
ConfigRef config = inputs.config();
|
||||
SubscriptionGuard guard = inputs.guard();
|
||||
|
||||
// fleetd #612 A-gaps (gap 2): moved in from Fleetd.main, immediately after cfg.validateAll()
|
||||
// there — the exact point main used to call these two, and still the first thing this
|
||||
// method does, before any socket or broker work below. See this class's javadoc and each
|
||||
// method's own for why they run here rather than in FleetConfig#validateAll() itself.
|
||||
Fleetd.reportRoleFallbackGaps(cfg);
|
||||
Fleetd.assertChartersNameOnlyRegisteredTools(cfg);
|
||||
|
||||
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
|
||||
? Path.of(cfg.herdrSocket())
|
||||
: UnixSocketHerdrClient.defaultSocketPath();
|
||||
|
||||
HerdrClient herdr = ports.connectHerdr(socket);
|
||||
HerdrClient memberHerdr = cfg.memberHerdrSocket() != null && !cfg.memberHerdrSocket().isBlank()
|
||||
? ports.connectHerdr(Path.of(cfg.memberHerdrSocket()))
|
||||
: herdr;
|
||||
AtomicReference<Supplier<Map<String, String>>> leadsRef = new AtomicReference<>(Map::of);
|
||||
HerdrRouter router = new HerdrRouter(herdr, memberHerdr,
|
||||
target -> leadsRef.get().get().containsKey(target));
|
||||
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
|
||||
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
|
||||
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
|
||||
Map<String, FleetConfig.Profile> claudeProfiles = new LinkedHashMap<>();
|
||||
Map<String, FleetConfig.Profile> opencodeProfiles = new LinkedHashMap<>();
|
||||
cfg.profiles().forEach((name, w) -> {
|
||||
if (w.isOpenCode()) {
|
||||
opencodeProfiles.put(name, w);
|
||||
} else {
|
||||
claudeProfiles.put(name, w);
|
||||
}
|
||||
});
|
||||
List<HerdrPeerLauncher> adapters = new ArrayList<>();
|
||||
// fleetd #175: the daemon's real ExhaustionSink can only be built once `sessions` exists
|
||||
// (below), but `sessions` needs `workers`, which needs the adapters built right here — a
|
||||
// genuine cycle. Break it exactly like liveCountRef below: a forwarding sink built now,
|
||||
// pointed at the real one once it exists.
|
||||
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
|
||||
ExhaustionSink forwardingExhaustionSink = Fleetd.forwardingExhaustionSink(exhaustionSinkRef);
|
||||
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
|
||||
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
|
||||
// unless opencode is the only kind configured.
|
||||
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
|
||||
adapters.add(Fleetd.claudeCodeLauncher(router.memberAgents(), router.memberSpaces(), guard,
|
||||
claudeProfiles, cfg, config));
|
||||
}
|
||||
if (!opencodeProfiles.isEmpty()) {
|
||||
adapters.add(Fleetd.openCodeLauncher(router.memberAgents(), router.memberSpaces(),
|
||||
opencodeProfiles, cfg, config, forwardingExhaustionSink));
|
||||
}
|
||||
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
|
||||
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
|
||||
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
|
||||
BackendQuarantine quarantine = BackendQuarantine.withEscalation(ports.nanoClock(),
|
||||
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
|
||||
// fleetd #201 Unit 5: one outage-cool-off tracker for the whole daemon, shared between the
|
||||
// launcher (checked at spawn, like `quarantine` above) and the backend-error sink wired in
|
||||
// below (written on a classified backend error).
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(ports.nanoClock());
|
||||
PeerLauncher workers = new CompositePeerLauncher(
|
||||
adapters,
|
||||
cfg.effectiveDefaultProfile(),
|
||||
config,
|
||||
profileName -> liveCountRef.get().apply(profileName),
|
||||
quarantine,
|
||||
outagePolicy);
|
||||
// fleetd #422 follow-up: say which of the three model-gate states the daemon booted into.
|
||||
log.info("model gate (fleetd #422): {}", Fleetd.modelGateCoverageLine(workers.modelGateState()));
|
||||
// CB-504: under supervision (launchd/systemd) fleetd can start before herdr's socket
|
||||
// exists. Wait, then degrade rather than die: serving with /healthz reporting "degraded" is
|
||||
// strictly more useful than exiting.
|
||||
Fleetd.HerdrAwaitOutcome herdrOutcome = Fleetd.awaitHerdr(herdr, ports.nanoClock(), Fleetd::sleepHerdrPoll);
|
||||
boolean herdrUp = Fleetd.logHerdrWaitOutcomeAndShouldReap(herdrOutcome);
|
||||
if (herdrUp) {
|
||||
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
|
||||
// with the previous process — reap those leaked orphans now, before we start serving.
|
||||
workers.reapOrphanWorkers();
|
||||
}
|
||||
|
||||
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
|
||||
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
|
||||
// CB-303 part 2: context cap is opt-in and disabled (0) when absent/null.
|
||||
int contextCap = 0;
|
||||
if (cfg.lifecycle() != null && cfg.lifecycle().contextCap() != null
|
||||
&& cfg.lifecycle().contextCap() > 0) {
|
||||
contextCap = cfg.lifecycle().contextCap();
|
||||
}
|
||||
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
|
||||
SessionManager sessions = new SessionManager(workers,
|
||||
new GitWorktrees(cfg.worktreeRoot(), cfg.worktreeGroup(), cfg.memberSkills()),
|
||||
ports.nanoClock(), contextCap, clearAfterTurn);
|
||||
liveCountRef.set(profileName -> Fleetd.liveSessionCount(sessions.roster(), profileName));
|
||||
|
||||
// Idle-sleep guard: hold an OS-level assertion against idle sleep while at least one
|
||||
// member is live. No-op (never constructed) off macOS or when idleSleepGuard.enabled is
|
||||
// explicitly false; the mechanism itself is additionally a no-op if 'caffeinate' cannot be
|
||||
// started, so this can never fail a spawn, a release, or startup.
|
||||
boolean idleSleepGuardEnabled = cfg.idleSleepGuard() == null || cfg.idleSleepGuard().isEnabled();
|
||||
final IdleSleepGuard idleSleepGuard;
|
||||
if (idleSleepGuardEnabled) {
|
||||
idleSleepGuard = new IdleSleepGuard(new CaffeinateSleepAssertionMechanism(), sessions::size);
|
||||
sessions.onAcquire(_ -> idleSleepGuard.recheck());
|
||||
sessions.onRelease(_ -> idleSleepGuard.recheck());
|
||||
} else {
|
||||
idleSleepGuard = null;
|
||||
}
|
||||
|
||||
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled. FIRST of the
|
||||
// recurring background loops to start — see this class's own javadoc for the full order.
|
||||
final SessionReaper reaper;
|
||||
if (cfg.lifecycle() != null
|
||||
&& cfg.lifecycle().idleTtlSeconds() != null
|
||||
&& cfg.lifecycle().idleTtlSeconds() > 0) {
|
||||
reaper = new SessionReaper(sessions, cfg.lifecycle().idleTtlSeconds());
|
||||
reaper.start();
|
||||
} else {
|
||||
reaper = null;
|
||||
}
|
||||
|
||||
// CB-530: every pane the config names as a lead, merged from `leaders:` and the legacy
|
||||
// singular pin. PrimaryRegistry below still tracks ONE terminal — it addresses the push
|
||||
// loop's nudges, which need a single destination — so it keeps the legacy pin.
|
||||
Map<String, String> leadTerminals = cfg.leaderTerminals();
|
||||
if (leadTerminals.size() > 1) {
|
||||
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
|
||||
}
|
||||
// CB-531/CB-579: discover leads by the tab labels the operator writes, one scanner per
|
||||
// configured lead's own exact `tab:` label.
|
||||
final Supplier<Map<String, String>> leads;
|
||||
var leaders = cfg.fleet().leaders();
|
||||
if (!leaders.isEmpty()) {
|
||||
Map<String, String> tabToName = new LinkedHashMap<>();
|
||||
leaders.forEach((name, leader) -> {
|
||||
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
|
||||
tabToName.put(leader.tab(), name);
|
||||
}
|
||||
});
|
||||
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
|
||||
// This must use the lead daemon: scanning member tabs would demote the lead to a worker.
|
||||
leads = new LeadTabScanner(herdr, tabToName, Set.of(),
|
||||
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), ports.nanoClock());
|
||||
log.info("lead scan: tabs {} host a lead (rescan every {}s, shared fleet space)",
|
||||
tabToName.keySet(), scanIntervalSeconds);
|
||||
} else {
|
||||
leads = () -> leadTerminals;
|
||||
}
|
||||
leadsRef.set(leads);
|
||||
|
||||
// CB-558: start any declared lead that is not already running. After the scanner is built,
|
||||
// and only when herdr answered — the launcher's whole safety property is that it can count
|
||||
// live leads first, and must never guess and risk a second orchestrator.
|
||||
if (herdrUp && !leaders.isEmpty()) {
|
||||
int launched = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg).ensureLeads();
|
||||
if (launched > 0) {
|
||||
log.info("lead auto-launch: {} lead(s) started", launched);
|
||||
}
|
||||
}
|
||||
|
||||
// CB-548: config-declared architect slots. Nothing here spawns a slot; the terminal → slot
|
||||
// binding is owned by the registry and empty at startup.
|
||||
MemberRegistry members = MemberRegistry.live(() -> config.get().fleet());
|
||||
sessions.setMemberLifecycle(members);
|
||||
if (!members.slots().isEmpty()) {
|
||||
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
|
||||
+ "spawn lifecycle binds a live terminal to it)",
|
||||
members.slots().size(), members.slots().keySet());
|
||||
}
|
||||
|
||||
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
|
||||
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
// CB-578 stage A: classify a completion-fallback scrape that matches a profile's configured
|
||||
// usage-limit refusal as BACKEND_EXHAUSTED rather than handing it back as a real answer.
|
||||
LiveExhaustedPatterns liveExhaustedPatterns = Fleetd.liveExhaustedPatterns(config);
|
||||
ExhaustedPatternLookup exhaustedPatterns = Fleetd.exhaustedPatternLookup(sessions::roster, liveExhaustedPatterns);
|
||||
// The startup coverage line still reports the boot-time snapshot only.
|
||||
Set<String> exhaustedConfiguredAtStartup = cfg.profiles().entrySet().stream()
|
||||
.filter(e -> e.getValue().hasExhaustedPattern())
|
||||
.map(Map.Entry::getKey)
|
||||
.collect(Collectors.toCollection(LinkedHashSet::new));
|
||||
log.info("backend-exhausted classification (CB-578 stage A): {}",
|
||||
Fleetd.exhaustedPatternCoverageLine(cfg.profiles().keySet(), exhaustedConfiguredAtStartup));
|
||||
// fleetd #201 Unit 5: classify a completion-fallback scrape that matches a profile's
|
||||
// configured backend-error refusal (a credential outage, a provider 5xx) as a backend error
|
||||
// rather than handing it back as a real answer.
|
||||
Map<String, Pattern> errorPatternsByProfile = new LinkedHashMap<>();
|
||||
cfg.profiles().forEach((name, profile) -> {
|
||||
if (profile.hasErrorPattern()) {
|
||||
errorPatternsByProfile.put(name, Pattern.compile(profile.errorPattern()));
|
||||
}
|
||||
});
|
||||
BackendErrorPatternLookup backendErrorPatterns = Fleetd.backendErrorPatternLookup(sessions::roster,
|
||||
errorPatternsByProfile);
|
||||
log.info("backend-error classification (fleetd #201 Unit 5): {}",
|
||||
Fleetd.errorPatternCoverageLine(cfg.profiles().keySet(), errorPatternsByProfile.keySet()));
|
||||
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
|
||||
// CREDENTIAL, so a profile sharing that credential is refused too, not just the one that
|
||||
// happened to report it.
|
||||
Map<String, String> quarantineReasonByCredential = new ConcurrentHashMap<>();
|
||||
// fleetd #175: point the forwarding sink handed to OpenCodeLauncher above at the real one,
|
||||
// now that `sessions` exists to resolve target -> session -> profile.
|
||||
ExhaustionSink exhaustionSink = Fleetd.publishExhaustionSink(exhaustionSinkRef, sessions, config,
|
||||
quarantine, quarantineReasonByCredential, cfg);
|
||||
// fleetd #201 Unit 5: the production BackendErrorSink needs `pushLoop` (built further below,
|
||||
// after `sessions`) to tell a lead about an incident or an unmapped target — the same
|
||||
// construction-order cycle `exhaustionSinkRef` breaks above, broken the same way: a mutable
|
||||
// holder set once `pushLoop` exists, read lazily from inside the lambda built here.
|
||||
AtomicReference<ReplyPushLoop> pushLoopRef = new AtomicReference<>();
|
||||
BackendErrorSink backendErrorSink = Fleetd.backendErrorSink(sessions, () -> config.get().profiles(),
|
||||
outagePolicy, pushLoopRef::get);
|
||||
AgentControl agents = router.memberAgents();
|
||||
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns,
|
||||
exhaustionSink, backendErrorPatterns, backendErrorSink, ports.nanoClock(),
|
||||
Fleetd.worktreeBranchLookup(sessions::roster));
|
||||
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
|
||||
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
|
||||
MemberPresence presence = sessions.asPresence();
|
||||
TurnListener turnListener = Fleetd.turnListener(completion, sessions);
|
||||
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads);
|
||||
// fleetd #556: registration is wired directly to `completion`, not folded into the
|
||||
// `turnListener` fan-out above — so it survives `sessions.onDelivered` (or any future
|
||||
// listener) throwing, regardless of call order.
|
||||
Injector injector = new Injector(router, turnListener, deliverable,
|
||||
presence::forget, Fleetd.turnRegistrar(completion));
|
||||
StatusPoller poller = new StatusPoller(router, injector, Injector.POLL_INTERVAL_MILLIS);
|
||||
poller.start(); // SECOND of the recurring background loops to start, after the reaper.
|
||||
|
||||
// CB-307: reply inbox. A broker: block selects the AMQP-backed durable adapter; absent (or
|
||||
// unusable), fleetd stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
|
||||
// connection, so keep the reference to close it in the ordered shutdown hook.
|
||||
final ReplyInbox replyInbox = Fleetd.selectReplyInbox(cfg.broker(), ports.environment(),
|
||||
ports.replyInboxOpener());
|
||||
// CB-637: this daemon's lead-to-lead mailbox on the SHARED coordination vhost — a separate
|
||||
// broker from the reply inbox by design. Absent a coordinator: block this is null and every
|
||||
// lead path below is simply not wired, exactly the behaviour before this ticket. It owns a
|
||||
// broker connection, so keep the reference for the ordered shutdown hook.
|
||||
final LeadChannelHandle leadMailbox = Fleetd.openLeadMailbox(cfg.coordinator(), ports.environment(),
|
||||
ports.leadMailboxOpener());
|
||||
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
|
||||
String pinnedPrimaryTerminal = cfg.primary() != null ? cfg.primary().terminal() : null;
|
||||
PrimaryRegistry primaryRegistry = new PrimaryRegistry(pinnedPrimaryTerminal);
|
||||
// CB-532: `primary.terminal` is superseded and no longer needed for either of its jobs.
|
||||
if (pinnedPrimaryTerminal != null && !pinnedPrimaryTerminal.isBlank()) {
|
||||
log.warn("primary.terminal is DEPRECATED (CB-532) and can be deleted: identity now comes "
|
||||
+ "from leaders:/leadScan:, and reply nudges follow the lead that delegated. "
|
||||
+ "It still works, and is still the fallback nudge destination when a restart "
|
||||
+ "has lost the delegation map. Its pushReminders/pushBackoffMs stay valid.");
|
||||
}
|
||||
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
|
||||
// open fleet_send. Uses its own lightweight scheduled executor, separate from the injector.
|
||||
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
|
||||
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
|
||||
var pushScheduler = ports.newScheduler("bridge-push-");
|
||||
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
|
||||
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
|
||||
Metrics metrics = FleetMetrics.create(sessions, replyInbox);
|
||||
var pushLoop = new ReplyPushLoop(primaryRegistry, router.leadAgents(), replyInbox,
|
||||
pushScheduler, maxReminders, backoffMs, metrics);
|
||||
// fleetd #201 Unit 5: point the forwarding holder captured by the backendErrorSink lambda
|
||||
// above at the real push loop, now that it exists.
|
||||
pushLoopRef.set(pushLoop);
|
||||
// CB-551: idle-lead heartbeat. Opt-in; absent `leadHeartbeat:` this is never constructed, so
|
||||
// an upgraded daemon cannot silently start spending subscription on nudging an idle lead.
|
||||
// THIRD of the recurring background loops to start (optional).
|
||||
final LeadHeartbeatLoop heartbeat;
|
||||
var heartbeatScheduler = ports.newScheduler("bridge-heartbeat-");
|
||||
if (cfg.leadHeartbeat() != null) {
|
||||
var hb = cfg.leadHeartbeat();
|
||||
var leadContextGauge = new LeadContextGauge();
|
||||
// fleetd #621: the context-high notice's own wording must track this same effective
|
||||
// value — LeadRollover.confirm(...) already gates the roll on it (LeadRollover.java:480),
|
||||
// and absent `leadRollover:` entirely the roll is unusable regardless (NOT_CONFIGURED),
|
||||
// so `true` (the FleetConfig.LeadRollover default) is the safe, byte-identical fallback.
|
||||
// Carried in from Fleetd.main when #612 Unit A merged main: #622 added this line to the
|
||||
// block Unit A had already moved here, so the merge would otherwise have silently
|
||||
// dropped it — with a fully green suite, because nothing pins it (see the follow-up issue).
|
||||
boolean requireOperatorConfirm = cfg.leadRollover() == null || cfg.leadRollover().requireOperatorConfirm();
|
||||
heartbeat = new LeadHeartbeatLoop(primaryRegistry, router.leadAgents(), replyInbox, sessions::roster,
|
||||
pushLoop, heartbeatScheduler, ports.nanoClock(),
|
||||
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
|
||||
metrics,
|
||||
Fleetd.leadContextSource(leadContextGauge, router.leadAgents(), leads,
|
||||
Fleetd.leadConfigDirLookup(() -> config.get().profiles(), leaders)),
|
||||
Boolean.TRUE.equals(hb.contextHighNudge()), requireOperatorConfirm);
|
||||
heartbeat.start();
|
||||
} else {
|
||||
heartbeat = null;
|
||||
heartbeatScheduler.shutdownNow();
|
||||
}
|
||||
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed.
|
||||
LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), config, leads);
|
||||
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
|
||||
pushLoop, metrics);
|
||||
|
||||
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
|
||||
// FOURTH of the recurring background loops to start (optional).
|
||||
final FleetHealthMonitor healthMonitor;
|
||||
var healthScheduler = ports.newScheduler("bridge-health-");
|
||||
if (cfg.health() != null && cfg.health().isEnabled()) {
|
||||
// CB-580: a member found GONE/NEVER_READY must fail whatever ticket is waiting on it.
|
||||
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
|
||||
ports.nanoClock(), ports.wallClockNanos(),
|
||||
cfg.health().intervalOrDefault(),
|
||||
cfg.health().workingSuspectAfterOrDefault(), Fleetd.healthFailTarget(messages));
|
||||
String coverage = FleetHealthMonitor.coverage(true,
|
||||
cfg.health().notifications() != null && cfg.health().notifications().configured());
|
||||
if ("detection-only".equals(coverage)) {
|
||||
log.warn("fleet health: {} (no notification sink configured)", coverage);
|
||||
} else {
|
||||
log.info("fleet health: {}", coverage);
|
||||
}
|
||||
healthMonitor.start();
|
||||
} else {
|
||||
healthMonitor = null;
|
||||
healthScheduler.shutdownNow();
|
||||
}
|
||||
|
||||
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
|
||||
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
|
||||
sessions.onAcquire(replyInbox::own);
|
||||
// CB-516: releasing a worker must fail whatever send was waiting on it.
|
||||
sessions.onRelease(Fleetd.releaseCleanup(messages, replyInbox, primaryRegistry));
|
||||
|
||||
// MCP server face (CB-105): fleet_send/fleet_reply/fleet_status, mounted at /mcp.
|
||||
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
|
||||
ConnectionIdentity identity = new ConnectionIdentity(
|
||||
new PaneLocator(herdr, memberHerdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
|
||||
|
||||
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
|
||||
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
|
||||
final CallerResolver callers;
|
||||
if (cfg.auth().tokenMode()) {
|
||||
String token = ports.environment().get(cfg.auth().tokenEnv());
|
||||
if (token == null || token.isBlank()) {
|
||||
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
|
||||
+ " is unset or empty — export it before starting fleetd");
|
||||
}
|
||||
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
|
||||
log.info("auth: token mode (bearer required for non-worker callers, env {})",
|
||||
cfg.auth().tokenEnv());
|
||||
} else {
|
||||
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
|
||||
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
|
||||
}
|
||||
|
||||
FleetMcp.QuarantineSource quarantineSource = Fleetd.quarantineSource(config, quarantine,
|
||||
liveExhaustedPatterns, quarantineReasonByCredential);
|
||||
FleetMcp.OutageSource outageSource = new FleetMcp.OutageSource(profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.effectiveCredentialId();
|
||||
}, outagePolicy);
|
||||
|
||||
FleetMcp.LoopHealthSource loopHealth = Fleetd.loopHealthSource(poller, reaper);
|
||||
FleetMcp mcp = new FleetMcp(messages, workers, sessions, identity, presence,
|
||||
primaryRegistry, callers, FleetMcp.AuthorizationMode.ENFORCED, metrics,
|
||||
Fleetd.capacitySource(config, cfg, profile -> liveCountRef.get().apply(profile)),
|
||||
Fleetd.healthCoverageSource(config),
|
||||
loopHealth,
|
||||
quarantineSource,
|
||||
leadMailbox,
|
||||
outageSource,
|
||||
new FleetMcp.LeadSeatSource(Fleetd.leadSeatLookup(() -> config.get().profiles(), leaders, leads)),
|
||||
Fleetd.leadConfigDirSource(() -> config.get().profiles(), leaders),
|
||||
cfg.coordinator() == null ? List.of() : cfg.coordinator().peers(),
|
||||
leadRollover);
|
||||
|
||||
// CB-637: the receive half. Only constructed when a lead mailbox actually opened.
|
||||
final LeadCoordLoop leadCoordLoop;
|
||||
final ScheduledExecutorService leadCoordSchedulerRef;
|
||||
if (leadMailbox != null) {
|
||||
var leadCoordScheduler = ports.newScheduler("bridge-leadcoord-");
|
||||
leadCoordLoop = new LeadCoordLoop(leadMailbox, router.leadAgents(), leads, leadCoordScheduler,
|
||||
LEAD_COORD_INTERVAL_MS);
|
||||
leadCoordLoop.start(); // FIFTH of the recurring background loops to start (optional).
|
||||
leadCoordSchedulerRef = leadCoordScheduler;
|
||||
} else {
|
||||
leadCoordLoop = null;
|
||||
leadCoordSchedulerRef = null;
|
||||
}
|
||||
|
||||
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed.
|
||||
// SIXTH of the recurring background loops to start (optional).
|
||||
final ConfigWatcher configWatcher;
|
||||
if (cfg.configReload() != null && cfg.configReload().isEnabled()) {
|
||||
configWatcher = new ConfigWatcher(config, cfg.configReload().intervalSeconds());
|
||||
configWatcher.start();
|
||||
} else {
|
||||
configWatcher = null;
|
||||
}
|
||||
|
||||
// CB-303 part 3: single ordered shutdown hook — see FleetdRuntime#close() for the statements
|
||||
// this used to be. Registered here, at the exact point `main` used to register it: after
|
||||
// configWatcher, before the Javalin app exists (see this class's own javadoc for why).
|
||||
FleetdRuntime runtime = new FleetdRuntime(cfg, sessions, router, poller, messages, pushLoop, heartbeat,
|
||||
leadCoordLoop, leadCoordSchedulerRef, healthMonitor, configWatcher, mcp, reaper, idleSleepGuard,
|
||||
replyInbox, leadMailbox, completion, injector);
|
||||
ports.addShutdownHook(runtime::close);
|
||||
|
||||
// CB-185: give FleetApp both daemons — /healthz must require both to answer and
|
||||
// GET /sessions must merge across both, or a down/unpolled member daemon is invisible.
|
||||
Javalin app = new FleetApp(herdr, memberHerdr, workers, sessions, messages, presence, mcp.servlet(),
|
||||
callers, metrics, deliverable,
|
||||
() -> MemberCredentialPolicyView.of(config.get().memberCredentials()),
|
||||
quarantineSource, outageSource, loopHealth).build();
|
||||
runtime.attachApp(app);
|
||||
ports.startHttp(app, cfg.bind().host(), cfg.bind().port()); // HTTP starts LAST, always.
|
||||
log.info("fleetd listening on {}:{}, herdr socket {}",
|
||||
cfg.bind().host(), cfg.bind().port(), socket);
|
||||
return runtime;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,166 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigWatcher;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.health.FleetHealthMonitor;
|
||||
import dev.ltms.fleet.herdr.HerdrRouter;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.inject.StatusPoller;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadCoordLoop;
|
||||
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
|
||||
import dev.ltms.fleet.msg.MessageService;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import dev.ltms.fleet.msg.ReplyPushLoop;
|
||||
import dev.ltms.fleet.power.IdleSleepGuard;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import dev.ltms.fleet.session.SessionReaper;
|
||||
import io.javalin.Javalin;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
|
||||
/**
|
||||
* fleetd #612 Unit A: the assembled daemon. {@link FleetdAssembly#assembleAndStart} builds exactly
|
||||
* one of these and hands it to {@code ResourcePorts.addShutdownHook}; {@link #close()} is the
|
||||
* single ordered shutdown, moved verbatim out of {@code Fleetd.main}'s old shutdown-hook
|
||||
* {@code Thread} body — same statements, same order, see that method's javadoc.
|
||||
*
|
||||
* <p><strong>Owns the real production objects, never a copy.</strong> Every package-private
|
||||
* accessor below returns the identical instance the running daemon is using. That is the entire
|
||||
* point of this class existing (see fleetd #612's problem statement): a test that inspected a
|
||||
* snapshot built alongside the real objects could pass while production silently received
|
||||
* something else — the exact shape of the #602/#606 defect this ticket exists to stop from
|
||||
* recurring one call site at a time. Nothing here is rebuilt or copied for a test's benefit.
|
||||
*/
|
||||
final class FleetdRuntime implements AutoCloseable {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(FleetdRuntime.class);
|
||||
|
||||
private final FleetConfig cfg;
|
||||
private final SessionManager sessions;
|
||||
private final HerdrRouter router;
|
||||
private final StatusPoller poller;
|
||||
private final MessageService messages;
|
||||
private final ReplyPushLoop pushLoop;
|
||||
private final LeadHeartbeatLoop heartbeat; // nullable — leadHeartbeat: opt-in
|
||||
private final LeadCoordLoop leadCoordLoop; // nullable — coordinator: opt-in
|
||||
private final ScheduledExecutorService leadCoordScheduler; // nullable, paired with leadCoordLoop
|
||||
private final FleetHealthMonitor healthMonitor; // nullable — health.enabled opt-in
|
||||
private final ConfigWatcher configWatcher; // nullable — configReload.enabled opt-in
|
||||
private final FleetMcp mcp;
|
||||
private final SessionReaper reaper; // nullable — lifecycle.idleTtlSeconds opt-in
|
||||
private final IdleSleepGuard idleSleepGuard; // nullable — idleSleepGuard.enabled: false
|
||||
private final ReplyInbox replyInbox;
|
||||
private final LeadChannelHandle leadMailbox; // nullable — coordinator: opt-in
|
||||
private final CompletionResolver completion;
|
||||
private final Injector injector;
|
||||
/**
|
||||
* Not final: {@code Fleetd.main}'s shutdown hook was registered <em>before</em> the Javalin
|
||||
* {@code FleetApp} was built and started — see {@link FleetdAssembly#assembleAndStart}'s javadoc
|
||||
* for why that order could not be preserved AND have this constructor take {@code app}. {@link
|
||||
* #attachApp} is called immediately after the real app is built, still before HTTP starts
|
||||
* listening, so this is set long before any test or caller could observe it unset.
|
||||
*/
|
||||
private Javalin app;
|
||||
|
||||
FleetdRuntime(FleetConfig cfg, SessionManager sessions, HerdrRouter router, StatusPoller poller,
|
||||
MessageService messages, ReplyPushLoop pushLoop, LeadHeartbeatLoop heartbeat,
|
||||
LeadCoordLoop leadCoordLoop, ScheduledExecutorService leadCoordScheduler,
|
||||
FleetHealthMonitor healthMonitor, ConfigWatcher configWatcher, FleetMcp mcp,
|
||||
SessionReaper reaper, IdleSleepGuard idleSleepGuard, ReplyInbox replyInbox,
|
||||
LeadChannelHandle leadMailbox, CompletionResolver completion, Injector injector) {
|
||||
this.cfg = cfg;
|
||||
this.sessions = sessions;
|
||||
this.router = router;
|
||||
this.poller = poller;
|
||||
this.messages = messages;
|
||||
this.pushLoop = pushLoop;
|
||||
this.heartbeat = heartbeat;
|
||||
this.leadCoordLoop = leadCoordLoop;
|
||||
this.leadCoordScheduler = leadCoordScheduler;
|
||||
this.healthMonitor = healthMonitor;
|
||||
this.configWatcher = configWatcher;
|
||||
this.mcp = mcp;
|
||||
this.reaper = reaper;
|
||||
this.idleSleepGuard = idleSleepGuard;
|
||||
this.replyInbox = replyInbox;
|
||||
this.leadMailbox = leadMailbox;
|
||||
this.completion = completion;
|
||||
this.injector = injector;
|
||||
}
|
||||
|
||||
/** See the {@link #app} field doc for why this is a late-bound setter rather than a constructor arg. */
|
||||
void attachApp(Javalin app) {
|
||||
this.app = app;
|
||||
}
|
||||
|
||||
// --- package-private accessors: the SAME instances this runtime owns, never a copy ---------
|
||||
|
||||
SessionManager sessions() { return sessions; }
|
||||
HerdrRouter router() { return router; }
|
||||
StatusPoller poller() { return poller; }
|
||||
MessageService messages() { return messages; }
|
||||
ReplyPushLoop pushLoop() { return pushLoop; }
|
||||
LeadHeartbeatLoop heartbeat() { return heartbeat; }
|
||||
LeadCoordLoop leadCoordLoop() { return leadCoordLoop; }
|
||||
FleetHealthMonitor healthMonitor() { return healthMonitor; }
|
||||
ConfigWatcher configWatcher() { return configWatcher; }
|
||||
FleetMcp mcp() { return mcp; }
|
||||
SessionReaper reaper() { return reaper; }
|
||||
IdleSleepGuard idleSleepGuard() { return idleSleepGuard; }
|
||||
ReplyInbox replyInbox() { return replyInbox; }
|
||||
LeadChannelHandle leadMailbox() { return leadMailbox; }
|
||||
CompletionResolver completion() { return completion; }
|
||||
Injector injector() { return injector; }
|
||||
Javalin app() { return app; }
|
||||
|
||||
/**
|
||||
* CB-303 part 3: the single ordered shutdown. Moved verbatim out of {@code Fleetd.main}'s
|
||||
* shutdown-hook {@code Thread} body (fleetd #612 Unit A) — drain sessions first while herdr is
|
||||
* still open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close
|
||||
* herdr last, exactly as before. {@code Fleetd.main} never calls this directly; it hands the
|
||||
* reference to {@code ResourcePorts.addShutdownHook} the moment this runtime exists, the same
|
||||
* point it used to register the hook {@code Thread} itself.
|
||||
*/
|
||||
@Override
|
||||
public void close() {
|
||||
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
|
||||
poller.stop();
|
||||
messages.close();
|
||||
pushLoop.close();
|
||||
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
|
||||
if (leadCoordLoop != null) leadCoordLoop.close(); // CB-637: stop delivering peer-lead messages
|
||||
if (leadCoordScheduler != null) leadCoordScheduler.shutdownNow();
|
||||
if (healthMonitor != null) healthMonitor.stop();
|
||||
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
|
||||
mcp.close();
|
||||
if (reaper != null) reaper.stop();
|
||||
// Idle-sleep guard: release unconditionally, even though sessions.close() above already
|
||||
// drained every session (and each release already drove the live count to 0, which
|
||||
// releases the guard's assertion on its own) — this is the backstop for a drain that was
|
||||
// itself interrupted or threw, so no caffeinate child ever outlives the daemon.
|
||||
if (idleSleepGuard != null) idleSleepGuard.close();
|
||||
// Release the broker connection last among message resources (no-op for the in-memory inbox).
|
||||
if (replyInbox instanceof AutoCloseable closeable) {
|
||||
try {
|
||||
closeable.close();
|
||||
} catch (Exception e) {
|
||||
log.debug("reply inbox close: {}", e.toString());
|
||||
}
|
||||
}
|
||||
// CB-637: the coordination connection goes with it — after the loop that reads it has
|
||||
// stopped, so no tick can be mid-ack against a closed channel.
|
||||
if (leadMailbox != null) {
|
||||
try {
|
||||
leadMailbox.close();
|
||||
} catch (Exception e) {
|
||||
log.debug("lead mailbox close: {}", e.toString());
|
||||
}
|
||||
}
|
||||
router.close();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,67 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import io.javalin.Javalin;
|
||||
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
/**
|
||||
* fleetd #612 Unit A: every boot-time side effect {@link FleetdAssembly#assembleAndStart} performs
|
||||
* that a real daemon must do for real, and a test must not — read the process environment, connect
|
||||
* a herdr client, open a broker (the reply inbox, the lead mailbox), read a clock, start a
|
||||
* background scheduler, register the JVM shutdown hook, and bind the HTTP server.
|
||||
*
|
||||
* <p>{@link #system()} is the one production implementation ({@code SystemResourcePorts}), wired
|
||||
* verbatim from what {@code Fleetd.main} used to call directly at each of these call sites. A test
|
||||
* builds its own implementation instead of receiving an inert default from this interface —
|
||||
* deliberately, there is no {@code ResourcePorts.none()}. fleetd #612's whole problem is a call
|
||||
* site quietly swapped for an inert variant that still compiles; adding one here, even for tests,
|
||||
* would hand a future edit to {@code FleetdAssembly} the exact compiling substitute this ticket
|
||||
* exists to rule out. A test that wants an inert resource writes its own fake and owns that
|
||||
* decision explicitly.
|
||||
*/
|
||||
public interface ResourcePorts {
|
||||
|
||||
/** The process environment. Production: {@link System#getenv()}. */
|
||||
Map<String, String> environment();
|
||||
|
||||
/**
|
||||
* Connect a herdr client bound to {@code socketPath}. Production returns a real
|
||||
* {@code UnixSocketHerdrClient} — connection-per-call, so this itself never touches the socket.
|
||||
*/
|
||||
HerdrClient connectHerdr(Path socketPath);
|
||||
|
||||
/** The reply-inbox AMQP opener (CB-307). Production: {@link Fleetd#replyInboxOpener()}. */
|
||||
Fleetd.AmqpOpener replyInboxOpener();
|
||||
|
||||
/** The lead-mailbox AMQP opener (CB-637). Production: {@link Fleetd#leadMailboxOpener()}. */
|
||||
Fleetd.LeadMailboxOpener leadMailboxOpener();
|
||||
|
||||
/** A monotonic elapsed-time clock. Production: {@link System#nanoTime()}. */
|
||||
LongSupplier nanoClock();
|
||||
|
||||
/**
|
||||
* A wall-clock reading, in nanoseconds. Production: {@code System.currentTimeMillis()}
|
||||
* converted to nanoseconds. Kept separate from {@link #nanoClock()} because {@link
|
||||
* dev.ltms.fleet.health.FleetHealthMonitor} needs both — one monotonic clock for elapsed-time
|
||||
* decisions, one wall clock to detect and correct for a macOS sleep freezing the monotonic one.
|
||||
*/
|
||||
LongSupplier wallClockNanos();
|
||||
|
||||
/** A dedicated single-thread scheduler; production names its (virtual) thread {@code purpose}. */
|
||||
ScheduledExecutorService newScheduler(String purpose);
|
||||
|
||||
/** Register a JVM shutdown hook that runs {@code hook} on JVM exit. */
|
||||
void addShutdownHook(Runnable hook);
|
||||
|
||||
/** Bind and start the HTTP server. */
|
||||
void startHttp(Javalin app, String host, int port);
|
||||
|
||||
/** The real production ports: a live herdr socket, a live broker, real threads, a real bind. */
|
||||
static ResourcePorts system() {
|
||||
return new SystemResourcePorts();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,66 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.UnixSocketHerdrClient;
|
||||
import io.javalin.Javalin;
|
||||
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
/**
|
||||
* fleetd #612 Unit A: the one production {@link ResourcePorts} — every method here is the exact
|
||||
* call {@code Fleetd.main} used to make directly at each of these sites before this ticket.
|
||||
* Package-private: obtained only through {@link ResourcePorts#system()}.
|
||||
*/
|
||||
final class SystemResourcePorts implements ResourcePorts {
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return System.getenv();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return UnixSocketHerdrClient.connect(socketPath, new ObjectMapper());
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return Fleetd.replyInboxOpener();
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return Fleetd.leadMailboxOpener();
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return () -> TimeUnit.MILLISECONDS.toNanos(System.currentTimeMillis());
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor(r -> Thread.ofVirtual().name(purpose).unstarted(r));
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
Runtime.getRuntime().addShutdownHook(new Thread(hook));
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
app.start(host, port);
|
||||
}
|
||||
}
|
||||
@@ -221,7 +221,8 @@ public final class CallerResolver {
|
||||
// The config/live binding names this pane as an architect slot's own. Same
|
||||
// unforgeable pane mapping; the live binding, never a request argument, decides.
|
||||
// Check the slot role too: this defence in depth prevents a bad lifecycle bind from
|
||||
// escalating a dev or reviewer into an architect. Checked before the worker fallback.
|
||||
// escalating a dev, hunter or reviewer into an architect. Checked before
|
||||
// the worker fallback.
|
||||
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
|
||||
}
|
||||
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
|
||||
|
||||
@@ -42,7 +42,8 @@ public interface MemberLifecycle {
|
||||
* Try to bind a newly spawned {@code terminal} into the role it was granted.
|
||||
*
|
||||
* @return the role this session actually holds: {@code role} unchanged for a role with no
|
||||
* live slot-binding semantics (dev, reviewer), or when the bind succeeded; a fallback
|
||||
* live slot-binding semantics (dev, hunter, reviewer), or when the bind
|
||||
* succeeded; a fallback
|
||||
* role — never {@code role} — when a slot-bound role (architect) could not be bound.
|
||||
* Callers must record THIS value on the session, never the requested {@code role}, so
|
||||
* a later roster read never reports a role the session does not hold (CB-619). In
|
||||
|
||||
@@ -20,7 +20,8 @@ import java.util.function.Supplier;
|
||||
*
|
||||
* <p>Two halves, split by who owns each:
|
||||
* <ul>
|
||||
* <li><b>slots</b> — read from {@code fleet.architects}/{@code developers}/{@code reviewers}
|
||||
* <li><b>slots</b> — read from {@code fleet.architects}/{@code developers}/
|
||||
* {@code hunters}/{@code reviewers}
|
||||
* (see {@link #slots()}), each carrying the {@code profile} reference the spawn lifecycle
|
||||
* reads when it stands the slot up. <strong>Live, since fleetd #424</strong>: {@link #live}
|
||||
* re-reads {@code fleet:} on every call, through a supplier the same shape as
|
||||
@@ -322,7 +323,7 @@ public final class MemberRegistry implements MemberLifecycle {
|
||||
* CB-619 / fleetd #123: refuse an architect acquire before anything spawns when no configured
|
||||
* slot carries {@code profile} — the config-gap case from the original defect report (a spawn
|
||||
* asked for {@code role=architect, profile=sonnet}, and {@code fleet.architects} carried only
|
||||
* {@code opus} and {@code sol}). A dev/reviewer acquire is always a no-op: those pools are
|
||||
* {@code opus} and {@code sol}). A dev/hunter/reviewer acquire is always a no-op: those pools are
|
||||
* placement candidates only (see {@code CompositePeerLauncher}), never a live identity binding,
|
||||
* so there is nothing here to refuse — an explicit profile outside the pool for those roles is a
|
||||
* documented operator override, not a defect.
|
||||
|
||||
@@ -31,7 +31,7 @@ import java.util.function.Supplier;
|
||||
* {@code placement:}, and an existing profile's {@code weight} / {@code maxLoad}. Both are
|
||||
* read through a supplier on {@code CompositePeerLauncher}, which is what makes them hot —
|
||||
* not the fact that they are config. Most of {@code fleet:} — every role pool
|
||||
* ({@code architects}/{@code developers}/{@code reviewers}), {@code charters}, and
|
||||
* ({@code architects}/{@code developers}/{@code hunters}/{@code reviewers}), {@code charters}, and
|
||||
* {@code tabLabel} — is read the same live way, through the same supplier
|
||||
* ({@code () -> config.get().fleet()}). {@code architects} in particular is hot for
|
||||
* <strong>two independent consumers</strong> (fleetd #424): {@code CompositePeerLauncher}
|
||||
@@ -208,7 +208,7 @@ import java.util.function.Supplier;
|
||||
* that is correctly hot and for a key nobody triaged. Three times now — {@code worktreeGroup} (#323),
|
||||
* {@code primary}/{@code configReload} (#326), and {@code fleet.leaders} sitting in the escape hatch
|
||||
* (#333) — the second kind hid among the first. A top-level coverage checker in the
|
||||
* {@link ConfigRefProfileCoverageTest} shape (one level up, over {@code FleetConfig} itself rather
|
||||
* {@code ConfigRefProfileCoverageTest} shape (one level up, over {@code FleetConfig} itself rather
|
||||
* than {@code FleetConfig.Profile}) proves this file's four classes exhaust the record's components
|
||||
* — see {@code ConfigRefTopLevelCoverageTest}. That test proves the record's <em>shape</em> is fully
|
||||
* triaged; it does NOT prove a {@code SPLIT_KEYS}/{@code COLD_KEYS}/{@code DEFERRED_KEYS} member has
|
||||
@@ -605,7 +605,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
+ "opened once and needs a restart; the broker URI env-var name kept out of a "
|
||||
+ "member's environment is read live on every spawn and already applied");
|
||||
}
|
||||
// fleetd #333: unlike health/coordinator above, most of `fleet:` (developers, reviewers,
|
||||
// fleetd #333: unlike health/coordinator above, most of `fleet:` (developers, hunters, reviewers,
|
||||
// charters, tabLabel) is genuinely hot — ConfigRefTest.aHotChangeIsAppliedAndRead-
|
||||
// ThroughGet and aCharterChangeIsHotAndReachesTheLiveConfig prove it reaches the live config
|
||||
// with no restart note. `architects` is hot too, and — since fleetd #424 — hot for BOTH of
|
||||
@@ -630,7 +630,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
+ "identity map and to auto-launch leads, and neither is rebuilt on reload, so a "
|
||||
+ "lead added, removed, or given a new tab: label needs a restart — until then it "
|
||||
+ "stays unrecognised, and a caller from its new tab resolves as a worker, not a "
|
||||
+ "lead; the rest of fleet: (developers, reviewers, charters, tabLabel) is read "
|
||||
+ "lead; the rest of fleet: (developers, hunters, reviewers, charters, tabLabel) is read "
|
||||
+ "live through the supplier on CompositePeerLauncher, and architects is read "
|
||||
+ "live through that same supplier for placement AND through a separate supplier "
|
||||
+ "on MemberRegistry for spawn-time identity — both already applied");
|
||||
|
||||
@@ -62,7 +62,7 @@ import java.util.regex.PatternSyntaxException;
|
||||
* @param fleet who the daemon may run and under which role (CB-557). One block replacing
|
||||
* the former {@code leaders:}, {@code members:}, {@code leadScan:} and
|
||||
* {@code defaultProfile:}. Role is the containing key — {@code leaders},
|
||||
* {@code architects}, {@code developers}, {@code reviewers} — and each entry
|
||||
* {@code architects}, {@code developers}, {@code hunters}, {@code reviewers} — and each entry
|
||||
* names the {@code profiles:} backend it runs on. See {@link Fleet}
|
||||
* @param leadHeartbeat opt-in idle-lead heartbeat (CB-551); {@code null} ⇒ off, and an upgraded
|
||||
* daemon never nudges an idle lead on its own initiative
|
||||
@@ -462,8 +462,8 @@ public record FleetConfig(
|
||||
* profile that does not opt in. Read live off the current config, so it is
|
||||
* HOT: a change takes effect on the next exhaustion classification / spawn,
|
||||
* no restart needed.
|
||||
* @param autoCompactWindow opt-in per-profile token window that forces a spawned member to
|
||||
* auto-compact its context at (Claude Code) or within (opencode) a bound the
|
||||
* @param autoCompactWindow opt-in per-profile token window that forces a launched Claude Code
|
||||
* session to auto-compact its context at, or an opencode session within, a bound the
|
||||
* operator chooses, instead of the backend's own default. {@code null} (the
|
||||
* default) leaves today's behaviour exactly — opencode already forces
|
||||
* {@code compaction.auto: true} unconditionally (CB-523) but has no absolute
|
||||
@@ -1200,6 +1200,7 @@ public record FleetConfig(
|
||||
* @param leaders panes that orchestrate rather than are orchestrated, keyed by lead name
|
||||
* @param architects profiles the {@code architect} role may run on
|
||||
* @param developers profiles the {@code dev} role may run on
|
||||
* @param hunters profiles the {@code hunter} role may run on
|
||||
* @param reviewers profiles the {@code reviewer} role may run on
|
||||
* @param charters optional launch-charter text keyed by singular role wire name
|
||||
* @param tabLabel template for a member tab's label; {@code {role}}, {@code {profile}},
|
||||
@@ -1210,6 +1211,7 @@ public record FleetConfig(
|
||||
public record Fleet(Map<String, Leader> leaders,
|
||||
Map<String, Slot> architects,
|
||||
Map<String, Slot> developers,
|
||||
Map<String, Slot> hunters,
|
||||
Map<String, Slot> reviewers,
|
||||
Map<String, String> charters,
|
||||
String tabLabel) {
|
||||
@@ -1226,6 +1228,7 @@ public record FleetConfig(
|
||||
leaders = unmodifiableOrEmpty(leaders);
|
||||
architects = unmodifiableOrEmpty(architects);
|
||||
developers = unmodifiableOrEmpty(developers);
|
||||
hunters = unmodifiableOrEmpty(hunters);
|
||||
reviewers = unmodifiableOrEmpty(reviewers);
|
||||
charters = unmodifiableOrEmpty(charters);
|
||||
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? DEFAULT_TAB_LABEL : tabLabel;
|
||||
@@ -1240,9 +1243,15 @@ public record FleetConfig(
|
||||
* constructor: the launcher reads {@code fleet.charters()} from the live config. Jackson
|
||||
* binds the canonical constructor, so this one cannot swallow an operator's YAML.
|
||||
*/
|
||||
public Fleet(Map<String, Leader> leaders, Map<String, Slot> architects,
|
||||
Map<String, Slot> developers, Map<String, Slot> reviewers,
|
||||
Map<String, String> charters, String tabLabel) {
|
||||
this(leaders, architects, developers, null, reviewers, charters, tabLabel);
|
||||
}
|
||||
|
||||
public Fleet(Map<String, Leader> leaders, Map<String, Slot> architects,
|
||||
Map<String, Slot> developers, Map<String, Slot> reviewers, String tabLabel) {
|
||||
this(leaders, architects, developers, reviewers, null, tabLabel);
|
||||
this(leaders, architects, developers, null, reviewers, null, tabLabel);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1264,6 +1273,7 @@ public record FleetConfig(
|
||||
return switch (role) {
|
||||
case ARCHITECT -> architects;
|
||||
case DEV -> developers;
|
||||
case HUNTER -> hunters;
|
||||
case REVIEWER -> reviewers;
|
||||
};
|
||||
}
|
||||
@@ -1319,9 +1329,20 @@ public record FleetConfig(
|
||||
* each such nudge costs the lead a turn just to read "nothing pending";
|
||||
* three is enough to tell it it may stand down without nagging forever, and
|
||||
* it is the bound that stops an idle fleet from being a subscription burner.
|
||||
* @param contextHighNudge fleetd #609: when {@code true}, an idle lead whose own {@code
|
||||
* LeadContextGauge} reading is {@code HIGH} gets a text notice telling it
|
||||
* to consider a handover, appended to whatever heartbeat nudge the loop
|
||||
* already sends. Default {@code false} ({@code null} also means off) — an
|
||||
* upgraded daemon must not silently start telling leads to hand over.
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record LeadHeartbeat(Integer idleAfterSeconds, Long backoffMs, Integer quietNudgeCap) {
|
||||
public record LeadHeartbeat(Integer idleAfterSeconds, Long backoffMs, Integer quietNudgeCap,
|
||||
Boolean contextHighNudge) {
|
||||
/** Convenience constructor for every call site that predates fleetd #609: no context notice. */
|
||||
public LeadHeartbeat(Integer idleAfterSeconds, Long backoffMs, Integer quietNudgeCap) {
|
||||
this(idleAfterSeconds, backoffMs, quietNudgeCap, null);
|
||||
}
|
||||
|
||||
public LeadHeartbeat {
|
||||
idleAfterSeconds = (idleAfterSeconds == null || idleAfterSeconds <= 0) ? 300 : idleAfterSeconds;
|
||||
backoffMs = (backoffMs == null || backoffMs <= 0) ? 60_000L : backoffMs;
|
||||
@@ -1670,7 +1691,7 @@ public record FleetConfig(
|
||||
* <p>A herdr pane runs a login shell that re-sources the operator's own secret store, so a
|
||||
* member inherits every credential the operator's shell holds — measured at 31 names on this
|
||||
* host, of which only one ({@code GITEA_ACCESS_TOKEN}) used to be blocked, and that block was a
|
||||
* single name hardcoded in {@link HerdrPeerLauncher} rather than driven by config (gitea issue
|
||||
* single name hardcoded in {@link dev.ltms.fleet.member.HerdrPeerLauncher} rather than driven by config (gitea issue
|
||||
* #82). This record replaces that hardcoded shadow with a config-driven one.
|
||||
*
|
||||
* <p><b>deny-by-default, not a deny-list.</b> A deny-list (block these specific names, let
|
||||
@@ -1854,6 +1875,7 @@ public record FleetConfig(
|
||||
rejectDuplicateMemberSlots(yaml);
|
||||
rejectNegativeMaxLoad(yaml);
|
||||
rejectAutoCompactWindowOutOfRange(yaml);
|
||||
warnConflictingAutoCompactWindows(yaml);
|
||||
rejectMalformedProfilePatterns(yaml);
|
||||
rejectUnknownKind(yaml);
|
||||
rejectUnknownAuthMode(yaml);
|
||||
@@ -1872,7 +1894,7 @@ public record FleetConfig(
|
||||
|
||||
/** The {@code fleet:} child blocks whose direct children are slot names. */
|
||||
private static final Set<String> FLEET_POOL_KEYS =
|
||||
Set.of("leaders", "architects", "developers", "reviewers");
|
||||
Set.of("leaders", "architects", "developers", "hunters", "reviewers");
|
||||
|
||||
/**
|
||||
* Reject a {@code fleet:} role pool whose slot names repeat (CB-548, re-homed by CB-557).
|
||||
@@ -1882,7 +1904,7 @@ public record FleetConfig(
|
||||
* daemon would never know. Jackson's YAML parser does not fail on duplicate mapping keys by
|
||||
* default, so duplicates are caught here, at parse time, before the map is built.
|
||||
*
|
||||
* <p>Only the four pools <em>directly under the top-level {@code fleet:}</em> are considered,
|
||||
* <p>Only the five pools <em>directly under the top-level {@code fleet:}</em> are considered,
|
||||
* and only their direct child keys (the slot names). A nested field elsewhere, even one also
|
||||
* named {@code developers:}, is ignored, so parsing of the rest of the config is unaffected.
|
||||
*
|
||||
@@ -2049,8 +2071,9 @@ public record FleetConfig(
|
||||
"defaultProfile", "a role pool under 'fleet:' — an unqualified spawn now names a role,"
|
||||
+ " and that role's pool supplies the candidate profiles",
|
||||
"architects", "'fleet.architects'",
|
||||
"members", "a role pool under 'fleet:' — 'fleet.architects', 'fleet.developers' or"
|
||||
+ " 'fleet.reviewers'; the role is the containing key, not a 'role:' field",
|
||||
"members", "a role pool under 'fleet:' — 'fleet.architects', 'fleet.developers',"
|
||||
+ " 'fleet.hunters' or 'fleet.reviewers'; the role is the containing key, not"
|
||||
+ " a 'role:' field",
|
||||
"leaders", "'fleet.leaders'",
|
||||
"leadScan", "'fleet.leaders.<name>.tabPrefix' and '.scanIntervalSeconds' — lead"
|
||||
+ " discovery is now configured on the lead it discovers");
|
||||
@@ -2203,6 +2226,73 @@ public record FleetConfig(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Warn (never refuse to start) about a Claude Code profile whose auto-compaction flag and
|
||||
* environment setting disagree.
|
||||
*
|
||||
* <p>Renamed from {@code rejectConflictingAutoCompactWindows} (fleetd #601 review, measured
|
||||
* 2026-09-22): that method threw {@link IllegalStateException}, so {@link #load(Path)} refused
|
||||
* to start on a config carrying this conflict. On this host, four profiles trip it, including
|
||||
* the lead's own profile and the one every worker spawns on — so the throw is not a rare edge
|
||||
* case. Under launchd, a throw inside {@code load()} is a restart loop, not an error an operator
|
||||
* reads once, and the config that would fix it ({@code fleetd/fleetd.yaml}) is gitignored, so
|
||||
* the cause is invisible on the host where it bites. A WARN gives the operator the same
|
||||
* information — which profiles, and now both values, so they can fix it without reading the
|
||||
* source — without ever taking the fleet down.
|
||||
*
|
||||
* <p>fleetd #618 measured which of the two inputs Claude Code actually follows when they
|
||||
* disagree: the environment variable wins, so {@code autoCompactWindow} is inert on a profile
|
||||
* that also sets the env var. This method only detects and reports the disagreement — it does
|
||||
* not correct it — see {@link dev.ltms.fleet.launch.ClaudeCodeArguments} for the full measured
|
||||
* precedence.
|
||||
*
|
||||
* <p>Equal values never warn: either input then produces the same session window, so there is
|
||||
* nothing to reconcile.
|
||||
*/
|
||||
static void warnConflictingAutoCompactWindows(String yaml) {
|
||||
Map<?, ?> raw;
|
||||
try {
|
||||
raw = YAML.readValue(yaml, Map.class);
|
||||
} catch (IOException | IllegalArgumentException e) {
|
||||
return;
|
||||
}
|
||||
if (raw == null || !(raw.get("profiles") instanceof Map<?, ?> profiles)) {
|
||||
return;
|
||||
}
|
||||
List<String> names = new ArrayList<>();
|
||||
List<String> detail = new ArrayList<>();
|
||||
for (Map.Entry<?, ?> entry : profiles.entrySet()) {
|
||||
if (!(entry.getValue() instanceof Map<?, ?> profile)
|
||||
|| !(profile.get("autoCompactWindow") instanceof Number window)
|
||||
|| !(profile.get("env") instanceof Map<?, ?> env)
|
||||
|| !env.containsKey("CLAUDE_CODE_AUTO_COMPACT_WINDOW")) {
|
||||
continue;
|
||||
}
|
||||
Object kind = profile.get("kind");
|
||||
boolean claudeCode = kind == null || String.valueOf(kind).isBlank()
|
||||
|| Profile.KIND_CLAUDE_CODE.equalsIgnoreCase(String.valueOf(kind));
|
||||
Object envValue = env.get("CLAUDE_CODE_AUTO_COMPACT_WINDOW");
|
||||
if (claudeCode && !String.valueOf(window).equals(String.valueOf(envValue))) {
|
||||
String name = String.valueOf(entry.getKey());
|
||||
names.add(name);
|
||||
detail.add(name + " (autoCompactWindow=" + window
|
||||
+ ", env.CLAUDE_CODE_AUTO_COMPACT_WINDOW=" + envValue + ")");
|
||||
}
|
||||
}
|
||||
if (names.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
names.sort(String::compareTo);
|
||||
detail.sort(String::compareTo);
|
||||
log.warn("Claude Code profile(s) {} set disagreeing autoCompactWindow and env."
|
||||
+ "CLAUDE_CODE_AUTO_COMPACT_WINDOW — the daemon starts anyway: {}. fleetd "
|
||||
+ "#618 measured that CLAUDE_CODE_AUTO_COMPACT_WINDOW wins, so "
|
||||
+ "autoCompactWindow is inert on these profiles. Set equal values on each "
|
||||
+ "to resolve this — do not just delete the env var, since that LOWERS the "
|
||||
+ "live window to autoCompactWindow's value rather than fixing anything.",
|
||||
names, String.join(", ", detail));
|
||||
}
|
||||
|
||||
/**
|
||||
* Reject a profile whose {@code errorPattern} (fleetd #201 Unit 5) or {@code exhaustedPattern}
|
||||
* (CB-578 stage A) is not a valid Java regex, naming the profile, the key, and the parser's own
|
||||
|
||||
@@ -50,7 +50,7 @@ import java.util.regex.Pattern;
|
||||
* so the scrape's herdr round-trip never stalls the status poller. The captured waiter is read on the
|
||||
* poller thread (before any next-turn delivery can overwrite it) and passed into the virtual thread.
|
||||
*/
|
||||
public final class CompletionResolver implements TurnListener {
|
||||
public final class CompletionResolver implements TurnListener, TurnRegistrar {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(CompletionResolver.class);
|
||||
|
||||
@@ -235,6 +235,26 @@ public final class CompletionResolver implements TurnListener {
|
||||
this.worktreeBranches = Objects.requireNonNull(worktreeBranches, "worktreeBranches");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #556: {@link TurnRegistrar}'s structural half of what {@link #captureBaseline} used to
|
||||
* do alone — record the waiter, no I/O. Called directly and unconditionally by the {@link
|
||||
* Injector} as part of delivering a turn, before {@link #onDelivered} ever runs, so this
|
||||
* registration cannot be skipped by a {@link TurnListener} throwing (from {@code onDelivered} or
|
||||
* any other callback). {@link #captureBaseline} still performs this exact check-and-put itself
|
||||
* as well — harmless and idempotent when it runs right after this — so a caller that only wires
|
||||
* the {@link TurnListener} path (e.g. an existing test that never mentions {@link TurnRegistrar})
|
||||
* keeps working unchanged.
|
||||
*/
|
||||
@Override
|
||||
public void register(String target, TurnToken token) {
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = token.waiter();
|
||||
if (waiter == null) {
|
||||
inFlight.remove(target); // no send is waiting on this delivery — nothing to resolve later
|
||||
return;
|
||||
}
|
||||
inFlight.put(target, new InFlight(waiter, null, nowNanos.getAsLong(), token.injectedText()));
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onDelivered(String target, TurnToken token) {
|
||||
// Capture the exact waiter this turn belongs to (CB-116) and snapshot the pane's pre-turn
|
||||
@@ -244,7 +264,14 @@ public final class CompletionResolver implements TurnListener {
|
||||
captureBaseline(target, token);
|
||||
}
|
||||
|
||||
/** Capture the in-flight turn: its waiter and pre-turn baseline (the testable core of {@link #onDelivered}). */
|
||||
/**
|
||||
* Capture the in-flight turn: its waiter and pre-turn baseline (the testable core of {@link
|
||||
* #onDelivered}). fleetd #556: this still does its own full waiter check and {@code inFlight.put}
|
||||
* — the same registration {@link #register} performs — so it keeps working standalone (as every
|
||||
* test calling it directly already does) even though the {@link Injector} now also calls {@link
|
||||
* #register} on its own, earlier and unconditionally. The two writes are idempotent with each
|
||||
* other; only the second (this one) carries a real baseline, since only this one pays for a scrape.
|
||||
*/
|
||||
void captureBaseline(String target, TurnToken token) {
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = token.waiter();
|
||||
if (waiter == null) {
|
||||
|
||||
@@ -2,6 +2,7 @@ package dev.ltms.fleet.inject;
|
||||
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.AgentStatus;
|
||||
import dev.ltms.fleet.herdr.HerdrException;
|
||||
import dev.ltms.fleet.herdr.HerdrRouter;
|
||||
import dev.ltms.fleet.msg.TurnToken;
|
||||
import org.slf4j.Logger;
|
||||
@@ -11,6 +12,7 @@ import java.util.ArrayDeque;
|
||||
import java.util.ArrayList;
|
||||
import java.util.Deque;
|
||||
import java.util.List;
|
||||
import java.util.Objects;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
@@ -45,6 +47,17 @@ import java.util.stream.Collectors;
|
||||
* turn — the pickup-grace path (a turn too fast to sample) unwedges the queue but does not fire
|
||||
* completion, since without a sampled {@code working} there is no trustworthy "the worker just
|
||||
* finished the task" signal to act on.
|
||||
*
|
||||
* <p><strong>Delivery honesty (fleetd #551).</strong> The one external send call
|
||||
* ({@link AgentControl#send}) is irreversible, and its own response can still fail after the text
|
||||
* has already reached the worker's pane — herdr replied (or may have), so the request was already
|
||||
* processed, but the reply itself then failed to parse or carried an error. A queued entry is
|
||||
* polled off the queue and marked {@code ATTEMPTED} <em>before</em> that call is made, not after,
|
||||
* so a failure from the call itself can never be recorded as a confident {@code NOT_DELIVERED} for
|
||||
* text that may already be sitting in the pane. {@code NOT_DELIVERED} stays reserved for the cases
|
||||
* where nothing was ever attempted — the readiness grace expiring, {@link #drop}, or a herdr
|
||||
* {@code *_not_found} error, which the rest of this codebase already treats as a confirmed absence
|
||||
* rather than a merely inconclusive failure (see {@code StatusPoller}, {@code AgentControl}).
|
||||
*/
|
||||
public final class Injector {
|
||||
|
||||
@@ -93,6 +106,11 @@ public final class Injector {
|
||||
private final AgentControl agents;
|
||||
private final HerdrRouter router;
|
||||
private final TurnListener turnListener;
|
||||
/**
|
||||
* fleetd #556: the Injector's own registration invariant, called directly and unconditionally —
|
||||
* never folded into {@link #turnListener}'s fan-out. See {@link TurnRegistrar}.
|
||||
*/
|
||||
private final TurnRegistrar registrar;
|
||||
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
|
||||
private final Consumer<String> forget; // CB-114: clear a gone worker's readiness/presence
|
||||
/**
|
||||
@@ -134,15 +152,43 @@ public final class Injector {
|
||||
*/
|
||||
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
|
||||
Consumer<String> forget) {
|
||||
this(agents, turnListener, ready, forget, System::currentTimeMillis);
|
||||
// fleetd #556: no explicit registrar named here, so fall back to turnListener itself when it
|
||||
// happens to also implement TurnRegistrar (true for CompletionResolver, the only production
|
||||
// TurnListener that needs CB-106 registration) — every existing caller of this overload keeps
|
||||
// working unchanged. A turnListener that does NOT implement TurnRegistrar (a fan-out object,
|
||||
// or a bare test lambda) gets TurnRegistrar.NOOP here, same as before this ticket.
|
||||
this(agents, turnListener, ready, forget,
|
||||
turnListener instanceof TurnRegistrar r ? r : TurnRegistrar.NOOP);
|
||||
}
|
||||
|
||||
/** Full constructor — for tests: an injectable wall-clock supplier (fleetd #501). */
|
||||
/**
|
||||
* fleetd #556: explicit-registrar overload. Use this whenever {@code turnListener} does not
|
||||
* itself implement {@link TurnRegistrar} — e.g. a fan-out object that also notifies unrelated
|
||||
* observers — so registration is wired directly to the one component that owns it, rather than
|
||||
* relying on {@code turnListener} happening to implement both interfaces.
|
||||
*/
|
||||
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
|
||||
Consumer<String> forget, TurnRegistrar registrar) {
|
||||
this(agents, turnListener, ready, forget, registrar, System::currentTimeMillis);
|
||||
}
|
||||
|
||||
/**
|
||||
* Package-private constructor — for tests: an injectable wall-clock supplier (fleetd #501),
|
||||
* with no explicit registrar (fleetd #556 fallback, same as the public 4-arg overload above).
|
||||
*/
|
||||
Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
|
||||
Consumer<String> forget, LongSupplier nowMillis) {
|
||||
this(agents, turnListener, ready, forget,
|
||||
turnListener instanceof TurnRegistrar r ? r : TurnRegistrar.NOOP, nowMillis);
|
||||
}
|
||||
|
||||
/** Full constructor — for tests: an explicit registrar and an injectable wall-clock supplier. */
|
||||
Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
|
||||
Consumer<String> forget, TurnRegistrar registrar, LongSupplier nowMillis) {
|
||||
this.agents = agents;
|
||||
this.router = null;
|
||||
this.turnListener = turnListener;
|
||||
this.registrar = Objects.requireNonNull(registrar, "registrar");
|
||||
this.ready = ready;
|
||||
this.forget = forget;
|
||||
this.nowMillis = nowMillis;
|
||||
@@ -150,15 +196,28 @@ public final class Injector {
|
||||
|
||||
public Injector(HerdrRouter router, TurnListener turnListener, Predicate<String> ready,
|
||||
Consumer<String> forget) {
|
||||
this(router, turnListener, ready, forget, System::currentTimeMillis);
|
||||
// fleetd #556: see the AgentControl overload above for the fallback rationale.
|
||||
this(router, turnListener, ready, forget,
|
||||
turnListener instanceof TurnRegistrar r ? r : TurnRegistrar.NOOP);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #556: explicit-registrar overload — production wiring ({@code Fleetd}) uses this, since
|
||||
* its {@code turnListener} is a fan-out object that also notifies {@code SessionManager} and does
|
||||
* not itself implement {@link TurnRegistrar}.
|
||||
*/
|
||||
public Injector(HerdrRouter router, TurnListener turnListener, Predicate<String> ready,
|
||||
Consumer<String> forget, TurnRegistrar registrar) {
|
||||
this(router, turnListener, ready, forget, registrar, System::currentTimeMillis);
|
||||
}
|
||||
|
||||
/** Full constructor — for tests: an injectable wall-clock supplier (fleetd #501). */
|
||||
Injector(HerdrRouter router, TurnListener turnListener, Predicate<String> ready,
|
||||
Consumer<String> forget, LongSupplier nowMillis) {
|
||||
Consumer<String> forget, TurnRegistrar registrar, LongSupplier nowMillis) {
|
||||
this.agents = null;
|
||||
this.router = router;
|
||||
this.turnListener = turnListener;
|
||||
this.registrar = Objects.requireNonNull(registrar, "registrar");
|
||||
this.ready = ready;
|
||||
this.forget = forget;
|
||||
this.nowMillis = nowMillis;
|
||||
@@ -170,8 +229,23 @@ public final class Injector {
|
||||
|
||||
/** The result of trying to remove an undelivered message from the injector. */
|
||||
public enum Cancellation {
|
||||
/** The message was still queued and this call removed it; the target saw nothing. */
|
||||
CANCELLED,
|
||||
/** The message reached the worker's pane and is confirmed delivered. */
|
||||
DELIVERED,
|
||||
/**
|
||||
* fleetd #551: the injector called {@link AgentControl#send} for this message and that call
|
||||
* threw before its outcome was known — the text may or may not have reached the pane. A
|
||||
* caller reporting this to an operator must say "uncertain", not "definitely not
|
||||
* delivered": reading it as a confident negative invites a resend of text that may already
|
||||
* be sitting in the pane (a double delivery), which is worse than the ambiguity itself.
|
||||
*/
|
||||
ATTEMPTED,
|
||||
/**
|
||||
* The message never reached the worker's pane — nothing was ever attempted for it (the
|
||||
* readiness grace expired, the target was dropped, or send failed with a herdr
|
||||
* {@code *_not_found} error, which is a confirmed absence, not merely inconclusive).
|
||||
*/
|
||||
NOT_DELIVERED
|
||||
}
|
||||
|
||||
@@ -194,7 +268,29 @@ public final class Injector {
|
||||
|
||||
/** A pending message and the future that completes when it has been delivered. */
|
||||
private static final class Pending {
|
||||
enum State { QUEUED, DELIVERED, NOT_DELIVERED, CANCELLED }
|
||||
enum State {
|
||||
/** Still in the target's queue, not yet attempted. */
|
||||
QUEUED,
|
||||
/**
|
||||
* fleetd #551: {@link AgentControl#send} has been called for this entry and its outcome
|
||||
* is not yet known — recorded BEFORE the call (see {@link #onStatus}), so a Throwable
|
||||
* from send() can never leave the entry either still QUEUED or falsely marked
|
||||
* NOT_DELIVERED. Upgraded to DELIVERED on success; left as ATTEMPTED on an ordinary
|
||||
* failure, since reaching the catch does not prove the text never reached the pane.
|
||||
*/
|
||||
ATTEMPTED,
|
||||
/** {@code send()} returned normally: the text is confirmed to have reached the pane. */
|
||||
DELIVERED,
|
||||
/**
|
||||
* Confirmed — not merely inconclusive — that nothing was ever sent for this entry: the
|
||||
* worker never became ready ({@link #onStatus}'s readiness-grace expiry), the target was
|
||||
* dropped ({@link #drop}), or send failed with a herdr {@code *_not_found} error. Never
|
||||
* written for an entry whose send() outcome is unknown; see ATTEMPTED.
|
||||
*/
|
||||
NOT_DELIVERED,
|
||||
/** Removed from the queue by {@link #cancel} before it was ever attempted. */
|
||||
CANCELLED
|
||||
}
|
||||
|
||||
final String target;
|
||||
final String text;
|
||||
@@ -286,7 +382,14 @@ public final class Injector {
|
||||
}
|
||||
|
||||
private static Cancellation cancellationOf(Pending p) {
|
||||
return p.state == Pending.State.DELIVERED ? Cancellation.DELIVERED : Cancellation.NOT_DELIVERED;
|
||||
// fleetd #551 (comment 17037): a two-way split on a now-three-way question folded ATTEMPTED
|
||||
// into NOT_DELIVERED with no compiler error and no failing test — the exact defect this
|
||||
// ticket exists to fix, one layer up. ATTEMPTED gets its own answer instead.
|
||||
return switch (p.state) {
|
||||
case DELIVERED -> Cancellation.DELIVERED;
|
||||
case ATTEMPTED -> Cancellation.ATTEMPTED;
|
||||
case QUEUED, NOT_DELIVERED, CANCELLED -> Cancellation.NOT_DELIVERED;
|
||||
};
|
||||
}
|
||||
|
||||
private static boolean isQuiescent(Target t) {
|
||||
@@ -306,7 +409,7 @@ public final class Injector {
|
||||
if (t == null) return;
|
||||
|
||||
Pending sent = null;
|
||||
RuntimeException sendError = null;
|
||||
Throwable sendError = null;
|
||||
boolean turnCompleted = false;
|
||||
boolean turnFailed = false;
|
||||
boolean resubmit = false;
|
||||
@@ -379,20 +482,43 @@ public final class Injector {
|
||||
if (p != null && ready.test(target)) {
|
||||
t.notReadySincePoll = 0;
|
||||
t.notReadySinceMillis = 0;
|
||||
// fleetd #551: poll and record BEFORE the irreversible send, not after.
|
||||
// The entry comes off the queue and its state is set to ATTEMPTED here,
|
||||
// unconditionally — so a Throwable escaping the send call below (caught or
|
||||
// not) can never leave the entry QUEUED at the head of t.queue (the fleetd
|
||||
// #546 hazard, since peek() alone would let the next onStatus round re-enter
|
||||
// this block and send the same text again), and no path can write a
|
||||
// confident DELIVERED or NOT_DELIVERED before we actually know which one
|
||||
// happened.
|
||||
t.queue.poll();
|
||||
p.state = Pending.State.ATTEMPTED;
|
||||
try {
|
||||
agentsFor(target).send(target, p.text());
|
||||
t.queue.poll();
|
||||
p.state = Pending.State.DELIVERED;
|
||||
t.awaitingPickup = true;
|
||||
t.awaitingCompletion = true;
|
||||
t.turnObserved = false;
|
||||
t.injectableSincePickup = 0;
|
||||
sent = p;
|
||||
} catch (RuntimeException e) {
|
||||
// Delivery failed at herdr; drop the poisoned message and surface it
|
||||
// rather than blocking the queue behind it.
|
||||
t.queue.poll();
|
||||
p.state = Pending.State.NOT_DELIVERED;
|
||||
} catch (Throwable e) {
|
||||
// fleetd #551: leave p.state == ATTEMPTED (recorded above, before the
|
||||
// call) rather than downgrading it to NOT_DELIVERED here — reaching this
|
||||
// catch does not prove the text never reached the pane. Three of the
|
||||
// four HerdrException throw sites in HerdrCodec fire only after herdr
|
||||
// has already replied (so it processed the request), and the fourth (a
|
||||
// transport IOException) leaves it genuinely unknown whether herdr even
|
||||
// received the bytes — see #551 comment 16867. The one exception is a
|
||||
// herdr `*_not_found` error: that family is already read as "definitely
|
||||
// absent, not merely inconclusive" everywhere else in this codebase
|
||||
// (StatusPoller, AgentControl's own retry, WorkspaceControl,
|
||||
// HerdrPeerLauncher, FleetApp, ReplyPushLoop) because it means the
|
||||
// target pane/agent does not exist at all, so nothing could have been
|
||||
// pasted anywhere — #551 keeps the new state consistent with that
|
||||
// existing vocabulary rather than inventing a second one.
|
||||
if (e instanceof HerdrException he && he.code() != null
|
||||
&& he.code().endsWith("_not_found")) {
|
||||
p.state = Pending.State.NOT_DELIVERED;
|
||||
}
|
||||
sent = p;
|
||||
sendError = e;
|
||||
}
|
||||
@@ -489,54 +615,182 @@ public final class Injector {
|
||||
|
||||
// Fire listeners / herdr calls after releasing the monitor so nothing runs on the poller
|
||||
// thread while it holds the target lock.
|
||||
if (resubmit) {
|
||||
try {
|
||||
agentsFor(target).submit(target); // nudge a raced Enter so the pending paste submits
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("resubmit to {} failed (will retry next poll): {}", target, e.getMessage());
|
||||
//
|
||||
// fleetd #553: wrapped in try/finally. By this point, if `sent != null`, the delivery has
|
||||
// already happened inside the monitor above — off the queue, p.state == DELIVERED, text
|
||||
// typed into the target's pane — so `sent`'s future MUST be completed one way or another,
|
||||
// in every path out of this region, or the caller (a blocking fleet_send, or an async
|
||||
// ticket) waits forever on a message it actually received. But the finally must not
|
||||
// swallow whatever escaped: a listener that throws is a defect in THAT listener, and it
|
||||
// must still reach StatusPoller's catch (Throwable) so log.error fires — converting a
|
||||
// loud listener bug into a silently orphaned future would be worse than the bug itself.
|
||||
//
|
||||
// There are TWO futures at stake here, not one: `sent.delivered()` (the delivery future)
|
||||
// and `sent.token().waiter()` (the rendezvous waiter a blocking fleet_send actually waits
|
||||
// on for the worker's ANSWER). fleetd #556: the waiter is registered by `registrar.register`
|
||||
// now — called directly and unconditionally, structural rather than delegated to a listener
|
||||
// callback — never by turnListener.onDelivered() alone (before this ticket, that WAS the
|
||||
// only path, and any TurnListener wired ahead of the one that mattered could throw and skip
|
||||
// it; #553 could only make the reachable listener behave, not remove that dependency).
|
||||
// Completing `sent.delivered()` without also calling `registrar.register` leaves the waiter
|
||||
// unregistered forever — a hang with a success receipt, which is worse than the plain hang
|
||||
// #553 was about. So the `finally` below does not just complete the future: on the path
|
||||
// where the `if (sent != null)` block never ran, it does that block's whole job —
|
||||
// register, then onDelivered (notification only, may throw), then complete.
|
||||
//
|
||||
// `sentHandled` is NOT allowed to move earlier than this: an earlier comment on #553
|
||||
// proposed running `if (sent != null)` first, before turnCompleted/turnFailed, and that
|
||||
// was withdrawn — `registrar.register` WRITES CompletionResolver's inFlight record for this
|
||||
// turn, onTurnComplete READS it for the PREVIOUS turn, and running register first makes
|
||||
// onTurnComplete resolve the NEW turn's waiter with the PREVIOUS turn's stale output
|
||||
// (CB-116). The `if (sent != null)` block stays last; the `finally` is a backstop for it,
|
||||
// not a replacement. fleetd #556 keeps this ordering unchanged — it only splits what used
|
||||
// to be one listener call (register + notify) into two calls in the same position.
|
||||
boolean sentHandled = false;
|
||||
try {
|
||||
if (resubmit) {
|
||||
try {
|
||||
agentsFor(target).submit(target); // nudge a raced Enter so the pending paste submits
|
||||
} catch (Throwable e) {
|
||||
// fleetd #553: widened from RuntimeException (same reasoning as #549 at :382) —
|
||||
// this is the FIRST block after the monitor, so an Error escaping it used to skip
|
||||
// every block below it, including the `sent` completion.
|
||||
log.debug("resubmit to {} failed (will retry next poll): {}", target, e.getMessage());
|
||||
}
|
||||
}
|
||||
}
|
||||
if (notReady != null) {
|
||||
// Worker never became available: forget its (never-set) readiness, unblock every queued
|
||||
// caller, and route the awaiting send through the same failure path as a stalled turn so
|
||||
// a blocking or async waiter resolves WORKER_FAILED rather than riding out the timeout.
|
||||
forget.accept(target);
|
||||
RuntimeException cause = new IllegalStateException(
|
||||
target + " never became available (no bridge MCP connection within the boot window)");
|
||||
for (Pending p : notReady) {
|
||||
p.delivered().completeExceptionally(cause);
|
||||
if (notReady != null) {
|
||||
// Worker never became available: unblock every queued caller FIRST — these messages'
|
||||
// fate (NOT_DELIVERED, off the queue) was already decided inside the monitor above, so
|
||||
// fleetd #553 completes every one of them before calling forget.accept or
|
||||
// turnListener.onTurnFailed below. Either of those is a listener/consumer callback and
|
||||
// can throw (an ordinary RuntimeException is enough — the same reasoning as the rest of
|
||||
// this ticket): completing the futures first means such a throw can no longer leave any
|
||||
// of them permanently pending, regardless of which one throws or in which order.
|
||||
RuntimeException cause = new IllegalStateException(
|
||||
target + " never became available (no bridge MCP connection within the boot window)");
|
||||
for (Pending p : notReady) {
|
||||
p.delivered().completeExceptionally(cause);
|
||||
}
|
||||
// route the awaiting send through the same failure path as a stalled turn so a blocking
|
||||
// or async waiter resolves WORKER_FAILED rather than riding out the timeout.
|
||||
forget.accept(target);
|
||||
turnListener.onTurnFailed(target);
|
||||
}
|
||||
turnListener.onTurnFailed(target);
|
||||
}
|
||||
if (turnCompleted) {
|
||||
if (startPostTurn) {
|
||||
boolean started = turnListener.onTurnCompleteWithPostAction(target);
|
||||
synchronized (t) {
|
||||
t.postTurnPending = false;
|
||||
if (started) {
|
||||
t.awaitingPostTurnPickup = true;
|
||||
t.injectableSincePostTurnPickup = 0;
|
||||
if (turnCompleted) {
|
||||
if (startPostTurn) {
|
||||
// fleetd #553: the listener call is wrapped so `t.postTurnPending` (set true inside
|
||||
// the monitor above, before this call) is always reset. Before this wrapping, a
|
||||
// RuntimeException from onTurnCompleteWithPostAction skipped the reset below,
|
||||
// permanently wedging the target: postTurnPending stayed true forever, so this
|
||||
// target's delivery guard (:376) would never again pass and no further message to it
|
||||
// would ever be delivered — a target-level lockup, not just one skipped future.
|
||||
// `started` defaults to false so an exception is treated as "the post-turn action did
|
||||
// not start" rather than falsely arming the post-turn pickup latch for an action that
|
||||
// never ran.
|
||||
boolean started = false;
|
||||
try {
|
||||
started = turnListener.onTurnCompleteWithPostAction(target);
|
||||
} finally {
|
||||
synchronized (t) {
|
||||
t.postTurnPending = false;
|
||||
if (started) {
|
||||
t.awaitingPostTurnPickup = true;
|
||||
t.injectableSincePostTurnPickup = 0;
|
||||
}
|
||||
if (t.queue.isEmpty() && !t.awaitingPostTurnPickup) {
|
||||
targets.remove(target, t);
|
||||
}
|
||||
}
|
||||
}
|
||||
if (t.queue.isEmpty() && !t.awaitingPostTurnPickup) {
|
||||
targets.remove(target, t);
|
||||
} else {
|
||||
turnListener.onTurnComplete(target);
|
||||
}
|
||||
}
|
||||
if (turnFailed) {
|
||||
turnListener.onTurnFailed(target);
|
||||
}
|
||||
if (sent != null) {
|
||||
// fleetd #553: record the INTENT to handle before any side effect, not the fact of
|
||||
// having completed it. If register()/onDelivered() throws part way through — e.g.
|
||||
// after CompletionResolver's captureBaseline has already done its inFlight.put — a
|
||||
// flag set only after this block would still read false, and the `finally` below
|
||||
// would re-run this block's job a SECOND time. A second captureBaseline runs later,
|
||||
// once onTurnComplete has already thrown and possibly scraped the pane, and can
|
||||
// snapshot a pane that already absorbed this turn's output — which means the pane's
|
||||
// tail never differs from that baseline again and CompletionResolver.resolve's own
|
||||
// suppression at :364-369 (`return` keeping the in-flight record) drops every future
|
||||
// completion for this turn, permanently. So this flag must go true FIRST.
|
||||
sentHandled = true;
|
||||
if (sendError != null) {
|
||||
log.warn("inject to {} failed, dropped message: {}", target, sendError.getMessage());
|
||||
sent.delivered().completeExceptionally(sendError);
|
||||
} else {
|
||||
// fleetd #556: register the waiter FIRST and unconditionally — this is the
|
||||
// Injector's own invariant, kept structural rather than delegated. Even if
|
||||
// turnListener.onDelivered below throws (a bug in an unrelated observer, e.g.
|
||||
// SessionManager's bookkeeping, or a fan-out wired in a different order), the
|
||||
// waiter is already registered and resolvable — a throwing TurnListener can no
|
||||
// longer take this invariant down with it.
|
||||
registrar.register(target, sent.token());
|
||||
// Notification only, from here down: baseline the pane's pre-turn content so a
|
||||
// misattributed completion (no new output) can't resolve this send with the
|
||||
// previous turn's stale answer (CB-115). Allowed to throw and to fail (the herdr
|
||||
// scrape inside already fails open with baseline == null on its own).
|
||||
turnListener.onDelivered(target, sent.token());
|
||||
sent.delivered().complete(null);
|
||||
}
|
||||
}
|
||||
} finally {
|
||||
// Backstop (fleetd #553, extended #556): if an earlier block in this try threw before the
|
||||
// `if (sent != null)` block above ran, `sentHandled` is still false here, and this does
|
||||
// that block's WHOLE job — not just the future completion. Skipping registrar.register()
|
||||
// here would leave sent.token().waiter() never registered with CompletionResolver, so a
|
||||
// blocking fleet_send would be told its message was delivered and then wait out its full
|
||||
// timeout for an answer that can never resolve — worse than the plain hang, because it now
|
||||
// looks like success. This block runs only when sendError == null: nothing was delivered
|
||||
// on the error path, so there is nothing to register.
|
||||
//
|
||||
// fleetd #553 (lead review, comment 16916): `sentHandled` guards ONLY the register/notify
|
||||
// re-call below, never the future completion outside this inner try. One boolean cannot
|
||||
// carry both meanings — "the turn was handled" and "the future has been dealt with" —
|
||||
// because they come apart exactly when onDelivered throws PART WAY THROUGH: sentHandled
|
||||
// is already true (set before the call, correctly — see the `if (sent != null)` block
|
||||
// above), so a guard on the outer `if` here would skip this whole recovery, including the
|
||||
// completion, and leave sent.delivered() pending forever even though the message really
|
||||
// was typed into the pane and taken off the queue. So `sent != null` alone gates whether
|
||||
// this target has anything to finish; `!sentHandled` gates only the register/notify
|
||||
// re-call inside. CompletableFuture.complete/completeExceptionally are idempotent — on the
|
||||
// ordinary path (sentHandled == true, no throw) the `if (sent != null)` block above has
|
||||
// already completed this future, so the calls below are a no-op returning false.
|
||||
//
|
||||
// This recovery is wrapped in its own try/catch(Throwable) that swallows only ITS OWN
|
||||
// throwable and logs at WARN — the future is still completed either way — while the
|
||||
// ORIGINAL throwable from the try block above is left alone to keep unwinding out of
|
||||
// this method to StatusPoller's catch (Throwable), so a listener bug stays loud.
|
||||
if (sent != null) {
|
||||
try {
|
||||
if (!sentHandled && sendError == null) {
|
||||
// fleetd #556: register first here too, same as the ordinary path above — this
|
||||
// recovery path is reached only when an EARLIER block (resubmit/turnCompleted/
|
||||
// turnFailed) threw before the ordinary path ever ran, so nothing has
|
||||
// registered this turn yet.
|
||||
registrar.register(target, sent.token());
|
||||
turnListener.onDelivered(target, sent.token());
|
||||
}
|
||||
} catch (Throwable recoveryError) {
|
||||
log.warn("fleetd #553 backstop: onDelivered failed for {} while recovering from "
|
||||
+ "an earlier listener failure; completing its delivery future "
|
||||
+ "anyway: {}",
|
||||
target, recoveryError.getMessage());
|
||||
} finally {
|
||||
// No-op (returns false) on the ordinary path, where the `if (sent != null)` block
|
||||
// above already completed this future — see the comment above this block.
|
||||
if (sendError != null) {
|
||||
sent.delivered().completeExceptionally(sendError);
|
||||
} else {
|
||||
sent.delivered().complete(null);
|
||||
}
|
||||
}
|
||||
} else {
|
||||
turnListener.onTurnComplete(target);
|
||||
}
|
||||
}
|
||||
if (turnFailed) {
|
||||
turnListener.onTurnFailed(target);
|
||||
}
|
||||
if (sent != null) {
|
||||
if (sendError != null) {
|
||||
log.warn("inject to {} failed, dropped message: {}", target, sendError.getMessage());
|
||||
sent.delivered().completeExceptionally(sendError);
|
||||
} else {
|
||||
// Baseline the pane's pre-turn content so a misattributed completion (no new output)
|
||||
// can't resolve this send with the previous turn's stale answer (CB-115).
|
||||
turnListener.onDelivered(target, sent.token());
|
||||
sent.delivered().complete(null);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,91 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
/**
|
||||
* A progress watchdog for a singleton background loop (fleetd #544): {@link StatusPoller} and
|
||||
* {@code dev.ltms.fleet.session.SessionReaper} each hold one. It tracks the monotonic timestamp
|
||||
* of the loop's last <em>completed</em> round and classifies health from that timestamp alone —
|
||||
* never from thread liveness. That choice is deliberate: measured on {@code main} at
|
||||
* {@code 7611b69}, {@code UnixSocketHerdrClient.call()} reads with no deadline anywhere in
|
||||
* {@code herdr/}, so a herdrd that accepts a connection and never answers parks the calling
|
||||
* loop's thread forever — the thread stays alive and {@code Thread.isAlive()} keeps reporting
|
||||
* {@code true} the whole time. A last-round timestamp that stops advancing catches that parked
|
||||
* case exactly the same way it catches a thread that died outright: either way, nothing marks a
|
||||
* new round complete.
|
||||
*
|
||||
* <p>Lives in {@code inject} (rather than {@code health}, which would be the more obvious home
|
||||
* for a health-reporting fact) because {@code health} already depends on {@code session}
|
||||
* ({@code HealthSnapshot}/{@code FleetHealthMonitor} read {@code MemberSession}), and
|
||||
* {@code session} already depends on {@code inject} ({@code SessionManager} implements
|
||||
* {@code TurnListener} and extends {@code MemberPresence}). Putting this class in {@code health}
|
||||
* would have made {@code SessionReaper} (in {@code session}) depend back on {@code health},
|
||||
* closing a cycle through {@code health -> session -> health} — {@code PackageCyclesTest} catches
|
||||
* exactly this. {@code inject} has no dependency on {@code health}, so {@code session} depending
|
||||
* on {@code inject} for this stays a one-way edge, same direction it already depends in.
|
||||
*
|
||||
* <p>Three states, not two (the fleetd #512 shape — one flag carrying two conditions that need
|
||||
* opposite handling). {@code stop()} also makes the loop go quiet, exactly like a crash or a
|
||||
* hang does, so a caller must say which one happened: {@link #markStoppedByCaller()} records
|
||||
* "this halt was on purpose," and {@link #state()} reports {@link State#STOPPED} for it
|
||||
* regardless of how stale the last round looks — an intentional stop is never an alarm.
|
||||
*/
|
||||
public final class LoopWatchdog {
|
||||
|
||||
/**
|
||||
* {@code RUNNING} — the loop is going and its last round finished recently.
|
||||
* {@code STALLED} — the loop was never stopped on purpose, and its last-completed-round
|
||||
* timestamp is older than the threshold. Covers both a dead loop and a parked one.
|
||||
* {@code STOPPED} — {@link #markStoppedByCaller()} was called. Not an alarm.
|
||||
*/
|
||||
public enum State { RUNNING, STALLED, STOPPED }
|
||||
|
||||
private final LongSupplier nowNanos;
|
||||
private final long staleAfterNanos;
|
||||
private volatile long lastRoundNanos;
|
||||
private volatile boolean stoppedByCaller;
|
||||
|
||||
/**
|
||||
* @param nowNanos a monotonic elapsed-time clock, e.g. {@code System::nanoTime} live, a
|
||||
* controllable stub in tests.
|
||||
* @param staleAfterNanos how long a last-completed-round timestamp may age before {@link #state()}
|
||||
* reports {@link State#STALLED}. Derive this from the loop's own poll
|
||||
* interval with generous headroom — see the caller's own derivation comment.
|
||||
*/
|
||||
public LoopWatchdog(LongSupplier nowNanos, long staleAfterNanos) {
|
||||
this.nowNanos = nowNanos;
|
||||
this.staleAfterNanos = staleAfterNanos;
|
||||
this.lastRoundNanos = nowNanos.getAsLong();
|
||||
}
|
||||
|
||||
/** Call once per completed round, from inside the loop. */
|
||||
public void recordRoundComplete() {
|
||||
lastRoundNanos = nowNanos.getAsLong();
|
||||
}
|
||||
|
||||
/**
|
||||
* Call from {@code start()} (and any future restart): clears a previous stop's mark and
|
||||
* resets the clock, so the loop gets a fresh grace period before its first round completes
|
||||
* rather than being judged against a timestamp left over from before it was (re)started.
|
||||
*/
|
||||
public void reset() {
|
||||
stoppedByCaller = false;
|
||||
lastRoundNanos = nowNanos.getAsLong();
|
||||
}
|
||||
|
||||
/**
|
||||
* Call from {@code stop()}: marks the halt as intentional, so {@link #state()} reports
|
||||
* {@link State#STOPPED} rather than {@link State#STALLED} no matter how stale the last round is.
|
||||
*/
|
||||
public void markStoppedByCaller() {
|
||||
stoppedByCaller = true;
|
||||
}
|
||||
|
||||
/** The current fact about this loop, per the three-state contract above. */
|
||||
public State state() {
|
||||
if (stoppedByCaller) {
|
||||
return State.STOPPED;
|
||||
}
|
||||
return (nowNanos.getAsLong() - lastRoundNanos >= staleAfterNanos) ? State.STALLED : State.RUNNING;
|
||||
}
|
||||
}
|
||||
@@ -8,6 +8,8 @@ import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
/**
|
||||
* Drives the {@link Injector} by sampling each active worker's {@code agent_status} and
|
||||
@@ -22,11 +24,25 @@ public final class StatusPoller {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(StatusPoller.class);
|
||||
|
||||
/**
|
||||
* How many multiples of {@code intervalMillis} the last-completed round may age before
|
||||
* {@link #health()} reports {@link LoopWatchdog.State#STALLED} (fleetd #544). A round this
|
||||
* loop runs is normally a small fraction of one interval — it only takes longer when a herdr
|
||||
* call is genuinely wedged (herdr's {@code UnixSocketHerdrClient} has no read deadline, so a
|
||||
* herdrd that accepts and never answers parks the thread forever) — so a wide multiplier
|
||||
* avoids false alarms from an ordinarily slow round. At the production interval
|
||||
* ({@link Injector#POLL_INTERVAL_MILLIS} = 250ms) this puts the threshold at 10s: long enough
|
||||
* that a transient slow poll never trips it, short enough that a stuck poller is visible well
|
||||
* before anything else downstream (delivery, `/healthz` consumers) would show a symptom.
|
||||
*/
|
||||
private static final long STALE_AFTER_INTERVAL_MULTIPLIER = 40;
|
||||
|
||||
private final AgentControl agents;
|
||||
private final HerdrRouter router;
|
||||
private final Injector injector;
|
||||
private final StatusRefiner refiner;
|
||||
private final long intervalMillis;
|
||||
private final LoopWatchdog watchdog;
|
||||
private volatile boolean running;
|
||||
private Thread thread;
|
||||
|
||||
@@ -36,14 +52,26 @@ public final class StatusPoller {
|
||||
|
||||
public StatusPoller(AgentControl agents, Injector injector, StatusRefiner refiner,
|
||||
long intervalMillis) {
|
||||
this(agents, injector, refiner, intervalMillis, System::nanoTime);
|
||||
}
|
||||
|
||||
/** Full constructor — for tests: an injectable monotonic clock (fleetd #544). */
|
||||
StatusPoller(AgentControl agents, Injector injector, StatusRefiner refiner,
|
||||
long intervalMillis, LongSupplier nowNanos) {
|
||||
this.agents = agents;
|
||||
this.router = null;
|
||||
this.injector = injector;
|
||||
this.refiner = refiner;
|
||||
this.intervalMillis = intervalMillis;
|
||||
this.watchdog = new LoopWatchdog(nowNanos, staleAfterNanos(intervalMillis));
|
||||
}
|
||||
|
||||
public StatusPoller(HerdrRouter router, Injector injector, long intervalMillis) {
|
||||
this(router, injector, intervalMillis, System::nanoTime);
|
||||
}
|
||||
|
||||
/** Full constructor — for tests: an injectable monotonic clock (fleetd #544). */
|
||||
StatusPoller(HerdrRouter router, Injector injector, long intervalMillis, LongSupplier nowNanos) {
|
||||
this.agents = null;
|
||||
this.router = router;
|
||||
this.injector = injector;
|
||||
@@ -53,45 +81,70 @@ public final class StatusPoller {
|
||||
// against the LEAD daemon even though this field points at the member one.
|
||||
this.refiner = new StatusRefiner(router.memberAgents());
|
||||
this.intervalMillis = intervalMillis;
|
||||
this.watchdog = new LoopWatchdog(nowNanos, staleAfterNanos(intervalMillis));
|
||||
}
|
||||
|
||||
private static long staleAfterNanos(long intervalMillis) {
|
||||
return TimeUnit.MILLISECONDS.toNanos(Math.max(intervalMillis, 1)) * STALE_AFTER_INTERVAL_MULTIPLIER;
|
||||
}
|
||||
|
||||
/**
|
||||
* This loop's progress fact (fleetd #544): {@code RUNNING}, {@code STALLED} (dead or parked —
|
||||
* indistinguishable from outside, and never observed on purpose), or {@code STOPPED}
|
||||
* ({@link #stop()} was called). See {@link LoopWatchdog} for why staleness — not thread
|
||||
* liveness — is the signal.
|
||||
*/
|
||||
public LoopWatchdog.State health() {
|
||||
return watchdog.state();
|
||||
}
|
||||
|
||||
/** Start the polling loop on a virtual thread. Idempotent. */
|
||||
public synchronized void start() {
|
||||
if (running) return;
|
||||
running = true;
|
||||
watchdog.reset(); // fleetd #544: a fresh grace period, not last run's stale mark/timestamp.
|
||||
thread = Thread.ofVirtual().name("status-poller").start(this::loop);
|
||||
log.info("status poller started (interval {}ms)", intervalMillis);
|
||||
}
|
||||
|
||||
private void loop() {
|
||||
while (running) {
|
||||
Set<String> active = injector.activeTargets();
|
||||
for (String target : active) {
|
||||
if (!running) return;
|
||||
try {
|
||||
// herdr's agent_status can misreport a settled worker as `unknown`; refine it
|
||||
// against the pane content before it drives delivery/completion (CB-115).
|
||||
// CB-185: refine THROUGH the same control the raw status came from — a router
|
||||
// splits lead/member targets across two herdr daemons, and reading a lead's pane
|
||||
// through the (fixed) member refiner never finds it, wedging that lead at UNKNOWN.
|
||||
AgentControl control = router != null ? router.agentsFor(target) : agents;
|
||||
AgentStatus status = refiner.refine(target, control.status(target), control);
|
||||
injector.onStatus(target, status);
|
||||
} catch (HerdrException e) {
|
||||
// The worker's agent is gone — stop trying and unblock its waiters.
|
||||
if (e.code() != null && e.code().endsWith("_not_found")) {
|
||||
log.debug("target {} gone; dropping its queue", target);
|
||||
injector.drop(target, e);
|
||||
} else {
|
||||
log.debug("status poll for {} failed (will retry): {}", target, e.getMessage());
|
||||
try {
|
||||
while (running) {
|
||||
Set<String> active = injector.activeTargets();
|
||||
for (String target : active) {
|
||||
if (!running) return;
|
||||
try {
|
||||
// herdr's agent_status can misreport a settled worker as `unknown`; refine it
|
||||
// against the pane content before it drives delivery/completion (CB-115).
|
||||
// CB-185: refine THROUGH the same control the raw status came from — a router
|
||||
// splits lead/member targets across two herdr daemons, and reading a lead's pane
|
||||
// through the (fixed) member refiner never finds it, wedging that lead at UNKNOWN.
|
||||
AgentControl control = router != null ? router.agentsFor(target) : agents;
|
||||
AgentStatus status = refiner.refine(target, control.status(target), control);
|
||||
injector.onStatus(target, status);
|
||||
} catch (HerdrException e) {
|
||||
// The worker's agent is gone — stop trying and unblock its waiters.
|
||||
if (e.code() != null && e.code().endsWith("_not_found")) {
|
||||
log.debug("target {} gone; dropping its queue", target);
|
||||
injector.drop(target, e);
|
||||
} else {
|
||||
log.debug("status poll for {} failed (will retry): {}", target, e.getMessage());
|
||||
}
|
||||
} catch (Throwable e) {
|
||||
log.error("unexpected failure polling {}; skipping this round", target, e);
|
||||
}
|
||||
} catch (RuntimeException e) {
|
||||
// Never let one target's unexpected error (e.g. an odd agent.get shape) kill
|
||||
// the single poller thread and stall injection for every worker.
|
||||
log.warn("unexpected error polling {}; skipping this round", target, e);
|
||||
}
|
||||
// fleetd #544: a round is "complete" only once every active target has been polled —
|
||||
// a target parked mid-for-loop (control.status(target) never returning) means this
|
||||
// line is never reached, so the watchdog goes stale exactly like a dead loop would.
|
||||
watchdog.recordRoundComplete();
|
||||
sleep();
|
||||
}
|
||||
sleep();
|
||||
} finally {
|
||||
if (running) {
|
||||
log.error("status poller loop exited unexpectedly; it can be restarted");
|
||||
}
|
||||
running = false;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -107,6 +160,7 @@ public final class StatusPoller {
|
||||
/** Stop the polling loop. Idempotent. */
|
||||
public synchronized void stop() {
|
||||
running = false;
|
||||
watchdog.markStoppedByCaller(); // fleetd #544: this halt is on purpose — health() must say STOPPED, not STALLED.
|
||||
if (thread != null) thread.interrupt();
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,29 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import dev.ltms.fleet.msg.TurnToken;
|
||||
|
||||
/**
|
||||
* fleetd #556: the {@link Injector}'s own invariant — every delivered turn has a registered
|
||||
* waiter — kept structural rather than delegated to a {@link TurnListener} callback that can
|
||||
* throw. The {@link Injector} calls {@link #register} directly and unconditionally as part of
|
||||
* delivering a turn, before it ever calls {@code turnListener.onDelivered}. A {@link TurnListener}
|
||||
* that throws from every callback (a bug in an unrelated observer, e.g. session bookkeeping, or a
|
||||
* fan-out object wired in the wrong order) can therefore never leave a delivered turn unregistered
|
||||
* — the shape #553 could not close, since that fix could only make the reachable listener behave,
|
||||
* never remove the Injector's dependency on a listener behaving at all.
|
||||
*
|
||||
* <p>Deliberately narrow: registration only, no I/O, nothing that scrapes a pane or can reasonably
|
||||
* fail. The CB-115 staleness baseline (a herdr round-trip, allowed to fail) stays a
|
||||
* {@link TurnListener#onDelivered} concern — fired by the {@link Injector} only after this call has
|
||||
* already run, so its own failure cannot un-register anything.
|
||||
*/
|
||||
@FunctionalInterface
|
||||
public interface TurnRegistrar {
|
||||
|
||||
/** Record {@code token}'s waiter as the turn currently in flight for {@code target}. */
|
||||
void register(String target, TurnToken token);
|
||||
|
||||
/** No-op registrar for callers that don't need CB-106 rendezvous tracking. */
|
||||
TurnRegistrar NOOP = (_, _) -> {
|
||||
};
|
||||
}
|
||||
@@ -0,0 +1,37 @@
|
||||
package dev.ltms.fleet.launch;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
|
||||
/** Arguments shared by every fleetd path that starts Claude Code. */
|
||||
public final class ClaudeCodeArguments {
|
||||
|
||||
private ClaudeCodeArguments() {
|
||||
}
|
||||
|
||||
/**
|
||||
* Append the configured Claude Code auto-compaction window when the profile opts in.
|
||||
*
|
||||
* <p>This flag and the environment variable {@code CLAUDE_CODE_AUTO_COMPACT_WINDOW} can
|
||||
* disagree, and fleetd #618 measured which one Claude Code actually follows: the environment
|
||||
* variable wins, ahead of this {@code --autocompact} flag, ahead of the settings file, ahead of
|
||||
* clientdata, the experiment, and the model default. So when a profile sets both, the flag this
|
||||
* method appends has NO effect — Claude Code reads {@code CLAUDE_CODE_AUTO_COMPACT_WINDOW}
|
||||
* first and never consults the flag. {@link FleetConfig#load(java.nio.file.Path)} only WARNS
|
||||
* when a Claude Code profile sets both to different values (see {@code
|
||||
* FleetConfig.warnConflictingAutoCompactWindows}) — it does not stop the daemon from starting,
|
||||
* and the launched session honours the env var, not this flag. Measured against Claude Code
|
||||
* 2.1.278 (fleetd #618) — a later version could reorder this precedence.
|
||||
*/
|
||||
public static List<String> withAutoCompactWindow(List<String> argv, FleetConfig.Profile profile) {
|
||||
if (profile.autoCompactWindow() == null) {
|
||||
return argv;
|
||||
}
|
||||
List<String> withAutoCompact = new ArrayList<>(argv);
|
||||
withAutoCompact.add("--autocompact");
|
||||
withAutoCompact.add(String.valueOf(profile.autoCompactWindow()));
|
||||
return withAutoCompact;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,313 @@
|
||||
package dev.ltms.fleet.lead;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.io.RandomAccessFile;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.atomic.AtomicInteger;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
/**
|
||||
* Reads how full a lead's own Claude Code context window is, from the transcript Claude Code
|
||||
* itself writes — never from the lead's pane (fleetd has {@code AgentControl.read} for that, and
|
||||
* this must not use it: a pane holds terminal text, not the structured usage numbers a transcript
|
||||
* carries, and scraping it would also race the lead's own rendering).
|
||||
*
|
||||
* <p><strong>Why this exists.</strong> A lead auto-compacts when its context fills — on the host
|
||||
* this was built for, that happened 30 times in one session, discarding roughly 250,000 tokens
|
||||
* and costing 46 seconds to 3 minutes each time, and fleetd had no way to see it coming. This
|
||||
* class is the first thing that looks.
|
||||
*
|
||||
* <p><strong>The route.</strong> Claude Code appends one JSON object per line to
|
||||
* {@code <configDir>/projects/<slug>/<sessionId>.jsonl}. {@code <slug>} is an undocumented,
|
||||
* internal encoding of the working directory — this class never derives it. Instead it lists the
|
||||
* one-level-deep subdirectories of {@code <configDir>/projects/} and looks for
|
||||
* {@code <sessionId>.jsonl} by name, so the slug rule can change without breaking this reader.
|
||||
*
|
||||
* <ul>
|
||||
* <li><strong>Live context</strong> is read off the last record in the read window that carries
|
||||
* a {@code message.usage} object: {@code input_tokens + cache_read_input_tokens +
|
||||
* cache_creation_input_tokens}. This is what actually fills the window — a plain
|
||||
* {@code input_tokens} count alone understates it once the conversation has any cached
|
||||
* prefix, which on a long-lived lead is always.</li>
|
||||
* <li><strong>Compaction history</strong> is a count of {@code subtype: "compact_boundary"}
|
||||
* records seen in the same read window — see {@link Reading#compactions()}. It is a count
|
||||
* within the window this reader actually looked at, not a lifetime total: a session with
|
||||
* more compactions than fit in {@link #TAIL_BYTES} of transcript will undercount. That
|
||||
* trade-off is deliberate — see {@link #TAIL_BYTES}.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p><strong>Three states, not two (OK / HIGH / UNKNOWN).</strong> Every path that cannot
|
||||
* positively establish the live token count — a missing file, an unreadable one, a peer that
|
||||
* is not a Claude backend, or every line in the read window failing to parse as JSON — returns
|
||||
* {@link State#UNKNOWN} with no token number, never a default "0" or "ok" that would read as
|
||||
* "this lead is fine" when the honest answer is "I could not look".
|
||||
*
|
||||
* <p><strong>A torn final line does not mean UNKNOWN.</strong> {@code fleet_list} reads this
|
||||
* transcript while Claude Code may be mid-write on it, so the last line in the window can be cut
|
||||
* off mid-flush — that is an ordinary, expected race, not a sign the format has changed. Earlier
|
||||
* this class treated ANY unparseable last line as UNKNOWN, on the theory that "if the format
|
||||
* changes, we should see UNKNOWN". That reasoning does not hold: a real format change makes
|
||||
* <em>every</em> line in the window unparseable, not only the last one written. So a single
|
||||
* malformed line (most often the final, torn one, but the check is not position-specific) is
|
||||
* simply skipped rather than treated as fatal, and the reading is built from whatever lines in the
|
||||
* window did parse. Only when <em>none</em> of them parse — the real format-change signal — does
|
||||
* this return {@link State#UNKNOWN}, still with no stale number standing in for "I could not
|
||||
* tell".
|
||||
*
|
||||
* <p><strong>Bounded cost.</strong> {@code fleet_list} is polled constantly, so every read is
|
||||
* capped two ways: {@link #TAIL_BYTES} bounds how much of the transcript is ever read from disk
|
||||
* (never the whole 52 MB a long-lived transcript reaches on the host this was measured on), and
|
||||
* {@link #DEFAULT_CACHE_TTL_MILLIS} bounds how often that bounded read actually happens — a burst
|
||||
* of {@code fleet_list} calls inside one TTL window reads the file once. One instance's cache is
|
||||
* keyed by {@code (configDir, sessionId)}, so it is safe to share across every lead a single
|
||||
* {@code fleet_list} call reports on.
|
||||
*/
|
||||
public final class LeadContextGauge {
|
||||
|
||||
private static final String PROJECTS_DIR = "projects";
|
||||
private static final ObjectMapper MAPPER = new ObjectMapper();
|
||||
|
||||
/**
|
||||
* How many trailing bytes of a transcript a single read ever pulls off disk. Chosen so one
|
||||
* read comfortably spans many recent turns — each usage or compact_boundary record is at most
|
||||
* a few KB — while staying nowhere near the 52 MB a long session's real transcript reaches on
|
||||
* the host this was built for; reading that whole file on every {@code fleet_list} call is
|
||||
* exactly the cost this bound exists to avoid. 2 MiB holds on the order of hundreds of recent
|
||||
* lines even when a turn's tool output is unusually large, which is far more than needed to
|
||||
* find the most recent usage record and any recent compaction.
|
||||
*/
|
||||
static final int TAIL_BYTES = 2 * 1024 * 1024;
|
||||
|
||||
/**
|
||||
* How long a {@link Reading} is served from cache before the file is read again.
|
||||
* {@code fleet_list} is called constantly (by design — it is the fleet's own status probe), so
|
||||
* without a TTL a burst of calls would re-read the transcript tail once per call. 5 seconds is
|
||||
* short enough that a caller watching for a state change never waits long, and long enough that
|
||||
* a poll loop calling every second or two only touches disk once per window.
|
||||
*/
|
||||
static final long DEFAULT_CACHE_TTL_MILLIS = 5_000;
|
||||
|
||||
/**
|
||||
* Live tokens at or above this count report {@link State#HIGH}. On the host this was measured
|
||||
* on, auto-compaction actually fires around 267,000–270,000 tokens, but the point of a HIGH
|
||||
* state is to warn before that happens, not at it — 200,000 is the standard Claude context
|
||||
* window size and a sensible built-in default: no config key is required to pick it, and a
|
||||
* lead crossing it is already deep enough into its window that a compaction is foreseeable.
|
||||
*/
|
||||
static final long HIGH_THRESHOLD_TOKENS = 200_000;
|
||||
|
||||
/** The only peer kind this reader understands ({@code Agent.agentType()}'s wire value). */
|
||||
private static final String CLAUDE_AGENT_TYPE = "claude";
|
||||
|
||||
public enum State { OK, HIGH, UNKNOWN }
|
||||
|
||||
/**
|
||||
* @param state {@link State#UNKNOWN} whenever {@code tokens} could not be established
|
||||
* @param tokens live context tokens, or {@code null} exactly when {@code state} is
|
||||
* {@link State#UNKNOWN}
|
||||
* @param compactions {@code compact_boundary} records seen in the read window (see class
|
||||
* javadoc) — {@code 0} both for "genuinely none seen" and for "unknown",
|
||||
* since a caller that already sees {@code state: UNKNOWN} has no reason to
|
||||
* trust this number either way
|
||||
*/
|
||||
public record Reading(State state, Long tokens, int compactions) {
|
||||
/**
|
||||
* fleetd #609: widened from package-private to public so {@code
|
||||
* dev.ltms.fleet.msg.LeadHeartbeatLoop.LeadContextSource.none()} (a different package) can
|
||||
* return the same inert "I could not look" reading the gauge itself uses, without inventing
|
||||
* a parallel unknown-reading constant. Behaviour of this class is otherwise unchanged.
|
||||
*/
|
||||
public static Reading unknown() {
|
||||
return new Reading(State.UNKNOWN, null, 0);
|
||||
}
|
||||
}
|
||||
|
||||
private record CacheEntry(Reading reading, long readAtMillis) {
|
||||
}
|
||||
|
||||
private final LongSupplier clock;
|
||||
private final long ttlMillis;
|
||||
private final Map<String, CacheEntry> cache = new ConcurrentHashMap<>();
|
||||
/** Test seam only (package-private) — counts real disk reads, i.e. cache misses. */
|
||||
private final AtomicInteger diskReads = new AtomicInteger();
|
||||
|
||||
public LeadContextGauge() {
|
||||
this(System::currentTimeMillis, DEFAULT_CACHE_TTL_MILLIS);
|
||||
}
|
||||
|
||||
/** Test seam: an injectable clock and TTL so cache expiry is provable without sleeping. */
|
||||
LeadContextGauge(LongSupplier clock, long ttlMillis) {
|
||||
this.clock = clock;
|
||||
this.ttlMillis = ttlMillis;
|
||||
}
|
||||
|
||||
/** How many times this instance has actually read a transcript off disk — test seam only. */
|
||||
int diskReadCount() {
|
||||
return diskReads.get();
|
||||
}
|
||||
|
||||
/**
|
||||
* @param configDir the lead's {@code CLAUDE_CONFIG_DIR}, or {@code null}/blank to use the
|
||||
* default {@code <user.home>/.claude} — the right answer for the common case
|
||||
* where the lead's profile sets no {@code configDir} override
|
||||
* @param sessionId the lead's own Claude session id ({@code Agent.sessionId()}), or
|
||||
* {@code null} when herdr has not resolved one yet
|
||||
* @param agentType the detected peer kind ({@code Agent.agentType()}); anything other than
|
||||
* {@code "claude"} (including {@code null}, meaning undetected) reports
|
||||
* {@link State#UNKNOWN} — this reader only understands Claude Code's own
|
||||
* transcript format
|
||||
*/
|
||||
public Reading read(String configDir, String sessionId, String agentType) {
|
||||
if (sessionId == null || sessionId.isBlank()) {
|
||||
return Reading.unknown();
|
||||
}
|
||||
if (!CLAUDE_AGENT_TYPE.equalsIgnoreCase(agentType)) {
|
||||
return Reading.unknown();
|
||||
}
|
||||
String base = (configDir == null || configDir.isBlank())
|
||||
? System.getProperty("user.home") + "/.claude"
|
||||
: configDir;
|
||||
String cacheKey = base + '\u0000' + sessionId;
|
||||
long now = clock.getAsLong();
|
||||
CacheEntry cached = cache.get(cacheKey);
|
||||
if (cached != null && now - cached.readAtMillis() < ttlMillis) {
|
||||
return cached.reading();
|
||||
}
|
||||
Reading fresh = readUncached(base, sessionId);
|
||||
cache.put(cacheKey, new CacheEntry(fresh, now));
|
||||
return fresh;
|
||||
}
|
||||
|
||||
private Reading readUncached(String base, String sessionId) {
|
||||
diskReads.incrementAndGet();
|
||||
Path file = findTranscript(base, sessionId);
|
||||
if (file == null) {
|
||||
return Reading.unknown();
|
||||
}
|
||||
TailRead tail;
|
||||
try {
|
||||
tail = tailBytes(file, TAIL_BYTES);
|
||||
} catch (IOException e) {
|
||||
return Reading.unknown();
|
||||
}
|
||||
return parse(tail);
|
||||
}
|
||||
|
||||
/**
|
||||
* Finds {@code <sessionId>.jsonl} under {@code <base>/projects/}, one level deep — never by
|
||||
* deriving the slug directory from a working directory (see class javadoc). Bounded to a
|
||||
* single {@code list()} of {@code projects/} itself: it never recurses further, so the cost is
|
||||
* the number of project directories, not the size of any transcript inside them.
|
||||
*/
|
||||
private Path findTranscript(String base, String sessionId) {
|
||||
Path projectsDir = Path.of(base, PROJECTS_DIR);
|
||||
if (!Files.isDirectory(projectsDir)) {
|
||||
return null;
|
||||
}
|
||||
String filename = sessionId + ".jsonl";
|
||||
Path direct = projectsDir.resolve(filename);
|
||||
if (Files.isRegularFile(direct)) {
|
||||
return direct;
|
||||
}
|
||||
try (var children = Files.list(projectsDir)) {
|
||||
return children.filter(Files::isDirectory)
|
||||
.map(dir -> dir.resolve(filename))
|
||||
.filter(Files::isRegularFile)
|
||||
.findFirst()
|
||||
.orElse(null);
|
||||
} catch (IOException e) {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/** Package-private (not {@code private}): {@link #tailBytes} is a test seam, see its javadoc. */
|
||||
record TailRead(byte[] bytes, boolean fromStart) {
|
||||
}
|
||||
|
||||
/**
|
||||
* Reads at most {@code maxBytes} trailing bytes of {@code file}. Package-private (not
|
||||
* {@code private}) so a test can assert directly on the returned array's length — "bytes
|
||||
* actually read", not on any parsed answer — without needing a file anywhere near
|
||||
* {@link #TAIL_BYTES} in size to prove the cap holds.
|
||||
*/
|
||||
static TailRead tailBytes(Path file, int maxBytes) throws IOException {
|
||||
try (RandomAccessFile raf = new RandomAccessFile(file.toFile(), "r")) {
|
||||
long length = raf.length();
|
||||
long start = Math.max(0, length - maxBytes);
|
||||
raf.seek(start);
|
||||
byte[] buf = new byte[(int) (length - start)];
|
||||
raf.readFully(buf);
|
||||
return new TailRead(buf, start == 0);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Parses the tail into a {@link Reading}. The first line is dropped unconditionally whenever
|
||||
* the tail is not the whole file (it starts mid-line, cut by {@link #TAIL_BYTES} — an expected
|
||||
* artefact of the bound, not a data problem). Every remaining line is then parsed on a
|
||||
* best-effort basis: a line that fails to parse (most often the last one, torn by a write this
|
||||
* read raced — see the "torn final line" section of the class javadoc) is skipped, not fatal.
|
||||
* Only when none of the remaining lines parse does this report {@link State#UNKNOWN}.
|
||||
*/
|
||||
private Reading parse(TailRead tail) {
|
||||
String text = new String(tail.bytes(), StandardCharsets.UTF_8);
|
||||
List<String> lines = new ArrayList<>(List.of(text.split("\n", -1)));
|
||||
if (!lines.isEmpty() && lines.get(lines.size() - 1).isEmpty()) {
|
||||
lines.remove(lines.size() - 1); // trailing newline leaves a phantom empty last element
|
||||
}
|
||||
if (!tail.fromStart() && !lines.isEmpty()) {
|
||||
lines.remove(0); // first line is a fragment cut by our own tail bound, not real data
|
||||
}
|
||||
if (lines.isEmpty()) {
|
||||
return Reading.unknown();
|
||||
}
|
||||
Long tokens = null;
|
||||
int compactions = 0;
|
||||
boolean anyLineParsed = false;
|
||||
for (String line : lines) {
|
||||
JsonNode node = tryParse(line);
|
||||
if (node == null) {
|
||||
// A malformed line — typically the last one, cut mid-flush by a write this read
|
||||
// raced — is skipped rather than treated as fatal. See the class javadoc's "torn
|
||||
// final line" section for why: a real format change makes EVERY line unparseable,
|
||||
// not only this one, and that case is still caught below by anyLineParsed.
|
||||
continue;
|
||||
}
|
||||
anyLineParsed = true;
|
||||
JsonNode usage = node.path("message").path("usage");
|
||||
if (usage.isObject()) {
|
||||
tokens = usage.path("input_tokens").asLong(0)
|
||||
+ usage.path("cache_read_input_tokens").asLong(0)
|
||||
+ usage.path("cache_creation_input_tokens").asLong(0);
|
||||
}
|
||||
if ("compact_boundary".equals(node.path("subtype").asText(null))) {
|
||||
compactions++;
|
||||
}
|
||||
}
|
||||
if (!anyLineParsed) {
|
||||
return Reading.unknown();
|
||||
}
|
||||
if (tokens == null) {
|
||||
return new Reading(State.UNKNOWN, null, compactions);
|
||||
}
|
||||
State state = tokens >= HIGH_THRESHOLD_TOKENS ? State.HIGH : State.OK;
|
||||
return new Reading(state, tokens, compactions);
|
||||
}
|
||||
|
||||
private JsonNode tryParse(String line) {
|
||||
try {
|
||||
return MAPPER.readTree(line);
|
||||
} catch (IOException e) {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -8,6 +8,7 @@ import dev.ltms.fleet.herdr.PendingCloseMarker;
|
||||
import dev.ltms.fleet.herdr.Tab;
|
||||
import dev.ltms.fleet.herdr.Workspace;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.launch.ClaudeCodeArguments;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
@@ -358,7 +359,8 @@ public final class LeadLauncher {
|
||||
}
|
||||
|
||||
/**
|
||||
* The lead's argv: the profile's own command, the model pin, and the bridge MCP mount.
|
||||
* The lead's argv: the profile's own command, the model and auto-compaction pins, and the bridge
|
||||
* MCP mount.
|
||||
*
|
||||
* <p>No {@code --append-system-prompt}. That flag carries the worker reply charter, and a lead
|
||||
* is not a worker — it reads its orchestration rules from the project's {@code CLAUDE.md} like
|
||||
@@ -380,7 +382,7 @@ public final class LeadLauncher {
|
||||
argv.add("--model");
|
||||
argv.add(profile.model());
|
||||
}
|
||||
return argv;
|
||||
return profile.isOpenCode() ? argv : ClaudeCodeArguments.withAutoCompactWindow(argv, profile);
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -10,6 +10,8 @@ import java.io.IOException;
|
||||
import java.io.UncheckedIOException;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Collections;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.Map;
|
||||
import java.util.UUID;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
@@ -25,9 +27,11 @@ import java.util.function.Supplier;
|
||||
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation
|
||||
* clears the lead's own pane and bootstraps a fresh session against that file.
|
||||
*
|
||||
* <p>This is the executor only. Nothing in this ticket wires an MCP tool onto {@link #open}/
|
||||
* {@link #confirm}/{@link #cancel} — that is a separate, later unit; until it lands, nothing calls
|
||||
* this class at all.
|
||||
* <p>This is the executor behind the {@code fleet_handover} MCP tool ({@code
|
||||
* dev.ltms.fleet.mcp.FleetMcp#handover}), which drives {@link #open}, {@link #confirm}, {@link
|
||||
* #cancel}, and {@link #status} from a tool call — wired in fleetd #480 Unit C. <strong>An earlier
|
||||
* version of this paragraph said nothing called this class at all; that stopped being true once
|
||||
* that unit landed, and this correction exists so the javadoc does not go on claiming it.</strong>
|
||||
*
|
||||
* <p><strong>{@code confirm()} cannot roll inline — a fleetd #480 correction.</strong> The first
|
||||
* version of this class called {@code agents.send(lead, "/clear")} directly from inside {@code
|
||||
@@ -181,6 +185,106 @@ public final class LeadRollover {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* How many tokens {@link #outcomes} remembers before it starts evicting the oldest — bounded
|
||||
* so a long-running daemon never grows this map without limit. Chosen generously rather than
|
||||
* tightly: production rolls are rare (this class's own ticket found exactly ONE completed roll
|
||||
* ever logged on this host), and each entry is a handful of short strings, so even a full cap
|
||||
* costs a few tens of kilobytes — nowhere near a reason to make it configurable. 200 entries
|
||||
* comfortably outlasts any operator's own memory of "did that roll I asked for actually
|
||||
* happen", which is the whole reason {@link #status} exists.
|
||||
*
|
||||
* <p><strong>This cap counts {@link RollState#IN_PROGRESS} entries exactly the same as
|
||||
* finished ones.</strong> There is only the one bounded map: {@link #confirm} writes an {@link
|
||||
* RollState#IN_PROGRESS} entry into {@link #outcomes} at hand-off, and the deferred
|
||||
* continuation later overwrites that SAME key with a terminal state — it never inserts a
|
||||
* second entry. An approved roll therefore occupies one slot in this map for its entire
|
||||
* lifetime, from the moment {@link #confirm} hands off, not only once it finishes; a
|
||||
* confirmed-but-not-yet-finished roll counts against the cap exactly like a finished one. The
|
||||
* alternative (a separate, uncapped in-flight map) would let a burst of confirmed-but-stuck
|
||||
* rolls grow without bound — the exact failure this cap exists to prevent — so it was rejected.
|
||||
*/
|
||||
static final int OUTCOME_HISTORY_CAP = 200;
|
||||
|
||||
/**
|
||||
* What is known about one token, right now — the answer {@link #status} gives. Distinguishes
|
||||
* three terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
|
||||
* that has been approved but has not finished yet, and two answers for a token that names no
|
||||
* active work at all: still pending confirmation, or nothing known about this token at all.
|
||||
*/
|
||||
public enum RollState {
|
||||
/**
|
||||
* {@code token} is still open: either {@link #open} was called and {@link #confirm} has not
|
||||
* been (or not successfully) yet, or a {@link #confirm} call failed one of its gate checks
|
||||
* and left the token pending for a retry — see {@link #confirm}'s javadoc ("token stays
|
||||
* pending"). Indistinguishable from a genuinely fresh request; a caller wanting to know
|
||||
* WHICH gate most recently refused should read the {@link RollDecision} that {@link
|
||||
* #confirm} itself returned, not this status. <strong>Never the state of an APPROVED
|
||||
* roll</strong> — see {@link #IN_PROGRESS}, which {@link #confirm} records at the moment it
|
||||
* hands off, before this token is even removed from the pending set.
|
||||
*/
|
||||
PENDING,
|
||||
/**
|
||||
* {@link #confirm} approved this roll and handed it to the deferred continuation, which has
|
||||
* not finished yet. Recorded by {@link #confirm} itself, at hand-off — <strong>before</strong>
|
||||
* {@code token} is removed from the pending set — so there is never a gap in which {@link
|
||||
* #status} could wrongly answer {@link #UNKNOWN} ("nothing was ever requested") for a roll
|
||||
* that is, in fact, actively running. This is not sticky: the deferred continuation
|
||||
* overwrites this same entry with a terminal state ({@link #ROLLED}, {@link
|
||||
* #TURN_NEVER_SETTLED}, {@link #CLEAR_NEVER_SETTLED}, or {@link #FAILED}) once it finishes
|
||||
* — including by throwing, which fleetd #615's catch in {@link #runRollover} now turns into
|
||||
* {@link #FAILED} instead of leaving this entry stuck forever.
|
||||
*/
|
||||
IN_PROGRESS,
|
||||
/**
|
||||
* {@link #confirm} was approved and the deferred continuation completed the entire roll:
|
||||
* the calling lead's turn settled, {@code /clear} was sent and settled, and {@code
|
||||
* bootstrapText} was sent.
|
||||
*/
|
||||
ROLLED,
|
||||
/**
|
||||
* {@link #confirm} was approved, but the calling lead's own turn never reached a boundary
|
||||
* (IDLE or DONE) within {@code turnSettleSeconds} — no {@code /clear} was ever sent, at
|
||||
* all. This is the branch the fleetd #480 correction exists to make safe, and the one this
|
||||
* status exists to make VISIBLE: before this, a lead that hit this case had no way to find
|
||||
* out, and would carry on believing it was about to be replaced. See this class's javadoc.
|
||||
*/
|
||||
TURN_NEVER_SETTLED,
|
||||
/**
|
||||
* {@link #confirm} was approved and {@code /clear} was sent, but the pane never re-settled
|
||||
* within {@code clearSettleSeconds} — {@code bootstrapText} was never sent.
|
||||
*/
|
||||
CLEAR_NEVER_SETTLED,
|
||||
/**
|
||||
* fleetd #615: the deferred continuation threw a {@link RuntimeException} — most likely a
|
||||
* {@link dev.ltms.fleet.herdr.HerdrException} out of one of the two unwrapped {@code
|
||||
* agents.send} calls in {@link #runRollover} — and the continuation thread died with it.
|
||||
* Before this state existed, that throw left {@link #outcomes} holding {@link #IN_PROGRESS}
|
||||
* forever, because the production {@code continuationRunner} is a bare virtual thread with
|
||||
* no uncaught-exception handler and nothing downstream of the throw ever ran to write a
|
||||
* terminal outcome. {@code detail} names the exception, so a reader has something to act on
|
||||
* — the same diagnostic style as {@link #TURN_NEVER_SETTLED} and {@link
|
||||
* #CLEAR_NEVER_SETTLED}. The roll is dead at this point and does not retry itself; a stuck
|
||||
* lead must {@link #open} a fresh request.
|
||||
*/
|
||||
FAILED,
|
||||
/**
|
||||
* {@code token} names nothing this instance currently knows about: never issued by {@link
|
||||
* #open}, dropped by {@link #cancel}, or aged out of {@link #outcomes}'s bounded history.
|
||||
* These three causes are not distinguished — all of them mean "there is nothing to tell
|
||||
* you", which is the entire content of a clean answer here.
|
||||
*/
|
||||
UNKNOWN
|
||||
}
|
||||
|
||||
/**
|
||||
* The answer {@link #status} gives for one token: a {@link RollState} and a human-readable
|
||||
* {@code detail}. For {@link RollState#TURN_NEVER_SETTLED}, {@code detail} names {@code
|
||||
* turnSettleSeconds} and its configured value explicitly, so a reader who sees this knows what
|
||||
* to raise.
|
||||
*/
|
||||
public record RollStatus(RollState state, String detail) {}
|
||||
|
||||
private final AgentControl agents;
|
||||
private final Supplier<FleetConfig.LeadRollover> configSupplier;
|
||||
/**
|
||||
@@ -201,6 +305,23 @@ public final class LeadRollover {
|
||||
*/
|
||||
private final Consumer<Runnable> continuationRunner;
|
||||
private final Map<String, PendingRollover> pending = new ConcurrentHashMap<>();
|
||||
/**
|
||||
* Finished tokens → what actually happened, for {@link #status}. Bounded by {@link
|
||||
* #OUTCOME_HISTORY_CAP}, oldest evicted first ({@code removeEldestEntry} on an insertion-order
|
||||
* {@link LinkedHashMap}). Wrapped in {@link Collections#synchronizedMap} because entries are
|
||||
* written from whatever thread {@code continuationRunner} runs the roll on (a fresh virtual
|
||||
* thread in production, the calling test thread under {@code Runnable::run}) and read from
|
||||
* whatever thread calls {@link #status} (the MCP handler thread) — a plain {@code
|
||||
* LinkedHashMap} is not safe for that, and {@code removeEldestEntry} additionally requires
|
||||
* external synchronization even for a thread-safe map that merely wraps it.
|
||||
*/
|
||||
private final Map<String, RollStatus> outcomes = Collections.synchronizedMap(
|
||||
new LinkedHashMap<>(16, 0.75f, false) {
|
||||
@Override
|
||||
protected boolean removeEldestEntry(Map.Entry<String, RollStatus> eldest) {
|
||||
return size() > OUTCOME_HISTORY_CAP;
|
||||
}
|
||||
});
|
||||
|
||||
/** Production constructor — wall clock, real sleep between settle polls, a real virtual thread. */
|
||||
public LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
|
||||
@@ -366,6 +487,15 @@ public final class LeadRollover {
|
||||
return docCheck;
|
||||
}
|
||||
|
||||
// Record IN_PROGRESS BEFORE removing from `pending` — see RollState#IN_PROGRESS and
|
||||
// OUTCOME_HISTORY_CAP's javadoc. This ordering means `token` is written into `outcomes`
|
||||
// while it is STILL present in `pending`; status() checks `outcomes` first (see that
|
||||
// method), so it reports IN_PROGRESS immediately, not the brief-but-real gap a
|
||||
// remove-then-put ordering would leave in which the token is in neither map.
|
||||
outcomes.put(token, new RollStatus(RollState.IN_PROGRESS,
|
||||
"confirm() approved this roll and handed it to the deferred continuation; it has "
|
||||
+ "not finished yet — still waiting for the calling turn to settle, for "
|
||||
+ "/clear to be sent and settle, or for bootstrapText to be sent"));
|
||||
pending.remove(token);
|
||||
log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends",
|
||||
token, callerTerminal);
|
||||
@@ -378,8 +508,42 @@ public final class LeadRollover {
|
||||
* entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the
|
||||
* four-step order. There is no result to return to by this point, so every outcome is logged
|
||||
* only.
|
||||
*
|
||||
* <p><strong>fleetd #615 — the whole body is wrapped in one {@code try}.</strong> The two {@code
|
||||
* agents.send} calls below are not wrapped individually: {@code send} → {@code agentCall} →
|
||||
* {@code herdr.call} can throw an unchecked {@link dev.ltms.fleet.herdr.HerdrException} (see
|
||||
* {@code AgentControl.java}), and the production {@code continuationRunner} is a bare virtual
|
||||
* thread with no uncaught-exception handler (see this class's public constructor). Before this
|
||||
* fix, either throw killed the continuation thread silently, leaving the {@link
|
||||
* RollState#IN_PROGRESS} entry {@link #confirm} wrote at hand-off stuck forever — {@link
|
||||
* #status} had no way to tell a dead roll from one still genuinely running. The {@code catch}
|
||||
* below is scoped to the method body rather than to each {@code send} call individually, so it
|
||||
* also covers anything else added to this continuation later, not just today's two call sites —
|
||||
* the same reasoning that put the write-a-terminal-outcome step at each of this method's other
|
||||
* exits (see the {@link RollState#TURN_NEVER_SETTLED} and {@link RollState#CLEAR_NEVER_SETTLED}
|
||||
* branches below) rather than inside the helpers that detect them.</p>
|
||||
*
|
||||
* <p>Only {@link RuntimeException} is caught, matching the local convention {@link
|
||||
* #waitUntilAtTurnBoundary} already set around its own {@code agents.status} call — not the
|
||||
* broader {@link Exception} or {@link Throwable}, which would also swallow something like an
|
||||
* {@link OutOfMemoryError} this continuation has no business handling.</p>
|
||||
*/
|
||||
private void runRollover(PendingRollover p, FleetConfig.LeadRollover cfg) {
|
||||
try {
|
||||
runRolloverUnguarded(p, cfg);
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("lead-rollover: continuation for token={} lead={} threw {} — the roll is dead; "
|
||||
+ "no further step in this continuation will run",
|
||||
p.token(), p.leadTerminal(), e.toString(), e);
|
||||
outcomes.put(p.token(), new RollStatus(RollState.FAILED,
|
||||
"the roll's continuation threw " + e.toString() + " — the roll is dead and will "
|
||||
+ "not retry itself; check the daemon log for the stack trace, then open() "
|
||||
+ "a fresh rollover request"));
|
||||
}
|
||||
}
|
||||
|
||||
/** The actual body of {@link #runRollover}, unwrapped — see that method's javadoc for the catch. */
|
||||
private void runRolloverUnguarded(PendingRollover p, FleetConfig.LeadRollover cfg) {
|
||||
String lead = p.leadTerminal();
|
||||
long rollStartMillis = nowMillis.getAsLong();
|
||||
TurnSettleResult turnResult = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds());
|
||||
@@ -393,6 +557,11 @@ public final class LeadRollover {
|
||||
+ "turn is still live and clearing it now would destroy live context "
|
||||
+ "(token={}, configured={}s elapsed={}ms)",
|
||||
lead, p.token(), cfg.turnSettleSeconds(), turnResult.elapsedMillis());
|
||||
outcomes.put(p.token(), new RollStatus(RollState.TURN_NEVER_SETTLED,
|
||||
"the calling lead's own turn never reached a boundary (IDLE or DONE) within "
|
||||
+ "turnSettleSeconds=" + cfg.turnSettleSeconds() + "s (measured elapsed="
|
||||
+ turnResult.elapsedMillis() + "ms) — no /clear was ever sent. If this "
|
||||
+ "keeps happening, raise turnSettleSeconds in fleetd.yaml"));
|
||||
return;
|
||||
}
|
||||
|
||||
@@ -411,11 +580,17 @@ public final class LeadRollover {
|
||||
+ "elapsed={}ms nudges={})",
|
||||
lead, p.token(), cfg.clearSettleSeconds(), clearResult.elapsedMillis(),
|
||||
clearResult.nudges());
|
||||
outcomes.put(p.token(), new RollStatus(RollState.CLEAR_NEVER_SETTLED,
|
||||
"/clear was sent, but the pane never re-settled within clearSettleSeconds="
|
||||
+ cfg.clearSettleSeconds() + "s (measured elapsed=" + clearResult.elapsedMillis()
|
||||
+ "ms, nudges=" + clearResult.nudges() + ") — bootstrapText was never sent"));
|
||||
return;
|
||||
}
|
||||
agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()));
|
||||
long rollElapsedMillis = nowMillis.getAsLong() - rollStartMillis;
|
||||
log.info("lead-rollover: rolled token={} lead={} elapsedMs={}", p.token(), lead, rollElapsedMillis);
|
||||
outcomes.put(p.token(), new RollStatus(RollState.ROLLED,
|
||||
"rolled successfully in " + rollElapsedMillis + "ms"));
|
||||
}
|
||||
|
||||
/** Drop a pending request without rolling. @return whether a pending request existed for {@code token} */
|
||||
@@ -423,6 +598,47 @@ public final class LeadRollover {
|
||||
return pending.remove(token) != null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Read-only: what is currently known about {@code token}. <strong>Never sends anything, never
|
||||
* schedules, cancels, or retries a roll</strong> — a caller may poll this as often as it likes
|
||||
* with no side effect at all, which is exactly why it exists: every failure past {@link
|
||||
* #confirm} used to be a {@code log.warn} a lead can never read (see this class's javadoc), and
|
||||
* this is the only route back.
|
||||
*
|
||||
* @param token the token {@link #open} returned; {@code null} or blank is a clean {@link
|
||||
* RollState#UNKNOWN}, never a {@link NullPointerException} — {@link #pending} is a
|
||||
* {@link ConcurrentHashMap}, which throws on a {@code null} key lookup, so this
|
||||
* short-circuits before ever reaching it
|
||||
* @return {@link RollState#IN_PROGRESS} for an approved roll whose continuation has not
|
||||
* finished yet, or a terminal state once it has (both read from {@link #outcomes} —
|
||||
* checked FIRST, see below); {@link RollState#PENDING} while {@code token} is still
|
||||
* open and has not yet been approved (including one left pending by a {@link #confirm}
|
||||
* gate refusal — see that method's javadoc); or {@link RollState#UNKNOWN} for a token
|
||||
* never issued, cancelled, or aged out of the bounded history
|
||||
*/
|
||||
public RollStatus status(String token) {
|
||||
if (token == null || token.isBlank()) {
|
||||
return new RollStatus(RollState.UNKNOWN, "no token given");
|
||||
}
|
||||
// `outcomes` is checked BEFORE `pending`, deliberately: `confirm` writes an IN_PROGRESS
|
||||
// entry into `outcomes` before it removes `token` from `pending` (see `confirm`'s own
|
||||
// comment at that call site), so for the brief window where a token is present in BOTH
|
||||
// maps, this order reports the more accurate answer (IN_PROGRESS, already approved) rather
|
||||
// than the stale one (PENDING, not yet approved) a pending-first check would give.
|
||||
RollStatus recorded = outcomes.get(token);
|
||||
if (recorded != null) {
|
||||
return recorded;
|
||||
}
|
||||
if (pending.containsKey(token)) {
|
||||
return new RollStatus(RollState.PENDING, "open() has been called for this token and "
|
||||
+ "it has not yet been confirmed — or a confirm() gate check failed and left it "
|
||||
+ "pending, so the same token may be retried once the problem is fixed");
|
||||
}
|
||||
return new RollStatus(RollState.UNKNOWN, "token names no pending or finished rollover "
|
||||
+ "request known to this instance — never issued, cancelled, or aged out of the "
|
||||
+ "bounded history (cap=" + OUTCOME_HISTORY_CAP + ")");
|
||||
}
|
||||
|
||||
/**
|
||||
* The three handover-file checks, in order: exists, not empty, fresh (modified after
|
||||
* {@link #open}'s timestamp and not older than {@code maxDocAgeSeconds}). Stats {@code
|
||||
|
||||
@@ -31,6 +31,20 @@ public final class ConnectionIdentity {
|
||||
this.cwds = cwds;
|
||||
}
|
||||
|
||||
/**
|
||||
* The {@link PaneLocator} this identity resolves callers against — fleetd #612 CB-185: lets a
|
||||
* test drive the exact {@link PaneLocator} a real assembly wired up (e.g. {@code
|
||||
* FleetdAssembly}'s {@code new ConnectionIdentity(new PaneLocator(herdr, memberHerdr), ...)})
|
||||
* directly with a chosen pid, bypassing the OS-dependent {@link PeerPidLookup} that {@link
|
||||
* #resolve} otherwise goes through. A full HTTP round trip cannot exercise this: {@code
|
||||
* LsofPeerPidLookup} excludes its own pid, and an in-process test client and server share one
|
||||
* JVM pid, so {@code pidForLocalPort} always returns {@code -1} and {@link PaneLocator} never
|
||||
* gets called at all.
|
||||
*/
|
||||
public PaneLocator panes() {
|
||||
return panes;
|
||||
}
|
||||
|
||||
/**
|
||||
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
|
||||
* primary / an off-host client), its {@code pid} (or {@code -1} if not resolvable), and whether
|
||||
|
||||
@@ -12,6 +12,7 @@ import dev.ltms.fleet.metrics.Metrics;
|
||||
import dev.ltms.fleet.inject.MemberPresence;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.herdr.HerdrException;
|
||||
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||
import dev.ltms.fleet.lead.LeadRollover;
|
||||
import dev.ltms.fleet.msg.LeadChannel;
|
||||
import dev.ltms.fleet.msg.LeadMessage;
|
||||
@@ -20,6 +21,7 @@ import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.peer.PeerUnreachableException;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.inject.LoopWatchdog;
|
||||
import dev.ltms.fleet.placement.PlacementException;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
@@ -105,14 +107,24 @@ public final class FleetMcp {
|
||||
* untested identity heuristic.
|
||||
*/
|
||||
private final boolean authorizationEnforced;
|
||||
/**
|
||||
* fleetd #612 CB-185: kept as a field (rather than only captured by the {@code
|
||||
* contextExtractor} closure built in the constructor) so a test can reach the exact {@link
|
||||
* ConnectionIdentity} — and, through {@link ConnectionIdentity#panes()}, the exact {@link
|
||||
* dev.ltms.fleet.herdr.PaneLocator} — that a real assembly wired up. See {@link #identity()}.
|
||||
*/
|
||||
private final ConnectionIdentity identity;
|
||||
private final Metrics metrics; // CB-502: null → auth failures not counted
|
||||
private final CapacitySource capacity;
|
||||
private final HealthCoverageSource healthCoverage;
|
||||
private final LoopHealthSource loopHealth;
|
||||
private final QuarantineSource quarantine;
|
||||
/** fleetd #201 Unit 5: SEPARATE from {@link #quarantine} — see {@link OutageSource}'s doc. */
|
||||
private final OutageSource outage;
|
||||
/** fleetd #176: SEPARATE from both of the above — see {@link LeadSeatSource}'s doc. */
|
||||
private final LeadSeatSource leadSeats;
|
||||
/** fleetd #602 gauge-wiring: see {@link LeadConfigDirSource}. */
|
||||
private final LeadConfigDirSource leadConfigDirs;
|
||||
/** CB-637: this daemon's lead-to-lead channel; {@code null} when no coordinator is configured. */
|
||||
private final LeadChannel leadChannel;
|
||||
/** fleetd #361: {@code coordinator.peers} — see {@link CoordinationSource}. Empty when unset. */
|
||||
@@ -124,6 +136,17 @@ public final class FleetMcp {
|
||||
* clean {@code NOT_CONFIGURED} refusal rather than throwing. See {@link #handover}.
|
||||
*/
|
||||
private final LeadRollover leadRollover;
|
||||
/**
|
||||
* "Lead context gauge": how full each lead's own Claude Code context window is, reported on
|
||||
* {@code fleet_list}'s {@code leads} rows (see {@link #contextView}). Built unconditionally, in
|
||||
* the field initializer rather than a constructor parameter — this is not a togglable feature
|
||||
* with an on/off config knob the way {@link OutageSource}/{@link LeadSeatSource} are: it needs
|
||||
* no config at all (see {@link LeadContextGauge}'s own javadoc for the built-in defaults), so
|
||||
* there is no "off" value to thread through every existing constructor call site. One instance
|
||||
* per daemon so its read cache (keyed by session, TTL'd) is actually shared across
|
||||
* {@code fleet_list} calls rather than rebuilt — and therefore useless — on every call.
|
||||
*/
|
||||
private final LeadContextGauge leadContextGauge = new LeadContextGauge();
|
||||
|
||||
/** Capacity facts used by {@code fleet_list}; production must supply the placement live count. */
|
||||
public record CapacitySource(Function<String, Integer> liveCount, Function<String, Integer> maxLoad,
|
||||
@@ -136,6 +159,15 @@ public final class FleetMcp {
|
||||
/** Coverage is supplied by the health wiring, not inferred from a missing dependency. */
|
||||
public record HealthCoverageSource(Supplier<String> value) { }
|
||||
|
||||
/** Progress states for fleetd's singleton background loops, read by {@code fleet_list} and {@code /healthz}. */
|
||||
public record LoopHealthSource(Supplier<LoopWatchdog.State> statusPoller,
|
||||
Supplier<LoopWatchdog.State> sessionReaper) {
|
||||
/** Inert source for callers that do not wire the background loops. */
|
||||
public static LoopHealthSource none() {
|
||||
return new LoopHealthSource(() -> LoopWatchdog.State.STOPPED, () -> LoopWatchdog.State.STOPPED);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-578 stage B quarantine facts used by {@code fleet_profiles}: a profile → credential id
|
||||
* lookup, plus the shared {@link BackendQuarantine} to read remaining cooldowns off.
|
||||
@@ -235,6 +267,29 @@ public final class FleetMcp {
|
||||
public static LeadSeatSource none() { return new LeadSeatSource(_ -> 0); }
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #602 gauge-wiring: a lead name's configured {@code CLAUDE_CONFIG_DIR} override, fed to
|
||||
* {@link LeadContextGauge#read} so {@code fleet_list}'s {@code context} row reads the transcript
|
||||
* directory the lead's own profile actually writes to, not always the built-in
|
||||
* {@code <user.home>/.claude} default.
|
||||
*
|
||||
* <p>The same idiom as {@link LeadSeatSource} — a {@code FleetMcp} constructor field, not a
|
||||
* lookup {@code contextView} performs itself, because {@code FleetMcp} holds no
|
||||
* {@link dev.ltms.fleet.config.FleetConfig} and {@code contextView} is {@code static}. See
|
||||
* {@code Fleetd.leadConfigDirLookup} for the derivation: the same
|
||||
* {@code fleet.leaders.<name>.profile} link {@link LeadSeatSource} already follows, resolved to
|
||||
* that profile's own {@code configDir:}.
|
||||
*
|
||||
* @param configDirFor lead name → {@code configDir}, or {@code null} when the lead's entry names
|
||||
* no profile, or that profile sets no {@code configDir} override — either
|
||||
* way {@link LeadContextGauge} then falls back to its own built-in default,
|
||||
* exactly as before this ticket
|
||||
*/
|
||||
public record LeadConfigDirSource(Function<String, String> configDirFor) {
|
||||
/** Inert source — every lead reads {@link LeadContextGauge}'s built-in default {@code configDir}. */
|
||||
public static LeadConfigDirSource none() { return new LeadConfigDirSource(_ -> null); }
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #361: peer-visibility facts for {@code fleet_list}'s {@code coordinator} row — this
|
||||
* daemon's own {@link LeadChannel} (for its self mailbox state and held messages) plus the
|
||||
@@ -328,21 +383,36 @@ public final class FleetMcp {
|
||||
* instead of throwing. See {@link #handover}.
|
||||
*/
|
||||
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
|
||||
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
|
||||
CallerResolver callers, AuthorizationMode authorizationMode, Metrics metrics,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage,
|
||||
LeadSeatSource leadSeats, List<String> peers, LeadRollover leadRollover) {
|
||||
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
|
||||
CallerResolver callers, AuthorizationMode authorizationMode, Metrics metrics,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage,
|
||||
LeadSeatSource leadSeats, List<String> peers, LeadRollover leadRollover) {
|
||||
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, authorizationMode, metrics,
|
||||
capacity, healthCoverage, LoopHealthSource.none(), quarantine, leadChannel, outage, leadSeats,
|
||||
LeadConfigDirSource.none(), peers, leadRollover);
|
||||
}
|
||||
|
||||
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
|
||||
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
|
||||
CallerResolver callers, AuthorizationMode authorizationMode, Metrics metrics,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage, LoopHealthSource loopHealth,
|
||||
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage,
|
||||
LeadSeatSource leadSeats, LeadConfigDirSource leadConfigDirs, List<String> peers,
|
||||
LeadRollover leadRollover) {
|
||||
Objects.requireNonNull(callers, "callers");
|
||||
this.authorizationEnforced = Objects.requireNonNull(authorizationMode, "authorizationMode")
|
||||
== AuthorizationMode.ENFORCED;
|
||||
this.identity = identity;
|
||||
this.leadChannel = leadChannel;
|
||||
this.peers = peers == null ? List.of() : List.copyOf(peers);
|
||||
this.capacity = capacity;
|
||||
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
|
||||
this.outage = Objects.requireNonNull(outage, "outage");
|
||||
this.leadSeats = Objects.requireNonNull(leadSeats, "leadSeats");
|
||||
this.leadConfigDirs = Objects.requireNonNull(leadConfigDirs, "leadConfigDirs");
|
||||
this.healthCoverage = healthCoverage;
|
||||
this.loopHealth = Objects.requireNonNull(loopHealth, "loopHealth");
|
||||
this.leadRollover = leadRollover;
|
||||
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
|
||||
this.transport = HttpServletStreamableServerTransportProvider.builder()
|
||||
@@ -474,8 +544,8 @@ public final class FleetMcp {
|
||||
(exchange, _) -> {
|
||||
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_list", Map.of()), null);
|
||||
if (denied != null) return denied;
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
|
||||
leadSeats, callers.leads(),
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine, outage,
|
||||
leadSeats, leadContextGauge, leadConfigDirs, callers.leads(),
|
||||
callerTerminal(exchange),
|
||||
new CoordinationSource(leadChannel, peers),
|
||||
coordinatorVisibleTo(principal(exchange)));
|
||||
@@ -693,6 +763,15 @@ public final class FleetMcp {
|
||||
return transport;
|
||||
}
|
||||
|
||||
/**
|
||||
* The {@link ConnectionIdentity} this server resolves every caller against — fleetd #612
|
||||
* CB-185: lets a test reach the exact {@link dev.ltms.fleet.herdr.PaneLocator} a real assembly
|
||||
* wired up (via {@link ConnectionIdentity#panes()}), rather than a copy built for the test.
|
||||
*/
|
||||
public ConnectionIdentity identity() {
|
||||
return identity;
|
||||
}
|
||||
|
||||
/** Mark a connected spawned member available for the injector readiness gate. */
|
||||
static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) {
|
||||
if (caller.isSpawnedMember()) {
|
||||
@@ -718,6 +797,46 @@ public final class FleetMcp {
|
||||
return server.listTools();
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #612 B3 — same reason as {@link #registeredTools()}: a test that must drive the REAL
|
||||
* {@link QuarantineSource} (and the real {@link BackendQuarantine} it wraps) this daemon was
|
||||
* assembled with, rather than scraping {@code FleetdAssembly.java}'s source text for the
|
||||
* constructor call that built it. Unlike {@link #registeredTools()}'s callers, that test cannot
|
||||
* live in this package: it also builds the {@code ResourcePorts} that drives
|
||||
* {@code FleetdAssembly.assembleAndStart}, and {@code ResourcePorts}' methods return
|
||||
* {@code Fleetd}-nested types that are only visible from package {@code dev.ltms.fleet} — so
|
||||
* this accessor is {@code public}, not package-private, to stay reachable from there. {@code
|
||||
* FleetdBackendQuarantineAssemblyTest} quarantines a credential twice through this exact
|
||||
* instance and checks the second cooldown is longer than the first — the one behavioural
|
||||
* difference {@link BackendQuarantine#withEscalation} and the flat two-argument constructor
|
||||
* actually produce.
|
||||
*/
|
||||
public QuarantineSource quarantineSource() {
|
||||
return quarantine;
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #612 B3 — as {@link #quarantineSource()}, {@code public} for the same cross-package
|
||||
* reason, for the real {@link LeadSeatSource} this daemon was assembled with. {@code
|
||||
* FleetdLeadSeatAssemblyTest} calls {@code seatsFor} on this exact instance and checks it
|
||||
* reports a live lead's seat, which {@link LeadSeatSource#none()} can never do (it is a
|
||||
* constant-zero function regardless of input).
|
||||
*/
|
||||
public LeadSeatSource leadSeatSource() {
|
||||
return leadSeats;
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #612 B3 — as {@link #quarantineSource()}, {@code public} for the same cross-package
|
||||
* reason, for the real {@link LeadRollover} (or {@code null}) this daemon was assembled with.
|
||||
* {@code FleetdLeadRolloverAssemblyTest} drives {@code open}/{@code confirm} on this exact
|
||||
* instance and waits for the real continuation to send {@code /clear} and {@code bootstrapText}
|
||||
* through the real {@code router.leadAgents()}.
|
||||
*/
|
||||
public LeadRollover leadRollover() {
|
||||
return leadRollover;
|
||||
}
|
||||
|
||||
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
|
||||
|
||||
/**
|
||||
@@ -805,6 +924,12 @@ public final class FleetMcp {
|
||||
+ "answered (turnId stale)");
|
||||
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> text("[no reply within " + timeout + "ms — worker "
|
||||
+ r.outcome().name().toLowerCase().replace("timed_out_", "") + "; retry or poll status]");
|
||||
// fleetd #571: delivery is unknown here — agent.prompt pastes and submits in one call,
|
||||
// so the message may already be sitting in the pane. Do not invite a blind retry the way
|
||||
// the case above does; a resend on this route can double-deliver the same brief.
|
||||
case TIMED_OUT_UNCONFIRMED -> text("[no reply within " + timeout + "ms — delivery unconfirmed; "
|
||||
+ "the message may already have reached the worker, so a retry risks sending it "
|
||||
+ "twice — poll status before resending]");
|
||||
};
|
||||
}
|
||||
|
||||
@@ -1219,14 +1344,16 @@ public final class FleetMcp {
|
||||
Map<String, Object> args) {
|
||||
String action = str(args, "action");
|
||||
if (isBlank(action)) {
|
||||
return error("action is required: \"open\", \"confirm\" or \"cancel\"");
|
||||
return error("action is required: \"open\", \"confirm\", \"cancel\" or \"status\"");
|
||||
}
|
||||
return switch (action) {
|
||||
case "open" -> handoverOpen(leadRollover, callerTerminal, str(args, "reason"));
|
||||
case "confirm" -> handoverConfirm(leadRollover, callerTerminal, str(args, "token"),
|
||||
truthy(args, "operatorConfirmed"));
|
||||
case "cancel" -> handoverCancel(leadRollover, str(args, "token"));
|
||||
default -> error("unknown action \"" + action + "\" — must be \"open\", \"confirm\" or \"cancel\"");
|
||||
case "status" -> handoverStatus(leadRollover, str(args, "token"));
|
||||
default -> error("unknown action \"" + action
|
||||
+ "\" — must be \"open\", \"confirm\", \"cancel\" or \"status\"");
|
||||
};
|
||||
}
|
||||
|
||||
@@ -1296,6 +1423,29 @@ public final class FleetMcp {
|
||||
return text(json(m));
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code action: "status"}. Read-only — see {@link LeadRollover#status}: never schedules,
|
||||
* cancels, or retries anything, and it is the only way for a lead to find out what happened to
|
||||
* a token past {@code confirm()}, since every outcome after that point is otherwise logged only
|
||||
* (see {@link LeadRollover}'s class javadoc).
|
||||
*/
|
||||
private static McpSchema.CallToolResult handoverStatus(LeadRollover leadRollover, String token) {
|
||||
if (leadRollover == null) {
|
||||
Map<String, Object> m = new LinkedHashMap<>();
|
||||
m.put("state", "NOT_CONFIGURED");
|
||||
m.put("detail", "leadRollover: is not configured");
|
||||
return text(json(m));
|
||||
}
|
||||
if (isBlank(token)) {
|
||||
return error("token is required for action \"status\"");
|
||||
}
|
||||
LeadRollover.RollStatus s = leadRollover.status(token);
|
||||
Map<String, Object> m = new LinkedHashMap<>();
|
||||
m.put("state", s.state().name());
|
||||
m.put("detail", s.detail());
|
||||
return text(json(m));
|
||||
}
|
||||
|
||||
/** The one shared {@code NOT_CONFIGURED} refusal shape for {@code open}/{@code confirm}. */
|
||||
private static McpSchema.CallToolResult notConfigured() {
|
||||
return refusalJson(false, "NOT_CONFIGURED", "leadRollover: is not configured");
|
||||
@@ -1546,25 +1696,42 @@ public final class FleetMcp {
|
||||
* @param selfTerm the calling pane's terminal id, or blank for a caller with no pane
|
||||
*/
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions,
|
||||
Map<String, String> leads, String selfTerm) {
|
||||
Map<String, String> leads, String selfTerm) {
|
||||
return listFleet(workers, sessions, null, CapacitySource.none(), new HealthCoverageSource(() -> "off"),
|
||||
QuarantineSource.none(), leads, selfTerm);
|
||||
LoopHealthSource.none(), QuarantineSource.none(), leads, selfTerm);
|
||||
}
|
||||
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, Map<String, String> leads, String selfTerm) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, leads, selfTerm,
|
||||
CoordinationSource.none());
|
||||
}
|
||||
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, Map<String, String> leads, String selfTerm) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, leads, selfTerm,
|
||||
LoopHealthSource loopHealth, QuarantineSource quarantine,
|
||||
Map<String, String> leads, String selfTerm) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine, leads, selfTerm,
|
||||
CoordinationSource.none());
|
||||
}
|
||||
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
LoopHealthSource loopHealth, QuarantineSource quarantine,
|
||||
Map<String, String> leads, String selfTerm,
|
||||
CoordinationSource coordination) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine,
|
||||
OutageSource.none(), LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
|
||||
}
|
||||
|
||||
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, OutageSource outage,
|
||||
Map<String, String> leads, String selfTerm) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
|
||||
LeadSeatSource.none(), leads, selfTerm, CoordinationSource.none());
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
|
||||
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, CoordinationSource.none(), false);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1580,8 +1747,8 @@ public final class FleetMcp {
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, Map<String, String> leads, String selfTerm,
|
||||
CoordinationSource coordination) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, OutageSource.none(),
|
||||
LeadSeatSource.none(), leads, selfTerm, coordination);
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, OutageSource.none(),
|
||||
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
|
||||
}
|
||||
|
||||
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
|
||||
@@ -1589,8 +1756,8 @@ public final class FleetMcp {
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, OutageSource outage,
|
||||
Map<String, String> leads, String selfTerm, CoordinationSource coordination) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
|
||||
LeadSeatSource.none(), leads, selfTerm, coordination);
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
|
||||
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1611,8 +1778,8 @@ public final class FleetMcp {
|
||||
QuarantineSource quarantine, OutageSource outage,
|
||||
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
|
||||
CoordinationSource coordination) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
|
||||
leadSeats, leads, selfTerm, coordination, false);
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
|
||||
leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1631,9 +1798,29 @@ public final class FleetMcp {
|
||||
* an explicit {@code true}
|
||||
*/
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, OutageSource outage,
|
||||
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, OutageSource outage,
|
||||
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
|
||||
CoordinationSource coordination, boolean callerIsPrimary) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine,
|
||||
outage, leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, callerIsPrimary);
|
||||
}
|
||||
|
||||
/**
|
||||
* The canonical implementation. {@code contextGauge} is the "lead context gauge" (see
|
||||
* {@link LeadContextGauge}) — every wrapper overload above passes a freshly constructed one,
|
||||
* which is correct for them (none of them exercise repeated calls where a shared cache would
|
||||
* matter); the one caller that matters for caching, {@code fleet_list}'s MCP handler, passes
|
||||
* its own single long-lived instance instead (see {@code FleetMcp}'s {@code leadContextGauge}
|
||||
* field).
|
||||
*/
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
LoopHealthSource loopHealth,
|
||||
QuarantineSource quarantine, OutageSource outage,
|
||||
LeadSeatSource leadSeats, LeadContextGauge contextGauge,
|
||||
LeadConfigDirSource leadConfigDirs,
|
||||
Map<String, String> leads, String selfTerm,
|
||||
CoordinationSource coordination, boolean callerIsPrimary) {
|
||||
try {
|
||||
Map<String, Agent> live = workers.list().stream()
|
||||
@@ -1642,7 +1829,8 @@ public final class FleetMcp {
|
||||
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
|
||||
List<Map<String, Object>> leadRows = leads.entrySet().stream()
|
||||
.sorted(Map.Entry.comparingByValue())
|
||||
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm))
|
||||
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm, contextGauge,
|
||||
leadConfigDirs))
|
||||
.toList();
|
||||
// fleetd #209: this is the caller-driven fleet_list read that actually reports
|
||||
// agentSessionId (via memberCapacityView -> SessionManager.rosterView), so it uses the
|
||||
@@ -1656,6 +1844,9 @@ public final class FleetMcp {
|
||||
Map<String, Object> result = new LinkedHashMap<>();
|
||||
result.put("leads", leadRows); result.put("members", out);
|
||||
result.put("healthCoverage", healthCoverage.value().get());
|
||||
result.put("loopHealth", Map.of(
|
||||
"statusPoller", loopHealth.statusPoller().get().name(),
|
||||
"sessionReaper", loopHealth.sessionReaper().get().name()));
|
||||
// fleetd #439: coordinator/coordinatorView is lead-to-lead coordination state and must
|
||||
// never reach a worker or an architect -- gate BEFORE assembling it, not after, so the
|
||||
// key is absent rather than present-and-empty.
|
||||
@@ -1934,15 +2125,25 @@ public final class FleetMcp {
|
||||
}
|
||||
|
||||
/**
|
||||
* One lead's row: its address, its name, and whether it can be reached right now.
|
||||
* One lead's row: its address, its name, whether it can be reached right now, and how full its
|
||||
* own Claude Code context window is.
|
||||
*
|
||||
* <p>{@code status} is herdr's live view, and {@code unknown} when herdr is not tracking that
|
||||
* pane as an agent — the honest answer, and the one that matters: a lead whose pane herdr cannot
|
||||
* see is a lead a {@code fleet_send} cannot be typed into. It is reported rather than hidden,
|
||||
* because a peer that has gone unreachable is exactly what the sender needs to know.
|
||||
*
|
||||
* <p>{@code context} is the "lead context gauge" (fleetd's context-usage visibility ticket):
|
||||
* {@code {state: "ok"|"high"|"unknown", tokens?: number, compactions: number}}, read from the
|
||||
* lead's transcript — see {@link LeadContextGauge}. Kept small on purpose (a token count, a
|
||||
* state, and a compaction count) rather than echoing the whole reading history: this is a
|
||||
* roster row a caller glances at, not a diagnostics dump. {@code tokens} is present only when
|
||||
* {@code state} is not {@code "unknown"} — never a stale or default number standing in for "I
|
||||
* could not tell".
|
||||
*/
|
||||
private static Map<String, Object> leadView(String terminal, String name, Agent live,
|
||||
String selfTerm) {
|
||||
String selfTerm, LeadContextGauge contextGauge,
|
||||
LeadConfigDirSource leadConfigDirs) {
|
||||
Map<String, Object> m = new LinkedHashMap<>();
|
||||
m.put("sessionId", terminal);
|
||||
m.put("name", name);
|
||||
@@ -1951,9 +2152,32 @@ public final class FleetMcp {
|
||||
if (terminal.equals(selfTerm)) {
|
||||
m.put("self", true);
|
||||
}
|
||||
String configDir = leadConfigDirs.configDirFor().apply(name);
|
||||
m.put("context", contextView(contextGauge, live, configDir));
|
||||
return m;
|
||||
}
|
||||
|
||||
/**
|
||||
* Reads the lead context gauge for one lead. {@code configDir} is this lead's configured
|
||||
* {@code CLAUDE_CONFIG_DIR} override (see {@link LeadConfigDirSource}), derived from
|
||||
* {@code fleet.leaders.<name>.profile} → that profile's own {@code configDir:} — or {@code null}
|
||||
* when the lead's entry names no profile, or that profile sets no override, in which case
|
||||
* {@link LeadContextGauge#read} falls back to its own built-in default
|
||||
* ({@code <user.home>/.claude}).
|
||||
*/
|
||||
private static Map<String, Object> contextView(LeadContextGauge contextGauge, Agent live, String configDir) {
|
||||
String sessionId = live == null ? null : live.sessionId();
|
||||
String agentType = live == null ? null : live.agentType();
|
||||
LeadContextGauge.Reading reading = contextGauge.read(configDir, sessionId, agentType);
|
||||
Map<String, Object> c = new LinkedHashMap<>();
|
||||
c.put("state", reading.state().name().toLowerCase());
|
||||
if (reading.tokens() != null) {
|
||||
c.put("tokens", reading.tokens());
|
||||
}
|
||||
c.put("compactions", reading.compactions());
|
||||
return c;
|
||||
}
|
||||
|
||||
/** {@code fleet_stop}: tear a worker down by its pane id. */
|
||||
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
|
||||
if (isBlank(paneId)) {
|
||||
@@ -2075,8 +2299,9 @@ public final class FleetMcp {
|
||||
return tool(FleetTool.SPAWN.wireName(),
|
||||
"Spawn a new off-subscription member session. A member has two independent attributes: "
|
||||
+ "role (what it is for) and profile (which backend it runs on). Pass role to pick "
|
||||
+ "the contract — 'dev' implements a unit and opens its own PR, 'reviewer' reviews a "
|
||||
+ "diff it did not write, 'architect' refines a ticket before anyone builds it; omit "
|
||||
+ "the contract — 'dev' implements a unit and opens its own PR, 'hunter' sweeps a "
|
||||
+ "scope without changing it, 'reviewer' reviews a diff it did not write, 'architect' "
|
||||
+ "refines a ticket before anyone builds it; omit "
|
||||
+ "it for 'dev'. Pass profile (from fleet_profiles) to pick the backend, or omit it "
|
||||
+ "for the default. The two are independent: a reviewer may run on the same profile "
|
||||
+ "as the dev it reviews. The member opens your current directory by default; pass "
|
||||
@@ -2093,7 +2318,7 @@ public final class FleetMcp {
|
||||
+ "one. Returns the member's sessionId (use with fleet_send) and paneId (use with "
|
||||
+ "fleet_stop).",
|
||||
objectSchema(Map.of(
|
||||
"role", stringProp("What the member is for: architect, dev or reviewer (default dev)"),
|
||||
"role", stringProp("What the member is for: architect, dev, hunter, or reviewer (default dev)"),
|
||||
"profile", stringProp("Which backend to run it on (omit for the default profile)"),
|
||||
"cwd", stringProp("Working directory for the member (omit to inherit yours)"),
|
||||
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
|
||||
@@ -2135,9 +2360,10 @@ public final class FleetMcp {
|
||||
+ "cannot reliably re-identify: some backends (e.g. opencode) resolve it from the "
|
||||
+ "member's working directory, which only uniquely identifies a member when it "
|
||||
+ "was spawned into its own fleetd-provisioned worktree (worktree:true/<slug>); a "
|
||||
+ "member spawned without one shares its directory with others and never reports "
|
||||
+ "an id, however long it runs (fleetd #249). An empty 'members' "
|
||||
+ "means no members are spawned; it says nothing about peers. When capacity "
|
||||
+ "member spawned without one shares its directory with others and never reports "
|
||||
+ "an id, however long it runs (fleetd #249). An empty 'members' "
|
||||
+ "means no members are spawned; it says nothing about peers. 'loopHealth' reports "
|
||||
+ "the RUNNING, STALLED, or STOPPED state of statusPoller and sessionReaper. When capacity "
|
||||
+ "facts are configured, a 'capacity' row per profile reports 'free' — the "
|
||||
+ "slots a fresh fleet_spawn on that profile will actually be granted right "
|
||||
+ "now (max(0, maxLoad - live)), the same check the spawn gate itself runs. A "
|
||||
@@ -2204,20 +2430,26 @@ public final class FleetMcp {
|
||||
return tool(FleetTool.HANDOVER.wireName(),
|
||||
"Replace your OWN lead session once its context is full: write a handover file, "
|
||||
+ "then use this to have fleetd clear your pane and bootstrap a fresh lead "
|
||||
+ "session against it. Three actions: 'open' (requests a token and the "
|
||||
+ "session against it. Four actions: 'open' (requests a token and the "
|
||||
+ "handoverPath you must write the handover file to before confirming), "
|
||||
+ "'confirm' (validates every gate and — only if every one passes — schedules "
|
||||
+ "the roll; it does NOT itself clear the pane, the roll runs once this call's "
|
||||
+ "own turn ends), and 'cancel' (drops a pending request without rolling). "
|
||||
+ "Primary-only. There is deliberately no terminal/session/leadTerminal "
|
||||
+ "parameter: the pane to roll is always resolved from YOUR OWN connection, "
|
||||
+ "never a value you pass, so you can only ever roll yourself — never another "
|
||||
+ "lead. Requires leadRollover: to be configured; when it is not, every action "
|
||||
+ "returns a clean refusal naming NOT_CONFIGURED instead of failing.",
|
||||
+ "own turn ends), 'cancel' (drops a pending request without rolling), and "
|
||||
+ "'status' (read-only: what happened to a token after 'confirm' — still "
|
||||
+ "running (approved but not finished yet), the roll completed, the calling "
|
||||
+ "turn never settled within turnSettleSeconds so no /clear was ever sent, or "
|
||||
+ "/clear itself never settled so bootstrapText was never sent; never "
|
||||
+ "schedules, cancels or retries anything). Primary-only. "
|
||||
+ "There is deliberately no terminal/session/leadTerminal parameter: the pane "
|
||||
+ "to roll is always resolved from YOUR OWN connection, never a value you "
|
||||
+ "pass, so you can only ever roll yourself — never another lead. Requires "
|
||||
+ "leadRollover: to be configured; when it is not, every action returns a "
|
||||
+ "clean refusal naming NOT_CONFIGURED instead of failing.",
|
||||
objectSchema(Map.of(
|
||||
"action", stringProp("\"open\", \"confirm\" or \"cancel\""),
|
||||
"action", stringProp("\"open\", \"confirm\", \"cancel\" or \"status\""),
|
||||
"reason", stringProp("Free-text audit note for \"open\" (optional, logged only)"),
|
||||
"token", stringProp("The token \"open\" returned — required for \"confirm\" and \"cancel\""),
|
||||
"token", stringProp("The token \"open\" returned — required for \"confirm\", "
|
||||
+ "\"cancel\" and \"status\""),
|
||||
"operatorConfirmed", Map.of("type", "boolean",
|
||||
"description", "For \"confirm\": your answer to \"has the human operator "
|
||||
+ "confirmed this wipe\" (default false; only consulted when "
|
||||
|
||||
@@ -8,6 +8,7 @@ import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.Agent;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.launch.ClaudeCodeArguments;
|
||||
import dev.ltms.fleet.peer.Capability;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import org.slf4j.Logger;
|
||||
@@ -298,7 +299,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
// has neither MCP nor a charter — session flags must be added into a list we own.
|
||||
List<String> argv = mutableArgv(argvWithFleet(cfg, spec));
|
||||
String agentSessionId = applySessionIdentity(argv, spec.sessionName(), spec.resumeSessionId());
|
||||
return new Launch(workerEnv, argvWithAutoCompact(argvWithModel(argv, cfg), cfg), agentSessionId);
|
||||
return new Launch(workerEnv, ClaudeCodeArguments.withAutoCompactWindow(argvWithModel(argv, cfg), cfg), agentSessionId);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -908,30 +909,6 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
return withModel;
|
||||
}
|
||||
|
||||
/**
|
||||
* Pin a bounded auto-compaction window on the command line via {@code --autocompact <tokens>},
|
||||
* opt-in per profile (CB-634's sibling ticket: a member that runs out of context dies mid-turn
|
||||
* and its {@code fleet_reply} — the whole point of the turn — is lost with it; opencode already
|
||||
* forces {@code compaction.auto: true} unconditionally, CB-523, but Claude Code has no equivalent
|
||||
* and runs at the backend's own default window).
|
||||
*
|
||||
* <p>Mirrors {@link #argvWithModel}: appended after it, so it survives the {@code ccs <profile>}
|
||||
* wrapper the same way {@code --model} does, and outranks env/settings and the operator's own
|
||||
* {@code argv}. Verified: {@code claude 2.1.241 --help} lists {@code --autocompact <auto|tokens>}
|
||||
* (either the literal {@code auto}, or an integer 100k–1M) — {@link FleetConfig#load} rejects a
|
||||
* configured value outside that band before this ever runs, so the flag Claude Code receives here
|
||||
* is always in range.
|
||||
*/
|
||||
private static List<String> argvWithAutoCompact(List<String> argv, FleetConfig.Profile cfg) {
|
||||
if (cfg.autoCompactWindow() == null) {
|
||||
return argv;
|
||||
}
|
||||
List<String> withAutoCompact = mutableArgv(argv);
|
||||
withAutoCompact.add("--autocompact");
|
||||
withAutoCompact.add(String.valueOf(cfg.autoCompactWindow()));
|
||||
return withAutoCompact;
|
||||
}
|
||||
|
||||
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
|
||||
|
||||
/** Spawn a worker for the default profile in the resolved default cwd. */
|
||||
|
||||
@@ -246,7 +246,7 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
* neither. So before this method existed with a requeue step, it dropped {@link #held}'s entries
|
||||
* for {@code target} while the broker still considered them outstanding: never acked, never
|
||||
* nacked, never requeued, and no longer reachable by {@link #peek} — permanently invisible. This
|
||||
* is unlike {@link #handleRecovery} and {@link #close()}, whose bare {@code held.clear()} is
|
||||
* is unlike {@link RecoveryListener#handleRecovery(Recoverable)} and {@link #close()}, whose bare {@code held.clear()} is
|
||||
* correct because each has already made the broker requeue (a real connection drop, or
|
||||
* {@code channel.close()} respectively) before clearing local state.
|
||||
*
|
||||
|
||||
@@ -0,0 +1,24 @@
|
||||
package dev.ltms.fleet.msg;
|
||||
|
||||
/**
|
||||
* fleetd #612 A-gaps (gap 1): a {@link LeadChannel} that its owner can also close.
|
||||
*
|
||||
* <p>{@link LeadChannel}'s own javadoc says plainly that {@code close()} is deliberately left out
|
||||
* of that interface — draining is a caller convenience nobody uses, and closing is the
|
||||
* <em>owner's</em> job. This interface is that owner's own, wider view: whoever opens the
|
||||
* coordination mailbox (the assembly that builds the daemon) also needs to close it from the
|
||||
* shutdown path, and a test standing in for a real broker connection needs a fake it can mark
|
||||
* closed, without ever holding a live connection. Every ordinary consumer ({@code FleetMcp},
|
||||
* {@link LeadCoordLoop}) keeps taking the narrower {@link LeadChannel} exactly as before — only
|
||||
* the owner speaks this wider one.
|
||||
*
|
||||
* <p>{@link LeadMailbox} is still the only production implementation. This only generalises the
|
||||
* TYPE its owner holds it as (previously the concrete class), so a test can substitute a fake
|
||||
* closeable channel instead of a real AMQP connection.
|
||||
*/
|
||||
public interface LeadChannelHandle extends LeadChannel, AutoCloseable {
|
||||
|
||||
/** Release the underlying connection. Declared with no checked exception, unlike the plain {@link AutoCloseable#close()}. */
|
||||
@Override
|
||||
void close();
|
||||
}
|
||||
@@ -2,6 +2,7 @@ package dev.ltms.fleet.msg;
|
||||
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.AgentStatus;
|
||||
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||
import dev.ltms.fleet.mcp.PrimaryRegistry;
|
||||
import dev.ltms.fleet.metrics.FleetMetrics;
|
||||
import dev.ltms.fleet.metrics.Metrics;
|
||||
@@ -13,6 +14,7 @@ import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
@@ -44,6 +46,15 @@ import java.util.function.Supplier;
|
||||
* loop stands down, so two competing injections never start two turns in the same pane
|
||||
* (constraint 6).</li>
|
||||
* </ol>
|
||||
*
|
||||
* <p><b>fleetd #609 — context-high notice.</b> Optionally ({@code contextHighNudge}, opt-in like the
|
||||
* loop itself), a tick that finds the lead's own {@link LeadContextGauge} reading at {@link
|
||||
* LeadContextGauge.State#HIGH} appends a text notice to whatever nudge it sends, telling the lead to
|
||||
* consider {@code fleet_handover}. This is text only — it never rolls a pane itself. It fires once per
|
||||
* HIGH stretch (a latch, cleared only by a later {@code OK} reading — {@code UNKNOWN} neither sets nor
|
||||
* clears it, since "I could not look" must not be read as "it got better"), and it never spends the
|
||||
* quiet-nudge budget: an idle, quiet, HIGH-context lead is exactly the case {@link Action#QUIET_DONE}
|
||||
* would otherwise swallow, and it is the one case most worth interrupting the quiet cap for.
|
||||
*/
|
||||
public final class LeadHeartbeatLoop {
|
||||
|
||||
@@ -63,11 +74,16 @@ public final class LeadHeartbeatLoop {
|
||||
private final long backoffMs;
|
||||
private final int quietNudgeCap;
|
||||
private final Metrics metrics; // CB-512 pattern: nullable — no registry in unit tests
|
||||
private final LeadContextSource contextSource; // fleetd #609
|
||||
private final boolean contextHighNudge; // fleetd #609: opt-in, like the loop itself
|
||||
private final boolean requireOperatorConfirm; // fleetd #621: mirrors leadRollover.requireOperatorConfirm
|
||||
|
||||
/** When the current idle stretch began (nanos), or {@link #NOT_IDLE}. Single scheduler thread only. */
|
||||
private long idleSinceNanos = NOT_IDLE;
|
||||
/** Consecutive nudges that found no pending fleet state. Single scheduler thread only. */
|
||||
private int quietCount = 0;
|
||||
/** fleetd #609: latched "already told this lead about this HIGH stretch". Single scheduler thread only. */
|
||||
private boolean contextNotified = false;
|
||||
|
||||
/** Constructor with an injectable clock and no metric registry (unit tests, or wiring that opts out). */
|
||||
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
||||
@@ -75,7 +91,7 @@ public final class LeadHeartbeatLoop {
|
||||
ScheduledExecutorService scheduler, LongSupplier clock,
|
||||
long idleAfterNanos, long backoffMs, int quietNudgeCap) {
|
||||
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
|
||||
idleAfterNanos, backoffMs, quietNudgeCap, null);
|
||||
idleAfterNanos, backoffMs, quietNudgeCap, null, LeadContextSource.none(), false);
|
||||
}
|
||||
|
||||
/** As above, with a metric registry (the CB-512 pattern) so nudge outcomes are counted. */
|
||||
@@ -83,6 +99,40 @@ public final class LeadHeartbeatLoop {
|
||||
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
|
||||
ScheduledExecutorService scheduler, LongSupplier clock,
|
||||
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics) {
|
||||
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
|
||||
idleAfterNanos, backoffMs, quietNudgeCap, metrics, LeadContextSource.none(), false);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #609: as above, plus the lead's own context source and whether a HIGH reading should
|
||||
* append a hand-over notice to the loop's nudge. Pass {@link LeadContextSource#none()} and
|
||||
* {@code false} to keep the pre-#609 behaviour exactly (both existing public constructors do).
|
||||
*
|
||||
* <p>fleetd #621: delegates to the full constructor with {@code requireOperatorConfirm=true} —
|
||||
* the pre-#621 wording ("ask the operator ... only the operator can approve the roll") assumed
|
||||
* the config default, so every caller of this overload keeps that text byte-identical.
|
||||
*/
|
||||
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
||||
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
|
||||
ScheduledExecutorService scheduler, LongSupplier clock,
|
||||
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics,
|
||||
LeadContextSource contextSource, boolean contextHighNudge) {
|
||||
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
|
||||
idleAfterNanos, backoffMs, quietNudgeCap, metrics, contextSource, contextHighNudge, true);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #621: as above, plus the daemon's effective {@code leadRollover.requireOperatorConfirm}
|
||||
* value — threaded into {@link #contextNotice(boolean, LeadContextGauge.Reading, boolean, boolean)}
|
||||
* so the notice's wording tracks the config the daemon actually enforces (see {@code
|
||||
* LeadRollover.confirm}) instead of always asserting the operator gate is on.
|
||||
*/
|
||||
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
||||
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
|
||||
ScheduledExecutorService scheduler, LongSupplier clock,
|
||||
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics,
|
||||
LeadContextSource contextSource, boolean contextHighNudge,
|
||||
boolean requireOperatorConfirm) {
|
||||
this.primaryRegistry = primaryRegistry;
|
||||
this.agents = agents;
|
||||
this.inbox = inbox;
|
||||
@@ -94,6 +144,20 @@ public final class LeadHeartbeatLoop {
|
||||
this.backoffMs = backoffMs;
|
||||
this.quietNudgeCap = quietNudgeCap;
|
||||
this.metrics = metrics;
|
||||
this.contextSource = contextSource;
|
||||
this.contextHighNudge = contextHighNudge;
|
||||
this.requireOperatorConfirm = requireOperatorConfirm;
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #609: one lead's own context reading, keyed by its terminal id — the same injected-source
|
||||
* idiom {@code FleetMcp.LeadSeatSource}/{@code FleetMcp.LeadConfigDirSource} already use.
|
||||
*/
|
||||
public record LeadContextSource(Function<String, LeadContextGauge.Reading> readingFor) {
|
||||
/** Inert source — every lead reads UNKNOWN, so the context notice can never fire. */
|
||||
public static LeadContextSource none() {
|
||||
return new LeadContextSource(_ -> LeadContextGauge.Reading.unknown());
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -126,8 +190,11 @@ public final class LeadHeartbeatLoop {
|
||||
STAND_DOWN
|
||||
}
|
||||
|
||||
/** The outcome of one decision: the action plus the state to persist for the next tick. */
|
||||
record Decision(Action action, Long idleSinceNanos, int quietCount) {}
|
||||
/**
|
||||
* The outcome of one decision: the action, the state to persist for the next tick, and (fleetd
|
||||
* #609) whether the lead has now been told about the current HIGH context stretch.
|
||||
*/
|
||||
record Decision(Action action, Long idleSinceNanos, int quietCount, boolean contextNotified) {}
|
||||
|
||||
/**
|
||||
* Pure decision function: given the current loop state and fleet/lead facts, return what to do
|
||||
@@ -142,34 +209,47 @@ public final class LeadHeartbeatLoop {
|
||||
* @param pushLoopActive whether {@link ReplyPushLoop} is currently nudging some target (constraint 6)
|
||||
* @param leadKnown whether a lead terminal is known to nudge at all
|
||||
* @param fleet a snapshot of the pending fleet state (constraint 5)
|
||||
* @param context fleetd #609: the lead's own {@link LeadContextGauge} reading for this tick
|
||||
* @param contextNotified fleetd #609: whether the lead has already been told about the current HIGH
|
||||
* stretch — a latch, carried forward by {@link #applyDecision}
|
||||
* @return the action to take and the state to persist
|
||||
*/
|
||||
Decision decide(long nowNanos, Long idleSinceNanos, int quietCount, AgentStatus status,
|
||||
boolean pushLoopActive, boolean leadKnown, FleetState fleet) {
|
||||
boolean pushLoopActive, boolean leadKnown, FleetState fleet,
|
||||
LeadContextGauge.State context, boolean contextNotified) {
|
||||
// fleetd #609: re-arm the latch only on a positive OK reading. UNKNOWN means "I could not
|
||||
// look", not "it got better" — re-arming on UNKNOWN would let a flapping gauge (a transcript
|
||||
// read that misses one tick) nudge a full lead again on every recovery, defeating the "once
|
||||
// per HIGH stretch" promise. Computed once, up front, so every gate below carries it forward
|
||||
// unchanged unless it is the gate that actually discharges it.
|
||||
boolean latch = context == LeadContextGauge.State.OK ? false : contextNotified;
|
||||
boolean contextHigh = contextHighNudge && context == LeadContextGauge.State.HIGH;
|
||||
|
||||
// Constraint 6: while ReplyPushLoop is actively nudging the lead, injecting a second,
|
||||
// competing prompt into the same pane would start a second turn — racing loops multiply
|
||||
// turns and context burn. Stand aside, and treat the active push as real state (re-arm the
|
||||
// quiet counter), because the reply that drove it is exactly the kind of new state that
|
||||
// should reset the cap.
|
||||
// should reset the cap. The context latch is untouched: standing down must not spend the
|
||||
// one notice this stretch gets.
|
||||
if (pushLoopActive) {
|
||||
return new Decision(Action.STAND_DOWN, idleSinceNanos, 0);
|
||||
return new Decision(Action.STAND_DOWN, idleSinceNanos, 0, latch);
|
||||
}
|
||||
// Constraint 2: a WORKING lead is making progress and must NOT be touched; an unreadable
|
||||
// status (read failure, or the agent is gone) is safest treated the same way — never inject
|
||||
// into a state we cannot read. Either way, reset the idle window and the quiet counter: the
|
||||
// lead was / may be active, so the next idle stretch must count its own quiet period fresh.
|
||||
if (status == null || !status.injectable()) {
|
||||
return new Decision(Action.LEAD_BUSY, null, 0);
|
||||
return new Decision(Action.LEAD_BUSY, null, 0, latch);
|
||||
}
|
||||
if (idleSinceNanos == null) {
|
||||
// The lead just became injectable — record the start of an idle stretch and wait out the
|
||||
// debounce quiet period before ever nudging (constraint 3).
|
||||
return new Decision(Action.WAIT_IDLE, nowNanos, quietCount);
|
||||
return new Decision(Action.WAIT_IDLE, nowNanos, quietCount, latch);
|
||||
}
|
||||
if (nowNanos - idleSinceNanos < idleAfterNanos) {
|
||||
// Still within the quiet period: the lead that just finished a turn sits momentarily idle
|
||||
// and must not be re-prompted into every natural pause.
|
||||
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
|
||||
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount, latch);
|
||||
}
|
||||
// Past the quiet period with an injectable lead: it is a genuine candidate for a nudge. Two
|
||||
// gating facts decide whether and how:
|
||||
@@ -177,23 +257,33 @@ public final class LeadHeartbeatLoop {
|
||||
// No lead terminal is known yet (e.g. the registry has not learned one) — there is nobody
|
||||
// to nudge. Keep waiting; the window stays open so discovery re-arms it without a fresh
|
||||
// quiet period.
|
||||
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
|
||||
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount, latch);
|
||||
}
|
||||
if (fleet.hasPending()) {
|
||||
// Real fleet state is waiting — a worker reply or a DONE session. This is new state, so
|
||||
// it resets the quiet counter (constraint 4) and the lead is nudged to go collect it.
|
||||
return new Decision(Action.INJECT, idleSinceNanos, 0);
|
||||
// it resets the quiet counter (constraint 4) and the lead is nudged to go collect it. The
|
||||
// nudge text carries the context notice too when contextHigh — see injectNudge/contextNotice
|
||||
// — so this route discharges the same duty and must set the latch.
|
||||
return new Decision(Action.INJECT, idleSinceNanos, 0, latch || contextHigh);
|
||||
}
|
||||
if (contextHigh && !latch) {
|
||||
// fleetd #609: the lead is idle, its context is full, and nothing is pending. This is the
|
||||
// one case the quiet cap would otherwise swallow, and it is exactly when the lead most
|
||||
// needs to hear it. Fire once per HIGH stretch, and do NOT spend the quiet budget on it:
|
||||
// this is an event notice, not a "are you still there" nudge.
|
||||
return new Decision(Action.INJECT, idleSinceNanos, quietCount, true);
|
||||
}
|
||||
if (quietCount < quietNudgeCap) {
|
||||
// Nothing is pending, but the cap is not exhausted: nudge anyway, telling the lead
|
||||
// exactly that nothing is waiting so it can choose to stand down rather than hunt
|
||||
// (constraint 5). Count it toward the consecutive-quiet cap.
|
||||
return new Decision(Action.INJECT, idleSinceNanos, quietCount + 1);
|
||||
// (constraint 5). Count it toward the consecutive-quiet cap. This nudge also carries the
|
||||
// context notice when contextHigh (already latched above, or being latched now).
|
||||
return new Decision(Action.INJECT, idleSinceNanos, quietCount + 1, latch || contextHigh);
|
||||
}
|
||||
// Nothing pending and the cap is exhausted: stop nudging until real state appears again
|
||||
// (constraint 4). The loop still ticks on backoff so a genuinely new reply or session change
|
||||
// re-arms it — QUIET_DONE stops injection, not observation.
|
||||
return new Decision(Action.QUIET_DONE, idleSinceNanos, quietCount);
|
||||
return new Decision(Action.QUIET_DONE, idleSinceNanos, quietCount, latch);
|
||||
}
|
||||
|
||||
// --- loop ----------------------------------------------------------------------------------
|
||||
@@ -203,62 +293,195 @@ public final class LeadHeartbeatLoop {
|
||||
* does not evaluate the lead's idle state before the fleet has settled.
|
||||
*/
|
||||
public void start() {
|
||||
log.info("idle-lead heartbeat: on — nudge lead after {}s idle (recheck {}ms, quiet cap {})",
|
||||
TimeUnit.NANOSECONDS.toSeconds(idleAfterNanos), backoffMs, quietNudgeCap);
|
||||
// fleetd #613: contextHighNudge added alongside the three settings already here — an
|
||||
// operator otherwise cannot tell from the boot log whether the #609 handover notice is
|
||||
// armed, and had to load the deployed jar's config to confirm it.
|
||||
log.info("idle-lead heartbeat: on — nudge lead after {}s idle (recheck {}ms, quiet cap {}, "
|
||||
+ "context-high nudge {})",
|
||||
TimeUnit.NANOSECONDS.toSeconds(idleAfterNanos), backoffMs, quietNudgeCap,
|
||||
contextHighNudge);
|
||||
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
|
||||
}
|
||||
|
||||
/** One loop tick, every {@link #backoffMs} — the thin scheduler around {@link #decide}. */
|
||||
private void tick() {
|
||||
/** One loop tick, every {@link #backoffMs} — the thin scheduler around {@link #decide}. Package-private
|
||||
* (mirroring {@link ReplyPushLoop#tick(String)}) so tests can drive it directly with a fake clock and a
|
||||
* fake {@link AgentControl} instead of racing the scheduler thread. */
|
||||
void tick() {
|
||||
boolean leadKnown = primaryRegistry.primaryTerminal().isPresent();
|
||||
FleetState fleet = snapshot(inbox, roster);
|
||||
AgentStatus status = AgentStatus.UNKNOWN;
|
||||
LeadContextGauge.Reading reading = LeadContextGauge.Reading.unknown();
|
||||
if (leadKnown) {
|
||||
String leadTerminal = primaryRegistry.primaryTerminal().orElseThrow();
|
||||
try {
|
||||
status = agents.status(primaryRegistry.primaryTerminal().orElseThrow());
|
||||
status = agents.status(leadTerminal);
|
||||
} catch (RuntimeException e) {
|
||||
// A failed status read degrades to "unknown" — decide() treats that like a busy lead
|
||||
// and never injects into a state it cannot read. Retry on the next backoff.
|
||||
log.debug("idle-heartbeat: status check failed for lead, will retry: {}", e.toString());
|
||||
}
|
||||
// fleetd #609: read the lead's own context regardless of status — decide() still gates on
|
||||
// status first (constraint 2), so this is harmless work on a WORKING lead and lets the
|
||||
// latch state stay accurate for whenever the lead does go idle.
|
||||
reading = contextSource.readingFor().apply(leadTerminal);
|
||||
}
|
||||
|
||||
Decision d = decide(clock.getAsLong(),
|
||||
idleSinceNanos == NOT_IDLE ? null : idleSinceNanos,
|
||||
quietCount, status, pushLoop.isActive(), leadKnown, fleet);
|
||||
quietCount, status, pushLoop.isActive(), leadKnown, fleet,
|
||||
reading.state(), contextNotified);
|
||||
applyDecision(d);
|
||||
switch (d.action()) {
|
||||
case INJECT -> injectNudge(fleet);
|
||||
case QUIET_DONE -> countNudge("exhausted");
|
||||
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN -> { /* nothing to inject, nothing to count */ }
|
||||
case INJECT -> injectNudge(d, fleet, reading);
|
||||
case QUIET_DONE -> {
|
||||
countNudge("exhausted");
|
||||
contextNotified = d.contextNotified();
|
||||
}
|
||||
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN -> contextNotified = d.contextNotified();
|
||||
}
|
||||
scheduleNext();
|
||||
}
|
||||
|
||||
/** Persist the state a decision returned, so the next tick starts from it. */
|
||||
/**
|
||||
* Persist the idle/quiet state a decision returned, so the next tick starts from it.
|
||||
*
|
||||
* <p>fleetd #609 review: the context latch ({@link #contextNotified}) is deliberately <em>not</em>
|
||||
* set here any more. Setting it from the decision unconditionally — before {@link #injectNudge} even
|
||||
* tries to send — is exactly the review's blocker: a decision to notify is not the same fact as "the
|
||||
* notice reached the pane". Every branch of {@link #tick} now assigns {@link #contextNotified} itself,
|
||||
* once it knows whether a send happened and whether it carried the notice (see {@link #injectNudge}).
|
||||
*/
|
||||
private void applyDecision(Decision d) {
|
||||
idleSinceNanos = d.idleSinceNanos() == null ? NOT_IDLE : d.idleSinceNanos();
|
||||
quietCount = d.quietCount();
|
||||
}
|
||||
|
||||
/** Send the nudge to the known lead. */
|
||||
private void injectNudge(FleetState fleet) {
|
||||
/**
|
||||
* Send the nudge to the known lead, with the fleetd #609 context notice appended when it applies, and
|
||||
* persist the context latch based on what actually happened this tick — not merely what {@code d}
|
||||
* chose to attempt.
|
||||
*/
|
||||
private void injectNudge(Decision d, FleetState fleet, LeadContextGauge.Reading reading) {
|
||||
// fleetd #609 review: build the notice from the latch as it stood BEFORE this tick's decision —
|
||||
// d.contextNotified() is the value to persist once delivery is confirmed, not the value the text
|
||||
// itself should be built from. Otherwise a HIGH stretch that is still latched would never see the
|
||||
// notice at all, defeating the very check this fixes.
|
||||
String notice = contextNotice(contextHighNudge, reading, contextNotified, requireOperatorConfirm);
|
||||
var lead = primaryRegistry.primaryTerminal();
|
||||
if (lead.isEmpty()) {
|
||||
return; // the lead disappeared between the decision and the injection
|
||||
}
|
||||
String leadTerminal = lead.get();
|
||||
boolean sent = lead.isPresent() && trySend(lead.get(), fleet.nudgeText() + notice, notice);
|
||||
// The latch becomes true only when all three hold: decide() chose to notify, a notice was
|
||||
// actually included in the text, and the send reached the pane without throwing. Whenever no
|
||||
// notice was attempted (disabled, not HIGH, or already latched), nothing was promised to the lead
|
||||
// this tick, so apply the decision's own carried-forward value unconditionally — that is how the
|
||||
// OK-only re-arm rule and STAND_DOWN's "don't burn the notice" rule keep working through this path
|
||||
// too. A lead that disappeared between the decision and the send (lead.isEmpty()) is treated the
|
||||
// same as a failed send: nothing reached the pane, so the latch must not be set.
|
||||
contextNotified = notice.isEmpty() ? d.contextNotified() : sent;
|
||||
}
|
||||
|
||||
/**
|
||||
* Attempt one herdr send and count its outcome. Returns whether {@code agents.send} returned without
|
||||
* throwing — the caller ({@link #injectNudge}) needs this to decide whether the fleetd #609 context
|
||||
* latch may be persisted as set.
|
||||
*/
|
||||
private boolean trySend(String leadTerminal, String text, String notice) {
|
||||
try {
|
||||
agents.send(leadTerminal, fleet.nudgeText());
|
||||
agents.send(leadTerminal, text);
|
||||
log.debug("idle-heartbeat: nudge sent to lead {} (quiet nudges so far in this stretch: {})",
|
||||
leadTerminal, quietCount);
|
||||
countNudge("sent");
|
||||
// fleetd #609: a nudge that carries the context notice is counted under its own outcome so
|
||||
// it is visible in /metrics — one count per nudge either way, never two.
|
||||
countNudge(notice.isEmpty() ? "sent" : "sent_context");
|
||||
return true;
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("idle-heartbeat: failed to nudge lead {}: {}", leadTerminal, e.toString());
|
||||
countNudge("failed");
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #609: the text appended to a nudge when the lead's own context is full — {@code ""}
|
||||
* whenever the notice does not apply, so callers can unconditionally append this without an extra
|
||||
* branch. Wording stays plain (CEFR B1) and honest about who actually gates the roll — see the
|
||||
* {@code requireOperatorConfirm} overload (fleetd #621) for which check that is. This loop only
|
||||
* ever prints text, it never calls {@code fleet_handover} itself.
|
||||
*
|
||||
* @param enabled the {@code leadHeartbeat.contextHighNudge} config flag
|
||||
* @param reading the lead's current {@link LeadContextGauge} reading
|
||||
* @return the notice text (starting with a leading space, to append directly after {@link
|
||||
* FleetState#nudgeText()}), or {@code ""} when disabled or the state is not {@code HIGH}
|
||||
*/
|
||||
static String contextNotice(boolean enabled, LeadContextGauge.Reading reading) {
|
||||
return contextNotice(enabled, reading, false);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #609 review: as {@link #contextNotice(boolean, LeadContextGauge.Reading)}, but also gated on
|
||||
* {@code alreadyNotified} — the context latch as it stood <em>before</em> the current tick's decision.
|
||||
* Without this gate, every pending-driven {@code INJECT} that lands while the context stays {@code
|
||||
* HIGH} would re-append the full notice on top of an already-latched stretch, making the notice's own
|
||||
* closing sentence ("You will not be told again until your context reads ok.") false. {@link
|
||||
* #injectNudge} is the only caller that passes a non-default {@code alreadyNotified}.
|
||||
*
|
||||
* <p>fleetd #621: delegates with {@code requireOperatorConfirm=true} — the pre-#621 default and the
|
||||
* value every existing caller of this overload (including every test written before #621) already
|
||||
* assumed, so the text this overload returns stays byte-identical.
|
||||
*
|
||||
* @param alreadyNotified whether the lead has already been told about the current HIGH stretch
|
||||
*/
|
||||
static String contextNotice(boolean enabled, LeadContextGauge.Reading reading, boolean alreadyNotified) {
|
||||
return contextNotice(enabled, reading, alreadyNotified, true);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #621: as {@link #contextNotice(boolean, LeadContextGauge.Reading, boolean)}, but the closing
|
||||
* instructions also track the daemon's effective {@code leadRollover.requireOperatorConfirm} value,
|
||||
* instead of always asserting that only the operator can approve the roll.
|
||||
*
|
||||
* <p>{@code LeadRollover.confirm(...)} already honours this flag: when it is {@code false}, the daemon
|
||||
* itself gates the roll on the three handover-file checks alone (exists, modified after the {@code
|
||||
* open()} request, and no older than {@code maxDocAgeSeconds}) and never consults {@code
|
||||
* operatorConfirmed}. Before this parameter existed, this notice told the lead to ask the operator
|
||||
* regardless — so a lead that followed its own instructions asked anyway, and setting the config knob
|
||||
* to {@code false} stopped the daemon refusing the roll without stopping the operator being
|
||||
* interrupted. This parameter is how the text is kept honest about which gate is actually live.
|
||||
*
|
||||
* @param requireOperatorConfirm the effective {@code leadRollover.requireOperatorConfirm} value
|
||||
*/
|
||||
static String contextNotice(boolean enabled, LeadContextGauge.Reading reading, boolean alreadyNotified,
|
||||
boolean requireOperatorConfirm) {
|
||||
if (!enabled || alreadyNotified || reading.state() != LeadContextGauge.State.HIGH) {
|
||||
return "";
|
||||
}
|
||||
StringBuilder sb = new StringBuilder(" Your own context is nearly full");
|
||||
String compactionWord = reading.compactions() == 1 ? "compaction" : "compactions";
|
||||
if (reading.tokens() != null) {
|
||||
sb.append(": ").append(reading.tokens()).append(" tokens used, ")
|
||||
.append(reading.compactions()).append(' ').append(compactionWord).append(" so far.");
|
||||
} else {
|
||||
// A HIGH reading always carries a non-null token count today: LeadContextGauge only
|
||||
// reaches HIGH by comparing a number against HIGH_THRESHOLD_TOKENS. That invariant
|
||||
// lives in another class and nothing asserts it, so this branch does not rely on it —
|
||||
// it drops the token clause rather than printing "null tokens".
|
||||
sb.append(" (").append(reading.compactions()).append(' ').append(compactionWord)
|
||||
.append(" so far).");
|
||||
}
|
||||
if (requireOperatorConfirm) {
|
||||
sb.append(" A fresh session would work better. To hand over: call fleet_handover(action=\"open\"), "
|
||||
+ "write the file it names, ask the operator, then call fleet_handover(action=\"confirm\", "
|
||||
+ "token, operatorConfirmed). Only the operator can approve the roll. You will not be told "
|
||||
+ "again until your context reads ok.");
|
||||
} else {
|
||||
sb.append(" A fresh session would work better. To hand over: call fleet_handover(action=\"open\"), "
|
||||
+ "write the file it names, then call fleet_handover(action=\"confirm\", token). Decide for "
|
||||
+ "yourself when to confirm: the roll goes through if the handover file exists, was "
|
||||
+ "changed after you opened it, and is not older than maxDocAgeSeconds. You will not be "
|
||||
+ "told again until your context reads ok.");
|
||||
}
|
||||
return sb.toString();
|
||||
}
|
||||
|
||||
/** Schedule the next tick on the scheduler thread pool. */
|
||||
private void scheduleNext() {
|
||||
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
|
||||
|
||||
@@ -63,7 +63,7 @@ import java.util.concurrent.TimeoutException;
|
||||
* which messages reached a lead. Any publish still awaiting its confirm is failed rather than left to idle out
|
||||
* the confirm timeout against a sequence number that means nothing on the new channel.
|
||||
*/
|
||||
public final class LeadMailbox implements LeadChannel, AutoCloseable {
|
||||
public final class LeadMailbox implements LeadChannelHandle {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(LeadMailbox.class);
|
||||
|
||||
|
||||
@@ -92,8 +92,31 @@ public final class MessageService {
|
||||
QUESTION,
|
||||
/** Timed out after the message was delivered — the worker is still working. */
|
||||
TIMED_OUT_WORKING,
|
||||
/** Timed out before delivery — the message is still queued for the worker. */
|
||||
/**
|
||||
* Timed out with no confirmed delivery, and the target saw nothing — the message will not
|
||||
* arrive later, so a caller may resend. {@link #send} reaches this outcome through {@link
|
||||
* Injector#cancel} reporting one of two routes: {@link Injector.Cancellation#CANCELLED}
|
||||
* means the message was still queued and this call removed it; {@link
|
||||
* Injector.Cancellation#NOT_DELIVERED} means nothing was ever sent — the queue was cleared
|
||||
* because the target never became ready or was abandoned, or the injector's call to the
|
||||
* target's terminal ({@link dev.ltms.fleet.herdr.AgentControl#send}) failed with a herdr
|
||||
* error that this codebase already treats as a confirmed absence. A third route,
|
||||
* {@link Injector.Cancellation#ATTEMPTED}, used to be folded into this same outcome
|
||||
* (fleetd #571) — it no longer is; see {@link #TIMED_OUT_UNCONFIRMED}.
|
||||
*/
|
||||
TIMED_OUT_QUEUED,
|
||||
/**
|
||||
* Timed out with delivery unknown. {@link #send} reaches this outcome when {@link
|
||||
* Injector#cancel} reports {@link Injector.Cancellation#ATTEMPTED} (fleetd #551): the call
|
||||
* to the target's terminal ({@link dev.ltms.fleet.herdr.AgentControl#send}) was made, but
|
||||
* this caller never observed whether it reached the pane. {@code agent.prompt} pastes
|
||||
* <em>and submits</em> in one call, so the target may already hold a complete, submitted
|
||||
* turn and be working on it right now — the same reality as {@link #TIMED_OUT_WORKING},
|
||||
* just not confirmed. The message may or may not have arrived. Treat this as neither a
|
||||
* confirmed delivery nor a confirmed absence: a caller that resends on this outcome risks a
|
||||
* double delivery — the same brief typed into the pane twice (fleetd #571).
|
||||
*/
|
||||
TIMED_OUT_UNCONFIRMED,
|
||||
/** Another send to this session was in flight for the whole window. */
|
||||
BUSY,
|
||||
/**
|
||||
@@ -299,12 +322,21 @@ public final class MessageService {
|
||||
*/
|
||||
private final ConcurrentHashMap<String, Boolean> strandedReplies = new ConcurrentHashMap<>();
|
||||
/**
|
||||
* Targets whose last send timed out with {@link Outcome#TIMED_OUT_QUEUED} (CB-640) — the
|
||||
* message never reached the {@link Injector} delivery window before the caller's deadline, so
|
||||
* it is still sitting in the injector's own per-target queue. Set where {@link #send} already
|
||||
* computes {@code wasDelivered} for that outcome; no new queue is kept here, only the fact.
|
||||
* Cleared the same way as {@link #strandedReplies}: the next accepted delivery for the target
|
||||
* ({@link #send} opening a fresh waiter) or a teardown ({@link #abandon}).
|
||||
* Targets whose last send timed out with no confirmed delivery (CB-640) — {@link #send} called
|
||||
* {@link Injector#cancel} and got back something other than {@code DELIVERED}. That covers
|
||||
* three histories, not one: {@link Injector.Cancellation#CANCELLED} — the message was still
|
||||
* queued and {@code cancel} removed it right there; {@link Injector.Cancellation#NOT_DELIVERED}
|
||||
* — nothing was ever sent, because the target never became ready, was torn down, or the call to
|
||||
* its terminal failed with a herdr error this codebase already treats as a confirmed absence; or
|
||||
* {@link Injector.Cancellation#ATTEMPTED} (fleetd #551) — the call to the target's terminal was
|
||||
* made and its outcome is unknown, so the target may already hold a complete, submitted turn.
|
||||
* Only the first two mean the message will not arrive later and the target saw nothing; on the
|
||||
* third it may already have arrived in full — and the caller sees a different outcome for it
|
||||
* ({@link Outcome#TIMED_OUT_UNCONFIRMED}, fleetd #571) than for the first two ({@link
|
||||
* Outcome#TIMED_OUT_QUEUED}). Set where {@link #send} already computes {@code wasDelivered} for
|
||||
* that outcome; no queue is kept here, only the fact that the send ended with no confirmed
|
||||
* delivery. Cleared the same way as {@link #strandedReplies}: the next accepted delivery for the
|
||||
* target ({@link #send} opening a fresh waiter) or a teardown ({@link #abandon}).
|
||||
*/
|
||||
private final ConcurrentHashMap<String, Boolean> queuedDeliveries = new ConcurrentHashMap<>();
|
||||
private final AtomicLong ticketSeq = new AtomicLong();
|
||||
@@ -391,12 +423,22 @@ public final class MessageService {
|
||||
|
||||
/**
|
||||
* Read-only delegation fact for fleet views (CB-640): {@code target}'s last send timed out
|
||||
* before the {@link Injector} ever delivered it — the caller saw
|
||||
* {@link Outcome#TIMED_OUT_QUEUED} (see the {@code TimeoutException} branch of {@link #send}),
|
||||
* and the message is still sitting in the injector's per-target queue waiting for the worker
|
||||
* to go idle. Distinct from {@link Outcome#TIMED_OUT_WORKING}, where delivery already happened
|
||||
* and only the reply is outstanding. Cleared the next time this target's delivery is accepted
|
||||
* or the target is abandoned — see {@link #queuedDeliveries}.
|
||||
* with no confirmed delivery — the caller saw {@link Outcome#TIMED_OUT_QUEUED} or {@link
|
||||
* Outcome#TIMED_OUT_UNCONFIRMED} (fleetd #571; see the {@code TimeoutException} branch of
|
||||
* {@link #send}). Despite the method's name, this is not proof that a message is sitting in a
|
||||
* queue: {@link Injector#cancel} reports this outcome through three routes. {@link
|
||||
* Injector.Cancellation#CANCELLED} means the message was still queued and got removed right
|
||||
* there. {@link Injector.Cancellation#NOT_DELIVERED} means nothing was ever sent — the target
|
||||
* never became ready, was torn down, or the call to its terminal failed with a herdr error this
|
||||
* codebase already treats as a confirmed absence. Only these two routes mean the message will
|
||||
* not arrive later, and both report {@code TIMED_OUT_QUEUED}. {@link
|
||||
* Injector.Cancellation#ATTEMPTED} (fleetd #551) means the call to the target's terminal was
|
||||
* made and its outcome is unknown: {@code agent.prompt} pastes <em>and submits</em> in one
|
||||
* call, so on this route the target may already hold a complete, submitted turn and be
|
||||
* working on it right now — it does NOT follow that the target saw nothing, and this route
|
||||
* reports {@code TIMED_OUT_UNCONFIRMED} instead. Distinct from {@link Outcome#TIMED_OUT_WORKING},
|
||||
* where delivery already happened and only the reply is outstanding. Cleared the next time this
|
||||
* target's delivery is accepted or the target is abandoned — see {@link #queuedDeliveries}.
|
||||
*/
|
||||
public boolean hasQueuedDelivery(String target) {
|
||||
return target != null && queuedDeliveries.containsKey(target);
|
||||
@@ -416,15 +458,15 @@ public final class MessageService {
|
||||
/**
|
||||
* Read-only delegation fact for fleet views (CB-640): an async ticket is still
|
||||
* {@link Phase#PENDING} against {@code target}, yet nothing is actually in flight for it — no
|
||||
* open rendezvous waiter ({@link #hasAcceptedDelivery}) and no message still sitting in the
|
||||
* injector's queue ({@link #hasQueuedDelivery}). A healthy PENDING ticket can briefly look this
|
||||
* way while its virtual thread has not yet been scheduled or is blocked on the session lock
|
||||
* behind another send to the same target, so this is a snapshot fact for the health classifier
|
||||
* to weigh across ticks, not proof on its own that the ticket is stuck. It also genuinely
|
||||
* persists — not just as a passing race — once an async {@code fleet_ask} lapses unanswered:
|
||||
* {@link #ask} clears the ticket's question and returns it to {@code PENDING}, but {@link #send}
|
||||
* already closed the forward waiter the instant the question surfaced, so the target has
|
||||
* neither an accepted nor a queued delivery left to show for it.
|
||||
* open rendezvous waiter ({@link #hasAcceptedDelivery}) and no record of a send that ended with
|
||||
* no confirmed delivery ({@link #hasQueuedDelivery}). A healthy PENDING ticket can briefly look
|
||||
* this way while its virtual thread has not yet been scheduled or is blocked on the session
|
||||
* lock behind another send to the same target, so this is a snapshot fact for the health
|
||||
* classifier to weigh across ticks, not proof on its own that the ticket is stuck. It also
|
||||
* genuinely persists — not just as a passing race — once an async {@code fleet_ask} lapses
|
||||
* unanswered: {@link #ask} clears the ticket's question and returns it to {@code PENDING}, but
|
||||
* {@link #send} already closed the forward waiter the instant the question surfaced, so the
|
||||
* target has neither an accepted nor a queued delivery left to show for it.
|
||||
*
|
||||
* <p><strong>Deliberately still {@code question == null} only (fleetd #275).</strong> This
|
||||
* method must not also report a still-{@link Phase#ASKING} task as orphaned: the worker may
|
||||
@@ -611,7 +653,7 @@ public final class MessageService {
|
||||
return switch (o) {
|
||||
case REPLIED -> "replied";
|
||||
case COMPLETED_UNREPLIED -> "completion_fallback";
|
||||
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> "timeout";
|
||||
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, TIMED_OUT_UNCONFIRMED, BUSY -> "timeout";
|
||||
case WORKER_FAILED -> "failed";
|
||||
case BACKEND_EXHAUSTED -> "backend_exhausted";
|
||||
case STALE_TURN, QUESTION -> null; // not a completed delegation
|
||||
@@ -713,7 +755,7 @@ public final class MessageService {
|
||||
* failure.
|
||||
*
|
||||
* <p><strong>Without {@code sweepAsking} on the release path, a target torn down while
|
||||
* genuinely {@code ASKING} was unrecoverable.</strong> {@link #resolveQuestion} had already
|
||||
* genuinely {@code ASKING} was unrecoverable.</strong> {@link Rendezvous#resolveQuestion(String, String, String)} had already
|
||||
* closed the forward waiter the instant the question surfaced (so the {@code waiter} branch
|
||||
* below finds nothing to fail), the {@code question == null} guard excluded the task from
|
||||
* {@code matching} (so the loop below skipped it too), and the worker's own {@code fleet_ask}
|
||||
@@ -940,6 +982,7 @@ public final class MessageService {
|
||||
} catch (TimeoutException e) {
|
||||
boolean wasDelivered = delivery.completion().isDone()
|
||||
&& !delivery.completion().isCompletedExceptionally();
|
||||
Injector.Cancellation cancellation = null;
|
||||
if (!wasDelivered) {
|
||||
if (timeoutCancellationRaceHookForTest != null) {
|
||||
// Test-only (fleetd #345): see the field's own javadoc.
|
||||
@@ -947,16 +990,27 @@ public final class MessageService {
|
||||
}
|
||||
// The target monitor makes cancellation atomic with onStatus picking this
|
||||
// Pending up. If pickup won, report TIMED_OUT_WORKING because the text landed.
|
||||
wasDelivered = injector.cancel(delivery) == Injector.Cancellation.DELIVERED;
|
||||
cancellation = injector.cancel(delivery);
|
||||
wasDelivered = cancellation == Injector.Cancellation.DELIVERED;
|
||||
}
|
||||
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
|
||||
if (!wasDelivered) {
|
||||
// CB-640: record that delivery did not happen for fleet health (see
|
||||
// queuedDeliveries). The exact Pending was cancelled, so it cannot arrive later.
|
||||
Outcome outcome;
|
||||
if (wasDelivered) {
|
||||
outcome = Outcome.TIMED_OUT_WORKING;
|
||||
} else if (cancellation == Injector.Cancellation.ATTEMPTED) {
|
||||
// fleetd #571: the call to the target's terminal was made and its outcome is
|
||||
// unknown — the message may already have arrived in full, so this must not
|
||||
// be reported as TIMED_OUT_QUEUED, which promises it never will.
|
||||
outcome = Outcome.TIMED_OUT_UNCONFIRMED;
|
||||
} else {
|
||||
// CB-640: record that delivery is not confirmed, for fleet health (see
|
||||
// queuedDeliveries). cancellation is CANCELLED (this call removed a
|
||||
// still-queued Pending) or NOT_DELIVERED (an earlier attempt already failed
|
||||
// with a confirmed absence) — both mean the target saw nothing.
|
||||
queuedDeliveries.put(target, Boolean.TRUE);
|
||||
outcome = Outcome.TIMED_OUT_QUEUED;
|
||||
}
|
||||
return recorded(new Reply(
|
||||
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
|
||||
return recorded(new Reply(outcome, null));
|
||||
} catch (ExecutionException e) {
|
||||
Throwable cause = e.getCause();
|
||||
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
|
||||
@@ -1096,21 +1150,33 @@ public final class MessageService {
|
||||
}
|
||||
try {
|
||||
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession);
|
||||
// #282: mirror send()'s registration (:802) so a SECOND fleet_ask inside this same
|
||||
// resumed turn can re-associate the async ticket with its new turnId via
|
||||
// markAsyncQuestion — without this, that second ask has no Task to attach to, and
|
||||
// markAsyncQuestion silently returns null.
|
||||
Task task = asyncTasksByTurn.get(turnId);
|
||||
if (task != null) {
|
||||
asyncTasksByWaiter.put(reply, task);
|
||||
}
|
||||
if (!rendezvous.answerAsk(turnId, content)) {
|
||||
asyncTasksByWaiter.remove(reply);
|
||||
rendezvous.close(workerSession, reply);
|
||||
return new Reply(Outcome.STALE_TURN, null); // lapsed between the lookup and the unblock
|
||||
}
|
||||
clearAsyncQuestion(turnId, false);
|
||||
// fleetd #575: this try used to open below, AFTER the Task lookup/registration and the
|
||||
// STALE_TURN early return that follows it — so that return was covered only by a
|
||||
// hand-rolled copy of the finally's own cleanup pair, not the finally itself. Widening the
|
||||
// try up to wrap the registration closes the gap structurally: every exit from here on,
|
||||
// STALE_TURN included, now runs through the one finally below exactly once, and the
|
||||
// duplicated pair is gone. This did not fix a live leak — see the ticket: neither
|
||||
// rendezvous.answerAsk nor clearAsyncQuestion(turnId, false) can throw, so nothing ever
|
||||
// actually left through the old gap uncovered — but #572 found this exact drift (one
|
||||
// finally asserted, an identical sibling not) on this same file, and the hand-rolled copy
|
||||
// was the wrong shape to keep regardless of whether it was ever exercised.
|
||||
try {
|
||||
// #282: mirror send()'s registration (:802) so a SECOND fleet_ask inside this same
|
||||
// resumed turn can re-associate the async ticket with its new turnId via
|
||||
// markAsyncQuestion — without this, that second ask has no Task to attach to, and
|
||||
// markAsyncQuestion silently returns null.
|
||||
Task task = asyncTasksByTurn.get(turnId);
|
||||
if (task != null) {
|
||||
asyncTasksByWaiter.put(reply, task);
|
||||
}
|
||||
if (answerAskLapseRaceHookForTest != null) {
|
||||
// Test-only (fleetd #575): see the field's own javadoc.
|
||||
answerAskLapseRaceHookForTest.run();
|
||||
}
|
||||
if (!rendezvous.answerAsk(turnId, content)) {
|
||||
return new Reply(Outcome.STALE_TURN, null); // lapsed between the lookup and the unblock
|
||||
}
|
||||
clearAsyncQuestion(turnId, false);
|
||||
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
|
||||
Reply result = new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
|
||||
// #282: this waiter can resolve with a FRESH question rather than a terminal reply —
|
||||
@@ -1610,6 +1676,30 @@ public final class MessageService {
|
||||
this.askTimeoutRaceHookForTest = hook;
|
||||
}
|
||||
|
||||
/**
|
||||
* Null in production; test seam for fleetd #575 — invoked from {@link #answer}, right after this
|
||||
* call's own {@code Task} registration and right before its {@code rendezvous.answerAsk(turnId,
|
||||
* content)} call. A test installs this to complete the SAME turnId's ask directly via {@link
|
||||
* Rendezvous#answerAsk} from inside that exact window, deterministically reproducing what a
|
||||
* second, concurrent {@code answer()} call racing to unblock the same ask can otherwise only
|
||||
* win by timing luck: this call's own {@code askSession(turnId)} lookup at the top already saw
|
||||
* the ask as open, but by the time it reaches {@code rendezvous.answerAsk} here, the other call
|
||||
* already completed it (or the worker's own {@code ask()} teardown already closed it) — so this
|
||||
* call must see {@code false} and return {@link Outcome#STALE_TURN}, exactly the "lapsed between
|
||||
* the lookup and the unblock" case named at that call site. Proves the #575 fix (widening this
|
||||
* method's try so a single finally covers this exit) does not change that outcome and still
|
||||
* cleans this call's own {@code reply} up exactly once.
|
||||
*/
|
||||
private volatile Runnable answerAskLapseRaceHookForTest;
|
||||
|
||||
/**
|
||||
* Test-only (fleetd #575): install {@link #answerAskLapseRaceHookForTest}. Package-private so the
|
||||
* test, in the same package, can reach it without widening any production API.
|
||||
*/
|
||||
void setAnswerAskLapseRaceHookForTest(Runnable hook) {
|
||||
this.answerAskLapseRaceHookForTest = hook;
|
||||
}
|
||||
|
||||
/** A new send must not open a waiter while an async ticket owns this worker's paused turn. */
|
||||
private boolean hasAsyncQuestion(String target) {
|
||||
return asyncTasksByTurn.values().stream().anyMatch(task -> target.equals(task.target));
|
||||
|
||||
@@ -44,7 +44,7 @@ import java.util.stream.Collectors;
|
||||
* still holds an unacked message ({@link #pendingReplies}), tickets not yet collected
|
||||
* ({@link #pendingTickets}), and open questions not yet answered or lapsed
|
||||
* ({@link #pendingQuestions}) — and sends at most one combined nudge per tick
|
||||
* ({@link #injectNudge(String, int, int, int)}). Work that arrives while the lead is busy is
|
||||
* ({@link #injectNudge(String, int, int, int, int, int)}). Work that arrives while the lead is busy is
|
||||
* never lost: it is re-read fresh on every tick until the lead is injectable or its own reminder
|
||||
* cap ({@link #maxReminders}) is reached — each source spends from its own budget, so one source
|
||||
* exhausting its cap does not stop nudges about the others (post-CB-590 regression fix; see
|
||||
|
||||
@@ -29,8 +29,8 @@ public enum MemberRole {
|
||||
* <p>Reads the repo and writes analysis. Never commits code and never opens a pull request —
|
||||
* an architect that starts implementing has stopped doing the job that makes it useful.
|
||||
*
|
||||
* <p>Architects are the one member kind declared in config, because a lead addresses the same
|
||||
* slots across many tickets and needs a stable name for them.
|
||||
* <p>Architects are the one member kind with live slot binding, because a lead addresses the
|
||||
* same slots across many tickets and needs a stable name for them.
|
||||
*/
|
||||
ARCHITECT,
|
||||
|
||||
@@ -43,6 +43,14 @@ public enum MemberRole {
|
||||
*/
|
||||
DEV,
|
||||
|
||||
/**
|
||||
* Sweeps an assigned package for defects and reports several ranked findings.
|
||||
*
|
||||
* <p>Never changes code, commits, or opens a pull request. A hunt gathers evidence, which can
|
||||
* include running the build, but leaves every fix to a later implementation unit.
|
||||
*/
|
||||
HUNTER,
|
||||
|
||||
/**
|
||||
* Reviews a diff it did not write and reports one structured finding.
|
||||
*
|
||||
@@ -59,7 +67,7 @@ public enum MemberRole {
|
||||
|
||||
/**
|
||||
* The {@code fleet:} block that holds this role's pool — {@code architects},
|
||||
* {@code developers}, {@code reviewers}.
|
||||
* {@code developers}, {@code hunters}, {@code reviewers}.
|
||||
*
|
||||
* <p>Plural, and not always the wire name: the pool of things a {@code dev} may run on reads
|
||||
* naturally as {@code developers:}. The wire name stays the singular {@code dev}, because that
|
||||
@@ -69,6 +77,7 @@ public enum MemberRole {
|
||||
return switch (this) {
|
||||
case ARCHITECT -> "architects";
|
||||
case DEV -> "developers";
|
||||
case HUNTER -> "hunters";
|
||||
case REVIEWER -> "reviewers";
|
||||
};
|
||||
}
|
||||
|
||||
@@ -94,6 +94,7 @@ public final class FleetApp {
|
||||
// .none() (the honest "feature not wired" view) for every constructor that does not pass one.
|
||||
private final FleetMcp.QuarantineSource quarantine;
|
||||
private final FleetMcp.OutageSource outage;
|
||||
private final FleetMcp.LoopHealthSource loopHealth;
|
||||
private final ObjectMapper mapper = new ObjectMapper();
|
||||
|
||||
/**
|
||||
@@ -155,7 +156,8 @@ public final class FleetApp {
|
||||
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics,
|
||||
Predicate<String> deliverable, Supplier<MemberCredentialPolicyView> memberCredentials) {
|
||||
this(herdr, memberHerdr, workers, sessions, messages, presence, mcpServlet, auth, metrics,
|
||||
deliverable, memberCredentials, FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none());
|
||||
deliverable, memberCredentials, FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
|
||||
FleetMcp.LoopHealthSource.none());
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -169,8 +171,18 @@ public final class FleetApp {
|
||||
public FleetApp(HerdrClient herdr, HerdrClient memberHerdr, PeerLauncher workers, SessionManager sessions,
|
||||
MessageService messages, MemberPresence presence,
|
||||
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics,
|
||||
Predicate<String> deliverable, Supplier<MemberCredentialPolicyView> memberCredentials,
|
||||
FleetMcp.QuarantineSource quarantine, FleetMcp.OutageSource outage) {
|
||||
Predicate<String> deliverable, Supplier<MemberCredentialPolicyView> memberCredentials,
|
||||
FleetMcp.QuarantineSource quarantine, FleetMcp.OutageSource outage) {
|
||||
this(herdr, memberHerdr, workers, sessions, messages, presence, mcpServlet, auth, metrics, deliverable,
|
||||
memberCredentials, quarantine, outage, FleetMcp.LoopHealthSource.none());
|
||||
}
|
||||
|
||||
public FleetApp(HerdrClient herdr, HerdrClient memberHerdr, PeerLauncher workers, SessionManager sessions,
|
||||
MessageService messages, MemberPresence presence,
|
||||
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics,
|
||||
Predicate<String> deliverable, Supplier<MemberCredentialPolicyView> memberCredentials,
|
||||
FleetMcp.QuarantineSource quarantine, FleetMcp.OutageSource outage,
|
||||
FleetMcp.LoopHealthSource loopHealth) {
|
||||
this.herdr = herdr;
|
||||
this.memberHerdr = memberHerdr != null ? memberHerdr : herdr;
|
||||
this.workers = workers;
|
||||
@@ -183,6 +195,7 @@ public final class FleetApp {
|
||||
this.memberCredentials = memberCredentials != null ? memberCredentials : MemberCredentialPolicyView::absent;
|
||||
this.quarantine = quarantine != null ? quarantine : FleetMcp.QuarantineSource.none();
|
||||
this.outage = outage != null ? outage : FleetMcp.OutageSource.none();
|
||||
this.loopHealth = loopHealth != null ? loopHealth : FleetMcp.LoopHealthSource.none();
|
||||
}
|
||||
|
||||
/** Wire routes onto a fresh, unstarted Javalin instance. Caller starts it. */
|
||||
@@ -291,15 +304,19 @@ public final class FleetApp {
|
||||
* spotted by comparing two numbers by eye.
|
||||
*/
|
||||
private void healthz(Context ctx) {
|
||||
HealthzResponse response = healthzResponse(herdr, memberHerdr, loopHealth);
|
||||
ctx.status(response.status()).json(response.body());
|
||||
}
|
||||
|
||||
record HealthzResponse(int status, Map<String, Object> body) { }
|
||||
|
||||
static HealthzResponse healthzResponse(HerdrClient herdr, HerdrClient memberHerdr,
|
||||
FleetMcp.LoopHealthSource loopHealth) {
|
||||
JsonNode pong;
|
||||
try {
|
||||
pong = herdr.call("ping");
|
||||
} catch (HerdrException e) {
|
||||
ctx.status(503).json(Map.of(
|
||||
"status", "degraded",
|
||||
"herdr", "unreachable",
|
||||
"detail", e.getMessage()));
|
||||
return;
|
||||
return degradedResponse("unreachable", e.getMessage(), loopHealth);
|
||||
}
|
||||
Map<String, Object> body = new LinkedHashMap<>();
|
||||
body.put("status", "ok");
|
||||
@@ -311,11 +328,11 @@ public final class FleetApp {
|
||||
try {
|
||||
memberPong = memberHerdr.call("ping");
|
||||
} catch (HerdrException e) {
|
||||
ctx.status(503).json(Map.of(
|
||||
return new HealthzResponse(503, Map.of(
|
||||
"status", "degraded",
|
||||
"herdr", "member unreachable",
|
||||
"detail", e.getMessage()));
|
||||
return;
|
||||
"detail", e.getMessage(),
|
||||
"loopHealth", loopHealthView(loopHealth)));
|
||||
}
|
||||
int leadProtocol = pong.path("protocol").asInt();
|
||||
int memberProtocol = memberPong.path("protocol").asInt();
|
||||
@@ -326,7 +343,20 @@ public final class FleetApp {
|
||||
body.put("protocolMismatch", true);
|
||||
}
|
||||
}
|
||||
ctx.status(200).json(body);
|
||||
body.put("loopHealth", loopHealthView(loopHealth));
|
||||
return new HealthzResponse(200, body);
|
||||
}
|
||||
|
||||
private static Map<String, String> loopHealthView(FleetMcp.LoopHealthSource loopHealth) {
|
||||
return Map.of("statusPoller", loopHealth.statusPoller().get().name(),
|
||||
"sessionReaper", loopHealth.sessionReaper().get().name());
|
||||
}
|
||||
|
||||
private static HealthzResponse degradedResponse(String herdr, String detail,
|
||||
FleetMcp.LoopHealthSource loopHealth) {
|
||||
return new HealthzResponse(503, Map.of(
|
||||
"status", "degraded", "herdr", herdr, "detail", detail,
|
||||
"loopHealth", loopHealthView(loopHealth)));
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -640,15 +670,31 @@ public final class FleetApp {
|
||||
}
|
||||
default -> ctx.status(202).json(Map.of(
|
||||
"sessionId", id,
|
||||
// fleetd #571 (ticket comment 17126): no `default` here on purpose. This switch
|
||||
// is an expression, so the compiler already demands every Outcome constant have
|
||||
// an arm — adding an 11th constant to Outcome is a compile error here, not a
|
||||
// silent fall-through. That is exactly the bug this ticket exists to fix:
|
||||
// `default -> "done"` used to sit here and would have told a REST caller the
|
||||
// delegation completed for TIMED_OUT_UNCONFIRMED, the one outcome where delivery
|
||||
// is unknown. REPLIED, COMPLETED_UNREPLIED, QUESTION and STALE_TURN can never
|
||||
// actually reach this inner switch — the outer switch above always dispatches
|
||||
// them first — but they still need an arm to keep this switch exhaustive.
|
||||
"status", switch (reply.outcome()) {
|
||||
case TIMED_OUT_WORKING -> "working";
|
||||
case TIMED_OUT_QUEUED -> "queued";
|
||||
// Delivery here is unknown, not merely still queued — see
|
||||
// Outcome#TIMED_OUT_UNCONFIRMED's own javadoc.
|
||||
case TIMED_OUT_UNCONFIRMED -> "unconfirmed";
|
||||
case BUSY -> "busy";
|
||||
case WORKER_FAILED -> "failed";
|
||||
case BACKEND_EXHAUSTED -> "backend_exhausted";
|
||||
default -> "done"; // unreachable (terminal outcomes handled above)
|
||||
case REPLIED, COMPLETED_UNREPLIED, QUESTION, STALE_TURN -> "done"; // unreachable
|
||||
},
|
||||
"detail", (reply.outcome() == MessageService.Outcome.WORKER_FAILED
|
||||
"detail", reply.outcome() == MessageService.Outcome.TIMED_OUT_UNCONFIRMED
|
||||
? "no reply within " + timeout + "ms; delivery is unconfirmed — the "
|
||||
+ "message may already have reached the worker, so a resend "
|
||||
+ "risks sending it twice; poll status first"
|
||||
: (reply.outcome() == MessageService.Outcome.WORKER_FAILED
|
||||
|| reply.outcome() == MessageService.Outcome.BACKEND_EXHAUSTED)
|
||||
&& reply.text() != null
|
||||
? reply.text()
|
||||
|
||||
@@ -871,10 +871,23 @@ public final class SessionManager implements TurnListener {
|
||||
// digest lets a lead tell at a glance whether all members got the same charter; the source
|
||||
// records whether a role charter was configured ("fleet.charters.<role>") or only the reply
|
||||
// charter was composed ("none").
|
||||
// #604: charterBytes rides along with charterSha256, not with charterSource — it is only
|
||||
// meaningful as the digest's companion (a length turns "they differ" into "by how much").
|
||||
// A member with no composed charter reports charterSource and nothing else, as before.
|
||||
//
|
||||
// Gating on the DIGEST rather than on the receipt is deliberate, and the two are not
|
||||
// always null together. CharterReceipt.compose() derives the digest with digestOf(), which
|
||||
// returns null for BLANK text, while the byte count is composed.getBytes().length, which
|
||||
// does not. So a whitespace-only charter (a blank fleet.charters.<role> on a profile with
|
||||
// no MCP, so no reply charter is appended) yields a null digest beside a non-zero size.
|
||||
// Reporting a size with no digest would say "they differ by N bytes" about an artifact we
|
||||
// cannot fingerprint, so this gate omits both. Never widen it to the receipt-level null
|
||||
// check without deciding what that case should report.
|
||||
if (session.charterReceipt() != null) {
|
||||
m.put("charterSource", session.charterReceipt().charterSource());
|
||||
if (session.charterReceipt().charterSha256() != null) {
|
||||
m.put("charterSha256", session.charterReceipt().charterSha256());
|
||||
m.put("charterBytes", session.charterReceipt().charterBytes());
|
||||
}
|
||||
}
|
||||
m.put("liveStatus", live == null ? "unknown" : live.status().name().toLowerCase());
|
||||
@@ -1074,6 +1087,29 @@ public final class SessionManager implements TurnListener {
|
||||
* members were live) and must still produce the line. Both {@link #drainSnapshot} passes (the
|
||||
* main snapshot and the straggler sweep) are folded into the one line: a caller reading two
|
||||
* lines could not tell a two-pass drain from two separate drains.
|
||||
*
|
||||
* <p><strong>Non-goal: this line must never move into a {@code finally} block, and this method
|
||||
* must never grow one around it.</strong> "Every time" above means every time the drain
|
||||
* <em>finishes</em>, not every time this method exits. The absence of the line is the signal
|
||||
* that the drain died, so a {@code finally} would destroy the signal and print confident
|
||||
* partial counts in the same edit — the line would appear after a drain that threw, carrying
|
||||
* whatever {@code tally} it had reached. Both halves of the value are lost at once. The line
|
||||
* has to be the last statement of the successful path and reachable only from it.
|
||||
*
|
||||
* <p>This is written down because it is the obvious review comment ("shouldn't we always log
|
||||
* the drain result?"), it sounds like thoroughness, and the paragraph above reads as an
|
||||
* invitation to it. Raised by the fleet01 lead on 2026-09-12, from their 2026-09-10 incident:
|
||||
* we only know that drain died on that host because it <em>threw</em>, and a
|
||||
* {@code NoClassDefFoundError} reached the JVM's uncaught handler. A drain that hung on one
|
||||
* session, or returned early on a condition rather than an exception, would leave no stack
|
||||
* trace, no {@code ERROR} token and no priority — only a missing line. That makes the loud
|
||||
* variant the one we have seen and the quiet variants the ones this line exists to catch.
|
||||
*
|
||||
* <p>Related: {@code released} and {@code abandoned} are counted incrementally inside {@link
|
||||
* #drainSnapshot}'s loop and folded with {@link DrainTally#plus}, rather than derived from a
|
||||
* collection read at the end, for the same reason. If a partial report is ever wanted it must
|
||||
* be a different line with a different verb. One line must not serve both, or a reader cannot
|
||||
* tell a finished drain from an interrupted one by its wording.
|
||||
*/
|
||||
void drainAll(long timeoutNanos) {
|
||||
long deadline = System.nanoTime() + timeoutNanos;
|
||||
|
||||
@@ -1,9 +1,11 @@
|
||||
package dev.ltms.fleet.session;
|
||||
|
||||
import dev.ltms.fleet.inject.LoopWatchdog;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
/**
|
||||
* Periodic virtual-thread reaper that tears down {@code READY}/{@code DONE} sessions which have
|
||||
@@ -14,6 +16,16 @@ public final class SessionReaper {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
|
||||
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
|
||||
/**
|
||||
* How many multiples of {@code intervalMillis} the last-completed round may age before
|
||||
* {@link #health()} reports {@link LoopWatchdog.State#STALLED} (fleetd #544). One round here
|
||||
* is {@code sessions.reapIdle} plus the (rare, best-effort) WIP sweep — both normally finish
|
||||
* in a small fraction of one interval. At the default 5s interval this puts the threshold at
|
||||
* 60s: generous enough that an occasional slow git call in the WIP sweep never trips it, short
|
||||
* enough that a genuinely wedged reap round (mirroring the herdr-read-with-no-deadline hang
|
||||
* {@code StatusPoller} guards against — see fleetd #544) is caught inside about a minute.
|
||||
*/
|
||||
private static final long STALE_AFTER_INTERVAL_MULTIPLIER = 12;
|
||||
/** CB-586: the refs/wip age floor — never sweep a snapshot younger than 24h (the CB-586 rule). */
|
||||
private static final long WIP_MIN_AGE_MILLIS = TimeUnit.HOURS.toMillis(24);
|
||||
/**
|
||||
@@ -25,6 +37,7 @@ public final class SessionReaper {
|
||||
private final SessionManager sessions;
|
||||
private final long idleTtlNanos;
|
||||
private final long intervalMillis;
|
||||
private final LoopWatchdog watchdog;
|
||||
private volatile boolean running;
|
||||
private Thread thread;
|
||||
/**
|
||||
@@ -45,29 +58,61 @@ public final class SessionReaper {
|
||||
|
||||
/** Construct a reaper with an explicit polling interval (useful for tests). */
|
||||
public SessionReaper(SessionManager sessions, long idleTtlSeconds, long intervalMillis) {
|
||||
this(sessions, idleTtlSeconds, intervalMillis, System::nanoTime);
|
||||
}
|
||||
|
||||
/** Full constructor — for tests: an injectable monotonic clock (fleetd #544). */
|
||||
SessionReaper(SessionManager sessions, long idleTtlSeconds, long intervalMillis, LongSupplier nowNanos) {
|
||||
this.sessions = sessions;
|
||||
this.idleTtlNanos = TimeUnit.SECONDS.toNanos(idleTtlSeconds);
|
||||
this.intervalMillis = intervalMillis;
|
||||
this.watchdog = new LoopWatchdog(nowNanos, staleAfterNanos(intervalMillis));
|
||||
}
|
||||
|
||||
private static long staleAfterNanos(long intervalMillis) {
|
||||
return TimeUnit.MILLISECONDS.toNanos(Math.max(intervalMillis, 1)) * STALE_AFTER_INTERVAL_MULTIPLIER;
|
||||
}
|
||||
|
||||
/**
|
||||
* This loop's progress fact (fleetd #544): {@code RUNNING}, {@code STALLED} (dead or parked —
|
||||
* indistinguishable from outside, and never observed on purpose), or {@code STOPPED}
|
||||
* ({@link #stop()} was called). See {@link LoopWatchdog} for why staleness — not thread
|
||||
* liveness — is the signal.
|
||||
*/
|
||||
public LoopWatchdog.State health() {
|
||||
return watchdog.state();
|
||||
}
|
||||
|
||||
/** Start the reaper loop on a virtual thread. Idempotent. */
|
||||
public synchronized void start() {
|
||||
if (running) return;
|
||||
running = true;
|
||||
watchdog.reset(); // fleetd #544: a fresh grace period, not last run's stale mark/timestamp.
|
||||
thread = Thread.ofVirtual().name("session-reaper").start(this::loop);
|
||||
log.info("session reaper started (idle ttl {}s, interval {}ms)",
|
||||
TimeUnit.NANOSECONDS.toSeconds(idleTtlNanos), intervalMillis);
|
||||
}
|
||||
|
||||
private void loop() {
|
||||
while (running) {
|
||||
try {
|
||||
sessions.reapIdle(idleTtlNanos);
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("session reaper iteration failed; continuing", e);
|
||||
try {
|
||||
while (running) {
|
||||
try {
|
||||
sessions.reapIdle(idleTtlNanos);
|
||||
} catch (Throwable e) {
|
||||
log.error("session reaper iteration failed; continuing", e);
|
||||
}
|
||||
maybeSweepWipRefs();
|
||||
// fleetd #544: a round is "complete" only after reapIdle AND the WIP-sweep gate have
|
||||
// both returned — a hang in either (e.g. a wedged git call) means this line is never
|
||||
// reached, so the watchdog goes stale exactly like a dead loop would.
|
||||
watchdog.recordRoundComplete();
|
||||
sleep();
|
||||
}
|
||||
maybeSweepWipRefs();
|
||||
sleep();
|
||||
} finally {
|
||||
if (running) {
|
||||
log.error("session reaper loop exited unexpectedly; it can be restarted");
|
||||
}
|
||||
running = false;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -90,8 +135,8 @@ public final class SessionReaper {
|
||||
log.info("refs/wip retention sweep deleted {} snapshot ref(s) older than 24h whose "
|
||||
+ "content was already reachable from main", deleted);
|
||||
}
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("refs/wip retention sweep failed; continuing", e);
|
||||
} catch (Throwable e) {
|
||||
log.error("refs/wip retention sweep failed; continuing", e);
|
||||
}
|
||||
// Set even when the sweep threw, so a broken repo is retried on the slow cadence rather
|
||||
// than hammering git on every 5-second iteration.
|
||||
@@ -111,6 +156,7 @@ public final class SessionReaper {
|
||||
/** Stop the reaper loop. Idempotent. */
|
||||
public synchronized void stop() {
|
||||
running = false;
|
||||
watchdog.markStoppedByCaller(); // fleetd #544: this halt is on purpose — health() must say STOPPED, not STALLED.
|
||||
if (thread != null) thread.interrupt();
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,222 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.PaneLocator;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.CopyOnWriteArrayList;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
|
||||
/**
|
||||
* fleetd #612 step 2, unit B2 (CB-185, identity half). Replaces the deleted
|
||||
* {@code FleetdConnectionIdentityConstructionTest}, which pinned this claim by reading {@code
|
||||
* Fleetd.java}'s source text for {@code "new PaneLocator(herdr, memberHerdr)"}. That claim moved
|
||||
* to {@code FleetdAssembly.java} (fleetd #612 Unit A) and is pinned here instead, by driving the
|
||||
* real {@link ConnectionIdentity} — via {@code runtime.mcp().identity()}, not a copy — that {@link
|
||||
* FleetdAssembly#assembleAndStart} built.
|
||||
*
|
||||
* <p><strong>What this guards against</strong> (from the deleted test's own javadoc): pinning
|
||||
* {@code PaneLocator} to {@code memberHerdr} alone leaves every LEAD's own MCP connection
|
||||
* unresolvable ({@code callerTerminal == null}) the moment {@code memberHerdrSocket} names a
|
||||
* second daemon, which breaks {@code fleet_reply}/{@code fleet_ask}/{@code fleet_whoami} for a
|
||||
* lead. {@code PaneLocatorTest} already proves {@link PaneLocator} itself can search two clients
|
||||
* given two — the gap this pins is that the assembly actually passes it two, and in the right
|
||||
* order (lead first).
|
||||
*
|
||||
* <p><strong>Why this cannot be driven through a real MCP/HTTP round trip.</strong> The natural
|
||||
* way to observe {@code ConnectionIdentity} would be a real {@code fleet_whoami} call over the
|
||||
* built {@code FleetMcp}, the way {@code FleetMcpContextExtractorTest} drives its own
|
||||
* hand-built one. That does not work for the REAL assembly, because {@code FleetdAssembly} wires
|
||||
* {@code ConnectionIdentity} with a hardcoded {@code new LsofPeerPidLookup()} (see {@code
|
||||
* FleetdAssembly.java:444}), and {@code LsofPeerPidLookup} explicitly excludes its own PID — see
|
||||
* its javadoc: "we exclude our own PID and take the other end". In a JUnit test the HTTP client
|
||||
* and the daemon under test run in the very same JVM, so the "client" and "server" ends of the
|
||||
* loopback connection ARE the same PID, and {@code pidForLocalPort} always returns {@code -1}
|
||||
* before {@link PaneLocator} is ever reached — proving nothing about which daemon(s) got searched.
|
||||
* This test instead reaches the real {@link PaneLocator} the assembly built (through {@link
|
||||
* ConnectionIdentity#panes()}, added for exactly this) and drives it with a chosen pid directly,
|
||||
* bypassing the OS-dependent PID lookup entirely — a legitimate substitute, since the pid lookup
|
||||
* is not what CB-185 is about.
|
||||
*/
|
||||
class FleetdAssemblyConnectionIdentityTest {
|
||||
|
||||
/** Same shape as {@code FleetdAssemblyLifecycleTest}'s fake, but keys {@code connectHerdr} by
|
||||
* socket path so the lead and member daemons can be two DIFFERENT {@link FakeHerdr}s. */
|
||||
private static final class TwoHerdrResourcePorts implements ResourcePorts {
|
||||
|
||||
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
|
||||
final CopyOnWriteArrayList<ScheduledExecutorService> schedulers = new CopyOnWriteArrayList<>();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
HerdrClient client = herdrsBySocket.get(socketPath);
|
||||
if (client == null) {
|
||||
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
|
||||
}
|
||||
return client;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> new dev.ltms.fleet.msg.InMemoryReplyInbox();
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||
schedulers.add(scheduler);
|
||||
return scheduler;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Deliberately never bind — this test never issues a real HTTP request.
|
||||
}
|
||||
}
|
||||
|
||||
private FleetdRuntime runtime;
|
||||
private TwoHerdrResourcePorts ports;
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
if (ports != null && ports.shutdownHook != null) {
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
|
||||
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
|
||||
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: "%s"
|
||||
memberHerdrSocket: "%s"
|
||||
lifecycle:
|
||||
idleTtlSeconds: 600
|
||||
health:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
""".formatted(LEAD_SOCKET, MEMBER_SOCKET));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
private FleetdRuntime assemble(Path dir, FakeHerdr lead, FakeHerdr member) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
ports = new TwoHerdrResourcePorts();
|
||||
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
|
||||
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
|
||||
runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
return runtime;
|
||||
}
|
||||
|
||||
/**
|
||||
* The pin. {@code lead} carries the one pane {@link FakeHerdr}'s canned {@code
|
||||
* pane.process_info} ties to {@link FakeHerdr#WORKER_PID} (pane {@code w2:p7}); {@code member}
|
||||
* reports NO panes at all ({@link FakeHerdr#withNoPanes()}) — modelling a second daemon that
|
||||
* simply does not host the caller's pane, exactly the CB-185 javadoc's scenario for a lead's
|
||||
* own connection. If {@code PaneLocator} only ever searches the member daemon (the bug), this
|
||||
* pid resolves to nothing, because the pane that owns it lives on the LEAD daemon the bug
|
||||
* skips.
|
||||
*/
|
||||
@Test
|
||||
void connectionIdentitySearchesTheLeadDaemonNotJustTheMemberOne(@TempDir Path dir) throws Exception {
|
||||
FakeHerdr lead = new FakeHerdr();
|
||||
FakeHerdr member = new FakeHerdr().withNoPanes();
|
||||
|
||||
assemble(dir, lead, member);
|
||||
|
||||
PaneLocator panes = runtime.mcp().identity().panes();
|
||||
PaneLocator.Lookup lookup = panes.terminalForPid(FakeHerdr.WORKER_PID);
|
||||
|
||||
assertEquals("term_a", lookup.terminal(),
|
||||
"the pane owning WORKER_PID lives on the LEAD daemon only (the member fake reports "
|
||||
+ "no panes) — PaneLocator must still find it, which is only possible if it "
|
||||
+ "searches the lead client and not just the member one");
|
||||
}
|
||||
|
||||
/**
|
||||
* The mirror control: when the pane instead lives ONLY on the member daemon (the lead reports
|
||||
* no panes), the lookup must still find it — proving the member client is genuinely searched
|
||||
* too, not merely tolerated as a second, always-losing argument.
|
||||
*/
|
||||
@Test
|
||||
void connectionIdentityAlsoSearchesTheMemberDaemon(@TempDir Path dir) throws Exception {
|
||||
FakeHerdr lead = new FakeHerdr().withNoPanes();
|
||||
FakeHerdr member = new FakeHerdr();
|
||||
|
||||
assemble(dir, lead, member);
|
||||
|
||||
PaneLocator panes = runtime.mcp().identity().panes();
|
||||
PaneLocator.Lookup lookup = panes.terminalForPid(FakeHerdr.WORKER_PID);
|
||||
|
||||
assertEquals("term_a", lookup.terminal(),
|
||||
"the pane owning WORKER_PID lives on the MEMBER daemon only — PaneLocator must "
|
||||
+ "find it there too");
|
||||
}
|
||||
|
||||
/** Sanity control: a pid nobody owns resolves to nothing on either daemon. */
|
||||
@Test
|
||||
void aPidNoPaneOwnsResolvesToNoTerminalOnEitherDaemon(@TempDir Path dir) throws Exception {
|
||||
FakeHerdr lead = new FakeHerdr();
|
||||
FakeHerdr member = new FakeHerdr();
|
||||
|
||||
assemble(dir, lead, member);
|
||||
|
||||
PaneLocator panes = runtime.mcp().identity().panes();
|
||||
PaneLocator.Lookup lookup = panes.terminalForPid(999_999L);
|
||||
|
||||
assertNull(lookup.terminal());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,211 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.msg.LeadChannel;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadMessage;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertSame;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 A-gaps (gap 1): {@code FleetdAssemblyLifecycleTest}'s own class javadoc says plainly
|
||||
* that it leaves {@code coordinator:} unset, so {@code leadMailbox} and {@code leadCoordLoop} stay
|
||||
* {@code null} throughout — the configured-coordinator path is never exercised by Unit A's own
|
||||
* test. This class drives that path instead: a real {@code coordinator:} block, a fake {@link
|
||||
* Fleetd.LeadMailboxOpener} returning a fake closeable channel (never a real broker connection),
|
||||
* and proof that {@link FleetdAssembly#assembleAndStart} both builds it and, on shutdown, closes it.
|
||||
*
|
||||
* <p>Made possible by generalising {@code Fleetd.LeadMailboxOpener}'s return type (and {@code
|
||||
* FleetdRuntime}'s field) from the concrete {@code LeadMailbox} to {@link LeadChannelHandle} — a
|
||||
* {@link LeadChannel} its owner can also close. {@code FleetMcp} and {@code LeadCoordLoop} already
|
||||
* consumed the narrower {@link LeadChannel}; this only widens the one seam that owns and closes it.
|
||||
*/
|
||||
class FleetdAssemblyCoordinatorLifecycleTest {
|
||||
|
||||
private static final String SELF_COORD_ID = "test-lead";
|
||||
|
||||
/** A fake {@link LeadChannelHandle}: never touches a broker, and records whether it was closed. */
|
||||
private static final class FakeLeadChannel implements LeadChannelHandle {
|
||||
volatile boolean closed = false;
|
||||
|
||||
@Override
|
||||
public void publish(String toCoordId, LeadMessage m) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<LeadMessage> peek() {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void ack(String msgId) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public String selfCoordId() {
|
||||
return SELF_COORD_ID;
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean heldDurable() {
|
||||
return true;
|
||||
}
|
||||
|
||||
@Override
|
||||
public MailboxState inspect(String coordId) {
|
||||
return MailboxState.unknown(coordId);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
closed = true;
|
||||
}
|
||||
}
|
||||
|
||||
/** Minimal fake {@link ResourcePorts}: a real herdr fake, a fake reply inbox, and a real, offered fake lead channel. */
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||
final FakeLeadChannel leadChannel = new FakeLeadChannel();
|
||||
String offeredUri;
|
||||
String offeredSelfId;
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> replyInbox;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
offeredUri = uri;
|
||||
offeredSelfId = selfCoordId;
|
||||
return leadChannel;
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// No real HTTP bind in a unit test.
|
||||
}
|
||||
}
|
||||
|
||||
private static final class SentinelReplyInbox implements ReplyInbox {
|
||||
@Override
|
||||
public void own(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void release(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean ack(String target, String msgId) {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
coordinator:
|
||||
uri: "amqp://fake-lead-broker/vh"
|
||||
selfId: "test-lead"
|
||||
""");
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@Test
|
||||
void configuredCoordinatorIsBuiltByTheAssemblyAndClosedOnShutdown(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
|
||||
// --- the assembly actually calls the configured opener and builds the coordinator path ---
|
||||
assertEquals("amqp://fake-lead-broker/vh", ports.offeredUri,
|
||||
"the assembly must open the mailbox at the configured broker uri");
|
||||
assertEquals(SELF_COORD_ID, ports.offeredSelfId,
|
||||
"the assembly must open the mailbox under the configured selfId");
|
||||
assertSame(ports.leadChannel, runtime.leadMailbox(),
|
||||
"FleetdRuntime must own the exact LeadChannelHandle the opener returned, not a copy");
|
||||
assertNotNull(runtime.leadCoordLoop(),
|
||||
"a configured coordinator: block must build the receiving LeadCoordLoop too");
|
||||
assertFalse(ports.leadChannel.closed, "the channel must still be open while the daemon is running");
|
||||
|
||||
// --- shutting the assembly down closes it -------------------------------------------------
|
||||
assertNotNull(ports.shutdownHook, "FleetdAssembly must have registered a shutdown hook");
|
||||
ports.shutdownHook.run();
|
||||
|
||||
assertTrue(ports.leadChannel.closed,
|
||||
"FleetdRuntime.close() must close the configured LeadChannelHandle");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,237 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.net.URI;
|
||||
import java.net.http.HttpClient;
|
||||
import java.net.http.HttpRequest;
|
||||
import java.net.http.HttpResponse;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.CopyOnWriteArrayList;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* fleetd #612 step 2, unit B2 (CB-185, {@code FleetApp} half). Replaces the deleted {@code
|
||||
* FleetdFleetAppConstructionTest}, which pinned this claim by reading {@code Fleetd.java}'s
|
||||
* source text for {@code "new FleetApp(herdr, memberHerdr, workers,"}. That claim moved to {@code
|
||||
* FleetdAssembly.java} (fleetd #612 Unit A) and is pinned here instead, by driving the real {@code
|
||||
* Javalin} app — via {@code runtime.app()}, not a copy — that {@link
|
||||
* FleetdAssembly#assembleAndStart} built and handed to {@link FleetdRuntime}.
|
||||
*
|
||||
* <p><strong>What this guards against</strong> (from the deleted test's own javadoc): constructing
|
||||
* {@code FleetApp} with the lead-only {@code herdr} client (dropping {@code memberHerdr}) makes
|
||||
* {@code GET /healthz} report green while the MEMBER daemon is down — so every spawn fails
|
||||
* invisibly — and silently drops every member workspace from {@code GET /sessions}. {@code
|
||||
* FleetAppTwoDaemonTest} already proves {@code FleetApp} itself merges/gates correctly given two
|
||||
* clients; the gap this pins is that the assembly actually passes it two.
|
||||
*
|
||||
* <p><strong>Both directions, not just one</strong> (fleetd #612 issue comment 17525): the deleted
|
||||
* guard's positive assertion required the exact pair {@code "new FleetApp(herdr, memberHerdr,
|
||||
* workers,"}, which does not survive EITHER daemon being dropped. An earlier version of this class
|
||||
* only proved the member-dropped direction, which left {@code new FleetApp(memberHerdr,
|
||||
* memberHerdr, ...)} — the symmetric bug, {@code /healthz} green while the LEAD daemon is down —
|
||||
* an undetected regression. {@link #healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp}
|
||||
* closes that.
|
||||
*
|
||||
* <p>Unlike the {@code ConnectionIdentity} half of CB-185 ({@code
|
||||
* FleetdAssemblyConnectionIdentityTest}), {@code /healthz} needs no caller identity at all, so
|
||||
* this test can bind {@link FleetdRuntime#app()} to a REAL ephemeral port (exactly {@code
|
||||
* FleetAppTwoDaemonTest} does for its own hand-built {@code FleetApp}) and drive it with a real
|
||||
* {@code HttpClient} — no accessor needed for this half.
|
||||
*
|
||||
* <p><strong>{@code GET /sessions} could not be driven the same way</strong>, so this class does
|
||||
* not pin the merge half of the deleted test's javadoc. {@code /sessions} requires
|
||||
* {@code Authz.Action.READ}, which — through the REAL assembly's real {@code
|
||||
* CallerResolver}/{@code ConnectionIdentity} (built with a hardcoded {@code
|
||||
* new LsofPeerPidLookup()}) — needs {@code Caller.resolved()}, i.e. a real positive pid from
|
||||
* {@code lsof}. {@code LsofPeerPidLookup} excludes its own pid (see its javadoc), and a JUnit
|
||||
* test's HTTP client and the daemon under test share one JVM pid, so the resolved pid is always
|
||||
* {@code -1} and every such request is refused as {@code ANONYMOUS} (fleetd #317's fail-closed
|
||||
* rule) before the route handler — and its {@code memberHerdr} merge — is ever reached. Verified
|
||||
* directly: driving {@code GET /sessions} here returns {@code 401 unauthenticated}, not the
|
||||
* merged body. {@code FleetAppTwoDaemonTest} avoids this because it builds {@code FleetApp} with
|
||||
* {@code callers: null}, which is not what the real assembly passes. The {@code /healthz} pin
|
||||
* below is what this class relies on for CB-185's {@code FleetApp} half; {@code
|
||||
* FleetAppTwoDaemonTest} remains the full behavioural proof that {@code FleetApp} itself merges
|
||||
* {@code /sessions} correctly once handed two clients.
|
||||
*/
|
||||
class FleetdAssemblyFleetAppTest {
|
||||
|
||||
private static final class TwoHerdrResourcePorts implements ResourcePorts {
|
||||
|
||||
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
|
||||
final CopyOnWriteArrayList<ScheduledExecutorService> schedulers = new CopyOnWriteArrayList<>();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
HerdrClient client = herdrsBySocket.get(socketPath);
|
||||
if (client == null) {
|
||||
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
|
||||
}
|
||||
return client;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> new dev.ltms.fleet.msg.InMemoryReplyInbox();
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||
schedulers.add(scheduler);
|
||||
return scheduler;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Deliberately never bind here — this test binds runtime.app() itself, for real, below.
|
||||
}
|
||||
}
|
||||
|
||||
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
|
||||
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
|
||||
|
||||
private final HttpClient http = HttpClient.newHttpClient();
|
||||
private FleetdRuntime runtime;
|
||||
private TwoHerdrResourcePorts ports;
|
||||
private Javalin boundApp;
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
if (boundApp != null) {
|
||||
boundApp.stop();
|
||||
}
|
||||
if (ports != null && ports.shutdownHook != null) {
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: "%s"
|
||||
memberHerdrSocket: "%s"
|
||||
lifecycle:
|
||||
idleTtlSeconds: 600
|
||||
health:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
""".formatted(LEAD_SOCKET, MEMBER_SOCKET));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
/** Assembles the real graph, then binds the real {@code Javalin app} to an ephemeral port. */
|
||||
private int assembleAndBind(Path dir, FakeHerdr lead, FakeHerdr member) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
ports = new TwoHerdrResourcePorts();
|
||||
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
|
||||
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
|
||||
runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
boundApp = runtime.app().start("127.0.0.1", 0);
|
||||
return boundApp.port();
|
||||
}
|
||||
|
||||
private HttpResponse<String> get(int port, String path) throws Exception {
|
||||
HttpRequest req = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path)).GET().build();
|
||||
return http.send(req, HttpResponse.BodyHandlers.ofString());
|
||||
}
|
||||
|
||||
/**
|
||||
* The pin. The MEMBER daemon is down; the LEAD daemon is healthy. If the assembly built
|
||||
* {@code FleetApp} with only the lead client (the bug: passing {@code herdr} where {@code
|
||||
* memberHerdr} is expected), the down member is invisible and {@code /healthz} stays 200.
|
||||
*/
|
||||
@Test
|
||||
void healthzGoesRedWhenTheMemberDaemonIsDownEvenThoughTheLeadIsUp(@TempDir Path dir) throws Exception {
|
||||
FakeHerdr lead = new FakeHerdr();
|
||||
FakeHerdr member = new FakeHerdr().healthy(false);
|
||||
|
||||
int port = assembleAndBind(dir, lead, member);
|
||||
|
||||
HttpResponse<String> res = get(port, "/healthz");
|
||||
assertEquals(503, res.statusCode(),
|
||||
"a down MEMBER daemon must not be masked by a healthy lead: " + res.body());
|
||||
}
|
||||
|
||||
/** Sanity control: both daemons healthy must still be green through the real assembly. */
|
||||
@Test
|
||||
void healthzIsGreenWhenBothDaemonsAreUp(@TempDir Path dir) throws Exception {
|
||||
FakeHerdr lead = new FakeHerdr();
|
||||
FakeHerdr member = new FakeHerdr();
|
||||
|
||||
int port = assembleAndBind(dir, lead, member);
|
||||
|
||||
assertEquals(200, get(port, "/healthz").statusCode());
|
||||
}
|
||||
|
||||
/**
|
||||
* The symmetric pin (fleetd #612 issue comment 17525): the LEAD daemon is down; the MEMBER
|
||||
* daemon is healthy. If the assembly built {@code FleetApp} with only the member client
|
||||
* (dropping {@code herdr} — the mirror of the bug above, {@code new FleetApp(memberHerdr,
|
||||
* memberHerdr, ...)}), the down LEAD is invisible and {@code /healthz} stays 200. Without this
|
||||
* case the pair above is one-directional and does not cover the deleted guard's positive
|
||||
* assertion (it required BOTH {@code herdr,} and {@code memberHerdr,} in that order).
|
||||
*/
|
||||
@Test
|
||||
void healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp(@TempDir Path dir) throws Exception {
|
||||
FakeHerdr lead = new FakeHerdr().healthy(false);
|
||||
FakeHerdr member = new FakeHerdr();
|
||||
|
||||
int port = assembleAndBind(dir, lead, member);
|
||||
|
||||
HttpResponse<String> res = get(port, "/healthz");
|
||||
assertEquals(503, res.statusCode(),
|
||||
"a down LEAD daemon must not be masked by a healthy member: " + res.body());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,258 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.inject.LoopWatchdog;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.CopyOnWriteArrayList;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 Unit A: proves {@link FleetdAssembly#assembleAndStart} — not a copy of its logic —
|
||||
* against a real {@link FleetConfig}, a {@link FakeHerdr} and a fake {@link ResourcePorts}, with no
|
||||
* real herdr socket, no real broker, and no real HTTP bind.
|
||||
*
|
||||
* <p><strong>The canonical order this test asserts against was recorded from {@code Fleetd.java}
|
||||
* BEFORE any code moved</strong> (fleetd #612 Unit A's mandated order of work), by reading the
|
||||
* original {@code main}'s body and its shutdown-hook {@code Thread}:
|
||||
*
|
||||
* <p>Start order: {@code SessionReaper.start()} → {@code StatusPoller.start()} →
|
||||
* {@code LeadHeartbeatLoop.start()} (opt-in) → {@code FleetHealthMonitor.start()} (opt-in) →
|
||||
* {@code LeadCoordLoop.start()} (opt-in) → {@code ConfigWatcher.start()} (opt-in) →
|
||||
* {@code app.start()} (HTTP), always last.
|
||||
*
|
||||
* <p>Close order (from the original shutdown hook body): {@code sessions.close(drainTimeoutSeconds)}
|
||||
* → {@code poller.stop()} → {@code messages.close()} → {@code pushLoop.close()} →
|
||||
* {@code heartbeat.close()} (if present) → {@code leadCoordLoop.close()} (if present) →
|
||||
* {@code leadCoordScheduler.shutdownNow()} (if present) → {@code healthMonitor.stop()} (if present)
|
||||
* → {@code configWatcher.stop()} (if present) → {@code mcp.close()} → {@code reaper.stop()} (if
|
||||
* present) → {@code idleSleepGuard.close()} (if present) → {@code replyInbox.close()} (if
|
||||
* {@code AutoCloseable}) → {@code leadMailbox.close()} (if present) → {@code router.close()}.
|
||||
*
|
||||
* <p>This test's config deliberately leaves {@code coordinator:} unset, so {@code leadMailbox} and
|
||||
* {@code leadCoordLoop} stay {@code null} throughout — the lead-mailbox resource-ledger criterion is
|
||||
* NOT exercised here; see the class-level caveat in the implementer's hand-off. {@code
|
||||
* idleSleepGuard.enabled: false} is set for the same kind of reason: it would otherwise try to spawn
|
||||
* a real {@code caffeinate} subprocess, which is not one of the resources the ticket's acceptance
|
||||
* criteria names (scheduler/inbox/mailbox/loop/MCP server/herdr router).
|
||||
*/
|
||||
class FleetdAssemblyLifecycleTest {
|
||||
|
||||
/** The one {@link HerdrClient} both {@code herdrSocket} and {@code memberHerdrSocket} resolve to. */
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||
final List<String> ledger = new CopyOnWriteArrayList<>();
|
||||
final List<ScheduledExecutorService> schedulers = new CopyOnWriteArrayList<>();
|
||||
final List<String> schedulerPurposes = new ArrayList<>();
|
||||
Runnable shutdownHook;
|
||||
Javalin startedApp;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
ledger.add("connectHerdr");
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
ledger.add("replyInboxOpener");
|
||||
return (uri, prefetch) -> replyInbox;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
// Never invoked: this test's config has no `coordinator:` block, so
|
||||
// Fleetd.openLeadMailbox returns null before calling the opener at all.
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
ledger.add("newScheduler:" + purpose);
|
||||
schedulerPurposes.add(purpose);
|
||||
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||
schedulers.add(scheduler);
|
||||
return scheduler;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
ledger.add("addShutdownHook");
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Deliberately never call app.start(host, port): no real HTTP bind in a unit test.
|
||||
ledger.add("startHttp");
|
||||
this.startedApp = app;
|
||||
}
|
||||
}
|
||||
|
||||
/** A fake {@link ReplyInbox} that is also {@link AutoCloseable}, so the ledger can prove it closes. */
|
||||
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
volatile boolean closed = false;
|
||||
|
||||
@Override
|
||||
public void own(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void release(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean ack(String target, String msgId) {
|
||||
return false;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
closed = true;
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
lifecycle:
|
||||
idleTtlSeconds: 600
|
||||
health:
|
||||
enabled: true
|
||||
intervalSeconds: 30
|
||||
leadHeartbeat:
|
||||
idleAfterSeconds: 600
|
||||
backoffMs: 15000
|
||||
quietNudgeCap: 5
|
||||
configReload:
|
||||
enabled: true
|
||||
intervalSeconds: 30
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
""");
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@Test
|
||||
void assemblesTheRealBootGraphWithoutTouchingAnyRealSocketBrokerOrPort(@TempDir Path dir)
|
||||
throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
|
||||
// --- start order: the herdr connect, then the recurring background loops, in the recorded
|
||||
// order, then the shutdown hook is registered, then (finally) HTTP "starts" -------------
|
||||
assertTrue(ports.ledger.indexOf("connectHerdr") < ports.ledger.indexOf("newScheduler:bridge-push-"),
|
||||
"herdr must connect before the push scheduler is created: " + ports.ledger);
|
||||
assertEquals(List.of("bridge-push-", "bridge-heartbeat-", "bridge-health-"), ports.schedulerPurposes,
|
||||
"the three always-created schedulers must be requested in exactly this order: "
|
||||
+ ports.schedulerPurposes);
|
||||
assertTrue(ports.ledger.indexOf("addShutdownHook") < ports.ledger.indexOf("startHttp"),
|
||||
"the shutdown hook must be registered before HTTP starts — the one statement that "
|
||||
+ "could not be reordered without changing FleetdRuntime's constructor shape, "
|
||||
+ "see FleetdAssembly's javadoc: " + ports.ledger);
|
||||
assertEquals(ports.ledger.size() - 1, ports.ledger.indexOf("startHttp"),
|
||||
"HTTP must start LAST of everything this fake observes: " + ports.ledger);
|
||||
assertNotNull(ports.startedApp, "FleetdAssembly must have built and handed off a real FleetApp");
|
||||
|
||||
// No real HTTP bind and no real herdr socket: this call returning at all, plus the ledger
|
||||
// above, is the proof — a real bind or a real UnixSocketHerdrClient.connect would have
|
||||
// thrown or hung against the sockets/ports this test never opened.
|
||||
assertNotNull(runtime.app(), "FleetdRuntime must own the same Javalin app that was built");
|
||||
assertEquals(ports.startedApp, runtime.app(), "attachApp must hand FleetdRuntime the SAME instance");
|
||||
|
||||
// --- the runtime owns the real, live objects — not a copy --------------------------------
|
||||
assertEquals(LoopWatchdog.State.RUNNING, runtime.reaper().health(),
|
||||
"SessionReaper must be running: lifecycle.idleTtlSeconds is configured");
|
||||
assertEquals(LoopWatchdog.State.RUNNING, runtime.poller().health(), "StatusPoller must be running");
|
||||
assertNotNull(runtime.heartbeat(), "leadHeartbeat: is configured, so the loop must be built and started");
|
||||
assertNotNull(runtime.healthMonitor(), "health.enabled: true, so the monitor must be built and started");
|
||||
assertNotNull(runtime.configWatcher(), "configReload.enabled: true, so the watcher must be built and started");
|
||||
// Documented gap (see class javadoc): no coordinator: block, so these stay null.
|
||||
assertEquals(null, runtime.leadCoordLoop(), "no coordinator: block — leadCoordLoop must stay unbuilt");
|
||||
assertEquals(null, runtime.leadMailbox(), "no coordinator: block — leadMailbox must stay unbuilt");
|
||||
assertEquals(runtime.replyInbox(), ports.replyInbox,
|
||||
"FleetdRuntime must own the exact ReplyInbox instance this fake's AmqpOpener returned");
|
||||
assertFalse(ports.herdr.closed, "herdr must still be open while the daemon is running");
|
||||
assertFalse(ports.replyInbox.closed, "the reply inbox must still be open while the daemon is running");
|
||||
for (ScheduledExecutorService scheduler : ports.schedulers) {
|
||||
assertFalse(scheduler.isShutdown(), "a scheduler must still be running while the daemon is up");
|
||||
}
|
||||
|
||||
// --- close order: invoke the captured shutdown-hook Runnable directly (no real JVM shutdown
|
||||
// happens in a unit test) and prove every resource this fake can observe is released --------
|
||||
assertNotNull(ports.shutdownHook, "FleetdAssembly must have registered a shutdown hook");
|
||||
ports.shutdownHook.run();
|
||||
|
||||
assertEquals(LoopWatchdog.State.STOPPED, runtime.reaper().health(), "SessionReaper must stop on close");
|
||||
assertEquals(LoopWatchdog.State.STOPPED, runtime.poller().health(), "StatusPoller must stop on close");
|
||||
assertTrue(ports.herdr.closed, "router.close() must close the herdr client last");
|
||||
assertTrue(ports.replyInbox.closed, "the AutoCloseable reply inbox must be closed");
|
||||
for (int i = 0; i < ports.schedulers.size(); i++) {
|
||||
assertTrue(ports.schedulers.get(i).isShutdown(),
|
||||
"scheduler for '" + ports.schedulerPurposes.get(i) + "' must be shut down by close(): "
|
||||
+ "LeadHeartbeatLoop/FleetHealthMonitor/ReplyPushLoop each call "
|
||||
+ "scheduler.shutdownNow() on the exact instance ports.newScheduler(...) handed them");
|
||||
}
|
||||
|
||||
// Calling the captured hook a second time must never happen for a real JVM shutdown hook,
|
||||
// but nothing above should have thrown — that already proves every accessed field's close()
|
||||
// tolerated running once, in the recorded order, without an exception escaping.
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,165 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import dev.ltms.fleet.testing.CapturedLog;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 A-gaps (gap 2): {@code Fleetd.java:185} used to call {@code
|
||||
* reportRoleFallbackGaps(cfg)} <em>outside</em> the boundary {@code FleetdAssembly.assembleAndStart}
|
||||
* — the assembly call itself sat at line 201, after it — so nothing that drives the assembly (the
|
||||
* seam {@code FleetdAssemblyLifecycleTest} exercises) could ever notice the call being deleted.
|
||||
* {@code Fleetd.main} still refuses a bad config end to end (see {@code
|
||||
* FleetdStartupValidationTest}), but that only pins {@code validateAll()} and {@code
|
||||
* assertChartersNameOnlyRegisteredTools} — a validator that <em>throws</em>. {@code
|
||||
* reportRoleFallbackGaps} only logs; nothing about {@code main} throwing or not throwing can
|
||||
* observe whether that particular call ran.
|
||||
*
|
||||
* <p>This drives {@link FleetdAssembly#assembleAndStart} directly — never a copy of its logic —
|
||||
* with a config that has no {@code fleet:} pools or charters configured for any role, so every role
|
||||
* trips both of {@code reportRoleFallbackGaps}' log branches, and asserts on the real log line a
|
||||
* {@code ListAppender} attached to the shared {@code Fleetd}/{@code FleetdAssembly} logger
|
||||
* captures. Deleting the call from {@code FleetdAssembly} (verified by hand, see the ticket) turns
|
||||
* this test red; deleting it from {@code Fleetd.main} instead (its old location) would not, which
|
||||
* is exactly the gap this test closes.
|
||||
*/
|
||||
class FleetdAssemblyRoleFallbackBoundaryTest {
|
||||
|
||||
/** Minimal fake {@link ResourcePorts}: enough for {@code assembleAndStart} to run with no real I/O. */
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> replyInbox;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// No real HTTP bind in a unit test.
|
||||
}
|
||||
}
|
||||
|
||||
private static final class SentinelReplyInbox implements ReplyInbox {
|
||||
@Override
|
||||
public void own(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void release(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean ack(String target, String msgId) {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
// Deliberately no `fleet:` block at all: every MemberRole has neither a pool nor a
|
||||
// charter, so reportRoleFallbackGaps' "role fallback: no fleet.<role>s: pool for ..."
|
||||
// branch is guaranteed to log something to assert on.
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
""");
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@Test
|
||||
void assembleAndStartItselfReportsTheRoleFallbackGaps(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
|
||||
try (var captured = CapturedLog.at(Fleetd.class, Level.INFO)) {
|
||||
FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
|
||||
String infoLines = captured.events().stream()
|
||||
.filter(e -> e.getLevel() == Level.INFO)
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.reduce("", (a, b) -> a + "\n" + b);
|
||||
assertTrue(infoLines.contains("role fallback: no fleet.<role>s: pool for"),
|
||||
() -> "FleetdAssembly.assembleAndStart itself must call reportRoleFallbackGaps "
|
||||
+ "(fleetd #612 A-gaps gap 2) — captured INFO lines: " + infoLines);
|
||||
} finally {
|
||||
if (ports.shutdownHook != null) {
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,14 +1,12 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.HerdrException;
|
||||
import dev.ltms.fleet.testing.CapturedLog;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.concurrent.atomic.AtomicBoolean;
|
||||
import java.util.function.LongSupplier;
|
||||
@@ -92,49 +90,42 @@ class FleetdAwaitHerdrTest {
|
||||
|
||||
@Test
|
||||
void answeredLogsNothingAndSaysReap() {
|
||||
ListAppender<ILoggingEvent> events = attach();
|
||||
try {
|
||||
try (CapturedLog log = attach()) {
|
||||
boolean shouldReap = Fleetd.logHerdrWaitOutcomeAndShouldReap(
|
||||
new Fleetd.HerdrAwaitOutcome(Fleetd.HerdrWaitResult.ANSWERED, 0L));
|
||||
|
||||
assertTrue(shouldReap, "only ANSWERED should tell main to reap orphan workers");
|
||||
assertEquals(0, events.list.size(), "the answered path logs nothing itself");
|
||||
} finally {
|
||||
detach(events);
|
||||
assertEquals(0, log.events().size(), "the answered path logs nothing itself");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void deadlinePassedLogsConfiguredAndMeasuredElapsedTogether() {
|
||||
ListAppender<ILoggingEvent> events = attach();
|
||||
try {
|
||||
try (CapturedLog log = attach()) {
|
||||
boolean shouldReap = Fleetd.logHerdrWaitOutcomeAndShouldReap(
|
||||
new Fleetd.HerdrAwaitOutcome(Fleetd.HerdrWaitResult.DEADLINE_PASSED, 30_500_000_000L));
|
||||
|
||||
assertFalse(shouldReap, "a deadline-passed wait must not tell main to reap");
|
||||
assertEquals(1, events.list.size());
|
||||
ILoggingEvent event = events.list.getFirst();
|
||||
assertEquals(1, log.events().size());
|
||||
ILoggingEvent event = log.events().getFirst();
|
||||
assertEquals(Level.WARN, event.getLevel());
|
||||
assertEquals("herdr did not answer within the configured wait (configured=30s "
|
||||
+ "elapsed=30500ms) — starting anyway; /healthz will report degraded until it "
|
||||
+ "comes up. Orphaned worker panes (if any) were NOT reaped.",
|
||||
event.getFormattedMessage());
|
||||
} finally {
|
||||
detach(events);
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void interruptedLogsItsOwnMessageAndNeverClaimsTheBudgetElapsed() {
|
||||
ListAppender<ILoggingEvent> events = attach();
|
||||
try {
|
||||
try (CapturedLog log = attach()) {
|
||||
// 3ms: the ticket's own example of "a few milliseconds in", not the 30s budget.
|
||||
boolean shouldReap = Fleetd.logHerdrWaitOutcomeAndShouldReap(
|
||||
new Fleetd.HerdrAwaitOutcome(Fleetd.HerdrWaitResult.INTERRUPTED, 3_000_000L));
|
||||
|
||||
assertFalse(shouldReap, "an interrupted wait must not tell main to reap");
|
||||
assertEquals(1, events.list.size());
|
||||
ILoggingEvent event = events.list.getFirst();
|
||||
assertEquals(1, log.events().size());
|
||||
ILoggingEvent event = log.events().getFirst();
|
||||
assertEquals(Level.WARN, event.getLevel());
|
||||
String message = event.getFormattedMessage();
|
||||
assertEquals("herdr wait was interrupted before the configured wait ran out "
|
||||
@@ -143,8 +134,6 @@ class FleetdAwaitHerdrTest {
|
||||
message);
|
||||
assertFalse(message.contains("did not answer"),
|
||||
"an interrupted wait must not be reported as if herdr failed to answer within the budget");
|
||||
} finally {
|
||||
detach(events);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -197,16 +186,7 @@ class FleetdAwaitHerdrTest {
|
||||
}
|
||||
}
|
||||
|
||||
private static ListAppender<ILoggingEvent> attach() {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
logger.setLevel(Level.DEBUG);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
return appender;
|
||||
}
|
||||
|
||||
private static void detach(ListAppender<ILoggingEvent> appender) {
|
||||
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
|
||||
private static CapturedLog attach() {
|
||||
return CapturedLog.at(Fleetd.class, Level.DEBUG);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,199 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.OptionalLong;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 B3 — replaces {@code FleetdBackendQuarantineWiringTest} (fleetd #466), a source-text
|
||||
* test that scraped {@code Fleetd.java} (now {@code FleetdAssembly.java}, moved there by fleetd #612
|
||||
* Unit A) for the {@code BackendQuarantine.withEscalation(...)} call, and separately asserted the
|
||||
* flat two-argument constructor's text was ABSENT. That proves the right method NAME appears in
|
||||
* source; it proves nothing about what the constructed object actually DOES.
|
||||
*
|
||||
* <p>This test instead drives the REAL {@link BackendQuarantine} the real {@link
|
||||
* FleetdAssembly#assembleAndStart} builds — reached through {@link
|
||||
* dev.ltms.fleet.mcp.FleetMcp#quarantineSource()} on the real, live {@code FleetMcp} {@code
|
||||
* FleetdRuntime} owns — and asserts the ONE behavioural difference {@code withEscalation} and the
|
||||
* flat constructor actually produce (see {@link BackendQuarantine}'s own class doc, "Mechanism"):
|
||||
* quarantining the same credential twice in a row, within one base cooldown of the first deadline,
|
||||
* must escalate the second cooldown past the first. A flat instance reports the identical cooldown
|
||||
* both times.
|
||||
*/
|
||||
class FleetdBackendQuarantineAssemblyTest {
|
||||
|
||||
/** Base cooldown used throughout — long enough that rounding never blurs the 2x escalation. */
|
||||
private static final int COOLDOWN_SECONDS = 100;
|
||||
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||
final AtomicLong clockNanos = new AtomicLong(0L);
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> replyInbox;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
// Controllable: the SAME LongSupplier instance BackendQuarantine.withEscalation(...) is
|
||||
// built with, so advancing clockNanos after assembly moves the quarantine tracker's own
|
||||
// clock, with no real sleep needed to observe escalation.
|
||||
return clockNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return clockNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
}
|
||||
}
|
||||
|
||||
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
@Override
|
||||
public void own(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void release(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean ack(String target, String msgId) {
|
||||
return false;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
quarantineCooldownSeconds: %d
|
||||
""".formatted(COOLDOWN_SECONDS));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled BackendQuarantine escalates a repeated exhaustion, "
|
||||
+ "which the flat two-argument constructor can never do")
|
||||
void assembledQuarantineEscalatesOnARepeatedExhaustion(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
|
||||
BackendQuarantine quarantine = runtime.mcp().quarantineSource().quarantine();
|
||||
|
||||
// First exhaustion, at clock=0: a fresh occurrence, blocked for exactly the base cooldown.
|
||||
quarantine.quarantine("cred-x");
|
||||
BackendQuarantine.Status first = quarantine.status("cred-x").orElseThrow(
|
||||
() -> new AssertionError("credential must be quarantined immediately after quarantine()"));
|
||||
assertEquals(1, first.repeatCount(), "the first call is repeat #1");
|
||||
assertEquals(COOLDOWN_SECONDS, first.remainingSeconds(),
|
||||
"a fresh quarantine blocks for exactly the base cooldown");
|
||||
|
||||
// Second exhaustion, arriving just after the first deadline — well within one base cooldown
|
||||
// of it, so this is a CONTINUATION of the same streak (repeat #2), not a fresh occurrence.
|
||||
long firstDeadlineNanos = COOLDOWN_SECONDS * 1_000_000_000L;
|
||||
ports.clockNanos.set(firstDeadlineNanos + 1);
|
||||
quarantine.quarantine("cred-x");
|
||||
BackendQuarantine.Status second = quarantine.status("cred-x").orElseThrow(
|
||||
() -> new AssertionError("credential must be quarantined immediately after the second "
|
||||
+ "quarantine() call"));
|
||||
assertEquals(2, second.repeatCount(), "the second call, arriving within one base cooldown of "
|
||||
+ "the first deadline, continues the streak as repeat #2");
|
||||
|
||||
// The one behavioural difference: withEscalation doubles the cooldown on repeat #2 (capped
|
||||
// well above this at 12x base), the flat two-argument constructor never grows past the base
|
||||
// cooldown no matter how many times quarantine() is called in a row.
|
||||
assertEquals(2 * COOLDOWN_SECONDS, second.remainingSeconds(),
|
||||
"withEscalation's default backoff doubles the cooldown on the second consecutive "
|
||||
+ "exhaustion — this is the exact call FleetdAssembly.java makes at the "
|
||||
+ "BackendQuarantine.withEscalation(...) call site");
|
||||
assertTrue(second.remainingSeconds() > first.remainingSeconds(),
|
||||
"the flat two-argument BackendQuarantine constructor would report the SAME remaining "
|
||||
+ "seconds both times — this inequality is what a mutation to the flat "
|
||||
+ "constructor at that call site must fail");
|
||||
|
||||
// Also confirm isQuarantined/remainingSeconds agree, exercising the accessors a real caller
|
||||
// (fleet_profiles / fleet_list, per BackendQuarantine's own class doc) actually reads.
|
||||
assertTrue(quarantine.isQuarantined("cred-x"));
|
||||
OptionalLong remaining = quarantine.remainingSeconds("cred-x");
|
||||
assertTrue(remaining.isPresent());
|
||||
assertEquals(2 * COOLDOWN_SECONDS, remaining.getAsLong());
|
||||
}
|
||||
}
|
||||
@@ -1,65 +0,0 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #466 follow-up: {@code Fleetd.main} builds the daemon's one {@code BackendQuarantine}
|
||||
* from {@link dev.ltms.fleet.placement.BackendQuarantine#withEscalation(java.util.function.LongSupplier,
|
||||
* long)} — the escalating factory — rather than the plain two-argument constructor, which is still a
|
||||
* flat cooldown (kept for backward compatibility, see that class's doc). {@code
|
||||
* BackendQuarantineTest} proves {@code withEscalation} itself escalates, is ceilinged, and resets;
|
||||
* it says nothing about which one {@code main} actually calls.
|
||||
*
|
||||
* <p>Measured directly: reverting {@code main} to {@code new BackendQuarantine(System::nanoTime,
|
||||
* TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()))} — the pre-#466 flat call — compiles
|
||||
* with 0 errors and leaves the entire 1608-test suite (including every {@code BackendQuarantineTest}
|
||||
* case) green, because no other test constructs its {@code BackendQuarantine} through {@code main};
|
||||
* every one of them builds its own instance directly. That silent regression is exactly the shape
|
||||
* {@link FleetdLeadSeatWiringTest} and {@link FleetdCompletionResolverWiringTest} already guard
|
||||
* against for their own constructor arguments — this is the same class of gap for fleetd #466's
|
||||
* factory choice, following their approach.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
||||
* BackendQuarantine} and never runs {@code main} — a green result here proves only that the exact
|
||||
* text {@code main} calls {@code BackendQuarantine.withEscalation(...)} rather than the flat
|
||||
* constructor. It does not prove that call actually executes at startup (no test here starts the
|
||||
* daemon), and it does not prove the escalation reaches a real backend or credential — only
|
||||
* {@code BackendQuarantineTest} proves the factory's own behaviour, and only a live daemon proves
|
||||
* the wiring runs.
|
||||
*/
|
||||
class FleetdBackendQuarantineWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] main's BackendQuarantine local is still built from BackendQuarantine.withEscalation(...)")
|
||||
void mainStillWiresTheEscalatingQuarantineFactory() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,\n"
|
||||
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
|
||||
"Fleetd.main's BackendQuarantine local must still be built from "
|
||||
+ "BackendQuarantine.withEscalation(System::nanoTime, "
|
||||
+ "TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds())). Reverting to the flat "
|
||||
+ "two-argument constructor (fleetd #466's measured regression) compiles with 0 errors "
|
||||
+ "and leaves the whole suite green, including every BackendQuarantineTest case that "
|
||||
+ "proves the escalation itself works — this source check is what must go red instead. "
|
||||
+ "A reverted daemon would go back to retrying a weekly subscription limit on every "
|
||||
+ "flat ~30-minute cooldown, about 336 times across the week.");
|
||||
|
||||
// Negative form of the same check: the pre-#466 flat call, if it ever reappears at this
|
||||
// declaration, must not be mistaken for the escalating one by a looser positive-only check.
|
||||
assertFalse(source.contains(
|
||||
"BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,\n"
|
||||
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
|
||||
"main's BackendQuarantine local must never regress to the flat two-argument constructor");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,84 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
|
||||
/**
|
||||
* fleetd #589 Group 2: {@link Fleetd#claudeCodeLauncher} is the factory that replaced {@code
|
||||
* main}'s inline {@code new ClaudeCodeLauncher(...)} call, whose 10th argument is the CB-596 {@code
|
||||
* memberCredentials} policy supplier ({@code () -> config.get().memberCredentials()}). Before this
|
||||
* ticket that argument was untestable wiring: replacing it with {@code () -> null} compiled with 0
|
||||
* errors and left every existing test green, since no existing test builds the exact object {@code
|
||||
* main} wires and then spawns it. {@code memberCredentials} being a lambda is never itself {@code
|
||||
* null}, so {@link dev.ltms.fleet.member.HerdrPeerLauncher#applyMemberCredentialPolicy} sees {@code
|
||||
* memberCredentials.get() == null} and silently shadows nothing — reopening the exact CB-592
|
||||
* exposure gap CB-596's policy closed (gitea issue #82).
|
||||
*
|
||||
* <p>This test drives the factory with a real {@link ConfigRef} carrying a {@code
|
||||
* memberCredentials:} block, spawns through the resulting launcher, and inspects what {@code
|
||||
* tab.create} actually carried — the same observable surface {@code ClaudeCodeLauncherTest}'s
|
||||
* {@code everyKnownNameNotAllowedIsShadowedWithTheSentinel} uses for the launcher's own credential
|
||||
* policy, applied here to prove {@code main}'s wiring reaches it.
|
||||
*/
|
||||
class FleetdClaudeCodeLauncherCredentialWiringTest {
|
||||
|
||||
private static final String YAML = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
ltms-local:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: coder
|
||||
memberCredentials:
|
||||
policy: deny-by-default
|
||||
known:
|
||||
- GITEA_ACCESS_TOKEN
|
||||
""";
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static Map<String, String> startEnv(FakeHerdr herdr) {
|
||||
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("main's memberCredentials wiring reaches ClaudeCodeLauncher: a known-but-not-allowed "
|
||||
+ "name is shadowed on spawn")
|
||||
void memberCredentialsWiringReachesClaudeCodeLauncher(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, YAML);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = ConfigRef.fixed(cfg);
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
ClaudeCodeLauncher launcher = Fleetd.claudeCodeLauncher(new AgentControl(herdr),
|
||||
new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
|
||||
cfg.profiles(), cfg, config);
|
||||
launcher.spawn();
|
||||
|
||||
String shadowed = startEnv(herdr).get("GITEA_ACCESS_TOKEN");
|
||||
assertNotNull(shadowed,
|
||||
"GITEA_ACCESS_TOKEN is 'known' but not 'allow'-ed in the loaded config — it must be "
|
||||
+ "explicitly shadowed on spawn; replacing the memberCredentials supplier with "
|
||||
+ "() -> null at the Fleetd.claudeCodeLauncher call site must fail this "
|
||||
+ "assertion, since a null policy shadows nothing");
|
||||
assertFalse(shadowed.isBlank(), "the overlay value must be non-blank");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,307 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.msg.TurnToken;
|
||||
import dev.ltms.fleet.placement.PlacementException;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
import dev.ltms.fleet.session.WorktreeRequest;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 step 2, Unit B1 — replaces {@code FleetdCompletionResolverWiringTest} (deleted in
|
||||
* this same commit), whose four tests read {@code Fleetd.java}'s source text and asserted the
|
||||
* {@code CompletionResolver} construction call still named the right arguments. That proved the
|
||||
* call site's spelling, never that the assembled resolver actually behaves differently when an
|
||||
* argument is dropped.
|
||||
*
|
||||
* <p>These tests drive {@link FleetdAssembly#assembleAndStart} — the real boot composition,
|
||||
* fleetd #612 Unit A — and read {@link FleetdRuntime#completion()}: the exact {@link
|
||||
* CompletionResolver} instance the assembled daemon uses, never a copy built alongside it for the
|
||||
* test's benefit. Two behaviours are pinned, matching the ticket's own two measured mutations:
|
||||
*
|
||||
* <ul>
|
||||
* <li>the 8th constructor argument ({@code Fleetd.worktreeBranchLookup(sessions::roster)}) —
|
||||
* {@link #assembledResolverReportsTheMembersWorktreeAndBranchInAFallbackReport}; and</li>
|
||||
* <li>the 5th/6th arguments ({@code backendErrorPatterns}, {@code backendErrorSink}, both
|
||||
* assigned from {@code Fleetd}'s extracted factories rather than an inline lambda) —
|
||||
* {@link #assembledResolverClassifiesAndCoolsOffOnAConfiguredBackendErrorPattern}.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>Both tests bypass {@link dev.ltms.fleet.inject.StatusPoller} and drive {@link
|
||||
* CompletionResolver#onDelivered} / {@link CompletionResolver#resolveBeforePostAction} directly —
|
||||
* the same public, synchronous entry points {@code CompletionResolverTest} uses — with a
|
||||
* hand-built {@link CompletableFuture} waiter, so no real poller loop or herdr status poll is
|
||||
* needed. The pane scrape comes from {@link FakeHerdr#readText}; the elapsed-time floor
|
||||
* ({@code CompletionResolver.MIN_TURN_NANOS}) is controlled via a fake, advanceable {@link
|
||||
* ResourcePorts#nanoClock()} rather than a real sleep.
|
||||
*/
|
||||
class FleetdCompletionResolverAssemblyTest {
|
||||
|
||||
/** Same shape as {@code FleetdAssemblyLifecycleTest}'s fake, plus a nanoClock this test can advance. */
|
||||
private static final class ControllableResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr;
|
||||
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L); // arbitrary non-zero start
|
||||
Runnable shutdownHook;
|
||||
|
||||
ControllableResourcePorts(FakeHerdr herdr) {
|
||||
this.herdr = herdr;
|
||||
}
|
||||
|
||||
void advanceSeconds(long seconds) {
|
||||
nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(seconds));
|
||||
}
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
// Never invoked: this test's config has no `broker:` block, so Fleetd.selectReplyInbox
|
||||
// returns the in-memory inbox before calling the opener at all.
|
||||
return (uri, prefetch) -> {
|
||||
throw new UnsupportedOperationException("replyInboxOpener must not be called — no broker: block");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
// Never invoked: no `coordinator:` block configured either.
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException("leadMailboxOpener must not be called — no coordinator: block");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return nowNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return nowNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Deliberately never bind a real port.
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir, String profilesYaml, String extraGuardHost,
|
||||
String worktreeRootYamlLine) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
%s
|
||||
profiles:
|
||||
%s
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- %s
|
||||
""".formatted(worktreeRootYamlLine == null ? "" : worktreeRootYamlLine, profilesYaml, extraGuardHost));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
private static void gitQuiet(Path cwd, String... args) throws Exception {
|
||||
List<String> cmd = new java.util.ArrayList<>(List.of("git"));
|
||||
cmd.addAll(List.of(args));
|
||||
Process p = new ProcessBuilder(cmd).directory(cwd.toFile()).redirectErrorStream(true).start();
|
||||
String out = new String(p.getInputStream().readAllBytes());
|
||||
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git timed out: git " + String.join(" ", args));
|
||||
assertEquals(0, p.exitValue(), "git " + String.join(" ", args) + " failed:\n" + out);
|
||||
}
|
||||
|
||||
private static Path initRepo(Path dir) throws Exception {
|
||||
Files.createDirectories(dir);
|
||||
gitQuiet(dir, "init", "-q", "-b", "main");
|
||||
gitQuiet(dir, "config", "user.email", "test@example.invalid");
|
||||
gitQuiet(dir, "config", "user.name", "Test");
|
||||
Files.writeString(dir.resolve("README.md"), "seed\n");
|
||||
gitQuiet(dir, "add", "README.md");
|
||||
gitQuiet(dir, "commit", "-q", "-m", "seed");
|
||||
return dir;
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #248's first measured mutation: replacing {@code CompletionResolver}'s 8th constructor
|
||||
* argument with the inert {@code _ -> null} compiles clean and leaves every existing test green
|
||||
* — it silently drops fleetd #241's fallback-report location. This drives the real assembled
|
||||
* resolver through a member echoing its own injected brief back (no {@code fleet_reply}), which
|
||||
* resolves via {@code noReportMessage(target)}, and proves the real member's {@code branch} —
|
||||
* only obtainable via {@code Fleetd.worktreeBranchLookup(sessions::roster)} reading the real,
|
||||
* worktree-provisioned {@link MemberSession} — appears in the reported text.
|
||||
*/
|
||||
@Test
|
||||
void assembledResolverReportsTheMembersWorktreeAndBranchInAFallbackReport(@TempDir Path dir) throws Exception {
|
||||
Path repo = initRepo(dir.resolve("repo"));
|
||||
FleetConfig cfg = writeConfig(dir, """
|
||||
wtprofile:
|
||||
baseUrl: http://wthost.local:8000
|
||||
model: sonnet
|
||||
""", "wthost.local", "worktreeRoot: " + dir.resolve("wts"));
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
ControllableResourcePorts ports = new ControllableResourcePorts(new FakeHerdr());
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
try {
|
||||
MemberSession session = runtime.sessions().acquire("wtprofile", repo.toString(), repo.toString(),
|
||||
null, new WorktreeRequest("fleetd-612-b1", null));
|
||||
String target = session.terminalId();
|
||||
String branch = session.branch();
|
||||
assertTrue(branch != null && branch.startsWith("worker/"),
|
||||
"sanity: a worktree-provisioned session must carry a real branch, got: " + branch);
|
||||
|
||||
CompletionResolver completion = runtime.completion();
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = new CompletableFuture<>();
|
||||
String echoedBrief = "z".repeat(450); // >= CompletionResolver.ECHO_MIN_CHARS normalised chars
|
||||
|
||||
ports.herdr.readText("idle, nothing yet");
|
||||
completion.onDelivered(target, new TurnToken(target, waiter, echoedBrief));
|
||||
|
||||
ports.herdr.readText(echoedBrief); // the pane just echoes the injected brief back — no real report
|
||||
ports.advanceSeconds(3); // clear CompletionResolver.MIN_TURN_NANOS (2s) without a real sleep
|
||||
completion.resolveBeforePostAction(target);
|
||||
|
||||
Rendezvous.Resolution resolution = waiter.getNow(null);
|
||||
assertTrue(resolution != null, "the waiter must have resolved synchronously");
|
||||
assertEquals(Rendezvous.Kind.COMPLETION, resolution.kind());
|
||||
assertTrue(resolution.text().contains(CompletionResolver.NO_REPORT_PREFIX),
|
||||
"sanity: must have gone down the noReportMessage sub-path: " + resolution.text());
|
||||
assertTrue(resolution.text().contains("branch=" + branch),
|
||||
"the assembled resolver must report the member's real branch (fleetd #241 via "
|
||||
+ "fleetd #248's worktreeBranchLookup wiring); got: " + resolution.text());
|
||||
} finally {
|
||||
if (ports.shutdownHook != null) ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #248's second measured mutation, and fleetd#201 Unit 5's own gap: replacing {@code
|
||||
* backendErrorPatterns}/{@code backendErrorSink} with {@code BackendErrorPatternLookup.legacy()}
|
||||
* / {@code BackendErrorSink.none()} compiles clean and leaves every existing behavioural test
|
||||
* green.
|
||||
*
|
||||
* <p>Classification proof: this test's profile configures {@code errorPattern: "credential
|
||||
* outage"} — text the built-in {@code (?i)\bAPI Error\s*:} fallback ({@code legacy()}'s only
|
||||
* behaviour) never matches. So a real {@code Fleetd.backendErrorPatternLookup(...)} wiring
|
||||
* classifies the send as {@code FAILED}; {@code legacy()} would fall through to the plain
|
||||
* completion path instead ({@code Kind.COMPLETION}).
|
||||
*
|
||||
* <p>Cool-off proof: two distinct targets on the same profile/credential each classified as a
|
||||
* backend error inside the 60s window must cool the credential off ({@link
|
||||
* dev.ltms.fleet.placement.BackendOutagePolicy}, fleetd#201 Unit 5) — observable two ways: (1)
|
||||
* the real {@code Fleetd.backendErrorSink(...)} marks each session {@code BACKEND_ERROR} (only
|
||||
* the real sink calls {@code sessions.onBackendError}; {@code BackendErrorSink.none()} never
|
||||
* does), and (2) a third explicit-profile spawn attempt is refused with a {@link
|
||||
* PlacementException} naming the cool-off — only reachable because the real sink's {@code
|
||||
* outagePolicy.record(...)} call actually ran.
|
||||
*/
|
||||
@Test
|
||||
void assembledResolverClassifiesAndCoolsOffOnAConfiguredBackendErrorPattern(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir, """
|
||||
coolprofile:
|
||||
baseUrl: http://coolhost.local:8000
|
||||
model: sonnet
|
||||
errorPattern: "credential outage"
|
||||
""", "coolhost.local", null);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
ControllableResourcePorts ports = new ControllableResourcePorts(new FakeHerdr());
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
try {
|
||||
MemberSession session1 = runtime.sessions().acquire("coolprofile", null, dir.toString(), null);
|
||||
MemberSession session2 = runtime.sessions().acquire("coolprofile", null, dir.toString(), null);
|
||||
String target1 = session1.terminalId();
|
||||
String target2 = session2.terminalId();
|
||||
assertTrue(!target1.equals(target2), "sanity: the two spawns must be distinct targets");
|
||||
|
||||
CompletionResolver completion = runtime.completion();
|
||||
|
||||
CompletableFuture<Rendezvous.Resolution> waiter1 = new CompletableFuture<>();
|
||||
ports.herdr.readText("idle 1");
|
||||
completion.onDelivered(target1, new TurnToken(target1, waiter1, null));
|
||||
ports.herdr.readText("credential outage: upstream 503");
|
||||
ports.advanceSeconds(3);
|
||||
completion.resolveBeforePostAction(target1);
|
||||
Rendezvous.Resolution resolution1 = waiter1.getNow(null);
|
||||
assertTrue(resolution1 != null, "target1's waiter must have resolved synchronously");
|
||||
assertEquals(Rendezvous.Kind.FAILED, resolution1.kind(),
|
||||
"a configured errorPattern the built-in fallback never matches must classify as "
|
||||
+ "a backend error, not a plain completion; got: " + resolution1);
|
||||
assertTrue(resolution1.text().contains("credential outage: upstream 503"), resolution1.text());
|
||||
|
||||
CompletableFuture<Rendezvous.Resolution> waiter2 = new CompletableFuture<>();
|
||||
ports.herdr.readText("idle 2");
|
||||
completion.onDelivered(target2, new TurnToken(target2, waiter2, null));
|
||||
ports.herdr.readText("credential outage: upstream 503 again");
|
||||
ports.advanceSeconds(3);
|
||||
completion.resolveBeforePostAction(target2);
|
||||
Rendezvous.Resolution resolution2 = waiter2.getNow(null);
|
||||
assertTrue(resolution2 != null, "target2's waiter must have resolved synchronously");
|
||||
assertEquals(Rendezvous.Kind.FAILED, resolution2.kind());
|
||||
|
||||
List<MemberSession> roster = runtime.sessions().roster();
|
||||
assertTrue(roster.stream().anyMatch(s -> target1.equals(s.terminalId())
|
||||
&& s.state() == MemberSession.State.BACKEND_ERROR),
|
||||
"the real backendErrorSink must have transitioned target1 to BACKEND_ERROR: " + roster);
|
||||
assertTrue(roster.stream().anyMatch(s -> target2.equals(s.terminalId())
|
||||
&& s.state() == MemberSession.State.BACKEND_ERROR),
|
||||
"the real backendErrorSink must have transitioned target2 to BACKEND_ERROR: " + roster);
|
||||
|
||||
PlacementException coolOff = assertThrows(PlacementException.class,
|
||||
() -> runtime.sessions().acquire("coolprofile", null, dir.toString(), null),
|
||||
"two distinct targets classified within the 60s window must cool the credential "
|
||||
+ "off (BackendOutagePolicy), refusing a third explicit-profile spawn");
|
||||
assertTrue(coolOff.getMessage().contains("cooling off"), coolOff.getMessage());
|
||||
} finally {
|
||||
if (ports.shutdownHook != null) ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,89 +0,0 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #248: this is the test that was actually missing. {@code Fleetd.main} builds its {@code
|
||||
* CompletionResolver} from an 8-argument constructor, and the ticket's own measurement proved two
|
||||
* ways to silently unwire it — both compiled with 0 errors and left every existing test green:
|
||||
*
|
||||
* <ul>
|
||||
* <li>replacing the worktree/branch argument (the 8th) with {@code _ -> null} — drops
|
||||
* fleetd#241's fallback-report location entirely;</li>
|
||||
* <li>replacing {@code backendErrorPatterns, backendErrorSink} (5th/6th) with {@code
|
||||
* BackendErrorPatternLookup.legacy(), BackendErrorSink.none()} — drops fleetd#201 Unit 5's
|
||||
* backend-error classification and cool-off entirely.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>Neither mutation could be caught by any test that constructs its own {@code
|
||||
* CompletionResolver} (every test before this one did exactly that) or by a test of {@link
|
||||
* Fleetd#worktreeBranchLookup}, {@link Fleetd#backendErrorPatternLookup}, or {@link
|
||||
* Fleetd#backendErrorSink} in isolation (see {@code FleetdWorktreeBranchLookupTest}, {@code
|
||||
* FleetdBackendErrorPatternLookupTest}, {@code FleetdBackendErrorSinkTest}) — those prove the
|
||||
* factories work, never that {@code main} still calls them. This class is a plain source-text
|
||||
* assertion on {@code Fleetd.java} — crude, but honest about what it checks, and it turns red the
|
||||
* instant the wiring is dropped, mirroring the same fallback shape {@link
|
||||
* FleetdFleetAppConstructionTest} already uses for a different constructor argument.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
||||
* CompletionResolver} and never runs {@code main}.
|
||||
*/
|
||||
class FleetdCompletionResolverWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] CompletionResolver's construction call still names backendErrorPatterns and backendErrorSink")
|
||||
void backendErrorArgumentsAreStillNamedAtTheCallSite() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"exhaustionSink, backendErrorPatterns, backendErrorSink, System::nanoTime,"),
|
||||
"CompletionResolver's construction call must still pass backendErrorPatterns and "
|
||||
+ "backendErrorSink as its 5th/6th arguments. Replacing them with "
|
||||
+ "BackendErrorPatternLookup.legacy()/BackendErrorSink.none() (fleetd #248's measured "
|
||||
+ "mutation) compiles with 0 errors and leaves every behavioural test green — this "
|
||||
+ "source check is what must go red instead.");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] CompletionResolver's construction call still passes worktreeBranchLookup(sessions::roster)")
|
||||
void worktreeBranchLookupIsStillPassedAtTheCallSite() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains("worktreeBranchLookup(sessions::roster)"),
|
||||
"CompletionResolver's construction call must still pass worktreeBranchLookup(sessions::roster) "
|
||||
+ "as its 8th (last) argument. Replacing it with the inert `_ -> null` (fleetd #248's "
|
||||
+ "other measured mutation) compiles with 0 errors and leaves every behavioural test "
|
||||
+ "green — this source check is what must go red instead.");
|
||||
assertFalse(source.contains("System::nanoTime,\n _ -> null"),
|
||||
"the worktree/branch argument must never regress to the inert `_ -> null` literal");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] backendErrorPatterns is assigned from the extracted backendErrorPatternLookup(...) factory")
|
||||
void backendErrorPatternsComesFromTheFactory() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"BackendErrorPatternLookup backendErrorPatterns = backendErrorPatternLookup(sessions::roster,"),
|
||||
"backendErrorPatterns must be assigned from Fleetd.backendErrorPatternLookup(...), not an "
|
||||
+ "inline lambda that a source check on the CompletionResolver call alone cannot see "
|
||||
+ "through");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] backendErrorSink is assigned from the extracted backendErrorSink(...) factory")
|
||||
void backendErrorSinkComesFromTheFactory() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"BackendErrorSink backendErrorSink = backendErrorSink(sessions, () -> config.get().profiles(),"),
|
||||
"backendErrorSink must be assigned from Fleetd.backendErrorSink(...), not an inline lambda "
|
||||
+ "that a source check on the CompletionResolver call alone cannot see through");
|
||||
}
|
||||
}
|
||||
@@ -1,31 +0,0 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* CB-185: {@code ConnectionIdentity} must resolve a caller's pane on EITHER herdr daemon (a
|
||||
* lead's MCP connection resolves against the lead daemon; a member's against the member daemon).
|
||||
* Pinning {@code PaneLocator} to {@code memberHerdr} alone — the bug this guards against — leaves
|
||||
* every lead's own connection unresolvable ({@code callerTerminal == null}) the moment
|
||||
* {@code memberHerdrSocket} names a second daemon, which breaks {@code fleet_reply}/{@code
|
||||
* fleet_ask} and {@code fleet_whoami} for a lead. A unit test on {@link
|
||||
* dev.ltms.fleet.herdr.PaneLocator} alone (see {@code PaneLocatorTest}) proves the class CAN
|
||||
* search two clients, but not that {@code Fleetd.main} actually wires it that way — hence this
|
||||
* source-level assertion, the same technique {@code FleetdHerdrControlConstructionTest} uses.
|
||||
*/
|
||||
class FleetdConnectionIdentityConstructionTest {
|
||||
@Test
|
||||
void connectionIdentitySearchesBothDaemonsNotJustTheMemberOne() throws Exception {
|
||||
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
assertFalse(source.contains("new PaneLocator(memberHerdr)"),
|
||||
"PaneLocator must not be pinned to the member daemon alone — a lead's own "
|
||||
+ "connection resolves against the LEAD daemon and would never be found");
|
||||
assertTrue(source.contains("new PaneLocator(herdr, memberHerdr)"),
|
||||
"PaneLocator must search the lead daemon first, then the member daemon");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,86 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
|
||||
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
|
||||
import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #589 Group 1: {@link Fleetd#exhaustedPatternLookup} is the factory that replaced {@code
|
||||
* main}'s inline lambda — resolve a herdr {@code target} to its session's profile, then to that
|
||||
* profile's live {@link LiveExhaustedPatterns#patternFor}. Same shape as {@link
|
||||
* Fleetd#worktreeBranchLookup} (which {@code FleetdWorktreeBranchLookupTest} pins the same way).
|
||||
*
|
||||
* <p>Before this ticket the lambda was built inline in {@code main} and untestable: replacing it
|
||||
* with {@code target -> null} — the exact shape of {@link ExhaustedPatternLookup#none()} — compiled
|
||||
* with 0 errors and left every existing test green. Per the ticket, this is the worst consequence
|
||||
* in the whole #589 sweep: a genuine usage-limit refusal would stop being classified as {@code
|
||||
* BACKEND_EXHAUSTED} and would be handed back to a waiting {@code fleet_send} as if it were real
|
||||
* completed work.
|
||||
*/
|
||||
class FleetdExhaustedPatternLookupWiringTest {
|
||||
|
||||
private static final String YAML = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: claude-opus-5
|
||||
exhaustedPattern: "usage limit"
|
||||
""";
|
||||
|
||||
private static MemberSession session(String terminal, String profile) {
|
||||
return new MemberSession("pane-" + terminal, terminal, profile, MemberRole.DEV,
|
||||
"/cwd", null, 0L, 0L, 0, MemberSession.State.READY, null, null);
|
||||
}
|
||||
|
||||
private static LiveExhaustedPatterns liveExhaustedPatterns(Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, YAML);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
return new LiveExhaustedPatterns(() -> ConfigRef.fixed(cfg).get().profiles());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a known target resolves through its session's profile to that profile's live pattern")
|
||||
void knownTargetResolvesThroughItsProfile(@TempDir Path dir) throws Exception {
|
||||
LiveExhaustedPatterns patterns = liveExhaustedPatterns(dir);
|
||||
ExhaustedPatternLookup lookup = Fleetd.exhaustedPatternLookup(
|
||||
() -> List.of(session("term1", "terra")), patterns);
|
||||
|
||||
Pattern resolved = lookup.patternFor("term1");
|
||||
|
||||
assertNotNull(resolved,
|
||||
"the lookup must resolve term1 -> profile 'terra' -> LiveExhaustedPatterns.patternFor("
|
||||
+ "'terra') — replacing the lambda body with 'target -> null' at the "
|
||||
+ "Fleetd.exhaustedPatternLookup call site must fail this assertion");
|
||||
assertTrue(resolved.matcher("the usage limit has been reached").find());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unknown target resolves to null, not a thrown exception")
|
||||
void unknownTargetResolvesToNull(@TempDir Path dir) throws Exception {
|
||||
LiveExhaustedPatterns patterns = liveExhaustedPatterns(dir);
|
||||
ExhaustedPatternLookup lookup = Fleetd.exhaustedPatternLookup(
|
||||
() -> List.of(session("term1", "terra")), patterns);
|
||||
|
||||
assertNull(lookup.patternFor("term_stranger"));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,62 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.concurrent.atomic.AtomicBoolean;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #589 Group 1: {@link Fleetd#forwardingExhaustionSink} is the factory that replaced
|
||||
* {@code main}'s inline {@code ExhaustionSink.forwardingTo(exhaustionSinkRef::get)} (fleetd #175's
|
||||
* construction-order break: the adapters need a sink before {@code sessions} exists to build the
|
||||
* real one). Before this ticket that call site was untestable wiring: replacing the supplier
|
||||
* argument with a hardcoded {@code () -> ExhaustionSink.none()} compiled with 0 errors and left
|
||||
* every existing test green, because no test builds the object {@code main} actually wires and
|
||||
* then mutates the reference afterward — every existing {@code ExhaustionSink.forwardingTo} caller
|
||||
* in this codebase reads and writes the SAME reference within one test, so a hardcoded-none supplier
|
||||
* and a correctly-forwarding one are indistinguishable to them.
|
||||
*
|
||||
* <p>This test builds the reference, builds the forwarder from it, and only THEN repoints the
|
||||
* reference at a spy sink — the discriminating order fleetd #175's whole design depends on
|
||||
* ({@code exhaustionSinkRef} starts at {@code none()} and is repointed once {@code sessions}
|
||||
* exists). A forwarder that captured a fixed target at construction time (the inert form) can never
|
||||
* see that later repoint.
|
||||
*/
|
||||
class FleetdExhaustionSinkForwardingWiringTest {
|
||||
|
||||
@Test
|
||||
@DisplayName("the forwarder reads the reference live: repointing it AFTER construction is honoured")
|
||||
void forwarderReadsTheReferenceLiveNotAFixedTargetCapturedAtConstruction() {
|
||||
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
|
||||
ExhaustionSink forwarder = Fleetd.forwardingExhaustionSink(exhaustionSinkRef);
|
||||
|
||||
AtomicBoolean spyCalled = new AtomicBoolean(false);
|
||||
exhaustionSinkRef.set((target, reason, profile) -> spyCalled.set(true));
|
||||
|
||||
forwarder.onExhausted("term_x", "usage limit reached", "terra");
|
||||
|
||||
assertTrue(spyCalled.get(),
|
||||
"forwardingExhaustionSink must delegate to whatever exhaustionSinkRef currently "
|
||||
+ "holds — hardcoding the supplier to () -> ExhaustionSink.none() at the "
|
||||
+ "Fleetd.forwardingExhaustionSink call site must fail this assertion, "
|
||||
+ "since the spy set into the reference after construction would never run");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("before any repoint, the forwarder is inert — it starts at none(), not a crash")
|
||||
void beforeAnyRepointTheForwarderIsInert() {
|
||||
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
|
||||
ExhaustionSink forwarder = Fleetd.forwardingExhaustionSink(exhaustionSinkRef);
|
||||
|
||||
AtomicBoolean spyCalled = new AtomicBoolean(false);
|
||||
forwarder.onExhausted("term_x", "usage limit reached", "terra");
|
||||
|
||||
assertFalse(spyCalled.get(), "nothing was ever wired to be called here — this only pins "
|
||||
+ "that the factory does not throw before a real sink is published");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,110 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.HashMap;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #589 Group 1: {@link Fleetd#publishExhaustionSink} is the factory that replaced {@code
|
||||
* main}'s previously untested two-statement sequence — build the real {@link
|
||||
* Fleetd#exhaustionSink}, then {@code exhaustionSinkRef.set(exhaustionSink)}. {@link
|
||||
* Fleetd#exhaustionSink} itself is already pinned by {@code FleetdExhaustionSinkWarningTest} (its
|
||||
* log text) — what was NEVER pinned is the {@code .set(...)} call: {@code main} could replace it
|
||||
* with {@code exhaustionSinkRef.set(ExhaustionSink.none())} and compile with 0 errors, leaving
|
||||
* every existing test green, because {@link Fleetd#exhaustionSink}'s own tests build and call the
|
||||
* sink directly, never through the reference {@code main} publishes it into.
|
||||
*
|
||||
* <p>This test proves the PUBLISHED reference — not a freshly rebuilt sink — is the one that
|
||||
* actually quarantines a credential, by reading {@link BackendQuarantine#isQuarantined} after
|
||||
* calling {@code exhaustionSinkRef.get().onExhausted(...)}, the same object {@link
|
||||
* Fleetd#forwardingExhaustionSink} forwards to in production.
|
||||
*/
|
||||
class FleetdExhaustionSinkPublishWiringTest {
|
||||
|
||||
private static final String YAML = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: claude-opus-5
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
private static SessionManager emptyRosterSessions() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
FleetConfig.Profile dummy = new FleetConfig.Profile(
|
||||
"dummy", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
|
||||
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(dummy.profile(), dummy), dummy.profile(), _ -> "tok");
|
||||
// Never acquires a session — publishExhaustionSink's built sink resolves target -> profile
|
||||
// via the profileHint fallback (fleetd #234), exactly like OpenCodeLauncher's real call
|
||||
// site does, so this never needs a populated roster.
|
||||
return new SessionManager(launcher);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("the published reference actually quarantines — not a rebuilt-but-never-set sink")
|
||||
void publishedReferenceActuallyQuarantines(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, YAML);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = ConfigRef.fixed(cfg);
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
Map<String, String> reasonByCredential = new HashMap<>();
|
||||
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
|
||||
|
||||
Fleetd.publishExhaustionSink(exhaustionSinkRef, emptyRosterSessions(), config, quarantine,
|
||||
reasonByCredential, cfg);
|
||||
exhaustionSinkRef.get().onExhausted("term_x", "The usage limit has been reached", "terra");
|
||||
|
||||
assertTrue(quarantine.isQuarantined("terra"),
|
||||
"publishExhaustionSink must repoint exhaustionSinkRef at the REAL sink — "
|
||||
+ "replacing the .set(...) call with exhaustionSinkRef.set(ExhaustionSink.none()) "
|
||||
+ "at the Fleetd.publishExhaustionSink call site must fail this assertion, "
|
||||
+ "since none()'s onExhausted does nothing");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("before publishing, the reference is still inert — no quarantine, no crash")
|
||||
void beforePublishingTheReferenceIsStillInert(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, YAML);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = ConfigRef.fixed(cfg);
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
|
||||
|
||||
exhaustionSinkRef.get().onExhausted("term_x", "The usage limit has been reached", "terra");
|
||||
|
||||
assertFalse(quarantine.isQuarantined("terra"),
|
||||
"nothing was published yet — this only pins the starting state the other test's "
|
||||
+ "assertion actually distinguishes from");
|
||||
}
|
||||
}
|
||||
@@ -1,29 +0,0 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* CB-185: {@code FleetApp} must be constructed with BOTH herdr clients (the lead's and the
|
||||
* member's), never the raw lead-only {@code herdr}. Passing only {@code herdr} — the bug this
|
||||
* guards against — makes {@code GET /healthz} green while the member daemon is down (so every
|
||||
* spawn fails invisibly) and silently drops every member workspace from {@code GET /sessions}.
|
||||
* A behavioural test on {@code FleetApp} alone (see {@code FleetAppTwoDaemonTest}) proves the
|
||||
* class merges/gates correctly when given two clients, but not that {@code Fleetd.main} actually
|
||||
* passes it two — hence this source-level assertion, mirroring
|
||||
* {@code FleetdHerdrControlConstructionTest}.
|
||||
*/
|
||||
class FleetdFleetAppConstructionTest {
|
||||
@Test
|
||||
void fleetAppIsConstructedWithBothHerdrDaemons() throws Exception {
|
||||
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
assertFalse(source.contains("new FleetApp(herdr, workers,"),
|
||||
"FleetApp must not be constructed with the lead-only herdr client");
|
||||
assertTrue(source.contains("new FleetApp(herdr, memberHerdr, workers,"),
|
||||
"FleetApp must be constructed with both the lead and the member herdr client");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,160 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #426: {@code FleetHealthMonitor.coverage} had zero references anywhere in the test
|
||||
* tree — not the method, not either output string, not the field it populates. {@code
|
||||
* FleetHealthMonitorCoverageTest} (package {@code dev.ltms.fleet.health}) pins the three-branch
|
||||
* method itself; that is the easy half.
|
||||
*
|
||||
* <p>The half that actually matters is this one: {@link Fleetd#healthCoverageSource} is the exact
|
||||
* call site {@code Fleetd.main} wires into {@code FleetMcp}'s constructor, and it is what feeds
|
||||
* {@code fleet_list}'s {@code healthCoverage} field (see {@code FleetMcp#listFleet}'s {@code
|
||||
* result.put("healthCoverage", healthCoverage.value().get())}). Measured precedent on fleetd #423
|
||||
* (for #415): swapping two arguments at a call site like this one — recreating #415's defect with
|
||||
* the keys exchanged — compiled with 0 errors and ran the ENTIRE suite (1506 tests) green. A
|
||||
* thoroughly-tested method proves nothing about whether the call site pairs its arguments correctly;
|
||||
* only a test that drives the call site itself can catch that.
|
||||
*
|
||||
* <p><strong>Why this is not driven through a real {@code Fleetd.main} the way #407 drives its five
|
||||
* reporters</strong> (keep the config invalid, assert on the log line emitted before {@code
|
||||
* cfg.validateAll()} throws): all five of #407's reporters run in {@code main} before {@code
|
||||
* validateAll()} (line ~171). {@link Fleetd#healthCoverageSource} is built during {@code FleetMcp}
|
||||
* construction, which happens only after {@code UnixSocketHerdrClient.connect} has already opened a
|
||||
* real herdr socket (line ~188) and after {@code sessions}/{@code workers} are constructed. Reaching
|
||||
* this call site by actually running {@code main} would require a real socket connect — banned by
|
||||
* this ticket's hard constraints — so #407's option 1 does not apply here. Instead {@link
|
||||
* Fleetd#healthCoverageSource} is extracted to a directly-callable package-private factory, the same
|
||||
* shape {@link Fleetd#capacitySource} and {@link Fleetd#quarantineSource} already use for the same
|
||||
* reason (see {@code FleetdCapacitySourceWiringTest}, the direct precedent this test follows).
|
||||
*
|
||||
* <p>Uses a real {@link FleetConfig#load} + {@link ConfigRef} (no socket, no port bind, no spawn,
|
||||
* nothing written outside {@code @TempDir}) so the fixture goes through the actual YAML parser and
|
||||
* {@code FleetConfig.Health}/{@code Notifications} records, not a hand-built stand-in that could
|
||||
* silently drift from what the parser actually produces.
|
||||
*/
|
||||
class FleetdHealthCoverageSourceWiringTest {
|
||||
|
||||
private static final String BASE = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
""";
|
||||
|
||||
private static final String HEALTH_DISABLED_WITH_WEBHOOK = BASE + """
|
||||
health:
|
||||
enabled: false
|
||||
notifications:
|
||||
mode: webhook
|
||||
""";
|
||||
|
||||
private static final String HEALTH_DETECTION_ONLY = BASE + """
|
||||
health:
|
||||
enabled: true
|
||||
""";
|
||||
|
||||
private static final String HEALTH_FULL = BASE + """
|
||||
health:
|
||||
enabled: true
|
||||
notifications:
|
||||
mode: webhook
|
||||
""";
|
||||
|
||||
private static final String HEALTH_ABSENT = BASE;
|
||||
|
||||
@Test
|
||||
@DisplayName("enabled: false reports off, even with a webhook configured")
|
||||
void disabledHealthReportsOff(@TempDir Path dir) throws Exception {
|
||||
FleetMcp.HealthCoverageSource source = sourceFor(dir, HEALTH_DISABLED_WITH_WEBHOOK);
|
||||
|
||||
assertEquals("off", source.value().get(),
|
||||
"health.enabled: false must report 'off' regardless of notifications — flipping "
|
||||
+ "the 'enabled' argument at the HealthCoverageSource call site would report "
|
||||
+ "'full' here instead");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("enabled with no notifications reports detection-only")
|
||||
void enabledWithoutNotificationsReportsDetectionOnly(@TempDir Path dir) throws Exception {
|
||||
FleetMcp.HealthCoverageSource source = sourceFor(dir, HEALTH_DETECTION_ONLY);
|
||||
|
||||
assertEquals("detection-only", source.value().get());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("enabled with a webhook configured reports full")
|
||||
void enabledWithNotificationsReportsFull(@TempDir Path dir) throws Exception {
|
||||
FleetMcp.HealthCoverageSource source = sourceFor(dir, HEALTH_FULL);
|
||||
|
||||
assertEquals("full", source.value().get(),
|
||||
"health.enabled: true with notifications.mode: webhook must report 'full' — "
|
||||
+ "swapping 'full' and 'detection-only' at the call site, or breaking the "
|
||||
+ "enabled/notificationConfigured argument pairing, would report "
|
||||
+ "'detection-only' here instead");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an absent health: block reports off")
|
||||
void absentHealthBlockReportsOff(@TempDir Path dir) throws Exception {
|
||||
FleetMcp.HealthCoverageSource source = sourceFor(dir, HEALTH_ABSENT);
|
||||
|
||||
assertEquals("off", source.value().get());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #426, the live-wiring half: {@code health.notifications} is a {@code
|
||||
* ConfigRef.SPLIT_KEYS} entry, and {@link Fleetd#healthCoverageSource} reads {@code
|
||||
* config.get().health()} live (not the frozen startup {@code cfg}) — exactly like {@link
|
||||
* Fleetd#capacitySource}'s {@code maxLoad} ({@code FleetdCapacitySourceWiringTest}'s {@code
|
||||
* reloadedMaxLoadStillChangesWhatFleetListReports}). A hot-reloaded notifications block must
|
||||
* change what {@code fleet_list} reports without a restart; a fix that froze the whole source
|
||||
* against the startup snapshot would silently break that and every other test above would stay
|
||||
* green, since none of them reload.
|
||||
*/
|
||||
@Test
|
||||
@DisplayName("a hot notifications reload still changes what fleet_list reports")
|
||||
void reloadedNotificationsStillChangeWhatFleetListReports(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, HEALTH_DETECTION_ONLY);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = new ConfigRef(file, cfg);
|
||||
|
||||
FleetMcp.HealthCoverageSource source = Fleetd.healthCoverageSource(config);
|
||||
assertEquals("detection-only", source.value().get(),
|
||||
"sanity: detection-only before any reload");
|
||||
|
||||
Files.writeString(file, HEALTH_FULL);
|
||||
assertTrue(config.reload().applied());
|
||||
// The live snapshot now carries the webhook — proves the reload really happened and this
|
||||
// test is not accidentally passing because nothing changed.
|
||||
assertTrue(config.get().health().notifications() != null
|
||||
&& config.get().health().notifications().configured(),
|
||||
"sanity: the reloaded config really carries a configured webhook");
|
||||
|
||||
assertEquals("full", source.value().get(),
|
||||
"the SAME HealthCoverageSource instance must reflect a reloaded notifications "
|
||||
+ "block without a restart — health.notifications is read live off "
|
||||
+ "config.get(), exactly like capacitySource's maxLoad");
|
||||
}
|
||||
|
||||
private static FleetMcp.HealthCoverageSource sourceFor(Path dir, String yaml) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, yaml);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = new ConfigRef(file, cfg);
|
||||
return Fleetd.healthCoverageSource(config);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,80 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.msg.MessageService;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.function.BiConsumer;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #589 Group 3, site {@code :591}. {@code Fleetd.main} wires {@link
|
||||
* dev.ltms.fleet.health.FleetHealthMonitor}'s {@code failTarget} callback with {@code
|
||||
* messages::abandon} — before this ticket that was an inline argument to {@code new
|
||||
* FleetHealthMonitor(...)}. Measured: replacing it with a no-op {@code BiConsumer} at the call
|
||||
* site compiles with 0 errors and leaves the full suite green, because nothing else in the tree
|
||||
* ever drives that specific constructor argument. In production it means a member the monitor
|
||||
* classifies {@code GONE}/{@code NEVER_READY} never has its pending ticket failed — the caller
|
||||
* keeps reporting {@code PENDING} for the full 30-minute async timeout instead of the immediate,
|
||||
* accurate failure CB-580 exists to give it.
|
||||
*
|
||||
* <p>This test calls {@link Fleetd#healthFailTarget} directly — never {@code FleetHealthMonitor}
|
||||
* or {@code main} — against a real {@link MessageService}, using the same {@code sendAsync} +
|
||||
* {@code poll} observable {@link MessageServiceTest} already relies on to pin {@code
|
||||
* MessageService.abandon} itself.
|
||||
*/
|
||||
class FleetdHealthFailTargetWiringTest {
|
||||
|
||||
private static final String T = "term_a";
|
||||
|
||||
@Test
|
||||
@DisplayName("Fleetd.healthFailTarget delegates to the real MessageService.abandon, not a no-op")
|
||||
void healthFailTargetDelegatesToMessagesAbandon() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("$ prompt");
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
Injector injector = new Injector(agents);
|
||||
MessageService messages = new MessageService(agents, injector, rendezvous);
|
||||
|
||||
BiConsumer<String, String> failTarget = Fleetd.healthFailTarget(messages);
|
||||
|
||||
String ticket = messages.sendAsync(T, "long task");
|
||||
awaitWaiting(rendezvous);
|
||||
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase());
|
||||
|
||||
failTarget.accept(T, "member unreachable (health monitor)");
|
||||
|
||||
MessageService.TaskView view = null;
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (System.currentTimeMillis() < deadline) {
|
||||
view = messages.poll(ticket);
|
||||
if (view.phase() != MessageService.Phase.PENDING) {
|
||||
break;
|
||||
}
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
}
|
||||
assertEquals(MessageService.Phase.FAILED, view == null ? null : view.phase(),
|
||||
"Fleetd.healthFailTarget(messages) must return messages::abandon — replacing it "
|
||||
+ "with a no-op BiConsumer at the Fleetd.healthFailTarget call site means "
|
||||
+ "this ticket is never failed and keeps polling as PENDING");
|
||||
assertTrue(view.detail() != null && view.detail().contains("member unreachable"),
|
||||
"the failure reason passed to failTarget.accept must reach MessageService.abandon "
|
||||
+ "and end up in the ticket's detail");
|
||||
}
|
||||
|
||||
private static void awaitWaiting(Rendezvous rendezvous) throws InterruptedException {
|
||||
long deadline = System.currentTimeMillis() + 2000;
|
||||
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting(T), "send should have opened its rendezvous waiter");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,100 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Map;
|
||||
import java.util.function.Function;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* fleetd #602 gauge-wiring: {@link Fleetd#leadConfigDirLookup} is the factory {@code Fleetd.main}
|
||||
* wires into {@code FleetMcp.LeadConfigDirSource} so {@code fleet_list}'s {@code context} row reads
|
||||
* the transcript directory a lead's OWN profile actually writes to, instead of always falling back
|
||||
* to the built-in {@code <user.home>/.claude} default (see {@code LeadContextGauge}).
|
||||
*
|
||||
* <p>{@code FleetMcpLeadContextGaugeWiringTest} proves the directory this factory returns is what
|
||||
* actually gets read; this class proves the factory's own matching logic — the same
|
||||
* {@code fleet.leaders.<name>.profile} link {@code Fleetd.leadSeatLookup} already follows (see
|
||||
* {@code FleetdLeadSeatLookupTest}), one step further to that profile's own {@code configDir:}.
|
||||
*/
|
||||
class FleetdLeadConfigDirLookupTest {
|
||||
|
||||
private static FleetConfig.Profile profileWithConfigDir(String name, String configDir) {
|
||||
return new FleetConfig.Profile(name, null, "claude-sonnet-5", configDir, null, null,
|
||||
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
|
||||
null, 3, true, null, null, null);
|
||||
}
|
||||
|
||||
private static FleetConfig.Leader leadOnProfile(String profile) {
|
||||
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a lead on a profile that sets configDir resolves to that directory")
|
||||
void leadOnAProfileWithConfigDirResolvesToIt() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", "/mnt/opus-claude"));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||
Function<String, String> lookup = Fleetd.leadConfigDirLookup(() -> profiles, leaders);
|
||||
|
||||
assertEquals("/mnt/opus-claude", lookup.apply("primary"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("changing the config from directory A to directory B changes what the lookup reports")
|
||||
void configChangeFromDirectoryAToDirectoryBChangesTheAnswer() {
|
||||
java.util.concurrent.atomic.AtomicReference<Map<String, FleetConfig.Profile>> profilesRef =
|
||||
new java.util.concurrent.atomic.AtomicReference<>(
|
||||
Map.of("opus", profileWithConfigDir("opus", "/mnt/dir-a")));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||
Function<String, String> lookup = Fleetd.leadConfigDirLookup(profilesRef::get, leaders);
|
||||
|
||||
assertEquals("/mnt/dir-a", lookup.apply("primary"), "must read directory A before the config changes");
|
||||
|
||||
profilesRef.set(Map.of("opus", profileWithConfigDir("opus", "/mnt/dir-b")));
|
||||
assertEquals("/mnt/dir-b", lookup.apply("primary"), "must read directory B once the LIVE config changes — "
|
||||
+ "a lookup that snapshotted the profile map at construction would still answer directory A here");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a lead entry with no `profile:` (recognise-only) resolves to null, not a thrown exception")
|
||||
void recogniseOnlyLeadWithNoProfileResolvesToNull() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", "/mnt/opus-claude"));
|
||||
FleetConfig.Leader recogniseOnly = new FleetConfig.Leader(null, "lead: primary", 1, "lead:", 10,
|
||||
"claude", "claude-sonnet-5");
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", recogniseOnly);
|
||||
Function<String, String> lookup = Fleetd.leadConfigDirLookup(() -> profiles, leaders);
|
||||
|
||||
assertNull(lookup.apply("primary"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a lead naming a profile that is not configured resolves to null, not a thrown exception")
|
||||
void leadOnAnUnconfiguredProfileResolvesToNull() {
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("ghost-profile"));
|
||||
Function<String, String> lookup = Fleetd.leadConfigDirLookup(Map::of, leaders);
|
||||
|
||||
assertNull(lookup.apply("primary"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a lead on a profile that sets no configDir override resolves to null")
|
||||
void leadOnAProfileWithNoConfigDirResolvesToNull() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", null));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||
Function<String, String> lookup = Fleetd.leadConfigDirLookup(() -> profiles, leaders);
|
||||
|
||||
assertNull(lookup.apply("primary"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unrecognised lead name resolves to null, not a thrown exception")
|
||||
void unrecognisedLeadNameResolvesToNull() {
|
||||
Function<String, String> lookup = Fleetd.leadConfigDirLookup(Map::of, Map.of());
|
||||
|
||||
assertNull(lookup.apply("ghost-lead"));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,98 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
|
||||
/**
|
||||
* fleetd #602 gauge-wiring follow-up (PR #606 review comment 17353): {@code Fleetd.main}'s {@code
|
||||
* LeadConfigDirSource} local used to be a bare {@code new FleetMcp.LeadConfigDirSource(
|
||||
* leadConfigDirLookup(...))} built inline, with nothing a test could call directly. Measured on
|
||||
* that shape: replacing the whole expression with {@code FleetMcp.LeadConfigDirSource.none()} at
|
||||
* the call site compiled with 0 errors and left the full 1822-test suite green — the daemon could
|
||||
* be changed to always report every lead's context as {@code UNKNOWN}, forever, and no test would
|
||||
* notice. That is the same hand-built-vs-config-wired shape as fleetd #561/#248/#426/#562
|
||||
* ({@code FleetdLoopHealthSourceWiringTest}).
|
||||
*
|
||||
* <p>{@code FleetMcpLeadContextGaugeWiringTest} and {@code FleetdLeadConfigDirLookupTest} both
|
||||
* predate this class and are both still correct — but neither can catch the mutation above. One
|
||||
* builds its own {@code FleetMcp} and hands it its own {@code LeadConfigDirSource}; the other
|
||||
* builds its own lookup and calls {@link Fleetd#leadConfigDirLookup} directly. Neither one ever
|
||||
* calls the thing {@code Fleetd.main} actually calls.
|
||||
*
|
||||
* <p>The fix extracts the inline {@code new} into {@link Fleetd#leadConfigDirSource}, a
|
||||
* package-private factory in the same style as {@link Fleetd#loopHealthSource}/{@link
|
||||
* Fleetd#capacitySource}/{@link Fleetd#healthCoverageSource} — which is exactly what makes it
|
||||
* directly callable here. This test calls that factory with real {@link FleetConfig.Profile}/
|
||||
* {@link FleetConfig.Leader} fixtures (the same shapes {@code FleetdLeadConfigDirLookupTest}
|
||||
* already uses) and asserts the returned source resolves a real {@code configDir} — a property
|
||||
* that would be false if {@link Fleetd#leadConfigDirSource} were mutated to {@code return
|
||||
* FleetMcp.LeadConfigDirSource.none();}. Measured: mutating exactly that line makes
|
||||
* {@link #resolvesTheRealConfiguredConfigDir()} fail ({@code expected: </mnt/opus-claude> but was:
|
||||
* <null>}); restoring it makes the whole suite green again.
|
||||
*
|
||||
* <p><b>What this class does not and cannot cover.</b> {@code main}'s own line —
|
||||
* {@code leadConfigDirSource(() -> config.get().profiles(), leaders)} — could itself be swapped
|
||||
* for a bare {@code FleetMcp.LeadConfigDirSource.none()}, bypassing this factory entirely. Measured:
|
||||
* that exact mutation compiles with 0 errors and leaves every test in this file, and the full
|
||||
* 1825-test suite, green. {@link Fleetd#loopHealthSource}'s own wiring test has the identical gap
|
||||
* for its own one-line call in {@code main} — no test in this codebase calls {@code Fleetd.main}
|
||||
* far enough to observe which factory call it made. This class narrows the gap from "nothing tests
|
||||
* the wiring" (the pre-extraction state this ticket found) to "the factory's own logic is pinned,
|
||||
* and main's call to it is a one-line, visually-verifiable delegation" — the same standard already
|
||||
* accepted for {@code loopHealthSource}/{@code capacitySource}/{@code healthCoverageSource}.
|
||||
*/
|
||||
class FleetdLeadConfigDirSourceWiringTest {
|
||||
|
||||
private static FleetConfig.Profile profileWithConfigDir(String name, String configDir) {
|
||||
return new FleetConfig.Profile(name, null, "claude-sonnet-5", configDir, null, null,
|
||||
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
|
||||
null, 3, true, null, null, null);
|
||||
}
|
||||
|
||||
private static FleetConfig.Leader leadOnProfile(String profile) {
|
||||
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("the returned source resolves the lead's REAL configured configDir, not a hardcoded null")
|
||||
void resolvesTheRealConfiguredConfigDir() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", "/mnt/opus-claude"));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||
|
||||
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
|
||||
|
||||
assertEquals("/mnt/opus-claude", source.configDirFor().apply("primary"),
|
||||
"the configDirFor function must delegate to the real leadConfigDirLookup — mutating "
|
||||
+ "Fleetd.leadConfigDirSource's own body to `return FleetMcp.LeadConfigDirSource.none();` "
|
||||
+ "must fail this assertion (measured: it does — see this test's class javadoc for the "
|
||||
+ "companion measurement on main's one-line call to this factory, which this assertion "
|
||||
+ "does not and structurally cannot cover)");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a lead on a profile with no configDir override still resolves to null, not a crash")
|
||||
void leadWithNoConfigDirOverrideResolvesToNull() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", null));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||
|
||||
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
|
||||
|
||||
assertNull(source.configDirFor().apply("primary"),
|
||||
"no configDir: override configured ⇒ null, so LeadContextGauge falls back to its own default");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unrecognised lead name resolves to null, not a thrown exception")
|
||||
void unrecognisedLeadNameResolvesToNull() {
|
||||
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(Map::of, Map.of());
|
||||
|
||||
assertNull(source.configDirFor().apply("ghost-lead"));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,156 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.HerdrException;
|
||||
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.function.Function;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
|
||||
/**
|
||||
* fleetd #609: {@link Fleetd#leadContextLookup} is the factory {@code Fleetd.main} wires into
|
||||
* {@code LeadHeartbeatLoop.LeadContextSource} so the idle-lead heartbeat can read a lead's own
|
||||
* {@link LeadContextGauge} reading by its terminal id. Three hops: terminal → lead name (unknown ⇒
|
||||
* UNKNOWN), lead name → configDir, and {@code agents.get(terminal)} for the live session id and
|
||||
* agent type the gauge itself needs — a herdr failure on that last hop must degrade to UNKNOWN, not
|
||||
* throw and kill the heartbeat's own tick.
|
||||
*
|
||||
* <p>{@code AgentControl.agentCall} resolves a {@code term_}-prefixed target's pane id via a first
|
||||
* {@code agent.list} round trip (see its own javadoc); this class's terminal id deliberately does
|
||||
* NOT start with {@code term_} so the stub {@link HerdrClient} below only needs to answer
|
||||
* {@code agent.get} — the one call this factory actually depends on.
|
||||
*/
|
||||
class FleetdLeadContextLookupTest {
|
||||
|
||||
private static final String LEAD_TERMINAL = "leadpane1";
|
||||
private static final String LEAD_NAME = "opus";
|
||||
private static final String SESSION_ID = "sess-609-happy-path";
|
||||
|
||||
private static String usageLine(long tokens) {
|
||||
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
|
||||
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
|
||||
}
|
||||
|
||||
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} carrying one usage record. */
|
||||
private static void writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
|
||||
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
|
||||
Files.createDirectories(projectDir);
|
||||
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
|
||||
}
|
||||
|
||||
/** An {@link AgentControl} whose every {@code agent.get} answers with the given session/type/status. */
|
||||
private static AgentControl agentControlStub(String sessionId, String agentType, String status) {
|
||||
HerdrClient client = new HerdrClient() {
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) {
|
||||
if (!"agent.get".equals(method)) {
|
||||
throw new HerdrException("stub has no canned response for " + method);
|
||||
}
|
||||
try {
|
||||
String sessionField = sessionId == null ? ""
|
||||
: ",\"agent_session\":{\"kind\":\"id\",\"value\":\"" + sessionId + "\"}";
|
||||
String agentField = agentType == null ? "null" : "\"" + agentType + "\"";
|
||||
return new ObjectMapper().readTree(("""
|
||||
{"type":"agent_info","agent":{"terminal_id":"%s","agent":%s,
|
||||
"agent_status":"%s"%s}}""")
|
||||
.formatted(LEAD_TERMINAL, agentField, status, sessionField));
|
||||
} catch (Exception e) {
|
||||
throw new HerdrException("stub decode failed", e);
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
};
|
||||
return new AgentControl(client);
|
||||
}
|
||||
|
||||
/** An {@link AgentControl} whose every herdr call fails — models a herdr hiccup mid-tick. */
|
||||
private static AgentControl throwingAgentControl() {
|
||||
HerdrClient client = new HerdrClient() {
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) {
|
||||
throw new HerdrException("herdr unreachable (stub)");
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
};
|
||||
return new AgentControl(client);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unrecognised terminal resolves to UNKNOWN, not a thrown exception")
|
||||
void unrecognisedTerminalResolvesToUnknown() {
|
||||
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
|
||||
new LeadContextGauge(), throwingAgentControl(), Map::of, name -> null);
|
||||
|
||||
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply("ghost-terminal"));
|
||||
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
|
||||
assertNull(reading.tokens());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("agents.get throwing degrades to UNKNOWN, no exception escapes")
|
||||
void agentsGetThrowingDegradesToUnknown() {
|
||||
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
|
||||
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
|
||||
new LeadContextGauge(), throwingAgentControl(), () -> liveLeadTerminals, name -> null);
|
||||
|
||||
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply(LEAD_TERMINAL));
|
||||
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
|
||||
"a herdr failure resolving the live agent must degrade to UNKNOWN, never kill the heartbeat's tick");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("the happy path resolves the configDir, sessionId and agentType through to the gauge")
|
||||
void happyPathPassesResolvedFactsThroughToTheGauge(@TempDir Path tmp) throws IOException {
|
||||
writeTranscript(tmp, SESSION_ID, 12_345);
|
||||
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
|
||||
AgentControl agents = agentControlStub(SESSION_ID, "claude", "idle");
|
||||
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
|
||||
new LeadContextGauge(), agents, () -> liveLeadTerminals,
|
||||
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
|
||||
|
||||
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
|
||||
|
||||
assertEquals(LeadContextGauge.State.OK, reading.state(),
|
||||
"the resolved configDir + the live agent's own sessionId/agentType must reach the gauge — "
|
||||
+ "a wrong hop anywhere in the chain would read no transcript and report UNKNOWN instead");
|
||||
assertEquals(12_345L, reading.tokens());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a lead whose agent type is not claude still resolves to UNKNOWN, never a crash")
|
||||
void nonClaudeAgentTypeResolvesToUnknown(@TempDir Path tmp) throws IOException {
|
||||
writeTranscript(tmp, SESSION_ID, 12_345);
|
||||
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
|
||||
AgentControl agents = agentControlStub(SESSION_ID, "opencode", "idle");
|
||||
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
|
||||
new LeadContextGauge(), agents, () -> liveLeadTerminals,
|
||||
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
|
||||
|
||||
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
|
||||
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
|
||||
"the agentType hop must reach the gauge too — a non-claude peer must not be misread as claude");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,108 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.HerdrException;
|
||||
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* fleetd #609: {@code Fleetd.main}'s {@code LeadHeartbeatLoop.LeadContextSource} local could be
|
||||
* swapped for a bare {@code LeadHeartbeatLoop.LeadContextSource.none()} at the call site — compiling
|
||||
* with 0 errors and leaving every pre-existing test green — exactly the shape #602/#606 already found
|
||||
* for {@code LeadConfigDirSource} (see {@code FleetdLeadConfigDirSourceWiringTest}'s own javadoc for
|
||||
* the measured version of that gap).
|
||||
*
|
||||
* <p>The fix follows the same pattern: {@link Fleetd#leadContextSource} is the extracted,
|
||||
* directly-callable factory {@code main} calls to build the source it hands {@code
|
||||
* LeadHeartbeatLoop}'s constructor. This test calls that exact factory and asserts it resolves a
|
||||
* REAL reading off a real transcript file — a property that would be false if {@link
|
||||
* Fleetd#leadContextSource} were mutated to {@code return LeadHeartbeatLoop.LeadContextSource.none();}.
|
||||
*
|
||||
* <p>What this class does not and cannot cover: {@code main}'s own one-line call to this factory
|
||||
* could itself be swapped for {@code LeadHeartbeatLoop.LeadContextSource.none()}, bypassing this
|
||||
* factory entirely — the same structural gap {@code FleetdLeadConfigDirSourceWiringTest} names for
|
||||
* its own factory, and for the same reason (no test in this codebase calls {@code Fleetd.main} far
|
||||
* enough to observe which factory call it made).
|
||||
*/
|
||||
class FleetdLeadContextSourceWiringTest {
|
||||
|
||||
/** Deliberately not {@code term_}-prefixed — see {@code FleetdLeadContextLookupTest}'s class doc. */
|
||||
private static final String LEAD_TERMINAL = "leadpane1";
|
||||
private static final String LEAD_NAME = "opus";
|
||||
private static final String SESSION_ID = "sess-609-wiring";
|
||||
|
||||
private static String usageLine(long tokens) {
|
||||
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
|
||||
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
|
||||
}
|
||||
|
||||
private static void writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
|
||||
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
|
||||
Files.createDirectories(projectDir);
|
||||
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
|
||||
}
|
||||
|
||||
private static AgentControl agentControlStub() {
|
||||
HerdrClient client = new HerdrClient() {
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) {
|
||||
if (!"agent.get".equals(method)) {
|
||||
throw new HerdrException("stub has no canned response for " + method);
|
||||
}
|
||||
try {
|
||||
return new ObjectMapper().readTree(("""
|
||||
{"type":"agent_info","agent":{"terminal_id":"%s","agent":"claude",
|
||||
"agent_status":"idle","agent_session":{"kind":"id","value":"%s"}}}""")
|
||||
.formatted(LEAD_TERMINAL, SESSION_ID));
|
||||
} catch (Exception e) {
|
||||
throw new HerdrException("stub decode failed", e);
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
};
|
||||
return new AgentControl(client);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("main's factory resolves a REAL reading, not the inert none() answer")
|
||||
void resolvesARealReadingNotTheInertNoneAnswer(@TempDir Path tmp) throws IOException {
|
||||
writeTranscript(tmp, SESSION_ID, 54_321);
|
||||
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
|
||||
|
||||
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
|
||||
agentControlStub(), () -> liveLeadTerminals, name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
|
||||
|
||||
LeadContextGauge.Reading reading = source.readingFor().apply(LEAD_TERMINAL);
|
||||
|
||||
assertEquals(LeadContextGauge.State.OK, reading.state(),
|
||||
"mutating Fleetd.leadContextSource's own body to `return LeadHeartbeatLoop.LeadContextSource.none();` "
|
||||
+ "must fail this assertion");
|
||||
assertEquals(54_321L, reading.tokens());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unrecognised lead terminal resolves to UNKNOWN, not a thrown exception")
|
||||
void unrecognisedTerminalResolvesToUnknown() {
|
||||
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
|
||||
agentControlStub(), Map::of, name -> null);
|
||||
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, source.readingFor().apply("ghost-terminal").state());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,48 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.net.ServerSocket;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #589 Group 3, site {@code :518}. {@code Fleetd.main} passes {@code LeadMailbox::open} as
|
||||
* the {@link Fleetd.LeadMailboxOpener} argument to {@code openLeadMailbox(...)} — before this
|
||||
* ticket that method reference was inline at the call site. Measured: replacing it with the inert
|
||||
* {@code (uri, selfCoordId, prefetch) -> null} compiles with 0 errors and leaves the full suite
|
||||
* green — {@code FleetdLeadMailboxSelectionTest} drives {@code openLeadMailbox} with its own
|
||||
* injected opener and never observes what {@code main} itself actually passes.
|
||||
*
|
||||
* <p>This test calls {@link Fleetd#leadMailboxOpener} directly and proves it is the real,
|
||||
* network-attempting opener rather than a stub: pointed at a guaranteed-closed local port, it must
|
||||
* throw, exactly mirroring {@link FleetdReplyInboxOpenerWiringTest} for the reply-inbox opener.
|
||||
* The inert form never attempts a connection and returns {@code null} without throwing, so it
|
||||
* fails this assertion silently.
|
||||
*/
|
||||
class FleetdLeadMailboxOpenerWiringTest {
|
||||
|
||||
@Test
|
||||
@DisplayName("Fleetd.leadMailboxOpener is the real LeadMailbox::open, not a stub that never connects")
|
||||
void leadMailboxOpenerAttemptsARealConnection() throws Exception {
|
||||
int closedPort;
|
||||
try (ServerSocket socket = new ServerSocket(0)) {
|
||||
closedPort = socket.getLocalPort();
|
||||
} // released immediately — connecting to it now is a guaranteed refusal, not a fluke
|
||||
|
||||
Fleetd.LeadMailboxOpener opener = Fleetd.leadMailboxOpener();
|
||||
|
||||
IllegalStateException thrown = assertThrows(IllegalStateException.class,
|
||||
() -> opener.open("amqp://user:pw@127.0.0.1:" + closedPort + "/coord", "coord-1", 50),
|
||||
"Fleetd.leadMailboxOpener() must be LeadMailbox::open — a real network attempt "
|
||||
+ "against a genuinely unreachable broker must throw. The inert form "
|
||||
+ "(uri, selfCoordId, prefetch) -> null never attempts a connection and "
|
||||
+ "returns null instead of throwing, so it would fail this assertion "
|
||||
+ "silently.");
|
||||
assertTrue(thrown.getMessage().contains("cannot connect to AMQP coordination broker"),
|
||||
"must be LeadMailbox.open's own real failure message, not a different exception "
|
||||
+ "shape standing in for it");
|
||||
}
|
||||
}
|
||||
@@ -1,15 +1,13 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadMailbox;
|
||||
import dev.ltms.fleet.testing.CapturedLog;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
@@ -38,7 +36,7 @@ class FleetdLeadMailboxSelectionTest {
|
||||
boolean unreachable;
|
||||
|
||||
@Override
|
||||
public LeadMailbox open(String uri, String selfCoordId, int prefetch) {
|
||||
public LeadChannelHandle open(String uri, String selfCoordId, int prefetch) {
|
||||
this.offeredUri = uri;
|
||||
this.offeredSelfId = selfCoordId;
|
||||
this.offeredPrefetch = prefetch;
|
||||
@@ -52,29 +50,22 @@ class FleetdLeadMailboxSelectionTest {
|
||||
}
|
||||
}
|
||||
|
||||
private static ListAppender<ILoggingEvent> captureFleetdLogs() {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
return appender;
|
||||
}
|
||||
|
||||
private static String joined(ListAppender<ILoggingEvent> appender, Level level) {
|
||||
return appender.list.stream().filter(e -> e.getLevel() == level)
|
||||
private static String joined(CapturedLog captured, Level level) {
|
||||
return captured.events().stream().filter(e -> e.getLevel() == level)
|
||||
.map(ILoggingEvent::getFormattedMessage).reduce("", (a, b) -> a + "\n" + b);
|
||||
}
|
||||
|
||||
@Test
|
||||
void noCoordinatorBlockLeavesTheFeatureOffSilently() {
|
||||
var appender = captureFleetdLogs();
|
||||
var opener = new RecordingOpener();
|
||||
try (var captured = CapturedLog.of(Fleetd.class)) {
|
||||
var opener = new RecordingOpener();
|
||||
|
||||
assertNull(Fleetd.openLeadMailbox(null, Map.of(), opener));
|
||||
assertNull(Fleetd.openLeadMailbox(null, Map.of(), opener));
|
||||
|
||||
assertNull(opener.offeredUri, "nothing configured means nothing is opened");
|
||||
assertEquals("", joined(appender, Level.WARN),
|
||||
"an opt-in feature nobody asked for must not warn on every boot");
|
||||
assertNull(opener.offeredUri, "nothing configured means nothing is opened");
|
||||
assertEquals("", joined(captured, Level.WARN),
|
||||
"an opt-in feature nobody asked for must not warn on every boot");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -114,31 +105,33 @@ class FleetdLeadMailboxSelectionTest {
|
||||
|
||||
@Test
|
||||
void warnsAndStaysOffWhenSelfIdIsMissing() {
|
||||
var appender = captureFleetdLogs();
|
||||
var opener = new RecordingOpener();
|
||||
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, null, null, null);
|
||||
try (var captured = CapturedLog.of(Fleetd.class)) {
|
||||
var opener = new RecordingOpener();
|
||||
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, null, null, null);
|
||||
|
||||
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener));
|
||||
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener));
|
||||
|
||||
assertNull(opener.offeredUri, "a mailbox with no owning coord-id has no queue to declare");
|
||||
String warns = joined(appender, Level.WARN);
|
||||
assertTrue(warns.contains("coordinator.selfId"), () -> "say which key is missing: " + warns);
|
||||
assertFalse(warns.contains(SECRET), () -> "the URI's password must never be logged: " + warns);
|
||||
assertNull(opener.offeredUri, "a mailbox with no owning coord-id has no queue to declare");
|
||||
String warns = joined(captured, Level.WARN);
|
||||
assertTrue(warns.contains("coordinator.selfId"), () -> "say which key is missing: " + warns);
|
||||
assertFalse(warns.contains(SECRET), () -> "the URI's password must never be logged: " + warns);
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void warnsAndStaysOffWhenTheBrokerIsUnreachableAtBoot() {
|
||||
var appender = captureFleetdLogs();
|
||||
var opener = new RecordingOpener();
|
||||
opener.unreachable = true;
|
||||
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, "mac-opus", null, null);
|
||||
try (var captured = CapturedLog.of(Fleetd.class)) {
|
||||
var opener = new RecordingOpener();
|
||||
opener.unreachable = true;
|
||||
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, "mac-opus", null, null);
|
||||
|
||||
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener),
|
||||
"a down coordination broker turns the feature off; it must never take the daemon down");
|
||||
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener),
|
||||
"a down coordination broker turns the feature off; it must never take the daemon down");
|
||||
|
||||
String warns = joined(appender, Level.WARN);
|
||||
assertTrue(warns.contains("coord.example"), () -> "name the host that failed: " + warns);
|
||||
assertFalse(warns.contains(SECRET), () -> "with credentials stripped: " + warns);
|
||||
assertTrue(warns.contains("Connection refused"), () -> "and the real reason: " + warns);
|
||||
String warns = joined(captured, Level.WARN);
|
||||
assertTrue(warns.contains("coord.example"), () -> "name the host that failed: " + warns);
|
||||
assertFalse(warns.contains(SECRET), () -> "with credentials stripped: " + warns);
|
||||
assertTrue(warns.contains("Connection refused"), () -> "and the real reason: " + warns);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,285 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.lead.LeadRollover;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
import static org.junit.jupiter.api.Assertions.fail;
|
||||
|
||||
/**
|
||||
* fleetd #612 B3 — replaces {@code FleetdLeadRolloverWiringTest} (fleetd #480). That class was a
|
||||
* source-text test scraping {@code Fleetd.java} (now {@code FleetdAssembly.java}, moved there by
|
||||
* fleetd #612 Unit A) with three methods: {@code unrelatedAnchorStillPresent} (a scaffold anchor,
|
||||
* not an independent claim — needs no replacement of its own), {@code
|
||||
* mainStillCallsTheLeadRolloverFactory} (the call-site pin replaced by {@link
|
||||
* #assembledLeadRolloverRunsTheRealClearAndBootstrapSequence}), and {@code
|
||||
* factoryGatesOnConfigPresence} (the absent-config claim replaced by {@link
|
||||
* #absentLeadRolloverConfigMeansNoRolloverIsBuilt} — a claim this ticket found was NOT actually
|
||||
* covered behaviourally anywhere else: {@code LeadRolloverTest}'s only related assertion is
|
||||
* vacuous, {@code assertNull(null)}, and never calls the real factory).
|
||||
*
|
||||
* <p><strong>fleetd #612 B3 correction (ticket comment 17553):</strong> the first version of this
|
||||
* test configured a single shared {@link FakeHerdr} for both the lead and member herdr sockets.
|
||||
* {@code FleetdAssembly.java:140-142} falls back to {@code memberHerdr = herdr} whenever no
|
||||
* distinct {@code memberHerdrSocket} is configured, so with one fake, {@code
|
||||
* router.leadAgents()} and {@code router.memberAgents()} wrapped the identical client — a
|
||||
* mutation swapping {@code Fleetd.leadRollover(cfg, router.leadAgents(), config, leads)} for
|
||||
* {@code ..., router.memberAgents(), ...} at {@code FleetdAssembly.java:408} was therefore
|
||||
* invisible to this test, even though the two are genuinely different daemons in production. This
|
||||
* version configures two distinct sockets and two distinct {@link FakeHerdr} instances (the same
|
||||
* pattern {@code FleetdAssemblyConnectionIdentityTest}, fleetd #612 B2, already uses to separate
|
||||
* lead from member) and asserts the roll's {@code /clear}/bootstrap sends land on the LEAD fake
|
||||
* and never on the MEMBER one.
|
||||
*/
|
||||
class FleetdLeadRolloverAssemblyTest {
|
||||
|
||||
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
|
||||
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
|
||||
|
||||
/** Keys {@code connectHerdr} by socket path so the lead and member daemons can be two
|
||||
* DIFFERENT {@link FakeHerdr}s — same shape as B2's {@code FleetdAssemblyConnectionIdentityTest
|
||||
* .TwoHerdrResourcePorts}. */
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
|
||||
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
|
||||
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
HerdrClient client = herdrsBySocket.get(socketPath);
|
||||
if (client == null) {
|
||||
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
|
||||
}
|
||||
return client;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> replyInbox;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
}
|
||||
}
|
||||
|
||||
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
@Override
|
||||
public void own(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void release(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean ack(String target, String msgId) {
|
||||
return false;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir, Path leadCwd) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: "%s"
|
||||
memberHerdrSocket: "%s"
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
cwd: "%s"
|
||||
leadRollover:
|
||||
handoverPath: handover.md
|
||||
requireOperatorConfirm: false
|
||||
""".formatted(LEAD_SOCKET, MEMBER_SOCKET, leadCwd.toString()));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled LeadRollover runs the full open/confirm/continuation "
|
||||
+ "sequence — /clear, then bootstrapText — through the real herdr router")
|
||||
void assembledLeadRolloverRunsTheRealClearAndBootstrapSequence(@TempDir Path dir) throws Exception {
|
||||
Path leadCwd = dir.resolve("lead-workspace");
|
||||
Files.createDirectories(leadCwd);
|
||||
FleetConfig cfg = writeConfig(dir, leadCwd);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
// Two DISTINCT fakes — one per configured socket — so leadAgents()/memberAgents() wrap
|
||||
// genuinely different clients, exactly like production when memberHerdrSocket is set.
|
||||
FakeHerdr lead = new FakeHerdr();
|
||||
lead.withTab("w2", "w2:t7", "lead: opus");
|
||||
FakeHerdr member = new FakeHerdr();
|
||||
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
|
||||
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
|
||||
LeadRollover rollover = runtime.mcp().leadRollover();
|
||||
assertNotNull(rollover, "leadRollover: is present in this test's config, so "
|
||||
+ "FleetdAssembly.assembleAndStart must have built a real LeadRollover through the "
|
||||
+ "Fleetd.leadRollover(...) call site — a mutation to `LeadRollover leadRollover = "
|
||||
+ "null;` at that call site can never pass this");
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open("term_a", "fleetd #612 B3 test");
|
||||
String expectedHandoverPath = leadCwd.resolve("handover.md").normalize().toString();
|
||||
assertEquals(expectedHandoverPath, pending.handoverPath());
|
||||
|
||||
// Ensure the handover file's mtime lands strictly AFTER open()'s requestedAtMillis —
|
||||
// LeadRollover.checkHandover refuses on mtime <= requestedAt (HANDOVER_STALE).
|
||||
Thread.sleep(50);
|
||||
Files.writeString(Path.of(pending.handoverPath()), "handover content for fleetd #612 B3");
|
||||
|
||||
LeadRollover.RollDecision decision = rollover.confirm("term_a", pending.token(), true);
|
||||
assertTrue(decision.accepted(), "confirm() must approve: requireOperatorConfirm is false, "
|
||||
+ "the caller terminal matches open()'s, and the handover file exists, is non-empty "
|
||||
+ "and fresh — got: " + decision);
|
||||
|
||||
// The production LeadRollover constructor always runs the post-confirm continuation on a
|
||||
// real virtual thread (see Fleetd.leadRollover, which never passes the package-private test
|
||||
// constructor), so this polls the real FleetMcp.leadRollover() instance's status(token)
|
||||
// until the real continuation finishes.
|
||||
LeadRollover.RollStatus status = pollUntilTerminal(rollover, pending.token());
|
||||
assertEquals(LeadRollover.RollState.ROLLED, status.state(),
|
||||
"the full happy path must complete: FakeHerdr's default agent status is 'idle', so "
|
||||
+ "the turn-boundary wait settles immediately and the post-/clear wait "
|
||||
+ "releases via its pickup-grace path — detail: " + status.detail());
|
||||
|
||||
// Prove the real herdr router actually sent BOTH messages, in order, to the real LEAD
|
||||
// pane — this is the one thing a source-text pin on the call site could never show.
|
||||
List<FakeHerdr.Call> prompts = lead.calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt"))
|
||||
.toList();
|
||||
assertTrue(prompts.size() >= 2, "expected at least a /clear send and a bootstrapText send "
|
||||
+ "on the LEAD daemon, got " + prompts.size() + " agent.prompt calls: " + prompts);
|
||||
assertEquals("/clear", ((Map<String, Object>) prompts.get(0).params()).get("text"),
|
||||
"the first send must be the literal /clear housekeeping command");
|
||||
Object secondText = ((Map<String, Object>) prompts.get(1).params()).get("text");
|
||||
assertTrue(secondText instanceof String && ((String) secondText).contains(expectedHandoverPath),
|
||||
"the second send must be the default bootstrapText naming the resolved handover "
|
||||
+ "path, got: " + secondText);
|
||||
|
||||
// fleetd #612 B3 correction: prove the roll never touches the MEMBER daemon. A mutation
|
||||
// swapping router.leadAgents() for router.memberAgents() at the real call site would move
|
||||
// both sends above onto `member` instead, which this assertion catches — the thing the
|
||||
// single-fake version of this test could never see, because both wrapped the same client.
|
||||
List<FakeHerdr.Call> memberPrompts = member.calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt"))
|
||||
.toList();
|
||||
assertTrue(memberPrompts.isEmpty(), "the roll must be wired to the LEAD daemon only — got "
|
||||
+ memberPrompts.size() + " agent.prompt call(s) on the MEMBER daemon instead: "
|
||||
+ memberPrompts);
|
||||
}
|
||||
|
||||
private static LeadRollover.RollStatus pollUntilTerminal(LeadRollover rollover, String token)
|
||||
throws InterruptedException {
|
||||
long deadline = System.nanoTime() + java.util.concurrent.TimeUnit.SECONDS.toNanos(10);
|
||||
while (System.nanoTime() < deadline) {
|
||||
LeadRollover.RollStatus status = rollover.status(token);
|
||||
if (status.state() != LeadRollover.RollState.PENDING
|
||||
&& status.state() != LeadRollover.RollState.IN_PROGRESS) {
|
||||
return status;
|
||||
}
|
||||
Thread.sleep(50);
|
||||
}
|
||||
fail("the real continuation did not reach a terminal state within 10s — last status: "
|
||||
+ rollover.status(token));
|
||||
throw new AssertionError("unreachable");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] Fleetd.leadRollover(...) returns null when leadRollover: is absent "
|
||||
+ "from config — the opt-in gate FleetdLeadRolloverWiringTest's "
|
||||
+ "factoryGatesOnConfigPresence pinned by source text alone")
|
||||
void absentLeadRolloverConfigMeansNoRolloverIsBuilt(@TempDir Path dir) throws Exception {
|
||||
Path yaml = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
""");
|
||||
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
|
||||
AgentControl agents = new AgentControl(new FakeHerdr());
|
||||
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, config, Map::of);
|
||||
|
||||
assertNull(rollover, "leadRollover: is absent from this config, so the factory's opt-in "
|
||||
+ "gate (`if (cfg.leadRollover() == null) return null;`) must fire and no "
|
||||
+ "LeadRollover must be constructed at all");
|
||||
}
|
||||
}
|
||||
@@ -1,89 +0,0 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #480 Unit A, hard requirement 6: pin {@code Fleetd.main}'s construction of {@link
|
||||
* dev.ltms.fleet.lead.LeadRollover} with a source-text assertion, mirroring {@code
|
||||
* FleetdCompletionResolverWiringTest}'s pattern — five log-only reporters in {@code Fleetd.main}
|
||||
* already survived mutation batteries this exact way (fleetd #415's extraction antidote note).
|
||||
*
|
||||
* <p>What this class still covers, and what it never claimed to. {@code LeadRolloverTest}
|
||||
* constructs its own {@code LeadRollover} directly (as every prior test of an extracted factory
|
||||
* does) with a hand-built lookup, so a mutation that deletes the {@code leadRollover(...)} call
|
||||
* from {@code main} — or replaces one of its arguments with something that still compiles, e.g.
|
||||
* {@code router.leadAgents()} swapped for {@code null}, or the whole assignment swapped for a bare
|
||||
* {@code null} literal — leaves every behavioural test green. This is a plain string read, guarded
|
||||
* by an unrelated anchor assertion so a broken or empty file read cannot pass as a real change.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
||||
* LeadRollover} and never runs {@code main}. It pins the {@code leadRollover(...)} CALL SITE's
|
||||
* argument list — that {@code main} still passes {@code leads} at all — never what the factory
|
||||
* DOES with that argument once inside its own body.
|
||||
*
|
||||
* <p><b>Correction (fleetd #480 relative-handover-path follow-up): that gap used to be real, and
|
||||
* now is not — but not here.</b> This class's javadoc previously claimed "no behavioural test can
|
||||
* catch this wiring dropping out" for the whole factory, including the lambda {@code
|
||||
* leadRollover(...)} builds internally (terminal → lead name → {@code Leader.cwd()}). That claim
|
||||
* was proven true at the time — mutating that lambda's body to {@code String leadName = null;}
|
||||
* (always "no lead found", which silently reintroduces the daemon-cwd bug this ticket fixes) left
|
||||
* the full suite green, {@code Tests run: 1669, Failures: 0}. It is no longer true: {@code
|
||||
* FleetdLeadRolloverWorkspaceLookupTest} now calls {@code Fleetd.leadRollover(...)} directly with a
|
||||
* real {@link dev.ltms.fleet.config.ConfigRef} built from a temp {@code fleetd.yaml}, and fails
|
||||
* against that exact one-line mutation. So: THIS class still covers only the call site's argument
|
||||
* list; {@code FleetdLeadRolloverWorkspaceLookupTest} is what now covers the lambda's body. Neither
|
||||
* one subsumes the other — keep both.
|
||||
*/
|
||||
class FleetdLeadRolloverWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] unrelated anchor: Fleetd.java still declares the Fleetd class")
|
||||
void unrelatedAnchorStillPresent() throws Exception {
|
||||
// Guards the two assertions below: without this, a bad read (empty string, wrong file,
|
||||
// truncated file) could vacuously fail to contain the leadRollover(...) call too, and a
|
||||
// test that only asserts "contains X" would report a false pass for the wrong reason if X
|
||||
// happened to match. Asserting an unrelated, structurally distant string first proves the
|
||||
// read actually pulled real file content.
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains("public final class Fleetd"),
|
||||
"sanity anchor failed — the file read did not return real Fleetd.java source; the "
|
||||
+ "leadRollover(...) wiring assertions below cannot be trusted until this passes");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] main still constructs LeadRollover via the leadRollover(...) factory, exactly as heartbeat is constructed")
|
||||
void mainStillCallsTheLeadRolloverFactory() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"LeadRollover leadRollover = leadRollover(cfg, router.leadAgents(), config, leads);"),
|
||||
"Fleetd.main must still assign `LeadRollover leadRollover = leadRollover(cfg, "
|
||||
+ "router.leadAgents(), config, leads);`. Dropping this call, or swapping one of "
|
||||
+ "its arguments for something that still compiles (e.g. null in place of "
|
||||
+ "router.leadAgents()), leaves every behavioural test green — this source check is "
|
||||
+ "what must go red instead. fleetd #480 correction 2 deliberately dropped "
|
||||
+ "primaryRegistry from this call — see LeadRollover's class javadoc for why a "
|
||||
+ "single-slot lookup was wrong here. The fleetd #480 relative-handover-path "
|
||||
+ "follow-up added `leads` (terminal → lead name) so the factory can resolve a "
|
||||
+ "relative handoverPath against the calling lead's own workspace.");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] the leadRollover(...) factory itself gates construction on cfg.leadRollover() != null")
|
||||
void factoryGatesOnConfigPresence() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains("if (cfg.leadRollover() == null) {"),
|
||||
"Fleetd.leadRollover(...) must refuse to construct a LeadRollover when the "
|
||||
+ "leadRollover: block is absent — an upgraded daemon must never silently acquire "
|
||||
+ "the ability to clear the lead's own pane. See LeadHeartbeatLoop's construction "
|
||||
+ "gate (cfg.leadHeartbeat() != null) for the pattern this mirrors.");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,172 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* fleetd #612 B3 — replaces {@code FleetdLeadSeatWiringTest} (fleetd #176), a source-text test that
|
||||
* scraped {@code Fleetd.java} (now {@code FleetdAssembly.java}, moved there by fleetd #612 Unit A)
|
||||
* for the exact {@code new FleetMcp.LeadSeatSource(Fleetd.leadSeatLookup(...))} constructor-call
|
||||
* text. That proves the right symbols appear in source; it proves nothing about what the daemon's
|
||||
* live {@code fleet_list} actually reports.
|
||||
*
|
||||
* <p>This test instead drives the REAL {@link FleetMcp.LeadSeatSource} the real {@link
|
||||
* FleetdAssembly#assembleAndStart} builds — including the REAL {@code LeadTabScanner} it wires
|
||||
* {@code Fleetd.leadSeatLookup} through — reached via {@link FleetMcp#leadSeatSource()} on the
|
||||
* live {@code FleetMcp} {@code FleetdRuntime} owns. It seeds one FakeHerdr tab labelled to match a
|
||||
* configured {@code fleet.leaders.opus.tab}, with a live agent already in it (FakeHerdr's own
|
||||
* default {@code agent.list}/{@code pane.list} entries for {@code term_a}/{@code w2:p7}/{@code
|
||||
* w2:t7} — no FakeHerdr change needed), and asserts the assembled seat source reports exactly the
|
||||
* seat {@link FleetMcp.LeadSeatSource#none()} (the inert stand-in) could never produce: 1, not 0.
|
||||
*/
|
||||
class FleetdLeadSeatAssemblyTest {
|
||||
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> replyInbox;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
}
|
||||
}
|
||||
|
||||
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
@Override
|
||||
public void own(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void release(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean ack(String target, String msgId) {
|
||||
return false;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
profile: sonnet
|
||||
profiles:
|
||||
sonnet:
|
||||
subscription: true
|
||||
argv: ["ccs", "sonnet"]
|
||||
""");
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled LeadSeatSource, backed by the real LeadTabScanner, "
|
||||
+ "reports a live lead's seat against its own subscription profile")
|
||||
void assembledLeadSeatSourceReportsALiveLeadsSeat(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
// Label FakeHerdr's own default pane's tab (term_a / w2:p7 / w2:t7, already carrying a live
|
||||
// agent) to match fleet.leaders.opus.tab exactly — no FakeHerdr change needed at all.
|
||||
ports.herdr.withTab("w2", "w2:t7", "lead: opus");
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
|
||||
FleetMcp.LeadSeatSource seatSource = runtime.mcp().leadSeatSource();
|
||||
assertEquals(1, seatSource.seatsFor().apply("sonnet"),
|
||||
"the real LeadTabScanner recognises the labelled tab as a live 'opus' lead on "
|
||||
+ "profile 'sonnet' (same credential, matched by Fleetd.leadSeatLookup), so "
|
||||
+ "subscription profile 'sonnet' must be charged one seat — "
|
||||
+ "FleetMcp.LeadSeatSource.none() (the inert stand-in this test's mutation "
|
||||
+ "swaps the call site for) always reports 0, whatever the input");
|
||||
|
||||
// A profile no lead is running on gets no seat charged — the same seat source, applied to
|
||||
// an input that must stay at the inert answer even on the real, non-inert instance.
|
||||
assertEquals(0, seatSource.seatsFor().apply("no-such-profile"));
|
||||
}
|
||||
}
|
||||
@@ -1,43 +0,0 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #176: {@code Fleetd.main} builds its {@code FleetMcp} from a 14-argument constructor whose
|
||||
* last argument is a {@code FleetMcp.LeadSeatSource} wrapping {@link Fleetd#leadSeatLookup}. That
|
||||
* argument is exactly the kind of wiring fleetd #248 warned about: dropping it (or swapping it for
|
||||
* the inert {@code FleetMcp.LeadSeatSource.none()}) compiles with 0 errors and leaves every test
|
||||
* that builds its own {@code FleetMcp}/{@code CapacitySource} directly — every test that predates
|
||||
* this ticket — green, because none of them go through {@code main} at all.
|
||||
*
|
||||
* <p>{@link FleetdLeadSeatLookupTest} proves the factory's own matching logic; this class is the
|
||||
* plain source-text assertion that proves {@code main} still passes its result in, mirroring
|
||||
* {@code FleetdCompletionResolverWiringTest}'s approach for the same class of gap.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a
|
||||
* {@code FleetMcp} and never runs {@code main}.
|
||||
*/
|
||||
class FleetdLeadSeatWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] FleetMcp's construction call still passes a LeadSeatSource built from leadSeatLookup(...)")
|
||||
void fleetMcpConstructionStillWiresLeadSeatLookup() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains("new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), "
|
||||
+ "leaders, leads))"),
|
||||
"FleetMcp's construction call must still pass a LeadSeatSource built from "
|
||||
+ "Fleetd.leadSeatLookup(...). Dropping it or swapping in "
|
||||
+ "FleetMcp.LeadSeatSource.none() (fleetd #176's would-be silent regression, the same "
|
||||
+ "shape as fleetd #248's measured mutations) compiles with 0 errors and leaves every "
|
||||
+ "existing behavioural test green — this source check is what must go red instead.");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,72 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #589 Group 1: {@link Fleetd#liveExhaustedPatterns} is the factory that replaced {@code
|
||||
* main}'s inline {@code new LiveExhaustedPatterns(() -> config.get().profiles())}. Before this
|
||||
* ticket, that supplier argument was untestable wiring: replacing it with a hardcoded {@code () ->
|
||||
* Map.of()} compiled with 0 errors and left every existing test green, because {@code
|
||||
* LiveExhaustedPatternsTest} builds its own instance directly with a hand-supplied map and never
|
||||
* goes through {@code main}'s call site.
|
||||
*
|
||||
* <p>Silently losing this wiring means every profile's {@code exhaustedPattern} stops being
|
||||
* recognised — {@link Fleetd#exhaustedPatternLookup} would never see a match, and a genuine
|
||||
* usage-limit refusal would be handed back to a waiting {@code fleet_send} as real completed work.
|
||||
*/
|
||||
class FleetdLiveExhaustedPatternsWiringTest {
|
||||
|
||||
private static final String YAML = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: claude-opus-5
|
||||
exhaustedPattern: "usage limit"
|
||||
gx:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
""";
|
||||
|
||||
private static ConfigRef loadConfig(Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, YAML);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
return ConfigRef.fixed(cfg);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a profile with a configured exhaustedPattern is armed, with a compiled matcher")
|
||||
void configuredProfileIsArmed(@TempDir Path dir) throws Exception {
|
||||
LiveExhaustedPatterns patterns = Fleetd.liveExhaustedPatterns(loadConfig(dir));
|
||||
|
||||
assertTrue(patterns.armed("terra"),
|
||||
"the config's live profiles() supplier must reach LiveExhaustedPatterns — hardcoding "
|
||||
+ "the supplier to () -> Map.of() at the Fleetd.liveExhaustedPatterns call "
|
||||
+ "site must fail this assertion");
|
||||
assertTrue(patterns.patternFor("terra").matcher("the usage limit has been reached").find());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a profile with no configured exhaustedPattern is not armed, but is still resolvable")
|
||||
void unconfiguredProfileIsNotArmed(@TempDir Path dir) throws Exception {
|
||||
LiveExhaustedPatterns patterns = Fleetd.liveExhaustedPatterns(loadConfig(dir));
|
||||
|
||||
assertFalse(patterns.armed("gx"), "'gx' has no exhaustedPattern configured");
|
||||
assertNull(patterns.patternFor("gx"));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,174 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.inject.LoopWatchdog;
|
||||
import dev.ltms.fleet.inject.StatusPoller;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import dev.ltms.fleet.peer.Capability;
|
||||
import dev.ltms.fleet.peer.PeerHandle;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import dev.ltms.fleet.session.SessionReaper;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* fleetd #562 follow-up (issue comment "HOLD on PR #579"): {@code Fleetd.main}'s {@code loopHealth}
|
||||
* local used to be a bare {@code new FleetMcp.LoopHealthSource(poller::health, ...)} built inline,
|
||||
* with nothing a test could call directly. Measured on that shape: replacing {@code
|
||||
* poller::health} with a constant {@code () -> LoopWatchdog.State.RUNNING} at the call site
|
||||
* compiled with 0 errors and left all 1771 existing tests green — the daemon could be changed to
|
||||
* always report the {@link StatusPoller} as {@code RUNNING}, so the watchdog could never fire and
|
||||
* a stalled poller would be invisible, while every test stayed green. That is exactly the false
|
||||
* negative this ticket exists to prevent.
|
||||
*
|
||||
* <p>The five tests PR #579 added ({@code FleetMcpTest}, {@code FleetAppTest}) all build their own
|
||||
* {@link FleetMcp.LoopHealthSource} directly with fixed lambdas — they prove the seam ({@code
|
||||
* LoopHealthSource} reports what it is given) and nothing about what {@code Fleetd.main} actually
|
||||
* gives it. This is the same hand-built-vs-config-wired shape as fleetd #561/#248/#426.
|
||||
*
|
||||
* <p>The fix extracts the inline {@code new} into {@link Fleetd#loopHealthSource}, a package-private
|
||||
* factory in the same style as {@link Fleetd#capacitySource} and {@link Fleetd#healthCoverageSource}
|
||||
* — which is exactly what makes it directly callable here. This test calls that factory with real
|
||||
* {@link StatusPoller}/{@link SessionReaper} instances (never started, so no herdr or git I/O
|
||||
* happens) and pins each half separately, plus the {@code reaper == null} branch: one invariant
|
||||
* wired at three places needs three assertions, not one combined check whose non-zero total could
|
||||
* hide a gap at any single place.
|
||||
*/
|
||||
class FleetdLoopHealthSourceWiringTest {
|
||||
|
||||
@Test
|
||||
@DisplayName("the statusPoller half reports the real poller's health, not a hardcoded state")
|
||||
void statusPollerHalfReflectsThePollersRealHealth() {
|
||||
// Stopped without ever being started — stop() still marks the watchdog STOPPED. A poller
|
||||
// that has never reported RUNNING is the discriminating case: if Fleetd.loopHealthSource
|
||||
// ever hardcoded RUNNING (the exact mutation this test exists to catch), this would fail.
|
||||
StatusPoller stoppedPoller = freshPoller();
|
||||
stoppedPoller.stop();
|
||||
SessionReaper unusedReaper = freshReaper(); // present only to satisfy the signature
|
||||
|
||||
FleetMcp.LoopHealthSource source = Fleetd.loopHealthSource(stoppedPoller, unusedReaper);
|
||||
|
||||
assertEquals(LoopWatchdog.State.STOPPED, source.statusPoller().get(),
|
||||
"the statusPoller supplier must delegate to the real poller's health() — "
|
||||
+ "replacing poller::health with a constant () -> RUNNING at the "
|
||||
+ "Fleetd.loopHealthSource call site must fail this assertion");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("the sessionReaper half reports the real reaper's health, not a hardcoded state")
|
||||
void sessionReaperHalfReflectsTheReapersRealHealth() {
|
||||
StatusPoller unusedPoller = freshPoller(); // present only to satisfy the signature
|
||||
SessionReaper stoppedReaper = freshReaper();
|
||||
stoppedReaper.stop();
|
||||
|
||||
FleetMcp.LoopHealthSource source = Fleetd.loopHealthSource(unusedPoller, stoppedReaper);
|
||||
|
||||
assertEquals(LoopWatchdog.State.STOPPED, source.sessionReaper().get(),
|
||||
"the sessionReaper supplier must delegate to the real reaper's health() — "
|
||||
+ "replacing reaper.health() with a constant at the "
|
||||
+ "Fleetd.loopHealthSource call site must fail this assertion");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a null reaper (idle ttl not configured) still reports STOPPED, not a crash")
|
||||
void nullReaperStillReportsStopped() {
|
||||
// SessionReaper is only constructed when lifecycle.idleTtlSeconds is configured (see the
|
||||
// `reaper` local in Fleetd.main) — a real deployment routinely passes null here. That null
|
||||
// check is real behaviour, not a simplification to delete: it must keep reporting STOPPED
|
||||
// rather than throwing a NullPointerException on the first fleet_list/healthz call.
|
||||
StatusPoller runningPoller = freshPoller();
|
||||
|
||||
FleetMcp.LoopHealthSource source = Fleetd.loopHealthSource(runningPoller, null);
|
||||
|
||||
assertEquals(LoopWatchdog.State.STOPPED, source.sessionReaper().get(),
|
||||
"reaper == null must still report STOPPED, exactly like an intentionally-stopped "
|
||||
+ "reaper would — do not delete this null check to simplify the wiring");
|
||||
}
|
||||
|
||||
/** Never started, so no herdr call is ever made; freshly constructed reports RUNNING. */
|
||||
private static StatusPoller freshPoller() {
|
||||
AgentControl agents = new AgentControl(new FakeHerdr());
|
||||
return new StatusPoller(agents, new Injector(agents), 1000);
|
||||
}
|
||||
|
||||
/** Never started, so no git/session I/O is ever made; freshly constructed reports RUNNING. */
|
||||
private static SessionReaper freshReaper() {
|
||||
return new SessionReaper(new SessionManager(new NeverSpawnsLauncher()), 60, 1000);
|
||||
}
|
||||
|
||||
/**
|
||||
* Same minimal shape as {@code FleetdBackendErrorSinkTest.NeverSpawnsLauncher} — every method
|
||||
* throws or returns an empty/no-op value, since a {@link SessionReaper} that is only ever
|
||||
* constructed and then stopped (never started) never calls any of them.
|
||||
*/
|
||||
private static final class NeverSpawnsLauncher implements PeerLauncher {
|
||||
@Override
|
||||
public Set<Capability> capabilities() {
|
||||
return Set.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<Capability> capabilitiesFor(String profileName) {
|
||||
return Set.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req) {
|
||||
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return Set.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public String defaultProfile() {
|
||||
return null;
|
||||
}
|
||||
|
||||
@Override
|
||||
public String effectiveCwd(SpawnRequest req) {
|
||||
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<String> parityOverlay(String profileName) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<?> list() {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public int reapOrphanWorkers() {
|
||||
return 0;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void stop(String id) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean clearContext(String id) {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,77 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||
import dev.ltms.fleet.member.OpenCodeLauncher;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
|
||||
/**
|
||||
* fleetd #589 Group 2: {@link Fleetd#openCodeLauncher} is the factory that replaced {@code main}'s
|
||||
* inline {@code new OpenCodeLauncher(...)} call — the {@code opencode} counterpart to {@link
|
||||
* Fleetd#claudeCodeLauncher}, extracted for the identical reason. Its {@code memberCredentials}
|
||||
* argument is the same {@code () -> config.get().memberCredentials()} supplier; replacing it with
|
||||
* {@code () -> null} compiled with 0 errors and left every existing test green before this ticket,
|
||||
* reopening the same CB-592 exposure gap CB-596's policy closed.
|
||||
*
|
||||
* <p>Same observable surface as {@code OpenCodeLauncherTest}'s own {@code memberCredentials} tests:
|
||||
* spawn through the launcher {@code main} actually wires and inspect what {@code tab.create}
|
||||
* carried.
|
||||
*/
|
||||
class FleetdOpenCodeLauncherCredentialWiringTest {
|
||||
|
||||
private static final String YAML = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
gemini:
|
||||
kind: opencode
|
||||
model: google/gemini-2.5-pro
|
||||
memberCredentials:
|
||||
policy: deny-by-default
|
||||
known:
|
||||
- GITEA_ACCESS_TOKEN
|
||||
""";
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static Map<String, String> startEnv(FakeHerdr herdr) {
|
||||
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("main's memberCredentials wiring reaches OpenCodeLauncher: a known-but-not-allowed "
|
||||
+ "name is shadowed on spawn")
|
||||
void memberCredentialsWiringReachesOpenCodeLauncher(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, YAML);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = ConfigRef.fixed(cfg);
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
OpenCodeLauncher launcher = Fleetd.openCodeLauncher(new AgentControl(herdr),
|
||||
new WorkspaceControl(herdr), cfg.profiles(), cfg, config, ExhaustionSink.none());
|
||||
launcher.spawn();
|
||||
|
||||
String shadowed = startEnv(herdr).get("GITEA_ACCESS_TOKEN");
|
||||
assertNotNull(shadowed,
|
||||
"GITEA_ACCESS_TOKEN is 'known' but not 'allow'-ed in the loaded config — it must be "
|
||||
+ "explicitly shadowed on spawn; replacing the memberCredentials supplier with "
|
||||
+ "() -> null at the Fleetd.openCodeLauncher call site must fail this "
|
||||
+ "assertion, since a null policy shadows nothing");
|
||||
assertFalse(shadowed.isBlank(), "the overlay value must be non-blank");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,107 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.mcp.PrimaryRegistry;
|
||||
import dev.ltms.fleet.msg.InMemoryReplyInbox;
|
||||
import dev.ltms.fleet.msg.MessageService;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.function.Consumer;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #589 Group 3, site {@code :611-631}. {@code Fleetd.main} wires {@code
|
||||
* sessions.onRelease(...)} with a lambda that calls three collaborators — {@code
|
||||
* messages.abandon}, {@code replyInbox.release}, and {@code primaryRegistry.forgetDelegation} —
|
||||
* before this ticket built inline inside {@code main}. Measured: replacing the whole lambda body
|
||||
* with {@code detail -> { }} compiles with 0 errors and leaves the full suite green, because each
|
||||
* collaborator is separately tested in isolation ({@code MessageServiceTest}, {@code
|
||||
* PrimaryRegistryTest}) but nothing before this ticket drove the lambda that calls all three from
|
||||
* {@code main}.
|
||||
*
|
||||
* <p>{@code MessageService.abandon}'s own javadoc documents the consequence: without this,
|
||||
* tearing a worker down leaves its rendezvous waiter open, so a blocking {@code fleet_send} keeps
|
||||
* blocking and an async one reports {@code PENDING} for a hardcoded thirty minutes on every {@code
|
||||
* fleet_stop} and every idle-reap.
|
||||
*
|
||||
* <p>This test calls {@link Fleetd#releaseCleanup} directly — never {@code SessionManager} or
|
||||
* {@code main} — against real {@link MessageService}, {@link InMemoryReplyInbox}, and {@link
|
||||
* PrimaryRegistry} instances, and asserts each collaborator's own observable effect: the pending
|
||||
* async ticket transitions to {@code FAILED} (abandon), the inbox no longer owns the target's
|
||||
* queue (release), and the recorded delegation is forgotten (forgetDelegation).
|
||||
*/
|
||||
class FleetdReleaseCleanupWiringTest {
|
||||
|
||||
private static final String T = "term_a";
|
||||
|
||||
@Test
|
||||
@DisplayName("Fleetd.releaseCleanup reaches messages.abandon, replyInbox.release, and primaryRegistry.forgetDelegation")
|
||||
void releaseCleanupReachesAllThreeCollaborators() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("$ prompt");
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
Injector injector = new Injector(agents);
|
||||
MessageService messages = new MessageService(agents, injector, rendezvous);
|
||||
InMemoryReplyInbox replyInbox = new InMemoryReplyInbox();
|
||||
PrimaryRegistry primaryRegistry = new PrimaryRegistry(null);
|
||||
|
||||
// Set up the "before" state each collaborator's own effect is measured against.
|
||||
String ticket = messages.sendAsync(T, "long task");
|
||||
awaitWaiting(rendezvous);
|
||||
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase(),
|
||||
"sanity: the async ticket is pending before cleanup runs");
|
||||
|
||||
replyInbox.own(T);
|
||||
replyInbox.publish(T, "msg-1", "hello");
|
||||
assertEquals(1, replyInbox.peek(T).size(),
|
||||
"sanity: the inbox owns T and holds one message before cleanup runs");
|
||||
|
||||
primaryRegistry.recordDelegation(T, "lead-1");
|
||||
assertEquals("lead-1", primaryRegistry.nudgeTargetFor(T).orElse(null),
|
||||
"sanity: the delegation is recorded before cleanup runs");
|
||||
|
||||
Consumer<SessionManager.ReleaseDetail> cleanup =
|
||||
Fleetd.releaseCleanup(messages, replyInbox, primaryRegistry);
|
||||
cleanup.accept(new SessionManager.ReleaseDetail(T, null, null, null, null));
|
||||
|
||||
MessageService.TaskView view = null;
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (System.currentTimeMillis() < deadline) {
|
||||
view = messages.poll(ticket);
|
||||
if (view.phase() != MessageService.Phase.PENDING) {
|
||||
break;
|
||||
}
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
}
|
||||
assertEquals(MessageService.Phase.FAILED, view == null ? null : view.phase(),
|
||||
"releaseCleanup must call messages.abandon(...) — an inert detail -> { } lambda "
|
||||
+ "leaves this ticket PENDING forever");
|
||||
assertTrue(view.detail() != null && view.detail().contains("released"),
|
||||
"the abandon reason must say the worker session was released");
|
||||
|
||||
assertTrue(replyInbox.peek(T).isEmpty(),
|
||||
"releaseCleanup must call replyInbox.release(...) — an inert lambda leaves the "
|
||||
+ "inbox still owning T with its message");
|
||||
|
||||
assertTrue(primaryRegistry.nudgeTargetFor(T).isEmpty(),
|
||||
"releaseCleanup must call primaryRegistry.forgetDelegation(...) — an inert lambda "
|
||||
+ "leaves the stale delegation in place");
|
||||
}
|
||||
|
||||
private static void awaitWaiting(Rendezvous rendezvous) throws InterruptedException {
|
||||
long deadline = System.currentTimeMillis() + 2000;
|
||||
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting(T), "send should have opened its rendezvous waiter");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,49 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.net.ServerSocket;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #589 Group 3, site {@code :512}. {@code Fleetd.main} passes {@code AmqpReplyInbox::open}
|
||||
* as the {@link Fleetd.AmqpOpener} argument to {@code selectReplyInbox(...)} — before this ticket
|
||||
* that method reference was inline at the call site. Measured: replacing it with the inert {@code
|
||||
* (uri, prefetch) -> new InMemoryReplyInbox()} compiles with 0 errors and leaves the full suite
|
||||
* green — {@code FleetdReplyInboxSelectionTest} drives {@code selectReplyInbox} with its own
|
||||
* injected opener (including one test that passes the real {@code AmqpReplyInbox::open}
|
||||
* explicitly) and never observes what {@code main} itself actually passes.
|
||||
*
|
||||
* <p>This test calls {@link Fleetd#replyInboxOpener} directly and proves it is the real, network-
|
||||
* attempting opener rather than a stub: pointed at a guaranteed-closed local port, it must throw —
|
||||
* the same shape {@code FleetdReplyInboxSelectionTest.aRealUnreachableBrokerFallsBackViaTheRealOpener}
|
||||
* already relies on for {@code AmqpReplyInbox.open} itself. The inert form never attempts a
|
||||
* connection and never throws, so it fails this assertion silently (by returning normally).
|
||||
*/
|
||||
class FleetdReplyInboxOpenerWiringTest {
|
||||
|
||||
@Test
|
||||
@DisplayName("Fleetd.replyInboxOpener is the real AmqpReplyInbox::open, not a stub that never connects")
|
||||
void replyInboxOpenerAttemptsARealConnection() throws Exception {
|
||||
int closedPort;
|
||||
try (ServerSocket socket = new ServerSocket(0)) {
|
||||
closedPort = socket.getLocalPort();
|
||||
} // released immediately — connecting to it now is a guaranteed refusal, not a fluke
|
||||
|
||||
Fleetd.AmqpOpener opener = Fleetd.replyInboxOpener();
|
||||
|
||||
IllegalStateException thrown = assertThrows(IllegalStateException.class,
|
||||
() -> opener.open("amqp://user:pw@127.0.0.1:" + closedPort + "/vh", 50),
|
||||
"Fleetd.replyInboxOpener() must be AmqpReplyInbox::open — a real network attempt "
|
||||
+ "against a genuinely unreachable broker must throw. The inert form "
|
||||
+ "(uri, prefetch) -> new InMemoryReplyInbox() never attempts a connection "
|
||||
+ "and never throws, so it would return normally here and fail this "
|
||||
+ "assertion silently.");
|
||||
assertTrue(thrown.getMessage().contains("cannot connect to AMQP broker"),
|
||||
"must be AmqpReplyInbox.open's own real failure message, not a different exception "
|
||||
+ "shape standing in for it");
|
||||
}
|
||||
}
|
||||
@@ -1,15 +1,13 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.msg.AmqpReplyInbox;
|
||||
import dev.ltms.fleet.msg.InMemoryReplyInbox;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import dev.ltms.fleet.testing.CapturedLog;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.net.ServerSocket;
|
||||
import java.util.List;
|
||||
@@ -50,43 +48,34 @@ class FleetdReplyInboxSelectionTest {
|
||||
}
|
||||
}
|
||||
|
||||
private static ListAppender<ILoggingEvent> attach() {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
private static CapturedLog attach() {
|
||||
// logback-test.xml pins dev.ltms.fleet to WARN; raise it so INFO selection lines are captured.
|
||||
logger.setLevel(Level.INFO);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
return appender;
|
||||
return CapturedLog.at(Fleetd.class, Level.INFO);
|
||||
}
|
||||
|
||||
private static void detach(ListAppender<ILoggingEvent> appender) {
|
||||
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
|
||||
}
|
||||
|
||||
private static void assertNoLogContains(ListAppender<ILoggingEvent> appender, String secret) {
|
||||
assertTrue(appender.list.stream().noneMatch(e -> e.getFormattedMessage().contains(secret)),
|
||||
private static void assertNoLogContains(List<ILoggingEvent> events, String secret) {
|
||||
assertTrue(events.stream().noneMatch(e -> e.getFormattedMessage().contains(secret)),
|
||||
"no log line may contain the resolved URI's password");
|
||||
}
|
||||
|
||||
@Test
|
||||
void uriEnvSetAndPresentSelectsAmqpWithTheResolvedUri() {
|
||||
FleetConfig.Broker broker = new FleetConfig.Broker(null, "LAVINMQ_URI", null);
|
||||
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), false, (opener, appender, inbox) -> {
|
||||
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), false, (opener, events, inbox) -> {
|
||||
assertEquals(RESOLVED_URI, opener.offeredUri,
|
||||
"the daemon must connect with the value resolved from uriEnv — selection, not just parse");
|
||||
assertEquals(opener.inbox, inbox, "the AMQP opener's inbox is what is selected");
|
||||
assertNoLogContains(appender, SECRET);
|
||||
assertNoLogContains(events, SECRET);
|
||||
});
|
||||
}
|
||||
|
||||
@Test
|
||||
void uriEnvSetButVariableMissingFallsBackToInMemoryAndWarns() {
|
||||
FleetConfig.Broker broker = new FleetConfig.Broker(null, "LAVINMQ_URI", null);
|
||||
recording(broker, Map.of(), false, (opener, appender, inbox) -> {
|
||||
recording(broker, Map.of(), false, (opener, events, inbox) -> {
|
||||
assertInstanceOf(InMemoryReplyInbox.class, inbox);
|
||||
assertNull(opener.offeredUri, "AMQP must never be attempted when the variable is missing");
|
||||
assertTrue(hasWarnContaining(appender, "LAVINMQ_URI") && hasWarnContaining(appender, "DISABLED"),
|
||||
assertTrue(hasWarnContaining(events, "LAVINMQ_URI") && hasWarnContaining(events, "DISABLED"),
|
||||
"a missing uriEnv variable must warn loudly, not fail silently");
|
||||
});
|
||||
}
|
||||
@@ -94,10 +83,10 @@ class FleetdReplyInboxSelectionTest {
|
||||
@Test
|
||||
void uriEnvSetButVariableBlankFallsBackToInMemoryAndWarns() {
|
||||
FleetConfig.Broker broker = new FleetConfig.Broker("amqp://user:lame@old:5672/", "LAVINMQ_URI", null);
|
||||
recording(broker, Map.of("LAVINMQ_URI", " "), false, (opener, appender, inbox) -> {
|
||||
recording(broker, Map.of("LAVINMQ_URI", " "), false, (opener, events, inbox) -> {
|
||||
assertInstanceOf(InMemoryReplyInbox.class, inbox);
|
||||
assertNull(opener.offeredUri, "a blank env value must not select AMQP, not even via the literal uri");
|
||||
assertTrue(hasWarnContaining(appender, "LAVINMQ_URI"),
|
||||
assertTrue(hasWarnContaining(events, "LAVINMQ_URI"),
|
||||
"a blank uriEnv value must warn, and must not fall back to the literal uri");
|
||||
});
|
||||
}
|
||||
@@ -106,22 +95,22 @@ class FleetdReplyInboxSelectionTest {
|
||||
void bothUriAndUriEnvSetUriEnvWinsDeterministically() {
|
||||
FleetConfig.Broker broker
|
||||
= new FleetConfig.Broker("amqp://user:oldpw@old.example:5672/", "LAVINMQ_URI", null);
|
||||
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), false, (opener, appender, inbox) -> {
|
||||
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), false, (opener, events, inbox) -> {
|
||||
assertEquals(RESOLVED_URI, opener.offeredUri,
|
||||
"uriEnv must win over uri, deterministically, every run");
|
||||
assertTrue(appender.list.stream().anyMatch(e -> e.getFormattedMessage().contains("broker.uri is ignored")),
|
||||
assertTrue(events.stream().anyMatch(e -> e.getFormattedMessage().contains("broker.uri is ignored")),
|
||||
"must log that the literal uri is ignored when uriEnv is set");
|
||||
assertNoLogContains(appender, SECRET);
|
||||
assertNoLogContains(appender, "oldpw");
|
||||
assertNoLogContains(events, SECRET);
|
||||
assertNoLogContains(events, "oldpw");
|
||||
});
|
||||
}
|
||||
|
||||
@Test
|
||||
void unreachableBrokerStartsDaemonWithInMemoryInboxAndLoudWarning() {
|
||||
FleetConfig.Broker broker = new FleetConfig.Broker(null, "LAVINMQ_URI", null);
|
||||
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), true, (opener, appender, inbox) -> {
|
||||
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), true, (opener, events, inbox) -> {
|
||||
assertInstanceOf(InMemoryReplyInbox.class, inbox, "an unreachable broker must NOT stop the daemon");
|
||||
String warn = appender.list.stream()
|
||||
String warn = events.stream()
|
||||
.filter(e -> e.getLevel() == Level.WARN)
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.reduce("", (a, b) -> a + "\n" + b)
|
||||
@@ -129,16 +118,16 @@ class FleetdReplyInboxSelectionTest {
|
||||
assertTrue(warn.contains("durable") && warn.contains("soft-state"),
|
||||
"the warning must say exactly what was lost: durable delivery off, replies soft-state");
|
||||
assertTrue(!warn.contains(SECRET), "the failing URI must be logged with credentials stripped");
|
||||
assertNoLogContains(appender, SECRET);
|
||||
assertNoLogContains(events, SECRET);
|
||||
});
|
||||
}
|
||||
|
||||
@Test
|
||||
void noBrokerConfiguredStaysQuietInMemory() {
|
||||
FleetConfig.Broker broker = null;
|
||||
recording(broker, Map.of(), false, (opener, appender, inbox) -> {
|
||||
recording(broker, Map.of(), false, (opener, events, inbox) -> {
|
||||
assertInstanceOf(InMemoryReplyInbox.class, inbox);
|
||||
assertTrue(appender.list.stream().noneMatch(e -> e.getLevel() == Level.WARN),
|
||||
assertTrue(events.stream().noneMatch(e -> e.getLevel() == Level.WARN),
|
||||
"no broker configured must keep the existing QUIET in-memory path — no warning");
|
||||
});
|
||||
}
|
||||
@@ -152,21 +141,18 @@ class FleetdReplyInboxSelectionTest {
|
||||
closedPort = s.getLocalPort();
|
||||
}
|
||||
FleetConfig.Broker broker = new FleetConfig.Broker(null, "LAVINMQ_URI", null);
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
try (CapturedLog log = attach()) {
|
||||
ReplyInbox inbox = Fleetd.selectReplyInbox(
|
||||
broker, Map.of("LAVINMQ_URI", "amqp://user:" + SECRET + "@127.0.0.1:" + closedPort + "/vh"),
|
||||
AmqpReplyInbox::open);
|
||||
assertInstanceOf(InMemoryReplyInbox.class, inbox,
|
||||
"a genuinely unreachable broker (real AmqpReplyInbox::open) must fall back to in-memory");
|
||||
} finally {
|
||||
detach(appender);
|
||||
assertNoLogContains(log.events(), SECRET);
|
||||
}
|
||||
assertNoLogContains(appender, SECRET);
|
||||
}
|
||||
|
||||
private boolean hasWarnContaining(ListAppender<ILoggingEvent> appender, String fragment) {
|
||||
return appender.list.stream().anyMatch(e ->
|
||||
private boolean hasWarnContaining(List<ILoggingEvent> events, String fragment) {
|
||||
return events.stream().anyMatch(e ->
|
||||
e.getLevel() == Level.WARN && e.getFormattedMessage().contains(fragment));
|
||||
}
|
||||
|
||||
@@ -174,18 +160,14 @@ class FleetdReplyInboxSelectionTest {
|
||||
private void recording(FleetConfig.Broker broker, Map<String, String> env, boolean unreachable, Check check) {
|
||||
RecordingAmqp opener = new RecordingAmqp();
|
||||
opener.unreachable = unreachable;
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
ReplyInbox inbox;
|
||||
try {
|
||||
inbox = Fleetd.selectReplyInbox(broker, env, opener);
|
||||
} finally {
|
||||
detach(appender);
|
||||
try (CapturedLog log = attach()) {
|
||||
ReplyInbox inbox = Fleetd.selectReplyInbox(broker, env, opener);
|
||||
check.run(opener, log.events(), inbox);
|
||||
}
|
||||
check.run(opener, appender, inbox);
|
||||
}
|
||||
|
||||
@FunctionalInterface
|
||||
private interface Check {
|
||||
void run(RecordingAmqp opener, ListAppender<ILoggingEvent> appender, ReplyInbox inbox);
|
||||
void run(RecordingAmqp opener, List<ILoggingEvent> events, ReplyInbox inbox);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,286 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
|
||||
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||
import dev.ltms.fleet.inject.TurnListener;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.msg.TurnToken;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #561: {@code Fleetd.turnListener} composes the completion resolver and the session
|
||||
* manager into one {@link TurnListener}. Each of the four callbacks below used to be two bare,
|
||||
* unguarded statements (completion first, then sessions) — a throw from the session half used to
|
||||
* skip nothing <em>after</em> it (there was nothing after it), but nothing enforced that the
|
||||
* completion half had to come first either, beyond call order in the source. {@code onDelivered}
|
||||
* had the identical shape and was fixed by #556 (moving its registration off the fan-out
|
||||
* entirely); these four callbacks cannot be fixed that way, so the fan-out itself is hardened
|
||||
* instead (see {@link Fleetd#turnListener} and {@code Fleetd.bothMustRun}/{@code
|
||||
* bothMustRunKeepingSecondResult}).
|
||||
*
|
||||
* <p>The invariant under test: <b>a throwing session-listener must not prevent the completion
|
||||
* resolver from being told the turn ended.</b> Each test below builds the exact production
|
||||
* composition ({@link Fleetd#turnListener}) from a real {@link CompletionResolver} and a fake
|
||||
* {@code sessions} half that throws, then asserts the completion half's effect (the captured
|
||||
* {@link Rendezvous} waiter resolving) happened anyway — never by inspecting call order directly.
|
||||
*/
|
||||
class FleetdTurnListenerCompositionTest {
|
||||
|
||||
/** Records which callbacks ran and can be told to throw from a chosen one. */
|
||||
private static final class RecordingSessions implements TurnListener {
|
||||
final Set<String> called = new LinkedHashSet<>();
|
||||
private final Set<String> throwing;
|
||||
|
||||
RecordingSessions(String... throwingMethods) {
|
||||
this.throwing = Set.of(throwingMethods);
|
||||
}
|
||||
|
||||
private void maybeThrow(String method) {
|
||||
called.add(method);
|
||||
if (throwing.contains(method)) {
|
||||
throw new IllegalStateException("boom: sessions." + method);
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
maybeThrow("onTurnComplete");
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean onTurnCompleteWithPostAction(String target) {
|
||||
maybeThrow("onTurnCompleteWithPostAction");
|
||||
return true;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onTurnFailed(String target) {
|
||||
maybeThrow("onTurnFailed");
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onTurnFailed(String target, String reason) {
|
||||
maybeThrow("onTurnFailedWithReason");
|
||||
}
|
||||
}
|
||||
|
||||
private static CompletionResolver newResolver(FakeHerdr herdr, Rendezvous rendezvous) {
|
||||
return new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||
}
|
||||
|
||||
/**
|
||||
* A resolver with a controllable clock, and the clock itself, so a test can register a turn
|
||||
* "delivered" at time 0 and then jump the clock past {@link CompletionResolver#MIN_TURN_NANOS}
|
||||
* before resolving it — otherwise {@code onTurnComplete}/{@code onTurnCompleteWithPostAction}
|
||||
* resolve within microseconds of registering in-test, well inside the fleetd#164 floor, and
|
||||
* get classified as a too-fast crash rather than a real completion. Mirrors the {@code
|
||||
* LongSupplier} clock-injection pattern fleetd#164's own tests use.
|
||||
*/
|
||||
private static CompletionResolver newResolverPastTheFloor(FakeHerdr herdr, Rendezvous rendezvous,
|
||||
AtomicLong clock) {
|
||||
LongSupplier nowNanos = clock::get;
|
||||
return new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none(), nowNanos);
|
||||
}
|
||||
|
||||
@Test
|
||||
void onTurnCompleteResolvesTheWaiterEvenWhenTheSessionHalfThrows() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ real answer\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
AtomicLong clock = new AtomicLong(0);
|
||||
CompletionResolver completion = newResolverPastTheFloor(herdr, rendezvous, clock);
|
||||
var waiter = rendezvous.open("term_a");
|
||||
completion.register("term_a", new TurnToken("term_a", waiter)); // delivered at clock=0
|
||||
clock.set(CompletionResolver.MIN_TURN_NANOS + 1_000_000); // past the too-fast floor
|
||||
|
||||
RecordingSessions sessions = new RecordingSessions("onTurnComplete");
|
||||
TurnListener composed = Fleetd.turnListener(completion, sessions);
|
||||
|
||||
IllegalStateException thrown = assertThrows(IllegalStateException.class,
|
||||
() -> composed.onTurnComplete("term_a"),
|
||||
"the session half's throw must still escape the composed listener");
|
||||
assertEquals("boom: sessions.onTurnComplete", thrown.getMessage());
|
||||
assertTrue(sessions.called.contains("onTurnComplete"), "the session half must have run");
|
||||
|
||||
// onTurnComplete resolves off-thread (a virtual thread) — wait on the waiter itself,
|
||||
// exactly like MessageServiceTest.completionFallbackIsNeverQueued does.
|
||||
Rendezvous.Resolution resolution = waiter.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(Rendezvous.Kind.COMPLETION, resolution.kind(),
|
||||
"the completion resolver must still have resolved the send, despite the session "
|
||||
+ "half throwing");
|
||||
assertEquals("real answer", resolution.text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void onTurnFailedResolvesTheWaiterAsFailedEvenWhenTheSessionHalfThrows() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ crash context\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver completion = newResolver(herdr, rendezvous);
|
||||
var waiter = rendezvous.open("term_b");
|
||||
completion.register("term_b", new TurnToken("term_b", waiter));
|
||||
|
||||
RecordingSessions sessions = new RecordingSessions("onTurnFailed");
|
||||
TurnListener composed = Fleetd.turnListener(completion, sessions);
|
||||
|
||||
IllegalStateException thrown = assertThrows(IllegalStateException.class,
|
||||
() -> composed.onTurnFailed("term_b"),
|
||||
"the session half's throw must still escape the composed listener");
|
||||
assertEquals("boom: sessions.onTurnFailed", thrown.getMessage());
|
||||
assertTrue(sessions.called.contains("onTurnFailed"), "the session half must have run");
|
||||
|
||||
Rendezvous.Resolution resolution = waiter.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(Rendezvous.Kind.FAILED, resolution.kind(),
|
||||
"the completion resolver must still have failed the send, despite the session "
|
||||
+ "half throwing");
|
||||
}
|
||||
|
||||
@Test
|
||||
void onTurnFailedWithReasonResolvesTheWaiterEvenWhenTheSessionHalfThrows() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver completion = newResolver(herdr, rendezvous);
|
||||
var waiter = rendezvous.open("term_c");
|
||||
completion.register("term_c", new TurnToken("term_c", waiter));
|
||||
|
||||
// fleetd #561: the composed listener calls sessions.onTurnFailed(target) — the ONE-arg
|
||||
// overload — for this reason-carrying event too (matching the pre-existing production
|
||||
// behaviour: SessionManager never overrides the two-arg overload either), so the
|
||||
// throwing key here is "onTurnFailed", not a distinct "...WithReason" one.
|
||||
RecordingSessions sessions = new RecordingSessions("onTurnFailed");
|
||||
TurnListener composed = Fleetd.turnListener(completion, sessions);
|
||||
|
||||
IllegalStateException thrown = assertThrows(IllegalStateException.class,
|
||||
() -> composed.onTurnFailed("term_c", "worker unreachable"),
|
||||
"the session half's throw must still escape the composed listener");
|
||||
assertEquals("boom: sessions.onTurnFailed", thrown.getMessage());
|
||||
assertTrue(sessions.called.contains("onTurnFailed"), "the session half must have run");
|
||||
|
||||
Rendezvous.Resolution resolution = waiter.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(Rendezvous.Kind.FAILED, resolution.kind(),
|
||||
"the completion resolver must still have failed the send, despite the session "
|
||||
+ "half throwing");
|
||||
assertEquals("worker unreachable", resolution.text(),
|
||||
"the explicit reason must still reach the resolved send");
|
||||
}
|
||||
|
||||
@Test
|
||||
void onTurnCompleteWithPostActionResolvesTheWaiterEvenWhenTheSessionHalfThrows() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ answer before reset\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
AtomicLong clock = new AtomicLong(0);
|
||||
CompletionResolver completion = newResolverPastTheFloor(herdr, rendezvous, clock);
|
||||
var waiter = rendezvous.open("term_d");
|
||||
completion.register("term_d", new TurnToken("term_d", waiter)); // delivered at clock=0
|
||||
clock.set(CompletionResolver.MIN_TURN_NANOS + 1_000_000); // past the too-fast floor
|
||||
|
||||
RecordingSessions sessions = new RecordingSessions("onTurnCompleteWithPostAction");
|
||||
TurnListener composed = Fleetd.turnListener(completion, sessions);
|
||||
|
||||
IllegalStateException thrown = assertThrows(IllegalStateException.class,
|
||||
() -> composed.onTurnCompleteWithPostAction("term_d"),
|
||||
"the session half's throw must still escape the composed listener");
|
||||
assertEquals("boom: sessions.onTurnCompleteWithPostAction", thrown.getMessage());
|
||||
assertTrue(sessions.called.contains("onTurnCompleteWithPostAction"), "the session half must have run");
|
||||
|
||||
// resolveBeforePostAction is synchronous by design (it must run before the context reset
|
||||
// can erase the pane) — the waiter is already resolved by the time the throw propagates.
|
||||
assertTrue(waiter.isDone(), "resolveBeforePostAction is synchronous — the send must "
|
||||
+ "already be resolved once the composed call returns (by throwing)");
|
||||
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
|
||||
assertEquals("answer before reset", waiter.getNow(null).text());
|
||||
}
|
||||
|
||||
/**
|
||||
* The mirror case (comment 17041's acceptance item 2): the completion half throws, and the
|
||||
* session half still ran. The production {@link CompletionResolver} is deliberately defensive
|
||||
* (its scrape reads are wrapped in {@code catch (RuntimeException)}, by the same fail-open
|
||||
* design {@link CompletionResolver#captureBaseline} documents) so it essentially never throws
|
||||
* synchronously in normal operation — forcing it to do so needs a genuinely-real trigger, not
|
||||
* a fabricated one. {@code ConcurrentHashMap.get(null)} is that trigger: passing a {@code null}
|
||||
* target makes {@code onTurnComplete}'s {@code inFlight.get(target)} throw a
|
||||
* {@link NullPointerException} before it ever starts its resolving thread — a real code path,
|
||||
* not a contrived one.
|
||||
*/
|
||||
@Test
|
||||
void sessionHalfStillRunsWhenTheCompletionHalfThrowsSynchronously() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver completion = newResolver(herdr, rendezvous);
|
||||
|
||||
RecordingSessions sessions = new RecordingSessions(); // throws from nothing
|
||||
TurnListener composed = Fleetd.turnListener(completion, sessions);
|
||||
|
||||
assertThrows(NullPointerException.class, () -> composed.onTurnComplete(null),
|
||||
"ConcurrentHashMap.get(null) inside CompletionResolver.onTurnComplete must still "
|
||||
+ "escape the composed listener");
|
||||
assertTrue(sessions.called.contains("onTurnComplete"),
|
||||
"the session half must still have run even though the completion half threw first");
|
||||
}
|
||||
|
||||
/**
|
||||
* The mirror case above only exercises {@code onTurnComplete}/{@code bothMustRun}. {@code
|
||||
* onTurnCompleteWithPostAction} is composed through the OTHER helper,
|
||||
* {@code bothMustRunKeepingSecondResult}, and nothing previously asserted that its session half
|
||||
* still runs when its completion half throws — that helper could be reverted to the pre-#561
|
||||
* broken shape (run the session half only if the completion half did not throw) and the suite
|
||||
* would still stay green. Uses the same real, non-fabricated trigger as the test above: a
|
||||
* {@code null} target makes {@code resolveBeforePostAction}'s {@code inFlight.get(target)}
|
||||
* throw a {@link NullPointerException} before {@code resolve} is ever entered.
|
||||
*/
|
||||
@Test
|
||||
void sessionHalfStillRunsWhenTheCompletionHalfThrowsSynchronouslyForPostAction() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver completion = newResolver(herdr, rendezvous);
|
||||
|
||||
RecordingSessions sessions = new RecordingSessions(); // throws from nothing
|
||||
TurnListener composed = Fleetd.turnListener(completion, sessions);
|
||||
|
||||
assertThrows(NullPointerException.class, () -> composed.onTurnCompleteWithPostAction(null),
|
||||
"ConcurrentHashMap.get(null) inside CompletionResolver.resolveBeforePostAction must "
|
||||
+ "still escape the composed listener");
|
||||
assertTrue(sessions.called.contains("onTurnCompleteWithPostAction"),
|
||||
"the session half must still have run even though the completion half threw first");
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code bothMustRun} must not just run both halves — it must not DROP a second failure when
|
||||
* both halves throw. Forces the completion half to throw (the same real {@code
|
||||
* inFlight.get(null)} NullPointerException trigger used above) while the session half throws a
|
||||
* distinct {@link IllegalStateException}, and asserts the completion half's throwable is what
|
||||
* escapes while the session half's throwable survives as a suppressed exception rather than
|
||||
* being silently discarded.
|
||||
*/
|
||||
@Test
|
||||
void bothFailuresEscapeWhenBothHalvesThrowDistinctExceptions() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver completion = newResolver(herdr, rendezvous);
|
||||
|
||||
RecordingSessions sessions = new RecordingSessions("onTurnComplete");
|
||||
TurnListener composed = Fleetd.turnListener(completion, sessions);
|
||||
|
||||
NullPointerException thrown = assertThrows(NullPointerException.class,
|
||||
() -> composed.onTurnComplete(null),
|
||||
"the completion half's throw (NPE from inFlight.get(null)) must be what escapes");
|
||||
assertTrue(sessions.called.contains("onTurnComplete"), "the session half must still have run");
|
||||
assertEquals(1, thrown.getSuppressed().length,
|
||||
"the session half's distinct failure must be recorded as suppressed, not dropped");
|
||||
assertEquals(IllegalStateException.class, thrown.getSuppressed()[0].getClass());
|
||||
assertEquals("boom: sessions.onTurnComplete", thrown.getSuppressed()[0].getMessage());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,72 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
|
||||
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||
import dev.ltms.fleet.inject.TurnRegistrar;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.msg.TurnToken;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #589 Group 3, site {@code :505}. {@code Fleetd.main} wires the {@link
|
||||
* dev.ltms.fleet.inject.Injector}'s {@link TurnRegistrar} with {@code completion::register} —
|
||||
* before this ticket that was an inline argument to {@code new Injector(...)}, so nothing could
|
||||
* pin it directly. Measured: replacing it with {@link TurnRegistrar#NOOP} at the call site
|
||||
* compiles with 0 errors and leaves the full suite green, because {@code onDelivered}'s own {@code
|
||||
* captureBaseline} performs the identical {@code inFlight} check-and-put a moment later on the
|
||||
* ordinary delivery path — the two are indistinguishable unless something reads the resolver
|
||||
* between {@code register} and {@code onDelivered}, or {@code onDelivered} never runs at all (the
|
||||
* gap fleetd #556 introduced {@link TurnRegistrar} to close).
|
||||
*
|
||||
* <p>This test calls {@link Fleetd#turnRegistrar} directly — never {@code Injector} or {@code
|
||||
* main} — and drives {@link CompletionResolver} entirely through its public API: {@link
|
||||
* TurnRegistrar#register} followed by {@link CompletionResolver#resolveBeforePostAction}, which
|
||||
* looks up the same {@code inFlight} entry {@code onTurnComplete} would. With the real registrar,
|
||||
* that entry exists and the waiter opened by {@link Rendezvous#open} resolves; with {@link
|
||||
* TurnRegistrar#NOOP} nothing was ever registered, {@code resolveBeforePostAction} finds no
|
||||
* in-flight turn, and the waiter is left exactly as it started — never done.
|
||||
*/
|
||||
class FleetdTurnRegistrarWiringTest {
|
||||
|
||||
private static final String T = "term_a";
|
||||
|
||||
@Test
|
||||
@DisplayName("Fleetd.turnRegistrar delegates to the real CompletionResolver, not a no-op")
|
||||
void turnRegistrarDelegatesToCompletionRegister() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ BUILD GREEN: 391 files\n❯ ");
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
// fleetd#164: an ever-advancing fake clock stands in for the real time a turn would take
|
||||
// between delivery and resolution, so the MIN_TURN_NANOS "too fast" floor never trips here —
|
||||
// see MessageServiceTest's resolverClock for the same technique.
|
||||
AtomicLong clock = new AtomicLong();
|
||||
CompletionResolver completion = new CompletionResolver(agents, rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none(),
|
||||
() -> clock.addAndGet(CompletionResolver.MIN_TURN_NANOS + 1));
|
||||
|
||||
TurnRegistrar registrar = Fleetd.turnRegistrar(completion);
|
||||
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(T);
|
||||
registrar.register(T, new TurnToken(T, waiter));
|
||||
|
||||
// Mirrors what Injector.onStatus's confirmed working->idle boundary would trigger via
|
||||
// CompletionResolver.onTurnComplete — resolveBeforePostAction is the public, synchronous
|
||||
// twin of that path and reads the exact same inFlight entry register() must have written.
|
||||
completion.resolveBeforePostAction(T);
|
||||
|
||||
assertTrue(waiter.isDone(),
|
||||
"Fleetd.turnRegistrar(completion) must return completion::register — replacing it "
|
||||
+ "with TurnRegistrar.NOOP at the Fleetd.turnRegistrar call site means this "
|
||||
+ "turn is never registered with CompletionResolver, so resolveBeforePostAction "
|
||||
+ "finds no in-flight turn and this waiter is never resolved");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,252 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #613: a {@code MemberRole} with no {@code fleet.<role>s:} pool falls back to
|
||||
* <em>every</em> configured profile ({@code FleetConfig#candidateProfiles}), and one with no
|
||||
* {@code fleet.charters.<role>:} entry runs with only the launcher's reply charter. Both are
|
||||
* deliberate, legitimate states — neither is refused by {@code validateMembers()} — but both were
|
||||
* silent at boot. On the host that opened this ticket, an unqualified {@code hunter} spawn silently
|
||||
* widened to all 8 configured profiles and its resolved first choice was {@code local}, a profile
|
||||
* every other pool on that same config gave weight 0 to.
|
||||
*
|
||||
* <p>{@link Fleetd#reportRoleFallbackGaps} must name every gapped role, and for a pool gap, the
|
||||
* exact resolved first-choice profile — that number, not the pool size, is what actually surprised
|
||||
* the operator. Mirrors {@link ExhaustedPatternGapReportTest}'s pattern: capture the real log via a
|
||||
* {@link ListAppender} rather than asserting on the call site's source text.
|
||||
*/
|
||||
class RoleFallbackGapReportTest {
|
||||
|
||||
private static FleetConfig load(Path dir, String yaml) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml);
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
/**
|
||||
* The level this logger had before {@link #attach()} raised it, so {@link #detach} can put it
|
||||
* back. {@code null} means "inherit from the parent" — the state this logger starts in.
|
||||
*/
|
||||
private static Level originalLevel;
|
||||
|
||||
/**
|
||||
* {@code reportRoleFallbackGaps} logs at INFO, and {@code logback-test.xml} sets
|
||||
* {@code dev.ltms.fleet} to WARN — so INFO events are dropped by the level check before any
|
||||
* appender sees them. Raise the level for the duration of the test, exactly like {@code
|
||||
* GitHostShapeReportTest#attach}.
|
||||
*/
|
||||
private static ListAppender<ILoggingEvent> attach() {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
originalLevel = logger.getLevel();
|
||||
logger.setLevel(Level.INFO);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
return appender;
|
||||
}
|
||||
|
||||
private static void detach(ListAppender<ILoggingEvent> appender) {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
logger.detachAppender(appender);
|
||||
logger.setLevel(originalLevel);
|
||||
}
|
||||
|
||||
private static List<String> infoMessages(ListAppender<ILoggingEvent> appender) {
|
||||
return appender.list.stream()
|
||||
.filter(e -> e.getLevel() == Level.INFO)
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.toList();
|
||||
}
|
||||
|
||||
/**
|
||||
* Reproduces the shape measured in the ticket: {@code dev}, {@code reviewer} and
|
||||
* {@code architect} each have a pool and a charter; {@code hunter} has neither. The pool-gap
|
||||
* line must name {@code hunter}, the profile count (3), and the resolved first choice
|
||||
* ({@code local}, the first profile in definition order) — and must not name the three healthy
|
||||
* roles. The charter-gap line must separately name only {@code hunter}.
|
||||
*/
|
||||
@Test
|
||||
void hunterWithNoPoolOrCharterIsNamedWithItsResolvedFirstChoice(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = load(dir, """
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
sonnet:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
terra:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
fleet:
|
||||
developers:
|
||||
a:
|
||||
profile: sonnet
|
||||
reviewers:
|
||||
b:
|
||||
profile: terra
|
||||
architects:
|
||||
c:
|
||||
profile: sonnet
|
||||
charters:
|
||||
dev: "dev charter text"
|
||||
reviewer: "reviewer charter text"
|
||||
architect: "architect charter text"
|
||||
""");
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
Fleetd.reportRoleFallbackGaps(cfg);
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
List<String> infos = infoMessages(appender);
|
||||
|
||||
String poolLine = infos.stream()
|
||||
.filter(m -> m.contains("no fleet.<role>s: pool"))
|
||||
.findFirst()
|
||||
.orElseThrow(() -> new AssertionError("expected a pool-gap INFO line: " + infos));
|
||||
assertTrue(poolLine.contains("hunter"), poolLine);
|
||||
assertTrue(poolLine.contains("3"), "must name the profile count the gap opens onto: " + poolLine);
|
||||
assertTrue(poolLine.contains("'local'"),
|
||||
"must name the resolved first-choice profile, the number that actually surprised "
|
||||
+ "the operator: " + poolLine);
|
||||
for (String healthy : List.of("dev", "reviewer", "architect")) {
|
||||
assertFalse(poolLine.contains(healthy),
|
||||
"pool-gap line must not name a role that has a pool: " + poolLine);
|
||||
}
|
||||
|
||||
String charterLine = infos.stream()
|
||||
.filter(m -> m.contains("no fleet.charters: entry"))
|
||||
.findFirst()
|
||||
.orElseThrow(() -> new AssertionError("expected a charter-gap INFO line: " + infos));
|
||||
assertTrue(charterLine.contains("hunter"), charterLine);
|
||||
for (String healthy : List.of("dev", "reviewer", "architect")) {
|
||||
assertFalse(charterLine.contains(healthy),
|
||||
"charter-gap line must not name a role that has a charter: " + charterLine);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The resolved first choice must be genuinely computed from definition order, not hardcoded —
|
||||
* reordering {@code profiles:} so a different entry comes first changes the reported choice.
|
||||
*/
|
||||
@Test
|
||||
void theResolvedFirstChoiceFollowsProfileDefinitionOrder(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = load(dir, """
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
local:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
fleet:
|
||||
developers:
|
||||
a:
|
||||
profile: sonnet
|
||||
charters:
|
||||
dev: "dev charter text"
|
||||
""");
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
Fleetd.reportRoleFallbackGaps(cfg);
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
String poolLine = infoMessages(appender).stream()
|
||||
.filter(m -> m.contains("no fleet.<role>s: pool"))
|
||||
.findFirst()
|
||||
.orElseThrow();
|
||||
// hunter, reviewer and architect all lack a pool here; each falls back to the full 2-profile
|
||||
// set and the first-choice is 'sonnet' because it is first in profiles: definition order.
|
||||
assertTrue(poolLine.contains("'sonnet'"), poolLine);
|
||||
assertFalse(poolLine.contains("'local'"), poolLine);
|
||||
}
|
||||
|
||||
/** A config with a pool and a charter for every role produces no role-fallback log at all. */
|
||||
@Test
|
||||
void everyRoleWithAPoolAndACharterProducesNoLogAtAll(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = load(dir, """
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
fleet:
|
||||
developers:
|
||||
a:
|
||||
profile: sonnet
|
||||
reviewers:
|
||||
b:
|
||||
profile: sonnet
|
||||
hunters:
|
||||
c:
|
||||
profile: sonnet
|
||||
architects:
|
||||
d:
|
||||
profile: sonnet
|
||||
charters:
|
||||
dev: "dev charter text"
|
||||
reviewer: "reviewer charter text"
|
||||
hunter: "hunter charter text"
|
||||
architect: "architect charter text"
|
||||
""");
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
Fleetd.reportRoleFallbackGaps(cfg);
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
assertTrue(appender.list.isEmpty(),
|
||||
"a config with no gaps must not print a per-role block: " + infoMessages(appender));
|
||||
}
|
||||
|
||||
/**
|
||||
* A config with only {@code profiles:} and no {@code fleet:} block at all must still be
|
||||
* reported (every role is gapped, both pool and charter) rather than throwing — this is the
|
||||
* exact shape {@code candidateProfiles}' fallback exists to keep starting.
|
||||
*/
|
||||
@Test
|
||||
void aConfigWithNoFleetBlockAtAllReportsEveryRoleGapped(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = load(dir, """
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
""");
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
Fleetd.reportRoleFallbackGaps(cfg);
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
List<String> infos = infoMessages(appender);
|
||||
String poolLine = infos.stream()
|
||||
.filter(m -> m.contains("no fleet.<role>s: pool"))
|
||||
.findFirst()
|
||||
.orElseThrow(() -> new AssertionError("expected a pool-gap INFO line: " + infos));
|
||||
String charterLine = infos.stream()
|
||||
.filter(m -> m.contains("no fleet.charters: entry"))
|
||||
.findFirst()
|
||||
.orElseThrow(() -> new AssertionError("expected a charter-gap INFO line: " + infos));
|
||||
for (String role : List.of("dev", "hunter", "reviewer", "architect")) {
|
||||
assertTrue(poolLine.contains(role), poolLine);
|
||||
assertTrue(charterLine.contains(role), charterLine);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,15 +1,12 @@
|
||||
package dev.ltms.fleet.auth;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.LoggerContext;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.fleet.testing.CapturedLog;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
@@ -24,28 +21,21 @@ import static org.junit.jupiter.api.Assertions.*;
|
||||
class AuditLogTest {
|
||||
|
||||
private final ObjectMapper mapper = new ObjectMapper();
|
||||
private ListAppender<ILoggingEvent> appender;
|
||||
private ch.qos.logback.classic.Logger auditLogger;
|
||||
private CapturedLog auditLog;
|
||||
|
||||
@BeforeEach
|
||||
void attach() {
|
||||
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
|
||||
auditLogger = ctx.getLogger("audit");
|
||||
appender = new ListAppender<>();
|
||||
appender.setContext(ctx);
|
||||
appender.start();
|
||||
auditLogger.addAppender(appender);
|
||||
auditLogger.setLevel(Level.INFO);
|
||||
auditLog = CapturedLog.at("audit", Level.INFO);
|
||||
}
|
||||
|
||||
@AfterEach
|
||||
void detach() {
|
||||
auditLogger.detachAppender(appender);
|
||||
auditLog.close();
|
||||
}
|
||||
|
||||
private JsonNode onlyRecord() throws Exception {
|
||||
assertEquals(1, appender.list.size(), "exactly one audit line expected");
|
||||
String line = appender.list.getFirst().getFormattedMessage();
|
||||
assertEquals(1, auditLog.events().size(), "exactly one audit line expected");
|
||||
String line = auditLog.events().getFirst().getFormattedMessage();
|
||||
return mapper.readTree(line); // throws if the line is not valid JSON
|
||||
}
|
||||
|
||||
|
||||
@@ -1,8 +1,11 @@
|
||||
package dev.ltms.fleet.config;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import dev.ltms.fleet.auth.MemberRegistry;
|
||||
import dev.ltms.fleet.msg.LeadMailbox;
|
||||
import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.testing.CapturedLog;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
@@ -93,6 +96,72 @@ class FleetConfigTest {
|
||||
"unset means off — today's behaviour, unchanged");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #601 review (measured 2026-09-22): this guard used to throw {@link
|
||||
* IllegalStateException} and refuse to start. On a host with the conflict configured, that
|
||||
* turned into a launchd restart loop with no readable cause, since {@code fleetd.yaml} is
|
||||
* gitignored. It must now WARN and let the daemon start, and the warning must carry both values
|
||||
* so an operator can fix the config without reading the source. Pinning the log line itself (via
|
||||
* {@link CapturedLog}) rather than a snippet of production source text — the latter is the
|
||||
* anti-pattern this repo avoids; the former is the actual observable behaviour a reader (or an
|
||||
* alert on the log) depends on.
|
||||
*/
|
||||
@Test
|
||||
void aClaudeProfileWithConflictingAutoCompactFlagAndEnvironmentWindowLoadsAndWarnsWithBothValues(
|
||||
@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("conflicting-auto-compact-window.yaml");
|
||||
Files.writeString(f, """
|
||||
profiles:
|
||||
claude-profile:
|
||||
autoCompactWindow: 250000
|
||||
env:
|
||||
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "300000"
|
||||
""");
|
||||
|
||||
FleetConfig cfg;
|
||||
List<String> warnings;
|
||||
try (CapturedLog log = CapturedLog.at(FleetConfig.class, Level.WARN)) {
|
||||
cfg = FleetConfig.load(f);
|
||||
warnings = log.events().stream().map(ILoggingEvent::getFormattedMessage).toList();
|
||||
}
|
||||
|
||||
assertEquals(250_000, cfg.profiles().get("claude-profile").autoCompactWindow(),
|
||||
"the disagreement is reported, not corrected — the flag value still loads as-is");
|
||||
assertEquals(1, warnings.size(), "exactly one warning for the one conflicting profile: " + warnings);
|
||||
String warning = warnings.get(0);
|
||||
assertTrue(warning.contains("claude-profile"), "names the offending profile: " + warning);
|
||||
assertTrue(warning.contains("autoCompactWindow=250000"), "names the flag value: " + warning);
|
||||
assertTrue(warning.contains("CLAUDE_CODE_AUTO_COMPACT_WINDOW=300000"), "names the env value: " + warning);
|
||||
}
|
||||
|
||||
/**
|
||||
* The negative probe paired with the test above (per fleetd #601 review): a warning that fires
|
||||
* on every load and a warning that never fires read the same from a single test, so both must be
|
||||
* checked. No conflict here — the flag and the env value agree — so no warning should be logged.
|
||||
*/
|
||||
@Test
|
||||
void aClaudeProfileWithEqualAutoCompactFlagAndEnvironmentWindowLoadsWithNoWarning(@TempDir Path dir)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("equal-auto-compact-window.yaml");
|
||||
Files.writeString(f, """
|
||||
profiles:
|
||||
claude-profile:
|
||||
autoCompactWindow: 250000
|
||||
env:
|
||||
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "250000"
|
||||
""");
|
||||
|
||||
FleetConfig cfg;
|
||||
List<ILoggingEvent> events;
|
||||
try (CapturedLog log = CapturedLog.at(FleetConfig.class, Level.WARN)) {
|
||||
cfg = FleetConfig.load(f);
|
||||
events = log.events();
|
||||
}
|
||||
|
||||
assertEquals(250_000, cfg.profiles().get("claude-profile").autoCompactWindow());
|
||||
assertTrue(events.isEmpty(), "equal values must not warn: " + events);
|
||||
}
|
||||
|
||||
// ── fleetd #201 Unit 5: errorPattern ────────────────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
@@ -345,7 +414,7 @@ class FleetConfigTest {
|
||||
IllegalStateException unknownError = assertThrows(IllegalStateException.class,
|
||||
() -> FleetConfig.load(unknown).validateCharters());
|
||||
assertTrue(unknownError.getMessage().contains("architetc"));
|
||||
assertTrue(unknownError.getMessage().contains("[architect, dev, reviewer]"));
|
||||
assertTrue(unknownError.getMessage().contains("[architect, dev, hunter, reviewer]"));
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -812,13 +881,17 @@ class FleetConfigTest {
|
||||
reviewers:
|
||||
b:
|
||||
profile: sonnet
|
||||
hunters:
|
||||
c:
|
||||
profile: sonnet
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertEquals(List.of("sonnet"), cfg.fleet().profilesFor(MemberRole.DEV));
|
||||
assertEquals(List.of("sonnet"), cfg.fleet().profilesFor(MemberRole.HUNTER));
|
||||
assertEquals(List.of("sonnet"), cfg.fleet().profilesFor(MemberRole.REVIEWER));
|
||||
assertTrue(cfg.fleet().profilesFor(MemberRole.ARCHITECT).isEmpty());
|
||||
assertEquals(List.of(MemberRole.DEV, MemberRole.REVIEWER), cfg.fleet().rolesConfigured());
|
||||
assertEquals(List.of(MemberRole.DEV, MemberRole.HUNTER, MemberRole.REVIEWER), cfg.fleet().rolesConfigured());
|
||||
}
|
||||
|
||||
/** The case the two axes exist for: one backend, two roles, and neither is a duplicate. */
|
||||
|
||||
@@ -0,0 +1,40 @@
|
||||
package dev.ltms.fleet.health;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* fleetd #426: {@link FleetHealthMonitor#coverage} had zero references anywhere in the test tree —
|
||||
* not the method, not either output string, not the field it populates. This pins the method's own
|
||||
* three branches directly.
|
||||
*
|
||||
* <p>This is the easy half. It proves the method words each combination correctly, but it proves
|
||||
* nothing about whether {@code Fleetd.java}'s two call sites pass the right argument in the right
|
||||
* position — see {@code FleetdHealthCoverageSourceWiringTest} (package {@code dev.ltms.fleet}) for
|
||||
* the half that actually guards the call site, following the measured fleetd #415 lesson that a
|
||||
* thoroughly-tested method and an untested argument pairing at its call site are different risks.
|
||||
*
|
||||
* <p><strong>The three output strings are load-bearing and must not change here.</strong> {@code
|
||||
* "detection-only"} is read live off a running daemon's {@code fleet_list} today (measured
|
||||
* 2026-09-12) — this test intentionally asserts the exact literal strings so a future edit to the
|
||||
* wording trips it here first.
|
||||
*/
|
||||
class FleetHealthMonitorCoverageTest {
|
||||
|
||||
@Test
|
||||
void disabledIsOffRegardlessOfNotificationConfig() {
|
||||
assertEquals("off", FleetHealthMonitor.coverage(false, false));
|
||||
assertEquals("off", FleetHealthMonitor.coverage(false, true));
|
||||
}
|
||||
|
||||
@Test
|
||||
void enabledWithoutNotificationsIsDetectionOnly() {
|
||||
assertEquals("detection-only", FleetHealthMonitor.coverage(true, false));
|
||||
}
|
||||
|
||||
@Test
|
||||
void enabledWithNotificationsIsFull() {
|
||||
assertEquals("full", FleetHealthMonitor.coverage(true, true));
|
||||
}
|
||||
}
|
||||
@@ -456,7 +456,11 @@ public final class FakeHerdr implements HerdrClient {
|
||||
}
|
||||
}
|
||||
|
||||
/** fleetd #612 Unit A: set by {@link #close()} so a test can prove a router/client actually closed this. */
|
||||
public volatile boolean closed = false;
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
closed = true;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,9 +1,7 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.LoggerContext;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
@@ -12,8 +10,8 @@ import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.msg.TestTurnTokens;
|
||||
import dev.ltms.fleet.msg.TurnToken;
|
||||
import dev.ltms.fleet.testing.CapturedLog;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.Set;
|
||||
import java.util.regex.Pattern;
|
||||
@@ -21,6 +19,7 @@ import java.util.regex.Pattern;
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/** Unit behaviour of the CB-106 completion resolver in isolation from the injector. */
|
||||
@@ -470,15 +469,7 @@ class CompletionResolverTest {
|
||||
// CB-564: this used to be a bare DEBUG "failed send to X via turn-stall fallback" — a symptom
|
||||
// with no cause, and below the level anyone watching for member health would see. A fail that
|
||||
// resolves a caller's blocked send is at least WARN and must carry the reason.
|
||||
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
|
||||
ch.qos.logback.classic.Logger resolverLog =
|
||||
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(CompletionResolver.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.setContext(ctx);
|
||||
appender.start();
|
||||
resolverLog.addAppender(appender);
|
||||
resolverLog.setLevel(Level.WARN);
|
||||
try {
|
||||
try (CapturedLog log = CapturedLog.at(CompletionResolver.class, Level.WARN)) {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||
@@ -486,7 +477,7 @@ class CompletionResolverTest {
|
||||
|
||||
resolver.fail("term_a", null);
|
||||
|
||||
String warn = appender.list.stream()
|
||||
String warn = log.events().stream()
|
||||
.filter(e -> e.getLevel().equals(Level.WARN))
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.findFirst()
|
||||
@@ -494,8 +485,6 @@ class CompletionResolverTest {
|
||||
assertTrue(warn.contains("term_a"), "the log names the target: " + warn);
|
||||
assertTrue(warn.contains("stuck on an error screen"), "the log carries the reason: " + warn);
|
||||
assertTrue(waiter.isDone());
|
||||
} finally {
|
||||
resolverLog.detachAppender(appender);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -531,6 +520,254 @@ class CompletionResolverTest {
|
||||
assertTrue(rendezvous.isWaiting("term_a"), "turn N+1 is still awaiting its own resolution");
|
||||
}
|
||||
|
||||
// --- fleetd #556 (comment 16908): name the invariant the two-arg remove(target, turn) idiom
|
||||
// exists to maintain — a superseded turn's terminal handling must not evict its successor's
|
||||
// registration — and test THAT, not the idiom's literal spelling. Three entry points reach a
|
||||
// conditional remove: resolve()'s plain completion path, resolve()'s echoed-brief/noReportMessage
|
||||
// sub-path, and fail(). This is a non-goal-to-break for fleetd #556's redesign: flattening any of
|
||||
// these removes to the one-arg form must turn one of these three red. -------------------------
|
||||
|
||||
@Test
|
||||
void aSupersededTurnsPlainCompletionMustNotEvictItsSuccessorsRegistration() {
|
||||
// Turn A is overtaken by turn B on the same target (a rapid back-to-back send): B's own
|
||||
// delivery overwrites the map entry A's delivery installed. A's late completion fallback
|
||||
// then finally fires and reaches resolve()'s ordinary (non-echoed) completion branch — the
|
||||
// two-arg remove(target, turnA) there must be a no-op once the map holds B, not a blind
|
||||
// evict.
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ A's pre-turn pane\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||
|
||||
var waiterA = rendezvous.open("term_a");
|
||||
var turnA = new CompletionResolver.InFlight(waiterA, null); // baseline null: fails open, never suppressed
|
||||
rendezvous.close("term_a", waiterA); // A's own send's finally, freeing the session for B
|
||||
var waiterB = rendezvous.open("term_a");
|
||||
resolver.captureBaseline("term_a", new TurnToken("term_a", waiterB)); // the injector registering B
|
||||
|
||||
herdr.readText("⏺ A's late, genuine answer\n❯ "); // what the pane shows when A's fallback fires
|
||||
resolver.resolve("term_a", turnA);
|
||||
|
||||
assertTrue(waiterA.isDone(), "A's own resolve must still complete A's waiter normally");
|
||||
assertEquals(Rendezvous.Kind.COMPLETION, waiterA.getNow(null).kind());
|
||||
|
||||
CompletionResolver.InFlight afterA = resolver.inFlight("term_a");
|
||||
assertNotNull(afterA, "B's registration must still be present — a one-arg remove(target) "
|
||||
+ "here would evict it even though the map no longer holds turnA");
|
||||
assertEquals(waiterB, afterA.waiter(), "the surviving entry must still be B's, untouched by A's "
|
||||
+ "terminal handling");
|
||||
assertTrue(rendezvous.resolve("term_a", "B replied"),
|
||||
"B must still resolve normally through the ordinary reply path");
|
||||
assertEquals(Rendezvous.Kind.REPLY, waiterB.getNow(null).kind());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aSupersededTurnsEchoedNoReportSubPathMustNotEvictItsSuccessorsRegistration() {
|
||||
// Same invariant, driven down resolve()'s OTHER branch: a pane that merely echoes the
|
||||
// injected brief back (no real report) resolves through noReportMessage(), not the plain
|
||||
// tail — comment 16908 calls this "a sub-path of the first, not a separate method", reaching
|
||||
// the same terminal remove() call site from different logic above it.
|
||||
String echoedBrief = "y".repeat(450); // >= CompletionResolver.ECHO_MIN_CHARS normalised chars
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ " + echoedBrief + "\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||
|
||||
var waiterA = rendezvous.open("term_a");
|
||||
// Back-dated deliveredAtNanos (like the 2-arg InFlight convenience ctor) so MIN_TURN_NANOS
|
||||
// never fires; injectedText == the pane's own content so echoesInjectedBrief() is true.
|
||||
var turnA = new CompletionResolver.InFlight(waiterA, null, Long.MIN_VALUE / 2, echoedBrief);
|
||||
rendezvous.close("term_a", waiterA);
|
||||
var waiterB = rendezvous.open("term_a");
|
||||
resolver.captureBaseline("term_a", new TurnToken("term_a", waiterB));
|
||||
|
||||
resolver.resolve("term_a", turnA);
|
||||
|
||||
assertTrue(waiterA.isDone(), "A's own resolve must still complete A's waiter via the echoed "
|
||||
+ "no-report branch");
|
||||
assertTrue(waiterA.getNow(null).text().contains(CompletionResolver.NO_REPORT_PREFIX),
|
||||
"sanity: A must actually have gone down the noReportMessage sub-path, not the plain one");
|
||||
|
||||
CompletionResolver.InFlight afterA = resolver.inFlight("term_a");
|
||||
assertNotNull(afterA, "B's registration must still be present after A's echoed-brief resolve");
|
||||
assertEquals(waiterB, afterA.waiter());
|
||||
assertTrue(rendezvous.resolve("term_a", "B replied"),
|
||||
"B must still resolve normally through the ordinary reply path");
|
||||
assertEquals(Rendezvous.Kind.REPLY, waiterB.getNow(null).kind());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aSupersededTurnsFailMustNotEvictItsSuccessorsRegistration() {
|
||||
// The fail() entry point (CB-109 wedge / CB-110 drop): turn A is failed while turn B has
|
||||
// already superseded it in the map. fail()'s own two-arg remove(target, turnA) must be a
|
||||
// no-op once the map holds B.
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(new FakeHerdr()), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||
|
||||
var waiterA = rendezvous.open("term_a");
|
||||
var turnA = new CompletionResolver.InFlight(waiterA, null);
|
||||
rendezvous.close("term_a", waiterA);
|
||||
var waiterB = rendezvous.open("term_a");
|
||||
resolver.captureBaseline("term_a", new TurnToken("term_a", waiterB));
|
||||
|
||||
resolver.fail("term_a", turnA, "A wedged in an unknown state");
|
||||
|
||||
assertTrue(waiterA.isDone(), "A's own fail must still complete A's waiter");
|
||||
assertEquals(Rendezvous.Kind.FAILED, waiterA.getNow(null).kind());
|
||||
|
||||
CompletionResolver.InFlight afterA = resolver.inFlight("term_a");
|
||||
assertNotNull(afterA, "B's registration must still be present — a one-arg remove(target) in "
|
||||
+ "fail() would evict it even though the map no longer holds turnA");
|
||||
assertEquals(waiterB, afterA.waiter());
|
||||
assertTrue(rendezvous.resolve("term_a", "B replied"),
|
||||
"B must still resolve normally through the ordinary reply path");
|
||||
assertEquals(Rendezvous.Kind.REPLY, waiterB.getNow(null).kind());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aSupersededDoneTurnMustNotEvictItsSuccessorsRegistration() {
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(new FakeHerdr()), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||
|
||||
var waiterA = rendezvous.open("term_a");
|
||||
var turnA = new CompletionResolver.InFlight(waiterA, null);
|
||||
assertTrue(rendezvous.resolve("term_a", "A replied"));
|
||||
rendezvous.close("term_a", waiterA);
|
||||
var waiterB = rendezvous.open("term_a");
|
||||
resolver.captureBaseline("term_a", new TurnToken("term_a", waiterB));
|
||||
|
||||
resolver.resolve("term_a", turnA);
|
||||
|
||||
assertSuccessorRegistrationSurvives(resolver, rendezvous, waiterB,
|
||||
"a done turn must not evict B from resolve()'s early return");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aSupersededExhaustedTurnMustNotEvictItsSuccessorsRegistration() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ usage limit has been reached\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
target -> Pattern.compile("usage limit has been reached"), ExhaustionSink.none());
|
||||
|
||||
var waiterA = rendezvous.open("term_a");
|
||||
var turnA = new CompletionResolver.InFlight(waiterA, null);
|
||||
rendezvous.close("term_a", waiterA);
|
||||
var waiterB = rendezvous.open("term_a");
|
||||
resolver.captureBaseline("term_a", new TurnToken("term_a", waiterB));
|
||||
|
||||
resolver.resolve("term_a", turnA);
|
||||
|
||||
assertSuccessorRegistrationSurvives(resolver, rendezvous, waiterB,
|
||||
"an exhausted turn must not evict B from resolve()'s exhausted branch");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aSupersededBackendErrorTurnMustNotEvictItsSuccessorsRegistration() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ API Error: 400 invalid request body\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||
|
||||
var waiterA = rendezvous.open("term_a");
|
||||
var turnA = new CompletionResolver.InFlight(waiterA, null);
|
||||
rendezvous.close("term_a", waiterA);
|
||||
var waiterB = rendezvous.open("term_a");
|
||||
resolver.captureBaseline("term_a", new TurnToken("term_a", waiterB));
|
||||
|
||||
resolver.resolve("term_a", turnA);
|
||||
|
||||
assertSuccessorRegistrationSurvives(resolver, rendezvous, waiterB,
|
||||
"a backend-error turn must not evict B from resolve()'s error branch");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aSupersededRawExhaustedTurnMustNotEvictItsSuccessorsRegistration() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("╭────\nusage limit has been reached");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
target -> Pattern.compile("usage limit has been reached"), ExhaustionSink.none());
|
||||
|
||||
var waiterA = rendezvous.open("term_a");
|
||||
var turnA = new CompletionResolver.InFlight(waiterA, null);
|
||||
rendezvous.close("term_a", waiterA);
|
||||
var waiterB = rendezvous.open("term_a");
|
||||
resolver.captureBaseline("term_a", new TurnToken("term_a", waiterB));
|
||||
|
||||
resolver.resolve("term_a", turnA);
|
||||
|
||||
assertSuccessorRegistrationSurvives(resolver, rendezvous, waiterB,
|
||||
"a raw exhausted turn must not evict B from the raw-scrape exhausted branch");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aSupersededRawBackendErrorTurnMustNotEvictItsSuccessorsRegistration() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("╭────\nAPI Error: 400 invalid request body");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||
|
||||
var waiterA = rendezvous.open("term_a");
|
||||
var turnA = new CompletionResolver.InFlight(waiterA, null);
|
||||
rendezvous.close("term_a", waiterA);
|
||||
var waiterB = rendezvous.open("term_a");
|
||||
resolver.captureBaseline("term_a", new TurnToken("term_a", waiterB));
|
||||
|
||||
resolver.resolve("term_a", turnA);
|
||||
|
||||
assertSuccessorRegistrationSurvives(resolver, rendezvous, waiterB,
|
||||
"a raw backend-error turn must not evict B from the raw-scrape error branch");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aSupersededDoneFailedTurnMustNotEvictItsSuccessorsRegistration() {
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(new FakeHerdr()), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||
|
||||
var waiterA = rendezvous.open("term_a");
|
||||
var turnA = new CompletionResolver.InFlight(waiterA, null);
|
||||
assertTrue(rendezvous.resolve("term_a", "A replied"));
|
||||
rendezvous.close("term_a", waiterA);
|
||||
var waiterB = rendezvous.open("term_a");
|
||||
resolver.captureBaseline("term_a", new TurnToken("term_a", waiterB));
|
||||
|
||||
resolver.fail("term_a", turnA);
|
||||
|
||||
assertSuccessorRegistrationSurvives(resolver, rendezvous, waiterB,
|
||||
"a done failed turn must not evict B from fail()'s early return");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aSupersededTooFastBackendErrorTurnMustNotEvictItsSuccessorsRegistration() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ API Error: 400 invalid request body\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
long[] clock = {10_000_000_000L};
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none(), () -> clock[0]);
|
||||
|
||||
var waiterA = rendezvous.open("term_a");
|
||||
var turnA = new CompletionResolver.InFlight(waiterA, null, clock[0]);
|
||||
rendezvous.close("term_a", waiterA);
|
||||
var waiterB = rendezvous.open("term_a");
|
||||
resolver.captureBaseline("term_a", new TurnToken("term_a", waiterB));
|
||||
clock[0] += CompletionResolver.MIN_TURN_NANOS - 1;
|
||||
|
||||
resolver.resolve("term_a", turnA);
|
||||
|
||||
assertSuccessorRegistrationSurvives(resolver, rendezvous, waiterB,
|
||||
"a too-fast backend-error turn must not evict B from failTooFast()");
|
||||
}
|
||||
|
||||
private static void assertSuccessorRegistrationSurvives(CompletionResolver resolver, Rendezvous rendezvous,
|
||||
Object waiterB, String message) {
|
||||
CompletionResolver.InFlight afterA = resolver.inFlight("term_a");
|
||||
assertNotNull(afterA, message + " — a one-arg remove(target) would remove B");
|
||||
assertEquals(waiterB, afterA.waiter(), message + " — the surviving record must belong to B");
|
||||
assertTrue(rendezvous.resolve("term_a", "B replied"), message + " — B must still resolve normally");
|
||||
}
|
||||
|
||||
// --- CB-578 stage A: backend-exhausted classification ---------------------------------
|
||||
|
||||
@Test
|
||||
|
||||
@@ -1,16 +1,18 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.LoggerContext;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.AgentStatus;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.HerdrException;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.msg.TestTurnTokens;
|
||||
import dev.ltms.fleet.msg.TurnToken;
|
||||
import dev.ltms.fleet.testing.CapturedLog;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
@@ -493,23 +495,15 @@ class InjectorTest {
|
||||
void readinessGraceExpiryIsLogged() {
|
||||
// CB-562: the grace-expiry path used to clear the queue silently, so a message that never
|
||||
// reached the worker's pane surfaced elsewhere as an unrelated turn-stall failure. Assert the
|
||||
// expiry now names the real cause. (ListAppender capture pattern mirrors AuditLogTest.)
|
||||
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
|
||||
ch.qos.logback.classic.Logger injectorLog =
|
||||
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.setContext(ctx);
|
||||
appender.start();
|
||||
injectorLog.addAppender(appender);
|
||||
injectorLog.setLevel(Level.WARN);
|
||||
try {
|
||||
// expiry now names the real cause. (CapturedLog pattern mirrors AuditLogTest.)
|
||||
try (CapturedLog log = CapturedLog.at(Injector.class, Level.WARN)) {
|
||||
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
|
||||
});
|
||||
inj.enqueue(T, "task", TestTurnTokens.inert(T));
|
||||
|
||||
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
String warn = appender.list.stream()
|
||||
String warn = log.events().stream()
|
||||
.filter(e -> e.getLevel().equals(Level.WARN))
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.findFirst()
|
||||
@@ -518,8 +512,6 @@ class InjectorTest {
|
||||
assertTrue(warn.contains("never reached"), "the log names the real cause: " + warn);
|
||||
assertTrue(warn.contains("1 queued message"),
|
||||
"the log carries the failed message count: " + warn);
|
||||
} finally {
|
||||
injectorLog.detachAppender(appender);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -533,22 +525,14 @@ class InjectorTest {
|
||||
// count and a labelled configured budget, using literal numbers (240, 60), never
|
||||
// READINESS_GRACE_POLLS or POLL_INTERVAL_MILLIS, so the assertion can't silently track a
|
||||
// constant change instead of catching a real regression.
|
||||
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
|
||||
ch.qos.logback.classic.Logger injectorLog =
|
||||
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.setContext(ctx);
|
||||
appender.start();
|
||||
injectorLog.addAppender(appender);
|
||||
injectorLog.setLevel(Level.WARN);
|
||||
try {
|
||||
try (CapturedLog log = CapturedLog.at(Injector.class, Level.WARN)) {
|
||||
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
|
||||
});
|
||||
inj.enqueue(T, "task", TestTurnTokens.inert(T));
|
||||
|
||||
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
String warn = appender.list.stream()
|
||||
String warn = log.events().stream()
|
||||
.filter(e -> e.getLevel().equals(Level.WARN))
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.findFirst()
|
||||
@@ -557,8 +541,6 @@ class InjectorTest {
|
||||
"must print the measured poll count as a plain number: " + warn);
|
||||
assertTrue(warn.contains("configured=240 polls/60s"),
|
||||
"must print the configured budget, clearly labelled: " + warn);
|
||||
} finally {
|
||||
injectorLog.detachAppender(appender);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -584,22 +566,14 @@ class InjectorTest {
|
||||
return readings[i];
|
||||
};
|
||||
|
||||
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
|
||||
ch.qos.logback.classic.Logger injectorLog =
|
||||
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.setContext(ctx);
|
||||
appender.start();
|
||||
injectorLog.addAppender(appender);
|
||||
injectorLog.setLevel(Level.WARN);
|
||||
try {
|
||||
try (CapturedLog log = CapturedLog.at(Injector.class, Level.WARN)) {
|
||||
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
|
||||
}, stubClock);
|
||||
inj.enqueue(T, "task", TestTurnTokens.inert(T));
|
||||
|
||||
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
String warn = appender.list.stream()
|
||||
String warn = log.events().stream()
|
||||
.filter(e -> e.getLevel().equals(Level.WARN))
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.findFirst()
|
||||
@@ -609,8 +583,6 @@ class InjectorTest {
|
||||
assertFalse(warn.contains("elapsed=60000ms"), "must not print "
|
||||
+ "READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS (240 * 250 = 60000ms) as if it "
|
||||
+ "were the measured elapsed time: " + warn);
|
||||
} finally {
|
||||
injectorLog.detachAppender(appender);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -647,21 +619,13 @@ class InjectorTest {
|
||||
// CB-564: a vanished worker used to drop its queue with no log at all — the only trace was
|
||||
// whatever failed downstream (e.g. a caller's send timing out with no clue why). Assert the
|
||||
// drop itself now names the cause and the number of messages it failed.
|
||||
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
|
||||
ch.qos.logback.classic.Logger injectorLog =
|
||||
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.setContext(ctx);
|
||||
appender.start();
|
||||
injectorLog.addAppender(appender);
|
||||
injectorLog.setLevel(Level.WARN);
|
||||
try {
|
||||
try (CapturedLog log = CapturedLog.at(Injector.class, Level.WARN)) {
|
||||
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, _ -> {
|
||||
});
|
||||
inj.enqueue(T, "orphan", TestTurnTokens.inert(T));
|
||||
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
|
||||
|
||||
String warn = appender.list.stream()
|
||||
String warn = log.events().stream()
|
||||
.filter(e -> e.getLevel().equals(Level.WARN))
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.findFirst()
|
||||
@@ -669,8 +633,6 @@ class InjectorTest {
|
||||
assertTrue(warn.contains(T), "the log names the target terminal: " + warn);
|
||||
assertTrue(warn.contains("1 message"), "the log carries the failed message count: " + warn);
|
||||
assertTrue(warn.contains("worker gone"), "the log carries the real cause: " + warn);
|
||||
} finally {
|
||||
injectorLog.detachAppender(appender);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -706,4 +668,783 @@ class InjectorTest {
|
||||
ExecutionException ex = assertThrows(ExecutionException.class, f::get);
|
||||
assertInstanceOf(HerdrException.class, ex.getCause());
|
||||
}
|
||||
|
||||
/**
|
||||
* A {@link HerdrClient} that throws a non-{@link RuntimeException} {@link Error} from {@code
|
||||
* agent.prompt} instead of delegating — the fleetd #546 case: a stray non-RuntimeException
|
||||
* throwable (e.g. a {@code NoClassDefFoundError}, fleetd #413) escaping the send seam at
|
||||
* {@code Injector.java:383}. Records every call it sees itself, including the ones it throws
|
||||
* for, since the delegate's own recording is never reached for {@code agent.prompt} — so a test
|
||||
* can assert on exactly what this fake actually received.
|
||||
*/
|
||||
private static final class ErrorOnPrompt implements HerdrClient {
|
||||
private final FakeHerdr delegate;
|
||||
private final List<FakeHerdr.Call> calls = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||
|
||||
private ErrorOnPrompt(FakeHerdr delegate) {
|
||||
this.delegate = delegate;
|
||||
}
|
||||
|
||||
List<FakeHerdr.Call> calls() {
|
||||
return calls;
|
||||
}
|
||||
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) {
|
||||
calls.add(new FakeHerdr.Call(method, params));
|
||||
if (method.equals("agent.prompt")) {
|
||||
throw new AssertionError("simulated non-RuntimeException send failure (fleetd #546)");
|
||||
}
|
||||
return delegate.call(method, params);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
delegate.close();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Mirrors StatusPoller's own per-target {@code catch (Throwable)} (fleetd #538 / PR #543): the
|
||||
* production polling loop already swallows whatever escapes one target's round and comes back
|
||||
* for the next one. A unit test that calls {@code onStatus} directly (bypassing StatusPoller)
|
||||
* needs the same survival so it can observe what a SECOND round does, regardless of whether
|
||||
* fleetd #546's fix is present.
|
||||
*/
|
||||
private static void pollOnceSurviving(Injector inj, String target, AgentStatus status) {
|
||||
try {
|
||||
inj.onStatus(target, status);
|
||||
} catch (Throwable ignored) {
|
||||
// matches StatusPoller.loop's own catch (Throwable) added by PR #543
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void anErrorFromSendRemovesTheMessageAndMarksItAttempted() {
|
||||
// fleetd #546, acceptance test 1: an Error (not a RuntimeException) escaping the send seam
|
||||
// must still be caught and the message dropped from the queue, never left QUEUED at the
|
||||
// head. fleetd #551 updates what it is marked: an AssertionError from the send call itself
|
||||
// proves nothing about whether the text reached the pane, so it is recorded as ATTEMPTED
|
||||
// (uncertain), not a confident NOT_DELIVERED — see anErrorFromSendMarksItNotDeliveredOnlyForANotFoundCode
|
||||
// below for the one case that still gets NOT_DELIVERED.
|
||||
ErrorOnPrompt throwing = new ErrorOnPrompt(new FakeHerdr());
|
||||
Injector inj = new Injector(new AgentControl(throwing));
|
||||
Injector.Delivery delivery = inj.enqueue(T, "brief", TestTurnTokens.inert(T));
|
||||
|
||||
assertDoesNotThrow(() -> inj.onStatus(T, AgentStatus.IDLE),
|
||||
"fleetd #546: an Error from the send seam must be caught inside onStatus, not "
|
||||
+ "escape it");
|
||||
|
||||
assertTrue(delivery.completion().isCompletedExceptionally(),
|
||||
"the delivery's future must surface the send failure");
|
||||
assertEquals(Injector.Cancellation.ATTEMPTED, inj.cancel(delivery),
|
||||
"fleetd #551: the message must be dropped and marked ATTEMPTED (not a confident "
|
||||
+ "NOT_DELIVERED), and never left QUEUED at the head of the queue");
|
||||
}
|
||||
|
||||
@Test
|
||||
void anErrorFromSendDoesNotRedeliverOnASecondRound() {
|
||||
// fleetd #546, acceptance test 2 — the test this ticket exists for. Before the fix, an
|
||||
// Error at the send seam left the message QUEUED (Injector.java:378 peeks, not polls), so a
|
||||
// second onStatus round re-entered the same try and sent the same text again: the member's
|
||||
// pane got the same brief typed into it twice.
|
||||
ErrorOnPrompt throwing = new ErrorOnPrompt(new FakeHerdr());
|
||||
Injector inj = new Injector(new AgentControl(throwing));
|
||||
inj.enqueue(T, "brief", TestTurnTokens.inert(T));
|
||||
|
||||
pollOnceSurviving(inj, T, AgentStatus.IDLE);
|
||||
pollOnceSurviving(inj, T, AgentStatus.IDLE);
|
||||
|
||||
long promptCalls = throwing.calls().stream().filter(c -> c.method().equals("agent.prompt")).count();
|
||||
assertEquals(1, promptCalls, "fleetd #546: the poisoned text must be sent exactly once — a "
|
||||
+ "second onStatus round must not re-enter send for the same message");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aHerdrExceptionFromSendStillSurfacesButNowReportsAttempted() {
|
||||
// fleetd #546, acceptance test 3 (control): HerdrException extends RuntimeException, so it
|
||||
// was already caught before #546's widening. The send-failure/surface behaviour is
|
||||
// unchanged by #551 — still dropped, still surfaced to the caller — but fleetd #551
|
||||
// deliberately changes WHAT it is marked: "send_failed" is not a herdr `*_not_found` code,
|
||||
// so this is exactly the response-half case #551 exists to fix. See
|
||||
// aHerdrExceptionAfterThePasteIsNeverRecordedAsConfidentlyNotDelivered for the acceptance
|
||||
// test this ticket was filed for.
|
||||
FakeHerdr failing = new FakeHerdr().agentSendFailsWith("send_failed");
|
||||
Injector inj = new Injector(new AgentControl(failing));
|
||||
Injector.Delivery delivery = inj.enqueue(T, "boom", TestTurnTokens.inert(T));
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
assertTrue(delivery.completion().isCompletedExceptionally(),
|
||||
"a HerdrException at the send seam must still surface to the caller, unchanged by "
|
||||
+ "fleetd #546's widening");
|
||||
assertEquals(Injector.Cancellation.ATTEMPTED, inj.cancel(delivery),
|
||||
"fleetd #551: a non-*_not_found HerdrException must now report ATTEMPTED, not a "
|
||||
+ "confident NOT_DELIVERED");
|
||||
}
|
||||
|
||||
// --- fleetd #551: the injector records the delivery attempt BEFORE the irreversible send, not
|
||||
// after, so a failure in the send's response window can never be written down as a confident
|
||||
// NOT_DELIVERED for text that may already be sitting in the worker's pane. ---
|
||||
|
||||
@Test
|
||||
void aHerdrExceptionAfterThePasteIsNeverRecordedAsConfidentlyNotDelivered() {
|
||||
// fleetd #551, the ticket's acceptance test 1, and comment 17037's added test ("a test
|
||||
// pinning that a caller cannot receive a confident NOT_DELIVERED for a message that reached
|
||||
// the paste"): a HerdrException thrown from the response half of the send seam — herdr
|
||||
// replied, with an error, so it definitely processed the request — must not leave a record
|
||||
// claiming the text was never delivered. Red before the fix: on main at ba2f4d1 this
|
||||
// asserted (and got) Injector.Cancellation.NOT_DELIVERED.
|
||||
FakeHerdr failing = new FakeHerdr().agentSendFailsWith("send_failed");
|
||||
Injector inj = new Injector(new AgentControl(failing));
|
||||
Injector.Delivery delivery = inj.enqueue(T, "brief", TestTurnTokens.inert(T));
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
assertTrue(delivery.completion().isCompletedExceptionally(),
|
||||
"the caller must still see the send failure");
|
||||
assertEquals(Injector.Cancellation.ATTEMPTED, inj.cancel(delivery),
|
||||
"fleetd #551: a HerdrException from the response half of the send seam must not be "
|
||||
+ "recorded as a confident NOT_DELIVERED — the text may already be sitting "
|
||||
+ "in the pane");
|
||||
}
|
||||
|
||||
@Test
|
||||
void ordinarySuccessStillReportsDeliveredExactlyOnce() {
|
||||
// fleetd #551, the ticket's acceptance test 2: the record-before-send reorder must not
|
||||
// change the ordinary success path — exactly one send, and the caller-facing Cancellation
|
||||
// for it is still DELIVERED, never left at the new ATTEMPTED value.
|
||||
Injector.Delivery delivery = injector.enqueue(T, "hello", TestTurnTokens.inert(T));
|
||||
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
assertEquals(List.of("hello"), sent(), "the text must be sent exactly once");
|
||||
assertTrue(delivery.completion().isDone() && !delivery.completion().isCompletedExceptionally(),
|
||||
"the ordinary success path must still complete normally");
|
||||
assertEquals(Injector.Cancellation.DELIVERED, injector.cancel(delivery),
|
||||
"fleetd #551: the ordinary success path must still report DELIVERED, unaffected by "
|
||||
+ "the record-before-send reorder");
|
||||
}
|
||||
|
||||
@Test
|
||||
void transportDownStillSurfacesTheErrorAndDropsTheEntry() {
|
||||
// fleetd #551, the ticket's acceptance test 3: the ordinary transport-down path (herdr
|
||||
// unreachable — a plain HerdrException with no error code) must be unaffected by the #551
|
||||
// reorder — the failure still surfaces to the caller, and the entry is not left sitting in
|
||||
// the queue for a second onStatus round to resend.
|
||||
FakeHerdr unreachable = new FakeHerdr().healthy(false);
|
||||
Injector inj = new Injector(new AgentControl(unreachable));
|
||||
Injector.Delivery delivery = inj.enqueue(T, "brief", TestTurnTokens.inert(T));
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
assertTrue(delivery.completion().isCompletedExceptionally(),
|
||||
"a transport-down failure must still surface to the caller");
|
||||
assertTrue(inj.activeTargets().isEmpty(),
|
||||
"the entry must not be left queued for a second round to resend");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aNotFoundCodeFromSendStillReportsAConfidentNotDelivered() {
|
||||
// fleetd #551 (comment 17037, point 3): a herdr `*_not_found` error means the target
|
||||
// pane/agent does not exist at all, so nothing could have been pasted anywhere — the
|
||||
// codebase already treats this family as a confirmed absence everywhere else (StatusPoller,
|
||||
// AgentControl's own retry). #551's new ATTEMPTED state must not swallow this case: it stays
|
||||
// a confident NOT_DELIVERED.
|
||||
FakeHerdr failing = new FakeHerdr().agentSendFailsWith("agent_not_found");
|
||||
Injector inj = new Injector(new AgentControl(failing));
|
||||
Injector.Delivery delivery = inj.enqueue(T, "brief", TestTurnTokens.inert(T));
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
assertTrue(delivery.completion().isCompletedExceptionally(),
|
||||
"an agent_not_found failure must still surface to the caller");
|
||||
assertEquals(Injector.Cancellation.NOT_DELIVERED, inj.cancel(delivery),
|
||||
"fleetd #551: a herdr *_not_found error is a confirmed absence, not merely "
|
||||
+ "inconclusive — it must stay NOT_DELIVERED, not the new ATTEMPTED");
|
||||
}
|
||||
|
||||
// --- fleetd #553: a throwable from any listener callback in onStatus must not skip the
|
||||
// delivered-future completion. The whole post-monitor region is now wrapped in try/finally. ---
|
||||
|
||||
@Test
|
||||
void theOriginalThrowableFromAListenerStillEscapesOnStatus() {
|
||||
// fleetd #553's control: the entire point of the finally backstop is that it completes a
|
||||
// future WITHOUT swallowing whatever escaped. A fix built the wrong way (e.g. catching
|
||||
// Throwable in the finally, or wrapping the region in try/catch instead of try/finally)
|
||||
// could make every other test in this group pass while still converting a loud listener
|
||||
// bug into a silent one — which the ticket calls a worse outcome than the bug itself. This
|
||||
// is deliberately its own test, not folded into another one's assertion.
|
||||
RuntimeException boom = new RuntimeException("fleetd #553 control");
|
||||
TurnListener throwing = new TurnListener() {
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
throw boom;
|
||||
}
|
||||
};
|
||||
Injector inj = new Injector(new AgentControl(herdr), throwing);
|
||||
inj.enqueue(T, "only", TestTurnTokens.inert(T));
|
||||
inj.onStatus(T, AgentStatus.IDLE); // deliver
|
||||
inj.onStatus(T, AgentStatus.WORKING); // pickup
|
||||
|
||||
RuntimeException thrown = assertThrows(RuntimeException.class,
|
||||
() -> inj.onStatus(T, AgentStatus.IDLE),
|
||||
"onStatus must still propagate the listener's own throwable, unmodified");
|
||||
assertSame(boom, thrown, "must be the EXACT throwable, not a wrapper or a different instance");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aRuntimeExceptionFromOnTurnCompleteStillCompletesTheNextDelivery() {
|
||||
// fleetd #553, acceptance test 1: onTurnComplete fires for "first"'s completed turn INSIDE
|
||||
// the same onStatus call that then peeks and delivers "second" — the turnCompleted check
|
||||
// (Injector.java, inside the released-pickup branch) clears awaitingCompletion just before
|
||||
// the delivery guard right below it is evaluated, so both happen in one round when there is
|
||||
// no post-turn action. Before fleetd #553, a RuntimeException thrown here unwound straight
|
||||
// out of onStatus and skipped the `if (sent != null)` block, leaving "second"'s future
|
||||
// pending forever even though it was already off the queue, marked DELIVERED, and typed
|
||||
// into the pane.
|
||||
RuntimeException boom = new RuntimeException("boom from onTurnComplete");
|
||||
TurnListener throwing = new TurnListener() {
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
throw boom;
|
||||
}
|
||||
};
|
||||
Injector inj = new Injector(new AgentControl(herdr), throwing);
|
||||
inj.enqueue(T, "first", TestTurnTokens.inert(T));
|
||||
Injector.Delivery second = inj.enqueue(T, "second", TestTurnTokens.inert(T));
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE); // delivers "first"
|
||||
inj.onStatus(T, AgentStatus.WORKING); // picked up
|
||||
|
||||
RuntimeException thrown = assertThrows(RuntimeException.class,
|
||||
() -> inj.onStatus(T, AgentStatus.IDLE),
|
||||
"the throwable from onTurnComplete must still escape onStatus");
|
||||
assertSame(boom, thrown);
|
||||
|
||||
assertEquals(List.of("first", "second"), sent(),
|
||||
"fleetd #553: \"second\" must still be sent even though onTurnComplete threw for "
|
||||
+ "\"first\"'s completion");
|
||||
assertTrue(second.completion().isDone() && !second.completion().isCompletedExceptionally(),
|
||||
"fleetd #553: \"second\"'s delivered future must still complete normally despite "
|
||||
+ "the throw");
|
||||
}
|
||||
|
||||
private static final int READINESS_GRACE_SAMPLES = 240; // Injector.READINESS_GRACE_POLLS is private
|
||||
|
||||
@Test
|
||||
void aRuntimeExceptionFromForgetStillCompletesTheReadinessFailureFutures() {
|
||||
// fleetd #553: forget.accept fires as part of the CB-114 readiness-grace path (production
|
||||
// wires it to presence::forget). Before fleetd #553, that block called forget.accept
|
||||
// BEFORE completing the queued messages' futures, so a RuntimeException from forget left
|
||||
// every one of them pending forever even though they were already marked NOT_DELIVERED and
|
||||
// dropped from the queue. fleetd #553 completes those futures first, so forget.accept
|
||||
// throwing afterward can no longer un-complete them.
|
||||
RuntimeException boom = new RuntimeException("boom from forget");
|
||||
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
|
||||
throw boom;
|
||||
});
|
||||
Injector.Delivery delivery = inj.enqueue(T, "task", TestTurnTokens.inert(T));
|
||||
|
||||
// Build the not-ready streak up to (but not past) the grace threshold.
|
||||
for (int i = 0; i < READINESS_GRACE_SAMPLES - 1; i++) inj.onStatus(T, AgentStatus.IDLE);
|
||||
assertFalse(delivery.completion().isDone(), "must not be resolved before the grace expires");
|
||||
|
||||
RuntimeException thrown = assertThrows(RuntimeException.class,
|
||||
() -> inj.onStatus(T, AgentStatus.IDLE),
|
||||
"the throwable from forget.accept must still escape onStatus");
|
||||
assertSame(boom, thrown);
|
||||
|
||||
assertTrue(delivery.completion().isCompletedExceptionally(),
|
||||
"fleetd #553: the queued message's future must still be completed even though "
|
||||
+ "forget.accept threw");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aRuntimeExceptionFromOnTurnFailedStillLeavesTheReadinessFailureFuturesCompleted() {
|
||||
// fleetd #553, acceptance test: same readiness-grace path as the forget test above, but the
|
||||
// throw comes from onTurnFailed instead. NOTE: onTurnFailed is the LAST statement in this
|
||||
// block (both before and after this ticket), so the queued messages' futures are already
|
||||
// completed by the time it runs, regardless of this fix — this test cannot be made to fail
|
||||
// against the pre-fleetd-553 code the way the onTurnComplete/forget/postAction tests can; I
|
||||
// could not find a construction where onTurnFailed's own throw is what stands between a
|
||||
// future and its completion (flagged to the lead via fleet_ask; no reply arrived before the
|
||||
// ~55s window closed, so recorded here instead). It still pins a real invariant this ticket
|
||||
// cares about — an ordinary RuntimeException from this callback must not un-complete a
|
||||
// future that was already decided — so it stays as a regression lock, not a bug-fix proof.
|
||||
RuntimeException boom = new RuntimeException("boom from onTurnFailed");
|
||||
TurnListener throwing = new TurnListener() {
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onTurnFailed(String target) {
|
||||
throw boom;
|
||||
}
|
||||
};
|
||||
Injector inj = new Injector(new AgentControl(herdr), throwing, _ -> false, _ -> {
|
||||
});
|
||||
Injector.Delivery delivery = inj.enqueue(T, "task", TestTurnTokens.inert(T));
|
||||
|
||||
for (int i = 0; i < READINESS_GRACE_SAMPLES - 1; i++) inj.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
RuntimeException thrown = assertThrows(RuntimeException.class,
|
||||
() -> inj.onStatus(T, AgentStatus.IDLE),
|
||||
"the throwable from onTurnFailed must still escape onStatus");
|
||||
assertSame(boom, thrown);
|
||||
|
||||
assertTrue(delivery.completion().isCompletedExceptionally(),
|
||||
"the queued message's future must remain completed despite onTurnFailed throwing");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aRuntimeExceptionFromOnTurnCompleteWithPostActionStillUnwedgesTheTarget() {
|
||||
// fleetd #553, acceptance test: before this ticket, a RuntimeException from
|
||||
// onTurnCompleteWithPostAction skipped the `t.postTurnPending = false` reset that used to
|
||||
// sit right after it (with nothing guarding it), permanently wedging the target —
|
||||
// postTurnPending stayed true forever, so the delivery guard never passed again and
|
||||
// "second" (already queued) was never delivered, no matter how many further onStatus
|
||||
// rounds ran. fleetd #553 wraps that call so the reset always runs.
|
||||
RuntimeException boom = new RuntimeException("boom from onTurnCompleteWithPostAction");
|
||||
class ThrowingPostTurn implements TurnListener {
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean hasPostTurnAction(String target) {
|
||||
return true;
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean onTurnCompleteWithPostAction(String target) {
|
||||
throw boom;
|
||||
}
|
||||
}
|
||||
Injector inj = new Injector(new AgentControl(herdr), new ThrowingPostTurn());
|
||||
inj.enqueue(T, "first", TestTurnTokens.inert(T));
|
||||
Injector.Delivery second = inj.enqueue(T, "second", TestTurnTokens.inert(T));
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE); // delivers "first"
|
||||
inj.onStatus(T, AgentStatus.WORKING); // picked up
|
||||
|
||||
RuntimeException thrown = assertThrows(RuntimeException.class,
|
||||
() -> inj.onStatus(T, AgentStatus.IDLE), // "first"'s turn completes; post-action throws
|
||||
"the throwable from onTurnCompleteWithPostAction must still escape onStatus");
|
||||
assertSame(boom, thrown);
|
||||
assertEquals(List.of("first"), sent(), "\"second\" must not be sent in the SAME round as the throw");
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE); // a later, unrelated round
|
||||
assertEquals(List.of("first", "second"), sent(),
|
||||
"fleetd #553: the target must not stay wedged — \"second\" must still be delivered "
|
||||
+ "once postTurnPending is reset despite the earlier throw");
|
||||
assertTrue(second.completion().isDone() && !second.completion().isCompletedExceptionally());
|
||||
}
|
||||
|
||||
/**
|
||||
* A {@link HerdrClient} that throws a non-{@link RuntimeException} {@link Error} from {@code
|
||||
* agent.send_keys} instead of delegating — the fleetd #553 case for the resubmit nudge: its own
|
||||
* catch (Injector.java, the first block after the monitor) was {@code RuntimeException}-only,
|
||||
* the same one-class-too-narrow shape #546 fixed at the send seam. Records every call it sees
|
||||
* itself, mirroring {@code ErrorOnPrompt} above, since the delegate's own recording is never
|
||||
* reached for {@code agent.send_keys}.
|
||||
*/
|
||||
private static final class ErrorOnSendKeys implements HerdrClient {
|
||||
private final FakeHerdr delegate;
|
||||
private final List<FakeHerdr.Call> calls = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||
|
||||
private ErrorOnSendKeys(FakeHerdr delegate) {
|
||||
this.delegate = delegate;
|
||||
}
|
||||
|
||||
List<FakeHerdr.Call> calls() {
|
||||
return calls;
|
||||
}
|
||||
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) {
|
||||
calls.add(new FakeHerdr.Call(method, params));
|
||||
if (method.equals("agent.send_keys")) {
|
||||
throw new AssertionError("simulated non-RuntimeException resubmit failure (fleetd #553)");
|
||||
}
|
||||
return delegate.call(method, params);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
delegate.close();
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void anErrorFromTheResubmitNudgeDoesNotPreventTheSentCompletion() {
|
||||
// fleetd #553, acceptance test: the resubmit nudge's own catch was RuntimeException-only —
|
||||
// the same one-class-too-narrow shape #546 fixed at the send seam (now :382). It sits FIRST
|
||||
// after the monitor, so before this ticket an Error escaping it skipped every block below in
|
||||
// THAT round, including any `sent` completion that round might otherwise produce. Widened to
|
||||
// Throwable (matching #549's own widening) so it can no longer escape onStatus at all — the
|
||||
// delivery that already completed at send time, and the worker's later pickup, are both
|
||||
// unaffected by the Error in between.
|
||||
ErrorOnSendKeys throwing = new ErrorOnSendKeys(new FakeHerdr());
|
||||
Injector inj = new Injector(new AgentControl(throwing));
|
||||
Injector.Delivery delivery = inj.enqueue(T, "task", TestTurnTokens.inert(T));
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE); // delivers "task"; awaiting pickup
|
||||
assertTrue(delivery.completion().isDone() && !delivery.completion().isCompletedExceptionally(),
|
||||
"the delivery completes at send time, before any resubmit nudge is attempted");
|
||||
|
||||
assertDoesNotThrow(() -> inj.onStatus(T, AgentStatus.IDLE), // still idle -> resubmit; Error thrown+caught
|
||||
"fleetd #553: an Error from the resubmit nudge must be caught inside onStatus, not "
|
||||
+ "escape it");
|
||||
|
||||
inj.onStatus(T, AgentStatus.WORKING); // the worker's real pickup must still resolve cleanly
|
||||
long promptCalls = throwing.calls().stream().filter(c -> c.method().equals("agent.prompt")).count();
|
||||
assertEquals(1, promptCalls,
|
||||
"sent exactly once, unaffected by the resubmit Error in between");
|
||||
}
|
||||
|
||||
@Test
|
||||
void ordinarySuccessStillCompletesExactlyOnceAfterOnDelivered() {
|
||||
// Acceptance: the ordinary success path is unchanged by fleetd #553 — the future completes
|
||||
// normally exactly once, and onDelivered still runs before it (CB-115's pane baseline).
|
||||
List<String> events = new ArrayList<>();
|
||||
TurnListener listener = new TurnListener() {
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onDelivered(String target, dev.ltms.fleet.msg.TurnToken token) {
|
||||
events.add("onDelivered");
|
||||
}
|
||||
};
|
||||
Injector inj = new Injector(new AgentControl(herdr), listener);
|
||||
CompletableFuture<Void> f = inj.enqueue(T, "hello", TestTurnTokens.inert(T)).completion();
|
||||
f.whenComplete((v, ex) -> events.add("completed"));
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
assertEquals(List.of("onDelivered", "completed"), events,
|
||||
"onDelivered must still run before the future completes, unchanged by fleetd #553");
|
||||
assertTrue(f.isDone() && !f.isCompletedExceptionally());
|
||||
}
|
||||
|
||||
// Acceptance: the ordinary failure path (a HerdrException at the send seam still completes the
|
||||
// future exceptionally with that exception) is already covered, unchanged, by the existing
|
||||
// aHerdrExceptionFromSendStillProducesNotDeliveredUnchanged test above (fleetd #546) — its
|
||||
// codepath is untouched by fleetd #553's try/finally, since sendError != null is set before the
|
||||
// try block begins and that branch never threw to begin with.
|
||||
|
||||
// --- fleetd #553, sharpened acceptance (ticket comments 16870/16884/16890/16903): completing
|
||||
// sent.delivered() is only HALF the job. See the two tests below. ---
|
||||
|
||||
@Test
|
||||
void aRuntimeExceptionFromOnTurnCompleteStillRegistersTheRendezvousWaiterForTheNextDelivery() {
|
||||
// fleetd #553's real invariant, not just the completion of sent.delivered(). There are TWO
|
||||
// futures at stake: sent.delivered() (the delivery future) and sent.token().waiter() (the
|
||||
// rendezvous waiter a blocking fleet_send actually waits on for the worker's ANSWER). The
|
||||
// waiter is registered only by turnListener.onDelivered() — production wires this to
|
||||
// CompletionResolver.captureBaseline, which does the inFlight.put() — and that call lives
|
||||
// inside the very `if (sent != null)` block a plain "complete the future" backstop does not
|
||||
// reach. A finally that only completes sent.delivered() converts a hang into a HANG WITH A
|
||||
// SUCCESS RECEIPT: the caller is told the send landed, then waits out its full timeout for
|
||||
// an answer that can never resolve, because CompletionResolver.resolve (:301-308) finds no
|
||||
// inFlight entry and returns silently. This test is red on that half-fix, and green only
|
||||
// once the finally also runs onDelivered() when the normal block never got the chance.
|
||||
//
|
||||
// The scenario: onTurnComplete throws while completing "first"'s turn, in the SAME onStatus
|
||||
// round that (per Injector.java :361-391) then peeks and delivers "second" — the normal
|
||||
// case, not a corner, since the :365 assignment is what lets the :376 delivery guard pass.
|
||||
// turnListener.onTurnComplete runs BEFORE the `if (sent != null)` block, so the throw here
|
||||
// means onDelivered for "second" is never reached on the normal path — only the finally
|
||||
// backstop can register it.
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(new FakeHerdr()), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||
RuntimeException boom = new RuntimeException("boom from onTurnComplete");
|
||||
TurnListener throwingOnTurnCompleteOnly = new TurnListener() {
|
||||
@Override
|
||||
public void onDelivered(String target, TurnToken token) {
|
||||
resolver.onDelivered(target, token); // the real registration, exactly as production wires it
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
throw boom; // "first"'s completion callback, thrown before "second" is delivered
|
||||
}
|
||||
};
|
||||
Injector inj = new Injector(new AgentControl(herdr), throwingOnTurnCompleteOnly);
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open("session-a");
|
||||
TurnToken secondToken = new TurnToken(T, waiter);
|
||||
inj.enqueue(T, "first", TestTurnTokens.inert(T));
|
||||
Injector.Delivery second = inj.enqueue(T, "second", secondToken);
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE); // delivers "first"
|
||||
inj.onStatus(T, AgentStatus.WORKING); // picked up
|
||||
|
||||
RuntimeException thrown = assertThrows(RuntimeException.class,
|
||||
() -> inj.onStatus(T, AgentStatus.IDLE), // "first" completes (throws); "second" delivered
|
||||
"the original throwable from onTurnComplete must still escape onStatus");
|
||||
assertSame(boom, thrown, "must be the EXACT throwable, not a wrapper or a different instance");
|
||||
|
||||
assertEquals(List.of("first", "second"), sent(),
|
||||
"\"second\" must still be sent even though onTurnComplete threw for \"first\"");
|
||||
assertTrue(second.completion().isDone() && !second.completion().isCompletedExceptionally(),
|
||||
"\"second\"'s delivery future must still complete despite the throw");
|
||||
|
||||
CompletionResolver.InFlight inFlight = resolver.inFlight(T);
|
||||
assertNotNull(inFlight,
|
||||
"fleetd #553: the finally must register the new turn's waiter (via onDelivered), not "
|
||||
+ "just complete its delivery future — otherwise a blocking fleet_send is told "
|
||||
+ "its message landed and then waits out the full timeout for an answer that "
|
||||
+ "can never resolve");
|
||||
assertSame(waiter, inFlight.waiter(),
|
||||
"the registered waiter must be exactly \"second\"'s waiter, not some other one");
|
||||
}
|
||||
|
||||
@Test
|
||||
void anOnDeliveredThrowAfterItsOwnRegistrationDoesNotRunASecondTime() {
|
||||
// fleetd #553 comment 16903: `sentHandled` must be set to true BEFORE onDelivered() runs,
|
||||
// not after. Setting it after would mean a throw from onDelivered() PART WAY THROUGH — e.g.
|
||||
// after CompletionResolver's own captureBaseline has already done its inFlight.put() — still
|
||||
// leaves the flag false, so the finally backstop (seeing "not handled") calls onDelivered() a
|
||||
// SECOND time. That second captureBaseline runs later, over a pane that may already have
|
||||
// absorbed this turn's output, and CompletionResolver.resolve's own baseline-equals-tail
|
||||
// suppression (:364-369, which deliberately keeps the in-flight record on a match) can then
|
||||
// drop every future completion for the turn, permanently. This assertion is red on a
|
||||
// `sentHandled = true` placed AFTER the onDelivered() call (onDelivered runs twice) and green
|
||||
// on the correct placement (runs exactly once) — regardless of what onDelivered itself did.
|
||||
AtomicInteger onDeliveredCalls = new AtomicInteger();
|
||||
RuntimeException boom = new RuntimeException("boom from onDelivered, after its own registration ran");
|
||||
TurnListener listener = new TurnListener() {
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onDelivered(String target, TurnToken token) {
|
||||
onDeliveredCalls.incrementAndGet(); // stand-in for CompletionResolver's inFlight.put
|
||||
throw boom; // then fail, as if a LATER step inside onDelivered blew up
|
||||
}
|
||||
};
|
||||
Injector inj = new Injector(new AgentControl(herdr), listener);
|
||||
inj.enqueue(T, "task", TestTurnTokens.inert(T));
|
||||
|
||||
RuntimeException thrown = assertThrows(RuntimeException.class,
|
||||
() -> inj.onStatus(T, AgentStatus.IDLE), // delivers "task"; onDelivered throws
|
||||
"the original throwable from onDelivered must still escape onStatus");
|
||||
assertSame(boom, thrown);
|
||||
|
||||
assertEquals(1, onDeliveredCalls.get(),
|
||||
"fleetd #553: onDelivered must be called exactly once — a `sentHandled` flag set "
|
||||
+ "AFTER the call (rather than before) would leave it false here and the "
|
||||
+ "finally backstop would call onDelivered a second time");
|
||||
}
|
||||
|
||||
@Test
|
||||
void anOnDeliveredThrowOnTheNormalPathStillCompletesTheDeliveryFuture() {
|
||||
// fleetd #553, lead review of PR #557 (ticket comment 16916): `sentHandled` must guard ONLY
|
||||
// the onDelivered RE-CALL in the finally, never the future completion alongside it. Gating
|
||||
// BOTH behind `!sentHandled` (the shape this test is red against) misses the one path the
|
||||
// flag's own correct placement creates: `sentHandled` is set to true FIRST, before the
|
||||
// onDelivered() call, inside the `if (sent != null)` block above (see the previous test —
|
||||
// that placement is right and must not change). So when onDelivered() itself throws on that
|
||||
// NORMAL path, `sentHandled` already reads true by the time control reaches the finally, and
|
||||
// a `sent != null && !sentHandled` guard around the WHOLE recovery — completion included —
|
||||
// skips it entirely. The message was typed into the target's pane and taken off the queue
|
||||
// inside the monitor, same as any other delivery, so its future is left pending forever: the
|
||||
// exact defect this ticket exists to close, just reached from a different throwing call.
|
||||
//
|
||||
// The fix splits the one flag's two jobs: `!sentHandled` keeps gating only the onDelivered
|
||||
// call (so the exactly-once guarantee in the test above still holds — CompletableFuture.
|
||||
// complete/completeExceptionally are idempotent, so completing unconditionally here is a
|
||||
// no-op on the ordinary path, where the `if (sent != null)` block already completed it.
|
||||
RuntimeException boom = new RuntimeException("boom from onDelivered on the normal path");
|
||||
TurnListener listener = new TurnListener() {
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onDelivered(String target, TurnToken token) {
|
||||
throw boom;
|
||||
}
|
||||
};
|
||||
Injector inj = new Injector(new AgentControl(herdr), listener);
|
||||
CompletableFuture<Void> delivered = inj.enqueue(T, "task", TestTurnTokens.inert(T)).completion();
|
||||
|
||||
RuntimeException thrown = assertThrows(RuntimeException.class,
|
||||
() -> inj.onStatus(T, AgentStatus.IDLE), // delivers "task"; onDelivered throws
|
||||
"the original throwable from onDelivered must still escape onStatus");
|
||||
assertSame(boom, thrown);
|
||||
|
||||
assertTrue(delivered.isDone(),
|
||||
"fleetd #553: \"task\" was actually delivered — typed into the pane and taken off "
|
||||
+ "the queue inside the monitor — so its delivery future must be completed on "
|
||||
+ "every path out of onStatus, including the one where onDelivered itself is "
|
||||
+ "what threw. Leaving it pending here is a hang, not a fix");
|
||||
}
|
||||
|
||||
// --- fleetd #556: registration is the Injector's own invariant now, structural rather than
|
||||
// delegated to a TurnListener callback that can throw. #553 could only make the ONE reachable
|
||||
// listener behave; this closes the shape itself. ---
|
||||
|
||||
@Test
|
||||
void aTurnListenerThatThrowsFromEveryCallbackStillLeavesTheDeliveredTurnRegisteredAndResolvable() {
|
||||
// fleetd #556 acceptance criterion 1 (the ticket's own, restated by comment 16966 as the
|
||||
// replacement for #561's order-dependent test): a TurnListener that throws from EVERY
|
||||
// callback it implements — nothing about it is safe to lean on — must still leave a
|
||||
// delivered turn registered, with its waiter still resolvable through the ordinary path.
|
||||
// Before this ticket, registration lived INSIDE onDelivered, one of the callbacks that
|
||||
// throws here — this test is red against that shape, because the only listener callback
|
||||
// this specific delivery scenario exercises (a first delivery, no previous turn to
|
||||
// complete) is exactly the one that throws.
|
||||
RuntimeException boom = new RuntimeException("fleetd #556: throws from every callback");
|
||||
TurnListener throwsEverywhere = new TurnListener() {
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
throw boom;
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean hasPostTurnAction(String target) {
|
||||
throw boom;
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean onTurnCompleteWithPostAction(String target) {
|
||||
throw boom;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onTurnFailed(String target) {
|
||||
throw boom;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onTurnFailed(String target, String reason) {
|
||||
throw boom;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onDelivered(String target, TurnToken token) {
|
||||
throw boom;
|
||||
}
|
||||
};
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver completion = new CompletionResolver(new AgentControl(new FakeHerdr()), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||
|
||||
Injector inj = new Injector(new AgentControl(herdr), throwsEverywhere, _ -> true, _ -> {
|
||||
}, completion::register);
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(T);
|
||||
inj.enqueue(T, "brief", new TurnToken(T, waiter));
|
||||
|
||||
RuntimeException thrown = assertThrows(RuntimeException.class,
|
||||
() -> inj.onStatus(T, AgentStatus.IDLE), // delivers "brief"; onDelivered throws
|
||||
"the throwing listener's own bug must still escape onStatus — a listener bug stays "
|
||||
+ "loud (fleetd #553's own guarantee, unchanged by this ticket)");
|
||||
assertSame(boom, thrown);
|
||||
|
||||
CompletionResolver.InFlight inFlight = completion.inFlight(T);
|
||||
assertNotNull(inFlight,
|
||||
"fleetd #556: the delivered turn must be registered even though the ONLY "
|
||||
+ "TurnListener wired throws from every single callback, including "
|
||||
+ "onDelivered — registration must not depend on that call succeeding");
|
||||
assertSame(waiter, inFlight.waiter(),
|
||||
"the registered entry must carry the exact waiter this turn's send opened");
|
||||
|
||||
assertFalse(waiter.isDone(), "sanity: nothing has resolved this waiter yet");
|
||||
assertTrue(rendezvous.resolve(T, "hello"),
|
||||
"the waiter registrar.register captured must be the real, live rendezvous waiter — "
|
||||
+ "still resolvable through the ordinary fleet_reply path, unaffected by the "
|
||||
+ "throwing listener");
|
||||
assertEquals(Rendezvous.Kind.REPLY, waiter.getNow(null).kind());
|
||||
}
|
||||
|
||||
@Test
|
||||
void theNewTurnsRegistrationRunsAfterThePreviousTurnsCompletionIsRead() {
|
||||
// fleetd #556 acceptance criterion 2 (CB-116 guard, ticket body + comment 16966): the
|
||||
// ordering constraint from fleetd #553 must survive this redesign — onTurnComplete reads
|
||||
// CompletionResolver's inFlight entry for the PREVIOUS turn, so the new turn's
|
||||
// registrar.register() call must not run before that read, or a redesign trades this
|
||||
// ticket's defect for the CB-116 cross-turn stale reply. Pinned here as a call-ORDER
|
||||
// assertion so a future redesign that hoists registrar.register() earlier (e.g. "for
|
||||
// simplicity", ahead of the turnCompleted block) goes red immediately, before it can ever
|
||||
// reach a real cross-turn scrape race.
|
||||
List<String> events = new ArrayList<>();
|
||||
TurnListener listener = new TurnListener() {
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
events.add("onTurnComplete:" + target);
|
||||
}
|
||||
};
|
||||
TurnRegistrar registrar = (target, token) -> events.add("register:" + target);
|
||||
|
||||
Injector inj = new Injector(new AgentControl(herdr), listener, _ -> true, _ -> {
|
||||
}, registrar);
|
||||
inj.enqueue(T, "first", TestTurnTokens.inert(T));
|
||||
inj.enqueue(T, "second", TestTurnTokens.inert(T));
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE); // delivers "first" -> register:term_a (not asserted below)
|
||||
inj.onStatus(T, AgentStatus.WORKING); // picked up
|
||||
events.clear();
|
||||
inj.onStatus(T, AgentStatus.IDLE); // "first" completes AND "second" delivers, same cycle
|
||||
|
||||
assertEquals(List.of("onTurnComplete:" + T, "register:" + T), events,
|
||||
"onTurnComplete for the previous turn (\"first\") must run before register for the "
|
||||
+ "new turn (\"second\") within the same onStatus cycle — reversing this order "
|
||||
+ "is the CB-116 cross-turn stale reply fleetd #553 fixed, and this redesign "
|
||||
+ "must not reopen it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aRuntimeExceptionFromOnTurnCompleteStillLeavesTheNextDeliveryRegisteredOnTheRecoveryPath() {
|
||||
// fleetd #556 rework (PR #566, comment 17009): there are TWO registrar.register call sites
|
||||
// in the delivery method — the ordinary path inside `if (sent != null)`, and the fleetd
|
||||
// #553 `finally` backstop, reached only when an earlier block (here, onTurnComplete for the
|
||||
// PREVIOUS turn) throws before the ordinary path ever runs. The existing #553 regression
|
||||
// test for this exact scenario
|
||||
// (aRuntimeExceptionFromOnTurnCompleteStillCompletesTheNextDelivery, above) asserts only
|
||||
// that "second"'s DELIVERED FUTURE completes — never that "second" is REGISTERED with
|
||||
// CompletionResolver — so a redesign that dropped registrar.register() from the backstop
|
||||
// passed every existing test while reopening this ticket's own defect on the one path
|
||||
// fleetd #553 exists for. This test closes that gap: on the recovery path, the backstop is
|
||||
// the ONLY thing that registers "second", so its waiter must still be resolvable afterward.
|
||||
RuntimeException boom = new RuntimeException("fleetd #556 rework: boom from onTurnComplete");
|
||||
TurnListener throwing = new TurnListener() {
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
throw boom;
|
||||
}
|
||||
};
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver completion = new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none());
|
||||
Injector inj = new Injector(new AgentControl(herdr), throwing, _ -> true, _ -> {
|
||||
}, completion::register);
|
||||
|
||||
inj.enqueue(T, "first", TestTurnTokens.inert(T));
|
||||
CompletableFuture<Rendezvous.Resolution> secondWaiter = rendezvous.open(T);
|
||||
inj.enqueue(T, "second", new TurnToken(T, secondWaiter));
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE); // delivers "first" (inert token — nothing to register)
|
||||
inj.onStatus(T, AgentStatus.WORKING); // picked up
|
||||
|
||||
RuntimeException thrown = assertThrows(RuntimeException.class,
|
||||
// "first" completes (onTurnComplete throws) before "second"'s ordinary delivery
|
||||
// path ever runs; only the finally backstop is left to register "second".
|
||||
() -> inj.onStatus(T, AgentStatus.IDLE),
|
||||
"the throwable from onTurnComplete must still escape onStatus");
|
||||
assertSame(boom, thrown);
|
||||
|
||||
CompletionResolver.InFlight inFlight = completion.inFlight(T);
|
||||
assertNotNull(inFlight,
|
||||
"fleetd #556: \"second\" must be registered by the fleetd #553 finally backstop "
|
||||
+ "even though onTurnComplete threw for \"first\"'s completion before the "
|
||||
+ "ordinary registration path ever ran for \"second\"");
|
||||
assertSame(secondWaiter, inFlight.waiter(),
|
||||
"the registered entry must carry \"second\"'s own waiter, not some other value");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,61 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/** fleetd #544: the three-state contract in isolation from either loop that composes it. */
|
||||
class LoopWatchdogTest {
|
||||
|
||||
private static final long STALE_AFTER_NANOS = 1000;
|
||||
|
||||
@Test
|
||||
void reportsRunningBeforeTheThresholdElapses() {
|
||||
AtomicLong clock = new AtomicLong(0);
|
||||
LoopWatchdog watchdog = new LoopWatchdog(clock::get, STALE_AFTER_NANOS);
|
||||
clock.set(STALE_AFTER_NANOS - 1);
|
||||
assertEquals(LoopWatchdog.State.RUNNING, watchdog.state());
|
||||
}
|
||||
|
||||
@Test
|
||||
void reportsStalledOnceTheLastRoundAgesPastTheThreshold() {
|
||||
AtomicLong clock = new AtomicLong(0);
|
||||
LoopWatchdog watchdog = new LoopWatchdog(clock::get, STALE_AFTER_NANOS);
|
||||
clock.set(STALE_AFTER_NANOS);
|
||||
assertEquals(LoopWatchdog.State.STALLED, watchdog.state());
|
||||
}
|
||||
|
||||
@Test
|
||||
void recordRoundCompleteResetsTheStaleClock() {
|
||||
AtomicLong clock = new AtomicLong(0);
|
||||
LoopWatchdog watchdog = new LoopWatchdog(clock::get, STALE_AFTER_NANOS);
|
||||
clock.set(STALE_AFTER_NANOS - 1);
|
||||
watchdog.recordRoundComplete(); // last round is now "now" again
|
||||
clock.set(2 * STALE_AFTER_NANOS - 2);
|
||||
assertEquals(LoopWatchdog.State.RUNNING, watchdog.state(),
|
||||
"a fresh round completion must push the staleness deadline forward");
|
||||
}
|
||||
|
||||
@Test
|
||||
void markStoppedByCallerReportsStoppedRegardlessOfStaleness() {
|
||||
AtomicLong clock = new AtomicLong(0);
|
||||
LoopWatchdog watchdog = new LoopWatchdog(clock::get, STALE_AFTER_NANOS);
|
||||
watchdog.markStoppedByCaller();
|
||||
clock.set(STALE_AFTER_NANOS * 1000);
|
||||
assertEquals(LoopWatchdog.State.STOPPED, watchdog.state(),
|
||||
"an intentional stop must never be reported as STALLED");
|
||||
}
|
||||
|
||||
@Test
|
||||
void resetClearsAPreviousStopAndTheStaleClock() {
|
||||
AtomicLong clock = new AtomicLong(0);
|
||||
LoopWatchdog watchdog = new LoopWatchdog(clock::get, STALE_AFTER_NANOS);
|
||||
watchdog.markStoppedByCaller();
|
||||
clock.set(STALE_AFTER_NANOS * 1000);
|
||||
watchdog.reset();
|
||||
assertEquals(LoopWatchdog.State.RUNNING, watchdog.state(),
|
||||
"reset() must clear both the stop mark and the stale timestamp");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,99 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.msg.TestTurnTokens;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.lang.reflect.Field;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.CountDownLatch;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicBoolean;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotSame;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
class StatusPollerResilienceTest {
|
||||
|
||||
@Test
|
||||
void anErrorForOneTargetDoesNotStopPollingTheNextTarget() throws Exception {
|
||||
FakeHerdr fake = new FakeHerdr().withAgent("worker", "term_b", "w2:p8", "w2:t8");
|
||||
CountDownLatch errorThrown = new CountDownLatch(1);
|
||||
AgentControl agents = new AgentControl(new ErrorOnceForFirstTarget(fake, errorThrown));
|
||||
Injector injector = new Injector(agents);
|
||||
StatusPoller poller = new StatusPoller(agents, injector, 1);
|
||||
poller.start();
|
||||
try {
|
||||
injector.enqueue("term_a", "first", TestTurnTokens.inert("term_a"));
|
||||
assertTrue(errorThrown.await(2, TimeUnit.SECONDS),
|
||||
"the first target must throw its test Error");
|
||||
|
||||
CompletableFuture<Void> delivered =
|
||||
injector.enqueue("term_b", "second", TestTurnTokens.inert("term_b")).completion();
|
||||
delivered.get(2, TimeUnit.SECONDS);
|
||||
} finally {
|
||||
poller.stop();
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void anAbnormalExitClearsRunningSoStartCreatesANewLoop() throws Exception {
|
||||
StatusPoller poller = new StatusPoller(new AgentControl(new FakeHerdr()), new Injector(new AgentControl(new FakeHerdr())), -1);
|
||||
poller.start();
|
||||
Thread first = threadOf(poller);
|
||||
first.join(2000);
|
||||
assertFalse(runningOf(poller), "an abnormal loop exit must clear running");
|
||||
|
||||
poller.start();
|
||||
Thread restarted = threadOf(poller);
|
||||
try {
|
||||
assertNotSame(first, restarted, "start() must create a new loop after an abnormal exit");
|
||||
restarted.join(2000);
|
||||
} finally {
|
||||
poller.stop();
|
||||
}
|
||||
}
|
||||
|
||||
private static Thread threadOf(StatusPoller poller) throws ReflectiveOperationException {
|
||||
Field field = StatusPoller.class.getDeclaredField("thread");
|
||||
field.setAccessible(true);
|
||||
return (Thread) field.get(poller);
|
||||
}
|
||||
|
||||
private static boolean runningOf(StatusPoller poller) throws ReflectiveOperationException {
|
||||
Field field = StatusPoller.class.getDeclaredField("running");
|
||||
field.setAccessible(true);
|
||||
return field.getBoolean(poller);
|
||||
}
|
||||
|
||||
private static final class ErrorOnceForFirstTarget implements HerdrClient {
|
||||
private final FakeHerdr delegate;
|
||||
private final CountDownLatch errorThrown;
|
||||
private final AtomicBoolean first = new AtomicBoolean(true);
|
||||
|
||||
private ErrorOnceForFirstTarget(FakeHerdr delegate, CountDownLatch errorThrown) {
|
||||
this.delegate = delegate;
|
||||
this.errorThrown = errorThrown;
|
||||
}
|
||||
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) {
|
||||
if (method.equals("agent.get") && params instanceof Map<?, ?> map
|
||||
&& "w2:p7".equals(map.get("target")) && first.compareAndSet(true, false)) {
|
||||
errorThrown.countDown();
|
||||
throw new AssertionError("test Error from the first poll target");
|
||||
}
|
||||
return delegate.call(method, params);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
delegate.close();
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,172 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.msg.TestTurnTokens;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.lang.reflect.Field;
|
||||
import java.util.concurrent.CountDownLatch;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #544: {@link StatusPoller#health()} — the progress watchdog, not a liveness check. Every
|
||||
* test drives {@link LoopWatchdog}'s clock explicitly (never real elapsed time), so the threshold
|
||||
* crossing is deterministic rather than a race against the poll interval.
|
||||
*/
|
||||
class StatusPollerWatchdogTest {
|
||||
|
||||
private static final long INTERVAL_MILLIS = 50;
|
||||
// Mirrors StatusPoller.staleAfterNanos(50): 50ms * the 40x multiplier = 2000ms.
|
||||
private static final long STALE_AFTER_NANOS = TimeUnit.MILLISECONDS.toNanos(INTERVAL_MILLIS) * 40;
|
||||
|
||||
@Test
|
||||
void aLoopParkedInAHerdrCallThatNeverReturnsReportsStalled() throws Exception {
|
||||
CountDownLatch entered = new CountDownLatch(1);
|
||||
AgentControl agents = new AgentControl(new BlocksOnAgentGet(entered));
|
||||
Injector injector = new Injector(agents);
|
||||
AtomicLong clock = new AtomicLong(0);
|
||||
StatusPoller poller = new StatusPoller(agents, injector, new StatusRefiner(agents),
|
||||
INTERVAL_MILLIS, clock::get);
|
||||
poller.start();
|
||||
try {
|
||||
injector.enqueue("w1:p1", "hello", TestTurnTokens.inert("w1:p1"));
|
||||
assertTrue(entered.await(2, TimeUnit.SECONDS),
|
||||
"the poller must have entered the blocking herdr call");
|
||||
|
||||
assertEquals(LoopWatchdog.State.RUNNING, poller.health(),
|
||||
"freshly parked, before the threshold elapses, must still read RUNNING");
|
||||
|
||||
clock.set(STALE_AFTER_NANOS + 1);
|
||||
assertEquals(LoopWatchdog.State.STALLED, poller.health(),
|
||||
"a loop parked mid-round past the staleness threshold must report STALLED");
|
||||
} finally {
|
||||
poller.stop();
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void aLoopThatDiedReportsStalled() throws Exception {
|
||||
AtomicLong clock = new AtomicLong(0);
|
||||
// intervalMillis=-1: Thread.sleep(-1) throws IllegalArgumentException, uncaught by loop(),
|
||||
// which kills the carrier thread — the same "died" shape StatusPollerResilienceTest uses.
|
||||
StatusPoller poller = new StatusPoller(new AgentControl(new IdleHerdr()), new Injector(new AgentControl(new IdleHerdr())),
|
||||
new StatusRefiner(new AgentControl(new IdleHerdr())), -1, clock::get);
|
||||
poller.start();
|
||||
Thread thread = threadOf(poller);
|
||||
thread.join(2000);
|
||||
assertFalse(thread.isAlive(), "the loop must have died from Thread.sleep(-1)");
|
||||
|
||||
// staleAfterNanos(-1) = TimeUnit.MILLISECONDS.toNanos(max(-1,1)) * 40 = 40ms.
|
||||
clock.set(TimeUnit.MILLISECONDS.toNanos(1) * 40 + 1);
|
||||
assertEquals(LoopWatchdog.State.STALLED, poller.health(),
|
||||
"a dead loop's last-completed round goes stale and must report STALLED");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aLoopStoppedOnPurposeReportsStoppedNotStalled() throws Exception {
|
||||
AtomicLong clock = new AtomicLong(0);
|
||||
StatusPoller poller = new StatusPoller(new AgentControl(new IdleHerdr()), new Injector(new AgentControl(new IdleHerdr())),
|
||||
new StatusRefiner(new AgentControl(new IdleHerdr())), INTERVAL_MILLIS, clock::get);
|
||||
poller.start();
|
||||
poller.stop();
|
||||
|
||||
// Advance the clock far past any staleness threshold — an intentional stop must never
|
||||
// be reported as STALLED no matter how old the last round looks (fleetd #512 shape).
|
||||
clock.set(STALE_AFTER_NANOS * 100);
|
||||
assertEquals(LoopWatchdog.State.STOPPED, poller.health(),
|
||||
"stop() must report STOPPED, never STALLED, however stale the last round looks");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aRestartedLoopReportsRunningAgainNotStoppedForever() throws Exception {
|
||||
// fleetd #544 review (issue comment #16944): stoppedByCaller is sticky, and reset() —
|
||||
// called only from start() — is the sole thing that clears it. start() is documented
|
||||
// idempotent and loop()'s own error line says "it can be restarted", so stop() followed
|
||||
// by start() is an anticipated path. Without the reset() call in start(), health() would
|
||||
// report STOPPED forever after a restart even though the loop is genuinely running again.
|
||||
AtomicLong clock = new AtomicLong(0);
|
||||
StatusPoller poller = new StatusPoller(new AgentControl(new IdleHerdr()), new Injector(new AgentControl(new IdleHerdr())),
|
||||
new StatusRefiner(new AgentControl(new IdleHerdr())), INTERVAL_MILLIS, clock::get);
|
||||
poller.start();
|
||||
poller.stop();
|
||||
poller.start();
|
||||
|
||||
assertEquals(LoopWatchdog.State.RUNNING, poller.health(),
|
||||
"an intentional stop must not outlive the restart that follows it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aHealthyLoopReportsRunning() throws Exception {
|
||||
// Real elapsed time on purpose, unlike the other tests here: a clock frozen at 0 would read
|
||||
// RUNNING even if recordRoundComplete() were never wired into loop() at all, since the
|
||||
// constructor's own initial timestamp already satisfies "not stale yet". Running for several
|
||||
// multiples of this instance's own threshold (5ms interval * 40 = 200ms) and still reading
|
||||
// RUNNING instead proves the loop's repeated round completions are the thing keeping the
|
||||
// staleness deadline pushed forward.
|
||||
long fastIntervalMillis = 5;
|
||||
StatusPoller poller = new StatusPoller(new AgentControl(new IdleHerdr()), new Injector(new AgentControl(new IdleHerdr())),
|
||||
new StatusRefiner(new AgentControl(new IdleHerdr())), fastIntervalMillis, System::nanoTime);
|
||||
poller.start();
|
||||
try {
|
||||
Thread.sleep(800);
|
||||
assertEquals(LoopWatchdog.State.RUNNING, poller.health(),
|
||||
"a healthy loop's repeated round completions must keep pushing the staleness deadline forward");
|
||||
} finally {
|
||||
poller.stop();
|
||||
}
|
||||
}
|
||||
|
||||
private static Thread threadOf(StatusPoller poller) throws ReflectiveOperationException {
|
||||
Field field = StatusPoller.class.getDeclaredField("thread");
|
||||
field.setAccessible(true);
|
||||
return (Thread) field.get(poller);
|
||||
}
|
||||
|
||||
/** Never has any active target, so the loop just idles and completes empty rounds. */
|
||||
private static final class IdleHerdr implements HerdrClient {
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) {
|
||||
throw new AssertionError("unexpected herdr call: " + method);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static final class BlocksOnAgentGet implements HerdrClient {
|
||||
private final CountDownLatch entered;
|
||||
|
||||
private BlocksOnAgentGet(CountDownLatch entered) {
|
||||
this.entered = entered;
|
||||
}
|
||||
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) {
|
||||
if ("agent.get".equals(method)) {
|
||||
entered.countDown();
|
||||
try {
|
||||
// Blocks forever, mirroring UnixSocketHerdrClient.readLine()'s deadline-free
|
||||
// read (fleetd #544) — a test double standing in for a parked socket, not a
|
||||
// real one.
|
||||
new CountDownLatch(1).await();
|
||||
} catch (InterruptedException e) {
|
||||
Thread.currentThread().interrupt();
|
||||
}
|
||||
throw new AssertionError("unreachable: only stop()'s interrupt reaches here");
|
||||
}
|
||||
throw new AssertionError("unexpected herdr call: " + method);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,248 @@
|
||||
package dev.ltms.fleet.lead;
|
||||
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.io.RandomAccessFile;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
import static org.junit.jupiter.api.Assumptions.assumeFalse;
|
||||
|
||||
/**
|
||||
* Ticket "lead context gauge" — fleetd could not see how full a lead's own Claude Code context
|
||||
* window was, so a lead auto-compacting was always a surprise. {@link LeadContextGauge} reads that
|
||||
* off the lead's own transcript. Five acceptance properties from the ticket, one test class each
|
||||
* (plus the three separate UNKNOWN cases the ticket calls out by name):
|
||||
* <ol>
|
||||
* <li>{@link #tokenCountTracksTheLastUsageRecordAndChangesWithIt()}</li>
|
||||
* <li>{@link #compactionCountTracksCompactBoundaryRecordsAndChangesWithIt()}</li>
|
||||
* <li>{@link #missingFileIsUnknown()}, {@link #unreadableFileIsUnknown()}</li>
|
||||
* <li>{@link #tornFinalLineFallsBackToTheLastGoodReading()},
|
||||
* {@link #everyLineUnparseableIsUnknown()} (fleetd #602 gauge-wiring, finding 2)</li>
|
||||
* <li>{@link #readNeverExceedsTheTailBound()}</li>
|
||||
* <li>{@link #secondReadInsideTtlDoesNotTouchDiskAgain()}</li>
|
||||
* </ol>
|
||||
*
|
||||
* <p>Every fixture lives under {@code @TempDir} — never the operator's real config directory (see
|
||||
* the ticket's hard constraint on this).
|
||||
*
|
||||
* <p><strong>fleetd #602 gauge-wiring, finding 2.</strong> The property above used to be "a
|
||||
* truncated or invalid LAST line reports UNKNOWN" — on the theory that a bad last line signals a
|
||||
* format change. That reasoning did not hold: {@code fleet_list} reads this transcript while Claude
|
||||
* Code may be mid-write on it, so a torn LAST line is an ordinary race, not a format change, and a
|
||||
* real format change makes EVERY line unparseable, not only the last one written. So the property
|
||||
* is now split in two: {@link #tornFinalLineFallsBackToTheLastGoodReading} (the torn-line case must
|
||||
* NOT destroy a good earlier reading) and its control, {@link #everyLineUnparseableIsUnknown} (only
|
||||
* when NOTHING in the window parses does this report UNKNOWN).
|
||||
*/
|
||||
class LeadContextGaugeTest {
|
||||
|
||||
private static final String SESSION_ID = "11111111-1111-1111-1111-111111111111";
|
||||
|
||||
/** One line Claude Code would write for a turn with the given live-context total. */
|
||||
private static String usageLine(long inputTokens, long cacheRead, long cacheCreation) {
|
||||
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
|
||||
+ "\"input_tokens\":" + inputTokens + ","
|
||||
+ "\"cache_read_input_tokens\":" + cacheRead + ","
|
||||
+ "\"cache_creation_input_tokens\":" + cacheCreation + "}}}";
|
||||
}
|
||||
|
||||
/** One line Claude Code writes when an auto-compaction happens. */
|
||||
private static String compactionLine() {
|
||||
return "{\"type\":\"system\",\"subtype\":\"compact_boundary\","
|
||||
+ "\"compactMetadata\":{\"preTokens\":250000,\"postTokens\":5000,"
|
||||
+ "\"trigger\":\"auto\",\"durationMs\":54000}}";
|
||||
}
|
||||
|
||||
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} and writes {@code lines}. */
|
||||
private static String writeTranscript(Path configDir, String sessionId, String... lines) throws IOException {
|
||||
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
|
||||
Files.createDirectories(projectDir);
|
||||
Path file = projectDir.resolve(sessionId + ".jsonl");
|
||||
StringBuilder sb = new StringBuilder();
|
||||
for (String line : lines) {
|
||||
sb.append(line).append('\n');
|
||||
}
|
||||
Files.writeString(file, sb.toString(), StandardCharsets.UTF_8);
|
||||
return configDir.toString();
|
||||
}
|
||||
|
||||
// --- property 1: live token count tracks the last usage record -----------------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("reported tokens equal the LAST usage record's total, and change when it does")
|
||||
void tokenCountTracksTheLastUsageRecordAndChangesWithIt(@TempDir Path tmp) throws IOException {
|
||||
String configDir = writeTranscript(tmp, SESSION_ID,
|
||||
usageLine(1_000, 0, 0),
|
||||
usageLine(40_000, 5_000, 3_000)); // last record: 48,000
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
|
||||
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude");
|
||||
assertEquals(48_000L, first.tokens(), "must total input+cache_read+cache_creation of the LAST usage record");
|
||||
assertEquals(LeadContextGauge.State.OK, first.state());
|
||||
|
||||
// Change N in the fixture (a fresh session id avoids the cache) -- the reported number
|
||||
// must change with it, not stay pinned to the first fixture's total.
|
||||
String otherSession = "22222222-2222-2222-2222-222222222222";
|
||||
writeTranscript(tmp, otherSession, usageLine(100_000, 50_000, 50_000)); // last record: 200,000
|
||||
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude");
|
||||
assertEquals(200_000L, second.tokens());
|
||||
assertTrue(second.tokens() != first.tokens(), "changing N in the fixture must change the reported number");
|
||||
}
|
||||
|
||||
// --- property 2: compaction count tracks compact_boundary records --------------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("reported compaction count equals K compact_boundary records, and changes with K")
|
||||
void compactionCountTracksCompactBoundaryRecordsAndChangesWithIt(@TempDir Path tmp) throws IOException {
|
||||
String sessionTwoCompactions = "33333333-3333-3333-3333-333333333333";
|
||||
writeTranscript(tmp, sessionTwoCompactions,
|
||||
usageLine(1_000, 0, 0),
|
||||
compactionLine(),
|
||||
usageLine(2_000, 0, 0),
|
||||
compactionLine(),
|
||||
usageLine(3_000, 0, 0));
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude");
|
||||
assertEquals(2, twoCompactions.compactions());
|
||||
|
||||
String sessionZeroCompactions = "44444444-4444-4444-4444-444444444444";
|
||||
writeTranscript(tmp, sessionZeroCompactions, usageLine(3_000, 0, 0));
|
||||
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude");
|
||||
assertEquals(0, zeroCompactions.compactions(), "changing K in the fixture must change the reported count");
|
||||
}
|
||||
|
||||
// --- property 3: two separate UNKNOWN cases -------------------------------------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("a missing transcript file reports UNKNOWN with no token number")
|
||||
void missingFileIsUnknown(@TempDir Path tmp) {
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
|
||||
assertNull(reading.tokens());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unreadable transcript file reports UNKNOWN with no token number")
|
||||
void unreadableFileIsUnknown(@TempDir Path tmp) throws IOException {
|
||||
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
|
||||
Path file = tmp.resolve("projects").resolve("some-project-slug").resolve(SESSION_ID + ".jsonl");
|
||||
assertTrue(file.toFile().setReadable(false), "test setup: must be able to revoke read permission");
|
||||
try {
|
||||
// setReadable(false) really did clear the read bit (asserted above), but that alone
|
||||
// does not prove the file is UNREADABLE: running as root (e.g. a CI container) ignores
|
||||
// the read bit and opens the file anyway. Files.isReadable checks what actually happens
|
||||
// on open, not the bit. When it still reports readable, this test cannot create the
|
||||
// condition it needs on this machine, so it skips honestly instead of asserting on a
|
||||
// state that was never reached. A skip here means "I could not set up the case", NOT
|
||||
// "the UNKNOWN behaviour is fine" -- it is not evidence either way.
|
||||
assumeFalse(Files.isReadable(file),
|
||||
"runs as root (CI container): the read bit does not stop root, so this case cannot be set up here");
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
|
||||
assertNull(reading.tokens());
|
||||
} finally {
|
||||
file.toFile().setReadable(true); // so @TempDir cleanup can delete it
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("fleetd #602 finding 2: a torn final line does not destroy a good earlier reading")
|
||||
void tornFinalLineFallsBackToTheLastGoodReading(@TempDir Path tmp) throws IOException {
|
||||
// Deliberately NOT via writeTranscript: that helper appends a trailing newline after every
|
||||
// line, including the last one, which would not model what a write caught mid-flush looks
|
||||
// like on disk. Claude Code appends and flushes one line at a time, so the torn line here has
|
||||
// no trailing newline at all -- exactly the shape a read racing an in-progress write sees.
|
||||
Path projectDir = tmp.resolve("projects").resolve("some-project-slug");
|
||||
Files.createDirectories(projectDir);
|
||||
Path file = projectDir.resolve(SESSION_ID + ".jsonl");
|
||||
String lastCompleteLine = usageLine(1_000, 2_000, 3_000); // last COMPLETE record: 6,000
|
||||
String tornLine = "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{\"input_tok";
|
||||
Files.writeString(file, lastCompleteLine + "\n" + tornLine, StandardCharsets.UTF_8);
|
||||
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
|
||||
assertEquals(LeadContextGauge.State.OK, reading.state(),
|
||||
"a torn final line must not turn a good earlier reading into UNKNOWN");
|
||||
assertEquals(6_000L, reading.tokens(),
|
||||
"must report the last COMPLETE line's total, not fail the whole read");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("fleetd #602 finding 2 control: when EVERY line is unparseable, the reading IS UNKNOWN")
|
||||
void everyLineUnparseableIsUnknown(@TempDir Path tmp) throws IOException {
|
||||
// The real signal a format change gives: not one bad line, but ALL of them. Without this
|
||||
// control, code that never reports UNKNOWN at all would still pass the torn-line property.
|
||||
String configDir = writeTranscript(tmp, SESSION_ID,
|
||||
"{this is not json at all",
|
||||
"neither is this{{{");
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
|
||||
"every line unparseable is the real format-change signal and must still report UNKNOWN");
|
||||
assertNull(reading.tokens());
|
||||
}
|
||||
|
||||
// --- property 4: the tail bound is actually enforced ----------------------------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("reading a transcript far larger than the tail bound reads no more than the bound")
|
||||
void readNeverExceedsTheTailBound(@TempDir Path tmp) throws IOException {
|
||||
int smallBound = 4_096; // exercise the mechanism without writing a multi-MB fixture
|
||||
Path file = tmp.resolve("big.jsonl");
|
||||
// One line far bigger than smallBound, repeated, so the whole file is many times the bound.
|
||||
String line = usageLine(42, 0, 0) + " ".repeat(500);
|
||||
StringBuilder sb = new StringBuilder();
|
||||
for (int i = 0; i < 50; i++) {
|
||||
sb.append(line).append('\n');
|
||||
}
|
||||
Files.writeString(file, sb.toString(), StandardCharsets.UTF_8);
|
||||
long fileLength = Files.size(file);
|
||||
assertTrue(fileLength > (long) smallBound * 5, "test setup: fixture must genuinely dwarf the bound");
|
||||
|
||||
var tail = LeadContextGauge.tailBytes(file, smallBound);
|
||||
assertEquals(smallBound, tail.bytes().length,
|
||||
"must read exactly the bound, not the whole " + fileLength + "-byte file — assert on bytes actually read");
|
||||
}
|
||||
|
||||
// --- property 5: the cache TTL means a burst of calls reads the file once -------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("the same (configDir, sessionId) read twice inside the TTL touches disk once")
|
||||
void secondReadInsideTtlDoesNotTouchDiskAgain(@TempDir Path tmp) throws IOException {
|
||||
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
|
||||
AtomicLong now = new AtomicLong(0);
|
||||
LeadContextGauge gauge = new LeadContextGauge(now::get, 5_000);
|
||||
|
||||
gauge.read(configDir, SESSION_ID, "claude");
|
||||
gauge.read(configDir, SESSION_ID, "claude"); // still inside the TTL window
|
||||
assertEquals(1, gauge.diskReadCount(), "two reads inside the TTL must touch disk once");
|
||||
|
||||
now.set(6_000); // past the TTL
|
||||
gauge.read(configDir, SESSION_ID, "claude");
|
||||
assertEquals(2, gauge.diskReadCount(), "a read past the TTL must touch disk again");
|
||||
}
|
||||
|
||||
// --- non-Claude peer / unresolved session -----------------------------------------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("a non-claude agent type or a null session id is UNKNOWN, never OK")
|
||||
void nonClaudePeerOrUnresolvedSessionIsUnknown(@TempDir Path tmp) throws IOException {
|
||||
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode").state());
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null).state());
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude").state());
|
||||
}
|
||||
}
|
||||
@@ -32,13 +32,26 @@ class LeadLauncherTest {
|
||||
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true, null);
|
||||
}
|
||||
|
||||
private static FleetConfig.Profile profileWithAutoCompactWindow(String kind) {
|
||||
return new FleetConfig.Profile(
|
||||
"opus", null, "claude-opus-5", null, "FLEETD_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms"), "tab", "fleet", null,
|
||||
"http://127.0.0.1:8765/mcp", null, null,
|
||||
null, null, kind, Map.of(), null, null, true, null, null,
|
||||
null, null, null, 250_000);
|
||||
}
|
||||
|
||||
private static FleetConfig configWith(FleetConfig.Leader lead) {
|
||||
return configWith(lead, opusProfile());
|
||||
}
|
||||
|
||||
private static FleetConfig configWith(FleetConfig.Leader lead, FleetConfig.Profile profile) {
|
||||
Map<String, FleetConfig.Leader> leaders = new LinkedHashMap<>();
|
||||
leaders.put("opus", lead);
|
||||
FleetConfig.Fleet fleet =
|
||||
new FleetConfig.Fleet(leaders, Map.of(), Map.of(), Map.of(), null);
|
||||
return new FleetConfig(
|
||||
null, null, Map.of("opus", opusProfile()), null, null, null, null, null,
|
||||
null, null, Map.of("opus", profile), null, null, null, null, null,
|
||||
null, null, fleet, null, "fixed", null).withDefaults();
|
||||
}
|
||||
|
||||
@@ -330,6 +343,36 @@ class LeadLauncherTest {
|
||||
"--model is appended last so it outranks the ccs wrapper (CB-533)");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aClaudeLeadPassesItsConfiguredAutoCompactWindowToClaudeCode() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile profile = profileWithAutoCompactWindow("claude-code");
|
||||
|
||||
launcher(herdr, configWith(lead("opus", "lead: opus", 1), profile)).ensureLeads();
|
||||
|
||||
List<String> args = startedArgs(herdr);
|
||||
assertEquals("250000", args.get(args.indexOf("--autocompact") + 1));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aClaudeLeadWithNoAutoCompactWindowGetsNoAutoCompactFlag() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
|
||||
|
||||
assertFalse(startedArgs(herdr).contains("--autocompact"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aNonClaudeLeadDoesNotGetAnAutoCompactFlag() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile profile = profileWithAutoCompactWindow("opencode");
|
||||
|
||||
launcher(herdr, configWith(lead("opus", "lead: opus", 1), profile)).ensureLeads();
|
||||
|
||||
assertFalse(startedArgs(herdr).contains("--autocompact"));
|
||||
}
|
||||
|
||||
/** A lead runs on the operator's subscription. Nothing may move it off. */
|
||||
@Test
|
||||
void theLeadEnvCarriesNoAnthropicBinding() {
|
||||
|
||||
@@ -1,23 +1,23 @@
|
||||
package dev.ltms.fleet.lead;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.HerdrException;
|
||||
import dev.ltms.fleet.testing.CapturedLog;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.LongSupplier;
|
||||
@@ -862,25 +862,16 @@ class LeadRolloverTest {
|
||||
|
||||
// ---- fleetd #494: the log lines must print MEASURED values, never the configured budget --
|
||||
|
||||
private static ListAppender<ILoggingEvent> attachLog() {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(LeadRollover.class);
|
||||
logger.setLevel(Level.DEBUG);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
return appender;
|
||||
private static CapturedLog attachLog() {
|
||||
return CapturedLog.at(LeadRollover.class, Level.DEBUG);
|
||||
}
|
||||
|
||||
private static void detachLog(ListAppender<ILoggingEvent> appender) {
|
||||
((Logger) LoggerFactory.getLogger(LeadRollover.class)).detachAppender(appender);
|
||||
}
|
||||
|
||||
private static ILoggingEvent lastEventContaining(ListAppender<ILoggingEvent> events, String substring) {
|
||||
return events.list.stream()
|
||||
private static ILoggingEvent lastEventContaining(List<ILoggingEvent> events, String substring) {
|
||||
return events.stream()
|
||||
.filter(e -> e.getFormattedMessage().contains(substring))
|
||||
.reduce((_, b) -> b)
|
||||
.orElseThrow(() -> new AssertionError("no log event contained \"" + substring
|
||||
+ "\"; got: " + events.list.stream().map(ILoggingEvent::getFormattedMessage).toList()));
|
||||
+ "\"; got: " + events.stream().map(ILoggingEvent::getFormattedMessage).toList()));
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -914,15 +905,14 @@ class LeadRolloverTest {
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(flipsAfterClear, config, () -> clock.addAndGet(500));
|
||||
|
||||
ListAppender<ILoggingEvent> events = attachLog();
|
||||
try {
|
||||
try (CapturedLog log = attachLog()) {
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "every synchronous gate passes; the refusal is logged "
|
||||
+ "only, deep inside the deferred continuation");
|
||||
|
||||
ILoggingEvent event = lastEventContaining(events, "NOT sending bootstrapText");
|
||||
ILoggingEvent event = lastEventContaining(log.events(), "NOT sending bootstrapText");
|
||||
assertEquals(Level.WARN, event.getLevel());
|
||||
String message = event.getFormattedMessage();
|
||||
assertTrue(message.contains("configured=1s"), "must label the configured budget: " + message);
|
||||
@@ -934,8 +924,6 @@ class LeadRolloverTest {
|
||||
+ message);
|
||||
assertFalse(message.contains("within 1s"), "must not present the configured budget as if "
|
||||
+ "it were the measured wait duration: " + message);
|
||||
} finally {
|
||||
detachLog(events);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -954,15 +942,14 @@ class LeadRolloverTest {
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, config, () -> clock.addAndGet(500));
|
||||
|
||||
ListAppender<ILoggingEvent> events = attachLog();
|
||||
try {
|
||||
try (CapturedLog log = attachLog()) {
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason()
|
||||
+ " / " + decision.detail());
|
||||
|
||||
ILoggingEvent event = lastEventContaining(events, "releasing rather than wedging the roll");
|
||||
ILoggingEvent event = lastEventContaining(log.events(), "releasing rather than wedging the roll");
|
||||
assertEquals(Level.WARN, event.getLevel(), "the grace-limit release must be WARN, not "
|
||||
+ "INFO — it is exactly the case that reported false success in the real incident "
|
||||
+ "this fix comes from (a roll that 'succeeded' after 438ms of a 20s budget)");
|
||||
@@ -982,8 +969,6 @@ class LeadRolloverTest {
|
||||
assertTrue(message.contains("after 8 consecutive IDLE/DONE polls (7 of those were nudged)"),
|
||||
"must print the measured poll count and nudge count as plain numbers, not the "
|
||||
+ "PICKUP_GRACE_POLLS constant standing in for either: " + message);
|
||||
} finally {
|
||||
detachLog(events);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1000,15 +985,14 @@ class LeadRolloverTest {
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, config, () -> clock.addAndGet(500));
|
||||
|
||||
ListAppender<ILoggingEvent> events = attachLog();
|
||||
try {
|
||||
try (CapturedLog log = attachLog()) {
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason()
|
||||
+ " / " + decision.detail());
|
||||
|
||||
ILoggingEvent event = lastEventContaining(events, "lead-rollover: rolled");
|
||||
ILoggingEvent event = lastEventContaining(log.events(), "lead-rollover: rolled");
|
||||
assertEquals(Level.INFO, event.getLevel());
|
||||
String message = event.getFormattedMessage();
|
||||
// fleetd #494 follow-up: waitUntilAtTurnBoundary now also reads the injected clock one
|
||||
@@ -1018,11 +1002,327 @@ class LeadRolloverTest {
|
||||
assertTrue(message.contains("elapsedMs=7000"), "must print the MEASURED elapsed time for "
|
||||
+ "the whole roll — with this fixture's advancing clock, the full roll (turn-settle "
|
||||
+ "wait + /clear wait + bootstrapText) took 7000ms: " + message);
|
||||
} finally {
|
||||
detachLog(events);
|
||||
}
|
||||
}
|
||||
|
||||
// ---- CB-... : LeadRollover#status makes the outcome of a confirmed roll readable -----------
|
||||
|
||||
@Test
|
||||
@DisplayName("[STATUS 1] after the calling turn never settles, status() reports "
|
||||
+ "TURN_NEVER_SETTLED for that token — this assertion could not even be written before "
|
||||
+ "status() existed")
|
||||
void statusReportsTurnNeverSettledAfterTheRollIsAbandoned() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
herdr.agentStatus("working"); // the calling lead's own pane — never goes idle in this test
|
||||
Path handover = writeHandover("handover contents");
|
||||
FleetConfig.LeadRollover config =
|
||||
new FleetConfig.LeadRollover(handover.toString(), true, 3600, 1 /*turnSettleSeconds*/, 20, "text");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, config, () -> clock.addAndGet(500));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
assertTrue(decision.accepted(), "every synchronous gate should pass; the refusal happens "
|
||||
+ "only inside the deferred continuation, which this test's synchronous runner has "
|
||||
+ "already run to completion by the time confirm() returns");
|
||||
assertEquals(0, promptCallCount(herdr), "sanity: /clear was never sent");
|
||||
|
||||
LeadRollover.RollStatus status = rollover.status(pending.token());
|
||||
assertEquals(LeadRollover.RollState.TURN_NEVER_SETTLED, status.state());
|
||||
assertTrue(status.detail().contains("turnSettleSeconds"), "the detail must name the knob a "
|
||||
+ "reader needs to raise: " + status.detail());
|
||||
assertTrue(status.detail().toLowerCase().contains("no /clear"), "the detail must say plainly "
|
||||
+ "that no /clear was ever sent: " + status.detail());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[STATUS 2] a roll that completes reports ROLLED for its token")
|
||||
void statusReportsRolledForACompletedRoll() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr(); // default idle — a full successful roll
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||
assertEquals(2, promptCallCount(herdr), "sanity: /clear then bootstrapText were both sent");
|
||||
|
||||
LeadRollover.RollStatus status = rollover.status(pending.token());
|
||||
assertEquals(LeadRollover.RollState.ROLLED, status.state());
|
||||
assertNotNull(status.detail());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[STATUS 3] a roll where /clear never settles reports CLEAR_NEVER_SETTLED — "
|
||||
+ "distinct from both ROLLED and TURN_NEVER_SETTLED")
|
||||
void statusReportsClearNeverSettledDistinctFromTheOtherTwoStates() throws IOException {
|
||||
// Idle until /clear is sent, then permanently working — the SECOND wait never settles.
|
||||
FakeHerdr fake = new FakeHerdr();
|
||||
HerdrClient flipsAfterClear = new HerdrClient() {
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) throws HerdrException {
|
||||
JsonNode result = fake.call(method, params);
|
||||
if ("agent.prompt".equals(method) && String.valueOf(params).contains("/clear")) {
|
||||
fake.agentStatus("working");
|
||||
}
|
||||
return result;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
fake.close();
|
||||
}
|
||||
};
|
||||
Path handover = writeHandover("handover contents");
|
||||
FleetConfig.LeadRollover config =
|
||||
new FleetConfig.LeadRollover(handover.toString(), true, 3600, 20, 1 /*clearSettleSeconds*/, "boot text");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(flipsAfterClear, config, () -> clock.addAndGet(500));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
assertTrue(decision.accepted(), "every synchronous gate passes; the refusal is logged only, "
|
||||
+ "deep inside the deferred continuation");
|
||||
assertEquals(1, promptCallCount(fake), "sanity: only /clear was sent, never bootstrapText");
|
||||
|
||||
LeadRollover.RollStatus status = rollover.status(pending.token());
|
||||
assertEquals(LeadRollover.RollState.CLEAR_NEVER_SETTLED, status.state());
|
||||
assertNotEquals(LeadRollover.RollState.ROLLED, status.state());
|
||||
assertNotEquals(LeadRollover.RollState.TURN_NEVER_SETTLED, status.state());
|
||||
assertTrue(status.detail().contains("clearSettleSeconds"), status.detail());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[STATUS 4] a token that was never issued, or was cancelled, gives a clean "
|
||||
+ "UNKNOWN answer rather than an exception or a false ROLLED")
|
||||
void statusOnUnissuedOrCancelledTokenIsCleanNotAnExceptionOrFalseSuccess() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
|
||||
|
||||
// never issued at all
|
||||
LeadRollover.RollStatus neverIssued = assertDoesNotThrow(() -> rollover.status("no-such-token"));
|
||||
assertEquals(LeadRollover.RollState.UNKNOWN, neverIssued.state());
|
||||
assertNotEquals(LeadRollover.RollState.ROLLED, neverIssued.state());
|
||||
|
||||
// null / blank must not throw either — pending is a ConcurrentHashMap, which throws on a
|
||||
// null-key lookup unless status() guards it first
|
||||
assertEquals(LeadRollover.RollState.UNKNOWN, assertDoesNotThrow(() -> rollover.status(null)).state());
|
||||
assertEquals(LeadRollover.RollState.UNKNOWN, assertDoesNotThrow(() -> rollover.status(" ")).state());
|
||||
|
||||
// opened, then cancelled
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
assertTrue(rollover.cancel(pending.token()));
|
||||
LeadRollover.RollStatus cancelled = assertDoesNotThrow(() -> rollover.status(pending.token()));
|
||||
assertEquals(LeadRollover.RollState.UNKNOWN, cancelled.state());
|
||||
assertNotEquals(LeadRollover.RollState.ROLLED, cancelled.state());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[STATUS 5] the bounded outcome history never grows past its cap")
|
||||
void statusHistoryDoesNotGrowPastItsCap() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr(); // default idle — every roll completes
|
||||
Path handover = writeHandover("handover contents"); // one file, reused by every roll below —
|
||||
// checkHandover only compares its mtime against each open()'s OWN requestedAtMillis (the
|
||||
// fake clock, in the low thousands), and the file's real (wall-clock) mtime is always far
|
||||
// larger than that, so freshness passes on every iteration without rewriting the file.
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), () -> clock.addAndGet(1));
|
||||
|
||||
int rolls = LeadRollover.OUTCOME_HISTORY_CAP + 50;
|
||||
String[] tokens = new String[rolls];
|
||||
for (int i = 0; i < rolls; i++) {
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full #" + i);
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
assertTrue(decision.accepted(), "roll #" + i + " should have been approved: " + decision.detail());
|
||||
tokens[i] = pending.token();
|
||||
}
|
||||
|
||||
assertEquals(LeadRollover.RollState.ROLLED, rollover.status(tokens[rolls - 1]).state(),
|
||||
"the most recently finished roll's outcome must still be in the bounded history");
|
||||
|
||||
// There is no direct size accessor for the bounded history, so boundedness is asserted
|
||||
// indirectly and behaviourally: the OLDEST finished roll's outcome must have been evicted
|
||||
// (reads back as a clean UNKNOWN, exactly like a token that was never issued) once more than
|
||||
// OUTCOME_HISTORY_CAP rolls have gone through this instance. If the cap were not enforced,
|
||||
// tokens[0] would still read back ROLLED here, and this assertion would fail.
|
||||
assertEquals(LeadRollover.RollState.UNKNOWN, rollover.status(tokens[0]).state(),
|
||||
"the oldest finished roll's outcome must have been evicted once the cap was "
|
||||
+ "exceeded — otherwise the bounded history is not actually bounded");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[STATUS 6] a status() call against a pending roll is read-only — no herdr call is "
|
||||
+ "made and the pending request is left untouched")
|
||||
void statusCallAgainstAPendingRollIsReadOnlyAndTouchesNothing() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
|
||||
LeadRollover.RollStatus status = rollover.status(pending.token());
|
||||
assertEquals(LeadRollover.RollState.PENDING, status.state());
|
||||
assertEquals(0, herdr.calls.size(), "status() must never make any herdr call at all — not "
|
||||
+ "just no agent.prompt — since it must never schedule, cancel, or retry anything");
|
||||
|
||||
// the pending request must be left exactly as it was: the SAME token can still be confirmed
|
||||
// afterwards, as if status() had never been called.
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
assertTrue(decision.accepted(), "status() must not have consumed or otherwise disturbed the "
|
||||
+ "pending request: " + decision.reason() + " / " + decision.detail());
|
||||
}
|
||||
|
||||
// ---- PR #600 review round 2: IN_PROGRESS — the gap between confirm() handing off and the ----
|
||||
// ---- continuation finishing must never read back as UNKNOWN ("nothing was ever requested") --
|
||||
|
||||
/**
|
||||
* A {@code continuationRunner} that CAPTURES the roll instead of running it, so a test can
|
||||
* observe {@link LeadRollover#status} in the window between {@link LeadRollover#confirm}
|
||||
* handing off and the roll actually finishing — the window the synchronous {@code
|
||||
* Runnable::run} runner used everywhere else in this class collapses to nothing. Call {@link
|
||||
* #runNext()} to finish exactly one held roll, once the test is done observing the in-flight
|
||||
* state.
|
||||
*/
|
||||
private static final class HoldingRunner implements java.util.function.Consumer<Runnable> {
|
||||
private final List<Runnable> held = new ArrayList<>();
|
||||
|
||||
@Override
|
||||
public void accept(Runnable runnable) {
|
||||
held.add(runnable);
|
||||
}
|
||||
|
||||
int heldCount() {
|
||||
return held.size();
|
||||
}
|
||||
|
||||
/** Runs (and removes) the oldest held roll — FIFO, matching confirm() call order. */
|
||||
void runNext() {
|
||||
held.remove(0).run();
|
||||
}
|
||||
}
|
||||
|
||||
private static LeadRollover newRolloverWithHoldingRunner(HerdrClient herdr,
|
||||
FleetConfig.LeadRollover config, LongSupplier nowMillis, HoldingRunner runner) {
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
return new LeadRollover(agents, () -> config, _ -> null, nowMillis, () -> { }, runner);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[IN-PROGRESS 1] between an approved confirm() and the continuation finishing, "
|
||||
+ "status() reports IN_PROGRESS — not UNKNOWN, and not PENDING")
|
||||
void statusReportsInProgressBetweenConfirmAndTheContinuationFinishing() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr(); // default idle — the held roll WOULD complete once run
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
HoldingRunner runner = new HoldingRunner();
|
||||
LeadRollover rollover = newRolloverWithHoldingRunner(herdr, cfg(handover.toString()),
|
||||
fixedClock(clock), runner);
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||
assertEquals(1, runner.heldCount(), "sanity: the roll must have been handed to the "
|
||||
+ "continuation runner and held there, not run yet");
|
||||
assertEquals(0, promptCallCount(herdr), "sanity: the held continuation has not run, so "
|
||||
+ "nothing has been sent to the pane yet");
|
||||
|
||||
LeadRollover.RollStatus status = rollover.status(pending.token());
|
||||
assertEquals(LeadRollover.RollState.IN_PROGRESS, status.state(),
|
||||
"a lead calling status() right after confirm() returned approved, while the roll "
|
||||
+ "is still running, must be told IN_PROGRESS — not UNKNOWN (\"nothing was "
|
||||
+ "ever requested\", which would wrongly invite it to call open() again "
|
||||
+ "mid-roll) and not PENDING (\"not yet approved\", which is simply false "
|
||||
+ "here): got " + status.state() + " / " + status.detail());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[IN-PROGRESS 2] once the continuation finishes, the same token reports its "
|
||||
+ "terminal state — IN_PROGRESS is not sticky")
|
||||
void inProgressStateIsNotStickyOnceTheContinuationFinishes() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr(); // default idle — the held roll completes once run
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
HoldingRunner runner = new HoldingRunner();
|
||||
LeadRollover rollover = newRolloverWithHoldingRunner(herdr, cfg(handover.toString()),
|
||||
fixedClock(clock), runner);
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||
assertEquals(LeadRollover.RollState.IN_PROGRESS, rollover.status(pending.token()).state(),
|
||||
"sanity: must be IN_PROGRESS before the held continuation is run");
|
||||
|
||||
runner.runNext(); // finish the held roll now
|
||||
|
||||
assertEquals(2, promptCallCount(herdr), "sanity: the roll actually ran to completion "
|
||||
+ "once released — /clear then bootstrapText");
|
||||
LeadRollover.RollStatus after = rollover.status(pending.token());
|
||||
assertEquals(LeadRollover.RollState.ROLLED, after.state(),
|
||||
"the SAME token must now report its terminal state — IN_PROGRESS must not still be "
|
||||
+ "reported once the roll has actually finished");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[IN-PROGRESS 4] IN_PROGRESS is distinct from every other RollState")
|
||||
void inProgressStateIsDistinctFromAllOtherStates() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
HoldingRunner runner = new HoldingRunner();
|
||||
LeadRollover rollover = newRolloverWithHoldingRunner(herdr, cfg(handover.toString()),
|
||||
fixedClock(clock), runner);
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||
|
||||
LeadRollover.RollState inProgress = rollover.status(pending.token()).state();
|
||||
assertEquals(LeadRollover.RollState.IN_PROGRESS, inProgress);
|
||||
for (LeadRollover.RollState other : LeadRollover.RollState.values()) {
|
||||
if (other == LeadRollover.RollState.IN_PROGRESS) {
|
||||
continue;
|
||||
}
|
||||
assertNotEquals(other, inProgress, "IN_PROGRESS must be distinct from " + other);
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[IN-PROGRESS 5] eviction counts IN_PROGRESS entries toward the cap exactly like "
|
||||
+ "finished ones — a burst of confirmed-but-not-yet-finished rolls still ages out")
|
||||
void evictionCountsInProgressEntriesTowardTheCap() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path handover = writeHandover("handover contents"); // reused by every roll — see
|
||||
// statusHistoryDoesNotGrowPastItsCap for why one shared file is enough for freshness.
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
HoldingRunner runner = new HoldingRunner(); // nothing run below — every roll stays IN_PROGRESS
|
||||
LeadRollover rollover = newRolloverWithHoldingRunner(herdr, cfg(handover.toString()),
|
||||
() -> clock.addAndGet(1), runner);
|
||||
|
||||
int rolls = LeadRollover.OUTCOME_HISTORY_CAP + 50;
|
||||
String[] tokens = new String[rolls];
|
||||
for (int i = 0; i < rolls; i++) {
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full #" + i);
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
assertTrue(decision.accepted(), "roll #" + i + " should have been approved: " + decision.detail());
|
||||
tokens[i] = pending.token();
|
||||
}
|
||||
assertEquals(rolls, runner.heldCount(), "sanity: none of these rolls have been run — every "
|
||||
+ "one of them is sitting in outcomes as IN_PROGRESS, not a separate uncapped map");
|
||||
|
||||
assertEquals(LeadRollover.RollState.IN_PROGRESS, rollover.status(tokens[rolls - 1]).state(),
|
||||
"the most recently confirmed (still in-flight) roll must still be in the bounded "
|
||||
+ "history");
|
||||
assertEquals(LeadRollover.RollState.UNKNOWN, rollover.status(tokens[0]).state(),
|
||||
"the oldest confirmed roll's IN_PROGRESS entry must have been evicted once the cap "
|
||||
+ "was exceeded, exactly like a finished entry would be — proving IN_PROGRESS "
|
||||
+ "entries share the SAME bounded map and count against the SAME cap, rather "
|
||||
+ "than living in a second, uncapped in-flight map");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #494 follow-up] the turn-settle timeout warn line prints the MEASURED "
|
||||
+ "elapsed time next to the configured budget, never the configured value alone")
|
||||
@@ -1040,8 +1340,7 @@ class LeadRolloverTest {
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, config, () -> clock.addAndGet(500));
|
||||
|
||||
ListAppender<ILoggingEvent> events = attachLog();
|
||||
try {
|
||||
try (CapturedLog log = attachLog()) {
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
@@ -1049,7 +1348,7 @@ class LeadRolloverTest {
|
||||
+ "only inside the deferred continuation, which this test's synchronous runner "
|
||||
+ "has already run to completion by the time confirm() returns");
|
||||
|
||||
ILoggingEvent event = lastEventContaining(events, "refusing to send /clear at all");
|
||||
ILoggingEvent event = lastEventContaining(log.events(), "refusing to send /clear at all");
|
||||
assertEquals(Level.WARN, event.getLevel());
|
||||
String message = event.getFormattedMessage();
|
||||
assertTrue(message.contains("configured=1s"), "must label the configured budget: " + message);
|
||||
@@ -1058,8 +1357,94 @@ class LeadRolloverTest {
|
||||
+ "1s(=1000ms) configured budget: " + message);
|
||||
assertFalse(message.contains("within 1s"), "must not present the configured budget as if "
|
||||
+ "it were the measured wait duration: " + message);
|
||||
} finally {
|
||||
detachLog(events);
|
||||
}
|
||||
}
|
||||
|
||||
// ---- fleetd #615: a HerdrException out of either unwrapped agents.send call must leave a ----
|
||||
// ---- TERMINAL FAILED outcome, never a stuck IN_PROGRESS ---------------------------------------
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #615 — 1] send() throwing on the /clear call leaves status(token) "
|
||||
+ "reporting FAILED, not stuck at IN_PROGRESS")
|
||||
void sendThrowingOnClearLeavesStatusReportingFailed() throws IOException {
|
||||
FakeHerdr fake = new FakeHerdr(); // default idle — the turn-settle wait passes immediately
|
||||
HerdrClient throwsOnClear = new HerdrClient() {
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) throws HerdrException {
|
||||
if ("agent.prompt".equals(method) && String.valueOf(params).contains("/clear")) {
|
||||
throw new HerdrException("simulated herdr transport failure sending /clear");
|
||||
}
|
||||
return fake.call(method, params);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
fake.close();
|
||||
}
|
||||
};
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(throwsOnClear, cfg(handover.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "every synchronous gate passes; the throw happens only "
|
||||
+ "inside the deferred continuation, which this test's synchronous runner has "
|
||||
+ "already run to completion by the time confirm() returns");
|
||||
|
||||
LeadRollover.RollStatus status = rollover.status(pending.token());
|
||||
assertEquals(LeadRollover.RollState.FAILED, status.state(),
|
||||
"a HerdrException out of the /clear send must leave a TERMINAL FAILED outcome — "
|
||||
+ "before fleetd #615's fix, the continuation thread died silently and "
|
||||
+ "status() was stuck reporting the IN_PROGRESS confirm() wrote at hand-off, "
|
||||
+ "forever: got " + status.state() + " / " + status.detail());
|
||||
assertNotEquals(LeadRollover.RollState.IN_PROGRESS, status.state());
|
||||
assertTrue(status.detail().contains("HerdrException"), "the detail must name the exception "
|
||||
+ "so an operator reading status() has something to act on: " + status.detail());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #615 — 2] send() throwing on the bootstrap-text call (after /clear "
|
||||
+ "succeeded and the pane settled) also leaves status(token) reporting FAILED — a "
|
||||
+ "DIFFERENT exit from the /clear-throw case above")
|
||||
void sendThrowingOnBootstrapTextLeavesStatusReportingFailed() throws IOException {
|
||||
FakeHerdr fake = new FakeHerdr(); // default idle throughout — both settle waits pass promptly
|
||||
HerdrClient throwsOnBootstrapText = new HerdrClient() {
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) throws HerdrException {
|
||||
if ("agent.prompt".equals(method) && String.valueOf(params).contains("read the handover file")) {
|
||||
throw new HerdrException("simulated herdr transport failure sending bootstrapText");
|
||||
}
|
||||
return fake.call(method, params);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
fake.close();
|
||||
}
|
||||
};
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(throwsOnBootstrapText, cfg(handover.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "every synchronous gate passes; the throw happens only "
|
||||
+ "inside the deferred continuation, which this test's synchronous runner has "
|
||||
+ "already run to completion by the time confirm() returns");
|
||||
assertEquals(1, promptCallCount(fake), "sanity: /clear was sent and settled — only the "
|
||||
+ "SECOND agent.prompt call (bootstrapText) threw");
|
||||
|
||||
LeadRollover.RollStatus status = rollover.status(pending.token());
|
||||
assertEquals(LeadRollover.RollState.FAILED, status.state(),
|
||||
"a HerdrException out of the bootstrapText send — a DIFFERENT exit from the /clear "
|
||||
+ "throw, reached only after /clear already succeeded and the pane already "
|
||||
+ "settled — must also leave a TERMINAL FAILED outcome, not a stuck "
|
||||
+ "IN_PROGRESS: got " + status.state() + " / " + status.detail());
|
||||
assertNotEquals(LeadRollover.RollState.IN_PROGRESS, status.state());
|
||||
assertTrue(status.detail().contains("HerdrException"), "the detail must name the exception "
|
||||
+ "so an operator reading status() has something to act on: " + status.detail());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -258,5 +258,51 @@ class FleetMcpHandoverTest {
|
||||
assertTrue(textOf(r).contains("\"cancelled\":true"), textOf(r));
|
||||
}
|
||||
|
||||
// --- new: action "status" — makes the outcome of a confirm() readable through the tool ----
|
||||
|
||||
@Test
|
||||
@DisplayName("status on a null LeadRollover is a clean NOT_CONFIGURED refusal, never a throw")
|
||||
void statusWithNullLeadRolloverRefusesCleanly() {
|
||||
McpSchema.CallToolResult r = assertDoesNotThrow(() -> FleetMcp.handover(null, LEAD,
|
||||
Map.of("action", "status", "token", "whatever")));
|
||||
assertFalse(r.isError());
|
||||
assertTrue(textOf(r).contains("NOT_CONFIGURED"), textOf(r));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("status on a token that was never opened reports UNKNOWN")
|
||||
void statusOnUnknownTokenReportsUnknown() {
|
||||
LeadRollover rollover = newRollover(tmp.resolve("h.md").toString());
|
||||
|
||||
McpSchema.CallToolResult r = FleetMcp.handover(rollover, LEAD,
|
||||
Map.of("action", "status", "token", "does-not-exist"));
|
||||
assertFalse(r.isError());
|
||||
assertTrue(textOf(r).contains("\"state\":\"UNKNOWN\""), textOf(r));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("status on a token that is still pending (opened, not confirmed) reports PENDING")
|
||||
void statusOnPendingTokenReportsPending() {
|
||||
LeadRollover rollover = newRollover(tmp.resolve("h.md").toString());
|
||||
String token = extractToken(textOf(FleetMcp.handover(rollover, LEAD, Map.of("action", "open"))));
|
||||
|
||||
McpSchema.CallToolResult r = FleetMcp.handover(rollover, LEAD,
|
||||
Map.of("action", "status", "token", token));
|
||||
assertFalse(r.isError());
|
||||
assertTrue(textOf(r).contains("\"state\":\"PENDING\""), textOf(r));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("status is registered on the tool's schema and the schema still names no caller-identity parameter")
|
||||
void statusActionIsAdvertisedOnTheSchema() {
|
||||
FleetMcp m = mcp(null);
|
||||
McpSchema.Tool tool = m.registeredTools().stream()
|
||||
.filter(t -> "fleet_handover".equals(t.name()))
|
||||
.findFirst()
|
||||
.orElseThrow(() -> new AssertionError("fleet_handover was not registered"));
|
||||
assertTrue(tool.description().contains("'status'"),
|
||||
"the tool's own description must advertise the 'status' action: " + tool.description());
|
||||
}
|
||||
|
||||
// --- acceptance 7 (wiring) is covered by FleetdLeadRolloverWiringTest, unchanged -----------
|
||||
}
|
||||
|
||||
@@ -0,0 +1,126 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #602 gauge-wiring: {@code FleetMcp}'s {@code fleet_list} handler shipped calling
|
||||
* {@code LeadContextGauge.read(null, ...)} unconditionally, so every lead whose profile sets a
|
||||
* {@code configDir:} override read the wrong transcript directory forever, with no error anywhere —
|
||||
* {@link LeadContextGauge}'s own 8 tests all passed because none of them exercised THIS wiring; each
|
||||
* one hands the gauge its own directory directly.
|
||||
*
|
||||
* <p>This class proves the directory {@code fleet.leaders.<name>.profile}'s own {@code configDir:}
|
||||
* names is the one {@code fleet_list}'s {@code context} row actually reads — not merely that SOME
|
||||
* directory got passed. If the wiring in {@code FleetMcp.leadView}/{@code contextView} is ever
|
||||
* reverted to a hardcoded {@code null}, {@link #configuredDirectoryDecidesWhichTranscriptIsRead()}
|
||||
* must go red: both directories below hold a REAL, DIFFERENT token count for the SAME session id,
|
||||
* so a hardcoded {@code null} (always reading the built-in default, where neither transcript lives)
|
||||
* would report {@code UNKNOWN} both times instead of the two distinct numbers this test asserts.
|
||||
*/
|
||||
class FleetMcpLeadContextGaugeWiringTest {
|
||||
|
||||
private static final String LEAD_TERMINAL = "term_lead";
|
||||
private static final String LEAD_NAME = "opus";
|
||||
/** {@code FakeHerdr.withAgent} always projects {@code agent_session.value} as {@code sess-<terminalId>}. */
|
||||
private static final String LEAD_SESSION_ID = "sess-" + LEAD_TERMINAL;
|
||||
|
||||
private static ClaudeCodeLauncher workerService(FakeHerdr h) {
|
||||
FleetConfig.Profile cfg = new FleetConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
|
||||
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> "tok");
|
||||
}
|
||||
|
||||
private static String textOf(McpSchema.CallToolResult r) {
|
||||
return ((McpSchema.TextContent) r.content().getFirst()).text();
|
||||
}
|
||||
|
||||
/** One line Claude Code would write for a turn with the given live-context total. */
|
||||
private static String usageLine(long tokens) {
|
||||
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
|
||||
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
|
||||
}
|
||||
|
||||
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} carrying one usage record. */
|
||||
private static String writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
|
||||
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
|
||||
Files.createDirectories(projectDir);
|
||||
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
|
||||
return configDir.toString();
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("the CONFIGURED directory decides which transcript is read, and follows a config change")
|
||||
void configuredDirectoryDecidesWhichTranscriptIsRead(@TempDir Path tmp) throws IOException {
|
||||
Path dirA = Files.createDirectory(tmp.resolve("dir-a"));
|
||||
Path dirB = Files.createDirectory(tmp.resolve("dir-b"));
|
||||
writeTranscript(dirA, LEAD_SESSION_ID, 11_000);
|
||||
writeTranscript(dirB, LEAD_SESSION_ID, 22_000);
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr().withAgent("lead-opus", LEAD_TERMINAL, "wL:p1", "wL:t1");
|
||||
SessionManager sessions = new SessionManager(workerService(herdr));
|
||||
LeadContextGauge contextGauge = new LeadContextGauge();
|
||||
AtomicReference<String> configuredDir = new AtomicReference<>(dirA.toString());
|
||||
FleetMcp.LeadConfigDirSource source = new FleetMcp.LeadConfigDirSource(name ->
|
||||
LEAD_NAME.equals(name) ? configuredDir.get() : null);
|
||||
|
||||
String firstRead = textOf(listFleet(herdr, sessions, contextGauge, source));
|
||||
assertTrue(firstRead.contains("\"tokens\":11000"),
|
||||
"config names directory A ⇒ fleet_list must report A's token count: " + firstRead);
|
||||
|
||||
// The config changes to name directory B instead -- the very next read must follow it.
|
||||
configuredDir.set(dirB.toString());
|
||||
String secondRead = textOf(listFleet(herdr, sessions, contextGauge, source));
|
||||
assertTrue(secondRead.contains("\"tokens\":22000"),
|
||||
"config now names directory B ⇒ fleet_list must report B's token count, not A's stale one: "
|
||||
+ secondRead);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a lead whose config names no directory still degrades to the built-in default, never throws")
|
||||
void aLeadWhoseConfigNamesNoDirectoryStillFallsBackWithoutThrowing() {
|
||||
FakeHerdr herdr = new FakeHerdr().withAgent("lead-opus", LEAD_TERMINAL, "wL:p1", "wL:t1");
|
||||
SessionManager sessions = new SessionManager(workerService(herdr));
|
||||
LeadContextGauge contextGauge = new LeadContextGauge();
|
||||
|
||||
McpSchema.CallToolResult res = listFleet(herdr, sessions, contextGauge, FleetMcp.LeadConfigDirSource.none());
|
||||
|
||||
assertNotEquals(Boolean.TRUE, res.isError(), "a lead with no configured configDir must degrade, never throw");
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"state\":\"unknown\""), "no transcript under the built-in default ⇒ UNKNOWN: " + out);
|
||||
assertFalse(out.contains("\"tokens\""), "UNKNOWN must never carry a stale/default token number: " + out);
|
||||
}
|
||||
|
||||
private static McpSchema.CallToolResult listFleet(FakeHerdr herdr, SessionManager sessions,
|
||||
LeadContextGauge contextGauge, FleetMcp.LeadConfigDirSource leadConfigDirs) {
|
||||
return FleetMcp.listFleet(workerService(herdr), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
|
||||
FleetMcp.LeadSeatSource.none(), contextGauge, leadConfigDirs,
|
||||
Map.of(LEAD_TERMINAL, LEAD_NAME), LEAD_TERMINAL, FleetMcp.CoordinationSource.none(), false);
|
||||
}
|
||||
}
|
||||
@@ -7,7 +7,9 @@ import dev.ltms.fleet.auth.Role;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.AgentStatus;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.inject.LoopWatchdog;
|
||||
import dev.ltms.fleet.herdr.PaneLocator;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
@@ -57,9 +59,10 @@ class FleetMcpTest {
|
||||
|
||||
private final FakeHerdr herdr = new FakeHerdr();
|
||||
private final AgentControl agents = new AgentControl(herdr);
|
||||
private final Injector injector = new Injector(agents);
|
||||
private final Rendezvous rendezvous = new Rendezvous();
|
||||
private final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||
private final MessageService messages = new MessageService(agents, new Injector(agents), rendezvous, inbox);
|
||||
private final MessageService messages = new MessageService(agents, injector, rendezvous, inbox);
|
||||
|
||||
@BeforeEach
|
||||
void setUp() {
|
||||
@@ -298,6 +301,35 @@ class FleetMcpTest {
|
||||
assertTrue(textOf(res).contains("no reply"), "got: " + textOf(res));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #571 (ticket CORRECTION 5): {@code formatReply}'s {@code TIMED_OUT_UNCONFIRMED} arm is
|
||||
* the one message whose whole job is to stop a caller retrying a delivery that may already have
|
||||
* arrived. Pin that its wording is actually distinct from the queued/working arm's retry
|
||||
* invitation — a mutation that swapped this arm's text for that one still passed every other
|
||||
* test in this suite, because nothing asserted the specific wording.
|
||||
*/
|
||||
@Test
|
||||
void sendTimesOutWithAnUnconfirmedNoteNotARetryInvitation() throws Exception {
|
||||
herdr.agentSendFailsWith("send_failed");
|
||||
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
|
||||
() -> FleetMcp.send(messages, T, "hi", 150L, null, Set.of()));
|
||||
long deadline = System.currentTimeMillis() + 2000;
|
||||
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting(T), "send should have opened its rendezvous waiter");
|
||||
injector.onStatus(T, AgentStatus.IDLE); // triggers the failing delivery attempt -> ATTEMPTED
|
||||
|
||||
McpSchema.CallToolResult res = send.get(5, TimeUnit.SECONDS);
|
||||
String text = textOf(res);
|
||||
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
|
||||
assertTrue(text.contains("delivery unconfirmed"), "got: " + text);
|
||||
assertFalse(text.contains("retry or poll status"),
|
||||
"an unconfirmed delivery must not carry the queued/working arm's retry invitation — "
|
||||
+ "a resend here can double-deliver the same brief: got " + text);
|
||||
}
|
||||
|
||||
@Test
|
||||
void sendRejectsMissingArgs() {
|
||||
assertTrue(FleetMcp.send(messages, null, "hi", null, null, Set.of()).isError());
|
||||
@@ -1053,6 +1085,36 @@ class FleetMcpTest {
|
||||
assertFalse(out.contains("quarantinedForSeconds"), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void loopHealthReportsStalledStatusPoller() {
|
||||
String out = loopHealth(LoopWatchdog.State.STALLED, LoopWatchdog.State.RUNNING);
|
||||
assertTrue(out.contains("\"statusPoller\":\"STALLED\""),
|
||||
"fleet_list must report a stalled StatusPoller: " + out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void loopHealthReportsStoppedSessionReaperAsStopped() {
|
||||
String out = loopHealth(LoopWatchdog.State.RUNNING, LoopWatchdog.State.STOPPED);
|
||||
assertTrue(out.contains("\"sessionReaper\":\"STOPPED\""),
|
||||
"fleet_list must report a deliberately stopped SessionReaper as STOPPED, not an alarm: " + out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void loopHealthReportsRunningStatusPoller() {
|
||||
String out = loopHealth(LoopWatchdog.State.RUNNING, LoopWatchdog.State.STOPPED);
|
||||
assertTrue(out.contains("\"statusPoller\":\"RUNNING\""),
|
||||
"fleet_list must report a running StatusPoller: " + out);
|
||||
}
|
||||
|
||||
private static String loopHealth(LoopWatchdog.State statusPoller, LoopWatchdog.State sessionReaper) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
return textOf(FleetMcp.listFleet(workerService(herdr, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||
new SessionManager(workerService(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"))), null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
new FleetMcp.LoopHealthSource(() -> statusPoller, () -> sessionReaper),
|
||||
FleetMcp.QuarantineSource.none(), Map.of(), ""));
|
||||
}
|
||||
|
||||
@Test
|
||||
void capacityIncludesConfiguredProfileWithoutMembers() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
@@ -1875,7 +1937,7 @@ class FleetMcpTest {
|
||||
null, null, null, null, null, null);
|
||||
|
||||
assertEquals(Boolean.TRUE, res.isError());
|
||||
assertTrue(textOf(res).contains("architect, dev, reviewer"), textOf(res));
|
||||
assertTrue(textOf(res).contains("architect, dev, hunter, reviewer"), textOf(res));
|
||||
}
|
||||
|
||||
// ── CB-619 / fleetd #123: a spawn asking for a role its profile has no slot for must be
|
||||
|
||||
@@ -556,6 +556,31 @@ class ClaudeCodeLauncherTest {
|
||||
"no --agent flag when the role has no agent-definition file");
|
||||
}
|
||||
|
||||
@Test
|
||||
void hunterRoleUsesItsAgentFileAndStopsUsingItWhenRemoved(@TempDir Path cwd) throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path agentFile = Files.createDirectories(cwd.resolve(".claude/agents")).resolve("hunter.md");
|
||||
Files.writeString(agentFile, "---\nname: hunter\n---\nSweep for defects.");
|
||||
FleetConfig.Profile cfg = new FleetConfig.Profile(
|
||||
"sonnet", "http://gx00.gw:8000", null, null, "FLEETD_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "fleetd-workers", "w #{n}", null, null, null);
|
||||
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
|
||||
svc.spawn(new SpawnRequest("sonnet", cwd.toString(), null, null, null, MemberRole.HUNTER));
|
||||
|
||||
List<String> args = spawnedArgs(herdr);
|
||||
int flag = args.indexOf("--agent");
|
||||
assertTrue(flag >= 0, "the hunter role reaches its agent-definition file: " + args);
|
||||
assertEquals("hunter", args.get(flag + 1));
|
||||
|
||||
Files.delete(agentFile);
|
||||
svc.spawn(new SpawnRequest("sonnet", cwd.toString(), null, null, null, MemberRole.HUNTER));
|
||||
|
||||
assertFalse(spawnedArgs(herdr).contains("--agent"),
|
||||
"the hunter role no longer gets an agent when its file is removed");
|
||||
}
|
||||
|
||||
private ClaudeCodeLauncher multiProfile(FakeHerdr herdr) {
|
||||
FleetConfig.Profile gx10 = new FleetConfig.Profile("gx10", "http://gx10.gw:8000", "coder",
|
||||
null, "FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers", "w #{n}", null, null, null);
|
||||
|
||||
@@ -1,17 +1,17 @@
|
||||
package dev.ltms.fleet.msg;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import com.rabbitmq.client.Channel;
|
||||
import com.rabbitmq.client.ConnectionFactory;
|
||||
import com.rabbitmq.client.impl.DefaultExceptionHandler;
|
||||
import dev.ltms.fleet.testing.CapturedLog;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.lang.reflect.Proxy;
|
||||
import java.util.List;
|
||||
import java.util.concurrent.atomic.AtomicInteger;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
@@ -32,21 +32,17 @@ class AmqpConnectionFailureLoggerTest {
|
||||
assertEquals(AmqpConnectionFailureLogger.REPLY_INBOX, inboxHandler.connectionName());
|
||||
assertEquals(AmqpConnectionFailureLogger.LEAD_MAILBOX, mailboxHandler.connectionName());
|
||||
|
||||
ListAppender<ILoggingEvent> inboxEvents = attach(AmqpReplyInbox.class);
|
||||
ListAppender<ILoggingEvent> mailboxEvents = attach(LeadMailbox.class);
|
||||
IllegalStateException inboxFailure = new IllegalStateException("inbox failure");
|
||||
IllegalStateException mailboxFailure = new IllegalStateException("mailbox failure");
|
||||
try {
|
||||
try (CapturedLog inboxLog = attach(AmqpReplyInbox.class);
|
||||
CapturedLog mailboxLog = attach(LeadMailbox.class)) {
|
||||
inboxHandler.handleUnexpectedConnectionDriverException(null, inboxFailure);
|
||||
mailboxHandler.handleConnectionRecoveryException(null, mailboxFailure);
|
||||
|
||||
assertError(inboxEvents, "AMQP connection fleetd-reply-inbox: An unexpected connection driver error occurred",
|
||||
assertError(inboxLog.events(), "AMQP connection fleetd-reply-inbox: An unexpected connection driver error occurred",
|
||||
inboxFailure, "inbox failure line");
|
||||
assertError(mailboxEvents, "AMQP connection fleetd-lead-mailbox: Caught an exception during connection recovery!",
|
||||
assertError(mailboxLog.events(), "AMQP connection fleetd-lead-mailbox: Caught an exception during connection recovery!",
|
||||
mailboxFailure, "mailbox recovery line");
|
||||
} finally {
|
||||
detach(AmqpReplyInbox.class, inboxEvents);
|
||||
detach(LeadMailbox.class, mailboxEvents);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -54,17 +50,14 @@ class AmqpConnectionFailureLoggerTest {
|
||||
void connectionResetKeepsForgivingHandlerWarningSemantics() {
|
||||
AmqpConnectionFailureLogger handler = new AmqpConnectionFailureLogger(
|
||||
AmqpConnectionFailureLogger.REPLY_INBOX, LoggerFactory.getLogger(AmqpReplyInbox.class));
|
||||
ListAppender<ILoggingEvent> events = attach(AmqpReplyInbox.class);
|
||||
try {
|
||||
try (CapturedLog log = attach(AmqpReplyInbox.class)) {
|
||||
handler.handleUnexpectedConnectionDriverException(null, new IOException("Connection reset"));
|
||||
assertEquals(1, events.list.size(), "the handler must still log a reset");
|
||||
ILoggingEvent event = events.list.getFirst();
|
||||
assertEquals(1, log.events().size(), "the handler must still log a reset");
|
||||
ILoggingEvent event = log.events().getFirst();
|
||||
assertEquals(Level.WARN, event.getLevel(), "ForgivingExceptionHandler logs connection resets at WARN");
|
||||
assertEquals("AMQP connection fleetd-reply-inbox: An unexpected connection driver error occurred "
|
||||
+ "(Exception message: Connection reset)", event.getFormattedMessage());
|
||||
assertTrue(event.getThrowableProxy() == null, "ForgivingExceptionHandler does not attach a reset stack trace");
|
||||
} finally {
|
||||
detach(AmqpReplyInbox.class, events);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -103,17 +96,8 @@ class AmqpConnectionFailureLoggerTest {
|
||||
"all exception-handling methods must remain inherited from DefaultExceptionHandler");
|
||||
}
|
||||
|
||||
private static ListAppender<ILoggingEvent> attach(Class<?> owner) {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(owner);
|
||||
logger.setLevel(Level.DEBUG);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
return appender;
|
||||
}
|
||||
|
||||
private static void detach(Class<?> owner, ListAppender<ILoggingEvent> appender) {
|
||||
((Logger) LoggerFactory.getLogger(owner)).detachAppender(appender);
|
||||
private static CapturedLog attach(Class<?> owner) {
|
||||
return CapturedLog.at(owner, Level.DEBUG);
|
||||
}
|
||||
|
||||
private static AmqpConnectionFailureLogger installedStrictHandler(ConnectionFactory factory, String connection) {
|
||||
@@ -123,9 +107,9 @@ class AmqpConnectionFailureLoggerTest {
|
||||
return assertInstanceOf(AmqpConnectionFailureLogger.class, factory.getExceptionHandler());
|
||||
}
|
||||
|
||||
private static void assertError(ListAppender<ILoggingEvent> events, String message, Throwable cause, String name) {
|
||||
assertEquals(1, events.list.size(), name);
|
||||
ILoggingEvent event = events.list.getFirst();
|
||||
private static void assertError(List<ILoggingEvent> events, String message, Throwable cause, String name) {
|
||||
assertEquals(1, events.size(), name);
|
||||
ILoggingEvent event = events.getFirst();
|
||||
assertEquals(Level.ERROR, event.getLevel(), name);
|
||||
assertEquals(message, event.getFormattedMessage(), name);
|
||||
assertEquals(cause.toString(), event.getThrowableProxy().getClassName() + ": "
|
||||
|
||||
@@ -8,7 +8,10 @@ import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.Timeout;
|
||||
|
||||
import java.lang.reflect.InvocationHandler;
|
||||
import java.lang.reflect.Method;
|
||||
import java.lang.reflect.Proxy;
|
||||
import java.io.IOException;
|
||||
import java.util.Map;
|
||||
import java.util.List;
|
||||
import java.util.concurrent.CopyOnWriteArrayList;
|
||||
import java.util.concurrent.CountDownLatch;
|
||||
@@ -20,6 +23,7 @@ import java.util.concurrent.atomic.AtomicReference;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* CB-528 follow-up: {@link AmqpReplyInbox#failPendingPublishesOnRecovery()} must not fail a publish
|
||||
@@ -197,11 +201,78 @@ class AmqpReplyInboxRecoveryRaceTest {
|
||||
+ elapsedMillis.get() + "ms");
|
||||
}
|
||||
|
||||
@Test
|
||||
void publishIOExceptionRemovesThePendingMessageId() throws Exception {
|
||||
AtomicLong seqCounter = new AtomicLong();
|
||||
Channel failing = fakeChannel(seqCounter, new CopyOnWriteArrayList<>(), new CopyOnWriteArrayList<>(),
|
||||
new AtomicReference<>(), new AtomicReference<>(), true);
|
||||
AmqpReplyInbox inbox = new AmqpReplyInbox(fakeConnection(failing, failing), AmqpReplyInbox.DEFAULT_PREFETCH);
|
||||
|
||||
try {
|
||||
org.junit.jupiter.api.Assertions.assertThrows(IllegalStateException.class,
|
||||
() -> inbox.publish("worker", "catch", "body"));
|
||||
assertEquals(0, pendingByMsgId(inbox).size(), "publish IOException must remove its msgId entry");
|
||||
} finally {
|
||||
inbox.close();
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void interruptedPublishRemovesThePendingMessageIdInFinally() throws Exception {
|
||||
InboxFixture fixture = new InboxFixture();
|
||||
Thread publish = fixture.startPublish("finally");
|
||||
fixture.awaitPublished("finally");
|
||||
publish.interrupt();
|
||||
publish.join(5_000);
|
||||
|
||||
assertEquals(0, pendingByMsgId(fixture.inbox).size(), "publish finally must remove its msgId entry");
|
||||
fixture.inbox.close();
|
||||
}
|
||||
|
||||
@Test
|
||||
void confirmResolutionRemovesThePendingMessageId() throws Exception {
|
||||
AmqpReplyInbox inbox = new InboxFixture().inbox;
|
||||
try {
|
||||
seedPending(inbox, 1, "confirm");
|
||||
invoke(inbox, "resolveConfirm", new Class<?>[] {long.class, boolean.class, boolean.class}, 1L, false, true);
|
||||
assertEquals(0, pendingByMsgId(inbox).size(), "confirm resolution must remove its msgId entry");
|
||||
} finally {
|
||||
inbox.close();
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void recoverySweepRemovesThePendingMessageId() throws Exception {
|
||||
AmqpReplyInbox inbox = new InboxFixture().inbox;
|
||||
try {
|
||||
seedPending(inbox, 1, "recovery");
|
||||
inbox.failPendingPublishesOnRecovery();
|
||||
assertEquals(0, pendingByMsgId(inbox).size(), "recovery sweep must remove its msgId entry");
|
||||
} finally {
|
||||
inbox.close();
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void closeRemovesThePendingMessageId() throws Exception {
|
||||
AmqpReplyInbox inbox = new InboxFixture().inbox;
|
||||
seedPending(inbox, 1, "close");
|
||||
inbox.close();
|
||||
|
||||
assertEquals(0, pendingByMsgId(inbox).size(), "close must remove its msgId entry");
|
||||
}
|
||||
|
||||
/** A {@link Proxy}-backed {@link Channel}: only the calls {@link AmqpReplyInbox} actually makes
|
||||
* are meaningfully implemented; everything else returns a harmless default. */
|
||||
private static Channel fakeChannel(AtomicLong seqCounter, List<Long> seqOrder, List<String> msgIdOrder,
|
||||
AtomicReference<ConfirmCallback> ackCallback,
|
||||
AtomicReference<ConfirmCallback> nackCallback) {
|
||||
AtomicReference<ConfirmCallback> ackCallback,
|
||||
AtomicReference<ConfirmCallback> nackCallback) {
|
||||
return fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback, false);
|
||||
}
|
||||
|
||||
private static Channel fakeChannel(AtomicLong seqCounter, List<Long> seqOrder, List<String> msgIdOrder,
|
||||
AtomicReference<ConfirmCallback> ackCallback,
|
||||
AtomicReference<ConfirmCallback> nackCallback, boolean failPublish) {
|
||||
InvocationHandler handler = (proxy, method, args) -> {
|
||||
String name = method.getName();
|
||||
if (name.equals("getNextPublishSeqNo")) {
|
||||
@@ -210,6 +281,9 @@ class AmqpReplyInboxRecoveryRaceTest {
|
||||
return value;
|
||||
}
|
||||
if (name.equals("basicPublish")) {
|
||||
if (failPublish) {
|
||||
throw new IOException("test publish failure");
|
||||
}
|
||||
AMQP.BasicProperties props = (AMQP.BasicProperties) args[3];
|
||||
msgIdOrder.add(props.getMessageId());
|
||||
return null;
|
||||
@@ -285,4 +359,60 @@ class AmqpReplyInboxRecoveryRaceTest {
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static Map<String, Object> pendingByMsgId(AmqpReplyInbox inbox) throws Exception {
|
||||
var field = AmqpReplyInbox.class.getDeclaredField("pendingByMsgId");
|
||||
field.setAccessible(true);
|
||||
return (Map<String, Object>) field.get(inbox);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static void seedPending(AmqpReplyInbox inbox, long seq, String msgId) throws Exception {
|
||||
Class<?> pendingType = Class.forName(AmqpReplyInbox.class.getName() + "$Pending");
|
||||
var constructor = pendingType.getDeclaredConstructor(String.class);
|
||||
constructor.setAccessible(true);
|
||||
Object pending = constructor.newInstance(msgId);
|
||||
var seqField = AmqpReplyInbox.class.getDeclaredField("pendingBySeq");
|
||||
seqField.setAccessible(true);
|
||||
((Map<Long, Object>) seqField.get(inbox)).put(seq, pending);
|
||||
pendingByMsgId(inbox).put(msgId, pending);
|
||||
}
|
||||
|
||||
private static void invoke(AmqpReplyInbox inbox, String name, Class<?>[] types, Object... args) throws Exception {
|
||||
Method method = AmqpReplyInbox.class.getDeclaredMethod(name, types);
|
||||
method.setAccessible(true);
|
||||
method.invoke(inbox, args);
|
||||
}
|
||||
|
||||
private static final class InboxFixture {
|
||||
final AtomicLong seqCounter = new AtomicLong();
|
||||
final List<Long> seqOrder = new CopyOnWriteArrayList<>();
|
||||
final List<String> msgIdOrder = new CopyOnWriteArrayList<>();
|
||||
final AtomicReference<ConfirmCallback> ackCallback = new AtomicReference<>();
|
||||
final AtomicReference<ConfirmCallback> nackCallback = new AtomicReference<>();
|
||||
final AmqpReplyInbox inbox = new AmqpReplyInbox(
|
||||
fakeConnection(fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback),
|
||||
fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback)),
|
||||
AmqpReplyInbox.DEFAULT_PREFETCH);
|
||||
|
||||
Thread startPublish(String msgId) {
|
||||
Thread thread = Thread.ofVirtual().start(() -> {
|
||||
try {
|
||||
inbox.publish("worker", msgId, "body");
|
||||
} catch (IllegalStateException ignored) {
|
||||
// Interrupting the confirm wait is the path under test.
|
||||
}
|
||||
});
|
||||
return thread;
|
||||
}
|
||||
|
||||
void awaitPublished(String msgId) throws InterruptedException {
|
||||
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(5);
|
||||
while (!msgIdOrder.contains(msgId) && System.nanoTime() < deadline) {
|
||||
Thread.sleep(10);
|
||||
}
|
||||
assertTrue(msgIdOrder.contains(msgId), "publish did not register " + msgId);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,27 +1,46 @@
|
||||
package dev.ltms.fleet.msg;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.AgentStatus;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||
import dev.ltms.fleet.mcp.PrimaryRegistry;
|
||||
import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* Unit tests for the CB-551 idle-lead heartbeat: the pure {@link LeadHeartbeatLoop#decide} decision
|
||||
* function, the nudge text, and the {@link LeadHeartbeatLoop#snapshot} fleet snapshot.
|
||||
* function, the nudge text, and the {@link LeadHeartbeatLoop#snapshot} fleet snapshot. Also covers
|
||||
* fleetd #609's context-high notice: the {@code context}/{@code contextNotified} parameters {@code
|
||||
* decide} gained, and the standalone {@link LeadHeartbeatLoop#contextNotice} text builder.
|
||||
*
|
||||
* <p>All decision tests call {@code decide} directly with explicit nanoTime values from an injected
|
||||
* clock — no sleeping, no scheduler races. This mirrors how {@code ReplyPushLoopTest} pins the pure
|
||||
* decision before exercising the loop.
|
||||
*
|
||||
* <p>Every pre-#609 test below passes {@code LeadContextGauge.State.UNKNOWN, false} for the two new
|
||||
* {@code decide} parameters — the same as a lead whose context could not be read and has never been
|
||||
* notified — so each one still pins exactly the behaviour it pinned before this ticket.
|
||||
*/
|
||||
class LeadHeartbeatLoopTest {
|
||||
|
||||
@@ -56,12 +75,18 @@ class LeadHeartbeatLoopTest {
|
||||
return new LeadHeartbeatLoop.FleetState(2, 1, 3, List.of("term_a", "term_b"));
|
||||
}
|
||||
|
||||
/** The loop under test; the scheduler is never invoked on the pure decide path. */
|
||||
/** The loop under test, context notice off; the scheduler is never invoked on the pure decide path. */
|
||||
private static LeadHeartbeatLoop loop(int quietCap) {
|
||||
return loop(quietCap, false);
|
||||
}
|
||||
|
||||
/** As above, with the fleetd #609 {@code contextHighNudge} flag set explicitly. */
|
||||
private static LeadHeartbeatLoop loop(int quietCap, boolean contextHighNudge) {
|
||||
return new LeadHeartbeatLoop(
|
||||
new PrimaryRegistry("term_lead"), null /*agents — unused on the decide path*/,
|
||||
null /*inbox*/, List::of, null /*pushLoop*/, null /*scheduler*/, () -> 0L,
|
||||
IDLE_AFTER_NANOS, 1_000L, quietCap);
|
||||
IDLE_AFTER_NANOS, 1_000L, quietCap, null,
|
||||
LeadHeartbeatLoop.LeadContextSource.none(), contextHighNudge);
|
||||
}
|
||||
|
||||
// ── (a) a working lead is never injected ───────────────────────────────────────────────────
|
||||
@@ -69,7 +94,8 @@ class LeadHeartbeatLoopTest {
|
||||
@Test
|
||||
void aWorkingLeadIsNeverInjected() {
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet());
|
||||
NOW, IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet(),
|
||||
LeadContextGauge.State.UNKNOWN, false);
|
||||
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(),
|
||||
"a WORKING lead is making progress and must not be touched");
|
||||
assertNull(d.idleSinceNanos(), "a busy lead resets the idle window");
|
||||
@@ -80,17 +106,71 @@ class LeadHeartbeatLoopTest {
|
||||
void anUnreadableStatusIsNeverInjectedEither() {
|
||||
// A failed status read (or a gone agent) must degrade to "do not inject", never hammer the pane.
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, IDLE_PAST, 0, AgentStatus.UNKNOWN, false, true, pendingFleet());
|
||||
NOW, IDLE_PAST, 0, AgentStatus.UNKNOWN, false, true, pendingFleet(),
|
||||
LeadContextGauge.State.UNKNOWN, false);
|
||||
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(),
|
||||
"never inject into a state the loop cannot read");
|
||||
}
|
||||
|
||||
// ── fleetd #613: the boot line names all four heartbeat settings ──────────────────────────
|
||||
|
||||
/**
|
||||
* fleetd #613: {@link LeadHeartbeatLoop#start()}'s boot line named only 3 of the 4 constructor
|
||||
* settings — {@code contextHighNudge} (fleetd #609) was missing, so an operator could not tell
|
||||
* from the log whether the handover notice was armed. Captures the real log via a
|
||||
* {@link ListAppender}, raising the logger's level past {@code logback-test.xml}'s
|
||||
* {@code dev.ltms.fleet -> WARN} override for the duration of the call — the same seam {@code
|
||||
* GitHostShapeReportTest#attach} uses for its own INFO-level boot line.
|
||||
*/
|
||||
private static String heartbeatBootLine(boolean contextHighNudge, ScheduledExecutorService scheduler) {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(LeadHeartbeatLoop.class);
|
||||
Level originalLevel = logger.getLevel();
|
||||
logger.setLevel(Level.INFO);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
try {
|
||||
LeadHeartbeatLoop l = new LeadHeartbeatLoop(
|
||||
new PrimaryRegistry("term_lead"), null /*agents*/, null /*inbox*/, List::of,
|
||||
null /*pushLoop*/, scheduler, () -> 0L, IDLE_AFTER_NANOS, 1_000L, 3, null,
|
||||
LeadHeartbeatLoop.LeadContextSource.none(), contextHighNudge);
|
||||
l.start();
|
||||
} finally {
|
||||
logger.detachAppender(appender);
|
||||
logger.setLevel(originalLevel);
|
||||
}
|
||||
return appender.list.stream()
|
||||
.filter(e -> e.getLevel() == Level.INFO)
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.filter(m -> m.startsWith("idle-lead heartbeat:"))
|
||||
.findFirst()
|
||||
.orElseThrow(() -> new AssertionError("expected the heartbeat boot line to be logged"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void theBootLineNamesContextHighNudgeWhenArmed() {
|
||||
String line = heartbeatBootLine(true, scheduler);
|
||||
assertTrue(line.contains("300s idle"), line);
|
||||
assertTrue(line.contains("recheck 1000ms"), line);
|
||||
assertTrue(line.contains("quiet cap 3"), line);
|
||||
assertTrue(line.contains("context-high nudge true"),
|
||||
"the boot line must name the 4th setting, contextHighNudge, alongside the other "
|
||||
+ "three: " + line);
|
||||
}
|
||||
|
||||
@Test
|
||||
void theBootLineNamesContextHighNudgeWhenOff() {
|
||||
String line = heartbeatBootLine(false, scheduler);
|
||||
assertTrue(line.contains("context-high nudge false"), line);
|
||||
}
|
||||
|
||||
// ── (b) an idle lead within the quiet period is not yet injected ───────────────────────────
|
||||
|
||||
@Test
|
||||
void justBecameIdleStartsTheDebounceWindow() {
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, null, 0, AgentStatus.IDLE, false, true, quietFleet());
|
||||
NOW, null, 0, AgentStatus.IDLE, false, true, quietFleet(),
|
||||
LeadContextGauge.State.UNKNOWN, false);
|
||||
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
|
||||
"the first injectable tick only records the start of the idle stretch");
|
||||
assertEquals(NOW, d.idleSinceNanos(), "the idle window opens at the moment the lead became injectable");
|
||||
@@ -99,7 +179,8 @@ class LeadHeartbeatLoopTest {
|
||||
@Test
|
||||
void idleWithinQuietPeriodIsNotInjected() {
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, IDLE_WITHIN, 0, AgentStatus.IDLE, false, true, quietFleet());
|
||||
NOW, IDLE_WITHIN, 0, AgentStatus.IDLE, false, true, quietFleet(),
|
||||
LeadContextGauge.State.UNKNOWN, false);
|
||||
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
|
||||
"a lead idle for 10s (< 300s) has just finished a turn — do not re-prompt it");
|
||||
}
|
||||
@@ -109,7 +190,8 @@ class LeadHeartbeatLoopTest {
|
||||
@Test
|
||||
void idlePastQuietPeriodIsInjected() {
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, true, quietFleet());
|
||||
NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, true, quietFleet(),
|
||||
LeadContextGauge.State.UNKNOWN, false);
|
||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
|
||||
"a lead continuously idle past the quiet period is the reason to nudge");
|
||||
}
|
||||
@@ -117,14 +199,17 @@ class LeadHeartbeatLoopTest {
|
||||
@Test
|
||||
void blockedAndDoneAreInjectableViewsOfIdle() {
|
||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, loop(3).decide(
|
||||
NOW, IDLE_PAST, 0, AgentStatus.BLOCKED, false, true, quietFleet()).action());
|
||||
NOW, IDLE_PAST, 0, AgentStatus.BLOCKED, false, true, quietFleet(),
|
||||
LeadContextGauge.State.UNKNOWN, false).action());
|
||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, loop(3).decide(
|
||||
NOW, IDLE_PAST, 0, AgentStatus.DONE, false, true, quietFleet()).action());
|
||||
NOW, IDLE_PAST, 0, AgentStatus.DONE, false, true, quietFleet(),
|
||||
LeadContextGauge.State.UNKNOWN, false).action());
|
||||
}
|
||||
|
||||
@Test
|
||||
void nothingIsInjectedWhenNoLeadIsKnown() {
|
||||
var d = loop(3).decide(NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, false, pendingFleet());
|
||||
var d = loop(3).decide(NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, false, pendingFleet(),
|
||||
LeadContextGauge.State.UNKNOWN, false);
|
||||
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
|
||||
"with no known lead there is nobody to nudge — keep waiting until one is discovered");
|
||||
assertEquals(IDLE_PAST, d.idleSinceNanos(), "the idle window stays open so discovery re-arms it");
|
||||
@@ -135,7 +220,8 @@ class LeadHeartbeatLoopTest {
|
||||
@Test
|
||||
void quietNudgeCapStopsTheLoopWhenNothingIsPending() {
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet());
|
||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
|
||||
LeadContextGauge.State.UNKNOWN, false);
|
||||
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, d.action(),
|
||||
"3 consecutive nothing-pending nudges have already happened — stop nagging an empty fleet");
|
||||
}
|
||||
@@ -145,7 +231,8 @@ class LeadHeartbeatLoopTest {
|
||||
@Test
|
||||
void newPendingStateResetsTheQuietCap() {
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet());
|
||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet(),
|
||||
LeadContextGauge.State.UNKNOWN, false);
|
||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
|
||||
"real state appearing re-arms the loop past an exhausted cap");
|
||||
assertEquals(0, d.quietCount(), "the pending state resets the consecutive-quiet counter");
|
||||
@@ -156,7 +243,8 @@ class LeadHeartbeatLoopTest {
|
||||
@Test
|
||||
void standsDownWhileReplyPushLoopIsActive() {
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet());
|
||||
NOW, IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet(),
|
||||
LeadContextGauge.State.UNKNOWN, false);
|
||||
assertEquals(LeadHeartbeatLoop.Action.STAND_DOWN, d.action(),
|
||||
"a second injection would start a competing turn — stand aside instead");
|
||||
assertEquals(0, d.quietCount(), "the active push is real state, so it re-arms the cap");
|
||||
@@ -213,4 +301,398 @@ class LeadHeartbeatLoopTest {
|
||||
assertFalse(fs.hasPending());
|
||||
assertEquals(0, fs.liveWorkers());
|
||||
}
|
||||
|
||||
// ── fleetd #609: context-high notice ───────────────────────────────────────────────────────
|
||||
|
||||
/** One gate scenario, keyed by name, over which the off-state/context-independence property is checked. */
|
||||
private record Gate(String name, Long idleSince, int quietCount, AgentStatus status,
|
||||
boolean pushLoopActive, boolean leadKnown, LeadHeartbeatLoop.FleetState fleet) {}
|
||||
|
||||
private static List<Gate> allEightGates() {
|
||||
return List.of(
|
||||
new Gate("1-standDown", IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet()),
|
||||
new Gate("2-leadBusy", IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet()),
|
||||
new Gate("3-justBecameIdle", null, 0, AgentStatus.IDLE, false, true, quietFleet()),
|
||||
new Gate("4-withinQuietPeriod", IDLE_WITHIN, 0, AgentStatus.IDLE, false, true, quietFleet()),
|
||||
new Gate("5-noLeadKnown", IDLE_PAST, 0, AgentStatus.IDLE, false, false, pendingFleet()),
|
||||
new Gate("6-pending", IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet()),
|
||||
new Gate("7-quietNotExhausted", IDLE_PAST, 0, AgentStatus.IDLE, false, true, quietFleet()),
|
||||
new Gate("8-quietExhausted", IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet()));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aOffFlagIgnoresContextAcrossAllEightGates() {
|
||||
// Property A: with contextHighNudge off, decide() must not depend on the context reading at
|
||||
// all — not just "usually agrees", but identical Action/idleSinceNanos/quietCount whatever
|
||||
// context state is passed, and contextNotified must stay false throughout. This is the proof
|
||||
// that fleetd #609 is opt-in: a daemon upgraded to carry this code, but never configuring
|
||||
// `contextHighNudge: true`, behaves exactly as it did before this ticket for every one of the
|
||||
// 8 gates the class javadoc numbers.
|
||||
LeadHeartbeatLoop offLoop = loop(3, false);
|
||||
for (Gate g : allEightGates()) {
|
||||
var withHigh = offLoop.decide(NOW, g.idleSince(), g.quietCount(), g.status(),
|
||||
g.pushLoopActive(), g.leadKnown(), g.fleet(), LeadContextGauge.State.HIGH, false);
|
||||
var withUnknown = offLoop.decide(NOW, g.idleSince(), g.quietCount(), g.status(),
|
||||
g.pushLoopActive(), g.leadKnown(), g.fleet(), LeadContextGauge.State.UNKNOWN, false);
|
||||
var withOk = offLoop.decide(NOW, g.idleSince(), g.quietCount(), g.status(),
|
||||
g.pushLoopActive(), g.leadKnown(), g.fleet(), LeadContextGauge.State.OK, false);
|
||||
|
||||
assertEquals(withUnknown.action(), withHigh.action(), g.name() + ": action must not depend on context");
|
||||
assertEquals(withUnknown.idleSinceNanos(), withHigh.idleSinceNanos(), g.name());
|
||||
assertEquals(withUnknown.quietCount(), withHigh.quietCount(), g.name());
|
||||
assertEquals(withUnknown.action(), withOk.action(), g.name() + ": nor on an OK reading");
|
||||
assertFalse(withHigh.contextNotified(), g.name() + ": contextNotified must stay false when the flag is off");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void bHighContextFiresEvenPastAnExhaustedQuietCapWithoutSpendingIt() {
|
||||
LeadHeartbeatLoop on = loop(3, true);
|
||||
LeadHeartbeatLoop.Decision d = on.decide(
|
||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
|
||||
LeadContextGauge.State.HIGH, false);
|
||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
|
||||
"an idle, quiet, HIGH-context lead is exactly the case the exhausted cap must not swallow");
|
||||
assertEquals(3, d.quietCount(), "the context notice is an event notice, not a quiet nudge — it must not "
|
||||
+ "spend or grow the consecutive-quiet budget");
|
||||
assertTrue(d.contextNotified(), "firing the notice sets the latch");
|
||||
}
|
||||
|
||||
@Test
|
||||
void cASecondTickWithTheLatchAlreadySetDoesNotTellTheLeadAgain() {
|
||||
LeadHeartbeatLoop on = loop(3, true);
|
||||
LeadHeartbeatLoop.Decision d = on.decide(
|
||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
|
||||
LeadContextGauge.State.HIGH, true);
|
||||
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, d.action(),
|
||||
"the lead was already told once about this HIGH stretch — telling it every tick would be nagging, "
|
||||
+ "not a notice");
|
||||
}
|
||||
|
||||
@Test
|
||||
void dFlappingBetweenHighAndUnknownNeverReInjectsWhileLatched() {
|
||||
LeadHeartbeatLoop on = loop(3, true);
|
||||
// The lead was already told once (latch = true from a prior HIGH tick).
|
||||
LeadHeartbeatLoop.Decision afterUnknown = on.decide(
|
||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
|
||||
LeadContextGauge.State.UNKNOWN, true);
|
||||
assertTrue(afterUnknown.contextNotified(),
|
||||
"UNKNOWN means 'I could not look', not 'it got better' — it must not clear the latch");
|
||||
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, afterUnknown.action(),
|
||||
"a still-latched, still-quiet tick must not inject just because the reading is UNKNOWN");
|
||||
|
||||
LeadHeartbeatLoop.Decision afterHighAgain = on.decide(
|
||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
|
||||
LeadContextGauge.State.HIGH, afterUnknown.contextNotified());
|
||||
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, afterHighAgain.action(),
|
||||
"HIGH -> UNKNOWN -> HIGH with the latch already set must never inject again — this is the "
|
||||
+ "flapping case that would otherwise cost a full lead its remaining turns");
|
||||
}
|
||||
|
||||
@Test
|
||||
void eAnOkReadingClearsTheLatchSoALaterHighInjectsAgain() {
|
||||
LeadHeartbeatLoop on = loop(3, true);
|
||||
LeadHeartbeatLoop.Decision afterOk = on.decide(
|
||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
|
||||
LeadContextGauge.State.OK, true);
|
||||
assertFalse(afterOk.contextNotified(), "an OK reading clears the latch — the lead's context recovered");
|
||||
|
||||
LeadHeartbeatLoop.Decision afterHighAgain = on.decide(
|
||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
|
||||
LeadContextGauge.State.HIGH, afterOk.contextNotified());
|
||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, afterHighAgain.action(),
|
||||
"the cleared latch lets a genuinely new HIGH stretch notify again");
|
||||
}
|
||||
|
||||
@Test
|
||||
void fStandingDownDoesNotBurnTheOneContextNotice() {
|
||||
LeadHeartbeatLoop on = loop(3, true);
|
||||
LeadHeartbeatLoop.Decision d = on.decide(
|
||||
NOW, IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet(),
|
||||
LeadContextGauge.State.HIGH, false);
|
||||
assertEquals(LeadHeartbeatLoop.Action.STAND_DOWN, d.action(), "ReplyPushLoop is active — stand aside");
|
||||
assertFalse(d.contextNotified(),
|
||||
"standing down must not spend the one notice this HIGH stretch gets — the latch stays clear so a "
|
||||
+ "later tick can still fire it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void gAWorkingLeadWithHighContextIsStillLeadBusy() {
|
||||
LeadHeartbeatLoop on = loop(3, true);
|
||||
LeadHeartbeatLoop.Decision d = on.decide(
|
||||
NOW, IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet(),
|
||||
LeadContextGauge.State.HIGH, false);
|
||||
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(), "the status gate wins over the context notice");
|
||||
}
|
||||
|
||||
@Test
|
||||
void hAPendingDrivenInjectWhileHighSetsTheLatchTooSoTheLeadIsNotToldTwice() {
|
||||
LeadHeartbeatLoop on = loop(3, true);
|
||||
LeadHeartbeatLoop.Decision d = on.decide(
|
||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet(),
|
||||
LeadContextGauge.State.HIGH, false);
|
||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action());
|
||||
assertTrue(d.contextNotified(),
|
||||
"the pending-driven nudge carries the context notice too (see contextNotice), so it must set the "
|
||||
+ "latch — otherwise the lead could be told twice by two different routes");
|
||||
}
|
||||
|
||||
// ── contextNotice text builder ─────────────────────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
void contextNoticeIsEmptyWhenDisabled() {
|
||||
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
|
||||
assertEquals("", LeadHeartbeatLoop.contextNotice(false, reading));
|
||||
}
|
||||
|
||||
@Test
|
||||
void contextNoticeIsEmptyWhenStateIsOk() {
|
||||
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.OK, 1_000L, 0);
|
||||
assertEquals("", LeadHeartbeatLoop.contextNotice(true, reading));
|
||||
}
|
||||
|
||||
@Test
|
||||
void contextNoticeIsEmptyWhenStateIsUnknown() {
|
||||
assertEquals("", LeadHeartbeatLoop.contextNotice(true, LeadContextGauge.Reading.unknown()));
|
||||
}
|
||||
|
||||
@Test
|
||||
void contextNoticeNamesTheTokenCountWhenHighAndEnabled() {
|
||||
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
|
||||
String notice = LeadHeartbeatLoop.contextNotice(true, reading);
|
||||
assertFalse(notice.isEmpty());
|
||||
assertTrue(notice.contains("260771"), notice);
|
||||
assertTrue(notice.contains("1 compaction"), notice);
|
||||
assertTrue(notice.contains("fleet_handover"), notice);
|
||||
assertTrue(notice.contains("operator"), notice);
|
||||
}
|
||||
|
||||
@Test
|
||||
void contextNoticeOmitsTheTokenClauseRatherThanPrintingNull() {
|
||||
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, null, 2);
|
||||
String notice = LeadHeartbeatLoop.contextNotice(true, reading);
|
||||
assertFalse(notice.isEmpty());
|
||||
assertFalse(notice.toLowerCase().contains("null"), notice);
|
||||
assertTrue(notice.contains("2 compactions"), notice);
|
||||
}
|
||||
|
||||
// ── fleetd #621: the notice must track the effective requireOperatorConfirm value ─────────────
|
||||
|
||||
@Test
|
||||
void contextNoticeKeepsAskingTheOperatorWhenRequireOperatorConfirmIsTrue() {
|
||||
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
|
||||
String notice = LeadHeartbeatLoop.contextNotice(true, reading, false, true);
|
||||
assertTrue(notice.contains("ask the operator"), notice);
|
||||
assertTrue(notice.contains("Only the operator can approve the roll"), notice);
|
||||
}
|
||||
|
||||
@Test
|
||||
void contextNoticeDropsTheOperatorAskWhenRequireOperatorConfirmIsFalse() {
|
||||
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
|
||||
String notice = LeadHeartbeatLoop.contextNotice(true, reading, false, false);
|
||||
assertFalse(notice.contains("ask the operator"), notice);
|
||||
assertFalse(notice.contains("Only the operator can approve the roll"), notice);
|
||||
assertTrue(notice.contains("fleet_handover"), notice);
|
||||
assertTrue(notice.contains("maxDocAgeSeconds"), notice);
|
||||
}
|
||||
|
||||
// ── fleetd #609 review: the latch must mean "the notice reached the pane" ────────────────────
|
||||
//
|
||||
// These four drive LeadHeartbeatLoop.tick() directly (package-private, same reasoning as
|
||||
// ReplyPushLoop#tick(String) being directly testable) against a real AgentControl wrapping a
|
||||
// FailableHerdrClient, so the send path (agents.send -> herdr -> possible throw) is exercised
|
||||
// for real rather than assumed from decide()'s Decision alone.
|
||||
|
||||
private static final String LEAD = "term_lead";
|
||||
private static final String WORKER = "term_w1";
|
||||
|
||||
/** A HIGH reading with a fixed token/compaction count, for the four tests below. */
|
||||
private static LeadContextGauge.Reading highReading() {
|
||||
return new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds a real {@link LeadHeartbeatLoop} wired to {@code herdr} via a real {@link AgentControl},
|
||||
* a mutable fake clock, and a mutable roster so a test can change fleet state between ticks. The
|
||||
* lead is always reported IDLE by {@code herdr}, so every tick's outcome is governed only by the
|
||||
* idle-window/quiet-cap/context gates under test.
|
||||
*/
|
||||
private static LeadHeartbeatLoop tickableLoop(FailableHerdrClient herdr, AtomicLong now,
|
||||
List<MemberSession>[] rosterBox, InMemoryReplyInbox inbox,
|
||||
int quietNudgeCap, ScheduledExecutorService scheduler) {
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
PrimaryRegistry registry = new PrimaryRegistry(LEAD);
|
||||
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, agents, inbox, scheduler, 5, 100_000);
|
||||
return new LeadHeartbeatLoop(registry, agents, inbox, () -> rosterBox[0], pushLoop, scheduler,
|
||||
now::get, IDLE_AFTER_NANOS, 100_000L, quietNudgeCap, null,
|
||||
new LeadHeartbeatLoop.LeadContextSource(t -> highReading()), true);
|
||||
}
|
||||
|
||||
@Test
|
||||
void iAFailedSendDoesNotConsumeTheNotice() {
|
||||
var herdr = new FailableHerdrClient(LEAD);
|
||||
var now = new AtomicLong(NOW);
|
||||
@SuppressWarnings("unchecked")
|
||||
List<MemberSession>[] rosterBox = new List[]{List.of()}; // quiet: nothing pending
|
||||
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler);
|
||||
|
||||
loop.tick(); // first injectable tick: only opens the idle window (WAIT_IDLE)
|
||||
now.addAndGet(TimeUnit.SECONDS.toNanos(400)); // now clearly past the quiet period
|
||||
|
||||
herdr.throwOnNextSend();
|
||||
loop.tick(); // quiet fleet, quiet cap exhausted (0), context HIGH, latch clear -> INJECT, send throws
|
||||
|
||||
assertEquals(0, herdr.sentTexts().size(), "the failed send must not have recorded any text");
|
||||
|
||||
loop.tick(); // same inputs — the latch must still be clear, so this must INJECT and send again
|
||||
assertEquals(1, herdr.sentTexts().size(),
|
||||
"a retried tick with the latch still clear must attempt the send again");
|
||||
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"),
|
||||
"the retried, successful send must carry the notice: " + herdr.sentTexts().get(0));
|
||||
}
|
||||
|
||||
@Test
|
||||
void jASuccessfulSendDoesConsumeIt() {
|
||||
var herdr = new FailableHerdrClient(LEAD);
|
||||
var now = new AtomicLong(NOW);
|
||||
@SuppressWarnings("unchecked")
|
||||
List<MemberSession>[] rosterBox = new List[]{List.of()};
|
||||
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler);
|
||||
|
||||
loop.tick(); // opens the idle window
|
||||
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
|
||||
|
||||
loop.tick(); // INJECT, send succeeds -> latch set
|
||||
assertEquals(1, herdr.sentTexts().size());
|
||||
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"), herdr.sentTexts().get(0));
|
||||
|
||||
loop.tick(); // same inputs — the lead was already told this stretch
|
||||
assertEquals(1, herdr.sentTexts().size(),
|
||||
"the next tick with the same inputs must not send a second notice");
|
||||
}
|
||||
|
||||
@Test
|
||||
void kAPendingDrivenInjectWithTheLatchAlreadySetSendsNoNotice() {
|
||||
var herdr = new FailableHerdrClient(LEAD);
|
||||
var now = new AtomicLong(NOW);
|
||||
@SuppressWarnings("unchecked")
|
||||
List<MemberSession>[] rosterBox = new List[]{List.of()};
|
||||
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler);
|
||||
|
||||
loop.tick();
|
||||
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
|
||||
loop.tick(); // latches the notice (quiet, HIGH, cap exhausted -> the forced context INJECT)
|
||||
assertEquals(1, herdr.sentTexts().size());
|
||||
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"));
|
||||
|
||||
// Now make the fleet have real pending state, so the NEXT INJECT is pending-driven, not the
|
||||
// forced-context route — with the latch already set from the tick above.
|
||||
inbox.own(WORKER);
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
rosterBox[0] = List.of(new MemberSession("p1", WORKER, "prof", MemberRole.DEV, "/cwd", null,
|
||||
0, 0, 0, MemberSession.State.READY, null, null));
|
||||
|
||||
loop.tick();
|
||||
|
||||
assertEquals(2, herdr.sentTexts().size(), "the pending-driven tick must still send a nudge");
|
||||
assertFalse(herdr.sentTexts().get(1).contains("Your own context is nearly full"),
|
||||
"a pending-driven INJECT while the latch is already set must carry no context notice: "
|
||||
+ herdr.sentTexts().get(1));
|
||||
}
|
||||
|
||||
@Test
|
||||
void lTheNoticeAppearsExactlyOnceAcrossThreeDifferentlyDrivenInjects() {
|
||||
var herdr = new FailableHerdrClient(LEAD);
|
||||
var now = new AtomicLong(NOW);
|
||||
@SuppressWarnings("unchecked")
|
||||
List<MemberSession>[] rosterBox = new List[]{List.of()};
|
||||
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||
// quietNudgeCap=1 so a still-not-exhausted quiet nudge is available as the "forced" route below,
|
||||
// distinct from both the pending-driven route and the exhausted-cap forced-context route.
|
||||
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 1, scheduler);
|
||||
|
||||
loop.tick(); // opens the idle window
|
||||
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
|
||||
|
||||
// 1) pending-driven INJECT: real fleet state present. Sets the latch and carries the notice.
|
||||
inbox.own(WORKER);
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
rosterBox[0] = List.of(new MemberSession("p1", WORKER, "prof", MemberRole.DEV, "/cwd", null,
|
||||
0, 0, 0, MemberSession.State.READY, null, null));
|
||||
loop.tick();
|
||||
assertEquals(1, herdr.sentTexts().size());
|
||||
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"), herdr.sentTexts().get(0));
|
||||
|
||||
// 2) "forced" INJECT: nothing pending, but the quiet cap (1) is not yet exhausted, so decide()
|
||||
// nudges anyway. The latch is already set, so no notice.
|
||||
inbox.ack(WORKER, "m1");
|
||||
rosterBox[0] = List.of();
|
||||
loop.tick();
|
||||
assertEquals(2, herdr.sentTexts().size(), "the quiet-cap-not-yet-exhausted nudge must still fire");
|
||||
assertFalse(herdr.sentTexts().get(1).contains("Your own context is nearly full"), herdr.sentTexts().get(1));
|
||||
|
||||
// 3) pending-driven INJECT again. Still latched, still no notice.
|
||||
inbox.publish(WORKER, "m2", "hello again");
|
||||
rosterBox[0] = List.of(new MemberSession("p1", WORKER, "prof", MemberRole.DEV, "/cwd", null,
|
||||
0, 0, 0, MemberSession.State.READY, null, null));
|
||||
loop.tick();
|
||||
assertEquals(3, herdr.sentTexts().size());
|
||||
assertFalse(herdr.sentTexts().get(2).contains("Your own context is nearly full"), herdr.sentTexts().get(2));
|
||||
|
||||
long noticeCount = herdr.sentTexts().stream()
|
||||
.filter(t -> t.contains("Your own context is nearly full")).count();
|
||||
assertEquals(1, noticeCount,
|
||||
"the notice text must appear exactly once across all three sends: " + herdr.sentTexts());
|
||||
}
|
||||
|
||||
/**
|
||||
* Fake herdr client for the four tests above: always reports {@code lead} as IDLE, records the
|
||||
* {@code text} of every {@code agent.prompt} call, and can be told to throw on the very next
|
||||
* {@code agent.prompt} call — standing in for one transient herdr send failure.
|
||||
*/
|
||||
private static final class FailableHerdrClient implements HerdrClient {
|
||||
private static final ObjectMapper MAPPER = new ObjectMapper();
|
||||
private final String lead;
|
||||
private final List<String> sentTexts = new ArrayList<>();
|
||||
private boolean throwOnNextSend = false;
|
||||
|
||||
FailableHerdrClient(String lead) {
|
||||
this.lead = lead;
|
||||
}
|
||||
|
||||
void throwOnNextSend() {
|
||||
throwOnNextSend = true;
|
||||
}
|
||||
|
||||
List<String> sentTexts() {
|
||||
return List.copyOf(sentTexts);
|
||||
}
|
||||
|
||||
@Override
|
||||
@SuppressWarnings("unchecked")
|
||||
public JsonNode call(String method, Object params) {
|
||||
if ("agent.get".equals(method)) {
|
||||
return MAPPER.createObjectNode()
|
||||
.set("agent", MAPPER.createObjectNode()
|
||||
.put("terminal_id", lead)
|
||||
.put("agent_status", "idle"));
|
||||
}
|
||||
if ("agent.prompt".equals(method)) {
|
||||
if (throwOnNextSend) {
|
||||
throwOnNextSend = false;
|
||||
throw new RuntimeException("simulated transient herdr send failure");
|
||||
}
|
||||
Map<String, Object> p = params instanceof Map ? (Map<String, Object>) params : Map.of();
|
||||
sentTexts.add(String.valueOf(p.get("text")));
|
||||
}
|
||||
return MAPPER.createObjectNode();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -12,13 +12,18 @@ import org.testcontainers.junit.jupiter.Testcontainers;
|
||||
import org.testcontainers.utility.DockerImageName;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.lang.reflect.InvocationHandler;
|
||||
import java.lang.reflect.Method;
|
||||
import java.lang.reflect.Proxy;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertInstanceOf;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
@@ -207,6 +212,31 @@ class LeadMailboxTest {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* {@link LeadMailbox#inspect} opens a third channel after the mailbox's consume and publish
|
||||
* channels. Limit this connection to three channels, then require a replacement channel after
|
||||
* the successful inspect. If inspect leaves its probe open, the broker refuses that replacement.
|
||||
*/
|
||||
@Test
|
||||
void inspectClosesItsSuccessfulProbeChannel() throws Exception {
|
||||
var factory = LeadMailbox.connectionFactory(uri());
|
||||
factory.setRequestedChannelMax(3);
|
||||
Connection connection = factory.newConnection();
|
||||
try (LeadMailbox mailbox = new LeadMailbox(connection, coordId("lead-inspect-probe-close"))) {
|
||||
LeadChannel.MailboxState state = mailbox.inspect(mailbox.selfCoordId());
|
||||
assertTrue(state.exists(), "the owned mailbox must be found before checking the probe channel");
|
||||
|
||||
Channel replacement = connection.createChannel();
|
||||
assertNotNull(replacement,
|
||||
"inspect must close its successful probe channel; the replacement channel was null");
|
||||
try {
|
||||
assertTrue(replacement.isOpen(), "the replacement channel must be open after inspect returns");
|
||||
} finally {
|
||||
replacement.close();
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #440: {@code heldDurable()} must be derived from what {@link LeadMailbox#own} actually
|
||||
* did against the real broker — a durable queue declare plus a manual-ack consumer — not a
|
||||
@@ -347,6 +377,72 @@ class LeadMailboxTest {
|
||||
() -> "expected AlreadyClosedException, got: " + thrown);
|
||||
}
|
||||
|
||||
@Test
|
||||
void publishIOExceptionRemovesThePendingMessageId() throws Exception {
|
||||
LeadMailbox mailbox = newMailbox(true);
|
||||
try {
|
||||
assertThrows(IllegalStateException.class,
|
||||
() -> mailbox.publish("target", new LeadMessage("catch", "from", "target", "body")));
|
||||
assertEquals(0, pendingByMsgId(mailbox).size(), "publish IOException must remove its msgId entry");
|
||||
} finally {
|
||||
mailbox.close();
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void interruptedPublishRemovesThePendingMessageIdInFinally() throws Exception {
|
||||
LeadMailbox mailbox = newMailbox(false);
|
||||
Thread publish = Thread.ofVirtual().start(() -> {
|
||||
try {
|
||||
mailbox.publish("target", new LeadMessage("finally", "from", "target", "body"));
|
||||
} catch (IllegalStateException ignored) {
|
||||
// Interrupting the confirm wait is the path under test.
|
||||
}
|
||||
});
|
||||
awaitPending(mailbox, "finally");
|
||||
publish.interrupt();
|
||||
publish.join(5_000);
|
||||
|
||||
try {
|
||||
assertEquals(0, pendingByMsgId(mailbox).size(), "publish finally must remove its msgId entry");
|
||||
} finally {
|
||||
mailbox.close();
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void confirmResolutionRemovesThePendingMessageId() throws Exception {
|
||||
LeadMailbox mailbox = newMailbox(false);
|
||||
try {
|
||||
seedPending(mailbox, 1, "confirm");
|
||||
invoke(mailbox, "resolveConfirm", new Class<?>[] {long.class, boolean.class, boolean.class}, 1L, false, true);
|
||||
assertEquals(0, pendingByMsgId(mailbox).size(), "confirm resolution must remove its msgId entry");
|
||||
} finally {
|
||||
mailbox.close();
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void recoverySweepRemovesThePendingMessageId() throws Exception {
|
||||
LeadMailbox mailbox = newMailbox(false);
|
||||
try {
|
||||
seedPending(mailbox, 1, "recovery");
|
||||
mailbox.failPendingPublishesOnRecovery();
|
||||
assertEquals(0, pendingByMsgId(mailbox).size(), "recovery sweep must remove its msgId entry");
|
||||
} finally {
|
||||
mailbox.close();
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void closeRemovesThePendingMessageId() throws Exception {
|
||||
LeadMailbox mailbox = newMailbox(false);
|
||||
seedPending(mailbox, 1, "close");
|
||||
mailbox.close();
|
||||
|
||||
assertEquals(0, pendingByMsgId(mailbox).size(), "close must remove its msgId entry");
|
||||
}
|
||||
|
||||
/** Poll peek until at least one message is held, or ~10s elapse (broker delivery is async). */
|
||||
@SuppressWarnings("BusyWait")
|
||||
private static List<LeadMessage> awaitPeek(LeadMailbox inbox) throws InterruptedException {
|
||||
@@ -372,4 +468,108 @@ class LeadMailboxTest {
|
||||
}
|
||||
return state;
|
||||
}
|
||||
|
||||
private static LeadMailbox newMailbox(boolean failPublish) {
|
||||
AtomicLong sequence = new AtomicLong();
|
||||
Channel consume = fakeChannel(sequence, false);
|
||||
Channel publish = fakeChannel(sequence, failPublish);
|
||||
return new LeadMailbox(fakeConnection(consume, publish), "self");
|
||||
}
|
||||
|
||||
private static Channel fakeChannel(AtomicLong sequence, boolean failPublish) {
|
||||
InvocationHandler handler = (proxy, method, args) -> {
|
||||
if (method.getName().equals("getNextPublishSeqNo")) {
|
||||
return sequence.incrementAndGet();
|
||||
}
|
||||
if (method.getName().equals("basicPublish") && failPublish) {
|
||||
throw new IOException("test publish failure");
|
||||
}
|
||||
if (method.getName().equals("equals")) {
|
||||
return proxy == args[0];
|
||||
}
|
||||
if (method.getName().equals("hashCode")) {
|
||||
return System.identityHashCode(proxy);
|
||||
}
|
||||
return defaultValue(method.getReturnType());
|
||||
};
|
||||
return (Channel) Proxy.newProxyInstance(LeadMailboxTest.class.getClassLoader(), new Class<?>[] {Channel.class}, handler);
|
||||
}
|
||||
|
||||
private static Connection fakeConnection(Channel first, Channel second) {
|
||||
AtomicLong calls = new AtomicLong();
|
||||
InvocationHandler handler = (proxy, method, args) -> {
|
||||
if (method.getName().equals("createChannel") && (args == null || args.length == 0)) {
|
||||
return calls.getAndIncrement() == 0 ? first : second;
|
||||
}
|
||||
if (method.getName().equals("equals")) {
|
||||
return proxy == args[0];
|
||||
}
|
||||
if (method.getName().equals("hashCode")) {
|
||||
return System.identityHashCode(proxy);
|
||||
}
|
||||
return defaultValue(method.getReturnType());
|
||||
};
|
||||
return (Connection) Proxy.newProxyInstance(LeadMailboxTest.class.getClassLoader(), new Class<?>[] {Connection.class}, handler);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static Map<String, Object> pendingByMsgId(LeadMailbox mailbox) throws Exception {
|
||||
var field = LeadMailbox.class.getDeclaredField("pendingByMsgId");
|
||||
field.setAccessible(true);
|
||||
return (Map<String, Object>) field.get(mailbox);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static void seedPending(LeadMailbox mailbox, long seq, String msgId) throws Exception {
|
||||
Class<?> pendingType = Class.forName(LeadMailbox.class.getName() + "$Pending");
|
||||
var constructor = pendingType.getDeclaredConstructor(String.class);
|
||||
constructor.setAccessible(true);
|
||||
Object pending = constructor.newInstance(msgId);
|
||||
var seqField = LeadMailbox.class.getDeclaredField("pendingBySeq");
|
||||
seqField.setAccessible(true);
|
||||
((Map<Long, Object>) seqField.get(mailbox)).put(seq, pending);
|
||||
pendingByMsgId(mailbox).put(msgId, pending);
|
||||
}
|
||||
|
||||
private static void invoke(LeadMailbox mailbox, String name, Class<?>[] types, Object... args) throws Exception {
|
||||
Method method = LeadMailbox.class.getDeclaredMethod(name, types);
|
||||
method.setAccessible(true);
|
||||
method.invoke(mailbox, args);
|
||||
}
|
||||
|
||||
private static void awaitPending(LeadMailbox mailbox, String msgId) throws Exception {
|
||||
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(5);
|
||||
while (!pendingByMsgId(mailbox).containsKey(msgId) && System.nanoTime() < deadline) {
|
||||
Thread.sleep(10);
|
||||
}
|
||||
assertTrue(pendingByMsgId(mailbox).containsKey(msgId), "publish did not register " + msgId);
|
||||
}
|
||||
|
||||
private static Object defaultValue(Class<?> type) {
|
||||
if (!type.isPrimitive() || type == void.class) {
|
||||
return null;
|
||||
}
|
||||
if (type == boolean.class) {
|
||||
return Boolean.FALSE;
|
||||
}
|
||||
if (type == long.class) {
|
||||
return 0L;
|
||||
}
|
||||
if (type == short.class) {
|
||||
return (short) 0;
|
||||
}
|
||||
if (type == byte.class) {
|
||||
return (byte) 0;
|
||||
}
|
||||
if (type == char.class) {
|
||||
return (char) 0;
|
||||
}
|
||||
if (type == double.class) {
|
||||
return 0.0d;
|
||||
}
|
||||
if (type == float.class) {
|
||||
return 0.0f;
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,168 @@
|
||||
package dev.ltms.fleet.msg;
|
||||
|
||||
import java.util.ArrayDeque;
|
||||
import java.util.ArrayList;
|
||||
import java.util.Collection;
|
||||
import java.util.Deque;
|
||||
import java.util.List;
|
||||
import java.util.concurrent.Callable;
|
||||
import java.util.concurrent.ExecutionException;
|
||||
import java.util.concurrent.Future;
|
||||
import java.util.concurrent.RejectedExecutionException;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.ScheduledFuture;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
/**
|
||||
* A {@link ScheduledExecutorService} for tests that never runs a task on a timer — it records what
|
||||
* {@link ReplyPushLoop} schedules and only ever runs it when the test itself calls
|
||||
* {@link #runDueTasks()} (fleetd #608).
|
||||
*
|
||||
* <p>The test this exists for ({@code anAlreadyCollectedTicketProducesNoNudge}) used to wire
|
||||
* {@link ReplyPushLoop} to a real {@code Executors.newSingleThreadScheduledExecutor()} and bet a
|
||||
* 300ms backoff was "wide" enough that the test's own work (collecting the ticket) always won the
|
||||
* race against the scheduler's own timer firing the next tick. On an idle machine that held; under
|
||||
* a loaded full-suite run it did not, and the test went red on perfectly correct code. Replacing
|
||||
* the timer with this fake removes the race rather than widening it: nothing here ever runs on a
|
||||
* schedule of its own, so no backoff value — however small — can make the tick fire before the test
|
||||
* is ready for it.
|
||||
*
|
||||
* <p><strong>Deliberately narrow.</strong> {@link ReplyPushLoop} calls exactly two methods on its
|
||||
* {@link ScheduledExecutorService} field — {@link #schedule(Runnable, long, TimeUnit)} (every tick,
|
||||
* including the very first one {@code startOrCoalesce} kicks off) and {@link #shutdownNow()} (on
|
||||
* {@code ReplyPushLoop.stop()}/{@code close()}) — verified by reading every {@code scheduler.} call
|
||||
* site in that class. Every other {@link ScheduledExecutorService} method throws
|
||||
* {@link UnsupportedOperationException} rather than silently doing the wrong thing, so a future
|
||||
* change to {@code ReplyPushLoop} that starts calling one of them fails this fake loudly, on the
|
||||
* very first test that exercises it, instead of being quietly mishandled.
|
||||
*/
|
||||
final class ManualScheduler implements ScheduledExecutorService {
|
||||
|
||||
private final Deque<Runnable> pending = new ArrayDeque<>();
|
||||
private boolean shutdown = false;
|
||||
|
||||
/**
|
||||
* Run every task pending as of the START of this call — a snapshot taken before anything runs.
|
||||
* {@link ReplyPushLoop#tick} always reschedules its own next tick before returning (see
|
||||
* {@code scheduleNext} in its {@code INJECT}/{@code WAIT_BUSY} branches), so without the
|
||||
* snapshot a single call here would recurse forever. Taking it up front means one call is
|
||||
* exactly one tick, deterministically, no matter what that tick itself goes on to schedule.
|
||||
*
|
||||
* @return how many tasks actually ran
|
||||
*/
|
||||
synchronized int runDueTasks() {
|
||||
List<Runnable> due = new ArrayList<>(pending);
|
||||
pending.clear();
|
||||
for (Runnable task : due) {
|
||||
task.run();
|
||||
}
|
||||
return due.size();
|
||||
}
|
||||
|
||||
/** How many tasks are currently queued, without running any of them. */
|
||||
synchronized int pendingCount() {
|
||||
return pending.size();
|
||||
}
|
||||
|
||||
@Override
|
||||
public synchronized ScheduledFuture<?> schedule(Runnable command, long delay, TimeUnit unit) {
|
||||
if (shutdown) {
|
||||
throw new RejectedExecutionException("ManualScheduler is shut down");
|
||||
}
|
||||
pending.add(command);
|
||||
return null; // ReplyPushLoop discards the return value of every schedule() call it makes.
|
||||
}
|
||||
|
||||
@Override
|
||||
public synchronized List<Runnable> shutdownNow() {
|
||||
shutdown = true;
|
||||
List<Runnable> left = new ArrayList<>(pending);
|
||||
pending.clear();
|
||||
return left;
|
||||
}
|
||||
|
||||
@Override
|
||||
public synchronized boolean isShutdown() {
|
||||
return shutdown;
|
||||
}
|
||||
|
||||
// --- everything below: not called by ReplyPushLoop, and not supported by this fake (see the
|
||||
// class javadoc's "deliberately narrow" note) -------------------------------------------------
|
||||
|
||||
@Override
|
||||
public ScheduledFuture<?> scheduleAtFixedRate(Runnable command, long initialDelay, long period, TimeUnit unit) {
|
||||
throw unsupported();
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledFuture<?> scheduleWithFixedDelay(Runnable command, long initialDelay, long delay, TimeUnit unit) {
|
||||
throw unsupported();
|
||||
}
|
||||
|
||||
@Override
|
||||
public <V> ScheduledFuture<V> schedule(Callable<V> callable, long delay, TimeUnit unit) {
|
||||
throw unsupported();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void shutdown() {
|
||||
throw unsupported();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean isTerminated() {
|
||||
throw unsupported();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean awaitTermination(long timeout, TimeUnit unit) {
|
||||
throw unsupported();
|
||||
}
|
||||
|
||||
@Override
|
||||
public <T> Future<T> submit(Callable<T> task) {
|
||||
throw unsupported();
|
||||
}
|
||||
|
||||
@Override
|
||||
public <T> Future<T> submit(Runnable task, T result) {
|
||||
throw unsupported();
|
||||
}
|
||||
|
||||
@Override
|
||||
public Future<?> submit(Runnable task) {
|
||||
throw unsupported();
|
||||
}
|
||||
|
||||
@Override
|
||||
public <T> List<Future<T>> invokeAll(Collection<? extends Callable<T>> tasks) {
|
||||
throw unsupported();
|
||||
}
|
||||
|
||||
@Override
|
||||
public <T> List<Future<T>> invokeAll(Collection<? extends Callable<T>> tasks, long timeout, TimeUnit unit) {
|
||||
throw unsupported();
|
||||
}
|
||||
|
||||
@Override
|
||||
public <T> T invokeAny(Collection<? extends Callable<T>> tasks) throws ExecutionException {
|
||||
throw unsupported();
|
||||
}
|
||||
|
||||
@Override
|
||||
public <T> T invokeAny(Collection<? extends Callable<T>> tasks, long timeout, TimeUnit unit)
|
||||
throws ExecutionException {
|
||||
throw unsupported();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void execute(Runnable command) {
|
||||
throw unsupported();
|
||||
}
|
||||
|
||||
private static UnsupportedOperationException unsupported() {
|
||||
return new UnsupportedOperationException(
|
||||
"ManualScheduler only supports schedule(Runnable, long, TimeUnit), shutdownNow() and isShutdown() "
|
||||
+ "— see the class javadoc");
|
||||
}
|
||||
}
|
||||
@@ -22,6 +22,7 @@ import java.util.concurrent.TimeUnit;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
@@ -505,6 +506,52 @@ class MessageServiceTest {
|
||||
"an answer to a turn that never existed (or already lapsed) is stale, not a hang");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #575. {@code answer()}'s STALE_TURN return for "the ask lapsed between the lookup and
|
||||
* the unblock" ({@code rendezvous.answerAsk(turnId, content)} returning {@code false} even though
|
||||
* this call's own {@code rendezvous.askSession(turnId)} check at the top saw the ask as open) used
|
||||
* to be covered only by a hand-rolled copy of the cleanup pair its own finally already runs, not
|
||||
* the finally itself — a structural gap, closed by widening the try up to cover the registration
|
||||
* above it. This test pins the exact race with {@link
|
||||
* MessageService#setAnswerAskLapseRaceHookForTest}, which fires right before this call's own
|
||||
* {@code rendezvous.answerAsk} call and completes the same turnId's ask directly — reproducing
|
||||
* what a second, concurrent {@code answer()} winning that race could otherwise only do by timing
|
||||
* luck. Proves the fix changed no behaviour on this path: it still returns {@code STALE_TURN},
|
||||
* and this call's own forward waiter is still closed exactly once (not left open, and not closed
|
||||
* twice — there is now only one cleanup site left to run).
|
||||
*/
|
||||
@Test
|
||||
void answerLosingTheRaceToAnAlreadyAnsweredAskStillReturnsStaleTurnAndCleansUpOnce() throws Exception {
|
||||
String ticket = messages.sendAsync(T, "long task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
|
||||
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
|
||||
String turnId = asking.turnId();
|
||||
assertNotNull(turnId, "an ASKING view carries the turnId to answer on");
|
||||
|
||||
messages.setAnswerAskLapseRaceHookForTest(() -> rendezvous.answerAsk(turnId, "raced in first"));
|
||||
try {
|
||||
MessageService.Reply r = messages.answer(turnId, "too late", 500);
|
||||
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
|
||||
"an ask already answered by the race must be seen as lapsed, not double-delivered");
|
||||
} finally {
|
||||
messages.setAnswerAskLapseRaceHookForTest(null);
|
||||
}
|
||||
|
||||
// Cleanup ran exactly once: the forward waiter THIS call opened is closed, not leaked.
|
||||
assertNull(rendezvous.currentWaiter(T),
|
||||
"the forward waiter this answer() call opened must be closed after a STALE_TURN return");
|
||||
|
||||
// The worker's own ask() call, unblocked by the hook's direct answerAsk, still completes
|
||||
// normally — the race this test simulates does not strand it.
|
||||
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome());
|
||||
assertEquals("raced in first", a.answer());
|
||||
}
|
||||
|
||||
// --- timeout, answer, poll, and lock-contention edges ----------------------------------
|
||||
|
||||
@Test
|
||||
@@ -554,6 +601,30 @@ class MessageServiceTest {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #571 (the acceptance test the ticket was filed for). The worker is idle so the injector
|
||||
* attempts delivery, but the {@code agent.prompt} call itself fails with a herdr error that is
|
||||
* not a confirmed absence (not a {@code *_not_found} code) — {@link Injector} marks the Pending
|
||||
* {@code ATTEMPTED} (fleetd #551), meaning the call was made and whether it reached the pane is
|
||||
* unknown. Before this fix, {@code send}'s {@code TimeoutException} branch collapsed
|
||||
* {@code ATTEMPTED} into {@code TIMED_OUT_QUEUED} — a promise that the message will never arrive,
|
||||
* which may already be false: {@code agent.prompt} pastes and submits in one call.
|
||||
*/
|
||||
@Test
|
||||
void sendTimesOutWithAttemptedDeliveryReportsUnconfirmedNotQueued() throws Exception {
|
||||
herdr.agentSendFailsWith("send_failed");
|
||||
CompletableFuture<MessageService.Reply> send =
|
||||
CompletableFuture.supplyAsync(() -> messages.send(T, "brief", 150));
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE); // triggers the failing delivery attempt → ATTEMPTED
|
||||
|
||||
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_UNCONFIRMED, r.outcome(),
|
||||
"an ATTEMPTED delivery must not collapse into TIMED_OUT_QUEUED — the message may "
|
||||
+ "already have arrived in full, and TIMED_OUT_QUEUED promises it never will");
|
||||
assertNull(r.text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void answerTimesOutWhenTheResumedWorkerNeverReplies() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
@@ -579,6 +650,183 @@ class MessageServiceTest {
|
||||
assertEquals("config.yaml", a.answer());
|
||||
}
|
||||
|
||||
// --- fleetd #572: answer() must release the session lock on EVERY exit, not just return it -----
|
||||
//
|
||||
// answer()'s session lock is released in an outer `finally` (MessageService.java:1217-1219) that
|
||||
// wraps its whole body. Removing that one line survives the entire suite: coverage is not the
|
||||
// gap, assertion is — every existing answer() test checks the RETURN VALUE, never that the lock
|
||||
// it took is reacquirable afterward. If it is not, the session is wedged forever with no
|
||||
// exception and no log line. Each test below drives answer() through one of its four exits
|
||||
// (normal reply, TIMED_OUT_WORKING, ExecutionException, InterruptedException) and then proves
|
||||
// reacquisition the only way that actually proves it: a bounded follow-up `send` on the SAME
|
||||
// session must not come back BUSY. A `send` only reports BUSY when `tryLock` itself timed out
|
||||
// (MessageService.java:926-928) — every other early return in `send` still passes through its own
|
||||
// lock-acquired `try`, so a non-BUSY probe result is specifically evidence the lock was free.
|
||||
|
||||
/** The REPLIED exit (the happy path) — the resumed worker's real {@code fleet_reply} arrives. */
|
||||
@Test
|
||||
void answerReleasesTheSessionLockAfterANormalReply() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
injector.onStatus(T, AgentStatus.WORKING);
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
|
||||
|
||||
CompletableFuture<MessageService.Reply> answer =
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
|
||||
ask.get(5, TimeUnit.SECONDS); // worker resumed with the answer
|
||||
|
||||
awaitWaiting(); // the answering call has (re)opened its own forward waiter
|
||||
assertTrue(rendezvous.resolve(T, "done"), "the worker's final reply resolves the answering send");
|
||||
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.REPLIED, done.outcome());
|
||||
|
||||
MessageService.Reply probe = messages.send(T, "probe after normal reply", 300);
|
||||
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
|
||||
"the session lock must be released after a normal REPLIED answer(), or this bounded "
|
||||
+ "follow-up send would come back BUSY instead of timing out on its own work");
|
||||
}
|
||||
|
||||
/**
|
||||
* The {@code TIMED_OUT_WORKING} exit. Same setup as {@link #answerTimesOutWhenTheResumedWorkerNeverReplies}
|
||||
* (which only checks {@code answer()}'s return value — exactly the assertion the fleetd #572
|
||||
* mutation survives), plus the reacquisition proof that test does not make.
|
||||
*
|
||||
* <p>{@code answer()} MUST run on its own thread here, not the test's main thread: {@code lock}
|
||||
* is a {@link java.util.concurrent.locks.ReentrantLock}, so a probe {@code send} issued by the
|
||||
* SAME thread that (under the mutation) still "holds" it would reenter for free and report a
|
||||
* false pass — reentrancy, not release. Measured while writing this test: with {@code answer()}
|
||||
* called inline, this test stayed green under the mutation while its three siblings correctly
|
||||
* went red.
|
||||
*/
|
||||
@Test
|
||||
void answerReleasesTheSessionLockAfterATimedOutWorkingReturn() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
injector.onStatus(T, AgentStatus.WORKING);
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
|
||||
|
||||
// The primary answers, unblocking the worker; the worker never sends its follow-up
|
||||
// fleet_reply, so the answering call rides out its short window as still-working.
|
||||
CompletableFuture<MessageService.Reply> answer =
|
||||
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 200));
|
||||
MessageService.Reply answered = answer.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answered.outcome(),
|
||||
"an answered worker that never replies times out as still working");
|
||||
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
|
||||
|
||||
MessageService.Reply probe = messages.send(T, "probe after timed-out-working", 300);
|
||||
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
|
||||
"the session lock must be released after a TIMED_OUT_WORKING answer(), or this bounded "
|
||||
+ "follow-up send would come back BUSY instead of timing out on its own work");
|
||||
}
|
||||
|
||||
/**
|
||||
* The {@code ExecutionException} exit. Forces it directly on answer()'s own reopened forward
|
||||
* waiter — {@code rendezvous.currentWaiter(T)} is the exact {@code CompletableFuture} its
|
||||
* {@code reply.get(...)} is blocked on — rather than trying to make a real worker fail, since the
|
||||
* failure mode under test is in answer()'s own wait, not in how it got triggered.
|
||||
*/
|
||||
@Test
|
||||
void answerReleasesTheSessionLockWhenTheReplyFutureFailsExceptionally() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
injector.onStatus(T, AgentStatus.WORKING);
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
|
||||
String turnId = q.turnId();
|
||||
|
||||
java.util.concurrent.atomic.AtomicReference<Throwable> caught =
|
||||
new java.util.concurrent.atomic.AtomicReference<>();
|
||||
Thread answerer = new Thread(() -> {
|
||||
try {
|
||||
messages.answer(turnId, "config.yaml", 5000);
|
||||
caught.set(new AssertionError("expected answer() to throw"));
|
||||
} catch (Throwable t) {
|
||||
caught.set(t);
|
||||
}
|
||||
});
|
||||
answerer.start();
|
||||
awaitWaiting(); // answer() re-opened its forward waiter and is about to block on it
|
||||
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(T);
|
||||
assertNotNull(waiter, "answer() should have registered a forward waiter for the worker session");
|
||||
waiter.completeExceptionally(new RuntimeException("boom"));
|
||||
|
||||
answerer.join(5000);
|
||||
assertFalse(answerer.isAlive(), "answer() should have left by throwing once its reply future failed");
|
||||
Throwable thrown = caught.get();
|
||||
assertNotNull(thrown, "answer() should have thrown");
|
||||
assertTrue(thrown instanceof RuntimeException, "a RuntimeException cause is rethrown as-is: " + thrown);
|
||||
assertEquals("boom", thrown.getMessage());
|
||||
|
||||
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
|
||||
|
||||
MessageService.Reply probe = messages.send(T, "probe after exceptional failure", 300);
|
||||
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
|
||||
"the session lock must be released when answer()'s reply future fails exceptionally, "
|
||||
+ "or this bounded follow-up send would come back BUSY instead of timing out on its own work");
|
||||
}
|
||||
|
||||
/**
|
||||
* The {@code InterruptedException} exit. Same shape as the {@code ExecutionException} case above,
|
||||
* but the thread blocked in {@code reply.get(...)} is interrupted instead of the future failing.
|
||||
*/
|
||||
@Test
|
||||
void answerReleasesTheSessionLockWhenInterrupted() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
injector.onStatus(T, AgentStatus.WORKING);
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
|
||||
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
|
||||
String turnId = q.turnId();
|
||||
|
||||
java.util.concurrent.atomic.AtomicReference<Throwable> caught =
|
||||
new java.util.concurrent.atomic.AtomicReference<>();
|
||||
Thread answerer = new Thread(() -> {
|
||||
try {
|
||||
messages.answer(turnId, "config.yaml", 5000);
|
||||
caught.set(new AssertionError("expected answer() to throw"));
|
||||
} catch (Throwable t) {
|
||||
caught.set(t);
|
||||
}
|
||||
});
|
||||
answerer.start();
|
||||
awaitWaiting(); // answer() re-opened its forward waiter and is about to block on it
|
||||
|
||||
answerer.interrupt();
|
||||
answerer.join(5000);
|
||||
assertFalse(answerer.isAlive(), "answer() should have left by throwing once its thread was interrupted");
|
||||
Throwable thrown = caught.get();
|
||||
assertNotNull(thrown, "answer() should have thrown");
|
||||
assertTrue(thrown instanceof IllegalStateException,
|
||||
"an interrupted wait is wrapped in IllegalStateException: " + thrown);
|
||||
|
||||
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
|
||||
|
||||
MessageService.Reply probe = messages.send(T, "probe after interruption", 300);
|
||||
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
|
||||
"the session lock must be released when answer()'s wait is interrupted, or this bounded "
|
||||
+ "follow-up send would come back BUSY instead of timing out on its own work");
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollReturnsNullForAnUnknownTicket() {
|
||||
assertNull(messages.poll("task-999999"), "a ticket that was never minted is unknown");
|
||||
@@ -1593,6 +1841,43 @@ class MessageServiceTest {
|
||||
return new PushWiring(service, leadHerdr, scheduler);
|
||||
}
|
||||
|
||||
/**
|
||||
* A {@link MessageService} wired to a real {@link ReplyPushLoop}, like {@link #wireWithPushLoop},
|
||||
* but backed by {@link ManualScheduler} instead of a real timer (fleetd #608): its tick never
|
||||
* fires on its own — a test drives it explicitly via {@link ManualScheduler#runDueTasks()}.
|
||||
*/
|
||||
private record ManualPushWiring(MessageService service, FakeHerdr leadHerdr, ManualScheduler scheduler,
|
||||
ReplyPushLoop pushLoop)
|
||||
implements AutoCloseable {
|
||||
@Override
|
||||
public void close() {
|
||||
scheduler.shutdownNow();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* As {@link #wireWithPushLoop}, but the schedule never runs on its own. {@code backoffMs} is
|
||||
* still threaded through to {@link ReplyPushLoop}'s constructor (it takes one), but nothing here
|
||||
* ever waits it out, so its value cannot affect anything a test built on this observes — see
|
||||
* {@code anAlreadyCollectedTicketProducesNoNudge}, which sets it to 1 to prove exactly that.
|
||||
*/
|
||||
private ManualPushWiring wireWithManualScheduler(int maxReminders, long backoffMs) {
|
||||
return wireWithManualScheduler(maxReminders, backoffMs, System::nanoTime);
|
||||
}
|
||||
|
||||
/** As above, with an injectable clock for tests that exercise terminal-ticket pruning. */
|
||||
private ManualPushWiring wireWithManualScheduler(int maxReminders, long backoffMs,
|
||||
java.util.function.LongSupplier nowNanos) {
|
||||
PrimaryRegistry registry = new PrimaryRegistry(null);
|
||||
registry.recordDelegation(T, LEAD);
|
||||
FakeHerdr leadHerdr = new FakeHerdr();
|
||||
AgentControl leadAgents = new AgentControl(leadHerdr);
|
||||
ManualScheduler scheduler = new ManualScheduler();
|
||||
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, leadAgents, inbox, scheduler, maxReminders, backoffMs);
|
||||
MessageService service = new MessageService(agents, injector, rendezvous, inbox, pushLoop, null, nowNanos);
|
||||
return new ManualPushWiring(service, leadHerdr, scheduler, pushLoop);
|
||||
}
|
||||
|
||||
private void awaitNudge(FakeHerdr leadHerdr) throws InterruptedException {
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!leadHerdr.called("agent.prompt") && System.currentTimeMillis() < deadline) {
|
||||
@@ -1634,16 +1919,46 @@ class MessageServiceTest {
|
||||
|
||||
@Test
|
||||
void anAlreadyCollectedTicketProducesNoNudge() throws Exception {
|
||||
try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: poll before the first tick fires
|
||||
String ticket = wiring.service().sendAsync(T, "long task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "async result"));
|
||||
// fleetd #608: backoff is 1ms — the most hostile value there is, the tick due immediately —
|
||||
// and the test still must pass, because with ManualScheduler the tick never runs on a timer
|
||||
// at all; it only runs when this test calls runDueTasks() below. The old version bet a 300ms
|
||||
// backoff was "wide" enough that collecting the ticket always won the race against a real
|
||||
// scheduler's own timer; that held on an idle machine and failed under a loaded full-suite
|
||||
// run — a false red on correct code, since nothing here forced that ordering, it only made
|
||||
// it likely.
|
||||
try (var wiring = wireWithManualScheduler(1, 1)) {
|
||||
java.util.concurrent.CountDownLatch terminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||
// finishAsyncTask's task.future.complete(...) runs every whenComplete registered on that
|
||||
// same future — including sendAsync's own hook that calls ReplyPushLoop.onTicketTerminal
|
||||
// — synchronously, before complete() returns (see finishAsyncTask's javadoc and fleetd
|
||||
// #399). This test-only hook fires right after that complete() call, on the very same
|
||||
// (async executor) thread, so waiting for it guarantees onTicketTerminal has already run
|
||||
// and the ticket is really sitting in the push loop's pending set — unlike waiting for
|
||||
// Phase.DONE via poll(), which CompletableFuture.complete() can make visible to another
|
||||
// thread before every whenComplete dependent has actually finished running (the same
|
||||
// publish-then-run-dependents gap awaitCompletionStamped exists to close elsewhere).
|
||||
String ticket;
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(terminalReached::countDown);
|
||||
try {
|
||||
ticket = wiring.service().sendAsync(T, "long task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "async result"));
|
||||
assertTrue(terminalReached.await(5, TimeUnit.SECONDS),
|
||||
"the ticket never reached its terminal phase");
|
||||
} finally {
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
|
||||
}
|
||||
|
||||
// Collect it — the exact action the nudge must never follow.
|
||||
MessageService.TaskView view = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.DONE);
|
||||
assertEquals(MessageService.Phase.DONE, view.phase());
|
||||
|
||||
Thread.sleep(400); // let the scheduled tick run — it must find nothing pending
|
||||
// Now run the pending tick explicitly. onTicketTerminal scheduled exactly one (the ticket
|
||||
// is the only thing this lead has ever had pending); it must find nothing left pending —
|
||||
// the collection above already removed it — and send no nudge.
|
||||
assertEquals(1, wiring.scheduler().runDueTasks(),
|
||||
"expected exactly the one tick onTicketTerminal scheduled");
|
||||
assertFalse(wiring.leadHerdr().called("agent.prompt"),
|
||||
"a ticket the lead already polled must never be nudged");
|
||||
}
|
||||
@@ -1651,24 +1966,35 @@ class MessageServiceTest {
|
||||
|
||||
@Test
|
||||
void severalAsyncTicketsFinishingTogetherProduceOneCoalescedNudge() throws Exception {
|
||||
try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: both tickets land before the tick fires
|
||||
String first = wiring.service().sendAsync(T, "first task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "first done"));
|
||||
// Settle without polling: poll() itself marks a ticket collected (that's the point of
|
||||
// anAlreadyCollectedTicketProducesNoNudge above) — using it here to detect completion
|
||||
// would collect the ticket before the coalescing this test checks ever gets a chance.
|
||||
Thread.sleep(100);
|
||||
// A 1ms backoff is due immediately. ManualScheduler still cannot run it until this test
|
||||
// explicitly calls runDueTasks(), so both terminal tickets join one scheduled tick.
|
||||
try (var wiring = wireWithManualScheduler(1, 1)) {
|
||||
String first;
|
||||
String second;
|
||||
java.util.concurrent.CountDownLatch firstTerminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(firstTerminalReached::countDown);
|
||||
try {
|
||||
first = wiring.service().sendAsync(T, "first task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "first done"));
|
||||
assertTrue(firstTerminalReached.await(5, TimeUnit.SECONDS),
|
||||
"the first ticket never reached its terminal phase");
|
||||
|
||||
String second = wiring.service().sendAsync(T, "second task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "second done"));
|
||||
Thread.sleep(100);
|
||||
java.util.concurrent.CountDownLatch secondTerminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(secondTerminalReached::countDown);
|
||||
second = wiring.service().sendAsync(T, "second task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "second done"));
|
||||
assertTrue(secondTerminalReached.await(5, TimeUnit.SECONDS),
|
||||
"the second ticket never reached its terminal phase");
|
||||
} finally {
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
|
||||
}
|
||||
|
||||
awaitNudge(wiring.leadHerdr());
|
||||
Thread.sleep(200); // settle — nothing more should arrive beyond the one coalesced nudge
|
||||
assertEquals(1, wiring.scheduler().runDueTasks(),
|
||||
"both terminal tickets must coalesce onto one scheduled tick");
|
||||
long nudgeCount = wiring.leadHerdr().calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt")).count();
|
||||
assertEquals(1, nudgeCount, "two tickets finishing together must produce ONE nudge, not two");
|
||||
@@ -1747,7 +2073,8 @@ class MessageServiceTest {
|
||||
|
||||
@Test
|
||||
void answeringAQuestionStopsFurtherNudgesAboutIt() throws Exception {
|
||||
try (var wiring = wireWithPushLoop(5, 50)) {
|
||||
// A 1ms backoff is due immediately, but ManualScheduler only ticks when this test asks it to.
|
||||
try (var wiring = wireWithManualScheduler(5, 1)) {
|
||||
String ticket = wiring.service().sendAsync(T, "task that asks");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
@@ -1755,8 +2082,9 @@ class MessageServiceTest {
|
||||
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
|
||||
() -> wiring.service().ask(T, "which config file?", 5000));
|
||||
MessageService.TaskView asking = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.ASKING);
|
||||
awaitQuestionPendingOn(wiring.pushLoop(), LEAD, asking.turnId());
|
||||
|
||||
awaitNudge(wiring.leadHerdr());
|
||||
assertEquals(1, wiring.scheduler().runDueTasks(), "the open question must have one scheduled tick");
|
||||
long callsBeforeAnswer = wiring.leadHerdr().calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt")).count();
|
||||
|
||||
@@ -1764,12 +2092,22 @@ class MessageServiceTest {
|
||||
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting();
|
||||
java.util.concurrent.CountDownLatch terminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(terminalReached::countDown);
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
try {
|
||||
assertTrue(terminalReached.await(5, TimeUnit.SECONDS),
|
||||
"the answered ticket never reached its terminal phase");
|
||||
} finally {
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
|
||||
}
|
||||
answer.get(5, TimeUnit.SECONDS);
|
||||
|
||||
// Let several more ticks (and the ticket's own now-legitimate terminal nudge) fire —
|
||||
// none of them may still name the question's turnId, which is closed.
|
||||
Thread.sleep(300);
|
||||
// Run the ticket's legitimate terminal nudge and every later scheduled tick through its
|
||||
// reminder cap. None may still name the closed question.
|
||||
for (int tick = 0; tick < 6; tick++) {
|
||||
assertEquals(1, wiring.scheduler().runDueTasks(), "expected one scheduled reminder tick");
|
||||
}
|
||||
boolean anyNamesClosedQuestion = wiring.leadHerdr().calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt"))
|
||||
.skip(callsBeforeAnswer)
|
||||
@@ -1852,35 +2190,45 @@ class MessageServiceTest {
|
||||
// decideTickets hits the cap and STOPs — activeLeads drops the lead, but (before the fix)
|
||||
// pendingTickets never drops the ticket. That is the exact "cap already STOPped" branch of
|
||||
// the bug report, reached deterministically rather than by timing it against a live tick.
|
||||
try (var wiring = wireWithPushLoop(1, 50, clock::get)) {
|
||||
String stale = wiring.service().sendAsync(T, "first task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "stale result"));
|
||||
// A 1ms backoff is due immediately, but ManualScheduler runs only the ticks below.
|
||||
try (var wiring = wireWithManualScheduler(1, 1, clock::get)) {
|
||||
String stale;
|
||||
String fresh;
|
||||
java.util.concurrent.CountDownLatch staleTerminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(staleTerminalReached::countDown);
|
||||
try {
|
||||
stale = wiring.service().sendAsync(T, "first task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "stale result"));
|
||||
assertTrue(staleTerminalReached.await(5, TimeUnit.SECONDS),
|
||||
"the stale ticket never reached its terminal phase");
|
||||
|
||||
// Let the reminder loop fire its one nudge and hit the cap (STOP removes it from
|
||||
// activeLeads; pendingTickets is untouched either way — that asymmetry is the bug).
|
||||
awaitNudge(wiring.leadHerdr());
|
||||
Thread.sleep(300);
|
||||
assertTrue(wiring.leadHerdr().lastCall("agent.prompt").params().toString().contains(stale),
|
||||
"sanity: the stale ticket's own reminder must have fired first");
|
||||
// Fire the stale ticket's one nudge, then its cap tick (STOP removes it from activeLeads;
|
||||
// pendingTickets is untouched either way — that asymmetry is the bug).
|
||||
assertEquals(1, wiring.scheduler().runDueTasks(), "the stale ticket must have one nudge tick");
|
||||
assertEquals(1, wiring.scheduler().runDueTasks(), "the stale ticket must have one cap tick");
|
||||
assertTrue(wiring.leadHerdr().lastCall("agent.prompt").params().toString().contains(stale),
|
||||
"sanity: the stale ticket's own reminder must have fired first");
|
||||
|
||||
// Cross the TTL — a real clock would need 10 minutes; the injected one does it instantly.
|
||||
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
|
||||
// Cross the TTL — a real clock would need 10 minutes; the injected one does it instantly.
|
||||
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
|
||||
|
||||
// A second, unrelated ticket to the same target/lead reaches sendAsync, which prunes.
|
||||
String fresh = wiring.service().sendAsync(T, "second task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "fresh result"));
|
||||
|
||||
// The fresh ticket restarts the (now-dormant) reminder loop with its own nudge.
|
||||
long before = wiring.leadHerdr().calls.stream().filter(c -> c.method().equals("agent.prompt")).count();
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (wiring.leadHerdr().calls.stream().filter(c -> c.method().equals("agent.prompt")).count() <= before
|
||||
&& System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(10);
|
||||
// A second, unrelated ticket to the same target/lead reaches sendAsync, which prunes.
|
||||
java.util.concurrent.CountDownLatch freshTerminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(freshTerminalReached::countDown);
|
||||
fresh = wiring.service().sendAsync(T, "second task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "fresh result"));
|
||||
assertTrue(freshTerminalReached.await(5, TimeUnit.SECONDS),
|
||||
"the fresh ticket never reached its terminal phase");
|
||||
} finally {
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
|
||||
}
|
||||
|
||||
// The fresh ticket restarts the now-dormant reminder loop with its own nudge.
|
||||
assertEquals(1, wiring.scheduler().runDueTasks(), "the fresh ticket must have one nudge tick");
|
||||
String latestNudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
|
||||
assertTrue(latestNudge.contains(fresh), "the fresh ticket's nudge must still arrive: " + latestNudge);
|
||||
assertFalse(latestNudge.contains(stale),
|
||||
|
||||
@@ -10,12 +10,13 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
class MemberRoleTest {
|
||||
|
||||
@Test
|
||||
void theThreeRolesAreArchitectDevAndReviewer() {
|
||||
assertEquals(3, MemberRole.values().length,
|
||||
void theFourRolesAreArchitectDevHunterAndReviewer() {
|
||||
assertEquals(4, MemberRole.values().length,
|
||||
"a new role changes the charter, the role file, the skill and the authz row — "
|
||||
+ "adding one is a deliberate act, so this count is meant to fail first");
|
||||
assertEquals("architect", MemberRole.ARCHITECT.wireName());
|
||||
assertEquals("dev", MemberRole.DEV.wireName());
|
||||
assertEquals("hunter", MemberRole.HUNTER.wireName());
|
||||
assertEquals("reviewer", MemberRole.REVIEWER.wireName());
|
||||
}
|
||||
|
||||
@@ -30,6 +31,7 @@ class MemberRoleTest {
|
||||
void parseIsCaseInsensitiveAndTrimsSurroundingSpace() {
|
||||
assertSame(MemberRole.ARCHITECT, MemberRole.parse("Architect"));
|
||||
assertSame(MemberRole.DEV, MemberRole.parse(" DEV "));
|
||||
assertSame(MemberRole.HUNTER, MemberRole.parse("HuNtEr"));
|
||||
assertSame(MemberRole.REVIEWER, MemberRole.parse("ReViEwEr"));
|
||||
}
|
||||
|
||||
@@ -38,7 +40,7 @@ class MemberRoleTest {
|
||||
IllegalArgumentException e =
|
||||
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse("archtiect"));
|
||||
assertTrue(e.getMessage().contains("archtiect"), e.getMessage());
|
||||
assertTrue(e.getMessage().contains("architect, dev, reviewer"),
|
||||
assertTrue(e.getMessage().contains("architect, dev, hunter, reviewer"),
|
||||
"a typo in config should be fixable from the message alone: " + e.getMessage());
|
||||
}
|
||||
|
||||
|
||||
@@ -8,6 +8,7 @@ import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.inject.LoopWatchdog;
|
||||
import dev.ltms.fleet.inject.StatusPoller;
|
||||
import dev.ltms.fleet.inject.MemberPresence;
|
||||
import dev.ltms.fleet.msg.InMemoryReplyInbox;
|
||||
@@ -165,6 +166,29 @@ class FleetAppTest {
|
||||
assertEquals("degraded", mapper.readTree(res.body()).get("status").asText());
|
||||
}
|
||||
|
||||
@Test
|
||||
void healthzKeepsOkStatusAndReportsLoopHealthInItsBody() {
|
||||
FleetMcp.LoopHealthSource loops = new FleetMcp.LoopHealthSource(
|
||||
() -> LoopWatchdog.State.RUNNING, () -> LoopWatchdog.State.STOPPED);
|
||||
|
||||
FleetApp.HealthzResponse ok = FleetApp.healthzResponse(new FakeHerdr(), new FakeHerdr(), loops);
|
||||
assertEquals(200, ok.status(), "a healthy herdr must keep /healthz at 200 regardless of loop states");
|
||||
assertEquals(Map.of("statusPoller", "RUNNING", "sessionReaper", "STOPPED"), ok.body().get("loopHealth"),
|
||||
"the /healthz body must report each loop state without making STOPPED an alarm");
|
||||
}
|
||||
|
||||
@Test
|
||||
void healthzKeepsDegradedStatusAndReportsLoopHealthInItsBody() {
|
||||
FleetMcp.LoopHealthSource loops = new FleetMcp.LoopHealthSource(
|
||||
() -> LoopWatchdog.State.RUNNING, () -> LoopWatchdog.State.STOPPED);
|
||||
FleetApp.HealthzResponse degraded = FleetApp.healthzResponse(new FakeHerdr().healthy(false),
|
||||
new FakeHerdr(), loops);
|
||||
assertEquals(503, degraded.status(),
|
||||
"an unreachable herdr must keep /healthz at 503 regardless of loop states");
|
||||
assertEquals(Map.of("statusPoller", "RUNNING", "sessionReaper", "STOPPED"),
|
||||
degraded.body().get("loopHealth"), "the degraded /healthz body must retain loop states");
|
||||
}
|
||||
|
||||
@Test
|
||||
void sessionsMapsWorkspaceList() throws Exception {
|
||||
int port = startHealthy();
|
||||
@@ -623,6 +647,28 @@ class FleetAppTest {
|
||||
assertTrue(herdr.called("agent.prompt"), "message was injected");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #571: the worker is idle, so the poller attempts delivery, but the {@code agent.prompt}
|
||||
* call itself fails with a herdr error that is not a confirmed absence — {@link
|
||||
* dev.ltms.fleet.inject.Injector} marks this {@code ATTEMPTED}, meaning the call was made and
|
||||
* whether it reached the pane is unknown. {@code writeReply}'s default arm must map this to its
|
||||
* own {@code "unconfirmed"} status, not silently fall through to {@code "done"} (which would
|
||||
* claim the delegation completed) nor collapse into {@code "queued"} (which would claim the
|
||||
* message will never arrive, when it may already be sitting in the pane).
|
||||
*/
|
||||
@Test
|
||||
void messageTimesOutUnconfirmedWhenDeliveryAttemptFails() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr().agentStatus("idle").agentSendFailsWith("send_failed");
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":250}");
|
||||
assertEquals(202, res.statusCode());
|
||||
JsonNode body = mapper.readTree(res.body());
|
||||
assertEquals("unconfirmed", body.get("status").asText(),
|
||||
"an ATTEMPTED delivery must report its own status, not \"queued\" or \"done\"");
|
||||
assertTrue(herdr.called("agent.prompt"), "delivery must have been attempted");
|
||||
}
|
||||
|
||||
@Test
|
||||
void messageRejectsBlankContent() throws Exception {
|
||||
int port = startHealthy();
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user