Compare commits
81 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 141ae3b04d | |||
| faefea14c4 | |||
| ea6896f2ef | |||
| d7f94cafa2 | |||
| 096f08c866 | |||
| 4eb720029c | |||
| 0db6d31dc2 | |||
| 158a2a84b5 | |||
| 41ebc9cf69 | |||
| 619769bb52 | |||
| 180c953c42 | |||
| 4af919af0c | |||
| 09c37061c1 | |||
| 2be287ea03 | |||
| c3b0406826 | |||
| e76fa1660b | |||
| 9dca376604 | |||
| 91be5079d6 | |||
| 26f198675c | |||
| 72f46d7c0c | |||
| 640f4d5f23 | |||
| 6edeb70bc4 | |||
| bc49d87cb8 | |||
| 14b169c410 | |||
| 2b52324d9a | |||
| b6147a39f6 | |||
| dbfc34cb6d | |||
| a20cb96730 | |||
| cda1a6a917 | |||
| 608e4496be | |||
| fa61dc587c | |||
| 526e3b7459 | |||
| b6b006c651 | |||
| 63eec8a0da | |||
| cbe872b538 | |||
| 7d9a807243 | |||
| 8915e40c7d | |||
| 3f7bc3815e | |||
| 6cb31a10e4 | |||
| 8368a274a0 | |||
| 17127efb88 | |||
| 203f034528 | |||
| 9ee16f5b85 | |||
| 388ef5a3c3 | |||
| be6c45ff78 | |||
| e99cb70a8b | |||
| 987ccef4c7 | |||
| 076cc43f7b | |||
| bad47a8444 | |||
| 955b9ea013 | |||
| 89cb8ff79b | |||
| fa62e9906d | |||
| d7390ccd37 | |||
| aa517ae0ec | |||
| 60496831c2 | |||
| 9a992d0f70 | |||
| 3762aca307 | |||
| 4e27bde2d7 | |||
| b85d9b0e46 | |||
| 856dfc6318 | |||
| 81c1d8e91c | |||
| d345b14e43 | |||
| 9640deeffc | |||
| ad3d81941f | |||
| 8b986a52e0 | |||
| f336bcef39 | |||
| 3ed7bfca67 | |||
| f5c6a0e4fc | |||
| a7aee5b982 | |||
| bf895616a5 | |||
| b874afb0af | |||
| 6eb34a654f | |||
| 7084d99b89 | |||
| ae7845c375 | |||
| 61115f6f61 | |||
| 4b9ebda1b3 | |||
| 1e68d7ee39 | |||
| 42820fbe75 | |||
| d91ff886da | |||
| 386e760a5c | |||
| 2e349139e9 |
@@ -0,0 +1,23 @@
|
|||||||
|
---
|
||||||
|
name: hunter
|
||||||
|
description: Sweep one assigned scope for defects and report ranked findings without changes.
|
||||||
|
---
|
||||||
|
|
||||||
|
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
|
||||||
|
|
||||||
|
You sweep the assigned package or scope for real defects. Read the full assigned scope before you
|
||||||
|
judge it. Report several ranked findings when the evidence supports them. Change nothing: do not
|
||||||
|
edit code, commit, push, or open a pull request.
|
||||||
|
|
||||||
|
You may run the build or tests to check a finding. Read the complete output and report the real
|
||||||
|
result. Do not hide failures with a pipe. State only checks you actually ran. The primary's IDE
|
||||||
|
tools are not yours. A mounted forge tool may use a blocked credential and fail by design.
|
||||||
|
|
||||||
|
Do only the assigned scope. Note anything outside it in one line and do not investigate it further.
|
||||||
|
Use `fleet_ask{question}` only when a decision belongs to the lead, such as an unclear requirement
|
||||||
|
or two defensible fixes. Do not ask about something you can decide by reading more code.
|
||||||
|
|
||||||
|
Your handoff must name the files you read, each ranked finding or `NO FINDINGS`, the checks you ran,
|
||||||
|
and any caveat for review.
|
||||||
|
|
||||||
|
The launcher provides the required bridge reply instructions for every member.
|
||||||
@@ -167,19 +167,59 @@ fails.
|
|||||||
- **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved
|
- **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved
|
||||||
from your connection, so you can only ever roll yourself.
|
from your connection, so you can only ever roll yourself.
|
||||||
- **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you
|
- **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you
|
||||||
are confident. Ask, wait for the answer, then pass what they said. `requireOperatorConfirm`
|
are confident. Ask, wait for the answer, then pass what they said.
|
||||||
defaults to `true` and this is the only thing standing between a judgement call and a wiped
|
|
||||||
session.
|
- **Whether you must ask at all depends on `leadRollover.requireOperatorConfirm`. Check it; do not
|
||||||
|
assume.** The default is `true` (`FleetConfig.java:1426`), and then `confirm` refuses unless you
|
||||||
|
also pass `operatorConfirmed: true`. **This host set it to `false` on 2026-09-22**, on the
|
||||||
|
operator's explicit grant, because they do not want to approve routine context rolls. Where it is
|
||||||
|
`false`, the three handover-file checks are the whole gate: the file must exist, be fresher than
|
||||||
|
`maxDocAgeSeconds`, and have been modified after the open request.
|
||||||
|
|
||||||
|
Read the live value rather than trusting this line:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
grep -A1 'requireOperatorConfirm' fleetd/fleetd.yaml
|
||||||
|
```
|
||||||
|
|
||||||
|
No match means the key is unset, so the default `true` applies and you must ask. The key is
|
||||||
|
**deferred, not hot** — it is read once at boot, so an edit does nothing until the daemon is
|
||||||
|
redeployed.
|
||||||
|
|
||||||
|
**Until fleetd #621 merges, the nudge text will tell you to ask the operator even where the
|
||||||
|
daemon no longer requires it.** `LeadHeartbeatLoop.contextNotice()` hardcodes "ask the operator"
|
||||||
|
and takes no config, so it cannot know. Trust the config value over the nudge text. Once #621 is
|
||||||
|
merged and deployed, the nudge matches the config and this warning can be deleted.
|
||||||
|
|
||||||
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
|
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
|
||||||
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
|
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
|
||||||
- **The bootstrap prompt has never yet landed, and the fix is unproven (fleetd #489).** The first
|
- **The bootstrap prompt works end to end. Measured 2026-09-22.** This used to say the fix was
|
||||||
real rollover, on 2026-09-12, joined `/clear` and the bootstrap text into one line and Claude Code
|
unproven (fleetd #489) and told you to expect a failure. That is no longer true. The daemon log
|
||||||
refused it as `Unknown command: /clearFresh`. The pane was never cleared and no context was lost,
|
now holds four `lead-rollover: rolled` lines, and three of them ran on 2026-09-22 at 10:01:43,
|
||||||
so the failure was safe — the roll simply did nothing. PR #490 fixed the cause and is deployed,
|
10:38:28 and 11:15:47. Each one cleared the old lead and started a fresh session against the
|
||||||
but no roll has bootstrapped a fresh session end to end yet. **Assume it may still fail, and tell
|
handover file, with the configured `bootstrapText` arriving as its first message. No context was
|
||||||
the operator so before you confirm.** The recovery is the same either way: the file is already
|
lost. The old `Unknown command: /clearFresh` failure from 2026-09-12 does not appear in the log
|
||||||
written, so the operator starts a session and points it at the file. That is why you write the
|
at all. Re-measure both numbers with:
|
||||||
file before you confirm, and never the other way round.
|
|
||||||
|
```bash
|
||||||
|
grep -c "lead-rollover: rolled" fleetd/fleetd.out # successful rolls
|
||||||
|
grep -c "lead-rollover:" fleetd/fleetd.out # positive control: must be larger
|
||||||
|
grep -c "Unknown command" fleetd/fleetd.out # the old failure: expect 0
|
||||||
|
```
|
||||||
|
|
||||||
|
Run the control line too. A broken pattern returns a clean `0` that reads exactly like good news.
|
||||||
|
If the first number stops growing across rolls, or `Unknown command` returns anything above 0,
|
||||||
|
the bootstrap has regressed and this paragraph is stale again.
|
||||||
|
|
||||||
|
**You still write the file before you confirm, and never the other way round.** That order is not
|
||||||
|
about the bootstrap being unreliable. It is what the daemon checks: the handover file must have
|
||||||
|
been modified *after* the open request, or `confirm` refuses it as stale.
|
||||||
|
|
||||||
|
- **One warning in the log is normal and is not a failure.** Every one of the three rolls above also
|
||||||
|
logged `/clear on term_… was never observed as WORKING after 8 consecutive IDLE/DONE polls —
|
||||||
|
releasing rather than wedging the roll`. The daemon could not see the pane go WORKING after
|
||||||
|
`/clear`, so it released instead of hanging. The roll then succeeded anyway. That is the safe
|
||||||
|
branch behaving correctly. Do not report it as a broken roll.
|
||||||
|
|
||||||
## Writing style
|
## Writing style
|
||||||
|
|
||||||
|
|||||||
@@ -20,6 +20,13 @@ fleetd.out
|
|||||||
fleetd/fleetd.out
|
fleetd/fleetd.out
|
||||||
logs/
|
logs/
|
||||||
|
|
||||||
|
# fleetd #635 follow-up — scripts/config-edit.sh's backup directory. No leading slash, so this is
|
||||||
|
# ignored at every depth: the real one lives under fleetd/ (also named in fleetd/.gitignore, next
|
||||||
|
# to the config it backs up), and scripts/test-config-edit.sh's own throwaway fixtures build one
|
||||||
|
# under the repo root while the suite runs. --config can point anywhere, so the directory name is
|
||||||
|
# ignored everywhere rather than only where the live daemon happens to use it.
|
||||||
|
.config-backups/
|
||||||
|
|
||||||
# fleetd #480: the lead rollover handover file. `leadRollover.handoverPath` points here, and the
|
# fleetd #480: the lead rollover handover file. `leadRollover.handoverPath` points here, and the
|
||||||
# outgoing lead rewrites it on every rollover. It is a snapshot of one moment's live state —
|
# outgoing lead rewrites it on every rollover. It is a snapshot of one moment's live state —
|
||||||
# unpushed branches, running builds, open questions — so it is stale the moment it is written and
|
# unpushed branches, running builds, open questions — so it is stale the moment it is written and
|
||||||
|
|||||||
@@ -141,12 +141,14 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
|
|||||||
|
|
||||||
**When a decision blocks you, consult architects — not the operator.** Spawn one or more architect
|
**When a decision blocks you, consult architects — not the operator.** Spawn one or more architect
|
||||||
members, give them the question and the evidence you have, and act on what they agree. They are
|
members, give them the question and the evidence you have, and act on what they agree. They are
|
||||||
authorized to settle it, not only to advise. If two of them still disagree after two rounds, they
|
authorized to settle it, not only to advise. Architects first form independent positions, then
|
||||||
return both positions and you decide. Go to the operator only for something outside the fleet's
|
compare them. If they still disagree after that comparison, they return both positions and their
|
||||||
authority: money, credentials, or a promise made to someone else. **Then write the decision on the
|
checked evidence; the lead decides. Go to the operator only for an action the fleet has no
|
||||||
ticket.** Taking the operator out of the loop also removes the signal they used to get, because
|
authority to take, such as spending money, granting access, or making a promise to someone else.
|
||||||
that signal was the block itself — work stopped, so they found out. A ticket comment replaces it,
|
**Then write the decision on the ticket.** Taking the operator out of the loop also removes the
|
||||||
and it reaches them whether or not they are at a terminal when you decide.
|
signal they used to get, because that signal was the block itself — work stopped, so they found
|
||||||
|
out. A ticket comment replaces it, and it reaches them whether or not they are at a terminal when
|
||||||
|
you decide.
|
||||||
|
|
||||||
| Intent | Tool |
|
| Intent | Tool |
|
||||||
|---|---|
|
|---|---|
|
||||||
@@ -284,6 +286,8 @@ must obey belongs in the charter, not here.
|
|||||||
it for a multi-finding sweep hands the worker two contradictory output contracts. That has
|
it for a multi-finding sweep hands the worker two contradictory output contracts. That has
|
||||||
already cost three workers' turns: each wrote a good report to its terminal and ended the turn
|
already cost three workers' turns: each wrote a good report to its terminal and ended the turn
|
||||||
with no `fleet_reply`, and the scrape returned the tail of the brief instead.
|
with no `fleet_reply`, and the scrape returned the tail of the brief instead.
|
||||||
|
Spawn `implementer` with role `dev`, `reviewer` with role `reviewer`, and `hunter` with role
|
||||||
|
`hunter`.
|
||||||
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
|
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
|
||||||
`port-to-opencode` (make an OpenCode session a participant in this workspace),
|
`port-to-opencode` (make an OpenCode session a participant in this workspace),
|
||||||
`fleets-status` (report every fleet that shares one LavinMQ instance),
|
`fleets-status` (report every fleet that shares one LavinMQ instance),
|
||||||
|
|||||||
@@ -7,6 +7,13 @@ dependency-reduced-pom.xml
|
|||||||
fleetd.yaml
|
fleetd.yaml
|
||||||
bridged.yaml
|
bridged.yaml
|
||||||
|
|
||||||
|
# fleetd #635 follow-up — scripts/config-edit.sh's backups of fleetd.yaml. A backup of a file
|
||||||
|
# that must never be committed inherits that requirement. The directory is the real protection
|
||||||
|
# (it keeps working even if the backup naming changes); the glob is a backstop for a stray
|
||||||
|
# backup written the old way, directly beside fleetd.yaml, or by an older copy of the script.
|
||||||
|
.config-backups/
|
||||||
|
fleetd.yaml.bak.*
|
||||||
|
|
||||||
# CB-505 audit trail + daemon stdout/stderr — runtime records, never source
|
# CB-505 audit trail + daemon stdout/stderr — runtime records, never source
|
||||||
logs/
|
logs/
|
||||||
|
|
||||||
|
|||||||
+27
-12
@@ -77,16 +77,25 @@ bind:
|
|||||||
# a lead turn nobody asked for), so upgrading the daemon must never switch it on for you. Absent
|
# a lead turn nobody asked for), so upgrading the daemon must never switch it on for you. Absent
|
||||||
# block = feature off, exactly as before.
|
# block = feature off, exactly as before.
|
||||||
#
|
#
|
||||||
# Three knobs, each with a default that errs on the side of not burning context:
|
# Four knobs, each with a default that errs on the side of not burning context:
|
||||||
# idleAfterSeconds: 300 # how long the lead must stay idle before the FIRST nudge (default 300 —
|
# idleAfterSeconds: 300 # how long the lead must stay idle before the FIRST nudge (default 300 —
|
||||||
# # absorbs normal post-turn pauses; re-prompting every pause burns context)
|
# # absorbs normal post-turn pauses; re-prompting every pause burns context)
|
||||||
# backoffMs: 60000 # re-check cadence / spacing between nudges past the quiet period (default 60000)
|
# backoffMs: 60000 # re-check cadence / spacing between nudges past the quiet period (default 60000)
|
||||||
# quietNudgeCap: 3 # cap on consecutive nudges that find NOTHING pending, then it stops
|
# quietNudgeCap: 3 # cap on consecutive nudges that find NOTHING pending, then it stops
|
||||||
# # until real state appears (default 3 — never nag an empty fleet forever)
|
# # until real state appears (default 3 — never nag an empty fleet forever)
|
||||||
|
# contextHighNudge: false # fleetd #609 — when true, an idle lead whose OWN Claude Code context
|
||||||
|
# # reads HIGH (see LeadContextGauge; fleet_list's context row) gets a text
|
||||||
|
# # notice appended to its nudge telling it to consider fleet_handover. Text
|
||||||
|
# # only — it never rolls a pane by itself, and only the operator can approve
|
||||||
|
# # a roll. Fires once per HIGH stretch (a later OK reading re-arms it), and
|
||||||
|
# # never spends the quietNudgeCap budget. Default false/absent = off, same
|
||||||
|
# # as every other knob here — an upgraded daemon must not start telling
|
||||||
|
# # leads to hand over on its own.
|
||||||
# leadHeartbeat:
|
# leadHeartbeat:
|
||||||
# idleAfterSeconds: 300
|
# idleAfterSeconds: 300
|
||||||
# backoffMs: 60000
|
# backoffMs: 60000
|
||||||
# quietNudgeCap: 3
|
# quietNudgeCap: 3
|
||||||
|
# contextHighNudge: false
|
||||||
|
|
||||||
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
|
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
|
||||||
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
|
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
|
||||||
@@ -200,13 +209,13 @@ herdrSocket: ~/.config/herdr/herdr.sock
|
|||||||
# e.g. `env DISPLAY=:10.0 idea {dir}`. Best-effort: a failure is logged, never
|
# e.g. `env DISPLAY=:10.0 idea {dir}`. Best-effort: a failure is logged, never
|
||||||
# fails the spawn. Omit to open the member's module by hand. There is no close
|
# fails the spawn. Omit to open the member's module by hand. There is no close
|
||||||
# half yet — an opened module stays open until the operator closes it.
|
# half yet — an opened module stays open until the operator closes it.
|
||||||
# autoCompactWindow → opt-in, default off. A bounded token window that forces a spawned member to
|
# autoCompactWindow → opt-in, default off. A bounded token window that forces a launched Claude Code
|
||||||
# compact its context instead of running on the backend's own default and dying
|
# lead or member to compact its context instead of running on the backend's own
|
||||||
# mid-turn (losing its fleet_reply — the whole point of the turn — with it).
|
# default. A member that runs out of context can die mid-turn and lose its fleet_reply.
|
||||||
# Validated at config load to [100000, 1000000] — the band Claude Code's own
|
# Validated at config load to [100000, 1000000] — the band Claude Code's own
|
||||||
# --autocompact flag accepts.
|
# --autocompact flag accepts.
|
||||||
# CROSS-BACKEND SEMANTICS DIFFER: on claude-code this is a launch-time
|
# CROSS-BACKEND SEMANTICS DIFFER: on claude-code this is a launch-time
|
||||||
# `--autocompact <tokens>` flag — the member compacts AT this window. opencode
|
# `--autocompact <tokens>` flag — the Claude Code session compacts AT this window. opencode
|
||||||
# has no equivalent flag (it only forces `compaction.auto: true`, unconditionally,
|
# has no equivalent flag (it only forces `compaction.auto: true`, unconditionally,
|
||||||
# already), so this is instead applied as the model's `limit.context` in the
|
# already), so this is instead applied as the model's `limit.context` in the
|
||||||
# generated opencode.json — the member compacts WITHIN this window, not exactly
|
# generated opencode.json — the member compacts WITHIN this window, not exactly
|
||||||
@@ -406,7 +415,7 @@ profiles:
|
|||||||
# ideMcpUrl: http://127.0.0.1:29170/index-mcp/streamable-http # opt-in (CB-634): IDE code intelligence, pinned to the worktree
|
# ideMcpUrl: http://127.0.0.1:29170/index-mcp/streamable-http # opt-in (CB-634): IDE code intelligence, pinned to the worktree
|
||||||
# ideProjectDir: fleetd # CB-634: module dir the IDE opens + the overlay pins (this repo's pom is in fleetd/)
|
# ideProjectDir: fleetd # CB-634: module dir the IDE opens + the overlay pins (this repo's pom is in fleetd/)
|
||||||
# ideOpenCommand: env DISPLAY=:10.0 idea {dir} # CB-634 auto-open: opens {dir} in the IDE at spawn; omit to open by hand
|
# ideOpenCommand: env DISPLAY=:10.0 idea {dir} # CB-634 auto-open: opens {dir} in the IDE at spawn; omit to open by hand
|
||||||
# autoCompactWindow: 250000 # opt-in: bound member context; claude-code compacts AT this, opencode within it (model limit.context)
|
# autoCompactWindow: 250000 # opt-in: bound Claude Code lead/member context; claude-code compacts AT this, opencode within it (model limit.context)
|
||||||
gx11: # a second backend, so `placement: weighted` has a choice
|
gx11: # a second backend, so `placement: weighted` has a choice
|
||||||
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
|
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
|
||||||
placement: tab
|
placement: tab
|
||||||
@@ -560,7 +569,7 @@ placement: weighted
|
|||||||
# older keys: `leaders:`, `members:`, `leadScan:` and `defaultProfile:`.
|
# older keys: `leaders:`, `members:`, `leadScan:` and `defaultProfile:`.
|
||||||
#
|
#
|
||||||
# A member is anything a lead spawns, and every member has two INDEPENDENT attributes:
|
# A member is anything a lead spawns, and every member has two INDEPENDENT attributes:
|
||||||
# role — which contract: architect, dev or reviewer. It picks the launch charter, the role
|
# role — which contract: architect, dev, hunter or reviewer. It picks the launch charter, the role
|
||||||
# file, the playbook skill and the authz row.
|
# file, the playbook skill and the authz row.
|
||||||
# profile — which backend: one of the `profiles:` keys above (model, CLI adapter, cost).
|
# profile — which backend: one of the `profiles:` keys above (model, CLI adapter, cost).
|
||||||
# They vary on their own. A reviewer may run on the same profile as the dev whose diff it reads,
|
# They vary on their own. A reviewer may run on the same profile as the dev whose diff it reads,
|
||||||
@@ -572,13 +581,13 @@ placement: weighted
|
|||||||
#
|
#
|
||||||
# Each pool lists the profiles that role MAY run on — these are pools, not identities. That is also
|
# Each pool lists the profiles that role MAY run on — these are pools, not identities. That is also
|
||||||
# what replaced `defaultProfile:`: an unqualified spawn names a role, and that role's pool supplies
|
# what replaced `defaultProfile:`: an unqualified spawn names a role, and that role's pool supplies
|
||||||
# the candidates, in definition order. A dev and a reviewer staying anonymous is exactly compatible
|
# the candidates, in definition order. A dev, hunter and reviewer staying anonymous is exactly
|
||||||
# with being listed here; the entry key just names the entry.
|
# compatible with being listed here; the entry key just names the entry.
|
||||||
fleet:
|
fleet:
|
||||||
# Optional launch-charter text, keyed only by the singular role wire names: architect, dev,
|
# Optional launch-charter text, keyed only by the singular role wire names: architect, dev,
|
||||||
# reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put secrets
|
# hunter, reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put
|
||||||
# here: a later launch step writes this text to a world-readable temp file, and ${ENV} interpolation
|
# secrets here: a later launch step writes this text to a world-readable temp file, and ${ENV}
|
||||||
# is deliberately not supported.
|
# interpolation is deliberately not supported.
|
||||||
charters:
|
charters:
|
||||||
architect: |-
|
architect: |-
|
||||||
You are an architect in this fleet. You refine work before anyone builds it:
|
You are an architect in this fleet. You refine work before anyone builds it:
|
||||||
@@ -589,6 +598,9 @@ fleet:
|
|||||||
dev: |-
|
dev: |-
|
||||||
You implement the one unit you were given, and nothing else. You test it,
|
You implement the one unit you were given, and nothing else. You test it,
|
||||||
commit it, and open your own pull request. You never merge.
|
commit it, and open your own pull request. You never merge.
|
||||||
|
hunter: |-
|
||||||
|
You sweep the assigned scope for real defects. You may run the build or tests
|
||||||
|
to check a finding. You change nothing, and report several ranked findings.
|
||||||
reviewer: |-
|
reviewer: |-
|
||||||
You review the diff you were given. You report bugs, risks and missing tests.
|
You review the diff you were given. You report bugs, risks and missing tests.
|
||||||
You do not change code.
|
You do not change code.
|
||||||
@@ -666,6 +678,9 @@ fleet:
|
|||||||
developers:
|
developers:
|
||||||
gx10:
|
gx10:
|
||||||
profile: gx10
|
profile: gx10
|
||||||
|
# hunters:
|
||||||
|
# gx10:
|
||||||
|
# profile: gx10 # a hunt may run checks, but never changes code
|
||||||
# reviewers:
|
# reviewers:
|
||||||
# gx10:
|
# gx10:
|
||||||
# profile: gx10 # the same backend may serve two roles; that is the point
|
# profile: gx10 # the same backend may serve two roles; that is the point
|
||||||
|
|||||||
@@ -0,0 +1,22 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 Unit A: everything {@code Fleetd.main} has ready once config is loaded, reported
|
||||||
|
* and validated — the exact point {@code main} used to keep going straight into socket and broker
|
||||||
|
* work. Building this (and handing it to {@link FleetdAssembly#assembleAndStart}) is the seam a
|
||||||
|
* test now has to drive the real boot composition without being {@code main} itself.
|
||||||
|
*
|
||||||
|
* @param cfg the boot-time {@link FleetConfig} snapshot every one-time wiring decision reads —
|
||||||
|
* see {@code Fleetd.main}'s own comment on why this must never be swapped for a live
|
||||||
|
* reference once loaded
|
||||||
|
* @param config the live {@link ConfigRef} the hot-reloadable paths read per use
|
||||||
|
* @param guard the same {@link SubscriptionGuard} {@code Fleetd.main} already used to assert the
|
||||||
|
* launching environment is clean, reused rather than rebuilt so the assembled
|
||||||
|
* launchers see the identical instance {@code main} already validated with
|
||||||
|
*/
|
||||||
|
record AssemblyInputs(FleetConfig cfg, ConfigRef config, SubscriptionGuard guard) {
|
||||||
|
}
|
||||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,536 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.auth.CallerResolver;
|
||||||
|
import dev.ltms.fleet.auth.MemberRegistry;
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.ConfigWatcher;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.health.FleetHealthMonitor;
|
||||||
|
import dev.ltms.fleet.herdr.AgentControl;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrRouter;
|
||||||
|
import dev.ltms.fleet.herdr.LeadTabScanner;
|
||||||
|
import dev.ltms.fleet.herdr.PaneLocator;
|
||||||
|
import dev.ltms.fleet.herdr.UnixSocketHerdrClient;
|
||||||
|
import dev.ltms.fleet.inject.BackendErrorPatternLookup;
|
||||||
|
import dev.ltms.fleet.inject.BackendErrorSink;
|
||||||
|
import dev.ltms.fleet.inject.CompletionResolver;
|
||||||
|
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
|
||||||
|
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||||
|
import dev.ltms.fleet.inject.Injector;
|
||||||
|
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
|
||||||
|
import dev.ltms.fleet.inject.MemberPresence;
|
||||||
|
import dev.ltms.fleet.inject.StatusPoller;
|
||||||
|
import dev.ltms.fleet.inject.TurnListener;
|
||||||
|
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||||
|
import dev.ltms.fleet.lead.LeadLauncher;
|
||||||
|
import dev.ltms.fleet.lead.LeadRollover;
|
||||||
|
import dev.ltms.fleet.mcp.ConnectionIdentity;
|
||||||
|
import dev.ltms.fleet.mcp.FleetMcp;
|
||||||
|
import dev.ltms.fleet.mcp.LsofPeerPidLookup;
|
||||||
|
import dev.ltms.fleet.mcp.LsofProcessCwdLookup;
|
||||||
|
import dev.ltms.fleet.mcp.PrimaryRegistry;
|
||||||
|
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||||
|
import dev.ltms.fleet.member.HerdrPeerLauncher;
|
||||||
|
import dev.ltms.fleet.member.MemberCredentialPolicyView;
|
||||||
|
import dev.ltms.fleet.metrics.FleetMetrics;
|
||||||
|
import dev.ltms.fleet.metrics.Metrics;
|
||||||
|
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||||
|
import dev.ltms.fleet.msg.LeadCoordLoop;
|
||||||
|
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
|
||||||
|
import dev.ltms.fleet.msg.MessageService;
|
||||||
|
import dev.ltms.fleet.msg.Rendezvous;
|
||||||
|
import dev.ltms.fleet.msg.ReplyInbox;
|
||||||
|
import dev.ltms.fleet.msg.ReplyPushLoop;
|
||||||
|
import dev.ltms.fleet.peer.PeerLauncher;
|
||||||
|
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||||
|
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||||
|
import dev.ltms.fleet.power.CaffeinateSleepAssertionMechanism;
|
||||||
|
import dev.ltms.fleet.power.IdleSleepGuard;
|
||||||
|
import dev.ltms.fleet.rest.FleetApp;
|
||||||
|
import dev.ltms.fleet.session.GitWorktrees;
|
||||||
|
import dev.ltms.fleet.session.SessionManager;
|
||||||
|
import dev.ltms.fleet.session.SessionReaper;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.slf4j.Logger;
|
||||||
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.ArrayList;
|
||||||
|
import java.util.LinkedHashMap;
|
||||||
|
import java.util.LinkedHashSet;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.Set;
|
||||||
|
import java.util.concurrent.ConcurrentHashMap;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.atomic.AtomicReference;
|
||||||
|
import java.util.function.Function;
|
||||||
|
import java.util.function.Predicate;
|
||||||
|
import java.util.function.Supplier;
|
||||||
|
import java.util.regex.Pattern;
|
||||||
|
import java.util.stream.Collectors;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 Unit A: the real boot assembly, extracted out of {@code Fleetd.main} so a test can
|
||||||
|
* drive it directly. {@link #assembleAndStart} is <em>the same statements {@code main} used to run
|
||||||
|
* inline</em>, in the same order, against a real {@link ResourcePorts} in production and a fake one
|
||||||
|
* in a test — see {@code FleetdAssemblyLifecycleTest}. {@code Fleetd.main} keeps config loading and
|
||||||
|
* {@code cfg.validateAll()}; everything from immediately after that call onward moved here —
|
||||||
|
* including, since fleetd #612 A-gaps (gap 2), the two post-validation reports ({@code
|
||||||
|
* reportRoleFallbackGaps}, {@code assertChartersNameOnlyRegisteredTools}) that Unit A originally
|
||||||
|
* left behind in {@code main}. Those two calls do no I/O themselves, but leaving them outside this
|
||||||
|
* method meant deleting either one compiled clean and left the whole suite green — nothing drove
|
||||||
|
* {@code main} itself, so nothing could notice. They run first here, in the same relative order,
|
||||||
|
* before the herdr socket or anything else that touches the outside world.
|
||||||
|
*
|
||||||
|
* <p><strong>Construction and start order is preserved exactly, on purpose.</strong> This is not
|
||||||
|
* rebuilt into "construct everything, then start everything" — that would change boot timing. The
|
||||||
|
* order recorded before any code moved (see the ticket and {@code FleetdAssemblyLifecycleTest}):
|
||||||
|
* {@code SessionReaper} starts first (if {@code lifecycle.idleTtlSeconds} is configured), then
|
||||||
|
* {@link StatusPoller}, then the optional {@link LeadHeartbeatLoop} and {@link FleetHealthMonitor},
|
||||||
|
* then the optional {@link LeadCoordLoop} and {@link ConfigWatcher}, and the HTTP server starts
|
||||||
|
* last of all. The close order (see {@link FleetdRuntime#close()}) is the mirror the original
|
||||||
|
* shutdown hook always used.
|
||||||
|
*
|
||||||
|
* <p><strong>One statement could not move without reordering startup.</strong> {@code Fleetd.main}
|
||||||
|
* registered its shutdown hook <em>before</em> building the Javalin {@code FleetApp} — the hook
|
||||||
|
* itself never touched {@code app} (it still doesn't; see {@link FleetdRuntime#close()}), but the
|
||||||
|
* hook needs a {@link FleetdRuntime} to close over, and the ticket asks that runtime to also own
|
||||||
|
* {@code FleetApp}. Building the runtime before the app exists and mutating it afterward (via
|
||||||
|
* {@link FleetdRuntime#attachApp}) preserves the exact original order — hook registered, then app
|
||||||
|
* built, then HTTP started — without moving the app's construction earlier or the hook's
|
||||||
|
* registration later. That is the one seam this ticket did not get to pin any other way.
|
||||||
|
*
|
||||||
|
* <p><strong>No inert variant.</strong> Deliberately, there is no overload of this method that
|
||||||
|
* accepts a smaller/optional {@link ResourcePorts} or defaults one internally. A future edit that
|
||||||
|
* wants to skip {@code FleetdAssembly} entirely and build its own graph is still possible — no
|
||||||
|
* static analysis stops that — but it cannot do so by quietly swapping this call for an inert
|
||||||
|
* substitute that still compiles, because none exists.
|
||||||
|
*/
|
||||||
|
final class FleetdAssembly {
|
||||||
|
|
||||||
|
private static final Logger log = LoggerFactory.getLogger(Fleetd.class);
|
||||||
|
|
||||||
|
/** CB-637: how often the lead coordination loop looks for peer messages — see {@code Fleetd}'s own constant. */
|
||||||
|
private static final long LEAD_COORD_INTERVAL_MS = 3_000L;
|
||||||
|
|
||||||
|
private FleetdAssembly() {
|
||||||
|
}
|
||||||
|
|
||||||
|
static FleetdRuntime assembleAndStart(AssemblyInputs inputs, ResourcePorts ports) {
|
||||||
|
FleetConfig cfg = inputs.cfg();
|
||||||
|
ConfigRef config = inputs.config();
|
||||||
|
SubscriptionGuard guard = inputs.guard();
|
||||||
|
|
||||||
|
// fleetd #612 A-gaps (gap 2): moved in from Fleetd.main, immediately after cfg.validateAll()
|
||||||
|
// there — the exact point main used to call these two, and still the first thing this
|
||||||
|
// method does, before any socket or broker work below. See this class's javadoc and each
|
||||||
|
// method's own for why they run here rather than in FleetConfig#validateAll() itself.
|
||||||
|
Fleetd.reportRoleFallbackGaps(cfg);
|
||||||
|
Fleetd.assertChartersNameOnlyRegisteredTools(cfg);
|
||||||
|
|
||||||
|
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
|
||||||
|
? Path.of(cfg.herdrSocket())
|
||||||
|
: UnixSocketHerdrClient.defaultSocketPath();
|
||||||
|
|
||||||
|
HerdrClient herdr = ports.connectHerdr(socket);
|
||||||
|
HerdrClient memberHerdr = cfg.memberHerdrSocket() != null && !cfg.memberHerdrSocket().isBlank()
|
||||||
|
? ports.connectHerdr(Path.of(cfg.memberHerdrSocket()))
|
||||||
|
: herdr;
|
||||||
|
AtomicReference<Supplier<Map<String, String>>> leadsRef = new AtomicReference<>(Map::of);
|
||||||
|
HerdrRouter router = new HerdrRouter(herdr, memberHerdr,
|
||||||
|
target -> leadsRef.get().get().containsKey(target));
|
||||||
|
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
|
||||||
|
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
|
||||||
|
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
|
||||||
|
Map<String, FleetConfig.Profile> claudeProfiles = new LinkedHashMap<>();
|
||||||
|
Map<String, FleetConfig.Profile> opencodeProfiles = new LinkedHashMap<>();
|
||||||
|
cfg.profiles().forEach((name, w) -> {
|
||||||
|
if (w.isOpenCode()) {
|
||||||
|
opencodeProfiles.put(name, w);
|
||||||
|
} else {
|
||||||
|
claudeProfiles.put(name, w);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
List<HerdrPeerLauncher> adapters = new ArrayList<>();
|
||||||
|
// fleetd #175: the daemon's real ExhaustionSink can only be built once `sessions` exists
|
||||||
|
// (below), but `sessions` needs `workers`, which needs the adapters built right here — a
|
||||||
|
// genuine cycle. Break it exactly like liveCountRef below: a forwarding sink built now,
|
||||||
|
// pointed at the real one once it exists.
|
||||||
|
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
|
||||||
|
ExhaustionSink forwardingExhaustionSink = Fleetd.forwardingExhaustionSink(exhaustionSinkRef);
|
||||||
|
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
|
||||||
|
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
|
||||||
|
// unless opencode is the only kind configured.
|
||||||
|
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
|
||||||
|
adapters.add(Fleetd.claudeCodeLauncher(router.memberAgents(), router.memberSpaces(), guard,
|
||||||
|
claudeProfiles, cfg, config));
|
||||||
|
}
|
||||||
|
if (!opencodeProfiles.isEmpty()) {
|
||||||
|
adapters.add(Fleetd.openCodeLauncher(router.memberAgents(), router.memberSpaces(),
|
||||||
|
opencodeProfiles, cfg, config, forwardingExhaustionSink));
|
||||||
|
}
|
||||||
|
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
|
||||||
|
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
|
||||||
|
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
|
||||||
|
BackendQuarantine quarantine = BackendQuarantine.withEscalation(ports.nanoClock(),
|
||||||
|
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
|
||||||
|
// fleetd #201 Unit 5: one outage-cool-off tracker for the whole daemon, shared between the
|
||||||
|
// launcher (checked at spawn, like `quarantine` above) and the backend-error sink wired in
|
||||||
|
// below (written on a classified backend error).
|
||||||
|
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(ports.nanoClock());
|
||||||
|
PeerLauncher workers = new CompositePeerLauncher(
|
||||||
|
adapters,
|
||||||
|
cfg.effectiveDefaultProfile(),
|
||||||
|
config,
|
||||||
|
profileName -> liveCountRef.get().apply(profileName),
|
||||||
|
quarantine,
|
||||||
|
outagePolicy);
|
||||||
|
// fleetd #422 follow-up: say which of the three model-gate states the daemon booted into.
|
||||||
|
log.info("model gate (fleetd #422): {}", Fleetd.modelGateCoverageLine(workers.modelGateState()));
|
||||||
|
// CB-504: under supervision (launchd/systemd) fleetd can start before herdr's socket
|
||||||
|
// exists. Wait, then degrade rather than die: serving with /healthz reporting "degraded" is
|
||||||
|
// strictly more useful than exiting.
|
||||||
|
Fleetd.HerdrAwaitOutcome herdrOutcome = Fleetd.awaitHerdr(herdr, ports.nanoClock(), ports.herdrPollWait());
|
||||||
|
boolean herdrUp = Fleetd.logHerdrWaitOutcomeAndShouldReap(herdrOutcome);
|
||||||
|
if (herdrUp) {
|
||||||
|
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
|
||||||
|
// with the previous process — reap those leaked orphans now, before we start serving.
|
||||||
|
workers.reapOrphanWorkers();
|
||||||
|
}
|
||||||
|
|
||||||
|
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
|
||||||
|
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
|
||||||
|
// CB-303 part 2: context cap is opt-in and disabled (0) when absent/null.
|
||||||
|
int contextCap = 0;
|
||||||
|
if (cfg.lifecycle() != null && cfg.lifecycle().contextCap() != null
|
||||||
|
&& cfg.lifecycle().contextCap() > 0) {
|
||||||
|
contextCap = cfg.lifecycle().contextCap();
|
||||||
|
}
|
||||||
|
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
|
||||||
|
SessionManager sessions = new SessionManager(workers,
|
||||||
|
new GitWorktrees(cfg.worktreeRoot(), cfg.worktreeGroup(), cfg.memberSkills()),
|
||||||
|
ports.nanoClock(), contextCap, clearAfterTurn);
|
||||||
|
liveCountRef.set(profileName -> Fleetd.liveSessionCount(sessions.roster(), profileName));
|
||||||
|
|
||||||
|
// Idle-sleep guard: hold an OS-level assertion against idle sleep while at least one
|
||||||
|
// member is live. No-op (never constructed) off macOS or when idleSleepGuard.enabled is
|
||||||
|
// explicitly false; the mechanism itself is additionally a no-op if 'caffeinate' cannot be
|
||||||
|
// started, so this can never fail a spawn, a release, or startup.
|
||||||
|
boolean idleSleepGuardEnabled = cfg.idleSleepGuard() == null || cfg.idleSleepGuard().isEnabled();
|
||||||
|
final IdleSleepGuard idleSleepGuard;
|
||||||
|
if (idleSleepGuardEnabled) {
|
||||||
|
idleSleepGuard = new IdleSleepGuard(new CaffeinateSleepAssertionMechanism(), sessions::size);
|
||||||
|
sessions.onAcquire(_ -> idleSleepGuard.recheck());
|
||||||
|
sessions.onRelease(_ -> idleSleepGuard.recheck());
|
||||||
|
} else {
|
||||||
|
idleSleepGuard = null;
|
||||||
|
}
|
||||||
|
|
||||||
|
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled. FIRST of the
|
||||||
|
// recurring background loops to start — see this class's own javadoc for the full order.
|
||||||
|
final SessionReaper reaper;
|
||||||
|
if (cfg.lifecycle() != null
|
||||||
|
&& cfg.lifecycle().idleTtlSeconds() != null
|
||||||
|
&& cfg.lifecycle().idleTtlSeconds() > 0) {
|
||||||
|
reaper = new SessionReaper(sessions, cfg.lifecycle().idleTtlSeconds());
|
||||||
|
reaper.start();
|
||||||
|
} else {
|
||||||
|
reaper = null;
|
||||||
|
}
|
||||||
|
|
||||||
|
// CB-530: every pane the config names as a lead, merged from `leaders:` and the legacy
|
||||||
|
// singular pin. PrimaryRegistry below still tracks ONE terminal — it addresses the push
|
||||||
|
// loop's nudges, which need a single destination — so it keeps the legacy pin.
|
||||||
|
Map<String, String> leadTerminals = cfg.leaderTerminals();
|
||||||
|
if (leadTerminals.size() > 1) {
|
||||||
|
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
|
||||||
|
}
|
||||||
|
// CB-531/CB-579: discover leads by the tab labels the operator writes, one scanner per
|
||||||
|
// configured lead's own exact `tab:` label.
|
||||||
|
final Supplier<Map<String, String>> leads;
|
||||||
|
var leaders = cfg.fleet().leaders();
|
||||||
|
if (!leaders.isEmpty()) {
|
||||||
|
Map<String, String> tabToName = new LinkedHashMap<>();
|
||||||
|
leaders.forEach((name, leader) -> {
|
||||||
|
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
|
||||||
|
tabToName.put(leader.tab(), name);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
|
||||||
|
// This must use the lead daemon: scanning member tabs would demote the lead to a worker.
|
||||||
|
leads = new LeadTabScanner(herdr, tabToName, Set.of(),
|
||||||
|
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), ports.nanoClock());
|
||||||
|
log.info("lead scan: tabs {} host a lead (rescan every {}s, shared fleet space)",
|
||||||
|
tabToName.keySet(), scanIntervalSeconds);
|
||||||
|
} else {
|
||||||
|
leads = () -> leadTerminals;
|
||||||
|
}
|
||||||
|
leadsRef.set(leads);
|
||||||
|
|
||||||
|
// CB-558: start any declared lead that is not already running. After the scanner is built,
|
||||||
|
// and only when herdr answered — the launcher's whole safety property is that it can count
|
||||||
|
// live leads first, and must never guess and risk a second orchestrator.
|
||||||
|
if (herdrUp && !leaders.isEmpty()) {
|
||||||
|
int launched = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg).ensureLeads();
|
||||||
|
if (launched > 0) {
|
||||||
|
log.info("lead auto-launch: {} lead(s) started", launched);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// CB-548: config-declared architect slots. Nothing here spawns a slot; the terminal → slot
|
||||||
|
// binding is owned by the registry and empty at startup.
|
||||||
|
MemberRegistry members = MemberRegistry.live(() -> config.get().fleet());
|
||||||
|
sessions.setMemberLifecycle(members);
|
||||||
|
if (!members.slots().isEmpty()) {
|
||||||
|
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
|
||||||
|
+ "spawn lifecycle binds a live terminal to it)",
|
||||||
|
members.slots().size(), members.slots().keySet());
|
||||||
|
}
|
||||||
|
|
||||||
|
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
|
||||||
|
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
|
||||||
|
Rendezvous rendezvous = new Rendezvous();
|
||||||
|
// CB-578 stage A: classify a completion-fallback scrape that matches a profile's configured
|
||||||
|
// usage-limit refusal as BACKEND_EXHAUSTED rather than handing it back as a real answer.
|
||||||
|
LiveExhaustedPatterns liveExhaustedPatterns = Fleetd.liveExhaustedPatterns(config);
|
||||||
|
ExhaustedPatternLookup exhaustedPatterns = Fleetd.exhaustedPatternLookup(sessions::roster, liveExhaustedPatterns);
|
||||||
|
// The startup coverage line still reports the boot-time snapshot only.
|
||||||
|
Set<String> exhaustedConfiguredAtStartup = cfg.profiles().entrySet().stream()
|
||||||
|
.filter(e -> e.getValue().hasExhaustedPattern())
|
||||||
|
.map(Map.Entry::getKey)
|
||||||
|
.collect(Collectors.toCollection(LinkedHashSet::new));
|
||||||
|
log.info("backend-exhausted classification (CB-578 stage A): {}",
|
||||||
|
Fleetd.exhaustedPatternCoverageLine(cfg.profiles().keySet(), exhaustedConfiguredAtStartup));
|
||||||
|
// fleetd #201 Unit 5: classify a completion-fallback scrape that matches a profile's
|
||||||
|
// configured backend-error refusal (a credential outage, a provider 5xx) as a backend error
|
||||||
|
// rather than handing it back as a real answer.
|
||||||
|
Map<String, Pattern> errorPatternsByProfile = new LinkedHashMap<>();
|
||||||
|
cfg.profiles().forEach((name, profile) -> {
|
||||||
|
if (profile.hasErrorPattern()) {
|
||||||
|
errorPatternsByProfile.put(name, Pattern.compile(profile.errorPattern()));
|
||||||
|
}
|
||||||
|
});
|
||||||
|
BackendErrorPatternLookup backendErrorPatterns = Fleetd.backendErrorPatternLookup(sessions::roster,
|
||||||
|
errorPatternsByProfile);
|
||||||
|
log.info("backend-error classification (fleetd #201 Unit 5): {}",
|
||||||
|
Fleetd.errorPatternCoverageLine(cfg.profiles().keySet(), errorPatternsByProfile.keySet()));
|
||||||
|
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
|
||||||
|
// CREDENTIAL, so a profile sharing that credential is refused too, not just the one that
|
||||||
|
// happened to report it.
|
||||||
|
Map<String, String> quarantineReasonByCredential = new ConcurrentHashMap<>();
|
||||||
|
// fleetd #175: point the forwarding sink handed to OpenCodeLauncher above at the real one,
|
||||||
|
// now that `sessions` exists to resolve target -> session -> profile.
|
||||||
|
ExhaustionSink exhaustionSink = Fleetd.publishExhaustionSink(exhaustionSinkRef, sessions, config,
|
||||||
|
quarantine, quarantineReasonByCredential, cfg);
|
||||||
|
// fleetd #201 Unit 5: the production BackendErrorSink needs `pushLoop` (built further below,
|
||||||
|
// after `sessions`) to tell a lead about an incident or an unmapped target — the same
|
||||||
|
// construction-order cycle `exhaustionSinkRef` breaks above, broken the same way: a mutable
|
||||||
|
// holder set once `pushLoop` exists, read lazily from inside the lambda built here.
|
||||||
|
AtomicReference<ReplyPushLoop> pushLoopRef = new AtomicReference<>();
|
||||||
|
BackendErrorSink backendErrorSink = Fleetd.backendErrorSink(sessions, () -> config.get().profiles(),
|
||||||
|
outagePolicy, pushLoopRef::get);
|
||||||
|
AgentControl agents = router.memberAgents();
|
||||||
|
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns,
|
||||||
|
exhaustionSink, backendErrorPatterns, backendErrorSink, ports.nanoClock(),
|
||||||
|
Fleetd.worktreeBranchLookup(sessions::roster));
|
||||||
|
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
|
||||||
|
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
|
||||||
|
MemberPresence presence = sessions.asPresence();
|
||||||
|
TurnListener turnListener = Fleetd.turnListener(completion, sessions);
|
||||||
|
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads);
|
||||||
|
// fleetd #556: registration is wired directly to `completion`, not folded into the
|
||||||
|
// `turnListener` fan-out above — so it survives `sessions.onDelivered` (or any future
|
||||||
|
// listener) throwing, regardless of call order.
|
||||||
|
Injector injector = new Injector(router, turnListener, deliverable,
|
||||||
|
presence::forget, Fleetd.turnRegistrar(completion));
|
||||||
|
StatusPoller poller = new StatusPoller(router, injector, Injector.POLL_INTERVAL_MILLIS);
|
||||||
|
poller.start(); // SECOND of the recurring background loops to start, after the reaper.
|
||||||
|
|
||||||
|
// CB-307: reply inbox. A broker: block selects the AMQP-backed durable adapter; absent (or
|
||||||
|
// unusable), fleetd stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
|
||||||
|
// connection, so keep the reference to close it in the ordered shutdown hook.
|
||||||
|
final ReplyInbox replyInbox = Fleetd.selectReplyInbox(cfg.broker(), ports.environment(),
|
||||||
|
ports.replyInboxOpener());
|
||||||
|
// CB-637: this daemon's lead-to-lead mailbox on the SHARED coordination vhost — a separate
|
||||||
|
// broker from the reply inbox by design. Absent a coordinator: block this is null and every
|
||||||
|
// lead path below is simply not wired, exactly the behaviour before this ticket. It owns a
|
||||||
|
// broker connection, so keep the reference for the ordered shutdown hook.
|
||||||
|
final LeadChannelHandle leadMailbox = Fleetd.openLeadMailbox(cfg.coordinator(), ports.environment(),
|
||||||
|
ports.leadMailboxOpener());
|
||||||
|
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
|
||||||
|
String pinnedPrimaryTerminal = cfg.primary() != null ? cfg.primary().terminal() : null;
|
||||||
|
PrimaryRegistry primaryRegistry = new PrimaryRegistry(pinnedPrimaryTerminal);
|
||||||
|
// CB-532: `primary.terminal` is superseded and no longer needed for either of its jobs.
|
||||||
|
if (pinnedPrimaryTerminal != null && !pinnedPrimaryTerminal.isBlank()) {
|
||||||
|
log.warn("primary.terminal is DEPRECATED (CB-532) and can be deleted: identity now comes "
|
||||||
|
+ "from leaders:/leadScan:, and reply nudges follow the lead that delegated. "
|
||||||
|
+ "It still works, and is still the fallback nudge destination when a restart "
|
||||||
|
+ "has lost the delegation map. Its pushReminders/pushBackoffMs stay valid.");
|
||||||
|
}
|
||||||
|
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
|
||||||
|
// open fleet_send. Uses its own lightweight scheduled executor, separate from the injector.
|
||||||
|
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
|
||||||
|
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
|
||||||
|
var pushScheduler = ports.newScheduler("bridge-push-");
|
||||||
|
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
|
||||||
|
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
|
||||||
|
Metrics metrics = FleetMetrics.create(sessions, replyInbox);
|
||||||
|
var pushLoop = new ReplyPushLoop(primaryRegistry, router.leadAgents(), replyInbox,
|
||||||
|
pushScheduler, maxReminders, backoffMs, metrics);
|
||||||
|
// fleetd #201 Unit 5: point the forwarding holder captured by the backendErrorSink lambda
|
||||||
|
// above at the real push loop, now that it exists.
|
||||||
|
pushLoopRef.set(pushLoop);
|
||||||
|
// CB-551: idle-lead heartbeat. Opt-in; absent `leadHeartbeat:` this is never constructed, so
|
||||||
|
// an upgraded daemon cannot silently start spending subscription on nudging an idle lead.
|
||||||
|
// THIRD of the recurring background loops to start (optional).
|
||||||
|
final LeadHeartbeatLoop heartbeat;
|
||||||
|
var heartbeatScheduler = ports.newScheduler("bridge-heartbeat-");
|
||||||
|
if (cfg.leadHeartbeat() != null) {
|
||||||
|
var hb = cfg.leadHeartbeat();
|
||||||
|
var leadContextGauge = new LeadContextGauge();
|
||||||
|
// fleetd #621: the context-high notice's own wording must track this same effective
|
||||||
|
// value — LeadRollover.confirm(...) already gates the roll on it (LeadRollover.java:480),
|
||||||
|
// and absent `leadRollover:` entirely the roll is unusable regardless (NOT_CONFIGURED),
|
||||||
|
// so `true` (the FleetConfig.LeadRollover default) is the safe, byte-identical fallback.
|
||||||
|
// Carried in from Fleetd.main when #612 Unit A merged main: #622 added this line to the
|
||||||
|
// block Unit A had already moved here, so the merge would otherwise have silently
|
||||||
|
// dropped it — with a fully green suite, because nothing pins it (see the follow-up issue).
|
||||||
|
boolean requireOperatorConfirm = cfg.leadRollover() == null || cfg.leadRollover().requireOperatorConfirm();
|
||||||
|
heartbeat = new LeadHeartbeatLoop(primaryRegistry, router.leadAgents(), replyInbox, sessions::roster,
|
||||||
|
pushLoop, heartbeatScheduler, ports.nanoClock(),
|
||||||
|
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
|
||||||
|
metrics,
|
||||||
|
Fleetd.leadContextSource(leadContextGauge, router.leadAgents(), leads,
|
||||||
|
Fleetd.leadConfigDirLookup(() -> config.get().profiles(), leaders)),
|
||||||
|
Boolean.TRUE.equals(hb.contextHighNudge()), requireOperatorConfirm);
|
||||||
|
heartbeat.start();
|
||||||
|
} else {
|
||||||
|
heartbeat = null;
|
||||||
|
heartbeatScheduler.shutdownNow();
|
||||||
|
}
|
||||||
|
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed.
|
||||||
|
LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), config, leads);
|
||||||
|
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
|
||||||
|
pushLoop, metrics);
|
||||||
|
|
||||||
|
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
|
||||||
|
// FOURTH of the recurring background loops to start (optional).
|
||||||
|
final FleetHealthMonitor healthMonitor;
|
||||||
|
var healthScheduler = ports.newScheduler("bridge-health-");
|
||||||
|
if (cfg.health() != null && cfg.health().isEnabled()) {
|
||||||
|
// CB-580: a member found GONE/NEVER_READY must fail whatever ticket is waiting on it.
|
||||||
|
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
|
||||||
|
ports.nanoClock(), ports.wallClockNanos(),
|
||||||
|
cfg.health().intervalOrDefault(),
|
||||||
|
cfg.health().workingSuspectAfterOrDefault(), Fleetd.healthFailTarget(messages));
|
||||||
|
String coverage = FleetHealthMonitor.coverage(true,
|
||||||
|
cfg.health().notifications() != null && cfg.health().notifications().configured());
|
||||||
|
if ("detection-only".equals(coverage)) {
|
||||||
|
log.warn("fleet health: {} (no notification sink configured)", coverage);
|
||||||
|
} else {
|
||||||
|
log.info("fleet health: {}", coverage);
|
||||||
|
}
|
||||||
|
healthMonitor.start();
|
||||||
|
} else {
|
||||||
|
healthMonitor = null;
|
||||||
|
healthScheduler.shutdownNow();
|
||||||
|
}
|
||||||
|
|
||||||
|
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
|
||||||
|
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
|
||||||
|
sessions.onAcquire(replyInbox::own);
|
||||||
|
// CB-516: releasing a worker must fail whatever send was waiting on it.
|
||||||
|
sessions.onRelease(Fleetd.releaseCleanup(messages, replyInbox, primaryRegistry));
|
||||||
|
|
||||||
|
// MCP server face (CB-105): fleet_send/fleet_reply/fleet_status, mounted at /mcp.
|
||||||
|
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
|
||||||
|
ConnectionIdentity identity = new ConnectionIdentity(
|
||||||
|
new PaneLocator(herdr, memberHerdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
|
||||||
|
|
||||||
|
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
|
||||||
|
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
|
||||||
|
final CallerResolver callers;
|
||||||
|
if (cfg.auth().tokenMode()) {
|
||||||
|
String token = ports.environment().get(cfg.auth().tokenEnv());
|
||||||
|
if (token == null || token.isBlank()) {
|
||||||
|
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
|
||||||
|
+ " is unset or empty — export it before starting fleetd");
|
||||||
|
}
|
||||||
|
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
|
||||||
|
log.info("auth: token mode (bearer required for non-worker callers, env {})",
|
||||||
|
cfg.auth().tokenEnv());
|
||||||
|
} else {
|
||||||
|
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
|
||||||
|
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
|
||||||
|
}
|
||||||
|
|
||||||
|
FleetMcp.QuarantineSource quarantineSource = Fleetd.quarantineSource(config, quarantine,
|
||||||
|
liveExhaustedPatterns, quarantineReasonByCredential);
|
||||||
|
FleetMcp.OutageSource outageSource = new FleetMcp.OutageSource(profile -> {
|
||||||
|
var configured = config.get().profiles().get(profile);
|
||||||
|
return configured == null ? null : configured.effectiveCredentialId();
|
||||||
|
}, outagePolicy);
|
||||||
|
|
||||||
|
FleetMcp.LoopHealthSource loopHealth = Fleetd.loopHealthSource(poller, reaper);
|
||||||
|
FleetMcp mcp = new FleetMcp(messages, workers, sessions, identity, presence,
|
||||||
|
primaryRegistry, callers, FleetMcp.AuthorizationMode.ENFORCED, metrics,
|
||||||
|
Fleetd.capacitySource(config, cfg, profile -> liveCountRef.get().apply(profile)),
|
||||||
|
Fleetd.healthCoverageSource(config),
|
||||||
|
loopHealth,
|
||||||
|
quarantineSource,
|
||||||
|
leadMailbox,
|
||||||
|
outageSource,
|
||||||
|
new FleetMcp.LeadSeatSource(Fleetd.leadSeatLookup(() -> config.get().profiles(), leaders, leads)),
|
||||||
|
Fleetd.leadConfigDirSource(() -> config.get().profiles(), leaders),
|
||||||
|
cfg.coordinator() == null ? List.of() : cfg.coordinator().peers(),
|
||||||
|
leadRollover);
|
||||||
|
|
||||||
|
// CB-637: the receive half. Only constructed when a lead mailbox actually opened.
|
||||||
|
final LeadCoordLoop leadCoordLoop;
|
||||||
|
final ScheduledExecutorService leadCoordSchedulerRef;
|
||||||
|
if (leadMailbox != null) {
|
||||||
|
var leadCoordScheduler = ports.newScheduler("bridge-leadcoord-");
|
||||||
|
leadCoordLoop = new LeadCoordLoop(leadMailbox, router.leadAgents(), leads, leadCoordScheduler,
|
||||||
|
LEAD_COORD_INTERVAL_MS);
|
||||||
|
leadCoordLoop.start(); // FIFTH of the recurring background loops to start (optional).
|
||||||
|
leadCoordSchedulerRef = leadCoordScheduler;
|
||||||
|
} else {
|
||||||
|
leadCoordLoop = null;
|
||||||
|
leadCoordSchedulerRef = null;
|
||||||
|
}
|
||||||
|
|
||||||
|
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed.
|
||||||
|
// SIXTH of the recurring background loops to start (optional).
|
||||||
|
final ConfigWatcher configWatcher;
|
||||||
|
if (cfg.configReload() != null && cfg.configReload().isEnabled()) {
|
||||||
|
configWatcher = new ConfigWatcher(config, cfg.configReload().intervalSeconds());
|
||||||
|
configWatcher.start();
|
||||||
|
} else {
|
||||||
|
configWatcher = null;
|
||||||
|
}
|
||||||
|
|
||||||
|
// CB-303 part 3: single ordered shutdown hook — see FleetdRuntime#close() for the statements
|
||||||
|
// this used to be. Registered here, at the exact point `main` used to register it: after
|
||||||
|
// configWatcher, before the Javalin app exists (see this class's own javadoc for why).
|
||||||
|
FleetdRuntime runtime = new FleetdRuntime(cfg, sessions, router, poller, messages, pushLoop, heartbeat,
|
||||||
|
leadCoordLoop, leadCoordSchedulerRef, healthMonitor, configWatcher, mcp, reaper, idleSleepGuard,
|
||||||
|
replyInbox, leadMailbox, completion, injector);
|
||||||
|
ports.addShutdownHook(runtime::close);
|
||||||
|
|
||||||
|
// CB-185: give FleetApp both daemons — /healthz must require both to answer and
|
||||||
|
// GET /sessions must merge across both, or a down/unpolled member daemon is invisible.
|
||||||
|
Javalin app = new FleetApp(herdr, memberHerdr, workers, sessions, messages, presence, mcp.servlet(),
|
||||||
|
callers, metrics, deliverable,
|
||||||
|
() -> MemberCredentialPolicyView.of(config.get().memberCredentials()),
|
||||||
|
quarantineSource, outageSource, loopHealth).build();
|
||||||
|
runtime.attachApp(app);
|
||||||
|
ports.startHttp(app, cfg.bind().host(), cfg.bind().port()); // HTTP starts LAST, always.
|
||||||
|
log.info("fleetd listening on {}:{}, herdr socket {}",
|
||||||
|
cfg.bind().host(), cfg.bind().port(), socket);
|
||||||
|
return runtime;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,166 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigWatcher;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.health.FleetHealthMonitor;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrRouter;
|
||||||
|
import dev.ltms.fleet.inject.CompletionResolver;
|
||||||
|
import dev.ltms.fleet.inject.Injector;
|
||||||
|
import dev.ltms.fleet.inject.StatusPoller;
|
||||||
|
import dev.ltms.fleet.mcp.FleetMcp;
|
||||||
|
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||||
|
import dev.ltms.fleet.msg.LeadCoordLoop;
|
||||||
|
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
|
||||||
|
import dev.ltms.fleet.msg.MessageService;
|
||||||
|
import dev.ltms.fleet.msg.ReplyInbox;
|
||||||
|
import dev.ltms.fleet.msg.ReplyPushLoop;
|
||||||
|
import dev.ltms.fleet.power.IdleSleepGuard;
|
||||||
|
import dev.ltms.fleet.session.SessionManager;
|
||||||
|
import dev.ltms.fleet.session.SessionReaper;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.slf4j.Logger;
|
||||||
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 Unit A: the assembled daemon. {@link FleetdAssembly#assembleAndStart} builds exactly
|
||||||
|
* one of these and hands it to {@code ResourcePorts.addShutdownHook}; {@link #close()} is the
|
||||||
|
* single ordered shutdown, moved verbatim out of {@code Fleetd.main}'s old shutdown-hook
|
||||||
|
* {@code Thread} body — same statements, same order, see that method's javadoc.
|
||||||
|
*
|
||||||
|
* <p><strong>Owns the real production objects, never a copy.</strong> Every package-private
|
||||||
|
* accessor below returns the identical instance the running daemon is using. That is the entire
|
||||||
|
* point of this class existing (see fleetd #612's problem statement): a test that inspected a
|
||||||
|
* snapshot built alongside the real objects could pass while production silently received
|
||||||
|
* something else — the exact shape of the #602/#606 defect this ticket exists to stop from
|
||||||
|
* recurring one call site at a time. Nothing here is rebuilt or copied for a test's benefit.
|
||||||
|
*/
|
||||||
|
final class FleetdRuntime implements AutoCloseable {
|
||||||
|
|
||||||
|
private static final Logger log = LoggerFactory.getLogger(FleetdRuntime.class);
|
||||||
|
|
||||||
|
private final FleetConfig cfg;
|
||||||
|
private final SessionManager sessions;
|
||||||
|
private final HerdrRouter router;
|
||||||
|
private final StatusPoller poller;
|
||||||
|
private final MessageService messages;
|
||||||
|
private final ReplyPushLoop pushLoop;
|
||||||
|
private final LeadHeartbeatLoop heartbeat; // nullable — leadHeartbeat: opt-in
|
||||||
|
private final LeadCoordLoop leadCoordLoop; // nullable — coordinator: opt-in
|
||||||
|
private final ScheduledExecutorService leadCoordScheduler; // nullable, paired with leadCoordLoop
|
||||||
|
private final FleetHealthMonitor healthMonitor; // nullable — health.enabled opt-in
|
||||||
|
private final ConfigWatcher configWatcher; // nullable — configReload.enabled opt-in
|
||||||
|
private final FleetMcp mcp;
|
||||||
|
private final SessionReaper reaper; // nullable — lifecycle.idleTtlSeconds opt-in
|
||||||
|
private final IdleSleepGuard idleSleepGuard; // nullable — idleSleepGuard.enabled: false
|
||||||
|
private final ReplyInbox replyInbox;
|
||||||
|
private final LeadChannelHandle leadMailbox; // nullable — coordinator: opt-in
|
||||||
|
private final CompletionResolver completion;
|
||||||
|
private final Injector injector;
|
||||||
|
/**
|
||||||
|
* Not final: {@code Fleetd.main}'s shutdown hook was registered <em>before</em> the Javalin
|
||||||
|
* {@code FleetApp} was built and started — see {@link FleetdAssembly#assembleAndStart}'s javadoc
|
||||||
|
* for why that order could not be preserved AND have this constructor take {@code app}. {@link
|
||||||
|
* #attachApp} is called immediately after the real app is built, still before HTTP starts
|
||||||
|
* listening, so this is set long before any test or caller could observe it unset.
|
||||||
|
*/
|
||||||
|
private Javalin app;
|
||||||
|
|
||||||
|
FleetdRuntime(FleetConfig cfg, SessionManager sessions, HerdrRouter router, StatusPoller poller,
|
||||||
|
MessageService messages, ReplyPushLoop pushLoop, LeadHeartbeatLoop heartbeat,
|
||||||
|
LeadCoordLoop leadCoordLoop, ScheduledExecutorService leadCoordScheduler,
|
||||||
|
FleetHealthMonitor healthMonitor, ConfigWatcher configWatcher, FleetMcp mcp,
|
||||||
|
SessionReaper reaper, IdleSleepGuard idleSleepGuard, ReplyInbox replyInbox,
|
||||||
|
LeadChannelHandle leadMailbox, CompletionResolver completion, Injector injector) {
|
||||||
|
this.cfg = cfg;
|
||||||
|
this.sessions = sessions;
|
||||||
|
this.router = router;
|
||||||
|
this.poller = poller;
|
||||||
|
this.messages = messages;
|
||||||
|
this.pushLoop = pushLoop;
|
||||||
|
this.heartbeat = heartbeat;
|
||||||
|
this.leadCoordLoop = leadCoordLoop;
|
||||||
|
this.leadCoordScheduler = leadCoordScheduler;
|
||||||
|
this.healthMonitor = healthMonitor;
|
||||||
|
this.configWatcher = configWatcher;
|
||||||
|
this.mcp = mcp;
|
||||||
|
this.reaper = reaper;
|
||||||
|
this.idleSleepGuard = idleSleepGuard;
|
||||||
|
this.replyInbox = replyInbox;
|
||||||
|
this.leadMailbox = leadMailbox;
|
||||||
|
this.completion = completion;
|
||||||
|
this.injector = injector;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** See the {@link #app} field doc for why this is a late-bound setter rather than a constructor arg. */
|
||||||
|
void attachApp(Javalin app) {
|
||||||
|
this.app = app;
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- package-private accessors: the SAME instances this runtime owns, never a copy ---------
|
||||||
|
|
||||||
|
SessionManager sessions() { return sessions; }
|
||||||
|
HerdrRouter router() { return router; }
|
||||||
|
StatusPoller poller() { return poller; }
|
||||||
|
MessageService messages() { return messages; }
|
||||||
|
ReplyPushLoop pushLoop() { return pushLoop; }
|
||||||
|
LeadHeartbeatLoop heartbeat() { return heartbeat; }
|
||||||
|
LeadCoordLoop leadCoordLoop() { return leadCoordLoop; }
|
||||||
|
FleetHealthMonitor healthMonitor() { return healthMonitor; }
|
||||||
|
ConfigWatcher configWatcher() { return configWatcher; }
|
||||||
|
FleetMcp mcp() { return mcp; }
|
||||||
|
SessionReaper reaper() { return reaper; }
|
||||||
|
IdleSleepGuard idleSleepGuard() { return idleSleepGuard; }
|
||||||
|
ReplyInbox replyInbox() { return replyInbox; }
|
||||||
|
LeadChannelHandle leadMailbox() { return leadMailbox; }
|
||||||
|
CompletionResolver completion() { return completion; }
|
||||||
|
Injector injector() { return injector; }
|
||||||
|
Javalin app() { return app; }
|
||||||
|
|
||||||
|
/**
|
||||||
|
* CB-303 part 3: the single ordered shutdown. Moved verbatim out of {@code Fleetd.main}'s
|
||||||
|
* shutdown-hook {@code Thread} body (fleetd #612 Unit A) — drain sessions first while herdr is
|
||||||
|
* still open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close
|
||||||
|
* herdr last, exactly as before. {@code Fleetd.main} never calls this directly; it hands the
|
||||||
|
* reference to {@code ResourcePorts.addShutdownHook} the moment this runtime exists, the same
|
||||||
|
* point it used to register the hook {@code Thread} itself.
|
||||||
|
*/
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
|
||||||
|
poller.stop();
|
||||||
|
messages.close();
|
||||||
|
pushLoop.close();
|
||||||
|
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
|
||||||
|
if (leadCoordLoop != null) leadCoordLoop.close(); // CB-637: stop delivering peer-lead messages
|
||||||
|
if (leadCoordScheduler != null) leadCoordScheduler.shutdownNow();
|
||||||
|
if (healthMonitor != null) healthMonitor.stop();
|
||||||
|
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
|
||||||
|
mcp.close();
|
||||||
|
if (reaper != null) reaper.stop();
|
||||||
|
// Idle-sleep guard: release unconditionally, even though sessions.close() above already
|
||||||
|
// drained every session (and each release already drove the live count to 0, which
|
||||||
|
// releases the guard's assertion on its own) — this is the backstop for a drain that was
|
||||||
|
// itself interrupted or threw, so no caffeinate child ever outlives the daemon.
|
||||||
|
if (idleSleepGuard != null) idleSleepGuard.close();
|
||||||
|
// Release the broker connection last among message resources (no-op for the in-memory inbox).
|
||||||
|
if (replyInbox instanceof AutoCloseable closeable) {
|
||||||
|
try {
|
||||||
|
closeable.close();
|
||||||
|
} catch (Exception e) {
|
||||||
|
log.debug("reply inbox close: {}", e.toString());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
// CB-637: the coordination connection goes with it — after the loop that reads it has
|
||||||
|
// stopped, so no tick can be mid-ack against a closed channel.
|
||||||
|
if (leadMailbox != null) {
|
||||||
|
try {
|
||||||
|
leadMailbox.close();
|
||||||
|
} catch (Exception e) {
|
||||||
|
log.debug("lead mailbox close: {}", e.toString());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
router.close();
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,80 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 Unit A: every boot-time side effect {@link FleetdAssembly#assembleAndStart} performs
|
||||||
|
* that a real daemon must do for real, and a test must not — read the process environment, connect
|
||||||
|
* a herdr client, open a broker (the reply inbox, the lead mailbox), read a clock, start a
|
||||||
|
* background scheduler, register the JVM shutdown hook, and bind the HTTP server.
|
||||||
|
*
|
||||||
|
* <p>{@link #system()} is the one production implementation ({@code SystemResourcePorts}), wired
|
||||||
|
* verbatim from what {@code Fleetd.main} used to call directly at each of these call sites. A test
|
||||||
|
* builds its own implementation instead of receiving an inert default from this interface —
|
||||||
|
* deliberately, there is no {@code ResourcePorts.none()}. fleetd #612's whole problem is a call
|
||||||
|
* site quietly swapped for an inert variant that still compiles; adding one here, even for tests,
|
||||||
|
* would hand a future edit to {@code FleetdAssembly} the exact compiling substitute this ticket
|
||||||
|
* exists to rule out. A test that wants an inert resource writes its own fake and owns that
|
||||||
|
* decision explicitly.
|
||||||
|
*/
|
||||||
|
public interface ResourcePorts {
|
||||||
|
|
||||||
|
/** The process environment. Production: {@link System#getenv()}. */
|
||||||
|
Map<String, String> environment();
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Connect a herdr client bound to {@code socketPath}. Production returns a real
|
||||||
|
* {@code UnixSocketHerdrClient} — connection-per-call, so this itself never touches the socket.
|
||||||
|
*/
|
||||||
|
HerdrClient connectHerdr(Path socketPath);
|
||||||
|
|
||||||
|
/** The reply-inbox AMQP opener (CB-307). Production: {@link Fleetd#replyInboxOpener()}. */
|
||||||
|
Fleetd.AmqpOpener replyInboxOpener();
|
||||||
|
|
||||||
|
/** The lead-mailbox AMQP opener (CB-637). Production: {@link Fleetd#leadMailboxOpener()}. */
|
||||||
|
Fleetd.LeadMailboxOpener leadMailboxOpener();
|
||||||
|
|
||||||
|
/** A monotonic elapsed-time clock. Production: {@link System#nanoTime()}. */
|
||||||
|
LongSupplier nanoClock();
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #629: the per-poll wait {@code FleetdAssembly#assembleAndStart} passes to {@code
|
||||||
|
* Fleetd#awaitHerdr} while polling for herdr's socket. Production: {@link
|
||||||
|
* Fleetd#sleepHerdrPoll()} — a real {@code Thread.sleep}. {@link #nanoClock()} alone is not
|
||||||
|
* enough to make {@code awaitHerdr}'s deadline controllable: the old call site passed {@code
|
||||||
|
* Fleetd::sleepHerdrPoll} directly, hardcoded, so a test that injected a fake clock still had
|
||||||
|
* to wait out the real sleep between each poll to ever reach the deadline — the clock looked
|
||||||
|
* injected and was not actually controllable. A test supplies a no-op that advances its own
|
||||||
|
* injected {@link #nanoClock()} instead, so the deadline becomes reachable without any real
|
||||||
|
* wall-clock time passing.
|
||||||
|
*/
|
||||||
|
Runnable herdrPollWait();
|
||||||
|
|
||||||
|
/**
|
||||||
|
* A wall-clock reading, in nanoseconds. Production: {@code System.currentTimeMillis()}
|
||||||
|
* converted to nanoseconds. Kept separate from {@link #nanoClock()} because {@link
|
||||||
|
* dev.ltms.fleet.health.FleetHealthMonitor} needs both — one monotonic clock for elapsed-time
|
||||||
|
* decisions, one wall clock to detect and correct for a macOS sleep freezing the monotonic one.
|
||||||
|
*/
|
||||||
|
LongSupplier wallClockNanos();
|
||||||
|
|
||||||
|
/** A dedicated single-thread scheduler; production names its (virtual) thread {@code purpose}. */
|
||||||
|
ScheduledExecutorService newScheduler(String purpose);
|
||||||
|
|
||||||
|
/** Register a JVM shutdown hook that runs {@code hook} on JVM exit. */
|
||||||
|
void addShutdownHook(Runnable hook);
|
||||||
|
|
||||||
|
/** Bind and start the HTTP server. */
|
||||||
|
void startHttp(Javalin app, String host, int port);
|
||||||
|
|
||||||
|
/** The real production ports: a live herdr socket, a live broker, real threads, a real bind. */
|
||||||
|
static ResourcePorts system() {
|
||||||
|
return new SystemResourcePorts();
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,71 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.herdr.UnixSocketHerdrClient;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 Unit A: the one production {@link ResourcePorts} — every method here is the exact
|
||||||
|
* call {@code Fleetd.main} used to make directly at each of these sites before this ticket.
|
||||||
|
* Package-private: obtained only through {@link ResourcePorts#system()}.
|
||||||
|
*/
|
||||||
|
final class SystemResourcePorts implements ResourcePorts {
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return System.getenv();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
return UnixSocketHerdrClient.connect(socketPath, new ObjectMapper());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
return Fleetd.replyInboxOpener();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
return Fleetd.leadMailboxOpener();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
return Fleetd::sleepHerdrPoll;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return () -> TimeUnit.MILLISECONDS.toNanos(System.currentTimeMillis());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
return Executors.newSingleThreadScheduledExecutor(r -> Thread.ofVirtual().name(purpose).unstarted(r));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
Runtime.getRuntime().addShutdownHook(new Thread(hook));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
app.start(host, port);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -221,7 +221,8 @@ public final class CallerResolver {
|
|||||||
// The config/live binding names this pane as an architect slot's own. Same
|
// The config/live binding names this pane as an architect slot's own. Same
|
||||||
// unforgeable pane mapping; the live binding, never a request argument, decides.
|
// unforgeable pane mapping; the live binding, never a request argument, decides.
|
||||||
// Check the slot role too: this defence in depth prevents a bad lifecycle bind from
|
// Check the slot role too: this defence in depth prevents a bad lifecycle bind from
|
||||||
// escalating a dev or reviewer into an architect. Checked before the worker fallback.
|
// escalating a dev, hunter or reviewer into an architect. Checked before
|
||||||
|
// the worker fallback.
|
||||||
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
|
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
|
||||||
}
|
}
|
||||||
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
|
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
|
||||||
|
|||||||
@@ -42,7 +42,8 @@ public interface MemberLifecycle {
|
|||||||
* Try to bind a newly spawned {@code terminal} into the role it was granted.
|
* Try to bind a newly spawned {@code terminal} into the role it was granted.
|
||||||
*
|
*
|
||||||
* @return the role this session actually holds: {@code role} unchanged for a role with no
|
* @return the role this session actually holds: {@code role} unchanged for a role with no
|
||||||
* live slot-binding semantics (dev, reviewer), or when the bind succeeded; a fallback
|
* live slot-binding semantics (dev, hunter, reviewer), or when the bind
|
||||||
|
* succeeded; a fallback
|
||||||
* role — never {@code role} — when a slot-bound role (architect) could not be bound.
|
* role — never {@code role} — when a slot-bound role (architect) could not be bound.
|
||||||
* Callers must record THIS value on the session, never the requested {@code role}, so
|
* Callers must record THIS value on the session, never the requested {@code role}, so
|
||||||
* a later roster read never reports a role the session does not hold (CB-619). In
|
* a later roster read never reports a role the session does not hold (CB-619). In
|
||||||
|
|||||||
@@ -20,7 +20,8 @@ import java.util.function.Supplier;
|
|||||||
*
|
*
|
||||||
* <p>Two halves, split by who owns each:
|
* <p>Two halves, split by who owns each:
|
||||||
* <ul>
|
* <ul>
|
||||||
* <li><b>slots</b> — read from {@code fleet.architects}/{@code developers}/{@code reviewers}
|
* <li><b>slots</b> — read from {@code fleet.architects}/{@code developers}/
|
||||||
|
* {@code hunters}/{@code reviewers}
|
||||||
* (see {@link #slots()}), each carrying the {@code profile} reference the spawn lifecycle
|
* (see {@link #slots()}), each carrying the {@code profile} reference the spawn lifecycle
|
||||||
* reads when it stands the slot up. <strong>Live, since fleetd #424</strong>: {@link #live}
|
* reads when it stands the slot up. <strong>Live, since fleetd #424</strong>: {@link #live}
|
||||||
* re-reads {@code fleet:} on every call, through a supplier the same shape as
|
* re-reads {@code fleet:} on every call, through a supplier the same shape as
|
||||||
@@ -322,7 +323,7 @@ public final class MemberRegistry implements MemberLifecycle {
|
|||||||
* CB-619 / fleetd #123: refuse an architect acquire before anything spawns when no configured
|
* CB-619 / fleetd #123: refuse an architect acquire before anything spawns when no configured
|
||||||
* slot carries {@code profile} — the config-gap case from the original defect report (a spawn
|
* slot carries {@code profile} — the config-gap case from the original defect report (a spawn
|
||||||
* asked for {@code role=architect, profile=sonnet}, and {@code fleet.architects} carried only
|
* asked for {@code role=architect, profile=sonnet}, and {@code fleet.architects} carried only
|
||||||
* {@code opus} and {@code sol}). A dev/reviewer acquire is always a no-op: those pools are
|
* {@code opus} and {@code sol}). A dev/hunter/reviewer acquire is always a no-op: those pools are
|
||||||
* placement candidates only (see {@code CompositePeerLauncher}), never a live identity binding,
|
* placement candidates only (see {@code CompositePeerLauncher}), never a live identity binding,
|
||||||
* so there is nothing here to refuse — an explicit profile outside the pool for those roles is a
|
* so there is nothing here to refuse — an explicit profile outside the pool for those roles is a
|
||||||
* documented operator override, not a defect.
|
* documented operator override, not a defect.
|
||||||
|
|||||||
@@ -31,7 +31,7 @@ import java.util.function.Supplier;
|
|||||||
* {@code placement:}, and an existing profile's {@code weight} / {@code maxLoad}. Both are
|
* {@code placement:}, and an existing profile's {@code weight} / {@code maxLoad}. Both are
|
||||||
* read through a supplier on {@code CompositePeerLauncher}, which is what makes them hot —
|
* read through a supplier on {@code CompositePeerLauncher}, which is what makes them hot —
|
||||||
* not the fact that they are config. Most of {@code fleet:} — every role pool
|
* not the fact that they are config. Most of {@code fleet:} — every role pool
|
||||||
* ({@code architects}/{@code developers}/{@code reviewers}), {@code charters}, and
|
* ({@code architects}/{@code developers}/{@code hunters}/{@code reviewers}), {@code charters}, and
|
||||||
* {@code tabLabel} — is read the same live way, through the same supplier
|
* {@code tabLabel} — is read the same live way, through the same supplier
|
||||||
* ({@code () -> config.get().fleet()}). {@code architects} in particular is hot for
|
* ({@code () -> config.get().fleet()}). {@code architects} in particular is hot for
|
||||||
* <strong>two independent consumers</strong> (fleetd #424): {@code CompositePeerLauncher}
|
* <strong>two independent consumers</strong> (fleetd #424): {@code CompositePeerLauncher}
|
||||||
@@ -605,7 +605,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
|||||||
+ "opened once and needs a restart; the broker URI env-var name kept out of a "
|
+ "opened once and needs a restart; the broker URI env-var name kept out of a "
|
||||||
+ "member's environment is read live on every spawn and already applied");
|
+ "member's environment is read live on every spawn and already applied");
|
||||||
}
|
}
|
||||||
// fleetd #333: unlike health/coordinator above, most of `fleet:` (developers, reviewers,
|
// fleetd #333: unlike health/coordinator above, most of `fleet:` (developers, hunters, reviewers,
|
||||||
// charters, tabLabel) is genuinely hot — ConfigRefTest.aHotChangeIsAppliedAndRead-
|
// charters, tabLabel) is genuinely hot — ConfigRefTest.aHotChangeIsAppliedAndRead-
|
||||||
// ThroughGet and aCharterChangeIsHotAndReachesTheLiveConfig prove it reaches the live config
|
// ThroughGet and aCharterChangeIsHotAndReachesTheLiveConfig prove it reaches the live config
|
||||||
// with no restart note. `architects` is hot too, and — since fleetd #424 — hot for BOTH of
|
// with no restart note. `architects` is hot too, and — since fleetd #424 — hot for BOTH of
|
||||||
@@ -630,7 +630,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
|||||||
+ "identity map and to auto-launch leads, and neither is rebuilt on reload, so a "
|
+ "identity map and to auto-launch leads, and neither is rebuilt on reload, so a "
|
||||||
+ "lead added, removed, or given a new tab: label needs a restart — until then it "
|
+ "lead added, removed, or given a new tab: label needs a restart — until then it "
|
||||||
+ "stays unrecognised, and a caller from its new tab resolves as a worker, not a "
|
+ "stays unrecognised, and a caller from its new tab resolves as a worker, not a "
|
||||||
+ "lead; the rest of fleet: (developers, reviewers, charters, tabLabel) is read "
|
+ "lead; the rest of fleet: (developers, hunters, reviewers, charters, tabLabel) is read "
|
||||||
+ "live through the supplier on CompositePeerLauncher, and architects is read "
|
+ "live through the supplier on CompositePeerLauncher, and architects is read "
|
||||||
+ "live through that same supplier for placement AND through a separate supplier "
|
+ "live through that same supplier for placement AND through a separate supplier "
|
||||||
+ "on MemberRegistry for spawn-time identity — both already applied");
|
+ "on MemberRegistry for spawn-time identity — both already applied");
|
||||||
|
|||||||
@@ -62,7 +62,7 @@ import java.util.regex.PatternSyntaxException;
|
|||||||
* @param fleet who the daemon may run and under which role (CB-557). One block replacing
|
* @param fleet who the daemon may run and under which role (CB-557). One block replacing
|
||||||
* the former {@code leaders:}, {@code members:}, {@code leadScan:} and
|
* the former {@code leaders:}, {@code members:}, {@code leadScan:} and
|
||||||
* {@code defaultProfile:}. Role is the containing key — {@code leaders},
|
* {@code defaultProfile:}. Role is the containing key — {@code leaders},
|
||||||
* {@code architects}, {@code developers}, {@code reviewers} — and each entry
|
* {@code architects}, {@code developers}, {@code hunters}, {@code reviewers} — and each entry
|
||||||
* names the {@code profiles:} backend it runs on. See {@link Fleet}
|
* names the {@code profiles:} backend it runs on. See {@link Fleet}
|
||||||
* @param leadHeartbeat opt-in idle-lead heartbeat (CB-551); {@code null} ⇒ off, and an upgraded
|
* @param leadHeartbeat opt-in idle-lead heartbeat (CB-551); {@code null} ⇒ off, and an upgraded
|
||||||
* daemon never nudges an idle lead on its own initiative
|
* daemon never nudges an idle lead on its own initiative
|
||||||
@@ -462,8 +462,8 @@ public record FleetConfig(
|
|||||||
* profile that does not opt in. Read live off the current config, so it is
|
* profile that does not opt in. Read live off the current config, so it is
|
||||||
* HOT: a change takes effect on the next exhaustion classification / spawn,
|
* HOT: a change takes effect on the next exhaustion classification / spawn,
|
||||||
* no restart needed.
|
* no restart needed.
|
||||||
* @param autoCompactWindow opt-in per-profile token window that forces a spawned member to
|
* @param autoCompactWindow opt-in per-profile token window that forces a launched Claude Code
|
||||||
* auto-compact its context at (Claude Code) or within (opencode) a bound the
|
* session to auto-compact its context at, or an opencode session within, a bound the
|
||||||
* operator chooses, instead of the backend's own default. {@code null} (the
|
* operator chooses, instead of the backend's own default. {@code null} (the
|
||||||
* default) leaves today's behaviour exactly — opencode already forces
|
* default) leaves today's behaviour exactly — opencode already forces
|
||||||
* {@code compaction.auto: true} unconditionally (CB-523) but has no absolute
|
* {@code compaction.auto: true} unconditionally (CB-523) but has no absolute
|
||||||
@@ -1200,6 +1200,7 @@ public record FleetConfig(
|
|||||||
* @param leaders panes that orchestrate rather than are orchestrated, keyed by lead name
|
* @param leaders panes that orchestrate rather than are orchestrated, keyed by lead name
|
||||||
* @param architects profiles the {@code architect} role may run on
|
* @param architects profiles the {@code architect} role may run on
|
||||||
* @param developers profiles the {@code dev} role may run on
|
* @param developers profiles the {@code dev} role may run on
|
||||||
|
* @param hunters profiles the {@code hunter} role may run on
|
||||||
* @param reviewers profiles the {@code reviewer} role may run on
|
* @param reviewers profiles the {@code reviewer} role may run on
|
||||||
* @param charters optional launch-charter text keyed by singular role wire name
|
* @param charters optional launch-charter text keyed by singular role wire name
|
||||||
* @param tabLabel template for a member tab's label; {@code {role}}, {@code {profile}},
|
* @param tabLabel template for a member tab's label; {@code {role}}, {@code {profile}},
|
||||||
@@ -1210,6 +1211,7 @@ public record FleetConfig(
|
|||||||
public record Fleet(Map<String, Leader> leaders,
|
public record Fleet(Map<String, Leader> leaders,
|
||||||
Map<String, Slot> architects,
|
Map<String, Slot> architects,
|
||||||
Map<String, Slot> developers,
|
Map<String, Slot> developers,
|
||||||
|
Map<String, Slot> hunters,
|
||||||
Map<String, Slot> reviewers,
|
Map<String, Slot> reviewers,
|
||||||
Map<String, String> charters,
|
Map<String, String> charters,
|
||||||
String tabLabel) {
|
String tabLabel) {
|
||||||
@@ -1226,6 +1228,7 @@ public record FleetConfig(
|
|||||||
leaders = unmodifiableOrEmpty(leaders);
|
leaders = unmodifiableOrEmpty(leaders);
|
||||||
architects = unmodifiableOrEmpty(architects);
|
architects = unmodifiableOrEmpty(architects);
|
||||||
developers = unmodifiableOrEmpty(developers);
|
developers = unmodifiableOrEmpty(developers);
|
||||||
|
hunters = unmodifiableOrEmpty(hunters);
|
||||||
reviewers = unmodifiableOrEmpty(reviewers);
|
reviewers = unmodifiableOrEmpty(reviewers);
|
||||||
charters = unmodifiableOrEmpty(charters);
|
charters = unmodifiableOrEmpty(charters);
|
||||||
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? DEFAULT_TAB_LABEL : tabLabel;
|
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? DEFAULT_TAB_LABEL : tabLabel;
|
||||||
@@ -1240,9 +1243,15 @@ public record FleetConfig(
|
|||||||
* constructor: the launcher reads {@code fleet.charters()} from the live config. Jackson
|
* constructor: the launcher reads {@code fleet.charters()} from the live config. Jackson
|
||||||
* binds the canonical constructor, so this one cannot swallow an operator's YAML.
|
* binds the canonical constructor, so this one cannot swallow an operator's YAML.
|
||||||
*/
|
*/
|
||||||
|
public Fleet(Map<String, Leader> leaders, Map<String, Slot> architects,
|
||||||
|
Map<String, Slot> developers, Map<String, Slot> reviewers,
|
||||||
|
Map<String, String> charters, String tabLabel) {
|
||||||
|
this(leaders, architects, developers, null, reviewers, charters, tabLabel);
|
||||||
|
}
|
||||||
|
|
||||||
public Fleet(Map<String, Leader> leaders, Map<String, Slot> architects,
|
public Fleet(Map<String, Leader> leaders, Map<String, Slot> architects,
|
||||||
Map<String, Slot> developers, Map<String, Slot> reviewers, String tabLabel) {
|
Map<String, Slot> developers, Map<String, Slot> reviewers, String tabLabel) {
|
||||||
this(leaders, architects, developers, reviewers, null, tabLabel);
|
this(leaders, architects, developers, null, reviewers, null, tabLabel);
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -1264,6 +1273,7 @@ public record FleetConfig(
|
|||||||
return switch (role) {
|
return switch (role) {
|
||||||
case ARCHITECT -> architects;
|
case ARCHITECT -> architects;
|
||||||
case DEV -> developers;
|
case DEV -> developers;
|
||||||
|
case HUNTER -> hunters;
|
||||||
case REVIEWER -> reviewers;
|
case REVIEWER -> reviewers;
|
||||||
};
|
};
|
||||||
}
|
}
|
||||||
@@ -1319,9 +1329,20 @@ public record FleetConfig(
|
|||||||
* each such nudge costs the lead a turn just to read "nothing pending";
|
* each such nudge costs the lead a turn just to read "nothing pending";
|
||||||
* three is enough to tell it it may stand down without nagging forever, and
|
* three is enough to tell it it may stand down without nagging forever, and
|
||||||
* it is the bound that stops an idle fleet from being a subscription burner.
|
* it is the bound that stops an idle fleet from being a subscription burner.
|
||||||
|
* @param contextHighNudge fleetd #609: when {@code true}, an idle lead whose own {@code
|
||||||
|
* LeadContextGauge} reading is {@code HIGH} gets a text notice telling it
|
||||||
|
* to consider a handover, appended to whatever heartbeat nudge the loop
|
||||||
|
* already sends. Default {@code false} ({@code null} also means off) — an
|
||||||
|
* upgraded daemon must not silently start telling leads to hand over.
|
||||||
*/
|
*/
|
||||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||||
public record LeadHeartbeat(Integer idleAfterSeconds, Long backoffMs, Integer quietNudgeCap) {
|
public record LeadHeartbeat(Integer idleAfterSeconds, Long backoffMs, Integer quietNudgeCap,
|
||||||
|
Boolean contextHighNudge) {
|
||||||
|
/** Convenience constructor for every call site that predates fleetd #609: no context notice. */
|
||||||
|
public LeadHeartbeat(Integer idleAfterSeconds, Long backoffMs, Integer quietNudgeCap) {
|
||||||
|
this(idleAfterSeconds, backoffMs, quietNudgeCap, null);
|
||||||
|
}
|
||||||
|
|
||||||
public LeadHeartbeat {
|
public LeadHeartbeat {
|
||||||
idleAfterSeconds = (idleAfterSeconds == null || idleAfterSeconds <= 0) ? 300 : idleAfterSeconds;
|
idleAfterSeconds = (idleAfterSeconds == null || idleAfterSeconds <= 0) ? 300 : idleAfterSeconds;
|
||||||
backoffMs = (backoffMs == null || backoffMs <= 0) ? 60_000L : backoffMs;
|
backoffMs = (backoffMs == null || backoffMs <= 0) ? 60_000L : backoffMs;
|
||||||
@@ -1854,6 +1875,7 @@ public record FleetConfig(
|
|||||||
rejectDuplicateMemberSlots(yaml);
|
rejectDuplicateMemberSlots(yaml);
|
||||||
rejectNegativeMaxLoad(yaml);
|
rejectNegativeMaxLoad(yaml);
|
||||||
rejectAutoCompactWindowOutOfRange(yaml);
|
rejectAutoCompactWindowOutOfRange(yaml);
|
||||||
|
warnConflictingAutoCompactWindows(yaml);
|
||||||
rejectMalformedProfilePatterns(yaml);
|
rejectMalformedProfilePatterns(yaml);
|
||||||
rejectUnknownKind(yaml);
|
rejectUnknownKind(yaml);
|
||||||
rejectUnknownAuthMode(yaml);
|
rejectUnknownAuthMode(yaml);
|
||||||
@@ -1872,7 +1894,7 @@ public record FleetConfig(
|
|||||||
|
|
||||||
/** The {@code fleet:} child blocks whose direct children are slot names. */
|
/** The {@code fleet:} child blocks whose direct children are slot names. */
|
||||||
private static final Set<String> FLEET_POOL_KEYS =
|
private static final Set<String> FLEET_POOL_KEYS =
|
||||||
Set.of("leaders", "architects", "developers", "reviewers");
|
Set.of("leaders", "architects", "developers", "hunters", "reviewers");
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Reject a {@code fleet:} role pool whose slot names repeat (CB-548, re-homed by CB-557).
|
* Reject a {@code fleet:} role pool whose slot names repeat (CB-548, re-homed by CB-557).
|
||||||
@@ -1882,7 +1904,7 @@ public record FleetConfig(
|
|||||||
* daemon would never know. Jackson's YAML parser does not fail on duplicate mapping keys by
|
* daemon would never know. Jackson's YAML parser does not fail on duplicate mapping keys by
|
||||||
* default, so duplicates are caught here, at parse time, before the map is built.
|
* default, so duplicates are caught here, at parse time, before the map is built.
|
||||||
*
|
*
|
||||||
* <p>Only the four pools <em>directly under the top-level {@code fleet:}</em> are considered,
|
* <p>Only the five pools <em>directly under the top-level {@code fleet:}</em> are considered,
|
||||||
* and only their direct child keys (the slot names). A nested field elsewhere, even one also
|
* and only their direct child keys (the slot names). A nested field elsewhere, even one also
|
||||||
* named {@code developers:}, is ignored, so parsing of the rest of the config is unaffected.
|
* named {@code developers:}, is ignored, so parsing of the rest of the config is unaffected.
|
||||||
*
|
*
|
||||||
@@ -2049,8 +2071,9 @@ public record FleetConfig(
|
|||||||
"defaultProfile", "a role pool under 'fleet:' — an unqualified spawn now names a role,"
|
"defaultProfile", "a role pool under 'fleet:' — an unqualified spawn now names a role,"
|
||||||
+ " and that role's pool supplies the candidate profiles",
|
+ " and that role's pool supplies the candidate profiles",
|
||||||
"architects", "'fleet.architects'",
|
"architects", "'fleet.architects'",
|
||||||
"members", "a role pool under 'fleet:' — 'fleet.architects', 'fleet.developers' or"
|
"members", "a role pool under 'fleet:' — 'fleet.architects', 'fleet.developers',"
|
||||||
+ " 'fleet.reviewers'; the role is the containing key, not a 'role:' field",
|
+ " 'fleet.hunters' or 'fleet.reviewers'; the role is the containing key, not"
|
||||||
|
+ " a 'role:' field",
|
||||||
"leaders", "'fleet.leaders'",
|
"leaders", "'fleet.leaders'",
|
||||||
"leadScan", "'fleet.leaders.<name>.tabPrefix' and '.scanIntervalSeconds' — lead"
|
"leadScan", "'fleet.leaders.<name>.tabPrefix' and '.scanIntervalSeconds' — lead"
|
||||||
+ " discovery is now configured on the lead it discovers");
|
+ " discovery is now configured on the lead it discovers");
|
||||||
@@ -2203,6 +2226,73 @@ public record FleetConfig(
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Warn (never refuse to start) about a Claude Code profile whose auto-compaction flag and
|
||||||
|
* environment setting disagree.
|
||||||
|
*
|
||||||
|
* <p>Renamed from {@code rejectConflictingAutoCompactWindows} (fleetd #601 review, measured
|
||||||
|
* 2026-09-22): that method threw {@link IllegalStateException}, so {@link #load(Path)} refused
|
||||||
|
* to start on a config carrying this conflict. On this host, four profiles trip it, including
|
||||||
|
* the lead's own profile and the one every worker spawns on — so the throw is not a rare edge
|
||||||
|
* case. Under launchd, a throw inside {@code load()} is a restart loop, not an error an operator
|
||||||
|
* reads once, and the config that would fix it ({@code fleetd/fleetd.yaml}) is gitignored, so
|
||||||
|
* the cause is invisible on the host where it bites. A WARN gives the operator the same
|
||||||
|
* information — which profiles, and now both values, so they can fix it without reading the
|
||||||
|
* source — without ever taking the fleet down.
|
||||||
|
*
|
||||||
|
* <p>fleetd #618 measured which of the two inputs Claude Code actually follows when they
|
||||||
|
* disagree: the environment variable wins, so {@code autoCompactWindow} is inert on a profile
|
||||||
|
* that also sets the env var. This method only detects and reports the disagreement — it does
|
||||||
|
* not correct it — see {@link dev.ltms.fleet.launch.ClaudeCodeArguments} for the full measured
|
||||||
|
* precedence.
|
||||||
|
*
|
||||||
|
* <p>Equal values never warn: either input then produces the same session window, so there is
|
||||||
|
* nothing to reconcile.
|
||||||
|
*/
|
||||||
|
static void warnConflictingAutoCompactWindows(String yaml) {
|
||||||
|
Map<?, ?> raw;
|
||||||
|
try {
|
||||||
|
raw = YAML.readValue(yaml, Map.class);
|
||||||
|
} catch (IOException | IllegalArgumentException e) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (raw == null || !(raw.get("profiles") instanceof Map<?, ?> profiles)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
List<String> names = new ArrayList<>();
|
||||||
|
List<String> detail = new ArrayList<>();
|
||||||
|
for (Map.Entry<?, ?> entry : profiles.entrySet()) {
|
||||||
|
if (!(entry.getValue() instanceof Map<?, ?> profile)
|
||||||
|
|| !(profile.get("autoCompactWindow") instanceof Number window)
|
||||||
|
|| !(profile.get("env") instanceof Map<?, ?> env)
|
||||||
|
|| !env.containsKey("CLAUDE_CODE_AUTO_COMPACT_WINDOW")) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
Object kind = profile.get("kind");
|
||||||
|
boolean claudeCode = kind == null || String.valueOf(kind).isBlank()
|
||||||
|
|| Profile.KIND_CLAUDE_CODE.equalsIgnoreCase(String.valueOf(kind));
|
||||||
|
Object envValue = env.get("CLAUDE_CODE_AUTO_COMPACT_WINDOW");
|
||||||
|
if (claudeCode && !String.valueOf(window).equals(String.valueOf(envValue))) {
|
||||||
|
String name = String.valueOf(entry.getKey());
|
||||||
|
names.add(name);
|
||||||
|
detail.add(name + " (autoCompactWindow=" + window
|
||||||
|
+ ", env.CLAUDE_CODE_AUTO_COMPACT_WINDOW=" + envValue + ")");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (names.isEmpty()) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
names.sort(String::compareTo);
|
||||||
|
detail.sort(String::compareTo);
|
||||||
|
log.warn("Claude Code profile(s) {} set disagreeing autoCompactWindow and env."
|
||||||
|
+ "CLAUDE_CODE_AUTO_COMPACT_WINDOW — the daemon starts anyway: {}. fleetd "
|
||||||
|
+ "#618 measured that CLAUDE_CODE_AUTO_COMPACT_WINDOW wins, so "
|
||||||
|
+ "autoCompactWindow is inert on these profiles. Set equal values on each "
|
||||||
|
+ "to resolve this — do not just delete the env var, since that LOWERS the "
|
||||||
|
+ "live window to autoCompactWindow's value rather than fixing anything.",
|
||||||
|
names, String.join(", ", detail));
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Reject a profile whose {@code errorPattern} (fleetd #201 Unit 5) or {@code exhaustedPattern}
|
* Reject a profile whose {@code errorPattern} (fleetd #201 Unit 5) or {@code exhaustedPattern}
|
||||||
* (CB-578 stage A) is not a valid Java regex, naming the profile, the key, and the parser's own
|
* (CB-578 stage A) is not a valid Java regex, naming the profile, the key, and the parser's own
|
||||||
|
|||||||
@@ -0,0 +1,37 @@
|
|||||||
|
package dev.ltms.fleet.launch;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
|
||||||
|
import java.util.ArrayList;
|
||||||
|
import java.util.List;
|
||||||
|
|
||||||
|
/** Arguments shared by every fleetd path that starts Claude Code. */
|
||||||
|
public final class ClaudeCodeArguments {
|
||||||
|
|
||||||
|
private ClaudeCodeArguments() {
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Append the configured Claude Code auto-compaction window when the profile opts in.
|
||||||
|
*
|
||||||
|
* <p>This flag and the environment variable {@code CLAUDE_CODE_AUTO_COMPACT_WINDOW} can
|
||||||
|
* disagree, and fleetd #618 measured which one Claude Code actually follows: the environment
|
||||||
|
* variable wins, ahead of this {@code --autocompact} flag, ahead of the settings file, ahead of
|
||||||
|
* clientdata, the experiment, and the model default. So when a profile sets both, the flag this
|
||||||
|
* method appends has NO effect — Claude Code reads {@code CLAUDE_CODE_AUTO_COMPACT_WINDOW}
|
||||||
|
* first and never consults the flag. {@link FleetConfig#load(java.nio.file.Path)} only WARNS
|
||||||
|
* when a Claude Code profile sets both to different values (see {@code
|
||||||
|
* FleetConfig.warnConflictingAutoCompactWindows}) — it does not stop the daemon from starting,
|
||||||
|
* and the launched session honours the env var, not this flag. Measured against Claude Code
|
||||||
|
* 2.1.278 (fleetd #618) — a later version could reorder this precedence.
|
||||||
|
*/
|
||||||
|
public static List<String> withAutoCompactWindow(List<String> argv, FleetConfig.Profile profile) {
|
||||||
|
if (profile.autoCompactWindow() == null) {
|
||||||
|
return argv;
|
||||||
|
}
|
||||||
|
List<String> withAutoCompact = new ArrayList<>(argv);
|
||||||
|
withAutoCompact.add("--autocompact");
|
||||||
|
withAutoCompact.add(String.valueOf(profile.autoCompactWindow()));
|
||||||
|
return withAutoCompact;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,313 @@
|
|||||||
|
package dev.ltms.fleet.lead;
|
||||||
|
|
||||||
|
import com.fasterxml.jackson.databind.JsonNode;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
|
||||||
|
import java.io.IOException;
|
||||||
|
import java.io.RandomAccessFile;
|
||||||
|
import java.nio.charset.StandardCharsets;
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.ArrayList;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.ConcurrentHashMap;
|
||||||
|
import java.util.concurrent.atomic.AtomicInteger;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reads how full a lead's own Claude Code context window is, from the transcript Claude Code
|
||||||
|
* itself writes — never from the lead's pane (fleetd has {@code AgentControl.read} for that, and
|
||||||
|
* this must not use it: a pane holds terminal text, not the structured usage numbers a transcript
|
||||||
|
* carries, and scraping it would also race the lead's own rendering).
|
||||||
|
*
|
||||||
|
* <p><strong>Why this exists.</strong> A lead auto-compacts when its context fills — on the host
|
||||||
|
* this was built for, that happened 30 times in one session, discarding roughly 250,000 tokens
|
||||||
|
* and costing 46 seconds to 3 minutes each time, and fleetd had no way to see it coming. This
|
||||||
|
* class is the first thing that looks.
|
||||||
|
*
|
||||||
|
* <p><strong>The route.</strong> Claude Code appends one JSON object per line to
|
||||||
|
* {@code <configDir>/projects/<slug>/<sessionId>.jsonl}. {@code <slug>} is an undocumented,
|
||||||
|
* internal encoding of the working directory — this class never derives it. Instead it lists the
|
||||||
|
* one-level-deep subdirectories of {@code <configDir>/projects/} and looks for
|
||||||
|
* {@code <sessionId>.jsonl} by name, so the slug rule can change without breaking this reader.
|
||||||
|
*
|
||||||
|
* <ul>
|
||||||
|
* <li><strong>Live context</strong> is read off the last record in the read window that carries
|
||||||
|
* a {@code message.usage} object: {@code input_tokens + cache_read_input_tokens +
|
||||||
|
* cache_creation_input_tokens}. This is what actually fills the window — a plain
|
||||||
|
* {@code input_tokens} count alone understates it once the conversation has any cached
|
||||||
|
* prefix, which on a long-lived lead is always.</li>
|
||||||
|
* <li><strong>Compaction history</strong> is a count of {@code subtype: "compact_boundary"}
|
||||||
|
* records seen in the same read window — see {@link Reading#compactions()}. It is a count
|
||||||
|
* within the window this reader actually looked at, not a lifetime total: a session with
|
||||||
|
* more compactions than fit in {@link #TAIL_BYTES} of transcript will undercount. That
|
||||||
|
* trade-off is deliberate — see {@link #TAIL_BYTES}.</li>
|
||||||
|
* </ul>
|
||||||
|
*
|
||||||
|
* <p><strong>Three states, not two (OK / HIGH / UNKNOWN).</strong> Every path that cannot
|
||||||
|
* positively establish the live token count — a missing file, an unreadable one, a peer that
|
||||||
|
* is not a Claude backend, or every line in the read window failing to parse as JSON — returns
|
||||||
|
* {@link State#UNKNOWN} with no token number, never a default "0" or "ok" that would read as
|
||||||
|
* "this lead is fine" when the honest answer is "I could not look".
|
||||||
|
*
|
||||||
|
* <p><strong>A torn final line does not mean UNKNOWN.</strong> {@code fleet_list} reads this
|
||||||
|
* transcript while Claude Code may be mid-write on it, so the last line in the window can be cut
|
||||||
|
* off mid-flush — that is an ordinary, expected race, not a sign the format has changed. Earlier
|
||||||
|
* this class treated ANY unparseable last line as UNKNOWN, on the theory that "if the format
|
||||||
|
* changes, we should see UNKNOWN". That reasoning does not hold: a real format change makes
|
||||||
|
* <em>every</em> line in the window unparseable, not only the last one written. So a single
|
||||||
|
* malformed line (most often the final, torn one, but the check is not position-specific) is
|
||||||
|
* simply skipped rather than treated as fatal, and the reading is built from whatever lines in the
|
||||||
|
* window did parse. Only when <em>none</em> of them parse — the real format-change signal — does
|
||||||
|
* this return {@link State#UNKNOWN}, still with no stale number standing in for "I could not
|
||||||
|
* tell".
|
||||||
|
*
|
||||||
|
* <p><strong>Bounded cost.</strong> {@code fleet_list} is polled constantly, so every read is
|
||||||
|
* capped two ways: {@link #TAIL_BYTES} bounds how much of the transcript is ever read from disk
|
||||||
|
* (never the whole 52 MB a long-lived transcript reaches on the host this was measured on), and
|
||||||
|
* {@link #DEFAULT_CACHE_TTL_MILLIS} bounds how often that bounded read actually happens — a burst
|
||||||
|
* of {@code fleet_list} calls inside one TTL window reads the file once. One instance's cache is
|
||||||
|
* keyed by {@code (configDir, sessionId)}, so it is safe to share across every lead a single
|
||||||
|
* {@code fleet_list} call reports on.
|
||||||
|
*/
|
||||||
|
public final class LeadContextGauge {
|
||||||
|
|
||||||
|
private static final String PROJECTS_DIR = "projects";
|
||||||
|
private static final ObjectMapper MAPPER = new ObjectMapper();
|
||||||
|
|
||||||
|
/**
|
||||||
|
* How many trailing bytes of a transcript a single read ever pulls off disk. Chosen so one
|
||||||
|
* read comfortably spans many recent turns — each usage or compact_boundary record is at most
|
||||||
|
* a few KB — while staying nowhere near the 52 MB a long session's real transcript reaches on
|
||||||
|
* the host this was built for; reading that whole file on every {@code fleet_list} call is
|
||||||
|
* exactly the cost this bound exists to avoid. 2 MiB holds on the order of hundreds of recent
|
||||||
|
* lines even when a turn's tool output is unusually large, which is far more than needed to
|
||||||
|
* find the most recent usage record and any recent compaction.
|
||||||
|
*/
|
||||||
|
static final int TAIL_BYTES = 2 * 1024 * 1024;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* How long a {@link Reading} is served from cache before the file is read again.
|
||||||
|
* {@code fleet_list} is called constantly (by design — it is the fleet's own status probe), so
|
||||||
|
* without a TTL a burst of calls would re-read the transcript tail once per call. 5 seconds is
|
||||||
|
* short enough that a caller watching for a state change never waits long, and long enough that
|
||||||
|
* a poll loop calling every second or two only touches disk once per window.
|
||||||
|
*/
|
||||||
|
static final long DEFAULT_CACHE_TTL_MILLIS = 5_000;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Live tokens at or above this count report {@link State#HIGH}. On the host this was measured
|
||||||
|
* on, auto-compaction actually fires around 267,000–270,000 tokens, but the point of a HIGH
|
||||||
|
* state is to warn before that happens, not at it — 200,000 is the standard Claude context
|
||||||
|
* window size and a sensible built-in default: no config key is required to pick it, and a
|
||||||
|
* lead crossing it is already deep enough into its window that a compaction is foreseeable.
|
||||||
|
*/
|
||||||
|
static final long HIGH_THRESHOLD_TOKENS = 200_000;
|
||||||
|
|
||||||
|
/** The only peer kind this reader understands ({@code Agent.agentType()}'s wire value). */
|
||||||
|
private static final String CLAUDE_AGENT_TYPE = "claude";
|
||||||
|
|
||||||
|
public enum State { OK, HIGH, UNKNOWN }
|
||||||
|
|
||||||
|
/**
|
||||||
|
* @param state {@link State#UNKNOWN} whenever {@code tokens} could not be established
|
||||||
|
* @param tokens live context tokens, or {@code null} exactly when {@code state} is
|
||||||
|
* {@link State#UNKNOWN}
|
||||||
|
* @param compactions {@code compact_boundary} records seen in the read window (see class
|
||||||
|
* javadoc) — {@code 0} both for "genuinely none seen" and for "unknown",
|
||||||
|
* since a caller that already sees {@code state: UNKNOWN} has no reason to
|
||||||
|
* trust this number either way
|
||||||
|
*/
|
||||||
|
public record Reading(State state, Long tokens, int compactions) {
|
||||||
|
/**
|
||||||
|
* fleetd #609: widened from package-private to public so {@code
|
||||||
|
* dev.ltms.fleet.msg.LeadHeartbeatLoop.LeadContextSource.none()} (a different package) can
|
||||||
|
* return the same inert "I could not look" reading the gauge itself uses, without inventing
|
||||||
|
* a parallel unknown-reading constant. Behaviour of this class is otherwise unchanged.
|
||||||
|
*/
|
||||||
|
public static Reading unknown() {
|
||||||
|
return new Reading(State.UNKNOWN, null, 0);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private record CacheEntry(Reading reading, long readAtMillis) {
|
||||||
|
}
|
||||||
|
|
||||||
|
private final LongSupplier clock;
|
||||||
|
private final long ttlMillis;
|
||||||
|
private final Map<String, CacheEntry> cache = new ConcurrentHashMap<>();
|
||||||
|
/** Test seam only (package-private) — counts real disk reads, i.e. cache misses. */
|
||||||
|
private final AtomicInteger diskReads = new AtomicInteger();
|
||||||
|
|
||||||
|
public LeadContextGauge() {
|
||||||
|
this(System::currentTimeMillis, DEFAULT_CACHE_TTL_MILLIS);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Test seam: an injectable clock and TTL so cache expiry is provable without sleeping. */
|
||||||
|
LeadContextGauge(LongSupplier clock, long ttlMillis) {
|
||||||
|
this.clock = clock;
|
||||||
|
this.ttlMillis = ttlMillis;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** How many times this instance has actually read a transcript off disk — test seam only. */
|
||||||
|
int diskReadCount() {
|
||||||
|
return diskReads.get();
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* @param configDir the lead's {@code CLAUDE_CONFIG_DIR}, or {@code null}/blank to use the
|
||||||
|
* default {@code <user.home>/.claude} — the right answer for the common case
|
||||||
|
* where the lead's profile sets no {@code configDir} override
|
||||||
|
* @param sessionId the lead's own Claude session id ({@code Agent.sessionId()}), or
|
||||||
|
* {@code null} when herdr has not resolved one yet
|
||||||
|
* @param agentType the detected peer kind ({@code Agent.agentType()}); anything other than
|
||||||
|
* {@code "claude"} (including {@code null}, meaning undetected) reports
|
||||||
|
* {@link State#UNKNOWN} — this reader only understands Claude Code's own
|
||||||
|
* transcript format
|
||||||
|
*/
|
||||||
|
public Reading read(String configDir, String sessionId, String agentType) {
|
||||||
|
if (sessionId == null || sessionId.isBlank()) {
|
||||||
|
return Reading.unknown();
|
||||||
|
}
|
||||||
|
if (!CLAUDE_AGENT_TYPE.equalsIgnoreCase(agentType)) {
|
||||||
|
return Reading.unknown();
|
||||||
|
}
|
||||||
|
String base = (configDir == null || configDir.isBlank())
|
||||||
|
? System.getProperty("user.home") + "/.claude"
|
||||||
|
: configDir;
|
||||||
|
String cacheKey = base + '\u0000' + sessionId;
|
||||||
|
long now = clock.getAsLong();
|
||||||
|
CacheEntry cached = cache.get(cacheKey);
|
||||||
|
if (cached != null && now - cached.readAtMillis() < ttlMillis) {
|
||||||
|
return cached.reading();
|
||||||
|
}
|
||||||
|
Reading fresh = readUncached(base, sessionId);
|
||||||
|
cache.put(cacheKey, new CacheEntry(fresh, now));
|
||||||
|
return fresh;
|
||||||
|
}
|
||||||
|
|
||||||
|
private Reading readUncached(String base, String sessionId) {
|
||||||
|
diskReads.incrementAndGet();
|
||||||
|
Path file = findTranscript(base, sessionId);
|
||||||
|
if (file == null) {
|
||||||
|
return Reading.unknown();
|
||||||
|
}
|
||||||
|
TailRead tail;
|
||||||
|
try {
|
||||||
|
tail = tailBytes(file, TAIL_BYTES);
|
||||||
|
} catch (IOException e) {
|
||||||
|
return Reading.unknown();
|
||||||
|
}
|
||||||
|
return parse(tail);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Finds {@code <sessionId>.jsonl} under {@code <base>/projects/}, one level deep — never by
|
||||||
|
* deriving the slug directory from a working directory (see class javadoc). Bounded to a
|
||||||
|
* single {@code list()} of {@code projects/} itself: it never recurses further, so the cost is
|
||||||
|
* the number of project directories, not the size of any transcript inside them.
|
||||||
|
*/
|
||||||
|
private Path findTranscript(String base, String sessionId) {
|
||||||
|
Path projectsDir = Path.of(base, PROJECTS_DIR);
|
||||||
|
if (!Files.isDirectory(projectsDir)) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
String filename = sessionId + ".jsonl";
|
||||||
|
Path direct = projectsDir.resolve(filename);
|
||||||
|
if (Files.isRegularFile(direct)) {
|
||||||
|
return direct;
|
||||||
|
}
|
||||||
|
try (var children = Files.list(projectsDir)) {
|
||||||
|
return children.filter(Files::isDirectory)
|
||||||
|
.map(dir -> dir.resolve(filename))
|
||||||
|
.filter(Files::isRegularFile)
|
||||||
|
.findFirst()
|
||||||
|
.orElse(null);
|
||||||
|
} catch (IOException e) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Package-private (not {@code private}): {@link #tailBytes} is a test seam, see its javadoc. */
|
||||||
|
record TailRead(byte[] bytes, boolean fromStart) {
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reads at most {@code maxBytes} trailing bytes of {@code file}. Package-private (not
|
||||||
|
* {@code private}) so a test can assert directly on the returned array's length — "bytes
|
||||||
|
* actually read", not on any parsed answer — without needing a file anywhere near
|
||||||
|
* {@link #TAIL_BYTES} in size to prove the cap holds.
|
||||||
|
*/
|
||||||
|
static TailRead tailBytes(Path file, int maxBytes) throws IOException {
|
||||||
|
try (RandomAccessFile raf = new RandomAccessFile(file.toFile(), "r")) {
|
||||||
|
long length = raf.length();
|
||||||
|
long start = Math.max(0, length - maxBytes);
|
||||||
|
raf.seek(start);
|
||||||
|
byte[] buf = new byte[(int) (length - start)];
|
||||||
|
raf.readFully(buf);
|
||||||
|
return new TailRead(buf, start == 0);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Parses the tail into a {@link Reading}. The first line is dropped unconditionally whenever
|
||||||
|
* the tail is not the whole file (it starts mid-line, cut by {@link #TAIL_BYTES} — an expected
|
||||||
|
* artefact of the bound, not a data problem). Every remaining line is then parsed on a
|
||||||
|
* best-effort basis: a line that fails to parse (most often the last one, torn by a write this
|
||||||
|
* read raced — see the "torn final line" section of the class javadoc) is skipped, not fatal.
|
||||||
|
* Only when none of the remaining lines parse does this report {@link State#UNKNOWN}.
|
||||||
|
*/
|
||||||
|
private Reading parse(TailRead tail) {
|
||||||
|
String text = new String(tail.bytes(), StandardCharsets.UTF_8);
|
||||||
|
List<String> lines = new ArrayList<>(List.of(text.split("\n", -1)));
|
||||||
|
if (!lines.isEmpty() && lines.get(lines.size() - 1).isEmpty()) {
|
||||||
|
lines.remove(lines.size() - 1); // trailing newline leaves a phantom empty last element
|
||||||
|
}
|
||||||
|
if (!tail.fromStart() && !lines.isEmpty()) {
|
||||||
|
lines.remove(0); // first line is a fragment cut by our own tail bound, not real data
|
||||||
|
}
|
||||||
|
if (lines.isEmpty()) {
|
||||||
|
return Reading.unknown();
|
||||||
|
}
|
||||||
|
Long tokens = null;
|
||||||
|
int compactions = 0;
|
||||||
|
boolean anyLineParsed = false;
|
||||||
|
for (String line : lines) {
|
||||||
|
JsonNode node = tryParse(line);
|
||||||
|
if (node == null) {
|
||||||
|
// A malformed line — typically the last one, cut mid-flush by a write this read
|
||||||
|
// raced — is skipped rather than treated as fatal. See the class javadoc's "torn
|
||||||
|
// final line" section for why: a real format change makes EVERY line unparseable,
|
||||||
|
// not only this one, and that case is still caught below by anyLineParsed.
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
anyLineParsed = true;
|
||||||
|
JsonNode usage = node.path("message").path("usage");
|
||||||
|
if (usage.isObject()) {
|
||||||
|
tokens = usage.path("input_tokens").asLong(0)
|
||||||
|
+ usage.path("cache_read_input_tokens").asLong(0)
|
||||||
|
+ usage.path("cache_creation_input_tokens").asLong(0);
|
||||||
|
}
|
||||||
|
if ("compact_boundary".equals(node.path("subtype").asText(null))) {
|
||||||
|
compactions++;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (!anyLineParsed) {
|
||||||
|
return Reading.unknown();
|
||||||
|
}
|
||||||
|
if (tokens == null) {
|
||||||
|
return new Reading(State.UNKNOWN, null, compactions);
|
||||||
|
}
|
||||||
|
State state = tokens >= HIGH_THRESHOLD_TOKENS ? State.HIGH : State.OK;
|
||||||
|
return new Reading(state, tokens, compactions);
|
||||||
|
}
|
||||||
|
|
||||||
|
private JsonNode tryParse(String line) {
|
||||||
|
try {
|
||||||
|
return MAPPER.readTree(line);
|
||||||
|
} catch (IOException e) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -8,6 +8,7 @@ import dev.ltms.fleet.herdr.PendingCloseMarker;
|
|||||||
import dev.ltms.fleet.herdr.Tab;
|
import dev.ltms.fleet.herdr.Tab;
|
||||||
import dev.ltms.fleet.herdr.Workspace;
|
import dev.ltms.fleet.herdr.Workspace;
|
||||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||||
|
import dev.ltms.fleet.launch.ClaudeCodeArguments;
|
||||||
import org.slf4j.Logger;
|
import org.slf4j.Logger;
|
||||||
import org.slf4j.LoggerFactory;
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
@@ -358,7 +359,8 @@ public final class LeadLauncher {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* The lead's argv: the profile's own command, the model pin, and the bridge MCP mount.
|
* The lead's argv: the profile's own command, the model and auto-compaction pins, and the bridge
|
||||||
|
* MCP mount.
|
||||||
*
|
*
|
||||||
* <p>No {@code --append-system-prompt}. That flag carries the worker reply charter, and a lead
|
* <p>No {@code --append-system-prompt}. That flag carries the worker reply charter, and a lead
|
||||||
* is not a worker — it reads its orchestration rules from the project's {@code CLAUDE.md} like
|
* is not a worker — it reads its orchestration rules from the project's {@code CLAUDE.md} like
|
||||||
@@ -380,7 +382,7 @@ public final class LeadLauncher {
|
|||||||
argv.add("--model");
|
argv.add("--model");
|
||||||
argv.add(profile.model());
|
argv.add(profile.model());
|
||||||
}
|
}
|
||||||
return argv;
|
return profile.isOpenCode() ? argv : ClaudeCodeArguments.withAutoCompactWindow(argv, profile);
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
|
|||||||
@@ -10,6 +10,8 @@ import java.io.IOException;
|
|||||||
import java.io.UncheckedIOException;
|
import java.io.UncheckedIOException;
|
||||||
import java.nio.file.Files;
|
import java.nio.file.Files;
|
||||||
import java.nio.file.Path;
|
import java.nio.file.Path;
|
||||||
|
import java.util.Collections;
|
||||||
|
import java.util.LinkedHashMap;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
import java.util.UUID;
|
import java.util.UUID;
|
||||||
import java.util.concurrent.ConcurrentHashMap;
|
import java.util.concurrent.ConcurrentHashMap;
|
||||||
@@ -25,9 +27,11 @@ import java.util.function.Supplier;
|
|||||||
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation
|
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation
|
||||||
* clears the lead's own pane and bootstraps a fresh session against that file.
|
* clears the lead's own pane and bootstraps a fresh session against that file.
|
||||||
*
|
*
|
||||||
* <p>This is the executor only. Nothing in this ticket wires an MCP tool onto {@link #open}/
|
* <p>This is the executor behind the {@code fleet_handover} MCP tool ({@code
|
||||||
* {@link #confirm}/{@link #cancel} — that is a separate, later unit; until it lands, nothing calls
|
* dev.ltms.fleet.mcp.FleetMcp#handover}), which drives {@link #open}, {@link #confirm}, {@link
|
||||||
* this class at all.
|
* #cancel}, and {@link #status} from a tool call — wired in fleetd #480 Unit C. <strong>An earlier
|
||||||
|
* version of this paragraph said nothing called this class at all; that stopped being true once
|
||||||
|
* that unit landed, and this correction exists so the javadoc does not go on claiming it.</strong>
|
||||||
*
|
*
|
||||||
* <p><strong>{@code confirm()} cannot roll inline — a fleetd #480 correction.</strong> The first
|
* <p><strong>{@code confirm()} cannot roll inline — a fleetd #480 correction.</strong> The first
|
||||||
* version of this class called {@code agents.send(lead, "/clear")} directly from inside {@code
|
* version of this class called {@code agents.send(lead, "/clear")} directly from inside {@code
|
||||||
@@ -181,6 +185,106 @@ public final class LeadRollover {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* How many tokens {@link #outcomes} remembers before it starts evicting the oldest — bounded
|
||||||
|
* so a long-running daemon never grows this map without limit. Chosen generously rather than
|
||||||
|
* tightly: production rolls are rare (this class's own ticket found exactly ONE completed roll
|
||||||
|
* ever logged on this host), and each entry is a handful of short strings, so even a full cap
|
||||||
|
* costs a few tens of kilobytes — nowhere near a reason to make it configurable. 200 entries
|
||||||
|
* comfortably outlasts any operator's own memory of "did that roll I asked for actually
|
||||||
|
* happen", which is the whole reason {@link #status} exists.
|
||||||
|
*
|
||||||
|
* <p><strong>This cap counts {@link RollState#IN_PROGRESS} entries exactly the same as
|
||||||
|
* finished ones.</strong> There is only the one bounded map: {@link #confirm} writes an {@link
|
||||||
|
* RollState#IN_PROGRESS} entry into {@link #outcomes} at hand-off, and the deferred
|
||||||
|
* continuation later overwrites that SAME key with a terminal state — it never inserts a
|
||||||
|
* second entry. An approved roll therefore occupies one slot in this map for its entire
|
||||||
|
* lifetime, from the moment {@link #confirm} hands off, not only once it finishes; a
|
||||||
|
* confirmed-but-not-yet-finished roll counts against the cap exactly like a finished one. The
|
||||||
|
* alternative (a separate, uncapped in-flight map) would let a burst of confirmed-but-stuck
|
||||||
|
* rolls grow without bound — the exact failure this cap exists to prevent — so it was rejected.
|
||||||
|
*/
|
||||||
|
static final int OUTCOME_HISTORY_CAP = 200;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* What is known about one token, right now — the answer {@link #status} gives. Distinguishes
|
||||||
|
* three terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
|
||||||
|
* that has been approved but has not finished yet, and two answers for a token that names no
|
||||||
|
* active work at all: still pending confirmation, or nothing known about this token at all.
|
||||||
|
*/
|
||||||
|
public enum RollState {
|
||||||
|
/**
|
||||||
|
* {@code token} is still open: either {@link #open} was called and {@link #confirm} has not
|
||||||
|
* been (or not successfully) yet, or a {@link #confirm} call failed one of its gate checks
|
||||||
|
* and left the token pending for a retry — see {@link #confirm}'s javadoc ("token stays
|
||||||
|
* pending"). Indistinguishable from a genuinely fresh request; a caller wanting to know
|
||||||
|
* WHICH gate most recently refused should read the {@link RollDecision} that {@link
|
||||||
|
* #confirm} itself returned, not this status. <strong>Never the state of an APPROVED
|
||||||
|
* roll</strong> — see {@link #IN_PROGRESS}, which {@link #confirm} records at the moment it
|
||||||
|
* hands off, before this token is even removed from the pending set.
|
||||||
|
*/
|
||||||
|
PENDING,
|
||||||
|
/**
|
||||||
|
* {@link #confirm} approved this roll and handed it to the deferred continuation, which has
|
||||||
|
* not finished yet. Recorded by {@link #confirm} itself, at hand-off — <strong>before</strong>
|
||||||
|
* {@code token} is removed from the pending set — so there is never a gap in which {@link
|
||||||
|
* #status} could wrongly answer {@link #UNKNOWN} ("nothing was ever requested") for a roll
|
||||||
|
* that is, in fact, actively running. This is not sticky: the deferred continuation
|
||||||
|
* overwrites this same entry with a terminal state ({@link #ROLLED}, {@link
|
||||||
|
* #TURN_NEVER_SETTLED}, {@link #CLEAR_NEVER_SETTLED}, or {@link #FAILED}) once it finishes
|
||||||
|
* — including by throwing, which fleetd #615's catch in {@link #runRollover} now turns into
|
||||||
|
* {@link #FAILED} instead of leaving this entry stuck forever.
|
||||||
|
*/
|
||||||
|
IN_PROGRESS,
|
||||||
|
/**
|
||||||
|
* {@link #confirm} was approved and the deferred continuation completed the entire roll:
|
||||||
|
* the calling lead's turn settled, {@code /clear} was sent and settled, and {@code
|
||||||
|
* bootstrapText} was sent.
|
||||||
|
*/
|
||||||
|
ROLLED,
|
||||||
|
/**
|
||||||
|
* {@link #confirm} was approved, but the calling lead's own turn never reached a boundary
|
||||||
|
* (IDLE or DONE) within {@code turnSettleSeconds} — no {@code /clear} was ever sent, at
|
||||||
|
* all. This is the branch the fleetd #480 correction exists to make safe, and the one this
|
||||||
|
* status exists to make VISIBLE: before this, a lead that hit this case had no way to find
|
||||||
|
* out, and would carry on believing it was about to be replaced. See this class's javadoc.
|
||||||
|
*/
|
||||||
|
TURN_NEVER_SETTLED,
|
||||||
|
/**
|
||||||
|
* {@link #confirm} was approved and {@code /clear} was sent, but the pane never re-settled
|
||||||
|
* within {@code clearSettleSeconds} — {@code bootstrapText} was never sent.
|
||||||
|
*/
|
||||||
|
CLEAR_NEVER_SETTLED,
|
||||||
|
/**
|
||||||
|
* fleetd #615: the deferred continuation threw a {@link RuntimeException} — most likely a
|
||||||
|
* {@link dev.ltms.fleet.herdr.HerdrException} out of one of the two unwrapped {@code
|
||||||
|
* agents.send} calls in {@link #runRollover} — and the continuation thread died with it.
|
||||||
|
* Before this state existed, that throw left {@link #outcomes} holding {@link #IN_PROGRESS}
|
||||||
|
* forever, because the production {@code continuationRunner} is a bare virtual thread with
|
||||||
|
* no uncaught-exception handler and nothing downstream of the throw ever ran to write a
|
||||||
|
* terminal outcome. {@code detail} names the exception, so a reader has something to act on
|
||||||
|
* — the same diagnostic style as {@link #TURN_NEVER_SETTLED} and {@link
|
||||||
|
* #CLEAR_NEVER_SETTLED}. The roll is dead at this point and does not retry itself; a stuck
|
||||||
|
* lead must {@link #open} a fresh request.
|
||||||
|
*/
|
||||||
|
FAILED,
|
||||||
|
/**
|
||||||
|
* {@code token} names nothing this instance currently knows about: never issued by {@link
|
||||||
|
* #open}, dropped by {@link #cancel}, or aged out of {@link #outcomes}'s bounded history.
|
||||||
|
* These three causes are not distinguished — all of them mean "there is nothing to tell
|
||||||
|
* you", which is the entire content of a clean answer here.
|
||||||
|
*/
|
||||||
|
UNKNOWN
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The answer {@link #status} gives for one token: a {@link RollState} and a human-readable
|
||||||
|
* {@code detail}. For {@link RollState#TURN_NEVER_SETTLED}, {@code detail} names {@code
|
||||||
|
* turnSettleSeconds} and its configured value explicitly, so a reader who sees this knows what
|
||||||
|
* to raise.
|
||||||
|
*/
|
||||||
|
public record RollStatus(RollState state, String detail) {}
|
||||||
|
|
||||||
private final AgentControl agents;
|
private final AgentControl agents;
|
||||||
private final Supplier<FleetConfig.LeadRollover> configSupplier;
|
private final Supplier<FleetConfig.LeadRollover> configSupplier;
|
||||||
/**
|
/**
|
||||||
@@ -201,6 +305,23 @@ public final class LeadRollover {
|
|||||||
*/
|
*/
|
||||||
private final Consumer<Runnable> continuationRunner;
|
private final Consumer<Runnable> continuationRunner;
|
||||||
private final Map<String, PendingRollover> pending = new ConcurrentHashMap<>();
|
private final Map<String, PendingRollover> pending = new ConcurrentHashMap<>();
|
||||||
|
/**
|
||||||
|
* Finished tokens → what actually happened, for {@link #status}. Bounded by {@link
|
||||||
|
* #OUTCOME_HISTORY_CAP}, oldest evicted first ({@code removeEldestEntry} on an insertion-order
|
||||||
|
* {@link LinkedHashMap}). Wrapped in {@link Collections#synchronizedMap} because entries are
|
||||||
|
* written from whatever thread {@code continuationRunner} runs the roll on (a fresh virtual
|
||||||
|
* thread in production, the calling test thread under {@code Runnable::run}) and read from
|
||||||
|
* whatever thread calls {@link #status} (the MCP handler thread) — a plain {@code
|
||||||
|
* LinkedHashMap} is not safe for that, and {@code removeEldestEntry} additionally requires
|
||||||
|
* external synchronization even for a thread-safe map that merely wraps it.
|
||||||
|
*/
|
||||||
|
private final Map<String, RollStatus> outcomes = Collections.synchronizedMap(
|
||||||
|
new LinkedHashMap<>(16, 0.75f, false) {
|
||||||
|
@Override
|
||||||
|
protected boolean removeEldestEntry(Map.Entry<String, RollStatus> eldest) {
|
||||||
|
return size() > OUTCOME_HISTORY_CAP;
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
/** Production constructor — wall clock, real sleep between settle polls, a real virtual thread. */
|
/** Production constructor — wall clock, real sleep between settle polls, a real virtual thread. */
|
||||||
public LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
|
public LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
|
||||||
@@ -366,6 +487,15 @@ public final class LeadRollover {
|
|||||||
return docCheck;
|
return docCheck;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Record IN_PROGRESS BEFORE removing from `pending` — see RollState#IN_PROGRESS and
|
||||||
|
// OUTCOME_HISTORY_CAP's javadoc. This ordering means `token` is written into `outcomes`
|
||||||
|
// while it is STILL present in `pending`; status() checks `outcomes` first (see that
|
||||||
|
// method), so it reports IN_PROGRESS immediately, not the brief-but-real gap a
|
||||||
|
// remove-then-put ordering would leave in which the token is in neither map.
|
||||||
|
outcomes.put(token, new RollStatus(RollState.IN_PROGRESS,
|
||||||
|
"confirm() approved this roll and handed it to the deferred continuation; it has "
|
||||||
|
+ "not finished yet — still waiting for the calling turn to settle, for "
|
||||||
|
+ "/clear to be sent and settle, or for bootstrapText to be sent"));
|
||||||
pending.remove(token);
|
pending.remove(token);
|
||||||
log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends",
|
log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends",
|
||||||
token, callerTerminal);
|
token, callerTerminal);
|
||||||
@@ -378,8 +508,42 @@ public final class LeadRollover {
|
|||||||
* entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the
|
* entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the
|
||||||
* four-step order. There is no result to return to by this point, so every outcome is logged
|
* four-step order. There is no result to return to by this point, so every outcome is logged
|
||||||
* only.
|
* only.
|
||||||
|
*
|
||||||
|
* <p><strong>fleetd #615 — the whole body is wrapped in one {@code try}.</strong> The two {@code
|
||||||
|
* agents.send} calls below are not wrapped individually: {@code send} → {@code agentCall} →
|
||||||
|
* {@code herdr.call} can throw an unchecked {@link dev.ltms.fleet.herdr.HerdrException} (see
|
||||||
|
* {@code AgentControl.java}), and the production {@code continuationRunner} is a bare virtual
|
||||||
|
* thread with no uncaught-exception handler (see this class's public constructor). Before this
|
||||||
|
* fix, either throw killed the continuation thread silently, leaving the {@link
|
||||||
|
* RollState#IN_PROGRESS} entry {@link #confirm} wrote at hand-off stuck forever — {@link
|
||||||
|
* #status} had no way to tell a dead roll from one still genuinely running. The {@code catch}
|
||||||
|
* below is scoped to the method body rather than to each {@code send} call individually, so it
|
||||||
|
* also covers anything else added to this continuation later, not just today's two call sites —
|
||||||
|
* the same reasoning that put the write-a-terminal-outcome step at each of this method's other
|
||||||
|
* exits (see the {@link RollState#TURN_NEVER_SETTLED} and {@link RollState#CLEAR_NEVER_SETTLED}
|
||||||
|
* branches below) rather than inside the helpers that detect them.</p>
|
||||||
|
*
|
||||||
|
* <p>Only {@link RuntimeException} is caught, matching the local convention {@link
|
||||||
|
* #waitUntilAtTurnBoundary} already set around its own {@code agents.status} call — not the
|
||||||
|
* broader {@link Exception} or {@link Throwable}, which would also swallow something like an
|
||||||
|
* {@link OutOfMemoryError} this continuation has no business handling.</p>
|
||||||
*/
|
*/
|
||||||
private void runRollover(PendingRollover p, FleetConfig.LeadRollover cfg) {
|
private void runRollover(PendingRollover p, FleetConfig.LeadRollover cfg) {
|
||||||
|
try {
|
||||||
|
runRolloverUnguarded(p, cfg);
|
||||||
|
} catch (RuntimeException e) {
|
||||||
|
log.warn("lead-rollover: continuation for token={} lead={} threw {} — the roll is dead; "
|
||||||
|
+ "no further step in this continuation will run",
|
||||||
|
p.token(), p.leadTerminal(), e.toString(), e);
|
||||||
|
outcomes.put(p.token(), new RollStatus(RollState.FAILED,
|
||||||
|
"the roll's continuation threw " + e.toString() + " — the roll is dead and will "
|
||||||
|
+ "not retry itself; check the daemon log for the stack trace, then open() "
|
||||||
|
+ "a fresh rollover request"));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** The actual body of {@link #runRollover}, unwrapped — see that method's javadoc for the catch. */
|
||||||
|
private void runRolloverUnguarded(PendingRollover p, FleetConfig.LeadRollover cfg) {
|
||||||
String lead = p.leadTerminal();
|
String lead = p.leadTerminal();
|
||||||
long rollStartMillis = nowMillis.getAsLong();
|
long rollStartMillis = nowMillis.getAsLong();
|
||||||
TurnSettleResult turnResult = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds());
|
TurnSettleResult turnResult = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds());
|
||||||
@@ -393,6 +557,11 @@ public final class LeadRollover {
|
|||||||
+ "turn is still live and clearing it now would destroy live context "
|
+ "turn is still live and clearing it now would destroy live context "
|
||||||
+ "(token={}, configured={}s elapsed={}ms)",
|
+ "(token={}, configured={}s elapsed={}ms)",
|
||||||
lead, p.token(), cfg.turnSettleSeconds(), turnResult.elapsedMillis());
|
lead, p.token(), cfg.turnSettleSeconds(), turnResult.elapsedMillis());
|
||||||
|
outcomes.put(p.token(), new RollStatus(RollState.TURN_NEVER_SETTLED,
|
||||||
|
"the calling lead's own turn never reached a boundary (IDLE or DONE) within "
|
||||||
|
+ "turnSettleSeconds=" + cfg.turnSettleSeconds() + "s (measured elapsed="
|
||||||
|
+ turnResult.elapsedMillis() + "ms) — no /clear was ever sent. If this "
|
||||||
|
+ "keeps happening, raise turnSettleSeconds in fleetd.yaml"));
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -411,11 +580,17 @@ public final class LeadRollover {
|
|||||||
+ "elapsed={}ms nudges={})",
|
+ "elapsed={}ms nudges={})",
|
||||||
lead, p.token(), cfg.clearSettleSeconds(), clearResult.elapsedMillis(),
|
lead, p.token(), cfg.clearSettleSeconds(), clearResult.elapsedMillis(),
|
||||||
clearResult.nudges());
|
clearResult.nudges());
|
||||||
|
outcomes.put(p.token(), new RollStatus(RollState.CLEAR_NEVER_SETTLED,
|
||||||
|
"/clear was sent, but the pane never re-settled within clearSettleSeconds="
|
||||||
|
+ cfg.clearSettleSeconds() + "s (measured elapsed=" + clearResult.elapsedMillis()
|
||||||
|
+ "ms, nudges=" + clearResult.nudges() + ") — bootstrapText was never sent"));
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()));
|
agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()));
|
||||||
long rollElapsedMillis = nowMillis.getAsLong() - rollStartMillis;
|
long rollElapsedMillis = nowMillis.getAsLong() - rollStartMillis;
|
||||||
log.info("lead-rollover: rolled token={} lead={} elapsedMs={}", p.token(), lead, rollElapsedMillis);
|
log.info("lead-rollover: rolled token={} lead={} elapsedMs={}", p.token(), lead, rollElapsedMillis);
|
||||||
|
outcomes.put(p.token(), new RollStatus(RollState.ROLLED,
|
||||||
|
"rolled successfully in " + rollElapsedMillis + "ms"));
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Drop a pending request without rolling. @return whether a pending request existed for {@code token} */
|
/** Drop a pending request without rolling. @return whether a pending request existed for {@code token} */
|
||||||
@@ -423,6 +598,47 @@ public final class LeadRollover {
|
|||||||
return pending.remove(token) != null;
|
return pending.remove(token) != null;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Read-only: what is currently known about {@code token}. <strong>Never sends anything, never
|
||||||
|
* schedules, cancels, or retries a roll</strong> — a caller may poll this as often as it likes
|
||||||
|
* with no side effect at all, which is exactly why it exists: every failure past {@link
|
||||||
|
* #confirm} used to be a {@code log.warn} a lead can never read (see this class's javadoc), and
|
||||||
|
* this is the only route back.
|
||||||
|
*
|
||||||
|
* @param token the token {@link #open} returned; {@code null} or blank is a clean {@link
|
||||||
|
* RollState#UNKNOWN}, never a {@link NullPointerException} — {@link #pending} is a
|
||||||
|
* {@link ConcurrentHashMap}, which throws on a {@code null} key lookup, so this
|
||||||
|
* short-circuits before ever reaching it
|
||||||
|
* @return {@link RollState#IN_PROGRESS} for an approved roll whose continuation has not
|
||||||
|
* finished yet, or a terminal state once it has (both read from {@link #outcomes} —
|
||||||
|
* checked FIRST, see below); {@link RollState#PENDING} while {@code token} is still
|
||||||
|
* open and has not yet been approved (including one left pending by a {@link #confirm}
|
||||||
|
* gate refusal — see that method's javadoc); or {@link RollState#UNKNOWN} for a token
|
||||||
|
* never issued, cancelled, or aged out of the bounded history
|
||||||
|
*/
|
||||||
|
public RollStatus status(String token) {
|
||||||
|
if (token == null || token.isBlank()) {
|
||||||
|
return new RollStatus(RollState.UNKNOWN, "no token given");
|
||||||
|
}
|
||||||
|
// `outcomes` is checked BEFORE `pending`, deliberately: `confirm` writes an IN_PROGRESS
|
||||||
|
// entry into `outcomes` before it removes `token` from `pending` (see `confirm`'s own
|
||||||
|
// comment at that call site), so for the brief window where a token is present in BOTH
|
||||||
|
// maps, this order reports the more accurate answer (IN_PROGRESS, already approved) rather
|
||||||
|
// than the stale one (PENDING, not yet approved) a pending-first check would give.
|
||||||
|
RollStatus recorded = outcomes.get(token);
|
||||||
|
if (recorded != null) {
|
||||||
|
return recorded;
|
||||||
|
}
|
||||||
|
if (pending.containsKey(token)) {
|
||||||
|
return new RollStatus(RollState.PENDING, "open() has been called for this token and "
|
||||||
|
+ "it has not yet been confirmed — or a confirm() gate check failed and left it "
|
||||||
|
+ "pending, so the same token may be retried once the problem is fixed");
|
||||||
|
}
|
||||||
|
return new RollStatus(RollState.UNKNOWN, "token names no pending or finished rollover "
|
||||||
|
+ "request known to this instance — never issued, cancelled, or aged out of the "
|
||||||
|
+ "bounded history (cap=" + OUTCOME_HISTORY_CAP + ")");
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* The three handover-file checks, in order: exists, not empty, fresh (modified after
|
* The three handover-file checks, in order: exists, not empty, fresh (modified after
|
||||||
* {@link #open}'s timestamp and not older than {@code maxDocAgeSeconds}). Stats {@code
|
* {@link #open}'s timestamp and not older than {@code maxDocAgeSeconds}). Stats {@code
|
||||||
|
|||||||
@@ -31,6 +31,20 @@ public final class ConnectionIdentity {
|
|||||||
this.cwds = cwds;
|
this.cwds = cwds;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The {@link PaneLocator} this identity resolves callers against — fleetd #612 CB-185: lets a
|
||||||
|
* test drive the exact {@link PaneLocator} a real assembly wired up (e.g. {@code
|
||||||
|
* FleetdAssembly}'s {@code new ConnectionIdentity(new PaneLocator(herdr, memberHerdr), ...)})
|
||||||
|
* directly with a chosen pid, bypassing the OS-dependent {@link PeerPidLookup} that {@link
|
||||||
|
* #resolve} otherwise goes through. A full HTTP round trip cannot exercise this: {@code
|
||||||
|
* LsofPeerPidLookup} excludes its own pid, and an in-process test client and server share one
|
||||||
|
* JVM pid, so {@code pidForLocalPort} always returns {@code -1} and {@link PaneLocator} never
|
||||||
|
* gets called at all.
|
||||||
|
*/
|
||||||
|
public PaneLocator panes() {
|
||||||
|
return panes;
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
|
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
|
||||||
* primary / an off-host client), its {@code pid} (or {@code -1} if not resolvable), and whether
|
* primary / an off-host client), its {@code pid} (or {@code -1} if not resolvable), and whether
|
||||||
|
|||||||
@@ -12,6 +12,7 @@ import dev.ltms.fleet.metrics.Metrics;
|
|||||||
import dev.ltms.fleet.inject.MemberPresence;
|
import dev.ltms.fleet.inject.MemberPresence;
|
||||||
import dev.ltms.fleet.inject.CompletionResolver;
|
import dev.ltms.fleet.inject.CompletionResolver;
|
||||||
import dev.ltms.fleet.herdr.HerdrException;
|
import dev.ltms.fleet.herdr.HerdrException;
|
||||||
|
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||||
import dev.ltms.fleet.lead.LeadRollover;
|
import dev.ltms.fleet.lead.LeadRollover;
|
||||||
import dev.ltms.fleet.msg.LeadChannel;
|
import dev.ltms.fleet.msg.LeadChannel;
|
||||||
import dev.ltms.fleet.msg.LeadMessage;
|
import dev.ltms.fleet.msg.LeadMessage;
|
||||||
@@ -106,6 +107,13 @@ public final class FleetMcp {
|
|||||||
* untested identity heuristic.
|
* untested identity heuristic.
|
||||||
*/
|
*/
|
||||||
private final boolean authorizationEnforced;
|
private final boolean authorizationEnforced;
|
||||||
|
/**
|
||||||
|
* fleetd #612 CB-185: kept as a field (rather than only captured by the {@code
|
||||||
|
* contextExtractor} closure built in the constructor) so a test can reach the exact {@link
|
||||||
|
* ConnectionIdentity} — and, through {@link ConnectionIdentity#panes()}, the exact {@link
|
||||||
|
* dev.ltms.fleet.herdr.PaneLocator} — that a real assembly wired up. See {@link #identity()}.
|
||||||
|
*/
|
||||||
|
private final ConnectionIdentity identity;
|
||||||
private final Metrics metrics; // CB-502: null → auth failures not counted
|
private final Metrics metrics; // CB-502: null → auth failures not counted
|
||||||
private final CapacitySource capacity;
|
private final CapacitySource capacity;
|
||||||
private final HealthCoverageSource healthCoverage;
|
private final HealthCoverageSource healthCoverage;
|
||||||
@@ -115,6 +123,8 @@ public final class FleetMcp {
|
|||||||
private final OutageSource outage;
|
private final OutageSource outage;
|
||||||
/** fleetd #176: SEPARATE from both of the above — see {@link LeadSeatSource}'s doc. */
|
/** fleetd #176: SEPARATE from both of the above — see {@link LeadSeatSource}'s doc. */
|
||||||
private final LeadSeatSource leadSeats;
|
private final LeadSeatSource leadSeats;
|
||||||
|
/** fleetd #602 gauge-wiring: see {@link LeadConfigDirSource}. */
|
||||||
|
private final LeadConfigDirSource leadConfigDirs;
|
||||||
/** CB-637: this daemon's lead-to-lead channel; {@code null} when no coordinator is configured. */
|
/** CB-637: this daemon's lead-to-lead channel; {@code null} when no coordinator is configured. */
|
||||||
private final LeadChannel leadChannel;
|
private final LeadChannel leadChannel;
|
||||||
/** fleetd #361: {@code coordinator.peers} — see {@link CoordinationSource}. Empty when unset. */
|
/** fleetd #361: {@code coordinator.peers} — see {@link CoordinationSource}. Empty when unset. */
|
||||||
@@ -126,6 +136,17 @@ public final class FleetMcp {
|
|||||||
* clean {@code NOT_CONFIGURED} refusal rather than throwing. See {@link #handover}.
|
* clean {@code NOT_CONFIGURED} refusal rather than throwing. See {@link #handover}.
|
||||||
*/
|
*/
|
||||||
private final LeadRollover leadRollover;
|
private final LeadRollover leadRollover;
|
||||||
|
/**
|
||||||
|
* "Lead context gauge": how full each lead's own Claude Code context window is, reported on
|
||||||
|
* {@code fleet_list}'s {@code leads} rows (see {@link #contextView}). Built unconditionally, in
|
||||||
|
* the field initializer rather than a constructor parameter — this is not a togglable feature
|
||||||
|
* with an on/off config knob the way {@link OutageSource}/{@link LeadSeatSource} are: it needs
|
||||||
|
* no config at all (see {@link LeadContextGauge}'s own javadoc for the built-in defaults), so
|
||||||
|
* there is no "off" value to thread through every existing constructor call site. One instance
|
||||||
|
* per daemon so its read cache (keyed by session, TTL'd) is actually shared across
|
||||||
|
* {@code fleet_list} calls rather than rebuilt — and therefore useless — on every call.
|
||||||
|
*/
|
||||||
|
private final LeadContextGauge leadContextGauge = new LeadContextGauge();
|
||||||
|
|
||||||
/** Capacity facts used by {@code fleet_list}; production must supply the placement live count. */
|
/** Capacity facts used by {@code fleet_list}; production must supply the placement live count. */
|
||||||
public record CapacitySource(Function<String, Integer> liveCount, Function<String, Integer> maxLoad,
|
public record CapacitySource(Function<String, Integer> liveCount, Function<String, Integer> maxLoad,
|
||||||
@@ -246,6 +267,29 @@ public final class FleetMcp {
|
|||||||
public static LeadSeatSource none() { return new LeadSeatSource(_ -> 0); }
|
public static LeadSeatSource none() { return new LeadSeatSource(_ -> 0); }
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #602 gauge-wiring: a lead name's configured {@code CLAUDE_CONFIG_DIR} override, fed to
|
||||||
|
* {@link LeadContextGauge#read} so {@code fleet_list}'s {@code context} row reads the transcript
|
||||||
|
* directory the lead's own profile actually writes to, not always the built-in
|
||||||
|
* {@code <user.home>/.claude} default.
|
||||||
|
*
|
||||||
|
* <p>The same idiom as {@link LeadSeatSource} — a {@code FleetMcp} constructor field, not a
|
||||||
|
* lookup {@code contextView} performs itself, because {@code FleetMcp} holds no
|
||||||
|
* {@link dev.ltms.fleet.config.FleetConfig} and {@code contextView} is {@code static}. See
|
||||||
|
* {@code Fleetd.leadConfigDirLookup} for the derivation: the same
|
||||||
|
* {@code fleet.leaders.<name>.profile} link {@link LeadSeatSource} already follows, resolved to
|
||||||
|
* that profile's own {@code configDir:}.
|
||||||
|
*
|
||||||
|
* @param configDirFor lead name → {@code configDir}, or {@code null} when the lead's entry names
|
||||||
|
* no profile, or that profile sets no {@code configDir} override — either
|
||||||
|
* way {@link LeadContextGauge} then falls back to its own built-in default,
|
||||||
|
* exactly as before this ticket
|
||||||
|
*/
|
||||||
|
public record LeadConfigDirSource(Function<String, String> configDirFor) {
|
||||||
|
/** Inert source — every lead reads {@link LeadContextGauge}'s built-in default {@code configDir}. */
|
||||||
|
public static LeadConfigDirSource none() { return new LeadConfigDirSource(_ -> null); }
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* fleetd #361: peer-visibility facts for {@code fleet_list}'s {@code coordinator} row — this
|
* fleetd #361: peer-visibility facts for {@code fleet_list}'s {@code coordinator} row — this
|
||||||
* daemon's own {@link LeadChannel} (for its self mailbox state and held messages) plus the
|
* daemon's own {@link LeadChannel} (for its self mailbox state and held messages) plus the
|
||||||
@@ -346,7 +390,7 @@ public final class FleetMcp {
|
|||||||
LeadSeatSource leadSeats, List<String> peers, LeadRollover leadRollover) {
|
LeadSeatSource leadSeats, List<String> peers, LeadRollover leadRollover) {
|
||||||
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, authorizationMode, metrics,
|
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, authorizationMode, metrics,
|
||||||
capacity, healthCoverage, LoopHealthSource.none(), quarantine, leadChannel, outage, leadSeats,
|
capacity, healthCoverage, LoopHealthSource.none(), quarantine, leadChannel, outage, leadSeats,
|
||||||
peers, leadRollover);
|
LeadConfigDirSource.none(), peers, leadRollover);
|
||||||
}
|
}
|
||||||
|
|
||||||
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
|
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
|
||||||
@@ -354,16 +398,19 @@ public final class FleetMcp {
|
|||||||
CallerResolver callers, AuthorizationMode authorizationMode, Metrics metrics,
|
CallerResolver callers, AuthorizationMode authorizationMode, Metrics metrics,
|
||||||
CapacitySource capacity, HealthCoverageSource healthCoverage, LoopHealthSource loopHealth,
|
CapacitySource capacity, HealthCoverageSource healthCoverage, LoopHealthSource loopHealth,
|
||||||
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage,
|
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage,
|
||||||
LeadSeatSource leadSeats, List<String> peers, LeadRollover leadRollover) {
|
LeadSeatSource leadSeats, LeadConfigDirSource leadConfigDirs, List<String> peers,
|
||||||
|
LeadRollover leadRollover) {
|
||||||
Objects.requireNonNull(callers, "callers");
|
Objects.requireNonNull(callers, "callers");
|
||||||
this.authorizationEnforced = Objects.requireNonNull(authorizationMode, "authorizationMode")
|
this.authorizationEnforced = Objects.requireNonNull(authorizationMode, "authorizationMode")
|
||||||
== AuthorizationMode.ENFORCED;
|
== AuthorizationMode.ENFORCED;
|
||||||
|
this.identity = identity;
|
||||||
this.leadChannel = leadChannel;
|
this.leadChannel = leadChannel;
|
||||||
this.peers = peers == null ? List.of() : List.copyOf(peers);
|
this.peers = peers == null ? List.of() : List.copyOf(peers);
|
||||||
this.capacity = capacity;
|
this.capacity = capacity;
|
||||||
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
|
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
|
||||||
this.outage = Objects.requireNonNull(outage, "outage");
|
this.outage = Objects.requireNonNull(outage, "outage");
|
||||||
this.leadSeats = Objects.requireNonNull(leadSeats, "leadSeats");
|
this.leadSeats = Objects.requireNonNull(leadSeats, "leadSeats");
|
||||||
|
this.leadConfigDirs = Objects.requireNonNull(leadConfigDirs, "leadConfigDirs");
|
||||||
this.healthCoverage = healthCoverage;
|
this.healthCoverage = healthCoverage;
|
||||||
this.loopHealth = Objects.requireNonNull(loopHealth, "loopHealth");
|
this.loopHealth = Objects.requireNonNull(loopHealth, "loopHealth");
|
||||||
this.leadRollover = leadRollover;
|
this.leadRollover = leadRollover;
|
||||||
@@ -498,7 +545,7 @@ public final class FleetMcp {
|
|||||||
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_list", Map.of()), null);
|
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_list", Map.of()), null);
|
||||||
if (denied != null) return denied;
|
if (denied != null) return denied;
|
||||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine, outage,
|
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine, outage,
|
||||||
leadSeats, callers.leads(),
|
leadSeats, leadContextGauge, leadConfigDirs, callers.leads(),
|
||||||
callerTerminal(exchange),
|
callerTerminal(exchange),
|
||||||
new CoordinationSource(leadChannel, peers),
|
new CoordinationSource(leadChannel, peers),
|
||||||
coordinatorVisibleTo(principal(exchange)));
|
coordinatorVisibleTo(principal(exchange)));
|
||||||
@@ -716,6 +763,15 @@ public final class FleetMcp {
|
|||||||
return transport;
|
return transport;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The {@link ConnectionIdentity} this server resolves every caller against — fleetd #612
|
||||||
|
* CB-185: lets a test reach the exact {@link dev.ltms.fleet.herdr.PaneLocator} a real assembly
|
||||||
|
* wired up (via {@link ConnectionIdentity#panes()}), rather than a copy built for the test.
|
||||||
|
*/
|
||||||
|
public ConnectionIdentity identity() {
|
||||||
|
return identity;
|
||||||
|
}
|
||||||
|
|
||||||
/** Mark a connected spawned member available for the injector readiness gate. */
|
/** Mark a connected spawned member available for the injector readiness gate. */
|
||||||
static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) {
|
static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) {
|
||||||
if (caller.isSpawnedMember()) {
|
if (caller.isSpawnedMember()) {
|
||||||
@@ -741,6 +797,46 @@ public final class FleetMcp {
|
|||||||
return server.listTools();
|
return server.listTools();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 B3 — same reason as {@link #registeredTools()}: a test that must drive the REAL
|
||||||
|
* {@link QuarantineSource} (and the real {@link BackendQuarantine} it wraps) this daemon was
|
||||||
|
* assembled with, rather than scraping {@code FleetdAssembly.java}'s source text for the
|
||||||
|
* constructor call that built it. Unlike {@link #registeredTools()}'s callers, that test cannot
|
||||||
|
* live in this package: it also builds the {@code ResourcePorts} that drives
|
||||||
|
* {@code FleetdAssembly.assembleAndStart}, and {@code ResourcePorts}' methods return
|
||||||
|
* {@code Fleetd}-nested types that are only visible from package {@code dev.ltms.fleet} — so
|
||||||
|
* this accessor is {@code public}, not package-private, to stay reachable from there. {@code
|
||||||
|
* FleetdBackendQuarantineAssemblyTest} quarantines a credential twice through this exact
|
||||||
|
* instance and checks the second cooldown is longer than the first — the one behavioural
|
||||||
|
* difference {@link BackendQuarantine#withEscalation} and the flat two-argument constructor
|
||||||
|
* actually produce.
|
||||||
|
*/
|
||||||
|
public QuarantineSource quarantineSource() {
|
||||||
|
return quarantine;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 B3 — as {@link #quarantineSource()}, {@code public} for the same cross-package
|
||||||
|
* reason, for the real {@link LeadSeatSource} this daemon was assembled with. {@code
|
||||||
|
* FleetdLeadSeatAssemblyTest} calls {@code seatsFor} on this exact instance and checks it
|
||||||
|
* reports a live lead's seat, which {@link LeadSeatSource#none()} can never do (it is a
|
||||||
|
* constant-zero function regardless of input).
|
||||||
|
*/
|
||||||
|
public LeadSeatSource leadSeatSource() {
|
||||||
|
return leadSeats;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 B3 — as {@link #quarantineSource()}, {@code public} for the same cross-package
|
||||||
|
* reason, for the real {@link LeadRollover} (or {@code null}) this daemon was assembled with.
|
||||||
|
* {@code FleetdLeadRolloverAssemblyTest} drives {@code open}/{@code confirm} on this exact
|
||||||
|
* instance and waits for the real continuation to send {@code /clear} and {@code bootstrapText}
|
||||||
|
* through the real {@code router.leadAgents()}.
|
||||||
|
*/
|
||||||
|
public LeadRollover leadRollover() {
|
||||||
|
return leadRollover;
|
||||||
|
}
|
||||||
|
|
||||||
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
|
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -1248,14 +1344,16 @@ public final class FleetMcp {
|
|||||||
Map<String, Object> args) {
|
Map<String, Object> args) {
|
||||||
String action = str(args, "action");
|
String action = str(args, "action");
|
||||||
if (isBlank(action)) {
|
if (isBlank(action)) {
|
||||||
return error("action is required: \"open\", \"confirm\" or \"cancel\"");
|
return error("action is required: \"open\", \"confirm\", \"cancel\" or \"status\"");
|
||||||
}
|
}
|
||||||
return switch (action) {
|
return switch (action) {
|
||||||
case "open" -> handoverOpen(leadRollover, callerTerminal, str(args, "reason"));
|
case "open" -> handoverOpen(leadRollover, callerTerminal, str(args, "reason"));
|
||||||
case "confirm" -> handoverConfirm(leadRollover, callerTerminal, str(args, "token"),
|
case "confirm" -> handoverConfirm(leadRollover, callerTerminal, str(args, "token"),
|
||||||
truthy(args, "operatorConfirmed"));
|
truthy(args, "operatorConfirmed"));
|
||||||
case "cancel" -> handoverCancel(leadRollover, str(args, "token"));
|
case "cancel" -> handoverCancel(leadRollover, str(args, "token"));
|
||||||
default -> error("unknown action \"" + action + "\" — must be \"open\", \"confirm\" or \"cancel\"");
|
case "status" -> handoverStatus(leadRollover, str(args, "token"));
|
||||||
|
default -> error("unknown action \"" + action
|
||||||
|
+ "\" — must be \"open\", \"confirm\", \"cancel\" or \"status\"");
|
||||||
};
|
};
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1325,6 +1423,29 @@ public final class FleetMcp {
|
|||||||
return text(json(m));
|
return text(json(m));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* {@code action: "status"}. Read-only — see {@link LeadRollover#status}: never schedules,
|
||||||
|
* cancels, or retries anything, and it is the only way for a lead to find out what happened to
|
||||||
|
* a token past {@code confirm()}, since every outcome after that point is otherwise logged only
|
||||||
|
* (see {@link LeadRollover}'s class javadoc).
|
||||||
|
*/
|
||||||
|
private static McpSchema.CallToolResult handoverStatus(LeadRollover leadRollover, String token) {
|
||||||
|
if (leadRollover == null) {
|
||||||
|
Map<String, Object> m = new LinkedHashMap<>();
|
||||||
|
m.put("state", "NOT_CONFIGURED");
|
||||||
|
m.put("detail", "leadRollover: is not configured");
|
||||||
|
return text(json(m));
|
||||||
|
}
|
||||||
|
if (isBlank(token)) {
|
||||||
|
return error("token is required for action \"status\"");
|
||||||
|
}
|
||||||
|
LeadRollover.RollStatus s = leadRollover.status(token);
|
||||||
|
Map<String, Object> m = new LinkedHashMap<>();
|
||||||
|
m.put("state", s.state().name());
|
||||||
|
m.put("detail", s.detail());
|
||||||
|
return text(json(m));
|
||||||
|
}
|
||||||
|
|
||||||
/** The one shared {@code NOT_CONFIGURED} refusal shape for {@code open}/{@code confirm}. */
|
/** The one shared {@code NOT_CONFIGURED} refusal shape for {@code open}/{@code confirm}. */
|
||||||
private static McpSchema.CallToolResult notConfigured() {
|
private static McpSchema.CallToolResult notConfigured() {
|
||||||
return refusalJson(false, "NOT_CONFIGURED", "leadRollover: is not configured");
|
return refusalJson(false, "NOT_CONFIGURED", "leadRollover: is not configured");
|
||||||
@@ -1601,7 +1722,7 @@ public final class FleetMcp {
|
|||||||
Map<String, String> leads, String selfTerm,
|
Map<String, String> leads, String selfTerm,
|
||||||
CoordinationSource coordination) {
|
CoordinationSource coordination) {
|
||||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine,
|
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine,
|
||||||
OutageSource.none(), LeadSeatSource.none(), leads, selfTerm, coordination, false);
|
OutageSource.none(), LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
|
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
|
||||||
@@ -1610,7 +1731,7 @@ public final class FleetMcp {
|
|||||||
QuarantineSource quarantine, OutageSource outage,
|
QuarantineSource quarantine, OutageSource outage,
|
||||||
Map<String, String> leads, String selfTerm) {
|
Map<String, String> leads, String selfTerm) {
|
||||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
|
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
|
||||||
LeadSeatSource.none(), leads, selfTerm, CoordinationSource.none(), false);
|
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, CoordinationSource.none(), false);
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -1627,7 +1748,7 @@ public final class FleetMcp {
|
|||||||
QuarantineSource quarantine, Map<String, String> leads, String selfTerm,
|
QuarantineSource quarantine, Map<String, String> leads, String selfTerm,
|
||||||
CoordinationSource coordination) {
|
CoordinationSource coordination) {
|
||||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, OutageSource.none(),
|
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, OutageSource.none(),
|
||||||
LeadSeatSource.none(), leads, selfTerm, coordination, false);
|
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
|
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
|
||||||
@@ -1636,7 +1757,7 @@ public final class FleetMcp {
|
|||||||
QuarantineSource quarantine, OutageSource outage,
|
QuarantineSource quarantine, OutageSource outage,
|
||||||
Map<String, String> leads, String selfTerm, CoordinationSource coordination) {
|
Map<String, String> leads, String selfTerm, CoordinationSource coordination) {
|
||||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
|
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
|
||||||
LeadSeatSource.none(), leads, selfTerm, coordination, false);
|
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -1658,7 +1779,7 @@ public final class FleetMcp {
|
|||||||
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
|
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
|
||||||
CoordinationSource coordination) {
|
CoordinationSource coordination) {
|
||||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
|
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
|
||||||
leadSeats, leads, selfTerm, coordination, false);
|
leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -1682,14 +1803,24 @@ public final class FleetMcp {
|
|||||||
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
|
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
|
||||||
CoordinationSource coordination, boolean callerIsPrimary) {
|
CoordinationSource coordination, boolean callerIsPrimary) {
|
||||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine,
|
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine,
|
||||||
outage, leadSeats, leads, selfTerm, coordination, callerIsPrimary);
|
outage, leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, callerIsPrimary);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The canonical implementation. {@code contextGauge} is the "lead context gauge" (see
|
||||||
|
* {@link LeadContextGauge}) — every wrapper overload above passes a freshly constructed one,
|
||||||
|
* which is correct for them (none of them exercise repeated calls where a shared cache would
|
||||||
|
* matter); the one caller that matters for caching, {@code fleet_list}'s MCP handler, passes
|
||||||
|
* its own single long-lived instance instead (see {@code FleetMcp}'s {@code leadContextGauge}
|
||||||
|
* field).
|
||||||
|
*/
|
||||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||||
LoopHealthSource loopHealth,
|
LoopHealthSource loopHealth,
|
||||||
QuarantineSource quarantine, OutageSource outage,
|
QuarantineSource quarantine, OutageSource outage,
|
||||||
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
|
LeadSeatSource leadSeats, LeadContextGauge contextGauge,
|
||||||
|
LeadConfigDirSource leadConfigDirs,
|
||||||
|
Map<String, String> leads, String selfTerm,
|
||||||
CoordinationSource coordination, boolean callerIsPrimary) {
|
CoordinationSource coordination, boolean callerIsPrimary) {
|
||||||
try {
|
try {
|
||||||
Map<String, Agent> live = workers.list().stream()
|
Map<String, Agent> live = workers.list().stream()
|
||||||
@@ -1698,7 +1829,8 @@ public final class FleetMcp {
|
|||||||
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
|
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
|
||||||
List<Map<String, Object>> leadRows = leads.entrySet().stream()
|
List<Map<String, Object>> leadRows = leads.entrySet().stream()
|
||||||
.sorted(Map.Entry.comparingByValue())
|
.sorted(Map.Entry.comparingByValue())
|
||||||
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm))
|
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm, contextGauge,
|
||||||
|
leadConfigDirs))
|
||||||
.toList();
|
.toList();
|
||||||
// fleetd #209: this is the caller-driven fleet_list read that actually reports
|
// fleetd #209: this is the caller-driven fleet_list read that actually reports
|
||||||
// agentSessionId (via memberCapacityView -> SessionManager.rosterView), so it uses the
|
// agentSessionId (via memberCapacityView -> SessionManager.rosterView), so it uses the
|
||||||
@@ -1993,15 +2125,25 @@ public final class FleetMcp {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* One lead's row: its address, its name, and whether it can be reached right now.
|
* One lead's row: its address, its name, whether it can be reached right now, and how full its
|
||||||
|
* own Claude Code context window is.
|
||||||
*
|
*
|
||||||
* <p>{@code status} is herdr's live view, and {@code unknown} when herdr is not tracking that
|
* <p>{@code status} is herdr's live view, and {@code unknown} when herdr is not tracking that
|
||||||
* pane as an agent — the honest answer, and the one that matters: a lead whose pane herdr cannot
|
* pane as an agent — the honest answer, and the one that matters: a lead whose pane herdr cannot
|
||||||
* see is a lead a {@code fleet_send} cannot be typed into. It is reported rather than hidden,
|
* see is a lead a {@code fleet_send} cannot be typed into. It is reported rather than hidden,
|
||||||
* because a peer that has gone unreachable is exactly what the sender needs to know.
|
* because a peer that has gone unreachable is exactly what the sender needs to know.
|
||||||
|
*
|
||||||
|
* <p>{@code context} is the "lead context gauge" (fleetd's context-usage visibility ticket):
|
||||||
|
* {@code {state: "ok"|"high"|"unknown", tokens?: number, compactions: number}}, read from the
|
||||||
|
* lead's transcript — see {@link LeadContextGauge}. Kept small on purpose (a token count, a
|
||||||
|
* state, and a compaction count) rather than echoing the whole reading history: this is a
|
||||||
|
* roster row a caller glances at, not a diagnostics dump. {@code tokens} is present only when
|
||||||
|
* {@code state} is not {@code "unknown"} — never a stale or default number standing in for "I
|
||||||
|
* could not tell".
|
||||||
*/
|
*/
|
||||||
private static Map<String, Object> leadView(String terminal, String name, Agent live,
|
private static Map<String, Object> leadView(String terminal, String name, Agent live,
|
||||||
String selfTerm) {
|
String selfTerm, LeadContextGauge contextGauge,
|
||||||
|
LeadConfigDirSource leadConfigDirs) {
|
||||||
Map<String, Object> m = new LinkedHashMap<>();
|
Map<String, Object> m = new LinkedHashMap<>();
|
||||||
m.put("sessionId", terminal);
|
m.put("sessionId", terminal);
|
||||||
m.put("name", name);
|
m.put("name", name);
|
||||||
@@ -2010,9 +2152,32 @@ public final class FleetMcp {
|
|||||||
if (terminal.equals(selfTerm)) {
|
if (terminal.equals(selfTerm)) {
|
||||||
m.put("self", true);
|
m.put("self", true);
|
||||||
}
|
}
|
||||||
|
String configDir = leadConfigDirs.configDirFor().apply(name);
|
||||||
|
m.put("context", contextView(contextGauge, live, configDir));
|
||||||
return m;
|
return m;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reads the lead context gauge for one lead. {@code configDir} is this lead's configured
|
||||||
|
* {@code CLAUDE_CONFIG_DIR} override (see {@link LeadConfigDirSource}), derived from
|
||||||
|
* {@code fleet.leaders.<name>.profile} → that profile's own {@code configDir:} — or {@code null}
|
||||||
|
* when the lead's entry names no profile, or that profile sets no override, in which case
|
||||||
|
* {@link LeadContextGauge#read} falls back to its own built-in default
|
||||||
|
* ({@code <user.home>/.claude}).
|
||||||
|
*/
|
||||||
|
private static Map<String, Object> contextView(LeadContextGauge contextGauge, Agent live, String configDir) {
|
||||||
|
String sessionId = live == null ? null : live.sessionId();
|
||||||
|
String agentType = live == null ? null : live.agentType();
|
||||||
|
LeadContextGauge.Reading reading = contextGauge.read(configDir, sessionId, agentType);
|
||||||
|
Map<String, Object> c = new LinkedHashMap<>();
|
||||||
|
c.put("state", reading.state().name().toLowerCase());
|
||||||
|
if (reading.tokens() != null) {
|
||||||
|
c.put("tokens", reading.tokens());
|
||||||
|
}
|
||||||
|
c.put("compactions", reading.compactions());
|
||||||
|
return c;
|
||||||
|
}
|
||||||
|
|
||||||
/** {@code fleet_stop}: tear a worker down by its pane id. */
|
/** {@code fleet_stop}: tear a worker down by its pane id. */
|
||||||
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
|
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
|
||||||
if (isBlank(paneId)) {
|
if (isBlank(paneId)) {
|
||||||
@@ -2134,8 +2299,9 @@ public final class FleetMcp {
|
|||||||
return tool(FleetTool.SPAWN.wireName(),
|
return tool(FleetTool.SPAWN.wireName(),
|
||||||
"Spawn a new off-subscription member session. A member has two independent attributes: "
|
"Spawn a new off-subscription member session. A member has two independent attributes: "
|
||||||
+ "role (what it is for) and profile (which backend it runs on). Pass role to pick "
|
+ "role (what it is for) and profile (which backend it runs on). Pass role to pick "
|
||||||
+ "the contract — 'dev' implements a unit and opens its own PR, 'reviewer' reviews a "
|
+ "the contract — 'dev' implements a unit and opens its own PR, 'hunter' sweeps a "
|
||||||
+ "diff it did not write, 'architect' refines a ticket before anyone builds it; omit "
|
+ "scope without changing it, 'reviewer' reviews a diff it did not write, 'architect' "
|
||||||
|
+ "refines a ticket before anyone builds it; omit "
|
||||||
+ "it for 'dev'. Pass profile (from fleet_profiles) to pick the backend, or omit it "
|
+ "it for 'dev'. Pass profile (from fleet_profiles) to pick the backend, or omit it "
|
||||||
+ "for the default. The two are independent: a reviewer may run on the same profile "
|
+ "for the default. The two are independent: a reviewer may run on the same profile "
|
||||||
+ "as the dev it reviews. The member opens your current directory by default; pass "
|
+ "as the dev it reviews. The member opens your current directory by default; pass "
|
||||||
@@ -2152,7 +2318,7 @@ public final class FleetMcp {
|
|||||||
+ "one. Returns the member's sessionId (use with fleet_send) and paneId (use with "
|
+ "one. Returns the member's sessionId (use with fleet_send) and paneId (use with "
|
||||||
+ "fleet_stop).",
|
+ "fleet_stop).",
|
||||||
objectSchema(Map.of(
|
objectSchema(Map.of(
|
||||||
"role", stringProp("What the member is for: architect, dev or reviewer (default dev)"),
|
"role", stringProp("What the member is for: architect, dev, hunter, or reviewer (default dev)"),
|
||||||
"profile", stringProp("Which backend to run it on (omit for the default profile)"),
|
"profile", stringProp("Which backend to run it on (omit for the default profile)"),
|
||||||
"cwd", stringProp("Working directory for the member (omit to inherit yours)"),
|
"cwd", stringProp("Working directory for the member (omit to inherit yours)"),
|
||||||
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
|
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
|
||||||
@@ -2264,20 +2430,26 @@ public final class FleetMcp {
|
|||||||
return tool(FleetTool.HANDOVER.wireName(),
|
return tool(FleetTool.HANDOVER.wireName(),
|
||||||
"Replace your OWN lead session once its context is full: write a handover file, "
|
"Replace your OWN lead session once its context is full: write a handover file, "
|
||||||
+ "then use this to have fleetd clear your pane and bootstrap a fresh lead "
|
+ "then use this to have fleetd clear your pane and bootstrap a fresh lead "
|
||||||
+ "session against it. Three actions: 'open' (requests a token and the "
|
+ "session against it. Four actions: 'open' (requests a token and the "
|
||||||
+ "handoverPath you must write the handover file to before confirming), "
|
+ "handoverPath you must write the handover file to before confirming), "
|
||||||
+ "'confirm' (validates every gate and — only if every one passes — schedules "
|
+ "'confirm' (validates every gate and — only if every one passes — schedules "
|
||||||
+ "the roll; it does NOT itself clear the pane, the roll runs once this call's "
|
+ "the roll; it does NOT itself clear the pane, the roll runs once this call's "
|
||||||
+ "own turn ends), and 'cancel' (drops a pending request without rolling). "
|
+ "own turn ends), 'cancel' (drops a pending request without rolling), and "
|
||||||
+ "Primary-only. There is deliberately no terminal/session/leadTerminal "
|
+ "'status' (read-only: what happened to a token after 'confirm' — still "
|
||||||
+ "parameter: the pane to roll is always resolved from YOUR OWN connection, "
|
+ "running (approved but not finished yet), the roll completed, the calling "
|
||||||
+ "never a value you pass, so you can only ever roll yourself — never another "
|
+ "turn never settled within turnSettleSeconds so no /clear was ever sent, or "
|
||||||
+ "lead. Requires leadRollover: to be configured; when it is not, every action "
|
+ "/clear itself never settled so bootstrapText was never sent; never "
|
||||||
+ "returns a clean refusal naming NOT_CONFIGURED instead of failing.",
|
+ "schedules, cancels or retries anything). Primary-only. "
|
||||||
|
+ "There is deliberately no terminal/session/leadTerminal parameter: the pane "
|
||||||
|
+ "to roll is always resolved from YOUR OWN connection, never a value you "
|
||||||
|
+ "pass, so you can only ever roll yourself — never another lead. Requires "
|
||||||
|
+ "leadRollover: to be configured; when it is not, every action returns a "
|
||||||
|
+ "clean refusal naming NOT_CONFIGURED instead of failing.",
|
||||||
objectSchema(Map.of(
|
objectSchema(Map.of(
|
||||||
"action", stringProp("\"open\", \"confirm\" or \"cancel\""),
|
"action", stringProp("\"open\", \"confirm\", \"cancel\" or \"status\""),
|
||||||
"reason", stringProp("Free-text audit note for \"open\" (optional, logged only)"),
|
"reason", stringProp("Free-text audit note for \"open\" (optional, logged only)"),
|
||||||
"token", stringProp("The token \"open\" returned — required for \"confirm\" and \"cancel\""),
|
"token", stringProp("The token \"open\" returned — required for \"confirm\", "
|
||||||
|
+ "\"cancel\" and \"status\""),
|
||||||
"operatorConfirmed", Map.of("type", "boolean",
|
"operatorConfirmed", Map.of("type", "boolean",
|
||||||
"description", "For \"confirm\": your answer to \"has the human operator "
|
"description", "For \"confirm\": your answer to \"has the human operator "
|
||||||
+ "confirmed this wipe\" (default false; only consulted when "
|
+ "confirmed this wipe\" (default false; only consulted when "
|
||||||
|
|||||||
@@ -8,6 +8,7 @@ import dev.ltms.fleet.guard.SubscriptionGuard;
|
|||||||
import dev.ltms.fleet.herdr.Agent;
|
import dev.ltms.fleet.herdr.Agent;
|
||||||
import dev.ltms.fleet.herdr.AgentControl;
|
import dev.ltms.fleet.herdr.AgentControl;
|
||||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||||
|
import dev.ltms.fleet.launch.ClaudeCodeArguments;
|
||||||
import dev.ltms.fleet.peer.Capability;
|
import dev.ltms.fleet.peer.Capability;
|
||||||
import dev.ltms.fleet.peer.PeerLauncher;
|
import dev.ltms.fleet.peer.PeerLauncher;
|
||||||
import org.slf4j.Logger;
|
import org.slf4j.Logger;
|
||||||
@@ -298,7 +299,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
|||||||
// has neither MCP nor a charter — session flags must be added into a list we own.
|
// has neither MCP nor a charter — session flags must be added into a list we own.
|
||||||
List<String> argv = mutableArgv(argvWithFleet(cfg, spec));
|
List<String> argv = mutableArgv(argvWithFleet(cfg, spec));
|
||||||
String agentSessionId = applySessionIdentity(argv, spec.sessionName(), spec.resumeSessionId());
|
String agentSessionId = applySessionIdentity(argv, spec.sessionName(), spec.resumeSessionId());
|
||||||
return new Launch(workerEnv, argvWithAutoCompact(argvWithModel(argv, cfg), cfg), agentSessionId);
|
return new Launch(workerEnv, ClaudeCodeArguments.withAutoCompactWindow(argvWithModel(argv, cfg), cfg), agentSessionId);
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -908,30 +909,6 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
|||||||
return withModel;
|
return withModel;
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
|
||||||
* Pin a bounded auto-compaction window on the command line via {@code --autocompact <tokens>},
|
|
||||||
* opt-in per profile (CB-634's sibling ticket: a member that runs out of context dies mid-turn
|
|
||||||
* and its {@code fleet_reply} — the whole point of the turn — is lost with it; opencode already
|
|
||||||
* forces {@code compaction.auto: true} unconditionally, CB-523, but Claude Code has no equivalent
|
|
||||||
* and runs at the backend's own default window).
|
|
||||||
*
|
|
||||||
* <p>Mirrors {@link #argvWithModel}: appended after it, so it survives the {@code ccs <profile>}
|
|
||||||
* wrapper the same way {@code --model} does, and outranks env/settings and the operator's own
|
|
||||||
* {@code argv}. Verified: {@code claude 2.1.241 --help} lists {@code --autocompact <auto|tokens>}
|
|
||||||
* (either the literal {@code auto}, or an integer 100k–1M) — {@link FleetConfig#load} rejects a
|
|
||||||
* configured value outside that band before this ever runs, so the flag Claude Code receives here
|
|
||||||
* is always in range.
|
|
||||||
*/
|
|
||||||
private static List<String> argvWithAutoCompact(List<String> argv, FleetConfig.Profile cfg) {
|
|
||||||
if (cfg.autoCompactWindow() == null) {
|
|
||||||
return argv;
|
|
||||||
}
|
|
||||||
List<String> withAutoCompact = mutableArgv(argv);
|
|
||||||
withAutoCompact.add("--autocompact");
|
|
||||||
withAutoCompact.add(String.valueOf(cfg.autoCompactWindow()));
|
|
||||||
return withAutoCompact;
|
|
||||||
}
|
|
||||||
|
|
||||||
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
|
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
|
||||||
|
|
||||||
/** Spawn a worker for the default profile in the resolved default cwd. */
|
/** Spawn a worker for the default profile in the resolved default cwd. */
|
||||||
|
|||||||
@@ -0,0 +1,24 @@
|
|||||||
|
package dev.ltms.fleet.msg;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 A-gaps (gap 1): a {@link LeadChannel} that its owner can also close.
|
||||||
|
*
|
||||||
|
* <p>{@link LeadChannel}'s own javadoc says plainly that {@code close()} is deliberately left out
|
||||||
|
* of that interface — draining is a caller convenience nobody uses, and closing is the
|
||||||
|
* <em>owner's</em> job. This interface is that owner's own, wider view: whoever opens the
|
||||||
|
* coordination mailbox (the assembly that builds the daemon) also needs to close it from the
|
||||||
|
* shutdown path, and a test standing in for a real broker connection needs a fake it can mark
|
||||||
|
* closed, without ever holding a live connection. Every ordinary consumer ({@code FleetMcp},
|
||||||
|
* {@link LeadCoordLoop}) keeps taking the narrower {@link LeadChannel} exactly as before — only
|
||||||
|
* the owner speaks this wider one.
|
||||||
|
*
|
||||||
|
* <p>{@link LeadMailbox} is still the only production implementation. This only generalises the
|
||||||
|
* TYPE its owner holds it as (previously the concrete class), so a test can substitute a fake
|
||||||
|
* closeable channel instead of a real AMQP connection.
|
||||||
|
*/
|
||||||
|
public interface LeadChannelHandle extends LeadChannel, AutoCloseable {
|
||||||
|
|
||||||
|
/** Release the underlying connection. Declared with no checked exception, unlike the plain {@link AutoCloseable#close()}. */
|
||||||
|
@Override
|
||||||
|
void close();
|
||||||
|
}
|
||||||
@@ -2,6 +2,7 @@ package dev.ltms.fleet.msg;
|
|||||||
|
|
||||||
import dev.ltms.fleet.herdr.AgentControl;
|
import dev.ltms.fleet.herdr.AgentControl;
|
||||||
import dev.ltms.fleet.herdr.AgentStatus;
|
import dev.ltms.fleet.herdr.AgentStatus;
|
||||||
|
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||||
import dev.ltms.fleet.mcp.PrimaryRegistry;
|
import dev.ltms.fleet.mcp.PrimaryRegistry;
|
||||||
import dev.ltms.fleet.metrics.FleetMetrics;
|
import dev.ltms.fleet.metrics.FleetMetrics;
|
||||||
import dev.ltms.fleet.metrics.Metrics;
|
import dev.ltms.fleet.metrics.Metrics;
|
||||||
@@ -13,6 +14,7 @@ import java.util.ArrayList;
|
|||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.concurrent.ScheduledExecutorService;
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
import java.util.concurrent.TimeUnit;
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.function.Function;
|
||||||
import java.util.function.LongSupplier;
|
import java.util.function.LongSupplier;
|
||||||
import java.util.function.Supplier;
|
import java.util.function.Supplier;
|
||||||
|
|
||||||
@@ -44,6 +46,15 @@ import java.util.function.Supplier;
|
|||||||
* loop stands down, so two competing injections never start two turns in the same pane
|
* loop stands down, so two competing injections never start two turns in the same pane
|
||||||
* (constraint 6).</li>
|
* (constraint 6).</li>
|
||||||
* </ol>
|
* </ol>
|
||||||
|
*
|
||||||
|
* <p><b>fleetd #609 — context-high notice.</b> Optionally ({@code contextHighNudge}, opt-in like the
|
||||||
|
* loop itself), a tick that finds the lead's own {@link LeadContextGauge} reading at {@link
|
||||||
|
* LeadContextGauge.State#HIGH} appends a text notice to whatever nudge it sends, telling the lead to
|
||||||
|
* consider {@code fleet_handover}. This is text only — it never rolls a pane itself. It fires once per
|
||||||
|
* HIGH stretch (a latch, cleared only by a later {@code OK} reading — {@code UNKNOWN} neither sets nor
|
||||||
|
* clears it, since "I could not look" must not be read as "it got better"), and it never spends the
|
||||||
|
* quiet-nudge budget: an idle, quiet, HIGH-context lead is exactly the case {@link Action#QUIET_DONE}
|
||||||
|
* would otherwise swallow, and it is the one case most worth interrupting the quiet cap for.
|
||||||
*/
|
*/
|
||||||
public final class LeadHeartbeatLoop {
|
public final class LeadHeartbeatLoop {
|
||||||
|
|
||||||
@@ -63,11 +74,16 @@ public final class LeadHeartbeatLoop {
|
|||||||
private final long backoffMs;
|
private final long backoffMs;
|
||||||
private final int quietNudgeCap;
|
private final int quietNudgeCap;
|
||||||
private final Metrics metrics; // CB-512 pattern: nullable — no registry in unit tests
|
private final Metrics metrics; // CB-512 pattern: nullable — no registry in unit tests
|
||||||
|
private final LeadContextSource contextSource; // fleetd #609
|
||||||
|
private final boolean contextHighNudge; // fleetd #609: opt-in, like the loop itself
|
||||||
|
private final boolean requireOperatorConfirm; // fleetd #621: mirrors leadRollover.requireOperatorConfirm
|
||||||
|
|
||||||
/** When the current idle stretch began (nanos), or {@link #NOT_IDLE}. Single scheduler thread only. */
|
/** When the current idle stretch began (nanos), or {@link #NOT_IDLE}. Single scheduler thread only. */
|
||||||
private long idleSinceNanos = NOT_IDLE;
|
private long idleSinceNanos = NOT_IDLE;
|
||||||
/** Consecutive nudges that found no pending fleet state. Single scheduler thread only. */
|
/** Consecutive nudges that found no pending fleet state. Single scheduler thread only. */
|
||||||
private int quietCount = 0;
|
private int quietCount = 0;
|
||||||
|
/** fleetd #609: latched "already told this lead about this HIGH stretch". Single scheduler thread only. */
|
||||||
|
private boolean contextNotified = false;
|
||||||
|
|
||||||
/** Constructor with an injectable clock and no metric registry (unit tests, or wiring that opts out). */
|
/** Constructor with an injectable clock and no metric registry (unit tests, or wiring that opts out). */
|
||||||
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
||||||
@@ -75,7 +91,7 @@ public final class LeadHeartbeatLoop {
|
|||||||
ScheduledExecutorService scheduler, LongSupplier clock,
|
ScheduledExecutorService scheduler, LongSupplier clock,
|
||||||
long idleAfterNanos, long backoffMs, int quietNudgeCap) {
|
long idleAfterNanos, long backoffMs, int quietNudgeCap) {
|
||||||
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
|
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
|
||||||
idleAfterNanos, backoffMs, quietNudgeCap, null);
|
idleAfterNanos, backoffMs, quietNudgeCap, null, LeadContextSource.none(), false);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** As above, with a metric registry (the CB-512 pattern) so nudge outcomes are counted. */
|
/** As above, with a metric registry (the CB-512 pattern) so nudge outcomes are counted. */
|
||||||
@@ -83,6 +99,40 @@ public final class LeadHeartbeatLoop {
|
|||||||
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
|
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
|
||||||
ScheduledExecutorService scheduler, LongSupplier clock,
|
ScheduledExecutorService scheduler, LongSupplier clock,
|
||||||
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics) {
|
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics) {
|
||||||
|
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
|
||||||
|
idleAfterNanos, backoffMs, quietNudgeCap, metrics, LeadContextSource.none(), false);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #609: as above, plus the lead's own context source and whether a HIGH reading should
|
||||||
|
* append a hand-over notice to the loop's nudge. Pass {@link LeadContextSource#none()} and
|
||||||
|
* {@code false} to keep the pre-#609 behaviour exactly (both existing public constructors do).
|
||||||
|
*
|
||||||
|
* <p>fleetd #621: delegates to the full constructor with {@code requireOperatorConfirm=true} —
|
||||||
|
* the pre-#621 wording ("ask the operator ... only the operator can approve the roll") assumed
|
||||||
|
* the config default, so every caller of this overload keeps that text byte-identical.
|
||||||
|
*/
|
||||||
|
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
||||||
|
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
|
||||||
|
ScheduledExecutorService scheduler, LongSupplier clock,
|
||||||
|
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics,
|
||||||
|
LeadContextSource contextSource, boolean contextHighNudge) {
|
||||||
|
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
|
||||||
|
idleAfterNanos, backoffMs, quietNudgeCap, metrics, contextSource, contextHighNudge, true);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #621: as above, plus the daemon's effective {@code leadRollover.requireOperatorConfirm}
|
||||||
|
* value — threaded into {@link #contextNotice(boolean, LeadContextGauge.Reading, boolean, boolean)}
|
||||||
|
* so the notice's wording tracks the config the daemon actually enforces (see {@code
|
||||||
|
* LeadRollover.confirm}) instead of always asserting the operator gate is on.
|
||||||
|
*/
|
||||||
|
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
||||||
|
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
|
||||||
|
ScheduledExecutorService scheduler, LongSupplier clock,
|
||||||
|
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics,
|
||||||
|
LeadContextSource contextSource, boolean contextHighNudge,
|
||||||
|
boolean requireOperatorConfirm) {
|
||||||
this.primaryRegistry = primaryRegistry;
|
this.primaryRegistry = primaryRegistry;
|
||||||
this.agents = agents;
|
this.agents = agents;
|
||||||
this.inbox = inbox;
|
this.inbox = inbox;
|
||||||
@@ -94,6 +144,20 @@ public final class LeadHeartbeatLoop {
|
|||||||
this.backoffMs = backoffMs;
|
this.backoffMs = backoffMs;
|
||||||
this.quietNudgeCap = quietNudgeCap;
|
this.quietNudgeCap = quietNudgeCap;
|
||||||
this.metrics = metrics;
|
this.metrics = metrics;
|
||||||
|
this.contextSource = contextSource;
|
||||||
|
this.contextHighNudge = contextHighNudge;
|
||||||
|
this.requireOperatorConfirm = requireOperatorConfirm;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #609: one lead's own context reading, keyed by its terminal id — the same injected-source
|
||||||
|
* idiom {@code FleetMcp.LeadSeatSource}/{@code FleetMcp.LeadConfigDirSource} already use.
|
||||||
|
*/
|
||||||
|
public record LeadContextSource(Function<String, LeadContextGauge.Reading> readingFor) {
|
||||||
|
/** Inert source — every lead reads UNKNOWN, so the context notice can never fire. */
|
||||||
|
public static LeadContextSource none() {
|
||||||
|
return new LeadContextSource(_ -> LeadContextGauge.Reading.unknown());
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -126,8 +190,11 @@ public final class LeadHeartbeatLoop {
|
|||||||
STAND_DOWN
|
STAND_DOWN
|
||||||
}
|
}
|
||||||
|
|
||||||
/** The outcome of one decision: the action plus the state to persist for the next tick. */
|
/**
|
||||||
record Decision(Action action, Long idleSinceNanos, int quietCount) {}
|
* The outcome of one decision: the action, the state to persist for the next tick, and (fleetd
|
||||||
|
* #609) whether the lead has now been told about the current HIGH context stretch.
|
||||||
|
*/
|
||||||
|
record Decision(Action action, Long idleSinceNanos, int quietCount, boolean contextNotified) {}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Pure decision function: given the current loop state and fleet/lead facts, return what to do
|
* Pure decision function: given the current loop state and fleet/lead facts, return what to do
|
||||||
@@ -142,34 +209,47 @@ public final class LeadHeartbeatLoop {
|
|||||||
* @param pushLoopActive whether {@link ReplyPushLoop} is currently nudging some target (constraint 6)
|
* @param pushLoopActive whether {@link ReplyPushLoop} is currently nudging some target (constraint 6)
|
||||||
* @param leadKnown whether a lead terminal is known to nudge at all
|
* @param leadKnown whether a lead terminal is known to nudge at all
|
||||||
* @param fleet a snapshot of the pending fleet state (constraint 5)
|
* @param fleet a snapshot of the pending fleet state (constraint 5)
|
||||||
|
* @param context fleetd #609: the lead's own {@link LeadContextGauge} reading for this tick
|
||||||
|
* @param contextNotified fleetd #609: whether the lead has already been told about the current HIGH
|
||||||
|
* stretch — a latch, carried forward by {@link #applyDecision}
|
||||||
* @return the action to take and the state to persist
|
* @return the action to take and the state to persist
|
||||||
*/
|
*/
|
||||||
Decision decide(long nowNanos, Long idleSinceNanos, int quietCount, AgentStatus status,
|
Decision decide(long nowNanos, Long idleSinceNanos, int quietCount, AgentStatus status,
|
||||||
boolean pushLoopActive, boolean leadKnown, FleetState fleet) {
|
boolean pushLoopActive, boolean leadKnown, FleetState fleet,
|
||||||
|
LeadContextGauge.State context, boolean contextNotified) {
|
||||||
|
// fleetd #609: re-arm the latch only on a positive OK reading. UNKNOWN means "I could not
|
||||||
|
// look", not "it got better" — re-arming on UNKNOWN would let a flapping gauge (a transcript
|
||||||
|
// read that misses one tick) nudge a full lead again on every recovery, defeating the "once
|
||||||
|
// per HIGH stretch" promise. Computed once, up front, so every gate below carries it forward
|
||||||
|
// unchanged unless it is the gate that actually discharges it.
|
||||||
|
boolean latch = context == LeadContextGauge.State.OK ? false : contextNotified;
|
||||||
|
boolean contextHigh = contextHighNudge && context == LeadContextGauge.State.HIGH;
|
||||||
|
|
||||||
// Constraint 6: while ReplyPushLoop is actively nudging the lead, injecting a second,
|
// Constraint 6: while ReplyPushLoop is actively nudging the lead, injecting a second,
|
||||||
// competing prompt into the same pane would start a second turn — racing loops multiply
|
// competing prompt into the same pane would start a second turn — racing loops multiply
|
||||||
// turns and context burn. Stand aside, and treat the active push as real state (re-arm the
|
// turns and context burn. Stand aside, and treat the active push as real state (re-arm the
|
||||||
// quiet counter), because the reply that drove it is exactly the kind of new state that
|
// quiet counter), because the reply that drove it is exactly the kind of new state that
|
||||||
// should reset the cap.
|
// should reset the cap. The context latch is untouched: standing down must not spend the
|
||||||
|
// one notice this stretch gets.
|
||||||
if (pushLoopActive) {
|
if (pushLoopActive) {
|
||||||
return new Decision(Action.STAND_DOWN, idleSinceNanos, 0);
|
return new Decision(Action.STAND_DOWN, idleSinceNanos, 0, latch);
|
||||||
}
|
}
|
||||||
// Constraint 2: a WORKING lead is making progress and must NOT be touched; an unreadable
|
// Constraint 2: a WORKING lead is making progress and must NOT be touched; an unreadable
|
||||||
// status (read failure, or the agent is gone) is safest treated the same way — never inject
|
// status (read failure, or the agent is gone) is safest treated the same way — never inject
|
||||||
// into a state we cannot read. Either way, reset the idle window and the quiet counter: the
|
// into a state we cannot read. Either way, reset the idle window and the quiet counter: the
|
||||||
// lead was / may be active, so the next idle stretch must count its own quiet period fresh.
|
// lead was / may be active, so the next idle stretch must count its own quiet period fresh.
|
||||||
if (status == null || !status.injectable()) {
|
if (status == null || !status.injectable()) {
|
||||||
return new Decision(Action.LEAD_BUSY, null, 0);
|
return new Decision(Action.LEAD_BUSY, null, 0, latch);
|
||||||
}
|
}
|
||||||
if (idleSinceNanos == null) {
|
if (idleSinceNanos == null) {
|
||||||
// The lead just became injectable — record the start of an idle stretch and wait out the
|
// The lead just became injectable — record the start of an idle stretch and wait out the
|
||||||
// debounce quiet period before ever nudging (constraint 3).
|
// debounce quiet period before ever nudging (constraint 3).
|
||||||
return new Decision(Action.WAIT_IDLE, nowNanos, quietCount);
|
return new Decision(Action.WAIT_IDLE, nowNanos, quietCount, latch);
|
||||||
}
|
}
|
||||||
if (nowNanos - idleSinceNanos < idleAfterNanos) {
|
if (nowNanos - idleSinceNanos < idleAfterNanos) {
|
||||||
// Still within the quiet period: the lead that just finished a turn sits momentarily idle
|
// Still within the quiet period: the lead that just finished a turn sits momentarily idle
|
||||||
// and must not be re-prompted into every natural pause.
|
// and must not be re-prompted into every natural pause.
|
||||||
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
|
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount, latch);
|
||||||
}
|
}
|
||||||
// Past the quiet period with an injectable lead: it is a genuine candidate for a nudge. Two
|
// Past the quiet period with an injectable lead: it is a genuine candidate for a nudge. Two
|
||||||
// gating facts decide whether and how:
|
// gating facts decide whether and how:
|
||||||
@@ -177,23 +257,33 @@ public final class LeadHeartbeatLoop {
|
|||||||
// No lead terminal is known yet (e.g. the registry has not learned one) — there is nobody
|
// No lead terminal is known yet (e.g. the registry has not learned one) — there is nobody
|
||||||
// to nudge. Keep waiting; the window stays open so discovery re-arms it without a fresh
|
// to nudge. Keep waiting; the window stays open so discovery re-arms it without a fresh
|
||||||
// quiet period.
|
// quiet period.
|
||||||
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
|
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount, latch);
|
||||||
}
|
}
|
||||||
if (fleet.hasPending()) {
|
if (fleet.hasPending()) {
|
||||||
// Real fleet state is waiting — a worker reply or a DONE session. This is new state, so
|
// Real fleet state is waiting — a worker reply or a DONE session. This is new state, so
|
||||||
// it resets the quiet counter (constraint 4) and the lead is nudged to go collect it.
|
// it resets the quiet counter (constraint 4) and the lead is nudged to go collect it. The
|
||||||
return new Decision(Action.INJECT, idleSinceNanos, 0);
|
// nudge text carries the context notice too when contextHigh — see injectNudge/contextNotice
|
||||||
|
// — so this route discharges the same duty and must set the latch.
|
||||||
|
return new Decision(Action.INJECT, idleSinceNanos, 0, latch || contextHigh);
|
||||||
|
}
|
||||||
|
if (contextHigh && !latch) {
|
||||||
|
// fleetd #609: the lead is idle, its context is full, and nothing is pending. This is the
|
||||||
|
// one case the quiet cap would otherwise swallow, and it is exactly when the lead most
|
||||||
|
// needs to hear it. Fire once per HIGH stretch, and do NOT spend the quiet budget on it:
|
||||||
|
// this is an event notice, not a "are you still there" nudge.
|
||||||
|
return new Decision(Action.INJECT, idleSinceNanos, quietCount, true);
|
||||||
}
|
}
|
||||||
if (quietCount < quietNudgeCap) {
|
if (quietCount < quietNudgeCap) {
|
||||||
// Nothing is pending, but the cap is not exhausted: nudge anyway, telling the lead
|
// Nothing is pending, but the cap is not exhausted: nudge anyway, telling the lead
|
||||||
// exactly that nothing is waiting so it can choose to stand down rather than hunt
|
// exactly that nothing is waiting so it can choose to stand down rather than hunt
|
||||||
// (constraint 5). Count it toward the consecutive-quiet cap.
|
// (constraint 5). Count it toward the consecutive-quiet cap. This nudge also carries the
|
||||||
return new Decision(Action.INJECT, idleSinceNanos, quietCount + 1);
|
// context notice when contextHigh (already latched above, or being latched now).
|
||||||
|
return new Decision(Action.INJECT, idleSinceNanos, quietCount + 1, latch || contextHigh);
|
||||||
}
|
}
|
||||||
// Nothing pending and the cap is exhausted: stop nudging until real state appears again
|
// Nothing pending and the cap is exhausted: stop nudging until real state appears again
|
||||||
// (constraint 4). The loop still ticks on backoff so a genuinely new reply or session change
|
// (constraint 4). The loop still ticks on backoff so a genuinely new reply or session change
|
||||||
// re-arms it — QUIET_DONE stops injection, not observation.
|
// re-arms it — QUIET_DONE stops injection, not observation.
|
||||||
return new Decision(Action.QUIET_DONE, idleSinceNanos, quietCount);
|
return new Decision(Action.QUIET_DONE, idleSinceNanos, quietCount, latch);
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- loop ----------------------------------------------------------------------------------
|
// --- loop ----------------------------------------------------------------------------------
|
||||||
@@ -203,62 +293,195 @@ public final class LeadHeartbeatLoop {
|
|||||||
* does not evaluate the lead's idle state before the fleet has settled.
|
* does not evaluate the lead's idle state before the fleet has settled.
|
||||||
*/
|
*/
|
||||||
public void start() {
|
public void start() {
|
||||||
log.info("idle-lead heartbeat: on — nudge lead after {}s idle (recheck {}ms, quiet cap {})",
|
// fleetd #613: contextHighNudge added alongside the three settings already here — an
|
||||||
TimeUnit.NANOSECONDS.toSeconds(idleAfterNanos), backoffMs, quietNudgeCap);
|
// operator otherwise cannot tell from the boot log whether the #609 handover notice is
|
||||||
|
// armed, and had to load the deployed jar's config to confirm it.
|
||||||
|
log.info("idle-lead heartbeat: on — nudge lead after {}s idle (recheck {}ms, quiet cap {}, "
|
||||||
|
+ "context-high nudge {})",
|
||||||
|
TimeUnit.NANOSECONDS.toSeconds(idleAfterNanos), backoffMs, quietNudgeCap,
|
||||||
|
contextHighNudge);
|
||||||
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
|
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
|
||||||
}
|
}
|
||||||
|
|
||||||
/** One loop tick, every {@link #backoffMs} — the thin scheduler around {@link #decide}. */
|
/** One loop tick, every {@link #backoffMs} — the thin scheduler around {@link #decide}. Package-private
|
||||||
private void tick() {
|
* (mirroring {@link ReplyPushLoop#tick(String)}) so tests can drive it directly with a fake clock and a
|
||||||
|
* fake {@link AgentControl} instead of racing the scheduler thread. */
|
||||||
|
void tick() {
|
||||||
boolean leadKnown = primaryRegistry.primaryTerminal().isPresent();
|
boolean leadKnown = primaryRegistry.primaryTerminal().isPresent();
|
||||||
FleetState fleet = snapshot(inbox, roster);
|
FleetState fleet = snapshot(inbox, roster);
|
||||||
AgentStatus status = AgentStatus.UNKNOWN;
|
AgentStatus status = AgentStatus.UNKNOWN;
|
||||||
|
LeadContextGauge.Reading reading = LeadContextGauge.Reading.unknown();
|
||||||
if (leadKnown) {
|
if (leadKnown) {
|
||||||
|
String leadTerminal = primaryRegistry.primaryTerminal().orElseThrow();
|
||||||
try {
|
try {
|
||||||
status = agents.status(primaryRegistry.primaryTerminal().orElseThrow());
|
status = agents.status(leadTerminal);
|
||||||
} catch (RuntimeException e) {
|
} catch (RuntimeException e) {
|
||||||
// A failed status read degrades to "unknown" — decide() treats that like a busy lead
|
// A failed status read degrades to "unknown" — decide() treats that like a busy lead
|
||||||
// and never injects into a state it cannot read. Retry on the next backoff.
|
// and never injects into a state it cannot read. Retry on the next backoff.
|
||||||
log.debug("idle-heartbeat: status check failed for lead, will retry: {}", e.toString());
|
log.debug("idle-heartbeat: status check failed for lead, will retry: {}", e.toString());
|
||||||
}
|
}
|
||||||
|
// fleetd #609: read the lead's own context regardless of status — decide() still gates on
|
||||||
|
// status first (constraint 2), so this is harmless work on a WORKING lead and lets the
|
||||||
|
// latch state stay accurate for whenever the lead does go idle.
|
||||||
|
reading = contextSource.readingFor().apply(leadTerminal);
|
||||||
}
|
}
|
||||||
|
|
||||||
Decision d = decide(clock.getAsLong(),
|
Decision d = decide(clock.getAsLong(),
|
||||||
idleSinceNanos == NOT_IDLE ? null : idleSinceNanos,
|
idleSinceNanos == NOT_IDLE ? null : idleSinceNanos,
|
||||||
quietCount, status, pushLoop.isActive(), leadKnown, fleet);
|
quietCount, status, pushLoop.isActive(), leadKnown, fleet,
|
||||||
|
reading.state(), contextNotified);
|
||||||
applyDecision(d);
|
applyDecision(d);
|
||||||
switch (d.action()) {
|
switch (d.action()) {
|
||||||
case INJECT -> injectNudge(fleet);
|
case INJECT -> injectNudge(d, fleet, reading);
|
||||||
case QUIET_DONE -> countNudge("exhausted");
|
case QUIET_DONE -> {
|
||||||
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN -> { /* nothing to inject, nothing to count */ }
|
countNudge("exhausted");
|
||||||
|
contextNotified = d.contextNotified();
|
||||||
|
}
|
||||||
|
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN -> contextNotified = d.contextNotified();
|
||||||
}
|
}
|
||||||
scheduleNext();
|
scheduleNext();
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Persist the state a decision returned, so the next tick starts from it. */
|
/**
|
||||||
|
* Persist the idle/quiet state a decision returned, so the next tick starts from it.
|
||||||
|
*
|
||||||
|
* <p>fleetd #609 review: the context latch ({@link #contextNotified}) is deliberately <em>not</em>
|
||||||
|
* set here any more. Setting it from the decision unconditionally — before {@link #injectNudge} even
|
||||||
|
* tries to send — is exactly the review's blocker: a decision to notify is not the same fact as "the
|
||||||
|
* notice reached the pane". Every branch of {@link #tick} now assigns {@link #contextNotified} itself,
|
||||||
|
* once it knows whether a send happened and whether it carried the notice (see {@link #injectNudge}).
|
||||||
|
*/
|
||||||
private void applyDecision(Decision d) {
|
private void applyDecision(Decision d) {
|
||||||
idleSinceNanos = d.idleSinceNanos() == null ? NOT_IDLE : d.idleSinceNanos();
|
idleSinceNanos = d.idleSinceNanos() == null ? NOT_IDLE : d.idleSinceNanos();
|
||||||
quietCount = d.quietCount();
|
quietCount = d.quietCount();
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Send the nudge to the known lead. */
|
/**
|
||||||
private void injectNudge(FleetState fleet) {
|
* Send the nudge to the known lead, with the fleetd #609 context notice appended when it applies, and
|
||||||
|
* persist the context latch based on what actually happened this tick — not merely what {@code d}
|
||||||
|
* chose to attempt.
|
||||||
|
*/
|
||||||
|
private void injectNudge(Decision d, FleetState fleet, LeadContextGauge.Reading reading) {
|
||||||
|
// fleetd #609 review: build the notice from the latch as it stood BEFORE this tick's decision —
|
||||||
|
// d.contextNotified() is the value to persist once delivery is confirmed, not the value the text
|
||||||
|
// itself should be built from. Otherwise a HIGH stretch that is still latched would never see the
|
||||||
|
// notice at all, defeating the very check this fixes.
|
||||||
|
String notice = contextNotice(contextHighNudge, reading, contextNotified, requireOperatorConfirm);
|
||||||
var lead = primaryRegistry.primaryTerminal();
|
var lead = primaryRegistry.primaryTerminal();
|
||||||
if (lead.isEmpty()) {
|
boolean sent = lead.isPresent() && trySend(lead.get(), fleet.nudgeText() + notice, notice);
|
||||||
return; // the lead disappeared between the decision and the injection
|
// The latch becomes true only when all three hold: decide() chose to notify, a notice was
|
||||||
}
|
// actually included in the text, and the send reached the pane without throwing. Whenever no
|
||||||
String leadTerminal = lead.get();
|
// notice was attempted (disabled, not HIGH, or already latched), nothing was promised to the lead
|
||||||
|
// this tick, so apply the decision's own carried-forward value unconditionally — that is how the
|
||||||
|
// OK-only re-arm rule and STAND_DOWN's "don't burn the notice" rule keep working through this path
|
||||||
|
// too. A lead that disappeared between the decision and the send (lead.isEmpty()) is treated the
|
||||||
|
// same as a failed send: nothing reached the pane, so the latch must not be set.
|
||||||
|
contextNotified = notice.isEmpty() ? d.contextNotified() : sent;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Attempt one herdr send and count its outcome. Returns whether {@code agents.send} returned without
|
||||||
|
* throwing — the caller ({@link #injectNudge}) needs this to decide whether the fleetd #609 context
|
||||||
|
* latch may be persisted as set.
|
||||||
|
*/
|
||||||
|
private boolean trySend(String leadTerminal, String text, String notice) {
|
||||||
try {
|
try {
|
||||||
agents.send(leadTerminal, fleet.nudgeText());
|
agents.send(leadTerminal, text);
|
||||||
log.debug("idle-heartbeat: nudge sent to lead {} (quiet nudges so far in this stretch: {})",
|
log.debug("idle-heartbeat: nudge sent to lead {} (quiet nudges so far in this stretch: {})",
|
||||||
leadTerminal, quietCount);
|
leadTerminal, quietCount);
|
||||||
countNudge("sent");
|
// fleetd #609: a nudge that carries the context notice is counted under its own outcome so
|
||||||
|
// it is visible in /metrics — one count per nudge either way, never two.
|
||||||
|
countNudge(notice.isEmpty() ? "sent" : "sent_context");
|
||||||
|
return true;
|
||||||
} catch (RuntimeException e) {
|
} catch (RuntimeException e) {
|
||||||
log.warn("idle-heartbeat: failed to nudge lead {}: {}", leadTerminal, e.toString());
|
log.warn("idle-heartbeat: failed to nudge lead {}: {}", leadTerminal, e.toString());
|
||||||
countNudge("failed");
|
countNudge("failed");
|
||||||
|
return false;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #609: the text appended to a nudge when the lead's own context is full — {@code ""}
|
||||||
|
* whenever the notice does not apply, so callers can unconditionally append this without an extra
|
||||||
|
* branch. Wording stays plain (CEFR B1) and honest about who actually gates the roll — see the
|
||||||
|
* {@code requireOperatorConfirm} overload (fleetd #621) for which check that is. This loop only
|
||||||
|
* ever prints text, it never calls {@code fleet_handover} itself.
|
||||||
|
*
|
||||||
|
* @param enabled the {@code leadHeartbeat.contextHighNudge} config flag
|
||||||
|
* @param reading the lead's current {@link LeadContextGauge} reading
|
||||||
|
* @return the notice text (starting with a leading space, to append directly after {@link
|
||||||
|
* FleetState#nudgeText()}), or {@code ""} when disabled or the state is not {@code HIGH}
|
||||||
|
*/
|
||||||
|
static String contextNotice(boolean enabled, LeadContextGauge.Reading reading) {
|
||||||
|
return contextNotice(enabled, reading, false);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #609 review: as {@link #contextNotice(boolean, LeadContextGauge.Reading)}, but also gated on
|
||||||
|
* {@code alreadyNotified} — the context latch as it stood <em>before</em> the current tick's decision.
|
||||||
|
* Without this gate, every pending-driven {@code INJECT} that lands while the context stays {@code
|
||||||
|
* HIGH} would re-append the full notice on top of an already-latched stretch, making the notice's own
|
||||||
|
* closing sentence ("You will not be told again until your context reads ok.") false. {@link
|
||||||
|
* #injectNudge} is the only caller that passes a non-default {@code alreadyNotified}.
|
||||||
|
*
|
||||||
|
* <p>fleetd #621: delegates with {@code requireOperatorConfirm=true} — the pre-#621 default and the
|
||||||
|
* value every existing caller of this overload (including every test written before #621) already
|
||||||
|
* assumed, so the text this overload returns stays byte-identical.
|
||||||
|
*
|
||||||
|
* @param alreadyNotified whether the lead has already been told about the current HIGH stretch
|
||||||
|
*/
|
||||||
|
static String contextNotice(boolean enabled, LeadContextGauge.Reading reading, boolean alreadyNotified) {
|
||||||
|
return contextNotice(enabled, reading, alreadyNotified, true);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #621: as {@link #contextNotice(boolean, LeadContextGauge.Reading, boolean)}, but the closing
|
||||||
|
* instructions also track the daemon's effective {@code leadRollover.requireOperatorConfirm} value,
|
||||||
|
* instead of always asserting that only the operator can approve the roll.
|
||||||
|
*
|
||||||
|
* <p>{@code LeadRollover.confirm(...)} already honours this flag: when it is {@code false}, the daemon
|
||||||
|
* itself gates the roll on the three handover-file checks alone (exists, modified after the {@code
|
||||||
|
* open()} request, and no older than {@code maxDocAgeSeconds}) and never consults {@code
|
||||||
|
* operatorConfirmed}. Before this parameter existed, this notice told the lead to ask the operator
|
||||||
|
* regardless — so a lead that followed its own instructions asked anyway, and setting the config knob
|
||||||
|
* to {@code false} stopped the daemon refusing the roll without stopping the operator being
|
||||||
|
* interrupted. This parameter is how the text is kept honest about which gate is actually live.
|
||||||
|
*
|
||||||
|
* @param requireOperatorConfirm the effective {@code leadRollover.requireOperatorConfirm} value
|
||||||
|
*/
|
||||||
|
static String contextNotice(boolean enabled, LeadContextGauge.Reading reading, boolean alreadyNotified,
|
||||||
|
boolean requireOperatorConfirm) {
|
||||||
|
if (!enabled || alreadyNotified || reading.state() != LeadContextGauge.State.HIGH) {
|
||||||
|
return "";
|
||||||
|
}
|
||||||
|
StringBuilder sb = new StringBuilder(" Your own context is nearly full");
|
||||||
|
String compactionWord = reading.compactions() == 1 ? "compaction" : "compactions";
|
||||||
|
if (reading.tokens() != null) {
|
||||||
|
sb.append(": ").append(reading.tokens()).append(" tokens used, ")
|
||||||
|
.append(reading.compactions()).append(' ').append(compactionWord).append(" so far.");
|
||||||
|
} else {
|
||||||
|
// A HIGH reading always carries a non-null token count today: LeadContextGauge only
|
||||||
|
// reaches HIGH by comparing a number against HIGH_THRESHOLD_TOKENS. That invariant
|
||||||
|
// lives in another class and nothing asserts it, so this branch does not rely on it —
|
||||||
|
// it drops the token clause rather than printing "null tokens".
|
||||||
|
sb.append(" (").append(reading.compactions()).append(' ').append(compactionWord)
|
||||||
|
.append(" so far).");
|
||||||
|
}
|
||||||
|
if (requireOperatorConfirm) {
|
||||||
|
sb.append(" A fresh session would work better. To hand over: call fleet_handover(action=\"open\"), "
|
||||||
|
+ "write the file it names, ask the operator, then call fleet_handover(action=\"confirm\", "
|
||||||
|
+ "token, operatorConfirmed). Only the operator can approve the roll. You will not be told "
|
||||||
|
+ "again until your context reads ok.");
|
||||||
|
} else {
|
||||||
|
sb.append(" A fresh session would work better. To hand over: call fleet_handover(action=\"open\"), "
|
||||||
|
+ "write the file it names, then call fleet_handover(action=\"confirm\", token). Decide for "
|
||||||
|
+ "yourself when to confirm: the roll goes through if the handover file exists, was "
|
||||||
|
+ "changed after you opened it, and is not older than maxDocAgeSeconds. You will not be "
|
||||||
|
+ "told again until your context reads ok.");
|
||||||
|
}
|
||||||
|
return sb.toString();
|
||||||
|
}
|
||||||
|
|
||||||
/** Schedule the next tick on the scheduler thread pool. */
|
/** Schedule the next tick on the scheduler thread pool. */
|
||||||
private void scheduleNext() {
|
private void scheduleNext() {
|
||||||
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
|
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
|
||||||
|
|||||||
@@ -63,7 +63,7 @@ import java.util.concurrent.TimeoutException;
|
|||||||
* which messages reached a lead. Any publish still awaiting its confirm is failed rather than left to idle out
|
* which messages reached a lead. Any publish still awaiting its confirm is failed rather than left to idle out
|
||||||
* the confirm timeout against a sequence number that means nothing on the new channel.
|
* the confirm timeout against a sequence number that means nothing on the new channel.
|
||||||
*/
|
*/
|
||||||
public final class LeadMailbox implements LeadChannel, AutoCloseable {
|
public final class LeadMailbox implements LeadChannelHandle {
|
||||||
|
|
||||||
private static final Logger log = LoggerFactory.getLogger(LeadMailbox.class);
|
private static final Logger log = LoggerFactory.getLogger(LeadMailbox.class);
|
||||||
|
|
||||||
|
|||||||
@@ -29,8 +29,8 @@ public enum MemberRole {
|
|||||||
* <p>Reads the repo and writes analysis. Never commits code and never opens a pull request —
|
* <p>Reads the repo and writes analysis. Never commits code and never opens a pull request —
|
||||||
* an architect that starts implementing has stopped doing the job that makes it useful.
|
* an architect that starts implementing has stopped doing the job that makes it useful.
|
||||||
*
|
*
|
||||||
* <p>Architects are the one member kind declared in config, because a lead addresses the same
|
* <p>Architects are the one member kind with live slot binding, because a lead addresses the
|
||||||
* slots across many tickets and needs a stable name for them.
|
* same slots across many tickets and needs a stable name for them.
|
||||||
*/
|
*/
|
||||||
ARCHITECT,
|
ARCHITECT,
|
||||||
|
|
||||||
@@ -43,6 +43,14 @@ public enum MemberRole {
|
|||||||
*/
|
*/
|
||||||
DEV,
|
DEV,
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Sweeps an assigned package for defects and reports several ranked findings.
|
||||||
|
*
|
||||||
|
* <p>Never changes code, commits, or opens a pull request. A hunt gathers evidence, which can
|
||||||
|
* include running the build, but leaves every fix to a later implementation unit.
|
||||||
|
*/
|
||||||
|
HUNTER,
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Reviews a diff it did not write and reports one structured finding.
|
* Reviews a diff it did not write and reports one structured finding.
|
||||||
*
|
*
|
||||||
@@ -59,7 +67,7 @@ public enum MemberRole {
|
|||||||
|
|
||||||
/**
|
/**
|
||||||
* The {@code fleet:} block that holds this role's pool — {@code architects},
|
* The {@code fleet:} block that holds this role's pool — {@code architects},
|
||||||
* {@code developers}, {@code reviewers}.
|
* {@code developers}, {@code hunters}, {@code reviewers}.
|
||||||
*
|
*
|
||||||
* <p>Plural, and not always the wire name: the pool of things a {@code dev} may run on reads
|
* <p>Plural, and not always the wire name: the pool of things a {@code dev} may run on reads
|
||||||
* naturally as {@code developers:}. The wire name stays the singular {@code dev}, because that
|
* naturally as {@code developers:}. The wire name stays the singular {@code dev}, because that
|
||||||
@@ -69,6 +77,7 @@ public enum MemberRole {
|
|||||||
return switch (this) {
|
return switch (this) {
|
||||||
case ARCHITECT -> "architects";
|
case ARCHITECT -> "architects";
|
||||||
case DEV -> "developers";
|
case DEV -> "developers";
|
||||||
|
case HUNTER -> "hunters";
|
||||||
case REVIEWER -> "reviewers";
|
case REVIEWER -> "reviewers";
|
||||||
};
|
};
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -871,10 +871,23 @@ public final class SessionManager implements TurnListener {
|
|||||||
// digest lets a lead tell at a glance whether all members got the same charter; the source
|
// digest lets a lead tell at a glance whether all members got the same charter; the source
|
||||||
// records whether a role charter was configured ("fleet.charters.<role>") or only the reply
|
// records whether a role charter was configured ("fleet.charters.<role>") or only the reply
|
||||||
// charter was composed ("none").
|
// charter was composed ("none").
|
||||||
|
// #604: charterBytes rides along with charterSha256, not with charterSource — it is only
|
||||||
|
// meaningful as the digest's companion (a length turns "they differ" into "by how much").
|
||||||
|
// A member with no composed charter reports charterSource and nothing else, as before.
|
||||||
|
//
|
||||||
|
// Gating on the DIGEST rather than on the receipt is deliberate, and the two are not
|
||||||
|
// always null together. CharterReceipt.compose() derives the digest with digestOf(), which
|
||||||
|
// returns null for BLANK text, while the byte count is composed.getBytes().length, which
|
||||||
|
// does not. So a whitespace-only charter (a blank fleet.charters.<role> on a profile with
|
||||||
|
// no MCP, so no reply charter is appended) yields a null digest beside a non-zero size.
|
||||||
|
// Reporting a size with no digest would say "they differ by N bytes" about an artifact we
|
||||||
|
// cannot fingerprint, so this gate omits both. Never widen it to the receipt-level null
|
||||||
|
// check without deciding what that case should report.
|
||||||
if (session.charterReceipt() != null) {
|
if (session.charterReceipt() != null) {
|
||||||
m.put("charterSource", session.charterReceipt().charterSource());
|
m.put("charterSource", session.charterReceipt().charterSource());
|
||||||
if (session.charterReceipt().charterSha256() != null) {
|
if (session.charterReceipt().charterSha256() != null) {
|
||||||
m.put("charterSha256", session.charterReceipt().charterSha256());
|
m.put("charterSha256", session.charterReceipt().charterSha256());
|
||||||
|
m.put("charterBytes", session.charterReceipt().charterBytes());
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
m.put("liveStatus", live == null ? "unknown" : live.status().name().toLowerCase());
|
m.put("liveStatus", live == null ? "unknown" : live.status().name().toLowerCase());
|
||||||
|
|||||||
@@ -0,0 +1,183 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import ch.qos.logback.classic.Level;
|
||||||
|
import ch.qos.logback.classic.Logger;
|
||||||
|
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||||
|
import ch.qos.logback.core.read.ListAppender;
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.msg.InMemoryReplyInbox;
|
||||||
|
import dev.ltms.fleet.msg.LeadChannel;
|
||||||
|
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||||
|
import dev.ltms.fleet.msg.LeadMessage;
|
||||||
|
import dev.ltms.fleet.msg.ReplyInbox;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.concurrent.atomic.AtomicInteger;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertSame;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 step 4 ranks 3 and 8: the assembled daemon must use the AMQP openers from {@link
|
||||||
|
* ResourcePorts}, and its startup report must describe the object the runtime actually owns. These
|
||||||
|
* fakes never open a socket.
|
||||||
|
*/
|
||||||
|
class FleetdAssemblyAmqpOpenersTest {
|
||||||
|
|
||||||
|
private static final String COORD_ID = "assembly-test";
|
||||||
|
|
||||||
|
private static final class DurableReplyInbox implements ReplyInbox {
|
||||||
|
@Override public void own(String target) { }
|
||||||
|
@Override public void release(String target) { }
|
||||||
|
@Override public void publish(String target, String msgId, String content) { }
|
||||||
|
@Override public List<InboxMessage> peek(String target) { return List.of(); }
|
||||||
|
@Override public boolean ack(String target, String msgId) { return false; }
|
||||||
|
}
|
||||||
|
|
||||||
|
private static final class DurableLeadMailbox implements LeadChannelHandle {
|
||||||
|
@Override public void publish(String toCoordId, LeadMessage message) { }
|
||||||
|
@Override public List<LeadMessage> peek() { return List.of(); }
|
||||||
|
@Override public void ack(String msgId) { }
|
||||||
|
@Override public String selfCoordId() { return COORD_ID; }
|
||||||
|
@Override public boolean heldDurable() { return true; }
|
||||||
|
@Override public MailboxState inspect(String coordId) { return MailboxState.unknown(coordId); }
|
||||||
|
@Override public void close() { }
|
||||||
|
}
|
||||||
|
|
||||||
|
private static final class RecordingPorts implements ResourcePorts {
|
||||||
|
final FakeHerdr herdr = new FakeHerdr();
|
||||||
|
final DurableReplyInbox replyInbox = new DurableReplyInbox();
|
||||||
|
final DurableLeadMailbox leadMailbox = new DurableLeadMailbox();
|
||||||
|
final AtomicInteger replyOpenCalls = new AtomicInteger();
|
||||||
|
final AtomicInteger mailboxOpenCalls = new AtomicInteger();
|
||||||
|
final boolean openSucceeds;
|
||||||
|
|
||||||
|
RecordingPorts(boolean openSucceeds) {
|
||||||
|
this.openSucceeds = openSucceeds;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override public Map<String, String> environment() { return Map.of(); }
|
||||||
|
@Override public HerdrClient connectHerdr(Path socketPath) { return herdr; }
|
||||||
|
@Override public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
return (uri, prefetch) -> {
|
||||||
|
replyOpenCalls.incrementAndGet();
|
||||||
|
if (!openSucceeds) throw new IllegalStateException("fake reply broker is down");
|
||||||
|
return replyInbox;
|
||||||
|
};
|
||||||
|
}
|
||||||
|
@Override public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
return (uri, selfId, prefetch) -> {
|
||||||
|
mailboxOpenCalls.incrementAndGet();
|
||||||
|
if (!openSucceeds) throw new IllegalStateException("fake coordination broker is down");
|
||||||
|
return leadMailbox;
|
||||||
|
};
|
||||||
|
}
|
||||||
|
@Override public LongSupplier nanoClock() { return System::nanoTime; }
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||||
|
return () -> {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
@Override public LongSupplier wallClockNanos() { return System::nanoTime; }
|
||||||
|
@Override public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
return Executors.newSingleThreadScheduledExecutor();
|
||||||
|
}
|
||||||
|
@Override public void addShutdownHook(Runnable hook) { }
|
||||||
|
@Override public void startHttp(Javalin app, String host, int port) { }
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||||
|
Files.createDirectories(dir);
|
||||||
|
Path config = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(config, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
idleSleepGuard:
|
||||||
|
enabled: false
|
||||||
|
broker:
|
||||||
|
uri: "amqp://fake-reply-broker/vh"
|
||||||
|
coordinator:
|
||||||
|
uri: "amqp://fake-coordination-broker/vh"
|
||||||
|
selfId: "assembly-test"
|
||||||
|
""");
|
||||||
|
return FleetConfig.load(config);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetdRuntime assemble(Path dir, RecordingPorts ports) throws Exception {
|
||||||
|
FleetConfig cfg = writeConfig(dir);
|
||||||
|
return FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, new ConfigRef(dir.resolve("fleetd.yaml"), cfg),
|
||||||
|
new SubscriptionGuard(cfg.guard().hostSet())), ports);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static boolean reportContains(ListAppender<ILoggingEvent> appender, String text) {
|
||||||
|
return appender.list.stream().map(ILoggingEvent::getFormattedMessage).anyMatch(message -> message.contains(text));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void assembledAmqpOpenersAndTheirReportsAgreeOnDurableAndFallbackStates(@TempDir Path dir) throws Exception {
|
||||||
|
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||||
|
Level oldLevel = logger.getLevel();
|
||||||
|
ListAppender<ILoggingEvent> reports = new ListAppender<>();
|
||||||
|
reports.start();
|
||||||
|
logger.setLevel(Level.INFO);
|
||||||
|
logger.addAppender(reports);
|
||||||
|
try {
|
||||||
|
RecordingPorts durablePorts = new RecordingPorts(true);
|
||||||
|
FleetdRuntime durable = assemble(dir.resolve("durable"), durablePorts);
|
||||||
|
try {
|
||||||
|
// Control: this fails loudly if the assembly did not run or used an inert opener.
|
||||||
|
assertEquals(1, durablePorts.replyOpenCalls.get(), "assembly must call replyInboxOpener once");
|
||||||
|
assertEquals(1, durablePorts.mailboxOpenCalls.get(), "assembly must call leadMailboxOpener once");
|
||||||
|
assertSame(durablePorts.replyInbox, durable.replyInbox(),
|
||||||
|
"the durable reply report must describe the exact inbox the runtime owns");
|
||||||
|
assertSame(durablePorts.leadMailbox, durable.leadMailbox(),
|
||||||
|
"the coordination-on report must describe the exact mailbox the runtime owns");
|
||||||
|
assertNotNull(durable.leadCoordLoop(), "a durable mailbox must start lead coordination");
|
||||||
|
assertTrue(reportContains(reports, "reply inbox: AMQP broker (durable)"));
|
||||||
|
assertTrue(reportContains(reports, "lead coordination: ON as coord-id " + COORD_ID));
|
||||||
|
} finally {
|
||||||
|
durable.close();
|
||||||
|
}
|
||||||
|
|
||||||
|
reports.list.clear();
|
||||||
|
RecordingPorts fallbackPorts = new RecordingPorts(false);
|
||||||
|
FleetdRuntime fallback = assemble(dir.resolve("fallback"), fallbackPorts);
|
||||||
|
try {
|
||||||
|
assertEquals(1, fallbackPorts.replyOpenCalls.get(), "assembly must call the failing reply opener once");
|
||||||
|
assertEquals(1, fallbackPorts.mailboxOpenCalls.get(), "assembly must call the failing mailbox opener once");
|
||||||
|
assertTrue(fallback.replyInbox() instanceof InMemoryReplyInbox,
|
||||||
|
"a failed reply opener must make the runtime own the in-memory fallback");
|
||||||
|
assertNull(fallback.leadMailbox(), "a failed mailbox opener must leave coordination off");
|
||||||
|
assertNull(fallback.leadCoordLoop(), "coordination must not start without a mailbox");
|
||||||
|
assertTrue(reportContains(reports, "reply inbox: in-memory (soft-state)"));
|
||||||
|
assertTrue(reportContains(reports, "lead-to-lead messaging is OFF"));
|
||||||
|
} finally {
|
||||||
|
fallback.close();
|
||||||
|
}
|
||||||
|
} finally {
|
||||||
|
logger.detachAppender(reports);
|
||||||
|
logger.setLevel(oldLevel);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,230 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.herdr.PaneLocator;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.AfterEach;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.LinkedHashMap;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.CopyOnWriteArrayList;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 step 2, unit B2 (CB-185, identity half). Replaces the deleted
|
||||||
|
* {@code FleetdConnectionIdentityConstructionTest}, which pinned this claim by reading {@code
|
||||||
|
* Fleetd.java}'s source text for {@code "new PaneLocator(herdr, memberHerdr)"}. That claim moved
|
||||||
|
* to {@code FleetdAssembly.java} (fleetd #612 Unit A) and is pinned here instead, by driving the
|
||||||
|
* real {@link ConnectionIdentity} — via {@code runtime.mcp().identity()}, not a copy — that {@link
|
||||||
|
* FleetdAssembly#assembleAndStart} built.
|
||||||
|
*
|
||||||
|
* <p><strong>What this guards against</strong> (from the deleted test's own javadoc): pinning
|
||||||
|
* {@code PaneLocator} to {@code memberHerdr} alone leaves every LEAD's own MCP connection
|
||||||
|
* unresolvable ({@code callerTerminal == null}) the moment {@code memberHerdrSocket} names a
|
||||||
|
* second daemon, which breaks {@code fleet_reply}/{@code fleet_ask}/{@code fleet_whoami} for a
|
||||||
|
* lead. {@code PaneLocatorTest} already proves {@link PaneLocator} itself can search two clients
|
||||||
|
* given two — the gap this pins is that the assembly actually passes it two, and in the right
|
||||||
|
* order (lead first).
|
||||||
|
*
|
||||||
|
* <p><strong>Why this cannot be driven through a real MCP/HTTP round trip.</strong> The natural
|
||||||
|
* way to observe {@code ConnectionIdentity} would be a real {@code fleet_whoami} call over the
|
||||||
|
* built {@code FleetMcp}, the way {@code FleetMcpContextExtractorTest} drives its own
|
||||||
|
* hand-built one. That does not work for the REAL assembly, because {@code FleetdAssembly} wires
|
||||||
|
* {@code ConnectionIdentity} with a hardcoded {@code new LsofPeerPidLookup()} (see {@code
|
||||||
|
* FleetdAssembly.java:444}), and {@code LsofPeerPidLookup} explicitly excludes its own PID — see
|
||||||
|
* its javadoc: "we exclude our own PID and take the other end". In a JUnit test the HTTP client
|
||||||
|
* and the daemon under test run in the very same JVM, so the "client" and "server" ends of the
|
||||||
|
* loopback connection ARE the same PID, and {@code pidForLocalPort} always returns {@code -1}
|
||||||
|
* before {@link PaneLocator} is ever reached — proving nothing about which daemon(s) got searched.
|
||||||
|
* This test instead reaches the real {@link PaneLocator} the assembly built (through {@link
|
||||||
|
* ConnectionIdentity#panes()}, added for exactly this) and drives it with a chosen pid directly,
|
||||||
|
* bypassing the OS-dependent PID lookup entirely — a legitimate substitute, since the pid lookup
|
||||||
|
* is not what CB-185 is about.
|
||||||
|
*/
|
||||||
|
class FleetdAssemblyConnectionIdentityTest {
|
||||||
|
|
||||||
|
/** Same shape as {@code FleetdAssemblyLifecycleTest}'s fake, but keys {@code connectHerdr} by
|
||||||
|
* socket path so the lead and member daemons can be two DIFFERENT {@link FakeHerdr}s. */
|
||||||
|
private static final class TwoHerdrResourcePorts implements ResourcePorts {
|
||||||
|
|
||||||
|
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
|
||||||
|
final CopyOnWriteArrayList<ScheduledExecutorService> schedulers = new CopyOnWriteArrayList<>();
|
||||||
|
Runnable shutdownHook;
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
HerdrClient client = herdrsBySocket.get(socketPath);
|
||||||
|
if (client == null) {
|
||||||
|
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
|
||||||
|
}
|
||||||
|
return client;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
return (uri, prefetch) -> new dev.ltms.fleet.msg.InMemoryReplyInbox();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
return (uri, selfCoordId, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException(
|
||||||
|
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||||
|
schedulers.add(scheduler);
|
||||||
|
return scheduler;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
this.shutdownHook = hook;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
// Deliberately never bind — this test never issues a real HTTP request.
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||||
|
return () -> {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private FleetdRuntime runtime;
|
||||||
|
private TwoHerdrResourcePorts ports;
|
||||||
|
|
||||||
|
@AfterEach
|
||||||
|
void tearDown() {
|
||||||
|
if (ports != null && ports.shutdownHook != null) {
|
||||||
|
ports.shutdownHook.run();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
|
||||||
|
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
herdrSocket: "%s"
|
||||||
|
memberHerdrSocket: "%s"
|
||||||
|
lifecycle:
|
||||||
|
idleTtlSeconds: 600
|
||||||
|
health:
|
||||||
|
enabled: false
|
||||||
|
broker:
|
||||||
|
uri: "amqp://fake-test-broker/vh"
|
||||||
|
""".formatted(LEAD_SOCKET, MEMBER_SOCKET));
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
private FleetdRuntime assemble(Path dir, FakeHerdr lead, FakeHerdr member) throws Exception {
|
||||||
|
FleetConfig cfg = writeConfig(dir);
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
ports = new TwoHerdrResourcePorts();
|
||||||
|
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
|
||||||
|
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
|
||||||
|
runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
return runtime;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The pin. {@code lead} carries the one pane {@link FakeHerdr}'s canned {@code
|
||||||
|
* pane.process_info} ties to {@link FakeHerdr#WORKER_PID} (pane {@code w2:p7}); {@code member}
|
||||||
|
* reports NO panes at all ({@link FakeHerdr#withNoPanes()}) — modelling a second daemon that
|
||||||
|
* simply does not host the caller's pane, exactly the CB-185 javadoc's scenario for a lead's
|
||||||
|
* own connection. If {@code PaneLocator} only ever searches the member daemon (the bug), this
|
||||||
|
* pid resolves to nothing, because the pane that owns it lives on the LEAD daemon the bug
|
||||||
|
* skips.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void connectionIdentitySearchesTheLeadDaemonNotJustTheMemberOne(@TempDir Path dir) throws Exception {
|
||||||
|
FakeHerdr lead = new FakeHerdr();
|
||||||
|
FakeHerdr member = new FakeHerdr().withNoPanes();
|
||||||
|
|
||||||
|
assemble(dir, lead, member);
|
||||||
|
|
||||||
|
PaneLocator panes = runtime.mcp().identity().panes();
|
||||||
|
PaneLocator.Lookup lookup = panes.terminalForPid(FakeHerdr.WORKER_PID);
|
||||||
|
|
||||||
|
assertEquals("term_a", lookup.terminal(),
|
||||||
|
"the pane owning WORKER_PID lives on the LEAD daemon only (the member fake reports "
|
||||||
|
+ "no panes) — PaneLocator must still find it, which is only possible if it "
|
||||||
|
+ "searches the lead client and not just the member one");
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The mirror control: when the pane instead lives ONLY on the member daemon (the lead reports
|
||||||
|
* no panes), the lookup must still find it — proving the member client is genuinely searched
|
||||||
|
* too, not merely tolerated as a second, always-losing argument.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void connectionIdentityAlsoSearchesTheMemberDaemon(@TempDir Path dir) throws Exception {
|
||||||
|
FakeHerdr lead = new FakeHerdr().withNoPanes();
|
||||||
|
FakeHerdr member = new FakeHerdr();
|
||||||
|
|
||||||
|
assemble(dir, lead, member);
|
||||||
|
|
||||||
|
PaneLocator panes = runtime.mcp().identity().panes();
|
||||||
|
PaneLocator.Lookup lookup = panes.terminalForPid(FakeHerdr.WORKER_PID);
|
||||||
|
|
||||||
|
assertEquals("term_a", lookup.terminal(),
|
||||||
|
"the pane owning WORKER_PID lives on the MEMBER daemon only — PaneLocator must "
|
||||||
|
+ "find it there too");
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Sanity control: a pid nobody owns resolves to nothing on either daemon. */
|
||||||
|
@Test
|
||||||
|
void aPidNoPaneOwnsResolvesToNoTerminalOnEitherDaemon(@TempDir Path dir) throws Exception {
|
||||||
|
FakeHerdr lead = new FakeHerdr();
|
||||||
|
FakeHerdr member = new FakeHerdr();
|
||||||
|
|
||||||
|
assemble(dir, lead, member);
|
||||||
|
|
||||||
|
PaneLocator panes = runtime.mcp().identity().panes();
|
||||||
|
PaneLocator.Lookup lookup = panes.terminalForPid(999_999L);
|
||||||
|
|
||||||
|
assertNull(lookup.terminal());
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,219 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.msg.LeadChannel;
|
||||||
|
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||||
|
import dev.ltms.fleet.msg.LeadMessage;
|
||||||
|
import dev.ltms.fleet.msg.ReplyInbox;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertSame;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 A-gaps (gap 1): {@code FleetdAssemblyLifecycleTest}'s own class javadoc says plainly
|
||||||
|
* that it leaves {@code coordinator:} unset, so {@code leadMailbox} and {@code leadCoordLoop} stay
|
||||||
|
* {@code null} throughout — the configured-coordinator path is never exercised by Unit A's own
|
||||||
|
* test. This class drives that path instead: a real {@code coordinator:} block, a fake {@link
|
||||||
|
* Fleetd.LeadMailboxOpener} returning a fake closeable channel (never a real broker connection),
|
||||||
|
* and proof that {@link FleetdAssembly#assembleAndStart} both builds it and, on shutdown, closes it.
|
||||||
|
*
|
||||||
|
* <p>Made possible by generalising {@code Fleetd.LeadMailboxOpener}'s return type (and {@code
|
||||||
|
* FleetdRuntime}'s field) from the concrete {@code LeadMailbox} to {@link LeadChannelHandle} — a
|
||||||
|
* {@link LeadChannel} its owner can also close. {@code FleetMcp} and {@code LeadCoordLoop} already
|
||||||
|
* consumed the narrower {@link LeadChannel}; this only widens the one seam that owns and closes it.
|
||||||
|
*/
|
||||||
|
class FleetdAssemblyCoordinatorLifecycleTest {
|
||||||
|
|
||||||
|
private static final String SELF_COORD_ID = "test-lead";
|
||||||
|
|
||||||
|
/** A fake {@link LeadChannelHandle}: never touches a broker, and records whether it was closed. */
|
||||||
|
private static final class FakeLeadChannel implements LeadChannelHandle {
|
||||||
|
volatile boolean closed = false;
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void publish(String toCoordId, LeadMessage m) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public List<LeadMessage> peek() {
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void ack(String msgId) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public String selfCoordId() {
|
||||||
|
return SELF_COORD_ID;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public boolean heldDurable() {
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public MailboxState inspect(String coordId) {
|
||||||
|
return MailboxState.unknown(coordId);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
closed = true;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Minimal fake {@link ResourcePorts}: a real herdr fake, a fake reply inbox, and a real, offered fake lead channel. */
|
||||||
|
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||||
|
|
||||||
|
final FakeHerdr herdr = new FakeHerdr();
|
||||||
|
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||||
|
final FakeLeadChannel leadChannel = new FakeLeadChannel();
|
||||||
|
String offeredUri;
|
||||||
|
String offeredSelfId;
|
||||||
|
Runnable shutdownHook;
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
return herdr;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
return (uri, prefetch) -> replyInbox;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
return (uri, selfCoordId, prefetch) -> {
|
||||||
|
offeredUri = uri;
|
||||||
|
offeredSelfId = selfCoordId;
|
||||||
|
return leadChannel;
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
return Executors.newSingleThreadScheduledExecutor();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
this.shutdownHook = hook;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
// No real HTTP bind in a unit test.
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||||
|
return () -> {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static final class SentinelReplyInbox implements ReplyInbox {
|
||||||
|
@Override
|
||||||
|
public void own(String target) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void release(String target) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void publish(String target, String msgId, String content) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public List<InboxMessage> peek(String target) {
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public boolean ack(String target, String msgId) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
idleSleepGuard:
|
||||||
|
enabled: false
|
||||||
|
coordinator:
|
||||||
|
uri: "amqp://fake-lead-broker/vh"
|
||||||
|
selfId: "test-lead"
|
||||||
|
""");
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void configuredCoordinatorIsBuiltByTheAssemblyAndClosedOnShutdown(@TempDir Path dir) throws Exception {
|
||||||
|
FleetConfig cfg = writeConfig(dir);
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||||
|
|
||||||
|
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
|
||||||
|
// --- the assembly actually calls the configured opener and builds the coordinator path ---
|
||||||
|
assertEquals("amqp://fake-lead-broker/vh", ports.offeredUri,
|
||||||
|
"the assembly must open the mailbox at the configured broker uri");
|
||||||
|
assertEquals(SELF_COORD_ID, ports.offeredSelfId,
|
||||||
|
"the assembly must open the mailbox under the configured selfId");
|
||||||
|
assertSame(ports.leadChannel, runtime.leadMailbox(),
|
||||||
|
"FleetdRuntime must own the exact LeadChannelHandle the opener returned, not a copy");
|
||||||
|
assertNotNull(runtime.leadCoordLoop(),
|
||||||
|
"a configured coordinator: block must build the receiving LeadCoordLoop too");
|
||||||
|
assertFalse(ports.leadChannel.closed, "the channel must still be open while the daemon is running");
|
||||||
|
|
||||||
|
// --- shutting the assembly down closes it -------------------------------------------------
|
||||||
|
assertNotNull(ports.shutdownHook, "FleetdAssembly must have registered a shutdown hook");
|
||||||
|
ports.shutdownHook.run();
|
||||||
|
|
||||||
|
assertTrue(ports.leadChannel.closed,
|
||||||
|
"FleetdRuntime.close() must close the configured LeadChannelHandle");
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,267 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.AfterEach;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.Timeout;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.net.URI;
|
||||||
|
import java.net.http.HttpClient;
|
||||||
|
import java.net.http.HttpRequest;
|
||||||
|
import java.net.http.HttpResponse;
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.LinkedHashMap;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.CopyOnWriteArrayList;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 step 2, unit B2 (CB-185, {@code FleetApp} half). Replaces the deleted {@code
|
||||||
|
* FleetdFleetAppConstructionTest}, which pinned this claim by reading {@code Fleetd.java}'s
|
||||||
|
* source text for {@code "new FleetApp(herdr, memberHerdr, workers,"}. That claim moved to {@code
|
||||||
|
* FleetdAssembly.java} (fleetd #612 Unit A) and is pinned here instead, by driving the real {@code
|
||||||
|
* Javalin} app — via {@code runtime.app()}, not a copy — that {@link
|
||||||
|
* FleetdAssembly#assembleAndStart} built and handed to {@link FleetdRuntime}.
|
||||||
|
*
|
||||||
|
* <p><strong>What this guards against</strong> (from the deleted test's own javadoc): constructing
|
||||||
|
* {@code FleetApp} with the lead-only {@code herdr} client (dropping {@code memberHerdr}) makes
|
||||||
|
* {@code GET /healthz} report green while the MEMBER daemon is down — so every spawn fails
|
||||||
|
* invisibly — and silently drops every member workspace from {@code GET /sessions}. {@code
|
||||||
|
* FleetAppTwoDaemonTest} already proves {@code FleetApp} itself merges/gates correctly given two
|
||||||
|
* clients; the gap this pins is that the assembly actually passes it two.
|
||||||
|
*
|
||||||
|
* <p><strong>Both directions, not just one</strong> (fleetd #612 issue comment 17525): the deleted
|
||||||
|
* guard's positive assertion required the exact pair {@code "new FleetApp(herdr, memberHerdr,
|
||||||
|
* workers,"}, which does not survive EITHER daemon being dropped. An earlier version of this class
|
||||||
|
* only proved the member-dropped direction, which left {@code new FleetApp(memberHerdr,
|
||||||
|
* memberHerdr, ...)} — the symmetric bug, {@code /healthz} green while the LEAD daemon is down —
|
||||||
|
* an undetected regression. {@link #healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp}
|
||||||
|
* closes that.
|
||||||
|
*
|
||||||
|
* <p>Unlike the {@code ConnectionIdentity} half of CB-185 ({@code
|
||||||
|
* FleetdAssemblyConnectionIdentityTest}), {@code /healthz} needs no caller identity at all, so
|
||||||
|
* this test can bind {@link FleetdRuntime#app()} to a REAL ephemeral port (exactly {@code
|
||||||
|
* FleetAppTwoDaemonTest} does for its own hand-built {@code FleetApp}) and drive it with a real
|
||||||
|
* {@code HttpClient} — no accessor needed for this half.
|
||||||
|
*
|
||||||
|
* <p><strong>{@code GET /sessions} could not be driven the same way</strong>, so this class does
|
||||||
|
* not pin the merge half of the deleted test's javadoc. {@code /sessions} requires
|
||||||
|
* {@code Authz.Action.READ}, which — through the REAL assembly's real {@code
|
||||||
|
* CallerResolver}/{@code ConnectionIdentity} (built with a hardcoded {@code
|
||||||
|
* new LsofPeerPidLookup()}) — needs {@code Caller.resolved()}, i.e. a real positive pid from
|
||||||
|
* {@code lsof}. {@code LsofPeerPidLookup} excludes its own pid (see its javadoc), and a JUnit
|
||||||
|
* test's HTTP client and the daemon under test share one JVM pid, so the resolved pid is always
|
||||||
|
* {@code -1} and every such request is refused as {@code ANONYMOUS} (fleetd #317's fail-closed
|
||||||
|
* rule) before the route handler — and its {@code memberHerdr} merge — is ever reached. Verified
|
||||||
|
* directly: driving {@code GET /sessions} here returns {@code 401 unauthenticated}, not the
|
||||||
|
* merged body. {@code FleetAppTwoDaemonTest} avoids this because it builds {@code FleetApp} with
|
||||||
|
* {@code callers: null}, which is not what the real assembly passes. The {@code /healthz} pin
|
||||||
|
* below is what this class relies on for CB-185's {@code FleetApp} half; {@code
|
||||||
|
* FleetAppTwoDaemonTest} remains the full behavioural proof that {@code FleetApp} itself merges
|
||||||
|
* {@code /sessions} correctly once handed two clients.
|
||||||
|
*
|
||||||
|
* <p><strong>fleetd #629 follow-up.</strong> The fix below (see {@link TwoHerdrResourcePorts})
|
||||||
|
* makes {@link #healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp}'s fake {@code
|
||||||
|
* nanoClock()} frozen unless {@code herdrPollWait()} itself advances it. That is a sharper pin
|
||||||
|
* than an assertion — if a future edit to {@code FleetdAssembly} ever bypasses {@code
|
||||||
|
* ports.herdrPollWait()} again (e.g. reverting to a hardcoded {@code Thread.sleep}), the clock
|
||||||
|
* never advances, {@code Fleetd#awaitHerdr}'s deadline is never reached, and this test hangs
|
||||||
|
* forever instead of failing — proven by deliberately reintroducing that exact regression while
|
||||||
|
* fixing this ticket. {@code @Timeout} turns that silent hang into a bounded, named test failure:
|
||||||
|
* {@code SEPARATE_THREAD} so JUnit's timeout governor can actually interrupt a thread stuck in a
|
||||||
|
* real {@code Thread.sleep} loop (the default {@code SAME_THREAD} mode cannot — it only measures
|
||||||
|
* elapsed time after the test method returns on its own, which never happens here). 10 seconds is
|
||||||
|
* roughly 150x the real passing times measured here (~0.06s), so a slow CI machine has no reason
|
||||||
|
* to flake, and it is still 3x faster than discovering the regression by burning a CI job's whole
|
||||||
|
* wall-clock budget.
|
||||||
|
*/
|
||||||
|
@Timeout(value = 10, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
|
||||||
|
class FleetdAssemblyFleetAppTest {
|
||||||
|
|
||||||
|
private static final class TwoHerdrResourcePorts implements ResourcePorts {
|
||||||
|
|
||||||
|
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
|
||||||
|
final CopyOnWriteArrayList<ScheduledExecutorService> schedulers = new CopyOnWriteArrayList<>();
|
||||||
|
// fleetd #629: a fake, advanceable clock — NOT System::nanoTime. awaitHerdr's poll wait
|
||||||
|
// (herdrPollWait() below) advances this on every poll instead of sleeping for real, so the
|
||||||
|
// down-lead test below reaches awaitHerdr's deadline without burning real wall-clock time.
|
||||||
|
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L); // arbitrary non-zero start
|
||||||
|
Runnable shutdownHook;
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
HerdrClient client = herdrsBySocket.get(socketPath);
|
||||||
|
if (client == null) {
|
||||||
|
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
|
||||||
|
}
|
||||||
|
return client;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
return (uri, prefetch) -> new dev.ltms.fleet.msg.InMemoryReplyInbox();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
return (uri, selfCoordId, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException(
|
||||||
|
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
return nowNanos::get;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// fleetd #629: advance the fake clock instead of a real Thread.sleep, so awaitHerdr's
|
||||||
|
// deadline is reached in real time regardless of the configured poll interval.
|
||||||
|
return () -> nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(1));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||||
|
schedulers.add(scheduler);
|
||||||
|
return scheduler;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
this.shutdownHook = hook;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
// Deliberately never bind here — this test binds runtime.app() itself, for real, below.
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
|
||||||
|
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
|
||||||
|
|
||||||
|
private final HttpClient http = HttpClient.newHttpClient();
|
||||||
|
private FleetdRuntime runtime;
|
||||||
|
private TwoHerdrResourcePorts ports;
|
||||||
|
private Javalin boundApp;
|
||||||
|
|
||||||
|
@AfterEach
|
||||||
|
void tearDown() {
|
||||||
|
if (boundApp != null) {
|
||||||
|
boundApp.stop();
|
||||||
|
}
|
||||||
|
if (ports != null && ports.shutdownHook != null) {
|
||||||
|
ports.shutdownHook.run();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
herdrSocket: "%s"
|
||||||
|
memberHerdrSocket: "%s"
|
||||||
|
lifecycle:
|
||||||
|
idleTtlSeconds: 600
|
||||||
|
health:
|
||||||
|
enabled: false
|
||||||
|
broker:
|
||||||
|
uri: "amqp://fake-test-broker/vh"
|
||||||
|
""".formatted(LEAD_SOCKET, MEMBER_SOCKET));
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Assembles the real graph, then binds the real {@code Javalin app} to an ephemeral port. */
|
||||||
|
private int assembleAndBind(Path dir, FakeHerdr lead, FakeHerdr member) throws Exception {
|
||||||
|
FleetConfig cfg = writeConfig(dir);
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
ports = new TwoHerdrResourcePorts();
|
||||||
|
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
|
||||||
|
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
|
||||||
|
runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
boundApp = runtime.app().start("127.0.0.1", 0);
|
||||||
|
return boundApp.port();
|
||||||
|
}
|
||||||
|
|
||||||
|
private HttpResponse<String> get(int port, String path) throws Exception {
|
||||||
|
HttpRequest req = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path)).GET().build();
|
||||||
|
return http.send(req, HttpResponse.BodyHandlers.ofString());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The pin. The MEMBER daemon is down; the LEAD daemon is healthy. If the assembly built
|
||||||
|
* {@code FleetApp} with only the lead client (the bug: passing {@code herdr} where {@code
|
||||||
|
* memberHerdr} is expected), the down member is invisible and {@code /healthz} stays 200.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void healthzGoesRedWhenTheMemberDaemonIsDownEvenThoughTheLeadIsUp(@TempDir Path dir) throws Exception {
|
||||||
|
FakeHerdr lead = new FakeHerdr();
|
||||||
|
FakeHerdr member = new FakeHerdr().healthy(false);
|
||||||
|
|
||||||
|
int port = assembleAndBind(dir, lead, member);
|
||||||
|
|
||||||
|
HttpResponse<String> res = get(port, "/healthz");
|
||||||
|
assertEquals(503, res.statusCode(),
|
||||||
|
"a down MEMBER daemon must not be masked by a healthy lead: " + res.body());
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Sanity control: both daemons healthy must still be green through the real assembly. */
|
||||||
|
@Test
|
||||||
|
void healthzIsGreenWhenBothDaemonsAreUp(@TempDir Path dir) throws Exception {
|
||||||
|
FakeHerdr lead = new FakeHerdr();
|
||||||
|
FakeHerdr member = new FakeHerdr();
|
||||||
|
|
||||||
|
int port = assembleAndBind(dir, lead, member);
|
||||||
|
|
||||||
|
assertEquals(200, get(port, "/healthz").statusCode());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The symmetric pin (fleetd #612 issue comment 17525): the LEAD daemon is down; the MEMBER
|
||||||
|
* daemon is healthy. If the assembly built {@code FleetApp} with only the member client
|
||||||
|
* (dropping {@code herdr} — the mirror of the bug above, {@code new FleetApp(memberHerdr,
|
||||||
|
* memberHerdr, ...)}), the down LEAD is invisible and {@code /healthz} stays 200. Without this
|
||||||
|
* case the pair above is one-directional and does not cover the deleted guard's positive
|
||||||
|
* assertion (it required BOTH {@code herdr,} and {@code memberHerdr,} in that order).
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp(@TempDir Path dir) throws Exception {
|
||||||
|
FakeHerdr lead = new FakeHerdr().healthy(false);
|
||||||
|
FakeHerdr member = new FakeHerdr();
|
||||||
|
|
||||||
|
int port = assembleAndBind(dir, lead, member);
|
||||||
|
|
||||||
|
HttpResponse<String> res = get(port, "/healthz");
|
||||||
|
assertEquals(503, res.statusCode(),
|
||||||
|
"a down LEAD daemon must not be masked by a healthy member: " + res.body());
|
||||||
|
}
|
||||||
|
}
|
||||||
+198
@@ -0,0 +1,198 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.health.FleetHealthMonitor;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.msg.MessageService;
|
||||||
|
import dev.ltms.fleet.msg.ReplyInbox;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.lang.reflect.Field;
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.function.BiConsumer;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 rank 7 — {@code FleetdAssembly.java:429} wires {@link FleetHealthMonitor}'s {@code
|
||||||
|
* failTarget} callback with {@code Fleetd.healthFailTarget(messages)}. {@link
|
||||||
|
* FleetdHealthFailTargetWiringTest} already pins that the FACTORY itself delegates to {@code
|
||||||
|
* messages::abandon}, but it calls {@code Fleetd.healthFailTarget} directly — it never drives {@code
|
||||||
|
* FleetdAssembly.assembleAndStart} and so cannot see whether the real call site at {@code :429}
|
||||||
|
* still passes it the real, assembled {@link MessageService}. Swapping that argument for a no-op
|
||||||
|
* {@code (a, b) -> {}} compiles clean and leaves the whole suite — including the factory-level test
|
||||||
|
* — green: a dead member's waiting ticket then sits {@code PENDING} for the full 30-minute async
|
||||||
|
* timeout instead of failing immediately.
|
||||||
|
*
|
||||||
|
* <p>This test assembles the real daemon with {@code health.enabled: true}, pulls the REAL {@code
|
||||||
|
* failTarget} {@link BiConsumer} out of the REAL, assembled {@link FleetHealthMonitor} (via
|
||||||
|
* reflection — the field is package-private to {@code dev.ltms.fleet.health}, and nothing public
|
||||||
|
* exposes it; {@code StatusPollerResilienceTest} already uses the same technique in this suite), and
|
||||||
|
* invokes it directly against the REAL {@link MessageService} {@link FleetdRuntime#messages()}
|
||||||
|
* returns. A no-op lambda swapped in at the call site leaves the ticket {@code PENDING} forever,
|
||||||
|
* which this test catches; the real one fails it.
|
||||||
|
*/
|
||||||
|
class FleetdAssemblyHealthFailTargetBehaviouralTest {
|
||||||
|
|
||||||
|
private static final String TARGET = "term_a";
|
||||||
|
|
||||||
|
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||||
|
final FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Runnable shutdownHook;
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
return herdr;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
return (uri, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException("no broker: block is configured");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
return (uri, selfCoordId, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException("no coordinator: block is configured");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||||
|
return () -> {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
return Executors.newSingleThreadScheduledExecutor();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
this.shutdownHook = hook;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
idleSleepGuard:
|
||||||
|
enabled: false
|
||||||
|
health:
|
||||||
|
enabled: true
|
||||||
|
intervalSeconds: 30
|
||||||
|
""");
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
@Test
|
||||||
|
@DisplayName("[BEHAVIOURAL] the real assembled FleetHealthMonitor's failTarget reaches the real "
|
||||||
|
+ "MessageService.abandon, not a no-op")
|
||||||
|
void assembledHealthFailTargetReachesRealMessages(@TempDir Path dir) throws Exception {
|
||||||
|
FleetConfig cfg = writeConfig(dir);
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||||
|
|
||||||
|
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
assertNotNull(ports.shutdownHook, "FleetdAssembly must have registered a shutdown hook");
|
||||||
|
|
||||||
|
// Surefire runs the whole suite in one JVM fork, so the scheduler/loops this assembly starts
|
||||||
|
// (SessionReaper, StatusPoller, the health monitor) must be torn down here, on the failure
|
||||||
|
// path too — hence the try/finally, not just a statement at the end of the happy path.
|
||||||
|
try {
|
||||||
|
FleetHealthMonitor healthMonitor = runtime.healthMonitor();
|
||||||
|
assertNotNull(healthMonitor, "health.enabled: true in this test's config, so "
|
||||||
|
+ "FleetdAssembly.assembleAndStart must have built a real FleetHealthMonitor");
|
||||||
|
|
||||||
|
Field field = FleetHealthMonitor.class.getDeclaredField("failTarget");
|
||||||
|
field.setAccessible(true);
|
||||||
|
BiConsumer<String, String> failTarget = (BiConsumer<String, String>) field.get(healthMonitor);
|
||||||
|
assertNotNull(failTarget, "FleetHealthMonitor's failTarget must never be null — the "
|
||||||
|
+ "constructor itself requires it");
|
||||||
|
|
||||||
|
MessageService messages = runtime.messages();
|
||||||
|
|
||||||
|
// --- loud control: prove the assembled MessageService is actually wired up and a ticket is
|
||||||
|
// genuinely PENDING before failTarget ever runs. If this fails, the test below would pass
|
||||||
|
// vacuously on a MessageService that never got a ticket in the first place. TARGET has no
|
||||||
|
// live agent behind it (no session was ever acquired), so nothing resolves this ticket on
|
||||||
|
// its own — it stays PENDING until failTarget (or a timeout) ends it.
|
||||||
|
String ticket = messages.sendAsync(TARGET, "long task");
|
||||||
|
MessageService.TaskView before = messages.poll(ticket);
|
||||||
|
assertEquals(MessageService.Phase.PENDING, before.phase(),
|
||||||
|
"control: the async ticket must be PENDING before failTarget runs");
|
||||||
|
|
||||||
|
failTarget.accept(TARGET, "member unreachable (health monitor)");
|
||||||
|
|
||||||
|
MessageService.TaskView after = awaitTerminal(messages, ticket);
|
||||||
|
assertEquals(MessageService.Phase.FAILED, after.phase(),
|
||||||
|
"FleetdAssembly.java:429 must pass Fleetd.healthFailTarget(messages) built from the "
|
||||||
|
+ "SAME assembled MessageService — a no-op BiConsumer at that call site leaves "
|
||||||
|
+ "this ticket PENDING for the full 30-minute async timeout instead of failing it");
|
||||||
|
assertTrue(after.detail() != null && after.detail().contains("member unreachable"),
|
||||||
|
"the failure reason passed to failTarget.accept must reach MessageService.abandon and "
|
||||||
|
+ "end up in the ticket's detail");
|
||||||
|
} finally {
|
||||||
|
// Proof the teardown actually ran, not just an assurance that a finally was added: the
|
||||||
|
// captured shutdown hook's close order (FleetdAssemblyLifecycleTest) closes the herdr
|
||||||
|
// client last, so ports.herdr.closed flips to true only if this hook really executed.
|
||||||
|
ports.shutdownHook.run();
|
||||||
|
assertTrue(ports.herdr.closed, "the captured shutdown hook must have run and closed herdr — "
|
||||||
|
+ "proof this test's assembled background loops/scheduler were torn down");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static MessageService.TaskView awaitTerminal(MessageService messages, String ticket)
|
||||||
|
throws InterruptedException {
|
||||||
|
long deadline = System.currentTimeMillis() + 5000;
|
||||||
|
MessageService.TaskView view = messages.poll(ticket);
|
||||||
|
while (view.phase() == MessageService.Phase.PENDING && System.currentTimeMillis() < deadline) {
|
||||||
|
//noinspection BusyWait
|
||||||
|
Thread.sleep(10);
|
||||||
|
view = messages.poll(ticket);
|
||||||
|
}
|
||||||
|
return view;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,266 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.inject.LoopWatchdog;
|
||||||
|
import dev.ltms.fleet.msg.ReplyInbox;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.ArrayList;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.CopyOnWriteArrayList;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 Unit A: proves {@link FleetdAssembly#assembleAndStart} — not a copy of its logic —
|
||||||
|
* against a real {@link FleetConfig}, a {@link FakeHerdr} and a fake {@link ResourcePorts}, with no
|
||||||
|
* real herdr socket, no real broker, and no real HTTP bind.
|
||||||
|
*
|
||||||
|
* <p><strong>The canonical order this test asserts against was recorded from {@code Fleetd.java}
|
||||||
|
* BEFORE any code moved</strong> (fleetd #612 Unit A's mandated order of work), by reading the
|
||||||
|
* original {@code main}'s body and its shutdown-hook {@code Thread}:
|
||||||
|
*
|
||||||
|
* <p>Start order: {@code SessionReaper.start()} → {@code StatusPoller.start()} →
|
||||||
|
* {@code LeadHeartbeatLoop.start()} (opt-in) → {@code FleetHealthMonitor.start()} (opt-in) →
|
||||||
|
* {@code LeadCoordLoop.start()} (opt-in) → {@code ConfigWatcher.start()} (opt-in) →
|
||||||
|
* {@code app.start()} (HTTP), always last.
|
||||||
|
*
|
||||||
|
* <p>Close order (from the original shutdown hook body): {@code sessions.close(drainTimeoutSeconds)}
|
||||||
|
* → {@code poller.stop()} → {@code messages.close()} → {@code pushLoop.close()} →
|
||||||
|
* {@code heartbeat.close()} (if present) → {@code leadCoordLoop.close()} (if present) →
|
||||||
|
* {@code leadCoordScheduler.shutdownNow()} (if present) → {@code healthMonitor.stop()} (if present)
|
||||||
|
* → {@code configWatcher.stop()} (if present) → {@code mcp.close()} → {@code reaper.stop()} (if
|
||||||
|
* present) → {@code idleSleepGuard.close()} (if present) → {@code replyInbox.close()} (if
|
||||||
|
* {@code AutoCloseable}) → {@code leadMailbox.close()} (if present) → {@code router.close()}.
|
||||||
|
*
|
||||||
|
* <p>This test's config deliberately leaves {@code coordinator:} unset, so {@code leadMailbox} and
|
||||||
|
* {@code leadCoordLoop} stay {@code null} throughout — the lead-mailbox resource-ledger criterion is
|
||||||
|
* NOT exercised here; see the class-level caveat in the implementer's hand-off. {@code
|
||||||
|
* idleSleepGuard.enabled: false} is set for the same kind of reason: it would otherwise try to spawn
|
||||||
|
* a real {@code caffeinate} subprocess, which is not one of the resources the ticket's acceptance
|
||||||
|
* criteria names (scheduler/inbox/mailbox/loop/MCP server/herdr router).
|
||||||
|
*/
|
||||||
|
class FleetdAssemblyLifecycleTest {
|
||||||
|
|
||||||
|
/** The one {@link HerdrClient} both {@code herdrSocket} and {@code memberHerdrSocket} resolve to. */
|
||||||
|
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||||
|
|
||||||
|
final FakeHerdr herdr = new FakeHerdr();
|
||||||
|
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||||
|
final List<String> ledger = new CopyOnWriteArrayList<>();
|
||||||
|
final List<ScheduledExecutorService> schedulers = new CopyOnWriteArrayList<>();
|
||||||
|
final List<String> schedulerPurposes = new ArrayList<>();
|
||||||
|
Runnable shutdownHook;
|
||||||
|
Javalin startedApp;
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
ledger.add("connectHerdr");
|
||||||
|
return herdr;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
ledger.add("replyInboxOpener");
|
||||||
|
return (uri, prefetch) -> replyInbox;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
// Never invoked: this test's config has no `coordinator:` block, so
|
||||||
|
// Fleetd.openLeadMailbox returns null before calling the opener at all.
|
||||||
|
return (uri, selfCoordId, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException(
|
||||||
|
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
ledger.add("newScheduler:" + purpose);
|
||||||
|
schedulerPurposes.add(purpose);
|
||||||
|
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||||
|
schedulers.add(scheduler);
|
||||||
|
return scheduler;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
ledger.add("addShutdownHook");
|
||||||
|
this.shutdownHook = hook;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
// Deliberately never call app.start(host, port): no real HTTP bind in a unit test.
|
||||||
|
ledger.add("startHttp");
|
||||||
|
this.startedApp = app;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||||
|
return () -> {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** A fake {@link ReplyInbox} that is also {@link AutoCloseable}, so the ledger can prove it closes. */
|
||||||
|
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
|
||||||
|
volatile boolean closed = false;
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void own(String target) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void release(String target) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void publish(String target, String msgId, String content) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public List<InboxMessage> peek(String target) {
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public boolean ack(String target, String msgId) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
closed = true;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
lifecycle:
|
||||||
|
idleTtlSeconds: 600
|
||||||
|
health:
|
||||||
|
enabled: true
|
||||||
|
intervalSeconds: 30
|
||||||
|
leadHeartbeat:
|
||||||
|
idleAfterSeconds: 600
|
||||||
|
backoffMs: 15000
|
||||||
|
quietNudgeCap: 5
|
||||||
|
configReload:
|
||||||
|
enabled: true
|
||||||
|
intervalSeconds: 30
|
||||||
|
idleSleepGuard:
|
||||||
|
enabled: false
|
||||||
|
broker:
|
||||||
|
uri: "amqp://fake-test-broker/vh"
|
||||||
|
""");
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void assemblesTheRealBootGraphWithoutTouchingAnyRealSocketBrokerOrPort(@TempDir Path dir)
|
||||||
|
throws Exception {
|
||||||
|
FleetConfig cfg = writeConfig(dir);
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||||
|
|
||||||
|
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
|
||||||
|
// --- start order: the herdr connect, then the recurring background loops, in the recorded
|
||||||
|
// order, then the shutdown hook is registered, then (finally) HTTP "starts" -------------
|
||||||
|
assertTrue(ports.ledger.indexOf("connectHerdr") < ports.ledger.indexOf("newScheduler:bridge-push-"),
|
||||||
|
"herdr must connect before the push scheduler is created: " + ports.ledger);
|
||||||
|
assertEquals(List.of("bridge-push-", "bridge-heartbeat-", "bridge-health-"), ports.schedulerPurposes,
|
||||||
|
"the three always-created schedulers must be requested in exactly this order: "
|
||||||
|
+ ports.schedulerPurposes);
|
||||||
|
assertTrue(ports.ledger.indexOf("addShutdownHook") < ports.ledger.indexOf("startHttp"),
|
||||||
|
"the shutdown hook must be registered before HTTP starts — the one statement that "
|
||||||
|
+ "could not be reordered without changing FleetdRuntime's constructor shape, "
|
||||||
|
+ "see FleetdAssembly's javadoc: " + ports.ledger);
|
||||||
|
assertEquals(ports.ledger.size() - 1, ports.ledger.indexOf("startHttp"),
|
||||||
|
"HTTP must start LAST of everything this fake observes: " + ports.ledger);
|
||||||
|
assertNotNull(ports.startedApp, "FleetdAssembly must have built and handed off a real FleetApp");
|
||||||
|
|
||||||
|
// No real HTTP bind and no real herdr socket: this call returning at all, plus the ledger
|
||||||
|
// above, is the proof — a real bind or a real UnixSocketHerdrClient.connect would have
|
||||||
|
// thrown or hung against the sockets/ports this test never opened.
|
||||||
|
assertNotNull(runtime.app(), "FleetdRuntime must own the same Javalin app that was built");
|
||||||
|
assertEquals(ports.startedApp, runtime.app(), "attachApp must hand FleetdRuntime the SAME instance");
|
||||||
|
|
||||||
|
// --- the runtime owns the real, live objects — not a copy --------------------------------
|
||||||
|
assertEquals(LoopWatchdog.State.RUNNING, runtime.reaper().health(),
|
||||||
|
"SessionReaper must be running: lifecycle.idleTtlSeconds is configured");
|
||||||
|
assertEquals(LoopWatchdog.State.RUNNING, runtime.poller().health(), "StatusPoller must be running");
|
||||||
|
assertNotNull(runtime.heartbeat(), "leadHeartbeat: is configured, so the loop must be built and started");
|
||||||
|
assertNotNull(runtime.healthMonitor(), "health.enabled: true, so the monitor must be built and started");
|
||||||
|
assertNotNull(runtime.configWatcher(), "configReload.enabled: true, so the watcher must be built and started");
|
||||||
|
// Documented gap (see class javadoc): no coordinator: block, so these stay null.
|
||||||
|
assertEquals(null, runtime.leadCoordLoop(), "no coordinator: block — leadCoordLoop must stay unbuilt");
|
||||||
|
assertEquals(null, runtime.leadMailbox(), "no coordinator: block — leadMailbox must stay unbuilt");
|
||||||
|
assertEquals(runtime.replyInbox(), ports.replyInbox,
|
||||||
|
"FleetdRuntime must own the exact ReplyInbox instance this fake's AmqpOpener returned");
|
||||||
|
assertFalse(ports.herdr.closed, "herdr must still be open while the daemon is running");
|
||||||
|
assertFalse(ports.replyInbox.closed, "the reply inbox must still be open while the daemon is running");
|
||||||
|
for (ScheduledExecutorService scheduler : ports.schedulers) {
|
||||||
|
assertFalse(scheduler.isShutdown(), "a scheduler must still be running while the daemon is up");
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- close order: invoke the captured shutdown-hook Runnable directly (no real JVM shutdown
|
||||||
|
// happens in a unit test) and prove every resource this fake can observe is released --------
|
||||||
|
assertNotNull(ports.shutdownHook, "FleetdAssembly must have registered a shutdown hook");
|
||||||
|
ports.shutdownHook.run();
|
||||||
|
|
||||||
|
assertEquals(LoopWatchdog.State.STOPPED, runtime.reaper().health(), "SessionReaper must stop on close");
|
||||||
|
assertEquals(LoopWatchdog.State.STOPPED, runtime.poller().health(), "StatusPoller must stop on close");
|
||||||
|
assertTrue(ports.herdr.closed, "router.close() must close the herdr client last");
|
||||||
|
assertTrue(ports.replyInbox.closed, "the AutoCloseable reply inbox must be closed");
|
||||||
|
for (int i = 0; i < ports.schedulers.size(); i++) {
|
||||||
|
assertTrue(ports.schedulers.get(i).isShutdown(),
|
||||||
|
"scheduler for '" + ports.schedulerPurposes.get(i) + "' must be shut down by close(): "
|
||||||
|
+ "LeadHeartbeatLoop/FleetHealthMonitor/ReplyPushLoop each call "
|
||||||
|
+ "scheduler.shutdownNow() on the exact instance ports.newScheduler(...) handed them");
|
||||||
|
}
|
||||||
|
|
||||||
|
// Calling the captured hook a second time must never happen for a real JVM shutdown hook,
|
||||||
|
// but nothing above should have thrown — that already proves every accessed field's close()
|
||||||
|
// tolerated running once, in the recorded order, without an exception escaping.
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,232 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.mcp.PrimaryRegistry;
|
||||||
|
import dev.ltms.fleet.msg.MessageService;
|
||||||
|
import dev.ltms.fleet.msg.ReplyInbox;
|
||||||
|
import dev.ltms.fleet.msg.ReplyPushLoop;
|
||||||
|
import dev.ltms.fleet.session.SessionManager;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.lang.reflect.Field;
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.function.Consumer;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 rank 6 — {@code FleetdAssembly.java:447} wires {@code
|
||||||
|
* sessions.onRelease(Fleetd.releaseCleanup(messages, replyInbox, primaryRegistry))}. {@link
|
||||||
|
* FleetdReleaseCleanupWiringTest} already pins that the FACTORY {@code Fleetd.releaseCleanup}
|
||||||
|
* itself reaches all three collaborators — but it calls the factory directly, never {@code
|
||||||
|
* FleetdAssembly.assembleAndStart}, so it cannot see whether the real call site at {@code :447}
|
||||||
|
* still registers it (as opposed to a no-op {@code detail -> { }}) or still passes it the REAL,
|
||||||
|
* assembled {@code messages}/{@code replyInbox}/{@code primaryRegistry}. Swapping the registered
|
||||||
|
* listener for a no-op at that call site compiles clean and leaves the whole suite — including the
|
||||||
|
* factory-level test — green: EVERY teardown then leaks a stuck rendezvous waiter, an unreleased
|
||||||
|
* reply-inbox consumer, and a stale lead binding, all three at once.
|
||||||
|
*
|
||||||
|
* <p>This test assembles the real daemon with {@code idleSleepGuard.enabled: false} — the ONLY
|
||||||
|
* other {@code onRelease} registration in {@code FleetdAssembly} (see {@code
|
||||||
|
* dev.ltms.fleet.power.IdleSleepGuard}'s own wiring at {@code FleetdAssembly.java:228}) — so the
|
||||||
|
* real {@link SessionManager}'s release-listener list holds exactly the one listener this call site
|
||||||
|
* registers. It pulls that REAL listener out via reflection (the list itself is private, like
|
||||||
|
* {@code StatusPollerResilienceTest}'s use of the same technique elsewhere in this suite), invokes
|
||||||
|
* it directly, and asserts all three collaborator effects against the REAL, assembled {@link
|
||||||
|
* MessageService} ({@link FleetdRuntime#messages()}), the REAL {@link ReplyInbox} ({@link
|
||||||
|
* FleetdRuntime#replyInbox()}), and the REAL {@link PrimaryRegistry} — reached through {@link
|
||||||
|
* FleetdRuntime#pushLoop()}, the only other accessor that was handed the same {@code
|
||||||
|
* primaryRegistry} instance ({@code FleetdAssembly.java:382}), since {@code FleetMcp} never exposes
|
||||||
|
* it.
|
||||||
|
*/
|
||||||
|
class FleetdAssemblyReleaseCleanupBehaviouralTest {
|
||||||
|
|
||||||
|
private static final String TARGET = "term_a";
|
||||||
|
|
||||||
|
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||||
|
final FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Runnable shutdownHook;
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
return herdr;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
return (uri, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException("no broker: block is configured");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
return (uri, selfCoordId, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException("no coordinator: block is configured");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||||
|
return () -> {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
return Executors.newSingleThreadScheduledExecutor();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
this.shutdownHook = hook;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
idleSleepGuard:
|
||||||
|
enabled: false
|
||||||
|
""");
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
@Test
|
||||||
|
@DisplayName("[BEHAVIOURAL] the real assembled release listener reaches messages.abandon, "
|
||||||
|
+ "replyInbox.release, AND primaryRegistry.forgetDelegation — all three leaks at once")
|
||||||
|
void assembledReleaseListenerReachesAllThreeCollaborators(@TempDir Path dir) throws Exception {
|
||||||
|
FleetConfig cfg = writeConfig(dir);
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||||
|
|
||||||
|
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
assertNotNull(ports.shutdownHook, "FleetdAssembly must have registered a shutdown hook");
|
||||||
|
|
||||||
|
// Surefire runs the whole suite in one JVM fork, so the scheduler/loops this assembly starts
|
||||||
|
// must be torn down here, on the failure path too — hence the try/finally, not just a
|
||||||
|
// statement at the end of the happy path.
|
||||||
|
try {
|
||||||
|
// --- reach into SessionManager's private release-listener list. idleSleepGuard.enabled:
|
||||||
|
// false above means FleetdAssembly.java:228 never registers, so this list must hold EXACTLY
|
||||||
|
// the one listener :447 registers.
|
||||||
|
Field listenersField = SessionManager.class.getDeclaredField("releaseListeners");
|
||||||
|
listenersField.setAccessible(true);
|
||||||
|
List<Consumer<SessionManager.ReleaseDetail>> releaseListeners =
|
||||||
|
(List<Consumer<SessionManager.ReleaseDetail>>) listenersField.get(runtime.sessions());
|
||||||
|
assertEquals(1, releaseListeners.size(), "control: with idleSleepGuard.enabled: false, "
|
||||||
|
+ "FleetdAssembly.java:447 must be the ONLY onRelease registration — a different "
|
||||||
|
+ "count means this test is no longer isolating the call site it claims to pin");
|
||||||
|
Consumer<SessionManager.ReleaseDetail> releaseListener = releaseListeners.get(0);
|
||||||
|
|
||||||
|
MessageService messages = runtime.messages();
|
||||||
|
ReplyInbox replyInbox = runtime.replyInbox();
|
||||||
|
|
||||||
|
// primaryRegistry is never exposed by FleetdRuntime directly — ReplyPushLoop is the other
|
||||||
|
// collaborator FleetdAssembly.java:382 hands the SAME instance to, so reach it from there.
|
||||||
|
Field primaryRegistryField = ReplyPushLoop.class.getDeclaredField("primaryRegistry");
|
||||||
|
primaryRegistryField.setAccessible(true);
|
||||||
|
PrimaryRegistry primaryRegistry = (PrimaryRegistry) primaryRegistryField.get(runtime.pushLoop());
|
||||||
|
assertNotNull(primaryRegistry, "control: the assembled ReplyPushLoop must hold a real "
|
||||||
|
+ "PrimaryRegistry instance");
|
||||||
|
|
||||||
|
// --- loud controls: set up the "before" state each collaborator's effect is measured
|
||||||
|
// against, against the REAL assembled objects. If any of these three fails, the test below
|
||||||
|
// would pass vacuously because the subject it claims to observe never existed in the first
|
||||||
|
// place.
|
||||||
|
String ticket = messages.sendAsync(TARGET, "long task");
|
||||||
|
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase(),
|
||||||
|
"control: the async ticket must be PENDING before the release listener runs");
|
||||||
|
|
||||||
|
replyInbox.own(TARGET);
|
||||||
|
replyInbox.publish(TARGET, "msg-1", "hello");
|
||||||
|
assertEquals(1, replyInbox.peek(TARGET).size(),
|
||||||
|
"control: the reply inbox must own TARGET and hold one message before the release "
|
||||||
|
+ "listener runs");
|
||||||
|
|
||||||
|
primaryRegistry.recordDelegation(TARGET, "lead-1");
|
||||||
|
assertEquals("lead-1", primaryRegistry.nudgeTargetFor(TARGET).orElse(null),
|
||||||
|
"control: the delegation must be recorded before the release listener runs");
|
||||||
|
|
||||||
|
// --- the one call under test: invoke the REAL, assembled release listener directly, the
|
||||||
|
// same way SessionManager.release(...) would on a real teardown.
|
||||||
|
releaseListener.accept(new SessionManager.ReleaseDetail(TARGET, null, null, null, null));
|
||||||
|
|
||||||
|
MessageService.TaskView after = awaitTerminal(messages, ticket);
|
||||||
|
assertEquals(MessageService.Phase.FAILED, after.phase(),
|
||||||
|
"FleetdAssembly.java:447 must register a listener that calls messages.abandon(...) "
|
||||||
|
+ "on the SAME assembled MessageService — an inert listener leaves this "
|
||||||
|
+ "ticket PENDING for the full 30-minute async timeout");
|
||||||
|
assertTrue(after.detail() != null && after.detail().contains("released"),
|
||||||
|
"the abandon reason must say the worker session was released");
|
||||||
|
|
||||||
|
assertTrue(replyInbox.peek(TARGET).isEmpty(),
|
||||||
|
"FleetdAssembly.java:447 must register a listener that calls replyInbox.release(...) "
|
||||||
|
+ "— an inert listener leaves the inbox still owning TARGET with its message");
|
||||||
|
|
||||||
|
assertTrue(primaryRegistry.nudgeTargetFor(TARGET).isEmpty(),
|
||||||
|
"FleetdAssembly.java:447 must register a listener that calls "
|
||||||
|
+ "primaryRegistry.forgetDelegation(...) — an inert listener leaves the stale "
|
||||||
|
+ "delegation in place");
|
||||||
|
} finally {
|
||||||
|
// Proof the teardown actually ran, not just an assurance that a finally was added: the
|
||||||
|
// captured shutdown hook's close order (FleetdAssemblyLifecycleTest) closes the herdr
|
||||||
|
// client last, so ports.herdr.closed flips to true only if this hook really executed.
|
||||||
|
ports.shutdownHook.run();
|
||||||
|
assertTrue(ports.herdr.closed, "the captured shutdown hook must have run and closed herdr — "
|
||||||
|
+ "proof this test's assembled background loops/scheduler were torn down");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static MessageService.TaskView awaitTerminal(MessageService messages, String ticket)
|
||||||
|
throws InterruptedException {
|
||||||
|
long deadline = System.currentTimeMillis() + 5000;
|
||||||
|
MessageService.TaskView view = messages.poll(ticket);
|
||||||
|
while (view.phase() == MessageService.Phase.PENDING && System.currentTimeMillis() < deadline) {
|
||||||
|
//noinspection BusyWait
|
||||||
|
Thread.sleep(10);
|
||||||
|
view = messages.poll(ticket);
|
||||||
|
}
|
||||||
|
return view;
|
||||||
|
}
|
||||||
|
}
|
||||||
+241
@@ -0,0 +1,241 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||||
|
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.lang.reflect.Field;
|
||||||
|
import java.lang.reflect.Method;
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #630, and fleetd #612's ranks for this call site. {@code FleetdAssembly.java:402}
|
||||||
|
* computes {@code requireOperatorConfirm} from the effective {@code leadRollover.requireOperatorConfirm}
|
||||||
|
* config, and {@code :409} threads it as the 14th argument into the full {@link LeadHeartbeatLoop}
|
||||||
|
* constructor. Measured on 26f1986: dropping that one argument so the 13-argument overload is
|
||||||
|
* selected instead (it delegates with {@code true} hardcoded — see that overload's own javadoc,
|
||||||
|
* fleetd #621) compiles with 0 errors and leaves all 1883 tests green, both with and without the
|
||||||
|
* argument. In production this means the daemon keeps starting and keeps nudging, but the
|
||||||
|
* context-high notice silently goes back to telling EVERY lead to ask the operator before a
|
||||||
|
* context roll — on a host that set {@code requireOperatorConfirm: false} specifically so it would
|
||||||
|
* not have to. That is the operator's own fix silently reverting, with a fully green suite.
|
||||||
|
*
|
||||||
|
* <p>{@code LeadHeartbeatLoopTest} already proves {@link LeadHeartbeatLoop}'s package-private
|
||||||
|
* {@code contextNotice(boolean, LeadContextGauge.Reading, boolean, boolean)} branches correctly on
|
||||||
|
* its own {@code requireOperatorConfirm} argument — that the METHOD works. It says nothing about
|
||||||
|
* which value {@code FleetdAssembly} actually passes into the constructed loop, so it is not
|
||||||
|
* reused here as coverage for the call site.
|
||||||
|
*
|
||||||
|
* <p>This test assembles the real daemon TWICE — once with {@code leadRollover.requireOperatorConfirm:
|
||||||
|
* false}, once with {@code true} — pulls the REAL {@code requireOperatorConfirm} field out of the
|
||||||
|
* REAL, assembled {@link LeadHeartbeatLoop} each time (reflection: the field, and {@code
|
||||||
|
* contextNotice} itself, are package-private to {@code dev.ltms.fleet.msg}, and nothing public
|
||||||
|
* exposes either — the same technique {@code StatusPollerResilienceTest} already uses in this
|
||||||
|
* suite), and calls the REAL {@code contextNotice} method with that field's value to produce the
|
||||||
|
* actual notice text the assembled loop would append to a nudge. Both directions are asserted: a
|
||||||
|
* one-directional test here would pass on a constant.
|
||||||
|
*/
|
||||||
|
class FleetdAssemblyRequireOperatorConfirmBehaviouralTest {
|
||||||
|
|
||||||
|
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||||
|
final FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Runnable shutdownHook;
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
return herdr;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
return (uri, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException("no broker: block is configured");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
return (uri, selfCoordId, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException("no coordinator: block is configured");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||||
|
return () -> {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
return Executors.newSingleThreadScheduledExecutor();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
this.shutdownHook = hook;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir, boolean requireOperatorConfirm) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
idleSleepGuard:
|
||||||
|
enabled: false
|
||||||
|
leadHeartbeat:
|
||||||
|
idleAfterSeconds: 600
|
||||||
|
backoffMs: 15000
|
||||||
|
quietNudgeCap: 5
|
||||||
|
leadRollover:
|
||||||
|
handoverPath: handover.md
|
||||||
|
requireOperatorConfirm: %s
|
||||||
|
""".formatted(requireOperatorConfirm));
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Carries both the assembled loop under test AND its {@link RecordingResourcePorts}, so the
|
||||||
|
* caller can tear the assembly down (this test assembles the real daemon TWICE — see the class
|
||||||
|
* javadoc — and each assembly needs its own teardown, not just the last one). */
|
||||||
|
private record Assembled(LeadHeartbeatLoop heartbeat, RecordingResourcePorts ports) {
|
||||||
|
}
|
||||||
|
|
||||||
|
private static Assembled assembleHeartbeat(Path dir, boolean requireOperatorConfirm) throws Exception {
|
||||||
|
FleetConfig cfg = writeConfig(dir, requireOperatorConfirm);
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||||
|
|
||||||
|
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
assertNotNull(ports.shutdownHook, "FleetdAssembly must have registered a shutdown hook");
|
||||||
|
|
||||||
|
LeadHeartbeatLoop heartbeat = runtime.heartbeat();
|
||||||
|
assertNotNull(heartbeat, "control: leadHeartbeat: is configured, so FleetdAssembly.assembleAndStart "
|
||||||
|
+ "must have built a real LeadHeartbeatLoop");
|
||||||
|
return new Assembled(heartbeat, ports);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Pulls the REAL {@code requireOperatorConfirm} field off the REAL, assembled loop. */
|
||||||
|
private static boolean assembledRequireOperatorConfirm(LeadHeartbeatLoop heartbeat) throws Exception {
|
||||||
|
Field field = LeadHeartbeatLoop.class.getDeclaredField("requireOperatorConfirm");
|
||||||
|
field.setAccessible(true);
|
||||||
|
return field.getBoolean(heartbeat);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Calls the REAL, package-private {@code contextNotice(boolean, Reading, boolean, boolean)} via reflection. */
|
||||||
|
private static String contextNotice(boolean enabled, LeadContextGauge.Reading reading, boolean alreadyNotified,
|
||||||
|
boolean requireOperatorConfirm) throws Exception {
|
||||||
|
Method method = LeadHeartbeatLoop.class.getDeclaredMethod("contextNotice", boolean.class,
|
||||||
|
LeadContextGauge.Reading.class, boolean.class, boolean.class);
|
||||||
|
method.setAccessible(true);
|
||||||
|
return (String) method.invoke(null, enabled, reading, alreadyNotified, requireOperatorConfirm);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[BEHAVIOURAL] the real assembled LeadHeartbeatLoop's context-high notice tracks "
|
||||||
|
+ "leadRollover.requireOperatorConfirm — BOTH directions")
|
||||||
|
void assembledRequireOperatorConfirmControlsNoticeWording(@TempDir Path dir) throws Exception {
|
||||||
|
LeadContextGauge.Reading highReading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH,
|
||||||
|
250_000L, 2);
|
||||||
|
|
||||||
|
String noticeFalse;
|
||||||
|
String noticeTrue;
|
||||||
|
|
||||||
|
// --- direction 1: requireOperatorConfirm: false -----------------------------------------
|
||||||
|
Path falseDir = dir.resolve("false");
|
||||||
|
Files.createDirectories(falseDir);
|
||||||
|
Assembled assembledFalse = assembleHeartbeat(falseDir, false);
|
||||||
|
// Surefire runs the whole suite in one JVM fork, so each assembly's scheduler/loops must be
|
||||||
|
// torn down here, on the failure path too — hence try/finally per assembly (this test
|
||||||
|
// assembles TWICE, so both need their own teardown, not just the last one).
|
||||||
|
try {
|
||||||
|
boolean fieldFalse = assembledRequireOperatorConfirm(assembledFalse.heartbeat());
|
||||||
|
assertFalse(fieldFalse, "FleetdAssembly.java:402/:409 must thread leadRollover."
|
||||||
|
+ "requireOperatorConfirm: false into the assembled LeadHeartbeatLoop's own field — "
|
||||||
|
+ "dropping the 14th constructor argument selects the 13-argument overload, which "
|
||||||
|
+ "hardcodes true regardless of config (fleetd #621), and this would read true instead");
|
||||||
|
|
||||||
|
noticeFalse = contextNotice(true, highReading, false, fieldFalse);
|
||||||
|
assertTrue(noticeFalse.contains("Decide for yourself when to confirm"),
|
||||||
|
"with requireOperatorConfirm: false, the assembled loop's own notice must tell the "
|
||||||
|
+ "lead it can decide for itself — got: " + noticeFalse);
|
||||||
|
assertFalse(noticeFalse.contains("ask the operator") || noticeFalse.contains("Only the operator"),
|
||||||
|
"with requireOperatorConfirm: false, the assembled loop's own notice must NOT ask the "
|
||||||
|
+ "operator — got: " + noticeFalse);
|
||||||
|
} finally {
|
||||||
|
// Proof the teardown actually ran, not just an assurance that a finally was added: the
|
||||||
|
// captured shutdown hook's close order (FleetdAssemblyLifecycleTest) closes the herdr
|
||||||
|
// client last, so ports.herdr.closed flips to true only if this hook really executed.
|
||||||
|
assembledFalse.ports().shutdownHook.run();
|
||||||
|
assertTrue(assembledFalse.ports().herdr.closed, "the captured shutdown hook must have run "
|
||||||
|
+ "and closed herdr — proof this assembly's background loops/scheduler were torn down");
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- direction 2: requireOperatorConfirm: true -------------------------------------------
|
||||||
|
Path trueDir = dir.resolve("true");
|
||||||
|
Files.createDirectories(trueDir);
|
||||||
|
Assembled assembledTrue = assembleHeartbeat(trueDir, true);
|
||||||
|
try {
|
||||||
|
boolean fieldTrue = assembledRequireOperatorConfirm(assembledTrue.heartbeat());
|
||||||
|
assertTrue(fieldTrue, "FleetdAssembly.java:402/:409 must thread leadRollover."
|
||||||
|
+ "requireOperatorConfirm: true into the assembled LeadHeartbeatLoop's own field");
|
||||||
|
|
||||||
|
noticeTrue = contextNotice(true, highReading, false, fieldTrue);
|
||||||
|
assertTrue(noticeTrue.contains("ask the operator") && noticeTrue.contains("Only the operator can approve the roll"),
|
||||||
|
"with requireOperatorConfirm: true, the assembled loop's own notice must ask the "
|
||||||
|
+ "operator — got: " + noticeTrue);
|
||||||
|
assertFalse(noticeTrue.contains("Decide for yourself when to confirm"),
|
||||||
|
"with requireOperatorConfirm: true, the assembled loop's own notice must NOT tell "
|
||||||
|
+ "the lead it can decide for itself — got: " + noticeTrue);
|
||||||
|
} finally {
|
||||||
|
assembledTrue.ports().shutdownHook.run();
|
||||||
|
assertTrue(assembledTrue.ports().herdr.closed, "the captured shutdown hook must have run "
|
||||||
|
+ "and closed herdr — proof this assembly's background loops/scheduler were torn down");
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- the two directions must actually differ: a constant return would pass both assertion
|
||||||
|
// blocks above vacuously if they happened to share wording, so compare them directly too.
|
||||||
|
assertTrue(!noticeFalse.equals(noticeTrue),
|
||||||
|
"the two directions must produce genuinely different notice text — got the same "
|
||||||
|
+ "text for both: " + noticeFalse);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,173 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import ch.qos.logback.classic.Level;
|
||||||
|
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.msg.ReplyInbox;
|
||||||
|
import dev.ltms.fleet.testing.CapturedLog;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 A-gaps (gap 2): {@code Fleetd.java:185} used to call {@code
|
||||||
|
* reportRoleFallbackGaps(cfg)} <em>outside</em> the boundary {@code FleetdAssembly.assembleAndStart}
|
||||||
|
* — the assembly call itself sat at line 201, after it — so nothing that drives the assembly (the
|
||||||
|
* seam {@code FleetdAssemblyLifecycleTest} exercises) could ever notice the call being deleted.
|
||||||
|
* {@code Fleetd.main} still refuses a bad config end to end (see {@code
|
||||||
|
* FleetdStartupValidationTest}), but that only pins {@code validateAll()} and {@code
|
||||||
|
* assertChartersNameOnlyRegisteredTools} — a validator that <em>throws</em>. {@code
|
||||||
|
* reportRoleFallbackGaps} only logs; nothing about {@code main} throwing or not throwing can
|
||||||
|
* observe whether that particular call ran.
|
||||||
|
*
|
||||||
|
* <p>This drives {@link FleetdAssembly#assembleAndStart} directly — never a copy of its logic —
|
||||||
|
* with a config that has no {@code fleet:} pools or charters configured for any role, so every role
|
||||||
|
* trips both of {@code reportRoleFallbackGaps}' log branches, and asserts on the real log line a
|
||||||
|
* {@code ListAppender} attached to the shared {@code Fleetd}/{@code FleetdAssembly} logger
|
||||||
|
* captures. Deleting the call from {@code FleetdAssembly} (verified by hand, see the ticket) turns
|
||||||
|
* this test red; deleting it from {@code Fleetd.main} instead (its old location) would not, which
|
||||||
|
* is exactly the gap this test closes.
|
||||||
|
*/
|
||||||
|
class FleetdAssemblyRoleFallbackBoundaryTest {
|
||||||
|
|
||||||
|
/** Minimal fake {@link ResourcePorts}: enough for {@code assembleAndStart} to run with no real I/O. */
|
||||||
|
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||||
|
|
||||||
|
final FakeHerdr herdr = new FakeHerdr();
|
||||||
|
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||||
|
Runnable shutdownHook;
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
return herdr;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
return (uri, prefetch) -> replyInbox;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
return (uri, selfCoordId, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException(
|
||||||
|
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
return Executors.newSingleThreadScheduledExecutor();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
this.shutdownHook = hook;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
// No real HTTP bind in a unit test.
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||||
|
return () -> {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static final class SentinelReplyInbox implements ReplyInbox {
|
||||||
|
@Override
|
||||||
|
public void own(String target) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void release(String target) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void publish(String target, String msgId, String content) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public List<InboxMessage> peek(String target) {
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public boolean ack(String target, String msgId) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
// Deliberately no `fleet:` block at all: every MemberRole has neither a pool nor a
|
||||||
|
// charter, so reportRoleFallbackGaps' "role fallback: no fleet.<role>s: pool for ..."
|
||||||
|
// branch is guaranteed to log something to assert on.
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
idleSleepGuard:
|
||||||
|
enabled: false
|
||||||
|
""");
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void assembleAndStartItselfReportsTheRoleFallbackGaps(@TempDir Path dir) throws Exception {
|
||||||
|
FleetConfig cfg = writeConfig(dir);
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||||
|
|
||||||
|
try (var captured = CapturedLog.at(Fleetd.class, Level.INFO)) {
|
||||||
|
FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
|
||||||
|
String infoLines = captured.events().stream()
|
||||||
|
.filter(e -> e.getLevel() == Level.INFO)
|
||||||
|
.map(ILoggingEvent::getFormattedMessage)
|
||||||
|
.reduce("", (a, b) -> a + "\n" + b);
|
||||||
|
assertTrue(infoLines.contains("role fallback: no fleet.<role>s: pool for"),
|
||||||
|
() -> "FleetdAssembly.assembleAndStart itself must call reportRoleFallbackGaps "
|
||||||
|
+ "(fleetd #612 A-gaps gap 2) — captured INFO lines: " + infoLines);
|
||||||
|
} finally {
|
||||||
|
if (ports.shutdownHook != null) {
|
||||||
|
ports.shutdownHook.run();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,207 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.msg.ReplyInbox;
|
||||||
|
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.OptionalLong;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 B3 — replaces {@code FleetdBackendQuarantineWiringTest} (fleetd #466), a source-text
|
||||||
|
* test that scraped {@code Fleetd.java} (now {@code FleetdAssembly.java}, moved there by fleetd #612
|
||||||
|
* Unit A) for the {@code BackendQuarantine.withEscalation(...)} call, and separately asserted the
|
||||||
|
* flat two-argument constructor's text was ABSENT. That proves the right method NAME appears in
|
||||||
|
* source; it proves nothing about what the constructed object actually DOES.
|
||||||
|
*
|
||||||
|
* <p>This test instead drives the REAL {@link BackendQuarantine} the real {@link
|
||||||
|
* FleetdAssembly#assembleAndStart} builds — reached through {@link
|
||||||
|
* dev.ltms.fleet.mcp.FleetMcp#quarantineSource()} on the real, live {@code FleetMcp} {@code
|
||||||
|
* FleetdRuntime} owns — and asserts the ONE behavioural difference {@code withEscalation} and the
|
||||||
|
* flat constructor actually produce (see {@link BackendQuarantine}'s own class doc, "Mechanism"):
|
||||||
|
* quarantining the same credential twice in a row, within one base cooldown of the first deadline,
|
||||||
|
* must escalate the second cooldown past the first. A flat instance reports the identical cooldown
|
||||||
|
* both times.
|
||||||
|
*/
|
||||||
|
class FleetdBackendQuarantineAssemblyTest {
|
||||||
|
|
||||||
|
/** Base cooldown used throughout — long enough that rounding never blurs the 2x escalation. */
|
||||||
|
private static final int COOLDOWN_SECONDS = 100;
|
||||||
|
|
||||||
|
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||||
|
|
||||||
|
final FakeHerdr herdr = new FakeHerdr();
|
||||||
|
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||||
|
final AtomicLong clockNanos = new AtomicLong(0L);
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
return herdr;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
return (uri, prefetch) -> replyInbox;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
return (uri, selfCoordId, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException(
|
||||||
|
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
// Controllable: the SAME LongSupplier instance BackendQuarantine.withEscalation(...) is
|
||||||
|
// built with, so advancing clockNanos after assembly moves the quarantine tracker's own
|
||||||
|
// clock, with no real sleep needed to observe escalation.
|
||||||
|
return clockNanos::get;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return clockNanos::get;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
return Executors.newSingleThreadScheduledExecutor();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||||
|
return () -> {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
|
||||||
|
@Override
|
||||||
|
public void own(String target) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void release(String target) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void publish(String target, String msgId, String content) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public List<InboxMessage> peek(String target) {
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public boolean ack(String target, String msgId) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
idleSleepGuard:
|
||||||
|
enabled: false
|
||||||
|
broker:
|
||||||
|
uri: "amqp://fake-test-broker/vh"
|
||||||
|
quarantineCooldownSeconds: %d
|
||||||
|
""".formatted(COOLDOWN_SECONDS));
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[BEHAVIOURAL] the real assembled BackendQuarantine escalates a repeated exhaustion, "
|
||||||
|
+ "which the flat two-argument constructor can never do")
|
||||||
|
void assembledQuarantineEscalatesOnARepeatedExhaustion(@TempDir Path dir) throws Exception {
|
||||||
|
FleetConfig cfg = writeConfig(dir);
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||||
|
|
||||||
|
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
|
||||||
|
BackendQuarantine quarantine = runtime.mcp().quarantineSource().quarantine();
|
||||||
|
|
||||||
|
// First exhaustion, at clock=0: a fresh occurrence, blocked for exactly the base cooldown.
|
||||||
|
quarantine.quarantine("cred-x");
|
||||||
|
BackendQuarantine.Status first = quarantine.status("cred-x").orElseThrow(
|
||||||
|
() -> new AssertionError("credential must be quarantined immediately after quarantine()"));
|
||||||
|
assertEquals(1, first.repeatCount(), "the first call is repeat #1");
|
||||||
|
assertEquals(COOLDOWN_SECONDS, first.remainingSeconds(),
|
||||||
|
"a fresh quarantine blocks for exactly the base cooldown");
|
||||||
|
|
||||||
|
// Second exhaustion, arriving just after the first deadline — well within one base cooldown
|
||||||
|
// of it, so this is a CONTINUATION of the same streak (repeat #2), not a fresh occurrence.
|
||||||
|
long firstDeadlineNanos = COOLDOWN_SECONDS * 1_000_000_000L;
|
||||||
|
ports.clockNanos.set(firstDeadlineNanos + 1);
|
||||||
|
quarantine.quarantine("cred-x");
|
||||||
|
BackendQuarantine.Status second = quarantine.status("cred-x").orElseThrow(
|
||||||
|
() -> new AssertionError("credential must be quarantined immediately after the second "
|
||||||
|
+ "quarantine() call"));
|
||||||
|
assertEquals(2, second.repeatCount(), "the second call, arriving within one base cooldown of "
|
||||||
|
+ "the first deadline, continues the streak as repeat #2");
|
||||||
|
|
||||||
|
// The one behavioural difference: withEscalation doubles the cooldown on repeat #2 (capped
|
||||||
|
// well above this at 12x base), the flat two-argument constructor never grows past the base
|
||||||
|
// cooldown no matter how many times quarantine() is called in a row.
|
||||||
|
assertEquals(2 * COOLDOWN_SECONDS, second.remainingSeconds(),
|
||||||
|
"withEscalation's default backoff doubles the cooldown on the second consecutive "
|
||||||
|
+ "exhaustion — this is the exact call FleetdAssembly.java makes at the "
|
||||||
|
+ "BackendQuarantine.withEscalation(...) call site");
|
||||||
|
assertTrue(second.remainingSeconds() > first.remainingSeconds(),
|
||||||
|
"the flat two-argument BackendQuarantine constructor would report the SAME remaining "
|
||||||
|
+ "seconds both times — this inequality is what a mutation to the flat "
|
||||||
|
+ "constructor at that call site must fail");
|
||||||
|
|
||||||
|
// Also confirm isQuarantined/remainingSeconds agree, exercising the accessors a real caller
|
||||||
|
// (fleet_profiles / fleet_list, per BackendQuarantine's own class doc) actually reads.
|
||||||
|
assertTrue(quarantine.isQuarantined("cred-x"));
|
||||||
|
OptionalLong remaining = quarantine.remainingSeconds("cred-x");
|
||||||
|
assertTrue(remaining.isPresent());
|
||||||
|
assertEquals(2 * COOLDOWN_SECONDS, remaining.getAsLong());
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,65 +0,0 @@
|
|||||||
package dev.ltms.fleet;
|
|
||||||
|
|
||||||
import java.nio.file.Files;
|
|
||||||
import java.nio.file.Path;
|
|
||||||
import org.junit.jupiter.api.DisplayName;
|
|
||||||
import org.junit.jupiter.api.Test;
|
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
|
||||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
|
||||||
|
|
||||||
/**
|
|
||||||
* fleetd #466 follow-up: {@code Fleetd.main} builds the daemon's one {@code BackendQuarantine}
|
|
||||||
* from {@link dev.ltms.fleet.placement.BackendQuarantine#withEscalation(java.util.function.LongSupplier,
|
|
||||||
* long)} — the escalating factory — rather than the plain two-argument constructor, which is still a
|
|
||||||
* flat cooldown (kept for backward compatibility, see that class's doc). {@code
|
|
||||||
* BackendQuarantineTest} proves {@code withEscalation} itself escalates, is ceilinged, and resets;
|
|
||||||
* it says nothing about which one {@code main} actually calls.
|
|
||||||
*
|
|
||||||
* <p>Measured directly: reverting {@code main} to {@code new BackendQuarantine(System::nanoTime,
|
|
||||||
* TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()))} — the pre-#466 flat call — compiles
|
|
||||||
* with 0 errors and leaves the entire 1608-test suite (including every {@code BackendQuarantineTest}
|
|
||||||
* case) green, because no other test constructs its {@code BackendQuarantine} through {@code main};
|
|
||||||
* every one of them builds its own instance directly. That silent regression is exactly the shape
|
|
||||||
* {@link FleetdLeadSeatWiringTest} and {@link FleetdCompletionResolverWiringTest} already guard
|
|
||||||
* against for their own constructor arguments — this is the same class of gap for fleetd #466's
|
|
||||||
* factory choice, following their approach.
|
|
||||||
*
|
|
||||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
|
||||||
* BackendQuarantine} and never runs {@code main} — a green result here proves only that the exact
|
|
||||||
* text {@code main} calls {@code BackendQuarantine.withEscalation(...)} rather than the flat
|
|
||||||
* constructor. It does not prove that call actually executes at startup (no test here starts the
|
|
||||||
* daemon), and it does not prove the escalation reaches a real backend or credential — only
|
|
||||||
* {@code BackendQuarantineTest} proves the factory's own behaviour, and only a live daemon proves
|
|
||||||
* the wiring runs.
|
|
||||||
*/
|
|
||||||
class FleetdBackendQuarantineWiringTest {
|
|
||||||
|
|
||||||
private static String fleetdSource() throws Exception {
|
|
||||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
|
||||||
}
|
|
||||||
|
|
||||||
@Test
|
|
||||||
@DisplayName("[SOURCE TEXT] main's BackendQuarantine local is still built from BackendQuarantine.withEscalation(...)")
|
|
||||||
void mainStillWiresTheEscalatingQuarantineFactory() throws Exception {
|
|
||||||
String source = fleetdSource();
|
|
||||||
assertTrue(source.contains(
|
|
||||||
"BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,\n"
|
|
||||||
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
|
|
||||||
"Fleetd.main's BackendQuarantine local must still be built from "
|
|
||||||
+ "BackendQuarantine.withEscalation(System::nanoTime, "
|
|
||||||
+ "TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds())). Reverting to the flat "
|
|
||||||
+ "two-argument constructor (fleetd #466's measured regression) compiles with 0 errors "
|
|
||||||
+ "and leaves the whole suite green, including every BackendQuarantineTest case that "
|
|
||||||
+ "proves the escalation itself works — this source check is what must go red instead. "
|
|
||||||
+ "A reverted daemon would go back to retrying a weekly subscription limit on every "
|
|
||||||
+ "flat ~30-minute cooldown, about 336 times across the week.");
|
|
||||||
|
|
||||||
// Negative form of the same check: the pre-#466 flat call, if it ever reappears at this
|
|
||||||
// declaration, must not be mistaken for the escalating one by a looser positive-only check.
|
|
||||||
assertFalse(source.contains(
|
|
||||||
"BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,\n"
|
|
||||||
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
|
|
||||||
"main's BackendQuarantine local must never regress to the flat two-argument constructor");
|
|
||||||
}
|
|
||||||
}
|
|
||||||
@@ -0,0 +1,84 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.AgentControl;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||||
|
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.Set;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #589 Group 2: {@link Fleetd#claudeCodeLauncher} is the factory that replaced {@code
|
||||||
|
* main}'s inline {@code new ClaudeCodeLauncher(...)} call, whose 10th argument is the CB-596 {@code
|
||||||
|
* memberCredentials} policy supplier ({@code () -> config.get().memberCredentials()}). Before this
|
||||||
|
* ticket that argument was untestable wiring: replacing it with {@code () -> null} compiled with 0
|
||||||
|
* errors and left every existing test green, since no existing test builds the exact object {@code
|
||||||
|
* main} wires and then spawns it. {@code memberCredentials} being a lambda is never itself {@code
|
||||||
|
* null}, so {@link dev.ltms.fleet.member.HerdrPeerLauncher#applyMemberCredentialPolicy} sees {@code
|
||||||
|
* memberCredentials.get() == null} and silently shadows nothing — reopening the exact CB-592
|
||||||
|
* exposure gap CB-596's policy closed (gitea issue #82).
|
||||||
|
*
|
||||||
|
* <p>This test drives the factory with a real {@link ConfigRef} carrying a {@code
|
||||||
|
* memberCredentials:} block, spawns through the resulting launcher, and inspects what {@code
|
||||||
|
* tab.create} actually carried — the same observable surface {@code ClaudeCodeLauncherTest}'s
|
||||||
|
* {@code everyKnownNameNotAllowedIsShadowedWithTheSentinel} uses for the launcher's own credential
|
||||||
|
* policy, applied here to prove {@code main}'s wiring reaches it.
|
||||||
|
*/
|
||||||
|
class FleetdClaudeCodeLauncherCredentialWiringTest {
|
||||||
|
|
||||||
|
private static final String YAML = """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
herdrSocket: ~/.config/herdr/herdr.sock
|
||||||
|
profiles:
|
||||||
|
ltms-local:
|
||||||
|
baseUrl: http://gx00.gw:8000
|
||||||
|
model: coder
|
||||||
|
memberCredentials:
|
||||||
|
policy: deny-by-default
|
||||||
|
known:
|
||||||
|
- GITEA_ACCESS_TOKEN
|
||||||
|
""";
|
||||||
|
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
private static Map<String, String> startEnv(FakeHerdr herdr) {
|
||||||
|
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("main's memberCredentials wiring reaches ClaudeCodeLauncher: a known-but-not-allowed "
|
||||||
|
+ "name is shadowed on spawn")
|
||||||
|
void memberCredentialsWiringReachesClaudeCodeLauncher(@TempDir Path dir) throws Exception {
|
||||||
|
Path file = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(file, YAML);
|
||||||
|
FleetConfig cfg = FleetConfig.load(file);
|
||||||
|
ConfigRef config = ConfigRef.fixed(cfg);
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
|
||||||
|
ClaudeCodeLauncher launcher = Fleetd.claudeCodeLauncher(new AgentControl(herdr),
|
||||||
|
new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
|
||||||
|
cfg.profiles(), cfg, config);
|
||||||
|
launcher.spawn();
|
||||||
|
|
||||||
|
String shadowed = startEnv(herdr).get("GITEA_ACCESS_TOKEN");
|
||||||
|
assertNotNull(shadowed,
|
||||||
|
"GITEA_ACCESS_TOKEN is 'known' but not 'allow'-ed in the loaded config — it must be "
|
||||||
|
+ "explicitly shadowed on spawn; replacing the memberCredentials supplier with "
|
||||||
|
+ "() -> null at the Fleetd.claudeCodeLauncher call site must fail this "
|
||||||
|
+ "assertion, since a null policy shadows nothing");
|
||||||
|
assertFalse(shadowed.isBlank(), "the overlay value must be non-blank");
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,315 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.inject.CompletionResolver;
|
||||||
|
import dev.ltms.fleet.msg.Rendezvous;
|
||||||
|
import dev.ltms.fleet.msg.TurnToken;
|
||||||
|
import dev.ltms.fleet.placement.PlacementException;
|
||||||
|
import dev.ltms.fleet.session.MemberSession;
|
||||||
|
import dev.ltms.fleet.session.WorktreeRequest;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.CompletableFuture;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 step 2, Unit B1 — replaces {@code FleetdCompletionResolverWiringTest} (deleted in
|
||||||
|
* this same commit), whose four tests read {@code Fleetd.java}'s source text and asserted the
|
||||||
|
* {@code CompletionResolver} construction call still named the right arguments. That proved the
|
||||||
|
* call site's spelling, never that the assembled resolver actually behaves differently when an
|
||||||
|
* argument is dropped.
|
||||||
|
*
|
||||||
|
* <p>These tests drive {@link FleetdAssembly#assembleAndStart} — the real boot composition,
|
||||||
|
* fleetd #612 Unit A — and read {@link FleetdRuntime#completion()}: the exact {@link
|
||||||
|
* CompletionResolver} instance the assembled daemon uses, never a copy built alongside it for the
|
||||||
|
* test's benefit. Two behaviours are pinned, matching the ticket's own two measured mutations:
|
||||||
|
*
|
||||||
|
* <ul>
|
||||||
|
* <li>the 8th constructor argument ({@code Fleetd.worktreeBranchLookup(sessions::roster)}) —
|
||||||
|
* {@link #assembledResolverReportsTheMembersWorktreeAndBranchInAFallbackReport}; and</li>
|
||||||
|
* <li>the 5th/6th arguments ({@code backendErrorPatterns}, {@code backendErrorSink}, both
|
||||||
|
* assigned from {@code Fleetd}'s extracted factories rather than an inline lambda) —
|
||||||
|
* {@link #assembledResolverClassifiesAndCoolsOffOnAConfiguredBackendErrorPattern}.</li>
|
||||||
|
* </ul>
|
||||||
|
*
|
||||||
|
* <p>Both tests bypass {@link dev.ltms.fleet.inject.StatusPoller} and drive {@link
|
||||||
|
* CompletionResolver#onDelivered} / {@link CompletionResolver#resolveBeforePostAction} directly —
|
||||||
|
* the same public, synchronous entry points {@code CompletionResolverTest} uses — with a
|
||||||
|
* hand-built {@link CompletableFuture} waiter, so no real poller loop or herdr status poll is
|
||||||
|
* needed. The pane scrape comes from {@link FakeHerdr#readText}; the elapsed-time floor
|
||||||
|
* ({@code CompletionResolver.MIN_TURN_NANOS}) is controlled via a fake, advanceable {@link
|
||||||
|
* ResourcePorts#nanoClock()} rather than a real sleep.
|
||||||
|
*/
|
||||||
|
class FleetdCompletionResolverAssemblyTest {
|
||||||
|
|
||||||
|
/** Same shape as {@code FleetdAssemblyLifecycleTest}'s fake, plus a nanoClock this test can advance. */
|
||||||
|
private static final class ControllableResourcePorts implements ResourcePorts {
|
||||||
|
|
||||||
|
final FakeHerdr herdr;
|
||||||
|
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L); // arbitrary non-zero start
|
||||||
|
Runnable shutdownHook;
|
||||||
|
|
||||||
|
ControllableResourcePorts(FakeHerdr herdr) {
|
||||||
|
this.herdr = herdr;
|
||||||
|
}
|
||||||
|
|
||||||
|
void advanceSeconds(long seconds) {
|
||||||
|
nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(seconds));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
return herdr;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
// Never invoked: this test's config has no `broker:` block, so Fleetd.selectReplyInbox
|
||||||
|
// returns the in-memory inbox before calling the opener at all.
|
||||||
|
return (uri, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException("replyInboxOpener must not be called — no broker: block");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
// Never invoked: no `coordinator:` block configured either.
|
||||||
|
return (uri, selfCoordId, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException("leadMailboxOpener must not be called — no coordinator: block");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
return nowNanos::get;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return nowNanos::get;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
return Executors.newSingleThreadScheduledExecutor();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
this.shutdownHook = hook;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
// Deliberately never bind a real port.
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||||
|
return () -> {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir, String profilesYaml, String extraGuardHost,
|
||||||
|
String worktreeRootYamlLine) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
idleSleepGuard:
|
||||||
|
enabled: false
|
||||||
|
%s
|
||||||
|
profiles:
|
||||||
|
%s
|
||||||
|
guard:
|
||||||
|
offSubscriptionHosts:
|
||||||
|
- %s
|
||||||
|
""".formatted(worktreeRootYamlLine == null ? "" : worktreeRootYamlLine, profilesYaml, extraGuardHost));
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static void gitQuiet(Path cwd, String... args) throws Exception {
|
||||||
|
List<String> cmd = new java.util.ArrayList<>(List.of("git"));
|
||||||
|
cmd.addAll(List.of(args));
|
||||||
|
Process p = new ProcessBuilder(cmd).directory(cwd.toFile()).redirectErrorStream(true).start();
|
||||||
|
String out = new String(p.getInputStream().readAllBytes());
|
||||||
|
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git timed out: git " + String.join(" ", args));
|
||||||
|
assertEquals(0, p.exitValue(), "git " + String.join(" ", args) + " failed:\n" + out);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static Path initRepo(Path dir) throws Exception {
|
||||||
|
Files.createDirectories(dir);
|
||||||
|
gitQuiet(dir, "init", "-q", "-b", "main");
|
||||||
|
gitQuiet(dir, "config", "user.email", "test@example.invalid");
|
||||||
|
gitQuiet(dir, "config", "user.name", "Test");
|
||||||
|
Files.writeString(dir.resolve("README.md"), "seed\n");
|
||||||
|
gitQuiet(dir, "add", "README.md");
|
||||||
|
gitQuiet(dir, "commit", "-q", "-m", "seed");
|
||||||
|
return dir;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #248's first measured mutation: replacing {@code CompletionResolver}'s 8th constructor
|
||||||
|
* argument with the inert {@code _ -> null} compiles clean and leaves every existing test green
|
||||||
|
* — it silently drops fleetd #241's fallback-report location. This drives the real assembled
|
||||||
|
* resolver through a member echoing its own injected brief back (no {@code fleet_reply}), which
|
||||||
|
* resolves via {@code noReportMessage(target)}, and proves the real member's {@code branch} —
|
||||||
|
* only obtainable via {@code Fleetd.worktreeBranchLookup(sessions::roster)} reading the real,
|
||||||
|
* worktree-provisioned {@link MemberSession} — appears in the reported text.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void assembledResolverReportsTheMembersWorktreeAndBranchInAFallbackReport(@TempDir Path dir) throws Exception {
|
||||||
|
Path repo = initRepo(dir.resolve("repo"));
|
||||||
|
FleetConfig cfg = writeConfig(dir, """
|
||||||
|
wtprofile:
|
||||||
|
baseUrl: http://wthost.local:8000
|
||||||
|
model: sonnet
|
||||||
|
""", "wthost.local", "worktreeRoot: " + dir.resolve("wts"));
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
ControllableResourcePorts ports = new ControllableResourcePorts(new FakeHerdr());
|
||||||
|
|
||||||
|
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
try {
|
||||||
|
MemberSession session = runtime.sessions().acquire("wtprofile", repo.toString(), repo.toString(),
|
||||||
|
null, new WorktreeRequest("fleetd-612-b1", null));
|
||||||
|
String target = session.terminalId();
|
||||||
|
String branch = session.branch();
|
||||||
|
assertTrue(branch != null && branch.startsWith("worker/"),
|
||||||
|
"sanity: a worktree-provisioned session must carry a real branch, got: " + branch);
|
||||||
|
|
||||||
|
CompletionResolver completion = runtime.completion();
|
||||||
|
CompletableFuture<Rendezvous.Resolution> waiter = new CompletableFuture<>();
|
||||||
|
String echoedBrief = "z".repeat(450); // >= CompletionResolver.ECHO_MIN_CHARS normalised chars
|
||||||
|
|
||||||
|
ports.herdr.readText("idle, nothing yet");
|
||||||
|
completion.onDelivered(target, new TurnToken(target, waiter, echoedBrief));
|
||||||
|
|
||||||
|
ports.herdr.readText(echoedBrief); // the pane just echoes the injected brief back — no real report
|
||||||
|
ports.advanceSeconds(3); // clear CompletionResolver.MIN_TURN_NANOS (2s) without a real sleep
|
||||||
|
completion.resolveBeforePostAction(target);
|
||||||
|
|
||||||
|
Rendezvous.Resolution resolution = waiter.getNow(null);
|
||||||
|
assertTrue(resolution != null, "the waiter must have resolved synchronously");
|
||||||
|
assertEquals(Rendezvous.Kind.COMPLETION, resolution.kind());
|
||||||
|
assertTrue(resolution.text().contains(CompletionResolver.NO_REPORT_PREFIX),
|
||||||
|
"sanity: must have gone down the noReportMessage sub-path: " + resolution.text());
|
||||||
|
assertTrue(resolution.text().contains("branch=" + branch),
|
||||||
|
"the assembled resolver must report the member's real branch (fleetd #241 via "
|
||||||
|
+ "fleetd #248's worktreeBranchLookup wiring); got: " + resolution.text());
|
||||||
|
} finally {
|
||||||
|
if (ports.shutdownHook != null) ports.shutdownHook.run();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #248's second measured mutation, and fleetd#201 Unit 5's own gap: replacing {@code
|
||||||
|
* backendErrorPatterns}/{@code backendErrorSink} with {@code BackendErrorPatternLookup.legacy()}
|
||||||
|
* / {@code BackendErrorSink.none()} compiles clean and leaves every existing behavioural test
|
||||||
|
* green.
|
||||||
|
*
|
||||||
|
* <p>Classification proof: this test's profile configures {@code errorPattern: "credential
|
||||||
|
* outage"} — text the built-in {@code (?i)\bAPI Error\s*:} fallback ({@code legacy()}'s only
|
||||||
|
* behaviour) never matches. So a real {@code Fleetd.backendErrorPatternLookup(...)} wiring
|
||||||
|
* classifies the send as {@code FAILED}; {@code legacy()} would fall through to the plain
|
||||||
|
* completion path instead ({@code Kind.COMPLETION}).
|
||||||
|
*
|
||||||
|
* <p>Cool-off proof: two distinct targets on the same profile/credential each classified as a
|
||||||
|
* backend error inside the 60s window must cool the credential off ({@link
|
||||||
|
* dev.ltms.fleet.placement.BackendOutagePolicy}, fleetd#201 Unit 5) — observable two ways: (1)
|
||||||
|
* the real {@code Fleetd.backendErrorSink(...)} marks each session {@code BACKEND_ERROR} (only
|
||||||
|
* the real sink calls {@code sessions.onBackendError}; {@code BackendErrorSink.none()} never
|
||||||
|
* does), and (2) a third explicit-profile spawn attempt is refused with a {@link
|
||||||
|
* PlacementException} naming the cool-off — only reachable because the real sink's {@code
|
||||||
|
* outagePolicy.record(...)} call actually ran.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void assembledResolverClassifiesAndCoolsOffOnAConfiguredBackendErrorPattern(@TempDir Path dir) throws Exception {
|
||||||
|
FleetConfig cfg = writeConfig(dir, """
|
||||||
|
coolprofile:
|
||||||
|
baseUrl: http://coolhost.local:8000
|
||||||
|
model: sonnet
|
||||||
|
errorPattern: "credential outage"
|
||||||
|
""", "coolhost.local", null);
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
ControllableResourcePorts ports = new ControllableResourcePorts(new FakeHerdr());
|
||||||
|
|
||||||
|
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
try {
|
||||||
|
MemberSession session1 = runtime.sessions().acquire("coolprofile", null, dir.toString(), null);
|
||||||
|
MemberSession session2 = runtime.sessions().acquire("coolprofile", null, dir.toString(), null);
|
||||||
|
String target1 = session1.terminalId();
|
||||||
|
String target2 = session2.terminalId();
|
||||||
|
assertTrue(!target1.equals(target2), "sanity: the two spawns must be distinct targets");
|
||||||
|
|
||||||
|
CompletionResolver completion = runtime.completion();
|
||||||
|
|
||||||
|
CompletableFuture<Rendezvous.Resolution> waiter1 = new CompletableFuture<>();
|
||||||
|
ports.herdr.readText("idle 1");
|
||||||
|
completion.onDelivered(target1, new TurnToken(target1, waiter1, null));
|
||||||
|
ports.herdr.readText("credential outage: upstream 503");
|
||||||
|
ports.advanceSeconds(3);
|
||||||
|
completion.resolveBeforePostAction(target1);
|
||||||
|
Rendezvous.Resolution resolution1 = waiter1.getNow(null);
|
||||||
|
assertTrue(resolution1 != null, "target1's waiter must have resolved synchronously");
|
||||||
|
assertEquals(Rendezvous.Kind.FAILED, resolution1.kind(),
|
||||||
|
"a configured errorPattern the built-in fallback never matches must classify as "
|
||||||
|
+ "a backend error, not a plain completion; got: " + resolution1);
|
||||||
|
assertTrue(resolution1.text().contains("credential outage: upstream 503"), resolution1.text());
|
||||||
|
|
||||||
|
CompletableFuture<Rendezvous.Resolution> waiter2 = new CompletableFuture<>();
|
||||||
|
ports.herdr.readText("idle 2");
|
||||||
|
completion.onDelivered(target2, new TurnToken(target2, waiter2, null));
|
||||||
|
ports.herdr.readText("credential outage: upstream 503 again");
|
||||||
|
ports.advanceSeconds(3);
|
||||||
|
completion.resolveBeforePostAction(target2);
|
||||||
|
Rendezvous.Resolution resolution2 = waiter2.getNow(null);
|
||||||
|
assertTrue(resolution2 != null, "target2's waiter must have resolved synchronously");
|
||||||
|
assertEquals(Rendezvous.Kind.FAILED, resolution2.kind());
|
||||||
|
|
||||||
|
List<MemberSession> roster = runtime.sessions().roster();
|
||||||
|
assertTrue(roster.stream().anyMatch(s -> target1.equals(s.terminalId())
|
||||||
|
&& s.state() == MemberSession.State.BACKEND_ERROR),
|
||||||
|
"the real backendErrorSink must have transitioned target1 to BACKEND_ERROR: " + roster);
|
||||||
|
assertTrue(roster.stream().anyMatch(s -> target2.equals(s.terminalId())
|
||||||
|
&& s.state() == MemberSession.State.BACKEND_ERROR),
|
||||||
|
"the real backendErrorSink must have transitioned target2 to BACKEND_ERROR: " + roster);
|
||||||
|
|
||||||
|
PlacementException coolOff = assertThrows(PlacementException.class,
|
||||||
|
() -> runtime.sessions().acquire("coolprofile", null, dir.toString(), null),
|
||||||
|
"two distinct targets classified within the 60s window must cool the credential "
|
||||||
|
+ "off (BackendOutagePolicy), refusing a third explicit-profile spawn");
|
||||||
|
assertTrue(coolOff.getMessage().contains("cooling off"), coolOff.getMessage());
|
||||||
|
} finally {
|
||||||
|
if (ports.shutdownHook != null) ports.shutdownHook.run();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,89 +0,0 @@
|
|||||||
package dev.ltms.fleet;
|
|
||||||
|
|
||||||
import java.nio.file.Files;
|
|
||||||
import java.nio.file.Path;
|
|
||||||
import org.junit.jupiter.api.DisplayName;
|
|
||||||
import org.junit.jupiter.api.Test;
|
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
|
||||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
|
||||||
|
|
||||||
/**
|
|
||||||
* fleetd #248: this is the test that was actually missing. {@code Fleetd.main} builds its {@code
|
|
||||||
* CompletionResolver} from an 8-argument constructor, and the ticket's own measurement proved two
|
|
||||||
* ways to silently unwire it — both compiled with 0 errors and left every existing test green:
|
|
||||||
*
|
|
||||||
* <ul>
|
|
||||||
* <li>replacing the worktree/branch argument (the 8th) with {@code _ -> null} — drops
|
|
||||||
* fleetd#241's fallback-report location entirely;</li>
|
|
||||||
* <li>replacing {@code backendErrorPatterns, backendErrorSink} (5th/6th) with {@code
|
|
||||||
* BackendErrorPatternLookup.legacy(), BackendErrorSink.none()} — drops fleetd#201 Unit 5's
|
|
||||||
* backend-error classification and cool-off entirely.</li>
|
|
||||||
* </ul>
|
|
||||||
*
|
|
||||||
* <p>Neither mutation could be caught by any test that constructs its own {@code
|
|
||||||
* CompletionResolver} (every test before this one did exactly that) or by a test of {@link
|
|
||||||
* Fleetd#worktreeBranchLookup}, {@link Fleetd#backendErrorPatternLookup}, or {@link
|
|
||||||
* Fleetd#backendErrorSink} in isolation (see {@code FleetdWorktreeBranchLookupTest}, {@code
|
|
||||||
* FleetdBackendErrorPatternLookupTest}, {@code FleetdBackendErrorSinkTest}) — those prove the
|
|
||||||
* factories work, never that {@code main} still calls them. This class is a plain source-text
|
|
||||||
* assertion on {@code Fleetd.java} — crude, but honest about what it checks, and it turns red the
|
|
||||||
* instant the wiring is dropped, mirroring the same fallback shape {@link
|
|
||||||
* FleetdFleetAppConstructionTest} already uses for a different constructor argument.
|
|
||||||
*
|
|
||||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
|
||||||
* CompletionResolver} and never runs {@code main}.
|
|
||||||
*/
|
|
||||||
class FleetdCompletionResolverWiringTest {
|
|
||||||
|
|
||||||
private static String fleetdSource() throws Exception {
|
|
||||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
|
||||||
}
|
|
||||||
|
|
||||||
@Test
|
|
||||||
@DisplayName("[SOURCE TEXT] CompletionResolver's construction call still names backendErrorPatterns and backendErrorSink")
|
|
||||||
void backendErrorArgumentsAreStillNamedAtTheCallSite() throws Exception {
|
|
||||||
String source = fleetdSource();
|
|
||||||
assertTrue(source.contains(
|
|
||||||
"exhaustionSink, backendErrorPatterns, backendErrorSink, System::nanoTime,"),
|
|
||||||
"CompletionResolver's construction call must still pass backendErrorPatterns and "
|
|
||||||
+ "backendErrorSink as its 5th/6th arguments. Replacing them with "
|
|
||||||
+ "BackendErrorPatternLookup.legacy()/BackendErrorSink.none() (fleetd #248's measured "
|
|
||||||
+ "mutation) compiles with 0 errors and leaves every behavioural test green — this "
|
|
||||||
+ "source check is what must go red instead.");
|
|
||||||
}
|
|
||||||
|
|
||||||
@Test
|
|
||||||
@DisplayName("[SOURCE TEXT] CompletionResolver's construction call still passes worktreeBranchLookup(sessions::roster)")
|
|
||||||
void worktreeBranchLookupIsStillPassedAtTheCallSite() throws Exception {
|
|
||||||
String source = fleetdSource();
|
|
||||||
assertTrue(source.contains("worktreeBranchLookup(sessions::roster)"),
|
|
||||||
"CompletionResolver's construction call must still pass worktreeBranchLookup(sessions::roster) "
|
|
||||||
+ "as its 8th (last) argument. Replacing it with the inert `_ -> null` (fleetd #248's "
|
|
||||||
+ "other measured mutation) compiles with 0 errors and leaves every behavioural test "
|
|
||||||
+ "green — this source check is what must go red instead.");
|
|
||||||
assertFalse(source.contains("System::nanoTime,\n _ -> null"),
|
|
||||||
"the worktree/branch argument must never regress to the inert `_ -> null` literal");
|
|
||||||
}
|
|
||||||
|
|
||||||
@Test
|
|
||||||
@DisplayName("[SOURCE TEXT] backendErrorPatterns is assigned from the extracted backendErrorPatternLookup(...) factory")
|
|
||||||
void backendErrorPatternsComesFromTheFactory() throws Exception {
|
|
||||||
String source = fleetdSource();
|
|
||||||
assertTrue(source.contains(
|
|
||||||
"BackendErrorPatternLookup backendErrorPatterns = backendErrorPatternLookup(sessions::roster,"),
|
|
||||||
"backendErrorPatterns must be assigned from Fleetd.backendErrorPatternLookup(...), not an "
|
|
||||||
+ "inline lambda that a source check on the CompletionResolver call alone cannot see "
|
|
||||||
+ "through");
|
|
||||||
}
|
|
||||||
|
|
||||||
@Test
|
|
||||||
@DisplayName("[SOURCE TEXT] backendErrorSink is assigned from the extracted backendErrorSink(...) factory")
|
|
||||||
void backendErrorSinkComesFromTheFactory() throws Exception {
|
|
||||||
String source = fleetdSource();
|
|
||||||
assertTrue(source.contains(
|
|
||||||
"BackendErrorSink backendErrorSink = backendErrorSink(sessions, () -> config.get().profiles(),"),
|
|
||||||
"backendErrorSink must be assigned from Fleetd.backendErrorSink(...), not an inline lambda "
|
|
||||||
+ "that a source check on the CompletionResolver call alone cannot see through");
|
|
||||||
}
|
|
||||||
}
|
|
||||||
@@ -1,31 +0,0 @@
|
|||||||
package dev.ltms.fleet;
|
|
||||||
|
|
||||||
import java.nio.file.Files;
|
|
||||||
import java.nio.file.Path;
|
|
||||||
import org.junit.jupiter.api.Test;
|
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
|
||||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
|
||||||
|
|
||||||
/**
|
|
||||||
* CB-185: {@code ConnectionIdentity} must resolve a caller's pane on EITHER herdr daemon (a
|
|
||||||
* lead's MCP connection resolves against the lead daemon; a member's against the member daemon).
|
|
||||||
* Pinning {@code PaneLocator} to {@code memberHerdr} alone — the bug this guards against — leaves
|
|
||||||
* every lead's own connection unresolvable ({@code callerTerminal == null}) the moment
|
|
||||||
* {@code memberHerdrSocket} names a second daemon, which breaks {@code fleet_reply}/{@code
|
|
||||||
* fleet_ask} and {@code fleet_whoami} for a lead. A unit test on {@link
|
|
||||||
* dev.ltms.fleet.herdr.PaneLocator} alone (see {@code PaneLocatorTest}) proves the class CAN
|
|
||||||
* search two clients, but not that {@code Fleetd.main} actually wires it that way — hence this
|
|
||||||
* source-level assertion, the same technique {@code FleetdHerdrControlConstructionTest} uses.
|
|
||||||
*/
|
|
||||||
class FleetdConnectionIdentityConstructionTest {
|
|
||||||
@Test
|
|
||||||
void connectionIdentitySearchesBothDaemonsNotJustTheMemberOne() throws Exception {
|
|
||||||
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
|
||||||
assertFalse(source.contains("new PaneLocator(memberHerdr)"),
|
|
||||||
"PaneLocator must not be pinned to the member daemon alone — a lead's own "
|
|
||||||
+ "connection resolves against the LEAD daemon and would never be found");
|
|
||||||
assertTrue(source.contains("new PaneLocator(herdr, memberHerdr)"),
|
|
||||||
"PaneLocator must search the lead daemon first, then the member daemon");
|
|
||||||
}
|
|
||||||
}
|
|
||||||
@@ -0,0 +1,218 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.inject.CompletionResolver;
|
||||||
|
import dev.ltms.fleet.msg.Rendezvous;
|
||||||
|
import dev.ltms.fleet.msg.TurnToken;
|
||||||
|
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||||
|
import dev.ltms.fleet.session.MemberSession;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.OptionalLong;
|
||||||
|
import java.util.concurrent.CompletableFuture;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 step 4, ranks 1 and 2 (publish side) — {@link FleetdAssembly} lines
|
||||||
|
* {@code liveExhaustedPatterns}/{@code exhaustedPatterns} (CB-578 stage A, the ticket's own "worst
|
||||||
|
* consequence in the whole sweep": a genuine usage-limit refusal handed back to a waiting caller
|
||||||
|
* AS REAL COMPLETED WORK) and {@code Fleetd.publishExhaustionSink(...)} (CB-578 stage B: the
|
||||||
|
* credential that hit the limit is never quarantined). None of these three lines is driven by an
|
||||||
|
* existing test through the real assembly: {@code FleetdExhaustedPatternLookupWiringTest} and
|
||||||
|
* {@code FleetdLiveExhaustedPatternsWiringTest} (fleetd #589) call {@code Fleetd.liveExhaustedPatterns}
|
||||||
|
* / {@code Fleetd.exhaustedPatternLookup} directly as factories, never through {@link
|
||||||
|
* FleetdAssembly#assembleAndStart} — they prove the FACTORY classifies correctly, never that THIS
|
||||||
|
* call site is the one that actually got wired into the running {@link CompletionResolver}. {@link
|
||||||
|
* FleetdBackendQuarantineAssemblyTest} drives {@code BackendQuarantine.withEscalation(...)}
|
||||||
|
* directly, a different call site from {@code publishExhaustionSink} here.
|
||||||
|
*
|
||||||
|
* <p>This test drives the REAL assembled {@link CompletionResolver} ({@link
|
||||||
|
* FleetdRuntime#completion()}) with a profile carrying a configured {@code exhaustedPattern},
|
||||||
|
* through a pane scrape that matches it, and asserts both halves of the production consequence:
|
||||||
|
* (1) the resolution is {@link Rendezvous.Kind#BACKEND_EXHAUSTED}, never a plain completion handed
|
||||||
|
* back as real work, and (2) the profile's credential is actually quarantined afterward, through
|
||||||
|
* the REAL {@link BackendQuarantine} the same assembly built ({@link
|
||||||
|
* FleetdRuntime#mcp()}{@code .quarantineSource().quarantine()}) — never a copy.
|
||||||
|
*
|
||||||
|
* <p>Same {@code ControllableResourcePorts} shape as {@code FleetdCompletionResolverAssemblyTest}:
|
||||||
|
* a fake, advanceable {@code nanoClock} so {@code CompletionResolver.MIN_TURN_NANOS} clears without
|
||||||
|
* a real sleep, and {@link FakeHerdr#readText} to drive the pane scrape.
|
||||||
|
*/
|
||||||
|
class FleetdExhaustedPatternAssemblyTest {
|
||||||
|
|
||||||
|
private static final class ControllableResourcePorts implements ResourcePorts {
|
||||||
|
|
||||||
|
final FakeHerdr herdr;
|
||||||
|
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L); // arbitrary non-zero start
|
||||||
|
Runnable shutdownHook;
|
||||||
|
|
||||||
|
ControllableResourcePorts(FakeHerdr herdr) {
|
||||||
|
this.herdr = herdr;
|
||||||
|
}
|
||||||
|
|
||||||
|
void advanceSeconds(long seconds) {
|
||||||
|
nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(seconds));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
return herdr;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
return (uri, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException("replyInboxOpener must not be called — no broker: block");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
return (uri, selfCoordId, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException("leadMailboxOpener must not be called — no coordinator: block");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||||
|
return () -> {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
return nowNanos::get;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return nowNanos::get;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
return Executors.newSingleThreadScheduledExecutor();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
this.shutdownHook = hook;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
// Deliberately never bind a real port.
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir, int cooldownSeconds) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
idleSleepGuard:
|
||||||
|
enabled: false
|
||||||
|
quarantineCooldownSeconds: %d
|
||||||
|
profiles:
|
||||||
|
exhaustprofile:
|
||||||
|
baseUrl: http://exhausthost.local:8000
|
||||||
|
model: sonnet
|
||||||
|
exhaustedPattern: "usage limit reached"
|
||||||
|
guard:
|
||||||
|
offSubscriptionHosts:
|
||||||
|
- exhausthost.local
|
||||||
|
""".formatted(cooldownSeconds));
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #589's own description of this gap ({@code Fleetd#exhaustedPatternLookup}'s javadoc):
|
||||||
|
* "the worst consequence in the whole #589 sweep" — a genuine usage-limit refusal stops being
|
||||||
|
* classified as {@code BACKEND_EXHAUSTED} and is handed back to a waiting {@code fleet_send} as
|
||||||
|
* if it were real completed work. Pins {@code FleetdAssembly}'s {@code liveExhaustedPatterns}
|
||||||
|
* AND {@code exhaustedPatterns} lines (rank 1) together with {@code publishExhaustionSink}
|
||||||
|
* (rank 2, the non-OpenCode half) in one flow: classify, then quarantine.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
@DisplayName("[BEHAVIOURAL] a scrape matching the profile's exhaustedPattern resolves "
|
||||||
|
+ "BACKEND_EXHAUSTED (never a plain completion) and quarantines the credential")
|
||||||
|
void assembledResolverClassifiesExhaustionAndQuarantinesTheCredential(@TempDir Path dir) throws Exception {
|
||||||
|
int cooldownSeconds = 120;
|
||||||
|
FleetConfig cfg = writeConfig(dir, cooldownSeconds);
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
ControllableResourcePorts ports = new ControllableResourcePorts(new FakeHerdr());
|
||||||
|
|
||||||
|
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
try {
|
||||||
|
MemberSession session = runtime.sessions().acquire("exhaustprofile", null, dir.toString(), null);
|
||||||
|
String target = session.terminalId();
|
||||||
|
|
||||||
|
CompletionResolver completion = runtime.completion();
|
||||||
|
CompletableFuture<Rendezvous.Resolution> waiter = new CompletableFuture<>();
|
||||||
|
|
||||||
|
ports.herdr.readText("idle, nothing yet");
|
||||||
|
completion.onDelivered(target, new TurnToken(target, waiter, null));
|
||||||
|
// The matched text must START the pane line (CompletionResolver.startsWithExhaustion) —
|
||||||
|
// no preceding sentence — for the quarantine side-effect to fire, same as production.
|
||||||
|
ports.herdr.readText("usage limit reached: try again in a few hours");
|
||||||
|
ports.advanceSeconds(3); // clear CompletionResolver.MIN_TURN_NANOS (2s), no real sleep
|
||||||
|
completion.resolveBeforePostAction(target);
|
||||||
|
|
||||||
|
// CONTROL: the waiter must have resolved synchronously at all — if the assembled
|
||||||
|
// CompletionResolver were never actually driven (e.g. a wiring break upstream silently
|
||||||
|
// left the resolver unreachable), this fails loudly before the real assertions below
|
||||||
|
// ever run, rather than passing on an untouched waiter.
|
||||||
|
Rendezvous.Resolution resolution = waiter.getNow(null);
|
||||||
|
assertTrue(resolution != null, "CONTROL: the waiter must have resolved synchronously — "
|
||||||
|
+ "if this is null, the assembled resolver was never actually exercised");
|
||||||
|
|
||||||
|
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, resolution.kind(),
|
||||||
|
"a scrape matching the profile's configured exhaustedPattern must classify as "
|
||||||
|
+ "BACKEND_EXHAUSTED, not a plain completion handed back as real work — "
|
||||||
|
+ "replacing FleetdAssembly's liveExhaustedPatterns/exhaustedPatterns "
|
||||||
|
+ "lines with their inert forms (Map.of() / target -> null) must fail "
|
||||||
|
+ "this assertion; got: " + resolution);
|
||||||
|
assertTrue(resolution.text().contains("usage limit reached"), resolution.text());
|
||||||
|
|
||||||
|
BackendQuarantine quarantine = runtime.mcp().quarantineSource().quarantine();
|
||||||
|
assertTrue(quarantine.isQuarantined("exhaustprofile"),
|
||||||
|
"the real publishExhaustionSink-built sink must have quarantined the profile's "
|
||||||
|
+ "credential (effectiveCredentialId() == the profile name here, no "
|
||||||
|
+ "credentialId configured) — replacing FleetdAssembly's "
|
||||||
|
+ "publishExhaustionSink call site with a hardcoded ExhaustionSink.none() "
|
||||||
|
+ "must fail this assertion, since nothing would ever call "
|
||||||
|
+ "quarantine.quarantine(...)");
|
||||||
|
OptionalLong remaining = quarantine.remainingSeconds("exhaustprofile");
|
||||||
|
assertTrue(remaining.isPresent() && remaining.getAsLong() > 0
|
||||||
|
&& remaining.getAsLong() <= cooldownSeconds,
|
||||||
|
"a fresh quarantine must block for at most the configured base cooldown: " + remaining);
|
||||||
|
} finally {
|
||||||
|
if (ports.shutdownHook != null) ports.shutdownHook.run();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,86 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
|
||||||
|
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
|
||||||
|
import dev.ltms.fleet.peer.MemberRole;
|
||||||
|
import dev.ltms.fleet.session.MemberSession;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.regex.Pattern;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #589 Group 1: {@link Fleetd#exhaustedPatternLookup} is the factory that replaced {@code
|
||||||
|
* main}'s inline lambda — resolve a herdr {@code target} to its session's profile, then to that
|
||||||
|
* profile's live {@link LiveExhaustedPatterns#patternFor}. Same shape as {@link
|
||||||
|
* Fleetd#worktreeBranchLookup} (which {@code FleetdWorktreeBranchLookupTest} pins the same way).
|
||||||
|
*
|
||||||
|
* <p>Before this ticket the lambda was built inline in {@code main} and untestable: replacing it
|
||||||
|
* with {@code target -> null} — the exact shape of {@link ExhaustedPatternLookup#none()} — compiled
|
||||||
|
* with 0 errors and left every existing test green. Per the ticket, this is the worst consequence
|
||||||
|
* in the whole #589 sweep: a genuine usage-limit refusal would stop being classified as {@code
|
||||||
|
* BACKEND_EXHAUSTED} and would be handed back to a waiting {@code fleet_send} as if it were real
|
||||||
|
* completed work.
|
||||||
|
*/
|
||||||
|
class FleetdExhaustedPatternLookupWiringTest {
|
||||||
|
|
||||||
|
private static final String YAML = """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
herdrSocket: ~/.config/herdr/herdr.sock
|
||||||
|
profiles:
|
||||||
|
terra:
|
||||||
|
baseUrl: http://gx00.gw:8000
|
||||||
|
model: claude-opus-5
|
||||||
|
exhaustedPattern: "usage limit"
|
||||||
|
""";
|
||||||
|
|
||||||
|
private static MemberSession session(String terminal, String profile) {
|
||||||
|
return new MemberSession("pane-" + terminal, terminal, profile, MemberRole.DEV,
|
||||||
|
"/cwd", null, 0L, 0L, 0, MemberSession.State.READY, null, null);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static LiveExhaustedPatterns liveExhaustedPatterns(Path dir) throws Exception {
|
||||||
|
Path file = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(file, YAML);
|
||||||
|
FleetConfig cfg = FleetConfig.load(file);
|
||||||
|
return new LiveExhaustedPatterns(() -> ConfigRef.fixed(cfg).get().profiles());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("a known target resolves through its session's profile to that profile's live pattern")
|
||||||
|
void knownTargetResolvesThroughItsProfile(@TempDir Path dir) throws Exception {
|
||||||
|
LiveExhaustedPatterns patterns = liveExhaustedPatterns(dir);
|
||||||
|
ExhaustedPatternLookup lookup = Fleetd.exhaustedPatternLookup(
|
||||||
|
() -> List.of(session("term1", "terra")), patterns);
|
||||||
|
|
||||||
|
Pattern resolved = lookup.patternFor("term1");
|
||||||
|
|
||||||
|
assertNotNull(resolved,
|
||||||
|
"the lookup must resolve term1 -> profile 'terra' -> LiveExhaustedPatterns.patternFor("
|
||||||
|
+ "'terra') — replacing the lambda body with 'target -> null' at the "
|
||||||
|
+ "Fleetd.exhaustedPatternLookup call site must fail this assertion");
|
||||||
|
assertTrue(resolved.matcher("the usage limit has been reached").find());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("an unknown target resolves to null, not a thrown exception")
|
||||||
|
void unknownTargetResolvesToNull(@TempDir Path dir) throws Exception {
|
||||||
|
LiveExhaustedPatterns patterns = liveExhaustedPatterns(dir);
|
||||||
|
ExhaustedPatternLookup lookup = Fleetd.exhaustedPatternLookup(
|
||||||
|
() -> List.of(session("term1", "terra")), patterns);
|
||||||
|
|
||||||
|
assertNull(lookup.patternFor("term_stranger"));
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,62 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
|
||||||
|
import java.util.concurrent.atomic.AtomicBoolean;
|
||||||
|
import java.util.concurrent.atomic.AtomicReference;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #589 Group 1: {@link Fleetd#forwardingExhaustionSink} is the factory that replaced
|
||||||
|
* {@code main}'s inline {@code ExhaustionSink.forwardingTo(exhaustionSinkRef::get)} (fleetd #175's
|
||||||
|
* construction-order break: the adapters need a sink before {@code sessions} exists to build the
|
||||||
|
* real one). Before this ticket that call site was untestable wiring: replacing the supplier
|
||||||
|
* argument with a hardcoded {@code () -> ExhaustionSink.none()} compiled with 0 errors and left
|
||||||
|
* every existing test green, because no test builds the object {@code main} actually wires and
|
||||||
|
* then mutates the reference afterward — every existing {@code ExhaustionSink.forwardingTo} caller
|
||||||
|
* in this codebase reads and writes the SAME reference within one test, so a hardcoded-none supplier
|
||||||
|
* and a correctly-forwarding one are indistinguishable to them.
|
||||||
|
*
|
||||||
|
* <p>This test builds the reference, builds the forwarder from it, and only THEN repoints the
|
||||||
|
* reference at a spy sink — the discriminating order fleetd #175's whole design depends on
|
||||||
|
* ({@code exhaustionSinkRef} starts at {@code none()} and is repointed once {@code sessions}
|
||||||
|
* exists). A forwarder that captured a fixed target at construction time (the inert form) can never
|
||||||
|
* see that later repoint.
|
||||||
|
*/
|
||||||
|
class FleetdExhaustionSinkForwardingWiringTest {
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("the forwarder reads the reference live: repointing it AFTER construction is honoured")
|
||||||
|
void forwarderReadsTheReferenceLiveNotAFixedTargetCapturedAtConstruction() {
|
||||||
|
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
|
||||||
|
ExhaustionSink forwarder = Fleetd.forwardingExhaustionSink(exhaustionSinkRef);
|
||||||
|
|
||||||
|
AtomicBoolean spyCalled = new AtomicBoolean(false);
|
||||||
|
exhaustionSinkRef.set((target, reason, profile) -> spyCalled.set(true));
|
||||||
|
|
||||||
|
forwarder.onExhausted("term_x", "usage limit reached", "terra");
|
||||||
|
|
||||||
|
assertTrue(spyCalled.get(),
|
||||||
|
"forwardingExhaustionSink must delegate to whatever exhaustionSinkRef currently "
|
||||||
|
+ "holds — hardcoding the supplier to () -> ExhaustionSink.none() at the "
|
||||||
|
+ "Fleetd.forwardingExhaustionSink call site must fail this assertion, "
|
||||||
|
+ "since the spy set into the reference after construction would never run");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("before any repoint, the forwarder is inert — it starts at none(), not a crash")
|
||||||
|
void beforeAnyRepointTheForwarderIsInert() {
|
||||||
|
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
|
||||||
|
ExhaustionSink forwarder = Fleetd.forwardingExhaustionSink(exhaustionSinkRef);
|
||||||
|
|
||||||
|
AtomicBoolean spyCalled = new AtomicBoolean(false);
|
||||||
|
forwarder.onExhausted("term_x", "usage limit reached", "terra");
|
||||||
|
|
||||||
|
assertFalse(spyCalled.get(), "nothing was ever wired to be called here — this only pins "
|
||||||
|
+ "that the factory does not throw before a real sink is published");
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,110 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.AgentControl;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||||
|
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||||
|
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||||
|
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||||
|
import dev.ltms.fleet.session.SessionManager;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.HashMap;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.Set;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.atomic.AtomicReference;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #589 Group 1: {@link Fleetd#publishExhaustionSink} is the factory that replaced {@code
|
||||||
|
* main}'s previously untested two-statement sequence — build the real {@link
|
||||||
|
* Fleetd#exhaustionSink}, then {@code exhaustionSinkRef.set(exhaustionSink)}. {@link
|
||||||
|
* Fleetd#exhaustionSink} itself is already pinned by {@code FleetdExhaustionSinkWarningTest} (its
|
||||||
|
* log text) — what was NEVER pinned is the {@code .set(...)} call: {@code main} could replace it
|
||||||
|
* with {@code exhaustionSinkRef.set(ExhaustionSink.none())} and compile with 0 errors, leaving
|
||||||
|
* every existing test green, because {@link Fleetd#exhaustionSink}'s own tests build and call the
|
||||||
|
* sink directly, never through the reference {@code main} publishes it into.
|
||||||
|
*
|
||||||
|
* <p>This test proves the PUBLISHED reference — not a freshly rebuilt sink — is the one that
|
||||||
|
* actually quarantines a credential, by reading {@link BackendQuarantine#isQuarantined} after
|
||||||
|
* calling {@code exhaustionSinkRef.get().onExhausted(...)}, the same object {@link
|
||||||
|
* Fleetd#forwardingExhaustionSink} forwards to in production.
|
||||||
|
*/
|
||||||
|
class FleetdExhaustionSinkPublishWiringTest {
|
||||||
|
|
||||||
|
private static final String YAML = """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
herdrSocket: ~/.config/herdr/herdr.sock
|
||||||
|
profiles:
|
||||||
|
terra:
|
||||||
|
baseUrl: http://gx00.gw:8000
|
||||||
|
model: claude-opus-5
|
||||||
|
guard:
|
||||||
|
offSubscriptionHosts:
|
||||||
|
- gx00.gw
|
||||||
|
""";
|
||||||
|
|
||||||
|
private static SessionManager emptyRosterSessions() {
|
||||||
|
FakeHerdr h = new FakeHerdr();
|
||||||
|
FleetConfig.Profile dummy = new FleetConfig.Profile(
|
||||||
|
"dummy", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
|
||||||
|
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||||
|
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
|
||||||
|
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(dummy.profile(), dummy), dummy.profile(), _ -> "tok");
|
||||||
|
// Never acquires a session — publishExhaustionSink's built sink resolves target -> profile
|
||||||
|
// via the profileHint fallback (fleetd #234), exactly like OpenCodeLauncher's real call
|
||||||
|
// site does, so this never needs a populated roster.
|
||||||
|
return new SessionManager(launcher);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("the published reference actually quarantines — not a rebuilt-but-never-set sink")
|
||||||
|
void publishedReferenceActuallyQuarantines(@TempDir Path dir) throws Exception {
|
||||||
|
Path file = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(file, YAML);
|
||||||
|
FleetConfig cfg = FleetConfig.load(file);
|
||||||
|
ConfigRef config = ConfigRef.fixed(cfg);
|
||||||
|
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||||
|
Map<String, String> reasonByCredential = new HashMap<>();
|
||||||
|
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
|
||||||
|
|
||||||
|
Fleetd.publishExhaustionSink(exhaustionSinkRef, emptyRosterSessions(), config, quarantine,
|
||||||
|
reasonByCredential, cfg);
|
||||||
|
exhaustionSinkRef.get().onExhausted("term_x", "The usage limit has been reached", "terra");
|
||||||
|
|
||||||
|
assertTrue(quarantine.isQuarantined("terra"),
|
||||||
|
"publishExhaustionSink must repoint exhaustionSinkRef at the REAL sink — "
|
||||||
|
+ "replacing the .set(...) call with exhaustionSinkRef.set(ExhaustionSink.none()) "
|
||||||
|
+ "at the Fleetd.publishExhaustionSink call site must fail this assertion, "
|
||||||
|
+ "since none()'s onExhausted does nothing");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("before publishing, the reference is still inert — no quarantine, no crash")
|
||||||
|
void beforePublishingTheReferenceIsStillInert(@TempDir Path dir) throws Exception {
|
||||||
|
Path file = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(file, YAML);
|
||||||
|
FleetConfig cfg = FleetConfig.load(file);
|
||||||
|
ConfigRef config = ConfigRef.fixed(cfg);
|
||||||
|
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||||
|
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
|
||||||
|
|
||||||
|
exhaustionSinkRef.get().onExhausted("term_x", "The usage limit has been reached", "terra");
|
||||||
|
|
||||||
|
assertFalse(quarantine.isQuarantined("terra"),
|
||||||
|
"nothing was published yet — this only pins the starting state the other test's "
|
||||||
|
+ "assertion actually distinguishes from");
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,29 +0,0 @@
|
|||||||
package dev.ltms.fleet;
|
|
||||||
|
|
||||||
import java.nio.file.Files;
|
|
||||||
import java.nio.file.Path;
|
|
||||||
import org.junit.jupiter.api.Test;
|
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
|
||||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
|
||||||
|
|
||||||
/**
|
|
||||||
* CB-185: {@code FleetApp} must be constructed with BOTH herdr clients (the lead's and the
|
|
||||||
* member's), never the raw lead-only {@code herdr}. Passing only {@code herdr} — the bug this
|
|
||||||
* guards against — makes {@code GET /healthz} green while the member daemon is down (so every
|
|
||||||
* spawn fails invisibly) and silently drops every member workspace from {@code GET /sessions}.
|
|
||||||
* A behavioural test on {@code FleetApp} alone (see {@code FleetAppTwoDaemonTest}) proves the
|
|
||||||
* class merges/gates correctly when given two clients, but not that {@code Fleetd.main} actually
|
|
||||||
* passes it two — hence this source-level assertion, mirroring
|
|
||||||
* {@code FleetdHerdrControlConstructionTest}.
|
|
||||||
*/
|
|
||||||
class FleetdFleetAppConstructionTest {
|
|
||||||
@Test
|
|
||||||
void fleetAppIsConstructedWithBothHerdrDaemons() throws Exception {
|
|
||||||
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
|
||||||
assertFalse(source.contains("new FleetApp(herdr, workers,"),
|
|
||||||
"FleetApp must not be constructed with the lead-only herdr client");
|
|
||||||
assertTrue(source.contains("new FleetApp(herdr, memberHerdr, workers,"),
|
|
||||||
"FleetApp must be constructed with both the lead and the member herdr client");
|
|
||||||
}
|
|
||||||
}
|
|
||||||
@@ -0,0 +1,100 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.function.Function;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #602 gauge-wiring: {@link Fleetd#leadConfigDirLookup} is the factory {@code Fleetd.main}
|
||||||
|
* wires into {@code FleetMcp.LeadConfigDirSource} so {@code fleet_list}'s {@code context} row reads
|
||||||
|
* the transcript directory a lead's OWN profile actually writes to, instead of always falling back
|
||||||
|
* to the built-in {@code <user.home>/.claude} default (see {@code LeadContextGauge}).
|
||||||
|
*
|
||||||
|
* <p>{@code FleetMcpLeadContextGaugeWiringTest} proves the directory this factory returns is what
|
||||||
|
* actually gets read; this class proves the factory's own matching logic — the same
|
||||||
|
* {@code fleet.leaders.<name>.profile} link {@code Fleetd.leadSeatLookup} already follows (see
|
||||||
|
* {@code FleetdLeadSeatLookupTest}), one step further to that profile's own {@code configDir:}.
|
||||||
|
*/
|
||||||
|
class FleetdLeadConfigDirLookupTest {
|
||||||
|
|
||||||
|
private static FleetConfig.Profile profileWithConfigDir(String name, String configDir) {
|
||||||
|
return new FleetConfig.Profile(name, null, "claude-sonnet-5", configDir, null, null,
|
||||||
|
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
|
||||||
|
null, 3, true, null, null, null);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig.Leader leadOnProfile(String profile) {
|
||||||
|
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("a lead on a profile that sets configDir resolves to that directory")
|
||||||
|
void leadOnAProfileWithConfigDirResolvesToIt() {
|
||||||
|
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", "/mnt/opus-claude"));
|
||||||
|
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||||
|
Function<String, String> lookup = Fleetd.leadConfigDirLookup(() -> profiles, leaders);
|
||||||
|
|
||||||
|
assertEquals("/mnt/opus-claude", lookup.apply("primary"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("changing the config from directory A to directory B changes what the lookup reports")
|
||||||
|
void configChangeFromDirectoryAToDirectoryBChangesTheAnswer() {
|
||||||
|
java.util.concurrent.atomic.AtomicReference<Map<String, FleetConfig.Profile>> profilesRef =
|
||||||
|
new java.util.concurrent.atomic.AtomicReference<>(
|
||||||
|
Map.of("opus", profileWithConfigDir("opus", "/mnt/dir-a")));
|
||||||
|
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||||
|
Function<String, String> lookup = Fleetd.leadConfigDirLookup(profilesRef::get, leaders);
|
||||||
|
|
||||||
|
assertEquals("/mnt/dir-a", lookup.apply("primary"), "must read directory A before the config changes");
|
||||||
|
|
||||||
|
profilesRef.set(Map.of("opus", profileWithConfigDir("opus", "/mnt/dir-b")));
|
||||||
|
assertEquals("/mnt/dir-b", lookup.apply("primary"), "must read directory B once the LIVE config changes — "
|
||||||
|
+ "a lookup that snapshotted the profile map at construction would still answer directory A here");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("a lead entry with no `profile:` (recognise-only) resolves to null, not a thrown exception")
|
||||||
|
void recogniseOnlyLeadWithNoProfileResolvesToNull() {
|
||||||
|
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", "/mnt/opus-claude"));
|
||||||
|
FleetConfig.Leader recogniseOnly = new FleetConfig.Leader(null, "lead: primary", 1, "lead:", 10,
|
||||||
|
"claude", "claude-sonnet-5");
|
||||||
|
Map<String, FleetConfig.Leader> leaders = Map.of("primary", recogniseOnly);
|
||||||
|
Function<String, String> lookup = Fleetd.leadConfigDirLookup(() -> profiles, leaders);
|
||||||
|
|
||||||
|
assertNull(lookup.apply("primary"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("a lead naming a profile that is not configured resolves to null, not a thrown exception")
|
||||||
|
void leadOnAnUnconfiguredProfileResolvesToNull() {
|
||||||
|
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("ghost-profile"));
|
||||||
|
Function<String, String> lookup = Fleetd.leadConfigDirLookup(Map::of, leaders);
|
||||||
|
|
||||||
|
assertNull(lookup.apply("primary"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("a lead on a profile that sets no configDir override resolves to null")
|
||||||
|
void leadOnAProfileWithNoConfigDirResolvesToNull() {
|
||||||
|
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", null));
|
||||||
|
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||||
|
Function<String, String> lookup = Fleetd.leadConfigDirLookup(() -> profiles, leaders);
|
||||||
|
|
||||||
|
assertNull(lookup.apply("primary"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("an unrecognised lead name resolves to null, not a thrown exception")
|
||||||
|
void unrecognisedLeadNameResolvesToNull() {
|
||||||
|
Function<String, String> lookup = Fleetd.leadConfigDirLookup(Map::of, Map.of());
|
||||||
|
|
||||||
|
assertNull(lookup.apply("ghost-lead"));
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,98 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.mcp.FleetMcp;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
|
||||||
|
import java.util.Map;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #602 gauge-wiring follow-up (PR #606 review comment 17353): {@code Fleetd.main}'s {@code
|
||||||
|
* LeadConfigDirSource} local used to be a bare {@code new FleetMcp.LeadConfigDirSource(
|
||||||
|
* leadConfigDirLookup(...))} built inline, with nothing a test could call directly. Measured on
|
||||||
|
* that shape: replacing the whole expression with {@code FleetMcp.LeadConfigDirSource.none()} at
|
||||||
|
* the call site compiled with 0 errors and left the full 1822-test suite green — the daemon could
|
||||||
|
* be changed to always report every lead's context as {@code UNKNOWN}, forever, and no test would
|
||||||
|
* notice. That is the same hand-built-vs-config-wired shape as fleetd #561/#248/#426/#562
|
||||||
|
* ({@code FleetdLoopHealthSourceWiringTest}).
|
||||||
|
*
|
||||||
|
* <p>{@code FleetMcpLeadContextGaugeWiringTest} and {@code FleetdLeadConfigDirLookupTest} both
|
||||||
|
* predate this class and are both still correct — but neither can catch the mutation above. One
|
||||||
|
* builds its own {@code FleetMcp} and hands it its own {@code LeadConfigDirSource}; the other
|
||||||
|
* builds its own lookup and calls {@link Fleetd#leadConfigDirLookup} directly. Neither one ever
|
||||||
|
* calls the thing {@code Fleetd.main} actually calls.
|
||||||
|
*
|
||||||
|
* <p>The fix extracts the inline {@code new} into {@link Fleetd#leadConfigDirSource}, a
|
||||||
|
* package-private factory in the same style as {@link Fleetd#loopHealthSource}/{@link
|
||||||
|
* Fleetd#capacitySource}/{@link Fleetd#healthCoverageSource} — which is exactly what makes it
|
||||||
|
* directly callable here. This test calls that factory with real {@link FleetConfig.Profile}/
|
||||||
|
* {@link FleetConfig.Leader} fixtures (the same shapes {@code FleetdLeadConfigDirLookupTest}
|
||||||
|
* already uses) and asserts the returned source resolves a real {@code configDir} — a property
|
||||||
|
* that would be false if {@link Fleetd#leadConfigDirSource} were mutated to {@code return
|
||||||
|
* FleetMcp.LeadConfigDirSource.none();}. Measured: mutating exactly that line makes
|
||||||
|
* {@link #resolvesTheRealConfiguredConfigDir()} fail ({@code expected: </mnt/opus-claude> but was:
|
||||||
|
* <null>}); restoring it makes the whole suite green again.
|
||||||
|
*
|
||||||
|
* <p><b>What this class does not and cannot cover.</b> {@code main}'s own line —
|
||||||
|
* {@code leadConfigDirSource(() -> config.get().profiles(), leaders)} — could itself be swapped
|
||||||
|
* for a bare {@code FleetMcp.LeadConfigDirSource.none()}, bypassing this factory entirely. Measured:
|
||||||
|
* that exact mutation compiles with 0 errors and leaves every test in this file, and the full
|
||||||
|
* 1825-test suite, green. {@link Fleetd#loopHealthSource}'s own wiring test has the identical gap
|
||||||
|
* for its own one-line call in {@code main} — no test in this codebase calls {@code Fleetd.main}
|
||||||
|
* far enough to observe which factory call it made. This class narrows the gap from "nothing tests
|
||||||
|
* the wiring" (the pre-extraction state this ticket found) to "the factory's own logic is pinned,
|
||||||
|
* and main's call to it is a one-line, visually-verifiable delegation" — the same standard already
|
||||||
|
* accepted for {@code loopHealthSource}/{@code capacitySource}/{@code healthCoverageSource}.
|
||||||
|
*/
|
||||||
|
class FleetdLeadConfigDirSourceWiringTest {
|
||||||
|
|
||||||
|
private static FleetConfig.Profile profileWithConfigDir(String name, String configDir) {
|
||||||
|
return new FleetConfig.Profile(name, null, "claude-sonnet-5", configDir, null, null,
|
||||||
|
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
|
||||||
|
null, 3, true, null, null, null);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig.Leader leadOnProfile(String profile) {
|
||||||
|
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("the returned source resolves the lead's REAL configured configDir, not a hardcoded null")
|
||||||
|
void resolvesTheRealConfiguredConfigDir() {
|
||||||
|
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", "/mnt/opus-claude"));
|
||||||
|
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||||
|
|
||||||
|
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
|
||||||
|
|
||||||
|
assertEquals("/mnt/opus-claude", source.configDirFor().apply("primary"),
|
||||||
|
"the configDirFor function must delegate to the real leadConfigDirLookup — mutating "
|
||||||
|
+ "Fleetd.leadConfigDirSource's own body to `return FleetMcp.LeadConfigDirSource.none();` "
|
||||||
|
+ "must fail this assertion (measured: it does — see this test's class javadoc for the "
|
||||||
|
+ "companion measurement on main's one-line call to this factory, which this assertion "
|
||||||
|
+ "does not and structurally cannot cover)");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("a lead on a profile with no configDir override still resolves to null, not a crash")
|
||||||
|
void leadWithNoConfigDirOverrideResolvesToNull() {
|
||||||
|
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithConfigDir("opus", null));
|
||||||
|
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||||
|
|
||||||
|
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
|
||||||
|
|
||||||
|
assertNull(source.configDirFor().apply("primary"),
|
||||||
|
"no configDir: override configured ⇒ null, so LeadContextGauge falls back to its own default");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("an unrecognised lead name resolves to null, not a thrown exception")
|
||||||
|
void unrecognisedLeadNameResolvesToNull() {
|
||||||
|
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(Map::of, Map.of());
|
||||||
|
|
||||||
|
assertNull(source.configDirFor().apply("ghost-lead"));
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,156 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import com.fasterxml.jackson.databind.JsonNode;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
import dev.ltms.fleet.herdr.AgentControl;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrException;
|
||||||
|
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.io.IOException;
|
||||||
|
import java.nio.charset.StandardCharsets;
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.function.Function;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #609: {@link Fleetd#leadContextLookup} is the factory {@code Fleetd.main} wires into
|
||||||
|
* {@code LeadHeartbeatLoop.LeadContextSource} so the idle-lead heartbeat can read a lead's own
|
||||||
|
* {@link LeadContextGauge} reading by its terminal id. Three hops: terminal → lead name (unknown ⇒
|
||||||
|
* UNKNOWN), lead name → configDir, and {@code agents.get(terminal)} for the live session id and
|
||||||
|
* agent type the gauge itself needs — a herdr failure on that last hop must degrade to UNKNOWN, not
|
||||||
|
* throw and kill the heartbeat's own tick.
|
||||||
|
*
|
||||||
|
* <p>{@code AgentControl.agentCall} resolves a {@code term_}-prefixed target's pane id via a first
|
||||||
|
* {@code agent.list} round trip (see its own javadoc); this class's terminal id deliberately does
|
||||||
|
* NOT start with {@code term_} so the stub {@link HerdrClient} below only needs to answer
|
||||||
|
* {@code agent.get} — the one call this factory actually depends on.
|
||||||
|
*/
|
||||||
|
class FleetdLeadContextLookupTest {
|
||||||
|
|
||||||
|
private static final String LEAD_TERMINAL = "leadpane1";
|
||||||
|
private static final String LEAD_NAME = "opus";
|
||||||
|
private static final String SESSION_ID = "sess-609-happy-path";
|
||||||
|
|
||||||
|
private static String usageLine(long tokens) {
|
||||||
|
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
|
||||||
|
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} carrying one usage record. */
|
||||||
|
private static void writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
|
||||||
|
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
|
||||||
|
Files.createDirectories(projectDir);
|
||||||
|
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** An {@link AgentControl} whose every {@code agent.get} answers with the given session/type/status. */
|
||||||
|
private static AgentControl agentControlStub(String sessionId, String agentType, String status) {
|
||||||
|
HerdrClient client = new HerdrClient() {
|
||||||
|
@Override
|
||||||
|
public JsonNode call(String method, Object params) {
|
||||||
|
if (!"agent.get".equals(method)) {
|
||||||
|
throw new HerdrException("stub has no canned response for " + method);
|
||||||
|
}
|
||||||
|
try {
|
||||||
|
String sessionField = sessionId == null ? ""
|
||||||
|
: ",\"agent_session\":{\"kind\":\"id\",\"value\":\"" + sessionId + "\"}";
|
||||||
|
String agentField = agentType == null ? "null" : "\"" + agentType + "\"";
|
||||||
|
return new ObjectMapper().readTree(("""
|
||||||
|
{"type":"agent_info","agent":{"terminal_id":"%s","agent":%s,
|
||||||
|
"agent_status":"%s"%s}}""")
|
||||||
|
.formatted(LEAD_TERMINAL, agentField, status, sessionField));
|
||||||
|
} catch (Exception e) {
|
||||||
|
throw new HerdrException("stub decode failed", e);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
}
|
||||||
|
};
|
||||||
|
return new AgentControl(client);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** An {@link AgentControl} whose every herdr call fails — models a herdr hiccup mid-tick. */
|
||||||
|
private static AgentControl throwingAgentControl() {
|
||||||
|
HerdrClient client = new HerdrClient() {
|
||||||
|
@Override
|
||||||
|
public JsonNode call(String method, Object params) {
|
||||||
|
throw new HerdrException("herdr unreachable (stub)");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
}
|
||||||
|
};
|
||||||
|
return new AgentControl(client);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("an unrecognised terminal resolves to UNKNOWN, not a thrown exception")
|
||||||
|
void unrecognisedTerminalResolvesToUnknown() {
|
||||||
|
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
|
||||||
|
new LeadContextGauge(), throwingAgentControl(), Map::of, name -> null);
|
||||||
|
|
||||||
|
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply("ghost-terminal"));
|
||||||
|
|
||||||
|
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
|
||||||
|
assertNull(reading.tokens());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("agents.get throwing degrades to UNKNOWN, no exception escapes")
|
||||||
|
void agentsGetThrowingDegradesToUnknown() {
|
||||||
|
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
|
||||||
|
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
|
||||||
|
new LeadContextGauge(), throwingAgentControl(), () -> liveLeadTerminals, name -> null);
|
||||||
|
|
||||||
|
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply(LEAD_TERMINAL));
|
||||||
|
|
||||||
|
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
|
||||||
|
"a herdr failure resolving the live agent must degrade to UNKNOWN, never kill the heartbeat's tick");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("the happy path resolves the configDir, sessionId and agentType through to the gauge")
|
||||||
|
void happyPathPassesResolvedFactsThroughToTheGauge(@TempDir Path tmp) throws IOException {
|
||||||
|
writeTranscript(tmp, SESSION_ID, 12_345);
|
||||||
|
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
|
||||||
|
AgentControl agents = agentControlStub(SESSION_ID, "claude", "idle");
|
||||||
|
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
|
||||||
|
new LeadContextGauge(), agents, () -> liveLeadTerminals,
|
||||||
|
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
|
||||||
|
|
||||||
|
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
|
||||||
|
|
||||||
|
assertEquals(LeadContextGauge.State.OK, reading.state(),
|
||||||
|
"the resolved configDir + the live agent's own sessionId/agentType must reach the gauge — "
|
||||||
|
+ "a wrong hop anywhere in the chain would read no transcript and report UNKNOWN instead");
|
||||||
|
assertEquals(12_345L, reading.tokens());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("a lead whose agent type is not claude still resolves to UNKNOWN, never a crash")
|
||||||
|
void nonClaudeAgentTypeResolvesToUnknown(@TempDir Path tmp) throws IOException {
|
||||||
|
writeTranscript(tmp, SESSION_ID, 12_345);
|
||||||
|
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
|
||||||
|
AgentControl agents = agentControlStub(SESSION_ID, "opencode", "idle");
|
||||||
|
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
|
||||||
|
new LeadContextGauge(), agents, () -> liveLeadTerminals,
|
||||||
|
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
|
||||||
|
|
||||||
|
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
|
||||||
|
|
||||||
|
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
|
||||||
|
"the agentType hop must reach the gauge too — a non-claude peer must not be misread as claude");
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,108 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import com.fasterxml.jackson.databind.JsonNode;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
import dev.ltms.fleet.herdr.AgentControl;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrException;
|
||||||
|
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||||
|
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.io.IOException;
|
||||||
|
import java.nio.charset.StandardCharsets;
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.Map;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #609: {@code Fleetd.main}'s {@code LeadHeartbeatLoop.LeadContextSource} local could be
|
||||||
|
* swapped for a bare {@code LeadHeartbeatLoop.LeadContextSource.none()} at the call site — compiling
|
||||||
|
* with 0 errors and leaving every pre-existing test green — exactly the shape #602/#606 already found
|
||||||
|
* for {@code LeadConfigDirSource} (see {@code FleetdLeadConfigDirSourceWiringTest}'s own javadoc for
|
||||||
|
* the measured version of that gap).
|
||||||
|
*
|
||||||
|
* <p>The fix follows the same pattern: {@link Fleetd#leadContextSource} is the extracted,
|
||||||
|
* directly-callable factory {@code main} calls to build the source it hands {@code
|
||||||
|
* LeadHeartbeatLoop}'s constructor. This test calls that exact factory and asserts it resolves a
|
||||||
|
* REAL reading off a real transcript file — a property that would be false if {@link
|
||||||
|
* Fleetd#leadContextSource} were mutated to {@code return LeadHeartbeatLoop.LeadContextSource.none();}.
|
||||||
|
*
|
||||||
|
* <p>What this class does not and cannot cover: {@code main}'s own one-line call to this factory
|
||||||
|
* could itself be swapped for {@code LeadHeartbeatLoop.LeadContextSource.none()}, bypassing this
|
||||||
|
* factory entirely — the same structural gap {@code FleetdLeadConfigDirSourceWiringTest} names for
|
||||||
|
* its own factory, and for the same reason (no test in this codebase calls {@code Fleetd.main} far
|
||||||
|
* enough to observe which factory call it made).
|
||||||
|
*/
|
||||||
|
class FleetdLeadContextSourceWiringTest {
|
||||||
|
|
||||||
|
/** Deliberately not {@code term_}-prefixed — see {@code FleetdLeadContextLookupTest}'s class doc. */
|
||||||
|
private static final String LEAD_TERMINAL = "leadpane1";
|
||||||
|
private static final String LEAD_NAME = "opus";
|
||||||
|
private static final String SESSION_ID = "sess-609-wiring";
|
||||||
|
|
||||||
|
private static String usageLine(long tokens) {
|
||||||
|
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
|
||||||
|
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
|
||||||
|
}
|
||||||
|
|
||||||
|
private static void writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
|
||||||
|
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
|
||||||
|
Files.createDirectories(projectDir);
|
||||||
|
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static AgentControl agentControlStub() {
|
||||||
|
HerdrClient client = new HerdrClient() {
|
||||||
|
@Override
|
||||||
|
public JsonNode call(String method, Object params) {
|
||||||
|
if (!"agent.get".equals(method)) {
|
||||||
|
throw new HerdrException("stub has no canned response for " + method);
|
||||||
|
}
|
||||||
|
try {
|
||||||
|
return new ObjectMapper().readTree(("""
|
||||||
|
{"type":"agent_info","agent":{"terminal_id":"%s","agent":"claude",
|
||||||
|
"agent_status":"idle","agent_session":{"kind":"id","value":"%s"}}}""")
|
||||||
|
.formatted(LEAD_TERMINAL, SESSION_ID));
|
||||||
|
} catch (Exception e) {
|
||||||
|
throw new HerdrException("stub decode failed", e);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
}
|
||||||
|
};
|
||||||
|
return new AgentControl(client);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("main's factory resolves a REAL reading, not the inert none() answer")
|
||||||
|
void resolvesARealReadingNotTheInertNoneAnswer(@TempDir Path tmp) throws IOException {
|
||||||
|
writeTranscript(tmp, SESSION_ID, 54_321);
|
||||||
|
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
|
||||||
|
|
||||||
|
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
|
||||||
|
agentControlStub(), () -> liveLeadTerminals, name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
|
||||||
|
|
||||||
|
LeadContextGauge.Reading reading = source.readingFor().apply(LEAD_TERMINAL);
|
||||||
|
|
||||||
|
assertEquals(LeadContextGauge.State.OK, reading.state(),
|
||||||
|
"mutating Fleetd.leadContextSource's own body to `return LeadHeartbeatLoop.LeadContextSource.none();` "
|
||||||
|
+ "must fail this assertion");
|
||||||
|
assertEquals(54_321L, reading.tokens());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("an unrecognised lead terminal resolves to UNKNOWN, not a thrown exception")
|
||||||
|
void unrecognisedTerminalResolvesToUnknown() {
|
||||||
|
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
|
||||||
|
agentControlStub(), Map::of, name -> null);
|
||||||
|
|
||||||
|
assertEquals(LeadContextGauge.State.UNKNOWN, source.readingFor().apply("ghost-terminal").state());
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -3,6 +3,7 @@ package dev.ltms.fleet;
|
|||||||
import ch.qos.logback.classic.Level;
|
import ch.qos.logback.classic.Level;
|
||||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||||
import dev.ltms.fleet.config.FleetConfig;
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||||
import dev.ltms.fleet.msg.LeadMailbox;
|
import dev.ltms.fleet.msg.LeadMailbox;
|
||||||
import dev.ltms.fleet.testing.CapturedLog;
|
import dev.ltms.fleet.testing.CapturedLog;
|
||||||
import org.junit.jupiter.api.Test;
|
import org.junit.jupiter.api.Test;
|
||||||
@@ -35,7 +36,7 @@ class FleetdLeadMailboxSelectionTest {
|
|||||||
boolean unreachable;
|
boolean unreachable;
|
||||||
|
|
||||||
@Override
|
@Override
|
||||||
public LeadMailbox open(String uri, String selfCoordId, int prefetch) {
|
public LeadChannelHandle open(String uri, String selfCoordId, int prefetch) {
|
||||||
this.offeredUri = uri;
|
this.offeredUri = uri;
|
||||||
this.offeredSelfId = selfCoordId;
|
this.offeredSelfId = selfCoordId;
|
||||||
this.offeredPrefetch = prefetch;
|
this.offeredPrefetch = prefetch;
|
||||||
|
|||||||
@@ -0,0 +1,293 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.AgentControl;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.lead.LeadRollover;
|
||||||
|
import dev.ltms.fleet.msg.ReplyInbox;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.LinkedHashMap;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
import static org.junit.jupiter.api.Assertions.fail;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 B3 — replaces {@code FleetdLeadRolloverWiringTest} (fleetd #480). That class was a
|
||||||
|
* source-text test scraping {@code Fleetd.java} (now {@code FleetdAssembly.java}, moved there by
|
||||||
|
* fleetd #612 Unit A) with three methods: {@code unrelatedAnchorStillPresent} (a scaffold anchor,
|
||||||
|
* not an independent claim — needs no replacement of its own), {@code
|
||||||
|
* mainStillCallsTheLeadRolloverFactory} (the call-site pin replaced by {@link
|
||||||
|
* #assembledLeadRolloverRunsTheRealClearAndBootstrapSequence}), and {@code
|
||||||
|
* factoryGatesOnConfigPresence} (the absent-config claim replaced by {@link
|
||||||
|
* #absentLeadRolloverConfigMeansNoRolloverIsBuilt} — a claim this ticket found was NOT actually
|
||||||
|
* covered behaviourally anywhere else: {@code LeadRolloverTest}'s only related assertion is
|
||||||
|
* vacuous, {@code assertNull(null)}, and never calls the real factory).
|
||||||
|
*
|
||||||
|
* <p><strong>fleetd #612 B3 correction (ticket comment 17553):</strong> the first version of this
|
||||||
|
* test configured a single shared {@link FakeHerdr} for both the lead and member herdr sockets.
|
||||||
|
* {@code FleetdAssembly.java:140-142} falls back to {@code memberHerdr = herdr} whenever no
|
||||||
|
* distinct {@code memberHerdrSocket} is configured, so with one fake, {@code
|
||||||
|
* router.leadAgents()} and {@code router.memberAgents()} wrapped the identical client — a
|
||||||
|
* mutation swapping {@code Fleetd.leadRollover(cfg, router.leadAgents(), config, leads)} for
|
||||||
|
* {@code ..., router.memberAgents(), ...} at {@code FleetdAssembly.java:408} was therefore
|
||||||
|
* invisible to this test, even though the two are genuinely different daemons in production. This
|
||||||
|
* version configures two distinct sockets and two distinct {@link FakeHerdr} instances (the same
|
||||||
|
* pattern {@code FleetdAssemblyConnectionIdentityTest}, fleetd #612 B2, already uses to separate
|
||||||
|
* lead from member) and asserts the roll's {@code /clear}/bootstrap sends land on the LEAD fake
|
||||||
|
* and never on the MEMBER one.
|
||||||
|
*/
|
||||||
|
class FleetdLeadRolloverAssemblyTest {
|
||||||
|
|
||||||
|
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
|
||||||
|
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
|
||||||
|
|
||||||
|
/** Keys {@code connectHerdr} by socket path so the lead and member daemons can be two
|
||||||
|
* DIFFERENT {@link FakeHerdr}s — same shape as B2's {@code FleetdAssemblyConnectionIdentityTest
|
||||||
|
* .TwoHerdrResourcePorts}. */
|
||||||
|
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||||
|
|
||||||
|
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
|
||||||
|
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
HerdrClient client = herdrsBySocket.get(socketPath);
|
||||||
|
if (client == null) {
|
||||||
|
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
|
||||||
|
}
|
||||||
|
return client;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
return (uri, prefetch) -> replyInbox;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
return (uri, selfCoordId, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException(
|
||||||
|
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
return Executors.newSingleThreadScheduledExecutor();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||||
|
return () -> {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
|
||||||
|
@Override
|
||||||
|
public void own(String target) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void release(String target) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void publish(String target, String msgId, String content) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public List<InboxMessage> peek(String target) {
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public boolean ack(String target, String msgId) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir, Path leadCwd) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
herdrSocket: "%s"
|
||||||
|
memberHerdrSocket: "%s"
|
||||||
|
idleSleepGuard:
|
||||||
|
enabled: false
|
||||||
|
broker:
|
||||||
|
uri: "amqp://fake-test-broker/vh"
|
||||||
|
fleet:
|
||||||
|
leaders:
|
||||||
|
opus:
|
||||||
|
tab: "lead: opus"
|
||||||
|
cwd: "%s"
|
||||||
|
leadRollover:
|
||||||
|
handoverPath: handover.md
|
||||||
|
requireOperatorConfirm: false
|
||||||
|
""".formatted(LEAD_SOCKET, MEMBER_SOCKET, leadCwd.toString()));
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
@Test
|
||||||
|
@DisplayName("[BEHAVIOURAL] the real assembled LeadRollover runs the full open/confirm/continuation "
|
||||||
|
+ "sequence — /clear, then bootstrapText — through the real herdr router")
|
||||||
|
void assembledLeadRolloverRunsTheRealClearAndBootstrapSequence(@TempDir Path dir) throws Exception {
|
||||||
|
Path leadCwd = dir.resolve("lead-workspace");
|
||||||
|
Files.createDirectories(leadCwd);
|
||||||
|
FleetConfig cfg = writeConfig(dir, leadCwd);
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||||
|
// Two DISTINCT fakes — one per configured socket — so leadAgents()/memberAgents() wrap
|
||||||
|
// genuinely different clients, exactly like production when memberHerdrSocket is set.
|
||||||
|
FakeHerdr lead = new FakeHerdr();
|
||||||
|
lead.withTab("w2", "w2:t7", "lead: opus");
|
||||||
|
FakeHerdr member = new FakeHerdr();
|
||||||
|
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
|
||||||
|
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
|
||||||
|
|
||||||
|
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
|
||||||
|
LeadRollover rollover = runtime.mcp().leadRollover();
|
||||||
|
assertNotNull(rollover, "leadRollover: is present in this test's config, so "
|
||||||
|
+ "FleetdAssembly.assembleAndStart must have built a real LeadRollover through the "
|
||||||
|
+ "Fleetd.leadRollover(...) call site — a mutation to `LeadRollover leadRollover = "
|
||||||
|
+ "null;` at that call site can never pass this");
|
||||||
|
|
||||||
|
LeadRollover.PendingRollover pending = rollover.open("term_a", "fleetd #612 B3 test");
|
||||||
|
String expectedHandoverPath = leadCwd.resolve("handover.md").normalize().toString();
|
||||||
|
assertEquals(expectedHandoverPath, pending.handoverPath());
|
||||||
|
|
||||||
|
// Ensure the handover file's mtime lands strictly AFTER open()'s requestedAtMillis —
|
||||||
|
// LeadRollover.checkHandover refuses on mtime <= requestedAt (HANDOVER_STALE).
|
||||||
|
Thread.sleep(50);
|
||||||
|
Files.writeString(Path.of(pending.handoverPath()), "handover content for fleetd #612 B3");
|
||||||
|
|
||||||
|
LeadRollover.RollDecision decision = rollover.confirm("term_a", pending.token(), true);
|
||||||
|
assertTrue(decision.accepted(), "confirm() must approve: requireOperatorConfirm is false, "
|
||||||
|
+ "the caller terminal matches open()'s, and the handover file exists, is non-empty "
|
||||||
|
+ "and fresh — got: " + decision);
|
||||||
|
|
||||||
|
// The production LeadRollover constructor always runs the post-confirm continuation on a
|
||||||
|
// real virtual thread (see Fleetd.leadRollover, which never passes the package-private test
|
||||||
|
// constructor), so this polls the real FleetMcp.leadRollover() instance's status(token)
|
||||||
|
// until the real continuation finishes.
|
||||||
|
LeadRollover.RollStatus status = pollUntilTerminal(rollover, pending.token());
|
||||||
|
assertEquals(LeadRollover.RollState.ROLLED, status.state(),
|
||||||
|
"the full happy path must complete: FakeHerdr's default agent status is 'idle', so "
|
||||||
|
+ "the turn-boundary wait settles immediately and the post-/clear wait "
|
||||||
|
+ "releases via its pickup-grace path — detail: " + status.detail());
|
||||||
|
|
||||||
|
// Prove the real herdr router actually sent BOTH messages, in order, to the real LEAD
|
||||||
|
// pane — this is the one thing a source-text pin on the call site could never show.
|
||||||
|
List<FakeHerdr.Call> prompts = lead.calls.stream()
|
||||||
|
.filter(c -> c.method().equals("agent.prompt"))
|
||||||
|
.toList();
|
||||||
|
assertTrue(prompts.size() >= 2, "expected at least a /clear send and a bootstrapText send "
|
||||||
|
+ "on the LEAD daemon, got " + prompts.size() + " agent.prompt calls: " + prompts);
|
||||||
|
assertEquals("/clear", ((Map<String, Object>) prompts.get(0).params()).get("text"),
|
||||||
|
"the first send must be the literal /clear housekeeping command");
|
||||||
|
Object secondText = ((Map<String, Object>) prompts.get(1).params()).get("text");
|
||||||
|
assertTrue(secondText instanceof String && ((String) secondText).contains(expectedHandoverPath),
|
||||||
|
"the second send must be the default bootstrapText naming the resolved handover "
|
||||||
|
+ "path, got: " + secondText);
|
||||||
|
|
||||||
|
// fleetd #612 B3 correction: prove the roll never touches the MEMBER daemon. A mutation
|
||||||
|
// swapping router.leadAgents() for router.memberAgents() at the real call site would move
|
||||||
|
// both sends above onto `member` instead, which this assertion catches — the thing the
|
||||||
|
// single-fake version of this test could never see, because both wrapped the same client.
|
||||||
|
List<FakeHerdr.Call> memberPrompts = member.calls.stream()
|
||||||
|
.filter(c -> c.method().equals("agent.prompt"))
|
||||||
|
.toList();
|
||||||
|
assertTrue(memberPrompts.isEmpty(), "the roll must be wired to the LEAD daemon only — got "
|
||||||
|
+ memberPrompts.size() + " agent.prompt call(s) on the MEMBER daemon instead: "
|
||||||
|
+ memberPrompts);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static LeadRollover.RollStatus pollUntilTerminal(LeadRollover rollover, String token)
|
||||||
|
throws InterruptedException {
|
||||||
|
long deadline = System.nanoTime() + java.util.concurrent.TimeUnit.SECONDS.toNanos(10);
|
||||||
|
while (System.nanoTime() < deadline) {
|
||||||
|
LeadRollover.RollStatus status = rollover.status(token);
|
||||||
|
if (status.state() != LeadRollover.RollState.PENDING
|
||||||
|
&& status.state() != LeadRollover.RollState.IN_PROGRESS) {
|
||||||
|
return status;
|
||||||
|
}
|
||||||
|
Thread.sleep(50);
|
||||||
|
}
|
||||||
|
fail("the real continuation did not reach a terminal state within 10s — last status: "
|
||||||
|
+ rollover.status(token));
|
||||||
|
throw new AssertionError("unreachable");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[BEHAVIOURAL] Fleetd.leadRollover(...) returns null when leadRollover: is absent "
|
||||||
|
+ "from config — the opt-in gate FleetdLeadRolloverWiringTest's "
|
||||||
|
+ "factoryGatesOnConfigPresence pinned by source text alone")
|
||||||
|
void absentLeadRolloverConfigMeansNoRolloverIsBuilt(@TempDir Path dir) throws Exception {
|
||||||
|
Path yaml = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(yaml, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
""");
|
||||||
|
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
|
||||||
|
AgentControl agents = new AgentControl(new FakeHerdr());
|
||||||
|
|
||||||
|
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, config, Map::of);
|
||||||
|
|
||||||
|
assertNull(rollover, "leadRollover: is absent from this config, so the factory's opt-in "
|
||||||
|
+ "gate (`if (cfg.leadRollover() == null) return null;`) must fire and no "
|
||||||
|
+ "LeadRollover must be constructed at all");
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,89 +0,0 @@
|
|||||||
package dev.ltms.fleet;
|
|
||||||
|
|
||||||
import java.nio.file.Files;
|
|
||||||
import java.nio.file.Path;
|
|
||||||
import org.junit.jupiter.api.DisplayName;
|
|
||||||
import org.junit.jupiter.api.Test;
|
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
|
||||||
|
|
||||||
/**
|
|
||||||
* fleetd #480 Unit A, hard requirement 6: pin {@code Fleetd.main}'s construction of {@link
|
|
||||||
* dev.ltms.fleet.lead.LeadRollover} with a source-text assertion, mirroring {@code
|
|
||||||
* FleetdCompletionResolverWiringTest}'s pattern — five log-only reporters in {@code Fleetd.main}
|
|
||||||
* already survived mutation batteries this exact way (fleetd #415's extraction antidote note).
|
|
||||||
*
|
|
||||||
* <p>What this class still covers, and what it never claimed to. {@code LeadRolloverTest}
|
|
||||||
* constructs its own {@code LeadRollover} directly (as every prior test of an extracted factory
|
|
||||||
* does) with a hand-built lookup, so a mutation that deletes the {@code leadRollover(...)} call
|
|
||||||
* from {@code main} — or replaces one of its arguments with something that still compiles, e.g.
|
|
||||||
* {@code router.leadAgents()} swapped for {@code null}, or the whole assignment swapped for a bare
|
|
||||||
* {@code null} literal — leaves every behavioural test green. This is a plain string read, guarded
|
|
||||||
* by an unrelated anchor assertion so a broken or empty file read cannot pass as a real change.
|
|
||||||
*
|
|
||||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
|
||||||
* LeadRollover} and never runs {@code main}. It pins the {@code leadRollover(...)} CALL SITE's
|
|
||||||
* argument list — that {@code main} still passes {@code leads} at all — never what the factory
|
|
||||||
* DOES with that argument once inside its own body.
|
|
||||||
*
|
|
||||||
* <p><b>Correction (fleetd #480 relative-handover-path follow-up): that gap used to be real, and
|
|
||||||
* now is not — but not here.</b> This class's javadoc previously claimed "no behavioural test can
|
|
||||||
* catch this wiring dropping out" for the whole factory, including the lambda {@code
|
|
||||||
* leadRollover(...)} builds internally (terminal → lead name → {@code Leader.cwd()}). That claim
|
|
||||||
* was proven true at the time — mutating that lambda's body to {@code String leadName = null;}
|
|
||||||
* (always "no lead found", which silently reintroduces the daemon-cwd bug this ticket fixes) left
|
|
||||||
* the full suite green, {@code Tests run: 1669, Failures: 0}. It is no longer true: {@code
|
|
||||||
* FleetdLeadRolloverWorkspaceLookupTest} now calls {@code Fleetd.leadRollover(...)} directly with a
|
|
||||||
* real {@link dev.ltms.fleet.config.ConfigRef} built from a temp {@code fleetd.yaml}, and fails
|
|
||||||
* against that exact one-line mutation. So: THIS class still covers only the call site's argument
|
|
||||||
* list; {@code FleetdLeadRolloverWorkspaceLookupTest} is what now covers the lambda's body. Neither
|
|
||||||
* one subsumes the other — keep both.
|
|
||||||
*/
|
|
||||||
class FleetdLeadRolloverWiringTest {
|
|
||||||
|
|
||||||
private static String fleetdSource() throws Exception {
|
|
||||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
|
||||||
}
|
|
||||||
|
|
||||||
@Test
|
|
||||||
@DisplayName("[SOURCE TEXT] unrelated anchor: Fleetd.java still declares the Fleetd class")
|
|
||||||
void unrelatedAnchorStillPresent() throws Exception {
|
|
||||||
// Guards the two assertions below: without this, a bad read (empty string, wrong file,
|
|
||||||
// truncated file) could vacuously fail to contain the leadRollover(...) call too, and a
|
|
||||||
// test that only asserts "contains X" would report a false pass for the wrong reason if X
|
|
||||||
// happened to match. Asserting an unrelated, structurally distant string first proves the
|
|
||||||
// read actually pulled real file content.
|
|
||||||
String source = fleetdSource();
|
|
||||||
assertTrue(source.contains("public final class Fleetd"),
|
|
||||||
"sanity anchor failed — the file read did not return real Fleetd.java source; the "
|
|
||||||
+ "leadRollover(...) wiring assertions below cannot be trusted until this passes");
|
|
||||||
}
|
|
||||||
|
|
||||||
@Test
|
|
||||||
@DisplayName("[SOURCE TEXT] main still constructs LeadRollover via the leadRollover(...) factory, exactly as heartbeat is constructed")
|
|
||||||
void mainStillCallsTheLeadRolloverFactory() throws Exception {
|
|
||||||
String source = fleetdSource();
|
|
||||||
assertTrue(source.contains(
|
|
||||||
"LeadRollover leadRollover = leadRollover(cfg, router.leadAgents(), config, leads);"),
|
|
||||||
"Fleetd.main must still assign `LeadRollover leadRollover = leadRollover(cfg, "
|
|
||||||
+ "router.leadAgents(), config, leads);`. Dropping this call, or swapping one of "
|
|
||||||
+ "its arguments for something that still compiles (e.g. null in place of "
|
|
||||||
+ "router.leadAgents()), leaves every behavioural test green — this source check is "
|
|
||||||
+ "what must go red instead. fleetd #480 correction 2 deliberately dropped "
|
|
||||||
+ "primaryRegistry from this call — see LeadRollover's class javadoc for why a "
|
|
||||||
+ "single-slot lookup was wrong here. The fleetd #480 relative-handover-path "
|
|
||||||
+ "follow-up added `leads` (terminal → lead name) so the factory can resolve a "
|
|
||||||
+ "relative handoverPath against the calling lead's own workspace.");
|
|
||||||
}
|
|
||||||
|
|
||||||
@Test
|
|
||||||
@DisplayName("[SOURCE TEXT] the leadRollover(...) factory itself gates construction on cfg.leadRollover() != null")
|
|
||||||
void factoryGatesOnConfigPresence() throws Exception {
|
|
||||||
String source = fleetdSource();
|
|
||||||
assertTrue(source.contains("if (cfg.leadRollover() == null) {"),
|
|
||||||
"Fleetd.leadRollover(...) must refuse to construct a LeadRollover when the "
|
|
||||||
+ "leadRollover: block is absent — an upgraded daemon must never silently acquire "
|
|
||||||
+ "the ability to clear the lead's own pane. See LeadHeartbeatLoop's construction "
|
|
||||||
+ "gate (cfg.leadHeartbeat() != null) for the pattern this mirrors.");
|
|
||||||
}
|
|
||||||
}
|
|
||||||
@@ -0,0 +1,180 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.mcp.FleetMcp;
|
||||||
|
import dev.ltms.fleet.msg.ReplyInbox;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 B3 — replaces {@code FleetdLeadSeatWiringTest} (fleetd #176), a source-text test that
|
||||||
|
* scraped {@code Fleetd.java} (now {@code FleetdAssembly.java}, moved there by fleetd #612 Unit A)
|
||||||
|
* for the exact {@code new FleetMcp.LeadSeatSource(Fleetd.leadSeatLookup(...))} constructor-call
|
||||||
|
* text. That proves the right symbols appear in source; it proves nothing about what the daemon's
|
||||||
|
* live {@code fleet_list} actually reports.
|
||||||
|
*
|
||||||
|
* <p>This test instead drives the REAL {@link FleetMcp.LeadSeatSource} the real {@link
|
||||||
|
* FleetdAssembly#assembleAndStart} builds — including the REAL {@code LeadTabScanner} it wires
|
||||||
|
* {@code Fleetd.leadSeatLookup} through — reached via {@link FleetMcp#leadSeatSource()} on the
|
||||||
|
* live {@code FleetMcp} {@code FleetdRuntime} owns. It seeds one FakeHerdr tab labelled to match a
|
||||||
|
* configured {@code fleet.leaders.opus.tab}, with a live agent already in it (FakeHerdr's own
|
||||||
|
* default {@code agent.list}/{@code pane.list} entries for {@code term_a}/{@code w2:p7}/{@code
|
||||||
|
* w2:t7} — no FakeHerdr change needed), and asserts the assembled seat source reports exactly the
|
||||||
|
* seat {@link FleetMcp.LeadSeatSource#none()} (the inert stand-in) could never produce: 1, not 0.
|
||||||
|
*/
|
||||||
|
class FleetdLeadSeatAssemblyTest {
|
||||||
|
|
||||||
|
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||||
|
|
||||||
|
final FakeHerdr herdr = new FakeHerdr();
|
||||||
|
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
return herdr;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
return (uri, prefetch) -> replyInbox;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
return (uri, selfCoordId, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException(
|
||||||
|
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return System::nanoTime;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
return Executors.newSingleThreadScheduledExecutor();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||||
|
return () -> {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
|
||||||
|
@Override
|
||||||
|
public void own(String target) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void release(String target) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void publish(String target, String msgId, String content) {
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public List<InboxMessage> peek(String target) {
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public boolean ack(String target, String msgId) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
idleSleepGuard:
|
||||||
|
enabled: false
|
||||||
|
broker:
|
||||||
|
uri: "amqp://fake-test-broker/vh"
|
||||||
|
fleet:
|
||||||
|
leaders:
|
||||||
|
opus:
|
||||||
|
tab: "lead: opus"
|
||||||
|
profile: sonnet
|
||||||
|
profiles:
|
||||||
|
sonnet:
|
||||||
|
subscription: true
|
||||||
|
argv: ["ccs", "sonnet"]
|
||||||
|
""");
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[BEHAVIOURAL] the real assembled LeadSeatSource, backed by the real LeadTabScanner, "
|
||||||
|
+ "reports a live lead's seat against its own subscription profile")
|
||||||
|
void assembledLeadSeatSourceReportsALiveLeadsSeat(@TempDir Path dir) throws Exception {
|
||||||
|
FleetConfig cfg = writeConfig(dir);
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||||
|
// Label FakeHerdr's own default pane's tab (term_a / w2:p7 / w2:t7, already carrying a live
|
||||||
|
// agent) to match fleet.leaders.opus.tab exactly — no FakeHerdr change needed at all.
|
||||||
|
ports.herdr.withTab("w2", "w2:t7", "lead: opus");
|
||||||
|
|
||||||
|
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
|
||||||
|
FleetMcp.LeadSeatSource seatSource = runtime.mcp().leadSeatSource();
|
||||||
|
assertEquals(1, seatSource.seatsFor().apply("sonnet"),
|
||||||
|
"the real LeadTabScanner recognises the labelled tab as a live 'opus' lead on "
|
||||||
|
+ "profile 'sonnet' (same credential, matched by Fleetd.leadSeatLookup), so "
|
||||||
|
+ "subscription profile 'sonnet' must be charged one seat — "
|
||||||
|
+ "FleetMcp.LeadSeatSource.none() (the inert stand-in this test's mutation "
|
||||||
|
+ "swaps the call site for) always reports 0, whatever the input");
|
||||||
|
|
||||||
|
// A profile no lead is running on gets no seat charged — the same seat source, applied to
|
||||||
|
// an input that must stay at the inert answer even on the real, non-inert instance.
|
||||||
|
assertEquals(0, seatSource.seatsFor().apply("no-such-profile"));
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,43 +0,0 @@
|
|||||||
package dev.ltms.fleet;
|
|
||||||
|
|
||||||
import java.nio.file.Files;
|
|
||||||
import java.nio.file.Path;
|
|
||||||
import org.junit.jupiter.api.DisplayName;
|
|
||||||
import org.junit.jupiter.api.Test;
|
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
|
||||||
|
|
||||||
/**
|
|
||||||
* fleetd #176: {@code Fleetd.main} builds its {@code FleetMcp} from a 14-argument constructor whose
|
|
||||||
* last argument is a {@code FleetMcp.LeadSeatSource} wrapping {@link Fleetd#leadSeatLookup}. That
|
|
||||||
* argument is exactly the kind of wiring fleetd #248 warned about: dropping it (or swapping it for
|
|
||||||
* the inert {@code FleetMcp.LeadSeatSource.none()}) compiles with 0 errors and leaves every test
|
|
||||||
* that builds its own {@code FleetMcp}/{@code CapacitySource} directly — every test that predates
|
|
||||||
* this ticket — green, because none of them go through {@code main} at all.
|
|
||||||
*
|
|
||||||
* <p>{@link FleetdLeadSeatLookupTest} proves the factory's own matching logic; this class is the
|
|
||||||
* plain source-text assertion that proves {@code main} still passes its result in, mirroring
|
|
||||||
* {@code FleetdCompletionResolverWiringTest}'s approach for the same class of gap.
|
|
||||||
*
|
|
||||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a
|
|
||||||
* {@code FleetMcp} and never runs {@code main}.
|
|
||||||
*/
|
|
||||||
class FleetdLeadSeatWiringTest {
|
|
||||||
|
|
||||||
private static String fleetdSource() throws Exception {
|
|
||||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
|
||||||
}
|
|
||||||
|
|
||||||
@Test
|
|
||||||
@DisplayName("[SOURCE TEXT] FleetMcp's construction call still passes a LeadSeatSource built from leadSeatLookup(...)")
|
|
||||||
void fleetMcpConstructionStillWiresLeadSeatLookup() throws Exception {
|
|
||||||
String source = fleetdSource();
|
|
||||||
assertTrue(source.contains("new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), "
|
|
||||||
+ "leaders, leads))"),
|
|
||||||
"FleetMcp's construction call must still pass a LeadSeatSource built from "
|
|
||||||
+ "Fleetd.leadSeatLookup(...). Dropping it or swapping in "
|
|
||||||
+ "FleetMcp.LeadSeatSource.none() (fleetd #176's would-be silent regression, the same "
|
|
||||||
+ "shape as fleetd #248's measured mutations) compiles with 0 errors and leaves every "
|
|
||||||
+ "existing behavioural test green — this source check is what must go red instead.");
|
|
||||||
}
|
|
||||||
}
|
|
||||||
@@ -0,0 +1,72 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #589 Group 1: {@link Fleetd#liveExhaustedPatterns} is the factory that replaced {@code
|
||||||
|
* main}'s inline {@code new LiveExhaustedPatterns(() -> config.get().profiles())}. Before this
|
||||||
|
* ticket, that supplier argument was untestable wiring: replacing it with a hardcoded {@code () ->
|
||||||
|
* Map.of()} compiled with 0 errors and left every existing test green, because {@code
|
||||||
|
* LiveExhaustedPatternsTest} builds its own instance directly with a hand-supplied map and never
|
||||||
|
* goes through {@code main}'s call site.
|
||||||
|
*
|
||||||
|
* <p>Silently losing this wiring means every profile's {@code exhaustedPattern} stops being
|
||||||
|
* recognised — {@link Fleetd#exhaustedPatternLookup} would never see a match, and a genuine
|
||||||
|
* usage-limit refusal would be handed back to a waiting {@code fleet_send} as real completed work.
|
||||||
|
*/
|
||||||
|
class FleetdLiveExhaustedPatternsWiringTest {
|
||||||
|
|
||||||
|
private static final String YAML = """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
herdrSocket: ~/.config/herdr/herdr.sock
|
||||||
|
profiles:
|
||||||
|
terra:
|
||||||
|
baseUrl: http://gx00.gw:8000
|
||||||
|
model: claude-opus-5
|
||||||
|
exhaustedPattern: "usage limit"
|
||||||
|
gx:
|
||||||
|
baseUrl: http://gx00.gw:8000
|
||||||
|
""";
|
||||||
|
|
||||||
|
private static ConfigRef loadConfig(Path dir) throws Exception {
|
||||||
|
Path file = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(file, YAML);
|
||||||
|
FleetConfig cfg = FleetConfig.load(file);
|
||||||
|
return ConfigRef.fixed(cfg);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("a profile with a configured exhaustedPattern is armed, with a compiled matcher")
|
||||||
|
void configuredProfileIsArmed(@TempDir Path dir) throws Exception {
|
||||||
|
LiveExhaustedPatterns patterns = Fleetd.liveExhaustedPatterns(loadConfig(dir));
|
||||||
|
|
||||||
|
assertTrue(patterns.armed("terra"),
|
||||||
|
"the config's live profiles() supplier must reach LiveExhaustedPatterns — hardcoding "
|
||||||
|
+ "the supplier to () -> Map.of() at the Fleetd.liveExhaustedPatterns call "
|
||||||
|
+ "site must fail this assertion");
|
||||||
|
assertTrue(patterns.patternFor("terra").matcher("the usage limit has been reached").find());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("a profile with no configured exhaustedPattern is not armed, but is still resolvable")
|
||||||
|
void unconfiguredProfileIsNotArmed(@TempDir Path dir) throws Exception {
|
||||||
|
LiveExhaustedPatterns patterns = Fleetd.liveExhaustedPatterns(loadConfig(dir));
|
||||||
|
|
||||||
|
assertFalse(patterns.armed("gx"), "'gx' has no exhaustedPattern configured");
|
||||||
|
assertNull(patterns.patternFor("gx"));
|
||||||
|
}
|
||||||
|
}
|
||||||
+197
@@ -0,0 +1,197 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||||
|
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||||
|
import dev.ltms.fleet.member.HerdrPeerLauncher;
|
||||||
|
import dev.ltms.fleet.peer.PeerLauncher;
|
||||||
|
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.lang.reflect.Field;
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.Executors;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #612 step 4, rank 2 (OpenCode half) — {@link FleetdAssembly}'s {@code
|
||||||
|
* forwardingExhaustionSink} line ({@code Fleetd.forwardingExhaustionSink(exhaustionSinkRef)}),
|
||||||
|
* handed to {@link dev.ltms.fleet.member.OpenCodeLauncher} so its fleetd #175 model-mismatch check
|
||||||
|
* can quarantine a credential before {@code sessions} exists to build the real sink (the
|
||||||
|
* construction-order cycle documented at that call site). The ticket calls this independent from
|
||||||
|
* {@code publishExhaustionSink} (pinned by {@link FleetdExhaustedPatternAssemblyTest}): a credential
|
||||||
|
* that hits a usage limit through THIS path is never quarantined if {@code forwardingExhaustionSink}
|
||||||
|
* is swapped for a hardcoded {@link ExhaustionSink#none()} at that call site — the OpenCode
|
||||||
|
* launcher's own quarantine check keeps compiling and keeps "running", but it permanently talks to
|
||||||
|
* a sink that does nothing, independent of whatever {@code publishExhaustionSink} does later.
|
||||||
|
*
|
||||||
|
* <p>{@code FleetdExhaustionSinkForwardingWiringTest} (fleetd #589) already proves {@code
|
||||||
|
* Fleetd.forwardingExhaustionSink(ref)} forwards to whatever {@code ref} holds — as a bare factory
|
||||||
|
* call, never through {@link FleetdAssembly#assembleAndStart}. It proves nothing about whether
|
||||||
|
* THIS call site is the one FleetdAssembly actually wires into the real {@code OpenCodeLauncher}
|
||||||
|
* it builds, which is exactly the #602/#606-shaped gap this ticket exists to close.
|
||||||
|
*
|
||||||
|
* <p>No accessor on {@link FleetdRuntime} reaches the adapter instances (by design — see that
|
||||||
|
* class's own javadoc: only the final collaborators it owns directly are exposed), so this test
|
||||||
|
* reaches the REAL, assembled {@code OpenCodeLauncher}'s {@code exhaustionSink} field the same way
|
||||||
|
* {@code SessionManager}/{@code CompositePeerLauncher} wire it internally: a short, targeted
|
||||||
|
* reflective walk ({@code SessionManager.launcher} → {@code CompositePeerLauncher.byProfile} →
|
||||||
|
* {@code OpenCodeLauncher.exhaustionSink}) onto the exact object the assembly built — never a copy,
|
||||||
|
* and never a read of the source text. Reflection is used the same way elsewhere in this suite
|
||||||
|
* (e.g. {@code StatusPollerWatchdogTest}) to reach a private collaborator a production constructor
|
||||||
|
* intentionally does not expose a public accessor for.
|
||||||
|
*/
|
||||||
|
class FleetdOpenCodeExhaustionForwardingAssemblyTest {
|
||||||
|
|
||||||
|
private static final class ControllableResourcePorts implements ResourcePorts {
|
||||||
|
|
||||||
|
final FakeHerdr herdr = new FakeHerdr();
|
||||||
|
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L);
|
||||||
|
Runnable shutdownHook;
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
return herdr;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
return (uri, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException("replyInboxOpener must not be called — no broker: block");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
return (uri, selfCoordId, prefetch) -> {
|
||||||
|
throw new UnsupportedOperationException("leadMailboxOpener must not be called — no coordinator: block");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||||
|
return () -> {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||||
|
};
|
||||||
|
}
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
return nowNanos::get;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
return nowNanos::get;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
return Executors.newSingleThreadScheduledExecutor();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
this.shutdownHook = hook;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
// Deliberately never bind a real port.
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig writeConfig(Path dir, int cooldownSeconds) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
idleSleepGuard:
|
||||||
|
enabled: false
|
||||||
|
quarantineCooldownSeconds: %d
|
||||||
|
profiles:
|
||||||
|
gemini:
|
||||||
|
kind: opencode
|
||||||
|
model: google/gemini-2.5-pro
|
||||||
|
""".formatted(cooldownSeconds));
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Reach a declared field by name on {@code target}'s runtime class, bypassing the access check. */
|
||||||
|
private static Object readField(Object target, Class<?> declaringClass, String fieldName) throws Exception {
|
||||||
|
Field field = declaringClass.getDeclaredField(fieldName);
|
||||||
|
field.setAccessible(true);
|
||||||
|
return field.get(target);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[BEHAVIOURAL] the real assembled OpenCodeLauncher's exhaustionSink field forwards "
|
||||||
|
+ "an onExhausted call into the real daemon's BackendQuarantine")
|
||||||
|
void assembledOpenCodeLauncherExhaustionSinkQuarantinesTheCredential(@TempDir Path dir) throws Exception {
|
||||||
|
int cooldownSeconds = 90;
|
||||||
|
FleetConfig cfg = writeConfig(dir, cooldownSeconds);
|
||||||
|
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||||
|
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||||
|
ControllableResourcePorts ports = new ControllableResourcePorts();
|
||||||
|
|
||||||
|
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||||
|
try {
|
||||||
|
PeerLauncher launcherField = (PeerLauncher) readField(runtime.sessions(),
|
||||||
|
runtime.sessions().getClass(), "launcher");
|
||||||
|
// CONTROL: the composite launcher must actually be the real production type with a
|
||||||
|
// "gemini" -> OpenCodeLauncher entry — if this fails, nothing below exercised the real
|
||||||
|
// assembly at all, rather than silently passing on an empty/wrong object.
|
||||||
|
assertTrue(launcherField instanceof CompositePeerLauncher,
|
||||||
|
"CONTROL: SessionManager.launcher must be the real CompositePeerLauncher the "
|
||||||
|
+ "assembly built, got: " + launcherField);
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
Map<String, HerdrPeerLauncher> byProfile = (Map<String, HerdrPeerLauncher>)
|
||||||
|
readField(launcherField, CompositePeerLauncher.class, "byProfile");
|
||||||
|
HerdrPeerLauncher adapter = byProfile.get("gemini");
|
||||||
|
assertTrue(adapter != null && adapter.getClass().getSimpleName().equals("OpenCodeLauncher"),
|
||||||
|
"CONTROL: the 'gemini' profile must resolve to a real OpenCodeLauncher adapter, "
|
||||||
|
+ "got: " + adapter);
|
||||||
|
|
||||||
|
ExhaustionSink sink = (ExhaustionSink) readField(adapter, adapter.getClass(), "exhaustionSink");
|
||||||
|
assertTrue(sink != null, "CONTROL: OpenCodeLauncher.exhaustionSink must never be null");
|
||||||
|
|
||||||
|
// The exact call OpenCodeLauncher.SessionAwareHandle#checkModelMatch makes on a real
|
||||||
|
// model mismatch (fleetd #175): target, reason, and its own already-known profile name.
|
||||||
|
sink.onExhausted("term_gemini_1", "opencode model mismatch (test)", "gemini");
|
||||||
|
|
||||||
|
BackendQuarantine quarantine = runtime.mcp().quarantineSource().quarantine();
|
||||||
|
assertTrue(quarantine.isQuarantined("gemini"),
|
||||||
|
"the real forwardingExhaustionSink-wired field must have delegated into the "
|
||||||
|
+ "published production sink, which quarantines the profile's credential "
|
||||||
|
+ "('gemini' here — no credentialId configured) — replacing "
|
||||||
|
+ "FleetdAssembly's forwardingExhaustionSink call site with a hardcoded "
|
||||||
|
+ "ExhaustionSink.none() must fail this assertion, since the field read "
|
||||||
|
+ "above would then BE the inert no-op and nothing would ever reach "
|
||||||
|
+ "quarantine.quarantine(...)");
|
||||||
|
assertEquals(cooldownSeconds, quarantine.remainingSeconds("gemini").orElseThrow(
|
||||||
|
() -> new AssertionError("credential must report a remaining cooldown")));
|
||||||
|
} finally {
|
||||||
|
if (ports.shutdownHook != null) ports.shutdownHook.run();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,77 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.ConfigRef;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.herdr.AgentControl;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||||
|
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||||
|
import dev.ltms.fleet.member.OpenCodeLauncher;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.Map;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #589 Group 2: {@link Fleetd#openCodeLauncher} is the factory that replaced {@code main}'s
|
||||||
|
* inline {@code new OpenCodeLauncher(...)} call — the {@code opencode} counterpart to {@link
|
||||||
|
* Fleetd#claudeCodeLauncher}, extracted for the identical reason. Its {@code memberCredentials}
|
||||||
|
* argument is the same {@code () -> config.get().memberCredentials()} supplier; replacing it with
|
||||||
|
* {@code () -> null} compiled with 0 errors and left every existing test green before this ticket,
|
||||||
|
* reopening the same CB-592 exposure gap CB-596's policy closed.
|
||||||
|
*
|
||||||
|
* <p>Same observable surface as {@code OpenCodeLauncherTest}'s own {@code memberCredentials} tests:
|
||||||
|
* spawn through the launcher {@code main} actually wires and inspect what {@code tab.create}
|
||||||
|
* carried.
|
||||||
|
*/
|
||||||
|
class FleetdOpenCodeLauncherCredentialWiringTest {
|
||||||
|
|
||||||
|
private static final String YAML = """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
herdrSocket: ~/.config/herdr/herdr.sock
|
||||||
|
profiles:
|
||||||
|
gemini:
|
||||||
|
kind: opencode
|
||||||
|
model: google/gemini-2.5-pro
|
||||||
|
memberCredentials:
|
||||||
|
policy: deny-by-default
|
||||||
|
known:
|
||||||
|
- GITEA_ACCESS_TOKEN
|
||||||
|
""";
|
||||||
|
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
private static Map<String, String> startEnv(FakeHerdr herdr) {
|
||||||
|
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("main's memberCredentials wiring reaches OpenCodeLauncher: a known-but-not-allowed "
|
||||||
|
+ "name is shadowed on spawn")
|
||||||
|
void memberCredentialsWiringReachesOpenCodeLauncher(@TempDir Path dir) throws Exception {
|
||||||
|
Path file = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(file, YAML);
|
||||||
|
FleetConfig cfg = FleetConfig.load(file);
|
||||||
|
ConfigRef config = ConfigRef.fixed(cfg);
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
|
||||||
|
OpenCodeLauncher launcher = Fleetd.openCodeLauncher(new AgentControl(herdr),
|
||||||
|
new WorkspaceControl(herdr), cfg.profiles(), cfg, config, ExhaustionSink.none());
|
||||||
|
launcher.spawn();
|
||||||
|
|
||||||
|
String shadowed = startEnv(herdr).get("GITEA_ACCESS_TOKEN");
|
||||||
|
assertNotNull(shadowed,
|
||||||
|
"GITEA_ACCESS_TOKEN is 'known' but not 'allow'-ed in the loaded config — it must be "
|
||||||
|
+ "explicitly shadowed on spawn; replacing the memberCredentials supplier with "
|
||||||
|
+ "() -> null at the Fleetd.openCodeLauncher call site must fail this "
|
||||||
|
+ "assertion, since a null policy shadows nothing");
|
||||||
|
assertFalse(shadowed.isBlank(), "the overlay value must be non-blank");
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,202 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.guard.GuardException;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import io.javalin.Javalin;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.function.LongSupplier;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #625: pins {@link dev.ltms.fleet.guard.SubscriptionGuard#assertPrimaryClean}'s call site
|
||||||
|
* in {@link Fleetd#main(String[])} — the ONE place it runs at startup, and the check behind the
|
||||||
|
* bridge charter's invariant 1 (never let the primary carry {@code ANTHROPIC_BASE_URL}). Nothing
|
||||||
|
* pinned it before this ticket: deleting {@code guard.assertPrimaryClean(...)} from {@code main}
|
||||||
|
* left the full suite green, because the call site read the real process environment ({@code
|
||||||
|
* System.getenv()}), which a test cannot taint from inside the JVM.
|
||||||
|
*
|
||||||
|
* <p>{@link Fleetd#main(String[], ResourcePorts)} (added by this ticket) is the literal production
|
||||||
|
* sequence — not a copy of it — driven here with a {@link ResourcePorts} whose {@link
|
||||||
|
* ResourcePorts#environment()} is a plain {@code Map} a test controls. The guard itself was
|
||||||
|
* already pinned by {@code SubscriptionGuardTest}, directly, with a {@code Map} — that proves the
|
||||||
|
* method's behaviour, not that {@code main} still calls it at the right point. This class pins the
|
||||||
|
* call site and, separately, the ORDER: the guard must still run before {@code cfg.validateAll()}
|
||||||
|
* and before {@code FleetdAssembly.assembleAndStart} touches a socket, a broker, or HTTP — not just
|
||||||
|
* be present somewhere in {@code main}.
|
||||||
|
*
|
||||||
|
* <p>Presence alone is not enough (a fix that pins only presence trades an invisible deletion for
|
||||||
|
* an invisible reordering), so each test below is built so that EITHER deleting the guard call OR
|
||||||
|
* moving it later makes the <em>same</em> test fail — with a different exception type than the one
|
||||||
|
* asserted, not a vacuous pass. See each test's own javadoc for how.
|
||||||
|
*/
|
||||||
|
class FleetdSubscriptionGuardOrderingTest {
|
||||||
|
|
||||||
|
private static final Map<String, String> TAINTED_ENV =
|
||||||
|
Map.of("ANTHROPIC_BASE_URL", "http://tainted.example");
|
||||||
|
private static final Map<String, String> CLEAN_ENV = Map.of("PATH", "/usr/bin");
|
||||||
|
|
||||||
|
/**
|
||||||
|
* A {@link ResourcePorts} whose {@link #environment()} is fixed to whatever the test hands it,
|
||||||
|
* and whose every other method refuses to be called at all. That refusal is the ordering pin:
|
||||||
|
* if {@code main} ever reaches {@link FleetdAssembly#assembleAndStart} before the guard has had
|
||||||
|
* a chance to throw, the very first thing the assembly does with {@code ports} is {@link
|
||||||
|
* #connectHerdr} — so a test that expects {@link GuardException} and instead observes {@link
|
||||||
|
* UnsupportedOperationException} has just caught the guard running too late (or not at all).
|
||||||
|
*/
|
||||||
|
private static final class FixedEnvPorts implements ResourcePorts {
|
||||||
|
private final Map<String, String> env;
|
||||||
|
|
||||||
|
FixedEnvPorts(Map<String, String> env) {
|
||||||
|
this.env = env;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Map<String, String> environment() {
|
||||||
|
return env;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public HerdrClient connectHerdr(Path socketPath) {
|
||||||
|
throw new UnsupportedOperationException("connectHerdr must not be called before the guard runs");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||||
|
throw new UnsupportedOperationException("replyInboxOpener must not be called before the guard runs");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||||
|
throw new UnsupportedOperationException("leadMailboxOpener must not be called before the guard runs");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier nanoClock() {
|
||||||
|
throw new UnsupportedOperationException("nanoClock must not be called before the guard runs");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public LongSupplier wallClockNanos() {
|
||||||
|
throw new UnsupportedOperationException("wallClockNanos must not be called before the guard runs");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledExecutorService newScheduler(String purpose) {
|
||||||
|
throw new UnsupportedOperationException("newScheduler must not be called before the guard runs");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void addShutdownHook(Runnable hook) {
|
||||||
|
throw new UnsupportedOperationException("addShutdownHook must not be called before the guard runs");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void startHttp(Javalin app, String host, int port) {
|
||||||
|
throw new UnsupportedOperationException("startHttp must not be called before the guard runs");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Runnable herdrPollWait() {
|
||||||
|
throw new UnsupportedOperationException("herdrPollWait must not be called before the guard runs");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Otherwise-invalid: {@code bind.host: 0.0.0.0} with no {@code auth.mode: token} fails {@code
|
||||||
|
* cfg.validateAll()} (CB-501's auth-exposure check — the same fixture {@code
|
||||||
|
* FleetdStartupValidationTest#mainRefusesANonLoopbackBindWithoutTokenMode} uses), with an
|
||||||
|
* {@link IllegalStateException}. That is deliberate: it is what {@code main} would throw INSTEAD
|
||||||
|
* of {@link GuardException} if the guard call were deleted, or moved to run after {@code
|
||||||
|
* validateAll()} — a different, distinguishable exception type.
|
||||||
|
*/
|
||||||
|
private static Path invalidConfig(Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 0.0.0.0
|
||||||
|
port: 8765
|
||||||
|
""");
|
||||||
|
return f;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Passes {@code cfg.validateAll()} cleanly — nothing here trips any of its checks. */
|
||||||
|
private static Path validConfig(Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 8765
|
||||||
|
""");
|
||||||
|
return f;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Pins the order against {@code cfg.validateAll()}. The environment is tainted and the config
|
||||||
|
* is otherwise invalid (see {@link #invalidConfig}). If the guard runs first (the required
|
||||||
|
* order), {@code main} throws {@link GuardException} before {@code validateAll()} is ever
|
||||||
|
* reached. If the guard were deleted, or reordered to run after {@code validateAll()}, {@code
|
||||||
|
* validateAll()} throws {@link IllegalStateException} instead and this assertion fails on the
|
||||||
|
* wrong exception type.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void mainRefusesATaintedEnvironmentBeforeValidatingTheConfig(@TempDir Path dir) throws Exception {
|
||||||
|
Path config = invalidConfig(dir);
|
||||||
|
FixedEnvPorts ports = new FixedEnvPorts(TAINTED_ENV);
|
||||||
|
|
||||||
|
GuardException ex = assertThrows(GuardException.class,
|
||||||
|
() -> Fleetd.main(new String[]{config.toString()}, ports));
|
||||||
|
assertTrue(ex.getMessage().contains("tainted"),
|
||||||
|
"expected the primary-taint message, got: " + ex.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Pins the order against {@code FleetdAssembly.assembleAndStart}. The environment is tainted
|
||||||
|
* and the config is otherwise VALID (see {@link #validConfig}), so {@code cfg.validateAll()}
|
||||||
|
* passes silently and the next thing that could possibly run is the assembly's first socket
|
||||||
|
* call. If the guard runs first (the required order), {@code main} throws {@link
|
||||||
|
* GuardException} before assembly starts. If the guard were deleted, or reordered to run after
|
||||||
|
* assembly begins touching {@code ports}, {@link FixedEnvPorts#connectHerdr} throws {@link
|
||||||
|
* UnsupportedOperationException} instead and this assertion fails on the wrong exception type.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void mainRefusesATaintedEnvironmentBeforeAssemblyTouchesAnyPort(@TempDir Path dir) throws Exception {
|
||||||
|
Path config = validConfig(dir);
|
||||||
|
FixedEnvPorts ports = new FixedEnvPorts(TAINTED_ENV);
|
||||||
|
|
||||||
|
GuardException ex = assertThrows(GuardException.class,
|
||||||
|
() -> Fleetd.main(new String[]{config.toString()}, ports));
|
||||||
|
assertTrue(ex.getMessage().contains("tainted"),
|
||||||
|
"expected the primary-taint message, got: " + ex.getMessage());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The CONTROL for the two tests above. Same otherwise-valid config, same {@link FixedEnvPorts}
|
||||||
|
* whose every method but {@code environment()} refuses to be called — but a CLEAN environment.
|
||||||
|
* Without this, a guard that always threw {@link GuardException} regardless of input (the
|
||||||
|
* opposite bug — e.g. the check inverted) would make the two tests above pass for the wrong
|
||||||
|
* reason: not because they actually drove a real taint through a real guard, but because
|
||||||
|
* anything would have thrown {@code GuardException}. Here, with nothing to taint, the guard
|
||||||
|
* must let {@code main} proceed into {@code cfg.validateAll()} and on into the real assembly,
|
||||||
|
* which reaches {@code ports.connectHerdr} — and THAT throws. A loud, positive assertion: if
|
||||||
|
* the boot path never actually ran this far, there is no {@link UnsupportedOperationException}
|
||||||
|
* to catch, only a quiet, unexpected hang or an unrelated early failure.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void mainProceedsPastTheGuardOnACleanEnvironment(@TempDir Path dir) throws Exception {
|
||||||
|
Path config = validConfig(dir);
|
||||||
|
FixedEnvPorts ports = new FixedEnvPorts(CLEAN_ENV);
|
||||||
|
|
||||||
|
UnsupportedOperationException ex = assertThrows(UnsupportedOperationException.class,
|
||||||
|
() -> Fleetd.main(new String[]{config.toString()}, ports));
|
||||||
|
assertTrue(ex.getMessage().contains("connectHerdr"),
|
||||||
|
"expected forward progress to reach the assembly's first port call, got: " + ex.getMessage());
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,252 @@
|
|||||||
|
package dev.ltms.fleet;
|
||||||
|
|
||||||
|
import ch.qos.logback.classic.Level;
|
||||||
|
import ch.qos.logback.classic.Logger;
|
||||||
|
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||||
|
import ch.qos.logback.core.read.ListAppender;
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.List;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #613: a {@code MemberRole} with no {@code fleet.<role>s:} pool falls back to
|
||||||
|
* <em>every</em> configured profile ({@code FleetConfig#candidateProfiles}), and one with no
|
||||||
|
* {@code fleet.charters.<role>:} entry runs with only the launcher's reply charter. Both are
|
||||||
|
* deliberate, legitimate states — neither is refused by {@code validateMembers()} — but both were
|
||||||
|
* silent at boot. On the host that opened this ticket, an unqualified {@code hunter} spawn silently
|
||||||
|
* widened to all 8 configured profiles and its resolved first choice was {@code local}, a profile
|
||||||
|
* every other pool on that same config gave weight 0 to.
|
||||||
|
*
|
||||||
|
* <p>{@link Fleetd#reportRoleFallbackGaps} must name every gapped role, and for a pool gap, the
|
||||||
|
* exact resolved first-choice profile — that number, not the pool size, is what actually surprised
|
||||||
|
* the operator. Mirrors {@link ExhaustedPatternGapReportTest}'s pattern: capture the real log via a
|
||||||
|
* {@link ListAppender} rather than asserting on the call site's source text.
|
||||||
|
*/
|
||||||
|
class RoleFallbackGapReportTest {
|
||||||
|
|
||||||
|
private static FleetConfig load(Path dir, String yaml) throws Exception {
|
||||||
|
Path f = dir.resolve("fleetd.yaml");
|
||||||
|
Files.writeString(f, yaml);
|
||||||
|
return FleetConfig.load(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The level this logger had before {@link #attach()} raised it, so {@link #detach} can put it
|
||||||
|
* back. {@code null} means "inherit from the parent" — the state this logger starts in.
|
||||||
|
*/
|
||||||
|
private static Level originalLevel;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* {@code reportRoleFallbackGaps} logs at INFO, and {@code logback-test.xml} sets
|
||||||
|
* {@code dev.ltms.fleet} to WARN — so INFO events are dropped by the level check before any
|
||||||
|
* appender sees them. Raise the level for the duration of the test, exactly like {@code
|
||||||
|
* GitHostShapeReportTest#attach}.
|
||||||
|
*/
|
||||||
|
private static ListAppender<ILoggingEvent> attach() {
|
||||||
|
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||||
|
originalLevel = logger.getLevel();
|
||||||
|
logger.setLevel(Level.INFO);
|
||||||
|
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||||
|
appender.start();
|
||||||
|
logger.addAppender(appender);
|
||||||
|
return appender;
|
||||||
|
}
|
||||||
|
|
||||||
|
private static void detach(ListAppender<ILoggingEvent> appender) {
|
||||||
|
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||||
|
logger.detachAppender(appender);
|
||||||
|
logger.setLevel(originalLevel);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static List<String> infoMessages(ListAppender<ILoggingEvent> appender) {
|
||||||
|
return appender.list.stream()
|
||||||
|
.filter(e -> e.getLevel() == Level.INFO)
|
||||||
|
.map(ILoggingEvent::getFormattedMessage)
|
||||||
|
.toList();
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reproduces the shape measured in the ticket: {@code dev}, {@code reviewer} and
|
||||||
|
* {@code architect} each have a pool and a charter; {@code hunter} has neither. The pool-gap
|
||||||
|
* line must name {@code hunter}, the profile count (3), and the resolved first choice
|
||||||
|
* ({@code local}, the first profile in definition order) — and must not name the three healthy
|
||||||
|
* roles. The charter-gap line must separately name only {@code hunter}.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void hunterWithNoPoolOrCharterIsNamedWithItsResolvedFirstChoice(@TempDir Path dir) throws Exception {
|
||||||
|
FleetConfig cfg = load(dir, """
|
||||||
|
profiles:
|
||||||
|
local:
|
||||||
|
baseUrl: http://gx00.gw:8000
|
||||||
|
sonnet:
|
||||||
|
baseUrl: https://llm.ltms.dev/v1
|
||||||
|
terra:
|
||||||
|
baseUrl: https://llm.ltms.dev/v1
|
||||||
|
fleet:
|
||||||
|
developers:
|
||||||
|
a:
|
||||||
|
profile: sonnet
|
||||||
|
reviewers:
|
||||||
|
b:
|
||||||
|
profile: terra
|
||||||
|
architects:
|
||||||
|
c:
|
||||||
|
profile: sonnet
|
||||||
|
charters:
|
||||||
|
dev: "dev charter text"
|
||||||
|
reviewer: "reviewer charter text"
|
||||||
|
architect: "architect charter text"
|
||||||
|
""");
|
||||||
|
|
||||||
|
ListAppender<ILoggingEvent> appender = attach();
|
||||||
|
try {
|
||||||
|
Fleetd.reportRoleFallbackGaps(cfg);
|
||||||
|
} finally {
|
||||||
|
detach(appender);
|
||||||
|
}
|
||||||
|
|
||||||
|
List<String> infos = infoMessages(appender);
|
||||||
|
|
||||||
|
String poolLine = infos.stream()
|
||||||
|
.filter(m -> m.contains("no fleet.<role>s: pool"))
|
||||||
|
.findFirst()
|
||||||
|
.orElseThrow(() -> new AssertionError("expected a pool-gap INFO line: " + infos));
|
||||||
|
assertTrue(poolLine.contains("hunter"), poolLine);
|
||||||
|
assertTrue(poolLine.contains("3"), "must name the profile count the gap opens onto: " + poolLine);
|
||||||
|
assertTrue(poolLine.contains("'local'"),
|
||||||
|
"must name the resolved first-choice profile, the number that actually surprised "
|
||||||
|
+ "the operator: " + poolLine);
|
||||||
|
for (String healthy : List.of("dev", "reviewer", "architect")) {
|
||||||
|
assertFalse(poolLine.contains(healthy),
|
||||||
|
"pool-gap line must not name a role that has a pool: " + poolLine);
|
||||||
|
}
|
||||||
|
|
||||||
|
String charterLine = infos.stream()
|
||||||
|
.filter(m -> m.contains("no fleet.charters: entry"))
|
||||||
|
.findFirst()
|
||||||
|
.orElseThrow(() -> new AssertionError("expected a charter-gap INFO line: " + infos));
|
||||||
|
assertTrue(charterLine.contains("hunter"), charterLine);
|
||||||
|
for (String healthy : List.of("dev", "reviewer", "architect")) {
|
||||||
|
assertFalse(charterLine.contains(healthy),
|
||||||
|
"charter-gap line must not name a role that has a charter: " + charterLine);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The resolved first choice must be genuinely computed from definition order, not hardcoded —
|
||||||
|
* reordering {@code profiles:} so a different entry comes first changes the reported choice.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void theResolvedFirstChoiceFollowsProfileDefinitionOrder(@TempDir Path dir) throws Exception {
|
||||||
|
FleetConfig cfg = load(dir, """
|
||||||
|
profiles:
|
||||||
|
sonnet:
|
||||||
|
baseUrl: https://llm.ltms.dev/v1
|
||||||
|
local:
|
||||||
|
baseUrl: http://gx00.gw:8000
|
||||||
|
fleet:
|
||||||
|
developers:
|
||||||
|
a:
|
||||||
|
profile: sonnet
|
||||||
|
charters:
|
||||||
|
dev: "dev charter text"
|
||||||
|
""");
|
||||||
|
|
||||||
|
ListAppender<ILoggingEvent> appender = attach();
|
||||||
|
try {
|
||||||
|
Fleetd.reportRoleFallbackGaps(cfg);
|
||||||
|
} finally {
|
||||||
|
detach(appender);
|
||||||
|
}
|
||||||
|
|
||||||
|
String poolLine = infoMessages(appender).stream()
|
||||||
|
.filter(m -> m.contains("no fleet.<role>s: pool"))
|
||||||
|
.findFirst()
|
||||||
|
.orElseThrow();
|
||||||
|
// hunter, reviewer and architect all lack a pool here; each falls back to the full 2-profile
|
||||||
|
// set and the first-choice is 'sonnet' because it is first in profiles: definition order.
|
||||||
|
assertTrue(poolLine.contains("'sonnet'"), poolLine);
|
||||||
|
assertFalse(poolLine.contains("'local'"), poolLine);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** A config with a pool and a charter for every role produces no role-fallback log at all. */
|
||||||
|
@Test
|
||||||
|
void everyRoleWithAPoolAndACharterProducesNoLogAtAll(@TempDir Path dir) throws Exception {
|
||||||
|
FleetConfig cfg = load(dir, """
|
||||||
|
profiles:
|
||||||
|
sonnet:
|
||||||
|
baseUrl: https://llm.ltms.dev/v1
|
||||||
|
fleet:
|
||||||
|
developers:
|
||||||
|
a:
|
||||||
|
profile: sonnet
|
||||||
|
reviewers:
|
||||||
|
b:
|
||||||
|
profile: sonnet
|
||||||
|
hunters:
|
||||||
|
c:
|
||||||
|
profile: sonnet
|
||||||
|
architects:
|
||||||
|
d:
|
||||||
|
profile: sonnet
|
||||||
|
charters:
|
||||||
|
dev: "dev charter text"
|
||||||
|
reviewer: "reviewer charter text"
|
||||||
|
hunter: "hunter charter text"
|
||||||
|
architect: "architect charter text"
|
||||||
|
""");
|
||||||
|
|
||||||
|
ListAppender<ILoggingEvent> appender = attach();
|
||||||
|
try {
|
||||||
|
Fleetd.reportRoleFallbackGaps(cfg);
|
||||||
|
} finally {
|
||||||
|
detach(appender);
|
||||||
|
}
|
||||||
|
|
||||||
|
assertTrue(appender.list.isEmpty(),
|
||||||
|
"a config with no gaps must not print a per-role block: " + infoMessages(appender));
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* A config with only {@code profiles:} and no {@code fleet:} block at all must still be
|
||||||
|
* reported (every role is gapped, both pool and charter) rather than throwing — this is the
|
||||||
|
* exact shape {@code candidateProfiles}' fallback exists to keep starting.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void aConfigWithNoFleetBlockAtAllReportsEveryRoleGapped(@TempDir Path dir) throws Exception {
|
||||||
|
FleetConfig cfg = load(dir, """
|
||||||
|
profiles:
|
||||||
|
sonnet:
|
||||||
|
baseUrl: https://llm.ltms.dev/v1
|
||||||
|
""");
|
||||||
|
|
||||||
|
ListAppender<ILoggingEvent> appender = attach();
|
||||||
|
try {
|
||||||
|
Fleetd.reportRoleFallbackGaps(cfg);
|
||||||
|
} finally {
|
||||||
|
detach(appender);
|
||||||
|
}
|
||||||
|
|
||||||
|
List<String> infos = infoMessages(appender);
|
||||||
|
String poolLine = infos.stream()
|
||||||
|
.filter(m -> m.contains("no fleet.<role>s: pool"))
|
||||||
|
.findFirst()
|
||||||
|
.orElseThrow(() -> new AssertionError("expected a pool-gap INFO line: " + infos));
|
||||||
|
String charterLine = infos.stream()
|
||||||
|
.filter(m -> m.contains("no fleet.charters: entry"))
|
||||||
|
.findFirst()
|
||||||
|
.orElseThrow(() -> new AssertionError("expected a charter-gap INFO line: " + infos));
|
||||||
|
for (String role : List.of("dev", "hunter", "reviewer", "architect")) {
|
||||||
|
assertTrue(poolLine.contains(role), poolLine);
|
||||||
|
assertTrue(charterLine.contains(role), charterLine);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,8 +1,11 @@
|
|||||||
package dev.ltms.fleet.config;
|
package dev.ltms.fleet.config;
|
||||||
|
|
||||||
|
import ch.qos.logback.classic.Level;
|
||||||
|
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||||
import dev.ltms.fleet.auth.MemberRegistry;
|
import dev.ltms.fleet.auth.MemberRegistry;
|
||||||
import dev.ltms.fleet.msg.LeadMailbox;
|
import dev.ltms.fleet.msg.LeadMailbox;
|
||||||
import dev.ltms.fleet.peer.MemberRole;
|
import dev.ltms.fleet.peer.MemberRole;
|
||||||
|
import dev.ltms.fleet.testing.CapturedLog;
|
||||||
import org.junit.jupiter.api.Test;
|
import org.junit.jupiter.api.Test;
|
||||||
import org.junit.jupiter.api.io.TempDir;
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
@@ -93,6 +96,72 @@ class FleetConfigTest {
|
|||||||
"unset means off — today's behaviour, unchanged");
|
"unset means off — today's behaviour, unchanged");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #601 review (measured 2026-09-22): this guard used to throw {@link
|
||||||
|
* IllegalStateException} and refuse to start. On a host with the conflict configured, that
|
||||||
|
* turned into a launchd restart loop with no readable cause, since {@code fleetd.yaml} is
|
||||||
|
* gitignored. It must now WARN and let the daemon start, and the warning must carry both values
|
||||||
|
* so an operator can fix the config without reading the source. Pinning the log line itself (via
|
||||||
|
* {@link CapturedLog}) rather than a snippet of production source text — the latter is the
|
||||||
|
* anti-pattern this repo avoids; the former is the actual observable behaviour a reader (or an
|
||||||
|
* alert on the log) depends on.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void aClaudeProfileWithConflictingAutoCompactFlagAndEnvironmentWindowLoadsAndWarnsWithBothValues(
|
||||||
|
@TempDir Path dir) throws Exception {
|
||||||
|
Path f = dir.resolve("conflicting-auto-compact-window.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
profiles:
|
||||||
|
claude-profile:
|
||||||
|
autoCompactWindow: 250000
|
||||||
|
env:
|
||||||
|
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "300000"
|
||||||
|
""");
|
||||||
|
|
||||||
|
FleetConfig cfg;
|
||||||
|
List<String> warnings;
|
||||||
|
try (CapturedLog log = CapturedLog.at(FleetConfig.class, Level.WARN)) {
|
||||||
|
cfg = FleetConfig.load(f);
|
||||||
|
warnings = log.events().stream().map(ILoggingEvent::getFormattedMessage).toList();
|
||||||
|
}
|
||||||
|
|
||||||
|
assertEquals(250_000, cfg.profiles().get("claude-profile").autoCompactWindow(),
|
||||||
|
"the disagreement is reported, not corrected — the flag value still loads as-is");
|
||||||
|
assertEquals(1, warnings.size(), "exactly one warning for the one conflicting profile: " + warnings);
|
||||||
|
String warning = warnings.get(0);
|
||||||
|
assertTrue(warning.contains("claude-profile"), "names the offending profile: " + warning);
|
||||||
|
assertTrue(warning.contains("autoCompactWindow=250000"), "names the flag value: " + warning);
|
||||||
|
assertTrue(warning.contains("CLAUDE_CODE_AUTO_COMPACT_WINDOW=300000"), "names the env value: " + warning);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The negative probe paired with the test above (per fleetd #601 review): a warning that fires
|
||||||
|
* on every load and a warning that never fires read the same from a single test, so both must be
|
||||||
|
* checked. No conflict here — the flag and the env value agree — so no warning should be logged.
|
||||||
|
*/
|
||||||
|
@Test
|
||||||
|
void aClaudeProfileWithEqualAutoCompactFlagAndEnvironmentWindowLoadsWithNoWarning(@TempDir Path dir)
|
||||||
|
throws Exception {
|
||||||
|
Path f = dir.resolve("equal-auto-compact-window.yaml");
|
||||||
|
Files.writeString(f, """
|
||||||
|
profiles:
|
||||||
|
claude-profile:
|
||||||
|
autoCompactWindow: 250000
|
||||||
|
env:
|
||||||
|
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "250000"
|
||||||
|
""");
|
||||||
|
|
||||||
|
FleetConfig cfg;
|
||||||
|
List<ILoggingEvent> events;
|
||||||
|
try (CapturedLog log = CapturedLog.at(FleetConfig.class, Level.WARN)) {
|
||||||
|
cfg = FleetConfig.load(f);
|
||||||
|
events = log.events();
|
||||||
|
}
|
||||||
|
|
||||||
|
assertEquals(250_000, cfg.profiles().get("claude-profile").autoCompactWindow());
|
||||||
|
assertTrue(events.isEmpty(), "equal values must not warn: " + events);
|
||||||
|
}
|
||||||
|
|
||||||
// ── fleetd #201 Unit 5: errorPattern ────────────────────────────────────────────────────────
|
// ── fleetd #201 Unit 5: errorPattern ────────────────────────────────────────────────────────
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -345,7 +414,7 @@ class FleetConfigTest {
|
|||||||
IllegalStateException unknownError = assertThrows(IllegalStateException.class,
|
IllegalStateException unknownError = assertThrows(IllegalStateException.class,
|
||||||
() -> FleetConfig.load(unknown).validateCharters());
|
() -> FleetConfig.load(unknown).validateCharters());
|
||||||
assertTrue(unknownError.getMessage().contains("architetc"));
|
assertTrue(unknownError.getMessage().contains("architetc"));
|
||||||
assertTrue(unknownError.getMessage().contains("[architect, dev, reviewer]"));
|
assertTrue(unknownError.getMessage().contains("[architect, dev, hunter, reviewer]"));
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -812,13 +881,17 @@ class FleetConfigTest {
|
|||||||
reviewers:
|
reviewers:
|
||||||
b:
|
b:
|
||||||
profile: sonnet
|
profile: sonnet
|
||||||
|
hunters:
|
||||||
|
c:
|
||||||
|
profile: sonnet
|
||||||
""");
|
""");
|
||||||
|
|
||||||
FleetConfig cfg = FleetConfig.load(f);
|
FleetConfig cfg = FleetConfig.load(f);
|
||||||
assertEquals(List.of("sonnet"), cfg.fleet().profilesFor(MemberRole.DEV));
|
assertEquals(List.of("sonnet"), cfg.fleet().profilesFor(MemberRole.DEV));
|
||||||
|
assertEquals(List.of("sonnet"), cfg.fleet().profilesFor(MemberRole.HUNTER));
|
||||||
assertEquals(List.of("sonnet"), cfg.fleet().profilesFor(MemberRole.REVIEWER));
|
assertEquals(List.of("sonnet"), cfg.fleet().profilesFor(MemberRole.REVIEWER));
|
||||||
assertTrue(cfg.fleet().profilesFor(MemberRole.ARCHITECT).isEmpty());
|
assertTrue(cfg.fleet().profilesFor(MemberRole.ARCHITECT).isEmpty());
|
||||||
assertEquals(List.of(MemberRole.DEV, MemberRole.REVIEWER), cfg.fleet().rolesConfigured());
|
assertEquals(List.of(MemberRole.DEV, MemberRole.HUNTER, MemberRole.REVIEWER), cfg.fleet().rolesConfigured());
|
||||||
}
|
}
|
||||||
|
|
||||||
/** The case the two axes exist for: one backend, two roles, and neither is a duplicate. */
|
/** The case the two axes exist for: one backend, two roles, and neither is a duplicate. */
|
||||||
|
|||||||
@@ -456,7 +456,11 @@ public final class FakeHerdr implements HerdrClient {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** fleetd #612 Unit A: set by {@link #close()} so a test can prove a router/client actually closed this. */
|
||||||
|
public volatile boolean closed = false;
|
||||||
|
|
||||||
@Override
|
@Override
|
||||||
public void close() {
|
public void close() {
|
||||||
|
closed = true;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,248 @@
|
|||||||
|
package dev.ltms.fleet.lead;
|
||||||
|
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.io.IOException;
|
||||||
|
import java.io.RandomAccessFile;
|
||||||
|
import java.nio.charset.StandardCharsets;
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
import static org.junit.jupiter.api.Assumptions.assumeFalse;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Ticket "lead context gauge" — fleetd could not see how full a lead's own Claude Code context
|
||||||
|
* window was, so a lead auto-compacting was always a surprise. {@link LeadContextGauge} reads that
|
||||||
|
* off the lead's own transcript. Five acceptance properties from the ticket, one test class each
|
||||||
|
* (plus the three separate UNKNOWN cases the ticket calls out by name):
|
||||||
|
* <ol>
|
||||||
|
* <li>{@link #tokenCountTracksTheLastUsageRecordAndChangesWithIt()}</li>
|
||||||
|
* <li>{@link #compactionCountTracksCompactBoundaryRecordsAndChangesWithIt()}</li>
|
||||||
|
* <li>{@link #missingFileIsUnknown()}, {@link #unreadableFileIsUnknown()}</li>
|
||||||
|
* <li>{@link #tornFinalLineFallsBackToTheLastGoodReading()},
|
||||||
|
* {@link #everyLineUnparseableIsUnknown()} (fleetd #602 gauge-wiring, finding 2)</li>
|
||||||
|
* <li>{@link #readNeverExceedsTheTailBound()}</li>
|
||||||
|
* <li>{@link #secondReadInsideTtlDoesNotTouchDiskAgain()}</li>
|
||||||
|
* </ol>
|
||||||
|
*
|
||||||
|
* <p>Every fixture lives under {@code @TempDir} — never the operator's real config directory (see
|
||||||
|
* the ticket's hard constraint on this).
|
||||||
|
*
|
||||||
|
* <p><strong>fleetd #602 gauge-wiring, finding 2.</strong> The property above used to be "a
|
||||||
|
* truncated or invalid LAST line reports UNKNOWN" — on the theory that a bad last line signals a
|
||||||
|
* format change. That reasoning did not hold: {@code fleet_list} reads this transcript while Claude
|
||||||
|
* Code may be mid-write on it, so a torn LAST line is an ordinary race, not a format change, and a
|
||||||
|
* real format change makes EVERY line unparseable, not only the last one written. So the property
|
||||||
|
* is now split in two: {@link #tornFinalLineFallsBackToTheLastGoodReading} (the torn-line case must
|
||||||
|
* NOT destroy a good earlier reading) and its control, {@link #everyLineUnparseableIsUnknown} (only
|
||||||
|
* when NOTHING in the window parses does this report UNKNOWN).
|
||||||
|
*/
|
||||||
|
class LeadContextGaugeTest {
|
||||||
|
|
||||||
|
private static final String SESSION_ID = "11111111-1111-1111-1111-111111111111";
|
||||||
|
|
||||||
|
/** One line Claude Code would write for a turn with the given live-context total. */
|
||||||
|
private static String usageLine(long inputTokens, long cacheRead, long cacheCreation) {
|
||||||
|
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
|
||||||
|
+ "\"input_tokens\":" + inputTokens + ","
|
||||||
|
+ "\"cache_read_input_tokens\":" + cacheRead + ","
|
||||||
|
+ "\"cache_creation_input_tokens\":" + cacheCreation + "}}}";
|
||||||
|
}
|
||||||
|
|
||||||
|
/** One line Claude Code writes when an auto-compaction happens. */
|
||||||
|
private static String compactionLine() {
|
||||||
|
return "{\"type\":\"system\",\"subtype\":\"compact_boundary\","
|
||||||
|
+ "\"compactMetadata\":{\"preTokens\":250000,\"postTokens\":5000,"
|
||||||
|
+ "\"trigger\":\"auto\",\"durationMs\":54000}}";
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} and writes {@code lines}. */
|
||||||
|
private static String writeTranscript(Path configDir, String sessionId, String... lines) throws IOException {
|
||||||
|
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
|
||||||
|
Files.createDirectories(projectDir);
|
||||||
|
Path file = projectDir.resolve(sessionId + ".jsonl");
|
||||||
|
StringBuilder sb = new StringBuilder();
|
||||||
|
for (String line : lines) {
|
||||||
|
sb.append(line).append('\n');
|
||||||
|
}
|
||||||
|
Files.writeString(file, sb.toString(), StandardCharsets.UTF_8);
|
||||||
|
return configDir.toString();
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- property 1: live token count tracks the last usage record -----------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("reported tokens equal the LAST usage record's total, and change when it does")
|
||||||
|
void tokenCountTracksTheLastUsageRecordAndChangesWithIt(@TempDir Path tmp) throws IOException {
|
||||||
|
String configDir = writeTranscript(tmp, SESSION_ID,
|
||||||
|
usageLine(1_000, 0, 0),
|
||||||
|
usageLine(40_000, 5_000, 3_000)); // last record: 48,000
|
||||||
|
LeadContextGauge gauge = new LeadContextGauge();
|
||||||
|
|
||||||
|
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude");
|
||||||
|
assertEquals(48_000L, first.tokens(), "must total input+cache_read+cache_creation of the LAST usage record");
|
||||||
|
assertEquals(LeadContextGauge.State.OK, first.state());
|
||||||
|
|
||||||
|
// Change N in the fixture (a fresh session id avoids the cache) -- the reported number
|
||||||
|
// must change with it, not stay pinned to the first fixture's total.
|
||||||
|
String otherSession = "22222222-2222-2222-2222-222222222222";
|
||||||
|
writeTranscript(tmp, otherSession, usageLine(100_000, 50_000, 50_000)); // last record: 200,000
|
||||||
|
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude");
|
||||||
|
assertEquals(200_000L, second.tokens());
|
||||||
|
assertTrue(second.tokens() != first.tokens(), "changing N in the fixture must change the reported number");
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- property 2: compaction count tracks compact_boundary records --------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("reported compaction count equals K compact_boundary records, and changes with K")
|
||||||
|
void compactionCountTracksCompactBoundaryRecordsAndChangesWithIt(@TempDir Path tmp) throws IOException {
|
||||||
|
String sessionTwoCompactions = "33333333-3333-3333-3333-333333333333";
|
||||||
|
writeTranscript(tmp, sessionTwoCompactions,
|
||||||
|
usageLine(1_000, 0, 0),
|
||||||
|
compactionLine(),
|
||||||
|
usageLine(2_000, 0, 0),
|
||||||
|
compactionLine(),
|
||||||
|
usageLine(3_000, 0, 0));
|
||||||
|
LeadContextGauge gauge = new LeadContextGauge();
|
||||||
|
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude");
|
||||||
|
assertEquals(2, twoCompactions.compactions());
|
||||||
|
|
||||||
|
String sessionZeroCompactions = "44444444-4444-4444-4444-444444444444";
|
||||||
|
writeTranscript(tmp, sessionZeroCompactions, usageLine(3_000, 0, 0));
|
||||||
|
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude");
|
||||||
|
assertEquals(0, zeroCompactions.compactions(), "changing K in the fixture must change the reported count");
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- property 3: two separate UNKNOWN cases -------------------------------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("a missing transcript file reports UNKNOWN with no token number")
|
||||||
|
void missingFileIsUnknown(@TempDir Path tmp) {
|
||||||
|
LeadContextGauge gauge = new LeadContextGauge();
|
||||||
|
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
|
||||||
|
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
|
||||||
|
assertNull(reading.tokens());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("an unreadable transcript file reports UNKNOWN with no token number")
|
||||||
|
void unreadableFileIsUnknown(@TempDir Path tmp) throws IOException {
|
||||||
|
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
|
||||||
|
Path file = tmp.resolve("projects").resolve("some-project-slug").resolve(SESSION_ID + ".jsonl");
|
||||||
|
assertTrue(file.toFile().setReadable(false), "test setup: must be able to revoke read permission");
|
||||||
|
try {
|
||||||
|
// setReadable(false) really did clear the read bit (asserted above), but that alone
|
||||||
|
// does not prove the file is UNREADABLE: running as root (e.g. a CI container) ignores
|
||||||
|
// the read bit and opens the file anyway. Files.isReadable checks what actually happens
|
||||||
|
// on open, not the bit. When it still reports readable, this test cannot create the
|
||||||
|
// condition it needs on this machine, so it skips honestly instead of asserting on a
|
||||||
|
// state that was never reached. A skip here means "I could not set up the case", NOT
|
||||||
|
// "the UNKNOWN behaviour is fine" -- it is not evidence either way.
|
||||||
|
assumeFalse(Files.isReadable(file),
|
||||||
|
"runs as root (CI container): the read bit does not stop root, so this case cannot be set up here");
|
||||||
|
LeadContextGauge gauge = new LeadContextGauge();
|
||||||
|
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
|
||||||
|
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
|
||||||
|
assertNull(reading.tokens());
|
||||||
|
} finally {
|
||||||
|
file.toFile().setReadable(true); // so @TempDir cleanup can delete it
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("fleetd #602 finding 2: a torn final line does not destroy a good earlier reading")
|
||||||
|
void tornFinalLineFallsBackToTheLastGoodReading(@TempDir Path tmp) throws IOException {
|
||||||
|
// Deliberately NOT via writeTranscript: that helper appends a trailing newline after every
|
||||||
|
// line, including the last one, which would not model what a write caught mid-flush looks
|
||||||
|
// like on disk. Claude Code appends and flushes one line at a time, so the torn line here has
|
||||||
|
// no trailing newline at all -- exactly the shape a read racing an in-progress write sees.
|
||||||
|
Path projectDir = tmp.resolve("projects").resolve("some-project-slug");
|
||||||
|
Files.createDirectories(projectDir);
|
||||||
|
Path file = projectDir.resolve(SESSION_ID + ".jsonl");
|
||||||
|
String lastCompleteLine = usageLine(1_000, 2_000, 3_000); // last COMPLETE record: 6,000
|
||||||
|
String tornLine = "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{\"input_tok";
|
||||||
|
Files.writeString(file, lastCompleteLine + "\n" + tornLine, StandardCharsets.UTF_8);
|
||||||
|
|
||||||
|
LeadContextGauge gauge = new LeadContextGauge();
|
||||||
|
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
|
||||||
|
assertEquals(LeadContextGauge.State.OK, reading.state(),
|
||||||
|
"a torn final line must not turn a good earlier reading into UNKNOWN");
|
||||||
|
assertEquals(6_000L, reading.tokens(),
|
||||||
|
"must report the last COMPLETE line's total, not fail the whole read");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("fleetd #602 finding 2 control: when EVERY line is unparseable, the reading IS UNKNOWN")
|
||||||
|
void everyLineUnparseableIsUnknown(@TempDir Path tmp) throws IOException {
|
||||||
|
// The real signal a format change gives: not one bad line, but ALL of them. Without this
|
||||||
|
// control, code that never reports UNKNOWN at all would still pass the torn-line property.
|
||||||
|
String configDir = writeTranscript(tmp, SESSION_ID,
|
||||||
|
"{this is not json at all",
|
||||||
|
"neither is this{{{");
|
||||||
|
LeadContextGauge gauge = new LeadContextGauge();
|
||||||
|
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
|
||||||
|
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
|
||||||
|
"every line unparseable is the real format-change signal and must still report UNKNOWN");
|
||||||
|
assertNull(reading.tokens());
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- property 4: the tail bound is actually enforced ----------------------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("reading a transcript far larger than the tail bound reads no more than the bound")
|
||||||
|
void readNeverExceedsTheTailBound(@TempDir Path tmp) throws IOException {
|
||||||
|
int smallBound = 4_096; // exercise the mechanism without writing a multi-MB fixture
|
||||||
|
Path file = tmp.resolve("big.jsonl");
|
||||||
|
// One line far bigger than smallBound, repeated, so the whole file is many times the bound.
|
||||||
|
String line = usageLine(42, 0, 0) + " ".repeat(500);
|
||||||
|
StringBuilder sb = new StringBuilder();
|
||||||
|
for (int i = 0; i < 50; i++) {
|
||||||
|
sb.append(line).append('\n');
|
||||||
|
}
|
||||||
|
Files.writeString(file, sb.toString(), StandardCharsets.UTF_8);
|
||||||
|
long fileLength = Files.size(file);
|
||||||
|
assertTrue(fileLength > (long) smallBound * 5, "test setup: fixture must genuinely dwarf the bound");
|
||||||
|
|
||||||
|
var tail = LeadContextGauge.tailBytes(file, smallBound);
|
||||||
|
assertEquals(smallBound, tail.bytes().length,
|
||||||
|
"must read exactly the bound, not the whole " + fileLength + "-byte file — assert on bytes actually read");
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- property 5: the cache TTL means a burst of calls reads the file once -------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("the same (configDir, sessionId) read twice inside the TTL touches disk once")
|
||||||
|
void secondReadInsideTtlDoesNotTouchDiskAgain(@TempDir Path tmp) throws IOException {
|
||||||
|
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
|
||||||
|
AtomicLong now = new AtomicLong(0);
|
||||||
|
LeadContextGauge gauge = new LeadContextGauge(now::get, 5_000);
|
||||||
|
|
||||||
|
gauge.read(configDir, SESSION_ID, "claude");
|
||||||
|
gauge.read(configDir, SESSION_ID, "claude"); // still inside the TTL window
|
||||||
|
assertEquals(1, gauge.diskReadCount(), "two reads inside the TTL must touch disk once");
|
||||||
|
|
||||||
|
now.set(6_000); // past the TTL
|
||||||
|
gauge.read(configDir, SESSION_ID, "claude");
|
||||||
|
assertEquals(2, gauge.diskReadCount(), "a read past the TTL must touch disk again");
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- non-Claude peer / unresolved session -----------------------------------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("a non-claude agent type or a null session id is UNKNOWN, never OK")
|
||||||
|
void nonClaudePeerOrUnresolvedSessionIsUnknown(@TempDir Path tmp) throws IOException {
|
||||||
|
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
|
||||||
|
LeadContextGauge gauge = new LeadContextGauge();
|
||||||
|
|
||||||
|
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode").state());
|
||||||
|
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null).state());
|
||||||
|
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude").state());
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -32,13 +32,26 @@ class LeadLauncherTest {
|
|||||||
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true, null);
|
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true, null);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private static FleetConfig.Profile profileWithAutoCompactWindow(String kind) {
|
||||||
|
return new FleetConfig.Profile(
|
||||||
|
"opus", null, "claude-opus-5", null, "FLEETD_WORKER_TOKEN",
|
||||||
|
List.of("ccs", "ltms"), "tab", "fleet", null,
|
||||||
|
"http://127.0.0.1:8765/mcp", null, null,
|
||||||
|
null, null, kind, Map.of(), null, null, true, null, null,
|
||||||
|
null, null, null, 250_000);
|
||||||
|
}
|
||||||
|
|
||||||
private static FleetConfig configWith(FleetConfig.Leader lead) {
|
private static FleetConfig configWith(FleetConfig.Leader lead) {
|
||||||
|
return configWith(lead, opusProfile());
|
||||||
|
}
|
||||||
|
|
||||||
|
private static FleetConfig configWith(FleetConfig.Leader lead, FleetConfig.Profile profile) {
|
||||||
Map<String, FleetConfig.Leader> leaders = new LinkedHashMap<>();
|
Map<String, FleetConfig.Leader> leaders = new LinkedHashMap<>();
|
||||||
leaders.put("opus", lead);
|
leaders.put("opus", lead);
|
||||||
FleetConfig.Fleet fleet =
|
FleetConfig.Fleet fleet =
|
||||||
new FleetConfig.Fleet(leaders, Map.of(), Map.of(), Map.of(), null);
|
new FleetConfig.Fleet(leaders, Map.of(), Map.of(), Map.of(), null);
|
||||||
return new FleetConfig(
|
return new FleetConfig(
|
||||||
null, null, Map.of("opus", opusProfile()), null, null, null, null, null,
|
null, null, Map.of("opus", profile), null, null, null, null, null,
|
||||||
null, null, fleet, null, "fixed", null).withDefaults();
|
null, null, fleet, null, "fixed", null).withDefaults();
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -330,6 +343,36 @@ class LeadLauncherTest {
|
|||||||
"--model is appended last so it outranks the ccs wrapper (CB-533)");
|
"--model is appended last so it outranks the ccs wrapper (CB-533)");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aClaudeLeadPassesItsConfiguredAutoCompactWindowToClaudeCode() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
FleetConfig.Profile profile = profileWithAutoCompactWindow("claude-code");
|
||||||
|
|
||||||
|
launcher(herdr, configWith(lead("opus", "lead: opus", 1), profile)).ensureLeads();
|
||||||
|
|
||||||
|
List<String> args = startedArgs(herdr);
|
||||||
|
assertEquals("250000", args.get(args.indexOf("--autocompact") + 1));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aClaudeLeadWithNoAutoCompactWindowGetsNoAutoCompactFlag() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
|
||||||
|
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
|
||||||
|
|
||||||
|
assertFalse(startedArgs(herdr).contains("--autocompact"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aNonClaudeLeadDoesNotGetAnAutoCompactFlag() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
FleetConfig.Profile profile = profileWithAutoCompactWindow("opencode");
|
||||||
|
|
||||||
|
launcher(herdr, configWith(lead("opus", "lead: opus", 1), profile)).ensureLeads();
|
||||||
|
|
||||||
|
assertFalse(startedArgs(herdr).contains("--autocompact"));
|
||||||
|
}
|
||||||
|
|
||||||
/** A lead runs on the operator's subscription. Nothing may move it off. */
|
/** A lead runs on the operator's subscription. Nothing may move it off. */
|
||||||
@Test
|
@Test
|
||||||
void theLeadEnvCarriesNoAnthropicBinding() {
|
void theLeadEnvCarriesNoAnthropicBinding() {
|
||||||
|
|||||||
@@ -16,6 +16,7 @@ import org.junit.jupiter.api.io.TempDir;
|
|||||||
import java.io.IOException;
|
import java.io.IOException;
|
||||||
import java.nio.file.Files;
|
import java.nio.file.Files;
|
||||||
import java.nio.file.Path;
|
import java.nio.file.Path;
|
||||||
|
import java.util.ArrayList;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.concurrent.atomic.AtomicLong;
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
import java.util.function.Function;
|
import java.util.function.Function;
|
||||||
@@ -1004,6 +1005,324 @@ class LeadRolloverTest {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ---- CB-... : LeadRollover#status makes the outcome of a confirmed roll readable -----------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[STATUS 1] after the calling turn never settles, status() reports "
|
||||||
|
+ "TURN_NEVER_SETTLED for that token — this assertion could not even be written before "
|
||||||
|
+ "status() existed")
|
||||||
|
void statusReportsTurnNeverSettledAfterTheRollIsAbandoned() throws IOException {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
herdr.agentStatus("working"); // the calling lead's own pane — never goes idle in this test
|
||||||
|
Path handover = writeHandover("handover contents");
|
||||||
|
FleetConfig.LeadRollover config =
|
||||||
|
new FleetConfig.LeadRollover(handover.toString(), true, 3600, 1 /*turnSettleSeconds*/, 20, "text");
|
||||||
|
AtomicLong clock = new AtomicLong(1_000);
|
||||||
|
LeadRollover rollover = newRollover(herdr, config, () -> clock.addAndGet(500));
|
||||||
|
|
||||||
|
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||||
|
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||||
|
assertTrue(decision.accepted(), "every synchronous gate should pass; the refusal happens "
|
||||||
|
+ "only inside the deferred continuation, which this test's synchronous runner has "
|
||||||
|
+ "already run to completion by the time confirm() returns");
|
||||||
|
assertEquals(0, promptCallCount(herdr), "sanity: /clear was never sent");
|
||||||
|
|
||||||
|
LeadRollover.RollStatus status = rollover.status(pending.token());
|
||||||
|
assertEquals(LeadRollover.RollState.TURN_NEVER_SETTLED, status.state());
|
||||||
|
assertTrue(status.detail().contains("turnSettleSeconds"), "the detail must name the knob a "
|
||||||
|
+ "reader needs to raise: " + status.detail());
|
||||||
|
assertTrue(status.detail().toLowerCase().contains("no /clear"), "the detail must say plainly "
|
||||||
|
+ "that no /clear was ever sent: " + status.detail());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[STATUS 2] a roll that completes reports ROLLED for its token")
|
||||||
|
void statusReportsRolledForACompletedRoll() throws IOException {
|
||||||
|
FakeHerdr herdr = new FakeHerdr(); // default idle — a full successful roll
|
||||||
|
Path handover = writeHandover("handover contents");
|
||||||
|
AtomicLong clock = new AtomicLong(1_000);
|
||||||
|
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
|
||||||
|
|
||||||
|
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||||
|
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||||
|
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||||
|
assertEquals(2, promptCallCount(herdr), "sanity: /clear then bootstrapText were both sent");
|
||||||
|
|
||||||
|
LeadRollover.RollStatus status = rollover.status(pending.token());
|
||||||
|
assertEquals(LeadRollover.RollState.ROLLED, status.state());
|
||||||
|
assertNotNull(status.detail());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[STATUS 3] a roll where /clear never settles reports CLEAR_NEVER_SETTLED — "
|
||||||
|
+ "distinct from both ROLLED and TURN_NEVER_SETTLED")
|
||||||
|
void statusReportsClearNeverSettledDistinctFromTheOtherTwoStates() throws IOException {
|
||||||
|
// Idle until /clear is sent, then permanently working — the SECOND wait never settles.
|
||||||
|
FakeHerdr fake = new FakeHerdr();
|
||||||
|
HerdrClient flipsAfterClear = new HerdrClient() {
|
||||||
|
@Override
|
||||||
|
public JsonNode call(String method, Object params) throws HerdrException {
|
||||||
|
JsonNode result = fake.call(method, params);
|
||||||
|
if ("agent.prompt".equals(method) && String.valueOf(params).contains("/clear")) {
|
||||||
|
fake.agentStatus("working");
|
||||||
|
}
|
||||||
|
return result;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
fake.close();
|
||||||
|
}
|
||||||
|
};
|
||||||
|
Path handover = writeHandover("handover contents");
|
||||||
|
FleetConfig.LeadRollover config =
|
||||||
|
new FleetConfig.LeadRollover(handover.toString(), true, 3600, 20, 1 /*clearSettleSeconds*/, "boot text");
|
||||||
|
AtomicLong clock = new AtomicLong(1_000);
|
||||||
|
LeadRollover rollover = newRollover(flipsAfterClear, config, () -> clock.addAndGet(500));
|
||||||
|
|
||||||
|
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||||
|
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||||
|
assertTrue(decision.accepted(), "every synchronous gate passes; the refusal is logged only, "
|
||||||
|
+ "deep inside the deferred continuation");
|
||||||
|
assertEquals(1, promptCallCount(fake), "sanity: only /clear was sent, never bootstrapText");
|
||||||
|
|
||||||
|
LeadRollover.RollStatus status = rollover.status(pending.token());
|
||||||
|
assertEquals(LeadRollover.RollState.CLEAR_NEVER_SETTLED, status.state());
|
||||||
|
assertNotEquals(LeadRollover.RollState.ROLLED, status.state());
|
||||||
|
assertNotEquals(LeadRollover.RollState.TURN_NEVER_SETTLED, status.state());
|
||||||
|
assertTrue(status.detail().contains("clearSettleSeconds"), status.detail());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[STATUS 4] a token that was never issued, or was cancelled, gives a clean "
|
||||||
|
+ "UNKNOWN answer rather than an exception or a false ROLLED")
|
||||||
|
void statusOnUnissuedOrCancelledTokenIsCleanNotAnExceptionOrFalseSuccess() throws IOException {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Path handover = writeHandover("handover contents");
|
||||||
|
AtomicLong clock = new AtomicLong(1_000);
|
||||||
|
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
|
||||||
|
|
||||||
|
// never issued at all
|
||||||
|
LeadRollover.RollStatus neverIssued = assertDoesNotThrow(() -> rollover.status("no-such-token"));
|
||||||
|
assertEquals(LeadRollover.RollState.UNKNOWN, neverIssued.state());
|
||||||
|
assertNotEquals(LeadRollover.RollState.ROLLED, neverIssued.state());
|
||||||
|
|
||||||
|
// null / blank must not throw either — pending is a ConcurrentHashMap, which throws on a
|
||||||
|
// null-key lookup unless status() guards it first
|
||||||
|
assertEquals(LeadRollover.RollState.UNKNOWN, assertDoesNotThrow(() -> rollover.status(null)).state());
|
||||||
|
assertEquals(LeadRollover.RollState.UNKNOWN, assertDoesNotThrow(() -> rollover.status(" ")).state());
|
||||||
|
|
||||||
|
// opened, then cancelled
|
||||||
|
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||||
|
assertTrue(rollover.cancel(pending.token()));
|
||||||
|
LeadRollover.RollStatus cancelled = assertDoesNotThrow(() -> rollover.status(pending.token()));
|
||||||
|
assertEquals(LeadRollover.RollState.UNKNOWN, cancelled.state());
|
||||||
|
assertNotEquals(LeadRollover.RollState.ROLLED, cancelled.state());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[STATUS 5] the bounded outcome history never grows past its cap")
|
||||||
|
void statusHistoryDoesNotGrowPastItsCap() throws IOException {
|
||||||
|
FakeHerdr herdr = new FakeHerdr(); // default idle — every roll completes
|
||||||
|
Path handover = writeHandover("handover contents"); // one file, reused by every roll below —
|
||||||
|
// checkHandover only compares its mtime against each open()'s OWN requestedAtMillis (the
|
||||||
|
// fake clock, in the low thousands), and the file's real (wall-clock) mtime is always far
|
||||||
|
// larger than that, so freshness passes on every iteration without rewriting the file.
|
||||||
|
AtomicLong clock = new AtomicLong(1_000);
|
||||||
|
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), () -> clock.addAndGet(1));
|
||||||
|
|
||||||
|
int rolls = LeadRollover.OUTCOME_HISTORY_CAP + 50;
|
||||||
|
String[] tokens = new String[rolls];
|
||||||
|
for (int i = 0; i < rolls; i++) {
|
||||||
|
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full #" + i);
|
||||||
|
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||||
|
assertTrue(decision.accepted(), "roll #" + i + " should have been approved: " + decision.detail());
|
||||||
|
tokens[i] = pending.token();
|
||||||
|
}
|
||||||
|
|
||||||
|
assertEquals(LeadRollover.RollState.ROLLED, rollover.status(tokens[rolls - 1]).state(),
|
||||||
|
"the most recently finished roll's outcome must still be in the bounded history");
|
||||||
|
|
||||||
|
// There is no direct size accessor for the bounded history, so boundedness is asserted
|
||||||
|
// indirectly and behaviourally: the OLDEST finished roll's outcome must have been evicted
|
||||||
|
// (reads back as a clean UNKNOWN, exactly like a token that was never issued) once more than
|
||||||
|
// OUTCOME_HISTORY_CAP rolls have gone through this instance. If the cap were not enforced,
|
||||||
|
// tokens[0] would still read back ROLLED here, and this assertion would fail.
|
||||||
|
assertEquals(LeadRollover.RollState.UNKNOWN, rollover.status(tokens[0]).state(),
|
||||||
|
"the oldest finished roll's outcome must have been evicted once the cap was "
|
||||||
|
+ "exceeded — otherwise the bounded history is not actually bounded");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[STATUS 6] a status() call against a pending roll is read-only — no herdr call is "
|
||||||
|
+ "made and the pending request is left untouched")
|
||||||
|
void statusCallAgainstAPendingRollIsReadOnlyAndTouchesNothing() throws IOException {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Path handover = writeHandover("handover contents");
|
||||||
|
AtomicLong clock = new AtomicLong(1_000);
|
||||||
|
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
|
||||||
|
|
||||||
|
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||||
|
|
||||||
|
LeadRollover.RollStatus status = rollover.status(pending.token());
|
||||||
|
assertEquals(LeadRollover.RollState.PENDING, status.state());
|
||||||
|
assertEquals(0, herdr.calls.size(), "status() must never make any herdr call at all — not "
|
||||||
|
+ "just no agent.prompt — since it must never schedule, cancel, or retry anything");
|
||||||
|
|
||||||
|
// the pending request must be left exactly as it was: the SAME token can still be confirmed
|
||||||
|
// afterwards, as if status() had never been called.
|
||||||
|
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||||
|
assertTrue(decision.accepted(), "status() must not have consumed or otherwise disturbed the "
|
||||||
|
+ "pending request: " + decision.reason() + " / " + decision.detail());
|
||||||
|
}
|
||||||
|
|
||||||
|
// ---- PR #600 review round 2: IN_PROGRESS — the gap between confirm() handing off and the ----
|
||||||
|
// ---- continuation finishing must never read back as UNKNOWN ("nothing was ever requested") --
|
||||||
|
|
||||||
|
/**
|
||||||
|
* A {@code continuationRunner} that CAPTURES the roll instead of running it, so a test can
|
||||||
|
* observe {@link LeadRollover#status} in the window between {@link LeadRollover#confirm}
|
||||||
|
* handing off and the roll actually finishing — the window the synchronous {@code
|
||||||
|
* Runnable::run} runner used everywhere else in this class collapses to nothing. Call {@link
|
||||||
|
* #runNext()} to finish exactly one held roll, once the test is done observing the in-flight
|
||||||
|
* state.
|
||||||
|
*/
|
||||||
|
private static final class HoldingRunner implements java.util.function.Consumer<Runnable> {
|
||||||
|
private final List<Runnable> held = new ArrayList<>();
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void accept(Runnable runnable) {
|
||||||
|
held.add(runnable);
|
||||||
|
}
|
||||||
|
|
||||||
|
int heldCount() {
|
||||||
|
return held.size();
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Runs (and removes) the oldest held roll — FIFO, matching confirm() call order. */
|
||||||
|
void runNext() {
|
||||||
|
held.remove(0).run();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static LeadRollover newRolloverWithHoldingRunner(HerdrClient herdr,
|
||||||
|
FleetConfig.LeadRollover config, LongSupplier nowMillis, HoldingRunner runner) {
|
||||||
|
AgentControl agents = new AgentControl(herdr);
|
||||||
|
return new LeadRollover(agents, () -> config, _ -> null, nowMillis, () -> { }, runner);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[IN-PROGRESS 1] between an approved confirm() and the continuation finishing, "
|
||||||
|
+ "status() reports IN_PROGRESS — not UNKNOWN, and not PENDING")
|
||||||
|
void statusReportsInProgressBetweenConfirmAndTheContinuationFinishing() throws IOException {
|
||||||
|
FakeHerdr herdr = new FakeHerdr(); // default idle — the held roll WOULD complete once run
|
||||||
|
Path handover = writeHandover("handover contents");
|
||||||
|
AtomicLong clock = new AtomicLong(1_000);
|
||||||
|
HoldingRunner runner = new HoldingRunner();
|
||||||
|
LeadRollover rollover = newRolloverWithHoldingRunner(herdr, cfg(handover.toString()),
|
||||||
|
fixedClock(clock), runner);
|
||||||
|
|
||||||
|
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||||
|
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||||
|
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||||
|
assertEquals(1, runner.heldCount(), "sanity: the roll must have been handed to the "
|
||||||
|
+ "continuation runner and held there, not run yet");
|
||||||
|
assertEquals(0, promptCallCount(herdr), "sanity: the held continuation has not run, so "
|
||||||
|
+ "nothing has been sent to the pane yet");
|
||||||
|
|
||||||
|
LeadRollover.RollStatus status = rollover.status(pending.token());
|
||||||
|
assertEquals(LeadRollover.RollState.IN_PROGRESS, status.state(),
|
||||||
|
"a lead calling status() right after confirm() returned approved, while the roll "
|
||||||
|
+ "is still running, must be told IN_PROGRESS — not UNKNOWN (\"nothing was "
|
||||||
|
+ "ever requested\", which would wrongly invite it to call open() again "
|
||||||
|
+ "mid-roll) and not PENDING (\"not yet approved\", which is simply false "
|
||||||
|
+ "here): got " + status.state() + " / " + status.detail());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[IN-PROGRESS 2] once the continuation finishes, the same token reports its "
|
||||||
|
+ "terminal state — IN_PROGRESS is not sticky")
|
||||||
|
void inProgressStateIsNotStickyOnceTheContinuationFinishes() throws IOException {
|
||||||
|
FakeHerdr herdr = new FakeHerdr(); // default idle — the held roll completes once run
|
||||||
|
Path handover = writeHandover("handover contents");
|
||||||
|
AtomicLong clock = new AtomicLong(1_000);
|
||||||
|
HoldingRunner runner = new HoldingRunner();
|
||||||
|
LeadRollover rollover = newRolloverWithHoldingRunner(herdr, cfg(handover.toString()),
|
||||||
|
fixedClock(clock), runner);
|
||||||
|
|
||||||
|
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||||
|
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||||
|
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||||
|
assertEquals(LeadRollover.RollState.IN_PROGRESS, rollover.status(pending.token()).state(),
|
||||||
|
"sanity: must be IN_PROGRESS before the held continuation is run");
|
||||||
|
|
||||||
|
runner.runNext(); // finish the held roll now
|
||||||
|
|
||||||
|
assertEquals(2, promptCallCount(herdr), "sanity: the roll actually ran to completion "
|
||||||
|
+ "once released — /clear then bootstrapText");
|
||||||
|
LeadRollover.RollStatus after = rollover.status(pending.token());
|
||||||
|
assertEquals(LeadRollover.RollState.ROLLED, after.state(),
|
||||||
|
"the SAME token must now report its terminal state — IN_PROGRESS must not still be "
|
||||||
|
+ "reported once the roll has actually finished");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[IN-PROGRESS 4] IN_PROGRESS is distinct from every other RollState")
|
||||||
|
void inProgressStateIsDistinctFromAllOtherStates() throws IOException {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Path handover = writeHandover("handover contents");
|
||||||
|
AtomicLong clock = new AtomicLong(1_000);
|
||||||
|
HoldingRunner runner = new HoldingRunner();
|
||||||
|
LeadRollover rollover = newRolloverWithHoldingRunner(herdr, cfg(handover.toString()),
|
||||||
|
fixedClock(clock), runner);
|
||||||
|
|
||||||
|
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||||
|
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||||
|
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||||
|
|
||||||
|
LeadRollover.RollState inProgress = rollover.status(pending.token()).state();
|
||||||
|
assertEquals(LeadRollover.RollState.IN_PROGRESS, inProgress);
|
||||||
|
for (LeadRollover.RollState other : LeadRollover.RollState.values()) {
|
||||||
|
if (other == LeadRollover.RollState.IN_PROGRESS) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
assertNotEquals(other, inProgress, "IN_PROGRESS must be distinct from " + other);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[IN-PROGRESS 5] eviction counts IN_PROGRESS entries toward the cap exactly like "
|
||||||
|
+ "finished ones — a burst of confirmed-but-not-yet-finished rolls still ages out")
|
||||||
|
void evictionCountsInProgressEntriesTowardTheCap() throws IOException {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Path handover = writeHandover("handover contents"); // reused by every roll — see
|
||||||
|
// statusHistoryDoesNotGrowPastItsCap for why one shared file is enough for freshness.
|
||||||
|
AtomicLong clock = new AtomicLong(1_000);
|
||||||
|
HoldingRunner runner = new HoldingRunner(); // nothing run below — every roll stays IN_PROGRESS
|
||||||
|
LeadRollover rollover = newRolloverWithHoldingRunner(herdr, cfg(handover.toString()),
|
||||||
|
() -> clock.addAndGet(1), runner);
|
||||||
|
|
||||||
|
int rolls = LeadRollover.OUTCOME_HISTORY_CAP + 50;
|
||||||
|
String[] tokens = new String[rolls];
|
||||||
|
for (int i = 0; i < rolls; i++) {
|
||||||
|
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full #" + i);
|
||||||
|
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||||
|
assertTrue(decision.accepted(), "roll #" + i + " should have been approved: " + decision.detail());
|
||||||
|
tokens[i] = pending.token();
|
||||||
|
}
|
||||||
|
assertEquals(rolls, runner.heldCount(), "sanity: none of these rolls have been run — every "
|
||||||
|
+ "one of them is sitting in outcomes as IN_PROGRESS, not a separate uncapped map");
|
||||||
|
|
||||||
|
assertEquals(LeadRollover.RollState.IN_PROGRESS, rollover.status(tokens[rolls - 1]).state(),
|
||||||
|
"the most recently confirmed (still in-flight) roll must still be in the bounded "
|
||||||
|
+ "history");
|
||||||
|
assertEquals(LeadRollover.RollState.UNKNOWN, rollover.status(tokens[0]).state(),
|
||||||
|
"the oldest confirmed roll's IN_PROGRESS entry must have been evicted once the cap "
|
||||||
|
+ "was exceeded, exactly like a finished entry would be — proving IN_PROGRESS "
|
||||||
|
+ "entries share the SAME bounded map and count against the SAME cap, rather "
|
||||||
|
+ "than living in a second, uncapped in-flight map");
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@DisplayName("[fleetd #494 follow-up] the turn-settle timeout warn line prints the MEASURED "
|
@DisplayName("[fleetd #494 follow-up] the turn-settle timeout warn line prints the MEASURED "
|
||||||
+ "elapsed time next to the configured budget, never the configured value alone")
|
+ "elapsed time next to the configured budget, never the configured value alone")
|
||||||
@@ -1040,4 +1359,92 @@ class LeadRolloverTest {
|
|||||||
+ "it were the measured wait duration: " + message);
|
+ "it were the measured wait duration: " + message);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ---- fleetd #615: a HerdrException out of either unwrapped agents.send call must leave a ----
|
||||||
|
// ---- TERMINAL FAILED outcome, never a stuck IN_PROGRESS ---------------------------------------
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[fleetd #615 — 1] send() throwing on the /clear call leaves status(token) "
|
||||||
|
+ "reporting FAILED, not stuck at IN_PROGRESS")
|
||||||
|
void sendThrowingOnClearLeavesStatusReportingFailed() throws IOException {
|
||||||
|
FakeHerdr fake = new FakeHerdr(); // default idle — the turn-settle wait passes immediately
|
||||||
|
HerdrClient throwsOnClear = new HerdrClient() {
|
||||||
|
@Override
|
||||||
|
public JsonNode call(String method, Object params) throws HerdrException {
|
||||||
|
if ("agent.prompt".equals(method) && String.valueOf(params).contains("/clear")) {
|
||||||
|
throw new HerdrException("simulated herdr transport failure sending /clear");
|
||||||
|
}
|
||||||
|
return fake.call(method, params);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
fake.close();
|
||||||
|
}
|
||||||
|
};
|
||||||
|
Path handover = writeHandover("handover contents");
|
||||||
|
AtomicLong clock = new AtomicLong(1_000);
|
||||||
|
LeadRollover rollover = newRollover(throwsOnClear, cfg(handover.toString()), fixedClock(clock));
|
||||||
|
|
||||||
|
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||||
|
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||||
|
|
||||||
|
assertTrue(decision.accepted(), "every synchronous gate passes; the throw happens only "
|
||||||
|
+ "inside the deferred continuation, which this test's synchronous runner has "
|
||||||
|
+ "already run to completion by the time confirm() returns");
|
||||||
|
|
||||||
|
LeadRollover.RollStatus status = rollover.status(pending.token());
|
||||||
|
assertEquals(LeadRollover.RollState.FAILED, status.state(),
|
||||||
|
"a HerdrException out of the /clear send must leave a TERMINAL FAILED outcome — "
|
||||||
|
+ "before fleetd #615's fix, the continuation thread died silently and "
|
||||||
|
+ "status() was stuck reporting the IN_PROGRESS confirm() wrote at hand-off, "
|
||||||
|
+ "forever: got " + status.state() + " / " + status.detail());
|
||||||
|
assertNotEquals(LeadRollover.RollState.IN_PROGRESS, status.state());
|
||||||
|
assertTrue(status.detail().contains("HerdrException"), "the detail must name the exception "
|
||||||
|
+ "so an operator reading status() has something to act on: " + status.detail());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("[fleetd #615 — 2] send() throwing on the bootstrap-text call (after /clear "
|
||||||
|
+ "succeeded and the pane settled) also leaves status(token) reporting FAILED — a "
|
||||||
|
+ "DIFFERENT exit from the /clear-throw case above")
|
||||||
|
void sendThrowingOnBootstrapTextLeavesStatusReportingFailed() throws IOException {
|
||||||
|
FakeHerdr fake = new FakeHerdr(); // default idle throughout — both settle waits pass promptly
|
||||||
|
HerdrClient throwsOnBootstrapText = new HerdrClient() {
|
||||||
|
@Override
|
||||||
|
public JsonNode call(String method, Object params) throws HerdrException {
|
||||||
|
if ("agent.prompt".equals(method) && String.valueOf(params).contains("read the handover file")) {
|
||||||
|
throw new HerdrException("simulated herdr transport failure sending bootstrapText");
|
||||||
|
}
|
||||||
|
return fake.call(method, params);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
fake.close();
|
||||||
|
}
|
||||||
|
};
|
||||||
|
Path handover = writeHandover("handover contents");
|
||||||
|
AtomicLong clock = new AtomicLong(1_000);
|
||||||
|
LeadRollover rollover = newRollover(throwsOnBootstrapText, cfg(handover.toString()), fixedClock(clock));
|
||||||
|
|
||||||
|
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||||
|
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||||
|
|
||||||
|
assertTrue(decision.accepted(), "every synchronous gate passes; the throw happens only "
|
||||||
|
+ "inside the deferred continuation, which this test's synchronous runner has "
|
||||||
|
+ "already run to completion by the time confirm() returns");
|
||||||
|
assertEquals(1, promptCallCount(fake), "sanity: /clear was sent and settled — only the "
|
||||||
|
+ "SECOND agent.prompt call (bootstrapText) threw");
|
||||||
|
|
||||||
|
LeadRollover.RollStatus status = rollover.status(pending.token());
|
||||||
|
assertEquals(LeadRollover.RollState.FAILED, status.state(),
|
||||||
|
"a HerdrException out of the bootstrapText send — a DIFFERENT exit from the /clear "
|
||||||
|
+ "throw, reached only after /clear already succeeded and the pane already "
|
||||||
|
+ "settled — must also leave a TERMINAL FAILED outcome, not a stuck "
|
||||||
|
+ "IN_PROGRESS: got " + status.state() + " / " + status.detail());
|
||||||
|
assertNotEquals(LeadRollover.RollState.IN_PROGRESS, status.state());
|
||||||
|
assertTrue(status.detail().contains("HerdrException"), "the detail must name the exception "
|
||||||
|
+ "so an operator reading status() has something to act on: " + status.detail());
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -258,5 +258,51 @@ class FleetMcpHandoverTest {
|
|||||||
assertTrue(textOf(r).contains("\"cancelled\":true"), textOf(r));
|
assertTrue(textOf(r).contains("\"cancelled\":true"), textOf(r));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- new: action "status" — makes the outcome of a confirm() readable through the tool ----
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("status on a null LeadRollover is a clean NOT_CONFIGURED refusal, never a throw")
|
||||||
|
void statusWithNullLeadRolloverRefusesCleanly() {
|
||||||
|
McpSchema.CallToolResult r = assertDoesNotThrow(() -> FleetMcp.handover(null, LEAD,
|
||||||
|
Map.of("action", "status", "token", "whatever")));
|
||||||
|
assertFalse(r.isError());
|
||||||
|
assertTrue(textOf(r).contains("NOT_CONFIGURED"), textOf(r));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("status on a token that was never opened reports UNKNOWN")
|
||||||
|
void statusOnUnknownTokenReportsUnknown() {
|
||||||
|
LeadRollover rollover = newRollover(tmp.resolve("h.md").toString());
|
||||||
|
|
||||||
|
McpSchema.CallToolResult r = FleetMcp.handover(rollover, LEAD,
|
||||||
|
Map.of("action", "status", "token", "does-not-exist"));
|
||||||
|
assertFalse(r.isError());
|
||||||
|
assertTrue(textOf(r).contains("\"state\":\"UNKNOWN\""), textOf(r));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("status on a token that is still pending (opened, not confirmed) reports PENDING")
|
||||||
|
void statusOnPendingTokenReportsPending() {
|
||||||
|
LeadRollover rollover = newRollover(tmp.resolve("h.md").toString());
|
||||||
|
String token = extractToken(textOf(FleetMcp.handover(rollover, LEAD, Map.of("action", "open"))));
|
||||||
|
|
||||||
|
McpSchema.CallToolResult r = FleetMcp.handover(rollover, LEAD,
|
||||||
|
Map.of("action", "status", "token", token));
|
||||||
|
assertFalse(r.isError());
|
||||||
|
assertTrue(textOf(r).contains("\"state\":\"PENDING\""), textOf(r));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("status is registered on the tool's schema and the schema still names no caller-identity parameter")
|
||||||
|
void statusActionIsAdvertisedOnTheSchema() {
|
||||||
|
FleetMcp m = mcp(null);
|
||||||
|
McpSchema.Tool tool = m.registeredTools().stream()
|
||||||
|
.filter(t -> "fleet_handover".equals(t.name()))
|
||||||
|
.findFirst()
|
||||||
|
.orElseThrow(() -> new AssertionError("fleet_handover was not registered"));
|
||||||
|
assertTrue(tool.description().contains("'status'"),
|
||||||
|
"the tool's own description must advertise the 'status' action: " + tool.description());
|
||||||
|
}
|
||||||
|
|
||||||
// --- acceptance 7 (wiring) is covered by FleetdLeadRolloverWiringTest, unchanged -----------
|
// --- acceptance 7 (wiring) is covered by FleetdLeadRolloverWiringTest, unchanged -----------
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,126 @@
|
|||||||
|
package dev.ltms.fleet.mcp;
|
||||||
|
|
||||||
|
import dev.ltms.fleet.config.FleetConfig;
|
||||||
|
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||||
|
import dev.ltms.fleet.herdr.AgentControl;
|
||||||
|
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||||
|
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||||
|
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||||
|
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||||
|
import dev.ltms.fleet.session.SessionManager;
|
||||||
|
import io.modelcontextprotocol.spec.McpSchema;
|
||||||
|
import org.junit.jupiter.api.DisplayName;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.junit.jupiter.api.io.TempDir;
|
||||||
|
|
||||||
|
import java.io.IOException;
|
||||||
|
import java.nio.charset.StandardCharsets;
|
||||||
|
import java.nio.file.Files;
|
||||||
|
import java.nio.file.Path;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.Set;
|
||||||
|
import java.util.concurrent.atomic.AtomicReference;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #602 gauge-wiring: {@code FleetMcp}'s {@code fleet_list} handler shipped calling
|
||||||
|
* {@code LeadContextGauge.read(null, ...)} unconditionally, so every lead whose profile sets a
|
||||||
|
* {@code configDir:} override read the wrong transcript directory forever, with no error anywhere —
|
||||||
|
* {@link LeadContextGauge}'s own 8 tests all passed because none of them exercised THIS wiring; each
|
||||||
|
* one hands the gauge its own directory directly.
|
||||||
|
*
|
||||||
|
* <p>This class proves the directory {@code fleet.leaders.<name>.profile}'s own {@code configDir:}
|
||||||
|
* names is the one {@code fleet_list}'s {@code context} row actually reads — not merely that SOME
|
||||||
|
* directory got passed. If the wiring in {@code FleetMcp.leadView}/{@code contextView} is ever
|
||||||
|
* reverted to a hardcoded {@code null}, {@link #configuredDirectoryDecidesWhichTranscriptIsRead()}
|
||||||
|
* must go red: both directories below hold a REAL, DIFFERENT token count for the SAME session id,
|
||||||
|
* so a hardcoded {@code null} (always reading the built-in default, where neither transcript lives)
|
||||||
|
* would report {@code UNKNOWN} both times instead of the two distinct numbers this test asserts.
|
||||||
|
*/
|
||||||
|
class FleetMcpLeadContextGaugeWiringTest {
|
||||||
|
|
||||||
|
private static final String LEAD_TERMINAL = "term_lead";
|
||||||
|
private static final String LEAD_NAME = "opus";
|
||||||
|
/** {@code FakeHerdr.withAgent} always projects {@code agent_session.value} as {@code sess-<terminalId>}. */
|
||||||
|
private static final String LEAD_SESSION_ID = "sess-" + LEAD_TERMINAL;
|
||||||
|
|
||||||
|
private static ClaudeCodeLauncher workerService(FakeHerdr h) {
|
||||||
|
FleetConfig.Profile cfg = new FleetConfig.Profile(
|
||||||
|
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
|
||||||
|
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||||
|
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
|
||||||
|
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> "tok");
|
||||||
|
}
|
||||||
|
|
||||||
|
private static String textOf(McpSchema.CallToolResult r) {
|
||||||
|
return ((McpSchema.TextContent) r.content().getFirst()).text();
|
||||||
|
}
|
||||||
|
|
||||||
|
/** One line Claude Code would write for a turn with the given live-context total. */
|
||||||
|
private static String usageLine(long tokens) {
|
||||||
|
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
|
||||||
|
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} carrying one usage record. */
|
||||||
|
private static String writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
|
||||||
|
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
|
||||||
|
Files.createDirectories(projectDir);
|
||||||
|
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
|
||||||
|
return configDir.toString();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("the CONFIGURED directory decides which transcript is read, and follows a config change")
|
||||||
|
void configuredDirectoryDecidesWhichTranscriptIsRead(@TempDir Path tmp) throws IOException {
|
||||||
|
Path dirA = Files.createDirectory(tmp.resolve("dir-a"));
|
||||||
|
Path dirB = Files.createDirectory(tmp.resolve("dir-b"));
|
||||||
|
writeTranscript(dirA, LEAD_SESSION_ID, 11_000);
|
||||||
|
writeTranscript(dirB, LEAD_SESSION_ID, 22_000);
|
||||||
|
|
||||||
|
FakeHerdr herdr = new FakeHerdr().withAgent("lead-opus", LEAD_TERMINAL, "wL:p1", "wL:t1");
|
||||||
|
SessionManager sessions = new SessionManager(workerService(herdr));
|
||||||
|
LeadContextGauge contextGauge = new LeadContextGauge();
|
||||||
|
AtomicReference<String> configuredDir = new AtomicReference<>(dirA.toString());
|
||||||
|
FleetMcp.LeadConfigDirSource source = new FleetMcp.LeadConfigDirSource(name ->
|
||||||
|
LEAD_NAME.equals(name) ? configuredDir.get() : null);
|
||||||
|
|
||||||
|
String firstRead = textOf(listFleet(herdr, sessions, contextGauge, source));
|
||||||
|
assertTrue(firstRead.contains("\"tokens\":11000"),
|
||||||
|
"config names directory A ⇒ fleet_list must report A's token count: " + firstRead);
|
||||||
|
|
||||||
|
// The config changes to name directory B instead -- the very next read must follow it.
|
||||||
|
configuredDir.set(dirB.toString());
|
||||||
|
String secondRead = textOf(listFleet(herdr, sessions, contextGauge, source));
|
||||||
|
assertTrue(secondRead.contains("\"tokens\":22000"),
|
||||||
|
"config now names directory B ⇒ fleet_list must report B's token count, not A's stale one: "
|
||||||
|
+ secondRead);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
@DisplayName("a lead whose config names no directory still degrades to the built-in default, never throws")
|
||||||
|
void aLeadWhoseConfigNamesNoDirectoryStillFallsBackWithoutThrowing() {
|
||||||
|
FakeHerdr herdr = new FakeHerdr().withAgent("lead-opus", LEAD_TERMINAL, "wL:p1", "wL:t1");
|
||||||
|
SessionManager sessions = new SessionManager(workerService(herdr));
|
||||||
|
LeadContextGauge contextGauge = new LeadContextGauge();
|
||||||
|
|
||||||
|
McpSchema.CallToolResult res = listFleet(herdr, sessions, contextGauge, FleetMcp.LeadConfigDirSource.none());
|
||||||
|
|
||||||
|
assertNotEquals(Boolean.TRUE, res.isError(), "a lead with no configured configDir must degrade, never throw");
|
||||||
|
String out = textOf(res);
|
||||||
|
assertTrue(out.contains("\"state\":\"unknown\""), "no transcript under the built-in default ⇒ UNKNOWN: " + out);
|
||||||
|
assertFalse(out.contains("\"tokens\""), "UNKNOWN must never carry a stale/default token number: " + out);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static McpSchema.CallToolResult listFleet(FakeHerdr herdr, SessionManager sessions,
|
||||||
|
LeadContextGauge contextGauge, FleetMcp.LeadConfigDirSource leadConfigDirs) {
|
||||||
|
return FleetMcp.listFleet(workerService(herdr), sessions, null,
|
||||||
|
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||||
|
FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
|
||||||
|
FleetMcp.LeadSeatSource.none(), contextGauge, leadConfigDirs,
|
||||||
|
Map.of(LEAD_TERMINAL, LEAD_NAME), LEAD_TERMINAL, FleetMcp.CoordinationSource.none(), false);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1937,7 +1937,7 @@ class FleetMcpTest {
|
|||||||
null, null, null, null, null, null);
|
null, null, null, null, null, null);
|
||||||
|
|
||||||
assertEquals(Boolean.TRUE, res.isError());
|
assertEquals(Boolean.TRUE, res.isError());
|
||||||
assertTrue(textOf(res).contains("architect, dev, reviewer"), textOf(res));
|
assertTrue(textOf(res).contains("architect, dev, hunter, reviewer"), textOf(res));
|
||||||
}
|
}
|
||||||
|
|
||||||
// ── CB-619 / fleetd #123: a spawn asking for a role its profile has no slot for must be
|
// ── CB-619 / fleetd #123: a spawn asking for a role its profile has no slot for must be
|
||||||
|
|||||||
@@ -556,6 +556,31 @@ class ClaudeCodeLauncherTest {
|
|||||||
"no --agent flag when the role has no agent-definition file");
|
"no --agent flag when the role has no agent-definition file");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void hunterRoleUsesItsAgentFileAndStopsUsingItWhenRemoved(@TempDir Path cwd) throws Exception {
|
||||||
|
FakeHerdr herdr = new FakeHerdr();
|
||||||
|
Path agentFile = Files.createDirectories(cwd.resolve(".claude/agents")).resolve("hunter.md");
|
||||||
|
Files.writeString(agentFile, "---\nname: hunter\n---\nSweep for defects.");
|
||||||
|
FleetConfig.Profile cfg = new FleetConfig.Profile(
|
||||||
|
"sonnet", "http://gx00.gw:8000", null, null, "FLEETD_WORKER_TOKEN",
|
||||||
|
List.of("claude"), "tab", "fleetd-workers", "w #{n}", null, null, null);
|
||||||
|
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||||
|
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||||
|
|
||||||
|
svc.spawn(new SpawnRequest("sonnet", cwd.toString(), null, null, null, MemberRole.HUNTER));
|
||||||
|
|
||||||
|
List<String> args = spawnedArgs(herdr);
|
||||||
|
int flag = args.indexOf("--agent");
|
||||||
|
assertTrue(flag >= 0, "the hunter role reaches its agent-definition file: " + args);
|
||||||
|
assertEquals("hunter", args.get(flag + 1));
|
||||||
|
|
||||||
|
Files.delete(agentFile);
|
||||||
|
svc.spawn(new SpawnRequest("sonnet", cwd.toString(), null, null, null, MemberRole.HUNTER));
|
||||||
|
|
||||||
|
assertFalse(spawnedArgs(herdr).contains("--agent"),
|
||||||
|
"the hunter role no longer gets an agent when its file is removed");
|
||||||
|
}
|
||||||
|
|
||||||
private ClaudeCodeLauncher multiProfile(FakeHerdr herdr) {
|
private ClaudeCodeLauncher multiProfile(FakeHerdr herdr) {
|
||||||
FleetConfig.Profile gx10 = new FleetConfig.Profile("gx10", "http://gx10.gw:8000", "coder",
|
FleetConfig.Profile gx10 = new FleetConfig.Profile("gx10", "http://gx10.gw:8000", "coder",
|
||||||
null, "FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers", "w #{n}", null, null, null);
|
null, "FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers", "w #{n}", null, null, null);
|
||||||
|
|||||||
@@ -1,27 +1,46 @@
|
|||||||
package dev.ltms.fleet.msg;
|
package dev.ltms.fleet.msg;
|
||||||
|
|
||||||
|
import ch.qos.logback.classic.Level;
|
||||||
|
import ch.qos.logback.classic.Logger;
|
||||||
|
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||||
|
import ch.qos.logback.core.read.ListAppender;
|
||||||
|
import com.fasterxml.jackson.databind.JsonNode;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
import dev.ltms.fleet.herdr.AgentControl;
|
||||||
import dev.ltms.fleet.herdr.AgentStatus;
|
import dev.ltms.fleet.herdr.AgentStatus;
|
||||||
|
import dev.ltms.fleet.herdr.HerdrClient;
|
||||||
|
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||||
import dev.ltms.fleet.mcp.PrimaryRegistry;
|
import dev.ltms.fleet.mcp.PrimaryRegistry;
|
||||||
import dev.ltms.fleet.peer.MemberRole;
|
import dev.ltms.fleet.peer.MemberRole;
|
||||||
import dev.ltms.fleet.session.MemberSession;
|
import dev.ltms.fleet.session.MemberSession;
|
||||||
import org.junit.jupiter.api.AfterEach;
|
import org.junit.jupiter.api.AfterEach;
|
||||||
import org.junit.jupiter.api.BeforeEach;
|
import org.junit.jupiter.api.BeforeEach;
|
||||||
import org.junit.jupiter.api.Test;
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.slf4j.LoggerFactory;
|
||||||
|
|
||||||
|
import java.util.ArrayList;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
import java.util.concurrent.Executors;
|
import java.util.concurrent.Executors;
|
||||||
import java.util.concurrent.ScheduledExecutorService;
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
import java.util.concurrent.TimeUnit;
|
import java.util.concurrent.TimeUnit;
|
||||||
|
import java.util.concurrent.atomic.AtomicLong;
|
||||||
|
|
||||||
import static org.junit.jupiter.api.Assertions.*;
|
import static org.junit.jupiter.api.Assertions.*;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Unit tests for the CB-551 idle-lead heartbeat: the pure {@link LeadHeartbeatLoop#decide} decision
|
* Unit tests for the CB-551 idle-lead heartbeat: the pure {@link LeadHeartbeatLoop#decide} decision
|
||||||
* function, the nudge text, and the {@link LeadHeartbeatLoop#snapshot} fleet snapshot.
|
* function, the nudge text, and the {@link LeadHeartbeatLoop#snapshot} fleet snapshot. Also covers
|
||||||
|
* fleetd #609's context-high notice: the {@code context}/{@code contextNotified} parameters {@code
|
||||||
|
* decide} gained, and the standalone {@link LeadHeartbeatLoop#contextNotice} text builder.
|
||||||
*
|
*
|
||||||
* <p>All decision tests call {@code decide} directly with explicit nanoTime values from an injected
|
* <p>All decision tests call {@code decide} directly with explicit nanoTime values from an injected
|
||||||
* clock — no sleeping, no scheduler races. This mirrors how {@code ReplyPushLoopTest} pins the pure
|
* clock — no sleeping, no scheduler races. This mirrors how {@code ReplyPushLoopTest} pins the pure
|
||||||
* decision before exercising the loop.
|
* decision before exercising the loop.
|
||||||
|
*
|
||||||
|
* <p>Every pre-#609 test below passes {@code LeadContextGauge.State.UNKNOWN, false} for the two new
|
||||||
|
* {@code decide} parameters — the same as a lead whose context could not be read and has never been
|
||||||
|
* notified — so each one still pins exactly the behaviour it pinned before this ticket.
|
||||||
*/
|
*/
|
||||||
class LeadHeartbeatLoopTest {
|
class LeadHeartbeatLoopTest {
|
||||||
|
|
||||||
@@ -56,12 +75,18 @@ class LeadHeartbeatLoopTest {
|
|||||||
return new LeadHeartbeatLoop.FleetState(2, 1, 3, List.of("term_a", "term_b"));
|
return new LeadHeartbeatLoop.FleetState(2, 1, 3, List.of("term_a", "term_b"));
|
||||||
}
|
}
|
||||||
|
|
||||||
/** The loop under test; the scheduler is never invoked on the pure decide path. */
|
/** The loop under test, context notice off; the scheduler is never invoked on the pure decide path. */
|
||||||
private static LeadHeartbeatLoop loop(int quietCap) {
|
private static LeadHeartbeatLoop loop(int quietCap) {
|
||||||
|
return loop(quietCap, false);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** As above, with the fleetd #609 {@code contextHighNudge} flag set explicitly. */
|
||||||
|
private static LeadHeartbeatLoop loop(int quietCap, boolean contextHighNudge) {
|
||||||
return new LeadHeartbeatLoop(
|
return new LeadHeartbeatLoop(
|
||||||
new PrimaryRegistry("term_lead"), null /*agents — unused on the decide path*/,
|
new PrimaryRegistry("term_lead"), null /*agents — unused on the decide path*/,
|
||||||
null /*inbox*/, List::of, null /*pushLoop*/, null /*scheduler*/, () -> 0L,
|
null /*inbox*/, List::of, null /*pushLoop*/, null /*scheduler*/, () -> 0L,
|
||||||
IDLE_AFTER_NANOS, 1_000L, quietCap);
|
IDLE_AFTER_NANOS, 1_000L, quietCap, null,
|
||||||
|
LeadHeartbeatLoop.LeadContextSource.none(), contextHighNudge);
|
||||||
}
|
}
|
||||||
|
|
||||||
// ── (a) a working lead is never injected ───────────────────────────────────────────────────
|
// ── (a) a working lead is never injected ───────────────────────────────────────────────────
|
||||||
@@ -69,7 +94,8 @@ class LeadHeartbeatLoopTest {
|
|||||||
@Test
|
@Test
|
||||||
void aWorkingLeadIsNeverInjected() {
|
void aWorkingLeadIsNeverInjected() {
|
||||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||||
NOW, IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet());
|
NOW, IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.UNKNOWN, false);
|
||||||
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(),
|
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(),
|
||||||
"a WORKING lead is making progress and must not be touched");
|
"a WORKING lead is making progress and must not be touched");
|
||||||
assertNull(d.idleSinceNanos(), "a busy lead resets the idle window");
|
assertNull(d.idleSinceNanos(), "a busy lead resets the idle window");
|
||||||
@@ -80,17 +106,71 @@ class LeadHeartbeatLoopTest {
|
|||||||
void anUnreadableStatusIsNeverInjectedEither() {
|
void anUnreadableStatusIsNeverInjectedEither() {
|
||||||
// A failed status read (or a gone agent) must degrade to "do not inject", never hammer the pane.
|
// A failed status read (or a gone agent) must degrade to "do not inject", never hammer the pane.
|
||||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||||
NOW, IDLE_PAST, 0, AgentStatus.UNKNOWN, false, true, pendingFleet());
|
NOW, IDLE_PAST, 0, AgentStatus.UNKNOWN, false, true, pendingFleet(),
|
||||||
|
LeadContextGauge.State.UNKNOWN, false);
|
||||||
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(),
|
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(),
|
||||||
"never inject into a state the loop cannot read");
|
"never inject into a state the loop cannot read");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ── fleetd #613: the boot line names all four heartbeat settings ──────────────────────────
|
||||||
|
|
||||||
|
/**
|
||||||
|
* fleetd #613: {@link LeadHeartbeatLoop#start()}'s boot line named only 3 of the 4 constructor
|
||||||
|
* settings — {@code contextHighNudge} (fleetd #609) was missing, so an operator could not tell
|
||||||
|
* from the log whether the handover notice was armed. Captures the real log via a
|
||||||
|
* {@link ListAppender}, raising the logger's level past {@code logback-test.xml}'s
|
||||||
|
* {@code dev.ltms.fleet -> WARN} override for the duration of the call — the same seam {@code
|
||||||
|
* GitHostShapeReportTest#attach} uses for its own INFO-level boot line.
|
||||||
|
*/
|
||||||
|
private static String heartbeatBootLine(boolean contextHighNudge, ScheduledExecutorService scheduler) {
|
||||||
|
Logger logger = (Logger) LoggerFactory.getLogger(LeadHeartbeatLoop.class);
|
||||||
|
Level originalLevel = logger.getLevel();
|
||||||
|
logger.setLevel(Level.INFO);
|
||||||
|
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||||
|
appender.start();
|
||||||
|
logger.addAppender(appender);
|
||||||
|
try {
|
||||||
|
LeadHeartbeatLoop l = new LeadHeartbeatLoop(
|
||||||
|
new PrimaryRegistry("term_lead"), null /*agents*/, null /*inbox*/, List::of,
|
||||||
|
null /*pushLoop*/, scheduler, () -> 0L, IDLE_AFTER_NANOS, 1_000L, 3, null,
|
||||||
|
LeadHeartbeatLoop.LeadContextSource.none(), contextHighNudge);
|
||||||
|
l.start();
|
||||||
|
} finally {
|
||||||
|
logger.detachAppender(appender);
|
||||||
|
logger.setLevel(originalLevel);
|
||||||
|
}
|
||||||
|
return appender.list.stream()
|
||||||
|
.filter(e -> e.getLevel() == Level.INFO)
|
||||||
|
.map(ILoggingEvent::getFormattedMessage)
|
||||||
|
.filter(m -> m.startsWith("idle-lead heartbeat:"))
|
||||||
|
.findFirst()
|
||||||
|
.orElseThrow(() -> new AssertionError("expected the heartbeat boot line to be logged"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void theBootLineNamesContextHighNudgeWhenArmed() {
|
||||||
|
String line = heartbeatBootLine(true, scheduler);
|
||||||
|
assertTrue(line.contains("300s idle"), line);
|
||||||
|
assertTrue(line.contains("recheck 1000ms"), line);
|
||||||
|
assertTrue(line.contains("quiet cap 3"), line);
|
||||||
|
assertTrue(line.contains("context-high nudge true"),
|
||||||
|
"the boot line must name the 4th setting, contextHighNudge, alongside the other "
|
||||||
|
+ "three: " + line);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void theBootLineNamesContextHighNudgeWhenOff() {
|
||||||
|
String line = heartbeatBootLine(false, scheduler);
|
||||||
|
assertTrue(line.contains("context-high nudge false"), line);
|
||||||
|
}
|
||||||
|
|
||||||
// ── (b) an idle lead within the quiet period is not yet injected ───────────────────────────
|
// ── (b) an idle lead within the quiet period is not yet injected ───────────────────────────
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void justBecameIdleStartsTheDebounceWindow() {
|
void justBecameIdleStartsTheDebounceWindow() {
|
||||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||||
NOW, null, 0, AgentStatus.IDLE, false, true, quietFleet());
|
NOW, null, 0, AgentStatus.IDLE, false, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.UNKNOWN, false);
|
||||||
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
|
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
|
||||||
"the first injectable tick only records the start of the idle stretch");
|
"the first injectable tick only records the start of the idle stretch");
|
||||||
assertEquals(NOW, d.idleSinceNanos(), "the idle window opens at the moment the lead became injectable");
|
assertEquals(NOW, d.idleSinceNanos(), "the idle window opens at the moment the lead became injectable");
|
||||||
@@ -99,7 +179,8 @@ class LeadHeartbeatLoopTest {
|
|||||||
@Test
|
@Test
|
||||||
void idleWithinQuietPeriodIsNotInjected() {
|
void idleWithinQuietPeriodIsNotInjected() {
|
||||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||||
NOW, IDLE_WITHIN, 0, AgentStatus.IDLE, false, true, quietFleet());
|
NOW, IDLE_WITHIN, 0, AgentStatus.IDLE, false, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.UNKNOWN, false);
|
||||||
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
|
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
|
||||||
"a lead idle for 10s (< 300s) has just finished a turn — do not re-prompt it");
|
"a lead idle for 10s (< 300s) has just finished a turn — do not re-prompt it");
|
||||||
}
|
}
|
||||||
@@ -109,7 +190,8 @@ class LeadHeartbeatLoopTest {
|
|||||||
@Test
|
@Test
|
||||||
void idlePastQuietPeriodIsInjected() {
|
void idlePastQuietPeriodIsInjected() {
|
||||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||||
NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, true, quietFleet());
|
NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.UNKNOWN, false);
|
||||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
|
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
|
||||||
"a lead continuously idle past the quiet period is the reason to nudge");
|
"a lead continuously idle past the quiet period is the reason to nudge");
|
||||||
}
|
}
|
||||||
@@ -117,14 +199,17 @@ class LeadHeartbeatLoopTest {
|
|||||||
@Test
|
@Test
|
||||||
void blockedAndDoneAreInjectableViewsOfIdle() {
|
void blockedAndDoneAreInjectableViewsOfIdle() {
|
||||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, loop(3).decide(
|
assertEquals(LeadHeartbeatLoop.Action.INJECT, loop(3).decide(
|
||||||
NOW, IDLE_PAST, 0, AgentStatus.BLOCKED, false, true, quietFleet()).action());
|
NOW, IDLE_PAST, 0, AgentStatus.BLOCKED, false, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.UNKNOWN, false).action());
|
||||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, loop(3).decide(
|
assertEquals(LeadHeartbeatLoop.Action.INJECT, loop(3).decide(
|
||||||
NOW, IDLE_PAST, 0, AgentStatus.DONE, false, true, quietFleet()).action());
|
NOW, IDLE_PAST, 0, AgentStatus.DONE, false, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.UNKNOWN, false).action());
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void nothingIsInjectedWhenNoLeadIsKnown() {
|
void nothingIsInjectedWhenNoLeadIsKnown() {
|
||||||
var d = loop(3).decide(NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, false, pendingFleet());
|
var d = loop(3).decide(NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, false, pendingFleet(),
|
||||||
|
LeadContextGauge.State.UNKNOWN, false);
|
||||||
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
|
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
|
||||||
"with no known lead there is nobody to nudge — keep waiting until one is discovered");
|
"with no known lead there is nobody to nudge — keep waiting until one is discovered");
|
||||||
assertEquals(IDLE_PAST, d.idleSinceNanos(), "the idle window stays open so discovery re-arms it");
|
assertEquals(IDLE_PAST, d.idleSinceNanos(), "the idle window stays open so discovery re-arms it");
|
||||||
@@ -135,7 +220,8 @@ class LeadHeartbeatLoopTest {
|
|||||||
@Test
|
@Test
|
||||||
void quietNudgeCapStopsTheLoopWhenNothingIsPending() {
|
void quietNudgeCapStopsTheLoopWhenNothingIsPending() {
|
||||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet());
|
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.UNKNOWN, false);
|
||||||
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, d.action(),
|
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, d.action(),
|
||||||
"3 consecutive nothing-pending nudges have already happened — stop nagging an empty fleet");
|
"3 consecutive nothing-pending nudges have already happened — stop nagging an empty fleet");
|
||||||
}
|
}
|
||||||
@@ -145,7 +231,8 @@ class LeadHeartbeatLoopTest {
|
|||||||
@Test
|
@Test
|
||||||
void newPendingStateResetsTheQuietCap() {
|
void newPendingStateResetsTheQuietCap() {
|
||||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet());
|
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet(),
|
||||||
|
LeadContextGauge.State.UNKNOWN, false);
|
||||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
|
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
|
||||||
"real state appearing re-arms the loop past an exhausted cap");
|
"real state appearing re-arms the loop past an exhausted cap");
|
||||||
assertEquals(0, d.quietCount(), "the pending state resets the consecutive-quiet counter");
|
assertEquals(0, d.quietCount(), "the pending state resets the consecutive-quiet counter");
|
||||||
@@ -156,7 +243,8 @@ class LeadHeartbeatLoopTest {
|
|||||||
@Test
|
@Test
|
||||||
void standsDownWhileReplyPushLoopIsActive() {
|
void standsDownWhileReplyPushLoopIsActive() {
|
||||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||||
NOW, IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet());
|
NOW, IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.UNKNOWN, false);
|
||||||
assertEquals(LeadHeartbeatLoop.Action.STAND_DOWN, d.action(),
|
assertEquals(LeadHeartbeatLoop.Action.STAND_DOWN, d.action(),
|
||||||
"a second injection would start a competing turn — stand aside instead");
|
"a second injection would start a competing turn — stand aside instead");
|
||||||
assertEquals(0, d.quietCount(), "the active push is real state, so it re-arms the cap");
|
assertEquals(0, d.quietCount(), "the active push is real state, so it re-arms the cap");
|
||||||
@@ -213,4 +301,398 @@ class LeadHeartbeatLoopTest {
|
|||||||
assertFalse(fs.hasPending());
|
assertFalse(fs.hasPending());
|
||||||
assertEquals(0, fs.liveWorkers());
|
assertEquals(0, fs.liveWorkers());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ── fleetd #609: context-high notice ───────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
/** One gate scenario, keyed by name, over which the off-state/context-independence property is checked. */
|
||||||
|
private record Gate(String name, Long idleSince, int quietCount, AgentStatus status,
|
||||||
|
boolean pushLoopActive, boolean leadKnown, LeadHeartbeatLoop.FleetState fleet) {}
|
||||||
|
|
||||||
|
private static List<Gate> allEightGates() {
|
||||||
|
return List.of(
|
||||||
|
new Gate("1-standDown", IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet()),
|
||||||
|
new Gate("2-leadBusy", IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet()),
|
||||||
|
new Gate("3-justBecameIdle", null, 0, AgentStatus.IDLE, false, true, quietFleet()),
|
||||||
|
new Gate("4-withinQuietPeriod", IDLE_WITHIN, 0, AgentStatus.IDLE, false, true, quietFleet()),
|
||||||
|
new Gate("5-noLeadKnown", IDLE_PAST, 0, AgentStatus.IDLE, false, false, pendingFleet()),
|
||||||
|
new Gate("6-pending", IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet()),
|
||||||
|
new Gate("7-quietNotExhausted", IDLE_PAST, 0, AgentStatus.IDLE, false, true, quietFleet()),
|
||||||
|
new Gate("8-quietExhausted", IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet()));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void aOffFlagIgnoresContextAcrossAllEightGates() {
|
||||||
|
// Property A: with contextHighNudge off, decide() must not depend on the context reading at
|
||||||
|
// all — not just "usually agrees", but identical Action/idleSinceNanos/quietCount whatever
|
||||||
|
// context state is passed, and contextNotified must stay false throughout. This is the proof
|
||||||
|
// that fleetd #609 is opt-in: a daemon upgraded to carry this code, but never configuring
|
||||||
|
// `contextHighNudge: true`, behaves exactly as it did before this ticket for every one of the
|
||||||
|
// 8 gates the class javadoc numbers.
|
||||||
|
LeadHeartbeatLoop offLoop = loop(3, false);
|
||||||
|
for (Gate g : allEightGates()) {
|
||||||
|
var withHigh = offLoop.decide(NOW, g.idleSince(), g.quietCount(), g.status(),
|
||||||
|
g.pushLoopActive(), g.leadKnown(), g.fleet(), LeadContextGauge.State.HIGH, false);
|
||||||
|
var withUnknown = offLoop.decide(NOW, g.idleSince(), g.quietCount(), g.status(),
|
||||||
|
g.pushLoopActive(), g.leadKnown(), g.fleet(), LeadContextGauge.State.UNKNOWN, false);
|
||||||
|
var withOk = offLoop.decide(NOW, g.idleSince(), g.quietCount(), g.status(),
|
||||||
|
g.pushLoopActive(), g.leadKnown(), g.fleet(), LeadContextGauge.State.OK, false);
|
||||||
|
|
||||||
|
assertEquals(withUnknown.action(), withHigh.action(), g.name() + ": action must not depend on context");
|
||||||
|
assertEquals(withUnknown.idleSinceNanos(), withHigh.idleSinceNanos(), g.name());
|
||||||
|
assertEquals(withUnknown.quietCount(), withHigh.quietCount(), g.name());
|
||||||
|
assertEquals(withUnknown.action(), withOk.action(), g.name() + ": nor on an OK reading");
|
||||||
|
assertFalse(withHigh.contextNotified(), g.name() + ": contextNotified must stay false when the flag is off");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void bHighContextFiresEvenPastAnExhaustedQuietCapWithoutSpendingIt() {
|
||||||
|
LeadHeartbeatLoop on = loop(3, true);
|
||||||
|
LeadHeartbeatLoop.Decision d = on.decide(
|
||||||
|
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.HIGH, false);
|
||||||
|
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
|
||||||
|
"an idle, quiet, HIGH-context lead is exactly the case the exhausted cap must not swallow");
|
||||||
|
assertEquals(3, d.quietCount(), "the context notice is an event notice, not a quiet nudge — it must not "
|
||||||
|
+ "spend or grow the consecutive-quiet budget");
|
||||||
|
assertTrue(d.contextNotified(), "firing the notice sets the latch");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void cASecondTickWithTheLatchAlreadySetDoesNotTellTheLeadAgain() {
|
||||||
|
LeadHeartbeatLoop on = loop(3, true);
|
||||||
|
LeadHeartbeatLoop.Decision d = on.decide(
|
||||||
|
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.HIGH, true);
|
||||||
|
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, d.action(),
|
||||||
|
"the lead was already told once about this HIGH stretch — telling it every tick would be nagging, "
|
||||||
|
+ "not a notice");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void dFlappingBetweenHighAndUnknownNeverReInjectsWhileLatched() {
|
||||||
|
LeadHeartbeatLoop on = loop(3, true);
|
||||||
|
// The lead was already told once (latch = true from a prior HIGH tick).
|
||||||
|
LeadHeartbeatLoop.Decision afterUnknown = on.decide(
|
||||||
|
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.UNKNOWN, true);
|
||||||
|
assertTrue(afterUnknown.contextNotified(),
|
||||||
|
"UNKNOWN means 'I could not look', not 'it got better' — it must not clear the latch");
|
||||||
|
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, afterUnknown.action(),
|
||||||
|
"a still-latched, still-quiet tick must not inject just because the reading is UNKNOWN");
|
||||||
|
|
||||||
|
LeadHeartbeatLoop.Decision afterHighAgain = on.decide(
|
||||||
|
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.HIGH, afterUnknown.contextNotified());
|
||||||
|
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, afterHighAgain.action(),
|
||||||
|
"HIGH -> UNKNOWN -> HIGH with the latch already set must never inject again — this is the "
|
||||||
|
+ "flapping case that would otherwise cost a full lead its remaining turns");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void eAnOkReadingClearsTheLatchSoALaterHighInjectsAgain() {
|
||||||
|
LeadHeartbeatLoop on = loop(3, true);
|
||||||
|
LeadHeartbeatLoop.Decision afterOk = on.decide(
|
||||||
|
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.OK, true);
|
||||||
|
assertFalse(afterOk.contextNotified(), "an OK reading clears the latch — the lead's context recovered");
|
||||||
|
|
||||||
|
LeadHeartbeatLoop.Decision afterHighAgain = on.decide(
|
||||||
|
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.HIGH, afterOk.contextNotified());
|
||||||
|
assertEquals(LeadHeartbeatLoop.Action.INJECT, afterHighAgain.action(),
|
||||||
|
"the cleared latch lets a genuinely new HIGH stretch notify again");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void fStandingDownDoesNotBurnTheOneContextNotice() {
|
||||||
|
LeadHeartbeatLoop on = loop(3, true);
|
||||||
|
LeadHeartbeatLoop.Decision d = on.decide(
|
||||||
|
NOW, IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.HIGH, false);
|
||||||
|
assertEquals(LeadHeartbeatLoop.Action.STAND_DOWN, d.action(), "ReplyPushLoop is active — stand aside");
|
||||||
|
assertFalse(d.contextNotified(),
|
||||||
|
"standing down must not spend the one notice this HIGH stretch gets — the latch stays clear so a "
|
||||||
|
+ "later tick can still fire it");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void gAWorkingLeadWithHighContextIsStillLeadBusy() {
|
||||||
|
LeadHeartbeatLoop on = loop(3, true);
|
||||||
|
LeadHeartbeatLoop.Decision d = on.decide(
|
||||||
|
NOW, IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet(),
|
||||||
|
LeadContextGauge.State.HIGH, false);
|
||||||
|
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(), "the status gate wins over the context notice");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void hAPendingDrivenInjectWhileHighSetsTheLatchTooSoTheLeadIsNotToldTwice() {
|
||||||
|
LeadHeartbeatLoop on = loop(3, true);
|
||||||
|
LeadHeartbeatLoop.Decision d = on.decide(
|
||||||
|
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet(),
|
||||||
|
LeadContextGauge.State.HIGH, false);
|
||||||
|
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action());
|
||||||
|
assertTrue(d.contextNotified(),
|
||||||
|
"the pending-driven nudge carries the context notice too (see contextNotice), so it must set the "
|
||||||
|
+ "latch — otherwise the lead could be told twice by two different routes");
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── contextNotice text builder ─────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void contextNoticeIsEmptyWhenDisabled() {
|
||||||
|
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
|
||||||
|
assertEquals("", LeadHeartbeatLoop.contextNotice(false, reading));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void contextNoticeIsEmptyWhenStateIsOk() {
|
||||||
|
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.OK, 1_000L, 0);
|
||||||
|
assertEquals("", LeadHeartbeatLoop.contextNotice(true, reading));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void contextNoticeIsEmptyWhenStateIsUnknown() {
|
||||||
|
assertEquals("", LeadHeartbeatLoop.contextNotice(true, LeadContextGauge.Reading.unknown()));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void contextNoticeNamesTheTokenCountWhenHighAndEnabled() {
|
||||||
|
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
|
||||||
|
String notice = LeadHeartbeatLoop.contextNotice(true, reading);
|
||||||
|
assertFalse(notice.isEmpty());
|
||||||
|
assertTrue(notice.contains("260771"), notice);
|
||||||
|
assertTrue(notice.contains("1 compaction"), notice);
|
||||||
|
assertTrue(notice.contains("fleet_handover"), notice);
|
||||||
|
assertTrue(notice.contains("operator"), notice);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void contextNoticeOmitsTheTokenClauseRatherThanPrintingNull() {
|
||||||
|
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, null, 2);
|
||||||
|
String notice = LeadHeartbeatLoop.contextNotice(true, reading);
|
||||||
|
assertFalse(notice.isEmpty());
|
||||||
|
assertFalse(notice.toLowerCase().contains("null"), notice);
|
||||||
|
assertTrue(notice.contains("2 compactions"), notice);
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── fleetd #621: the notice must track the effective requireOperatorConfirm value ─────────────
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void contextNoticeKeepsAskingTheOperatorWhenRequireOperatorConfirmIsTrue() {
|
||||||
|
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
|
||||||
|
String notice = LeadHeartbeatLoop.contextNotice(true, reading, false, true);
|
||||||
|
assertTrue(notice.contains("ask the operator"), notice);
|
||||||
|
assertTrue(notice.contains("Only the operator can approve the roll"), notice);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void contextNoticeDropsTheOperatorAskWhenRequireOperatorConfirmIsFalse() {
|
||||||
|
var reading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
|
||||||
|
String notice = LeadHeartbeatLoop.contextNotice(true, reading, false, false);
|
||||||
|
assertFalse(notice.contains("ask the operator"), notice);
|
||||||
|
assertFalse(notice.contains("Only the operator can approve the roll"), notice);
|
||||||
|
assertTrue(notice.contains("fleet_handover"), notice);
|
||||||
|
assertTrue(notice.contains("maxDocAgeSeconds"), notice);
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── fleetd #609 review: the latch must mean "the notice reached the pane" ────────────────────
|
||||||
|
//
|
||||||
|
// These four drive LeadHeartbeatLoop.tick() directly (package-private, same reasoning as
|
||||||
|
// ReplyPushLoop#tick(String) being directly testable) against a real AgentControl wrapping a
|
||||||
|
// FailableHerdrClient, so the send path (agents.send -> herdr -> possible throw) is exercised
|
||||||
|
// for real rather than assumed from decide()'s Decision alone.
|
||||||
|
|
||||||
|
private static final String LEAD = "term_lead";
|
||||||
|
private static final String WORKER = "term_w1";
|
||||||
|
|
||||||
|
/** A HIGH reading with a fixed token/compaction count, for the four tests below. */
|
||||||
|
private static LeadContextGauge.Reading highReading() {
|
||||||
|
return new LeadContextGauge.Reading(LeadContextGauge.State.HIGH, 260_771L, 1);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Builds a real {@link LeadHeartbeatLoop} wired to {@code herdr} via a real {@link AgentControl},
|
||||||
|
* a mutable fake clock, and a mutable roster so a test can change fleet state between ticks. The
|
||||||
|
* lead is always reported IDLE by {@code herdr}, so every tick's outcome is governed only by the
|
||||||
|
* idle-window/quiet-cap/context gates under test.
|
||||||
|
*/
|
||||||
|
private static LeadHeartbeatLoop tickableLoop(FailableHerdrClient herdr, AtomicLong now,
|
||||||
|
List<MemberSession>[] rosterBox, InMemoryReplyInbox inbox,
|
||||||
|
int quietNudgeCap, ScheduledExecutorService scheduler) {
|
||||||
|
AgentControl agents = new AgentControl(herdr);
|
||||||
|
PrimaryRegistry registry = new PrimaryRegistry(LEAD);
|
||||||
|
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, agents, inbox, scheduler, 5, 100_000);
|
||||||
|
return new LeadHeartbeatLoop(registry, agents, inbox, () -> rosterBox[0], pushLoop, scheduler,
|
||||||
|
now::get, IDLE_AFTER_NANOS, 100_000L, quietNudgeCap, null,
|
||||||
|
new LeadHeartbeatLoop.LeadContextSource(t -> highReading()), true);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void iAFailedSendDoesNotConsumeTheNotice() {
|
||||||
|
var herdr = new FailableHerdrClient(LEAD);
|
||||||
|
var now = new AtomicLong(NOW);
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
List<MemberSession>[] rosterBox = new List[]{List.of()}; // quiet: nothing pending
|
||||||
|
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||||
|
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler);
|
||||||
|
|
||||||
|
loop.tick(); // first injectable tick: only opens the idle window (WAIT_IDLE)
|
||||||
|
now.addAndGet(TimeUnit.SECONDS.toNanos(400)); // now clearly past the quiet period
|
||||||
|
|
||||||
|
herdr.throwOnNextSend();
|
||||||
|
loop.tick(); // quiet fleet, quiet cap exhausted (0), context HIGH, latch clear -> INJECT, send throws
|
||||||
|
|
||||||
|
assertEquals(0, herdr.sentTexts().size(), "the failed send must not have recorded any text");
|
||||||
|
|
||||||
|
loop.tick(); // same inputs — the latch must still be clear, so this must INJECT and send again
|
||||||
|
assertEquals(1, herdr.sentTexts().size(),
|
||||||
|
"a retried tick with the latch still clear must attempt the send again");
|
||||||
|
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"),
|
||||||
|
"the retried, successful send must carry the notice: " + herdr.sentTexts().get(0));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void jASuccessfulSendDoesConsumeIt() {
|
||||||
|
var herdr = new FailableHerdrClient(LEAD);
|
||||||
|
var now = new AtomicLong(NOW);
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
List<MemberSession>[] rosterBox = new List[]{List.of()};
|
||||||
|
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||||
|
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler);
|
||||||
|
|
||||||
|
loop.tick(); // opens the idle window
|
||||||
|
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
|
||||||
|
|
||||||
|
loop.tick(); // INJECT, send succeeds -> latch set
|
||||||
|
assertEquals(1, herdr.sentTexts().size());
|
||||||
|
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"), herdr.sentTexts().get(0));
|
||||||
|
|
||||||
|
loop.tick(); // same inputs — the lead was already told this stretch
|
||||||
|
assertEquals(1, herdr.sentTexts().size(),
|
||||||
|
"the next tick with the same inputs must not send a second notice");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void kAPendingDrivenInjectWithTheLatchAlreadySetSendsNoNotice() {
|
||||||
|
var herdr = new FailableHerdrClient(LEAD);
|
||||||
|
var now = new AtomicLong(NOW);
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
List<MemberSession>[] rosterBox = new List[]{List.of()};
|
||||||
|
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||||
|
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler);
|
||||||
|
|
||||||
|
loop.tick();
|
||||||
|
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
|
||||||
|
loop.tick(); // latches the notice (quiet, HIGH, cap exhausted -> the forced context INJECT)
|
||||||
|
assertEquals(1, herdr.sentTexts().size());
|
||||||
|
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"));
|
||||||
|
|
||||||
|
// Now make the fleet have real pending state, so the NEXT INJECT is pending-driven, not the
|
||||||
|
// forced-context route — with the latch already set from the tick above.
|
||||||
|
inbox.own(WORKER);
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
rosterBox[0] = List.of(new MemberSession("p1", WORKER, "prof", MemberRole.DEV, "/cwd", null,
|
||||||
|
0, 0, 0, MemberSession.State.READY, null, null));
|
||||||
|
|
||||||
|
loop.tick();
|
||||||
|
|
||||||
|
assertEquals(2, herdr.sentTexts().size(), "the pending-driven tick must still send a nudge");
|
||||||
|
assertFalse(herdr.sentTexts().get(1).contains("Your own context is nearly full"),
|
||||||
|
"a pending-driven INJECT while the latch is already set must carry no context notice: "
|
||||||
|
+ herdr.sentTexts().get(1));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void lTheNoticeAppearsExactlyOnceAcrossThreeDifferentlyDrivenInjects() {
|
||||||
|
var herdr = new FailableHerdrClient(LEAD);
|
||||||
|
var now = new AtomicLong(NOW);
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
List<MemberSession>[] rosterBox = new List[]{List.of()};
|
||||||
|
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||||
|
// quietNudgeCap=1 so a still-not-exhausted quiet nudge is available as the "forced" route below,
|
||||||
|
// distinct from both the pending-driven route and the exhausted-cap forced-context route.
|
||||||
|
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 1, scheduler);
|
||||||
|
|
||||||
|
loop.tick(); // opens the idle window
|
||||||
|
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
|
||||||
|
|
||||||
|
// 1) pending-driven INJECT: real fleet state present. Sets the latch and carries the notice.
|
||||||
|
inbox.own(WORKER);
|
||||||
|
inbox.publish(WORKER, "m1", "hello");
|
||||||
|
rosterBox[0] = List.of(new MemberSession("p1", WORKER, "prof", MemberRole.DEV, "/cwd", null,
|
||||||
|
0, 0, 0, MemberSession.State.READY, null, null));
|
||||||
|
loop.tick();
|
||||||
|
assertEquals(1, herdr.sentTexts().size());
|
||||||
|
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"), herdr.sentTexts().get(0));
|
||||||
|
|
||||||
|
// 2) "forced" INJECT: nothing pending, but the quiet cap (1) is not yet exhausted, so decide()
|
||||||
|
// nudges anyway. The latch is already set, so no notice.
|
||||||
|
inbox.ack(WORKER, "m1");
|
||||||
|
rosterBox[0] = List.of();
|
||||||
|
loop.tick();
|
||||||
|
assertEquals(2, herdr.sentTexts().size(), "the quiet-cap-not-yet-exhausted nudge must still fire");
|
||||||
|
assertFalse(herdr.sentTexts().get(1).contains("Your own context is nearly full"), herdr.sentTexts().get(1));
|
||||||
|
|
||||||
|
// 3) pending-driven INJECT again. Still latched, still no notice.
|
||||||
|
inbox.publish(WORKER, "m2", "hello again");
|
||||||
|
rosterBox[0] = List.of(new MemberSession("p1", WORKER, "prof", MemberRole.DEV, "/cwd", null,
|
||||||
|
0, 0, 0, MemberSession.State.READY, null, null));
|
||||||
|
loop.tick();
|
||||||
|
assertEquals(3, herdr.sentTexts().size());
|
||||||
|
assertFalse(herdr.sentTexts().get(2).contains("Your own context is nearly full"), herdr.sentTexts().get(2));
|
||||||
|
|
||||||
|
long noticeCount = herdr.sentTexts().stream()
|
||||||
|
.filter(t -> t.contains("Your own context is nearly full")).count();
|
||||||
|
assertEquals(1, noticeCount,
|
||||||
|
"the notice text must appear exactly once across all three sends: " + herdr.sentTexts());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Fake herdr client for the four tests above: always reports {@code lead} as IDLE, records the
|
||||||
|
* {@code text} of every {@code agent.prompt} call, and can be told to throw on the very next
|
||||||
|
* {@code agent.prompt} call — standing in for one transient herdr send failure.
|
||||||
|
*/
|
||||||
|
private static final class FailableHerdrClient implements HerdrClient {
|
||||||
|
private static final ObjectMapper MAPPER = new ObjectMapper();
|
||||||
|
private final String lead;
|
||||||
|
private final List<String> sentTexts = new ArrayList<>();
|
||||||
|
private boolean throwOnNextSend = false;
|
||||||
|
|
||||||
|
FailableHerdrClient(String lead) {
|
||||||
|
this.lead = lead;
|
||||||
|
}
|
||||||
|
|
||||||
|
void throwOnNextSend() {
|
||||||
|
throwOnNextSend = true;
|
||||||
|
}
|
||||||
|
|
||||||
|
List<String> sentTexts() {
|
||||||
|
return List.copyOf(sentTexts);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
public JsonNode call(String method, Object params) {
|
||||||
|
if ("agent.get".equals(method)) {
|
||||||
|
return MAPPER.createObjectNode()
|
||||||
|
.set("agent", MAPPER.createObjectNode()
|
||||||
|
.put("terminal_id", lead)
|
||||||
|
.put("agent_status", "idle"));
|
||||||
|
}
|
||||||
|
if ("agent.prompt".equals(method)) {
|
||||||
|
if (throwOnNextSend) {
|
||||||
|
throwOnNextSend = false;
|
||||||
|
throw new RuntimeException("simulated transient herdr send failure");
|
||||||
|
}
|
||||||
|
Map<String, Object> p = params instanceof Map ? (Map<String, Object>) params : Map.of();
|
||||||
|
sentTexts.add(String.valueOf(p.get("text")));
|
||||||
|
}
|
||||||
|
return MAPPER.createObjectNode();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
}
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,168 @@
|
|||||||
|
package dev.ltms.fleet.msg;
|
||||||
|
|
||||||
|
import java.util.ArrayDeque;
|
||||||
|
import java.util.ArrayList;
|
||||||
|
import java.util.Collection;
|
||||||
|
import java.util.Deque;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.concurrent.Callable;
|
||||||
|
import java.util.concurrent.ExecutionException;
|
||||||
|
import java.util.concurrent.Future;
|
||||||
|
import java.util.concurrent.RejectedExecutionException;
|
||||||
|
import java.util.concurrent.ScheduledExecutorService;
|
||||||
|
import java.util.concurrent.ScheduledFuture;
|
||||||
|
import java.util.concurrent.TimeUnit;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* A {@link ScheduledExecutorService} for tests that never runs a task on a timer — it records what
|
||||||
|
* {@link ReplyPushLoop} schedules and only ever runs it when the test itself calls
|
||||||
|
* {@link #runDueTasks()} (fleetd #608).
|
||||||
|
*
|
||||||
|
* <p>The test this exists for ({@code anAlreadyCollectedTicketProducesNoNudge}) used to wire
|
||||||
|
* {@link ReplyPushLoop} to a real {@code Executors.newSingleThreadScheduledExecutor()} and bet a
|
||||||
|
* 300ms backoff was "wide" enough that the test's own work (collecting the ticket) always won the
|
||||||
|
* race against the scheduler's own timer firing the next tick. On an idle machine that held; under
|
||||||
|
* a loaded full-suite run it did not, and the test went red on perfectly correct code. Replacing
|
||||||
|
* the timer with this fake removes the race rather than widening it: nothing here ever runs on a
|
||||||
|
* schedule of its own, so no backoff value — however small — can make the tick fire before the test
|
||||||
|
* is ready for it.
|
||||||
|
*
|
||||||
|
* <p><strong>Deliberately narrow.</strong> {@link ReplyPushLoop} calls exactly two methods on its
|
||||||
|
* {@link ScheduledExecutorService} field — {@link #schedule(Runnable, long, TimeUnit)} (every tick,
|
||||||
|
* including the very first one {@code startOrCoalesce} kicks off) and {@link #shutdownNow()} (on
|
||||||
|
* {@code ReplyPushLoop.stop()}/{@code close()}) — verified by reading every {@code scheduler.} call
|
||||||
|
* site in that class. Every other {@link ScheduledExecutorService} method throws
|
||||||
|
* {@link UnsupportedOperationException} rather than silently doing the wrong thing, so a future
|
||||||
|
* change to {@code ReplyPushLoop} that starts calling one of them fails this fake loudly, on the
|
||||||
|
* very first test that exercises it, instead of being quietly mishandled.
|
||||||
|
*/
|
||||||
|
final class ManualScheduler implements ScheduledExecutorService {
|
||||||
|
|
||||||
|
private final Deque<Runnable> pending = new ArrayDeque<>();
|
||||||
|
private boolean shutdown = false;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Run every task pending as of the START of this call — a snapshot taken before anything runs.
|
||||||
|
* {@link ReplyPushLoop#tick} always reschedules its own next tick before returning (see
|
||||||
|
* {@code scheduleNext} in its {@code INJECT}/{@code WAIT_BUSY} branches), so without the
|
||||||
|
* snapshot a single call here would recurse forever. Taking it up front means one call is
|
||||||
|
* exactly one tick, deterministically, no matter what that tick itself goes on to schedule.
|
||||||
|
*
|
||||||
|
* @return how many tasks actually ran
|
||||||
|
*/
|
||||||
|
synchronized int runDueTasks() {
|
||||||
|
List<Runnable> due = new ArrayList<>(pending);
|
||||||
|
pending.clear();
|
||||||
|
for (Runnable task : due) {
|
||||||
|
task.run();
|
||||||
|
}
|
||||||
|
return due.size();
|
||||||
|
}
|
||||||
|
|
||||||
|
/** How many tasks are currently queued, without running any of them. */
|
||||||
|
synchronized int pendingCount() {
|
||||||
|
return pending.size();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public synchronized ScheduledFuture<?> schedule(Runnable command, long delay, TimeUnit unit) {
|
||||||
|
if (shutdown) {
|
||||||
|
throw new RejectedExecutionException("ManualScheduler is shut down");
|
||||||
|
}
|
||||||
|
pending.add(command);
|
||||||
|
return null; // ReplyPushLoop discards the return value of every schedule() call it makes.
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public synchronized List<Runnable> shutdownNow() {
|
||||||
|
shutdown = true;
|
||||||
|
List<Runnable> left = new ArrayList<>(pending);
|
||||||
|
pending.clear();
|
||||||
|
return left;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public synchronized boolean isShutdown() {
|
||||||
|
return shutdown;
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- everything below: not called by ReplyPushLoop, and not supported by this fake (see the
|
||||||
|
// class javadoc's "deliberately narrow" note) -------------------------------------------------
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledFuture<?> scheduleAtFixedRate(Runnable command, long initialDelay, long period, TimeUnit unit) {
|
||||||
|
throw unsupported();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ScheduledFuture<?> scheduleWithFixedDelay(Runnable command, long initialDelay, long delay, TimeUnit unit) {
|
||||||
|
throw unsupported();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public <V> ScheduledFuture<V> schedule(Callable<V> callable, long delay, TimeUnit unit) {
|
||||||
|
throw unsupported();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void shutdown() {
|
||||||
|
throw unsupported();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public boolean isTerminated() {
|
||||||
|
throw unsupported();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public boolean awaitTermination(long timeout, TimeUnit unit) {
|
||||||
|
throw unsupported();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public <T> Future<T> submit(Callable<T> task) {
|
||||||
|
throw unsupported();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public <T> Future<T> submit(Runnable task, T result) {
|
||||||
|
throw unsupported();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Future<?> submit(Runnable task) {
|
||||||
|
throw unsupported();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public <T> List<Future<T>> invokeAll(Collection<? extends Callable<T>> tasks) {
|
||||||
|
throw unsupported();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public <T> List<Future<T>> invokeAll(Collection<? extends Callable<T>> tasks, long timeout, TimeUnit unit) {
|
||||||
|
throw unsupported();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public <T> T invokeAny(Collection<? extends Callable<T>> tasks) throws ExecutionException {
|
||||||
|
throw unsupported();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public <T> T invokeAny(Collection<? extends Callable<T>> tasks, long timeout, TimeUnit unit)
|
||||||
|
throws ExecutionException {
|
||||||
|
throw unsupported();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void execute(Runnable command) {
|
||||||
|
throw unsupported();
|
||||||
|
}
|
||||||
|
|
||||||
|
private static UnsupportedOperationException unsupported() {
|
||||||
|
return new UnsupportedOperationException(
|
||||||
|
"ManualScheduler only supports schedule(Runnable, long, TimeUnit), shutdownNow() and isShutdown() "
|
||||||
|
+ "— see the class javadoc");
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1841,6 +1841,43 @@ class MessageServiceTest {
|
|||||||
return new PushWiring(service, leadHerdr, scheduler);
|
return new PushWiring(service, leadHerdr, scheduler);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* A {@link MessageService} wired to a real {@link ReplyPushLoop}, like {@link #wireWithPushLoop},
|
||||||
|
* but backed by {@link ManualScheduler} instead of a real timer (fleetd #608): its tick never
|
||||||
|
* fires on its own — a test drives it explicitly via {@link ManualScheduler#runDueTasks()}.
|
||||||
|
*/
|
||||||
|
private record ManualPushWiring(MessageService service, FakeHerdr leadHerdr, ManualScheduler scheduler,
|
||||||
|
ReplyPushLoop pushLoop)
|
||||||
|
implements AutoCloseable {
|
||||||
|
@Override
|
||||||
|
public void close() {
|
||||||
|
scheduler.shutdownNow();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* As {@link #wireWithPushLoop}, but the schedule never runs on its own. {@code backoffMs} is
|
||||||
|
* still threaded through to {@link ReplyPushLoop}'s constructor (it takes one), but nothing here
|
||||||
|
* ever waits it out, so its value cannot affect anything a test built on this observes — see
|
||||||
|
* {@code anAlreadyCollectedTicketProducesNoNudge}, which sets it to 1 to prove exactly that.
|
||||||
|
*/
|
||||||
|
private ManualPushWiring wireWithManualScheduler(int maxReminders, long backoffMs) {
|
||||||
|
return wireWithManualScheduler(maxReminders, backoffMs, System::nanoTime);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** As above, with an injectable clock for tests that exercise terminal-ticket pruning. */
|
||||||
|
private ManualPushWiring wireWithManualScheduler(int maxReminders, long backoffMs,
|
||||||
|
java.util.function.LongSupplier nowNanos) {
|
||||||
|
PrimaryRegistry registry = new PrimaryRegistry(null);
|
||||||
|
registry.recordDelegation(T, LEAD);
|
||||||
|
FakeHerdr leadHerdr = new FakeHerdr();
|
||||||
|
AgentControl leadAgents = new AgentControl(leadHerdr);
|
||||||
|
ManualScheduler scheduler = new ManualScheduler();
|
||||||
|
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, leadAgents, inbox, scheduler, maxReminders, backoffMs);
|
||||||
|
MessageService service = new MessageService(agents, injector, rendezvous, inbox, pushLoop, null, nowNanos);
|
||||||
|
return new ManualPushWiring(service, leadHerdr, scheduler, pushLoop);
|
||||||
|
}
|
||||||
|
|
||||||
private void awaitNudge(FakeHerdr leadHerdr) throws InterruptedException {
|
private void awaitNudge(FakeHerdr leadHerdr) throws InterruptedException {
|
||||||
long deadline = System.currentTimeMillis() + 3000;
|
long deadline = System.currentTimeMillis() + 3000;
|
||||||
while (!leadHerdr.called("agent.prompt") && System.currentTimeMillis() < deadline) {
|
while (!leadHerdr.called("agent.prompt") && System.currentTimeMillis() < deadline) {
|
||||||
@@ -1882,16 +1919,46 @@ class MessageServiceTest {
|
|||||||
|
|
||||||
@Test
|
@Test
|
||||||
void anAlreadyCollectedTicketProducesNoNudge() throws Exception {
|
void anAlreadyCollectedTicketProducesNoNudge() throws Exception {
|
||||||
try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: poll before the first tick fires
|
// fleetd #608: backoff is 1ms — the most hostile value there is, the tick due immediately —
|
||||||
String ticket = wiring.service().sendAsync(T, "long task");
|
// and the test still must pass, because with ManualScheduler the tick never runs on a timer
|
||||||
awaitWaiting();
|
// at all; it only runs when this test calls runDueTasks() below. The old version bet a 300ms
|
||||||
injectDelivery();
|
// backoff was "wide" enough that collecting the ticket always won the race against a real
|
||||||
assertTrue(rendezvous.resolve(T, "async result"));
|
// scheduler's own timer; that held on an idle machine and failed under a loaded full-suite
|
||||||
|
// run — a false red on correct code, since nothing here forced that ordering, it only made
|
||||||
|
// it likely.
|
||||||
|
try (var wiring = wireWithManualScheduler(1, 1)) {
|
||||||
|
java.util.concurrent.CountDownLatch terminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||||
|
// finishAsyncTask's task.future.complete(...) runs every whenComplete registered on that
|
||||||
|
// same future — including sendAsync's own hook that calls ReplyPushLoop.onTicketTerminal
|
||||||
|
// — synchronously, before complete() returns (see finishAsyncTask's javadoc and fleetd
|
||||||
|
// #399). This test-only hook fires right after that complete() call, on the very same
|
||||||
|
// (async executor) thread, so waiting for it guarantees onTicketTerminal has already run
|
||||||
|
// and the ticket is really sitting in the push loop's pending set — unlike waiting for
|
||||||
|
// Phase.DONE via poll(), which CompletableFuture.complete() can make visible to another
|
||||||
|
// thread before every whenComplete dependent has actually finished running (the same
|
||||||
|
// publish-then-run-dependents gap awaitCompletionStamped exists to close elsewhere).
|
||||||
|
String ticket;
|
||||||
|
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(terminalReached::countDown);
|
||||||
|
try {
|
||||||
|
ticket = wiring.service().sendAsync(T, "long task");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
assertTrue(rendezvous.resolve(T, "async result"));
|
||||||
|
assertTrue(terminalReached.await(5, TimeUnit.SECONDS),
|
||||||
|
"the ticket never reached its terminal phase");
|
||||||
|
} finally {
|
||||||
|
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Collect it — the exact action the nudge must never follow.
|
||||||
MessageService.TaskView view = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.DONE);
|
MessageService.TaskView view = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.DONE);
|
||||||
assertEquals(MessageService.Phase.DONE, view.phase());
|
assertEquals(MessageService.Phase.DONE, view.phase());
|
||||||
|
|
||||||
Thread.sleep(400); // let the scheduled tick run — it must find nothing pending
|
// Now run the pending tick explicitly. onTicketTerminal scheduled exactly one (the ticket
|
||||||
|
// is the only thing this lead has ever had pending); it must find nothing left pending —
|
||||||
|
// the collection above already removed it — and send no nudge.
|
||||||
|
assertEquals(1, wiring.scheduler().runDueTasks(),
|
||||||
|
"expected exactly the one tick onTicketTerminal scheduled");
|
||||||
assertFalse(wiring.leadHerdr().called("agent.prompt"),
|
assertFalse(wiring.leadHerdr().called("agent.prompt"),
|
||||||
"a ticket the lead already polled must never be nudged");
|
"a ticket the lead already polled must never be nudged");
|
||||||
}
|
}
|
||||||
@@ -1899,24 +1966,35 @@ class MessageServiceTest {
|
|||||||
|
|
||||||
@Test
|
@Test
|
||||||
void severalAsyncTicketsFinishingTogetherProduceOneCoalescedNudge() throws Exception {
|
void severalAsyncTicketsFinishingTogetherProduceOneCoalescedNudge() throws Exception {
|
||||||
try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: both tickets land before the tick fires
|
// A 1ms backoff is due immediately. ManualScheduler still cannot run it until this test
|
||||||
String first = wiring.service().sendAsync(T, "first task");
|
// explicitly calls runDueTasks(), so both terminal tickets join one scheduled tick.
|
||||||
awaitWaiting();
|
try (var wiring = wireWithManualScheduler(1, 1)) {
|
||||||
injectDelivery();
|
String first;
|
||||||
assertTrue(rendezvous.resolve(T, "first done"));
|
String second;
|
||||||
// Settle without polling: poll() itself marks a ticket collected (that's the point of
|
java.util.concurrent.CountDownLatch firstTerminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||||
// anAlreadyCollectedTicketProducesNoNudge above) — using it here to detect completion
|
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(firstTerminalReached::countDown);
|
||||||
// would collect the ticket before the coalescing this test checks ever gets a chance.
|
try {
|
||||||
Thread.sleep(100);
|
first = wiring.service().sendAsync(T, "first task");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
assertTrue(rendezvous.resolve(T, "first done"));
|
||||||
|
assertTrue(firstTerminalReached.await(5, TimeUnit.SECONDS),
|
||||||
|
"the first ticket never reached its terminal phase");
|
||||||
|
|
||||||
String second = wiring.service().sendAsync(T, "second task");
|
java.util.concurrent.CountDownLatch secondTerminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||||
awaitWaiting();
|
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(secondTerminalReached::countDown);
|
||||||
injectDelivery();
|
second = wiring.service().sendAsync(T, "second task");
|
||||||
assertTrue(rendezvous.resolve(T, "second done"));
|
awaitWaiting();
|
||||||
Thread.sleep(100);
|
injectDelivery();
|
||||||
|
assertTrue(rendezvous.resolve(T, "second done"));
|
||||||
|
assertTrue(secondTerminalReached.await(5, TimeUnit.SECONDS),
|
||||||
|
"the second ticket never reached its terminal phase");
|
||||||
|
} finally {
|
||||||
|
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
|
||||||
|
}
|
||||||
|
|
||||||
awaitNudge(wiring.leadHerdr());
|
assertEquals(1, wiring.scheduler().runDueTasks(),
|
||||||
Thread.sleep(200); // settle — nothing more should arrive beyond the one coalesced nudge
|
"both terminal tickets must coalesce onto one scheduled tick");
|
||||||
long nudgeCount = wiring.leadHerdr().calls.stream()
|
long nudgeCount = wiring.leadHerdr().calls.stream()
|
||||||
.filter(c -> c.method().equals("agent.prompt")).count();
|
.filter(c -> c.method().equals("agent.prompt")).count();
|
||||||
assertEquals(1, nudgeCount, "two tickets finishing together must produce ONE nudge, not two");
|
assertEquals(1, nudgeCount, "two tickets finishing together must produce ONE nudge, not two");
|
||||||
@@ -1995,7 +2073,8 @@ class MessageServiceTest {
|
|||||||
|
|
||||||
@Test
|
@Test
|
||||||
void answeringAQuestionStopsFurtherNudgesAboutIt() throws Exception {
|
void answeringAQuestionStopsFurtherNudgesAboutIt() throws Exception {
|
||||||
try (var wiring = wireWithPushLoop(5, 50)) {
|
// A 1ms backoff is due immediately, but ManualScheduler only ticks when this test asks it to.
|
||||||
|
try (var wiring = wireWithManualScheduler(5, 1)) {
|
||||||
String ticket = wiring.service().sendAsync(T, "task that asks");
|
String ticket = wiring.service().sendAsync(T, "task that asks");
|
||||||
awaitWaiting();
|
awaitWaiting();
|
||||||
injectDelivery();
|
injectDelivery();
|
||||||
@@ -2003,8 +2082,9 @@ class MessageServiceTest {
|
|||||||
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
|
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
|
||||||
() -> wiring.service().ask(T, "which config file?", 5000));
|
() -> wiring.service().ask(T, "which config file?", 5000));
|
||||||
MessageService.TaskView asking = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.ASKING);
|
MessageService.TaskView asking = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.ASKING);
|
||||||
|
awaitQuestionPendingOn(wiring.pushLoop(), LEAD, asking.turnId());
|
||||||
|
|
||||||
awaitNudge(wiring.leadHerdr());
|
assertEquals(1, wiring.scheduler().runDueTasks(), "the open question must have one scheduled tick");
|
||||||
long callsBeforeAnswer = wiring.leadHerdr().calls.stream()
|
long callsBeforeAnswer = wiring.leadHerdr().calls.stream()
|
||||||
.filter(c -> c.method().equals("agent.prompt")).count();
|
.filter(c -> c.method().equals("agent.prompt")).count();
|
||||||
|
|
||||||
@@ -2012,12 +2092,22 @@ class MessageServiceTest {
|
|||||||
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
|
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
|
||||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||||
awaitWaiting();
|
awaitWaiting();
|
||||||
|
java.util.concurrent.CountDownLatch terminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||||
|
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(terminalReached::countDown);
|
||||||
assertTrue(rendezvous.resolve(T, "done"));
|
assertTrue(rendezvous.resolve(T, "done"));
|
||||||
|
try {
|
||||||
|
assertTrue(terminalReached.await(5, TimeUnit.SECONDS),
|
||||||
|
"the answered ticket never reached its terminal phase");
|
||||||
|
} finally {
|
||||||
|
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
|
||||||
|
}
|
||||||
answer.get(5, TimeUnit.SECONDS);
|
answer.get(5, TimeUnit.SECONDS);
|
||||||
|
|
||||||
// Let several more ticks (and the ticket's own now-legitimate terminal nudge) fire —
|
// Run the ticket's legitimate terminal nudge and every later scheduled tick through its
|
||||||
// none of them may still name the question's turnId, which is closed.
|
// reminder cap. None may still name the closed question.
|
||||||
Thread.sleep(300);
|
for (int tick = 0; tick < 6; tick++) {
|
||||||
|
assertEquals(1, wiring.scheduler().runDueTasks(), "expected one scheduled reminder tick");
|
||||||
|
}
|
||||||
boolean anyNamesClosedQuestion = wiring.leadHerdr().calls.stream()
|
boolean anyNamesClosedQuestion = wiring.leadHerdr().calls.stream()
|
||||||
.filter(c -> c.method().equals("agent.prompt"))
|
.filter(c -> c.method().equals("agent.prompt"))
|
||||||
.skip(callsBeforeAnswer)
|
.skip(callsBeforeAnswer)
|
||||||
@@ -2100,35 +2190,45 @@ class MessageServiceTest {
|
|||||||
// decideTickets hits the cap and STOPs — activeLeads drops the lead, but (before the fix)
|
// decideTickets hits the cap and STOPs — activeLeads drops the lead, but (before the fix)
|
||||||
// pendingTickets never drops the ticket. That is the exact "cap already STOPped" branch of
|
// pendingTickets never drops the ticket. That is the exact "cap already STOPped" branch of
|
||||||
// the bug report, reached deterministically rather than by timing it against a live tick.
|
// the bug report, reached deterministically rather than by timing it against a live tick.
|
||||||
try (var wiring = wireWithPushLoop(1, 50, clock::get)) {
|
// A 1ms backoff is due immediately, but ManualScheduler runs only the ticks below.
|
||||||
String stale = wiring.service().sendAsync(T, "first task");
|
try (var wiring = wireWithManualScheduler(1, 1, clock::get)) {
|
||||||
awaitWaiting();
|
String stale;
|
||||||
injectDelivery();
|
String fresh;
|
||||||
assertTrue(rendezvous.resolve(T, "stale result"));
|
java.util.concurrent.CountDownLatch staleTerminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||||
|
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(staleTerminalReached::countDown);
|
||||||
|
try {
|
||||||
|
stale = wiring.service().sendAsync(T, "first task");
|
||||||
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
|
assertTrue(rendezvous.resolve(T, "stale result"));
|
||||||
|
assertTrue(staleTerminalReached.await(5, TimeUnit.SECONDS),
|
||||||
|
"the stale ticket never reached its terminal phase");
|
||||||
|
|
||||||
// Let the reminder loop fire its one nudge and hit the cap (STOP removes it from
|
// Fire the stale ticket's one nudge, then its cap tick (STOP removes it from activeLeads;
|
||||||
// activeLeads; pendingTickets is untouched either way — that asymmetry is the bug).
|
// pendingTickets is untouched either way — that asymmetry is the bug).
|
||||||
awaitNudge(wiring.leadHerdr());
|
assertEquals(1, wiring.scheduler().runDueTasks(), "the stale ticket must have one nudge tick");
|
||||||
Thread.sleep(300);
|
assertEquals(1, wiring.scheduler().runDueTasks(), "the stale ticket must have one cap tick");
|
||||||
assertTrue(wiring.leadHerdr().lastCall("agent.prompt").params().toString().contains(stale),
|
assertTrue(wiring.leadHerdr().lastCall("agent.prompt").params().toString().contains(stale),
|
||||||
"sanity: the stale ticket's own reminder must have fired first");
|
"sanity: the stale ticket's own reminder must have fired first");
|
||||||
|
|
||||||
// Cross the TTL — a real clock would need 10 minutes; the injected one does it instantly.
|
// Cross the TTL — a real clock would need 10 minutes; the injected one does it instantly.
|
||||||
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
|
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
|
||||||
|
|
||||||
// A second, unrelated ticket to the same target/lead reaches sendAsync, which prunes.
|
// A second, unrelated ticket to the same target/lead reaches sendAsync, which prunes.
|
||||||
String fresh = wiring.service().sendAsync(T, "second task");
|
java.util.concurrent.CountDownLatch freshTerminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||||
awaitWaiting();
|
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(freshTerminalReached::countDown);
|
||||||
injectDelivery();
|
fresh = wiring.service().sendAsync(T, "second task");
|
||||||
assertTrue(rendezvous.resolve(T, "fresh result"));
|
awaitWaiting();
|
||||||
|
injectDelivery();
|
||||||
// The fresh ticket restarts the (now-dormant) reminder loop with its own nudge.
|
assertTrue(rendezvous.resolve(T, "fresh result"));
|
||||||
long before = wiring.leadHerdr().calls.stream().filter(c -> c.method().equals("agent.prompt")).count();
|
assertTrue(freshTerminalReached.await(5, TimeUnit.SECONDS),
|
||||||
long deadline = System.currentTimeMillis() + 3000;
|
"the fresh ticket never reached its terminal phase");
|
||||||
while (wiring.leadHerdr().calls.stream().filter(c -> c.method().equals("agent.prompt")).count() <= before
|
} finally {
|
||||||
&& System.currentTimeMillis() < deadline) {
|
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
|
||||||
Thread.sleep(10);
|
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// The fresh ticket restarts the now-dormant reminder loop with its own nudge.
|
||||||
|
assertEquals(1, wiring.scheduler().runDueTasks(), "the fresh ticket must have one nudge tick");
|
||||||
String latestNudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
|
String latestNudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
|
||||||
assertTrue(latestNudge.contains(fresh), "the fresh ticket's nudge must still arrive: " + latestNudge);
|
assertTrue(latestNudge.contains(fresh), "the fresh ticket's nudge must still arrive: " + latestNudge);
|
||||||
assertFalse(latestNudge.contains(stale),
|
assertFalse(latestNudge.contains(stale),
|
||||||
|
|||||||
@@ -10,12 +10,13 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
|
|||||||
class MemberRoleTest {
|
class MemberRoleTest {
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void theThreeRolesAreArchitectDevAndReviewer() {
|
void theFourRolesAreArchitectDevHunterAndReviewer() {
|
||||||
assertEquals(3, MemberRole.values().length,
|
assertEquals(4, MemberRole.values().length,
|
||||||
"a new role changes the charter, the role file, the skill and the authz row — "
|
"a new role changes the charter, the role file, the skill and the authz row — "
|
||||||
+ "adding one is a deliberate act, so this count is meant to fail first");
|
+ "adding one is a deliberate act, so this count is meant to fail first");
|
||||||
assertEquals("architect", MemberRole.ARCHITECT.wireName());
|
assertEquals("architect", MemberRole.ARCHITECT.wireName());
|
||||||
assertEquals("dev", MemberRole.DEV.wireName());
|
assertEquals("dev", MemberRole.DEV.wireName());
|
||||||
|
assertEquals("hunter", MemberRole.HUNTER.wireName());
|
||||||
assertEquals("reviewer", MemberRole.REVIEWER.wireName());
|
assertEquals("reviewer", MemberRole.REVIEWER.wireName());
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -30,6 +31,7 @@ class MemberRoleTest {
|
|||||||
void parseIsCaseInsensitiveAndTrimsSurroundingSpace() {
|
void parseIsCaseInsensitiveAndTrimsSurroundingSpace() {
|
||||||
assertSame(MemberRole.ARCHITECT, MemberRole.parse("Architect"));
|
assertSame(MemberRole.ARCHITECT, MemberRole.parse("Architect"));
|
||||||
assertSame(MemberRole.DEV, MemberRole.parse(" DEV "));
|
assertSame(MemberRole.DEV, MemberRole.parse(" DEV "));
|
||||||
|
assertSame(MemberRole.HUNTER, MemberRole.parse("HuNtEr"));
|
||||||
assertSame(MemberRole.REVIEWER, MemberRole.parse("ReViEwEr"));
|
assertSame(MemberRole.REVIEWER, MemberRole.parse("ReViEwEr"));
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -38,7 +40,7 @@ class MemberRoleTest {
|
|||||||
IllegalArgumentException e =
|
IllegalArgumentException e =
|
||||||
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse("archtiect"));
|
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse("archtiect"));
|
||||||
assertTrue(e.getMessage().contains("archtiect"), e.getMessage());
|
assertTrue(e.getMessage().contains("archtiect"), e.getMessage());
|
||||||
assertTrue(e.getMessage().contains("architect, dev, reviewer"),
|
assertTrue(e.getMessage().contains("architect, dev, hunter, reviewer"),
|
||||||
"a typo in config should be fixable from the message alone: " + e.getMessage());
|
"a typo in config should be fixable from the message alone: " + e.getMessage());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -328,9 +328,10 @@ class SessionManagerTest {
|
|||||||
void rosterViewExposesTheCharterReceiptButNeverTheCharterText() {
|
void rosterViewExposesTheCharterReceiptButNeverTheCharterText() {
|
||||||
// The roster (fleet_list and GET /members both render through rosterView) must let a lead
|
// The roster (fleet_list and GET /members both render through rosterView) must let a lead
|
||||||
// see which charter a member got, without ever carrying the charter prose itself (CB-571).
|
// see which charter a member got, without ever carrying the charter prose itself (CB-571).
|
||||||
|
String composed = "role charter\n\nreply";
|
||||||
|
CharterReceipt receipt = CharterReceipt.compose(MemberRole.DEV, "prof", "role charter", composed);
|
||||||
MemberSession s = new MemberSession("p1", "term1", "prof", MemberRole.DEV, "/cwd", null,
|
MemberSession s = new MemberSession("p1", "term1", "prof", MemberRole.DEV, "/cwd", null,
|
||||||
0, 0, 0, MemberSession.State.READY, null, null,
|
0, 0, 0, MemberSession.State.READY, null, null, receipt, null);
|
||||||
CharterReceipt.compose(MemberRole.DEV, "prof", "role charter", "role charter\n\nreply"), null);
|
|
||||||
|
|
||||||
Map<String, Object> view = SessionManager.rosterView(s, null);
|
Map<String, Object> view = SessionManager.rosterView(s, null);
|
||||||
|
|
||||||
@@ -338,10 +339,49 @@ class SessionManagerTest {
|
|||||||
"the config key that supplied the role charter is reported");
|
"the config key that supplied the role charter is reported");
|
||||||
assertEquals(CharterReceipt.digestOf("role charter\n\nreply"), view.get("charterSha256"),
|
assertEquals(CharterReceipt.digestOf("role charter\n\nreply"), view.get("charterSha256"),
|
||||||
"the digest of the exact composed charter bytes is reported");
|
"the digest of the exact composed charter bytes is reported");
|
||||||
|
// #604: the byte count rides alongside the digest, and must match what the receipt itself
|
||||||
|
// carries (not a hardcoded literal) so a bug that reads the wrong field is caught.
|
||||||
|
assertEquals(receipt.charterBytes(), view.get("charterBytes"),
|
||||||
|
"the exact composed byte count is reported, read from the receipt");
|
||||||
|
assertEquals(composed.getBytes(java.nio.charset.StandardCharsets.UTF_8).length, view.get("charterBytes"),
|
||||||
|
"the byte count is the real UTF-8 length of the composed charter");
|
||||||
assertFalse(view.values().toString().contains("role charter"),
|
assertFalse(view.values().toString().contains("role charter"),
|
||||||
"the roster row must not embed the charter text itself");
|
"the roster row must not embed the charter text itself");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void rosterViewOmitsCharterBytesAndDigestWhenNoCharterWasComposed() {
|
||||||
|
// #604: a member with no role charter and no reply charter (composed == null) still reports
|
||||||
|
// charterSource ("none"), but neither a digest nor a size — the digest is absent, and the
|
||||||
|
// size only ever accompanies a real digest. This must not be satisfiable by code that always
|
||||||
|
// writes charterBytes.
|
||||||
|
CharterReceipt receipt = CharterReceipt.compose(MemberRole.DEV, "prof", null, null);
|
||||||
|
MemberSession s = new MemberSession("p1", "term1", "prof", MemberRole.DEV, "/cwd", null,
|
||||||
|
0, 0, 0, MemberSession.State.READY, null, null, receipt, null);
|
||||||
|
|
||||||
|
Map<String, Object> view = SessionManager.rosterView(s, null);
|
||||||
|
|
||||||
|
assertEquals(CharterReceipt.NO_SOURCE, view.get("charterSource"),
|
||||||
|
"no configured role or reply charter reports the explicit \"none\" source");
|
||||||
|
assertFalse(view.containsKey("charterSha256"), "no digest is reported when no charter was composed");
|
||||||
|
assertFalse(view.containsKey("charterBytes"), "no byte count is reported when no charter was composed");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void rosterViewOmitsCharterFieldsEntirelyWhenTheReceiptItselfIsAbsent() {
|
||||||
|
// #604 acceptance criterion 3: charterReceipt() can be null on its own (a session recorded
|
||||||
|
// before CB-571, or a launcher that never composed one) — the outer null-guard must still
|
||||||
|
// suppress charterSource, charterSha256 AND charterBytes together.
|
||||||
|
MemberSession s = new MemberSession("p1", "term1", "prof", MemberRole.DEV, "/cwd", null,
|
||||||
|
0, 0, 0, MemberSession.State.READY, null, null, null, null);
|
||||||
|
|
||||||
|
Map<String, Object> view = SessionManager.rosterView(s, null);
|
||||||
|
|
||||||
|
assertFalse(view.containsKey("charterSource"), "no charter fields at all when the receipt is null");
|
||||||
|
assertFalse(view.containsKey("charterSha256"), "no charter fields at all when the receipt is null");
|
||||||
|
assertFalse(view.containsKey("charterBytes"), "no charter fields at all when the receipt is null");
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void aNullTerminalFromThePrimaryIsANoOpEvenWithSessionsRegistered() {
|
void aNullTerminalFromThePrimaryIsANoOpEvenWithSessionsRegistered() {
|
||||||
// The primary resolves to a Principal with no terminal, and FleetMcp's context extractor
|
// The primary resolves to a Principal with no terminal, and FleetMcp's context extractor
|
||||||
|
|||||||
Executable
+732
@@ -0,0 +1,732 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
#
|
||||||
|
# The one auditable way to edit the live fleetd.yaml.
|
||||||
|
#
|
||||||
|
# fleetd ticket #635 — why this exists at all: fleetd.yaml is gitignored and holds the live
|
||||||
|
# fleet's settings. A bad raw edit reaches a daemon that is already serving, so a direct `Edit`
|
||||||
|
# on it is refused by policy. This script is the allow-listed alternative, and it is not just
|
||||||
|
# convenience — it is the thing a raw file write can never give you: a backup, a parse check
|
||||||
|
# BEFORE the file is installed, and the daemon's own reload verdict read back afterwards. An
|
||||||
|
# edit to a live config is not finished when the bytes are written. It is finished when the
|
||||||
|
# daemon has said what it did with them.
|
||||||
|
#
|
||||||
|
# What the daemon says, and how this script finds it — measured against `ConfigRef.java` on
|
||||||
|
# fleetd commit 158a2a8, 2026-10-01:
|
||||||
|
#
|
||||||
|
# 1. `ConfigRef` re-reads fleetd.yaml only when the WATCHER sees the mtime move (every 10s by
|
||||||
|
# default — read the real interval out of the daemon's own startup line, "config watch: ...
|
||||||
|
# re-read when it changes (every Ns)"). So a verdict never appears before the next tick.
|
||||||
|
# 2. `ConfigRef.Outcome.summary()` logs exactly one of five strings (ConfigRef.java:371-391):
|
||||||
|
# config reload refused — <error message>
|
||||||
|
# config reload refused — these keys cannot change under a running daemon: <keys>. ...
|
||||||
|
# config reloaded
|
||||||
|
# config reloaded; these changes need a restart to take effect: <keys>
|
||||||
|
# config reloaded; partially live — <key: detail | ...>
|
||||||
|
# A parse/validation failure logs a DIFFERENT line instead, before any summary ever runs
|
||||||
|
# (ConfigRef.java:425): "config reload from <path> refused, keeping the running config:
|
||||||
|
# <message>". This script recognises both shapes of refusal.
|
||||||
|
# 3. The em dash in those strings is a real multi-byte character — match the stable prefix
|
||||||
|
# "config reload refused" (or "...refused, keeping the running config" for the parse-failure
|
||||||
|
# shape), never the dash itself.
|
||||||
|
# 4. A cold-key change (bind/herdrSocket/memberHerdrSocket/broker/auth) throws away the WHOLE
|
||||||
|
# reload — the running config keeps every old value, not only the cold one.
|
||||||
|
# 5. A deferred/split change IS applied (current.set(fresh) runs) — "needs a restart" is a
|
||||||
|
# SUCCESS with a follow-up, never a failure.
|
||||||
|
#
|
||||||
|
# Four outcomes, and they stay four (see the exit code table below). The one most likely to be
|
||||||
|
# gotten wrong is "cannot tell" (exit 5): the daemon may be down, or the watcher may be stalled,
|
||||||
|
# and folding that into either "refused" or "applied" is worse than never checking at all,
|
||||||
|
# because a caller then acts on a verdict nobody actually read. So exit 5 never restores — a
|
||||||
|
# visible, recoverable edit beats an invisible revert of a GOOD edit.
|
||||||
|
#
|
||||||
|
# Usage:
|
||||||
|
# scripts/config-edit.sh --check
|
||||||
|
# scripts/config-edit.sh --set <yq-path>=<value> [--set ...]
|
||||||
|
# scripts/config-edit.sh --from <candidate.yaml>
|
||||||
|
# scripts/config-edit.sh --dry-run --set <yq-path>=<value>
|
||||||
|
# scripts/config-edit.sh --restore
|
||||||
|
#
|
||||||
|
# `--set .a.b=` (an empty value — a forgotten typo) is REFUSED, not accepted as "clear the
|
||||||
|
# field": a null value falls back to its default rather than erroring, which is silent, not
|
||||||
|
# safe. To clear a key on purpose, write a literal null: `--set .a.b=null`. Every other value
|
||||||
|
# is always written as a YAML string (via yq's strenv(), never spliced into the expression), so
|
||||||
|
# there is currently no --set spelling for the literal three-character STRING "null" itself — use
|
||||||
|
# --from for that rare case.
|
||||||
|
#
|
||||||
|
# Overrides (so this is drivable with no daemon — see scripts/test-config-edit.sh):
|
||||||
|
# --config <path> default: fleetd/fleetd.yaml
|
||||||
|
# --log <path> default: fleetd/fleetd.out
|
||||||
|
# --wait-seconds <n> default: 4x the watch interval this script reads out of --log (10 -> 40)
|
||||||
|
#
|
||||||
|
# Exit codes (the --check/--restore/usage-error paths are reported separately, see below):
|
||||||
|
# 0 applied; verdict read; clean
|
||||||
|
# 3 applied; verdict read; needs a restart (deferred or split keys named)
|
||||||
|
# 4 REFUSED by the daemon; backup restored (and the restore's own verdict reported if seen)
|
||||||
|
# 5 CANNOT TELL — no verdict line inside the wait window. Nothing is restored.
|
||||||
|
#
|
||||||
|
# Never prints a secret. fleetd.yaml keeps credentials out by indirection (broker.uriEnv,
|
||||||
|
# gitTokenEnv) but this script does not rely on that staying true: every diff it prints is piped
|
||||||
|
# through `redact`, which (a) blanks the userinfo of any `scheme://user:pass@host` and (b) masks
|
||||||
|
# the whole value on any line whose key looks like a credential. See `redact` below.
|
||||||
|
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||||
|
SELF="$REPO/scripts/config-edit.sh"
|
||||||
|
|
||||||
|
CONFIG="$REPO/fleetd/fleetd.yaml"
|
||||||
|
LOG="$REPO/fleetd/fleetd.out"
|
||||||
|
WAIT_SECONDS_OVERRIDE=""
|
||||||
|
FALLBACK_PORT=8765
|
||||||
|
|
||||||
|
MODE=""
|
||||||
|
DRY_RUN=0
|
||||||
|
SETS=()
|
||||||
|
FROM_FILE=""
|
||||||
|
|
||||||
|
# fleetd #635 follow-up — a signal (or any early exit while a candidate is still uninstalled) must
|
||||||
|
# not leave a `.config-edit.XXXXXX` file sitting beside the live config forever. CAND is global
|
||||||
|
# (never a function-local) on purpose: this ONE trap, set once, covers every path that ever
|
||||||
|
# creates a candidate — run_edit and dry_run_diff both assign it, and clear it back to "" once the
|
||||||
|
# file is consumed (installed, or explicitly removed), so a later, unrelated exit never retries a
|
||||||
|
# path that already served its purpose.
|
||||||
|
CAND=""
|
||||||
|
cleanup_candidate() { [ -n "$CAND" ] && rm -f "$CAND" 2>/dev/null; return 0; }
|
||||||
|
trap cleanup_candidate EXIT INT TERM
|
||||||
|
|
||||||
|
say() { printf '\n\033[1m== %s\033[0m\n' "$*"; }
|
||||||
|
ok() { printf ' ok %s\n' "$*"; }
|
||||||
|
warn() { printf ' WARN %s\n' "$*"; }
|
||||||
|
die() { printf '\n FAIL %s\n\n' "$*" >&2; exit 1; }
|
||||||
|
|
||||||
|
set_mode() {
|
||||||
|
local new="$1"
|
||||||
|
if [ -n "$MODE" ] && [ "$MODE" != "$new" ]; then
|
||||||
|
die "cannot combine --$MODE and --$new in one invocation"
|
||||||
|
fi
|
||||||
|
MODE="$new"
|
||||||
|
}
|
||||||
|
|
||||||
|
while [ $# -gt 0 ]; do
|
||||||
|
case "$1" in
|
||||||
|
--check) set_mode check; shift ;;
|
||||||
|
--restore) set_mode restore; shift ;;
|
||||||
|
--set)
|
||||||
|
[ $# -ge 2 ] || die "--set requires <yq-path>=<value>"
|
||||||
|
set_mode set
|
||||||
|
SETS+=("$2")
|
||||||
|
shift 2 ;;
|
||||||
|
--from)
|
||||||
|
[ $# -ge 2 ] || die "--from requires a candidate file path"
|
||||||
|
set_mode from
|
||||||
|
FROM_FILE="$2"
|
||||||
|
shift 2 ;;
|
||||||
|
--dry-run) DRY_RUN=1; shift ;;
|
||||||
|
--config)
|
||||||
|
[ $# -ge 2 ] || die "--config requires a path"
|
||||||
|
CONFIG="$2"; shift 2 ;;
|
||||||
|
--log)
|
||||||
|
[ $# -ge 2 ] || die "--log requires a path"
|
||||||
|
LOG="$2"; shift 2 ;;
|
||||||
|
--wait-seconds)
|
||||||
|
[ $# -ge 2 ] || die "--wait-seconds requires a number of seconds"
|
||||||
|
WAIT_SECONDS_OVERRIDE="$2"; shift 2 ;;
|
||||||
|
-h|--help) sed -n '3,70p' "$SELF"; exit 0 ;;
|
||||||
|
*) echo "unknown option: $1 (try --help)" >&2; exit 2 ;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
|
||||||
|
[ -n "$MODE" ] || die "no action given — use --check, --set, --from, or --restore (see --help)"
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------------------------- redaction
|
||||||
|
#
|
||||||
|
# Two independent passes, applied to every diff this script ever prints:
|
||||||
|
# 1. `scheme://user:pass@host` -> `scheme://<redacted>@host`, globally (the `g` flag matters —
|
||||||
|
# a line can carry more than one URI).
|
||||||
|
# 2. Any line whose key looks like TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY,
|
||||||
|
# matched case-insensitively against the key text (uriEnv, gitTokenEnv, ... are camelCase,
|
||||||
|
# not SCREAMING_CASE) has its whole value blanked, diff marker and indentation kept so the
|
||||||
|
# shape of the change is still visible. Deliberately conservative: a false-positive
|
||||||
|
# redaction on an unrelated line costs nothing, an unredacted secret is a security defect
|
||||||
|
# (acceptance criterion 7).
|
||||||
|
#
|
||||||
|
# fleetd #635 follow-up (ticket comment 17670, defect 7) — a masked key line is not the whole
|
||||||
|
# story: a YAML block scalar (`|`, `|-`, `>`, `>-`, ...) puts the VALUE on the lines that follow
|
||||||
|
# the key, each indented deeper than it. The key-name match above only ever sees the key line
|
||||||
|
# itself, so those continuation lines used to flow straight through unredacted while the key line
|
||||||
|
# right above them printed a reassuring "<redacted>" — an incomplete redactor that looks complete
|
||||||
|
# is worse than one that visibly does nothing, because it stops a reviewer from looking further.
|
||||||
|
# The fix is structural, not another name to match: once a key line is masked, every following
|
||||||
|
# line indented STRICTLY DEEPER than that key is masked too, by indentation alone, until the
|
||||||
|
# indentation returns to the key's own level or shallower. This needs no knowledge of the key's
|
||||||
|
# name, so it covers a block scalar under any masked key — but ONLY while that key's own line is
|
||||||
|
# itself inside the hunk being printed. `diff -u` prints just three lines of context, so a block
|
||||||
|
# scalar's body often reaches this function with its key line left out; there is then nothing to
|
||||||
|
# anchor to, `masked` is never set, and the body prints in full. A blank line inside a block
|
||||||
|
# scalar loses the anchor the same way, because a blank diff line measures as indent 0. Both are
|
||||||
|
# measured and filed as fleetd #639 — do not read this paragraph as a guarantee that a masked
|
||||||
|
# key's value can never be printed.
|
||||||
|
#
|
||||||
|
# `redact` is always fed `diff -u` output, and every line of a unified diff starts with exactly
|
||||||
|
# one of ' ', '+', '-' (the three body markers; '@'/'-'/'+' for the three header-line kinds too).
|
||||||
|
# That one leading character is NOT part of the YAML indentation, and must be stripped before
|
||||||
|
# indentation is measured or a key is matched — otherwise a changed ('+' or '-') line reads one
|
||||||
|
# column shallower than it really is, and either wrongly escapes a continuation mask or wrongly
|
||||||
|
# ends one early. Tabs are out of scope: YAML forbids them for indentation, and this is a bounded
|
||||||
|
# fix, not a YAML parser.
|
||||||
|
redact() {
|
||||||
|
local line prefix content indent lead key
|
||||||
|
local masked=0 masked_indent=0 saved_nocasematch=0
|
||||||
|
shopt -q nocasematch && saved_nocasematch=1
|
||||||
|
shopt -s nocasematch
|
||||||
|
sed -E 's#://[^@]*@#://<redacted>@#g' | while IFS= read -r line || [ -n "$line" ]; do
|
||||||
|
case "$line" in
|
||||||
|
[\ +-]*) prefix="${line:0:1}"; content="${line:1}" ;;
|
||||||
|
*) prefix=""; content="$line" ;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
indent=0
|
||||||
|
while [ "${content:$indent:1}" = " " ]; do indent=$((indent + 1)); done
|
||||||
|
|
||||||
|
if [ "$masked" = 1 ] && [ "$indent" -gt "$masked_indent" ]; then
|
||||||
|
printf '%s%*s<redacted>\n' "$prefix" "$indent" ""
|
||||||
|
continue
|
||||||
|
fi
|
||||||
|
masked=0
|
||||||
|
|
||||||
|
if [[ "$content" =~ ^([[:space:]]*)([A-Za-z0-9_.-]+:) ]]; then
|
||||||
|
lead="${BASH_REMATCH[1]}"
|
||||||
|
key="${BASH_REMATCH[2]}"
|
||||||
|
if [[ "$key" =~ (TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY) ]]; then
|
||||||
|
printf '%s%s%s <redacted>\n' "$prefix" "$lead" "$key"
|
||||||
|
masked=1
|
||||||
|
masked_indent="$indent"
|
||||||
|
continue
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
printf '%s\n' "$line"
|
||||||
|
done
|
||||||
|
[ "$saved_nocasematch" = 1 ] || shopt -u nocasematch
|
||||||
|
}
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------------------------- the probe
|
||||||
|
#
|
||||||
|
# Probe the SOCKET, never `pgrep`/`ps -f` — both print argv, and argv holds `NAME=value`, making
|
||||||
|
# either a credential channel. The port comes from the config's own `bind.port`; 8765 is only a
|
||||||
|
# fallback when that key is absent or the file does not parse yet.
|
||||||
|
resolve_port() {
|
||||||
|
local file="$1" port
|
||||||
|
if [ -f "$file" ] && command -v yq >/dev/null 2>&1; then
|
||||||
|
port="$(yq eval '.bind.port' "$file" 2>/dev/null || true)"
|
||||||
|
else
|
||||||
|
port=""
|
||||||
|
fi
|
||||||
|
case "$port" in
|
||||||
|
''|null) echo "$FALLBACK_PORT" ;;
|
||||||
|
*) echo "$port" ;;
|
||||||
|
esac
|
||||||
|
}
|
||||||
|
|
||||||
|
daemon_listening() {
|
||||||
|
local port="$1"
|
||||||
|
if command -v nc >/dev/null 2>&1; then
|
||||||
|
nc -z -w1 127.0.0.1 "$port" 2>/dev/null
|
||||||
|
else
|
||||||
|
( exec 3<>"/dev/tcp/127.0.0.1/$port" ) 2>/dev/null
|
||||||
|
fi
|
||||||
|
}
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------------- the log marker
|
||||||
|
#
|
||||||
|
# Take the log's line count BEFORE touching anything. Every later read of "what did the daemon
|
||||||
|
# say" starts strictly after this mark, so a refusal from hours ago can never be mistaken for
|
||||||
|
# this edit's verdict. Same approach as scripts/redeploy-fleetd.sh's RESTART_MARK.
|
||||||
|
log_mark() {
|
||||||
|
local file="$1"
|
||||||
|
if [ -f "$file" ]; then
|
||||||
|
wc -l < "$file" 2>/dev/null || echo 0
|
||||||
|
else
|
||||||
|
echo 0
|
||||||
|
fi
|
||||||
|
}
|
||||||
|
|
||||||
|
read_verdict_after_marker() {
|
||||||
|
local file="$1" mark="$2"
|
||||||
|
[ -f "$file" ] || return 0
|
||||||
|
tail -n "+$((mark + 1))" "$file" 2>/dev/null || true
|
||||||
|
}
|
||||||
|
|
||||||
|
# Classifies one log LINE. Echoes one of: refused | clean | needs-restart | none. Always
|
||||||
|
# succeeds (every branch ends in `echo`), so it is safe to call from inside `$( )`.
|
||||||
|
classify_verdict_line() {
|
||||||
|
local line="$1"
|
||||||
|
case "$line" in
|
||||||
|
*'config reload refused'*) echo refused ;;
|
||||||
|
*'config reload from '*'refused, keeping the running config'*) echo refused ;;
|
||||||
|
*'config reloaded'*)
|
||||||
|
case "$line" in
|
||||||
|
*'need a restart'*|*'partially live'*) echo needs-restart ;;
|
||||||
|
*) echo clean ;;
|
||||||
|
esac ;;
|
||||||
|
*) echo none ;;
|
||||||
|
esac
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
|
scan_region_for_verdict() {
|
||||||
|
local region="$1" line kind
|
||||||
|
[ -n "$region" ] || return 1
|
||||||
|
while IFS= read -r line || [ -n "$line" ]; do
|
||||||
|
kind="$(classify_verdict_line "$line")"
|
||||||
|
if [ "$kind" != "none" ]; then
|
||||||
|
VERDICT_KIND="$kind"
|
||||||
|
VERDICT_LINE="$line"
|
||||||
|
return 0
|
||||||
|
fi
|
||||||
|
done <<< "$region"
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
# Sets VERDICT_KIND/VERDICT_LINE and returns 0 on the first verdict line found after $mark;
|
||||||
|
# returns 1 (VERDICT_KIND=none) if none appeared inside $wait_s seconds. Checks once before each
|
||||||
|
# sleep AND once more after the last sleep, the same boundary idiom
|
||||||
|
# scripts/redeploy-fleetd.sh's wait_for_daemon_exit/wait_for_new_pid already use.
|
||||||
|
wait_for_verdict() {
|
||||||
|
local log="$1" mark="$2" wait_s="$3" _i region
|
||||||
|
VERDICT_KIND="none"
|
||||||
|
VERDICT_LINE=""
|
||||||
|
for _i in $(seq "$wait_s"); do
|
||||||
|
region="$(read_verdict_after_marker "$log" "$mark")"
|
||||||
|
scan_region_for_verdict "$region" && return 0
|
||||||
|
sleep 1
|
||||||
|
done
|
||||||
|
region="$(read_verdict_after_marker "$log" "$mark")"
|
||||||
|
scan_region_for_verdict "$region" && return 0
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
last_verdict_line() {
|
||||||
|
local file="$1" line out=""
|
||||||
|
[ -f "$file" ] || return 0
|
||||||
|
while IFS= read -r line || [ -n "$line" ]; do
|
||||||
|
if [ "$(classify_verdict_line "$line")" != "none" ]; then
|
||||||
|
out="$line"
|
||||||
|
fi
|
||||||
|
done < "$file"
|
||||||
|
printf '%s' "$out"
|
||||||
|
}
|
||||||
|
|
||||||
|
default_wait_seconds() {
|
||||||
|
local log="$1" interval=""
|
||||||
|
if [ -f "$log" ]; then
|
||||||
|
interval="$(grep -F 'config watch:' "$log" 2>/dev/null | tail -1 \
|
||||||
|
| sed -E 's/.*\(every ([0-9]+)s\).*/\1/' || true)"
|
||||||
|
fi
|
||||||
|
case "$interval" in
|
||||||
|
''|*[!0-9]*) interval=10 ;;
|
||||||
|
esac
|
||||||
|
echo $((interval * 4))
|
||||||
|
}
|
||||||
|
|
||||||
|
# ----------------------------------------------------------------------------------- the backup
|
||||||
|
#
|
||||||
|
# Timestamped, never pruned — "keep backups" per the ticket. A pid suffix avoids a same-second
|
||||||
|
# collision between two invocations.
|
||||||
|
#
|
||||||
|
# fleetd #635 follow-up — lands under a DEDICATED, gitignored directory beside the config
|
||||||
|
# (<dir>/.config-backups/), never beside the config file itself. The whole reason fleetd.yaml is
|
||||||
|
# gitignored is that it must never be committed, and a backup of it inherits that requirement — a
|
||||||
|
# bare `fleetd.yaml.bak.*` next to a tracked directory is one `git add -A`/`git add .` away from
|
||||||
|
# committing the live config. A directory beats a glob on its own: the glob only protects today's
|
||||||
|
# naming, a location keeps working even if the naming changes later. (See .gitignore for the glob
|
||||||
|
# kept anyway, as a backstop for a stray backup written the old way.)
|
||||||
|
BACKUP_DIRNAME=".config-backups"
|
||||||
|
|
||||||
|
backup_dir_for() {
|
||||||
|
local src="$1"
|
||||||
|
printf '%s/%s' "$(dirname "$src")" "$BACKUP_DIRNAME"
|
||||||
|
}
|
||||||
|
|
||||||
|
backup_config() {
|
||||||
|
local src="$1" ts backup dir base
|
||||||
|
dir="$(backup_dir_for "$src")"
|
||||||
|
mkdir -p "$dir" \
|
||||||
|
|| die "could not create the backup directory $dir — refusing to edit without a backup. The live config at $src was NOT touched."
|
||||||
|
ts="$(date -u +%Y%m%dT%H%M%S)Z"
|
||||||
|
base="$(basename "$src")"
|
||||||
|
backup="${dir}/${base}.bak.${ts}.$$"
|
||||||
|
cp "$src" "$backup" \
|
||||||
|
|| die "could not create a backup at $backup — refusing to edit without one. The live config at $src was NOT touched."
|
||||||
|
printf '%s' "$backup"
|
||||||
|
}
|
||||||
|
|
||||||
|
newest_backup() {
|
||||||
|
local cfg="$1" dir base
|
||||||
|
dir="$(backup_dir_for "$cfg")"
|
||||||
|
base="$(basename "$cfg")"
|
||||||
|
ls -t "${dir}/${base}".bak.* 2>/dev/null | head -1 || true
|
||||||
|
}
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------------------------ the file mode
|
||||||
|
#
|
||||||
|
# fleetd #635 follow-up — `mv` from a mktemp candidate carries mktemp's 0600 onto the live path
|
||||||
|
# forever (measured: 644 -> 600 after one --set), and a restore does not undo it either, because
|
||||||
|
# `cp` onto an EXISTING file keeps the DESTINATION's mode, not the source's. Capture the live
|
||||||
|
# file's mode before anything touches it, and reapply it to whatever lands on that path
|
||||||
|
# afterwards — the candidate before install, and the config again after a restore — so an edit
|
||||||
|
# changes the file's CONTENT only, never its permissions. BSD `stat -f '%Lp'` first (matches this
|
||||||
|
# project's dev machine), GNU `stat -c '%a'` as the fallback. Prints nothing when the file does
|
||||||
|
# not exist yet, so apply_mode then does nothing and a first-ever edit falls back to the normal
|
||||||
|
# umask default rather than inventing a number.
|
||||||
|
file_mode() {
|
||||||
|
local file="$1"
|
||||||
|
[ -f "$file" ] || return 0
|
||||||
|
stat -f '%Lp' "$file" 2>/dev/null || stat -c '%a' "$file" 2>/dev/null || true
|
||||||
|
}
|
||||||
|
|
||||||
|
apply_mode() {
|
||||||
|
local file="$1" mode="$2"
|
||||||
|
[ -n "$mode" ] || return 0
|
||||||
|
chmod "$mode" "$file" 2>/dev/null || true
|
||||||
|
}
|
||||||
|
|
||||||
|
# --------------------------------------------------------------------------- candidate builders
|
||||||
|
#
|
||||||
|
# Never edit the live file in place. Each builder fills $1 (a temp file already sitting in the
|
||||||
|
# SAME directory as the live config, so the later `mv` install is a rename, not a cross-device
|
||||||
|
# copy — see run_edit).
|
||||||
|
# fleetd #635 follow-up — a forgotten value (`--set .a.b=`, a plausible typo) must never be
|
||||||
|
# accepted as "clear the field". `*=*` alone cannot tell "--set .a.b=" from "--set .a.b=7" apart
|
||||||
|
# — both contain an `=` — so the guard has to look at the VALUE, not the shape of the argument.
|
||||||
|
# An empty value refuses outright: nothing is installed, and the message names the likely cause
|
||||||
|
# AND the two ways to actually mean it (clear on purpose, or an intentional empty string via
|
||||||
|
# --from). Measured against the real daemon loader: a quoted empty string reads back as a null
|
||||||
|
# field (`quoted empty -> OK int=null`), and a null numeric field FALLS BACK TO ITS DEFAULT rather
|
||||||
|
# than erroring — so this is not a cosmetic nit, it is the one shape of edit that widens capacity
|
||||||
|
# silently instead of failing loudly, which is exactly what this script exists to catch.
|
||||||
|
#
|
||||||
|
# A deliberate clear needs its own spelling, because `""` and YAML `null` are NOT the same value
|
||||||
|
# to the loader (`""` is a valid empty String; `null` means absent, and an Integer field reads
|
||||||
|
# either the same way — null — but a String field would keep `""` as a real value). `--set
|
||||||
|
# .a.b=null` is that spelling: it writes a literal, unquoted `null` via yq, never the string
|
||||||
|
# "null" through strenv(). One consequence worth knowing: there is currently no --set spelling
|
||||||
|
# for the three-character STRING "null" itself (it collides with the clear spelling) — use
|
||||||
|
# --from for that rare case.
|
||||||
|
apply_set_pairs() {
|
||||||
|
local cand="$1" kv path value
|
||||||
|
shift
|
||||||
|
for kv in "$@"; do
|
||||||
|
case "$kv" in
|
||||||
|
*=*) : ;;
|
||||||
|
*) die "--set expects <yq-path>=<value>, got: '$kv'" ;;
|
||||||
|
esac
|
||||||
|
path="${kv%%=*}"
|
||||||
|
path="${path#.}"
|
||||||
|
value="${kv#*=}"
|
||||||
|
if [ -z "$value" ]; then
|
||||||
|
die "--set '$kv' has an EMPTY value — refusing. Nothing was installed. A forgotten value
|
||||||
|
would NULL the field, and a null value falls back to its default rather than erroring —
|
||||||
|
silent, not safe. Did you mean --set .${path}=null to clear it on purpose, or --from a
|
||||||
|
file if you need a genuinely empty string?"
|
||||||
|
fi
|
||||||
|
# fleetd #635 follow-up (ticket comment 17673, defect 8) — these two failure messages used to
|
||||||
|
# echo the full "$kv" (path=value, exactly as the operator typed it), unredacted. The operator
|
||||||
|
# already has the value, so a terminal is not where this leaks — the risk is where the output
|
||||||
|
# goes NEXT: this fleet pastes command output into tickets, PRs and fleet_reply bodies, and a
|
||||||
|
# failure is exactly when someone copies it to ask for help. Print the PATH, which is what's
|
||||||
|
# needed to fix the command, and never the value. $kv is not key:value-shaped YAML, so piping
|
||||||
|
# it through redact would just pass it straight through — a false sense of coverage, the same
|
||||||
|
# mistake as defect 7.
|
||||||
|
if [ "$value" = "null" ]; then
|
||||||
|
yq eval -i ".${path} = null" "$cand" \
|
||||||
|
|| die "yq could not clear --set '.${path}=null' — nothing was installed. The live config is unchanged."
|
||||||
|
continue
|
||||||
|
fi
|
||||||
|
CONFIG_EDIT_SET_VALUE="$value" yq eval -i ".${path} = strenv(CONFIG_EDIT_SET_VALUE)" "$cand" \
|
||||||
|
|| die "yq could not apply --set '.${path}=<value>' — nothing was installed. The live config is unchanged."
|
||||||
|
done
|
||||||
|
}
|
||||||
|
|
||||||
|
build_from_set() {
|
||||||
|
local cand="$1"
|
||||||
|
cp "$CONFIG" "$cand"
|
||||||
|
apply_set_pairs "$cand" "${SETS[@]}"
|
||||||
|
}
|
||||||
|
|
||||||
|
build_from_file() {
|
||||||
|
local cand="$1"
|
||||||
|
[ -f "$FROM_FILE" ] || die "--from file not found: $FROM_FILE"
|
||||||
|
cp "$FROM_FILE" "$cand"
|
||||||
|
}
|
||||||
|
|
||||||
|
parse_check() {
|
||||||
|
yq eval '.' "$1" >/dev/null 2>&1
|
||||||
|
}
|
||||||
|
|
||||||
|
install_candidate() {
|
||||||
|
local cand="$1" live="$2"
|
||||||
|
mv -f "$cand" "$live"
|
||||||
|
}
|
||||||
|
|
||||||
|
# -------------------------------------------------------------------------------- the report path
|
||||||
|
#
|
||||||
|
# Prints the literal command the operator (or a test) can run to restore the backup by hand — the
|
||||||
|
# absolute path to THIS script plus the overrides actually in force, so it works from any cwd.
|
||||||
|
restore_command_line() {
|
||||||
|
printf '%q --restore --config %q --log %q --wait-seconds %q' "$SELF" "$CONFIG" "$LOG" "$WAIT_SECONDS"
|
||||||
|
}
|
||||||
|
|
||||||
|
# State 4 only: restore the pre-edit backup, then wait for a SECOND verdict confirming the
|
||||||
|
# restore itself reloaded cleanly. Never claims a restore it did not observe — if the second wait
|
||||||
|
# also times out, it says so plainly rather than reporting "restored" as though confirmed.
|
||||||
|
restore_and_confirm() {
|
||||||
|
local backup="$1" mark2 orig_mode
|
||||||
|
orig_mode="$(file_mode "$CONFIG")"
|
||||||
|
mark2="$(log_mark "$LOG")"
|
||||||
|
cp "$backup" "$CONFIG" \
|
||||||
|
|| die "could not restore $backup onto $CONFIG — the live config is left as the REFUSED edit. Fix this by hand immediately: cp \"$backup\" \"$CONFIG\""
|
||||||
|
apply_mode "$CONFIG" "$orig_mode"
|
||||||
|
ok "restored from $backup"
|
||||||
|
if wait_for_verdict "$LOG" "$mark2" "$WAIT_SECONDS"; then
|
||||||
|
case "$VERDICT_KIND" in
|
||||||
|
refused) warn "the RESTORE was also refused by the daemon: $VERDICT_LINE" ;;
|
||||||
|
*) ok "restore confirmed: $VERDICT_LINE" ;;
|
||||||
|
esac
|
||||||
|
else
|
||||||
|
warn "the restore is on disk, but no confirming verdict line appeared within ${WAIT_SECONDS}s"
|
||||||
|
warn "cannot confirm the restore reloaded cleanly — check $LOG by hand"
|
||||||
|
fi
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
|
# The four-outcome decision. Echoed as a function so run_edit/restore_mode share one place that
|
||||||
|
# can return 0/3/4/5 — never duplicated, never re-worded between the two callers.
|
||||||
|
report_outcome() {
|
||||||
|
local mark="$1" backup="$2" kind line
|
||||||
|
|
||||||
|
say "waiting for the daemon's verdict (up to ${WAIT_SECONDS}s)"
|
||||||
|
if wait_for_verdict "$LOG" "$mark" "$WAIT_SECONDS"; then
|
||||||
|
kind="$VERDICT_KIND"; line="$VERDICT_LINE"
|
||||||
|
else
|
||||||
|
kind="none"
|
||||||
|
fi
|
||||||
|
|
||||||
|
case "$kind" in
|
||||||
|
clean)
|
||||||
|
ok "daemon verdict: $line"
|
||||||
|
say "result: applied cleanly"
|
||||||
|
return 0 ;;
|
||||||
|
needs-restart)
|
||||||
|
ok "daemon verdict: $line"
|
||||||
|
say "result: applied — a restart is needed for the change(s) named above"
|
||||||
|
return 3 ;;
|
||||||
|
refused)
|
||||||
|
warn "daemon verdict: $line"
|
||||||
|
say "result: REFUSED — restoring the backup"
|
||||||
|
restore_and_confirm "$backup"
|
||||||
|
return 4 ;;
|
||||||
|
none)
|
||||||
|
warn "no verdict line appeared within ${WAIT_SECONDS}s after $LOG line $mark"
|
||||||
|
warn "CANNOT TELL whether the daemon applied this edit, refused it, or is simply down."
|
||||||
|
warn "Nothing was restored — the edit is still on disk at $CONFIG."
|
||||||
|
echo
|
||||||
|
echo " backup: $backup"
|
||||||
|
echo " to restore it by hand:"
|
||||||
|
echo " $(restore_command_line)"
|
||||||
|
return 5 ;;
|
||||||
|
esac
|
||||||
|
}
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------------------------- the modes
|
||||||
|
check_mode() {
|
||||||
|
say "config-edit --check"
|
||||||
|
if [ -f "$CONFIG" ]; then
|
||||||
|
if parse_check "$CONFIG"; then
|
||||||
|
ok "config parses: $CONFIG"
|
||||||
|
else
|
||||||
|
warn "config does NOT parse as valid YAML: $CONFIG"
|
||||||
|
fi
|
||||||
|
else
|
||||||
|
warn "no config file at $CONFIG"
|
||||||
|
fi
|
||||||
|
|
||||||
|
local port
|
||||||
|
port="$(resolve_port "$CONFIG")"
|
||||||
|
if daemon_listening "$port"; then
|
||||||
|
ok "daemon is listening on 127.0.0.1:$port"
|
||||||
|
else
|
||||||
|
warn "no daemon detected listening on 127.0.0.1:$port"
|
||||||
|
fi
|
||||||
|
|
||||||
|
ok "watch interval assumed: $(( $(default_wait_seconds "$LOG") / 4 ))s (derives --wait-seconds default of $(default_wait_seconds "$LOG")s)"
|
||||||
|
|
||||||
|
local verdict
|
||||||
|
verdict="$(last_verdict_line "$LOG")"
|
||||||
|
if [ -n "$verdict" ]; then
|
||||||
|
ok "last verdict in log: $verdict"
|
||||||
|
else
|
||||||
|
warn "no reload verdict line found in $LOG"
|
||||||
|
fi
|
||||||
|
|
||||||
|
local backup
|
||||||
|
backup="$(newest_backup "$CONFIG")"
|
||||||
|
if [ -n "$backup" ]; then
|
||||||
|
ok "newest backup: $backup"
|
||||||
|
else
|
||||||
|
warn "no backups found for $CONFIG"
|
||||||
|
fi
|
||||||
|
|
||||||
|
if command -v yq >/dev/null 2>&1; then
|
||||||
|
ok "yq: $(yq --version 2>&1)"
|
||||||
|
else
|
||||||
|
warn "yq not found on PATH"
|
||||||
|
fi
|
||||||
|
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
|
# Shared by --set and --from: backup, build, parse-check, redacted diff, install, await verdict.
|
||||||
|
run_edit() {
|
||||||
|
local builder="$1"
|
||||||
|
[ -f "$CONFIG" ] || die "no config at $CONFIG — nothing to edit"
|
||||||
|
|
||||||
|
local mark orig_mode
|
||||||
|
mark="$(log_mark "$LOG")"
|
||||||
|
orig_mode="$(file_mode "$CONFIG")"
|
||||||
|
|
||||||
|
say "probe"
|
||||||
|
local port
|
||||||
|
port="$(resolve_port "$CONFIG")"
|
||||||
|
if daemon_listening "$port"; then
|
||||||
|
ok "daemon appears to be listening on 127.0.0.1:$port"
|
||||||
|
else
|
||||||
|
warn "no daemon detected listening on 127.0.0.1:$port — a verdict may never appear"
|
||||||
|
fi
|
||||||
|
|
||||||
|
say "backup"
|
||||||
|
local backup
|
||||||
|
backup="$(backup_config "$CONFIG")"
|
||||||
|
ok "backup: $backup"
|
||||||
|
|
||||||
|
say "candidate"
|
||||||
|
local cand
|
||||||
|
CAND="$(mktemp "$(dirname "$CONFIG")/.config-edit.XXXXXX")" \
|
||||||
|
|| die "could not create a candidate temp file next to $CONFIG"
|
||||||
|
cand="$CAND"
|
||||||
|
if ! "$builder" "$cand"; then
|
||||||
|
rm -f "$cand"; CAND=""
|
||||||
|
die "could not build the candidate — nothing was installed. The live config at $CONFIG is unchanged."
|
||||||
|
fi
|
||||||
|
|
||||||
|
if ! parse_check "$cand"; then
|
||||||
|
rm -f "$cand"; CAND=""
|
||||||
|
die "candidate does not parse as valid YAML — nothing was installed. The live config at $CONFIG is unchanged."
|
||||||
|
fi
|
||||||
|
ok "candidate parses"
|
||||||
|
|
||||||
|
apply_mode "$cand" "$orig_mode"
|
||||||
|
|
||||||
|
say "change (redacted)"
|
||||||
|
diff -u "$backup" "$cand" | redact || true
|
||||||
|
|
||||||
|
say "install"
|
||||||
|
install_candidate "$cand" "$CONFIG" \
|
||||||
|
|| die "could not install the candidate onto $CONFIG — the live config was NOT changed. The validated candidate is sitting at $cand; investigate before retrying."
|
||||||
|
CAND=""
|
||||||
|
ok "installed: $CONFIG"
|
||||||
|
|
||||||
|
local rc=0
|
||||||
|
report_outcome "$mark" "$backup" || rc=$?
|
||||||
|
return "$rc"
|
||||||
|
}
|
||||||
|
|
||||||
|
dry_run_diff() {
|
||||||
|
local builder="$1"
|
||||||
|
[ -f "$CONFIG" ] || die "no config at $CONFIG — nothing to diff against"
|
||||||
|
local cand
|
||||||
|
CAND="$(mktemp "$(dirname "$CONFIG")/.config-edit.XXXXXX")" \
|
||||||
|
|| die "could not create a candidate temp file next to $CONFIG"
|
||||||
|
cand="$CAND"
|
||||||
|
if ! "$builder" "$cand"; then
|
||||||
|
rm -f "$cand"; CAND=""
|
||||||
|
die "could not build the candidate — this was a --dry-run, nothing would have been installed either"
|
||||||
|
fi
|
||||||
|
if ! parse_check "$cand"; then
|
||||||
|
rm -f "$cand"; CAND=""
|
||||||
|
die "candidate does not parse as valid YAML — this was a --dry-run, nothing would have been installed either"
|
||||||
|
fi
|
||||||
|
say "dry run — diff (redacted), nothing installed"
|
||||||
|
diff -u "$CONFIG" "$cand" | redact || true
|
||||||
|
rm -f "$cand"; CAND=""
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
|
restore_mode() {
|
||||||
|
[ -f "$CONFIG" ] || die "no config at $CONFIG to restore onto"
|
||||||
|
local backup dir base
|
||||||
|
backup="$(newest_backup "$CONFIG")"
|
||||||
|
if [ -z "$backup" ]; then
|
||||||
|
# fleetd #635 follow-up (ticket comment 17664) — this message must name the directory the
|
||||||
|
# code actually searches (backup_dir_for, same as newest_backup), not the old beside-the-
|
||||||
|
# config glob. A backup written the OLD way is real and NOT searched any more — say so and
|
||||||
|
# give the one-line recovery command — but do NOT make the search itself look there; that
|
||||||
|
# would be a behaviour change nobody asked for. The message is the only thing being fixed.
|
||||||
|
dir="$(backup_dir_for "$CONFIG")"
|
||||||
|
base="$(basename "$CONFIG")"
|
||||||
|
die "no backup found matching ${dir}/${base}.bak.* — nothing to restore.
|
||||||
|
A backup written the OLD way, directly beside the config (${CONFIG}.bak.*), is NOT
|
||||||
|
searched — that location was retired so a backup of a file that must never be committed
|
||||||
|
cannot sit next to a tracked directory. If one exists there, recover it by hand:
|
||||||
|
cp ${CONFIG}.bak.<timestamp>.<pid> $CONFIG"
|
||||||
|
fi
|
||||||
|
[ -f "$backup" ] || die "backup candidate $backup vanished"
|
||||||
|
|
||||||
|
say "restore"
|
||||||
|
ok "restoring $backup onto $CONFIG"
|
||||||
|
local mark orig_mode
|
||||||
|
mark="$(log_mark "$LOG")"
|
||||||
|
orig_mode="$(file_mode "$CONFIG")"
|
||||||
|
cp "$backup" "$CONFIG" || die "could not copy $backup onto $CONFIG"
|
||||||
|
apply_mode "$CONFIG" "$orig_mode"
|
||||||
|
ok "installed: $CONFIG"
|
||||||
|
|
||||||
|
local rc=0
|
||||||
|
report_outcome "$mark" "$backup" || rc=$?
|
||||||
|
return "$rc"
|
||||||
|
}
|
||||||
|
|
||||||
|
# -------------------------------------------------------------------------------------- dispatch
|
||||||
|
|
||||||
|
if [ -n "$WAIT_SECONDS_OVERRIDE" ]; then
|
||||||
|
WAIT_SECONDS="$WAIT_SECONDS_OVERRIDE"
|
||||||
|
else
|
||||||
|
WAIT_SECONDS="$(default_wait_seconds "$LOG")"
|
||||||
|
fi
|
||||||
|
|
||||||
|
RC=0
|
||||||
|
case "$MODE" in
|
||||||
|
check)
|
||||||
|
check_mode || RC=$?
|
||||||
|
;;
|
||||||
|
set)
|
||||||
|
[ "${#SETS[@]}" -gt 0 ] || die "--set requires at least one <yq-path>=<value>"
|
||||||
|
if [ "$DRY_RUN" = 1 ]; then
|
||||||
|
dry_run_diff build_from_set || RC=$?
|
||||||
|
else
|
||||||
|
run_edit build_from_set || RC=$?
|
||||||
|
fi
|
||||||
|
;;
|
||||||
|
from)
|
||||||
|
[ -n "$FROM_FILE" ] || die "--from requires a candidate file path"
|
||||||
|
if [ "$DRY_RUN" = 1 ]; then
|
||||||
|
dry_run_diff build_from_file || RC=$?
|
||||||
|
else
|
||||||
|
run_edit build_from_file || RC=$?
|
||||||
|
fi
|
||||||
|
;;
|
||||||
|
restore)
|
||||||
|
restore_mode || RC=$?
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
exit "$RC"
|
||||||
+131
-23
@@ -50,6 +50,14 @@
|
|||||||
# absence is the real signal, because a drain that dies on its first session prints nothing
|
# absence is the real signal, because a drain that dies on its first session prints nothing
|
||||||
# else either. Warns loudly; never fails the redeploy, because by the time this is detectable
|
# else either. Warns loudly; never fails the redeploy, because by the time this is detectable
|
||||||
# the new daemon is already up and healthy.
|
# the new daemon is already up and healthy.
|
||||||
|
# 10. fleetd #603 — the same shape as trap 3 above, through a different door: the step that waited
|
||||||
|
# for the NEW process to appear gave it its own short, fixed 10s budget, then hard-`die`d,
|
||||||
|
# while the health check right after it waits a full $HEALTH_WAIT (60s) for the same daemon to
|
||||||
|
# answer. Under launchd, `launchctl load` returns as soon as launchd accepts the job, before the
|
||||||
|
# java process exists, and on a slow host that took longer than 10s — so the script died with
|
||||||
|
# "no process appeared" on a deploy that had fully succeeded. The pid poll now shares
|
||||||
|
# $HEALTH_WAIT instead of a separate, shorter budget, and a miss there falls through to the
|
||||||
|
# health check (the truer signal: is it actually answering?) instead of killing the run.
|
||||||
#
|
#
|
||||||
# Usage:
|
# Usage:
|
||||||
# scripts/redeploy-fleetd.sh # build, confirm, restart, verify
|
# scripts/redeploy-fleetd.sh # build, confirm, restart, verify
|
||||||
@@ -79,7 +87,9 @@ OUT="$MODULE/fleetd.out"
|
|||||||
PATTERN='target/fleetd.jar'
|
PATTERN='target/fleetd.jar'
|
||||||
HEALTH='http://127.0.0.1:8765/healthz'
|
HEALTH='http://127.0.0.1:8765/healthz'
|
||||||
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
|
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
|
||||||
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start
|
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start — fleetd #603: also the pid-
|
||||||
|
# poll budget below (wait_for_new_pid/await_daemon_started), so the two checks
|
||||||
|
# share one named budget instead of the pid poll holding its own shorter one
|
||||||
|
|
||||||
# CB-594: the launchd agent this script must not fight with (see trap 6 above).
|
# CB-594: the launchd agent this script must not fight with (see trap 6 above).
|
||||||
LAUNCHD_LABEL='dev.ltms.fleetd'
|
LAUNCHD_LABEL='dev.ltms.fleetd'
|
||||||
@@ -169,7 +179,55 @@ hash256() {
|
|||||||
# on PATH). "absent" must never be the answer for a file that exists — that conflation, on Linux,
|
# on PATH). "absent" must never be the answer for a file that exists — that conflation, on Linux,
|
||||||
# was the whole defect this ticket fixes.
|
# was the whole defect this ticket fixes.
|
||||||
jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
|
jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
|
||||||
running_pid() { pgrep -f "$PATTERN" || true; }
|
|
||||||
|
# fleetd #593 — `pgrep -f "$PATTERN"` matches ANY process whose full command line CONTAINS the
|
||||||
|
# pattern text, and that is not the same thing as "is the daemon". A shell that merely embeds the
|
||||||
|
# pattern as literal text — a human typing this exact investigation by hand, an ssh-shaped
|
||||||
|
# `sh -c '...; ...'`, a pipeline, or any other non-exec'ing shell that never replaced itself with
|
||||||
|
# the pattern-holding command — still shows up in that match, and it is the INSTRUMENT, not the
|
||||||
|
# daemon. Measured live on this Mac: `sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 30' &`
|
||||||
|
# leaves a real `sh` process alive (it forks for the `sleep`, it does not exec into it) whose own
|
||||||
|
# `ps -o args` is `sh -c echo "target/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
|
||||||
|
# matches that line right alongside the real `java -jar target/fleetd.jar` process. `pgrep -c`
|
||||||
|
# (an in-one-call count) does not exist on BSD/macOS at all, so this cannot be fixed by switching
|
||||||
|
# pgrep flags — it has to filter what pgrep already found, after the fact, in a way that still
|
||||||
|
# runs on BSD.
|
||||||
|
#
|
||||||
|
# fleetd #593 CORRECTION 1 — the first cut of this filter kept everything whose `comm` was NOT a
|
||||||
|
# shell name (a denylist: sh/bash/zsh/dash/ksh). Two holes in that, both the same false-positive
|
||||||
|
# shape the ticket exists to remove in the first place:
|
||||||
|
# 1. a pid `pgrep` just listed can exit before the `ps -o comm=` lookup runs; on a gone pid `ps`
|
||||||
|
# prints nothing, `comm` ends up empty, and an empty string matches none of the denied shell
|
||||||
|
# names — so a pid that no longer exists was still counted.
|
||||||
|
# 2. the denylist only knows the shells someone thought to name. `ssh`, `perl`, `python3`,
|
||||||
|
# `ruby`, `tail` — anything else that carries the pattern in its own argv — was still
|
||||||
|
# counted right along with the real daemon, and the ticket names `ssh` as a live route.
|
||||||
|
# Both close with the same change: allowlist `comm = java` instead of denying shells. Measured on
|
||||||
|
# the live daemon: `pid=30224 comm=java`. An empty comm (hole 1) is not `java` either, so it is
|
||||||
|
# excluded for free — no separate "is this pid still alive" check needed.
|
||||||
|
#
|
||||||
|
# The objection, because it is real: an allowlist can UNDER-count. If fleetd ever stops being
|
||||||
|
# launched as `java -jar ...` — a native image, a renamed launcher — `running_pid()` silently
|
||||||
|
# returns nothing and `assert_single_daemon` stops noticing a second daemon at all. For a guard,
|
||||||
|
# that false-negative direction is the worse one to be wrong in. This is not a new assumption,
|
||||||
|
# though: `PATTERN='target/fleetd.jar'` two lines up already assumes the daemon is a jar, which
|
||||||
|
# is only ever run by `java`. If that launch method changes, `PATTERN` stops matching anything
|
||||||
|
# before this allowlist would ever get the chance to be wrong — the allowlist rides on the same
|
||||||
|
# assumption that is already load-bearing, it does not add a new one. Whoever changes the launch
|
||||||
|
# method needs to update both `PATTERN` and this allowlist together.
|
||||||
|
running_pid() {
|
||||||
|
local pid comm out=''
|
||||||
|
for pid in $(pgrep -f "$PATTERN" 2>/dev/null || true); do
|
||||||
|
comm="$(ps -o comm= -p "$pid" 2>/dev/null || true)"
|
||||||
|
comm="${comm##*/}"
|
||||||
|
comm="${comm#-}"
|
||||||
|
# Allowlist, not a denylist of wrappers — see the CORRECTION 1 comment above. Anything that
|
||||||
|
# is not literally `java` is excluded, including an empty comm from a pid that already exited.
|
||||||
|
[ "$comm" = java ] || continue
|
||||||
|
out="$out$pid"$'\n'
|
||||||
|
done
|
||||||
|
printf '%s' "$out"
|
||||||
|
}
|
||||||
|
|
||||||
# fleetd #493 — three small, independently testable pieces of "never build into the path a
|
# fleetd #493 — three small, independently testable pieces of "never build into the path a
|
||||||
# running process holds":
|
# running process holds":
|
||||||
@@ -507,8 +565,11 @@ assert_single_daemon() {
|
|||||||
die "more than one fleetd process is running after this restart (pids: $(printf '%s' "$pids" | tr '\n' ' ')).
|
die "more than one fleetd process is running after this restart (pids: $(printf '%s' "$pids" | tr '\n' ' ')).
|
||||||
This is the exact failure a racing supervisor produces: the OLD jar was revived by its
|
This is the exact failure a racing supervisor produces: the OLD jar was revived by its
|
||||||
supervisor while this script started a NEW copy. Two daemons on one herdr session kill
|
supervisor while this script started a NEW copy. Two daemons on one herdr session kill
|
||||||
each other's members. Investigate with 'pgrep -f \"$PATTERN\"' and stop the wrong one by
|
each other's members. Investigate with 'ps -eo pid,comm,args | grep -F \"$PATTERN\"' and
|
||||||
hand — do not assume either pid is the one you want."
|
check the COMM column of each hit yourself before acting — a bare 'pgrep -f \"$PATTERN\"'
|
||||||
|
(fleetd #593) can match the very shell you type it into, not just the daemon, so it is not
|
||||||
|
safe remediation advice on its own. Stop the wrong one by hand — do not assume either pid
|
||||||
|
is the one you want."
|
||||||
fi
|
fi
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1016,6 +1077,67 @@ $(tail -30 "$out_file" 2>/dev/null)"
|
|||||||
fi
|
fi
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# fleetd #603 — the dual of wait_for_daemon_exit above: waits for a pid to APPEAR instead of
|
||||||
|
# disappear. Used to be a bare `for _ in $(seq 10)` sitting directly in the main flow, with its own
|
||||||
|
# short, fixed budget that had nothing to do with $HEALTH_WAIT (60s) — the budget the health check
|
||||||
|
# right after it gets for the very same daemon. Under launchd, `launchctl load` returns as soon as
|
||||||
|
# launchd accepts the job, before the java process exists, and on a slow host that took longer than
|
||||||
|
# 10s — so the script died with "no process appeared" on a deploy that had fully succeeded (the
|
||||||
|
# operator confirmed pid, healthz, and a fresh log line all present, by hand, right afterwards).
|
||||||
|
# Sharing $HEALTH_WAIT here removes the extra, shorter magic number without inventing a new one.
|
||||||
|
wait_for_new_pid() {
|
||||||
|
local timeout="$1" _i
|
||||||
|
for _i in $(seq "$timeout"); do
|
||||||
|
[ -n "$(running_pid)" ] && return 0
|
||||||
|
sleep 1
|
||||||
|
done
|
||||||
|
[ -n "$(running_pid)" ]
|
||||||
|
}
|
||||||
|
|
||||||
|
# fleetd #603 — the pid-appeared check and the healthz check, folded into one decision. Same shape
|
||||||
|
# as swap_if_built/refuse_drain_gate/report_shutdown_drain above (#521/#528/#512): the main flow
|
||||||
|
# calls this ONE function unconditionally, so there is no bare guard left for a future edit to
|
||||||
|
# invert independently of it. That matters more here than for most of those: a plain grep of this
|
||||||
|
# script's source cannot tell "a pid miss falls through to the health check" from "a pid miss still
|
||||||
|
# dies" apart, because both read as the same two lines of text with only the runtime branch
|
||||||
|
# changed — a source-text test could pass on either behavior. A real behavioural test on this
|
||||||
|
# function is the only thing that can actually tell them apart, which is why one exists below.
|
||||||
|
#
|
||||||
|
# wait_for_new_pid's miss is no longer fatal by itself: it falls through to the health check, which
|
||||||
|
# is direct proof the new daemon is up (/healthz answers 200) rather than a proxy for it (a process
|
||||||
|
# merely existing under a name running_pid() recognises). A genuine failure still dies here: it
|
||||||
|
# misses the pid poll AND the health poll, and report_health's own die() still prints the tail of
|
||||||
|
# $out_file, exactly as before this fix.
|
||||||
|
#
|
||||||
|
# Sets NEW_PID (global — the caller's "pid N, jar ..." result line reads it afterwards) and
|
||||||
|
# HEALTH_BODY/HEALTH_CODE (globals, the same reason report_health already needs them handed back).
|
||||||
|
# Only ever called as a bare statement in the main flow below, never from inside a `$( )`: a die()
|
||||||
|
# reached from inside a command substitution only kills that subshell, not the whole script, which
|
||||||
|
# would silently turn a genuine failure back into a false "succeeded" exit (see poll_health_body's
|
||||||
|
# own `|| true` idiom for the same hazard from the other direction).
|
||||||
|
await_daemon_started() {
|
||||||
|
local health_wait="$1" old_pid="$2" health_url="$3" out_file="$4"
|
||||||
|
NEW_PID=""
|
||||||
|
if wait_for_new_pid "$health_wait"; then
|
||||||
|
NEW_PID="$(running_pid)"
|
||||||
|
[ "$NEW_PID" != "${old_pid:-}" ] || die "pid unchanged ($NEW_PID) — the old daemon never died"
|
||||||
|
ok "started, pid $NEW_PID"
|
||||||
|
else
|
||||||
|
warn "no process matched $PATTERN within ${health_wait}s of starting — falling through to the health check, which is the more truthful signal"
|
||||||
|
fi
|
||||||
|
|
||||||
|
HEALTH_BODY="$(poll_health_body "$health_url" "$health_wait")" || true
|
||||||
|
HEALTH_CODE="000"
|
||||||
|
[ -n "$HEALTH_BODY" ] || HEALTH_CODE="$(curl -s -o /dev/null -w '%{http_code}' --max-time 2 "$health_url" 2>/dev/null || echo 000)"
|
||||||
|
report_health "$HEALTH_BODY" "$HEALTH_CODE" "$out_file" "$health_wait"
|
||||||
|
|
||||||
|
# By now /healthz has answered (report_health above would have died otherwise), so the daemon is
|
||||||
|
# confirmed up even if the pid poll never matched it — see running_pid()'s own comment on
|
||||||
|
# under-counting if the launch method ever stops being a plain `java -jar`. Fill NEW_PID in for the
|
||||||
|
# result line rather than leave it blank on an otherwise fully successful redeploy.
|
||||||
|
[ -n "$NEW_PID" ] || NEW_PID="$(running_pid)"
|
||||||
|
}
|
||||||
|
|
||||||
# Ticket item 8 — `HAD_OLD_PID=0; [ -n "$OLD_PID" ] && HAD_OLD_PID=1`. Measured safe under `set -e`
|
# Ticket item 8 — `HAD_OLD_PID=0; [ -n "$OLD_PID" ] && HAD_OLD_PID=1`. Measured safe under `set -e`
|
||||||
# at both bash 3.2.57 and 5.x (see the header comment trap 9 discussion in the ticket) — not a `set
|
# at both bash 3.2.57 and 5.x (see the header comment trap 9 discussion in the ticket) — not a `set
|
||||||
# -e` hazard, but still an untested computation feeding report_shutdown_drain's own four-way
|
# -e` hazard, but still an untested computation feeding report_shutdown_drain's own four-way
|
||||||
@@ -1234,28 +1356,14 @@ say "start"
|
|||||||
# there now, so this line is the only thing left in the main flow to get wrong.
|
# there now, so this line is the only thing left in the main flow to get wrong.
|
||||||
dispatch_start "$SUPERVISOR_KIND"
|
dispatch_start "$SUPERVISOR_KIND"
|
||||||
|
|
||||||
for _ in $(seq 10); do
|
|
||||||
NEW_PID="$(running_pid)"
|
|
||||||
[ -n "$NEW_PID" ] && break
|
|
||||||
sleep 1
|
|
||||||
done
|
|
||||||
[ -n "${NEW_PID:-}" ] || die "no process appeared. Last lines of $OUT:
|
|
||||||
$(tail -20 "$OUT" 2>/dev/null)"
|
|
||||||
[ "$NEW_PID" != "${OLD_PID:-}" ] || die "pid unchanged ($NEW_PID) — the old daemon never died"
|
|
||||||
ok "started, pid $NEW_PID"
|
|
||||||
|
|
||||||
# ------------------------------------------------------------------ verify
|
# ------------------------------------------------------------------ verify
|
||||||
|
|
||||||
say "verify"
|
say "verify"
|
||||||
|
|
||||||
# fleetd #555: poll_health_body/health_is_up/report_health above. HEALTH_CODE is only ever
|
# fleetd #603: await_daemon_started above folds the pid-appeared check and the healthz check into
|
||||||
# consulted by report_health when the body came back empty; `|| true` on both assignments is the
|
# one decision — see its own comment for why a bare `if` here, split back into the old two pieces,
|
||||||
# same "an absent/failing command substitution must not kill the script under set -e" idiom the
|
# would put back an untestable branch this ticket exists to close.
|
||||||
# swap/drain helpers already rely on (see poll_health_body's own comment).
|
await_daemon_started "$HEALTH_WAIT" "$OLD_PID" "$HEALTH" "$OUT"
|
||||||
HEALTH_BODY="$(poll_health_body "$HEALTH" "$HEALTH_WAIT")" || true
|
|
||||||
HEALTH_CODE="000"
|
|
||||||
[ -n "$HEALTH_BODY" ] || HEALTH_CODE="$(curl -s -o /dev/null -w '%{http_code}' --max-time 2 "$HEALTH" 2>/dev/null || echo 000)"
|
|
||||||
report_health "$HEALTH_BODY" "$HEALTH_CODE" "$OUT" "$HEALTH_WAIT"
|
|
||||||
|
|
||||||
# A fresh listening line, strictly after the restart mark. An old daemon that never died would
|
# A fresh listening line, strictly after the restart mark. An old daemon that never died would
|
||||||
# otherwise let an old line pass for a new one.
|
# otherwise let an old line pass for a new one.
|
||||||
@@ -1310,7 +1418,7 @@ report_shutdown_drain "$FRESH_LOG" "$HAD_OLD_PID"
|
|||||||
assert_single_daemon "$(running_pid)"
|
assert_single_daemon "$(running_pid)"
|
||||||
|
|
||||||
say "result"
|
say "result"
|
||||||
ok "pid $NEW_PID, jar $(jar_id)"
|
ok "pid ${NEW_PID:-unknown}, jar $(jar_id)"
|
||||||
if [ "$REDEPLOY_AMQP_CHECK_SKIPPED" -eq 1 ]; then
|
if [ "$REDEPLOY_AMQP_CHECK_SKIPPED" -eq 1 ]; then
|
||||||
# fleetd #552: the fourth reader of the fresh-log region. Without this branch REDEPLOY_ERROR_COUNT
|
# fleetd #552: the fourth reader of the fresh-log region. Without this branch REDEPLOY_ERROR_COUNT
|
||||||
# stays at its untouched 0 (classify_amqp_connection_errors never got a file to read) and
|
# stays at its untouched 0 (classify_amqp_connection_errors never got a file to read) and
|
||||||
|
|||||||
Executable
+585
@@ -0,0 +1,585 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Self-contained checks for scripts/config-edit.sh — fleetd ticket #635.
|
||||||
|
#
|
||||||
|
# Drives the REAL config-edit.sh as a subprocess against a FIXTURE config and a FIXTURE log in a
|
||||||
|
# throwaway temp directory this file creates and removes. Never touches fleetd/fleetd.yaml or
|
||||||
|
# fleetd/fleetd.out, and never starts, stops, or contacts a daemon — there is no daemon here, so
|
||||||
|
# each test PLAYS the daemon: it starts config-edit.sh in the background (it is waiting on the
|
||||||
|
# log), appends the verdict line it wants, then collects the real exit code.
|
||||||
|
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||||
|
EDIT="$ROOT/scripts/config-edit.sh"
|
||||||
|
TMP="$(mktemp -d "$ROOT/.config-edit-test.XXXXXX")"
|
||||||
|
trap 'rm -rf "$TMP"' EXIT
|
||||||
|
|
||||||
|
fail() {
|
||||||
|
printf 'FAIL: %s\n' "$*" >&2
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
assert_equals() {
|
||||||
|
local expected="$1" actual="$2" description="$3"
|
||||||
|
[ "$expected" = "$actual" ] || fail "$description: expected $expected, got $actual"
|
||||||
|
}
|
||||||
|
|
||||||
|
assert_contains() {
|
||||||
|
local needle="$1" text="$2" description="$3"
|
||||||
|
printf '%s' "$text" | grep -qF -- "$needle" || fail "$description: missing [$needle]"
|
||||||
|
}
|
||||||
|
|
||||||
|
assert_not_contains() {
|
||||||
|
local needle="$1" text="$2" description="$3"
|
||||||
|
if printf '%s' "$text" | grep -qF -- "$needle"; then
|
||||||
|
fail "$description: must NOT contain [$needle], but it does"
|
||||||
|
fi
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
|
# A fresh fixture pair per test: $1/fleetd.yaml (the config) and $1/fleetd.out (the log), plus a
|
||||||
|
# small wait-seconds budget so no test takes long. Returns the fixture dir via stdout.
|
||||||
|
new_fixture() {
|
||||||
|
local dir
|
||||||
|
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
|
||||||
|
cat > "$dir/fleetd.yaml" <<'YAML'
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 19999
|
||||||
|
broker:
|
||||||
|
uri: amqp://user:hunter2@host/vhost
|
||||||
|
profiles:
|
||||||
|
sonnet:
|
||||||
|
weight: 3
|
||||||
|
maxLoad: 5
|
||||||
|
YAML
|
||||||
|
: > "$dir/fleetd.out"
|
||||||
|
printf '%s' "$dir"
|
||||||
|
}
|
||||||
|
|
||||||
|
# Runs config-edit.sh in the background against $dir's fixtures, with the given extra args, and
|
||||||
|
# a short --wait-seconds. Sets RUN_PID. Caller appends to $dir/fleetd.out (or not, for the
|
||||||
|
# silence test) and then calls collect_run to block for the exit code.
|
||||||
|
start_run() {
|
||||||
|
local dir="$1" wait_s="$2"; shift 2
|
||||||
|
(
|
||||||
|
# config-edit.sh deliberately exits 3/4/5 on several of these tests. This subshell inherits
|
||||||
|
# the parent's `set -e`, and without disabling it here the FIRST nonzero exit would kill the
|
||||||
|
# subshell before the `echo $? > rc` line ever ran — the real code would never reach the file.
|
||||||
|
set +e
|
||||||
|
"$EDIT" --config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds "$wait_s" "$@" \
|
||||||
|
> "$dir/stdout.log" 2>&1
|
||||||
|
echo $? > "$dir/rc"
|
||||||
|
) &
|
||||||
|
RUN_PID=$!
|
||||||
|
}
|
||||||
|
|
||||||
|
collect_run() {
|
||||||
|
local dir="$1"
|
||||||
|
# wait echoes back the backgrounded subshell's own exit status (here, deliberately 3/4/5 on
|
||||||
|
# several tests) — under `set -e` a bare nonzero `wait` would abort this whole test script, so
|
||||||
|
# it is neutralized with `|| true`; the real code is read from $dir/rc right after.
|
||||||
|
wait "$RUN_PID" || true
|
||||||
|
RUN_OUTPUT="$(cat "$dir/stdout.log")"
|
||||||
|
RUN_RC="$(cat "$dir/rc")"
|
||||||
|
}
|
||||||
|
|
||||||
|
# -------------------------------------------------------------- acceptance criterion 1: refusal
|
||||||
|
test_refusal_restores_byte_for_byte() {
|
||||||
|
local dir
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
|
||||||
|
|
||||||
|
start_run "$dir" 5 --set '.broker.uri=amqp://changed@host/x'
|
||||||
|
sleep 1
|
||||||
|
printf 'config reload refused — these keys cannot change under a running daemon: broker. Restart fleetd to apply them.\n' >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
|
||||||
|
assert_equals 4 "$RUN_RC" "refusal exit code"
|
||||||
|
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|
||||||
|
|| fail "refusal must restore the config byte for byte onto the pre-edit backup"
|
||||||
|
}
|
||||||
|
|
||||||
|
# -------------------------------------------------------------- acceptance criterion 2: clean
|
||||||
|
test_clean_reload_keeps_the_edit() {
|
||||||
|
local dir
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
|
||||||
|
start_run "$dir" 5 --set '.profiles.sonnet.weight=7'
|
||||||
|
sleep 1
|
||||||
|
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
|
||||||
|
assert_equals 0 "$RUN_RC" "clean reload exit code"
|
||||||
|
assert_equals "7" "$(yq eval '.profiles.sonnet.weight' "$dir/fleetd.yaml")" "clean reload live value"
|
||||||
|
}
|
||||||
|
|
||||||
|
# ----------------------------------------------------- acceptance criterion 3: deferred != clean
|
||||||
|
test_deferred_reload_is_told_apart_from_clean() {
|
||||||
|
local dir
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
|
||||||
|
start_run "$dir" 5 --set '.profiles.sonnet.weight=9'
|
||||||
|
sleep 1
|
||||||
|
printf 'config reloaded; these changes need a restart to take effect: profiles\n' >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
|
||||||
|
assert_equals 3 "$RUN_RC" "deferred reload exit code"
|
||||||
|
[ "$RUN_RC" != 0 ] || fail "deferred reload must not report exit 0"
|
||||||
|
assert_contains "profiles" "$RUN_OUTPUT" "deferred reload names the key"
|
||||||
|
assert_contains "restart" "$RUN_OUTPUT" "deferred reload says a restart is needed"
|
||||||
|
}
|
||||||
|
|
||||||
|
# -------------------------------------------------------------- acceptance criterion 4: silence
|
||||||
|
test_silence_is_its_own_answer() {
|
||||||
|
local dir
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
|
||||||
|
start_run "$dir" 2 --set '.profiles.sonnet.weight=11'
|
||||||
|
# Feed the log nothing.
|
||||||
|
collect_run "$dir"
|
||||||
|
|
||||||
|
assert_equals 5 "$RUN_RC" "silence exit code"
|
||||||
|
assert_equals "11" "$(yq eval '.profiles.sonnet.weight' "$dir/fleetd.yaml")" "the edited value must still be on disk"
|
||||||
|
assert_contains '--restore' "$RUN_OUTPUT" "silence prints the --restore command"
|
||||||
|
|
||||||
|
local backup restore_cmd
|
||||||
|
backup="$(ls -t "$dir"/.config-backups/fleetd.yaml.bak.* | head -1)"
|
||||||
|
[ -n "$backup" ] || fail "silence must still have taken a backup"
|
||||||
|
|
||||||
|
restore_cmd="$(printf '%s\n' "$RUN_OUTPUT" | grep -F -- '--restore --config' | sed -E 's/^[[:space:]]*//')"
|
||||||
|
[ -n "$restore_cmd" ] || fail "could not find the printed --restore invocation in the output"
|
||||||
|
# Running this --restore invocation installs the backup, then itself waits for a confirming
|
||||||
|
# verdict that this fixture never feeds — so it legitimately exits 5 ("cannot tell") here, same
|
||||||
|
# as any edit with no daemon on the other end. Only a usage/internal error (1 or 2) is a real
|
||||||
|
# failure of the command itself; the actual assertion is the byte-for-byte cmp below.
|
||||||
|
local restore_rc=0
|
||||||
|
eval "$restore_cmd" > "$dir/restore.log" 2>&1 || restore_rc=$?
|
||||||
|
case "$restore_rc" in
|
||||||
|
0|3|4|5) : ;;
|
||||||
|
*) fail "the printed --restore command errored out (exit $restore_rc): $(cat "$dir/restore.log")" ;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
cmp -s "$dir/fleetd.yaml" "$backup" \
|
||||||
|
|| fail "running the printed --restore command must put the file back to the original backup"
|
||||||
|
}
|
||||||
|
|
||||||
|
# --------------------------------------------------------- acceptance criterion 5: bad candidate
|
||||||
|
test_broken_candidate_never_reaches_live_path() {
|
||||||
|
local dir rc=0
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
printf 'foo: [unclosed\n' > "$dir/broken.yaml"
|
||||||
|
|
||||||
|
"$EDIT" --from "$dir/broken.yaml" --config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
|
||||||
|
> "$dir/stdout.log" 2>&1 || rc=$?
|
||||||
|
|
||||||
|
[ "$rc" -ne 0 ] || fail "a broken --from candidate must exit non-zero"
|
||||||
|
cmp -s "$dir/fleetd.yaml" <(new_fixture_yaml) \
|
||||||
|
|| fail "the broken candidate must never reach the live fixture config"
|
||||||
|
}
|
||||||
|
new_fixture_yaml() {
|
||||||
|
cat <<'YAML'
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 19999
|
||||||
|
broker:
|
||||||
|
uri: amqp://user:hunter2@host/vhost
|
||||||
|
profiles:
|
||||||
|
sonnet:
|
||||||
|
weight: 3
|
||||||
|
maxLoad: 5
|
||||||
|
YAML
|
||||||
|
}
|
||||||
|
|
||||||
|
# -------------------------------------------------------------------- acceptance criterion 6
|
||||||
|
test_marker_skips_lines_before_it() {
|
||||||
|
local dir
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
printf 'config reload refused — something ancient\n' > "$dir/fleetd.out"
|
||||||
|
|
||||||
|
start_run "$dir" 5 --set '.profiles.sonnet.weight=5'
|
||||||
|
sleep 1
|
||||||
|
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
|
||||||
|
assert_equals 0 "$RUN_RC" "a stale refusal before the marker must not be read as this edit's verdict"
|
||||||
|
}
|
||||||
|
|
||||||
|
# ------------------------------------------------------------------- acceptance criterion 7 (+13)
|
||||||
|
# fleetd #635 follow-up (ticket comment 17659) — the two assertions below this comment were the
|
||||||
|
# WHOLE test before the follow-up, and both are negative-only: they pass just as happily when the
|
||||||
|
# diff is never printed at all as when it is printed and correctly redacted. A mutant that deletes
|
||||||
|
# `diff -u "$backup" "$cand" | redact` from the edit path survives them, because an absent output
|
||||||
|
# contains neither "hunter2" nor "user:" either — see the mutation-and-revert proof in the reply.
|
||||||
|
# Criterion 13 is the fix: a LOUD positive control that only passes when a diff was demonstrably
|
||||||
|
# printed AND the redaction demonstrably ran on real content, not merely that nothing leaked.
|
||||||
|
test_redaction_holds() {
|
||||||
|
local dir
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
|
||||||
|
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
|
||||||
|
sleep 1
|
||||||
|
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
|
||||||
|
assert_equals 0 "$RUN_RC" "redaction-case reload exit code"
|
||||||
|
assert_not_contains "hunter2" "$RUN_OUTPUT" "full output must never contain the password"
|
||||||
|
assert_not_contains "user:" "$RUN_OUTPUT" "full output must never contain the userinfo"
|
||||||
|
# acceptance criterion 13 — positive control: the diff's default 3-line context around the
|
||||||
|
# changed "weight" key also covers the fixture's "uri:" line, so a genuinely-printed, genuinely-
|
||||||
|
# redacted diff must contain BOTH the redaction marker and the changed key's name. A test that
|
||||||
|
# only ever asserts absence cannot tell "redacted" from "never printed" apart; this can.
|
||||||
|
assert_contains "<redacted>" "$RUN_OUTPUT" "the redaction must be PROVEN to have run on real content, not merely absent"
|
||||||
|
assert_contains "weight" "$RUN_OUTPUT" "a diff must have been demonstrably printed at all"
|
||||||
|
}
|
||||||
|
|
||||||
|
# ------------------------------------------------------- acceptance criterion 9: forgotten value
|
||||||
|
# `--set .a.b=` is a plausible typo (the value simply forgotten), and it must be refused outright
|
||||||
|
# rather than silently nulling the field — a null numeric field falls back to its default, which
|
||||||
|
# widens capacity instead of failing loudly. No background verdict feeder here: a refused --set
|
||||||
|
# must never even reach the daemon, so this never starts a background run at all.
|
||||||
|
test_forgotten_value_refuses_and_installs_nothing() {
|
||||||
|
local dir rc=0
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
|
||||||
|
|
||||||
|
"$EDIT" --set '.profiles.sonnet.maxLoad=' \
|
||||||
|
--config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
|
||||||
|
> "$dir/stdout.log" 2>&1 || rc=$?
|
||||||
|
RUN_OUTPUT="$(cat "$dir/stdout.log")"
|
||||||
|
|
||||||
|
[ "$rc" -ne 0 ] || fail "an empty --set value must exit non-zero, got 0"
|
||||||
|
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|
||||||
|
|| fail "an empty --set value must install nothing — the live fixture changed"
|
||||||
|
assert_contains "EMPTY value" "$RUN_OUTPUT" "the refusal must name the empty value"
|
||||||
|
}
|
||||||
|
|
||||||
|
# ---------------------------------------------------------- acceptance criterion 10: explicit null
|
||||||
|
# `--set .a.b=null` is the deliberate-clear spelling, and it must write a REAL yaml null, never
|
||||||
|
# the string "''" — those are different values to the daemon's loader (fleetd ticket #635's
|
||||||
|
# follow-up comment measured `""` reading back as a null field anyway, which is exactly why the
|
||||||
|
# two forms must not collapse onto each other: `--set path=` refuses instead of silently reaching
|
||||||
|
# this same null outcome through the back door). Read the RAW line with grep, never only through
|
||||||
|
# `yq` — `yq eval` reports `null` for both an actual null and a missing/absent key, so it cannot
|
||||||
|
# tell "wrote null" apart from "wrote nothing"; only the literal line on disk can.
|
||||||
|
test_explicit_null_writes_bare_null_not_empty_string() {
|
||||||
|
local dir
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
|
||||||
|
start_run "$dir" 5 --set '.profiles.sonnet.maxLoad=null'
|
||||||
|
sleep 1
|
||||||
|
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
|
||||||
|
assert_equals 0 "$RUN_RC" "explicit null clear exit code"
|
||||||
|
local raw_line
|
||||||
|
raw_line="$(grep -E 'maxLoad' "$dir/fleetd.yaml")"
|
||||||
|
assert_contains "null" "$raw_line" "the installed line must spell a bare null"
|
||||||
|
assert_not_contains '""' "$raw_line" "the installed line must NOT be a quoted empty string"
|
||||||
|
}
|
||||||
|
|
||||||
|
# ------------------------------------------------------- acceptance criterion 11: backup never committable
|
||||||
|
# A backup of fleetd.yaml inherits fleetd.yaml's own "never commit this" requirement (fleetd #635
|
||||||
|
# follow-up, ticket comment 17655). Proves two things: the backup lands somewhere `git
|
||||||
|
# check-ignore` reports as ignored (equivalently, a path `git status --porcelain` never lists as
|
||||||
|
# untracked), AND that --restore still finds and uses it from that location.
|
||||||
|
test_backup_is_never_committable() {
|
||||||
|
local dir backup
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
|
||||||
|
|
||||||
|
start_run "$dir" 5 --set '.profiles.sonnet.weight=55'
|
||||||
|
sleep 1
|
||||||
|
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
assert_equals 0 "$RUN_RC" "setup edit exit code"
|
||||||
|
|
||||||
|
backup="$(ls -t "$dir"/.config-backups/fleetd.yaml.bak.* 2>/dev/null | head -1)"
|
||||||
|
[ -n "$backup" ] || fail "no backup found under .config-backups/ — did the location change?"
|
||||||
|
|
||||||
|
git -C "$ROOT" check-ignore -q -- "$backup" \
|
||||||
|
|| fail "the backup at $backup is NOT gitignored — it would survive a git add -A"
|
||||||
|
if git -C "$ROOT" status --porcelain -- "$backup" 2>/dev/null | grep -q '^??'; then
|
||||||
|
fail "git status still lists the backup as untracked: $backup"
|
||||||
|
fi
|
||||||
|
|
||||||
|
start_run "$dir" 5 --restore
|
||||||
|
sleep 1
|
||||||
|
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
assert_equals 0 "$RUN_RC" "--restore after the backup-location change exit code"
|
||||||
|
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|
||||||
|
|| fail "--restore from the new backup location must still put the file back byte for byte"
|
||||||
|
}
|
||||||
|
|
||||||
|
# --------------------------------------------------------- acceptance criterion 12: file mode
|
||||||
|
# `mv` from a mktemp candidate carries mktemp's 0600 forever, and a plain `cp` onto an existing
|
||||||
|
# file keeps the DESTINATION's mode rather than the source's, so a restore does not undo the
|
||||||
|
# narrowing either (fleetd #635 follow-up, ticket comment 17657). Proves the mode survives an edit
|
||||||
|
# AND a subsequent restore, from two different starting points — 644 is the common case, 600
|
||||||
|
# proves the fix PRESERVES whatever mode was there rather than hardcoding 644.
|
||||||
|
test_file_mode_survives_edit_and_restore() {
|
||||||
|
local dir want got
|
||||||
|
for want in 644 600; do
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
chmod "$want" "$dir/fleetd.yaml"
|
||||||
|
|
||||||
|
start_run "$dir" 5 --set ".profiles.sonnet.weight=${want}"
|
||||||
|
sleep 1
|
||||||
|
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
assert_equals 0 "$RUN_RC" "mode-preservation setup edit exit code ($want)"
|
||||||
|
got="$(stat -f '%Lp' "$dir/fleetd.yaml" 2>/dev/null || stat -c '%a' "$dir/fleetd.yaml")"
|
||||||
|
assert_equals "$want" "$got" "mode must survive a --set ($want)"
|
||||||
|
|
||||||
|
start_run "$dir" 5 --restore
|
||||||
|
sleep 1
|
||||||
|
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
assert_equals 0 "$RUN_RC" "mode-preservation restore exit code ($want)"
|
||||||
|
got="$(stat -f '%Lp' "$dir/fleetd.yaml" 2>/dev/null || stat -c '%a' "$dir/fleetd.yaml")"
|
||||||
|
assert_equals "$want" "$got" "mode must survive a --restore ($want)"
|
||||||
|
done
|
||||||
|
}
|
||||||
|
|
||||||
|
# ----------------------------------------- acceptance criterion 14: restore message names the real directory
|
||||||
|
# fleetd #635 follow-up (ticket comment 17664, defect 6) — the --restore "no backup found"
|
||||||
|
# message used to print the OLD beside-the-config glob even though newest_backup had already
|
||||||
|
# moved to searching the managed directory. Proves BOTH directions: the not-found message names
|
||||||
|
# the directory actually searched (not merely that it says SOMETHING), and that a real backup
|
||||||
|
# sitting in that directory still lets --restore succeed — otherwise the fix could regress into
|
||||||
|
# a message that is always printed regardless of whether a backup exists.
|
||||||
|
test_restore_message_names_the_searched_directory() {
|
||||||
|
local dir rc=0
|
||||||
|
|
||||||
|
# Direction 1: no backup anywhere — the message must name .config-backups/, not the bare
|
||||||
|
# beside-the-config glob the OLD code printed.
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
"$EDIT" --restore --config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
|
||||||
|
> "$dir/stdout.log" 2>&1 || rc=$?
|
||||||
|
RUN_OUTPUT="$(cat "$dir/stdout.log")"
|
||||||
|
|
||||||
|
assert_equals 1 "$rc" "--restore with no backup anywhere exit code"
|
||||||
|
assert_contains ".config-backups/fleetd.yaml.bak.*" "$RUN_OUTPUT" \
|
||||||
|
"the not-found message must name the directory actually searched, not the old beside-the-config glob"
|
||||||
|
|
||||||
|
# Direction 2: a real backup IS present in .config-backups/ — --restore must still succeed, so
|
||||||
|
# the message fix cannot have turned into one that prints regardless of whether a backup exists.
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
|
||||||
|
start_run "$dir" 5 --set '.profiles.sonnet.weight=77'
|
||||||
|
sleep 1
|
||||||
|
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
assert_equals 0 "$RUN_RC" "setup edit exit code for criterion 14's second half"
|
||||||
|
|
||||||
|
start_run "$dir" 5 --restore
|
||||||
|
sleep 1
|
||||||
|
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
assert_equals 0 "$RUN_RC" "--restore with a real backup present must still succeed"
|
||||||
|
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|
||||||
|
|| fail "--restore with a real backup present must put the file back byte for byte"
|
||||||
|
}
|
||||||
|
|
||||||
|
# ------------------------- acceptance criterion 15a: block-scalar continuation lines are redacted
|
||||||
|
# fleetd #635 follow-up (ticket comment 17670, defect 7) — redact() used to look only AT the key
|
||||||
|
# line. A YAML block scalar (`|`) puts its value on the lines that FOLLOW the key, each indented
|
||||||
|
# deeper than it, so the real secret flowed through untouched while the key line right above it
|
||||||
|
# printed a reassuring "<redacted>" — worse than no redaction, because the marker stops a reader
|
||||||
|
# from looking further. The edited key here ("retries") sits directly next to the block scalar,
|
||||||
|
# well inside diff -u's default 3-line context window, so the printed hunk is GUARANTEED to
|
||||||
|
# include the secret's lines — placing the edit further away would let this pass today even
|
||||||
|
# without the fix, proving nothing (the ticket comment's own warning, from the lead's first
|
||||||
|
# reproduction attempt). The positive control runs FIRST: without it, "the secret never entered
|
||||||
|
# the diff at all" would pass identically to "it entered and was correctly redacted".
|
||||||
|
new_fixture_block_scalar() {
|
||||||
|
local dir
|
||||||
|
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
|
||||||
|
cat > "$dir/fleetd.yaml" <<'YAML'
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 19999
|
||||||
|
broker:
|
||||||
|
uri: amqp://user:hunter2@host/vhost
|
||||||
|
auth:
|
||||||
|
token: |
|
||||||
|
FAKELEAK-BLOCK-SCALAR
|
||||||
|
retries: 1
|
||||||
|
profiles:
|
||||||
|
sonnet:
|
||||||
|
weight: 3
|
||||||
|
maxLoad: 5
|
||||||
|
YAML
|
||||||
|
: > "$dir/fleetd.out"
|
||||||
|
printf '%s' "$dir"
|
||||||
|
}
|
||||||
|
|
||||||
|
test_block_scalar_continuation_is_redacted() {
|
||||||
|
local dir
|
||||||
|
dir="$(new_fixture_block_scalar)"
|
||||||
|
|
||||||
|
start_run "$dir" 5 --set '.auth.retries=2'
|
||||||
|
sleep 1
|
||||||
|
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
|
||||||
|
assert_equals 0 "$RUN_RC" "block-scalar case reload exit code"
|
||||||
|
# Positive control FIRST: the key's own (masked) line must really be in the printed diff, or the
|
||||||
|
# negative assertion right after proves nothing — see the comment above this test.
|
||||||
|
assert_contains "token:" "$RUN_OUTPUT" "block-scalar case: the key's line must be in the printed diff"
|
||||||
|
assert_contains "<redacted>" "$RUN_OUTPUT" "block-scalar case: redaction must be proven to have run on real content"
|
||||||
|
assert_not_contains "FAKELEAK-BLOCK-SCALAR" "$RUN_OUTPUT" "block-scalar case: the block scalar's VALUE must never leak"
|
||||||
|
}
|
||||||
|
|
||||||
|
# ------------------------------------- acceptance criterion 15b: "passphrase" is also recognised
|
||||||
|
# "passphrase" was in none of TOKEN|SECRET|PASSWORD|PASSWD|CREDENTIAL|URI|_KEY (ticket comment
|
||||||
|
# 17670). This is a plain key:value line, not a block scalar — kept in its OWN fixture and OWN
|
||||||
|
# function, separate from criterion 15a, so that a failure in one case can never mask a failure in
|
||||||
|
# the other (a single combined test would abort under `set -e` at its first failing assertion,
|
||||||
|
# and the second case would then never even run).
|
||||||
|
new_fixture_passphrase() {
|
||||||
|
local dir
|
||||||
|
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
|
||||||
|
cat > "$dir/fleetd.yaml" <<'YAML'
|
||||||
|
bind:
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 19999
|
||||||
|
broker:
|
||||||
|
uri: amqp://user:hunter2@host/vhost
|
||||||
|
auth:
|
||||||
|
passphrase: FAKELEAK-PASSPHRASE
|
||||||
|
retries: 1
|
||||||
|
profiles:
|
||||||
|
sonnet:
|
||||||
|
weight: 3
|
||||||
|
maxLoad: 5
|
||||||
|
YAML
|
||||||
|
: > "$dir/fleetd.out"
|
||||||
|
printf '%s' "$dir"
|
||||||
|
}
|
||||||
|
|
||||||
|
test_passphrase_key_is_redacted() {
|
||||||
|
local dir
|
||||||
|
dir="$(new_fixture_passphrase)"
|
||||||
|
|
||||||
|
start_run "$dir" 5 --set '.auth.retries=2'
|
||||||
|
sleep 1
|
||||||
|
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
|
||||||
|
assert_equals 0 "$RUN_RC" "passphrase case reload exit code"
|
||||||
|
assert_contains "passphrase:" "$RUN_OUTPUT" "passphrase case: the key's line must be in the printed diff"
|
||||||
|
assert_contains "<redacted>" "$RUN_OUTPUT" "passphrase case: redaction must be proven to have run on real content"
|
||||||
|
assert_not_contains "FAKELEAK-PASSPHRASE" "$RUN_OUTPUT" "passphrase case: the passphrase VALUE must never leak"
|
||||||
|
}
|
||||||
|
|
||||||
|
# ----------------------------------- acceptance criterion 16: a failing --set must not echo value
|
||||||
|
# fleetd #635 follow-up (ticket comment 17673, defect 8) — apply_set_pairs used to echo the FULL
|
||||||
|
# "$kv" (path=value, exactly as typed) in its yq-failure messages, so a broken --set with a
|
||||||
|
# secret-looking value printed that value right back out. The path alone is what the positive
|
||||||
|
# control proves is still there — it is what the operator needs to fix their command — and the
|
||||||
|
# negative assertion proves the value itself never appears. Kept to exactly this one failure
|
||||||
|
# shape (an invalid yq path/expression), matching the ticket's own reproduction.
|
||||||
|
test_failing_set_does_not_echo_its_value() {
|
||||||
|
local dir rc=0
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
|
||||||
|
|
||||||
|
"$EDIT" --dry-run --set '.broker.["bad=FAKELEAK-SETVALUE' \
|
||||||
|
--config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
|
||||||
|
> "$dir/stdout.log" 2>&1 || rc=$?
|
||||||
|
RUN_OUTPUT="$(cat "$dir/stdout.log")"
|
||||||
|
|
||||||
|
[ "$rc" -ne 0 ] || fail "a --set with an invalid yq expression must exit non-zero, got 0"
|
||||||
|
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|
||||||
|
|| fail "a failing --set must install nothing — the live fixture changed"
|
||||||
|
# Positive control FIRST: the path must still be in the message, or the negative assertion right
|
||||||
|
# after proves nothing (the message could simply have disappeared entirely).
|
||||||
|
assert_contains '.broker.["bad' "$RUN_OUTPUT" "the failure message must still name the PATH"
|
||||||
|
assert_not_contains "FAKELEAK-SETVALUE" "$RUN_OUTPUT" "the failure message must NEVER echo the VALUE"
|
||||||
|
}
|
||||||
|
|
||||||
|
# dry-run must never touch the live file and must still redact.
|
||||||
|
test_dry_run_never_installs_and_redacts() {
|
||||||
|
local dir before
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
before="$(cat "$dir/fleetd.yaml")"
|
||||||
|
|
||||||
|
"$EDIT" --dry-run --set '.profiles.sonnet.weight=99' \
|
||||||
|
--config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
|
||||||
|
> "$dir/stdout.log" 2>&1
|
||||||
|
local rc=$?
|
||||||
|
RUN_OUTPUT="$(cat "$dir/stdout.log")"
|
||||||
|
|
||||||
|
assert_equals 0 "$rc" "dry-run exit code"
|
||||||
|
assert_equals "$before" "$(cat "$dir/fleetd.yaml")" "dry-run must never write the live config"
|
||||||
|
assert_not_contains "hunter2" "$RUN_OUTPUT" "dry-run diff must also be redacted"
|
||||||
|
assert_contains "99" "$RUN_OUTPUT" "dry-run diff must show the candidate value"
|
||||||
|
# Same positive-control reasoning as acceptance criterion 13, applied to the dry-run diff path.
|
||||||
|
assert_contains "<redacted>" "$RUN_OUTPUT" "the dry-run diff's redaction must be PROVEN to have run, not merely absent"
|
||||||
|
}
|
||||||
|
|
||||||
|
# --check is read-only and always exits 0, even against a dead "daemon".
|
||||||
|
test_check_is_read_only_and_exits_zero() {
|
||||||
|
local dir before rc=0
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
before="$(cat "$dir/fleetd.yaml")"
|
||||||
|
|
||||||
|
"$EDIT" --check --config "$dir/fleetd.yaml" --log "$dir/fleetd.out" \
|
||||||
|
> "$dir/stdout.log" 2>&1 || rc=$?
|
||||||
|
|
||||||
|
assert_equals 0 "$rc" "--check exit code"
|
||||||
|
assert_equals "$before" "$(cat "$dir/fleetd.yaml")" "--check must never modify the config"
|
||||||
|
}
|
||||||
|
|
||||||
|
test_refusal_shape_from_parse_failure_wording_is_recognised() {
|
||||||
|
local dir
|
||||||
|
dir="$(new_fixture)"
|
||||||
|
|
||||||
|
start_run "$dir" 5 --set '.profiles.sonnet.weight=6'
|
||||||
|
sleep 1
|
||||||
|
printf 'config reload from %s refused, keeping the running config: boom\n' "$dir/fleetd.yaml" >> "$dir/fleetd.out"
|
||||||
|
collect_run "$dir"
|
||||||
|
|
||||||
|
assert_equals 4 "$RUN_RC" "the parse-failure refusal shape must also exit 4, not be read as silence"
|
||||||
|
}
|
||||||
|
|
||||||
|
echo "== acceptance criterion 1: refusal restores byte for byte =="
|
||||||
|
test_refusal_restores_byte_for_byte
|
||||||
|
echo "== acceptance criterion 2: clean reload keeps the edit =="
|
||||||
|
test_clean_reload_keeps_the_edit
|
||||||
|
echo "== acceptance criterion 3: deferred reload told apart from clean =="
|
||||||
|
test_deferred_reload_is_told_apart_from_clean
|
||||||
|
echo "== acceptance criterion 4: silence is its own answer =="
|
||||||
|
test_silence_is_its_own_answer
|
||||||
|
echo "== acceptance criterion 5: broken candidate never reaches the live path =="
|
||||||
|
test_broken_candidate_never_reaches_live_path
|
||||||
|
echo "== acceptance criterion 6: the marker works =="
|
||||||
|
test_marker_skips_lines_before_it
|
||||||
|
echo "== acceptance criterion 7 (+13: redaction is proven to have run) =="
|
||||||
|
test_redaction_holds
|
||||||
|
echo "== acceptance criterion 9: a forgotten value refuses and installs nothing =="
|
||||||
|
test_forgotten_value_refuses_and_installs_nothing
|
||||||
|
echo "== acceptance criterion 10: an explicit clear writes a bare null =="
|
||||||
|
test_explicit_null_writes_bare_null_not_empty_string
|
||||||
|
echo "== acceptance criterion 11: a backup is never committable =="
|
||||||
|
test_backup_is_never_committable
|
||||||
|
echo "== acceptance criterion 12: the file mode survives an edit and a restore =="
|
||||||
|
test_file_mode_survives_edit_and_restore
|
||||||
|
echo "== acceptance criterion 14: the restore message names the directory actually searched =="
|
||||||
|
test_restore_message_names_the_searched_directory
|
||||||
|
echo "== acceptance criterion 15a: a block scalar's continuation lines are redacted =="
|
||||||
|
test_block_scalar_continuation_is_redacted
|
||||||
|
echo "== acceptance criterion 15b: a passphrase key is also recognised =="
|
||||||
|
test_passphrase_key_is_redacted
|
||||||
|
echo "== acceptance criterion 16: a failing --set must not echo its value =="
|
||||||
|
test_failing_set_does_not_echo_its_value
|
||||||
|
echo "== extra: dry-run never installs, and redacts =="
|
||||||
|
test_dry_run_never_installs_and_redacts
|
||||||
|
echo "== extra: --check is read-only and always exits 0 =="
|
||||||
|
test_check_is_read_only_and_exits_zero
|
||||||
|
echo "== extra: the parse-failure refusal shape is also recognised =="
|
||||||
|
test_refusal_shape_from_parse_failure_wording_is_recognised
|
||||||
|
|
||||||
|
printf 'PASS: config-edit acceptance criteria\n'
|
||||||
@@ -433,6 +433,130 @@ test_assert_single_daemon_rejects_two_pids() {
|
|||||||
printf '%s' "$output" | grep -qF '4343' || fail "refusal message does not list the pids it found"
|
printf '%s' "$output" | grep -qF '4343' || fail "refusal message does not list the pids it found"
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# fleetd #593 instance 2 — `running_pid()` used to be a bare `pgrep -f "$PATTERN"`, which matches
|
||||||
|
# ANY process whose full command line contains the pattern TEXT, including a shell that merely
|
||||||
|
# embeds it as literal text rather than being the daemon. Measured live on this Mac: `pgrep -c`
|
||||||
|
# (a one-call count) does not exist on BSD at all, and `bash -c "<single command>"` execs in place
|
||||||
|
# so no parent shell survives to hold the pattern — which is exactly why the defect did not
|
||||||
|
# reproduce from a plain script and needs a wrapper shaped like this instead. A `sh -c '...; ...'`
|
||||||
|
# with MORE THAN ONE statement does not get that exec-in-place treatment: the shell forks a child
|
||||||
|
# for the second statement and stays alive itself, holding the whole `-c` string — pattern text
|
||||||
|
# included — in its own `ps -o args`, for as long as it runs. That is the same shape an
|
||||||
|
# `ssh host "…; …"` wrapper or a hand-typed pipeline leaves behind. Before the fix this test would
|
||||||
|
# have found the wrapper's pid in running_pid()'s output; it must not.
|
||||||
|
test_running_pid_excludes_self_matching_wrapper_shell() {
|
||||||
|
local before after wrapper_pid
|
||||||
|
before="$(running_pid)"
|
||||||
|
sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' &
|
||||||
|
wrapper_pid=$!
|
||||||
|
sleep 0.3
|
||||||
|
after="$(running_pid)"
|
||||||
|
kill "$wrapper_pid" 2>/dev/null || true
|
||||||
|
wait "$wrapper_pid" 2>/dev/null || true
|
||||||
|
[ "$after" = "$before" ] \
|
||||||
|
|| fail "running_pid() counted a self-matching wrapper shell (pid $wrapper_pid, holding the pattern as literal text in its own argv, not the daemon): before=[$before] after=[$after]"
|
||||||
|
}
|
||||||
|
|
||||||
|
# fleetd #593 CORRECTION 1 — the round-1 version of this test gave its standin an argv[0]
|
||||||
|
# containing the pattern text (via `exec -a`) and left `comm` as whatever that override produced,
|
||||||
|
# which was never `java`. That was fine for a denylist-of-shells filter, but the allowlist below
|
||||||
|
# now requires `comm = java` specifically, so the standin here must actually carry that comm, not
|
||||||
|
# just avoid being a shell. `exec -a java` overrides argv[0] to `java` while the process itself
|
||||||
|
# stays a genuine, harmless `sh`; combining it with the same non-exec'ing multi-statement shape
|
||||||
|
# the wrapper-shell test above uses keeps the pattern text in the process's own `ps -o args` for
|
||||||
|
# as long as it runs. Measured live on this Mac (BSD/macOS: `ps -o comm=` here reflects argv[0]):
|
||||||
|
# `comm=java`, `args` contains the pattern, `pgrep -f "$PATTERN"` finds it. Copying a real system
|
||||||
|
# binary into a scratch path and executing it from there was tried first, for a more literal
|
||||||
|
# stand-in daemon, and the OS killed it outright (SIGKILL, exit 137 — almost certainly a
|
||||||
|
# code-signing check on a relocated binary); `exec -a` needs no binary of its own and nothing
|
||||||
|
# under a scratch directory, and it is the technique CORRECTION 1 names as the right one.
|
||||||
|
#
|
||||||
|
# This is the one live-process test in this file whose result could differ on Linux: Linux sets
|
||||||
|
# `comm` from the actually-executed binary's own path, not from `exec -a`'s argv[0] override (BSD
|
||||||
|
# ties `comm` to argv[0], which is what makes this technique work here) — so on Linux this
|
||||||
|
# specific fixture might report `comm=sh`, not `comm=java`, even though the REAL daemon (a literal
|
||||||
|
# `java -jar target/fleetd.jar` process, never fabricated) is unaffected either way. I could not
|
||||||
|
# verify this fixture's behavior on Linux, so test_running_pid_counts_a_pid_whose_comm_is_java
|
||||||
|
# below backstops the same claim (the allowlist admits a pid whose comm is `java`) with a stubbed
|
||||||
|
# `ps`, which is identical bash on every platform and carries no such platform question.
|
||||||
|
test_running_pid_finds_a_real_java_named_second_process() {
|
||||||
|
local before after standin_pid
|
||||||
|
before="$(running_pid)"
|
||||||
|
( exec -a java sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' ) &
|
||||||
|
standin_pid=$!
|
||||||
|
sleep 0.3
|
||||||
|
after="$(running_pid)"
|
||||||
|
kill "$standin_pid" 2>/dev/null || true
|
||||||
|
wait "$standin_pid" 2>/dev/null || true
|
||||||
|
printf '%s\n' "$after" | grep -qxF "$standin_pid" \
|
||||||
|
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) whose own argv holds the pattern: before=[$before] after=[$after]"
|
||||||
|
}
|
||||||
|
|
||||||
|
# fleetd #593 CORRECTION 1, hole 2 — the round-1 filter denied known shell names (sh/bash/zsh/
|
||||||
|
# dash/ksh) and counted everything else. `ssh`, `perl`, `python3`, `ruby`, `tail` — anything not on
|
||||||
|
# that list, carrying the pattern in its own argv — was still counted right alongside the real
|
||||||
|
# daemon, and the ticket names `ssh` as a live route. Stubbing `pgrep`/`ps` (rather than spawning a
|
||||||
|
# real perl/ssh process) pins the exact discriminator this correction is about — comm, not the
|
||||||
|
# caller's shape — deterministically on every platform, with no dependency on perl/python3/ruby
|
||||||
|
# being installed in whatever environment runs this suite, and no dependency on how a given OS
|
||||||
|
# derives `comm` for a fabricated process (see the comment above
|
||||||
|
# test_running_pid_finds_a_real_java_named_second_process for why that matters here).
|
||||||
|
test_running_pid_drops_a_pid_whose_comm_is_not_java() {
|
||||||
|
pgrep() { printf '4242\n'; }
|
||||||
|
ps() { printf 'perl\n'; }
|
||||||
|
local found
|
||||||
|
found="$(running_pid)"
|
||||||
|
unset -f pgrep ps
|
||||||
|
[ -z "$found" ] \
|
||||||
|
|| fail "running_pid() counted pid 4242 whose comm is 'perl', not 'java' — denying known shell names does not exclude a non-shell wrapper such as ssh or perl (fleetd #593 CORRECTION 1): found=[$found]"
|
||||||
|
}
|
||||||
|
|
||||||
|
# fleetd #593 CORRECTION 1, hole 1 — pgrep can list a pid that exits before the following
|
||||||
|
# `ps -o comm=` lookup runs; on a gone pid `ps` prints nothing, so `comm` comes back empty. Under
|
||||||
|
# the round-1 denylist an empty string matched none of the denied shell names, so the dead pid was
|
||||||
|
# still counted — the exact false-positive shape the ticket exists to remove, just rarer. The
|
||||||
|
# allowlist fixes this for free: an empty comm is not `java` either.
|
||||||
|
test_running_pid_drops_a_pid_that_exited_before_the_comm_lookup() {
|
||||||
|
pgrep() { printf '4242\n'; }
|
||||||
|
ps() { :; } # a pid that no longer exists: the real `ps -p <gone>` prints nothing and this mirrors that
|
||||||
|
local found
|
||||||
|
found="$(running_pid)"
|
||||||
|
unset -f pgrep ps
|
||||||
|
[ -z "$found" ] \
|
||||||
|
|| fail "running_pid() counted pid 4242 whose comm lookup came back empty (the pid had already exited before the lookup ran) — an empty comm must not pass the allowlist (fleetd #593 CORRECTION 1): found=[$found]"
|
||||||
|
}
|
||||||
|
|
||||||
|
# The positive backstop for both stubbed tests above, and for
|
||||||
|
# test_running_pid_finds_a_real_java_named_second_process on whatever platform that live fixture
|
||||||
|
# does not itself carry comm=java: the allowlist must still ADMIT the one comm value the real
|
||||||
|
# daemon actually has. Measured on the real, currently-running daemon on this Mac: `comm=java`.
|
||||||
|
test_running_pid_counts_a_pid_whose_comm_is_java() {
|
||||||
|
pgrep() { printf '4242\n'; }
|
||||||
|
ps() { printf 'java\n'; }
|
||||||
|
local found
|
||||||
|
found="$(running_pid)"
|
||||||
|
unset -f pgrep ps
|
||||||
|
printf '%s\n' "$found" | grep -qxF '4242' \
|
||||||
|
|| fail "running_pid() did not count pid 4242 whose comm is 'java' — the daemon's own name must pass the allowlist: found=[$found]"
|
||||||
|
}
|
||||||
|
|
||||||
|
# fleetd #593 instance 3 — assert_single_daemon's refusal message used to tell the operator to
|
||||||
|
# "Investigate with 'pgrep -f \"\$PATTERN\"'", which — typed by hand or over ssh — is precisely the
|
||||||
|
# self-matching invocation instance 2 above fixes. A source-text check, the same technique
|
||||||
|
# test_no_error_lines_message_gated_by_drain_state uses: this is prose inside a die() call, never
|
||||||
|
# reached by sourcing (the SOURCED guard stops before the main flow, and this text only prints
|
||||||
|
# from inside a call assert_single_daemon makes when it is already refusing).
|
||||||
|
test_die_message_does_not_recommend_bare_pgrep_as_remediation() {
|
||||||
|
local src="$ROOT/scripts/redeploy-fleetd.sh" block bad
|
||||||
|
block="$(grep -A6 -F 'racing supervisor produces' "$src" || true)"
|
||||||
|
[ -n "$block" ] || fail "could not find the assert_single_daemon refusal message in redeploy-fleetd.sh"
|
||||||
|
bad="$(printf '%s' "$block" | grep -F "Investigate with 'pgrep -f" || true)"
|
||||||
|
[ -z "$bad" ] \
|
||||||
|
|| fail "assert_single_daemon's die message still hands the operator a bare 'pgrep -f \"\$PATTERN\"' as remediation (fleetd #593) — that is exactly the self-matching invocation"
|
||||||
|
printf '%s' "$block" | grep -qF 'fleetd #593' \
|
||||||
|
|| fail "assert_single_daemon's die message does not say in words that a pattern can match the caller (fleetd #593)"
|
||||||
|
}
|
||||||
|
|
||||||
# fleetd #511 — jar_id()'s no-argument default was unpinned by any test: nothing proved it reports
|
# fleetd #511 — jar_id()'s no-argument default was unpinned by any test: nothing proved it reports
|
||||||
# $JAR (the live path) rather than $JAR_STAGED. Both halves matter, so this pins both: the bare call
|
# $JAR (the live path) rather than $JAR_STAGED. Both halves matter, so this pins both: the bare call
|
||||||
# must hash the live jar, and an explicit path argument must hash THAT file, not fall back to $JAR.
|
# must hash the live jar, and an explicit path argument must hash THAT file, not fall back to $JAR.
|
||||||
@@ -646,6 +770,137 @@ test_wait_for_daemon_exit_times_out_if_pid_never_clears() {
|
|||||||
source "$ROOT/scripts/redeploy-fleetd.sh" # restore the real running_pid/sleep for later tests
|
source "$ROOT/scripts/redeploy-fleetd.sh" # restore the real running_pid/sleep for later tests
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# fleetd #603 — wait_for_new_pid is the dual of wait_for_daemon_exit above: it must not report
|
||||||
|
# success while running_pid() still answers empty, and must report success the moment a pid
|
||||||
|
# appears. Same counter-file idiom as test_wait_for_daemon_exit_returns_true_once_pid_clears above,
|
||||||
|
# for the same reason (running_pid() runs inside a `$(...)` subshell on every call).
|
||||||
|
test_wait_for_new_pid_returns_true_once_pid_appears() {
|
||||||
|
local counter_file="$TMP/wait-new-pid-calls" final_calls
|
||||||
|
printf '0' > "$counter_file"
|
||||||
|
running_pid() {
|
||||||
|
local n
|
||||||
|
n="$(cat "$counter_file")"
|
||||||
|
n=$((n + 1))
|
||||||
|
printf '%s' "$n" > "$counter_file"
|
||||||
|
if [ "$n" -lt 3 ]; then printf ''; else printf '4242'; fi
|
||||||
|
}
|
||||||
|
sleep() { :; }
|
||||||
|
wait_for_new_pid 10 || fail "wait_for_new_pid did not report success once the pid appeared"
|
||||||
|
final_calls="$(cat "$counter_file")"
|
||||||
|
[ "$final_calls" -ge 3 ] || fail "wait_for_new_pid returned before actually re-checking running_pid"
|
||||||
|
source "$ROOT/scripts/redeploy-fleetd.sh" # restore the real running_pid/sleep for later tests
|
||||||
|
}
|
||||||
|
|
||||||
|
test_wait_for_new_pid_times_out_if_pid_never_appears() {
|
||||||
|
local rc=0
|
||||||
|
running_pid() { printf ''; }
|
||||||
|
sleep() { :; }
|
||||||
|
wait_for_new_pid 3 || rc=$?
|
||||||
|
[ "$rc" -ne 0 ] || fail "wait_for_new_pid reported success while the pid never appeared"
|
||||||
|
source "$ROOT/scripts/redeploy-fleetd.sh" # restore the real running_pid/sleep for later tests
|
||||||
|
}
|
||||||
|
|
||||||
|
# fleetd #603 — await_daemon_started folds the pid-appeared check and the healthz check into one
|
||||||
|
# decision (see its own comment in redeploy-fleetd.sh for why a source-text grep cannot tell the two
|
||||||
|
# possible behaviors apart here). These two tests are the ticket's own acceptance criteria, run
|
||||||
|
# together in this one suite invocation so neither can be satisfied by code that never fails at all:
|
||||||
|
#
|
||||||
|
# 1. a slow start must still succeed — running_pid mimics a process that does not appear until
|
||||||
|
# well after the OLD, buggy 10-second budget, but does appear, and healthz answers.
|
||||||
|
# 2. a genuine failure must still fail, and still print the log tail — running_pid and the health
|
||||||
|
# check both report nothing at all, ever.
|
||||||
|
test_await_daemon_started_slow_pid_then_healthy_succeeds() {
|
||||||
|
source "$ROOT/scripts/redeploy-fleetd.sh"
|
||||||
|
stub_die_recorder
|
||||||
|
local counter_file="$TMP/await-slow-pid-calls" output
|
||||||
|
printf '0' > "$counter_file"
|
||||||
|
running_pid() {
|
||||||
|
local n
|
||||||
|
n="$(cat "$counter_file")"
|
||||||
|
n=$((n + 1))
|
||||||
|
printf '%s' "$n" > "$counter_file"
|
||||||
|
# Stays empty well past the old 10-second budget, then appears — the exact shape #603 reports.
|
||||||
|
if [ "$n" -lt 12 ]; then printf ''; else printf '4242'; fi
|
||||||
|
}
|
||||||
|
sleep() { :; }
|
||||||
|
poll_health_body() { printf '{"status":"ok"}'; return 0; }
|
||||||
|
# NOT `output="$(await_daemon_started ...)"`: that would run the call in a subshell, and
|
||||||
|
# NEW_PID — a plain global assignment inside the function, by design (see its own comment) — would
|
||||||
|
# die with that subshell instead of reaching this test's own shell. Redirect to a file instead, the
|
||||||
|
# same hazard await_daemon_started's own comment warns die() itself is subject to.
|
||||||
|
await_daemon_started 60 "" "http://ignored/healthz" "$TMP/await-slow-pid.out" \
|
||||||
|
> "$TMP/await-slow-pid-output.log" 2>&1
|
||||||
|
output="$(cat "$TMP/await-slow-pid-output.log")"
|
||||||
|
[ "$DIED_CALLED" = 0 ] \
|
||||||
|
|| fail "await_daemon_started must not die on a slow-but-real start: $DIED_MESSAGE"
|
||||||
|
assert_equals "4242" "$NEW_PID" "await_daemon_started NEW_PID after a slow-but-real start"
|
||||||
|
printf '%s' "$output" | grep -qF 'started, pid 4242' \
|
||||||
|
|| fail "await_daemon_started did not report the pid once it finally appeared"
|
||||||
|
source "$ROOT/scripts/redeploy-fleetd.sh"
|
||||||
|
}
|
||||||
|
|
||||||
|
test_await_daemon_started_never_appears_dies_with_log_tail() {
|
||||||
|
source "$ROOT/scripts/redeploy-fleetd.sh"
|
||||||
|
stub_die_recorder
|
||||||
|
local out_file="$TMP/await-never-appears.out"
|
||||||
|
printf 'boot line one\nboot line two\n' > "$out_file"
|
||||||
|
running_pid() { printf ''; }
|
||||||
|
sleep() { :; }
|
||||||
|
poll_health_body() { return 1; }
|
||||||
|
await_daemon_started 2 "" "http://127.0.0.1:1/healthz" "$out_file" > /dev/null 2>&1
|
||||||
|
[ "$DIED_CALLED" = 1 ] \
|
||||||
|
|| fail "await_daemon_started must die when the daemon never appears and never becomes healthy"
|
||||||
|
printf '%s' "$DIED_MESSAGE" | grep -qF 'never answered' \
|
||||||
|
|| fail "await_daemon_started die message does not say healthz never answered"
|
||||||
|
printf '%s' "$DIED_MESSAGE" | grep -qF 'boot line two' \
|
||||||
|
|| fail "await_daemon_started die message does not include the log tail"
|
||||||
|
source "$ROOT/scripts/redeploy-fleetd.sh"
|
||||||
|
}
|
||||||
|
|
||||||
|
# fleetd #603 review — the path the fall-through actually exists for, and the one gap a lead
|
||||||
|
# mutation found in the first version of this test file: running_pid() NEVER finds anything (as its
|
||||||
|
# own doc comment says it eventually will, once the daemon stops being launched as a plain
|
||||||
|
# `java -jar` its allowlist recognises), while /healthz answers anyway. Neither of the two tests
|
||||||
|
# above drives this: the slow-pid test has the pid appear, so the `else` branch never runs, and the
|
||||||
|
# never-appears test fails BOTH checks, so it dies either way and cannot tell which branch fired.
|
||||||
|
# This must not die, must warn (so the operator is told the pid could not be identified), and must
|
||||||
|
# leave NEW_PID empty — the honest "could not establish this" answer, never a guessed pid, which is
|
||||||
|
# what the final result line's `${NEW_PID:-unknown}` fallback exists to print truthfully.
|
||||||
|
#
|
||||||
|
# Proof this actually pins the behavior, not just the source text (paste from a real run, not
|
||||||
|
# claimed): reverting the `warn` below back to `die "no process appeared"` (the old fleetd #603
|
||||||
|
# defect, reintroduced) turns this one test red —
|
||||||
|
# FAIL: await_daemon_started must not die when the pid is never found but healthz answers
|
||||||
|
# — and restoring `warn` turns the whole suite green again. Both halves observed, not asserted.
|
||||||
|
test_await_daemon_started_pid_never_found_but_healthy_warns_and_survives() {
|
||||||
|
source "$ROOT/scripts/redeploy-fleetd.sh"
|
||||||
|
stub_die_recorder
|
||||||
|
local out_file="$TMP/await-pid-never-found.out" output
|
||||||
|
printf 'boot line\n' > "$out_file"
|
||||||
|
running_pid() { printf ''; }
|
||||||
|
sleep() { :; }
|
||||||
|
poll_health_body() { printf '{"status":"ok"}'; return 0; }
|
||||||
|
# NOT `output="$(await_daemon_started ...)"` — see the slow-pid test above for why that would
|
||||||
|
# drop NEW_PID's assignment in a subshell instead of reaching this test's own shell.
|
||||||
|
await_daemon_started 3 "" "http://ignored/healthz" "$out_file" \
|
||||||
|
> "$TMP/await-pid-never-found-output.log" 2>&1
|
||||||
|
output="$(cat "$TMP/await-pid-never-found-output.log")"
|
||||||
|
[ "$DIED_CALLED" = 0 ] \
|
||||||
|
|| fail "await_daemon_started must not die when the pid is never found but healthz answers: $DIED_MESSAGE"
|
||||||
|
printf '%s' "$output" | grep -qF 'falling through to the health check' \
|
||||||
|
|| fail "await_daemon_started did not warn that the pid could not be identified"
|
||||||
|
assert_equals "" "$NEW_PID" \
|
||||||
|
"await_daemon_started NEW_PID when the pid is never found but healthz answers — must stay empty, never a guessed pid"
|
||||||
|
source "$ROOT/scripts/redeploy-fleetd.sh"
|
||||||
|
}
|
||||||
|
|
||||||
|
test_await_daemon_started_call_site_present() {
|
||||||
|
local src="$ROOT/scripts/redeploy-fleetd.sh" call_line
|
||||||
|
call_line="$(grep -Fn 'await_daemon_started "$HEALTH_WAIT" "$OLD_PID" "$HEALTH" "$OUT"' "$src" | head -1 | cut -d: -f1 || true)"
|
||||||
|
[ -n "$call_line" ] \
|
||||||
|
|| fail "could not find the main flow's await_daemon_started call site in redeploy-fleetd.sh"
|
||||||
|
}
|
||||||
|
|
||||||
# fleetd #521 — the swap step's guard, at two levels.
|
# fleetd #521 — the swap step's guard, at two levels.
|
||||||
#
|
#
|
||||||
# The first two tests call the predicate should_swap() directly. They pin its logic, and that is all
|
# The first two tests call the predicate should_swap() directly. They pin its logic, and that is all
|
||||||
@@ -1226,9 +1481,11 @@ test_poll_health_body_returns_nonzero_when_unreachable() {
|
|||||||
|
|
||||||
test_report_health_call_site_present() {
|
test_report_health_call_site_present() {
|
||||||
local src="$ROOT/scripts/redeploy-fleetd.sh" call_line
|
local src="$ROOT/scripts/redeploy-fleetd.sh" call_line
|
||||||
call_line="$(grep -Fn 'report_health "$HEALTH_BODY" "$HEALTH_CODE" "$OUT" "$HEALTH_WAIT"' "$src" | head -1 | cut -d: -f1 || true)"
|
# fleetd #603 — moved from a literal main-flow call into await_daemon_started (see its own
|
||||||
|
# comment above for why); this now finds the call inside that function instead.
|
||||||
|
call_line="$(grep -Fn 'report_health "$HEALTH_BODY" "$HEALTH_CODE" "$out_file" "$health_wait"' "$src" | head -1 | cut -d: -f1 || true)"
|
||||||
[ -n "$call_line" ] \
|
[ -n "$call_line" ] \
|
||||||
|| fail "could not find the main flow's report_health call site in redeploy-fleetd.sh"
|
|| fail "could not find await_daemon_started's report_health call site in redeploy-fleetd.sh"
|
||||||
}
|
}
|
||||||
|
|
||||||
# fleetd #555 item 8 — `HAD_OLD_PID=0; [ -n "$OLD_PID" ] && HAD_OLD_PID=1`. Measured safe under
|
# fleetd #555 item 8 — `HAD_OLD_PID=0; [ -n "$OLD_PID" ] && HAD_OLD_PID=1`. Measured safe under
|
||||||
@@ -1894,6 +2151,12 @@ test_require_drivable_supervisor_accepts_known_kinds
|
|||||||
test_count_daemon_pids
|
test_count_daemon_pids
|
||||||
test_assert_single_daemon_accepts_one_pid
|
test_assert_single_daemon_accepts_one_pid
|
||||||
test_assert_single_daemon_rejects_two_pids
|
test_assert_single_daemon_rejects_two_pids
|
||||||
|
test_running_pid_excludes_self_matching_wrapper_shell
|
||||||
|
test_running_pid_finds_a_real_java_named_second_process
|
||||||
|
test_running_pid_drops_a_pid_whose_comm_is_not_java
|
||||||
|
test_running_pid_drops_a_pid_that_exited_before_the_comm_lookup
|
||||||
|
test_running_pid_counts_a_pid_whose_comm_is_java
|
||||||
|
test_die_message_does_not_recommend_bare_pgrep_as_remediation
|
||||||
test_jar_id_defaults_to_live_and_reports_explicit_path
|
test_jar_id_defaults_to_live_and_reports_explicit_path
|
||||||
test_hash256_computes_a_real_sha256
|
test_hash256_computes_a_real_sha256
|
||||||
test_jar_id_reports_absent_for_missing_file
|
test_jar_id_reports_absent_for_missing_file
|
||||||
@@ -1911,6 +2174,12 @@ test_require_no_build_jar_dies_when_absent
|
|||||||
test_require_no_build_jar_accepts_present_jar
|
test_require_no_build_jar_accepts_present_jar
|
||||||
test_wait_for_daemon_exit_returns_true_once_pid_clears
|
test_wait_for_daemon_exit_returns_true_once_pid_clears
|
||||||
test_wait_for_daemon_exit_times_out_if_pid_never_clears
|
test_wait_for_daemon_exit_times_out_if_pid_never_clears
|
||||||
|
test_wait_for_new_pid_returns_true_once_pid_appears
|
||||||
|
test_wait_for_new_pid_times_out_if_pid_never_appears
|
||||||
|
test_await_daemon_started_slow_pid_then_healthy_succeeds
|
||||||
|
test_await_daemon_started_never_appears_dies_with_log_tail
|
||||||
|
test_await_daemon_started_pid_never_found_but_healthy_warns_and_survives
|
||||||
|
test_await_daemon_started_call_site_present
|
||||||
test_swap_ordered_after_wait_and_before_start
|
test_swap_ordered_after_wait_and_before_start
|
||||||
test_drain_gate_abort_message_says_no_no_build
|
test_drain_gate_abort_message_says_no_no_build
|
||||||
test_drain_gate_refusal_build_ran_staged_present
|
test_drain_gate_refusal_build_ran_staged_present
|
||||||
|
|||||||
Reference in New Issue
Block a user