Compare commits
51 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| eb568ff451 | |||
| ac790e4cce | |||
| 6b5f3f472f | |||
| 38dec72152 | |||
| 0e8bfb74fc | |||
| 51f7b0a3ca | |||
| 21c539f22e | |||
| bd2774b5f1 | |||
| c50f5b2d61 | |||
| 01a840cc14 | |||
| c4deef08be | |||
| c796eac09c | |||
| e897e5257b | |||
| 2afa3652bb | |||
| 80092ff359 | |||
| 9d37f3aa29 | |||
| eaf89abaf6 | |||
| d895f02bc1 | |||
| 43206cac2f | |||
| 8bba3a8184 | |||
| 5cf3ca9a89 | |||
| ac474981e4 | |||
| ba04b2359b | |||
| 2e5b63f6f6 | |||
| 3437d6313d | |||
| 743377d6cd | |||
| 5952d559c7 | |||
| 205ad823b0 | |||
| d654ccb818 | |||
| a89dcc9b7e | |||
| 0d7b4fb026 | |||
| 321d8dcbb5 | |||
| ef8c97871e | |||
| 26bafe824b | |||
| 838a701109 | |||
| e5eb3534c7 | |||
| 959c83534f | |||
| 31b028e860 | |||
| 7840e9adf6 | |||
| c3672f5472 | |||
| bbf68f3e3c | |||
| 776743cbe2 | |||
| c935b181dd | |||
| ee932fd85b | |||
| fe2e5ede34 | |||
| c325054242 | |||
| 7662e2d0c8 | |||
| 4877992a70 | |||
| 1c051c4e47 | |||
| adb7a67880 | |||
| 723fe494e9 |
@@ -188,6 +188,18 @@ must obey belongs in the charter, not here.
|
||||
- **This repo is the bridge.** The daemon is `fleetd`, its MCP mount is `http://127.0.0.1:8765/mcp`,
|
||||
and the code behind the rules above is `mcp/FleetMcp` (tools), `auth/Authz` (the role table),
|
||||
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
|
||||
- **`fleet_profiles`/`fleet_list` report two separate outage states, and they are not the same
|
||||
thing.** *Quarantined* (CB-578) means the backend told us it is out of capacity — a long,
|
||||
1800s-default cooldown. *Cooling off* (fleetd #201/#227) means a profile's credential threw two
|
||||
distinct backend errors (a non-exhaustion failure such as an HTTP 5xx) within 60 seconds — a
|
||||
short, fixed 60s cooldown, not configurable per profile. Each check runs independently, so a
|
||||
profile can show both at once. In the JSON: a cooling profile carries `credentialId` and
|
||||
`coolingOffForSeconds`; a quarantined profile carries `quarantinedForSeconds`; a profile hit by
|
||||
both carries all three fields, and either state alone already sets that profile's `free` to `0`.
|
||||
A `fleet_spawn` naming a cooling-off profile is refused before it ever reaches the backend
|
||||
adapter, with a message naming the credential and the remaining seconds ("cooling off after
|
||||
repeated backend errors") — distinct wording from a quarantine refusal, so don't conflate the
|
||||
two when reading a spawn failure.
|
||||
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
|
||||
`reviewer` (scoped review → one structured finding). Name one in every delegation.
|
||||
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
|
||||
@@ -195,11 +207,22 @@ must obey belongs in the charter, not here.
|
||||
`fleets-status` (report every fleet that shares one LavinMQ instance).
|
||||
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
|
||||
(a submodule with its own remote).
|
||||
- **Flows and the error model** — rendezvous, `fleet_ask`, detached delivery, turn-done fallback —
|
||||
are diagrammed in `docs/MCP-Contract.md` **§6 only**. The rest of that page is a pre-build design
|
||||
doc whose tool names, parameter names and REST paths never caught up with the code, so do not use
|
||||
it as the tool reference (CB-609). Section 6 is kept out of this file because this file loads into
|
||||
every session's context.
|
||||
- **A provisioned worktree neutralizes `.mcp.json`, `opencode.json` and `.autoenv`** — the repo's
|
||||
committed copies would otherwise mount the primary's IDE and forge servers (fleetd #134). The
|
||||
worktree's copy of each is a stub, **not** the repo's real file, so a worker that reads one and
|
||||
reports what it found is reporting on the stub. The daemon logs a per-spawn summary, but the
|
||||
worker cannot see that log. From inside its own worktree a worker — or a lead debugging one —
|
||||
reads the list with `git config --worktree --get-all fleet.neutralizedConfig`, and the
|
||||
consequence with `git config --worktree --get fleet.neutralizedConfigNote`. Never brief a worker
|
||||
to edit one of these files: the edit cannot be committed, and it will not tell you so.
|
||||
- **Flows and the error model** — rendezvous, `fleet_ask`, detached delivery, the turn-done
|
||||
fallback and status gating — are diagrammed in `docs/MCP-Contract.md`. That page is now flows
|
||||
only: its pre-build tool catalogue, parameter tables and REST paths were deleted rather than
|
||||
corrected, because a hand-maintained second copy of the tool surface is what drifted for a month
|
||||
while this line pointed every session at it (CB-609 / #114). **The live MCP schema is the tool
|
||||
reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if
|
||||
that page names a `fleet_*` tool the server does not register. The flows are kept out of this
|
||||
file because this file loads into every session's context.
|
||||
|
||||
### Redeploying the daemon — the lead may do this (primary only)
|
||||
|
||||
|
||||
@@ -0,0 +1,531 @@
|
||||
# CB-201 and CB-227 refinement
|
||||
|
||||
Date: 2026-09-03
|
||||
|
||||
## Decision
|
||||
|
||||
#201 and #227 are one delivery program, but they are not one implementation unit.
|
||||
|
||||
#201 has a real seam: `CompletionResolver` can publish a typed backend-error event only after its
|
||||
waiter resolution wins. #227 can consume that event without knowing any pane text. The classifier
|
||||
must land before the final #227 wiring. However, the policy engine, roster state, and lead nudge can
|
||||
be built in parallel with the classifier.
|
||||
|
||||
I propose five units. Units 1 to 4 own separate files and can run in parallel. Unit 5 owns all
|
||||
composition files and lands after them. It also depends on the #234 defect 2 fix named in the task.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
U1["Unit 1: typed backend-error classification"] --> U5["Unit 5: wire policy, spawn gate, and fleet views"]
|
||||
U2["Unit 2: credential outage policy"] --> U5
|
||||
U3["Unit 3: lead outage nudge"] --> U5
|
||||
U4["Unit 4: durable member outcome"] --> U5
|
||||
D234["#234 defect 2: fail-loud target resolution"] --> U5
|
||||
```
|
||||
|
||||
*Figure 1. Four file-disjoint foundations feed one composition unit.*
|
||||
|
||||
This split keeps `Fleetd.java` under one owner. It also keeps every other production file under one
|
||||
unit in this plan.
|
||||
|
||||
## Evidence checked in the current branch
|
||||
|
||||
I read both issue pages in full. Each page reports zero comments.
|
||||
|
||||
| Evidence | What the code says now |
|
||||
|---|---|
|
||||
| `inject/CompletionResolver.java:229-237` | A turn below two seconds fails before normal scrape classification. A matching fast backend error is therefore only a generic failure today. |
|
||||
| `inject/CompletionResolver.java:260-276` and `:332-367` | #211 already added raw-screen classification when `lastAssistantBlock` is empty. The “dead-code question” in #201 is stale on this branch. |
|
||||
| `inject/CompletionResolver.java:288-317` | Exhaustion wins before the hard-coded `API Error:` match. A backend error then goes through generic `fail(...)`. |
|
||||
| `inject/CompletionResolver.java:311-316` | The code admits that the pattern is a heuristic. A member report which quotes an API error may match it. |
|
||||
| `inject/CompletionResolver.java:449-467` | Startup coverage exists only for `exhaustedPattern`. |
|
||||
| `inject/ExhaustedPatternLookup.java:13-25` | The current lookup and explicit `none()` value are a good shape for the new classifier seam. |
|
||||
| `Fleetd.java:196-207` | One `BackendQuarantine` is shared by placement and the exhaustion sink. Its cooldown comes from `quarantineCooldownSeconds`. |
|
||||
| `Fleetd.java:322-363` | Pattern compilation, target-to-profile lookup, and the live `ExhaustionSink` are composed in `Fleetd.main`. The sink on this branch still ends in `.ifPresent(...)`. This plan assumes #234 replaces that silent path. |
|
||||
| `placement/BackendQuarantine.java:60-87` | A repeated exhaustion restarts one long quarantine. The store is credential-keyed and uses an injected monotonic clock. |
|
||||
| `member/CompositePeerLauncher.java:260-317` | Explicit and policy-selected spawns have separate gates. Both paths must learn about outage cool-off. |
|
||||
| `member/CompositePeerLauncher.java:347-379` | Exhaustion refusal already checks a credential for explicit spawns and filters policy candidates. Its error text says “exhausted”. |
|
||||
| `placement/PlacementContext.java:10-22` and `PlacementPolicyUtil.java:14-83` | Automatic placement has only one transient exclusion set named `quarantined`. Reusing it would make outage errors say “backend exhausted”. |
|
||||
| `mcp/FleetMcp.java:913-1025` | `fleet_list` sets `free: 0` and adds `credentialId` plus `quarantinedForSeconds` when quarantine is active. |
|
||||
| `session/MemberSession.java:51-59` | The roster has `DONE` and generic `FAILED`, but no backend-error state or stored reason. |
|
||||
| `session/SessionManager.java:695-773` | A normal boundary moves `BUSY` to `DONE`. A failure moves any non-released session to `FAILED`. The async completion resolver can race the `DONE` update. |
|
||||
| `session/SessionManager.java:648-687` | `rosterView` reports the session state, but it reports no terminal reason. |
|
||||
| `msg/MessageService.java:922-940` | CB-588 already nudges for every terminal async ticket, including failures. Current code would report failed tickets, but it would not report one correlated outage. |
|
||||
| `msg/ReplyPushLoop.java:20-48` | Replies, terminal tickets, and questions share one per-lead schedule. This prevents two push sources from injecting competing turns. |
|
||||
| `msg/ReplyPushLoop.java:305-395` | Each push entry point resolves the owning lead through `PrimaryRegistry`. Missing ownership is logged and the durable or pending item remains the backstop. |
|
||||
| `msg/ReplyPushLoop.java:496-547` | One tick builds one combined nudge. Pending items have separate reminder counts. |
|
||||
| `health/FleetHealthMonitor.java:91-143` | Health is a slow periodic observer of members and message-layer facts. It does not receive completion classifications. |
|
||||
| `health/FleetHealthMonitor.java:206-208` | `healthCoverage` means health enabled plus webhook configured. It does not describe lead-pane alerts. |
|
||||
| `Fleetd.java:465-486` | Health stays `detection-only` without the webhook notification setting. |
|
||||
|
||||
I also read the related unit tests for `CompletionResolver`, `ReplyPushLoop`, `BackendQuarantine`,
|
||||
`CompositePeerLauncher`, `PlacementPolicyUtil`, `SessionManager`, `MessageService`, and `FleetMcp`.
|
||||
|
||||
I did not inspect the in-progress #234 branch. I only used the two measured facts in the task. No
|
||||
peer architect was named, so I did not exchange a design with one.
|
||||
|
||||
## Required behaviour
|
||||
|
||||
The policy should use these first values:
|
||||
|
||||
- Threshold: **2** classified backend errors.
|
||||
- Window: **60 seconds**, measured from the first error to the second.
|
||||
- Cool-off: **60 seconds**, starting when the threshold is reached.
|
||||
- Correlation key: `credentialId`, never profile name and never error text.
|
||||
- Incident rule: one active incident per credential. Errors during its cool-off do not extend it and
|
||||
do not create more lead notices.
|
||||
- Rearm rule: after cool-off ends, two fresh errors are needed for another incident.
|
||||
|
||||
Two errors are the smallest threshold which protects the honest one-turn failure. A 60-second window
|
||||
fits the measured two-member outage. A 60-second cool-off blocks immediate repeat spawns without
|
||||
turning a short backend fault into the default 1,800-second exhaustion quarantine.
|
||||
|
||||
A single classified error still fails its send and marks its member `backend_error`. It does not
|
||||
cool a credential and does not send an outage notice. This is what “a single error changes nothing”
|
||||
must mean at the credential level. It cannot mean that the failed member still looks successful.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant R1 as Resolver for member A
|
||||
participant R2 as Resolver for member B
|
||||
participant P as Outage policy
|
||||
participant S as Spawn gate
|
||||
participant N as Lead push loop
|
||||
participant L as Lead pane
|
||||
|
||||
R1->>P: backend error for credential C
|
||||
Note over P: Count 1, no cool-off
|
||||
R2->>P: backend error for credential C within 60s
|
||||
P->>P: Start one 60s incident
|
||||
P->>S: Credential C is cooling off
|
||||
P->>N: Queue one incident notice
|
||||
N->>L: Inject when lead is idle, blocked, or done
|
||||
L->>S: Request another spawn on credential C
|
||||
S-->>L: Refuse and report remaining cool-off
|
||||
```
|
||||
|
||||
*Figure 2. The second independent classification creates the fleet-level event.*
|
||||
|
||||
Against the 2026-09-01 case, the second failed member would start cool-off. `fleet_list` would show
|
||||
zero free capacity and both members as `backend_error`. The push loop would inject one outage notice
|
||||
even if the lead had not polled either ticket yet. The design reports the outage. It does not recover
|
||||
uncommitted work from the members.
|
||||
|
||||
## Unit 1 — Typed backend-error classification
|
||||
|
||||
### Scope
|
||||
|
||||
Replace the direct hard-coded check inside `CompletionResolver` with a lookup and a sink. Keep the
|
||||
public send result as a failed send. The typed internal event is the seam #227 consumes.
|
||||
|
||||
The lookup returns the pattern for a target. The sink receives the target, matched line, and full
|
||||
failure reason. It fires only after `Rendezvous.resolveFailure(...)` wins for that exact captured
|
||||
waiter. This copies the race rule already used by `ExhaustionSink`.
|
||||
|
||||
The classifier must run in all three current paths:
|
||||
|
||||
1. a normal non-empty assistant block;
|
||||
2. the #211 raw scrape fallback;
|
||||
3. a turn inside `MIN_TURN_NANOS`, before it becomes a generic too-fast failure.
|
||||
|
||||
In every path, the order stays: stale-baseline guard, exhaustion, backend error, then generic
|
||||
failure or completion. A fast turn still fails when no configured pattern matches.
|
||||
|
||||
Keep `(?i)\bAPI Error\s*:` as a compatibility pattern for profiles without `errorPattern` until the
|
||||
operator config is updated. Do not call this full coverage. Startup reporting in Unit 5 must name
|
||||
profiles using this weaker legacy default.
|
||||
|
||||
### Files owned
|
||||
|
||||
- Add `fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorPatternLookup.java`.
|
||||
- Add `fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorSink.java`.
|
||||
- Change `fleetd/src/main/java/dev/ltms/fleet/inject/CompletionResolver.java`.
|
||||
- Change `fleetd/src/test/java/dev/ltms/fleet/inject/CompletionResolverTest.java`.
|
||||
|
||||
No other unit may edit these files.
|
||||
|
||||
### Acceptance criteria
|
||||
|
||||
1. A target-specific error pattern matches a normal assistant block and resolves the send as failed.
|
||||
2. The same match calls `BackendErrorSink` exactly once after the waiter resolution wins.
|
||||
3. A late classification which loses to `fleet_reply` does not call the sink.
|
||||
4. An exhausted line that also matches the generic error pattern stays `BACKEND_EXHAUSTED`. It calls
|
||||
only `ExhaustionSink`.
|
||||
5. A raw pane with leading Terminal User Interface (TUI) chrome and no assistant marker still uses
|
||||
the #211 fallback and calls the backend-error sink.
|
||||
6. A matching error inside the two-second floor is typed and sent to the sink. A non-matching fast
|
||||
turn stays a generic failure.
|
||||
7. An unchanged delivery baseline which contains old backend-error text is suppressed. It never
|
||||
increments outage evidence.
|
||||
8. A non-match keeps the existing completion result and text.
|
||||
9. Constructors used by current callers keep compiling. They use the legacy default lookup and an
|
||||
explicit inert sink until Unit 5 supplies the production objects.
|
||||
10. Unit tests pass. The developer runs the focused test first, then `mvn clean install` from
|
||||
`fleetd/`.
|
||||
|
||||
### Dependencies
|
||||
|
||||
None. Unit 1 can run with Units 2, 3, and 4.
|
||||
|
||||
Unit 5 depends on its new lookup, sink, and constructor.
|
||||
|
||||
### What to report back
|
||||
|
||||
- The exact classifier order in all three paths.
|
||||
- The focused test command and result.
|
||||
- The test which proves a losing waiter race does not publish an event.
|
||||
- The test which proves a fast matching failure is typed.
|
||||
- The final `mvn clean install` result.
|
||||
- Any constructor kept only for transition and where Unit 5 replaces it.
|
||||
|
||||
## Unit 2 — Credential outage policy
|
||||
|
||||
### Scope
|
||||
|
||||
Build a small credential-keyed state machine. It accepts already-classified backend-error events.
|
||||
It does not read pane text, profiles, sessions, or lead state.
|
||||
|
||||
Use an injected monotonic clock. A call records `credentialId`, target, and reason. It returns a new
|
||||
incident only on the threshold crossing. The incident contains a stable event id, credential id,
|
||||
the distinct affected targets, evidence count, window, and remaining cool-off.
|
||||
|
||||
This class owns both correlation and short cool-off. Keeping them together makes threshold crossing
|
||||
and the cool-off deadline one atomic state change.
|
||||
|
||||
### Files owned
|
||||
|
||||
- Add `fleetd/src/main/java/dev/ltms/fleet/placement/BackendOutagePolicy.java`.
|
||||
- Add `fleetd/src/test/java/dev/ltms/fleet/placement/BackendOutagePolicyTest.java`.
|
||||
|
||||
No other unit may edit these files.
|
||||
|
||||
### Acceptance criteria
|
||||
|
||||
1. One error creates no incident and no cool-off.
|
||||
2. Two errors for one credential within 60 seconds create exactly one incident and a 60-second
|
||||
cool-off.
|
||||
3. Two errors more than 60 seconds apart do not create an incident.
|
||||
4. The exact 60-second boundary has a pinned result. Use inclusive `<= 60s` so scheduler delay does
|
||||
not discard evidence at the boundary.
|
||||
5. Different credentials never share evidence.
|
||||
6. Different profiles which supply the same credential id do share evidence. The policy itself only
|
||||
sees the credential id.
|
||||
7. More errors during active cool-off do not extend its deadline and do not return another incident.
|
||||
8. After expiry, old evidence is cleared. Two fresh errors are needed to create the next incident.
|
||||
9. Remaining seconds round up, matching `BackendQuarantine` reporting.
|
||||
10. Concurrent second and third errors cannot return two incidents.
|
||||
11. The focused tests and `mvn clean install` pass.
|
||||
|
||||
### Dependencies
|
||||
|
||||
None. Unit 2 can run with Units 1, 3, and 4.
|
||||
|
||||
Unit 5 depends on the policy API.
|
||||
|
||||
### What to report back
|
||||
|
||||
- The state transition table and locking method.
|
||||
- The exact threshold, window, cool-off, and boundary rule.
|
||||
- The test which proves one incident under concurrent calls.
|
||||
- The focused test and `mvn clean install` results.
|
||||
|
||||
## Unit 3 — Lead outage nudge
|
||||
|
||||
### Scope
|
||||
|
||||
Add backend incidents as a fourth pending source in `ReplyPushLoop`. Do not create another scheduler
|
||||
or call `AgentControl.send` from `Fleetd`. The existing combined per-lead schedule is the control
|
||||
which prevents competing injected turns.
|
||||
|
||||
The entry point takes an incident id, affected worker targets, credential id, affected profile
|
||||
names, and remaining cool-off. It resolves distinct owning leads through `PrimaryRegistry`.
|
||||
|
||||
Each `(incidentId, lead)` item is one-shot. It waits while the lead is not injectable. After one
|
||||
successful `agents.send`, remove it. A send exception keeps it pending for a bounded retry. It never
|
||||
uses the repeated reminder behaviour of an uncollected ticket.
|
||||
|
||||
Also add a fail-loud entry point for a classified target that Unit 5 cannot map to a credential. It
|
||||
uses `PrimaryRegistry.nudgeTargetFor(target)` and says that correlation could not run. If no lead is
|
||||
known, log at `WARN`, not `DEBUG`.
|
||||
|
||||
### Files owned
|
||||
|
||||
- Change `fleetd/src/main/java/dev/ltms/fleet/msg/ReplyPushLoop.java`.
|
||||
- Change `fleetd/src/test/java/dev/ltms/fleet/msg/ReplyPushLoopTest.java`.
|
||||
|
||||
No other unit may edit these files.
|
||||
|
||||
### Acceptance criteria
|
||||
|
||||
1. One incident affecting two workers owned by one lead causes one successful pane injection.
|
||||
2. Two affected workers owned by two leads cause one successful injection per affected lead. This
|
||||
is one notice per event per lead, not one notice per member.
|
||||
3. Repeating the same incident id is idempotent.
|
||||
4. A busy or unknown lead is not injected. The item stays pending until the lead becomes injectable
|
||||
or its attempt cap is reached.
|
||||
5. After one successful injection, later ticks do not mention that incident again.
|
||||
6. A failed `agents.send` is retried within the existing bound. A successful retry still gives only
|
||||
one successful send.
|
||||
7. A pending ticket and an outage incident for one lead appear in one combined nudge, not two
|
||||
competing turns.
|
||||
8. The text names the credential, profiles, affected workers, and remaining cool-off. It tells the
|
||||
lead to run `fleet_list`.
|
||||
9. An unmapped target produces a direct warning notice when a lead is known. If no lead is known,
|
||||
the code logs a `WARN` naming the target and reason.
|
||||
10. `stop()` clears incident state as it clears other push state.
|
||||
11. Existing reply, ticket, and question tests stay green. The focused tests and
|
||||
`mvn clean install` pass.
|
||||
|
||||
### Dependencies
|
||||
|
||||
None. The API uses plain values, not the Unit 2 incident class. This lets Unit 3 run in parallel.
|
||||
|
||||
Unit 5 adapts the Unit 2 incident into this entry point.
|
||||
|
||||
### What to report back
|
||||
|
||||
- The exact one-shot and retry rules.
|
||||
- The test showing one combined nudge with a failed ticket.
|
||||
- The test showing one successful send for two affected workers.
|
||||
- The focused test and `mvn clean install` results.
|
||||
|
||||
## Unit 4 — Durable member backend-error outcome
|
||||
|
||||
### Scope
|
||||
|
||||
Make a classified backend failure remain visible after its ticket is collected or expires.
|
||||
|
||||
Add `BACKEND_ERROR` to `MemberSession.State`. Add a nullable failure detail to `MemberSession` and
|
||||
render it as `failureReason` in `SessionManager.rosterView`. Add
|
||||
`SessionManager.onBackendError(target, reason)`.
|
||||
|
||||
The transition must handle both completion orderings:
|
||||
|
||||
- `BUSY -> BACKEND_ERROR` when classification wins before the normal completion state update;
|
||||
- `DONE -> BACKEND_ERROR` when the async resolver runs after `SessionManager.onTurnComplete`.
|
||||
|
||||
It must use a compare-and-set retry or another atomic update. `RELEASED` must never return to the
|
||||
roster. A backend-error member is terminal and cannot accept another delivery.
|
||||
|
||||
### Files owned
|
||||
|
||||
- Change `fleetd/src/main/java/dev/ltms/fleet/session/MemberSession.java`.
|
||||
- Change `fleetd/src/main/java/dev/ltms/fleet/session/SessionManager.java`.
|
||||
- Change `fleetd/src/test/java/dev/ltms/fleet/session/SessionManagerTest.java`.
|
||||
|
||||
No other unit may edit these files.
|
||||
|
||||
### Acceptance criteria
|
||||
|
||||
1. `onBackendError` moves a `BUSY` member to `BACKEND_ERROR` and stores the reason.
|
||||
2. It also moves `DONE` to `BACKEND_ERROR`, covering the resolver race.
|
||||
3. A later normal `onTurnComplete` cannot change `BACKEND_ERROR` back to `DONE`.
|
||||
4. A released or unknown member is not recreated. The unknown case logs at `WARN` and returns an
|
||||
explicit false result to its caller.
|
||||
5. `onDelivered` refuses a `BACKEND_ERROR` member, just as it refuses generic `FAILED`.
|
||||
6. `rosterView` reports `state: backend_error` and `failureReason` after the send ticket is gone.
|
||||
7. Ordinary members do not gain a blank or invented `failureReason` field.
|
||||
8. Existing constructors keep source compatibility for tests and adapters.
|
||||
9. The focused tests and `mvn clean install` pass.
|
||||
|
||||
### Dependencies
|
||||
|
||||
None. Unit 4 can run with Units 1, 2, and 3.
|
||||
|
||||
Unit 5 calls the new session method from the production sink.
|
||||
|
||||
### What to report back
|
||||
|
||||
- The two race orderings and the tests for both.
|
||||
- The exact roster JSON shape.
|
||||
- The unknown-target result and log level.
|
||||
- The focused test and `mvn clean install` results.
|
||||
|
||||
## Unit 5 — Production wiring, spawn gate, and fleet views
|
||||
|
||||
### Scope
|
||||
|
||||
Compose Units 1 to 4 in production. This is the only unit which edits `Fleetd.java`.
|
||||
|
||||
Add per-profile `errorPattern` config beside `exhaustedPattern`. Compile both once at startup. A
|
||||
configured pattern wins over the legacy default. Report configured profiles and legacy-default
|
||||
profiles separately at startup. A bad regex must stop startup with the profile and key in the
|
||||
message.
|
||||
|
||||
Wire one production `BackendErrorSink` with this order:
|
||||
|
||||
1. mark the member `backend_error` with its reason;
|
||||
2. resolve the profile and its current `effectiveCredentialId()` through the fail-loud #234 seam;
|
||||
3. record the error in `BackendOutagePolicy`;
|
||||
4. on a new incident, submit one event to `ReplyPushLoop`.
|
||||
|
||||
If target metadata cannot be resolved, do not end in `Optional.ifPresent`. Log an error and call the
|
||||
Unit 3 unmapped-target notice. The failed send still reaches its ticket through CB-588.
|
||||
|
||||
Teach both spawn paths about a separate cool-off source. Exhaustion quarantine has priority when
|
||||
both states are active. Automatic placement needs a distinct `coolingOff` set so its refusal does
|
||||
not say “exhausted”.
|
||||
|
||||
Extend the MCP (Model Context Protocol) views:
|
||||
|
||||
- A cooling profile has `free: 0`, `credentialId`, and `coolingOffForSeconds` in `fleet_list`.
|
||||
- It does not have `quarantinedForSeconds` unless exhaustion quarantine is also active.
|
||||
- `fleet_profiles` has a separate `coolingOff` map, not an entry in `quarantined`.
|
||||
- A direct spawn refusal says the credential is cooling off after repeated backend errors and gives
|
||||
the remaining seconds.
|
||||
|
||||
Do not change `FleetHealthMonitor.coverage`. It still describes the periodic health webhook path.
|
||||
Lead-pane outage delivery is a separate capability.
|
||||
|
||||
### Files owned
|
||||
|
||||
- `fleetd/src/main/java/dev/ltms/fleet/Fleetd.java`
|
||||
- `fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java`
|
||||
- `fleetd/src/main/java/dev/ltms/fleet/config/ConfigRef.java`
|
||||
- `fleetd/src/main/java/dev/ltms/fleet/member/CompositePeerLauncher.java`
|
||||
- `fleetd/src/main/java/dev/ltms/fleet/placement/PlacementContext.java`
|
||||
- `fleetd/src/main/java/dev/ltms/fleet/placement/PlacementPolicyUtil.java`
|
||||
- `fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java`
|
||||
- `fleetd/src/test/java/dev/ltms/fleet/config/FleetConfigTest.java`
|
||||
- `fleetd/src/test/java/dev/ltms/fleet/config/ConfigRefTest.java`
|
||||
- `fleetd/src/test/java/dev/ltms/fleet/member/CompositePeerLauncherTest.java`
|
||||
- `fleetd/src/test/java/dev/ltms/fleet/placement/PlacementPolicyTest.java`
|
||||
- `fleetd/src/test/java/dev/ltms/fleet/mcp/FleetMcpTest.java`
|
||||
- Add `fleetd/src/test/java/dev/ltms/fleet/BackendOutageFlowTest.java`.
|
||||
- `fleetd/fleetd.example.yaml`
|
||||
- `CLAUDE.md`
|
||||
|
||||
No earlier unit edits these files.
|
||||
|
||||
The lead, not a worker, must update `wiki/7-Use-Cases.md`, `wiki/9-Implementation.md`, and
|
||||
`wiki/11-Features.md`. Project rules forbid workers from committing `wiki/`. The portable block in
|
||||
`CLAUDE.md` and `wiki/7-Use-Cases.md` must remain byte-identical.
|
||||
|
||||
### Acceptance criteria
|
||||
|
||||
1. `errorPattern` binds per profile. Blank uses the legacy default and is reported as degraded
|
||||
coverage. A config reload which changes it is reported as deferred because patterns are compiled
|
||||
at startup.
|
||||
2. A malformed `errorPattern` stops startup and names `profiles.<name>.errorPattern`.
|
||||
3. The real production sink never silently drops an unknown target. A test captures its error log
|
||||
and the fallback notice call.
|
||||
4. One real backend-error classification through `CompletionResolver` marks only its member. It does
|
||||
not cool the credential and does not send an outage notice.
|
||||
5. Two real classifications for one credential within 60 seconds start one incident.
|
||||
6. The integration test then calls the real explicit-profile spawn gate. It is refused before any
|
||||
adapter spawn call, with “cooling off” and remaining seconds in the message.
|
||||
7. The same test calls an automatic placement path. A cooling candidate is skipped. If every
|
||||
candidate is cooling, the error names cool-off rather than exhaustion.
|
||||
8. Profiles sharing the credential are all blocked. A profile on another credential stays usable.
|
||||
9. `fleet_list` from the same fixture shows both members as `backend_error`, preserves each failure
|
||||
reason, and reports `free: 0`, the credential, and `coolingOffForSeconds`.
|
||||
10. `fleet_profiles` reports cool-off separately from quarantine.
|
||||
11. The real `ReplyPushLoop` receives one incident and makes one successful lead-pane send. Existing
|
||||
failed-ticket notice content may share that same combined send.
|
||||
12. Exhaustion still wins when a line matches both patterns. A simultaneous exhaustion quarantine
|
||||
also wins in spawn errors and fleet views.
|
||||
13. After the 60-second cool-off, spawn is allowed again. A new incident needs two fresh errors.
|
||||
14. `healthCoverage` has the same value before and after this change for the same health config.
|
||||
15. `fleetd.example.yaml` explains `errorPattern`, the legacy fallback, 2/60/60 policy, and the
|
||||
difference between cool-off and exhaustion quarantine.
|
||||
16. `CLAUDE.md` tells leads how `fleet_profiles` and `fleet_list` report cool-off. The lead later
|
||||
applies the matching wiki updates and runs the documented byte-sync check.
|
||||
17. The developer records the new end-to-end test failing before implementation, then passing. The
|
||||
focused suites and `mvn clean install` pass.
|
||||
|
||||
### Dependencies
|
||||
|
||||
Unit 5 starts only after Units 1 to 4 are merged or rebased into its branch. It also starts after the
|
||||
#234 defect 2 fix lands, because both areas touch the same target-resolution control path.
|
||||
|
||||
### What to report back
|
||||
|
||||
- The exact commits used for Units 1 to 4 and #234.
|
||||
- The startup coverage line with one configured and one legacy-default profile.
|
||||
- The red test output before the implementation and its green result after.
|
||||
- The explicit and automatic spawn refusal text.
|
||||
- Sample `fleet_list` and `fleet_profiles` JSON for cool-off and exhaustion.
|
||||
- The number and text of lead-pane sends in the real-path test.
|
||||
- The focused test commands and final `mvn clean install` result.
|
||||
- The exact `CLAUDE.md` change and the wiki edits the lead must apply.
|
||||
|
||||
## File ownership summary
|
||||
|
||||
| Area | Unit | Shared edit risk |
|
||||
|---|---:|---|
|
||||
| Completion classification | 1 | Only Unit 1 edits `CompletionResolver` and its test. |
|
||||
| Correlation and cool-off state | 2 | New files only. |
|
||||
| Lead push scheduling | 3 | Only Unit 3 edits `ReplyPushLoop` and its test. |
|
||||
| Member terminal state | 4 | Only Unit 4 edits `MemberSession`, `SessionManager`, and their test. |
|
||||
| Main composition, config, placement, MCP views, shipped prompt | 5 | Only Unit 5 edits `Fleetd`, `FleetConfig`, `CompositePeerLauncher`, placement context, `FleetMcp`, and `CLAUDE.md`. |
|
||||
| Wiki propagation | Lead after Unit 5 | Workers do not commit the wiki submodule. |
|
||||
|
||||
## What I would not build
|
||||
|
||||
1. **Do not reuse `BackendQuarantine` for outages.** Its repeat call restarts a long credential
|
||||
quarantine. Its fields and errors say “exhausted”. That is wrong for a short outage.
|
||||
2. **Do not merge the exhaustion and generic error patterns.** Exhaustion must win because it has a
|
||||
different policy and duration.
|
||||
3. **Do not group by error string.** One outage can produce different text. The shared operational
|
||||
limit is the credential.
|
||||
4. **Do not mark a profile unusable until config changes.** The current classifier cannot safely
|
||||
tell a permanent malformed request from a transient service fault. A permanent state would need
|
||||
a stronger error taxonomy first.
|
||||
5. **Do not quarantine on the first generic backend error.** That would turn one bad request or one
|
||||
false pattern match into a fleet-wide capacity loss.
|
||||
6. **Do not add this to `FleetHealthMonitor`.** The monitor samples slow member health. The exact
|
||||
backend event already exists at completion resolution, and moving it to polling would lose type
|
||||
and time.
|
||||
7. **Do not add another direct lead injector.** `ReplyPushLoop` already owns status gating,
|
||||
per-lead coalescing, retry bounds, and heartbeat stand-down.
|
||||
8. **Do not change `healthCoverage` to `full`.** That field still means a webhook notification sink
|
||||
exists for periodic health. A backend outage nudge does not make every health event visible.
|
||||
9. **Do not persist incident history across daemon restart in this work.** Existing exhaustion
|
||||
quarantine is also in memory. A 60-second state does not justify a new durable store.
|
||||
10. **Do not build work recovery.** The PR-body survival story proves why checkpoint-first work is
|
||||
useful, but these tickets are about detection, capacity, and signalling.
|
||||
11. **Do not remove the legacy `API Error:` fallback in the first release.** Doing so would turn an
|
||||
unedited config back into a false successful completion. Report it as degraded coverage instead.
|
||||
12. **Do not reorder or add the old `visibleTurn` fallback from #201.** #211 already implemented the
|
||||
narrow raw-scrape fallback at `CompletionResolver.classifyRawScrapeFallback`.
|
||||
|
||||
## Riskiest assumption and cheapest experiment
|
||||
|
||||
The riskiest assumption is that a configured error regex means “the backend failed this turn”. The
|
||||
current code and test already show the counterexample: a worker may quote `API Error:` while writing
|
||||
a valid report. Two such false matches on one credential would now remove capacity for 60 seconds.
|
||||
|
||||
The cheapest experiment is a replay corpus before Unit 5 ships:
|
||||
|
||||
1. Save the full pane text from the measured 2026-09-01 outage.
|
||||
2. Produce one safe failure per backend with a disposable invalid endpoint or request.
|
||||
3. Save one valid member report which quotes each error line.
|
||||
4. Replay all samples through the real `CompletionResolver` test fixture.
|
||||
5. Require outage samples to match and quoted-report samples not to match after assistant-block
|
||||
extraction and baseline checks.
|
||||
|
||||
This costs no outage deployment and no real sleep. If quoted reports still match, narrow the profile
|
||||
patterns before enabling correlation. Do not raise the threshold to hide a bad classifier.
|
||||
|
||||
## Sequencing with three developers
|
||||
|
||||
First wave:
|
||||
|
||||
1. Developer A: Unit 1, typed classification.
|
||||
2. Developer B: Unit 2, credential outage policy.
|
||||
3. Developer C: Unit 3, lead outage nudge.
|
||||
|
||||
As soon as one slot is free, start Unit 4. It is file-disjoint from every first-wave unit. Merge and
|
||||
review Units 1 to 4 independently.
|
||||
|
||||
Start Unit 5 only after all four foundations and #234 are available. Unit 5 is the only high-conflict
|
||||
integration branch, so no other active unit should touch its file list.
|
||||
|
||||
## Checks performed for this refinement
|
||||
|
||||
- Read issue #201 and issue #227 through their Gitea pages. Both showed zero comments.
|
||||
- Read the source and tests named in the evidence section.
|
||||
- Ran `git status --short --branch`; the branch was clean before this document was added.
|
||||
- Ran `git log --oneline -12` to identify the branch base.
|
||||
- I did not run Maven because this change adds only a design document.
|
||||
- Rendered both Mermaid blocks with `npx @mermaid-js/mermaid-cli`; both commands succeeded.
|
||||
+123
-324
@@ -1,311 +1,162 @@
|
||||
# MCP Contract — `fleetd`'s unified gateway
|
||||
# MCP flows and error model — `fleetd`
|
||||
|
||||
> **Status: 🔴 HISTORICAL DESIGN — do NOT use as the tool reference.** Written 2026-07-14, before
|
||||
> any MCP code existed. The system shipped and this page never caught up, so **its tool names,
|
||||
> parameter names and REST paths are wrong today**. Audited 2026-08-17; the specific drift:
|
||||
> **What this page is.** The **flows**: how a delegation, a clarification, a detached task and a
|
||||
> silent member each travel through `fleetd`. These shapes are what shipped, and they are hard to
|
||||
> read off the code because they span the MCP face, the rendezvous registry, the `Injector` and
|
||||
> herdr.
|
||||
>
|
||||
> - **Tools it names that do not exist:** `fleet_read`, `fleet_cancel`.
|
||||
> - **Shipped tools it omits:** `fleet_poll`, `fleet_ack`, `fleet_profiles`, `fleet_whoami`.
|
||||
> - **Parameter names are wrong nearly everywhere** — it says `message`/`target`/`timeout_seconds`/
|
||||
> `block` where the code takes `content`/`sessionId`/`timeoutMs`/`wait`; `text` where
|
||||
> `fleet_reply` takes `content`; `target` where `fleet_stop` takes `paneId`.
|
||||
> - **REST paths are wrong:** it says `POST /workers` and `DELETE /workers/{paneId}`; the daemon
|
||||
> serves `POST /members` and `DELETE /members/{paneId}`.
|
||||
> **What this page is NOT: a tool reference.** It deliberately holds no tool catalogue, no
|
||||
> parameter tables and no REST paths. **The live MCP schema is the authority** — each tool's own
|
||||
> description and parameters, as mounted — with the intent→tool table in `CLAUDE.md` as the short
|
||||
> form.
|
||||
>
|
||||
> **The authoritative tool surface is the live MCP schema** (each tool's own description and
|
||||
> parameters, as mounted), with the intent→tool table in `CLAUDE.md` as the short form. Both were
|
||||
> checked against `mcp/FleetMcp.java` on 2026-08-17 and are accurate.
|
||||
> That absence is the fix for fleetd #114 (CB-609), and it is worth stating why. This page used to
|
||||
> carry a full tool catalogue written in July 2026, before any MCP code existed. The code shipped;
|
||||
> the page did not follow. By August it named two tools that do not exist, omitted five that do,
|
||||
> had the wrong name for nearly every parameter, pointed at REST paths the daemon does not serve,
|
||||
> and — worst — still described an identity model (*"any connection that does not map to a known
|
||||
> worker is treated as a primary"*) that was a real privilege bug, fixed since by the ancestry
|
||||
> walk in fleetd #161. Every one of those errors is the same error: **a second, hand-maintained
|
||||
> copy of something the code already states**. So the second copy is gone rather than corrected.
|
||||
> Only the flows remain, because a flow is a shape rather than a name, and shapes are what this
|
||||
> page was ever good for.
|
||||
>
|
||||
> What is still worth reading here is **§6 — the flows and the error model** (rendezvous,
|
||||
> `fleet_ask`, detached delivery, the turn-done fallback). The shapes it describes are the ones
|
||||
> that shipped; only the names around them drifted. Rewriting this page is tracked as **CB-609**.
|
||||
|
||||
`fleetd` is the **sole communication gateway** for every Claude session in the bridge. Both
|
||||
the **primary** (Opus, on subscription) and every **worker** (off-subscription Claude Code)
|
||||
mount the *same* MCP server with a single `claude mcp add` line, and talk only through its
|
||||
tools. No Claude session ever addresses a broker, a peer, or the network directly.
|
||||
|
||||
This document defines every MCP tool that face must expose, who may call it, its blocking
|
||||
semantics, and how it maps onto the code already in the tree.
|
||||
> The names that do appear below are checked by `McpContractDocTest`, which fails if this page
|
||||
> names a `fleet_*` tool the server does not register. That test is the whole reason it is safe to
|
||||
> write a tool name here at all.
|
||||
|
||||
---
|
||||
|
||||
## 1. Design constraints (non-negotiable)
|
||||
## 1. Rendezvous flows
|
||||
|
||||
These come from the project's core invariants and bound every decision below.
|
||||
### 1.1 Delegation — happy path
|
||||
|
||||
1. **One server, both roles.** The primary and all workers mount an identical server. The
|
||||
catalog must serve both, and `fleetd` must decide *who is calling* from the connection —
|
||||
never from a caller-supplied argument that could be spoofed.
|
||||
2. **Subscription-safe by construction.** No MCP tool ever reads, sets, or forwards
|
||||
`ANTHROPIC_BASE_URL`. Mounting the bridge cannot move a session off subscription.
|
||||
Enforced today by [`SubscriptionGuard`](1-Architecture).
|
||||
3. **Blocking rendezvous, no busy-poll.** The primary consumes a worker's reply through a
|
||||
*single* MCP call that `fleetd` holds open — never a cross-turn poll loop that would burn
|
||||
subscription quota.
|
||||
4. **Status-gated delivery.** Anything that puts text into a worker flows through the existing
|
||||
[`Injector`](1-Architecture): delivered only when the worker is `idle`/`blocked`, at most
|
||||
one message per turn.
|
||||
5. **`fleetd` owns policy; herdr owns PTYs.** MCP tools express *intent*; `fleetd`
|
||||
translates it into guard checks, rendezvous bookkeeping, and herdr `agent.*` calls.
|
||||
|
||||
---
|
||||
|
||||
## 2. Topology
|
||||
|
||||
Both faces live in the one daemon. The **north face** is MCP (this document); the **south
|
||||
face** is the herdr Unix socket. REST/SSE remains only for non-Claude clients and dashboards.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
OPUS["Opus — primary<br/>(Claude Code, env CLEAN)<br/>MCP client"]
|
||||
subgraph BD["fleetd — standalone daemon"]
|
||||
MCP["MCP server (north face)<br/>fleet_send · fleet_reply<br/>fleet_ask · fleet_status · lifecycle"]
|
||||
RDV["rendezvous registry<br/>(blocking-call waiters)"]
|
||||
INJ["Injector + StatusPoller<br/>(status-gated writer)"]
|
||||
SOCK["herdr socket client (south face)"]
|
||||
MCP --> RDV
|
||||
RDV --> INJ
|
||||
INJ --> SOCK
|
||||
MCP --> SOCK
|
||||
end
|
||||
HERDR["herdr<br/>panes · agent-status"]
|
||||
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
|
||||
|
||||
OPUS -->|"fleet_send (blocks)"| MCP
|
||||
W -.->|"fleet_reply / fleet_ask"| MCP
|
||||
SOCK -->|"agent.start · agent.send<br/>agent.get · pane.close"| HERDR
|
||||
HERDR -->|"drives PTY"| W
|
||||
|
||||
classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
|
||||
classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
|
||||
class OPUS,W ext
|
||||
class MCP,RDV,INJ,SOCK core
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Identity & addressing
|
||||
|
||||
Because the same server is mounted by everyone, `fleetd` resolves the caller's role on every
|
||||
request — this is the linchpin of the whole contract and has no code yet.
|
||||
|
||||
- **Workers are known.** `fleetd` spawns every worker
|
||||
([`WorkerService`](1-Architecture)) and records its herdr session UUID / `terminal_id` on
|
||||
the returned [`Agent`]. When a call arrives on a connection that maps to a known worker,
|
||||
the caller is *that* worker — so **workers never pass a target**; routing is implicit.
|
||||
- **The primary is "not a worker".** Any connection that does not map to a known worker is
|
||||
treated as a primary. It addresses workers **explicitly** by `target` — a session UUID,
|
||||
a `terminal_id`, or a friendly `profile` name.
|
||||
- **Turn correlation.** A blocking `fleet_send` registers a *waiter* keyed by worker
|
||||
identity. A worker's later `fleet_reply` / `fleet_ask` on the same identity resolves that
|
||||
waiter. A `turn_id` is minted per exchange so a clarification round-trip
|
||||
(§6.2) rejoins the right turn.
|
||||
|
||||
---
|
||||
|
||||
## 4. Transport
|
||||
|
||||
`fleetd` is a long-lived daemon serving **multiple** concurrent clients (one primary + N
|
||||
workers), so a per-client stdio child is the wrong shape. The recommended transport is
|
||||
**streamable-HTTP / SSE** on the same bind as the REST face:
|
||||
|
||||
```bash
|
||||
# identical on primary and every worker
|
||||
claude mcp add --transport http fleetd http://127.0.0.1:8080/mcp
|
||||
```
|
||||
|
||||
This adds an MCP-server dependency the pom does not yet carry. See [Open decisions](#10-open-decisions).
|
||||
|
||||
---
|
||||
|
||||
## 5. Tool catalog
|
||||
|
||||
| Tool | Caller | Blocks? | Backing (exists today?) |
|
||||
|---|---|---|---|
|
||||
| [`fleet_send`](#fleet_send) | primary | yes (default) | `Injector.enqueue` ✅ · rendezvous registry ❌ (CB-104) |
|
||||
| [`fleet_reply`](#fleet_reply) | worker | no | rendezvous ❌ · pane injection via `Injector` ✅ |
|
||||
| [`fleet_ask`](#fleet_ask) | worker | yes | reverse rendezvous ❌ |
|
||||
| [`fleet_status`](#fleet_status) | either | no | `AgentControl.status` ✅ · `Injector.activeTargets` ✅ |
|
||||
| [`fleet_spawn`](#lifecycle) | primary | no | `WorkerService.spawn` ✅ (`POST /workers`) |
|
||||
| [`fleet_list`](#lifecycle) | either | no | `WorkerService.list` ✅ (`/agents`) |
|
||||
| [`fleet_stop`](#lifecycle) | primary | no | `WorkerService.stop` ✅ (`DELETE /workers/{paneId}`) |
|
||||
| [`fleet_read`](#fleet_read) | primary | no | `AgentControl.read` ✅ |
|
||||
| [`fleet_cancel`](#fleet_cancel) | primary | no | — ❌ (future) |
|
||||
|
||||
### Core: delegation & rendezvous
|
||||
|
||||
#### `fleet_send`
|
||||
*(primary → worker — the headline tool, CB-104)*
|
||||
|
||||
- **Params:** `message` (required); `target` (optional — defaults to the sole worker / default
|
||||
profile); `timeout_seconds` (default 600); `block` (default `true`); `auto_spawn`
|
||||
(default `true`); `turn_id` (optional — supplied when answering a worker's `fleet_ask`).
|
||||
- **Blocking (`block:true`):** enqueue `message` via the `Injector`, then hold the call open
|
||||
until exactly one of:
|
||||
- worker calls `fleet_reply` → `{ outcome:"reply", text }`
|
||||
- worker calls `fleet_ask` → `{ outcome:"question", text, turn_id }`
|
||||
- worker's `agent_status` reaches done/idle with no reply → `{ outcome:"turn_done", text:<terminal tail> }`
|
||||
- deadline elapses → `{ outcome:"timeout" }`
|
||||
- worker gone → error `worker_gone`
|
||||
- **Detached (`block:false`):** enqueue and return `{ outcome:"dispatched", dispatch_id }`
|
||||
immediately. The eventual reply is injected into the primary's idle pane (§6.3), or drained
|
||||
via `fleet_status` on a split-host primary.
|
||||
|
||||
#### `fleet_reply`
|
||||
*(worker → primary)*
|
||||
|
||||
- **Params:** `text` (required); `final` (default `true`).
|
||||
- **Behavior:** resolve the primary waiter registered against this worker with `text`. If no
|
||||
waiter exists (detached delegation), `fleetd` **injects the primary's idle pane** instead.
|
||||
Returns `{ delivered:true, mode:"resolved"|"injected" }`. No `target` — identity is implicit.
|
||||
|
||||
#### `fleet_ask`
|
||||
*(worker → primary — the reverse rendezvous)*
|
||||
|
||||
- **Params:** `question` (required); `timeout_seconds`.
|
||||
- **Behavior:** blocks the *worker's* call. Surfaces the question to the primary (resolving its
|
||||
open `fleet_send` with `outcome:"question"`, or injecting its pane). When the primary
|
||||
answers — a `fleet_send` carrying the matching `turn_id` — that unblocks this call and
|
||||
returns `{ answer }` to the worker, which continues **in the same turn**.
|
||||
|
||||
### Worker lifecycle
|
||||
<a id="lifecycle"></a>
|
||||
Thin adapters over [`WorkerService`](1-Architecture) — parity with the existing REST routes.
|
||||
|
||||
- **`fleet_spawn`** — `{ profile? }` → worker view (`sessionId`, `terminalId`, `paneId`,
|
||||
`status`). Guard-checked; a boundary breach returns error `subscription_boundary` (the
|
||||
REST `403`).
|
||||
- **`fleet_list`** — no params → all workers + `agent_status`. Read-only, either role.
|
||||
- **`fleet_stop`** — `{ target }` → tears down the pane and its dedicated tab. Idempotent.
|
||||
|
||||
### Observability
|
||||
|
||||
#### `fleet_status`
|
||||
*(either role — the README's 4th named tool)*
|
||||
|
||||
- **Params:** `target?`.
|
||||
- **Behavior:** per-worker `agent_status`, queue depth (`Injector.activeTargets`), whether a
|
||||
rendezvous is open, and ids. For the *calling* session it also reports/drains **pending
|
||||
messages addressed to me** — the path a split-host primary's `Stop`-hook uses to wake and
|
||||
collect replies without being injectable. Read-only, non-blocking.
|
||||
|
||||
#### `fleet_read`
|
||||
*(primary)*
|
||||
|
||||
- **Params:** `target`; `source` ∈ `visible | recent | recent_unwrapped | detection`.
|
||||
- **Behavior:** returns the worker's terminal text so the primary can peek at a *detached*
|
||||
worker's progress. Adapter over `AgentControl.read`.
|
||||
|
||||
### Control (future)
|
||||
|
||||
#### `fleet_cancel`
|
||||
*(primary)*
|
||||
|
||||
- **Params:** `target`. Interrupt the worker's current turn / abandon the rendezvous. No
|
||||
backing code yet.
|
||||
|
||||
---
|
||||
|
||||
## 6. Rendezvous flows
|
||||
|
||||
### 6.1 Delegation — happy path
|
||||
|
||||
One blocking call, zero polls.
|
||||
One blocking call, zero polls. The lead's call is held open by `fleetd` until the member answers.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant P as Primary (Opus)
|
||||
participant B as fleetd (MCP + Injector)
|
||||
participant P as "Lead (primary)"
|
||||
participant B as "fleetd (MCP + Injector)"
|
||||
participant H as herdr
|
||||
participant W as Worker (Claude)
|
||||
participant W as "Member"
|
||||
|
||||
P->>B: fleet_send("do X", target=w) — blocks
|
||||
B->>B: register waiter(w)
|
||||
B->>H: agent.send(w, "do X") (idle window)
|
||||
H-->>W: prompt injected
|
||||
W->>W: works the turn
|
||||
W->>B: fleet_reply("result")
|
||||
B->>B: resolve waiter(w)
|
||||
B-->>P: { outcome:"reply", text:"result" }
|
||||
P->>B: "fleet_send{sessionId, content} — blocks"
|
||||
B->>B: "register waiter(sessionId)"
|
||||
B->>H: "agent.send — only in an injectable window"
|
||||
H-->>W: "prompt injected"
|
||||
W->>W: "works the turn"
|
||||
W->>B: "fleet_reply{content}"
|
||||
B->>B: "resolve waiter"
|
||||
B-->>P: "{ outcome: reply }"
|
||||
```
|
||||
|
||||
### 6.2 Clarification — reverse rendezvous (`fleet_ask`)
|
||||
**The cap that matters:** a blocking `fleet_send` is bounded by the *caller's own* MCP client
|
||||
timeout, about 60 seconds — not by the task. Anything slower than that must use the detached flow
|
||||
in §1.3, or the lead's call returns while the member is still working.
|
||||
|
||||
The worker pauses mid-turn to ask; the primary answers; the worker resumes in the same turn.
|
||||
### 1.2 Clarification — reverse rendezvous
|
||||
|
||||
The member pauses mid-turn to ask, the lead answers, and the member resumes **the same turn** with
|
||||
its context intact.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant P as Primary
|
||||
participant P as "Lead"
|
||||
participant B as fleetd
|
||||
participant W as Worker
|
||||
participant W as "Member"
|
||||
|
||||
P->>B: fleet_send("do X", target=w) — blocks
|
||||
B-->>W: "do X" (injected)
|
||||
W->>B: fleet_ask("which config?") — worker blocks
|
||||
B-->>P: { outcome:"question", text:"which config?", turn_id }
|
||||
P->>B: fleet_send("config.yaml", target=w, turn_id) — blocks again
|
||||
B-->>W: resolve fleet_ask → { answer:"config.yaml" }
|
||||
W->>W: resumes same turn
|
||||
W->>B: fleet_reply("done")
|
||||
B-->>P: { outcome:"reply", text:"done" }
|
||||
P->>B: "fleet_send{sessionId, content} — blocks"
|
||||
B-->>W: "content injected"
|
||||
W->>B: "fleet_ask{question} — member blocks"
|
||||
B-->>P: "{ outcome: question, turnId }"
|
||||
P->>B: "fleet_send{turnId, content} — answers THIS turn"
|
||||
B-->>W: "fleet_ask returns the answer"
|
||||
W->>W: "resumes the same turn"
|
||||
W->>B: "fleet_reply{content}"
|
||||
B-->>P: "{ outcome: reply }"
|
||||
```
|
||||
|
||||
### 6.3 Detached delegation — pane injection
|
||||
**Answer with `turnId`, never `sessionId`.** A `sessionId` send starts a new turn; it does not
|
||||
resolve the waiting `fleet_ask`.
|
||||
|
||||
The primary does not block; the reply arrives later in its idle pane.
|
||||
**The window is about 55 seconds and no nudge extends it.** So never brief a member to "ask me":
|
||||
decide before delegating, or give the member an explicit default to fall back on.
|
||||
|
||||
### 1.3 Detached delegation — the lead does not block
|
||||
|
||||
The lead gets a ticket immediately and collects the answer later. This is the flow for any real
|
||||
task, because of the ~60s cap in §1.1.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant P as Primary
|
||||
participant P as "Lead"
|
||||
participant B as fleetd
|
||||
participant W as Worker
|
||||
participant W as "Member"
|
||||
|
||||
P->>B: fleet_send("do X", target=w, block=false)
|
||||
B-->>P: { outcome:"dispatched", dispatch_id }
|
||||
P->>P: continues its own work
|
||||
W->>B: fleet_reply("result")
|
||||
Note over B: no waiter → detached path
|
||||
B->>B: Injector.enqueue(primary_pane, "result")
|
||||
B-->>P: injected into idle pane (status-gated)
|
||||
P->>B: "fleet_send{sessionId, content, wait:false}"
|
||||
B-->>P: "accepted — ticket"
|
||||
P->>P: "continues its own work"
|
||||
W->>B: "fleet_reply{content}"
|
||||
Note over B: "no waiter is blocked — the reply is held"
|
||||
B->>B: "nudge the lead's own pane (status-gated)"
|
||||
P->>B: "fleet_poll{ticket}"
|
||||
B-->>P: "the member's report"
|
||||
P->>B: "fleet_ack{target, msgId}"
|
||||
```
|
||||
|
||||
### 6.4 Uncooperative worker — turn-done fallback
|
||||
A terminal ticket nudges the lead's pane by itself, so a detached task does not need watching. The
|
||||
nudge needs an injectable lead pane and is capped, so it is a convenience rather than a guarantee.
|
||||
|
||||
A worker that never calls `fleet_reply` still returns a result: `fleetd` reads its terminal
|
||||
tail when the turn completes.
|
||||
### 1.4 The member never replies — turn-done fallback
|
||||
|
||||
A member that ends its turn without `fleet_reply` still produces something: `fleetd` reads its
|
||||
pane tail. This is a **fallback, not a channel** — it is lossy in three separate ways, and every
|
||||
one of them has produced a wrong answer in practice.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant P as Primary
|
||||
participant P as "Lead"
|
||||
participant B as fleetd
|
||||
participant W as Worker
|
||||
participant W as "Member"
|
||||
|
||||
P->>B: fleet_send("do X", target=w) — blocks
|
||||
B-->>W: "do X" (injected)
|
||||
W->>W: works, never calls fleet_reply
|
||||
B->>B: StatusPoller sees agent_status → idle/done
|
||||
B->>B: AgentControl.read(w, "recent")
|
||||
B-->>P: { outcome:"turn_done", text:<terminal tail> }
|
||||
P->>B: "fleet_send — blocks or detaches"
|
||||
B-->>W: "content injected"
|
||||
W->>W: "works, never calls fleet_reply"
|
||||
B->>B: "StatusPoller sees the turn end"
|
||||
B->>B: "read the pane tail"
|
||||
B->>B: "classify: exhausted? echoed brief? real report?"
|
||||
B-->>P: "{ outcome: turn_done } or a named failure"
|
||||
```
|
||||
|
||||
The three ways it goes wrong, and what each looks like now:
|
||||
|
||||
| What happened | What the lead used to get | What it gets today |
|
||||
|---|---|---|
|
||||
| The report is longer than the scrape window | The **end** silently cut off | Still clipped, but marked partial |
|
||||
| The member never started — spent credential | The lead's **own brief** echoed back as a report | A named failure: backend exhausted |
|
||||
| The member is simply slow | A tail of work in progress | Unchanged — read it as a hint, not a result |
|
||||
|
||||
The echoed-brief case is the one to remember: it reads as a long, on-topic report with nothing in
|
||||
it from the member. It is suppressed now, but the general rule stands — **check the member's
|
||||
worktree with `git log` before believing a report you did not watch arrive.**
|
||||
|
||||
---
|
||||
|
||||
## 7. Status gating
|
||||
## 2. Status gating
|
||||
|
||||
Delivery only happens in a safe window. This is the state machine the `Injector` already
|
||||
enforces via `AgentStatus.injectable()`; MCP `fleet_send` is simply its producer.
|
||||
Delivery only happens in a safe window. `fleet_send` is a producer for the `Injector`, which
|
||||
already enforces this through `AgentStatus.injectable()`.
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> IDLE
|
||||
IDLE --> WORKING: message delivered / picks up
|
||||
WORKING --> IDLE: turn done
|
||||
WORKING --> BLOCKED: awaits input
|
||||
BLOCKED --> WORKING: input delivered
|
||||
IDLE --> UNKNOWN: detection glitch
|
||||
BLOCKED --> UNKNOWN: detection glitch
|
||||
UNKNOWN --> IDLE: re-detected
|
||||
IDLE --> WORKING: "message delivered, picked up"
|
||||
WORKING --> IDLE: "turn done"
|
||||
WORKING --> BLOCKED: "awaits input"
|
||||
BLOCKED --> WORKING: "input delivered"
|
||||
IDLE --> UNKNOWN: "detection glitch"
|
||||
BLOCKED --> UNKNOWN: "detection glitch"
|
||||
UNKNOWN --> IDLE: "re-detected"
|
||||
|
||||
note right of IDLE
|
||||
injectable — deliver head of FIFO
|
||||
@@ -321,69 +172,17 @@ stateDiagram-v2
|
||||
end note
|
||||
```
|
||||
|
||||
At most one message is delivered per turn: after a send the `Injector` waits for a `WORKING`
|
||||
pickup before delivering the next, with a `PICKUP_GRACE_POLLS` fallback for turns faster than
|
||||
the poll interval. A herdr `events.subscribe` stream can later replace the sampling without
|
||||
touching this state machine.
|
||||
**At most one message per turn.** After a send, the `Injector` waits for a `WORKING` pickup before
|
||||
delivering the next, with a grace-poll fallback for turns that finish faster than the poll
|
||||
interval.
|
||||
|
||||
---
|
||||
Two consequences a lead feels directly:
|
||||
|
||||
## 8. Error model
|
||||
- **A second send to a busy member never lands.** It reports as queued and times out. The member
|
||||
is fine; the message simply waits, and then restarts the member when it next goes idle.
|
||||
- **A spawned member is not deliverable until it has mounted the MCP.** Until then a send waits on
|
||||
that gate for about 60 seconds and then fails without ever reaching the pane.
|
||||
|
||||
| Condition | `fleet_send` result | Notes |
|
||||
|---|---|---|
|
||||
| Worker replies | `{ outcome:"reply" }` | normal |
|
||||
| Worker asks | `{ outcome:"question", turn_id }` | answer with `fleet_send(turn_id)` |
|
||||
| Turn ends, no reply | `{ outcome:"turn_done" }` | terminal tail as text |
|
||||
| Deadline elapsed | `{ outcome:"timeout" }` | message may still be queued/delivered |
|
||||
| Worker vanished | error `worker_gone` | `Injector.drop` fails the queued future |
|
||||
| Guard breach on spawn | error `subscription_boundary` | REST `403` parity |
|
||||
| Delivery failed at herdr | error, message dropped | poisoned message not left blocking the FIFO |
|
||||
|
||||
`fleet_reply` from a worker with no open waiter is **not** an error — it falls through to
|
||||
detached pane injection (§6.3).
|
||||
|
||||
---
|
||||
|
||||
## 9. Mapping to existing code
|
||||
|
||||
The MCP face is a thin adapter layer; nearly every capability already exists behind the REST
|
||||
seam. Only the **rendezvous registry** and the **caller-identity resolver** are new.
|
||||
|
||||
| MCP tool | Existing collaborator | New work |
|
||||
|---|---|---|
|
||||
| `fleet_send` | `Injector.enqueue`, `AgentControl.send` | waiter registry, timeout, outcome mux (CB-104) |
|
||||
| `fleet_reply` / `fleet_ask` | `Injector` (pane injection) | reverse rendezvous, identity resolver |
|
||||
| `fleet_status` | `AgentControl.status`, `Injector.activeTargets` | pending-drain projection |
|
||||
| `fleet_spawn` / `list` / `stop` | `WorkerService.{spawn,list,stop}` | MCP adapter only |
|
||||
| `fleet_read` | `AgentControl.read` | MCP adapter only |
|
||||
|
||||
Because the REST routes in `FleetApp` already exercise the collaborators, MCP tools are
|
||||
validated by **parity** against those routes, not by re-testing behavior.
|
||||
|
||||
---
|
||||
|
||||
## 10. Open decisions
|
||||
|
||||
1. **`fleet_ask` direction.** This page defines it as *worker-asks-primary* (a genuine reverse
|
||||
channel, matching the "inject the primary's pane" language). The alternative — a synonym for
|
||||
a blocking primary→worker send — is weaker and produces different plumbing. **Recommend
|
||||
worker-asks-primary.**
|
||||
2. **Detached delivery shape.** A `block:false` param on `fleet_send` (keeps the catalog
|
||||
small) vs. a separate `fleet_dispatch` tool. **Recommend the param.**
|
||||
3. **Auto-spawn on send.** `fleet_send` provisions a worker per profile when none exists
|
||||
(simplest primary UX) vs. requiring an explicit `fleet_spawn` first. **Recommend
|
||||
auto-spawn, defaulting on.**
|
||||
4. **Transport & SDK.** Streamable-HTTP/SSE co-located with the REST bind (recommended) vs.
|
||||
stdio. Requires choosing a Java MCP server SDK and adding it to the pom.
|
||||
|
||||
---
|
||||
|
||||
## 11. Implementation staging
|
||||
|
||||
- **CB-104** — blocking `fleet_send` + rendezvous registry + caller-identity resolver
|
||||
(the producer that finally drives the inert `StatusPoller`).
|
||||
- **CB-1xx** — `fleet_reply` / `fleet_ask` reverse rendezvous + detached pane injection.
|
||||
- **CB-1xx** — lifecycle + observability adapters (`fleet_spawn/list/stop/status/read`).
|
||||
- **CB-1xx** — transport wiring + `claude mcp add` docs; parity tests vs. REST.
|
||||
- **Later** — `fleet_cancel`; swap `StatusPoller` for herdr `events.subscribe`.
|
||||
`UNKNOWN` is deliberately neither injectable nor a pickup. A pane whose status cannot be read is
|
||||
not a pane that is safe to write to — see fleetd #176 for what happens when a gate treats an
|
||||
unreadable pane as a ready one.
|
||||
|
||||
@@ -172,7 +172,11 @@ herdrSocket: ~/.config/herdr/herdr.sock
|
||||
# skills/MCP/hooks. Omit to leave the worker on the host default.
|
||||
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
|
||||
# worktree sees the same local config (CB-301-ext). Omit for the default set:
|
||||
# [.env, .envrc]. (.claude/settings.local.json is NOT in the default — it
|
||||
# [.env] only (CB-148). .envrc is left out of the default on purpose: it is
|
||||
# executable shell that direnv runs on every cd, so copying it carries
|
||||
# behaviour into the worker, not just values, unlike .env. An operator who
|
||||
# wants it copied can still write parityOverlay: [.env, .envrc] explicitly.
|
||||
# (.claude/settings.local.json is NOT in the default — it
|
||||
# pre-approves IDE/tool grants a member must not hold ambiently; CB-525/CB-634.)
|
||||
#
|
||||
# Do NOT add .mcp.json (CB-525). A worker's tools are whatever its launcher
|
||||
@@ -206,6 +210,39 @@ herdrSocket: ~/.config/herdr/herdr.sock
|
||||
# profile quarantines alone, under its own name, exactly as if the field did
|
||||
# not exist. Cooldown length is the top-level quarantineCooldownSeconds below.
|
||||
# HOT: read live at every spawn/exhaustion check — no restart needed.
|
||||
# errorPattern → fleetd #201 / #227: regex matched against a completion-fallback scrape to
|
||||
# classify a turn that ended with no fleet_reply as a BACKEND ERROR — a
|
||||
# credential outage or a provider 5xx — rather than a real answer or a
|
||||
# usage-limit exhaustion (exhaustedPattern above always wins when a line
|
||||
# matches both). Opt-in. Omit it and this profile falls back to fleetd's
|
||||
# built-in legacy pattern `(?i)\bAPI Error\s*:` — classification still
|
||||
# happens, just without a profile-specific match; every backend words its
|
||||
# failure differently, so a hardcoded sentence would only ever match one
|
||||
# of them.
|
||||
# DEFERRED: compiled once into a startup pattern map, same as exhaustedPattern
|
||||
# — editing it needs a daemon restart.
|
||||
# # errorPattern: "503 Service Unavailable" # opt-in: classify a backend outage
|
||||
#
|
||||
# What happens once a match fires (BackendOutagePolicy, credentialId-keyed,
|
||||
# SEPARATE from the CB-578 stage B quarantine above and never merged with it):
|
||||
# - threshold 2 — TWO DISTINCT TARGETS (never raw events) on the same
|
||||
# effective credential inside a 60-second window start an "incident" and a
|
||||
# 60-second cool-off for that credential. One member repeating the same
|
||||
# classified line twice never cools anything off — a real outage hits
|
||||
# every target on that credential, so requiring a second, independent
|
||||
# target loses nothing against the case this guards against, while
|
||||
# protecting against a heuristic misfire on one flaky member.
|
||||
# - a fresh error while a credential is already cooling off is ignored
|
||||
# outright: it neither extends the 60s deadline nor starts a new incident.
|
||||
# - `fleet_list`/`fleet_profiles` report a cooling credential with
|
||||
# `coolingOffForSeconds` (never `quarantinedForSeconds`, unless CB-578
|
||||
# exhaustion quarantine is ALSO independently active for the same
|
||||
# credential — the two checks can both fire at once). A spawn onto a
|
||||
# cooling profile is refused with a message naming the credential and
|
||||
# remaining seconds — "cooling off", never "exhausted", so an operator can
|
||||
# tell a short transient fault from a spent subscription at a glance.
|
||||
# - the lead gets ONE nudge per incident (not one per affected target), via
|
||||
# the same push loop that already delivers ticket/question reminders.
|
||||
# env → extra environment for this profile's workers, as a literal key/value map
|
||||
# (CB-511). Use it to give workers a toolchain.
|
||||
#
|
||||
@@ -269,13 +306,43 @@ profiles:
|
||||
# GOTCHA 2 — `maxLoad` is the ONLY throttle you have here. There is no metering, no budget
|
||||
# and no refusal on cost; the cap on live members is the single thing standing between a
|
||||
# fan-out and your monthly limit. Set it deliberately and keep it small.
|
||||
#
|
||||
# GOTCHA 3 (fleetd #176) — `maxLoad` counts members, never the lead itself. The lead is a live
|
||||
# `claude` session on this SAME account (a lead is never moved off-subscription, whatever its
|
||||
# own profile says), so it already holds one seat before any member spawns. If a lead's
|
||||
# `fleet.leaders.<name>.profile` names THIS profile — or ANY OTHER `subscription: true`
|
||||
# profile that shares this one's account (see THE SENTINEL, just below, next to
|
||||
# `credentialId:`) — `fleet_list`'s `free` for this profile subtracts that lead's live
|
||||
# seat(s) automatically; see `profile:` under THE FLEET below. If no lead entry names a
|
||||
# profile sharing this account, fleetd has no way to know a lead holds a seat here, and `free`
|
||||
# will overstate what a fresh `fleet_spawn` actually gets by exactly the seats the lead is
|
||||
# quietly holding.
|
||||
#
|
||||
# THE SENTINEL (fleetd #176 stage 2, correcting an inert stage 1 fix): every `subscription:
|
||||
# true` profile that leaves `credentialId` unset shares ONE implicit account-wide credential
|
||||
# id with every other such profile on this host — because a subscription profile doesn't
|
||||
# authenticate with a credential of its own, it authenticates as the operator's own Claude
|
||||
# login, and there is exactly one of those. So on a typical host, `opus` (the lead's profile)
|
||||
# and `sonnet` (the members' profile) are linked automatically, with NOTHING to set here — that
|
||||
# is what makes GOTCHA 3 above work without also writing matching `credentialId:` values on
|
||||
# both. This linkage is not just cosmetic: it is the same key `BackendQuarantine`/cool-off use,
|
||||
# so a usage-limit hit on `opus` now quarantines `sonnet` too (and vice versa) — correct, since
|
||||
# they are one Claude account, but worth knowing before you wonder why an unrelated-looking
|
||||
# profile went quarantined.
|
||||
#
|
||||
# WHEN TO OVERRIDE — set explicit, DIFFERENT `credentialId:` values on two `subscription: true`
|
||||
# profiles only when they are genuinely two separate Claude logins on the same host (a real,
|
||||
# if unusual, setup). An explicit `credentialId` always wins over the sentinel, so this is the
|
||||
# one way to keep two subscription profiles from being treated as one account for lead-seat
|
||||
# counting AND for quarantine/cool-off grouping alike.
|
||||
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
|
||||
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
|
||||
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
|
||||
# credentialId: shared-openai # opt-in: quarantine together with every other profile sharing this id (CB-578)
|
||||
# errorPattern: "503 Service Unavailable" # opt-in: classify a backend outage (fleetd #201/#227) — see the key doc above
|
||||
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
|
||||
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
|
||||
# parityOverlay: [".env", ".envrc"] # the default; never add .mcp.json or .claude/settings.local.json — see above
|
||||
# parityOverlay: [".env"] # the default; add ".envrc" explicitly if you want it copied too (CB-148) — never add .mcp.json or .claude/settings.local.json — see above
|
||||
# ideMcpUrl: http://127.0.0.1:29170/index-mcp/streamable-http # opt-in (CB-634): IDE code intelligence, pinned to the worktree
|
||||
# ideProjectDir: fleetd # CB-634: module dir the IDE opens + the overlay pins (this repo's pom is in fleetd/)
|
||||
# ideOpenCommand: env DISPLAY=:10.0 idea {dir} # CB-634 auto-open: opens {dir} in the IDE at spawn; omit to open by hand
|
||||
@@ -369,6 +436,12 @@ placement: weighted
|
||||
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
|
||||
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
|
||||
# restart. Editing this needs a daemon restart to take effect.
|
||||
#
|
||||
# This does NOT govern the fleetd #201 / #227 backend-error cool-off documented under errorPattern
|
||||
# above — that mechanism is a separate, shorter-lived, NOT-configurable policy (threshold 2 distinct
|
||||
# targets, 60-second window, 60-second cool-off), on purpose: it exists to survive a brief transient
|
||||
# fault, not to replace this 30-minute exhaustion quarantine. Do not conflate the two when reading
|
||||
# fleet_list/fleet_profiles — coolingOffForSeconds and quarantinedForSeconds are independent facts.
|
||||
# quarantineCooldownSeconds: 1800
|
||||
|
||||
# Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an
|
||||
@@ -395,8 +468,10 @@ placement: weighted
|
||||
# stage B — baked once into the quarantine tracker built at startup), ADDING or
|
||||
# REMOVING a profile (a new backend needs its own launcher, and launchers are built
|
||||
# once), AND an existing profile's launch settings — model, baseUrl, argv, env,
|
||||
# configDir, mcpUrl, tabLabel, exhaustedPattern. The launcher takes a copy of
|
||||
# `profiles:` at startup and resolves every spawn out of that copy, so those never
|
||||
# configDir, mcpUrl, tabLabel, exhaustedPattern, errorPattern (fleetd #201 / #227 —
|
||||
# compiled once into a startup pattern map the same way exhaustedPattern is). The
|
||||
# launcher takes a copy of `profiles:` at startup and resolves every spawn out of
|
||||
# that copy, so those never
|
||||
# reach a launch until you restart. The reload logs them by name rather than
|
||||
# pretending they applied.
|
||||
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
|
||||
@@ -457,6 +532,22 @@ fleet:
|
||||
# recognised: give it a `profile:` and the daemon launches the shortfall when fewer than
|
||||
# `instances` are live. Omit `profile:` and it is recognise-only, as before.
|
||||
#
|
||||
# `profile:` has a SECOND job as of fleetd #176, even for a recognise-only lead you never want
|
||||
# auto-launched: it is also how fleetd learns which account this lead's own session shares. A
|
||||
# `subscription: true` profile bills the operator's Claude account, and the lead itself is always
|
||||
# a live `claude` session on that same account — `maxLoad` never counted that seat. If a lead
|
||||
# entry here names a profile that shares a worker profile's account, `fleet_list`'s `free` for
|
||||
# that worker profile subtracts the lead's live seat(s) automatically. "Shares the account" is
|
||||
# decided by matching `effectiveCredentialId()`, which (fleetd #176 stage 2 — see THE SENTINEL,
|
||||
# next to `credentialId:`, in THE WORKERS above) means: an explicit, matching `credentialId:` on
|
||||
# both, OR — the common case, needing NO extra config — both being `subscription: true` with
|
||||
# `credentialId` left unset, since those all share one implicit account-wide id. A lead on `opus`
|
||||
# and workers on `sonnet` link automatically this way; they do NOT need the same profile name.
|
||||
# Setting `profile:` on an already-running, recognise-only lead is safe — the daemon only launches
|
||||
# the SHORTFALL below `instances`, so naming a profile here does not, by itself, start anything.
|
||||
# Omit it and fleetd has no way to derive the sharing — there is no other reliable signal on the
|
||||
# daemon's side — so that lead's seat goes uncounted, exactly as before this ticket.
|
||||
#
|
||||
# `tab:` (CB-579) is REQUIRED and is the only field identity depends on — the exact label of the
|
||||
# tab hosting the lead, matched case-insensitively. Label the tab yourself and put that same
|
||||
# string here, and the pane is recognised on the next rescan. Reopen the tab later, or the session
|
||||
|
||||
@@ -13,6 +13,8 @@ import dev.ltms.fleet.lead.LeadLauncher;
|
||||
import dev.ltms.fleet.herdr.PaneLocator;
|
||||
import dev.ltms.fleet.herdr.UnixSocketHerdrClient;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.inject.BackendErrorPatternLookup;
|
||||
import dev.ltms.fleet.inject.BackendErrorSink;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
|
||||
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||
@@ -49,7 +51,9 @@ import dev.ltms.fleet.session.SessionReaper;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||
import dev.ltms.fleet.member.HerdrPeerLauncher;
|
||||
import dev.ltms.fleet.member.MemberCredentialPolicyView;
|
||||
import dev.ltms.fleet.member.OpenCodeLauncher;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import io.javalin.Javalin;
|
||||
import org.slf4j.Logger;
|
||||
@@ -61,6 +65,7 @@ import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Optional;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
@@ -174,7 +179,13 @@ public final class Fleetd {
|
||||
// genuine cycle. Break it exactly like liveCountRef below: a forwarding sink built now,
|
||||
// pointed at the real one once it exists.
|
||||
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
|
||||
ExhaustionSink forwardingExhaustionSink = (target, reason) -> exhaustionSinkRef.get().onExhausted(target, reason);
|
||||
// fleetd #234, round 4: routed through the shared ExhaustionSink.forwardingTo factory
|
||||
// rather than written inline here — not because a lambda at this call site is unsafe
|
||||
// anymore (it is not: the 3-arg overload is now the interface's single abstract method, so
|
||||
// there is no 2-arg overload left for any lambda to silently bind to instead), but so a
|
||||
// test can call the exact same object this line builds, instead of asserting a copy of its
|
||||
// shape (round 3's lesson).
|
||||
ExhaustionSink forwardingExhaustionSink = ExhaustionSink.forwardingTo(exhaustionSinkRef::get);
|
||||
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
|
||||
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
|
||||
// unless opencode is the only kind configured.
|
||||
@@ -199,12 +210,18 @@ public final class Fleetd {
|
||||
// here, at startup, and a config reload only changes it for a daemon restart.
|
||||
BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,
|
||||
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
|
||||
// fleetd #201 Unit 5: one outage-cool-off tracker for the whole daemon, shared between the
|
||||
// launcher (checked at spawn, like `quarantine` above) and the backend-error sink wired in
|
||||
// below (written on a classified backend error). A SEPARATE, shorter-lived mechanism from
|
||||
// `quarantine` — see BackendOutagePolicy's class doc — never merged with it.
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(System::nanoTime);
|
||||
PeerLauncher workers = new CompositePeerLauncher(
|
||||
adapters,
|
||||
cfg.effectiveDefaultProfile(),
|
||||
config,
|
||||
profileName -> liveCountRef.get().apply(profileName),
|
||||
quarantine);
|
||||
quarantine,
|
||||
outagePolicy);
|
||||
// CB-504: under supervision (launchd/systemd) fleetd can start before herdr's socket
|
||||
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
|
||||
// the first thing that actually talks to herdr, so without this wait a boot-order race
|
||||
@@ -335,7 +352,29 @@ public final class Fleetd {
|
||||
.map(session -> exhaustedPatternsByProfile.get(session.profile()))
|
||||
.orElse(null);
|
||||
log.info("backend-exhausted classification (CB-578 stage A): {}",
|
||||
CompletionResolver.coverage(cfg.profiles().keySet(), exhaustedPatternsByProfile.keySet()));
|
||||
CompletionResolver.coverage("exhaustedPattern", cfg.profiles().keySet(),
|
||||
exhaustedPatternsByProfile.keySet()));
|
||||
// fleetd #201 Unit 5: classify a completion-fallback scrape that matches a profile's
|
||||
// configured backend-error refusal (a credential outage, a provider 5xx) as a backend error
|
||||
// rather than handing it back as a real answer. Compiled once at startup, keyed by profile
|
||||
// name, mirroring exhaustedPatternsByProfile above — a profile with no configured
|
||||
// errorPattern is simply absent here, so CompletionResolver falls back to its built-in
|
||||
// narrow {@code (?i)\bAPI Error\s*:} compatibility pattern for that profile's targets
|
||||
// (BackendErrorPatternLookup#legacy's contract — see backendErrorPatterns below).
|
||||
Map<String, Pattern> errorPatternsByProfile = new LinkedHashMap<>();
|
||||
cfg.profiles().forEach((name, profile) -> {
|
||||
if (profile.hasErrorPattern()) {
|
||||
errorPatternsByProfile.put(name, Pattern.compile(profile.errorPattern()));
|
||||
}
|
||||
});
|
||||
// fleetd #248: extracted to a static factory (see backendErrorPatternLookup below) so a
|
||||
// test can prove main() actually PASSES this into CompletionResolver, not only that the
|
||||
// lookup itself behaves correctly — the exact gap fleetd #248 exists to close.
|
||||
BackendErrorPatternLookup backendErrorPatterns = backendErrorPatternLookup(sessions::roster,
|
||||
errorPatternsByProfile);
|
||||
log.info("backend-error classification (fleetd #201 Unit 5): {}",
|
||||
CompletionResolver.coverage("errorPattern", cfg.profiles().keySet(),
|
||||
errorPatternsByProfile.keySet()));
|
||||
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
|
||||
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
|
||||
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
|
||||
@@ -345,22 +384,68 @@ public final class Fleetd {
|
||||
// model-mismatch check — it needs nothing profile-specific from the caller beyond `target`
|
||||
// (a herdr terminal id) and `reason`, so reusing it here is exactly "the existing
|
||||
// ExhaustionSink path", not a new mechanism.
|
||||
ExhaustionSink exhaustionSink = (target, reason) -> sessions.roster().stream()
|
||||
.filter(session -> target.equals(session.terminalId()))
|
||||
.findFirst()
|
||||
.map(MemberSession::profile)
|
||||
.map(profileName -> config.get().profiles().get(profileName))
|
||||
.ifPresent(profile -> {
|
||||
String credentialId = profile.effectiveCredentialId();
|
||||
quarantine.quarantine(credentialId);
|
||||
log.warn("credential '{}' quarantined for {}s (profile '{}'): {}", credentialId,
|
||||
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
|
||||
});
|
||||
//
|
||||
// fleetd #234: that check fires from SessionAwareHandle.agentSessionId(), which runs during
|
||||
// SessionManager.acquire() BEFORE this session is registered in sessions.roster() — so the
|
||||
// roster-only lookup below used to find nothing, .ifPresent silently no-op'd, and the
|
||||
// ERROR the check had just logged ("quarantining this profile's credential") was a lie:
|
||||
// nothing was quarantined, and nothing said so. Two changes: (1) OpenCodeLauncher now
|
||||
// passes its OWN profile name via ExhaustionSink's 3-arg overload — it already has the
|
||||
// FleetConfig.Profile in hand and does not need the roster at all — used here as a
|
||||
// fallback whenever the roster lookup misses; (2) if a profile still cannot be resolved
|
||||
// (neither the roster nor the hint names a configured one), this logs loudly at ERROR
|
||||
// instead of silently doing nothing — a control that cannot act must say so.
|
||||
// fleetd #234, round 4: the 3-arg overload is now ExhaustionSink's single abstract method,
|
||||
// so this is safely a lambda — there is no separate 2-arg overload left for it to bind to
|
||||
// instead and silently drop profileHint (that was rounds 1-3's whole hazard).
|
||||
ExhaustionSink exhaustionSink = (target, reason, profileHint) -> {
|
||||
String profileName = sessions.roster().stream()
|
||||
.filter(session -> target.equals(session.terminalId()))
|
||||
.findFirst()
|
||||
.map(MemberSession::profile)
|
||||
.orElse(profileHint);
|
||||
FleetConfig.Profile profile = profileName == null ? null : config.get().profiles().get(profileName);
|
||||
if (profile == null) {
|
||||
log.error("quarantine requested for target '{}' ({}) but no profile could be "
|
||||
+ "resolved — the target is not (yet) in the roster, and {} — "
|
||||
+ "credential NOT quarantined (fleetd #234)",
|
||||
target, reason,
|
||||
profileHint == null ? "no profile hint was given"
|
||||
: "the hinted profile '" + profileHint + "' is not configured");
|
||||
return;
|
||||
}
|
||||
String credentialId = profile.effectiveCredentialId();
|
||||
quarantine.quarantine(credentialId);
|
||||
log.warn("credential '{}' quarantined for {}s (profile '{}'): {}", credentialId,
|
||||
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
|
||||
};
|
||||
// fleetd #175: point the forwarding sink handed to OpenCodeLauncher above at the real one,
|
||||
// now that `sessions` exists to resolve target -> session -> profile.
|
||||
exhaustionSinkRef.set(exhaustionSink);
|
||||
// fleetd #201 Unit 5: the production BackendErrorSink needs `pushLoop` (built further below,
|
||||
// after `sessions`) to tell a lead about an incident or an unmapped target — the same
|
||||
// construction-order cycle `exhaustionSinkRef` breaks above, broken the same way: a mutable
|
||||
// holder set once `pushLoop` exists, read lazily from inside the lambda built here.
|
||||
AtomicReference<ReplyPushLoop> pushLoopRef = new AtomicReference<>();
|
||||
// fleetd #248: extracted to a static factory (see backendErrorSink below), public rather
|
||||
// than package-private like the other two factories here, so
|
||||
// dev.ltms.fleet.inject.BackendOutageFlowTest can exercise the REAL production sink
|
||||
// directly instead of a hand-mirrored copy of this lambda — that copy was precisely the
|
||||
// gap fleetd #248 exists to close (see that test's class doc for the history).
|
||||
BackendErrorSink backendErrorSink = backendErrorSink(sessions, () -> config.get().profiles(),
|
||||
outagePolicy, pushLoopRef::get);
|
||||
AgentControl agents = router.memberAgents();
|
||||
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns, exhaustionSink);
|
||||
// Both fleetd#201 Unit 5 (backend-error patterns + sink) and fleetd#241 (the worktree/branch
|
||||
// lookup the fallback report names) land on this one call. The full constructor takes both,
|
||||
// so neither feature is dropped; nowNanos must be passed explicitly to reach it.
|
||||
//
|
||||
// fleetd #248: every argument built specifically for this call (backendErrorPatterns and
|
||||
// backendErrorSink above, and the worktree/branch lookup right here) now comes from a
|
||||
// static factory tested on its own; FleetdCompletionResolverWiringTest source-asserts that
|
||||
// THIS call actually passes them, which is the coverage that was missing before.
|
||||
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns,
|
||||
exhaustionSink, backendErrorPatterns, backendErrorSink, System::nanoTime,
|
||||
worktreeBranchLookup(sessions::roster));
|
||||
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
|
||||
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
|
||||
MemberPresence presence = sessions.asPresence();
|
||||
@@ -442,6 +527,9 @@ public final class Fleetd {
|
||||
Metrics metrics = FleetMetrics.create(sessions, replyInbox);
|
||||
var pushLoop = new ReplyPushLoop(primaryRegistry, router.leadAgents(), replyInbox,
|
||||
pushScheduler, maxReminders, backoffMs, metrics);
|
||||
// fleetd #201 Unit 5: point the forwarding holder captured by the backendErrorSink lambda
|
||||
// above at the real push loop, now that it exists.
|
||||
pushLoopRef.set(pushLoop);
|
||||
// CB-551: idle-lead heartbeat. Opt-in; absent `leadHeartbeat:` this is never constructed, so
|
||||
// an upgraded daemon cannot silently start spending subscription on nudging an idle lead.
|
||||
// It has its own single-thread scheduler and holds its own scheduler shutdown via close().
|
||||
@@ -549,7 +637,12 @@ public final class Fleetd {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.effectiveCredentialId();
|
||||
}, quarantine),
|
||||
leadMailbox);
|
||||
leadMailbox,
|
||||
new FleetMcp.OutageSource(profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.effectiveCredentialId();
|
||||
}, outagePolicy),
|
||||
new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), leaders, leads)));
|
||||
|
||||
// CB-637: the receive half. Only constructed when a lead mailbox actually opened — with no
|
||||
// coordinator (or an unreachable one) there is nothing to deliver, so no scheduler is
|
||||
@@ -617,8 +710,11 @@ public final class Fleetd {
|
||||
|
||||
// CB-185: give FleetApp both daemons — /healthz must require both to answer and
|
||||
// GET /sessions must merge across both, or a down/unpolled member daemon is invisible.
|
||||
// fleetd #111: live (re-read-per-request) memberCredentials view for GET /member-credentials —
|
||||
// same hot-reload shape as the memberCredentials supplier passed to ClaudeCodeLauncher above.
|
||||
Javalin app = new FleetApp(herdr, memberHerdr, workers, sessions, messages, presence, mcp.servlet(),
|
||||
callers, metrics, deliverable).build();
|
||||
callers, metrics, deliverable,
|
||||
() -> MemberCredentialPolicyView.of(config.get().memberCredentials())).build();
|
||||
app.start(cfg.bind().host(), cfg.bind().port());
|
||||
log.info("fleetd listening on {}:{}, herdr socket {}",
|
||||
cfg.bind().host(), cfg.bind().port(), socket);
|
||||
@@ -647,6 +743,185 @@ public final class Fleetd {
|
||||
return target -> presence.isPresent(target) || leads.get().containsKey(target);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #248: package-private factory for the member worktree/branch lookup {@link
|
||||
* CompletionResolver} uses to name a fallback report's worktree and branch (fleetd#241).
|
||||
*
|
||||
* <p>Before this ticket the lookup was an anonymous lambda built inline inside {@code main}'s
|
||||
* {@code CompletionResolver} constructor call — provably untested wiring, the whole reason
|
||||
* fleetd #248 exists: dropping that one argument (passing {@code _ -> null} instead) compiled
|
||||
* clean and left every test green. Extracted here, {@code main} now calls this factory instead
|
||||
* of building the lambda inline, and a source assertion on that call site
|
||||
* ({@code FleetdCompletionResolverWiringTest}) proves the argument is still actually passed.
|
||||
*
|
||||
* <p>Takes the roster as a plain {@link Supplier} — not a {@link SessionManager} — so this is
|
||||
* directly testable with a hand-built session list; no real {@code SessionManager} (launcher,
|
||||
* worktrees, …) needs constructing. Follows the same {@code static} factory pattern as
|
||||
* {@link #deliverableTo} above.
|
||||
*
|
||||
* @param roster the live member roster, normally {@code sessions::roster}
|
||||
*/
|
||||
static Function<String, CompletionResolver.WorktreeBranch> worktreeBranchLookup(
|
||||
Supplier<List<MemberSession>> roster) {
|
||||
return target -> roster.get().stream()
|
||||
.filter(session -> target.equals(session.terminalId()))
|
||||
.findFirst()
|
||||
.map(session -> new CompletionResolver.WorktreeBranch(session.worktree(), session.branch()))
|
||||
.orElse(null);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #176: per-profile factory for {@link FleetMcp.LeadSeatSource} — how many seats a
|
||||
* profile's own live LEAD session(s) hold on the same Claude subscription.
|
||||
*
|
||||
* <p>{@code maxLoad} counts only members; the lead itself is a live {@code claude} session that
|
||||
* is never moved off-subscription ({@code LeadLauncher} strips {@code ANTHROPIC_BASE_URL}/
|
||||
* {@code AUTH_TOKEN} from a lead's env whatever its profile says), so a {@code subscription:
|
||||
* true} profile's real ceiling is lower than its configured {@code maxLoad} by exactly the
|
||||
* number of lead seats sharing that same account.
|
||||
*
|
||||
* <p><b>The derivation, and why this route was chosen over a new config key.</b> The link is
|
||||
* {@code fleet.leaders.<name>.profile} — the field the operator already sets to name which
|
||||
* {@code profiles:} entry a lead runs on (CB-557; see {@code fleetd.example.yaml}) — matched
|
||||
* against the profile passed in here via {@link FleetConfig.Profile#effectiveCredentialId()},
|
||||
* the same grouping key {@link dev.ltms.fleet.placement.BackendQuarantine} already uses to say
|
||||
* two profiles share one account. Nothing new is added to the config schema: this reuses a field
|
||||
* that already exists and already means "the profile this lead's own session runs on". A lead
|
||||
* entry that names no {@code profile:} (recognise-only, CB-558) says nothing about which account
|
||||
* it shares, and there is no other reliable signal on the daemon's side to derive that from — so
|
||||
* such a lead contributes no seats, exactly as before this ticket. Making that lead's seat count
|
||||
* requires the operator to add one line (`profile: sonnet` under its {@code fleet.leaders} entry)
|
||||
* — a config statement, not a code change, and the smallest one available given the field
|
||||
* already exists for a closely related purpose.
|
||||
*
|
||||
* <p>Only counts leads {@code liveLeadTerminals} currently reports — CB-531's live tab scan (or
|
||||
* the legacy {@code primary.terminal} pin) — never every configured lead: an entry whose
|
||||
* {@code instances} nobody has actually started is not really competing for a seat, and must not
|
||||
* shrink capacity for one that was never live.
|
||||
*
|
||||
* @param profiles the live profile map, normally {@code () -> config.get().profiles()}
|
||||
* in {@code main} — hot, like every other {@code maxLoad}/
|
||||
* {@code credentialId} read {@link FleetMcp.CapacitySource} already does
|
||||
* @param leaders {@code fleet.leaders}, read once at startup like the rest of that
|
||||
* block ({@code fleetd.example.yaml} notes it is not hot) — passed as a
|
||||
* plain map, never re-read from {@code config.get()}
|
||||
* @param liveLeadTerminals terminal_id → lead name for every CURRENTLY recognised lead, normally
|
||||
* the same supplier {@link dev.ltms.fleet.auth.CallerResolver#leads()}
|
||||
* and {@code LeadCoordLoop} already consult
|
||||
*/
|
||||
static Function<String, Integer> leadSeatLookup(Supplier<Map<String, FleetConfig.Profile>> profiles,
|
||||
Map<String, FleetConfig.Leader> leaders, Supplier<Map<String, String>> liveLeadTerminals) {
|
||||
return profileName -> {
|
||||
FleetConfig.Profile target = profiles.get().get(profileName);
|
||||
if (target == null || !target.isSubscription()) {
|
||||
return 0;
|
||||
}
|
||||
String targetCredential = target.effectiveCredentialId();
|
||||
int seats = 0;
|
||||
for (String leadName : liveLeadTerminals.get().values()) {
|
||||
FleetConfig.Leader lead = leaders.get(leadName);
|
||||
if (lead == null || lead.profile() == null || lead.profile().isBlank()) {
|
||||
continue;
|
||||
}
|
||||
FleetConfig.Profile leadProfile = profiles.get().get(lead.profile());
|
||||
if (leadProfile != null && targetCredential.equals(leadProfile.effectiveCredentialId())) {
|
||||
seats++;
|
||||
}
|
||||
}
|
||||
return seats;
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #248 / fleetd#201 Unit 5: package-private factory for the per-target backend-error
|
||||
* pattern lookup {@link CompletionResolver} classifies a pane scrape against. Closes over the
|
||||
* live roster (to resolve a target to a profile) and {@code errorPatternsByProfile} (each
|
||||
* profile's configured {@code errorPattern}, already compiled by the caller — the same map also
|
||||
* feeds the coverage log next to where this is called) — nothing else, so it is directly
|
||||
* testable. See {@link #worktreeBranchLookup} above for why this ticket exists and why the
|
||||
* factory takes a roster {@link Supplier} rather than a {@link SessionManager}.
|
||||
*
|
||||
* @param roster the live member roster, normally {@code sessions::roster}
|
||||
* @param errorPatternsByProfile every profile that has an {@code errorPattern} configured,
|
||||
* keyed by profile name
|
||||
*/
|
||||
static BackendErrorPatternLookup backendErrorPatternLookup(Supplier<List<MemberSession>> roster,
|
||||
Map<String, Pattern> errorPatternsByProfile) {
|
||||
return target -> roster.get().stream()
|
||||
.filter(session -> target.equals(session.terminalId()))
|
||||
.findFirst()
|
||||
.map(session -> errorPatternsByProfile.get(session.profile()))
|
||||
.orElse(null);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #248 / fleetd#201 Unit 5: factory for the production {@link BackendErrorSink} — the
|
||||
* collaborator {@link CompletionResolver} notifies when a pane-scrape classification actually
|
||||
* resolves a waiter as a backend error. Order: (1) mark the member BACKEND_ERROR; (2) resolve
|
||||
* profile/credential through the roster — fail loud (never {@code Optional.ifPresent}, the
|
||||
* fleetd #234 lesson applied to this sink) and notify the lead via {@code
|
||||
* onBackendTargetUnmapped} when it cannot be resolved; (3) record the error in {@code
|
||||
* outagePolicy}; (4) on a NEW incident (the {@code record()} call that actually crosses the
|
||||
* threshold), tell the lead via {@code onBackendIncident}.
|
||||
*
|
||||
* <p>{@code public}, unlike {@link #worktreeBranchLookup} and {@link #backendErrorPatternLookup}
|
||||
* above: {@code dev.ltms.fleet.inject.BackendOutageFlowTest} exercises this exact object as the
|
||||
* real, wired production path, replacing what its own class doc used to call out as a
|
||||
* hand-mirrored copy of this lambda ("mirrors {@code Fleetd.main}'s {@code backendErrorSink}
|
||||
* lambda line-for-line") — that copy proved only itself, never that {@code main} still wires the
|
||||
* real thing. That was precisely the gap fleetd #248 exists to close.
|
||||
*
|
||||
* @param sessions the session registry; both read (roster) and written (onBackendError)
|
||||
* @param profiles the live profile map, normally {@code () -> config.get().profiles()} in
|
||||
* {@code main}, or a fixed test map via {@code () -> profiles}
|
||||
* @param pushLoop the lead-nudge loop, read lazily: {@code main} builds this sink before the
|
||||
* real {@link ReplyPushLoop} exists (a genuine construction-order cycle, broken
|
||||
* the same way {@code exhaustionSinkRef} is a few lines above it), so a
|
||||
* {@link Supplier} reads whatever {@code main} has filled in by the time a real
|
||||
* backend error fires
|
||||
*/
|
||||
public static BackendErrorSink backendErrorSink(SessionManager sessions,
|
||||
Supplier<Map<String, FleetConfig.Profile>> profiles, BackendOutagePolicy outagePolicy,
|
||||
Supplier<ReplyPushLoop> pushLoop) {
|
||||
return (target, matchedLine, reason) -> {
|
||||
sessions.onBackendError(target, reason);
|
||||
|
||||
String profileName = sessions.roster().stream()
|
||||
.filter(session -> target.equals(session.terminalId()))
|
||||
.findFirst()
|
||||
.map(MemberSession::profile)
|
||||
.orElse(null);
|
||||
FleetConfig.Profile profile = profileName == null ? null : profiles.get().get(profileName);
|
||||
if (profile == null) {
|
||||
log.error("backend error on target '{}' ({}) but no profile could be resolved — the "
|
||||
+ "target is not (yet) in the roster — no cool-off applied (fleetd #201 Unit 5)",
|
||||
target, reason);
|
||||
ReplyPushLoop loop = pushLoop.get();
|
||||
if (loop != null) {
|
||||
loop.onBackendTargetUnmapped(target, reason);
|
||||
}
|
||||
return;
|
||||
}
|
||||
String credentialId = profile.effectiveCredentialId();
|
||||
Optional<BackendOutagePolicy.Incident> incident = outagePolicy.record(credentialId, target, reason);
|
||||
incident.ifPresent(inc -> {
|
||||
List<String> affectedProfiles = profiles.get().values().stream()
|
||||
.filter(p -> credentialId.equals(p.effectiveCredentialId()))
|
||||
.map(FleetConfig.Profile::profile)
|
||||
.sorted()
|
||||
.toList();
|
||||
log.warn("credential '{}' cooling off for {}s after backend errors on {} distinct "
|
||||
+ "target(s) (profile '{}'): {}", credentialId,
|
||||
inc.remainingCoolOffSeconds(), inc.evidenceCount(), profile.profile(), reason);
|
||||
ReplyPushLoop loop = pushLoop.get();
|
||||
if (loop != null) {
|
||||
loop.onBackendIncident(inc.id(), inc.targets(), credentialId, affectedProfiles,
|
||||
(int) inc.remainingCoolOffSeconds());
|
||||
}
|
||||
});
|
||||
};
|
||||
}
|
||||
|
||||
/** Injection seam for {@link #selectReplyInbox}: production binds {@link AmqpReplyInbox#open}. */
|
||||
@FunctionalInterface
|
||||
interface AmqpOpener {
|
||||
@@ -795,7 +1070,7 @@ public final class Fleetd {
|
||||
|
||||
/**
|
||||
* CB-594: which env vars the loaded config actually needs, and why — every non-{@code
|
||||
* subscription} profile's {@code tokenEnv} (a subscription profile never reads one, see
|
||||
* subscription} profile's explicitly configured {@code tokenEnv} (a subscription profile never reads one, see
|
||||
* {@link FleetConfig.Profile#isSubscription()}), plus every profile's {@code gitTokenEnv}
|
||||
* where set (opt-in), plus a configured {@code broker.uriEnv} (CB-151). Derived from the
|
||||
* config, not hard-coded, so a new profile is covered for free. A var required by more than one
|
||||
@@ -812,7 +1087,9 @@ public final class Fleetd {
|
||||
static Map<String, List<String>> requiredSecretEnvVars(FleetConfig cfg) {
|
||||
Map<String, List<String>> requiredBy = new LinkedHashMap<>();
|
||||
cfg.profiles().forEach((name, profile) -> {
|
||||
if (!profile.isSubscription()) {
|
||||
// All kinds are checked when tokenEnv is explicit. OpenCode may use provider credentials,
|
||||
// but an explicit tokenEnv still declares a required host secret for its configured provider.
|
||||
if (!profile.isSubscription() && profile.hasTokenEnv()) {
|
||||
requiredBy.computeIfAbsent(profile.tokenEnv(), _ -> new ArrayList<>())
|
||||
.add("profile '" + name + "' tokenEnv");
|
||||
}
|
||||
@@ -838,7 +1115,7 @@ public final class Fleetd {
|
||||
* <p>A missing entry only warns — it must never refuse to start. A daemon that boots and says
|
||||
* what is wrong is strictly more useful than one that will not boot at all.
|
||||
*/
|
||||
private static void reportRequiredSecrets(FleetConfig cfg) {
|
||||
static void reportRequiredSecrets(FleetConfig cfg) {
|
||||
Map<String, List<String>> requiredBy = requiredSecretEnvVars(cfg);
|
||||
if (requiredBy.isEmpty()) {
|
||||
log.info("startup secrets: no profile references a token env var — nothing to check");
|
||||
@@ -872,12 +1149,14 @@ public final class Fleetd {
|
||||
* #requiredSecretEnvVars} is exposed for {@link #reportRequiredSecrets}'s own test.
|
||||
*/
|
||||
static void reportMemberCredentialsGap(FleetConfig cfg) {
|
||||
FleetConfig.MemberCredentials creds = cfg.memberCredentials();
|
||||
if (creds != null && !creds.known().isEmpty()) {
|
||||
// fleetd #111: the counts below come from MemberCredentialPolicyView, the same class the
|
||||
// live GET /member-credentials endpoint reads — one place computes them, not two.
|
||||
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(cfg.memberCredentials());
|
||||
if (view.present()) {
|
||||
log.info("memberCredentials: policy={}, {} known name(s), {} allowed — blocking {} on "
|
||||
+ "every spawn{}",
|
||||
creds.policy(), creds.known().size(), creds.allow().size(), creds.blockedSet().size(),
|
||||
creds.isAllowList()
|
||||
view.policy(), view.knownCount(), view.allowedCount(), view.blockedCount(),
|
||||
cfg.memberCredentials().isAllowList()
|
||||
? " (allow-list: known/allow are reporting only — the control is the derived ZDOTDIR scrub)"
|
||||
: "");
|
||||
return;
|
||||
|
||||
@@ -43,7 +43,9 @@ import java.util.function.Supplier;
|
||||
* which is constructed once), <em>and an existing profile's launch settings</em> —
|
||||
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl},
|
||||
* {@code exhaustedPattern} (CB-578 stage A — compiled once into {@code Fleetd.main}'s
|
||||
* pattern map at startup), and the rest. {@code credentialId} (CB-578 stage B) is NOT on
|
||||
* pattern map at startup), {@code errorPattern} (fleetd #201 Unit 5 — compiled once into
|
||||
* {@code Fleetd.main}'s backend-error pattern map at startup, the same way), and the rest.
|
||||
* {@code credentialId} (CB-578 stage B) is NOT on
|
||||
* this list — it is read live off the config supplier at every quarantine check and
|
||||
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so it is hot instead.
|
||||
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} at construction and resolves
|
||||
@@ -292,6 +294,10 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
// CB-578 stage B: exhaustedPattern is compiled once into Fleetd.main's pattern map
|
||||
// at startup (see ExhaustedPatternLookup wiring) — a reload never re-reads it, so a
|
||||
// changed pattern must be reported as deferred, exactly like model/baseUrl/argv.
|
||||
&& Objects.equals(a.exhaustedPattern(), b.exhaustedPattern());
|
||||
&& Objects.equals(a.exhaustedPattern(), b.exhaustedPattern())
|
||||
// fleetd #201 Unit 5: errorPattern is compiled once into Fleetd.main's backend-error
|
||||
// pattern map at startup (see BackendErrorPatternLookup wiring), the same way
|
||||
// exhaustedPattern is — a reload never re-reads it either.
|
||||
&& Objects.equals(a.errorPattern(), b.errorPattern());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -24,6 +24,8 @@ import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.regex.Pattern;
|
||||
import java.util.regex.PatternSyntaxException;
|
||||
|
||||
/**
|
||||
* {@code fleetd} configuration, loaded from a YAML file (see
|
||||
@@ -90,9 +92,15 @@ import java.util.Set;
|
||||
* fleetd's own process, so fleetd's own {@code $SHELL} says nothing about what
|
||||
* that pane runs. There is no channel to ask herdr for another user's shell, so
|
||||
* this must be told, never guessed. {@code null}/blank (or a value not ending
|
||||
* in {@code zsh}) is treated the same as "not zsh": the {@code
|
||||
* memberCredentials.policy: allow-list} ZDOTDIR scrub is skipped in favour of
|
||||
* the CB-596 sentinel overlay — a degraded control, never a refusal to spawn.
|
||||
* in {@code zsh}) is treated the same as "not zsh". fleetd #155: what that means
|
||||
* now depends on {@code memberCredentials.policy}. Under {@code deny-by-default}
|
||||
* it stays a degraded control, never a refusal to spawn — the pane-creation
|
||||
* overlay is unaffected by shell type, so the spawn proceeds with a WARN naming
|
||||
* the shell. Under {@code allow-list} the spawn is REFUSED instead: that policy's
|
||||
* whole point is a control a sourced file cannot undo, so silently falling back
|
||||
* to the weaker overlay would be the same "control silently does nothing" defect
|
||||
* #155 exists to remove — configure this field (or move the account to zsh, or
|
||||
* switch policy back to {@code deny-by-default}) to unblock the spawn.
|
||||
* When {@code memberHerdrSocket} is NOT configured this field is never
|
||||
* consulted at all; fleetd keeps reading its own {@code $SHELL}, exactly as
|
||||
* before this field existed.
|
||||
@@ -240,7 +248,8 @@ public record FleetConfig(
|
||||
* @param configDir {@code CLAUDE_CONFIG_DIR} so the worker inherits the profile's
|
||||
* skills/MCP/hooks (may be {@code null})
|
||||
* @param tokenEnv name of the host env var holding the worker's auth token; its value
|
||||
* is injected as {@code ANTHROPIC_AUTH_TOKEN} (never stored in config)
|
||||
* is injected as {@code ANTHROPIC_AUTH_TOKEN} (never stored in config).
|
||||
* {@code null}/blank means this profile needs no token.
|
||||
* @param argv launch command; defaults to {@code ["claude"]}
|
||||
* @param placement where a worker lands: {@code "tab"} (default — its own tab in the
|
||||
* worker space) or {@code "pane"} (legacy — split the focused tab)
|
||||
@@ -256,7 +265,13 @@ public record FleetConfig(
|
||||
* @param cwd fixed working directory for this profile's workers (CB-112 "told otherwise");
|
||||
* {@code null}/blank → inherit the primary's cwd, else the daemon's
|
||||
* @param parityOverlay repo-relative paths copied primary→worktree for config parity; null/empty
|
||||
* defaults to a sensible set of local config files.
|
||||
* defaults to {@code [.env]} only (CB-148 point 2). {@code .env} is data — a
|
||||
* copy of it can only carry values. {@code .envrc} is executable shell that
|
||||
* {@code direnv} evaluates on every {@code cd}, so copying it moves behaviour
|
||||
* into the worker, not just values, and that is a different risk from copying
|
||||
* data. It is deliberately left out of the default for that reason. The knob
|
||||
* still supports it: an operator who wants it copied writes
|
||||
* {@code parityOverlay: [.env, .envrc]} explicitly and owns that choice.
|
||||
* <p><strong>Every overlay path must be gitignored or tracked-and-skipped.</strong>
|
||||
* CB-576 made {@code release()} preserve a worktree that {@code git status
|
||||
* --porcelain} reports as dirty, and it deliberately counts untracked files —
|
||||
@@ -266,11 +281,10 @@ public record FleetConfig(
|
||||
* worktrees then accumulate with no error anywhere.
|
||||
* <p>Checked on 2026-08-15 (CB-581): inert as configured. Tracked overlay
|
||||
* files carry {@code --skip-worktree} so {@code --porcelain} cannot see them,
|
||||
* {@code fleetd.yaml} is gitignored, and the default pair {@code .env} /
|
||||
* {@code .envrc} does not exist in this repo. Note the default applies to
|
||||
* <em>every</em> profile, so creating either file at the repo root is enough
|
||||
* to make it live. Add a new overlay path to {@code .gitignore} in the same
|
||||
* change that adds it here.
|
||||
* {@code fleetd.yaml} is gitignored, and {@code .env} does not exist in this
|
||||
* repo. Note the default applies to <em>every</em> profile, so creating
|
||||
* {@code .env} at the repo root is enough to make it live. Add a new overlay
|
||||
* path to {@code .gitignore} in the same change that adds it here.
|
||||
* <p>CB-578 stage C raised the stakes: a preserved dirty worktree is now also
|
||||
* committed to {@code refs/wip/<branch>} via {@code git add -A}. The overlay
|
||||
* carries the primary's own environment files into the worktree, so a
|
||||
@@ -355,6 +369,16 @@ public record FleetConfig(
|
||||
* {@code limit.context} instead, which bounds the window opencode compacts
|
||||
* <em>within</em>, and only when {@code model} resolves to a
|
||||
* {@code provider/model} pair.
|
||||
* @param errorPattern regex matched against a completion-fallback scrape (fleetd #201 Unit 5)
|
||||
* to classify a turn that ended with no {@code fleet_reply} as a backend
|
||||
* error (credential outage, provider 5xx) rather than a real answer or an
|
||||
* exhausted usage limit. {@code null}/blank ⇒ this profile relies on
|
||||
* {@link dev.ltms.fleet.inject.CompletionResolver}'s built-in
|
||||
* {@code (?i)\bAPI Error\s*:} compatibility pattern instead — classification
|
||||
* still happens, just without a profile-specific match (reported separately
|
||||
* at startup as legacy-default coverage, never full coverage). Compiled once
|
||||
* at startup ({@code Fleetd.main}), like {@code exhaustedPattern}, so a
|
||||
* reload of this key is DEFERRED (see {@code ConfigRef}).
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record Profile(String profile, String baseUrl, String model,
|
||||
@@ -373,7 +397,8 @@ public record FleetConfig(
|
||||
String ideMcpUrl,
|
||||
String ideProjectDir,
|
||||
String ideOpenCommand,
|
||||
Integer autoCompactWindow) {
|
||||
Integer autoCompactWindow,
|
||||
String errorPattern) {
|
||||
|
||||
/** Peer kind spawned by {@link dev.ltms.fleet.member.ClaudeCodeLauncher} (the default). */
|
||||
public static final String KIND_CLAUDE_CODE = "claude-code";
|
||||
@@ -389,7 +414,7 @@ public record FleetConfig(
|
||||
? (KIND_CLAUDE_CODE.equals(k) ? List.of("claude") : List.of(k))
|
||||
: List.copyOf(argv);
|
||||
kind = k;
|
||||
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "FLEETD_WORKER_TOKEN" : tokenEnv;
|
||||
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? null : tokenEnv;
|
||||
placement = (placement == null || placement.isBlank()) ? "tab" : placement.toLowerCase();
|
||||
workspace = (workspace == null || workspace.isBlank()) ? "fleet" : workspace;
|
||||
// CB-557: no per-profile default any more. A label is generated from the member's ROLE
|
||||
@@ -407,7 +432,7 @@ public record FleetConfig(
|
||||
// servers do not exist there to be enabled. This is defence in depth, not a live fix: the
|
||||
// worker stays isolated only because this separate mechanism already removes the servers.
|
||||
parityOverlay = (parityOverlay == null || parityOverlay.isEmpty())
|
||||
? List.of(".env", ".envrc")
|
||||
? List.of(".env")
|
||||
: List.copyOf(parityOverlay);
|
||||
// gitTokenEnv stays null when unset (opt-in). gitHostEnv defaults so operators enabling
|
||||
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
|
||||
@@ -447,6 +472,32 @@ public record FleetConfig(
|
||||
// there is no blank-string form to normalize (it's an Integer), and the [100000, 1000000]
|
||||
// range is enforced eagerly at config load (rejectAutoCompactWindowOutOfRange), naming the
|
||||
// profile, rather than silently clamped here. A profile that never sets it keeps null.
|
||||
// errorPattern stays null when unset/blank (opt-in) — same rule as exhaustedPattern: no
|
||||
// defaulting, no vendor wording. A profile that never sets it relies on
|
||||
// CompletionResolver's built-in compatibility pattern instead (never "off").
|
||||
errorPattern = (errorPattern == null || errorPattern.isBlank()) ? null : errorPattern;
|
||||
}
|
||||
|
||||
/**
|
||||
* Backward-compatible constructor without the fleetd #201 {@code errorPattern} field — the
|
||||
* profile relies on {@code CompletionResolver}'s built-in compatibility pattern instead
|
||||
* (legacy-default coverage). This is the shape the canonical constructor had before the
|
||||
* field was added; every pre-existing Java call site (and any YAML that omits the key) keeps
|
||||
* compiling and behaving identically. Jackson binds the canonical (longest) constructor, so
|
||||
* YAML omitting {@code errorPattern:} still lands here as {@code null} via that path too.
|
||||
*/
|
||||
public Profile(String profile, String baseUrl, String model,
|
||||
String configDir, String tokenEnv, List<String> argv,
|
||||
String placement, String workspace, String tabLabel, String mcpUrl,
|
||||
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
|
||||
String kind, Map<String, String> env, Float weight, Integer maxLoad,
|
||||
Boolean subscription, String exhaustedPattern, String credentialId,
|
||||
String ideMcpUrl, String ideProjectDir, String ideOpenCommand,
|
||||
Integer autoCompactWindow) {
|
||||
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
|
||||
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad,
|
||||
subscription, exhaustedPattern, credentialId, ideMcpUrl, ideProjectDir, ideOpenCommand,
|
||||
autoCompactWindow, null);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -491,7 +542,8 @@ public record FleetConfig(
|
||||
public Profile withProfile(String p) {
|
||||
return new Profile(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
|
||||
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, subscription,
|
||||
exhaustedPattern, credentialId, ideMcpUrl, ideProjectDir, ideOpenCommand, autoCompactWindow);
|
||||
exhaustedPattern, credentialId, ideMcpUrl, ideProjectDir, ideOpenCommand, autoCompactWindow,
|
||||
errorPattern);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -589,14 +641,55 @@ public record FleetConfig(
|
||||
}
|
||||
|
||||
/**
|
||||
* The credential group this profile quarantines with (CB-578 stage B): the configured
|
||||
* {@link #credentialId} when set, else this profile's own name — so an unconfigured profile
|
||||
* quarantines alone, exactly as it did before this field existed. Two profiles that set the
|
||||
* same non-blank {@code credentialId} share one quarantine: a {@code BACKEND_EXHAUSTED}
|
||||
* classification on either one quarantines both.
|
||||
* True when this profile configures its own fleetd #201 Unit 5 backend-error classification
|
||||
* pattern. {@code false} means this profile relies on {@code CompletionResolver}'s built-in
|
||||
* {@code (?i)\bAPI Error\s*:} compatibility pattern instead (legacy-default coverage).
|
||||
*/
|
||||
public boolean hasErrorPattern() {
|
||||
return errorPattern != null;
|
||||
}
|
||||
|
||||
/**
|
||||
* The shared credential id every {@code subscription: true} profile falls back to when it
|
||||
* sets no explicit {@link #credentialId} (fleetd #176 stage 2, correcting an inert first cut
|
||||
* of that ticket). A subscription profile has no credential of its own to fall back to its
|
||||
* name for: it authenticates as the operator's own Claude login, and a host has exactly one
|
||||
* of those, whatever names the operator gives the profiles running on it. Falling back to the
|
||||
* profile's own name (the way an ordinary off-subscription profile does) would keep two
|
||||
* subscription profiles on one login apart from each other, which is the opposite of what
|
||||
* "one account" means.
|
||||
*
|
||||
* <p>Measured live and what it broke: a lead on profile {@code opus}, members on profile
|
||||
* {@code sonnet}, same Claude login, neither setting {@code credentialId}. Before this
|
||||
* sentinel, {@code opus.effectiveCredentialId()} was {@code "opus"} and {@code sonnet
|
||||
* .effectiveCredentialId()} was {@code "sonnet"} — so fleetd #176's lead-seat matcher (and,
|
||||
* this sentinel now also fixes, {@code CompositePeerLauncher}'s quarantine/cool-off spawn
|
||||
* refusal and {@code BackendOutagePolicy}'s incident grouping) silently never linked them: the
|
||||
* fix shipped, and stayed inert on the one host it was written for.
|
||||
*/
|
||||
public static final String SUBSCRIPTION_CREDENTIAL_ID = "<subscription>";
|
||||
|
||||
/**
|
||||
* The credential group this profile quarantines with (CB-578 stage B; extended fleetd #176
|
||||
* stage 2 — see {@link #SUBSCRIPTION_CREDENTIAL_ID}): the configured {@link #credentialId}
|
||||
* when set — that always wins, so an operator with two separate Claude logins on one host can
|
||||
* still keep them apart. Otherwise, a {@code subscription: true} profile falls back to
|
||||
* {@link #SUBSCRIPTION_CREDENTIAL_ID} rather than its own name; an ordinary off-subscription
|
||||
* profile falls back to its own name, exactly as it did before this field existed, so an
|
||||
* unconfigured off-subscription profile still quarantines alone.
|
||||
*
|
||||
* <p>A {@code BACKEND_EXHAUSTED} (or repeated backend-error) classification on any profile
|
||||
* sharing the result quarantines/cools off every profile that shares it — including, now,
|
||||
* every {@code subscription: true} profile with no explicit {@code credentialId}. That is
|
||||
* intended, not incidental: one Claude subscription hitting a usage limit really does take out
|
||||
* every profile running on it, the same way {@code credentialId: openai-shared} already lets
|
||||
* two OpenAI-backed profiles share one quarantine.
|
||||
*/
|
||||
public String effectiveCredentialId() {
|
||||
return (credentialId == null || credentialId.isBlank()) ? profile : credentialId;
|
||||
if (credentialId != null && !credentialId.isBlank()) {
|
||||
return credentialId;
|
||||
}
|
||||
return isSubscription() ? SUBSCRIPTION_CREDENTIAL_ID : profile;
|
||||
}
|
||||
|
||||
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
|
||||
@@ -604,6 +697,11 @@ public record FleetConfig(
|
||||
return gitTokenEnv != null && !gitTokenEnv.isBlank();
|
||||
}
|
||||
|
||||
/** True when this profile explicitly names a host token environment variable. */
|
||||
public boolean hasTokenEnv() {
|
||||
return tokenEnv != null && !tokenEnv.isBlank();
|
||||
}
|
||||
|
||||
/**
|
||||
* True when this profile's {@code env:} block names a worker-side Anthropic binding variable.
|
||||
* Those two keys are the adapter's, never the operator's: on a claude-code worker
|
||||
@@ -886,7 +984,15 @@ public record FleetConfig(
|
||||
* survives restarts of the agent inside it — so identity is now the tab label alone.
|
||||
*
|
||||
* @param profile the {@code profiles:} entry to launch this lead on when one must
|
||||
* be created; {@code null} ⇒ recognise-only, never create
|
||||
* be created; {@code null} ⇒ recognise-only, never create.
|
||||
* <p>fleetd #176: also the field {@code Fleetd.leadSeatLookup} reads
|
||||
* to learn which account this lead's own live session shares — set it
|
||||
* (safely, even on an already-running recognise-only lead: naming a
|
||||
* profile here never starts anything beyond {@code instances}) so a
|
||||
* {@code subscription: true} worker profile sharing its
|
||||
* {@code effectiveCredentialId()} has this lead's seat subtracted from
|
||||
* {@code fleet_list}'s {@code free}. {@code null} here also means this
|
||||
* lead's seat cannot be derived and is not counted.
|
||||
* @param tab the exact tab label hosting this lead, matched case-insensitively;
|
||||
* the only field identity depends on. Required — a lead with no
|
||||
* {@code tab} can never be discovered, launched or not
|
||||
@@ -1383,6 +1489,7 @@ public record FleetConfig(
|
||||
rejectDuplicateMemberSlots(yaml);
|
||||
rejectNegativeMaxLoad(yaml);
|
||||
rejectAutoCompactWindowOutOfRange(yaml);
|
||||
rejectMalformedErrorPattern(yaml);
|
||||
rejectUnknownKind(yaml);
|
||||
rejectUnknownAuthMode(yaml);
|
||||
rejectUnknownPlacement(yaml);
|
||||
@@ -1731,6 +1838,51 @@ public record FleetConfig(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Reject a profile whose {@code errorPattern} (fleetd #201 Unit 5) is not a valid Java regex,
|
||||
* naming the profile, the key, and the parser's own message.
|
||||
*
|
||||
* <p>Unset/{@code null} means "use {@code CompletionResolver}'s built-in {@code (?i)\bAPI
|
||||
* Error\s*:} compatibility pattern" and passes silently. A profile that DOES set the key gets it
|
||||
* compiled once at daemon startup ({@code Fleetd.main}, mirroring {@code exhaustedPattern}) — an
|
||||
* uncaught {@link java.util.regex.PatternSyntaxException} there crashes startup without naming
|
||||
* which profile or key is at fault. Validate eagerly here instead, at config load, the same
|
||||
* "fail loud at load, not lazily later" reasoning as {@link #rejectAutoCompactWindowOutOfRange}.
|
||||
*
|
||||
* @param yaml the raw config text
|
||||
* @throws IllegalStateException when any profile's {@code errorPattern} fails to compile
|
||||
*/
|
||||
static void rejectMalformedErrorPattern(String yaml) {
|
||||
Map<?, ?> raw;
|
||||
try {
|
||||
raw = YAML.readValue(yaml, Map.class);
|
||||
} catch (IOException | IllegalArgumentException e) {
|
||||
return; // a malformed file is reported by the real parse, not here
|
||||
}
|
||||
if (raw == null || !(raw.get("profiles") instanceof Map<?, ?> profiles)) {
|
||||
return;
|
||||
}
|
||||
List<String> bad = new ArrayList<>();
|
||||
for (Map.Entry<?, ?> e : profiles.entrySet()) {
|
||||
if (!(e.getValue() instanceof Map<?, ?> p)) {
|
||||
continue;
|
||||
}
|
||||
if (!(p.get("errorPattern") instanceof String pattern) || pattern.isBlank()) {
|
||||
continue;
|
||||
}
|
||||
try {
|
||||
Pattern.compile(pattern);
|
||||
} catch (PatternSyntaxException ex) {
|
||||
bad.add("profiles." + e.getKey() + ".errorPattern (\"" + pattern + "\"): " + ex.getMessage());
|
||||
}
|
||||
}
|
||||
bad.sort(String::compareTo);
|
||||
if (!bad.isEmpty()) {
|
||||
throw new IllegalStateException("refusing to start: malformed errorPattern — "
|
||||
+ String.join("; ", bad));
|
||||
}
|
||||
}
|
||||
|
||||
/** The peer kinds this build has an adapter for — {@link Profile#kind()}'s only valid values. */
|
||||
private static final Set<String> KNOWN_KINDS = Set.of(Profile.KIND_CLAUDE_CODE, Profile.KIND_OPENCODE);
|
||||
|
||||
|
||||
@@ -0,0 +1,35 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
/**
|
||||
* Per-target lookup for a profile's configured backend-error pattern (fleetd#201 / #227): how
|
||||
* {@link CompletionResolver} recognizes a backend that failed a turn outright (a credential
|
||||
* outage, a provider 5xx) from whatever ends up on the member's pane, apart from a genuine
|
||||
* completion.
|
||||
*
|
||||
* <p>The pattern is always profile config, never a vendor string in Java source: every backend
|
||||
* words its failure differently, so a hardcoded sentence would only ever match one of them. A
|
||||
* target with no configured pattern is not "off" the way {@link ExhaustedPatternLookup#none()}
|
||||
* is — {@link CompletionResolver} falls back to its narrow {@code (?i)\bAPI Error\s*:}
|
||||
* compatibility pattern instead, so classification still happens, just without a profile-specific
|
||||
* match. {@link #legacy()} is the explicit stand-in existing (pre-fleetd#201) callers pass to get
|
||||
* exactly that fallback-only behavior until a caller wires a real, profile-driven lookup
|
||||
* (fleetd#201 Unit 5).
|
||||
*/
|
||||
@FunctionalInterface
|
||||
public interface BackendErrorPatternLookup {
|
||||
|
||||
/** The compiled pattern configured for {@code target}'s profile, or {@code null} if none. */
|
||||
Pattern patternFor(String target);
|
||||
|
||||
/**
|
||||
* Legacy lookup — no profile has a configured pattern, so {@link CompletionResolver} classifies
|
||||
* every target using only its built-in {@code (?i)\bAPI Error\s*:} compatibility pattern. The
|
||||
* explicit stand-in existing constructors pass so their behavior is unchanged until a caller
|
||||
* wires a real, profile-driven lookup.
|
||||
*/
|
||||
static BackendErrorPatternLookup legacy() {
|
||||
return target -> null;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,35 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
/**
|
||||
* Notified when {@link CompletionResolver} actually delivers a typed backend-error classification
|
||||
* to a waiting send (fleetd#201 / #227) — never on a race that lost. {@link CompletionResolver}
|
||||
* calls this only after {@code Rendezvous.resolveFailure} returns {@code true} for that exact
|
||||
* waiter, mirroring the win-only race rule {@link ExhaustionSink} already uses.
|
||||
*
|
||||
* <p>The public send result is unchanged by this classification — it is still a failed send
|
||||
* ({@code Rendezvous.Kind#FAILED}); this sink is the internal seam a later stage (fleetd#201 Unit
|
||||
* 5, the policy that cools a credential off after two such failures in 60 seconds and tells the
|
||||
* lead) consumes. {@link CompletionResolver} knows only {@code target} (a herdr terminal id); it
|
||||
* has no notion of profiles or credentials, so mapping {@code target} to whatever should be
|
||||
* quarantined is entirely the sink's job — exactly like {@link ExhaustionSink}.
|
||||
*/
|
||||
@FunctionalInterface
|
||||
public interface BackendErrorSink {
|
||||
|
||||
/**
|
||||
* @param target the herdr terminal id whose turn was classified as a backend error
|
||||
* @param matchedLine the single pane line that matched the backend-error pattern
|
||||
* @param reason the full failure reason carried by the classification (the matched line
|
||||
* plus whatever pane context the classifying path attaches)
|
||||
*/
|
||||
void onBackendError(String target, String matchedLine, String reason);
|
||||
|
||||
/**
|
||||
* Inert sink — nothing happens on a backend-error classification. The explicit stand-in a
|
||||
* caller (or a test not exercising this feature) passes instead of a defaulting overload,
|
||||
* exactly like {@link ExhaustionSink#none()}.
|
||||
*/
|
||||
static BackendErrorSink none() {
|
||||
return (target, matchedLine, reason) -> { };
|
||||
}
|
||||
}
|
||||
@@ -14,6 +14,7 @@ import java.util.TreeSet;
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.function.Function;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
/**
|
||||
@@ -85,11 +86,29 @@ public final class CompletionResolver implements TurnListener {
|
||||
private static final String CLIPPED_PANE_TAIL_MARKER =
|
||||
"[Pane tail clipped: member did not call fleet_reply.]";
|
||||
|
||||
/** A pane echo must be this large before it can replace a completion report. */
|
||||
static final int ECHO_MIN_CHARS = 400;
|
||||
|
||||
/** Normalised TUI chrome may add this many characters to an otherwise echoed brief. */
|
||||
static final int MAX_ECHO_EXCESS_CHARS = 160;
|
||||
|
||||
/** The explicit result returned instead of a lead's echoed injected brief. */
|
||||
public static final String NO_REPORT_PREFIX = "[no report — the member ended its turn without fleet_reply, "
|
||||
+ "and the pane still shows the injected brief. Nothing was produced on the pane. Check the "
|
||||
+ "member's worktree and branch for committed work before re-delegating.";
|
||||
|
||||
private final AgentControl agents;
|
||||
private final Rendezvous rendezvous;
|
||||
private final ExhaustedPatternLookup exhaustedPatterns;
|
||||
private final ExhaustionSink exhaustionSink;
|
||||
private final BackendErrorPatternLookup backendErrorPatterns;
|
||||
private final BackendErrorSink backendErrorSink;
|
||||
private final LongSupplier nowNanos;
|
||||
private final Function<String, WorktreeBranch> worktreeBranches;
|
||||
|
||||
/** Known member location, used only to guide a lead after an echoed brief. */
|
||||
public record WorktreeBranch(String worktree, String branch) {
|
||||
}
|
||||
|
||||
/**
|
||||
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
|
||||
@@ -106,7 +125,8 @@ public final class CompletionResolver implements TurnListener {
|
||||
* the other half of the {@link #MIN_TURN_NANOS} floor check, compared against a fresh reading at
|
||||
* resolution time.
|
||||
*/
|
||||
record InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline, long deliveredAtNanos) {
|
||||
record InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline, long deliveredAtNanos,
|
||||
String injectedText) {
|
||||
|
||||
/**
|
||||
* Convenience for tests exercising scrape/suppression logic that don't care about turn
|
||||
@@ -114,7 +134,11 @@ public final class CompletionResolver implements TurnListener {
|
||||
* Not used by production code — {@link #captureBaseline} always records a real reading.
|
||||
*/
|
||||
InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline) {
|
||||
this(waiter, baseline, Long.MIN_VALUE / 2);
|
||||
this(waiter, baseline, Long.MIN_VALUE / 2, null);
|
||||
}
|
||||
|
||||
InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline, long deliveredAtNanos) {
|
||||
this(waiter, baseline, deliveredAtNanos, null);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -129,10 +153,55 @@ public final class CompletionResolver implements TurnListener {
|
||||
* classification actually resolves a waiter. Required for the same
|
||||
* reason as {@code exhaustedPatterns} — pass {@link ExhaustionSink#none()}
|
||||
* to opt out.
|
||||
*
|
||||
* <p>Transition constructor (fleetd#201 Unit 1): delegates to the full constructor below with
|
||||
* {@link BackendErrorPatternLookup#legacy()} and {@link BackendErrorSink#none()}, so every
|
||||
* existing caller keeps today's behavior — the narrow built-in {@code API Error:} match, no
|
||||
* sink notified — until a caller wires a real, profile-driven backend-error lookup and sink
|
||||
* (fleetd#201 Unit 5).
|
||||
*/
|
||||
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
|
||||
ExhaustionSink exhaustionSink) {
|
||||
this(agents, rendezvous, exhaustedPatterns, exhaustionSink, System::nanoTime);
|
||||
ExhaustionSink exhaustionSink) {
|
||||
this(agents, rendezvous, exhaustedPatterns, exhaustionSink,
|
||||
BackendErrorPatternLookup.legacy(), BackendErrorSink.none(), System::nanoTime);
|
||||
}
|
||||
|
||||
/** Production constructor with a lookup for the member worktree and branch. */
|
||||
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
|
||||
ExhaustionSink exhaustionSink, Function<String, WorktreeBranch> worktreeBranches) {
|
||||
this(agents, rendezvous, exhaustedPatterns, exhaustionSink,
|
||||
BackendErrorPatternLookup.legacy(), BackendErrorSink.none(), System::nanoTime, worktreeBranches);
|
||||
}
|
||||
|
||||
/**
|
||||
* Transition constructor (fleetd#201 Unit 1): same legacy backend-error defaults as the 4-arg
|
||||
* constructor above, but with the injectable clock. Kept so existing fleetd#164 timing tests
|
||||
* (and any other caller that wants a controllable clock but not the backend-error feature)
|
||||
* keep compiling unchanged.
|
||||
*/
|
||||
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
|
||||
ExhaustionSink exhaustionSink, LongSupplier nowNanos) {
|
||||
this(agents, rendezvous, exhaustedPatterns, exhaustionSink,
|
||||
BackendErrorPatternLookup.legacy(), BackendErrorSink.none(), nowNanos);
|
||||
}
|
||||
|
||||
/**
|
||||
* Production constructor (fleetd#201 Unit 1): adds the target-keyed backend-error pattern
|
||||
* lookup and sink alongside the existing exhaustion pair. Required, like {@code exhaustedPatterns}
|
||||
* and {@code exhaustionSink} — pass {@link BackendErrorPatternLookup#legacy()} and
|
||||
* {@link BackendErrorSink#none()} to opt out.
|
||||
*
|
||||
* @param backendErrorPatterns fleetd#201: per-target lookup for a profile's configured
|
||||
* backend-error pattern; a target with none configured is matched
|
||||
* against the built-in compatibility pattern instead (never "off").
|
||||
* @param backendErrorSink fleetd#201: notified when a backend-error classification actually
|
||||
* resolves a waiter — never on a race that lost.
|
||||
*/
|
||||
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
|
||||
ExhaustionSink exhaustionSink, BackendErrorPatternLookup backendErrorPatterns,
|
||||
BackendErrorSink backendErrorSink) {
|
||||
this(agents, rendezvous, exhaustedPatterns, exhaustionSink, backendErrorPatterns, backendErrorSink,
|
||||
System::nanoTime, _ -> null);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -145,12 +214,25 @@ public final class CompletionResolver implements TurnListener {
|
||||
* package.
|
||||
*/
|
||||
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
|
||||
ExhaustionSink exhaustionSink, LongSupplier nowNanos) {
|
||||
ExhaustionSink exhaustionSink, BackendErrorPatternLookup backendErrorPatterns,
|
||||
BackendErrorSink backendErrorSink, LongSupplier nowNanos) {
|
||||
this(agents, rendezvous, exhaustedPatterns, exhaustionSink, backendErrorPatterns, backendErrorSink,
|
||||
nowNanos, _ -> null);
|
||||
}
|
||||
|
||||
/** Full constructor with injectable clock and member location lookup. */
|
||||
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
|
||||
ExhaustionSink exhaustionSink, BackendErrorPatternLookup backendErrorPatterns,
|
||||
BackendErrorSink backendErrorSink, LongSupplier nowNanos,
|
||||
Function<String, WorktreeBranch> worktreeBranches) {
|
||||
this.agents = agents;
|
||||
this.rendezvous = rendezvous;
|
||||
this.exhaustedPatterns = Objects.requireNonNull(exhaustedPatterns, "exhaustedPatterns");
|
||||
this.exhaustionSink = Objects.requireNonNull(exhaustionSink, "exhaustionSink");
|
||||
this.backendErrorPatterns = Objects.requireNonNull(backendErrorPatterns, "backendErrorPatterns");
|
||||
this.backendErrorSink = Objects.requireNonNull(backendErrorSink, "backendErrorSink");
|
||||
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
|
||||
this.worktreeBranches = Objects.requireNonNull(worktreeBranches, "worktreeBranches");
|
||||
}
|
||||
|
||||
@Override
|
||||
@@ -180,7 +262,7 @@ public final class CompletionResolver implements TurnListener {
|
||||
baseline = null; // fail open: no baseline ⇒ no suppression
|
||||
log.debug("delivery baseline for {} failed: {}", target, e.getMessage());
|
||||
}
|
||||
inFlight.put(target, new InFlight(waiter, baseline, nowNanos.getAsLong()));
|
||||
inFlight.put(target, new InFlight(waiter, baseline, nowNanos.getAsLong(), token.injectedText()));
|
||||
}
|
||||
|
||||
/** The turn currently baselined for {@code target}, or {@code null} — a test hook for the captureBaseline path. */
|
||||
@@ -232,7 +314,7 @@ public final class CompletionResolver implements TurnListener {
|
||||
// screen, since that is usually the backend's own error.
|
||||
long elapsedNanos = nowNanos.getAsLong() - turn.deliveredAtNanos();
|
||||
if (elapsedNanos < MIN_TURN_NANOS) {
|
||||
fail(target, turn, tooFastReason(target, elapsedNanos));
|
||||
failTooFast(target, turn, waiter, elapsedNanos);
|
||||
return;
|
||||
}
|
||||
String tail;
|
||||
@@ -302,21 +384,33 @@ public final class CompletionResolver implements TurnListener {
|
||||
}
|
||||
return;
|
||||
}
|
||||
// fleetd#164 (part 2): a scrape that read cleanly and produced content still isn't a real
|
||||
// reply when that content is the backend's own rejection (e.g. an HTTP 400 before the worker
|
||||
// did any work). Classify it as a failure naming the member, rather than handing the caller a
|
||||
// scrape that reads like a completed answer.
|
||||
String backendError = firstMatchingLine(assistantBlock, BACKEND_ERROR);
|
||||
// fleetd#164 (part 2) / fleetd#201: a scrape that read cleanly and produced content still
|
||||
// isn't a real reply when that content is the backend's own rejection (e.g. an HTTP 400
|
||||
// before the worker did any work). Classify it as a failure naming the member, rather than
|
||||
// handing the caller a scrape that reads like a completed answer, and — only on the
|
||||
// resolution that actually wins the race, mirroring the exhaustion sink above — notify the
|
||||
// typed backend-error sink so a later stage can act on repeated failures.
|
||||
String backendError = firstMatchingLine(assistantBlock, backendErrorPatternOrFallback(target));
|
||||
if (backendError != null) {
|
||||
// Carry the whole scrape, not just the matched line. The pattern is a heuristic: a member
|
||||
// that forgot fleet_reply while reporting *about* a backend error matches it too. Failing
|
||||
// is still right — the caller must not read a scrape as an answer — but dropping the rest
|
||||
// of the pane would destroy the report, which is the same defect fleetd#164 is about.
|
||||
fail(target, turn, "member " + target + " ended on a backend error: " + backendError
|
||||
+ "\n--- pane tail ---\n" + tail);
|
||||
String reason = "member " + target + " ended on a backend error: " + backendError
|
||||
+ "\n--- pane tail ---\n" + tail;
|
||||
if (rendezvous.resolveFailure(waiter, reason)) {
|
||||
inFlight.remove(target, turn);
|
||||
log.warn("failing send to {} via turn-stall fallback: {}", target, reason);
|
||||
// fleetd#201 Unit 1: only on the resolution that actually won the race — a late
|
||||
// duplicate must never double-count one backend failure.
|
||||
backendErrorSink.onBackendError(target, backendError, reason);
|
||||
}
|
||||
return;
|
||||
}
|
||||
String completion = clipped ? tail + "\n" + CLIPPED_PANE_TAIL_MARKER : tail;
|
||||
if (echoesInjectedBrief(tail, turn.injectedText())) {
|
||||
completion = noReportMessage(target) + (clipped ? "\n" + CLIPPED_PANE_TAIL_MARKER : "");
|
||||
}
|
||||
if (rendezvous.resolveCompletion(waiter, completion)) {
|
||||
inFlight.remove(target, turn);
|
||||
if (clipped) {
|
||||
@@ -329,6 +423,40 @@ public final class CompletionResolver implements TurnListener {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* A full echoed brief is at least 400 normalised characters. A scrape that contains the brief may
|
||||
* add no more than 160 normalised characters of TUI chrome. This accepts harmless status text, but
|
||||
* preserves a real report that restates the full brief before adding substantive content.
|
||||
*/
|
||||
static boolean echoesInjectedBrief(String scrape, String injectedText) {
|
||||
String normalScrape = normalize(scrape);
|
||||
String normalInjected = normalize(injectedText);
|
||||
if (normalScrape.length() < ECHO_MIN_CHARS || normalInjected.length() < ECHO_MIN_CHARS) {
|
||||
return false;
|
||||
}
|
||||
if (normalInjected.contains(normalScrape)) {
|
||||
return true;
|
||||
}
|
||||
return normalScrape.contains(normalInjected)
|
||||
&& normalScrape.length() - normalInjected.length() <= MAX_ECHO_EXCESS_CHARS;
|
||||
}
|
||||
|
||||
private static String normalize(String text) {
|
||||
return text == null ? "" : text.toLowerCase().replaceAll("[^a-z0-9]+", "");
|
||||
}
|
||||
|
||||
private String noReportMessage(String target) {
|
||||
WorktreeBranch location = worktreeBranches.apply(target);
|
||||
if (location == null || (location.worktree() == null && location.branch() == null)) {
|
||||
return NO_REPORT_PREFIX + "]";
|
||||
}
|
||||
String locationText = location.worktree() == null ? "" : " worktree=" + location.worktree();
|
||||
if (location.branch() != null) {
|
||||
locationText += " branch=" + location.branch();
|
||||
}
|
||||
return NO_REPORT_PREFIX + locationText + "]";
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd#211: the raw-scrape fallback classification, run only when {@link #lastAssistantBlock}
|
||||
* found nothing usable (see the call site in {@link #resolve}). Mirrors the two classifications
|
||||
@@ -353,14 +481,20 @@ public final class CompletionResolver implements TurnListener {
|
||||
}
|
||||
return true;
|
||||
}
|
||||
String backendError = firstMatchingLine(raw, BACKEND_ERROR);
|
||||
String backendError = firstMatchingLine(raw, backendErrorPatternOrFallback(target));
|
||||
if (backendError != null) {
|
||||
// Carry the pane, not just the matched line — the same fleetd#164 rule the normal path
|
||||
// above applies. Here it matters more, not less: the trimmed block was empty, so the raw
|
||||
// scrape is the ONLY copy of whatever the member managed to say. Clipped to the same cap
|
||||
// the normal path uses, since a raw screen has no boundary trimming to bound it.
|
||||
fail(target, turn, "member " + target + " ended on a backend error: " + backendError
|
||||
+ "\n--- pane tail ---\n" + clip(raw));
|
||||
String reason = "member " + target + " ended on a backend error: " + backendError
|
||||
+ "\n--- pane tail ---\n" + clip(raw);
|
||||
if (rendezvous.resolveFailure(waiter, reason)) {
|
||||
inFlight.remove(target, turn);
|
||||
log.warn("failing send to {} via turn-stall fallback from the raw scrape: {}", target, reason);
|
||||
// fleetd#201 Unit 1: only on the resolution that actually won the race.
|
||||
backendErrorSink.onBackendError(target, backendError, reason);
|
||||
}
|
||||
return true;
|
||||
}
|
||||
return false;
|
||||
@@ -402,22 +536,53 @@ public final class CompletionResolver implements TurnListener {
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd#164: the failure reason for a turn that settled inside {@link #MIN_TURN_NANOS} — names
|
||||
* the member and both timings, and appends whatever the pane shows (usually the backend's own
|
||||
* error) so the caller sees the cause, not just "it failed".
|
||||
* fleetd#164 (floor) / fleetd#201 (classification): fail a turn that settled inside
|
||||
* {@link #MIN_TURN_NANOS} — a crash signature (e.g. a backend HTTP 400 before the worker did
|
||||
* anything) that a bare {@code BUSY -> DONE} transition cannot be told apart from a genuinely
|
||||
* fast completion. Runs the same backend-error classification the normal and raw-scrape paths
|
||||
* apply, against whatever is on screen right now: a match is a typed failure that notifies
|
||||
* {@link #backendErrorSink} (only on the resolution that wins the race); a non-match stays the
|
||||
* original generic too-fast failure, naming the member and both timings, with whatever the pane
|
||||
* shows appended so the caller sees the cause, not just "it failed".
|
||||
*/
|
||||
private String tooFastReason(String target, long elapsedNanos) {
|
||||
private void failTooFast(String target, InFlight turn, CompletableFuture<Rendezvous.Resolution> waiter,
|
||||
long elapsedNanos) {
|
||||
String scrape;
|
||||
try {
|
||||
scrape = clip(agents.read(target, SCRAPE_SOURCE));
|
||||
scrape = agents.read(target, SCRAPE_SOURCE);
|
||||
} catch (RuntimeException e) {
|
||||
scrape = "";
|
||||
}
|
||||
String reason = String.format(
|
||||
String clippedScrape = clip(scrape);
|
||||
String baseReason = String.format(
|
||||
"member %s went BUSY -> DONE in %dms (floor %dms) — too fast to be real work, most "
|
||||
+ "likely a backend error before any work started",
|
||||
target, elapsedNanos / 1_000_000, MIN_TURN_NANOS / 1_000_000);
|
||||
return scrape.isBlank() ? reason : reason + ": " + scrape;
|
||||
String backendError = firstMatchingLine(scrape, backendErrorPatternOrFallback(target));
|
||||
if (backendError != null) {
|
||||
String reason = baseReason + ": " + clippedScrape;
|
||||
if (rendezvous.resolveFailure(waiter, reason)) {
|
||||
inFlight.remove(target, turn);
|
||||
log.warn("failing send to {} via turn-stall fallback: {}", target, reason);
|
||||
// fleetd#201 Unit 1: only on the resolution that actually won the race.
|
||||
backendErrorSink.onBackendError(target, backendError, reason);
|
||||
}
|
||||
return;
|
||||
}
|
||||
fail(target, turn, clippedScrape.isBlank() ? baseReason : baseReason + ": " + clippedScrape);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd#201: the pattern to classify a backend error against for {@code target} — its
|
||||
* profile's configured {@link BackendErrorPatternLookup} entry when there is one, else the
|
||||
* built-in {@link #BACKEND_ERROR} compatibility pattern. A target with no configured pattern is
|
||||
* never "off": it always falls back to this narrow default, exactly like the classification
|
||||
* behaved before fleetd#201 (Unit 5 reports a target relying on this fallback separately, as
|
||||
* legacy-default coverage rather than full per-profile coverage).
|
||||
*/
|
||||
private Pattern backendErrorPatternOrFallback(String target) {
|
||||
Pattern configured = backendErrorPatterns.patternFor(target);
|
||||
return configured != null ? configured : BACKEND_ERROR;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -455,9 +620,9 @@ public final class CompletionResolver implements TurnListener {
|
||||
* @param allProfiles every configured profile name
|
||||
* @param configuredProfiles the subset of {@code allProfiles} that carry an exhausted pattern
|
||||
*/
|
||||
public static String coverage(Set<String> allProfiles, Set<String> configuredProfiles) {
|
||||
public static String coverage(String patternKey, Set<String> allProfiles, Set<String> configuredProfiles) {
|
||||
if (configuredProfiles.isEmpty()) {
|
||||
return "off (no profile has an exhaustedPattern configured; profiles: " + sorted(allProfiles) + ")";
|
||||
return "off (no profile has an " + patternKey + " configured; profiles: " + sorted(allProfiles) + ")";
|
||||
}
|
||||
Set<String> unconfigured = new TreeSet<>(allProfiles);
|
||||
unconfigured.removeAll(configuredProfiles);
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
* Notified when {@link CompletionResolver} actually delivers a {@code BACKEND_EXHAUSTED}
|
||||
* classification to a waiting send (CB-578 stage B) — never on a race that lost (see
|
||||
@@ -10,22 +12,81 @@ package dev.ltms.fleet.inject;
|
||||
* profiles or credentials, so mapping {@code target} to whatever should be quarantined is entirely
|
||||
* the sink's job — see {@code Fleetd.main}'s wiring, which resolves target → session → profile →
|
||||
* {@code effectiveCredentialId()} and calls {@code BackendQuarantine.quarantine} on it.
|
||||
*
|
||||
* <p><strong>The 3-arg overload is the single abstract method — fleetd #234, round 4.</strong> Two
|
||||
* earlier rounds each shipped a caller that silently dropped the profile hint (see {@link
|
||||
* #onExhausted(String, String, String)}): a lambda written against this interface can only ever
|
||||
* implement whichever overload is abstract, and while the 2-arg form held that position, EVERY
|
||||
* lambda site — a call-site forwarder in {@code Fleetd.java}, a hand-built test double — bound to
|
||||
* it and silently inherited the profile-dropping default, whether or not its author remembered the
|
||||
* hazard. Making the 3-arg form abstract instead removes the shape entirely: a lambda declared
|
||||
* against this interface today is <em>forced</em> by the compiler to take {@code (target, reason,
|
||||
* profile)}, so there is no overload left for it to bind to that can drop the hint. This is a type
|
||||
* change, not a test — it holds even for a caller that has never heard of fleetd #234.
|
||||
*/
|
||||
@FunctionalInterface
|
||||
public interface ExhaustionSink {
|
||||
|
||||
/**
|
||||
* @param target the herdr terminal id whose turn was classified {@code BACKEND_EXHAUSTED}
|
||||
* @param reason the matched-line reason carried by the classification
|
||||
* @param target the herdr terminal id whose turn was classified {@code BACKEND_EXHAUSTED}
|
||||
* @param reason the matched-line reason carried by the classification
|
||||
* @param profile the profile the caller already knows should be quarantined, or {@code null}
|
||||
* when the caller has no better answer than {@code target} alone (fleetd #234):
|
||||
* a caller whose {@code target} is not yet resolvable through whatever roster
|
||||
* the sink's implementation consults — {@link
|
||||
* dev.ltms.fleet.member.OpenCodeLauncher}'s model-mismatch check fires from
|
||||
* {@code SessionAwareHandle.agentSessionId()}, which runs during {@code
|
||||
* SessionManager.acquire()} <em>before</em> that session is registered, so a
|
||||
* target -> session -> profile lookup finds nothing at that point. That launcher
|
||||
* already has its own {@code FleetConfig.Profile} in hand and does not need the
|
||||
* roster to know which profile to quarantine, so it supplies this directly
|
||||
* instead of leaving the sink to guess.
|
||||
*/
|
||||
void onExhausted(String target, String reason);
|
||||
void onExhausted(String target, String reason, String profile);
|
||||
|
||||
/**
|
||||
* Convenience for a caller with no profile to offer — every existing call site that predates
|
||||
* the hint (fleetd #234): {@link dev.ltms.fleet.inject.CompletionResolver}'s two call sites
|
||||
* always call with a {@code target} that IS live in the roster at the time of the call, so they
|
||||
* need no hint and keep working exactly as before, unchanged by this default.
|
||||
*
|
||||
* @param target as {@link #onExhausted(String, String, String)}
|
||||
* @param reason as {@link #onExhausted(String, String, String)}
|
||||
*/
|
||||
default void onExhausted(String target, String reason) {
|
||||
onExhausted(target, reason, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Inert sink — nothing happens on exhaustion. The explicit stand-in a caller (or a test not
|
||||
* exercising this feature) passes instead of a defaulting overload, exactly like
|
||||
* {@link ExhaustedPatternLookup#none()}.
|
||||
* {@link ExhaustedPatternLookup#none()}. Safe as a lambda: the 3-arg form is now the interface's
|
||||
* single abstract method, so a lambda here has no other overload to silently bind to instead —
|
||||
* it simply does nothing with all three arguments.
|
||||
*/
|
||||
static ExhaustionSink none() {
|
||||
return (target, reason) -> { };
|
||||
return (target, reason, profile) -> { };
|
||||
}
|
||||
|
||||
/**
|
||||
* A sink that forwards to whatever {@code target} currently supplies (fleetd #234). Exists to
|
||||
* break a genuine construction-order cycle: {@code Fleetd.main} builds its adapters (including
|
||||
* {@link dev.ltms.fleet.member.OpenCodeLauncher}) before {@code sessions} exists, so it cannot
|
||||
* hand them the real sink yet — it hands them a forwarder pointed at an {@code
|
||||
* AtomicReference<ExhaustionSink>} that starts at {@link #none()} and gets {@code .set()} to the
|
||||
* real sink once {@code sessions} is built. {@code target} is evaluated on every call, never
|
||||
* cached, so the forwarder keeps working after the reference is repointed.
|
||||
*
|
||||
* <p>Now safe as a one-line lambda (round 4): forwarding the single 3-arg abstract method
|
||||
* forwards everything a caller can supply — there is no separate 2-arg overload left for a
|
||||
* forwarder to bind to instead and silently lose the hint. Kept as a named factory rather than
|
||||
* written inline at each call site anyway, so a test can call the exact object {@code
|
||||
* Fleetd.java} builds instead of asserting a rebuilt copy of its shape (round 3's lesson).
|
||||
*
|
||||
* @param target supplies the sink to forward to, evaluated fresh on every call
|
||||
* @return a sink whose call delegates to {@code target.get()}
|
||||
*/
|
||||
static ExhaustionSink forwardingTo(Supplier<ExhaustionSink> target) {
|
||||
return (t, r, p) -> target.get().onExhausted(t, r, p);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -10,12 +10,14 @@ import dev.ltms.fleet.herdr.Agent;
|
||||
import dev.ltms.fleet.metrics.FleetMetrics;
|
||||
import dev.ltms.fleet.metrics.Metrics;
|
||||
import dev.ltms.fleet.inject.MemberPresence;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.herdr.HerdrException;
|
||||
import dev.ltms.fleet.msg.LeadChannel;
|
||||
import dev.ltms.fleet.msg.LeadMessage;
|
||||
import dev.ltms.fleet.msg.MessageService;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.peer.PeerUnreachableException;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementException;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
@@ -92,6 +94,10 @@ public final class FleetMcp {
|
||||
private final CapacitySource capacity;
|
||||
private final HealthCoverageSource healthCoverage;
|
||||
private final QuarantineSource quarantine;
|
||||
/** fleetd #201 Unit 5: SEPARATE from {@link #quarantine} — see {@link OutageSource}'s doc. */
|
||||
private final OutageSource outage;
|
||||
/** fleetd #176: SEPARATE from both of the above — see {@link LeadSeatSource}'s doc. */
|
||||
private final LeadSeatSource leadSeats;
|
||||
/** CB-637: this daemon's lead-to-lead channel; {@code null} when no coordinator is configured. */
|
||||
private final LeadChannel leadChannel;
|
||||
|
||||
@@ -115,6 +121,40 @@ public final class FleetMcp {
|
||||
public static QuarantineSource none() { return new QuarantineSource(_ -> null, BackendQuarantine.none()); }
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #201 Unit 5 cool-off facts used by {@code fleet_list}/{@code fleet_profiles}: a
|
||||
* profile → credential id lookup, plus the shared {@link BackendOutagePolicy} to read remaining
|
||||
* cool-offs off. A SEPARATE source from {@link QuarantineSource} — never merged into it — so a
|
||||
* credential that is cooling off after repeated backend errors is reported independently of
|
||||
* whether it is ALSO under CB-578 stage B exhaustion quarantine; the two checks can both fire
|
||||
* for the same profile, and when they do, {@code fleet_list}/{@code fleet_profiles} report both
|
||||
* (see {@link #capacityView}/{@link #profiles}).
|
||||
*/
|
||||
public record OutageSource(Function<String, String> credentialIdFor, BackendOutagePolicy outagePolicy) {
|
||||
/** Inert source — no profile is ever reported cooling off. Explicit stand-in, not a default. */
|
||||
public static OutageSource none() {
|
||||
return new OutageSource(_ -> null, new BackendOutagePolicy(System::nanoTime));
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #176: the seats a profile's own live LEAD session(s) hold on the same Claude
|
||||
* subscription — the third reason (alongside {@link QuarantineSource} and {@link OutageSource})
|
||||
* {@code free} can overstate what a fresh {@code fleet_spawn} would actually get.
|
||||
*
|
||||
* <p>{@code maxLoad} counts only <em>members</em>, never the lead itself. But a
|
||||
* {@code subscription: true} profile bills the operator's own Claude account, and the lead is
|
||||
* always a live {@code claude} session on that same account (it is never moved off-subscription
|
||||
* — see {@code LeadLauncher}). So a fan-out that fills every member slot still leaves the lead's
|
||||
* own seat unaccounted for, and the daemon reports a slot that was never really free. See
|
||||
* {@code Fleetd.leadSeatLookup} for how the count is derived — from {@code fleet.leaders.<name>
|
||||
* .profile} and each profile's {@code effectiveCredentialId()}, never a hardcoded constant.
|
||||
*/
|
||||
public record LeadSeatSource(Function<String, Integer> seatsFor) {
|
||||
/** Inert source — no profile is ever reported as sharing a seat with a lead. */
|
||||
public static LeadSeatSource none() { return new LeadSeatSource(_ -> 0); }
|
||||
}
|
||||
|
||||
/**
|
||||
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
|
||||
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
|
||||
@@ -129,7 +169,7 @@ public final class FleetMcp {
|
||||
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine) {
|
||||
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, metrics, capacity,
|
||||
healthCoverage, quarantine, null);
|
||||
healthCoverage, quarantine, null, OutageSource.none(), LeadSeatSource.none());
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -142,9 +182,43 @@ public final class FleetMcp {
|
||||
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
|
||||
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, LeadChannel leadChannel) {
|
||||
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, metrics, capacity,
|
||||
healthCoverage, quarantine, leadChannel, OutageSource.none(), LeadSeatSource.none());
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, with fleetd #201 Unit 5 cool-off facts for {@code fleet_list}/{@code fleet_profiles}
|
||||
* (see {@link OutageSource}).
|
||||
*
|
||||
* @param outage required — pass {@link OutageSource#none()} for a caller that does not want the
|
||||
* feature, never a defaulting overload (the same rule {@code quarantine} follows).
|
||||
*/
|
||||
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
|
||||
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
|
||||
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage) {
|
||||
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, metrics, capacity,
|
||||
healthCoverage, quarantine, leadChannel, outage, LeadSeatSource.none());
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, with fleetd #176 lead-seat facts (see {@link LeadSeatSource}). This is what
|
||||
* {@code Fleetd.main} actually wires up.
|
||||
*
|
||||
* @param leadSeats required — pass {@link LeadSeatSource#none()} for a caller that does not want
|
||||
* the feature, never a defaulting overload (the same rule {@code quarantine} and
|
||||
* {@code outage} follow).
|
||||
*/
|
||||
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
|
||||
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
|
||||
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage,
|
||||
LeadSeatSource leadSeats) {
|
||||
this.leadChannel = leadChannel;
|
||||
this.capacity = capacity;
|
||||
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
|
||||
this.outage = Objects.requireNonNull(outage, "outage");
|
||||
this.leadSeats = Objects.requireNonNull(leadSeats, "leadSeats");
|
||||
this.healthCoverage = healthCoverage;
|
||||
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
|
||||
this.transport = HttpServletStreamableServerTransportProvider.builder()
|
||||
@@ -273,8 +347,8 @@ public final class FleetMcp {
|
||||
(exchange, _) -> {
|
||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
|
||||
if (denied != null) return denied;
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine,
|
||||
callers == null ? Map.of() : callers.leads(),
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
|
||||
leadSeats, callers == null ? Map.of() : callers.leads(),
|
||||
callerTerminal(exchange),
|
||||
leadChannel == null ? null : leadChannel.selfCoordId());
|
||||
};
|
||||
@@ -289,7 +363,7 @@ public final class FleetMcp {
|
||||
(exchange, _) -> {
|
||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
|
||||
if (denied != null) return denied;
|
||||
return profiles(workers, quarantine);
|
||||
return profiles(workers, quarantine, outage);
|
||||
};
|
||||
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> whoamiHandler =
|
||||
(exchange, _) -> {
|
||||
@@ -543,8 +617,9 @@ public final class FleetMcp {
|
||||
case REPLIED -> text(r.text());
|
||||
// The worker's turn finished but it never called fleet_reply — hand back the scraped
|
||||
// transcript tail, flagged so the primary knows it isn't a structured reply.
|
||||
case COMPLETED_UNREPLIED -> text(
|
||||
"[worker finished without a structured fleet_reply — transcript tail follows]\n" + r.text());
|
||||
case COMPLETED_UNREPLIED -> text(r.text().startsWith(CompletionResolver.NO_REPORT_PREFIX)
|
||||
? r.text()
|
||||
: "[worker finished without a structured fleet_reply — transcript tail follows]\n" + r.text());
|
||||
// The worker ran the turn then wedged (CB-109) — surface the error context.
|
||||
case WORKER_FAILED -> text("[worker failed — turn ended in an unrecoverable state]\n" + r.text());
|
||||
// The backend refused on a subscription usage limit (CB-578 stage A) — the worker's
|
||||
@@ -659,6 +734,7 @@ public final class FleetMcp {
|
||||
}
|
||||
return switch (v.phase()) {
|
||||
case DONE -> text(v.replySource() != null && v.replySource().equals("transcript")
|
||||
&& !v.reply().startsWith(CompletionResolver.NO_REPORT_PREFIX)
|
||||
? "[done — worker finished without a structured fleet_reply; transcript tail follows]\n" + v.reply()
|
||||
: v.reply());
|
||||
case PENDING -> text("[pending — " + v.detail() + "]");
|
||||
@@ -872,25 +948,47 @@ public final class FleetMcp {
|
||||
* has ever been quarantined gets exactly the pre-stage-B shape.
|
||||
*/
|
||||
static McpSchema.CallToolResult profiles(PeerLauncher workers, QuarantineSource quarantine) {
|
||||
return profiles(workers, quarantine, OutageSource.none());
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, plus a SEPARATE {@code coolingOff} map (fleetd #201 Unit 5) — never merged into
|
||||
* {@code quarantined} — for a profile whose credential is cooling off after repeated backend
|
||||
* errors ({@link BackendOutagePolicy}). The two checks are independent: a profile can appear in
|
||||
* both maps at once when it is both exhaustion-quarantined AND cooling off.
|
||||
*/
|
||||
static McpSchema.CallToolResult profiles(PeerLauncher workers, QuarantineSource quarantine, OutageSource outage) {
|
||||
Map<String, Object> result = new LinkedHashMap<>();
|
||||
result.put("profiles", workers.profiles());
|
||||
result.put("default", workers.defaultProfile() == null ? "" : workers.defaultProfile());
|
||||
Map<String, Object> quarantined = new LinkedHashMap<>();
|
||||
Map<String, Object> coolingOff = new LinkedHashMap<>();
|
||||
for (String profile : workers.profiles()) {
|
||||
String credentialId = quarantine.credentialIdFor().apply(profile);
|
||||
if (credentialId == null) {
|
||||
continue;
|
||||
if (credentialId != null) {
|
||||
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
|
||||
Map<String, Object> row = new LinkedHashMap<>();
|
||||
row.put("credentialId", credentialId);
|
||||
row.put("quarantinedForSeconds", remaining);
|
||||
quarantined.put(profile, row);
|
||||
});
|
||||
}
|
||||
String outageCredentialId = outage.credentialIdFor().apply(profile);
|
||||
if (outageCredentialId != null) {
|
||||
outage.outagePolicy().remainingCoolOffSeconds(outageCredentialId).ifPresent(remaining -> {
|
||||
Map<String, Object> row = new LinkedHashMap<>();
|
||||
row.put("credentialId", outageCredentialId);
|
||||
row.put("coolingOffForSeconds", remaining);
|
||||
coolingOff.put(profile, row);
|
||||
});
|
||||
}
|
||||
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
|
||||
Map<String, Object> row = new LinkedHashMap<>();
|
||||
row.put("credentialId", credentialId);
|
||||
row.put("quarantinedForSeconds", remaining);
|
||||
quarantined.put(profile, row);
|
||||
});
|
||||
}
|
||||
if (!quarantined.isEmpty()) {
|
||||
result.put("quarantined", quarantined);
|
||||
}
|
||||
if (!coolingOff.isEmpty()) {
|
||||
result.put("coolingOff", coolingOff);
|
||||
}
|
||||
return text(json(result));
|
||||
}
|
||||
|
||||
@@ -929,6 +1027,15 @@ public final class FleetMcp {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, leads, selfTerm, null);
|
||||
}
|
||||
|
||||
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, OutageSource outage,
|
||||
Map<String, String> leads, String selfTerm) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
|
||||
LeadSeatSource.none(), leads, selfTerm, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, additionally reporting this daemon's own lead coordination id (CB-637) when one is
|
||||
* configured and its channel opened. There is no peer-discovery surface yet — a lead addresses a
|
||||
@@ -942,6 +1049,25 @@ public final class FleetMcp {
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, Map<String, String> leads, String selfTerm,
|
||||
String selfCoordId) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, OutageSource.none(),
|
||||
LeadSeatSource.none(), leads, selfTerm, selfCoordId);
|
||||
}
|
||||
|
||||
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, OutageSource outage,
|
||||
Map<String, String> leads, String selfTerm, String selfCoordId) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
|
||||
LeadSeatSource.none(), leads, selfTerm, selfCoordId);
|
||||
}
|
||||
|
||||
/** As above, plus fleetd #176 lead-seat facts (see {@link LeadSeatSource}). */
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, OutageSource outage,
|
||||
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
|
||||
String selfCoordId) {
|
||||
try {
|
||||
Map<String, Agent> live = workers.list().stream()
|
||||
.map(Agent.class::cast)
|
||||
@@ -968,7 +1094,7 @@ public final class FleetMcp {
|
||||
}
|
||||
if (capacity.available()) result.put("capacity", profiles.stream()
|
||||
.map(profile -> capacityView(profile, capacity.liveCount(), capacity.maxLoad(), roster, messages,
|
||||
capacity.clock().getAsLong(), quarantine)).toList());
|
||||
capacity.clock().getAsLong(), quarantine, outage, leadSeats)).toList());
|
||||
return text(json(result));
|
||||
} catch (HerdrException e) {
|
||||
return error("herdr error listing the fleet: " + e.getMessage());
|
||||
@@ -1001,19 +1127,39 @@ public final class FleetMcp {
|
||||
* reports, reusing {@link QuarantineSource} rather than a second lookup. Both new keys are
|
||||
* added only when the profile is actually quarantined, so an ordinary fleet's rows are
|
||||
* byte-identical to before this change.
|
||||
*
|
||||
* <p>fleetd #201 Unit 5: a profile whose credential is cooling off (a SEPARATE, independent
|
||||
* check from quarantine — see {@link OutageSource}) also forces {@code free} to 0 and carries
|
||||
* {@code credentialId}/{@code coolingOffForSeconds}, but never {@code quarantinedForSeconds} —
|
||||
* that key is added only when exhaustion quarantine is ALSO active for this profile, since the
|
||||
* two checks are independent and either, both, or neither can be true.
|
||||
*
|
||||
* <p>fleetd #176: {@code maxLoad} counts panes, not subscription seats — it never counted the
|
||||
* lead's own seat on a {@code subscription: true} profile's account. {@link LeadSeatSource}
|
||||
* reports that count (0 for a non-subscription profile, or when no live lead shares its
|
||||
* credential), and it is subtracted from {@code free} the same way {@code live} already is —
|
||||
* {@code maxLoad} itself is left untouched, so the row still reports the configured cap. The
|
||||
* {@code leadSeats} key is added only when the count is positive, for the same
|
||||
* byte-identical-when-unused reason as the quarantine/cool-off keys above.
|
||||
*/
|
||||
private static Map<String, Object> capacityView(String profile, Function<String, Integer> liveCount,
|
||||
Function<String, Integer> maxLoad, List<MemberSession> roster,
|
||||
MessageService messages, long nowNanos, QuarantineSource quarantine) {
|
||||
MessageService messages, long nowNanos, QuarantineSource quarantine,
|
||||
OutageSource outage, LeadSeatSource leadSeats) {
|
||||
Integer cap = maxLoad.apply(profile);
|
||||
int live = liveCount.apply(profile);
|
||||
int leadSeatCount = leadSeats.seatsFor().apply(profile);
|
||||
int reclaimable = (int) roster.stream().filter(s -> profile.equals(s.profile()))
|
||||
.filter(s -> (s.state() == MemberSession.State.READY || s.state() == MemberSession.State.DONE))
|
||||
.filter(s -> messages == null || (!messages.hasAcceptedDelivery(s.terminalId()) && !messages.hasInboxMessage(s.terminalId())))
|
||||
.count();
|
||||
Map<String, Object> row = new LinkedHashMap<>();
|
||||
row.put("profile", profile); row.put("maxLoad", cap); row.put("live", live);
|
||||
row.put("free", cap == null ? null : Math.max(0, cap - live)); row.put("reclaimable", reclaimable);
|
||||
row.put("free", cap == null ? null : Math.max(0, cap - live - leadSeatCount));
|
||||
row.put("reclaimable", reclaimable);
|
||||
if (leadSeatCount > 0) {
|
||||
row.put("leadSeats", leadSeatCount);
|
||||
}
|
||||
String credentialId = quarantine.credentialIdFor().apply(profile);
|
||||
if (credentialId != null) {
|
||||
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
|
||||
@@ -1022,6 +1168,14 @@ public final class FleetMcp {
|
||||
row.put("quarantinedForSeconds", remaining);
|
||||
});
|
||||
}
|
||||
String outageCredentialId = outage.credentialIdFor().apply(profile);
|
||||
if (outageCredentialId != null) {
|
||||
outage.outagePolicy().remainingCoolOffSeconds(outageCredentialId).ifPresent(remaining -> {
|
||||
row.put("free", 0);
|
||||
row.put("credentialId", outageCredentialId);
|
||||
row.put("coolingOffForSeconds", remaining);
|
||||
});
|
||||
}
|
||||
return row;
|
||||
}
|
||||
|
||||
@@ -1161,7 +1315,11 @@ public final class FleetMcp {
|
||||
+ "worktree:<ticket-slug> to provision an isolated git worktree. Pass resumeSessionId "
|
||||
+ "to relaunch onto a prior conversation instead of starting cold — this requires an "
|
||||
+ "explicit profile whose backend supports it (fleet_list shows agentSessionId for "
|
||||
+ "resumable members), and is refused otherwise rather than silently starting fresh. "
|
||||
+ "resumable members; it is absent for a member fleetd cannot reliably re-identify, "
|
||||
+ "e.g. an opencode member spawned without a worktree), and is refused otherwise "
|
||||
+ "rather than silently starting fresh. For an opencode profile, resumeSessionId "
|
||||
+ "itself also requires worktree:true/<slug> on THIS spawn — without one fleetd can "
|
||||
+ "never re-verify which conversation it actually resumed (fleetd #249). "
|
||||
+ "sessionName gives the member a display name in its own UI when the backend supports "
|
||||
+ "one. Returns the member's sessionId (use with fleet_send) and paneId (use with "
|
||||
+ "fleet_stop).",
|
||||
@@ -1172,7 +1330,7 @@ public final class FleetMcp {
|
||||
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
|
||||
"ticket", stringProp("Ticket slug when worktree:true"),
|
||||
"sessionName", stringProp("Logical display name for the member's own session, when its backend supports one"),
|
||||
"resumeSessionId", stringProp("A prior member's agentSessionId (from fleet_list) to resume — requires an explicit profile that supports it")),
|
||||
"resumeSessionId", stringProp("A prior member's agentSessionId (from fleet_list) to resume — requires an explicit profile that supports it, and (for opencode) a worktree on this spawn too")),
|
||||
List.of()));
|
||||
}
|
||||
|
||||
@@ -1193,9 +1351,14 @@ public final class FleetMcp {
|
||||
+ "discover a peer lead without being told its address. 'members' are the "
|
||||
+ "sessions delegated to — each with sessionId, paneId, role (architect/dev/"
|
||||
+ "reviewer), profile (the backend it runs on), state, optional "
|
||||
+ "worktree/branch/owner/agentSessionId (the id to pass as fleet_spawn's "
|
||||
+ "resumeSessionId to relaunch onto that same conversation, when the backend "
|
||||
+ "supports it), and live herdr status. An empty 'members' "
|
||||
+ "worktree/branch/owner/agentSessionId, and live herdr status. agentSessionId, "
|
||||
+ "when present, is the id to pass as fleet_spawn's resumeSessionId to relaunch "
|
||||
+ "onto that same conversation. It is ABSENT — not a guess — for a member fleetd "
|
||||
+ "cannot reliably re-identify: some backends (e.g. opencode) resolve it from the "
|
||||
+ "member's working directory, which only uniquely identifies a member when it "
|
||||
+ "was spawned into its own fleetd-provisioned worktree (worktree:true/<slug>); a "
|
||||
+ "member spawned without one shares its directory with others and never reports "
|
||||
+ "an id, however long it runs (fleetd #249). An empty 'members' "
|
||||
+ "means no members are spawned; it says nothing about peers. When capacity "
|
||||
+ "facts are configured, a 'capacity' row per profile also reports free: 0 for "
|
||||
+ "a quarantined profile's credential (see fleet_profiles), whatever its "
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
package dev.ltms.fleet.member;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.fasterxml.jackson.databind.node.ObjectNode;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.Agent;
|
||||
@@ -14,6 +17,8 @@ import java.io.IOException;
|
||||
import java.io.UncheckedIOException;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.nio.file.StandardCopyOption;
|
||||
import java.nio.file.attribute.PosixFileAttributeView;
|
||||
import java.util.EnumSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
@@ -47,6 +52,21 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(ClaudeCodeLauncher.class);
|
||||
|
||||
/** JSON codec for the additive workspace-trust seed (fleetd #149) — Jackson's default settings. */
|
||||
private static final ObjectMapper TRUST_JSON = new ObjectMapper();
|
||||
|
||||
/**
|
||||
* Serialises every {@link #seedTrustDialog} read-modify-write for the whole daemon process.
|
||||
* Two claude-code spawns starting at once are normal (fleetd runs several members in parallel
|
||||
* routinely) and both would otherwise read the same {@code .claude.json}, add their own entry
|
||||
* to their own in-memory copy, and write — the second write wins and the first spawn's trust
|
||||
* entry silently disappears. A single process-wide lock is enough because every spawn on this
|
||||
* daemon runs in this one JVM; it does not protect against a second daemon process or the
|
||||
* operator's own Claude Code process writing at the same instant, which {@link #writeAtomically}
|
||||
* covers instead (each writer only ever sees a fully-old or fully-new file, never a torn one).
|
||||
*/
|
||||
private static final Object TRUST_JSON_LOCK = new Object();
|
||||
|
||||
private final SubscriptionGuard guard;
|
||||
|
||||
/**
|
||||
@@ -258,11 +278,15 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
workerEnv.remove("ANTHROPIC_AUTH_TOKEN");
|
||||
} else {
|
||||
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
|
||||
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", env.apply(cfg.tokenEnv()));
|
||||
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", resolveEnv(cfg.tokenEnv()));
|
||||
}
|
||||
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
|
||||
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
|
||||
applyGitToken(workerEnv, cfg);
|
||||
// fleetd #149: seed the workspace-trust entry BEFORE this spawn ever reaches herdr — see
|
||||
// seedTrustDialog for why, and isProvisionedWorktree for why this is gated to a worktree
|
||||
// fleetd itself provisioned (never a real checkout, never an un-configured fallback cwd).
|
||||
seedTrustDialog(cfg.configDir(), spec.cwd());
|
||||
|
||||
// CB-547a: Claude Code can MINT its own session id, so fleetd chooses it — a fresh spawn
|
||||
// gets a UUID we pass as --session-id and return from agentSessionId(), so the resume
|
||||
@@ -412,12 +436,12 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
*/
|
||||
private static void writeIdeOverlay(String cwd, String projectPath) {
|
||||
try {
|
||||
Path dotGit = Path.of(cwd, ".git");
|
||||
if (!Files.isRegularFile(dotGit)) {
|
||||
if (!isProvisionedWorktree(cwd)) {
|
||||
// Not a provisioned worktree (primary's real checkout has a .git directory, or the
|
||||
// cwd is not a repo at all). Never write into it.
|
||||
return;
|
||||
}
|
||||
Path dotGit = Path.of(cwd, ".git");
|
||||
// The overlay FILE lives at the worktree root (claude-code's cwd), but its CONTENT pins
|
||||
// project_path to the module dir the IDE opened (projectPath), not the worktree root.
|
||||
Files.writeString(Path.of(cwd, "CLAUDE.local.md"), PeerLauncher.ideOverlayText(projectPath));
|
||||
@@ -447,6 +471,173 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #149: pre-seed the workspace-trust entry for {@code cwd} in Claude Code's own config
|
||||
* file, BEFORE this launch ever reaches herdr (called from {@link #buildLaunch}, which always
|
||||
* runs before the base class starts the process). Claude Code asks an interactive, un-timed
|
||||
* "Is this a project you created or one you trust?" the first time it starts in a directory it
|
||||
* has not seen before, and every {@code worktree: true} spawn lands in a brand-new directory —
|
||||
* so without this seed the member sits on that dialog forever, never mounts the bridge MCP, and
|
||||
* never calls {@code fleet_reply}. herdr reports it as healthy the whole time ({@code
|
||||
* agent_status: blocked}, {@code interactive_ready: true}), so nothing else catches it. Measured
|
||||
* live on fleet01 2026-08-23: an unseeded fresh cwd sat on the dialog indefinitely; a seeded one
|
||||
* reached {@code idle} clean.
|
||||
*
|
||||
* <p>This is not a new grant — the operator already trusted this repo by configuring the
|
||||
* profile against it, and a worktree is a checkout of that same repo.
|
||||
*
|
||||
* <p>The key Claude Code reads is per-project, in {@code .claude.json}: {@code
|
||||
* projects.<cwd>.hasTrustDialogAccepted}. The file lives at {@code <configDir>/.claude.json}
|
||||
* when the profile sets {@code CLAUDE_CONFIG_DIR} (mirrors this launcher's own env var above),
|
||||
* else the default {@code ~/.claude.json} — the same file Claude Code itself would read either
|
||||
* way, so this seeds exactly what the spawned peer is about to open.
|
||||
*
|
||||
* <p><b>Additive, not a rewrite.</b> {@code .claude.json} is large (tens of KB, dozens of
|
||||
* projects) and Claude Code itself rewrites it while running, so this reads the file as a JSON
|
||||
* tree (missing or unreadable → treated as an empty object) and changes only
|
||||
* {@code projects.<cwd>.hasTrustDialogAccepted} (that key alone — see fleetd #247) —
|
||||
* every other top-level key and every other project entry is written back untouched. Only the
|
||||
* one project entry for {@code cwd} is replaced/created; an existing entry for a DIFFERENT cwd
|
||||
* (or the operator's own project history) is never touched.
|
||||
*
|
||||
* <p>Best-effort, like {@link #writeIdeOverlay}: a failure here (unwritable configDir, a
|
||||
* corrupt existing file, …) must never fail the spawn — it is logged at debug and swallowed. A
|
||||
* peer that starts without the seed still starts; it just may hit the dialog fleetd #149
|
||||
* describes.
|
||||
*
|
||||
* <p><b>Gated to a provisioned worktree</b> ({@link HerdrPeerLauncher#isProvisionedWorktree}) — see that
|
||||
* method's javadoc for the incident that made this gate mandatory, not optional: this must
|
||||
* never run against a real checkout or an un-configured fallback cwd, only the exact
|
||||
* always-fresh-directory population fleetd #149 describes.
|
||||
*
|
||||
* <p><b>Atomic and lock-protected.</b> {@code .claude.json} is a live file — Claude Code itself
|
||||
* rewrites it while running, and this daemon routinely spawns several members at once, each
|
||||
* calling this method for its own cwd. Every write goes through {@link #writeAtomically} (a
|
||||
* sibling-temp-file + {@code ATOMIC_MOVE}, never a truncate-in-place) so a crash mid-write or a
|
||||
* concurrent reader never observes a half-written file, and through {@link #TRUST_JSON_LOCK} so
|
||||
* two concurrent spawns' entries both survive instead of the second write silently discarding
|
||||
* the first. Both exist because of a real incident: see {@link HerdrPeerLauncher#isProvisionedWorktree}'s javadoc
|
||||
* and {@link #writeAtomically}'s javadoc.
|
||||
*
|
||||
* @param configDir the profile's {@code CLAUDE_CONFIG_DIR} ({@code cfg.configDir()}), or
|
||||
* {@code null}/blank to target the default {@code ~/.claude.json}
|
||||
* @param cwd the spawn's resolved working directory — the exact key Claude Code will look
|
||||
* up for itself once it starts there
|
||||
*/
|
||||
private static void seedTrustDialog(String configDir, String cwd) {
|
||||
if (!isProvisionedWorktree(cwd)) {
|
||||
return;
|
||||
}
|
||||
Path target = (configDir == null || configDir.isBlank())
|
||||
? Path.of(System.getProperty("user.home"), ".claude.json")
|
||||
: Path.of(configDir, ".claude.json");
|
||||
synchronized (TRUST_JSON_LOCK) {
|
||||
try {
|
||||
if (target.getParent() != null) {
|
||||
Files.createDirectories(target.getParent());
|
||||
}
|
||||
ObjectNode root = null;
|
||||
if (Files.isRegularFile(target)) {
|
||||
JsonNode existing = TRUST_JSON.readTree(target.toFile());
|
||||
if (existing instanceof ObjectNode existingObject) {
|
||||
root = existingObject;
|
||||
}
|
||||
}
|
||||
if (root == null) {
|
||||
root = TRUST_JSON.createObjectNode();
|
||||
}
|
||||
JsonNode projectsNode = root.get("projects");
|
||||
ObjectNode projects = projectsNode instanceof ObjectNode projectsObject
|
||||
? projectsObject : TRUST_JSON.createObjectNode();
|
||||
if (!(projectsNode instanceof ObjectNode)) {
|
||||
root.set("projects", projects);
|
||||
}
|
||||
JsonNode projectNode = projects.get(cwd);
|
||||
ObjectNode project = projectNode instanceof ObjectNode projectObject
|
||||
? projectObject : TRUST_JSON.createObjectNode();
|
||||
if (!(projectNode instanceof ObjectNode)) {
|
||||
projects.set(cwd, project);
|
||||
}
|
||||
// fleetd #247: ONLY hasTrustDialogAccepted. We used to write
|
||||
// hasCompletedProjectOnboarding beside it; do not put it back. Measured on
|
||||
// 2026-09-03, minutes after a live spawn seeded this file: 28 of 28 project
|
||||
// entries carried hasTrustDialogAccepted and 0 of 28 carried the onboarding key
|
||||
// — including the 27 entries Claude Code wrote for itself. Claude Code
|
||||
// normalises the whole file when it saves and drops that key every time, so
|
||||
// writing it achieved nothing except making the next reader think it mattered.
|
||||
// The member reached idle with the trust flag alone, which is the only outcome
|
||||
// this seed exists for. If a future Claude Code needs the second flag the
|
||||
// symptom returns as the trust dialog fleetd #149 describes — re-measure then,
|
||||
// do not restore it on a guess.
|
||||
project.put("hasTrustDialogAccepted", true);
|
||||
writeAtomically(target, TRUST_JSON.writerWithDefaultPrettyPrinter().writeValueAsString(root));
|
||||
} catch (Exception e) {
|
||||
log.debug("cannot seed workspace-trust entry for cwd '{}' into '{}'", cwd, target, e);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Write {@code content} to {@code target} atomically: serialise to a sibling temp file in the
|
||||
* <strong>same directory</strong> as {@code target} (an atomic move is only guaranteed within
|
||||
* one filesystem — a different directory could mean a different filesystem), then
|
||||
* {@link StandardCopyOption#ATOMIC_MOVE} it into place. A reader — Claude Code itself, or
|
||||
* another {@code seedTrustDialog} call — only ever observes the fully-old file or the
|
||||
* fully-new one, never a truncated or half-written one.
|
||||
*
|
||||
* <p><b>fleetd #149 incident.</b> The original implementation used
|
||||
* {@code Files.writeString(target, content)} directly, which truncates {@code target} in place
|
||||
* before writing the replacement bytes. Combined with an ungated {@code cwd} (see
|
||||
* {@link HerdrPeerLauncher#isProvisionedWorktree}'s javadoc), a mutation-testing run hit that truncation window
|
||||
* against the operator's real {@code ~/.claude.json} and left it at 178 bytes. The gate closes
|
||||
* <em>which file</em> this can ever target; this closes <em>how</em> the target is written, so
|
||||
* that even a legitimate write against a real, live, concurrently-read {@code .claude.json}
|
||||
* cannot leave it observably empty or partial.
|
||||
*
|
||||
* <p>Preserves {@code target}'s existing POSIX permissions (Claude Code ships {@code
|
||||
* .claude.json} as {@code 0600}) when the filesystem reports them; a freshly created temp file
|
||||
* already defaults to owner-only permissions on a POSIX filesystem, so a first-ever write (no
|
||||
* existing {@code target}) is no less private without this. On a non-POSIX filesystem (e.g.
|
||||
* Windows) the permission copy is a silent no-op rather than a failure.
|
||||
*
|
||||
* <p>Package-visible (not {@code private}) so a test can drive it directly with a concurrent
|
||||
* reader thread and prove the torn-file property this method exists for — the LOCK in
|
||||
* {@link #seedTrustDialog} already fully serialises every call this launcher itself makes, so a
|
||||
* test that only ever goes through {@code seedTrustDialog}/{@code spawn()} could never observe
|
||||
* a torn file regardless of whether this method is atomic; it would be proving the lock, not
|
||||
* this method. Atomicity's actual job is protecting against a writer the lock cannot reach at
|
||||
* all — a second daemon process, or the operator's own live Claude Code — so the test for it
|
||||
* has to reach this method on its own.
|
||||
*/
|
||||
static void writeAtomically(Path target, String content) throws IOException {
|
||||
Path parent = target.getParent();
|
||||
Path tmp = Files.createTempFile(parent, target.getFileName() + ".", ".tmp");
|
||||
try {
|
||||
Files.writeString(tmp, content);
|
||||
copyPosixPermissionsIfPresent(target, tmp);
|
||||
Files.move(tmp, target, StandardCopyOption.ATOMIC_MOVE, StandardCopyOption.REPLACE_EXISTING);
|
||||
} catch (IOException e) {
|
||||
Files.deleteIfExists(tmp);
|
||||
throw e;
|
||||
}
|
||||
}
|
||||
|
||||
/** Copy {@code target}'s POSIX permissions onto {@code tmp}, or no-op where either is unsupported. */
|
||||
private static void copyPosixPermissionsIfPresent(Path target, Path tmp) {
|
||||
try {
|
||||
if (!Files.isRegularFile(target)) {
|
||||
return; // nothing to inherit from — first-ever write, temp file's own default stands
|
||||
}
|
||||
PosixFileAttributeView view = Files.getFileAttributeView(target, PosixFileAttributeView.class);
|
||||
if (view == null) {
|
||||
return; // non-POSIX filesystem — nothing this JVM can read/set here
|
||||
}
|
||||
Files.setPosixFilePermissions(tmp, Files.getPosixFilePermissions(target));
|
||||
} catch (IOException e) {
|
||||
log.debug("cannot preserve permissions of '{}' onto its replacement", target, e);
|
||||
}
|
||||
}
|
||||
|
||||
/** {@code s}, or {@code null} when {@code s} is null/blank — the charter-presence test used above. */
|
||||
private static String nonBlank(String s) {
|
||||
return (s == null || s.isBlank()) ? null : s;
|
||||
|
||||
@@ -10,6 +10,7 @@ import dev.ltms.fleet.peer.PeerHandle;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.PeerUnreachableException;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementCandidate;
|
||||
import dev.ltms.fleet.placement.PlacementContext;
|
||||
@@ -92,6 +93,18 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
/** CB-578 stage B: credential cooldown, checked before an explicit spawn and filtered into placement. */
|
||||
private final BackendQuarantine quarantine;
|
||||
|
||||
/**
|
||||
* fleetd #201 Unit 5: credential cool-off after repeated backend errors, checked before an
|
||||
* explicit spawn and filtered into placement — a SEPARATE, shorter-lived source from
|
||||
* {@link #quarantine}. A never-{@code record}-called instance is naturally inert (its
|
||||
* {@code remainingCoolOffSeconds} always returns empty), so back-compat constructors that predate
|
||||
* this feature share one fixed instance rather than needing a {@code none()} sentinel.
|
||||
*/
|
||||
private final BackendOutagePolicy outagePolicy;
|
||||
|
||||
/** The shared inert instance back-compat constructors wire in — never {@code record}-called. */
|
||||
private static final BackendOutagePolicy NO_OUTAGE_POLICY = new BackendOutagePolicy(System::nanoTime);
|
||||
|
||||
/**
|
||||
* CB-557: the role pools an unqualified spawn draws its candidates from. A supplier that yields
|
||||
* {@code null}, and an empty pool for a role, both fall back to every configured profile — the
|
||||
@@ -130,7 +143,8 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
Map<String, FleetConfig.Profile> profileConfigs,
|
||||
PlacementPolicy placementPolicy,
|
||||
Function<String, Integer> liveCount) {
|
||||
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, null, BackendQuarantine.none());
|
||||
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, null,
|
||||
BackendQuarantine.none(), NO_OUTAGE_POLICY);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -148,13 +162,14 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
PlacementPolicy placementPolicy,
|
||||
Function<String, Integer> liveCount,
|
||||
FleetConfig.Fleet fleet) {
|
||||
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet, BackendQuarantine.none());
|
||||
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet,
|
||||
BackendQuarantine.none(), NO_OUTAGE_POLICY);
|
||||
}
|
||||
|
||||
/**
|
||||
* Production constructor with role pools and quarantine (CB-578 stage B). The full-featured
|
||||
* non-reloading form; {@link #CompositePeerLauncher(List, String, Supplier, Function, BackendQuarantine)}
|
||||
* is what {@code Fleetd.main} actually wires up.
|
||||
* Production constructor with role pools and quarantine (CB-578 stage B), no cool-off (fleetd
|
||||
* #201 Unit 5). Kept for callers that predate the cool-off feature; use the 8-arg overload below
|
||||
* to wire a real {@link BackendOutagePolicy}.
|
||||
*
|
||||
* @param quarantine required — pass {@link BackendQuarantine#none()} for a caller that does not
|
||||
* want the feature, never a defaulting overload (CB-578 stage B's own rule).
|
||||
@@ -166,17 +181,42 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
Function<String, Integer> liveCount,
|
||||
FleetConfig.Fleet fleet,
|
||||
BackendQuarantine quarantine) {
|
||||
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet, quarantine,
|
||||
NO_OUTAGE_POLICY);
|
||||
}
|
||||
|
||||
/**
|
||||
* Production constructor with role pools, quarantine (CB-578 stage B), and cool-off (fleetd
|
||||
* #201 Unit 5). The full-featured non-reloading form;
|
||||
* {@link #CompositePeerLauncher(List, String, Supplier, Function, BackendQuarantine, BackendOutagePolicy)}
|
||||
* is what {@code Fleetd.main} actually wires up.
|
||||
*
|
||||
* @param quarantine required — pass {@link BackendQuarantine#none()} for a caller that does not
|
||||
* want the feature, never a defaulting overload (CB-578 stage B's own rule).
|
||||
* @param outagePolicy required — pass a fresh, never-{@code record}-called {@link
|
||||
* BackendOutagePolicy} for a caller that does not want the feature; same
|
||||
* "explicit opt-out, never a silent default" rule as {@code quarantine}.
|
||||
*/
|
||||
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
|
||||
String defaultProfile,
|
||||
Map<String, FleetConfig.Profile> profileConfigs,
|
||||
PlacementPolicy placementPolicy,
|
||||
Function<String, Integer> liveCount,
|
||||
FleetConfig.Fleet fleet,
|
||||
BackendQuarantine quarantine,
|
||||
BackendOutagePolicy outagePolicy) {
|
||||
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
|
||||
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
|
||||
// differ from one JVM run to the next.
|
||||
this(delegates, defaultProfile,
|
||||
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
|
||||
constant(placementPolicy), liveCount, constant(fleet), quarantine);
|
||||
constant(placementPolicy), liveCount, constant(fleet), quarantine, outagePolicy);
|
||||
}
|
||||
|
||||
/**
|
||||
* Production constructor that re-reads its placement inputs per spawn (CB-559), so a config
|
||||
* reload retargets the next member without a restart.
|
||||
* reload retargets the next member without a restart. No cool-off (fleetd #201 Unit 5); use the
|
||||
* 6-arg overload below to wire a real {@link BackendOutagePolicy}.
|
||||
*
|
||||
* @param config the live configuration — read at every spawn, never captured
|
||||
* @param quarantine required — CB-578 stage B; pass {@link BackendQuarantine#none()} to opt out
|
||||
@@ -186,12 +226,31 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
Supplier<FleetConfig> config,
|
||||
Function<String, Integer> liveCount,
|
||||
BackendQuarantine quarantine) {
|
||||
this(delegates, defaultProfile, config, liveCount, quarantine, NO_OUTAGE_POLICY);
|
||||
}
|
||||
|
||||
/**
|
||||
* Production constructor that re-reads its placement inputs per spawn (CB-559), with cool-off
|
||||
* (fleetd #201 Unit 5). This is what {@code Fleetd.main} actually wires up.
|
||||
*
|
||||
* @param config the live configuration — read at every spawn, never captured
|
||||
* @param quarantine required — CB-578 stage B; pass {@link BackendQuarantine#none()} to opt out
|
||||
* @param outagePolicy required — fleetd #201 Unit 5; pass a fresh, never-{@code record}-called
|
||||
* {@link BackendOutagePolicy} to opt out
|
||||
*/
|
||||
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
|
||||
String defaultProfile,
|
||||
Supplier<FleetConfig> config,
|
||||
Function<String, Integer> liveCount,
|
||||
BackendQuarantine quarantine,
|
||||
BackendOutagePolicy outagePolicy) {
|
||||
this(delegates, defaultProfile,
|
||||
() -> config.get().profiles(),
|
||||
() -> PlacementPolicies.fromName(config.get().placement()),
|
||||
liveCount,
|
||||
() -> config.get().fleet(),
|
||||
quarantine);
|
||||
quarantine,
|
||||
outagePolicy);
|
||||
}
|
||||
|
||||
/** The all-suppliers form every other constructor funnels into. */
|
||||
@@ -201,9 +260,11 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
Supplier<PlacementPolicy> placementPolicy,
|
||||
Function<String, Integer> liveCount,
|
||||
Supplier<FleetConfig.Fleet> fleet,
|
||||
BackendQuarantine quarantine) {
|
||||
BackendQuarantine quarantine,
|
||||
BackendOutagePolicy outagePolicy) {
|
||||
this.fleet = fleet;
|
||||
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
|
||||
this.outagePolicy = Objects.requireNonNull(outagePolicy, "outagePolicy");
|
||||
if (delegates.isEmpty()) {
|
||||
throw new IllegalArgumentException("at least one peer adapter must be configured");
|
||||
}
|
||||
@@ -266,7 +327,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
// the charter makes explicit-profile spawns the normal path — so skipping the check
|
||||
// here would leave the cap dead config in real operation.
|
||||
HerdrPeerLauncher d = route(requestedProfile);
|
||||
// Checked in this order so exhaustion quarantine wins when both are active: quarantine
|
||||
// throws first and short-circuits before the cool-off check ever runs (fleetd #201 Unit 5).
|
||||
enforceNotQuarantined(requestedProfile);
|
||||
enforceNotCoolingOff(requestedProfile);
|
||||
enforceMaxLoad(requestedProfile);
|
||||
PeerHandle handle = d.spawn(req);
|
||||
spawnedBy.put(handle.id(), d);
|
||||
@@ -283,7 +347,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
|
||||
// call, so re-deriving it per retry would only cost work, never change the answer.
|
||||
Set<String> quarantined = quarantinedProfiles(candidates);
|
||||
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable, quarantined);
|
||||
// fleetd #201 Unit 5: a distinct set from quarantined — see PlacementContext.coolingOff.
|
||||
Set<String> coolingOff = coolingOffProfiles(candidates);
|
||||
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable,
|
||||
quarantined, coolingOff);
|
||||
|
||||
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
|
||||
for (int attempt = 0; attempt < maxAttempts; attempt++) {
|
||||
@@ -313,7 +380,8 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
chosen.profile(), e.getMessage());
|
||||
unreachable.add(chosen.profile());
|
||||
// Update the context for the next selection so the policy excludes this profile.
|
||||
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable, quarantined);
|
||||
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable,
|
||||
quarantined, coolingOff);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -379,6 +447,31 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
.collect(Collectors.toSet());
|
||||
}
|
||||
|
||||
/**
|
||||
* Refuse an explicit-profile spawn whose credential is cooling off after repeated backend errors
|
||||
* (fleetd #201 Unit 5 — {@link BackendOutagePolicy}): a SEPARATE, shorter-lived source from
|
||||
* {@link #enforceNotQuarantined}'s exhaustion quarantine. Checked after quarantine so exhaustion
|
||||
* wins when both are active — see the call site in {@link #spawn}.
|
||||
*
|
||||
* @throws PlacementException naming the profile, its credential, and the remaining cool-off
|
||||
*/
|
||||
private void enforceNotCoolingOff(String profile) {
|
||||
String credentialId = credentialIdFor(profile);
|
||||
outagePolicy.remainingCoolOffSeconds(credentialId).ifPresent(remaining -> {
|
||||
throw new PlacementException("worker profile '" + profile + "' credential '" + credentialId
|
||||
+ "' is cooling off after repeated backend errors; ~" + remaining
|
||||
+ "s remaining — refusing spawn");
|
||||
});
|
||||
}
|
||||
|
||||
/** The subset of {@code candidates} whose credential is currently cooling off (fleetd #201 Unit 5). */
|
||||
private Set<String> coolingOffProfiles(List<PlacementCandidate> candidates) {
|
||||
return candidates.stream()
|
||||
.map(PlacementCandidate::profile)
|
||||
.filter(p -> outagePolicy.remainingCoolOffSeconds(credentialIdFor(p)).isPresent())
|
||||
.collect(Collectors.toSet());
|
||||
}
|
||||
|
||||
private void enforceMaxLoad(String profile) {
|
||||
// Absent config, or a config whose maxLoad normalized to null (ABSENT ⇒ unlimited at load),
|
||||
// means no cap — never cap what wasn't configured. Note "non-positive ⇒ unlimited" was true
|
||||
|
||||
@@ -184,7 +184,11 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
*/
|
||||
private final ConcurrentMap<String, Path> zdotdirByPane = new ConcurrentHashMap<>();
|
||||
|
||||
/** Guards {@link #warnNonZsh} to one WARN per launcher instance, not one per spawn. */
|
||||
/**
|
||||
* Guards {@link #warnNonZsh} to one WARN per launcher instance, not one per spawn. fleetd #155:
|
||||
* only the {@code policy: deny-by-default} path still warns on a non-zsh shell — the {@code
|
||||
* policy: allow-list} path refuses the spawn instead (see {@link #applyEnvironmentAllowListPolicy}).
|
||||
*/
|
||||
private final AtomicBoolean nonZshShellWarned = new AtomicBoolean();
|
||||
/** Live config provides URI environment names that must never enter member panes. */
|
||||
private final Supplier<FleetConfig> config;
|
||||
@@ -358,6 +362,42 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
return Files.isRegularFile(candidate) ? candidate : null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether {@code cwd} is a fleetd-provisioned git worktree — signalled the same way
|
||||
* {@code ClaudeCodeLauncher#writeIdeOverlay} already gates on: a {@code .git} that is a
|
||||
* <strong>regular file</strong> holding a {@code gitdir:} pointer, as opposed to a real
|
||||
* checkout's {@code .git} <strong>directory</strong>. {@code null}/blank never qualifies.
|
||||
*
|
||||
* <p>Shared by every write (and, since fleetd #249, every identity read) that must land only
|
||||
* in a worktree fleetd itself created for a member — never in a real checkout, an arbitrary
|
||||
* configured directory, or (see the incident below) the daemon's own fallback cwd. Package-
|
||||
* private (not {@code protected}) on purpose: {@link ClaudeCodeLauncher} and
|
||||
* {@link OpenCodeLauncher} both call it, and same-package visibility is enough — no subclass
|
||||
* outside this package needs it.
|
||||
*
|
||||
* <p><b>fleetd #149 incident.</b> {@code ClaudeCodeLauncher#seedTrustDialog} originally ran
|
||||
* unconditionally on any non-blank {@code cwd}. Most of that launcher's OWN tests spawn a
|
||||
* profile with no {@code cwd} configured, so the base class's {@code resolveCwd} falls
|
||||
* through to the real {@code user.dir} — and with no {@code configDir} either (also the
|
||||
* common case in that file's fixtures), the seed's target falls through the same way to the
|
||||
* real {@code ~/.claude.json}. Running this repo's own test suite corrupted the operator's
|
||||
* actual config file (it shrank from ~72 KB to a single seeded entry) the first time a
|
||||
* mutation happened to make the write non-additive. Gating both cwd-targeted writes on "this
|
||||
* is a worktree fleetd provisioned" — exactly the population fleetd #149 describes
|
||||
* ({@code worktree: true} always lands in a brand-new directory) — makes that class of write
|
||||
* impossible against a real checkout or an untouched fallback cwd, in production or in tests.
|
||||
*
|
||||
* <p><b>fleetd #249.</b> The same reasoning extends to a READ: {@code
|
||||
* OpenCodeSessionDiscovery#sessionIdForDirectory} keys on {@code directory}, a heuristic that
|
||||
* is only reliable when the directory is unique to this member — i.e., exactly the population
|
||||
* this gate identifies. {@link OpenCodeLauncher} uses it to withhold {@code agentSessionId()}
|
||||
* (report absence rather than a guess) and to refuse a {@code resumeSessionId} spawn that
|
||||
* cannot be resolved reliably going forward.
|
||||
*/
|
||||
static boolean isProvisionedWorktree(String cwd) {
|
||||
return cwd != null && !cwd.isBlank() && Files.isRegularFile(Path.of(cwd, ".git"));
|
||||
}
|
||||
|
||||
// --- profile surface -----------------------------------------------------------------------
|
||||
|
||||
/** The configured peer profile names (what {@code spawn(profile)} accepts). */
|
||||
@@ -1219,6 +1259,11 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
if (!creds.isAllowList()) {
|
||||
overlayBlockedCredentials(workerEnv, creds);
|
||||
logCredentialGap(creds, null);
|
||||
// fleetd #155: this overlay itself does not depend on the shell (it lands in the
|
||||
// pane-creation env map before any shell runs), so the spawn is never refused here —
|
||||
// only policy=allow-list's ZDOTDIR scrub needs a login shell to run at all. Still worth
|
||||
// telling the operator: the stronger post-shell control is unavailable on this shell.
|
||||
warnNonZsh(memberLoginShell());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1271,10 +1316,16 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
* dev.ltms.fleet.session.Worktrees#shareWithGroup} already uses, reused rather than
|
||||
* inventing a second group key. Either one missing means the scrub cannot be guaranteed
|
||||
* reachable by the member, which is the same "cannot guarantee the scrub runs" case as a
|
||||
* non-zsh shell, so it gets the identical fallback.</li>
|
||||
* non-zsh shell — that one still gets the fallback (see fleetd #155's javadoc note below
|
||||
* on why a missing worktreeRoot/worktreeGroup is a different defect, out of that ticket's
|
||||
* scope).</li>
|
||||
* </ul>
|
||||
* In every branch: never refuse to spawn. A degraded credential control must not become an
|
||||
* outage for an opt-in feature.
|
||||
* fleetd #155: the one exception to "never refuse to spawn" is a non-zsh login shell, right
|
||||
* below. The operator asked for {@code policy: allow-list} specifically — its whole point is a
|
||||
* control a sourced file cannot undo — so silently degrading to the weaker overlay is exactly
|
||||
* the "silently does nothing" failure this ticket exists to remove. Every other branch in this
|
||||
* method (worktreeRoot/worktreeGroup missing) keeps the old "degrade, never refuse" behaviour;
|
||||
* that gap is real but is fleetd #213's scope, not this one.
|
||||
*/
|
||||
private Path applyEnvironmentAllowListPolicy(FleetConfig.Profile cfg, Launch launch) {
|
||||
FleetConfig.MemberCredentials creds = memberCredentials == null ? null : memberCredentials.get();
|
||||
@@ -1288,20 +1339,21 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
// must not even be called on this path, only the explicit memberLoginShell: config can
|
||||
// answer it. With memberHerdrSocket absent, nothing here changes: fleetd's own $SHELL is
|
||||
// still the input, exactly as before this fix.
|
||||
String loginShell = memberHerdrSocket ? configuredMemberLoginShell() : resolveEnv("SHELL");
|
||||
String loginShell = memberLoginShell();
|
||||
boolean zsh = isZshShell(loginShell);
|
||||
if (!zsh) {
|
||||
// A non-zsh login shell ignores ZDOTDIR entirely: NO scrub would run, so pretending
|
||||
// otherwise would be worse than saying so. Warn loudly and fall back to the CB-596
|
||||
// sentinel overlay over the enumerated known: names — weaker (a sourced file can undo
|
||||
// it), but strictly better than nothing. Deliberately no "allowed N of M" line here: the
|
||||
// scrub this count describes does not run on this path, so printing it would tell an
|
||||
// operator that a fraction of names were blocked when the real number blocked is zero.
|
||||
// logCredentialGap's WARN (below) is the only signal for this path.
|
||||
warnNonZsh(loginShell);
|
||||
overlayBlockedCredentials(launch.env(), creds);
|
||||
logCredentialGap(creds, null);
|
||||
return null;
|
||||
// fleetd #155: a non-zsh login shell ignores ZDOTDIR entirely — NO scrub would run. The
|
||||
// operator explicitly asked for policy=allow-list's blocking control, so degrading to
|
||||
// the weaker CB-596 overlay and spawning anyway would be the same "control silently does
|
||||
// nothing" defect this ticket exists to close. Refuse instead — the caller (FleetMcp.spawn)
|
||||
// catches IllegalArgumentException and surfaces the message to the operator.
|
||||
throw new IllegalArgumentException(
|
||||
"memberCredentials policy=allow-list requires the member's login shell to be "
|
||||
+ "zsh, so the ZDOTDIR scrub can run after it — refusing to spawn under "
|
||||
+ "login shell '" + (loginShell == null ? "<unset>" : loginShell)
|
||||
+ "'. Configure memberLoginShell: as a zsh path (only read when "
|
||||
+ "memberHerdrSocket is set), move the member's OS account onto zsh, or "
|
||||
+ "set memberCredentials.policy: deny-by-default instead.");
|
||||
}
|
||||
Path parentDir;
|
||||
String group = null;
|
||||
@@ -1356,6 +1408,20 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
return cfg == null ? null : cfg.memberLoginShell();
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #155: the member pane's login shell, using the same {@code memberHerdrSocket} routing
|
||||
* {@link #applyEnvironmentAllowListPolicy} already used — fleetd's own {@code $SHELL} decides it
|
||||
* when {@code memberHerdrSocket} is absent (today's only mode); the configured {@code
|
||||
* memberLoginShell:} decides it otherwise, since fleetd's own {@code $SHELL} names a different
|
||||
* user's shell once member panes run under a different OS user. Shared by both {@link
|
||||
* #applyEnvironmentAllowListPolicy} (allow-list: refuses on non-zsh) and {@link
|
||||
* #applyMemberCredentialPolicy} (deny-by-default: warns on non-zsh) so the two policies agree on
|
||||
* what "the member's shell" means.
|
||||
*/
|
||||
private String memberLoginShell() {
|
||||
return memberHerdrSocketConfigured() ? configuredMemberLoginShell() : resolveEnv("SHELL");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #213: {@code worktreeRoot}, as the ZDOTDIR scrub's parent directory under {@code
|
||||
* memberHerdrSocket}, or {@code null} when unconfigured — the same "cannot guarantee the scrub
|
||||
@@ -1440,15 +1506,23 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
private static final String SSH_AUTH_SOCK = MemberEnvAllowList.SSH_AUTH_SOCK;
|
||||
|
||||
/**
|
||||
* CB-633: a non-zsh login shell means the allow-list control CANNOT run — say so once per
|
||||
* launcher instance, naming the shell, instead of failing silently.
|
||||
* fleetd #155: called ONLY from the {@code policy: deny-by-default} branch of {@link
|
||||
* #applyMemberCredentialPolicy} — under {@code policy: allow-list}, a non-zsh shell now REFUSES
|
||||
* the spawn instead (see {@link #applyEnvironmentAllowListPolicy}), since that policy's whole
|
||||
* point is a control a sourced file cannot undo. deny-by-default's own overlay does not depend
|
||||
* on the shell (it lands in the pane-creation env map before any shell runs), so this is a
|
||||
* heads-up, not a defect report: say once per launcher instance, naming the shell, that the
|
||||
* stronger allow-list control is unavailable here — never silently.
|
||||
*/
|
||||
private void warnNonZsh(String shell) {
|
||||
if (isZshShell(shell)) {
|
||||
return;
|
||||
}
|
||||
if (nonZshShellWarned.compareAndSet(false, true)) {
|
||||
log.warn("memberCredentials policy=allow-list: member login shell '{}' is NOT zsh — "
|
||||
+ "ZDOTDIR scrubbing cannot run, so members' inherited environment is "
|
||||
+ "UNPROTECTED beyond the enumerated known: fallback. Move herdr onto a "
|
||||
+ "zsh account or switch policy back to deny-by-default.",
|
||||
log.warn("memberCredentials policy=deny-by-default: member login shell '{}' is NOT zsh — "
|
||||
+ "the pane-creation credential overlay still applies here (it does not "
|
||||
+ "depend on the shell), but policy=allow-list's stronger post-shell "
|
||||
+ "ZDOTDIR scrub is unavailable on this shell.",
|
||||
shell == null ? "<unset>" : shell);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,75 @@
|
||||
package dev.ltms.fleet.member;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* fleetd #111 (CB-608): a single, testable read of the {@code memberCredentials:} policy — names
|
||||
* and counts only, never a value. The daemon never holds a credential's <em>value</em> in the
|
||||
* first place (only the names configured under {@code known:}/{@code allow:}), so there is
|
||||
* nothing here to redact by construction; the point of this class is that it is the ONE place
|
||||
* that turns a policy into names-and-counts, so nothing else hand-counts a second time.
|
||||
*
|
||||
* <p>Before this class, {@link dev.ltms.fleet.Fleetd#reportMemberCredentialsGap} computed these
|
||||
* same counts inline for the startup log line, and {@code scripts/probe-member-credentials.sh}
|
||||
* carried its own hardcoded {@code NAMES} array that the live policy could grow past silently
|
||||
* (#111) — the exact "hand-maintained second copy drifts" shape #114 fixed for the tool
|
||||
* catalogue. Both now read this class: the startup log via {@link
|
||||
* dev.ltms.fleet.Fleetd#reportMemberCredentialsGap}, and a live daemon via the {@code
|
||||
* GET /member-credentials} REST endpoint ({@link dev.ltms.fleet.rest.FleetApp}), which the probe
|
||||
* script fetches instead of carrying its own list.
|
||||
*
|
||||
* @param present policy configured with at least one {@code known} name. {@code false} for an
|
||||
* absent or empty {@code memberCredentials:} block — represented honestly as "no
|
||||
* policy", never as "nothing blocked" (an empty {@link #blocked} could otherwise be
|
||||
* misread as a clean bill of health).
|
||||
* @param policy the normalized policy mode ({@link FleetConfig.MemberCredentials#policy()}), or
|
||||
* {@code null} when {@link #present} is {@code false}.
|
||||
* @param known every name the policy declares, in configured order. Names only, never a value.
|
||||
* @param allowed the subset of {@link #known} explicitly let through. Names only.
|
||||
* @param blocked {@link #known} minus {@link #allowed} — the names an actual spawn shadows. Names
|
||||
* only.
|
||||
*/
|
||||
public record MemberCredentialPolicyView(boolean present, String policy, List<String> known,
|
||||
List<String> allowed, List<String> blocked) {
|
||||
|
||||
private static final MemberCredentialPolicyView ABSENT =
|
||||
new MemberCredentialPolicyView(false, null, List.of(), List.of(), List.of());
|
||||
|
||||
public MemberCredentialPolicyView {
|
||||
known = known == null ? List.of() : List.copyOf(known);
|
||||
allowed = allowed == null ? List.of() : List.copyOf(allowed);
|
||||
blocked = blocked == null ? List.of() : List.copyOf(blocked);
|
||||
}
|
||||
|
||||
/** The honest "no policy configured" view. */
|
||||
public static MemberCredentialPolicyView absent() {
|
||||
return ABSENT;
|
||||
}
|
||||
|
||||
/**
|
||||
* Build the view straight from the live config. {@code creds} may be {@code null} (no {@code
|
||||
* memberCredentials:} block at all) — treated the same as a present-but-empty block, exactly
|
||||
* like {@link dev.ltms.fleet.Fleetd#reportMemberCredentialsGap} already did.
|
||||
*/
|
||||
public static MemberCredentialPolicyView of(FleetConfig.MemberCredentials creds) {
|
||||
if (creds == null || creds.known().isEmpty()) {
|
||||
return ABSENT;
|
||||
}
|
||||
return new MemberCredentialPolicyView(true, creds.policy(), creds.known(), creds.allow(),
|
||||
List.copyOf(creds.blockedSet()));
|
||||
}
|
||||
|
||||
public int knownCount() {
|
||||
return known.size();
|
||||
}
|
||||
|
||||
public int allowedCount() {
|
||||
return allowed.size();
|
||||
}
|
||||
|
||||
public int blockedCount() {
|
||||
return blocked.size();
|
||||
}
|
||||
}
|
||||
@@ -22,6 +22,7 @@ import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.atomic.AtomicBoolean;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
import java.util.function.BooleanSupplier;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.LongSupplier;
|
||||
@@ -671,13 +672,29 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
/** Add lazy on-disk session discovery to the base handle. */
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req) {
|
||||
String cwd = effectiveCwd(req);
|
||||
// fleetd #249: refuse rather than silently resume into unverifiable territory. opencode's
|
||||
// `-s <id>` flag itself resumes precisely — the resolved id is what fails, not the resume —
|
||||
// but resolvedSessionId() below can never confirm (or later re-report) this handle's own
|
||||
// identity without a fleetd-provisioned worktree (isProvisionedWorktree(cwd)), because the
|
||||
// directory is shared and sessionIdForDirectory's "most recently updated row" heuristic can
|
||||
// pick a sibling's session. Refusing here, before anything spawns, beats letting the member
|
||||
// start and only then discovering fleetd can never again verify who it actually is.
|
||||
if (req.resumeSessionId() != null && !req.resumeSessionId().isBlank()
|
||||
&& !isProvisionedWorktree(cwd)) {
|
||||
throw new IllegalArgumentException("resumeSessionId requires a fleetd-provisioned "
|
||||
+ "worktree for an opencode profile — without one, this member's cwd is shared "
|
||||
+ "with other sessions, so fleetd can never reliably confirm (now or later) which "
|
||||
+ "conversation it is actually running (fleetd #249). Pass fleet_spawn{worktree:"
|
||||
+ "<ticket-slug>} to resume this member.");
|
||||
}
|
||||
PeerHandle inner = super.spawn(req);
|
||||
// fleetd #175: the same profile config buildLaunch resolved for this spawn (requireProfile
|
||||
// is deterministic on req.profileName(), so re-resolving here costs a map lookup, not a
|
||||
// second decision) — SessionAwareHandle needs cfg.model() to know what THIS session should
|
||||
// be running.
|
||||
FleetConfig.Profile cfg = requireProfile(req.profileName());
|
||||
return new SessionAwareHandle(inner, discovery, effectiveCwd(req), cfg,
|
||||
return new SessionAwareHandle(inner, discovery, cwd, cfg,
|
||||
this::memberHerdrSocketConfigured, discoveryUnavailableWarned, exhaustionSink);
|
||||
}
|
||||
|
||||
@@ -706,6 +723,28 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
private final ExhaustionSink exhaustionSink;
|
||||
/** CAS'd true the first (and only) time a model mismatch is reported for this handle. */
|
||||
private final AtomicBoolean modelMismatchReported = new AtomicBoolean();
|
||||
/**
|
||||
* The session id, once {@link OpenCodeSessionDiscovery#sessionIdForDirectory} first
|
||||
* resolves a non-null answer for this handle (fleetd #234). Sticky on purpose: {@code
|
||||
* directory} is a shared-cwd heuristic (see {@link OpenCodeSessionDiscovery}'s class
|
||||
* javadoc) that can start returning a DIFFERENT row once another session shares the same
|
||||
* directory and writes a newer one — re-deriving it on every call would let this handle's
|
||||
* identity silently drift to a sibling's session. Once resolved, this IS the answer, and
|
||||
* {@link #checkModelMatch} reads only the row this id names, never "whatever is newest in
|
||||
* the directory right now."
|
||||
*/
|
||||
private final AtomicReference<String> resolvedSessionId = new AtomicReference<>();
|
||||
/**
|
||||
* fleetd #249: whether {@link #cwd} is a fleetd-provisioned git worktree
|
||||
* ({@link HerdrPeerLauncher#isProvisionedWorktree}), computed once at spawn time since
|
||||
* {@code cwd} never changes for this handle. When {@code false} the directory is shared
|
||||
* with other sessions (the default no-worktree spawn inherits the lead's own cwd), so
|
||||
* {@link OpenCodeSessionDiscovery#sessionIdForDirectory}'s "most recently updated row for
|
||||
* this directory" heuristic can and does pick another session's row — see that class's
|
||||
* javadoc. {@link #agentSessionId()} refuses to guess in that case: it reports absent
|
||||
* rather than a possibly-foreign id.
|
||||
*/
|
||||
private final boolean worktreeProvisioned;
|
||||
|
||||
SessionAwareHandle(PeerHandle delegate, OpenCodeSessionDiscovery discovery, String cwd,
|
||||
FleetConfig.Profile cfg,
|
||||
@@ -719,6 +758,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
this.discoveryUnavailable = discoveryUnavailable;
|
||||
this.discoveryUnavailableWarned = discoveryUnavailableWarned;
|
||||
this.exhaustionSink = exhaustionSink;
|
||||
this.worktreeProvisioned = isProvisionedWorktree(cwd);
|
||||
}
|
||||
|
||||
@Override
|
||||
@@ -748,6 +788,9 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
// built from (see OpenCodeLauncher#defaultDiscoveryRoot's javadoc for the full
|
||||
// reasoning). Scanning fleetd's own $HOME under that config would only ever find "no
|
||||
// row" and read as "resume unsupported" — declare it unavailable instead, once, loudly.
|
||||
// Checked before the fleetd #249 worktree gate below: this OS-user mismatch makes
|
||||
// discovery unusable regardless of whether cwd happens to be a provisioned worktree, so
|
||||
// it earns the one-time WARN either way.
|
||||
if (discoveryUnavailable.getAsBoolean()) {
|
||||
if (discoveryUnavailableWarned.compareAndSet(false, true)) {
|
||||
log.warn("opencode session discovery unavailable: memberHerdrSocket is "
|
||||
@@ -759,17 +802,37 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
}
|
||||
return null;
|
||||
}
|
||||
// fleetd #249: cwd is shared with other sessions unless fleetd itself provisioned this
|
||||
// worktree, and sessionIdForDirectory's directory-keyed heuristic cannot tell this
|
||||
// member's row apart from a sibling's in that case (measured: a three-day-old row from
|
||||
// a different profile). Refuse to guess — absent is the honest answer, and it is what
|
||||
// this codebase already returns elsewhere for absent evidence (fleetd #175's UNKNOWN).
|
||||
// No WARN here: unlike discoveryUnavailable above, this is the ordinary, expected shape
|
||||
// of the large majority of spawns (no worktree requested), not a configuration gap.
|
||||
if (!worktreeProvisioned) {
|
||||
return null;
|
||||
}
|
||||
// fleetd #234: once resolved, stay resolved. Re-deriving from `directory` on every call
|
||||
// would let this handle's identity drift to a sibling session that later shares the
|
||||
// same cwd and writes a newer row — see resolvedSessionId's javadoc.
|
||||
String cached = resolvedSessionId.get();
|
||||
if (cached != null) {
|
||||
return cached;
|
||||
}
|
||||
// Lazy + retried, never a spawn-time blocker: opencode writes the session record only
|
||||
// when the session is first persisted, so null here is the correct interim answer and
|
||||
// the caller re-calls later (each call re-scans, picking up a record that has since
|
||||
// appeared).
|
||||
String id = discovery.sessionIdForDirectory(cwd);
|
||||
if (id != null) {
|
||||
resolvedSessionId.compareAndSet(null, id);
|
||||
}
|
||||
// fleetd #175: check on the SAME tick — while the caller (SessionManager's late-resolve
|
||||
// step) is still re-polling because the id is unknown, the row this id came from (once
|
||||
// it exists) is exactly the row that also carries the actual model. Once id resolves,
|
||||
// the caller stops calling agentSessionId() for this session, so this is naturally a
|
||||
// once-only check that happens right when the row first appears.
|
||||
checkModelMatch();
|
||||
checkModelMatch(id);
|
||||
return id;
|
||||
}
|
||||
|
||||
@@ -777,16 +840,25 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
* Verify the live opencode session is running the model {@link #cfg} requested (fleetd
|
||||
* #175) and, on a real mismatch, log an ERROR and quarantine through {@link
|
||||
* #exhaustionSink}. A no-op when there is nothing to compare against — no model configured,
|
||||
* already reported once for this handle, or the actual model is still UNKNOWN (no row yet,
|
||||
* unreadable database, or unparseable evidence). UNKNOWN must never be treated as a
|
||||
* mismatch: that is the single most important safety rule here — a false positive would
|
||||
* already reported once for this handle, {@code sessionId} itself is not resolved yet
|
||||
* (fleetd #234: absent evidence, not a mismatch), or the actual model is still UNKNOWN (no
|
||||
* row yet, unreadable database, or unparseable evidence). UNKNOWN must never be treated as
|
||||
* a mismatch: that is the single most important safety rule here — a false positive would
|
||||
* quarantine a perfectly working profile's credential.
|
||||
*
|
||||
* @param sessionId the id {@link #agentSessionId()} just resolved (or had cached) for THIS
|
||||
* handle — the model is read back for this exact session (fleetd #234's
|
||||
* {@link OpenCodeSessionDiscovery#actualModelForSessionId}), never
|
||||
* re-derived from {@code directory}
|
||||
*/
|
||||
private void checkModelMatch() {
|
||||
private void checkModelMatch(String sessionId) {
|
||||
if (modelMismatchReported.get() || cfg.model() == null || cfg.model().isBlank()) {
|
||||
return;
|
||||
}
|
||||
OpenCodeSessionDiscovery.ActualModel actual = discovery.actualModelForDirectory(cwd);
|
||||
if (sessionId == null || sessionId.isBlank()) {
|
||||
return; // id not resolved yet — UNKNOWN, never a mismatch (fleetd #175's rule)
|
||||
}
|
||||
OpenCodeSessionDiscovery.ActualModel actual = discovery.actualModelForSessionId(sessionId);
|
||||
if (actual == null) {
|
||||
return; // UNKNOWN evidence — never a mismatch
|
||||
}
|
||||
@@ -817,10 +889,16 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
+ "falls back to a default model, which may be a PAID credential "
|
||||
+ "(fleetd #175); quarantining this profile's credential",
|
||||
cfg.profile(), cfg.model(), actualDisplay);
|
||||
// fleetd #234: pass our OWN profile name too. This check fires from agentSessionId(),
|
||||
// called during SessionManager.acquire() BEFORE this session is registered in
|
||||
// sessions.roster() — a roster-only sink (Fleetd's target -> session -> profile lookup)
|
||||
// finds nothing at this point and silently no-ops (defect 2). We already know exactly
|
||||
// which profile to quarantine without the roster; the sink is passed it explicitly.
|
||||
exhaustionSink.onExhausted(delegate.terminalId(),
|
||||
"opencode model mismatch: profile '" + cfg.profile() + "' requested '"
|
||||
+ cfg.model() + "' but the live session is running '" + actualDisplay
|
||||
+ "' (fleetd #175)");
|
||||
+ "' (fleetd #175)",
|
||||
cfg.profile());
|
||||
}
|
||||
|
||||
@Override
|
||||
|
||||
@@ -32,11 +32,19 @@ import java.util.concurrent.atomic.AtomicBoolean;
|
||||
* here, so a layout change, or a switch to the HTTP server, changes exactly one class and nothing
|
||||
* in {@link OpenCodeLauncher}.
|
||||
*
|
||||
* <p>The determinism that makes this useful is structural, not a guess: every fleetd worker runs
|
||||
* in its own unique git worktree, so the row's {@code directory} (its project root) equals the
|
||||
* worker's cwd identifies <em>its</em> session unambiguously. We match on {@code directory} rather
|
||||
* than diffing {@code opencode session list} before/after — that races under concurrent spawns, and
|
||||
* the CLI listing does not even show the directory.
|
||||
* <p><strong>The {@code directory} match is a heuristic, not an identity — fleetd #234.</strong> A
|
||||
* worktree is opt-in: {@code fleet_spawn} only provisions one when the caller passes {@code
|
||||
* worktree:}; the default spawn inherits the lead's own cwd, which every other worker spawned the
|
||||
* same way (and every past session ever run there) shares. {@code directory} therefore does
|
||||
* <em>not</em> identify a session unambiguously in general — only in the special case of a fresh,
|
||||
* unique worktree does "most recently updated row for this directory" reliably mean "this worker's
|
||||
* own row." {@link #sessionIdForDirectory} still has to use this heuristic (the id has to come from
|
||||
* somewhere, and nothing else is available at this layer — see that method's javadoc), but a caller
|
||||
* that already holds a resolved id must never re-derive evidence about that same session via
|
||||
* {@code directory} again; see {@link #actualModelForSessionId}, which looks up by {@code id}
|
||||
* instead for exactly this reason. We match on {@code directory} rather than diffing
|
||||
* {@code opencode session list} before/after — that races under concurrent spawns, and the CLI
|
||||
* listing does not even show the directory.
|
||||
*
|
||||
* <p>All reads are best-effort and never throw: a missing or unreadable database, a query that
|
||||
* fails, or a directory with no row yet all yield {@code null}, and the caller (the session
|
||||
@@ -89,8 +97,16 @@ final class OpenCodeSessionDiscovery {
|
||||
/**
|
||||
* The opencode session id whose row references {@code directory} (the worker's cwd), or
|
||||
* {@code null} when no row matches yet. When several rows share the directory — e.g. repeated
|
||||
* spawns into the same worktree — the row with the highest {@code time_updated} wins: it is
|
||||
* the session the pane most likely corresponds to.
|
||||
* spawns into the same worktree, OR several workers sharing one cwd because none of them was
|
||||
* given a worktree (fleetd #234) — the row with the highest {@code time_updated} wins: it is
|
||||
* the session the pane most likely corresponds to. That "most likely" is a real caveat, not a
|
||||
* formality: when the directory is shared, this can and does pick another session's row (see
|
||||
* the class javadoc). This is the one place fleetd resolves an opencode session id at all —
|
||||
* nothing else is available at this layer to disambiguate further (no {@code opencode session
|
||||
* list} entry names the directory, and diffing before/after races under concurrent spawns) — so
|
||||
* the heuristic stays here unchanged. What must never happen is a SECOND, independent piece of
|
||||
* evidence about the same session being re-derived via {@code directory} once an id has already
|
||||
* come out of this method; see {@link #actualModelForSessionId}.
|
||||
*
|
||||
* <p>Never throws: a missing {@code opencode.db}, a locked/unreadable database, a query
|
||||
* failure, or a directory that has not been persisted yet all resolve to {@code null} rather
|
||||
@@ -134,24 +150,36 @@ final class OpenCodeSessionDiscovery {
|
||||
}
|
||||
|
||||
/**
|
||||
* The model opencode actually ran the {@code directory}'s most-recent session on (fleetd
|
||||
* #175), read from the same row {@link #sessionIdForDirectory} matches — but via its own
|
||||
* query and its own connection, deliberately kept independent so a database whose schema
|
||||
* predates the {@code model} column (or any other read failure on this column alone) can
|
||||
* never take {@link #sessionIdForDirectory}'s id resolution down with it. That would be a
|
||||
* regression of the id-resolution feature #209 shipped; this method degrades on its own.
|
||||
* The model opencode actually ran the session {@code sessionId} on (fleetd #175/#234), read by
|
||||
* primary-key lookup — the ONE row that id names, and no other. Deliberately keyed on
|
||||
* {@code id} rather than {@code directory}: two independent {@code WHERE directory = ? ORDER BY
|
||||
* time_updated DESC LIMIT 1} queries (one for the id, one for the model) can each pick a
|
||||
* DIFFERENT row once more than one session shares a directory (fleetd #234 — the default
|
||||
* no-worktree spawn shares the lead's cwd with every other worker and every past session ever
|
||||
* run there), silently comparing a profile's requested model against a session that is not even
|
||||
* the one whose id was returned. Keying on {@code id} instead makes that impossible: the model
|
||||
* read back is always the SAME session {@link #sessionIdForDirectory} (or a cached copy of its
|
||||
* answer) already resolved.
|
||||
*
|
||||
* <p>Kept as its own query and its own connection, independent from {@link
|
||||
* #sessionIdForDirectory}: a database whose schema predates the {@code model} column (or any
|
||||
* other read failure on this column alone) can never take id resolution down with it. That
|
||||
* would be a regression of the id-resolution feature #209 shipped; this method degrades on its
|
||||
* own.
|
||||
*
|
||||
* <p>Never throws, and every failure mode — no matching row, a missing/unreadable database, a
|
||||
* missing {@code model} column, a null/blank {@code model} value, or JSON that does not parse
|
||||
* into {@code {"id": "...", "providerID": "..."}} with a non-blank {@code id} — resolves to
|
||||
* {@code null}. That is UNKNOWN evidence, not a mismatch signal: the caller must never
|
||||
* quarantine a profile on the strength of a {@code null} here.
|
||||
* quarantine a profile on the strength of a {@code null} here. A blank/null {@code sessionId}
|
||||
* (the id is not resolved yet) is UNKNOWN too, for the same reason — never call this with one.
|
||||
*
|
||||
* @param directory the worker's cwd, as resolved for this spawn
|
||||
* @param sessionId the session id already resolved by {@link #sessionIdForDirectory} for this
|
||||
* spawn — never re-derived from {@code directory} here
|
||||
* @return the actual model, or {@code null} when unknown
|
||||
*/
|
||||
ActualModel actualModelForDirectory(String directory) {
|
||||
if (directory == null || directory.isBlank()) {
|
||||
ActualModel actualModelForSessionId(String sessionId) {
|
||||
if (sessionId == null || sessionId.isBlank()) {
|
||||
return null;
|
||||
}
|
||||
if (!Files.isRegularFile(databasePath)) {
|
||||
@@ -159,10 +187,10 @@ final class OpenCodeSessionDiscovery {
|
||||
// exact condition — do not double-log it here.
|
||||
return null;
|
||||
}
|
||||
String sql = "SELECT model FROM session WHERE directory = ? ORDER BY time_updated DESC LIMIT 1";
|
||||
String sql = "SELECT model FROM session WHERE id = ?";
|
||||
try (Connection connection = openReadOnly();
|
||||
PreparedStatement statement = connection.prepareStatement(sql)) {
|
||||
statement.setString(1, directory);
|
||||
statement.setString(1, sessionId);
|
||||
try (ResultSet rows = statement.executeQuery()) {
|
||||
if (rows.next()) {
|
||||
return parseModel(rows.getString("model"));
|
||||
|
||||
@@ -746,7 +746,7 @@ public final class MessageService {
|
||||
if (task != null) {
|
||||
asyncTasksByWaiter.put(reply, task);
|
||||
}
|
||||
TurnToken token = new TurnToken(target, reply);
|
||||
TurnToken token = new TurnToken(target, reply, content);
|
||||
// The send has won the lock; the accepted-delivery hook records delegator ownership
|
||||
// here (CB-548). It runs BEFORE enqueue so a throwing hook — onAccepted is now a
|
||||
// public callback — fails the send without queuing a message that would orphan.
|
||||
|
||||
@@ -9,8 +9,10 @@ import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.Collection;
|
||||
import java.util.HashSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
@@ -68,6 +70,12 @@ public final class ReplyPushLoop {
|
||||
static final String QUESTIONS_NUDGE_FORMAT =
|
||||
"%d workers are paused on a question — run fleet_poll(ticket=...) for each, then answer "
|
||||
+ "with fleet_send(turnId=..., content=...): %s";
|
||||
static final String BACKEND_INCIDENT_NUDGE_FORMAT =
|
||||
"Backend credential %s is cooling for %d remaining seconds; affected profiles: %s; "
|
||||
+ "affected workers: %s. Run fleet_list to see more.";
|
||||
static final String BACKEND_TARGET_UNMAPPED_NUDGE_FORMAT =
|
||||
"A backend error on worker %s could not be mapped to a credential (%s) — no cool-off "
|
||||
+ "was applied. Run fleet_list to check that worker.";
|
||||
|
||||
private final PrimaryRegistry primaryRegistry;
|
||||
private final AgentControl agents;
|
||||
@@ -93,6 +101,13 @@ public final class ReplyPushLoop {
|
||||
* how depleted an older, still-open question's count is.
|
||||
*/
|
||||
private final ConcurrentHashMap<String, PendingQuestion> pendingQuestions = new ConcurrentHashMap<>();
|
||||
/** Backend incidents awaiting one successful delivery, keyed by incident id and owning lead. */
|
||||
private final ConcurrentHashMap<IncidentLead, PendingIncident> pendingIncidents = new ConcurrentHashMap<>();
|
||||
/** Incident/lead pairs already delivered. They make repeated incident reports one-shot. */
|
||||
private final Set<IncidentLead> deliveredIncidents = ConcurrentHashMap.newKeySet();
|
||||
private final ConcurrentHashMap<UnmappedTargetLead, PendingUnmappedTarget> pendingUnmappedTargets =
|
||||
new ConcurrentHashMap<>();
|
||||
private final Set<UnmappedTargetLead> deliveredUnmappedTargets = ConcurrentHashMap.newKeySet();
|
||||
/** CB-590: leads with an active combined reminder schedule (replies and/or tickets and/or questions). */
|
||||
private final ConcurrentHashMap<String, Boolean> activeLeads = new ConcurrentHashMap<>();
|
||||
|
||||
@@ -178,7 +193,20 @@ public final class ReplyPushLoop {
|
||||
* tracked per question, not per lead per source).
|
||||
*/
|
||||
private record PendingQuestion(String turnId, String ticket, String target, String lead,
|
||||
String question, int nudgeCount) {
|
||||
String question, int nudgeCount) {
|
||||
}
|
||||
|
||||
private record IncidentLead(String incidentId, String lead) {
|
||||
}
|
||||
|
||||
private record PendingIncident(IncidentLead key, String credential, List<String> profiles,
|
||||
List<String> targets, int remainingCoolOffSeconds, int nudgeCount) {
|
||||
}
|
||||
|
||||
private record UnmappedTargetLead(String target, String reason, String lead) {
|
||||
}
|
||||
|
||||
private record PendingUnmappedTarget(UnmappedTargetLead key, int nudgeCount) {
|
||||
}
|
||||
|
||||
/** Questions still open for {@code lead}, snapshotted fresh for one tick. */
|
||||
@@ -192,6 +220,24 @@ public final class ReplyPushLoop {
|
||||
.collect(Collectors.toUnmodifiableSet());
|
||||
}
|
||||
|
||||
private List<PendingIncident> pendingIncidentsFor(String lead) {
|
||||
return pendingIncidents.values().stream().filter(i -> lead.equals(i.key().lead())).toList();
|
||||
}
|
||||
|
||||
private Set<IncidentLead> pendingIncidentKeysFor(String lead) {
|
||||
return pendingIncidentsFor(lead).stream().map(PendingIncident::key)
|
||||
.collect(Collectors.toUnmodifiableSet());
|
||||
}
|
||||
|
||||
private List<PendingUnmappedTarget> pendingUnmappedTargetsFor(String lead) {
|
||||
return pendingUnmappedTargets.values().stream().filter(i -> lead.equals(i.key().lead())).toList();
|
||||
}
|
||||
|
||||
private Set<UnmappedTargetLead> pendingUnmappedTargetKeysFor(String lead) {
|
||||
return pendingUnmappedTargetsFor(lead).stream().map(PendingUnmappedTarget::key)
|
||||
.collect(Collectors.toUnmodifiableSet());
|
||||
}
|
||||
|
||||
/**
|
||||
* The reply-source reminder count {@link #decide} should see for {@code lead} on this tick:
|
||||
* the <em>minimum</em> nudge count among the reply targets currently pending for it (CB-598).
|
||||
@@ -236,6 +282,15 @@ public final class ReplyPushLoop {
|
||||
return min == Integer.MAX_VALUE ? 0 : min;
|
||||
}
|
||||
|
||||
private int minIncidentNudgeCountFor(String lead) {
|
||||
return pendingIncidentsFor(lead).stream().mapToInt(PendingIncident::nudgeCount).min().orElse(0);
|
||||
}
|
||||
|
||||
private int minUnmappedTargetNudgeCountFor(String lead) {
|
||||
return pendingUnmappedTargetsFor(lead).stream().mapToInt(PendingUnmappedTarget::nudgeCount)
|
||||
.min().orElse(0);
|
||||
}
|
||||
|
||||
/**
|
||||
* Pure decision function: examine everything pending for {@code lead} — reply targets and
|
||||
* tickets alike — and return what the loop should do.
|
||||
@@ -272,17 +327,35 @@ public final class ReplyPushLoop {
|
||||
* @param questionReminderCount the lowest nudge count among questions open for this lead
|
||||
*/
|
||||
Action decide(String lead, int replyReminderCount, int ticketReminderCount, int questionReminderCount) {
|
||||
return decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount,
|
||||
minIncidentNudgeCountFor(lead));
|
||||
}
|
||||
|
||||
/** As above, with backend incidents as a fourth, independently bounded source. */
|
||||
Action decide(String lead, int replyReminderCount, int ticketReminderCount, int questionReminderCount,
|
||||
int incidentReminderCount) {
|
||||
return decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount,
|
||||
incidentReminderCount, minUnmappedTargetNudgeCountFor(lead));
|
||||
}
|
||||
|
||||
/** As above, with unmapped backend targets as a fifth, independently bounded source. */
|
||||
Action decide(String lead, int replyReminderCount, int ticketReminderCount, int questionReminderCount,
|
||||
int incidentReminderCount, int unmappedTargetReminderCount) {
|
||||
boolean hasReplyWork = !pendingReplyTargetsFor(lead).isEmpty();
|
||||
boolean hasTicketWork = !pendingTicketIdsFor(lead).isEmpty();
|
||||
boolean hasQuestionWork = !pendingQuestionTurnIdsFor(lead).isEmpty();
|
||||
if (!hasReplyWork && !hasTicketWork && !hasQuestionWork) {
|
||||
boolean hasIncidentWork = !pendingIncidentKeysFor(lead).isEmpty();
|
||||
boolean hasUnmappedTargetWork = !pendingUnmappedTargetKeysFor(lead).isEmpty();
|
||||
if (!hasReplyWork && !hasTicketWork && !hasQuestionWork && !hasIncidentWork && !hasUnmappedTargetWork) {
|
||||
log.debug("push: nothing pending for lead {}, stopping reminder", lead);
|
||||
return Action.STOP;
|
||||
}
|
||||
boolean replyEligible = hasReplyWork && replyReminderCount < maxReminders;
|
||||
boolean ticketEligible = hasTicketWork && ticketReminderCount < maxReminders;
|
||||
boolean questionEligible = hasQuestionWork && questionReminderCount < maxReminders;
|
||||
if (!replyEligible && !ticketEligible && !questionEligible) {
|
||||
boolean incidentEligible = hasIncidentWork && incidentReminderCount < maxReminders;
|
||||
boolean unmappedTargetEligible = hasUnmappedTargetWork && unmappedTargetReminderCount < maxReminders;
|
||||
if (!replyEligible && !ticketEligible && !questionEligible && !incidentEligible && !unmappedTargetEligible) {
|
||||
log.debug("push: reminder cap ({}) reached for lead {} on every source with pending work, stopping",
|
||||
maxReminders, lead);
|
||||
countNudge("exhausted");
|
||||
@@ -394,6 +467,48 @@ public final class ReplyPushLoop {
|
||||
pendingQuestions.remove(turnId);
|
||||
}
|
||||
|
||||
/**
|
||||
* Queue a one-shot backend credential outage notice for every distinct lead that owns an
|
||||
* affected worker. A successful injection records its {@code (incidentId, lead)} key, so a
|
||||
* repeat report never reminds that lead again.
|
||||
*/
|
||||
public void onBackendIncident(String incidentId, Collection<String> targets, String credential,
|
||||
Collection<String> profiles, int remainingCoolOffSeconds) {
|
||||
Map<String, List<String>> targetsByLead = new ConcurrentHashMap<>();
|
||||
for (String target : targets) {
|
||||
var lead = primaryRegistry.nudgeTargetFor(target);
|
||||
if (lead.isEmpty()) {
|
||||
log.warn("push: backend incident {} has no known lead for target {}", incidentId, target);
|
||||
continue;
|
||||
}
|
||||
targetsByLead.computeIfAbsent(lead.get(), _ -> new ArrayList<>()).add(target);
|
||||
}
|
||||
List<String> profileNames = profiles.stream().sorted().toList();
|
||||
for (var entry : targetsByLead.entrySet()) {
|
||||
IncidentLead key = new IncidentLead(incidentId, entry.getKey());
|
||||
if (deliveredIncidents.contains(key)) continue;
|
||||
pendingIncidents.putIfAbsent(key, new PendingIncident(key, credential, profileNames,
|
||||
entry.getValue().stream().sorted().toList(), remainingCoolOffSeconds, 0));
|
||||
startOrCoalesce(entry.getKey());
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Tell the owning lead that a classified backend error could not be tied to a credential.
|
||||
* Without an owning lead, emit a warning because no control can act on the target.
|
||||
*/
|
||||
public void onBackendTargetUnmapped(String target, String reason) {
|
||||
var lead = primaryRegistry.nudgeTargetFor(target);
|
||||
if (lead.isEmpty()) {
|
||||
log.warn("push: backend target {} could not map to a credential: {}", target, reason);
|
||||
return;
|
||||
}
|
||||
UnmappedTargetLead key = new UnmappedTargetLead(target, reason, lead.get());
|
||||
if (deliveredUnmappedTargets.contains(key)) return;
|
||||
pendingUnmappedTargets.putIfAbsent(key, new PendingUnmappedTarget(key, 0));
|
||||
startOrCoalesce(lead.get());
|
||||
}
|
||||
|
||||
// --- the schedule ----------------------------------------------------------------------------
|
||||
|
||||
/** Start a reminder schedule for {@code lead}, or join the one already running. */
|
||||
@@ -425,18 +540,25 @@ public final class ReplyPushLoop {
|
||||
Set<String> repliesBefore = pendingReplyTargetsFor(lead);
|
||||
Set<String> ticketsBefore = pendingTicketIdsFor(lead);
|
||||
Set<String> questionsBefore = pendingQuestionTurnIdsFor(lead);
|
||||
Set<IncidentLead> incidentsBefore = pendingIncidentKeysFor(lead);
|
||||
Set<UnmappedTargetLead> unmappedTargetsBefore = pendingUnmappedTargetKeysFor(lead);
|
||||
int replyReminderCount = minReplyNudgeCountFor(lead);
|
||||
int ticketReminderCount = minTicketNudgeCountFor(lead);
|
||||
int questionReminderCount = minQuestionNudgeCountFor(lead);
|
||||
var action = decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount);
|
||||
int incidentReminderCount = minIncidentNudgeCountFor(lead);
|
||||
int unmappedTargetReminderCount = minUnmappedTargetNudgeCountFor(lead);
|
||||
var action = decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount,
|
||||
incidentReminderCount, unmappedTargetReminderCount);
|
||||
switch (action) {
|
||||
case INJECT -> {
|
||||
injectNudge(lead, replyReminderCount, ticketReminderCount, questionReminderCount);
|
||||
injectNudge(lead, replyReminderCount, ticketReminderCount, questionReminderCount,
|
||||
incidentReminderCount, unmappedTargetReminderCount);
|
||||
scheduleNext(lead);
|
||||
}
|
||||
// Re-check after the configured backoff; the lead may become injectable soon.
|
||||
case WAIT_BUSY -> scheduleNext(lead);
|
||||
case STOP -> stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore);
|
||||
case STOP -> stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore, incidentsBefore,
|
||||
unmappedTargetsBefore);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -480,11 +602,25 @@ public final class ReplyPushLoop {
|
||||
* decision-to-release window reclaims the schedule slot exactly like a raced-in reply or ticket.
|
||||
*/
|
||||
void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore,
|
||||
Set<String> questionsBefore) {
|
||||
Set<String> questionsBefore) {
|
||||
stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore, pendingIncidentKeysFor(lead));
|
||||
}
|
||||
|
||||
private void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore,
|
||||
Set<String> questionsBefore, Set<IncidentLead> incidentsBefore) {
|
||||
stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore, incidentsBefore,
|
||||
pendingUnmappedTargetKeysFor(lead));
|
||||
}
|
||||
|
||||
private void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore,
|
||||
Set<String> questionsBefore, Set<IncidentLead> incidentsBefore,
|
||||
Set<UnmappedTargetLead> unmappedTargetsBefore) {
|
||||
activeLeads.remove(lead);
|
||||
boolean racedIn = pendingReplyTargetsFor(lead).stream().anyMatch(t -> !repliesBefore.contains(t))
|
||||
|| pendingTicketIdsFor(lead).stream().anyMatch(t -> !ticketsBefore.contains(t))
|
||||
|| pendingQuestionTurnIdsFor(lead).stream().anyMatch(t -> !questionsBefore.contains(t));
|
||||
|| pendingQuestionTurnIdsFor(lead).stream().anyMatch(t -> !questionsBefore.contains(t))
|
||||
|| pendingIncidentKeysFor(lead).stream().anyMatch(i -> !incidentsBefore.contains(i))
|
||||
|| pendingUnmappedTargetKeysFor(lead).stream().anyMatch(i -> !unmappedTargetsBefore.contains(i));
|
||||
if (racedIn && activeLeads.putIfAbsent(lead, Boolean.TRUE) == null) {
|
||||
log.debug("push: new work for lead {} raced the reminder loop's stop — restarting", lead);
|
||||
scheduleNext(lead);
|
||||
@@ -495,18 +631,22 @@ public final class ReplyPushLoop {
|
||||
|
||||
/** Send one combined nudge covering everything currently pending for {@code lead}. */
|
||||
private void injectNudge(String lead, int replyReminderCount, int ticketReminderCount,
|
||||
int questionReminderCount) {
|
||||
int questionReminderCount, int incidentReminderCount,
|
||||
int unmappedTargetReminderCount) {
|
||||
// Re-read rather than threading it down from decide(): a reply can drain, a ticket be
|
||||
// collected, or a question be answered (or another arrive), between the decision and the
|
||||
// injection.
|
||||
Set<String> replyTargets = pendingReplyTargetsFor(lead);
|
||||
List<PendingTicket> tickets = pendingTicketsFor(lead);
|
||||
List<PendingQuestion> questions = pendingQuestionsFor(lead);
|
||||
if (replyTargets.isEmpty() && tickets.isEmpty() && questions.isEmpty()) {
|
||||
List<PendingIncident> incidents = pendingIncidentsFor(lead);
|
||||
List<PendingUnmappedTarget> unmappedTargets = pendingUnmappedTargetsFor(lead);
|
||||
if (replyTargets.isEmpty() && tickets.isEmpty() && questions.isEmpty() && incidents.isEmpty()
|
||||
&& unmappedTargets.isEmpty()) {
|
||||
log.debug("push: pending work for lead {} drained before the nudge could be sent", lead);
|
||||
return;
|
||||
}
|
||||
String nudge = formatNudge(replyTargets, tickets, questions);
|
||||
String nudge = formatNudge(replyTargets, tickets, questions, incidents, unmappedTargets);
|
||||
try {
|
||||
agents.send(lead, nudge);
|
||||
log.debug("push: nudge sent to lead {} (reply {}/{}, ticket {}/{}, question {}/{}; "
|
||||
@@ -515,6 +655,16 @@ public final class ReplyPushLoop {
|
||||
questionReminderCount + 1, maxReminders,
|
||||
replyTargets.size(), tickets.size(), questions.size());
|
||||
countNudge("delivered");
|
||||
for (PendingIncident incident : incidents) {
|
||||
if (pendingIncidents.remove(incident.key(), incident)) {
|
||||
deliveredIncidents.add(incident.key());
|
||||
}
|
||||
}
|
||||
for (PendingUnmappedTarget unmappedTarget : unmappedTargets) {
|
||||
if (pendingUnmappedTargets.remove(unmappedTarget.key(), unmappedTarget)) {
|
||||
deliveredUnmappedTargets.add(unmappedTarget.key());
|
||||
}
|
||||
}
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("push: failed to nudge lead {} (reply {}/{}, ticket {}/{}, question {}/{}): {}",
|
||||
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
|
||||
@@ -526,12 +676,13 @@ public final class ReplyPushLoop {
|
||||
// item already at or over the cap keeps riding along in the text (still pending, still
|
||||
// named) but its extra bumps here are inert: decide() already treats it as ineligible once
|
||||
// its count reaches maxReminders.
|
||||
bumpNudgeCounts(replyTargets, tickets, questions);
|
||||
bumpNudgeCounts(replyTargets, tickets, questions, incidents, unmappedTargets);
|
||||
}
|
||||
|
||||
/** Record that every one of these items was just named in a sent (or attempted) nudge. */
|
||||
private void bumpNudgeCounts(Set<String> replyTargets, List<PendingTicket> tickets,
|
||||
List<PendingQuestion> questions) {
|
||||
List<PendingQuestion> questions, List<PendingIncident> incidents,
|
||||
List<PendingUnmappedTarget> unmappedTargets) {
|
||||
for (String target : replyTargets) {
|
||||
pendingReplies.computeIfPresent(target, (t, e) -> new ReplyEntry(e.lead(), e.nudgeCount() + 1));
|
||||
}
|
||||
@@ -544,6 +695,14 @@ public final class ReplyPushLoop {
|
||||
new PendingQuestion(e.turnId(), e.ticket(), e.target(), e.lead(), e.question(),
|
||||
e.nudgeCount() + 1));
|
||||
}
|
||||
for (PendingIncident incident : incidents) {
|
||||
pendingIncidents.computeIfPresent(incident.key(), (id, e) -> new PendingIncident(e.key(),
|
||||
e.credential(), e.profiles(), e.targets(), e.remainingCoolOffSeconds(), e.nudgeCount() + 1));
|
||||
}
|
||||
for (PendingUnmappedTarget unmappedTarget : unmappedTargets) {
|
||||
pendingUnmappedTargets.computeIfPresent(unmappedTarget.key(), (id, e) ->
|
||||
new PendingUnmappedTarget(e.key(), e.nudgeCount() + 1));
|
||||
}
|
||||
}
|
||||
|
||||
/** Schedule the next tick on the scheduler thread pool. */
|
||||
@@ -556,7 +715,8 @@ public final class ReplyPushLoop {
|
||||
|
||||
/** Render everything pending for one lead as a single nudge line. */
|
||||
private static String formatNudge(Set<String> replyTargets, List<PendingTicket> tickets,
|
||||
List<PendingQuestion> questions) {
|
||||
List<PendingQuestion> questions, List<PendingIncident> incidents,
|
||||
List<PendingUnmappedTarget> unmappedTargets) {
|
||||
List<String> parts = new ArrayList<>();
|
||||
if (!replyTargets.isEmpty()) {
|
||||
parts.add(formatRepliesNudge(replyTargets));
|
||||
@@ -567,6 +727,12 @@ public final class ReplyPushLoop {
|
||||
if (!questions.isEmpty()) {
|
||||
parts.add(formatQuestionsNudge(questions));
|
||||
}
|
||||
if (!incidents.isEmpty()) {
|
||||
parts.addAll(incidents.stream().map(ReplyPushLoop::formatBackendIncidentNudge).toList());
|
||||
}
|
||||
if (!unmappedTargets.isEmpty()) {
|
||||
parts.addAll(unmappedTargets.stream().map(ReplyPushLoop::formatUnmappedTargetNudge).toList());
|
||||
}
|
||||
return String.join(" | ", parts);
|
||||
}
|
||||
|
||||
@@ -606,6 +772,16 @@ public final class ReplyPushLoop {
|
||||
return QUESTIONS_NUDGE_FORMAT.formatted(pending.size(), ids);
|
||||
}
|
||||
|
||||
private static String formatBackendIncidentNudge(PendingIncident incident) {
|
||||
return BACKEND_INCIDENT_NUDGE_FORMAT.formatted(incident.credential(), incident.remainingCoolOffSeconds(),
|
||||
String.join(", ", incident.profiles()), String.join(", ", incident.targets()));
|
||||
}
|
||||
|
||||
private static String formatUnmappedTargetNudge(PendingUnmappedTarget unmappedTarget) {
|
||||
return BACKEND_TARGET_UNMAPPED_NUDGE_FORMAT.formatted(unmappedTarget.key().target(),
|
||||
unmappedTarget.key().reason());
|
||||
}
|
||||
|
||||
// --- lifecycle -----------------------------------------------------------------------------
|
||||
|
||||
/**
|
||||
@@ -627,6 +803,10 @@ public final class ReplyPushLoop {
|
||||
pendingReplies.clear();
|
||||
pendingTickets.clear();
|
||||
pendingQuestions.clear();
|
||||
pendingIncidents.clear();
|
||||
deliveredIncidents.clear();
|
||||
pendingUnmappedTargets.clear();
|
||||
deliveredUnmappedTargets.clear();
|
||||
}
|
||||
|
||||
/** @see #stop() */
|
||||
|
||||
@@ -9,12 +9,19 @@ import java.util.concurrent.CompletableFuture;
|
||||
public final class TurnToken {
|
||||
private final String target;
|
||||
private final CompletableFuture<Rendezvous.Resolution> waiter;
|
||||
private final String injectedText;
|
||||
|
||||
public TurnToken(String target, CompletableFuture<Rendezvous.Resolution> waiter) {
|
||||
this(target, waiter, null);
|
||||
}
|
||||
|
||||
public TurnToken(String target, CompletableFuture<Rendezvous.Resolution> waiter, String injectedText) {
|
||||
this.target = target;
|
||||
this.waiter = waiter;
|
||||
this.injectedText = injectedText;
|
||||
}
|
||||
|
||||
public String target() { return target; }
|
||||
public CompletableFuture<Rendezvous.Resolution> waiter() { return waiter; }
|
||||
public String injectedText() { return injectedText; }
|
||||
}
|
||||
|
||||
@@ -1,55 +1,68 @@
|
||||
package dev.ltms.fleet.placement;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
|
||||
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps and
|
||||
* reachability so that a pre-existing config behaves identically after upgrade.
|
||||
*
|
||||
* <p>Two exceptions walk past the default instead of returning it unconditionally:
|
||||
* <p>Three exceptions walk past the default instead of returning it unconditionally:
|
||||
* <ul>
|
||||
* <li>Quarantine (CB-578 stage B): a quarantined default is a credential that just refused on
|
||||
* a usage limit, not a transient capacity or reachability concern.
|
||||
* <li>Cooling off (fleetd #201 Unit 5): a credential cooling off after repeated backend errors
|
||||
* ({@code BackendOutagePolicy}) — a separate, shorter-lived source from quarantine. When a
|
||||
* profile is both quarantined and cooling off, only the quarantine reason is reported
|
||||
* (exhaustion takes priority), matching {@code CompositePeerLauncher}'s explicit-spawn order.
|
||||
* <li>Weight 0 (CB-554): {@code fixed} is still automatic selection, so a profile the operator
|
||||
* marked "never auto-select me" ({@code weight <= 0}) must be skipped here exactly as
|
||||
* {@code weighted}/{@code round-robin} skip it — an explicit {@code fleet_spawn} naming
|
||||
* the profile is unaffected, only this automatic fallback walk.
|
||||
* </ul>
|
||||
* A fleet where nothing is ever quarantined or weight-0 never exercises either path, so today's
|
||||
* behaviour is unchanged.
|
||||
* A fleet where nothing is ever quarantined, cooling off, or weight-0 never exercises any of these
|
||||
* paths, so today's behaviour is unchanged.
|
||||
*/
|
||||
final class FixedPlacementPolicy implements PlacementPolicy {
|
||||
|
||||
@Override
|
||||
public PlacementCandidate select(PlacementContext ctx) {
|
||||
String d = ctx.defaultProfile();
|
||||
if (d != null && !d.isBlank() && !ctx.quarantined().contains(d) && !weightExcluded(ctx, d)) {
|
||||
if (d != null && !d.isBlank() && !ctx.quarantined().contains(d) && !ctx.coolingOff().contains(d)
|
||||
&& !weightExcluded(ctx, d)) {
|
||||
return new PlacementCandidate(d, null, 1.0f, null);
|
||||
}
|
||||
for (PlacementCandidate c : ctx.candidates()) {
|
||||
if (!ctx.quarantined().contains(c.profile()) && !c.excluded()) {
|
||||
if (!ctx.quarantined().contains(c.profile()) && !ctx.coolingOff().contains(c.profile())
|
||||
&& !c.excluded()) {
|
||||
return new PlacementCandidate(c.profile(), null, c.weight(), c.maxLoad());
|
||||
}
|
||||
}
|
||||
if (d != null && !d.isBlank()) {
|
||||
boolean dQuarantined = ctx.quarantined().contains(d);
|
||||
// Exhaustion quarantine takes priority: reported only when quarantine is absent, so the
|
||||
// message never claims "cooling off" for a profile that is really backend-exhausted.
|
||||
boolean dCoolingOff = !dQuarantined && ctx.coolingOff().contains(d);
|
||||
boolean dWeightExcluded = weightExcluded(ctx, d);
|
||||
if (dQuarantined && dWeightExcluded) {
|
||||
throw new PlacementException("worker profile '" + d + "' is quarantined (backend "
|
||||
+ "exhausted) and has weight 0 (excluded from automatic selection), and no "
|
||||
+ "available candidate remains");
|
||||
}
|
||||
if (dWeightExcluded) {
|
||||
throw new PlacementException("worker profile '" + d + "' has weight 0 (excluded "
|
||||
+ "from automatic selection) and no available candidate remains");
|
||||
}
|
||||
if (dQuarantined) {
|
||||
throw new PlacementException("worker profile '" + d + "' is quarantined (backend "
|
||||
+ "exhausted) and no un-quarantined candidate is available");
|
||||
if (dQuarantined || dCoolingOff || dWeightExcluded) {
|
||||
List<String> reasons = new ArrayList<>();
|
||||
if (dQuarantined) {
|
||||
reasons.add("is quarantined (backend exhausted)");
|
||||
}
|
||||
if (dCoolingOff) {
|
||||
reasons.add("is cooling off after repeated backend errors");
|
||||
}
|
||||
if (dWeightExcluded) {
|
||||
reasons.add("has weight 0 (excluded from automatic selection)");
|
||||
}
|
||||
throw new PlacementException("worker profile '" + d + "' "
|
||||
+ String.join(" and ", reasons) + ", and no available candidate remains");
|
||||
}
|
||||
}
|
||||
if (!ctx.candidates().isEmpty()) {
|
||||
throw new PlacementException(
|
||||
"all worker profiles are excluded from automatic selection (quarantined or weight-0)");
|
||||
throw new PlacementException("all worker profiles are excluded from automatic "
|
||||
+ "selection (quarantined, cooling off, or weight-0)");
|
||||
}
|
||||
throw new PlacementException("no worker profiles configured");
|
||||
}
|
||||
|
||||
@@ -14,10 +14,19 @@ import java.util.function.Function;
|
||||
* @param quarantined profiles whose credential is currently quarantined (CB-578 stage B) — a
|
||||
* {@code BACKEND_EXHAUSTED} classification put it, or a profile it shares a
|
||||
* credential with, on cooldown. Filtered the same way as {@code unreachable}.
|
||||
* @param coolingOff profiles whose credential is currently cooling off after repeated backend
|
||||
* errors (fleetd #201 Unit 5 — {@code BackendOutagePolicy}), a SEPARATE,
|
||||
* shorter-lived source from {@code quarantined}: a credential outage cools off
|
||||
* even when no member was ever exhausted. Deliberately its own set rather than
|
||||
* merged into {@code quarantined} — {@link PlacementPolicyUtil} needs to tell
|
||||
* the two apart so its refusal message says "cooling off", not "exhausted",
|
||||
* when only this one is active. A profile can be in both sets at once; when it
|
||||
* is, exhaustion quarantine is reported (it takes priority).
|
||||
*/
|
||||
public record PlacementContext(String defaultProfile,
|
||||
List<PlacementCandidate> candidates,
|
||||
Function<String, Integer> liveCount,
|
||||
Set<String> unreachable,
|
||||
Set<String> quarantined) {
|
||||
Set<String> quarantined,
|
||||
Set<String> coolingOff) {
|
||||
}
|
||||
|
||||
@@ -14,14 +14,16 @@ final class PlacementPolicyUtil {
|
||||
/**
|
||||
* Candidates that are not weight-excluded (CB-554: explicit {@code weight <= 0}, checked
|
||||
* first because it is a static config choice rather than transient state), not
|
||||
* known-unreachable, not quarantined (CB-578 stage B), and have not reached their maxLoad.
|
||||
* A {@code null} maxLoad means unlimited.
|
||||
* known-unreachable, not quarantined (CB-578 stage B), not cooling off after repeated backend
|
||||
* errors (fleetd #201 Unit 5 — a separate, shorter-lived source from quarantine), and have not
|
||||
* reached their maxLoad. A {@code null} maxLoad means unlimited.
|
||||
*/
|
||||
static List<PlacementCandidate> available(PlacementContext ctx) {
|
||||
List<PlacementCandidate> out = new ArrayList<>();
|
||||
for (PlacementCandidate c : ctx.candidates()) {
|
||||
if (c.excluded() || ctx.unreachable().contains(c.profile())
|
||||
|| ctx.quarantined().contains(c.profile())) {
|
||||
|| ctx.quarantined().contains(c.profile())
|
||||
|| ctx.coolingOff().contains(c.profile())) {
|
||||
continue;
|
||||
}
|
||||
Integer cap = c.maxLoad();
|
||||
@@ -38,21 +40,27 @@ final class PlacementPolicyUtil {
|
||||
|
||||
/**
|
||||
* Build a clear exception describing why every candidate was dropped: all weight-0, all
|
||||
* quarantined, all at capacity, all unreachable, or a mix. Each candidate is counted into
|
||||
* exactly one bucket (weight-excluded takes priority) so a candidate excluded for more than
|
||||
* one reason is never double-counted.
|
||||
* quarantined, all cooling off, all at capacity, all unreachable, or a mix. Each candidate is
|
||||
* counted into exactly one bucket (weight-excluded first, then quarantined, then cooling off)
|
||||
* so a candidate excluded for more than one reason is never double-counted — a candidate that is
|
||||
* both quarantined (CB-578 stage B, backend exhausted) and cooling off (fleetd #201 Unit 5,
|
||||
* repeated backend errors) counts only as quarantined, matching {@code CompositePeerLauncher}'s
|
||||
* explicit-spawn ordering: exhaustion quarantine takes priority when both are active.
|
||||
*/
|
||||
static PlacementException emptyException(PlacementContext ctx) {
|
||||
int weightExcluded = 0;
|
||||
int atCap = 0;
|
||||
int unreachable = 0;
|
||||
int quarantined = 0;
|
||||
int coolingOff = 0;
|
||||
for (PlacementCandidate c : ctx.candidates()) {
|
||||
Integer cap = c.maxLoad();
|
||||
if (c.excluded()) {
|
||||
weightExcluded++;
|
||||
} else if (ctx.quarantined().contains(c.profile())) {
|
||||
quarantined++;
|
||||
} else if (ctx.coolingOff().contains(c.profile())) {
|
||||
coolingOff++;
|
||||
} else if (ctx.unreachable().contains(c.profile())) {
|
||||
unreachable++;
|
||||
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
|
||||
@@ -71,6 +79,10 @@ final class PlacementPolicyUtil {
|
||||
if (quarantined == total) {
|
||||
return new PlacementException("all worker profiles are quarantined (backend exhausted)");
|
||||
}
|
||||
if (coolingOff == total) {
|
||||
return new PlacementException(
|
||||
"all worker profiles are cooling off after repeated backend errors");
|
||||
}
|
||||
if (atCap == total) {
|
||||
return new PlacementException("all worker profiles are at maxLoad");
|
||||
}
|
||||
@@ -79,7 +91,9 @@ final class PlacementPolicyUtil {
|
||||
}
|
||||
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
|
||||
+ unreachable + " unreachable, " + quarantined + " quarantined, "
|
||||
+ coolingOff + " cooling off, "
|
||||
+ weightExcluded + " weight-0, "
|
||||
+ (total - atCap - unreachable - quarantined - weightExcluded) + " remaining");
|
||||
+ (total - atCap - unreachable - quarantined - coolingOff - weightExcluded)
|
||||
+ " remaining");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -13,6 +13,7 @@ import dev.ltms.fleet.herdr.Agent;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.HerdrException;
|
||||
import dev.ltms.fleet.inject.MemberPresence;
|
||||
import dev.ltms.fleet.member.MemberCredentialPolicyView;
|
||||
import dev.ltms.fleet.peer.PeerUnreachableException;
|
||||
import dev.ltms.fleet.placement.PlacementException;
|
||||
import dev.ltms.fleet.msg.MessageService;
|
||||
@@ -32,6 +33,7 @@ import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.Predicate;
|
||||
import java.util.function.Supplier;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
/**
|
||||
@@ -64,6 +66,10 @@ public final class FleetApp {
|
||||
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
|
||||
private final CallerResolver auth; // CB-501: null → authz not enforced (legacy behaviour)
|
||||
private final Metrics metrics; // CB-502: null → /metrics not exposed
|
||||
// fleetd #111: re-read per request, same hot-reload shape as every other live config read —
|
||||
// absent() (the honest "no policy configured" view) for every constructor that does not wire
|
||||
// a real one, so existing legacy call sites keep building without knowing this field exists.
|
||||
private final Supplier<MemberCredentialPolicyView> memberCredentials;
|
||||
private final ObjectMapper mapper = new ObjectMapper();
|
||||
|
||||
/**
|
||||
@@ -111,6 +117,19 @@ public final class FleetApp {
|
||||
MessageService messages, MemberPresence presence,
|
||||
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics,
|
||||
Predicate<String> deliverable) {
|
||||
this(herdr, memberHerdr, workers, sessions, messages, presence, mcpServlet, auth, metrics,
|
||||
deliverable, MemberCredentialPolicyView::absent);
|
||||
}
|
||||
|
||||
/**
|
||||
* @param memberCredentials live {@code memberCredentials:} policy view (fleetd #111), re-read
|
||||
* per request for {@code GET /member-credentials}; production wiring
|
||||
* passes the same hot-reload shape as every other live config read
|
||||
*/
|
||||
public FleetApp(HerdrClient herdr, HerdrClient memberHerdr, PeerLauncher workers, SessionManager sessions,
|
||||
MessageService messages, MemberPresence presence,
|
||||
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics,
|
||||
Predicate<String> deliverable, Supplier<MemberCredentialPolicyView> memberCredentials) {
|
||||
this.herdr = herdr;
|
||||
this.memberHerdr = memberHerdr != null ? memberHerdr : herdr;
|
||||
this.workers = workers;
|
||||
@@ -120,6 +139,7 @@ public final class FleetApp {
|
||||
this.mcpServlet = mcpServlet;
|
||||
this.auth = auth;
|
||||
this.metrics = metrics;
|
||||
this.memberCredentials = memberCredentials != null ? memberCredentials : MemberCredentialPolicyView::absent;
|
||||
}
|
||||
|
||||
/** Wire routes onto a fresh, unstarted Javalin instance. Caller starts it. */
|
||||
@@ -148,6 +168,7 @@ public final class FleetApp {
|
||||
app.get("/agents", this::agents);
|
||||
app.get("/members", this::listMembers); // CB-304: registry roster + live herdr status
|
||||
app.get("/profiles", this::profiles); // configured backend profiles
|
||||
app.get("/member-credentials", this::memberCredentials); // fleetd #111: policy names + counts, never a value
|
||||
app.post("/members", this::spawnMember); // optional ?role=&profile= or {"role":…,"profile":…}
|
||||
app.delete("/members/{paneId}", this::stopMember);
|
||||
app.post("/sessions/{id}/message", this::sendMessage); // fleet_send (primary; blocking, wait:false, or answer via turnId)
|
||||
@@ -346,6 +367,29 @@ public final class FleetApp {
|
||||
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile()));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #111 (CB-608): the live {@code memberCredentials:} policy as names and counts —
|
||||
* NEVER a value. The daemon does not hold a credential's value in the first place (only the
|
||||
* name it is configured under), so there is nothing to redact here beyond what {@link
|
||||
* MemberCredentialPolicyView} already omits by construction. This is the source
|
||||
* {@code scripts/probe-member-credentials.sh} reads instead of carrying its own hardcoded
|
||||
* name list, which is exactly what let the list drift silently behind the real policy.
|
||||
*/
|
||||
private void memberCredentials(Context ctx) {
|
||||
if (!allow(ctx, Authz.Action.READ, null)) {
|
||||
return;
|
||||
}
|
||||
MemberCredentialPolicyView view = memberCredentials.get();
|
||||
ctx.status(200).json(Map.of(
|
||||
"present", view.present(),
|
||||
"policy", view.policy() == null ? "" : view.policy(),
|
||||
"known", view.known(),
|
||||
"allowed", view.allowed(),
|
||||
"knownCount", view.knownCount(),
|
||||
"allowedCount", view.allowedCount(),
|
||||
"blockedCount", view.blockedCount()));
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn a guard-checked worker. An optional {@code profile} (query param or {@code {"profile":…}}
|
||||
* body) picks which configured profile; omitted → the default. 403 if the base_url would breach
|
||||
|
||||
@@ -426,19 +426,75 @@ public final class GitWorktrees implements Worktrees {
|
||||
* neutralized copy from ever showing up as a local modification the worker might commit. A config
|
||||
* the repo does not carry is skipped silently — no stub is invented for a file the repo does not
|
||||
* have, and one missing file must never fail provisioning.
|
||||
*
|
||||
* <p>fleetd #134. The neutralization above is correct and stays unconditional — the defect was
|
||||
* that it was invisible on both sides. Neither the daemon's own log nor the worker sitting in the
|
||||
* worktree could tell a stub from the repo's real file: a real worker read a 3-byte {@code {}}
|
||||
* where the repo's {@code opencode.json} is 30+ lines, and reported — truthfully from what it
|
||||
* could see, and wrongly — that a mount key did not exist. Two fixes, aimed at two different
|
||||
* readers:
|
||||
* <ul>
|
||||
* <li>the daemon operator reads {@link #log}, so the summary below names the denominator, what
|
||||
* was neutralized, and why anything was not — the same shape {@code overlayParity} reports
|
||||
* its own copy in;
|
||||
* <li>the worker reads its own worktree, not the daemon's log, so the same fact is recorded a
|
||||
* second time in worktree-scoped git config ({@code fleet.neutralizedConfig} /
|
||||
* {@code fleet.neutralizedConfigNote}, readable with {@code git config --worktree --get-all
|
||||
* fleet.neutralizedConfig}) rather than as a file in the working tree. A working-tree file
|
||||
* would show up in {@code git status} for the worker to trip on or commit; worktree-scoped
|
||||
* config lives in {@code .git/worktrees/<nonce>/config.worktree} and can never appear there.
|
||||
* This reuses the exact mechanism {@link #configureEnvironmentCredentialHelper} and
|
||||
* {@link #configureHttpsUrlRewriteForSshOrigin} already use for other worktree-local state.
|
||||
* </ul>
|
||||
*/
|
||||
private void isolateToolSurface(String worktreePath) {
|
||||
Path root = Path.of(worktreePath).toAbsolutePath().normalize();
|
||||
List<String> neutralized = new ArrayList<>();
|
||||
List<String> skipped = new ArrayList<>();
|
||||
for (WorktreeHostileConfig cfg : WORKTREE_HOSTILE_CONFIGS) {
|
||||
neutralize(root, worktreePath, cfg);
|
||||
if (neutralize(root, worktreePath, cfg)) {
|
||||
neutralized.add(cfg.file());
|
||||
} else {
|
||||
skipped.add(cfg.file() + " absent");
|
||||
}
|
||||
}
|
||||
String detail = neutralized.isEmpty() ? String.join(", ", skipped)
|
||||
: skipped.isEmpty() ? String.join(", ", neutralized)
|
||||
: String.join(", ", neutralized) + " (" + String.join(", ", skipped) + ")";
|
||||
log.info("tool-surface isolation: neutralized {} of {} configs: {} — the worktree copy is a "
|
||||
+ "stub, not the repo's file; edit the real file in the primary checkout instead",
|
||||
neutralized.size(), WORKTREE_HOSTILE_CONFIGS.size(), detail);
|
||||
recordNeutralizedConfigForWorker(worktreePath, neutralized);
|
||||
}
|
||||
|
||||
private void neutralize(Path root, String worktreePath, WorktreeHostileConfig cfg) {
|
||||
/**
|
||||
* The worker-readable half of fleetd #134: record which files were neutralized where the worker
|
||||
* itself can read it, without a working-tree file that would show up in {@code git status}.
|
||||
* Worktree-scoped git config is per-worktree, lives under {@code .git/worktrees/<nonce>/} rather
|
||||
* than the working tree, and this repo already relies on the same mechanism (and the same
|
||||
* {@code extensions.worktreeConfig} enablement) for the credential helper and the SSH→HTTPS
|
||||
* rewrite — see {@link #configureEnvironmentCredentialHelper}.
|
||||
*/
|
||||
private void recordNeutralizedConfigForWorker(String worktreePath, List<String> neutralized) {
|
||||
if (neutralized.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
exec("git", "-C", worktreePath, "config", "extensions.worktreeConfig", "true");
|
||||
for (String file : neutralized) {
|
||||
exec("git", "-C", worktreePath, "config", "--worktree", "--add", "fleet.neutralizedConfig", file);
|
||||
}
|
||||
exec("git", "-C", worktreePath, "config", "--worktree", "fleet.neutralizedConfigNote",
|
||||
"the worktree copy of each fleet.neutralizedConfig path is a stub, not the repo's "
|
||||
+ "committed file; edit the real file from the primary checkout instead");
|
||||
}
|
||||
|
||||
/** @return true if {@code cfg} was neutralized (present, or created because {@link
|
||||
* WorktreeHostileConfig#createIfAbsent()}); false if the repo does not carry it and it was
|
||||
* left alone. */
|
||||
private boolean neutralize(Path root, String worktreePath, WorktreeHostileConfig cfg) {
|
||||
Path target = root.resolve(cfg.file());
|
||||
if (!Files.exists(target) && !cfg.createIfAbsent()) {
|
||||
log.debug("{} absent in the worktree — skipping (repo does not carry it)", cfg.file());
|
||||
return;
|
||||
return false;
|
||||
}
|
||||
try {
|
||||
Files.writeString(target, cfg.stub());
|
||||
@@ -449,7 +505,7 @@ public final class GitWorktrees implements Worktrees {
|
||||
if (isTracked(root, cfg.file())) {
|
||||
exec("git", "-C", worktreePath, "update-index", "--skip-worktree", cfg.file());
|
||||
}
|
||||
log.debug("neutralized {} — worker tool surface is launcher-mounted only", cfg.file());
|
||||
return true;
|
||||
}
|
||||
|
||||
@Override
|
||||
@@ -486,25 +542,46 @@ public final class GitWorktrees implements Worktrees {
|
||||
}
|
||||
Path srcRoot = Path.of(repoRoot).toAbsolutePath().normalize();
|
||||
Path dstRoot = Path.of(worktreePath).toAbsolutePath().normalize();
|
||||
List<String> copied = new ArrayList<>();
|
||||
List<String> skipped = new ArrayList<>();
|
||||
List<String> neutralized = new ArrayList<>();
|
||||
for (String rel : overlay) {
|
||||
Path src = srcRoot.resolve(rel).normalize();
|
||||
if (!Files.exists(src)) {
|
||||
log.debug("parity overlay source missing — skipping {}", rel);
|
||||
skipped.add(rel + " absent");
|
||||
continue;
|
||||
}
|
||||
Path dst = dstRoot.resolve(rel).normalize();
|
||||
try {
|
||||
Files.createDirectories(dst.getParent());
|
||||
Files.copy(src, dst, StandardCopyOption.REPLACE_EXISTING, StandardCopyOption.COPY_ATTRIBUTES);
|
||||
log.debug("copied parity overlay {}", rel);
|
||||
copied.add(rel);
|
||||
} catch (IOException e) {
|
||||
throw new WorktreeException("cannot copy overlay " + rel + ": " + e.getMessage(), e);
|
||||
}
|
||||
if (isTracked(dstRoot, rel)) {
|
||||
exec("git", "-C", worktreePath, "update-index", "--skip-worktree", rel);
|
||||
log.debug("marked overlay --skip-worktree {}", rel);
|
||||
neutralized.add(rel);
|
||||
}
|
||||
}
|
||||
// CB-148 point 3: a bare count ("copied 1") hides which candidates were even considered — the
|
||||
// same shape of under-reporting this repo has been bitten by before. Name the denominator
|
||||
// (every configured candidate), what was actually copied, and — for anything not copied —
|
||||
// why, so a spawn's overlay outcome is legible from the log alone, no filesystem dig required.
|
||||
String detail = copied.isEmpty() ? String.join(", ", skipped)
|
||||
: skipped.isEmpty() ? String.join(", ", copied)
|
||||
: String.join(", ", copied) + " (" + String.join(", ", skipped) + ")";
|
||||
log.info("parity overlay: copied {} of {} candidates: {}", copied.size(), overlay.size(), detail);
|
||||
// CB-134: a neutralized tracked file is otherwise a silent trap — a worker edits it, git
|
||||
// ignores the change with no error, and nothing anywhere said the file could not be
|
||||
// committed from this worktree. Name every file marked --skip-worktree here, with the
|
||||
// consequence stated in the message itself, rather than adding a worktree-local marker
|
||||
// file: acceptance criterion 1 requires the worktree hold exactly the configured overlay
|
||||
// set and nothing else, so an extra marker would itself violate the fix.
|
||||
if (!neutralized.isEmpty()) {
|
||||
log.info("parity overlay marked --skip-worktree (cannot be committed from this worktree): {}",
|
||||
String.join(", ", neutralized));
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
|
||||
@@ -46,7 +46,8 @@ public record MemberSession(
|
||||
String worktree,
|
||||
String branch,
|
||||
CharterReceipt charterReceipt,
|
||||
String agentSessionId) {
|
||||
String agentSessionId,
|
||||
String failureReason) {
|
||||
|
||||
/** One-shot worker lifecycle states. */
|
||||
public enum State {
|
||||
@@ -54,6 +55,7 @@ public record MemberSession(
|
||||
READY,
|
||||
BUSY,
|
||||
DONE,
|
||||
BACKEND_ERROR,
|
||||
FAILED,
|
||||
RELEASED
|
||||
}
|
||||
@@ -69,25 +71,35 @@ public record MemberSession(
|
||||
long lastActivityAtNanos, int turnCount, State state,
|
||||
String worktree, String branch) {
|
||||
this(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
||||
lastActivityAtNanos, turnCount, state, worktree, branch, null, null);
|
||||
lastActivityAtNanos, turnCount, state, worktree, branch, null, null, null);
|
||||
}
|
||||
|
||||
/** Backward-compatible shape without a backend failure reason. */
|
||||
public MemberSession(String paneId, String terminalId, String profile, MemberRole role,
|
||||
String cwd, String ownerTerminal, long spawnedAtNanos,
|
||||
long lastActivityAtNanos, int turnCount, State state,
|
||||
String worktree, String branch, CharterReceipt charterReceipt,
|
||||
String agentSessionId) {
|
||||
this(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
||||
lastActivityAtNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId, null);
|
||||
}
|
||||
|
||||
/** Return a copy of this session in {@code state}. */
|
||||
public MemberSession withState(State state) {
|
||||
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
||||
lastActivityAtNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId);
|
||||
lastActivityAtNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId, failureReason);
|
||||
}
|
||||
|
||||
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
|
||||
public MemberSession withActivity(long nowNanos) {
|
||||
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
||||
nowNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId);
|
||||
nowNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId, failureReason);
|
||||
}
|
||||
|
||||
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
|
||||
public MemberSession bumpTurn(long nowNanos) {
|
||||
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
||||
nowNanos, turnCount + 1, state, worktree, branch, charterReceipt, agentSessionId);
|
||||
nowNanos, turnCount + 1, state, worktree, branch, charterReceipt, agentSessionId, failureReason);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -98,6 +110,12 @@ public record MemberSession(
|
||||
*/
|
||||
public MemberSession withAgentSessionId(String agentSessionId) {
|
||||
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
||||
lastActivityAtNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId);
|
||||
lastActivityAtNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId, failureReason);
|
||||
}
|
||||
|
||||
/** Return a copy with the durable backend failure detail. */
|
||||
public MemberSession withFailureReason(String failureReason) {
|
||||
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
||||
lastActivityAtNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId, failureReason);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -658,6 +658,9 @@ public final class SessionManager implements TurnListener {
|
||||
m.put("profile", session.profile());
|
||||
m.put("role", session.role() == null ? "dev" : session.role().wireName());
|
||||
m.put("state", session.state().name().toLowerCase());
|
||||
if (session.failureReason() != null) {
|
||||
m.put("failureReason", session.failureReason());
|
||||
}
|
||||
if (session.worktree() != null) {
|
||||
m.put("worktree", session.worktree());
|
||||
}
|
||||
@@ -718,6 +721,36 @@ public final class SessionManager implements TurnListener {
|
||||
completeTurn(target, false);
|
||||
}
|
||||
|
||||
/**
|
||||
* Record that a member backend failed after a delegated ticket expired. This CAS loop accepts
|
||||
* either side of the completion race: {@code BUSY -> BACKEND_ERROR} or
|
||||
* {@code DONE -> BACKEND_ERROR}. A released or unknown session is never recreated.
|
||||
*
|
||||
* @return {@code true} when the error is recorded on a live session
|
||||
*/
|
||||
public boolean onBackendError(String target, String reason) {
|
||||
while (true) {
|
||||
MemberSession current = findByTerminal(target);
|
||||
if (current == null) {
|
||||
log.warn("backend error for unknown member terminal={}", target);
|
||||
return false;
|
||||
}
|
||||
if (current.state() == MemberSession.State.RELEASED) {
|
||||
return false;
|
||||
}
|
||||
if (current.state() == MemberSession.State.BACKEND_ERROR) {
|
||||
return true;
|
||||
}
|
||||
MemberSession updated = current.withState(MemberSession.State.BACKEND_ERROR)
|
||||
.withFailureReason(reason).withActivity(nowNanos.getAsLong());
|
||||
if (replace(current, updated)) {
|
||||
log.warn("member terminal={} pane={} transitioned {} -> BACKEND_ERROR: {}",
|
||||
target, current.paneId(), current.state(), reason);
|
||||
return true;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean hasPostTurnAction(String target) {
|
||||
if (!clearAfterTurn) return false;
|
||||
@@ -736,10 +769,9 @@ public final class SessionManager implements TurnListener {
|
||||
if (current == null || current.state() != MemberSession.State.BUSY) return false;
|
||||
long now = nowNanos.getAsLong();
|
||||
MemberSession updated = current.withState(MemberSession.State.DONE).withActivity(now);
|
||||
if (replace(current, updated)) {
|
||||
log.debug("session transitioned terminal={} pane={} BUSY -> DONE turn={}",
|
||||
target, current.paneId(), updated.turnCount());
|
||||
}
|
||||
if (!replace(current, updated)) return false;
|
||||
log.debug("session transitioned terminal={} pane={} BUSY -> DONE turn={}",
|
||||
target, current.paneId(), updated.turnCount());
|
||||
if (contextCap > 0 && updated.turnCount() >= contextCap) {
|
||||
release(current.paneId());
|
||||
return false;
|
||||
|
||||
@@ -0,0 +1,59 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.inject.BackendErrorPatternLookup;
|
||||
import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
|
||||
/**
|
||||
* fleetd #248 / fleetd#201 Unit 5: {@link Fleetd#backendErrorPatternLookup} is the factory that
|
||||
* replaced the local lambda {@code Fleetd.main} used to build {@code backendErrorPatterns} — one
|
||||
* of the two arguments {@code CompletionResolver} lost cleanly (0 compile errors, every test still
|
||||
* green) when this ticket's measurement dropped it alongside {@code backendErrorSink}. This class
|
||||
* proves the factory's own behaviour; {@code FleetdCompletionResolverWiringTest} proves {@code
|
||||
* main} still passes its result into {@code CompletionResolver}.
|
||||
*/
|
||||
class FleetdBackendErrorPatternLookupTest {
|
||||
|
||||
private static MemberSession session(String terminal, String profile) {
|
||||
return new MemberSession("pane-" + terminal, terminal, profile, MemberRole.DEV,
|
||||
"/cwd", null, 0L, 0L, 0, MemberSession.State.READY, null, null);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a target on a profile with a configured pattern resolves to that pattern")
|
||||
void configuredProfileResolves() {
|
||||
Map<String, Pattern> byProfile = Map.of("terra", Pattern.compile("(?i)503"));
|
||||
BackendErrorPatternLookup lookup =
|
||||
Fleetd.backendErrorPatternLookup(() -> List.of(session("term1", "terra")), byProfile);
|
||||
|
||||
assertEquals("(?i)503", lookup.patternFor("term1").pattern());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a target on a profile with no configured pattern resolves to null")
|
||||
void unconfiguredProfileResolvesToNull() {
|
||||
Map<String, Pattern> byProfile = Map.of("terra", Pattern.compile("x"));
|
||||
BackendErrorPatternLookup lookup =
|
||||
Fleetd.backendErrorPatternLookup(() -> List.of(session("term1", "sol")), byProfile);
|
||||
|
||||
assertNull(lookup.patternFor("term1"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unknown target resolves to null")
|
||||
void unknownTargetResolvesToNull() {
|
||||
BackendErrorPatternLookup lookup =
|
||||
Fleetd.backendErrorPatternLookup(List::of, Map.of("terra", Pattern.compile("x")));
|
||||
|
||||
assertNull(lookup.patternFor("term_stranger"));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,229 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.inject.BackendErrorSink;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||
import dev.ltms.fleet.mcp.PrimaryRegistry;
|
||||
import dev.ltms.fleet.msg.InMemoryReplyInbox;
|
||||
import dev.ltms.fleet.msg.ReplyPushLoop;
|
||||
import dev.ltms.fleet.peer.Capability;
|
||||
import dev.ltms.fleet.peer.PeerHandle;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.CopyOnWriteArrayList;
|
||||
import java.util.concurrent.CountDownLatch;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #248 / fleetd#201 Unit 5: {@link Fleetd#backendErrorSink} is the factory that replaced
|
||||
* the local lambda {@code Fleetd.main} used to build {@code backendErrorSink} — the other half of
|
||||
* the pair this ticket's measurement dropped cleanly (0 compile errors, every test still green).
|
||||
*
|
||||
* <p>Before this ticket, the closest thing to coverage was {@code
|
||||
* dev.ltms.fleet.inject.BackendOutageFlowTest}, whose own class doc said it "mirrors {@code
|
||||
* Fleetd.main}'s {@code backendErrorSink} lambda line-for-line" — a hand-copy that proves itself,
|
||||
* never that {@code main} still wires the real thing. This class exercises the actual production
|
||||
* factory instead. {@code FleetdCompletionResolverWiringTest} proves {@code main} still passes its
|
||||
* result into {@code CompletionResolver}.
|
||||
*/
|
||||
class FleetdBackendErrorSinkTest {
|
||||
|
||||
private final List<ScheduledExecutorService> schedulers = new ArrayList<>();
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
schedulers.forEach(ScheduledExecutorService::shutdownNow);
|
||||
}
|
||||
|
||||
private static FleetConfig.Profile stubWorker(String profile, String credentialId) {
|
||||
return new FleetConfig.Profile(profile, "http://gx00.gw:8000", "coder",
|
||||
null, "FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"w #{n}", null, null, null, null, null, null, null, null, null,
|
||||
null, null, credentialId, null);
|
||||
}
|
||||
|
||||
private static Map<String, FleetConfig.Profile> orderedProfiles() {
|
||||
Map<String, FleetConfig.Profile> m = new LinkedHashMap<>();
|
||||
m.put("terra", stubWorker("terra", "shared-openai"));
|
||||
m.put("sol", stubWorker("sol", "shared-openai"));
|
||||
return m;
|
||||
}
|
||||
|
||||
/** Minimal recording {@code HerdrClient} for the LEAD pane — mirrors ReplyPushLoopTest's own. */
|
||||
private static final class RecordingLeadClient implements HerdrClient {
|
||||
private static final ObjectMapper MAPPER = new ObjectMapper();
|
||||
private final List<Object> prompts = new CopyOnWriteArrayList<>();
|
||||
volatile CountDownLatch sendLatch = new CountDownLatch(1);
|
||||
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) {
|
||||
if ("agent.get".equals(method)) {
|
||||
return MAPPER.createObjectNode().set("agent", MAPPER.createObjectNode()
|
||||
.put("terminal_id", "term_primary").put("agent_status", "idle"));
|
||||
}
|
||||
if ("agent.prompt".equals(method)) {
|
||||
prompts.add(params);
|
||||
sendLatch.countDown();
|
||||
}
|
||||
return MAPPER.createObjectNode();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
|
||||
int sendCount() {
|
||||
return prompts.size();
|
||||
}
|
||||
}
|
||||
|
||||
/** A {@link PeerLauncher} that never actually spawns — enough to construct a bare {@link SessionManager}. */
|
||||
private static final class NeverSpawnsLauncher implements PeerLauncher {
|
||||
@Override
|
||||
public Set<Capability> capabilities() {
|
||||
return Set.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<Capability> capabilitiesFor(String profileName) {
|
||||
return Set.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req) {
|
||||
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return Set.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public String defaultProfile() {
|
||||
return null;
|
||||
}
|
||||
|
||||
@Override
|
||||
public String effectiveCwd(SpawnRequest req) {
|
||||
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<String> parityOverlay(String profileName) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<?> list() {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public int reapOrphanWorkers() {
|
||||
return 0;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void stop(String id) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean clearContext(String id) {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a target with no resolvable profile logs and returns without recording an incident (never throws)")
|
||||
void unresolvableProfileDoesNotRecordOrThrow() {
|
||||
SessionManager sessions = new SessionManager(new NeverSpawnsLauncher());
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(() -> 0L);
|
||||
RecordingLeadClient leadClient = new RecordingLeadClient();
|
||||
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||
PrimaryRegistry registry = new PrimaryRegistry(null);
|
||||
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||
schedulers.add(scheduler);
|
||||
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, new AgentControl(leadClient), inbox, scheduler, 3, 50);
|
||||
|
||||
BackendErrorSink sink = Fleetd.backendErrorSink(sessions, Map::of, outagePolicy, () -> pushLoop);
|
||||
sink.onBackendError("term_unmapped", "matched line", "503 Service Unavailable");
|
||||
|
||||
assertTrue(outagePolicy.remainingCoolOffSeconds("shared-openai").isEmpty(),
|
||||
"no credential is ever resolvable here, so nothing must be recorded");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("two distinct targets classified through the real factory start an incident and cool the credential")
|
||||
void twoDistinctTargetsStartAnIncident() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr()
|
||||
.readText("⏺ 503 Service Unavailable: upstream credential rejected\n❯ ");
|
||||
Map<String, FleetConfig.Profile> profiles = orderedProfiles();
|
||||
AtomicLong clockNanos = new AtomicLong(0L);
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(clockNanos::get);
|
||||
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr),
|
||||
new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
|
||||
profiles, "terra", _ -> "tok");
|
||||
CompositePeerLauncher workers = new CompositePeerLauncher(List.of(adapter), "terra", profiles,
|
||||
PlacementPolicies.weighted(), _ -> 0, null, BackendQuarantine.none(), outagePolicy);
|
||||
SessionManager sessions = new SessionManager(workers);
|
||||
MemberSession s1 = sessions.acquire("terra", null, null, null);
|
||||
MemberSession s2 = sessions.acquire("terra", null, null, null);
|
||||
|
||||
PrimaryRegistry registry = new PrimaryRegistry(null);
|
||||
registry.recordDelegation(s1.terminalId(), "term_primary");
|
||||
registry.recordDelegation(s2.terminalId(), "term_primary");
|
||||
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||
inbox.own(s1.terminalId());
|
||||
inbox.own(s2.terminalId());
|
||||
RecordingLeadClient leadClient = new RecordingLeadClient();
|
||||
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||
schedulers.add(scheduler);
|
||||
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, new AgentControl(leadClient), inbox, scheduler, 3, 50);
|
||||
AtomicReference<ReplyPushLoop> pushLoopRef = new AtomicReference<>(pushLoop);
|
||||
|
||||
// The exact object under test: Fleetd's real production factory, not a hand copy.
|
||||
BackendErrorSink sink = Fleetd.backendErrorSink(sessions, () -> profiles, outagePolicy, pushLoopRef::get);
|
||||
|
||||
sink.onBackendError(s1.terminalId(), "matched line", "503 Service Unavailable");
|
||||
assertTrue(outagePolicy.remainingCoolOffSeconds("shared-openai").isEmpty(),
|
||||
"one distinct target must not start a cool-off");
|
||||
|
||||
sink.onBackendError(s2.terminalId(), "matched line", "503 Service Unavailable");
|
||||
|
||||
assertTrue(leadClient.sendLatch.await(3, TimeUnit.SECONDS),
|
||||
"the second distinct target must cross the threshold and nudge the lead");
|
||||
var remaining = outagePolicy.remainingCoolOffSeconds("shared-openai");
|
||||
assertTrue(remaining.isPresent(), "two distinct targets must start a cool-off");
|
||||
assertEquals(1, leadClient.sendCount());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,89 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #248: this is the test that was actually missing. {@code Fleetd.main} builds its {@code
|
||||
* CompletionResolver} from an 8-argument constructor, and the ticket's own measurement proved two
|
||||
* ways to silently unwire it — both compiled with 0 errors and left every existing test green:
|
||||
*
|
||||
* <ul>
|
||||
* <li>replacing the worktree/branch argument (the 8th) with {@code _ -> null} — drops
|
||||
* fleetd#241's fallback-report location entirely;</li>
|
||||
* <li>replacing {@code backendErrorPatterns, backendErrorSink} (5th/6th) with {@code
|
||||
* BackendErrorPatternLookup.legacy(), BackendErrorSink.none()} — drops fleetd#201 Unit 5's
|
||||
* backend-error classification and cool-off entirely.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>Neither mutation could be caught by any test that constructs its own {@code
|
||||
* CompletionResolver} (every test before this one did exactly that) or by a test of {@link
|
||||
* Fleetd#worktreeBranchLookup}, {@link Fleetd#backendErrorPatternLookup}, or {@link
|
||||
* Fleetd#backendErrorSink} in isolation (see {@code FleetdWorktreeBranchLookupTest}, {@code
|
||||
* FleetdBackendErrorPatternLookupTest}, {@code FleetdBackendErrorSinkTest}) — those prove the
|
||||
* factories work, never that {@code main} still calls them. This class is a plain source-text
|
||||
* assertion on {@code Fleetd.java} — crude, but honest about what it checks, and it turns red the
|
||||
* instant the wiring is dropped, mirroring the same fallback shape {@link
|
||||
* FleetdFleetAppConstructionTest} already uses for a different constructor argument.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
||||
* CompletionResolver} and never runs {@code main}.
|
||||
*/
|
||||
class FleetdCompletionResolverWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] CompletionResolver's construction call still names backendErrorPatterns and backendErrorSink")
|
||||
void backendErrorArgumentsAreStillNamedAtTheCallSite() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"exhaustionSink, backendErrorPatterns, backendErrorSink, System::nanoTime,"),
|
||||
"CompletionResolver's construction call must still pass backendErrorPatterns and "
|
||||
+ "backendErrorSink as its 5th/6th arguments. Replacing them with "
|
||||
+ "BackendErrorPatternLookup.legacy()/BackendErrorSink.none() (fleetd #248's measured "
|
||||
+ "mutation) compiles with 0 errors and leaves every behavioural test green — this "
|
||||
+ "source check is what must go red instead.");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] CompletionResolver's construction call still passes worktreeBranchLookup(sessions::roster)")
|
||||
void worktreeBranchLookupIsStillPassedAtTheCallSite() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains("worktreeBranchLookup(sessions::roster)"),
|
||||
"CompletionResolver's construction call must still pass worktreeBranchLookup(sessions::roster) "
|
||||
+ "as its 8th (last) argument. Replacing it with the inert `_ -> null` (fleetd #248's "
|
||||
+ "other measured mutation) compiles with 0 errors and leaves every behavioural test "
|
||||
+ "green — this source check is what must go red instead.");
|
||||
assertFalse(source.contains("System::nanoTime,\n _ -> null"),
|
||||
"the worktree/branch argument must never regress to the inert `_ -> null` literal");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] backendErrorPatterns is assigned from the extracted backendErrorPatternLookup(...) factory")
|
||||
void backendErrorPatternsComesFromTheFactory() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"BackendErrorPatternLookup backendErrorPatterns = backendErrorPatternLookup(sessions::roster,"),
|
||||
"backendErrorPatterns must be assigned from Fleetd.backendErrorPatternLookup(...), not an "
|
||||
+ "inline lambda that a source check on the CompletionResolver call alone cannot see "
|
||||
+ "through");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] backendErrorSink is assigned from the extracted backendErrorSink(...) factory")
|
||||
void backendErrorSinkComesFromTheFactory() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"BackendErrorSink backendErrorSink = backendErrorSink(sessions, () -> config.get().profiles(),"),
|
||||
"backendErrorSink must be assigned from Fleetd.backendErrorSink(...), not an inline lambda "
|
||||
+ "that a source check on the CompletionResolver call alone cannot see through");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,156 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Map;
|
||||
import java.util.function.Function;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* fleetd #176: {@link Fleetd#leadSeatLookup} is the factory {@code Fleetd.main} wires into {@code
|
||||
* FleetMcp.LeadSeatSource} so {@code fleet_list}'s {@code free} can subtract the seat(s) a
|
||||
* {@code subscription: true} profile's own live LEAD session holds on that same account —
|
||||
* {@code maxLoad} never counted the lead, only members. {@code FleetdLeadSeatWiringTest} proves
|
||||
* {@code main} still passes this factory's result in; this class proves the factory's own matching
|
||||
* logic: subscription-only, credential-matched, and counting only CURRENTLY LIVE leads.
|
||||
*/
|
||||
class FleetdLeadSeatLookupTest {
|
||||
|
||||
private static FleetConfig.Profile subscriptionProfile(String name, String credentialId) {
|
||||
return new FleetConfig.Profile(name, null, "claude-sonnet-5", null, null, null,
|
||||
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
|
||||
null, 3, true, null, credentialId, null);
|
||||
}
|
||||
|
||||
private static FleetConfig.Profile offSubscriptionProfile(String name, String credentialId) {
|
||||
return new FleetConfig.Profile(name, "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
|
||||
null, "tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
|
||||
null, 2, false, null, credentialId, null);
|
||||
}
|
||||
|
||||
private static FleetConfig.Leader leadOnProfile(String profile) {
|
||||
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a live lead sharing the target profile's credential counts as one seat")
|
||||
void liveLeadSharingCredentialCountsAsOneSeat() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
|
||||
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
|
||||
() -> Map.of("term_primary", "primary"));
|
||||
|
||||
assertEquals(1, lookup.apply("sonnet"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("no live lead names this profile ⇒ zero seats, exactly as before this ticket")
|
||||
void noLiveLeadOnTheProfileCountsAsZero() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
|
||||
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders, Map::of);
|
||||
|
||||
assertEquals(0, lookup.apply("sonnet"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a lead entry with no `profile:` (recognise-only) contributes no seats — cannot be derived")
|
||||
void recogniseOnlyLeadWithNoProfileContributesNothing() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
|
||||
FleetConfig.Leader recogniseOnly = new FleetConfig.Leader(null, "lead: primary", 1, "lead:", 10,
|
||||
"claude", "claude-sonnet-5");
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", recogniseOnly);
|
||||
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
|
||||
() -> Map.of("term_primary", "primary"));
|
||||
|
||||
assertEquals(0, lookup.apply("sonnet"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a non-subscription profile never has a lead seat subtracted, whatever the credential match")
|
||||
void nonSubscriptionProfileIsNeverAdjusted() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of(
|
||||
"terra", offSubscriptionProfile("terra", "shared-openai"),
|
||||
"sonnet", subscriptionProfile("sonnet", "shared-openai"));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
|
||||
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
|
||||
() -> Map.of("term_primary", "primary"));
|
||||
|
||||
assertEquals(0, lookup.apply("terra"), "terra is not subscription:true, so it must never be adjusted");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("explicit, different credentialIds still separate two subscription profiles (post fleetd #176 "
|
||||
+ "stage 2 sentinel)")
|
||||
void differentCredentialIsNotCounted() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of(
|
||||
"sonnet", subscriptionProfile("sonnet", "claude-account-a"),
|
||||
"opus", subscriptionProfile("opus", "claude-account-b"));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
|
||||
() -> Map.of("term_primary", "primary"));
|
||||
|
||||
assertEquals(0, lookup.apply("sonnet"), "different accounts must never be conflated into one seat count "
|
||||
+ "— an explicit credentialId on both sides must still win over the subscription sentinel, so an "
|
||||
+ "operator with two separate Claude logins on one host can keep them apart");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #176 stage 2 — the exact live shape that shipped inert: a lead on subscription profile
|
||||
* {@code opus}, members on a DIFFERENTLY NAMED subscription profile {@code sonnet}, same Claude
|
||||
* login, and NEITHER profile sets {@code credentialId}. Every other test in this class puts the
|
||||
* lead on the SAME profile name as the target, which happened to keep working even with the old
|
||||
* fall-back-to-profile-name {@code effectiveCredentialId()} — this is the one that did not, and
|
||||
* its absence is what let the stage-1 fix ship without ever catching the bug it was filed for.
|
||||
*/
|
||||
@Test
|
||||
@DisplayName("[LIVE SHAPE] lead on a DIFFERENT subscription profile, same account, neither sets "
|
||||
+ "credentialId ⇒ still counts as a seat")
|
||||
void leadOnADifferentSubscriptionProfileSameAccountStillCountsAsASeat() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of(
|
||||
"opus", subscriptionProfile("opus", null),
|
||||
"sonnet", subscriptionProfile("sonnet", null));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
|
||||
() -> Map.of("term_primary", "primary"));
|
||||
|
||||
assertEquals(1, lookup.apply("sonnet"), "opus and sonnet are both subscription:true with no explicit "
|
||||
+ "credentialId, so they share one Claude login and the lead's live seat on opus must be charged "
|
||||
+ "against sonnet too — this is the live host's actual shape (fleetd #176 stage 2)");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("two live instances of the same lead count as two seats")
|
||||
void twoLiveInstancesOfTheSameLeadCountAsTwoSeats() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
|
||||
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
|
||||
() -> Map.of("term_a", "primary", "term_b", "primary"));
|
||||
|
||||
assertEquals(2, lookup.apply("sonnet"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unconfigured target profile resolves to zero, not a thrown exception")
|
||||
void unconfiguredTargetProfileIsZero() {
|
||||
Function<String, Integer> lookup = Fleetd.leadSeatLookup(Map::of, Map.of(), Map::of);
|
||||
|
||||
assertEquals(0, lookup.apply("ghost"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("live leads are read through the supplier on every call, not snapshotted")
|
||||
void liveLeadsAreReadThroughOnEveryCall() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
|
||||
java.util.Map<String, String> live = new java.util.HashMap<>();
|
||||
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders, () -> live);
|
||||
|
||||
assertEquals(0, lookup.apply("sonnet"));
|
||||
live.put("term_primary", "primary");
|
||||
assertEquals(1, lookup.apply("sonnet"));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,43 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #176: {@code Fleetd.main} builds its {@code FleetMcp} from a 14-argument constructor whose
|
||||
* last argument is a {@code FleetMcp.LeadSeatSource} wrapping {@link Fleetd#leadSeatLookup}. That
|
||||
* argument is exactly the kind of wiring fleetd #248 warned about: dropping it (or swapping it for
|
||||
* the inert {@code FleetMcp.LeadSeatSource.none()}) compiles with 0 errors and leaves every test
|
||||
* that builds its own {@code FleetMcp}/{@code CapacitySource} directly — every test that predates
|
||||
* this ticket — green, because none of them go through {@code main} at all.
|
||||
*
|
||||
* <p>{@link FleetdLeadSeatLookupTest} proves the factory's own matching logic; this class is the
|
||||
* plain source-text assertion that proves {@code main} still passes its result in, mirroring
|
||||
* {@code FleetdCompletionResolverWiringTest}'s approach for the same class of gap.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a
|
||||
* {@code FleetMcp} and never runs {@code main}.
|
||||
*/
|
||||
class FleetdLeadSeatWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] FleetMcp's construction call still passes a LeadSeatSource built from leadSeatLookup(...)")
|
||||
void fleetMcpConstructionStillWiresLeadSeatLookup() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains("new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), "
|
||||
+ "leaders, leads))"),
|
||||
"FleetMcp's construction call must still pass a LeadSeatSource built from "
|
||||
+ "Fleetd.leadSeatLookup(...). Dropping it or swapping in "
|
||||
+ "FleetMcp.LeadSeatSource.none() (fleetd #176's would-be silent regression, the same "
|
||||
+ "shape as fleetd #248's measured mutations) compiles with 0 errors and leaves every "
|
||||
+ "existing behavioural test green — this source check is what must go red instead.");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,64 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.function.Function;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
|
||||
/**
|
||||
* fleetd #248: {@link Fleetd#worktreeBranchLookup} is the factory that replaced the anonymous
|
||||
* lambda {@code Fleetd.main} used to build inline, as the 8th (last) argument to {@code
|
||||
* CompletionResolver}'s constructor. Before this ticket that argument was untestable wiring:
|
||||
* replacing it with {@code _ -> null} compiled clean and every existing test stayed green, because
|
||||
* every existing test builds its own {@code CompletionResolver} directly rather than going through
|
||||
* {@code main}. This class proves the factory's own behaviour; {@code
|
||||
* FleetdCompletionResolverWiringTest} proves {@code main} still passes it in.
|
||||
*/
|
||||
class FleetdWorktreeBranchLookupTest {
|
||||
|
||||
private static MemberSession session(String terminal, String worktree, String branch) {
|
||||
return new MemberSession("pane-" + terminal, terminal, "terra", MemberRole.DEV,
|
||||
"/cwd", null, 0L, 0L, 0, MemberSession.State.READY, worktree, branch);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a known target resolves to its session's worktree and branch")
|
||||
void knownTargetResolves() {
|
||||
Function<String, CompletionResolver.WorktreeBranch> lookup =
|
||||
Fleetd.worktreeBranchLookup(() -> List.of(session("term1", "/wt/worker_x", "worker/x")));
|
||||
|
||||
CompletionResolver.WorktreeBranch resolved = lookup.apply("term1");
|
||||
|
||||
assertEquals("/wt/worker_x", resolved.worktree());
|
||||
assertEquals("worker/x", resolved.branch());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unknown target resolves to null, not a thrown exception")
|
||||
void unknownTargetResolvesToNull() {
|
||||
Function<String, CompletionResolver.WorktreeBranch> lookup =
|
||||
Fleetd.worktreeBranchLookup(() -> List.of(session("term1", "/wt/worker_x", "worker/x")));
|
||||
|
||||
assertNull(lookup.apply("term_stranger"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("the roster is read through the supplier on every call, not snapshotted")
|
||||
void rosterIsReadThroughOnEveryCall() {
|
||||
List<MemberSession> roster = new ArrayList<>();
|
||||
Function<String, CompletionResolver.WorktreeBranch> lookup = Fleetd.worktreeBranchLookup(() -> roster);
|
||||
|
||||
assertNull(lookup.apply("term_late"));
|
||||
roster.add(session("term_late", "/wt/late", "worker/late"));
|
||||
|
||||
assertEquals("/wt/late", lookup.apply("term_late").worktree());
|
||||
}
|
||||
}
|
||||
@@ -1,8 +1,12 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
@@ -26,19 +30,66 @@ class RequiredSecretEnvVarsTest {
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
private static ListAppender<ILoggingEvent> attach() {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
return appender;
|
||||
}
|
||||
|
||||
private static void detach(ListAppender<ILoggingEvent> appender) {
|
||||
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
|
||||
}
|
||||
|
||||
@Test
|
||||
void collectsATokenEnvPerNonSubscriptionProfile(@TempDir Path dir) throws Exception {
|
||||
void explicitlyConfiguredUnsetTokenEnvWarnsWithCurrentWording(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = load(dir, """
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
tokenEnv: AI_GATEWAY_TOKEN
|
||||
tokenEnv: CB115_MISSING_TOKEN_ENV_819367
|
||||
""");
|
||||
|
||||
Map<String, List<String>> required = Fleetd.requiredSecretEnvVars(cfg);
|
||||
|
||||
assertTrue(required.containsKey("AI_GATEWAY_TOKEN"));
|
||||
assertEquals(List.of("profile 'local' tokenEnv"), required.get("AI_GATEWAY_TOKEN"));
|
||||
assertTrue(required.containsKey("CB115_MISSING_TOKEN_ENV_819367"));
|
||||
assertEquals(List.of("profile 'local' tokenEnv"), required.get("CB115_MISSING_TOKEN_ENV_819367"));
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
Fleetd.reportRequiredSecrets(cfg);
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
assertTrue(appender.list.stream().anyMatch(e ->
|
||||
e.getLevel() == ch.qos.logback.classic.Level.WARN
|
||||
&& e.getFormattedMessage().contains("startup secret CB115_MISSING_TOKEN_ENV_819367: MISSING")
|
||||
&& e.getFormattedMessage().contains("profile 'local' tokenEnv")),
|
||||
"an explicitly configured but unset tokenEnv must retain the startup warning");
|
||||
}
|
||||
|
||||
@Test
|
||||
void profileWithoutTokenEnvDoesNotRequireTheOldDefault(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = load(dir, """
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
""");
|
||||
|
||||
assertTrue(Fleetd.requiredSecretEnvVars(cfg).isEmpty(),
|
||||
"an absent tokenEnv means this profile needs no token, not FLEETD_WORKER_TOKEN");
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
Fleetd.reportRequiredSecrets(cfg);
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
assertFalse(appender.list.stream().anyMatch(e -> e.getLevel() == ch.qos.logback.classic.Level.WARN),
|
||||
"a profile without tokenEnv must not produce a startup secret warning");
|
||||
}
|
||||
|
||||
@Test
|
||||
|
||||
@@ -386,6 +386,54 @@ class ConfigRefTest {
|
||||
assertEquals("rate limit exceeded", ref.get().profiles().get("sonnet").exhaustedPattern());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #201 Unit 5: errorPattern is compiled once into Fleetd.main's backend-error pattern
|
||||
* map at startup (see BackendErrorPatternLookup), the same way exhaustedPattern is above — a
|
||||
* reload never re-reads it either, so a changed value must be reported deferred.
|
||||
*/
|
||||
@Test
|
||||
void changingAProfilesErrorPatternIsReportedAsDeferred(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet
|
||||
errorPattern: "credential outage"
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""");
|
||||
ConfigRef ref = refFor(f);
|
||||
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet
|
||||
errorPattern: "provider 5xx"
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""");
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertEquals(1, out.deferred().size(), out.deferred().toString());
|
||||
assertTrue(out.deferred().getFirst().contains("sonnet"), out.deferred().toString());
|
||||
assertTrue(out.deferred().getFirst().contains("launch settings"), out.deferred().toString());
|
||||
// The snapshot still carries the new value — a restart is what makes it take effect.
|
||||
assertEquals("provider 5xx", ref.get().profiles().get("sonnet").errorPattern());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aFixedRefHasNoFileAndRefusesToReload() {
|
||||
FleetConfig cfg = new FleetConfig(null, null, null, null, null, null,
|
||||
|
||||
@@ -6,11 +6,18 @@ import dev.ltms.fleet.peer.MemberRole;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.ParameterizedType;
|
||||
import java.lang.reflect.RecordComponent;
|
||||
import java.lang.reflect.Type;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.HashSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.TreeMap;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
@@ -86,6 +93,85 @@ class FleetConfigTest {
|
||||
"unset means off — today's behaviour, unchanged");
|
||||
}
|
||||
|
||||
// ── fleetd #201 Unit 5: errorPattern ────────────────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
void aProfileWithAnErrorPatternBindsAndHasErrorPatternIsTrue(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("error-pattern.yaml");
|
||||
Files.writeString(f, """
|
||||
profiles:
|
||||
ltms-local:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
errorPattern: "(?i)\\\\bcredential outage\\\\b"
|
||||
""");
|
||||
|
||||
FleetConfig.Profile w = FleetConfig.load(f).profiles().get("ltms-local");
|
||||
assertTrue(w.hasErrorPattern(), "errorPattern binds and enables fleetd #201 Unit 5 classification");
|
||||
assertEquals("(?i)\\bcredential outage\\b", w.errorPattern());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aProfileWithNoErrorPatternLeavesItNullAndHasErrorPatternIsFalse(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("no-error-pattern.yaml");
|
||||
Files.writeString(f, """
|
||||
profiles:
|
||||
ltms-local:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
""");
|
||||
|
||||
FleetConfig.Profile w = FleetConfig.load(f).profiles().get("ltms-local");
|
||||
assertNull(w.errorPattern(), "unset means 'use the built-in fallback' — never off");
|
||||
assertFalse(w.hasErrorPattern());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aBlankErrorPatternNormalizesToNullJustLikeUnset(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("blank-error-pattern.yaml");
|
||||
Files.writeString(f, """
|
||||
profiles:
|
||||
ltms-local:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
errorPattern: " "
|
||||
""");
|
||||
|
||||
FleetConfig.Profile w = FleetConfig.load(f).profiles().get("ltms-local");
|
||||
assertNull(w.errorPattern());
|
||||
assertFalse(w.hasErrorPattern());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aProfileWithAMalformedErrorPatternIsRejectedAtLoadNamingTheProfileAndKey(@TempDir Path dir)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("malformed-error-pattern.yaml");
|
||||
Files.writeString(f, """
|
||||
profiles:
|
||||
ltms-local:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
errorPattern: "(unterminated["
|
||||
""");
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class, () -> FleetConfig.load(f));
|
||||
assertTrue(e.getMessage().contains("ltms-local"), "the offending profile is named: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("errorPattern"), "the offending key is named: " + e.getMessage());
|
||||
}
|
||||
|
||||
@Test
|
||||
void withProfileCarriesErrorPatternThrough(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("with-profile-error-pattern.yaml");
|
||||
Files.writeString(f, """
|
||||
profiles:
|
||||
ltms-local:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
errorPattern: "credential outage"
|
||||
""");
|
||||
|
||||
// The compact constructor defaults `profile` to the map key via withProfile(name) — see
|
||||
// loadsFullConfig above. errorPattern must survive that same rebuild.
|
||||
FleetConfig.Profile w = FleetConfig.load(f).profiles().get("ltms-local");
|
||||
assertEquals("ltms-local", w.profile());
|
||||
assertEquals("credential outage", w.errorPattern());
|
||||
}
|
||||
|
||||
@Test
|
||||
void appliesDefaultsForMissingSections(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("minimal.yaml");
|
||||
@@ -1287,9 +1373,18 @@ class FleetConfigTest {
|
||||
}
|
||||
|
||||
/**
|
||||
* Every optional knob the example documents must bind under the exact spelling used there.
|
||||
* Keep this list in step with {@code fleetd.example.yaml}: a rename that updates the record
|
||||
* but not the example (or vice versa) fails here instead of silently no-op'ing in production.
|
||||
* A spot check that the knobs listed below bind under the exact spelling the example uses —
|
||||
* it asserts real VALUES arrive in the record, which no name-matching guard can do.
|
||||
*
|
||||
* <p><b>This is NOT a coverage guard, and must not be read as one</b> (fleetd #113). The list
|
||||
* inside it is hand-written, so it only ever covers what someone remembered to add. Coverage
|
||||
* — "is every key the code reads documented, and does every documented key bind?" — comes
|
||||
* from {@link #everyNestedConfigKeyIsDocumentedInTheExample} and
|
||||
* {@link #everyLiveKeyInTheExampleBindsToARecordComponent}, both of which derive their key
|
||||
* set from the record tree and therefore cannot drift.
|
||||
*
|
||||
* <p>Adding a knob here is optional. Leaving one out is not a coverage gap, because the two
|
||||
* derived guards above already fail on an undocumented or unbindable key.
|
||||
*/
|
||||
@Test
|
||||
void everyOptionalKnobDocumentedInTheExampleBinds(@TempDir Path dir) throws Exception {
|
||||
@@ -1313,6 +1408,7 @@ class FleetConfigTest {
|
||||
weight: 0.5
|
||||
maxLoad: 2
|
||||
exhaustedPattern: "usage limit has been reached"
|
||||
errorPattern: "credential outage"
|
||||
placement: weighted
|
||||
lifecycle:
|
||||
idleTtlSeconds: 300
|
||||
@@ -1350,6 +1446,8 @@ class FleetConfigTest {
|
||||
assertEquals(2, w.maxLoad(), "maxLoad binds as an integer");
|
||||
assertTrue(w.hasExhaustedPattern(), "exhaustedPattern binds and enables the CB-578 stage A classification");
|
||||
assertEquals("usage limit has been reached", w.exhaustedPattern());
|
||||
assertTrue(w.hasErrorPattern(), "errorPattern binds and enables fleetd #201 Unit 5 classification");
|
||||
assertEquals("credential outage", w.errorPattern());
|
||||
|
||||
assertEquals("weighted", cfg.placement(), "placement binds at the top level");
|
||||
assertEquals(300, cfg.lifecycle().idleTtlSeconds());
|
||||
@@ -1443,6 +1541,230 @@ class FleetConfigTest {
|
||||
return p.matcher(yaml).find();
|
||||
}
|
||||
|
||||
/**
|
||||
* The nested half of {@link #everyKnownTopLevelKeyIsDocumentedInTheExample} (fleetd #113).
|
||||
*
|
||||
* <p>That guard walks {@link FleetConfig#KNOWN_TOP_LEVEL_KEYS} and anchors its regex at
|
||||
* column 0, so it sees ONLY top-level keys. Every nested key — {@code profiles.<name>.model},
|
||||
* {@code health.paneProbeIntervalSeconds} and a hundred others — is outside its scope, and
|
||||
* neither its name nor its output says so. A green run then reads as "the example documents
|
||||
* the schema" when most of the schema was never looked at.
|
||||
*
|
||||
* <p>This walks the record tree rather than a name list, so a key added to any nested record
|
||||
* is covered the moment it compiles, with no edit here. That is the point: a hand-maintained
|
||||
* second copy of a list always drifts from the thing it mirrors.
|
||||
*
|
||||
* <p><b>Scope, stated on purpose</b> (fleetd #113 criterion 3 — every check reports what it
|
||||
* did and did not look at):
|
||||
* <ul>
|
||||
* <li>It checks each key NAME appears somewhere in the example as a YAML key, live or
|
||||
* commented out. It does NOT check the key sits at the right path.</li>
|
||||
* <li>It does NOT check a documented key is read by anything. {@code paneProbeIntervalSeconds}
|
||||
* is parsed into {@link FleetConfig.Health} and used nowhere, and this guard passes it.
|
||||
* Proving a key is live code needs a call graph, which this is not.</li>
|
||||
* </ul>
|
||||
*/
|
||||
@Test
|
||||
void everyNestedConfigKeyIsDocumentedInTheExample() throws Exception {
|
||||
Path example = Path.of("fleetd.example.yaml");
|
||||
assertTrue(Files.exists(example), "fleetd.example.yaml must ship next to the pom");
|
||||
String text = Files.readString(example);
|
||||
|
||||
Map<String, String> pathByName = configKeyPaths();
|
||||
|
||||
// The denominator. An under-counting walk passes every subset check vacuously, which is
|
||||
// the exact shape fleetd #113 collects — so the walk has to prove it descended at all.
|
||||
// The floor is DERIVED, not a literal: the nested walk must find substantially more keys
|
||||
// than the top-level set the old guard used, or it has not gone below the first level.
|
||||
int topLevel = FleetConfig.KNOWN_TOP_LEVEL_KEYS.size();
|
||||
assertTrue(pathByName.size() > topLevel * 2,
|
||||
"the record walk found " + pathByName.size() + " config key(s) against "
|
||||
+ topLevel + " top-level key(s) — it has stopped descending into the "
|
||||
+ "nested records, so this guard would pass vacuously. Fix the walk "
|
||||
+ "before trusting a green run.");
|
||||
|
||||
List<String> undocumented = pathByName.entrySet().stream()
|
||||
.filter(e -> !keyDocumentedAnywhere(text, e.getKey()))
|
||||
.map(Map.Entry::getValue)
|
||||
.sorted()
|
||||
.toList();
|
||||
|
||||
assertTrue(undocumented.isEmpty(), () -> "checked " + pathByName.size()
|
||||
+ " config key(s) that FleetConfig can bind; " + undocumented.size()
|
||||
+ " appear nowhere in fleetd.example.yaml: " + undocumented
|
||||
+ " — document each one there, commented out if optional. fleetd.yaml is "
|
||||
+ "gitignored, so the example is the only committed description of the schema.");
|
||||
}
|
||||
|
||||
/**
|
||||
* The other direction: a LIVE key in the example that {@link FleetConfig} cannot bind. That is
|
||||
* a key an operator would copy into {@code fleetd.yaml} expecting it to do something, where it
|
||||
* would be silently ignored.
|
||||
*
|
||||
* <p><b>Scope, stated on purpose:</b> only live (uncommented) keys are checked. Most of the
|
||||
* example is commented-out prose, and that prose contains lines like {@code # mode: token}
|
||||
* that are indistinguishable from keys by text alone. Parsing them would produce false
|
||||
* failures, so they are deliberately out of scope — and saying so here is the point, rather
|
||||
* than letting a reader assume the whole file was validated.
|
||||
*/
|
||||
@Test
|
||||
void everyLiveKeyInTheExampleBindsToARecordComponent() throws Exception {
|
||||
Path example = Path.of("fleetd.example.yaml");
|
||||
String text = Files.readString(example);
|
||||
|
||||
List<List<String>> paths = liveKeyPaths(text);
|
||||
assertTrue(paths.size() >= 20,
|
||||
"only " + paths.size() + " live key path(s) were parsed out of the example — the "
|
||||
+ "parser is not seeing the file, so this guard would pass vacuously.");
|
||||
|
||||
List<String> unbindable = paths.stream()
|
||||
.filter(path -> !pathBinds(path))
|
||||
.map(path -> String.join(".", path))
|
||||
.distinct()
|
||||
.sorted()
|
||||
.toList();
|
||||
|
||||
assertTrue(unbindable.isEmpty(), () -> "checked " + paths.size()
|
||||
+ " live key path(s) in fleetd.example.yaml; " + unbindable.size()
|
||||
+ " bind to nothing in FleetConfig: " + unbindable
|
||||
+ " — an operator copying one of these into fleetd.yaml gets silence, not an error.");
|
||||
}
|
||||
|
||||
/**
|
||||
* Every configuration key {@link FleetConfig} can bind, at every depth, as
|
||||
* {@code name -> a dotted path to one place it appears}. Derived from the record components,
|
||||
* so it cannot drift from the code.
|
||||
*/
|
||||
private static Map<String, String> configKeyPaths() {
|
||||
Map<String, String> out = new TreeMap<>();
|
||||
collectConfigKeys(FleetConfig.class, "", new HashSet<>(), out);
|
||||
return out;
|
||||
}
|
||||
|
||||
private static void collectConfigKeys(Class<?> type, String prefix, Set<String> seen,
|
||||
Map<String, String> out) {
|
||||
if (!type.isRecord() || !seen.add(type.getName())) {
|
||||
return;
|
||||
}
|
||||
for (RecordComponent rc : type.getRecordComponents()) {
|
||||
String path = prefix.isEmpty() ? rc.getName() : prefix + "." + rc.getName();
|
||||
out.putIfAbsent(rc.getName(), path);
|
||||
Class<?> nested = rc.getType();
|
||||
if (nested.isRecord()) {
|
||||
collectConfigKeys(nested, path, seen, out);
|
||||
} else if (Map.class.isAssignableFrom(nested) || List.class.isAssignableFrom(nested)) {
|
||||
Class<?> element = elementRecord(rc);
|
||||
if (element != null) {
|
||||
String childPrefix = Map.class.isAssignableFrom(nested)
|
||||
? path + ".<name>" : path + "[]";
|
||||
collectConfigKeys(element, childPrefix, seen, out);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/** The record type inside a {@code Map<String, X>} or {@code List<X>} component, else null. */
|
||||
private static Class<?> elementRecord(RecordComponent rc) {
|
||||
if (rc.getGenericType() instanceof ParameterizedType pt) {
|
||||
Type[] args = pt.getActualTypeArguments();
|
||||
if (args.length > 0 && args[args.length - 1] instanceof Class<?> c && c.isRecord()) {
|
||||
return c;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/**
|
||||
* True when {@code key} is documented in the example, in either of the two conventions that
|
||||
* file actually uses:
|
||||
* <ol>
|
||||
* <li>as a YAML key at any indentation, live or commented out ({@code key:}); or</li>
|
||||
* <li>in a prose block that describes a section's sub-keys, one per line, as
|
||||
* {@code # key → what it does}.</li>
|
||||
* </ol>
|
||||
*
|
||||
* <p>The second form is not decoration. {@code broker.uri} is documented ONLY that way, on
|
||||
* purpose: writing it out as a copy-pasteable {@code uri: amqp://user:pass@host} invites an
|
||||
* operator to paste a password into a file, which is the very thing {@code uriEnv} exists to
|
||||
* avoid. A guard that demanded the key form would push the file toward doing that. So this
|
||||
* encodes the convention the example really uses rather than imposing a new one.
|
||||
*/
|
||||
private static boolean keyDocumentedAnywhere(String yaml, String key) {
|
||||
String quoted = Pattern.quote(key);
|
||||
Pattern asYamlKey = Pattern.compile("(?m)^\\s*(?:#\\s*)?" + quoted + ":");
|
||||
Pattern asProseEntry = Pattern.compile("(?m)^\\s*#\\s*" + quoted + "\\s+\u2192");
|
||||
return asYamlKey.matcher(yaml).find() || asProseEntry.matcher(yaml).find();
|
||||
}
|
||||
|
||||
/** Every live (uncommented) key in {@code yaml}, as a path from the document root. */
|
||||
private static List<List<String>> liveKeyPaths(String yaml) {
|
||||
Pattern keyLine = Pattern.compile("^(\\s*)([A-Za-z][A-Za-z0-9_]*):(\\s.*)?$");
|
||||
List<String> stack = new ArrayList<>();
|
||||
List<Integer> indents = new ArrayList<>();
|
||||
List<List<String>> paths = new ArrayList<>();
|
||||
for (String line : yaml.split("\n", -1)) {
|
||||
if (line.isBlank() || line.stripLeading().startsWith("#")) {
|
||||
continue;
|
||||
}
|
||||
Matcher m = keyLine.matcher(line);
|
||||
if (!m.matches()) {
|
||||
continue;
|
||||
}
|
||||
int indent = m.group(1).length();
|
||||
while (!indents.isEmpty() && indents.get(indents.size() - 1) >= indent) {
|
||||
indents.remove(indents.size() - 1);
|
||||
stack.remove(stack.size() - 1);
|
||||
}
|
||||
indents.add(indent);
|
||||
stack.add(m.group(2));
|
||||
paths.add(List.copyOf(stack));
|
||||
}
|
||||
return paths;
|
||||
}
|
||||
|
||||
/** True when a dotted YAML path resolves to something {@link FleetConfig} can bind. */
|
||||
private static boolean pathBinds(List<String> path) {
|
||||
Class<?> type = FleetConfig.class;
|
||||
boolean nextSegmentIsAFreeFormName = false;
|
||||
for (int i = 0; i < path.size(); i++) {
|
||||
if (nextSegmentIsAFreeFormName) {
|
||||
nextSegmentIsAFreeFormName = false;
|
||||
continue;
|
||||
}
|
||||
RecordComponent rc = componentNamed(type, path.get(i));
|
||||
if (rc == null) {
|
||||
return false;
|
||||
}
|
||||
Class<?> t = rc.getType();
|
||||
if (t.isRecord()) {
|
||||
type = t;
|
||||
} else if (Map.class.isAssignableFrom(t)) {
|
||||
Class<?> element = elementRecord(rc);
|
||||
if (element == null) {
|
||||
return true; // Map<String,String>: its entries are data, not schema
|
||||
}
|
||||
type = element;
|
||||
nextSegmentIsAFreeFormName = true;
|
||||
} else {
|
||||
// A scalar or a list of scalars: nothing may legitimately nest under it.
|
||||
return i == path.size() - 1;
|
||||
}
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
private static RecordComponent componentNamed(Class<?> type, String name) {
|
||||
if (type == null || !type.isRecord()) {
|
||||
return null;
|
||||
}
|
||||
for (RecordComponent rc : type.getRecordComponents()) {
|
||||
if (rc.getName().equals(name)) {
|
||||
return rc;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
@Test
|
||||
void placementDefaultsToFixedForExistingConfigs(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("no-placement.yaml");
|
||||
@@ -1781,7 +2103,7 @@ class FleetConfigTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void parityOverlayDefaultsToEnvFilesOnly(@TempDir Path dir) throws Exception {
|
||||
void parityOverlayDefaultsToDotEnvOnly(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("no-overlay.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
@@ -1792,9 +2114,11 @@ class FleetConfigTest {
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertEquals(List.of(".env", ".envrc"),
|
||||
assertEquals(List.of(".env"),
|
||||
cfg.profiles().get("gx10").parityOverlay(),
|
||||
"the default parity overlay is the env files; settings.local.json is no longer copied by default");
|
||||
"the default parity overlay is .env only (CB-148): .envrc is executable shell that "
|
||||
+ "direnv runs on every cd, so it is no longer copied by default; settings.local.json "
|
||||
+ "is also not copied by default");
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -1815,6 +2139,25 @@ class FleetConfigTest {
|
||||
"an operator's explicit list survives verbatim — the default only changes when unset");
|
||||
}
|
||||
|
||||
@Test
|
||||
void parityOverlayExplicitEnvrcOptInStillWorks(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("explicit-envrc-overlay.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
parityOverlay: [".env", ".envrc"]
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertEquals(List.of(".env", ".envrc"),
|
||||
cfg.profiles().get("gx10").parityOverlay(),
|
||||
"an operator can still opt into copying .envrc explicitly (CB-148); it is only "
|
||||
+ "dropped from the unset default, not removed as a capability");
|
||||
}
|
||||
|
||||
// ── CB-542: subscription:true must not smuggle an unguarded endpoint via env: ───────────────
|
||||
|
||||
@Test
|
||||
|
||||
@@ -49,6 +49,7 @@ public final class FakeHerdr implements HerdrClient {
|
||||
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
|
||||
private String pinnedStartTerminal;
|
||||
private String pinnedStartPane;
|
||||
private Runnable onAgentStart; // fires the instant agent.start is called — see onAgentStart(Runnable)
|
||||
private volatile int agentGetOkCalls = Integer.MAX_VALUE; // how many agent.get calls succeed first
|
||||
private volatile String agentGetFailCode = null; // error code every agent.get call after that reports
|
||||
|
||||
@@ -152,6 +153,18 @@ public final class FakeHerdr implements HerdrClient {
|
||||
return this;
|
||||
}
|
||||
|
||||
/**
|
||||
* Run {@code hook} synchronously the instant an {@code agent.start} call reaches this fake —
|
||||
* i.e. the instant the peer PROCESS would start against a real herdr daemon. A test uses this
|
||||
* to assert something is already true at that exact point (rather than merely true once
|
||||
* {@code spawn()} returns), e.g. fleetd #149's trust-dialog seed having already been written to
|
||||
* disk before the process herdr would launch ever starts.
|
||||
*/
|
||||
public FakeHerdr onAgentStart(Runnable hook) {
|
||||
this.onAgentStart = hook;
|
||||
return this;
|
||||
}
|
||||
|
||||
|
||||
/**
|
||||
* Seed a named agent into {@code agent.list} (e.g. an orphaned worker for CB-117 reaper tests).
|
||||
@@ -250,6 +263,9 @@ public final class FakeHerdr implements HerdrClient {
|
||||
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
|
||||
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
|
||||
case "agent.start" -> {
|
||||
if (onAgentStart != null) {
|
||||
onAgentStart.run();
|
||||
}
|
||||
// Protocol 19: kind and pane_id are required — reject like the real daemon.
|
||||
java.util.Map<?, ?> p = params instanceof java.util.Map<?, ?> m ? m : java.util.Map.of();
|
||||
for (String required : new String[]{"kind", "pane_id"}) {
|
||||
|
||||
@@ -0,0 +1,272 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.fleet.Fleetd;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.mcp.PrimaryRegistry;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||
import dev.ltms.fleet.msg.InMemoryReplyInbox;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.msg.ReplyPushLoop;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementException;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Optional;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.CopyOnWriteArrayList;
|
||||
import java.util.concurrent.CountDownLatch;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* fleetd #201 / #227 Unit 5: proves the WIRED production path — a real {@link CompletionResolver}
|
||||
* pane-scrape classification, through the exact {@code BackendErrorSink} ordering {@code Fleetd.main}
|
||||
* builds (mirrored here line for line since the lambda itself lives inline in {@code Fleetd.main}),
|
||||
* into a real {@link BackendOutagePolicy}, a real {@link CompositePeerLauncher} spawn gate, and a
|
||||
* real {@link ReplyPushLoop} lead nudge — never by calling {@link BackendOutagePolicy#record} or the
|
||||
* sink directly, which is what the finer-grained unit tests in
|
||||
* {@code dev.ltms.fleet.member.CompositePeerLauncherTest} and {@code ReplyPushLoopTest} do for their
|
||||
* own concerns. This class exists to catch a wiring mistake those unit tests cannot see: each of them
|
||||
* hands the collaborator a value it assumes was computed correctly one layer up.
|
||||
*
|
||||
* <p>{@code fleet_list}/{@code fleet_profiles} rendering of the resulting cool-off (criteria 11/12)
|
||||
* is proven separately in {@code dev.ltms.fleet.mcp.FleetMcpTest} — {@code FleetMcp}'s view methods
|
||||
* are package-private to {@code dev.ltms.fleet.mcp} and this class needs {@code CompletionResolver}'s
|
||||
* package-private {@code InFlight}, so the two cannot share one test class. What matters for THIS
|
||||
* test is only that the same {@link BackendOutagePolicy} object the real scrape path populates is the
|
||||
* one {@code FleetMcp} reads from — proven by asserting directly against
|
||||
* {@link BackendOutagePolicy#remainingCoolOffSeconds} after the real classification below.
|
||||
*/
|
||||
class BackendOutageFlowTest {
|
||||
|
||||
private final List<ScheduledExecutorService> schedulers = new ArrayList<>();
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
schedulers.forEach(ScheduledExecutorService::shutdownNow);
|
||||
}
|
||||
|
||||
private static FleetConfig.Profile stubWorker(String profile, String credentialId) {
|
||||
return new FleetConfig.Profile(profile, "http://gx00.gw:8000", "coder",
|
||||
null, "FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"w #{n}", null, null, null, null, null, null, null, null, null,
|
||||
null, null, credentialId, null);
|
||||
}
|
||||
|
||||
/** Minimal recording {@code HerdrClient} for the LEAD pane — mirrors ReplyPushLoopTest's own. */
|
||||
private static final class RecordingLeadClient implements HerdrClient {
|
||||
private static final ObjectMapper MAPPER = new ObjectMapper();
|
||||
private final List<Object> prompts = new CopyOnWriteArrayList<>();
|
||||
volatile CountDownLatch sendLatch = new CountDownLatch(1);
|
||||
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) {
|
||||
if ("agent.get".equals(method)) {
|
||||
return MAPPER.createObjectNode().set("agent", MAPPER.createObjectNode()
|
||||
.put("terminal_id", "term_primary").put("agent_status", "idle"));
|
||||
}
|
||||
if ("agent.prompt".equals(method)) {
|
||||
prompts.add(params);
|
||||
sendLatch.countDown();
|
||||
}
|
||||
return MAPPER.createObjectNode();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
|
||||
int sendCount() {
|
||||
return prompts.size();
|
||||
}
|
||||
|
||||
String lastText() {
|
||||
return ((Map<?, ?>) prompts.get(prompts.size() - 1)).get("text").toString();
|
||||
}
|
||||
}
|
||||
|
||||
/** Every collaborator the flow wires together, built exactly once per test. */
|
||||
private final class Flow {
|
||||
final FakeHerdr herdr = new FakeHerdr()
|
||||
.readText("⏺ 503 Service Unavailable: upstream credential rejected\n❯ ");
|
||||
final Map<String, FleetConfig.Profile> profiles = orderedProfiles();
|
||||
final AtomicLong clockNanos = new AtomicLong(0L);
|
||||
final BackendOutagePolicy outagePolicy = new BackendOutagePolicy(clockNanos::get);
|
||||
final CompositePeerLauncher workers;
|
||||
final SessionManager sessions;
|
||||
final MemberSession s1;
|
||||
final MemberSession s2;
|
||||
final PrimaryRegistry registry = new PrimaryRegistry(null);
|
||||
final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||
final RecordingLeadClient leadClient = new RecordingLeadClient();
|
||||
final ReplyPushLoop pushLoop;
|
||||
final CompletionResolver resolver;
|
||||
final Rendezvous rendezvous = new Rendezvous();
|
||||
|
||||
Flow() {
|
||||
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr),
|
||||
new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
|
||||
profiles, "terra", _ -> "tok");
|
||||
workers = new CompositePeerLauncher(List.of(adapter), "terra", profiles,
|
||||
PlacementPolicies.weighted(), _ -> 0, null, BackendQuarantine.none(), outagePolicy);
|
||||
sessions = new SessionManager(workers);
|
||||
// Two REAL member sessions on "terra" — the roster lookup the production sink does.
|
||||
s1 = sessions.acquire("terra", null, null, null);
|
||||
s2 = sessions.acquire("terra", null, null, null);
|
||||
registry.recordDelegation(s1.terminalId(), "term_primary");
|
||||
registry.recordDelegation(s2.terminalId(), "term_primary");
|
||||
inbox.own(s1.terminalId());
|
||||
inbox.own(s2.terminalId());
|
||||
|
||||
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||
schedulers.add(scheduler);
|
||||
pushLoop = new ReplyPushLoop(registry, new AgentControl(leadClient), inbox, scheduler, 3, 50);
|
||||
AtomicReference<ReplyPushLoop> pushLoopRef = new AtomicReference<>(pushLoop);
|
||||
|
||||
// fleetd #248 follow-up: this used to be a 30-line hand-copy of Fleetd.main's
|
||||
// backendErrorSink lambda, with a comment promising it mirrored production "EXACTLY".
|
||||
// That promise is exactly the problem: a copy proves the copy. Editing or deleting the
|
||||
// real sink left this whole flow test green, because it never touched the real sink.
|
||||
// #248 made Fleetd.backendErrorSink public precisely so a cross-package test could
|
||||
// drive the real object, so this now calls it. Every assertion below is about
|
||||
// production code again.
|
||||
BackendErrorSink backendErrorSink =
|
||||
Fleetd.backendErrorSink(sessions, () -> profiles, outagePolicy, pushLoopRef::get);
|
||||
|
||||
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
|
||||
resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(),
|
||||
ExhaustionSink.none(), patterns, backendErrorSink);
|
||||
}
|
||||
|
||||
/** Drives one real pane-scrape classification for {@code target} through {@code resolver}. */
|
||||
void classifyBackendErrorOn(String target) {
|
||||
var waiter = rendezvous.open(target);
|
||||
resolver.resolve(target, new CompletionResolver.InFlight(waiter, null));
|
||||
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
|
||||
"a classified backend error still resolves the blocked send as a failure");
|
||||
}
|
||||
}
|
||||
|
||||
private static Map<String, FleetConfig.Profile> orderedProfiles() {
|
||||
Map<String, FleetConfig.Profile> m = new LinkedHashMap<>();
|
||||
m.put("terra", stubWorker("terra", "shared-openai"));
|
||||
m.put("sol", stubWorker("sol", "shared-openai"));
|
||||
return m;
|
||||
}
|
||||
|
||||
private long agentStartCalls(FakeHerdr herdr) {
|
||||
return herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count();
|
||||
}
|
||||
|
||||
// --- criteria 4/5: two DISTINCT targets start an incident; one target never does ------------
|
||||
|
||||
@Test
|
||||
void aSingleClassifiedBackendErrorOnOneTargetNeverCoolsTheCredential() {
|
||||
Flow flow = new Flow();
|
||||
|
||||
flow.classifyBackendErrorOn(flow.s1.terminalId());
|
||||
|
||||
assertTrue(flow.outagePolicy.remainingCoolOffSeconds("shared-openai").isEmpty(),
|
||||
"one distinct target's classified error must not start a cool-off");
|
||||
assertEquals(0, flow.leadClient.sendCount(), "no incident yet, so no lead nudge");
|
||||
assertDoesNotThrow(() -> flow.workers.spawn(new SpawnRequest("terra", null, null)),
|
||||
"terra must still be spawnable after only one target's error");
|
||||
}
|
||||
|
||||
@Test
|
||||
void twoDistinctTargetsClassifiedThroughTheRealScrapeStartAnIncidentAndCoolTheCredential() throws Exception {
|
||||
Flow flow = new Flow();
|
||||
|
||||
flow.classifyBackendErrorOn(flow.s1.terminalId());
|
||||
assertTrue(flow.outagePolicy.remainingCoolOffSeconds("shared-openai").isEmpty());
|
||||
|
||||
flow.classifyBackendErrorOn(flow.s2.terminalId());
|
||||
|
||||
assertTrue(flow.leadClient.sendLatch.await(3, TimeUnit.SECONDS),
|
||||
"the second distinct target must cross the threshold and nudge the lead");
|
||||
Thread.sleep(150);
|
||||
var remaining = flow.outagePolicy.remainingCoolOffSeconds("shared-openai");
|
||||
assertTrue(remaining.isPresent(), "two distinct targets must start a cool-off");
|
||||
assertEquals(60L, remaining.getAsLong(),
|
||||
"this is the exact fact FleetMcp.capacityView/profiles reads as coolingOffForSeconds "
|
||||
+ "(see FleetMcpTest for the rendering side)");
|
||||
}
|
||||
|
||||
// --- criterion 13: the incident reaches the lead through ReplyPushLoop exactly once ----------
|
||||
|
||||
@Test
|
||||
void theIncidentSendsExactlyOneNudgeToTheLeadNamingBothCredentialAndTargets() throws Exception {
|
||||
Flow flow = new Flow();
|
||||
|
||||
flow.classifyBackendErrorOn(flow.s1.terminalId());
|
||||
flow.classifyBackendErrorOn(flow.s2.terminalId());
|
||||
|
||||
assertTrue(flow.leadClient.sendLatch.await(3, TimeUnit.SECONDS));
|
||||
Thread.sleep(200);
|
||||
assertEquals(1, flow.leadClient.sendCount(), "one incident must send exactly one nudge");
|
||||
String nudge = flow.leadClient.lastText();
|
||||
assertTrue(nudge.contains("shared-openai"), nudge);
|
||||
assertTrue(nudge.contains(flow.s1.terminalId()) && nudge.contains(flow.s2.terminalId()), nudge);
|
||||
assertTrue(nudge.contains("60"), nudge);
|
||||
}
|
||||
|
||||
// --- criterion 6: explicit spawn is refused BEFORE the adapter is ever called -----------------
|
||||
|
||||
@Test
|
||||
void explicitSpawnOnACooledCredentialIsRefusedBeforeReachingTheAdapter() throws Exception {
|
||||
Flow flow = new Flow();
|
||||
flow.classifyBackendErrorOn(flow.s1.terminalId());
|
||||
flow.classifyBackendErrorOn(flow.s2.terminalId());
|
||||
assertTrue(flow.leadClient.sendLatch.await(3, TimeUnit.SECONDS));
|
||||
|
||||
long startsBefore = agentStartCalls(flow.herdr);
|
||||
PlacementException e = assertThrows(PlacementException.class,
|
||||
() -> flow.workers.spawn(new SpawnRequest("terra", null, null)));
|
||||
|
||||
assertTrue(e.getMessage().contains("cooling off"), e.getMessage());
|
||||
assertTrue(e.getMessage().contains("shared-openai"), e.getMessage());
|
||||
assertEquals(startsBefore, agentStartCalls(flow.herdr),
|
||||
"the refusal must happen before the adapter is ever called — no new agent.start");
|
||||
}
|
||||
|
||||
// --- criteria 7/9: automatic placement skips every profile sharing the cooled credential -----
|
||||
|
||||
@Test
|
||||
void automaticPlacementRefusesAndNamesCoolingOffWhenEveryCandidateSharesTheCredential() throws Exception {
|
||||
Flow flow = new Flow();
|
||||
flow.classifyBackendErrorOn(flow.s1.terminalId());
|
||||
flow.classifyBackendErrorOn(flow.s2.terminalId());
|
||||
assertTrue(flow.leadClient.sendLatch.await(3, TimeUnit.SECONDS));
|
||||
|
||||
PlacementException e = assertThrows(PlacementException.class,
|
||||
() -> flow.workers.spawn(new SpawnRequest(null, null, null)),
|
||||
"terra and sol both share the cooled-off credential — nothing is available");
|
||||
assertTrue(e.getMessage().contains("cooling off"), e.getMessage());
|
||||
assertThrows(PlacementException.class, () -> flow.workers.spawn(new SpawnRequest("sol", null, null)),
|
||||
"sol shares terra's credential, so it must be locked out too");
|
||||
}
|
||||
}
|
||||
@@ -4,8 +4,11 @@ import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.LoggerContext;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.msg.TestTurnTokens;
|
||||
import dev.ltms.fleet.msg.TurnToken;
|
||||
@@ -193,6 +196,26 @@ class CompletionResolverTest {
|
||||
assertEquals("complete report", waiter.getNow(null).text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void suppressesABareEchoWithOnlyTuiChrome() {
|
||||
String injected = "Load the implementer skill. You own fleetd #999. ".repeat(12);
|
||||
String scrape = injected + "\nDev auto - GPT-5.6 Terra OpenAI";
|
||||
|
||||
assertTrue(CompletionResolver.echoesInjectedBrief(scrape, injected));
|
||||
}
|
||||
|
||||
@Test
|
||||
void pinsTheMaximumTuiChromeExcess() {
|
||||
String injected = "a".repeat(CompletionResolver.ECHO_MIN_CHARS);
|
||||
String underMargin = injected + "b".repeat(CompletionResolver.MAX_ECHO_EXCESS_CHARS);
|
||||
String overMargin = injected + "b".repeat(CompletionResolver.MAX_ECHO_EXCESS_CHARS + 1);
|
||||
|
||||
assertTrue(CompletionResolver.echoesInjectedBrief(underMargin, injected),
|
||||
"the configured excess itself remains an echoed brief");
|
||||
assertFalse(CompletionResolver.echoesInjectedBrief(overMargin, injected),
|
||||
"one character beyond the excess must preserve the scrape as a real report");
|
||||
}
|
||||
|
||||
@Test
|
||||
void resolvesSynchronouslyBeforePostTurnContextClearing() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ previous answer\n❯ ");
|
||||
@@ -499,7 +522,7 @@ class CompletionResolverTest {
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
|
||||
java.util.List<String> notified = new java.util.ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
|
||||
ExhaustionSink sink = (target, reason, profile) -> notified.add(target + ": " + reason);
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
@@ -520,7 +543,7 @@ class CompletionResolverTest {
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
|
||||
java.util.List<String> notified = new java.util.ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
|
||||
ExhaustionSink sink = (target, reason, profile) -> notified.add(target + ": " + reason);
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
@@ -648,7 +671,7 @@ class CompletionResolverTest {
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
|
||||
java.util.List<String> notified = new java.util.ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
|
||||
ExhaustionSink sink = (target, reason, profile) -> notified.add(target + ": " + reason);
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
@@ -701,7 +724,7 @@ class CompletionResolverTest {
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
|
||||
java.util.List<String> notified = new java.util.ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
|
||||
ExhaustionSink sink = (target, reason, profile) -> notified.add(target + ": " + reason);
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
@@ -728,7 +751,7 @@ class CompletionResolverTest {
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
|
||||
java.util.List<String> notified = new java.util.ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
|
||||
ExhaustionSink sink = (target, reason, profile) -> notified.add(target + ": " + reason);
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
@@ -745,18 +768,234 @@ class CompletionResolverTest {
|
||||
@Test
|
||||
void coverageIsOffWhenNoProfileHasAPatternConfigured() {
|
||||
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [terra])",
|
||||
CompletionResolver.coverage(Set.of("terra"), Set.of()));
|
||||
CompletionResolver.coverage("exhaustedPattern", Set.of("terra"), Set.of()));
|
||||
}
|
||||
|
||||
@Test
|
||||
void coverageIsFullWhenEveryProfileHasAPatternConfigured() {
|
||||
assertEquals("full (all profiles configured: [gx10, terra])",
|
||||
CompletionResolver.coverage(Set.of("terra", "gx10"), Set.of("terra", "gx10")));
|
||||
CompletionResolver.coverage("exhaustedPattern", Set.of("terra", "gx10"), Set.of("terra", "gx10")));
|
||||
}
|
||||
|
||||
@Test
|
||||
void coverageIsPartialAndNamesWhichProfilesAreConfigured() {
|
||||
assertEquals("partial (configured: [terra]; not configured: [gx10])",
|
||||
CompletionResolver.coverage(Set.of("terra", "gx10"), Set.of("terra")));
|
||||
CompletionResolver.coverage("exhaustedPattern", Set.of("terra", "gx10"), Set.of("terra")));
|
||||
}
|
||||
|
||||
/**
|
||||
* Found live on 2026-09-03, reading a real boot log rather than a test. {@code coverage} is
|
||||
* shared by two call sites — CB-578's {@code exhaustedPattern} line and fleetd#201 Unit 5's
|
||||
* {@code errorPattern} line — but its "off" branch hard-coded the word {@code exhaustedPattern}.
|
||||
* So a daemon with no {@code errorPattern} anywhere printed "no profile has an exhaustedPattern
|
||||
* configured" directly beneath a line reporting that two profiles DO have one. Both lines were
|
||||
* individually defensible and together they were nonsense, and the message sent an operator to
|
||||
* set the wrong key.
|
||||
*
|
||||
* <p>Every earlier test here passed the exhaustion case only, so none of them could see it. This
|
||||
* one pins that the message names the key the caller actually meant.
|
||||
*/
|
||||
@Test
|
||||
void coverageNamesTheConfigKeyItsCallerMeansRatherThanAlwaysSayingExhaustedPattern() {
|
||||
assertEquals("off (no profile has an errorPattern configured; profiles: [gx10, terra])",
|
||||
CompletionResolver.coverage("errorPattern", Set.of("terra", "gx10"), Set.of()));
|
||||
}
|
||||
|
||||
// --- fleetd#201 Unit 1: target-keyed backend-error pattern + typed sink ----------------------
|
||||
|
||||
@Test
|
||||
void aConfiguredBackendErrorPatternClassifiesAMatchAsAFailureAndNotifiesTheSinkOnce() {
|
||||
String block = "⏺ 503 Service Unavailable: upstream credential rejected\n❯ ";
|
||||
FakeHerdr herdr = new FakeHerdr().readText(block);
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
|
||||
java.util.List<String> notified = new java.util.ArrayList<>();
|
||||
BackendErrorSink sink = (target, matchedLine, reason) -> notified.add(target + ": " + matchedLine);
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none(), patterns, sink);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
||||
|
||||
assertTrue(waiter.isDone(), "a configured-pattern match still resolves the blocked send");
|
||||
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
|
||||
assertEquals(1, notified.size(), "the sink fires exactly once for the winning classification");
|
||||
assertEquals("term_a: 503 Service Unavailable: upstream credential rejected", notified.get(0));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aLosingBackendErrorClassificationNeverNotifiesTheSink() {
|
||||
// The waiter was already resolved (e.g. by the worker's own fleet_reply) before this scrape
|
||||
// landed — resolveFailure loses the race and must return false, so the sink must not fire.
|
||||
String block = "⏺ 503 Service Unavailable: upstream credential rejected\n❯ ";
|
||||
FakeHerdr herdr = new FakeHerdr().readText(block);
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
|
||||
java.util.List<String> notified = new java.util.ArrayList<>();
|
||||
BackendErrorSink sink = (target, matchedLine, reason) -> notified.add(target + ": " + matchedLine);
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none(), patterns, sink);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
var turn = new CompletionResolver.InFlight(waiter, null);
|
||||
assertTrue(rendezvous.resolveCompletion(waiter, "already replied")); // fleet_reply won first
|
||||
|
||||
resolver.resolve("term_a", turn);
|
||||
|
||||
assertTrue(notified.isEmpty(), "a classification that loses the race must not fire the sink");
|
||||
assertEquals("already replied", waiter.getNow(null).text(), "the earlier resolution stands untouched");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aBackendErrorClassificationThatLosesARealRaceDuringTheScrapeNeverNotifiesTheSink() {
|
||||
// The test above pre-resolves the waiter BEFORE calling resolve(), so it is caught by
|
||||
// resolve()'s own top-of-method isDone() guard before ever reaching the win-gated
|
||||
// resolveFailure call this unit adds — it proves the outcome, not the specific gate. This
|
||||
// test forces the actual race window fleetd#201 must be safe in: a genuine fleet_reply lands
|
||||
// WHILE the herdr scrape for this turn is in flight, i.e. AFTER resolve() has already passed
|
||||
// the top guard (waiter not done yet) and committed to reading the pane, but BEFORE it
|
||||
// reaches the "if (rendezvous.resolveFailure(...))" line below. The side effect is attached
|
||||
// to the one real I/O step resolve() performs between those two points: the herdr
|
||||
// "agent.read" call itself.
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
var waiter = rendezvous.open("term_a");
|
||||
ObjectMapper mapper = new ObjectMapper();
|
||||
HerdrClient racingDuringScrape = new HerdrClient() {
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) {
|
||||
if ("agent.read".equals(method)) {
|
||||
// The real fleet_reply that wins the race, landing mid-scrape.
|
||||
assertTrue(rendezvous.resolve("term_a", "a genuine fleet_reply landed first"),
|
||||
"the racing reply must land while the waiter is still open");
|
||||
}
|
||||
try {
|
||||
return mapper.readTree(mapper.writeValueAsString(java.util.Map.of(
|
||||
"type", "agent_read",
|
||||
"read", java.util.Map.of("text",
|
||||
"⏺ 503 Service Unavailable: upstream credential rejected\n❯ "))));
|
||||
} catch (Exception e) {
|
||||
throw new RuntimeException(e);
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
};
|
||||
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
|
||||
java.util.List<String> notified = new java.util.ArrayList<>();
|
||||
BackendErrorSink sink = (target, matchedLine, reason) -> notified.add(target + ": " + matchedLine);
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(racingDuringScrape), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none(), patterns, sink);
|
||||
|
||||
var turn = new CompletionResolver.InFlight(waiter, null); // not done yet — passes the top guard
|
||||
resolver.resolve("term_a", turn);
|
||||
|
||||
assertTrue(notified.isEmpty(),
|
||||
"a classification that loses a real race during the scrape must not fire the sink");
|
||||
assertEquals(Rendezvous.Kind.REPLY, waiter.getNow(null).kind(), "the genuine fleet_reply stands");
|
||||
assertEquals("a genuine fleet_reply landed first", waiter.getNow(null).text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void exhaustionKeepsWinningOverBackendErrorEvenWhenBothPatternsMatchTheSameLine() {
|
||||
// A line matching both a configured exhausted pattern AND a configured backend-error pattern
|
||||
// must classify BACKEND_EXHAUSTED and call only the ExhaustionSink — the ordering fleetd#578
|
||||
// already relies on must not change.
|
||||
String block = "⏺ The usage limit has been reached. Try again later.\n❯ ";
|
||||
FakeHerdr herdr = new FakeHerdr().readText(block);
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
ExhaustedPatternLookup exhausted = target -> Pattern.compile("usage limit has been reached");
|
||||
BackendErrorPatternLookup backendErrors = target -> Pattern.compile("(?i)usage limit");
|
||||
java.util.List<String> exhaustedNotified = new java.util.ArrayList<>();
|
||||
java.util.List<String> backendErrorNotified = new java.util.ArrayList<>();
|
||||
ExhaustionSink exhaustionSink = (target, reason, profile) -> exhaustedNotified.add(target);
|
||||
BackendErrorSink backendErrorSink = (target, matchedLine, reason) -> backendErrorNotified.add(target);
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
exhausted, exhaustionSink, backendErrors, backendErrorSink);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
||||
|
||||
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, waiter.getNow(null).kind(),
|
||||
"exhaustion must keep winning over backend error");
|
||||
assertEquals(1, exhaustedNotified.size(), "the exhaustion sink fires");
|
||||
assertTrue(backendErrorNotified.isEmpty(), "the backend-error sink must never fire for this turn");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aConfiguredPatternAlsoClassifiesTheRawScrapeFallbackAndNotifiesTheSink() {
|
||||
// No ⏺ marker and leading TUI chrome ⇒ lastAssistantBlock() yields "", so classification must
|
||||
// fall back to the raw scrape (fleetd#211) — and it must use the configured pattern too.
|
||||
String block = """
|
||||
╭──────────────────────────────────────╮
|
||||
503 Service Unavailable: upstream credential rejected
|
||||
""";
|
||||
FakeHerdr herdr = new FakeHerdr().readText(block);
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
|
||||
java.util.List<String> notified = new java.util.ArrayList<>();
|
||||
BackendErrorSink sink = (target, matchedLine, reason) -> notified.add(target + ": " + matchedLine);
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none(), patterns, sink);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
||||
|
||||
assertTrue(waiter.isDone(), "a raw-scrape configured-pattern match still resolves the blocked send");
|
||||
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
|
||||
"classified from the raw scrape even though the trimmed block was empty");
|
||||
assertEquals(1, notified.size(), "the sink fires exactly once from the raw-scrape path");
|
||||
assertTrue(notified.get(0).contains("503 Service Unavailable"), notified.get(0));
|
||||
}
|
||||
|
||||
// --- fleetd#201 Unit 1: classification inside the fleetd#164 MIN_TURN_NANOS floor -------------
|
||||
|
||||
@Test
|
||||
void aMatchingErrorInsideTheFloorIsTypedAndNotifiesTheSink() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ 503 Service Unavailable: upstream credential rejected\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
long[] clock = {10_000_000_000L};
|
||||
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
|
||||
java.util.List<String> notified = new java.util.ArrayList<>();
|
||||
BackendErrorSink sink = (target, matchedLine, reason) -> notified.add(target + ": " + matchedLine);
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none(), patterns, sink, () -> clock[0]);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
var turn = new CompletionResolver.InFlight(waiter, null, clock[0]); // delivered "now"
|
||||
clock[0] += CompletionResolver.MIN_TURN_NANOS - 1; // inside the floor
|
||||
|
||||
resolver.resolve("term_a", turn);
|
||||
|
||||
assertTrue(waiter.isDone());
|
||||
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
|
||||
assertEquals(1, notified.size(), "a matching error inside the floor notifies the sink");
|
||||
assertTrue(notified.get(0).contains("503 Service Unavailable"), notified.get(0));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aNonMatchInsideTheFloorStaysGenericAndNeverNotifiesTheSink() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ still starting up\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
long[] clock = {10_000_000_000L};
|
||||
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
|
||||
java.util.List<String> notified = new java.util.ArrayList<>();
|
||||
BackendErrorSink sink = (target, matchedLine, reason) -> notified.add(target + ": " + matchedLine);
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none(), patterns, sink, () -> clock[0]);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
var turn = new CompletionResolver.InFlight(waiter, null, clock[0]); // delivered "now"
|
||||
clock[0] += CompletionResolver.MIN_TURN_NANOS - 1; // inside the floor
|
||||
|
||||
resolver.resolve("term_a", turn);
|
||||
|
||||
assertTrue(waiter.isDone());
|
||||
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
|
||||
"still a failure — the floor itself, not the pattern, is why");
|
||||
assertTrue(waiter.getNow(null).text().contains("too fast to be real work"),
|
||||
"a non-match inside the floor stays the generic too-fast reason: " + waiter.getNow(null).text());
|
||||
assertTrue(notified.isEmpty(), "a non-match must never notify the typed sink");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,41 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* fleetd #234. Round 3: calls the REAL production factory, {@link ExhaustionSink#forwardingTo},
|
||||
* rather than rebuilding its shape locally — a round-2 version of this test built its own copy of
|
||||
* the forwarder inline, so mutating {@code Fleetd.java}'s actual forwarder back into a lambda left
|
||||
* that copy untouched and the test kept passing while production had regressed. Calling the shared
|
||||
* factory here means the same object under test IS the object {@code Fleetd.java} builds via
|
||||
* {@code ExhaustionSink.forwardingTo(exhaustionSinkRef::get)}.
|
||||
*
|
||||
* <p>Round 4: the 3-arg overload is now {@link ExhaustionSink}'s single abstract method, so {@code
|
||||
* forwardingTo} itself is a one-line lambda and every sink below can safely be one too — there is
|
||||
* no 2-arg overload left for any of them to silently bind to instead. This test still catches a
|
||||
* regression inside {@code forwardingTo}'s body (e.g. one that calls the 2-arg default and drops
|
||||
* the hint that way) because it still goes through the shared factory rather than a rebuilt copy.
|
||||
*/
|
||||
class ExhaustionSinkForwardingHazardTest {
|
||||
|
||||
@Test
|
||||
void forwardingToDeliversTheProfileHintToWhateverSinkTheSupplierCurrentlyReturns() {
|
||||
AtomicReference<String> hintSeenByRealSink = new AtomicReference<>("NEVER CALLED");
|
||||
ExhaustionSink real = (target, reason, profileHint) -> hintSeenByRealSink.set(profileHint);
|
||||
|
||||
// Same shape Fleetd.java uses: a reference that starts at none() and is repointed later.
|
||||
AtomicReference<ExhaustionSink> ref = new AtomicReference<>(ExhaustionSink.none());
|
||||
ExhaustionSink forwarding = ExhaustionSink.forwardingTo(ref::get);
|
||||
ref.set(real);
|
||||
|
||||
forwarding.onExhausted("term_x", "model mismatch", "gx");
|
||||
|
||||
assertEquals("gx", hintSeenByRealSink.get(),
|
||||
"ExhaustionSink.forwardingTo must deliver the profile hint to the sink the supplier "
|
||||
+ "currently returns — the exact object Fleetd.java's forwarder is built from");
|
||||
}
|
||||
}
|
||||
@@ -20,6 +20,7 @@ import dev.ltms.fleet.session.MemberSession;
|
||||
import dev.ltms.fleet.session.WorktreeRequest;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
@@ -500,6 +501,25 @@ class FleetMcpTest {
|
||||
assertTrue(out.contains("\"quarantinedForSeconds\":1800"), out);
|
||||
}
|
||||
|
||||
/** fleetd #201 Unit 5: {@code coolingOff} is a SEPARATE map from {@code quarantined}. */
|
||||
@Test
|
||||
void profilesReportsACoolingOffCredentialInASeparateMap() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(() -> 0L);
|
||||
outagePolicy.record("shared-openai", "t1", "API Error: rate limited");
|
||||
outagePolicy.record("shared-openai", "t2", "API Error: rate limited"); // 2nd distinct target starts the incident
|
||||
FleetMcp.OutageSource outage = new FleetMcp.OutageSource(
|
||||
profile -> "ltms-local".equals(profile) ? "shared-openai" : null, outagePolicy);
|
||||
McpSchema.CallToolResult res = FleetMcp.profiles(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), FleetMcp.QuarantineSource.none(), outage);
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"coolingOff\""), out);
|
||||
assertTrue(out.contains("shared-openai"), out);
|
||||
assertTrue(out.contains("\"coolingOffForSeconds\":60"), out);
|
||||
assertFalse(out.contains("\"quarantined\""), "nothing is exhaustion-quarantined here: " + out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void listReportsTrackedWorkers() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
@@ -631,6 +651,123 @@ class FleetMcpTest {
|
||||
assertEquals(2, out.split("\"free\":0", -1).length - 1, out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #176 stage 2 (correcting the inert stage 1): two {@code subscription: true} profiles,
|
||||
* {@code opus} and {@code sonnet}, neither setting an explicit {@code credentialId} — the exact
|
||||
* shape measured on the live Mac fleet. This is INTENDED, not a regression: a real Claude usage
|
||||
* limit on the one login behind both profiles really does take out every profile running on it,
|
||||
* the same way {@code credentialId: openai-shared} already lets two OpenAI-backed profiles share
|
||||
* one quarantine (see {@code everyProfileSharingTheQuarantinedCredentialReportsZeroFree} above).
|
||||
* The {@code credentialIdFor} function here is built the same way {@code Fleetd.main} wires it —
|
||||
* {@code profile -> profiles.get(profile).effectiveCredentialId()} — so this proves the actual
|
||||
* config-driven behaviour, not just {@code capacityView}'s arithmetic with a hand-picked string.
|
||||
*/
|
||||
@Test
|
||||
void quarantiningOneSubscriptionProfileZeroesFreeOnTheOtherSharingTheSameAccount() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of(
|
||||
"opus", new FleetConfig.Profile("opus", null, "claude-opus-4", null, null, null,
|
||||
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
|
||||
null, 3, true, null, null, null),
|
||||
"sonnet", new FleetConfig.Profile("sonnet", null, "claude-sonnet-5", null, null, null,
|
||||
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
|
||||
null, 3, true, null, null, null));
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(20));
|
||||
quarantine.quarantine(FleetConfig.Profile.SUBSCRIPTION_CREDENTIAL_ID);
|
||||
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(profile -> {
|
||||
FleetConfig.Profile configured = profiles.get(profile);
|
||||
return configured == null ? null : configured.effectiveCredentialId();
|
||||
}, quarantine);
|
||||
|
||||
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 3,
|
||||
() -> Set.of("opus", "sonnet"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
source, Map.of(), ""));
|
||||
|
||||
assertEquals(2, out.split("\"free\":0", -1).length - 1,
|
||||
"opus and sonnet share one Claude login with neither setting credentialId, so quarantining "
|
||||
+ "opus's account must also zero sonnet's free — this is intended, not a side effect: "
|
||||
+ out);
|
||||
assertEquals(2, out.split("\"credentialId\":\"" + FleetConfig.Profile.SUBSCRIPTION_CREDENTIAL_ID + "\"", -1)
|
||||
.length - 1, out);
|
||||
}
|
||||
|
||||
/** fleetd #201 Unit 5: cool-off forces {@code free:0} but never adds {@code quarantinedForSeconds}. */
|
||||
@Test
|
||||
void coolingOffProfileReportsZeroFreeButNeverQuarantinedForSeconds() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(() -> 0L);
|
||||
outagePolicy.record("shared-openai", "t1", "API Error: rate limited");
|
||||
outagePolicy.record("shared-openai", "t2", "API Error: rate limited");
|
||||
FleetMcp.OutageSource outage = new FleetMcp.OutageSource(
|
||||
profile -> "terra".equals(profile) ? "shared-openai" : null, outagePolicy);
|
||||
|
||||
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 2,
|
||||
() -> Set.of("terra"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), outage, Map.of(), ""));
|
||||
|
||||
assertTrue(out.contains("\"profile\":\"terra\""), out);
|
||||
assertTrue(out.contains("\"free\":0"), out);
|
||||
assertTrue(out.contains("\"credentialId\":\"shared-openai\""), out);
|
||||
assertTrue(out.contains("\"coolingOffForSeconds\":60"), out);
|
||||
assertFalse(out.contains("quarantinedForSeconds"),
|
||||
"exhaustion quarantine was never active — this key must not appear: " + out);
|
||||
}
|
||||
|
||||
/**
|
||||
* The two checks are independent — a profile can be BOTH exhaustion-quarantined AND cooling off
|
||||
* for the same credential at once, and both maps/fields report it simultaneously.
|
||||
*/
|
||||
@Test
|
||||
void bothQuarantineAndCoolingOffCanFireForTheSameProfileAtOnce() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(20));
|
||||
quarantine.quarantine("shared-openai");
|
||||
FleetMcp.QuarantineSource quarantineSource = new FleetMcp.QuarantineSource(
|
||||
profile -> "terra".equals(profile) ? "shared-openai" : null, quarantine);
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(() -> 0L);
|
||||
outagePolicy.record("shared-openai", "t1", "API Error: rate limited");
|
||||
outagePolicy.record("shared-openai", "t2", "API Error: rate limited");
|
||||
FleetMcp.OutageSource outage = new FleetMcp.OutageSource(
|
||||
profile -> "terra".equals(profile) ? "shared-openai" : null, outagePolicy);
|
||||
|
||||
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 2,
|
||||
() -> Set.of("terra"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
quarantineSource, outage, Map.of(), ""));
|
||||
|
||||
assertTrue(out.contains("\"free\":0"), out);
|
||||
assertTrue(out.contains("\"quarantinedForSeconds\":1200"), out);
|
||||
assertTrue(out.contains("\"coolingOffForSeconds\":60"), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void listAndProfilesAgreeOnWhatIsCoolingOff() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(() -> 0L);
|
||||
outagePolicy.record("shared-openai", "t1", "API Error: rate limited");
|
||||
outagePolicy.record("shared-openai", "t2", "API Error: rate limited");
|
||||
FleetMcp.OutageSource outage = new FleetMcp.OutageSource(
|
||||
profile -> "ltms-local".equals(profile) ? "shared-openai" : null, outagePolicy);
|
||||
var workers = workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
String listOut = textOf(FleetMcp.listFleet(workers, sessions, null,
|
||||
new FleetMcp.CapacitySource(profile -> 0, profile -> 2, () -> Set.of("ltms-local"), () -> 0),
|
||||
new FleetMcp.HealthCoverageSource(() -> "off"), FleetMcp.QuarantineSource.none(), outage,
|
||||
Map.of(), ""));
|
||||
String profilesOut = textOf(FleetMcp.profiles(workers, FleetMcp.QuarantineSource.none(), outage));
|
||||
|
||||
assertTrue(listOut.contains("\"free\":0"), listOut);
|
||||
assertTrue(listOut.contains("\"credentialId\":\"shared-openai\""), listOut);
|
||||
assertTrue(profilesOut.contains("\"coolingOff\""), profilesOut);
|
||||
assertTrue(profilesOut.contains("shared-openai"), profilesOut);
|
||||
}
|
||||
|
||||
@Test
|
||||
void listAndProfilesAgreeOnWhatIsQuarantined() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
@@ -672,6 +809,68 @@ class FleetMcpTest {
|
||||
assertFalse(out.contains("quarantinedForSeconds"), out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #176: this is the exact shape measured on the Mac fleet — {@code maxLoad:3, live:2},
|
||||
* where one of the "free" three is really the lead's own seat on the same subscription. The old
|
||||
* formula ({@code max(0, cap - live)}) reported {@code free:1}; the real ceiling is {@code 0}
|
||||
* (two members plus the lead's own seat already fill all three).
|
||||
*/
|
||||
@Test
|
||||
void leadSeatSubtractsFromFreeTheSameWayLiveDoes() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
FleetMcp.LeadSeatSource leadSeats = new FleetMcp.LeadSeatSource(
|
||||
profile -> "sonnet".equals(profile) ? 1 : 0);
|
||||
|
||||
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||
sessions, null, new FleetMcp.CapacitySource(profile -> 2, profile -> 3,
|
||||
() -> Set.of("sonnet"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), leadSeats, Map.of(), "", null));
|
||||
|
||||
assertTrue(out.contains("\"maxLoad\":3"), "maxLoad itself must be left untouched: " + out);
|
||||
assertTrue(out.contains("\"live\":2"), out);
|
||||
assertTrue(out.contains("\"free\":0"), "2 live + 1 lead seat fills all 3: " + out);
|
||||
assertTrue(out.contains("\"leadSeats\":1"), out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #176: the OTHER measurement in the issue — a completely idle fleet still overstates
|
||||
* {@code free} by the lead's own seat. {@code maxLoad:3, live:0} must report {@code free:2}, the
|
||||
* real fan-out ceiling, not {@code 3}.
|
||||
*/
|
||||
@Test
|
||||
void leadSeatLowersFreeOnAnOtherwiseIdleSubscriptionProfile() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
FleetMcp.LeadSeatSource leadSeats = new FleetMcp.LeadSeatSource(
|
||||
profile -> "sonnet".equals(profile) ? 1 : 0);
|
||||
|
||||
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 3,
|
||||
() -> Set.of("sonnet"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), leadSeats, Map.of(), "", null));
|
||||
|
||||
assertTrue(out.contains("\"live\":0"), out);
|
||||
assertTrue(out.contains("\"free\":2"), "an idle fleet's real ceiling is 3 minus the lead's own seat: " + out);
|
||||
assertTrue(out.contains("\"leadSeats\":1"), out);
|
||||
}
|
||||
|
||||
/** A profile with no lead seats reported must be byte-identical to before this ticket. */
|
||||
@Test
|
||||
void zeroLeadSeatsOmitsTheKeyAndLeavesFreeUnchanged() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
|
||||
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 2,
|
||||
() -> Set.of("terra"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", null));
|
||||
|
||||
assertTrue(out.contains("\"free\":2"), out);
|
||||
assertFalse(out.contains("leadSeats"), "no lead shares this profile's credential: " + out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void listReportsLeadsAndFlagsTheCallersOwnRow() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
|
||||
@@ -0,0 +1,121 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.Set;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #114 (CB-609): the guard that lets {@code docs/MCP-Contract.md} name a tool at all.
|
||||
*
|
||||
* <p>That page was written in July 2026, before any MCP code existed, and then did not follow the
|
||||
* code. By August it named two tools that had never been built, omitted five that shipped, had the
|
||||
* wrong name for nearly every parameter, and still described a caller-identity rule that was a
|
||||
* privilege bug by then. Nothing failed, because nothing checked it — and {@code CLAUDE.md} sends
|
||||
* every session in the fleet to that page.
|
||||
*
|
||||
* <p>The fix was to delete the tool catalogue rather than correct it: a hand-maintained second copy
|
||||
* of the tool surface is the defect, not the particular errors it had accumulated. What survives is
|
||||
* the flows, which are shapes rather than names. But the flows still have to say {@code fleet_send}
|
||||
* somewhere to be readable, and that is exactly the sentence that rots. This test is what makes it
|
||||
* safe to write.
|
||||
*
|
||||
* <p><b>It checks source text, not behaviour.</b> It reads the Markdown and reads {@link FleetMcp}'s
|
||||
* source, and it only catches a name in the doc that the server does not register. It cannot catch a
|
||||
* flow that describes the wrong order, or a parameter name in prose — those are not name-shaped. The
|
||||
* doc's own header carries that caveat for its readers.
|
||||
*/
|
||||
class McpContractDocTest {
|
||||
|
||||
/** Tests run with the module directory as cwd, so the repo-root doc is one level up. */
|
||||
private static final Path DOC = Path.of("../docs/MCP-Contract.md");
|
||||
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
|
||||
|
||||
private static Set<String> matches(Path file, String regex) throws Exception {
|
||||
Matcher m = Pattern.compile(regex).matcher(Files.readString(file));
|
||||
Set<String> found = new LinkedHashSet<>();
|
||||
while (m.find()) {
|
||||
found.add(m.group(1));
|
||||
}
|
||||
return found;
|
||||
}
|
||||
|
||||
/** Every {@code fleet_*} the doc mentions, in prose or in a diagram. */
|
||||
private static Set<String> toolsNamedInTheDoc() throws Exception {
|
||||
return matches(DOC, "(fleet_[a-z_]+)");
|
||||
}
|
||||
|
||||
/** Every tool {@link FleetMcp} actually registers, read from its {@code tool("…")} calls. */
|
||||
private static Set<String> toolsTheServerRegisters() throws Exception {
|
||||
return matches(MCP_SOURCE, "tool\\(\"(fleet_[a-z_]+)\"");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] every fleet_* tool named in MCP-Contract.md is one the server registers")
|
||||
void theDocNamesNoToolThatDoesNotExist() throws Exception {
|
||||
Set<String> registered = toolsTheServerRegisters();
|
||||
Set<String> named = toolsNamedInTheDoc();
|
||||
|
||||
Set<String> unknown = new LinkedHashSet<>(named);
|
||||
unknown.removeAll(registered);
|
||||
|
||||
assertTrue(unknown.isEmpty(),
|
||||
"docs/MCP-Contract.md names " + unknown + ", which FleetMcp does not register. "
|
||||
+ "Checked " + named.size() + " name(s) in the doc against " + registered.size()
|
||||
+ " registered tool(s): " + registered + ". This is the fleetd #114 defect "
|
||||
+ "recurring — the doc named fleet_read and fleet_cancel for weeks after the "
|
||||
+ "code shipped without them. Either fix the name or drop it from the page; do "
|
||||
+ "NOT weaken this test.");
|
||||
}
|
||||
|
||||
/**
|
||||
* The denominator guard. The check above passes trivially if the doc stops naming any tool at
|
||||
* all — an empty set is a subset of everything. A checker that can silently check nothing is the
|
||||
* fleetd #113 shape, so this pins that the doc really is still describing the flows, and that
|
||||
* the registration scrape really did find the server's tools.
|
||||
*/
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] the doc/server name check is not vacuous — both sides found names")
|
||||
void theCheckActuallyHasSomethingToCheck() throws Exception {
|
||||
Set<String> registered = toolsTheServerRegisters();
|
||||
Set<String> named = toolsNamedInTheDoc();
|
||||
|
||||
assertTrue(registered.size() >= 10,
|
||||
"scraped only " + registered.size() + " tool registrations from FleetMcp (" + registered
|
||||
+ "); the server registers eleven, so the tool(\"…\") scrape has stopped matching "
|
||||
+ "and the check above is now vacuous");
|
||||
assertTrue(named.size() >= 4,
|
||||
"docs/MCP-Contract.md names only " + named.size() + " fleet_* tool(s) (" + named + "). "
|
||||
+ "The flows describe delegation, clarification, detached delivery and the "
|
||||
+ "turn-done fallback, so it should name several. Too few means the page has been "
|
||||
+ "gutted and this test is guarding nothing.");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #114's actual lesson. The catalogue was deleted on purpose; a well-meaning "let me just
|
||||
* document the tools here" restores the exact second copy that drifted for a month.
|
||||
*/
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] MCP-Contract.md still says it is not the tool reference")
|
||||
void theDocStillDisclaimsBeingTheToolReference() throws Exception {
|
||||
String doc = Files.readString(DOC);
|
||||
assertTrue(doc.contains("**What this page is NOT: a tool reference.**"),
|
||||
"docs/MCP-Contract.md must keep saying it is not the tool reference. That sentence is "
|
||||
+ "the fix for fleetd #114: the page carried a hand-maintained tool catalogue that "
|
||||
+ "drifted from the code for a month while CLAUDE.md pointed every session at it.");
|
||||
assertEquals(0, countTables(doc.substring(0, doc.indexOf("## 1. Rendezvous flows"))),
|
||||
"the header of docs/MCP-Contract.md must not grow a tool/parameter table — that is the "
|
||||
+ "second copy fleetd #114 deleted");
|
||||
}
|
||||
|
||||
private static int countTables(String markdown) {
|
||||
return (int) markdown.lines().filter(l -> l.strip().startsWith("|")).count();
|
||||
}
|
||||
}
|
||||
@@ -4,6 +4,9 @@ import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.fasterxml.jackson.databind.node.ObjectNode;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.GuardException;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
@@ -16,6 +19,8 @@ import dev.ltms.fleet.peer.PeerHandle;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.PeerUnreachableException;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import org.junit.jupiter.api.AfterAll;
|
||||
import org.junit.jupiter.api.BeforeAll;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
@@ -23,11 +28,17 @@ import org.slf4j.LoggerFactory;
|
||||
import java.io.IOException;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.nio.file.attribute.PosixFileAttributeView;
|
||||
import java.nio.file.attribute.PosixFilePermissions;
|
||||
import java.util.HashSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.TreeSet;
|
||||
import java.util.UUID;
|
||||
import java.util.concurrent.CountDownLatch;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicBoolean;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.Supplier;
|
||||
@@ -139,18 +150,25 @@ class ClaudeCodeLauncherTest {
|
||||
// every ide_* call to the member's own worktree via the charter.
|
||||
|
||||
/** A profile carrying an ideMcpUrl (plus optional bridge mcpUrl and cwd). ideMcpUrl is the last record component. */
|
||||
private FleetConfig.Profile ideProfile(String mcpUrl, String ideMcpUrl, String cwd) {
|
||||
/**
|
||||
* {@code configDir} is first and mandatory on purpose (fleetd #258). A profile that sets no
|
||||
* {@code configDir} sends {@code seedTrustDialog}'s write to the operator's real
|
||||
* {@code ~/.claude.json}, and the fleetd #149 gate does not stop that when the fixture's
|
||||
* {@code cwd} is worktree-shaped — which every IDE-overlay fixture's is. Pass a {@code @TempDir}
|
||||
* whenever {@code cwd} has a {@code .git} FILE; {@code null} is only safe when it does not.
|
||||
*/
|
||||
private FleetConfig.Profile ideProfile(String configDir, String mcpUrl, String ideMcpUrl, String cwd) {
|
||||
return new FleetConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", configDir, "FLEETD_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "fleetd-workers", "w #{n}", mcpUrl, cwd, null,
|
||||
null, null, null, null, null, null, null, null, null, ideMcpUrl);
|
||||
}
|
||||
|
||||
/** As {@link #ideProfile} but carrying the CB-634 auto-open fields (module subdir + open command). */
|
||||
private FleetConfig.Profile ideProfileModule(String ideMcpUrl, String cwd, String ideProjectDir,
|
||||
String ideOpenCommand) {
|
||||
private FleetConfig.Profile ideProfileModule(String configDir, String ideMcpUrl, String cwd,
|
||||
String ideProjectDir, String ideOpenCommand) {
|
||||
return new FleetConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", configDir, "FLEETD_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "fleetd-workers", "w #{n}", null, cwd, null,
|
||||
null, null, null, null, null, null, null, null, null, ideMcpUrl, ideProjectDir, ideOpenCommand);
|
||||
}
|
||||
@@ -163,7 +181,7 @@ class ClaudeCodeLauncherTest {
|
||||
@Test
|
||||
void mountsIdeMcpAsSecondServerWhenIdeMcpUrlSet() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
launcher(herdr, ideProfile("http://127.0.0.1:8765/mcp",
|
||||
launcher(herdr, ideProfile(null, "http://127.0.0.1:8765/mcp",
|
||||
"http://127.0.0.1:29170/index-mcp/streamable-http", null)).spawn();
|
||||
|
||||
List<String> args = spawnedArgs(herdr);
|
||||
@@ -178,7 +196,8 @@ class ClaudeCodeLauncherTest {
|
||||
@Test
|
||||
void ideMcpUrlAloneStillEmitsTheMount() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
launcher(herdr, ideProfile(null, "http://127.0.0.1:29170/index-mcp/streamable-http", null)).spawn();
|
||||
launcher(herdr, ideProfile(null, null, "http://127.0.0.1:29170/index-mcp/streamable-http", null))
|
||||
.spawn();
|
||||
|
||||
List<String> args = spawnedArgs(herdr);
|
||||
assertTrue(args.contains("--mcp-config"),
|
||||
@@ -193,7 +212,7 @@ class ClaudeCodeLauncherTest {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
String roleCharter = "You review changes.";
|
||||
String worktree = "/tmp/.fleet-worktrees/rev-1";
|
||||
FleetConfig.Profile cfg = ideProfile("http://127.0.0.1:8765/mcp",
|
||||
FleetConfig.Profile cfg = ideProfile(null, "http://127.0.0.1:8765/mcp",
|
||||
"http://127.0.0.1:29170/index-mcp/streamable-http", worktree);
|
||||
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
|
||||
@@ -218,6 +237,62 @@ class ClaudeCodeLauncherTest {
|
||||
}, "the --append-system-prompt-file path must be a readable file");
|
||||
}
|
||||
|
||||
// ---- fleetd #258: no fixture in this class may seed the operator's real ~/.claude.json ----
|
||||
//
|
||||
// seedTrustDialog writes <configDir>/.claude.json, or ~/.claude.json when the profile sets no
|
||||
// configDir. The fleetd #149 gate (isProvisionedWorktree) closes the case where a fixture leaves
|
||||
// cwd unset and it falls back to user.dir. It does NOT close the case where a fixture builds a
|
||||
// worktree-shaped @TempDir on purpose — the gate opens, and a null configDir still points the
|
||||
// write at the real home. Two IDE-overlay fixtures did exactly that on EVERY run; by 2026-09-03
|
||||
// the operator's ~/.claude.json carried 116 dead JUnit temp paths, none of which still existed.
|
||||
//
|
||||
// DIFFERENTIAL, not absolute: it snapshots the temp-dir project keys already present and fails
|
||||
// only on keys this class ADDS. An absolute check would fail on any host still carrying the
|
||||
// historical entries, and a check that fails for a reason nobody can fix gets deleted, not fixed.
|
||||
|
||||
private static Set<String> tempProjectKeysBefore;
|
||||
|
||||
@BeforeAll
|
||||
static void snapshotTempProjectKeysInTheDefaultClaudeJson() {
|
||||
tempProjectKeysBefore = tempProjectKeysInDefaultClaudeJson();
|
||||
}
|
||||
|
||||
@AfterAll
|
||||
static void noFixtureSeededTheDefaultClaudeJson() {
|
||||
Set<String> added = new TreeSet<>(tempProjectKeysInDefaultClaudeJson());
|
||||
added.removeAll(tempProjectKeysBefore);
|
||||
assertTrue(added.isEmpty(),
|
||||
"a fixture in this class seeded the DEFAULT .claude.json (the operator's real file "
|
||||
+ "when user.home is not redirected) with " + added.size() + " temp-dir "
|
||||
+ "project entry/entries: " + added + ". Give that fixture's profile a "
|
||||
+ "@TempDir configDir — see ideProfile's javadoc.");
|
||||
}
|
||||
|
||||
/**
|
||||
* Project keys under the JVM temp dir in {@code <user.home>/.claude.json}, or an empty set when
|
||||
* the file is absent or unreadable. Only key NAMES are read; nothing in the operator's file is
|
||||
* copied, asserted on, or written back.
|
||||
*/
|
||||
private static Set<String> tempProjectKeysInDefaultClaudeJson() {
|
||||
Set<String> keys = new TreeSet<>();
|
||||
Path target = Path.of(System.getProperty("user.home"), ".claude.json");
|
||||
if (!Files.isRegularFile(target)) {
|
||||
return keys;
|
||||
}
|
||||
String tmp = System.getProperty("java.io.tmpdir");
|
||||
try {
|
||||
JsonNode projects = new ObjectMapper().readTree(target.toFile()).path("projects");
|
||||
projects.fieldNames().forEachRemaining(name -> {
|
||||
if (name.startsWith(tmp) || name.contains("/junit-")) {
|
||||
keys.add(name);
|
||||
}
|
||||
});
|
||||
} catch (IOException e) {
|
||||
return keys; // unreadable file proves nothing either way
|
||||
}
|
||||
return keys;
|
||||
}
|
||||
|
||||
// CB-634: the IDE guidance is delivered as a CLAUDE.local.md overlay (written only into a
|
||||
// provisioned worktree — cwd with a `.git` FILE) and registered in the repository's COMMON
|
||||
// info/exclude. git reads a worktree's excludes from the common dir, not the per-worktree
|
||||
@@ -233,8 +308,8 @@ class ClaudeCodeLauncherTest {
|
||||
Files.writeString(worktree.resolve(".git"), "gitdir: " + gitDir);
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
launcher(herdr, ideProfile(null, "http://127.0.0.1:29170/index-mcp/streamable-http",
|
||||
worktree.toString())).spawn();
|
||||
launcher(herdr, ideProfile(root.toString(), null,
|
||||
"http://127.0.0.1:29170/index-mcp/streamable-http", worktree.toString())).spawn();
|
||||
|
||||
Path overlay = worktree.resolve("CLAUDE.local.md");
|
||||
assertTrue(Files.exists(overlay), "the overlay is written beside the project's CLAUDE.md");
|
||||
@@ -260,8 +335,9 @@ class ClaudeCodeLauncherTest {
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
// ideProjectDir "fleetd" ⇒ the pin is <worktree>/fleetd, not <worktree>.
|
||||
launcher(herdr, ideProfileModule("http://127.0.0.1:29170/index-mcp/streamable-http",
|
||||
worktree.toString(), "fleetd", null)).spawn();
|
||||
launcher(herdr, ideProfileModule(root.toString(),
|
||||
"http://127.0.0.1:29170/index-mcp/streamable-http", worktree.toString(), "fleetd", null))
|
||||
.spawn();
|
||||
|
||||
Path overlay = worktree.resolve("CLAUDE.local.md");
|
||||
assertTrue(Files.exists(overlay), "the overlay file still lives at the worktree root");
|
||||
@@ -296,8 +372,8 @@ class ClaudeCodeLauncherTest {
|
||||
Files.createDirectories(worktree.resolve(".git"));
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
launcher(herdr, ideProfile(null, "http://127.0.0.1:29170/index-mcp/streamable-http",
|
||||
worktree.toString())).spawn();
|
||||
launcher(herdr, ideProfile(root.toString(), null,
|
||||
"http://127.0.0.1:29170/index-mcp/streamable-http", worktree.toString())).spawn();
|
||||
|
||||
assertFalse(Files.exists(worktree.resolve("CLAUDE.local.md")),
|
||||
"the safety gate refuses to write into a non-worktree cwd (.git directory)");
|
||||
@@ -583,6 +659,24 @@ class ClaudeCodeLauncherTest {
|
||||
assertNull(env.get("GITEA_HOST"), "no forge host without a granted token");
|
||||
}
|
||||
|
||||
@Test
|
||||
void noTokenEnvSpawnsWithoutInjectingAnAuthToken() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile cfg = new FleetConfig.Profile(
|
||||
"local-direct", "http://gx00.gw:8000", "coder", null, null,
|
||||
List.of("claude"), "tab", "fleetd-workers", "w #{n}", null, null, null);
|
||||
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), System::getenv);
|
||||
|
||||
assertDoesNotThrow(() -> {
|
||||
launcher.spawn();
|
||||
},
|
||||
"a profile that needs no token must still spawn against its configured endpoint");
|
||||
Map<String, String> env = startEnv(herdr);
|
||||
assertFalse(env.containsKey("ANTHROPIC_AUTH_TOKEN"),
|
||||
"an unconfigured tokenEnv must not inject an auth-token key");
|
||||
}
|
||||
|
||||
// --- CB-117 orphan reap: the pure predicate --------------------------------
|
||||
|
||||
@Test
|
||||
@@ -1383,14 +1477,16 @@ class ClaudeCodeLauncherTest {
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-633 follow-up (#192): the mirror of the zsh test above. On a non-zsh login shell {@code
|
||||
* ZDOTDIR} is ignored, so no scrub ever runs — the report must keep the WARN wording (a name here
|
||||
* really is inherited unblocked) rather than claiming a scrub protects it. This is the trap PR
|
||||
* #174 fell into the other direction: keying the wording on the shell, not on {@code
|
||||
* creds.isAllowList()}, is what keeps this branch correct.
|
||||
* fleetd #155 (was CB-633 follow-up (#192) "allowListPolicyOnNonZshKeepsTheWarnWording"): on a
|
||||
* non-zsh login shell {@code ZDOTDIR} is ignored, so no scrub could ever run there. Before #155
|
||||
* the launcher degraded to the weaker overlay and spawned anyway; that is exactly the "control
|
||||
* silently does nothing" failure #155 exists to close, since {@code policy: allow-list} is the
|
||||
* operator explicitly asking for a control a sourced file cannot undo. Now the spawn is REFUSED
|
||||
* instead — this test asserts the refusal, naming the shell, on the real spawn path ({@link
|
||||
* ClaudeCodeLauncher#spawn}), not merely on the launcher method that computes it.
|
||||
*/
|
||||
@Test
|
||||
void allowListPolicyOnNonZshKeepsTheWarnWording() {
|
||||
void allowListPolicyOnNonZshRefusesTheSpawn() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile cfg = new FleetConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
|
||||
@@ -1404,27 +1500,12 @@ class ClaudeCodeLauncherTest {
|
||||
0, System::currentTimeMillis, () -> {}, null, () -> creds,
|
||||
() -> Set.of("AI_GATEWAY_TOKEN", "A_BRAND_NEW_SECRET_TOKEN"));
|
||||
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
try {
|
||||
svc.spawn();
|
||||
} finally {
|
||||
logger.detachAppender(appender);
|
||||
}
|
||||
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, svc::spawn,
|
||||
"policy=allow-list on a non-zsh shell must refuse the spawn, not silently degrade "
|
||||
+ "to the weaker overlay");
|
||||
|
||||
assertTrue(appender.list.stream().anyMatch(e ->
|
||||
e.getFormattedMessage().contains("memberCredentials gap")
|
||||
&& e.getFormattedMessage().contains("UNBLOCKED")
|
||||
&& e.getFormattedMessage().contains("A_BRAND_NEW_SECRET_TOKEN")),
|
||||
"no scrub runs on a non-zsh shell, so the WARN wording must be kept — got: "
|
||||
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
|
||||
assertFalse(appender.list.stream().anyMatch(e ->
|
||||
e.getFormattedMessage().contains("memberCredentials gap")
|
||||
&& e.getFormattedMessage().toLowerCase(java.util.Locale.ROOT).contains("scrub")),
|
||||
"nothing is scrubbed on this path, so the report must not claim otherwise — got: "
|
||||
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
|
||||
assertTrue(e.getMessage().contains("allow-list") && e.getMessage().contains("/bin/bash"),
|
||||
"the refusal must name the policy and the actual shell, got: " + e.getMessage());
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -2080,4 +2161,437 @@ class ClaudeCodeLauncherTest {
|
||||
assertTrue(charterFile.getFileName().toString().startsWith("fleetd-role-charter-"),
|
||||
"same file-naming scheme as before this fix (no wrapping directory): " + charterFile);
|
||||
}
|
||||
|
||||
// --- fleetd #149: workspace-trust dialog seed ------------------------------------------------
|
||||
//
|
||||
// Claude Code asks an interactive, un-timed "Is this a project you created or one you trust?"
|
||||
// the first time it starts in a directory it has not seen. Every worktree: true spawn lands in
|
||||
// a brand-new directory, so without a seed the member sits on that dialog forever — herdr still
|
||||
// reports it healthy (agent_status: blocked, interactive_ready: true) — and never mounts the
|
||||
// bridge MCP or calls fleet_reply. These tests start the REAL launcher (only the herdr transport
|
||||
// is faked) so the seed is proven to run inside buildLaunch, before agent.start (the
|
||||
// process-starting call) ever fires — a test that only checked the JSON writer in isolation
|
||||
// would prove nothing about whether the launcher actually calls it at the right time.
|
||||
//
|
||||
// INCIDENT: the seed originally ran on ANY non-blank cwd. Running this file's own test suite —
|
||||
// most of whose fixtures spawn with no cwd/configDir set, so both fall back to the real
|
||||
// user.dir / ~/.claude.json — corrupted the operator's actual ~/.claude.json (it shrank from
|
||||
// ~72 KB to a single seeded entry) the first time a manual mutation run made the write
|
||||
// non-additive. The fix restricts the seed to isProvisionedWorktree(cwd) — a real .git FILE
|
||||
// (not directory) — exactly the writeIdeOverlay gate already used for the same category of
|
||||
// risk. Every test below marks its own worktree fixture with that .git file, and
|
||||
// seedTrustDialogNeverWritesWhenCwdIsNotAProvisionedWorktree is the regression test for the
|
||||
// incident itself.
|
||||
|
||||
private FleetConfig.Profile trustProfile(String configDir, String cwd) {
|
||||
return new FleetConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", configDir, "FLEETD_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "fleetd-workers", "w #{n}", "http://127.0.0.1:8765/mcp",
|
||||
cwd, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Give {@code dir} the exact signature {@link ClaudeCodeLauncher#isProvisionedWorktree} (and
|
||||
* {@code writeIdeOverlay} before it) checks for: a {@code .git} REGULAR FILE, never a
|
||||
* directory. The content is never parsed by the trust seed, so any {@code gitdir:} pointer is
|
||||
* fine.
|
||||
*/
|
||||
private static void markAsProvisionedWorktree(Path dir) throws IOException {
|
||||
Files.writeString(dir.resolve(".git"), "gitdir: /tmp/not-a-real-gitdir");
|
||||
}
|
||||
|
||||
@Test
|
||||
void seedsWorkspaceTrustForTheCwdBeforeTheProcessStarts(
|
||||
@TempDir Path configDir, @TempDir Path worktree) throws Exception {
|
||||
markAsProvisionedWorktree(worktree);
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path claudeJson = configDir.resolve(".claude.json");
|
||||
AtomicReference<Boolean> seededBeforeStart = new AtomicReference<>(false);
|
||||
herdr.onAgentStart(() -> {
|
||||
try {
|
||||
if (!Files.exists(claudeJson)) {
|
||||
return;
|
||||
}
|
||||
JsonNode root = new ObjectMapper().readTree(claudeJson.toFile());
|
||||
JsonNode project = root.path("projects").path(worktree.toString());
|
||||
seededBeforeStart.set(project.path("hasTrustDialogAccepted").asBoolean(false));
|
||||
} catch (IOException e) {
|
||||
seededBeforeStart.set(false);
|
||||
}
|
||||
});
|
||||
|
||||
FleetConfig.Profile cfg = trustProfile(configDir.toString(), worktree.toString());
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null)
|
||||
.spawn();
|
||||
|
||||
assertTrue(herdr.called("agent.start"), "the hook must actually have fired during the spawn");
|
||||
assertEquals(Boolean.TRUE, seededBeforeStart.get(),
|
||||
"the trust entry for the cwd must already exist at the instant agent.start (the "
|
||||
+ "process-starting herdr call) fires — not merely once spawn() returns");
|
||||
}
|
||||
|
||||
@Test
|
||||
void seedTrustDialogWritesOnlyTheTrustFlagAndNotTheOnboardingKey(
|
||||
@TempDir Path configDir, @TempDir Path worktree) throws Exception {
|
||||
markAsProvisionedWorktree(worktree);
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile cfg = trustProfile(configDir.toString(), worktree.toString());
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null)
|
||||
.spawn();
|
||||
|
||||
Path claudeJson = configDir.resolve(".claude.json");
|
||||
assertTrue(Files.exists(claudeJson), "seeded into <configDir>/.claude.json");
|
||||
JsonNode project = new ObjectMapper().readTree(claudeJson.toFile())
|
||||
.path("projects").path(worktree.toString());
|
||||
assertTrue(project.path("hasTrustDialogAccepted").asBoolean(false));
|
||||
// fleetd #247: the onboarding key must NOT be written. Claude Code strips it on every
|
||||
// save (measured: 0 of 28 live entries had it, including its own), so writing it only
|
||||
// adds a contested key to a file two processes share. This assertion is the guard that
|
||||
// stops it coming back as a plausible-looking "completeness" fix.
|
||||
assertFalse(project.has("hasCompletedProjectOnboarding"),
|
||||
"hasCompletedProjectOnboarding must not be written — Claude Code drops it on "
|
||||
+ "every save, and the member reaches idle on hasTrustDialogAccepted alone");
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 3: seeding is additive. An existing {@code .claude.json} carries the operator's own
|
||||
* project history and unrelated top-level settings — the seed must change only
|
||||
* {@code projects.<cwd>} for THIS cwd and leave everything else, including a different project's
|
||||
* own unrelated data, exactly as it was.
|
||||
*/
|
||||
@Test
|
||||
void seedTrustDialogIsAdditiveAndPreservesUnknownKeysAndOtherProjects(
|
||||
@TempDir Path configDir, @TempDir Path worktree) throws Exception {
|
||||
markAsProvisionedWorktree(worktree);
|
||||
Path claudeJson = configDir.resolve(".claude.json");
|
||||
Files.writeString(claudeJson, """
|
||||
{
|
||||
"numStartups": 42,
|
||||
"oauthAccount": {"emailAddress": "operator@example.com"},
|
||||
"projects": {
|
||||
"/some/other/project": {
|
||||
"hasTrustDialogAccepted": true,
|
||||
"mcpServers": {"foo": {"type": "stdio", "command": "foo-mcp"}}
|
||||
}
|
||||
}
|
||||
}
|
||||
""");
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile cfg = trustProfile(configDir.toString(), worktree.toString());
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null)
|
||||
.spawn();
|
||||
|
||||
JsonNode root = new ObjectMapper().readTree(claudeJson.toFile());
|
||||
assertEquals(42, root.path("numStartups").asInt(), "unrelated top-level key survives untouched");
|
||||
assertEquals("operator@example.com", root.path("oauthAccount").path("emailAddress").asText(),
|
||||
"an unrelated nested top-level key survives untouched");
|
||||
|
||||
JsonNode other = root.path("projects").path("/some/other/project");
|
||||
assertTrue(other.path("hasTrustDialogAccepted").asBoolean(false),
|
||||
"a different project's own trust entry survives");
|
||||
assertEquals("foo-mcp", other.path("mcpServers").path("foo").path("command").asText(),
|
||||
"a different project's own unrelated nested data survives");
|
||||
|
||||
JsonNode mine = root.path("projects").path(worktree.toString());
|
||||
assertTrue(mine.path("hasTrustDialogAccepted").asBoolean(false));
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 2: where the profile sets no {@code configDir}, the seed goes to the default
|
||||
* {@code ~/.claude.json}. {@code user.home} is redirected to a {@code @TempDir} for the
|
||||
* duration of this test and restored in a {@code finally} — the real operator {@code
|
||||
* ~/.claude.json} must never be touched by a test.
|
||||
*/
|
||||
@Test
|
||||
void seedTrustDialogTargetsDefaultClaudeJsonWhenConfigDirIsUnset(
|
||||
@TempDir Path fakeHome, @TempDir Path worktree) throws Exception {
|
||||
markAsProvisionedWorktree(worktree);
|
||||
String originalHome = System.getProperty("user.home");
|
||||
System.setProperty("user.home", fakeHome.toString());
|
||||
try {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile cfg = trustProfile(null, worktree.toString());
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null)
|
||||
.spawn();
|
||||
|
||||
Path claudeJson = fakeHome.resolve(".claude.json");
|
||||
assertTrue(Files.exists(claudeJson),
|
||||
"no configDir set — the default target is ~/.claude.json, here the redirected fake home");
|
||||
JsonNode project = new ObjectMapper().readTree(claudeJson.toFile())
|
||||
.path("projects").path(worktree.toString());
|
||||
assertTrue(project.path("hasTrustDialogAccepted").asBoolean(false));
|
||||
} finally {
|
||||
System.setProperty("user.home", originalHome);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Regression test for the fleetd #149 incident itself: a cwd that is NOT a fleetd-provisioned
|
||||
* worktree (no {@code .git} FILE — the exact shape a real checkout, or an un-configured
|
||||
* fallback cwd, has) must never be written to, however {@code configDir} is set. This is the
|
||||
* fix for the exact defect that corrupted the operator's real {@code ~/.claude.json}.
|
||||
*/
|
||||
@Test
|
||||
void seedTrustDialogNeverWritesWhenCwdIsNotAProvisionedWorktree(
|
||||
@TempDir Path configDir, @TempDir Path plainCwd) {
|
||||
// plainCwd deliberately carries NO .git file — the same shape a real checkout's cwd
|
||||
// fallback has.
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile cfg = trustProfile(configDir.toString(), plainCwd.toString());
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null)
|
||||
.spawn();
|
||||
|
||||
assertFalse(Files.exists(configDir.resolve(".claude.json")),
|
||||
"a non-worktree cwd must never get a .claude.json written for it — this is the fix "
|
||||
+ "for the incident where writing unconditionally corrupted the operator's own "
|
||||
+ "real ~/.claude.json via this file's own no-cwd/no-configDir test fixtures");
|
||||
}
|
||||
|
||||
/**
|
||||
* Same regression, for the default (no {@code configDir}) path — the exact combination (no
|
||||
* {@code configDir}, no worktree-shaped {@code cwd}) that hit the operator's real
|
||||
* {@code ~/.claude.json} during the incident. {@code user.home} is still redirected to a
|
||||
* {@code @TempDir} out of caution, so even a reintroduced bug here cannot touch the real file.
|
||||
*/
|
||||
@Test
|
||||
void seedTrustDialogNeverWritesToDefaultHomeWhenCwdIsNotAProvisionedWorktree(
|
||||
@TempDir Path fakeHome, @TempDir Path plainCwd) {
|
||||
String originalHome = System.getProperty("user.home");
|
||||
System.setProperty("user.home", fakeHome.toString());
|
||||
try {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile cfg = trustProfile(null, plainCwd.toString());
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null)
|
||||
.spawn();
|
||||
|
||||
assertFalse(Files.exists(fakeHome.resolve(".claude.json")),
|
||||
"the exact incident combination — no configDir, non-worktree cwd — must never "
|
||||
+ "write, even to the (redirected) default ~/.claude.json");
|
||||
} finally {
|
||||
System.setProperty("user.home", originalHome);
|
||||
}
|
||||
}
|
||||
|
||||
// --- fleetd #149 review round 2: the write must be atomic and lock-protected -----------------
|
||||
//
|
||||
// The first round fixed WHICH cwd this can ever target. This round fixes HOW the target is
|
||||
// written: .claude.json is large (tens of KB, dozens of projects on a real host), Claude Code
|
||||
// itself rewrites it while running, and this daemon spawns several members in parallel as a
|
||||
// matter of routine. The pre-fix Files.writeString(target, content) truncates target in place
|
||||
// before writing the replacement bytes — a crash mid-write, or another writer's read landing in
|
||||
// that window, loses data. Two more failure modes follow directly: (a) a crash/kill mid-write
|
||||
// leaves target truncated, and (b) two concurrent spawns racing a naive read-modify-write let
|
||||
// the second writer's write silently discard the first spawn's entry. The fix is
|
||||
// ClaudeCodeLauncher.writeAtomically (sibling temp file + ATOMIC_MOVE, package-visible for the
|
||||
// test below that proves it directly) plus TRUST_JSON_LOCK (a process-wide lock serialising
|
||||
// every seedTrustDialog call this launcher itself makes).
|
||||
|
||||
/**
|
||||
* Requested test 1: an existing, large-ish (not just a two-key fixture) .claude.json must never
|
||||
* collapse. Asserts on the restored KEY SET — not merely that the result still parses as JSON,
|
||||
* which the incident's 178-byte file also did.
|
||||
*/
|
||||
@Test
|
||||
void seedTrustDialogPreservesALargeExistingFileWithoutCollapsing(
|
||||
@TempDir Path configDir, @TempDir Path worktree) throws Exception {
|
||||
markAsProvisionedWorktree(worktree);
|
||||
Path claudeJson = configDir.resolve(".claude.json");
|
||||
|
||||
ObjectMapper mapper = new ObjectMapper();
|
||||
ObjectNode root = mapper.createObjectNode();
|
||||
root.put("numStartups", 4200);
|
||||
root.put("firstStartTime", "2025-01-01T00:00:00.000Z");
|
||||
root.putObject("oauthAccount").put("emailAddress", "operator@example.com");
|
||||
Set<String> otherProjectPaths = new HashSet<>();
|
||||
ObjectNode projects = root.putObject("projects");
|
||||
for (int i = 0; i < 30; i++) {
|
||||
String path = "/Users/operator/code/project-" + i;
|
||||
otherProjectPaths.add(path);
|
||||
ObjectNode project = projects.putObject(path);
|
||||
project.put("hasTrustDialogAccepted", true);
|
||||
project.putObject("mcpServers").putObject("server-" + i).put("command", "server-" + i + "-mcp");
|
||||
}
|
||||
String before = mapper.writerWithDefaultPrettyPrinter().writeValueAsString(root);
|
||||
Files.writeString(claudeJson, before);
|
||||
long sizeBefore = Files.size(claudeJson);
|
||||
assertTrue(sizeBefore > 4096,
|
||||
"fixture must actually be large-ish to be a meaningful proof: " + sizeBefore + " bytes");
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile cfg = trustProfile(configDir.toString(), worktree.toString());
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null)
|
||||
.spawn();
|
||||
|
||||
long sizeAfter = Files.size(claudeJson);
|
||||
assertTrue(sizeAfter >= sizeBefore,
|
||||
"the file must never collapse below its pre-seed size — before=" + sizeBefore
|
||||
+ " after=" + sizeAfter + " bytes (the incident shrank ~72 KB to 178 bytes)");
|
||||
|
||||
JsonNode after = mapper.readTree(claudeJson.toFile());
|
||||
assertEquals(4200, after.path("numStartups").asInt(), "unrelated top-level key survives");
|
||||
assertEquals("operator@example.com", after.path("oauthAccount").path("emailAddress").asText());
|
||||
JsonNode afterProjects = after.path("projects");
|
||||
for (String path : otherProjectPaths) {
|
||||
assertTrue(afterProjects.path(path).path("hasTrustDialogAccepted").asBoolean(false),
|
||||
"pre-existing project entry " + path + " must survive");
|
||||
}
|
||||
assertEquals(otherProjectPaths.size() + 1, afterProjects.size(),
|
||||
"exactly one NEW project entry (this worktree's) must be added, none dropped");
|
||||
assertTrue(afterProjects.path(worktree.toString()).path("hasTrustDialogAccepted").asBoolean(false));
|
||||
}
|
||||
|
||||
/**
|
||||
* Requested test 2: two concurrent spawns for DIFFERENT cwd values sharing one configDir must
|
||||
* both end up present in the final file — a naive concurrent read-modify-write would let the
|
||||
* second writer's read (taken before the first writer's write lands) silently discard the
|
||||
* first. A {@link CountDownLatch} — not a sleep — lines both threads up at the starting line so
|
||||
* this does not depend on scheduling luck to be meaningful.
|
||||
*
|
||||
* <p>This is a test of {@code TRUST_JSON_LOCK}, not of {@code writeAtomically}: the lock fully
|
||||
* serialises every {@code seedTrustDialog} call this launcher itself makes, so this test would
|
||||
* pass even without atomicity. See {@link #writeAtomicallyNeverExposesATornFileToAConcurrentReader}
|
||||
* for the test that exercises atomicity specifically.
|
||||
*/
|
||||
@Test
|
||||
void concurrentSeedsForDifferentCwdsBothSurvive(
|
||||
@TempDir Path configDir, @TempDir Path worktreeA, @TempDir Path worktreeB) throws Exception {
|
||||
markAsProvisionedWorktree(worktreeA);
|
||||
markAsProvisionedWorktree(worktreeB);
|
||||
|
||||
CountDownLatch ready = new CountDownLatch(2);
|
||||
CountDownLatch go = new CountDownLatch(1);
|
||||
AtomicReference<Exception> failureA = new AtomicReference<>();
|
||||
AtomicReference<Exception> failureB = new AtomicReference<>();
|
||||
|
||||
Thread ta = new Thread(spawnTask(configDir, worktreeA, ready, go, failureA), "spawn-a");
|
||||
Thread tb = new Thread(spawnTask(configDir, worktreeB, ready, go, failureB), "spawn-b");
|
||||
ta.start();
|
||||
tb.start();
|
||||
|
||||
assertTrue(ready.await(5, TimeUnit.SECONDS), "both threads must reach the starting line");
|
||||
go.countDown();
|
||||
ta.join(5000);
|
||||
tb.join(5000);
|
||||
assertFalse(ta.isAlive(), "spawn A must finish within the timeout");
|
||||
assertFalse(tb.isAlive(), "spawn B must finish within the timeout");
|
||||
assertNull(failureA.get(), "spawn A must not throw: " + failureA.get());
|
||||
assertNull(failureB.get(), "spawn B must not throw: " + failureB.get());
|
||||
|
||||
JsonNode root = new ObjectMapper().readTree(configDir.resolve(".claude.json").toFile());
|
||||
assertTrue(root.path("projects").path(worktreeA.toString())
|
||||
.path("hasTrustDialogAccepted").asBoolean(false),
|
||||
"worktree A's entry must survive the race");
|
||||
assertTrue(root.path("projects").path(worktreeB.toString())
|
||||
.path("hasTrustDialogAccepted").asBoolean(false),
|
||||
"worktree B's entry must survive the race");
|
||||
}
|
||||
|
||||
private Runnable spawnTask(Path configDir, Path worktree, CountDownLatch ready, CountDownLatch go,
|
||||
AtomicReference<Exception> failure) {
|
||||
return () -> {
|
||||
try {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile cfg = trustProfile(configDir.toString(), worktree.toString());
|
||||
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(new AgentControl(herdr),
|
||||
new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
|
||||
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
ready.countDown();
|
||||
go.await();
|
||||
launcher.spawn();
|
||||
} catch (Exception e) {
|
||||
failure.set(e);
|
||||
}
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Requested test 3: the atomic write must preserve {@code .claude.json}'s existing {@code 0600}
|
||||
* permissions, not silently widen them via a fresh temp file's own defaults landing on top of a
|
||||
* file that had different (e.g. group-readable) permissions. Skips rather than fails on a
|
||||
* filesystem with no POSIX permissions (e.g. Windows) — the same {@code assumeTrue} pattern this
|
||||
* file already uses for {@code memberHerdrSocketWithWorktreeRootAndGroupPutsCharterUnderWorktreeRootAndSharesIt}.
|
||||
*/
|
||||
@Test
|
||||
void seedTrustDialogPreservesExisting0600Permissions(
|
||||
@TempDir Path configDir, @TempDir Path worktree) throws Exception {
|
||||
markAsProvisionedWorktree(worktree);
|
||||
Path claudeJson = configDir.resolve(".claude.json");
|
||||
Files.writeString(claudeJson, "{}");
|
||||
assumeTrue(Files.getFileAttributeView(claudeJson, PosixFileAttributeView.class) != null,
|
||||
"no POSIX permissions on this filesystem — skipping rather than failing");
|
||||
Files.setPosixFilePermissions(claudeJson, PosixFilePermissions.fromString("rw-------"));
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile cfg = trustProfile(configDir.toString(), worktree.toString());
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null)
|
||||
.spawn();
|
||||
|
||||
assertEquals("rw-------", PosixFilePermissions.toString(Files.getPosixFilePermissions(claudeJson)),
|
||||
"the atomic write must preserve .claude.json's existing 0600 permissions, not widen "
|
||||
+ "them via a fresh temp file's own defaults");
|
||||
}
|
||||
|
||||
/**
|
||||
* Direct proof that {@link ClaudeCodeLauncher#writeAtomically} — not {@code TRUST_JSON_LOCK} —
|
||||
* is what keeps a concurrent reader of {@code target} from ever observing a truncated or
|
||||
* partially-written file. Calls {@code writeAtomically} directly (package-visible for exactly
|
||||
* this test) rather than going through {@code seedTrustDialog}/{@code spawn()}, because {@link
|
||||
* #concurrentSeedsForDifferentCwdsBothSurvive} above is protected by the lock and would pass
|
||||
* even without atomicity — it proves the lock, not the atomic move. This test proves the atomic
|
||||
* move specifically: a reader racing a writer OUTSIDE that lock (a second daemon process, or the
|
||||
* operator's own live Claude Code — exactly what the lock cannot reach) must still never see a
|
||||
* torn file.
|
||||
*
|
||||
* <p>The new content is made large (tens of MB) so a naive truncate-then-write has a real,
|
||||
* non-instantaneous window for the busy-poll reader thread to land in — this is inherently a
|
||||
* race, not a guaranteed-deterministic assertion, but it uses no {@code Thread.sleep} and
|
||||
* reliably reproduced the torn read when run against the pre-fix
|
||||
* {@code Files.writeString(target, content)} implementation (see the PR's mutation table).
|
||||
*/
|
||||
@Test
|
||||
void writeAtomicallyNeverExposesATornFileToAConcurrentReader(@TempDir Path dir) throws Exception {
|
||||
Path target = dir.resolve(".claude.json");
|
||||
String oldContent = "{\"marker\":\"OLD\"}";
|
||||
Files.writeString(target, oldContent);
|
||||
String newContent = "{\"marker\":\"NEW\",\"pad\":\"" + "x".repeat(20_000_000) + "\"}";
|
||||
|
||||
AtomicReference<String> tornRead = new AtomicReference<>();
|
||||
AtomicBoolean stop = new AtomicBoolean(false);
|
||||
Thread reader = new Thread(() -> {
|
||||
while (!stop.get()) {
|
||||
try {
|
||||
String seen = Files.readString(target);
|
||||
if (!seen.equals(oldContent) && !seen.equals(newContent)) {
|
||||
tornRead.compareAndSet(null, "torn read of length " + seen.length() + ": "
|
||||
+ seen.substring(0, Math.min(seen.length(), 80)));
|
||||
stop.set(true);
|
||||
}
|
||||
} catch (IOException ignored) {
|
||||
// ATOMIC_MOVE guarantees the path always resolves to old-or-new content once
|
||||
// readable at all — a transient "briefly missing during the rename" is fine to
|
||||
// ignore and keep sampling.
|
||||
}
|
||||
}
|
||||
});
|
||||
reader.start();
|
||||
try {
|
||||
ClaudeCodeLauncher.writeAtomically(target, newContent);
|
||||
} finally {
|
||||
stop.set(true);
|
||||
reader.join(5000);
|
||||
}
|
||||
|
||||
assertNull(tornRead.get(), "a concurrent reader must never observe a partially-written file: "
|
||||
+ tornRead.get());
|
||||
assertEquals(newContent, Files.readString(target), "the final content must be the new content");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -17,6 +17,7 @@ import dev.ltms.fleet.peer.PeerHandle;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.PeerUnreachableException;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementException;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
@@ -911,4 +912,185 @@ class CompositePeerLauncherTest {
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("claude", null, null)));
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("gemini", null, null)));
|
||||
}
|
||||
|
||||
// ── fleetd #201 Unit 5: repeated backend errors cool a credential off — a SEPARATE, shorter- ──
|
||||
// ── lived mechanism from CB-578 stage B quarantine above, never merged with it ───────────────
|
||||
|
||||
@Test
|
||||
void explicitSpawnOntoACoolingOffProfileIsRefusedNamingTheCredentialAndRemainingTime() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"sol", stubWorker("sol", "shared-openai"),
|
||||
"terra", stubWorker("terra", "shared-openai"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(() -> 0L);
|
||||
// Two DISTINCT targets on the same credential — evidenceCount() counts distinct targets,
|
||||
// never raw events, so one target repeating a classified error can never mint an incident.
|
||||
outagePolicy.record("shared-openai", "term-1", "API Error: 500");
|
||||
assertTrue(outagePolicy.record("shared-openai", "term-2", "API Error: 500").isPresent(),
|
||||
"the second distinct target crosses the threshold and starts the cool-off");
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), outagePolicy);
|
||||
|
||||
PlacementException e = assertThrows(PlacementException.class,
|
||||
() -> composite.spawn(new SpawnRequest("sol", null, null)));
|
||||
assertTrue(e.getMessage().contains("sol"), "message names the profile: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("shared-openai"), "message names the credential: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("cooling off"), "message says cooling off, not exhausted: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("60"), "message names roughly when it lifts: " + e.getMessage());
|
||||
assertEquals(0, adapter.spawnCount("sol"), "the cooling-off profile is never delegated to");
|
||||
}
|
||||
|
||||
/**
|
||||
* A single error from one target must never cool a credential off — the threshold is on
|
||||
* distinct targets. This proves the spawn gate really reads {@link BackendOutagePolicy}'s
|
||||
* result rather than reacting to any recorded event.
|
||||
*/
|
||||
@Test
|
||||
void explicitSpawnIsNotRefusedAfterOnlyOneBackendErrorOnOneTarget() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"sol", stubWorker("sol", "shared-openai"),
|
||||
"terra", stubWorker("terra", "shared-openai"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(() -> 0L);
|
||||
assertTrue(outagePolicy.record("shared-openai", "term-1", "API Error: 500").isEmpty(),
|
||||
"one distinct target is below THRESHOLD=2 — no incident yet");
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), outagePolicy);
|
||||
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("sol", null, null)),
|
||||
"one target's error never cools the credential off");
|
||||
assertEquals(1, adapter.spawnCount("sol"));
|
||||
}
|
||||
|
||||
/**
|
||||
* sol and terra share one OpenAI credential; cooling off because of an outage classified via
|
||||
* ONE of them must lock out the other too, exactly like CB-578 stage B quarantine does.
|
||||
*/
|
||||
@Test
|
||||
void twoProfilesSharingACredentialAreBothCoolingOffByOneOutageIncident() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"sol", stubWorker("sol", "shared-openai"),
|
||||
"terra", stubWorker("terra", "shared-openai"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(() -> 0L);
|
||||
outagePolicy.record("shared-openai", "term-1", "API Error: 500");
|
||||
outagePolicy.record("shared-openai", "term-2", "API Error: 500");
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), outagePolicy);
|
||||
|
||||
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("sol", null, null)));
|
||||
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("terra", null, null)),
|
||||
"terra shares sol's credential, so it must be locked out too");
|
||||
assertEquals(0, adapter.spawnCount("sol"));
|
||||
assertEquals(0, adapter.spawnCount("terra"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void placementSkipsACoolingOffProfileAndRoutesToAnotherOne() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"sol", stubWorker("sol", "shared-openai"),
|
||||
"b", stubWorker("b"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(() -> 0L);
|
||||
outagePolicy.record("shared-openai", "term-1", "API Error: 500");
|
||||
outagePolicy.record("shared-openai", "term-2", "API Error: 500");
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||
PlacementPolicies.weighted(), _ -> 0, null, BackendQuarantine.none(), outagePolicy);
|
||||
|
||||
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
|
||||
assertEquals("b", h.profile(), "sol is cooling off, so an unqualified spawn must land on b");
|
||||
assertEquals(0, adapter.spawnCount("sol"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void automaticPlacementNamesCoolingOffWhenEveryCandidateIsCoolingOff() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"sol", stubWorker("sol", "shared-openai"),
|
||||
"terra", stubWorker("terra", "shared-openai"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(() -> 0L);
|
||||
outagePolicy.record("shared-openai", "term-1", "API Error: 500");
|
||||
outagePolicy.record("shared-openai", "term-2", "API Error: 500");
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||
PlacementPolicies.weighted(), _ -> 0, null, BackendQuarantine.none(), outagePolicy);
|
||||
|
||||
// PlacementPolicy.select throws uncaught out of spawn() when available() is empty — never
|
||||
// wrapped in PeerUnreachableException, which is only for a delegate that actually refused.
|
||||
PlacementException e = assertThrows(PlacementException.class,
|
||||
() -> composite.spawn(new SpawnRequest(null, null, null)),
|
||||
"both candidates share the cooling-off credential — nothing is available");
|
||||
assertTrue(e.getMessage().contains("cooling off"),
|
||||
"message says cooling off, not exhausted, since nothing here is quarantined: " + e.getMessage());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aCoolOffLiftsOnTheInjectedClockAndTheProfileBecomesSpawnableAgain() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"sol", stubWorker("sol", "shared-openai"),
|
||||
"b", stubWorker("b"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
|
||||
AtomicLong nowNanos = new AtomicLong(0L);
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(nowNanos::get);
|
||||
outagePolicy.record("shared-openai", "term-1", "API Error: 500");
|
||||
outagePolicy.record("shared-openai", "term-2", "API Error: 500");
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), outagePolicy);
|
||||
|
||||
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("sol", null, null)),
|
||||
"still inside the 60s cool-off");
|
||||
|
||||
nowNanos.set(TimeUnit.SECONDS.toNanos(61));
|
||||
|
||||
PeerHandle h = composite.spawn(new SpawnRequest("sol", null, null));
|
||||
assertEquals("sol", h.profile(), "the cool-off expired on the injected clock — sol is spawnable again");
|
||||
assertEquals(1, adapter.spawnCount("sol"));
|
||||
}
|
||||
|
||||
/**
|
||||
* A credential that is BOTH exhaustion-quarantined (CB-578 stage B) and cooling off (fleetd
|
||||
* #201 Unit 5) at once — the two checks are independent and can both be true for one profile.
|
||||
* Exhaustion quarantine must win: the refusal names quarantine, never cooling off, because
|
||||
* {@code enforceNotQuarantined} is checked (and throws) before {@code enforceNotCoolingOff} ever
|
||||
* runs.
|
||||
*/
|
||||
@Test
|
||||
void quarantineWinsOverCoolingOffWhenBothAreActiveOnTheSameCredential() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"sol", stubWorker("sol", "shared-openai"),
|
||||
"terra", stubWorker("terra", "shared-openai"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
quarantine.quarantine("shared-openai");
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(() -> 0L);
|
||||
outagePolicy.record("shared-openai", "term-1", "API Error: 500");
|
||||
outagePolicy.record("shared-openai", "term-2", "API Error: 500");
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, quarantine, outagePolicy);
|
||||
|
||||
PlacementException e = assertThrows(PlacementException.class,
|
||||
() -> composite.spawn(new SpawnRequest("sol", null, null)));
|
||||
assertTrue(e.getMessage().contains("quarantined"),
|
||||
"exhaustion quarantine takes priority in the message: " + e.getMessage());
|
||||
assertFalse(e.getMessage().contains("cooling off"),
|
||||
"cooling off is never mentioned when quarantine already refused the spawn: " + e.getMessage());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aFleetWithNoBackendErrorsEverRecordedBehavesExactlyAsBeforeCoolOffExisted() {
|
||||
// A never-record()-called BackendOutagePolicy is naturally inert — the 2/5/6-arg
|
||||
// constructors used throughout this file all wire this in implicitly (NO_OUTAGE_POLICY).
|
||||
// This test pins that an explicit fresh instance also never refuses a spawn, for any profile.
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
PeerLauncher composite = composite(herdr);
|
||||
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("claude", null, null)));
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("gemini", null, null)));
|
||||
}
|
||||
}
|
||||
|
||||
+66
-38
@@ -29,6 +29,7 @@ import java.util.function.Supplier;
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
import static org.junit.jupiter.api.Assumptions.assumeTrue;
|
||||
|
||||
@@ -99,18 +100,26 @@ class HerdrPeerLauncherAllowListWiringTest {
|
||||
}
|
||||
|
||||
/**
|
||||
* A non-zsh shell cannot read {@code ZDOTDIR} at all. The launcher must fall back rather than
|
||||
* generate a directory nothing will ever read — a directory that would look like protection.
|
||||
* fleetd #155: a non-zsh shell cannot read {@code ZDOTDIR} at all, so under {@code
|
||||
* policy: allow-list} — the policy the operator picked specifically for a control a sourced file
|
||||
* cannot undo — the launcher must REFUSE the spawn rather than silently generate a directory
|
||||
* nothing will ever read (protection theatre) or fall back to the weaker overlay (exactly the
|
||||
* "control silently does nothing" defect this ticket exists to close). Real path: through {@link
|
||||
* HerdrPeerLauncher#spawn}, the method the daemon actually calls.
|
||||
*/
|
||||
@Test
|
||||
void aNonZshShellGeneratesNothingAndFallsBack() {
|
||||
void aNonZshShellUnderAllowListPolicyRefusesTheSpawn() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/bash");
|
||||
|
||||
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
|
||||
IllegalArgumentException e = assertThrows(IllegalArgumentException.class,
|
||||
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
|
||||
"policy=allow-list on a non-zsh shell must refuse the spawn, not silently degrade");
|
||||
|
||||
assertTrue(e.getMessage().contains("allow-list") && e.getMessage().contains("/bin/bash"),
|
||||
"the refusal must name the policy and the actual shell, got: " + e.getMessage());
|
||||
assertFalse(launcher.env.containsKey("ZDOTDIR"),
|
||||
"bash ignores ZDOTDIR; setting it would be protection theatre");
|
||||
"a refused spawn must not have generated (or wired in) a scrub directory: " + launcher.env);
|
||||
}
|
||||
|
||||
private static Supplier<FleetConfig.MemberCredentials> allowList() {
|
||||
@@ -246,11 +255,14 @@ class HerdrPeerLauncherAllowListWiringTest {
|
||||
* "allowed N of M" line — which describes what the scrub does — must not be printed there either.
|
||||
* Before this fix the line was logged BEFORE the zsh gate, so a non-zsh host printed e.g.
|
||||
* "allowed 1 of 3" while blocking nothing at all, telling an operator a control ran when it did
|
||||
* not. Real path: goes through {@link HerdrPeerLauncher#spawn}, same as the sibling test above,
|
||||
* with the shell fixed to bash so the fallback branch is the one exercised.
|
||||
* not. fleetd #155: the spawn itself is now refused on this path (see {@code
|
||||
* aNonZshShellUnderAllowListPolicyRefusesTheSpawn}) rather than falling back — this test's own
|
||||
* concern still holds under the refusal: the "allowed N of M" line describes a scrub that never
|
||||
* ran here, so it must still never appear. Real path: goes through {@link
|
||||
* HerdrPeerLauncher#spawn}, same as the sibling test above, with the shell fixed to bash.
|
||||
*/
|
||||
@Test
|
||||
void noAllowedCountLineIsEmittedOnTheNonZshFallbackPath() {
|
||||
void noAllowedCountLineIsEmittedOnTheNonZshRefusalPath() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Set<String> hostEnvNames = Set.of(INJECTED, "SOME_UNRELATED_NAME", "ANOTHER_UNRELATED_NAME");
|
||||
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/bash", () -> hostEnvNames);
|
||||
@@ -262,7 +274,9 @@ class HerdrPeerLauncherAllowListWiringTest {
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
try {
|
||||
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
|
||||
"policy=allow-list on a non-zsh shell must refuse the spawn");
|
||||
} finally {
|
||||
logger.detachAppender(appender);
|
||||
logger.setLevel(original);
|
||||
@@ -310,13 +324,20 @@ class HerdrPeerLauncherAllowListWiringTest {
|
||||
* different OS user — {@link HerdrPeerLauncher#hostEnvNames} describes fleetd's own process, not
|
||||
* that user's. Neither "inherits them UNBLOCKED" nor "scrub blanks them" is evidence-backed
|
||||
* there, so neither may print; the single unknown-environment WARN must, naming the config key.
|
||||
*
|
||||
* <p>fleetd #155: {@code memberLoginShell} is now given explicitly as zsh, so the spawn reaches
|
||||
* this WARN through the (still-degrading, not refusing) missing-{@code worktreeRoot}/{@code
|
||||
* worktreeGroup} fallback rather than through the zsh gate, which now refuses instead of falling
|
||||
* back — see {@code aNonZshShellUnderAllowListPolicyRefusesTheSpawn}. This test's own concern
|
||||
* (the unknown-environment WARN) is orthogonal to which fallback reached {@code
|
||||
* logCredentialGap}, so it still holds.
|
||||
*/
|
||||
@Test
|
||||
void gapDetectorReportsUnknownInsteadOfAConclusionWhenMemberHerdrSocketIsConfigured() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
|
||||
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
|
||||
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock"));
|
||||
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", "/bin/zsh"));
|
||||
|
||||
List<String> messages = spawnAndCaptureLogs(launcher);
|
||||
|
||||
@@ -340,8 +361,9 @@ class HerdrPeerLauncherAllowListWiringTest {
|
||||
void theUnknownEnvironmentWarnFiresOnceNotOncePerSpawn() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
|
||||
// fleetd #155: memberLoginShell given explicitly as zsh — see the sibling test above for why.
|
||||
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
|
||||
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock"));
|
||||
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", "/bin/zsh"));
|
||||
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
|
||||
Level original = logger.getLevel();
|
||||
@@ -389,42 +411,44 @@ class HerdrPeerLauncherAllowListWiringTest {
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #213 defect 1, acceptance criterion 1: {@code memberHerdrSocket} configured and {@code
|
||||
* memberLoginShell} configured as non-zsh must fall back to the sentinel overlay exactly like a
|
||||
* non-zsh {@code $SHELL} does today — and the "generated ZDOTDIR" INFO must not appear, since no
|
||||
* scrub actually runs. The WiringLauncher's own {@code env("SHELL")} is deliberately set to
|
||||
* {@code /bin/zsh} — the OPPOSITE of what {@code memberLoginShell} says — so a launcher that
|
||||
* (incorrectly) fell back to fleetd's own {@code $SHELL} here would wrongly pass the gate and
|
||||
* fail this test.
|
||||
* fleetd #213 defect 1, acceptance criterion 1 — updated by fleetd #155: {@code
|
||||
* memberHerdrSocket} configured and {@code memberLoginShell} configured as non-zsh must now
|
||||
* REFUSE the spawn (not fall back to the sentinel overlay — see {@code
|
||||
* aNonZshShellUnderAllowListPolicyRefusesTheSpawn}'s javadoc for why). The WiringLauncher's own
|
||||
* {@code env("SHELL")} is deliberately set to {@code /bin/zsh} — the OPPOSITE of what {@code
|
||||
* memberLoginShell} says — so a launcher that (incorrectly) fell back to fleetd's own {@code
|
||||
* $SHELL} here would wrongly pass the gate and fail this test.
|
||||
*/
|
||||
@Test
|
||||
void memberHerdrSocketWithNonZshMemberLoginShellFallsBackToTheOverlay() {
|
||||
void memberHerdrSocketWithNonZshMemberLoginShellRefusesTheSpawn() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
WiringLauncher launcher = new WiringLauncher(herdr, allowListWithKnown(List.of("SOME_TOKEN")),
|
||||
"/bin/zsh", null,
|
||||
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", "/bin/bash"));
|
||||
|
||||
List<String> messages = spawnAndCaptureLogs(launcher);
|
||||
IllegalArgumentException e = assertThrows(IllegalArgumentException.class,
|
||||
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
|
||||
"a non-zsh configured memberLoginShell under policy=allow-list must refuse the spawn");
|
||||
|
||||
assertEquals("blocked-by-fleetd-cb596-see-gitea-issue-82", launcher.env.get("SOME_TOKEN"),
|
||||
"a non-zsh memberLoginShell must fall back to the CB-596 sentinel overlay, exactly "
|
||||
+ "like a non-zsh $SHELL does when memberHerdrSocket is absent");
|
||||
assertTrue(e.getMessage().contains("/bin/bash"),
|
||||
"the refusal must name the configured memberLoginShell, got: " + e.getMessage());
|
||||
assertFalse(launcher.env.containsKey("SOME_TOKEN"),
|
||||
"a refused spawn must not have touched the pane env at all: " + launcher.env);
|
||||
assertFalse(launcher.env.containsKey("ZDOTDIR"),
|
||||
"no scrub directory may be generated when the configured member login shell is not zsh");
|
||||
assertFalse(messages.stream().anyMatch(m -> m.contains("generated ZDOTDIR")),
|
||||
"the 'generated ZDOTDIR' INFO must not appear when the scrub never runs — got: " + messages);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #213 defect 1, acceptance criterion 2: {@code memberHerdrSocket} configured and NO
|
||||
* {@code memberLoginShell} configured must fall back exactly like criterion 1 above — AND
|
||||
* fleetd's own {@code $SHELL} must never even be consulted (not merely "not decisive"). The
|
||||
* fixture's {@code env} function reports {@code /bin/zsh} for {@code SHELL} — a value that would
|
||||
* WRONGLY pass the zsh gate if the fix regressed to reading it — while flagging whether it was
|
||||
* ever asked for at all, so this test fails loudly on either kind of regression.
|
||||
* fleetd #213 defect 1, acceptance criterion 2 — updated by fleetd #155: {@code
|
||||
* memberHerdrSocket} configured and NO {@code memberLoginShell} configured must now REFUSE the
|
||||
* spawn exactly like criterion 1 above — AND fleetd's own {@code $SHELL} must never even be
|
||||
* consulted (not merely "not decisive"). The fixture's {@code env} function reports {@code
|
||||
* /bin/zsh} for {@code SHELL} — a value that would WRONGLY pass the zsh gate if the fix
|
||||
* regressed to reading it — while flagging whether it was ever asked for at all, so this test
|
||||
* fails loudly on either kind of regression.
|
||||
*/
|
||||
@Test
|
||||
void memberHerdrSocketWithNoMemberLoginShellFallsBackAndNeverConsultsFleetdsOwnShell() {
|
||||
void memberHerdrSocketWithNoMemberLoginShellRefusesAndNeverConsultsFleetdsOwnShell() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AtomicBoolean shellQueried = new AtomicBoolean(false);
|
||||
Function<String, String> env = name -> {
|
||||
@@ -437,15 +461,18 @@ class HerdrPeerLauncherAllowListWiringTest {
|
||||
WiringLauncher launcher = new WiringLauncher(herdr, allowListWithKnown(List.of("SOME_TOKEN")), env,
|
||||
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", null));
|
||||
|
||||
List<String> messages = spawnAndCaptureLogs(launcher);
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
|
||||
"no configured memberLoginShell under policy=allow-list must refuse the spawn, same "
|
||||
+ "as an explicit non-zsh one");
|
||||
|
||||
assertFalse(shellQueried.get(), "fleetd's own $SHELL must never be consulted once "
|
||||
+ "memberHerdrSocket is configured — only memberLoginShell: may decide the gate");
|
||||
assertEquals("blocked-by-fleetd-cb596-see-gitea-issue-82", launcher.env.get("SOME_TOKEN"),
|
||||
"no memberLoginShell configured must fall back to the sentinel overlay, same as a "
|
||||
assertFalse(launcher.env.containsKey("SOME_TOKEN"),
|
||||
"a refused spawn must not have touched the pane env at all, same as a "
|
||||
+ "configured non-zsh shell");
|
||||
assertFalse(messages.stream().anyMatch(m -> m.contains("generated ZDOTDIR")),
|
||||
"no scrub may run without a configured memberLoginShell — got: " + messages);
|
||||
assertFalse(launcher.env.containsKey("ZDOTDIR"),
|
||||
"no scrub may run without a configured memberLoginShell");
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -700,7 +727,8 @@ class HerdrPeerLauncherAllowListWiringTest {
|
||||
/**
|
||||
* fleetd #213: as above, plus {@code worktreeRoot:}/{@code worktreeGroup:} — both required for
|
||||
* the ZDOTDIR scrub to run at all once {@code memberHerdrSocket} is configured; either missing
|
||||
* falls back to the sentinel overlay, same as a non-zsh {@code memberLoginShell}.
|
||||
* falls back to the sentinel overlay (fleetd #155 left this branch alone — it is not a shell
|
||||
* problem, so it is not this ticket's refusal).
|
||||
*/
|
||||
private static FleetConfig configWithMemberHerdrSocketRootAndGroup(String memberHerdrSocket,
|
||||
String memberLoginShell, String worktreeRoot, String worktreeGroup) {
|
||||
|
||||
@@ -0,0 +1,99 @@
|
||||
package dev.ltms.fleet.member;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #111 (CB-608): {@link MemberCredentialPolicyView} is the one place that turns a {@code
|
||||
* memberCredentials:} policy into names-and-counts, so the startup log line and {@code
|
||||
* GET /member-credentials} cannot drift apart. These tests pin: the counts always match the
|
||||
* policy that produced them, an absent/empty policy is represented honestly (never as "nothing
|
||||
* blocked"), and the view carries names only — no value ever flows through it, because it is
|
||||
* built only from {@link FleetConfig.MemberCredentials}, which itself never holds a value.
|
||||
*/
|
||||
class MemberCredentialPolicyViewTest {
|
||||
|
||||
@Test
|
||||
void nullPolicyIsAbsentNotClean() {
|
||||
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(null);
|
||||
|
||||
assertFalse(view.present(), "a null policy must be reported as absent");
|
||||
assertEquals(0, view.knownCount());
|
||||
assertEquals(0, view.allowedCount());
|
||||
assertEquals(0, view.blockedCount());
|
||||
assertTrue(view.known().isEmpty());
|
||||
assertTrue(view.allowed().isEmpty());
|
||||
assertTrue(view.blocked().isEmpty());
|
||||
}
|
||||
|
||||
@Test
|
||||
void emptyKnownListIsAbsentEvenWithAPolicyModeSet() {
|
||||
// A memberCredentials: block can be present in YAML with policy: set but known: empty —
|
||||
// that must still read as "no policy configured", the same as a fully absent block,
|
||||
// because zero known names means the daemon blocks nothing either way.
|
||||
FleetConfig.MemberCredentials creds =
|
||||
new FleetConfig.MemberCredentials("deny-by-default", List.of(), List.of());
|
||||
|
||||
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(creds);
|
||||
|
||||
assertFalse(view.present());
|
||||
assertEquals(0, view.knownCount());
|
||||
}
|
||||
|
||||
@Test
|
||||
void countsMatchARealPolicyExactly() {
|
||||
FleetConfig.MemberCredentials creds = new FleetConfig.MemberCredentials(
|
||||
"deny-by-default",
|
||||
List.of("AI_GATEWAY_TOKEN"),
|
||||
List.of("AI_GATEWAY_TOKEN", "GITEA_ACCESS_TOKEN", "WORKER_GITEA_TOKEN"));
|
||||
|
||||
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(creds);
|
||||
|
||||
assertTrue(view.present());
|
||||
assertEquals("deny-by-default", view.policy());
|
||||
assertEquals(List.of("AI_GATEWAY_TOKEN", "GITEA_ACCESS_TOKEN", "WORKER_GITEA_TOKEN"), view.known());
|
||||
assertEquals(List.of("AI_GATEWAY_TOKEN"), view.allowed());
|
||||
assertEquals(3, view.knownCount());
|
||||
assertEquals(1, view.allowedCount());
|
||||
// known minus allowed — the two names actually shadowed on a spawn.
|
||||
assertEquals(2, view.blockedCount());
|
||||
assertTrue(view.blocked().containsAll(List.of("GITEA_ACCESS_TOKEN", "WORKER_GITEA_TOKEN")));
|
||||
}
|
||||
|
||||
@Test
|
||||
void presentPolicyThatBlocksNothingIsStillDistinctFromAbsent() {
|
||||
// known == allow => blockedCount is 0, exactly like an absent policy's blockedCount — the
|
||||
// two must still be told apart by `present`, or a reader cannot tell "policy configured,
|
||||
// nothing currently blocked" from "no policy at all".
|
||||
FleetConfig.MemberCredentials creds = new FleetConfig.MemberCredentials(
|
||||
"deny-by-default", List.of("X"), List.of("X"));
|
||||
|
||||
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(creds);
|
||||
|
||||
assertTrue(view.present());
|
||||
assertEquals(1, view.knownCount());
|
||||
assertEquals(0, view.blockedCount());
|
||||
assertFalse(MemberCredentialPolicyView.absent().present());
|
||||
}
|
||||
|
||||
@Test
|
||||
void namesPassThroughUnchangedNeverAValue() {
|
||||
// The view is built only from FleetConfig.MemberCredentials, which itself carries names,
|
||||
// never values (see its javadoc) — so there is no code path here that could substitute a
|
||||
// secret's value for its name. This pins the identity: what goes into `known`/`allow` is
|
||||
// exactly what comes out, character for character.
|
||||
List<String> known = List.of("SOME_TOKEN_NAME", "ANOTHER_NAME");
|
||||
FleetConfig.MemberCredentials creds =
|
||||
new FleetConfig.MemberCredentials("deny-by-default", List.of(), known);
|
||||
|
||||
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(creds);
|
||||
|
||||
assertEquals(known, view.known());
|
||||
}
|
||||
}
|
||||
@@ -17,6 +17,7 @@ import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.peer.PeerHandle;
|
||||
import dev.ltms.fleet.peer.PeerUnreachableException;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import org.junit.jupiter.api.Test;
|
||||
@@ -34,6 +35,7 @@ import java.util.Optional;
|
||||
import java.util.concurrent.ExecutorService;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.Future;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
@@ -52,6 +54,19 @@ class OpenCodeLauncherTest {
|
||||
null, null, gitTokenEnv, null, FleetConfig.Profile.KIND_OPENCODE);
|
||||
}
|
||||
|
||||
/**
|
||||
* A profile carrying an explicit {@code credentialId} (fleetd #234 defect 2) — distinct from
|
||||
* the profile's own name, so a test can assert on the credential precisely rather than relying
|
||||
* on {@code effectiveCredentialId()}'s profile-name fallback.
|
||||
*/
|
||||
private static FleetConfig.Profile opencodeCfgWithCredential(String profileName, String model,
|
||||
String credentialId) {
|
||||
return new FleetConfig.Profile(profileName, null, model, null, "FLEETD_WORKER_TOKEN",
|
||||
List.of("opencode"), "tab", "fleetd-workers", "opencode: {model} #{n}", null,
|
||||
null, List.of(), null, null, FleetConfig.Profile.KIND_OPENCODE, Map.of(), 1.0f,
|
||||
null, false, null, credentialId, null);
|
||||
}
|
||||
|
||||
/** Gate-disabled launcher whose per-spawn config dirs land under an inspectable temp root. */
|
||||
private static OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, FleetConfig.Profile cfg) {
|
||||
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
@@ -308,11 +323,39 @@ class OpenCodeLauncherTest {
|
||||
|
||||
// --- CB-547: resume + post-hoc session discovery --------------------------------------------
|
||||
|
||||
/**
|
||||
* Give {@code dir} the exact signature {@link HerdrPeerLauncher#isProvisionedWorktree} checks
|
||||
* for: a {@code .git} REGULAR FILE, never a directory. Content is never parsed by that gate, so
|
||||
* any {@code gitdir:} pointer is fine. Mirrors {@code ClaudeCodeLauncherTest}'s helper of the
|
||||
* same shape (fleetd #249).
|
||||
*/
|
||||
private static void markAsProvisionedWorktree(Path dir) throws IOException {
|
||||
Files.writeString(dir.resolve(".git"), "gitdir: /tmp/not-a-real-gitdir");
|
||||
}
|
||||
|
||||
/**
|
||||
* A fresh subdirectory of {@code configRoot}, marked as a provisioned worktree (fleetd #249),
|
||||
* for tests that predate this gate and stood in a bare {@code "/work/dir"} string as their
|
||||
* member's cwd — a directory that never existed on disk and, post-#249, would never pass
|
||||
* {@link HerdrPeerLauncher#isProvisionedWorktree} either. Those tests are about the model
|
||||
* mismatch / late-resolve machinery (fleetd #175/#234/#209), not about the worktree gate
|
||||
* itself, so they need a cwd the gate accepts without changing what each test demonstrates.
|
||||
*/
|
||||
private static String provisionedWorkDir(Path configRoot) throws IOException {
|
||||
Path dir = Files.createDirectories(configRoot.resolve("work-dir"));
|
||||
markAsProvisionedWorktree(dir);
|
||||
return dir.toString();
|
||||
}
|
||||
|
||||
@Test
|
||||
void aResumeSpawnPassesTheSessionIdAsDashS(@TempDir Path root) {
|
||||
void aResumeSpawnIntoAProvisionedWorktreePassesTheSessionIdAsDashS(@TempDir Path root,
|
||||
@TempDir Path worktree)
|
||||
throws Exception {
|
||||
markAsProvisionedWorktree(worktree);
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", null, null))
|
||||
.spawn(new SpawnRequest(null, null, null, null, "ses_41b79fc90ffeI9E8uZv6VprUn2"));
|
||||
.spawn(new SpawnRequest(null, worktree.toString(), null, null,
|
||||
"ses_41b79fc90ffeI9E8uZv6VprUn2"));
|
||||
|
||||
List<String> args = startArgs(herdr);
|
||||
int s = args.indexOf("-s");
|
||||
@@ -321,6 +364,27 @@ class OpenCodeLauncherTest {
|
||||
"the resume target id follows -s");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #249 acceptance criterion 3: without a fleetd-provisioned worktree, the member's cwd
|
||||
* is shared with other sessions, so fleetd can never reliably confirm (now or later via {@link
|
||||
* OpenCodeSessionDiscovery}) which conversation it is actually running. Refuse the spawn itself
|
||||
* rather than silently launching opencode's {@code -s <id>} into unverifiable territory.
|
||||
*/
|
||||
@Test
|
||||
void aResumeSpawnWithoutAProvisionedWorktreeIsRefused(@TempDir Path root) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
OpenCodeLauncher launcher = service(herdr, root,
|
||||
opencodeCfg("google/gemini-2.5-pro", null, null));
|
||||
|
||||
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, () ->
|
||||
launcher.spawn(new SpawnRequest(null, null, null, null,
|
||||
"ses_41b79fc90ffeI9E8uZv6VprUn2")));
|
||||
|
||||
assertTrue(e.getMessage().contains("worktree"), e.getMessage());
|
||||
assertFalse(herdr.called("agent.start"),
|
||||
"the refusal must happen before anything spawns — no pane, no process");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aFreshSpawnCarriesNoSessionFlag(@TempDir Path root) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
@@ -333,25 +397,59 @@ class OpenCodeLauncherTest {
|
||||
|
||||
@Test
|
||||
void theHandleDiscoversTheSessionIdForTheWorkersCwdOnlyAfterItAppears(@TempDir Path root,
|
||||
@TempDir Path discRoot)
|
||||
@TempDir Path discRoot,
|
||||
@TempDir Path worktree)
|
||||
throws Exception {
|
||||
markAsProvisionedWorktree(worktree);
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
OpenCodeLauncher launcher = new OpenCodeLauncher(new AgentControl(herdr),
|
||||
new WorkspaceControl(herdr), Map.of("gemini", opencodeCfg(null, null, null)),
|
||||
"gemini", _ -> null, 0, System::currentTimeMillis, () -> { }, root, discRoot);
|
||||
|
||||
PeerHandle handle = launcher.spawn(new SpawnRequest(null, "/work/dir", null));
|
||||
PeerHandle handle = launcher.spawn(new SpawnRequest(null, worktree.toString(), null));
|
||||
|
||||
// opencode writes the record only when the session is first persisted — the instant the
|
||||
// pane is ready it does not exist, so agentSessionId() is null (never a spawn failure).
|
||||
assertNull(handle.agentSessionId(), "no record yet → null, not a spawn-time block");
|
||||
// Once the record appears (here: same cwd), lazy discovery resolves it — the handle's
|
||||
// session id matches its own worktree, not another's.
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_resolved", "/work/dir", 1000L);
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_resolved", worktree.toString(), 1000L);
|
||||
assertEquals("ses_resolved", handle.agentSessionId(),
|
||||
"agentSessionId() re-scans and picks up a record that has since been written");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #249 acceptance criterion 1, exercised through the real caller path (the handle
|
||||
* {@code fleet_list} actually reads), not {@link OpenCodeSessionDiscovery} directly. Without a
|
||||
* fleetd-provisioned worktree the member's cwd is shared — the default no-worktree spawn
|
||||
* inherits the lead's own long-lived cwd — so even once a matching row appears (here:
|
||||
* simulating another profile's session that happens to share the directory) the handle must
|
||||
* report absence rather than guess. Measured real-world case (2026-09-03): the row it would
|
||||
* otherwise pick was three days old and belonged to a different profile.
|
||||
*/
|
||||
@Test
|
||||
void theHandleNeverReportsAnIdForANonProvisionedCwdEvenAfterARowAppears(@TempDir Path root,
|
||||
@TempDir Path discRoot)
|
||||
throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
OpenCodeLauncher launcher = new OpenCodeLauncher(new AgentControl(herdr),
|
||||
new WorkspaceControl(herdr), Map.of("gemini", opencodeCfg(null, null, null)),
|
||||
"gemini", _ -> null, 0, System::currentTimeMillis, () -> { }, root, discRoot);
|
||||
// No markAsProvisionedWorktree — this cwd has no .git file, the shared-cwd shape a
|
||||
// no-worktree spawn (or a real checkout) actually has.
|
||||
String sharedCwd = root.resolve("shared-cwd").toString();
|
||||
|
||||
PeerHandle handle = launcher.spawn(new SpawnRequest(null, sharedCwd, null));
|
||||
|
||||
assertNull(handle.agentSessionId(), "no record yet → null, same as the provisioned case");
|
||||
// A row for this exact directory now appears — e.g. a sibling member, or a stale session
|
||||
// from days earlier, sharing the same unprovisioned cwd.
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_someone_elses", sharedCwd, 1000L);
|
||||
assertNull(handle.agentSessionId(),
|
||||
"a non-provisioned cwd must NEVER report an id, even once a row for it exists — "
|
||||
+ "the row could belong to any other session sharing this directory");
|
||||
}
|
||||
|
||||
@Test
|
||||
void foreignWorkerMatchesOpencodePrefixButNotClaude() {
|
||||
String nonce = "abc123";
|
||||
@@ -905,16 +1003,17 @@ class OpenCodeLauncherTest {
|
||||
@Test
|
||||
void theRealSessionManagerLateResolvePathCatchesAModelMismatch(@TempDir Path configRoot,
|
||||
@TempDir Path discRoot) throws Exception {
|
||||
String workDir = provisionedWorkDir(configRoot);
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
// xf's real shape (fleetd #175): weight:80, model "opencode/nemotron-3-ultra-free", no
|
||||
// credentialId — the profile that actually escaped the fleet's accounting.
|
||||
FleetConfig.Profile cfg = opencodeCfg("opencode/nemotron-3-ultra-free", null, null);
|
||||
List<String> exhausted = new ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason) -> exhausted.add(target + "|" + reason);
|
||||
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(target + "|" + reason);
|
||||
OpenCodeLauncher launcher = serviceWithSink(herdr, configRoot, discRoot, cfg, sink);
|
||||
|
||||
SessionManager sessions = new SessionManager(launcher);
|
||||
MemberSession acquired = sessions.acquire(cfg.profile(), "/work/dir", null, null);
|
||||
MemberSession acquired = sessions.acquire(cfg.profile(), workDir, null, null);
|
||||
|
||||
// Real late-resolve path, driven BEFORE opencode has written its session row — same shape
|
||||
// as production the instant a pane goes ready.
|
||||
@@ -925,7 +1024,7 @@ class OpenCodeLauncherTest {
|
||||
|
||||
// opencode writes its row late, running gpt-5.6-sol (a PAID credential) instead of the
|
||||
// withdrawn free model the profile actually asked for — the exact fleetd #175 scenario.
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
|
||||
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
|
||||
|
||||
// Drive the SAME real late-resolve path again: sessions.get() -> resolveAgentSessionId ->
|
||||
@@ -944,12 +1043,13 @@ class OpenCodeLauncherTest {
|
||||
@Test
|
||||
void aProviderPrefixedModelMatchingBothIdAndProviderIsNotAMismatch(@TempDir Path configRoot,
|
||||
@TempDir Path discRoot) throws Exception {
|
||||
String workDir = provisionedWorkDir(configRoot);
|
||||
List<String> exhausted = new ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason) -> exhausted.add(reason);
|
||||
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
|
||||
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
|
||||
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
|
||||
.spawn(new SpawnRequest(null, "/work/dir", null));
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
|
||||
.spawn(new SpawnRequest(null, workDir, null));
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
|
||||
"{\"id\":\"gpt-5.6-terra\",\"providerID\":\"openai\"}");
|
||||
|
||||
assertEquals("ses_x", handle.agentSessionId());
|
||||
@@ -960,12 +1060,13 @@ class OpenCodeLauncherTest {
|
||||
@Test
|
||||
void aGxProviderPrefixedModelMatchingBothIdAndProviderIsNotAMismatch(@TempDir Path configRoot,
|
||||
@TempDir Path discRoot) throws Exception {
|
||||
String workDir = provisionedWorkDir(configRoot);
|
||||
List<String> exhausted = new ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason) -> exhausted.add(reason);
|
||||
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
|
||||
FleetConfig.Profile cfg = opencodeCfg("gx/deepseek-v4-flash", null, null);
|
||||
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
|
||||
.spawn(new SpawnRequest(null, "/work/dir", null));
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
|
||||
.spawn(new SpawnRequest(null, workDir, null));
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
|
||||
"{\"id\":\"deepseek-v4-flash\",\"providerID\":\"gx\"}");
|
||||
|
||||
assertEquals("ses_x", handle.agentSessionId());
|
||||
@@ -982,12 +1083,13 @@ class OpenCodeLauncherTest {
|
||||
@Test
|
||||
void aMissingProviderIdInTheEvidenceIsUnknownNotAMismatchWhenTheIdMatches(
|
||||
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
|
||||
String workDir = provisionedWorkDir(configRoot);
|
||||
List<String> exhausted = new ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason) -> exhausted.add(reason);
|
||||
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
|
||||
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
|
||||
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
|
||||
.spawn(new SpawnRequest(null, "/work/dir", null));
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
|
||||
.spawn(new SpawnRequest(null, workDir, null));
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
|
||||
"{\"id\":\"gpt-5.6-terra\"}");
|
||||
|
||||
assertEquals("ses_x", handle.agentSessionId());
|
||||
@@ -1003,12 +1105,13 @@ class OpenCodeLauncherTest {
|
||||
@Test
|
||||
void aMissingProviderIdInTheEvidenceStillCatchesARealIdMismatch(
|
||||
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
|
||||
String workDir = provisionedWorkDir(configRoot);
|
||||
List<String> exhausted = new ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason) -> exhausted.add(reason);
|
||||
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
|
||||
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
|
||||
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
|
||||
.spawn(new SpawnRequest(null, "/work/dir", null));
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
|
||||
.spawn(new SpawnRequest(null, workDir, null));
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
|
||||
"{\"id\":\"gpt-5.6-sol\"}");
|
||||
|
||||
assertEquals("ses_x", handle.agentSessionId());
|
||||
@@ -1027,12 +1130,13 @@ class OpenCodeLauncherTest {
|
||||
@Test
|
||||
void aBareModelWithNoProviderPrefixMatchesOnIdAloneAndIsNotAMismatch(@TempDir Path configRoot,
|
||||
@TempDir Path discRoot) throws Exception {
|
||||
String workDir = provisionedWorkDir(configRoot);
|
||||
List<String> exhausted = new ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason) -> exhausted.add(reason);
|
||||
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
|
||||
FleetConfig.Profile cfg = opencodeCfg("deepseek-v4-flash", null, null);
|
||||
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
|
||||
.spawn(new SpawnRequest(null, "/work/dir", null));
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
|
||||
.spawn(new SpawnRequest(null, workDir, null));
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
|
||||
"{\"id\":\"deepseek-v4-flash\",\"providerID\":\"gx\"}");
|
||||
|
||||
assertEquals("ses_x", handle.agentSessionId());
|
||||
@@ -1049,8 +1153,9 @@ class OpenCodeLauncherTest {
|
||||
@Test
|
||||
void aRealIdMismatchLogsAnErrorNamingBothModelsAndQuarantinesThroughTheSink(
|
||||
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
|
||||
String workDir = provisionedWorkDir(configRoot);
|
||||
List<String> exhausted = new ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason) -> exhausted.add(target + "|" + reason);
|
||||
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(target + "|" + reason);
|
||||
FleetConfig.Profile cfg = opencodeCfg("opencode/nemotron-3-ultra-free", null, null);
|
||||
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(OpenCodeLauncher.class);
|
||||
@@ -1060,8 +1165,8 @@ class OpenCodeLauncherTest {
|
||||
PeerHandle handle;
|
||||
try {
|
||||
handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
|
||||
.spawn(new SpawnRequest(null, "/work/dir", null));
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
|
||||
.spawn(new SpawnRequest(null, workDir, null));
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
|
||||
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
|
||||
assertEquals("ses_x", handle.agentSessionId());
|
||||
} finally {
|
||||
@@ -1096,11 +1201,12 @@ class OpenCodeLauncherTest {
|
||||
@Test
|
||||
void unknownOrUnparseableModelEvidenceNeverQuarantines(@TempDir Path configRoot,
|
||||
@TempDir Path discRoot) throws Exception {
|
||||
String workDir = provisionedWorkDir(configRoot);
|
||||
List<String> exhausted = new ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason) -> exhausted.add(reason);
|
||||
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
|
||||
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
|
||||
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
|
||||
.spawn(new SpawnRequest(null, "/work/dir", null));
|
||||
.spawn(new SpawnRequest(null, workDir, null));
|
||||
|
||||
// No row yet at all.
|
||||
assertNull(handle.agentSessionId());
|
||||
@@ -1122,15 +1228,216 @@ class OpenCodeLauncherTest {
|
||||
@Test
|
||||
void aProfileWithNoConfiguredModelIsNeverCheckedForAMismatch(@TempDir Path configRoot,
|
||||
@TempDir Path discRoot) throws Exception {
|
||||
String workDir = provisionedWorkDir(configRoot);
|
||||
List<String> exhausted = new ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason) -> exhausted.add(reason);
|
||||
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
|
||||
FleetConfig.Profile cfg = opencodeCfg(null, null, null);
|
||||
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
|
||||
.spawn(new SpawnRequest(null, "/work/dir", null));
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
|
||||
.spawn(new SpawnRequest(null, workDir, null));
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
|
||||
"{\"id\":\"anything-at-all\",\"providerID\":\"anyone\"}");
|
||||
|
||||
assertEquals("ses_x", handle.agentSessionId());
|
||||
assertTrue(exhausted.isEmpty(), "no model configured → nothing to compare: " + exhausted);
|
||||
}
|
||||
|
||||
// --- fleetd #234 defect 1: key the model check on the RESOLVED id, not the shared directory --
|
||||
|
||||
/**
|
||||
* The exact shape fleetd #234 reported: {@code fleet_spawn} with no {@code worktree:} shares
|
||||
* the lead's cwd across every worker, so more than one session row can exist for the SAME
|
||||
* {@code directory}. Once THIS handle's own session id is resolved, a sibling member spawned
|
||||
* later into the same shared directory — writing a NEWER, unrelated row — must never make the
|
||||
* already-resolved session look mismatched. Today's code re-derives "the newest row in this
|
||||
* directory" on every call (both for the id AND, independently, for the model), so it would
|
||||
* pick up the sibling's row on the second call and flag a false mismatch AND flip the returned
|
||||
* id. The fix (fleetd #234) makes the id sticky once resolved and reads the model back for
|
||||
* exactly that id (see {@link OpenCodeSessionDiscovery#actualModelForSessionId}) — never
|
||||
* "whatever is newest in the directory right now."
|
||||
*/
|
||||
@Test
|
||||
void modelCheckReadsTheResolvedSessionsOwnRowNotWhateverIsNewestInTheSharedDirectory(
|
||||
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
|
||||
String workDir = provisionedWorkDir(configRoot);
|
||||
List<String> exhausted = new ArrayList<>();
|
||||
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
|
||||
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
|
||||
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
|
||||
.spawn(new SpawnRequest(null, workDir, null));
|
||||
|
||||
// Our own session's row, correctly matching the profile's requested model.
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_ours", workDir, 1000L,
|
||||
"{\"id\":\"gpt-5.6-terra\",\"providerID\":\"openai\"}");
|
||||
assertEquals("ses_ours", handle.agentSessionId(), "resolves to our own session");
|
||||
assertTrue(exhausted.isEmpty(), "matching model → no mismatch on first resolve: " + exhausted);
|
||||
|
||||
// A sibling member, spawned later into the SAME shared directory (no worktree, fleetd
|
||||
// #234's default), writes a newer row running a totally different model.
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_sibling", workDir, 9000L,
|
||||
"{\"id\":\"deepseek-v4-flash\",\"providerID\":\"gx\"}");
|
||||
|
||||
assertEquals("ses_ours", handle.agentSessionId(),
|
||||
"the session id, once resolved, must not flip to a sibling sharing the directory");
|
||||
assertTrue(exhausted.isEmpty(),
|
||||
"a sibling's later, unrelated row in the same shared directory must never be read "
|
||||
+ "as OUR session's model: " + exhausted);
|
||||
}
|
||||
|
||||
// --- fleetd #234 defect 2: the quarantine the ERROR announces must actually happen -----------
|
||||
|
||||
/**
|
||||
* fleetd #234: {@link OpenCodeLauncher.SessionAwareHandle#checkModelMatch} fires from {@code
|
||||
* agentSessionId()}, which {@code SessionManager.acquire()} calls to build the very first
|
||||
* {@code MemberSession} record — BEFORE that session is put into the registry {@code
|
||||
* sessions.roster()} reads. A sink that resolves {@code target -> profile} ONLY through the
|
||||
* roster (today's {@code Fleetd.java} code, before this fix) therefore finds nothing at this
|
||||
* exact moment and silently does not quarantine, even though it just logged an ERROR saying it
|
||||
* would. This test drives the REAL path — {@code SessionManager.acquire()} — not the sink
|
||||
* directly, because the bug is entirely about this ordering; a direct-sink test cannot see it
|
||||
* (and is exactly why #175's own test suite, which only ever called the sink directly or after
|
||||
* registration, never caught this).
|
||||
*
|
||||
* <p>The sink under test mirrors {@code Fleetd.main()}'s real wiring after the fix: resolve
|
||||
* via the roster first (unchanged for {@code CompletionResolver}'s two call sites), falling
|
||||
* back to the profile hint {@link OpenCodeLauncher} now supplies via {@link
|
||||
* ExhaustionSink#onExhausted(String, String, String)} when the roster lookup misses.
|
||||
*/
|
||||
@Test
|
||||
void aSpawnTimeModelMismatchActuallyQuarantinesTheCredentialThroughTheRealAcquirePath(
|
||||
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
|
||||
String workDir = provisionedWorkDir(configRoot);
|
||||
FleetConfig.Profile cfg = opencodeCfgWithCredential(
|
||||
"terra", "opencode/nemotron-3-ultra-free", "openai-shared");
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of(cfg.profile(), cfg);
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.SECONDS.toNanos(1800));
|
||||
|
||||
// fleetd #234, round 4: safely a lambda now — the 3-arg overload is the interface's single
|
||||
// abstract method, so there is no 2-arg overload left to bind to instead.
|
||||
ExhaustionSink sink = (target, reason, profileHint) -> {
|
||||
FleetConfig.Profile profile = profileHint == null ? null : profiles.get(profileHint);
|
||||
if (profile != null) {
|
||||
quarantine.quarantine(profile.effectiveCredentialId());
|
||||
}
|
||||
};
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
OpenCodeLauncher launcher = serviceWithSink(herdr, configRoot, discRoot, cfg, sink);
|
||||
SessionManager sessions = new SessionManager(launcher);
|
||||
|
||||
// The mismatching row exists BEFORE the spawn — reproducing fleetd #234's exact timing:
|
||||
// opencode's session table already carries evidence by the moment acquire() first asks.
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
|
||||
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
|
||||
|
||||
assertFalse(quarantine.isQuarantined("openai-shared"), "nothing quarantined before the spawn");
|
||||
|
||||
// The real production entrypoint: acquire() builds the MemberSession by calling
|
||||
// handle.agentSessionId() BEFORE registry.put() runs.
|
||||
MemberSession acquired = sessions.acquire(cfg.profile(), workDir, null, null);
|
||||
|
||||
assertEquals("ses_x", acquired.agentSessionId(), "the id itself still resolves correctly");
|
||||
assertTrue(quarantine.isQuarantined("openai-shared"),
|
||||
"the mismatch fires DURING acquire(), before roster registration, and must still "
|
||||
+ "reach the quarantine via the profile hint — not silently no-op");
|
||||
}
|
||||
|
||||
/**
|
||||
* The other half of the same proof: a sink that resolves {@code target -> profile} ONLY
|
||||
* through the roster (i.e. ignores the profile hint entirely, ~today's pre-fix {@code
|
||||
* Fleetd.java}) drops the SAME spawn-time mismatch silently — the credential is never
|
||||
* quarantined even though {@link OpenCodeLauncher} logged the mismatch ERROR. This is the
|
||||
* failure fleetd #234 reported, reproduced through the real {@code SessionManager.acquire()}
|
||||
* path rather than asserted by inspecting the fix.
|
||||
*/
|
||||
@Test
|
||||
void aRosterOnlySinkSilentlyDropsTheSpawnTimeQuarantine(@TempDir Path configRoot,
|
||||
@TempDir Path discRoot) throws Exception {
|
||||
String workDir = provisionedWorkDir(configRoot);
|
||||
FleetConfig.Profile cfg = opencodeCfgWithCredential(
|
||||
"terra", "opencode/nemotron-3-ultra-free", "openai-shared");
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.SECONDS.toNanos(1800));
|
||||
|
||||
// Deliberately ignores the profile hint — the pre-fix shape: only a roster lookup (modelled
|
||||
// here as always empty, since acquire() has not registered the session yet either way).
|
||||
ExhaustionSink rosterOnlySink = (target, reason, profile) -> { };
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
OpenCodeLauncher launcher = serviceWithSink(herdr, configRoot, discRoot, cfg, rosterOnlySink);
|
||||
SessionManager sessions = new SessionManager(launcher);
|
||||
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
|
||||
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
|
||||
|
||||
MemberSession acquired = sessions.acquire(cfg.profile(), workDir, null, null);
|
||||
|
||||
assertEquals("ses_x", acquired.agentSessionId(), "the id itself still resolves correctly");
|
||||
assertFalse(quarantine.isQuarantined("openai-shared"),
|
||||
"a roster-only sink cannot see this target yet — the quarantine silently never "
|
||||
+ "happens, which is exactly fleetd #234 defect 2");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #234, round 2: the two tests above inject a sink DIRECTLY into {@link OpenCodeLauncher},
|
||||
* which is not what {@code Fleetd.main} actually does. Production has one more hop: {@code
|
||||
* Fleetd.java} builds the adapters (including {@code OpenCodeLauncher}) before {@code sessions}
|
||||
* exists — a genuine construction-order cycle — so it hands the launcher a <em>forwarding</em>
|
||||
* sink pointed at an {@code AtomicReference<ExhaustionSink>}, and only later {@code .set(...)}s
|
||||
* that reference to the real sink once {@code sessions} is built.
|
||||
*
|
||||
* <p><strong>Round 3:</strong> a round-2 version of this test built its own copy of that
|
||||
* forwarder's shape as an anonymous class. Mutating {@code Fleetd.java}'s ACTUAL forwarder back
|
||||
* into a broken lambda left this test's own copy untouched, so it kept passing while production
|
||||
* had regressed to the exact bug being fixed. This version instead calls {@link
|
||||
* ExhaustionSink#forwardingTo}, the same factory {@code Fleetd.java} calls — the identical
|
||||
* object, not a rebuilt copy of its shape — so a regression at either the {@code Fleetd.java}
|
||||
* call site or inside {@code forwardingTo} itself has nowhere left to hide. See {@code
|
||||
* ExhaustionSinkForwardingHazardTest} for the same factory exercised in isolation.
|
||||
*/
|
||||
@Test
|
||||
void theSpawnTimeQuarantineSurvivesTheFleetdStyleForwardingHop(@TempDir Path configRoot,
|
||||
@TempDir Path discRoot) throws Exception {
|
||||
String workDir = provisionedWorkDir(configRoot);
|
||||
FleetConfig.Profile cfg = opencodeCfgWithCredential(
|
||||
"terra", "opencode/nemotron-3-ultra-free", "openai-shared");
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of(cfg.profile(), cfg);
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.SECONDS.toNanos(1800));
|
||||
|
||||
// Fleetd.java:176 — the forwarding sink is built BEFORE the real one can exist, and the
|
||||
// launcher below is constructed against this forwarder, exactly like Fleetd.main. Calling
|
||||
// the SAME factory Fleetd.java calls — ExhaustionSink.forwardingTo — rather than rebuilding
|
||||
// the forwarder's shape here is the whole point (fleetd #234, round 3): a round-2 version of
|
||||
// this test built its own copy, so mutating Fleetd.java's real forwarder back into a lambda
|
||||
// left this test untouched. Routed through the shared factory, a regression at either the
|
||||
// Fleetd.java call site (reverting to a hand-written lambda) or inside the factory body
|
||||
// itself has nowhere left to hide from this test.
|
||||
java.util.concurrent.atomic.AtomicReference<ExhaustionSink> exhaustionSinkRef =
|
||||
new java.util.concurrent.atomic.AtomicReference<>(ExhaustionSink.none());
|
||||
ExhaustionSink forwardingExhaustionSink = ExhaustionSink.forwardingTo(exhaustionSinkRef::get);
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
OpenCodeLauncher launcher = serviceWithSink(herdr, configRoot, discRoot, cfg, forwardingExhaustionSink);
|
||||
SessionManager sessions = new SessionManager(launcher);
|
||||
|
||||
// Fleetd.java:359-390 — the real sink is only built and pointed to AFTER `sessions` exists,
|
||||
// same order as production. Safely a lambda (round 4): see the note on the forwarder above.
|
||||
ExhaustionSink realSink = (target, reason, profileHint) -> {
|
||||
FleetConfig.Profile profile = profileHint == null ? null : profiles.get(profileHint);
|
||||
if (profile != null) {
|
||||
quarantine.quarantine(profile.effectiveCredentialId());
|
||||
}
|
||||
};
|
||||
exhaustionSinkRef.set(realSink);
|
||||
|
||||
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", workDir, 1000L,
|
||||
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
|
||||
|
||||
assertFalse(quarantine.isQuarantined("openai-shared"), "nothing quarantined before the spawn");
|
||||
|
||||
MemberSession acquired = sessions.acquire(cfg.profile(), workDir, null, null);
|
||||
|
||||
assertEquals("ses_x", acquired.agentSessionId(), "the id itself still resolves correctly");
|
||||
assertTrue(quarantine.isQuarantined("openai-shared"),
|
||||
"the profile hint must survive the Fleetd-style forwarding hop and reach the real "
|
||||
+ "sink — a lambda forwarder drops it and this must go red");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -146,14 +146,14 @@ class OpenCodeSessionDiscoveryTest {
|
||||
"an unreadable database resolves to null, not an exception");
|
||||
}
|
||||
|
||||
// --- fleetd #175: actualModelForDirectory / the model JSON column ---------------------------
|
||||
// --- fleetd #175/#234: actualModelForSessionId / the model JSON column -----------------------
|
||||
|
||||
@Test
|
||||
void parsesTheModelJsonIntoProviderAndId(@TempDir Path root) throws Exception {
|
||||
writeRecord(root, "ses_aaa", "/w/a", 1000L, "{\"id\":\"gpt-5.6-terra\",\"providerID\":\"openai\"}");
|
||||
|
||||
OpenCodeSessionDiscovery.ActualModel actual =
|
||||
new OpenCodeSessionDiscovery(root).actualModelForDirectory("/w/a");
|
||||
new OpenCodeSessionDiscovery(root).actualModelForSessionId("ses_aaa");
|
||||
|
||||
assertNotNull(actual, "a well-formed model JSON parses");
|
||||
assertEquals("gpt-5.6-terra", actual.id());
|
||||
@@ -161,28 +161,39 @@ class OpenCodeSessionDiscoveryTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void prefersTheModelOfTheMostRecentlyUpdatedRow(@TempDir Path root) throws Exception {
|
||||
writeRecord(root, "ses_old", "/w/a", 1000L, "{\"id\":\"old-model\",\"providerID\":\"openai\"}");
|
||||
writeRecord(root, "ses_new", "/w/a", 5000L, "{\"id\":\"new-model\",\"providerID\":\"openai\"}");
|
||||
void looksUpByIdEvenWhenAnotherRowInTheSameDirectoryIsNewer(@TempDir Path root) throws Exception {
|
||||
writeRecord(root, "ses_ours", "/work/dir", 1000L, "{\"id\":\"gpt-5.6-terra\",\"providerID\":\"openai\"}");
|
||||
writeRecord(root, "ses_sibling", "/work/dir", 9000L, "{\"id\":\"deepseek-v4-flash\",\"providerID\":\"gx\"}");
|
||||
|
||||
assertEquals("new-model",
|
||||
new OpenCodeSessionDiscovery(root).actualModelForDirectory("/w/a").id(),
|
||||
"the model of the row with the highest time_updated wins, same as the id");
|
||||
OpenCodeSessionDiscovery discovery = new OpenCodeSessionDiscovery(root);
|
||||
assertEquals("gpt-5.6-terra", discovery.actualModelForSessionId("ses_ours").id(),
|
||||
"querying by id reads OUR row, not the directory's newest row");
|
||||
assertEquals("deepseek-v4-flash", discovery.actualModelForSessionId("ses_sibling").id(),
|
||||
"each id resolves to its own row independently of time_updated ordering");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aNonMatchingDirectoryYieldsUnknownModelRatherThanAMismatch(@TempDir Path root) throws Exception {
|
||||
void anUnknownSessionIdYieldsUnknownModelRatherThanAMismatch(@TempDir Path root) throws Exception {
|
||||
writeRecord(root, "ses_aaa", "/w/a", 1000L, "{\"id\":\"x\",\"providerID\":\"y\"}");
|
||||
|
||||
assertNull(new OpenCodeSessionDiscovery(root).actualModelForDirectory("/w/other"),
|
||||
"no row for this cwd yet → unknown, not a wrong model");
|
||||
assertNull(new OpenCodeSessionDiscovery(root).actualModelForSessionId("ses_no_such_row"),
|
||||
"no row for this id yet → unknown, not a wrong model");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aBlankOrNullSessionIdYieldsUnknownModel(@TempDir Path root) throws Exception {
|
||||
writeRecord(root, "ses_aaa", "/w/a", 1000L, "{\"id\":\"x\",\"providerID\":\"y\"}");
|
||||
|
||||
OpenCodeSessionDiscovery discovery = new OpenCodeSessionDiscovery(root);
|
||||
assertNull(discovery.actualModelForSessionId(null));
|
||||
assertNull(discovery.actualModelForSessionId(" "));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aNullModelColumnYieldsUnknownWithoutThrowing(@TempDir Path root) throws Exception {
|
||||
writeRecord(root, "ses_aaa", "/w/a", 1000L, null);
|
||||
|
||||
assertNull(new OpenCodeSessionDiscovery(root).actualModelForDirectory("/w/a"),
|
||||
assertNull(new OpenCodeSessionDiscovery(root).actualModelForSessionId("ses_aaa"),
|
||||
"a row with no model value yet is unknown, not a mismatch");
|
||||
}
|
||||
|
||||
@@ -190,7 +201,7 @@ class OpenCodeSessionDiscoveryTest {
|
||||
void unparseableModelJsonYieldsUnknownWithoutThrowing(@TempDir Path root) throws Exception {
|
||||
writeRecord(root, "ses_aaa", "/w/a", 1000L, "this is not json");
|
||||
|
||||
assertNull(new OpenCodeSessionDiscovery(root).actualModelForDirectory("/w/a"),
|
||||
assertNull(new OpenCodeSessionDiscovery(root).actualModelForSessionId("ses_aaa"),
|
||||
"JSON that fails to parse resolves to unknown, never an exception");
|
||||
}
|
||||
|
||||
@@ -198,13 +209,13 @@ class OpenCodeSessionDiscoveryTest {
|
||||
void modelJsonMissingIdYieldsUnknown(@TempDir Path root) throws Exception {
|
||||
writeRecord(root, "ses_aaa", "/w/a", 1000L, "{\"providerID\":\"openai\"}");
|
||||
|
||||
assertNull(new OpenCodeSessionDiscovery(root).actualModelForDirectory("/w/a"),
|
||||
assertNull(new OpenCodeSessionDiscovery(root).actualModelForSessionId("ses_aaa"),
|
||||
"no id in the JSON → unknown, since id is what a caller actually compares");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aMissingDatabaseYieldsUnknownModelWithoutThrowing(@TempDir Path root) {
|
||||
assertNull(new OpenCodeSessionDiscovery(root).actualModelForDirectory("/w/a"));
|
||||
assertNull(new OpenCodeSessionDiscovery(root).actualModelForSessionId("ses_aaa"));
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -212,7 +223,7 @@ class OpenCodeSessionDiscoveryTest {
|
||||
* shape a real opencode upgrade/downgrade could produce. This must degrade to UNKNOWN for the
|
||||
* model, and — the property that actually matters — must NOT take id resolution down with it.
|
||||
* A combined single query for both columns would fail this test; that is why
|
||||
* {@link OpenCodeSessionDiscovery#actualModelForDirectory} runs its own independent query.
|
||||
* {@link OpenCodeSessionDiscovery#actualModelForSessionId} runs its own independent query.
|
||||
*/
|
||||
@Test
|
||||
void aMissingModelColumnYieldsUnknownButIdResolutionStillWorks(@TempDir Path root) throws Exception {
|
||||
@@ -234,7 +245,7 @@ class OpenCodeSessionDiscoveryTest {
|
||||
OpenCodeSessionDiscovery discovery = new OpenCodeSessionDiscovery(root);
|
||||
assertEquals("ses_aaa", discovery.sessionIdForDirectory("/w/a"),
|
||||
"id resolution must survive a database with no model column at all");
|
||||
assertNull(discovery.actualModelForDirectory("/w/a"),
|
||||
assertNull(discovery.actualModelForSessionId("ses_aaa"),
|
||||
"no model column → unknown, not a throw and not a mismatch");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -56,7 +56,11 @@ class MessageServiceTest {
|
||||
|
||||
/** Run {@code send} on a background thread; the current thread drives the worker's turn. */
|
||||
private CompletableFuture<MessageService.Reply> sendAsync() {
|
||||
return CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 5000));
|
||||
return sendAsync("do the task");
|
||||
}
|
||||
|
||||
private CompletableFuture<MessageService.Reply> sendAsync(String content) {
|
||||
return CompletableFuture.supplyAsync(() -> messages.send(T, content, 5000));
|
||||
}
|
||||
|
||||
private void awaitWaiting() throws InterruptedException {
|
||||
@@ -86,6 +90,94 @@ class MessageServiceTest {
|
||||
assertTrue(reply.completed(), "a scraped completion still counts as completed");
|
||||
}
|
||||
|
||||
@Test
|
||||
void completionFallbackReplacesAnEchoedInjectedBriefWithNoReportOutcome() throws Exception {
|
||||
String brief = "Implement the requested change. ".repeat(20);
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync(brief);
|
||||
awaitWaiting();
|
||||
|
||||
herdr.readText("$ prompt");
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
injector.onStatus(T, AgentStatus.WORKING);
|
||||
herdr.readText("⏺ " + brief + "\n❯ ");
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, reply.outcome());
|
||||
assertEquals(CompletionResolver.NO_REPORT_PREFIX + "]", reply.text(),
|
||||
"the real injector -> completion fallback path must not return the lead's brief");
|
||||
}
|
||||
|
||||
@Test
|
||||
void completionFallbackKeepsARealReportThatRestatesTheWholeBrief() throws Exception {
|
||||
String brief = "Load the implementer skill. You own fleetd #999. ".repeat(12);
|
||||
String report = brief + "\n\n## Report\n"
|
||||
+ ("I implemented the fix in GitWorktrees.java, added five tests, ran mvn clean install "
|
||||
+ "and got 1168 tests with 0 failures. Commit 321d8dc pushed. ").repeat(8);
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync(brief);
|
||||
awaitWaiting();
|
||||
|
||||
herdr.readText("$ prompt");
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
injector.onStatus(T, AgentStatus.WORKING);
|
||||
herdr.readText("⏺ " + report + "\n❯ ");
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, reply.outcome());
|
||||
assertEquals(report.strip(), reply.text(), "a report that restates the whole brief must survive unchanged");
|
||||
}
|
||||
|
||||
@Test
|
||||
void completionFallbackNamesTheKnownWorktreeAndBranchForAnEchoedBrief() throws Exception {
|
||||
FakeHerdr localHerdr = new FakeHerdr();
|
||||
Rendezvous localRendezvous = new Rendezvous();
|
||||
java.util.concurrent.atomic.AtomicLong localClock = new java.util.concurrent.atomic.AtomicLong();
|
||||
CompletionResolver localCompletion = new CompletionResolver(new AgentControl(localHerdr), localRendezvous,
|
||||
ExhaustedPatternLookup.none(), ExhaustionSink.none(),
|
||||
dev.ltms.fleet.inject.BackendErrorPatternLookup.legacy(),
|
||||
dev.ltms.fleet.inject.BackendErrorSink.none(),
|
||||
() -> localClock.addAndGet(CompletionResolver.MIN_TURN_NANOS + 1),
|
||||
target -> new CompletionResolver.WorktreeBranch("/tmp/member-worktree", "worker/cb241"));
|
||||
Injector localInjector = new Injector(new AgentControl(localHerdr), localCompletion);
|
||||
MessageService localMessages = new MessageService(new AgentControl(localHerdr), localInjector,
|
||||
localRendezvous, new InMemoryReplyInbox());
|
||||
String brief = "Implement the requested change. ".repeat(20);
|
||||
CompletableFuture<MessageService.Reply> send =
|
||||
CompletableFuture.supplyAsync(() -> localMessages.send(T, brief, 5000));
|
||||
long deadline = System.currentTimeMillis() + 2000;
|
||||
while (!localRendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(localRendezvous.isWaiting(T));
|
||||
|
||||
localHerdr.readText("$ prompt");
|
||||
localInjector.onStatus(T, AgentStatus.IDLE);
|
||||
localInjector.onStatus(T, AgentStatus.WORKING);
|
||||
localHerdr.readText("⏺ " + brief + "\n❯ ");
|
||||
localInjector.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
assertEquals(CompletionResolver.NO_REPORT_PREFIX + " worktree=/tmp/member-worktree "
|
||||
+ "branch=worker/cb241]", send.get(5, TimeUnit.SECONDS).text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void completionFallbackKeepsTheClippedMarkerWhenAnEchoedBriefIsTooLong() throws Exception {
|
||||
String brief = "a".repeat(4_001); // CompletionResolver's 4,000-character scrape cap
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync(brief);
|
||||
awaitWaiting();
|
||||
|
||||
herdr.readText("$ prompt");
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
injector.onStatus(T, AgentStatus.WORKING);
|
||||
herdr.readText("⏺ " + brief + "\n❯ ");
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
assertEquals(CompletionResolver.NO_REPORT_PREFIX + "]\n"
|
||||
+ "[Pane tail clipped: member did not call fleet_reply.]",
|
||||
send.get(5, TimeUnit.SECONDS).text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void backendErrorScrapeThroughMessageServiceFailsInsteadOfBecomingReplyText() throws Exception {
|
||||
// fleetd#164 (part 2 addendum): a scrape that reads cleanly but is only the backend's own
|
||||
|
||||
@@ -43,6 +43,7 @@ class ReplyPushLoopTest {
|
||||
private static final String PRIMARY = "term_primary";
|
||||
private static final String WORKER = "term_worker";
|
||||
private static final String WORKER2 = "term_worker2";
|
||||
private static final String OTHER_PRIMARY = "term_other_primary";
|
||||
private static final ObjectMapper MAPPER = new ObjectMapper();
|
||||
|
||||
private PrimaryRegistry registry;
|
||||
@@ -529,6 +530,106 @@ class ReplyPushLoopTest {
|
||||
assertTrue(nudge.contains("task-1"), "the ticket must not be dropped: " + nudge);
|
||||
}
|
||||
|
||||
// --- CB-201: backend credential outage nudges -----------------------------------------------
|
||||
|
||||
@Test
|
||||
void oneBackendIncidentForTwoWorkersOfOneLeadSendsOnceAndIsOneShot() throws Exception {
|
||||
var rec = recordingClient();
|
||||
agents = new AgentControl(rec);
|
||||
registry.recordDelegation(WORKER, PRIMARY);
|
||||
registry.recordDelegation(WORKER2, PRIMARY);
|
||||
var loop = loop(2, 100);
|
||||
|
||||
loop.onBackendIncident("incident-1", List.of(WORKER, WORKER2), "cred-a",
|
||||
List.of("terra", "codex"), 47);
|
||||
|
||||
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one lead gets one outage injection");
|
||||
Thread.sleep(200);
|
||||
assertEquals(1, rec.sendCount(), "two workers for one lead must send only one outage notice");
|
||||
String nudge = rec.sentParams().getFirst().getValue().toString();
|
||||
assertTrue(nudge.contains("cred-a"));
|
||||
assertTrue(nudge.contains("codex, terra"));
|
||||
assertTrue(nudge.contains(WORKER) && nudge.contains(WORKER2));
|
||||
assertTrue(nudge.contains("47 remaining seconds"));
|
||||
assertTrue(nudge.contains("fleet_list"));
|
||||
|
||||
loop.onBackendIncident("incident-1", List.of(WORKER, WORKER2), "cred-a",
|
||||
List.of("terra", "codex"), 47);
|
||||
Thread.sleep(250);
|
||||
assertEquals(1, rec.sendCount(), "a delivered incident id must never be sent again");
|
||||
}
|
||||
|
||||
@Test
|
||||
void oneBackendIncidentSendsOnceToEachAffectedLead() throws Exception {
|
||||
var rec = recordingClient();
|
||||
rec.sendLatch = new CountDownLatch(2);
|
||||
agents = new AgentControl(rec);
|
||||
registry.recordDelegation(WORKER, PRIMARY);
|
||||
registry.recordDelegation(WORKER2, OTHER_PRIMARY);
|
||||
var loop = loop(1, 50);
|
||||
|
||||
loop.onBackendIncident("incident-2", List.of(WORKER, WORKER2), "cred-a", List.of("terra"), 60);
|
||||
|
||||
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "each affected lead gets its one notice");
|
||||
assertEquals(2, rec.sendCount());
|
||||
}
|
||||
|
||||
@Test
|
||||
void backendIncidentWaitsForBusyLeadThenCombinesWithFailedTicket() throws Exception {
|
||||
var rec = new BusyThenIdleHerdrClient(2);
|
||||
agents = new AgentControl(rec);
|
||||
var loop = loop(1, 50);
|
||||
|
||||
loop.onTicketTerminal("task-failed", WORKER, true);
|
||||
loop.onBackendIncident("incident-3", List.of(WORKER), "cred-b", List.of("terra"), 31);
|
||||
|
||||
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "busy lead should receive deferred combined nudge");
|
||||
assertEquals(1, rec.sendCount(), "ticket and outage must use the same injection");
|
||||
String nudge = rec.sentParams().getFirst().getValue().toString();
|
||||
assertTrue(nudge.contains("task-failed") && nudge.contains("cred-b"),
|
||||
"the one nudge must name both the failed ticket and outage: " + nudge);
|
||||
}
|
||||
|
||||
@Test
|
||||
void failedBackendIncidentSendRetriesWithinTheExistingReminderCap() throws Exception {
|
||||
var rec = new ThrowOnceHerdrClient();
|
||||
agents = new AgentControl(rec);
|
||||
var loop = loop(2, 50);
|
||||
|
||||
loop.onBackendIncident("incident-4", List.of(WORKER), "cred-c", List.of("terra"), 20);
|
||||
|
||||
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "failed first send must retry under maxReminders");
|
||||
assertEquals(2, rec.promptAttempts.get(), "one failed send and one retry use the existing cap");
|
||||
}
|
||||
|
||||
@Test
|
||||
void unmappedBackendTargetUsesTheKnownLeadSchedule() throws Exception {
|
||||
var rec = recordingClient();
|
||||
agents = new AgentControl(rec);
|
||||
var loop = loop(1, 50);
|
||||
|
||||
loop.onBackendTargetUnmapped(WORKER, "no credential matched");
|
||||
|
||||
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS));
|
||||
String nudge = ((Map<?, ?>) rec.sentParams().getFirst().getValue()).get("text").toString();
|
||||
assertEquals("A backend error on worker term_worker could not be mapped to a credential "
|
||||
+ "(no credential matched) — no cool-off was applied. "
|
||||
+ "Run fleet_list to check that worker.",
|
||||
nudge);
|
||||
}
|
||||
|
||||
@Test
|
||||
void stopClearsPendingBackendIncidents() {
|
||||
agents = agentWithStatus("idle");
|
||||
var loop = loop(1, 100_000);
|
||||
loop.onBackendIncident("incident-stop", List.of(WORKER), "cred-stop", List.of("terra"), 60);
|
||||
|
||||
loop.stop();
|
||||
|
||||
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 0, 0),
|
||||
"stop must clear a pending backend incident as well as older sources");
|
||||
}
|
||||
|
||||
// --- CB-590 follow-up: per-source reminder budgets — the regression this round exists for ---
|
||||
|
||||
@Test
|
||||
@@ -928,6 +1029,31 @@ class ReplyPushLoopTest {
|
||||
return new RecordingHerdrClient();
|
||||
}
|
||||
|
||||
/** Fails its first prompt call, then records the retry. */
|
||||
private static final class ThrowOnceHerdrClient implements HerdrClient {
|
||||
private final AtomicInteger promptAttempts = new AtomicInteger();
|
||||
private final CountDownLatch sendLatch = new CountDownLatch(1);
|
||||
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) {
|
||||
if ("agent.get".equals(method)) {
|
||||
return MAPPER.createObjectNode().set("agent", MAPPER.createObjectNode()
|
||||
.put("terminal_id", PRIMARY).put("agent_status", "idle"));
|
||||
}
|
||||
if ("agent.prompt".equals(method) && promptAttempts.getAndIncrement() == 0) {
|
||||
throw new IllegalStateException("first prompt fails");
|
||||
}
|
||||
if ("agent.prompt".equals(method)) {
|
||||
sendLatch.countDown();
|
||||
}
|
||||
return MAPPER.createObjectNode();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Thread-safe fake that reports {@code working} (not injectable) for its first
|
||||
* {@code busyChecks} status calls, then {@code idle} forever after — used to prove work queued
|
||||
|
||||
@@ -30,7 +30,14 @@ class PlacementPolicyTest {
|
||||
private static PlacementContext ctx(List<PlacementCandidate> candidates,
|
||||
Function<String, Integer> liveCount,
|
||||
Set<String> unreachable, Set<String> quarantined) {
|
||||
return new PlacementContext("b", candidates, liveCount, unreachable, quarantined);
|
||||
return ctx(candidates, liveCount, unreachable, quarantined, Set.of());
|
||||
}
|
||||
|
||||
private static PlacementContext ctx(List<PlacementCandidate> candidates,
|
||||
Function<String, Integer> liveCount,
|
||||
Set<String> unreachable, Set<String> quarantined,
|
||||
Set<String> coolingOff) {
|
||||
return new PlacementContext("b", candidates, liveCount, unreachable, quarantined, coolingOff);
|
||||
}
|
||||
|
||||
private static PlacementContext ctx(List<PlacementCandidate> candidates,
|
||||
@@ -52,14 +59,14 @@ class PlacementPolicyTest {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext(null,
|
||||
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
||||
noSessions(), Set.of(), Set.of());
|
||||
noSessions(), Set.of(), Set.of(), Set.of());
|
||||
assertEquals("a", policy.select(ctx).profile());
|
||||
}
|
||||
|
||||
@Test
|
||||
void fixedThrowsWhenNoProfilesAndNoDefault() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext(null, List.of(), noSessions(), Set.of(), Set.of());
|
||||
PlacementContext ctx = new PlacementContext(null, List.of(), noSessions(), Set.of(), Set.of(), Set.of());
|
||||
assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||
}
|
||||
|
||||
@@ -68,7 +75,7 @@ class PlacementPolicyTest {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
||||
noSessions(), Set.of(), Set.of("b"));
|
||||
noSessions(), Set.of(), Set.of("b"), Set.of());
|
||||
assertEquals("a", policy.select(ctx).profile(),
|
||||
"the default 'b' is quarantined, so fixed falls through to the first un-quarantined candidate");
|
||||
}
|
||||
@@ -78,7 +85,7 @@ class PlacementPolicyTest {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
||||
noSessions(), Set.of(), Set.of("a", "b"));
|
||||
noSessions(), Set.of(), Set.of("a", "b"), Set.of());
|
||||
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
|
||||
}
|
||||
@@ -211,6 +218,67 @@ class PlacementPolicyTest {
|
||||
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
|
||||
}
|
||||
|
||||
// --- fleetd #201 Unit 5: coolingOff excludes a candidate, SEPARATE from quarantined --------
|
||||
|
||||
@Test
|
||||
void roundRobinSkipsCoolingOffProfiles() {
|
||||
PlacementPolicy policy = PlacementPolicies.roundRobin();
|
||||
List<PlacementCandidate> candidates = List.of(
|
||||
PlacementCandidate.profile("a"),
|
||||
PlacementCandidate.profile("b"));
|
||||
PlacementContext ctx = ctx(candidates, noSessions(), Set.of(), Set.of(), Set.of("a"));
|
||||
for (int i = 0; i < 5; i++) {
|
||||
assertEquals("b", policy.select(ctx).profile(), "a is cooling off, so every pick lands on b");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void weightedSkipsCoolingOffProfile() {
|
||||
PlacementPolicy policy = PlacementPolicies.weighted();
|
||||
List<PlacementCandidate> candidates = List.of(
|
||||
PlacementCandidate.profile("a", 1.0f, null),
|
||||
PlacementCandidate.profile("b", 1.0f, null));
|
||||
PlacementContext ctx = ctx(candidates, noSessions(), Set.of(), Set.of(), Set.of("a"));
|
||||
for (int i = 0; i < 5; i++) {
|
||||
assertEquals("b", policy.select(ctx).profile(), "a is cooling off, so every pick lands on b");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void weightedThrowsWhenAllCoolingOff() {
|
||||
PlacementPolicy policy = PlacementPolicies.weighted();
|
||||
List<PlacementCandidate> candidates = List.of(
|
||||
PlacementCandidate.profile("a"),
|
||||
PlacementCandidate.profile("b"));
|
||||
PlacementException e = assertThrows(PlacementException.class,
|
||||
() -> policy.select(ctx(candidates, noSessions(), Set.of(), Set.of(), Set.of("a", "b"))));
|
||||
assertTrue(e.getMessage().contains("cooling off"), e.getMessage());
|
||||
assertFalse(e.getMessage().contains("quarantined"),
|
||||
"nothing here is quarantined — the message must say cooling off, not exhausted: " + e.getMessage());
|
||||
}
|
||||
|
||||
/**
|
||||
* A candidate that is BOTH quarantined (CB-578 stage B) and cooling off (fleetd #201 Unit 5)
|
||||
* counts only as quarantined — exhaustion takes priority, matching
|
||||
* {@code CompositePeerLauncher}'s explicit-spawn check order.
|
||||
*/
|
||||
@Test
|
||||
void quarantinedAndCoolingOffIsNotDoubleCounted() {
|
||||
PlacementPolicy policy = PlacementPolicies.weighted();
|
||||
List<PlacementCandidate> candidates = List.of(
|
||||
PlacementCandidate.profile("a"),
|
||||
PlacementCandidate.profile("b"));
|
||||
// a is BOTH quarantined and cooling off; b is quarantined only. If a were double-bucketed
|
||||
// as cooling-off instead of quarantined, quarantined would undercount to 1 of 2 candidates
|
||||
// and this would fall through to the generic "no worker profile available: N quarantined, M
|
||||
// cooling off, ..." message instead — which also happens to contain the substring
|
||||
// "quarantined", so a loose contains() check here would pass either way. Pinning the exact
|
||||
// "all ... quarantined" message is what actually proves the two are not double-counted.
|
||||
PlacementContext ctx = ctx(candidates, noSessions(), Set.of(), Set.of("a", "b"), Set.of("a"));
|
||||
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||
assertEquals("all worker profiles are quarantined (backend exhausted)", e.getMessage());
|
||||
}
|
||||
|
||||
@Test
|
||||
void weightedThrowsWhenAllUnreachable() {
|
||||
PlacementPolicy policy = PlacementPolicies.weighted();
|
||||
@@ -320,7 +388,7 @@ class PlacementPolicyTest {
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a", 1.0f, null),
|
||||
PlacementCandidate.profile("b", 0.0f, null)),
|
||||
noSessions(), Set.of(), Set.of());
|
||||
noSessions(), Set.of(), Set.of(), Set.of());
|
||||
assertEquals("a", policy.select(ctx).profile(),
|
||||
"the default 'b' has weight 0, so fixed falls through to the first available candidate");
|
||||
}
|
||||
@@ -331,7 +399,7 @@ class PlacementPolicyTest {
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a", 0.0f, null),
|
||||
PlacementCandidate.profile("b", 0.0f, null)),
|
||||
noSessions(), Set.of(), Set.of());
|
||||
noSessions(), Set.of(), Set.of(), Set.of());
|
||||
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||
assertTrue(e.getMessage().contains("weight 0"), e.getMessage());
|
||||
}
|
||||
|
||||
@@ -0,0 +1,121 @@
|
||||
package dev.ltms.fleet.rest;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.Locale;
|
||||
import java.util.Set;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #252: the REST surface has no supported operator entry point of its own — it is the
|
||||
* documented fallback for when the MCP mount drops (see the operator wiki's REST page), and
|
||||
* nothing has ever checked a hand-written route list against it. Proof it drifts: the ticket
|
||||
* itself was filed with 14 routes, and {@link FleetApp} registers 15 — {@code GET
|
||||
* /member-credentials} (fleetd #111) shipped hours before the ticket and was already missing from
|
||||
* its list.
|
||||
*
|
||||
* <p>This test is modelled on {@code McpContractDocTest} (fleetd #114 / CB-609), which solved the
|
||||
* same shape of problem for the MCP tool catalogue: read the source text with a regex instead of
|
||||
* trusting a maintained copy. Here the "doc" is an explicit inventory written directly in this
|
||||
* test rather than a separate Markdown file, because the operator wiki page lives in a submodule
|
||||
* a worker cannot read reliably (see {@code CLAUDE.md} → Project addendum). Keeping the expected
|
||||
* list in the test means it still fails loudly the moment {@link FleetApp} changes, which is the
|
||||
* property that actually matters; a human keeps the wiki page in sync using the failure message
|
||||
* below as the diff.
|
||||
*
|
||||
* <p><b>It checks source text, not behaviour.</b> It reads {@link FleetApp}'s source for {@code
|
||||
* app.<verb>("<path>")} registrations and does not boot a server. It cannot catch a route that is
|
||||
* registered through some other mechanism entirely (a filter, a redirect) — only ones shaped like
|
||||
* the {@code app.get/post/delete/put/patch(...)} calls every route here actually uses.
|
||||
*/
|
||||
class RestRouteInventoryTest {
|
||||
|
||||
/** Tests run with the module directory as cwd. */
|
||||
private static final Path REST_SOURCE = Path.of("src/main/java/dev/ltms/fleet/rest/FleetApp.java");
|
||||
|
||||
/**
|
||||
* The REST surface as verified against {@link FleetApp} on 2026-09-03 (fleetd #252). Update
|
||||
* this list AND the operator wiki's REST-surface entry together whenever a route is added,
|
||||
* removed, or renamed — never one without the other.
|
||||
*
|
||||
* <p>{@code /mcp} is deliberately excluded: it is a raw Jetty {@code ServletHolder} mount
|
||||
* (see {@code FleetApp.build()}, around the {@code modifyServletContextHandler} call), not an
|
||||
* {@code app.<verb>(...)} route, so it is a different registration mechanism and this test's
|
||||
* regex does not — and should not — see it. If {@code /mcp} ever moves to a Javalin route,
|
||||
* add it here explicitly rather than relying on the regex to pick it up by accident.
|
||||
*/
|
||||
private static final Set<String> EXPECTED_ROUTES = Set.of(
|
||||
"GET /healthz",
|
||||
"GET /metrics",
|
||||
"GET /sessions",
|
||||
"GET /agents",
|
||||
"GET /members",
|
||||
"GET /profiles",
|
||||
"GET /member-credentials",
|
||||
"POST /members",
|
||||
"DELETE /members/{paneId}",
|
||||
"POST /sessions/{id}/message",
|
||||
"POST /sessions/{id}/reply",
|
||||
"GET /sessions/{id}/replies",
|
||||
"POST /sessions/{id}/ask",
|
||||
"GET /sessions/{id}/status",
|
||||
"GET /tasks/{ticket}"
|
||||
);
|
||||
|
||||
/**
|
||||
* Every {@code app.<verb>("<path>")} call in {@link FleetApp}'s source, as {@code "VERB path"}.
|
||||
* This matches inside an {@code if (...) { ... }} block just as well as a top-level statement —
|
||||
* the regex only looks for the call shape, not its surrounding control flow — which is what
|
||||
* catches {@code GET /metrics} (registered conditionally on {@code metrics != null}).
|
||||
*/
|
||||
private static Set<String> routesTheServerRegisters() throws Exception {
|
||||
String source = Files.readString(REST_SOURCE);
|
||||
Matcher m = Pattern.compile("app\\.(get|post|delete|put|patch)\\(\\s*\"([^\"]+)\"").matcher(source);
|
||||
Set<String> found = new LinkedHashSet<>();
|
||||
while (m.find()) {
|
||||
found.add(m.group(1).toUpperCase(Locale.ROOT) + " " + m.group(2));
|
||||
}
|
||||
return found;
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] FleetApp registers exactly the documented REST route inventory")
|
||||
void theRegisteredRoutesMatchTheExpectedInventory() throws Exception {
|
||||
Set<String> actual = routesTheServerRegisters();
|
||||
|
||||
Set<String> added = new LinkedHashSet<>(actual);
|
||||
added.removeAll(EXPECTED_ROUTES);
|
||||
Set<String> removed = new LinkedHashSet<>(EXPECTED_ROUTES);
|
||||
removed.removeAll(actual);
|
||||
|
||||
assertTrue(added.isEmpty() && removed.isEmpty(),
|
||||
"FleetApp's registered REST routes no longer match this test's expected inventory. "
|
||||
+ "Added (in FleetApp, not in this test): " + added + ". "
|
||||
+ "Removed (in this test, not in FleetApp): " + removed + ". "
|
||||
+ "Update EXPECTED_ROUTES in RestRouteInventoryTest AND the operator wiki's "
|
||||
+ "REST-surface entry together — this is the fleetd #252 defect: the route "
|
||||
+ "list drifted for a month with nothing checking it. Do NOT weaken this test.");
|
||||
}
|
||||
|
||||
/**
|
||||
* The denominator guard (same shape as {@code McpContractDocTest}'s vacuity check). Pins that
|
||||
* the regex really is still finding registrations, so a scrape that silently stops matching
|
||||
* can't make the check above pass by finding nothing on both sides.
|
||||
*/
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] the route scrape is not vacuous — it found the expected count")
|
||||
void theScrapeActuallyFoundRoutes() throws Exception {
|
||||
Set<String> actual = routesTheServerRegisters();
|
||||
assertTrue(actual.size() >= EXPECTED_ROUTES.size(),
|
||||
"scraped only " + actual.size() + " route registration(s) from FleetApp (" + actual
|
||||
+ "), but this test expects at least " + EXPECTED_ROUTES.size()
|
||||
+ "; the app.<verb>(\"...\") scrape has stopped matching and the check above "
|
||||
+ "is now vacuous");
|
||||
}
|
||||
}
|
||||
@@ -666,6 +666,103 @@ class GitWorktreesTest {
|
||||
assertEquals("", status(Path.of(wt), ".autoenv"), ".autoenv still shows as modified");
|
||||
}
|
||||
|
||||
// ---- fleetd #134: isolateToolSurface must report what it neutralized (to the daemon operator's
|
||||
// log) and record it where the worker itself can read it (worktree-scoped git config), without
|
||||
// ever showing up in the worker's own `git status`. All drive the real provisioning path,
|
||||
// GitWorktrees#add, per criterion 5 — the whole provisioned worktree is what's under test here. ----
|
||||
|
||||
private static Path initRepoWithAllThreeConfigs(Path dir) throws Exception {
|
||||
Files.createDirectories(dir);
|
||||
git(dir, "init", "-q", "-b", "main");
|
||||
git(dir, "config", "user.email", "test@example.invalid");
|
||||
git(dir, "config", "user.name", "Test");
|
||||
Files.writeString(dir.resolve(".mcp.json"), WITH_SERVERS);
|
||||
Files.writeString(dir.resolve("opencode.json"), OPENCODE_WITH_FILE_REF);
|
||||
Files.writeString(dir.resolve(".autoenv"), AUTOENV_WITH_DIRECTIVE);
|
||||
Files.writeString(dir.resolve("README.md"), "seed\n");
|
||||
git(dir, "add", ".mcp.json", "opencode.json", ".autoenv", "README.md");
|
||||
git(dir, "commit", "-q", "-m", "seed");
|
||||
return dir;
|
||||
}
|
||||
|
||||
/** Criterion 1, all three present: the summary names the denominator and every neutralized file. */
|
||||
@Test
|
||||
void isolateToolSurfaceLogsAllThreeConfigsNeutralized(@TempDir Path tmp) throws Exception {
|
||||
reportingLogger.setLevel(Level.INFO);
|
||||
Path repo = initRepoWithAllThreeConfigs(tmp.resolve("repo"));
|
||||
|
||||
new GitWorktrees(tmp.resolve("wts").toString()).add(repo.toString(), "cb-134-log-all", "HEAD");
|
||||
|
||||
assertTrue(capturedMessages().contains(
|
||||
"tool-surface isolation: neutralized 3 of 3 configs: .mcp.json, opencode.json, "
|
||||
+ ".autoenv — the worktree copy is a stub, not the repo's file; edit the "
|
||||
+ "real file in the primary checkout instead"),
|
||||
"expected the all-neutralized summary line, got:\n" + capturedMessages());
|
||||
}
|
||||
|
||||
/** Criterion 1, two absent: the summary must still name the denominator and say why. */
|
||||
@Test
|
||||
void isolateToolSurfaceLogsAbsentConfigsWithReason(@TempDir Path tmp) throws Exception {
|
||||
reportingLogger.setLevel(Level.INFO);
|
||||
Path repo = initRepo(tmp.resolve("repo")); // only .mcp.json + README committed
|
||||
|
||||
new GitWorktrees(tmp.resolve("wts").toString()).add(repo.toString(), "cb-134-log-partial", "HEAD");
|
||||
|
||||
assertTrue(capturedMessages().contains(
|
||||
"tool-surface isolation: neutralized 1 of 3 configs: .mcp.json (opencode.json "
|
||||
+ "absent, .autoenv absent) — the worktree copy is a stub, not the repo's "
|
||||
+ "file; edit the real file in the primary checkout instead"),
|
||||
"expected the partial summary line, got:\n" + capturedMessages());
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 2. The daemon's own log is invisible to the worker process — it never reads fleetd's
|
||||
* stdout. This is the mechanism the worker itself can query, from inside its own worktree, to
|
||||
* learn "this file is neutralized here, the repo's real file differs" instead of trusting what it
|
||||
* just read on disk.
|
||||
*/
|
||||
@Test
|
||||
void aWorkerCanDiscoverNeutralizedConfigsFromWorktreeScopedGitConfig(@TempDir Path tmp) throws Exception {
|
||||
Path repo = initRepoWithAllThreeConfigs(tmp.resolve("repo"));
|
||||
|
||||
String wt = new GitWorktrees(tmp.resolve("wts").toString())
|
||||
.add(repo.toString(), "cb-134-discover", "HEAD");
|
||||
|
||||
String recorded = gitOutput(Path.of(wt), "config", "--worktree", "--get-all", "fleet.neutralizedConfig");
|
||||
Set<String> files = new HashSet<>();
|
||||
for (String line : recorded.split("\\R")) {
|
||||
if (!line.isBlank()) {
|
||||
files.add(line.trim());
|
||||
}
|
||||
}
|
||||
assertEquals(Set.of(".mcp.json", "opencode.json", ".autoenv"), files,
|
||||
"the worker-readable record must name every neutralized file: " + recorded);
|
||||
|
||||
String note = gitOutput(Path.of(wt), "config", "--worktree", "--get", "fleet.neutralizedConfigNote").trim();
|
||||
assertTrue(note.contains("stub"), "note must say the worktree copy is a stub: " + note);
|
||||
assertTrue(note.contains("primary checkout"),
|
||||
"note must state the consequence — where to edit the real file instead: " + note);
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 3. Whatever fleetd#134's worker-discovery mechanism writes must never appear as
|
||||
* untracked or modified in the worker's own `git status` — a worker that sees a stray file either
|
||||
* commits it by mistake or burns a turn asking about it. This runs the full porcelain status, not
|
||||
* a single-file check, so any leftover file anywhere in the worktree would fail it.
|
||||
*/
|
||||
@Test
|
||||
void aProvisionedWorktreeHasCleanGitStatusDespiteNeutralizedConfigRecordkeeping(@TempDir Path tmp)
|
||||
throws Exception {
|
||||
Path repo = initRepoWithAllThreeConfigs(tmp.resolve("repo"));
|
||||
|
||||
String wt = new GitWorktrees(tmp.resolve("wts").toString())
|
||||
.add(repo.toString(), "cb-134-clean-status", "HEAD");
|
||||
|
||||
assertEquals("", fullStatus(Path.of(wt)),
|
||||
"a freshly provisioned worktree must show a clean `git status --porcelain`, including "
|
||||
+ "after the worker-readable neutralized-config record was written");
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-578 stage C, acceptance criterion 1. A dirty worktree — a tracked edit plus a brand-new
|
||||
* untracked file, exactly the shape lost in CB-576 — must land in {@code refs/wip/<branch>}'s
|
||||
@@ -1177,4 +1274,112 @@ class GitWorktreesTest {
|
||||
+ "the refusal must happen before `git worktree add` ever runs");
|
||||
}
|
||||
}
|
||||
|
||||
// ---- fleetd #134 / #148 point 3: overlayParity must report what it did, and marking a tracked
|
||||
// file --skip-worktree must say the file can no longer be committed from this worktree. Drives
|
||||
// overlayParity directly against a real worktree (git worktree add, no GitWorktrees#add) so these
|
||||
// tests are independent of origin/credential-helper provisioning, which is not under test here. ----
|
||||
|
||||
/** A bare worktree, sibling to {@code repo}, created with plain git — the target overlayParity
|
||||
* copies into. Deliberately not {@link GitWorktrees#add}: that method does unrelated
|
||||
* provisioning (origin rewrite, .mcp.json neutralization, credential helper) that would only
|
||||
* add noise to the log assertions below. */
|
||||
private static Path bareWorktree(Path repo, Path wtDir, String branch) throws Exception {
|
||||
git(repo, "worktree", "add", "-q", wtDir.toString(), "-b", branch, "HEAD");
|
||||
return wtDir;
|
||||
}
|
||||
|
||||
/** Acceptance criterion 1: only the configured candidates land in the worktree — nothing else
|
||||
* from the source tree leaks in alongside them. */
|
||||
@Test
|
||||
void overlayParityCopiesExactlyTheConfiguredFilesAndNothingElse(@TempDir Path tmp) throws Exception {
|
||||
Path repo = initRepo(tmp.resolve("repo"));
|
||||
Files.writeString(repo.resolve(".env"), "A=1\n");
|
||||
Files.writeString(repo.resolve(".envrc"), "export A=1\n");
|
||||
Files.writeString(repo.resolve("not-overlaid.txt"), "must not be copied\n");
|
||||
Path wt = bareWorktree(repo, tmp.resolve("wt"), "cb134-exact");
|
||||
|
||||
new GitWorktrees(tmp.resolve("wts").toString())
|
||||
.overlayParity(repo.toString(), wt.toString(), List.of(".env", ".envrc"));
|
||||
|
||||
assertEquals("A=1\n", Files.readString(wt.resolve(".env")));
|
||||
assertEquals("export A=1\n", Files.readString(wt.resolve(".envrc")));
|
||||
assertFalse(Files.exists(wt.resolve("not-overlaid.txt")),
|
||||
"overlayParity must copy only the configured candidates, not the whole source tree");
|
||||
}
|
||||
|
||||
/** Criterion 2: both candidates present — the summary line names both and the denominator. */
|
||||
@Test
|
||||
void overlayParityLogsBothCopiedWhenBothCandidatesArePresent(@TempDir Path tmp) throws Exception {
|
||||
reportingLogger.setLevel(Level.INFO);
|
||||
Path repo = initRepo(tmp.resolve("repo"));
|
||||
Files.writeString(repo.resolve(".env"), "A=1\n");
|
||||
Files.writeString(repo.resolve(".envrc"), "export A=1\n");
|
||||
Path wt = bareWorktree(repo, tmp.resolve("wt"), "cb134-both");
|
||||
|
||||
new GitWorktrees(tmp.resolve("wts").toString())
|
||||
.overlayParity(repo.toString(), wt.toString(), List.of(".env", ".envrc"));
|
||||
|
||||
assertTrue(capturedMessages().contains("parity overlay: copied 2 of 2 candidates: .env, .envrc"),
|
||||
"expected the both-copied summary line, got:\n" + capturedMessages());
|
||||
}
|
||||
|
||||
/** Criterion 2: one candidate present, one absent — the summary must name the copied file, the
|
||||
* denominator, and why the other candidate was not copied. */
|
||||
@Test
|
||||
void overlayParityLogsOneCopiedOneAbsent(@TempDir Path tmp) throws Exception {
|
||||
reportingLogger.setLevel(Level.INFO);
|
||||
Path repo = initRepo(tmp.resolve("repo"));
|
||||
Files.writeString(repo.resolve(".env"), "A=1\n");
|
||||
// .envrc deliberately not created — the absent candidate.
|
||||
Path wt = bareWorktree(repo, tmp.resolve("wt"), "cb134-partial");
|
||||
|
||||
new GitWorktrees(tmp.resolve("wts").toString())
|
||||
.overlayParity(repo.toString(), wt.toString(), List.of(".env", ".envrc"));
|
||||
|
||||
assertTrue(Files.exists(wt.resolve(".env")));
|
||||
assertFalse(Files.exists(wt.resolve(".envrc")));
|
||||
assertTrue(capturedMessages().contains("parity overlay: copied 1 of 2 candidates: .env (.envrc absent)"),
|
||||
"expected the copied/absent summary line, got:\n" + capturedMessages());
|
||||
}
|
||||
|
||||
/** Criterion 3: a tracked candidate is marked --skip-worktree, and that must be named in the log
|
||||
* with the consequence spelled out — a worker editing it afterward finds git ignoring the
|
||||
* change, silently, unless this line told it so beforehand. */
|
||||
@Test
|
||||
void overlayParityLogsSkipWorktreeConsequenceForATrackedFile(@TempDir Path tmp) throws Exception {
|
||||
reportingLogger.setLevel(Level.INFO);
|
||||
Path repo = initRepo(tmp.resolve("repo"));
|
||||
Files.writeString(repo.resolve(".env"), "A=1\n");
|
||||
git(repo, "add", ".env");
|
||||
git(repo, "commit", "-q", "-m", "track env");
|
||||
Path wt = bareWorktree(repo, tmp.resolve("wt"), "cb134-tracked");
|
||||
// Change the source after the worktree checkout, so the overlay copy actually overwrites it.
|
||||
Files.writeString(repo.resolve(".env"), "A=2\n");
|
||||
|
||||
new GitWorktrees(tmp.resolve("wts").toString())
|
||||
.overlayParity(repo.toString(), wt.toString(), List.of(".env"));
|
||||
|
||||
assertEquals("A=2\n", Files.readString(wt.resolve(".env")));
|
||||
assertEquals("", status(wt, ".env"),
|
||||
"the skip-worktree'd file must not show as modified even though its content changed");
|
||||
assertTrue(capturedMessages().contains(
|
||||
"parity overlay marked --skip-worktree (cannot be committed from this worktree): .env"),
|
||||
"expected the skip-worktree consequence line, got:\n" + capturedMessages());
|
||||
}
|
||||
|
||||
/** Criterion 5: null and empty overlay lists return quietly — no exception, no log noise. */
|
||||
@Test
|
||||
void overlayParityWithNoCandidatesLogsNothing(@TempDir Path tmp) throws Exception {
|
||||
reportingLogger.setLevel(Level.INFO);
|
||||
Path repo = initRepo(tmp.resolve("repo"));
|
||||
Path wt = bareWorktree(repo, tmp.resolve("wt"), "cb134-empty");
|
||||
GitWorktrees worktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||
|
||||
worktrees.overlayParity(repo.toString(), wt.toString(), null);
|
||||
worktrees.overlayParity(repo.toString(), wt.toString(), List.of());
|
||||
|
||||
assertTrue(reportingAppender.list.isEmpty(),
|
||||
"a null/empty overlay must log nothing, got:\n" + capturedMessages());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -300,6 +300,151 @@ class SessionManagerTest {
|
||||
assertTrue(sessions.roster().contains(updated), "FAILED is still in acquired-minus-released roster");
|
||||
}
|
||||
|
||||
@Test
|
||||
void backendErrorWinsBeforeNormalCompletionFromBusy() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
sessions.asPresence().markPresent(session.terminalId());
|
||||
sessions.onDelivered(session.terminalId(), TestTurnTokens.inert(session.terminalId()));
|
||||
|
||||
assertTrue(sessions.onBackendError(session.terminalId(), "backend exited"));
|
||||
sessions.onTurnComplete(session.terminalId());
|
||||
|
||||
MemberSession updated = sessions.get(session.paneId()).orElseThrow();
|
||||
assertEquals(MemberSession.State.BACKEND_ERROR, updated.state(),
|
||||
"normal completion must not restore DONE after a backend error from BUSY");
|
||||
assertEquals("backend exited", updated.failureReason());
|
||||
}
|
||||
|
||||
@Test
|
||||
void backendErrorWinsAfterNormalCompletionFromDone() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
sessions.asPresence().markPresent(session.terminalId());
|
||||
sessions.onDelivered(session.terminalId(), TestTurnTokens.inert(session.terminalId()));
|
||||
sessions.onTurnComplete(session.terminalId());
|
||||
|
||||
assertTrue(sessions.onBackendError(session.terminalId(), "backend exited"));
|
||||
|
||||
MemberSession updated = sessions.get(session.paneId()).orElseThrow();
|
||||
assertEquals(MemberSession.State.BACKEND_ERROR, updated.state(),
|
||||
"backend error must replace DONE when normal completion won first");
|
||||
assertEquals("backend exited", updated.failureReason());
|
||||
}
|
||||
|
||||
@Test
|
||||
void backendErrorMemberCannotReceiveAnotherDelivery() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
sessions.asPresence().markPresent(session.terminalId());
|
||||
sessions.onDelivered(session.terminalId(), TestTurnTokens.inert(session.terminalId()));
|
||||
assertTrue(sessions.onBackendError(session.terminalId(), "backend exited"));
|
||||
|
||||
sessions.onDelivered(session.terminalId(), TestTurnTokens.inert(session.terminalId()));
|
||||
|
||||
assertEquals(MemberSession.State.BACKEND_ERROR, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"onDelivered must refuse a backend-error member");
|
||||
}
|
||||
|
||||
@Test
|
||||
void losingCompletionDoesNotReleaseOrClearABackendErrorMember() {
|
||||
assertLosingCompletionLeavesBackendErrorIntact(1, false,
|
||||
"a stale completion must not release a backend-error member at the context cap");
|
||||
assertLosingCompletionLeavesBackendErrorIntact(0, true,
|
||||
"a stale completion must not clear the context of a backend-error member");
|
||||
}
|
||||
|
||||
private void assertLosingCompletionLeavesBackendErrorIntact(int contextCap, boolean clearAfterTurn,
|
||||
String protectedAction) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Profile cfg = new FleetConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher delegate = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
ClearContextSpyLauncher launcher = new ClearContextSpyLauncher(delegate);
|
||||
java.util.concurrent.atomic.AtomicReference<SessionManager> manager = new java.util.concurrent.atomic.AtomicReference<>();
|
||||
java.util.concurrent.atomic.AtomicReference<String> terminal = new java.util.concurrent.atomic.AtomicReference<>();
|
||||
java.util.concurrent.atomic.AtomicBoolean armed = new java.util.concurrent.atomic.AtomicBoolean();
|
||||
SessionManager sessions = new SessionManager(launcher, new RecordingWorktrees(), () -> {
|
||||
if (armed.compareAndSet(true, false)) {
|
||||
manager.get().onBackendError(terminal.get(), "backend exited during completion");
|
||||
}
|
||||
return 1;
|
||||
}, contextCap, clearAfterTurn);
|
||||
manager.set(sessions);
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
terminal.set(session.terminalId());
|
||||
sessions.asPresence().markPresent(session.terminalId());
|
||||
sessions.onDelivered(session.terminalId(), TestTurnTokens.inert(session.terminalId()));
|
||||
armed.set(true);
|
||||
|
||||
assertFalse(sessions.onTurnCompleteWithPostAction(session.terminalId()),
|
||||
"the completion CAS loses after the injected backend error");
|
||||
|
||||
assertTrue(sessions.get(session.paneId()).isPresent(), protectedAction);
|
||||
MemberSession updated = sessions.get(session.paneId()).orElseThrow();
|
||||
assertEquals(MemberSession.State.BACKEND_ERROR, updated.state());
|
||||
assertEquals("backend exited during completion", updated.failureReason());
|
||||
assertFalse(herdr.called("pane.close"), protectedAction);
|
||||
assertEquals(0, launcher.clearContextCalls(), protectedAction);
|
||||
}
|
||||
|
||||
@Test
|
||||
void backendErrorForUnknownTargetIsWarnedAndDoesNotCreateASession() {
|
||||
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
|
||||
ch.qos.logback.classic.Logger sessionLog =
|
||||
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.setContext(ctx);
|
||||
appender.start();
|
||||
sessionLog.addAppender(appender);
|
||||
try {
|
||||
SessionManager sessions = sessionManager(new FakeHerdr());
|
||||
|
||||
assertFalse(sessions.onBackendError("term_missing", "backend exited"));
|
||||
|
||||
assertTrue(sessions.roster().isEmpty(), "unknown target must not create a session");
|
||||
assertTrue(appender.list.stream().anyMatch(e -> e.getLevel().equals(Level.WARN)
|
||||
&& e.getFormattedMessage().contains("term_missing")),
|
||||
"unknown target is logged at WARN");
|
||||
} finally {
|
||||
sessionLog.detachAppender(appender);
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void backendErrorForReleasedTargetDoesNotRecreateTheSession() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
sessions.release(session.paneId());
|
||||
|
||||
assertFalse(sessions.onBackendError(session.terminalId(), "backend exited"));
|
||||
assertTrue(sessions.get(session.paneId()).isEmpty(), "released session must stay absent");
|
||||
}
|
||||
|
||||
@Test
|
||||
void rosterViewIncludesFailureReasonOnlyForBackendError() {
|
||||
MemberSession failed = new MemberSession("pane-error", "term-error", "prof", MemberRole.DEV,
|
||||
"/cwd", null, 0, 0, 1, MemberSession.State.BACKEND_ERROR, null, null,
|
||||
null, null, "backend exited");
|
||||
MemberSession ordinary = new MemberSession("pane-ready", "term-ready", "prof", MemberRole.DEV,
|
||||
"/cwd", null, 0, 0, 0, MemberSession.State.READY, null, null);
|
||||
|
||||
Map<String, Object> failedView = SessionManager.rosterView(failed, null);
|
||||
Map<String, Object> ordinaryView = SessionManager.rosterView(ordinary, null);
|
||||
|
||||
assertEquals("backend_error", failedView.get("state"));
|
||||
assertEquals("backend exited", failedView.get("failureReason"));
|
||||
assertFalse(ordinaryView.containsKey("failureReason"),
|
||||
"ordinary rows must not gain a blank failureReason field");
|
||||
}
|
||||
|
||||
@Test
|
||||
void onTurnFailedIsLoggedAtWarnWithThePriorState() {
|
||||
// CB-564: this transition used to be a bare DEBUG "session marked failed" — a symptom with no
|
||||
@@ -1004,6 +1149,76 @@ class SessionManagerTest {
|
||||
}
|
||||
}
|
||||
|
||||
/** Delegates all launcher work while recording context clears. */
|
||||
private static final class ClearContextSpyLauncher implements PeerLauncher {
|
||||
private final PeerLauncher delegate;
|
||||
private int clearContextCalls;
|
||||
|
||||
ClearContextSpyLauncher(PeerLauncher delegate) {
|
||||
this.delegate = delegate;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<Capability> capabilities() {
|
||||
return delegate.capabilities();
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<Capability> capabilitiesFor(String profileName) {
|
||||
return delegate.capabilitiesFor(profileName);
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req) {
|
||||
return delegate.spawn(req);
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return delegate.profiles();
|
||||
}
|
||||
|
||||
@Override
|
||||
public String defaultProfile() {
|
||||
return delegate.defaultProfile();
|
||||
}
|
||||
|
||||
@Override
|
||||
public String effectiveCwd(SpawnRequest req) {
|
||||
return delegate.effectiveCwd(req);
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<String> parityOverlay(String profileName) {
|
||||
return delegate.parityOverlay(profileName);
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<?> list() {
|
||||
return delegate.list();
|
||||
}
|
||||
|
||||
@Override
|
||||
public int reapOrphanWorkers() {
|
||||
return delegate.reapOrphanWorkers();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void stop(String id) {
|
||||
delegate.stop(id);
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean clearContext(String id) {
|
||||
clearContextCalls++;
|
||||
return delegate.clearContext(id);
|
||||
}
|
||||
|
||||
int clearContextCalls() {
|
||||
return clearContextCalls;
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void acquireRecordsTheAgentSessionIdFromTheHandle() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
@@ -140,7 +140,7 @@ class WorktreeSessionManagerTest {
|
||||
void worktreeAcquireRunsParityOverlayWithProfileDefaults() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
|
||||
.track(".envrc")
|
||||
.track(".env")
|
||||
.exists(".claude/settings.local.json");
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
|
||||
@@ -151,14 +151,15 @@ class WorktreeSessionManagerTest {
|
||||
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
|
||||
assertNotNull(overlay);
|
||||
assertEquals("/repo", overlay.repoRoot());
|
||||
assertEquals(List.of(".env", ".envrc"),
|
||||
overlay.requested(), "default parity overlay is used when unset");
|
||||
assertEquals(List.of(".env"),
|
||||
overlay.requested(), "default parity overlay is used when unset (CB-148: .env only, "
|
||||
+ ".envrc is no longer defaulted because it is executable shell direnv runs on cd)");
|
||||
assertFalse(overlay.requested().contains(".mcp.json"),
|
||||
"CB-525: replicating the primary's MCP config gives a worker the primary's IDE "
|
||||
+ "servers, which navigate its edits out of its own worktree");
|
||||
assertEquals(List.of(".envrc"), overlay.copied(),
|
||||
assertEquals(List.of(".env"), overlay.copied(),
|
||||
"existing paths are copied; missing paths are skipped");
|
||||
assertEquals(List.of(".envrc"), overlay.skipWorktree(),
|
||||
assertEquals(List.of(".env"), overlay.skipWorktree(),
|
||||
"tracked copied paths are --skip-worktree'd");
|
||||
}
|
||||
|
||||
|
||||
@@ -22,8 +22,34 @@
|
||||
# A prefix of a short secret is most of the secret, and it would end up pasted into a ticket. The
|
||||
# hash answers every question the prefix was for — is it set, is it the same value as over there,
|
||||
# is it the CB-592 sentinel — and answers none of the ones it should not.
|
||||
# * It never writes anywhere, never contacts the network, and never touches secrets.sh, which is
|
||||
# the operator's file.
|
||||
# * It never writes anywhere, never contacts the network except the daemon's own REST port (see
|
||||
# below), and never touches secrets.sh, which is the operator's file.
|
||||
#
|
||||
# WHERE THE NAME LIST COMES FROM (fleetd #111 / CB-608)
|
||||
#
|
||||
# Earlier versions of this script carried their own hardcoded NAMES array, recorded by hand on
|
||||
# 2026-08-16. The live memberCredentials: policy in fleetd.yaml grew past that list, and this probe
|
||||
# never noticed — it kept checking the same 31 names, printed a clean-looking table, and exited 0.
|
||||
# A verification tool that silently under-reports the thing it verifies is worse than no tool at
|
||||
# all, because its "clean" output gets taken as proof rather than treated with the suspicion an
|
||||
# absent tool would get.
|
||||
#
|
||||
# The fix is the same one #114 used for the drifted tool catalogue: delete the hand-maintained copy
|
||||
# rather than update it. This script now fetches the policy's name list from the daemon itself, at
|
||||
# `GET /member-credentials` (dev.ltms.fleet.member.MemberCredentialPolicyView via FleetApp) — names
|
||||
# and counts only, the same way the daemon's own startup log line is computed, from the SAME class.
|
||||
# If fleetd adds a name to memberCredentials.known tomorrow, this probe checks it tomorrow too,
|
||||
# with no edit here required. There is no local fallback list. See fetch_policy() below for what
|
||||
# happens when the daemon cannot be reached — it is a hard failure, on purpose (see next section).
|
||||
#
|
||||
# WHY AN UNREACHABLE DAEMON IS A HARD FAILURE, NOT A DEGRADED RUN
|
||||
#
|
||||
# An empty (or short) name list passes every subset check trivially — a probe that checked zero
|
||||
# names would print "0 of 0 names are set" and look identical to a clean bill of health. That trap
|
||||
# has bitten this project twice in one week (see docs/memory — "silent defaults disable features"
|
||||
# and "a test on the seam does not prove the caller"). So the denominator is guarded explicitly:
|
||||
# this script refuses to proceed unless it got a policy with at least one known name, and it refuses
|
||||
# just as hard if the count it fetched does not match the count it is about to check.
|
||||
#
|
||||
# HOW TO RUN IT
|
||||
#
|
||||
@@ -32,34 +58,22 @@
|
||||
# 2. For the comparison row, in your OWN shell — a lead, not a member:
|
||||
# bash scripts/probe-member-credentials.sh --allow-outside-member
|
||||
#
|
||||
# Both readings need the daemon's REST port reachable (default http://127.0.0.1:8765; override with
|
||||
# FLEETD_HOST). That is normally true in every pane this script is meant to run in.
|
||||
#
|
||||
# The two outputs side by side are the finding: any name whose hash matches between them is a
|
||||
# credential the member holds in full.
|
||||
#
|
||||
set -uo pipefail
|
||||
|
||||
# The names ${SHARED_ENV}/tools/secrets.sh exports, recorded on 2026-08-16 (issue #82). Names only —
|
||||
# this list contains no values and never should. If secrets.sh gains a name, this list goes stale and
|
||||
# the probe silently stops asking about it; that staleness is itself part of what #82's criterion 4
|
||||
# has to solve, so it is called out in the summary rather than hidden.
|
||||
NAMES=(
|
||||
AI_GATEWAY_TOKEN BESZEL_ADMIN_EMAIL BESZEL_ADMIN_PASSWORD
|
||||
BESZEL_HUB_URL BESZEL_KEY BESZEL_UNIVERSAL_TOKEN
|
||||
BRAIN_MCP_TOKEN CF_ACCOUNT_ID CF_API_TOKEN
|
||||
CF_USER_TOKEN CONFLUENCE_API_TOKEN CONFLUENCE_USERNAME
|
||||
CONTEXT7_TOKEN GITEA_HOST GITLAB_OAUTH_CLIENT_SECRET
|
||||
GITLAB_PERSONAL_ACCESS_TOKEN GRAFANA_ADMIN_PASSWORD GRAFANA_ADMIN_USER
|
||||
HASS_TOKEN HW_PASSWORD HW_USER
|
||||
LTMS_API_KEY MEMORY_MCP_TOKEN METRICS_PUSH_TOKEN
|
||||
OPENCODE_AUTOMODE_MODEL TELEGRAM_BOT_TOKEN TELEGRAM_CHAT_ID
|
||||
TS_API_KEY TS_AUTHKEY WORKER_GITEA_TOKEN
|
||||
GITEA_ACCESS_TOKEN
|
||||
)
|
||||
FLEETD_HOST="${FLEETD_HOST:-http://127.0.0.1:8765}"
|
||||
POLICY_URL="${FLEETD_HOST%/}/member-credentials"
|
||||
|
||||
allow_outside=0
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--allow-outside-member) allow_outside=1 ;;
|
||||
-h|--help) sed -n '2,40p' "$0"; exit 0 ;;
|
||||
-h|--help) sed -n '2,60p' "$0"; exit 0 ;;
|
||||
*) echo "unknown argument: $arg" >&2; exit 2 ;;
|
||||
esac
|
||||
done
|
||||
@@ -75,6 +89,107 @@ EOF
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# --- fetch the policy from the daemon (fleetd #111) — no local fallback, ever ------------------
|
||||
#
|
||||
# Prefer jq (a real JSON parser); fall back to python3 (present on every host this has run on so
|
||||
# far); if neither exists, fail loudly rather than guess at the JSON with grep/sed, which is exactly
|
||||
# the kind of "looks like it worked" degradation this ticket exists to remove.
|
||||
#
|
||||
# NOTE: jq's `//` alternative operator treats `false` AND `0` as "missing" and substitutes the
|
||||
# default — so `.present // empty` silently turns a real `"present": false` into an empty string
|
||||
# ("unknown"), not the false it actually is. Every extraction below reads its field directly
|
||||
# instead, so a genuine false/0 is reported as exactly that, not swallowed into "unknown".
|
||||
if ! command -v jq >/dev/null 2>&1 && ! command -v python3 >/dev/null 2>&1; then
|
||||
echo "refusing to run: neither jq nor python3 is on PATH, and this probe will not guess at JSON" \
|
||||
"with grep/sed. Install one of them, or run from a shell that has one." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
POLICY_JSON="$(curl -fsS --max-time 5 "$POLICY_URL" 2>/dev/null)"
|
||||
CURL_STATUS=$?
|
||||
if [ "$CURL_STATUS" -ne 0 ] || [ -z "$POLICY_JSON" ]; then
|
||||
cat >&2 <<EOF
|
||||
refusing to run: could not fetch the memberCredentials policy from $POLICY_URL (curl exit $CURL_STATUS).
|
||||
|
||||
This probe has NO built-in name list any more (fleetd #111) — it only checks what the live daemon
|
||||
reports, so an unreachable daemon means it cannot check anything at all. It will not fall back to a
|
||||
guessed or empty list, because an empty list would pass every check trivially and look clean.
|
||||
|
||||
Fix: confirm fleetd is up (curl \${FLEETD_HOST:-http://127.0.0.1:8765}/healthz) and that
|
||||
FLEETD_HOST (if set) points at it, then re-run.
|
||||
EOF
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# One parse pass: line 1 = present (true/false/null), line 2 = policy mode (possibly blank),
|
||||
# lines 3-5 = knownCount/allowedCount/blockedCount, remaining lines = the known[] names. A single
|
||||
# pass avoids re-parsing (and re-risking a truthiness bug) five separate times.
|
||||
if command -v jq >/dev/null 2>&1; then
|
||||
mapfile -t _FIELDS < <(printf '%s' "$POLICY_JSON" | jq -r '
|
||||
(.present | tostring),
|
||||
(.policy // ""),
|
||||
(.knownCount // 0 | tostring),
|
||||
(.allowedCount // 0 | tostring),
|
||||
(.blockedCount // 0 | tostring),
|
||||
(.known[]? // empty)')
|
||||
else
|
||||
mapfile -t _FIELDS < <(printf '%s' "$POLICY_JSON" | python3 - <<'PY'
|
||||
import json, sys
|
||||
data = json.load(sys.stdin)
|
||||
print(str(data.get("present")))
|
||||
print(data.get("policy") or "")
|
||||
print(data.get("knownCount") if data.get("knownCount") is not None else 0)
|
||||
print(data.get("allowedCount") if data.get("allowedCount") is not None else 0)
|
||||
print(data.get("blockedCount") if data.get("blockedCount") is not None else 0)
|
||||
for n in (data.get("known") or []):
|
||||
print(n)
|
||||
PY
|
||||
)
|
||||
fi
|
||||
|
||||
PRESENT="${_FIELDS[0]:-null}"
|
||||
POLICY_MODE="${_FIELDS[1]:-}"
|
||||
KNOWN_COUNT_REPORTED="${_FIELDS[2]:-0}"
|
||||
ALLOWED_COUNT_REPORTED="${_FIELDS[3]:-0}"
|
||||
BLOCKED_COUNT_REPORTED="${_FIELDS[4]:-0}"
|
||||
NAMES=("${_FIELDS[@]:5}")
|
||||
|
||||
# knownCount must be a plain non-negative integer for the arithmetic guard below — a malformed or
|
||||
# unparseable response must fail loudly, not be coerced into a number that happens to compare true.
|
||||
case "$KNOWN_COUNT_REPORTED" in
|
||||
''|*[!0-9]*)
|
||||
echo "refusing to run: knownCount in the response ('$KNOWN_COUNT_REPORTED') is not a plain" \
|
||||
"non-negative integer — the response could not be parsed as expected." >&2
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
|
||||
# --- guard the denominator explicitly — never proceed on a zero/short count ---------------------
|
||||
#
|
||||
# This is the exact trap named in the ticket: an empty (or truncated) NAMES array passes every
|
||||
# subsequent "is it set" check vacuously and prints a table that LOOKS complete. So this is checked
|
||||
# before anything else runs, with a message that says why, not just that it failed.
|
||||
if [ "${#NAMES[@]}" -eq 0 ] || [ "$KNOWN_COUNT_REPORTED" -eq 0 ]; then
|
||||
cat >&2 <<EOF
|
||||
refusing to run: the policy fetched from $POLICY_URL contains 0 known names (present=${PRESENT:-unknown}).
|
||||
|
||||
Either memberCredentials: is absent/empty on the running daemon (nothing is protected — see fleetd's
|
||||
own startup warning), or the response could not be parsed. Either way, checking zero names would
|
||||
print a clean-looking table for a policy that protects nothing, or for a probe that read nothing.
|
||||
This is refused rather than reported as a pass.
|
||||
EOF
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [ "${#NAMES[@]}" -ne "$KNOWN_COUNT_REPORTED" ]; then
|
||||
cat >&2 <<EOF
|
||||
refusing to run: the policy reports knownCount=$KNOWN_COUNT_REPORTED but the known[] array this probe
|
||||
parsed has ${#NAMES[@]} entries. That mismatch means the JSON was not parsed correctly, and this
|
||||
probe will not check a name list it cannot trust to be complete.
|
||||
EOF
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Prefer sha256sum (Linux), fall back to shasum (macOS). If neither exists, report presence and
|
||||
# length only — degraded, but never a value.
|
||||
hasher=""
|
||||
@@ -95,12 +210,13 @@ else
|
||||
where="NOT a member — comparison reading only"
|
||||
fi
|
||||
|
||||
echo "CB-596 credential probe"
|
||||
echo "CB-596 credential probe (fleetd #111: names sourced live from $POLICY_URL)"
|
||||
echo "reading from : $where"
|
||||
echo "shell : ${SHELL:-unknown}"
|
||||
echo "hash : ${hasher:-none available — lengths only}"
|
||||
# Only printed so the two readings can be told apart when they are pasted side by side.
|
||||
echo "host : $(hostname 2>/dev/null || echo unknown)"
|
||||
echo "policy : mode=${POLICY_MODE:-unknown} known=$KNOWN_COUNT_REPORTED allowed=${ALLOWED_COUNT_REPORTED:-?} blocked=${BLOCKED_COUNT_REPORTED:-?}"
|
||||
echo
|
||||
printf '%-30s %-7s %6s %s\n' "NAME" "STATE" "LEN" "SHA256-12"
|
||||
printf '%-30s %-7s %6s %s\n' "------------------------------" "-------" "------" "------------"
|
||||
@@ -118,6 +234,7 @@ done
|
||||
|
||||
echo
|
||||
echo "$set_count of ${#NAMES[@]} names are set in this shell."
|
||||
echo "policy contains $KNOWN_COUNT_REPORTED name(s); this run checked ${#NAMES[@]} — they match."
|
||||
echo
|
||||
cat <<'EOF'
|
||||
How to read this:
|
||||
@@ -129,7 +246,8 @@ How to read this:
|
||||
most urgent thing on this page.
|
||||
* AI_GATEWAY_TOKEN matching is expected and correct, not a leak: fleetd.yaml names it in
|
||||
`tokenEnv:` for the local and gx profiles, so a member reaching the gateway is by design.
|
||||
* A name that is set here but is NOT in the list above will not appear at all. The list was
|
||||
recorded on 2026-08-16 and does not update itself. Anything added to secrets.sh since then is
|
||||
invisible to this probe — which is the same gap issue #82 criterion 4 asks to close properly.
|
||||
* The name list above is fetched live from the running daemon's memberCredentials: policy
|
||||
(fleetd #111) — it is never hand-maintained here, so it cannot go stale the way the old
|
||||
hardcoded list did. If the daemon's policy changes, the next run of this script reflects it
|
||||
with no edit to this file.
|
||||
EOF
|
||||
|
||||
Reference in New Issue
Block a user