2e138a199b
Two changes ship together here.
1. One shared herdr workspace. The lead and every worker now live in one
workspace called "fleet", so the operator sees one "session" with many
windows, not two. Before, the lead sat in a "leads" workspace and workers
in "bridged-workers", which read as two sessions. The lead is still told
apart from workers by its exact tab label ("lead: <name>"), so putting them
in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
for split layouts; Fleetd now passes an empty exclude set.
2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
launchd/systemd units, module dir, and MCP mount).
- Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
- Log line, comments, docs, and CLAUDE.md updated to say fleetd.
- Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
- Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
- Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
still read as a fallback, and still gitignored.
- MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
server name is "fleet". The mount name in the local .mcp.json becomes
"fleet" (gitignored, not in this commit).
- Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.
Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.
Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.
The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).
949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
78 lines
3.3 KiB
Markdown
78 lines
3.3 KiB
Markdown
# Bridge conversation test (`e2e/`)
|
|
|
|
A standard, repeatable **live** end-to-end test of the two-way channel: it drives a real
|
|
multi-turn conversation between a primary and an off-subscription worker **through the
|
|
running `fleetd` daemon**, captures the full transcript, and grades the channel.
|
|
|
|
This is the committed form of the ad-hoc channel test that discovered the CB-115 gaps
|
|
(herdr `unknown` misclassification wedging delivery, dirty completion scrapes, and workers
|
|
never calling `fleet_reply` in conversation). Run it after any change to the injector,
|
|
status handling, completion/failure paths, or the worker reply charter.
|
|
|
|
## What it exercises
|
|
|
|
Each turn goes through the whole gateway exactly as a primary Opus session would — async
|
|
fire-and-poll (`POST /sessions/{id}/message {"wait":false}` → `GET /tasks/{ticket}`), so it
|
|
also validates the path that beats the caller's MCP timeout. It never sets
|
|
`ANTHROPIC_BASE_URL` and never talks to herdr directly, so it is **subscription-safe by
|
|
construction** — it only calls the bridge's loopback REST face.
|
|
|
|
```mermaid
|
|
sequenceDiagram
|
|
participant T as conversation_test.py
|
|
participant B as fleetd (REST)
|
|
participant W as worker (off-sub)
|
|
T->>B: POST /workers (spawn)
|
|
T->>B: GET /sessions/{id}/status (await ready)
|
|
loop each turn
|
|
T->>B: POST /sessions/{id}/message {wait:false}
|
|
B-->>T: ticket
|
|
B->>W: inject prompt (status-gated)
|
|
W-->>B: fleet_reply
|
|
T->>B: GET /tasks/{ticket} (poll)
|
|
B-->>T: done + reply
|
|
end
|
|
T->>B: DELETE /workers/{pane} (stop)
|
|
```
|
|
|
|
## Prerequisites
|
|
|
|
- `fleetd` is running (default REST on `http://127.0.0.1:8765`) with at least one worker
|
|
profile configured and its backend reachable.
|
|
- herdr is up (the daemon needs it).
|
|
- Python 3 (standard library only — no pip installs).
|
|
|
|
## Run
|
|
|
|
```bash
|
|
# spawn the default-profile worker, run the built-in 5-turn conversation, grade, clean up
|
|
python3 e2e/conversation_test.py
|
|
|
|
# pick a profile / reuse a live worker / use your own prompts
|
|
python3 e2e/conversation_test.py --profile ollama
|
|
python3 e2e/conversation_test.py --tid term_abc123 --keep-worker
|
|
python3 e2e/conversation_test.py --prompts my_prompts.txt --out /tmp/run1
|
|
```
|
|
|
|
A prompts file is one prompt per line; blank lines and `#` comments are ignored.
|
|
|
|
## Output & grading
|
|
|
|
- Writes `transcript.md` (in `--out`, default `e2e/`) — every turn's prompt, worker reply,
|
|
latency, resolution source, and observed status transitions, with an inline `> **GAP**`
|
|
note on any non-clean turn.
|
|
- Prints a per-turn line and an overall summary, and **exits non-zero** if any turn wedged,
|
|
failed, or returned empty — so it is CI-usable.
|
|
|
|
Per-turn grade:
|
|
|
|
| Grade | Meaning |
|
|
|------------|---------------------------------------------------------------------|
|
|
| `OK` | delivered and resolved by an explicit `fleet_reply` (`source=reply`) |
|
|
| `DEGRADED` | delivered and answered, but resolved via completion-scrape fallback |
|
|
| `EMPTY` | turn completed but the reply was empty |
|
|
| `FAILED` | the worker's turn ended in failure (`phase=failed`) |
|
|
| `WEDGE` | never resolved within the poll window (delivery wedge / lost turn) |
|
|
|
|
`PASS` requires every turn to be `OK` or `DEGRADED`; a clean run is every turn `OK`.
|