Files
Dai Ha 2e138a199b CB-634: one shared "fleet" workspace + rename bridged -> fleetd cutover
Two changes ship together here.

1. One shared herdr workspace. The lead and every worker now live in one
   workspace called "fleet", so the operator sees one "session" with many
   windows, not two. Before, the lead sat in a "leads" workspace and workers
   in "bridged-workers", which read as two sessions. The lead is still told
   apart from workers by its exact tab label ("lead: <name>"), so putting them
   in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
   for split layouts; Fleetd now passes an empty exclude set.

2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
   launchd/systemd units, module dir, and MCP mount).
   - Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
   - Log line, comments, docs, and CLAUDE.md updated to say fleetd.
   - Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
     bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
   - Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
     bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
   - Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
     still read as a fallback, and still gitignored.
   - MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
     server name is "fleet". The mount name in the local .mcp.json becomes
     "fleet" (gitignored, not in this commit).
   - Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
     BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.

Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.

Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.

The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).

949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
2026-08-25 04:01:08 +02:00

78 lines
3.3 KiB
Markdown

# Bridge conversation test (`e2e/`)
A standard, repeatable **live** end-to-end test of the two-way channel: it drives a real
multi-turn conversation between a primary and an off-subscription worker **through the
running `fleetd` daemon**, captures the full transcript, and grades the channel.
This is the committed form of the ad-hoc channel test that discovered the CB-115 gaps
(herdr `unknown` misclassification wedging delivery, dirty completion scrapes, and workers
never calling `fleet_reply` in conversation). Run it after any change to the injector,
status handling, completion/failure paths, or the worker reply charter.
## What it exercises
Each turn goes through the whole gateway exactly as a primary Opus session would — async
fire-and-poll (`POST /sessions/{id}/message {"wait":false}` → `GET /tasks/{ticket}`), so it
also validates the path that beats the caller's MCP timeout. It never sets
`ANTHROPIC_BASE_URL` and never talks to herdr directly, so it is **subscription-safe by
construction** — it only calls the bridge's loopback REST face.
```mermaid
sequenceDiagram
participant T as conversation_test.py
participant B as fleetd (REST)
participant W as worker (off-sub)
T->>B: POST /workers (spawn)
T->>B: GET /sessions/{id}/status (await ready)
loop each turn
T->>B: POST /sessions/{id}/message {wait:false}
B-->>T: ticket
B->>W: inject prompt (status-gated)
W-->>B: fleet_reply
T->>B: GET /tasks/{ticket} (poll)
B-->>T: done + reply
end
T->>B: DELETE /workers/{pane} (stop)
```
## Prerequisites
- `fleetd` is running (default REST on `http://127.0.0.1:8765`) with at least one worker
profile configured and its backend reachable.
- herdr is up (the daemon needs it).
- Python 3 (standard library only — no pip installs).
## Run
```bash
# spawn the default-profile worker, run the built-in 5-turn conversation, grade, clean up
python3 e2e/conversation_test.py
# pick a profile / reuse a live worker / use your own prompts
python3 e2e/conversation_test.py --profile ollama
python3 e2e/conversation_test.py --tid term_abc123 --keep-worker
python3 e2e/conversation_test.py --prompts my_prompts.txt --out /tmp/run1
```
A prompts file is one prompt per line; blank lines and `#` comments are ignored.
## Output & grading
- Writes `transcript.md` (in `--out`, default `e2e/`) — every turn's prompt, worker reply,
latency, resolution source, and observed status transitions, with an inline `> **GAP**`
note on any non-clean turn.
- Prints a per-turn line and an overall summary, and **exits non-zero** if any turn wedged,
failed, or returned empty — so it is CI-usable.
Per-turn grade:
| Grade | Meaning |
|------------|---------------------------------------------------------------------|
| `OK` | delivered and resolved by an explicit `fleet_reply` (`source=reply`) |
| `DEGRADED` | delivered and answered, but resolved via completion-scrape fallback |
| `EMPTY` | turn completed but the reply was empty |
| `FAILED` | the worker's turn ended in failure (`phase=failed`) |
| `WEDGE` | never resolved within the poll window (delivery wedge / lost turn) |
`PASS` requires every turn to be `OK` or `DEGRADED`; a clean run is every turn `OK`.