Files
fleetd/e2e/README.md
Dai Ha 5f0ec034d9 e2e: standard bridge conversation test harness
A repeatable multi-turn primary↔worker conversation driven entirely through the
bridge's loopback REST face (async fire-and-poll) — never sets ANTHROPIC_BASE_URL
and never touches herdr, so it is subscription-safe by construction. Records every
turn to a transcript, grades each (OK / DEGRADED / EMPTY / FAILED / WEDGE), and
exits non-zero if any turn fails to deliver-and-reply, so it is CI-usable.

This is the repeatable form of the ad-hoc channel test that surfaced the CB-115 and
CB-116 gaps.
2026-07-16 08:27:50 +02:00

78 lines
3.3 KiB
Markdown

# Bridge conversation test (`e2e/`)
A standard, repeatable **live** end-to-end test of the two-way channel: it drives a real
multi-turn conversation between a primary and an off-subscription worker **through the
running `bridged` daemon**, captures the full transcript, and grades the channel.
This is the committed form of the ad-hoc channel test that discovered the CB-115 gaps
(herdr `unknown` misclassification wedging delivery, dirty completion scrapes, and workers
never calling `bridge_reply` in conversation). Run it after any change to the injector,
status handling, completion/failure paths, or the worker reply charter.
## What it exercises
Each turn goes through the whole gateway exactly as a primary Opus session would — async
fire-and-poll (`POST /sessions/{id}/message {"wait":false}` → `GET /tasks/{ticket}`), so it
also validates the path that beats the caller's MCP timeout. It never sets
`ANTHROPIC_BASE_URL` and never talks to herdr directly, so it is **subscription-safe by
construction** — it only calls the bridge's loopback REST face.
```mermaid
sequenceDiagram
participant T as conversation_test.py
participant B as bridged (REST)
participant W as worker (off-sub)
T->>B: POST /workers (spawn)
T->>B: GET /sessions/{id}/status (await ready)
loop each turn
T->>B: POST /sessions/{id}/message {wait:false}
B-->>T: ticket
B->>W: inject prompt (status-gated)
W-->>B: bridge_reply
T->>B: GET /tasks/{ticket} (poll)
B-->>T: done + reply
end
T->>B: DELETE /workers/{pane} (stop)
```
## Prerequisites
- `bridged` is running (default REST on `http://127.0.0.1:8765`) with at least one worker
profile configured and its backend reachable.
- herdr is up (the daemon needs it).
- Python 3 (standard library only — no pip installs).
## Run
```bash
# spawn the default-profile worker, run the built-in 5-turn conversation, grade, clean up
python3 e2e/conversation_test.py
# pick a profile / reuse a live worker / use your own prompts
python3 e2e/conversation_test.py --profile ollama
python3 e2e/conversation_test.py --tid term_abc123 --keep-worker
python3 e2e/conversation_test.py --prompts my_prompts.txt --out /tmp/run1
```
A prompts file is one prompt per line; blank lines and `#` comments are ignored.
## Output & grading
- Writes `transcript.md` (in `--out`, default `e2e/`) — every turn's prompt, worker reply,
latency, resolution source, and observed status transitions, with an inline `> **GAP**`
note on any non-clean turn.
- Prints a per-turn line and an overall summary, and **exits non-zero** if any turn wedged,
failed, or returned empty — so it is CI-usable.
Per-turn grade:
| Grade | Meaning |
|------------|---------------------------------------------------------------------|
| `OK` | delivered and resolved by an explicit `bridge_reply` (`source=reply`) |
| `DEGRADED` | delivered and answered, but resolved via completion-scrape fallback |
| `EMPTY` | turn completed but the reply was empty |
| `FAILED` | the worker's turn ended in failure (`phase=failed`) |
| `WEDGE` | never resolved within the poll window (delivery wedge / lost turn) |
`PASS` requires every turn to be `OK` or `DEGRADED`; a clean run is every turn `OK`.