Files
fleetd/e2e/README.md
T
Dai Ha 76e4b577a0
CI / build (pull_request) Successful in 1m6s
CI / contract (pull_request) Successful in 1m6s
CB-622: rename bridge_* tool names to fleet_* in Markdown docs
Rename the eleven MCP tool names (bridge_ack/ask/list/poll/profiles/reply/
send/spawn/status/stop/whoami) to their fleet_* names across the Markdown
documentation. fleet_* is written as the normal name; one deprecation note
in README.md says bridge_* still works for one release.

docs/MCP-Contract.md is renamed only inside section 6 (lines 210-293):
sections 1-5 and 7-11 are stale pre-build design text (CB-609) and are
deliberately left with old names so dead text does not look maintained.

Also renames e2e/bridge_ask_transcript.md to e2e/fleet_ask_transcript.md
to match its content. CLAUDE.md, wiki/, plugin/skills/setup/SKILL.md and
.claude/skills/port-to-opencode/SKILL.md are owned by other units and are
untouched.
2026-08-22 21:52:24 +02:00

3.3 KiB

Bridge conversation test (e2e/)

A standard, repeatable live end-to-end test of the two-way channel: it drives a real multi-turn conversation between a primary and an off-subscription worker through the running bridged daemon, captures the full transcript, and grades the channel.

This is the committed form of the ad-hoc channel test that discovered the CB-115 gaps (herdr unknown misclassification wedging delivery, dirty completion scrapes, and workers never calling fleet_reply in conversation). Run it after any change to the injector, status handling, completion/failure paths, or the worker reply charter.

What it exercises

Each turn goes through the whole gateway exactly as a primary Opus session would — async fire-and-poll (POST /sessions/{id}/message {"wait":false} → GET /tasks/{ticket}), so it also validates the path that beats the caller's MCP timeout. It never sets ANTHROPIC_BASE_URL and never talks to herdr directly, so it is subscription-safe by construction — it only calls the bridge's loopback REST face.

sequenceDiagram
    participant T as conversation_test.py
    participant B as bridged (REST)
    participant W as worker (off-sub)
    T->>B: POST /workers  (spawn)
    T->>B: GET /sessions/{id}/status  (await ready)
    loop each turn
        T->>B: POST /sessions/{id}/message {wait:false}
        B-->>T: ticket
        B->>W: inject prompt (status-gated)
        W-->>B: fleet_reply
        T->>B: GET /tasks/{ticket}  (poll)
        B-->>T: done + reply
    end
    T->>B: DELETE /workers/{pane}  (stop)

Prerequisites

  • bridged is running (default REST on http://127.0.0.1:8765) with at least one worker profile configured and its backend reachable.
  • herdr is up (the daemon needs it).
  • Python 3 (standard library only — no pip installs).

Run

# spawn the default-profile worker, run the built-in 5-turn conversation, grade, clean up
python3 e2e/conversation_test.py

# pick a profile / reuse a live worker / use your own prompts
python3 e2e/conversation_test.py --profile ollama
python3 e2e/conversation_test.py --tid term_abc123 --keep-worker
python3 e2e/conversation_test.py --prompts my_prompts.txt --out /tmp/run1

A prompts file is one prompt per line; blank lines and # comments are ignored.

Output & grading

  • Writes transcript.md (in --out, default e2e/) — every turn's prompt, worker reply, latency, resolution source, and observed status transitions, with an inline > **GAP** note on any non-clean turn.
  • Prints a per-turn line and an overall summary, and exits non-zero if any turn wedged, failed, or returned empty — so it is CI-usable.

Per-turn grade:

Grade Meaning
OK delivered and resolved by an explicit fleet_reply (source=reply)
DEGRADED delivered and answered, but resolved via completion-scrape fallback
EMPTY turn completed but the reply was empty
FAILED the worker's turn ended in failure (phase=failed)
WEDGE never resolved within the poll window (delivery wedge / lost turn)

PASS requires every turn to be OK or DEGRADED; a clean run is every turn OK.