Files
fleetd/e2e
Dai Ha 3aea4e1ec6
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m30s
CB-622 lead follow-up: the three changes no worker could make
1. opencode.json — mount key bridged -> fleetd. Unit C could not do this:
   the file is neutralized by the worktree overlay, so every worker sees a
   stub and correctly reported it held no mount key. Tracked as CB-628.

2. e2e/bridge_ask_transcript.md keeps its name. It is a dated record of a
   run on 2026-07-16 that really did call bridge_ask, and its first line
   says so. The harness now writes fleet_ask_transcript.md for new runs;
   its header already says 'Live fleet_ask', so the two now agree.

3. docs/MCP-Contract.md line 20 — the historical banner describes section
   6, and section 6 now says fleet_ask.
2026-08-22 21:59:23 +02:00
..

Bridge conversation test (e2e/)

A standard, repeatable live end-to-end test of the two-way channel: it drives a real multi-turn conversation between a primary and an off-subscription worker through the running bridged daemon, captures the full transcript, and grades the channel.

This is the committed form of the ad-hoc channel test that discovered the CB-115 gaps (herdr unknown misclassification wedging delivery, dirty completion scrapes, and workers never calling fleet_reply in conversation). Run it after any change to the injector, status handling, completion/failure paths, or the worker reply charter.

What it exercises

Each turn goes through the whole gateway exactly as a primary Opus session would — async fire-and-poll (POST /sessions/{id}/message {"wait":false} → GET /tasks/{ticket}), so it also validates the path that beats the caller's MCP timeout. It never sets ANTHROPIC_BASE_URL and never talks to herdr directly, so it is subscription-safe by construction — it only calls the bridge's loopback REST face.

sequenceDiagram
    participant T as conversation_test.py
    participant B as bridged (REST)
    participant W as worker (off-sub)
    T->>B: POST /workers  (spawn)
    T->>B: GET /sessions/{id}/status  (await ready)
    loop each turn
        T->>B: POST /sessions/{id}/message {wait:false}
        B-->>T: ticket
        B->>W: inject prompt (status-gated)
        W-->>B: fleet_reply
        T->>B: GET /tasks/{ticket}  (poll)
        B-->>T: done + reply
    end
    T->>B: DELETE /workers/{pane}  (stop)

Prerequisites

  • bridged is running (default REST on http://127.0.0.1:8765) with at least one worker profile configured and its backend reachable.
  • herdr is up (the daemon needs it).
  • Python 3 (standard library only — no pip installs).

Run

# spawn the default-profile worker, run the built-in 5-turn conversation, grade, clean up
python3 e2e/conversation_test.py

# pick a profile / reuse a live worker / use your own prompts
python3 e2e/conversation_test.py --profile ollama
python3 e2e/conversation_test.py --tid term_abc123 --keep-worker
python3 e2e/conversation_test.py --prompts my_prompts.txt --out /tmp/run1

A prompts file is one prompt per line; blank lines and # comments are ignored.

Output & grading

  • Writes transcript.md (in --out, default e2e/) — every turn's prompt, worker reply, latency, resolution source, and observed status transitions, with an inline > **GAP** note on any non-clean turn.
  • Prints a per-turn line and an overall summary, and exits non-zero if any turn wedged, failed, or returned empty — so it is CI-usable.

Per-turn grade:

Grade Meaning
OK delivered and resolved by an explicit fleet_reply (source=reply)
DEGRADED delivered and answered, but resolved via completion-scrape fallback
EMPTY turn completed but the reply was empty
FAILED the worker's turn ended in failure (phase=failed)
WEDGE never resolved within the poll window (delivery wedge / lost turn)

PASS requires every turn to be OK or DEGRADED; a clean run is every turn OK.