Files
fleetd/e2e/README.md
Dai Ha 2e138a199b CB-634: one shared "fleet" workspace + rename bridged -> fleetd cutover
Two changes ship together here.

1. One shared herdr workspace. The lead and every worker now live in one
   workspace called "fleet", so the operator sees one "session" with many
   windows, not two. Before, the lead sat in a "leads" workspace and workers
   in "bridged-workers", which read as two sessions. The lead is still told
   apart from workers by its exact tab label ("lead: <name>"), so putting them
   in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
   for split layouts; Fleetd now passes an empty exclude set.

2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
   launchd/systemd units, module dir, and MCP mount).
   - Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
   - Log line, comments, docs, and CLAUDE.md updated to say fleetd.
   - Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
     bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
   - Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
     bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
   - Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
     still read as a fallback, and still gitignored.
   - MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
     server name is "fleet". The mount name in the local .mcp.json becomes
     "fleet" (gitignored, not in this commit).
   - Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
     BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.

Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.

Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.

The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).

949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
2026-08-25 04:01:08 +02:00

3.3 KiB

Bridge conversation test (e2e/)

A standard, repeatable live end-to-end test of the two-way channel: it drives a real multi-turn conversation between a primary and an off-subscription worker through the running fleetd daemon, captures the full transcript, and grades the channel.

This is the committed form of the ad-hoc channel test that discovered the CB-115 gaps (herdr unknown misclassification wedging delivery, dirty completion scrapes, and workers never calling fleet_reply in conversation). Run it after any change to the injector, status handling, completion/failure paths, or the worker reply charter.

What it exercises

Each turn goes through the whole gateway exactly as a primary Opus session would — async fire-and-poll (POST /sessions/{id}/message {"wait":false} → GET /tasks/{ticket}), so it also validates the path that beats the caller's MCP timeout. It never sets ANTHROPIC_BASE_URL and never talks to herdr directly, so it is subscription-safe by construction — it only calls the bridge's loopback REST face.

sequenceDiagram
    participant T as conversation_test.py
    participant B as fleetd (REST)
    participant W as worker (off-sub)
    T->>B: POST /workers  (spawn)
    T->>B: GET /sessions/{id}/status  (await ready)
    loop each turn
        T->>B: POST /sessions/{id}/message {wait:false}
        B-->>T: ticket
        B->>W: inject prompt (status-gated)
        W-->>B: fleet_reply
        T->>B: GET /tasks/{ticket}  (poll)
        B-->>T: done + reply
    end
    T->>B: DELETE /workers/{pane}  (stop)

Prerequisites

  • fleetd is running (default REST on http://127.0.0.1:8765) with at least one worker profile configured and its backend reachable.
  • herdr is up (the daemon needs it).
  • Python 3 (standard library only — no pip installs).

Run

# spawn the default-profile worker, run the built-in 5-turn conversation, grade, clean up
python3 e2e/conversation_test.py

# pick a profile / reuse a live worker / use your own prompts
python3 e2e/conversation_test.py --profile ollama
python3 e2e/conversation_test.py --tid term_abc123 --keep-worker
python3 e2e/conversation_test.py --prompts my_prompts.txt --out /tmp/run1

A prompts file is one prompt per line; blank lines and # comments are ignored.

Output & grading

  • Writes transcript.md (in --out, default e2e/) — every turn's prompt, worker reply, latency, resolution source, and observed status transitions, with an inline > **GAP** note on any non-clean turn.
  • Prints a per-turn line and an overall summary, and exits non-zero if any turn wedged, failed, or returned empty — so it is CI-usable.

Per-turn grade:

Grade Meaning
OK delivered and resolved by an explicit fleet_reply (source=reply)
DEGRADED delivered and answered, but resolved via completion-scrape fallback
EMPTY turn completed but the reply was empty
FAILED the worker's turn ended in failure (phase=failed)
WEDGE never resolved within the poll window (delivery wedge / lost turn)

PASS requires every turn to be OK or DEGRADED; a clean run is every turn OK.