6
10 Cross Host Messaging
Dai Ha edited this page 2026-08-31 10:12:16 +07:00
This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

10. Cross-Host Messaging & Broker Topology

Scope. This chapter is the broker design for cross-host operation — the CB-308 federation and CB-500 multi-tier fabric — built on the as-built single-host reply inbox shipped by CB-307 (msg.AmqpReplyInbox, verified at 9-Implementation). It answers three questions: what are the cross-host communication use cases, how the AMQP broker is laid out, and which exchange + queue belongs to each entity. Where a piece exists today it is marked [as-built]; the rest is [proposed] (§9 draws the line).

One piece of cross-host traffic already ships, and it is not in the design below. Lead-to-lead coordination between daemons on different hosts works today, over a shared AMQP vhost configured by a coordinator: block. A lead sends to a peer with fleet_send{coordId}, and its own coord-id is reported by fleet_list. The code is msg/LeadMailbox and msg/LeadCoordLoop, wired in Fleetd.java:405 and Fleetd.java:546-551; the config record is FleetConfig.java:74-78 and 99-101.

That path deliberately does not use the exchange topology in §3. It is a direct durable mailbox per lead, not a federated fabric, and it carries coordination only — leads divide the map and share findings, they never assign each other work. Read §3 onward as the design for agent traffic across hosts, which is still proposed.

The one rule that shapes everything: the broker moves messages and presence, never keystrokes. Delivery into a peer is always a local herdr injection (an agent) or a local MCP pull (a primary/main) done by the gateway co-located with that peer. The broker only carries the middle hop between gateways. This is the same asymmetry CB-307 closes on one host, stretched across hosts by CB-308 — see 1. Architecture and CB-500 §11.

1. The cross-host communication use cases

Every cross-host interaction is one of these. Each is a message flow between two entities (a primary/main, an orchestrator, or a worker/sandboxed agent) that may live on different hosts.

# Use case Flow Kind
U1 Delegate + reply — a main on host A tasks a worker/sandboxed agent on host B, awaits its answer A → B, then B → A SEND → REPLY
U2 Mid-turn ask — a worker on B pauses its turn to ask the main on A, resumes on the answer B → A, then A → B (turn-scoped) ASK → ANSWER
U3 Stranded / late reply — B replies after A's blocking send timed out (or was never open); held durably until A pulls B → A (held) REPLY (deferred)
U4 Cross-host spawn — A requests a new agent on host B; gateway B launches it locally and announces it A → gateway B (control) SPAWN
U5 Roster / presence — every gateway announces its local agents so all build one union "who/where/status" view each gateway → all PRESENCE
U6 Lead ↔ architect — the human-driven lead engages either independent advisory architect L ↔ A SEND/REPLY
U7 Parallel architect advice — the lead sends the same brief to Claude Sonnet 5 and GPT-5.6/opencode, then compares independent results L → A1, A2 SEND/REPLY
U8 Broadcast / group announce — a lead publishes once to all workers (or a named group); the broker copies the message into every bound inbox L → all BROADCAST

U1–U3 are the CB-307 rendezvous semantics, now spanning hosts. U4–U5 are the net-new federation control plane. U6–U7 reuse the same per-entity inbox as U1 — a lead and an architect are just message-addressable entities with their own inbox. The two architect model families are deliberate: independent agreement is evidence rather than correlated echo. U8 is for identical announcements ("everyone: stop", "everyone: report status") — tailored task briefs keep the per-worker fan-out pattern (Team), which works unchanged across hosts because each send routes to its recipient's inbox wherever it lives.

2. Entities & identity

flowchart LR
    subgraph e["Message-addressable entities (each owns ONE inbox)"]
        arch["architect<br/>(MCP client)"]
        main["human-driven lead<br/>(MCP client)"]
        wrk["worker / sandboxed agent<br/>(herdr pane)"]
    end
    gw["gateway = fleetd<br/>(one per host)"]
    arch -->|"delivered by local MCP pull"| gw
    main -->|"delivered by local MCP pull"| gw
    wrk -->|"delivered by local herdr inject"| gw
    gw -->|"consumes its local entities' inboxes,<br/>publishes to remote inboxes"| broker["BROKER (AMQP)"]
    classDef bus fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
    class broker bus

Figure 1 — the gateway is the only AMQP client. Entities never touch the broker: an agent is injected into over local herdr, a main/orchestrator pulls over local MCP. The gateway consumes the inboxes of its co-located entities and publishes to the inboxes of remote ones.

Global id. The routing key is the entity's global id — a UUID minted at spawn (CB-308), not the herdr paneId (which is host-local and meaningless off-host). Host is directory metadata, carried in the roster, never in the routing key. [as-built: keyed by paneId, single host] → [proposed: keyed by globalId].

2.1 The message envelope [proposed]

Single-host, the envelope was deliberately descoped (CB-201): connection identity answers who is talking on every call. A broker hop has no connection identity, so cross-host the envelope returns — minimal, and carried in AMQP headers, not a JSON wrapper:

Field Meaning
v envelope schema version (§10.3)
kind SEND · REPLY · ASK · ANSWER · TOO_LATE · NO_WAITER · ABANDONED · FAILED · INJECTED · BROADCAST · PRESENCE · SPAWN
msgId unique per message — the dedup key
traceId minted at the flow's first send, carried by every hop (§10.4)
from / to sender / recipient global ids
turnId conversation flows only (ASK / ANSWER / TOO_LATE)
spawnId SPAWN only — doubles as the new worker's gid (CB-308 §7.8)
ts / expiresAt publish time / gateway-enforced expiry — on every kind (also bounds replay)
sig signature by the publishing gateway

The body stays opaque bytes (UTF-8 text today). sig covers the exact body bytes plus the canonical header subset above — signing headers-plus-body keeps JSON canonicalization out of the security path entirely. Verification at the receiving gateway, in order: signature valid → the claimed from lives on the signing host per signed presence (§8 step 5) → to matches the routing key (else a captured message could be replayed into a different inbox). Any check fails → basicReject(requeue=false) → fleet.dlx, loud log, never a wedge. Replay is bounded by expiresAt plus a bounded seen-set — the at-most-once forward path (§6) has no broker redelivery for a replayed duplicate to hide behind.

3. Broker topology — exchanges

Five exchanges. Point-to-point messaging, presence, and control are separated so each has its own durability and fan-out semantics.

Exchange Type Durable Purpose Status
fleet.msg topic yes All entity→entity messages (U1–U3, U6, U7) and broadcast (U8). Routing key = recipient globalId, or broadcast.all / broadcast.<group> for U8. [proposed]*
fleet.roster topic yes Presence heartbeats (U5). Routing key = roster.<host>.<globalId>. [proposed]
fleet.control direct yes Cross-host spawn/stop/lifecycle (U4). Routing key = target host id. [proposed]
fleet.dlx fanout yes Dead-letter sink for poison messages off any inbox. [proposed]
fleet.delay x-delayed-message yes Optional broker-driven remind/backoff — re-publishes into fleet.msg after a delay. Native on LavinMQ; a plugin on RabbitMQ — stays optional so "RabbitMQ by URI swap" holds. [proposed]

* [as-built, with a semantic migration] CB-307 publishes to the default exchange ("") with routing key = queue name — but note what that queue means today: agent.<target>.inbox is keyed by the worker a primary-bound reply came from (sender-keyed), drained per-worker. This chapter keys the inbox by recipient (invariant 2). Same name shape, opposite meaning — so CB-308 is a migration, not a rename. Migrated queues take a version-suffixed name (agent.<gid>.inbox.v2): AMQP refuses to redeclare an existing durable queue with new arguments (PRECONDITION_FAILED — a crash loop on an in-place upgrade from v1.0.0), and the suffix keeps old sender-keyed and new recipient-keyed queues apart while both exist. The per-worker drain surface (fleet_poll(target)) survives by filtering on the envelope's from field (§2.1). This is also where CB-201's story completes: descoping the envelope was right on one host (connection identity routes everything) and wrong across hosts (a broker hop has none) — the envelope returns as §2.1.

flowchart TB
    subgraph exch["Exchanges"]
        msg["fleet.msg<br/>(topic)"]
        ros["fleet.roster<br/>(topic)"]
        ctl["fleet.control<br/>(direct)"]
        dlx["fleet.dlx<br/>(fanout)"]
    end
    subgraph gwA["gateway A (host A)"]
        inA["agent.&lt;gidA&gt;.inbox<br/>(durable)"]
        ctlA["control.hostA<br/>(durable)"]
        rosA["roster.hostA<br/>(exclusive, transient)"]
    end
    subgraph gwB["gateway B (host B)"]
        inB["agent.&lt;gidB&gt;.inbox<br/>(durable)"]
        ctlB["control.hostB<br/>(durable)"]
        rosB["roster.hostB<br/>(exclusive, transient)"]
    end
    dlq["fleet.dlq<br/>(durable)"]
    msg -->|"key = gidA"| inA
    msg -->|"key = gidB"| inB
    ctl -->|"key = hostA"| ctlA
    ctl -->|"key = hostB"| ctlB
    ros -->|"key roster.#"| rosA
    ros -->|"key roster.#"| rosB
    inA -.->|"poison"| dlx
    inB -.->|"poison"| dlx
    dlx --> dlq
    classDef q fill:#2f855a,stroke:#22543d,color:#ffffff;
    classDef dead fill:#9b2c2c,stroke:#742a2a,color:#ffffff;
    class inA,inB,ctlA,ctlB,rosA,rosB q
    class dlq,dlx dead

Figure 2 — exchanges (top) route to per-entity and per-gateway queues (green). Inbox queues dead-letter poison messages to fleet.dlx → fleet.dlq (red). Each gateway consumes only the queues in its own box.

4. Queues per entity

The core rule: one inbox per message-addressable entity, consumed by exactly one gateway — the one co-located with that entity. Gateways additionally own a control queue and a roster queue.

Owner (entity/scope) Queue Bound to (exchange · key) Sole consumer Durability Notes
Worker / sandboxed agent gid agent.<gid>.inbox fleet.msg · <gid> + broadcast.all / broadcast.<group> (U8) the agent's local gateway → herdr inject durable [proposed] — recipient-keyed .v2 migration of CB-307's sender-keyed queue (§3 footnote); survives idle and daemon bounce; broadcast = extra bindings on the same queue
Primary / main gid agent.<gid>.inbox fleet.msg · <gid> its local gateway → MCP pull (resolve open send, else hold + nudge) durable identical pattern — a main is just an entity with an inbox (U6/U7); per-worker drain filters on envelope from
Orchestrator gid agent.<gid>.inbox fleet.msg · <gid> its local gateway → MCP pull durable top MCP client; same fabric
Gateway (host) control.<host> fleet.control · <host> that gateway durable cross-host spawn/stop (U4)
Gateway (host) roster.<host> fleet.roster · roster.# that gateway transient (exclusive, auto-delete) presence is soft-state — union view, rebuilt from heartbeats
Fleet (shared) fleet.dlq fleet.dlx ops / redelivery tooling durable poison messages after N redeliveries

Why the roster queue is transient while inboxes are durable: the broker owns message durability, not who/where/status. Presence is rebuilt from heartbeats on reconnect; a missed heartbeat expires an entry (CB-303 TTL thinking). This preserves the persistence boundary — fleetd stays soft-state; only messages are durable. See 1. Architecture.

5. Invariants

  1. Single consumer per inbox — broker-enforced. Only the co-located gateway consumes agent.<gid>.inbox → per-recipient FIFO ordering, and no double-injection. Consumers are opened with AMQP's exclusive flag, so a misconfigured second gateway fails loudly at connect time instead of silently splitting the stream (two consumers on one queue round-robin — each sees half). Fan-out to many recipients is always copies into many inboxes (U8), never two consumers on one queue.
  2. Publish by recipient, not by host. A sender publishes to fleet.msg with key = recipient globalId; the broker routes to whichever gateway holds that inbox. Senders are oblivious to the recipient's host — the roster resolves existence, the broker resolves location.
  3. The final hop is a pull for MCP clients. For a primary/main/orchestrator the broker makes the middle hop lossless + ordered + idempotent, but the last hop is still fleet_poll / drain — the gateway holds the reply until the client pulls (consume-and-hold). The broker does not dissolve the MCP asymmetry (§ CB-308 "the one thing the broker does NOT dissolve").
  4. Keystrokes never traverse the broker. Only messages + presence. Injection into an agent is a local herdr write by its owning gateway (CB-500 §11).
  5. Two ack models — matched to what a duplicate costs. Reply/pull path (into a main): at-least-once — ack after drain; a redelivered reply is benign and deduped by msgId. Forward path (a task brief into a worker): at-most-once — ack before the local inject; a crash in the window loses the inject, which surfaces as a visible failure at the sender (loud, recoverable) instead of a silent duplicate injection (quiet, dangerous). The worker's gateway publishes an INJECTED confirmation once the inject lands; no INJECTED within a bound = fast explicit failure, not a send-timeout-long blind window. Poison messages are explicitly rejected (basicReject(requeue=false)) → fleet.dlx — nothing is ever silently requeued forever.
  6. Signed messages. Every cross-host message is signed by its publishing gateway; the receiver verifies the signature and that the claimed sender lives on the signing host (roster check). This extends the single-host rule — identity comes from the connection, never an argument — across the broker: cross-host, identity comes from the key. Decision record: docs/CB-308-Multi-Host-Federation.md §7.

6. Message lifecycle & ack semantics [as-built for the reply path; forward path proposed]

Two lifecycles, split by what a duplicate would cost (invariant 5). The consumer runs with a basicQos prefetch bound, so the backlog stays on the queue — where the caps and expiry of §10.1 can act — instead of draining into gateway heap.

sequenceDiagram
    autonumber
    participant SND as sender gateway
    participant Q as agent.GID.inbox.v2
    participant OWN as owning gateway
    participant E as entity (agent/main)
    rect rgb(235, 244, 255)
    Note over SND,E: REPLY / pull path — at-least-once (as-built shape)
    SND->>Q: publish (confirms + mandatory, §8 step 6)
    OWN->>Q: consume (prefetch-bounded, dedup by msgId)
    OWN->>E: hold for MCP pull
    E-->>OWN: drained
    OWN->>Q: basicAck — a bounce BEFORE ack → redelivered, deduped
    end
    rect rgb(255, 245, 235)
    Note over SND,E: FORWARD path (brief into a worker) — at-most-once
    SND->>Q: publish (confirms + mandatory)
    OWN->>Q: consume
    OWN->>Q: basicAck FIRST — a redelivered brief can never double-inject
    OWN->>E: status-gated herdr inject
    OWN->>SND: publish INJECTED (correlated to msgId)
    Note over SND: no INJECTED within bound → loud, fast failure (retry is the sender's call)
    end

Figure 3 — the reply path keeps CB-307's consume-and-hold with deferred ack (AmqpReplyInbox today, 9-Implementation); the forward path inverts the ack order so a crash costs a visible loss, never a silent duplicate injection. INJECTED closes the blind window between ack and inject.

7. Use-case walkthroughs

7.1 U1 — cross-host delegate + reply

sequenceDiagram
    autonumber
    participant MA as main (host A)
    participant GA as gateway A
    participant BR as fleet.msg
    participant GB as gateway B
    participant W as worker (host B)
    MA->>GA: fleet_send(gidW, content)
    GA->>GA: roster - gidW local? NO
    GA->>BR: publish(key = gidW)
    BR->>GB: route to agent.gidW.inbox
    GB->>W: inject via B's LOCAL herdr
    W-->>GB: fleet_reply(to gidMA)
    GB->>BR: publish(key = gidMA, durable)
    BR->>GA: route to agent.gidMA.inbox
    Note over GA: held until MA pulls (MA is an MCP client)
    MA->>GA: blocking send resolves / fleet_poll
    GA-->>MA: reply

Figure 4 — U1. Both injection points (into W on B, and the drain into MA on A) are local; only the two middle hops cross the broker. U3 (stranded reply) is the same picture where step 8's "held" outlives A's send window and MA collects it later by fleet_poll(gidW) / drain.

7.2 U4 — cross-host spawn (the control plane)

sequenceDiagram
    autonumber
    participant MA as main (host A)
    participant GA as gateway A
    participant CX as fleet.control
    participant GB as gateway B
    participant SL as SandboxLauncher (B)
    MA->>GA: fleet_spawn(profile, host = B)
    GA->>CX: publish(key = hostB, SpawnRequest)
    CX->>GB: route to control.hostB
    GB->>SL: spawn locally (into a local sandbox)
    Note over SL: CB-306 readiness gate — wait until injectable
    SL-->>GB: PeerHandle(new gid)
    GB->>GB: announce on fleet.roster (roster.hostB.newGid)
    Note over MA: gid appears in every gateway's union roster,<br/>then U1 addresses it normally

Figure 5 — U4. Spawn is a control message to host B's gateway, which runs the launcher locally (CB-500 §11 — a sandbox is spawned by its own host's gateway). The readiness gate is more valuable here: the far side wants a positive "ready" before anyone sends. The new agent enters the roster and is then reachable by U1.

7.3 U5 — presence federation

flowchart LR
    subgraph gwA["gateway A"]
        hbA["heartbeat local agents<br/>roster.hostA.*"]
        viewA["union rosterView"]
    end
    subgraph gwB["gateway B"]
        hbB["heartbeat local agents<br/>roster.hostB.*"]
        viewB["union rosterView"]
    end
    ros["fleet.roster (topic)"]
    hbA -->|"publish"| ros
    hbB -->|"publish"| ros
    ros -->|"roster.# → roster.hostA"| viewA
    ros -->|"roster.# → roster.hostB"| viewB
    classDef bus fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
    class ros bus

Figure 6 — U5. Each gateway publishes heartbeats for its local agents and binds roster.# to see everyone's; both converge on the same eventually-consistent union view (CB-304 rosterView, federated). Transient queues + heartbeat TTL keep it soft-state.

7.4 U2 — cross-host ask (live-only, with expiry)

Asks are the one flow that is deliberately not durable — the single-host rule kept: only terminal replies are queued, a live conversation is not. Cross-host, ASK/ANSWER traverse fleet.msg carrying the asking turn's turn_id and an expiresAt that the gateways enforce at delivery time — a broker per-message TTL cannot expire a message already consumed into the held map, so broker TTL is only a backstop for a queue nobody is consuming. Two arrival checks preserve the single-host semantics: gateway A checks for an open waiter on arrival and publishes NO_WAITER straight back if there is none (the worker learns "nobody listening" in one round trip, as it does synchronously today), and gateway B injects an ANSWER only if that turn is still waiting:

sequenceDiagram
    autonumber
    participant W as worker (host B, parked mid-turn)
    participant GB as gateway B
    participant BR as fleet.msg
    participant GA as gateway A
    participant MA as main (host A)
    W->>GB: fleet_ask(question)
    GB->>BR: publish ASK (key = gidMA, expiresAt, turn_id)
    BR->>GA: route
    GA->>MA: surfaces on the open send / drain (local pull)
    MA->>GA: answer on turn_id
    GA->>BR: publish ANSWER (key = gidW, expiresAt, turn_id)
    BR->>GB: route
    alt turn still waiting
        GB->>W: inject answer — same turn resumes
    else turn gone (timeout / completed / recycled)
        GB--xW: NOT injected
        GB->>BR: publish TOO_LATE notice (key = gidMA)
    end

Figure 7 — U2 across hosts. The one failure case — an answer outliving its turn — is loud, not weird: the stale answer is dropped (never injected into an unrelated turn) and the main is told.

7.5 U8 — broadcast

The main publishes once with key broadcast.all (or broadcast.<group>); the broker copies the message into every inbox bound to that key, and each copy then follows the normal local delivery rules (Figure 3) at its own gateway. A worker that joins later starts receiving broadcasts the moment its inbox binds the key. Replies to a broadcast are just N ordinary U1 replies. Use it for identical announcements; per-worker briefs stay on the N-send fan-out (Team).

8. Setup checklist (per gateway)

On boot, a gateway declares and wires exactly its own slice:

  1. Connect to the broker URI over TLS (amqps://…), with this gateway's own broker login. LavinMQ default; RabbitMQ by URI swap. [as-built for plain amqp://; TLS + per-gateway login proposed]
  2. Declare the shared exchanges fleet.msg, fleet.roster, fleet.control, fleet.dlx (idempotent). [proposed]
  3. For each local entity: declare agent.<gid>.inbox.v2 (durable, DLX = fleet.dlx, capped: x-max-length + message TTL, x-expires as the orphan-GC backstop), bind to fleet.msg keys <gid> and broadcast.all (+ its groups), start a manual-ack, exclusive consumer with a basicQos prefetch bound → inject/hold. The .v2 suffix is the migration seam: AMQP refuses to redeclare an existing durable queue with new arguments (PRECONDITION_FAILED). [proposed; the unsuffixed, uncapped queue is as-built]
  4. Declare control.<thisHost> (durable), bind to fleet.control key <thisHost>, consume → run the launcher locally. [proposed]
  5. Declare roster.<thisHost> (exclusive, auto-delete), bind to fleet.roster key roster.#, consume → maintain the union view; start heartbeating local agents and a host-level entry (roster.<thisHost>._gateway, carrying this gateway's profile list — an agentless host must still be visible as a spawn target). All presence messages are signed: invariant 6's roster check rests on them. [proposed]
  6. Open a separate publish channel with publisher confirms, the mandatory flag, and a return listener — an unroutable publish is a fast error, not a black hole, and confirms never serialize the consume/ack channel. Ordering caveat: a return arrives before the confirm, so "confirmed" ≠ "routed" — check the returned-set at confirm time. mandatory is false only for BROADCAST (an empty group is legal silence). [proposed]
  7. On session stop / reap: delete the worker's inbox queue — its broadcast.* bindings die with it, so no broadcasts to the dead; x-expires collects queues orphaned by a crashed gateway. [proposed]
  8. Automatic connection + topology recovery re-declares queues and re-attaches consumers after a broker blip (dedup by msgId prevents double-queue). [as-built] Exclusive-consumer caveat: after a gateway crash the broker holds the stale lock until its heartbeat times out — keep the broker heartbeat short (~10s) so takeover is quick. [proposed]

9. As-built (CB-307) vs proposed (CB-308 / CB-500)

Piece Today [as-built] Cross-host [proposed]
Per-entity inbox agent.<target>.inbox durable, manual-ack consume-and-hold, dedup by msgId, auto-recovery — on the default exchange, keyed by the sender of a primary-bound reply recipient-keyed agent.<gid>.inbox.v2 on fleet.msg — a semantic migration (§3 footnote); per-worker drain preserved via envelope from
Consumer model one daemon consumes all inboxes one gateway per host, sole consumer of its local entities' inboxes
Presence in-process roster (CB-304) fleet.roster + roster.<host> transient queues → federated union
Spawn local call in-process fleet.control + control.<host> cross-host control message
Poison handling none fleet.dlx → fleet.dlq
Remind/backoff in-JVM ReplyPushLoop (CB-307 Stage 3), keyed by target rekeyed to envelope from once the inbox flips; optionally broker-driven via fleet.delay
Identity in key paneId globalId (UUID); host in roster metadata
Broadcast none (N separate sends) broadcast.* bindings on every inbox (U8) — publish once, the broker copies
Sender authenticity connection identity (loopback) per-gateway message signing + roster host check
Inbox bounds unbounded x-max-length + TTL → fleet.dlq, with a metric
Consumer enforcement convention AMQP exclusive consumers — a second gateway fails loudly
Multi-primary (U6/U7) single-slot PrimaryRegistry, one nudge target CB-500's delta, not CB-308's — N message-addressable MCP clients per gateway, per-client nudge targets

Lead-to-lead is the exception, and it already shipped. The table above is about agent traffic. Cross-host lead-to-lead coordination is real today and took a different route: a durable mailbox per lead on a shared vhost, addressed by coordId, with no exchange topology, no presence plane and no global id. It solved a smaller problem, so it needed a smaller mechanism.

Piece Lead-to-lead today [as-built]
Address coordId, from the coordinator: block; a lead's own is in fleet_list
Transport one shared AMQP vhost, separate from the per-fleet vhost
Send fleet_send{coordId} — mutually exclusive with sessionId and turnId
Receive msg/LeadCoordLoop, status-gated like any other delivery
Carries coordination only — never a task, never a brief

Net: the inbox half is real and already multi-host-ready on a shared broker, and lead-to-lead runs across hosts today. What cross-host agent traffic still needs is a presence plane, a control plane, dead-lettering, and a global id — with no change to the delivery asymmetry or the consume-and-hold contract.

10. Hardening rules (design review 2026-08-10) [proposed]

Cross-cutting rules settled in a design review of this chapter + CB-308. The sender-side decisions (message signing, profile ownership, repo provisioning, ask expiry, spawn dedup) are recorded in docs/CB-308-Multi-Host-Federation.md §7; the broker-level rules live here:

  1. Inbox caps — made real by prefetch. Every inbox carries x-max-length and a message TTL, and the consumer runs with a basicQos prefetch bound — without it, consume-and-hold drains the queue into gateway heap and the caps guard an empty queue (an as-built gap, ticketed). Overflow and expiry dead-letter to fleet.dlq (never vanish) and a metric fires. A gateway sweeper acks held messages past their expiresAt, so one abandoned sender's backlog cannot wedge a shared inbox behind a full prefetch window (head-of-line). A DLQ consumer reports every dead-lettered forward brief back to its sender as FAILED — the caps must not become a new silent loss channel — and fleet.dlq itself is capped (anything with broker write access could otherwise fill it).
  2. Private broker, TLS. Gateways connect over amqps:// with per-gateway logins. Written rule: the broker's disk holds readable task text (the audit log deliberately does not) — so the broker runs on a machine inside the trust circle, never a shared or rented one.
  3. Schema version. Every message carries a schema version. Within a major version, unknown fields are ignored — rolling upgrades across mixed-version gateways just work. A newer-major message is refused to the DLQ with a loud log line, never half-parsed.
  4. Trace id. The first send of a flow mints a trace id; every subsequent hop (send, ask, answer, reply, spawn) carries it, and every log/audit line prints it. Debugging a three-host flow = search one id on each host.
  5. Broker outage. Same-host traffic never touches the broker (the routing fork) and keeps working with the broker down — a written promise. A remote send while the broker is unreachable fails fast with a clear error; the gateway never buffers on the broker's behalf (it stays soft-state, so a crash cannot lose messages it claimed to deliver). Auto-reconnect re-declares the topology; remote hosts read as unknown in the roster meanwhile.
  • 1. Architecture — the two modes, the invariants, the MCP asymmetry.
  • 8. Roadmap — CB-307 (delivery), CB-308 (federation), CB-500 (multi-tier) staging.
  • 9. Implementation — msg.AmqpReplyInbox / ReplyPushLoop as-built.

Design chapter for the cross-host broker fabric. As-built pieces verified against msg.AmqpReplyInbox at wiki-source parity; proposed pieces track CB-308 (docs/CB-308-Multi-Host-Federation.md) and CB-500 (docs/CB-500-Multi-Tier-Coordination.md).