diff --git a/11-Features.md b/11-Features.md index d0d05a3..449f688 100644 --- a/11-Features.md +++ b/11-Features.md @@ -51,6 +51,7 @@ six weeks, and the table alone will not carry it. | [Durable reply inbox](#durable-reply-inbox) | `broker:` | CB-307 | `msg/AmqpReplyInbox` | | [Set a member's auto-compact window](#set-a-members-auto-compact-window) | per-profile `autoCompactWindow:` | CB-636 | `member/ClaudeCodeLauncher`, `member/OpenCodeLauncher` | | [Leads on different hosts talk over a shared broker](#leads-on-different-hosts-talk-over-a-shared-broker) | `coordinator:` | CB-637 | `msg/LeadMailbox`, `msg/LeadCoordLoop` | +| [Draining an inbox is lead-only, on both entry paths](#draining-an-inbox-is-lead-only-on-both-entry-paths) | (always on, with `auth:`) | #272 | `mcp/FleetMcp.pollAction` | | [Broker password out of the config](#broker-password-out-of-the-config) | `broker.uriEnv:` | CB-635 | `config/FleetConfig` | | [An unreachable broker does not stop the daemon](#an-unreachable-broker-does-not-stop-the-daemon) | (always on) | CB-635 | `Fleetd.selectReplyInbox` | | [Reject overlapping rendezvous](#reject-overlapping-rendezvous) | automatic | CB-548 | `msg/Rendezvous` | @@ -430,6 +431,30 @@ real task's runtime. Without a durable inbox, a reply arriving after that window writing into someone else's broker. Also see the known hole: an async ticket that times out while the session is still BUSY currently discards the later completion rather than parking it. +## Draining an inbox is lead-only, on both entry paths + +**What.** `fleet_poll{target}` empties that session's reply inbox — the replies are removed and a +second call returns nothing. It is therefore gated as a drain, exactly like `fleet_ack`, and only +the primary may call it. `fleet_poll{ticket}` is unaffected: it observes an async delegation and +changes nothing, so a worker or an architect may still poll a ticket it owns. + +**On.** Always on, whenever authorization is enforced (an `auth:` block / a real `CallerResolver`). +Nothing to configure. + +**Why.** Until fleetd #272 the MCP handler passed a constant `READ` for both branches, and `READ` is +open to every authenticated role. Any worker could read a peer's `sessionId` out of `fleet_list` and +destroy the replies that peer had queued for the primary — an unrecoverable loss, since a drained +reply is gone. The REST path had always checked `DRAIN`, and the wiki had always described the tool +as lead-only; the MCP gate was the one that disagreed. + +**Gotcha.** An **architect** cannot drain either, even though it may `fleet_send`. That is +deliberate and matches `fleet_ack`: delegating is not a lifecycle right. If an architect delegates +with `wait:false` it still polls its own **ticket**, which is the branch that stayed open. + +The general rule this came from is worth keeping: when one tool name covers two operations, the gate +belongs **inside** the branch that picks between them, not above it. `fleet_poll` was the only +handler in the server with that shape — the other ten were audited and are correct. + ## Broker password out of the config **What.** `broker.uriEnv:` names an environment variable that holds the AMQP URI, instead of writing @@ -2542,11 +2567,24 @@ that made it unknowable. One WARN per launcher, not per spawn. this path is byte-identical to before, pinned by a test. **Why it exists.** The detector enumerates **fleetd's own** environment, on the assumption that the -member pane's login shell exports the same set. Under `memberHerdrSocket:` the pane belongs to a -different OS user, with a different `$HOME` and a different secret store, so that assumption is -simply false. The old output would then report on the wrong process — and could call a gap clean -while the member user exported something dangerous. There is no channel to read another user's -environment, so the honest answer is the only correct one. +member pane's login shell exports the same set. Under `memberHerdrSocket:` panes are routed to a +second herdr, and **fleetd has no channel to confirm what OS user that herdr runs as** — it may be +the same user or a different one, with a different `$HOME` and a different secret store. Either way +the assumption is no longer something fleetd can check, so the old output was reporting on a process +it could not see, and could call a gap clean while the member exported something dangerous. There is +no channel to read another user's environment, so "unknown" is the only honest answer. + +Note the shape of that sentence: the fix is *not* "the member runs as a different user". Saying that +would be the same mistake pointing the other way — fleetd cannot confirm a different uid any more +than it can confirm the same one (fleetd #269). What changed is the **claim**, from a false +certainty to an accurate unknown. + +**The count line beside it says whose environment it counted.** The INFO line +`member credentials: allowed N of M` reads as a statement about the member's pane, and under +`memberHerdrSocket:` it is not one — the counts come from fleetd's own process. Since fleetd #276 the +line says so in place, because the caveat used to live only in a javadoc the operator reading +`fleetd.out` never sees. The counts themselves are unchanged and still real. With +`memberHerdrSocket:` absent the wording is byte-identical to before. **The gotcha: this fixes the report, not the control.** Two related defects in the same area are open under **#213** — the ZDOTDIR scrub decides whether it can run by reading *fleetd's* `$SHELL`,