Features: draining an inbox is lead-only (#272); the count line names whose env it counted (#276)

Also corrects this page's own version of the overclaim #269 removed from the
code: the entry asserted that under memberHerdrSocket the member pane belongs
to a different OS user. fleetd cannot confirm that either way — that is the
whole reason the report says unknown. Fixing the code and leaving the wiki
asserting the old claim is the same one-way repair the entry warns about.
Dai Ha
2026-09-04 10:03:32 +07:00
parent 2837d76bd9
commit f384546d30
+43 -5
@@ -51,6 +51,7 @@ six weeks, and the table alone will not carry it.
| [Durable reply inbox](#durable-reply-inbox) | `broker:` | CB-307 | `msg/AmqpReplyInbox` |
| [Set a member's auto-compact window](#set-a-members-auto-compact-window) | per-profile `autoCompactWindow:` | CB-636 | `member/ClaudeCodeLauncher`, `member/OpenCodeLauncher` |
| [Leads on different hosts talk over a shared broker](#leads-on-different-hosts-talk-over-a-shared-broker) | `coordinator:` | CB-637 | `msg/LeadMailbox`, `msg/LeadCoordLoop` |
| [Draining an inbox is lead-only, on both entry paths](#draining-an-inbox-is-lead-only-on-both-entry-paths) | (always on, with `auth:`) | #272 | `mcp/FleetMcp.pollAction` |
| [Broker password out of the config](#broker-password-out-of-the-config) | `broker.uriEnv:` | CB-635 | `config/FleetConfig` |
| [An unreachable broker does not stop the daemon](#an-unreachable-broker-does-not-stop-the-daemon) | (always on) | CB-635 | `Fleetd.selectReplyInbox` |
| [Reject overlapping rendezvous](#reject-overlapping-rendezvous) | automatic | CB-548 | `msg/Rendezvous` |
@@ -430,6 +431,30 @@ real task's runtime. Without a durable inbox, a reply arriving after that window
writing into someone else's broker. Also see the known hole: an async ticket that times out while
the session is still BUSY currently discards the later completion rather than parking it.
## Draining an inbox is lead-only, on both entry paths
**What.** `fleet_poll{target}` empties that session's reply inbox — the replies are removed and a
second call returns nothing. It is therefore gated as a drain, exactly like `fleet_ack`, and only
the primary may call it. `fleet_poll{ticket}` is unaffected: it observes an async delegation and
changes nothing, so a worker or an architect may still poll a ticket it owns.
**On.** Always on, whenever authorization is enforced (an `auth:` block / a real `CallerResolver`).
Nothing to configure.
**Why.** Until fleetd #272 the MCP handler passed a constant `READ` for both branches, and `READ` is
open to every authenticated role. Any worker could read a peer's `sessionId` out of `fleet_list` and
destroy the replies that peer had queued for the primary — an unrecoverable loss, since a drained
reply is gone. The REST path had always checked `DRAIN`, and the wiki had always described the tool
as lead-only; the MCP gate was the one that disagreed.
**Gotcha.** An **architect** cannot drain either, even though it may `fleet_send`. That is
deliberate and matches `fleet_ack`: delegating is not a lifecycle right. If an architect delegates
with `wait:false` it still polls its own **ticket**, which is the branch that stayed open.
The general rule this came from is worth keeping: when one tool name covers two operations, the gate
belongs **inside** the branch that picks between them, not above it. `fleet_poll` was the only
handler in the server with that shape — the other ten were audited and are correct.
## Broker password out of the config
**What.** `broker.uriEnv:` names an environment variable that holds the AMQP URI, instead of writing
@@ -2542,11 +2567,24 @@ that made it unknowable. One WARN per launcher, not per spawn.
this path is byte-identical to before, pinned by a test.
**Why it exists.** The detector enumerates **fleetd's own** environment, on the assumption that the
member pane's login shell exports the same set. Under `memberHerdrSocket:` the pane belongs to a
different OS user, with a different `$HOME` and a different secret store, so that assumption is
simply false. The old output would then report on the wrong process — and could call a gap clean
while the member user exported something dangerous. There is no channel to read another user's
environment, so the honest answer is the only correct one.
member pane's login shell exports the same set. Under `memberHerdrSocket:` panes are routed to a
second herdr, and **fleetd has no channel to confirm what OS user that herdr runs as** — it may be
the same user or a different one, with a different `$HOME` and a different secret store. Either way
the assumption is no longer something fleetd can check, so the old output was reporting on a process
it could not see, and could call a gap clean while the member exported something dangerous. There is
no channel to read another user's environment, so "unknown" is the only honest answer.
Note the shape of that sentence: the fix is *not* "the member runs as a different user". Saying that
would be the same mistake pointing the other way — fleetd cannot confirm a different uid any more
than it can confirm the same one (fleetd #269). What changed is the **claim**, from a false
certainty to an accurate unknown.
**The count line beside it says whose environment it counted.** The INFO line
`member credentials: allowed N of M` reads as a statement about the member's pane, and under
`memberHerdrSocket:` it is not one — the counts come from fleetd's own process. Since fleetd #276 the
line says so in place, because the caveat used to live only in a javadoc the operator reading
`fleetd.out` never sees. The counts themselves are unchanged and still real. With
`memberHerdrSocket:` absent the wording is byte-identical to before.
**The gotcha: this fixes the report, not the control.** Two related defects in the same area are
open under **#213** — the ZDOTDIR scrub decides whether it can run by reading *fleetd's* `$SHELL`,