Features: the fleet_reply lead refusal and the scrub receipt fix

Dai Ha
2026-09-10 09:19:55 +07:00
parent 6d8a2d2cb4
commit 0f88cb4573
+75
@@ -4902,3 +4902,78 @@ configurations. The list is the source of truth; nothing else can be.
boot — which is the point, but it is better to know before a restart than during one.
fleetd #398.
## `fleet_reply` tells a lead which tool to use instead
**What.** A lead that calls `fleet_reply` is refused, and the refusal names both working routes:
```
fleet_reply has no route to a peer lead. Use fleet_send{coordId: ...} for a peer on another
daemon or fleet_send{sessionId: ...} for a peer on this host. fleet_reply resolves a member's
blocked fleet_send, and a peer's coord-id message is durable and non-blocking, so there is
nothing for it to resolve.
```
**The knob.** None. The check runs on every `fleet_reply` call and needs no configuration.
**Why it exists.** A lead that gets a message from a peer reaches for "reply" — the word matches
what it is doing. But `fleet_reply` exists to resolve a *member's* blocked `fleet_send`, and a peer
lead never has one open. So the call used to be accepted and the message went into a worker inbox
nobody would ever drain. The peer waited on nothing, and no error said so.
The old refusal was worse than none, because it only covered the case where the caller could not be
identified at all. A lead **is** identified, so it sailed past that check.
The message names the routes on purpose. A refusal that says only "not allowed" makes the lead
guess, and the two right answers depend on where the peer lives — `coordId` across daemons,
`sessionId` on this host.
**Gotchas.**
- **The order of the two checks matters, and is now pinned.** An unidentified caller gets the
"workers only" message even if its role is `PRIMARY`. That is deliberate: "we do not know who you
are" is the more useful thing to hear first.
- **A queued reply is still a success for a worker.** This refusal is about leads only. A worker
whose `fleet_reply` finds no open send still gets its reply stored in the inbox, and that is
normal.
- Role comes from the connection, never from an argument, so a caller cannot present itself as a
worker to get around this.
fleetd #391.
## The credential-scrub receipt now measures the blank, not the attempt
**What.** The startup scrub that clears inherited credentials from a member's shell reports which
names it actually blanked. It now checks each parameter's **value** after trying, instead of
trusting the exit status of the attempt.
```
allowed 41 of 57 failed 3
<name> ← blanked, confirmed empty
!<name> ← could not be blanked
```
**The knob.** None. The receipt is part of the scrub and always runs.
**Why it exists.** The old loop counted a name as blanked when `eval "export ${n}="` returned 0.
zsh has integer parameters — `SECONDS RANDOM SHLVL HISTSIZE COLUMNS LINES USERNAME` — and on those
an empty assignment is **coerced to a number** rather than failing. So `eval` returns 0 and the
value is unchanged. Measured across 10 names: **7 false receipts.**
The general rule, which cost two tickets to learn: **an attempt's exit status is not a measurement
of its effect.** The fix verifies afterwards with `[[ -z "${(P)n}" ]]` and only then counts it.
**Gotchas.**
- **`eval` is required, not stylistic.** `export UID=` is a *fatal* zsh parameter error that aborts
the whole sourced file — it once killed the scrub at name 42 of 57 and left the operator's own
exports untouched. Only `eval "export ${n}=" 2>/dev/null` contains that. The same applies to
`EUID GID EGID PPID LINENO`.
- **The `!` names are the ones to read.** They are shell parameters that cannot be blanked, not
credentials that leaked. A long `!` list is normal; a *shrinking* allowed count is the warning.
- **The scrub is zsh-only**, like the secret store it defends against. It lives in `.zshenv`, the
one file zsh always reads.
- The receipt reports the member's shell, so a claim about it measured from inside an agent's
`zsh -c` child is measuring the wrong process.
fleetd #394, #400.