fleetd #368: a stale lead binding no longer swallows the reply nudge

Dai Ha
2026-09-06 20:17:25 +07:00
parent fd7824d28f
commit 42b66db99f
+9
@@ -777,6 +777,15 @@ recorded delegation the daemon nudges **nobody** rather than guessing, and deliv
result is worse than a quiet inbox — but it does mean a post-restart reply may need an explicit
poll.
**Second gotcha, fixed in fleetd #368.** A binding was only ever dropped when the **worker** was
released. The map is keyed by the worker, so a lead that was closed, crashed or relaunched left its
bindings behind, pointing at a pane that no longer exists. A stale entry is not null, so it beat the
single-lead fallback above every time, and the nudge was sent into nothing. The push loop now checks
the recorded lead with `agents.status` before trusting it, and forgets a lead that is really gone so
resolution reaches the fallback. Only `agent_not_found` counts as gone: any other failure — a socket
blip, a decode error — is treated as live, because a wrong guess there unbinds a working lead
permanently, while a wrong guess the other way costs one retry on the next tick.
## Unknown config keys are named
**What.** A top-level key in `fleetd.yaml` that this build does not understand is logged as a WARN