From 4710a1c4354f3bda719bddac215e8a774f08c4e9 Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Fri, 4 Sep 2026 14:06:19 +0700 Subject: [PATCH] =?UTF-8?q?Features:=20#317=20=E2=80=94=20an=20unresolved?= =?UTF-8?q?=20caller=20is=20refused,=20not=20promoted?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- 11-Features.md | 26 ++++++++++++++++++++++++++ 1 file changed, 26 insertions(+) diff --git a/11-Features.md b/11-Features.md index ee217fa..b755c78 100644 --- a/11-Features.md +++ b/11-Features.md @@ -3860,3 +3860,29 @@ retry loop early when the policy returns an already-unreachable profile only mak the same dead profile, because only the policy chooses what comes next. The failure message now counts *distinct* candidates, so "tried 1 distinct candidate(s)" is the visible symptom of this whole class of problem. + +## A caller whose identity cannot be resolved is refused, not treated as the primary + +**What.** In `loopback-trust` mode, fleetd grants the primary role only to a caller whose operating +system process id was actually found. A caller whose peer-PID lookup failed is now refused as +`ANONYMOUS`. `ConnectionIdentity.Caller` carries a `resolved()` predicate that says which case it is. + +**On.** Always on in `loopback-trust`, which is the mode you get when `fleetd.yaml` has no `auth:` +block — this fleet's live configuration. `token` mode was never affected. + +**Why it exists.** `LsofPeerPidLookup` returns `-1` for every failure, including the silent one +where `lsof` runs fine and simply reports no matching process. No pane matches `-1`, so the caller +arrived at the resolver with a `null` terminal — the same shape the real primary has, because no +pane owns the primary either. The two were indistinguishable, and both were granted `SPAWN`, `STOP`, +`SEND` and `DRAIN`. A worker whose lookup failed became the lead. `PaneLocator`'s javadoc had already +named this exact escalation for a different trigger, and CB-161's ancestry walk closed that one — but +the walk needs a candidate pid to walk, and a failed lookup has none. + +**One thing to know for maintenance.** The signal that separates the two cases was already in the +data and simply not read: the primary has a real pid and no pane; an unresolved caller has neither. +That is why `resolved()` lives on the record next to the sentinel rather than as a `pid > 0` test +copied into the resolver — one rule in two copies is what let #305 drift. This change is +deliberately fail-closed: if `lsof` ever fails for the *primary's own* connection, the primary is +refused until its next call resolves. That costs availability and it is the right trade, because the +old behaviour spent it on a silent escalation instead. A new DEBUG line in `LsofPeerPidLookup` now +names the previously-silent no-match case, so a refusal that does happen can be diagnosed.