diff --git a/11-Features.md b/11-Features.md index ee217fa..b755c78 100644 --- a/11-Features.md +++ b/11-Features.md @@ -3860,3 +3860,29 @@ retry loop early when the policy returns an already-unreachable profile only mak the same dead profile, because only the policy chooses what comes next. The failure message now counts *distinct* candidates, so "tried 1 distinct candidate(s)" is the visible symptom of this whole class of problem. + +## A caller whose identity cannot be resolved is refused, not treated as the primary + +**What.** In `loopback-trust` mode, fleetd grants the primary role only to a caller whose operating +system process id was actually found. A caller whose peer-PID lookup failed is now refused as +`ANONYMOUS`. `ConnectionIdentity.Caller` carries a `resolved()` predicate that says which case it is. + +**On.** Always on in `loopback-trust`, which is the mode you get when `fleetd.yaml` has no `auth:` +block — this fleet's live configuration. `token` mode was never affected. + +**Why it exists.** `LsofPeerPidLookup` returns `-1` for every failure, including the silent one +where `lsof` runs fine and simply reports no matching process. No pane matches `-1`, so the caller +arrived at the resolver with a `null` terminal — the same shape the real primary has, because no +pane owns the primary either. The two were indistinguishable, and both were granted `SPAWN`, `STOP`, +`SEND` and `DRAIN`. A worker whose lookup failed became the lead. `PaneLocator`'s javadoc had already +named this exact escalation for a different trigger, and CB-161's ancestry walk closed that one — but +the walk needs a candidate pid to walk, and a failed lookup has none. + +**One thing to know for maintenance.** The signal that separates the two cases was already in the +data and simply not read: the primary has a real pid and no pane; an unresolved caller has neither. +That is why `resolved()` lives on the record next to the sentinel rather than as a `pid > 0` test +copied into the resolver — one rule in two copies is what let #305 drift. This change is +deliberately fail-closed: if `lsof` ever fails for the *primary's own* connection, the primary is +refused until its next call resolves. That costs availability and it is the right trade, because the +old behaviour spent it on a silent escalation instead. A new DEBUG line in `LsofPeerPidLookup` now +names the previously-silent no-match case, so a refusal that does happen can be diagnosed.