Features: #317 — an unresolved caller is refused, not promoted

Dai Ha
2026-09-04 14:06:19 +07:00
parent 7b99569d5e
commit 4710a1c435
+26
@@ -3860,3 +3860,29 @@ retry loop early when the policy returns an already-unreachable profile only mak
the same dead profile, because only the policy chooses what comes next. The failure message now
counts *distinct* candidates, so "tried 1 distinct candidate(s)" is the visible symptom of this
whole class of problem.
## A caller whose identity cannot be resolved is refused, not treated as the primary
**What.** In `loopback-trust` mode, fleetd grants the primary role only to a caller whose operating
system process id was actually found. A caller whose peer-PID lookup failed is now refused as
`ANONYMOUS`. `ConnectionIdentity.Caller` carries a `resolved()` predicate that says which case it is.
**On.** Always on in `loopback-trust`, which is the mode you get when `fleetd.yaml` has no `auth:`
block — this fleet's live configuration. `token` mode was never affected.
**Why it exists.** `LsofPeerPidLookup` returns `-1` for every failure, including the silent one
where `lsof` runs fine and simply reports no matching process. No pane matches `-1`, so the caller
arrived at the resolver with a `null` terminal — the same shape the real primary has, because no
pane owns the primary either. The two were indistinguishable, and both were granted `SPAWN`, `STOP`,
`SEND` and `DRAIN`. A worker whose lookup failed became the lead. `PaneLocator`'s javadoc had already
named this exact escalation for a different trigger, and CB-161's ancestry walk closed that one — but
the walk needs a candidate pid to walk, and a failed lookup has none.
**One thing to know for maintenance.** The signal that separates the two cases was already in the
data and simply not read: the primary has a real pid and no pane; an unresolved caller has neither.
That is why `resolved()` lives on the record next to the sentinel rather than as a `pid > 0` test
copied into the resolver — one rule in two copies is what let #305 drift. This change is
deliberately fail-closed: if `lsof` ever fails for the *primary's own* connection, the primary is
refused until its next call resolves. That costs availability and it is the right trade, because the
old behaviour spent it on a silent escalation instead. A new DEBUG line in `LsofPeerPidLookup` now
names the previously-silent no-match case, so a refusal that does happen can be diagnosed.