Features: #317 — an unresolved caller is refused, not promoted
+26
@@ -3860,3 +3860,29 @@ retry loop early when the policy returns an already-unreachable profile only mak
|
||||
the same dead profile, because only the policy chooses what comes next. The failure message now
|
||||
counts *distinct* candidates, so "tried 1 distinct candidate(s)" is the visible symptom of this
|
||||
whole class of problem.
|
||||
|
||||
## A caller whose identity cannot be resolved is refused, not treated as the primary
|
||||
|
||||
**What.** In `loopback-trust` mode, fleetd grants the primary role only to a caller whose operating
|
||||
system process id was actually found. A caller whose peer-PID lookup failed is now refused as
|
||||
`ANONYMOUS`. `ConnectionIdentity.Caller` carries a `resolved()` predicate that says which case it is.
|
||||
|
||||
**On.** Always on in `loopback-trust`, which is the mode you get when `fleetd.yaml` has no `auth:`
|
||||
block — this fleet's live configuration. `token` mode was never affected.
|
||||
|
||||
**Why it exists.** `LsofPeerPidLookup` returns `-1` for every failure, including the silent one
|
||||
where `lsof` runs fine and simply reports no matching process. No pane matches `-1`, so the caller
|
||||
arrived at the resolver with a `null` terminal — the same shape the real primary has, because no
|
||||
pane owns the primary either. The two were indistinguishable, and both were granted `SPAWN`, `STOP`,
|
||||
`SEND` and `DRAIN`. A worker whose lookup failed became the lead. `PaneLocator`'s javadoc had already
|
||||
named this exact escalation for a different trigger, and CB-161's ancestry walk closed that one — but
|
||||
the walk needs a candidate pid to walk, and a failed lookup has none.
|
||||
|
||||
**One thing to know for maintenance.** The signal that separates the two cases was already in the
|
||||
data and simply not read: the primary has a real pid and no pane; an unresolved caller has neither.
|
||||
That is why `resolved()` lives on the record next to the sentinel rather than as a `pid > 0` test
|
||||
copied into the resolver — one rule in two copies is what let #305 drift. This change is
|
||||
deliberately fail-closed: if `lsof` ever fails for the *primary's own* connection, the primary is
|
||||
refused until its next call resolves. That costs availability and it is the right trade, because the
|
||||
old behaviour spent it on a silent escalation instead. A new DEBUG line in `LsofPeerPidLookup` now
|
||||
names the previously-silent no-match case, so a refusal that does happen can be diagnosed.
|
||||
|
||||
Reference in New Issue
Block a user