A worker connecting from 127.0.0.2 is resolved as the primary, and gains spawn/stop/send/drain #305

Closed
opened 2026-09-04 07:58:08 +02:00 by ltms · 1 comment
Owner

Privilege escalation. Reachable today on the Linux fleet host, in the default auth mode. Found by a delegated hunter; I verified every step of it myself, including measuring the kernel behaviour it depends on.

The two definitions of "loopback"

ConnectionIdentity accepts three exact strings:

private static boolean isLoopback(String addr) {
    return "127.0.0.1".equals(addr) || "::1".equals(addr) || "0:0:0:0:0:0:0:1".equals(addr);
}

CallerResolver keeps its own copy, and that one is wider:

return remoteAddr.equals("127.0.0.1") || remoteAddr.equals("::1")
        || remoteAddr.equals("0:0:0:0:0:0:0:1") || remoteAddr.startsWith("127.");

One rule, two copies, and they drifted. That is the #284 / #297 shape again — but this time the drift is between an identity resolver and an authorization check, so the cost is not a wrong number in a report.

Why the drift escalates

ConnectionIdentity.resolve returns before the PID lookup when its check fails:

public Caller resolve(String remoteAddr, int remotePort) {
    if (!isLoopback(remoteAddr)) {
        return new Caller(null, -1);   // no PID lookup, so no pane, so no terminal
    }
    long pid = pids.pidForLocalPort(remotePort);
    return new Caller(panes.terminalForPid(pid), pid);
}

Then CallerResolver.resolve reads a null terminal as "not a worker", and under loopback-trust a same-host non-worker is the primary:

// loopback-trust: same-host callers that are not workers are the primary.
return isLoopback(remoteAddr) ? Principal.primary(c.pid()) : Principal.anonymous();

So for a source address of 127.0.0.2: identity resolution is skipped, the caller has no terminal, and the wider check here promotes it. The escalation happens precisely because the identity side was the stricter of the two.

Being strict in ConnectionIdentity is not the safe direction. That predicate does not decide whether a caller is trusted; it decides whether the caller's identity is resolved at all. Every address it excludes is an address on which a worker silently becomes the lead. It also sits in front of the PID ancestry walk, so the fix that stopped a member's curl child being resolved as primary is bypassed too — this hole is wider than that one was.

Authz grants PRIMARY spawn, stop, send and drain. A worker reaching this gets all four.

Reachability — measured, not argued

Both doors pass the raw socket address straight in, so the resolver is the real caller path, not a seam:

  • FleetApp.java:201-203 — auth.resolve(ctx.req().getRemoteAddr(), ctx.req().getRemotePort(), …)
  • FleetMcp.java:239-245 — callers.resolve(req.getRemoteAddr(), req.getRemotePort(), …)

The remaining question is whether the OS will deliver a connection whose source is 127.0.0.2. I measured it on both hosts rather than assuming:

macOS (this workstation) — not reachable by default. 127.0.0.2 is not assigned to lo0, so the source bind fails:

$ curl -s --max-time 5 --interface 127.0.0.2 http://127.0.0.1:8765/healthz
curl exit 45   # interface could not be used

Linux (fleet01) — fully reachable.

$ curl -s --max-time 3 --interface 127.0.0.2 http://127.0.0.1:1/
rc=7            # 7 = the source bind SUCCEEDED, the connection was refused (port 1 is closed)
$ ip -4 addr show lo
1: lo: <LOOPBACK,UP,LOWER_UP> ...
    inet 127.0.0.1/8 scope host lo

rc=7 rather than rc=45 is the whole proof: the kernel accepted 127.0.0.2 as a source. The /8 prefix is why — on Linux the entire range is local by default. A worker there needs one flag: curl --interface 127.0.0.2 http://127.0.0.1:8765/members -X POST.

The mode is not exotic either. auth: is commented out in fleetd.example.yaml, so loopback-trust is the default, and the live daemon logs it at every boot:

12:41:34.151 INFO  [main] dev.ltms.fleet.Fleetd - auth: loopback-trust (any loopback non-worker caller is the primary)

The fix

One predicate, on ConnectionIdentity, called by CallerResolver. Sharing the inputs would not have helped here — both already read the same address. Only sharing the computation does.

The shared predicate accepts all of 127.0.0.0/8 plus the IPv4-mapped IPv6 form. Widening a check normally weakens it; widening this one closes the hole, because resolution is what demotes a worker.

A second, milder bug the same fix closes

Neither copy handled ::ffff:127.0.0.1, the IPv4-mapped form Jetty can report. A legitimate primary on that form resolved to ANONYMOUS and was refused everything. That is the safe direction, so it is a much weaker bug than the escalation — but it is real, and the shared predicate fixes it too.

Test coverage before this ticket

Every existing test used 127.0.0.1. ConnectionIdentityTest and CallerResolverTest both, and FleetAppAuthTest binds and targets 127.0.0.1 as well. The gate was never exercised on any other loopback address, which is exactly why a drift between two copies of it survived.

I am not adding a transport-level test that binds a real 127.0.0.2 source: it cannot run on macOS, so it would fail for every developer on this workstation. Recording that decision here rather than leaving the absence unexplained.

Privilege escalation. Reachable today on the Linux fleet host, in the default auth mode. Found by a delegated hunter; **I verified every step of it myself**, including measuring the kernel behaviour it depends on. ## The two definitions of "loopback" `ConnectionIdentity` accepts three exact strings: ```java private static boolean isLoopback(String addr) { return "127.0.0.1".equals(addr) || "::1".equals(addr) || "0:0:0:0:0:0:0:1".equals(addr); } ``` `CallerResolver` keeps its own copy, and that one is wider: ```java return remoteAddr.equals("127.0.0.1") || remoteAddr.equals("::1") || remoteAddr.equals("0:0:0:0:0:0:0:1") || remoteAddr.startsWith("127."); ``` One rule, two copies, and they drifted. That is the #284 / #297 shape again — but this time the drift is between an identity resolver and an authorization check, so the cost is not a wrong number in a report. ## Why the drift escalates `ConnectionIdentity.resolve` returns before the PID lookup when its check fails: ```java public Caller resolve(String remoteAddr, int remotePort) { if (!isLoopback(remoteAddr)) { return new Caller(null, -1); // no PID lookup, so no pane, so no terminal } long pid = pids.pidForLocalPort(remotePort); return new Caller(panes.terminalForPid(pid), pid); } ``` Then `CallerResolver.resolve` reads a null terminal as "not a worker", and under loopback-trust a same-host non-worker is the primary: ```java // loopback-trust: same-host callers that are not workers are the primary. return isLoopback(remoteAddr) ? Principal.primary(c.pid()) : Principal.anonymous(); ``` So for a source address of `127.0.0.2`: identity resolution is **skipped**, the caller has no terminal, and the wider check here promotes it. The escalation happens precisely because the identity side was the stricter of the two. **Being strict in `ConnectionIdentity` is not the safe direction.** That predicate does not decide whether a caller is trusted; it decides whether the caller's identity is resolved at all. Every address it excludes is an address on which a worker silently becomes the lead. It also sits *in front of* the PID ancestry walk, so the fix that stopped a member's `curl` child being resolved as primary is bypassed too — this hole is wider than that one was. `Authz` grants `PRIMARY` spawn, stop, send and drain. A worker reaching this gets all four. ## Reachability — measured, not argued Both doors pass the raw socket address straight in, so the resolver is the real caller path, not a seam: - `FleetApp.java:201-203` — `auth.resolve(ctx.req().getRemoteAddr(), ctx.req().getRemotePort(), …)` - `FleetMcp.java:239-245` — `callers.resolve(req.getRemoteAddr(), req.getRemotePort(), …)` The remaining question is whether the OS will deliver a connection whose source is `127.0.0.2`. I measured it on both hosts rather than assuming: **macOS (this workstation) — not reachable by default.** `127.0.0.2` is not assigned to `lo0`, so the source bind fails: ``` $ curl -s --max-time 5 --interface 127.0.0.2 http://127.0.0.1:8765/healthz curl exit 45 # interface could not be used ``` **Linux (fleet01) — fully reachable.** ``` $ curl -s --max-time 3 --interface 127.0.0.2 http://127.0.0.1:1/ rc=7 # 7 = the source bind SUCCEEDED, the connection was refused (port 1 is closed) $ ip -4 addr show lo 1: lo: <LOOPBACK,UP,LOWER_UP> ... inet 127.0.0.1/8 scope host lo ``` `rc=7` rather than `rc=45` is the whole proof: the kernel accepted `127.0.0.2` as a source. The `/8` prefix is why — on Linux the entire range is local by default. A worker there needs one flag: `curl --interface 127.0.0.2 http://127.0.0.1:8765/members -X POST`. The mode is not exotic either. `auth:` is commented out in `fleetd.example.yaml`, so loopback-trust is the default, and the live daemon logs it at every boot: ``` 12:41:34.151 INFO [main] dev.ltms.fleet.Fleetd - auth: loopback-trust (any loopback non-worker caller is the primary) ``` ## The fix One predicate, on `ConnectionIdentity`, called by `CallerResolver`. Sharing the inputs would not have helped here — both already read the same address. Only sharing the computation does. The shared predicate accepts all of `127.0.0.0/8` plus the IPv4-mapped IPv6 form. Widening a check normally weakens it; widening this one closes the hole, because resolution is what *demotes* a worker. ## A second, milder bug the same fix closes Neither copy handled `::ffff:127.0.0.1`, the IPv4-mapped form Jetty can report. A legitimate primary on that form resolved to `ANONYMOUS` and was refused everything. That is the safe direction, so it is a much weaker bug than the escalation — but it is real, and the shared predicate fixes it too. ## Test coverage before this ticket Every existing test used `127.0.0.1`. `ConnectionIdentityTest` and `CallerResolverTest` both, and `FleetAppAuthTest` binds and targets `127.0.0.1` as well. The gate was never exercised on any other loopback address, which is exactly why a drift between two copies of it survived. I am not adding a transport-level test that binds a real `127.0.0.2` source: it cannot run on macOS, so it would fail for every developer on this workstation. Recording that decision here rather than leaving the absence unexplained.
ltms closed this issue 2026-09-04 07:58:27 +02:00
Author
Owner

Fixed in 9379f92, on main.

Mutation proof

Reverted only the two source files (git apply -R), kept the tests. The failure text states the escalation better than prose can:

[ERROR] Tests run: 35, Failures: 2 -- in dev.ltms.fleet.auth.CallerResolverTest
org.opentest4j.AssertionFailedError: a worker must stay a worker from source 127.0.0.2
    ==> expected: <WORKER> but was: <PRIMARY>
org.opentest4j.AssertionFailedError: source ::ffff:127.0.0.1
    ==> expected: <PRIMARY> but was: <ANONYMOUS>
[ERROR] Tests run: 5, Failures: 1 -- in dev.ltms.fleet.mcp.ConnectionIdentityTest
org.opentest4j.AssertionFailedError: expected: <term_a> but was: <null>

The third line is the mechanism: without the fix a worker at 127.0.0.2 resolves to no terminal at all, and the two lines above are what that costs.

Restored, then mvn clean install: Tests run: 1313, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. main was at 1310; this adds the 3 tests above.

Exposure of the running daemon

The live daemon on this workstation is still on the pre-fix jar eca358254637, and it is not exploitable there. macOS does not assign 127.0.0.2 to lo0, so the source bind fails before a packet is sent (curl exit 45, measured). I am not doing an emergency restart: four members are mid-task, a restart drops their tickets and reports, and there is no exposure here to race. It goes out with the next redeploy.

fleet01 is the host where this matters, since its kernel accepts the source bind. Its daemon is deployed manually and separately, and I have not touched it.

Credit and one correction to the report

The hunter found this, traced both files, and got the direction of harm right — including that the harm is more-permission rather than less. It also spotted that no test exercised any loopback address other than 127.0.0.1, which is why the drift survived.

One thing it could not do: it wrote "I did not run a live source-bound request", and that turned out to be the load-bearing question. The whole finding rests on whether an OS delivers a connection from 127.0.0.2, and the answer differs by host — no on macOS, yes on Linux. Reading the Java alone cannot settle it. I measured both.

That is the right split of work. The report was accurate about what it had and had not checked, which is what made it quick to verify rather than something to re-derive.

Fixed in `9379f92`, on `main`. ## Mutation proof Reverted only the two source files (`git apply -R`), kept the tests. The failure text states the escalation better than prose can: ``` [ERROR] Tests run: 35, Failures: 2 -- in dev.ltms.fleet.auth.CallerResolverTest org.opentest4j.AssertionFailedError: a worker must stay a worker from source 127.0.0.2 ==> expected: <WORKER> but was: <PRIMARY> org.opentest4j.AssertionFailedError: source ::ffff:127.0.0.1 ==> expected: <PRIMARY> but was: <ANONYMOUS> [ERROR] Tests run: 5, Failures: 1 -- in dev.ltms.fleet.mcp.ConnectionIdentityTest org.opentest4j.AssertionFailedError: expected: <term_a> but was: <null> ``` The third line is the mechanism: without the fix a worker at `127.0.0.2` resolves to no terminal at all, and the two lines above are what that costs. Restored, then `mvn clean install`: `Tests run: 1313, Failures: 0, Errors: 0, Skipped: 0` — `BUILD SUCCESS`. `main` was at 1310; this adds the 3 tests above. ## Exposure of the running daemon The live daemon on this workstation is still on the pre-fix jar `eca358254637`, and **it is not exploitable there**. macOS does not assign `127.0.0.2` to `lo0`, so the source bind fails before a packet is sent (`curl` exit 45, measured). I am not doing an emergency restart: four members are mid-task, a restart drops their tickets and reports, and there is no exposure here to race. It goes out with the next redeploy. **fleet01 is the host where this matters**, since its kernel accepts the source bind. Its daemon is deployed manually and separately, and I have not touched it. ## Credit and one correction to the report The hunter found this, traced both files, and got the direction of harm right — including that the harm is more-permission rather than less. It also spotted that no test exercised any loopback address other than `127.0.0.1`, which is why the drift survived. One thing it could not do: it wrote "I did not run a live source-bound request", and that turned out to be the load-bearing question. The whole finding rests on whether an OS delivers a connection from `127.0.0.2`, and the answer differs by host — no on macOS, yes on Linux. Reading the Java alone cannot settle it. I measured both. That is the right split of work. The report was accurate about what it had and had not checked, which is what made it quick to verify rather than something to re-derive.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#305