a pane member's grandchild process may be resolved as primary — possible worker→primary escalation #161

Closed
opened 2026-08-23 14:47:20 +02:00 by ltms · 2 comments
Owner

Security concern found while running the #160 experiment. Plausible from the code; not confirmed in practice — see below.

The mechanism

Identity is resolved from the pid that opened the connection:

  • ConnectionIdentity.resolve → PaneLocator.terminalForPid(pid)
  • PaneLocator.paneOwnsPid matches the pid against a pane's shell_pid or its foreground_processes pids — nothing else.

If the pid matches a pane, you are that pane's worker. If it matches no pane, CallerResolver falls through to loopback-trust and resolves you as primary.

A worker's own MCP client is the pane's foreground process, so a worker calling normally is correctly a worker. But a grandchild the worker spawns — python3, curl, any helper that opens its own connection to 127.0.0.1:8765 — is neither the shell pid nor a foreground pid. It matches no pane. It is resolved as primary.

If that holds, a worker escalates to primary — gaining fleet_spawn / fleet_stop / fleet_send over the whole fleet — just by shelling out. No exploit, an ordinary child process.

What I actually measured

  • A plain Python process (not in any pane) is resolved as primary by the live daemon, and a lead-only fleet_stop from it succeeded. So the "no pane ⇒ primary" half is confirmed.
  • I could not confirm the worker half. I asked a real worker member to run the same probe from inside its pane; its own local safety classifier refused to run the script, so the call never reached the daemon. The in-pane path is therefore unproven.

So: the daemon-side behaviour is confirmed (non-pane ⇒ primary); the worker-can-trigger-it half is inferred from process ancestry and not demonstrated end to end.

Why it matters

The whole authz model rests on "identity comes from the connection and is unforgeable". That is true for the pane's foreground process and not obviously true for its children. The invariant in the portable CLAUDE.md block — "Identity comes from the connection, never an argument" — is sound, but "the connection" is really "the connecting pid", and pid→pane is a foreground-only match.

How to confirm

Run, from inside a worker pane, a process that opens its own MCP session and calls a lead-only tool. If it is allowed, the escalation is real. The #160 probe scripts do exactly this; the obstacle was the worker backend's classifier, not the daemon. A backend without that classifier (or a direct child spawned by the launcher for the test) would settle it.

Possible directions, once confirmed

  • Walk the process ancestry in paneOwnsPid: a pid whose ancestor is a pane's shell pid belongs to that pane. This closes the gap in the direction of more things being workers, which is the safe direction.
  • Or require a pane-bound token for worker calls, though that fights the "nothing spoofable" design.

The ancestry walk is the natural fix and matches what the PaneLocator docstring already claims it does — it says "the agent pane whose process tree contains it", but the code only checks direct foreground pids, not the tree.

Related

  • #160 — surfaced here; the manager must never let a member mount it, for the same root reason.
  • The portable-block invariant 3 on identity.
Security concern found while running the #160 experiment. **Plausible from the code; not confirmed in practice — see below.** ## The mechanism Identity is resolved from the pid that opened the connection: - `ConnectionIdentity.resolve` → `PaneLocator.terminalForPid(pid)` - `PaneLocator.paneOwnsPid` matches the pid against a pane's **`shell_pid`** or its **`foreground_processes`** pids — nothing else. If the pid matches a pane, you are that pane's worker. If it matches no pane, `CallerResolver` falls through to loopback-trust and resolves you as **primary**. A worker's own MCP client is the pane's foreground process, so a worker calling normally is correctly a worker. But a **grandchild** the worker spawns — `python3`, `curl`, any helper that opens its own connection to `127.0.0.1:8765` — is neither the shell pid nor a foreground pid. It matches no pane. It is resolved as **primary**. If that holds, a worker escalates to primary — gaining `fleet_spawn` / `fleet_stop` / `fleet_send` over the whole fleet — just by shelling out. No exploit, an ordinary child process. ## What I actually measured - A plain Python process (not in any pane) **is** resolved as `primary` by the live daemon, and a lead-only `fleet_stop` from it **succeeded**. So the "no pane ⇒ primary" half is confirmed. - I could **not** confirm the worker half. I asked a real worker member to run the same probe from inside its pane; its own local safety classifier refused to run the script, so the call never reached the daemon. The in-pane path is therefore unproven. So: the daemon-side behaviour is confirmed (non-pane ⇒ primary); the worker-can-trigger-it half is inferred from process ancestry and not demonstrated end to end. ## Why it matters The whole authz model rests on "identity comes from the connection and is unforgeable". That is true for the pane's foreground process and not obviously true for its children. The invariant in the portable CLAUDE.md block — *"Identity comes from the connection, never an argument"* — is sound, but "the connection" is really "the connecting pid", and pid→pane is a foreground-only match. ## How to confirm Run, from **inside** a worker pane, a process that opens its own MCP session and calls a lead-only tool. If it is allowed, the escalation is real. The #160 probe scripts do exactly this; the obstacle was the worker backend's classifier, not the daemon. A backend without that classifier (or a direct child spawned by the launcher for the test) would settle it. ## Possible directions, once confirmed - Walk the process ancestry in `paneOwnsPid`: a pid whose **ancestor** is a pane's shell pid belongs to that pane. This closes the gap in the direction of *more* things being workers, which is the safe direction. - Or require a pane-bound token for worker calls, though that fights the "nothing spoofable" design. The ancestry walk is the natural fix and matches what the `PaneLocator` docstring already *claims* it does — it says "the agent pane whose process tree contains it", but the code only checks direct foreground pids, not the tree. ## Related - #160 — surfaced here; the manager must never let a member mount it, for the same root reason. - The portable-block invariant 3 on identity.
Author
Owner

Fixed and merged to main as 08ce9ae. Closing.

The unconfirmed half is now settled

This ticket said the daemon-side behaviour was confirmed (no pane ⇒ primary) but the worker-can-trigger-it half was inferred from process ancestry and never demonstrated. Reading the code settles it without needing the blocked probe: paneOwnsPid matched a pid only against a pane's shell_pid and its foreground_processes, while the class javadoc had always claimed it found "the agent pane whose process tree contains it". The javadoc described the intended behaviour; the code never implemented it.

So the mechanism was exactly as described, and the ancestry walk is the fix this ticket already identified as natural.

What shipped

terminalForPid now builds the caller's ancestor set once — itself plus its parent chain, walked through a new injectable ParentResolver seam, bounded at 32 generations with a cycle guard — and matches any ancestor against each pane's pids. The set is computed once per resolution and reused across both herdr clients on the CB-185 two-daemon path.

This only ever adds matches, which is the safe direction: the failure mode of the fix is a member correctly restricted; the failure mode of the bug is a member acting as the lead.

On the over-matching risk, which was the thing worth worrying about

The real danger of this fix is the opposite error: if the lead's own tree ever reached a pane's shell_pid, the lead would be resolved as a worker and every orchestration call would be refused — the fleet would stop working.

A reviewer checked this empirically on this host rather than by reading alone. ps -eo pid,ppid,command shows every pane's shell spawned as a direct, sibling child of the single herdr server process — no pane is nested inside another pane's tree. The walk goes strictly upward, so a caller's ancestry can only ever contain the shell_pid of the one pane hosting it. And a lead that resolves to its own pane is already the designed path: CallerResolver maps that terminal back to PRIMARY through the leaders: registry (CB-530), which this change does not touch.

Not proven for a nested-pane topology, because none exists here.

Verification

  • mvn clean install unpiped: 1043 tests, exit 0, no [ERROR] lines.
  • 6 new tests. The two that matter were sabotage-proven: the grandchild case fails without the walk (expected: <term_x> but was: <null>), and the cycle case hangs without the guard. A call-counting resolver pins the "walked once" claim at 3 lookups rather than 6.
  • The test that a pid matching no pane still resolves to null is the one protecting the lead, and it is in the suite.

Noted, not fixed

terminalForPid still issues one pane.process_info call per pane on every resolution. The class javadoc already flags a spawn-time pid→terminal cache as the obvious optimisation. Unchanged here — this ticket was about correctness, and identity resolution is now doing strictly more work per call, so that cache is a little more attractive than it was.

Fixed and merged to `main` as `08ce9ae`. Closing. ## The unconfirmed half is now settled This ticket said the daemon-side behaviour was confirmed (no pane ⇒ primary) but the worker-can-trigger-it half was inferred from process ancestry and never demonstrated. Reading the code settles it without needing the blocked probe: `paneOwnsPid` matched a pid **only** against a pane's `shell_pid` and its `foreground_processes`, while the class javadoc had always claimed it found "the agent pane whose process tree contains it". The javadoc described the intended behaviour; the code never implemented it. So the mechanism was exactly as described, and the ancestry walk is the fix this ticket already identified as natural. ## What shipped `terminalForPid` now builds the caller's ancestor set once — itself plus its parent chain, walked through a new injectable `ParentResolver` seam, bounded at 32 generations with a cycle guard — and matches any ancestor against each pane's pids. The set is computed once per resolution and reused across both herdr clients on the CB-185 two-daemon path. This only ever **adds** matches, which is the safe direction: the failure mode of the fix is a member correctly restricted; the failure mode of the bug is a member acting as the lead. ## On the over-matching risk, which was the thing worth worrying about The real danger of this fix is the opposite error: if the **lead's** own tree ever reached a pane's `shell_pid`, the lead would be resolved as a worker and every orchestration call would be refused — the fleet would stop working. A reviewer checked this empirically on this host rather than by reading alone. `ps -eo pid,ppid,command` shows every pane's shell spawned as a **direct, sibling child** of the single herdr server process — no pane is nested inside another pane's tree. The walk goes strictly upward, so a caller's ancestry can only ever contain the `shell_pid` of the one pane hosting it. And a lead that resolves to its own pane is already the designed path: `CallerResolver` maps that terminal back to `PRIMARY` through the `leaders:` registry (CB-530), which this change does not touch. Not proven for a nested-pane topology, because none exists here. ## Verification - `mvn clean install` unpiped: **1043 tests**, exit 0, no `[ERROR]` lines. - 6 new tests. The two that matter were sabotage-proven: the grandchild case fails without the walk (`expected: <term_x> but was: <null>`), and the cycle case **hangs** without the guard. A call-counting resolver pins the "walked once" claim at 3 lookups rather than 6. - The test that a pid matching no pane still resolves to `null` is the one protecting the lead, and it is in the suite. ## Noted, not fixed `terminalForPid` still issues one `pane.process_info` call per pane on every resolution. The class javadoc already flags a spawn-time `pid→terminal` cache as the obvious optimisation. Unchanged here — this ticket was about correctness, and identity resolution is now doing strictly more work per call, so that cache is a little more attractive than it was.
ltms closed this issue 2026-08-31 09:27:32 +02:00
Author
Owner

Confirmed live on the running daemon — the half this ticket could never prove

The fix is deployed (pid 19293, jar 509df0608da2, 1049 tests) and I have now run the end-to-end probe this ticket described but could not complete.

The original attempt failed because the worker's own local safety classifier refused to run the script, so the call never reached the daemon. I used an opencode member instead, which has no such classifier, and made the probe read-only: a single fleet_whoami rather than a lead-only mutation. That removes both the obstacle and the risk of a probe that actually spawns or stops something if the escalation is real.

The probe

A member ran, from its own shell, a plain curl — a grandchild process, neither the pane's shell_pid nor a foreground pid — which completed its own MCP handshake to 127.0.0.1:8765 and called fleet_whoami:

data: {"jsonrpc":"2.0","id":2,"result":{"content":[{"type":"text","text":
  "{\"role\":\"worker\",\"sessionId\":\"term_65a52c58aaccf29\",
    \"paneId\":\"f46ccd8e-0009-495a-8adb-7ed11b7a42f1\",\"profile\":\"gx\",
    \"state\":\"busy\",\"owner\":\"term_65a4c36d11af27\"}"}],"isError":false}}

"role":"worker". Before this fix that same connection would have matched no pane and been resolved as the primary.

The member also cross-checked that it is bound to its own pane rather than merely to some worker: the sessionId, paneId and owner from the curl probe are identical to those from its own mounted bridge connection. So the daemon traced the grandchild back to the exact right pane.

And the lead is still the lead

The opposite failure — over-matching demoting the primary, which would refuse every orchestration call — did not happen. fleet_whoami on the lead's own connection still answers primary with the ancestry walk live, and this session has since spawned, delegated to, polled and stopped members normally.

So both directions are now demonstrated on the live daemon, not inferred:

caller before after
member's own agent process worker worker
member's grandchild (curl) primary ← the hole worker
the lead primary primary

Nothing further outstanding on this ticket.

## Confirmed live on the running daemon — the half this ticket could never prove The fix is deployed (pid 19293, jar `509df0608da2`, 1049 tests) and I have now run the end-to-end probe this ticket described but could not complete. The original attempt failed because the worker's own local safety classifier refused to run the script, so the call never reached the daemon. I used an **opencode** member instead, which has no such classifier, and made the probe read-only: a single `fleet_whoami` rather than a lead-only mutation. That removes both the obstacle and the risk of a probe that actually spawns or stops something if the escalation is real. ## The probe A member ran, **from its own shell**, a plain `curl` — a grandchild process, neither the pane's `shell_pid` nor a foreground pid — which completed its own MCP handshake to `127.0.0.1:8765` and called `fleet_whoami`: ``` data: {"jsonrpc":"2.0","id":2,"result":{"content":[{"type":"text","text": "{\"role\":\"worker\",\"sessionId\":\"term_65a52c58aaccf29\", \"paneId\":\"f46ccd8e-0009-495a-8adb-7ed11b7a42f1\",\"profile\":\"gx\", \"state\":\"busy\",\"owner\":\"term_65a4c36d11af27\"}"}],"isError":false}} ``` **`"role":"worker"`.** Before this fix that same connection would have matched no pane and been resolved as the primary. The member also cross-checked that it is bound to *its own* pane rather than merely to some worker: the `sessionId`, `paneId` and `owner` from the curl probe are identical to those from its own mounted bridge connection. So the daemon traced the grandchild back to the exact right pane. ## And the lead is still the lead The opposite failure — over-matching demoting the primary, which would refuse every orchestration call — did not happen. `fleet_whoami` on the lead's own connection still answers `primary` with the ancestry walk live, and this session has since spawned, delegated to, polled and stopped members normally. So both directions are now demonstrated on the live daemon, not inferred: | caller | before | after | |---|---|---| | member's own agent process | worker | worker | | **member's grandchild (curl)** | **primary** ← the hole | **worker** | | the lead | primary | primary | Nothing further outstanding on this ticket.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#161