blocking SSH_AUTH_SOCK is not a control: the forge key is an unencrypted file the member can read #184

Open
opened 2026-08-28 01:31:50 +02:00 by ltms · 5 comments
Owner

Found while proving the memberCredentials.policy: allow-list flip on 2026-08-28. The scrub works. The thing it was believed to be protecting is reachable another way.

What was believed

#110 and the sshAuthSock: block decision rest on this premise: SSH_AUTH_SOCK is a live handle to the operator's ssh-agent, so blocking it stops a member signing with the operator's keys. I stated, in a review comment on #177 and in a brief, that ~/.ssh holds no private key file at all and the forge identity exists only inside the agent.

That was wrong. I had looked only in ~/.ssh.

What is actually true

A live opencode member, spawned under policy: allow-list, reported:

SSH_AUTH_SOCK: empty
AWS_SECRET_ACCESS_KEY: empty
GITEA_ACCESS_TOKEN: empty
WORKER_GITEA_TOKEN: set (40 chars)

…and then pushed to the forge over SSH successfully.

The reason:

$ ssh -G git.ltms.dev
user git
port 2224
identityfile <a path outside ~/.ssh>

~/.ssh/config is four Include lines pointing into a shared-env directory, and the included files set IdentityFile for the forge. Checking that file:

forge private key file: EXISTS on disk
readable by this user  : yes
permissions            : -rw------- dai.ha
passphrase-protected   : NO — usable with no passphrase

A member runs as the same OS user. So it reads the key file directly. The agent is irrelevant to it.

Why this matters

  1. sshAuthSock: block is cosmetic on this host. It blanks a variable that SSH does not need. A member has full SSH access to the forge — and to every other host those included configs define an identity for, which is a large list.
  2. It is the same class of defect as #157, which is what makes it worth a ticket rather than a note. The memberCredentials scrub removes environment variables. It cannot remove a file. #157 was a token in .git/config; this is a key in ~/Drive00/.... Both walk straight past a control that only ever looked at the environment.
  3. The premise propagated into decisions. The urgency argument for the #157 HTTPS work ("once the agent is blocked a member cannot push at all") was built on it. That work is still correct — it removes the token from git config, which was the real defect — but it was never the thing standing between a member and the forge.

The general shape, again

The enumeration keeps being of the wrong thing. #596 enumerated a file and missed an environment. This enumerated an environment and missed a file. A control that names a channel will always be walked around by the channel nobody named.

The useful question is not "which variables do we blank?" but "what can a process running as this user reach?" — environment, files, sockets, keychain, and the ssh_config graph.

Suggested direction

Nothing here is a quick fix, and I have not implemented any of it.

  1. Report the gap rather than imply a guarantee. At spawn under allow-list, resolve ssh -G for the configured forge host and, if it yields a readable identityfile, log a WARN saying SSH access is NOT blocked by this policy. The daemon should not let an operator believe sshAuthSock: block did something it did not.
  2. Decide whether member isolation is actually achievable as a same-user process. It may not be. A member that runs as the operator's user can read anything the operator can. If real isolation is wanted, the answer is a separate OS user or a container, not a longer variable list. That is a design decision, not a bug fix — and worth taking before more effort goes into env-var scrubbing.
  3. Until then, say plainly in the docs what the scrub does and does not cover, so the next decision is not made on the premise this ticket just disproved.

What the flip DID achieve — not nothing

Worth keeping in view. Under allow-list the daemon blanked 27 credential-shaped variables, measured:

AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, BESZEL_ADMIN_PASSWORD, BESZEL_KEY,
BESZEL_UNIVERSAL_TOKEN, CF_API_TOKEN, CF_USER_TOKEN, CONFLUENCE_API_TOKEN,
GITEA_ACCESS_TOKEN, GITLAB_OAUTH_CLIENT_SECRET, GITLAB_PERSONAL_ACCESS_TOKEN,
GRAFANA_ADMIN_PASSWORD, HASS_TOKEN, HW_PASSWORD, JENKINS_MCP_AUTH, LLM_KEY_N8N,
LTMS_API_KEY, METRICS_PUSH_TOKEN, N8N_ENCRYPTION_KEY, N8N_OWNER_PASSWORD,
N8N_WEBHOOK_TOKEN, SSH_AUTH_SOCK, TELEGRAM_BOT_TOKEN, TS_API_KEY, TS_AUTHKEY,
BRAIN_MCP_TOKEN, MEMORY_MCP_TOKEN

That includes the AWS AdministratorAccess pair from #144 — a genuine, large reduction. This ticket is about not overstating it.

Related

  • #144 — the credential-exposure line.
  • #110 — the SSH_AUTH_SOCK decision this disproves the premise of.
  • #157 / #182 — a credential in a file, walking past an environment-only control.
Found while proving the `memberCredentials.policy: allow-list` flip on 2026-08-28. The scrub works. The thing it was believed to be protecting is reachable another way. ## What was believed #110 and the `sshAuthSock: block` decision rest on this premise: `SSH_AUTH_SOCK` is a live handle to the operator's ssh-agent, so blocking it stops a member signing with the operator's keys. I stated, in a review comment on #177 and in a brief, that `~/.ssh` holds **no private key file at all** and the forge identity exists **only** inside the agent. That was wrong. I had looked only in `~/.ssh`. ## What is actually true A live `opencode` member, spawned under `policy: allow-list`, reported: ``` SSH_AUTH_SOCK: empty AWS_SECRET_ACCESS_KEY: empty GITEA_ACCESS_TOKEN: empty WORKER_GITEA_TOKEN: set (40 chars) ``` …and then **pushed to the forge over SSH successfully.** The reason: ``` $ ssh -G git.ltms.dev user git port 2224 identityfile <a path outside ~/.ssh> ``` `~/.ssh/config` is four `Include` lines pointing into a shared-env directory, and the included files set `IdentityFile` for the forge. Checking that file: ``` forge private key file: EXISTS on disk readable by this user : yes permissions : -rw------- dai.ha passphrase-protected : NO — usable with no passphrase ``` A member runs as the same OS user. So it reads the key file directly. The agent is irrelevant to it. ## Why this matters 1. **`sshAuthSock: block` is cosmetic on this host.** It blanks a variable that SSH does not need. A member has full SSH access to the forge — and to every other host those included configs define an identity for, which is a large list. 2. **It is the same class of defect as #157**, which is what makes it worth a ticket rather than a note. The `memberCredentials` scrub removes **environment variables**. It cannot remove a **file**. #157 was a token in `.git/config`; this is a key in `~/Drive00/...`. Both walk straight past a control that only ever looked at the environment. 3. **The premise propagated into decisions.** The urgency argument for the #157 HTTPS work ("once the agent is blocked a member cannot push at all") was built on it. That work is still correct — it removes the token from git config, which was the real defect — but it was never the thing standing between a member and the forge. ## The general shape, again The enumeration keeps being of the wrong thing. #596 enumerated a file and missed an environment. This enumerated an environment and missed a file. A control that names a channel will always be walked around by the channel nobody named. The useful question is not "which variables do we blank?" but **"what can a process running as this user reach?"** — environment, files, sockets, keychain, and the ssh_config graph. ## Suggested direction Nothing here is a quick fix, and I have not implemented any of it. 1. **Report the gap rather than imply a guarantee.** At spawn under `allow-list`, resolve `ssh -G` for the configured forge host and, if it yields a readable `identityfile`, log a WARN saying SSH access is NOT blocked by this policy. The daemon should not let an operator believe `sshAuthSock: block` did something it did not. 2. **Decide whether member isolation is actually achievable as a same-user process.** It may not be. A member that runs as the operator's user can read anything the operator can. If real isolation is wanted, the answer is a separate OS user or a container, not a longer variable list. That is a design decision, not a bug fix — and worth taking before more effort goes into env-var scrubbing. 3. Until then, say plainly in the docs what the scrub does and does not cover, so the next decision is not made on the premise this ticket just disproved. ## What the flip DID achieve — not nothing Worth keeping in view. Under `allow-list` the daemon blanked 27 credential-shaped variables, measured: ``` AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, BESZEL_ADMIN_PASSWORD, BESZEL_KEY, BESZEL_UNIVERSAL_TOKEN, CF_API_TOKEN, CF_USER_TOKEN, CONFLUENCE_API_TOKEN, GITEA_ACCESS_TOKEN, GITLAB_OAUTH_CLIENT_SECRET, GITLAB_PERSONAL_ACCESS_TOKEN, GRAFANA_ADMIN_PASSWORD, HASS_TOKEN, HW_PASSWORD, JENKINS_MCP_AUTH, LLM_KEY_N8N, LTMS_API_KEY, METRICS_PUSH_TOKEN, N8N_ENCRYPTION_KEY, N8N_OWNER_PASSWORD, N8N_WEBHOOK_TOKEN, SSH_AUTH_SOCK, TELEGRAM_BOT_TOKEN, TS_API_KEY, TS_AUTHKEY, BRAIN_MCP_TOKEN, MEMORY_MCP_TOKEN ``` That includes the AWS `AdministratorAccess` pair from #144 — a genuine, large reduction. This ticket is about not overstating it. ## Related - #144 — the credential-exposure line. - #110 — the `SSH_AUTH_SOCK` decision this disproves the premise of. - #157 / #182 — a credential in a file, walking past an environment-only control.
Author
Owner

I put this to an architect as a design question — "what is the honest security boundary we can actually build, and what should we stop claiming?" — rather than as an implementation job. Its answer is below, checked and endorsed. I have started the top-ranked item.

1. What a same-uid member can reach today

The honest threat model: a same-user member is a trusted local process. It is not a sandbox.

Channel Path in Control today Real or theatre?
Environment The pane's login shell sources the operator's files; the member gets their exports. policy: allow-list runs a late ZDOTDIR scrub. Real for ambient inheritance. It does not protect the values — the member can source the same files again.
Files The member reads any owner-readable file. Proven path: ssh_config → IdentityFile → passphrase-free key. Also CLI auth stores and secret scripts. Worktree provisioning neutralizes selected config files. Not a boundary. File modes do not separate processes with the same uid. The member can use an absolute path to the original.
Sockets Any Unix socket the uid may open, including an ssh-agent socket whose path it learns. sshAuthSock: block blanks one env var. Theatre. It closes the default inheritance path only. It does not revoke access, and it leaves key files open.
Git config A linked worktree shares repo config; the same user reads global config, credential stores and other checkouts. Provisioning clears inherited credential.helper and rewrites a parsed SSH origin to HTTPS. Real routing, not enforcement. The member can change config, use another checkout, or call ssh directly.
Process table argv is world-visible; same-uid process inspection may expose more. fleetd keeps some secrets out of argv and redacts logs. No boundary. Secrets must never go in argv — and cross-uid isolation would not make argv safe either.

The concrete bad path, start to finish: the scrub blanks SSH_AUTH_SOCK → git starts ssh → ssh reads user config → it opens the readable IdentityFile → the forge accepts the operator's key. The gate closes lead → child environment inheritance. It does not close member → filesystem → credential. That is the one-way-gate shape again: it closes only the direction the original incident came from.

2. Is any control possible without changing uid?

No. Not through ordinary Unix permissions. Environment cleanup, worktrees, config rewrites and tool permissions reduce accidents. A member that can run commands bypasses all of them.

A real boundary inside one uid needs a different OS mechanism — mandatory access control, a container with narrow mounts, or a VM — denying the operator's home, lead sockets and process inspection, allowing only the worktree and scoped credentials. That is a new execution model, not a better name list.

The existing path is #185's second herdr under a different user. That protects files, sockets, process environment and process memory through OS permissions. It does not protect argv, and it does not protect shared git integrity where worktreeGroup grants access to shared objects and refs — a separate member-owned clone would be needed for that.

3. What we must stop claiming — started

The shipped fleetd.example.yaml currently says:

"Blocking it breaks git over SSH inside members (push/fetch authenticate as you)"

That is measured false. On 2026-08-28 a member with SSH_AUTH_SOCK blanked pushed to the forge over SSH successfully. The live fleetd.yaml was corrected at the time; the committed example never was, so the false claim is the one every new operator reads.

Delegated now. The corrected wording says what the setting does (omits the inherited agent path from the member environment), what it does not do (does not deny same-uid access to the socket, does not block readable key files, git over SSH may still work), and that the block stays because it is free — but is not a control.

The wiki needs the same treatment: memberCredentials controls ambient environment propagation, not credential access; a worktree isolates changes and normal git routing, not filesystem access; same-user mode assumes trusted members.

The architect also suggests a later rename to sshAgentEnv: omit|inherit, since block states more than it does. I have not done that — it is a schema change and needs its own ticket.

4. Ranked by value gained over work needed

  1. Write down the truth, and warn about same-uid at startup. Best ratio. Closes false assurance; closes no credential path. In progress.
  2. Document and operate memberHerdrSocket with a different user (#185). The smallest existing option that creates a real boundary. Acceptance: the operator's home and lead sockets unreadable; only scoped member inputs readable; the shared-git limit stated.
  3. Make the normal git route fail closed when the HTTPS rewrite fails. Stops accidental use of the operator's ssh identity on ordinary worktree pushes. Direct ssh stays open, so label it attribution and hygiene, not isolation.
  4. A sandbox/container/VM launcher. High value, high work. Its acceptance test must cover every outward direction: files, sockets, process inspection, network, writable mounts.

Scope of the architect's work

Read-only. No file changed, no test run, and it said plainly that no peer architect was reachable, so this is one checked position rather than an agreed one.

Keeping this open for items 2–4. Item 1 will be linked when it lands.

I put this to an architect as a design question — "what is the honest security boundary we can actually build, and what should we stop claiming?" — rather than as an implementation job. Its answer is below, checked and endorsed. I have started the top-ranked item. ## 1. What a same-uid member can reach today The honest threat model: **a same-user member is a trusted local process. It is not a sandbox.** | Channel | Path in | Control today | Real or theatre? | |---|---|---|---| | Environment | The pane's login shell sources the operator's files; the member gets their exports. | `policy: allow-list` runs a late `ZDOTDIR` scrub. | **Real** for ambient inheritance. It does not protect the values — the member can source the same files again. | | Files | The member reads any owner-readable file. Proven path: `ssh_config` → `IdentityFile` → passphrase-free key. Also CLI auth stores and secret scripts. | Worktree provisioning neutralizes selected config files. | **Not a boundary.** File modes do not separate processes with the same uid. The member can use an absolute path to the original. | | Sockets | Any Unix socket the uid may open, including an ssh-agent socket whose path it learns. | `sshAuthSock: block` blanks one env var. | **Theatre.** It closes the default inheritance path only. It does not revoke access, and it leaves key files open. | | Git config | A linked worktree shares repo config; the same user reads global config, credential stores and other checkouts. | Provisioning clears inherited `credential.helper` and rewrites a parsed SSH origin to HTTPS. | **Real routing, not enforcement.** The member can change config, use another checkout, or call `ssh` directly. | | Process table | argv is world-visible; same-uid process inspection may expose more. | fleetd keeps some secrets out of argv and redacts logs. | **No boundary.** Secrets must never go in argv — and cross-uid isolation would not make argv safe either. | The concrete bad path, start to finish: the scrub blanks `SSH_AUTH_SOCK` → git starts ssh → ssh reads user config → it opens the readable `IdentityFile` → the forge accepts the operator's key. **The gate closes lead → child environment inheritance. It does not close member → filesystem → credential.** That is the one-way-gate shape again: it closes only the direction the original incident came from. ## 2. Is any control possible without changing uid? **No.** Not through ordinary Unix permissions. Environment cleanup, worktrees, config rewrites and tool permissions reduce **accidents**. A member that can run commands bypasses all of them. A real boundary inside one uid needs a different OS mechanism — mandatory access control, a container with narrow mounts, or a VM — denying the operator's home, lead sockets and process inspection, allowing only the worktree and scoped credentials. That is a new execution model, not a better name list. The existing path is #185's second herdr under a different user. That protects files, sockets, process environment and process memory through OS permissions. It does **not** protect argv, and it does not protect shared git integrity where `worktreeGroup` grants access to shared objects and refs — a separate member-owned clone would be needed for that. ## 3. What we must stop claiming — started The shipped `fleetd.example.yaml` currently says: > "Blocking it breaks git over SSH inside members (push/fetch authenticate as you)" That is **measured false**. On 2026-08-28 a member with `SSH_AUTH_SOCK` blanked pushed to the forge over SSH successfully. The live `fleetd.yaml` was corrected at the time; the committed example never was, so the false claim is the one every new operator reads. Delegated now. The corrected wording says what the setting does (omits the inherited agent path from the member environment), what it does not do (does not deny same-uid access to the socket, does not block readable key files, git over SSH may still work), and that the block stays because it is free — but is not a control. The wiki needs the same treatment: `memberCredentials` controls ambient **environment propagation**, not credential **access**; a worktree isolates changes and normal git routing, not filesystem access; same-user mode assumes **trusted** members. The architect also suggests a later rename to `sshAgentEnv: omit|inherit`, since `block` states more than it does. I have not done that — it is a schema change and needs its own ticket. ## 4. Ranked by value gained over work needed 1. **Write down the truth, and warn about same-uid at startup.** Best ratio. Closes false assurance; closes no credential path. **In progress.** 2. **Document and operate `memberHerdrSocket` with a different user** (#185). The smallest existing option that creates a real boundary. Acceptance: the operator's home and lead sockets unreadable; only scoped member inputs readable; the shared-git limit stated. 3. **Make the normal git route fail closed when the HTTPS rewrite fails.** Stops accidental use of the operator's ssh identity on ordinary worktree pushes. Direct ssh stays open, so label it attribution and hygiene, not isolation. 4. **A sandbox/container/VM launcher.** High value, high work. Its acceptance test must cover every outward direction: files, sockets, process inspection, network, writable mounts. ## Scope of the architect's work Read-only. No file changed, no test run, and it said plainly that no peer architect was reachable, so this is one checked position rather than an agreed one. Keeping this open for items 2–4. Item 1 will be linked when it lands.
Author
Owner

Item 1 progress — the "write down the truth" half is done, the startup warning is in flight.

Done: the config no longer lies

PR #264, merged as b5ddbe5, with a follow-up in a2b8caf. The sshAuthSock entry in fleetd.example.yaml now says what the setting does (omits the inherited agent path), what it does not do (does not deny same-uid access to the socket, does not block readable key files, git over SSH may still work), and where the real boundary is.

The false claim it replaced — "Blocking it breaks git over SSH inside members" — had been in the shipped example since before the measurement that disproved it. The live fleetd.yaml was corrected on 2026-08-28; the committed example was not, so the false version is the one every new operator has been reading for a week.

One thing worth recording about the fix itself. The first cut removed the false claim and also removed a true one next to it: that SSH_AUTH_SOCK is a live handle to the agent, so a member holding it can sign with every key the agent holds. Without that sentence the entry reads as if the knob does not matter — and an operator has no reason left not to set it to allow.

So correcting an overclaim had produced an underclaim. I restored the sentence and split the two ideas: blocking it does not contain a member, but allowing it hands one a signing capability for no gain, so keep the block. The entry now also carries the measurement and the original mistake (looking only in ~/.ssh, which holds nothing but four Include lines).

Done: the wiki states the frame

Added a Features entry, "What a member can actually reach — the honest boundary". It carries the per-channel table from the design above, the proven path to the operator's key, and the plain statement that no meaningful boundary exists inside one uid.

It also names the recurring shape, because this is the third time it has come up here: a gate written after an incident closes only the direction that incident came from. Ask which states open it, not just which it blocked.

In flight: the startup line

A member is now working on the second half of item 1 — one INFO line at startup saying members run as the same OS user, what follows from that, and that memberHerdrSocket is the existing route to a real boundary.

The part I was careful to brief explicitly: when memberHerdrSocket IS set, the line must change and must not claim a boundary fleetd cannot verify. fleetd cannot see the uid of a process on the other end of a socket, so it can say members are routed to a separate herdr and that whether that herdr runs as a different user is the operator's to confirm. Replacing one false assurance with a new one would defeat the whole ticket.

Still open, unchanged

Items 2, 3 and 4 from the ranked list. Item 3 (make the normal git route fail closed when the HTTPS rewrite fails) is the next one with a real cost/benefit case, and it is attribution and hygiene rather than isolation — direct ssh stays open whatever we do there.

The sshAgentEnv: omit|inherit rename is not done. It is a schema change and wants its own ticket.

Item 1 progress — the "write down the truth" half is done, the startup warning is in flight. ## Done: the config no longer lies PR #264, merged as `b5ddbe5`, with a follow-up in `a2b8caf`. The `sshAuthSock` entry in `fleetd.example.yaml` now says what the setting does (omits the inherited agent path), what it does not do (does not deny same-uid access to the socket, does not block readable key files, git over SSH may still work), and where the real boundary is. The false claim it replaced — *"Blocking it breaks git over SSH inside members"* — had been in the shipped example since before the measurement that disproved it. The live `fleetd.yaml` was corrected on 2026-08-28; the committed example was not, so the false version is the one every new operator has been reading for a week. **One thing worth recording about the fix itself.** The first cut removed the false claim and also removed a true one next to it: that `SSH_AUTH_SOCK` is a live handle to the agent, so a member holding it can sign with every key the agent holds. Without that sentence the entry reads as if the knob does not matter — and an operator has no reason left not to set it to `allow`. So correcting an **overclaim** had produced an **underclaim**. I restored the sentence and split the two ideas: blocking it does not contain a member, but allowing it hands one a signing capability for no gain, so keep the block. The entry now also carries the measurement and the original mistake (looking only in `~/.ssh`, which holds nothing but four `Include` lines). ## Done: the wiki states the frame Added a Features entry, "What a member can actually reach — the honest boundary". It carries the per-channel table from the design above, the proven path to the operator's key, and the plain statement that no meaningful boundary exists inside one uid. It also names the recurring shape, because this is the third time it has come up here: **a gate written after an incident closes only the direction that incident came from.** Ask which states open it, not just which it blocked. ## In flight: the startup line A member is now working on the second half of item 1 — one INFO line at startup saying members run as the same OS user, what follows from that, and that `memberHerdrSocket` is the existing route to a real boundary. The part I was careful to brief explicitly: **when `memberHerdrSocket` IS set, the line must change and must not claim a boundary fleetd cannot verify.** fleetd cannot see the uid of a process on the other end of a socket, so it can say members are routed to a separate herdr and that whether that herdr runs as a different user is the operator's to confirm. Replacing one false assurance with a new one would defeat the whole ticket. ## Still open, unchanged Items 2, 3 and 4 from the ranked list. Item 3 (make the normal git route fail closed when the HTTPS rewrite fails) is the next one with a real cost/benefit case, and it is attribution and hygiene rather than isolation — direct `ssh` stays open whatever we do there. The `sshAgentEnv: omit|inherit` rename is not done. It is a schema change and wants its own ticket.
Author
Owner

Item 1 is done — merged as fa97f59 (PR #265). fleetd now logs the member trust model at startup, and the message changes when memberHerdrSocket is set. The unset branch says plainly that members run as the same OS user and can read any file that user can read, whatever memberCredentials says.

Also shipped earlier under this ticket: b5ddbe5 + a2b8caf removed the false claim that blocking SSH_AUTH_SOCK breaks git over SSH, and kept the true half — the socket is a live handle to the agent, so a member holding it can sign with every key the agent holds.

New item 5 — HerdrPeerLauncher makes the exact claim this ticket removes

Found by the worker on #265, confirmed by me in the code.

HerdrPeerLauncher.warnUnknownMemberEnvironment logs:

memberHerdrSocket is configured, so member panes run under a different OS user than fleetd's own process [...] The credential gap for member panes is UNKNOWN, not clean.

fleetd cannot see the uid at the other end of a unix socket. The "so" is an assumption, not a fact. The same claim is stated as a definition in the memberHerdrSocketConfigured() javadoc (line ~1630) and in the class javadoc (line ~120).

Why this matters, and which direction is harmful. If an operator points memberHerdrSocket at a second herdr running as the same user — a reasonable config, two herdr instances for pane isolation — then members inherit fleetd's own environment. The real credential gap is exactly the count the WARN just told the operator to disregard. A known exposure is reported as an unknown one. That is an under-report, and it is the direction that costs something.

It also now contradicts the line merged in this ticket: the startup report says "fleetd cannot see that herdr's uid", and this WARN says "so member panes run under a different OS user". Both go to the same log.

The fix is wording, not behaviour. The branch should still be taken — reporting on the wrong process is still the risk — but it must say "fleetd cannot confirm the uid, so treat the member environment as unknown", not "members run as a different user". Blast radius is small: no test asserts that WARN's text.

Items 2–4 remain open.

**Item 1 is done** — merged as `fa97f59` (PR #265). `fleetd` now logs the member trust model at startup, and the message changes when `memberHerdrSocket` is set. The unset branch says plainly that members run as the same OS user and can read any file that user can read, whatever `memberCredentials` says. Also shipped earlier under this ticket: `b5ddbe5` + `a2b8caf` removed the false claim that blocking `SSH_AUTH_SOCK` breaks git over SSH, and kept the true half — the socket is a live handle to the agent, so a member holding it can sign with every key the agent holds. ## New item 5 — `HerdrPeerLauncher` makes the exact claim this ticket removes Found by the worker on #265, confirmed by me in the code. `HerdrPeerLauncher.warnUnknownMemberEnvironment` logs: > memberHerdrSocket is configured, **so member panes run under a different OS user than fleetd's own process** [...] The credential gap for member panes is UNKNOWN, not clean. `fleetd` cannot see the uid at the other end of a unix socket. The "so" is an assumption, not a fact. The same claim is stated as a definition in the `memberHerdrSocketConfigured()` javadoc (line ~1630) and in the class javadoc (line ~120). **Why this matters, and which direction is harmful.** If an operator points `memberHerdrSocket` at a second herdr running as the *same* user — a reasonable config, two herdr instances for pane isolation — then members inherit `fleetd`'s own environment. The real credential gap is exactly the count the WARN just told the operator to disregard. A *known* exposure is reported as an *unknown* one. That is an under-report, and it is the direction that costs something. It also now contradicts the line merged in this ticket: the startup report says "fleetd cannot see that herdr's uid", and this WARN says "so member panes run under a different OS user". Both go to the same log. The fix is wording, not behaviour. The branch should still be taken — reporting on the wrong process is still the risk — but it must say *"fleetd cannot confirm the uid, so treat the member environment as unknown"*, not *"members run as a different user"*. Blast radius is small: no test asserts that WARN's text. Items 2–4 remain open.
Author
Owner

Split out #266 — memberCredentials.sshAuthSock's value names (block/allow) make the same kind of claim this ticket is about, one layer down. block names an effect fleetd does not have: it omits the variable, it does not deny same-user access to the socket.

Kept out of this ticket because it is a config rename needing a read-both shim, not a prose fix. It carries a real risk this ticket does not: the live config on this host uses the old spelling, so the shim is load-bearing.

Item 5 (the HerdrPeerLauncher uid claim) is delegated and in progress.

Split out #266 — `memberCredentials.sshAuthSock`'s value names (`block`/`allow`) make the same kind of claim this ticket is about, one layer down. `block` names an effect fleetd does not have: it omits the variable, it does not deny same-user access to the socket. Kept out of this ticket because it is a config rename needing a read-both shim, not a prose fix. It carries a real risk this ticket does not: the live config on this host uses the old spelling, so the shim is load-bearing. Item 5 (the `HerdrPeerLauncher` uid claim) is delegated and in progress.
Author
Owner

Item 5 is done — merged as 27aefbf (PR #269). All four sites in HerdrPeerLauncher now say fleetd cannot confirm what OS user the second herdr runs as, instead of asserting it does. No behaviour change. The WARN still states the honest conclusion, so this did not trade an overclaim for an underclaim.

I re-ran the mutation proof myself rather than trusting the report: reintroducing the old false claim turns the new test red. Restored, git diff --exit-code clean, full build 1265 tests green at that point (1269 after the later merges).

On the 10 further "same shape" sites the implementer reported

I checked three of them in the code — GitWorktrees.shareRootWithGroup, EnvAllowListScrub.shareWithGroup, and ClaudeCodeLauncher.writeCharterFile. I am not filing them, and I think that is the right call.

Every one of them assumes a different OS user and therefore does more accommodation: share the group, make worktreeRoot traversable, keep the charter out of fleetd's 0700 tmpdir. If the assumption is wrong and the member is the same user, the extra work is harmless — it was already reachable. The claim is inaccurate, but being wrong costs nothing.

That is the opposite of the WARN this ticket just fixed, where the wrong assumption caused a real exposure to be reported as unknown. The direction is what makes it a defect, not the phrasing.

Worth recording because two workers disagreed here. A separate audit of exactly this shape, run in parallel, read EnvAllowListScrub in full and GitWorktrees in part and reported no findings — it was applying a "being wrong must cost something" filter. The implementer's brief asked it to report the shape without repeating that filter, so it matched on the phrase. The audit was right; the difference came from my briefing, not from either worker's care.

The remaining sites are still inaccurate, and a future maintainer could lean on the premise in the harmful direction. Cheap mitigation rather than a fan-out: soften them whenever those files are next touched, matching Fleetd.reportMemberTrustModel.

Items 2–4 remain open. #266 (the sshAuthSock value rename) is done and closed.

**Item 5 is done** — merged as `27aefbf` (PR #269). All four sites in `HerdrPeerLauncher` now say fleetd cannot confirm what OS user the second herdr runs as, instead of asserting it does. No behaviour change. The WARN still states the honest conclusion, so this did not trade an overclaim for an underclaim. I re-ran the mutation proof myself rather than trusting the report: reintroducing the old false claim turns the new test red. Restored, `git diff --exit-code` clean, full build 1265 tests green at that point (1269 after the later merges). ## On the 10 further "same shape" sites the implementer reported I checked three of them in the code — `GitWorktrees.shareRootWithGroup`, `EnvAllowListScrub.shareWithGroup`, and `ClaudeCodeLauncher.writeCharterFile`. **I am not filing them, and I think that is the right call.** Every one of them assumes a different OS user and therefore does *more* accommodation: share the group, make `worktreeRoot` traversable, keep the charter out of fleetd's 0700 tmpdir. If the assumption is wrong and the member is the same user, the extra work is harmless — it was already reachable. The claim is inaccurate, but being wrong costs nothing. That is the opposite of the WARN this ticket just fixed, where the wrong assumption caused a real exposure to be reported as unknown. **The direction is what makes it a defect, not the phrasing.** Worth recording because two workers disagreed here. A separate audit of exactly this shape, run in parallel, read `EnvAllowListScrub` in full and `GitWorktrees` in part and reported **no findings** — it was applying a "being wrong must cost something" filter. The implementer's brief asked it to report the shape without repeating that filter, so it matched on the phrase. The audit was right; the difference came from my briefing, not from either worker's care. The remaining sites are still inaccurate, and a future maintainer could lean on the premise in the harmful direction. Cheap mitigation rather than a fan-out: soften them whenever those files are next touched, matching `Fleetd.reportMemberTrustModel`. Items 2–4 remain open. #266 (the `sshAuthSock` value rename) is done and closed.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#184