seedTrustDialog writes the trust seed to fleetd's own home when memberHerdrSocket is set, so the member never sees it #285

Closed
opened 2026-09-04 05:39:50 +02:00 by ltms · 1 comment
Owner

Found by an audit comparing the launchers against each other. The auditor was honest about its confidence, and I have kept that distinction below.

The asymmetry

ClaudeCodeLauncher.seedTrustDialog (:548-632) writes the workspace-trust seed to configDir/.claude.json, or to ~/.claude.json — fleetd's own home — when configDir is unset. It gates only on isProvisionedWorktree(cwd) (:549). It never checks memberHerdrSocketConfigured(), and being static it structurally cannot.

Its sibling one method below does check. writeCharterFile (:788) tests memberHerdrSocketConfigured() at :790, routes the file under worktreeRoot, shares it with EnvAllowListScrub.shareWithGroup (:813), and refuses the spawn outright when worktreeRoot/worktreeGroup are not set. OpenCodeLauncher.writeConfig/configParentDir (:422-559) do the same for opencode.json.

So of the member-facing cross-process file writes in this class, seedTrustDialog is the one that skips the check.

It is made worse by writeAtomically → copyPosixPermissionsIfPresent (:696-710), which preserves the target's existing 0600 — Claude Code's own default — so even a shared directory would not help without a permission change.

The path and the consequence

A claude profile with memberHerdrSocket: set and configDir: unset, spawned with worktree:true. buildLaunch calls seedTrustDialog unconditionally at :290, before herdr is touched.

Under memberHerdrSocket the member runs as a different OS user with its own $HOME. The seed lands in fleetd's home, not the member's. The member never sees it, hits Claude Code's un-timed interactive trust dialog, never mounts the bridge MCP, and the readiness gate times it out — which is the fleetd #149 incident this method exists to prevent, brought back for this one config combination.

The symptom an operator would see is the familiar useless one: "did not reach injectable state", with nothing pointing at the trust dialog.

Confidence — read this before you start

High on the code path. I confirmed the contrast between seedTrustDialog and writeCharterFile myself.

Not confirmed on live reachability. Other javadoc in HerdrPeerLauncher describes "memberHerdrSocket ABSENT" as today's only live mode, so the running fleet may never take this path. I did not verify that, and the live fleetd.yaml is not in git.

So step 1 is the reachability question, not the fix. Establish whether a claude profile can actually be spawned with memberHerdrSocket set and configDir unset. If the config makes that combination impossible, close this with the explanation and write no code. If it is possible only in a mode nobody runs, say that too and let me decide — the feature has had three fixes for this same seam already (#213, #219, #222), which is the argument for fixing it even if it is not live today.

If it is reachable

Make seedTrustDialog follow the same rule as writeCharterFile: under memberHerdrSocket, place the file where the member's OS user can read it and share it with the configured group, or refuse the spawn with a message naming what is missing. A refusal is much better than the silent timeout this produces now — the same reasoning as fleetd #155.

Note it will have to stop being static to see that state. That is a real change to the class; say what you changed and confirm the non-memberHerdrSocket path behaves exactly as before.

Also reported, not fixed

  • writeCharterFile (claude) and writeConfig (opencode) both create per-spawn temp files and directories that rely only on deleteOnExit. spawnInternal explicitly deletes the CB-633 ZDOTDIR directory on a post-buildLaunch spawn failure (HerdrPeerLauncher.java:519-528); these two do not. For a long-running daemon deleteOnExit is effectively never. Symmetric across both launchers, disk clutter only. Do not fix it in this ticket — mention it in your report and I will file it separately if it is worth it.
  • writeIdeOverlay also skips memberHerdrSocketConfigured(), but writes into cwd (the worktree), which provisioning may already share with the member's user. Unverified, outside this scope.
Found by an audit comparing the launchers against each other. The auditor was honest about its confidence, and I have kept that distinction below. ## The asymmetry `ClaudeCodeLauncher.seedTrustDialog` (`:548-632`) writes the workspace-trust seed to `configDir/.claude.json`, or to `~/.claude.json` — **fleetd's own home** — when `configDir` is unset. It gates only on `isProvisionedWorktree(cwd)` (`:549`). It never checks `memberHerdrSocketConfigured()`, and being `static` it structurally cannot. Its sibling one method below does check. `writeCharterFile` (`:788`) tests `memberHerdrSocketConfigured()` at `:790`, routes the file under `worktreeRoot`, shares it with `EnvAllowListScrub.shareWithGroup` (`:813`), and refuses the spawn outright when `worktreeRoot`/`worktreeGroup` are not set. `OpenCodeLauncher.writeConfig`/`configParentDir` (`:422-559`) do the same for `opencode.json`. So of the member-facing cross-process file writes in this class, `seedTrustDialog` is the one that skips the check. It is made worse by `writeAtomically` → `copyPosixPermissionsIfPresent` (`:696-710`), which preserves the target's existing `0600` — Claude Code's own default — so even a shared directory would not help without a permission change. ## The path and the consequence A `claude` profile with `memberHerdrSocket:` set and `configDir:` unset, spawned with `worktree:true`. `buildLaunch` calls `seedTrustDialog` unconditionally at `:290`, before herdr is touched. Under `memberHerdrSocket` the member runs as a **different OS user with its own `$HOME`**. The seed lands in fleetd's home, not the member's. The member never sees it, hits Claude Code's un-timed interactive trust dialog, never mounts the bridge MCP, and the readiness gate times it out — which is the fleetd #149 incident this method exists to prevent, brought back for this one config combination. The symptom an operator would see is the familiar useless one: "did not reach injectable state", with nothing pointing at the trust dialog. ## Confidence — read this before you start **High on the code path.** I confirmed the contrast between `seedTrustDialog` and `writeCharterFile` myself. **Not confirmed on live reachability.** Other javadoc in `HerdrPeerLauncher` describes "`memberHerdrSocket` ABSENT" as today's only live mode, so the running fleet may never take this path. I did not verify that, and the live `fleetd.yaml` is not in git. **So step 1 is the reachability question, not the fix.** Establish whether a `claude` profile can actually be spawned with `memberHerdrSocket` set and `configDir` unset. If the config makes that combination impossible, close this with the explanation and write no code. If it is possible only in a mode nobody runs, say that too and let me decide — the feature has had three fixes for this same seam already (#213, #219, #222), which is the argument for fixing it even if it is not live today. ## If it is reachable Make `seedTrustDialog` follow the same rule as `writeCharterFile`: under `memberHerdrSocket`, place the file where the member's OS user can read it and share it with the configured group, or refuse the spawn with a message naming what is missing. A refusal is much better than the silent timeout this produces now — the same reasoning as fleetd #155. Note it will have to stop being `static` to see that state. That is a real change to the class; say what you changed and confirm the non-`memberHerdrSocket` path behaves exactly as before. ## Also reported, not fixed - `writeCharterFile` (claude) and `writeConfig` (opencode) both create per-spawn temp files and directories that rely only on `deleteOnExit`. `spawnInternal` explicitly deletes the CB-633 ZDOTDIR directory on a post-`buildLaunch` spawn failure (`HerdrPeerLauncher.java:519-528`); these two do not. For a long-running daemon `deleteOnExit` is effectively never. Symmetric across both launchers, disk clutter only. **Do not fix it in this ticket** — mention it in your report and I will file it separately if it is worth it. - `writeIdeOverlay` also skips `memberHerdrSocketConfigured()`, but writes into `cwd` (the worktree), which provisioning may already share with the member's user. Unverified, outside this scope.
Author
Owner

Merged as 75b1508, corrected in 94f50e5. Real merge built green at 1296 tests.

The defect is real and reachable in code, as the worker reported. seedTrustDialog was static, so it could not see memberHerdrSocketConfigured() at all — while writeCharterFile, one method below, already refuses a spawn when it cannot place a file where a different-uid member can read it. Same file, same problem, one guard.

My mutations

Removed the refusal block:

seedTrustDialogUnderMemberHerdrSocketRefusesWhenConfigDirUnset:2801
  expected: <java.lang.IllegalStateException> but was: <java.io.UncheckedIOException>
seedTrustDialogUnderMemberHerdrSocketRefusesWhenWorktreeGroupUnset:2830
  expected: <java.lang.IllegalStateException> but was: <java.lang.NullPointerException>

Removed the group share:

seedTrustDialogUnderMemberHerdrSocketSharesTheFileWithTheConfiguredGroup:2774
  ... a 0600 file (Claude Code's own default) is unreadable by the member's different OS user
  ==> expected: <rw-r-----> but was: <rw------->

Both restored; tree clean.

What I checked beyond the tests

The refusal-not-degradation call is right, and matches the sibling. A member with no readable trust seed is not "slightly worse" — it sits on the interactive dialog forever and never calls fleet_reply, which is the exact failure the whole workspace-trust feature exists to prevent.

The test that matters most is the regression one:
seedTrustDialogPreservesExisting0600PermissionsWhenMemberHerdrSocketAbsentEvenWithALiveConfigSupplier. It uses a live, non-null config supplier rather than config == null, which is what every other seed test in that file uses — so it actually exercises the path the live fleet takes today and proves nothing about it changed. With memberHerdrSocket absent, this fix is inert.

The bug test redirects user.home to a @TempDir before driving the old code path, so a reintroduced defect still cannot touch the operator's real ~/.claude.json. That matters here: this file has already been destroyed once by a mutation test.

I checked the control-flow rewrite (return inside the retry loop became a written flag) is behaviour-preserving, and that the added return in the swallow-all catch is what keeps a failed write from being followed by a chgrp.

My correction — one helper, not two

The worker added a private shareTrustJsonWithGroup to ClaudeCodeLauncher. Not reusing EnvAllowListScrub.shareWithGroup was right — that one takes a directory and rewrites every entry in it, which would have chmod'd all of configDir. But the copy dropped something. shareWithGroup catches two exception types:

} catch (IOException e) {          ...
} catch (UnsupportedOperationException e) {   // "this filesystem does not support POSIX group ownership"

The copy caught only IOException. view.setGroup throws UnsupportedOperationException on a filesystem without POSIX group ownership, and that is not an IOException — so the copy would have thrown a raw runtime exception where its sibling gives the operator a sentence naming the cause.

Moved into EnvAllowListScrub as shareFileWithGroup(Path, String), sitting next to shareWithGroup, reusing the same private setGroupAndPermissions and carrying both catches. One copy of the chgrp+chmod rule, one error vocabulary.

Not verified

I have not run this on a live memberHerdrSocket fleet — no such fleet is running here. The worker said the same about its own finding, and that is the honest limit on both sides: this is proven in code and in tests, not in production.

Wiki entry added to A claude-code member no longer blocks forever on the workspace-trust dialog as its third gotcha.

Merged as `75b1508`, corrected in `94f50e5`. Real merge built green at **1296 tests**. The defect is real and reachable in code, as the worker reported. `seedTrustDialog` was `static`, so it could not see `memberHerdrSocketConfigured()` at all — while `writeCharterFile`, one method below, already refuses a spawn when it cannot place a file where a different-uid member can read it. Same file, same problem, one guard. ## My mutations **Removed the refusal block:** ``` seedTrustDialogUnderMemberHerdrSocketRefusesWhenConfigDirUnset:2801 expected: <java.lang.IllegalStateException> but was: <java.io.UncheckedIOException> seedTrustDialogUnderMemberHerdrSocketRefusesWhenWorktreeGroupUnset:2830 expected: <java.lang.IllegalStateException> but was: <java.lang.NullPointerException> ``` **Removed the group share:** ``` seedTrustDialogUnderMemberHerdrSocketSharesTheFileWithTheConfiguredGroup:2774 ... a 0600 file (Claude Code's own default) is unreadable by the member's different OS user ==> expected: <rw-r-----> but was: <rw-------> ``` Both restored; tree clean. ## What I checked beyond the tests The refusal-not-degradation call is right, and matches the sibling. A member with no readable trust seed is not "slightly worse" — it sits on the interactive dialog forever and never calls `fleet_reply`, which is the exact failure the whole workspace-trust feature exists to prevent. The test that matters most is the regression one: `seedTrustDialogPreservesExisting0600PermissionsWhenMemberHerdrSocketAbsentEvenWithALiveConfigSupplier`. It uses a **live, non-null** config supplier rather than `config == null`, which is what every other seed test in that file uses — so it actually exercises the path the live fleet takes today and proves nothing about it changed. With `memberHerdrSocket` absent, this fix is inert. The bug test redirects `user.home` to a `@TempDir` before driving the old code path, so a reintroduced defect still cannot touch the operator's real `~/.claude.json`. That matters here: this file has already been destroyed once by a mutation test. I checked the control-flow rewrite (`return` inside the retry loop became a `written` flag) is behaviour-preserving, and that the added `return` in the swallow-all catch is what keeps a failed write from being followed by a chgrp. ## My correction — one helper, not two The worker added a private `shareTrustJsonWithGroup` to `ClaudeCodeLauncher`. Not reusing `EnvAllowListScrub.shareWithGroup` was right — that one takes a *directory* and rewrites every entry in it, which would have chmod'd all of `configDir`. But the copy dropped something. `shareWithGroup` catches **two** exception types: ```java } catch (IOException e) { ... } catch (UnsupportedOperationException e) { // "this filesystem does not support POSIX group ownership" ``` The copy caught only `IOException`. `view.setGroup` throws `UnsupportedOperationException` on a filesystem without POSIX group ownership, and that is not an `IOException` — so the copy would have thrown a raw runtime exception where its sibling gives the operator a sentence naming the cause. Moved into `EnvAllowListScrub` as `shareFileWithGroup(Path, String)`, sitting next to `shareWithGroup`, reusing the same private `setGroupAndPermissions` and carrying both catches. One copy of the chgrp+chmod rule, one error vocabulary. ## Not verified I have **not** run this on a live `memberHerdrSocket` fleet — no such fleet is running here. The worker said the same about its own finding, and that is the honest limit on both sides: this is proven in code and in tests, not in production. Wiki entry added to *A claude-code member no longer blocks forever on the workspace-trust dialog* as its third gotcha.
ltms closed this issue 2026-09-04 06:14:08 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#285