Run members on a second herdr under a dedicated user (optional; single-herdr stays the default) #185

Open
opened 2026-08-28 04:05:56 +02:00 by ltms · 3 comments
Owner

Follow-up to #184. A member runs as the operator's OS user, so it reads the operator's forge ssh key — a readable, passphrase-free file. Env scrubbing cannot fix that. This ticket is the fix that can: run member panes under a different uid.

Measured end to end on fleet01 on 2026-08-28. The host was restored to its original state afterwards.

The one-line answer

fleetd never forks a process. It sends {cwd, env, argv} to herdr, and herdr forks the PTY as its own uid. herdr's socket API has no user/uid/run-as parameter on workspace.create, tab.create, pane.split or agent.start — confirmed in herdr's own Socket API docs, not only in our client. So the only way to change a member's uid is a second herdr server running as that user.

What herdr supports (docs)

Socket path resolution order: --session <name> → HERDR_SOCKET_PATH → HERDR_SESSION → default ~/.config/herdr/herdr.sock. Named sessions live at ~/.config/herdr/sessions/<name>/herdr.sock. fleet01 already runs a named session (fleet01).

fleetd already has the knob: FleetConfig.herdrSocket, plus a HERDR_SOCKET_PATH override in UnixSocketHerdrClient.defaultSocketPath().

Measured on fleet01 (herdr 0.8.0)

E1 — two servers, two users, two sockets, at once. PASS.

srw------- 1 ltms     fleetipc  /run/fleet/A.sock
srw------- 1 fleetmbr fleetipc  /run/fleet/B.sock

Each server also mints a -client.sock beside its api socket.

E2 — the socket is mode 0600. This is a real blocker.

$ HERDR_SOCKET_PATH=/run/fleet/B.sock herdr workspace list      # as ltms
Error: Os { code: 13, kind: PermissionDenied }

$ sudo chmod g+rw /run/fleet/B.sock ; retry
{"result":{"type":"workspace_list","workspaces":[]}}            # works

umask 007 before starting the server does not change it (verified with a third server) — herdr sets 0600 explicitly. No config key for socket mode was found. So the member herdr's start unit must chmod g+rw after the socket appears, which is a race window and an operational wart. Worth asking herdr upstream for a socket-mode or socket-group option.

E3 — the isolation works. This is the #184 fix. PASS.

$ sudo -u fleetmbr cat /home/ltms/exp-secret.txt
cat: /home/ltms/exp-secret.txt: Permission denied
$ sudo -u fleetmbr ls /home/ltms/
ls: cannot open directory '/home/ltms/': Permission denied

And the pane really is that user — herdr pane read on server B returned the prompt fleetmbr@fleet01:~$, with cwd: /home/fleetmbr.

E4 — id collision is guaranteed, not merely possible. This is the biggest code consequence.

Workspace, tab and pane ids are per-daemon sequential counters. Both servers were live with the same ids at the same moment:

server A: workspaces [('w1','~'), ('w2','exp-A')]   panes [w1:p3, w1:p1, w2:p1]
server B: workspaces [('w1','exp-B')]               panes [w1:p1]      <-- same id, different pane, different user

terminal_id is different in shape — a timestamp prefix plus random suffix (term_65a11bc2731862 vs term_65a11c25eba401) — and stayed distinct. But that is probabilistic, not structural, and the prefix collides for panes created in the same second.

fleet_stop{paneId} takes a paneId. Under two daemons w1:p1 is ambiguous. Pane ids appear in the MCP surface and in the roster, so they must become daemon-qualified.

Design

Not two global clients — a router keyed by target. Fleetd.java:151 builds one UnixSocketHerdrClient and hands agents + spaces to about fifteen consumers, which split by who they act on:

Side Consumers
lead herdr (operator uid) LeadTabScanner (identity), ReplyPushLoop, LeadHeartbeatLoop, LeadCoordLoop, LeadLauncher
member herdr (member uid) HerdrPeerLauncher + both launchers, StatusPoller, StatusRefiner, CompletionResolver, PaneLocator, FleetHealthMonitor
both — need the router Injector, MessageService, FleetApp

Single-herdr stays the default. A new optional key:

herdrSocket: ~/.config/herdr/herdr.sock   # lead + members
memberHerdrSocket:                        # absent => same client for everything

Absent means the router returns one client for every target, so behaviour is byte-identical to today. The Mac runs that way. Two-herdr mode is opt-in per host.

Risks, ordered

  1. Lead identity. LeadTabScanner.scan() runs workspace.list + tab.list + pane.list on one client. It must be pinned to the lead's herdr. Point it at the member herdr and the lead is demoted to worker, and every orchestration call is refused.
  2. Pane/tab/workspace id collision — proven above. Qualify the ids. Do not rely on terminal_id entropy.
  3. Worktree ownership. GitWorktrees shells git worktree add via ProcessBuilder, so it runs as fleetd's uid and creates the tree owned by the operator. The member then writes there as another user, and committing writes into the operator's .git/objects and .git/refs regardless. Needs a shared group and setgid. This isolates credentials, not the repo — say so plainly rather than implying more.
  4. HerdrPeerLauncher.hostEnvNames reads fleetd's own environment and assumes the pane's login shell exports the same set. Under a second user that is false, so CB-596's gap detector reports on the wrong environment.
  5. CB-634's shared fleet workspace cannot span two daemons. Only affects two-herdr mode; members get their own workspace there.
  6. Socket mode race — see E2.

The payoff, beyond #184

Today the member's login shell re-sources the operator's secrets.sh, which is the whole reason memberCredentials exists. A second user has nothing to source. tab.create{env} becomes the only path a credential can reach a member — the spawn boundary we kept trying to build inside a sourced file. The scrub stops being the control and becomes defence in depth.

Staging

  1. Router + memberHerdrSocket (absent = today). Qualify pane ids. Pin LeadTabScanner.
  2. Fix hostEnvNames for the two-user case.
  3. Worktree group/setgid provisioning.
  4. Roll out on fleet01 first — it already runs headless and has no GUI session to worry about.

Related

  • #184 — why this exists.
  • #110 — the sshAuthSock: block decision whose premise #184 disproved.
  • #157 / #182 — credentials reachable past an environment-only control.
Follow-up to #184. A member runs as the operator's OS user, so it reads the operator's forge ssh key — a readable, passphrase-free file. Env scrubbing cannot fix that. This ticket is the fix that can: run member panes under a different uid. Measured end to end on fleet01 on 2026-08-28. The host was restored to its original state afterwards. ## The one-line answer fleetd never forks a process. It sends `{cwd, env, argv}` to herdr, and herdr forks the PTY as **its own** uid. herdr's socket API has no user/uid/run-as parameter on `workspace.create`, `tab.create`, `pane.split` or `agent.start` — confirmed in herdr's own Socket API docs, not only in our client. So the only way to change a member's uid is **a second herdr server running as that user**. ## What herdr supports (docs) Socket path resolution order: `--session <name>` → `HERDR_SOCKET_PATH` → `HERDR_SESSION` → default `~/.config/herdr/herdr.sock`. Named sessions live at `~/.config/herdr/sessions/<name>/herdr.sock`. fleet01 already runs a named session (`fleet01`). fleetd already has the knob: `FleetConfig.herdrSocket`, plus a `HERDR_SOCKET_PATH` override in `UnixSocketHerdrClient.defaultSocketPath()`. ## Measured on fleet01 (herdr 0.8.0) **E1 — two servers, two users, two sockets, at once. PASS.** ``` srw------- 1 ltms fleetipc /run/fleet/A.sock srw------- 1 fleetmbr fleetipc /run/fleet/B.sock ``` Each server also mints a `-client.sock` beside its api socket. **E2 — the socket is mode 0600. This is a real blocker.** ``` $ HERDR_SOCKET_PATH=/run/fleet/B.sock herdr workspace list # as ltms Error: Os { code: 13, kind: PermissionDenied } $ sudo chmod g+rw /run/fleet/B.sock ; retry {"result":{"type":"workspace_list","workspaces":[]}} # works ``` `umask 007` before starting the server does **not** change it (verified with a third server) — herdr sets 0600 explicitly. No config key for socket mode was found. So the member herdr's start unit must `chmod g+rw` **after** the socket appears, which is a race window and an operational wart. Worth asking herdr upstream for a socket-mode or socket-group option. **E3 — the isolation works. This is the #184 fix. PASS.** ``` $ sudo -u fleetmbr cat /home/ltms/exp-secret.txt cat: /home/ltms/exp-secret.txt: Permission denied $ sudo -u fleetmbr ls /home/ltms/ ls: cannot open directory '/home/ltms/': Permission denied ``` And the pane really is that user — `herdr pane read` on server B returned the prompt `fleetmbr@fleet01:~$`, with `cwd: /home/fleetmbr`. **E4 — id collision is guaranteed, not merely possible. This is the biggest code consequence.** Workspace, tab and pane ids are **per-daemon sequential counters**. Both servers were live with the same ids at the same moment: ``` server A: workspaces [('w1','~'), ('w2','exp-A')] panes [w1:p3, w1:p1, w2:p1] server B: workspaces [('w1','exp-B')] panes [w1:p1] <-- same id, different pane, different user ``` `terminal_id` is different in shape — a timestamp prefix plus random suffix (`term_65a11bc2731862` vs `term_65a11c25eba401`) — and stayed distinct. But that is probabilistic, not structural, and the prefix collides for panes created in the same second. **`fleet_stop{paneId}` takes a paneId.** Under two daemons `w1:p1` is ambiguous. Pane ids appear in the MCP surface and in the roster, so they must become daemon-qualified. ## Design Not two global clients — a **router keyed by target**. `Fleetd.java:151` builds one `UnixSocketHerdrClient` and hands `agents` + `spaces` to about fifteen consumers, which split by who they act on: | Side | Consumers | |---|---| | lead herdr (operator uid) | `LeadTabScanner` (identity), `ReplyPushLoop`, `LeadHeartbeatLoop`, `LeadCoordLoop`, `LeadLauncher` | | member herdr (member uid) | `HerdrPeerLauncher` + both launchers, `StatusPoller`, `StatusRefiner`, `CompletionResolver`, `PaneLocator`, `FleetHealthMonitor` | | both — need the router | `Injector`, `MessageService`, `FleetApp` | **Single-herdr stays the default.** A new optional key: ```yaml herdrSocket: ~/.config/herdr/herdr.sock # lead + members memberHerdrSocket: # absent => same client for everything ``` Absent means the router returns one client for every target, so behaviour is byte-identical to today. The Mac runs that way. Two-herdr mode is opt-in per host. ## Risks, ordered 1. **Lead identity.** `LeadTabScanner.scan()` runs `workspace.list` + `tab.list` + `pane.list` on one client. It must be pinned to the **lead's** herdr. Point it at the member herdr and the lead is demoted to worker, and every orchestration call is refused. 2. **Pane/tab/workspace id collision** — proven above. Qualify the ids. Do not rely on `terminal_id` entropy. 3. **Worktree ownership.** `GitWorktrees` shells `git worktree add` via `ProcessBuilder`, so it runs as fleetd's uid and creates the tree owned by the operator. The member then writes there as another user, and committing writes into the operator's `.git/objects` and `.git/refs` regardless. Needs a shared group and setgid. **This isolates credentials, not the repo** — say so plainly rather than implying more. 4. **`HerdrPeerLauncher.hostEnvNames`** reads fleetd's own environment and assumes the pane's login shell exports the same set. Under a second user that is false, so CB-596's gap detector reports on the wrong environment. 5. **CB-634's shared `fleet` workspace** cannot span two daemons. Only affects two-herdr mode; members get their own workspace there. 6. **Socket mode race** — see E2. ## The payoff, beyond #184 Today the member's login shell re-sources the operator's `secrets.sh`, which is the whole reason `memberCredentials` exists. A second user has nothing to source. `tab.create{env}` becomes the **only** path a credential can reach a member — the spawn boundary we kept trying to build inside a sourced file. The scrub stops being the control and becomes defence in depth. ## Staging 1. Router + `memberHerdrSocket` (absent = today). Qualify pane ids. Pin `LeadTabScanner`. 2. Fix `hostEnvNames` for the two-user case. 3. Worktree group/setgid provisioning. 4. Roll out on fleet01 first — it already runs headless and has no GUI session to worry about. ## Related - #184 — why this exists. - #110 — the `sshAuthSock: block` decision whose premise #184 disproved. - #157 / #182 — credentials reachable past an environment-only control.
Author
Owner

Must be fixed BEFORE memberHerdrSocket is ever set — running list

Two PRs have now landed or been reviewed for this epic, and both had the same shape of defect:
correct today only because both herdr clients are the same object, wrong the moment the second
daemon is configured.
Every test passed in each case. Keeping the list in one place so nothing
enables the feature on top of a half-routed daemon.

Landed

  • #187 — stop() / list() pane-id ambiguity. Merged as a237fbf. Includes three lead-review
    fixes: stop() no longer forgets the owner when the delegate refuses, list() is keyed by
    (owning daemon, pane id) instead of the raw pane id, and the class javadoc no longer states the
    single-connection premise as fact. 993 tests green.

Open, blocking

  • #186 — three unrouted consumers. Sent back to a member. In short:

    1. PaneLocator is pinned to the member daemon, so a lead's connection resolves to no terminal.
      FleetMcp.reply and FleetMcp.ask both refuse when callerTerminal == null, so a lead can no
      longer answer a peer lead, and fleet_whoami loses its leader name.
    2. StatusRefiner is pinned to the member daemon while the status call next to it is routed per
      target. A lead whose raw status reads UNKNOWN gets refined against the wrong daemon and stays
      UNKNOWN, which wedges status-gated delivery to that lead.
    3. FleetApp never goes through the router: /healthz stays green while the MEMBER daemon is
      down (and then every spawn fails), and GET /sessions drops every member workspace.
  • Pane ownership does not survive a restart. CompositePeerLauncher.spawnedBy is in memory. With
    two daemons, every pane that outlives the daemon is unowned, so fleet_stop on it refuses — by
    design, since guessing would close a pane on an arbitrary daemon. Recovering ownership at boot is
    its own unit and it is a prerequisite, not a nicety.

  • The socket is mode 0600. Cross-user access needs a post-start chmod g+rw, and umask does
    not change it. Measured on herdr 0.8.0. An upstream request is drafted; until then any start unit
    has to poll for the socket and relax it, which is a race on every start.

Lower severity, do not lose

  • HerdrPeerLauncher.stop has the same before-the-call removal that #187 just fixed one level up:
    paneByAgentId.remove(idOrPane) runs before agents.close(paneId). A failed stop there falls back
    to raw-pane addressing rather than refusing, so it degrades instead of breaking.

The pattern worth naming

Three of these were found by asking one question of each consumer: which daemon does this call go
to, and is that the daemon that owns the thing it is asking about?
None was found by a test.
A test that builds its own object graph proves the seam works; it says nothing about whether the
production wiring uses it. Any further PR on this epic should assert against the real wiring.

## Must be fixed BEFORE `memberHerdrSocket` is ever set — running list Two PRs have now landed or been reviewed for this epic, and both had the same shape of defect: **correct today only because both herdr clients are the same object, wrong the moment the second daemon is configured.** Every test passed in each case. Keeping the list in one place so nothing enables the feature on top of a half-routed daemon. ### Landed - **#187 — `stop()` / `list()` pane-id ambiguity.** Merged as `a237fbf`. Includes three lead-review fixes: `stop()` no longer forgets the owner when the delegate refuses, `list()` is keyed by (owning daemon, pane id) instead of the raw pane id, and the class javadoc no longer states the single-connection premise as fact. 993 tests green. ### Open, blocking - **#186 — three unrouted consumers.** Sent back to a member. In short: 1. `PaneLocator` is pinned to the member daemon, so a lead's connection resolves to no terminal. `FleetMcp.reply` and `FleetMcp.ask` both refuse when `callerTerminal == null`, so a lead can no longer answer a peer lead, and `fleet_whoami` loses its leader name. 2. `StatusRefiner` is pinned to the member daemon while the status call next to it is routed per target. A lead whose raw status reads UNKNOWN gets refined against the wrong daemon and stays UNKNOWN, which wedges status-gated delivery to that lead. 3. `FleetApp` never goes through the router: `/healthz` stays green while the MEMBER daemon is down (and then every spawn fails), and `GET /sessions` drops every member workspace. - **Pane ownership does not survive a restart.** `CompositePeerLauncher.spawnedBy` is in memory. With two daemons, every pane that outlives the daemon is unowned, so `fleet_stop` on it refuses — by design, since guessing would close a pane on an arbitrary daemon. Recovering ownership at boot is its own unit and it is a prerequisite, not a nicety. - **The socket is mode 0600.** Cross-user access needs a post-start `chmod g+rw`, and `umask` does not change it. Measured on herdr 0.8.0. An upstream request is drafted; until then any start unit has to poll for the socket and relax it, which is a race on every start. ### Lower severity, do not lose - `HerdrPeerLauncher.stop` has the same before-the-call removal that #187 just fixed one level up: `paneByAgentId.remove(idOrPane)` runs before `agents.close(paneId)`. A failed stop there falls back to raw-pane addressing rather than refusing, so it degrades instead of breaking. ### The pattern worth naming Three of these were found by asking one question of each consumer: *which daemon does this call go to, and is that the daemon that owns the thing it is asking about?* None was found by a test. A test that builds its own object graph proves the seam works; it says nothing about whether the production wiring uses it. Any further PR on this epic should assert against the real wiring.
ltms closed this issue 2026-08-31 04:30:57 +02:00
ltms reopened this issue 2026-08-31 04:32:47 +02:00
Author
Owner

Reopened. PR #196 auto-closed this on merge, but it only cleared the two stage-1 blockers — the ticket has four stages and three are untouched.

Done (merged as 63c19dc):

  1. stop() after a restart. spawnedBy is in-memory only, so a daemon restart made every surviving member un-stoppable under two daemons. CompositePeerLauncher.probeOwner now asks each distinct daemon which one knows the pane: one match routes and caches, no match is treated as already-gone, more than one match is a genuine ambiguity (per-daemon counters — E4 above) and is refused with a message that now says so.
    • Review round 2 caught a second bug inside that fix: probeOwner called list() unguarded, so one unreachable daemon made panes on a different healthy daemon un-stoppable too — blocker 1 resurfacing through a new door. Each list() is now wrapped per daemon.
  2. /healthz masking a broken member daemon. It pinged the member daemon and discarded the answer, so a member herdr with a mismatched protocol looked green while every spawn failed. It now reports the member daemon's version/protocol under a separate member key, and sets protocolMismatch: true when the two protocol numbers differ. The herdr key still carries the lead daemon's values unchanged, because redeploy-fleetd.sh and rename-checkout.sh both read this endpoint.

Still open — the reason this ticket stays open:

  • Stage 1 remainder: qualify pane ids. fleet_stop{paneId} still takes a bare w1:p1. probeOwner makes the ambiguous case fail loudly instead of closing the wrong pane, which is a safety fix, not the design fix. The ids in the MCP surface and the roster are still not daemon-qualified.
  • Stage 2: hostEnvNames still reads fleetd's own environment (risk 4). Under a second user that is the wrong environment, so CB-596's gap detector reports on a set the member never had. This now interacts with the #192 work merged today: the gap report distinguishes names the scrub blanks from names the derived allow-list keeps, and both halves are computed from hostEnvNames.
  • Stage 3: worktree group/setgid provisioning (risk 3). Untouched. Worth repeating the ticket's own caveat — this isolates credentials, not the repo.
  • Stage 4: rollout on fleet01. Not started. memberHerdrSocket: is unset on both hosts, so two-herdr mode has never run outside the fleet01 experiment.
  • E2 socket-mode race. Still needs the upstream ask about a socket-mode or socket-group option. A draft exists locally but has not been posted.

Verified by the lead: merged with #194 onto an integration branch off main, full mvn clean install green at 1025 tests.

Reopened. PR #196 auto-closed this on merge, but it only cleared the **two stage-1 blockers** — the ticket has four stages and three are untouched. **Done (merged as `63c19dc`):** 1. **`stop()` after a restart.** `spawnedBy` is in-memory only, so a daemon restart made every surviving member un-stoppable under two daemons. `CompositePeerLauncher.probeOwner` now asks each distinct daemon which one knows the pane: one match routes and caches, no match is treated as already-gone, more than one match is a genuine ambiguity (per-daemon counters — E4 above) and is refused with a message that now says so. - Review round 2 caught a second bug inside that fix: `probeOwner` called `list()` unguarded, so **one unreachable daemon made panes on a different healthy daemon un-stoppable too** — blocker 1 resurfacing through a new door. Each `list()` is now wrapped per daemon. 2. **`/healthz` masking a broken member daemon.** It pinged the member daemon and discarded the answer, so a member herdr with a mismatched protocol looked green while every spawn failed. It now reports the member daemon's version/protocol under a separate `member` key, and sets `protocolMismatch: true` when the two protocol numbers differ. The `herdr` key still carries the **lead** daemon's values unchanged, because `redeploy-fleetd.sh` and `rename-checkout.sh` both read this endpoint. **Still open — the reason this ticket stays open:** - **Stage 1 remainder: qualify pane ids.** `fleet_stop{paneId}` still takes a bare `w1:p1`. `probeOwner` makes the ambiguous case *fail loudly* instead of closing the wrong pane, which is a safety fix, not the design fix. The ids in the MCP surface and the roster are still not daemon-qualified. - **Stage 2: `hostEnvNames`** still reads fleetd's own environment (risk 4). Under a second user that is the wrong environment, so CB-596's gap detector reports on a set the member never had. This now interacts with the #192 work merged today: the gap report distinguishes names the scrub blanks from names the derived allow-list keeps, and both halves are computed from `hostEnvNames`. - **Stage 3: worktree group/setgid provisioning** (risk 3). Untouched. Worth repeating the ticket's own caveat — this isolates **credentials, not the repo**. - **Stage 4: rollout on fleet01.** Not started. `memberHerdrSocket:` is unset on both hosts, so two-herdr mode has never run outside the fleet01 experiment. - **E2 socket-mode race.** Still needs the upstream ask about a socket-mode or socket-group option. A draft exists locally but has not been posted. Verified by the lead: merged with #194 onto an integration branch off main, full `mvn clean install` green at 1025 tests.
Author
Owner

Stage 3 in progress, and four more files that assume fleetd's uid owns everything it writes

Stage 3 (worktree group/setgid provisioning) is implemented in PR #208: an opt-in worktreeGroup:
key, a Worktrees.shareWithGroup seam, and a SessionManager call placed after overlayParity
— because overlayParity copies more files into the worktree after add() returns, so sharing any
earlier leaves every overlay file operator-owned and unwritable. A test pins that order.

Two defects found in review and fixed before merge, both of which would have failed every
provisioning spawn rather than just the two-user case:

  • .git/logs was passed to chgrp unguarded. It does not exist when core.logAllRefUpdates=false
    or before the first ref update, and chgrp on a missing path exits non-zero — surfacing as a
    WorktreeException blaming a group that is in fact fine. Every path is now skipped when absent.
  • repoRoot + "/.git" was hardcoded. That is a file, not a directory, when the checkout is
    itself a linked worktree — the very thing this codebase creates for every member. Now resolved
    with git rev-parse --git-common-dir (relative or absolute, resolved against repoRoot).

The stage-3 fix-up does not reach four other writes

Found while reviewing #208. Reported, not fixed — these belong to stage 4, not to that PR.

1. The role charter goes to the system temp directory, which on macOS is per-user 0700.

ClaudeCodeLauncher.writeCharterFile (ClaudeCodeLauncher.java:454) does
Files.createTempFile("fleetd-role-charter-", ".md") and passes the path to the member as a launch
flag. That path is outside the worktree, so worktreeGroup cannot reach it. I checked what the
directory actually is:

$ ls -ld "$TMPDIR"
drwx------@ 219 dai.ha  staff  /var/folders/wf/ljcnfwxj2pd1nxtd63tv2df40000gq/T/
$ java -XshowSettings:properties -version 2>&1 | grep java.io.tmpdir
    java.io.tmpdir = /var/folders/wf/…/T/

0700, owned by the operator. A member running as a different user cannot even traverse it, so
it never reads its charter — and the charter is what carries the reply contract. This is a hard
blocker for two-user mode, not a permissions nicety. The same applies to
EnvAllowListScrub.write (EnvAllowListScrub.java:103), which creates its ZDOTDIR with
Files.createTempDirectory(parentDir, …): the member's own login shell, as another uid, has to read
what is generated there.

2. Two launcher writes land after the fix-up has already run.

ClaudeCodeLauncher.writeIdeOverlay (CLAUDE.local.md into the worktree root) and its companion
write to <commonDir>/info/exclude both happen at spawn time, i.e. after
SessionManager has called shareWithGroup. Those files are fleetd-uid-owned with only
umask-derived group access. They are read-only from the member's side today, so 0644 happens to be
enough — but nothing states that, and the moment one becomes member-writable it breaks silently.

This is the same ordering trap as overlayParity, one layer further out: the permission pass has
to be the last thing that touches the tree, and spawn is after it.

3. Not checked. OpenCodeLauncher's equivalent config/charter writes. The shape is likely the
same.

What this means for staging

Stage 3 alone does not make two-user mode work. Item 1 above has to be solved first, and it is not a
chmod — a per-user 0700 temp directory cannot be opened up safely. The charter and the scrub
directory need to move somewhere both uids can reach, most likely under the worktree (which stage 3
already shares) or a purpose-made directory owned by the group.

## Stage 3 in progress, and four more files that assume fleetd's uid owns everything it writes Stage 3 (worktree group/setgid provisioning) is implemented in PR #208: an opt-in `worktreeGroup:` key, a `Worktrees.shareWithGroup` seam, and a `SessionManager` call placed **after** `overlayParity` — because `overlayParity` copies more files into the worktree after `add()` returns, so sharing any earlier leaves every overlay file operator-owned and unwritable. A test pins that order. Two defects found in review and fixed before merge, both of which would have failed **every** provisioning spawn rather than just the two-user case: - `.git/logs` was passed to `chgrp` unguarded. It does not exist when `core.logAllRefUpdates=false` or before the first ref update, and `chgrp` on a missing path exits non-zero — surfacing as a `WorktreeException` blaming a group that is in fact fine. Every path is now skipped when absent. - `repoRoot + "/.git"` was hardcoded. That is a **file**, not a directory, when the checkout is itself a linked worktree — the very thing this codebase creates for every member. Now resolved with `git rev-parse --git-common-dir` (relative or absolute, resolved against `repoRoot`). ### The stage-3 fix-up does not reach four other writes Found while reviewing #208. Reported, not fixed — these belong to stage 4, not to that PR. **1. The role charter goes to the system temp directory, which on macOS is per-user 0700.** `ClaudeCodeLauncher.writeCharterFile` (`ClaudeCodeLauncher.java:454`) does `Files.createTempFile("fleetd-role-charter-", ".md")` and passes the path to the member as a launch flag. That path is outside the worktree, so `worktreeGroup` cannot reach it. I checked what the directory actually is: ``` $ ls -ld "$TMPDIR" drwx------@ 219 dai.ha staff /var/folders/wf/ljcnfwxj2pd1nxtd63tv2df40000gq/T/ $ java -XshowSettings:properties -version 2>&1 | grep java.io.tmpdir java.io.tmpdir = /var/folders/wf/…/T/ ``` `0700`, owned by the operator. **A member running as a different user cannot even traverse it**, so it never reads its charter — and the charter is what carries the reply contract. This is a hard blocker for two-user mode, not a permissions nicety. The same applies to `EnvAllowListScrub.write` (`EnvAllowListScrub.java:103`), which creates its `ZDOTDIR` with `Files.createTempDirectory(parentDir, …)`: the member's own login shell, as another uid, has to read what is generated there. **2. Two launcher writes land after the fix-up has already run.** `ClaudeCodeLauncher.writeIdeOverlay` (`CLAUDE.local.md` into the worktree root) and its companion write to `<commonDir>/info/exclude` both happen at **spawn** time, i.e. after `SessionManager` has called `shareWithGroup`. Those files are fleetd-uid-owned with only umask-derived group access. They are read-only from the member's side today, so `0644` happens to be enough — but nothing states that, and the moment one becomes member-writable it breaks silently. This is the same ordering trap as `overlayParity`, one layer further out: **the permission pass has to be the last thing that touches the tree, and spawn is after it.** **3. Not checked.** `OpenCodeLauncher`'s equivalent config/charter writes. The shape is likely the same. ### What this means for staging Stage 3 alone does not make two-user mode work. Item 1 above has to be solved first, and it is not a `chmod` — a per-user `0700` temp directory cannot be opened up safely. The charter and the scrub directory need to move somewhere both uids can reach, most likely under the worktree (which stage 3 already shares) or a purpose-made directory owned by the group.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#185