Features: worktreeGroup, and opencode session ids from opencode.db

Two shipped capabilities that had landed nowhere, plus a correction to the
memberHerdrSocket entry: worktree uid is no longer the open blocker there, but
the per-user 0700 system temp directory holding the role charter now is, and
that one no chmod fixes.
Dai Ha
2026-08-31 21:49:05 +07:00
parent 4912b7acaf
commit 34593af1a1
+103 -1
@@ -67,6 +67,8 @@ six weeks, and the table alone will not carry it.
| [See which charter a member got](#see-which-charter-a-member-got) | automatic | CB-571 | `peer/CharterReceipt` |
| [Redeploy the daemon safely](#redeploy-the-daemon-safely) | run the script | — | `scripts/redeploy-fleetd.sh` |
| [Run members on a second herdr daemon](#memberherdrsocket--run-members-on-a-second-herdr-daemon) | `memberHerdrSocket:` | CB-185 | `herdr/HerdrRouter` |
| [Let a member under another OS user write its worktree](#worktreegroup--let-a-member-under-another-os-user-write-its-worktree) | `worktreeGroup:` | CB-185 | `session/GitWorktrees` |
| [Resume an opencode member's prior session](#opencode-session-ids-come-from-opencodedb) | automatic | CB-206 | `member/OpenCodeSessionDiscovery` |
Nearly every knob above lives in one file, on one profile:
@@ -2283,7 +2285,24 @@ access still needs a post-start `chmod g+rw` — a race on every start, and an u
not been made. Pane ids are still **not** daemon-qualified: `fleet_stop{paneId}` takes a bare
`w1:p1`, so the probe makes a collision fail loudly instead of silently closing the wrong pane, which
is a safety net, not the design fix. `hostEnvNames` still reads fleetd's own environment, which is
the wrong environment under a second user. Worktrees are still created as the operator's uid.
the wrong environment under a second user.
Worktree ownership now has an answer — [`worktreeGroup`](#worktreegroup--let-a-member-under-another-os-user-write-its-worktree),
below — but a **new** blocker replaced it, and it is the harder one. `ClaudeCodeLauncher` writes the
role charter with `Files.createTempFile`, and on macOS the system temp directory is per-user:
```
$ ls -ld "$TMPDIR"
drwx------@ 219 dai.ha staff /var/folders/wf/…/T/
```
A member running as another user cannot even traverse it, so it never reads its charter — and the
charter carries the reply contract. `EnvAllowListScrub` has the same problem: it generates the
member's `ZDOTDIR` there, and the member's own login shell has to read it. No `chmod` fixes a
per-user `0700` directory safely; both writes have to move somewhere both uids can reach. Two
launcher writes (`CLAUDE.local.md` and `<commonDir>/info/exclude`) also land at spawn time, i.e.
*after* the group fix-up has run — harmless today because the member only reads them, but nothing
says so.
**The shape to watch for.** herdr's workspace, tab and pane ids are **per-daemon sequential
counters**. Two daemons really do both hold `w1:p1`, pointing at different panes owned by different
@@ -2320,6 +2339,89 @@ member daemon's protocol, not the lead's, that decides whether a spawn works.
---
## `worktreeGroup` — let a member under another OS user write its worktree
**What.** An optional top-level key naming an **OS group**. When set, every provisioned worktree and
the repository's git store are made group-writable, so a member spawned under a different OS user
(see [`memberHerdrSocket`](#memberherdrsocket--run-members-on-a-second-herdr-daemon)) can write its
own worktree, its per-worktree git metadata, and the objects its commits create. Absent — the
default — nothing runs and no process is spawned.
```yaml
worktreeGroup: fleet-workers # omit for the single-user default
```
**On.** Add the key and restart. The operator running fleetd must already be a member of that group,
or every provisioning spawn fails loudly, naming the group.
**Why.** `GitWorktrees` shells `git worktree add`, so it runs as **fleetd's** uid and everything it
creates is owned by the operator at `0644`/`0755`. A member running as another user cannot write any
of it. That user is the whole point of #185: a member currently reads the operator's forge ssh key
straight off the filesystem, and no environment scrub can stop that, because the key is a file
(#184). This key is the git half of the fix.
**It isolates credentials, not the repository.** A member in the group can still write the
operator's git objects and refs. Say that plainly rather than implying more — the two uids share one
repository by design, which is what keeps the lead's `refs/wip/*` safety net working.
**Gotchas, all three found the hard way.**
*The share pass must run after `overlayParity`, never inside `add()`.* `overlayParity` copies more
files into the worktree after `add()` returns, so sharing any earlier leaves every overlay file
operator-owned and unwritable — with the whole suite still green. A test pins the order.
*Missing paths are skipped, not passed to `chgrp`.* `.git/logs` does not exist with
`core.logAllRefUpdates=false` or before the first ref update, and `packed-refs` does not exist until
refs are packed. `chgrp` on a missing path exits non-zero, which would fail **every** spawn with a
message blaming a group that is fine.
*The git directory is asked for, never assumed.* `<repoRoot>/.git` is a **file**, not a directory,
when the checkout is itself a linked worktree — which is what this code creates for every member. It
resolves `git rev-parse --git-common-dir` against `repoRoot`, because git answers relatively for an
ordinary checkout and absolutely for a linked one.
**Cost.** The fix-up re-runs on every spawn, deliberately. `core.sharedRepository=group` governs only
what git writes *afterwards*; the walk is what covers everything already on disk. Three walks of the
object store per spawn — about 3000 files in this repo, well under a second, growing with the repo.
It is not redundant work to optimise away.
---
## opencode session ids come from `opencode.db`
**What.** `fleet_list` reports `agentSessionId` for an opencode member, and
`fleet_spawn{resumeSessionId}` can therefore relaunch one onto its prior conversation. Automatic —
no knob.
**On.** Nothing to turn on. It needs opencode's own storage root
(`~/.local/share/opencode`) to be readable, which it is.
**Why it is an entry at all: it was silently dead.** opencode moved its session store from a
one-JSON-file-per-session tree to a SQLite database in **January 2026**, and
`OpenCodeSessionDiscovery` went on scanning the frozen tree. It therefore returned `null` for every
member, forever. `agentSessionId` was never known, so `resumeSessionId` was unreachable for every
opencode profile — most of the fleet — and nothing said so, through 57 member spawns. It now reads
`<storageRoot>/opencode.db`, matching the `session` table's `directory` column against the worker's
cwd and taking the highest `time_updated`.
**The gotcha is what the silence cost.** A discovery that finds nothing and a store that is empty
look identical. That is why a missing database is now a **WARN** naming the path searched — it means
the layout moved again — while a present database with no matching row stays at DEBUG, because that
is the normal answer right after a spawn.
**Read-only, and pinned as such.** opencode may be running and writing that database (WAL mode, 841MB
here), so the connection is opened with `SQLITE_OPEN_READONLY`. Beware the obvious test for this:
making the file unwritable proves **nothing**, because SQLite silently downgrades a read-write open
of an unwritable file to read-only. The test that works asks the connection to `INSERT` and requires
the refusal.
**One dependency.** `org.xerial:sqlite-jdbc`, which ships bundled native libraries and takes the
shaded jar to about 28.6MB. Chosen over shelling out to the `sqlite3` binary so there is no
undeclared external-binary requirement on a headless host.
---
## Backfill status
This page was started after the fact, so it is **not yet complete**. Entries above are written from