agentSessionId is resolved once at spawn and frozen, so an opencode member never reports one #209

Closed
opened 2026-08-31 16:57:12 +02:00 by ltms · 1 comment
Owner

Follow-up to #206. The SQLite discovery fixed in #206/#207 works, but the value it produces can never reach the roster. fleet_list still shows no agentSessionId for any opencode member, so resumeSessionId remains unusable for that backend.

The defect

SessionManager calls handle.agentSessionId() exactly once, at spawn, and freezes the answer in an immutable MemberSession record:

  • SessionManager.java:509 (worktree path) and SessionManager.java:206 (plain path)
  • MemberSession is a record; none of its withers (withState, withActivity, bumpTurn) ever recomputes agentSessionId — they all copy the frozen value through.
  • rosterView (SessionManager.java:578) reads that frozen field.

For opencode the value at spawn time is always null, because opencode has not yet written the session row when the pane is created. That is the same asynchrony #206 documented. OpenCodeLauncher.SessionAwareHandle.agentSessionId() says so in its own javadoc:

Lazy + retried, never a spawn-time blocker: opencode writes the session record only when the session is first persisted, so null here is the correct interim answer and the caller re-calls later (each call re-scans, picking up a record that has since appeared).

No caller re-calls. SessionManager does not retain the PeerHandle at all — the registry maps paneId to MemberSession, and the handle goes out of scope at the end of the spawn method. So "the caller re-calls later" describes a caller that does not exist.

Live evidence (2026-08-31, daemon pid 97696, jar 9ec5fab0f136)

Spawned an opencode member on a fresh worktree, so any id must come from a new row rather than a pre-existing one:

fleet_spawn{profile: "gx", role: "dev", worktree: "cb206-live-probe", ticket: "cb206-live-probe"}

The row exists in ~/.local/share/opencode/opencode.db:

ses_fa7b15930ffeOSrh436Vln7mcn|/Users/dai.ha/LTMS/.bridged-worktrees/4bde6e-1

The member finished its turn and replied. fleet_list for the same member:

{"sessionId":"term_65a58f25502e82c","paneId":"c1f42324-...","profile":"gx","role":"dev",
 "state":"done","worktree":"/Users/dai.ha/LTMS/.bridged-worktrees/4bde6e-1",
 "branch":"worker/cb206-live-probe-61cfb4-1","owner":"term_65a4c36d11af27",
 "charterSource":"none","liveStatus":"done","reclaimable":true,"idleForSeconds":198}

The directory in the database and the worktree in the roster are the same string. There is no agentSessionId field.

Why the tests did not catch it

Every test asserts on the seam — OpenCodeSessionDiscovery.sessionIdForDirectory(...) and SessionAwareHandle.agentSessionId() — both of which answer correctly. None starts a real spawn, waits for the row to appear, and then reads the roster. This is the same family as #175/#203 and the SessionReaper sweep: the work is right and the caller never reaches it.

Suggested fix

Retain the handle and re-resolve while the stored id is still null, sticky once found:

  1. Add ConcurrentHashMap<String /*paneId*/, PeerHandle> handles to SessionManager; populate it on both spawn paths; remove the entry in release(...) next to registry.remove(paneId).
  2. Add MemberSession.withAgentSessionId(String).
  3. Add a private resolve step: if session.agentSessionId() != null, return it unchanged; otherwise look up the handle, call agentSessionId(), and on a non-null answer CAS-replace the registry entry (the existing registry.replace(expected, updated) helper at SessionManager.java:814) and return the updated session.
  4. Run that step from roster() and from find(paneId), so fleet_list and fleet_status both see it. Work is bounded: only members whose id is still null do any I/O, and each one stops doing it as soon as the row appears.
  5. Resolve once more in release(...) before notifyReleased(new ReleaseDetail(..., removed.agentSessionId())) at SessionManager.java:307, so a released member's detail carries the id it now has. That is the CB-584 criterion-5 path, and it is frozen-null today for exactly the same reason.

Acceptance

  • A test that starts a real spawn against a fake launcher whose handle returns null first and an id on a later call, then asserts roster()/rosterView reports the id. Removing the fix must make it fail — watch it fail before keeping it.
  • A test that release(...) carries a late-resolved id into ReleaseDetail.
  • Live: spawn an opencode member on a fresh worktree, wait for the row, and see agentSessionId in fleet_list.
Follow-up to #206. The SQLite discovery fixed in #206/#207 works, but the value it produces can never reach the roster. `fleet_list` still shows no `agentSessionId` for any opencode member, so `resumeSessionId` remains unusable for that backend. ## The defect `SessionManager` calls `handle.agentSessionId()` exactly once, at spawn, and freezes the answer in an immutable `MemberSession` record: - `SessionManager.java:509` (worktree path) and `SessionManager.java:206` (plain path) - `MemberSession` is a record; none of its withers (`withState`, `withActivity`, `bumpTurn`) ever recomputes `agentSessionId` — they all copy the frozen value through. - `rosterView` (`SessionManager.java:578`) reads that frozen field. For opencode the value at spawn time is **always** null, because opencode has not yet written the session row when the pane is created. That is the same asynchrony #206 documented. `OpenCodeLauncher.SessionAwareHandle.agentSessionId()` says so in its own javadoc: > Lazy + retried, never a spawn-time blocker: opencode writes the session record only when the session is first persisted, so null here is the correct interim answer and **the caller re-calls later** (each call re-scans, picking up a record that has since appeared). No caller re-calls. `SessionManager` does not retain the `PeerHandle` at all — the registry maps paneId to `MemberSession`, and the handle goes out of scope at the end of the spawn method. So "the caller re-calls later" describes a caller that does not exist. ## Live evidence (2026-08-31, daemon pid 97696, jar 9ec5fab0f136) Spawned an opencode member on a **fresh** worktree, so any id must come from a new row rather than a pre-existing one: `fleet_spawn{profile: "gx", role: "dev", worktree: "cb206-live-probe", ticket: "cb206-live-probe"}` The row exists in `~/.local/share/opencode/opencode.db`: ``` ses_fa7b15930ffeOSrh436Vln7mcn|/Users/dai.ha/LTMS/.bridged-worktrees/4bde6e-1 ``` The member finished its turn and replied. `fleet_list` for the same member: ```json {"sessionId":"term_65a58f25502e82c","paneId":"c1f42324-...","profile":"gx","role":"dev", "state":"done","worktree":"/Users/dai.ha/LTMS/.bridged-worktrees/4bde6e-1", "branch":"worker/cb206-live-probe-61cfb4-1","owner":"term_65a4c36d11af27", "charterSource":"none","liveStatus":"done","reclaimable":true,"idleForSeconds":198} ``` The `directory` in the database and the `worktree` in the roster are the same string. There is no `agentSessionId` field. ## Why the tests did not catch it Every test asserts on the seam — `OpenCodeSessionDiscovery.sessionIdForDirectory(...)` and `SessionAwareHandle.agentSessionId()` — both of which answer correctly. None starts a real spawn, waits for the row to appear, and then reads the roster. This is the same family as #175/#203 and the `SessionReaper` sweep: the work is right and the caller never reaches it. ## Suggested fix Retain the handle and re-resolve while the stored id is still null, sticky once found: 1. Add `ConcurrentHashMap<String /*paneId*/, PeerHandle> handles` to `SessionManager`; populate it on both spawn paths; remove the entry in `release(...)` next to `registry.remove(paneId)`. 2. Add `MemberSession.withAgentSessionId(String)`. 3. Add a private resolve step: if `session.agentSessionId() != null`, return it unchanged; otherwise look up the handle, call `agentSessionId()`, and on a non-null answer CAS-replace the registry entry (the existing `registry.replace(expected, updated)` helper at `SessionManager.java:814`) and return the updated session. 4. Run that step from `roster()` and from `find(paneId)`, so `fleet_list` and `fleet_status` both see it. Work is bounded: only members whose id is still null do any I/O, and each one stops doing it as soon as the row appears. 5. Resolve once more in `release(...)` **before** `notifyReleased(new ReleaseDetail(..., removed.agentSessionId()))` at `SessionManager.java:307`, so a released member's detail carries the id it now has. That is the CB-584 criterion-5 path, and it is frozen-null today for exactly the same reason. ## Acceptance - A test that starts a real spawn against a fake launcher whose handle returns null first and an id on a later call, then asserts `roster()`/`rosterView` reports the id. Removing the fix must make it fail — watch it fail before keeping it. - A test that `release(...)` carries a late-resolved id into `ReleaseDetail`. - Live: spawn an opencode member on a fresh worktree, wait for the row, and see `agentSessionId` in `fleet_list`.
Author
Owner

Proven live on the running daemon (pid 79448, jar ed3eab5c8baa, redeployed from 966c58a today).

I spawned a fresh opencode member on a brand-new worktree, so any id reported had to be produced during this test — a stale value could not explain the result.

At spawn, fleet_list had no agentSessionId (opencode has not written its session row yet). Once the row appeared:

sqlite3 -readonly ~/.local/share/opencode/opencode.db
  select id, directory from session where directory like '%154133-1%';
ses_fa75aad35ffeQTF5ZwRqq3Sqhg|/Users/dai.ha/LTMS/.bridged-worktrees/154133-1

fleet_list then reported:

"agentSessionId": "ses_fa75aad35ffeQTF5ZwRqq3Sqhg"

Exact match. The MCP tool and GET /members both resolve it now; before the fix the value was read once at spawn and frozen, so for opencode it was null forever.

GET /members checked at the same time (this also closes out #199 live):

top-level keys: ['members', 'wipRefs', 'workers']
members rows: 1
workers alias rows: 1
alias identical: True
  agentSessionId = ses_fa75aad35ffeQTF5ZwRqq3Sqhg

This unblocks #175 — resumeSessionId now has a real id to resume onto for opencode members.

**Proven live on the running daemon** (pid 79448, jar `ed3eab5c8baa`, redeployed from `966c58a` today). I spawned a fresh opencode member on a brand-new worktree, so any id reported had to be produced during this test — a stale value could not explain the result. At spawn, `fleet_list` had no `agentSessionId` (opencode has not written its session row yet). Once the row appeared: ``` sqlite3 -readonly ~/.local/share/opencode/opencode.db select id, directory from session where directory like '%154133-1%'; ses_fa75aad35ffeQTF5ZwRqq3Sqhg|/Users/dai.ha/LTMS/.bridged-worktrees/154133-1 ``` `fleet_list` then reported: ``` "agentSessionId": "ses_fa75aad35ffeQTF5ZwRqq3Sqhg" ``` Exact match. The MCP tool and `GET /members` both resolve it now; before the fix the value was read once at spawn and frozen, so for opencode it was `null` forever. `GET /members` checked at the same time (this also closes out #199 live): ``` top-level keys: ['members', 'wipRefs', 'workers'] members rows: 1 workers alias rows: 1 alias identical: True agentSessionId = ses_fa75aad35ffeQTF5ZwRqq3Sqhg ``` This unblocks #175 — `resumeSessionId` now has a real id to resume onto for opencode members.
ltms closed this issue 2026-08-31 18:29:04 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#209