Files
fleetd/docs/CB-301-Session-Manager.md
Dai Ha 19b10e3216 docs: CB-301 session-manager design spec (one-shot, no reuse; recycle in scope)
The as-built audit surfaced that WorkerService keeps no registry of what it
spawned (its own Javadoc: 'there is no registry; list() only asks herdr').
CB-301 adds a SessionManager wrapping WorkerService: an authoritative in-daemon
roster with a per-session lifecycle FSM (SPAWNING/READY/BUSY/DONE/RELEASED/
FAILED), deterministic release, and recycle (= release + fresh acquire, no
reuse). Leaves clean seams for CB-302 (checkpoint on release), CB-303 (idle_ttl/
context_cap/drain policy over roster), CB-304 (bridge_list reads roster).
2026-07-16 19:10:13 +02:00

6.8 KiB

CB-301 — Session Manager (one-shot, no reuse)

Status: design spec for review → delegate implementation. Grounded in: WorkerService, Injector/StatusPoller/TurnListener, MessageService, BridgeMcp, BridgedApp (see wiki 9. Implementation).

Problem

WorkerService is stateless about what it spawned. Its own Javadoc says it plainly:

"there is no registry; list() only asks herdr." — WorkerService.reapOrphanWorkers (line 288)

Consequences today:

  • The daemon cannot answer "which workers did I spawn, in what lifecycle state, owned by whom, since when?" without shelling to herdr for a raw agent list (no state, no ownership, no age).
  • Cleanup of a worker that outlived its owning process depends entirely on the boot-time name-nonce reaper (CB-117) — there is no live, authoritative roster during a run.
  • bridge_list (CB-304) can only surface herdr's view, not a bridge-owned roster.
  • There is no seam for per-session policy (checkpoint on teardown → CB-302; idle_ttl / context_cap / drain → CB-303).

Goal & non-goals

Goal. Introduce a SessionManager that owns an authoritative in-daemon registry of the worker sessions this daemon process spawned, tracks each one's lifecycle state, and tears each down deterministically. It becomes the single source of truth for the roster and the seam CB-302/303/304 build on.

Non-goals (explicit — reuse policy chosen: one-shot, no reuse).

  • No pooling / no reuse. Every delegated task gets a fresh worker; a finished worker is torn down, never handed to a later task. No "warm idle" pool, no role@profile keying.
  • No auto-teardown timing. When a one-shot worker is released (immediately on turn completion vs after an idle grace) is CB-303. CB-301 provides the mechanism (release) and the registry; CB-303 sets the policy.
  • No checkpoint content. Writing STATE.md + commit on teardown is CB-302; CB-301 only exposes the release hook it will attach to.

"Recycle" under no-reuse is simply release + fresh acquire — a helper, not a pool operation.

Design

SessionManager wraps WorkerService (does not replace it). WorkerService keeps doing the subscription-guarded spawn/teardown mechanics; SessionManager adds the registry, lifecycle, and ownership on top.

Package: new dev.ltms.bridged.session — keeps the registry/lifecycle concern separate from the worker spawn mechanics. Holds SessionManager + WorkerSession.

recycle is IN SCOPE for CB-301 (decided): implement recycle(paneId, …) = release the old session then acquire a fresh one, asserting a new distinct paneId (the no-reuse invariant). It is a thin convenience over the two primitives, shipped now so the no-reuse teardown+respawn path is covered by a test from day one.

WorkerSession (record or small mutable holder)

Field Source Notes
paneId Agent.paneId() registry key
terminalId Agent.terminalId() for status/identity joins
profile spawn arg which profile spawned it
cwd resolved cwd the worker's working dir
ownerTerminal caller identity (nullable) the primary/turn that requested it; null = daemon/anon
spawnedAtNanos System.nanoTime() age basis for CB-303 (monotonic; no wall clock in tests)
state lifecycle FSM see below

State is held in a ConcurrentHashMap<String /*paneId*/, WorkerSession>.

Lifecycle state machine (one-shot)

SPAWNING --ready(MCP present)--> READY
READY --onDelivered--> BUSY
BUSY --onTurnComplete--> DONE
BUSY --onTurnFailed--> FAILED
READY|DONE|FAILED --release()--> RELEASED (deregistered)
SPAWNING|READY|BUSY|DONE --vanished/drop--> FAILED
  • Transitions are driven by hooks the manager already has access to: WorkerPresence.markPresent → READY; TurnListener.onDelivered/onTurnComplete/onTurnFailed (the manager implements or decorates TurnListener) → BUSY/DONE/FAILED.
  • RELEASED sessions are removed from the registry (teardown is terminal).
  • Any state → FAILED on drop (worker vanished / injector drop), mirroring Injector.

API

final class SessionManager {
    WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal);
    void          release(String paneId);          // deterministic teardown + deregister
    WorkerSession recycle(String paneId, ...);      // release + acquire (no-reuse convenience)
    Optional<WorkerSession> get(String paneId);
    List<WorkerSession> roster();                   // bridge-owned view (CB-304 consumes this)
    // lifecycle hooks (package-private): onReady/onDelivered/onComplete/onFailed(target)
}
  • acquire = workerService.spawn(profile, requestedCwd, callerCwd) → register SPAWNING.
  • release = workerService.stop(paneId) → deregister. Idempotent (already-gone tolerated, matching WorkerService.stop).
  • roster joins the registry with live herdr status for a truthful "roster + live" (CB-304).

Integration points

  • Bridged.main — construct SessionManager(workerService, ...); wire it as/decorating the TurnListener alongside CompletionResolver so it sees turn boundaries, and give it the WorkerPresence signal for READY.
  • BridgeMcp.spawn / BridgedApp.spawnWorker — route spawn through SessionManager.acquire (carry callerTerminal as ownerTerminal). bridge_stop / DELETE /workers/{paneId} → SessionManager.release.
  • bridge_list / GET /sessions (CB-304 later) — read SessionManager.roster().
  • MessageService — no change required for one-shot; a later CB-303 auto-release hook can call release from onTurnComplete under policy.

Acceptance (tests, no live herdr — fakes as elsewhere)

  1. acquire registers a SPAWNING session with the right owner/profile/cwd; a second acquire yields a distinct paneId and a distinct session (no reuse).
  2. Presence signal moves SPAWNING → READY; a delivered turn moves READY → BUSY → DONE.
  3. release tears the worker down via WorkerService.stop and removes it from roster(); a second release on the same paneId is a harmless no-op.
  4. onTurnFailed / drop moves the session to FAILED and it is absent from the live roster.
  5. recycle produces a new paneId and the old one is gone (no-reuse invariant).
  6. roster() reflects exactly the sessions acquired-minus-released, joined with live status.

Seams left open (deliberately)

  • CB-302 — attach a checkpoint step (STATE.md + commit) to the release path.
  • CB-303 — a policy loop over roster() using spawnedAtNanos/state to auto-release on idle_ttl, or drain on context_cap.
  • CB-304 — bridge_list reads roster() for a bridge-owned roster + live join.