Files
fleetd/docs/CB-301-Session-Manager.md
Dai Ha ecc590f344 CB-632 unit 5: rename the daemon and classes in the docs prose
Part of #145 (CB-632). Documentation only, plus one internal literal.

Unit 1 renamed the package and classes, which left every doc describing
classes that no longer exist. This fixes the prose across README.md,
docs/ and bridged/docs/ -- 18 files.

Renamed: dev.ltms.bridged -> dev.ltms.fleet, the five class names, and
"bridged" where it names the daemon as a product rather than a path.

Also renamed two literals, because a doc that disagrees with the code is
worse than one that is out of date:

  - bridged-local-noauth -> fleetd-local-noauth. A placeholder apiKey
    OpenCodeLauncher sends when a profile resolves no token, to a local
    endpoint that does not check it. No test asserts the old string.
  - the vnd.ltms.bridged.* media type in the M4 design doc. It appears
    in no Java file, so nothing implements it yet.

Deliberately NOT renamed, because each is still literally true today and
changes only at the cutover:

  - paths: bridged/, bridged.yaml, bridged.example.yaml, bridged.jar,
    .bridged-worktrees, deploy/dev.ltms.bridged.plist,
    scripts/redeploy-bridged.sh, bridged-launchd-wrapper.sh
  - bridged_* metric names -- renaming these after the monitoring is
    wired would break dashboard continuity, so they move before it is
  - bridge_* MCP tool names, which answer alongside fleet_* on purpose
  - BRIDGED_* environment variables, read by a file outside this repo

Method note: perl, not sed. BSD sed has no \b and no lookaround, and a
word-boundary expression there fails silently. The prose replace uses
(?<![\w./-])bridged(?![\w./-]) so it cannot touch a path or an
identifier, then every remaining hit was read by hand.

Verified: mvn clean install green, 51 classes, 878 tests, 0 failures.
2026-08-23 06:46:34 +02:00

6.3 KiB

CB-301 — Session Manager (one-shot, no reuse)

Status: design spec for review → delegate implementation. Grounded in: WorkerService, Injector/StatusPoller/TurnListener, MessageService, FleetMcp, FleetApp (see wiki 9. Implementation).

Problem

WorkerService is stateless about what it spawned. Its own Javadoc says it plainly:

"there is no registry; list() only asks herdr." — WorkerService.reapOrphanWorkers (line 288)

Consequences today:

  • The daemon cannot answer "which workers did I spawn, in what lifecycle state, owned by whom, since when?" without shelling to herdr for a raw agent list (no state, no ownership, no age).
  • Cleanup of a worker that outlived its owning process depends entirely on the boot-time name-nonce reaper (CB-117) — there is no live, authoritative roster during a run.
  • fleet_list (CB-304) can only surface herdr's view, not a bridge-owned roster.
  • There is no seam for per-session policy (checkpoint on teardown → CB-302; idle_ttl / context_cap / drain → CB-303).

Goal & non-goals

Goal. Introduce a SessionManager that owns an authoritative in-daemon registry of the worker sessions this daemon process spawned, tracks each one's lifecycle state, and tears each down deterministically. It becomes the single source of truth for the roster and the seam CB-302/303/304 build on.

Non-goals (explicit — reuse policy chosen: one-shot, no reuse).

  • No pooling / no reuse. Every delegated task gets a fresh worker; a finished worker is torn down, never handed to a later task. No "warm idle" pool, no role@profile keying.
  • No auto-teardown timing. When a one-shot worker is released (immediately on turn completion vs after an idle grace) is CB-303. CB-301 provides the mechanism (release) and the registry; CB-303 sets the policy.
  • No checkpoint content. Writing STATE.md + commit on teardown is CB-302; CB-301 only exposes the release hook it will attach to.

Under no-reuse, a released session is terminal. A new acquire always creates a fresh session.

Design

SessionManager wraps WorkerService (does not replace it). WorkerService keeps doing the subscription-guarded spawn/teardown mechanics; SessionManager adds the registry, lifecycle, and ownership on top.

Package: new dev.ltms.fleet.session — keeps the registry/lifecycle concern separate from the worker spawn mechanics. Holds SessionManager + WorkerSession.

WorkerSession (record or small mutable holder)

Field Source Notes
paneId Agent.paneId() registry key
terminalId Agent.terminalId() for status/identity joins
profile spawn arg which profile spawned it
cwd resolved cwd the worker's working dir
ownerTerminal caller identity (nullable) the primary/turn that requested it; null = daemon/anon
spawnedAtNanos System.nanoTime() age basis for CB-303 (monotonic; no wall clock in tests)
state lifecycle FSM see below

State is held in a ConcurrentHashMap<String /*paneId*/, WorkerSession>.

Lifecycle state machine (one-shot)

SPAWNING --ready(MCP present)--> READY
READY --onDelivered--> BUSY
BUSY --onTurnComplete--> DONE
BUSY --onTurnFailed--> FAILED
READY|DONE|FAILED --release()--> RELEASED (deregistered)
SPAWNING|READY|BUSY|DONE --vanished/drop--> FAILED
  • Transitions are driven by hooks the manager already has access to: WorkerPresence.markPresent → READY; TurnListener.onDelivered/onTurnComplete/onTurnFailed (the manager implements or decorates TurnListener) → BUSY/DONE/FAILED.
  • RELEASED sessions are removed from the registry (teardown is terminal).
  • Any state → FAILED on drop (worker vanished / injector drop), mirroring Injector.

API

final class SessionManager {
    WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal);
    void          release(String paneId);          // deterministic teardown + deregister
    Optional<WorkerSession> get(String paneId);
    List<WorkerSession> roster();                   // bridge-owned view (CB-304 consumes this)
    // lifecycle hooks (package-private): onReady/onDelivered/onComplete/onFailed(target)
}
  • acquire = workerService.spawn(profile, requestedCwd, callerCwd) → register SPAWNING.
  • release = workerService.stop(paneId) → deregister. Idempotent (already-gone tolerated, matching WorkerService.stop).
  • roster joins the registry with live herdr status for a truthful "roster + live" (CB-304).

Integration points

  • Fleetd.main — construct SessionManager(workerService, ...); wire it as/decorating the TurnListener alongside CompletionResolver so it sees turn boundaries, and give it the WorkerPresence signal for READY.
  • FleetMcp.spawn / FleetApp.spawnWorker — route spawn through SessionManager.acquire (carry callerTerminal as ownerTerminal). fleet_stop / DELETE /workers/{paneId} → SessionManager.release.
  • fleet_list / GET /sessions (CB-304 later) — read SessionManager.roster().
  • MessageService — no change required for one-shot; a later CB-303 auto-release hook can call release from onTurnComplete under policy.

Acceptance (tests, no live herdr — fakes as elsewhere)

  1. acquire registers a SPAWNING session with the right owner/profile/cwd; a second acquire yields a distinct paneId and a distinct session (no reuse).
  2. Presence signal moves SPAWNING → READY; a delivered turn moves READY → BUSY → DONE.
  3. release tears the worker down via WorkerService.stop and removes it from roster(); a second release on the same paneId is a harmless no-op.
  4. onTurnFailed / drop moves the session to FAILED and it is absent from the live roster.
  5. roster() reflects exactly the sessions acquired-minus-released, joined with live status.

Seams left open (deliberately)

  • CB-302 — attach a checkpoint step (STATE.md + commit) to the release path.
  • CB-303 — a policy loop over roster() using spawnedAtNanos/state to auto-release on idle_ttl, or drain on context_cap.
  • CB-304 — fleet_list reads roster() for a bridge-owned roster + live join.