Files
fleetd/docs/CB-301-Session-Manager.md
T
Dai Ha 76e4b577a0
CI / build (pull_request) Successful in 1m6s
CI / contract (pull_request) Successful in 1m6s
CB-622: rename bridge_* tool names to fleet_* in Markdown docs
Rename the eleven MCP tool names (bridge_ack/ask/list/poll/profiles/reply/
send/spawn/status/stop/whoami) to their fleet_* names across the Markdown
documentation. fleet_* is written as the normal name; one deprecation note
in README.md says bridge_* still works for one release.

docs/MCP-Contract.md is renamed only inside section 6 (lines 210-293):
sections 1-5 and 7-11 are stale pre-build design text (CB-609) and are
deliberately left with old names so dead text does not look maintained.

Also renames e2e/bridge_ask_transcript.md to e2e/fleet_ask_transcript.md
to match its content. CLAUDE.md, wiki/, plugin/skills/setup/SKILL.md and
.claude/skills/port-to-opencode/SKILL.md are owned by other units and are
untouched.
2026-08-22 21:52:24 +02:00

125 lines
6.3 KiB
Markdown

# CB-301 — Session Manager (one-shot, no reuse)
**Status:** design spec for review → delegate implementation.
**Grounded in:** `WorkerService`, `Injector`/`StatusPoller`/`TurnListener`, `MessageService`,
`BridgeMcp`, `BridgedApp` (see [wiki 9. Implementation](../wiki/9-Implementation.md)).
## Problem
`WorkerService` is **stateless about what it spawned**. Its own Javadoc says it plainly:
> "there is no registry; `list()` only asks herdr." — `WorkerService.reapOrphanWorkers` (line 288)
Consequences today:
- The daemon cannot answer "which workers did *I* spawn, in what lifecycle state, owned by whom,
since when?" without shelling to herdr for a raw agent list (no state, no ownership, no age).
- Cleanup of a worker that outlived its owning process depends entirely on the boot-time
name-nonce **reaper** (CB-117) — there is no live, authoritative roster during a run.
- `fleet_list` (CB-304) can only surface herdr's view, not a bridge-owned roster.
- There is no seam for per-session policy (checkpoint on teardown → CB-302; idle_ttl /
context_cap / drain → CB-303).
## Goal & non-goals
**Goal.** Introduce a `SessionManager` that owns an authoritative in-daemon registry of the worker
sessions this daemon process spawned, tracks each one's lifecycle state, and tears each down
deterministically. It becomes the single source of truth for the roster and the seam CB-302/303/304
build on.
**Non-goals (explicit — reuse policy chosen: one-shot, no reuse).**
- **No pooling / no reuse.** Every delegated task gets a fresh worker; a finished worker is torn
down, never handed to a later task. No "warm idle" pool, no `role@profile` keying.
- **No auto-teardown *timing*.** *When* a one-shot worker is released (immediately on turn
completion vs after an idle grace) is CB-303. CB-301 provides the **mechanism** (`release`) and
the registry; CB-303 sets the policy.
- **No checkpoint content.** Writing `STATE.md` + commit on teardown is CB-302; CB-301 only exposes
the release hook it will attach to.
Under no-reuse, a released session is terminal. A new `acquire` always creates a fresh session.
## Design
`SessionManager` **wraps** `WorkerService` (does not replace it). `WorkerService` keeps doing the
subscription-guarded spawn/teardown mechanics; `SessionManager` adds the registry, lifecycle, and
ownership on top.
**Package:** new `dev.ltms.bridged.session` — keeps the registry/lifecycle concern separate from
the `worker` spawn mechanics. Holds `SessionManager` + `WorkerSession`.
### `WorkerSession` (record or small mutable holder)
| Field | Source | Notes |
|---|---|---|
| `paneId` | `Agent.paneId()` | registry key |
| `terminalId` | `Agent.terminalId()` | for status/identity joins |
| `profile` | spawn arg | which profile spawned it |
| `cwd` | resolved cwd | the worker's working dir |
| `ownerTerminal` | caller identity (nullable) | the primary/turn that requested it; `null` = daemon/anon |
| `spawnedAtNanos` | `System.nanoTime()` | age basis for CB-303 (monotonic; no wall clock in tests) |
| `state` | lifecycle FSM | see below |
State is held in a `ConcurrentHashMap<String /*paneId*/, WorkerSession>`.
### Lifecycle state machine (one-shot)
```
SPAWNING --ready(MCP present)--> READY
READY --onDelivered--> BUSY
BUSY --onTurnComplete--> DONE
BUSY --onTurnFailed--> FAILED
READY|DONE|FAILED --release()--> RELEASED (deregistered)
SPAWNING|READY|BUSY|DONE --vanished/drop--> FAILED
```
- Transitions are driven by hooks the manager already has access to:
`WorkerPresence.markPresent` → `READY`; `TurnListener.onDelivered/onTurnComplete/onTurnFailed`
(the manager implements or decorates `TurnListener`) → `BUSY`/`DONE`/`FAILED`.
- `RELEASED` sessions are removed from the registry (teardown is terminal).
- Any state → `FAILED` on drop (worker vanished / injector `drop`), mirroring `Injector`.
### API
```java
final class SessionManager {
WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal);
void release(String paneId); // deterministic teardown + deregister
Optional<WorkerSession> get(String paneId);
List<WorkerSession> roster(); // bridge-owned view (CB-304 consumes this)
// lifecycle hooks (package-private): onReady/onDelivered/onComplete/onFailed(target)
}
```
- `acquire` = `workerService.spawn(profile, requestedCwd, callerCwd)` → register `SPAWNING`.
- `release` = `workerService.stop(paneId)` → deregister. Idempotent (already-gone tolerated, matching
`WorkerService.stop`).
- `roster` joins the registry with live herdr status for a truthful "roster + live" (CB-304).
### Integration points
- **`Bridged.main`** — construct `SessionManager(workerService, ...)`; wire it as/decorating the
`TurnListener` alongside `CompletionResolver` so it sees turn boundaries, and give it the
`WorkerPresence` signal for `READY`.
- **`BridgeMcp.spawn` / `BridgedApp.spawnWorker`** — route spawn through `SessionManager.acquire`
(carry `callerTerminal` as `ownerTerminal`). **`fleet_stop` / `DELETE /workers/{paneId}`** →
`SessionManager.release`.
- **`fleet_list` / `GET /sessions` (CB-304 later)** — read `SessionManager.roster()`.
- **`MessageService`** — no change required for one-shot; a later CB-303 auto-release hook can call
`release` from `onTurnComplete` under policy.
## Acceptance (tests, no live herdr — fakes as elsewhere)
1. `acquire` registers a `SPAWNING` session with the right owner/profile/cwd; a second `acquire`
yields a **distinct** paneId and a **distinct** session (no reuse).
2. Presence signal moves `SPAWNING → READY`; a delivered turn moves `READY → BUSY → DONE`.
3. `release` tears the worker down via `WorkerService.stop` and removes it from `roster()`;
a second `release` on the same paneId is a harmless no-op.
4. `onTurnFailed` / drop moves the session to `FAILED` and it is absent from the live roster.
5. `roster()` reflects exactly the sessions acquired-minus-released, joined with live status.
## Seams left open (deliberately)
- **CB-302** — attach a checkpoint step (`STATE.md` + commit) to the `release` path.
- **CB-303** — a policy loop over `roster()` using `spawnedAtNanos`/state to auto-`release` on
`idle_ttl`, or drain on `context_cap`.
- **CB-304** — `fleet_list` reads `roster()` for a bridge-owned roster + live join.