diff --git a/11-Features.md b/11-Features.md index 0682d2b..f2f83ef 100644 --- a/11-Features.md +++ b/11-Features.md @@ -45,6 +45,8 @@ six weeks, and the table alone will not carry it. | [Worktree-hostile config isolation](#worktree-hostile-config-isolation) | automatic | CB-543 | `session/GitWorktrees` | | [Worker opens its own PR](#worker-opens-its-own-pr) | `gitTokenEnv:` / `gitHostEnv:` | CB-302 | `worker/HerdrPeerLauncher` | | [Session lifecycle caps](#session-lifecycle-caps) | `lifecycle:` | CB-303 | `session/SessionManager` | +| [Keep a worktree that still holds work](#keep-a-worktree-that-still-holds-work) | automatic | CB-576 | `session/GitWorktrees` | +| [Fail a ticket when its member dies](#fail-a-ticket-when-its-member-dies) | `health.enabled: true` | CB-580 | `health/FleetHealthMonitor` | | [Durable reply inbox](#durable-reply-inbox) | `broker:` | CB-307 | `msg/AmqpReplyInbox` | | [Reject overlapping rendezvous](#reject-overlapping-rendezvous) | automatic | CB-548 | `msg/Rendezvous` | | [Pin an opencode endpoint](#pin-an-opencode-endpoint) | profile `baseUrl:` | CB-508 | `worker/OpenCodeLauncher` | @@ -964,6 +966,60 @@ classifier `false` means "no fault" rather than "not known yet". --- +## Keep a worktree that still holds work + +**What.** Before a finished session's git worktree is deleted, bridged checks whether it still holds +uncommitted changes. If it does, the directory is kept and a WARN names its path, the pane and the +release cause. A clean worktree is removed as before. + +**On.** Always on, for every worktree-backed member. There is no knob. + +**Why.** A member's uncommitted work exists in exactly one place — its worktree — so deleting it is +loss with no copy and no error. CB-544 already protected the shutdown drain for this reason, but left +the ordinary `COMPLETED` release deleting with `--force`. That gap fired: the idle reaper released two +members and deleted both worktrees, and only luck decided the work had already been pushed. A worker +that ends a turn without committing — because it stopped to ask a question, or refused the turn — is +the normal case, not the rare one. + +**Gotcha.** The check is `git status --porcelain` with **no** `--untracked-files=no`, so an untracked +file counts as dirty. That is deliberate: the work at risk in the original incident was a new file that +was never `git add`ed, and ignoring untracked files would have missed exactly it. The cost is that a +profile whose parity overlay ever copies an untracked, non-gitignored file would make *every* release +preserve, and worktrees would pile up silently. Inert today — tracked overlay files carry +`--skip-worktree` so `--porcelain` cannot see them, and `bridged.yaml` is gitignored — but it is a real +constraint on `overlayParity`, tracked in CB-581. + +Second gotcha: `hasUncommitted` tolerates a worktree that is already gone and reports it clean. It has +to. It runs inside `SessionManager.release()` *after* the registry entry is dropped and *before* the +pane is stopped, so throwing there would orphan a live pane and strand a `bridge_send` caller on a +rendezvous nothing resolves. Anything added to that window needs the same tolerance. + +--- + +## Fail a ticket when its member dies + +**What.** When fleet health sees a member reach a terminal state — `GONE` or `NEVER_READY` — every +ticket waiting on that member is failed straight away, naming the state as the reason, instead of +staying `PENDING` until something else notices. + +**On.** The same `health:` block that turns on [health watching](#watch-the-fleets-health). No separate +key. + +**Why.** Detection without action just moves the silence. A lead that fires `bridge_send{wait:false}` +and polls its ticket gets `pending` forever when the member behind it is already gone — the failure is +known inside the daemon and invisible to the only caller who cares. Routing it through CB-568's +existing idempotent target-wide failure means the outcome is also *counted*, so a dead delegation stops +being invisible to `/metrics`. + +**Gotcha.** It fires on the **transition** into the terminal state, not on every tick. An earlier +attempt put the call outside the transition guard, so a member that stayed `GONE` had the failure +operation invoked once per interval for as long as it remained in the roster; that commit was rejected. +The flip side is the honest limitation: the new state is recorded *before* the bounded retries run, so +if all three attempts throw, the tickets stay pending and no later tick retries. That path logs at WARN +and has its own test — it is a known edge, not an oversight. The retries also carry no backoff. + +--- + ## Backfill status This page was started after the fact, so it is **not yet complete**. Entries above are written from