Features: keep a dirty worktree (CB-576), fail a ticket on terminal health (CB-580)

Dai Ha
2026-08-15 09:06:33 +02:00
parent 891bd9fc0b
commit 8781f01338
+56
@@ -45,6 +45,8 @@ six weeks, and the table alone will not carry it.
| [Worktree-hostile config isolation](#worktree-hostile-config-isolation) | automatic | CB-543 | `session/GitWorktrees` |
| [Worker opens its own PR](#worker-opens-its-own-pr) | `gitTokenEnv:` / `gitHostEnv:` | CB-302 | `worker/HerdrPeerLauncher` |
| [Session lifecycle caps](#session-lifecycle-caps) | `lifecycle:` | CB-303 | `session/SessionManager` |
| [Keep a worktree that still holds work](#keep-a-worktree-that-still-holds-work) | automatic | CB-576 | `session/GitWorktrees` |
| [Fail a ticket when its member dies](#fail-a-ticket-when-its-member-dies) | `health.enabled: true` | CB-580 | `health/FleetHealthMonitor` |
| [Durable reply inbox](#durable-reply-inbox) | `broker:` | CB-307 | `msg/AmqpReplyInbox` |
| [Reject overlapping rendezvous](#reject-overlapping-rendezvous) | automatic | CB-548 | `msg/Rendezvous` |
| [Pin an opencode endpoint](#pin-an-opencode-endpoint) | profile `baseUrl:` | CB-508 | `worker/OpenCodeLauncher` |
@@ -964,6 +966,60 @@ classifier `false` means "no fault" rather than "not known yet".
---
## Keep a worktree that still holds work
**What.** Before a finished session's git worktree is deleted, bridged checks whether it still holds
uncommitted changes. If it does, the directory is kept and a WARN names its path, the pane and the
release cause. A clean worktree is removed as before.
**On.** Always on, for every worktree-backed member. There is no knob.
**Why.** A member's uncommitted work exists in exactly one place — its worktree — so deleting it is
loss with no copy and no error. CB-544 already protected the shutdown drain for this reason, but left
the ordinary `COMPLETED` release deleting with `--force`. That gap fired: the idle reaper released two
members and deleted both worktrees, and only luck decided the work had already been pushed. A worker
that ends a turn without committing — because it stopped to ask a question, or refused the turn — is
the normal case, not the rare one.
**Gotcha.** The check is `git status --porcelain` with **no** `--untracked-files=no`, so an untracked
file counts as dirty. That is deliberate: the work at risk in the original incident was a new file that
was never `git add`ed, and ignoring untracked files would have missed exactly it. The cost is that a
profile whose parity overlay ever copies an untracked, non-gitignored file would make *every* release
preserve, and worktrees would pile up silently. Inert today — tracked overlay files carry
`--skip-worktree` so `--porcelain` cannot see them, and `bridged.yaml` is gitignored — but it is a real
constraint on `overlayParity`, tracked in CB-581.
Second gotcha: `hasUncommitted` tolerates a worktree that is already gone and reports it clean. It has
to. It runs inside `SessionManager.release()` *after* the registry entry is dropped and *before* the
pane is stopped, so throwing there would orphan a live pane and strand a `bridge_send` caller on a
rendezvous nothing resolves. Anything added to that window needs the same tolerance.
---
## Fail a ticket when its member dies
**What.** When fleet health sees a member reach a terminal state — `GONE` or `NEVER_READY` — every
ticket waiting on that member is failed straight away, naming the state as the reason, instead of
staying `PENDING` until something else notices.
**On.** The same `health:` block that turns on [health watching](#watch-the-fleets-health). No separate
key.
**Why.** Detection without action just moves the silence. A lead that fires `bridge_send{wait:false}`
and polls its ticket gets `pending` forever when the member behind it is already gone — the failure is
known inside the daemon and invisible to the only caller who cares. Routing it through CB-568's
existing idempotent target-wide failure means the outcome is also *counted*, so a dead delegation stops
being invisible to `/metrics`.
**Gotcha.** It fires on the **transition** into the terminal state, not on every tick. An earlier
attempt put the call outside the transition guard, so a member that stayed `GONE` had the failure
operation invoked once per interval for as long as it remained in the roster; that commit was rejected.
The flip side is the honest limitation: the new state is recorded *before* the bounded retries run, so
if all three attempts throw, the tickets stay pending and no later tick retries. That path logs at WARN
and has its own test — it is a known edge, not an oversight. The retries also carry no backoff.
---
## Backfill status
This page was started after the fact, so it is **not yet complete**. Entries above are written from