Features: keep a dirty worktree (CB-576), fail a ticket on terminal health (CB-580)
+56
@@ -45,6 +45,8 @@ six weeks, and the table alone will not carry it.
|
||||
| [Worktree-hostile config isolation](#worktree-hostile-config-isolation) | automatic | CB-543 | `session/GitWorktrees` |
|
||||
| [Worker opens its own PR](#worker-opens-its-own-pr) | `gitTokenEnv:` / `gitHostEnv:` | CB-302 | `worker/HerdrPeerLauncher` |
|
||||
| [Session lifecycle caps](#session-lifecycle-caps) | `lifecycle:` | CB-303 | `session/SessionManager` |
|
||||
| [Keep a worktree that still holds work](#keep-a-worktree-that-still-holds-work) | automatic | CB-576 | `session/GitWorktrees` |
|
||||
| [Fail a ticket when its member dies](#fail-a-ticket-when-its-member-dies) | `health.enabled: true` | CB-580 | `health/FleetHealthMonitor` |
|
||||
| [Durable reply inbox](#durable-reply-inbox) | `broker:` | CB-307 | `msg/AmqpReplyInbox` |
|
||||
| [Reject overlapping rendezvous](#reject-overlapping-rendezvous) | automatic | CB-548 | `msg/Rendezvous` |
|
||||
| [Pin an opencode endpoint](#pin-an-opencode-endpoint) | profile `baseUrl:` | CB-508 | `worker/OpenCodeLauncher` |
|
||||
@@ -964,6 +966,60 @@ classifier `false` means "no fault" rather than "not known yet".
|
||||
|
||||
---
|
||||
|
||||
## Keep a worktree that still holds work
|
||||
|
||||
**What.** Before a finished session's git worktree is deleted, bridged checks whether it still holds
|
||||
uncommitted changes. If it does, the directory is kept and a WARN names its path, the pane and the
|
||||
release cause. A clean worktree is removed as before.
|
||||
|
||||
**On.** Always on, for every worktree-backed member. There is no knob.
|
||||
|
||||
**Why.** A member's uncommitted work exists in exactly one place — its worktree — so deleting it is
|
||||
loss with no copy and no error. CB-544 already protected the shutdown drain for this reason, but left
|
||||
the ordinary `COMPLETED` release deleting with `--force`. That gap fired: the idle reaper released two
|
||||
members and deleted both worktrees, and only luck decided the work had already been pushed. A worker
|
||||
that ends a turn without committing — because it stopped to ask a question, or refused the turn — is
|
||||
the normal case, not the rare one.
|
||||
|
||||
**Gotcha.** The check is `git status --porcelain` with **no** `--untracked-files=no`, so an untracked
|
||||
file counts as dirty. That is deliberate: the work at risk in the original incident was a new file that
|
||||
was never `git add`ed, and ignoring untracked files would have missed exactly it. The cost is that a
|
||||
profile whose parity overlay ever copies an untracked, non-gitignored file would make *every* release
|
||||
preserve, and worktrees would pile up silently. Inert today — tracked overlay files carry
|
||||
`--skip-worktree` so `--porcelain` cannot see them, and `bridged.yaml` is gitignored — but it is a real
|
||||
constraint on `overlayParity`, tracked in CB-581.
|
||||
|
||||
Second gotcha: `hasUncommitted` tolerates a worktree that is already gone and reports it clean. It has
|
||||
to. It runs inside `SessionManager.release()` *after* the registry entry is dropped and *before* the
|
||||
pane is stopped, so throwing there would orphan a live pane and strand a `bridge_send` caller on a
|
||||
rendezvous nothing resolves. Anything added to that window needs the same tolerance.
|
||||
|
||||
---
|
||||
|
||||
## Fail a ticket when its member dies
|
||||
|
||||
**What.** When fleet health sees a member reach a terminal state — `GONE` or `NEVER_READY` — every
|
||||
ticket waiting on that member is failed straight away, naming the state as the reason, instead of
|
||||
staying `PENDING` until something else notices.
|
||||
|
||||
**On.** The same `health:` block that turns on [health watching](#watch-the-fleets-health). No separate
|
||||
key.
|
||||
|
||||
**Why.** Detection without action just moves the silence. A lead that fires `bridge_send{wait:false}`
|
||||
and polls its ticket gets `pending` forever when the member behind it is already gone — the failure is
|
||||
known inside the daemon and invisible to the only caller who cares. Routing it through CB-568's
|
||||
existing idempotent target-wide failure means the outcome is also *counted*, so a dead delegation stops
|
||||
being invisible to `/metrics`.
|
||||
|
||||
**Gotcha.** It fires on the **transition** into the terminal state, not on every tick. An earlier
|
||||
attempt put the call outside the transition guard, so a member that stayed `GONE` had the failure
|
||||
operation invoked once per interval for as long as it remained in the roster; that commit was rejected.
|
||||
The flip side is the honest limitation: the new state is recorded *before* the bounded retries run, so
|
||||
if all three attempts throw, the tickets stay pending and no later tick retries. That path logs at WARN
|
||||
and has its own test — it is a known edge, not an oversight. The retries also carry no backoff.
|
||||
|
||||
---
|
||||
|
||||
## Backfill status
|
||||
|
||||
This page was started after the fact, so it is **not yet complete**. Entries above are written from
|
||||
|
||||
Reference in New Issue
Block a user