diff --git a/11-Features.md b/11-Features.md index e2d2268..ff8da49 100644 --- a/11-Features.md +++ b/11-Features.md @@ -1813,6 +1813,65 @@ one source of truth for what a valid policy name is. --- +## Old work-in-progress snapshots clean themselves up + +**What.** The daemon now deletes a `refs/wip/` snapshot once it is safe to — and `GET +/members` reports how many are left and roughly what they cost: + +```json +{ "workers": [ ... ], "wipRefs": { "count": 12, "costBytes": 4183042 } } +``` + +The rule has **two** conditions and both must hold: + +1. The snapshot commit's **tree is already reachable from `main`**. The content exists somewhere + else, so deleting the ref loses nothing. +2. The snapshot is **older than 24 hours**. + +Every deletion logs the ref name, the commit sha, the age, and the exact `git update-ref` command +that puts it back: + +``` +pruned snapshot ref refs/wip/worker/cb578c-92885c-1 commit=4a91f0c (age 51h): its tree is +already reachable from main, so the work is preserved; recover from reflog via +git update-ref refs/wip/worker/cb578c-92885c-1 4a91f0c +``` + +**On.** Always on; nothing to configure. The sweep rides the existing session reaper loop and runs +every 6 hours, plus once on the first pass after a restart. A fleet that has never snapshotted +anything sees no change, and `wipRefs` is simply absent from `/members` when no worktree session has +told the daemon which repo to look in. + +**Why.** CB-578 stage C commits a dirty worktree to `refs/wip/` on release, so a worker's +uncommitted work is never lost. Nothing deleted those refs. A ref is a garbage-collection root, so +every snapshot pinned its whole tree and `git gc` could never reclaim any of it. Branch names are +unique per session, so the refs pile up rather than replace each other. It was never an error — it +would have shown up as a slow `git gc` and a large `.git`, months later. + +Reachability is the safety floor, not the age. A snapshot exists *because* the work was not committed +anywhere else, so a plain time-based sweep would throw away the only copy — exactly the failure the +snapshots were built to stop. + +**Gotcha — the sweep shipped dead, and every test still passed.** `SessionReaper` held the last sweep +time in a `long` that started at `Long.MIN_VALUE` to mean "never yet". That sentinel cannot be +compared by subtraction: `System.nanoTime()` is positive, so `now - Long.MIN_VALUE` overflows to a +large negative number, the "has 6 hours passed?" gate read it as *swept moments ago*, and the method +returned **before** the assignment that would have fixed the field. The sweep never ran once, for the +life of the process, with nothing in the log to say so. + +The unit tests missed it because they all called the sweep method directly and walked around the +gate. Two lessons worth keeping: **a "never yet" sentinel belongs in its own boolean, not in a +magic value of the same field**, and **a test that calls the seam does not prove the caller reaches +it**. The test that now guards this starts the real reaper loop and requires a prune call to arrive. + +**Second gotcha — `costBytes` is a rough figure, not disk usage.** It sums the uncompressed sizes of +every blob in every snapshot's tree. Objects shared between snapshots, or shared with `main`, are +counted once per snapshot. So it reads much larger than the space actually reclaimable, and it is +useful for spotting growth over time, not for capacity planning. Nothing pushes `refs/wip/*` to the +forge; they are deliberately local and invisible to `git branch`. + +--- + ## Backfill status This page was started after the fact, so it is **not yet complete**. Entries above are written from