CB-586: refs/wip snapshots are swept once they are safe to drop

New entry for the retention rule (reachable-from-main AND older than 24h),
the /members wipRefs{count,costBytes} census, and two gotchas: the sweep
shipped dead because a Long.MIN_VALUE sentinel overflowed the interval gate,
and costBytes double-counts shared objects so it is not disk usage.
Dai Ha
2026-08-16 20:15:05 +02:00
parent 697e4f3de2
commit 7c4f42c58f
+59
@@ -1813,6 +1813,65 @@ one source of truth for what a valid policy name is.
---
## Old work-in-progress snapshots clean themselves up
**What.** The daemon now deletes a `refs/wip/<branch>` snapshot once it is safe to — and `GET
/members` reports how many are left and roughly what they cost:
```json
{ "workers": [ ... ], "wipRefs": { "count": 12, "costBytes": 4183042 } }
```
The rule has **two** conditions and both must hold:
1. The snapshot commit's **tree is already reachable from `main`**. The content exists somewhere
else, so deleting the ref loses nothing.
2. The snapshot is **older than 24 hours**.
Every deletion logs the ref name, the commit sha, the age, and the exact `git update-ref` command
that puts it back:
```
pruned snapshot ref refs/wip/worker/cb578c-92885c-1 commit=4a91f0c (age 51h): its tree is
already reachable from main, so the work is preserved; recover from reflog via
git update-ref refs/wip/worker/cb578c-92885c-1 4a91f0c
```
**On.** Always on; nothing to configure. The sweep rides the existing session reaper loop and runs
every 6 hours, plus once on the first pass after a restart. A fleet that has never snapshotted
anything sees no change, and `wipRefs` is simply absent from `/members` when no worktree session has
told the daemon which repo to look in.
**Why.** CB-578 stage C commits a dirty worktree to `refs/wip/<branch>` on release, so a worker's
uncommitted work is never lost. Nothing deleted those refs. A ref is a garbage-collection root, so
every snapshot pinned its whole tree and `git gc` could never reclaim any of it. Branch names are
unique per session, so the refs pile up rather than replace each other. It was never an error — it
would have shown up as a slow `git gc` and a large `.git`, months later.
Reachability is the safety floor, not the age. A snapshot exists *because* the work was not committed
anywhere else, so a plain time-based sweep would throw away the only copy — exactly the failure the
snapshots were built to stop.
**Gotcha — the sweep shipped dead, and every test still passed.** `SessionReaper` held the last sweep
time in a `long` that started at `Long.MIN_VALUE` to mean "never yet". That sentinel cannot be
compared by subtraction: `System.nanoTime()` is positive, so `now - Long.MIN_VALUE` overflows to a
large negative number, the "has 6 hours passed?" gate read it as *swept moments ago*, and the method
returned **before** the assignment that would have fixed the field. The sweep never ran once, for the
life of the process, with nothing in the log to say so.
The unit tests missed it because they all called the sweep method directly and walked around the
gate. Two lessons worth keeping: **a "never yet" sentinel belongs in its own boolean, not in a
magic value of the same field**, and **a test that calls the seam does not prove the caller reaches
it**. The test that now guards this starts the real reaper loop and requires a prune call to arrive.
**Second gotcha — `costBytes` is a rough figure, not disk usage.** It sums the uncompressed sizes of
every blob in every snapshot's tree. Objects shared between snapshots, or shared with `main`, are
counted once per snapshot. So it reads much larger than the space actually reclaimable, and it is
useful for spotting growth over time, not for capacity planning. Nothing pushes `refs/wip/*` to the
forge; they are deliberately local and invisible to `git branch`.
---
## Backfill status
This page was started after the fact, so it is **not yet complete**. Entries above are written from