fleetd #359: dead lead tabs are cleaned up without closing a live one

Dai Ha
2026-09-06 12:10:06 +07:00
parent 4f9596be6d
commit 707d2a2f18
+34
@@ -4575,3 +4575,37 @@ everything else.
`.git/info/exclude` was rejected as the mechanism, and this is worth knowing before someone tries
it again: from a linked worktree it resolves to the **common** git dir, so it would have hidden the
seeded paths in the primary checkout and every sibling worktree too.
---
## Dead lead tabs are cleaned up, and a live one is never closed
**What it does.** On startup, `LeadLauncher.ensureLeads()` now looks for tabs that carry a lead's
configured label but have no running agent. It does not close them straight away. It renames the
tab, appending ` [fleetd:pending-close]`, and leaves it open. Only a *later* reconcile that still
finds the same tab dead actually closes it. `LeadTabScanner` does the matching cross-check on the
read side: it joins a labelled tab to a terminal only when `agent.list` says that terminal is
live, and it grants one grace scan to a terminal it already knew was live.
**The knob that turns it on.** None — it is always on for any lead slot declared under
`fleet.leaders.*`. The marker suffix is a constant in `PendingCloseMarker`, and every place that
matches a tab label strips it first, so a flagged tab is still recognised as that lead's tab.
**Why it exists.** Two separate defects met here. `LeadTabScanner` joined labelled tabs straight
to terminals with **no liveness check at all**, and its javadoc excused that ("a stale name costs
nothing here"). It cost plenty: `LeadCoordLoop` reads that map to choose which pane a peer lead's
message is delivered into, so a dead tab was a valid candidate. `LeadLauncher` had the opposite
problem — no cleanup path whatsoever, so every reconcile that found 0 live leads created another
tab and left the old one behind. Restart the daemon a few times and the tab bar fills up.
**One thing to know for maintenance.** The two-reading rule is not caution for its own sake. The
evidence that opened this ticket was fleet01's own daemon log: `agent.list` reported **0 live**
while `ps` showed one real `claude` process. A first cut of this fix closed tabs on that single
reading, which would have closed the operator's live lead rather than tidying a spare tab. So the
accepted cost is stated plainly: `ensureLeads()` runs at **startup only**, so the second reading
arrives at the next restart, and a tab whose agent dies mid-session stays flagged and open until
then. That is deliberate. The bug is about repeated restarts, and one leftover tab is much cheaper
than closing a live session on evidence that has already been seen to lie.
Measured on merge: 1412 tests green; making `PendingCloseMarker.strip()` the identity function
fails 4 tests. fleetd #359.