Dead lead tabs accumulate and make lead-to-lead coordination permanently undeliverable #359

Closed
opened 2026-09-05 02:40:58 +02:00 by ltms · 1 comment
Owner

What happens

Every LeadLauncher auto-launch creates a new herdr tab with the configured label
(fleet.leaders.<name>.tab). The tab of the previous lead is never closed or reused. herdr session
state survives a herdr server restart, so the dead tabs stay.

LeadTabScanner matches tabs by that exact label, so it reports every dead tab as a live lead.

LeadCoordLoop.resolveLocalLead() then cannot resolve a target:

  1. no lead is named after coordinator.selfId (the name comes from the config key, not the host), and
  2. the "sole lead" fallback needs known.size() == 1.

So it takes step 3 and holds the message. Nothing is lost, but nothing is ever delivered either.

Measured on fleet01, 2026-09-05

restart 1   lead panes: {term_65ab0ce04298e3=opus, term_65ab0cf22ad394=opus}          2 known
restart 2   (fleetd only)                                                             2 known
restart 3   lead panes: {term_...e47873=opus, term_...e5d584=opus, term_...f00b45=opus}  3 known

with, every 3 seconds:

WARN  LeadCoordLoop - lead coordination: 3 leads are known and none is named "fleet01"
      - cannot tell which pane a peer message is for; name one lead after coordinator.selfId to fix this
DEBUG LeadCoordLoop - lead coordination: 1 message(s) waiting but no local lead pane to deliver to

One real claude process was running the whole time (ps confirmed a single
claude --mcp-config ... --model claude-opus-5). The other two tabs were dead shells.

The only way I could clear it was to stop herdr and archive
~/.config/herdr/sessions/<name>/, then restart. After that: lead panes: {term_65ab1a061dfa33=opus},
and the held message delivered immediately — delivered message 29ffdea3-... from mac to lead term_65ab1a061dfa33. The message had survived three restarts, so durability is fine; only
resolution was broken.

Why this matters more than it looks

LeadTabScanner's own javadoc already names the hazard:

a label left behind by a session that has since died would read as a live lead forever

Nothing acts on it. The failure is silent and delayed: coordination works on a fresh host, then
stops for good after the second daemon restart. The WARN suggests renaming a lead after
selfId, but that is not a safe fix here — with two tabs sharing one name, step 1 returns
whichever it finds first, which may be the dead pane. The message would then be typed into a dead
shell and acked, turning "held forever" into "silently lost". Making the advice safe depends on
fixing the duplicates.

Suggested direction (candidate, not a decided fix)

Any one of these would close it; I have not tested them:

  • have LeadLauncher reuse an existing tab with the configured label instead of creating another;
  • have it close/relabel the tab it is replacing when it finds a dead lead;
  • have LeadTabScanner drop tabs whose pane has no live process, so dead labels stop counting;
  • in resolveLocalLead, treat "several known, all the same name, one live" as unambiguous.

The last one alone is not enough — the tabs would still pile up.

Notes

  • coordinator.selfId should stay a host identity, not a lead profile name. Both fleets here
    run a lead called opus; a profile-named coord-id would create one lead.opus.inbox queue that
    two daemons consume off a shared vhost.
  • Reproduce by restarting fleetd twice on a host with an auto-launched lead and reading the
    lead panes: line.
## What happens Every `LeadLauncher` auto-launch creates a **new** herdr tab with the configured label (`fleet.leaders.<name>.tab`). The tab of the previous lead is never closed or reused. herdr session state survives a herdr server restart, so the dead tabs stay. `LeadTabScanner` matches tabs by that exact label, so it reports every dead tab as a live lead. `LeadCoordLoop.resolveLocalLead()` then cannot resolve a target: 1. no lead is named after `coordinator.selfId` (the name comes from the config key, not the host), and 2. the "sole lead" fallback needs `known.size() == 1`. So it takes step 3 and holds the message. Nothing is lost, but nothing is ever delivered either. ## Measured on fleet01, 2026-09-05 ``` restart 1 lead panes: {term_65ab0ce04298e3=opus, term_65ab0cf22ad394=opus} 2 known restart 2 (fleetd only) 2 known restart 3 lead panes: {term_...e47873=opus, term_...e5d584=opus, term_...f00b45=opus} 3 known ``` with, every 3 seconds: ``` WARN LeadCoordLoop - lead coordination: 3 leads are known and none is named "fleet01" - cannot tell which pane a peer message is for; name one lead after coordinator.selfId to fix this DEBUG LeadCoordLoop - lead coordination: 1 message(s) waiting but no local lead pane to deliver to ``` One real `claude` process was running the whole time (`ps` confirmed a single `claude --mcp-config ... --model claude-opus-5`). The other two tabs were dead shells. The only way I could clear it was to stop herdr and archive `~/.config/herdr/sessions/<name>/`, then restart. After that: `lead panes: {term_65ab1a061dfa33=opus}`, and the held message delivered immediately — `delivered message 29ffdea3-... from mac to lead term_65ab1a061dfa33`. The message had survived three restarts, so durability is fine; only resolution was broken. ## Why this matters more than it looks `LeadTabScanner`'s own javadoc already names the hazard: > *a label left behind by a session that has since died would read as a live lead forever* Nothing acts on it. The failure is silent and delayed: coordination works on a fresh host, then stops for good after the second daemon restart. The WARN suggests renaming a lead after `selfId`, but that is **not** a safe fix here — with two tabs sharing one name, step 1 returns whichever it finds first, which may be the dead pane. The message would then be typed into a dead shell and acked, turning "held forever" into "silently lost". Making the advice safe depends on fixing the duplicates. ## Suggested direction (candidate, not a decided fix) Any one of these would close it; I have not tested them: - have `LeadLauncher` reuse an existing tab with the configured label instead of creating another; - have it close/relabel the tab it is replacing when it finds a dead lead; - have `LeadTabScanner` drop tabs whose pane has no live process, so dead labels stop counting; - in `resolveLocalLead`, treat "several known, all the same name, one live" as unambiguous. The last one alone is not enough — the tabs would still pile up. ## Notes - `coordinator.selfId` should stay a **host** identity, not a lead profile name. Both fleets here run a lead called `opus`; a profile-named coord-id would create one `lead.opus.inbox` queue that two daemons consume off a shared vhost. - Reproduce by restarting fleetd twice on a host with an auto-launched lead and reading the `lead panes:` line.
Author
Owner

Fixed by PR #367, merged as 6c61355 on main.

Both halves of the defect are closed:

  • LeadTabScanner.scan() now cross-checks agent.list before joining a
    labelled tab to a terminal, so a dead tab is no longer a delivery candidate
    for LeadCoordLoop.
  • LeadLauncher.ensureLeads() now has a cleanup path at all, so repeated
    restarts stop leaving a tab behind each time.

Neither trusts one reading. A dead tab is renamed with a
[fleetd:pending-close] marker and left open; only a later reconcile that
still finds it dead closes it. The scanner grants one grace scan to a terminal
it already knew was live. That caution came straight from this ticket's own
evidence: agent.list reported 0 live on fleet01 while ps showed one real
claude, so a single-reading close would have killed the operator's live lead.

Verified on the merge: 1412 tests, 0 failures. Mutation on merge — making
PendingCloseMarker.strip() the identity function fails 4 tests.

Known limit, accepted: ensureLeads() runs at startup only, so a tab whose
agent dies mid-session stays flagged and open until the next restart.

Wiki: 11-Features.md → "Dead lead tabs are cleaned up, and a live one is
never closed".

A separate finding from the sweep is filed as #368 (PrimaryRegistry
delegation binding outlives the lead).

Fixed by PR #367, merged as `6c61355` on `main`. Both halves of the defect are closed: - `LeadTabScanner.scan()` now cross-checks `agent.list` before joining a labelled tab to a terminal, so a dead tab is no longer a delivery candidate for `LeadCoordLoop`. - `LeadLauncher.ensureLeads()` now has a cleanup path at all, so repeated restarts stop leaving a tab behind each time. Neither trusts one reading. A dead tab is renamed with a ` [fleetd:pending-close]` marker and left open; only a later reconcile that still finds it dead closes it. The scanner grants one grace scan to a terminal it already knew was live. That caution came straight from this ticket's own evidence: `agent.list` reported 0 live on fleet01 while `ps` showed one real `claude`, so a single-reading close would have killed the operator's live lead. Verified on the merge: 1412 tests, 0 failures. Mutation on merge — making `PendingCloseMarker.strip()` the identity function fails 4 tests. Known limit, accepted: `ensureLeads()` runs at startup only, so a tab whose agent dies mid-session stays flagged and open until the next restart. Wiki: `11-Features.md` → "Dead lead tabs are cleaned up, and a live one is never closed". A separate finding from the sweep is filed as #368 (`PrimaryRegistry` delegation binding outlives the lead).
ltms closed this issue 2026-09-06 07:10:47 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#359