Files
fleetd/bridged
kevin cf4ad186ab
CI / build (push) Successful in 1m30s
CB-516: fail a delegation when its worker session is released
Tearing a worker down left the send that was waiting on it stranded. The
rendezvous waiter stayed open, so a blocking bridge_send kept blocking and an
async one kept reporting PENDING until ASYNC_TIMEOUT_MS — thirty minutes —
even though the worker provably no longer existed and the delegation could
never complete.

Observed repeatedly this session: a task sitting at
{"phase":"pending","detail":"worker unknown"} long after I had personally
deleted the pane. The maddening part is that poll() already HAD the evidence —
it calls liveStatus() to build that detail string, gets back "unknown", and
reports PENDING anyway. It also never reached /metrics under any outcome, so a
stalled delegation was invisible to both the task view and the dashboard. The
only way I ever diagnosed one was reading the worker's pane over the herdr
socket by hand.

Fixed at the choke point rather than by guessing from status strings:
SessionManager.release() is the single path every teardown funnels through
(REST stop, MCP stop, idle-TTL reaper, recycle, shutdown drain), so it now
notifies a release listener with the terminalId, and Bridged.main wires that to
MessageService.abandon(). Abandon fails the open waiter with a real reason.

Deliberately NOT done by inferring "gone" from liveStatus(): that method
collapses a vanished pane, a wedged worker and a herdr hiccup into the same
"unknown" string, so acting on it would fail live delegations during a
transient blip. An explicit lifecycle signal cannot be ambiguous.

Notified before launcher.stop() so a blocked caller fails fast, and wrapped so
a listener failure can never prevent the teardown it is reacting to.

Resolving as a failure rather than letting it time out also means the outcome
is counted — a torn-down delegation now shows up as
sends_total{outcome="failed"} instead of nothing at all.

353 tests (was 346): abandon fails a waiting send / is a no-op with no waiter /
never clobbers a send the worker already answered, an abandoned async task
polls FAILED rather than PENDING, release notifies with the right terminal,
releasing an unknown pane notifies nobody, and a throwing listener does not
block teardown.

Verified live on the running daemon, reproducing the original scenario:
  async send        -> {"phase":"pending","detail":"worker working"}
  DELETE the worker -> 204
  poll              -> {"phase":"failed","detail":"the worker session was
                        released before it replied"}
  /metrics          -> bridged_sends_total{outcome="failed"} 1
Previously that poll returned PENDING for thirty minutes and the counter stayed
empty.

NOT addressed here, and worth deciding separately: a worker that simply never
replies (rather than being released) still rides out the full 30-minute
ASYNC_TIMEOUT_MS. That is a policy question about long-running delegations, not
a correctness bug.
2026-08-01 23:57:45 +07:00
..
2026-08-01 21:34:59 +07:00