A daemon-side crash in the completion thread is silently misattributed to the worker #412

Open
opened 2026-09-10 04:38:21 +02:00 by ltms · 1 comment
Owner

The incident

On fleet01, three members were declared useless and two were stopped for "burning 40 minutes and
returning nothing". The workers were fine. The daemon was destroying their replies:

Exception in thread "completion-term_65b17156023f61f"
java.lang.NoClassDefFoundError: dev/ltms/fleet/msg/Rendezvous$Resolution
    at dev.ltms.fleet.msg.Rendezvous.resolveCompletion(Rendezvous.java:219)
    at dev.ltms.fleet.inject.CompletionResolver.resolve(CompletionResolver.java:417)
    at dev.ltms.fleet.inject.CompletionResolver.lambda$onTurnComplete$0(CompletionResolver.java:278)

Three exceptions, matching exactly the three members that reached completion. The fourth was killed
before it got there.

Why this is the worst possible misattribution

The operator-visible symptom of a daemon-side crash is identical to a slow or stuck worker:

What the lead sees Worker stalled Completion thread threw
fleet_poll{ticket} timed_out_working timed_out_working
fleet_poll{target} empty inbox empty inbox
fleet_status working working
fleet_stop succeeds reports timeout, actually succeeds

Nothing distinguishes them. It was found by grepping the journal for a completion-fallback scrape,
on a hunch.

So a bug in the daemon makes the lead distrust the workers. That is worse than a loud crash: it
sends the operator to fix the wrong thing, and the conclusion it produces ("opencode members here
produce nothing") is durable — it gets written into notes and shapes later decisions. It cost two
members that were working correctly.

The defect

Exception in thread "completion-<terminal>" means the throwable reached the thread's default
handler. Nothing caught it, so:

  • the ticket is never failed, it just never completes
  • the reply is lost with no record against the ticket
  • the member's turn is over, so the reply cannot be regenerated

A NoClassDefFoundError is an Error, not an Exception, so a catch (Exception e) anywhere on
that path would not have helped. This needs Throwable.

The goal

A daemon-side failure while resolving a completion must fail the ticket with its cause, not hang
it.
The lead must be able to tell "the worker did not answer" from "the daemon broke while
handling the answer".

Requirements:

  • The completion path catches Throwable, records the ticket as FAILED, and carries the throwable's
    class and message as the failure reason.
  • fleet_poll{ticket} on such a ticket returns that reason. A lead reading it must be able to see
    the fault is daemon-side without reading a log file.
  • The throwable is logged at ERROR with the terminal id and the ticket.
  • Catching must not swallow: after recording, rethrow or log the full stack. A silent catch here
    would convert a visible thread-death into an invisible one, which is a worse outcome than today.

Do not try to recover or retry the completion. The member's turn is finished and its output is
gone; the only honest action is to fail the ticket and say why.

Acceptance

  • A test that injects a Throwable (use an Error, not an Exception — that distinction is the
    whole point) into the completion path and asserts the ticket reaches FAILED with a reason naming
    the cause.
  • A test that the reason is visible through fleet_poll{ticket}.
  • The existing happy path is unchanged.
  • Grep the other long-lived daemon threads for the same shape and report what you find without
    fixing it: any Thread or executor task whose body can throw past its own boundary and whose death
    is not visible to a caller. The push loop and the reaper are the obvious candidates.

Related, and the reason this went unnoticed for so long

  • #410 — liveStatus: working cannot distinguish a stuck member from a working one, so it agreed
    with the wrong diagnosis here.
  • #412 — the trigger was a jar rewritten under the live JVM. That is a separate defect, and fixing it
    does not fix this one: any daemon-side throwable on this path has the same silent outcome.

Reported by the fleet01 lead, who found it, retracted their earlier conclusion about the workers, and
asked for the ticket.

## The incident On fleet01, three members were declared useless and two were stopped for "burning 40 minutes and returning nothing". The workers were fine. The daemon was destroying their replies: ``` Exception in thread "completion-term_65b17156023f61f" java.lang.NoClassDefFoundError: dev/ltms/fleet/msg/Rendezvous$Resolution at dev.ltms.fleet.msg.Rendezvous.resolveCompletion(Rendezvous.java:219) at dev.ltms.fleet.inject.CompletionResolver.resolve(CompletionResolver.java:417) at dev.ltms.fleet.inject.CompletionResolver.lambda$onTurnComplete$0(CompletionResolver.java:278) ``` Three exceptions, matching exactly the three members that reached completion. The fourth was killed before it got there. ## Why this is the worst possible misattribution The operator-visible symptom of a daemon-side crash is **identical** to a slow or stuck worker: | What the lead sees | Worker stalled | Completion thread threw | |---|---|---| | `fleet_poll{ticket}` | `timed_out_working` | `timed_out_working` | | `fleet_poll{target}` | empty inbox | empty inbox | | `fleet_status` | `working` | `working` | | `fleet_stop` | succeeds | reports timeout, actually succeeds | Nothing distinguishes them. It was found by grepping the journal for a completion-fallback scrape, on a hunch. **So a bug in the daemon makes the lead distrust the workers.** That is worse than a loud crash: it sends the operator to fix the wrong thing, and the conclusion it produces ("opencode members here produce nothing") is durable — it gets written into notes and shapes later decisions. It cost two members that were working correctly. ## The defect `Exception in thread "completion-<terminal>"` means the throwable reached the thread's default handler. Nothing caught it, so: - the ticket is never failed, it just never completes - the reply is lost with no record against the ticket - the member's turn is over, so the reply cannot be regenerated A `NoClassDefFoundError` is an `Error`, not an `Exception`, so a `catch (Exception e)` anywhere on that path would not have helped. This needs `Throwable`. ## The goal **A daemon-side failure while resolving a completion must fail the ticket with its cause, not hang it.** The lead must be able to tell "the worker did not answer" from "the daemon broke while handling the answer". Requirements: - The completion path catches `Throwable`, records the ticket as FAILED, and carries the throwable's class and message as the failure reason. - `fleet_poll{ticket}` on such a ticket returns that reason. A lead reading it must be able to see the fault is daemon-side without reading a log file. - The throwable is logged at ERROR with the terminal id and the ticket. - Catching must not swallow: after recording, rethrow or log the full stack. A silent catch here would convert a visible thread-death into an invisible one, which is a worse outcome than today. Do **not** try to recover or retry the completion. The member's turn is finished and its output is gone; the only honest action is to fail the ticket and say why. ## Acceptance - A test that injects a `Throwable` (use an `Error`, not an `Exception` — that distinction is the whole point) into the completion path and asserts the ticket reaches FAILED with a reason naming the cause. - A test that the reason is visible through `fleet_poll{ticket}`. - The existing happy path is unchanged. - Grep the other long-lived daemon threads for the same shape and **report** what you find without fixing it: any `Thread` or executor task whose body can throw past its own boundary and whose death is not visible to a caller. The push loop and the reaper are the obvious candidates. ## Related, and the reason this went unnoticed for so long - #410 — `liveStatus: working` cannot distinguish a stuck member from a working one, so it agreed with the wrong diagnosis here. - #412 — the trigger was a jar rewritten under the live JVM. That is a separate defect, and fixing it does not fix this one: any daemon-side throwable on this path has the same silent outcome. Reported by the fleet01 lead, who found it, retracted their earlier conclusion about the workers, and asked for the ticket.
Author
Owner

Cross-reference fix: the jar-rewrite trigger is #413, not #412. #412 is this ticket — I wrote the
reference before the number was assigned.

The two are deliberately separate, and the split matters for whoever picks either up:

  • #413 stops one trigger. Even with it fixed, any daemon-side throwable on the completion path
    still hangs the ticket silently.
  • This ticket stops the silence, whatever the trigger. That is the more valuable half: a jar
    rewrite is one way to get an Error onto that thread, and it will not be the last.

Fix this one first if you only do one.

Cross-reference fix: the jar-rewrite trigger is **#413**, not #412. #412 is this ticket — I wrote the reference before the number was assigned. The two are deliberately separate, and the split matters for whoever picks either up: - **#413** stops one *trigger*. Even with it fixed, any daemon-side throwable on the completion path still hangs the ticket silently. - **This ticket** stops the *silence*, whatever the trigger. That is the more valuable half: a jar rewrite is one way to get an `Error` onto that thread, and it will not be the last. Fix this one first if you only do one.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#412