Features: #329 — the async catch that could not report, and the ticket it hid

Dai Ha
2026-09-04 15:36:12 +07:00
parent 3dce3d1889
commit 13ef00e027
+47
@@ -4095,3 +4095,50 @@ citation is true. "Compared in `changedDeferredKeys`" and "read live off `config
about other files that a reflection test over one record cannot inspect. That disclosure is not
modesty — it is what made the two gaps in #333 findable in minutes. Hold any checker in this repo to
the same standard: say what it does not cover, next to what it does.
---
## An exception thrown after a ticket resolves now reaches the log
**What it does.** `sendAsync`'s executor used to end with a bare
`catch (Throwable t) { task.future.completeExceptionally(t); }`. That future is already completed by
then, because `finishAsyncTask` completes it on its first line. `completeExceptionally` on a
completed future returns `false` and does nothing. The exception simply vanished. The catch now
checks that return value and logs at `error` with the ticket and the target when it is `false`.
**On.** Always on (fleetd #329). Nothing to configure — look for
`async send task-N -> <target> threw after its ticket was already resolved` in `fleetd.out`.
**Why it exists.** This was not one bug, it was a blind spot over the whole async region. Measured
during #324: a deliberately broken `finishAsyncTask` threw **19 real NullPointerExceptions on the
ordinary path** while the suite reported **1340 tests green and not one log line**. Any defect that
throws after the future completes produced a passing build and a silent daemon. The daemon could not
tell an operator, and no test could tell a developer.
**One thing to know for maintenance.** When a mutation you expect to fail passes, that can mean the
failure is **invisible**, not that the code is unpinned. Add one temporary `log.error` and re-run
before you conclude anything. That is the only reason this was found.
---
## A worker's real reply completes its async ticket more often
**What it does.** `answer()` used to look the task up a second time, by `turnId`, when completing the
async ticket. That second lookup raced `ask()`'s own unlocked timeout cleanup, so a ticket could stay
`PENDING` forever after the worker had actually replied. `answer()` now completes from the `Task` it
already holds from its first lookup.
**On.** Always on (fleetd #329).
**Why it exists.** The failure was silent and it lied to the lead. `fleet_poll{ticket}` reported
`PENDING` for good, and teardown later resolved the ticket as `WORKER_FAILED` — "session released
before it replied" — long after the worker had replied. A lead reading that is told something false
about its own worker.
**One thing to know for maintenance.** **This narrows the window; it does not close it.** `ask()`
drops the `asyncTasksByTurn` entry in its `catch` (`clearAsyncQuestion(turnId, true)`) but closes the
ask later, in its `finally`. Between those two the ask is still answerable and the entry is already
gone, so `answer()`'s own first lookup returns `null` and the ticket is stranded one step earlier in
the same race. Measured on 2026-09-04: a probe firing only that first half printed
`answer=REPLIED phase=PENDING reply=null`. Open as fleetd #334, and the comment in `answer()` says so.
Do not read that `task != null` guard as complete.