Async ticket TTL runs from creation, so a long task's report is destroyed on arrival #197

Closed
opened 2026-08-31 04:25:47 +02:00 by ltms · 0 comments
Owner

What happened

I lost two workers' complete reports today. Both had finished successfully. Both members were still
alive and done. fleet_poll{ticket} said unknown ticket: task-1 (never issued, or expired) and
fleet_poll{target} returned [].

The work itself survived only because the workers had pushed branches. Had the task been a question,
an analysis, or anything not committed to git, the whole run would have been wasted.

Root cause

MessageService.pruneTerminalTickets (fleetd/src/main/java/dev/ltms/fleet/msg/MessageService.java:825):

long cutoff = nowNanos.getAsLong() - TICKET_TTL_NANOS;
tasks.entrySet().removeIf(e -> {
    Task t = e.getValue();
    boolean expired = t.future.isDone() && t.createdNanos < cutoff;
    ...
});

TICKET_TTL_NANOS is 10 minutes (line 62). The cutoff is compared against t.createdNanos — the
moment the ticket was issued — not the moment it completed.

So the collection window is not "10 minutes to collect". It is 10 minutes minus however long the
task took
. Concretely:

task runtime time left to collect the report
1 min 9 min
9 min 1 min
10 min 0 — pruned on the first sweep after it completes
20 min 0 — the report is discarded the moment it lands

Any wait:false delegation that runs longer than the TTL loses its report unconditionally. That is
the normal case for real work: all three of today's workers ran well past 10 minutes.

Why the inbox does not save it

The reply lives only in Task.future. Pruning the map drops it. Nothing spills a pruned ticket's
reply into the target's reply inbox, which is why fleet_poll{target} came back empty rather than
holding the answer. The existing note that "a timed-out ticket's real answer lands in the reply
inbox" is true for a timed-out blocking send, not for a pruned async ticket.

Why the CB-588 nudge does not save it either

The nudge fires when the ticket goes terminal, so in principle the lead is told to collect. But the
lead may be mid-turn on something else, and the reminder cap is 5. With the TTL already exhausted by
the task's own runtime, the nudge can arrive after — or within seconds of — the prune that destroys
what it is pointing at. Today the nudge for one ticket arrived and the report was already gone.

Suggested fix

Measure the TTL from completion, not creation. Record completedNanos when the future resolves
and prune on t.completedNanos < cutoff. That gives every ticket the full documented window whatever
its runtime, and still bounds tasks.

Worth considering alongside it:

  • Do not prune an uncollected ticket at all while its member is still alive and done. The
    member is holding a pane and a worktree; the few hundred bytes of its report are not the thing
    bounding memory here.
  • On prune, push the reply into the target's reply inbox instead of discarding it, so
    fleet_poll{target} is the honest fallback the error message implies.
  • The poll error text should distinguish "never issued" from "expired — the report was discarded".
    Those tell the operator completely different things, and right now they are the same string.

Reproduce

  1. fleet_send{sessionId, content, wait:false} with a task that takes more than 10 minutes.
  2. Wait for it to finish.
  3. fleet_poll{ticket} -> unknown ticket. fleet_poll{target} -> [].

Observed 2026-08-31 on main 2d55b0b, daemon pid 25577.

## What happened I lost two workers' complete reports today. Both had finished successfully. Both members were still alive and `done`. `fleet_poll{ticket}` said `unknown ticket: task-1 (never issued, or expired)` and `fleet_poll{target}` returned `[]`. The work itself survived only because the workers had pushed branches. Had the task been a question, an analysis, or anything not committed to git, the whole run would have been wasted. ## Root cause `MessageService.pruneTerminalTickets` (`fleetd/src/main/java/dev/ltms/fleet/msg/MessageService.java:825`): ```java long cutoff = nowNanos.getAsLong() - TICKET_TTL_NANOS; tasks.entrySet().removeIf(e -> { Task t = e.getValue(); boolean expired = t.future.isDone() && t.createdNanos < cutoff; ... }); ``` `TICKET_TTL_NANOS` is 10 minutes (line 62). The cutoff is compared against **`t.createdNanos`** — the moment the ticket was *issued* — not the moment it completed. So the collection window is not "10 minutes to collect". It is **10 minutes minus however long the task took**. Concretely: | task runtime | time left to collect the report | |---|---| | 1 min | 9 min | | 9 min | 1 min | | 10 min | **0 — pruned on the first sweep after it completes** | | 20 min | **0 — the report is discarded the moment it lands** | Any `wait:false` delegation that runs longer than the TTL loses its report unconditionally. That is the normal case for real work: all three of today's workers ran well past 10 minutes. ## Why the inbox does not save it The reply lives only in `Task.future`. Pruning the map drops it. Nothing spills a pruned ticket's reply into the target's reply inbox, which is why `fleet_poll{target}` came back empty rather than holding the answer. The existing note that "a timed-out ticket's real answer lands in the reply inbox" is true for a *timed-out blocking send*, not for a *pruned async ticket*. ## Why the CB-588 nudge does not save it either The nudge fires when the ticket goes terminal, so in principle the lead is told to collect. But the lead may be mid-turn on something else, and the reminder cap is 5. With the TTL already exhausted by the task's own runtime, the nudge can arrive after — or within seconds of — the prune that destroys what it is pointing at. Today the nudge for one ticket arrived and the report was already gone. ## Suggested fix Measure the TTL from **completion**, not creation. Record `completedNanos` when the future resolves and prune on `t.completedNanos < cutoff`. That gives every ticket the full documented window whatever its runtime, and still bounds `tasks`. Worth considering alongside it: - **Do not prune an uncollected ticket at all while its member is still alive and `done`.** The member is holding a pane and a worktree; the few hundred bytes of its report are not the thing bounding memory here. - **On prune, push the reply into the target's reply inbox** instead of discarding it, so `fleet_poll{target}` is the honest fallback the error message implies. - The `poll` error text should distinguish "never issued" from "expired — the report was discarded". Those tell the operator completely different things, and right now they are the same string. ## Reproduce 1. `fleet_send{sessionId, content, wait:false}` with a task that takes more than 10 minutes. 2. Wait for it to finish. 3. `fleet_poll{ticket}` -> `unknown ticket`. `fleet_poll{target}` -> `[]`. Observed 2026-08-31 on main `2d55b0b`, daemon pid 25577.
ltms closed this issue 2026-08-31 05:10:46 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#197