#197: measure the async ticket TTL from completion, not from creation #198
Reference in New Issue
Block a user
Delete Branch "fix/cb-197-ticket-ttl-from-completion"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Fixes #197.
The bug
pruneTerminalTicketscompared the cutoff againstcreatedNanos:So the window to collect a reply was
TTLminus however long the task ran, notTTL:Every delegation that runs longer than the TTL loses its report unconditionally. That is the normal
case here: three workers in one session ran well past ten minutes and two of their complete reports
were destroyed. Only their pushed branches saved the work.
There is no fallback. The reply lives only in
Task.future, so pruning the map discards it, andfleet_poll{target}returns[]rather than holding it.The fix
TaskstampscompletedNanosfrom awhenCompletehook registered in its constructor, and theprune measures from that. Registering it in the constructor means every completion path stamps
it — a reply, the CB-106 completion fallback, a timeout, a failure, an abandon on teardown — without
each one having to remember to.
The stamp is a boxed
Longrather than alongwith a sentinel, on purpose:System.nanoTimemayreturn any value, zero and negatives included, so no numeric sentinel can mean "not stamped yet". A
task whose future is done but whose stamp has not landed is left alone; the next sweep collects it.
createdNanoshad no other reader, so it is removed rather than left as a field an inspection wouldflag.
What this does not change
The TTL still bounds
tasks. An uncollected finished ticket is still evicted once the TTL passessince it finished, and that is pinned by its own test — measuring from completion would be a
memory leak if a finished ticket were then kept forever.
Verification
Full
mvn clean install: Tests run: 1032, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS.Proved by revert. Restoring the old comparison (a creation stamp taken from the same injected clock,
so the sabotage is faithful rather than mixing a real clock with the test's fake one) fails the new
test:
Restored afterwards and confirmed zero
TEMP SABOTAGEmarkers remain.Not verified: IDE inspections.
fleetdis not currently open in IntelliJ on this machine, soide_diagnosticscould not run. The Maven build covers compile errors but not inspections.Left for a follow-up, deliberately not in this PR
Two further improvements from the issue, neither needed to stop the data loss:
fleet_poll{target}becomes the honest fallback the error message already implies.polldistinguish "never issued" from "expired". Today both return the same string, and theytell an operator completely different things.
pruneTerminalTickets compared the cutoff against createdNanos, so the real window to collect a reply was "TTL minus however long the task ran". A delegation that ran longer than the 10-minute TTL was already past the cutoff the moment it finished, so the next prune destroyed its reply. That is the normal case here, not an edge case. Real delegated work runs well past ten minutes. Three workers in one session did, and two of their complete reports were lost. The reply lives only in Task.future, so pruning it discards the worker's whole report, and fleet_poll{target} returns [] rather than holding it — there is no fallback. Task now stamps completedNanos from a whenComplete hook registered in its constructor, so every completion path stamps it (a reply, the completion fallback, a timeout, a failure, an abandon on teardown) without each one having to remember to. The stamp is a boxed Long, not a long with a sentinel: System.nanoTime may return any value, so no number can mean "not stamped yet". A task that is done but not yet stamped is left for the next sweep. createdNanos had no other reader and is removed. The TTL still bounds tasks — an uncollected finished ticket is still evicted once the TTL passes since it finished. Both halves are pinned by a test, and the first one was proved by restoring the old comparison and watching it fail. Tests run: 1032, Failures: 0, Errors: 0, Skipped: 0