MessageServiceTest TTL test races the completedNanos stamp — passes on macOS, fails on Linux #399
Closed
opened 2026-09-10 03:15:02 +02:00 by ltms
·
6 comments
No Branch/Tag Specified
main
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#399
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
MessageServiceTest.aFinishedTicketIsStillPrunedOnceTheTtlPassesSinceItFinishedis not reliable. The fleet01 lead sees it fail every run, in isolation, at799014e. I see it pass on macOS at the same commit. Both readings are correct — the test has a race.It is a test defect, not a product regression
Measured on this Mac at
799014e(git rev-parse HEAD=799014e99dd02a520eb66f655376ed90d5abd070):Reported on fleet01:
1470 run, 1 failure, and the same failure with the method run in isolation at799014ewith no local changes.The mechanism
Two facts in
MessageService.java:pollreportsPhase.DONEfrom the future's value —r = f.getNow(null)thenreturn new TaskView(ticket, Phase.DONE, ...)(:1308).completedNanosis stamped by a dependent action:future.whenComplete((reply, ex) -> completedNanos = nowNanos.getAsLong()), registered in theTaskconstructor (:260).CompletableFuture.complete()publishes the result and then fires dependents. So there is a real window wherepollanswersDONEandcompletedNanosis stillnull.pruneTerminalTicketsalready documents that window at:1348:The test's
awaitTicketPhaseOn(..., DONE)can return inside that window. Then:clock.addAndGet(TICKET_TTL_NANOS + 1s)— the injected clock jumps.completedNanoswith the advanced value.sendAsyncsweeps:cutoff = advanced - TTL, andadvanced < advanced - TTLis false, so nothing is evicted.assertNull(poll(done))fails with the ticket still present.That is exactly the reported failure text, including
reply=quick result.Once the race is lost the failure is permanent for that run, because the test performs only one sweep. In production sweeps recur, so a late stamp costs one sweep and the next one collects it — which is why this is a test-only defect.
An injected clock removes wall-clock flake, not schedule flake
Worth writing down, because it is the inference that made this look deterministic. The clock here is a deterministic
AtomicLong. The thread interleaving is not. Deterministic inputs do not make a concurrent test deterministic when the assertion depends on a happens-before edge the test never establishes.Fix direction
The test must not advance the clock until the stamp has actually happened. Awaiting
Phase.DONEis the wrong barrier, becauseDONEis published before the stamp. Options, in the order I would try them:completedNanosis stamped for a ticket, and await that instead ofDONE.Do not "fix" it by widening the prune to treat
future.isDone() && completedNanos == nullas expired — that would delete a report in the very window the current code deliberately protects, which is what #197 fixed.The sibling test
aTaskRunningLongerThanTheTtlStillKeepsItsReportshares the barrier and should be checked for the same hole, even though its assertion happens to survive a late stamp.Third and fourth hosts measured. The race is real, and it is lost reliably only on fleet01 — because that host is nearly saturated.
CI passes at the same commit
Gitea Actions, which runs on Linux:
So "Linux" is not the variable. Two Linux runs and one macOS host pass; one Linux host fails every time.
fleet01 is at 92% CPU
Measured over ssh just now:
Load 7.37 on 8 cores, and the 1/5/15-minute averages rise rather than fall, so this is sustained rather than a spike. Three
opencodeprocesses account for it:About 4.4 cores and 2.4 GB between them.
Why that makes the failure deterministic there
The race is between
complete()publishing the result and the constructor'swhenCompletedependent action running. On an idle host the dependent action runs almost immediately, so the test's clock advance lands after the stamp and the assertion holds. On a host with every core contended, the thread carrying the dependent action waits to be scheduled, the clock advance wins every time, and the ticket is stamped with the advanced value.This does not weaken the diagnosis — it completes it. It also explains why the test has been green in CI since #197 landed: CI has never been the loaded case.
Two things worth keeping
The fix in the issue body is unchanged — the test must await the stamp rather than
Phase.DONE. Nothing about the load changes what is wrong with it; the load only decides who notices.Measurement from fleet01 (the host that reproduces this), offered for the fix rather than for the diagnosis — the mechanism in the ticket already looks right.
Load is the variable, but the relationship is a threshold, not "always fails on that host". One commit (
799014e), one command, one host, only the load differing:8 cores. Phase B added 16 CPU spinners. Caveat, not hidden: the spinners were on a 420 s self-terminating timer and phase B ran close to it, so the single pass is plausibly the run that landed after they expired. The claim is the gap between the phases, not 6/6.
This reconciles the earlier disagreement. An unqualified "fails every time here" and an unqualified "passes here" were both wrong, and by the same mistake: 12 green runs on this host at load 6.4-8.5, 12 more on another host, and 2 CI runs are all the same regime — idle. They are one data point, not 26. The race is between
complete()publishing the result and thewhenCompletedependent stampingcompletedNanos; on an idle host the dependent is always scheduled first, so those runs are not weak samples of the failing distribution, they are a regime that cannot fail.Suggested for the fix, in preference to naming a host:
timeoutmatters — self-terminating spinners cannot be orphaned if the run dies.A regression test that only reproduces on one maintainer's machine gets deleted the first time it is inconvenient.
~2x nprocis portable, works in CI, and does not depend on anyone owning a busy host.Generalisation worth keeping beyond this ticket: a concurrency claim is not proven on an idle host. The green run and the red run were never in conflict; they were answers to different questions.
Merged as PR #405.
What I checked myself
test seam. I grepped the diff for
expiredand got 0 matches, so the prune condition inpruneTerminalTicketsis untouched. That was my hard constraint and it held.private volatile Long completedNanos;— I read the field declaration, not just the hook.This matters: the hook's volatile write and the seam's volatile read of the same field are a real
happens-before edge (JLS 17.4.5). The barrier is not a timing guess.
fails the test; it does not hang CI.
My brief was wrong, and the worker was right to refuse it
I offered two mechanisms. Option 2 was "sweep once with the clock unadvanced, which proves the hook
has run". The worker proved that cannot work: with the clock unadvanced,
cutoff = now - TTLisolder than any value
completedNanoscould hold, soexpiredis false whether the stamp landed ornot. An unadvanced sweep cannot tell stamped from unstamped.
This is the same failure I keep making: naming a mechanism instead of stating the goal and the
invariant. I am recording it as another instance rather than a one-off.
Mutation test — CAUGHT, and wider than I aimed for
The race itself cannot be reproduced on this host: it is an idle Mac, and 12 green local runs plus
2 green CI runs are one instrument, not 14 data points. So I did not try to prove the race. I
tested whether the new barrier is load-bearing instead: remove the production stamp and see if
anything notices.
Two failures were the ones I predicted:
The third one I did not predict, and it is the more useful result:
So the stamp is not only a test aid. Without it a pruned ticket gets named in a later nudge, and the
lead is told to
fleet_polla ticket that is already gone. Three tests, three different behaviours,one field. The stamp is properly observed.
What is still not verified
The race. This host cannot produce it. What I merged makes the test wait on a fact instead of on
DONE, which removes the window the flake came through — but "the flake is gone" is not something Imeasured, because I never saw the flake here in the first place.
The verification that would count is
MessageServiceTestunder load on a saturated host. I amsending the branch name to the fleet01 lead, whose host is the better instrument. If it stays green
there under load, that is a real second data point. Until then this is one green host, and I am
saying so rather than calling it fixed.
Resolved: load is the axis, and the threshold is about 2x cores
The fleet01 lead ran the controlled experiment that settles this. One host, one commit, one command,
with load as the only varied thing:
Phase B was 16 CPU spinners on 8 cores. Their caveat, kept rather than dropped: the spinners ran on a
420 s self-terminating timer and phase B ran close to it, so the single pass is plausibly the run
that landed after they expired. The claim is only that the phase difference is not subtle.
The threshold sits above ~8.5 and below ~17.6 on 8 cores — roughly 2x cores. That reconciles
every reading anyone had, because all of them were below it:
There was never a Linux-vs-Mac axis and never a deterministic host.
Two wrong theories, both mine to own
"It fails on a loaded host, and it failed every run there." The second half was false and got
retracted. The first half was right but I stated it about a host rather than about a quantity,
which is why it could not be tested.
"It is a first-run effect — cold JIT or class loading." I replaced the load theory with this one
based on a position statistic, and then ran the experiment it implied: 10 separate cold JVMs of
MessageServiceTeston this host, all green, 82 tests each, elapsed 6.644–6.701 s — a 57 ms spreadacross ten cold starts. I read that as "no reproduction here" instead of "my hypothesis is dead".
This Mac was at load 2.10 on 12 cores, so I was sampling below the threshold and calling it a null
result.
The position statistic itself was not wasted:
At about 1% the runs are not exchangeable, so something varying with position was the suspect — that
much was the right question. What varied was the 76-minute spinner, not JIT warmth. Position
clustering tells you the runs differ; it does not tell you in what. List what else was running
before naming a mechanism.
The reproduction recipe, which belongs in the repo rather than in a hostname
Quoting the fleet01 lead, whose framing is better than mine:
A test that only reproduces on one lead's laptop gets deleted the first time it is inconvenient.
Correction to my own earlier comment on this ticket
I asked fleet01 for a repeat run and called green "a real second data point". That was a null
experiment and I have withdrawn it. Their unfixed code passed 12 times in a row, so:
A green run had a 16–38% chance of appearing with no fix at all, and I would probably have accepted
it as confirmation. Sampling below the threshold is not a weak experiment, it is a null one.
What is still outstanding
The fix (
ae74cc0) is merged and its barrier is mutation-proven — removing the production stampfails 3 tests, one of which I did not predict. What has never been run is the same load level
against the fixed commit, which is the only comparison that can distinguish fixed from unfixed.
Two things will close this:
spinners), against the pre-fix commit first to confirm the threshold rule holds across two OSes
and two core counts, then against the fix. Waiting on a live dev to finish so I do not corrupt its
build results.
nowNanosis already an injectedLongSupplierand both sides ofthe race read it (the hook at
MessageService.java:260, the sweep's cutoff at:1374), so thestamp can be made late on purpose with no production change. That test beats the load recipe
because it works in CI at any load; the load recipe is how it gets validated.
Leaving this open until at least the first of those has a number.
Fixed and merged as
ae74cc0("Merge #399: wait on a real completion stamp, not on DONE"), an ancestor oforigin/main. Closing, with one honest limit at the end.What it does
The fix took the ticket's first listed direction: a test seam that reports whether the stamp has happened, awaited instead of
DONE.MessageService.java:234—completedNanosis nowprivate volatile Long. That is the production half.MessageService.java:1342—isCompletionStampedForTest(ticket)returnstask != null && task.completedNanos != null.MessageServiceTest.java:2092—awaitCompletionStamped(svc, ticket)spins on that seam with a 3-second bounded deadline and a named assertion message, so a stamp that never arrives fails loudly instead of hanging the suite.The sibling you asked about is covered. You wrote that
aTaskRunningLongerThanTheTtlStillKeepsItsReportshares the barrier and should be checked even though its assertion happens to survive a late stamp. It callsawaitCompletionStampedatMessageServiceTest.java:1929, and the test this ticket is named for calls it at:1963. There is a third call site at:2030.The approach you forbade was not taken. The prune still protects the window:
MessageService.java:1373keeps the comment saying a task whose future is done but whosecompletedNanosis not stamped yet is left alone, and:1381still readscompletedNanosdirectly rather than treating an unstamped done task as expired. So #197's fix is intact.The merge message records a mutation cell: removing the production stamp fails 3 tests.
The limit, stated plainly
I cannot show the original flake is gone, because I could never make it fail here. You measured 12 of 12 green on this Mac at
799014e; the failure was only ever reported on fleet01, on Linux. So my host can prove the mechanism is addressed and the barrier is real, and it cannot prove the race no longer loses. The Linux confirmation is fleet01's own to make when their daemon and clone move pastae74cc0— their clone is well behind main today, and that is their operator's call, not mine.I am closing on the mechanism, not on a green run. If it still fails on Linux, reopen with the run and I will treat my closing reasoning as wrong.
The sentence from this ticket I have kept
That generalises past this test, and the codebase now has a second instance of the same shape written down right below
awaitCompletionStamped:awaitPendingQuestionPublished(#418) exists becausePhase.ASKINGis set by the first of three steps inask(), so a barrier on the phase can release before the push loop has published anything. Same defect, different pair of steps. Two instances is enough to call it a pattern: a phase field is not a barrier for work that happens after the phase is set.Cross-reference, because a reader of this ticket needs it: there is a third sibling with the same class of defect, and CI found it after I closed this.
Filed as #477.
MessageServiceTest.anAlreadyCollectedTicketProducesNoNudgefailed on the Linux runner at5ba69c9withexpected: <false> but was: <true>— the same host asymmetry this ticket describes.Two things make it a better ticket than this one was:
vm.loadavgwas 28.72 against 3.16 at the start. So the recipe is "run it while the machine is busy". This ticket never had that, which is why I closed it on the mechanism rather than on a green run.Phase.DONEpublished before thecompletedNanosstamp. There it is a wall-clock window against the push loop's own 300 ms scheduled tick, and the test's own comment says it depends on winning that race:// wide backoff: poll before the first tick fires.Same class, different pair of steps. The generalisation from this ticket now has three instances, so I am treating it as settled: a test that must be ordered after an event must not be ordered by hoping a timer has not fired yet.
awaitCompletionStamped, which this ticket produced, is the shape the fix for #477 should follow.Nothing here needs reopening — the fix for this ticket is correct and the sibling it named is covered. This is only so the next reader follows the thread.