The #399 completion-stamp race is verifiable deterministically — nowNanos is already injectable #409
Closed
opened 2026-09-10 04:27:34 +02:00 by ltms
·
2 comments
No Branch/Tag Specified
main
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#409
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Why this ticket exists
#399 is fixed and merged (
ae74cc0), but nobody can verify the fix, and the reason is ameasurement problem rather than a code problem.
The evidence we have is a flake. On fleet01 at
799014e, in time order:So the unfixed code passes 12 times in a row. That makes a passing run worthless as evidence
about the fix:
I confirmed the same uselessness from the other side on this host: 10 separate cold JVMs of
MessageServiceTestat current main, all green, 82 tests each, elapsed 6.644-6.701 s. Ten greenruns that tell me nothing, because the unfixed code would very likely have produced them too.
A probabilistic test of a race needs ~20-40 runs to say anything. A deterministic one needs one.
The goal
MessageServiceTestshould contain a test that fails on the pre-#399 ordering and passes on thecurrent ordering, every time, on any host, with no repeat runs and no dependence on machine load.
The invariant to pin is the ordering itself:
Why this is reachable — the seam already exists
MessageServicetakes its clock as an injectedLongSupplier(MessageService.java:342), andboth sides of the race read from that one supplier:
MessageService.java:259-260— the task'screatedNanos, and the hook:future.whenComplete((reply, ex) -> completedNanos = nowNanos.getAsLong());MessageService.java:1374— the sweep:long cutoff = nowNanos.getAsLong() - TICKET_TTL_NANOS;So a test-supplied clock can observe and influence when the hook stamps, without touching
production code. No new production seam is needed. Confirm that before designing anything — if it
turns out a new seam is required, say so in your reply and stop; do not add one on your own
judgment.
What makes the old ordering fail
MessageServiceTest.java:1955carries the explanation in a comment added by the fix:That is the mechanism to weaponise: make the hook stamp late but bounded, so that a test which
orders on
DONEproceeds, advances the clock, sweeps, and then receives a stamp carrying theadvanced value. The eviction is then missed and the bug is hidden behind a false pass.
Note the direction: the pre-fix failure mode here is a false pass, not a red test. A test that
merely goes red on the old code is not sufficient evidence you reproduced this — check which way it
fails and say so.
This mechanism is a candidate, not an instruction. A bounded delay inside the supplier is the
obvious lever, but a latch, a counting supplier that stalls on the Nth call, or something else may
be cleaner. Pick what makes the test readable and reliable, and explain the choice. If the
mechanism cannot work, say why — that answer is worth as much as the test.
Constraints
expiredexpression inpruneTerminalTickets, and do not changethe TTL. This ticket adds test coverage; it is not a behaviour change.
change, stop and report instead.
Thread.sleepas the synchronisation primitive in the assertion path. A sleep long enoughto be reliable is slow, and a short one reintroduces the flake this ticket exists to remove.
Sleeping inside the injected clock to widen the window is a different thing and is acceptable.
lands.
awaitCompletionStampedbarrier and both of its call sites.Acceptance
MessageServiceTestthat fails deterministically when the barrier call atMessageServiceTest.java:1955is removed, and passes with it present. Demonstrate bothdirections by actually running them, and paste the two results.
mvn clean installgreen. Report the realTests run:line, unpiped.determined which.
Report back
The two run results (barrier removed / barrier present), the mechanism you chose and why, the full
build line, and anything you found that has this same shape elsewhere — report it, do not fix it.
Context: fleetd #399, merged as PR #405. The barrier is already known to be load-bearing: removing
the production stamp fails 3 tests. What is not pinned is the ordering guarantee itself.
Correction to this ticket's premise — a reproduction recipe now exists
Posted after the brief went out, so read this if you are working the ticket. Nothing in the goal,
the constraints or the acceptance criteria changes. What changes is one claim I made in the framing,
and one new tool you can use to check your own work.
What I got wrong above
I wrote "a probabilistic test of a race needs ~20-40 runs to say anything." That was derived from
runs taken below the load threshold, where the race essentially never fires. The fleet01 lead then
ran the controlled A/B — one host, one commit, one command, load as the only varied thing:
Phase B was 16 CPU spinners on 8 cores. The threshold is roughly 2x cores. Above it, six runs are
plenty; below it, forty prove nothing. So the honest version of my sentence is: sampling below the
threshold is a null experiment, not a weak one.
I also told you I ran 10 cold JVMs here, all green, and implied that meant the race would not
reproduce on this host. That was wrong for the same reason — this Mac was at load 2.10 on 12 cores.
Do not take those 10 green runs as evidence about anything.
Why the ticket still stands, unchanged
A deterministic test is still strictly better than the load recipe:
So build the test as briefed. The recipe below is not an alternative to it.
What the recipe gives you — a way to check your own test is honest
This is the useful part for you. Once your deterministic test exists, you have a cross-check
available that did not exist when I wrote the brief:
If your test is pinning the real defect, it must go red on exactly the code that goes red under
~2x-cores load. If it goes green where load goes red, your test is pinning something else.
That is a much stronger self-check than "it fails when I remove the barrier", because removing the
barrier is a change you chose. You are not required to run the load comparison — I will run it on
this host (12 cores, ~24 spinners) once you are finished, since loading the machine while you are
building would corrupt your results. But if your test passes when you expect it to fail, or you are
unsure whether you reproduced the right thing, say so in your reply rather than tuning the test
until it goes red. An unsure answer is more useful to me than a confident wrong one.
One detail from the load data that matters for acceptance criterion 4
Phase B counts red runs. The brief tells you the pre-fix failure mode is a false pass, and both
are true — they are different call sites.
MessageServiceTest.java:1920carries a comment saying itsown assertion happens to survive a late stamp (the clock is not advanced further before its sweep),
while
:1955is the one where a late stamp records the advanced time and hides a real eviction.So when you answer criterion 4, say which call site you reproduced and which failure mode you
saw there. "It fails" is not specific enough to tell whether you hit the false-pass site or the red
site.
Merged as part of PR #414. Ruling on the contradiction the worker flagged: they were right, and the fault was in my brief.
The contradiction was mine, not a misreading
My ticket said two things that cannot both hold for one test:
Those describe two different test shapes, and I wrote them as if they described one.
A test that asserts survival cannot go red under this race.
aTaskRunningLongerThanTheTtlStillKeepsItsReportassertsassertNotNull, so a late stamp leaves it green - it passes for the wrong reason. That is the false pass, and it is unfalsifiable by construction.A test that asserts eviction is the only shape the race can turn red, because eviction is the direction a missing stamp blocks. So criterion 1 could only ever be met by the eviction shape.
The worker read the false-pass sentence as describing the sibling test, picked the eviction shape mirroring
aFinishedTicketIsStillPrunedOnceTheTtlPassesSinceItFinished, and said so in the PR body in case they had misread my intent. They had not. My brief held a requirement and its own counter-example side by side.What I checked myself before merging
I did not promote their "3/3" to a fact:
Tests run: 83, Failures: 0Failures: 1, only this test, identical assertionI also checked the claim their comment asserts, in the code:
pruneTerminalTicketshas exactly one caller (sendAsync:1270) andpollnever prunes. So betweendelayArmed.set(true)and the completion hook, no othernowNanosreader runs, and the one-shot delay cannot be consumed by the wrong call site. That claim was load-bearing - if another site ate the delay, the hook would stamp promptly, the barrier would become a no-op, and the test would go green while pinning nothing.The green direction has no timing budget to lose:
awaitCompletionStampedallows 3000ms against a 300ms sleep, and the test thread blocks there while the async thread sleeps. Only the manual red-direction experiment spends the 300ms window, and it needs microseconds of it.Scope limit, recorded so nobody over-reads the test
This pins the test-harness invariant, not a production guarantee. Production deliberately leaves a done-but-unstamped ticket alone (
MessageService.java:1369-1371- "the next sweep collects it"). In the red direction the ticket survives becausecompletedNanoswas stillnullat the sweep, and the test runs no second sweep. So the test proves: a test that advances the clock and sweeps without waiting for the stamp gets a wrong answer. That is exactly what this ticket asked for. It is not evidence that production evicts correctly under a late stamp.For my own future briefs
This is the sixth time a defect came from my own ticket text rather than the code. The specific error here: I described the failure mode and the acceptance criterion from two different test shapes without naming which shape each applied to. State the shape the criterion is about, or the worker has to guess which of my two sentences to obey - and a worker who guesses right still loses a turn asking.
Closing.