Ticket ids restart at task-1 on every daemon restart, so a handover or ticket comment naming an id can silently resolve to a different task #719
Open
opened 2026-10-04 07:59:20 +02:00 by ltms
·
4 comments
No Branch/Tag Specified
main
worker/702-4f5c7f-2
worker/715-5c43fc-1
worker/721-70f9ea-5
worker/718-99362b-2
worker/task-15-af0d10-12
worker/task-16-50a702-13
worker/task-12-4d0479-9
worker/task-13-823ce2-10
worker/705-ticket-owner-af9928-8
worker/703-list-collaborators-9c06c2-7
worker/669-example-truth-0b303d-6
worker/669-collab-deliverability-9ba859-3
worker/669-collab-reload-report-2a21bd-4
worker/669-7e80a6-1
worker/669-unit-d-efbbd7-1
worker/669-1b786a-1
worker/669-1d1d9f-1
worker/692-4afb9d-2
worker/689-02fced-13
worker/693-cf23fa-14
worker/677-fix-lead-collision-f69073-12
worker/638-fix-overmask-dbb1bf-11
worker/675-5b7478-4
worker/669-unit-a-70cc8f-3
worker/677-8cdaaf-5
worker/638-a7b391-1
worker/683-4536d6-2
worker/651-a75bbe-8
worker/680-20607d-7
worker/664-c12e95-3
worker/668-08534d-4
worker/672-0f2469-2
worker/670-7d1022-1
worker/661-ac7c28-2
worker/664-37fb9b-3
worker/663-remove-3arg-read-3f6783-1
worker/659-remove-dead-backcompat-ba5e6f-1
worker/637-revision-60a488-23
worker/656-redact-regression-tests-892903-19
worker/637-context-gauge-threshold-466eb5-16
worker/639-redact-line-numbers-de4ac4-17
worker/641-set-reformat-guard-6f96a4-18
worker/642-herdr-guard-scope-5de0e4-15
worker/650-javadoc-scope-95f3b3-14
worker/612-01e9f7-13
worker/612-a-r4-quarantine-outage-7ab0e8-5
worker/612-a-r9-r11-capacity-coverage-peers-cfcc79-7
worker/612-a-r10-loophealth-ccc872-8
worker/612-a-r12-turnregistrar-9e3bb7-9
worker/612-a-r5-leadconfigdir-9e70cf-6
lead/config-edit-redact-anchor-wording
worker/config-edit-seam-ca8dc1-1
worker/612-r67-630-lifecycle-290b8d-3
worker/629-625-ports-seams-da7d5d-4
worker/612-r12-exhaustion-f37cd7-1
worker/612-r38-amqp-24b083-2
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#719
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Measured, twice, today
1. Observed live. Before today's redeploy the ticket counter had reached
task-19. I restartedthe daemon with
scripts/redeploy-fleetd.sh --yes(new pid 13614, jar101654e0192b). The verynext
fleet_send{wait:false}returned:2. Confirmed in the source. The counter has exactly two references in the whole file:
A plain
AtomicLongfield. Nothing persists it and nothing restores it. So ids restart at 1 andare reused after every restart.
Why it matters, and why the obvious reassurance is not enough
This is not a cross-session read hole. The
tasksmap is in-memory too, so the oldTaskobjects die with the daemon. A reused id refers to a fresh task carrying its own
creatorTerminal,and #705's gate still checks it. I want that stated plainly so nobody re-opens this as a security
issue.
The real failure is a silent wrong answer to the lead that wrote the id down, and the ownership
gate does not catch it precisely because the owner is the same session:
task-15, and writes "polltask-15" into.handover/HANDOVER.mdorinto a ticket comment — which this project's own charter requires, because corrections and
state go on the ticket where they can be read later.
task-15, andownsTicketpasses — same creator — soit receives a different unit's report with no warning.
If a different session created the new
task-15the gate refuses, which is a loud and safefailure. The dangerous case is the common one: one long-lived lead terminal across a restart.
This bit me in a near-miss shape today. The handover I inherited named
task-15andtask-16as"poll these first". Had I redeployed before collecting them — and the handover's own §8 ordering is
what stopped me — those ids would have pointed at nothing, or later at something else.
Suggested shapes, cheapest first
task-<boot>-<n>. A stale id from a previous boot then fails to resolve instead of resolving tothe wrong task. This is the smallest change and it converts a silent wrong answer into a clean
miss.
polldistinguish "unknown ticket" from "that ticket belonged to a previous daemon" — a thirdstate, rather than folding "gone" and "never existed" together.
resolving to nothing with a generic message.
Option 1 or 2. Both are small. Prefer whichever gives the caller a message that names the cause,
because the whole defect is that the current behaviour explains nothing.
Not measured
I did not drive the collision end to end — that needs the counter walked back up past a recorded id
after a restart, with the same creator terminal. The argument above is read from the two greps and
the one observed reset. The reset itself is measured; the collision is reasoned from it.
Found while writing the lead handover, when I noticed the ticket id I was about to write down had
just been reset by my own redeploy.
Withdrawing my own option 2 — it does not work
I filed this an hour ago and listed three shapes. Option 2 cannot detect the failure it was
supposed to detect. Working it through:
Option 2 was "record the boot id alongside each task and have
polldistinguish a stale ticket froman unknown one". But the caller holds nothing except the ticket string. In the failure case the
lead polls
task-15during boot B, and a newtask-15exists in boot B. That task's recorded bootid is B, which matches the current boot, so the check passes and the wrong report is still returned
silently. Recording the boot on the task only helps if the thing being compared came from the
caller — and it did not.
The general rule, which is the useful part: staleness must be carried by the identifier the caller
holds. Anything recorded only on the server side is invisible to a caller presenting an old id.
That leaves two shapes, and both work for the opposite reason:
task-<boot>-<n>. A boot-A id presented in boot B cannotmatch any boot-B id, so it resolves to nothing. The staleness is in the string the caller holds.
task-15is then never reissued, so the old id resolves to nothing.Staleness is carried by the id too, just by never colliding instead of by being tagged.
Shape 1 is a visible format change. The ticket id appears in the
fleet_sendreceipt text, inCLAUDE.md's intent table, in every handover, and in ticket comments. So it needs awiki/11-Features.mdentry and a scan for anything that parses or asserts the
task-Nform. Shape 2 keeps the format andneeds durable storage this daemon does not currently have for tickets.
I am not delegating this yet — the choice between a visible format change and new persistence is a
design call, and it touches
MessageServicewhere #718 is about to add a scrape test. Keeping itmine for now.
Two things in the original body remain correct and measured: the reset itself (observed,
task-19→task-1across today's redeploy) and the two-lineticketSeqgrep. Only the remedy list was wrong.Decision: option 1, the per-boot nonce. It is far cheaper than I claimed.
In my previous comment I said shape 1 is "a visible format change" needing a scan for anything
that parses or asserts the
task-Nform, and implied that was a real cost. I then measured it, andthe cost is close to zero.
Nothing in production parses a ticket id. There is exactly one reference to the literal:
That is the mint itself. No parser, no splitter, no consumer of its shape.
No test asserts the minted format. Three test files mention the literal, 52 occurrences in all,
and I checked what they are:
ReplyPushLoopTest— 50 occurrences, all fabricated input strings ("task-1","task-2")handed to the loop as arbitrary ticket ids. The mint is not involved, so the format can change
without touching any of them.
MessageServiceTest:888—assertNull(messages.poll("task-999999"), "a ticket that was never minted is unknown"). A deliberately never-minted id; it stays never-minted under any format.FleetMcpTest:293—FleetMcp.poll(messages, "task-999", null). Same shape.So zero assertions on the minted format. The change is one production line plus new tests for
the new behaviour.
What ships
task-<boot>-<n>, where<boot>is generated once perMessageServiceinstance. Keep it short; it is read aloud in handovers and typed by hand.MessageServiceinstances mint disjoint id spaces, so an id from one neverresolves in the other. That is the property this ticket exists for, and it is testable without a
restart: construct two instances and cross-poll.
rather than resolving. Include the control: the same id resolves correctly in its own
instance, so a test that "fails to resolve" cannot pass because the plumbing is broken.
wiki/11-Features.mdentry. The ticket id is operator-visible — it appears in thefleet_sendreceipt text — and the project rule is that anything an operator can observe earnsone entry: what it does, the knob, why it exists, the gotcha. The why here is the silent
wrong answer in the body above.
What does not ship
Persistence. Shape 2 (persist the counter) also works, but it needs durable storage for tickets that
this daemon does not have, and it buys nothing extra: a stale id resolves to nothing either way. The
nonce gets the same safety for one line.
One thing to decide while implementing, not before
Whether
fleet_pollshould name the cause — "that ticket belonged to a previous daemon" — or justreport unknown. Naming it is better if the boot can be recognised from the id, and the whole defect
is that today's behaviour explains nothing. Do not let it become a second code path that can
disagree with the first; one lookup, one answer.
Ready to delegate. Holding only until #718 lands, because that unit is in flight against
MessageServiceand I want the merges serialised.Lead. This stopped being theoretical — it happened to me this session, and I can show the numbers.
What I measured
The daemon restarted under me when I redeployed after merging #716 and #717:
The handover file I inherited was written at 07:15, before that restart. It names the tickets the previous lead left in flight:
After the restart the counter began again at 1. The three units I have running right now are task-6, task-7, task-8.
And the counter is confirmed to have no persistence — two references in the whole codebase:
The part that is worse than a lost ticket
A stale id does not stay a 404. It becomes a wrong answer.
Right now
task-15returns 404, because the counter has only reached 8. The documented REST behaviour for an unknown ticket is a 404 (FleetApp.taskStatus: "404 for an unknown/expired ticket"). That is the safe failure, and it is the one everybody assumes.But the counter is climbing. Seven more delegations in this session and it reaches 15. At that moment
fleet_poll{ticket:"task-15"}stops returning 404 and starts returning a real task that is not the one the handover meant — my architect's work on #702, say, handed to whoever asked for a unit from the previous session. No error, no warning, and the reply text looks plausible because it is a genuine reply to a genuine delegation.So the failure mode is not "the ticket is gone". It is "the id silently rebinds to someone else's work", and it becomes reachable simply by doing more work after a restart. That is a sharper edge than the ticket body describes, and it is the reason this should not sit.
This confirms the design rule I withdrew option 2 for
Earlier on this ticket I proposed recording a boot id server-side and rejecting a ticket from an older boot. I withdrew it because it cannot work: the caller holds only the string
task-15, so there is nothing in what it presents that a server-side boot id can be compared against. The new same-numbered task simply carries the current boot id and matches.The measurement above is the evidence for that rule: staleness has to be carried by the identifier the caller holds. A nonce in the id itself —
task-7f3a21-15or similar — makes a pre-restart id unmatchable by construction, because the nonce is minted once per boot and the old string carries the old one. Nothing server-side can recover that information after the fact.Scheduling
Still queued behind #721, which is in flight and also edits
MessageService. Two units editing one file in parallel is how a clean auto-merge turns into a broken build, so these serialise. #721 touchespendingAsk(around:1763); this touches the mint at:1303. Different methods, same file — I am not relying on that to make them mergeable.One thing I have not measured
I have not checked whether anything other than a ticket string has this shape — an id a caller holds across a restart and presents back.
turnIdis the obvious candidate, since #715 establishes it issession + "#" + askSeq.incrementAndGet()andaskSeqis also a per-RendezvousAtomicLong. A session id changes per pane, so a staleturnIdprobably cannot collide the same way, but I have not read it carefully and I am not claiming it.Lead. The collision is now measured end to end. It happened to this ticket's own comment.
The body says "I did not drive the collision end to end" and the last comment says the same. That gap
is closed, and the evidence is this ticket.
The two readings
The comment above, written at 08:32:02 +02:00, records:
That was the boot of pid 13614, started 07:53:13.
The daemon has restarted since. The current listener:
Started 10:27:52, which is after that comment. So the counter reset again.
At 17:54 today I delegated #727 with
fleet_send{wait:false}. The receipt:The id does not 404 — it answers
So the string
task-6resolves right now, and it resolves to my #727 unit. It does not resolve tothe unit the comment above meant. Someone reading this ticket's own history, following its own
instruction to poll
task-6, gets a different worker's work with no warning.That is the exact failure the body predicted, and it took no special setup. It needed only a restart
and six more delegations.
What this adds, and what it does not
Added: the collision is observed, not reasoned. The previous comment argued it from one reset plus two
greps and said the collision itself was inferred. It is now a measurement.
Not added, and I want to be exact: the two
task-6tasks were created by different lead sessions,so I have not shown
ownsTicketpassing for a reused id within one terminal. The body's sharpest case— one long-lived lead terminal across a restart, where the ownership gate cannot help because the
creator matches — is still reasoned rather than driven. What is now measured is the id reuse and the
resolution to a different unit.
Also worth stating plainly: the old tasks died with their daemon, so nothing leaked between sessions.
The harm is a wrong answer to whoever holds the written-down id, exactly as the body framed it.
This is a reason not to leave it queued
The ticket is held behind #718 and #721 because both edit
MessageService, and that serialisation isstill right. But the defect is reachable today, by the normal working pattern this project's charter
requires — put state on the ticket, redeploy after a merge. Those two habits combine into this.
Design is unchanged: option 1, the per-boot nonce, as decided above. Nothing here argues for a
different remedy; it only removes the last "not measured" line from the case for it.
Re-measure with
ps -o lstart= -p $(lsof -nP -iTCP:8765 -sTCP:LISTEN -t)against the timestamp of anycomment naming a
task-N. A daemon start later than the comment means that comment's ids are unsafeto poll. Once ids carry a boot nonce this check stops being needed, and this comment can go.