Ticket ids restart at task-1 on every daemon restart, so a handover or ticket comment naming an id can silently resolve to a different task #719

Open
opened 2026-10-04 07:59:20 +02:00 by ltms · 4 comments
Owner

Measured, twice, today

1. Observed live. Before today's redeploy the ticket counter had reached task-19. I restarted
the daemon with scripts/redeploy-fleetd.sh --yes (new pid 13614, jar 101654e0192b). The very
next fleet_send{wait:false} returned:

accepted — task delegated. Poll fleet_poll with ticket=task-1

2. Confirmed in the source. The counter has exactly two references in the whole file:

$ grep -n 'ticketSeq' fleetd/src/main/java/dev/ltms/fleet/msg/MessageService.java
351:    private final AtomicLong ticketSeq = new AtomicLong();
1303:        String ticket = "task-" + ticketSeq.incrementAndGet();

A plain AtomicLong field. Nothing persists it and nothing restores it. So ids restart at 1 and
are reused after every restart.

Why it matters, and why the obvious reassurance is not enough

This is not a cross-session read hole. The tasks map is in-memory too, so the old Task
objects die with the daemon. A reused id refers to a fresh task carrying its own creatorTerminal,
and #705's gate still checks it. I want that stated plainly so nobody re-opens this as a security
issue.

The real failure is a silent wrong answer to the lead that wrote the id down, and the ownership
gate does not catch it precisely because the owner is the same session:

  1. A lead delegates, gets task-15, and writes "poll task-15" into .handover/HANDOVER.md or
    into a ticket comment — which this project's own charter requires, because corrections and
    state go on the ticket where they can be read later.
  2. The daemon restarts. A redeploy is routine here and the charter tells the lead to do it.
  3. The counter climbs back past 15 as new work is delegated by the same lead terminal.
  4. The lead follows its own handover, polls task-15, and ownsTicket passes — same creator — so
    it receives a different unit's report with no warning.

If a different session created the new task-15 the gate refuses, which is a loud and safe
failure. The dangerous case is the common one: one long-lived lead terminal across a restart.

This bit me in a near-miss shape today. The handover I inherited named task-15 and task-16 as
"poll these first". Had I redeployed before collecting them — and the handover's own §8 ordering is
what stopped me — those ids would have pointed at nothing, or later at something else.

Suggested shapes, cheapest first

  1. Make the id unforgeable across boots. Prefix the counter with a per-boot nonce, e.g.
    task-<boot>-<n>. A stale id from a previous boot then fails to resolve instead of resolving to
    the wrong task. This is the smallest change and it converts a silent wrong answer into a clean
    miss.
  2. Keep the numbering but refuse a stale id. Record the boot id alongside each task and have
    poll distinguish "unknown ticket" from "that ticket belonged to a previous daemon" — a third
    state, rather than folding "gone" and "never existed" together.
  3. Persist the counter. More work, and it buys less: it stops reuse but still leaves a stale id
    resolving to nothing with a generic message.

Option 1 or 2. Both are small. Prefer whichever gives the caller a message that names the cause,
because the whole defect is that the current behaviour explains nothing.

Not measured

I did not drive the collision end to end — that needs the counter walked back up past a recorded id
after a restart, with the same creator terminal. The argument above is read from the two greps and
the one observed reset. The reset itself is measured; the collision is reasoned from it.

Found while writing the lead handover, when I noticed the ticket id I was about to write down had
just been reset by my own redeploy.

## Measured, twice, today **1. Observed live.** Before today's redeploy the ticket counter had reached `task-19`. I restarted the daemon with `scripts/redeploy-fleetd.sh --yes` (new pid 13614, jar `101654e0192b`). The very next `fleet_send{wait:false}` returned: ``` accepted — task delegated. Poll fleet_poll with ticket=task-1 ``` **2. Confirmed in the source.** The counter has exactly two references in the whole file: ``` $ grep -n 'ticketSeq' fleetd/src/main/java/dev/ltms/fleet/msg/MessageService.java 351: private final AtomicLong ticketSeq = new AtomicLong(); 1303: String ticket = "task-" + ticketSeq.incrementAndGet(); ``` A plain `AtomicLong` field. Nothing persists it and nothing restores it. So ids restart at 1 and are **reused** after every restart. ## Why it matters, and why the obvious reassurance is not enough This is **not** a cross-session read hole. The `tasks` map is in-memory too, so the old `Task` objects die with the daemon. A reused id refers to a fresh task carrying its own `creatorTerminal`, and #705's gate still checks it. I want that stated plainly so nobody re-opens this as a security issue. The real failure is a **silent wrong answer to the lead that wrote the id down**, and the ownership gate does not catch it precisely because the owner is the same session: 1. A lead delegates, gets `task-15`, and writes "poll `task-15`" into `.handover/HANDOVER.md` or into a ticket comment — which this project's own charter *requires*, because corrections and state go on the ticket where they can be read later. 2. The daemon restarts. A redeploy is routine here and the charter tells the lead to do it. 3. The counter climbs back past 15 as new work is delegated **by the same lead terminal**. 4. The lead follows its own handover, polls `task-15`, and `ownsTicket` passes — same creator — so it receives a **different unit's report** with no warning. If a *different* session created the new `task-15` the gate refuses, which is a loud and safe failure. The dangerous case is the common one: one long-lived lead terminal across a restart. This bit me in a near-miss shape today. The handover I inherited named `task-15` and `task-16` as "poll these first". Had I redeployed before collecting them — and the handover's own §8 ordering is what stopped me — those ids would have pointed at nothing, or later at something else. ## Suggested shapes, cheapest first 1. **Make the id unforgeable across boots.** Prefix the counter with a per-boot nonce, e.g. `task-<boot>-<n>`. A stale id from a previous boot then fails to resolve instead of resolving to the wrong task. This is the smallest change and it converts a silent wrong answer into a clean miss. 2. **Keep the numbering but refuse a stale id.** Record the boot id alongside each task and have `poll` distinguish "unknown ticket" from "that ticket belonged to a previous daemon" — a third state, rather than folding "gone" and "never existed" together. 3. **Persist the counter.** More work, and it buys less: it stops reuse but still leaves a stale id resolving to nothing with a generic message. Option 1 or 2. Both are small. Prefer whichever gives the caller a message that names the cause, because the whole defect is that the current behaviour explains nothing. ## Not measured I did not drive the collision end to end — that needs the counter walked back up past a recorded id after a restart, with the same creator terminal. The argument above is read from the two greps and the one observed reset. The reset itself is measured; the collision is reasoned from it. Found while writing the lead handover, when I noticed the ticket id I was about to write down had just been reset by my own redeploy.
Author
Owner

Withdrawing my own option 2 — it does not work

I filed this an hour ago and listed three shapes. Option 2 cannot detect the failure it was
supposed to detect.
Working it through:

Option 2 was "record the boot id alongside each task and have poll distinguish a stale ticket from
an unknown one". But the caller holds nothing except the ticket string. In the failure case the
lead polls task-15 during boot B, and a new task-15 exists in boot B. That task's recorded boot
id is B, which matches the current boot, so the check passes and the wrong report is still returned
silently. Recording the boot on the task only helps if the thing being compared came from the
caller — and it did not.

The general rule, which is the useful part: staleness must be carried by the identifier the caller
holds.
Anything recorded only on the server side is invisible to a caller presenting an old id.

That leaves two shapes, and both work for the opposite reason:

  1. Put a per-boot nonce in the id — task-<boot>-<n>. A boot-A id presented in boot B cannot
    match any boot-B id, so it resolves to nothing. The staleness is in the string the caller holds.
  2. Persist the counter — task-15 is then never reissued, so the old id resolves to nothing.
    Staleness is carried by the id too, just by never colliding instead of by being tagged.

Shape 1 is a visible format change. The ticket id appears in the fleet_send receipt text, in
CLAUDE.md's intent table, in every handover, and in ticket comments. So it needs a wiki/11-Features.md
entry and a scan for anything that parses or asserts the task-N form. Shape 2 keeps the format and
needs durable storage this daemon does not currently have for tickets.

I am not delegating this yet — the choice between a visible format change and new persistence is a
design call, and it touches MessageService where #718 is about to add a scrape test. Keeping it
mine for now.

Two things in the original body remain correct and measured: the reset itself (observed, task-19 →
task-1 across today's redeploy) and the two-line ticketSeq grep. Only the remedy list was wrong.

## Withdrawing my own option 2 — it does not work I filed this an hour ago and listed three shapes. **Option 2 cannot detect the failure it was supposed to detect.** Working it through: Option 2 was "record the boot id alongside each task and have `poll` distinguish a stale ticket from an unknown one". But the caller holds nothing except the ticket **string**. In the failure case the lead polls `task-15` during boot B, and a *new* `task-15` exists in boot B. That task's recorded boot id is B, which matches the current boot, so the check passes and the wrong report is still returned silently. Recording the boot on the **task** only helps if the thing being compared came from the **caller** — and it did not. The general rule, which is the useful part: **staleness must be carried by the identifier the caller holds.** Anything recorded only on the server side is invisible to a caller presenting an old id. That leaves two shapes, and both work for the opposite reason: 1. **Put a per-boot nonce in the id** — `task-<boot>-<n>`. A boot-A id presented in boot B cannot match any boot-B id, so it resolves to nothing. The staleness is in the string the caller holds. 2. **Persist the counter** — `task-15` is then never reissued, so the old id resolves to nothing. Staleness is carried by the id too, just by never colliding instead of by being tagged. Shape 1 is a **visible format change**. The ticket id appears in the `fleet_send` receipt text, in `CLAUDE.md`'s intent table, in every handover, and in ticket comments. So it needs a `wiki/11-Features.md` entry and a scan for anything that parses or asserts the `task-N` form. Shape 2 keeps the format and needs durable storage this daemon does not currently have for tickets. I am not delegating this yet — the choice between a visible format change and new persistence is a design call, and it touches `MessageService` where #718 is about to add a scrape test. Keeping it mine for now. Two things in the original body remain correct and measured: the reset itself (observed, `task-19` → `task-1` across today's redeploy) and the two-line `ticketSeq` grep. Only the remedy list was wrong.
Author
Owner

Decision: option 1, the per-boot nonce. It is far cheaper than I claimed.

In my previous comment I said shape 1 is "a visible format change" needing a scan for anything
that parses or asserts the task-N form, and implied that was a real cost. I then measured it, and
the cost is close to zero.

Nothing in production parses a ticket id. There is exactly one reference to the literal:

$ grep -rn '"task-' fleetd/src/main/java --include='*.java'
MessageService.java:1303:   String ticket = "task-" + ticketSeq.incrementAndGet();

That is the mint itself. No parser, no splitter, no consumer of its shape.

No test asserts the minted format. Three test files mention the literal, 52 occurrences in all,
and I checked what they are:

  • ReplyPushLoopTest — 50 occurrences, all fabricated input strings ("task-1", "task-2")
    handed to the loop as arbitrary ticket ids. The mint is not involved, so the format can change
    without touching any of them.
  • MessageServiceTest:888 — assertNull(messages.poll("task-999999"), "a ticket that was never minted is unknown"). A deliberately never-minted id; it stays never-minted under any format.
  • FleetMcpTest:293 — FleetMcp.poll(messages, "task-999", null). Same shape.

So zero assertions on the minted format. The change is one production line plus new tests for
the new behaviour.

What ships

  1. Mint the id with a per-boot nonce — task-<boot>-<n>, where <boot> is generated once per
    MessageService instance. Keep it short; it is read aloud in handovers and typed by hand.
  2. A test that two MessageService instances mint disjoint id spaces, so an id from one never
    resolves in the other. That is the property this ticket exists for, and it is testable without a
    restart: construct two instances and cross-poll.
  3. A test that an unknown-but-well-formed id from another boot returns the unknown-ticket outcome
    rather than resolving. Include the control: the same id resolves correctly in its own
    instance, so a test that "fails to resolve" cannot pass because the plumbing is broken.
  4. A wiki/11-Features.md entry. The ticket id is operator-visible — it appears in the
    fleet_send receipt text — and the project rule is that anything an operator can observe earns
    one entry: what it does, the knob, why it exists, the gotcha. The why here is the silent
    wrong answer in the body above.

What does not ship

Persistence. Shape 2 (persist the counter) also works, but it needs durable storage for tickets that
this daemon does not have, and it buys nothing extra: a stale id resolves to nothing either way. The
nonce gets the same safety for one line.

One thing to decide while implementing, not before

Whether fleet_poll should name the cause — "that ticket belonged to a previous daemon" — or just
report unknown. Naming it is better if the boot can be recognised from the id, and the whole defect
is that today's behaviour explains nothing. Do not let it become a second code path that can
disagree with the first; one lookup, one answer.

Ready to delegate. Holding only until #718 lands, because that unit is in flight against
MessageService and I want the merges serialised.

## Decision: option 1, the per-boot nonce. It is far cheaper than I claimed. In my previous comment I said shape 1 is "a **visible format change**" needing a scan for anything that parses or asserts the `task-N` form, and implied that was a real cost. I then measured it, and the cost is close to zero. **Nothing in production parses a ticket id.** There is exactly one reference to the literal: ``` $ grep -rn '"task-' fleetd/src/main/java --include='*.java' MessageService.java:1303: String ticket = "task-" + ticketSeq.incrementAndGet(); ``` That is the mint itself. No parser, no splitter, no consumer of its shape. **No test asserts the minted format.** Three test files mention the literal, 52 occurrences in all, and I checked what they are: - `ReplyPushLoopTest` — 50 occurrences, all **fabricated input strings** (`"task-1"`, `"task-2"`) handed to the loop as arbitrary ticket ids. The mint is not involved, so the format can change without touching any of them. - `MessageServiceTest:888` — `assertNull(messages.poll("task-999999"), "a ticket that was never minted is unknown")`. A deliberately never-minted id; it stays never-minted under any format. - `FleetMcpTest:293` — `FleetMcp.poll(messages, "task-999", null)`. Same shape. So **zero assertions on the minted format**. The change is one production line plus new tests for the new behaviour. ### What ships 1. Mint the id with a per-boot nonce — `task-<boot>-<n>`, where `<boot>` is generated once per `MessageService` instance. Keep it short; it is read aloud in handovers and typed by hand. 2. A test that two `MessageService` instances mint **disjoint** id spaces, so an id from one never resolves in the other. That is the property this ticket exists for, and it is testable without a restart: construct two instances and cross-poll. 3. A test that an unknown-but-well-formed id from another boot returns the unknown-ticket outcome rather than resolving. Include the **control**: the same id resolves correctly in its own instance, so a test that "fails to resolve" cannot pass because the plumbing is broken. 4. A `wiki/11-Features.md` entry. The ticket id is operator-visible — it appears in the `fleet_send` receipt text — and the project rule is that anything an operator can observe earns one entry: what it does, the knob, **why it exists**, the gotcha. The *why* here is the silent wrong answer in the body above. ### What does not ship Persistence. Shape 2 (persist the counter) also works, but it needs durable storage for tickets that this daemon does not have, and it buys nothing extra: a stale id resolves to nothing either way. The nonce gets the same safety for one line. ### One thing to decide while implementing, not before Whether `fleet_poll` should **name** the cause — "that ticket belonged to a previous daemon" — or just report unknown. Naming it is better if the boot can be recognised from the id, and the whole defect is that today's behaviour explains nothing. Do not let it become a second code path that can disagree with the first; one lookup, one answer. Ready to delegate. Holding only until #718 lands, because that unit is in flight against `MessageService` and I want the merges serialised.
Author
Owner

Lead. This stopped being theoretical — it happened to me this session, and I can show the numbers.

What I measured

The daemon restarted under me when I redeployed after merging #716 and #717:

$ ps -o lstart=,etime= -p 13614
Sun Oct  4 07:53:13 2026     38:13

The handover file I inherited was written at 07:15, before that restart. It names the tickets the previous lead left in flight:

$ grep -o 'task-[0-9]*' .handover/HANDOVER.md | sort -u
task-11 task-12 task-13 task-15 task-16

After the restart the counter began again at 1. The three units I have running right now are task-6, task-7, task-8.

And the counter is confirmed to have no persistence — two references in the whole codebase:

$ grep -rn 'ticketSeq' fleetd/src/main/java fleetd/src/test/java
MessageService.java:351:   private final AtomicLong ticketSeq = new AtomicLong();
MessageService.java:1303:  String ticket = "task-" + ticketSeq.incrementAndGet();

The part that is worse than a lost ticket

A stale id does not stay a 404. It becomes a wrong answer.

Right now task-15 returns 404, because the counter has only reached 8. The documented REST behaviour for an unknown ticket is a 404 (FleetApp.taskStatus: "404 for an unknown/expired ticket"). That is the safe failure, and it is the one everybody assumes.

But the counter is climbing. Seven more delegations in this session and it reaches 15. At that moment fleet_poll{ticket:"task-15"} stops returning 404 and starts returning a real task that is not the one the handover meant — my architect's work on #702, say, handed to whoever asked for a unit from the previous session. No error, no warning, and the reply text looks plausible because it is a genuine reply to a genuine delegation.

So the failure mode is not "the ticket is gone". It is "the id silently rebinds to someone else's work", and it becomes reachable simply by doing more work after a restart. That is a sharper edge than the ticket body describes, and it is the reason this should not sit.

This confirms the design rule I withdrew option 2 for

Earlier on this ticket I proposed recording a boot id server-side and rejecting a ticket from an older boot. I withdrew it because it cannot work: the caller holds only the string task-15, so there is nothing in what it presents that a server-side boot id can be compared against. The new same-numbered task simply carries the current boot id and matches.

The measurement above is the evidence for that rule: staleness has to be carried by the identifier the caller holds. A nonce in the id itself — task-7f3a21-15 or similar — makes a pre-restart id unmatchable by construction, because the nonce is minted once per boot and the old string carries the old one. Nothing server-side can recover that information after the fact.

Scheduling

Still queued behind #721, which is in flight and also edits MessageService. Two units editing one file in parallel is how a clean auto-merge turns into a broken build, so these serialise. #721 touches pendingAsk (around :1763); this touches the mint at :1303. Different methods, same file — I am not relying on that to make them mergeable.

One thing I have not measured

I have not checked whether anything other than a ticket string has this shape — an id a caller holds across a restart and presents back. turnId is the obvious candidate, since #715 establishes it is session + "#" + askSeq.incrementAndGet() and askSeq is also a per-Rendezvous AtomicLong. A session id changes per pane, so a stale turnId probably cannot collide the same way, but I have not read it carefully and I am not claiming it.

Lead. **This stopped being theoretical — it happened to me this session, and I can show the numbers.** ## What I measured The daemon restarted under me when I redeployed after merging #716 and #717: ``` $ ps -o lstart=,etime= -p 13614 Sun Oct 4 07:53:13 2026 38:13 ``` The handover file I inherited was written at **07:15**, before that restart. It names the tickets the previous lead left in flight: ``` $ grep -o 'task-[0-9]*' .handover/HANDOVER.md | sort -u task-11 task-12 task-13 task-15 task-16 ``` After the restart the counter began again at 1. The three units I have running right now are **task-6, task-7, task-8**. And the counter is confirmed to have no persistence — two references in the whole codebase: ``` $ grep -rn 'ticketSeq' fleetd/src/main/java fleetd/src/test/java MessageService.java:351: private final AtomicLong ticketSeq = new AtomicLong(); MessageService.java:1303: String ticket = "task-" + ticketSeq.incrementAndGet(); ``` ## The part that is worse than a lost ticket A stale id does not stay a 404. It becomes a **wrong answer**. Right now `task-15` returns 404, because the counter has only reached 8. The documented REST behaviour for an unknown ticket is a 404 (`FleetApp.taskStatus`: *"404 for an unknown/expired ticket"*). That is the safe failure, and it is the one everybody assumes. But the counter is climbing. Seven more delegations in this session and it reaches 15. At that moment `fleet_poll{ticket:"task-15"}` stops returning 404 and starts returning **a real task that is not the one the handover meant** — my architect's work on #702, say, handed to whoever asked for a unit from the previous session. No error, no warning, and the reply text looks plausible because it is a genuine reply to a genuine delegation. So the failure mode is not "the ticket is gone". It is "the id silently rebinds to someone else's work", and it becomes reachable simply by doing more work after a restart. That is a sharper edge than the ticket body describes, and it is the reason this should not sit. ## This confirms the design rule I withdrew option 2 for Earlier on this ticket I proposed recording a boot id server-side and rejecting a ticket from an older boot. I withdrew it because it cannot work: the caller holds only the string `task-15`, so there is nothing in what it presents that a server-side boot id can be compared against. The new same-numbered task simply carries the current boot id and matches. The measurement above is the evidence for that rule: **staleness has to be carried by the identifier the caller holds.** A nonce in the id itself — `task-7f3a21-15` or similar — makes a pre-restart id unmatchable by construction, because the nonce is minted once per boot and the old string carries the old one. Nothing server-side can recover that information after the fact. ## Scheduling Still queued behind #721, which is in flight and also edits `MessageService`. Two units editing one file in parallel is how a clean auto-merge turns into a broken build, so these serialise. #721 touches `pendingAsk` (around `:1763`); this touches the mint at `:1303`. Different methods, same file — I am not relying on that to make them mergeable. ## One thing I have not measured I have not checked whether anything **other** than a ticket string has this shape — an id a caller holds across a restart and presents back. `turnId` is the obvious candidate, since #715 establishes it is `session + "#" + askSeq.incrementAndGet()` and `askSeq` is also a per-`Rendezvous` `AtomicLong`. A session id changes per pane, so a stale `turnId` probably cannot collide the same way, but I have not read it carefully and I am not claiming it.
Author
Owner

Lead. The collision is now measured end to end. It happened to this ticket's own comment.

The body says "I did not drive the collision end to end" and the last comment says the same. That gap
is closed, and the evidence is this ticket.

The two readings

The comment above, written at 08:32:02 +02:00, records:

The three units I have running right now are task-6, task-7, task-8.

That was the boot of pid 13614, started 07:53:13.

The daemon has restarted since. The current listener:

$ lsof -nP -iTCP:8765 -sTCP:LISTEN -t
28368
$ ps -o pid=,lstart=,etime= -p 28368
28368 Sun Oct  4 10:27:52 2026     07:28:08

Started 10:27:52, which is after that comment. So the counter reset again.

At 17:54 today I delegated #727 with fleet_send{wait:false}. The receipt:

accepted — task delegated. Poll fleet_poll with ticket=task-6

The id does not 404 — it answers

$ fleet_poll{ticket:"task-6"}
[pending — worker working]

So the string task-6 resolves right now, and it resolves to my #727 unit. It does not resolve to
the unit the comment above meant. Someone reading this ticket's own history, following its own
instruction to poll task-6, gets a different worker's work with no warning.

That is the exact failure the body predicted, and it took no special setup. It needed only a restart
and six more delegations.

What this adds, and what it does not

Added: the collision is observed, not reasoned. The previous comment argued it from one reset plus two
greps and said the collision itself was inferred. It is now a measurement.

Not added, and I want to be exact: the two task-6 tasks were created by different lead sessions,
so I have not shown ownsTicket passing for a reused id within one terminal. The body's sharpest case
— one long-lived lead terminal across a restart, where the ownership gate cannot help because the
creator matches — is still reasoned rather than driven. What is now measured is the id reuse and the
resolution to a different unit.

Also worth stating plainly: the old tasks died with their daemon, so nothing leaked between sessions.
The harm is a wrong answer to whoever holds the written-down id, exactly as the body framed it.

This is a reason not to leave it queued

The ticket is held behind #718 and #721 because both edit MessageService, and that serialisation is
still right. But the defect is reachable today, by the normal working pattern this project's charter
requires — put state on the ticket, redeploy after a merge. Those two habits combine into this.

Design is unchanged: option 1, the per-boot nonce, as decided above. Nothing here argues for a
different remedy; it only removes the last "not measured" line from the case for it.

Re-measure with ps -o lstart= -p $(lsof -nP -iTCP:8765 -sTCP:LISTEN -t) against the timestamp of any
comment naming a task-N. A daemon start later than the comment means that comment's ids are unsafe
to poll. Once ids carry a boot nonce this check stops being needed, and this comment can go.

Lead. **The collision is now measured end to end. It happened to this ticket's own comment.** The body says "I did not drive the collision end to end" and the last comment says the same. That gap is closed, and the evidence is this ticket. ## The two readings The comment above, written at **08:32:02 +02:00**, records: > The three units I have running right now are **task-6, task-7, task-8**. That was the boot of pid 13614, started 07:53:13. The daemon has restarted since. The current listener: ``` $ lsof -nP -iTCP:8765 -sTCP:LISTEN -t 28368 $ ps -o pid=,lstart=,etime= -p 28368 28368 Sun Oct 4 10:27:52 2026 07:28:08 ``` Started **10:27:52**, which is after that comment. So the counter reset again. At 17:54 today I delegated #727 with `fleet_send{wait:false}`. The receipt: ``` accepted — task delegated. Poll fleet_poll with ticket=task-6 ``` ## The id does not 404 — it answers ``` $ fleet_poll{ticket:"task-6"} [pending — worker working] ``` So the string `task-6` **resolves right now**, and it resolves to my #727 unit. It does not resolve to the unit the comment above meant. Someone reading this ticket's own history, following its own instruction to poll `task-6`, gets a different worker's work with no warning. That is the exact failure the body predicted, and it took no special setup. It needed only a restart and six more delegations. ## What this adds, and what it does not Added: the collision is observed, not reasoned. The previous comment argued it from one reset plus two greps and said the collision itself was inferred. It is now a measurement. Not added, and I want to be exact: the two `task-6` tasks were created by **different lead sessions**, so I have not shown `ownsTicket` passing for a reused id within one terminal. The body's sharpest case — one long-lived lead terminal across a restart, where the ownership gate cannot help because the creator matches — is still reasoned rather than driven. What is now measured is the id reuse and the resolution to a different unit. Also worth stating plainly: the old tasks died with their daemon, so nothing leaked between sessions. The harm is a wrong answer to whoever holds the written-down id, exactly as the body framed it. ## This is a reason not to leave it queued The ticket is held behind #718 and #721 because both edit `MessageService`, and that serialisation is still right. But the defect is reachable today, by the normal working pattern this project's charter requires — put state on the ticket, redeploy after a merge. Those two habits combine into this. Design is unchanged: option 1, the per-boot nonce, as decided above. Nothing here argues for a different remedy; it only removes the last "not measured" line from the case for it. Re-measure with `ps -o lstart= -p $(lsof -nP -iTCP:8765 -sTCP:LISTEN -t)` against the timestamp of any comment naming a `task-N`. A daemon start later than the comment means that comment's ids are unsafe to poll. Once ids carry a boot nonce this check stops being needed, and this comment can go.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#719