turnId has the same per-boot reuse shape as the ticket id #719 fixed, but its reachability depends on a premise I have not checked #729

Closed
opened 2026-10-04 18:13:16 +02:00 by ltms · 2 comments
Owner

The shape

#719 fixed ticket ids restarting at task-1 on every daemon restart. A worker flagged turnId as
having the same shape while implementing it. I read the code and it does, in part.

Measured 2026-10-04 on main:

$ grep -n 'askSeq\|turnId =' fleetd/src/main/java/dev/ltms/fleet/msg/Rendezvous.java
103:    private final AtomicLong askSeq = new AtomicLong();
189:            String turnId = openAsksBySession.computeIfAbsent(session, _ -> {
190:                String newTurnId = session + "#" + askSeq.incrementAndGet();

So turnId is <session>#<n>, and askSeq is a per-Rendezvous AtomicLong with nothing persisting
it. The counter half resets on every daemon restart, exactly like ticketSeq did.

Why this is weaker than #719, not equal to it

A ticket id was entirely minted by fleetd, so a reset made the whole id reusable. A turnId is two
parts, and only the second part resets:

  • <n> — resets to 0 on every restart. Reusable.
  • <session> — a herdr terminal_id. herdr is a separate process and a fleetd restart does not
    restart it
    , so herdr keeps issuing fresh terminal ids and existing ones keep their value.

A collision therefore needs the same session string to appear again after the restart, not just the
counter to climb back. That is a real extra condition, and it is why this is filed separately rather
than folded into #719.

The premise I have not checked

For a collision, a member terminal must survive a fleetd restart and be re-adopted into the new
daemon's roster
, then ask again. I have not read how the roster is rebuilt after a restart, and I have
not driven it. If a surviving terminal is never re-adopted — if every post-restart member gets a brand
new terminal id — then the <session> half never repeats and this is not reachable at all.

Someone must establish that before writing code. A fix for an unreachable path is cost with no
benefit, and I am not going to guess. This is the one question the ticket turns on.

If it is reachable, who gets hurt

The same party as in #719: whoever wrote the id down. A lead answers a member's question with
fleet_send{turnId, content}. A lead that recorded a turnId — in .handover/HANDOVER.md or a ticket
comment, which this project's charter tells it to do — could, after a restart, send an answer to a
different open ask.

The rendezvous state is in memory, so immediately after a restart there are no open asks and a stale
turnId fails safely. The dangerous window opens later, once askSeq has climbed back to the same
number and the same session has asked again. That is the #719 pattern precisely: a stale id that starts
as a clean miss and becomes a wrong answer as normal work continues.

#715's ownership gate narrows this but does not close it, for the same reason it did not close #719: the
gate compares the creator, and the dangerous case is one long-lived lead whose creator check passes.

Remedy, if the premise holds

Mirror #719: fold a per-boot nonce into the id so staleness is carried by the string the caller holds.
MessageService now does this with a ticketBootNonce field, and turnId can take the same shape.

Do not record the boot id server-side and compare it. #719 tried that design and withdrew it: the
caller presents only the string, so a server-side boot id has nothing to compare against. The general
rule from that ticket applies unchanged — staleness must be carried by the identifier the caller
holds.

Measured vs not

  • Measured: the turnId format and that askSeq is an unpersisted per-Rendezvous AtomicLong. Both
    from the grep above.
  • Measured elsewhere: the identical ticket-id defect, driven end to end on #719.
  • Not measured: whether a member terminal id can recur in a later daemon boot. Without that, the
    collision cannot happen, and this ticket should be closed rather than implemented.
  • Not measured: whether any lead has ever written a turnId into a handover. The charter encourages
    writing state down, but I did not grep the handover files for the # form.

Found while reviewing PR #728 (#719). Not urgent, and it must not be bundled into that PR.

## The shape #719 fixed ticket ids restarting at `task-1` on every daemon restart. A worker flagged `turnId` as having the same shape while implementing it. I read the code and it does, in part. Measured 2026-10-04 on `main`: ``` $ grep -n 'askSeq\|turnId =' fleetd/src/main/java/dev/ltms/fleet/msg/Rendezvous.java 103: private final AtomicLong askSeq = new AtomicLong(); 189: String turnId = openAsksBySession.computeIfAbsent(session, _ -> { 190: String newTurnId = session + "#" + askSeq.incrementAndGet(); ``` So `turnId` is `<session>#<n>`, and `askSeq` is a per-`Rendezvous` `AtomicLong` with nothing persisting it. The counter half resets on every daemon restart, exactly like `ticketSeq` did. ## Why this is weaker than #719, not equal to it A ticket id was **entirely** minted by fleetd, so a reset made the whole id reusable. A `turnId` is two parts, and only the second part resets: - `<n>` — resets to 0 on every restart. Reusable. - `<session>` — a herdr `terminal_id`. **herdr is a separate process and a fleetd restart does not restart it**, so herdr keeps issuing fresh terminal ids and existing ones keep their value. A collision therefore needs the **same session string** to appear again after the restart, not just the counter to climb back. That is a real extra condition, and it is why this is filed separately rather than folded into #719. ## The premise I have not checked For a collision, a member terminal must **survive a fleetd restart and be re-adopted into the new daemon's roster**, then ask again. I have not read how the roster is rebuilt after a restart, and I have not driven it. If a surviving terminal is never re-adopted — if every post-restart member gets a brand new terminal id — then the `<session>` half never repeats and this is not reachable at all. **Someone must establish that before writing code.** A fix for an unreachable path is cost with no benefit, and I am not going to guess. This is the one question the ticket turns on. ## If it is reachable, who gets hurt The same party as in #719: whoever wrote the id down. A lead answers a member's question with `fleet_send{turnId, content}`. A lead that recorded a `turnId` — in `.handover/HANDOVER.md` or a ticket comment, which this project's charter tells it to do — could, after a restart, send an answer to a **different** open ask. The rendezvous state is in memory, so immediately after a restart there are no open asks and a stale `turnId` fails safely. The dangerous window opens later, once `askSeq` has climbed back to the same number and the same session has asked again. That is the #719 pattern precisely: a stale id that starts as a clean miss and becomes a wrong answer as normal work continues. #715's ownership gate narrows this but does not close it, for the same reason it did not close #719: the gate compares the **creator**, and the dangerous case is one long-lived lead whose creator check passes. ## Remedy, if the premise holds Mirror #719: fold a per-boot nonce into the id so staleness is carried by the string the caller holds. `MessageService` now does this with a `ticketBootNonce` field, and `turnId` can take the same shape. Do **not** record the boot id server-side and compare it. #719 tried that design and withdrew it: the caller presents only the string, so a server-side boot id has nothing to compare against. The general rule from that ticket applies unchanged — **staleness must be carried by the identifier the caller holds.** ## Measured vs not - Measured: the `turnId` format and that `askSeq` is an unpersisted per-`Rendezvous` `AtomicLong`. Both from the grep above. - Measured elsewhere: the identical ticket-id defect, driven end to end on #719. - **Not measured:** whether a member terminal id can recur in a later daemon boot. Without that, the collision cannot happen, and this ticket should be closed rather than implemented. - **Not measured:** whether any lead has ever written a `turnId` into a handover. The charter encourages writing state down, but I did not grep the handover files for the `#` form. Found while reviewing PR #728 (#719). Not urgent, and it must not be bundled into that PR.
Author
Owner

Decision: fix it, rather than keep waiting on the gating fact

This ticket was gated on establishing whether a member terminal id can recur in a later daemon boot,
with "if it cannot, close this instead". I tried to establish it and it cannot be answered from this
repo.

fleetd never mints a terminal id. The only place the string appears in src/main/java is a prefix
test:

// AgentControl.java:69
if (target == null || !target.startsWith("term_")) {
    return target;
}

The ids come from herdr, and herdr's source is not checked in here (ls -d ../herdr* → no match). I
also tried to read the property off the ids themselves, and that failed: term_65d0665bc875586,
term_65c806b784dc28 and term_65910edceb7f267 decode to 458529771086108038, 28648903822072872 and
457415450716271207, which are not epoch seconds, millis, micros or nanos in any reading I tried. The
three I minted today do share a prefix and do increase monotonically, which hints at a counter or a
clock, but a hint is not the property. I did not establish it, and I am not claiming it either way.

So the gate cannot be satisfied, and leaving the ticket open waits on a fact nobody here can read.

What is being done

Apply #719's fix. Rendezvous.java:103 holds a per-instance AtomicLong askSeq that restarts at 1
on every daemon boot, and :190 mints session + "#" + askSeq.incrementAndGet(). Folding in a
per-instance nonce, exactly as MessageService.ticketBootNonce now does for ticket ids, makes the
question moot: an id minted by one instance can never match another instance's id space, whatever
herdr does with terminal ids.

That is one field, one import and one changed expression. Cheaper than owning an open question about
a component we cannot read.

What the risk actually is, stated narrowly

For this to bite, boot B must mint the same terminal id and reach the same sequence number,
and answerAsk would then resolve an unrelated waiter with no error. A stale turnId from boot A
cannot resolve against boot B on its own, because boot B's asks map starts empty. So this is narrow
— but it is the same shape as #719, which was also narrow and was also worth three lines.

Delegated with mutation evidence required, including a check that the test is not vacuous in the way
#719's first version was. The worker is also sweeping src/main/java for any other identifier built
from a per-instance counter with no per-boot component, and will report those without fixing them.

## Decision: fix it, rather than keep waiting on the gating fact This ticket was gated on establishing whether a member terminal id can recur in a later daemon boot, with "if it cannot, close this instead". I tried to establish it and **it cannot be answered from this repo.** fleetd never mints a terminal id. The only place the string appears in `src/main/java` is a prefix test: ```java // AgentControl.java:69 if (target == null || !target.startsWith("term_")) { return target; } ``` The ids come from herdr, and herdr's source is not checked in here (`ls -d ../herdr*` → no match). I also tried to read the property off the ids themselves, and that failed: `term_65d0665bc875586`, `term_65c806b784dc28` and `term_65910edceb7f267` decode to 458529771086108038, 28648903822072872 and 457415450716271207, which are not epoch seconds, millis, micros or nanos in any reading I tried. The three I minted today do share a prefix and do increase monotonically, which hints at a counter or a clock, but a hint is not the property. **I did not establish it, and I am not claiming it either way.** So the gate cannot be satisfied, and leaving the ticket open waits on a fact nobody here can read. ### What is being done Apply #719's fix. `Rendezvous.java:103` holds a per-instance `AtomicLong askSeq` that restarts at 1 on every daemon boot, and `:190` mints `session + "#" + askSeq.incrementAndGet()`. Folding in a per-instance nonce, exactly as `MessageService.ticketBootNonce` now does for ticket ids, makes the question moot: an id minted by one instance can never match another instance's id space, whatever herdr does with terminal ids. That is one field, one import and one changed expression. Cheaper than owning an open question about a component we cannot read. ### What the risk actually is, stated narrowly For this to bite, boot B must mint the *same terminal id* **and** reach the *same sequence number*, and `answerAsk` would then resolve an unrelated waiter with no error. A stale turnId from boot A cannot resolve against boot B on its own, because boot B's `asks` map starts empty. So this is narrow — but it is the same shape as #719, which was also narrow and was also worth three lines. Delegated with mutation evidence required, including a check that the test is not vacuous in the way #719's first version was. The worker is also sweeping `src/main/java` for any other identifier built from a per-instance counter with no per-boot component, and will report those without fixing them.
Author
Owner

Fixed and merged as 8bb2aa0 — closing

turnId now mints as session + "#" + askBootNonce + "-" + askSeq.incrementAndGet(), e.g.
term_a#f3a9c1-1. An id minted by one Rendezvous instance can no longer match another instance's id
space, so the gating question — whether a member terminal id can recur in a later daemon boot — no
longer has to be answered for this to be safe.

PR #732, merged locally and pushed; origin/main confirmed at 8bb2aa0be45ffe543f393ef7caf1e21a4bfcb9b9
by ref.

Verified by me, not taken from the report

  • Build on the pushed tree in a throwaway worktree: BUILD SUCCESS, 2057 tests, 0 failures,
    RendezvousTest 20 green. The merged tree hash is byte-identical to that branch tree, so no separate
    post-merge build was needed — that is a measurement, not an assumption.
  • Nothing parses a turnId. This was the real risk of changing the format and the report only
    covered the prefix. Every turnId use in src/main/java grepped for
    split|substring|indexOf|parse|lastIndexOf|replace|matches|charAt: all three hits are prose in
    comments. It is an opaque key in all six files that touch it.
  • The test is not vacuous: other is driven to the same sequence number, and the positive control
    asserts the minting instance still resolves its own id. With the nonce removed and the
    sequence-driving line deleted, the test falsely passed — which is the proof that line is
    load-bearing.

The gating question is still unanswered, and that is now fine

I could not establish whether a herdr terminal id can recur. fleetd never mints one
(AgentControl.java:69 only tests the term_ prefix), herdr's source is not in this repo, and the ids
do not decode as any epoch unit I tried. The fix makes the question moot rather than answering it.
Nobody should read this ticket as having settled it.

Follow-up filed separately

The sweep for other identifiers with the same shape turned up three, and reachability was established
for none of them. They are filed as their own ticket rather than left in this one's history:
BackendOutagePolicy.java:107, UnixSocketHerdrClient.java:69, HerdrPeerLauncher.java:558.

## Fixed and merged as `8bb2aa0` — closing `turnId` now mints as `session + "#" + askBootNonce + "-" + askSeq.incrementAndGet()`, e.g. `term_a#f3a9c1-1`. An id minted by one `Rendezvous` instance can no longer match another instance's id space, so the gating question — whether a member terminal id can recur in a later daemon boot — no longer has to be answered for this to be safe. PR #732, merged locally and pushed; `origin/main` confirmed at `8bb2aa0be45ffe543f393ef7caf1e21a4bfcb9b9` by ref. ### Verified by me, not taken from the report - Build on the pushed tree in a throwaway worktree: `BUILD SUCCESS`, 2057 tests, 0 failures, `RendezvousTest` 20 green. The merged tree hash is byte-identical to that branch tree, so no separate post-merge build was needed — that is a measurement, not an assumption. - **Nothing parses a `turnId`.** This was the real risk of changing the format and the report only covered the prefix. Every `turnId` use in `src/main/java` grepped for `split|substring|indexOf|parse|lastIndexOf|replace|matches|charAt`: all three hits are prose in comments. It is an opaque key in all six files that touch it. - The test is not vacuous: `other` is driven to the same sequence number, and the positive control asserts the minting instance still resolves its own id. With the nonce removed *and* the sequence-driving line deleted, the test falsely passed — which is the proof that line is load-bearing. ### The gating question is still unanswered, and that is now fine I could not establish whether a herdr terminal id can recur. fleetd never mints one (`AgentControl.java:69` only tests the `term_` prefix), herdr's source is not in this repo, and the ids do not decode as any epoch unit I tried. The fix makes the question moot rather than answering it. Nobody should read this ticket as having settled it. ### Follow-up filed separately The sweep for other identifiers with the same shape turned up three, and reachability was established for none of them. They are filed as their own ticket rather than left in this one's history: `BackendOutagePolicy.java:107`, `UnixSocketHerdrClient.java:69`, `HerdrPeerLauncher.java:558`.
ltms closed this issue 2026-10-04 19:00:56 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#729