After #726 ships, a lead handover will orphan every in-flight delegation: the fresh lead gets a new terminal id and ownsTicket refuses it #737

Closed
opened 2026-10-04 19:36:17 +02:00 by ltms · 14 comments
Owner

Found while deciding whether to roll this lead with two delegations running. This is a regression #726 unit 2 will introduce, not a defect in today's code — so it must be settled before unit 2 is deployed.

The two facts, read in the code

1. A named lead's ticket reads are checked against its terminal id.

// msg/MessageService.java:1463-1465
private static boolean ownsTicket(Task task, String callerTerminal) {
    return callerTerminal == null || callerTerminal.equals(task.creatorTerminal);
}
// msg/MessageService.java:1417-1419
if (!ownsTicket(task, callerTerminal)) {
    return new TaskView(ticket, Phase.FAILED, null, null,
            "forbidden: this ticket was created by a different session", null);
}

A named lead carries a real terminal, not null:

// auth/CallerResolver.java:308
return Principal.leader(lead, c.terminal(), c.pid());

// auth/Principal.java:55-57
public static Principal leader(String name, String terminal, long pid) {
    return new Principal(Role.PRIMARY, terminal, pid, name);
}

So the callerTerminal == null escape in ownsTicket is for the unnamed primary (Principal.primary(pid)), never for a configured lead.

2. #726 unit 2 gives the fresh lead a new terminal id.

That is the whole point of the ticket: end the old claude process, close its pane, create a new tab and start a new agent. A herdr terminal id belongs to the pane, and LeadLauncher.launch() creates a fresh tab every time. The new lead resolves through the same CallerResolver.java:308 rung — same lead name, different c.terminal().

The consequence

ownsTicket compares the thing that changes and ignores the thing that does not. After a roll, the fresh lead calling fleet_poll{ticket} on any ticket the previous lead created gets:

forbidden: this ticket was created by a different session

A lead rolls precisely when its context is full, which is usually mid-work with workers running. So the failure fires exactly when it costs most: every in-flight delegation's report becomes uncollectable, and the worker's whole turn is lost. The workers themselves are fine — they keep working and fleet_reply normally. Nobody can read the answer.

This falsifies a conclusion already recorded as safe

#705's comment 18617 answered a required question with:

A lead keeps the same terminal id across fleet_handover. The roll sends /clear and the bootstrap text to p.leadTerminal() — the same terminal recorded at open(), never a new pane […] So a lead's pre-handover tickets stay readable afterwards. Good: the fix does not break the rollover path.

That was true when written and true today, because the roll is a /clear into the same pane. Unit 2 is the change that makes it false. This is the "a fix can create reachability" shape from the other direction: a correctness fix elsewhere turns a previously-safe statement into a defect, and the statement reads as settled so nobody re-checks it.

REST is closed too, and one escape hatch exists

GET /tasks/{ticket} threads the caller terminal as well (PR #716), so dropping to REST from the lead's own pane hits the same refusal — its pid maps to its pane, so it resolves as the same named lead.

The hatch: a session that is not in a lead-labelled tab resolves as the unnamed primary, whose terminal is null, so ownsTicket returns true for every ticket. An ordinary terminal on this host can therefore collect an orphaned ticket. That is a real workaround and worth knowing, but it is a manual rescue by an operator, not a fix — and it depends on an authorization property (ANONYMOUS is denied TASK_READ, so null means "the primary", not "could not tell") that #705's comment already flagged as pinned by nothing.

Fix shapes, not equivalent

  1. Hand the tickets over with the pane. LeadRollover tells MessageService to reassign creatorTerminal from the old terminal to the new one, as part of the roll. Smallest change that keeps the ownership model intact. The reassignment is an authorization-relevant write, so it must be reachable only from the roll and must name the old terminal it is replacing, not accept an arbitrary pair.
  2. Key ownership on the stable identity. A configured lead has a name that survives the roll; the terminal does not. Checking the name for a named lead (and the terminal for everyone else) fixes the root cause rather than patching the one path. Larger, and it touches the #705 option-2 security fix, so it needs its own care — in particular it must not let lead A read lead B's tickets, which that fix deliberately stopped.
  3. Document the loss and drain before rolling. Make the roll refuse, or warn, while the lead has uncollected tickets. Cheapest, and it turns a silent loss into a visible one, but it makes a context-full lead unable to roll exactly when it needs to.

My own reading is that 1 or 3 must ship with or before unit 2 reaching the live daemon, and 2 is the better long-term answer. I am not deciding it here — the next lead should settle it, consulting architects if it is not obvious, and record the decision on this ticket.

Scope note for #726 unit 2

The unit 2 worker has not been asked to fix this, and should not widen its scope. Its brief is write-once and the correction is on #726's ticket instead.

Not measured

I have not rolled a lead with unit 2's code, because unit 2 is not merged. Everything above is read from the four code sites quoted, at main = 428a12a. Nobody has observed the refusal, and the statement it falsifies was itself read from code rather than measured.

Found while deciding whether to roll this lead with two delegations running. **This is a regression #726 unit 2 will introduce, not a defect in today's code** — so it must be settled before unit 2 is deployed. ## The two facts, read in the code **1. A named lead's ticket reads are checked against its terminal id.** ```java // msg/MessageService.java:1463-1465 private static boolean ownsTicket(Task task, String callerTerminal) { return callerTerminal == null || callerTerminal.equals(task.creatorTerminal); } ``` ```java // msg/MessageService.java:1417-1419 if (!ownsTicket(task, callerTerminal)) { return new TaskView(ticket, Phase.FAILED, null, null, "forbidden: this ticket was created by a different session", null); } ``` A named lead carries a **real** terminal, not null: ```java // auth/CallerResolver.java:308 return Principal.leader(lead, c.terminal(), c.pid()); // auth/Principal.java:55-57 public static Principal leader(String name, String terminal, long pid) { return new Principal(Role.PRIMARY, terminal, pid, name); } ``` So the `callerTerminal == null` escape in `ownsTicket` is for the **unnamed** primary (`Principal.primary(pid)`), never for a configured lead. **2. #726 unit 2 gives the fresh lead a new terminal id.** That is the whole point of the ticket: end the old `claude` process, close its pane, create a new tab and start a new agent. A herdr terminal id belongs to the pane, and `LeadLauncher.launch()` creates a fresh tab every time. The new lead resolves through the same `CallerResolver.java:308` rung — same lead *name*, **different** `c.terminal()`. ## The consequence `ownsTicket` compares the thing that changes and ignores the thing that does not. After a roll, the fresh lead calling `fleet_poll{ticket}` on any ticket the previous lead created gets: > `forbidden: this ticket was created by a different session` A lead rolls **precisely when its context is full**, which is usually mid-work with workers running. So the failure fires exactly when it costs most: every in-flight delegation's report becomes uncollectable, and the worker's whole turn is lost. The workers themselves are fine — they keep working and `fleet_reply` normally. Nobody can read the answer. ## This falsifies a conclusion already recorded as safe #705's comment 18617 answered a required question with: > **A lead keeps the same terminal id across `fleet_handover`.** The roll sends `/clear` and the bootstrap text to `p.leadTerminal()` — the same terminal recorded at `open()`, never a new pane […] So a lead's pre-handover tickets stay readable afterwards. Good: the fix does not break the rollover path. That was **true when written and true today**, because the roll is a `/clear` into the same pane. Unit 2 is the change that makes it false. This is the "a fix can create reachability" shape from the other direction: a correctness fix elsewhere turns a previously-safe statement into a defect, and the statement reads as settled so nobody re-checks it. ## REST is closed too, and one escape hatch exists `GET /tasks/{ticket}` threads the caller terminal as well (PR #716), so dropping to REST from the lead's own pane hits the same refusal — its pid maps to its pane, so it resolves as the same named lead. The hatch: a session that is **not** in a lead-labelled tab resolves as the unnamed primary, whose terminal is null, so `ownsTicket` returns true for every ticket. An ordinary terminal on this host can therefore collect an orphaned ticket. That is a real workaround and worth knowing, but it is a manual rescue by an operator, not a fix — and it depends on an authorization property (`ANONYMOUS` is denied `TASK_READ`, so null means "the primary", not "could not tell") that #705's comment already flagged as pinned by nothing. ## Fix shapes, not equivalent 1. **Hand the tickets over with the pane.** `LeadRollover` tells `MessageService` to reassign `creatorTerminal` from the old terminal to the new one, as part of the roll. Smallest change that keeps the ownership model intact. The reassignment is an authorization-relevant write, so it must be reachable only from the roll and must name the old terminal it is replacing, not accept an arbitrary pair. 2. **Key ownership on the stable identity.** A configured lead has a **name** that survives the roll; the terminal does not. Checking the name for a named lead (and the terminal for everyone else) fixes the root cause rather than patching the one path. Larger, and it touches the #705 option-2 security fix, so it needs its own care — in particular it must not let lead A read lead B's tickets, which that fix deliberately stopped. 3. **Document the loss and drain before rolling.** Make the roll refuse, or warn, while the lead has uncollected tickets. Cheapest, and it turns a silent loss into a visible one, but it makes a context-full lead unable to roll exactly when it needs to. My own reading is that 1 or 3 must ship **with or before** unit 2 reaching the live daemon, and 2 is the better long-term answer. I am not deciding it here — the next lead should settle it, consulting architects if it is not obvious, and record the decision on this ticket. ## Scope note for #726 unit 2 The unit 2 worker has **not** been asked to fix this, and should not widen its scope. Its brief is write-once and the correction is on #726's ticket instead. ## Not measured I have not rolled a lead with unit 2's code, because unit 2 is not merged. Everything above is read from the four code sites quoted, at `main` = `428a12a`. Nobody has observed the refusal, and the statement it falsifies was itself read from code rather than measured.
Author
Owner

Delegated to two architects, and the question is wider than the issue body says

Two architects are working this now, independently, with an identical brief — sol (term_65d07475d124a8e) and opus (term_65d0747d6e8fe8f). Neither has seen the other's answer.

I widened the question, because the issue body under-scopes it. I enumerated exactly one channel — tickets — and then wrote a conclusion about handover in general. That is a mistake with a name here: enumerate one channel, conclude about all of them. The real question is not "do tickets break" but:

What does a lead lose when its terminal id changes?

So the brief asks for an enumeration of every read and write keyed on a lead's terminal, each ruled in or out with a file:line, covering at least Task.creatorTerminal, the per-session reply inbox, rendezvous entries and fleet_ask turnIds, PrimaryRegistry, LeadRollover's own post-roll p.leadTerminal() uses, the coordinator's held lead-to-lead mail, and any terminal match in Authz or Injector. Tickets may turn out to be the only one. That would be a finding; assuming it is not.

The brief also asks each architect to check my three code reads rather than take them, and to mark every claim as read-in-the-code, inferred, or not-checked. I told them explicitly that finding my reading wrong is a useful answer.

On the three fix shapes, each must answer five things that are where a fix of this kind goes wrong: whether it keeps lead A out of lead B's tickets (#705 option 2 closed that deliberately, and reopening it would be worse than this bug), whether a reassignment stays reachable only from the roll, who owns the state when a roll dies half way, whether the unnamed-primary null escape has to change, and whether Task can even see a lead name today.

I am not asking them for a verdict. On #726 two architects agreed on a conclusion while both had the grace-scan mechanism backwards, and their agreement made that harder to catch rather than easier. So the brief asks for mechanisms, and I will read those rather than count votes. If they disagree after their independent pass, I will exchange their positions and ask each to compare; if they still disagree, the decision is mine and I will record it here with the reason.

Nothing is decided yet. This comment records who is working it and what they were actually asked.

## Delegated to two architects, and the question is wider than the issue body says Two architects are working this now, independently, with an identical brief — `sol` (`term_65d07475d124a8e`) and `opus` (`term_65d0747d6e8fe8f`). Neither has seen the other's answer. **I widened the question, because the issue body under-scopes it.** I enumerated exactly one channel — tickets — and then wrote a conclusion about handover in general. That is a mistake with a name here: enumerate one channel, conclude about all of them. The real question is not "do tickets break" but: > **What does a lead lose when its terminal id changes?** So the brief asks for an enumeration of every read and write keyed on a lead's terminal, each ruled in or out with a `file:line`, covering at least `Task.creatorTerminal`, the per-session reply inbox, rendezvous entries and `fleet_ask` turnIds, `PrimaryRegistry`, `LeadRollover`'s own post-roll `p.leadTerminal()` uses, the coordinator's held lead-to-lead mail, and any terminal match in `Authz` or `Injector`. Tickets may turn out to be the only one. That would be a finding; assuming it is not. The brief also asks each architect to **check my three code reads rather than take them**, and to mark every claim as read-in-the-code, inferred, or not-checked. I told them explicitly that finding my reading wrong is a useful answer. On the three fix shapes, each must answer five things that are where a fix of this kind goes wrong: whether it keeps lead A out of lead B's tickets (#705 option 2 closed that deliberately, and reopening it would be worse than this bug), whether a reassignment stays reachable only from the roll, who owns the state when a roll dies half way, whether the unnamed-primary `null` escape has to change, and whether `Task` can even see a lead name today. **I am not asking them for a verdict.** On #726 two architects agreed on a conclusion while both had the grace-scan mechanism backwards, and their agreement made that harder to catch rather than easier. So the brief asks for mechanisms, and I will read those rather than count votes. If they disagree after their independent pass, I will exchange their positions and ask each to compare; if they still disagree, the decision is mine and I will record it here with the reason. Nothing is decided yet. This comment records who is working it and what they were actually asked.
Author
Owner

Architect A (sol) — eight things break, not one. My issue body was badly under-scoped.

I have not re-verified these line references myself yet, and the second architect is still working. This comment records architect A's report so it is durable. Treat every row as architect A's reading until a second pair of eyes confirms it. Its own claim-marking (read in the code / inferred / not checked) is preserved below.

It checked at 428a12a.

Breaks when a lead's terminal id changes

Thing Mechanism (architect A's reading) Why it breaks
Async ticket owner Task.creatorTerminal written at MessageService.java:285-296 and :1331-1334; refused at :1412-1420 via ownsTicket :1463-1465 the fresh lead cannot poll a predecessor's ticket. Both MCP (FleetMcp.java:1225-1240) and REST (FleetApp.java:901-907) pass the new terminal into the same gate
Pending question in fleet_status same ownsTicket gate in pendingAsk at MessageService.java:1793-1797 the fresh lead keeps the worker status but loses the question, ticket and turnId block
fleet_ask answer Rendezvous.openAsk copies ownerOf(workerSession) at Rendezvous.java:195-203; MessageService.answer requires an exact terminal match at :1175-1182 the turnId survives, its stored owner does not — the answer is refused as NOT_TURN_OWNER
Nudge routing PrimaryRegistry.leadByTarget at :30-36,79-84; nudgeTargetFor at :103-105 every worker the old lead delegated keeps routing nudges to the dead terminal
Pending reminders ReplyPushLoop stores the lead terminal in ReplyEntry / PendingTicket / PendingQuestion at :96-114,178-206; tick sends at :394-405,753-776 entries are never re-keyed, a failed status read returns WAIT_BUSY, so it retries the dead address forever. resolveLiveLead is not called by that tick
Idle heartbeat / context warning LeadHeartbeatLoop.tick reads only PrimaryRegistry.primaryTerminal() at :309-327, sends at :364-397 breaks until the singleton moves — and a pinned singleton (PrimaryRegistry.java:58-60) never recovers by itself
Rollover single-flight rollingByTerminal keyed by the old terminal at LeadRollover.java:313-322, claimed :505-513, released in runRollover's finally :572-589 inferred for unit 2: the fresh terminal is a different map key, so a second roll can start while the predecessor's continuation still holds the old key
Queued local messages Injector.targets keyed by target terminal at :124, enqueue :342-358 a peer message already queued to the old lead address is never moved; the peer must refresh fleet_list and retry

Survives

Thing Why
Forward rendezvous + final worker reply keyed by the worker session (Rendezvous.java:96-103, :308-310), which does not change. The reply still completes the ticket; only reading it then fails at ownsTicket
Worker reply inbox keyed by worker target (InMemoryReplyInbox.java:9-24, AmqpReplyInbox.java:32-40,85-86), and DRAIN is primary-wide (Authz.java:103-108)
Coordinator held mail LeadMailbox owns lead.<selfCoordId>.inbox (:197-208); LeadCoordLoop re-reads the live terminal→name map each tick (:194-220)
LeadRollover pending request confirm runs before the old pane dies; status(token) is token-keyed (:666-686)
Authz no hidden owner state — the ticket and ask breaks are in MessageService and Rendezvous, not the role table

Corrections to what I wrote in this issue

  • The reply inbox does not break. I listed it as a candidate; it is keyed by worker target and stays drainable.
  • The coordinator's stable key is the configured daemon selfCoordId, not "the host". My wording was loose.
  • My null-escape claim needs a qualifier. "The escape is for the unnamed primary only" is true at an authorized production entry point, but ownsTicket itself cannot tell why its argument is null — safety depends on Authz refusing anonymous callers first. #705 comment 18617 already flagged that coupling as unpinned.
  • My three quoted code reads were confirmed correct at this commit.
  • "Unit 2 gives a new terminal" is not readable in current main, because unit 2 is not merged. Architect A accepted it as inferred from #726's design and comments, not as a fact read from the branch. That is the right way to hold it, and I should have written it that way.

Architect A's proposed fix — option 2, broadened

Not just a creatorLeadName field. One typed, immutable owner identity, built only from the connection-resolved Principal and never from a request argument:

  • NamedLead(name) for a configured lead
  • Session(terminal) for an architect or other terminal-bound caller
  • UnnamedPrimary for the off-pane primary

Then resolve name → current terminal late, at each send, rather than storing an address. Applied to Task, Rendezvous.Owner, PrimaryRegistry.leadByTarget, ReplyPushLoop's schedules, the heartbeat, and LeadRollover's single-flight key.

Two deliberate restrictions in its design, both of which I think are the right instincts:

  • No transferTicket(old, new) primitive. The owner is immutable from creation; only the address is resolved late. That answers my "does this become a general owner-change primitive" question with "there is nothing to reach".
  • An architect keeps terminal ownership, not name ownership, so a replacement architect cannot inherit its predecessor's tickets. Only Role.PRIMARY with a non-null configured name gets NamedLead.

On lead A versus lead B: NamedLead("A") matches only NamedLead("A"), and the name comes from CallerResolver.java:302-308, not the request — so #705's boundary holds.

On old tickets: no migration needed. Tasks live only in an in-memory map (MessageService.java:322-328) and ticket ids carry a per-boot nonce (:357-365), and loading new code requires a restart, which builds a new MessageService. So no ticket from the old class can exist inside the new one.

Sequencing

Architect A: may merge before, must not deploy before. Its reason is that a warning is not enough, because a worker can finish or raise a question during the roll, after any pre-roll drain check. That matches what I wrote, with a better reason than I gave.

What architect A did not check

No unit 2 code (not merged in its worktree), no live rollover, no OpenCode restart behaviour, no fleetd.yaml, no wiki/, and no Maven run — it changed no code. Its final broad rg sweep did not run (rg is not installed in the worktree); the table comes from the editor Grep and Read tools, which did run.

Still to do

Architect B (opus) is still working and has not seen this. I will read its mechanism rather than count votes, post it, and then adjudicate. One item needs carrying to #726 regardless: the rollover single-flight claim unit 3 just merged is keyed on the lead terminal, which is exactly what unit 2 changes.

## Architect A (`sol`) — eight things break, not one. My issue body was badly under-scoped. **I have not re-verified these line references myself yet, and the second architect is still working.** This comment records architect A's report so it is durable. Treat every row as architect A's reading until a second pair of eyes confirms it. Its own claim-marking (read in the code / inferred / not checked) is preserved below. It checked at `428a12a`. ### Breaks when a lead's terminal id changes | Thing | Mechanism (architect A's reading) | Why it breaks | |---|---|---| | Async ticket owner | `Task.creatorTerminal` written at `MessageService.java:285-296` and `:1331-1334`; refused at `:1412-1420` via `ownsTicket` `:1463-1465` | the fresh lead cannot poll a predecessor's ticket. Both MCP (`FleetMcp.java:1225-1240`) and REST (`FleetApp.java:901-907`) pass the new terminal into the same gate | | Pending question in `fleet_status` | same `ownsTicket` gate in `pendingAsk` at `MessageService.java:1793-1797` | the fresh lead keeps the worker status but **loses the question, ticket and `turnId` block** | | `fleet_ask` answer | `Rendezvous.openAsk` copies `ownerOf(workerSession)` at `Rendezvous.java:195-203`; `MessageService.answer` requires an exact terminal match at `:1175-1182` | the `turnId` survives, its stored owner does not — the answer is refused as `NOT_TURN_OWNER` | | Nudge routing | `PrimaryRegistry.leadByTarget` at `:30-36,79-84`; `nudgeTargetFor` at `:103-105` | every worker the old lead delegated keeps routing nudges to the dead terminal | | Pending reminders | `ReplyPushLoop` stores the lead terminal in `ReplyEntry` / `PendingTicket` / `PendingQuestion` at `:96-114,178-206`; tick sends at `:394-405,753-776` | entries are never re-keyed, a failed status read returns `WAIT_BUSY`, so it retries the dead address forever. `resolveLiveLead` is **not** called by that tick | | Idle heartbeat / context warning | `LeadHeartbeatLoop.tick` reads only `PrimaryRegistry.primaryTerminal()` at `:309-327`, sends at `:364-397` | breaks until the singleton moves — and **a pinned singleton (`PrimaryRegistry.java:58-60`) never recovers by itself** | | Rollover single-flight | `rollingByTerminal` keyed by the old terminal at `LeadRollover.java:313-322`, claimed `:505-513`, released in `runRollover`'s `finally` `:572-589` | **inferred for unit 2:** the fresh terminal is a different map key, so a second roll can start while the predecessor's continuation still holds the old key | | Queued local messages | `Injector.targets` keyed by target terminal at `:124`, enqueue `:342-358` | a peer message already queued to the old lead address is never moved; the peer must refresh `fleet_list` and retry | ### Survives | Thing | Why | |---|---| | Forward rendezvous + final worker reply | keyed by the **worker** session (`Rendezvous.java:96-103`, `:308-310`), which does not change. The reply still completes the ticket; only *reading* it then fails at `ownsTicket` | | Worker reply inbox | keyed by worker target (`InMemoryReplyInbox.java:9-24`, `AmqpReplyInbox.java:32-40,85-86`), and `DRAIN` is primary-wide (`Authz.java:103-108`) | | Coordinator held mail | `LeadMailbox` owns `lead.<selfCoordId>.inbox` (`:197-208`); `LeadCoordLoop` re-reads the live terminal→name map each tick (`:194-220`) | | `LeadRollover` pending request | `confirm` runs before the old pane dies; `status(token)` is token-keyed (`:666-686`) | | `Authz` | no hidden owner state — the ticket and ask breaks are in `MessageService` and `Rendezvous`, not the role table | ### Corrections to what I wrote in this issue - **The reply inbox does not break.** I listed it as a candidate; it is keyed by worker target and stays drainable. - **The coordinator's stable key is the configured daemon `selfCoordId`, not "the host".** My wording was loose. - **My `null`-escape claim needs a qualifier.** "The escape is for the unnamed primary only" is true at an authorized production entry point, but `ownsTicket` itself cannot tell *why* its argument is null — safety depends on `Authz` refusing anonymous callers first. #705 comment 18617 already flagged that coupling as unpinned. - My three quoted code reads were confirmed correct at this commit. - **"Unit 2 gives a new terminal" is not readable in current `main`**, because unit 2 is not merged. Architect A accepted it as inferred from #726's design and comments, not as a fact read from the branch. That is the right way to hold it, and I should have written it that way. ### Architect A's proposed fix — option 2, broadened Not just a `creatorLeadName` field. One typed, immutable owner identity, built only from the connection-resolved `Principal` and never from a request argument: - `NamedLead(name)` for a configured lead - `Session(terminal)` for an architect or other terminal-bound caller - `UnnamedPrimary` for the off-pane primary Then resolve `name → current terminal` **late**, at each send, rather than storing an address. Applied to `Task`, `Rendezvous.Owner`, `PrimaryRegistry.leadByTarget`, `ReplyPushLoop`'s schedules, the heartbeat, and `LeadRollover`'s single-flight key. Two deliberate restrictions in its design, both of which I think are the right instincts: - **No `transferTicket(old, new)` primitive.** The owner is immutable from creation; only the *address* is resolved late. That answers my "does this become a general owner-change primitive" question with "there is nothing to reach". - **An architect keeps terminal ownership**, not name ownership, so a replacement architect cannot inherit its predecessor's tickets. Only `Role.PRIMARY` with a non-null configured name gets `NamedLead`. On lead A versus lead B: `NamedLead("A")` matches only `NamedLead("A")`, and the name comes from `CallerResolver.java:302-308`, not the request — so #705's boundary holds. On old tickets: no migration needed. Tasks live only in an in-memory map (`MessageService.java:322-328`) and ticket ids carry a per-boot nonce (`:357-365`), and loading new code requires a restart, which builds a new `MessageService`. So no ticket from the old class can exist inside the new one. ### Sequencing Architect A: **may merge before, must not deploy before.** Its reason is that a warning is not enough, because a worker can finish or raise a question during the roll, after any pre-roll drain check. That matches what I wrote, with a better reason than I gave. ### What architect A did not check No unit 2 code (not merged in its worktree), no live rollover, no OpenCode restart behaviour, no `fleetd.yaml`, no `wiki/`, and no Maven run — it changed no code. Its final broad `rg` sweep did not run (`rg` is not installed in the worktree); the table comes from the editor Grep and Read tools, which did run. ### Still to do Architect B (`opus`) is still working and has not seen this. I will read its mechanism rather than count votes, post it, and then adjudicate. **One item needs carrying to #726 regardless:** the rollover single-flight claim unit 3 just merged is keyed on the lead terminal, which is exactly what unit 2 changes.
Author
Owner

Decision: option 2, role-prefixed. Both architects chose it. And one claim in my issue body is wrong — correcting it first.

Architect B (opus) has reported. I re-read every line it corrected, in the tree now at main = 11998cd, and I am quoting what I read rather than relaying it.


First: my escape-hatch paragraph is wrong. Do not use it.

I wrote in this issue:

The hatch: a session that is not in a lead-labelled tab resolves as the unnamed primary, whose terminal is null, so ownsTicket returns true for every ticket. An ordinary terminal on this host can therefore collect an orphaned ticket.

That is false, and it would make an operator's rescue fail silently. Any herdr pane gets a terminal, and an unmatched pane does not fall through to the unnamed primary — it hits the pane floor:

// auth/CallerResolver.java:326  (at main = 11998cd)
return Principal.observer(c.terminal(), c.pid()); // unforgeable; never token-gated

A non-null terminal, so ownsTicket refuses it. The unnamed primary is reached only when no pane resolves at all — a plain terminal application that is not a herdr pane. An operator on this host is usually already inside a herdr pane, which is exactly where my stated rescue does not work.

And it is now closed harder than that. #705 option 1 merged a few minutes ago (11998cd), narrowing that floor from WORKER to OBSERVER, and OBSERVER is not in the TASK_READ gate at all:

// auth/Authz.java:149
case TASK_READ -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();

So a herdr pane that is not a lead, member, architect or collaborator is now refused at the role gate, before any owner comparison. Architect B flagged that the hatch line should not be relied on long-term; it is already gone.


Where the two architects agree, and where they do not

Both independently chose option 2, both rejected a transfer primitive, and both said the same thing about sequencing: free to merge, coupled to deploy. I am not treating that agreement as a check — on #726 two architects agreed while both had a mechanism backwards — so I read their mechanisms, and the mechanisms differ in useful ways.

They disagree on exactly one row: unit 3's single-flight claim

  • A: breaks — the fresh terminal is a different map key, so a second roll can start while the predecessor's continuation still holds the old key.
  • B: survives — it is claimed and released on the same old terminal, in runRollover's finally, and nothing re-reads it later.

They are both right, about different properties, and neither said which property it was answering. I read the code: the claim is taken with putIfAbsent(p.leadTerminal(), token) and released with the two-argument remove(p.leadTerminal(), p.token()) — the same key both times. So B is right that the release is correct and nothing leaks. A is right that it stops being a per-lead lock: once the terminal changes mid-roll, a second confirm() from the fresh lead lands on a different key and is not excluded.

So this is not a contradiction but an under-specification: the claim does what it was built to do (release reliably) and no longer does a thing nobody had asked it to do yet (exclude a second roll of the same lead). That is a real, narrow gap, and it belongs in the fix — B's design already closes it by keying single-flight on the lead identity rather than the pane.

This is why the brief asked for mechanisms. Two bare verdicts here would have been a coin toss.

B found a third break neither I nor A ranked correctly

fleet_send{turnId} — answering a worker's fleet_ask — breaks through a different gate:

// msg/Rendezvous.java:91-93  (javadoc above it, verbatim)
// Whether callerTerminal matches owner. A null owner means no owner was ever recorded, and
// that state matches no caller, not even one whose own terminal is null — "no record" and
// "recorded as the unnamed primary" are different states.
public static boolean permits(Owner owner, String callerTerminal) {
    return owner != null && java.util.Objects.equals(owner.terminal(), callerTerminal);
}

B's severity argument is right and I had it wrong. A refused poll loses a report. A refused answer loses control of live work: the worker's ask window is about 55 seconds, nothing extends it, so it resumes unanswered and takes whatever default it had. That can produce wrong code, not a missing message.

And note what the javadoc already says: ownsTicket and Owner.permits treat null differently on purpose. Any fix touches both and must not quietly align them. The ask path's fail-closed default is deliberate and documented at the line.

B closed the escape I was half relying on, with a mechanism I can confirm from this session

B's claim: the ungated fleet_poll{target} inbox drain cannot rescue an in-flight delegation, because reply() gives the text to the open waiter and never publishes to the inbox:

// msg/MessageService.java:560
if (rendezvous.resolve(session, content)) {
    count(FleetMetrics.REPLIES, "path", "rendezvous");
    return ReplyOutcome.RESOLVED_SEND; // a live send took it — unchanged fast path
}
// ... only when no waiter took it:
// msg/MessageService.java:602
inbox.publish(session, UUID.randomUUID().toString(), content);

and the forward waiter stays open for ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L (:56). So for 30 minutes there is no route: the ticket is refused and the inbox is empty.

I can corroborate this from my own session today, and the two readings differ in the axis that matters:

Member Ticket state when I polled fleet_poll{target} fleet_poll{ticket}
#705 implementer timed_out_working (past 30 min) returned the full report failed
architect A still live returned [] returned the full report

That is exactly B's mechanism, observed twice, varying ticket-expiry rather than repeating one case. It also means the stranding is real but time-bounded, as B said.


The decision

Option 2, with architect B's restriction, which is not optional.

Give Task and Rendezvous.Owner a stable, role-prefixed owner key derived only from the connection-resolved Principal:

  • Role.PRIMARY with a non-null name → "leader:" + name
  • everything else unchanged → "worker:" + terminal, "architect:" + slot, "collaborator:" + name
  • unnamed primary (null name, null terminal) → null, so today's escape is untouched

Three things I am adopting because the reasoning behind them is sound, not because an architect said them:

  1. Prefix, never a bare name. A bare name would let an architect slot named opus collide with the lead named opus. And key only a configured lead on its name — an architect must keep terminal ownership, or a newly spawned architect in the same slot inherits its predecessor's tickets and open turns. That is the one way option 2 could end up worse than the bug.
  2. Do not reuse Principal.describe(). It returns the constant "primary" for the unnamed primary, which would destroy the null escape, and its javadoc scopes it to audit lines. Add a separate ownerKey().
  3. No transfer primitive at all. Ownership is derived per call, as today; only the address is resolved late. There is nothing for another caller to reach, so there is nothing to guard.

The argument that actually decided it is B's failed-roll case, and it is the one I would have got wrong on my own. Option 1 is a one-shot write that can only run after a new terminal exists. If the bounded relaunch exhausts its retries there is no new terminal, the roll is terminally FAILED with no retry, and the state stays owned by a dead terminal — unreadable by every terminal-bearing caller, permanently. Option 1 delivers nothing in exactly the case I asked about. Option 2 has no event to miss: ownership is leader:opus whether a pane exists or not, so the next session to hold that lead's tab label can read it.

Put plainly: after a failed roll, option 1 leaves the state owned by nobody; option 2 leaves it owned by the next session in that lead's tab.

B also noted there is no migration surface at all — tasks is an in-memory map and ticket ids carry a per-boot nonce, so new code can only arrive via a restart that empties it. The change that looks larger on paper has zero compatibility cost.

Option 3 survives as a warning, not a refusal

fleet_handover{action:"open"} should report how many tickets and open ask-turns the caller owns. Useful whichever fix lands. It must not refuse: a lead rolls when its context is full, which is the worst possible moment to be told it may not, and "drain first" is not achievable in bounded time when a worker is mid-brief against a 30-minute clock.

Sequencing — and I am not accepting "deploy with the loss documented"

Unit 2 may merge before this fix. The fix must be in the jar before unit 2 is.

B's reason for refusing the documented-caveat option is better than mine, so I am taking it: the daemon actively pushes the lead toward the unsafe action. LeadHeartbeatLoop carries a context-high nudge, and the live bootstrapText is already written for a handover. So the system tells the lead to hand over precisely when its context is full — and a context-full lead is the least reliable reader of a documented caveat. A mitigation that lives in the judgement of the agent with the least context left is not a mitigation.

That is also the shape this whole ticket came from: a safe answer recorded on #705 expired when another ticket shipped. Deploying unit 2 with a second written caveat would recreate the same failure one level up.


What is NOT settled, and what nobody has measured

  • B's widening argument is load-bearing and unproven, and B said so. Option 2 means a lead's pane replaced outside fleet_handover — an operator closing the tab and opening a new one with the same label — also inherits the tickets. B argues this is acceptable because whoever can set a lead's tab label already gets Role.PRIMARY from CallerResolver.java:308, which already grants fleet_spawn and fleet_stop, so inheriting tickets is strictly less than the label already grants. I find that convincing and I have not tested it. The untested case B named is a mislabelled tab with no agent behind it. Whoever implements this should probe it before relying on the argument.
  • Nobody has seen any of these refusals happen. Unit 2 is not merged, no roll has run under it, and both architects ran zero tests — correctly, as neither wrote code. Every break is read from a gate plus the fact that the terminal changes.
  • Neither architect read unit 2's diff, because it does not exist yet. "The fresh lead gets a new terminal" remains inferred from #726's design and from LeadLauncher creating a fresh tab. If unit 2 somehow reuses the pane, this whole ticket is moot.
  • creatorTerminal's rename will turn a wide set of tests red — MessageService.poll(String) alone is reported to have 44 test callers. That figure is from #705 comment 18676 and I have not counted it myself.
  • B's "degrades, self-heals" verdicts for ReplyPushLoop and LeadHeartbeatLoop depend on herdr answering agent_not_found for a closed pane's terminal. That is not verified for pane.close — the reported evidence is for /exit, a different teardown.
  • A's rg sweep did not run (rg is not installed in a worktree); its table came from editor Grep and Read, which did.

Implementation units, for whoever picks this up

Roughly A's split, with B's restriction folded in. Not yet briefed, and deliberately not started tonight:

  1. Identity: add Principal.ownerKey(), role-prefixed, lead-name only for PRIMARY with a non-null name. Leave describe() alone. Cover named lead, architect, collaborator, unnamed primary and anonymous.
  2. Messaging: use it in Task, Rendezvous.Owner, poll, pendingAsk and answer. Acceptance: lead A's old terminal creates, A's new terminal polls and answers, lead B is refused both. Keep the two null rules distinct and say in one place why.
  3. Routing: PrimaryRegistry, ReplyPushLoop and the heartbeat keep a stable lead identity and resolve the terminal at send time. Acceptance: a reminder pending before the terminal changes reaches the new terminal.
  4. Rollover: key single-flight on the lead identity, keep the old pane address separately as the thing to close. Acceptance: a failed roll leaves A's state owned by A, B refused, and a later A can continue.
  5. Warning: open reports owned tickets and open asks.
  6. A fake that returns the same terminal proves nothing here. Every test must actually change the terminal.

The one item to carry to #726 now: unit 3's single-flight claim is keyed on the lead terminal, which is what unit 2 changes.

## Decision: option 2, role-prefixed. Both architects chose it. And one claim in my issue body is wrong — correcting it first. Architect B (`opus`) has reported. I re-read every line it corrected, in the tree now at `main` = `11998cd`, and I am quoting what I read rather than relaying it. --- ## First: my escape-hatch paragraph is wrong. Do not use it. I wrote in this issue: > The hatch: a session that is **not** in a lead-labelled tab resolves as the unnamed primary, whose terminal is null, so `ownsTicket` returns true for every ticket. An ordinary terminal on this host can therefore collect an orphaned ticket. **That is false, and it would make an operator's rescue fail silently.** Any herdr pane gets a terminal, and an unmatched pane does not fall through to the unnamed primary — it hits the pane floor: ```java // auth/CallerResolver.java:326 (at main = 11998cd) return Principal.observer(c.terminal(), c.pid()); // unforgeable; never token-gated ``` A non-null terminal, so `ownsTicket` refuses it. The unnamed primary is reached only when **no pane resolves at all** — a plain terminal application that is not a herdr pane. An operator on this host is usually already inside a herdr pane, which is exactly where my stated rescue does not work. **And it is now closed harder than that.** #705 option 1 merged a few minutes ago (`11998cd`), narrowing that floor from `WORKER` to `OBSERVER`, and `OBSERVER` is not in the `TASK_READ` gate at all: ```java // auth/Authz.java:149 case TASK_READ -> caller.isPrimary() || caller.isWorker() || caller.isArchitect(); ``` So a herdr pane that is not a lead, member, architect or collaborator is now refused at the **role** gate, before any owner comparison. Architect B flagged that the hatch line should not be relied on long-term; it is already gone. --- ## Where the two architects agree, and where they do not **Both independently chose option 2**, both rejected a transfer primitive, and both said the same thing about sequencing: free to merge, coupled to deploy. I am not treating that agreement as a check — on #726 two architects agreed while both had a mechanism backwards — so I read their mechanisms, and the mechanisms differ in useful ways. ### They disagree on exactly one row: unit 3's single-flight claim - **A:** *breaks* — the fresh terminal is a different map key, so a second roll can start while the predecessor's continuation still holds the old key. - **B:** *survives* — it is claimed and released on the same old terminal, in `runRollover`'s `finally`, and nothing re-reads it later. **They are both right, about different properties, and neither said which property it was answering.** I read the code: the claim is taken with `putIfAbsent(p.leadTerminal(), token)` and released with the two-argument `remove(p.leadTerminal(), p.token())` — the same key both times. So B is right that **the release is correct and nothing leaks**. A is right that **it stops being a per-lead lock**: once the terminal changes mid-roll, a second `confirm()` from the fresh lead lands on a different key and is not excluded. So this is not a contradiction but an under-specification: the claim does what it was built to do (release reliably) and no longer does a thing nobody had asked it to do yet (exclude a second roll of the same *lead*). That is a real, narrow gap, and it belongs in the fix — B's design already closes it by keying single-flight on the lead identity rather than the pane. This is why the brief asked for mechanisms. Two bare verdicts here would have been a coin toss. ### B found a third break neither I nor A ranked correctly `fleet_send{turnId}` — answering a worker's `fleet_ask` — breaks through a **different** gate: ```java // msg/Rendezvous.java:91-93 (javadoc above it, verbatim) // Whether callerTerminal matches owner. A null owner means no owner was ever recorded, and // that state matches no caller, not even one whose own terminal is null — "no record" and // "recorded as the unnamed primary" are different states. public static boolean permits(Owner owner, String callerTerminal) { return owner != null && java.util.Objects.equals(owner.terminal(), callerTerminal); } ``` **B's severity argument is right and I had it wrong.** A refused poll loses a *report*. A refused answer loses *control of live work*: the worker's ask window is about 55 seconds, nothing extends it, so it resumes **unanswered** and takes whatever default it had. That can produce wrong code, not a missing message. And note what the javadoc already says: `ownsTicket` and `Owner.permits` treat `null` **differently on purpose**. Any fix touches both and must not quietly align them. The ask path's fail-closed default is deliberate and documented at the line. ### B closed the escape I was half relying on, with a mechanism I can confirm from this session B's claim: the ungated `fleet_poll{target}` inbox drain cannot rescue an in-flight delegation, because `reply()` gives the text to the open waiter and never publishes to the inbox: ```java // msg/MessageService.java:560 if (rendezvous.resolve(session, content)) { count(FleetMetrics.REPLIES, "path", "rendezvous"); return ReplyOutcome.RESOLVED_SEND; // a live send took it — unchanged fast path } // ... only when no waiter took it: // msg/MessageService.java:602 inbox.publish(session, UUID.randomUUID().toString(), content); ``` and the forward waiter stays open for `ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L` (`:56`). So for 30 minutes there is **no route**: the ticket is refused and the inbox is empty. **I can corroborate this from my own session today, and the two readings differ in the axis that matters:** | Member | Ticket state when I polled | `fleet_poll{target}` | `fleet_poll{ticket}` | |---|---|---|---| | #705 implementer | `timed_out_working` (past 30 min) | **returned the full report** | failed | | architect A | still live | **returned `[]`** | **returned the full report** | That is exactly B's mechanism, observed twice, varying ticket-expiry rather than repeating one case. It also means the stranding is real but time-bounded, as B said. --- ## The decision **Option 2, with architect B's restriction, which is not optional.** Give `Task` and `Rendezvous.Owner` a stable, **role-prefixed** owner key derived only from the connection-resolved `Principal`: - `Role.PRIMARY` with a non-null name → `"leader:" + name` - everything else unchanged → `"worker:" + terminal`, `"architect:" + slot`, `"collaborator:" + name` - unnamed primary (null name, null terminal) → `null`, so today's escape is untouched Three things I am adopting because the reasoning behind them is sound, not because an architect said them: 1. **Prefix, never a bare name.** A bare name would let an architect slot named `opus` collide with the lead named `opus`. And **key only a configured lead on its name** — an architect must keep terminal ownership, or a newly spawned architect in the same slot inherits its predecessor's tickets and open turns. That is the one way option 2 could end up worse than the bug. 2. **Do not reuse `Principal.describe()`.** It returns the constant `"primary"` for the unnamed primary, which would destroy the null escape, and its javadoc scopes it to audit lines. Add a separate `ownerKey()`. 3. **No transfer primitive at all.** Ownership is derived per call, as today; only the *address* is resolved late. There is nothing for another caller to reach, so there is nothing to guard. **The argument that actually decided it** is B's failed-roll case, and it is the one I would have got wrong on my own. Option 1 is a one-shot write that can only run *after* a new terminal exists. If the bounded relaunch exhausts its retries there is no new terminal, the roll is terminally `FAILED` with no retry, and the state stays owned by a dead terminal — unreadable by every terminal-bearing caller, permanently. Option 1 delivers nothing in exactly the case I asked about. Option 2 has no event to miss: ownership is `leader:opus` whether a pane exists or not, so the next session to hold that lead's tab label can read it. Put plainly: **after a failed roll, option 1 leaves the state owned by nobody; option 2 leaves it owned by the next session in that lead's tab.** B also noted there is no migration surface at all — `tasks` is an in-memory map and ticket ids carry a per-boot nonce, so new code can only arrive via a restart that empties it. The change that looks larger on paper has zero compatibility cost. ### Option 3 survives as a warning, not a refusal `fleet_handover{action:"open"}` should **report** how many tickets and open ask-turns the caller owns. Useful whichever fix lands. It must not refuse: a lead rolls when its context is full, which is the worst possible moment to be told it may not, and "drain first" is not achievable in bounded time when a worker is mid-brief against a 30-minute clock. ### Sequencing — and I am not accepting "deploy with the loss documented" Unit 2 **may merge** before this fix. The fix **must be in the jar before unit 2 is**. B's reason for refusing the documented-caveat option is better than mine, so I am taking it: the daemon *actively pushes the lead toward the unsafe action*. `LeadHeartbeatLoop` carries a context-high nudge, and the live `bootstrapText` is already written for a handover. So the system tells the lead to hand over precisely when its context is full — and a context-full lead is the least reliable reader of a documented caveat. **A mitigation that lives in the judgement of the agent with the least context left is not a mitigation.** That is also the shape this whole ticket came from: a safe answer recorded on #705 expired when another ticket shipped. Deploying unit 2 with a second written caveat would recreate the same failure one level up. --- ## What is NOT settled, and what nobody has measured - **B's widening argument is load-bearing and unproven, and B said so.** Option 2 means a lead's pane replaced *outside* `fleet_handover` — an operator closing the tab and opening a new one with the same label — also inherits the tickets. B argues this is acceptable because whoever can set a lead's tab label already gets `Role.PRIMARY` from `CallerResolver.java:308`, which already grants `fleet_spawn` and `fleet_stop`, so inheriting tickets is strictly less than the label already grants. **I find that convincing and I have not tested it.** The untested case B named is a mislabelled tab with no agent behind it. Whoever implements this should probe it before relying on the argument. - **Nobody has seen any of these refusals happen.** Unit 2 is not merged, no roll has run under it, and both architects ran zero tests — correctly, as neither wrote code. Every break is read from a gate plus the fact that the terminal changes. - **Neither architect read unit 2's diff**, because it does not exist yet. "The fresh lead gets a new terminal" remains inferred from #726's design and from `LeadLauncher` creating a fresh tab. If unit 2 somehow reuses the pane, this whole ticket is moot. - `creatorTerminal`'s rename will turn a wide set of tests red — `MessageService.poll(String)` alone is reported to have 44 test callers. That figure is from #705 comment 18676 and **I have not counted it myself.** - B's "degrades, self-heals" verdicts for `ReplyPushLoop` and `LeadHeartbeatLoop` depend on herdr answering `agent_not_found` for a closed pane's terminal. **That is not verified for `pane.close`** — the reported evidence is for `/exit`, a different teardown. - A's `rg` sweep did not run (`rg` is not installed in a worktree); its table came from editor Grep and Read, which did. ## Implementation units, for whoever picks this up Roughly A's split, with B's restriction folded in. Not yet briefed, and deliberately not started tonight: 1. **Identity:** add `Principal.ownerKey()`, role-prefixed, lead-name only for `PRIMARY` with a non-null name. Leave `describe()` alone. Cover named lead, architect, collaborator, unnamed primary and anonymous. 2. **Messaging:** use it in `Task`, `Rendezvous.Owner`, `poll`, `pendingAsk` and `answer`. Acceptance: lead A's old terminal creates, A's *new* terminal polls **and answers**, lead B is refused both. Keep the two null rules distinct and say in one place why. 3. **Routing:** `PrimaryRegistry`, `ReplyPushLoop` and the heartbeat keep a stable lead identity and resolve the terminal *at send time*. Acceptance: a reminder pending before the terminal changes reaches the new terminal. 4. **Rollover:** key single-flight on the lead identity, keep the old pane address separately as the thing to close. Acceptance: a failed roll leaves A's state owned by A, B refused, and a later A can continue. 5. **Warning:** `open` reports owned tickets and open asks. 6. A fake that returns the same terminal proves nothing here. Every test must actually change the terminal. The one item to carry to #726 now: unit 3's single-flight claim is keyed on the lead terminal, which is what unit 2 changes.
Author
Owner

Units 1+2 delegated as one. And one addition to the design that neither architect proposed.

Delegated to term_65d0789fa36a791 (profile sol, branch worker/737-owner-key-ff061f-10), covering the identity unit and the messaging unit together. I combined them because unit 1 alone is about thirty lines and is meaningless without unit 2.

Blast radius, measured rather than quoted

Architect B relayed a figure of 44 test callers from #705 comment 18676. I did not reproduce that number and I am not repeating it. What I measured on main just now:

What main/java test/java
creatorTerminal refs 2 files, 12 refs total incl. tests MessageServiceTest only
Owner.of / Owner.permits / Rendezvous.Owner 21 refs across both —
sendAsync( call sites 16 (FleetMcp 5, MessageService 10, FleetApp 1) 108, of which 90 are in MessageServiceTest.java

So the production change is genuinely small and the churn is large but mechanical and concentrated in a single test file. That is what made one unit the right size rather than two.

The addition: ANONYMOUS gets its own key, and that pins the open #705 coupling

Both architects said to leave the null escape alone, and both separately flagged that nothing pins the coupling that makes it safe — namely that Authz denies ANONYMOUS the TASK_READ action, so an anonymous caller never reaches ownsTicket. #705 comments 18617 and 18676 both record it as unpinned.

I think that is a null doing two opposite jobs, and the fix costs one table row. Today null means both:

  • "the unnamed primary" → may read every ticket
  • "this caller could not be identified" → may read nothing

One symbol, two states needing opposite handling. So ownerKey() returns null only for the unnamed primary, and "anonymous" for ANONYMOUS — a value that matches no stored owner, because no anonymous caller can ever create a ticket.

The gate then fails closed on its own. It no longer depends on Authz refusing anonymous first, which means the thing #705 recorded as unpinned stops needing a pin: it is no longer load-bearing. The brief requires a test that proves the refusal without relying on Authz, and a mutation (anonymous returns null again) that must kill it.

This is a small widening of the unit's scope beyond what the architects proposed. I am recording it here rather than only in the brief, because it changes a safety property and the next reader should see the reasoning, not just the code.

What the brief withholds, deliberately

The worker is told not to touch PrimaryRegistry, ReplyPushLoop, LeadHeartbeatLoop, LeadRollover's single-flight key, or Authz's role table — those are the routing and rollover units. LeadRollover especially, because #726 unit 2 is editing that file right now.

It is also told there is no transfer primitive, with the instruction to stop and re-read the decision if it finds itself adding a setter.

Required acceptance, so it cannot be faked green

The headline test must use a terminal that actually differs across the roll. A fake returning the same terminal would pass while proving nothing, and that is the single most likely way this unit comes back green and wrong. Six mutations are required, each with a named test that must die, including "a lead is keyed on its terminal again" and "an architect is keyed on its slot name".

Remaining units, not yet briefed

  1. Routing — PrimaryRegistry, ReplyPushLoop, the heartbeat: hold a stable lead identity, resolve the terminal at send time.
  2. Rollover — key single-flight on the lead identity, keep the old pane address separately as the thing to close. Must land after #726 unit 2 merges, to avoid colliding in LeadRollover.
  3. Warning — fleet_handover{open} reports owned tickets and open asks.

Deployment gate, restated

The running daemon is on a jar that predates all of today's merges. I am not redeploying until this fix and #726 unit 2 are both in the tree, for the reason in the decision comment: the daemon's own context-high nudge pushes a lead toward a handover exactly when its context is full, so a documented caveat is not a mitigation.

## Units 1+2 delegated as one. And one addition to the design that neither architect proposed. Delegated to `term_65d0789fa36a791` (profile `sol`, branch `worker/737-owner-key-ff061f-10`), covering the identity unit and the messaging unit together. I combined them because unit 1 alone is about thirty lines and is meaningless without unit 2. ### Blast radius, measured rather than quoted Architect B relayed a figure of 44 test callers from #705 comment 18676. I did not reproduce that number and I am not repeating it. What I measured on `main` just now: | What | main/java | test/java | |---|---|---| | `creatorTerminal` refs | 2 files, 12 refs total incl. tests | `MessageServiceTest` only | | `Owner.of` / `Owner.permits` / `Rendezvous.Owner` | 21 refs across both | — | | `sendAsync(` call sites | **16** (`FleetMcp` 5, `MessageService` 10, `FleetApp` 1) | **108**, of which **90 are in `MessageServiceTest.java`** | So the production change is genuinely small and the churn is large but mechanical and concentrated in a single test file. That is what made one unit the right size rather than two. ### The addition: `ANONYMOUS` gets its own key, and that pins the open #705 coupling Both architects said to leave the `null` escape alone, and both separately flagged that nothing pins the coupling that makes it safe — namely that `Authz` denies `ANONYMOUS` the `TASK_READ` action, so an anonymous caller never reaches `ownsTicket`. #705 comments 18617 and 18676 both record it as unpinned. **I think that is a `null` doing two opposite jobs, and the fix costs one table row.** Today `null` means both: - "the unnamed primary" → **may read every ticket** - "this caller could not be identified" → **may read nothing** One symbol, two states needing opposite handling. So `ownerKey()` returns `null` **only** for the unnamed primary, and `"anonymous"` for `ANONYMOUS` — a value that matches no stored owner, because no anonymous caller can ever create a ticket. The gate then fails closed on its own. It no longer depends on `Authz` refusing anonymous first, which means the thing #705 recorded as unpinned stops needing a pin: it is no longer load-bearing. The brief requires a test that proves the refusal **without** relying on `Authz`, and a mutation (anonymous returns `null` again) that must kill it. This is a small widening of the unit's scope beyond what the architects proposed. I am recording it here rather than only in the brief, because it changes a safety property and the next reader should see the reasoning, not just the code. ### What the brief withholds, deliberately The worker is told **not** to touch `PrimaryRegistry`, `ReplyPushLoop`, `LeadHeartbeatLoop`, `LeadRollover`'s single-flight key, or `Authz`'s role table — those are the routing and rollover units. `LeadRollover` especially, because #726 unit 2 is editing that file right now. It is also told there is **no transfer primitive**, with the instruction to stop and re-read the decision if it finds itself adding a setter. ### Required acceptance, so it cannot be faked green The headline test must use a terminal that **actually differs** across the roll. A fake returning the same terminal would pass while proving nothing, and that is the single most likely way this unit comes back green and wrong. Six mutations are required, each with a named test that must die, including "a lead is keyed on its terminal again" and "an architect is keyed on its slot name". ### Remaining units, not yet briefed 3. **Routing** — `PrimaryRegistry`, `ReplyPushLoop`, the heartbeat: hold a stable lead identity, resolve the terminal at send time. 4. **Rollover** — key single-flight on the lead identity, keep the old pane address separately as the thing to close. Must land after #726 unit 2 merges, to avoid colliding in `LeadRollover`. 5. **Warning** — `fleet_handover{open}` reports owned tickets and open asks. ### Deployment gate, restated The running daemon is on a jar that predates all of today's merges. **I am not redeploying until this fix and #726 unit 2 are both in the tree**, for the reason in the decision comment: the daemon's own context-high nudge pushes a lead toward a handover exactly when its context is full, so a documented caveat is not a mitigation.
Author
Owner

Unit 3 (routing) is read and specified, but NOT delegated yet. It would collide with #726 unit 2.

I read the routing path rather than briefing it from the architects' summary, and it is both narrower and sharper than "these classes hold stale lead addresses". One finding changes the unit, and one changes the schedule.

The schedule first: do not delegate this until #726 unit 2 merges

The fix needs the live-lead map injected into the registry or the loops, and that wiring lives in Fleetd.java and FleetdAssembly.java. #726 unit 2 is editing both of those files right now — it rewires Fleetd.leadRollover and hoists the LeadLauncher at FleetdAssembly.java:304. Two workers in those files is a guaranteed conflict for no gain, so unit 3 waits. Nothing else in the #737 list is disjoint from it either, so unit 3 is next in line after unit 2 lands.

What leadByTarget does on a roll: it already self-heals

ReplyPushLoop.resolveLiveLead (:443-452) already probes the recorded lead before trusting it, and forgets the binding when herdr affirmatively answers agent_not_found (fleetd #368). After a roll the old pane is genuinely closed, so a worker's stale binding is dropped on the first tick and resolution retries. That half is fine and must not be touched. isLive narrowing death to the one affirmative code is also correct and is exactly the care #359 asked for — leave it alone.

The actual defect: the fallback is returned without being probed

private Optional<String> resolveLiveLead(String target) {
    Optional<String> lead = primaryRegistry.nudgeTargetFor(target);
    if (lead.isEmpty() || isLive(lead.get())) {
        return lead;
    }
    primaryRegistry.forgetDelegation(target);
    return primaryRegistry.nudgeTargetFor(target);   // never probed
}

The second nudgeTargetFor falls through to PrimaryRegistry's single terminal slot. That slot holds the dead lead's terminal too, and it is returned with no liveness check at all. The method is named resolveLiveLead and its second return is not a live lead.

So the sequence on a roll is: drop the stale per-worker binding correctly, then hand back a different stale address unchecked.

Why the single slot does not recover on its own

recordPrimarySingleton (FleetMcp.java:728) is reached from exactly two call sites, FleetMcp.java:467 and :546 — the send and spawn handlers. Measured with grep -n 'recordPrimarySingleton\|registry.record(' mcp/FleetMcp.java. So the slot is refreshed only when a lead sends or spawns.

A freshly bootstrapped lead reads its handover file first. Until it calls fleet_send or fleet_spawn, the slot still names the pane that was closed. During that window:

  • a worker's reply nudges a dead terminal, so the new lead is never told a reply arrived
  • LeadHeartbeatLoop (:310, :315, :370) reads primaryTerminal() for the context-pressure notice and aims it at the same dead pane

Neither loses data — the durable inbox still holds the reply and fleet_poll still works. What is lost is every push notification, which is the whole point of the feature. The new lead has to already suspect there is something to collect.

The shape of the fix

The daemon already knows which terminals are live leads: LeadTabScanner, the same Supplier<Map<String, String>> of terminal to name that Fleetd.leadRollover is handed. The single slot is a value learned from call traffic while an authoritative source sits right next to it. So:

  1. Record the delegating lead's name alongside its terminal when the name resolves, and resolve name to current terminal at nudge time. A name survives a roll; a terminal does not.
  2. Keep the learned terminal as the fallback for a primary the scanner cannot see — an unnamed primary, off-host, or non-herdr. That path is real and must keep working.
  3. Probe the fallback. Whatever resolveLiveLead returns must have passed isLive, or the method is misnamed.
  4. Have LeadHeartbeatLoop ask for the current terminal rather than the learned one.

Scope limits for whoever takes this

  • Do not touch ownerKey or MessageService. Unit 1+2 is building a stable identity for comparison — "may this caller read this ticket". This unit needs resolution — "which pane do I nudge now". Only a lead can be rolled, so only the lead case needs the indirection, and reusing the comparison key here would drag an architect and collaborator case into a question that does not have one. Two deliberately separate notions; say so in the javadoc so the next reader does not merge them.
  • Do not weaken isLive. Any RuntimeException other than agent_not_found must keep meaning "still live".
  • The required mutation: return the fallback unprobed again, and a test must die. That is the defect, so that is the mutation that proves the fix.

One thing I have not checked

I have not run a roll with a worker mid-delegation and watched where the nudge went. Everything above is read from the code, not observed. The window is real on the code's own terms, but I am not reporting a measured incident.

## Unit 3 (routing) is read and specified, but NOT delegated yet. It would collide with #726 unit 2. I read the routing path rather than briefing it from the architects' summary, and it is both narrower and sharper than "these classes hold stale lead addresses". One finding changes the unit, and one changes the schedule. ### The schedule first: do not delegate this until #726 unit 2 merges The fix needs the live-lead map injected into the registry or the loops, and that wiring lives in `Fleetd.java` and `FleetdAssembly.java`. **#726 unit 2 is editing both of those files right now** — it rewires `Fleetd.leadRollover` and hoists the `LeadLauncher` at `FleetdAssembly.java:304`. Two workers in those files is a guaranteed conflict for no gain, so unit 3 waits. Nothing else in the #737 list is disjoint from it either, so unit 3 is next in line after unit 2 lands. ### What `leadByTarget` does on a roll: it already self-heals `ReplyPushLoop.resolveLiveLead` (`:443-452`) already probes the recorded lead before trusting it, and forgets the binding when herdr affirmatively answers `agent_not_found` (fleetd #368). After a roll the old pane is genuinely closed, so a worker's stale binding is dropped on the first tick and resolution retries. **That half is fine and must not be touched.** `isLive` narrowing death to the one affirmative code is also correct and is exactly the care #359 asked for — leave it alone. ### The actual defect: the fallback is returned without being probed ```java private Optional<String> resolveLiveLead(String target) { Optional<String> lead = primaryRegistry.nudgeTargetFor(target); if (lead.isEmpty() || isLive(lead.get())) { return lead; } primaryRegistry.forgetDelegation(target); return primaryRegistry.nudgeTargetFor(target); // never probed } ``` The second `nudgeTargetFor` falls through to `PrimaryRegistry`'s single `terminal` slot. **That slot holds the dead lead's terminal too**, and it is returned with no liveness check at all. The method is named `resolveLiveLead` and its second return is not a live lead. So the sequence on a roll is: drop the stale per-worker binding correctly, then hand back a different stale address unchecked. ### Why the single slot does not recover on its own `recordPrimarySingleton` (`FleetMcp.java:728`) is reached from exactly **two** call sites, `FleetMcp.java:467` and `:546` — the send and spawn handlers. Measured with `grep -n 'recordPrimarySingleton\|registry.record(' mcp/FleetMcp.java`. So the slot is refreshed only when a lead sends or spawns. A freshly bootstrapped lead reads its handover file first. Until it calls `fleet_send` or `fleet_spawn`, the slot still names the pane that was closed. During that window: - a worker's reply nudges a dead terminal, so the new lead is never told a reply arrived - `LeadHeartbeatLoop` (`:310`, `:315`, `:370`) reads `primaryTerminal()` for the context-pressure notice and aims it at the same dead pane Neither loses data — the durable inbox still holds the reply and `fleet_poll` still works. What is lost is every push notification, which is the whole point of the feature. The new lead has to already suspect there is something to collect. ### The shape of the fix The daemon already knows which terminals are live leads: `LeadTabScanner`, the same `Supplier<Map<String, String>>` of terminal to name that `Fleetd.leadRollover` is handed. **The single slot is a value learned from call traffic while an authoritative source sits right next to it.** So: 1. Record the delegating lead's **name** alongside its terminal when the name resolves, and resolve name to current terminal at nudge time. A name survives a roll; a terminal does not. 2. Keep the learned terminal as the fallback for a primary the scanner cannot see — an unnamed primary, off-host, or non-herdr. That path is real and must keep working. 3. **Probe the fallback.** Whatever `resolveLiveLead` returns must have passed `isLive`, or the method is misnamed. 4. Have `LeadHeartbeatLoop` ask for the current terminal rather than the learned one. ### Scope limits for whoever takes this - **Do not touch `ownerKey` or `MessageService`.** Unit 1+2 is building a stable identity for *comparison* — "may this caller read this ticket". This unit needs *resolution* — "which pane do I nudge now". Only a lead can be rolled, so only the lead case needs the indirection, and reusing the comparison key here would drag an architect and collaborator case into a question that does not have one. Two deliberately separate notions; say so in the javadoc so the next reader does not merge them. - **Do not weaken `isLive`.** Any `RuntimeException` other than `agent_not_found` must keep meaning "still live". - The required mutation: **return the fallback unprobed again, and a test must die.** That is the defect, so that is the mutation that proves the fix. ### One thing I have not checked I have not run a roll with a worker mid-delegation and watched where the nudge went. Everything above is read from the code, not observed. The window is real on the code's own terms, but I am not reporting a measured incident.
Author
Owner

Units 1+2: the member died on a usage limit — the work is safe, and here is where it is (2026-10-04)

Read this before you pick units 1+2 up. The work is nearly done and it is not committed. Do not start it again from nothing.

What happened

The member on profile sol stopped mid-task. Its pane went quiet and the bridge recorded state: done, which is a turn boundary and not a finish. The real cause came from the ticket poll:

fleet_poll{ticket: task-39745b-14}
→ failed — backend exhausted (usage limit)

fleet_profiles then showed why it takes two profiles with it:

"quarantined": {
  "sol":   {"credentialId": "openai-shared", "quarantinedForSeconds": 1745, "model": "openai/gpt-5.6-sol"},
  "terra": {"credentialId": "openai-shared", "quarantinedForSeconds": 1745, "model": "openai/gpt-5.6-terra"}
}

sol and terra share the credential openai-shared, so one exhaustion quarantined both for about 29 minutes. I stopped the pane, because an exhausted credential cannot do any more work and a live pane could write into the tree while I was building it.

Where the work is

Branch worker/737-owner-key-ff061f-10, in the worktree /Users/dai.ha/LTMS/.bridged-worktrees/ac9bb2-10.

  • 0 commits. Everything is uncommitted working-tree change.
  • 11 files changed, 259 insertions, 179 deletions, plus one new file fleetd/src/test/java/dev/ltms/fleet/auth/PrincipalTest.java (31 lines, untracked).
  • Main code touched: auth/Principal.java, mcp/FleetMcp.java, msg/MessageService.java, msg/Rendezvous.java, rest/FleetApp.java.
  • The worktree HEAD is e3050ef. origin/main is now aabecce, so the branch is four commits behind.

The resumable member id, if anyone wants to continue that conversation once the quarantine lifts: ses_ef7ec1b66ffesitLBWRn0Kk5Er. It needs profile sol again — a resume cannot cross backends.

The core of it is already right

Principal.ownerKey() is implemented, and it handles the ANONYMOUS case this ticket asked for:

public String ownerKey() {
    return switch (role) {
        case PRIMARY -> name == null ? null : prefixed("leader", name);
        case WORKER -> prefixed("worker", terminal);
        case ARCHITECT -> prefixed("architect", terminal);
        case COLLABORATOR -> prefixed("collaborator", name);
        case OBSERVER -> prefixed("observer", terminal);
        case ANONYMOUS -> "anonymous";
    };
}

The unnamed primary returning null is deliberate and matches the design on this ticket: it keeps the message layer's primary-wide ticket rule for a caller that carries no name.

I have not reviewed the other ten files yet, and I am not calling this correct. I am recording that it exists and looks coherent.

What still has to happen

  1. Merge origin/main into the branch first. The branch is four commits behind. Those four commits touch session/SessionManager.java, .claude/skills/handover/SKILL.md and the wiki pointer — none of the five main files above — so no conflict is expected. Expected is not measured: do the merge, then build.
  2. A build of the merge, not of the branch. A branch's own green build does not prove the merge compiles when the branch is behind.
  3. The mutation evidence this ticket asks for. None of it has been produced yet.
  4. Commit, push, open the PR.

One hygiene note for whoever reuses that worktree

Delete fleetd/target/surefire-reports before trusting a test count. A reused worktree keeps XML from earlier runs, including for test classes that no longer exist, so the count reads high. I cleared it before my own build.

An aside that cost us this turn

exhaustionDetectionArmed is true for only sol and terra:

{"opus": false, "sonnet": false, "local-direct": false, "local": false,
 "sol": true, "xf": false, "terra": true, "gx": false}

So a usage limit on any other profile can never be classified or quarantined, however many times it happens. This loss was detected because it landed on an armed profile. The same failure on sonnet would surface as an opaque dead member, and a lead would likely respawn straight back into it. That is a separate concern from this ticket and I am not fixing it here — noting it so it is written down somewhere.

## Units 1+2: the member died on a usage limit — the work is safe, and here is where it is (2026-10-04) **Read this before you pick units 1+2 up.** The work is nearly done and it is **not committed**. Do not start it again from nothing. ### What happened The member on profile `sol` stopped mid-task. Its pane went quiet and the bridge recorded `state: done`, which is a turn boundary and not a finish. The real cause came from the ticket poll: ``` fleet_poll{ticket: task-39745b-14} → failed — backend exhausted (usage limit) ``` `fleet_profiles` then showed why it takes two profiles with it: ```json "quarantined": { "sol": {"credentialId": "openai-shared", "quarantinedForSeconds": 1745, "model": "openai/gpt-5.6-sol"}, "terra": {"credentialId": "openai-shared", "quarantinedForSeconds": 1745, "model": "openai/gpt-5.6-terra"} } ``` `sol` and `terra` share the credential `openai-shared`, so one exhaustion quarantined both for about 29 minutes. I stopped the pane, because an exhausted credential cannot do any more work and a live pane could write into the tree while I was building it. ### Where the work is Branch `worker/737-owner-key-ff061f-10`, in the worktree `/Users/dai.ha/LTMS/.bridged-worktrees/ac9bb2-10`. - **0 commits.** Everything is uncommitted working-tree change. - **11 files changed, 259 insertions, 179 deletions**, plus one new file `fleetd/src/test/java/dev/ltms/fleet/auth/PrincipalTest.java` (31 lines, untracked). - Main code touched: `auth/Principal.java`, `mcp/FleetMcp.java`, `msg/MessageService.java`, `msg/Rendezvous.java`, `rest/FleetApp.java`. - The worktree HEAD is `e3050ef`. `origin/main` is now `aabecce`, so **the branch is four commits behind**. The resumable member id, if anyone wants to continue that conversation once the quarantine lifts: `ses_ef7ec1b66ffesitLBWRn0Kk5Er`. It needs profile `sol` again — a resume cannot cross backends. ### The core of it is already right `Principal.ownerKey()` is implemented, and it handles the `ANONYMOUS` case this ticket asked for: ```java public String ownerKey() { return switch (role) { case PRIMARY -> name == null ? null : prefixed("leader", name); case WORKER -> prefixed("worker", terminal); case ARCHITECT -> prefixed("architect", terminal); case COLLABORATOR -> prefixed("collaborator", name); case OBSERVER -> prefixed("observer", terminal); case ANONYMOUS -> "anonymous"; }; } ``` The unnamed primary returning `null` is deliberate and matches the design on this ticket: it keeps the message layer's primary-wide ticket rule for a caller that carries no name. I have **not** reviewed the other ten files yet, and I am not calling this correct. I am recording that it exists and looks coherent. ### What still has to happen 1. **Merge `origin/main` into the branch first.** The branch is four commits behind. Those four commits touch `session/SessionManager.java`, `.claude/skills/handover/SKILL.md` and the wiki pointer — none of the five main files above — so no conflict is expected. Expected is not measured: do the merge, then build. 2. **A build of the merge, not of the branch.** A branch's own green build does not prove the merge compiles when the branch is behind. 3. **The mutation evidence this ticket asks for.** None of it has been produced yet. 4. Commit, push, open the PR. ### One hygiene note for whoever reuses that worktree Delete `fleetd/target/surefire-reports` before trusting a test count. A reused worktree keeps XML from earlier runs, including for test classes that no longer exist, so the count reads high. I cleared it before my own build. ### An aside that cost us this turn `exhaustionDetectionArmed` is `true` for only `sol` and `terra`: ```json {"opus": false, "sonnet": false, "local-direct": false, "local": false, "sol": true, "xf": false, "terra": true, "gx": false} ``` So a usage limit on any other profile can never be classified or quarantined, however many times it happens. This loss was detected because it landed on an armed profile. The same failure on `sonnet` would surface as an opaque dead member, and a lead would likely respawn straight back into it. That is a separate concern from this ticket and I am not fixing it here — noting it so it is written down somewhere.
Author
Owner

Lead review of units 1+2's main-code diff (2026-10-04)

I read all five main files myself. I said earlier I had not, so here is the result. No blocker. One finding worth carrying, and two things done right that are easy to get wrong.

Base for this review: the branch fast-forwarded to aabecce, mvn clean install green — 2094 tests, 0 failures, 0 errors, 0 skipped, across 177 classes. That is main's 2089 plus the 5 in the new PrincipalTest, and the class count is main's 176 plus one. It reconciles from both ends.

Done right: the three null states stay three

This is the part I expected to be wrong, and it is not.

// Rendezvous.Owner
public static Owner of(String ownerKey) {
    return ownerKey == null ? UNNAMED_PRIMARY : new Owner(ownerKey);
}
public static boolean permits(Owner owner, String callerOwner) {
    return owner != null && Objects.equals(owner.ownerKey(), callerOwner);
}
// MessageService
private static boolean ownsTicket(Task task, String callerOwner) {
    return callerOwner == null || callerOwner.equals(task.creatorOwner);
}

Three states, not two, and they are kept apart:

state meaning result
Owner is null no owner was ever recorded matches nobody
Owner is UNNAMED_PRIMARY (key null) recorded as the unnamed primary matches the unnamed primary
caller key is null at ownsTicket the unnamed primary polling may read every ticket

So permits and ownsTicket treat null in opposite directions on purpose, and the diff says so in the javadoc. That is correct, and it is the distinction a single sentinel usually loses.

ANONYMOUS -> "anonymous" is also the right call. A non-null key means an anonymous caller can never fall into the null-means-unnamed-primary rule, even though Authz already refuses it earlier. Two gates, same answer.

Done right: the primary's key is its NAME

case PRIMARY -> name == null ? null : prefixed("leader", name);

This is the whole unit. The key survives a handover because the name does and the terminal does not. COLLABORATOR uses its name for the same reason. WORKER, ARCHITECT and OBSERVER use the terminal, which is right — they have no configured name, and none of them outlives its pane.

The finding: the type-level guard is on 1 of 6 entry points

sendAsync takes a Principal and derives the key inside, and its javadoc gives the reason:

The key is derived here from the resolved principal so callers cannot pass a terminal address where an owner identity is required.

That reason is good. It is just not applied anywhere else:

944:  public Reply send(String target, String content, long timeoutMillis, String callerOwner)
960:  public Reply send(..., Runnable onAccepted, String callerOwner)
1173: public Reply answer(String turnId, String content, long timeoutMillis, String callerOwner)
1328: public String sendAsync(String target, String content, Runnable onAccepted, Principal creator)   <-- guarded
1410: public TaskView poll(String ticket, String callerOwner)
1790: public PendingAsk pendingAsk(String workerSession, String callerOwner)

Five of six take a bare String, so a future call site writing caller.terminal() instead of caller.ownerKey() compiles. The javadoc on the guarded one reads as if the hazard is handled, which makes the other five easier to get wrong, not harder. Same shape as a grant applied at one gate of two.

Why this is not a blocker: every confusion direction fails closed. I worked through each one:

  • poll(ticket, caller.terminal()) — a terminal compared against a stored leader:opus never matches, so the caller is refused its own ticket. Loud.
  • answer(turnId, ..., caller.terminal()) — same, permits returns false and the turn is not resolved. Loud.
  • a terminal recorded as a ticket's owner — the ticket becomes readable by nobody but the unnamed primary. Loud, and it strands one ticket.

None of them widens access. That is the safe direction and it is why I am not holding the PR for it.

What I suggest, and am not asking the current member to do (its brief says change no behaviour): take the Principal overload pattern to send, answer, poll and pendingAsk as a follow-up, or drop the claim from sendAsync's javadoc so it does not describe a protection the API mostly lacks. One fact, one place — right now the sentence is true of one method and reads as true of the layer.

Correctly left alone

recordPrimarySingleton and primaryRegistry.recordDelegation still take the terminal:

recordPrimarySingleton(primaryRegistry, callerTerminal, caller);
Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, callerTerminal);

That is right for this unit. The registry's job is nudge routing to a live pane, and a pane is addressed by terminal. Making it owner-keyed is unit 3's problem, not a miss here.

Still outstanding before I merge

The mutation evidence. It has not been produced, and a green 2094 proves the code runs, not that the tests would notice it breaking. A member is on that now.

## Lead review of units 1+2's main-code diff (2026-10-04) I read all five main files myself. I said earlier I had not, so here is the result. **No blocker.** One finding worth carrying, and two things done right that are easy to get wrong. Base for this review: the branch fast-forwarded to `aabecce`, `mvn clean install` green — 2094 tests, 0 failures, 0 errors, 0 skipped, across 177 classes. That is `main`'s 2089 plus the 5 in the new `PrincipalTest`, and the class count is `main`'s 176 plus one. It reconciles from both ends. ### Done right: the three null states stay three This is the part I expected to be wrong, and it is not. ```java // Rendezvous.Owner public static Owner of(String ownerKey) { return ownerKey == null ? UNNAMED_PRIMARY : new Owner(ownerKey); } public static boolean permits(Owner owner, String callerOwner) { return owner != null && Objects.equals(owner.ownerKey(), callerOwner); } ``` ```java // MessageService private static boolean ownsTicket(Task task, String callerOwner) { return callerOwner == null || callerOwner.equals(task.creatorOwner); } ``` Three states, not two, and they are kept apart: | state | meaning | result | |---|---|---| | `Owner` is `null` | no owner was ever recorded | matches nobody | | `Owner` is `UNNAMED_PRIMARY` (key `null`) | recorded as the unnamed primary | matches the unnamed primary | | caller key is `null` at `ownsTicket` | the unnamed primary polling | may read every ticket | So `permits` and `ownsTicket` treat `null` in **opposite** directions on purpose, and the diff says so in the javadoc. That is correct, and it is the distinction a single sentinel usually loses. `ANONYMOUS -> "anonymous"` is also the right call. A non-null key means an anonymous caller can never fall into the null-means-unnamed-primary rule, even though `Authz` already refuses it earlier. Two gates, same answer. ### Done right: the primary's key is its NAME ```java case PRIMARY -> name == null ? null : prefixed("leader", name); ``` This is the whole unit. The key survives a handover because the name does and the terminal does not. `COLLABORATOR` uses its name for the same reason. `WORKER`, `ARCHITECT` and `OBSERVER` use the terminal, which is right — they have no configured name, and none of them outlives its pane. ### The finding: the type-level guard is on 1 of 6 entry points `sendAsync` takes a `Principal` and derives the key inside, and its javadoc gives the reason: > The key is derived here from the resolved principal so callers cannot pass a terminal address where an owner identity is required. That reason is good. It is just not applied anywhere else: ``` 944: public Reply send(String target, String content, long timeoutMillis, String callerOwner) 960: public Reply send(..., Runnable onAccepted, String callerOwner) 1173: public Reply answer(String turnId, String content, long timeoutMillis, String callerOwner) 1328: public String sendAsync(String target, String content, Runnable onAccepted, Principal creator) <-- guarded 1410: public TaskView poll(String ticket, String callerOwner) 1790: public PendingAsk pendingAsk(String workerSession, String callerOwner) ``` Five of six take a bare `String`, so a future call site writing `caller.terminal()` instead of `caller.ownerKey()` compiles. The javadoc on the guarded one reads as if the hazard is handled, which makes the other five easier to get wrong, not harder. Same shape as a grant applied at one gate of two. **Why this is not a blocker: every confusion direction fails closed.** I worked through each one: - `poll(ticket, caller.terminal())` — a terminal compared against a stored `leader:opus` never matches, so the caller is refused its own ticket. Loud. - `answer(turnId, ..., caller.terminal())` — same, `permits` returns false and the turn is not resolved. Loud. - a terminal recorded as a ticket's owner — the ticket becomes readable by nobody but the unnamed primary. Loud, and it strands one ticket. None of them widens access. That is the safe direction and it is why I am not holding the PR for it. **What I suggest, and am not asking the current member to do** (its brief says change no behaviour): take the `Principal` overload pattern to `send`, `answer`, `poll` and `pendingAsk` as a follow-up, or drop the claim from `sendAsync`'s javadoc so it does not describe a protection the API mostly lacks. One fact, one place — right now the sentence is true of one method and reads as true of the layer. ### Correctly left alone `recordPrimarySingleton` and `primaryRegistry.recordDelegation` still take the terminal: ```java recordPrimarySingleton(primaryRegistry, callerTerminal, caller); Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, callerTerminal); ``` That is right for this unit. The registry's job is nudge routing to a live pane, and a pane is addressed by terminal. Making it owner-keyed is unit 3's problem, not a miss here. ### Still outstanding before I merge The mutation evidence. It has not been produced, and a green 2094 proves the code runs, not that the tests would notice it breaking. A member is on that now.
Author
Owner

Review of PR #741: one finding, adjudicated — not a blocker, but it names a real gap

A reviewer looked at the ownership semantics of PR #741. It returned one finding, high: ownsTicket treats a null caller key as a wildcard, so a caller with no key reads every ticket.

I checked the claim myself. Here is what is true, what is not, and what I am doing about it.

The finding is real, but it is NOT introduced by this PR

The reviewer quoted MessageService.java:1463 and the name callerTerminal. Those are main's line and main's name, so it reviewed main, not the diff. The behaviour it describes is the same on both sides:

$ git show origin/main:.../MessageService.java | grep -n "callerTerminal == null"
1464:        return callerTerminal == null || callerTerminal.equals(task.creatorTerminal);

$ git show origin/worker/737-owner-key-ff061f-10:.../MessageService.java | grep -n "callerOwner == null"
1461:        return callerOwner == null || callerOwner.equals(task.creatorOwner);

So the PR changes what the key is (terminal → owner key) and keeps the null rule exactly as it was. The PR's own javadoc states the rule on purpose: "A null caller key is the unnamed primary and may read every ticket."

So this is not a regression and does not block the merge. The fix the reviewer asks for — "replace terminal-based ownership with a stable owner key derived from role+identity" — is what this PR already did. Only the null treatment is left.

One part of the finding is wrong

The reviewer wrote that an "unauthenticated caller" can poll tickets. It cannot. Authz.permits refuses before poll is ever reached:

if (caller == null || caller.isAnonymous()) {
    return false; // authenticated as nothing ⇒ authorized for nothing
}

Both surfaces gate first: FleetMcp.java:1246 sits behind the TASK_READ check, and FleetApp.java:903 calls allow(...) before :907. An anonymous caller also does not even get a null key — Principal.ownerKey() maps ANONYMOUS to the string "anonymous".

What the gap actually is, measured

In production, a null owner key has exactly one source. I checked every .poll( in main sources on the branch:

$ git grep -n "\.poll(" origin/worker/737-owner-key-ff061f-10 -- 'fleetd/src/main/java'
.../mcp/FleetMcp.java:1246:        MessageService.TaskView v = messages.poll(ticket, callerOwner);
.../rest/FleetApp.java:907:        ... messages.poll(ctx.pathParam("ticket"), caller == null ? null : caller.ownerKey());
.../inject/Injector.java:493:                            t.queue.poll();     // a queue, not this class

The single-argument poll(String ticket) overload has no production caller. It is a test seam. So the only live route to a null key is Principal.ownerKey() returning null, and that happens for one case only:

case PRIMARY -> name == null ? null : prefixed("leader", name);

That is the unnamed primary: a caller the daemon resolved as primary with no pane name — token auth, or loopback trust. Every other role returns a non-null key.

So the gap is: an unnamed primary reads every lead's ticket, including the full reply text. The live-turn path refuses the same caller — Rendezvous.Owner.permits needs an exact match and never treats null as a wildcard. Async tickets are therefore less protected than live turns, which is the asymmetry the reviewer spotted. That part stands.

The shape of it: one null, two meanings

The null key carries two different states that need opposite handling:

who passes null what it means what it should get
an internal caller, through poll(String) "identity is not relevant here" no check
an unnamed primary, through ownerKey() "a real caller with no name" the same check as everyone else

Today both get "no check". This is the same conflated-sentinel shape as #497: one symbol, two states, one handling. The fix is a third state — keep the internal bypass explicit, and stop the unnamed primary sharing its value.

How this relates to #705

#705 offers three fix shapes for ticket walking. Its option 2 is "give tickets an owner check", and that is what this ticket is. This finding is the hole left inside option 2: the owner check now exists, and one caller still skips it. #705 can close on option 1 or 2 without this being fixed, so I am keeping the two separate rather than merging the tickets.

Decision

  • PR #741 is not blocked by this. It fixes the "too closed" side it was written for and leaves the null rule untouched and documented.
  • The null wildcard becomes #737 unit 6, specified below with the rest of the units.
  • I am not changing it inside #741. Widening the scope of a merged-ready PR to cover a pre-existing hole is how a verified diff stops being verified.

Severity, stated plainly: lower than the reviewer's high. An unnamed primary already holds SPAWN, STOP, SEND and DRAIN over every session on this host, so it can reach most of the same content by other routes. It is still worth closing, because the live-turn path already refuses it and the two paths should agree.

I checked the code in this comment myself, on origin/main at aabecce and on origin/worker/737-owner-key-ff061f-10. I did not run a live probe from an unnamed primary.

## Review of PR #741: one finding, adjudicated — not a blocker, but it names a real gap A reviewer looked at the ownership semantics of PR #741. It returned one finding, `high`: `ownsTicket` treats a `null` caller key as a wildcard, so a caller with no key reads every ticket. I checked the claim myself. Here is what is true, what is not, and what I am doing about it. ### The finding is real, but it is NOT introduced by this PR The reviewer quoted `MessageService.java:1463` and the name `callerTerminal`. Those are `main`'s line and `main`'s name, so it reviewed `main`, not the diff. The behaviour it describes is the same on both sides: ``` $ git show origin/main:.../MessageService.java | grep -n "callerTerminal == null" 1464: return callerTerminal == null || callerTerminal.equals(task.creatorTerminal); $ git show origin/worker/737-owner-key-ff061f-10:.../MessageService.java | grep -n "callerOwner == null" 1461: return callerOwner == null || callerOwner.equals(task.creatorOwner); ``` So the PR changes **what** the key is (terminal → owner key) and keeps the `null` rule exactly as it was. The PR's own javadoc states the rule on purpose: *"A `null` caller key is the unnamed primary and may read every ticket."* **So this is not a regression and does not block the merge.** The fix the reviewer asks for — "replace terminal-based ownership with a stable owner key derived from role+identity" — is what this PR already did. Only the `null` treatment is left. ### One part of the finding is wrong The reviewer wrote that an *"unauthenticated caller"* can poll tickets. It cannot. `Authz.permits` refuses before `poll` is ever reached: ```java if (caller == null || caller.isAnonymous()) { return false; // authenticated as nothing ⇒ authorized for nothing } ``` Both surfaces gate first: `FleetMcp.java:1246` sits behind the `TASK_READ` check, and `FleetApp.java:903` calls `allow(...)` before `:907`. An anonymous caller also does not even get a `null` key — `Principal.ownerKey()` maps `ANONYMOUS` to the string `"anonymous"`. ### What the gap actually is, measured In production, a `null` owner key has exactly one source. I checked every `.poll(` in main sources on the branch: ``` $ git grep -n "\.poll(" origin/worker/737-owner-key-ff061f-10 -- 'fleetd/src/main/java' .../mcp/FleetMcp.java:1246: MessageService.TaskView v = messages.poll(ticket, callerOwner); .../rest/FleetApp.java:907: ... messages.poll(ctx.pathParam("ticket"), caller == null ? null : caller.ownerKey()); .../inject/Injector.java:493: t.queue.poll(); // a queue, not this class ``` The single-argument `poll(String ticket)` overload has **no production caller**. It is a test seam. So the only live route to a `null` key is `Principal.ownerKey()` returning `null`, and that happens for one case only: ```java case PRIMARY -> name == null ? null : prefixed("leader", name); ``` That is the **unnamed primary**: a caller the daemon resolved as primary with no pane name — token auth, or loopback trust. Every other role returns a non-null key. So the gap is: **an unnamed primary reads every lead's ticket, including the full reply text.** The live-turn path refuses the same caller — `Rendezvous.Owner.permits` needs an exact match and never treats `null` as a wildcard. Async tickets are therefore less protected than live turns, which is the asymmetry the reviewer spotted. That part stands. ### The shape of it: one `null`, two meanings The `null` key carries two different states that need opposite handling: | who passes `null` | what it means | what it should get | |---|---|---| | an internal caller, through `poll(String)` | "identity is not relevant here" | no check | | an unnamed primary, through `ownerKey()` | "a real caller with no name" | the same check as everyone else | Today both get "no check". This is the same conflated-sentinel shape as #497: one symbol, two states, one handling. The fix is a third state — keep the internal bypass explicit, and stop the unnamed primary sharing its value. ### How this relates to #705 #705 offers three fix shapes for ticket walking. Its **option 2** is "give tickets an owner check", and that is what this ticket is. This finding is the hole left *inside* option 2: the owner check now exists, and one caller still skips it. #705 can close on option 1 or 2 without this being fixed, so I am keeping the two separate rather than merging the tickets. ### Decision - PR #741 is **not blocked by this**. It fixes the "too closed" side it was written for and leaves the `null` rule untouched and documented. - The `null` wildcard becomes **#737 unit 6**, specified below with the rest of the units. - I am not changing it inside #741. Widening the scope of a merged-ready PR to cover a pre-existing hole is how a verified diff stops being verified. Severity, stated plainly: lower than the reviewer's `high`. An unnamed primary already holds `SPAWN`, `STOP`, `SEND` and `DRAIN` over every session on this host, so it can reach most of the same content by other routes. It is still worth closing, because the live-turn path already refuses it and the two paths should agree. I checked the code in this comment myself, on `origin/main` at `aabecce` and on `origin/worker/737-owner-key-ff061f-10`. I did not run a live probe from an unnamed primary.
Author
Owner

Unit 6 — stop the unnamed primary sharing the internal bypass

Comes from the PR #741 review above. Do this after units 1+2 are merged, because it edits the same method.

What changes

MessageService.ownsTicket must stop treating a null owner key as "read everything". Today one null means two things: an internal caller that has no identity to check, and a real authenticated caller that has no name. Give them two different values.

Shape to use — a named constant for the internal bypass, so the caller says what it means:

  • the single-argument poll(String ticket) overload passes an explicit "no check" marker, not null. It has no production caller (measured above), so this is a test seam saying so in code.
  • an owner key of null from Principal.ownerKey() is then just another key. An unnamed primary matches only a ticket it created itself, the same rule every other role already follows.

Do not simply delete the null branch. A ticket created by an unnamed primary records creatorOwner == null, so the comparison must still treat two nulls as a match, or a token-auth primary loses its own tickets. That is the bug this whole ticket exists to prevent, in a new place.

Apply the same change to pendingAsk (:1793 on the branch), which calls ownsTicket with the same key.

Why

Rendezvous.Owner.permits already refuses a null caller for a named owner, so the live-turn path and the ticket path disagree about the same caller. One of them is wrong, and the stricter one is right.

Tests

  • an unnamed primary polls a ticket a named lead created ⇒ refused, with the existing forbidden: reason and no reply text.
  • an unnamed primary polls a ticket it created ⇒ allowed. This is the positive control. Without it, the first test passes just as well if the method refuses everyone.
  • the internal bypass still reads any ticket, driven through the one-argument overload.
  • the same three for pendingAsk.

Mutation evidence, not a green build

  • make the bypass marker compare equal to a real key ⇒ the "named lead's ticket" test dies.
  • make two null creators compare unequal ⇒ the "its own ticket" test dies.

Show each one red, then revert and show it green. A mutation that kills nothing means the test is vacuous.

Out of scope

Authz's TASK_READ row is not part of this. Whether an unconfigured pane should hold TASK_READ at all is #705, and the two fixes are independent.

## Unit 6 — stop the unnamed primary sharing the internal bypass Comes from the PR #741 review above. Do this **after** units 1+2 are merged, because it edits the same method. ### What changes `MessageService.ownsTicket` must stop treating a `null` owner key as "read everything". Today one `null` means two things: an internal caller that has no identity to check, and a real authenticated caller that has no name. Give them two different values. Shape to use — a named constant for the internal bypass, so the caller says what it means: - the single-argument `poll(String ticket)` overload passes an explicit "no check" marker, not `null`. It has **no production caller** (measured above), so this is a test seam saying so in code. - an owner key of `null` from `Principal.ownerKey()` is then just another key. An unnamed primary matches only a ticket it created itself, the same rule every other role already follows. Do not simply delete the `null` branch. A ticket created by an unnamed primary records `creatorOwner == null`, so the comparison must still treat two `null`s as a match, or a token-auth primary loses its own tickets. That is the bug this whole ticket exists to prevent, in a new place. Apply the same change to `pendingAsk` (`:1793` on the branch), which calls `ownsTicket` with the same key. ### Why `Rendezvous.Owner.permits` already refuses a `null` caller for a named owner, so the live-turn path and the ticket path disagree about the same caller. One of them is wrong, and the stricter one is right. ### Tests - an unnamed primary polls a ticket a named lead created ⇒ refused, with the existing `forbidden:` reason and no reply text. - an unnamed primary polls a ticket **it** created ⇒ allowed. This is the positive control. Without it, the first test passes just as well if the method refuses everyone. - the internal bypass still reads any ticket, driven through the one-argument overload. - the same three for `pendingAsk`. ### Mutation evidence, not a green build - make the bypass marker compare equal to a real key ⇒ the "named lead's ticket" test dies. - make two `null` creators compare unequal ⇒ the "its own ticket" test dies. Show each one red, then revert and show it green. A mutation that kills nothing means the test is vacuous. ### Out of scope `Authz`'s `TASK_READ` row is not part of this. Whether an unconfigured pane should hold `TASK_READ` at all is #705, and the two fixes are independent.
Author
Owner

Unit 6 addendum — two existing tests pin the behaviour you are changing

Units 1+2 are merged as cd1f04c. While checking the merged tree I found that the null wildcard is pinned by two tests. Whoever takes unit 6 will see them go red and must invert them, not weaken or delete them:

$ git grep -n "unnamedPrimary\|UnnamedPrimary" -- 'fleetd/src/test/*'
.../msg/MessageServiceTest.java:951:    void unnamedPrimaryStillReadsAnyTicket()
.../rest/FleetAppAuthTest.java:217:    void restPollRefusesADifferentWorkerButAllowsTheCreatorAndTheUnnamedPrimary()

Both currently assert the wildcard is correct. After unit 6 they must assert the opposite: an unnamed primary reads only a ticket it created. Rename each one to say what it now protects.

Two more that must keep passing unchanged, because they are the reason the null branch cannot simply be deleted:

.../auth/AuthzTest.java:43:    void anAnonymousCallerIsRefusedAnswerSoItCanNeverBeMistakenForTheUnnamedPrimary()
.../auth/CallerResolverTest.java:367:    void anUnnamedPrimaryWithNoPaneOwnsNothing()

This is the trap in this unit. A worker that changes ownsTicket and then finds two red tests has an easy wrong move available: edit the two tests until they are green again. That turns a security fix into a test weakening, and the build stays green either way. Invert them deliberately, and say in the reply what each one asserted before and after.

## Unit 6 addendum — two existing tests pin the behaviour you are changing Units 1+2 are merged as `cd1f04c`. While checking the merged tree I found that the `null` wildcard is **pinned by two tests**. Whoever takes unit 6 will see them go red and must invert them, not weaken or delete them: ``` $ git grep -n "unnamedPrimary\|UnnamedPrimary" -- 'fleetd/src/test/*' .../msg/MessageServiceTest.java:951: void unnamedPrimaryStillReadsAnyTicket() .../rest/FleetAppAuthTest.java:217: void restPollRefusesADifferentWorkerButAllowsTheCreatorAndTheUnnamedPrimary() ``` Both currently assert the wildcard is correct. After unit 6 they must assert the opposite: an unnamed primary reads **only** a ticket it created. Rename each one to say what it now protects. Two more that must keep passing unchanged, because they are the reason the `null` branch cannot simply be deleted: ``` .../auth/AuthzTest.java:43: void anAnonymousCallerIsRefusedAnswerSoItCanNeverBeMistakenForTheUnnamedPrimary() .../auth/CallerResolverTest.java:367: void anUnnamedPrimaryWithNoPaneOwnsNothing() ``` **This is the trap in this unit.** A worker that changes `ownsTicket` and then finds two red tests has an easy wrong move available: edit the two tests until they are green again. That turns a security fix into a test weakening, and the build stays green either way. Invert them deliberately, and say in the reply what each one asserted before and after.
Author
Owner

Unit 4 addendum — the rollover single-flight key

Unit 6 is merged (6ab3a81 + lead follow-up 7467ffa, PR #744 closed). Unit 4's one-line spec needs expanding before it can be delegated, because the obvious implementation introduces a worse bug than the one it fixes. Writing it here so the worker can re-read it.

The defect, stated precisely

rollingByTerminal (LeadRollover.java:372) is the single-flight lock for a lead roll. It is keyed on p.leadTerminal() — the pane address. A roll replaces the pane, so the fresh lead has a different terminal. The lock therefore stops being a per-lead lock: once the fresh lead is up, a second open()+confirm() from it lands on a different map key and is approved while the first roll's continuation is still running.

The two architects disagreed on exactly this row, and both were right about different properties:

  • the claim at :569 and the releases at :601/:647 all use p.leadTerminal(), so the release is correct and nothing leaks;
  • and it no longer excludes a second roll of the same lead.

The fix is the second half only. Do not "fix" the release — it is not broken.

The hazard that makes this harder than it looks

leadNameForTerminal is wired in Fleetd.java:962 as:

Function<String, String> leadNameForTerminal = terminal -> liveLeadTerminals.get().get(terminal);

It reads the live lead roster. The roll kills the old pane, so by the time runRollover's finally runs, leadNameForTerminal.apply(p.leadTerminal()) returns null.

So the key must not be computed independently at the claim site and the release sites. If it is, the release computes a different key from the claim, the claim is never removed, and that lead can never be rolled again — confirm() answers ROLL_ALREADY_RUNNING forever. Resolve the key once and carry it to both release sites.

What must change, and what must not

Change: key the single-flight claim on the lead's configured name.

Keep, because these are genuinely about the pane:

  • :552 NOT_YOUR_ROLLOVER must keep comparing p.leadTerminal() to callerTerminal. That gate means "only the exact pane that opened this request may confirm it". Do not touch it.
  • runRollover's use of p.leadTerminal() from :653 onwards as the pane to tear down stays as it is.

Fallback: if the lead name is null or blank (the terminal is not in the live roster), fall back to keying on the terminal. That is at least as strict as today, must not throw, and must not skip the claim.

Rename the field so it stops saying Terminal once it is keyed by name.

Acceptance criteria

  1. A second confirm() for the same lead, while the first roll's continuation is in flight, is refused ROLL_ALREADY_RUNNING even when the second request was opened from a different terminal for that lead. New behaviour, needs a new test.
  2. The claim is still released on the success path and the thrown-exception path, proven with the old terminal already absent from the live lead roster — drive leadNameForTerminal to return null for the old terminal at release time, then show a later confirm() for that lead is approved, not refused. Without this test the regression above is invisible.
  3. The continuationRunner-rejected release at :601 still works; its existing test stays green.
  4. NOT_YOUR_ROLLOVER behaviour unchanged, existing test green.
  5. Two named mutations on the real file, each shown RED then GREEN:
    • make the release recompute the key from leadNameForTerminal instead of using the carried key → criterion 2's test must fail;
    • revert the claim key to p.leadTerminal() → criterion 1's test must fail.
  6. Full mvn clean install run from fleetd/ — there is no root pom, and running it one directory up fails with no POM in this directory. Unpiped, full output read, rm -rf target/surefire-reports first.

Implementation shape is the worker's call. Two that work: a token → resolved-key map cleaned up on the same paths that release the claim; or carry the resolved key on PendingRollover, resolved in open() at :462 while the lead is certainly still live. PendingRollover is consumed in production only at FleetMcp.java:1514, so adding a component is contained — but check handoverOpen does not put it into the tool response.

Out of scope

MessageService, Principal, FleetMcp's ownership javadoc, ReplyPushLoop, LeadHeartbeatLoop. Unit 5 also edits LeadRollover, so it is deliberately not running at the same time.

## Unit 4 addendum — the rollover single-flight key Unit 6 is merged (`6ab3a81` + lead follow-up `7467ffa`, PR #744 closed). Unit 4's one-line spec needs expanding before it can be delegated, because the obvious implementation introduces a worse bug than the one it fixes. Writing it here so the worker can re-read it. ### The defect, stated precisely `rollingByTerminal` (`LeadRollover.java:372`) is the single-flight lock for a lead roll. It is keyed on `p.leadTerminal()` — the **pane address**. A roll replaces the pane, so the fresh lead has a different terminal. The lock therefore stops being a per-**lead** lock: once the fresh lead is up, a second `open()`+`confirm()` from it lands on a different map key and is approved while the first roll's continuation is still running. The two architects disagreed on exactly this row, and **both were right about different properties**: - the claim at `:569` and the releases at `:601`/`:647` all use `p.leadTerminal()`, so the release is correct and nothing leaks; - and it no longer excludes a second roll of the same lead. The fix is the second half only. **Do not "fix" the release** — it is not broken. ### The hazard that makes this harder than it looks `leadNameForTerminal` is wired in `Fleetd.java:962` as: ```java Function<String, String> leadNameForTerminal = terminal -> liveLeadTerminals.get().get(terminal); ``` It reads the **live** lead roster. The roll kills the old pane, so by the time `runRollover`'s `finally` runs, `leadNameForTerminal.apply(p.leadTerminal())` returns **null**. So the key must **not** be computed independently at the claim site and the release sites. If it is, the release computes a different key from the claim, the claim is never removed, and that lead can **never be rolled again** — `confirm()` answers `ROLL_ALREADY_RUNNING` forever. Resolve the key **once** and carry it to both release sites. ### What must change, and what must not Change: key the single-flight claim on the lead's configured **name**. Keep, because these are genuinely about the pane: - `:552` `NOT_YOUR_ROLLOVER` must keep comparing `p.leadTerminal()` to `callerTerminal`. That gate means "only the exact pane that opened this request may confirm it". Do not touch it. - `runRollover`'s use of `p.leadTerminal()` from `:653` onwards as the pane to tear down stays as it is. Fallback: if the lead name is null or blank (the terminal is not in the live roster), fall back to keying on the terminal. That is at least as strict as today, must not throw, and must not skip the claim. Rename the field so it stops saying `Terminal` once it is keyed by name. ### Acceptance criteria 1. A second `confirm()` for the same lead, while the first roll's continuation is in flight, is refused `ROLL_ALREADY_RUNNING` **even when the second request was opened from a different terminal** for that lead. New behaviour, needs a new test. 2. The claim is still released on the success path **and** the thrown-exception path, proven **with the old terminal already absent from the live lead roster** — drive `leadNameForTerminal` to return null for the old terminal at release time, then show a later `confirm()` for that lead is approved, not refused. Without this test the regression above is invisible. 3. The `continuationRunner`-rejected release at `:601` still works; its existing test stays green. 4. `NOT_YOUR_ROLLOVER` behaviour unchanged, existing test green. 5. Two named mutations on the real file, each shown RED then GREEN: - make the release recompute the key from `leadNameForTerminal` instead of using the carried key → criterion 2's test must fail; - revert the claim key to `p.leadTerminal()` → criterion 1's test must fail. 6. Full `mvn clean install` run from `fleetd/` — **there is no root pom**, and running it one directory up fails with `no POM in this directory`. Unpiped, full output read, `rm -rf target/surefire-reports` first. Implementation shape is the worker's call. Two that work: a token → resolved-key map cleaned up on the same paths that release the claim; or carry the resolved key on `PendingRollover`, resolved in `open()` at `:462` while the lead is certainly still live. `PendingRollover` is consumed in production only at `FleetMcp.java:1514`, so adding a component is contained — but check `handoverOpen` does not put it into the tool response. ### Out of scope `MessageService`, `Principal`, `FleetMcp`'s ownership javadoc, `ReplyPushLoop`, `LeadHeartbeatLoop`. Unit 5 also edits `LeadRollover`, so it is deliberately **not** running at the same time.
Author
Owner

Unit 5 addendum — fleet_handover{open} reports outstanding tickets and open asks

Units 3 and 6 are merged (d438a74 + 803c91e, 6ab3a81 + 7467ffa; PRs #744 and #745 closed). Unit 4 is running. This is unit 5's full spec.

First, a correction to comment 2's quoted code

Comment 2 quotes Rendezvous.Owner.permits as:

public static boolean permits(Owner owner, String callerTerminal) {
    return owner != null && java.util.Objects.equals(owner.terminal(), callerTerminal);
}

That is not the deployed code. Rendezvous.java:78-92 on main today is:

public record Owner(String ownerKey) {
    public static final Owner UNNAMED_PRIMARY = new Owner(null);
    ...
    public static boolean permits(Owner owner, String callerOwner) {
        return owner != null && java.util.Objects.equals(owner.ownerKey(), callerOwner);
    }
}

It is keyed on the owner key, not the terminal. This matters directly for unit 5, because I started writing this spec around the hypothesis that a roll makes a pending ask unanswerable — the fresh lead has a new terminal, so a terminal-keyed gate would refuse it. That hypothesis is false, and I only caught it by reading the file instead of trusting the quote. Comment 2's severity argument about a refused answer still stands for a genuinely different caller; it does not apply to a rolled lead.

Why unit 5 is needed — stated correctly

A lead roll restarts the lead's process, not fleetd. The daemon keeps its tasks map and its rendezvous registrations across the roll. Both are gated on the caller's owner key: MessageService.ownsTicket and Rendezvous.Owner.permits each compare Principal.ownerKey(), which for a named lead is leader:<name> and does not change when the pane changes.

So the fresh session keeps the authority to poll those tickets and answer those asks. What it loses is the knowledge — the ticket ids and turnIds lived in the outgoing session's context and nowhere else. fleet_handover{open} is the one call made while the outgoing session still holds both, so it is the right place to list them for the handover file.

Do not implement this as a warning that a roll "drops" tickets or asks. It does not.

Scope

MessageService.java, the handover/handoverOpen path in FleetMcp.java, and their tests.

Not in scope: LeadRollover.java. Unit 4 is editing that file concurrently. If the change seems to belong there, stop and ask rather than editing it.

Required change

  1. MessageService: a public method listing the delegations a given owner key created that are not yet collected, plus the open asks it may answer. Follow the existing tasks.values() iteration pattern (:509, :648, :802, :1808). Return a record — formatting belongs in the MCP layer. Filter with the same ownership rule ownsTicket already uses; do not introduce a second rule that can drift from it. Do not special-case INTERNAL_NO_OWNER_CHECK here.
  2. FleetMcp: thread the caller's owner key into handover(...)/handoverOpen(...). The established pattern is principal(exchange).ownerKey() — see :525 and :536. Do not build a terminal-to-owner lookup; the principal already carries it.
  3. handoverOpen's JSON gains the outstanding tickets (ticket id, phase, target) and the open asks (ticket id, turnId, worker session). token, handoverPath and requestedAtMillis stay present and unchanged — the handover skill reads them.
  4. An unnamed primary is already refused earlier in handoverOpen by the isBlank(callerTerminal) check. Leave that refusal exactly as it is.

Acceptance criteria

  1. open() from a lead that created two async tickets reports both, with their phases.
  2. open() from a lead whose worker is paused in fleet_ask reports that ask's turnId.
  3. A lead does not see another lead's tickets or asks. Use two named leads with different names — not a named lead and an unnamed one, which would pass for the wrong reason.
  4. token, handoverPath, requestedAtMillis still present and unchanged; existing handoverOpen tests green.
  5. Empty case: a lead with no tickets and no asks gets the response with empty collections, not absent keys and not an error.
  6. Two named mutations, each RED then GREEN:
    • drop the owner filter from the new listing method → criterion 3's test must fail;
    • return an empty list unconditionally → criteria 1 and 2 must fail, and criterion 5's test must still pass. If criterion 5 also fails under this mutation it is not testing the empty case, it is testing nothing.
  7. Full mvn clean install from fleetd/ — there is no root pom. Unpiped, whole output read, rm -rf target/surefire-reports first.
## Unit 5 addendum — `fleet_handover{open}` reports outstanding tickets and open asks Units 3 and 6 are merged (`d438a74` + `803c91e`, `6ab3a81` + `7467ffa`; PRs #744 and #745 closed). Unit 4 is running. This is unit 5's full spec. ### First, a correction to comment 2's quoted code Comment 2 quotes `Rendezvous.Owner.permits` as: ```java public static boolean permits(Owner owner, String callerTerminal) { return owner != null && java.util.Objects.equals(owner.terminal(), callerTerminal); } ``` **That is not the deployed code.** `Rendezvous.java:78-92` on `main` today is: ```java public record Owner(String ownerKey) { public static final Owner UNNAMED_PRIMARY = new Owner(null); ... public static boolean permits(Owner owner, String callerOwner) { return owner != null && java.util.Objects.equals(owner.ownerKey(), callerOwner); } } ``` It is keyed on the **owner key**, not the terminal. This matters directly for unit 5, because I started writing this spec around the hypothesis that a roll makes a pending ask unanswerable — the fresh lead has a new terminal, so a terminal-keyed gate would refuse it. **That hypothesis is false**, and I only caught it by reading the file instead of trusting the quote. Comment 2's severity argument about a refused answer still stands for a genuinely different caller; it does not apply to a rolled lead. ### Why unit 5 is needed — stated correctly A lead roll restarts the **lead's** process, not `fleetd`. The daemon keeps its `tasks` map and its rendezvous registrations across the roll. Both are gated on the caller's owner key: `MessageService.ownsTicket` and `Rendezvous.Owner.permits` each compare `Principal.ownerKey()`, which for a named lead is `leader:<name>` and does **not** change when the pane changes. So the fresh session keeps the **authority** to poll those tickets and answer those asks. What it loses is the **knowledge** — the ticket ids and turnIds lived in the outgoing session's context and nowhere else. `fleet_handover{open}` is the one call made while the outgoing session still holds both, so it is the right place to list them for the handover file. Do **not** implement this as a warning that a roll "drops" tickets or asks. It does not. ### Scope `MessageService.java`, the `handover`/`handoverOpen` path in `FleetMcp.java`, and their tests. **Not in scope: `LeadRollover.java`.** Unit 4 is editing that file concurrently. If the change seems to belong there, stop and ask rather than editing it. ### Required change 1. **`MessageService`**: a public method listing the delegations a given owner key created that are not yet collected, plus the open asks it may answer. Follow the existing `tasks.values()` iteration pattern (`:509`, `:648`, `:802`, `:1808`). Return a record — formatting belongs in the MCP layer. Filter with the **same** ownership rule `ownsTicket` already uses; do not introduce a second rule that can drift from it. Do not special-case `INTERNAL_NO_OWNER_CHECK` here. 2. **`FleetMcp`**: thread the caller's owner key into `handover(...)`/`handoverOpen(...)`. The established pattern is `principal(exchange).ownerKey()` — see `:525` and `:536`. Do not build a terminal-to-owner lookup; the principal already carries it. 3. **`handoverOpen`'s JSON** gains the outstanding tickets (ticket id, phase, target) and the open asks (ticket id, turnId, worker session). `token`, `handoverPath` and `requestedAtMillis` stay present and unchanged — the `handover` skill reads them. 4. An unnamed primary is already refused earlier in `handoverOpen` by the `isBlank(callerTerminal)` check. Leave that refusal exactly as it is. ### Acceptance criteria 1. `open()` from a lead that created two async tickets reports both, with their phases. 2. `open()` from a lead whose worker is paused in `fleet_ask` reports that ask's `turnId`. 3. **A lead does not see another lead's tickets or asks.** Use two *named* leads with different names — not a named lead and an unnamed one, which would pass for the wrong reason. 4. `token`, `handoverPath`, `requestedAtMillis` still present and unchanged; existing `handoverOpen` tests green. 5. Empty case: a lead with no tickets and no asks gets the response with **empty collections**, not absent keys and not an error. 6. Two named mutations, each RED then GREEN: - drop the owner filter from the new listing method → criterion 3's test must fail; - return an empty list unconditionally → criteria 1 and 2 must fail, **and criterion 5's test must still pass**. If criterion 5 also fails under this mutation it is not testing the empty case, it is testing nothing. 7. Full `mvn clean install` from `fleetd/` — there is no root pom. Unpiped, whole output read, `rm -rf target/surefire-reports` first.
Author
Owner

Unit 5 correction — a DONE, uncollected ticket is the case that matters most

This supersedes the filter choice in the "Unit 5 addendum" comment. Everything else in that comment stands.

Unit 5 implemented "not yet collected" as !task.future.isDone(), so a DONE or FAILED ticket is excluded. The worker flagged this itself and gave its reasoning:

a DONE/FAILED ticket sitting in the map waiting for TTL pruning is deliberately excluded, since the lead already has authority to poll it anytime regardless of this feature

That reasoning conflates authority with knowledge, which is the exact distinction the unit 5 addendum drew. Authority survives the roll, because the owner key is leader:<name>. Knowledge does not — the ticket id lived only in the outgoing session's context. So "the lead can poll it anytime" is true and beside the point: after the roll it does not know what to poll.

And the excluded case is the perishable one. pruneTerminalTickets's own javadoc says so (msg/MessageService.java:1511-1533):

The TTL runs from completion, not from creation (#197). [...] the reply lives only in future, so pruning it discards the worker's whole report with nothing to fall back on.

The prune removes any task where future.isDone() && completedNanos < cutoff. So a finished, uncollected ticket is destroyed on a timer, and its report is unrecoverable. A PENDING ticket, by contrast, is not going anywhere — the worker is still running.

The filter is backwards for the purpose. The tickets most worth writing into a handover file are the finished ones nobody has read yet.

Required change

Report every ticket the caller owns that is still present in tasks, with its real phase — PENDING, ASKING, DONE, FAILED. Drop the !task.future.isDone() condition. Keep using ownsTicket as the only ownership rule.

A ticket already collected still sits in tasks until the TTL prunes it, so it will appear too. That is acceptable and better than the alternative: MessageService has no collected flag to filter on — collection is recorded in ReplyPushLoop via ticketCollected(...), and reaching into the push loop from here to hide a row is not worth a new coupling. The phase already tells the reader what it is looking at.

Extra acceptance criterion

  1. A ticket that has completed but has not been polled appears in outstandingTickets with a terminal phase. Add a mutation for it: restore the !task.future.isDone() filter and show this new test goes RED, then revert to GREEN.

Criteria 1-6 from the addendum are unchanged, and the existing tests for them must stay green.

## Unit 5 correction — a DONE, uncollected ticket is the case that matters most This supersedes the filter choice in the "Unit 5 addendum" comment. Everything else in that comment stands. Unit 5 implemented "not yet collected" as `!task.future.isDone()`, so a DONE or FAILED ticket is excluded. The worker flagged this itself and gave its reasoning: > a DONE/FAILED ticket sitting in the map waiting for TTL pruning is deliberately excluded, since the lead already has authority to poll it anytime regardless of this feature **That reasoning conflates authority with knowledge**, which is the exact distinction the unit 5 addendum drew. Authority survives the roll, because the owner key is `leader:<name>`. Knowledge does not — the ticket id lived only in the outgoing session's context. So "the lead can poll it anytime" is true and beside the point: after the roll it does not know *what* to poll. And the excluded case is the perishable one. `pruneTerminalTickets`'s own javadoc says so (`msg/MessageService.java:1511-1533`): > The TTL runs from **completion**, not from creation (#197). [...] the reply lives only in `future`, so pruning it discards the worker's whole report with nothing to fall back on. The prune removes any task where `future.isDone() && completedNanos < cutoff`. So a finished, uncollected ticket is destroyed on a timer, and its report is unrecoverable. A PENDING ticket, by contrast, is not going anywhere — the worker is still running. **The filter is backwards for the purpose.** The tickets most worth writing into a handover file are the finished ones nobody has read yet. ### Required change Report every ticket the caller owns that is still present in `tasks`, with its real phase — `PENDING`, `ASKING`, `DONE`, `FAILED`. Drop the `!task.future.isDone()` condition. Keep using `ownsTicket` as the only ownership rule. A ticket already collected still sits in `tasks` until the TTL prunes it, so it will appear too. That is acceptable and better than the alternative: `MessageService` has no `collected` flag to filter on — collection is recorded in `ReplyPushLoop` via `ticketCollected(...)`, and reaching into the push loop from here to hide a row is not worth a new coupling. The phase already tells the reader what it is looking at. ### Extra acceptance criterion 7. A ticket that has completed but has not been polled appears in `outstandingTickets` with a terminal phase. Add a mutation for it: restore the `!task.future.isDone()` filter and show this new test goes RED, then revert to GREEN. Criteria 1-6 from the addendum are unchanged, and the existing tests for them must stay green.
Author
Owner

Shipped and deployed — closing

The decision this ticket asked the next lead to settle: fix shape 2 (key ownership on the stable identity), plus the reporting half of fix shape 3. Shape 1 (reassign creatorTerminal during the roll) was not taken, because shape 2 removes the reason for it.

Six units merged to main:

Unit Commit What it did
1+2 efd9cdb key tickets and turns on a stable owner, not a terminal
3 f4176ae, 803c91e probe the fallback terminal and resolve it by lead name
4 8a1d73b key the rollover single-flight claim on the lead's name
5 70a735b, d41aff4 fleet_handover{open} reports outstandingTickets and openAsks
6 6ab3a81, 7467ffa stop the unnamed primary sharing the internal bypass

A named lead now owns a ticket through Principal.ownerKey(), which is leader:<name>. That value does not change when the pane changes, so the refusal this ticket predicted cannot fire. main is at 7f0c4a8 with 2118 tests and 0 failures.

Docs updated with the code: wiki/9-Implementation.md (the ownership rules and the line refs), wiki/11-Features.md (the rollover cycle and a third gotcha), and .claude/skills/handover/SKILL.md step 1, which now tells the outgoing lead to copy outstandingTickets and openAsks into the handover file.

It is deployed, and I checked the jar itself

scripts/redeploy-fleetd.sh --yes ran today. The daemon now running is pid 18094, started Mon Oct 5 06:11:16 2026, and it holds fleetd/run/fleetd.jar open, hash 93215c19c8e6. I read the shipped code out of that exact file rather than trusting git log:

$ javap -p -cp <unpacked run/fleetd.jar> dev.ltms.fleet.msg.MessageService | grep -iE 'outstanding|ownsTicket|INTERNAL_NO_OWNER'
  static final java.lang.String INTERNAL_NO_OWNER_CHECK;
  private static boolean ownsTicket(dev.ltms.fleet.msg.MessageService$Task, java.lang.String);
  public dev.ltms.fleet.msg.MessageService$Outstanding outstanding(java.lang.String);

$ javap -p -cp … 'dev.ltms.fleet.lead.LeadRollover$PendingRollover' | grep rolloverKey
  private final java.lang.String rolloverKey;
  public java.lang.String rolloverKey();

#726 unit 2 — the real process restart, which is the change that would have introduced this regression — is also in that jar (4ffe49f). So the regression and its fix went live together, which is what this ticket asked for.

What is NOT measured

Nobody has rolled a lead on the restart path. So the end-to-end claim — a fresh lead, new terminal, same name, collects a ticket the previous lead created — is still read from code and covered only by unit tests. I have not seen it happen.

I tried a non-destructive probe of the deployed behaviour: fleet_handover{action:"open"} followed straight by cancel, which creates a token and changes nothing else. My own permission classifier refused it as modifying a shared resource, and I did not route around it. The real proof needs a deliberate roll of a live lead, which is the operator's call, not mine.

That gap is already tracked in #595, which records that the roll has never been observed working end to end. I am not filing a second ticket for it. The matching gotcha in wiki/11-Features.md stays as it is, because it is still true.

Follow-ups found while doing this, filed separately

  • #740 — a validate* method is wired by its name, so a rename silently removes a boot check.
  • Three stale cross-class javadoc references were found and fixed in this ticket alone (FleetApp.java:885, LeadRollover.java:96, and one wiki bullet). All the same shape: the definition was updated, the prose pointing at it was not. A hunter sweep for {@code X#method} references that no longer resolve would probably find more.

Closing.

## Shipped and deployed — closing The decision this ticket asked the next lead to settle: **fix shape 2** (key ownership on the stable identity), plus the reporting half of **fix shape 3**. Shape 1 (reassign `creatorTerminal` during the roll) was not taken, because shape 2 removes the reason for it. Six units merged to `main`: | Unit | Commit | What it did | |---|---|---| | 1+2 | `efd9cdb` | key tickets and turns on a stable owner, not a terminal | | 3 | `f4176ae`, `803c91e` | probe the fallback terminal and resolve it by lead name | | 4 | `8a1d73b` | key the rollover single-flight claim on the lead's name | | 5 | `70a735b`, `d41aff4` | `fleet_handover{open}` reports `outstandingTickets` and `openAsks` | | 6 | `6ab3a81`, `7467ffa` | stop the unnamed primary sharing the internal bypass | A named lead now owns a ticket through `Principal.ownerKey()`, which is `leader:<name>`. That value does not change when the pane changes, so the refusal this ticket predicted cannot fire. `main` is at `7f0c4a8` with 2118 tests and 0 failures. Docs updated with the code: `wiki/9-Implementation.md` (the ownership rules and the line refs), `wiki/11-Features.md` (the rollover cycle and a third gotcha), and `.claude/skills/handover/SKILL.md` step 1, which now tells the outgoing lead to copy `outstandingTickets` and `openAsks` into the handover file. ### It is deployed, and I checked the jar itself `scripts/redeploy-fleetd.sh --yes` ran today. The daemon now running is pid 18094, started `Mon Oct 5 06:11:16 2026`, and it holds `fleetd/run/fleetd.jar` open, hash `93215c19c8e6`. I read the shipped code out of that exact file rather than trusting `git log`: ``` $ javap -p -cp <unpacked run/fleetd.jar> dev.ltms.fleet.msg.MessageService | grep -iE 'outstanding|ownsTicket|INTERNAL_NO_OWNER' static final java.lang.String INTERNAL_NO_OWNER_CHECK; private static boolean ownsTicket(dev.ltms.fleet.msg.MessageService$Task, java.lang.String); public dev.ltms.fleet.msg.MessageService$Outstanding outstanding(java.lang.String); $ javap -p -cp … 'dev.ltms.fleet.lead.LeadRollover$PendingRollover' | grep rolloverKey private final java.lang.String rolloverKey; public java.lang.String rolloverKey(); ``` #726 unit 2 — the real process restart, which is the change that would have introduced this regression — is also in that jar (`4ffe49f`). So the regression and its fix went live together, which is what this ticket asked for. ### What is NOT measured **Nobody has rolled a lead on the restart path.** So the end-to-end claim — a fresh lead, new terminal, same name, collects a ticket the previous lead created — is still read from code and covered only by unit tests. I have not seen it happen. I tried a non-destructive probe of the deployed behaviour: `fleet_handover{action:"open"}` followed straight by `cancel`, which creates a token and changes nothing else. My own permission classifier refused it as modifying a shared resource, and I did not route around it. The real proof needs a deliberate roll of a live lead, which is the operator's call, not mine. That gap is already tracked in #595, which records that the roll has never been observed working end to end. I am not filing a second ticket for it. The matching gotcha in `wiki/11-Features.md` stays as it is, because it is still true. ### Follow-ups found while doing this, filed separately - #740 — a `validate*` method is wired by its name, so a rename silently removes a boot check. - Three stale cross-class javadoc references were found and fixed in this ticket alone (`FleetApp.java:885`, `LeadRollover.java:96`, and one wiki bullet). All the same shape: the definition was updated, the prose pointing at it was not. A `hunter` sweep for `{@code X#method}` references that no longer resolve would probably find more. Closing.
ltms closed this issue 2026-10-05 06:18:18 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#737