A turnId is a guessable counter and answer() takes no caller identity, so any architect can answer a turn opened for someone else — the same shape as #705, on a write path #715

Closed
opened 2026-10-04 07:03:14 +02:00 by ltms · 8 comments
Owner

Found by the implementer on #705 as an out-of-scope observation, then checked by me in the main
clone.

The two facts

A turnId is predictable. Rendezvous.java:148:

String newTurnId = session + "#" + askSeq.incrementAndGet();

askSeq is one AtomicLong per Rendezvous (:77). So a turn id is the target session id plus a
small integer — term_abc#1, term_abc#2, and so on. A caller that knows a session id, which any
READ holder does today, can enumerate turn ids.

answer() takes no caller. MessageService.java:1150:

public Reply answer(String turnId, String content, long timeoutMillis) {
    String workerSession = rendezvous.askSession(turnId);

It resolves the worker from the turn id and proceeds. There is no check that the caller is the
session whose blocked fleet_send owns that turn.

And the action is granted broadly. Authz.java:118:

case ANSWER -> caller.isPrimary() || caller.isArchitect();

What that adds up to

Any architect can answer a fleet_ask that was opened for a different architect, or for the
primary. A worker is not in the attacker set here, so this is narrower than #705.

It is also worse in kind. #705 is a read: you learn another session's reply. This is a
write: the answer is injected into the worker's resumed turn, so whoever answers steers what
that worker does next. A lead's worker can be redirected by an architect the lead never involved.

Same shape as #705

An identifier is being used as if it were an authorization token. #705's ticket ids are a plain
counter with no owner check; these turn ids are a plain counter with no owner check. PR #712 fixed
the ticket case by recording the creating caller's terminal on the Task and comparing it on
poll. The same shape of fix applies: record the asking caller's counterpart on the rendezvous
entry and compare it in answer.

Decide one thing before implementing: a fleet_ask is answered by the lead that delegated the
turn
, which is recorded in PrimaryRegistry, not by whoever happens to hold ANSWER. So the
comparison is probably against the recorded delegator, not against the architect that called.

Not measured

Nobody has driven this from a second live architect session. All three facts above are read from
the source at the lines quoted. In particular I have not confirmed that a second architect can
reach fleet_send{turnId} for another architect's turn end to end — only that no code on the path
would stop it.

Found by the implementer on #705 as an out-of-scope observation, then checked by me in the main clone. ## The two facts **A `turnId` is predictable.** `Rendezvous.java:148`: ```java String newTurnId = session + "#" + askSeq.incrementAndGet(); ``` `askSeq` is one `AtomicLong` per `Rendezvous` (`:77`). So a turn id is the target session id plus a small integer — `term_abc#1`, `term_abc#2`, and so on. A caller that knows a session id, which any `READ` holder does today, can enumerate turn ids. **`answer()` takes no caller.** `MessageService.java:1150`: ```java public Reply answer(String turnId, String content, long timeoutMillis) { String workerSession = rendezvous.askSession(turnId); ``` It resolves the worker from the turn id and proceeds. There is no check that the caller is the session whose blocked `fleet_send` owns that turn. **And the action is granted broadly.** `Authz.java:118`: ```java case ANSWER -> caller.isPrimary() || caller.isArchitect(); ``` ## What that adds up to Any architect can answer a `fleet_ask` that was opened for a different architect, or for the primary. A worker is not in the attacker set here, so this is narrower than #705. It is also worse in kind. #705 is a **read**: you learn another session's reply. This is a **write**: the answer is injected into the worker's resumed turn, so whoever answers steers what that worker does next. A lead's worker can be redirected by an architect the lead never involved. ## Same shape as #705 An identifier is being used as if it were an authorization token. #705's ticket ids are a plain counter with no owner check; these turn ids are a plain counter with no owner check. PR #712 fixed the ticket case by recording the creating caller's terminal on the `Task` and comparing it on `poll`. The same shape of fix applies: record the asking caller's counterpart on the rendezvous entry and compare it in `answer`. Decide one thing before implementing: a `fleet_ask` is answered by *the lead that delegated the turn*, which is recorded in `PrimaryRegistry`, not by whoever happens to hold `ANSWER`. So the comparison is probably against the recorded delegator, not against the architect that called. ## Not measured Nobody has driven this from a second live architect session. All three facts above are read from the source at the lines quoted. In particular I have **not** confirmed that a second architect can reach `fleet_send{turnId}` for another architect's turn end to end — only that no code on the path would stop it.
Author
Owner

Design call — lead, 2026-10-04

I read the code in the main clone at c468953 and decided this. All line numbers below are from
that commit.

Decision 1 — do NOT compare against PrimaryRegistry. Reject that route.

The issue text above suggests comparing the caller against the recorded delegator in
PrimaryRegistry. I checked that class and I am rejecting it, for two reasons.

It is keyed by the wrong thing. PrimaryRegistry holds
ConcurrentHashMap<String, String> leadByTarget (:36), keyed by the target session, not by
the turn. recordDelegation(target, leadTerminal) (:79) overwrites the entry for that target, and
forgetDelegation(target) (:87) clears it. So its lifetime is the lifetime of the newest
delegation to that worker, not the lifetime of the turn we want to protect.

It is a nudge-routing map, not an authorization record. Its only production reader is
nudgeTargetFor(target) (:103). If the gate reads that map, the gate reads a different source
than the delivery did. We already have a defect of exactly that shape on file — a receipt that reads
a different source than the behaviour reads. I do not want to add a second one on a write path.

Decision 2 — use #712's shape: record the owner on the turn, compare it in answer.

PR #712 fixed the ticket case by recording the creating caller's terminal on the Task and
comparing it on poll. The same shape fits here, and it is exact rather than approximate.

The reason it is exact: the question reaches the lead through the waiter. ask() does
rendezvous.currentWaiter(workerSession) (MessageService.java:1057), then
resolveQuestion(...) (:1059) completes that waiter. The waiter was registered by
rendezvous.open(target) on the sender's behalf. So the session that opened the waiter is the
one and only session that ever learns the turnId. That is the owner. One writer, one reader, and
the turn's own lifetime.

rendezvous.open( has exactly two call sites — I enumerated them with
grep -rn "rendezvous\.open(" .:

  • MessageService.java:970 — the forward send
  • MessageService.java:1161 — the resumed-turn waiter inside answer itself

What the unit has to do

  1. Give the blocking send chain a creatorTerminal. It has none today. I checked the
    signatures: send at :929, :945 and :950 take no terminal, while sendAsync at :1302
    already takes one (that was #712). This is the missing half.

  2. Stamp the owner on the ask when the turn is minted — the fresh() branch of
    Rendezvous.openAsk only. A coalesced duplicate ask keeps the first owner; it does not
    re-stamp.

  3. Compare in answer, with the same null rule as ownsTicket. ownsTicket
    (MessageService.java:1434) is callerTerminal == null || callerTerminal.equals(...). The null
    allowance is for the unnamed primary, which carries no herdr pane. That rule carries over here
    unchanged, because an architect always has a real terminal.

  4. Return a refusal that is distinct from STALE_TURN. A caller must be able to tell "this turn
    is not yours" from "this turn lapsed". Today answer returns STALE_TURN for an unknown
    turnId (:1152), and reusing it would hide the refusal.

  5. Fix both gates. messages.answer(turnId, content, timeout) is called from two places, and
    neither passes a caller:

    • FleetMcp.java:905
    • FleetApp.java:691

    Point 5 is the part most likely to be missed. #705 had this exact shape: the MCP handler was
    fixed and the REST door stayed open, which is what task-15 is closing now. A check at one gate of
    two is not a check.

The coupling this inherits, written down so it is not rediscovered

The null allowance in point 3 is safe only because Authz.java:118 is
case ANSWER -> caller.isPrimary() || caller.isArchitect(), so an ANONYMOUS caller never reaches
answer at all. Nothing in the code pins that. Granting ANONYMOUS the ANSWER action would
silently open this gate. This is the same unpinned coupling that #705's fix depends on, now on a
second path.

Still not measured

I did not drive this from a second live architect session, and nor did the issue author. Every
claim above is read from the source at the lines quoted. So "any architect can answer another
architect's turn" remains a reading of the code, not an observed event. The unit should include a
behavioural test that makes it observed, with a control half that passes for the real owner.

## Design call — lead, 2026-10-04 I read the code in the main clone at `c468953` and decided this. All line numbers below are from that commit. ### Decision 1 — do NOT compare against `PrimaryRegistry`. Reject that route. The issue text above suggests comparing the caller against the recorded delegator in `PrimaryRegistry`. I checked that class and I am rejecting it, for two reasons. **It is keyed by the wrong thing.** `PrimaryRegistry` holds `ConcurrentHashMap<String, String> leadByTarget` (`:36`), keyed by the **target session**, not by the turn. `recordDelegation(target, leadTerminal)` (`:79`) overwrites the entry for that target, and `forgetDelegation(target)` (`:87`) clears it. So its lifetime is the lifetime of the newest delegation to that worker, not the lifetime of the turn we want to protect. **It is a nudge-routing map, not an authorization record.** Its only production reader is `nudgeTargetFor(target)` (`:103`). If the gate reads that map, the gate reads a different source than the delivery did. We already have a defect of exactly that shape on file — a receipt that reads a different source than the behaviour reads. I do not want to add a second one on a write path. ### Decision 2 — use #712's shape: record the owner on the turn, compare it in `answer`. PR #712 fixed the ticket case by recording the creating caller's terminal on the `Task` and comparing it on `poll`. The same shape fits here, and it is exact rather than approximate. The reason it is exact: the question reaches the lead through **the waiter**. `ask()` does `rendezvous.currentWaiter(workerSession)` (`MessageService.java:1057`), then `resolveQuestion(...)` (`:1059`) completes that waiter. The waiter was registered by `rendezvous.open(target)` on the **sender's** behalf. So the session that opened the waiter is the one and only session that ever learns the `turnId`. That is the owner. One writer, one reader, and the turn's own lifetime. `rendezvous.open(` has exactly two call sites — I enumerated them with `grep -rn "rendezvous\.open(" .`: - `MessageService.java:970` — the forward `send` - `MessageService.java:1161` — the resumed-turn waiter inside `answer` itself ### What the unit has to do 1. **Give the blocking `send` chain a `creatorTerminal`.** It has none today. I checked the signatures: `send` at `:929`, `:945` and `:950` take no terminal, while `sendAsync` at `:1302` already takes one (that was #712). This is the missing half. 2. **Stamp the owner on the ask when the turn is minted** — the `fresh()` branch of `Rendezvous.openAsk` only. A coalesced duplicate ask keeps the first owner; it does not re-stamp. 3. **Compare in `answer`, with the same null rule as `ownsTicket`.** `ownsTicket` (`MessageService.java:1434`) is `callerTerminal == null || callerTerminal.equals(...)`. The null allowance is for the unnamed primary, which carries no herdr pane. That rule carries over here unchanged, because an architect always has a real terminal. 4. **Return a refusal that is distinct from `STALE_TURN`.** A caller must be able to tell "this turn is not yours" from "this turn lapsed". Today `answer` returns `STALE_TURN` for an unknown `turnId` (`:1152`), and reusing it would hide the refusal. 5. **Fix both gates.** `messages.answer(turnId, content, timeout)` is called from two places, and neither passes a caller: - `FleetMcp.java:905` - `FleetApp.java:691` Point 5 is the part most likely to be missed. #705 had this exact shape: the MCP handler was fixed and the REST door stayed open, which is what task-15 is closing now. A check at one gate of two is not a check. ### The coupling this inherits, written down so it is not rediscovered The null allowance in point 3 is safe **only because** `Authz.java:118` is `case ANSWER -> caller.isPrimary() || caller.isArchitect()`, so an `ANONYMOUS` caller never reaches `answer` at all. Nothing in the code pins that. Granting `ANONYMOUS` the `ANSWER` action would silently open this gate. This is the same unpinned coupling that #705's fix depends on, now on a second path. ### Still not measured I did not drive this from a second live architect session, and nor did the issue author. Every claim above is read from the source at the lines quoted. So "any architect can answer another architect's turn" remains a reading of the code, not an observed event. The unit should include a behavioural test that makes it observed, with a control half that passes for the real owner.
Author
Owner

Design call, measured in the main clone at 9a64d42

The previous lead left this as "needs a design call first: compare against the recorded delegator in
PrimaryRegistry, not against whoever holds ANSWER." I worked that through. The direction is
right but the named field is wrong
, and the obvious version of the fix would break the architect
role.

How guessable a turnId actually is

// Rendezvous.java:148
String newTurnId = session + "#" + askSeq.incrementAndGet();

So a turnId is term_65cfd7f2d6b4270#3 — the target's terminal id plus a small counter. It is
not a bare sequential id like a ticket, which weakens the original framing: you must know a session
id before you can guess a turn.

That changes who the realistic attacker is. #710 now hides the members and leads arrays from a
worker — I confirmed that live today, a worker's fleet_list returns only healthCoverage,
loopHealth and capacity. So a worker can no longer enumerate session ids from the bridge at all.
But membersVisibleTo admits an architect, and ANSWER is caller.isPrimary() || caller.isArchitect(). So the reachable case is an architect answering a turn belonging to a
delegation it has nothing to do with.

Why ownerTerminal is the wrong field — this is the part that matters

The recorded spawn owner does exist, and it is on the session, not in PrimaryRegistry:

MemberSession.java:22   @param ownerTerminal  the caller that requested this worker (null = daemon/anon)
FleetMcp.java:1431      m.put("owner", s.ownerTerminal());

But case SPAWN -> caller.isPrimary(). Only a lead can ever spawn, so ownerTerminal is always
a lead's terminal or null. Gating answer on it would mean an architect can never answer any ask —
which deletes a capability the role table grants on purpose. The Authz comment states that intent:

Resolving a worker's blocked question is part of delegating to it, open to the same two roles that
may stand up that delegation in the first place.

An architect cannot spawn but can SEND, so an architect answering an ask from a member a lead
spawned is intended, not an abuse. A spawn-ownership check would forbid the normal case and allow
nothing new.

The field to compare against instead

Compare against who sent the brief that is currently running, not who spawned the pane. That
record already exists, because #705 added it: Task.creatorTerminal, the field ownsTicket reads.
Whoever delegated the turn is the party whose question it is, and that is true whether they are the
lead or an architect.

Two traps this fix must not walk into, both already paid for here

  1. Do not add a convenience overload that defaults the caller to null. That is exactly #718:
    MessageService.poll(String) forwards with null, and ownsTicket treats null as "skip the
    check", so the short form fails open. If answer gains a caller parameter, change the
    signature — do not leave a shorter one behind for a future call site to find.
  2. Decide the "no recorded sender" case deliberately, and say which way it fails. A blocking
    fleet_send and the REST path do not necessarily leave a Task. If the check is
    caller == null || caller.equals(recorded) it fails open and buys nothing. If it is a strict
    equality against a null record, every such rendezvous becomes unanswerable — which is the
    same break task-15 found in #705, where gating the ticket read alone would have closed a hole and
    broken the documented REST fallback in one commit. Enumerate the paths that open an ask with no
    recorded sender before writing the comparison, and state the chosen behaviour for each.

Still not measured

I have not driven the hijack end to end: two leads, one member, an architect answering the other
delegation's turn. Everything above is read from Rendezvous.java:148, the Authz table, the
SPAWN row and MemberSession's javadoc. The reachability argument rests on an architect being able
to read members, which I confirmed from membersVisibleTo, not from a live architect session.

Not delegating yet: trap 2 needs the path enumeration done first, and it touches MessageService
where #718's scrape test is in flight.

## Design call, measured in the main clone at `9a64d42` The previous lead left this as "needs a design call first: compare against the recorded delegator in `PrimaryRegistry`, not against whoever holds `ANSWER`." I worked that through. **The direction is right but the named field is wrong**, and the obvious version of the fix would break the architect role. ### How guessable a `turnId` actually is ```java // Rendezvous.java:148 String newTurnId = session + "#" + askSeq.incrementAndGet(); ``` So a `turnId` is `term_65cfd7f2d6b4270#3` — the **target's terminal id** plus a small counter. It is not a bare sequential id like a ticket, which weakens the original framing: you must know a session id before you can guess a turn. That changes who the realistic attacker is. #710 now hides the `members` **and** `leads` arrays from a worker — I confirmed that live today, a worker's `fleet_list` returns only `healthCoverage`, `loopHealth` and `capacity`. So a worker can no longer enumerate session ids from the bridge at all. But `membersVisibleTo` admits an **architect**, and `ANSWER` is `caller.isPrimary() || caller.isArchitect()`. So the reachable case is an architect answering a turn belonging to a delegation it has nothing to do with. ### Why `ownerTerminal` is the wrong field — this is the part that matters The recorded spawn owner does exist, and it is on the session, not in `PrimaryRegistry`: ``` MemberSession.java:22 @param ownerTerminal the caller that requested this worker (null = daemon/anon) FleetMcp.java:1431 m.put("owner", s.ownerTerminal()); ``` But `case SPAWN -> caller.isPrimary()`. **Only a lead can ever spawn**, so `ownerTerminal` is always a lead's terminal or `null`. Gating `answer` on it would mean an architect can never answer any ask — which deletes a capability the role table grants on purpose. The `Authz` comment states that intent: > Resolving a worker's blocked question is part of delegating to it, open to the same two roles that > may stand up that delegation in the first place. An architect cannot spawn but can `SEND`, so an architect answering an ask from a member a lead spawned is **intended**, not an abuse. A spawn-ownership check would forbid the normal case and allow nothing new. ### The field to compare against instead Compare against **who sent the brief that is currently running**, not who spawned the pane. That record already exists, because #705 added it: `Task.creatorTerminal`, the field `ownsTicket` reads. Whoever delegated the turn is the party whose question it is, and that is true whether they are the lead or an architect. ### Two traps this fix must not walk into, both already paid for here 1. **Do not add a convenience overload that defaults the caller to `null`.** That is exactly #718: `MessageService.poll(String)` forwards with `null`, and `ownsTicket` treats `null` as "skip the check", so the short form fails **open**. If `answer` gains a caller parameter, change the signature — do not leave a shorter one behind for a future call site to find. 2. **Decide the "no recorded sender" case deliberately, and say which way it fails.** A blocking `fleet_send` and the REST path do not necessarily leave a `Task`. If the check is `caller == null || caller.equals(recorded)` it fails open and buys nothing. If it is a strict equality against a `null` record, **every such rendezvous becomes unanswerable** — which is the same break task-15 found in #705, where gating the ticket read alone would have closed a hole and broken the documented REST fallback in one commit. Enumerate the paths that open an ask with no recorded sender **before** writing the comparison, and state the chosen behaviour for each. ### Still not measured I have not driven the hijack end to end: two leads, one member, an architect answering the other delegation's turn. Everything above is read from `Rendezvous.java:148`, the `Authz` table, the `SPAWN` row and `MemberSession`'s javadoc. The reachability argument rests on an architect being able to read `members`, which I confirmed from `membersVisibleTo`, not from a live architect session. Not delegating yet: trap 2 needs the path enumeration done first, and it touches `MessageService` where #718's scrape test is in flight.
Author
Owner

Correcting my own comment above: the turnId is handed out, not guessed — see #721

In the comment above I wrote that a turnId is session + "#" + askSeq.incrementAndGet(), and
concluded:

It is not a bare sequential id like a ticket, which weakens the original framing: you must know a
session id before you can guess a turn.

That mitigation does not hold. An architect reviewing #705 found that fleet_status has no owner
check, and I verified it on both the MCP and REST surfaces. FleetMcp.status returns the pending
ask's question text, its turnId and its ticket for any sessionId the caller names, gated
only on TASK_READ. Details and the quoted code are on #721.

So there is no guessing step at all:

Step Call Grant
1 fleet_list → every member's sessionId READ (membersVisibleTo admits an architect)
2 fleet_status{sessionId} → that member's turnId + question text TASK_READ
3 fleet_send{turnId, content} ANSWER = primary or architect

This makes #715 more serious than its title says, not less, and it changes the scoping:

  • The title's word "guessable" is now the least of it. The id is published to any TASK_READ holder.
  • #721 and this ticket want the same comparison against Task.creatorTerminal, on the read side
    and the write side. They are probably one change, and #721 is the half that leaks content, so it
    leads.
  • Everything else in my comment above stands, in particular that MemberSession.ownerTerminal is the
    wrong field — SPAWN is primary-only, so gating on it would forbid the normal architect case and
    permit nothing new — and that the "no recorded sender" case must be decided explicitly, because one
    direction fails open (#718's shape) and the other breaks the REST path (the break task-15 caught in
    #705).

Still not measured, unchanged from above: nobody has driven the hijack from a live architect session.

## Correcting my own comment above: the `turnId` is handed out, not guessed — see #721 In the comment above I wrote that a `turnId` is `session + "#" + askSeq.incrementAndGet()`, and concluded: > It is not a bare sequential id like a ticket, which weakens the original framing: you must know a > session id before you can guess a turn. **That mitigation does not hold.** An architect reviewing #705 found that `fleet_status` has no owner check, and I verified it on both the MCP and REST surfaces. `FleetMcp.status` returns the pending ask's **question text**, its `turnId` **and** its `ticket` for any `sessionId` the caller names, gated only on `TASK_READ`. Details and the quoted code are on **#721**. So there is no guessing step at all: | Step | Call | Grant | |---|---|---| | 1 | `fleet_list` → every member's `sessionId` | `READ` (`membersVisibleTo` admits an architect) | | 2 | `fleet_status{sessionId}` → that member's `turnId` + question text | `TASK_READ` | | 3 | `fleet_send{turnId, content}` | `ANSWER` = primary or architect | This makes #715 **more** serious than its title says, not less, and it changes the scoping: - The title's word "guessable" is now the least of it. The id is published to any `TASK_READ` holder. - #721 and this ticket want the **same** comparison against `Task.creatorTerminal`, on the read side and the write side. They are probably one change, and #721 is the half that leaks content, so it leads. - Everything else in my comment above stands, in particular that `MemberSession.ownerTerminal` is the wrong field — `SPAWN` is primary-only, so gating on it would forbid the normal architect case and permit nothing new — and that the "no recorded sender" case must be decided explicitly, because one direction fails open (#718's shape) and the other breaks the REST path (the break task-15 caught in #705). Still not measured, unchanged from above: nobody has driven the hijack from a live architect session.
Author
Owner

Decision — settled by an architect (profile sol) at 6e06058, and I accept it. This supersedes the "decide one thing before implementing" paragraph in the ticket body, which guessed wrong.

The decision

The owner of a turn is the caller whose accepted delegation started the running turn. Record that owner on the forward rendezvous; when a fresh fleet_ask opens, copy it onto the ask turn; answer() then compares the answering caller against that stored snapshot.

Do not look ownership up at answer time from Task, MemberSession, or PrimaryRegistry.

One authorization rule covers both paths. They differ only in the capture input:

  • blocking send — the caller identity passed into send()
  • async send — the existing Task.creatorTerminal

My own ticket body was wrong about the mechanism

The body says "the comparison is probably against the recorded delegator" in PrimaryRegistry. The architect rejected that with reasons I checked:

  • PrimaryRegistry is keyed by worker session, not turn, and a later accepted send overwrites the value (PrimaryRegistry.java:69-84). A blocking send returns QUESTION and releases its session lock (MessageService.java:989-1036) while its ask stays open — so a second send can change a session-scoped routing value while the older turn is still answerable.
  • It also drops null owners (:79-82), which loses the unnamed-primary case entirely.
  • Its production purpose is nudge routing (nudgeTargetFor(), :93-105), not authorization.

So the owner has to be snapshotted per turn, not looked up per session.

Three corrections to premises, all of which I verified myself

1. The async callers do not always pass a real terminal. This is the one that matters most, and I had asserted the opposite on #721:

// FleetApp.java:698
String ticket = messages.sendAsync(id, content, null, caller == null ? null : caller.terminal());
// Principal.java:35-37 — the unnamed primary has no terminal by construction
public static Principal primary(long pid) { return new Principal(Role.PRIMARY, null, pid); }

So Task.creatorTerminal == null is a valid production state.

2. The blocking waiter does not expose the sender. Rendezvous.waiters holds only a CompletableFuture<Resolution> (Rendezvous.java:73). The entry represents the blocked send but carries no sender identity. So there is nothing to read today, and this is why the fix needs a new field rather than a lookup.

3. Neither adapter passes a caller into blocking send or into answer. I confirmed all four sites:

FleetMcp.java:917:  messages.send(sessionId, content, timeout, onAccepted)   // no caller
FleetApp.java:704:  messages.send(id, content, timeout)                      // no caller
FleetMcp.java:933   and   FleetApp.java:691                                  // answer, no caller

That makes this a four-file unit, not the one- or two-file change the body implies.

And one addition to my own reasoning about MemberSession.ownerTerminal. I had ruled it out on #721 on the grounds that SPAWN is primary-only, so gating on it "would forbid the normal architect case while permitting nothing new". The architect agrees but shows the second half is too weak: a spawn-owner lead would keep the right to answer every later turn on that member, including one an architect delegated. So it does not permit nothing new — it permits a stale right. Rejecting it is still correct, for a stronger reason.

The null rule — three states, not two

This is the part I would have got wrong, and it is worth stating plainly because it is the opposite of what #721 does.

state meaning rule
known terminal owner a terminal-bearing caller delegated it must match exactly
known unnamed-primary owner the authenticated no-pane primary delegated it that primary may answer
no owner record the system cannot prove who owns the turn fail closed for everyone, including the unnamed primary

So the implementation needs record presence kept separate from the nullable terminal value. ownsTicket's two-state callerTerminal == null || equals(...) is adequate for #721's read path and is not adequate here: a missing record must not inherit the unnamed-primary exception. Collapsing "unknown" into "unnamed primary" is the trap.

Also: keep ownership refusal distinct from STALE_TURN, and run the ownership check before taking the session lock or opening the resumed waiter (MessageService.java:1155-1161), so a refused caller cannot consume the ask, clear question state, or delay the real owner.

Unit shape

One vertical unit, not split — splitting the core ownership record from either adapter leaves a working bypass. Files: Rendezvous.java, MessageService.java, FleetMcp.java, FleetApp.java. No Authz policy change; the ANSWER role gate stays as the outer check and gets a test pinning that coupling.

Scheduling: this rebases after #721. Both edit MessageService, and #721 is in flight now. The architect flagged the same thing independently.

Full acceptance properties are in the architect's reply and I will put them in the brief when I delegate. The two I most want pinned: a refused answer performs no rendezvous, task, or question cleanup, and no production answer overload can silently substitute a null caller — the second is the #718 shape, one file over.

Limits

The architect read the code and changed nothing (git status --short empty at 6e06058). It ran no build and no live hijack, and did not read #721's unmerged branch. The end-to-end hijack is still not demonstrated from a live second architect session — that was unverified when I filed this and it still is.

**Decision — settled by an architect (profile `sol`) at `6e06058`, and I accept it.** This supersedes the "decide one thing before implementing" paragraph in the ticket body, which guessed wrong. ## The decision **The owner of a turn is the caller whose accepted delegation started the running turn.** Record that owner on the forward rendezvous; when a *fresh* `fleet_ask` opens, copy it onto the ask turn; `answer()` then compares the answering caller against that stored snapshot. Do **not** look ownership up at answer time from `Task`, `MemberSession`, or `PrimaryRegistry`. One authorization rule covers both paths. They differ only in the capture input: - **blocking send** — the caller identity passed into `send()` - **async send** — the existing `Task.creatorTerminal` ## My own ticket body was wrong about the mechanism The body says *"the comparison is probably against the recorded delegator"* in `PrimaryRegistry`. The architect rejected that with reasons I checked: - `PrimaryRegistry` is keyed by **worker session, not turn**, and a later accepted send overwrites the value (`PrimaryRegistry.java:69-84`). A blocking send returns `QUESTION` and releases its session lock (`MessageService.java:989-1036`) while its ask stays open — so a second send can change a session-scoped routing value while the older turn is still answerable. - It also drops null owners (`:79-82`), which loses the unnamed-primary case entirely. - Its production purpose is nudge routing (`nudgeTargetFor()`, `:93-105`), not authorization. So the owner has to be snapshotted **per turn**, not looked up per session. ## Three corrections to premises, all of which I verified myself **1. The async callers do not always pass a real terminal.** This is the one that matters most, and I had asserted the opposite on #721: ```java // FleetApp.java:698 String ticket = messages.sendAsync(id, content, null, caller == null ? null : caller.terminal()); // Principal.java:35-37 — the unnamed primary has no terminal by construction public static Principal primary(long pid) { return new Principal(Role.PRIMARY, null, pid); } ``` So `Task.creatorTerminal == null` is a valid production state. **2. The blocking waiter does not expose the sender.** `Rendezvous.waiters` holds only a `CompletableFuture<Resolution>` (`Rendezvous.java:73`). The entry represents the blocked send but carries no sender identity. So there is nothing to read today, and this is why the fix needs a new field rather than a lookup. **3. Neither adapter passes a caller into blocking send or into answer.** I confirmed all four sites: ```java FleetMcp.java:917: messages.send(sessionId, content, timeout, onAccepted) // no caller FleetApp.java:704: messages.send(id, content, timeout) // no caller FleetMcp.java:933 and FleetApp.java:691 // answer, no caller ``` That makes this a four-file unit, not the one- or two-file change the body implies. **And one addition to my own reasoning about `MemberSession.ownerTerminal`.** I had ruled it out on #721 on the grounds that `SPAWN` is primary-only, so gating on it "would forbid the normal architect case while permitting nothing new". The architect agrees but shows the second half is too weak: a spawn-owner lead would keep the right to answer **every later turn on that member, including one an architect delegated**. So it does not permit nothing new — it permits a stale right. Rejecting it is still correct, for a stronger reason. ## The null rule — three states, not two This is the part I would have got wrong, and it is worth stating plainly because it is the opposite of what #721 does. | state | meaning | rule | |---|---|---| | known terminal owner | a terminal-bearing caller delegated it | must match exactly | | known unnamed-primary owner | the authenticated no-pane primary delegated it | that primary may answer | | **no owner record** | the system cannot prove who owns the turn | **fail closed for everyone, including the unnamed primary** | So the implementation needs **record presence kept separate from the nullable terminal value**. `ownsTicket`'s two-state `callerTerminal == null || equals(...)` is adequate for #721's *read* path and is **not** adequate here: a missing record must not inherit the unnamed-primary exception. Collapsing "unknown" into "unnamed primary" is the trap. Also: keep ownership refusal **distinct** from `STALE_TURN`, and run the ownership check **before** taking the session lock or opening the resumed waiter (`MessageService.java:1155-1161`), so a refused caller cannot consume the ask, clear question state, or delay the real owner. ## Unit shape **One vertical unit**, not split — splitting the core ownership record from either adapter leaves a working bypass. Files: `Rendezvous.java`, `MessageService.java`, `FleetMcp.java`, `FleetApp.java`. No `Authz` policy change; the `ANSWER` role gate stays as the outer check and gets a test pinning that coupling. **Scheduling: this rebases after #721.** Both edit `MessageService`, and #721 is in flight now. The architect flagged the same thing independently. Full acceptance properties are in the architect's reply and I will put them in the brief when I delegate. The two I most want pinned: *a refused answer performs no rendezvous, task, or question cleanup*, and *no production answer overload can silently substitute a null caller* — the second is the #718 shape, one file over. ## Limits The architect read the code and changed nothing (`git status --short` empty at `6e06058`). It ran no build and no live hijack, and did not read #721's unmerged branch. The end-to-end hijack is still **not** demonstrated from a live second architect session — that was unverified when I filed this and it still is.
Author
Owner

Do not build this on PrimaryRegistry — measured 2026-10-04 at 28a1f3d

There is already a map from a member to the lead that delegated to it, and it looks like exactly what
this ticket needs. It is not. Writing this down before an implementer finds it and reuses it.

mcp/PrimaryRegistry.java:36 holds leadByTarget, a ConcurrentHashMap<String, String> keyed by the
member's target, written by :79 recordDelegation(target, leadTerminal) and cleared by
:87 forgetDelegation(target). It is wired through MessageService.send's onAccepted hook (CB-548).

Three reasons it cannot serve as the authorization gate:

  1. One route records it, the other does not. The only production writer of recordDelegation is
    mcp/FleetMcp.java:490, which covers both the blocking send (:494) and the async one (:493).
    The REST blocking send at rest/FleetApp.java:704 calls messages.send(id, content, timeout) and
    passes no onAccepted, so it records no delegator at all. A gate reading this map would have
    nothing to compare for a REST-delegated ask. This is the same one-line-two-routes shape as #721.
  2. The record is erased during normal operation. forgetDelegation has two production callers:
    Fleetd.java:700 and msg/ReplyPushLoop.java:450. The second clears it when the delegating lead
    looks dead, which has nothing to do with whether an ask is owned. A gate built on it would change
    its answer for a reason unrelated to ownership.
  3. It is keyed by target, not by turn. It is last-writer-wins per member, so it cannot distinguish
    two successive delegations to the same member — which is the case the gate exists to separate.

So the record this ticket needs is its own, and msg/Rendezvous.java:63's AskWaiter is still the
right home — it is already per-turnId, which is the key the gate needs. The record is constructed at
exactly one site, :150, so adding a field has one place to fill it.

One observation I am deliberately not filing as a defect

FleetApp.java:704 recording no delegator may well be correct: a blocking caller is holding the
connection, so there is no reply to nudge anywhere. I did not find a path where its absence causes a
wrong outcome, so I am not calling it a bug. Noting it only because an implementer reading
recordDelegation's single call site should know the asymmetry is already there and is not theirs to
fix in this ticket.

## Do not build this on `PrimaryRegistry` — measured 2026-10-04 at `28a1f3d` There is already a map from a member to the lead that delegated to it, and it looks like exactly what this ticket needs. It is not. Writing this down before an implementer finds it and reuses it. `mcp/PrimaryRegistry.java:36` holds `leadByTarget`, a `ConcurrentHashMap<String, String>` keyed by the member's target, written by `:79 recordDelegation(target, leadTerminal)` and cleared by `:87 forgetDelegation(target)`. It is wired through `MessageService.send`'s `onAccepted` hook (CB-548). Three reasons it cannot serve as the authorization gate: 1. **One route records it, the other does not.** The only production writer of `recordDelegation` is `mcp/FleetMcp.java:490`, which covers both the blocking send (`:494`) and the async one (`:493`). The REST blocking send at `rest/FleetApp.java:704` calls `messages.send(id, content, timeout)` and passes **no** `onAccepted`, so it records no delegator at all. A gate reading this map would have nothing to compare for a REST-delegated ask. This is the same one-line-two-routes shape as #721. 2. **The record is erased during normal operation.** `forgetDelegation` has two production callers: `Fleetd.java:700` and `msg/ReplyPushLoop.java:450`. The second clears it when the *delegating lead* looks dead, which has nothing to do with whether an ask is owned. A gate built on it would change its answer for a reason unrelated to ownership. 3. **It is keyed by target, not by turn.** It is last-writer-wins per member, so it cannot distinguish two successive delegations to the same member — which is the case the gate exists to separate. So the record this ticket needs is its own, and `msg/Rendezvous.java:63`'s `AskWaiter` is still the right home — it is already per-`turnId`, which is the key the gate needs. The record is constructed at exactly one site, `:150`, so adding a field has one place to fill it. ### One observation I am deliberately *not* filing as a defect `FleetApp.java:704` recording no delegator may well be correct: a blocking caller is holding the connection, so there is no reply to nudge anywhere. I did not find a path where its absence causes a wrong outcome, so I am not calling it a bug. Noting it only because an implementer reading `recordDelegation`'s single call site should know the asymmetry is already there and is not theirs to fix in this ticket.
Author
Owner

Lead review of dd18bd1 — the fix is right. One deletion before I merge.

I read the production diff at dd18bd1. The design is implemented correctly and the two traps I most cared about are both avoided.

The three-state null rule is right. Rendezvous.Owner.permits is

return owner != null && java.util.Objects.equals(owner.terminal(), callerTerminal);

so a null owner — no record at all — matches nobody, including a caller whose own terminal is null. Owner.UNNAMED_PRIMARY is a recorded owner with a null terminal and does match the unnamed primary. "No record" and "recorded as the unnamed primary" stay distinct, which is the part that would have been easy to collapse.

No null-defaulting convenience overload survived on the gated methods. The 3-arg answer(...) and the 3-arg send(...) are gone rather than kept as forwarders. That was the #718 shape and it is not here.

The check runs early enough. answer resolves the session, returns STALE_TURN for an unknown turn, then checks ownership — all before sessionLocks.computeIfAbsent, before tryLock, and before rendezvous.open. So a refused answer performs no rendezvous, lock, task or question work. NOT_TURN_OWNER is a separate outcome and is excluded from the completed-delegation mapping alongside STALE_TURN and QUESTION.

The resumed waiter inherits the same owner (MessageService.java:1182 passes owner, not the answering caller), so answering a turn does not transfer ownership of it.

The one change I want: delete Rendezvous.open(String)

Rendezvous kept a 1-arg overload:

public CompletableFuture<Resolution> open(String session) {
    return open(session, null);
}

That is a convenience overload which defaults the owner to null, which is what the brief asked you not to leave behind. Two things make it a smaller problem than #718, and I want to be accurate about that:

  • it fails closed, not open — a null owner makes permits refuse everyone, so a misuse produces a loud unanswerable turn rather than a silent bypass;
  • it has no caller at all. I checked production and tests with git grep -nE '\b(rendezvous|rv|r)\.open\([^,)]*\)' across fleetd/src and got nothing, with a positive control confirming the pattern does find the 2-arg form.

So it is dead code whose only behaviour is to produce the fail-closed state. #718 handled the equivalent case one file over by pinning that no production caller uses the short form, but deleting is strictly better here: a method that does not exist needs no pin and cannot be found by a future caller.

Please delete the 1-arg open(String) overload and re-run mvn -o clean install. If the build then fails because something does call it, stop and tell me what — that would mean my grep was wrong and I would rather know.

Nothing else. Both production files and the test additions otherwise look right to me, and I will do my own build and mutation pass on the final commit.

## Lead review of `dd18bd1` — the fix is right. One deletion before I merge. I read the production diff at `dd18bd1`. The design is implemented correctly and the two traps I most cared about are both avoided. **The three-state null rule is right.** `Rendezvous.Owner.permits` is ```java return owner != null && java.util.Objects.equals(owner.terminal(), callerTerminal); ``` so a `null` *owner* — no record at all — matches nobody, including a caller whose own terminal is `null`. `Owner.UNNAMED_PRIMARY` is a *recorded* owner with a null terminal and does match the unnamed primary. "No record" and "recorded as the unnamed primary" stay distinct, which is the part that would have been easy to collapse. **No null-defaulting convenience overload survived on the gated methods.** The 3-arg `answer(...)` and the 3-arg `send(...)` are gone rather than kept as forwarders. That was the #718 shape and it is not here. **The check runs early enough.** `answer` resolves the session, returns `STALE_TURN` for an unknown turn, then checks ownership — all before `sessionLocks.computeIfAbsent`, before `tryLock`, and before `rendezvous.open`. So a refused answer performs no rendezvous, lock, task or question work. `NOT_TURN_OWNER` is a separate outcome and is excluded from the completed-delegation mapping alongside `STALE_TURN` and `QUESTION`. **The resumed waiter inherits the same owner** (`MessageService.java:1182` passes `owner`, not the answering caller), so answering a turn does not transfer ownership of it. ### The one change I want: delete `Rendezvous.open(String)` `Rendezvous` kept a 1-arg overload: ```java public CompletableFuture<Resolution> open(String session) { return open(session, null); } ``` That is a convenience overload which defaults the owner to `null`, which is what the brief asked you not to leave behind. Two things make it a smaller problem than #718, and I want to be accurate about that: - it fails **closed**, not open — a `null` owner makes `permits` refuse everyone, so a misuse produces a loud unanswerable turn rather than a silent bypass; - **it has no caller at all.** I checked production and tests with `git grep -nE '\b(rendezvous|rv|r)\.open\([^,)]*\)'` across `fleetd/src` and got nothing, with a positive control confirming the pattern does find the 2-arg form. So it is dead code whose only behaviour is to produce the fail-closed state. #718 handled the equivalent case one file over by pinning that no production caller uses the short form, but deleting is strictly better here: a method that does not exist needs no pin and cannot be found by a future caller. **Please delete the 1-arg `open(String)` overload** and re-run `mvn -o clean install`. If the build then fails because something does call it, stop and tell me what — that would mean my grep was wrong and I would rather know. Nothing else. Both production files and the test additions otherwise look right to me, and I will do my own build and mutation pass on the final commit.
Author
Owner

Correcting comment 18794: my grep was broken, so "delete the overload" was wrong advice

The implementer refused that instruction and checked my measurement instead of following it. It was right to. Do not act on comment 18794's deletion request.

What was wrong

I wrote that Rendezvous.open(String) has no caller, measured with:

git grep -nE '\b(rendezvous|rv|r)\.open\([^,)]*\)' fleetd/src

\b is a GNU-grep escape. git grep -E uses a POSIX ERE engine where it is not a word boundary, so the whole pattern silently matches nothing. I verified this myself on one file that definitely contains the text:

git grep -cE 'rendezvous\.open\('   …/CompletionResolverTest.java   → 64
git grep -cE '\brendezvous\.open\(' …/CompletionResolverTest.java   → no matches, exit 1

Same file, same text, one escape apart. My zero was the pattern failing, not the search succeeding.

And my "positive control" did not control anything. It checked that a different pattern — \.open\([^,)]*,, with no \b — still found the 2-arg form. So it exercised a pattern that could not fail in the way the real one did. A control has to run the same mechanism that might be broken, or it only proves that some other pattern works. That is the part I got wrong twice: once in the measurement, once in believing I had guarded it.

The real count

81 call sites of the 1-arg form, across 7 files, all test-only:

File Calls
inject/CompletionResolverTest.java 64
msg/RendezvousTest.java 7
FleetdTurnListenerCompositionTest.java 4
inject/InjectorTest.java 3
FleetdTurnRegistrarWiringTest.java 1
inject/BackendOutageFlowTest.java 1
msg/MessageServiceTest.java 1

Two small corrections to the implementer's own figures, which I checked: it reported 82 across 8 files and included health/FleetHealthMonitorTest.java, which has no call. And the single hit my grep found under src/main is a javadoc reference, {@code rendezvous.open()} at health/FleetHealthMonitor.java:95, not a call site. So its central claim stands: no production caller, and deleting the overload would break compilation across six files outside #715's scope.

Decision: keep the overload, and pin it the way #718 pinned poll(String)

Deleting is off the table — it is an 81-site migration across unrelated files, which is not this ticket.

The overload stays, and it is materially safer than #718's residual: a null owner makes permits refuse everyone, so a future production misuse produces a loud unanswerable turn, not a silent bypass. What it still lacks is anything stopping that misuse from being introduced, which is exactly what #718 added for poll(String).

So: a source-scrape test asserting no production caller uses the 1-arg form — with a mandatory positive control, for the reason this comment exists. A scrape test whose pattern is broken reads as a clean pass, which is the same failure I just made, promoted into the suite where it would sit unnoticed. The control must prove the pattern finds a real call, by matching one of the 81 known test call sites.

## Correcting comment 18794: my grep was broken, so "delete the overload" was wrong advice The implementer refused that instruction and checked my measurement instead of following it. It was right to. **Do not act on comment 18794's deletion request.** ### What was wrong I wrote that `Rendezvous.open(String)` has no caller, measured with: ``` git grep -nE '\b(rendezvous|rv|r)\.open\([^,)]*\)' fleetd/src ``` `\b` is a GNU-grep escape. `git grep -E` uses a POSIX ERE engine where it is not a word boundary, so **the whole pattern silently matches nothing**. I verified this myself on one file that definitely contains the text: ``` git grep -cE 'rendezvous\.open\(' …/CompletionResolverTest.java → 64 git grep -cE '\brendezvous\.open\(' …/CompletionResolverTest.java → no matches, exit 1 ``` Same file, same text, one escape apart. My zero was the pattern failing, not the search succeeding. **And my "positive control" did not control anything.** It checked that a *different* pattern — `\.open\([^,)]*,`, with no `\b` — still found the 2-arg form. So it exercised a pattern that could not fail in the way the real one did. A control has to run the same mechanism that might be broken, or it only proves that some other pattern works. That is the part I got wrong twice: once in the measurement, once in believing I had guarded it. ### The real count 81 call sites of the 1-arg form, across 7 files, **all test-only**: | File | Calls | |---|---| | `inject/CompletionResolverTest.java` | 64 | | `msg/RendezvousTest.java` | 7 | | `FleetdTurnListenerCompositionTest.java` | 4 | | `inject/InjectorTest.java` | 3 | | `FleetdTurnRegistrarWiringTest.java` | 1 | | `inject/BackendOutageFlowTest.java` | 1 | | `msg/MessageServiceTest.java` | 1 | Two small corrections to the implementer's own figures, which I checked: it reported 82 across 8 files and included `health/FleetHealthMonitorTest.java`, which has **no** call. And the single hit my grep found under `src/main` is a javadoc reference, `{@code rendezvous.open()}` at `health/FleetHealthMonitor.java:95`, not a call site. So its central claim stands: **no production caller**, and deleting the overload would break compilation across six files outside #715's scope. ### Decision: keep the overload, and pin it the way #718 pinned `poll(String)` Deleting is off the table — it is an 81-site migration across unrelated files, which is not this ticket. The overload stays, and it is materially safer than #718's residual: a `null` owner makes `permits` refuse **everyone**, so a future production misuse produces a loud unanswerable turn, not a silent bypass. What it still lacks is anything stopping that misuse from being introduced, which is exactly what #718 added for `poll(String)`. So: a source-scrape test asserting no **production** caller uses the 1-arg form — **with a mandatory positive control**, for the reason this comment exists. A scrape test whose pattern is broken reads as a clean pass, which is the same failure I just made, promoted into the suite where it would sit unnoticed. The control must prove the pattern finds a real call, by matching one of the 81 known test call sites.
Author
Owner

Lead verification of 2374de2 — my own build and mutation pass

I said in comment 18794 that I would do this myself. Here is what I ran and what it showed.

Merge identity. I merged origin/worker/715-5c43fc-1 onto origin/main (ed4f4b0) in a throwaway
worktree, no conflicts. The merge tree and the branch tree are the same object
(ac0147a5bfe2f6d1a199768525783b1c616bc331 both ways), so the worker's green build does cover this
merge. I built it again anyway.

Baseline: Tests run: 2040, Failures: 0, Errors: 0, Skipped: 0 — the same count the worker
reported.

Mutations

Each one was line-anchored, gated on a compile before the suite, then restored and checked
byte-identical against HEAD (git diff --stat empty every time).

# Mutation Result
M1 Owner.permits → return true KILLED by 7 tests in 4 classes: RendezvousTest, MessageServiceTest, FleetMcpTest, FleetAppAuthTest
M2 ownership failure returns STALE_TURN, not NOT_TURN_OWNER KILLED by 4 tests, including the REST mapping (403 became 409)
M3 resumed waiter opened with Owner.of(callerTerminal) instead of owner survived — equivalent mutant, no evidence (see below)
M3b resumed waiter opened with null owner KILLED by secondFleetAskInTheSameResumedTurnDoesNotKillTheAsyncTicket
M4b the new scan test's anchor broken to rendezvous.openNOPE( KILLED on the CONTROL, with the right message
M5 openAsk stamps Owner.UNNAMED_PRIMARY instead of ownerOf(session) KILLED by 11 tests (9 failures + 2 errors) in 4 classes

M3 is my own bad mutation, not a gap in the tests. By the time line 1182 runs, permits has
already proved owner.terminal() equals callerTerminal, so Owner.of(callerTerminal) is owner.
The mutation cannot change behaviour, so its survival is not evidence of anything. M3b is the
mutation I should have written first, and it is killed.

M4b is the one I most wanted. It fails with:

control failed: the scanner found only 0 one-argument rendezvous.open( call(s) in
src/test/java/dev/ltms/fleet/inject/CompletionResolverTest.java, which is known to hold many
-- the matching logic itself is broken: []

That is exactly the failure mode of my wrong git grep in comment 18794: a pattern that matches
nothing reads as a clean production result. The control fires first and names the real cause, so
that mistake can no longer pass as evidence. This test now pins the lesson in CI.

Two corrections to my own method, recorded because they both produced vacuous numbers:

  1. My first M4 attempt targeted the wrong line, and mvn compile does not compile test sources —
    so its exit=0 proved nothing and the mvn test failure that followed was a compile error, not a
    failing control. A mutation in a test file must be gated on mvn test-compile.
  2. baseline exit=$? after mvn … | tail -25 reads tail's status, not Maven's. The load-bearing
    evidence for the baseline is the surefire aggregate, and I re-ran mvn -o clean install with the
    exit code captured directly.

One limitation of the new test, not a blocker

scanFileForOpenCalls anchors on the literal receiver rendezvous.open(. A future production call
through a differently-named reference (rv.open() or a bare open(session) inside Rendezvous
itself would not be seen. That matches the current convention everywhere in src/main/java, and the
arity classification is the part that mattered, so I am not holding the merge for it. Noting it so
nobody later reads a green result as broader than it is.

Verdict: merging. The fail-open overload is gone from production, the three-state rule
(no record / unnamed primary / named terminal) is pinned by direct assertions, and all four
entry points — unit, service, MCP and REST — fail when the gate is removed.

## Lead verification of `2374de2` — my own build and mutation pass I said in comment 18794 that I would do this myself. Here is what I ran and what it showed. **Merge identity.** I merged `origin/worker/715-5c43fc-1` onto `origin/main` (`ed4f4b0`) in a throwaway worktree, no conflicts. The merge tree and the branch tree are the **same** object (`ac0147a5bfe2f6d1a199768525783b1c616bc331` both ways), so the worker's green build does cover this merge. I built it again anyway. **Baseline:** `Tests run: 2040, Failures: 0, Errors: 0, Skipped: 0` — the same count the worker reported. ### Mutations Each one was line-anchored, gated on a compile **before** the suite, then restored and checked byte-identical against `HEAD` (`git diff --stat` empty every time). | # | Mutation | Result | |---|---|---| | M1 | `Owner.permits` → `return true` | **KILLED** by 7 tests in 4 classes: `RendezvousTest`, `MessageServiceTest`, `FleetMcpTest`, `FleetAppAuthTest` | | M2 | ownership failure returns `STALE_TURN`, not `NOT_TURN_OWNER` | **KILLED** by 4 tests, including the REST mapping (403 became 409) | | M3 | resumed waiter opened with `Owner.of(callerTerminal)` instead of `owner` | **survived — equivalent mutant, no evidence** (see below) | | M3b | resumed waiter opened with `null` owner | **KILLED** by `secondFleetAskInTheSameResumedTurnDoesNotKillTheAsyncTicket` | | M4b | the new scan test's anchor broken to `rendezvous.openNOPE(` | **KILLED on the CONTROL**, with the right message | | M5 | `openAsk` stamps `Owner.UNNAMED_PRIMARY` instead of `ownerOf(session)` | **KILLED** by 11 tests (9 failures + 2 errors) in 4 classes | **M3 is my own bad mutation, not a gap in the tests.** By the time line 1182 runs, `permits` has already proved `owner.terminal()` equals `callerTerminal`, so `Owner.of(callerTerminal)` *is* `owner`. The mutation cannot change behaviour, so its survival is not evidence of anything. M3b is the mutation I should have written first, and it is killed. **M4b is the one I most wanted.** It fails with: ``` control failed: the scanner found only 0 one-argument rendezvous.open( call(s) in src/test/java/dev/ltms/fleet/inject/CompletionResolverTest.java, which is known to hold many -- the matching logic itself is broken: [] ``` That is exactly the failure mode of my wrong `git grep` in comment 18794: a pattern that matches nothing reads as a clean production result. The control fires first and names the real cause, so that mistake can no longer pass as evidence. This test now pins the lesson in CI. **Two corrections to my own method, recorded because they both produced vacuous numbers:** 1. My first M4 attempt targeted the wrong line, and `mvn compile` **does not compile test sources** — so its `exit=0` proved nothing and the `mvn test` failure that followed was a compile error, not a failing control. A mutation in a test file must be gated on `mvn test-compile`. 2. `baseline exit=$?` after `mvn … | tail -25` reads `tail`'s status, not Maven's. The load-bearing evidence for the baseline is the surefire aggregate, and I re-ran `mvn -o clean install` with the exit code captured directly. ### One limitation of the new test, not a blocker `scanFileForOpenCalls` anchors on the literal receiver `rendezvous.open(`. A future production call through a differently-named reference (`rv.open(`) or a bare `open(session)` inside `Rendezvous` itself would not be seen. That matches the current convention everywhere in `src/main/java`, and the arity classification is the part that mattered, so I am not holding the merge for it. Noting it so nobody later reads a green result as broader than it is. **Verdict: merging.** The fail-open overload is gone from production, the three-state rule (`no record` / `unnamed primary` / `named terminal`) is pinned by direct assertions, and all four entry points — unit, service, MCP and REST — fail when the gate is removed.
ltms closed this issue 2026-10-04 10:13:02 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#715