Any unconfigured herdr pane resolves to WORKER, and WORKER holds TASK_READ — so an unlisted tab can walk every ticket and read other sessions' delegation replies #705

Open
opened 2026-10-04 05:35:35 +02:00 by ltms · 7 comments
Owner

Split out of #669, where it is stated in the body but has never had its own ticket.

The hole

CallerResolver's resolution ladder (auth/CallerResolver.java:43):

A loopback peer PID that maps to any other herdr pane ⇒ Role#WORKER. This is unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured regardless of auth mode, so enabling auth never breaks the fleet.

So every pane that is not a live spawned member, not a configured lead, not a bound architect slot and not a configured collaborator falls to WORKER. That includes a tab a person opened for something unrelated.

And WORKER holds TASK_READ (auth/Authz.java:145):

case TASK_READ -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();

The code comment immediately above it already states the consequence — it is the stated reason a collaborator is excluded from the same row:

Ticket polling and session status, open to every role READ is open to except a collaborator: ticket ids are a sequential counter with no owner check, so a holder could walk every ticket and read another session's delegation reply.

That reasoning applies with equal force to the worker fallback, which is reachable without any configuration at all. The narrower role was given the stricter rule; the catch-all default was not.

#669's body says the same thing in passing:

leaving them unlisted makes them a worker, which is worse than it sounds — the worker fallback carries TASK_READ, so any unlisted human tab could already poll any ticket and read another member's delegation reply.

Why this is not the same as #702

#702 is a narrow timing window during a member's release, and only with placement: pane. This is the steady state for any pane the operator never configured, with no window and no special placement.

What makes it exploitable rather than theoretical

Ticket ids are a plain sequential counter with no owner check. So the holder does not need to guess or discover anything — task-1, task-2, … enumerates the fleet's delegation traffic. A delegation reply is exactly the content most likely to carry work product, file contents and findings.

Not yet measured

I have not run a live probe of this. I read the resolver ladder and the Authz row; I did not open an unconfigured tab and poll a ticket from it. Anyone picking this up should prove it end to end first, and should expect to need a positive control — a probe that returns nothing proves nothing until you have seen the same call succeed from a session you know holds TASK_READ.

The fix is a design choice, not a one-liner

Three shapes, and they are not equivalent:

  1. Narrow the fallback. A new bottom rung — READ/METRICS only, no TASK_READ, no SEND. Strictly narrower than both WORKER and COLLABORATOR, and it does not widen anything.
  2. Give tickets an owner check. Fixes the root cause rather than the role table, and would also let COLLABORATOR hold TASK_READ safely for its own tickets. Larger change.
  3. Default unconfigured panes to COLLABORATOR. Removes TASK_READ but adds SEND, so any unconfigured pane could inject text into a lead's pane. A trade, not a strict improvement — noted here so it is not mistaken for the obvious answer.

Option 1 is the smallest safe step and does not foreclose option 2. This interacts with the named-peer mesh question being designed under #669, so the two should land in a consistent order — but option 1 is defensible on its own, whatever the mesh decision turns out to be.

Split out of #669, where it is stated in the body but has never had its own ticket. ## The hole `CallerResolver`'s resolution ladder (`auth/CallerResolver.java:43`): > A loopback peer PID that maps to any other herdr pane ⇒ `Role#WORKER`. This is unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured regardless of auth mode, so enabling auth never breaks the fleet. So **every** pane that is not a live spawned member, not a configured lead, not a bound architect slot and not a configured collaborator falls to `WORKER`. That includes a tab a person opened for something unrelated. And `WORKER` holds `TASK_READ` (`auth/Authz.java:145`): ```java case TASK_READ -> caller.isPrimary() || caller.isWorker() || caller.isArchitect(); ``` The code comment immediately above it already states the consequence — it is the stated reason a **collaborator** is excluded from the same row: > Ticket polling and session status, open to every role READ is open to except a collaborator: ticket ids are a sequential counter with no owner check, so a holder could walk every ticket and read another session's delegation reply. That reasoning applies with equal force to the worker fallback, which is reachable **without any configuration at all**. The narrower role was given the stricter rule; the catch-all default was not. #669's body says the same thing in passing: > leaving them unlisted makes them a worker, which is worse than it sounds — the worker fallback carries `TASK_READ`, so any unlisted human tab could already poll any ticket and read another member's delegation reply. ## Why this is not the same as #702 #702 is a narrow timing window during a member's release, and only with `placement: pane`. This is the steady state for any pane the operator never configured, with no window and no special placement. ## What makes it exploitable rather than theoretical Ticket ids are a plain sequential counter with no owner check. So the holder does not need to guess or discover anything — `task-1`, `task-2`, … enumerates the fleet's delegation traffic. A delegation reply is exactly the content most likely to carry work product, file contents and findings. ## Not yet measured I have **not** run a live probe of this. I read the resolver ladder and the `Authz` row; I did not open an unconfigured tab and poll a ticket from it. Anyone picking this up should prove it end to end first, and should expect to need a positive control — a probe that returns nothing proves nothing until you have seen the same call succeed from a session you know holds `TASK_READ`. ## The fix is a design choice, not a one-liner Three shapes, and they are not equivalent: 1. **Narrow the fallback.** A new bottom rung — `READ`/`METRICS` only, no `TASK_READ`, no `SEND`. Strictly narrower than both `WORKER` and `COLLABORATOR`, and it does not widen anything. 2. **Give tickets an owner check.** Fixes the root cause rather than the role table, and would also let `COLLABORATOR` hold `TASK_READ` safely for its own tickets. Larger change. 3. **Default unconfigured panes to `COLLABORATOR`.** Removes `TASK_READ` but **adds** `SEND`, so any unconfigured pane could inject text into a lead's pane. A trade, not a strict improvement — noted here so it is not mistaken for the obvious answer. Option 1 is the smallest safe step and does not foreclose option 2. This interacts with the named-peer mesh question being designed under #669, so the two should land in a consistent order — but option 1 is defensible on its own, whatever the mesh decision turns out to be.
Author
Owner

A measurement that may make option 1 unsafe — with an architect now

The ticket prefers option 1, a new bottom rung below WORKER holding only READ/METRICS, and
calls it "the smallest safe step [that] does not foreclose option 2". I agree with the diagnosis.
I am not yet sure option 1 is safe, for a reason the ticket does not cover.

The WORKER floor looks load-bearing for a spawned member's boot.

  • SessionManager.java:235 calls launcher.spawn(req). The registry entry is written only at
    :244, after spawn returns.
  • Inside spawn, HerdrPeerLauncher.waitUntilInjectableOrThrow (around
    HerdrPeerLauncher.java:1034) polls herdr until the pane is injectable.
  • The live fleetd.yaml on this host sets spawnReadyTimeoutMs: 20000.
  • So a spawned member can be up and serving MCP, and absent from the registry, for up to 20
    seconds. CallerResolver asks the roster first (#669 Unit D), so during that window the member
    falls through to the floor.
  • FleetMcp.java:789-792 markSpawnedMemberPresent is the only writer into MemberPresence. I
    proved that by enumerating writers rather than readers: grep -rn markPresent returned 4 hits
    and one was an override calling super.
  • It fires only when Principal.isSpawnedMember() is true, and Principal.java:113 defines that
    as WORKER || ARCHITECT.

If a booting member's first MCP call resolved as the new bottom rung, presence would not be marked
for that call. If nothing later marks it, the lead's first fleet_send to that member would sit on
the injector readiness gate for about 60 seconds and fail without a keystroke reaching the pane.
That is the same silent, slow shape as the collaborator defect just fixed in #706 — allowed at one
gate, dropped at another.

What is not yet established. I have not shown that a member actually makes an MCP call inside
that window. It is possible every first call lands after registration, in which case the worry is
theoretical and option 1 is clean. I did not measure it, and I am not going to assume it in either
direction — assuming it this session is how I got two things wrong already on this same code.

Now with an architect (ticket task-8), asked to settle five things: whether the window is
really reachable; if it is, which of four fixes is right, including whether registering before the
injectable wait is even possible given that the paneId comes back from spawn(); whether the
ticket's option 2 (give tickets an owner check) dodges the problem entirely by fixing the root
cause instead of the role table; the landing order so no intermediate state is either exploitable
or broken; and a unit list. I asked it to form its own position and to say plainly where it
disagrees with the reading above.

The probe is still not run. The ticket's own "Not yet measured" section still stands: nobody
has opened an unconfigured tab and polled a ticket from it. That needs a positive control, because
a probe returning nothing proves nothing until the same call is seen to succeed from a session
known to hold TASK_READ. I have not done it either, and I am not treating the hole as confirmed
end to end.

## A measurement that may make option 1 unsafe — with an architect now The ticket prefers option 1, a new bottom rung below `WORKER` holding only `READ`/`METRICS`, and calls it "the smallest safe step [that] does not foreclose option 2". I agree with the diagnosis. I am not yet sure option 1 is safe, for a reason the ticket does not cover. **The `WORKER` floor looks load-bearing for a spawned member's boot.** - `SessionManager.java:235` calls `launcher.spawn(req)`. The registry entry is written only at `:244`, after spawn returns. - Inside spawn, `HerdrPeerLauncher.waitUntilInjectableOrThrow` (around `HerdrPeerLauncher.java:1034`) polls herdr until the pane is injectable. - The live `fleetd.yaml` on this host sets `spawnReadyTimeoutMs: 20000`. - So a spawned member can be up and serving MCP, and absent from the registry, for up to 20 seconds. `CallerResolver` asks the roster first (#669 Unit D), so during that window the member falls through to the floor. - `FleetMcp.java:789-792` `markSpawnedMemberPresent` is the only writer into `MemberPresence`. I proved that by enumerating writers rather than readers: `grep -rn markPresent` returned 4 hits and one was an override calling `super`. - It fires only when `Principal.isSpawnedMember()` is true, and `Principal.java:113` defines that as `WORKER || ARCHITECT`. If a booting member's first MCP call resolved as the new bottom rung, presence would not be marked for that call. If nothing later marks it, the lead's first `fleet_send` to that member would sit on the injector readiness gate for about 60 seconds and fail without a keystroke reaching the pane. That is the same silent, slow shape as the collaborator defect just fixed in #706 — allowed at one gate, dropped at another. **What is not yet established.** I have not shown that a member actually makes an MCP call inside that window. It is possible every first call lands after registration, in which case the worry is theoretical and option 1 is clean. I did not measure it, and I am not going to assume it in either direction — assuming it this session is how I got two things wrong already on this same code. **Now with an architect** (ticket `task-8`), asked to settle five things: whether the window is really reachable; if it is, which of four fixes is right, including whether registering before the injectable wait is even possible given that the paneId comes back *from* `spawn()`; whether the ticket's option 2 (give tickets an owner check) dodges the problem entirely by fixing the root cause instead of the role table; the landing order so no intermediate state is either exploitable or broken; and a unit list. I asked it to form its own position and to say plainly where it disagrees with the reading above. **The probe is still not run.** The ticket's own "Not yet measured" section still stands: nobody has opened an unconfigured tab and polled a ticket from it. That needs a positive control, because a probe returning nothing proves nothing until the same call is seen to succeed from a session known to hold `TASK_READ`. I have not done it either, and I am not treating the hole as confirmed end to end.
Author
Owner

Decision: option 2 first, and alone. Option 1 is optional hardening, not the fix.

The architect disagreed with me and it is right. I checked its correction in the code myself
before accepting it, because I had already been wrong twice this session on this same path.

My step 4 was wrong

I said a spawned member "can be up and serving MCP, but absent from the registry, for up to 20
seconds". That read spawnReadyTimeoutMs: 20000 as a duration the member spends serving MCP. It
is a timeout on reaching idle, and idle comes before the MCP connect, not after.

MemberPresence's own javadoc says it, and I read it: herdr's agent_status "reports idle for a
worker whose Claude is still booting", and presence is "populated from the MCP transport … its
initialize is the first such contact". So waitUntilInjectableOrThrow returns while Claude is
still booting, registry.put follows a few in-memory statements later, and the member's first MCP
call arrives after that. The order is the opposite of what my worry assumed. The whole reason
MemberPresence exists is that idle is an unreliable readiness signal — if MCP connected first
there would be no boot window to guard at all.

The architect also found a measurement rather than only a reading: a member can only reach READY
through transitionByTerminal, which scans the registry, and onDelivered refuses BUSY unless
the state is already READY or DONE. I verified that gate at SessionManager.java:911. All
three live members reported state: "busy", so each had a post-registration MCP contact. Under an
OBSERVER floor that same contact resolves WORKER from the roster and marks presence exactly as
it does today.

It was careful about what this does not prove: it did not prove the microsecond race impossible,
only that it does not fire in practice and that no live member depended on an in-window contact.
That is the right distinction and it is enough to act on.

But my worry was right about a different case, and that is the valuable finding

A member pane that outlives a daemon restart really is absent from the registry, indefinitely.
FleetMcp.java:1303 documents it in whoami's own javadoc: "A worker the registry has no record
of — one that outlived a daemon restart". Today that pane hits the floor, resolves WORKER, and
marks presence, so a new daemon can deliver to it again. Under option 1 it would resolve
OBSERVER, presence would never be marked, and Fleetd.deliverableTo would refuse it forever.
I verified the floor returns Principal.worker(...) and read that javadoc.

That is the same shape as the collaborator defect merged in #706 an hour ago: allowed at one gate,
dropped at another, silently. So the failure mode I feared is real — it just lives on the restart
path, not the spawn path. It is a precondition on option 1 and not on option 2.

Why option 2 wins, and this is the part that decided it

Option 1 does not fix the hole. It shrinks the set of principals that can exploit it. After option
1, every legitimately spawned member still holds TASK_READ and can still walk task-1,
task-2, … and read other members' delegation replies. A spawned member is a language model
running an arbitrary brief, which is not a principal I would trust with every other delegation's
reply body. Option 2 closes it for everyone, and it has no contact with the boot path at all.

Option 1 also does not cover all of TASK_READ: FleetMcp.java:1124 maps fleet_status to it as
well. Option 2 does not close that either, but a live status is a much smaller harm than a reply
body.

Order

Option 2 alone leaves no exploitable intermediate state. Option 1 without the presence change is
the worst kind of intermediate state — possibly broken rather than exploitable, and nothing in the
suite would catch it — so that state must not exist even for one commit. If option 1 ever ships, it
ships with the presence condition widened in the same commit, splitting the floor's two jobs: it
currently both grants authorization and supplies the liveness signal that marks presence, and
option 1 narrows only the first while silently taking the second with it.

Delegated

Unit A (the owner check) is going out now. Its acceptance criteria include the two the architect
named that I would not have thought of: an unnamed primary has a null terminal and must still
read any ticket, and the refusal branch itself needs a control, so that the "another terminal is
refused" test cannot pass merely because everything is refused.

Two things are explicitly not checked and must not be assumed: whether a lead keeps its
terminal id across fleet_handover, which nobody has read LeadRollover to confirm; and the
restart-with-live-panes path, which is reasoned from code only. And after this, lead A can no
longer poll lead B's ticket. I judge that correct and intended, but it is a behaviour change rather
than a side effect, so it is recorded here.

The end-to-end probe from an unconfigured tab is still not run, by anyone. The ticket's "Not
yet measured" section stands.

## Decision: option 2 first, and alone. Option 1 is optional hardening, not the fix. The architect disagreed with me and it is right. I checked its correction in the code myself before accepting it, because I had already been wrong twice this session on this same path. ### My step 4 was wrong I said a spawned member "can be up and serving MCP, but absent from the registry, for up to 20 seconds". That read `spawnReadyTimeoutMs: 20000` as a duration the member spends serving MCP. It is a timeout on reaching `idle`, and `idle` comes **before** the MCP connect, not after. `MemberPresence`'s own javadoc says it, and I read it: herdr's `agent_status` "reports `idle` for a worker whose Claude is still booting", and presence is "populated from the MCP transport … its `initialize` is the first such contact". So `waitUntilInjectableOrThrow` returns while Claude is still booting, `registry.put` follows a few in-memory statements later, and the member's first MCP call arrives after that. The order is the opposite of what my worry assumed. The whole reason `MemberPresence` exists is that `idle` is an unreliable readiness signal — if MCP connected first there would be no boot window to guard at all. The architect also found a measurement rather than only a reading: a member can only reach `READY` through `transitionByTerminal`, which scans the registry, and `onDelivered` refuses `BUSY` unless the state is already `READY` or `DONE`. I verified that gate at `SessionManager.java:911`. All three live members reported `state: "busy"`, so each had a post-registration MCP contact. Under an `OBSERVER` floor that same contact resolves `WORKER` from the roster and marks presence exactly as it does today. It was careful about what this does not prove: it did not prove the microsecond race impossible, only that it does not fire in practice and that no live member depended on an in-window contact. That is the right distinction and it is enough to act on. ### But my worry was right about a different case, and that is the valuable finding A member pane that **outlives a daemon restart** really is absent from the registry, indefinitely. `FleetMcp.java:1303` documents it in `whoami`'s own javadoc: "A worker the registry has no record of — one that outlived a daemon restart". Today that pane hits the floor, resolves `WORKER`, and marks presence, so a new daemon can deliver to it again. Under option 1 it would resolve `OBSERVER`, presence would never be marked, and `Fleetd.deliverableTo` would refuse it forever. I verified the floor returns `Principal.worker(...)` and read that javadoc. That is the same shape as the collaborator defect merged in #706 an hour ago: allowed at one gate, dropped at another, silently. So the failure mode I feared is real — it just lives on the restart path, not the spawn path. It is a precondition on option 1 and not on option 2. ### Why option 2 wins, and this is the part that decided it Option 1 does not fix the hole. It shrinks the set of principals that can exploit it. After option 1, **every legitimately spawned member still holds `TASK_READ`** and can still walk `task-1`, `task-2`, … and read other members' delegation replies. A spawned member is a language model running an arbitrary brief, which is not a principal I would trust with every other delegation's reply body. Option 2 closes it for everyone, and it has no contact with the boot path at all. Option 1 also does not cover all of `TASK_READ`: `FleetMcp.java:1124` maps `fleet_status` to it as well. Option 2 does not close that either, but a live status is a much smaller harm than a reply body. ### Order Option 2 alone leaves no exploitable intermediate state. Option 1 without the presence change is the worst kind of intermediate state — possibly broken rather than exploitable, and nothing in the suite would catch it — so that state must not exist even for one commit. If option 1 ever ships, it ships with the presence condition widened in the same commit, splitting the floor's two jobs: it currently both grants authorization and supplies the liveness signal that marks presence, and option 1 narrows only the first while silently taking the second with it. ### Delegated Unit A (the owner check) is going out now. Its acceptance criteria include the two the architect named that I would not have thought of: an unnamed primary has a `null` terminal and must still read any ticket, and the refusal branch itself needs a control, so that the "another terminal is refused" test cannot pass merely because everything is refused. Two things are explicitly **not** checked and must not be assumed: whether a lead keeps its terminal id across `fleet_handover`, which nobody has read `LeadRollover` to confirm; and the restart-with-live-panes path, which is reasoned from code only. And after this, lead A can no longer poll lead B's ticket. I judge that correct and intended, but it is a behaviour change rather than a side effect, so it is recorded here. The end-to-end probe from an unconfigured tab is **still not run**, by anyone. The ticket's "Not yet measured" section stands.
Author
Owner

Unit A merged in 25d53e6 — and this ticket stays open

PR #712 closes the MCP door. MessageService.Task now records the terminal of the caller whose
fleet_send{wait:false} created the ticket, and poll(ticket, callerTerminal) refuses a different
terminal with a FAILED view carrying no reply text.

Verified by me in a throwaway worktree, output to a file and not piped: exit 0, BUILD SUCCESS,
Tests run: 2008, Failures: 0. The pushed tree is byte-identical to the tree I built.

Two things are still missing, and I am not calling this fixed.

1. The REST door is still open

FleetApp.java:898:

MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"));

That is the no-check overload. The route is gated on TASK_READ
(FleetApp.java:81), and Authz.java:145 grants TASK_READ to a worker. A worker runs on the
same host as the daemon, so curl http://127.0.0.1:8765/tasks/task-1 reaches it. The exact
caller this ticket is about can still walk every ticket, through REST.

I checked both lines myself. The implementer reported this unprompted, which was the right call.

2. The handler wiring is not pinned — measured, not suspected

The three new tests prove the refusal at the seam, MessageService.poll(ticket, callerTerminal).
Nothing proves the MCP handler passes the caller's real terminal into it.

I replaced callerTerminal(exchange) with null at FleetMcp.java:527 using a line-anchored
sed, confirmed mvn compile stayed green so the mutation was live rather than a compile error,
and ran the full suite:

Tests run: 2008, Failures: 0, Errors: 0, Skipped: 0
BUILD SUCCESS

The mutation survives. The whole fix can be switched off by one argument and no test notices.
I restored the file and confirmed it byte-identical before merging.

The repo has the idiom already: FleetMcpAuthzTest scrapes the handler block and asserts the
argument, with a control assertion so it cannot pass by failing to find the block.
theFleetListHandlerActuallyConsultsCollaboratorsVisibleTo is the model.

Why I merged anyway

It is a strict improvement and it breaks nothing. Blocking one because it is not yet total would be
the wrong trade. But a half-closed door must not be recorded as a closed one, so this ticket stays
open until both items above are done. They are one follow-up unit.

One thing worth keeping from the review

ownsTicket lets any caller with a null terminal read every ticket. That is safe only
because
Authz.java:145 denies ANONYMOUS the TASK_READ action, so an anonymous caller never
reaches poll. I enumerated it: CALLER_TERMINAL has exactly one writer, on every call, and every
non-primary Principal in CallerResolver is built with a real terminal — the only terminal-less
constructions are Principal.primary(...) and Principal.anonymous(). So null does not conflate
"the primary" with "I could not tell".

Nothing pins that coupling. Granting ANONYMOUS any read action would silently open the ticket
gate, in a different file from the one that looks like it holds the rule.

Answers to the two questions I required

  1. A lead keeps the same terminal id across fleet_handover. The roll sends /clear and the
    bootstrap text to p.leadTerminal() — the same terminal recorded at open(), never a new pane,
    and confirm() requires p.leadTerminal().equals(callerTerminal) first. So a lead's pre-handover
    tickets stay readable afterwards. Good: the fix does not break the rollover path.
  2. A ticket whose creating pane is gone becomes unreadable to every terminal-bearing caller;
    only the unnamed primary can still read it. This does not hurt the common case, because
    abandon() already fails an open task when its target is torn down. The narrow exposure is a
    named lead whose pane is replaced outside fleet_handover. The implementer could not settle
    whether that is reachable in production config, because fleetd.yaml is gitignored and absent
    from a worker's worktree. Neither have I.

Related, filed separately

#715 — the same "an identifier used as an authorization token" shape on fleet_send{turnId}, which
is a write path. Also reported by this implementer.

Option 1 (the OBSERVER floor) remains undelegated and still must ship with the presence split in
the same commit.

## Unit A merged in `25d53e6` — and this ticket stays open PR #712 closes the **MCP** door. `MessageService.Task` now records the terminal of the caller whose `fleet_send{wait:false}` created the ticket, and `poll(ticket, callerTerminal)` refuses a different terminal with a `FAILED` view carrying no reply text. Verified by me in a throwaway worktree, output to a file and not piped: exit 0, `BUILD SUCCESS`, `Tests run: 2008, Failures: 0`. The pushed tree is byte-identical to the tree I built. **Two things are still missing, and I am not calling this fixed.** ### 1. The REST door is still open `FleetApp.java:898`: ```java MessageService.TaskView v = messages.poll(ctx.pathParam("ticket")); ``` That is the no-check overload. The route is gated on `TASK_READ` (`FleetApp.java:81`), and `Authz.java:145` grants `TASK_READ` to a worker. A worker runs on the same host as the daemon, so `curl http://127.0.0.1:8765/tasks/task-1` reaches it. **The exact caller this ticket is about can still walk every ticket, through REST.** I checked both lines myself. The implementer reported this unprompted, which was the right call. ### 2. The handler wiring is not pinned — measured, not suspected The three new tests prove the refusal at the **seam**, `MessageService.poll(ticket, callerTerminal)`. Nothing proves the MCP handler passes the caller's real terminal into it. I replaced `callerTerminal(exchange)` with `null` at `FleetMcp.java:527` using a line-anchored `sed`, confirmed `mvn compile` stayed green so the mutation was live rather than a compile error, and ran the full suite: ``` Tests run: 2008, Failures: 0, Errors: 0, Skipped: 0 BUILD SUCCESS ``` **The mutation survives.** The whole fix can be switched off by one argument and no test notices. I restored the file and confirmed it byte-identical before merging. The repo has the idiom already: `FleetMcpAuthzTest` scrapes the handler block and asserts the argument, with a control assertion so it cannot pass by failing to find the block. `theFleetListHandlerActuallyConsultsCollaboratorsVisibleTo` is the model. ### Why I merged anyway It is a strict improvement and it breaks nothing. Blocking one because it is not yet total would be the wrong trade. But a half-closed door must not be recorded as a closed one, so this ticket stays open until both items above are done. They are one follow-up unit. ### One thing worth keeping from the review `ownsTicket` lets any caller with a `null` terminal read every ticket. That is safe **only because** `Authz.java:145` denies `ANONYMOUS` the `TASK_READ` action, so an anonymous caller never reaches `poll`. I enumerated it: `CALLER_TERMINAL` has exactly one writer, on every call, and every non-primary `Principal` in `CallerResolver` is built with a real terminal — the only terminal-less constructions are `Principal.primary(...)` and `Principal.anonymous()`. So `null` does not conflate "the primary" with "I could not tell". **Nothing pins that coupling.** Granting `ANONYMOUS` any read action would silently open the ticket gate, in a different file from the one that looks like it holds the rule. ### Answers to the two questions I required 1. **A lead keeps the same terminal id across `fleet_handover`.** The roll sends `/clear` and the bootstrap text to `p.leadTerminal()` — the same terminal recorded at `open()`, never a new pane, and `confirm()` requires `p.leadTerminal().equals(callerTerminal)` first. So a lead's pre-handover tickets stay readable afterwards. Good: the fix does not break the rollover path. 2. **A ticket whose creating pane is gone** becomes unreadable to every terminal-bearing caller; only the unnamed primary can still read it. This does not hurt the common case, because `abandon()` already fails an open task when its target is torn down. The narrow exposure is a named lead whose pane is replaced *outside* `fleet_handover`. The implementer could not settle whether that is reachable in production config, because `fleetd.yaml` is gitignored and absent from a worker's worktree. Neither have I. ### Related, filed separately #715 — the same "an identifier used as an authorization token" shape on `fleet_send{turnId}`, which is a **write** path. Also reported by this implementer. Option 1 (the `OBSERVER` floor) remains undelegated and still must ship with the presence split in the same commit.
Author
Owner

Fixed on both entry paths. Closing.

  • MCP — PR #712, merged as a9a37af. fleet_poll{ticket} threads the caller's terminal.
  • REST — PR #716, merged as 9a64d42. GET /tasks/{ticket} threads it too, and the
    wait:false send path now records a creator terminal.

The REST half was two problems, not one

Closing the read door alone would have broken a working path. FleetApp.sendMessage's wait:false
branch called the sendAsync overload that records no creator, so a REST-created ticket carried
creatorTerminal == null. With the new check, ownsTicket evaluates "term_x".equals(null) for a
terminal-bearing caller, which is false — so the ticket's own creator would have been refused, and
driving the fleet over REST is a documented fallback for when the MCP mount drops. Both halves had to
ship together, and they did.

The mutation that mattered

Before this work, replacing callerTerminal(exchange) with null at FleetMcp.java:527 compiled
green and the full suite still passed — the fix could be switched off and nothing noticed. I re-ran
that exact mutation on the merged tree myself, confirmed it was live with mvn -o compile first, and
it now kills theFleetPollHandlerActuallyThreadsCallerTerminalIntoPoll. Two further mutations on the
REST lines kill one and three tests. Every file restored byte-identical.

Merged build: Tests run: 2018, Failures: 0, BUILD SUCCESS, against a main baseline of 2008 that I
measured myself.

The coupling this fix rests on, which nothing pins

ownsTicket never refuses a caller with no terminal, because that is the unnamed primary and it must
keep reading every ticket. That is safe only because Authz.java:145 denies ANONYMOUS the
TASK_READ action, so an anonymous caller never reaches the lookup. Granting ANONYMOUS any
ticket-read action would reopen this silently. It is written into the new wiki/11-Features.md entry
as a gotcha, since it is the kind of thing a later change breaks without noticing.

Follow-up

#718 — MessageService.poll(String) still exists and defaults the terminal to null, which skips
the check. It has no production caller now, but 44 test callers, so it stays and gets pinned by a
scrape test instead. #718 also covers a second instance of the same shape at FleetMcp.java:975.

## Fixed on both entry paths. Closing. - **MCP** — PR #712, merged as `a9a37af`. `fleet_poll{ticket}` threads the caller's terminal. - **REST** — PR #716, merged as `9a64d42`. `GET /tasks/{ticket}` threads it too, and the `wait:false` send path now records a creator terminal. ### The REST half was two problems, not one Closing the read door alone would have broken a working path. `FleetApp.sendMessage`'s `wait:false` branch called the `sendAsync` overload that records no creator, so a REST-created ticket carried `creatorTerminal == null`. With the new check, `ownsTicket` evaluates `"term_x".equals(null)` for a terminal-bearing caller, which is false — so **the ticket's own creator would have been refused**, and driving the fleet over REST is a documented fallback for when the MCP mount drops. Both halves had to ship together, and they did. ### The mutation that mattered Before this work, replacing `callerTerminal(exchange)` with `null` at `FleetMcp.java:527` compiled green and the full suite still passed — the fix could be switched off and nothing noticed. I re-ran that exact mutation on the merged tree myself, confirmed it was live with `mvn -o compile` first, and it now kills `theFleetPollHandlerActuallyThreadsCallerTerminalIntoPoll`. Two further mutations on the REST lines kill one and three tests. Every file restored byte-identical. Merged build: `Tests run: 2018, Failures: 0`, BUILD SUCCESS, against a `main` baseline of 2008 that I measured myself. ### The coupling this fix rests on, which nothing pins `ownsTicket` never refuses a caller with no terminal, because that is the unnamed primary and it must keep reading every ticket. That is safe **only** because `Authz.java:145` denies `ANONYMOUS` the `TASK_READ` action, so an anonymous caller never reaches the lookup. Granting `ANONYMOUS` any ticket-read action would reopen this silently. It is written into the new `wiki/11-Features.md` entry as a gotcha, since it is the kind of thing a later change breaks without noticing. ### Follow-up #718 — `MessageService.poll(String)` still exists and defaults the terminal to `null`, which skips the check. It has no production caller now, but 44 test callers, so it stays and gets pinned by a scrape test instead. #718 also covers a second instance of the same shape at `FleetMcp.java:975`.
ltms closed this issue 2026-10-04 07:51:01 +02:00
ltms reopened this issue 2026-10-04 07:51:19 +02:00
Author
Owner

Correction: I closed this a moment ago and have reopened it. Option 2 is done; option 1 is not.

My previous comment is accurate about what shipped, but closing the ticket on it was wrong. This
ticket offers three fix shapes, and PRs #712 and #716 implement option 2 — give tickets an owner
check. Option 1, the narrowed bottom rung (the OBSERVER floor), has not been done and is still
an open decision.

What is actually closed

The reported exploit. An unconfigured pane still resolves to WORKER and still holds TASK_READ, so
it can still call fleet_poll{ticket} — but it now reads only tickets its own terminal created, and
it creates none. So walking task-1, task-2, … returns nothing.

What is still open, and why it is worth keeping open

Option 1 is defence that does not depend on the owner check staying correct. Right now the only thing
standing between an unconfigured tab and other sessions' delegation replies is one comparison in
ownsTicket, and #718 exists precisely because that comparison can be bypassed by reaching for a
convenience overload that compiles and passes. A narrowed floor would mean an unconfigured pane never
holds TASK_READ at all, so the overload hazard and the owner check would both have to fail before
anything leaked.

It also still carries the constraint recorded against it: the floor must ship together with the
presence split in the same commit.
Otherwise a pane that outlived a daemon restart becomes
permanently undeliverable. That coupling is the reason this is a design decision and not a one-line
role-table edit, and it is why I am not delegating it yet.

Also still true

The "not yet measured" section above still stands in full. Nobody has opened an unconfigured tab and
polled a ticket from it, before or after this fix. What is proven is the owner check itself, by unit
test and by three mutations I ran on the merged tree. The resolver-ladder claim that an unconfigured
pane becomes a WORKER is still read from the source, not observed.

## Correction: I closed this a moment ago and have reopened it. Option 2 is done; option 1 is not. My previous comment is accurate about what shipped, but closing the ticket on it was wrong. This ticket offers three fix shapes, and PRs #712 and #716 implement **option 2** — give tickets an owner check. **Option 1, the narrowed bottom rung (the `OBSERVER` floor), has not been done** and is still an open decision. ### What is actually closed The reported exploit. An unconfigured pane still resolves to `WORKER` and still holds `TASK_READ`, so it can still call `fleet_poll{ticket}` — but it now reads only tickets its own terminal created, and it creates none. So walking `task-1`, `task-2`, … returns nothing. ### What is still open, and why it is worth keeping open Option 1 is defence that does not depend on the owner check staying correct. Right now the only thing standing between an unconfigured tab and other sessions' delegation replies is one comparison in `ownsTicket`, and #718 exists precisely because that comparison can be bypassed by reaching for a convenience overload that compiles and passes. A narrowed floor would mean an unconfigured pane never holds `TASK_READ` at all, so the overload hazard and the owner check would both have to fail before anything leaked. It also still carries the constraint recorded against it: **the floor must ship together with the presence split in the same commit.** Otherwise a pane that outlived a daemon restart becomes permanently undeliverable. That coupling is the reason this is a design decision and not a one-line role-table edit, and it is why I am not delegating it yet. ### Also still true The "not yet measured" section above still stands in full. Nobody has opened an unconfigured tab and polled a ticket from it, before or after this fix. What is proven is the owner check itself, by unit test and by three mutations I ran on the merged tree. The resolver-ladder claim that an unconfigured pane becomes a `WORKER` is still read from the source, not observed.
Author
Owner

Decision on option 1, after consulting two architects — and the inherited constraint was wrong

I put the question to two architect members on different profiles (sol and opus), with the same
brief and the same evidence, and deliberately withheld my own suspicion so their positions would be
independent. I then measured the crux myself. Recording the decision here per the charter.

The two positions

They returned opposite headline verdicts and then described the same mechanism:

sol opus
Coupling claim "True, but not for the likely Authz reason" "False as stated"
REPLY/ASK survive under OBSERVER? Yes Yes
What actually breaks inbound fleet_send to the surviving pane inbound fleet_send to the surviving pane
Cause presence marking requires isSpawnedMember(); the injector gates on presence identical
Required fix "readiness split" + registration reconciliation; would reject a role-only PR one predicate: add isObserver() to the presence gate
Ship OBSERVER? Yes, defence in depth, not urgent Yes, defence in depth, not next

Both independently reached the ownsSession reasoning I had suspected and withheld. Both
independently found a launch-before-registration window that neither I nor the previous lead had
named. That is genuine corroboration: two different models, two separate readings.

Verdict: the inherited constraint is false as stated, and must not go into a commit

The handover said an OBSERVER floor must ship with a "presence split" or a restart-surviving pane
becomes "permanently undeliverable". Three things are wrong with that, all of which I checked
myself:

  1. "Undeliverable" is the wrong axis. Delivery is gated on AgentStatus.injectable() — the
    pane's terminal status — not on the target's authz role.
  2. The pane can still report. case REPLY, ASK -> caller.ownsSession(targetSession) names no
    role, and ownsSession is terminal != null && terminal.equals(sessionId)
    (Principal.java:133-135). An observer carries its pane's terminal, so it still ends its turn.
  3. A restart already strands the exchange regardless. tasks, asyncTasksByWaiter,
    asyncTasksByTurn, strandedReplies and ticketSeq are all plain in-memory maps, and the
    roster is a plain ConcurrentHashMap with no persistence. After a restart the lead has no
    address for that pane from any tool.

But the constraint pointed at something real, and credit to both architects for finding it. The
loss is the lead's ability to push a new turn, and it is caused by one line:

// FleetMcp.java:832-837
/** Mark a connected spawned member available for the injector readiness gate. */
static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) {
    if (caller.isSpawnedMember()) {          // == WORKER || ARCHITECT  (Principal.java:113-115)
        presence.markPresent(caller.terminal());
    }
}
// Fleetd.java:238-241
return target -> presence.isPresent(target) || leads.get().containsKey(target)
        || collaborators.get().containsKey(target);

An observer is neither a worker nor an architect, so it would never mark presence, and the injector
would never consider it ready.

Where I ruled against sol, by measurement

sol additionally required registration reconciliation for the launch race, and would reject a
role-only PR without it. I checked that path and it is pre-existing and unchanged by this work:

// SessionManager.java:1256-1261 — PresenceFleet.markPresent
super.markPresent(terminal);
sessions.onReady(terminal);

// SessionManager.java:898-899
void onReady(String t) { transitionByTerminal(t, SPAWNING, READY); }

// SessionManager.java:1234-1235
MemberSession current = findByTerminal(terminalId);
if (current == null || current.state() != from) return;   // no registry entry => no-op

So a pre-registration MCP contact marks presence and silently loses the SPAWNING -> READY
transition — today, for a WORKER, exactly as it would for an OBSERVER. Once the observer is
added to the presence gate, that path is byte-for-byte what it is now. It does have a real
consequence (FleetMcp.java:2204 uses state() == READY for seat accounting), so it is worth
fixing — but it is not a prerequisite for this change, and bundling it would hide a pre-existing
defect inside a role addition. Filed separately.

With that removed, the two positions agree on the required behaviour. sol's own wording — "a
successful bridge MCP contact from an observer terminal must mark that terminal safe for injection"
— is opus's one predicate.

Decision

OBSERVER ships, but not next. It is defence in depth, not a fix for a live hole: option 2 is
merged and closed the reported exploit. Its real justification is narrower and still worth it — the
floor currently makes a false identity claim, telling the daemon that anything in a pane is a
spawned member.

Ship #721 first. opus found, and I verified on both the MCP and REST surfaces, that
fleet_status has no owner check and returns another session's pending question text, its
turnId and its ticket to any TASK_READ holder. That is a live content leak, it is cheaper than
this change, and it is independent of this decision. It also voids the mitigation I had claimed on
#715. That ticket now leads.

Required in the same commit as the floor (the union both architects agree on):

  1. Role.OBSERVER, Principal.observer(terminal, pid), isObserver(), and a describe() case —
    describe() has no default, so a missing case is a compile error. Keep that.
  2. Only the final herdr-pane fallback in CallerResolver.resolve changes. The roster rung must
    keep returning WORKER/ARCHITECT — a careless diff breaks that half, and it is the half that
    matters.
  3. Authz: add isObserver() to case READ, METRICS and to nothing else. REPLY/ASK need
    no edit.
  4. Add the observer to the presence gate at FleetMcp.java:834 and rename the method, since it is
    no longer only spawned members.
  5. An explicit isObserver() branch in fleet_whoami before if (!caller.isWorker())
    (FleetMcp.java:1402). Today an observer would reach the lead branch by elimination. The
    output happens to be correct only because m.put("leader", ...) is guarded on a non-null name —
    correct by luck is not correct.
  6. Correct or delete the Fleetd.java:229-234 paragraph claiming presence "doubles as the member
    roster's availability signal". opus enumerated every isPresent reader in main and found
    none that makes it true. A comment asserting a false invariant is what turned this into a "hard
    constraint" in the first place.
  7. CLAUDE.md + the byte-identical wiki template + a wiki/11-Features.md entry. This is a role
    table and ConnectionIdentity change, so the project's own "prompt is part of the product" table
    makes all three mandatory. The fallback ladder has no row for observer.
  8. Tests: an OBSERVER row for every action in AuthzTest; that an observer may REPLY/ASK
    for its own pane and not another's; and a CallerResolverTest case asserting both that an
    unknown pane resolves OBSERVER and that a registered member still resolves
    WORKER/ARCHITECT.

May follow later: narrowing fleet_profiles for an observer; a persisted roster.

Cost warning, which I am recording rather than discounting

opus noted that every role added here has needed follow-up: COLLABORATOR took #669, #703, #710
and then d2db8c7 as a review fix on the PR that merged yesterday. Four corrections for one role.
Treat the diff estimate as a lower bound.

Not measured

Neither architect ran a live probe, and nor did I. Nobody has opened an unconfigured tab and
attempted these calls. sol ran a 173-test subset green and changed nothing; opus ran no build.
Neither read AuthzTest in full, so the number of assertions a new enum constant turns red is
unknown. And nobody has verified that a Claude Code member re-establishes its MCP connection after
a daemon restart
— opus flagged that its whole restart argument assumes it does. If that
assumption is false, window 2 is moot under either floor, and that is one probe worth running before
anyone writes code for item 4.

## Decision on option 1, after consulting two architects — and the inherited constraint was wrong I put the question to two architect members on different profiles (`sol` and `opus`), with the same brief and the same evidence, and deliberately withheld my own suspicion so their positions would be independent. I then measured the crux myself. Recording the decision here per the charter. ### The two positions They returned **opposite headline verdicts** and then described **the same mechanism**: | | `sol` | `opus` | |---|---|---| | Coupling claim | "**True**, but not for the likely `Authz` reason" | "**False as stated**" | | `REPLY`/`ASK` survive under `OBSERVER`? | Yes | Yes | | What actually breaks | inbound `fleet_send` to the surviving pane | inbound `fleet_send` to the surviving pane | | Cause | presence marking requires `isSpawnedMember()`; the injector gates on presence | identical | | Required fix | "readiness split" + registration reconciliation; *would reject a role-only PR* | one predicate: add `isObserver()` to the presence gate | | Ship `OBSERVER`? | Yes, defence in depth, not urgent | Yes, defence in depth, not next | Both independently reached the `ownsSession` reasoning I had suspected and withheld. Both independently found a **launch-before-registration** window that neither I nor the previous lead had named. That is genuine corroboration: two different models, two separate readings. ### Verdict: the inherited constraint is false as stated, and must not go into a commit The handover said an `OBSERVER` floor must ship with a "presence split" or a restart-surviving pane becomes "**permanently undeliverable**". Three things are wrong with that, all of which I checked myself: 1. **"Undeliverable" is the wrong axis.** Delivery is gated on `AgentStatus.injectable()` — the pane's terminal status — not on the target's authz role. 2. **The pane can still report.** `case REPLY, ASK -> caller.ownsSession(targetSession)` names no role, and `ownsSession` is `terminal != null && terminal.equals(sessionId)` (`Principal.java:133-135`). An observer carries its pane's terminal, so it still ends its turn. 3. **A restart already strands the exchange regardless.** `tasks`, `asyncTasksByWaiter`, `asyncTasksByTurn`, `strandedReplies` and `ticketSeq` are all plain in-memory maps, and the roster is a plain `ConcurrentHashMap` with no persistence. After a restart the lead has no address for that pane from any tool. **But the constraint pointed at something real**, and credit to both architects for finding it. The loss is the lead's ability to **push a new turn**, and it is caused by one line: ```java // FleetMcp.java:832-837 /** Mark a connected spawned member available for the injector readiness gate. */ static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) { if (caller.isSpawnedMember()) { // == WORKER || ARCHITECT (Principal.java:113-115) presence.markPresent(caller.terminal()); } } ``` ```java // Fleetd.java:238-241 return target -> presence.isPresent(target) || leads.get().containsKey(target) || collaborators.get().containsKey(target); ``` An observer is neither a worker nor an architect, so it would never mark presence, and the injector would never consider it ready. ### Where I ruled against `sol`, by measurement `sol` additionally required **registration reconciliation** for the launch race, and would reject a role-only PR without it. I checked that path and it is **pre-existing and unchanged** by this work: ```java // SessionManager.java:1256-1261 — PresenceFleet.markPresent super.markPresent(terminal); sessions.onReady(terminal); // SessionManager.java:898-899 void onReady(String t) { transitionByTerminal(t, SPAWNING, READY); } // SessionManager.java:1234-1235 MemberSession current = findByTerminal(terminalId); if (current == null || current.state() != from) return; // no registry entry => no-op ``` So a pre-registration MCP contact marks presence and silently loses the `SPAWNING -> READY` transition — **today, for a `WORKER`, exactly as it would for an `OBSERVER`.** Once the observer is added to the presence gate, that path is byte-for-byte what it is now. It does have a real consequence (`FleetMcp.java:2204` uses `state() == READY` for seat accounting), so it is worth fixing — but it is **not a prerequisite for this change**, and bundling it would hide a pre-existing defect inside a role addition. Filed separately. With that removed, the two positions agree on the required behaviour. `sol`'s own wording — "a successful bridge MCP contact from an observer terminal must mark that terminal safe for injection" — *is* `opus`'s one predicate. ### Decision **`OBSERVER` ships, but not next.** It is defence in depth, not a fix for a live hole: option 2 is merged and closed the reported exploit. Its real justification is narrower and still worth it — the floor currently makes a **false identity claim**, telling the daemon that anything in a pane is a spawned member. **Ship #721 first.** `opus` found, and I verified on both the MCP and REST surfaces, that `fleet_status` has no owner check and returns another session's pending **question text**, its `turnId` and its `ticket` to any `TASK_READ` holder. That is a live content leak, it is cheaper than this change, and it is independent of this decision. It also voids the mitigation I had claimed on #715. That ticket now leads. **Required in the same commit as the floor** (the union both architects agree on): 1. `Role.OBSERVER`, `Principal.observer(terminal, pid)`, `isObserver()`, and a `describe()` case — `describe()` has no `default`, so a missing case is a compile error. Keep that. 2. **Only** the final herdr-pane fallback in `CallerResolver.resolve` changes. The roster rung must keep returning `WORKER`/`ARCHITECT` — a careless diff breaks that half, and it is the half that matters. 3. `Authz`: add `isObserver()` to `case READ, METRICS` and to **nothing else**. `REPLY`/`ASK` need no edit. 4. Add the observer to the presence gate at `FleetMcp.java:834` and rename the method, since it is no longer only spawned members. 5. An explicit `isObserver()` branch in `fleet_whoami` **before** `if (!caller.isWorker())` (`FleetMcp.java:1402`). Today an observer would reach the **lead** branch by elimination. The output happens to be correct only because `m.put("leader", ...)` is guarded on a non-null name — correct by luck is not correct. 6. Correct or delete the `Fleetd.java:229-234` paragraph claiming presence "doubles as the member roster's availability signal". `opus` enumerated every `isPresent` reader in `main` and found none that makes it true. A comment asserting a false invariant is what turned this into a "hard constraint" in the first place. 7. `CLAUDE.md` + the byte-identical wiki template + a `wiki/11-Features.md` entry. This is a role table and `ConnectionIdentity` change, so the project's own "prompt is part of the product" table makes all three mandatory. The fallback ladder has no row for `observer`. 8. Tests: an `OBSERVER` row for **every** action in `AuthzTest`; that an observer may `REPLY`/`ASK` for its own pane and not another's; and a `CallerResolverTest` case asserting both that an unknown pane resolves `OBSERVER` **and** that a registered member still resolves `WORKER`/`ARCHITECT`. **May follow later:** narrowing `fleet_profiles` for an observer; a persisted roster. ### Cost warning, which I am recording rather than discounting `opus` noted that every role added here has needed follow-up: `COLLABORATOR` took #669, #703, #710 and then `d2db8c7` as a review fix on the PR that merged yesterday. Four corrections for one role. Treat the diff estimate as a lower bound. ### Not measured Neither architect ran a live probe, and nor did I. Nobody has opened an unconfigured tab and attempted these calls. `sol` ran a 173-test subset green and changed nothing; `opus` ran no build. Neither read `AuthzTest` in full, so the number of assertions a new enum constant turns red is unknown. And **nobody has verified that a Claude Code member re-establishes its MCP connection after a daemon restart** — `opus` flagged that its whole restart argument assumes it does. If that assumption is false, window 2 is moot under either floor, and that is one probe worth running before anyone writes code for item 4.
Author
Owner

Refining item 6 of the decision — half that paragraph is true, so do not delete it wholesale

I attributed the claim about Fleetd.java:229-234 to the architect rather than measuring it, and
item 6 said "correct or delete the paragraph". I have now checked it myself, and the instruction was
too blunt: the paragraph makes two claims and only one is false.

The paragraph:

Neither a lead nor a collaborator is ever enrolled in MemberPresence — FleetMcp marks presence
for every spawned member (worker and architect), deliberately, since that map doubles as the
member roster's availability signal and a lead or collaborator counted there would show up as an
available member
. So without the second and third disjuncts a lead or collaborator is permanently
un-deliverable: every send to one sat on the gate for READINESS_GRACE_POLLS (~60s) and then
failed having never been typed into the pane.

Claim 1 — the stated reason — is false. Every reader of MemberPresence.isPresent in
fleetd/src/main/java, after filtering out unrelated Optional.isPresent() hits:

Fleetd.java:240          presence.isPresent(target) || leads... || collaborators...   <- the gate
FleetApp.java:138        presence::isPresent   (passed as the same deliverability predicate)
FleetApp.java:150        presence::isPresent   (same, the auth/metrics constructor)
MemberPresence.java:30   the definition itself

Three call sites, all one predicate: deliverability. Nothing reads presence as "an available member
in the roster."
Seat and capacity accounting come from elsewhere — reclaimable() keys on
session.state() == READY (FleetMcp.java:2204), and the roster itself comes from
sessions.roster(). So "a lead or collaborator counted there would show up as an available member"
has no reader behind it. It may have been true when written.

Claim 2 — the consequence — is true and load-bearing. Without the leads and collaborators
disjuncts a lead or collaborator really is undeliverable, because the gate is exactly those three
disjuncts. That sentence must stay.

Revised item 6

Correct only the causal clause — the "since that map doubles as…" reason. Keep the sentence about
the second and third disjuncts. And when you remove the false reason, say what actually pins the
exclusion now, or the next reader will reinstate it: per this project's comment rules, name the
constraint and what breaks if someone changes the line, not the history.

This also matters beyond a comment. That false reason is why the handover recorded a "hard
constraint"
in the first place: a reader looking at it concludes that adding anyone to presence has
roster side effects, so the one-predicate fix looks unsafe and a "presence split" looks mandatory. A
comment asserting an invariant nothing enforces cost us a wrong constraint, two architect turns and a
round of my own verification. It is a free test case — the claim was checkable in one grep.

## Refining item 6 of the decision — half that paragraph is true, so do not delete it wholesale I attributed the claim about `Fleetd.java:229-234` to the architect rather than measuring it, and item 6 said "correct or delete the paragraph". I have now checked it myself, and the instruction was too blunt: the paragraph makes **two** claims and only one is false. The paragraph: > Neither a lead nor a collaborator is ever enrolled in `MemberPresence` — `FleetMcp` marks presence > for every spawned member (worker and architect), deliberately, since **that map doubles as the > member roster's availability signal and a lead or collaborator counted there would show up as an > available member**. So without the second and third disjuncts a lead or collaborator is permanently > un-deliverable: every send to one sat on the gate for `READINESS_GRACE_POLLS` (~60s) and then > failed having never been typed into the pane. **Claim 1 — the stated reason — is false.** Every reader of `MemberPresence.isPresent` in `fleetd/src/main/java`, after filtering out unrelated `Optional.isPresent()` hits: ``` Fleetd.java:240 presence.isPresent(target) || leads... || collaborators... <- the gate FleetApp.java:138 presence::isPresent (passed as the same deliverability predicate) FleetApp.java:150 presence::isPresent (same, the auth/metrics constructor) MemberPresence.java:30 the definition itself ``` Three call sites, all one predicate: deliverability. **Nothing reads presence as "an available member in the roster."** Seat and capacity accounting come from elsewhere — `reclaimable()` keys on `session.state() == READY` (`FleetMcp.java:2204`), and the roster itself comes from `sessions.roster()`. So "a lead or collaborator counted there would show up as an available member" has no reader behind it. It may have been true when written. **Claim 2 — the consequence — is true and load-bearing.** Without the `leads` and `collaborators` disjuncts a lead or collaborator really is undeliverable, because the gate is exactly those three disjuncts. That sentence must stay. ### Revised item 6 Correct **only the causal clause** — the "since that map doubles as…" reason. Keep the sentence about the second and third disjuncts. And when you remove the false reason, say what actually pins the exclusion now, or the next reader will reinstate it: per this project's comment rules, name the constraint and what breaks if someone changes the line, not the history. This also matters beyond a comment. That false reason is **why the handover recorded a "hard constraint"** in the first place: a reader looking at it concludes that adding anyone to presence has roster side effects, so the one-predicate fix looks unsafe and a "presence split" looks mandatory. A comment asserting an invariant nothing enforces cost us a wrong constraint, two architect turns and a round of my own verification. It is a free test case — the claim was checkable in one `grep`.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#705