ConnectionIdentity's Caller record carried the same wrong rule as the two
places PR #760 fixed: it called the terminal a worker's, and read a null
terminal as the primary. The brief for #760 named only two of the three spots.
MemberPresence pointed at FleetMcp.markTrackedCallerPresent by name inside
{@code}, which the compiler does not check, and inject has no dependency on
mcp so a {@link} would add a cross-package reference. States the principle
instead, which stays true whichever roles qualify. Drops a ticket key.
paneRole now reads CallerResolver#boundToArchitectSlot (made public, no
second definition) so a slot-bound pane with no live member reports
"architect", matching what sendableObserverTarget already allowed as a
SEND target.
panesVisibleTo now admits an observer, since an observer holds SEND to
another observer pane. Its panes rows are filtered to
sendableObserverTarget and reduced to sessionId/label/status/role/
deliverable; every other caller's rows are unchanged.
Role.PRIMARY, ConnectionIdentity (class javadoc + callerTerminal), and
MemberPresence's class javadoc each state a role model the code no longer
implements. Comment-only change.
Three claims in the canonical block went false when the observer SEND grant
merged, and this file is the instruction surface the bridge ships:
- the observer role definition said "never SEND"
- invariant 3 listed send as "lead, architect, or collaborator"
- the hand-opened-pane row said such a pane "cannot fleet_send back"
The grant is narrow. Authz permits an observer's SEND only when
CallerResolver.sendableObserverTarget() classifies the target as an observer
too, so a lead, a collaborator, an architect slot and a spawned member are
each refused. TASK_READ stays denied, so #705 remains closed.
Measured on the merged tree: mvn clean install exit 0, 2137 tests from 178
surefire files, 0 failures. Mutating away the grant's call site
(FleetMcp.java:474) kills exactly FleetMcpObserverSendDeliveryTest, so the
behaviour is pinned and not only the predicate.
A workspace.list/tab.list failure in the label scan no longer costs the
caller the agent roster (GET /agents) or the leads/members/capacity/
coordinator rows (fleet_list) that never needed it. Both call sites now
fall back to an empty label map on HerdrException, so a pane row still
renders with label:null instead of the whole response failing.
Also cuts four comments down to the current contract, per the project's
comment rule: dropped the reviewer-facing justification from Fleetd's
deliverableTo javadoc, panesVisibleTo's javadoc, the fleet_list handler's
inline comment, and the panesVisible assembly-gate comment, and removed
the two fragments describing what a test must do.
FleetMcpObserverSendDeliveryTest already kills the same mutation end to
end (it asserts the exact attributed text herdr receives), and the code
quality rule in CLAUDE.md caps new source-text tests in this file at the
existing count.
Authz.SEND now grants an observer a narrow path: it may reach only a
target that CallerResolver.sendableObserverTarget() would itself
resolve as OBSERVER, never a lead, a collaborator, or a live spawned
member's terminal. This mirrors the existing collaborator SEND clause
rather than adding an unconditional caller.isObserver() grant, which
fleetd #705 already rejected as too broad.
Because the receiving pane cannot otherwise tell an observer's SEND
apart from a human paste, FleetMcp.attributeIfObserver prefixes the
delivered text with the sender's own daemon-resolved terminal on both
the MCP and REST entry paths, for exactly this one new path.
FleetAppAuthTest's start() helper never wired a spawnedMemberRole, so
its "worker" fixture actually resolved as OBSERVER under the real
CallerResolver -- invisible before because OBSERVER and WORKER shared
the same (zero) SEND grant. Granting OBSERVER a real SEND surfaced it:
two SEND-denial tests started passing for the wrong caller. Fixed the
fixture to resolve term_a as a live DEV, matching the helper's own
documented contract.
GET /agents now carries each agent's tab label, merged in from the same
herdr daemon(s) the roster is drawn from. fleet_list gains a panes array
with the same label, the sessionId fleet_send takes as a target, the
role the daemon resolves that pane as, and the deliverable gate the
injector itself enforces (Fleetd#deliverableTo, now public). Gated like
leads/collaborators (primary, architect, collaborator), not bare READ,
since a tab label and a member's cwd are not roster facts every READ
caller may see.
The lead's intent->tool table had no row for messaging an unconfigured pane,
and the user-scope instruction file said outright that the fleet has no route
to an interactive session unless an operator registers it as a collaborator.
That claim is false and it is load-bearing: a session reading it concludes the
exchange is impossible and stops, which is what happened here.
Delivery is gated on presence, not on SEND. contextExtractor runs on every MCP
request including initialize, markTrackedCallerPresent enrols an observer into
MemberPresence, and deliverableTo tests presence before the lead and
collaborator maps. So connecting the server is the enrolment, and fleet_reply
is gated on owning your own pane, which every pane does.
Add the table row, and add the enrolment side of the deliverability gate to the
flows page next to the existing "a spawned member is not deliverable until it
has mounted the MCP" bullet, which is the same gate read the other way.
Measured on a real pane, not a fake: trinotes answered with no fleet config, no
restart, and its fleet_* tools still deferred and unloaded.
errorPatternCoverageLine pointed at exhaustedPatternCoverageLine's javadoc
"for the measured swap mutation this pairing guards against". That narrative
was removed from the destination, so the pointer led nowhere.
The pairing with BUILT_IN_DEFAULT is the contract and it stays. What the
mutation proved belongs in history, not in the comment.
ignoreCycle() used to exempt every dependency between two packages, in both
directions, for the whole package. That meant a brand new dependency added
later between an already-excepted pair (auth/mcp, mcp/msg, inject/msg,
metrics/msg, msg/session) was silently exempted too, exactly where the
msg package makes the gate matter most.
Replace the package-wide ignore with a frozen baseline of the 45 exact
origin-class -> target-class edges that exist today between those five
pairs, and ignore only those via SliceRule.ignoreDependency(String, String).
A new dependency between a baselined pair is not in that set, so it is no
longer ignored and the existing beFreeOfCycles() check (or, when the new
edge alone would not form a cycle, a dedicated set-equality check) fails
and names the exact origin class, target class and package pair.
The set-equality check also fails on a baseline entry whose dependency no
longer exists in the code, so a removed edge cannot rot in the baseline
and mask the pair's eligibility for the ticket #131 removal steps. Rewrote
the javadoc to describe only the current contract.
Drop 2 dead *WiringTest names from 3 comment sites (renamed to
*AssemblyTest), rewriting each sentence to state the code's guarantee
instead of naming a test class. Reattach 5 javadoc blocks that were
orphaned behind a second /** block to the member they actually
describe, trimming history/evidence text down to the current contract
per the project's comment rule. Comment-only; no production logic
changed.
Two architects reviewed the codebase independently on separate backends and
both returned MIXED, not "a mess": 0 public mutable fields, 12 extends of
which 8 are exceptions, 120 records, and 2 files over 1000 code lines rather
than the 9 a total-line count suggests. Both rejected a Clean Code section,
a SOLID list, a pattern catalogue, class and method length limits, and a
coverage gate as text that would change no behaviour.
The comment rule they were briefed against was never in this repo. It lives
in the operator's personal global config, so it was not in git, never
reviewed, and not versioned with the code. Its "goes in the ADR" line was
unfollowable because this project has no ADR, which sent load-bearing
knowledge to a destination that does not exist. Rule 2 redirects it to
docs/<subject>.md, which exists.
Rule 4 covers what neither architect ranked first and both described: a
testing problem relieved by reshaping production code. 54 static factories
on Fleetd, 7 volatile race hooks in MessageService, 8 tests asserting on
main source as text, and a 452-line composition root across 8 tickets.
Rule 4's judgement half is marked as having no mechanism. A cap stops a
count growing; no test distinguishes a good decomposition from a bad one.
Verified: canonical block still byte-identical with wiki/7-Use-Cases.md.
fleet_handover{open} now returns outstandingTickets and openAsks. The
successor keeps the authority to poll and answer them but not the ids, and
an uncollected terminal ticket loses its reply at the ticket TTL.
A DONE/FAILED ticket nobody has polled yet is destroyed on a timer by
pruneTerminalTickets' completion-based TTL, while a PENDING ticket is not
going anywhere. Drop the !future.isDone() filter so outstanding() reports
every ticket the caller owns that is still in tasks, with its real phase.
rollingByTerminal keyed the single-flight lock on p.leadTerminal(), the pane
address. A roll replaces the pane, so a second roll of the same lead opened
from the new terminal landed on a different map key and could run concurrent
with the first roll's still-in-flight continuation.
PendingRollover now carries rolloverKey, resolved once in open() from
leadNameForTerminal while the lead is certainly still live, falling back to
the terminal itself when the name resolves null or blank. confirm()'s claim
and both release sites (the continuationRunner-rejection catch and
runRollover's finally) use the carried key instead of recomputing it, since
leadNameForTerminal no longer resolves the old terminal by release time.
NOT_YOUR_ROLLOVER and the pane-teardown calls stay keyed on p.leadTerminal(),
unchanged.
rollingByTerminal is renamed rollingByLead to match.
MessageService.outstanding(callerOwner) lists the caller's non-terminal
delegations and the subset paused in fleet_ask, filtered by the same
ownsTicket rule poll() already uses. FleetMcp threads the caller's owner
key into handover/handoverOpen and adds outstandingTickets/openAsks to
the open() JSON, alongside the unchanged token/handoverPath/requestedAtMillis.
The heartbeat loop now resolves its nudge target through
PrimaryRegistry#currentPrimaryTerminal(), so the sentence describing what a
background loop with no caller uses named the raw accessor instead.
The comment described the pre-#737 rule, where a null caller key read every
ticket. The pendingAsk gate now matches the unnamed primary's key like any
other, so the exemption it named no longer exists.
ReplyPushLoop.resolveLiveLead dropped a dead per-target delegation and
retried PrimaryRegistry.nudgeTargetFor, but returned that fallback without
checking isLive. The fallback is now probed the same way the first lead is,
and the method returns empty rather than trust a dead terminal.
PrimaryRegistry records a delegating lead's name alongside its learned
terminal and resolves the name back to its current terminal at nudge time,
through a name-to-terminal lookup backed by the live lead-tab scan. A name
with no current match falls back to the terminal that was actually learned,
so an unnamed primary, an off-host lead, or a non-herdr lead keeps working
exactly as before. LeadHeartbeatLoop now reads the resolved current terminal
instead of the raw learned one.
ownsTicket treated a null callerOwner as "read everything", conflating
the internal no-check test seam with a real unnamed primary's owner
key. Give them two different values: a named INTERNAL_NO_OWNER_CHECK
marker for the test-only bypass, and null as just another owner key
that must equal the ticket's recorded creatorOwner (including a
null-to-null match, so an unnamed primary still owns its own tickets).
Applies to both ownsTicket call sites, poll and pendingAsk.
The roll now ends the pane and launches a fresh process instead of typing
/clear, and that jar is deployed, so the skill's self-dated "a change is
coming" note had to go.
Measured on 2026-10-05 before writing:
grep -c "lead-rollover: rolled" fleetd/fleetd.out -> 20
grep -c "lead-rollover:" fleetd/fleetd.out -> 86 (control)
<tail from the last "fleetd listening"> | grep -c "lead-rollover:" -> 0
grep -c 'RELAUNCH_NEVER_READY\|RELAUNCH_NOT_RECOGNISED\|OLD_PANE_NEVER_DIED' -> 0
So all 20 recorded rolls ran under /clear and the restart path has never
executed. The section says that rather than implying the old numbers
describe it.
Also:
- name all eight RollState outcomes, with what each one guarantees
- state that relaunchReadySeconds bounds each of two waits, not the pair
- drop the "never observed as WORKING after 8 consecutive IDLE/DONE polls"
paragraph: grep finds that wait is deleted, so it cannot appear
- drop the #621 warning: contextNotice now takes requireOperatorConfirm
- split the surprise bullets into their own section
The javadoc on RollState.IN_PROGRESS enumerates the terminal states the entry
can be overwritten with, and omitted RELAUNCH_NEVER_READY. That state is
reachable at LeadRollover.java:710, so the list told a reader a state could not
occur when it can. Found by a reviewer on PR #742, outside its assigned scope.
Comment only; no behaviour change.
The assembled daemon's own boot-time orphan-worker reap makes a real call on
the member fake before any roll starts. Clear the member fake's recorded
calls once assembly finishes and before the roll begins, so the assertion
measures calls made since the roll started rather than the whole process's
lifetime, and reword its message to say so.
LeadRollover already sends every call through the lead-bound AgentControl and
WorkspaceControl it receives at construction, so no production code needed a
routing fix. Add a regression test that drives a full relaunch to the point
where recognition times out and asserts bootstrapText still lands on the lead
daemon and never on the member daemon, the one path the existing assembly test
never reaches.
LeadRollover's deferred continuation now ends the old lead's pane, relaunches
a fresh one, and bootstraps it, instead of sending /clear into the same
process. The relaunch step runs two separate bounded waits instead of one
combined check: a readiness wait (the fresh pane reaches a real turn
boundary) is the safety gate and withholds bootstrapText on timeout
(RELAUNCH_NEVER_READY); a recognition wait (the fresh terminal shows up in
the live-lead map) is bookkeeping only, so a timeout there still lets
bootstrapText go out (RELAUNCH_NOT_RECOGNISED). clearSettleSeconds is
retired in favor of relaunchReadySeconds (default 45), which bounds both
waits. Updates FleetConfig/FleetMcp operator-facing text to match.
Add Principal.ownerKey(): role-prefixed, keyed on name for a named lead and
a collaborator (survives a handover's terminal change), on terminal for a
worker, architect and observer, and on a distinct "anonymous" value for an
unauthenticated caller so that case no longer relies on Authz refusing it
first. The unnamed primary keeps a null key, preserving its primary-wide
ticket rule.
Thread that key through Task.creatorOwner, Rendezvous.Owner, poll,
pendingAsk and answer in place of a raw terminal, in both the MCP and REST
surfaces, so a named lead whose terminal changes can still poll and answer
its own delegations while a different lead is refused both.
Mutation evidence (each one-line change killed a named test, then reverted
to green):
- Principal.ownerKey() PRIMARY case made unconditional (dropped the
null-name guard) -> ownerKeyCoversEveryRole dies:
"expected: <null> but was: <leader:null>"
- OBSERVER case changed to use the "worker" prefix -> ownerKeyCoversEveryRole
dies: "expected: <observer:term_observer> but was: <worker:term_observer>"
- prefixed() changed to drop the role prefix entirely -> both
ownerKeyCoversEveryRole and rolePrefixesKeepLeadAndArchitectKeysDistinct
die on a lead/architect key collision: "expected: <leader:opus> but was:
<opus>"
- PRIMARY case changed to key on terminal instead of name ->
rolePrefixesKeepLeadAndArchitectKeysDistinct and ownerKeyCoversEveryRole
die: "expected: <leader:opus> but was: <leader:term_lead>"
- ARCHITECT case changed to key on the slot name instead of terminal ->
same two tests die: "expected: <architect:opus> but was: <architect:design>"
All five mutations were caught by the existing test suite; no test needed
adding.
SessionManager.releaseRemoved() tore down a member's registry row and pane
but never cleared it from MemberPresence, so a terminal stayed marked
"present" for the daemon's lifetime after release/idle-reap/shutdown drain.
Clear it in the method's unconditional finally block, alongside the other
must-always-run teardown step, so every release path (explicit release,
the idle reaper's releaseIfCurrent, and a shutdown drain) forgets it the
same way, and a throw from the dirty-worktree check does not skip it.
MemberPresence.forget(null) throws NullPointerException (verified empirically:
ConcurrentHashMap.remove(null) NPEs on key.hashCode()), so the new call guards
on a non-null, non-blank terminal id rather than relying on forget to no-op.
Unit 2 replaces the /clear continuation with a real process restart, so every
paragraph in the skill that describes /clear goes false the moment the new jar
is deployed. The code is not merged yet, and a merge is not a deployment, so
rewriting those paragraphs now would hand a lead doing a handover tonight a
document that does not match the daemon it is talking to.
Add a dated note instead. It states that the /clear text stays accurate while
the old jar runs, and gives a test a lead can apply with no shell: the
fleet_handover tool description is served by the running daemon, so if it still
says "clear your pane", the old behaviour is live. It also names the two things
that change, including the one that doubles as a second indicator -- the
"never observed as WORKING after 8 consecutive IDLE/DONE polls" warning cannot
appear once the wait that logs it is deleted. The note names the condition for
deleting itself.
Re-measure the roll evidence while here. The skill recorded four
"lead-rollover: rolled" lines from 2026-09-22; the log now holds 20, against a
control of 86 "lead-rollover:" lines, and "Unknown command" still returns 0.
Add the elapsed spread (median 16507 ms, max 48261 ms, two above 45000 ms) with
the caveat that it times the whole roll and is dominated by the wait for the
calling turn to end, so a slow roll is not a failed one.
Markdown only, no code touched, so no build was run.
The comment said ticket ids are a sequential counter with no owner check, so a
holder could walk every ticket and read another session's reply. PRs #712 and
#716 added that owner check: MessageService.ownsTicket compares a ticket's
creatorTerminal to the caller on every read.
The rule is still right, so only the reason changes. This matters now because
fleetd #737 is deciding ticket ownership across a lead handover, and a reader
who believed the old text could delete the TASK_READ restriction on the grounds
that its stated reason no longer applies.
Comment-only. mvn -o clean install: Tests run: 2083, Failures: 0, Errors: 0,
BUILD SUCCESS. Flagged by the #705 option-1 worker as out of its scope, which
was the right call.
Adds Role.OBSERVER as the bottom rung CallerResolver falls to when a
herdr pane matches no live roster entry, lead, architect slot, or
collaborator tab. An observer may only READ/METRICS and REPLY/ASK on
its own pane. Widens the presence gate so an observer's MCP contact
still marks it deliverable, matching what already happens for a
worker or architect, so a pane that outlives a daemon restart is not
left permanently undeliverable.
Ships as defence in depth alongside the already-merged ticket-owner
check (#712/#716), which closed the reachable exploit this ticket
reported.
A member whose MCP contact lands between launcher.spawn() and registry.put()
had its presence marked, but the SPAWNING -> READY transition that markPresent
triggers found no registry entry yet and silently did nothing. The mark then
persisted while registration left the session in SPAWNING, with nothing to
retry the transition. That left the session undeliverable to reclaim/seat
accounting even though it was present and deliverable.
Add SessionManager.reconcilePresence, called right after registry.put in both
the plain-spawn and worktree-spawn paths, to retry the transition for a
terminal already marked present. One private helper serves both call sites.
Tests cover both orderings (contact-then-register and register-then-contact)
for both spawn paths, plus a terminal never marked present staying in
SPAWNING. The contact-then-register tests use a new PresenceRacingLauncher
test double that marks presence from inside spawn(), before acquire()'s own
registry.put runs.
confirm() takes the per-lead-terminal claim before handing the roll to
continuationRunner, and the only release path was runRollover's own
finally. If continuationRunner.accept itself throws, runRollover never
starts, so that finally never runs, and nothing else ever writes
rollingByTerminal — the claim is held forever and the terminal can never
be rolled again. This differs from fleetd #615, which covers a throw
INSIDE the continuation (runRollover already catches that and still
releases the claim) — this is a throw from the hand-off itself, which
is not reachable with today's virtual-thread runner but would be with
a bounded executor's RejectedExecutionException.
confirm() now catches that throw, releases the claim, and overwrites
the IN_PROGRESS outcome with a terminal FAILED one, matching how a
throw inside the continuation is already surfaced.