A send to a pane that never frees up was accepted, then polled as
"pending — worker working" forever, with no WARN and no timeout. That
text is only ever correct for a message the injector already handed
off; a message still sitting in its queue now reports as queued, with
the target's live status and how long it has been waiting.
Injector.enqueue() now stamps Pending.enqueuedAtMillis so
queuedWaitMillis(target) can tell "never attempted" apart from
"delivered, now being worked on". MessageService.pendingDetail() uses
it to pick the poll wording. FleetMcp.sendAsync() adds a best-effort
warning to the accept text when the target is not injectable at send
time, per the brief's stated preference for a warning over a hard
refusal (a short busy spell is normal).
Injector gets a new bound, QUEUE_WAIT_GRACE_POLLS (4800 polls, ~20min
at the 250ms poll interval), mirroring READINESS_GRACE_POLLS: a
message that sits at the head of the queue with no delivery attempt
for that long fails via onTurnFailed, the same path a readiness
timeout uses, without touching presence — the target is busy, not
gone.
Fixed InjectorTest.readinessGraceExpiryLogsTheMeasuredElapsedTimeNotArithmeticOnConstants's
stub clock, which now sees one extra, legitimate nowMillis read from
enqueue()'s new stamp.
Tests: aQueuedButNeverInjectedMessageDoesNotPollAsWorking +
aDeliveredMessageStillPollsAsWorkerWorking (poll wording, with
positive control); failsAQueuedMessageWhoseTargetNeverFreesUp +
aTargetThatFreesUpBeforeTheQueueWaitGraceIsDeliveredNormally (timeout,
with positive control).
The flag tested leader.tab() and was named for it. It now tests whether
fleet.leaders has any entry at all, so the old name states a condition the
code no longer checks.
The ticket's correction comment pointed out the refusal message still
implied removing a lead's tab helps, when only placement: tab does.
Restate the message so the lead and collaborator remedies are not
conflated, and keep the variable name the correction specified.
The guard's lead half now fires whenever fleet.leaders has any entry,
since a lead's tab is always labelled by the fixed LEAD_TAB_LABEL
constant regardless of its own deprecated tab: field. The refusal
message no longer advises removing a lead's tab:, which cannot
satisfy the guard any more.
The redeploy skill's check 4 and the script's closing hint both told the
operator that a lead is found by fleet.leaders.*.tab. Identity is now the
fixed 'lead' label together with the lead's configured workspace, so both
would have sent a reader to a key that no longer decides anything.
The lead tab label becomes a fixed constant (Leader.LEAD_TAB_LABEL = "lead");
fleet.leaders.<name>.tab is now optional legacy, matched case-insensitively
alongside the constant via Leader.acceptedLabels(). The uniqueness boundary
between leads moves from the exact tab text to the workspace: FleetConfig
refuses two leaders that share a workspace, LeadTabScanner indexes lead
labels per space (collaborators stay space-agnostic), and
LeadLauncher.leadNameOf/countLeads require both the accepted label and the
lead's own space to match, so a legacy-labelled tab in the wrong space never
counts and a daemon restart never double-spawns a second lead next to a live
one. Config validation also refuses a fleet.tabLabel template or a
collaborator tab that can render as the fixed lead label.
Adds workspaceLabel next to workspaceId on fleet_list's panes row, read
from herdr's workspace.list via a new PaneLocator.workspaceLabelsByWorkspaceId().
An observer's reduced row still excludes it; a herdr failure or an unknown
workspaceId yields workspaceLabel: null without costing the rest of fleet_list.
The gate matched the box line as "│ >", which appears zero times on a current
Claude Code pane. EMPTY was unreachable, so every lead nudge was held forever.
Match the box as a line *starting* with "❯" or with "│ >", keeping the older
bordered layout readable. A marker further along a line is transcript text — a
caret the operator quoted — so it no longer counts, and the last matching line
is still the live box because the detection region carries scrollback above it.
Look for the generating marker only from the box line down, for the same
reason: an earlier turn's "esc to interrupt" survives in that scrollback, and
holding on it would be the same unreachable-EMPTY failure by another route.
Fixtures: add IDLE_PROMPT_CARET / DRAFTED_PROMPT_CARET from a live pane and
point every lead-pane fake at them. The bordered constants stay, now covering
the older client. Against the old marker these fixtures fail 55 tests across
PromptBoxTest and the three loop tests, which is the production bug reproduced.
cd fleetd && mvn clean install -> BUILD SUCCESS, MVN_EXIT=0,
Tests run: 2164, Failures: 0, Errors: 0, Skipped: 0 (179 surefire XML files).
herdr's agent.prompt pastes AND submits in one call, so a nudge arriving
while the operator is mid-sentence submitted their unfinished line with the
nudge glued to it. AgentStatus.injectable() cannot see this: it describes
the agent, and an idle agent reports the same status whether its input box
is empty or holds a half-typed line.
New herdr/PromptBox reads the pane's `detection` region — the same region
StatusRefiner uses, and the one the input box is drawn in — and clears a
delivery only when the box is positively empty. A box with characters, a
pane it cannot recognise, and a failed read all hold, because a held nudge
is recoverable and a submitted half-line is not. Whitespace and a cursor
block count as empty. After 20 consecutive holds for one target it logs one
warning, so a box that never clears is visible rather than silent; the
warning repeats only after the box has cleared again.
Wired into the three paths that nudge a lead's own pane:
- ReplyPushLoop.decide -> WAIT_BUSY (the pending work is re-read next tick)
- LeadHeartbeatLoop.tick -> a new DRAFT_HELD action that spends neither the
quiet budget nor the one context notice per HIGH stretch
- LeadCoordLoop.tick -> the peer message stays held and unacked
Not wired into inject/Injector: no human types into a spawned member's pane,
so it would buy nothing and cost a herdr agent.read per member poll. The
heartbeat reads the pane only for a tick that would otherwise send.
Tests. FakeHerdr gains detectionText(), because one readText cannot be both
a worker's transcript (what the completion scrape reads) and a lead's empty
prompt; it falls back to readText so no existing fixture changes meaning.
Fixtures for the lead-nudge paths now state what their pane shows, since the
behaviour depends on it. 11 new behavioural tests across the three loops plus
InjectorTest, and 13 for the classifier; all 10 loop tests were run against
the unpatched loops first and fail there.
mvn clean install: BUILD SUCCESS, Tests run: 2161, Failures: 0, Errors: 0,
Skipped: 0 (summed from target/surefire-reports).
Finding 4: the comment above fleet_reply's handler claimed the authz check
asks whether the caller is a worker at all. It actually checks terminal
ownership (Authz.java REPLY/ASK -> caller.ownsSession), which is why an
observer can reply on its own pane with no role test involved.
Finding 5: fleet_list's tool description hardcoded 'architect/dev/reviewer',
missing hunter. Added MemberRole.wireNames() (pulled out of parse()'s error
message builder, which now calls it too) and used it in the description so
the list can't drift again.
The table row for an unconfigured pane told every session "Neither
fleet_list nor ListAgents lists these". PR #762 made that false for
fleet_list, and a stale note of this shape is the worst kind: it tells a
future session it cannot do the thing at the moment doing it is the job.
Two edits:
- The observer definition now says where an observer finds a target id,
which is the one thing it could not learn before.
- The table row names the panes array, says the row is full for a lead, an
architect or a collaborator and filtered for an observer, and keeps the
herdr tab list join as the fallback. ListAgents still lists none of them.
wiki/7-Use-Cases.md regenerated from this block in the wiki submodule at
4872227; the sync check prints in sync: True.
ConnectionIdentity's Caller record carried the same wrong rule as the two
places PR #760 fixed: it called the terminal a worker's, and read a null
terminal as the primary. The brief for #760 named only two of the three spots.
MemberPresence pointed at FleetMcp.markTrackedCallerPresent by name inside
{@code}, which the compiler does not check, and inject has no dependency on
mcp so a {@link} would add a cross-package reference. States the principle
instead, which stays true whichever roles qualify. Drops a ticket key.
paneRole now reads CallerResolver#boundToArchitectSlot (made public, no
second definition) so a slot-bound pane with no live member reports
"architect", matching what sendableObserverTarget already allowed as a
SEND target.
panesVisibleTo now admits an observer, since an observer holds SEND to
another observer pane. Its panes rows are filtered to
sendableObserverTarget and reduced to sessionId/label/status/role/
deliverable; every other caller's rows are unchanged.
Role.PRIMARY, ConnectionIdentity (class javadoc + callerTerminal), and
MemberPresence's class javadoc each state a role model the code no longer
implements. Comment-only change.
Three claims in the canonical block went false when the observer SEND grant
merged, and this file is the instruction surface the bridge ships:
- the observer role definition said "never SEND"
- invariant 3 listed send as "lead, architect, or collaborator"
- the hand-opened-pane row said such a pane "cannot fleet_send back"
The grant is narrow. Authz permits an observer's SEND only when
CallerResolver.sendableObserverTarget() classifies the target as an observer
too, so a lead, a collaborator, an architect slot and a spawned member are
each refused. TASK_READ stays denied, so #705 remains closed.
Measured on the merged tree: mvn clean install exit 0, 2137 tests from 178
surefire files, 0 failures. Mutating away the grant's call site
(FleetMcp.java:474) kills exactly FleetMcpObserverSendDeliveryTest, so the
behaviour is pinned and not only the predicate.
A workspace.list/tab.list failure in the label scan no longer costs the
caller the agent roster (GET /agents) or the leads/members/capacity/
coordinator rows (fleet_list) that never needed it. Both call sites now
fall back to an empty label map on HerdrException, so a pane row still
renders with label:null instead of the whole response failing.
Also cuts four comments down to the current contract, per the project's
comment rule: dropped the reviewer-facing justification from Fleetd's
deliverableTo javadoc, panesVisibleTo's javadoc, the fleet_list handler's
inline comment, and the panesVisible assembly-gate comment, and removed
the two fragments describing what a test must do.
FleetMcpObserverSendDeliveryTest already kills the same mutation end to
end (it asserts the exact attributed text herdr receives), and the code
quality rule in CLAUDE.md caps new source-text tests in this file at the
existing count.
Authz.SEND now grants an observer a narrow path: it may reach only a
target that CallerResolver.sendableObserverTarget() would itself
resolve as OBSERVER, never a lead, a collaborator, or a live spawned
member's terminal. This mirrors the existing collaborator SEND clause
rather than adding an unconditional caller.isObserver() grant, which
fleetd #705 already rejected as too broad.
Because the receiving pane cannot otherwise tell an observer's SEND
apart from a human paste, FleetMcp.attributeIfObserver prefixes the
delivered text with the sender's own daemon-resolved terminal on both
the MCP and REST entry paths, for exactly this one new path.
FleetAppAuthTest's start() helper never wired a spawnedMemberRole, so
its "worker" fixture actually resolved as OBSERVER under the real
CallerResolver -- invisible before because OBSERVER and WORKER shared
the same (zero) SEND grant. Granting OBSERVER a real SEND surfaced it:
two SEND-denial tests started passing for the wrong caller. Fixed the
fixture to resolve term_a as a live DEV, matching the helper's own
documented contract.
GET /agents now carries each agent's tab label, merged in from the same
herdr daemon(s) the roster is drawn from. fleet_list gains a panes array
with the same label, the sessionId fleet_send takes as a target, the
role the daemon resolves that pane as, and the deliverable gate the
injector itself enforces (Fleetd#deliverableTo, now public). Gated like
leads/collaborators (primary, architect, collaborator), not bare READ,
since a tab label and a member's cwd are not roster facts every READ
caller may see.
The lead's intent->tool table had no row for messaging an unconfigured pane,
and the user-scope instruction file said outright that the fleet has no route
to an interactive session unless an operator registers it as a collaborator.
That claim is false and it is load-bearing: a session reading it concludes the
exchange is impossible and stops, which is what happened here.
Delivery is gated on presence, not on SEND. contextExtractor runs on every MCP
request including initialize, markTrackedCallerPresent enrols an observer into
MemberPresence, and deliverableTo tests presence before the lead and
collaborator maps. So connecting the server is the enrolment, and fleet_reply
is gated on owning your own pane, which every pane does.
Add the table row, and add the enrolment side of the deliverability gate to the
flows page next to the existing "a spawned member is not deliverable until it
has mounted the MCP" bullet, which is the same gate read the other way.
Measured on a real pane, not a fake: trinotes answered with no fleet config, no
restart, and its fleet_* tools still deferred and unloaded.
errorPatternCoverageLine pointed at exhaustedPatternCoverageLine's javadoc
"for the measured swap mutation this pairing guards against". That narrative
was removed from the destination, so the pointer led nowhere.
The pairing with BUILT_IN_DEFAULT is the contract and it stays. What the
mutation proved belongs in history, not in the comment.
ignoreCycle() used to exempt every dependency between two packages, in both
directions, for the whole package. That meant a brand new dependency added
later between an already-excepted pair (auth/mcp, mcp/msg, inject/msg,
metrics/msg, msg/session) was silently exempted too, exactly where the
msg package makes the gate matter most.
Replace the package-wide ignore with a frozen baseline of the 45 exact
origin-class -> target-class edges that exist today between those five
pairs, and ignore only those via SliceRule.ignoreDependency(String, String).
A new dependency between a baselined pair is not in that set, so it is no
longer ignored and the existing beFreeOfCycles() check (or, when the new
edge alone would not form a cycle, a dedicated set-equality check) fails
and names the exact origin class, target class and package pair.
The set-equality check also fails on a baseline entry whose dependency no
longer exists in the code, so a removed edge cannot rot in the baseline
and mask the pair's eligibility for the ticket #131 removal steps. Rewrote
the javadoc to describe only the current contract.
Drop 2 dead *WiringTest names from 3 comment sites (renamed to
*AssemblyTest), rewriting each sentence to state the code's guarantee
instead of naming a test class. Reattach 5 javadoc blocks that were
orphaned behind a second /** block to the member they actually
describe, trimming history/evidence text down to the current contract
per the project's comment rule. Comment-only; no production logic
changed.
Two architects reviewed the codebase independently on separate backends and
both returned MIXED, not "a mess": 0 public mutable fields, 12 extends of
which 8 are exceptions, 120 records, and 2 files over 1000 code lines rather
than the 9 a total-line count suggests. Both rejected a Clean Code section,
a SOLID list, a pattern catalogue, class and method length limits, and a
coverage gate as text that would change no behaviour.
The comment rule they were briefed against was never in this repo. It lives
in the operator's personal global config, so it was not in git, never
reviewed, and not versioned with the code. Its "goes in the ADR" line was
unfollowable because this project has no ADR, which sent load-bearing
knowledge to a destination that does not exist. Rule 2 redirects it to
docs/<subject>.md, which exists.
Rule 4 covers what neither architect ranked first and both described: a
testing problem relieved by reshaping production code. 54 static factories
on Fleetd, 7 volatile race hooks in MessageService, 8 tests asserting on
main source as text, and a 452-line composition root across 8 tickets.
Rule 4's judgement half is marked as having no mechanism. A cap stops a
count growing; no test distinguishes a good decomposition from a bad one.
Verified: canonical block still byte-identical with wiki/7-Use-Cases.md.
fleet_handover{open} now returns outstandingTickets and openAsks. The
successor keeps the authority to poll and answer them but not the ids, and
an uncollected terminal ticket loses its reply at the ticket TTL.
A DONE/FAILED ticket nobody has polled yet is destroyed on a timer by
pruneTerminalTickets' completion-based TTL, while a PENDING ticket is not
going anywhere. Drop the !future.isDone() filter so outstanding() reports
every ticket the caller owns that is still in tasks, with its real phase.
rollingByTerminal keyed the single-flight lock on p.leadTerminal(), the pane
address. A roll replaces the pane, so a second roll of the same lead opened
from the new terminal landed on a different map key and could run concurrent
with the first roll's still-in-flight continuation.
PendingRollover now carries rolloverKey, resolved once in open() from
leadNameForTerminal while the lead is certainly still live, falling back to
the terminal itself when the name resolves null or blank. confirm()'s claim
and both release sites (the continuationRunner-rejection catch and
runRollover's finally) use the carried key instead of recomputing it, since
leadNameForTerminal no longer resolves the old terminal by release time.
NOT_YOUR_ROLLOVER and the pane-teardown calls stay keyed on p.leadTerminal(),
unchanged.
rollingByTerminal is renamed rollingByLead to match.
MessageService.outstanding(callerOwner) lists the caller's non-terminal
delegations and the subset paused in fleet_ask, filtered by the same
ownsTicket rule poll() already uses. FleetMcp threads the caller's owner
key into handover/handoverOpen and adds outstandingTickets/openAsks to
the open() JSON, alongside the unchanged token/handoverPath/requestedAtMillis.
The heartbeat loop now resolves its nudge target through
PrimaryRegistry#currentPrimaryTerminal(), so the sentence describing what a
background loop with no caller uses named the raw accessor instead.
The comment described the pre-#737 rule, where a null caller key read every
ticket. The pendingAsk gate now matches the unnamed primary's key like any
other, so the exemption it named no longer exists.
ReplyPushLoop.resolveLiveLead dropped a dead per-target delegation and
retried PrimaryRegistry.nudgeTargetFor, but returned that fallback without
checking isLive. The fallback is now probed the same way the first lead is,
and the method returns empty rather than trust a dead terminal.
PrimaryRegistry records a delegating lead's name alongside its learned
terminal and resolves the name back to its current terminal at nudge time,
through a name-to-terminal lookup backed by the live lead-tab scan. A name
with no current match falls back to the terminal that was actually learned,
so an unnamed primary, an off-host lead, or a non-herdr lead keeps working
exactly as before. LeadHeartbeatLoop now reads the resolved current terminal
instead of the raw learned one.
ownsTicket treated a null callerOwner as "read everything", conflating
the internal no-check test seam with a real unnamed primary's owner
key. Give them two different values: a named INTERNAL_NO_OWNER_CHECK
marker for the test-only bypass, and null as just another owner key
that must equal the ticket's recorded creatorOwner (including a
null-to-null match, so an unnamed primary still owns its own tickets).
Applies to both ownsTicket call sites, poll and pendingAsk.
The roll now ends the pane and launches a fresh process instead of typing
/clear, and that jar is deployed, so the skill's self-dated "a change is
coming" note had to go.
Measured on 2026-10-05 before writing:
grep -c "lead-rollover: rolled" fleetd/fleetd.out -> 20
grep -c "lead-rollover:" fleetd/fleetd.out -> 86 (control)
<tail from the last "fleetd listening"> | grep -c "lead-rollover:" -> 0
grep -c 'RELAUNCH_NEVER_READY\|RELAUNCH_NOT_RECOGNISED\|OLD_PANE_NEVER_DIED' -> 0
So all 20 recorded rolls ran under /clear and the restart path has never
executed. The section says that rather than implying the old numbers
describe it.
Also:
- name all eight RollState outcomes, with what each one guarantees
- state that relaunchReadySeconds bounds each of two waits, not the pair
- drop the "never observed as WORKING after 8 consecutive IDLE/DONE polls"
paragraph: grep finds that wait is deleted, so it cannot appear
- drop the #621 warning: contextNotice now takes requireOperatorConfirm
- split the surprise bullets into their own section
The javadoc on RollState.IN_PROGRESS enumerates the terminal states the entry
can be overwritten with, and omitted RELAUNCH_NEVER_READY. That state is
reachable at LeadRollover.java:710, so the list told a reader a state could not
occur when it can. Found by a reviewer on PR #742, outside its assigned scope.
Comment only; no behaviour change.
The assembled daemon's own boot-time orphan-worker reap makes a real call on
the member fake before any roll starts. Clear the member fake's recorded
calls once assembly finishes and before the roll begins, so the assertion
measures calls made since the roll started rather than the whole process's
lifetime, and reword its message to say so.
LeadRollover already sends every call through the lead-bound AgentControl and
WorkspaceControl it receives at construction, so no production code needed a
routing fix. Add a regression test that drives a full relaunch to the point
where recognition times out and asserts bootstrapText still lands on the lead
daemon and never on the member daemon, the one path the existing assembly test
never reaches.