Commit Graph

1269 Commits

Author SHA1 Message Date
Dai Ha 1cc0322782 fleetd #780: tell the truth about a queued-but-undelivered send
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 2m10s
A send to a pane that never frees up was accepted, then polled as
"pending — worker working" forever, with no WARN and no timeout. That
text is only ever correct for a message the injector already handed
off; a message still sitting in its queue now reports as queued, with
the target's live status and how long it has been waiting.

Injector.enqueue() now stamps Pending.enqueuedAtMillis so
queuedWaitMillis(target) can tell "never attempted" apart from
"delivered, now being worked on". MessageService.pendingDetail() uses
it to pick the poll wording. FleetMcp.sendAsync() adds a best-effort
warning to the accept text when the target is not injectable at send
time, per the brief's stated preference for a warning over a hard
refusal (a short busy spell is normal).

Injector gets a new bound, QUEUE_WAIT_GRACE_POLLS (4800 polls, ~20min
at the 250ms poll interval), mirroring READINESS_GRACE_POLLS: a
message that sits at the head of the queue with no delivery attempt
for that long fails via onTurnFailed, the same path a readiness
timeout uses, without touching presence — the target is busy, not
gone.

Fixed InjectorTest.readinessGraceExpiryLogsTheMeasuredElapsedTimeNotArithmeticOnConstants's
stub clock, which now sees one extra, legitimate nowMillis read from
enqueue()'s new stamp.

Tests: aQueuedButNeverInjectedMessageDoesNotPollAsWorking +
aDeliveredMessageStillPollsAsWorkerWorking (poll wording, with
positive control); failsAQueuedMessageWhoseTargetNeverFreesUp +
aTargetThatFreesUpBeforeTheQueueWaitGraceIsDeliveredNormally (timeout,
with positive control).
2026-10-05 19:54:37 +02:00
Dai Ha c043d149cf fleetd #775: name the trigger flag for what it now reads
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 1m0s
CI / build (push) Failing after 1m57s
The flag tested leader.tab() and was named for it. It now tests whether
fleet.leaders has any entry at all, so the old name states a condition the
code no longer checks.
2026-10-05 14:21:19 +02:00
Dai Ha fc786d0f67 fleetd #775: say which remedy applies to a lead vs a collaborator
The ticket's correction comment pointed out the refusal message still
implied removing a lead's tab helps, when only placement: tab does.
Restate the message so the lead and collaborator remedies are not
conflated, and keep the variable name the correction specified.
2026-10-05 14:19:17 +02:00
Dai Ha b314cb4d51 fleetd #775: pane-placement guard trigger keys on lead existence, not tab:
The guard's lead half now fires whenever fleet.leaders has any entry,
since a lead's tab is always labelled by the fixed LEAD_TAB_LABEL
constant regardless of its own deprecated tab: field. The refusal
message no longer advises removing a lead's tab:, which cannot
satisfy the guard any more.
2026-10-05 14:19:17 +02:00
Dai Ha a3d296f639 fleetd #770: the lead identity check is the label AND the space
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 2m11s
The redeploy skill's check 4 and the script's closing hint both told the
operator that a lead is found by fleet.leaders.*.tab. Identity is now the
fixed 'lead' label together with the lead's configured workspace, so both
would have sent a reader to a key that no longer decides anything.
2026-10-05 14:01:54 +02:00
Dai Ha 79423787d0 fleetd #770: lead identity keys on the space, not the tab label
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Failing after 1m57s
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 53s
CI / build (push) Failing after 1m52s
The lead tab label becomes a fixed constant (Leader.LEAD_TAB_LABEL = "lead");
fleet.leaders.<name>.tab is now optional legacy, matched case-insensitively
alongside the constant via Leader.acceptedLabels(). The uniqueness boundary
between leads moves from the exact tab text to the workspace: FleetConfig
refuses two leaders that share a workspace, LeadTabScanner indexes lead
labels per space (collaborators stay space-agnostic), and
LeadLauncher.leadNameOf/countLeads require both the accepted label and the
lead's own space to match, so a legacy-labelled tab in the wrong space never
counts and a daemon restart never double-spawns a second lead next to a live
one. Config validation also refuses a fleet.tabLabel template or a
collaborator tab that can render as the fixed lead label.
2026-10-05 13:49:14 +02:00
Dai Ha 364b229db9 fleetd #771 step 1: report each pane's workspace label in fleet_list
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Failing after 1m57s
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 54s
CI / build (push) Failing after 1m50s
Adds workspaceLabel next to workspaceId on fleet_list's panes row, read
from herdr's workspace.list via a new PaneLocator.workspaceLabelsByWorkspaceId().
An observer's reduced row still excludes it; a herdr failure or an unknown
workspaceId yields workspaceLabel: null without costing the rest of fleet_list.
2026-10-05 11:37:41 +02:00
Dai Ha ac436efefb fleetd #761: invariant 4 — a lead's own pane needs an empty input box
CI / shell-tests (push) Failing after 11s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m58s
2026-10-05 11:17:33 +02:00
Dai Ha 7dad054045 Merge remote-tracking branch 'origin/worker/761-draft-aware-lead-nudge-d05850-4' into vfy/761
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 1m0s
CI / build (push) Failing after 2m16s
2026-10-05 11:09:40 +02:00
Dai Ha 4e414c475f fleetd #761: read the input box marker the live TUI actually draws
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 2m1s
The gate matched the box line as "│ >", which appears zero times on a current
Claude Code pane. EMPTY was unreachable, so every lead nudge was held forever.

Match the box as a line *starting* with "❯" or with "│ >", keeping the older
bordered layout readable. A marker further along a line is transcript text — a
caret the operator quoted — so it no longer counts, and the last matching line
is still the live box because the detection region carries scrollback above it.

Look for the generating marker only from the box line down, for the same
reason: an earlier turn's "esc to interrupt" survives in that scrollback, and
holding on it would be the same unreachable-EMPTY failure by another route.

Fixtures: add IDLE_PROMPT_CARET / DRAFTED_PROMPT_CARET from a live pane and
point every lead-pane fake at them. The bordered constants stay, now covering
the older client. Against the old marker these fixtures fail 55 tests across
PromptBoxTest and the three loop tests, which is the production bug reproduced.

cd fleetd && mvn clean install -> BUILD SUCCESS, MVN_EXIT=0,
Tests run: 2164, Failures: 0, Errors: 0, Skipped: 0 (179 surefire XML files).
2026-10-05 11:07:40 +02:00
Dai Ha 9084667493 fleetd #761: hold a lead nudge while the operator's prompt box holds a draft
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Failing after 2m13s
herdr's agent.prompt pastes AND submits in one call, so a nudge arriving
while the operator is mid-sentence submitted their unfinished line with the
nudge glued to it. AgentStatus.injectable() cannot see this: it describes
the agent, and an idle agent reports the same status whether its input box
is empty or holds a half-typed line.

New herdr/PromptBox reads the pane's `detection` region — the same region
StatusRefiner uses, and the one the input box is drawn in — and clears a
delivery only when the box is positively empty. A box with characters, a
pane it cannot recognise, and a failed read all hold, because a held nudge
is recoverable and a submitted half-line is not. Whitespace and a cursor
block count as empty. After 20 consecutive holds for one target it logs one
warning, so a box that never clears is visible rather than silent; the
warning repeats only after the box has cleared again.

Wired into the three paths that nudge a lead's own pane:
  - ReplyPushLoop.decide -> WAIT_BUSY (the pending work is re-read next tick)
  - LeadHeartbeatLoop.tick -> a new DRAFT_HELD action that spends neither the
    quiet budget nor the one context notice per HIGH stretch
  - LeadCoordLoop.tick -> the peer message stays held and unacked

Not wired into inject/Injector: no human types into a spawned member's pane,
so it would buy nothing and cost a herdr agent.read per member poll. The
heartbeat reads the pane only for a tick that would otherwise send.

Tests. FakeHerdr gains detectionText(), because one readText cannot be both
a worker's transcript (what the completion scrape reads) and a lead's empty
prompt; it falls back to readText so no existing fixture changes meaning.
Fixtures for the lead-nudge paths now state what their pane shows, since the
behaviour depends on it. 11 new behavioural tests across the three loops plus
InjectorTest, and 13 for the classifier; all 10 loop tests were run against
the unpatched loops first and fail there.

mvn clean install: BUILD SUCCESS, Tests run: 2161, Failures: 0, Errors: 0,
Skipped: 0 (summed from target/surefire-reports).
2026-10-05 10:49:50 +02:00
Dai Ha 2289e94223 fleetd #759: fix the fleet_reply comment and the hand-copied role list
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Failing after 1m59s
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 1m59s
Finding 4: the comment above fleet_reply's handler claimed the authz check
asks whether the caller is a worker at all. It actually checks terminal
ownership (Authz.java REPLY/ASK -> caller.ownsSession), which is why an
observer can reply on its own pane with no role test involved.

Finding 5: fleet_list's tool description hardcoded 'architect/dev/reviewer',
missing hunter. Added MemberRole.wireNames() (pulled out of parse()'s error
message builder, which now calls it too) and used it in the description so
the list can't drift again.
2026-10-05 10:44:55 +02:00
Dai Ha 92adfcfae5 fleetd #756/#758: the canonical block no longer says panes are unlistable
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 2m13s
The table row for an unconfigured pane told every session "Neither
fleet_list nor ListAgents lists these". PR #762 made that false for
fleet_list, and a stale note of this shape is the worst kind: it tells a
future session it cannot do the thing at the moment doing it is the job.

Two edits:

- The observer definition now says where an observer finds a target id,
  which is the one thing it could not learn before.
- The table row names the panes array, says the row is full for a lead, an
  architect or a collaborator and filtered for an observer, and keeps the
  herdr tab list join as the fallback. ListAgents still lists none of them.

wiki/7-Use-Cases.md regenerated from this block in the wiki submodule at
4872227; the sync check prints in sync: True.
2026-10-05 10:39:32 +02:00
Dai Ha 5f7f388e69 Merge remote-tracking branch 'origin/worker/756-758-observer-pane-discovery-7e6ffd-1' into vfy/762
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 54s
CI / build (push) Failing after 1m56s
2026-10-05 10:29:49 +02:00
Dai Ha 39accf73e6 fleetd #759: fix the third copy of the role claim, and drop two references that rot
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m58s
ConnectionIdentity's Caller record carried the same wrong rule as the two
places PR #760 fixed: it called the terminal a worker's, and read a null
terminal as the primary. The brief for #760 named only two of the three spots.

MemberPresence pointed at FleetMcp.markTrackedCallerPresent by name inside
{@code}, which the compiler does not check, and inject has no dependency on
mcp so a {@link} would add a cross-package reference. States the principle
instead, which stays true whichever roles qualify. Drops a ticket key.
2026-10-05 10:26:36 +02:00
Dai Ha 5c2f296bc3 fleetd #756/#758: panes array reads the architect-slot gate and widens to observer
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 2m9s
paneRole now reads CallerResolver#boundToArchitectSlot (made public, no
second definition) so a slot-bound pane with no live member reports
"architect", matching what sendableObserverTarget already allowed as a
SEND target.

panesVisibleTo now admits an observer, since an observer holds SEND to
another observer pane. Its panes rows are filtered to
sendableObserverTarget and reduced to sessionId/label/status/role/
deliverable; every other caller's rows are unchanged.
2026-10-05 10:24:00 +02:00
Dai Ha 0f2ec7a6b5 fleetd #759: fix three role-model comments that describe the pre-CB-501 rule
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Failing after 2m7s
Role.PRIMARY, ConnectionIdentity (class javadoc + callerTerminal), and
MemberPresence's class javadoc each state a role model the code no longer
implements. Comment-only change.
2026-10-05 10:18:01 +02:00
Dai Ha 02eff4c532 fleetd #743: an observer may now send to another observer
CI / shell-tests (push) Failing after 13s
CI / contract (push) Successful in 51s
CI / build (push) Failing after 2m4s
Three claims in the canonical block went false when the observer SEND grant
merged, and this file is the instruction surface the bridge ships:

  - the observer role definition said "never SEND"
  - invariant 3 listed send as "lead, architect, or collaborator"
  - the hand-opened-pane row said such a pane "cannot fleet_send back"

The grant is narrow. Authz permits an observer's SEND only when
CallerResolver.sendableObserverTarget() classifies the target as an observer
too, so a lead, a collaborator, an architect slot and a spawned member are
each refused. TASK_READ stays denied, so #705 remains closed.

Measured on the merged tree: mvn clean install exit 0, 2137 tests from 178
surefire files, 0 failures. Mutating away the grant's call site
(FleetMcp.java:474) kills exactly FleetMcpObserverSendDeliveryTest, so the
behaviour is pinned and not only the predicate.
2026-10-05 09:42:43 +02:00
Dai Ha 92b6d9406b Merge remote-tracking branch 'origin/worker/743-observer-send-4706db-6' 2026-10-05 09:37:14 +02:00
Dai Ha c4498607e1 Merge remote-tracking branch 'origin/worker/743-pane-discovery-ad5b75-5' 2026-10-05 09:37:14 +02:00
Dai Ha e42eab5b4c fleetd #743: pin GET /agents' tab-label merge with a positive test
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 2m5s
2026-10-05 09:34:15 +02:00
Dai Ha b1d2cb48ac fleetd #743: make the pane-label scan best-effort, trim justification comments
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m56s
A workspace.list/tab.list failure in the label scan no longer costs the
caller the agent roster (GET /agents) or the leads/members/capacity/
coordinator rows (fleet_list) that never needed it. Both call sites now
fall back to an empty label map on HerdrException, so a pane row still
renders with label:null instead of the whole response failing.

Also cuts four comments down to the current contract, per the project's
comment rule: dropped the reviewer-facing justification from Fleetd's
deliverableTo javadoc, panesVisibleTo's javadoc, the fleet_list handler's
inline comment, and the panesVisible assembly-gate comment, and removed
the two fragments describing what a test must do.
2026-10-05 09:19:16 +02:00
Dai Ha 001367d82c fleetd #743: drop the redundant source-scrape test for attributeIfObserver
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Failing after 2m4s
FleetMcpObserverSendDeliveryTest already kills the same mutation end to
end (it asserts the exact attributed text herdr receives), and the code
quality rule in CLAUDE.md caps new source-text tests in this file at the
existing count.
2026-10-05 09:16:29 +02:00
Dai Ha 9fdcaaa8fd fleetd #743: let one observer pane message another, and nothing else
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Failing after 2m0s
Authz.SEND now grants an observer a narrow path: it may reach only a
target that CallerResolver.sendableObserverTarget() would itself
resolve as OBSERVER, never a lead, a collaborator, or a live spawned
member's terminal. This mirrors the existing collaborator SEND clause
rather than adding an unconditional caller.isObserver() grant, which
fleetd #705 already rejected as too broad.

Because the receiving pane cannot otherwise tell an observer's SEND
apart from a human paste, FleetMcp.attributeIfObserver prefixes the
delivered text with the sender's own daemon-resolved terminal on both
the MCP and REST entry paths, for exactly this one new path.

FleetAppAuthTest's start() helper never wired a spawnedMemberRole, so
its "worker" fixture actually resolved as OBSERVER under the real
CallerResolver -- invisible before because OBSERVER and WORKER shared
the same (zero) SEND grant. Granting OBSERVER a real SEND surfaced it:
two SEND-denial tests started passing for the wrong caller. Fixed the
fixture to resolve term_a as a live DEV, matching the helper's own
documented contract.
2026-10-05 09:03:16 +02:00
Dai Ha 459a523e2c fleetd #743: expose herdr pane discovery (tab labels) over REST and MCP
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m0s
CI / build (pull_request) Failing after 2m15s
GET /agents now carries each agent's tab label, merged in from the same
herdr daemon(s) the roster is drawn from. fleet_list gains a panes array
with the same label, the sessionId fleet_send takes as a target, the
role the daemon resolves that pane as, and the deliverable gate the
injector itself enforces (Fleetd#deliverableTo, now public). Gated like
leads/collaborators (primary, architect, collaborator), not bare READ,
since a tab label and a member's cwd are not roster facts every READ
caller may see.
2026-10-05 08:58:58 +02:00
Dai Ha 291dc02c77 fleetd #743: document that a hand-opened pane is reachable with no config
The lead's intent->tool table had no row for messaging an unconfigured pane,
and the user-scope instruction file said outright that the fleet has no route
to an interactive session unless an operator registers it as a collaborator.
That claim is false and it is load-bearing: a session reading it concludes the
exchange is impossible and stops, which is what happened here.

Delivery is gated on presence, not on SEND. contextExtractor runs on every MCP
request including initialize, markTrackedCallerPresent enrols an observer into
MemberPresence, and deliverableTo tests presence before the lead and
collaborator maps. So connecting the server is the enrolment, and fleet_reply
is gated on owning your own pane, which every pane does.

Add the table row, and add the enrolment side of the deliverability gate to the
flows page next to the existing "a spawned member is not deliverable until it
has mounted the MCP" bullet, which is the same gate read the other way.

Measured on a real pane, not a fake: trinotes answered with no fleet config, no
restart, and its fleet_* tools still deferred and unloaded.
2026-10-05 08:56:01 +02:00
Dai Ha be835aa259 Merge branch 'worker/749-edge-baseline-28d1a0-3'
CI / shell-tests (push) Failing after 12s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 2m4s
2026-10-05 07:47:18 +02:00
Dai Ha 80506c79d0 fleetd #748: drop a cross-reference the comment cleanup left dangling
errorPatternCoverageLine pointed at exhaustedPatternCoverageLine's javadoc
"for the measured swap mutation this pairing guards against". That narrative
was removed from the destination, so the pointer led nowhere.

The pairing with BUILT_IN_DEFAULT is the contract and it stays. What the
mutation proved belongs in history, not in the comment.
2026-10-05 07:46:59 +02:00
Dai Ha 6677ec8c63 fleetd #749: pin PackageCyclesTest's exceptions to exact edges, not whole packages
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m8s
CI / build (pull_request) Failing after 2m25s
ignoreCycle() used to exempt every dependency between two packages, in both
directions, for the whole package. That meant a brand new dependency added
later between an already-excepted pair (auth/mcp, mcp/msg, inject/msg,
metrics/msg, msg/session) was silently exempted too, exactly where the
msg package makes the gate matter most.

Replace the package-wide ignore with a frozen baseline of the 45 exact
origin-class -> target-class edges that exist today between those five
pairs, and ignore only those via SliceRule.ignoreDependency(String, String).
A new dependency between a baselined pair is not in that set, so it is no
longer ignored and the existing beFreeOfCycles() check (or, when the new
edge alone would not form a cycle, a dedicated set-equality check) fails
and names the exact origin class, target class and package pair.

The set-equality check also fails on a baseline entry whose dependency no
longer exists in the code, so a removed edge cannot rot in the baseline
and mask the pair's eligibility for the ticket #131 removal steps. Rewrote
the javadoc to describe only the current contract.
2026-10-05 07:45:51 +02:00
Dai Ha ca0c965932 Merge branch 'worker/748-dead-comment-refs-f42ac5-4' 2026-10-05 07:44:44 +02:00
Dai Ha e8ab933cbc fleetd #748: fix dead test-class references and orphaned javadoc blocks
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 2m1s
Drop 2 dead *WiringTest names from 3 comment sites (renamed to
*AssemblyTest), rewriting each sentence to state the code's guarantee
instead of naming a test class. Reattach 5 javadoc blocks that were
orphaned behind a second /** block to the member they actually
describe, trimming history/evidence text down to the current contract
per the project's comment rule. Comment-only; no production logic
changed.
2026-10-05 07:42:55 +02:00
Dai Ha 2adb950a12 fleetd #748: five code-quality rules enter the repo
Two architects reviewed the codebase independently on separate backends and
both returned MIXED, not "a mess": 0 public mutable fields, 12 extends of
which 8 are exceptions, 120 records, and 2 files over 1000 code lines rather
than the 9 a total-line count suggests. Both rejected a Clean Code section,
a SOLID list, a pattern catalogue, class and method length limits, and a
coverage gate as text that would change no behaviour.

The comment rule they were briefed against was never in this repo. It lives
in the operator's personal global config, so it was not in git, never
reviewed, and not versioned with the code. Its "goes in the ADR" line was
unfollowable because this project has no ADR, which sent load-bearing
knowledge to a destination that does not exist. Rule 2 redirects it to
docs/<subject>.md, which exists.

Rule 4 covers what neither architect ranked first and both described: a
testing problem relieved by reshaping production code. 54 static factories
on Fleetd, 7 volatile race hooks in MessageService, 8 tests asserting on
main source as text, and a 452-line composition root across 8 tickets.

Rule 4's judgement half is marked as having no mechanism. A cap stops a
count growing; no test distinguishes a good decomposition from a bad one.

Verified: canonical block still byte-identical with wiki/7-Use-Cases.md.
2026-10-05 07:36:20 +02:00
Dai Ha 7f0c4a8464 fleetd #737: the handover skill carries the outstanding tickets forward
CI / shell-tests (push) Failing after 12s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 2m7s
fleet_handover{open} now returns outstandingTickets and openAsks. The
successor keeps the authority to poll and answer them but not the ids, and
an uncollected terminal ticket loses its reply at the ticket TTL.
2026-10-05 06:09:26 +02:00
Dai Ha 682991a846 Merge branch 'worker/737-9c61d3-4'
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 50s
CI / build (push) Failing after 1m52s
2026-10-05 06:07:33 +02:00
Dai Ha abe617c48c Merge branch 'worker/737-a263f3-3'
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m3s
CI / build (push) Failing after 2m20s
2026-10-05 06:01:49 +02:00
Dai Ha 886ce1521d Merge remote-tracking branch 'origin/main' into worker/737-9c61d3-4
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Failing after 1m59s
2026-10-05 06:00:49 +02:00
Dai Ha d41aff4012 fleetd #737 unit 5 correction: outstanding() reports terminal-phase tickets too
A DONE/FAILED ticket nobody has polled yet is destroyed on a timer by
pruneTerminalTickets' completion-based TTL, while a PENDING ticket is not
going anywhere. Drop the !future.isDone() filter so outstanding() reports
every ticket the caller owns that is still in tasks, with its real phase.
2026-10-05 06:00:43 +02:00
Dai Ha 8a1d73b39e fleetd #737 unit 4: key the rollover single-flight claim on the lead's name
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 1m59s
rollingByTerminal keyed the single-flight lock on p.leadTerminal(), the pane
address. A roll replaces the pane, so a second roll of the same lead opened
from the new terminal landed on a different map key and could run concurrent
with the first roll's still-in-flight continuation.

PendingRollover now carries rolloverKey, resolved once in open() from
leadNameForTerminal while the lead is certainly still live, falling back to
the terminal itself when the name resolves null or blank. confirm()'s claim
and both release sites (the continuationRunner-rejection catch and
runRollover's finally) use the carried key instead of recomputing it, since
leadNameForTerminal no longer resolves the old terminal by release time.
NOT_YOUR_ROLLOVER and the pane-teardown calls stay keyed on p.leadTerminal(),
unchanged.

rollingByTerminal is renamed rollingByLead to match.
2026-10-05 05:57:59 +02:00
Dai Ha 70a735b638 fleetd #737 unit 5: fleet_handover{open} reports outstanding tickets and open asks
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 2m10s
MessageService.outstanding(callerOwner) lists the caller's non-terminal
delegations and the subset paused in fleet_ask, filtered by the same
ownsTicket rule poll() already uses. FleetMcp threads the caller's owner
key into handover/handoverOpen and adds outstandingTickets/openAsks to
the open() JSON, alongside the unchanged token/handoverPath/requestedAtMillis.
2026-10-05 05:56:06 +02:00
Dai Ha 803c91ea6c fleetd #737 unit 3: name the resolving accessor in LeadRollover's javadoc
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m0s
CI / build (push) Failing after 2m7s
The heartbeat loop now resolves its nudge target through
PrimaryRegistry#currentPrimaryTerminal(), so the sentence describing what a
background loop with no caller uses named the raw accessor instead.
2026-10-05 05:37:32 +02:00
Dai Ha d438a74575 Merge branch 'worker/737-20d1d9-1' 2026-10-05 05:37:32 +02:00
Dai Ha 7467ffa252 fleetd #737 unit 6: drop the stale unnamed-primary claim from the REST status comment
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 2m9s
The comment described the pre-#737 rule, where a null caller key read every
ticket. The pendingAsk gate now matches the unnamed primary's key like any
other, so the exemption it named no longer exists.
2026-10-05 05:32:59 +02:00
Dai Ha f4176ae455 fleetd #737 unit 3: probe the fallback terminal and resolve it by lead name
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Failing after 1m55s
ReplyPushLoop.resolveLiveLead dropped a dead per-target delegation and
retried PrimaryRegistry.nudgeTargetFor, but returned that fallback without
checking isLive. The fallback is now probed the same way the first lead is,
and the method returns empty rather than trust a dead terminal.

PrimaryRegistry records a delegating lead's name alongside its learned
terminal and resolves the name back to its current terminal at nudge time,
through a name-to-terminal lookup backed by the live lead-tab scan. A name
with no current match falls back to the terminal that was actually learned,
so an unnamed primary, an off-host lead, or a non-herdr lead keeps working
exactly as before. LeadHeartbeatLoop now reads the resolved current terminal
instead of the raw learned one.
2026-10-05 05:29:06 +02:00
Dai Ha 6ab3a81af7 fleetd #737 unit 6: stop the unnamed primary sharing the internal bypass
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Failing after 1m59s
ownsTicket treated a null callerOwner as "read everything", conflating
the internal no-check test seam with a real unnamed primary's owner
key. Give them two different values: a named INTERNAL_NO_OWNER_CHECK
marker for the test-only bypass, and null as just another owner key
that must equal the ticket's recorded creatorOwner (including a
null-to-null match, so an unnamed primary still owns its own tickets).
Applies to both ownsTicket call sites, poll and pendingAsk.
2026-10-05 05:24:37 +02:00
Dai Ha 8fa8d18c97 fleetd #726: rewrite the handover skill for the restart roll
The roll now ends the pane and launches a fresh process instead of typing
/clear, and that jar is deployed, so the skill's self-dated "a change is
coming" note had to go.

Measured on 2026-10-05 before writing:

  grep -c "lead-rollover: rolled" fleetd/fleetd.out   -> 20
  grep -c "lead-rollover:" fleetd/fleetd.out          -> 86 (control)
  <tail from the last "fleetd listening"> | grep -c "lead-rollover:" -> 0
  grep -c 'RELAUNCH_NEVER_READY\|RELAUNCH_NOT_RECOGNISED\|OLD_PANE_NEVER_DIED' -> 0

So all 20 recorded rolls ran under /clear and the restart path has never
executed. The section says that rather than implying the old numbers
describe it.

Also:
- name all eight RollState outcomes, with what each one guarantees
- state that relaunchReadySeconds bounds each of two waits, not the pair
- drop the "never observed as WORKING after 8 consecutive IDLE/DONE polls"
  paragraph: grep finds that wait is deleted, so it cannot appear
- drop the #621 warning: contextNotice now takes requireOperatorConfirm
- split the surprise bullets into their own section
2026-10-05 05:04:52 +02:00
Dai Ha 9d653e86df fleetd #726: name RELAUNCH_NEVER_READY in the IN_PROGRESS terminal-state list
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 2m7s
The javadoc on RollState.IN_PROGRESS enumerates the terminal states the entry
can be overwritten with, and omitted RELAUNCH_NEVER_READY. That state is
reachable at LeadRollover.java:710, so the list told a reader a state could not
occur when it can. Found by a reviewer on PR #742, outside its assigned scope.

Comment only; no behaviour change.
2026-10-04 21:44:03 +02:00
Dai Ha f6d1131d7a Merge remote-tracking branch 'origin/worker/726-unit2-75cb13-4' 2026-10-04 21:43:34 +02:00
Dai Ha a0505dc614 fleetd #726 unit 2: scope the member-daemon assertion to the roll itself
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 2m0s
The assembled daemon's own boot-time orphan-worker reap makes a real call on
the member fake before any roll starts. Clear the member fake's recorded
calls once assembly finishes and before the roll begins, so the assertion
measures calls made since the roll started rather than the whole process's
lifetime, and reword its message to say so.
2026-10-04 21:24:49 +02:00
Dai Ha d2f30f1654 fleetd #726 unit 2: cover the bootstrapText relaunch send with a dedicated test
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 1m46s
LeadRollover already sends every call through the lead-bound AgentControl and
WorkspaceControl it receives at construction, so no production code needed a
routing fix. Add a regression test that drives a full relaunch to the point
where recognition times out and asserts bootstrapText still lands on the lead
daemon and never on the member daemon, the one path the existing assembly test
never reaches.
2026-10-04 21:17:15 +02:00
Dai Ha cd1f04cbb4 Merge remote-tracking branch 'origin/worker/737-owner-key-ff061f-10'
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 53s
CI / build (push) Failing after 1m48s
2026-10-04 21:11:20 +02:00