Commit Graph

1168 Commits

Author SHA1 Message Date
Dai Ha a9a37af957 Merge PR #712: fleetd #705 — gate fleet_poll's ticket lookup by the creating caller
Closes the MCP door only. The REST door (GET /tasks/{ticket}) still calls the
no-check poll overload, so #705 stays open.

Verified in a throwaway worktree, not piped: Tests run: 2008, BUILD SUCCESS.
2026-10-04 07:00:35 +02:00
Dai Ha 11eccc3a1b fleetd #705: gate fleet_poll's ticket lookup by the creating caller's terminal
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m45s
A ticket id is a plain sequential counter, so any session holding
TASK_READ could walk task-1, task-2, ... and read another session's
delegation reply. sendAsync now records the creating caller's terminal
on the Task, and poll refuses a caller whose terminal differs from it.
A caller with no terminal (the unnamed primary) is still allowed
through regardless, since it never carries a herdr pane to compare.
2026-10-04 06:50:26 +02:00
Dai Ha 4939d40362 Merge PR #709: fleetd #703 — fleet_list reports collaborators, narrowed by role
CI / shell-tests (push) Failing after 12s
CI / contract (push) Successful in 48s
CI / build (push) Failing after 1m53s
A collaborator could find and message a lead, but a lead had nowhere to read a
collaborator's sessionId, so the channel only worked once the collaborator
spoke first. fleet_list now carries a collaborators array, and
CallerResolver.collaborators() has its first caller outside the resolver.

Visible to the primary, an architect and a collaborator; absent for a worker.
The rule is that the roles which may see a collaborator row are the roles that
can act on it, and a worker cannot send to a collaborator at all. Narrowed in
the payload the way coordinatorVisibleTo already does it, because fleet_list is
gated on READ and a worker holds READ.

Each row carries the registry name and the sessionId only. A collaborator is
never spawned and has no profile, so a lead row's context and seat fields have
no meaning for it.

mvn clean install: exit 0, BUILD SUCCESS, Tests run: 2005, Failures: 0.

Nothing yet pins the handler to collaboratorsVisibleTo, unlike the coordinator
flag, so a literal passed there would regress silently. Tracked in #710.
2026-10-04 06:42:05 +02:00
Dai Ha e7a7711a4e fleetd #703: fleet_list reports a collaborators array so a lead can discover one
CI / shell-tests (pull_request) Failing after 11s
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Failing after 1m54s
Threads CallerResolver#collaborators() into fleet_list's canonical listFleet
overload and adds a collaborators array (name, sessionId), visible only to
the primary, an architect, and a collaborator -- the same roles Authz grants
SEND to a named peer -- never a worker. Gives CallerResolver#collaborators()
its first real caller outside the resolver itself.
2026-10-04 06:37:48 +02:00
Dai Ha 1fd2cfa716 Merge PR #708: fleetd #669 — fleetd.example.yaml promised a defence that is not wired
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m10s
CI / build (push) Failing after 2m15s
The example file said two things stop the tab-name convention from becoming a
way to claim leadership, and named workspace exclusion as the first. Production
builds the scanner with an empty excludedWorkspaceLabels, so that defence does
not exist. It now names the two that do: the startup refusal of a colliding
tabPrefix/tabLabel, and CallerResolver asking the live spawned-member roster
before any tab map.

It also said a lead's workspace default is "leads" and must not be a member
workspace. The default is "fleet", the same space the members use, so that
advice was against the shipped shape.

Adds the test nobody wrote: a lead is still discovered when its workspace is
the member workspace, with an empty exclusion set. Corrects a test javadoc that
called itself the guard that matters while covering a parameter production never
passes.

mvn clean install: exit 0, BUILD SUCCESS, Tests run: 2000, Failures: 0.
2026-10-04 06:35:49 +02:00
Dai Ha 3e8e314656 Merge PR #706 and PR #707: fleetd #669 — a collaborator is deliverable, and a collaborator config change is reported
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 1m0s
CI / build (push) Failing after 1m46s
Two follow-ups to #669, verified together in one throwaway worktree.

PR #707 (f820737): ConfigRef reports that a fleet.collaborators change needs a
restart. The map is read once at startup and a reload does not rebuild it, so
before this an operator saw a clean reload and no live effect.

PR #706 (20fc42b): deliverableTo opens for a collaborator. A collaborator is
never in MemberPresence and never found by the lead scan, so every send to one
was authorized, then held on the injector gate for ~60s and dropped without a
keystroke reaching the pane.

mvn clean install on the merged tree: exit 0, BUILD SUCCESS, Tests run: 1999,
Failures: 0, Errors: 0. ConfigRefTest 32, FleetDeliverabilityTest 9.
2026-10-04 06:19:47 +02:00
Dai Ha 20fc42b572 fleetd #669 follow-up: deliverableTo also opens for a collaborator
CI / shell-tests (pull_request) Failing after 11s
CI / contract (pull_request) Successful in 59s
CI / build (pull_request) Failing after 1m48s
A collaborator's terminal was authorized at Authz but never deliverable: it
is never enrolled in MemberPresence and never discovered by the lead scan,
so a send to a collaborator sat on the injector's readiness gate for
~60s and failed, never typed into the pane. deliverableTo now takes a
third collaborators supplier, read through on each call like the lead
supplier, and FleetdAssembly wires the existing collaboratorTerminals
supplier into it.
2026-10-04 06:13:03 +02:00
Dai Ha f82073717a fleetd #669: report collaborator reload restart
CI / shell-tests (pull_request) Failing after 11s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 2m1s
2026-10-04 06:12:37 +02:00
Dai Ha 16b52fac6c fleetd.example.yaml: the collaborator block works now (#669)
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m49s
The comment above the collaborators: example said the block 'is parsed and
validated today; nothing yet recognises or addresses the tab it names'. That
was true when Unit B landed the config shape. Units C, D and E have since
merged and deployed, so the tab is recognised, the role is authorized, and
the pane routes to the lead herdr daemon.

This is the worst place for that sentence to go stale: it sits in the file an
operator copies to turn the feature on, and it tells them the block does
nothing. Replaced with what the role may and may not do, taken from the rows
in auth/Authz.java rather than from the wiki, plus the #703 limit that a lead
cannot discover a collaborator.

FleetConfigTest (which loads this file) green at 170 tests.
2026-10-04 02:30:28 +02:00
Dai Ha 13b6ae2628 Merge PR #704: fleetd #669 Unit E — route a collaborator's pane to the lead herdr daemon
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 58s
CI / build (push) Failing after 2m5s
A configured collaborator's terminal now routes to the lead herdr daemon
rather than the member one. FleetdAssembly ORs a second terminal map into
the predicate it hands HerdrRouter, and the predicate's field is renamed
isLead -> routeToLead so its name matches the widened contract.

Verified in the lead's own throwaway worktree: mvn clean install green at
1992 tests, 0 failures (174 surefire report files, fresh worktree). The
fix rests on one test, so I reproduced the mutation myself rather than
taking the worker's word: dropping the collaborator clause from the
predicate turns FleetdAssemblyCollaboratorHerdrRoutingTest red at runtime
with two distinct AgentControl identities, while HerdrRouterTest's 4 tests
stay green. That confirms both the kill and that the two-distinct-fakes
setup is actually discriminating.

Also checked by hand: both terminal maps come from LeadTabScanner's one
byKind helper and are keyed terminal_id -> name, so containsKey(target) is
correct for both; no agentsFor call runs between router construction and
the ref being set, so the publication window is harmless and matches the
one leadsRef already had.

One reviewer fanned out against the diff, briefed from the diff rather
than the implementer's rationale. It reported no issue.
2026-10-04 02:19:44 +02:00
Dai Ha 7e838ba8b9 fleetd #669 Unit E: route a collaborator's pane to the lead herdr daemon
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Failing after 1m57s
HerdrRouter.agentsFor picked the member daemon for any terminal the lead
predicate did not recognize, so a configured collaborator's pane (opened
by a person, exactly like a lead's) was routed to the member herdr
daemon instead of the lead one.

FleetdAssembly now combines the leads map and the collaborator-terminals
map into the predicate it hands HerdrRouter. HerdrRouter's isLead field
and constructor parameter are renamed to routeToLead, with its javadoc
naming the real contract: true for any terminal whose pane lives in the
lead daemon, lead or collaborator.
2026-10-04 02:11:17 +02:00
Dai Ha c3e3554bde CLAUDE.md: the collaborator role, after #669 Unit D
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m50s
The block is the instruction surface this repo ships, and Unit D made
four of its statements false. A session can now resolve as
collaborator, read "Which role am I?", and find no section telling it
what it may do.

Each claim below was read from the merged tree at b5bc5d4.

  - fleet_whoami returns four roles, not three. FleetMcp.whoami returns
    before the lead branch, so a collaborator gets its registry name
    and its own sessionId, and no leader key.
  - The fallback ladder said only that it cannot separate a worker from
    an architect. It also fires for no collaborator at all: every rung
    detects a spawned member, and nothing launched a collaborator. It
    therefore falls to "act as a worker", which is the safe direction
    but leaves it unable to learn what it is without asking.
  - Invariant 3 said send is lead or architect. Authz SEND also allows
    a collaborator when the target passes knownLeadOrCollaborator, so
    it may reach a lead or another collaborator and no spawned member.
  - The invariants heading said "both roles"; there are now four.

New Collaborator section, placed after the member turn contract because
a collaborator must be told that contract is not its own. It owes no
fleet_reply: nothing delegates to it. Its two surprising limits are
written down rather than left to be discovered, and both were measured
in Authz: it cannot reach a worker, and it cannot read a ticket,
because TASK_READ is withheld where READ is not.

The intent table gains a row for messaging a collaborator, and that row
states the gap instead of implying discovery works: fleet_list emits
only "leads" and "members", and CallerResolver.collaborators() has no
caller outside the resolver, so a lead cannot find a collaborator
unless the collaborator tells it its sessionId.

The wiki template is byte-identical again, verified with the project's
own check in the main clone, which printed False before the propagation
and True after.
2026-10-04 01:58:03 +02:00
Dai Ha b5bc5d4ab5 Merge PR #701: fleetd #669 Unit D — the resolver, a live spawned member outranks every tab map
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m55s
2026-10-04 01:43:17 +02:00
Dai Ha d83972ede2 fleetd #669 Unit D correction: confirm live slot role before granting a spawned architect
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 1m44s
The new spawned-member roster step returned Principal.architect from the roster's
own MemberRole.ARCHITECT alone, never reconfirming against memberSlotRoles. A slot
revoked after the bind kept granting ARCHITECT to the already-bound session,
regressing fleetd #424's "config governs what a bound slot still grants" half. The
roster now only answers that the pane is a live spawned member; a confirmed live
slot role still decides whether that grants ARCHITECT, falling through to WORKER
otherwise — mirroring the existing architect-slot step a few lines below.

Also corrects LeadTabScanner.buildTabIndex's javadoc: a colliding lead/collaborator
key resolves to the lead because the lead entry is put last (overwriting), not
because it is put first.
2026-10-04 01:37:13 +02:00
Dai Ha 02c6909546 fleetd #669 Unit D: a live spawned member outranks every tab map
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Failing after 1m43s
CallerResolver.resolve() now checks the spawned-member roster first, ahead
of every tab map, so a live member's own role wins over a lead or
collaborator tab naming the same terminal (closes #661 at the resolver
level). Adds collaborator resolution (Role.COLLABORATOR) and a real
knownLeadOrCollaborator() classifier, wired into both FleetMcp.denyFor and
FleetApp.allow in place of the inert NO_KNOWN_LEAD_OR_COLLABORATOR stand-in.

LeadTabScanner is generalized to match lead and collaborator tabs in one
pass, keeping every existing liveness/caching/grace-scan property for both
kinds. FleetdAssembly wires the collaborator tab map and a roster-backed
spawned-member lookup; a collaborator-only fleet (no leaders configured)
falls back to a 10s scan interval, matching FleetConfig.Leader's own default.
2026-10-04 01:13:25 +02:00
Dai Ha b92a669ddc Merge PR #699: fleetd #669 Unit C — the COLLABORATOR role and its authorization row
CI / shell-tests (push) Failing after 11s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m49s
2026-10-04 00:34:13 +02:00
Dai Ha bb29b001e4 fleetd #669 Unit C: add the COLLABORATOR role and its authorization row
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Failing after 1m48s
Adds Role.COLLABORATOR, Principal.collaborator(), and the Authz.permits grant:
READ/METRICS open, REPLY/ASK via ownership, SEND limited to a configured lead
or collaborator via a new classifier parameter, everything else denied.
isSpawnedMember() stays WORKER || ARCHITECT. fleet_whoami reports role and
collaborator name with no leader key.

Per a same-day ticket correction, the existing 3-argument Authz.permits is
kept (fail-closed default via the new NO_KNOWN_LEAD_OR_COLLABORATOR
classifier) rather than deleted, so the 47 pre-existing call sites in
AuthzTest/CallerResolverTest are untouched; both production gates
(FleetMcp#denyFor, FleetApp#allow via the new permitsFor seam) call the
4-argument form explicitly with the shared constant.

Mutation-verified: removing the classifier conjunct from the SEND arm kills
exactly 3 tests (AuthzTest, FleetMcpAuthzTest, FleetAppAuthTest), one per
gate, no cascade.

Full mvn clean install at this commit: 1974 tests, 0 failures (baseline at
f0ff252 was 1964; +10 are the new collaborator-matrix tests). Independently
counted from target/surefire-reports/*.xml after rm -rf: 173 report files,
aggregate tests=1974 failures=0 errors=0.
2026-10-04 00:22:27 +02:00
Dai Ha f0ff25221e Merge PR #697: fleetd #669 Unit B — recognise-only fleet.collaborators config block
CI / shell-tests (push) Failing after 14s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 2m2s
2026-10-04 00:02:10 +02:00
Dai Ha 780cb342ad fleetd #669: pin the fleet-wide tabLabel-vs-collaborator-tab check
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Failing after 1m58s
Adds a test driving validateLeadTabPrefixes()'s fleet-wide
fleet.tabLabel branch directly against a collaborator tab, with no
profile override involved. Every existing collaborator test exercised
only the profile-override branch below it, leaving the fleet-wide
branch without a pinning test.
2026-10-03 23:55:47 +02:00
Dai Ha 5c563f02c8 fleetd #669: comments-only fixes — drop history/narration and premature behaviour claims
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Failing after 1m54s
The duplicate-collaborator test's javadoc narrated the before/after and a
timestamp; replaced with the present-tense behaviour it protects, per the
project's code-comment rule.

The Collaborator record javadoc and fleetd.example.yaml both claimed the
daemon "learns to address" a collaborator tab. Nothing does yet — that is
Units C-E. Dropped the claim from the record's contract and the example
now says the block is parsed and validated today, nothing more.

Also drops a confidence marker and a cross-class pointer from the
Collaborator record javadoc. No logic, assertion, or file outside this
PR's existing three changed.
2026-10-03 23:47:16 +02:00
Dai Ha ad593c9bb9 fleetd #669: add collaborators to the duplicate-slot-key guard
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 2m3s
FLEET_POOL_KEYS was missing "collaborators", so a duplicated
fleet.collaborators.<name> key took the skipValue branch in
rejectDuplicateSlotsInPools and silently collapsed last-wins, unlike
the other five pools. Measured: before this change, a test loading a
config with a duplicated collaborator name observed no exception;
after, it fails exactly like the existing architect-pool case.

Also updates the javadoc's "five pools" count to six.
2026-10-03 23:41:13 +02:00
Dai Ha 736fd9cf4b Merge PR #696: fleetd #692 — bound the unbounded userinfo mask in redact()
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m46s
2026-10-03 23:33:37 +02:00
Dai Ha 05244a82b3 fleetd #669 Unit B: recognise-only fleet.collaborators config block
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Failing after 1m46s
Adds fleet.collaborators.<name>.tab (a Collaborator record with only a tab
field — no profile, no instances, no kind; never auto-launched) and widens
the two startup refusals that stop it being a privilege hole:

- validatePanePlacementAgainstLeadTabs() now fires on either a lead tab or
  a collaborator tab, and no longer early-returns when fleet.leaders is
  empty but fleet.collaborators is not.
- validateLeadTabPrefixes() now also refuses a tabLabel template able to
  render as a collaborator tab, two collaborators sharing one exact tab,
  and a collaborator tab equal to a lead tab (case-insensitively).

validateMembers() refuses a collaborator with no (or blank) tab, next to
the existing leader-must-name-a-tab check.

Does not touch Authz, CallerResolver, LeadTabScanner, MemberRole or
FleetMcp — those are later units per the ticket's plan.
2026-10-03 23:33:28 +02:00
Dai Ha ef4996a01e fleetd #692: bound the unbounded userinfo mask at redact()'s line 273
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 2m9s
redact()'s final fallthrough still used the unbounded [^@]* class #638
removed from the verdict-line masker, so a diff line with a URL that has
no userinfo plus a later @ elsewhere had its text between them silently
deleted. Share one bounded implementation, mask_url_userinfo(), between
redact() and mask_verdict_userinfo() so the bound lives in one place.
2026-10-03 23:27:15 +02:00
Dai Ha edbd8d816a Merge PR #694: fleetd #689 — check SEND before reading the request body in sendMessage
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 44s
CI / build (push) Failing after 1m53s
2026-10-03 22:58:17 +02:00
Dai Ha 9425a9b696 fleetd #689: pin the answerGatePasses call site via the audit trail
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Failing after 2m3s
A unit test on the extracted helper proves the helper, not the call
site in sendMessage. allow() logs an AuditLog.allowed() entry for
every granted non-READ/METRICS/TASK_READ action, so a granted turnId
request must log both SEND and ANSWER, and a granted plain request
must log SEND alone. Verified this goes red when the call site is
deleted from sendMessage, and restores to a clean diff.
2026-10-03 22:56:04 +02:00
Dai Ha d0688c8a60 Merge PR #695: fleetd #693 — pin the lead-tab guard's case-insensitivity; fleetd #676 — drop the last stale validator count
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 1m47s
2026-10-03 22:54:04 +02:00
Dai Ha 2c467c2553 fleetd #693, #676: pin lead-tab case-insensitivity, drop stale validator count
CI / shell-tests (pull_request) Failing after 20s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m40s
#693: add a FleetConfigTest case for two fleet.leaders tabs differing only
in case — mutating equalsIgnoreCase to equals left this uncaught before.

#676: rename theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames
to theSweepRunsEveryValidateMethodOnAnUnrelatedClass, since there are
eight validators now, not six, and the count had drifted into the name
and its class-javadoc {@link}.
2026-10-03 22:51:23 +02:00
Dai Ha c6430d8edd fleetd #689: check SEND before reading the request body in sendMessage
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 1m56s
Authorize twice: the coarse SEND grant first, with no body read, then
parse the body, then check ANSWER too when turnId is present. Restores
the pre-#687 ordering (no attacker-controlled body parse before the
gate) while keeping the SEND/ANSWER split #687 introduced.
2026-10-03 22:47:25 +02:00
Dai Ha bfee23acc3 Merge PR #691: fleetd #677 — refuse two leads sharing one exact tab (supersedes #686)
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m37s
2026-10-03 22:45:23 +02:00
Dai Ha 7dec74f1b4 Merge PR #690: fleetd #638 — stop the verdict userinfo mask from crossing / or whitespace (supersedes #685)
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 54s
CI / build (push) Failing after 1m58s
2026-10-03 22:39:42 +02:00
Dai Ha 482598e2a6 fleetd #677: refuse two leads that share one exact tab
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Failing after 2m17s
validateLeadTabPrefixes() compared each lead's tab against the member
tabLabel template, but never against another lead's tab. Two leads
configured with the same exact tab loaded cleanly, even though exact
tab is the only thing lead identity is matched on, so only one of them
could ever be found.

Add an independent pass, case-insensitive, that refuses when two
fleet.leaders entries share one exact tab, with its own exception so
the message stays accurate for this relation.

Decided not to also refuse a lead's tab starting with a sibling's
tabPrefix: tabPrefix plays no role in identity resolution (only the
exact tab does), and the default tabPrefix is "lead:", the same
string used as the conventional lead tab prefix throughout this
codebase's own fixtures (e.g. "lead: opus" / "lead: sol"). Refusing
that case would reject the standard multi-lead setup with no matching
identity hazard.
2026-10-03 22:36:44 +02:00
Dai Ha 5f5d16fbd4 fleetd #638: stop the verdict userinfo mask from crossing / or whitespace
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Failing after 1m57s
mask_verdict_userinfo's character class [^@]* crossed a '/' or a space, so a
verdict line with a URI that has no userinfo plus a later @ (e.g. an email
address in diagnostic prose) had everything between them destroyed. Restrict
the class to [^@/[:space:]]* so the match stops at the end of the URI.

Adds the uncovered-direction test: a URI with no userinfo plus a later @ in
the same line must pass through byte for byte. Also drops the design-rationale
sentence from the helper's comment (now in the PR description).
2026-10-03 22:34:02 +02:00
Dai Ha 7df7985a16 Merge remote-tracking branch 'refs/remotes/pr/686' into worker/677-fix-lead-collision-f69073-12 2026-10-03 22:29:29 +02:00
Dai Ha 724b35b46e Merge remote-tracking branch 'refs/remotes/pr/685' into worker/638-fix-overmask-dbb1bf-11 2026-10-03 22:29:14 +02:00
Dai Ha 2eb2d6112e Merge PR #688: fleetd #675 — pin three unpinned FleetdAssembly constructor args
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 1m35s
2026-10-03 22:27:32 +02:00
Dai Ha 7c458e8bf2 Merge PR #687: fleetd #669 Unit A — split SEND and READ into their real call shapes (closes #678) 2026-10-03 22:27:32 +02:00
Dai Ha 133f03e428 Merge PR #684: fleetd #683 — decouple the completion-fallback test's two 5s budgets 2026-10-03 22:27:27 +02:00
Dai Ha 804279175d fleetd #675: pin assembly loop timing defaults
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 1m42s
2026-10-03 22:12:46 +02:00
Dai Ha 9dea289975 fleetd #669 Unit A: split SEND and READ into their real call shapes
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m6s
CI / build (pull_request) Failing after 2m7s
SEND covered three different call shapes under one action (local
sessionId delivery, the coordId cross-host broker route, and the
turnId answer-a-blocked-worker form). READ covered both roster/
profile/identity observation and ticket-polling/session-status.
Split each into its own Authz.Action — SEND/COORD_SEND/ANSWER and
READ/TASK_READ — with every new action granted to exactly who held
the combined action before, on both the MCP and REST entry paths.

Also fixes fleetd #678's Authz.java comment: READ no longer claims
"the roster carries no secrets" for ticket replies and pending
questions, because those now live under TASK_READ.
2026-10-03 22:12:13 +02:00
Dai Ha cb4a6869b9 fleetd #677: guard exact lead tab collisions
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Failing after 1m56s
2026-10-03 22:10:29 +02:00
Dai Ha 28b45d97e5 fleetd #638: mask userinfo in the daemon verdict line before it prints
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m7s
CI / build (pull_request) Failing after 2m4s
Apply the userinfo-only rewrite (sed -E 's#://[^@]*@#://<redacted>@#g') to every
path that prints config-edit.sh's $VERDICT_LINE: restore_and_confirm's two
prints, report_outcome's shared local copy (covering its clean/needs-restart/
refused branches), and check_mode's "last verdict in log" line via
last_verdict_line. Deliberately not routed through redact() — that function's
key:value masking does not match this line's prose, and the rest of the line
(e.g. the pattern quoted in a parse-failure refusal) is the detail an operator
needs to fix the refusal.

No current refusal message echoes a URI, token or password, so this is a guard
against a future validator doing so, not a fix for an observed leak.
2026-10-03 22:06:54 +02:00
Dai Ha 1a397e962e fleetd #683: decouple the completion-fallback test's send budget from its own setup clock
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Failing after 1m53s
completionFallbackResolvesATurnThatNeverCalledFleetReply gave messages.send a 5000 ms budget
that started ticking the instant sendAsync() ran, then raced that same clock against
awaitWaiting()'s own 2000 ms deadline plus several onStatus/readText calls before asserting with
send.get(5, SECONDS). On a loaded machine the setup could eat enough of the 5000 ms that the
production call expired first, returning TIMED_OUT_QUEUED instead of COMPLETED_UNREPLIED.

Give this one test's send a 30 000 ms budget (sendAsync(content, timeoutMillis)) so the setup can
never compete with it; send.get(5, SECONDS) stays the one clock the test depends on. A new
regression test injects a deterministic 5500 ms delay in the same spot and proves the budget is no
longer the binding constraint — reverting it to 5000 ms turns that test red with the same
TIMED_OUT_QUEUED mismatch, confirmed by mutation.
2026-10-03 22:06:26 +02:00
Dai Ha 209e1231ea Merge PR #682: fleetd #651 — raise turnSettleSeconds default to 300, correct the settle javadoc
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 1m54s
The deferred roll waits for the calling lead's own turn to end. The old 20s bound
was shorter than one ordinary closing message (measured elapsed=20394ms), so a
lead that followed the handover skill's instruction to say its goodbye in the
same turn landed in the refusal branch. 300 matches idleAfterSeconds' existing
written reason: five minutes absorbs a normal pause without stalling.

The javadoc said both waits are for the pane to report an injectable state.
AgentStatus.injectable() accepts BLOCKED; the settle check requires IDLE or DONE.
The wording named the wrong predicate.
2026-10-03 21:43:25 +02:00
Dai Ha 5d8b9d365c fleetd #651: raise turnSettleSeconds default from 20 to 300
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Failing after 1m55s
A lead that runs the documented handover procedure writes its goodbye
message in the same turn as fleet_handover confirm. The deferred roll
then waits turnSettleSeconds for that same pane to reach IDLE or DONE.
An ordinary goodbye turn took 20394ms, so a 20s budget was too small
and the roll refused itself with TURN_NEVER_SETTLED. The refusal
branch is correct; only the budget was wrong.

300 is not derived from that single 20394ms observation. It matches
leadHeartbeat.idleAfterSeconds (also default 300), the only other
constant in this codebase answering "how long may a lead legitimately
be mid-turn", whose own javadoc reasons that 5 minutes absorbs normal
pauses without stalling. The two constants answer the same question
and should not disagree by a factor of fifteen. The wait stays bounded
on purpose: an unbounded wait would let a lead whose turn never ends
park a continuation and hold a pending token forever, which is harder
to notice than a logged refusal.

Also corrects FleetConfig's javadoc for turnSettleSeconds and
clearSettleSeconds, which described both waits as waiting for the
pane to report an "injectable" state again. The code requires IDLE or
DONE specifically and excludes BLOCKED (a live turn merely paused),
which is the exact state where sending /clear would destroy context.
LeadRollover.java's javadoc already states this correctly; only
FleetConfig.java's was wrong.

Adds FleetConfigTest coverage for the default's resolution: unset,
positive, and <= 0 all resolve as expected. Coverage for the two
turn-boundary properties (settles-in-time vs. never-settles, and
BLOCKED is not treated as settled) already existed in
LeadRolloverTest and needed no change — confirmed by mutating the
default back to 20 and the IDLE||DONE check to also accept BLOCKED;
both mutations were caught by existing or new tests, then reverted.
2026-10-03 21:35:46 +02:00
Dai Ha a6aeda39e7 Merge PR #681: fleetd #680 — pin the jar-path split, check the plist's jar, widen the locator
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 46s
CI / build (push) Failing after 1m50s
Verified by the lead before merge, each mutation re-run independently:
- three-dot diff: 1 commit, 3 files, nothing dragged in; clean trial merge
- suite exit 0, 134 test functions (5 new); the 3 FAIL and 5 mktemp lines are
  the documented deliberate self-test output, unchanged from baseline
- JAR moved under target/        -> RED  (the core invariant is pinned now)
- JAR moved elsewhere outside it -> PASS (not a refuse-everything test)
- PATTERN narrowed back to run/  -> RED
- plist-jar mismatch die removed -> RED
- comm=java allowlist removed    -> RED (widening PATTERN did not weaken #593)
- mvn clean install from fleetd/: Tests run: 1929, Failures: 0, 172 reports

Measured separately: during a full Maven build no extra java process matches
the widened 'fleetd.jar' pattern, so assert_single_daemon does not false-
positive while a worker builds.
2026-10-03 21:34:14 +02:00
Dai Ha 31b3c24caa fleetd #651: tell a rolling lead that surviving its goodbye means the roll refused
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 47s
CI / build (push) Failing after 2m9s
A lead calls fleet_handover confirm from inside its own turn, and the skill
tells it to write its goodbye in that same turn. The deferred roll then waits
turnSettleSeconds for that pane to reach IDLE or DONE. Measured on 2026-10-02,
an ordinary goodbye turn took 20394ms against a 20s budget, so the roll wrote
TURN_NEVER_SETTLED and sent no /clear. confirm() had already returned accepted,
so there was no caller left to tell, and the lead carried on believing it had
been replaced.

The refusal is the correct branch — clearing a live turn would destroy context.
The gap is that nobody is told to look afterwards.

fleet_handover{action:"status", token} already reports the outcome: FleetMcp
dispatches it to LeadRollover.status, which reads the same outcomes map the
refusal writes into. Nothing new is needed in the daemon for this half.

The signal costs nothing, because a successful roll clears the lead. A lead
that is still running after its goodbye already knows the roll failed. That is
more reliable than warning it to keep the goodbye short, which would make the
feature worst exactly when it matters most.

The live budget on this host was raised to 300s separately, in fleetd.yaml,
which is hot: the config-watcher reloaded at 21:28:00 with no "needs a restart"
clause, four seconds after the edit, with no daemon restart. The code default
is a separate change.
2026-10-03 21:29:21 +02:00
Dai Ha d105da978d fleetd #680: pin the JAR/BUILD_JAR split, check the plist's jar path, and widen the daemon-locator pattern
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Failing after 2m28s
Part 1: asserts JAR and BUILD_JAR as the script sources them (no test assigns
them first), so reverting JAR to a path under target/ now fails the suite.

Part 2: check_jar_path_matches_plist reads the installed launchd plist's
ProgramArguments and refuses when its jar path does not resolve to $JAR,
mirroring check_log_path_matches_plist. Wired into report_supervisor_state's
launchd branch; unaffected when no plist is installed.

Part 3 (added to the ticket after the brief, by comment): PATTERN narrowed to
'fleetd.jar' so running_pid()/assert_single_daemon see a daemon regardless of
which build layout (run/ or target/) its jar sits under. The comm=java
allowlist still excludes a self-matching shell. Brought
.claude/skills/fleets-status/SKILL.md's pgrep pattern into agreement.

Verified: bash scripts/test-redeploy-fleetd.sh exits 0. Mutation both
directions for part 1 (JAR under target/ -> suite fails; JAR elsewhere ->
suite passes), a positive control for part 2 (neutralizing the mismatch
check makes the new test fail), and a positive control for part 3 (narrowing
PATTERN back to run/fleetd.jar makes the new target/-dir test fail). mvn -o
clean install: Tests run: 1929, Failures: 0, Errors: 0, Skipped: 0, BUILD
SUCCESS, 172 surefire report files.
2026-10-03 21:28:59 +02:00
Dai Ha 5051a06443 Merge PR #679: fleetd #664 — run the daemon from fleetd/run/fleetd.jar
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m28s
CI / build (push) Failing after 1m52s
Verified by the lead before merge:
- three-dot diff: 1 commit, 9 files, nothing dragged in
- trial merge in a throwaway worktree: no conflicts
- mvn clean install from fleetd/: Tests run: 1929, Failures: 0, 172 report files
- scripts/test-redeploy-fleetd.sh: exit 0; its 3 FAIL and 5 mktemp lines are
  deliberate self-test output, confirmed identical on unmerged origin/main
- run/ is correctly ignored by the new fleetd/.gitignore rule
- the build writes target/fleetd.jar and does not create run/
2026-10-03 21:13:33 +02:00
Dai Ha a134eccc57 fleetd #664: run the daemon from fleetd/run/fleetd.jar, not target/
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Failing after 2m58s
Separate Maven's output path (target/fleetd.jar) from the daemon's runtime
path (run/fleetd.jar), so a clean/install in the main clone can no longer
reach the jar a running daemon holds open. redeploy-fleetd.sh now swaps the
built jar onto the runtime path with a same-filesystem mv, only after the
old daemon is confirmed gone; --check reports the built and running jars
as two separately labelled hash+mtime facts. Updates every process-locator
pattern and runtime-path reference found by git grep, with a positive
control added for both running_pid()'s PATTERN and fleets-status's pgrep
pattern. fleetd/run/ is gitignored.
2026-10-03 21:06:42 +02:00