fleetd #669 Unit D: the caller resolver — spawned member outranks every tab map #701

Closed
agent wants to merge 0 commits from worker/669-unit-d-efbbd7-1 into main
Member

Unit D of fleetd #669 — the caller resolver

Implements the spec in issue comment 18450 (re-checked against 18455, posted after it, which is out of scope for Unit D and does not contradict it).

Resolution order in CallerResolver.resolve()

  1. terminal() == null tail — unchanged.
  2. Spawned-member roster (new, checked first): a live member's own role wins, consulting no tab map. MemberRole.ARCHITECT -> Role.ARCHITECT (resolving the bound slot's name), DEV/HUNTER/REVIEWER -> Role.WORKER.
  3. Lead tab map -> Role.PRIMARY — unchanged.
  4. Architect slot binding -> Role.ARCHITECT — unchanged; now only reached by a binding with no live member session.
  5. Collaborator tab map (new) -> Role.COLLABORATOR.
  6. Worker fallback — unchanged.

The roster lookup is injected as Function<String, MemberRole> (not a list), built in FleetdAssembly from SessionManager.roster() — the non-resolving roster meant for a hot path — never rosterResolved().

LeadTabScanner generalized, not duplicated

Matches lead and collaborator tabs in one pass: a merged tabToEntry index (Entry(name, kind)), the same liveness cross-check (agent.list), the same one-scan grace period, the same TTL cache, and the same "a failed scan keeps the previous answer" behaviour, for both kinds. get() still returns lead-only entries (so FleetdAssemblyLeadTabScannerExclusionTest's reflection on the leads field is unaffected); a new collaborators() method returns collaborator-only entries from the same cache.

Classifier wired for real

Authz.NO_KNOWN_LEAD_OR_COLLABORATOR stays as the fail-closed default, but both FleetMcp.denyFor and FleetApp.allow now pass CallerResolver.knownLeadOrCollaborator() — built from the exact same lead/collaborator maps resolve() reads, so a target that classifier calls known is one resolve() would actually resolve as a lead or collaborator.

Scanner-interval decision (collaborator-only fleet)

FleetConfig.Collaborator carries no scanIntervalSeconds (unlike Leader). When fleet.leaders is empty but fleet.collaborators is not, FleetdAssembly falls back to 10 seconds — the same default FleetConfig.Leader's own compact constructor uses — rather than inventing a second number for the same kind of scan.

Acceptance criteria

  1. Ordering: CallerResolverTest.aSpawnedMemberWinsOverALeadTabForTheSamePane — a terminal in both the roster and the lead map resolves as its member role.
  2. Mandatory mutation test, performed and reverted: removed the spawned-member step from resolve(). Exactly one test went red: dev.ltms.fleet.auth.CallerResolverTest.aSpawnedMemberWinsOverALeadTabForTheSamePane (47 other CallerResolverTest cases stayed green). Reverted; full suite re-confirmed green.
  3. CallerResolverTest.aConfiguredCollaboratorTabResolvesToCollaboratorCarryingItsName + scanner-level LeadTabScannerTest.aConfiguredCollaboratorTabIsReportedByCollaboratorsNotByGet.
  4. LeadTabScannerTest.aDeadCollaboratorTabIsNotReported.
  5. Regression: all pre-existing CallerResolverTest/LeadTabScannerTest cases pass unchanged (backward-compatible constructor overloads — new params default to no-ops).
  6. Both MCP and REST, over real HTTP for REST: FleetMcpAuthzTest.aCollaboratorMaySendToAKnownLeadOrCollaboratorButNotToASpawnedMembersTerminal and FleetAppAuthTest.aCollaboratorMaySendToAKnownLeadOrCollaboratorButNotToASpawnedMembersTerminalOverRest (POSTs to a real Javalin server: 202 to a known lead, 403 to a terminal standing in for a spawned member).
  7. mvn clean install in fleetd/: BUILD SUCCESS, Tests run: 1988, Failures: 0, Errors: 0, Skipped: 0. Baseline measured earlier in this session on untouched origin/main: 1974 tests, 0 failures — so this PR adds 14 new tests (confirmed by git diff | grep -c '@Test' on the added lines, which also counts 14).

Out of scope, noted not fixed

The Unit C implementer's shape sweep (and issue comment 18455) already named and filed as #700 an audit-skip asymmetry in FleetMcp.denyFor/FleetApp.allow (READ/TASK_READ/METRICS skip the allowed audit log line). Untouched here, per the Unit D brief.

Build run in my own worktree, not the primary's checkout.

## Unit D of fleetd #669 — the caller resolver Implements the spec in issue comment 18450 (re-checked against 18455, posted after it, which is out of scope for Unit D and does not contradict it). ### Resolution order in `CallerResolver.resolve()` 1. `terminal() == null` tail — unchanged. 2. **Spawned-member roster** (new, checked first): a live member's own role wins, consulting no tab map. `MemberRole.ARCHITECT` -> `Role.ARCHITECT` (resolving the bound slot's name), `DEV`/`HUNTER`/`REVIEWER` -> `Role.WORKER`. 3. Lead tab map -> `Role.PRIMARY` — unchanged. 4. Architect slot binding -> `Role.ARCHITECT` — unchanged; now only reached by a binding with no live member session. 5. **Collaborator tab map** (new) -> `Role.COLLABORATOR`. 6. Worker fallback — unchanged. The roster lookup is injected as `Function<String, MemberRole>` (not a list), built in `FleetdAssembly` from `SessionManager.roster()` — the non-resolving roster meant for a hot path — never `rosterResolved()`. ### `LeadTabScanner` generalized, not duplicated Matches lead and collaborator tabs in **one pass**: a merged `tabToEntry` index (`Entry(name, kind)`), the same liveness cross-check (`agent.list`), the same one-scan grace period, the same TTL cache, and the same "a failed scan keeps the previous answer" behaviour, for both kinds. `get()` still returns lead-only entries (so `FleetdAssemblyLeadTabScannerExclusionTest`'s reflection on the `leads` field is unaffected); a new `collaborators()` method returns collaborator-only entries from the same cache. ### Classifier wired for real `Authz.NO_KNOWN_LEAD_OR_COLLABORATOR` stays as the fail-closed default, but both `FleetMcp.denyFor` and `FleetApp.allow` now pass `CallerResolver.knownLeadOrCollaborator()` — built from the exact same lead/collaborator maps `resolve()` reads, so a target that classifier calls known is one `resolve()` would actually resolve as a lead or collaborator. ### Scanner-interval decision (collaborator-only fleet) `FleetConfig.Collaborator` carries no `scanIntervalSeconds` (unlike `Leader`). When `fleet.leaders` is empty but `fleet.collaborators` is not, `FleetdAssembly` falls back to **10 seconds** — the same default `FleetConfig.Leader`'s own compact constructor uses — rather than inventing a second number for the same kind of scan. ### Acceptance criteria 1. Ordering: `CallerResolverTest.aSpawnedMemberWinsOverALeadTabForTheSamePane` — a terminal in both the roster and the lead map resolves as its member role. 2. **Mandatory mutation test, performed and reverted**: removed the spawned-member step from `resolve()`. Exactly one test went red: `dev.ltms.fleet.auth.CallerResolverTest.aSpawnedMemberWinsOverALeadTabForTheSamePane` (47 other CallerResolverTest cases stayed green). Reverted; full suite re-confirmed green. 3. `CallerResolverTest.aConfiguredCollaboratorTabResolvesToCollaboratorCarryingItsName` + scanner-level `LeadTabScannerTest.aConfiguredCollaboratorTabIsReportedByCollaboratorsNotByGet`. 4. `LeadTabScannerTest.aDeadCollaboratorTabIsNotReported`. 5. Regression: all pre-existing `CallerResolverTest`/`LeadTabScannerTest` cases pass unchanged (backward-compatible constructor overloads — new params default to no-ops). 6. Both MCP and REST, over real HTTP for REST: `FleetMcpAuthzTest.aCollaboratorMaySendToAKnownLeadOrCollaboratorButNotToASpawnedMembersTerminal` and `FleetAppAuthTest.aCollaboratorMaySendToAKnownLeadOrCollaboratorButNotToASpawnedMembersTerminalOverRest` (POSTs to a real Javalin server: 202 to a known lead, 403 to a terminal standing in for a spawned member). 7. `mvn clean install` in `fleetd/`: **BUILD SUCCESS**, `Tests run: 1988, Failures: 0, Errors: 0, Skipped: 0`. Baseline measured earlier in this session on untouched `origin/main`: 1974 tests, 0 failures — so this PR adds 14 new tests (confirmed by `git diff | grep -c '@Test'` on the added lines, which also counts 14). ### Out of scope, noted not fixed The Unit C implementer's shape sweep (and issue comment 18455) already named and filed as #700 an audit-skip asymmetry in `FleetMcp.denyFor`/`FleetApp.allow` (READ/TASK_READ/METRICS skip the `allowed` audit log line). Untouched here, per the Unit D brief. Build run in my own worktree, not the primary's checkout.
agent added 1 commit 2026-10-04 01:13:46 +02:00
fleetd #669 Unit D: a live spawned member outranks every tab map
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Failing after 1m43s
02c6909546
CallerResolver.resolve() now checks the spawned-member roster first, ahead
of every tab map, so a live member's own role wins over a lead or
collaborator tab naming the same terminal (closes #661 at the resolver
level). Adds collaborator resolution (Role.COLLABORATOR) and a real
knownLeadOrCollaborator() classifier, wired into both FleetMcp.denyFor and
FleetApp.allow in place of the inert NO_KNOWN_LEAD_OR_COLLABORATOR stand-in.

LeadTabScanner is generalized to match lead and collaborator tabs in one
pass, keeping every existing liveness/caching/grace-scan property for both
kinds. FleetdAssembly wires the collaborator tab map and a roster-backed
spawned-member lookup; a collaborator-only fleet (no leaders configured)
falls back to a 10s scan interval, matching FleetConfig.Leader's own default.
agent added 1 commit 2026-10-04 01:37:17 +02:00
fleetd #669 Unit D correction: confirm live slot role before granting a spawned architect
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 1m44s
d83972ede2
The new spawned-member roster step returned Principal.architect from the roster's
own MemberRole.ARCHITECT alone, never reconfirming against memberSlotRoles. A slot
revoked after the bind kept granting ARCHITECT to the already-bound session,
regressing fleetd #424's "config governs what a bound slot still grants" half. The
roster now only answers that the pane is a live spawned member; a confirmed live
slot role still decides whether that grants ARCHITECT, falling through to WORKER
otherwise — mirroring the existing architect-slot step a few lines below.

Also corrects LeadTabScanner.buildTabIndex's javadoc: a colliding lead/collaborator
key resolves to the lead because the lead entry is put last (overwriting), not
because it is put first.
ltms closed this pull request 2026-10-04 01:43:56 +02:00
Some checks are pending
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 1m44s

Pull request closed

Sign in to join this conversation.