Auto-register local fleets by space and drop fleet.leaders.*.workspace — but make primary per-space FIRST #779

Open
opened 2026-10-05 19:08:41 +02:00 by ltms · 1 comment
Owner

Operator direction, 2026-10-05: "fleetd.yaml workspace: fleet is redundant, fleetd should track local fleet ids and auto register/unregister." Together with the boundary rule from the same conversation — one space is one fleet; a leader and its members share the space name — the space already is the fleet id, so declaring it in config repeats what herdr can be asked.

The goal is right. The order of the two steps is not optional, and doing them in the wrong order is a privilege escalation, not a regression.

workspace: is doing two jobs, and only one of them is redundant

Job Redundant?
naming which space holds the lead tab yes — discoverable from herdr: the space containing a tab labelled lead
being the privilege boundary that stops another fleet's lead resolving as this daemon's primary no — it is the only thing doing this today

Why the second job is load-bearing — measured

Principal carries no space:

public record Principal(Role role, String terminal, long pid, String name) { ... }

Principal.java:18. Authz.java contains zero occurrences of space, workspace or fleet(. So isPrimary() is a daemon-wide boolean: a primary is primary over the whole daemon, not over one fleet.

Right now exactly one pane can hold it, because (space=fleet, label=lead) matches one tab. Measured today with three lead-labelled tabs live:

tab space role
w2:tY fleet lead
wA:t1 trinotes observer
wB:t1 anki observer

Auto-register every lead-labelled space before scoping primary, and those three become three global primaries. Each could then fleet_spawn, fleet_stop, fleet_drain and fleet_handover against the other fleets' members, because the gate it passes has no idea which fleet asked. Authz.java:133 is case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary();.

This is the same asymmetry as #770's cutover, where renaming the tab was always safe and deleting tab: first demoted the lead and launched a second one. The safe step goes first.

Required order

  1. Make the fleet id a first-class field and scope primary rights to it. Add the fleet (space) to Principal, resolved from the connection like every other identity field — never an argument (invariant 3). Change the lead-only gates to "primary of the fleet that owns the target". A lead with no members of its own can then do nothing to anyone else's.
  2. Then auto-register: a space holding a live lead tab is a fleet; register it, and unregister when that tab stops hosting a live agent.
  3. Then fleet.leaders.<name>.workspace becomes optional, and finally removed.

Step 1 is safe on its own and ships value immediately — it closes the escalation that #777 describes from the collaborator side. Steps 2 and 3 are inert until it lands.

What auto-register must not do

  • Do not key liveness on the label. A label left by a dead session reads as a live fleet forever. The scanner already solved this once in #359 by cross-checking every labelled tab against agent.list and dropping tabs with no agent running. Auto-unregister must use that same signal, not the label's presence.
  • Do not let an absent declaration widen. Dropping workspace: must not make a fleet match any space; see the recorded shape an absent config block widens, it does not empty. The replacement for the declaration is discovery, not a wildcard.

Two things that break at the second fleet, and need a decision

  • coordinator.selfId is the host, not the fleet. It reads mac here. LeadCoordLoop.resolveLocalLead() matches a lead whose name equals selfCoordId(), else falls back to the sole lead (LeadCoordLoop.java:228-230), else WARNs and returns null. That fallback is correct today and becomes a coin toss with two fleets. Peer-lead mail needs a host+fleet address. (It refuses rather than guesses, so it is safe meanwhile — checked today.)
  • Capacity is daemon-wide. fleet_list's capacity array reports one maxLoad/free per profile for the whole daemon. Several fleets would draw on one pool with no quota. Needs a policy: per-fleet quota, or documented first-come.

Supersedes a standing constraint

My own notes have carried "not to be started: multiple fleets on one host", from the experiment where two daemons on one herdr session killed each other's members. This direction is different and is the supported shape: one daemon, many fleets, keyed by space. The two-daemon finding still stands and is not what this asks for.

Not measured

I have not driven a second fleet to watch a second pane resolve as LEAD — the escalation above is read from Principal.java:18, the absence of any space term in Authz.java, and Authz.java:133. Nothing is mis-scoped today: fleet_list reports one lead and collaborators: [].

Operator direction, 2026-10-05: *"fleetd.yaml `workspace: fleet` is redundant, fleetd should track local fleet ids and auto register/unregister."* Together with the boundary rule from the same conversation — **one space is one fleet; a leader and its members share the space name** — the space already *is* the fleet id, so declaring it in config repeats what herdr can be asked. The goal is right. **The order of the two steps is not optional**, and doing them in the wrong order is a privilege escalation, not a regression. ## `workspace:` is doing two jobs, and only one of them is redundant | Job | Redundant? | |---|---| | naming which space holds the lead tab | **yes** — discoverable from herdr: the space containing a tab labelled `lead` | | being the **privilege boundary** that stops another fleet's lead resolving as this daemon's primary | **no** — it is the only thing doing this today | ## Why the second job is load-bearing — measured `Principal` carries no space: ```java public record Principal(Role role, String terminal, long pid, String name) { ... } ``` `Principal.java:18`. `Authz.java` contains **zero** occurrences of `space`, `workspace` or `fleet(`. So `isPrimary()` is a daemon-wide boolean: a primary is primary over the whole daemon, not over one fleet. Right now exactly one pane can hold it, because `(space=fleet, label=lead)` matches one tab. Measured today with three `lead`-labelled tabs live: | tab | space | role | |---|---|---| | `w2:tY` | `fleet` | `lead` | | `wA:t1` | `trinotes` | `observer` | | `wB:t1` | `anki` | `observer` | Auto-register every `lead`-labelled space **before** scoping primary, and those three become three global primaries. Each could then `fleet_spawn`, `fleet_stop`, `fleet_drain` and `fleet_handover` against the other fleets' members, because the gate it passes has no idea which fleet asked. `Authz.java:133` is `case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary();`. This is the same asymmetry as #770's cutover, where renaming the tab was always safe and deleting `tab:` first demoted the lead *and* launched a second one. The safe step goes first. ## Required order 1. **Make the fleet id a first-class field and scope primary rights to it.** Add the fleet (space) to `Principal`, resolved from the connection like every other identity field — never an argument (invariant 3). Change the lead-only gates to "primary **of the fleet that owns the target**". A lead with no members of its own can then do nothing to anyone else's. 2. **Then** auto-register: a space holding a live `lead` tab is a fleet; register it, and unregister when that tab stops hosting a live agent. 3. **Then** `fleet.leaders.<name>.workspace` becomes optional, and finally removed. Step 1 is safe on its own and ships value immediately — it closes the escalation that #777 describes from the collaborator side. Steps 2 and 3 are inert until it lands. ## What auto-register must not do - **Do not key liveness on the label.** A label left by a dead session reads as a live fleet forever. The scanner already solved this once in #359 by cross-checking every labelled tab against `agent.list` and dropping tabs with no agent running. Auto-unregister must use that same signal, not the label's presence. - **Do not let an absent declaration widen.** Dropping `workspace:` must not make a fleet match *any* space; see the recorded shape *an absent config block widens, it does not empty*. The replacement for the declaration is discovery, not a wildcard. ## Two things that break at the second fleet, and need a decision - **`coordinator.selfId` is the host, not the fleet.** It reads `mac` here. `LeadCoordLoop.resolveLocalLead()` matches a lead whose name equals `selfCoordId()`, else falls back to the sole lead (`LeadCoordLoop.java:228-230`), else WARNs and returns null. That fallback is correct today and becomes a coin toss with two fleets. Peer-lead mail needs a host+fleet address. (It refuses rather than guesses, so it is safe meanwhile — checked today.) - **Capacity is daemon-wide.** `fleet_list`'s `capacity` array reports one `maxLoad`/`free` per profile for the whole daemon. Several fleets would draw on one pool with no quota. Needs a policy: per-fleet quota, or documented first-come. ## Supersedes a standing constraint My own notes have carried "not to be started: multiple fleets on one host", from the experiment where two daemons on one herdr session killed each other's members. This direction is different and is the supported shape: **one daemon, many fleets, keyed by space.** The two-daemon finding still stands and is not what this asks for. ## Not measured I have not driven a second fleet to watch a second pane resolve as LEAD — the escalation above is read from `Principal.java:18`, the absence of any space term in `Authz.java`, and `Authz.java:133`. Nothing is mis-scoped today: `fleet_list` reports one lead and `collaborators: []`.
Author
Owner

Sequencing: step 1 and #777 must be one unit, not two

I read the code to size step 1 and found that it and #777 change the same two members of LeadTabScanner. Splitting them guarantees a merge collision, so they should go to one worker.

Where the space already is. The scan loop holds the workspace in hand while it classifies each tab:

for (... ws ...) {
    if (ws.workspaceId() == null || excludedWorkspaceLabels.contains(ws.label())) { ... }
    Map<String, String> leadLabelsHere = leadLabelsFor(ws.label());
    ...
        Entry entry = entryOf(tab.label(), leadLabelsHere);

So ws.label() is available at the point each Entry is built. Entry is private record Entry(String name, Kind kind) — adding the space is a one-field change, and it is the same field both tickets need.

What each ticket needs from it.

  • #777 needs entryOf to key the collaborator lookup by space. Today the lead branch reads leadLabelsHere, which is space-scoped, and the very next line reads collaboratorTabToName, which is flat:

    String leadName = leadLabelsHere.get(normalized);
    if (leadName != null) return new Entry(leadName, Kind.LEAD);
    String collaboratorName = collaboratorTabToName.get(normalized);   // no space key
    
  • step 1 needs the space to travel out of the scanner, through leadTerminals, into Principal.leader(...), so Authz can compare the caller's fleet with the target's.

Both land in entryOf and in Entry. One worker, one PR.

Why step 1 is still the harder half. The scanner change is small; the reach is not. CallerResolver.resolve builds the lead principal from a Map<terminal, leadName>:

String lead = leadTerminals.get().get(c.terminal());
if (lead != null) {
    return Principal.leader(lead, c.terminal(), c.pid());
}

Principal is record Principal(Role role, String terminal, long pid, String name) — no space field — and Authz has zero references to a space today, so isPrimary() is daemon-wide. That is the whole reason a second space cannot be a second fleet right now.

The rule the change must encode, and the trap in it. One space is one fleet, so a lead's rights stop at its own space. But the rights are not uniform, and getting this backwards would break exactly what the operator asked for:

Verb Across fleets
SPAWN, STOP, DRAIN, HANDOVER refuse — a lead must never stop another fleet's member
SEND lead → lead allow — this is cross-team leader coordination, and it is the point of the ticket
SEND lead → a member of another fleet refuse — that is delegating into someone else's fleet
READ, METRICS allow — seeing the host is not acting on it

Authz line 133 is case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary();. A naive "same fleet" clause added to the whole switch would take the lead→lead SEND with it and leave the feature dead on arrival. That failure would be silent in the obvious direction too: with only one registered space on this host, a test that never builds a second fleet cannot tell a correct rule from one that refuses everything cross-fleet. Any change here needs a test with two fleets in it, and a positive case asserting the lead→lead send is allowed — not only negative cases asserting the refusals.

Ordering against the work in flight. Authz is being edited for #778 right now, so this unit cannot start until that merges. #782 is confined to PromptBox and does not touch either.

## Sequencing: step 1 and #777 must be one unit, not two I read the code to size step 1 and found that it and #777 change the same two members of `LeadTabScanner`. Splitting them guarantees a merge collision, so they should go to one worker. **Where the space already is.** The scan loop holds the workspace in hand while it classifies each tab: ```java for (... ws ...) { if (ws.workspaceId() == null || excludedWorkspaceLabels.contains(ws.label())) { ... } Map<String, String> leadLabelsHere = leadLabelsFor(ws.label()); ... Entry entry = entryOf(tab.label(), leadLabelsHere); ``` So `ws.label()` is available at the point each `Entry` is built. `Entry` is `private record Entry(String name, Kind kind)` — adding the space is a one-field change, and it is the same field both tickets need. **What each ticket needs from it.** - **#777** needs `entryOf` to key the collaborator lookup by space. Today the lead branch reads `leadLabelsHere`, which is space-scoped, and the very next line reads `collaboratorTabToName`, which is flat: ```java String leadName = leadLabelsHere.get(normalized); if (leadName != null) return new Entry(leadName, Kind.LEAD); String collaboratorName = collaboratorTabToName.get(normalized); // no space key ``` - **step 1** needs the space to travel out of the scanner, through `leadTerminals`, into `Principal.leader(...)`, so `Authz` can compare the caller's fleet with the target's. Both land in `entryOf` and in `Entry`. One worker, one PR. **Why step 1 is still the harder half.** The scanner change is small; the reach is not. `CallerResolver.resolve` builds the lead principal from a `Map<terminal, leadName>`: ```java String lead = leadTerminals.get().get(c.terminal()); if (lead != null) { return Principal.leader(lead, c.terminal(), c.pid()); } ``` `Principal` is `record Principal(Role role, String terminal, long pid, String name)` — no space field — and `Authz` has zero references to a space today, so `isPrimary()` is daemon-wide. That is the whole reason a second space cannot be a second fleet right now. **The rule the change must encode, and the trap in it.** One space is one fleet, so a lead's rights stop at its own space. But the rights are not uniform, and getting this backwards would break exactly what the operator asked for: | Verb | Across fleets | |---|---| | `SPAWN`, `STOP`, `DRAIN`, `HANDOVER` | refuse — a lead must never stop another fleet's member | | `SEND` lead → lead | **allow** — this is cross-team leader coordination, and it is the point of the ticket | | `SEND` lead → a member of another fleet | refuse — that is delegating into someone else's fleet | | `READ`, `METRICS` | allow — seeing the host is not acting on it | `Authz` line 133 is `case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary();`. A naive "same fleet" clause added to the whole switch would take the lead→lead `SEND` with it and leave the feature dead on arrival. That failure would be silent in the obvious direction too: with only one registered space on this host, a test that never builds a second fleet cannot tell a correct rule from one that refuses everything cross-fleet. Any change here needs a test with **two** fleets in it, and a positive case asserting the lead→lead send is allowed — not only negative cases asserting the refusals. **Ordering against the work in flight.** `Authz` is being edited for #778 right now, so this unit cannot start until that merges. #782 is confined to `PromptBox` and does not touch either.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#779