Lead identity must key on the space, not the tab label: one host, many teams, each space's first tab named "lead" #770

Open
opened 2026-10-05 11:29:13 +02:00 by ltms · 4 comments
Owner

Operator requirement, 2026-10-05:

for a space, lead tab is the first tab of space and always named "lead", no "lead: opus" like now. This is to prepare for new model in which one host can have multiple team

This is not a rename. It inverts the current identity rule

Today the discriminator is the label, and the space is explicitly not consulted. Measured in the code:

  • herdr/LeadTabScanner.buildTabIndex builds one flat Map<String, Entry> keyed by the normalised tab label, merging leads and collaborators into a single global index. A lookup is "a plain map hit" on that label.
  • The workspace appears only as excludedWorkspaceLabels — a set of spaces never scanned. It is a filter, never part of the key.
  • config/FleetConfig.Leader.workspace defaults to DEFAULT_WORKSPACE = "fleet", and the javadoc states the intent plainly: all leads share one space so "the operator sees one 'session' with many tabs", and "it tells a lead from a member by the exact tab label".
  • FleetConfig validation rejects two fleet.leaders entries that share one exact tab, case-insensitively.

So under the new convention every team's lead tab is called lead, and: the scanner's map has one key for all of them, and config validation refuses the second team outright. The current design has the uniqueness burden on the label, and the requirement removes exactly that uniqueness.

The new model moves the burden to the space: (workspace, "lead") is unique, "lead" alone is not.

What has to change

  1. Key lead identity on the space. LeadTabScanner's index becomes per-workspace — (workspaceLabel, tabLabel) → Entry, or a per-space index. Every caller that resolves a terminal to a lead follows.
  2. Invert the workspace's role. It stops being an exclusion list and becomes the team boundary. excludedWorkspaceLabels needs re-examining: under one-space-per-team, a space either belongs to a team or is not ours, and "the shared fleet space" stops being the normal case.
  3. Relax and re-tighten the duplicate-tab validation. Two leaders may share the label lead in different spaces, and must still be refused in the same space.
  4. Decide what "first tab of the space" means for identity. Nothing implements a positional rule today. Three options, and they are not equivalent: the label alone identifies the lead; the position alone does; or both must agree. If both, the behaviour when they disagree has to be specified — and a tab can be reordered or closed at runtime, so position is mutable state, not a stable key. My reading is that position should be a convention the launcher creates and validation can warn about, while the label stays the match. But that is a decision, not an obvious call; record it here.
  5. Re-check the prefix collision. Leader.tabPrefix defaults to "lead:", and validation refuses a member tab template that "could match a lead-tab naming convention". A bare lead does not start with lead:, so today's prefix check would not protect the new label. Whatever the new convention is, the member-template check must cover it, or a member tab could be named so that it resolves as a lead.

Operational hazard — order matters, and getting it wrong demotes a live lead

A lead is found by its configured tab label. Rename the tab before fleetd.yaml names the new label and that pane stops matching any lead, so it resolves as a worker and every orchestration call is refused. The redeploy-fleetd skill already warns about this for restarts; a rename hits it without any restart.

So: change fleet.leaders.<name>.tab first, then rename the tab. My own tab is lead: opus right now, so this applies to this very session.

I cannot do the config half — writing fleetd.yaml is refused to me by the permission classifier, and the operator owns that file.

Instruction surface and recorded facts that go false

Two things currently documented as true stop being true, and both are load-bearing:

  • "a session in a herdr pane is demoted to worker by any fleetd not naming its tab" — still true, but the tab is no longer unique on the host.
  • "the tab label is the only test; workspace is NOT consulted" — this becomes exactly wrong.

So the change must also update CLAUDE.md's fleet_whoami / fallback-ladder paragraph, the redeploy-fleetd skill's identity check, and wiki/11-Features.md. Under this repo's "the prompt is part of the product" rule that is part of the change, not follow-up.

Suggested shape

Land it in two steps so no team is ever half-identified:

  1. Make identity space-aware while (workspace, label) still accepts today's lead: opus. No behaviour change for a single team; the validation and the scanner key move.
  2. Then switch the convention to the bare lead and add the first-tab rule in the launcher, with the config edit ahead of the tab rename.

Blocked on nothing except the decision in point 4 and the config edit.

Operator requirement, 2026-10-05: > for a space, lead tab is the first tab of space and always named "lead", no "lead: opus" like now. This is to prepare for new model in which one host can have multiple team ## This is not a rename. It inverts the current identity rule Today the discriminator **is** the label, and the space is explicitly not consulted. Measured in the code: - `herdr/LeadTabScanner.buildTabIndex` builds one flat `Map<String, Entry>` keyed by the **normalised tab label**, merging leads and collaborators into a single global index. A lookup is "a plain map hit" on that label. - The workspace appears only as `excludedWorkspaceLabels` — a set of spaces never scanned. It is a filter, never part of the key. - `config/FleetConfig.Leader.workspace` defaults to `DEFAULT_WORKSPACE = "fleet"`, and the javadoc states the intent plainly: all leads share one space so "the operator sees one 'session' with many tabs", and "it tells a lead from a member by the exact tab label". - `FleetConfig` validation **rejects two `fleet.leaders` entries that share one exact tab**, case-insensitively. So under the new convention every team's lead tab is called `lead`, and: the scanner's map has one key for all of them, and config validation refuses the second team outright. The current design has the uniqueness burden on the label, and the requirement removes exactly that uniqueness. The new model moves the burden to the **space**: `(workspace, "lead")` is unique, `"lead"` alone is not. ## What has to change 1. **Key lead identity on the space.** `LeadTabScanner`'s index becomes per-workspace — `(workspaceLabel, tabLabel) → Entry`, or a per-space index. Every caller that resolves a terminal to a lead follows. 2. **Invert the workspace's role.** It stops being an exclusion list and becomes the team boundary. `excludedWorkspaceLabels` needs re-examining: under one-space-per-team, a space either belongs to a team or is not ours, and "the shared `fleet` space" stops being the normal case. 3. **Relax and re-tighten the duplicate-tab validation.** Two leaders may share the label `lead` in *different* spaces, and must still be refused in the *same* space. 4. **Decide what "first tab of the space" means for identity.** Nothing implements a positional rule today. Three options, and they are not equivalent: the label alone identifies the lead; the position alone does; or both must agree. If both, the behaviour when they disagree has to be specified — and a tab can be reordered or closed at runtime, so position is mutable state, not a stable key. My reading is that position should be a **convention the launcher creates and validation can warn about**, while the label stays the match. But that is a decision, not an obvious call; record it here. 5. **Re-check the prefix collision.** `Leader.tabPrefix` defaults to `"lead:"`, and validation refuses a member tab template that "could match a lead-tab naming convention". A bare `lead` does not start with `lead:`, so today's prefix check would not protect the new label. Whatever the new convention is, the member-template check must cover it, or a member tab could be named so that it resolves as a lead. ## Operational hazard — order matters, and getting it wrong demotes a live lead A lead is found by its configured tab label. Rename the tab before `fleetd.yaml` names the new label and that pane stops matching any lead, so it resolves as a **worker** and every orchestration call is refused. The `redeploy-fleetd` skill already warns about this for restarts; a rename hits it without any restart. So: change `fleet.leaders.<name>.tab` **first**, then rename the tab. My own tab is `lead: opus` right now, so this applies to this very session. I cannot do the config half — writing `fleetd.yaml` is refused to me by the permission classifier, and the operator owns that file. ## Instruction surface and recorded facts that go false Two things currently documented as true stop being true, and both are load-bearing: - "a session in a herdr pane is demoted to worker by any fleetd not naming its tab" — still true, but the tab is no longer unique on the host. - "the tab label is the only test; workspace is NOT consulted" — this becomes exactly wrong. So the change must also update `CLAUDE.md`'s `fleet_whoami` / fallback-ladder paragraph, the `redeploy-fleetd` skill's identity check, and `wiki/11-Features.md`. Under this repo's "the prompt is part of the product" rule that is part of the change, not follow-up. ## Suggested shape Land it in two steps so no team is ever half-identified: 1. Make identity space-aware while `(workspace, label)` still accepts today's `lead: opus`. No behaviour change for a single team; the validation and the scanner key move. 2. Then switch the convention to the bare `lead` and add the first-tab rule in the launcher, with the config edit ahead of the tab rename. Blocked on nothing except the decision in point 4 and the config edit.
Author
Owner

The addressing half is now #771: "with cross spaces communicate, we need to identify the space -> tab (names)".

The two must be decided together. If lead identity keys on (space, "lead") then (space, tab) is exactly the address, and space has to mean the same thing in both places. #771 also turned up one concrete gap that blocks everything else and needs no design decision: fleet_list's panes rows report workspaceId: "w2" but never the space name, although herdr workspace list has it (label: "fleet"). I have delegated that one additive change under #771; it is correct under either identity model.

Point 4 here (what "first tab" means for identity) and point 4 of #771 (whether a name address is a task or coordination channel) are the two decisions still open.

The addressing half is now #771: *"with cross spaces communicate, we need to identify the space -> tab (names)"*. The two must be decided together. If lead identity keys on `(space, "lead")` then `(space, tab)` is exactly the address, and `space` has to mean the same thing in both places. #771 also turned up one concrete gap that blocks everything else and needs no design decision: `fleet_list`'s `panes` rows report `workspaceId: "w2"` but never the space **name**, although `herdr workspace list` has it (`label: "fleet"`). I have delegated that one additive change under #771; it is correct under either identity model. Point 4 here (what "first tab" means for identity) and point 4 of #771 (whether a name address is a task or coordination channel) are the two decisions still open.
Author
Owner

Operator decision, 2026-10-05 — point 4 is settled

that leader name become a fix contract -> we dont need further config for it, start #770 with this

So the lead tab label is a constant in code, not a config key. Two things follow, and together they answer point 4:

  1. The label is the identity, and it is always lead. Position is not identity. The operator's words settle this: a name is the contract, and a name is stable while a tab position is not — a tab can be reordered or closed at runtime. "First tab of the space" becomes what the launcher aims for when it creates the tab, and nothing matches on it.
  2. fleet.leaders.<name>.tab stops being required. It is not needed to find a lead any more, so it is deprecated.

This also removes the config edit the ticket said was blocking, and with it the hazard in the section above: no fleetd.yaml change has to land before the code.

The one way this change can do real damage

lead/LeadLauncher.leadNameOf matches a tab with Leader.tabLabel(), and countLeads uses it to decide how many leads are already live. ensureLeads() then launches the shortfall.

So if tabLabel() starts returning lead while the live tab is still called lead: opus, countLeads counts 0 live leads and the daemon launches a second Opus lead into the fleet space. Two orchestrators on one fleet is the failure this class was written to prevent.

Measured now, read-only:

herdr workspace list   → w2  label "fleet"   tab_count 1
herdr tab list -w w2   → w2:tY | 'lead: opus'
fleet_whoami           → {"role":"primary","leader":"opus","sessionId":"term_65d106559b02e1"}

So the live state really is the state that would trip this.

How the cutover stays safe

The accepted labels for a lead become a set per space: the constant lead, plus the deprecated tab: while it is still configured. Both the scanner and countLeads read that same set.

That makes every order safe. Before the rename, lead: opus still matches. After it, lead matches. There is no moment when the live lead matches nothing, and no second lead is ever launched. Deleting the tab: key later shrinks the set to the constant, and that is the whole of step 2.

Scope now in flight

One implementer unit, worktree 770-lead-tab-is-a-fixed-contract:

  • config/FleetConfig.Leader — a LEAD_TAB_LABEL = "lead" constant, tab: optional, an accepted-labels accessor.
  • validation — two leaders in one space refused (replacing "two leaders share one exact tab"); a member tabLabel template that renders as lead refused; a collaborator tab named lead refused.
  • herdr/LeadTabScanner — the lead index keyed on (space label, tab label). Collaborators keep today's space-agnostic exact match; nothing in the requirement asks to change them.
  • FleetdAssembly — builds the per-space index, and warns once per lead that still carries tab:.
  • lead/LeadLauncher — countLeads counts per space, against the accepted-labels set.

Not in this unit, and tracked separately: msg/LeadCoordLoop.resolveLocalLead() picks the sole lead when no name matches coordinator.selfId. That is correct for one team and becomes a guess for several. It needs the space as a discriminator too, and it is a different file and a different decision, so it is not bundled here.

Instruction surface — the lead's own half

Two recorded facts go false and both are load-bearing, so under the prompt is part of the product they are part of this change, not follow-up:

  • CLAUDE.md — "the tab label is the only test; workspace is NOT consulted" becomes exactly wrong.
  • the redeploy-fleetd skill, check 4 — it names fleet.leaders.*.tab as what identity matches.
  • wiki/11-Features.md — one entry for the new convention.

The CLAUDE.md canonical block must stay byte-identical with the wiki template, and only the lead can run that check, so this half is mine and not the worker's.

## Operator decision, 2026-10-05 — point 4 is settled > that leader name become a fix contract -> we dont need further config for it, start #770 with this So the lead tab label is a **constant in code**, not a config key. Two things follow, and together they answer point 4: 1. **The label is the identity**, and it is always `lead`. Position is not identity. The operator's words settle this: a *name* is the contract, and a name is stable while a tab position is not — a tab can be reordered or closed at runtime. "First tab of the space" becomes what the launcher aims for when it creates the tab, and nothing matches on it. 2. **`fleet.leaders.<name>.tab` stops being required.** It is not needed to find a lead any more, so it is deprecated. This also removes the config edit the ticket said was blocking, and with it the hazard in the section above: no `fleetd.yaml` change has to land before the code. ## The one way this change can do real damage `lead/LeadLauncher.leadNameOf` matches a tab with `Leader.tabLabel()`, and `countLeads` uses it to decide how many leads are already live. `ensureLeads()` then launches the shortfall. So if `tabLabel()` starts returning `lead` while the live tab is still called `lead: opus`, `countLeads` counts **0** live leads and the daemon **launches a second Opus lead** into the `fleet` space. Two orchestrators on one fleet is the failure this class was written to prevent. Measured now, read-only: ``` herdr workspace list → w2 label "fleet" tab_count 1 herdr tab list -w w2 → w2:tY | 'lead: opus' fleet_whoami → {"role":"primary","leader":"opus","sessionId":"term_65d106559b02e1"} ``` So the live state really is the state that would trip this. ## How the cutover stays safe The accepted labels for a lead become a **set per space**: the constant `lead`, plus the deprecated `tab:` while it is still configured. Both the scanner and `countLeads` read that same set. That makes every order safe. Before the rename, `lead: opus` still matches. After it, `lead` matches. There is no moment when the live lead matches nothing, and no second lead is ever launched. Deleting the `tab:` key later shrinks the set to the constant, and that is the whole of step 2. ## Scope now in flight One `implementer` unit, worktree `770-lead-tab-is-a-fixed-contract`: - `config/FleetConfig.Leader` — a `LEAD_TAB_LABEL = "lead"` constant, `tab:` optional, an accepted-labels accessor. - validation — two leaders in **one space** refused (replacing "two leaders share one exact tab"); a member `tabLabel` template that renders as `lead` refused; a collaborator tab named `lead` refused. - `herdr/LeadTabScanner` — the lead index keyed on `(space label, tab label)`. Collaborators keep today's space-agnostic exact match; nothing in the requirement asks to change them. - `FleetdAssembly` — builds the per-space index, and warns once per lead that still carries `tab:`. - `lead/LeadLauncher` — `countLeads` counts per space, against the accepted-labels set. Not in this unit, and tracked separately: `msg/LeadCoordLoop.resolveLocalLead()` picks the sole lead when no name matches `coordinator.selfId`. That is correct for one team and becomes a guess for several. It needs the space as a discriminator too, and it is a different file and a different decision, so it is not bundled here. ## Instruction surface — the lead's own half Two recorded facts go false and both are load-bearing, so under *the prompt is part of the product* they are part of this change, not follow-up: - `CLAUDE.md` — "the tab label is the only test; workspace is NOT consulted" becomes exactly wrong. - the `redeploy-fleetd` skill, check 4 — it names `fleet.leaders.*.tab` as what identity matches. - `wiki/11-Features.md` — one entry for the new convention. The `CLAUDE.md` canonical block must stay byte-identical with the wiki template, and only the lead can run that check, so this half is mine and not the worker's.
Author
Owner

The code half is merged and live — 7942378, docs at a3d296f

Deployed and verified, not just merged. New jar 2f245a1fec31, pid 23954, fleetd listening at 13:57:33, no ERROR lines since restart.

13:57:30.762 WARN  lead 'opus' (fleet.leaders.opus) still configures tab: "lead: opus" — deprecated,
                   the lead tab label is now fixed to 'lead'
13:57:30.767 INFO  lead/collaborator scan: space per lead {opus=fleet}
13:57:34.095 INFO  lead/collaborator panes: {term_65d106559b02e1=Entry[name=opus, kind=LEAD]}

fleet_whoami            → {"role":"primary","leader":"opus","sessionId":"term_65d106559b02e1"}
herdr tab list -w w2    → tab count 1,  w2:tY 'lead: opus'

That last line is the check that mattered. countLeads found the live lead through its legacy label and ensureLeads() launched nothing — had it matched the constant alone it would have counted zero and started a second orchestrator into the fleet space.

Points 1, 2, 3 and 5 are done. Point 4 is settled.

1. Identity keys on the space — LeadTabScanner holds a per-space index; LeadLauncher.leadNameOf takes the space.

2. The workspace's role is inverted — it is the team boundary now. excludedWorkspaceLabels was left alone deliberately, and the reason is worth recording rather than treating as an omission: production passes Set.of(), and it must stay empty, because populating it would skip the lead's own space. The exclusion mechanism is now redundant rather than wrong — the space is part of the key, so nothing needs filtering out. Removing the parameter would be a separate change with its own test cost, and it buys no behaviour.

3. Duplicate validation relaxed and re-tightened — two leaders in different spaces with no tab: at all are accepted; two in the same space are refused, and the message says why.

4. "First tab" is a convention, not identity — settled by the operator's own words. A name is a stable contract; a tab position is mutable state. So the label matches and the launcher creates the tab first in its space. Nothing matches on position, and no validation warns about it.

5. The prefix collision is covered — a member tabLabel template that can render as lead is refused, and so is a collaborator whose tab is lead. The default template {role}: {profile} #{n} cannot render it, so this is a guard against a future override rather than a live problem.

Instruction surface — done, and smaller than this ticket predicted

I claimed above that CLAUDE.md's "the tab label is the only test" sentence would go false. That was wrong, and I measured it rather than acting on it: that sentence is in my own session memory, not in the canonical block. The block's only tab claim is about fleet.collaborators.<name>.tab, which this change does not touch. The sync check still prints in sync: True.

What actually needed changing, both in a3d296f:

  • .claude/skills/redeploy-fleetd/SKILL.md check 4 — it told the reader a lead is found by fleet.leaders.*.tab.
  • scripts/redeploy-fleetd.sh closing hint — the same sentence, printed at the end of every redeploy. The shell suite still exits 0.
  • wiki/11-Features.md — one new entry, plus a gotcha added to the existing "Startup refuses placement: pane while a lead names a tab" entry, because #775 makes half of it false.

Still open on this ticket, and both are the operator's

  1. Rename w2:tY from lead: opus to lead. Renaming a fleet pane's tab is a herdr control-plane action, which invariant 5 reserves. Safe at any time — the accepted-label set means there is no window where the lead matches nothing.
  2. Then delete tab: "lead: opus" from fleetd.yaml. Blocked on #775, not on the rename. Dropping that key is what disarms validatePanePlacementAgainstLeadTabs, which this ticket's own change caused. A fix is in flight.

Not in this ticket

msg/LeadCoordLoop.resolveLocalLead() still falls back to "the sole lead" when no lead's name matches coordinator.selfId. Correct for one team, a guess for several. It needs the space as a discriminator and it is a separate decision.

## The code half is merged and live — `7942378`, docs at `a3d296f` Deployed and verified, not just merged. New jar `2f245a1fec31`, pid 23954, `fleetd listening` at 13:57:33, no ERROR lines since restart. ``` 13:57:30.762 WARN lead 'opus' (fleet.leaders.opus) still configures tab: "lead: opus" — deprecated, the lead tab label is now fixed to 'lead' 13:57:30.767 INFO lead/collaborator scan: space per lead {opus=fleet} 13:57:34.095 INFO lead/collaborator panes: {term_65d106559b02e1=Entry[name=opus, kind=LEAD]} fleet_whoami → {"role":"primary","leader":"opus","sessionId":"term_65d106559b02e1"} herdr tab list -w w2 → tab count 1, w2:tY 'lead: opus' ``` That last line is the check that mattered. `countLeads` found the live lead through its legacy label and `ensureLeads()` launched nothing — had it matched the constant alone it would have counted zero and started a second orchestrator into the `fleet` space. ## Points 1, 2, 3 and 5 are done. Point 4 is settled. **1. Identity keys on the space** — `LeadTabScanner` holds a per-space index; `LeadLauncher.leadNameOf` takes the space. **2. The workspace's role is inverted** — it is the team boundary now. `excludedWorkspaceLabels` was left alone deliberately, and the reason is worth recording rather than treating as an omission: production passes `Set.of()`, and it must stay empty, because populating it would skip the lead's own space. The exclusion mechanism is now redundant rather than wrong — the space is part of the key, so nothing needs filtering out. Removing the parameter would be a separate change with its own test cost, and it buys no behaviour. **3. Duplicate validation relaxed and re-tightened** — two leaders in *different* spaces with no `tab:` at all are accepted; two in the *same* space are refused, and the message says why. **4. "First tab" is a convention, not identity** — settled by the operator's own words. A name is a stable contract; a tab position is mutable state. So the label matches and the launcher creates the tab first in its space. Nothing matches on position, and no validation warns about it. **5. The prefix collision is covered** — a member `tabLabel` template that can render as `lead` is refused, and so is a collaborator whose `tab` is `lead`. The default template `{role}: {profile} #{n}` cannot render it, so this is a guard against a future override rather than a live problem. ## Instruction surface — done, and smaller than this ticket predicted I claimed above that `CLAUDE.md`'s "the tab label is the only test" sentence would go false. **That was wrong**, and I measured it rather than acting on it: that sentence is in my own session memory, not in the canonical block. The block's only tab claim is about `fleet.collaborators.<name>.tab`, which this change does not touch. The sync check still prints `in sync: True`. What actually needed changing, both in `a3d296f`: - `.claude/skills/redeploy-fleetd/SKILL.md` check 4 — it told the reader a lead is found by `fleet.leaders.*.tab`. - `scripts/redeploy-fleetd.sh` closing hint — the same sentence, printed at the end of every redeploy. The shell suite still exits 0. - `wiki/11-Features.md` — one new entry, plus a gotcha added to the existing *"Startup refuses `placement: pane` while a lead names a tab"* entry, because #775 makes half of it false. ## Still open on this ticket, and both are the operator's 1. **Rename `w2:tY` from `lead: opus` to `lead`.** Renaming a fleet pane's tab is a herdr control-plane action, which invariant 5 reserves. Safe at any time — the accepted-label set means there is no window where the lead matches nothing. 2. **Then delete `tab: "lead: opus"` from `fleetd.yaml`.** **Blocked on #775**, not on the rename. Dropping that key is what disarms `validatePanePlacementAgainstLeadTabs`, which this ticket's own change caused. A fix is in flight. ## Not in this ticket `msg/LeadCoordLoop.resolveLocalLead()` still falls back to "the sole lead" when no lead's name matches `coordinator.selfId`. Correct for one team, a guess for several. It needs the space as a discriminator and it is a separate decision.
Author
Owner

Step 2 done — this host is fully migrated

The operator renamed tab w2:tY to lead and authorized the config edit. tab: "lead: opus" is out of fleetd.yaml. Daemon restarted: pid 47129, jar 72801148bc3b (unchanged — config only), fleetd listening 16:44:19, no ERROR lines.

Order mattered, and I checked it before acting rather than after

Deleting tab: while the tab was still labelled lead: opus would have been a self-inflicted outage, so I measured the two states first:

acceptedLabels live tab label result
tab: present [lead, lead: opus] lead: opus match
tab: deleted, tab not yet renamed [lead] lead: opus no match

No match means both failures at once: the lead is demoted and refuses every orchestration call, and countLeads returns 0 against instances: 1, so ensureLeads() launches a second orchestrator into the same space. So the sequence is strictly rename first, then delete — and the rename is safe at any time, because lead is accepted either way.

I proved the stripped config was valid before it went live, by loading it through the real FleetConfig.load and validatePanePlacementAgainstLeadTabs against the deployed jar:

load: OK
validatePanePlacementAgainstLeadTabs: did NOT refuse   (correct — every profile is placement: tab)
lead.tab()          = null
lead.tabLabel()     = lead
lead.acceptedLabels = [lead]
lead.workspace()    = fleet

That is also the first live exercise of #775 and step 2 together: the re-armed guard sees a lead with no tab: and correctly does not refuse, because no profile is pane-placed.

Verified after the restart

fleet_whoami                          → {"role":"primary","leader":"opus",...}
space per lead                        → {opus=fleet}
lead/collaborator panes               → {term_65d106559b02e1=Entry[name=opus, kind=LEAD]}
w2 tab count                          → 1   (label 'lead')  — no second lead launched
deprecation WARN 'still configures tab' → 0

My first attempt at that last count said STILL PRESENT, and it was wrong: my log window reached back 60 lines and spanned two boots. The control caught it — fleet health: detection-only counted 2, and there is exactly one per boot. Re-anchored the window between the last two fleetd listening lines (97362–97396), where both controls read 1 and the deprecation WARN reads 0. fleetd.out carries no date, so a line-count window is not a boot.

tabPrefix: "lead:" stays, and it is not leftover

It has exactly one reader, and that reader is still doing real work: FleetConfig.java:2809 refuses a member tabLabel template that collides with a lead's tab or starts with that prefix. With tab: gone, leader.tab() is null and the template check short-circuits, but the prefix check still guards against a future member label like lead: …. Keeping it costs nothing and removing it would drop a live guard.

The legacy branch in acceptedLabels() is deliberately NOT removed yet

Step 2 as filed also wanted the legacy label branch deleted. I am leaving it, and the reason is not caution for its own sake:

if (normalizedTab == null || normalizedTab.equals(LEAD_TAB_LABEL)) {
    return List.of(LEAD_TAB_LABEL);
}
return List.of(LEAD_TAB_LABEL, normalizedTab);   // ← this line stays

That branch is what makes the cutover order-free, and this host is not the only host. fleet01 runs its own fleetd from its own checkout with its own config. If that config still sets tab: and fleet01 later deploys a main without this branch, its lead matches nothing and the same double failure happens there — demoted lead plus a second orchestrator — with nobody watching for it.

The deprecation WARN is the migration signal, so the branch should outlive the WARN being acted on everywhere, not the WARN being added. Removing it is a separate ticket, and its real precondition is "every host that deploys this repo has no tab: in its config", which is a fact about other machines and cannot be checked from here. I am messaging fleet01's lead with the hazard and the required order.

Open on this ticket

Nothing on this host. Remaining work is the cross-host migration above, plus the item already noted as out of scope: msg/LeadCoordLoop.resolveLocalLead() still falls back to "the sole lead" when no name matches coordinator.selfId, which is a guess once a host runs several teams.

## Step 2 done — this host is fully migrated The operator renamed tab `w2:tY` to `lead` and authorized the config edit. `tab: "lead: opus"` is out of `fleetd.yaml`. Daemon restarted: pid 47129, jar `72801148bc3b` (unchanged — config only), `fleetd listening` 16:44:19, no ERROR lines. ### Order mattered, and I checked it before acting rather than after Deleting `tab:` while the tab was still labelled `lead: opus` would have been a self-inflicted outage, so I measured the two states first: | | `acceptedLabels` | live tab label | result | |---|---|---|---| | `tab:` present | `[lead, lead: opus]` | `lead: opus` | match | | `tab:` deleted, **tab not yet renamed** | `[lead]` | `lead: opus` | **no match** | No match means both failures at once: the lead is demoted and refuses every orchestration call, **and** `countLeads` returns 0 against `instances: 1`, so `ensureLeads()` launches a second orchestrator into the same space. So the sequence is strictly **rename first, then delete** — and the rename is safe at any time, because `lead` is accepted either way. I proved the stripped config was valid before it went live, by loading it through the real `FleetConfig.load` and `validatePanePlacementAgainstLeadTabs` against the deployed jar: ``` load: OK validatePanePlacementAgainstLeadTabs: did NOT refuse (correct — every profile is placement: tab) lead.tab() = null lead.tabLabel() = lead lead.acceptedLabels = [lead] lead.workspace() = fleet ``` That is also the first live exercise of #775 and step 2 together: the re-armed guard sees a lead with no `tab:` and correctly does not refuse, because no profile is pane-placed. ### Verified after the restart ``` fleet_whoami → {"role":"primary","leader":"opus",...} space per lead → {opus=fleet} lead/collaborator panes → {term_65d106559b02e1=Entry[name=opus, kind=LEAD]} w2 tab count → 1 (label 'lead') — no second lead launched deprecation WARN 'still configures tab' → 0 ``` My first attempt at that last count said **STILL PRESENT**, and it was wrong: my log window reached back 60 lines and spanned two boots. The control caught it — `fleet health: detection-only` counted **2**, and there is exactly one per boot. Re-anchored the window between the last two `fleetd listening` lines (97362–97396), where both controls read 1 and the deprecation WARN reads 0. `fleetd.out` carries no date, so a line-count window is not a boot. ### `tabPrefix: "lead:"` stays, and it is not leftover It has exactly one reader, and that reader is still doing real work: `FleetConfig.java:2809` refuses a member `tabLabel` template that collides with a lead's tab or starts with that prefix. With `tab:` gone, `leader.tab()` is null and the template check short-circuits, but the **prefix** check still guards against a future member label like `lead: …`. Keeping it costs nothing and removing it would drop a live guard. ### The legacy branch in `acceptedLabels()` is deliberately NOT removed yet Step 2 as filed also wanted the legacy label branch deleted. I am leaving it, and the reason is not caution for its own sake: ```java if (normalizedTab == null || normalizedTab.equals(LEAD_TAB_LABEL)) { return List.of(LEAD_TAB_LABEL); } return List.of(LEAD_TAB_LABEL, normalizedTab); // ← this line stays ``` That branch is what makes the cutover order-free, and **this host is not the only host**. `fleet01` runs its own `fleetd` from its own checkout with its own config. If that config still sets `tab:` and fleet01 later deploys a `main` without this branch, its lead matches nothing and the same double failure happens there — demoted lead plus a second orchestrator — with nobody watching for it. The deprecation WARN is the migration signal, so the branch should outlive the WARN being acted on everywhere, not the WARN being added. Removing it is a separate ticket, and its real precondition is *"every host that deploys this repo has no `tab:` in its config"*, which is a fact about other machines and cannot be checked from here. I am messaging fleet01's lead with the hazard and the required order. ### Open on this ticket Nothing on this host. Remaining work is the cross-host migration above, plus the item already noted as out of scope: `msg/LeadCoordLoop.resolveLocalLead()` still falls back to "the sole lead" when no name matches `coordinator.selfId`, which is a guess once a host runs several teams.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#770