A role with no pool and no charter starts anyway, and nothing at boot says so #613

Closed
opened 2026-09-20 12:30:38 +02:00 by ltms · 2 comments
Owner

Found live on this host on 2026-09-20 while running the fleetd #609 unit. fleet_list reported my hunter member with charterSource: "none" and charterBytes: 799, while the dev members on the same host reported charterSource: "fleet.charters.dev" and charterBytes: 1775.

The cause was that our fleetd.yaml had no fleet.hunters: pool and no fleet.charters.hunter: entry. Both gaps are silent, and they have different consequences.

What I measured

I ran the live daemon's own loader (java -cp fleetd/target/fleetd.jar against a small FleetConfig.load script), on the config before my fix and after it. profiles: has 8 entries.

before after
candidateProfiles(HUNTER) [local, local-direct, gx, opus, sonnet, sol, terra, xf] [sonnet, terra, local]
defaultProfileFor(HUNTER) local sonnet
charterFor(HUNTER) null 973 bytes

validateMembers() returned ok in both cases.

Why the pool half is the worse one

A missing pool is not a failure. It is documented as "unconstrained" in FleetConfig.candidateProfiles (FleetConfig.java:1817) and in CompositePeerLauncher.poolFor (CompositePeerLauncher.java:598-603): an empty pool falls back to every configured profile. That fallback is deliberate and I am not asking for it to change — it is what lets a config with profiles: and no fleet: still spawn.

The problem is what it does in a config that does declare pools for the other three roles. An unqualified role: hunter spawn silently widened to all 8 profiles, and the fixed-policy first choice became local — a profile this same config gives weight 0 to in every other pool, so it is disabled everywhere an operator looks. Nobody chose that. It fell out of definition order in profiles:.

Why the charter half is less bad

A missing charter already has a working instrument. HerdrPeerLauncher.logCharterReceipt (HerdrPeerLauncher.java:549) logs charterSource= on every spawn, and SessionManager.java:887 puts it on the fleet_list roster row. That is exactly how I found this. So the charter gap is observable after the fact, per member.

Still, the member ran with only the launcher's 799-byte reply charter, so it had no role contract at all. It behaved, because the brief named the hunter skill, but the charter is the layer that reaches every backend and the skill is not.

What I think the fix is

Not a refusal. "Unconstrained" is legitimate, and refusing to start would break the no-fleet:-block case that the fallback exists for.

A boot log line is enough: after validateMembers(), list the roles in MemberRole.values() that have an empty pool, and separately the ones with no charterFor, naming what each will do instead — "hunter has no fleet.hunters: pool, so an unqualified spawn may land on any of 8 profiles, first local". The operator then sees it once per restart, at the moment they can act on it, instead of discovering it from a roster row weeks later.

Worth checking whether the same line should name the resolved first-choice profile per role, since that is the number that surprised me, not the pool membership.

Not part of this issue

I have already added fleet.hunters: and fleet.charters.hunter: to this host's fleetd.yaml. That file is gitignored, so that change is invisible to the repo and it takes effect only at the next daemon restart — which is why I am recording the measurement here. It fixes one host. It does not fix the silence.

Found live on this host on 2026-09-20 while running the fleetd #609 unit. `fleet_list` reported my hunter member with `charterSource: "none"` and `charterBytes: 799`, while the dev members on the same host reported `charterSource: "fleet.charters.dev"` and `charterBytes: 1775`. The cause was that our `fleetd.yaml` had no `fleet.hunters:` pool and no `fleet.charters.hunter:` entry. Both gaps are silent, and they have different consequences. ## What I measured I ran the live daemon's own loader (`java -cp fleetd/target/fleetd.jar` against a small `FleetConfig.load` script), on the config before my fix and after it. `profiles:` has 8 entries. | | before | after | |---|---|---| | `candidateProfiles(HUNTER)` | `[local, local-direct, gx, opus, sonnet, sol, terra, xf]` | `[sonnet, terra, local]` | | `defaultProfileFor(HUNTER)` | `local` | `sonnet` | | `charterFor(HUNTER)` | `null` | 973 bytes | `validateMembers()` returned ok in both cases. ## Why the pool half is the worse one A missing pool is not a failure. It is documented as "unconstrained" in `FleetConfig.candidateProfiles` (`FleetConfig.java:1817`) and in `CompositePeerLauncher.poolFor` (`CompositePeerLauncher.java:598-603`): an empty pool falls back to **every** configured profile. That fallback is deliberate and I am not asking for it to change — it is what lets a config with `profiles:` and no `fleet:` still spawn. The problem is what it does in a config that *does* declare pools for the other three roles. An unqualified `role: hunter` spawn silently widened to all 8 profiles, and the `fixed`-policy first choice became `local` — a profile this same config gives weight 0 to in every other pool, so it is disabled everywhere an operator looks. Nobody chose that. It fell out of definition order in `profiles:`. ## Why the charter half is less bad A missing charter already has a working instrument. `HerdrPeerLauncher.logCharterReceipt` (`HerdrPeerLauncher.java:549`) logs `charterSource=` on every spawn, and `SessionManager.java:887` puts it on the `fleet_list` roster row. That is exactly how I found this. So the charter gap is observable after the fact, per member. Still, the member ran with only the launcher's 799-byte reply charter, so it had no role contract at all. It behaved, because the brief named the `hunter` skill, but the charter is the layer that reaches every backend and the skill is not. ## What I think the fix is Not a refusal. "Unconstrained" is legitimate, and refusing to start would break the no-`fleet:`-block case that the fallback exists for. A boot log line is enough: after `validateMembers()`, list the roles in `MemberRole.values()` that have an empty pool, and separately the ones with no `charterFor`, naming what each will do instead — "hunter has no `fleet.hunters:` pool, so an unqualified spawn may land on any of 8 profiles, first `local`". The operator then sees it once per restart, at the moment they can act on it, instead of discovering it from a roster row weeks later. Worth checking whether the same line should name the resolved first-choice profile per role, since that is the number that surprised me, not the pool membership. ## Not part of this issue I have already added `fleet.hunters:` and `fleet.charters.hunter:` to this host's `fleetd.yaml`. That file is gitignored, so that change is invisible to the repo and it takes effect only at the next daemon restart — which is why I am recording the measurement here. It fixes one host. It does not fix the silence.
Author
Owner

Both halves are now live on the Mac host, and the before/after is measurable in one log file.

The daemon was redeployed at 17:45 (jar 597af9057197 to fddfdd6c9794, pid 25112 to 59847). Then I spawned a hunter with no profile argument — the unqualified case this issue is about.

Pool

It landed on sonnet. Before the config fix, defaultProfileFor(HUNTER) resolved to local, so the same call would have picked a weight-0 profile.

The boot log also now lists the pool:

member slots: 12 configured [architect:opus, architect:sol, dev:sonnet, dev:terra, dev:xf,
  dev:local, hunter:sonnet, hunter:terra, hunter:local, reviewer:sonnet, reviewer:terra,
  reviewer:local]

Charter

Two spawned role=hunter profile=sonnet lines, same host, same role, same profile, one before the fix and one after:

16:55:07  charterSource=none                   charterBytes=799
17:46:09  charterSource=fleet.charters.hunter  charterBytes=1774

1774 is 973 (the hunter charter) + 2 (the \n\n separator) + 799 (the launcher reply charter), which matches how HerdrPeerLauncher.spawnInternal composes them.

This does not close the issue

What is fixed is one host's config. What this issue asks for is the boot line, and nothing about that has changed: a role with no pool still starts silently, and the resolved first-choice profile still appears nowhere.

One more line worth adding while someone is in there

The heartbeat's own boot line has the same shape of gap:

idle-lead heartbeat: on — nudge lead after 300s idle (recheck 60000ms, quiet cap 0)

It names three of the four settings. contextHighNudge (fleetd #609) is missing, so the log cannot tell an operator whether the handover notice is armed. I had to load the config with the deployed jar to confirm it was true. Same fix, same place: print the value that decides the behaviour, not just the ones that happen to be numbers.

Both halves are now live on the Mac host, and the before/after is measurable in one log file. The daemon was redeployed at 17:45 (jar `597af9057197` to `fddfdd6c9794`, pid 25112 to 59847). Then I spawned a hunter with **no profile argument** — the unqualified case this issue is about. ### Pool It landed on `sonnet`. Before the config fix, `defaultProfileFor(HUNTER)` resolved to `local`, so the same call would have picked a weight-0 profile. The boot log also now lists the pool: ``` member slots: 12 configured [architect:opus, architect:sol, dev:sonnet, dev:terra, dev:xf, dev:local, hunter:sonnet, hunter:terra, hunter:local, reviewer:sonnet, reviewer:terra, reviewer:local] ``` ### Charter Two `spawned role=hunter profile=sonnet` lines, same host, same role, same profile, one before the fix and one after: ``` 16:55:07 charterSource=none charterBytes=799 17:46:09 charterSource=fleet.charters.hunter charterBytes=1774 ``` 1774 is 973 (the hunter charter) + 2 (the `\n\n` separator) + 799 (the launcher reply charter), which matches how `HerdrPeerLauncher.spawnInternal` composes them. ### This does not close the issue What is fixed is one host's config. What this issue asks for is the boot line, and nothing about that has changed: a role with no pool still starts silently, and the resolved first-choice profile still appears nowhere. ### One more line worth adding while someone is in there The heartbeat's own boot line has the same shape of gap: ``` idle-lead heartbeat: on — nudge lead after 300s idle (recheck 60000ms, quiet cap 0) ``` It names three of the four settings. `contextHighNudge` (fleetd #609) is missing, so the log cannot tell an operator whether the handover notice is armed. I had to load the config with the deployed jar to confirm it was `true`. Same fix, same place: print the value that decides the behaviour, not just the ones that happen to be numbers.
Author
Owner

Fixed and merged as 9ee16f5 (PR #616)

Both halves landed: the role-fallback boot report, and contextHighNudge added to the lead-heartbeat boot line.

Verified by me, not taken from the worker's report

  • Own build in a scratch worktree, mvn -o clean install, unpiped: Tests run: 1870, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. main was 1864, so +6 (4 new RoleFallbackGapReportTest, 2 new LeadHeartbeatLoopTest).
  • Right accessor. The report tests fleet().profilesFor(role).isEmpty(). That matters: profilesFor (FleetConfig.java:1292) returns the raw pool and can be empty, while candidateProfiles (:1828) applies the fallback and therefore never is. Using the latter would have made the report a permanent no-op that still looked correct.
  • Tests pin behaviour, not source text. Both files attach a real ch.qos.logback.core.read.ListAppender and assert on getFormattedMessage().
  • No behaviour change. Fallback logic untouched, and a profiles:-only config with no fleet: block still loads.
  • Nothing forbidden committed — four files, all under fleetd/src. No fleetd.yaml, no .mcp.json, no wiki pointer.

Dogfooded against the real config, with a control

Ran the merged method through the built jar against this host's live fleetd/fleetd.yaml:

--- invoking against the LIVE config ---
--- done ---

Nothing. That zero proves nothing on its own, so I paired it with a positive control — a profiles:-only config with two profiles and no fleet: block:

INFO  dev.ltms.fleet.Fleetd - role fallback: no fleet.<role>s: pool for [architect (may land on
any of 2 profile(s), first choice 'alpha'), dev (...), hunter (...), reviewer (...)] — an
unqualified spawn of that role falls back to every configured profile (deliberate; see
FleetConfig#candidateProfiles)
INFO  dev.ltms.fleet.Fleetd - role fallback: no fleet.charters: entry for [architect, dev,
hunter, reviewer] — that role runs with only the launcher's reply charter, no role contract

So the report works and INFO prints at runtime. The silence on the live config is a real "no gaps on this host" — the fleet.hunters: and fleet.charters.hunter: entries added on 2026-09-20 closed the gap this ticket was opened about. The line will now say so at boot on any host that still has it.

Two things this does NOT do

  1. The call site in Fleetd.main is not pinned. I measured it: replacing reportRoleFallbackGaps(cfg) with a comment leaves the suite fully green (1870, 0 failures), while making the method body inert turns it red (2 failures, 1 error). The method is tested; its invocation is not. That is #612's shape and I have recorded it there as a 17th site rather than papering over it with a source-text assertion. Merged with that gap open and tracked.
  2. The report fires at startup only, not on reload. fleet pools and charters are hot — they can change without a restart — so an operator who introduces a gap by reloading will not see the line until the next restart. The ticket asked for "once per restart" and that is what shipped, so this is a scope note, not a defect. Worth a follow-up if it bites.

Not merged, still true

The daemon is still running the jar built from 076cc43. A merge is not a deployment, so these lines will not appear in the live log until the next redeploy.

Closing this ticket. The two items above are tracked where they belong.

## Fixed and merged as `9ee16f5` (PR #616) Both halves landed: the role-fallback boot report, and `contextHighNudge` added to the lead-heartbeat boot line. ### Verified by me, not taken from the worker's report - **Own build** in a scratch worktree, `mvn -o clean install`, unpiped: `Tests run: 1870, Failures: 0, Errors: 0, Skipped: 0` — BUILD SUCCESS. `main` was 1864, so +6 (4 new `RoleFallbackGapReportTest`, 2 new `LeadHeartbeatLoopTest`). - **Right accessor.** The report tests `fleet().profilesFor(role).isEmpty()`. That matters: `profilesFor` (`FleetConfig.java:1292`) returns the raw pool and can be empty, while `candidateProfiles` (`:1828`) applies the fallback and therefore never is. Using the latter would have made the report a permanent no-op that still looked correct. - **Tests pin behaviour, not source text.** Both files attach a real `ch.qos.logback.core.read.ListAppender` and assert on `getFormattedMessage()`. - **No behaviour change.** Fallback logic untouched, and a `profiles:`-only config with no `fleet:` block still loads. - **Nothing forbidden committed** — four files, all under `fleetd/src`. No `fleetd.yaml`, no `.mcp.json`, no `wiki` pointer. ### Dogfooded against the real config, with a control Ran the merged method through the built jar against this host's live `fleetd/fleetd.yaml`: ``` --- invoking against the LIVE config --- --- done --- ``` Nothing. That zero proves nothing on its own, so I paired it with a positive control — a `profiles:`-only config with two profiles and no `fleet:` block: ``` INFO dev.ltms.fleet.Fleetd - role fallback: no fleet.<role>s: pool for [architect (may land on any of 2 profile(s), first choice 'alpha'), dev (...), hunter (...), reviewer (...)] — an unqualified spawn of that role falls back to every configured profile (deliberate; see FleetConfig#candidateProfiles) INFO dev.ltms.fleet.Fleetd - role fallback: no fleet.charters: entry for [architect, dev, hunter, reviewer] — that role runs with only the launcher's reply charter, no role contract ``` So the report works and INFO prints at runtime. The silence on the live config is a real "no gaps on this host" — the `fleet.hunters:` and `fleet.charters.hunter:` entries added on 2026-09-20 closed the gap this ticket was opened about. The line will now say so at boot on any host that still has it. ### Two things this does NOT do 1. **The call site in `Fleetd.main` is not pinned.** I measured it: replacing `reportRoleFallbackGaps(cfg)` with a comment leaves the suite fully green (1870, 0 failures), while making the method body inert turns it red (2 failures, 1 error). The method is tested; its invocation is not. That is #612's shape and I have recorded it there as a 17th site rather than papering over it with a source-text assertion. Merged with that gap open and tracked. 2. **The report fires at startup only, not on reload.** `fleet` pools and charters are hot — they can change without a restart — so an operator who introduces a gap by reloading will not see the line until the next restart. The ticket asked for "once per restart" and that is what shipped, so this is a scope note, not a defect. Worth a follow-up if it bites. ### Not merged, still true The daemon is still running the jar built from `076cc43`. A merge is not a deployment, so these lines will not appear in the live log until the next redeploy. Closing this ticket. The two items above are tracked where they belong.
ltms closed this issue 2026-09-22 05:18:25 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#613