Features: fleet_list enumerates the startup profile set (#416)

Dai Ha
2026-09-10 11:36:14 +07:00
parent 0f88cb4573
commit 4d6a46258e
+42
@@ -4977,3 +4977,45 @@ of its effect.** The fix verifies afterwards with `[[ -z "${(P)n}" ]]` and only
`zsh -c` child is measuring the wrong process.
fleetd #394, #400.
## `fleet_list` no longer advertises a profile that `fleet_spawn` will refuse
**What.** The `capacity` rows in `fleet_list` now list the profiles the daemon can really spawn on —
the set it read at startup. Before, they listed the set in the **live** config, so a profile added by
a hot reload showed up with free slots and every spawn onto it failed.
```
# operator adds profile "ghost" to fleetd.yaml and the config reloads
fleet_list -> capacity: [... {profile: ghost, maxLoad: 3, live: 0, free: 3}]
fleet_spawn{profile: "ghost"} -> refused: unknown worker profile
```
**The knob.** None. This is the `capacity` reporting in `fleet_list`, and it always runs.
**Why it exists.** `profiles` is a "both ways" key. Per-profile tunables — `maxLoad`, `weight`,
`credentialId` — are **hot**, so a reload changes them at once. The profile **set** is **frozen**,
because `HerdrPeerLauncher` takes `Map.copyOf(profiles)` once when it is built and never looks
again. So "profiles is deferred" and "this live read is fine" are both true, of different halves of
the same key. The old wiring read the live `keySet()` and the hot `maxLoad()` through one lambda,
which looked consistent and was half wrong.
The direction of the error is what made it worth a ticket: it **overstated** a capability. A lead
reading `free: 3` had no way to tell that number from a real one, and only found out by spawning.
**Gotchas.**
- **`maxLoad` is still hot, and must stay hot.** The fix must freeze only the set. A test pins that:
it changes `maxLoad` from 3 to 9 by reload and asserts the **same** `CapacitySource` instance
reports the new number. Freezing both would be the mirror-image regression.
- **The right rule was already written three lines below,** for `coordinator.peers`: read from the
same snapshot the collaborator itself was opened from, rather than the live config, so the report
follows one rule instead of half hot-reloading.
- **A fresh-daemon test can never catch this.** On a daemon that has not reloaded, the live config
and the startup snapshot are the same object, so both the wrong and the right wiring pass. The
test has to perform a real reload and assert the reload applied, or it proves nothing.
- **Both directions need a test.** A single test that only adds a profile passes if the source
returns a permanently empty set. The second test — a startup profile *is* listed — is what makes
the first one load-bearing.
fleetd #416. Found by the fleet01 lead on a live daemon; the same shape as #404, where a status
field read a different source than the behaviour it described.