diff --git a/11-Features.md b/11-Features.md index c672b6b..8973144 100644 --- a/11-Features.md +++ b/11-Features.md @@ -4977,3 +4977,45 @@ of its effect.** The fix verifies afterwards with `[[ -z "${(P)n}" ]]` and only `zsh -c` child is measuring the wrong process. fleetd #394, #400. + +## `fleet_list` no longer advertises a profile that `fleet_spawn` will refuse + +**What.** The `capacity` rows in `fleet_list` now list the profiles the daemon can really spawn on — +the set it read at startup. Before, they listed the set in the **live** config, so a profile added by +a hot reload showed up with free slots and every spawn onto it failed. + +``` +# operator adds profile "ghost" to fleetd.yaml and the config reloads +fleet_list -> capacity: [... {profile: ghost, maxLoad: 3, live: 0, free: 3}] +fleet_spawn{profile: "ghost"} -> refused: unknown worker profile +``` + +**The knob.** None. This is the `capacity` reporting in `fleet_list`, and it always runs. + +**Why it exists.** `profiles` is a "both ways" key. Per-profile tunables — `maxLoad`, `weight`, +`credentialId` — are **hot**, so a reload changes them at once. The profile **set** is **frozen**, +because `HerdrPeerLauncher` takes `Map.copyOf(profiles)` once when it is built and never looks +again. So "profiles is deferred" and "this live read is fine" are both true, of different halves of +the same key. The old wiring read the live `keySet()` and the hot `maxLoad()` through one lambda, +which looked consistent and was half wrong. + +The direction of the error is what made it worth a ticket: it **overstated** a capability. A lead +reading `free: 3` had no way to tell that number from a real one, and only found out by spawning. + +**Gotchas.** + +- **`maxLoad` is still hot, and must stay hot.** The fix must freeze only the set. A test pins that: + it changes `maxLoad` from 3 to 9 by reload and asserts the **same** `CapacitySource` instance + reports the new number. Freezing both would be the mirror-image regression. +- **The right rule was already written three lines below,** for `coordinator.peers`: read from the + same snapshot the collaborator itself was opened from, rather than the live config, so the report + follows one rule instead of half hot-reloading. +- **A fresh-daemon test can never catch this.** On a daemon that has not reloaded, the live config + and the startup snapshot are the same object, so both the wrong and the right wiring pass. The + test has to perform a real reload and assert the reload applied, or it proves nothing. +- **Both directions need a test.** A single test that only adds a profile passes if the source + returns a permanently empty set. The second test — a startup profile *is* listed — is what makes + the first one load-bearing. + +fleetd #416. Found by the fleet01 lead on a live daemon; the same shape as #404, where a status +field read a different source than the behaviour it described.