Features: fleet_list enumerates the startup profile set (#416)
+42
@@ -4977,3 +4977,45 @@ of its effect.** The fix verifies afterwards with `[[ -z "${(P)n}" ]]` and only
|
||||
`zsh -c` child is measuring the wrong process.
|
||||
|
||||
fleetd #394, #400.
|
||||
|
||||
## `fleet_list` no longer advertises a profile that `fleet_spawn` will refuse
|
||||
|
||||
**What.** The `capacity` rows in `fleet_list` now list the profiles the daemon can really spawn on —
|
||||
the set it read at startup. Before, they listed the set in the **live** config, so a profile added by
|
||||
a hot reload showed up with free slots and every spawn onto it failed.
|
||||
|
||||
```
|
||||
# operator adds profile "ghost" to fleetd.yaml and the config reloads
|
||||
fleet_list -> capacity: [... {profile: ghost, maxLoad: 3, live: 0, free: 3}]
|
||||
fleet_spawn{profile: "ghost"} -> refused: unknown worker profile
|
||||
```
|
||||
|
||||
**The knob.** None. This is the `capacity` reporting in `fleet_list`, and it always runs.
|
||||
|
||||
**Why it exists.** `profiles` is a "both ways" key. Per-profile tunables — `maxLoad`, `weight`,
|
||||
`credentialId` — are **hot**, so a reload changes them at once. The profile **set** is **frozen**,
|
||||
because `HerdrPeerLauncher` takes `Map.copyOf(profiles)` once when it is built and never looks
|
||||
again. So "profiles is deferred" and "this live read is fine" are both true, of different halves of
|
||||
the same key. The old wiring read the live `keySet()` and the hot `maxLoad()` through one lambda,
|
||||
which looked consistent and was half wrong.
|
||||
|
||||
The direction of the error is what made it worth a ticket: it **overstated** a capability. A lead
|
||||
reading `free: 3` had no way to tell that number from a real one, and only found out by spawning.
|
||||
|
||||
**Gotchas.**
|
||||
|
||||
- **`maxLoad` is still hot, and must stay hot.** The fix must freeze only the set. A test pins that:
|
||||
it changes `maxLoad` from 3 to 9 by reload and asserts the **same** `CapacitySource` instance
|
||||
reports the new number. Freezing both would be the mirror-image regression.
|
||||
- **The right rule was already written three lines below,** for `coordinator.peers`: read from the
|
||||
same snapshot the collaborator itself was opened from, rather than the live config, so the report
|
||||
follows one rule instead of half hot-reloading.
|
||||
- **A fresh-daemon test can never catch this.** On a daemon that has not reloaded, the live config
|
||||
and the startup snapshot are the same object, so both the wrong and the right wiring pass. The
|
||||
test has to perform a real reload and assert the reload applied, or it proves nothing.
|
||||
- **Both directions need a test.** A single test that only adds a profile passes if the source
|
||||
returns a permanently empty set. The second test — a startup profile *is* listed — is what makes
|
||||
the first one load-bearing.
|
||||
|
||||
fleetd #416. Found by the fleet01 lead on a live daemon; the same shape as #404, where a status
|
||||
field read a different source than the behaviour it described.
|
||||
|
||||
Reference in New Issue
Block a user