Sweep for the REVERSE config mismatch: a boot snapshot read where the key is hot #417

Closed
opened 2026-09-10 05:00:41 +02:00 by ltms · 2 comments
Owner

Why this sweep exists

We have now found three defects of one shape — a site reads the live config for a key that is deferred, so a status field promises something the behaviour cannot deliver:

  • #404 — exhaustionDetectionArmed read the live config while detection read the startup snapshot.
  • #416 — CapacitySource reads the live profile keySet(), so fleet_list advertises free slots for a profile the spawn gate refuses.
  • #400 (different mechanism, same direction of harm) — a receipt measured the attempt rather than the effect.

Every one was found by looking in the same direction: live read, deferred key.

Nobody has looked the other way. A site that reads the boot snapshot for a key that is genuinely hot is the mirror defect: the operator edits fleetd.yaml, the reload log correctly says nothing needs a restart, and that one site silently keeps the old value forever. There is no error and no warning, because from the config layer's point of view nothing was deferred.

This is the a-one-way-gate-is-not-a-gate lesson applied to config: a gate built after an incident closes only the direction the incident came from. We have hardened one direction three times.

What is hot, and therefore what would be a defect

From the ConfigRef class javadoc:

  • ConfigRef.java:79-81 — credentialId "is NOT [deferred] ... exactly like weight / maxLoad, so it is hot instead."
  • ConfigRef.java:86-88 — profiles is read "both ways at different sites, so the key does not fit any class above as a whole." Per-profile tunables are hot; the profile set is frozen (#416).

So: a site holding maxLoad, weight or credentialId from the boot snapshot is a candidate defect. DEFERRED_KEYS is at ConfigRef.java:227-229; COLD_KEYS and SPLIT_KEYS are just above it. Anything not in those sets is hot, and a frozen read of it is suspect.

Scope

Read-only sweep. Change nothing. Primary target is the wiring in Fleetd.java — every lambda, Supplier, Function and record constructor argument handed to a long-lived object — plus any collaborator that captures a config value in a field at construction.

For each candidate, report:

  1. file:line and the exact expression.
  2. Which key it reads, and whether that key is hot, deferred, cold or split — cite the set or the javadoc line.
  3. Is the bad state reachable? Name the path: what does an operator edit, and what then observably misbehaves. A frozen read of a hot key that nothing ever consults again is not a defect — say so and rank it low. See a-defect-on-paper-is-not-a-reachable-defect.
  4. Which direction the failure points: does it overstate a capability (dangerous) or understate it (annoying)?
  5. Your confidence, and what you did not check.

Rank by reachability multiplied by harm, worst first. A ranked list of real ones beats a long list of pattern matches.

Explicitly out of scope

  • Do not fix anything, and do not open a PR that changes production code.
  • coordinator and broker wiring is correct as-is: Fleetd.java:687-690 documents that coordinator.peers deliberately reads the boot snapshot because the mailbox was opened from it. Confirm the reasoning holds, but this is not a finding.
  • #416 is already filed and being fixed; the live keySet() is not a new finding.

Report back

A ranked list of findings in the hunter format. If the honest answer is "no reachable instance of the reverse defect exists", that is a real and useful result — say it plainly and show what you searched, including the searches that returned nothing.

## Why this sweep exists We have now found three defects of one shape — a site reads the **live** config for a key that is **deferred**, so a status field promises something the behaviour cannot deliver: - **#404** — `exhaustionDetectionArmed` read the live config while detection read the startup snapshot. - **#416** — `CapacitySource` reads the live profile `keySet()`, so `fleet_list` advertises free slots for a profile the spawn gate refuses. - **#400** (different mechanism, same direction of harm) — a receipt measured the attempt rather than the effect. Every one was found by looking in the same direction: *live read, deferred key.* **Nobody has looked the other way.** A site that reads the **boot snapshot** for a key that is genuinely **hot** is the mirror defect: the operator edits `fleetd.yaml`, the reload log correctly says nothing needs a restart, and that one site silently keeps the old value forever. There is no error and no warning, because from the config layer's point of view nothing was deferred. This is the [[a-one-way-gate-is-not-a-gate]] lesson applied to config: a gate built after an incident closes only the direction the incident came from. We have hardened one direction three times. ## What is hot, and therefore what would be a defect From the `ConfigRef` class javadoc: - `ConfigRef.java:79-81` — `credentialId` "is NOT [deferred] ... exactly like `weight` / `maxLoad`, so it is hot instead." - `ConfigRef.java:86-88` — `profiles` is read "**both** ways at different sites, so the key does not fit any class above as a whole." Per-profile tunables are hot; the profile **set** is frozen (#416). So: a site holding `maxLoad`, `weight` or `credentialId` from the boot snapshot is a candidate defect. `DEFERRED_KEYS` is at `ConfigRef.java:227-229`; `COLD_KEYS` and `SPLIT_KEYS` are just above it. Anything **not** in those sets is hot, and a frozen read of it is suspect. ## Scope Read-only sweep. **Change nothing.** Primary target is the wiring in `Fleetd.java` — every lambda, `Supplier`, `Function` and record constructor argument handed to a long-lived object — plus any collaborator that captures a config value in a field at construction. For each candidate, report: 1. `file:line` and the exact expression. 2. Which key it reads, and whether that key is hot, deferred, cold or split — cite the set or the javadoc line. 3. **Is the bad state reachable?** Name the path: what does an operator edit, and what then observably misbehaves. A frozen read of a hot key that nothing ever consults again is not a defect — say so and rank it low. See [[a-defect-on-paper-is-not-a-reachable-defect]]. 4. Which direction the failure points: does it overstate a capability (dangerous) or understate it (annoying)? 5. Your confidence, and what you did not check. Rank by reachability multiplied by harm, worst first. A ranked list of real ones beats a long list of pattern matches. ## Explicitly out of scope - Do not fix anything, and do not open a PR that changes production code. - `coordinator` and `broker` wiring is **correct as-is**: `Fleetd.java:687-690` documents that `coordinator.peers` deliberately reads the boot snapshot because the mailbox was opened from it. Confirm the reasoning holds, but this is not a finding. - #416 is already filed and being fixed; the live `keySet()` is not a new finding. ## Report back A ranked list of findings in the `hunter` format. If the honest answer is "no reachable instance of the reverse defect exists", that is a real and useful result — say it plainly and show what you searched, including the searches that returned nothing.
Author
Owner

Sweep done. The reverse defect exists, twice. Both filed:

  • #424 — MemberRegistry freezes fleet.architects, so revoking an architect slot does not revoke it. This is the dangerous direction: it overstates a capability, and an architect is a materially more privileged role than a worker.
  • #425 — cfg.effectiveDefaultProfile() is frozen at boot and reported by fleet_profiles as "default", while placement already reads the live pool. Plus a stale profile name used for worktree overlay provisioning.

Verified here before filing, rather than promoted from the report:

  • Fleetd.java:121 — cfg is FleetConfig.load(configPath), the boot snapshot.
  • Fleetd.java:355 — new MemberRegistry(cfg.fleet()), and grep -c 'config\.get()\|Supplier' MemberRegistry.java returns 0. No supplier, no re-read.
  • ConfigRef.java:29-37 — the javadoc claims every role pool including architects "is read the same live way, through the same supplier".
  • MemberRegistry.java:258-295 — requireSlotFor and reserve both match against the frozen slotsFor(ARCHITECT) map on the spawn path.
  • FleetMcp.java:1150 / CompositePeerLauncher.java:76,511-515 — the frozen field is what fleet_profiles reports; poolFor(...).getFirst() is what placement uses.

One thing the sweep found that the brief did not ask for, and it is the more useful half. The reload does not merely omit this — at ConfigRef.java:543-551 it says, verbatim:

the rest of fleet: (architects, developers, reviewers, charters, tabLabel) is read live through the supplier on CompositePeerLauncher and already applied

That sentence is true of the placement consumer and false of the identity consumer. So fleet.architects is a split key inside the already-split fleet: key, and the daemon asserts the wrong half out loud. #404's lesson one level up: a claim that is correct about one consumer, applied to a key with two.

The structural gap, which outlives both tickets. ConfigRefTopLevelCoverageTest, ConfigRefTopLevelReportingCoverageTest and ConfigRefProfileCoverageTest all check ConfigRef's key sets against FleetConfig's record shape. None of them looks at a consumer. A key's class is a claim about every site that reads it, and nothing tests it that way — which is exactly why three tickets in this direction and two in the reverse all had to be found by a person reading code. Worth its own ticket; not filing it yet, because I want to think about what such a checker could actually assert without becoming a second hand-maintained list.

Correctly ruled not findings, and I agree with both:

  • coordinator.peers at Fleetd.java:684-688 reads the boot snapshot deliberately, matching the mailbox opened from the same snapshot at :531. No half-hot-reload gap.
  • HerdrPeerLauncher's own per-adapter defaultProfile field is the same frozen shape but is dead in production — the composite always routes spawn() with an explicit, live-resolved name. Noted in #425 as out of scope rather than counted.

Closing: the sweep produced what it was for. The two fixes are #424 and #425.

Sweep done. **The reverse defect exists, twice.** Both filed: - **#424** — `MemberRegistry` freezes `fleet.architects`, so revoking an architect slot does not revoke it. This is the dangerous direction: it overstates a capability, and an architect is a materially more privileged role than a worker. - **#425** — `cfg.effectiveDefaultProfile()` is frozen at boot and reported by `fleet_profiles` as `"default"`, while placement already reads the live pool. Plus a stale profile name used for worktree overlay provisioning. Verified here before filing, rather than promoted from the report: - `Fleetd.java:121` — `cfg` is `FleetConfig.load(configPath)`, the boot snapshot. - `Fleetd.java:355` — `new MemberRegistry(cfg.fleet())`, and `grep -c 'config\.get()\|Supplier' MemberRegistry.java` returns **0**. No supplier, no re-read. - `ConfigRef.java:29-37` — the javadoc claims every role pool including `architects` "is read the same live way, through the same supplier". - `MemberRegistry.java:258-295` — `requireSlotFor` and `reserve` both match against the frozen `slotsFor(ARCHITECT)` map on the spawn path. - `FleetMcp.java:1150` / `CompositePeerLauncher.java:76,511-515` — the frozen field is what `fleet_profiles` reports; `poolFor(...).getFirst()` is what placement uses. **One thing the sweep found that the brief did not ask for, and it is the more useful half.** The reload does not merely omit this — at `ConfigRef.java:543-551` it says, verbatim: > the rest of fleet: (architects, developers, reviewers, charters, tabLabel) is read live through the supplier on CompositePeerLauncher and **already applied** That sentence is true of the placement consumer and false of the identity consumer. So `fleet.architects` is a split key **inside** the already-split `fleet:` key, and the daemon asserts the wrong half out loud. #404's lesson one level up: a claim that is correct about one consumer, applied to a key with two. **The structural gap, which outlives both tickets.** `ConfigRefTopLevelCoverageTest`, `ConfigRefTopLevelReportingCoverageTest` and `ConfigRefProfileCoverageTest` all check `ConfigRef`'s key sets against `FleetConfig`'s **record shape**. None of them looks at a **consumer**. A key's class is a claim about every site that reads it, and nothing tests it that way — which is exactly why three tickets in this direction and two in the reverse all had to be found by a person reading code. Worth its own ticket; not filing it yet, because I want to think about what such a checker could actually assert without becoming a second hand-maintained list. Correctly ruled **not** findings, and I agree with both: - `coordinator.peers` at `Fleetd.java:684-688` reads the boot snapshot deliberately, matching the mailbox opened from the same snapshot at `:531`. No half-hot-reload gap. - `HerdrPeerLauncher`'s own per-adapter `defaultProfile` field is the same frozen shape but is dead in production — the composite always routes `spawn()` with an explicit, live-resolved name. Noted in #425 as out of scope rather than counted. Closing: the sweep produced what it was for. The two fixes are #424 and #425.
ltms closed this issue 2026-09-10 06:46:48 +02:00
Author
Owner

Filed the structural gap as #427 — I said above I wanted to think about what such a checker could actually assert first, so here is the answer.

The concrete thing I did not see when I wrote that comment: the frozen half is invisible to search. Every hunt for this class of bug, mine included, runs some form of grep 'config\.get()' — which by construction can only find the live half. The frozen half is a constructor argument or a field and matches no idiom anyone would think to grep for. That is why three live-read instances got fixed while #424 and #425 sat unfound in the same file.

So the goal in #427 is narrower and more achievable than "check every consumer against the classification": make a frozen read explicit, named and greppable, so it is a decision somebody wrote down rather than the default that happens when you pass a value.

Two candidates there, not exclusive — an ArchUnit rule forbidding a config value held in a field (which inverts the default, and depends on #131 landing ArchUnit), and the cheap half, renaming Fleetd.main's cfg local so one grep enumerates every frozen read in the file. #427 explicitly accepts "this cannot be checked mechanically without a hand-maintained list" as a real result, because a checker that needs a human to keep it truthful has moved the problem rather than solved it.

Filed the structural gap as **#427** — I said above I wanted to think about what such a checker could actually assert first, so here is the answer. The concrete thing I did not see when I wrote that comment: **the frozen half is invisible to search.** Every hunt for this class of bug, mine included, runs some form of `grep 'config\.get()'` — which by construction can only find the *live* half. The frozen half is a constructor argument or a field and matches no idiom anyone would think to grep for. That is why three live-read instances got fixed while #424 and #425 sat unfound in the same file. So the goal in #427 is narrower and more achievable than "check every consumer against the classification": **make a frozen read explicit, named and greppable, so it is a decision somebody wrote down rather than the default that happens when you pass a value.** Two candidates there, not exclusive — an ArchUnit rule forbidding a config value held in a field (which inverts the default, and depends on #131 landing ArchUnit), and the cheap half, renaming `Fleetd.main`'s `cfg` local so one grep enumerates every frozen read in the file. #427 explicitly accepts "this cannot be checked mechanically without a hand-maintained list" as a real result, because a checker that needs a human to keep it truthful has moved the problem rather than solved it.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#417