LeadContextGauge's HIGH threshold is a fixed 200k, but autoCompactWindow is per-profile and may be as low as 100k — the warning can never fire #637

Closed
opened 2026-10-01 18:18:06 +02:00 by ltms · 3 comments
Owner

What

LeadContextGauge.HIGH_THRESHOLD_TOKENS is a hardcoded 200_000 (LeadContextGauge.java:106). The thing it warns about — auto-compaction — is configured per profile by autoCompactWindow, and the config validator accepts any value from 100_000 up (FleetConfig.java:2184).

So the threshold is an absolute number, while the event it predicts is a configurable one. When the window is at or below the threshold, the HIGH state cannot fire before a compaction. The warning is not merely late; it is unreachable.

Measured on 158a2a8

LeadContextGauge.java:106   static final long HIGH_THRESHOLD_TOKENS = 200_000;
LeadContextGauge.java:302   State state = tokens >= HIGH_THRESHOLD_TOKENS ? State.HIGH : State.OK;
FleetConfig.java:2184       static final int AUTO_COMPACT_WINDOW_MIN = 100_000;
FleetConfig.java:2186       static final int AUTO_COMPACT_WINDOW_MAX = 1_000_000;

Any profile configured between 100000 and 200000 is accepted by the validator and compacts before tokens can reach 200_000.

Not currently reached — this is latent, not live

Every profile in the live fleetd/fleetd.yaml sets autoCompactWindow: 250000 (8 profiles, lines 25, 74, 100, 110, 124, 172, 194, 228). All are above the threshold, so nothing is broken right now. The defect is reachable by a config edit the validator allows, not by today's config.

For the four Claude Code profiles the yaml key is also inert, because CLAUDE_CODE_AUTO_COMPACT_WINDOW: "300000" in each env: block wins over it (measured, #618). Their effective window is 300000, which gives the gauge about 68k of headroom. That headroom is real but accidental.

Why the current justification does not cover it

The javadoc at LeadContextGauge.java:99-105 already knows the measured figure:

On the host this was measured on, auto-compaction actually fires around 267,000–270,000 tokens, but the point of a HIGH state is to warn before that happens, not at it — 200,000 is the standard Claude context window size and a sensible built-in default.

Two things are true here and only one is a problem.

"Warn before, not at" is right, and wanting no config key for it is right. But the number is anchored to the model's context window, which is not the quantity that ends a session. The quantity that ends a session is the auto-compact window. Those two happened to be close on the host that was measured, so the choice worked. They are not the same thing, and one of them is configurable.

This is the same mistake recorded in the opus lead's handover of 2026-10-01: a lead read "23% of a 1M window" and concluded no handover was needed, having used the model's context window as the denominator when the auto-compact window was the one that mattered. Same number, wrong denominator, inverted conclusion. The gauge encodes that same substitution as a constant.

Suggested fix

Derive the threshold from the window it is warning about, keeping a built-in default so no config key is needed:

  • Resolve the lead profile's effective window. For a Claude Code profile that must include CLAUDE_CODE_AUTO_COMPACT_WINDOW from the env: block, because it beats the yaml key — reading only autoCompactWindow would compute from an inert number, which is how this defect would survive its own fix.
  • Set HIGH at a fraction of that window, so the warn margin scales. The current pairing (200k of 300k) is two thirds.
  • Fall back to today's 200_000 when the window cannot be resolved, so behaviour is unchanged when there is nothing better to use.

Acceptance criteria

  1. A profile with an effective window of 100000 reports HIGH strictly below 100000, at some margin. Today it never reports HIGH before compacting. This assertion must be red before the fix.
  2. A profile whose yaml autoCompactWindow and CLAUDE_CODE_AUTO_COMPACT_WINDOW disagree computes from the env var. Build the fixture so the two values give different thresholds, or the test passes whichever one the code reads and proves nothing.
  3. With no window resolvable, the threshold is still 200_000, so existing behaviour is preserved.
  4. A test that pins the OK/HIGH boundary must assert on both sides of it. An assertion that only checks "HIGH at some high number" passes a threshold of zero.

Found how

Noticed while reading LeadContextGauge during the lead handover of 2026-10-01; the gauge had just fired on this lead at ~242k, usefully but for the wrong reason. Written up as a finding in that handover and filed here rather than acted on, since it needs a decision about where the threshold comes from.

## What `LeadContextGauge.HIGH_THRESHOLD_TOKENS` is a hardcoded `200_000` (`LeadContextGauge.java:106`). The thing it warns about — auto-compaction — is configured per profile by `autoCompactWindow`, and the config validator accepts any value from `100_000` up (`FleetConfig.java:2184`). So the threshold is an absolute number, while the event it predicts is a configurable one. When the window is at or below the threshold, the HIGH state cannot fire before a compaction. The warning is not merely late; it is unreachable. ## Measured on `158a2a8` ``` LeadContextGauge.java:106 static final long HIGH_THRESHOLD_TOKENS = 200_000; LeadContextGauge.java:302 State state = tokens >= HIGH_THRESHOLD_TOKENS ? State.HIGH : State.OK; FleetConfig.java:2184 static final int AUTO_COMPACT_WINDOW_MIN = 100_000; FleetConfig.java:2186 static final int AUTO_COMPACT_WINDOW_MAX = 1_000_000; ``` Any profile configured between `100000` and `200000` is accepted by the validator and compacts before `tokens` can reach `200_000`. ## Not currently reached — this is latent, not live Every profile in the live `fleetd/fleetd.yaml` sets `autoCompactWindow: 250000` (8 profiles, lines 25, 74, 100, 110, 124, 172, 194, 228). All are above the threshold, so nothing is broken right now. The defect is reachable by a config edit the validator allows, not by today's config. For the four Claude Code profiles the yaml key is also inert, because `CLAUDE_CODE_AUTO_COMPACT_WINDOW: "300000"` in each `env:` block wins over it (measured, #618). Their effective window is 300000, which gives the gauge about 68k of headroom. That headroom is real but accidental. ## Why the current justification does not cover it The javadoc at `LeadContextGauge.java:99-105` already knows the measured figure: > On the host this was measured on, auto-compaction actually fires around 267,000–270,000 tokens, but the point of a HIGH state is to warn before that happens, not at it — 200,000 is the standard Claude context window size and a sensible built-in default. Two things are true here and only one is a problem. "Warn before, not at" is right, and wanting no config key for it is right. But the number is anchored to the **model's context window**, which is not the quantity that ends a session. The quantity that ends a session is the **auto-compact window**. Those two happened to be close on the host that was measured, so the choice worked. They are not the same thing, and one of them is configurable. This is the same mistake recorded in the `opus` lead's handover of 2026-10-01: a lead read "23% of a 1M window" and concluded no handover was needed, having used the model's context window as the denominator when the auto-compact window was the one that mattered. Same number, wrong denominator, inverted conclusion. The gauge encodes that same substitution as a constant. ## Suggested fix Derive the threshold from the window it is warning about, keeping a built-in default so no config key is needed: - Resolve the lead profile's effective window. For a Claude Code profile that must include `CLAUDE_CODE_AUTO_COMPACT_WINDOW` from the `env:` block, because it beats the yaml key — reading only `autoCompactWindow` would compute from an inert number, which is how this defect would survive its own fix. - Set HIGH at a fraction of that window, so the warn margin scales. The current pairing (200k of 300k) is two thirds. - Fall back to today's `200_000` when the window cannot be resolved, so behaviour is unchanged when there is nothing better to use. ## Acceptance criteria 1. A profile with an effective window of `100000` reports HIGH strictly below `100000`, at some margin. Today it never reports HIGH before compacting. This assertion must be red before the fix. 2. A profile whose yaml `autoCompactWindow` and `CLAUDE_CODE_AUTO_COMPACT_WINDOW` disagree computes from the **env var**. Build the fixture so the two values give different thresholds, or the test passes whichever one the code reads and proves nothing. 3. With no window resolvable, the threshold is still `200_000`, so existing behaviour is preserved. 4. A test that pins the OK/HIGH boundary must assert on both sides of it. An assertion that only checks "HIGH at some high number" passes a threshold of zero. ## Found how Noticed while reading `LeadContextGauge` during the lead handover of 2026-10-01; the gauge had just fired on this lead at ~242k, usefully but for the wrong reason. Written up as a finding in that handover and filed here rather than acted on, since it needs a decision about where the threshold comes from.
Author
Owner

Re-read on 2026-10-01 after #618 item 3 raised four profiles to 300000. Still latent, and the raise moved the margin the safe way. No fix is owed yet.

What I measured, at main = 141ae3b:

  • LeadContextGauge.java:106 — static final long HIGH_THRESHOLD_TOKENS = 200_000; (unchanged), used at LeadContextGauge.java:302: tokens >= HIGH_THRESHOLD_TOKENS ? State.HIGH : State.OK.
  • FleetConfig.java:2184 — AUTO_COMPACT_WINDOW_MIN = 100_000, and FleetConfig.java:2186 — AUTO_COMPACT_WINDOW_MAX = 1_000_000. So validation still accepts a window below the fixed gauge threshold.
  • The live fleetd/fleetd.yaml, each autoCompactWindow mapped to the block it sits in:
    profile kind autoCompactWindow line
    local claude-code 300000 25
    local-direct claude-code 300000 74
    gx opencode 250000 100
    opus claude-code 300000 110
    sonnet claude-code 300000 124
    sol opencode 250000 172
    terra opencode 250000 194
    xf opencode 250000 228

So the danger band is [100000, 200000) and nothing is in it. Every profile is at 250000 or above, so HIGH fires before compaction in all eight cases.

For the two lead-capable profiles specifically — opus and sonnet, each leadSeats: 1 per fleet_list — the headroom between the HIGH warning and compaction grew from 50,000 to 100,000 tokens. Raising the window cannot bring a profile into the danger band; only lowering one can.

What is still a real defect: the gauge threshold is a hardcoded constant while the window it must stay below is per-profile and validated down to 100,000. A future profile set anywhere in [100000, 200000) passes validation and silently loses the HIGH warning entirely. That is the ticket, and it is unchanged.

One thing I did not check: whether the gauge reads the per-profile window at all, or only ever compares against the constant. I read the threshold and its one comparison site, not the whole gauge. Anyone fixing this should start there.

Re-read on 2026-10-01 after #618 item 3 raised four profiles to `300000`. **Still latent, and the raise moved the margin the safe way.** No fix is owed yet. What I measured, at `main` = `141ae3b`: - `LeadContextGauge.java:106` — `static final long HIGH_THRESHOLD_TOKENS = 200_000;` (unchanged), used at `LeadContextGauge.java:302`: `tokens >= HIGH_THRESHOLD_TOKENS ? State.HIGH : State.OK`. - `FleetConfig.java:2184` — `AUTO_COMPACT_WINDOW_MIN = 100_000`, and `FleetConfig.java:2186` — `AUTO_COMPACT_WINDOW_MAX = 1_000_000`. So validation still accepts a window below the fixed gauge threshold. - The live `fleetd/fleetd.yaml`, each `autoCompactWindow` mapped to the block it sits in: | profile | kind | autoCompactWindow | line | |---|---|---|---| | `local` | claude-code | 300000 | 25 | | `local-direct` | claude-code | 300000 | 74 | | `gx` | opencode | 250000 | 100 | | `opus` | claude-code | 300000 | 110 | | `sonnet` | claude-code | 300000 | 124 | | `sol` | opencode | 250000 | 172 | | `terra` | opencode | 250000 | 194 | | `xf` | opencode | 250000 | 228 | **So the danger band is `[100000, 200000)` and nothing is in it.** Every profile is at 250000 or above, so HIGH fires before compaction in all eight cases. For the two lead-capable profiles specifically — `opus` and `sonnet`, each `leadSeats: 1` per `fleet_list` — the headroom between the HIGH warning and compaction **grew from 50,000 to 100,000 tokens**. Raising the window cannot bring a profile into the danger band; only lowering one can. **What is still a real defect:** the gauge threshold is a hardcoded constant while the window it must stay below is per-profile and validated down to 100,000. A future profile set anywhere in `[100000, 200000)` passes validation and silently loses the HIGH warning entirely. That is the ticket, and it is unchanged. **One thing I did not check:** whether the gauge reads the per-profile window at all, or only ever compares against the constant. I read the threshold and its one comparison site, not the whole gauge. Anyone fixing this should start there.
Author
Owner

Lead adjudication of PR 657 — one revision round before merge

PR 657 is good work and I am not asking for a redesign. Wiring the two real consumers instead of leaving the fix latent was the right call. Three reviewers looked at it on separate dimensions, and I verified every finding myself. Two things must change before merge; one is recorded and deliberately not fixed.

Merge build, measured by me, on PR 657 merged into main at 3fab743: mvn -o clean install → Tests run: 1918, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS, exit 0. 1918 = 1904 + 14, so nothing was lost in the merge.


Finding 1 — MUST FIX. fleet_list's half of this fix is completely unguarded.

Fleetd.java:1079. Replace leadContextWindowLookup(profiles, leaders) with _ -> null in leadConfigDirSource and the whole suite stays green:

mvn -o install   →  Tests run: 1918, Failures: 0, Errors: 0, Skipped: 0
                    BUILD SUCCESS, exit 0

A reviewer found this against one test class. I re-ran it against all 1918 tests, which is the number that matters: fleet_list could report every lead's context against the fixed 200,000 fallback forever, and nothing notices.

Paired kill, so this is not "the tests never ran". Mutating HIGH_THRESHOLD_FRACTION to 1.0 kills 2 of 3 tests in LeadContextGaugeHighThresholdTest:

expected: <HIGH> but was: <OK>   (effectiveWindowOf100000ReportsHighStrictlyBelow100000:46)
expected: <HIGH> but was: <OK>   (boundarySitsStrictlyBetweenOkAndHighOnBothSides:65)

So the gauge's arithmetic is pinned and the tests do run. What is unpinned is the wiring that feeds it.

This is the exact defect class the javadoc 15 lines above the mutated line already describes (fleetd #602, PR #606 comment 17353): a source hand-built inline, with nothing a test could call. That javadoc records that LeadConfigDirSource.none() at the call site once compiled with 0 errors against a green 1822-test suite. leadConfigDirSource was extracted specifically so a test could call it — and this PR added a second argument to it without pinning that argument. FleetdLeadConfigDirSourceWiringTest is the precedent for the shape of the fix.

FleetdLeadContextWindowLookupTest tests the detached lookup, which is necessary but not sufficient — a test that supplies its own dependency says nothing about the producer.

Finding 2 — MUST FIX. The gauge's result cache ignores the threshold.

LeadContextGauge.java:195. cacheKey = base + '\u0000' + sessionId, TTL 5s, but Reading.state() now depends on the caller's effectiveWindowTokens. I reproduced it:

PROBE first(window=100000)=HIGH second(window=1000000)=HIGH
                                                       ^ must be OK
AssertionFailedError: a 1,000,000 window must report OK at 90,000 tokens
                      regardless of the previous call ==> expected: <OK> but was: <HIGH>

With the control that makes it sound: a fresh gauge returns HIGH for a 100,000 window and OK for a 1,000,000 window at the same 90,000 tokens. So the threshold logic is correct and the instrument can tell the two apart — the stale answer comes from the cache, not from broken arithmetic.

The three new tests cannot see this: each builds new LeadContextGauge() with its own @TempDir, so no two calls ever share a cache entry.

On reachability, I want the record straight, because a reviewer and I disagreed and we were each half right. A reviewer argued it is unreachable for two reasons: the two call sites use separate gauge instances (FleetdAssembly.java:394 and FleetMcp.java:149 — I verified this, it is true), and autoCompactWindow reload is deferred so a lead's window never changes while the daemon runs.

The second reason has the right conclusion and the wrong evidence. They cited FleetConfig.java:487, which documents exhaustedPattern, not autoCompactWindow. The real evidence is ConfigRef.java:114 (autoCompactWindow listed among the deferred keys, fleetd #323 instance 1) and ConfigRef.java:719-723.

But deferred is defined at ConfigRef.java:81 as "accepted into the new snapshot, but the wiring built at startup is not rebuilt" — and leadContextWindowLookup reads config.get().profiles(), the live snapshot. So after a reload the resolved window does change, within one gauge instance, and two fleet_list calls 2 seconds apart would straddle it. Narrow, but not blocked.

Fold the threshold into the cache key. The cost is one lost cache hit when a session's resolved window actually changes, which is already a cold path.

Finding 3 — recorded, do NOT fix in this PR.

FleetMcp.java:301 and the two narrow Fleetd overloads are dead production surface. My counts, not a reviewer's: the only production construction is Fleetd.java:1078 using the two-argument form; FleetdAssembly.java:407 resolves to the wide leadContextSource; and narrow leadContextSource:1165 has no caller at all, with the only path into narrow leadContextLookup:1127 being that dead wrapper's own body at :1168.

A reviewer reported 7 old-form test call sites in 3 files. My pattern also catches the unqualified new LeadConfigDirSource( and measures 8 across 4 files — they missed FleetMcpLeadContextGaugeWiringTest.java. Use my number if anyone acts on this.

Leaving it is a judgement call, not an oversight. Removing dead surface is right, but it means editing 8 test call sites in the same change as a correctness fix, and I would rather the correctness fix land clean. Filed separately rather than bundled.

One thing worth saying about it: every hazard in this PR lives in a back-compat form. The narrow read(configDir, sessionId, agentType) overload passes null, which both silently restores the old fixed threshold and is what makes finding 2 reachable. That is a stronger reason to remove the narrow forms than "dead code" was on its own.


Also noted, not a blocker

The window is resolved from the live snapshot, while the launched Claude session is still running on the --autocompact flag it was spawned with. After a deferred reload those disagree, so the gauge would scale against a window the session is not actually on until it is respawned. leadConfigDirLookup already has exactly this property, so this PR inherits the behaviour rather than introducing it. Out of scope here; say so if you think it deserves its own ticket.

What I did not check myself

I did not re-run the author's RED-before-fix measurement for the 4-arg read overload. The report describes a compile error as the RED, which is a weaker form than a failing assertion — a signature that does not exist yet cannot distinguish "the behaviour is wrong" from "the behaviour has no code path". I am accepting it because findings 1 and 2 are now pinned by tests that do assert behaviour.

## Lead adjudication of PR 657 — one revision round before merge PR 657 is good work and I am not asking for a redesign. Wiring the two real consumers instead of leaving the fix latent was the right call. Three reviewers looked at it on separate dimensions, and I verified every finding myself. **Two things must change before merge; one is recorded and deliberately not fixed.** Merge build, measured by me, on PR 657 merged into `main` at `3fab743`: `mvn -o clean install` → `Tests run: 1918, Failures: 0, Errors: 0, Skipped: 0`, `BUILD SUCCESS`, exit 0. 1918 = 1904 + 14, so nothing was lost in the merge. --- ### Finding 1 — MUST FIX. `fleet_list`'s half of this fix is completely unguarded. `Fleetd.java:1079`. Replace `leadContextWindowLookup(profiles, leaders)` with `_ -> null` in `leadConfigDirSource` and the **whole suite stays green**: ``` mvn -o install → Tests run: 1918, Failures: 0, Errors: 0, Skipped: 0 BUILD SUCCESS, exit 0 ``` A reviewer found this against one test class. I re-ran it against all 1918 tests, which is the number that matters: `fleet_list` could report every lead's context against the fixed 200,000 fallback forever, and nothing notices. **Paired kill, so this is not "the tests never ran".** Mutating `HIGH_THRESHOLD_FRACTION` to `1.0` kills 2 of 3 tests in `LeadContextGaugeHighThresholdTest`: ``` expected: <HIGH> but was: <OK> (effectiveWindowOf100000ReportsHighStrictlyBelow100000:46) expected: <HIGH> but was: <OK> (boundarySitsStrictlyBetweenOkAndHighOnBothSides:65) ``` So the gauge's arithmetic **is** pinned and the tests do run. What is unpinned is the wiring that feeds it. This is the exact defect class the javadoc 15 lines above the mutated line already describes (fleetd #602, PR #606 comment 17353): a source hand-built inline, with nothing a test could call. That javadoc records that `LeadConfigDirSource.none()` at the call site once compiled with 0 errors against a green 1822-test suite. `leadConfigDirSource` was extracted *specifically* so a test could call it — and this PR added a second argument to it without pinning that argument. `FleetdLeadConfigDirSourceWiringTest` is the precedent for the shape of the fix. `FleetdLeadContextWindowLookupTest` tests the detached lookup, which is necessary but not sufficient — a test that supplies its own dependency says nothing about the producer. ### Finding 2 — MUST FIX. The gauge's result cache ignores the threshold. `LeadContextGauge.java:195`. `cacheKey = base + '\u0000' + sessionId`, TTL 5s, but `Reading.state()` now depends on the caller's `effectiveWindowTokens`. I reproduced it: ``` PROBE first(window=100000)=HIGH second(window=1000000)=HIGH ^ must be OK AssertionFailedError: a 1,000,000 window must report OK at 90,000 tokens regardless of the previous call ==> expected: <OK> but was: <HIGH> ``` **With the control that makes it sound:** a *fresh* gauge returns `HIGH` for a 100,000 window and `OK` for a 1,000,000 window at the same 90,000 tokens. So the threshold logic is correct and the instrument can tell the two apart — the stale answer comes from the cache, not from broken arithmetic. The three new tests cannot see this: each builds `new LeadContextGauge()` with its own `@TempDir`, so no two calls ever share a cache entry. **On reachability, I want the record straight, because a reviewer and I disagreed and we were each half right.** A reviewer argued it is unreachable for two reasons: the two call sites use separate gauge instances (`FleetdAssembly.java:394` and `FleetMcp.java:149` — I verified this, it is true), and `autoCompactWindow` reload is deferred so a lead's window never changes while the daemon runs. The second reason has the right conclusion and the wrong evidence. They cited `FleetConfig.java:487`, which documents `exhaustedPattern`, not `autoCompactWindow`. The real evidence is `ConfigRef.java:114` (`autoCompactWindow` listed among the deferred keys, fleetd #323 instance 1) and `ConfigRef.java:719-723`. But *deferred* is defined at `ConfigRef.java:81` as "accepted into the new snapshot, but the wiring built at startup is not rebuilt" — and `leadContextWindowLookup` reads `config.get().profiles()`, the live snapshot. So after a reload the resolved window **does** change, within one gauge instance, and two `fleet_list` calls 2 seconds apart would straddle it. Narrow, but not blocked. Fold the threshold into the cache key. The cost is one lost cache hit when a session's resolved window actually changes, which is already a cold path. ### Finding 3 — recorded, do NOT fix in this PR. `FleetMcp.java:301` and the two narrow `Fleetd` overloads are dead production surface. My counts, not a reviewer's: the only production construction is `Fleetd.java:1078` using the two-argument form; `FleetdAssembly.java:407` resolves to the **wide** `leadContextSource`; and narrow `leadContextSource:1165` has no caller at all, with the only path into narrow `leadContextLookup:1127` being that dead wrapper's own body at `:1168`. A reviewer reported 7 old-form test call sites in 3 files. My pattern also catches the unqualified `new LeadConfigDirSource(` and measures **8 across 4 files** — they missed `FleetMcpLeadContextGaugeWiringTest.java`. Use my number if anyone acts on this. Leaving it is a judgement call, not an oversight. Removing dead surface is right, but it means editing 8 test call sites in the same change as a correctness fix, and I would rather the correctness fix land clean. **Filed separately rather than bundled.** One thing worth saying about it: every hazard in this PR lives in a back-compat form. The narrow `read(configDir, sessionId, agentType)` overload passes `null`, which both silently restores the old fixed threshold and is what makes finding 2 reachable. That is a stronger reason to remove the narrow forms than "dead code" was on its own. --- ### Also noted, not a blocker The window is resolved from the **live** snapshot, while the launched Claude session is still running on the `--autocompact` flag it was spawned with. After a deferred reload those disagree, so the gauge would scale against a window the session is not actually on until it is respawned. `leadConfigDirLookup` already has exactly this property, so this PR inherits the behaviour rather than introducing it. Out of scope here; say so if you think it deserves its own ticket. ### What I did not check myself I did not re-run the author's RED-before-fix measurement for the 4-arg `read` overload. The report describes a compile error as the RED, which is a weaker form than a failing assertion — a signature that does not exist yet cannot distinguish "the behaviour is wrong" from "the behaviour has no code path". I am accepting it because findings 1 and 2 are now pinned by tests that *do* assert behaviour.
Author
Owner

Done and merged into main as 136bec8, via PR #660 (PR #657 closed unmerged, superseded).

Both gaps from my adjudication comment are fixed and I re-killed both mutations myself on the merged tree:

  • The fleet_list window wiring is now pinned by FleetdLeadConfigDirSourceWindowWiringTest. Mutating Fleetd.java:1079 to _ -> null now fails with expected: <250000> but was: <null>. Before this round that same mutation left all 1918 tests green.
  • LeadContextGauge's cache key now folds in the derived highThreshold. Restoring the old base + NUL + sessionId key fails the new TTL test with expected: <OK> but was: <HIGH>.

Merged tree: mvn -o install gave Tests run: 1922, Failures: 0, Errors: 0, Skipped: 0, exit 0. bash scripts/test-config-edit.sh gave PASS, exit 0.

Still open and deliberately out of scope here: #659 removes the dead back-compat overloads this change left behind. It was blocked on this landing and is now unblocked.

A redeploy is owed, because this change touches fleetd/src/main/. A merge is not a deployment.

Done and merged into `main` as `136bec8`, via PR #660 (PR #657 closed unmerged, superseded). Both gaps from my adjudication comment are fixed and I re-killed both mutations myself on the merged tree: - The `fleet_list` window wiring is now pinned by `FleetdLeadConfigDirSourceWindowWiringTest`. Mutating `Fleetd.java:1079` to `_ -> null` now fails with `expected: <250000> but was: <null>`. Before this round that same mutation left all 1918 tests green. - `LeadContextGauge`'s cache key now folds in the derived `highThreshold`. Restoring the old `base + NUL + sessionId` key fails the new TTL test with `expected: <OK> but was: <HIGH>`. Merged tree: `mvn -o install` gave `Tests run: 1922, Failures: 0, Errors: 0, Skipped: 0`, exit 0. `bash scripts/test-config-edit.sh` gave `PASS`, exit 0. Still open and deliberately out of scope here: #659 removes the dead back-compat overloads this change left behind. It was blocked on this landing and is now unblocked. A redeploy is owed, because this change touches `fleetd/src/main/`. A merge is not a deployment.
ltms closed this issue 2026-10-03 16:31:04 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#637