CB-617: pass auto-compact window to leads #601

Merged
ltms merged 2 commits from worker/lead-autocompact-5f1ab2-3 into main 2026-09-22 05:22:57 +02:00
Member

Routes Claude profiles' autoCompactWindow to fleetd-launched leads, rejects conflicting environment settings, and updates comments. Tests: mvn -q clean install passed.

Routes Claude profiles' autoCompactWindow to fleetd-launched leads, rejects conflicting environment settings, and updates comments. Tests: mvn -q clean install passed.
Owner

Do not merge this onto the Mac fleet as it stands

Measured today, 2026-09-20. This PR is mergeable: true and had no comments, so nothing on it recorded the hazard. Writing it down before someone merges it on a clean read of the diff.

What the PR adds

FleetConfig.load(...) now calls rejectConflictingAutoCompactWindows(yaml), which throws:

IllegalStateException("refusing to start: Claude Code profile(s) [...] ... CLAUDE_CODE_AUTO_COMPACT_WINDOW; remove one key or set equal values.")

when a Claude Code profile sets both autoCompactWindow and the CLAUDE_CODE_AUTO_COMPACT_WINDOW env var to values that are not equal.

Why that is a problem here

Four profiles in this host's fleetd.yaml do exactly that, and the two values differ — 250000 against "300000":

profile autoCompactWindow CLAUDE_CODE_AUTO_COMPACT_WINDOW
local line 25 line 32
local-direct line 74 line 80
opus line 110 line 118
sonnet line 124 line 165

opus is the lead's own profile and sonnet is the one every worker spawns on, so this is not a corner of the config nobody uses.

The refusal happens inside load(...), which means the daemon does not start. Under launchd that is a restart loop, not a clean error an operator sees once. The whole fleet goes down, and the config that would fix it is gitignored, so the failure reaches a host where the repo cannot show you the cause.

Why it is not the PR's fault

The validation itself is reasonable — Claude Code gives the env var priority over the flag, so setting both to different numbers means the YAML is lying about what the session will do. The problem is sequencing, not correctness.

What has to happen first

Reconcile the four profiles in fleetd/fleetd.yaml — drop one key per profile, or set both to the same number — then merge this. The config file is gitignored, so that step cannot ride along in this PR and cannot be verified by CI. It has to be done on each host by hand, and this PR should not merge until the hosts it will run on are known to be clean.

Worth adding to the PR before it lands: make the refusal name the offending profiles and both values in the message, so an operator reading a crash-looping daemon's log can fix it without the source.

I have not checked fleet01's config for the same conflict. Someone should, before this merges.

## Do not merge this onto the Mac fleet as it stands Measured today, 2026-09-20. This PR is `mergeable: true` and had no comments, so nothing on it recorded the hazard. Writing it down before someone merges it on a clean read of the diff. ### What the PR adds `FleetConfig.load(...)` now calls `rejectConflictingAutoCompactWindows(yaml)`, which throws: ``` IllegalStateException("refusing to start: Claude Code profile(s) [...] ... CLAUDE_CODE_AUTO_COMPACT_WINDOW; remove one key or set equal values.") ``` when a Claude Code profile sets **both** `autoCompactWindow` and the `CLAUDE_CODE_AUTO_COMPACT_WINDOW` env var to values that are not equal. ### Why that is a problem here Four profiles in this host's `fleetd.yaml` do exactly that, and the two values differ — `250000` against `"300000"`: | profile | `autoCompactWindow` | `CLAUDE_CODE_AUTO_COMPACT_WINDOW` | |---|---|---| | `local` | line 25 | line 32 | | `local-direct` | line 74 | line 80 | | `opus` | line 110 | line 118 | | `sonnet` | line 124 | line 165 | `opus` is the lead's own profile and `sonnet` is the one every worker spawns on, so this is not a corner of the config nobody uses. The refusal happens inside `load(...)`, which means the daemon does not start. Under launchd that is a restart loop, not a clean error an operator sees once. The whole fleet goes down, and the config that would fix it is gitignored, so the failure reaches a host where the repo cannot show you the cause. ### Why it is not the PR's fault The validation itself is reasonable — Claude Code gives the env var priority over the flag, so setting both to different numbers means the YAML is lying about what the session will do. The problem is sequencing, not correctness. ### What has to happen first Reconcile the four profiles in `fleetd/fleetd.yaml` — drop one key per profile, or set both to the same number — **then** merge this. The config file is gitignored, so that step cannot ride along in this PR and cannot be verified by CI. It has to be done on each host by hand, and this PR should not merge until the hosts it will run on are known to be clean. Worth adding to the PR before it lands: make the refusal name the offending profiles and both values in the message, so an operator reading a crash-looping daemon's log can fix it without the source. I have not checked fleet01's config for the same conflict. Someone should, before this merges.
Owner

Re-measured 2026-09-22, with a control — and one claim above is not checked

I re-ran this today before deciding what to do with the PR. Two things: the hazard is real and I can now show it with a control, and the comment above states a fact about Claude Code that nobody here has measured.

The refusal fires. Measured, with a control.

I built this PR branch in a scratch worktree and ran its own FleetConfig.load(Path) against this host's live fleetd/fleetd.yaml. Then I ran the deployed jar's loader against the same file, as a control.

=== PR601 jar vs live config ===
REFUSED: IllegalStateException
MESSAGE: refusing to start: Claude Code profile(s) [local, local-direct, opus, sonnet] set
disagreeing autoCompactWindow and env.CLAUDE_CODE_AUTO_COMPACT_WINDOW; remove one key or set
equal values.

=== DEPLOYED jar (main) vs live config — control ===
LOADED OK

The control matters. Without it, "the loader refuses" could have been a broken config or a bad probe. The same file, same probe, loads fine on main and is refused on this branch. So the refusal comes from this PR and nothing else.

The four profiles and their lines, re-measured today, unchanged from the report above: local 25/32, local-direct 74/80, opus 110/118, sonnet 124/165. All are 250000 against "300000".

The message already names the profiles

The comment above asks for the refusal to "name the offending profiles and both values". Half of that is already done — the message names [local, local-direct, opus, sonnet]. It does not print the two values. Only the values half is still missing, so please do not re-add the profile list.

The claim that is not checked: which input wins?

The comment above says "Claude Code gives the env var priority over the flag". This PR's own javadoc in ClaudeCodeArguments says the same thing.

But this host's fleetd/fleetd.yaml line 25 says the opposite, in a comment:

autoCompactWindow: 250000  # CB-636: --autocompact FLAG (outranks the env var below)

So this repo now asserts both directions as fact, in two places. One of them is wrong. I have not measured which, and I could not find a safe way to measure it from here. What I did check: the flag is real in the installed CLI (Claude Code 2.1.278), and both numbers are in range.

--autocompact <auto|tokens>    Auto-compact window size (auto, or 100k–1M tokens)

The help does not state precedence.

This is why "just reconcile the config first" is not yet safe. The two outcomes are different:

  • If the env var wins, the fleet is compacting at 300000 today. Deleting the env var to fix the conflict would silently move every lead and worker to 250000.
  • If the flag wins, the fleet is at 250000 today, and deleting the env var changes nothing.

Reconciling before we know which is true is a coin flip on live behaviour, on the lead's own profile and on every worker's.

What I think should change in the PR

The sequencing problem goes away if the guard is not fatal. A throw inside load() under launchd is a restart loop, and the config that would fix it is gitignored, so the cause is invisible on the host where it bites. A WARN at boot, naming the profiles and both values, gives the operator exactly the same information and cannot take the fleet down. It also makes this PR safe to merge before every host is reconciled, which removes the cross-host coordination step entirely.

The feature itself — passing --autocompact to leads, not just members — is good and I want it. It serves the lead-context work directly. It is only the guard's severity that blocks it.

Not merging this yet. Tracking the precedence question as the thing that has to be settled before any host's config is edited.

## Re-measured 2026-09-22, with a control — and one claim above is not checked I re-ran this today before deciding what to do with the PR. Two things: the hazard is real and I can now show it with a control, and the comment above states a fact about Claude Code that nobody here has measured. ### The refusal fires. Measured, with a control. I built this PR branch in a scratch worktree and ran its own `FleetConfig.load(Path)` against this host's live `fleetd/fleetd.yaml`. Then I ran the **deployed** jar's loader against the same file, as a control. ``` === PR601 jar vs live config === REFUSED: IllegalStateException MESSAGE: refusing to start: Claude Code profile(s) [local, local-direct, opus, sonnet] set disagreeing autoCompactWindow and env.CLAUDE_CODE_AUTO_COMPACT_WINDOW; remove one key or set equal values. === DEPLOYED jar (main) vs live config — control === LOADED OK ``` The control matters. Without it, "the loader refuses" could have been a broken config or a bad probe. The same file, same probe, loads fine on `main` and is refused on this branch. So the refusal comes from this PR and nothing else. The four profiles and their lines, re-measured today, unchanged from the report above: `local` 25/32, `local-direct` 74/80, `opus` 110/118, `sonnet` 124/165. All are `250000` against `"300000"`. ### The message already names the profiles The comment above asks for the refusal to "name the offending profiles and both values". Half of that is already done — the message names `[local, local-direct, opus, sonnet]`. It does **not** print the two values. Only the values half is still missing, so please do not re-add the profile list. ### The claim that is not checked: which input wins? The comment above says "Claude Code gives the env var priority over the flag". This PR's own javadoc in `ClaudeCodeArguments` says the same thing. But this host's `fleetd/fleetd.yaml` line 25 says the opposite, in a comment: ```yaml autoCompactWindow: 250000 # CB-636: --autocompact FLAG (outranks the env var below) ``` So this repo now asserts both directions as fact, in two places. One of them is wrong. **I have not measured which, and I could not find a safe way to measure it from here.** What I did check: the flag is real in the installed CLI (Claude Code 2.1.278), and both numbers are in range. ``` --autocompact <auto|tokens> Auto-compact window size (auto, or 100k–1M tokens) ``` The help does not state precedence. **This is why "just reconcile the config first" is not yet safe.** The two outcomes are different: - If the **env var** wins, the fleet is compacting at `300000` today. Deleting the env var to fix the conflict would silently move every lead and worker to `250000`. - If the **flag** wins, the fleet is at `250000` today, and deleting the env var changes nothing. Reconciling before we know which is true is a coin flip on live behaviour, on the lead's own profile and on every worker's. ### What I think should change in the PR The sequencing problem goes away if the guard is not fatal. A `throw` inside `load()` under launchd is a restart loop, and the config that would fix it is gitignored, so the cause is invisible on the host where it bites. A **WARN at boot**, naming the profiles and both values, gives the operator exactly the same information and cannot take the fleet down. It also makes this PR safe to merge before every host is reconciled, which removes the cross-host coordination step entirely. The feature itself — passing `--autocompact` to leads, not just members — is good and I want it. It serves the lead-context work directly. It is only the guard's severity that blocks it. Not merging this yet. Tracking the precedence question as the thing that has to be settled before any host's config is edited.
agent added 2 commits 2026-09-22 05:15:46 +02:00
CB-617 review: warn instead of refuse on conflicting autoCompactWindow
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 2m9s
be6c45ff78
rejectConflictingAutoCompactWindows threw and stopped fleetd from starting when a
Claude Code profile's autoCompactWindow flag and CLAUDE_CODE_AUTO_COMPACT_WINDOW
env var disagreed. Under launchd that is a restart loop, and the config that
would fix it (fleetd.yaml) is gitignored, so the cause is invisible on the host
where it bites (measured live: 4 profiles on this host trip it, including the
lead's own profile and the one every worker spawns on).

Renamed to warnConflictingAutoCompactWindows: it now logs a WARN naming each
offending profile with BOTH values (autoCompactWindow=... and
env.CLAUDE_CODE_AUTO_COMPACT_WINDOW=...) instead of throwing, so the daemon
starts and an operator can fix the config without reading the source. Equal
values still load silently.

Also reworded ClaudeCodeArguments' javadoc, which stated as fact that the env
var takes precedence over the flag. That was never measured, and this host's
own fleetd.yaml comment asserts the opposite — the javadoc no longer picks a
side.
agent force-pushed worker/lead-autocompact-5f1ab2-3 from 792ea85d78 to be6c45ff78 2026-09-22 05:15:46 +02:00 Compare
ltms merged commit 17127efb88 into main 2026-09-22 05:22:57 +02:00
Owner

Merged. The start-up hazard is gone — measured with the same probe that refused before.

The guard now warns instead of throwing, so this no longer has to wait for every host's config to be reconciled first. That also removes the cross-host coordination step I flagged earlier.

The decisive check

I re-ran the exact probe from my 2026-09-22 comment — the built jar's own FleetConfig.load(Path) against this host's live fleetd/fleetd.yaml. Before, it returned REFUSED. Now:

WARN  d.l.fleet.config.FleetConfig - Claude Code profile(s) [local, local-direct, opus, sonnet]
set disagreeing autoCompactWindow and env.CLAUDE_CODE_AUTO_COMPACT_WINDOW — the daemon starts
anyway. Fix by removing one key or setting equal values on each: local (autoCompactWindow=250000,
env.CLAUDE_CODE_AUTO_COMPACT_WINDOW=300000), local-direct (...), opus (...), sonnet (...).
Which input Claude Code actually follows when they disagree is not verified here.
LOADED OK

LOADED OK. Paired with a negative control — a config with no conflict loads and prints no WARN at all, so this is not a warning that always fires.

The message does what was asked: keeps the existing profile list, adds both values per profile, and says plainly that the precedence is not verified. That last sentence is the part I care most about — the warning no longer teaches a reader something nobody has measured.

Also verified by me

  • My own build, mvn -o clean install in a scratch worktree, unpiped: Tests run: 1869, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS.
  • The javadoc no longer picks a side. ClaudeCodeArguments now records that the two inputs can disagree, that this file used to claim the env var wins, that nobody measured it, and that fleetd.yaml asserts the opposite. That is the right outcome: the contradiction is documented rather than silently resolved in one direction.
  • The old throw-assertion test was replaced, not deleted. It pinned "an IllegalStateException is thrown on conflict". That can no longer go wrong because throwing is no longer the behaviour; a WARN-assertion test pins the new behaviour, using the repo's existing CapturedLog helper rather than a hand-rolled appender.
  • Nothing forbidden committed. No fleetd.yaml, no .mcp.json, no wiki pointer.
  • The branch was 24 behind main and was rebased cleanly before merge.

Still open, and deliberately not answered here

Which input Claude Code actually honours is still unmeasured. Nothing in this PR settles it and nothing should be read as settling it. Until someone measures it:

  • do not "reconcile" a host's config by deleting one of the two keys — if the env var is the one that wins, deleting it moves that profile from 300000 to 250000;
  • treat both the fleetd.yaml:25 comment and any doc claiming the opposite as unverified.

One behaviour change to know about at redeploy

This PR's actual feature is that leads now get --autocompact too, where before only members did. So after the next redeploy a newly launched lead is passed the same flag its workers already received. Whichever input the backend honours, leads and members now resolve it the same way instead of differing — which is the point of the change. No lead currently running is affected; this applies at the next lead launch.

The daemon is still running the jar built from 076cc43, so none of this is live yet.

## Merged. The start-up hazard is gone — measured with the same probe that refused before. The guard now warns instead of throwing, so this no longer has to wait for every host's config to be reconciled first. That also removes the cross-host coordination step I flagged earlier. ### The decisive check I re-ran the **exact probe** from my 2026-09-22 comment — the built jar's own `FleetConfig.load(Path)` against this host's live `fleetd/fleetd.yaml`. Before, it returned `REFUSED`. Now: ``` WARN d.l.fleet.config.FleetConfig - Claude Code profile(s) [local, local-direct, opus, sonnet] set disagreeing autoCompactWindow and env.CLAUDE_CODE_AUTO_COMPACT_WINDOW — the daemon starts anyway. Fix by removing one key or setting equal values on each: local (autoCompactWindow=250000, env.CLAUDE_CODE_AUTO_COMPACT_WINDOW=300000), local-direct (...), opus (...), sonnet (...). Which input Claude Code actually follows when they disagree is not verified here. LOADED OK ``` `LOADED OK`. Paired with a negative control — a config with no conflict loads and prints no WARN at all, so this is not a warning that always fires. The message does what was asked: keeps the existing profile list, adds **both values per profile**, and says plainly that the precedence is not verified. That last sentence is the part I care most about — the warning no longer teaches a reader something nobody has measured. ### Also verified by me - **My own build**, `mvn -o clean install` in a scratch worktree, unpiped: `Tests run: 1869, Failures: 0, Errors: 0, Skipped: 0` — BUILD SUCCESS. - **The javadoc no longer picks a side.** `ClaudeCodeArguments` now records that the two inputs can disagree, that this file used to claim the env var wins, that nobody measured it, and that `fleetd.yaml` asserts the opposite. That is the right outcome: the contradiction is documented rather than silently resolved in one direction. - **The old throw-assertion test was replaced, not deleted.** It pinned "an `IllegalStateException` is thrown on conflict". That can no longer go wrong because throwing is no longer the behaviour; a WARN-assertion test pins the new behaviour, using the repo's existing `CapturedLog` helper rather than a hand-rolled appender. - **Nothing forbidden committed.** No `fleetd.yaml`, no `.mcp.json`, no `wiki` pointer. - The branch was 24 behind `main` and was rebased cleanly before merge. ### Still open, and deliberately not answered here **Which input Claude Code actually honours is still unmeasured.** Nothing in this PR settles it and nothing should be read as settling it. Until someone measures it: - do not "reconcile" a host's config by deleting one of the two keys — if the env var is the one that wins, deleting it moves that profile from `300000` to `250000`; - treat both the `fleetd.yaml:25` comment and any doc claiming the opposite as unverified. ### One behaviour change to know about at redeploy This PR's actual feature is that **leads** now get `--autocompact` too, where before only members did. So after the next redeploy a newly launched lead is passed the same flag its workers already received. Whichever input the backend honours, leads and members now resolve it the same way instead of differing — which is the point of the change. No lead currently running is affected; this applies at the next lead launch. The daemon is still running the jar built from `076cc43`, so none of this is live yet.
Sign in to join this conversation.