fleetd: report a lead's live context usage in fleet_list #602

Merged
ltms merged 4 commits from worker/lead-context-gauge-ad404f-1 into main 2026-09-20 11:38:31 +02:00
Owner

Why

fleetd could not see how full a lead's context window is. On this host a lead auto-compacted 30 times in one session. Each compaction throws away about 250,000 tokens and costs between 46s and 3m16s. Nothing could see it coming, so nothing could hand over in time.

What it does

LeadContextGauge reads the transcript file Claude Code writes for itself. It never reads the lead's pane, so it does not touch the control plane invariant 5 protects.

Three design points worth reading:

  1. It finds the transcript by name, not by slug. It lists <configDir>/projects/ one level deep and looks for <sessionId>.jsonl. The project-slug rule is an undocumented internal of Claude Code, so deriving it would break the day that rule changes.
  2. Three states, not two: OK, HIGH, UNKNOWN. Every path that cannot get a real token count returns UNKNOWN with no number. A gauge that says "fine" when it could not look is worse than no gauge.
  3. Bounded twice. TAIL_BYTES (2 MiB) caps what is read off disk. A 5 second cache TTL caps how often that read happens, because fleet_list is polled constantly.

How this was produced

The member that wrote this ended on a backend error (DNS ENOTFOUND) with the work uncommitted and unpushed in its worktree. I recovered it and committed it myself.

Verified by me before committing, in the worktree:

mvn -o clean install exit: 0
reports=144 tests=1813 failures=0 errors=0 skipped=0
LeadContextGaugeTest: tests="8" errors="0" skipped="0" failures="0"

KNOWN DEFECT — do not merge and call this shipped

I read the wiring after the tests went green and found the gauge is inert on this fleet.

FleetMcp.contextView calls:

LeadContextGauge.Reading reading = contextGauge.read(null, sessionId, agentType);

A null configDir falls back to <user.home>/.claude. But this host's lead profile sets a configDir override in fleetd.yaml. Measured just now:

jsonl files under the profile configDir: 19
~/.claude/projects EXISTS
my session found under the fallback: 0
fleet.leaders.opus.profile: opus

So the gauge would deploy and report UNKNOWN forever, for every lead, with no error anywhere.

All 8 tests pass and none of them says anything about this. Each test builds its own LeadContextGauge and hands it a @TempDir. A test proves the instantiation it creates, and nothing about the one the production call site creates. That is the shape to look for, not just this line.

The fix is available: fleet.leaders.<name>.profile resolves to the profile that carries the real configDir.

Second finding — the last line may be half-written

The gauge reads a file another process is appending to. A partial final line does not parse, and the current code answers UNKNOWN for the whole read. That makes the gauge flap between a real number and UNKNOWN at random.

The rationale given for that choice was "a format change should show as UNKNOWN". It does not hold: a real format change makes every line unparseable, not just the last one, so dropping an unparseable final line still reports UNKNOWN when the format truly changes.

Both findings go to a follow-up unit. This PR is the recovered work, recorded as it was written.

## Why fleetd could not see how full a lead's context window is. On this host a lead auto-compacted 30 times in one session. Each compaction throws away about 250,000 tokens and costs between 46s and 3m16s. Nothing could see it coming, so nothing could hand over in time. ## What it does `LeadContextGauge` reads the transcript file Claude Code writes for itself. It never reads the lead's pane, so it does not touch the control plane invariant 5 protects. Three design points worth reading: 1. **It finds the transcript by name, not by slug.** It lists `<configDir>/projects/` one level deep and looks for `<sessionId>.jsonl`. The project-slug rule is an undocumented internal of Claude Code, so deriving it would break the day that rule changes. 2. **Three states, not two: `OK`, `HIGH`, `UNKNOWN`.** Every path that cannot get a real token count returns `UNKNOWN` with no number. A gauge that says "fine" when it could not look is worse than no gauge. 3. **Bounded twice.** `TAIL_BYTES` (2 MiB) caps what is read off disk. A 5 second cache TTL caps how often that read happens, because `fleet_list` is polled constantly. ## How this was produced The member that wrote this ended on a backend error (DNS `ENOTFOUND`) with the work uncommitted and unpushed in its worktree. I recovered it and committed it myself. Verified by me before committing, in the worktree: ``` mvn -o clean install exit: 0 reports=144 tests=1813 failures=0 errors=0 skipped=0 LeadContextGaugeTest: tests="8" errors="0" skipped="0" failures="0" ``` ## KNOWN DEFECT — do not merge and call this shipped I read the wiring after the tests went green and found the gauge is **inert on this fleet**. `FleetMcp.contextView` calls: ```java LeadContextGauge.Reading reading = contextGauge.read(null, sessionId, agentType); ``` A `null` `configDir` falls back to `<user.home>/.claude`. But this host's lead profile sets a `configDir` override in `fleetd.yaml`. Measured just now: ``` jsonl files under the profile configDir: 19 ~/.claude/projects EXISTS my session found under the fallback: 0 fleet.leaders.opus.profile: opus ``` So the gauge would deploy and report `UNKNOWN` forever, for every lead, with no error anywhere. **All 8 tests pass and none of them says anything about this.** Each test builds its own `LeadContextGauge` and hands it a `@TempDir`. A test proves the instantiation it creates, and nothing about the one the production call site creates. That is the shape to look for, not just this line. The fix is available: `fleet.leaders.<name>.profile` resolves to the profile that carries the real `configDir`. ### Second finding — the last line may be half-written The gauge reads a file another process is appending to. A partial final line does not parse, and the current code answers `UNKNOWN` for the whole read. That makes the gauge flap between a real number and `UNKNOWN` at random. The rationale given for that choice was "a format change should show as `UNKNOWN`". It does not hold: a real format change makes *every* line unparseable, not just the last one, so dropping an unparseable **final** line still reports `UNKNOWN` when the format truly changes. Both findings go to a follow-up unit. This PR is the recovered work, recorded as it was written.
ltms added 1 commit 2026-09-20 11:06:28 +02:00
fleetd: report a lead's live context usage in fleet_list
CI / shell-tests (pull_request) Failing after 7s
CI / build (pull_request) Failing after 1m37s
CI / contract (pull_request) Successful in 2m7s
3ed7bfca67
fleetd had no way to see how full a lead's context window is. On this host a
lead auto-compacted 30 times in one session, discarding roughly 250,000 tokens
and costing 46s to 3m16s each time, and nothing could see it coming.

LeadContextGauge reads the transcript Claude Code itself writes, never the
lead's pane. It finds <sessionId>.jsonl by NAME under <configDir>/projects/
rather than deriving the project slug, which is an undocumented internal.

Three states, not two: OK, HIGH, UNKNOWN. Every path that cannot positively
establish a token count reports UNKNOWN with no number, so a lead is never
told it is fine when the honest answer is "I could not look".

Bounded two ways: TAIL_BYTES caps bytes read off disk, and a 5s cache TTL caps
how often that read happens, because fleet_list is polled constantly.

Recovered by the lead: the authoring member ended on a backend error (DNS
ENOTFOUND) with this work uncommitted and unpushed in its worktree. Verified
before committing: mvn -o clean install exit 0, 1813 tests, 0 failures,
0 errors, 0 skipped, 144 reports; LeadContextGaugeTest 8/8.

KNOWN INCOMPLETE - see the PR. The fleet_list call site passes configDir=null,
which falls back to ~/.claude, but this host's lead profile sets configDir to
an override. Measured: 19 transcripts under the real configDir, 0 under the
fallback. The gauge is therefore INERT on this fleet until that is wired.
ltms added 3 commits 2026-09-20 11:38:23 +02:00
FleetMcp.contextView hardcoded LeadContextGauge.read(null, ...), so a lead whose
profile sets its own CLAUDE_CONFIG_DIR always read the wrong transcript directory
and reported UNKNOWN forever, with no error anywhere.

- Add FleetMcp.LeadConfigDirSource (same idiom as LeadSeatSource) and thread it
  through the constructor / listFleet overload chain / leadView / contextView.
- Add Fleetd.leadConfigDirLookup, wired at construction, following the same
  fleet.leaders.<name>.profile link leadSeatLookup already uses, one step
  further to that profile's own configDir.
- LeadContextGauge.parse: a single unparseable line (typically the final one,
  torn by a write this read raced) is now skipped rather than forcing UNKNOWN;
  only when every line in the read window fails to parse does it report
  UNKNOWN, which is the real format-change signal.
- Tests: FleetdLeadConfigDirLookupTest (lookup logic), FleetMcpLeadContextGaugeWiringTest
  (end-to-end: config naming directory A vs B decides which is read; a lead with
  no configured dir degrades without throwing), and two replacement properties in
  LeadContextGaugeTest for the torn-line fix plus its all-unparseable control.
Extract the inline new FleetMcp.LeadConfigDirSource(leadConfigDirLookup(...))
construction in Fleetd.main into a package-private factory,
Fleetd.leadConfigDirSource, mirroring loopHealthSource/capacitySource/
healthCoverageSource. Add FleetdLeadConfigDirSourceWiringTest, which calls the
factory directly with real Profile/Leader fixtures and asserts the returned
source resolves a real configDir -- a property that is false if the factory's
body is mutated to return LeadConfigDirSource.none().

Neither FleetMcpLeadContextGaugeWiringTest nor FleetdLeadConfigDirLookupTest
could catch main losing this wiring: each builds its own instance instead of
calling what main calls. This closes that gap at the factory level, matching
the standard already accepted for loopHealthSource's own wiring test.
Merge #606: wire the lead context gauge to the real configDir, and stop flapping on a torn line
CI / shell-tests (pull_request) Failing after 6s
CI / build (pull_request) Failing after 1m21s
CI / contract (pull_request) Successful in 1m50s
3762aca307
Fixes the two defects recorded in #602's own body.

1. FleetMcp.contextView passed configDir=null, so the gauge read <user.home>/
   .claude while this host's lead profile sets an override. Measured before the
   fix: 19 transcripts under the real directory, 0 under the fallback. The gauge
   would have deployed green and reported UNKNOWN forever, for every lead.
   Now threaded via FleetMcp.LeadConfigDirSource, built by Fleetd
   .leadConfigDirSource, following fleet.leaders.<name>.profile to that
   profile's configDir and reading config.get() live inside the lambda.

2. A torn final line no longer means UNKNOWN. fleet_list reads a transcript
   Claude Code may be mid-write on, so the last line can be cut. The old code
   treated that as fatal, which would make the gauge flap at random. The stated
   reason ("a format change should show as UNKNOWN") does not hold: a real
   format change makes EVERY line unparseable, and that case is still caught.

Verified by the lead on the combined state with current main merged in, not on
the branch alone: mvn -o clean install exit 0, 147 reports, 1841 tests, 0
failures, 0 errors, 0 skipped.

Mutation-checked independently by the lead:
  Fleetd.leadConfigDirSource body -> none()   -> KILLED (1 failure)
  FleetMcp.contextView configDir -> null      -> KILLED (worker-measured)

KNOWN RESIDUAL, documented rather than overclaimed. main's own one-line call to
leadConfigDirSource could be swapped for none() and the suite stays green. Every
member of this wiring-test family (loopHealthSource, capacitySource,
healthCoverageSource) has the identical gap — no test runs Fleetd.main far enough
to observe which factory it called. Filed separately as a class-wide problem
rather than patched here. The empirical close is the dogfood check after redeploy.
ltms merged commit 9a992d0f70 into main 2026-09-20 11:38:31 +02:00
Sign in to join this conversation.