fleetd: report a lead's live context usage in fleet_list #602
Reference in New Issue
Block a user
Delete Branch "worker/lead-context-gauge-ad404f-1"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Why
fleetd could not see how full a lead's context window is. On this host a lead auto-compacted 30 times in one session. Each compaction throws away about 250,000 tokens and costs between 46s and 3m16s. Nothing could see it coming, so nothing could hand over in time.
What it does
LeadContextGaugereads the transcript file Claude Code writes for itself. It never reads the lead's pane, so it does not touch the control plane invariant 5 protects.Three design points worth reading:
<configDir>/projects/one level deep and looks for<sessionId>.jsonl. The project-slug rule is an undocumented internal of Claude Code, so deriving it would break the day that rule changes.OK,HIGH,UNKNOWN. Every path that cannot get a real token count returnsUNKNOWNwith no number. A gauge that says "fine" when it could not look is worse than no gauge.TAIL_BYTES(2 MiB) caps what is read off disk. A 5 second cache TTL caps how often that read happens, becausefleet_listis polled constantly.How this was produced
The member that wrote this ended on a backend error (DNS
ENOTFOUND) with the work uncommitted and unpushed in its worktree. I recovered it and committed it myself.Verified by me before committing, in the worktree:
KNOWN DEFECT — do not merge and call this shipped
I read the wiring after the tests went green and found the gauge is inert on this fleet.
FleetMcp.contextViewcalls:A
nullconfigDirfalls back to<user.home>/.claude. But this host's lead profile sets aconfigDiroverride infleetd.yaml. Measured just now:So the gauge would deploy and report
UNKNOWNforever, for every lead, with no error anywhere.All 8 tests pass and none of them says anything about this. Each test builds its own
LeadContextGaugeand hands it a@TempDir. A test proves the instantiation it creates, and nothing about the one the production call site creates. That is the shape to look for, not just this line.The fix is available:
fleet.leaders.<name>.profileresolves to the profile that carries the realconfigDir.Second finding — the last line may be half-written
The gauge reads a file another process is appending to. A partial final line does not parse, and the current code answers
UNKNOWNfor the whole read. That makes the gauge flap between a real number andUNKNOWNat random.The rationale given for that choice was "a format change should show as
UNKNOWN". It does not hold: a real format change makes every line unparseable, not just the last one, so dropping an unparseable final line still reportsUNKNOWNwhen the format truly changes.Both findings go to a follow-up unit. This PR is the recovered work, recorded as it was written.