Should a lead be rolled on a timer to pick up instruction changes? Measured: the roll is the expensive answer, it contradicts a stated invariant, and it has never been observed working end-to-end #595

Open
opened 2026-09-19 10:00:33 +02:00 by ltms · 0 comments
Owner

The question

The operator asked on 2026-09-19: "how about periodically restart the leader (for update catch
up)"
.

The underlying problem is real and I hit it in the same session. A lead reads CLAUDE.md once,
at session start, and holds it in context. I changed the canonical block and pushed a639969
while my own session was running. My context still holds the old block, and so does every other
running lead's.
Nothing re-reads that file, and nothing tells a lead it went stale.

So the goal — catch up on instruction changes — is worth solving. This ticket is about whether a
periodic roll is the way, and my measured answer is no.

Measured state, 2026-09-19

leadRollover: is configured on the Mac — handoverPath: .handover/HANDOVER.md, plus a
bootstrapText. So the mechanism exists here and is not hypothetical.

A timer is explicitly excluded by design. LeadRollover's class javadoc (:65-76):

The ticket that first defined this class required "no timer, no scheduler, no background thread"
so that nothing but an explicit confirm call could ever cause a /clear. That invariant is
about INITIATIVE … A recurring heartbeat or timer that could decide on its own initiative to
roll a pane is still, and will always be, absent from this class.

grep for schedule|interval|periodic|cron in that file returns only the single-shot
continuation and the settle polls. There is no scheduler to configure.

requireOperatorConfirm defaults to true (FleetConfig.java:1405). A timer has no human to
ask, so automating it means passing operatorConfirmed: true with nobody having confirmed. The
handover skill is explicit that this is the one gate between a judgement call and a wiped session:
"Do not pass true because you are confident."

The roll has never been observed working end-to-end. #489 measured it failing live on
2026-09-12: /clear and bootstrapText concatenated into /clearFresh lead session… and Claude
Code answered Unknown command. The fix is in the source and in the running jar (jar built
Sep 13 07:02, LeadRollover.java last modified Sep 12 09:44), and waitForClearPickupAndSettle
now detects a real WORKING -> IDLE/DONE pickup boundary with a bounded nudge. But #480's
criterion 6 has still never been observed passing. The documented failure mode is:
"If the bootstrap prompt never lands, your context is gone and no fresh session starts."

That is the decisive point. Automating an unproven destructive operation is the one thing not to
do
, and the grace-release path at PICKUP_GRACE_POLLS deliberately trades a possibly-unsubmitted
/clear for progress — the same path that reported false success in the original incident.

Why a roll is the wrong tool for this particular job

A roll destroys the lead's whole context and replaces it with one handover file. That is the right
trade when context is full, which is what #480 was built for. For catching up on an
instruction change it throws away everything to deliver a few KB of diff.

A timer also fires when nothing changed, and stays silent when something does. The real trigger is
not elapsed time — it is the instruction surface changing: CLAUDE.md, fleet.charters,
.claude/agents/*.md, .claude/skills/**. That is a git commit and a config reload, and the
daemon already sees the second one (it logged config reloaded at 14:55:43 today).

Recommendation, in order

  1. Notify, do not restart. When the instruction surface changes, send the lead one message
    naming the commit and what changed. A lead can re-read CLAUDE.md perfectly well on demand —
    it just never learns that it should. This is cheap, non-destructive, and needs no new gate.
    Blocked on #594: the existing push-into-a-lead's-pane path gates on injectable(), which
    accepts BLOCKED, so today that notice can land inside an approval prompt.
  2. Keep the roll lead-initiated, for context pressure. That is #480 working as designed.
  3. If it is ever automated, gate it on a signal, not a clock — instruction surface changed
    AND lead idle (never BLOCKED) AND no live members AND the roll proven end-to-end at least
    once. Elapsed time is none of those.

Before anything here is built

Do one attended roll and answer #480 criterion 6. It is a single measurement, it costs one
lead session, and every option above depends on knowing whether the bootstrap prompt lands. Until
then the honest status is "fixed in source, never seen working".

Related: #480, #489 (the measured failure), #486 (settle poll unbounded), #491 (relative
handoverPath), #482 (a context reset changes the session id), #594 (the BLOCKED gate that
blocks recommendation 1).

## The question The operator asked on 2026-09-19: *"how about periodically restart the leader (for update catch up)"*. The underlying problem is real and I hit it in the same session. A lead reads `CLAUDE.md` once, at session start, and holds it in context. I changed the canonical block and pushed `a639969` while my own session was running. **My context still holds the old block, and so does every other running lead's.** Nothing re-reads that file, and nothing tells a lead it went stale. So the goal — catch up on instruction changes — is worth solving. This ticket is about whether a periodic roll is the way, and my measured answer is no. ## Measured state, 2026-09-19 **`leadRollover:` is configured on the Mac** — `handoverPath: .handover/HANDOVER.md`, plus a `bootstrapText`. So the mechanism exists here and is not hypothetical. **A timer is explicitly excluded by design.** `LeadRollover`'s class javadoc (`:65-76`): > The ticket that first defined this class required "no timer, no scheduler, no background thread" > so that nothing but an explicit `confirm` call could ever cause a `/clear`. That invariant is > about INITIATIVE … **A recurring heartbeat or timer that could decide on its own initiative to > roll a pane is still, and will always be, absent from this class.** `grep` for `schedule|interval|periodic|cron` in that file returns only the single-shot continuation and the settle polls. There is no scheduler to configure. **`requireOperatorConfirm` defaults to `true`** (`FleetConfig.java:1405`). A timer has no human to ask, so automating it means passing `operatorConfirmed: true` with nobody having confirmed. The handover skill is explicit that this is the one gate between a judgement call and a wiped session: *"Do not pass `true` because you are confident."* **The roll has never been observed working end-to-end.** #489 measured it failing live on 2026-09-12: `/clear` and `bootstrapText` concatenated into `/clearFresh lead session…` and Claude Code answered `Unknown command`. The fix is in the source and in the running jar (jar built Sep 13 07:02, `LeadRollover.java` last modified Sep 12 09:44), and `waitForClearPickupAndSettle` now detects a real `WORKING -> IDLE/DONE` pickup boundary with a bounded nudge. But #480's criterion 6 has still never been *observed* passing. The documented failure mode is: *"If the bootstrap prompt never lands, your context is gone and no fresh session starts."* That is the decisive point. **Automating an unproven destructive operation is the one thing not to do**, and the grace-release path at `PICKUP_GRACE_POLLS` deliberately trades a possibly-unsubmitted `/clear` for progress — the same path that reported false success in the original incident. ## Why a roll is the wrong tool for this particular job A roll destroys the lead's whole context and replaces it with one handover file. That is the right trade when context is **full**, which is what #480 was built for. For catching up on an instruction change it throws away everything to deliver a few KB of diff. A timer also fires when nothing changed, and stays silent when something does. The real trigger is not elapsed time — it is **the instruction surface changing**: `CLAUDE.md`, `fleet.charters`, `.claude/agents/*.md`, `.claude/skills/**`. That is a git commit and a config reload, and the daemon already sees the second one (it logged `config reloaded` at 14:55:43 today). ## Recommendation, in order 1. **Notify, do not restart.** When the instruction surface changes, send the lead one message naming the commit and what changed. A lead can re-read `CLAUDE.md` perfectly well on demand — it just never learns that it should. This is cheap, non-destructive, and needs no new gate. **Blocked on #594**: the existing push-into-a-lead's-pane path gates on `injectable()`, which accepts `BLOCKED`, so today that notice can land inside an approval prompt. 2. **Keep the roll lead-initiated, for context pressure.** That is #480 working as designed. 3. **If it is ever automated, gate it on a signal, not a clock** — instruction surface changed AND lead idle (never `BLOCKED`) AND no live members AND the roll proven end-to-end at least once. Elapsed time is none of those. ## Before anything here is built **Do one attended roll and answer #480 criterion 6.** It is a single measurement, it costs one lead session, and every option above depends on knowing whether the bootstrap prompt lands. Until then the honest status is "fixed in source, never seen working". Related: #480, #489 (the measured failure), #486 (settle poll unbounded), #491 (relative `handoverPath`), #482 (a context reset changes the session id), #594 (the `BLOCKED` gate that blocks recommendation 1).
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#595