Features: a lead session can replace itself when its context fills up

fleetd #480 (#483, #484, #485). Covers what the cycle is, the leadRollover:
knob and its defaults, why the three design decisions were made the way they
were (the lead asks and nothing watches it; /clear in the same pane rather than
stop-and-relaunch; the operator confirms too), and six gotchas.

Two of those gotchas are measured facts that cost real debugging: /clear routed
through Injector wedges the pane for ever because it produces no turn boundary,
and the settle waits must require IDLE or DONE rather than injectable(), which
also accepts BLOCKED — a live turn paused on an approval prompt.

The last gotcha records what is NOT proven: nothing has yet shown end-to-end
that a freshly cleared pane picks up the bootstrap prompt. Stated with the
bound on the damage if it does not, rather than left out.

7-Use-Cases.md carries the matching fleet_handover row in the portable
CLAUDE.md block template, byte-identical with the project's own copy.
Dai Ha
2026-09-11 07:30:09 +07:00
parent be8f9a97ec
commit 1eadfbc6e5
2 changed files with 85 additions and 0 deletions
+84
@@ -5421,3 +5421,87 @@ quarantine line should not have to work out which of several configured model na
inert lambda leaves the whole suite green. Tracked in fleetd #460.
fleetd #446.
## A lead session can replace itself when its context fills up
**What.** A lead (primary) session runs out of context and has to be replaced by a fresh one. Until
now that was entirely manual: the lead wrote a handover file, told the operator where it was, and
the operator started a new session by hand and pointed it at the file. Now the lead can ask fleetd
to do the swap.
The cycle has three steps, and the order matters:
1. `fleet_handover{action: "open"}` — fleetd records a token and tells the lead the `handoverPath`
it must write to.
2. The lead writes the handover file. The `handover` skill is the procedure for what goes in it.
3. `fleet_handover{action: "confirm", token, operatorConfirmed}` — fleetd checks every gate and, if
all of them pass, schedules the roll: wait for the lead's own turn to end, send `/clear` to its
pane, wait again, then send a bootstrap prompt naming the handover file. A fresh session reads
the file and carries on.
`{action: "cancel", token}` drops a pending request without rolling.
**The knob.** A new top-level `leadRollover:` block in `fleetd.yaml`. It is **opt-in and off by
default** — with the block absent, `fleet_handover` is still registered but every action answers a
clean refusal naming `NOT_CONFIGURED`.
```yaml
leadRollover:
handoverPath: /path/to/HANDOVER.md # required when the block is present
requireOperatorConfirm: true # default true
maxDocAgeSeconds: 3600 # default 3600
turnSettleSeconds: 20 # default 20
clearSettleSeconds: 20 # default 20
bootstrapText: "…" # default names handoverPath
```
Every field is read fresh on each call, so the values are hot. **Adding the block where it was
absent at boot still needs a restart**, because `Fleetd.main` only constructs the executor when the
block is present in the startup snapshot. That is an existence fact, not a stale-value one — the
same way a brand-new `profiles:` entry needs a restart while an existing profile's fields do not.
**Why it exists.** A lead that fills its context is the one agent nobody else can replace: it holds
the orchestration state, and a worker cannot restart its own lead. The whole cycle previously
stopped dead waiting for a human to notice. Three design choices were deliberate and are worth not
re-litigating:
- **The lead asks; nothing watches it.** There is no timer, no heartbeat, and no background loop
that can decide on its own that a lead should be replaced. Only an explicit `confirm()` call that
passes every gate can ever cause a `/clear`. A context-pressure detector was considered and
rejected — the cost of a false positive is a destroyed live session.
- **`/clear` in the same pane, not stop-and-relaunch.** Relaunching would lose the pane's identity,
and a lead is pinned to its terminal by a tab label (`fleet.leaders.<name>.tab`), so a new pane is
a new lead as far as the daemon is concerned.
- **The operator confirms too.** `requireOperatorConfirm` defaults to `true`, so the lead's own
judgement is not enough to wipe a session.
**Gotchas.**
- **Write the file *after* `open`, never before.** `confirm` refuses with `HANDOVER_STALE` unless
the file's modified time is later than the `open` request. That check exists so a leftover file
from a previous session can never be accepted as this session's handover, and it means the
obvious order — write the file, then ask — is the wrong one.
- **`confirm` returning `accepted` does NOT mean the pane has been cleared.** It means every gate
passed and the roll is scheduled. The roll itself runs after the calling turn ends, and it may
still refuse at that point; those outcomes are logged only, because there is no caller left to
answer. Look for `lead-rollover:` lines in the daemon log.
- **A lead can only ever roll itself.** The tool has no terminal, session or lead parameter of any
kind. The pane comes from the caller's own connection. This daemon can hold more than one
labelled lead tab, and an earlier draft that looked the pane up in a single-slot registry let one
lead clear another lead's pane.
- **`/clear` must never go through `Injector`.** It produces no turn boundary, so the Injector's
turn never completes and every later message to that pane queues behind it for ever — the pane is
wedged. `LeadRollover` calls `agents.send(...)` directly, exactly like
`ClaudeCodeLauncher#clearContext`. Measured, not assumed.
- **The settle waits require `IDLE` or `DONE`, not `injectable()`.** `injectable()` also accepts
`BLOCKED`, which is a live turn paused on an approval prompt. Treating that as settled sent
`/clear` into an open prompt mid-turn.
- **One acceptance criterion is still unproven.** Nothing has yet demonstrated end-to-end that a
freshly cleared pane picks up the bootstrap prompt. `/clear` arriving as a real slash command was
proven live; what follows it was not, because no shipped tool or endpoint does a non-`Injector`
send on a throwaway pane, and the feature only ever acts on the caller's own lead pane. If the
bootstrap does not land, the lead's context is gone and no fresh session starts — but the handover
file is written before `confirm` is even accepted, so the operator can always recover by starting
a session by hand, which is exactly the manual path this feature replaces.
fleetd #480 (PRs #483, #484, #485). Related: #486.
+1
@@ -447,6 +447,7 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
| Tear down a member | `fleet_stop{paneId}` |
| Replace your OWN lead session when its context is full | `fleet_handover{action:"open", reason?}` → write the handover file it names → `fleet_handover{action:"confirm", token, operatorConfirmed}`. Primary-only. **In that order**: the file must be modified *after* `open`, or `confirm` refuses it as stale. There is no terminal parameter — the pane is always your own, so you can never roll another lead. `{action:"cancel", token}` drops a pending request |
### Lead ↔ lead — coordinate, never delegate