A context reset changes the peer's session id, and fleetd keeps reporting the dead one — so resumeSessionId would resume the conversation clearAfterTurn just discarded #482

Open
opened 2026-09-11 01:22:51 +02:00 by ltms · 0 comments
Owner

What I measured

Found while probing /clear for #480. Measured 2026-09-11 on the Mac, against 7b97aae.

One member, term_65b28fe1aad1a1ad, spawned on the sonnet profile. I sent it /clear. Claude
Code does not clear a conversation in place — it ends the session and starts a new one with a
new id.
Both transcripts are still on disk and I re-checked them in the turn that wrote this:

7ddac757-8fd9-4ad6-bb29-51caecfd6c2a.jsonl   born 06:04:49   last written 06:05:04
d8a2face-4898-4029-9e16-333dad14b34c.jsonl   born 06:05:04   last written 06:08:32

The two meet exactly at 06:05:04, and grep -l 'command-name>/clear' matches only
d8a2face — the new session carries the /clear envelope as its opening entry.

fleetd went on reporting the dead id. fleet_list reported
"agentSessionId":"7ddac757-…" for that member after the clear, and two ticket failures for it
still name agentSessionId=7ddac757-…. The live session was d8a2face.

Why this matters

fleet_list's agentSessionId is documented as "the id to pass as fleet_spawn's
resumeSessionId to relaunch onto that same conversation"
. session/SessionManager.java:270
(requireResumeCapability) gates resume on the profile declaring SESSION_RESUME, and
member/ClaudeCodeLauncher.java:319 (applySessionIdentity) turns the id into -r <id>.

So after a context reset, a resume onto the reported id relaunches the conversation the reset
deliberately threw away
. The failure is quiet: the resume succeeds, the flag is accepted, and
the member comes back holding exactly the context an operator turned on clearAfterTurn to get
rid of. Nothing reports a mismatch, because nothing re-reads the id.

This is the same shape as
#427 and the "a status field must read the source
the behaviour reads" rule: the behaviour moved and the field did not follow.

Reachability — honest scope

  • clearAfterTurn defaults to false (config/FleetConfig.java:868, asserted at
    config/FleetConfigTest.java:268: "context clearing is opt-in"). So on a default config this
    is inert today.
  • I did not measure the clearAfterTurn path itself. I measured a /clear I sent by hand.
    Both routes end in the same herdr agent.prompt call
    (member/ClaudeCodeLauncher.java:972), so I expect the same result — but expecting is not
    measuring, and I am not claiming it. Whoever takes this should drive the real
    clearAfterTurn path and check agentSessionId before and after.
  • The /clear I sent went through Injector, which also wedged the pane (see #480). That is a
    separate defect and does not affect this one: the id was already stale the moment the new
    session was born.

Suggested fix direction, not a decision

Two options, and I have not picked one:

  1. Re-read the id after a reset. clearContext knows it just reset the peer; it could clear
    or refresh the recorded agentSessionId.
  2. Report it as unknown rather than stale. Absent is safer than wrong here: fleet_list
    already documents agentSessionId as legitimately absent for a member fleetd cannot
    re-identify (fleetd #249), and requireResumeCapability refuses a resume rather than
    silently starting fresh. A wrong id is the one outcome that gate cannot catch.

Option 2 is cheaper and fails closed. Option 1 is better if a post-reset resume is ever wanted.

Note

Not urgent, because the feature that triggers it is off by default. Filed so it is not
rediscovered: it cost me an hour of confusion while probing #480, and the evidence is only
visible if you read the transcript files rather than the daemon's own report.

## What I measured Found while probing `/clear` for #480. Measured 2026-09-11 on the Mac, against `7b97aae`. One member, `term_65b28fe1aad1a1ad`, spawned on the `sonnet` profile. I sent it `/clear`. Claude Code does not clear a conversation in place — **it ends the session and starts a new one with a new id.** Both transcripts are still on disk and I re-checked them in the turn that wrote this: ``` 7ddac757-8fd9-4ad6-bb29-51caecfd6c2a.jsonl born 06:04:49 last written 06:05:04 d8a2face-4898-4029-9e16-333dad14b34c.jsonl born 06:05:04 last written 06:08:32 ``` The two meet exactly at 06:05:04, and `grep -l 'command-name>/clear'` matches **only** `d8a2face` — the new session carries the `/clear` envelope as its opening entry. **fleetd went on reporting the dead id.** `fleet_list` reported `"agentSessionId":"7ddac757-…"` for that member after the clear, and two ticket failures for it still name `agentSessionId=7ddac757-…`. The live session was `d8a2face`. ## Why this matters `fleet_list`'s `agentSessionId` is documented as *"the id to pass as `fleet_spawn`'s `resumeSessionId` to relaunch onto that same conversation"*. `session/SessionManager.java:270` (`requireResumeCapability`) gates resume on the profile declaring `SESSION_RESUME`, and `member/ClaudeCodeLauncher.java:319` (`applySessionIdentity`) turns the id into `-r <id>`. So after a context reset, a resume onto the reported id relaunches **the conversation the reset deliberately threw away**. The failure is quiet: the resume succeeds, the flag is accepted, and the member comes back holding exactly the context an operator turned on `clearAfterTurn` to get rid of. Nothing reports a mismatch, because nothing re-reads the id. This is the same shape as [#427](https://git.ltms.dev/fleet/fleetd/issues/427) and the "a status field must read the source the behaviour reads" rule: the behaviour moved and the field did not follow. ## Reachability — honest scope - **`clearAfterTurn` defaults to `false`** (`config/FleetConfig.java:868`, asserted at `config/FleetConfigTest.java:268`: *"context clearing is opt-in"*). So on a default config this is inert today. - **I did not measure the `clearAfterTurn` path itself.** I measured a `/clear` I sent by hand. Both routes end in the same herdr `agent.prompt` call (`member/ClaudeCodeLauncher.java:972`), so I expect the same result — but expecting is not measuring, and I am not claiming it. Whoever takes this should drive the real `clearAfterTurn` path and check `agentSessionId` before and after. - The `/clear` I sent went through `Injector`, which also wedged the pane (see #480). That is a separate defect and does not affect this one: the id was already stale the moment the new session was born. ## Suggested fix direction, not a decision Two options, and I have not picked one: 1. **Re-read the id after a reset.** `clearContext` knows it just reset the peer; it could clear or refresh the recorded `agentSessionId`. 2. **Report it as unknown rather than stale.** Absent is safer than wrong here: `fleet_list` already documents `agentSessionId` as legitimately absent for a member fleetd cannot re-identify (fleetd #249), and `requireResumeCapability` refuses a resume rather than silently starting fresh. A wrong id is the one outcome that gate cannot catch. Option 2 is cheaper and fails closed. Option 1 is better if a post-reset resume is ever wanted. ## Note Not urgent, because the feature that triggers it is off by default. Filed so it is not rediscovered: it cost me an hour of confusion while probing #480, and the evidence is only visible if you read the transcript files rather than the daemon's own report.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#482