Files
fleetd/.claude/skills/handover/SKILL.md
T
Dai Ha 7d9a807243
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 53s
CI / build (push) Failing after 2m33s
handover skill: the rollover bootstrap is proven, and requireOperatorConfirm is per-host
Two bullets in the handover skill were telling every outgoing lead something
that is no longer true.

1. The skill said the bootstrap prompt "has never yet landed, and the fix is
   unproven (fleetd #489)", and told the lead to warn the operator it may fail.
   Measured today from fleetd/fleetd.out:

     grep -c "lead-rollover: rolled" -> 4
     grep -c "lead-rollover:"        -> 16   (positive control)
     grep -c "Unknown command"       -> 0

   Three of the four rolls ran on 2026-09-22 (10:01:43, 10:38:28, 11:15:47).
   Each cleared the old lead and bootstrapped a fresh one against the handover
   file. The old "Unknown command: /clearFresh" failure does not appear at all.
   The paragraph now carries the measured result and the three re-measure
   commands, including the control line, because a broken grep pattern returns
   a clean 0 that reads like good news.

   It also records that the "/clear was never observed as WORKING ... releasing
   rather than wedging the roll" WARN accompanies every successful roll. That
   is the safe branch, not a failure, and it was being misread as one.

2. The skill said requireOperatorConfirm "defaults to true and this is the only
   thing standing between a judgement call and a wiped session", which reads as
   if asking is always required. The default is still true
   (FleetConfig.java:1426), but this host set it to false on 2026-09-22 on the
   operator's explicit grant. The bullet now says to read the live value rather
   than assume, and notes the key is deferred, not hot.

   It also warns that until fleetd #621 merges, LeadHeartbeatLoop.contextNotice()
   still hardcodes "ask the operator" and takes no config, so the nudge text and
   the config disagree. Trust the config. That warning names the ticket that
   removes it.

Documentation only. No code or test changes.
2026-09-22 11:24:03 +07:00

13 KiB
Raw Blame History

name, description
name description
handover Procedure for an outgoing lead to write the handover file that a fresh lead session inherits. Load this when your context is filling up and you are about to be replaced, whether you hand off by hand or fleetd does it for you. The file is the new lead's only inheritance — follow it exactly.

Handover — write the file the next lead depends on

A lead session fills up its context and has to be replaced by a fresh one. The outgoing lead writes a handover file, and the new session reads that file and carries on.

There are two ways to hand off, and the file is the same either way.

  • By hand. You write the file, then tell the operator where it is. The operator starts the new session and points it at the file. This always works.
  • With fleet_handover (fleetd #480, merged 2026-09-11). You ask fleetd to do the swap: it checks the file, clears your pane, and tells the fresh session to read it. This needs leadRollover: in fleetd.yaml; without it every action answers a clean refusal naming NOT_CONFIGURED, and you fall back to the manual path. Section 11 below is the procedure.

Nothing else in this skill changes between the two. Only who performs the swap changes.

The new lead's only inheritance is that file. It does not see your conversation, your plan, or your screen. If the file is thin or wrong, the new lead re-derives what you already knew, and that wastes hours. Writing a good handover file is real work. It is not paperwork you rush through at the end of a session.

This skill is the procedure for writing it. Every rule below earned its place because a past handover got it wrong.

1. Confirm you are the right session to write this

Run fleet_whoami first. It must answer primary. Only a primary (lead) session writes a handover file. A worker's job ends with its own pull request, not a fleet-wide handoff.

2. Every number needs a command, run in this turn

A number is a claim: a count, a commit hash, a process id, a percentage, a queue depth. Before you write one, run the command that produces it — now, in this turn, against the live state.

Never take a number from:

  • earlier in your own conversation — the state has moved since then,
  • a peer lead's report — that is their measurement, not yours,
  • your own memory of an earlier session.

Put the command, or its real output, next to the number. That lets the next lead re-run it and check it still matches. If you cannot measure something yourself, say so instead of guessing: "the fleet01 lead reports 91 commits behind; I have not checked this myself."

3. Say what you measured and what you did not

Mark every claim as one of two things:

  • "I checked this myself, in the code or on this host, at <time>."
  • "I did not check this myself; <who> reported it."

Never present someone else's measurement as your own. This matters most for cross-host claims — a peer lead's daemon, a worker's report, or something the operator said earlier that you cannot re-verify from here.

4. Record open decisions, and who owns them

List three things:

  • what the operator actually asked for, in their own words where you have them,
  • what is still unanswered,
  • any question you decided yourself instead of asking, with your reason.

Write the decision so it cannot be mistaken for the operator's instruction. Say plainly: "the operator never answered X; I decided Y, because Z." Without this, the next lead either silently reopens a closed question or assumes the operator chose something they never did.

5. Record live hazards

List anything that will break if the next lead does the obvious thing next. This includes:

  • unpushed commits or unmerged branches,
  • a build, a spawn, or a redeploy still running,
  • code merged to main but not yet redeployed to the live daemon,
  • any trap that looks safe and is not — say what goes wrong and why, not only that something is "tricky."

6. Record what is explicitly not owed

List work that is finished, and work that another party has said they do not want touched. Name who said so and when. Without this line, the next lead re-does closed work or reopens a question a peer already declined to revisit.

7. Open the file with three re-measurement commands

The file's own first section must give the next lead three concrete commands to run before acting on anything else in the file:

  1. confirm role — for example fleet_whoami,
  2. confirm the state of the working tree — for example git status and git rev-list origin/main..HEAD,
  3. read the live fleet — for example fleet_list.

Record what each command answered when you wrote the file, and tell the reader to run it again rather than trust your answer. The point of this section is that the reader checks live state before acting on any claim in the rest of the file, including yours.

8. Stamp the file with time and commit

Near the top of the file, write:

  • the date and time you wrote it,
  • the commit the tree was on (git rev-parse HEAD),
  • whether the tree was clean (git status).

Without this, nobody can tell how old the file is, or which code it describes.

9. State plainly that the file goes stale fast

Say near the top: re-measure anything you act on. The file goes stale the moment anyone merges a branch, spawns a member, or restarts the daemon. Everything in the file is a snapshot of one moment, not a live fact.

10. What to leave out

Do not include:

  • narration of how the session felt, or how hard something was,
  • anything the repo already records — code structure, git history, or a rule already written in CLAUDE.md. Point at it instead of repeating it,
  • advice that is only true for the session that is ending — a half-open terminal, a local variable, a train of thought with no state behind it.

A handover file is a record of state and decisions. It is not a diary.

11. Using fleet_handover (only if leadRollover: is configured)

Run the three steps in this order. The order is not a style choice — the wrong order is refused.

  1. fleet_handover{action: "open", reason: "<why now>"}. It returns a token and the handoverPath you must write to. Nothing has happened to your pane yet.

    Write to exactly that path, and do not resolve it yourself. It is always absolute, even when the operator configured a relative handoverPath: fleetd resolves a relative one against your own workspace before it hands it to you. The daemon and your pane can run in different directories, so a path you resolve yourself can point at a different file from the one the daemon will check.

    Check that the path is ignored by git before you write to it (#491). A relative handoverPath resolves inside YOUR workspace, which is usually a repository — and usually not the fleetd one, so an ignore rule added to fleetd does not protect it. Run grep -n handover <your workspace>/.gitignore. No output means the file you are about to write will show up as untracked content in that repo. The file is a snapshot of live state and must never be committed, so tell the operator rather than committing it or silently editing their .gitignore.

  2. Write the handover file at that path, following sections 1–10 above.

  3. Ask the operator, then fleet_handover{action: "confirm", token, operatorConfirmed: true}.

Why that order: confirm refuses with HANDOVER_STALE unless the file was modified after the open request. That check stops a leftover file from an earlier session being accepted as this one's handover. So writing the file first and then calling open — the obvious order — always fails.

{action: "cancel", token} drops a pending request without rolling.

Things that will surprise you:

  • accepted does not mean your pane has been cleared. It means every gate passed and the roll is scheduled to run once your current turn ends. Say your goodbye in the same turn — you will not get another one.

  • There is no terminal or session parameter, on purpose. The pane is always your own, resolved from your connection, so you can only ever roll yourself.

  • operatorConfirmed is your report of what a human told you. Do not pass true because you are confident. Ask, wait for the answer, then pass what they said.

  • Whether you must ask at all depends on leadRollover.requireOperatorConfirm. Check it; do not assume. The default is true (FleetConfig.java:1426), and then confirm refuses unless you also pass operatorConfirmed: true. This host set it to false on 2026-09-22, on the operator's explicit grant, because they do not want to approve routine context rolls. Where it is false, the three handover-file checks are the whole gate: the file must exist, be fresher than maxDocAgeSeconds, and have been modified after the open request.

    Read the live value rather than trusting this line:

    grep -A1 'requireOperatorConfirm' fleetd/fleetd.yaml
    

    No match means the key is unset, so the default true applies and you must ask. The key is deferred, not hot — it is read once at boot, so an edit does nothing until the daemon is redeployed.

    Until fleetd #621 merges, the nudge text will tell you to ask the operator even where the daemon no longer requires it. LeadHeartbeatLoop.contextNotice() hardcodes "ask the operator" and takes no config, so it cannot know. Trust the config value over the nudge text. Once #621 is merged and deployed, the nudge matches the config and this warning can be deleted.

  • The roll can still refuse after confirm returns, and by then there is no caller to tell. Those outcomes are logged only, as lead-rollover: lines in the daemon log.

  • The bootstrap prompt works end to end. Measured 2026-09-22. This used to say the fix was unproven (fleetd #489) and told you to expect a failure. That is no longer true. The daemon log now holds four lead-rollover: rolled lines, and three of them ran on 2026-09-22 at 10:01:43, 10:38:28 and 11:15:47. Each one cleared the old lead and started a fresh session against the handover file, with the configured bootstrapText arriving as its first message. No context was lost. The old Unknown command: /clearFresh failure from 2026-09-12 does not appear in the log at all. Re-measure both numbers with:

    grep -c "lead-rollover: rolled" fleetd/fleetd.out   # successful rolls
    grep -c "lead-rollover:" fleetd/fleetd.out          # positive control: must be larger
    grep -c "Unknown command" fleetd/fleetd.out         # the old failure: expect 0
    

    Run the control line too. A broken pattern returns a clean 0 that reads exactly like good news. If the first number stops growing across rolls, or Unknown command returns anything above 0, the bootstrap has regressed and this paragraph is stale again.

    You still write the file before you confirm, and never the other way round. That order is not about the bootstrap being unreliable. It is what the daemon checks: the handover file must have been modified after the open request, or confirm refuses it as stale.

  • One warning in the log is normal and is not a failure. Every one of the three rolls above also logged /clear on term_… was never observed as WORKING after 8 consecutive IDLE/DONE polls — releasing rather than wedging the roll. The daemon could not see the pane go WORKING after /clear, so it released instead of hanging. The roll then succeeded anyway. That is the safe branch behaving correctly. Do not report it as a broken roll.

Writing style

Write in plain English. Use everyday words, one idea per sentence, and active voice. Keep every class, method, file, flag, and config key exactly as it appears in the code — replacing a precise term with a vague one makes the sentence wrong, not simpler. Explain an abbreviation the first time you use it.

If a diagram genuinely helps, put it in the .md file as a fenced ```mermaid block with no hardcoded colors, so it stays readable on light and dark backgrounds. Quote any label that has brackets, colons, or slashes.

Template

# Handover — <fleet name> lead session, <date and time>

Written at commit `<output of git rev-parse HEAD>`. Tree was <clean, or dirty: `<git status
summary>`>. Re-measure anything you act on — this file goes stale the moment anyone merges,
spawns, or restarts.

## 0. Do these three things first

1. Confirm your role: `fleet_whoami` — must answer `primary`. (Answered `<result>` at `<time>`.)
2. Confirm tree state: `git status`, `git rev-list origin/main..HEAD`. (`<result>` at `<time>`.)
3. Read the live fleet: `fleet_list`. (`<result>` at `<time>`.)

## 1. What the operator asked for

<the live instructions, in their words where you have them; what is still open; any decision
you made yourself, and why>

## 2. Open decisions, and who owns them

<one line per decision: who owns it, what is unanswered>

## 3. Live hazards

<one entry per hazard: what breaks, and why, if the next lead does the obvious thing>

## 4. What is not owed

<finished work, and work another party has declined; name who said so and when>