From 7d9a80724392575594bb823c06224583dc2f3a45 Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Tue, 22 Sep 2026 11:24:03 +0700 Subject: [PATCH] handover skill: the rollover bootstrap is proven, and requireOperatorConfirm is per-host Two bullets in the handover skill were telling every outgoing lead something that is no longer true. 1. The skill said the bootstrap prompt "has never yet landed, and the fix is unproven (fleetd #489)", and told the lead to warn the operator it may fail. Measured today from fleetd/fleetd.out: grep -c "lead-rollover: rolled" -> 4 grep -c "lead-rollover:" -> 16 (positive control) grep -c "Unknown command" -> 0 Three of the four rolls ran on 2026-09-22 (10:01:43, 10:38:28, 11:15:47). Each cleared the old lead and bootstrapped a fresh one against the handover file. The old "Unknown command: /clearFresh" failure does not appear at all. The paragraph now carries the measured result and the three re-measure commands, including the control line, because a broken grep pattern returns a clean 0 that reads like good news. It also records that the "/clear was never observed as WORKING ... releasing rather than wedging the roll" WARN accompanies every successful roll. That is the safe branch, not a failure, and it was being misread as one. 2. The skill said requireOperatorConfirm "defaults to true and this is the only thing standing between a judgement call and a wiped session", which reads as if asking is always required. The default is still true (FleetConfig.java:1426), but this host set it to false on 2026-09-22 on the operator's explicit grant. The bullet now says to read the live value rather than assume, and notes the key is deferred, not hot. It also warns that until fleetd #621 merges, LeadHeartbeatLoop.contextNotice() still hardcodes "ask the operator" and takes no config, so the nudge text and the config disagree. Trust the config. That warning names the ticket that removes it. Documentation only. No code or test changes. --- .claude/skills/handover/SKILL.md | 62 ++++++++++++++++++++++++++------ 1 file changed, 51 insertions(+), 11 deletions(-) diff --git a/.claude/skills/handover/SKILL.md b/.claude/skills/handover/SKILL.md index 1b63437..8422d0e 100644 --- a/.claude/skills/handover/SKILL.md +++ b/.claude/skills/handover/SKILL.md @@ -167,19 +167,59 @@ fails. - **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved from your connection, so you can only ever roll yourself. - **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you - are confident. Ask, wait for the answer, then pass what they said. `requireOperatorConfirm` - defaults to `true` and this is the only thing standing between a judgement call and a wiped - session. + are confident. Ask, wait for the answer, then pass what they said. + +- **Whether you must ask at all depends on `leadRollover.requireOperatorConfirm`. Check it; do not + assume.** The default is `true` (`FleetConfig.java:1426`), and then `confirm` refuses unless you + also pass `operatorConfirmed: true`. **This host set it to `false` on 2026-09-22**, on the + operator's explicit grant, because they do not want to approve routine context rolls. Where it is + `false`, the three handover-file checks are the whole gate: the file must exist, be fresher than + `maxDocAgeSeconds`, and have been modified after the open request. + + Read the live value rather than trusting this line: + + ```bash + grep -A1 'requireOperatorConfirm' fleetd/fleetd.yaml + ``` + + No match means the key is unset, so the default `true` applies and you must ask. The key is + **deferred, not hot** — it is read once at boot, so an edit does nothing until the daemon is + redeployed. + + **Until fleetd #621 merges, the nudge text will tell you to ask the operator even where the + daemon no longer requires it.** `LeadHeartbeatLoop.contextNotice()` hardcodes "ask the operator" + and takes no config, so it cannot know. Trust the config value over the nudge text. Once #621 is + merged and deployed, the nudge matches the config and this warning can be deleted. + - **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell. Those outcomes are logged only, as `lead-rollover:` lines in the daemon log. -- **The bootstrap prompt has never yet landed, and the fix is unproven (fleetd #489).** The first - real rollover, on 2026-09-12, joined `/clear` and the bootstrap text into one line and Claude Code - refused it as `Unknown command: /clearFresh`. The pane was never cleared and no context was lost, - so the failure was safe — the roll simply did nothing. PR #490 fixed the cause and is deployed, - but no roll has bootstrapped a fresh session end to end yet. **Assume it may still fail, and tell - the operator so before you confirm.** The recovery is the same either way: the file is already - written, so the operator starts a session and points it at the file. That is why you write the - file before you confirm, and never the other way round. +- **The bootstrap prompt works end to end. Measured 2026-09-22.** This used to say the fix was + unproven (fleetd #489) and told you to expect a failure. That is no longer true. The daemon log + now holds four `lead-rollover: rolled` lines, and three of them ran on 2026-09-22 at 10:01:43, + 10:38:28 and 11:15:47. Each one cleared the old lead and started a fresh session against the + handover file, with the configured `bootstrapText` arriving as its first message. No context was + lost. The old `Unknown command: /clearFresh` failure from 2026-09-12 does not appear in the log + at all. Re-measure both numbers with: + + ```bash + grep -c "lead-rollover: rolled" fleetd/fleetd.out # successful rolls + grep -c "lead-rollover:" fleetd/fleetd.out # positive control: must be larger + grep -c "Unknown command" fleetd/fleetd.out # the old failure: expect 0 + ``` + + Run the control line too. A broken pattern returns a clean `0` that reads exactly like good news. + If the first number stops growing across rolls, or `Unknown command` returns anything above 0, + the bootstrap has regressed and this paragraph is stale again. + + **You still write the file before you confirm, and never the other way round.** That order is not + about the bootstrap being unreliable. It is what the daemon checks: the handover file must have + been modified *after* the open request, or `confirm` refuses it as stale. + +- **One warning in the log is normal and is not a failure.** Every one of the three rolls above also + logged `/clear on term_… was never observed as WORKING after 8 consecutive IDLE/DONE polls — + releasing rather than wedging the roll`. The daemon could not see the pane go WORKING after + `/clear`, so it released instead of hanging. The roll then succeeded anyway. That is the safe + branch behaving correctly. Do not report it as a broken roll. ## Writing style