diff --git a/.claude/skills/redeploy-fleetd/SKILL.md b/.claude/skills/redeploy-fleetd/SKILL.md new file mode 100644 index 0000000..18b0e17 --- /dev/null +++ b/.claude/skills/redeploy-fleetd/SKILL.md @@ -0,0 +1,72 @@ +--- +name: redeploy-fleetd +description: Rebuild and restart the live fleetd daemon after a merge (lead / primary only). Load this before redeploying — it holds the script, the drain step, the permission grant, and the five checks that have each gone wrong here before. Workers must never do this. +--- + +### Redeploying the daemon — the lead may do this (primary only) + +**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a +feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped" +about code the live daemon has never loaded is a false report. The lead **may and should** redeploy +rather than hand the job back to the operator. + +Workers must never do this. A worker has no business restarting the daemon it is talking through, +and stopping it kills the worker's own channel mid-turn. + +**Use the script — do not hand-roll the steps.** + +```bash +scripts/redeploy-fleetd.sh --check # report state, change nothing +scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify +scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked) +scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk +``` + +`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only +when you just built and nothing changed since. It gives up the protection in the next paragraph: no +build runs, so a stale or missing jar is not caught early. The script still checks the file is there +and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is +old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing +degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification +time, so you can see for yourself whether the jar is missing or older than the code you mean to ship. + +It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the +old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a +line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check` +first — it is read-only and reports whether the forge token resolves, which nothing else tells you. + +The script encodes the five things below, each of which has gone wrong here before. Read them anyway: +if the script is unavailable or a step fails, this is what it was protecting you from. + +1. **Login shell, or workers silently lose their forge token.** The daemon inherits + `WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from + `${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the + daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing + logs this at startup — the script's `--check` is the only thing that reports it, and it checks + whether the name resolves without ever printing the value. +2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything + you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and + rendezvous, and a member's report is not recoverable once its ticket is gone. +3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do + it. The startup log names which keys it accepted and which it deferred — read those lines rather + than assuming. +4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The + lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is + demoted to worker, which refuses every orchestration call. +5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of + `fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from + the outside. + +**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier +refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that: +it is one auditable command, so the operator allow-lists it once instead of approving a stop and a +start every time. The rule lives in the operator's Claude Code settings: + +```json +{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } } +``` + +Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by +running the stop and start as separate commands — that is exactly the approval the script replaced. +Say what you were going to run and why, and let the operator decide. + diff --git a/CLAUDE.md b/CLAUDE.md index 1e99cbc..d8d9487 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -257,8 +257,9 @@ must obey belongs in the charter, not here. already cost three workers' turns: each wrote a good report to its terminal and ended the turn with no `fleet_reply`, and the scrape returned the tail of the brief instead. - **Primary-side skills** (not delegation playbooks — a worker cannot use them): - `port-to-opencode` (make an OpenCode session a participant in this workspace) and - `fleets-status` (report every fleet that shares one LavinMQ instance). + `port-to-opencode` (make an OpenCode session a participant in this workspace), + `fleets-status` (report every fleet that shares one LavinMQ instance) and + `redeploy-fleetd` (rebuild and restart the live daemon after a merge). - **This repo is also a Claude Code marketplace, and ships a plugin.** `.claude-plugin/marketplace.json` points at `plugin/`, which carries the MCP mount and the `setup` skill (`/claude-bridge:setup` — make any project bridge-ready). It was added in CB-527 and then went @@ -292,60 +293,14 @@ must obey belongs in the charter, not here. ### Redeploying the daemon — the lead may do this (primary only) **A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a -feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped" -about code the live daemon has never loaded is a false report. The lead **may and should** redeploy -rather than hand the job back to the operator. +feature merged to `main` does nothing until the daemon is rebuilt and restarted. The lead **may and +should** redeploy rather than hand the job back to the operator. **Workers must never do this** — a +worker has no business restarting the daemon it is talking through, and stopping it kills the +worker's own channel mid-turn. -Workers must never do this. A worker has no business restarting the daemon it is talking through, -and stopping it kills the worker's own channel mid-turn. - -**Use the script — do not hand-roll the steps.** - -```bash -scripts/redeploy-fleetd.sh --check # report state, change nothing -scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify -scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked) -``` - -It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the -old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a -line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check` -first — it is read-only and reports whether the forge token resolves, which nothing else tells you. - -The script encodes the five things below, each of which has gone wrong here before. Read them anyway: -if the script is unavailable or a step fails, this is what it was protecting you from. - -1. **Login shell, or workers silently lose their forge token.** The daemon inherits - `WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from - `${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the - daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing - logs this at startup — the script's `--check` is the only thing that reports it, and it checks - whether the name resolves without ever printing the value. -2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything - you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and - rendezvous, and a member's report is not recoverable once its ticket is gone. -3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do - it. The startup log names which keys it accepted and which it deferred — read those lines rather - than assuming. -4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The - lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is - demoted to worker, which refuses every orchestration call. -5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of - `fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from - the outside. - -**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier -refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that: -it is one auditable command, so the operator allow-lists it once instead of approving a stop and a -start every time. The rule lives in the operator's Claude Code settings: - -```json -{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } } -``` - -Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by -running the stop and start as separate commands — that is exactly the approval the script replaced. -Say what you were going to run and why, and let the operator decide. +**Load the `redeploy-fleetd` skill before you redeploy.** It holds `scripts/redeploy-fleetd.sh` +and its flags, the drain step, the operator's permission grant, and the five checks that have each +gone wrong here before. Do not hand-roll the steps from memory. ### The prompt is part of the product — update it with the code (mandatory)