docs: move the redeploy procedure out of CLAUDE.md into a skill
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m34s

CLAUDE.md loads into every session. The daemon-redeploy procedure is
needed only when someone redeploys, so it paid for context it did not
use: 58 lines, about 918 est. tokens, every session.

The procedure now lives in .claude/skills/redeploy-fleetd/SKILL.md,
which loads only when invoked. The moved text is byte-identical to
what was removed, plus one new paragraph documenting --no-build
(scripts/redeploy-fleetd.sh:34,62,266) — the script and the wiki
already had that flag, CLAUDE.md never did.

CLAUDE.md keeps an 11-line pointer, because two rules must stay
resident: a merge is not a deployment, and workers must never
redeploy. A session learns it needs the procedure before it needs
the skill, then the next line names the skill.

Also updates the addendum's primary-side skill list, as this file's
own rule for .claude/skills/** changes requires.

CLAUDE.md: 34,442 -> 31,591 chars.
Canonical block untouched — the wiki sync check still prints True.
This commit is contained in:
Dai Ha
2026-09-11 05:48:06 +07:00
parent 49a404ddf3
commit 7b97aae85b
2 changed files with 82 additions and 55 deletions
+72
View File
@@ -0,0 +1,72 @@
---
name: redeploy-fleetd
description: Rebuild and restart the live fleetd daemon after a merge (lead / primary only). Load this before redeploying — it holds the script, the drain step, the permission grant, and the five checks that have each gone wrong here before. Workers must never do this.
---
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-fleetd.sh --check # report state, change nothing
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
```
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
+10 -55
View File
@@ -257,8 +257,9 @@ must obey belongs in the charter, not here.
already cost three workers' turns: each wrote a good report to its terminal and ended the turn
with no `fleet_reply`, and the scrape returned the tail of the brief instead.
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
`port-to-opencode` (make an OpenCode session a participant in this workspace) and
`fleets-status` (report every fleet that shares one LavinMQ instance).
`port-to-opencode` (make an OpenCode session a participant in this workspace),
`fleets-status` (report every fleet that shares one LavinMQ instance) and
`redeploy-fleetd` (rebuild and restart the live daemon after a merge).
- **This repo is also a Claude Code marketplace, and ships a plugin.** `.claude-plugin/marketplace.json`
points at `plugin/`, which carries the MCP mount and the `setup` skill
(`/claude-bridge:setup` — make any project bridge-ready). It was added in CB-527 and then went
@@ -292,60 +293,14 @@ must obey belongs in the charter, not here.
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
feature merged to `main` does nothing until the daemon is rebuilt and restarted. The lead **may and
should** redeploy rather than hand the job back to the operator. **Workers must never do this** — a
worker has no business restarting the daemon it is talking through, and stopping it kills the
worker's own channel mid-turn.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-fleetd.sh --check # report state, change nothing
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
```
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
**Load the `redeploy-fleetd` skill before you redeploy.** It holds `scripts/redeploy-fleetd.sh`
and its flags, the drain step, the operator's permission grant, and the five checks that have each
gone wrong here before. Do not hand-roll the steps from memory.
### The prompt is part of the product — update it with the code (mandatory)