The lead may redeploy the daemon — write down the procedure and its traps

A merge is not a deployment: the running bridged holds the jar it started
with, so merged code does nothing until the daemon is rebuilt and restarted.
Calling that work shipped is a false report. This makes the redeploy the
lead's job rather than something handed back to the operator, and records the
five things that have gone wrong doing it here — chiefly that starting the
daemon from a non-login shell empties WORKER_GITEA_TOKEN, which nothing logs
and which only surfaces later as workers that cannot open a PR.

Goes in the project addendum, not the canonical block; the block is unchanged
and still byte-identical with the wiki template.
This commit is contained in:
Dai Ha
2026-08-15 13:54:38 +02:00
parent 3a10f6ad17
commit 23af5dfdff
+45
View File
@@ -189,6 +189,51 @@ charter, not here.
are diagrammed in `docs/MCP-Contract.md` §6, kept out of this file because it loads into every
session's context.
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `bridged` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
```bash
# 1. Build. Never pipe mvn — a pipe's exit code hides BUILD FAILURE.
mvn -f bridged/pom.xml clean install
# 2. Find and stop the running daemon.
pgrep -f 'bridged/target/bridged.jar'
kill <pid>
# 3. Start it again FROM A LOGIN SHELL, detached, with cwd = bridged/.
(cd bridged && zsh -lc 'nohup java -jar target/bridged.jar >> bridged.out 2>&1 &')
```
Five things to get right, each of which has gone wrong here before:
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — that gap is why the rule has to be remembered here.
2. **Drain live members first.** `bridge_list`, then `bridge_stop` each member, and collect anything
you still want with `bridge_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `bridge_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `bridged listening` line at the end of
`bridged/bridged.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
If a step is refused by the command classifier, do not try to route around it. Say what you were
going to run and why, and ask the operator to run it with `!`.
### The prompt is part of the product — update it with the code (mandatory)
This repo *is* the bridge, so the canonical block above is not documentation about someone else's