CB-594: the launchd unit cannot carry the secrets the fleet needs, so supervision and a working fleet are mutually exclusive today #80

Closed
opened 2026-08-16 16:48:10 +02:00 by ltms · 2 comments
Owner

Found while checking release-1 close-out state on 2026-08-16. The daemon was down — it had died some time after 09:46 with no shutdown line — and nothing restarted it.

The state on the host

CB-504 shipped deploy/dev.ltms.bridged.plist. That file has RunAtLoad and a KeepAlive backstop, so it would have caught this. It is not installed:

ls ~/Library/LaunchAgents/ | grep bridged   ->  nothing

So the supervision unit exists in git and does not exist on the machine. That alone is a one-line install. The reason it was never installed is the real ticket.

The contradiction

CLAUDE.md §Redeploying the daemon, rule 1:

Login shell, or workers silently lose their forge token. The daemon inherits WORKER_GITEA_TOKEN from the shell that starts it, and that comes from ${SHARED_ENV}/tools/secrets.sh. Start it from a non-login shell and the variable is empty, the daemon starts fine, and the failure appears much later as workers that cannot open a PR.

deploy/dev.ltms.bridged.plist, lines 52-56:

Worker/API tokens are NOT set here: this file is committed. Export them from a private launchd override or a wrapper script.

launchd does not source a login shell. The plist even says so itself about PATH (line 47). So a daemon started by this unit gets:

  • no WORKER_GITEA_TOKEN → every worker silently fails to open a PR;
  • no AI_GATEWAY_TOKEN → every gateway-backed profile (local, gx) cannot authenticate.

Nothing logs either at startup. scripts/redeploy-bridged.sh --check is the only thing that reports it, and it only checks the login shell, not the environment the daemon actually got.

So today the operator picks one: a supervised daemon with a broken fleet, or a working fleet with no supervision. Neither is a release state. The "private override or wrapper script" the plist points at was never written.

Why this is worth a ticket rather than a one-liner

Three things have to agree, and right now none of them knows about the others:

  1. the plist (a CHANGEME template, never instantiated);
  2. scripts/redeploy-bridged.sh, which starts the daemon by hand and would fight a KeepAlive unit;
  3. ${SHARED_ENV}/tools/secrets.sh, which is the operator's file and must not be edited.

Acceptance criteria

  1. A supervised daemon starts with WORKER_GITEA_TOKEN and AI_GATEWAY_TOKEN resolved, without either secret appearing in a committed file. A wrapper that execs through a login shell is the obvious shape; pick one and write down why.
  2. bridged logs at startup which required secrets resolved and which did not — by name and by whether they are non-empty, never by value. Today an empty token is invisible until a worker fails hours later.
  3. scripts/redeploy-bridged.sh and the supervision unit do not fight. A redeploy must not be undone by KeepAlive, and a crash must still be restarted. State the interaction in the script's own output.
  4. scripts/redeploy-bridged.sh --check reports whether supervision is installed and loaded, not only whether a process is alive.
  5. The install steps are exact for this host — no CHANGEME left for a human to interpret. A wrong path here fails silently at boot.
  6. wiki/11-Features.md entry: what it does · the knob · why · the gotcha. The gotcha is this contradiction.

Not in scope

  • Editing ${SHARED_ENV}/tools/secrets.sh. It is the operator's file.
  • The systemd unit (deploy/bridged.service). Same defect almost certainly applies, but the Linux gateways it targets are CB-308 work.
Found while checking release-1 close-out state on 2026-08-16. The daemon was **down** — it had died some time after 09:46 with no shutdown line — and nothing restarted it. ## The state on the host CB-504 shipped `deploy/dev.ltms.bridged.plist`. That file has `RunAtLoad` and a `KeepAlive` backstop, so it would have caught this. It is **not installed**: ``` ls ~/Library/LaunchAgents/ | grep bridged -> nothing ``` So the supervision unit exists in git and does not exist on the machine. That alone is a one-line install. The reason it was never installed is the real ticket. ## The contradiction `CLAUDE.md` §Redeploying the daemon, rule 1: > **Login shell, or workers silently lose their forge token.** The daemon inherits `WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from `${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the daemon starts fine, and the failure appears much later as workers that cannot open a PR. `deploy/dev.ltms.bridged.plist`, lines 52-56: > Worker/API tokens are NOT set here: this file is committed. Export them from a private launchd override or a wrapper script. **launchd does not source a login shell.** The plist even says so itself about `PATH` (line 47). So a daemon started by this unit gets: - no `WORKER_GITEA_TOKEN` → every worker silently fails to open a PR; - no `AI_GATEWAY_TOKEN` → every gateway-backed profile (`local`, `gx`) cannot authenticate. Nothing logs either at startup. `scripts/redeploy-bridged.sh --check` is the only thing that reports it, and it only checks the *login shell*, not the environment the daemon actually got. So today the operator picks one: a supervised daemon with a broken fleet, or a working fleet with no supervision. Neither is a release state. The "private override or wrapper script" the plist points at was never written. ## Why this is worth a ticket rather than a one-liner Three things have to agree, and right now none of them knows about the others: 1. the plist (a `CHANGEME` template, never instantiated); 2. `scripts/redeploy-bridged.sh`, which starts the daemon by hand and would fight a `KeepAlive` unit; 3. `${SHARED_ENV}/tools/secrets.sh`, which is the operator's file and must not be edited. ## Acceptance criteria 1. A supervised daemon starts with `WORKER_GITEA_TOKEN` and `AI_GATEWAY_TOKEN` resolved, without either secret appearing in a committed file. A wrapper that execs through a login shell is the obvious shape; pick one and write down why. 2. `bridged` logs at startup **which** required secrets resolved and which did not — by name and by whether they are non-empty, never by value. Today an empty token is invisible until a worker fails hours later. 3. `scripts/redeploy-bridged.sh` and the supervision unit do not fight. A redeploy must not be undone by `KeepAlive`, and a crash must still be restarted. State the interaction in the script's own output. 4. `scripts/redeploy-bridged.sh --check` reports whether supervision is installed and loaded, not only whether a process is alive. 5. The install steps are exact for this host — no `CHANGEME` left for a human to interpret. A wrong path here fails silently at boot. 6. `wiki/11-Features.md` entry: what it does · the knob · **why** · the gotcha. The gotcha is this contradiction. ## Not in scope - Editing `${SHARED_ENV}/tools/secrets.sh`. It is the operator's file. - The systemd unit (`deploy/bridged.service`). Same defect almost certainly applies, but the Linux gateways it targets are CB-308 work.
ltms added this to the 1.1 — single-host close-out milestone 2026-08-16 16:48:10 +02:00
ltms closed this issue 2026-08-16 17:32:57 +02:00
Author
Owner

Delivered and merged to main as PR #86.

The contradiction this ticket described is resolved. scripts/bridged-launchd-wrapper.sh execs one login shell in place, so a launchd-started daemon inherits ${SHARED_ENV}/tools/secrets.sh exactly as a hand-started one does. Supervision and a working fleet are no longer mutually exclusive.

The plist is filled in with this host's real paths — all three verified to exist, no CHANGEME left. redeploy-bridged.sh now detects whether the agent is loaded and switches stop/start to launchctl unload -w / load -w, falling back to the existing kill + nohup when it is not. That branch matters more than it looks: the implementer measured that a SIGTERMed daemon exits 143 even when its shutdown hook completes, so under KeepAlive{SuccessfulExit: false} a bare kill would have made launchd restart the old jar, racing the script's own restart.

Bridged also now reports at startup which required secret env vars resolved and which are MISSING, by name only. The required set is derived from the config — each non-subscription profile's tokenEnv and each profile's gitTokenEnv — not hard-coded. This closes the silent-failure mode in the ticket, where an empty token produced a healthy-looking daemon and surfaced hours later as workers unable to open a PR.

Verified by the lead: 813 tests on the branch, and 822 after trial-merging onto current main, both unpiped, both BUILD SUCCESS. A reviewer confirmed live that the unsupervised path is unchanged, that launchctl list exits 113 for an absent agent, that the wrapper round-trips arguments with spaces and quotes as a single exec chain, and that neither branch of the secret report can print a value.

The agent is still not installed, on purpose. That is the operator's call, and #91 (CB-600) lists what should land first — chiefly that the script computes its log path from its own location while the plist hard-codes one, so a mismatch would make the post-restart check report "ok" while the daemon crash-loops.

Install command, when the operator decides to:

cp deploy/dev.ltms.bridged.plist ~/Library/LaunchAgents/
launchctl load -w ~/Library/LaunchAgents/dev.ltms.bridged.plist
launchctl list | grep bridged
Delivered and merged to `main` as PR #86. The contradiction this ticket described is resolved. `scripts/bridged-launchd-wrapper.sh` execs one login shell in place, so a launchd-started daemon inherits `${SHARED_ENV}/tools/secrets.sh` exactly as a hand-started one does. Supervision and a working fleet are no longer mutually exclusive. The plist is filled in with this host's real paths — all three verified to exist, no `CHANGEME` left. `redeploy-bridged.sh` now detects whether the agent is loaded and switches stop/start to `launchctl unload -w` / `load -w`, falling back to the existing kill + nohup when it is not. That branch matters more than it looks: the implementer measured that a `SIGTERM`ed daemon exits 143 even when its shutdown hook completes, so under `KeepAlive{SuccessfulExit: false}` a bare `kill` would have made launchd restart the **old** jar, racing the script's own restart. `Bridged` also now reports at startup which required secret env vars resolved and which are MISSING, by name only. The required set is derived from the config — each non-subscription profile's `tokenEnv` and each profile's `gitTokenEnv` — not hard-coded. This closes the silent-failure mode in the ticket, where an empty token produced a healthy-looking daemon and surfaced hours later as workers unable to open a PR. Verified by the lead: 813 tests on the branch, and 822 after trial-merging onto current main, both unpiped, both BUILD SUCCESS. A reviewer confirmed live that the unsupervised path is unchanged, that `launchctl list` exits 113 for an absent agent, that the wrapper round-trips arguments with spaces and quotes as a single exec chain, and that neither branch of the secret report can print a value. **The agent is still not installed, on purpose.** That is the operator's call, and #91 (CB-600) lists what should land first — chiefly that the script computes its log path from its own location while the plist hard-codes one, so a mismatch would make the post-restart check report "ok" while the daemon crash-loops. Install command, when the operator decides to: ```bash cp deploy/dev.ltms.bridged.plist ~/Library/LaunchAgents/ launchctl load -w ~/Library/LaunchAgents/dev.ltms.bridged.plist launchctl list | grep bridged ```
Author
Owner

Criterion 6 (the wiki/11-Features.md entry) is now done too — I closed this before writing it, which was premature. Wiki commit f7dd817.

Two entries came out of this ticket rather than one, because the startup secret report is a separate capability with its own knob and its own gotcha:

  • Supervise the daemon without breaking the fleet — the wrapper, the install commands, why a bare kill was wrong (exit 143 even with a clean shutdown hook), and the log-path gotcha that CB-600 must fix first.
  • See at startup which secrets the daemon actually got — why it warns instead of refusing to boot, and that it prints names and set/MISSING only, never a value.

I also corrected the backfill list: the launchd half of CB-504 is catalogued now, but deploy/bridged.service has not been checked for the same login-shell defect. That is left on the list rather than quietly marked done.

Criterion 6 (the `wiki/11-Features.md` entry) is now done too — I closed this before writing it, which was premature. Wiki commit `f7dd817`. Two entries came out of this ticket rather than one, because the startup secret report is a separate capability with its own knob and its own gotcha: - **Supervise the daemon without breaking the fleet** — the wrapper, the install commands, why a bare `kill` was wrong (exit 143 even with a clean shutdown hook), and the log-path gotcha that CB-600 must fix first. - **See at startup which secrets the daemon actually got** — why it warns instead of refusing to boot, and that it prints names and set/MISSING only, never a value. I also corrected the backfill list: the launchd half of CB-504 is catalogued now, but `deploy/bridged.service` has **not** been checked for the same login-shell defect. That is left on the list rather than quietly marked done.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#80