CB-600: make it safe to install the launchd agent
- redeploy-bridged.sh now refuses (not warns) a supervised restart when its computed log path disagrees with the loaded plist's StandardOutPath — otherwise every post-restart check reads the wrong file and can report a clean restart while the daemon crash-loops. The check is a pure, testable function; the script gained a source-for-test guard so it can be exercised without installing the agent or touching launchd. - a failed 'launchctl load' after a successful 'unload' now retries once and, on ultimate failure, tells the operator the agent is stopped AND disabled plus the exact recovery command, instead of leaving that silently worse than the pre-redeploy state. - the plist documents honestly that the crash loop launchd retries is unbounded (ThrottleInterval only paces it), and what actually stops it. - fixed the requiredSecretEnvVars javadoc: the auth.tokenEnv startup throw is ~370 lines below its call site, not a few lines above it, and only fires in auth.mode: token.
This commit is contained in:
@@ -82,8 +82,25 @@
|
||||
<key>RunAtLoad</key>
|
||||
<true/>
|
||||
|
||||
<!-- Restart on crash, but not in a tight loop if the config is bad (bridged fails fast on a
|
||||
non-loopback bind without token auth — that is a config error, not a transient one). -->
|
||||
<!--
|
||||
CB-600 — read this before assuming ThrottleInterval bounds anything. It paces restarts to at
|
||||
most one per 10s; it does NOT cap how many times launchd retries. If bridged fails fast on
|
||||
every start — a bad bridged.yaml, for example auth.mode: token with the token env var unset,
|
||||
which throws in main() before the daemon ever binds a port — launchd restarts it forever,
|
||||
once every 10s, until a human intervenes. LaunchAgents have no "give up after N attempts"
|
||||
primitive, so this is not something a config change here can fix.
|
||||
|
||||
That loop stops only two ways: (1) `launchctl unload -w ~/Library/LaunchAgents/dev.ltms.bridged.plist`,
|
||||
or (2) the underlying cause gets fixed, so the process starts successfully and stays up (no
|
||||
more exits to restart). scripts/redeploy-bridged.sh does not add a third way — it does not
|
||||
make bridged self-disable on a config error, on purpose: a fail-fast exit path that
|
||||
sometimes decides "this is unrecoverable, stop trying" is one more thing that can misfire,
|
||||
and a wrongly self-disabled daemon needs the exact same manual `launchctl load -w` recovery
|
||||
this comment already names — so it buys nothing an operator watching for the crash loop
|
||||
doesn't already have, at the cost of a new way to be silently down. Watch for it with
|
||||
`launchctl list dev.ltms.bridged` (a high restart count) or by tailing bridged.out for the
|
||||
same startup error repeating every ~10s.
|
||||
-->
|
||||
<key>KeepAlive</key>
|
||||
<dict>
|
||||
<key>SuccessfulExit</key>
|
||||
|
||||
Reference in New Issue
Block a user