a134eccc57
Separate Maven's output path (target/fleetd.jar) from the daemon's runtime path (run/fleetd.jar), so a clean/install in the main clone can no longer reach the jar a running daemon holds open. redeploy-fleetd.sh now swaps the built jar onto the runtime path with a same-filesystem mv, only after the old daemon is confirmed gone; --check reports the built and running jars as two separately labelled hash+mtime facts. Updates every process-locator pattern and runtime-path reference found by git grep, with a positive control added for both running_pid()'s PATTERN and fleets-status's pgrep pattern. fleetd/run/ is gitignored.
128 lines
6.8 KiB
Plaintext
128 lines
6.8 KiB
Plaintext
<?xml version="1.0" encoding="UTF-8"?>
|
|
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
|
|
<!--
|
|
CB-504 / CB-594 — launchd agent for fleetd (macOS).
|
|
|
|
This is the real supervision target today: the dogfooded daemon runs on macOS, where there is
|
|
no systemd. A systemd unit ships alongside (deploy/fleetd.service) for the Linux gateways
|
|
CB-308 introduces.
|
|
|
|
Install:
|
|
cp deploy/dev.ltms.fleetd.plist ~/Library/LaunchAgents/
|
|
launchctl load -w ~/Library/LaunchAgents/dev.ltms.fleetd.plist
|
|
launchctl list | grep fleetd
|
|
|
|
The paths below are already filled in for this host (resolved 2026-08-16 from
|
|
`/usr/libexec/java_home`... except that reported the system Applet-plugin JVM, not the jenv-
|
|
managed JDK 25 actually used to build/run fleetd, so JAVA_HOME here is the real one:
|
|
`JENV_VERSION=25.0.3 java -XshowSettings:properties -version 2>&1 | grep java.home`; `which mvn`;
|
|
`echo $HOME`). If this file is copied to a different host, re-resolve all three paths and check
|
|
no placeholder path is left behind; scripts/redeploy-fleetd.sh's check mode does not (and
|
|
cannot) check this file for you.
|
|
|
|
CB-594 — launchd cannot run a login shell (see the PATH comment on EnvironmentVariables below,
|
|
and scripts/fleetd-launchd-wrapper.sh for the fix): ProgramArguments below execs THAT wrapper,
|
|
not java directly, so WORKER_GITEA_TOKEN and AI_GATEWAY_TOKEN still get sourced from
|
|
${SHARED_ENV}/tools/secrets.sh even though launchd itself never sources anything.
|
|
|
|
Note on ordering: launchd has no "start after herdr" primitive for user agents, and neither
|
|
does systemd in a way that survives a socket appearing late. fleetd retries the herdr socket
|
|
on startup instead, so an agent that comes up before herdr converges rather than dying — that
|
|
retry is the actual fix; KeepAlive below is the backstop.
|
|
|
|
CB-594 — KeepAlive vs. scripts/redeploy-fleetd.sh: a bare SIGTERM makes this JVM exit 143 even
|
|
with its shutdown hook running to completion (measured, see the CB-594 report), which
|
|
SuccessfulExit:false below reads as a crash and races to restart the OLD jar. The redeploy
|
|
script now detects a loaded agent and uses `launchctl unload`/`load` instead of a raw kill, so
|
|
only one supervisor ever touches the process at a time — read that script's own output on a
|
|
redeploy for the confirmation.
|
|
-->
|
|
<plist version="1.0">
|
|
<dict>
|
|
<key>Label</key>
|
|
<string>dev.ltms.fleetd</string>
|
|
|
|
<key>ProgramArguments</key>
|
|
<array>
|
|
<string>/Users/dai.ha/LTMS/claude-bridge/scripts/fleetd-launchd-wrapper.sh</string>
|
|
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin/java</string>
|
|
<string>-jar</string>
|
|
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/run/fleetd.jar</string>
|
|
<string>fleetd.yaml</string>
|
|
</array>
|
|
|
|
<!-- Config path in ProgramArguments is relative, so the working directory must be the module. -->
|
|
<key>WorkingDirectory</key>
|
|
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd</string>
|
|
|
|
<key>EnvironmentVariables</key>
|
|
<dict>
|
|
<key>JAVA_HOME</key>
|
|
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home</string>
|
|
<key>HERDR_SOCKET_PATH</key>
|
|
<string>/Users/dai.ha/.config/herdr/herdr.sock</string>
|
|
<!--
|
|
PATH matters more than it looks (CB-511): fleetd propagates its own PATH to every worker
|
|
it spawns, so this line decides whether the fleet can run a build at all. launchd does NOT
|
|
source .zprofile/.zshrc, so without this the daemon (and therefore every worker) gets a
|
|
bare /usr/bin:/bin and no JDK or Maven. Keep the toolchain entries first.
|
|
-->
|
|
<key>PATH</key>
|
|
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin:/Users/dai.ha/Softwares/apache-maven/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin</string>
|
|
<!--
|
|
Worker/API tokens are NOT set here: this file is committed. CB-594 —
|
|
scripts/fleetd-launchd-wrapper.sh (named in ProgramArguments above) is what supplies
|
|
them, by execing a login shell that sources ${SHARED_ENV}/tools/secrets.sh before the
|
|
daemon itself starts. fleetd also reads the API token from the env var named by
|
|
auth.tokenEnv (default FLEETD_API_TOKEN) and only in auth.mode: token — the wrapper
|
|
covers that one too, since it is the same login shell.
|
|
-->
|
|
</dict>
|
|
|
|
<key>RunAtLoad</key>
|
|
<true/>
|
|
|
|
<!--
|
|
CB-600 — read this before assuming ThrottleInterval bounds anything. It paces restarts to at
|
|
most one per 10s; it does NOT cap how many times launchd retries. If fleetd fails fast on
|
|
every start — a bad fleetd.yaml, for example auth.mode: token with the token env var unset,
|
|
which throws in main() before the daemon ever binds a port — launchd restarts it forever,
|
|
once every 10s, until a human intervenes. LaunchAgents have no "give up after N attempts"
|
|
primitive, so this is not something a config change here can fix.
|
|
|
|
That loop stops only two ways: (1) `launchctl unload -w ~/Library/LaunchAgents/dev.ltms.fleetd.plist`,
|
|
or (2) the underlying cause gets fixed, so the process starts successfully and stays up (no
|
|
more exits to restart). scripts/redeploy-fleetd.sh does not add a third way — it does not
|
|
make fleetd self-disable on a config error, on purpose: a fail-fast exit path that
|
|
sometimes decides "this is unrecoverable, stop trying" is one more thing that can misfire,
|
|
and a wrongly self-disabled daemon needs the exact same manual `launchctl load -w` recovery
|
|
this comment already names — so it buys nothing an operator watching for the crash loop
|
|
doesn't already have, at the cost of a new way to be silently down. Watch for it with
|
|
`launchctl list dev.ltms.fleetd` (a high restart count) or by tailing fleetd.out for the
|
|
same startup error repeating every ~10s.
|
|
-->
|
|
<key>KeepAlive</key>
|
|
<dict>
|
|
<key>SuccessfulExit</key>
|
|
<false/>
|
|
</dict>
|
|
<key>ThrottleInterval</key>
|
|
<integer>10</integer>
|
|
|
|
<!--
|
|
CB-594 — same file scripts/redeploy-fleetd.sh already tails ($BRIDGED/fleetd.out), and both
|
|
streams point at it, not two separate log files: the script's fresh-line / ERROR-count checks
|
|
after a restart read this one path regardless of whether launchd or the script started the
|
|
process, and a stdout/stderr split would make half of what happens during a launchd-driven
|
|
restart invisible to it.
|
|
-->
|
|
<key>StandardOutPath</key>
|
|
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/fleetd.out</string>
|
|
<key>StandardErrorPath</key>
|
|
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/fleetd.out</string>
|
|
|
|
<key>ProcessType</key>
|
|
<string>Background</string>
|
|
</dict>
|
|
</plist>
|