2e138a199b
Two changes ship together here.
1. One shared herdr workspace. The lead and every worker now live in one
workspace called "fleet", so the operator sees one "session" with many
windows, not two. Before, the lead sat in a "leads" workspace and workers
in "bridged-workers", which read as two sessions. The lead is still told
apart from workers by its exact tab label ("lead: <name>"), so putting them
in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
for split layouts; Fleetd now passes an empty exclude set.
2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
launchd/systemd units, module dir, and MCP mount).
- Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
- Log line, comments, docs, and CLAUDE.md updated to say fleetd.
- Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
- Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
- Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
still read as a fallback, and still gitignored.
- MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
server name is "fleet". The mount name in the local .mcp.json becomes
"fleet" (gitignored, not in this commit).
- Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.
Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.
Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.
The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).
949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
128 lines
6.8 KiB
Plaintext
128 lines
6.8 KiB
Plaintext
<?xml version="1.0" encoding="UTF-8"?>
|
|
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
|
|
<!--
|
|
CB-504 / CB-594 — launchd agent for fleetd (macOS).
|
|
|
|
This is the real supervision target today: the dogfooded daemon runs on macOS, where there is
|
|
no systemd. A systemd unit ships alongside (deploy/fleetd.service) for the Linux gateways
|
|
CB-308 introduces.
|
|
|
|
Install:
|
|
cp deploy/dev.ltms.fleetd.plist ~/Library/LaunchAgents/
|
|
launchctl load -w ~/Library/LaunchAgents/dev.ltms.fleetd.plist
|
|
launchctl list | grep fleetd
|
|
|
|
The paths below are already filled in for this host (resolved 2026-08-16 from
|
|
`/usr/libexec/java_home`... except that reported the system Applet-plugin JVM, not the jenv-
|
|
managed JDK 25 actually used to build/run fleetd, so JAVA_HOME here is the real one:
|
|
`JENV_VERSION=25.0.3 java -XshowSettings:properties -version 2>&1 | grep java.home`; `which mvn`;
|
|
`echo $HOME`). If this file is copied to a different host, re-resolve all three paths and check
|
|
no placeholder path is left behind; scripts/redeploy-fleetd.sh's check mode does not (and
|
|
cannot) check this file for you.
|
|
|
|
CB-594 — launchd cannot run a login shell (see the PATH comment on EnvironmentVariables below,
|
|
and scripts/fleetd-launchd-wrapper.sh for the fix): ProgramArguments below execs THAT wrapper,
|
|
not java directly, so WORKER_GITEA_TOKEN and AI_GATEWAY_TOKEN still get sourced from
|
|
${SHARED_ENV}/tools/secrets.sh even though launchd itself never sources anything.
|
|
|
|
Note on ordering: launchd has no "start after herdr" primitive for user agents, and neither
|
|
does systemd in a way that survives a socket appearing late. fleetd retries the herdr socket
|
|
on startup instead, so an agent that comes up before herdr converges rather than dying — that
|
|
retry is the actual fix; KeepAlive below is the backstop.
|
|
|
|
CB-594 — KeepAlive vs. scripts/redeploy-fleetd.sh: a bare SIGTERM makes this JVM exit 143 even
|
|
with its shutdown hook running to completion (measured, see the CB-594 report), which
|
|
SuccessfulExit:false below reads as a crash and races to restart the OLD jar. The redeploy
|
|
script now detects a loaded agent and uses `launchctl unload`/`load` instead of a raw kill, so
|
|
only one supervisor ever touches the process at a time — read that script's own output on a
|
|
redeploy for the confirmation.
|
|
-->
|
|
<plist version="1.0">
|
|
<dict>
|
|
<key>Label</key>
|
|
<string>dev.ltms.fleetd</string>
|
|
|
|
<key>ProgramArguments</key>
|
|
<array>
|
|
<string>/Users/dai.ha/LTMS/claude-bridge/scripts/fleetd-launchd-wrapper.sh</string>
|
|
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin/java</string>
|
|
<string>-jar</string>
|
|
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/target/fleetd.jar</string>
|
|
<string>fleetd.yaml</string>
|
|
</array>
|
|
|
|
<!-- Config path in ProgramArguments is relative, so the working directory must be the module. -->
|
|
<key>WorkingDirectory</key>
|
|
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd</string>
|
|
|
|
<key>EnvironmentVariables</key>
|
|
<dict>
|
|
<key>JAVA_HOME</key>
|
|
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home</string>
|
|
<key>HERDR_SOCKET_PATH</key>
|
|
<string>/Users/dai.ha/.config/herdr/herdr.sock</string>
|
|
<!--
|
|
PATH matters more than it looks (CB-511): fleetd propagates its own PATH to every worker
|
|
it spawns, so this line decides whether the fleet can run a build at all. launchd does NOT
|
|
source .zprofile/.zshrc, so without this the daemon (and therefore every worker) gets a
|
|
bare /usr/bin:/bin and no JDK or Maven. Keep the toolchain entries first.
|
|
-->
|
|
<key>PATH</key>
|
|
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin:/Users/dai.ha/Softwares/apache-maven/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin</string>
|
|
<!--
|
|
Worker/API tokens are NOT set here: this file is committed. CB-594 —
|
|
scripts/fleetd-launchd-wrapper.sh (named in ProgramArguments above) is what supplies
|
|
them, by execing a login shell that sources ${SHARED_ENV}/tools/secrets.sh before the
|
|
daemon itself starts. fleetd also reads the API token from the env var named by
|
|
auth.tokenEnv (default FLEETD_API_TOKEN) and only in auth.mode: token — the wrapper
|
|
covers that one too, since it is the same login shell.
|
|
-->
|
|
</dict>
|
|
|
|
<key>RunAtLoad</key>
|
|
<true/>
|
|
|
|
<!--
|
|
CB-600 — read this before assuming ThrottleInterval bounds anything. It paces restarts to at
|
|
most one per 10s; it does NOT cap how many times launchd retries. If fleetd fails fast on
|
|
every start — a bad fleetd.yaml, for example auth.mode: token with the token env var unset,
|
|
which throws in main() before the daemon ever binds a port — launchd restarts it forever,
|
|
once every 10s, until a human intervenes. LaunchAgents have no "give up after N attempts"
|
|
primitive, so this is not something a config change here can fix.
|
|
|
|
That loop stops only two ways: (1) `launchctl unload -w ~/Library/LaunchAgents/dev.ltms.fleetd.plist`,
|
|
or (2) the underlying cause gets fixed, so the process starts successfully and stays up (no
|
|
more exits to restart). scripts/redeploy-fleetd.sh does not add a third way — it does not
|
|
make fleetd self-disable on a config error, on purpose: a fail-fast exit path that
|
|
sometimes decides "this is unrecoverable, stop trying" is one more thing that can misfire,
|
|
and a wrongly self-disabled daemon needs the exact same manual `launchctl load -w` recovery
|
|
this comment already names — so it buys nothing an operator watching for the crash loop
|
|
doesn't already have, at the cost of a new way to be silently down. Watch for it with
|
|
`launchctl list dev.ltms.fleetd` (a high restart count) or by tailing fleetd.out for the
|
|
same startup error repeating every ~10s.
|
|
-->
|
|
<key>KeepAlive</key>
|
|
<dict>
|
|
<key>SuccessfulExit</key>
|
|
<false/>
|
|
</dict>
|
|
<key>ThrottleInterval</key>
|
|
<integer>10</integer>
|
|
|
|
<!--
|
|
CB-594 — same file scripts/redeploy-fleetd.sh already tails ($BRIDGED/fleetd.out), and both
|
|
streams point at it, not two separate log files: the script's fresh-line / ERROR-count checks
|
|
after a restart read this one path regardless of whether launchd or the script started the
|
|
process, and a stdout/stderr split would make half of what happens during a launchd-driven
|
|
restart invisible to it.
|
|
-->
|
|
<key>StandardOutPath</key>
|
|
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/fleetd.out</string>
|
|
<key>StandardErrorPath</key>
|
|
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/fleetd.out</string>
|
|
|
|
<key>ProcessType</key>
|
|
<string>Background</string>
|
|
</dict>
|
|
</plist>
|