fleetd #360: fix systemd units that silently disabled caller identity and the credential scrub
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 2m12s

deploy/fleetd.service started clean on fleet01 but broke the daemon in three ways nothing
logs: ProtectSystem/ProtectHome/ProtectKernelTunables/ProtectControlGroups each put the unit
in its own mount namespace, which blinds fleetd's lsof-based caller lookup and falls every
caller back to ANONYMOUS; PrivateTmp=true silently no-ops the credential scrub the member
pane depends on; and running java directly from ExecStart skips the login shell that sources
the daemon's secrets, so it boots with empty credentials.

Replace the unit with the version verified working on fleet01 for a day, and add the
deploy/herdr.service companion unit it was already depending on via After=/Wants= but which
did not exist in the repo. Add deploy/herdr-inner.sh as the login-shell template
herdr.service's ExecStart wraps in a pty.

Add SystemdUnitSafetyTest (fleetd/src/test/java/dev/ltms/fleet/deploy) to read both unit
files from disk and fail if a forbidden mount-namespacing directive is active, PrivateTmp is
true, or fleetd.service's ExecStart does not go through a login shell -- the only guard
possible for a unit file with no compile step.
This commit is contained in:
Dai Ha
2026-09-06 19:45:23 +07:00
parent 6c61355f8f
commit 380eb63277
4 changed files with 267 additions and 46 deletions
+46 -46
View File
@@ -1,72 +1,72 @@
# CB-504 — systemd unit for fleetd (Linux).
#
# The macOS launchd agent (deploy/dev.ltms.fleetd.plist) is the supervision target for the
# current single-host deployment. This unit exists for the per-host gateways CB-308 introduces,
# which will run on Linux.
# fleetd #360: the previous version of this file started clean and broke the daemon in three ways
# that nothing logs (see the DO NOT block and the ExecStart/PrivateTmp comments below for what and
# why). The unit below, plus its companion deploy/herdr.service, is the version that has actually
# run on fleet01 without those failures. Do not "improve" it back toward the old shape without
# re-reading why each line is the way it is.
#
# Install (user service — fleetd drives the user's herdr, not a system daemon):
# mkdir -p ~/.config/systemd/user
# cp deploy/fleetd.service ~/.config/systemd/user/
# # edit ExecStart / WorkingDirectory / Environment below, then:
# cp deploy/fleetd.service deploy/herdr.service ~/.config/systemd/user/
# # edit WorkingDirectory / ExecStart below for your host's paths and java location
# systemctl --user daemon-reload
# systemctl --user enable --now fleetd
# systemctl --user enable --now herdr fleetd
# journalctl --user -u fleetd -f
#
# Secrets (AI_GATEWAY_TOKEN, WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI, ...) are not set
# here and need no systemd drop-in: ExecStart runs a login shell, so they come from wherever your
# login shell already sources them (this host: ~/.fleet/secrets.sh via ~/.zprofile). If a token is
# missing there, fleetd still starts — the daemon reports every secret a configured profile
# references, by name, never by value:
# journalctl --user -u fleetd | grep 'startup secret'
# A resolved one logs "startup secret NAME: set (profile 'x' tokenEnv)"; a missing one logs
# "startup secret NAME: MISSING" at WARN and the daemon starts anyway — the first visible symptom
# is a member that cannot open a pull request, hours later and in a different component.
[Unit]
Description=fleetd — claude-bridge message server
Description=fleetd — fleet message server
Documentation=https://git.ltms.dev/fleet/fleetd/wiki
# Ordering only: herdr is a user process and its socket may appear after us. This is advisory —
# fleetd retries the herdr socket rather than exiting, which is what actually makes a late
# socket survivable. Do NOT add Requires=: a herdr restart must not take fleetd down with it.
# Ordering only. fleetd retries the herdr socket rather than exiting, which is what actually makes
# a late socket survivable. Do NOT add Requires=: a herdr restart must not take fleetd down too.
After=herdr.service
Wants=herdr.service
[Service]
Type=simple
WorkingDirectory=%h/src/claude-bridge/fleetd
ExecStart=/usr/lib/jvm/temurin-25-jdk/bin/java -jar target/fleetd.jar fleetd.yaml
WorkingDirectory=%h/LTMS/fleetd/fleetd
Environment=HERDR_SOCKET_PATH=%h/.config/herdr/herdr.sock
# PATH matters more than it looks (CB-511): fleetd propagates its own PATH to every worker it
# spawns, so this line decides whether the fleet can run a build at all. systemd does not source a
# login shell, so without it the daemon — and every worker — gets a bare default with no JDK/Maven.
Environment=PATH=/usr/lib/jvm/temurin-25-jdk/bin:/usr/share/maven/bin:/usr/local/bin:/usr/bin:/bin
# Secrets are NOT set here — this file is committed. Put ALL three tokens in a private drop-in
# that systemd reads with restrictive permissions. In `systemctl --user edit fleetd`, add:
# [Service]
# Environment=FLEETD_API_TOKEN=...
# Environment=WORKER_GITEA_TOKEN=...
# Environment=AI_GATEWAY_TOKEN=...
# FLEETD_API_TOKEN protects fleetd's API. WORKER_GITEA_TOKEN lets members open pull requests; if
# it is missing, fleetd still starts, but a member fails when it later tries to open a pull request.
# AI_GATEWAY_TOKEN authenticates gateway profiles; if it is missing, fleetd still starts, but a
# gateway profile later returns HTTP 401. Or, put the same three variables in a 0600 file and add:
# EnvironmentFile=%h/.config/fleetd/env
# After starting, check which of them actually resolved. The daemon reports every secret a
# configured profile references, by name, never by value:
# journalctl --user -u fleetd | grep 'startup secret'
# A resolved one logs "startup secret NAME: set (profile 'x' tokenEnv)". A missing one logs
# "startup secret NAME: MISSING" at WARN — and the daemon starts anyway, which is the whole
# problem: without this grep the first sign is a member that cannot open a pull request, hours
# later and in a different component.
# Note what the report can and cannot tell you. It lists only names some profile actually
# references (tokenEnv, gitTokenEnv, and the broker uriEnv). A secret nothing references is never
# reported, because nothing needs it.
# A LOGIN shell, not java directly. Every secret this daemon needs (AI_GATEWAY_TOKEN,
# WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI) lives in ~/.fleet/secrets.sh, which only
# ~/.zprofile sources. systemd runs no login shell. Started any other way the daemon boots fine
# and looks healthy, and the failure appears hours later as a member that cannot open a pull
# request. exec keeps it one process, so systemd tracks the right PID.
# This also avoids a SECOND copy of the secrets in a systemd drop-in: one source of truth.
ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"
# PrivateTmp MUST stay false -- see herdr.service. fleetd creates the member ZDOTDIR scrub dir and
# the opencode config dir under java.io.tmpdir, and the member pane (a herdr child, a different
# unit) has to read them. A private /tmp turns the credential scrub into a silent no-op.
PrivateTmp=false
Restart=on-failure
RestartSec=10s
# A bad config (e.g. a non-loopback bind without token auth) makes fleetd fail fast by design.
# Give up rather than restart-loop on a permanent error.
# A bad config makes fleetd fail fast by design. Give up rather than restart-loop forever.
StartLimitBurst=5
StartLimitIntervalSec=120
# The daemon reads the repo, writes worktrees, and talks to a Unix socket — it needs no more.
# DO NOT add ProtectSystem=, ProtectHome=, ProtectKernelTunables= or ProtectControlGroups=.
# Measured on fleet01 2026-09-05: each of those gives the unit its own mount namespace, and
# fleetd resolves a caller role by running lsof to find the loopback peer PID
# (mcp/LsofPeerPidLookup). Inside such a namespace lsof returns nothing, every caller falls back
# to ANONYMOUS, and the primary is refused every orchestration call with
# "unauthenticated: anonymous may not SPAWN".
# The daemon still starts, healthz still returns ok and the secrets still resolve - the only
# symptom is that the fleet cannot be driven at all. Verified by bisecting the directives:
# no sandbox 3 lsof lines | ProtectSystem=strict 0 | ProtectHome=read-only 0
# ProtectKernelTunables 0 | ProtectControlGroups 0 | RestrictSUIDSGID 3 | NoNewPrivileges 3
# The two below add no mount namespace and are safe.
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=read-write
ProtectKernelTunables=true
ProtectControlGroups=true
RestrictSUIDSGID=true
StandardOutput=journal
+19
View File
@@ -0,0 +1,19 @@
#!/bin/zsh
# fleetd #360 — template for the script deploy/herdr.service's ExecStart wraps in a pty.
#
# `script -qfec <this> /dev/null` needs a real command to run, and that command has to be a LOGIN
# shell script: herdr itself needs the same secrets fleetd.service's login shell picks up (this
# host: ~/.fleet/secrets.sh via ~/.zprofile), because members it spawns inherit its environment.
# systemd's own Environment= lines in herdr.service are not enough for that -- they set TERM and a
# bare PATH so the pty starts at all, nothing more.
#
# Copy this file to the path deploy/herdr.service's ExecStart names
# (%h/LTMS/fleetd/fleetd-run/herdr-inner.sh by default) and `chmod +x` it. Not committed under
# that path itself because the session name below is host-specific.
# A 0x0 pty makes every pane spawn fail with "ghostty error -2" (see herdr-multi-instance-facts /
# fleet01-headless-herdr-standup) -- give it a real size before herdr ever touches it.
stty rows 50 cols 200
# -l: login shell, so herdr and everything it spawns gets the real secrets and PATH.
exec zsh -lc 'exec herdr --session <name>'
+40
View File
@@ -0,0 +1,40 @@
# fleetd #360 — systemd unit for herdr (Linux), the terminal multiplexer fleetd drives.
#
# This is fleetd.service's companion: fleetd.service's After=/Wants=herdr.service assumes this
# unit exists. Before this ticket it did not, so on a fresh host fleetd started against a herdr
# that systemd never supervised at all.
#
# Install: see deploy/fleetd.service's header comment (both units install the same way).
#
# ExecStart below runs deploy/herdr-inner.sh (copy the template of that name from this directory
# to the path in ExecStart, or point ExecStart at wherever you keep it, and make it executable).
# It is a separate file rather than an inline command because it must itself be a login shell (see
# its own header for why) and systemd's ExecStart does not run one.
[Unit]
Description=herdr terminal multiplexer (fleet session)
Documentation=https://git.ltms.dev/fleet/fleetd/wiki
[Service]
Type=simple
# script(1) gives herdr a real pty. Without it the client reports a 0x0 window and every pane
# spawn fails with "ghostty error -2" -- which surfaces as a fleetd spawn failure, not a herdr one.
ExecStart=/usr/bin/script -qfec %h/LTMS/fleetd/fleetd-run/herdr-inner.sh /dev/null
StandardInput=null
Environment=TERM=xterm-256color
Environment=PATH=%h/.local/bin:/usr/local/bin:/usr/bin:/bin
# PrivateTmp MUST stay false. fleetd writes the member ZDOTDIR scrub dir and the opencode config
# dir under its own java.io.tmpdir, and the member pane -- a child of THIS process -- has to read
# them. A private /tmp here silently breaks the credential scrub instead of failing loudly.
PrivateTmp=false
Restart=on-failure
RestartSec=5s
StandardOutput=journal
StandardError=journal
SyslogIdentifier=herdr
[Install]
WantedBy=default.target