dfb70871b4
PR #370 shipped both units but its install block stopped at `systemctl --user enable --now`. Without lingering a user manager starts at your first login and stops at your last logout, so the units do not come back after a reboot -- which is the whole reason this ticket moved fleet01 off the setsid scripts. It is easy to miss because leaving it out looks like success: `systemctl --user enable` reports "enabled" and both units run while you stay logged in. The issue named this and the PR did not carry it over. fleet01 itself is fine -- measured `Linger=yes`, both units `enabled`. This is about the next host that follows these instructions. Comment only; SystemdUnitSafetyTest still 8 green (a commented line is not an active directive).
86 lines
4.7 KiB
Desktop File
86 lines
4.7 KiB
Desktop File
# CB-504 — systemd unit for fleetd (Linux).
|
|
#
|
|
# fleetd #360: the previous version of this file started clean and broke the daemon in three ways
|
|
# that nothing logs (see the DO NOT block and the ExecStart/PrivateTmp comments below for what and
|
|
# why). The unit below, plus its companion deploy/herdr.service, is the version that has actually
|
|
# run on fleet01 without those failures. Do not "improve" it back toward the old shape without
|
|
# re-reading why each line is the way it is.
|
|
#
|
|
# Install (user service — fleetd drives the user's herdr, not a system daemon):
|
|
# mkdir -p ~/.config/systemd/user
|
|
# cp deploy/fleetd.service deploy/herdr.service ~/.config/systemd/user/
|
|
# # edit WorkingDirectory / ExecStart below for your host's paths and java location
|
|
# systemctl --user daemon-reload
|
|
# systemctl --user enable --now herdr fleetd
|
|
# loginctl enable-linger $USER # REQUIRED -- see below
|
|
# journalctl --user -u fleetd -f
|
|
#
|
|
# `loginctl enable-linger` is not optional and is easy to miss, because leaving it out looks like
|
|
# success: `systemctl --user enable` reports "enabled" and both units run for as long as you stay
|
|
# logged in. A user manager without lingering starts at your first login and stops at your last
|
|
# logout, so the fleet simply does not come back after a reboot -- which is the whole reason to
|
|
# use systemd here rather than the setsid scripts these units replaced. Check it with
|
|
# `loginctl show-user $USER -p Linger`; the answer must be `Linger=yes`.
|
|
#
|
|
# Secrets (AI_GATEWAY_TOKEN, WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI, ...) are not set
|
|
# here and need no systemd drop-in: ExecStart runs a login shell, so they come from wherever your
|
|
# login shell already sources them (this host: ~/.fleet/secrets.sh via ~/.zprofile). If a token is
|
|
# missing there, fleetd still starts — the daemon reports every secret a configured profile
|
|
# references, by name, never by value:
|
|
# journalctl --user -u fleetd | grep 'startup secret'
|
|
# A resolved one logs "startup secret NAME: set (profile 'x' tokenEnv)"; a missing one logs
|
|
# "startup secret NAME: MISSING" at WARN and the daemon starts anyway — the first visible symptom
|
|
# is a member that cannot open a pull request, hours later and in a different component.
|
|
|
|
[Unit]
|
|
Description=fleetd — fleet message server
|
|
Documentation=https://git.ltms.dev/fleet/fleetd/wiki
|
|
# Ordering only. fleetd retries the herdr socket rather than exiting, which is what actually makes
|
|
# a late socket survivable. Do NOT add Requires=: a herdr restart must not take fleetd down too.
|
|
After=herdr.service
|
|
Wants=herdr.service
|
|
|
|
[Service]
|
|
Type=simple
|
|
WorkingDirectory=%h/LTMS/fleetd/fleetd
|
|
|
|
# A LOGIN shell, not java directly. Every secret this daemon needs (AI_GATEWAY_TOKEN,
|
|
# WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI) lives in ~/.fleet/secrets.sh, which only
|
|
# ~/.zprofile sources. systemd runs no login shell. Started any other way the daemon boots fine
|
|
# and looks healthy, and the failure appears hours later as a member that cannot open a pull
|
|
# request. exec keeps it one process, so systemd tracks the right PID.
|
|
# This also avoids a SECOND copy of the secrets in a systemd drop-in: one source of truth.
|
|
ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"
|
|
|
|
# PrivateTmp MUST stay false -- see herdr.service. fleetd creates the member ZDOTDIR scrub dir and
|
|
# the opencode config dir under java.io.tmpdir, and the member pane (a herdr child, a different
|
|
# unit) has to read them. A private /tmp turns the credential scrub into a silent no-op.
|
|
PrivateTmp=false
|
|
|
|
Restart=on-failure
|
|
RestartSec=10s
|
|
# A bad config makes fleetd fail fast by design. Give up rather than restart-loop forever.
|
|
StartLimitBurst=5
|
|
StartLimitIntervalSec=120
|
|
|
|
# DO NOT add ProtectSystem=, ProtectHome=, ProtectKernelTunables= or ProtectControlGroups=.
|
|
# Measured on fleet01 2026-09-05: each of those gives the unit its own mount namespace, and
|
|
# fleetd resolves a caller role by running lsof to find the loopback peer PID
|
|
# (mcp/LsofPeerPidLookup). Inside such a namespace lsof returns nothing, every caller falls back
|
|
# to ANONYMOUS, and the primary is refused every orchestration call with
|
|
# "unauthenticated: anonymous may not SPAWN".
|
|
# The daemon still starts, healthz still returns ok and the secrets still resolve - the only
|
|
# symptom is that the fleet cannot be driven at all. Verified by bisecting the directives:
|
|
# no sandbox 3 lsof lines | ProtectSystem=strict 0 | ProtectHome=read-only 0
|
|
# ProtectKernelTunables 0 | ProtectControlGroups 0 | RestrictSUIDSGID 3 | NoNewPrivileges 3
|
|
# The two below add no mount namespace and are safe.
|
|
NoNewPrivileges=true
|
|
RestrictSUIDSGID=true
|
|
|
|
StandardOutput=journal
|
|
StandardError=journal
|
|
SyslogIdentifier=fleetd
|
|
|
|
[Install]
|
|
WantedBy=default.target
|