Compare commits

...

13 Commits

Author SHA1 Message Date
Dai Ha 6442a583ae t355: give the withDefaults() coverage test a value for idleSleepGuard
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 2m6s
FleetConfigWithDefaultsPreservesEveryComponentTest was added on main after
the sleep-guard commit was cherry-picked, specifically to catch this exact
rebase hazard (a withDefaults() call silently rebinding to a stale-arity
back-compat constructor). Its baseValues() map didn't know about the new
idleSleepGuard component yet, so the coverage test itself failed the
name-drift check. Add a real, non-null value for it, consistent with how
the sibling ConfigRefTopLevelReportingCoverageTest already covers it.
2026-09-07 20:25:33 +07:00
Dai Ha 24b96d29ae fleetd must not let the host idle-sleep while members are live
Adds a small IdleSleepGuard (dev.ltms.fleet.power) that holds a macOS
caffeinate -i child while at least one fleet member is live, and
releases it once none are. It hangs off SessionManager's existing
onAcquire/onRelease hooks and SessionManager#size() rather than
tracking members a second way. New idleSleepGuard: config block,
on by default, following the FleetConfig.Health/ConfigReload pattern.
2026-09-07 20:22:11 +07:00
Dai Ha 4721771052 charter: a peer reads your project addendum
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m47s
Fourth item under Lead <-> lead. The addendum is instruction surface -- every
future session on that host obeys it, and a wrong one is obeyed as faithfully as
a right one -- but nothing in the block said to have anyone check it.

Evidence is one addendum, written on fleet01 this week, and it carried two
defects that its author did not see and a non-author did. First: it described a
forge permission wall as if it were policy, so a token regrade would have made
it tell a session NOT to merge at the moment merging became its job. Second,
found only because the first was raised: it DID carry the 'this is a dated
measurement' caveat, at the bottom, after the prohibition -- so a session that
read the instruction and stopped had already taken it as policy. The caveat sat
downstream of the thing it qualified.

Neither is a writing slip. Both are the author being unable to see their own
qualifier placement, which is what a second reader is for.

Deliberately conditional: not every operator runs two leads, so the rule ends
with what a lone lead can still do -- ask which sentence goes false first, and
whether a reader reaches the caveat before acting.

Weaker evidence than the step 8 amendment (n=1 addendum, 2 defects, versus a
measured 405). Recorded as such so it can be dropped if it does not earn the
context it costs.
2026-09-07 05:03:58 +07:00
Dai Ha 2f71a30bd7 charter: say what step 8 means when the forge refuses the merge
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 1m35s
Step 8 said 'then merge' and nothing about a refusal. fleet01 is the first host
to hit that: on akb/kb its lead gets 405 'User not allowed to merge PR' from the
API, and main is protected, so merging locally and pushing is refused too. Both
routes shut. The charter was telling a lead to do something the forge would not
let it do, and that gap was latent in every copy of the block.

The amendment keeps the rule that the merge decision is never delegated, while
admitting the mechanical merge may not be the lead's to make.

The second sentence is the one that earns its keep, and it came from the fleet01
lead rather than from me: never call a PR 'ready to merge' without having read
the diff. A refusal is exactly when that shortcut is tempting, because no action
is left that forces the lead to look. Without it, a refused merge quietly turns
step 8 from 'read it yourself, then merge' into 'forward the reviewer's verdict'
— the proxy-delegation the same step forbids two lines earlier.

Propagated to the wiki template in the same turn; the sync check in this file's
addendum reports 'in sync: True'.
2026-09-07 05:01:31 +07:00
Dai Ha 22cdebbdbe fleetd #369: state hermeticGitEnv's real scope, measured
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 1m32s
The comment claimed no test in this class can reach the real machine's home
directory. Measured on the merge: 53 of the 58 new GitWorktrees(...)
constructions here pass no env override, and stripping the override from
seedingGitWorktrees leaves the class green under the poison command that the
same comment cites as proof. Say what it covers and what it does not.
2026-09-06 20:32:21 +07:00
Dai Ha 154971c2b8 Merge #372: GitWorktreesTest no longer reads the operator's real git config
fleetd #369. Every git subprocess the test class starts now goes through one
gitProcessBuilder factory that applies the hermetic environment. Before this,
gitOutput set GIT_CONFIG_GLOBAL/SYSTEM/TERMINAL_PROMPT but not XDG_CONFIG_HOME,
and status/fullStatus set nothing at all -- so the tests inherited the JVM's
whole real environment, including the operator's default excludes file. That
file applies with no core.excludesFile configured at all, and /dev/null for the
global config does not stop it.

Round 1 pinned this with a call-site count: exactly 2 literal
new ProcessBuilder( occurrences. That catches a NEW helper built the old way,
but it is a proxy, not the property. I measured the gap -- deleting
pb.environment().putAll(hermeticEnv()) from inside the factory left every call
site unchanged, the count stayed 2, and 1413 tests stayed green. Round 2 added
gitProcessBuilderCarriesTheFullHermeticEnvironment, which asserts on what the
factory actually hands to ProcessBuilder#start(). Both checks are kept: they
catch different regressions.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1414, Failures: 0, Errors: 0, BUILD SUCCESS
  poison control on the PRE-FIX file:
    XDG_CONFIG_HOME=<dir with a '*' git/ignore> mvn test -Dtest=GitWorktreesTest
    -> tests=59 failures=56, so the poison genuinely reaches these tests

Mutation run on merge, on a half neither the worker nor the reviewer touched --
made seedingGitWorktrees pass null instead of hermeticGitEnv(tmp), which strips
the hermetic environment from the PRODUCTION GitWorktrees instances rather than
from the test's own subprocesses: tests=61 failures=0 unpoisoned AND poisoned.
That half is unpinned. It is fleetd #362's protection, not this ticket's, and it
guards a different path -- the Java-side XDG read in
previouslyEffectiveExcludesFileContent, which no assertion observes. Out of
scope here; filed as a follow-up rather than held against this PR.

One javadoc sentence corrected in the merge: hermeticGitEnv claimed "no test in
this class can reach the real machine's home directory". 53 of the 58
new GitWorktrees(...) constructions in this file pass no env override at all, so
the claim is true of the 5 seeding sites and of every test-started subprocess,
not of the class.
2026-09-06 20:31:01 +07:00
Dai Ha fef287c346 Merge #371: a stale lead binding no longer swallows a worker's reply nudge
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m55s
fleetd #368. PrimaryRegistry.forgetDelegation fires only when a WORKER is
released, never when the delegating LEAD goes away. The map is keyed by the
worker, so a closed, crashed or relaunched lead left its bindings behind. A
stale non-null entry then beat nudgeTargetFor's single-primary fallback every
time -- and that fallback's own javadoc argues it is correct precisely in the
case the stale entry was hiding.

ReplyPushLoop now probes the recorded lead with agents.status before trusting
it, at all five entry points, and a lead that is really gone is forgotten so
resolution reaches the fallback.

Round 1 caught any RuntimeException and treated it as death, and death here
calls forgetDelegation -- destructive and permanent on ONE reading. That is the
#359 mistake repeated two days later: a socket blip or a decode error on a live
lead would silently unbind it forever. I measured the breadth was unpinned
(narrowing it left 1414 green), and round 2 narrowed it to the one affirmative
signal AgentControl.agentCall itself uses, agent_not_found. Every other failure
is treated as live, because guessing wrong costs one extra retry next tick while
guessing wrong the other way is unrecoverable.

The fallback it lands on is live, not stale: PrimaryRegistry.record overwrites
the single slot on every orchestration-side MCP call, and neither host pins
primary.terminal, so a relaunched lead re-registers on its first tool call.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1415, Failures: 0, Errors: 0, BUILD SUCCESS

Mutation run on merge, on a half the worker did not touch -- dropped the
forgetDelegation call while still returning the fallback, so behaviour on the
first nudge is identical and only the self-healing is lost: 2 failures,
BUILD FAILURE. The cleanup is pinned, not just the fallback.
2026-09-06 20:13:07 +07:00
Dai Ha d6ef0c8013 fleetd #368 review: only agent_not_found may forget a lead binding
CI / build (pull_request) Successful in 1m22s
CI / contract (pull_request) Successful in 1m22s
isLive treated any RuntimeException from the liveness probe as "the lead is
gone", which forgetDelegation then acted on destructively and permanently.
That made a transient herdr hiccup (socket blip, decode error) on a perfectly
live lead indistinguishable from the lead actually being dead — the same
one-bad-reading mistake #359 shipped a guard against for lead-tab liveness.

Narrow isLive to match AgentControl.agentCall's own rule: only an affirmative
HerdrException("agent_not_found") counts as gone. Every other failure is
treated as still live and the binding is left alone.

Adds aTransientLivenessFailureMustNotForgetABindingToAStillLiveLead, which
fails with the bare RuntimeException catch and passes with the narrowed one.
2026-09-06 20:08:18 +07:00
Dai Ha dfb70871b4 fleetd #360 follow-up: the install block must say loginctl enable-linger
CI / contract (push) Successful in 43s
CI / build (push) Successful in 1m54s
PR #370 shipped both units but its install block stopped at
`systemctl --user enable --now`. Without lingering a user manager starts at
your first login and stops at your last logout, so the units do not come back
after a reboot -- which is the whole reason this ticket moved fleet01 off the
setsid scripts.

It is easy to miss because leaving it out looks like success: `systemctl --user
enable` reports "enabled" and both units run while you stay logged in. The
issue named this and the PR did not carry it over.

fleet01 itself is fine -- measured `Linger=yes`, both units `enabled`. This is
about the next host that follows these instructions.

Comment only; SystemdUnitSafetyTest still 8 green (a commented line is not an
active directive).
2026-09-06 20:08:05 +07:00
Dai Ha 4ee7b16929 Merge #370: ship systemd units that do not silently disable the daemon
CI / contract (push) Successful in 48s
CI / build (push) Successful in 2m1s
fleetd #360. deploy/fleetd.service shipped four mount-namespacing directives
(ProtectSystem, ProtectHome, ProtectKernelTunables, ProtectControlGroups),
PrivateTmp=true, and an ExecStart that ran java directly. Each one starts green
and breaks the daemon in a way nothing logs: lsof goes blind so every caller is
resolved ANONYMOUS and refused; the member ZDOTDIR scrub becomes a no-op; every
credential is empty. deploy/herdr.service did not exist at all, though
fleetd.service's After=/Wants= already named it.

Both units are now the ones running on fleet01, comments included -- the
bisected lsof counts and the reasons live in the files, because the next person
to 'harden' this needs the reason, not the rule.

A unit file has no compile step, so SystemdUnitSafetyTest reads both units plus
deploy/herdr-inner.sh and fails on an active forbidden directive, on
PrivateTmp=true, on an ExecStart that skips the login shell, on a herdr-inner.sh
that does not exec a login shell, and on one that does not set a non-zero pty
size. Each message names the consequence. A vacuity guard pins that all three
files exist and that the DO-NOT-add comment block still mentions every forbidden
directive, so 'comment survives, directive does not' is actually exercised.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1420, Failures: 0, Errors: 0, BUILD SUCCESS

Round 1 shipped herdr-inner.sh untested; I mutated it (dropped both the login
shell and the stty sizing) and got 6 green. Round 2 added those two checks.

Mutation run on merge, on a half the worker never touched -- ProtectHome=read-only
added to herdr.service, the unit it only ever mutated fleetd.service for:
Tests run: 8, Failures: 1, BUILD FAILURE. The parameterisation really covers
both files.
2026-09-06 20:04:48 +07:00
Dai Ha 8557289dc0 fleetd #360 review round 2: extend SystemdUnitSafetyTest to herdr-inner.sh
CI / build (pull_request) Successful in 1m27s
CI / contract (pull_request) Successful in 1m30s
herdr.service's ExecStart only names deploy/herdr-inner.sh, so that script was the only
place herdr's login-shell and pty-size properties lived, and nothing was reading it -- the
same silent-at-startup shape the ticket was about, one file further down the chain. A
mutation dropping both the login shell and the stty sizing left SystemdUnitSafetyTest green.

Add herdrInnerScriptUsesALoginShell and herdrInnerScriptSetsANonZeroPtySize, extend the
vacuity guard to cover herdr-inner.sh too, and tighten execStartUsesALoginShell's check from
a bare contains("-lc") substring match to the same login-shell-invocation regex the new
checks use.
2026-09-06 19:54:36 +07:00
Dai Ha 5af786d135 fleetd #368: a stale lead delegation binding must not shadow the primary fallback
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 1m55s
PrimaryRegistry.forgetDelegation only fires when a WORKER is released, never when
the delegating LEAD terminal itself disappears (closed, crashed, or relaunched).
A stale, non-null leadByTarget entry always beat nudgeTargetFor's single-primary
fallback, so a dead lead silently swallowed every reply nudge for its workers.

Fix: ReplyPushLoop now verifies (via the same agents.status check decide() already
uses every tick) that a recorded delegating lead is actually live before trusting
it. A dead lead is treated as if never recorded — self-healing the binding
(mirroring AgentControl.paneByTerminal's self-heal on agent_not_found) and falling
through to PrimaryRegistry's existing fallback.
2026-09-06 19:54:26 +07:00
Dai Ha 380eb63277 fleetd #360: fix systemd units that silently disabled caller identity and the credential scrub
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 2m12s
deploy/fleetd.service started clean on fleet01 but broke the daemon in three ways nothing
logs: ProtectSystem/ProtectHome/ProtectKernelTunables/ProtectControlGroups each put the unit
in its own mount namespace, which blinds fleetd's lsof-based caller lookup and falls every
caller back to ANONYMOUS; PrivateTmp=true silently no-ops the credential scrub the member
pane depends on; and running java directly from ExecStart skips the login shell that sources
the daemon's secrets, so it boots with empty credentials.

Replace the unit with the version verified working on fleet01 for a day, and add the
deploy/herdr.service companion unit it was already depending on via After=/Wants= but which
did not exist in the repo. Add deploy/herdr-inner.sh as the login-shell template
herdr.service's ExecStart wraps in a pty.

Add SystemdUnitSafetyTest (fleetd/src/test/java/dev/ltms/fleet/deploy) to read both unit
files from disk and fail if a forbidden mount-namespacing directive is active, PrivateTmp is
true, or fleetd.service's ExecStart does not go through a login shell -- the only guard
possible for a unit file with no compile step.
2026-09-06 19:45:23 +07:00
23 changed files with 1330 additions and 62 deletions
+11
View File
@@ -94,6 +94,12 @@ below are the procedure — run them in order, every task, not only the big ones
8. **Adjudicate, merge, tear down — yours alone.** Read the diff yourself: fully if it is small,
targeted at the reported findings and the risky paths if it is large. Reviewer findings direct
your attention; they never substitute for it. Then merge, then `fleet_stop{paneId}`.
**If the forge refuses you the merge** — a protected branch, a token without the grant — the
adjudication is still yours. Read the diff, decide, and hand the operator a merge-ready queue
with the refusal quoted. Never report a PR as merged, and never call one "ready to merge"
without having read the diff yourself. A refusal is exactly when that shortcut is tempting,
because no action is left that forces you to look, and taking it turns this step into
forwarding a reviewer's verdict — which is delegating the merge by proxy, two lines above.
**Steps 3 and 4 are separate on purpose** — spawning and sending in one loop is how parallel work
silently becomes serial, and it is the most common way this layer is wasted. For the same reason,
@@ -140,6 +146,11 @@ The traffic between leads is coordination and nothing else:
3. **Verify a peer exactly as you verify yourself.** Peer status buys nothing: check the claim
against the code, and re-run the build. A peer's correction gets the same treatment — right or
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
4. **Ask a peer to read your project addendum.** Your addendum is instruction surface: every future
session on your host obeys it, and a wrong one is obeyed just as faithfully as a right one. The
author is the worst reader of their own qualifier placement — measured here, one addendum carried
two defects and a non-author found both. If you have no peer, at least re-read it asking "which
sentence goes false first, and would a reader reach the caveat before acting?"
Being messaged by a peer does not make you its worker: answer with `fleet_reply`, and push back on
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
+54 -46
View File
@@ -1,72 +1,80 @@
# CB-504 — systemd unit for fleetd (Linux).
#
# The macOS launchd agent (deploy/dev.ltms.fleetd.plist) is the supervision target for the
# current single-host deployment. This unit exists for the per-host gateways CB-308 introduces,
# which will run on Linux.
# fleetd #360: the previous version of this file started clean and broke the daemon in three ways
# that nothing logs (see the DO NOT block and the ExecStart/PrivateTmp comments below for what and
# why). The unit below, plus its companion deploy/herdr.service, is the version that has actually
# run on fleet01 without those failures. Do not "improve" it back toward the old shape without
# re-reading why each line is the way it is.
#
# Install (user service — fleetd drives the user's herdr, not a system daemon):
# mkdir -p ~/.config/systemd/user
# cp deploy/fleetd.service ~/.config/systemd/user/
# # edit ExecStart / WorkingDirectory / Environment below, then:
# cp deploy/fleetd.service deploy/herdr.service ~/.config/systemd/user/
# # edit WorkingDirectory / ExecStart below for your host's paths and java location
# systemctl --user daemon-reload
# systemctl --user enable --now fleetd
# systemctl --user enable --now herdr fleetd
# loginctl enable-linger $USER # REQUIRED -- see below
# journalctl --user -u fleetd -f
#
# `loginctl enable-linger` is not optional and is easy to miss, because leaving it out looks like
# success: `systemctl --user enable` reports "enabled" and both units run for as long as you stay
# logged in. A user manager without lingering starts at your first login and stops at your last
# logout, so the fleet simply does not come back after a reboot -- which is the whole reason to
# use systemd here rather than the setsid scripts these units replaced. Check it with
# `loginctl show-user $USER -p Linger`; the answer must be `Linger=yes`.
#
# Secrets (AI_GATEWAY_TOKEN, WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI, ...) are not set
# here and need no systemd drop-in: ExecStart runs a login shell, so they come from wherever your
# login shell already sources them (this host: ~/.fleet/secrets.sh via ~/.zprofile). If a token is
# missing there, fleetd still starts — the daemon reports every secret a configured profile
# references, by name, never by value:
# journalctl --user -u fleetd | grep 'startup secret'
# A resolved one logs "startup secret NAME: set (profile 'x' tokenEnv)"; a missing one logs
# "startup secret NAME: MISSING" at WARN and the daemon starts anyway — the first visible symptom
# is a member that cannot open a pull request, hours later and in a different component.
[Unit]
Description=fleetd — claude-bridge message server
Description=fleetd — fleet message server
Documentation=https://git.ltms.dev/fleet/fleetd/wiki
# Ordering only: herdr is a user process and its socket may appear after us. This is advisory —
# fleetd retries the herdr socket rather than exiting, which is what actually makes a late
# socket survivable. Do NOT add Requires=: a herdr restart must not take fleetd down with it.
# Ordering only. fleetd retries the herdr socket rather than exiting, which is what actually makes
# a late socket survivable. Do NOT add Requires=: a herdr restart must not take fleetd down too.
After=herdr.service
Wants=herdr.service
[Service]
Type=simple
WorkingDirectory=%h/src/claude-bridge/fleetd
ExecStart=/usr/lib/jvm/temurin-25-jdk/bin/java -jar target/fleetd.jar fleetd.yaml
WorkingDirectory=%h/LTMS/fleetd/fleetd
Environment=HERDR_SOCKET_PATH=%h/.config/herdr/herdr.sock
# PATH matters more than it looks (CB-511): fleetd propagates its own PATH to every worker it
# spawns, so this line decides whether the fleet can run a build at all. systemd does not source a
# login shell, so without it the daemon — and every worker — gets a bare default with no JDK/Maven.
Environment=PATH=/usr/lib/jvm/temurin-25-jdk/bin:/usr/share/maven/bin:/usr/local/bin:/usr/bin:/bin
# Secrets are NOT set here — this file is committed. Put ALL three tokens in a private drop-in
# that systemd reads with restrictive permissions. In `systemctl --user edit fleetd`, add:
# [Service]
# Environment=FLEETD_API_TOKEN=...
# Environment=WORKER_GITEA_TOKEN=...
# Environment=AI_GATEWAY_TOKEN=...
# FLEETD_API_TOKEN protects fleetd's API. WORKER_GITEA_TOKEN lets members open pull requests; if
# it is missing, fleetd still starts, but a member fails when it later tries to open a pull request.
# AI_GATEWAY_TOKEN authenticates gateway profiles; if it is missing, fleetd still starts, but a
# gateway profile later returns HTTP 401. Or, put the same three variables in a 0600 file and add:
# EnvironmentFile=%h/.config/fleetd/env
# After starting, check which of them actually resolved. The daemon reports every secret a
# configured profile references, by name, never by value:
# journalctl --user -u fleetd | grep 'startup secret'
# A resolved one logs "startup secret NAME: set (profile 'x' tokenEnv)". A missing one logs
# "startup secret NAME: MISSING" at WARN — and the daemon starts anyway, which is the whole
# problem: without this grep the first sign is a member that cannot open a pull request, hours
# later and in a different component.
# Note what the report can and cannot tell you. It lists only names some profile actually
# references (tokenEnv, gitTokenEnv, and the broker uriEnv). A secret nothing references is never
# reported, because nothing needs it.
# A LOGIN shell, not java directly. Every secret this daemon needs (AI_GATEWAY_TOKEN,
# WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI) lives in ~/.fleet/secrets.sh, which only
# ~/.zprofile sources. systemd runs no login shell. Started any other way the daemon boots fine
# and looks healthy, and the failure appears hours later as a member that cannot open a pull
# request. exec keeps it one process, so systemd tracks the right PID.
# This also avoids a SECOND copy of the secrets in a systemd drop-in: one source of truth.
ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"
# PrivateTmp MUST stay false -- see herdr.service. fleetd creates the member ZDOTDIR scrub dir and
# the opencode config dir under java.io.tmpdir, and the member pane (a herdr child, a different
# unit) has to read them. A private /tmp turns the credential scrub into a silent no-op.
PrivateTmp=false
Restart=on-failure
RestartSec=10s
# A bad config (e.g. a non-loopback bind without token auth) makes fleetd fail fast by design.
# Give up rather than restart-loop on a permanent error.
# A bad config makes fleetd fail fast by design. Give up rather than restart-loop forever.
StartLimitBurst=5
StartLimitIntervalSec=120
# The daemon reads the repo, writes worktrees, and talks to a Unix socket — it needs no more.
# DO NOT add ProtectSystem=, ProtectHome=, ProtectKernelTunables= or ProtectControlGroups=.
# Measured on fleet01 2026-09-05: each of those gives the unit its own mount namespace, and
# fleetd resolves a caller role by running lsof to find the loopback peer PID
# (mcp/LsofPeerPidLookup). Inside such a namespace lsof returns nothing, every caller falls back
# to ANONYMOUS, and the primary is refused every orchestration call with
# "unauthenticated: anonymous may not SPAWN".
# The daemon still starts, healthz still returns ok and the secrets still resolve - the only
# symptom is that the fleet cannot be driven at all. Verified by bisecting the directives:
# no sandbox 3 lsof lines | ProtectSystem=strict 0 | ProtectHome=read-only 0
# ProtectKernelTunables 0 | ProtectControlGroups 0 | RestrictSUIDSGID 3 | NoNewPrivileges 3
# The two below add no mount namespace and are safe.
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=read-write
ProtectKernelTunables=true
ProtectControlGroups=true
RestrictSUIDSGID=true
StandardOutput=journal
+19
View File
@@ -0,0 +1,19 @@
#!/bin/zsh
# fleetd #360 — template for the script deploy/herdr.service's ExecStart wraps in a pty.
#
# `script -qfec <this> /dev/null` needs a real command to run, and that command has to be a LOGIN
# shell script: herdr itself needs the same secrets fleetd.service's login shell picks up (this
# host: ~/.fleet/secrets.sh via ~/.zprofile), because members it spawns inherit its environment.
# systemd's own Environment= lines in herdr.service are not enough for that -- they set TERM and a
# bare PATH so the pty starts at all, nothing more.
#
# Copy this file to the path deploy/herdr.service's ExecStart names
# (%h/LTMS/fleetd/fleetd-run/herdr-inner.sh by default) and `chmod +x` it. Not committed under
# that path itself because the session name below is host-specific.
# A 0x0 pty makes every pane spawn fail with "ghostty error -2" (see herdr-multi-instance-facts /
# fleet01-headless-herdr-standup) -- give it a real size before herdr ever touches it.
stty rows 50 cols 200
# -l: login shell, so herdr and everything it spawns gets the real secrets and PATH.
exec zsh -lc 'exec herdr --session <name>'
+40
View File
@@ -0,0 +1,40 @@
# fleetd #360 — systemd unit for herdr (Linux), the terminal multiplexer fleetd drives.
#
# This is fleetd.service's companion: fleetd.service's After=/Wants=herdr.service assumes this
# unit exists. Before this ticket it did not, so on a fresh host fleetd started against a herdr
# that systemd never supervised at all.
#
# Install: see deploy/fleetd.service's header comment (both units install the same way).
#
# ExecStart below runs deploy/herdr-inner.sh (copy the template of that name from this directory
# to the path in ExecStart, or point ExecStart at wherever you keep it, and make it executable).
# It is a separate file rather than an inline command because it must itself be a login shell (see
# its own header for why) and systemd's ExecStart does not run one.
[Unit]
Description=herdr terminal multiplexer (fleet session)
Documentation=https://git.ltms.dev/fleet/fleetd/wiki
[Service]
Type=simple
# script(1) gives herdr a real pty. Without it the client reports a 0x0 window and every pane
# spawn fails with "ghostty error -2" -- which surfaces as a fleetd spawn failure, not a herdr one.
ExecStart=/usr/bin/script -qfec %h/LTMS/fleetd/fleetd-run/herdr-inner.sh /dev/null
StandardInput=null
Environment=TERM=xterm-256color
Environment=PATH=%h/.local/bin:/usr/local/bin:/usr/bin:/bin
# PrivateTmp MUST stay false. fleetd writes the member ZDOTDIR scrub dir and the opencode config
# dir under its own java.io.tmpdir, and the member pane -- a child of THIS process -- has to read
# them. A private /tmp here silently breaks the credential scrub instead of failing loudly.
PrivateTmp=false
Restart=on-failure
RestartSec=5s
StandardOutput=journal
StandardError=journal
SyslogIdentifier=herdr
[Install]
WantedBy=default.target
+8
View File
@@ -110,6 +110,14 @@ bind:
# notifications:
# mode: disabled
# Idle-sleep guard: while at least one member is live, hold an OS-level assertion against idle
# sleep (macOS only — a `caffeinate -i` child; a no-op elsewhere or if caffeinate is missing), so
# an unattended host does not idle-sleep out from under a member's long turn. Unlike health/
# configReload above, this is ON BY DEFAULT — omitting the block entirely leaves it enabled, the
# same as `enabled: true`. Uncomment only to turn it off:
# idleSleepGuard:
# enabled: false
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
@@ -55,6 +55,8 @@ import dev.ltms.fleet.member.MemberCredentialPolicyView;
import dev.ltms.fleet.member.OpenCodeLauncher;
import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.power.CaffeinateSleepAssertionMechanism;
import dev.ltms.fleet.power.IdleSleepGuard;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -253,6 +255,26 @@ public final class Fleetd {
System::nanoTime, contextCap, clearAfterTurn);
liveCountRef.set(profileName -> liveSessionCount(sessions.roster(), profileName));
// Idle-sleep guard: hold an OS-level assertion against idle sleep while at least one
// member is live, so an unattended host does not idle-sleep out from under a member's
// long turn (see FleetConfig.IdleSleepGuard / dev.ltms.fleet.power.IdleSleepGuard for the
// measurement that motivated this). Opt-out via idleSleepGuard.enabled: false; on by
// default. Hangs off SessionManager's own onAcquire/onRelease hooks (CB-520/CB-516,
// previously wired only to the reply inbox) and SessionManager#size() — the exact registry
// fleet_list's live/capacity numbers are themselves computed from — rather than tracking
// members a second way. No-op (never constructed) off macOS or when idleSleepGuard.enabled
// is explicitly false; the mechanism itself is additionally a no-op if 'caffeinate' cannot
// be started, so this can never fail a spawn, a release, or startup.
boolean idleSleepGuardEnabled = cfg.idleSleepGuard() == null || cfg.idleSleepGuard().isEnabled();
final IdleSleepGuard idleSleepGuard;
if (idleSleepGuardEnabled) {
idleSleepGuard = new IdleSleepGuard(new CaffeinateSleepAssertionMechanism(), sessions::size);
sessions.onAcquire(_ -> idleSleepGuard.recheck());
sessions.onRelease(_ -> idleSleepGuard.recheck());
} else {
idleSleepGuard = null;
}
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
final SessionReaper reaper;
if (cfg.lifecycle() != null
@@ -704,6 +726,11 @@ public final class Fleetd {
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
mcp.close();
if (reaper != null) reaper.stop();
// Idle-sleep guard: release unconditionally, even though sessions.close() above already
// drained every session (and each release already drove the live count to 0, which
// releases the guard's assertion on its own) — this is the backstop for a drain that was
// itself interrupted or threw, so no caffeinate child ever outlives the daemon.
if (idleSleepGuard != null) idleSleepGuard.close();
// Release the broker connection last among message resources (no-op for the in-memory inbox).
if (replyInbox instanceof AutoCloseable closeable) {
try {
@@ -37,6 +37,10 @@ import java.util.function.Supplier;
* makes {@code fleet:} split rather than hot — see below.</li>
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
* {@code idleSleepGuard:} ({@code Fleetd.java} reads it once, at startup, to decide whether
* to construct an {@code IdleSleepGuard} and wire {@code SessionManager}'s
* {@code onAcquire}/{@code onRelease} hooks to it — neither is rebuilt on reload, so a
* running daemon keeps whatever this was at startup regardless of a later edit),
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code quarantineCooldownSeconds}
* (CB-578 stage B — baked once into the {@code BackendQuarantine} built at startup),
* {@code guard:}, {@code worktreeRoot:}, {@code worktreeGroup:} and {@code memberSkills:}
@@ -131,9 +135,9 @@ import java.util.function.Supplier;
* </ul>
*
* <p><strong>The denominator, measured on 2026-09-04 (fleetd #330; recounted for fleetd #333);
* recounted again for fleetd #362.</strong> {@code FleetConfig} has 23 top-level record components:
* 5 cold, 12 deferred, 3 split, 3 hot-excluded. Three of them are named nowhere in this file, and
* the reason is the same for all
* recounted again for fleetd #362, and again after {@code idleSleepGuard:} was added.</strong>
* {@code FleetConfig} has 24 top-level record components: 5 cold, 13 deferred, 3 split, 3
* hot-excluded. Three of them are named nowhere in this file, and the reason is the same for all
* three: {@code placement}, {@code memberCredentials} and {@code memberLoginShell} are
* <strong>hot</strong> and correctly absent — all three are read live off {@code config.get()}
* (placement through the {@code CompositePeerLauncher} supplier the Hot bullet names;
@@ -215,7 +219,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
static final Set<String> DEFERRED_KEYS = Set.of(
"guard", "worktreeRoot", "worktreeGroup", "memberSkills", "primary", "configReload",
"leadHeartbeat", "lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs",
"quarantineCooldownSeconds", "profiles");
"quarantineCooldownSeconds", "profiles", "idleSleepGuard");
private final Path path;
private final AtomicReference<FleetConfig> current;
@@ -427,6 +431,14 @@ public final class ConfigRef implements Supplier<FleetConfig> {
if (!Objects.equals(old.configReload(), fresh.configReload())) {
changed.add("configReload");
}
// Fleetd.java reads cfg.idleSleepGuard() once, at startup, to decide whether to construct
// an IdleSleepGuard at all and wire SessionManager's onAcquire/onRelease hooks to it —
// neither is rebuilt on reload, so a running daemon keeps whatever this was at startup
// (armed or not) regardless of a later edit here. Not cold: nothing already-open goes
// inconsistent with the new value, an armed-or-not guard just keeps its original answer.
if (!Objects.equals(old.idleSleepGuard(), fresh.idleSleepGuard())) {
changed.add("idleSleepGuard");
}
if (!Objects.equals(old.spawnReadyTimeoutMs(), fresh.spawnReadyTimeoutMs())
|| !Objects.equals(old.spawnReadyPollMs(), fresh.spawnReadyPollMs())) {
changed.add("spawnReady*");
@@ -119,6 +119,11 @@ import java.util.regex.PatternSyntaxException;
* subdirectory of this directory is copied wholesale, with no per-file
* allowlist — do not park scratch files or drafts alongside the real skill
* folders, they will be copied into every provisioned worktree too.
* @param idleSleepGuard opt-in-by-default: hold an OS-level assertion against idle sleep while at
* least one member is live, so an unattended host does not sleep out from
* under a member's long turn. {@code null} (the block omitted) behaves the
* same as an explicit {@code enabled: true}; set {@code enabled: false} to
* turn it off. See {@link dev.ltms.fleet.power.IdleSleepGuard}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record FleetConfig(
@@ -144,7 +149,22 @@ public record FleetConfig(
Coordinator coordinator,
String worktreeGroup,
String memberLoginShell,
String memberSkills) {
String memberSkills,
IdleSleepGuard idleSleepGuard) {
/** Back-compat form before the {@code idleSleepGuard:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
String memberLoginShell, String memberSkills) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
memberLoginShell, memberSkills, null);
}
/** Back-compat form before the {@code memberSkills} key was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
@@ -157,7 +177,7 @@ public record FleetConfig(
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
memberLoginShell, null);
memberLoginShell, null, null);
}
/** Back-compat form before the {@code memberLoginShell} key was added. */
@@ -169,7 +189,7 @@ public record FleetConfig(
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup, null);
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup, null, null);
}
/** Back-compat form before the {@code worktreeGroup} key was added. */
@@ -1279,6 +1299,25 @@ public record FleetConfig(
}
}
/**
* Hold an OS-level assertion against idle sleep while at least one member is live (see
* {@link dev.ltms.fleet.power.IdleSleepGuard}).
*
* <p>Unlike most opt-in blocks in this file, this one defaults to <em>on</em>: an unattended
* host idle-sleeping mid-turn is a correctness problem (a dropped AMQP link, a frozen member),
* not a convenience, so the safer default is armed. An operator who wants the previous
* behaviour (no assertion held, ever) sets {@code enabled: false} explicitly.
*
* @param enabled {@code false} turns the guard off; {@code null} (the block omitted
* entirely) or {@code true} leaves it on
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record IdleSleepGuard(Boolean enabled) {
public boolean isEnabled() {
return !Boolean.FALSE.equals(enabled);
}
}
/**
* The terminal → lead-name map seeded from the legacy singular {@code primary:} pin (CB-530).
*
@@ -1529,7 +1568,8 @@ public record FleetConfig(
"bind", "herdrSocket", "memberHerdrSocket", "profiles", "guard", "worktreeRoot",
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds",
"memberCredentials", "coordinator", "worktreeGroup", "memberLoginShell", "memberSkills");
"memberCredentials", "coordinator", "worktreeGroup", "memberLoginShell", "memberSkills",
"idleSleepGuard");
/** Load and validate config from {@code path}. */
public static FleetConfig load(Path path) {
@@ -2209,9 +2249,15 @@ public record FleetConfig(
// memberSkills is left as-is (fleetd #362), like worktreeGroup/memberLoginShell: null/blank
// is "off", and there is no sane non-null default — the daemon may not even run from a
// checkout that ships its own .claude/skills/.
// idleSleepGuard is left as-is, like leadHeartbeat/configReload above, but for the opposite
// reason: it is on by default already (its own isEnabled() treats null the same as
// enabled: true — see its javadoc), so defaulting the block here would change nothing a
// reader observes and would only obscure that "block omitted" and "block present and
// enabled" are deliberately the same outcome.
return new FleetConfig(b, herdrSocket, memberHerdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
quarantineCooldown, mc, coordinator, worktreeGroup, memberLoginShell, memberSkills);
quarantineCooldown, mc, coordinator, worktreeGroup, memberLoginShell, memberSkills,
idleSleepGuard);
}
/**
@@ -2,6 +2,7 @@ package dev.ltms.fleet.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
@@ -13,6 +14,7 @@ import java.util.Collection;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
@@ -375,6 +377,80 @@ public final class ReplyPushLoop {
return Action.WAIT_BUSY;
}
/**
* Resolve who to nudge about {@code target}, the way every public entry point below wants it:
* {@link PrimaryRegistry#nudgeTargetFor}, but only after checking the delegating lead it names
* is still actually there (fleetd #368).
*
* <p><strong>The bug this closes.</strong> {@code PrimaryRegistry.forgetDelegation} is wired to
* exactly one event — a worker's release — because that is the only teardown the daemon already
* observes for a session in this map. Nothing removes a binding when the LEAD half goes away: a
* lead that is closed, crashes, or is relaunched leaves {@code leadByTarget} entries pointing at
* a terminal that no longer exists. {@code nudgeTargetFor} falls back to the single known
* primary only when the map holds nothing for {@code target} — a stale non-null entry beats the
* fallback every time, which is exactly backwards: the fallback's own javadoc argues it is safe
* precisely in the case a stale entry now hides.
*
* <p><strong>The fix.</strong> Before trusting a recorded delegation, probe the lead the same
* way {@link #decide} already does every tick ({@code agents.status}) — cheap, since it is a
* local herdr round-trip, and it is the same signal {@code AgentControl.paneByTerminal} already
* trusts to tell a genuinely dead target from a live one. A lead that fails the probe is treated
* as if it had never been recorded: the stale entry is forgotten (self-healing, exactly like
* {@code AgentControl.paneByTerminal} already does on {@code agent_not_found}) and resolution is
* retried, which now reaches the fallback {@code nudgeTargetFor} was built to reach — the same
* empty-map state its javadoc already argues is correct.
*
* <p><strong>fleetd #368 review — only a positive "gone" reading forgets the binding.</strong>
* The first version of this method treated <em>any</em> {@code RuntimeException} from the probe
* as death, which is the #359 mistake repeated: a transient socket blip or a codec error on a
* perfectly live lead would silently and permanently unbind it, with no re-record ever coming.
* That is destructive on one bad reading, exactly what #359 shipped a two-reading guard to avoid
* for the analogous lead-tab-liveness question. {@link #isLive} now matches
* {@code AgentControl.agentCall}'s own narrower rule (see its {@code agent_not_found} check): only
* that specific, affirmative "herdr has no such agent" signal counts as gone. Every other failure
* — timeout, transport error, a decode error — is treated as still live and the binding is left
* alone, because guessing wrong here is unrecoverable while guessing "live" merely costs one more
* retry on the next tick, which {@link #decide} already tolerates.
*/
private Optional<String> resolveLiveLead(String target) {
Optional<String> lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty() || isLive(lead.get())) {
return lead;
}
log.debug("push: lead {} delegated to for {} is no longer live, forgetting the stale binding "
+ "and falling back", lead.get(), target);
primaryRegistry.forgetDelegation(target);
return primaryRegistry.nudgeTargetFor(target);
}
/**
* Whether {@code lead} should still be trusted: {@code false} only when herdr affirmatively
* reports the terminal gone ({@code agent_not_found}), never on a merely inconclusive failure.
*
* <p>fleetd #368 review: an earlier version returned {@code false} for any {@code RuntimeException},
* which made a transient herdr hiccup on a live lead indistinguishable from the lead actually
* being dead — and the caller's response to {@code false} ({@code forgetDelegation}) is
* destructive and permanent. Narrowed to the one code {@code AgentControl.agentCall} itself
* already treats as a genuine, resolvable absence (see its {@code agent_not_found} handling) —
* every other {@code RuntimeException} is treated as "still live" and the binding survives to be
* probed again next time, which costs nothing worse than one more retry.
*/
private boolean isLive(String lead) {
try {
agents.status(lead);
return true;
} catch (RuntimeException e) {
boolean gone = e instanceof HerdrException he && "agent_not_found".equals(he.code());
if (gone) {
log.debug("push: lead {} no longer exists ({})", lead, e.toString());
} else {
log.debug("push: liveness check for lead {} was inconclusive ({}); treating as live "
+ "rather than risk destroying a live binding", lead, e.toString());
}
return !gone;
}
}
// --- public entrypoints ----------------------------------------------------------------------
/**
@@ -385,7 +461,7 @@ public final class ReplyPushLoop {
* backstop until a lead is recorded.
*/
public void onReplyQueued(String target) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}, skipping reminder", target);
return;
@@ -413,7 +489,7 @@ public final class ReplyPushLoop {
* @param failed whether the ticket ended in a failure phase rather than {@code DONE}
*/
public void onTicketTerminal(String ticket, String target, boolean failed) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on ticket {} (target {}), skipping nudge",
ticket, target);
@@ -448,7 +524,7 @@ public final class ReplyPushLoop {
* @param question the question text
*/
public void onQuestionOpened(String ticket, String target, String turnId, String question) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}'s question (turnId {}), skipping nudge",
target, turnId);
@@ -476,7 +552,7 @@ public final class ReplyPushLoop {
Collection<String> profiles, int remainingCoolOffSeconds) {
Map<String, List<String>> targetsByLead = new ConcurrentHashMap<>();
for (String target : targets) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.warn("push: backend incident {} has no known lead for target {}", incidentId, target);
continue;
@@ -498,7 +574,7 @@ public final class ReplyPushLoop {
* Without an owning lead, emit a warning because no control can act on the target.
*/
public void onBackendTargetUnmapped(String target, String reason) {
var lead = primaryRegistry.nudgeTargetFor(target);
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.warn("push: backend target {} could not map to a credential: {}", target, reason);
return;
@@ -0,0 +1,93 @@
package dev.ltms.fleet.power;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.util.Locale;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicBoolean;
/**
* Holds macOS idle sleep off by keeping a {@code caffeinate -i} child process alive for the life
* of the returned {@link SleepAssertion}.
*
* <p>{@code -i} asserts only against <em>idle</em> sleep — it does not stop the lid closing or an
* operator-requested sleep from taking effect. That is deliberate: this class exists to stop an
* unattended host from sleeping out from under a member's long turn, never to override the
* operator. {@code -s}/{@code -d} (which also block system/display sleep on demand) are
* intentionally not used here.
*
* <p>{@link #acquire()} never throws. It returns {@code null} — a no-op — off macOS, and again if
* starting the {@code caffeinate} child fails for any reason (binary missing, process table full,
* …); either case is logged once at INFO, not on every occurrence, so a daemon that runs for
* weeks with the tool unavailable does not fill its log.
*/
public final class CaffeinateSleepAssertionMechanism implements SleepAssertionMechanism {
private static final Logger log = LoggerFactory.getLogger(CaffeinateSleepAssertionMechanism.class);
private final AtomicBoolean loggedOnce = new AtomicBoolean(false);
/** {@code true} when running on macOS, the only platform {@code caffeinate} ships on. */
public static boolean isSupportedPlatform() {
return isSupportedPlatform(System.getProperty("os.name"));
}
/** Package-visible so a test can drive the platform check without touching a real property. */
static boolean isSupportedPlatform(String osName) {
return osName != null && osName.toLowerCase(Locale.ROOT).contains("mac");
}
@Override
public SleepAssertion acquire() {
if (!isSupportedPlatform()) {
logOnce("not running on macOS (os.name={}); the idle-sleep guard is a no-op on this platform",
System.getProperty("os.name"));
return null;
}
try {
Process process = new ProcessBuilder("caffeinate", "-i")
.redirectOutput(ProcessBuilder.Redirect.DISCARD)
.redirectError(ProcessBuilder.Redirect.DISCARD)
.start();
return new CaffeinateAssertion(process);
} catch (IOException | RuntimeException e) {
logOnce("could not start 'caffeinate -i' ({}); the host may idle-sleep while members are live",
e.toString());
return null;
}
}
private void logOnce(String format, Object arg) {
if (loggedOnce.compareAndSet(false, true)) {
log.info("idle-sleep guard: " + format, arg);
}
}
/** Wraps the live {@code caffeinate} child; {@link #close} force-destroys it, idempotently. */
private static final class CaffeinateAssertion implements SleepAssertion {
private final Process process;
CaffeinateAssertion(Process process) {
this.process = process;
}
@Override
public void close() {
if (!process.isAlive()) {
return;
}
process.destroy();
try {
if (!process.waitFor(2, TimeUnit.SECONDS)) {
process.destroyForcibly();
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
process.destroyForcibly();
}
}
}
}
@@ -0,0 +1,105 @@
package dev.ltms.fleet.power;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.function.IntSupplier;
/**
* Holds an OS-level assertion against idle sleep for exactly as long as at least one fleet
* member is live.
*
* <p><strong>Why this exists:</strong> a fleetd host was measured idle-sleeping after as little
* as one minute of inactivity (its {@code pmset -g custom} reports {@code sleep 1} on battery).
* Overnight the daemon's AMQP link to the broker dropped 13 times, and cross-checking every drop
* minute against {@code pmset -g log} found a sleep or wake event in the same minute or the one
* before, every time. The AMQP churn is only the visible symptom — the real problem is that a
* member mid-turn freezes with the host, and a long turn with nobody typing is exactly the case
* that goes idle.
*
* <p><strong>How it tracks "live":</strong> this is driven by {@code SessionManager}'s existing
* {@code onAcquire}/{@code onRelease} lifecycle hooks (added for CB-520/CB-516, previously wired
* to nothing but the reply inbox) rather than a second member count kept in parallel. Wire it as:
* <pre>{@code
* IdleSleepGuard guard = new IdleSleepGuard(mechanism, sessions::size);
* sessions.onAcquire(_ -> guard.recheck());
* sessions.onRelease(_ -> guard.recheck());
* }</pre>
* Every acquire/release event re-reads {@code SessionManager#size()} — the same registry {@code
* fleet_list}'s live/capacity numbers are themselves computed from — and only an actual 0→1 or
* 1→0 crossing touches the OS. A listener exception is already caught and logged by {@code
* SessionManager} itself (it must never let a listener failure block the acquire/release it is
* reacting to), so {@link #recheck()} does not need its own top-level try/catch to honor that.
*
* <p><strong>Failure posture:</strong> every method here is safe to call whether or not {@link
* SleepAssertionMechanism#acquire()} actually works. A mechanism that returns {@code null} (wrong
* platform, missing tool, spawn failure) simply means this guard never holds anything — it never
* throws and never blocks a spawn, a release, or shutdown.
*/
public final class IdleSleepGuard implements AutoCloseable {
private static final Logger log = LoggerFactory.getLogger(IdleSleepGuard.class);
private final SleepAssertionMechanism mechanism;
private final IntSupplier liveCount;
private final Object lock = new Object();
private SleepAssertion held;
public IdleSleepGuard(SleepAssertionMechanism mechanism, IntSupplier liveCount) {
this.mechanism = mechanism;
this.liveCount = liveCount;
}
/**
* Re-read the live count and acquire or release the held assertion to match: nothing held and
* at least one member live ⇒ acquire; something held and no member live ⇒ release. A steady
* count (still zero, still positive) is a no-op either way, so a single spawn or release only
* ever touches the OS on the crossing, not on every call.
*/
public void recheck() {
synchronized (lock) {
int live = liveCount.getAsInt();
if (live > 0 && held == null) {
held = mechanism.acquire();
if (held != null) {
log.debug("idle-sleep guard armed: {} live member(s)", live);
}
} else if (live == 0 && held != null) {
releaseHeldLocked();
}
}
}
/** {@code true} while an assertion is actually held. Exposed for tests. */
boolean isHeld() {
synchronized (lock) {
return held != null;
}
}
/**
* Release whatever is held, if anything. Idempotent and safe to call at any time, including
* repeatedly — a daemon shutdown hook calls this unconditionally so no assertion (and no
* {@code caffeinate} child) survives the process, even if the drain that would otherwise have
* driven the live count to zero was itself interrupted or threw.
*/
@Override
public void close() {
synchronized (lock) {
if (held != null) {
releaseHeldLocked();
}
}
}
/** Caller must hold {@link #lock}. */
private void releaseHeldLocked() {
try {
held.close();
} catch (RuntimeException e) {
log.warn("idle-sleep guard: failed to release its assertion cleanly: {}", e.toString());
} finally {
held = null;
}
}
}
@@ -0,0 +1,11 @@
package dev.ltms.fleet.power;
/**
* A held OS-level assertion against idle sleep. {@link #close} must be idempotent — safe to call
* more than once — and must never throw, matching {@link IdleSleepGuard}'s "never break the
* fleet" contract.
*/
public interface SleepAssertion extends AutoCloseable {
@Override
void close();
}
@@ -0,0 +1,21 @@
package dev.ltms.fleet.power;
/**
* The OS mechanism {@link IdleSleepGuard} uses to hold and release an idle-sleep assertion. This
* is the seam a test exercises instead of the real effect (a live {@code caffeinate} child) — see
* {@code IdleSleepGuardTest}.
*
* <p>Implementations must never throw. Every failure — wrong platform, missing tool, a spawn
* error — must show up as {@link #acquire()} returning {@code null}, so a caller can treat "no
* assertion held" and "the mechanism could not be used" identically and the fleet keeps running
* either way.
*/
public interface SleepAssertionMechanism {
/**
* Acquire a fresh assertion against idle sleep, or {@code null} when this mechanism is not
* usable right now (wrong platform, the tool is missing, the child process could not start).
* Never throws.
*/
SleepAssertion acquire();
}
@@ -108,6 +108,7 @@ class ConfigRefTopLevelReportingCoverageTest {
v.put("worktreeGroup", "group-a");
v.put("memberLoginShell", null);
v.put("memberSkills", "/skills/a");
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(true));
assertNamesMatchComponents(v);
return v;
}
@@ -149,6 +150,7 @@ class ConfigRefTopLevelReportingCoverageTest {
v.put("worktreeGroup", "group-b");
v.put("memberLoginShell", null);
v.put("memberSkills", "/skills/b");
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(false));
assertNamesMatchComponents(v);
return v;
}
@@ -2629,4 +2629,58 @@ class FleetConfigTest {
"with no pool to choose from, every configured profile is a candidate and the "
+ "first one wins");
}
// ── idle-sleep guard: default-on config block ───────────────────────────────────────────────
@Test
void idleSleepGuardIsOnByDefaultWhenTheBlockIsEntirelyAbsent(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8080
""");
FleetConfig cfg = FleetConfig.load(f);
assertNull(cfg.idleSleepGuard(), "an absent block parses to null, unlike most other blocks here");
// The block itself is absent, but the FEATURE stays on: Fleetd treats a null block the
// same as enabled: true (see FleetConfig.idleSleepGuard's javadoc) — this test only pins
// the parse result, the on-by-default behaviour is Fleetd's own null check.
}
@Test
void idleSleepGuardExplicitlyEnabledIsOn(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
idleSleepGuard:
enabled: true
""");
FleetConfig cfg = FleetConfig.load(f);
assertTrue(cfg.idleSleepGuard().isEnabled());
}
@Test
void idleSleepGuardExplicitlyDisabledIsOff(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
idleSleepGuard:
enabled: false
""");
FleetConfig cfg = FleetConfig.load(f);
assertFalse(cfg.idleSleepGuard().isEnabled());
}
@Test
void idleSleepGuardBlockPresentButEmptyDefaultsToEnabled(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
idleSleepGuard: {}
""");
FleetConfig cfg = FleetConfig.load(f);
assertTrue(cfg.idleSleepGuard().isEnabled(),
"unlike ConfigReload/Health, this block defaults to ON even when present but empty");
}
}
@@ -46,7 +46,7 @@ import static org.junit.jupiter.api.Assertions.assertEquals;
* comments document that it only ever REPLACES a component when the incoming value is {@code null}
* (or blank, for {@code placement}) — {@code broker}/{@code primary}/{@code leadHeartbeat}/
* {@code configReload}/{@code coordinator}/{@code worktreeGroup}/{@code memberLoginShell}/
* {@code memberSkills} are left as-is unconditionally, and {@code bind}/{@code guard}/{@code lifecycle}/{@code auth}/
* {@code memberSkills}/{@code idleSleepGuard} are left as-is unconditionally, and {@code bind}/{@code guard}/{@code lifecycle}/{@code auth}/
* {@code fleet}/{@code quarantineCooldownSeconds}/{@code memberCredentials}/{@code placement} are
* replaced only on null/blank input. A value that is never null or blank going in must therefore
* never change coming out, for every current component. No exclusion is needed today.
@@ -96,6 +96,7 @@ class FleetConfigWithDefaultsPreservesEveryComponentTest {
v.put("worktreeGroup", "group-guard");
v.put("memberLoginShell", "/bin/zsh");
v.put("memberSkills", "/skills/guard");
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(true));
assertNamesMatchComponents(v);
return v;
}
@@ -0,0 +1,232 @@
package dev.ltms.fleet.deploy;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.List;
import java.util.regex.Pattern;
import java.util.stream.Stream;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.params.ParameterizedTest;
import org.junit.jupiter.params.provider.MethodSource;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #360: {@code deploy/fleetd.service} started clean on fleet01 and broke the daemon in
* three ways that nothing logs at startup.
*
* <ul>
* <li>{@code ProtectSystem=}, {@code ProtectHome=}, {@code ProtectKernelTunables=} and
* {@code ProtectControlGroups=} each give the unit its own mount namespace. fleetd resolves
* a caller's role by running {@code lsof} to find the loopback peer PID (see
* {@code mcp/LsofPeerPidLookup}). Inside such a namespace {@code lsof} returns nothing,
* every caller falls back to ANONYMOUS, and the primary is refused every orchestration call
* with "unauthenticated: anonymous may not SPAWN" -- while healthz still reports ok.
* Measured by bisection on fleet01 2026-09-05 (lsof line count): no sandbox 3,
* {@code ProtectSystem=strict} 0, {@code ProtectHome=read-only} 0,
* {@code ProtectKernelTunables} 0, {@code ProtectControlGroups} 0,
* {@code RestrictSUIDSGID} 3, {@code NoNewPrivileges} 3 -- so only the first four are
* forbidden.
* <li>{@code PrivateTmp=true} gives the unit its own {@code /tmp}. fleetd writes the member
* ZDOTDIR credential-scrub directory and the opencode config directory under
* {@code java.io.tmpdir}, and the member pane -- a child of a *different* unit
* (herdr.service) -- has to read them back. A private {@code /tmp} on either unit turns the
* credential scrub into a silent no-op.
* <li>Running {@code java} directly from {@code ExecStart} skips the login shell. Every secret
* fleetd needs (AI_GATEWAY_TOKEN, WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI) lives in
* a file only the login shell sources; systemd runs no login shell on its own. Started that
* way the daemon boots fine with empty credentials, and the failure appears hours later as a
* member that cannot open a pull request.
* </ul>
*
* <p>Review round 2 on fleetd #360 found the same shape one file further down the chain:
* {@code herdr.service}'s {@code ExecStart} only names {@code deploy/herdr-inner.sh}, so that
* script is the ONLY place two more of these properties live, and nothing was reading it.
*
* <ul>
* <li>If the script does not exec a login shell, herdr -- and every member pane it spawns as a
* child -- starts with none of the secrets only a login shell sources, the same silent
* empty-credentials failure as fleetd.service's {@code ExecStart}, one process further away.
* <li>If the script does not set a real, non-zero pty size before starting herdr (with
* {@code stty}), herdr reports a 0x0 window and every pane spawn fails with
* "ghostty error -2" -- which surfaces as a fleetd spawn failure, not a herdr one.
* </ul>
*
* <p>None of these five failures makes the unit (or the script) fail to start, or makes
* {@code /healthz} report unhealthy, so nothing short of reading the files catches a regression.
* This test is that read.
*
* <p><b>It checks source text, not behaviour</b> -- it cannot start systemd or fork an mount
* namespace in a build sandbox. It parses the same two files an operator would install and fails
* if a forbidden directive is active, exactly the way {@code McpContractDocTest} guards
* {@code docs/MCP-Contract.md} against naming a tool that does not exist.
*/
class SystemdUnitSafetyTest {
/** Tests run with the module directory (fleetd/) as cwd; deploy/ is the repo-root sibling. */
private static final Path FLEETD_SERVICE = Path.of("../deploy/fleetd.service");
private static final Path HERDR_SERVICE = Path.of("../deploy/herdr.service");
private static final Path HERDR_INNER_SCRIPT = Path.of("../deploy/herdr-inner.sh");
private static final List<String> NAMESPACING_DIRECTIVES = List.of(
"ProtectSystem", "ProtectHome", "ProtectKernelTunables", "ProtectControlGroups");
/**
* A shell invoked with a login flag, e.g. {@code zsh -lc '...'} or {@code /bin/sh -l}. The
* property under test is the {@code -l}, not the specific shell or the rest of its flags, so
* this matches a token ending in "sh" followed by a flag cluster containing "l" -- tighter
* than a bare {@code contains("-lc")}, which a "-lc" anywhere in the line would also satisfy.
*/
private static final Pattern LOGIN_SHELL_INVOCATION =
Pattern.compile("(?:^|\\s)\\S*sh\\s+-[A-Za-z]*l[A-Za-z]*\\b");
private static final Pattern POSITIVE_ROWS = Pattern.compile("\\brows\\s+([1-9]\\d*)\\b");
private static final Pattern POSITIVE_COLS = Pattern.compile("\\bcols\\s+([1-9]\\d*)\\b");
static Stream<Path> bothUnits() {
return Stream.of(FLEETD_SERVICE, HERDR_SERVICE);
}
private static boolean invokesALoginShell(List<String> activeLines) {
return activeLines.stream().anyMatch(l -> LOGIN_SHELL_INVOCATION.matcher(l).find());
}
/** An active {@code stty} line naming both a positive row count and a positive column count. */
private static boolean setsANonZeroPtySize(List<String> activeLines) {
return activeLines.stream()
.filter(l -> l.startsWith("stty"))
.anyMatch(l -> POSITIVE_ROWS.matcher(l).find() && POSITIVE_COLS.matcher(l).find());
}
/** Lines that are actually in force: comments and blank lines don't count. */
private static List<String> activeLines(Path unit) throws Exception {
List<String> active = new ArrayList<>();
for (String line : Files.readAllLines(unit)) {
String stripped = line.strip();
if (!stripped.isEmpty() && !stripped.startsWith("#")) {
active.add(stripped);
}
}
return active;
}
@ParameterizedTest
@MethodSource("bothUnits")
@DisplayName("[SOURCE TEXT] no active mount-namespacing directive -- it blinds lsof and turns every caller ANONYMOUS")
void doesNotActivateAMountNamespace(Path unit) throws Exception {
List<String> active = activeLines(unit);
for (String directive : NAMESPACING_DIRECTIVES) {
Pattern activeDirective = Pattern.compile("^" + Pattern.quote(directive) + "\\s*=");
List<String> hits = active.stream().filter(l -> activeDirective.matcher(l).find()).toList();
assertTrue(hits.isEmpty(),
unit + " sets " + directive + " (" + hits + "). That directive gives the unit its "
+ "own mount namespace; inside it, fleetd's lsof-based caller lookup "
+ "(mcp/LsofPeerPidLookup) returns nothing, so every MCP caller falls back to "
+ "ANONYMOUS and the primary is refused every orchestration call with "
+ "\"unauthenticated: anonymous may not SPAWN\" -- while the daemon still "
+ "starts and /healthz still reports ok. Measured on fleet01 2026-09-05 "
+ "(fleetd #360). A commented-out mention in the file's own DO-NOT-add block "
+ "is fine; an active directive is not.");
}
}
@ParameterizedTest
@MethodSource("bothUnits")
@DisplayName("[SOURCE TEXT] PrivateTmp is not true -- it silently no-ops the credential scrub")
void privateTmpIsNotTrue(Path unit) throws Exception {
List<String> active = activeLines(unit);
Pattern privateTmpTrue = Pattern.compile("^PrivateTmp\\s*=\\s*true\\b");
boolean hasPrivateTmpTrue = active.stream().anyMatch(l -> privateTmpTrue.matcher(l).find());
assertFalse(hasPrivateTmpTrue,
unit + " sets PrivateTmp=true. fleetd writes the member ZDOTDIR credential-scrub "
+ "directory and the opencode config directory under java.io.tmpdir, and the "
+ "member pane -- a child of a DIFFERENT unit -- has to read them back. A "
+ "private /tmp on either unit turns the credential scrub into a silent no-op: "
+ "no error, no log line, the scrub just never happens (fleetd #360).");
}
@Test
@DisplayName("[SOURCE TEXT] fleetd.service's ExecStart goes through a login shell -- otherwise every secret is empty")
void execStartUsesALoginShell() throws Exception {
List<String> active = activeLines(FLEETD_SERVICE);
List<String> execStartLines = active.stream().filter(l -> l.startsWith("ExecStart=")).toList();
assertTrue(execStartLines.size() == 1,
FLEETD_SERVICE + " must have exactly one active ExecStart= line; found "
+ execStartLines.size() + ": " + execStartLines);
String execStart = execStartLines.get(0);
assertTrue(LOGIN_SHELL_INVOCATION.matcher(execStart).find(),
FLEETD_SERVICE + "'s ExecStart (" + execStart + ") does not run a login shell (a shell "
+ "invoked with a \"-l\" flag, e.g. \"zsh -lc\"). Every secret fleetd needs "
+ "(AI_GATEWAY_TOKEN, WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI) lives in a "
+ "file only the login shell sources; systemd runs no login shell on its own. "
+ "Running java directly boots fine with every credential empty, and the failure "
+ "surfaces hours later as a member that cannot open a pull request (fleetd #360).");
}
@Test
@DisplayName("[SOURCE TEXT] herdr-inner.sh execs a login shell -- otherwise herdr and every member it spawns start with empty credentials")
void herdrInnerScriptUsesALoginShell() throws Exception {
List<String> active = activeLines(HERDR_INNER_SCRIPT);
assertTrue(invokesALoginShell(active),
HERDR_INNER_SCRIPT + " does not exec a login shell (a shell invoked with a \"-l\" flag, "
+ "e.g. \"zsh -lc\"). herdr.service's ExecStart only names this script, so this "
+ "is the ONLY place herdr's login-shell property lives. Without it, herdr -- and "
+ "every member pane it spawns as a child of herdr -- starts with none of the "
+ "secrets that only a login shell sources (this host: ~/.fleet/secrets.sh via "
+ "~/.zprofile), and the failure surfaces hours later as a member with no "
+ "credentials at all, not just fleetd (fleetd #360).");
}
@Test
@DisplayName("[SOURCE TEXT] herdr-inner.sh sets a non-zero pty size before starting herdr -- otherwise every pane spawn fails with \"ghostty error -2\"")
void herdrInnerScriptSetsANonZeroPtySize() throws Exception {
List<String> active = activeLines(HERDR_INNER_SCRIPT);
assertTrue(setsANonZeroPtySize(active),
HERDR_INNER_SCRIPT + " does not run \"stty rows N cols M\" with both N and M positive "
+ "before starting herdr. Without a real pty size, herdr reports a 0x0 window and "
+ "every pane spawn fails with \"ghostty error -2\" -- which surfaces as a fleetd "
+ "spawn failure, not a herdr one, and gives no hint that the actual cause is this "
+ "script (fleetd #360).");
}
/**
* The denominator guard, same shape as {@code McpContractDocTest}'s: a check that scans for a
* forbidden pattern passes trivially if it is handed nothing to scan. Pin that all three files
* exist, are non-trivial, and that the DO-NOT block's commented mentions are still there -- so
* the "comment survives, directive doesn't" distinction above is actually being exercised.
*/
@Test
@DisplayName("[SOURCE TEXT] the namespacing check is not vacuous -- all three files exist and the DO-NOT block still mentions every forbidden directive in a comment")
void theCheckActuallyHasSomethingToCheck() throws Exception {
assertTrue(Files.size(FLEETD_SERVICE) > 200,
FLEETD_SERVICE + " is missing or unexpectedly small -- the checks above would pass "
+ "vacuously against an empty or absent file");
assertTrue(Files.size(HERDR_SERVICE) > 200,
HERDR_SERVICE + " is missing or unexpectedly small -- the checks above would pass "
+ "vacuously against an empty or absent file");
assertTrue(Files.size(HERDR_INNER_SCRIPT) > 50,
HERDR_INNER_SCRIPT + " is missing or unexpectedly small -- the login-shell and "
+ "pty-size checks above would either pass vacuously or fail with an opaque "
+ "\"file not found\" instead of naming the actual consequence (empty "
+ "credentials / \"ghostty error -2\") against an empty or absent file");
String fleetdService = Files.readString(FLEETD_SERVICE);
for (String directive : NAMESPACING_DIRECTIVES) {
assertTrue(fleetdService.contains(directive),
FLEETD_SERVICE + " no longer mentions " + directive + " anywhere, not even in the "
+ "DO-NOT-add comment block that explains why it must stay out. That comment "
+ "is the whole point of fleetd #360 -- it is what stops the next edit from "
+ "re-adding the directive without knowing why.");
}
}
}
@@ -4,6 +4,7 @@ import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
@@ -44,6 +45,7 @@ class ReplyPushLoopTest {
private static final String WORKER = "term_worker";
private static final String WORKER2 = "term_worker2";
private static final String OTHER_PRIMARY = "term_other_primary";
private static final String DEAD_LEAD = "term_dead_lead";
private static final ObjectMapper MAPPER = new ObjectMapper();
private PrimaryRegistry registry;
@@ -207,6 +209,101 @@ class ReplyPushLoopTest {
"exactly " + cap + " agent.prompt calls (cap=" + cap + ")");
}
// --- fleetd #368: a lead's delegation binding must not outlive the lead ---------------------
/**
* The bug: {@code PrimaryRegistry.forgetDelegation} is wired to a worker's release, never to
* the delegating lead's own disappearance, so a lead that closed, crashed, or was relaunched
* leaves {@code leadByTarget} pointing at a terminal herdr no longer knows. Before the fix,
* {@code onReplyQueued} took that stale, non-null entry at face value — {@code nudgeTargetFor}
* only ever falls back to the pinned primary when the map holds nothing for the target — so
* the nudge's only schedule ran against the dead terminal forever and the live primary never
* heard about the reply through this path.
*
* <p>This drives {@link ReplyPushLoop#onReplyQueued(String)} itself (not {@code PrimaryRegistry}
* directly), because the registry lookup was never the defect — the caller trusting it without
* checking liveness was. A test that only asserted on {@code PrimaryRegistry.nudgeTargetFor}
* would pass whether or not {@code ReplyPushLoop} ever adopted the fix.
*/
@Test
void aStaleLeadBindingFallsBackToTheLiveLeadInsteadOfNudgingADeadTerminal() throws Exception {
// PRIMARY is the single known (pinned) lead — set up in @BeforeEach via `registry`.
// DEAD_LEAD is a second lead that once delegated to WORKER and is now gone: herdr reports
// agent_not_found for it, exactly as it would for a closed/crashed/relaunched terminal.
registry.recordDelegation(WORKER, DEAD_LEAD);
var rec = new DeadLeadHerdrClient(DEAD_LEAD);
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
loop(1, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"the nudge should still reach the live primary, not silently vanish with the dead lead");
assertEquals(List.of(PRIMARY), rec.promptTargets(),
"the nudge must be sent to the live primary, never to the dead lead's terminal");
assertEquals(PRIMARY, registry.nudgeTargetFor(WORKER).orElseThrow(),
"the stale binding must be forgotten (self-healed) once found dead, exactly like "
+ "AgentControl.paneByTerminal already does on agent_not_found");
}
/**
* Same dead binding, but with no pinned primary to fall back to (the multi-lead, no-fallback
* case {@code PrimaryRegistry.nudgeTargetFor}'s own javadoc already covers): the loop must
* never nudge the dead terminal, and must not spin — no schedule starts at all once the stale
* binding resolves to empty, same as if the map had never held an entry for this target.
*/
@Test
void aStaleLeadBindingWithNoFallbackNeverNudgesTheDeadTerminal() throws Exception {
var unpinned = new PrimaryRegistry(null);
unpinned.recordDelegation(WORKER, DEAD_LEAD);
var rec = new DeadLeadHerdrClient(DEAD_LEAD);
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
var loop = new ReplyPushLoop(unpinned, agents, inbox, scheduler, 1, 50);
loop.onReplyQueued(WORKER);
Thread.sleep(200);
assertEquals(0, rec.sendCount(), "no lead is live to nudge, so nothing should ever be sent");
assertTrue(unpinned.nudgeTargetFor(WORKER).isEmpty(),
"the stale binding must be forgotten even when there is no fallback to hand back");
}
/**
* fleetd #368 review, must-fix: the first version of {@code isLive} treated <em>any</em>
* {@code RuntimeException} from the liveness probe as "the lead is gone" — indistinguishable
* from a transient herdr hiccup (a socket blip, a decode error) on a lead that is actually
* still live. The consequence of that misdiagnosis is destructive and permanent
* ({@code forgetDelegation}), which is the exact #359 mistake repeated two days later: a single
* bad reading must never destroy a live binding. This pins the narrower rule — only an
* affirmative {@code agent_not_found} may forget a binding; a merely inconclusive failure must
* leave the binding alone, and the lead must still be nudged once the probe recovers.
*/
@Test
void aTransientLivenessFailureMustNotForgetABindingToAStillLiveLead() throws Exception {
// OTHER_PRIMARY is delegated to and genuinely live — its FIRST agent.get call fails with a
// transient, non-agent_not_found HerdrException (a transport-level failure, code null,
// exactly what a socket blip looks like), then succeeds on every call after.
registry.recordDelegation(WORKER, OTHER_PRIMARY);
var rec = new FlakyThenLiveHerdrClient(OTHER_PRIMARY);
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
loop(2, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"the nudge must still reach the live lead once the transient failure clears");
assertEquals(List.of(OTHER_PRIMARY), rec.promptTargets(),
"the nudge must go to the lead that was only transiently unreachable, not the "
+ "unrelated pinned primary");
assertEquals(OTHER_PRIMARY, registry.nudgeTargetFor(WORKER).orElseThrow(),
"a merely transient failure must not forget the binding to a lead that is actually "
+ "still live");
}
// --- nudge format --------------------------------------------------------------------------
@Test
@@ -1098,4 +1195,101 @@ class ReplyPushLoopTest {
public void close() {
}
}
/**
* Fake herdr client for fleetd #368: {@code deadTarget} is a terminal herdr genuinely no
* longer knows about — {@code agent.get} fails with {@code agent_not_found} exactly as
* {@code AgentControl.agentCall} expects for a real dead/closed pane (see its javadoc). Every
* other target reports {@code idle} (injectable). Records the {@code target} named by every
* {@code agent.prompt} call, so a test can prove which terminal actually got nudged.
*/
private static final class DeadLeadHerdrClient implements HerdrClient {
private final String deadTarget;
private final List<String> promptTargets = Collections.synchronizedList(new ArrayList<>());
volatile CountDownLatch sendLatch = new CountDownLatch(1);
DeadLeadHerdrClient(String deadTarget) {
this.deadTarget = deadTarget;
}
@Override
@SuppressWarnings("unchecked")
public JsonNode call(String method, Object params) {
Map<String, Object> p = params instanceof Map ? (Map<String, Object>) params : Map.of();
if ("agent.get".equals(method)) {
String target = String.valueOf(p.get("target"));
if (deadTarget.equals(target)) {
throw new HerdrException("no such agent: " + target, "agent_not_found", null);
}
return MAPPER.createObjectNode()
.set("agent", MAPPER.createObjectNode()
.put("terminal_id", target)
.put("agent_status", "idle"));
}
if ("agent.prompt".equals(method)) {
promptTargets.add(String.valueOf(p.get("target")));
sendLatch.countDown();
}
return MAPPER.createObjectNode();
}
List<String> promptTargets() {
return List.copyOf(promptTargets);
}
long sendCount() {
return promptTargets.size();
}
@Override
public void close() {
}
}
/**
* Fake herdr client for fleetd #368 review: {@code flakyTarget}'s FIRST {@code agent.get} call
* fails with a transient, non-{@code agent_not_found} {@code HerdrException} — a transport-level
* failure (code {@code null}), exactly what a socket blip or a decode error on a perfectly live
* lead looks like — then succeeds ({@code idle}) on every call after. Used to prove a merely
* inconclusive failure must not be treated as the lead being gone.
*/
private static final class FlakyThenLiveHerdrClient implements HerdrClient {
private final String flakyTarget;
private final AtomicInteger getCalls = new AtomicInteger();
private final List<String> promptTargets = Collections.synchronizedList(new ArrayList<>());
volatile CountDownLatch sendLatch = new CountDownLatch(1);
FlakyThenLiveHerdrClient(String flakyTarget) {
this.flakyTarget = flakyTarget;
}
@Override
@SuppressWarnings("unchecked")
public JsonNode call(String method, Object params) {
Map<String, Object> p = params instanceof Map ? (Map<String, Object>) params : Map.of();
if ("agent.get".equals(method)) {
String target = String.valueOf(p.get("target"));
if (flakyTarget.equals(target) && getCalls.getAndIncrement() == 0) {
throw new HerdrException("herdr socket read timed out"); // transport failure, code == null
}
return MAPPER.createObjectNode()
.set("agent", MAPPER.createObjectNode()
.put("terminal_id", target)
.put("agent_status", "idle"));
}
if ("agent.prompt".equals(method)) {
promptTargets.add(String.valueOf(p.get("target")));
sendLatch.countDown();
}
return MAPPER.createObjectNode();
}
List<String> promptTargets() {
return List.copyOf(promptTargets);
}
@Override
public void close() {
}
}
}
@@ -0,0 +1,54 @@
package dev.ltms.fleet.power;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Platform-detection unit tests for {@link CaffeinateSleepAssertionMechanism}.
*
* <p>This deliberately never calls {@link CaffeinateSleepAssertionMechanism#acquire()} itself —
* doing so on a real macOS machine would actually start a live {@code caffeinate} child and hold
* a real idle-sleep assertion, which the ticket this class exists for explicitly forbids testing
* with. Instead this exercises the pure {@code isSupportedPlatform(String)} predicate that
* {@code acquire()} consults before ever touching {@link ProcessBuilder} — so it proves the
* platform check itself is correct on any CI OS, but it does <strong>not</strong> prove that a
* real {@code caffeinate -i} spawn succeeds or that its child is torn down correctly; that half is
* exercised indirectly by {@link IdleSleepGuardTest} against a {@link FakeSleepAssertionMechanism}
* instead, which is the seam invariant 2/3 in the ticket call for.
*/
class CaffeinateSleepAssertionMechanismTest {
@Test
void macOsNamesAreSupported() {
assertTrue(CaffeinateSleepAssertionMechanism.isSupportedPlatform("Mac OS X"));
assertTrue(CaffeinateSleepAssertionMechanism.isSupportedPlatform("macOS"));
assertTrue(CaffeinateSleepAssertionMechanism.isSupportedPlatform("MAC OS X"));
}
@Test
void nonMacNamesAreNotSupported() {
assertFalse(CaffeinateSleepAssertionMechanism.isSupportedPlatform("Linux"));
assertFalse(CaffeinateSleepAssertionMechanism.isSupportedPlatform("Windows 11"));
}
@Test
void nullOsNameIsNotSupported() {
assertFalse(CaffeinateSleepAssertionMechanism.isSupportedPlatform(null));
}
/**
* The overload {@code isSupportedPlatform()} (no args) reads the JVM's real {@code os.name} —
* proves the wiring is live, without asserting a specific answer (this suite itself must pass
* on both macOS and Linux CI).
*/
@Test
void noArgOverloadReadsRealSystemProperty() {
boolean expected = CaffeinateSleepAssertionMechanism
.isSupportedPlatform(System.getProperty("os.name"));
boolean actual = CaffeinateSleepAssertionMechanism.isSupportedPlatform();
assertEquals(expected, actual);
}
}
@@ -0,0 +1,56 @@
package dev.ltms.fleet.power;
import java.util.concurrent.CopyOnWriteArrayList;
import java.util.concurrent.atomic.AtomicInteger;
/**
* Recording fake {@link SleepAssertionMechanism} — the seam behind the real OS effect (a live
* {@code caffeinate} child process). No test in this package ever spawns that real process; every
* assertion here is against this fake's own call log instead.
*
* <p>Each acquired {@link FakeAssertion} records its own {@code close()} calls, and every
* acquired instance is kept in {@link #acquired} so a test can inspect all of them, including
* ones {@link IdleSleepGuard} has already released.
*/
final class FakeSleepAssertionMechanism implements SleepAssertionMechanism {
/** Every {@link FakeAssertion} this mechanism has ever handed out, in order. */
final CopyOnWriteArrayList<FakeAssertion> acquired = new CopyOnWriteArrayList<>();
private final AtomicInteger acquireCalls = new AtomicInteger();
private volatile boolean unavailable = false;
/** Make the next (and every subsequent) {@link #acquire()} return {@code null}, like a missing tool. */
void makeUnavailable() {
unavailable = true;
}
int acquireCallCount() {
return acquireCalls.get();
}
@Override
public SleepAssertion acquire() {
acquireCalls.incrementAndGet();
if (unavailable) {
return null;
}
FakeAssertion a = new FakeAssertion();
acquired.add(a);
return a;
}
/** A held fake assertion; records how many times {@code close()} was actually called. */
static final class FakeAssertion implements SleepAssertion {
private final AtomicInteger closeCalls = new AtomicInteger();
int closeCallCount() {
return closeCalls.get();
}
@Override
public void close() {
closeCalls.incrementAndGet();
}
}
}
@@ -0,0 +1,118 @@
package dev.ltms.fleet.power;
import org.junit.jupiter.api.Test;
import java.util.concurrent.atomic.AtomicInteger;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* {@link IdleSleepGuard} against a {@link FakeSleepAssertionMechanism} — the seam that stands in
* for a real {@code caffeinate} child process. No test in this class ever spawns a real OS
* process or asserts against real idle sleep; every assertion is against the fake's call log
* (how many times {@code acquire()}/{@code close()} were actually called). That proves the
* <em>orchestration</em> — when the guard decides to hold or release an assertion, and that it
* never throws — but it does <strong>not</strong> prove that {@code caffeinate -i} itself
* actually stops macOS from idle-sleeping; that half is outside what a unit test can safely
* exercise (see {@link CaffeinateSleepAssertionMechanismTest}'s class doc).
*/
class IdleSleepGuardTest {
@Test
void acquiresOnZeroToOneAndReleasesOnOneToZero() {
FakeSleepAssertionMechanism mechanism = new FakeSleepAssertionMechanism();
AtomicInteger liveCount = new AtomicInteger(0);
IdleSleepGuard guard = new IdleSleepGuard(mechanism, liveCount::get);
assertFalse(guard.isHeld(), "nothing held before any member is live");
liveCount.set(1);
guard.recheck();
assertTrue(guard.isHeld(), "an assertion must be held once a member is live");
assertEquals(1, mechanism.acquired.size());
assertEquals(0, mechanism.acquired.get(0).closeCallCount());
liveCount.set(0);
guard.recheck();
assertFalse(guard.isHeld(), "the assertion must be released once the last member goes");
assertEquals(1, mechanism.acquired.get(0).closeCallCount(), "the SAME held assertion must be closed");
}
@Test
void steadyLiveCountDoesNotReacquireOrRerelease() {
FakeSleepAssertionMechanism mechanism = new FakeSleepAssertionMechanism();
AtomicInteger liveCount = new AtomicInteger(2);
IdleSleepGuard guard = new IdleSleepGuard(mechanism, liveCount::get);
guard.recheck(); // 0 -> 2 crossing: acquires
guard.recheck(); // still 2: must be a no-op
guard.recheck(); // still 2: must be a no-op
assertEquals(1, mechanism.acquireCallCount(), "only the crossing touches the mechanism");
liveCount.set(1); // 2 -> 1: still > 0, still a no-op
guard.recheck();
assertTrue(guard.isHeld());
assertEquals(0, mechanism.acquired.get(0).closeCallCount());
assertEquals(1, mechanism.acquireCallCount());
}
/**
* Invariant 2: a missing/unavailable mechanism must never throw, and the guard must simply
* hold nothing. {@link FakeSleepAssertionMechanism#makeUnavailable()} makes {@code acquire()}
* return {@code null}, exactly like {@link CaffeinateSleepAssertionMechanism} does off macOS
* or when the {@code caffeinate} binary is missing.
*/
@Test
void unavailableMechanismNeverThrowsAndHoldsNothing() {
FakeSleepAssertionMechanism mechanism = new FakeSleepAssertionMechanism();
mechanism.makeUnavailable();
AtomicInteger liveCount = new AtomicInteger(1);
IdleSleepGuard guard = new IdleSleepGuard(mechanism, liveCount::get);
guard.recheck(); // must not throw
assertFalse(guard.isHeld(), "acquire() returned null, so nothing is held");
assertEquals(1, mechanism.acquireCallCount());
// still must not throw or leak on release, even though nothing was ever actually held
liveCount.set(0);
guard.recheck();
assertFalse(guard.isHeld());
guard.close(); // teardown with nothing held must also be a safe no-op
}
/**
* Invariant 3 (teardown). This is the test the mutation testing step removes the production
* release call to fail: with {@code releaseHeldLocked()} not invoked from {@link
* IdleSleepGuard#close()}, the held fake assertion's {@code close()} would never be called and
* this assertion would fail.
*/
@Test
void closeReleasesAHeldAssertionEvenWithoutAZeroCrossing() {
FakeSleepAssertionMechanism mechanism = new FakeSleepAssertionMechanism();
AtomicInteger liveCount = new AtomicInteger(1);
IdleSleepGuard guard = new IdleSleepGuard(mechanism, liveCount::get);
guard.recheck();
assertTrue(guard.isHeld());
guard.close();
assertFalse(guard.isHeld(), "close() must release whatever is held, independent of live count");
assertEquals(1, mechanism.acquired.get(0).closeCallCount());
}
@Test
void closeIsIdempotent() {
FakeSleepAssertionMechanism mechanism = new FakeSleepAssertionMechanism();
AtomicInteger liveCount = new AtomicInteger(1);
IdleSleepGuard guard = new IdleSleepGuard(mechanism, liveCount::get);
guard.recheck();
guard.close();
guard.close(); // must not throw, must not double-release
assertEquals(1, mechanism.acquired.get(0).closeCallCount());
}
}
@@ -0,0 +1,72 @@
package dev.ltms.fleet.power;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.SessionManager;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Proves the wiring {@code Fleetd.main} actually performs — {@code
* sessions.onAcquire(_ -> guard.recheck())} / {@code sessions.onRelease(_ -> guard.recheck())} —
* not just {@link IdleSleepGuard}'s own orchestration logic in isolation
* ({@link IdleSleepGuardTest} already covers that in isolation, which on its own would not catch
* a wiring gap — e.g. an {@code onAcquire} call typo'd to a no-op lambda, or the listener wired to
* the wrong SessionManager instance — see fleetd's own "a test on the seam does not prove the
* caller" lesson). This test builds a real {@link SessionManager} exactly as
* {@code SessionManagerTest} does (a {@link FakeHerdr}-backed {@link ClaudeCodeLauncher}, no live
* herdr process), wires it to an {@link IdleSleepGuard} the same two lines {@code Fleetd.main}
* uses, and drives real {@link SessionManager#acquire} / {@link SessionManager#release} calls.
*/
class IdleSleepGuardWiringTest {
private SessionManager sessionManager(FakeHerdr herdr) {
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
return new SessionManager(workers);
}
@Test
void acquiringAndReleasingRealSessionsDrivesTheGuardThroughTheSameWiringFleetdUses() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
FakeSleepAssertionMechanism mechanism = new FakeSleepAssertionMechanism();
IdleSleepGuard guard = new IdleSleepGuard(mechanism, sessions::size);
// The exact two lines Fleetd.main wires up.
sessions.onAcquire(_ -> guard.recheck());
sessions.onRelease(_ -> guard.recheck());
assertFalse(guard.isHeld(), "no member yet: nothing held");
MemberSession a = sessions.acquire("ltms-local", "/a", "/caller", "ownerA");
assertTrue(guard.isHeld(), "0 -> 1: the first live member must arm the guard");
MemberSession b = sessions.acquire("ltms-local", "/b", "/caller", "ownerB");
assertEquals(1, mechanism.acquireCallCount(), "2nd member: still just 1 live-to-2 step, no new acquire");
sessions.release(a.paneId());
assertTrue(guard.isHeld(), "one member still live: the guard must stay armed");
assertEquals(0, mechanism.acquired.get(0).closeCallCount());
sessions.release(b.paneId());
assertFalse(guard.isHeld(), "1 -> 0: the last member releasing must disarm the guard");
assertEquals(1, mechanism.acquired.get(0).closeCallCount());
}
}
@@ -1529,7 +1529,15 @@ class GitWorktreesTest {
* an empty, machine-independent {@code XDG_CONFIG_HOME} (so the fallback resolves to a file that
* provably does not exist) plus the same {@code GIT_CONFIG_GLOBAL}/{@code GIT_CONFIG_SYSTEM}/
* {@code GIT_TERMINAL_PROMPT} isolation the {@link #git}/{@link #gitOutput} helpers already use
* for repo setup — so no test in this class can reach the real machine's home directory.
* for repo setup.
*
* <p>Scope, measured on the fleetd #369 merge and narrower than an earlier version of this
* comment claimed: this protects the 5 {@link #seedingGitWorktrees} sites plus — through
* {@link #gitProcessBuilder} — every {@code git} subprocess the TEST itself starts. It does
* NOT cover the other 53 {@code new GitWorktrees(...)} constructions in this file, which pass
* no env override, so a production instance built that way still inherits the JVM's real
* environment. Stripping this override from {@code seedingGitWorktrees} leaves the class green
* both with and without the poison command above, so that half is currently unpinned.
*/
private static Map<String, String> hermeticGitEnv(Path tmp) {
return Map.of(