Lead-verified: merged with #194 onto an integration branch off main, full `mvn clean install` green at 1025 tests.
Review round 2 fixed the stale ambiguous-pane message and, more importantly, a real abort: probeOwner called list() unguarded, so one unreachable daemon made panes on a different healthy daemon un-stoppable too — the same bug blocker 1 exists to fix, through a new door. Worker proved it by removing the guard and quoting the failure.
Lead-verified: merged with #196 onto an integration branch off main, full `mvn clean install` green at 1025 tests. Deny-by-default WARN confirmed byte-identical to main.
Review round 2 fixed a defect I found in round 1: the new INFO asserted that gap names were not on the derived allow-list without ever checking, so a name derived from a profile's tokenEnv/gitTokenEnv/env: would be reported as safe while the member actually inherited it. Now split with MemberEnvAllowList.keeps — the same predicate the generated scrub evaluates.
Lead review of PR #196 found two issues in CompositePeerLauncher.probeOwner:
1. The "more than one daemon claims this pane" throw kept the old
pre-fix message ("no owning herdr daemon was recorded"), which was
only true of the code it replaced. Reworded to say what actually
happened: N configured herdr daemons report this pane, so it is
genuinely ambiguous. Updated the one test pinning the old string.
2. probeOwner let list() propagate straight out of the probe loop, so
one unreachable daemon aborted the whole probe and made a pane on a
DIFFERENT, healthy daemon un-stoppable too — resurrecting the exact
bug blocker 1 fixes. Now catches HerdrException per daemon, logs the
exception class only, and treats that daemon as not knowing the pane
so probing continues. New test proves this: verified it fails with
the try/catch removed (HerdrException propagates and the stop that
should succeed via the healthy daemon throws instead), then restored.
Full mvn clean install: 1019 tests, 0 failures, 0 errors.
Lead review on PR #194 found that the allow-list INFO wording claimed the
whole gap ("credential-shaped names on neither known: nor allow:") is blanked
by the scrub, without checking that against effectiveAllowed. effectiveAllowed
is a SUPERSET of known+allow — MemberEnvAllowList.derive also unions in every
profile's gitTokenEnv/gitHostEnv/tokenEnv/env: keys, and derivedAllowedNames
further unions in the spawn's own env keys — so a gap name can still be kept
by the derived list (e.g. a profile's tokenEnv names it) and reach the member
unblocked while the INFO said "no member pane keeps them". That inversion is
exactly what #192 exists to remove.
logCredentialGap now splits the gap with MemberEnvAllowList.keeps (the same
predicate the generated scrub itself evaluates, so this cannot drift from
what the scrub does): names it keeps get a WARN, guarded by the same
unprotectedGapLogged flag as the deny-by-default case (same severity — a name
reaching a member unprotected is equally serious either way); names it
blanks keep the existing INFO, guarded by allowListGapLogged. The
deny-by-default WARN text and the non-zsh fallback are untouched.
Added allowListWarnsWhenTheDerivedAllowListKeepsAnUncoveredName and
allowListSplitsAMixedGapBetweenTheWarnAndTheInfo to ClaudeCodeLauncherTest.
1. CompositePeerLauncher.stop() was permanently un-stoppable for any
member that survived a daemon restart, because spawnedBy is in-memory
only. On a cache miss with more than one configured herdr daemon, probe
each distinct daemon's agent.list() for the pane instead of refusing
outright: exactly one owner routes and caches; zero owners is treated
as already-stopped (a no-op, matching the tolerance HerdrPeerLauncher
already gives an already-gone pane); more than one owner is the
genuine per-daemon-pane-id ambiguity and still throws.
2. FleetApp#healthz always reported the LEAD daemon's herdr version/
protocol even when a second (member) daemon was configured, so a
member-daemon protocol mismatch was invisible behind a green
/healthz while every spawn silently failed. Added a separate "member"
key alongside the unchanged "herdr" key, and a "protocolMismatch"
flag when the two differ. Verified scripts/redeploy-fleetd.sh and
scripts/rename-checkout.sh only check the HTTP status code and print
the body verbatim — neither parses a specific field — so adding a key
is safe.
Both fixes are covered by tests written to fail without the fix
(verified by reverting each fix and watching the new tests fail, then
restoring). Full `mvn clean install`: 1018 tests, 0 failures, 0 errors.
The memberHerdrSocket section exists at 11-Features.md:2174; what was absent was
its row in the index table. My omission when I added the section. Wiki fixed at
b24965c.
logCredentialGap(creds) always emitted the WARN wording ("every member pane
inherits them UNBLOCKED"), even under memberCredentials.policy: allow-list on
a zsh login shell, where the generated ZDOTDIR scrub genuinely blanks the
name. The line reported the control working as though it were a hole.
Pass an effectiveAllowed set instead: null keeps the WARN (deny-by-default,
and the allow-list non-zsh fallback, where nothing is ever scrubbed); the
derived allow-list set (only reachable after applyEnvironmentAllowListPolicy's
own zsh gate) selects a new INFO wording that says the scrub will blank the
name instead of claiming it is inherited unblocked.
Also split the single credentialGapLogged AtomicBoolean into two guards
(unprotectedGapLogged / allowListGapLogged) — one per report kind. Since
memberCredentials is a live, re-read-per-spawn supplier, a shared flag let a
harmless allow-list INFO on one spawn permanently suppress a later spawn's
real deny-by-default WARN after a policy reload.
Fixes gitea #192.
memberHerdrSocket splits lead operations from member operations onto two herdr
daemons. Three seams still assumed one shared daemon and broke silently when the
two clients differ (all three collapse to today's behaviour when they are the
same object):
1. ConnectionIdentity's PaneLocator was pinned to the member daemon only, so a
lead's own MCP connection (which lives on the LEAD daemon) resolved to
terminal == null, breaking fleet_reply/fleet_ask/fleet_whoami for a lead.
PaneLocator now searches the lead client first, then the member client.
2. StatusPoller's StatusRefiner was pinned to the member daemon, so refining an
UNKNOWN status for a lead target read the wrong daemon's pane content and
never left UNKNOWN, wedging status-gated delivery to that lead forever.
StatusRefiner gained a refine(target, raw, control) overload and the poller
now refines through the same AgentControl the raw status was sampled from.
3. FleetApp was constructed with the raw lead-only herdr client, so /healthz
stayed green while the member daemon was down (every spawn then fails
invisibly) and GET /sessions silently dropped every member workspace.
FleetApp now takes both clients: healthz requires both to answer, sessions
merges workspaces from both.
Each fix has a test proven to fail without it (verified by reverting the
production change and re-running): FleetdConnectionIdentityConstructionTest /
FleetdFleetAppConstructionTest assert the actual Fleetd.java wiring (the same
technique as FleetdHerdrControlConstructionTest); StatusPollerRoutingTest and
the new PaneLocatorTest/FleetAppTwoDaemonTest cases exercise the real
production classes end to end rather than a hand-built object graph.
Three fixes on top of the pane-id PR, from my own read and the reviewer's:
- stop() removed the spawnedBy record BEFORE the delegate accepted the stop. A
delegate that threw left the pane alive with its owner forgotten, so the retry
fell into the ambiguous branch and refused the id for good. Remove after.
- list() deduplicated on the raw pane id. Pane ids are per-daemon counters, so
two daemons can each hold w1:p1 on different panes, and one of the two real
agents was silently dropped from fleet_list and every view built on it. The
key is now (owning daemon, pane id). Delegates sharing one daemon still
collapse, which is what the dedupe was for.
- The class javadoc still stated the single-herdr-connection premise as fact,
next to the bullet this PR had just corrected for stop(). Fixed there too.
Also drops a redundantly qualified java.util.Collections.
Tests: 993 run, 0 failures, BUILD SUCCESS.
The skill said SSH to fleet01 is denied, so every report wrote 'not
reachable' for that fleet's daemon facts. That is true only for the user
dai.ha. The host alias fleet01 maps to user ltms and key auth works.
Checked 2026-08-28 while measuring #185: ssh fleet01 connects, and ltms
has passwordless sudo there. So fleet01's PID, uptime, jar and /healthz
can be reported over SSH even though its REST port is unreachable.
The javadoc said a member cannot authenticate at all once memberCredentials
blocks SSH_AUTH_SOCK, "there is no private key file on this host, only an
ssh-agent socket". That is wrong, and it was written after looking only in
~/.ssh, which holds nothing but Include lines.
Measured: ssh -G git.ltms.dev resolves an IdentityFile under the shared-env
directory. That file exists, is readable by this user, and has no passphrase.
A live member with SSH_AUTH_SOCK blanked pushed to the forge over SSH.
The rewrite itself is unchanged and still worth having. Only its stated reason
was wrong: it routes a member through its own scoped token instead of the
operator's ssh identity, which is what makes a member's pushes attributable
and revocable. It is not what stands between a member and the forge.
The four log lines added with the worktree HTTPS rewrite echoed the origin
URL verbatim, and one of them echoed the ssh:// authority, which carries
user-info. An ssh authority is normally just git@, so in practice this
changes nothing -- but a remote URL is not obviously a credential channel,
and that is precisely why one has leaked here three times (#157, #182).
Redact at the log call, not after it surprises someone.
Git never consults a credential.helper for an SSH transport, so #177's helper was inert on this repo — whose origin is ssh://. Once allow-list policy blocks SSH_AUTH_SOCK, a member on an SSH origin cannot authenticate at all: there is no private key file on this host, only an agent socket.
A worktree-scoped `url.<https>.insteadOf <ssh>` gives the member HTTPS for fetch and push while the primary checkout keeps SSH untouched. Host and port are parsed from the origin, never hardcoded — a test with a synthetic host proves it. The scp-like shorthand is left alone deliberately, since its host:path split is defined by ssh_config aliases rather than URI syntax.
Verified by the lead in an independent worktree: Tests run: 986, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
The member credential scrub removes environment variables. It cannot remove a token written into git config inside the repo the member works in, so `git remote -v` handed a member a credential it was deliberately not given.
Provisioning now strips HTTPS user info from the origin before `git worktree add`, refuses the worktree if user info survives, and configures a per-worktree credential helper that reads WORKER_GITEA_TOKEN at call time. Nothing is persisted.
The helper emits BOTH username and password, and resets the inherited helper list first. An earlier revision emitted only `username=`, which made git fall through to the next helper — on a Mac that is osxkeychain, so a member would have authenticated with the operator's stored credential while every test passed and `git remote -v` looked clean. See #182.
`worktreeCredentialHelperCompletesWithoutUsingAnInheritedHelper` plants a synthetic operator helper in an isolated global config and proves the worktree helper wins. The worker confirmed it fails when the reset is removed. All credential tests pin GIT_CONFIG_GLOBAL and GIT_CONFIG_SYSTEM so they can neither read nor write real credentials.
Verified by the lead in an independent worktree: Tests run: 973, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
A turn that dies on a backend error produces the same working -> idle transition as a real one, just faster and with nothing on screen. The resolver accepted that as a completed turn and handed the caller HTTP 200 with an empty reply, so a lost turn and a successful empty answer were indistinguishable.
Now: an empty or unreadable scrape fails, naming the member; and a BUSY -> DONE inside MIN_TURN_NANOS (2s) fails as a crash signature.
One existing test encoded the bug — it asserted a failed scrape resolved as a success carrying "" — and has been inverted rather than worked around.
Verified by the lead in an independent worktree: Tests run: 973, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
`policy: allow-list` silently ignored every name an operator wrote under `allow:` unless a profile happened to carry it too, so turning the policy on would have blanked credentials working members depend on. Derivation now unions the operator's list.
`SSH_AUTH_SOCK` stays governed only by `sshAuthSock`, even when listed under `allow:` — it is a live handle to the operator's ssh-agent, not a value.
Adds one INFO line per allow-list spawn, `member credentials: allowed N of M`, emitted only after the shell gate so it can never report coverage on a path where the scrub does not run.
Verified by the lead in an independent worktree: Tests run: 976, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
No behaviour change. #154 supposed that ownership drains a whole queue into the in-memory `held` map, making `x-max-length` and per-message TTL decorative. Measurement says otherwise: `basicConsume` is manual-ack, `deliverCallback` acks only duplicates, and `basicQos` is set on the one shared channel before any consumer starts — so total `held` is bounded by the prefetch window across all targets.
Adds a fake-broker test that drives the real `own()` path and fails if receipt ever starts acking, plus a javadoc line naming the prefetch window at the point of first mention.
Verified by the lead in an independent worktree: Tests run: 971, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
CompletionResolver.resolve() used to hand the caller a successful "" reply whenever a
turn's scrape came back empty (whether the read failed, or genuinely produced nothing),
making a lost turn indistinguishable from a real empty answer. It also had no way to
tell a crashed backend's near-instant BUSY -> DONE transition apart from a genuine
completion.
Add MIN_TURN_NANOS (2s), a named floor below which a completed turn is treated as a
crash signature and failed rather than resolved as a reply. Fail on any empty scrape
(read failure or a clean-but-empty read) instead of resolving with "". Both failures
name the member and carry whatever is on the pane for context.
Thread an injectable LongSupplier clock through CompletionResolver (matching the
SessionManager/MessageService nowNanos pattern) so the floor is testable without a
real sleep.
The coverage line was logged before the zsh gate, so a non-zsh
spawn (where nothing is scrubbed — overlayBlockedCredentials is the
fallback instead) printed 'allowed N of M' as if the derived
allow-list scrub had run. Move the log after the gate so it only
fires on the path that actually generates the ZDOTDIR scrub; the
non-zsh fallback keeps logCredentialGap's WARN as its only signal.
Added a test proving no 'allowed N of M' line is emitted on the
non-zsh fallback, through the real HerdrPeerLauncher#spawn path.
Share one deliverability predicate between Injector and FleetApp, so the status endpoint reports the same answer the injector acts on instead of re-deriving it from one of that predicate's two inputs.
Verified by the lead in an independent worktree: Tests run: 971, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
MemberEnvAllowList.derive only ever looked at profile fields, so
memberCredentials.allow: was silently ignored under
policy: allow-list — turning the policy on would have blanked
credentials working members already depended on.
- derive(profiles, configuredAllow) unions memberCredentials.allow
into the derived set, with SSH_AUTH_SOCK explicitly excluded from
that union (it stays governed only by sshAuthSock: allow).
- HerdrPeerLauncher threads MemberCredentials.allowSet() into the
derivation instead of calling the profiles-only overload.
- Added a per-spawn INFO log 'member credentials: allowed N of M'
(N/M from the daemon's own env, the existing hostEnvNames proxy),
never logging a blocked name or a value.
CB-640 published the three message-layer facts and CB-641 wired the herdr
and time ones. This joins them, so every HealthSnapshot field now carries
real evidence and the NOT_YET_OBSERVED placeholder is gone. That constant
is what made 8 of the 9 fault states unreachable, GONE and NEVER_READY
included, which is why CB-580's failTarget never fired.
hasOrphanedDelegation is a true snapshot, but it can read true for one
tick during an ordinary race: an async ticket exists before its virtual
thread reaches rendezvous.open, so for that instant nothing is accepted or
queued behind it. decide maps the field straight to DELEGATION_ORPHANED
with no smoothing, so one racy read would log a fault that clears on the
next tick. The monitor now requires two consecutive observations. That
costs one interval on a real orphan and removes the false positive.
Two tests drive real ticks against a genuinely orphaned ticket (an
unanswered fleet_ask that lapsed back to PENDING), not the seam: one tick
reports nothing, two report once, and a single clean tick in between
resets the streak.
Also correct two config comments. paneProbeIntervalSeconds is parsed and
read by nothing, so its "minimum 60" note promised a floor that does not
exist.
970 tests green.
hasQueuedDelivery/hasStrandedReply/hasOrphanedDelegation surface three of the
message-layer facts FleetHealthMonitor needs but currently hardcodes to
NOT_YET_OBSERVED. Additive only — no existing public method's signature or
behavior changes.
scrubWritesAnAllowedNofMReport spawned /bin/zsh with no guard, so it
errored on the Gitea CI runner (Linux ARM64 container, no zsh) while
the two sibling zsh tests already skipped there via assumeTrue. Add the
same guard so CI skips instead of failing; the Mac build still runs it.
The alias removal left bridge_* tool names in prose. Fix them:
- README no longer claims the old bridge_* names still answer (they were removed).
- pom + LeadTabScanner comments name fleet_* tools.
- FleetMcp comment no longer mentions the removed deprecated twin.
- docs/MCP-Contract.md and e2e swept bridge_* -> fleet_*; e2e ask files renamed.
The historical mcp__bridge__* mount-name note in CLAUDE.md is kept on purpose.
949 tests pass.
Two changes ship together here.
1. One shared herdr workspace. The lead and every worker now live in one
workspace called "fleet", so the operator sees one "session" with many
windows, not two. Before, the lead sat in a "leads" workspace and workers
in "bridged-workers", which read as two sessions. The lead is still told
apart from workers by its exact tab label ("lead: <name>"), so putting them
in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
for split layouts; Fleetd now passes an empty exclude set.
2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
launchd/systemd units, module dir, and MCP mount).
- Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
- Log line, comments, docs, and CLAUDE.md updated to say fleetd.
- Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
- Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
- Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
still read as a fallback, and still gitignored.
- MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
server name is "fleet". The mount name in the local .mcp.json becomes
"fleet" (gitignored, not in this commit).
- Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.
Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.
Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.
The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).
949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
- CB-634: deliver IDE guidance as an on-disk CLAUDE.local.md / opencode overlay,
pinned to the module dir, best-effort auto-open (opt-in per profile via ideMcpUrl).
- CB-636: per-profile autoCompactWindow -> --autocompact for claude-code,
provider.<p>.models.<m>.limit.context for opencode. Range-validated [100000,1000000].
- CB-637: cross-host lead-to-lead over a shared AMQP coordination vhost
(coordinator: block, fleet_send{coordId}, LeadMailbox + LeadCoordLoop).
952 unit tests + 5 LeadMailbox contract tests green. Live E2E verified on the Mac daemon.
The lead-to-lead wiring shipped with CB-635 in its comments, but CB-635 is
already the broker.uriEnv / unreachable-broker work. Relabel the mailbox +
fleet_send{coordId} + receive loop to CB-637 so a ticket number names one
feature. Add the cross-host peer-lead row to the primary intent->tool table
(kept byte-identical with the wiki template).