Commit Graph

13 Commits

Author SHA1 Message Date
Dai Ha e4973eb8a4 Classify recovered AMQP redeploy errors
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m51s
2026-09-05 05:29:53 +07:00
Dai Ha 51f7b0a3ca fleetd #111: probe reads the live memberCredentials policy, no hardcoded name list
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 2m22s
scripts/probe-member-credentials.sh carried its own hand-maintained NAMES array
(31 names, recorded 2026-08-16), so a name added later to fleetd.yaml's
memberCredentials.known was never checked and the probe still exited 0 with a
clean-looking table. Same drift shape as #114's tool catalogue.

- New dev.ltms.fleet.member.MemberCredentialPolicyView: the single place that
  turns a MemberCredentials policy into names + counts (never a value). Reused
  by Fleetd.reportMemberCredentialsGap (startup log line) and by the new
  GET /member-credentials REST endpoint (FleetApp), so the two can no longer
  drift apart the way the probe and the policy did.
- FleetApp gains one route + handler + a Supplier<MemberCredentialPolicyView>
  constructor param (legacy constructors default to ::absent, so existing call
  sites are unaffected).
- probe-member-credentials.sh now fetches its name list from
  GET /member-credentials instead of carrying one. No local fallback: an
  unreachable daemon, an empty/absent policy, or a knownCount/known[] length
  mismatch all refuse with a non-zero exit rather than silently checking zero
  names. Prints "policy contains N; this run checked N" so the two numbers are
  visibly equal.
2026-09-03 13:21:35 +07:00
Dai Ha 2e138a199b CB-634: one shared "fleet" workspace + rename bridged -> fleetd cutover
Two changes ship together here.

1. One shared herdr workspace. The lead and every worker now live in one
   workspace called "fleet", so the operator sees one "session" with many
   windows, not two. Before, the lead sat in a "leads" workspace and workers
   in "bridged-workers", which read as two sessions. The lead is still told
   apart from workers by its exact tab label ("lead: <name>"), so putting them
   in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
   for split layouts; Fleetd now passes an empty exclude set.

2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
   launchd/systemd units, module dir, and MCP mount).
   - Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
   - Log line, comments, docs, and CLAUDE.md updated to say fleetd.
   - Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
     bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
   - Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
     bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
   - Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
     still read as a fallback, and still gitignored.
   - MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
     server name is "fleet". The mount name in the local .mcp.json becomes
     "fleet" (gitignored, not in this commit).
   - Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
     BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.

Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.

Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.

The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).

949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
2026-08-25 04:01:08 +02:00
Dai Ha 3f4ac2b24e CB-635: --check reports whether broker.uriEnv resolves in a login shell
CI / contract (push) Successful in 46s
CI / build (push) Failing after 1m30s
An empty uriEnv no longer stops the daemon (#152), so the failure is quiet: bridged
starts, falls back to the in-memory reply inbox, and held reports stop surviving a
restart. --check is the only thing that says so before the fact. The var name is read
out of bridged.yaml so a renamed key cannot make the check lie.
2026-08-23 14:03:08 +02:00
Dai Ha 51bdec22e7 CB-624: report three more rename surfaces as report-only
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Successful in 1m14s
--check gains section 6 counting ~/.claude.json (the projects entry keyed
by the old absolute path), ~/.config/herdr/session.json, and JetBrains
recentProjects.xml / trusted-paths.xml across every IntelliJIdea* version.
Apply mode's closing summary lists the same three with manual follow-ups.

None of the three is rewritten automatically, on purpose: ~/.claude.json
is global live Claude Code config, herdr session.json is live process
state, and JetBrains rewrites its own files when the project is reopened
at the new path. Header comment states the reason for each so it does not
read as an oversight.

JetBrains stores these paths as its $USER_HOME$ macro rather than a
literal absolute path, so the counter matches both forms and says which
form the hits used.
2026-08-23 06:07:16 +02:00
Dai Ha d43f670285 CB-624: add scripts/rename-checkout.sh to rename the checkout safely
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m32s
Renames ~/LTMS/claude-bridge -> ~/LTMS/fleetd as one auditable command,
modeled on redeploy-bridged.sh. --check reports every surface holding the
old absolute path (checkout files, launchd plist, Claude Code project
state, worktree .git pointers, running daemon). Apply mode stops the
daemon first (launchctl-aware), moves the checkout and the Claude Code
project-state slug dir derived from both paths, repairs worktree gitdir
pointers, rewrites bridged.yaml and the installed plist if they exist,
restarts, and verifies /healthz plus a fresh 'bridged listening' line
anchored to a pre-stop marker.
2026-08-23 05:55:04 +02:00
Dai Ha 03473a286b CB-622: rename bridge_* tools to fleet_* in scripts, e2e tests and config; mount name bridged -> fleetd
CI / build (pull_request) Successful in 58s
CI / contract (pull_request) Successful in 1m9s
2026-08-22 21:49:15 +02:00
Dai Ha 837fed7690 CB-596: the credential probe, as one auditable command
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m39s
Issue #82 step 1 is a measurement, and the classifier refuses an ad-hoc pipeline
that enumerates credential names inside a member — correctly. This is the seam:
one file the operator reads once and then runs, instead of approving a shell
pipeline they have to take on trust.

It never prints a credential value or any part of one. #82's criterion 1 asked
for a 6-character prefix; this prints a truncated SHA-256 instead. A prefix of a
short secret is most of the secret and would end up pasted into a ticket, while
the hash answers every question the prefix was for — is it set, is it the same
value as over there, is it the CB-592 sentinel.

Refuses to run unless BRIDGED_MEMBER=1, since the finding is what a MEMBER holds;
--allow-outside-member takes the comparison reading and labels it as such.

I have not run the reading path. That is the operator's call, which is the whole
point of the ticket.
2026-08-16 19:01:42 +02:00
Dai Ha cec48832be CB-600: make it safe to install the launchd agent
CI / build (pull_request) Failing after 1m21s
CI / contract (pull_request) Successful in 1m26s
- redeploy-bridged.sh now refuses (not warns) a supervised restart when
  its computed log path disagrees with the loaded plist's StandardOutPath
  — otherwise every post-restart check reads the wrong file and can
  report a clean restart while the daemon crash-loops. The check is a
  pure, testable function; the script gained a source-for-test guard so
  it can be exercised without installing the agent or touching launchd.
- a failed 'launchctl load' after a successful 'unload' now retries once
  and, on ultimate failure, tells the operator the agent is stopped AND
  disabled plus the exact recovery command, instead of leaving that
  silently worse than the pre-redeploy state.
- the plist documents honestly that the crash loop launchd retries is
  unbounded (ThrottleInterval only paces it), and what actually stops it.
- fixed the requiredSecretEnvVars javadoc: the auth.tokenEnv startup
  throw is ~370 lines below its call site, not a few lines above it, and
  only fires in auth.mode: token.
2026-08-16 18:08:58 +02:00
Dai Ha 3ba6d6784c CB-594: make supervision and a working fleet possible at the same time
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m40s
Adds scripts/bridged-launchd-wrapper.sh so the launchd-run daemon still gets
WORKER_GITEA_TOKEN/AI_GATEWAY_TOKEN by execing through a login shell (launchd
never sources secrets.sh itself). bridged now logs at startup which required
token env vars (derived from each profile's tokenEnv/gitTokenEnv, not a
hand-written list) resolved or are MISSING, by name only. Fills in the real
paths in deploy/dev.ltms.bridged.plist for this host and points it at the
wrapper. scripts/redeploy-bridged.sh now detects a loaded launchd agent and
uses launchctl unload/load instead of a raw kill+nohup, because a bare
SIGTERM exits this JVM at 143 (measured) which KeepAlive.SuccessfulExit=false
reads as a crash and would race the script's own restart; --check reports
installed/loaded state and stays read-only.
2026-08-16 17:13:17 +02:00
Dai Ha f0095bf8b2 CB-591: plan the move onto the LLM/MCP gateway, and check AI_GATEWAY_TOKEN
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m38s
The gateway (llm.ltms.dev) replaced Bifrost on 2026-08-15 and serves an Anthropic
surface and an OpenAI surface, so both member kinds can point at it. The plan is in
docs/CB-591-Gateway-Migration.md; gitea #76 tracks the work.

The opencode half needs no code: OpenCodeLauncher already pins an OpenAI-compatible
endpoint (CB-508), so baseUrl + tokenEnv + provider/model is a config change. That
matters more than it looks — every opencode member today is sol or terra, and both sit
on one OpenAI account via credentialId: openai-shared, so an exhaustion on either locks
out both. A gateway-backed opencode profile is free and off that credential, which
retires a single point of failure rather than only adding capacity.

Also extends the redeploy script's --check to AI_GATEWAY_TOKEN. A profile's tokenEnv is
resolved from the DAEMON's own environment by HerdrPeerLauncher.resolveEnv, so a token
added to secrets.sh after the daemon started is simply absent: the launcher injects an
empty token and the gateway answers 401, long after the restart and with nothing tying
the two together. That is the same trap as WORKER_GITEA_TOKEN, and it gets the same
login-shell check that never prints the value.
2026-08-15 17:06:31 +02:00
Dai Ha 5fe02b7c98 Fix redeploy script reporting failure on a successful restart
CI / contract (push) Successful in 45s
CI / build (push) Successful in 59s
Found by running it. The script launched the daemon with a relative jar path
(cwd is bridged/) but detected it with an absolute one, so pgrep never matched.
The daemon restarted correctly and booted clean, and the script still failed
with 'no process appeared' — the worst shape of bug for a deploy tool, because
it invites a second restart on a daemon that is already healthy.

Detection now matches both path forms, and the launch uses the absolute path so
ps names which checkout is running.
2026-08-15 15:18:44 +02:00
Dai Ha a1052f4fd1 Add scripts/redeploy-bridged.sh so the lead can deploy in one command
CI / build (push) Successful in 57s
CI / contract (push) Successful in 1m23s
The lead already owned the redeploy, but the command classifier refuses a bare
kill on the daemon, so in practice every deploy still needed the operator to
approve a stop and a start by hand. A single script is the seam that fixes
that: the operator allow-lists one auditable command instead of two ad-hoc
ones.

It also stops the procedure from living only in a checklist people read after
things go wrong. It builds before it stops anything, so a failed build never
leaves the fleet down; waits for the old process to exit instead of assuming;
polls /healthz; and anchors its log checks to a line marker taken before the
restart, so old errors cannot be misread as new ones.

The check with no log line anywhere in the daemon is the reason --check exists:
bridged inherits WORKER_GITEA_TOKEN from the shell that starts it, and starting
from a non-login shell empties it. The daemon boots fine, healthz is green, and
the failure only appears later as workers that cannot open a PR. --check tests
whether the name resolves and never prints the value.
2026-08-15 15:14:21 +02:00