08ce9aef11ca6d23cc300399a4682db9aaef1b02
7 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2e138a199b |
CB-634: one shared "fleet" workspace + rename bridged -> fleetd cutover
Two changes ship together here.
1. One shared herdr workspace. The lead and every worker now live in one
workspace called "fleet", so the operator sees one "session" with many
windows, not two. Before, the lead sat in a "leads" workspace and workers
in "bridged-workers", which read as two sessions. The lead is still told
apart from workers by its exact tab label ("lead: <name>"), so putting them
in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
for split layouts; Fleetd now passes an empty exclude set.
2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
launchd/systemd units, module dir, and MCP mount).
- Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
- Log line, comments, docs, and CLAUDE.md updated to say fleetd.
- Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
- Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
- Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
still read as a fallback, and still gitignored.
- MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
server name is "fleet". The mount name in the local .mcp.json becomes
"fleet" (gitignored, not in this commit).
- Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.
Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.
Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.
The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).
949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
|
||
|
|
555715ced9 |
CB-623: point every path at fleet/fleetd after the org transfer
The repo moved lms/claude-bridge -> fleet/claude-bridge -> fleet/fleetd. Gitea redirects hold, so most of this is not urgent, but one line was a real break: the implementer skill posts a worker's PR to a hardcoded repo path, so every worker PR would have gone to the old address. .claude/skills/implementer/SKILL.md the worker PR endpoint (functional) .gitmodules wiki submodule URL CLAUDE.md + wiki/7-Use-Cases.md the canonical block, kept byte-identical README.md clone command and wiki link deploy/bridged.service Documentation= plugin/.claude-plugin/plugin.json homepage + repository docs/*.md issue and wiki links The wiki is not a separate repo. /repos/lms/claude-bridge.wiki returns 404 and lms owned no .wiki entity, so the wiki moved with the repo; both the old and the new wiki SSH URLs resolve to the same sha. Ticket step 4 assumed a second transfer that does not exist. |
||
|
|
cec48832be |
CB-600: make it safe to install the launchd agent
- redeploy-bridged.sh now refuses (not warns) a supervised restart when its computed log path disagrees with the loaded plist's StandardOutPath — otherwise every post-restart check reads the wrong file and can report a clean restart while the daemon crash-loops. The check is a pure, testable function; the script gained a source-for-test guard so it can be exercised without installing the agent or touching launchd. - a failed 'launchctl load' after a successful 'unload' now retries once and, on ultimate failure, tells the operator the agent is stopped AND disabled plus the exact recovery command, instead of leaving that silently worse than the pre-redeploy state. - the plist documents honestly that the crash loop launchd retries is unbounded (ThrottleInterval only paces it), and what actually stops it. - fixed the requiredSecretEnvVars javadoc: the auth.tokenEnv startup throw is ~370 lines below its call site, not a few lines above it, and only fires in auth.mode: token. |
||
|
|
3ba6d6784c |
CB-594: make supervision and a working fleet possible at the same time
Adds scripts/bridged-launchd-wrapper.sh so the launchd-run daemon still gets WORKER_GITEA_TOKEN/AI_GATEWAY_TOKEN by execing through a login shell (launchd never sources secrets.sh itself). bridged now logs at startup which required token env vars (derived from each profile's tokenEnv/gitTokenEnv, not a hand-written list) resolved or are MISSING, by name only. Fills in the real paths in deploy/dev.ltms.bridged.plist for this host and points it at the wrapper. scripts/redeploy-bridged.sh now detects a loaded launchd agent and uses launchctl unload/load instead of a raw kill+nohup, because a bare SIGTERM exits this JVM at 143 (measured) which KeepAlive.SuccessfulExit=false reads as a crash and would race the script's own restart; --check reports installed/loaded state and stays read-only. |
||
|
|
b67b1585c2 |
CB-517: deploy LavinMQ as a pinned, durable, self-restarting broker
CI / build (push) Successful in 1m23s
The CB-307 durable ReplyInbox needs an AMQP broker, but the one behind it
was run ad hoc and had simply vanished from the host — which takes the
whole daemon with it, since AmqpReplyInbox.open throws and Bridged.java:187
does not guard it. A missing broker is a hard startup failure, not a
degraded mode, so 'how the broker runs' is part of the system, not a local
detail.
Pinned to 2.9.1 (:latest would move the broker under a running daemon),
data on a named volume (held-but-unacked replies are the entire point of
Stage 2 — a plain 'compose down' would discard exactly what durability
protects), and restart: unless-stopped so it comes back after a reboot
instead of disappearing again.
Ports are bound to 127.0.0.1 deliberately: LavinMQ ships a default
guest/guest account, which is only acceptable while nothing off-host can
reach it.
Verified by driving the production AmqpReplyInbox against this deployment
(publish/peek/dedup/FIFO/ack, then reconnect): 8/8 including redelivery of
the unacked message. That pairing had never been exercised — the
@Tag("contract") test runs against a RabbitMQ container, and is excluded
from the default build, so mvn clean install covers the broker path zero
times.
|
||
|
|
22ad24db6c |
CB-511: give workers a toolchain — propagate the daemon PATH, add profile env:
CI / build (push) Successful in 1m17s
Workers could not run `mvn` or `java`. Every delegated task that asked for a build came back "mvn is not on PATH", and the worker was right. Root cause: HerdrPeerLauncher seeded the worker environment with an EMPTY map, so bridged passed only the vars it explicitly set (OPENCODE_CONFIG, GITEA_TOKEN, ANTHROPIC_*) and never PATH. herdr merges that map into its own process env, so a worker inherited whatever PATH the herdr SERVER was started with. On this host that server (pid 79870, PPID 1) had been up since Jul 4 with a PATH containing neither the JDK nor Maven. Confirmed on a live worker: its PATH was byte-identical to herdr's, and the only var bridged had contributed was OPENCODE_CONFIG. The failure was invisible and non-deterministic: the fleet's capabilities depended on how a long-lived daemon happened to be launched weeks earlier. There are three herdr processes on this box with three different PATHs; the one owning the socket is the one without a toolchain. bridged itself HAD Maven on PATH the whole time — it just never passed it on. It also quietly contradicted the project's own principle that "a worker is a full peer of the primary", and the implementer skill's instruction to build, commit and open a PR. Every delegation so far has depended on the primary running the build gate. Fix: baseEnv(cfg) seeds each worker with the daemon's own PATH, then applies the profile's new optional env: map. Adapter-specific vars are layered on top and therefore win — that ordering is load-bearing, not incidental: it stops an env: entry from overwriting ANTHROPIC_BASE_URL and slipping past SubscriptionGuard, which is checked against the profile's baseUrl alone. Pinned by a test. Because the default is now the daemon's PATH, both supervision units set PATH explicitly — launchd and systemd do not source a login shell, so under CB-504 the daemon (and every worker) would otherwise get a bare /usr/bin:/bin and this bug would silently return in production. 324 tests (was 321): daemon-PATH propagation, profile env: passthrough including an explicit PATH override, and the guard-bypass ordering. Verified live: daemon restarted, worker spawned, and asked to run the tools — "Apache Maven 3.9.16", "java version 25.0.2". Previously both were absent. |
||
|
|
9daf1ec5ba |
CB-5xx: Stage 5 hardening — auth, authz+audit, metrics, CI, supervision
Closes out single-host before the cross-host work. Sequenced BEFORE CB-308 deliberately: federation's own gating concern is the trust model, and it inherits whatever identity shape lands here. The finding this stage is built around: bridged had exactly ONE security control, the loopback bind. ConnectionIdentity resolves a worker from its connection (unforgeable), but every caller that was not a recognised worker pane fell through to being treated as the PRIMARY -- the most privileged role on the bus. Latent today; load-bearing the moment a bind widens. CB-501 auth: - Role/Principal/CallerResolver: connection identity first, bearer token second, ANONYMOUS third. Inverts the old default so absence of identity means nothing, not everything. - Worker identity is never token-gated, so enabling auth cannot lock the fleet out of bridge_reply. - Constant-time token compare (MessageDigest.isEqual). - validateAuthExposure(): a non-loopback bind under loopback-trust now REFUSES TO START. Makes the dangerous config unrepresentable rather than merely documented. - TLS terminates at a reverse proxy by design (D3), not in the JVM. CB-505 authz + audit, enforced on BOTH entry paths: - The docs describe MCP as "a thin adapter over the REST core"; at code level it is not. BridgeMcp calls MessageService directly, and /mcp is a raw servlet on Jetty's context handler that never traverses Javalin's before filter. Enforcing only at REST would have left /mcp open. - Load-bearing rule is own-session-only: a worker may reply/ask only as itself. Structurally true over MCP already; over REST the session id in the URL path had simply been trusted. - Audit: JSON lines to a dedicated appender, additivity=false. Never records message content -- this bus carries source and prompts. CB-502 metrics: zero new dependencies. A ~150-line Prometheus text renderer instead of the specced Micrometer, because this pom already hand-pins jackson-annotations to reconcile Jackson 2/3, imports a Jetty BOM against skew, and carries four accepted-CVE advisories -- and CLAUDE.md's mandated dependency CVE gate could not be run (no JetBrains MCP server connected). Instrumented at MessageService, the single funnel both surfaces share. CB-503 CI: .gitea/workflows/ci.yml against the already-running Gitea runner. Needs no contract-exclusion flag -- the pom's default-excludes profile already sets excludedGroups=contract, so plain `mvn clean install` IS the mock-socket surface. Provisions JDK 25 explicitly (runner default-jdk is older). CB-504 supervision: launchd agent (the real target -- this host is macOS, there is no systemd) plus a systemd unit for the Linux gateways CB-308 adds. Ordering directives are advisory, so the actual fix is that startup now waits up to 30s for the herdr socket and then serves degraded, instead of crashing into a restart loop on a boot-order race. Also fixes drift found while surveying: - bridged.example.yaml documented spawn_ready_timeout_ms in snake_case; config binds via plain Jackson with ignoreUnknown, so uncommenting it would have been silently dropped and the default kept. Now camelCase, with a test that loads the shipped example and one that pins every documented knob's spelling -- no test had ever loaded that file. - Added the 6 shipped-but-undocumented knobs (worktreeRoot, parityOverlay, gitTokenEnv, gitHostEnv, configDir, primary:). - README "Next" listed bridge_ask and session lifecycle as upcoming; both shipped long ago. - docs/CB-301-ext and docs/CB-402 status headers said "design"/"pre- implementation" for work already merged. 307 unit/acceptance tests green (was 266), mvn clean install BUILD SUCCESS. Note: CLAUDE.md's per-file ide_diagnostics gate and the pom Mend.io CVE check could not be run -- no JetBrains/intellij-index MCP server is connected this session. mvn clean install is the only gate that ran. |