f8522edacd0300ce33a9f52d33fd18e1e2ace70e
3 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b67b1585c2 |
CB-517: deploy LavinMQ as a pinned, durable, self-restarting broker
CI / build (push) Successful in 1m23s
The CB-307 durable ReplyInbox needs an AMQP broker, but the one behind it
was run ad hoc and had simply vanished from the host — which takes the
whole daemon with it, since AmqpReplyInbox.open throws and Bridged.java:187
does not guard it. A missing broker is a hard startup failure, not a
degraded mode, so 'how the broker runs' is part of the system, not a local
detail.
Pinned to 2.9.1 (:latest would move the broker under a running daemon),
data on a named volume (held-but-unacked replies are the entire point of
Stage 2 — a plain 'compose down' would discard exactly what durability
protects), and restart: unless-stopped so it comes back after a reboot
instead of disappearing again.
Ports are bound to 127.0.0.1 deliberately: LavinMQ ships a default
guest/guest account, which is only acceptable while nothing off-host can
reach it.
Verified by driving the production AmqpReplyInbox against this deployment
(publish/peek/dedup/FIFO/ack, then reconnect): 8/8 including redelivery of
the unacked message. That pairing had never been exercised — the
@Tag("contract") test runs against a RabbitMQ container, and is excluded
from the default build, so mvn clean install covers the broker path zero
times.
|
||
|
|
22ad24db6c |
CB-511: give workers a toolchain — propagate the daemon PATH, add profile env:
CI / build (push) Successful in 1m17s
Workers could not run `mvn` or `java`. Every delegated task that asked for a build came back "mvn is not on PATH", and the worker was right. Root cause: HerdrPeerLauncher seeded the worker environment with an EMPTY map, so bridged passed only the vars it explicitly set (OPENCODE_CONFIG, GITEA_TOKEN, ANTHROPIC_*) and never PATH. herdr merges that map into its own process env, so a worker inherited whatever PATH the herdr SERVER was started with. On this host that server (pid 79870, PPID 1) had been up since Jul 4 with a PATH containing neither the JDK nor Maven. Confirmed on a live worker: its PATH was byte-identical to herdr's, and the only var bridged had contributed was OPENCODE_CONFIG. The failure was invisible and non-deterministic: the fleet's capabilities depended on how a long-lived daemon happened to be launched weeks earlier. There are three herdr processes on this box with three different PATHs; the one owning the socket is the one without a toolchain. bridged itself HAD Maven on PATH the whole time — it just never passed it on. It also quietly contradicted the project's own principle that "a worker is a full peer of the primary", and the implementer skill's instruction to build, commit and open a PR. Every delegation so far has depended on the primary running the build gate. Fix: baseEnv(cfg) seeds each worker with the daemon's own PATH, then applies the profile's new optional env: map. Adapter-specific vars are layered on top and therefore win — that ordering is load-bearing, not incidental: it stops an env: entry from overwriting ANTHROPIC_BASE_URL and slipping past SubscriptionGuard, which is checked against the profile's baseUrl alone. Pinned by a test. Because the default is now the daemon's PATH, both supervision units set PATH explicitly — launchd and systemd do not source a login shell, so under CB-504 the daemon (and every worker) would otherwise get a bare /usr/bin:/bin and this bug would silently return in production. 324 tests (was 321): daemon-PATH propagation, profile env: passthrough including an explicit PATH override, and the guard-bypass ordering. Verified live: daemon restarted, worker spawned, and asked to run the tools — "Apache Maven 3.9.16", "java version 25.0.2". Previously both were absent. |
||
|
|
9daf1ec5ba |
CB-5xx: Stage 5 hardening — auth, authz+audit, metrics, CI, supervision
Closes out single-host before the cross-host work. Sequenced BEFORE CB-308 deliberately: federation's own gating concern is the trust model, and it inherits whatever identity shape lands here. The finding this stage is built around: bridged had exactly ONE security control, the loopback bind. ConnectionIdentity resolves a worker from its connection (unforgeable), but every caller that was not a recognised worker pane fell through to being treated as the PRIMARY -- the most privileged role on the bus. Latent today; load-bearing the moment a bind widens. CB-501 auth: - Role/Principal/CallerResolver: connection identity first, bearer token second, ANONYMOUS third. Inverts the old default so absence of identity means nothing, not everything. - Worker identity is never token-gated, so enabling auth cannot lock the fleet out of bridge_reply. - Constant-time token compare (MessageDigest.isEqual). - validateAuthExposure(): a non-loopback bind under loopback-trust now REFUSES TO START. Makes the dangerous config unrepresentable rather than merely documented. - TLS terminates at a reverse proxy by design (D3), not in the JVM. CB-505 authz + audit, enforced on BOTH entry paths: - The docs describe MCP as "a thin adapter over the REST core"; at code level it is not. BridgeMcp calls MessageService directly, and /mcp is a raw servlet on Jetty's context handler that never traverses Javalin's before filter. Enforcing only at REST would have left /mcp open. - Load-bearing rule is own-session-only: a worker may reply/ask only as itself. Structurally true over MCP already; over REST the session id in the URL path had simply been trusted. - Audit: JSON lines to a dedicated appender, additivity=false. Never records message content -- this bus carries source and prompts. CB-502 metrics: zero new dependencies. A ~150-line Prometheus text renderer instead of the specced Micrometer, because this pom already hand-pins jackson-annotations to reconcile Jackson 2/3, imports a Jetty BOM against skew, and carries four accepted-CVE advisories -- and CLAUDE.md's mandated dependency CVE gate could not be run (no JetBrains MCP server connected). Instrumented at MessageService, the single funnel both surfaces share. CB-503 CI: .gitea/workflows/ci.yml against the already-running Gitea runner. Needs no contract-exclusion flag -- the pom's default-excludes profile already sets excludedGroups=contract, so plain `mvn clean install` IS the mock-socket surface. Provisions JDK 25 explicitly (runner default-jdk is older). CB-504 supervision: launchd agent (the real target -- this host is macOS, there is no systemd) plus a systemd unit for the Linux gateways CB-308 adds. Ordering directives are advisory, so the actual fix is that startup now waits up to 30s for the herdr socket and then serves degraded, instead of crashing into a restart loop on a boot-order race. Also fixes drift found while surveying: - bridged.example.yaml documented spawn_ready_timeout_ms in snake_case; config binds via plain Jackson with ignoreUnknown, so uncommenting it would have been silently dropped and the default kept. Now camelCase, with a test that loads the shipped example and one that pins every documented knob's spelling -- no test had ever loaded that file. - Added the 6 shipped-but-undocumented knobs (worktreeRoot, parityOverlay, gitTokenEnv, gitHostEnv, configDir, primary:). - README "Next" listed bridge_ask and session lifecycle as upcoming; both shipped long ago. - docs/CB-301-ext and docs/CB-402 status headers said "design"/"pre- implementation" for work already merged. 307 unit/acceptance tests green (was 266), mvn clean install BUILD SUCCESS. Note: CLAUDE.md's per-file ide_diagnostics gate and the pom Mend.io CVE check could not be run -- no JetBrains/intellij-index MCP server is connected this session. mvn clean install is the only gate that ran. |