9 Commits

Author SHA1 Message Date
Dai Ha 3da44eed63 fleetd #550: replace shasum with a portable hash256 helper, add a Linux CI job for the shell suite
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Successful in 1m46s
jar_id() in redeploy-fleetd.sh called shasum directly, which does not exist on GNU coreutils
Linux (Debian/Ubuntu/etc.) — there it silently reported an existing jar as "absent" with exit 0,
because the missing command made `cut` succeed on empty input and pipefail's failure was then
swallowed by the `|| echo "absent"` fallback. The shell test suite hit the same tool at
test-redeploy-fleetd.sh:298-299 and died at exit 127 with zero FAIL lines printed — the same shape
as a clean pass on the one channel anyone would check.

Adds one hash256() helper (prefer sha256sum, fall back to shasum -a 256, same idiom already used
in probe-member-credentials.sh) and points jar_id and the test suite's own reference hash at it.
jar_id now has three distinct answers instead of two: absent, a hash, or "unhashable" when neither
hasher is on PATH — "absent" is never used for a file that exists.

Adds a CI job (shell-tests) that runs scripts/test-redeploy-fleetd.sh on ubuntu-latest, gated on
the step's own exit code rather than a FAIL-line count, since a suite that dies before running is
exactly what a green run also looks like by that count.

New tests: test_jar_id_reports_unhashable_when_no_hasher_on_path (stubbed PATH with neither
hasher) and test_no_unguarded_macos_only_hasher_calls (a shape check across every script under
scripts/, not named lines — #545 already showed this idiom spreading from two sites to six).
2026-09-12 15:06:15 +07:00
Dai Ha 90253f832d fleetd #459: lint Javadoc references in CI
CI / contract (pull_request) Successful in 1m30s
CI / build (pull_request) Successful in 2m9s
2026-09-12 13:36:37 +07:00
Dai Ha d4f93a7b13 fleetd #449: fix stale herdr protocol 14 javadocs/assertion, diagnose and fix the timing-raced AgentControlContractTest, select contract tests by tag in CI
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 2m9s
- HerdrClient.java, HerdrCodec.java, HerdrContractTest.java: the herdr port to
  protocol 19 (CB-521) left the client javadoc and the contract test's own
  assertion still saying protocol 14 / herdr 0.7.0. Updated to 19 / 0.8.0 and
  renamed pingReturnsProtocol14 -> pingReturnsProtocol19. Verified the
  assertion is real by temporarily changing the expected value to 20 (fails),
  then restoring 19 (passes).

- AgentControlContractTest.java: tabCreateInjectsEnvIntoTheSeedShell was
  failing, not skipping, on a host with a live herdr socket. Diagnosed with a
  temporary instrumented run (not committed) that polled the pane every
  200ms before and after sending input: the seed shell reliably takes ~2.5s
  to reach its prompt (measured 3x), while the test's fixed 1000ms sleep
  raced that startup. Input typed too early was swallowed by the shell's own
  startup, leaving the typed line followed by the "Restored session" banner
  and no command output — indistinguishable at a glance from the env map
  never reaching the shell. Once the shell was actually ready, the injected
  env value showed up in ~200ms, ruling out an env-seam defect. Replaced both
  fixed sleeps with bounded polling on the actual conditions (pane text
  settling, then the expected output appearing). Ran the fixed test 3x
  standalone, all green.

- .gitea/workflows/ci.yml: the "Contract tests" step ran exactly one class by
  name (-Dtest=AmqpReplyInboxContractTest), silently excluding every other
  @Tag("contract") test from CI including the herdr ones above -- which is
  how the stale protocol 14 assertion went unnoticed. Changed to
  -Dgroups=contract, which selects the whole tagged group and picks up
  future contract tests automatically.
2026-09-10 17:14:09 +07:00
Dai Ha 2e138a199b CB-634: one shared "fleet" workspace + rename bridged -> fleetd cutover
Two changes ship together here.

1. One shared herdr workspace. The lead and every worker now live in one
   workspace called "fleet", so the operator sees one "session" with many
   windows, not two. Before, the lead sat in a "leads" workspace and workers
   in "bridged-workers", which read as two sessions. The lead is still told
   apart from workers by its exact tab label ("lead: <name>"), so putting them
   in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
   for split layouts; Fleetd now passes an empty exclude set.

2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
   launchd/systemd units, module dir, and MCP mount).
   - Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
   - Log line, comments, docs, and CLAUDE.md updated to say fleetd.
   - Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
     bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
   - Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
     bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
   - Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
     still read as a fallback, and still gitignored.
   - MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
     server name is "fleet". The mount name in the local .mcp.json becomes
     "fleet" (gitignored, not in this commit).
   - Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
     BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.

Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.

Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.

The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).

949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
2026-08-25 04:01:08 +02:00
Dai Ha 8d2893b67e CB-ci: drop host port mapping for rabbitmq service in contract job
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m33s
2026-08-23 05:35:58 +02:00
Dai Ha c1173346ef CB-521: make the AMQP contract test runnable locally and in CI 2026-08-09 05:56:45 +02:00
kevin 6da2a71050 CI: drop upload-artifact — unsupported on this Gitea instance
CI / build (push) Successful in 1m44s
Run 2 built clean (311 tests, BUILD SUCCESS, Maven 3.6.3 on Java 25.0.4 —
JAVA_HOME from setup-java correctly beat the JRE apt pulled in) but the job
still went red on the artifact step:

  GHESNotSupportedError: @actions/artifact v2.0.0+, upload-artifact@v4+ and
  download-artifact@v4+ are not currently supported on GHES.

Gitea Actions presents as GHES, so v4 artifact upload cannot work here. The
artifact was unretrievable regardless, so replace it with a failure-only step
that cats the failing surefire .txt reports into the job log, where they are
readable. Guarded with 'exit 0' so the dump itself can never mask the real
failure.
2026-07-29 23:26:21 +07:00
kevin e32ac39faf CI: install Maven — setup-java provides the JDK only
CI / build (push) Failing after 2m1s
First CI run failed at 'Build and test' with exit code 127 (command not
found): actions/setup-java@v4 provisions a JDK but not Maven, and the runner
image has no mvn on PATH. The sibling lms/alms workflow apt-installs both;
this workflow switched to setup-java for JDK 25 (the image's default-jdk is
too old for maven.compiler.release=25) and dropped the maven install with it.

Adds an explicit Maven install with Apache Maven 3.9.16 (2bdd9fddda4b155ebf8000e807eb73fd829a51d5)
Maven home: /Users/appbuilder/Tool/apache-maven-3.9.16
Java version: 25.0.2, vendor: Oracle Corporation, runtime: /Users/appbuilder/Tool/jdk-25.0.2.jdk/Contents/Home
Default locale: en_VN, platform encoding: UTF-8
OS name: "mac os x", version: "26.5.1", arch: "aarch64", family: "mac" so the log proves which JDK
it resolved — JAVA_HOME from setup-java must win over the JRE apt drags in.
2026-07-29 23:21:55 +07:00
kevin 9daf1ec5ba CB-5xx: Stage 5 hardening — auth, authz+audit, metrics, CI, supervision
Closes out single-host before the cross-host work. Sequenced BEFORE CB-308
deliberately: federation's own gating concern is the trust model, and it
inherits whatever identity shape lands here.

The finding this stage is built around: bridged had exactly ONE security
control, the loopback bind. ConnectionIdentity resolves a worker from its
connection (unforgeable), but every caller that was not a recognised worker
pane fell through to being treated as the PRIMARY -- the most privileged role
on the bus. Latent today; load-bearing the moment a bind widens.

CB-501 auth:
- Role/Principal/CallerResolver: connection identity first, bearer token
  second, ANONYMOUS third. Inverts the old default so absence of identity
  means nothing, not everything.
- Worker identity is never token-gated, so enabling auth cannot lock the
  fleet out of bridge_reply.
- Constant-time token compare (MessageDigest.isEqual).
- validateAuthExposure(): a non-loopback bind under loopback-trust now
  REFUSES TO START. Makes the dangerous config unrepresentable rather than
  merely documented.
- TLS terminates at a reverse proxy by design (D3), not in the JVM.

CB-505 authz + audit, enforced on BOTH entry paths:
- The docs describe MCP as "a thin adapter over the REST core"; at code level
  it is not. BridgeMcp calls MessageService directly, and /mcp is a raw
  servlet on Jetty's context handler that never traverses Javalin's before
  filter. Enforcing only at REST would have left /mcp open.
- Load-bearing rule is own-session-only: a worker may reply/ask only as
  itself. Structurally true over MCP already; over REST the session id in the
  URL path had simply been trusted.
- Audit: JSON lines to a dedicated appender, additivity=false. Never records
  message content -- this bus carries source and prompts.

CB-502 metrics: zero new dependencies. A ~150-line Prometheus text renderer
instead of the specced Micrometer, because this pom already hand-pins
jackson-annotations to reconcile Jackson 2/3, imports a Jetty BOM against
skew, and carries four accepted-CVE advisories -- and CLAUDE.md's mandated
dependency CVE gate could not be run (no JetBrains MCP server connected).
Instrumented at MessageService, the single funnel both surfaces share.

CB-503 CI: .gitea/workflows/ci.yml against the already-running Gitea runner.
Needs no contract-exclusion flag -- the pom's default-excludes profile
already sets excludedGroups=contract, so plain `mvn clean install` IS the
mock-socket surface. Provisions JDK 25 explicitly (runner default-jdk is older).

CB-504 supervision: launchd agent (the real target -- this host is macOS,
there is no systemd) plus a systemd unit for the Linux gateways CB-308 adds.
Ordering directives are advisory, so the actual fix is that startup now waits
up to 30s for the herdr socket and then serves degraded, instead of crashing
into a restart loop on a boot-order race.

Also fixes drift found while surveying:
- bridged.example.yaml documented spawn_ready_timeout_ms in snake_case; config
  binds via plain Jackson with ignoreUnknown, so uncommenting it would have
  been silently dropped and the default kept. Now camelCase, with a test that
  loads the shipped example and one that pins every documented knob's
  spelling -- no test had ever loaded that file.
- Added the 6 shipped-but-undocumented knobs (worktreeRoot, parityOverlay,
  gitTokenEnv, gitHostEnv, configDir, primary:).
- README "Next" listed bridge_ask and session lifecycle as upcoming; both
  shipped long ago.
- docs/CB-301-ext and docs/CB-402 status headers said "design"/"pre-
  implementation" for work already merged.

307 unit/acceptance tests green (was 266), mvn clean install BUILD SUCCESS.
Note: CLAUDE.md's per-file ide_diagnostics gate and the pom Mend.io CVE check
could not be run -- no JetBrains/intellij-index MCP server is connected this
session. mvn clean install is the only gate that ran.
2026-07-29 22:29:26 +07:00