Commit Graph

453 Commits

Author SHA1 Message Date
Dai Ha a052975420 #206: pin the read-only open with a test that actually fails without it
CI / build (pull_request) Successful in 1m18s
CI / contract (pull_request) Successful in 1m44s
setReadOnly(true) is the whole thing keeping fleetd out of the operator's
live 841MB opencode.db, and no test failed when it was removed.

The obvious test does not work. Making the database file unwritable and
checking the read still succeeds passes either way, because SQLite silently
downgrades a read-write open of an unwritable file to read-only. I wrote that
test, watched it pass with the flag removed, and threw it away.

What works: extract a package-private openReadOnly(), then ask that connection
to INSERT and require the refusal. Watched red with the flag removed, green
with it restored.

Also switches the test's INSERT helper to a PreparedStatement -- hand-escaped
SQL in a test is a pattern that gets copied into main code.
2026-08-31 21:46:07 +07:00
Dai Ha 3743789e8d fleetd #206: read opencode session ids from opencode.db (SQLite), not the frozen JSON tree
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 1m44s
opencode migrated its session store to SQLite in January 2026; the JSON tree under
storage/session/<projectID>/ses_*.json stopped being written, so OpenCodeSessionDiscovery
returned null for every member forever, and fleet_spawn{resumeSessionId} was unreachable.

- Add org.xerial:sqlite-jdbc 3.53.4.0, opened read-only (SQLiteConfig.setReadOnly), so it
  never disturbs a live opencode process writing the WAL-mode database.
- Rewrite sessionIdForDirectory to run a parameterized SELECT ... WHERE directory = ?
  ORDER BY time_updated DESC LIMIT 1 against the session table. Still never throws: a
  missing database, a locked/corrupt one, or no matching row all return null.
- Log the silence that let this go unnoticed: WARN once per instance when opencode.db
  itself is missing (the layout moved again), DEBUG when it exists but no row matches
  yet (the normal interim answer right after a spawn).
- Replace the JSON-fixture tests with a synthetic-SQLite-db fixture; delete the tests
  that only proved the old JSON scan worked.
2026-08-31 15:52:06 +07:00
Dai Ha 388aba7632 #137: an answered turn's reply completes its own ticket, instead of a false failure
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m17s
A wait:false delegation whose worker used fleet_ask ended with fleet_poll{ticket}
reporting 'the worker session was released before it replied' -- naming a
worktree, a branch and a snapshot commit, so it read as lost work. The worker had
in fact replied in full.

The ticket guessed the ask rendezvous detached the turn. The real cause is
narrower: answer() (behind fleet_send{turnId}) waits only for the lead's own
bounded MCP call window. A resumed turn doing real work -- edits, a build, a
push, a PR -- routinely outlives it. On timeout answer()'s finally closed the
waiter, so the worker's later fleet_reply found none and fell to the session
inbox, leaving the ticket's future unresolved until fleet_stop forced it FAILED.

reply() now looks for the async task parked on this exact answered turn and
completes it with the real reply. That is safe against completing the wrong
ticket: answer() calls clearAsyncQuestion(turnId, false), so the task keeps its
turnId and stays in asyncTasksByTurn, and hasAsyncQuestion therefore still makes
send() return BUSY for a second async send to that target. At most one candidate
task can exist per target.

abandon() keeps an independent check: if a reply was stranded, a released session
reports REPLIED with that text rather than a failure -- so the recovery hint that
implies lost work never prints once a reply exists.

Verified before merging: build green unpiped, and both new tests drive the full
delegation path (send async, ask, answer, reply, poll the ticket) rather than
handing a Reply to a sink, which is the trap this ticket called out.

Co-authored-by: fleetd worker <worker@ltms.dev>
2026-08-31 14:24:54 +07:00
Dai Ha ad587eafa3 #172: keep the broker URI, password and all, out of every member pane
CI / contract (push) Successful in 1m10s
CI / build (push) Successful in 1m15s
broker.uriEnv names an environment variable holding amqp://user:password@host,
and it was reaching every member. Its name is not credential-shaped -- no TOKEN,
KEY or SECRET in it -- so every name-pattern heuristic missed it, and it sat on
neither credential list.

fleetd already knows the name: the operator wrote it in broker.uriEnv. So derive
the exclusion from the config rather than hoping an operator also remembers to
deny it. coordinator.uriEnv has the same shape and is excluded too; on this host
both resolve to the same variable.

Excluded even when the operator lists the name under memberCredentials.allow:,
following the SSH_AUTH_SOCK precedent. There is no override, because a member has
no legitimate use for the broker password.

Reviewer finding, recorded rather than overstated: this is only a hard guarantee
under policy: allow-list, where the ZDOTDIR scrub runs after the pane's shell has
sourced the operator's chain. Under deny-list the name is removed from the
pre-shell env only, and a login shell re-exports it. That is deny-list's existing
weakness rather than a regression here, but the javadoc now says so plainly
instead of implying a guarantee that path cannot give.

Co-authored-by: fleetd worker <worker@ltms.dev>
2026-08-31 14:21:48 +07:00
Dai Ha 08ce9aef11 #161: resolve a pane by process ancestry, closing a worker->primary escalation
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m52s
PaneLocator's javadoc always claimed it found 'the agent pane whose process
tree contains' a pid. It did not: paneOwnsPid matched only the pane's shell_pid
and its foreground_processes. A process a member spawned -- python3, curl, any
helper opening its own connection to 127.0.0.1:8765 -- matched no pane, so
CallerResolver fell through to loopback-trust and resolved it as the PRIMARY.
A member escalated to lead by shelling out.

terminalForPid now builds the caller's ancestor set once (bounded at 32
generations, with a cycle guard) and matches any ancestor against a pane's pids.
The set is reused across both herdr clients on the CB-185 two-daemon path.

This only ever ADDS matches, which is the safe direction: the failure mode of
the fix is a member correctly restricted, while the failure mode of the bug is a
member acting as the lead. The no-match case still returns null, so the lead --
which maps to a pane named by leaders: -- still resolves as primary.

Ancestry is walked through a new ParentResolver seam so the tests drive it from
a fake pid->parent map rather than spawning real processes.

Co-authored-by: fleetd worker <worker@ltms.dev>
2026-08-31 14:15:39 +07:00
Dai Ha 4ac688b6d9 #164: classify a backend-error scrape as WORKER_FAILED, carrying the whole pane
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m22s
main already shipped the core of #164 in 3bfa828: the MIN_TURN_NANOS floor and
the hard fail on an empty or unreadable scrape. This adds the one case that was
still resolving as a success -- a scrape that reads cleanly but whose content is
the backend's own rejection (e.g. "API Error: 400 invalid request body").

The BACKEND_ERROR pattern is deliberately narrow. A growing list of ad-hoc error
strings rots as backends change their wording, and broader backend-error
surfacing is #164 point 3.

Because the pattern is a heuristic, it also matches a member that forgot
fleet_reply while reporting *about* a backend error. So the failure reason
carries the whole pane tail, not just the matched line: a genuine backend error
reads as before, and a false positive keeps its report instead of losing it.

Checked before merging: 3bfa828 is an ancestor of main; the branch was current
with main; mvn clean install green unpiped (1037 tests, 0 [ERROR] lines); and
each of the 5 new tests fails with the fix commented out.

Co-authored-by: fleetd worker <worker@ltms.dev>
2026-08-31 10:56:40 +07:00
ltms a814d1ef00 #197: measure the async ticket TTL from completion, not from creation
CI / build (push) Successful in 1m15s
CI / contract (push) Successful in 1m17s
Lead-authored and lead-verified: full mvn clean install green at 1032 tests, and the new test proved by restoring the old comparison and watching it fail.
2026-08-31 05:10:46 +02:00
Dai Ha ea98856130 #197: measure the async ticket TTL from completion, not from creation
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 1m32s
pruneTerminalTickets compared the cutoff against createdNanos, so the real
window to collect a reply was "TTL minus however long the task ran". A
delegation that ran longer than the 10-minute TTL was already past the cutoff
the moment it finished, so the next prune destroyed its reply.

That is the normal case here, not an edge case. Real delegated work runs well
past ten minutes. Three workers in one session did, and two of their complete
reports were lost. The reply lives only in Task.future, so pruning it discards
the worker's whole report, and fleet_poll{target} returns [] rather than
holding it — there is no fallback.

Task now stamps completedNanos from a whenComplete hook registered in its
constructor, so every completion path stamps it (a reply, the completion
fallback, a timeout, a failure, an abandon on teardown) without each one having
to remember to. The stamp is a boxed Long, not a long with a sentinel:
System.nanoTime may return any value, so no number can mean "not stamped yet".
A task that is done but not yet stamped is left for the next sweep.

createdNanos had no other reader and is removed.

The TTL still bounds tasks — an uncollected finished ticket is still evicted
once the TTL passes since it finished. Both halves are pinned by a test, and
the first one was proved by restoring the old comparison and watching it fail.

Tests run: 1032, Failures: 0, Errors: 0, Skipped: 0
2026-08-31 10:10:16 +07:00
ltms a49e96835a CB-189: cover every remote, both URLs, and any non-SSH scheme in the credential check
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m11s
Lead-verified: merged onto current main (which already carries #194 and #196), full `mvn clean install` green at 1030 tests. Confirmed no plain `exec` call that reads a remote URL remains.

Review round 2 closed the half-fix: execRedacted was added but applied only to the new calls, leaving the three pre-existing URL readers (lines 143, 171, 334) still copying stdout into exception messages — the exact hole CB-189 named.

Kept the check before the origin strip, on the worker's reasoning: the credential really is in the config at that moment, so WARN-found followed by INFO-fixed is the full audit trail, whereas moving it after would silence the origin case entirely.
2026-08-31 04:32:14 +02:00
ltms 63c19dcba7 CB-185: fix two blockers to switching on memberHerdrSocket
CI / build (push) Successful in 1m31s
CI / contract (push) Successful in 1m19s
Lead-verified: merged with #194 onto an integration branch off main, full `mvn clean install` green at 1025 tests.

Review round 2 fixed the stale ambiguous-pane message and, more importantly, a real abort: probeOwner called list() unguarded, so one unreachable daemon made panes on a different healthy daemon un-stoppable too — the same bug blocker 1 exists to fix, through a new door. Worker proved it by removing the guard and quoting the failure.
2026-08-31 04:30:57 +02:00
ltms 2823349c8e CB-192: fix false credential-gap WARN under allow-list+zsh, split its log guard
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m11s
Lead-verified: merged with #196 onto an integration branch off main, full `mvn clean install` green at 1025 tests. Deny-by-default WARN confirmed byte-identical to main.

Review round 2 fixed a defect I found in round 1: the new INFO asserted that gap names were not on the derived allow-list without ever checking, so a name derived from a profile's tokenEnv/gitTokenEnv/env: would be reported as safe while the member actually inherited it. Now split with MemberEnvAllowList.keeps — the same predicate the generated scrub evaluates.
2026-08-31 04:30:48 +02:00
Dai Ha 6fc301d62c CB-189 review fix: redact the three pre-existing URL-reading exec calls
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m42s
Review found that execRedacted was applied only to the new remote-enumeration code
and left three pre-existing calls reading remote.origin.url through the plain,
unredacted exec: removeUserInfoFromHttpsOrigin, requireCredentialFreeHttpsOrigin, and
configureHttpsUrlRewriteForSshOrigin. A non-zero exit or timeout on any of those could
still have copied the credentialed URL into a WorktreeException message. Switches all
three to execRedacted; the set-url write in removeUserInfoFromHttpsOrigin is left on
plain exec with a comment explaining why (it writes the already-stripped URL, not a
read).

Widens the shared exec(Map, boolean, String...) overload to package-private, the same
test-seam pattern already used by the afterWorktreeAdded constructor parameter, and
adds a test that drives it directly with a synthetic failing command whose stdout
carries a marker (passed via env, not argv, so the always-printed command line can't
carry it) and asserts the marker never reaches the exception message.
2026-08-31 09:29:45 +07:00
Dai Ha bc99d64786 CB-185: fix ambiguous-pane message and unreachable-daemon abort in probeOwner (review)
CI / build (pull_request) Successful in 1m6s
CI / contract (pull_request) Successful in 1m5s
Lead review of PR #196 found two issues in CompositePeerLauncher.probeOwner:

1. The "more than one daemon claims this pane" throw kept the old
   pre-fix message ("no owning herdr daemon was recorded"), which was
   only true of the code it replaced. Reworded to say what actually
   happened: N configured herdr daemons report this pane, so it is
   genuinely ambiguous. Updated the one test pinning the old string.

2. probeOwner let list() propagate straight out of the probe loop, so
   one unreachable daemon aborted the whole probe and made a pane on a
   DIFFERENT, healthy daemon un-stoppable too — resurrecting the exact
   bug blocker 1 fixes. Now catches HerdrException per daemon, logs the
   exception class only, and treats that daemon as not knowing the pane
   so probing continues. New test proves this: verified it fails with
   the try/catch removed (HerdrException propagates and the stop that
   should succeed via the healthy daemon throws instead), then restored.

Full mvn clean install: 1019 tests, 0 failures, 0 errors.
2026-08-31 09:28:29 +07:00
Dai Ha 615af4ed0a CB-192 review fix: split the allow-list gap by what the scrub actually keeps
CI / build (pull_request) Successful in 1m5s
CI / contract (pull_request) Successful in 1m20s
Lead review on PR #194 found that the allow-list INFO wording claimed the
whole gap ("credential-shaped names on neither known: nor allow:") is blanked
by the scrub, without checking that against effectiveAllowed. effectiveAllowed
is a SUPERSET of known+allow — MemberEnvAllowList.derive also unions in every
profile's gitTokenEnv/gitHostEnv/tokenEnv/env: keys, and derivedAllowedNames
further unions in the spawn's own env keys — so a gap name can still be kept
by the derived list (e.g. a profile's tokenEnv names it) and reach the member
unblocked while the INFO said "no member pane keeps them". That inversion is
exactly what #192 exists to remove.

logCredentialGap now splits the gap with MemberEnvAllowList.keeps (the same
predicate the generated scrub itself evaluates, so this cannot drift from
what the scrub does): names it keeps get a WARN, guarded by the same
unprotectedGapLogged flag as the deny-by-default case (same severity — a name
reaching a member unprotected is equally serious either way); names it
blanks keep the existing INFO, guarded by allowListGapLogged. The
deny-by-default WARN text and the non-zsh fallback are untouched.

Added allowListWarnsWhenTheDerivedAllowListKeepsAnUncoveredName and
allowListSplitsAMixedGapBetweenTheWarnAndTheInfo to ClaudeCodeLauncherTest.
2026-08-31 09:28:10 +07:00
Dai Ha 045d229728 CB-185: fix two blockers to switching on memberHerdrSocket (#185)
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m39s
1. CompositePeerLauncher.stop() was permanently un-stoppable for any
   member that survived a daemon restart, because spawnedBy is in-memory
   only. On a cache miss with more than one configured herdr daemon, probe
   each distinct daemon's agent.list() for the pane instead of refusing
   outright: exactly one owner routes and caches; zero owners is treated
   as already-stopped (a no-op, matching the tolerance HerdrPeerLauncher
   already gives an already-gone pane); more than one owner is the
   genuine per-daemon-pane-id ambiguity and still throws.

2. FleetApp#healthz always reported the LEAD daemon's herdr version/
   protocol even when a second (member) daemon was configured, so a
   member-daemon protocol mismatch was invisible behind a green
   /healthz while every spawn silently failed. Added a separate "member"
   key alongside the unchanged "herdr" key, and a "protocolMismatch"
   flag when the two differ. Verified scripts/redeploy-fleetd.sh and
   scripts/rename-checkout.sh only check the HTTP status code and print
   the body verbatim — neither parses a specific field — so adding a key
   is safe.

Both fixes are covered by tests written to fail without the fix
(verified by reverting each fix and watching the new tests fail, then
restoring). Full `mvn clean install`: 1018 tests, 0 failures, 0 errors.
2026-08-31 09:20:14 +07:00
Dai Ha d1fd5700f5 CB-189: cover every remote, both URLs, and any non-SSH scheme in the credential check
CI / build (pull_request) Successful in 1m7s
CI / contract (pull_request) Successful in 1m7s
GitWorktrees only ever inspected origin's HTTPS fetch URL for embedded credentials. A
credential on any other remote, on a pushurl, or on a plain http:// URL passed through
unreported. Adds an additive, reporting-only check that enumerates every remote and both
its fetch and push URLs, flagging non-empty user-info on any non-SSH-family scheme.

The existing origin/https strip-and-refuse behaviour is untouched. The new check is
wrapped so it can never abort a provision, and on failure logs only the exception's
class, never its message, since the enumerating `git remote` call is not redacted.
Also adds execRedacted, an exec variant that never copies captured stdout into a
WorktreeException message, for commands whose stdout may itself be a credentialed URL.
2026-08-31 09:17:54 +07:00
Dai Ha 2d55b0b9a5 #168: correct the audit's MISSING claim — the feature is documented, its index row was not
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m42s
The memberHerdrSocket section exists at 11-Features.md:2174; what was absent was
its row in the index table. My omission when I added the section. Wiki fixed at
b24965c.
2026-08-31 09:17:41 +07:00
Dai Ha d89ae94a2e CB-192: fix false credential-gap WARN under allow-list+zsh, split its log guard
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 1m7s
logCredentialGap(creds) always emitted the WARN wording ("every member pane
inherits them UNBLOCKED"), even under memberCredentials.policy: allow-list on
a zsh login shell, where the generated ZDOTDIR scrub genuinely blanks the
name. The line reported the control working as though it were a hole.

Pass an effectiveAllowed set instead: null keeps the WARN (deny-by-default,
and the allow-list non-zsh fallback, where nothing is ever scrubbed); the
derived allow-list set (only reachable after applyEnvironmentAllowListPolicy's
own zsh gate) selects a new INFO wording that says the scrub will blank the
name instead of claiming it is inherited unblocked.

Also split the single credentialGapLogged AtomicBoolean into two guards
(unprotectedGapLogged / allowListGapLogged) — one per report kind. Since
memberCredentials is a live, re-read-per-spawn supplier, a shared flag let a
harmless allow-list INFO on one spawn permanently suppress a later spawn's
real deny-by-default WARN after a policy reload.

Fixes gitea #192.
2026-08-31 09:17:20 +07:00
ltms e6193c4098 Merge pull request '#168: audit current wiki snapshot' (#193) from worker/cb-168-wiki-audit-3ef3d1-3 into main
CI / build (push) Successful in 1m6s
CI / contract (push) Successful in 1m5s
2026-08-31 04:17:07 +02:00
Dai Ha b66f0677ed #168: audit current wiki snapshot
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m37s
2026-08-31 09:14:22 +07:00
ltms 23ada1981e CB-185: route members to a separate herdr daemon (#186)
CI / build (push) Successful in 1m6s
CI / contract (push) Successful in 10m30s
2026-08-29 01:24:31 +02:00
ltms a22480c117 CB-185: route PaneLocator, StatusRefiner and FleetApp to the right herdr daemon (#188)
CI / build (pull_request) Successful in 1m3s
CI / contract (pull_request) Successful in 1m7s
2026-08-29 01:24:25 +02:00
Dai Ha 24f404f989 CB-185: fix three connection-identity/status/health gaps a second herdr daemon exposes
memberHerdrSocket splits lead operations from member operations onto two herdr
daemons. Three seams still assumed one shared daemon and broke silently when the
two clients differ (all three collapse to today's behaviour when they are the
same object):

1. ConnectionIdentity's PaneLocator was pinned to the member daemon only, so a
   lead's own MCP connection (which lives on the LEAD daemon) resolved to
   terminal == null, breaking fleet_reply/fleet_ask/fleet_whoami for a lead.
   PaneLocator now searches the lead client first, then the member client.

2. StatusPoller's StatusRefiner was pinned to the member daemon, so refining an
   UNKNOWN status for a lead target read the wrong daemon's pane content and
   never left UNKNOWN, wedging status-gated delivery to that lead forever.
   StatusRefiner gained a refine(target, raw, control) overload and the poller
   now refines through the same AgentControl the raw status was sampled from.

3. FleetApp was constructed with the raw lead-only herdr client, so /healthz
   stayed green while the member daemon was down (every spawn then fails
   invisibly) and GET /sessions silently dropped every member workspace.
   FleetApp now takes both clients: healthz requires both to answer, sessions
   merges workspaces from both.

Each fix has a test proven to fail without it (verified by reverting the
production change and re-running): FleetdConnectionIdentityConstructionTest /
FleetdFleetAppConstructionTest assert the actual Fleetd.java wiring (the same
technique as FleetdHerdrControlConstructionTest); StatusPollerRoutingTest and
the new PaneLocatorTest/FleetAppTwoDaemonTest cases exercise the real
production classes end to end rather than a hand-built object graph.
2026-08-29 06:20:32 +07:00
ltms a237fbff9d #185: refuse an unowned paneId when more than one herdr daemon could own it (#187)
CI / build (push) Successful in 1m4s
CI / contract (push) Successful in 1m17s
2026-08-29 01:10:21 +02:00
Dai Ha 31d5516991 #185 review: keep the stop owner on failure, key list() by daemon
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m30s
Three fixes on top of the pane-id PR, from my own read and the reviewer's:

- stop() removed the spawnedBy record BEFORE the delegate accepted the stop. A
  delegate that threw left the pane alive with its owner forgotten, so the retry
  fell into the ambiguous branch and refused the id for good. Remove after.
- list() deduplicated on the raw pane id. Pane ids are per-daemon counters, so
  two daemons can each hold w1:p1 on different panes, and one of the two real
  agents was silently dropped from fleet_list and every view built on it. The
  key is now (owning daemon, pane id). Delegates sharing one daemon still
  collapse, which is what the dedupe was for.
- The class javadoc still stated the single-herdr-connection premise as fact,
  next to the bullet this PR had just corrected for stop(). Fixed there too.

Also drops a redundantly qualified java.util.Collections.

Tests: 993 run, 0 failures, BUILD SUCCESS.
2026-08-29 06:09:43 +07:00
Ha Trong Dai 17af61e8dd CB-185: route message status by target
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m40s
2026-08-28 09:45:39 +07:00
Ha Trong Dai 6af87b6ad6 CB-185: share routed herdr controls
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m40s
2026-08-28 09:42:45 +07:00
Ha Trong Dai 5ba05d0bdb #185: count herdr owners for stop fallback
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m8s
2026-08-28 09:42:24 +07:00
Ha Trong Dai fc655e78c2 CB-185: route members to separate herdr
CI / contract (pull_request) Successful in 37s
CI / build (pull_request) Successful in 1m28s
2026-08-28 09:37:44 +07:00
Ha Trong Dai 25726a5ae7 #185: reject ambiguous unowned pane ids 2026-08-28 09:36:39 +07:00
Dai Ha 11c3ff67b6 fleets-status: fleet01 IS ssh-reachable; correct the 'denied' claim
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m12s
The skill said SSH to fleet01 is denied, so every report wrote 'not
reachable' for that fleet's daemon facts. That is true only for the user
dai.ha. The host alias fleet01 maps to user ltms and key auth works.

Checked 2026-08-28 while measuring #185: ssh fleet01 connects, and ltms
has passwordless sudo there. So fleet01's PID, uptime, jar and /healthz
can be reported over SSH even though its REST port is unreachable.
2026-08-28 09:06:54 +07:00
Dai Ha d867c87100 #184: correct the false ssh-agent premise in the URL-rewrite javadoc
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m37s
The javadoc said a member cannot authenticate at all once memberCredentials
blocks SSH_AUTH_SOCK, "there is no private key file on this host, only an
ssh-agent socket". That is wrong, and it was written after looking only in
~/.ssh, which holds nothing but Include lines.

Measured: ssh -G git.ltms.dev resolves an IdentityFile under the shared-env
directory. That file exists, is readable by this user, and has no passphrase.
A live member with SSH_AUTH_SOCK blanked pushed to the forge over SSH.

The rewrite itself is unchanged and still worth having. Only its stated reason
was wrong: it routes a member through its own scoped token instead of the
operator's ssh identity, which is what makes a member's pushes attributable
and revocable. It is not what stands between a member and the forge.
2026-08-28 06:35:46 +07:00
Dai Ha b0c4cedfab #157: redact remote-URL user-info before it reaches a log
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m25s
The four log lines added with the worktree HTTPS rewrite echoed the origin
URL verbatim, and one of them echoed the ssh:// authority, which carries
user-info. An ssh authority is normally just git@, so in practice this
changes nothing -- but a remote URL is not obviously a credential channel,
and that is precisely why one has leaked here three times (#157, #182).

Redact at the log call, not after it surprises someone.
2026-08-28 06:24:08 +07:00
ltms 85417d5215 #157: rewrite an SSH origin to HTTPS inside the provisioned worktree only
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m37s
Git never consults a credential.helper for an SSH transport, so #177's helper was inert on this repo — whose origin is ssh://. Once allow-list policy blocks SSH_AUTH_SOCK, a member on an SSH origin cannot authenticate at all: there is no private key file on this host, only an agent socket.

A worktree-scoped `url.<https>.insteadOf <ssh>` gives the member HTTPS for fetch and push while the primary checkout keeps SSH untouched. Host and port are parsed from the origin, never hardcoded — a test with a synthetic host proves it. The scp-like shorthand is left alone deliberately, since its host:path split is defined by ssh_config aliases rather than URI syntax.

Verified by the lead in an independent worktree: Tests run: 986, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-08-28 01:22:57 +02:00
Dai Ha 4accc746bd #157: rewrite SSH origin to HTTPS in the worktree so the credential helper is reachable
CI / build (pull_request) Successful in 1m10s
CI / contract (pull_request) Successful in 1m16s
2026-08-28 06:20:10 +07:00
ltms 21c4c8cbef #157: keep the forge token out of git config, via an environment credential helper
CI / build (push) Successful in 1m4s
CI / contract (push) Successful in 47s
The member credential scrub removes environment variables. It cannot remove a token written into git config inside the repo the member works in, so `git remote -v` handed a member a credential it was deliberately not given.

Provisioning now strips HTTPS user info from the origin before `git worktree add`, refuses the worktree if user info survives, and configures a per-worktree credential helper that reads WORKER_GITEA_TOKEN at call time. Nothing is persisted.

The helper emits BOTH username and password, and resets the inherited helper list first. An earlier revision emitted only `username=`, which made git fall through to the next helper — on a Mac that is osxkeychain, so a member would have authenticated with the operator's stored credential while every test passed and `git remote -v` looked clean. See #182.

`worktreeCredentialHelperCompletesWithoutUsingAnInheritedHelper` plants a synthetic operator helper in an isolated global config and proves the worktree helper wins. The worker confirmed it fails when the reset is removed. All credential tests pin GIT_CONFIG_GLOBAL and GIT_CONFIG_SYSTEM so they can neither read nor write real credentials.

Verified by the lead in an independent worktree: Tests run: 973, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-08-28 01:10:38 +02:00
ltms 430f5b0dae #164: never resolve a send with an empty scrape or a sub-floor turn
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m27s
A turn that dies on a backend error produces the same working -> idle transition as a real one, just faster and with nothing on screen. The resolver accepted that as a completed turn and handed the caller HTTP 200 with an empty reply, so a lost turn and a successful empty answer were indistinguishable.

Now: an empty or unreadable scrape fails, naming the member; and a BUSY -> DONE inside MIN_TURN_NANOS (2s) fails as a crash signature.

One existing test encoded the bug — it asserted a failed scrape resolved as a success carrying "" — and has been inverted rather than worked around.

Verified by the lead in an independent worktree: Tests run: 973, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-08-28 01:10:27 +02:00
ltms 42731833d0 CB-633: union memberCredentials.allow into the member env allow-list
CI / contract (push) Successful in 42s
CI / build (push) Successful in 1m38s
`policy: allow-list` silently ignored every name an operator wrote under `allow:` unless a profile happened to carry it too, so turning the policy on would have blanked credentials working members depend on. Derivation now unions the operator's list.

`SSH_AUTH_SOCK` stays governed only by `sshAuthSock`, even when listed under `allow:` — it is a live handle to the operator's ssh-agent, not a value.

Adds one INFO line per allow-list spawn, `member credentials: allowed N of M`, emitted only after the shell gate so it can never report coverage on a path where the scrub does not run.

Verified by the lead in an independent worktree: Tests run: 976, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-08-28 01:10:18 +02:00
ltms 65acf066ad #154: pin the AMQP reply inbox prefetch bound
CI / contract (push) Successful in 1m26s
CI / build (push) Successful in 1m39s
No behaviour change. #154 supposed that ownership drains a whole queue into the in-memory `held` map, making `x-max-length` and per-message TTL decorative. Measurement says otherwise: `basicConsume` is manual-ack, `deliverCallback` acks only duplicates, and `basicQos` is set on the one shared channel before any consumer starts — so total `held` is bounded by the prefetch window across all targets.

Adds a fake-broker test that drives the real `own()` path and fails if receipt ever starts acking, plus a javadoc line naming the prefetch window at the point of first mention.

Verified by the lead in an independent worktree: Tests run: 971, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-08-28 01:06:43 +02:00
Dai Ha ee8f570fd7 #157: isolate worktree credential helpers
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 2m46s
2026-08-28 06:06:12 +07:00
Dai Ha fa3f910d44 #154: pin AMQP reply inbox prefetch
CI / build (pull_request) Successful in 2m16s
CI / contract (pull_request) Successful in 2m18s
2026-08-28 06:03:54 +07:00
Dai Ha 46ac6e4e38 #157: use environment git credential helper
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Successful in 1m37s
2026-08-28 06:02:09 +07:00
Dai Ha 3bfa82839b fleetd#164: an empty or suspiciously fast scrape must fail, never resolve as a success
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Successful in 1m20s
CompletionResolver.resolve() used to hand the caller a successful "" reply whenever a
turn's scrape came back empty (whether the read failed, or genuinely produced nothing),
making a lost turn indistinguishable from a real empty answer. It also had no way to
tell a crashed backend's near-instant BUSY -> DONE transition apart from a genuine
completion.

Add MIN_TURN_NANOS (2s), a named floor below which a completed turn is treated as a
crash signature and failed rather than resolved as a reply. Fail on any empty scrape
(read failure or a clean-but-empty read) instead of resolving with "". Both failures
name the member and carry whatever is on the pane for context.

Thread an injectable LongSupplier clock through CompletionResolver (matching the
SessionManager/MessageService nowNanos pattern) so the floor is testable without a
real sleep.
2026-08-28 06:01:00 +07:00
Dai Ha 65ccf2e4ad CB-633 follow-up: only log allowed N of M when the scrub actually runs
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 1m44s
The coverage line was logged before the zsh gate, so a non-zsh
spawn (where nothing is scrubbed — overlayBlockedCredentials is the
fallback instead) printed 'allowed N of M' as if the derived
allow-list scrub had run. Move the log after the gate so it only
fires on the path that actually generates the ZDOTDIR scrub; the
non-zsh fallback keeps logCredentialGap's WARN as its only signal.

Added a test proving no 'allowed N of M' line is emitted on the
non-zsh fallback, through the real HerdrPeerLauncher#spawn path.
2026-08-28 06:00:38 +07:00
ltms 7a3b27f76f #150: report lead readiness from the delivery gate
CI / contract (push) Successful in 1m6s
CI / build (push) Successful in 1m7s
Share one deliverability predicate between Injector and FleetApp, so the status endpoint reports the same answer the injector acts on instead of re-deriving it from one of that predicate's two inputs.

Verified by the lead in an independent worktree: Tests run: 971, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-08-28 00:59:44 +02:00
Dai Ha 82e7be564c CB-633 follow-up: union memberCredentials.allow into the derived env allow-list
CI / build (pull_request) Successful in 1m7s
CI / contract (pull_request) Successful in 1m16s
MemberEnvAllowList.derive only ever looked at profile fields, so
memberCredentials.allow: was silently ignored under
policy: allow-list — turning the policy on would have blanked
credentials working members already depended on.

- derive(profiles, configuredAllow) unions memberCredentials.allow
  into the derived set, with SSH_AUTH_SOCK explicitly excluded from
  that union (it stays governed only by sshAuthSock: allow).
- HerdrPeerLauncher threads MemberCredentials.allowSet() into the
  derivation instead of calling the profiles-only overload.
- Added a per-spawn INFO log 'member credentials: allowed N of M'
  (N/M from the daemon's own env, the existing hostEnvNames proxy),
  never logging a blocked name or a value.
2026-08-28 05:56:27 +07:00
Dai Ha c5e24197bf #150: report lead readiness from delivery gate
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m37s
2026-08-28 05:54:40 +07:00
Dai Ha bcb402b688 #157: convert forge worktree origins to SSH
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m11s
2026-08-28 05:54:15 +07:00
Dai Ha 7c4170ff6d CB-643: join the message-layer evidence to the health monitor
CI / build (push) Successful in 1m8s
CI / contract (push) Successful in 1m9s
CB-640 published the three message-layer facts and CB-641 wired the herdr
and time ones. This joins them, so every HealthSnapshot field now carries
real evidence and the NOT_YET_OBSERVED placeholder is gone. That constant
is what made 8 of the 9 fault states unreachable, GONE and NEVER_READY
included, which is why CB-580's failTarget never fired.

hasOrphanedDelegation is a true snapshot, but it can read true for one
tick during an ordinary race: an async ticket exists before its virtual
thread reaches rendezvous.open, so for that instant nothing is accepted or
queued behind it. decide maps the field straight to DELEGATION_ORPHANED
with no smoothing, so one racy read would log a fault that clears on the
next tick. The monitor now requires two consecutive observations. That
costs one interval on a real orphan and removes the false positive.

Two tests drive real ticks against a genuinely orphaned ticket (an
unanswered fleet_ask that lapsed back to PENDING), not the seam: one tick
reports nothing, two report once, and a single clean tick in between
resets the streak.

Also correct two config comments. paneProbeIntervalSeconds is parsed and
read by nothing, so its "minimum 60" note promised a floor that does not
exist.

970 tests green.
2026-08-27 22:15:39 +07:00
Dai Ha a6095743f0 Merge CB-641: wire herdr and time evidence into the fleet health monitor 2026-08-27 22:06:50 +07:00