Compare commits
8 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 0efe1567c0 | |||
| 2124e043ce | |||
| e4f3620acb | |||
| 032a59a34d | |||
| e689090024 | |||
| 1cc34888fd | |||
| 0331ecd5d3 | |||
| 831a918c30 |
@@ -81,9 +81,9 @@ below are the procedure — run them in order, every task, not only the big ones
|
||||
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{ticket, msgId}`. Answer a worker's `bridge_ask`
|
||||
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
|
||||
`bridge_status`, never by reading its terminal.
|
||||
6. **Verify yourself.** Re-run the build and the checks. A worker mounts only the bridge MCP and
|
||||
cannot run your other tooling, and a piped command (`… | tail`) hides failures behind a zero
|
||||
exit — never promote a worker's "clean" to a fact.
|
||||
6. **Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
|
||||
forge tools it appears to have hold a blocked credential and fail, and a piped command
|
||||
(`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a fact.
|
||||
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
|
||||
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
|
||||
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
|
||||
@@ -155,9 +155,13 @@ you.
|
||||
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
|
||||
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
|
||||
lead with its end cut off.
|
||||
5. **Report honestly.** State only what you actually ran and its real output, including failures.
|
||||
You mount **only** the bridge MCP — the primary's other servers (IDE, forge, docs) are not yours,
|
||||
so never claim the result of a check you had no way to run.
|
||||
5. **Report honestly.** State only what you actually ran and its real output, including failures,
|
||||
and never claim the result of a check you had no way to run. **Measure your own tools; do not
|
||||
assume them.** What you mount depends on your backend: an opencode member gets the bridge and
|
||||
nothing else, while a Claude Code member also inherits the operator's user-scope MCP servers,
|
||||
which the bridge never chose for you. Two rules follow. The primary's IDE tooling is still not
|
||||
yours, whatever you see. And **a mounted tool is not a working tool** — the forge server you may
|
||||
find there holds a deliberately blocked credential and fails every call, by design.
|
||||
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
|
||||
project marks as not-yours-to-commit.
|
||||
|
||||
|
||||
@@ -781,10 +781,43 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
* PATH} seeding already depends on the overlay reliably replacing an inherited value (see its
|
||||
* javadoc), and that is only demonstrated for a non-blank value, so this reuses the same,
|
||||
* proven-reliable shape rather than the unverified one.
|
||||
*
|
||||
* <p><b>MEASURED ON A LIVE PANE, 2026-08-15: this sentinel alone does NOT hold.</b> The overlay
|
||||
* itself works — {@code GITEA_TOKEN} is injected here, is exported by no shell file, and does
|
||||
* reach the pane. The sentinel loses one step later. A herdr pane runs a <em>login</em> shell,
|
||||
* {@code ~/.zprofile} sources {@code ${SHARED_ENV}/tools/secrets.sh}, and that file does a plain
|
||||
* unconditional {@code export GITEA_ACCESS_TOKEN=...}. A login shell overwrites a value already
|
||||
* in the environment, so the real admin token is put back over this sentinel before the member
|
||||
* process ever starts. That defeat applies to <em>every</em> name {@code secrets.sh} exports,
|
||||
* and no launcher-side overlay can win against it.
|
||||
*
|
||||
* <p>So this constant is not the control on its own — {@link #MEMBER_MARKER} is the other half.
|
||||
* Keeping the sentinel is still worth it: it is correct for any peer kind whose pane does not
|
||||
* start a login shell, and it makes the intent explicit at the one place every adapter passes.
|
||||
*/
|
||||
private static final String BLOCKED_GITEA_ACCESS_TOKEN =
|
||||
"blocked-by-bridged-cb592-see-gitea-issue-77";
|
||||
|
||||
/**
|
||||
* CB-592: marks a pane as a bridged member so a shell startup file can decline to export
|
||||
* operator-only credentials into it (gitea issue #77).
|
||||
*
|
||||
* <p>This name is deliberately one that {@code secrets.sh} never exports, which is exactly why
|
||||
* it survives the login shell that wipes {@link #BLOCKED_GITEA_ACCESS_TOKEN}. The mechanism is
|
||||
* measured, not assumed: {@code GITEA_TOKEN} is injected the same way, is absent from a login
|
||||
* shell of its own, and was observed set inside a live member pane.
|
||||
*
|
||||
* <p>It is a no-op until the operator guards the export, which is a one-line change in a file
|
||||
* this repo does not own and must not edit unasked:
|
||||
*
|
||||
* <pre>{@code
|
||||
* [ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...
|
||||
* }</pre>
|
||||
*
|
||||
* <p>Setting the marker now costs nothing and means that edit is the whole remaining fix.
|
||||
*/
|
||||
static final String MEMBER_MARKER = "BRIDGED_MEMBER";
|
||||
|
||||
/**
|
||||
* Seed a worker's environment (CB-511): the daemon's own {@code PATH}, then the profile's
|
||||
* {@code env:} entries, then the CB-592 admin-token shadow.
|
||||
@@ -802,10 +835,11 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
|
||||
* {@code baseUrl} and nothing else.
|
||||
*
|
||||
* <p>The CB-592 shadow is put in <em>last</em>, after the profile's own {@code env:}, so no
|
||||
* profile — present or future — can restore the admin token by naming it in config. This is
|
||||
* the one place the shadow is applied: every {@code buildLaunch} in every adapter calls this
|
||||
* first, so a new profile, and a peer kind not yet written, gets it for free.
|
||||
* <p>The CB-592 shadow and marker are put in <em>last</em>, after the profile's own
|
||||
* {@code env:}, so no profile — present or future — can restore the admin token, or hide that
|
||||
* the pane is a member, by naming either in config. This is the one place both are applied:
|
||||
* every {@code buildLaunch} in every adapter calls this first, so a new profile, and a peer
|
||||
* kind not yet written, gets them for free.
|
||||
*/
|
||||
protected Map<String, String> baseEnv(BridgedConfig.Profile cfg) {
|
||||
Map<String, String> workerEnv = new LinkedHashMap<>();
|
||||
@@ -817,6 +851,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
workerEnv.putAll(cfg.env());
|
||||
}
|
||||
workerEnv.put("GITEA_ACCESS_TOKEN", BLOCKED_GITEA_ACCESS_TOKEN);
|
||||
workerEnv.put(MEMBER_MARKER, "1");
|
||||
return workerEnv;
|
||||
}
|
||||
|
||||
|
||||
@@ -234,7 +234,14 @@ public final class BridgedApp {
|
||||
List<Map<String, Object>> out = sessions.roster().stream()
|
||||
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
|
||||
.toList();
|
||||
ctx.status(200).json(Map.of("workers", out));
|
||||
Map<String, Object> body = new LinkedHashMap<>();
|
||||
body.put("workers", out);
|
||||
// CB-586: operator visibility for the refs/wip snapshot store without shelling into the
|
||||
// repo — how many snapshot refs exist and roughly what they cost. Present only once a
|
||||
// worktree session has established the repo, so a never-snapshotted fleet reports nothing.
|
||||
sessions.wipRefs().ifPresent(st -> body.put("wipRefs",
|
||||
Map.of("count", st.count(), "costBytes", st.costBytes())));
|
||||
ctx.status(200).json(body);
|
||||
}
|
||||
|
||||
/** The configured worker profiles and which one a no-argument spawn uses. */
|
||||
|
||||
@@ -12,9 +12,12 @@ import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.nio.file.StandardCopyOption;
|
||||
import java.security.SecureRandom;
|
||||
import java.util.ArrayList;
|
||||
import java.util.HashSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Optional;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.stream.Collectors;
|
||||
@@ -299,6 +302,129 @@ public final class GitWorktrees implements Worktrees {
|
||||
return index;
|
||||
}
|
||||
|
||||
/**
|
||||
* One {@code refs/wip/<branch>} snapshot ref as read by {@link #listWipRefs}: its full ref name,
|
||||
* the snapshot commit's sha, and that commit's committer time in unix millis (the age of the
|
||||
* snapshot — a snapshot is written once and never rewritten, so the commit date is the ref's).
|
||||
*/
|
||||
private record WipRef(String refName, String sha, long committerMillis) {
|
||||
String branch() {
|
||||
return refName.substring("refs/wip/".length());
|
||||
}
|
||||
}
|
||||
|
||||
@Override
|
||||
public WipRefStats wipRefs(String repoRoot) {
|
||||
List<WipRef> refs = listWipRefs(repoRoot);
|
||||
long costBytes = 0;
|
||||
for (WipRef ref : refs) {
|
||||
costBytes += treeSize(repoRoot, ref.sha());
|
||||
}
|
||||
return new WipRefStats(refs.size(), costBytes);
|
||||
}
|
||||
|
||||
@Override
|
||||
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
|
||||
// The rule is documented on Worktrees#pruneWipRefs: delete only a snapshot whose tree
|
||||
// content is already reachable from main AND that is older than minAgeMillis. Reachability
|
||||
// is the floor that keeps a worker's last copy; the age floor keeps a just-written snapshot
|
||||
// from being swept while a lead may still be looking at it.
|
||||
List<WipRef> refs = listWipRefs(repoRoot);
|
||||
if (refs.isEmpty()) {
|
||||
return 0;
|
||||
}
|
||||
long nowMillis = System.currentTimeMillis();
|
||||
// Resolve what main carries once per sweep, not once per ref.
|
||||
Set<String> mainObjects = reachableObjectsFromMain(repoRoot);
|
||||
int deleted = 0;
|
||||
for (WipRef ref : refs) {
|
||||
long ageMillis = nowMillis - ref.committerMillis();
|
||||
if (ageMillis <= minAgeMillis) {
|
||||
continue; // too recent — never swept, even if it looks recoverable (CB-586)
|
||||
}
|
||||
String tree = exec("git", "-C", repoRoot, "rev-parse", ref.sha() + "^{tree}").trim();
|
||||
if (!mainObjects.contains(tree)) {
|
||||
// Last copy of the snapshot's content — the worker's work exists nowhere else.
|
||||
// Never delete automatically (CB-586 criterion 2).
|
||||
continue;
|
||||
}
|
||||
exec("git", "-C", repoRoot, "update-ref", "-d", ref.refName());
|
||||
deleted++;
|
||||
log.info("pruned snapshot ref refs/wip/{} commit={} (age {}h): its tree is already "
|
||||
+ "reachable from main, so the work is preserved; recover from reflog via "
|
||||
+ "git update-ref refs/wip/{} {}",
|
||||
ref.branch(), ref.sha(), TimeUnit.MILLISECONDS.toHours(ageMillis),
|
||||
ref.branch(), ref.sha());
|
||||
}
|
||||
return deleted;
|
||||
}
|
||||
|
||||
/**
|
||||
* Every {@code refs/wip/*} ref (see {@link WipRef}). The committer date is read as a unix
|
||||
* count of seconds and converted to millis. {@code %00} (NUL) separates the fields because a
|
||||
* branch name may contain spaces.
|
||||
*/
|
||||
private List<WipRef> listWipRefs(String repoRoot) {
|
||||
String out = exec("git", "-C", repoRoot, "for-each-ref",
|
||||
"--format=%(refname)%00%(objectname)%00%(committerdate:unix)", "refs/wip/");
|
||||
List<WipRef> refs = new ArrayList<>();
|
||||
for (String line : out.split("\\R")) {
|
||||
if (line.isBlank()) {
|
||||
continue;
|
||||
}
|
||||
String[] parts = line.split("\u0000", -1);
|
||||
if (parts.length == 3 && !parts[1].isBlank()) {
|
||||
refs.add(new WipRef(parts[0], parts[1], Long.parseLong(parts[2]) * 1000L));
|
||||
}
|
||||
}
|
||||
return refs;
|
||||
}
|
||||
|
||||
/**
|
||||
* The set of object shas reachable from {@code main}, or an empty set when {@code main} cannot
|
||||
* be resolved. An empty set is the safe direction: the retention sweep then concludes nothing
|
||||
* is recoverable, so it deletes nothing — a repo with no {@code main} must never cause a
|
||||
* worker's last copy of a snapshot to be dropped on a reachability misreading.
|
||||
*/
|
||||
private Set<String> reachableObjectsFromMain(String repoRoot) {
|
||||
if (exitCode("git", "-C", repoRoot, "rev-parse", "--verify", "main") != 0) {
|
||||
log.debug("refs/wip retention: no 'main' ref in {} — treating nothing as reachable", repoRoot);
|
||||
return Set.of();
|
||||
}
|
||||
String out = exec("git", "-C", repoRoot, "rev-list", "--objects", "main");
|
||||
Set<String> objects = new HashSet<>();
|
||||
for (String line : out.split("\\R")) {
|
||||
if (line.isBlank()) {
|
||||
continue;
|
||||
}
|
||||
int sp = line.indexOf(' ');
|
||||
objects.add(sp < 0 ? line : line.substring(0, sp));
|
||||
}
|
||||
return objects;
|
||||
}
|
||||
|
||||
/** Approximate cost of a snapshot: the sum of every blob's size in its committed tree. */
|
||||
private long treeSize(String repoRoot, String sha) {
|
||||
String out = exec("git", "-C", repoRoot, "ls-tree", "-r", "-l", sha);
|
||||
long total = 0;
|
||||
for (String line : out.split("\\R")) {
|
||||
if (line.isBlank()) {
|
||||
continue;
|
||||
}
|
||||
// ls-tree -l row: "<mode> <type> <object> <size>\t<path>"; the size is only numeric for
|
||||
// blobs (trees read "-"), so gate on the type token and take the 4th whitespace field.
|
||||
String[] parts = line.split("\\s+");
|
||||
if (parts.length >= 4 && "blob".equals(parts[1])) {
|
||||
try {
|
||||
total += Long.parseLong(parts[3]);
|
||||
} catch (NumberFormatException ignored) {
|
||||
// a '-' size (or any anomaly) contributes nothing to the rough figure
|
||||
}
|
||||
}
|
||||
}
|
||||
return total;
|
||||
}
|
||||
|
||||
/** Resolve the directory that will hold per-session worktree checkouts. */
|
||||
private Path resolveRoot(String repoRoot) {
|
||||
if (configuredRoot != null && !configuredRoot.isBlank()) {
|
||||
|
||||
@@ -53,6 +53,14 @@ public final class SessionManager implements TurnListener {
|
||||
private final int contextCap;
|
||||
private final boolean clearAfterTurn;
|
||||
private volatile MemberLifecycle memberLifecycle = MemberLifecycle.NONE;
|
||||
/**
|
||||
* CB-586: the repo root the fleet actually works in, remembered the first time a worktree
|
||||
* session is spawned (worktrees are checkouts of it). {@code refs/wip/*} live there, and this
|
||||
* single cached value is what the snapshot retention sweep and the operator-visible census run
|
||||
* against. The daemon is bridged into one project at a time, so "the first worktree's repo" is
|
||||
* the repo; {@code null} until any worktree is spawned, meaning nothing to sweep or measure.
|
||||
*/
|
||||
private volatile String fleetRepoRoot;
|
||||
|
||||
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
|
||||
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||
@@ -450,6 +458,11 @@ public final class SessionManager implements TurnListener {
|
||||
// The non-worktree path always used this chain; only this branch was missed.
|
||||
String repoRoot = worktrees.repoRoot(
|
||||
launcher.effectiveCwd(new SpawnRequest(preResolvedProfile, requestedCwd, callerCwd)));
|
||||
if (fleetRepoRoot == null) {
|
||||
// CB-586: remember the repo whose worktrees the fleet spawns — its refs/wip/* are the
|
||||
// snapshot store the retention sweep and the operator census operate on.
|
||||
fleetRepoRoot = repoRoot;
|
||||
}
|
||||
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
|
||||
String path = null;
|
||||
PeerHandle handle;
|
||||
@@ -740,6 +753,26 @@ public final class SessionManager implements TurnListener {
|
||||
return registry.size();
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-586: the operator-visible census of {@code refs/wip/*} in the repo the fleet works in —
|
||||
* how many snapshot refs exist and roughly what they cost. Empty (no repo known) until at
|
||||
* least one worktree session has been spawned, exactly so a fleet that has never snapshotted
|
||||
* anything surfaces nothing new, as it did before CB-586.
|
||||
*/
|
||||
public Optional<Worktrees.WipRefStats> wipRefs() {
|
||||
String repo = fleetRepoRoot;
|
||||
return repo == null ? Optional.empty() : Optional.of(worktrees.wipRefs(repo));
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-586: run the snapshot retention sweep in the fleet's repo (a no-op until a worktree has
|
||||
* been spawned, which establishes the repo). Returns how many {@code refs/wip/*} it deleted.
|
||||
*/
|
||||
public int sweepWipRefs(long minAgeMillis) {
|
||||
String repo = fleetRepoRoot;
|
||||
return repo == null ? 0 : worktrees.pruneWipRefs(repo, minAgeMillis);
|
||||
}
|
||||
|
||||
/**
|
||||
* The registered session owning {@code terminalId}, or {@code null} if none does.
|
||||
*
|
||||
|
||||
@@ -14,12 +14,20 @@ public final class SessionReaper {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
|
||||
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
|
||||
/** CB-586: the refs/wip age floor — never sweep a snapshot younger than 24h (the CB-586 rule). */
|
||||
private static final long WIP_MIN_AGE_MILLIS = TimeUnit.HOURS.toMillis(24);
|
||||
/**
|
||||
* CB-586: how often the retention sweep runs. Given the 24h age floor, running it every few
|
||||
* hours means a ref is dropped within hours of becoming eligible, never within minutes.
|
||||
*/
|
||||
private static final long WIP_SWEEP_INTERVAL_NANOS = TimeUnit.HOURS.toNanos(6);
|
||||
|
||||
private final SessionManager sessions;
|
||||
private final long idleTtlNanos;
|
||||
private final long intervalMillis;
|
||||
private volatile boolean running;
|
||||
private Thread thread;
|
||||
private volatile long lastWipSweepNanos = Long.MIN_VALUE;
|
||||
|
||||
/** Construct a reaper with the default 5-second polling interval. */
|
||||
public SessionReaper(SessionManager sessions, long idleTtlSeconds) {
|
||||
@@ -49,10 +57,32 @@ public final class SessionReaper {
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("session reaper iteration failed; continuing", e);
|
||||
}
|
||||
maybeSweepWipRefs();
|
||||
sleep();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-586: run the refs/wip retention sweep on a slow cadence (hours, not the per-iteration
|
||||
* millisecond loop). Best-effort — a failure must never take the idle-reap loop down with it.
|
||||
*/
|
||||
private void maybeSweepWipRefs() {
|
||||
long now = System.nanoTime();
|
||||
if (now - lastWipSweepNanos < WIP_SWEEP_INTERVAL_NANOS) {
|
||||
return;
|
||||
}
|
||||
try {
|
||||
int deleted = sessions.sweepWipRefs(WIP_MIN_AGE_MILLIS);
|
||||
if (deleted > 0) {
|
||||
log.info("refs/wip retention sweep deleted {} snapshot ref(s) older than 24h whose "
|
||||
+ "content was already reachable from main", deleted);
|
||||
}
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("refs/wip retention sweep failed; continuing", e);
|
||||
}
|
||||
lastWipSweepNanos = now;
|
||||
}
|
||||
|
||||
private void sleep() {
|
||||
try {
|
||||
Thread.sleep(intervalMillis);
|
||||
|
||||
@@ -54,4 +54,48 @@ public interface Worktrees {
|
||||
* tolerance — a worktree that is gone holds nothing to snapshot)
|
||||
*/
|
||||
Optional<String> snapshot(String worktreePath, String branch, String message);
|
||||
|
||||
/**
|
||||
* CB-586: how many {@code refs/wip/*} snapshot refs exist in {@code repoRoot} and roughly what
|
||||
* they cost. This is the operator-visible surface for the snapshot growth CB-578 stage C left
|
||||
* behind — counts of refs alone hide that each one pins a whole tree for {@code git gc}.
|
||||
*
|
||||
* @param repoRoot the repository to scan
|
||||
* @return count of snapshot refs, and {@code costBytes} = the approximate total working-tree
|
||||
* size of every snapshot's committed content (summed per ref, so shared objects are
|
||||
* counted once per ref that carries them)
|
||||
*/
|
||||
WipRefStats wipRefs(String repoRoot);
|
||||
|
||||
/**
|
||||
* CB-586: run the {@code refs/wip/*} retention sweep and return how many refs it deleted.
|
||||
*
|
||||
* <p>The retention rule is <em>reachability plus an age floor</em>. A snapshot ref is deleted
|
||||
* only when <strong>both</strong> hold:
|
||||
* <ol>
|
||||
* <li>its commit's <em>tree content</em> is already reachable from {@code main} — the work
|
||||
* the snapshot preserved has been recovered, so dropping the ref loses nothing; and</li>
|
||||
* <li>the ref is older than {@code minAgeMillis} — a very recent snapshot is never swept
|
||||
* while a lead may still be looking at it.</li>
|
||||
* </ol>
|
||||
*
|
||||
* <p>Reachability is the safety property. A snapshot exists precisely because the work was not
|
||||
* committed anywhere else, so a snapshot whose content is <em>not</em> reachable from
|
||||
* {@code main} is the <strong>last copy</strong> of a worker's work and must never be deleted
|
||||
* automatically — that is the failure CB-576 and CB-578 stage C were built to stop. Age alone
|
||||
* must never drive a deletion, because age-based sweeping is exactly how the last copy gets
|
||||
* destroyed. (Both numbers and the rule are CB-586's decision; this method only implements it.)
|
||||
*
|
||||
* <p>Every deletion logs the ref name and the commit sha, so an operator who finds they lost
|
||||
* the wrong thing can still recover it from git's reflog.
|
||||
*
|
||||
* @param repoRoot the repository whose {@code refs/wip/*} to sweep
|
||||
* @param minAgeMillis the age floor; a ref younger than this is never touched
|
||||
* @return the number of snapshot refs deleted
|
||||
*/
|
||||
int pruneWipRefs(String repoRoot, long minAgeMillis);
|
||||
|
||||
/** CB-586: the operator-visible census of {@code refs/wip/*} in one repository. */
|
||||
record WipRefStats(int count, long costBytes) {
|
||||
}
|
||||
}
|
||||
|
||||
@@ -685,6 +685,12 @@ class ClaudeCodeLauncherTest {
|
||||
* explicit (non-blank) GITEA_ACCESS_TOKEN to herdr on every spawn, whatever the profile is, so
|
||||
* a future baseEnv refactor cannot silently drop it and reopen the leak. Asserted against what
|
||||
* tab.create's params actually carry, not an internal map built in the test (gitea #77).
|
||||
*
|
||||
* <p>Scope, measured on a live pane 2026-08-15: this pins what the launcher SENDS, and that is
|
||||
* all it can pin. It does not prove the value survives, and it does not: the pane runs a login
|
||||
* shell, ~/.zprofile sources secrets.sh, and its unconditional `export GITEA_ACCESS_TOKEN=...`
|
||||
* puts the real token back over this sentinel. Closing that needs the operator to guard the
|
||||
* export on BRIDGED_MEMBER — see everySpawnMarksThePaneAsAMember below.
|
||||
*/
|
||||
@Test
|
||||
void everySpawnShadowsTheAdminGiteaAccessToken() {
|
||||
@@ -716,6 +722,39 @@ class ClaudeCodeLauncherTest {
|
||||
"a profile's own env: must not be able to smuggle the admin token back in");
|
||||
}
|
||||
|
||||
/**
|
||||
* The half of CB-592 that can actually survive the pane's login shell. BRIDGED_MEMBER is a name
|
||||
* secrets.sh never exports, so nothing overwrites it — measured: GITEA_TOKEN is injected the
|
||||
* same way, is absent from a login shell of its own, and was observed set inside a live member
|
||||
* pane. It lets the operator guard the admin export with
|
||||
* `[ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...`, which is the whole fix.
|
||||
* Pinned here so a refactor cannot drop the marker and quietly un-guard every member (#77).
|
||||
*/
|
||||
@Test
|
||||
void everySpawnMarksThePaneAsAMember() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, List.of("claude"), null).spawn();
|
||||
|
||||
assertEquals("1", startEnv(herdr).get("BRIDGED_MEMBER"),
|
||||
"every member pane must be marked, or a shell file cannot tell it apart from the operator's");
|
||||
}
|
||||
|
||||
/** A profile must not be able to hide that its pane is a member, for the same reason as above. */
|
||||
@Test
|
||||
void aProfileEnvEntryCannotClearTheMemberMarker() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
|
||||
null, Map.of("BRIDGED_MEMBER", ""), null, null);
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||
_ -> null).spawn();
|
||||
|
||||
assertEquals("1", startEnv(herdr).get("BRIDGED_MEMBER"),
|
||||
"a profile's own env: must not be able to unmark its pane");
|
||||
}
|
||||
|
||||
// ── CB-533: the model is pinned on the command line, not only in the environment ────────────
|
||||
|
||||
/** A launcher for a profile identical but for its {@code model:} — the only variable here. */
|
||||
|
||||
@@ -27,11 +27,15 @@ public final class FakeWorktrees implements Worktrees {
|
||||
public record SnapshotCall(String worktreePath, String branch, String message) {
|
||||
}
|
||||
|
||||
public record PruneCall(String repoRoot, long minAgeMillis) {
|
||||
}
|
||||
|
||||
private final List<AddCall> addCalls = new CopyOnWriteArrayList<>();
|
||||
private final List<RemoveCall> removeCalls = new CopyOnWriteArrayList<>();
|
||||
private final List<OverlayCall> overlayCalls = new CopyOnWriteArrayList<>();
|
||||
private final List<RepoRootCall> repoRootCalls = new CopyOnWriteArrayList<>();
|
||||
private final List<SnapshotCall> snapshotCalls = new CopyOnWriteArrayList<>();
|
||||
private final List<PruneCall> pruneCalls = new CopyOnWriteArrayList<>();
|
||||
private final Set<String> existingPaths = ConcurrentHashMap.newKeySet();
|
||||
private final Set<String> trackedPaths = ConcurrentHashMap.newKeySet();
|
||||
private final AtomicLong snapshotSeq = new AtomicLong();
|
||||
@@ -40,6 +44,8 @@ public final class FakeWorktrees implements Worktrees {
|
||||
private volatile boolean dirty = false;
|
||||
private volatile String repoRoot = "/repo";
|
||||
private volatile String prefix = "/worktrees";
|
||||
private volatile WipRefStats wipRefs = new WipRefStats(0, 0L);
|
||||
private volatile int pruneResult = 0;
|
||||
|
||||
public FakeWorktrees withRepoRoot(String root) {
|
||||
this.repoRoot = root;
|
||||
@@ -82,6 +88,18 @@ public final class FakeWorktrees implements Worktrees {
|
||||
return this;
|
||||
}
|
||||
|
||||
/** Configure the value returned by {@link #wipRefs}. */
|
||||
public FakeWorktrees withWipRefs(WipRefStats stats) {
|
||||
this.wipRefs = stats;
|
||||
return this;
|
||||
}
|
||||
|
||||
/** Configure the value returned by {@link #pruneWipRefs}. */
|
||||
public FakeWorktrees withPruneResult(int deleted) {
|
||||
this.pruneResult = deleted;
|
||||
return this;
|
||||
}
|
||||
|
||||
@Override
|
||||
public String add(String repoRoot, String branch, String baseRef) {
|
||||
addCalls.add(new AddCall(repoRoot, branch, baseRef));
|
||||
@@ -135,6 +153,17 @@ public final class FakeWorktrees implements Worktrees {
|
||||
return Optional.of("wip" + snapshotSeq.incrementAndGet());
|
||||
}
|
||||
|
||||
@Override
|
||||
public WipRefStats wipRefs(String repoRoot) {
|
||||
return wipRefs;
|
||||
}
|
||||
|
||||
@Override
|
||||
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
|
||||
pruneCalls.add(new PruneCall(repoRoot, minAgeMillis));
|
||||
return pruneResult;
|
||||
}
|
||||
|
||||
public List<AddCall> addCalls() {
|
||||
return List.copyOf(addCalls);
|
||||
}
|
||||
@@ -167,6 +196,10 @@ public final class FakeWorktrees implements Worktrees {
|
||||
return List.copyOf(snapshotCalls);
|
||||
}
|
||||
|
||||
public List<PruneCall> pruneCalls() {
|
||||
return List.copyOf(pruneCalls);
|
||||
}
|
||||
|
||||
public SnapshotCall lastSnapshot() {
|
||||
return snapshotCalls.isEmpty() ? null : snapshotCalls.getLast();
|
||||
}
|
||||
|
||||
@@ -3,6 +3,7 @@ package dev.ltms.bridged.session;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.HashSet;
|
||||
@@ -138,6 +139,53 @@ class GitWorktreesTest {
|
||||
return out;
|
||||
}
|
||||
|
||||
/** Write {@code content} as a blob into the object database; returns its sha. */
|
||||
private static String blobOf(Path cwd, String content) throws Exception {
|
||||
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "hash-object", "-w", "--stdin")
|
||||
.redirectErrorStream(true).start();
|
||||
p.getOutputStream().write(content.getBytes(StandardCharsets.UTF_8));
|
||||
p.getOutputStream().close();
|
||||
String out = new String(p.getInputStream().readAllBytes()).trim();
|
||||
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git hash-object timed out");
|
||||
assertEquals(0, p.exitValue(), "git hash-object failed:\n" + out);
|
||||
return out;
|
||||
}
|
||||
|
||||
/** Build a single-file tree object from {@code blob}; returns the tree's sha. */
|
||||
private static String treeOf(Path cwd, String path, String blob) throws Exception {
|
||||
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "mktree")
|
||||
.redirectErrorStream(true).start();
|
||||
p.getOutputStream().write(("100644 blob " + blob + "\t" + path + "\n").getBytes(StandardCharsets.UTF_8));
|
||||
p.getOutputStream().close();
|
||||
String out = new String(p.getInputStream().readAllBytes()).trim();
|
||||
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git mktree timed out");
|
||||
assertEquals(0, p.exitValue(), "git mktree failed:\n" + out);
|
||||
return out;
|
||||
}
|
||||
|
||||
/** {@code git commit-tree} rooted at {@code tree} with a chosen committer date; returns the sha. */
|
||||
private static String commitTree(Path cwd, String tree, String parent, String committerDate,
|
||||
String message) throws Exception {
|
||||
ProcessBuilder pb = new ProcessBuilder("git", "-C", cwd.toString(), "commit-tree",
|
||||
tree, "-p", parent, "-m", message);
|
||||
pb.environment().put("GIT_COMMITTER_DATE", committerDate);
|
||||
Process p = pb.redirectErrorStream(true).start();
|
||||
String out = new String(p.getInputStream().readAllBytes()).trim();
|
||||
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git commit-tree timed out");
|
||||
assertEquals(0, p.exitValue(), "git commit-tree failed:\n" + out);
|
||||
return out;
|
||||
}
|
||||
|
||||
/** {@code git update-ref <ref> <sha>} — create the snapshot ref directly. */
|
||||
private static void updateRef(Path cwd, String ref, String sha) throws Exception {
|
||||
git(cwd, "update-ref", ref, sha);
|
||||
}
|
||||
|
||||
/** True when {@code ref} exists in the repo (for-each-ref on a missing ref is empty, not an error). */
|
||||
private static boolean refExists(Path cwd, String ref) throws Exception {
|
||||
return !forEachRef(cwd, ref).trim().isEmpty();
|
||||
}
|
||||
|
||||
/**
|
||||
* The heart of CB-525: a provisioned worktree must not inherit the primary's MCP servers. Without
|
||||
* the isolation step the checked-out {@code .mcp.json} carries them in, and a worker navigating
|
||||
@@ -474,4 +522,104 @@ class GitWorktreesTest {
|
||||
assertThrows(WorktreeException.class, () -> gitWorktrees.snapshot(wt, branch, "test snapshot"),
|
||||
"an unresolvable real index must fail loudly, not silently snapshot from an empty index");
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-586, criterion 2. A snapshot whose content is NOT reachable from {@code main} is the last
|
||||
* copy of a worker's work, and must never be deleted automatically — even when it is old and
|
||||
* even when the caller passes a zero age floor. Uses the real snapshot path on a dirty worktree,
|
||||
* so the unreachable tree is exactly the shape CB-576/CB-578 stage C exist to protect.
|
||||
*/
|
||||
@Test
|
||||
void anUnreachableSnapshotIsNeverPruned(@TempDir Path tmp) throws Exception {
|
||||
Path repo = initRepo(tmp.resolve("repo"));
|
||||
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||
String branch = "cb-586-unreachable";
|
||||
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
|
||||
Files.writeString(Path.of(wt).resolve("worker-draft.txt"), "work that exists nowhere else\n");
|
||||
|
||||
Optional<String> ref = gitWorktrees.snapshot(wt, branch, "snapshot with unreachable content");
|
||||
assertTrue(ref.isPresent());
|
||||
|
||||
// Age floor 0 makes age a non-issue: only reachability can save it — and it must.
|
||||
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), 0),
|
||||
"the unreachable snapshot is the last copy and must not be pruned");
|
||||
assertTrue(refExists(repo, "refs/wip/" + branch),
|
||||
"an unreachable snapshot must survive the sweep");
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-586, criterion 1 (the reachable half). A snapshot whose tree content IS already reachable
|
||||
* from {@code main} and which is older than the age floor is pure duplication — the work is
|
||||
* recovered — so it must be pruned.
|
||||
*/
|
||||
@Test
|
||||
void aReachableSnapshotOlderThanTheFloorIsPruned(@TempDir Path tmp) throws Exception {
|
||||
Path repo = initRepo(tmp.resolve("repo"));
|
||||
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||
|
||||
// A snapshot whose tree is exactly main's current tree: fully reachable from main.
|
||||
String mainTree = revParse(repo, "main^{tree}");
|
||||
String old = commitTree(repo, mainTree, revParse(repo, "HEAD"), "2020-01-01T00:00:00", "snapshot");
|
||||
updateRef(repo, "refs/wip/recovered", old);
|
||||
|
||||
assertEquals(1, gitWorktrees.pruneWipRefs(repo.toString(), TimeUnit.HOURS.toMillis(24)),
|
||||
"an old, main-reachable snapshot must be pruned");
|
||||
assertFalse(refExists(repo, "refs/wip/recovered"),
|
||||
"the reachable snapshot's ref must be gone after the sweep");
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-586, the age floor. A snapshot whose content IS reachable from {@code main} but which is
|
||||
* younger than the age floor must not be swept — a lead may still be looking at it.
|
||||
*/
|
||||
@Test
|
||||
void aReachableButRecentSnapshotIsNotPruned(@TempDir Path tmp) throws Exception {
|
||||
Path repo = initRepo(tmp.resolve("repo"));
|
||||
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||
|
||||
// Reachable from main, but committed "now" — a fresh snapshot. The 24h floor must protect it.
|
||||
String mainTree = revParse(repo, "main^{tree}");
|
||||
String fresh = commitTree(repo, mainTree, revParse(repo, "HEAD"),
|
||||
"2038-01-01T00:00:00", "snapshot just taken");
|
||||
updateRef(repo, "refs/wip/fresh", fresh);
|
||||
|
||||
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), TimeUnit.HOURS.toMillis(24)),
|
||||
"a recent snapshot must be kept even when reachable");
|
||||
assertTrue(refExists(repo, "refs/wip/fresh"),
|
||||
"the recent reachable snapshot must survive the sweep");
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-586, criterion 5. A fleet that has never snapshotted anything has no {@code refs/wip/*},
|
||||
* so a sweep is a no-op and the census reports none — identical to before CB-586 existed.
|
||||
*/
|
||||
@Test
|
||||
void aFleetWithNoSnapshotsPrunesNothingAndReportsNothing(@TempDir Path tmp) throws Exception {
|
||||
Path repo = initRepo(tmp.resolve("repo"));
|
||||
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||
|
||||
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), 0),
|
||||
"no snapshot refs means nothing to prune");
|
||||
Worktrees.WipRefStats stats = gitWorktrees.wipRefs(repo.toString());
|
||||
assertEquals(0, stats.count(), "a never-snapshotted fleet has zero refs/wip refs");
|
||||
assertEquals(0L, stats.costBytes(), "a never-snapshotted fleet costs zero bytes");
|
||||
}
|
||||
|
||||
/** CB-586, criterion 4: the census reports how many refs exist and roughly what they cost. */
|
||||
@Test
|
||||
void wipRefsReportsCountAndCost(@TempDir Path tmp) throws Exception {
|
||||
Path repo = initRepo(tmp.resolve("repo"));
|
||||
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
|
||||
|
||||
String blob = blobOf(repo, "a recoverable snapshot's worth of content");
|
||||
String tree = treeOf(repo, "snapshot.txt", blob);
|
||||
updateRef(repo, "refs/wip/one", commitTree(repo, tree, revParse(repo, "HEAD"),
|
||||
"2020-01-01T00:00:00", "snapshot"));
|
||||
updateRef(repo, "refs/wip/two", commitTree(repo, tree, revParse(repo, "HEAD"),
|
||||
"2020-01-02T00:00:00", "snapshot"));
|
||||
|
||||
Worktrees.WipRefStats stats = gitWorktrees.wipRefs(repo.toString());
|
||||
assertEquals(2, stats.count(), "two snapshot refs are reported");
|
||||
assertTrue(stats.costBytes() > 0, "the cost of the snapshots is a positive byte count");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -138,6 +138,16 @@ class SessionManagerTest {
|
||||
return java.util.Optional.of("wip" + snapshotSeq.incrementAndGet());
|
||||
}
|
||||
|
||||
@Override
|
||||
public WipRefStats wipRefs(String repoRoot) {
|
||||
return new WipRefStats(0, 0L);
|
||||
}
|
||||
|
||||
@Override
|
||||
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
|
||||
return 0;
|
||||
}
|
||||
|
||||
List<String> removeCalls() {
|
||||
return List.copyOf(removeCalls);
|
||||
}
|
||||
|
||||
@@ -444,4 +444,37 @@ class WorktreeSessionManagerTest {
|
||||
"a failed snapshot leaves no ref to report");
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-586, criteria 4 and 1. Once a worktree session is spawned the repo is known, so the
|
||||
* operator census and the retention sweep delegate to that repo's {@code refs/wip/*}.
|
||||
*/
|
||||
@Test
|
||||
void wipRefsAndSweepDelegateToTheFleetRepoOnceKnown() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
|
||||
.withWipRefs(new Worktrees.WipRefStats(3, 42L)).withPruneResult(2);
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
|
||||
assertTrue(sessions.wipRefs().isEmpty(),
|
||||
"no worktree spawned yet means no repo is known and nothing to report");
|
||||
assertEquals(0, sessions.sweepWipRefs(TimeUnit.HOURS.toMillis(24)),
|
||||
"no worktree spawned yet means the sweep is a no-op");
|
||||
|
||||
sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("cb-586", null));
|
||||
|
||||
Worktrees.WipRefStats stats = sessions.wipRefs().orElseThrow();
|
||||
assertEquals(3, stats.count(), "the census comes from the fleet repo");
|
||||
assertEquals(42L, stats.costBytes(), "the cost comes from the fleet repo");
|
||||
assertEquals(2, sessions.sweepWipRefs(TimeUnit.HOURS.toMillis(24)),
|
||||
"the sweep runs against the fleet repo");
|
||||
|
||||
List<FakeWorktrees.PruneCall> prunes = worktrees.pruneCalls();
|
||||
assertEquals(1, prunes.size(), "the no-op short-circuits before reaching the seam, so only "
|
||||
+ "the repo-known sweep issues a call");
|
||||
assertEquals("/repo", prunes.getFirst().repoRoot(), "the sweep targets the fleet repo");
|
||||
assertEquals(TimeUnit.HOURS.toMillis(24), prunes.getFirst().minAgeMillis(),
|
||||
"the caller's age floor is passed through");
|
||||
}
|
||||
|
||||
}
|
||||
|
||||
@@ -1,6 +1,11 @@
|
||||
# CB-591 — move the fleet onto the LLM and MCP gateway
|
||||
|
||||
**Status:** plan, not started · **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway)
|
||||
**Status: DONE — the fleet is on the gateway as of 2026-08-15.** `local` runs on `/anthropic` and
|
||||
`gx` on `/v1`, both at `weight: 100`; `local-direct` stays at `weight: 0` as the escape hatch. Getting
|
||||
here took a revert and two upstream fixes — see §7.1, which is the useful part of this document. One
|
||||
risk is **accepted rather than solved**: a stream cut by any mid-response timer arrives as HTTP 200
|
||||
with no terminator, and our third-party members cannot detect it (§7.2).
|
||||
· **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway)
|
||||
· **Upstream issue:** [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31)
|
||||
|
||||
The gateway went live on 2026-08-15 and replaced Bifrost. This plan says what that means for a
|
||||
@@ -154,12 +159,21 @@ speaks the Anthropic protocol, so `/anthropic` is both correct and the only safe
|
||||
This is the exact failure shape this repo keeps hitting: it compiles, it answers, it looks healthy,
|
||||
and a capability is quietly off. Treat it as a `silent-default` risk, not a config preference.
|
||||
|
||||
**Open question for 3b, stated as open.** The opencode profile must use the OpenAI surface, because
|
||||
that is the only thing an OpenAI-compatible provider can speak. vLLM serves that API natively, so I
|
||||
*expect* it to be a passthrough as well — but the wiki documents the translation trap only for the
|
||||
Anthropic path, and I have not checked how `reasoning_content` behaves through `/v1`. Verify it on
|
||||
first spawn (§7) rather than assuming. If reasoning is dropped there, that is a limitation of the
|
||||
opencode profile, not a reason to abandon it — opencode members do dev work, not deep reasoning.
|
||||
**Open question for 3b — ANSWERED, 2026-08-15.** The worry was that the OpenAI surface might drop
|
||||
reasoning the way the wiki documents for a mis-declared Anthropic backend. It does not. Checked at
|
||||
the API before any profile was switched:
|
||||
|
||||
| surface | request | result |
|
||||
|---|---|---|
|
||||
| `/anthropic/v1/messages` | `deepseek-v4-flash`, 64 tokens | 200, response carries a real `"type":"thinking"` block |
|
||||
| `/v1/chat/completions` | same | 200, message carries a populated `reasoning_content` (and a `reasoning` field) |
|
||||
| `/v1/models` | — | 200, exactly `["deepseek-v4-flash"]` — the exact-name trap is clear |
|
||||
| `/v1/models`, **no token** | — | **401** — Caddy is gating, as designed |
|
||||
|
||||
So reasoning survives on **both** surfaces, and the `/anthropic` choice for `local` is about protocol
|
||||
correctness rather than a repair for a known loss. The last row matters on its own: the wiki warns
|
||||
the gateway's own `SecurityPolicy` fails open, so it is worth knowing the proxy in front really does
|
||||
refuse an unauthenticated request here.
|
||||
|
||||
---
|
||||
|
||||
@@ -290,6 +304,160 @@ Merging config is not proving it. The checks, in order:
|
||||
|
||||
---
|
||||
|
||||
## 7.1 What the live run actually found — 2026-08-15
|
||||
|
||||
U1–U2c were done, the daemon restarted onto them, and both new profiles were spawned for real. The
|
||||
migration was then **reverted**. This section is the result, so none of it has to be re-derived.
|
||||
|
||||
### The blocker
|
||||
|
||||
`llm.ltms.dev` answers **HTTP 413 Request Entity Too Large** above **32 KiB (32768 bytes)**, on both
|
||||
surfaces:
|
||||
|
||||
```
|
||||
/v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200
|
||||
/v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413
|
||||
```
|
||||
|
||||
32 KiB is far below one real agent turn.
|
||||
|
||||
**Root cause — confirmed by the systems/vms side, 2026-08-15.** My guess that it was a Caddy
|
||||
`request_body max_size` was **wrong**. It is Envoy, inside `aigw` on `llm.vm`. Envoy Gateway defaults
|
||||
a listener's `per_connection_buffer_limit_bytes` to **32768**, and the AI Gateway buffers the *whole*
|
||||
request body before it can route on the model name — so that default is not a network tuning knob
|
||||
here, it is a hard ceiling on prompt size. Read out of the live Envoy `config_dump`:
|
||||
|
||||
```
|
||||
listener default/llm/http per_connection_buffer_limit_bytes: 32768
|
||||
```
|
||||
|
||||
Nobody chose 32 KiB; it was inherited from the default. Both TLS edges are innocent: the same
|
||||
boundary reproduces on the LAN path and the internet path, and both 413s carry an `x-llm-consumer`
|
||||
header their auth proxy sets only *after* authenticating — so the body cleared both edges and the
|
||||
auth. Directly on `llm.vm`, `aigw` 413s at 39 KB while the vLLM backend accepts the same 39 KB and
|
||||
answers 200.
|
||||
|
||||
**Do not plan around 32 KiB.** The intended ceiling is far higher. Their fix — a `ClientTrafficPolicy`
|
||||
setting `bufferLimit: 8Mi` — is written but **not deployed** as of this note, pending their operator's
|
||||
approval. I have not re-tested and will not until they confirm, so as not to measure a half-changed
|
||||
system. Fixed in **systems/vms**, not here.
|
||||
|
||||
### The part worth remembering
|
||||
|
||||
Two members were spawned at the same moment with the same message:
|
||||
|
||||
| | `local` (claude-code, `/anthropic`) | `gx` (opencode, `/v1`) |
|
||||
|---|---|---|
|
||||
| READY → BUSY | 19:07:26 | 19:07:45 |
|
||||
| BUSY → DONE | **19:08:51 (66s)** | **never — 10+ min, ticket FAILED** |
|
||||
|
||||
**`local` passed.** It passed only because the probe was three trivial questions in a fresh session,
|
||||
so the request fit under 32 KiB. The profile looked healthy and was a landmine set to fire on the
|
||||
first turn that reads a file.
|
||||
|
||||
So §7's checklist was not wrong, it was **too easy**. Any future run of it must use a task that reads
|
||||
a real file. A liveness probe proves the token and the URL; it does not prove the path.
|
||||
|
||||
`gx` did not fail loudly either. Reproduced outside the bridge by running `opencode` by hand with the
|
||||
launcher's own generated config:
|
||||
|
||||
```
|
||||
Error: Request Entity Too Large
|
||||
...compacts context, retries...
|
||||
Error: Request Entity Too Large
|
||||
```
|
||||
|
||||
opencode **catches the 413, compacts, and retries — indefinitely**. A member that fails loudly costs
|
||||
one turn; this one costs the whole task and is indistinguishable from a slow worker.
|
||||
|
||||
> **Diagnosing a stuck opencode member.** Do not read its pane. The launcher writes its config to a
|
||||
> temp dir and passes it as `OPENCODE_CONFIG` — find it with
|
||||
> `ls -dt /var/folders/*/*/T/bridged-opencode-* | head -1`, check the provider block and the key's
|
||||
> length and prefix (never its value), then reproduce with `opencode run --auto -m <provider>/<model>`
|
||||
> using the same `OPENCODE_CONFIG`. That is what turned "it hangs" into a one-line error.
|
||||
|
||||
### What checked out, and needs no re-testing
|
||||
|
||||
- Token accepted on both surfaces. **Unauthenticated → 401**, so the Caddy proxy really does gate —
|
||||
the wiki's "SecurityPolicy fails open" warning is about the gateway itself, not the edge.
|
||||
- `/v1/models` returns exactly `["deepseek-v4-flash"]`, so trap 3 is clear.
|
||||
- **Reasoning survives both surfaces** — see §3b above.
|
||||
- The launcher's generated opencode provider block is correct, carrying a real 48-character `llmk-`
|
||||
key rather than the `bridged-local-noauth` placeholder.
|
||||
- `SubscriptionGuard` accepted `llm.ltms.dev` after the allowlist edit and the restart: `local`
|
||||
spawned without throwing, which is the check that catches a missed restart.
|
||||
|
||||
### Resolution — both ceilings fixed, migration completed
|
||||
|
||||
systems/vms fixed both, and each was re-checked from this side rather than taken on trust:
|
||||
|
||||
| ceiling | was | now | our own check |
|
||||
|---|---|---|---|
|
||||
| listener buffer | 32 KiB | 32 Mi | 1.2 MB body → **200** (was 413) |
|
||||
| LLM route timeout | 60s | 86400s | the request that truncated: **101s, `message_stop` present, 4000/4000** |
|
||||
|
||||
The timeout moved in two steps on 2026-08-15: 60s → 1800s, then 1800s → **86400s (24 hours)** after
|
||||
the truncation risk below was discussed. They tried `request: 0s` first, which removes the
|
||||
total-duration timer completely. It works, but on an `AIGatewayRoute` the **idle timeout is derived
|
||||
from the request timeout**, so `0s` also removed any bound on a stalled connection. 86400s keeps a
|
||||
reaper for dead connections while putting the truncation timer out of practical reach.
|
||||
|
||||
Neither was deliberate. The 32 KiB was Envoy Gateway's default `per_connection_buffer_limit_bytes`;
|
||||
the 60s was Envoy AI Gateway's own documented default. The 60s bounded **generation** as well as
|
||||
prompt size — a tiny prompt with a long answer returned 504 at 60.05s.
|
||||
|
||||
Two configuration facts worth keeping, from their bisection:
|
||||
|
||||
- **`ClientTrafficPolicy` is honoured in standalone `aigw run`; `BackendTrafficPolicy` is NOT.** A
|
||||
`BackendTrafficPolicy` setting `requestTimeout` is accepted, logs nothing, and leaves the routes
|
||||
unchanged (upstream `envoyproxy/gateway#9513`). What works is `timeouts: {request: …}` on each
|
||||
`AIGatewayRoute` rule. Nothing from the outside distinguishes the two — the same silent-default
|
||||
shape as their `SecurityPolicy` caveat.
|
||||
- In that stack, "the config was accepted" proves nothing. Read the live `config_dump`.
|
||||
|
||||
## 7.2 The risk we accepted, and why we could not remove it
|
||||
|
||||
Raising the timeout made the failure **rare, not impossible**, and the residual failure is silent.
|
||||
|
||||
On a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked encoding *cleanly* instead of
|
||||
resetting the connection, so the client receives what looks like a complete transfer
|
||||
(`envoyproxy/envoy#17186` — acknowledged as a bug in 2021, closed by a stale bot, never fixed). The
|
||||
December 2025 fix `envoyproxy/envoy#42269` changes locally-originated resets from `NO_ERROR` to
|
||||
`INTERNAL_ERROR`, but it is **HTTP/2 only** and SSE clients here speak HTTP/1.1.
|
||||
|
||||
Measured on our side while the timeout was still 60s:
|
||||
|
||||
```
|
||||
HTTP 200 61.07s 141992 bytes
|
||||
message_stop 0 message_delta 0 error events 0
|
||||
emitted 2473 of 4000, ending on a WELL-FORMED SSE frame
|
||||
```
|
||||
|
||||
A syntactically valid stream that simply stops. Any timer firing mid-stream — route timeout, idle
|
||||
timeout, `max_stream_duration` — fails this same way.
|
||||
|
||||
**The recommended defence does not transfer to us.** The right fix is to treat a stream with no
|
||||
`message_stop` / `[DONE]` / `finish_reason` as failed. We cannot: our members are Claude Code and
|
||||
opencode, third-party clients whose SSE parsing we do not own, and there is no seam to insert the
|
||||
check. Whether either detects a missing terminator is unverified — and opencode's handling of the 413
|
||||
(swallow, compact, retry forever, never surface an error) does not suggest it is strict.
|
||||
|
||||
So the honest statement of our position:
|
||||
|
||||
> Gateway traffic is acceptable at 86400s because a single request would have to run for 24 hours to
|
||||
> trip the bug — **not** because we could detect it if it did.
|
||||
|
||||
At 86400s our **own** limit binds first, which is the ordering we want. `MessageService.ASYNC_TIMEOUT_MS`
|
||||
caps a turn at 30 minutes, so a runaway request ends as a clean `FAILED` ticket that we raised, rather
|
||||
than as a silently truncated `200` that we cannot see. While the gateway sat at 1800s the two numbers
|
||||
were equal and did not nest, so a gateway-side stall could have been misread as a bug in our own ticket
|
||||
handling. That ambiguity is now gone.
|
||||
|
||||
**If a member ever returns a confident but truncated answer, suspect this before anything in our own
|
||||
code.** That is the whole reason this section exists.
|
||||
|
||||
---
|
||||
|
||||
## 8. Related
|
||||
|
||||
- [CB-589 / #74](https://git.ltms.dev/lms/claude-bridge/issues/74) — cost-first placement and a
|
||||
|
||||
Reference in New Issue
Block a user