The eleven tool names become fleet_* throughout the portable block. The
fallback ladder keeps mcp__bridge__* unchanged and that is deliberate: both
launchers still mount the server as "bridge", so a spawned member really
does see that prefix. Renaming the mount is a separate surface CB-622 did
not touch.
Separately, a pre-existing defect: the ladder quoted the reply charter as
"You are an off-subscription worker in the claude-bridge fleet", while
REPLY_CHARTER says "You are a spawned member in the claude-bridge fleet".
A member matching that quote found nothing, so the ladder's first rung
could never fire. Now quoted verbatim.
wiki/7-Use-Cases.md is advanced to the matching template commit; the sync
check prints in sync: True.
1. opencode.json — mount key bridged -> fleetd. Unit C could not do this:
the file is neutralized by the worktree overlay, so every worker sees a
stub and correctly reported it held no mount key. Tracked as CB-628.
2. e2e/bridge_ask_transcript.md keeps its name. It is a dated record of a
run on 2026-07-16 that really did call bridge_ask, and its first line
says so. The harness now writes fleet_ask_transcript.md for new runs;
its header already says 'Live fleet_ask', so the two now agree.
3. docs/MCP-Contract.md line 20 — the historical banner describes section
6, and section 6 now says fleet_ask.
Lead-verified: 27 .md files, 165 occurrences, no wiki/, no CLAUDE.md. The docs/MCP-Contract.md boundary held exactly — every hunk falls in 223-286, inside section 6 (210-293). The transcript rename is reverted by the lead in a follow-up: that file is a dated record of a run that really did call bridge_ask.
Lead-verified: 9 files, +68/-68, nothing under wiki/, root .mcp.json untouched. The worker's opencode.json finding was correct about its worktree and led to CB-628 (#134); the tracked file is fixed separately by the lead.
Lead-verified: both names reach the same handler instance; warn-once uses a per-name Set, not a numeric sentinel; Authz keys on an action enum so the deprecated names keep their authorization. 878 tests green in a local integration with #132, #133 and the lead's opencode.json fix. Live MCP verification is still outstanding and is the lead's step after redeploy.
fleet_* is now the documented tool name for all eleven MCP tools; each
bridge_* twin is registered against the exact same handler (no logic
duplication) and its description leads with a DEPRECATED notice. A
bridge_* call logs one WARN naming the old and new name, once per name
for the life of the process (a Set, not a numeric sentinel).
REPLY_CHARTER in HerdrPeerLauncher now tells a spawned member to call
fleet_reply — the one rule that must survive with no repo checkout.
All other bridge_* string literals across mcp/, Javadoc, and tests were
renamed to fleet_* for consistency with the new documented name.
Rename the eleven MCP tool names (bridge_ack/ask/list/poll/profiles/reply/
send/spawn/status/stop/whoami) to their fleet_* names across the Markdown
documentation. fleet_* is written as the normal name; one deprecation note
in README.md says bridge_* still works for one release.
docs/MCP-Contract.md is renamed only inside section 6 (lines 210-293):
sections 1-5 and 7-11 are stale pre-build design text (CB-609) and are
deliberately left with old names so dead text does not look maintained.
Also renames e2e/bridge_ask_transcript.md to e2e/fleet_ask_transcript.md
to match its content. CLAUDE.md, wiki/, plugin/skills/setup/SKILL.md and
.claude/skills/port-to-opencode/SKILL.md are owned by other units and are
untouched.
Two defects in CB-617 that only a live spawn could find. Both were shipped
green: every test passed because every test read the argv we built, and none
ran the binary that has to accept it.
1. Claude Code refuses to start when both --append-system-prompt and
--append-system-prompt-file are on the command line:
Error: Cannot use both --append-system-prompt and
--append-system-prompt-file. Please use only one.
CB-617 put the role charter on the file flag and left the reply charter on
the inline flag, so every Claude-profile spawn with a role charter died at
launch. The pane exited on its own and bridged reported it as
spawn_timeout ("did not reach injectable state within 20000ms"), which
hides the real cause.
Both charters now go in the one file, role charter first and reply charter
last — last is where the reply rule must sit, because it is the rule that
must survive. A member with only a reply charter keeps the proven inline
flag, which is also the only form that reaches a member with no repo
checkout.
2. A Claude Code agent definition needs `name:` in its frontmatter. Ours had
only `description:`, so the files were skipped and --agent architect failed
with "not found. Available agents: claude, Explore, ...". Added to all
three. The OpenCode files take their name from the filename and are
unchanged.
Checked on this host, in this repo, with the real binary:
claude --model claude-sonnet-5 --agent architect \
--append-system-prompt-file /tmp/combined.md -p '...'
-> ROLEOK, REPLYOK, yes
so the agent definition, the role charter and the reply charter all compose.
The updated test now asserts the constraint that actually binds: with a role
charter present, --append-system-prompt must be absent, and the file must open
with the role charter and end with the reply charter.
874 tests pass.
herdr refused to shell-encode a multi-line inline --append-system-prompt
argument (invalid_agent_argument), which broke any Claude Code profile with
a configured multi-line fleet.charters.<role>. Split delivery: the role
charter now writes to a temp file and mounts via
--append-system-prompt-file; the one-line REPLY_CHARTER keeps its inline
--append-system-prompt delivery, since it must reach a member with no repo
checkout. Also pass --agent <role> when a role's agent-definition file
exists under the worker's cwd (.claude/agents/<role>.md for claude-code,
.opencode/agent/<role>.md for opencode); absent either file, nothing extra
is added and the member still spawns.
The base's LaunchSpec now carries roleCharter/replyCharter/cwd alongside
the existing composed charter field, so CharterReceipt keeps fingerprinting
the same composed text it always did -- unchanged, since the digest covers
the full logical charter content regardless of how it is delivered.
A page audit found the dead host still named in four places outside the wiki.
ollama.ltms.dev no longer exists; the one front door is llm.ltms.dev, and the
path differs by backend - /anthropic for claude-code, /v1 for opencode.
Anyone following README.md today points a member at nothing.
README.md also now links the new wiki chapter 13, the operator guide, since
the README's own next step used to be the never-written Setup page.
Not touched: SubscriptionGuard.java:14 still names ollama.ltms.dev, but only
in a javadoc line describing the Stage-1 example allowlist. It is a comment
about history, not a live default, so it stays until that class is next
edited for its own reasons.
profiles.<name>.subscription is read in five places (BridgedConfig record +
isSubscription + validateSubscriptionProfiles, ClaudeCodeLauncher, and the
startup secret check that deliberately skips it) and was in bridged.example.yaml
nowhere. bridged.yaml is gitignored, so the example is the only place an operator
can learn a key exists - which meant a fresh host had no way to discover the one
switch that moves cost onto the operator's own subscription. Two profiles here
have it set.
Documents what changes when it is true, that it is mutually exclusive with
baseUrl and refused at load, that the startup secret check skips such profiles
so a clean secrets report says nothing about them, and that maxLoad is the only
throttle against the operator's plan - there is no metering or budget refusal.
94 config tests pass; the example still loads.
The instruction surface named a parameter the tool refuses. bridge_ack requires
target and msgId (BridgeMcp.java:1001) and errors with 'target and msgId are
required' otherwise, but step 5 of the primary procedure said bridge_ack{ticket,
msgId}. The table two sections below it was already correct, so the file
contradicted itself and the wrong half was in the numbered steps a lead follows.
The same line is fixed in the byte-identical wiki template.
docs/MCP-Contract.md still opened with 'Greenfield - no MCP code exists yet' and
CLAUDE.md pointed every session at it. Audited against BridgeMcp.java: it names
two tools that do not exist, omits four that do, gets nearly every parameter name
wrong, and uses /workers paths the daemon does not serve. Its status banner now
says so and names the live schema as the authority; CLAUDE.md's pointer is
narrowed to section 6, which is the part that did survive. Rewrite is CB-609.
Replaces CB-592's single hardcoded GITEA_ACCESS_TOKEN shadow with a
memberCredentials: block: policy + allow + known, validated at config load
in the shape CB-606 established.
What this does and does not do, because the distinction matters:
IT DOES enumerate, report and validate. No credential name is hardcoded in
Java any more. An unrecognised policy: refuses at load. A credential-shaped
env var on neither known: nor allow: is named in a WARN. An absent block
warns loudly at startup, because absence now blocks nothing at all and there
is no hardcoded fallback left.
IT DOES NOT, on its own, block anything the login shell re-exports. The env
overlay is applied at pane creation and the pane's login shell runs after it.
An exec-time fix was attempted and has no seam: herdr protocol 19's
agent.start takes a fixed kind plus trailing CLI args, with no env map and no
argv[0] control. The enforcement therefore lives in the operator's secrets.sh,
which sets BRIDGED_MEMBER-guarded sentinels for the 29 blocked names.
Verified by me, not taken on report: my own unpiped mvn clean install gives
870 tests, BUILD SUCCESS; loading the live gitignored bridged.yaml with the
new code yields 34 known / 5 allow / 29 blocked with an empty intersection;
and that blocked set diffs clean against the secrets.sh block, so the two
lists agree today.
herdr protocol 19's agent.start takes a fixed `kind` (herdr resolves the
executable) plus trailing CLI args for that binary; only tab.create/pane.split
accept an env map, and that IS the round-1 pane-creation overlay already
shipped. There is no argv/env control point that runs after the pane's login
shell and before the agent process starts, so the proposed `env NAME=value ...`
argv prefix cannot be implemented against this API. Documented the finding and
corrected bridged.example.yaml's round-1 comments, which had overclaimed that
the overlay survives the login shell.
Kept everything else: added a startup WARN (Bridged.reportMemberCredentialsGap)
when memberCredentials: is absent or its known: list is empty, so CB-592's
protection loss is never silent, mirroring CB-594's reportRequiredSecrets.
Replaces CB-592's single hardcoded GITEA_ACCESS_TOKEN shadow in HerdrPeerLauncher
with BridgedConfig.MemberCredentials (memberCredentials: policy/allow/known).
Every known name not also allowed is overlaid with a sentinel; an unrecognized
policy value refuses at load; a credential-shaped host env var on neither list
is logged as a gap (name only, never a value). bridged.yaml is gitignored, so
the 31 measured names + 4-name allowlist ship as a commented block in
bridged.example.yaml for the operator to apply live.
SessionReaper.lastWipSweepNanos started at Long.MIN_VALUE as a "never swept
yet" sentinel. That sentinel cannot be compared by subtraction. nanoTime() is
positive on this platform, so `now - Long.MIN_VALUE` wraps to a large negative
number, the gate `delta < WIP_SWEEP_INTERVAL_NANOS` reads it as "swept moments
ago", and the method returns before the assignment that would have fixed the
field. The sweep never ran once, for the life of the process, and nothing in
the log said so.
Measured:
System.nanoTime() = 31305820625625 (positive)
now - Long.MIN_VALUE = -9223340731034150183
interval (6h in nanos) = 21600000000000
gate 'delta < interval' -> true => returns early, every iteration, forever
Fix: a separate `sweptOnce` boolean holds "never yet", so the subtraction only
runs once both operands come from nanoTime. The first pass always sweeps — a
restart is a fine moment for it, the 24h age floor keeps it safe, and the
feature becomes observable right after a redeploy instead of six hours later.
The existing tests all passed because they call SessionManager.sweepWipRefs
directly, which walks around the gate. The new test asserts through the reaper
loop instead: it spawns a worktree session, starts the reaper, and requires a
real prune call at the seam with the 24h floor intact. Removing the fix makes
it fail with "it never reached the seam".
860 tests, mvn clean install, BUILD SUCCESS.
A refs/wip snapshot ref is deleted only when both hold: its commit's tree is
already reachable from main, and it is older than 24h. Reachability is the safety
floor — a snapshot exists because the work was committed nowhere else, so an
unreachable one is the last copy and is never swept. Every deletion logs the ref
and the sha.
/members gains wipRefs{count,costBytes} so the growth is visible.
Verified against real git, not only the fakes: a recoverable+old ref is deleted,
a recoverable+young one survives the age floor, and the last copy survives. A repo
with no main deletes nothing, and a repo with no snapshots is a clean no-op.
Closes#67
The prompt is part of the product: CB-582 changed what bridge_status returns, so
the primary's step 5 was no longer the whole truth.
The second half matters more than the first. A nudge makes a lead more likely to
notice an ask; it does not widen the ~55s window, which is bounded by the
WORKER's own MCP client timeout, not by anything the daemon chooses. Without that
sentence a lead reads 'the ask now nudges me' as 'asking works now' and briefs a
worker to ask — which is the failure CB-582 was filed about.
Add the CB-586 retention rule to GitWorktrees and drive it from the reaper:
a snapshot is deleted only when its tree content is already reachable from
main AND the ref is older than 24h. Reachability keeps the last copy of a
worker's work; the age floor stops a fresh snapshot being swept while a
lead is still looking at it. Every deletion logs the ref name and commit
sha so it is recoverable from the reflog. The /members response gains a
wipRefs{count,costBytes} census the operator can read without shelling
into the repo.
Issue #82 step 1 is a measurement, and the classifier refuses an ad-hoc pipeline
that enumerates credential names inside a member — correctly. This is the seam:
one file the operator reads once and then runs, instead of approving a shell
pipeline they have to take on trust.
It never prints a credential value or any part of one. #82's criterion 1 asked
for a 6-character prefix; this prints a truncated SHA-256 instead. A prefix of a
short secret is most of the secret and would end up pasted into a ticket, while
the hash answers every question the prefix was for — is it set, is it the same
value as over there, is it the CB-592 sentinel.
Refuses to run unless BRIDGED_MEMBER=1, since the finding is what a MEMBER holds;
--allow-outside-member takes the comparison reading and labels it as such.
I have not run the reading path. That is the operator's call, which is the whole
point of the ticket.
Three fields had CB-604's shape — lower-cased, compared against one string,
never checked against the valid set. The auth.mode one was the worst: a typo of
'token' silently behaved as loopback-trust, and validateAuthExposure() only fires
on a non-loopback bind, so a loopback bind hid it end to end. The daemon started
clean and authenticated nobody while the operator believed token mode was on.
All three now refuse at config load, naming the value and the accepted set. The
top-level placement policy was already validated but only lazily at first spawn;
it now calls PlacementPolicies.fromName eagerly at load, so a bad name cannot
start a daemon that merely looks healthy.
Verified by probing the real BridgedConfig.load with eleven values, including the
critical auth.mode typo on a loopback bind, and by loading the live gitignored
bridged.yaml — which the worker cannot see and so could not check.
Closes#106
An auth.mode typo (e.g. "toekn") used to silently fall back to loopback-trust with no signal
anywhere — validateAuthExposure() only checks the pairing on a non-loopback bind, so on the
common loopback bind the daemon started cleanly and authenticated nobody. Per-profile placement
had the same shape, falling back to legacy pane placement. The top-level placement policy name
was already validated by PlacementPolicies.fromName, but only lazily at first spawn through
CompositePeerLauncher's Supplier; it is now checked eagerly at load, calling fromName itself as
the single source of truth.
Found reviewing CB-582 before the merge, not by the implementer.
ask() clears its question on three paths — no-waiter, timed out, and (from
answer()) answered. It can also leave by throwing: an interrupt while blocked on
the answer, or an ExecutionException from the answer future. Those run only the
finally block, which tore down the rendezvous turn but not the push loop's copy.
The result was a question that stayed pending for good: named in every nudge
until it hit its own cap, then left in pendingQuestions with no remover at all.
Teardown now happens where the rendezvous teardown already happens, so the two
cannot drift apart again. Closing a turnId that was never pending is a no-op, so
the normal paths are unaffected.
The new test fails on the pre-fix code with expected: <STOP> but was: <INJECT>.
An async (wait:false) delegation opens a ~55s reverse-rendezvous window when its
worker calls bridge_ask. A lead polling on its normal minutes-long cadence never
sees that window, so the worker times out and proceeds without an answer.
The question is now a third source in the CB-588 per-lead push schedule, with its
own per-item nudge count (CB-598's shape), and it is surfaced by bridge_status and
by REST /sessions/{id}/status and /tasks/{ticket} (which previously dropped turnId
on an ASKING phase, so a REST caller could see the question but not answer it).
Closes#61
kind: was lower-cased and compared against one string, so a typo like
'opencod' was accepted and routed to the claude-code adapter. With argv:
unset the launch command became the misspelled string itself, and the
daemon tried to run a program named after the typo. Nothing said a word
until the spawn failed.
It now throws at load, naming the profile, the bad value and the
accepted set - matching rejectNegativeMaxLoad and the duplicate-adapter
check, which already treat routing mistakes as fatal.
Verified here by probing the real BridgedConfig.load with five values:
opencod refused with the full message; opencode, OpenCode, claude-code
and an absent kind all accepted with the right adapter. 832 tests,
BUILD SUCCESS, unpiped.
Closes#102
bridge_ask blocks the worker's turn for ~55s by default (BridgedApp.java,
BridgeMcp.java) — a value deliberately kept just under the worker's own MCP
client's ~60s call cap so the daemon can return a clean timeout before the
client severs the call, not a value that can usefully be widened. A lead
following the charter's wait:false + poll cadence is minutes away, so the
window closes long before a poll would ever see the question — and until now
bridge_poll on such a ticket just read as ordinary "pending" progress.
bridge_poll(ticket) already surfaced Phase.ASKING with the question and
turnId (CB-205); this ships the two pieces that were still missing:
- The lead's own pane is now nudged the instant a question opens, reusing
the CB-588 ReplyPushLoop push mechanism (a third source alongside queued
replies and terminal tickets) rather than a new path. The nudge is capped
by the loop's existing maxReminders budget, and stops the moment the
question is answered or lapses.
- bridge_status(sessionId) and REST GET /sessions/{id}/status now also show
an open question and how to answer it, via a new
MessageService.pendingAsk() lookup — covering the case where a lead checks
status directly rather than the ticket.
- The REST /tasks/{ticket} endpoint was silently missing turnId on an ASKING
phase (only the MCP layer's formatted text carried it) — fixed as part of
making the state genuinely visible over both surfaces.
An unanswered question still behaves as today: the worker proceeds and its
reply says the ask went unanswered — not a hard failure.
Most of issue #65 was already shipped in 5d5b3bd - MemberSession
records agentSessionId, PeerHandle.agentSessionId() has no default, the
roster exposes it, and bridge_spawn accepts sessionName and
resumeSessionId. That commit left one item for follow-up: carrying the
id on ReleaseDetail.
This does that item. CB-578 stage C already lets a lead re-dispatch onto
the same worktree after a failure; without the session id that is a cold
start. With it, the work and the thread both survive.
Also updates the bridge_spawn and bridge_list rows in the CLAUDE.md
intent table, which the earlier commit missed.
Verified here: trial merge onto main builds 830 tests BUILD SUCCESS,
unpiped. Confirmed against the code that items 1-4 really were already
on main, so the scope-down is correct rather than work skipped.
Closes#65
The weight ratio does not express a preference order. weighted spreads
spawns across every profile with a free slot, so paid spawns happen
while the free box is idle - and a profile at maxLoad freezes its score,
so it can lose the next pick after a slot frees.
The real fix is a cost-first policy (CB-589). This documents the
workaround and its trap next to the key, because bridged.yaml is
gitignored: a fresh host starts without the workaround and quietly pays,
with nothing to tell the operator why.
Comments only. BridgedConfigTest: 85 tests, BUILD SUCCESS.
Closes the one piece issue #65 deliberately left out of 5d5b3bd: a
released member's agentSessionId now rides alongside worktree, branch
and snapshotRef on ReleaseDetail, and Bridged's onRelease handler
names it in the abandon reason, so a lead can resume the member's
conversation instead of only re-dispatching a fresh one onto the same
files.
Also updates the CLAUDE.md bridge_spawn/bridge_list table row, which
5d5b3bd shipped the sessionName/resumeSessionId/agentSessionId surface
for but never updated.
Three gaps that only bite once the agent is loaded, plus one wrong
comment.
The script computed its log path from its own location while the plist
hard-codes one. Run from a different checkout, every post-restart check
would read the wrong file and report a clean restart while the daemon
crash-looped. It now compares the two and fails, not warns.
A failed 'launchctl load' after a successful 'unload -w' left the agent
stopped AND persistently disabled - worse than before the redeploy. It
now retries once, then dies naming the exact recovery command.
The plist now says plainly that ThrottleInterval paces restarts but does
not bound them, and what actually stops the loop.
Verified here: ran the script with --check from the merged tree and it
behaves exactly as before, so the unsupervised path - my only restart
route - is intact. Exercised the log-path check against match, mismatch
and missing-plist fixtures using a truncated copy with no mutating code
in it: ok/1/1. 829 tests BUILD SUCCESS.
Closes#91
bridged.yaml is gitignored, so bridged.example.yaml is the only
committed description of the config schema. Two tests already covered
example -> code; nothing covered code -> example, so a brand-new key
could ship undocumented and no test would notice.
A new test compares BridgedConfig.KNOWN_TOP_LEVEL_KEYS against the
example scanned as TEXT, so a key documented only as a comment counts as
documented. That is what makes the guard correct rather than annoying:
most of the example is commented on purpose.
Verified here: added an undocumented key and watched the test fail with
an actionable message naming it; then documented that key as a comment
only and watched it pass. Probe reverted, tree clean.
Closes#96
The reminder count was one counter per lead per source, carried forward
across ticks. A counter carried forward has no memory of which item it
counted, so work arriving during the backoff window inherited an
already-capped count and was never named in a nudge.
tick() now recomputes each source's count fresh from the minimum count
among the items actually pending, tracked per item. A fresh item keeps
its source eligible; an older capped item still rides along in the text
without spending more budget. decide() is unchanged.
Verified here: read the diff; the bumped set is exactly the set named in
the nudge, and the empty early-return skips the bump. Trial merge onto
main builds 826 tests BUILD SUCCESS, unpiped.
Closes#87
Background loops call the fake from their own scheduler threads while a
test polls called() from the test thread. The list was a plain
ArrayList, so a nudge landing mid-stream threw
ConcurrentModificationException out of called().
It surfaced while I was verifying CB-598, which nudges more often, but
the race is on main today and is unrelated to that change.
824 tests, BUILD SUCCESS.
- redeploy-bridged.sh now refuses (not warns) a supervised restart when
its computed log path disagrees with the loaded plist's StandardOutPath
— otherwise every post-restart check reads the wrong file and can
report a clean restart while the daemon crash-loops. The check is a
pure, testable function; the script gained a source-for-test guard so
it can be exercised without installing the agent or touching launchd.
- a failed 'launchctl load' after a successful 'unload' now retries once
and, on ultimate failure, tells the operator the agent is stopped AND
disabled plus the exact recovery command, instead of leaving that
silently worse than the pre-redeploy state.
- the plist documents honestly that the crash loop launchd retries is
unbounded (ThrottleInterval only paces it), and what actually stops it.
- fixed the requiredSecretEnvVars javadoc: the auth.tokenEnv startup
throw is ~370 lines below its call site, not a few lines above it, and
only fires in auth.mode: token.
The test asserted one of two interleavings that are both correct, and
steered toward it with a 5 ms Thread.sleep. Under load the other
interleaving happened and main went red on a correct implementation.
The head start is now a latch counted down from inside the sweep's
guarded loop, so the ordering is guaranteed, not likely. Test file only;
the production guard is unchanged.
Verified here: 822 tests BUILD SUCCESS; 10/10 passes while a full clean
install ran in parallel; 5/5 failures with the guard removed, so the test
still catches the bug it exists for.
Closes#95
BridgedConfig.KNOWN_TOP_LEVEL_KEYS is now package-private so a test can assert
every key the parser accepts appears in bridged.example.yaml — live or
commented-out, since the file is gitignored and the example is the only
committed description of the config schema. The existing tests only checked
the example->code direction; this adds code->example.
My premise for this ticket was wrong and the worker corrected it. I reported five whole sections missing from the example; nothing was missing. My comparison script only counted uncommented lines, so every section documented as a commented-out example looked absent. Earlier tickets had each updated the example alongside their feature.
What it found instead is more useful than what I asked for — real inaccuracies, found by tracing each field through the parser and its consumers:
- `health.workingSuspectAfterSeconds` and `paneProbeIntervalSeconds` documented enforced minimums that do not exist. I checked: both names appear **only** in the `Health` record declaration and are read by nothing. Only `intervalSeconds` is clamped, and it is silently raised to 15 rather than rejected.
- `notifications.mode: webhook` only flips what `bridge_list` reports as `healthCoverage`. It sends no webhook — "webhook" appears in one `configured()` boolean and there is no delivery code in the repo.
- `lifecycle.clearAfterTurn` was undocumented, and is a no-op for any peer kind other than claude-code.
- The reload doc claimed the whole `fleet:` block is hot; `fleet.leaders` is built once at startup and is not rebuilt, so a change is silently accepted and does nothing until a restart.
- The `fleet.leaders` demotion consequence is now stated next to the block itself: an unmatched pane is silently an ordinary worker and every orchestration call it makes is refused, with no startup error.
Documenting a knob as dead is worth more than documenting it as working. Someone tuning `workingSuspectAfterSeconds` would otherwise have concluded their monitor was broken.
Comments only — no parsing or production code touched. Verified by the lead: parses cleanly under the project's own snakeyaml 1.30, top-level live keys `[bind, herdrSocket, profiles, placement, fleet, guard]`, the rest correctly commented examples. Both dead-knob claims verified by grep against `src/main` rather than taken on the worker's word.
`PlacementException extends IllegalStateException`, and neither spawn path caught that type, so it escaped to Javalin's default handler as a bare `500 Server Error` with a text/plain body — while every other failure on the same endpoint returned structured JSON. The reason existed and was good, but only in the daemon log.
I hit this live while orchestrating: asked for a member on a full profile, got a blank 500, guessed another profile, got a blank 500 again, and spent two round trips learning things the daemon already knew.
Both surfaces now catch it. REST returns 503 with `{"error":"no_capacity","detail":...}`; MCP returns the same reason in the `isError` shape it already uses for every other spawn failure. 503 is right because the request was valid and will likely succeed later — the caller did nothing wrong, so 400 would have been a lie.
The other throw sites all funnel through the same type, so quarantine cooldowns, weight-0 exclusion, and the all-at-cap / all-quarantined / all-unreachable messages now reach callers too. That last group matters most: those three distinguish "wait a moment" from "your backends are gone", and all three used to arrive as the identical blank 500.
Tests assert the caller can read the *reason*, not merely that the status changed — one per surface.
Verified by the lead: `mvn -f bridged/pom.xml clean install` unpiped, exit code captured — Tests run: 824, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS.
No exception message was reworded. This change delivers messages that were already written.
Work that arrived during the ~15s push_backoff_ms window between two
ticks landed in the pending map before the next tick's start-of-tick
snapshot, so a shared per-lead-per-source counter (carried forward via
scheduleNext(lead, count+1, ...)) already treated it as exhausted
backlog even though no nudge had ever named it. ReplyPushLoop.tick now
recomputes each source's reminder count fresh every tick as the
minimum nudge count among that source's currently pending items, so a
freshly-arrived item (count 0) keeps its source eligible regardless of
how depleted an older, still-undrained sibling's count is. decide()
itself is unchanged.
PlacementException extends IllegalStateException, which neither BridgedApp
nor BridgeMcp's spawn catch blocks handled, so a maxLoad/quarantine/
all-exhausted refusal fell through to a blank 500 on REST and lost its
message on MCP. Catch it on both surfaces, ahead of the unrelated
IllegalArgumentException(unknown_profile) mapping, and return its message
structured: REST as {"error":"no_capacity","detail":...} with status 503,
MCP as an isError result prefixed "no capacity: ...".
Audited bridged.example.yaml against BridgedConfig's KNOWN_TOP_LEVEL_KEYS and
found every top-level key already documented (broker, health, lifecycle,
configReload, quarantineCooldownSeconds, fleet.leaders/architects/reviewers
included) — CB-573/CB-566/CB-559/CB-579/CB-527/528 each updated the example
alongside their feature. What was actually wrong:
- health.workingSuspectAfterSeconds/paneProbeIntervalSeconds claimed enforced
minimums (300/60) that don't exist in code — only intervalSeconds is
clamped (floor 15); the other two are parsed but never read anywhere.
- notifications.mode: webhook was undocumented as only flipping the
healthCoverage label bridge_list reports — no webhook is ever sent.
- lifecycle.clearAfterTurn was missing entirely.
- the HOT bullet under configReload claimed the whole fleet: block reloads
live, but ConfigRef's own javadoc carves out fleet.leaders as needing a
restart with no deferred-list warning — added that exception.
- fleet.leaders' demotion consequence (unmatched tab -> silent WORKER
demotion, no startup error) is now stated inline next to the block, not
just implied by the multi-lead rationale higher up.
The launchd unit was a CHANGEME template that had never been installed, and it could not have worked if it were: launchd does not source a login shell, so the daemon would have started with no forge or gateway token, and the failure would only appear much later as workers unable to open a PR.
Four parts: a wrapper that execs one login shell in place so the job inherits the secret store; a startup report naming which required secret env vars resolved and which are MISSING, by name only, never a value; the plist filled in with this host's real verified paths; and `redeploy-bridged.sh` detecting the agent and switching stop/start to `launchctl unload -w` / `load -w`, falling back to the existing kill + nohup when it is not installed.
Verified by the lead. My own unpiped build of the branch: 813 tests, BUILD SUCCESS. Trial-merged onto current main (which had moved twice) and rebuilt: Tests run: 822, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. All three host paths in the plist exist; no CHANGEME remains; the secret report leaks no values.
An independent reviewer, briefed only from the diff, verified the parts that matter today and said merge. It confirmed live that the unsupervised path is unchanged, that `launchctl list` exits 113 for an absent agent so the detection reads correctly, that the wrapper round-trips arguments containing spaces and quotes and stays a single exec chain, and that neither branch of the secret report can print a value.
Its two open findings only bite once the agent is actually loaded, which has not happened and is the operator's call. Filed as #91 — that must land before the agent is ever installed. The important one is that the script computes its log path from its own location while the plist hard-codes an absolute one; if they ever disagree, the post-restart ERROR check reads the wrong file and reports "ok" while the daemon crash-loops.
The agent is deliberately NOT installed by this merge. Nothing here changes how the daemon runs today.