4bfab6b71805153d4aa9a14577e49af70ecee4d3
27 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4bfab6b718 |
fleetd #480 follow-up: resolve a relative leadRollover.handoverPath against the calling lead's workspace
LeadRollover.open() now resolves handoverPath to an absolute path exactly once, against the calling lead's fleet.leaders.<name>.cwd (falling back to the daemon's own user.dir when that lead has none configured), matching the LeadLauncher#launch precedent. PendingRollover stores only the resolved absolute path, so checkHandover's exists/empty/fresh checks, the path handed back to the lead in the fleet_handover open response, and the default bootstrapText sentence all see the same absolute location instead of a value resolved against whatever directory the daemon process happened to start in. FleetConfig.LeadRollover.bootstrapText is no longer defaulted in the compact constructor (it would otherwise still bake in the raw, possibly-relative handoverPath); a new bootstrapTextFor (resolvedHandoverPath) method builds the default sentence from the resolved path instead. Fleetd.leadRollover(...) gains a required liveLeadTerminals parameter to build the terminal to lead-name to Leader.cwd lookup, read live through the existing `leads` supplier and ConfigRef on every call, never off a startup snapshot. |
||
|
|
a94262271b |
fleetd #480 correction round: defer the roll, and gate it on caller identity
Two defects found after the fact, both from the original brief, both fixed here. 1. confirm() is called FROM the calling lead's own turn, so its pane is still WORKING and can never report injectable inside that same call. The old confirm() sent /clear before polling for that — the poll always timed out, but only after /clear had already fired and queued, destroying the lead's context with no fresh session ever started and a refusal return that lied about what had happened. Fix: confirm() now only validates and, if every gate passes, hands a one-shot continuation to a new continuationRunner (a real virtual thread in production, Runnable::run in tests) and returns RollDecision.approved() immediately - "scheduled", not "rolled". The continuation itself does the actual work, once the calling turn has ended: wait for the SAME pane to report injectable again (new turnSettleSeconds config key, default 20) - if this never happens, /clear is NEVER sent, at all - then /clear, then wait again (clearSettleSeconds, as before), then bootstrapText. The "no timer/scheduler, only confirm() can roll" invariant is restated precisely in LeadRollover's class javadoc: it is about initiative, not synchronicity - a single-shot continuation of an already-approved confirm() call still satisfies it; a recurring background loop would not. 2. confirm() resolved the pane to clear via PrimaryRegistry.primaryTerminal(), a single-slot lookup that is correct for a background loop with no caller but wrong here: on a daemon with more than one labelled lead tab, lead X's confirm() could clear lead Y's pane, violating the charter's "identity comes from the connection, never an argument" invariant. Fix: open() and confirm() now take the caller's terminal id as a parameter (resolved by the MCP layer from the connection - the later MCP-tool unit must pass it in, never accept it as a request field). confirm() refuses with a new NOT_YOUR_ROLLOVER reason unless it matches the terminal open() recorded. LeadRollover no longer depends on PrimaryRegistry at all. Also: renamed RollResult to RollDecision (rolled -> accepted) to reflect the new meaning - approved and scheduled, not necessarily cleared yet. Added turnSettleSeconds to the leadRollover: config block (documented in fleetd.example.yaml alongside the existing keys) and updated Fleetd.java's leadRollover(...) factory to drop the primaryRegistry parameter, with FleetdLeadRolloverWiringTest's source-text pin updated to match. New tests: turnThatNeverSettlesSendsNoClearAtAll (the branch that matters most - a turn that never ends means /clear is never sent) and aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover (NOT_YOUR_ROLLOVER), plus a settle-after-clear timeout test and an open() input-validation test. LeadRolloverTest: 11 -> 14 tests. |
||
|
|
5c12865c25 |
fleetd #480 Unit A: lead rollover core (config block + executor)
Adds the opt-in leadRollover: config block and LeadRollover, the executor a later unit's MCP tool will call. A lead writes a handover file, then open() records a token and confirm() verifies it (exists, non-empty, fresh) and an operator confirmation before clearing the lead's own pane via /clear (sent directly through AgentControl, bypassing Injector, same as ClaudeCodeLauncher#clearContext) and bootstrapping a fresh session. Nothing but an explicit confirm() call can ever roll a pane - no timer, no heartbeat, no background thread anywhere in this class. Wired into Fleetd.java exactly like LeadHeartbeatLoop: constructed only when leadRollover: is present at startup, and nothing calls it yet - the MCP tool is a separate, later unit. Classified leadRollover: as HOT in ConfigRef (joins placement/ memberCredentials/memberLoginShell/models): the executor holds Supplier<FleetConfig.LeadRollover> and reads every field fresh per call, unlike LeadHeartbeatLoop's frozen final fields. The one caveat: the object's construction is still gated on presence in the startup config snapshot, so a freshly-added block needs a restart before anything exists to call. Tests: LeadRolloverTest (14 cases covering the 6 hard requirements - no object without the config block, only confirm() can roll, missing/empty/ stale handover file each refuse by name, requireOperatorConfirm gating, and the injected wall-clock supplier) and FleetdLeadRolloverWiringTest (source- text pin on Fleetd.main's construction call, mirroring FleetdCompletionResolverWiringTest). Also updated the existing FleetConfigValidateAllTest, FleetConfigWithDefaultsPreservesEveryComponentTest, ConfigRefTopLevelCoverageTest and ConfigRefTopLevelReportingCoverageTest to account for the new record component. |
||
|
|
17052bb515 |
Merge #471: seeded skills reach an opencode member, and instructions[] stops depending on write order (fleetd #393)
memberSkills: copied skill folders into every provisioned worktree's .claude/skills/ and stopped there. Claude Code reads that directory natively; opencode never does. So the feature was INERT for opencode members rather than broken: the copy succeeded, the files were correct, and nothing ever read them. No test failed because there was nothing to fail - the feature worked at the only layer it implemented. An opencode member's only channel for static guidance is the instructions[] array in its generated config. OpenCodeLauncher now adds each seeded skill's SKILL.md there. A folder with no SKILL.md is never delivered and the log names it. The second half is an ordering hazard fleet01 found by reading, and that I then measured. Three writers append to instructions[]: the role charter, the seeded skills, and the IDE rules. The charter used putArray (CREATE-OR-REPLACE) while the other two used withArray (get-or-create). That was safe only because the charter ran first against an empty array - an undeclared constraint that nothing tested. Measured on the earlier merge: making the skills writer use putArray left 1603 tests green while silently deleting the charter entry, so an opencode member would launch with no role contract at all. Worse than the bug being fixed, and invisible. All three writers now use withArray, and three tests pin the array's CONTENTS (never its size - a size assertion passes when putArray swaps two entries for two others) across the combinations that matter: charter-only, charter+ide, charter+skills+ide. WHY NOTHING CAUGHT IT, MEASURED RATHER THAN ASSUMED. fleet01 first said the missing axis was the COMBINATION of writers, then revised that to a stronger claim: that nothing asserted the charter reaches an opencode member at all. I checked the second claim on this merge and it is FALSE. Deleting the charter writer outright fails 6 tests, and 3 of those existed before #393: writesRemoteMcpConfigAndCharterInstructionsWhenMcpUrlSet, roleCharterWithoutMcpOrCustomProviderStillWritesAConfig, and aPinnedEndpointAndTheFleetMcpCoexistInOneConfig. The charter reaching a member was already pinned. So their FIRST diagnosis was the right one. The surviving mutation did not delete the charter write; it made a LATER writer replace the whole array. Every pre-existing test had exactly one writer active, and with one writer putArray and withArray are indistinguishable. The gap was never "is the charter delivered" - it was "are two writers ever active at once", which is the combination axis. Recording this because the stronger claim is the more quotable one and it would have sent the next reader looking for a hole that is not there. One honest residue: the charter writer's own idiom cannot be pinned. Flipping it back to putArray leaves the suite green, and always will, because it runs first against an empty array where the two idioms are equivalent. The edit removes an undeclared constraint for the next person to add a writer; it is not a change any test can detect. The comment in the source claimed two named tests cover it - that was wrong, and the commit after this one corrects it rather than leaving a false claim beside the code. What this does NOT do: opencode has no equivalent of Claude Code's Skill tool, so the content is static system-prompt text present from spawn, not something a member can invoke by name. This closes the DELIVERY gap and cannot close the ACTIVATION gap. That residue is opencode's design. Numbers and my own mutation battery are on the ticket, measured on this merge commit. |
||
|
|
d4a2cd720c |
fleetd #393: deliver memberSkills to opencode members, and stop overclaiming seeding success
GitWorktrees.seedSkills copies memberSkills:-seeded skill folders into every provisioned worktree's .claude/skills/ and logged "skill seeding: N of M" as if that were success — but .claude/skills/ is a Claude Code CLI convention. opencode has no such discovery, so a kind: opencode member never actually read a seeded skill even though the log said N of M succeeded. Two changes, both required: 1. Deliver it. OpenCodeLauncher.skillInstructionFiles scans <cwd>/.claude/skills/*/SKILL.md at spawn time (the one point the launcher knows both the kind and the cwd) and appends each to the generated opencode.json's instructions[] array, the same channel already used for the member charter and IDE rules. A skill folder with no SKILL.md is named and skipped rather than silently dropped. 2. Stop claiming it where the claim can't be verified. GitWorktrees.seedSkills' log now says explicitly that consumption depends on the member's kind and points at the launcher's own log; OpenCodeLauncher logs its own kind-aware "skill delivery: M of N ..." line once the kind is actually known, naming any folder it could not turn into an instructions[] entry. fleetd.example.yaml's memberSkills: doc previously claimed "Claude Code members only; an opencode member reads a different path (.opencode/agent) this key does not touch" — false as of this fix, corrected to name both kinds and how each consumes it. Tests: OpenCodeLauncherTest gains two cases driving the real GitWorktrees#add seeding path (not a hand-built fixture) into an opencode-kind spawn — one asserting a seeded skill's SKILL.md lands in instructions[] plus the honest log line, one covering a skill folder without SKILL.md (delivered skills still flow, the malformed one is named in the log and excluded from instructions[]). ClaudeCodeLauncher is untouched — its native .claude/skills/ discovery already worked and is out of scope. mvn -B clean test: Tests run: 1603, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS |
||
|
|
5a467e1f8b |
fleetd #466: escalate BackendQuarantine's cooldown on repeated exhaustion
A flat 30-minute quarantine retries a weekly subscription limit about 336 times before the window resets. BackendQuarantine now doubles the cooldown on each consecutive exhaustion of the same credential (no more than one base cooldown after the previous quarantine's deadline), capped at 12x the base cooldown (~6h at the 1800s default), and resets back to the base cooldown once a base-cooldown's worth of quiet has passed with no further exhaustion. The flat two-argument constructor is unchanged (equivalent to multiplier 1.0 / ceiling == base), so all ~20 existing call sites keep their current shape and behaviour. Production wiring (Fleetd.main) switches to the new BackendQuarantine.withEscalation factory. This only touches the exhaustion path (BackendQuarantine's one production caller is Fleetd.exhaustionSink, fired on BACKEND_EXHAUSTED alone) and never the separate, unescalated cooling-off mechanism (BackendOutagePolicy, fixed 60s) that guards against a transient backend-error storm. |
||
|
|
45aca9eb3e |
fleetd #446: make exhaustedPattern hot, name the fix in the warning, report it in fleet_profiles
The model gate can be turned off at runtime (models.allow[].enabled: false, hot since fleetd #422), but the usage-limit detector it's meant to react to was compiled once at Fleetd.main startup into a frozen Map<String,Pattern> — arming or disarming exhaustedPattern needed a daemon restart. Backwards for a feature meant to react live. - New LiveExhaustedPatterns: reads exhaustedPattern off the live config supplier per lookup (matching CompositePeerLauncher#models0's live-supplier pattern), caching compiled Pattern objects by PROFILE NAME (not pattern text — pattern text would grow unboundedly as an operator tunes a regex across reloads; profile names are bounded by the small, restart-gated set of configured profiles). patternFor() backs CompletionResolver's classification; armed() backs fleet_profiles' exhaustionDetectionArmed — both read the same object, the fleetd #404 single-accessor rule CompositePeerLauncher.modelGateState() established for the model gate. - Fleetd.java: on BACKEND_EXHAUSTED, log a WARNING naming the profile's model and the exact fix (enabled: false under models.allow, hot, no restart; remove it again once the window resets) — or, when the profile has no model: configured, say quarantine is the only thing keeping spawns off it. - FleetMcp.QuarantineSource gains modelFor/reasonFor; fleet_profiles' quarantined rows gain model/reason fields so a lead can see why without reading the daemon log. capacityView (fleet_list) intentionally untouched — scoped to fleet_profiles only. - ConfigRef/FleetConfig docs + fleetd.example.yaml updated: exhaustedPattern moves from Deferred to Hot. errorPattern stays deferred on purpose (out of scope for this ticket). - Tests: LiveExhaustedPatternsTest (new, unit-level hotness/caching proof), FleetdExhaustionDetectionArmedWiringTest (rewritten — same QuarantineSource object, read before and after a reload, asserts the answer flips with no restart), ConfigRefTest/ConfigRefProfileCoverageTest updated for the new Hot/Deferred classification. |
||
|
|
e7b33fe3a0 |
fleetd: central allow-list of usable models (models: + validateModels())
Add an optional top-level `models:` block (Models{allow: List<ModelEntry>})
naming the models any profiles: entry may use. Absent/empty allow: keeps
today's behaviour exactly (no check, no warning). When configured,
FleetConfig.validateModels() fails config load (and reload, via ConfigRef)
naming both the model and the profile, if any profile's model: is outside
the list. The check is one-way: editing profiles: alone can never widen
what is permitted, only models.allow: can.
Wired into Fleetd.main() alongside the other validateXxx() calls, and into
ConfigRef.reload()/DEFERRED_KEYS so a bad edit can't slip in through a
reload either. Each ModelEntry is its own record (not a bare string) so a
later unit can add per-model on/off or load-limit state without changing
the YAML shape. One flat string namespace covers both a bare Claude id and
an opencode provider-prefixed id.
|
||
|
|
24b96d29ae |
fleetd must not let the host idle-sleep while members are live
Adds a small IdleSleepGuard (dev.ltms.fleet.power) that holds a macOS caffeinate -i child while at least one fleet member is live, and releases it once none are. It hangs off SessionManager's existing onAcquire/onRelease hooks and SessionManager#size() rather than tracking members a second way. New idleSleepGuard: config block, on by default, following the FleetConfig.Health/ConfigReload pattern. |
||
|
|
92c0f164f1 |
Merge #366: seed the bridge's worker skills into every provisioned worktree
fleetd #362 item 3. A member spawned against a repo that does not ship its own .claude/skills/ could not load implementer, reviewer or hunter at all. Every brief starts with "Load the <name> skill", and outside this repo that line was silently a no-op. memberSkills: <dir> now copies those folders into each provisioned worktree, skipping any name the target repo already ships. Two review rounds, both about the same hazard: core.excludesFile is single-valued, so pointing it at fleetd's own file would SHADOW the operator's. It now composes instead of replacing, and the XDG default excludes file is carried forward too. Verified on this merge, not taken from the worker's report: mvn clean install -> Tests run: 1402, Failures: 0, Errors: 0, BUILD SUCCESS Two mutations run on merge: drop the XDG fallback -> 1 failure, BUILD FAILURE remove the composition itself -> 2 failures, BUILD FAILURE Both directions are pinned. |
||
|
|
f84824ee29 |
fleetd #362 review fix: compose skill-seeding excludes with the operator's own excludesFile
core.excludesFile is single-valued, so pointing it at fleetd's own seeded-skill exclude file with --replace-all at worktree scope was SHADOWING whatever excludesFile the worktree already resolved (an operator's global config, most commonly) instead of adding to it. This repo's own .gitignore does not ignore target/ — only an operator's global excludesFile does — so every worker's `mvn clean install` would make target/ show up as untracked, and CB-576's deliberately untracked-inclusive hasUncommitted would then read every such worktree as dirty forever, so it is never cleaned up. excludeSeededSkillsFromGitStatus now reads whatever core.excludesFile resolves to BEFORE writing anything (falling back to git's own $XDG_CONFIG_HOME/git/ignore default when the key is unset entirely, per gitignore(5)), and writes that content into fleetd's own exclude file ahead of the seeded skill patterns, so every operator-configured pattern keeps applying inside the seeded worktree. Proven with a new test, seedSkillsComposesWithAnAlreadyEffectiveGlobalExcludesFile, which isolates a synthetic "operator's global config" via a new gitEnv test seam on GitWorktrees (GIT_CONFIG_GLOBAL pointed at a throwaway temp file, never the real machine's config) and drives the real add() path end to end. Also documents (FleetConfig javadoc + fleetd.example.yaml) that memberSkills copies every non-hidden subdirectory of its source wholesale, with no per-file allowlist. |
||
|
|
7c684e40d3 |
fleetd #362 (item 3): seed .claude/skills/ into provisioned worktrees
Add memberSkills: <dir> to FleetConfig. GitWorktrees#add copies each skill folder from that directory into <worktree>/.claude/skills/ so a member spawned against ANY repo — not only one that already ships its own skills — can load a bridge skill (e.g. implementer). A skill the target repo already carries is never overwritten. Every seeded path is hidden from `git status` in that worktree ONLY, via a --worktree-scoped core.excludesFile pointing at a file under the worktree's own private git dir (outside the working tree, so it can never be committed) — not the shared .git/info/exclude, which a linked worktree resolves to the repo's common git dir and would otherwise leak visibility changes into the primary checkout and every sibling worktree. Proven with a real `git status --porcelain` in GitWorktreesTest, not by reasoning. Seeding is best-effort like the existing overlayParity/isolateToolSurface steps: a missing/unreadable source or a copy/exclude failure is logged and skipped, never fails the spawn. memberSkills is triaged as a DEFERRED config key in ConfigRef (baked once into GitWorktrees at startup, like worktreeGroup), with its own changedDeferredKeys branch and coverage-test entries. |
||
|
|
0df34f3220 |
fleetd #361: close the lead-coordination visibility gap
Lead-to-lead AMQP coordination had a send half with tools and a receive
half without. This closes three blind spots:
- LeadChannel gains inspect(coordId) -> MailboxState(exists, pending,
consumers), implemented in LeadMailbox with a throwaway probe channel
(never the long-lived publish/consume channels) so a passive-declare
404 on a missing queue can never take down publish() on the same
instance.
- FleetConfig.Coordinator gains peers: List<String> (defaults to empty,
blank entries dropped) so a daemon can declare which peer coord-ids
it expects to reach.
- fleet_list reports coordination state via a new CoordinationSource
(own coord-id, own mailbox state, held messages as msgId/from/preview
only, and one row per configured peer with reachability/pending/
consumers), following the existing OutageSource/QuarantineSource
"Source record with none()" idiom instead of growing listFleet's
overload chain by another positional parameter. Every peer probe is
bounded by a 1.5s timeout on a virtual-thread pool and degrades to
absent rather than ever slowing or failing fleet_list.
- fleet_send{coordId}'s success text now says "durably confirmed by the
broker" instead of "delivered", and warns (while still reporting
success) when the target mailbox has zero consumers attached.
Tests: hermetic unit tests for exists/absent/zero-consumer/old-config-
no-peers-key/new fleet_list shape using FakeLeadChannel, plus a
@Tag("contract") LeadMailboxTest.inspectingAMissingMailboxNeverBreaks
PublishOnTheSameInstance proving the invariant against a real broker.
|
||
|
|
d42c2bc204 | fleetd #266: rename SSH agent environment setting | ||
|
|
a2b8caf6b5 |
#184: keep the reason SSH_AUTH_SOCK matters, and the measurement
The correction removed a false claim (blocking the socket breaks git over SSH) but took a true one with it: the socket is a live handle to the agent, so a member holding it can sign with every key the agent holds. Without that, the entry reads as if the setting does not matter, and an operator has no reason left not to set it to allow. Fixing an overclaim must not leave an underclaim. Also record the measurement and the mistake behind the old claim, so the next person does not re-argue it from scratch. |
||
|
|
d223a93039 | fleetd #184: correct sshAuthSock guidance | ||
|
|
279d6f5fbd |
fleetd #257: free must stop subtracting leadSeatCount
fleet_list's free row subtracted leadSeatCount, but the real spawn gate (CompositePeerLauncher#enforceMaxLoad) only ever compares live against maxLoad and never reads leadSeatCount. So free could report 0 while a fleet_spawn on that exact profile still succeeded, and a lead trusting free:0 gave up on capacity the gate would still grant. free now always equals max(0, maxLoad - live); leadSeats stays in the row as an informational fact, never subtracted. Documented in the fleet_list tool description and fleetd.example.yaml. |
||
|
|
c50f5b2d61 |
fleetd #176 stage 2: make effectiveCredentialId() subscription-aware
Stage 1's lead-seat matcher (leadSeatLookup) was correct but inert on the live host: the lead runs on profile 'opus', members on 'sonnet', both subscription:true with no explicit credentialId. Because effectiveCredentialId() fell back to the profile's own name, opus and sonnet never matched even though they share one Claude login, so the matcher charged zero seats. FleetConfig.Profile.effectiveCredentialId() now falls back to a shared sentinel (SUBSCRIPTION_CREDENTIAL_ID = "<subscription>") instead of the profile name when subscription:true and credentialId is unset. An explicit credentialId still wins, so two separate Claude logins on one host can still be kept apart. This is also BackendQuarantine's and BackendOutagePolicy's grouping key and CompositePeerLauncher's spawn-time enforcement key, so the fix also links quarantine/cool-off across subscription profiles sharing an account -- intentional: one usage limit really does take out every profile on that login, mirroring credentialId: openai-shared already doing this for off-subscription profiles. Every caller was reviewed; none wants "this exact profile" over "this account". Tests added: - FleetdLeadSeatLookupTest: the live shape itself (lead on a DIFFERENT subscription profile than the target, same account, neither sets credentialId) -- the case stage 1's suite never covered - FleetMcpTest: quarantining one subscription profile's shared account zeroes free on another sharing it, via the same effectiveCredentialId()-driven wiring Fleetd.main uses Mutation-tested: reverting the subscription branch to the old fall-back-to-profile-name behavior sends both new tests RED with 0 compile errors; reverting the mutation restores byte-identical (diff -q) source and green tests. fleetd.example.yaml's fleetd #176 notes are rewritten for the sentinel semantics and when to override it with an explicit credentialId. |
||
|
|
c796eac09c |
fleetd #176: subtract the lead's own subscription seat from free
maxLoad counted panes, never subscription seats: a subscription:true profile's lead is itself a live claude session on that same account, so free overstated capacity by the lead's own seat (measured free:1 with a real ceiling of 0, and free:3 on an idle fleet with a real ceiling of 2). Add FleetMcp.LeadSeatSource (same shape as QuarantineSource/ OutageSource) and Fleetd.leadSeatLookup, which derives the seat count from fleet.leaders.<name>.profile matched against the target profile by effectiveCredentialId() - no hardcoded "-1", and no new config key: profile: already exists for this exact "which account does this lead share" question. maxLoad itself is left untouched; only free (and a new, additive-only leadSeats field) changes. Exhaustion quarantine (cause 2 in the ticket) already forced free to 0 via the same BackendQuarantine capacityView already reads - confirmed by reading the exhaustionSink wiring, no code change needed there. |
||
|
|
8bba3a8184 |
fleetd #148 (point 2): drop .envrc from the default parityOverlay
.env is data; .envrc is executable shell that direnv runs on every cd, so copying it into a worker moves behaviour, not just values. The default parityOverlay is now [.env] only. The knob is unchanged: an operator who wants .envrc copied can still write parityOverlay: [.env, .envrc] explicitly. Updates FleetConfig's default and javadoc, fleetd.example.yaml's two mentions of the default, and the FleetConfigTest coverage: renamed the default test, added parityOverlayExplicitEnvrcOptInStillWorks to prove the .envrc opt-in escape hatch still works, and fixed WorktreeSessionManagerTest#worktreeAcquireRunsParityOverlayWithProfileDefaults which also hardcoded the old default. |
||
|
|
ba04b2359b |
fleetd #201/#227 unit 5: wire errorPattern, cool-off spawn gate, and fleet views
Wires the already-merged units into production: - Per-profile errorPattern config (beside exhaustedPattern), compiled once at startup; falls back to the legacy (?i)\bAPI Error\s*: pattern when unset. Startup logs configured-vs-legacy coverage, same as exhaustedPattern. - One production BackendErrorSink in Fleetd.java: mark backend_error on the session, resolve the profile's credential fail-loud (never Optional.ifPresent), record it in BackendOutagePolicy, and push a lead nudge on a new incident. - CompositePeerLauncher's explicit and automatic spawn paths both refuse a cooling-off credential; exhaustion quarantine wins when both are active. PlacementContext gets a separate coolingOff set so refusal text says "cooling off", never "exhausted". - fleet_list/fleet_profiles report coolingOffForSeconds as an independent fact from quarantinedForSeconds; both can appear together. - fleetd.example.yaml documents errorPattern and the 2/60/60 cool-off policy; CLAUDE.md tells leads how to read the two independent outage states. Also: FixedPlacementPolicy.java, not in the original file list, needed the same coolingOff filtering as PlacementPolicyUtil (it does its own inline candidate filtering rather than delegating). 1190 tests, 0 failures (up from the 1163 baseline); BUILD SUCCESS. |
||
|
|
0373b6c41b |
#213: fix ZDOTDIR credential scrub gate + directory under memberHerdrSocket
The memberCredentials.policy: allow-list ZDOTDIR scrub decided zsh-vs-not using fleetd's own process $SHELL and wrote the generated scrub dir into fleetd's own java.io.tmpdir. Under memberHerdrSocket: (member panes run as a different OS user than fleetd's own process) this silently protects nothing: the wrong shell decides the gate, and the directory can be unreachable to the member. - New FleetConfig.memberLoginShell: the member OS user's login shell, only ever read when memberHerdrSocket: is configured; fleetd's own $SHELL keeps deciding everything when memberHerdrSocket: is absent (byte-identical to before). - HerdrPeerLauncher.applyEnvironmentAllowListPolicy: memberHerdrSocket + memberLoginShell not configured/non-zsh falls back to the CB-596 sentinel overlay (warn loudly, never refuse to spawn). memberHerdrSocket + zsh memberLoginShell generates the ZDOTDIR under the configured worktreeRoot instead of java.io.tmpdir, shared with the existing worktreeGroup (reused, not a new key). - EnvAllowListScrub: new generate(parentDir, allowedNames, group) overload shares the generated directory via pure-Java POSIX group ownership (rwxr-x--- dir, rw-r----- files) — no external process spawn. - fleetd.example.yaml documents memberLoginShell: and worktreeGroup:'s reuse for the scrub directory (the live fleetd.yaml is gitignored). 4 new tests in HerdrPeerLauncherAllowListWiringTest cover the acceptance criteria; 3 of the 4 were watched failing against the pre-fix code. |
||
|
|
8067ee4ec4 |
fleetd #185: opt-in worktreeGroup config for group-shared worktrees
Adds worktreeGroup (top-level FleetConfig key), Worktrees.shareWithGroup (GitWorktrees impl: git config core.sharedRepository group + one-time chgrp/chmod g+rwX/setgid fix-up over the worktree, .git/objects, refs, logs, worktrees, and packed-refs when present), and wires SessionManager to call it AFTER overlayParity so overlay files are covered too. Off by default (byte-identical behaviour when unset). Documents the credentials-not-repository caveat in the javadoc and example config. |
||
|
|
fc655e78c2 | CB-185: route members to separate herdr | ||
|
|
7c4170ff6d |
CB-643: join the message-layer evidence to the health monitor
CB-640 published the three message-layer facts and CB-641 wired the herdr and time ones. This joins them, so every HealthSnapshot field now carries real evidence and the NOT_YET_OBSERVED placeholder is gone. That constant is what made 8 of the 9 fault states unreachable, GONE and NEVER_READY included, which is why CB-580's failTarget never fired. hasOrphanedDelegation is a true snapshot, but it can read true for one tick during an ordinary race: an async ticket exists before its virtual thread reaches rendezvous.open, so for that instant nothing is accepted or queued behind it. decide maps the field straight to DELEGATION_ORPHANED with no smoothing, so one racy read would log a fault that clears on the next tick. The monitor now requires two consecutive observations. That costs one interval on a real orphan and removes the false positive. Two tests drive real ticks against a genuinely orphaned ticket (an unanswered fleet_ask that lapsed back to PENDING), not the seam: one tick reports nothing, two report once, and a single clean tick in between resets the streak. Also correct two config comments. paneProbeIntervalSeconds is parsed and read by nothing, so its "minimum 60" note promised a floor that does not exist. 970 tests green. |
||
|
|
66e776d178 | CB-641: wire herdr health evidence | ||
|
|
2e138a199b |
CB-634: one shared "fleet" workspace + rename bridged -> fleetd cutover
Two changes ship together here.
1. One shared herdr workspace. The lead and every worker now live in one
workspace called "fleet", so the operator sees one "session" with many
windows, not two. Before, the lead sat in a "leads" workspace and workers
in "bridged-workers", which read as two sessions. The lead is still told
apart from workers by its exact tab label ("lead: <name>"), so putting them
in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
for split layouts; Fleetd now passes an empty exclude set.
2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
launchd/systemd units, module dir, and MCP mount).
- Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
- Log line, comments, docs, and CLAUDE.md updated to say fleetd.
- Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
- Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
- Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
still read as a fallback, and still gitignored.
- MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
server name is "fleet". The mount name in the local .mcp.json becomes
"fleet" (gitignored, not in this commit).
- Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.
Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.
Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.
The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).
949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
|