2026-08-28 - 2026-09-28
Overview
106 Pull requests merged by 2 users
Merged
#620 fleetd #612 Unit A: extract main's boot composition into FleetdAssembly/FleetdRuntime
Merged
#628 fleetd #612 B3: behavioural replacements for lead-seat, quarantine, lead-rollover guards
Merged
#626 fleetd #612 step 2 unit B2: behavioural replacements for the CB-185 pair
Merged
#627 fleetd #612 Unit B1: CompletionResolver behavioural test (replaces source-text guard)
Merged
#624 fleetd #612 A-gaps: coordinator path + reportRoleFallbackGaps inside the assembly boundary
Merged
#623 #608 replace MessageService timing sleeps
Merged
#622 fleetd #621: make the context-roll notice obey requireOperatorConfirm
Merged
#619 fleetd #618: state the measured auto-compact precedence
Merged
#601 CB-617: pass auto-compact window to leads
Merged
#617 fleetd #615: write FAILED instead of leaving status(token) stuck at IN_PROGRESS
Merged
#616 fleetd #613: boot log for role fallback gaps + heartbeat contextHighNudge
Merged
#614 fleetd CI: unreadableFileIsUnknown is red on main under root
Merged
#610 fleetd #609: nudge an idle lead to hand over when its own context reads HIGH
Merged
#611 fleetd #608: make anAlreadyCollectedTicketProducesNoNudge deterministic
Merged
#602 fleetd: report a lead's live context usage in fleet_list
Merged
#606 fleetd #602 gauge-wiring: thread a lead's configured configDir into the context gauge
Merged
#607 fleetd #603: share HEALTH_WAIT between the pid poll and the health check
Merged
#605 #604 item 1: fleet_list reports charterBytes alongside charterSha256
Merged
#599 fleetd #589 (Groups 1 & 2): pin 6 main() wiring sites with named factories
Merged
#597 fleetd #593 (pid-count half): running_pid() no longer matches the caller
Merged
#598 #589 Group 3: wiring-test 5 sites below line 500 in Fleetd.main()
Merged
#596 #568: add hunter member role
Merged
#583 #582: assert pending message-id cleanup
Merged
#574 fleetd #572: pin answer()'s session-lock release across all four exits
Merged
#573 fleetd #567: assert inspect closes probe channel
Merged
#570 fleetd #561: harden the completion/session TurnListener fan-out
Merged
#569 fleetd #551: record the delivery attempt before the irreversible send
Merged
#565 fleetd #555: lift 8 main-flow decisions into tested predicate/dispatch functions
Merged
#566 fleetd #556: make turn registration structural, independent of any TurnListener
Merged
#564 fleetd #513: fix TIMED_OUT_QUEUED javadoc — cancelled, not queued
Merged
#558 CLAUDE.md: the ticket is the pull channel, and the brief is write-once
Merged
#559 fleetd #544: progress watchdog for StatusPoller and SessionReaper loops
Merged
#560 fleetd #552: warn instead of aborting when the post-restart mktemp fails
Merged
#557 fleetd #553: register the rendezvous waiter in onStatus's finally backstop
Merged
#554 fleetd #550: portable hash256 helper for jar_id, Linux CI job for the shell suite
Merged
#549 fleetd #546: widen Injector's delivery catch to Throwable
Merged
#548 fleetd #545: fix mktemp -t templates for GNU coreutils, split unclear-supervisor detail
Merged
#547 fleetd #512 follow-up: record why the drain-complete line must not move into a finally
Merged
#541 fleetd #504 item 1: stop the false ok on the loaded-but-not-running stop path
Merged
#543 fleetd #538: recover polling loops after errors
Merged
#542 fleetd #426: pin FleetHealthMonitor.coverage and its HealthCoverageSource call site
Merged
#539 fleetd #459: lint Javadoc references in CI
Merged
#540 fleetd #537: pin CapturedLog.close()'s appender-detach and setLevel-immunity contracts
Merged
#536 fleetd #535: convert FleetdLeadMailboxSelectionTest to CapturedLog
Merged
#534 fleetd #512 part 2: detect a died shutdown drain the ERROR count is blind to
Merged
#533 fleetd #529: promote CapturedLog to a shared test helper, close the logger-level leak
Merged
#532 fleetd #528: pin drain_gate_refusal's call site, not just the predicate
Merged
#531 charter: separate the blocked forge MCP server from the working GITEA_TOKEN
Merged
#527 fleetd #525: restore the logger level, not just the appender, in SessionManagerTest
Merged
#524 fleetd #518: FleetMcp caller resolution is an explicit choice, and tested for real
Merged
#526 fleetd #521: extract should_swap so the swap guard can't be silently disabled
Merged
#523 fleetd #519: test policy probe guards
Merged
#522 fleetd #512 (part 1): log a positive completion line when drainAll finishes
Merged
#520 fleetd #517: pin the drain-gate abort branch and jar_id absent case
Merged
#516 fleetd #500: honest refusal when the policy parse fails, not just when it's empty
Merged
#515 fleetd #509: pin the pane-scan completeness fold; harden legacyPrincipal
Merged
#514 fleetd #511: fix wrong --no-build wording in drain-gate abort, pin jar_id() default
Merged
#510 fleetd #493: never build into the path a running daemon holds
Merged
#508 fleetd #505: refuse (not promote) a caller whose pane scan errored
Merged
#503 fleetd #501: readiness-grace expiry logs measured elapsed time and poll counter
Merged
#499 fleetd #492 follow-up: detect_supervisor must never read "could not tell" as "none"
Merged
#502 fleetd #498: awaitHerdr distinguishes deadline-passed from interrupted, with measured elapsed time
Merged
#496 fleetd #494: log measured elapsed time, not the configured budget, on lead-rollover failure/success
Merged
#490 fleetd #489: nudge the /clear submit keystroke before bootstrapText
Merged
#485 fleetd #480 Unit C: fleet_handover MCP tool
Merged
#484 fleetd #480 Unit E: BLOCKED is not a settled turn boundary
Merged
#483 fleetd #480 Unit A: lead rollover core (config block + executor)
Merged
#481 fleetd #480 Unit B: add handover skill for lead handoff
Merged
#461 #455: document herdr test socket exception
Merged
#456 fleetd #453: document PeerLauncher.defaultProfileFor/place override obligation
Merged
#452 fleetd #449: fix stale herdr protocol 14, diagnose AgentControlContractTest timing race, CI selects contract tests by tag
Merged
#448 fleetd #437: fleet_ack errors instead of claiming success on a miss
Merged
#451 fleetd #450: make PeerLauncher.spawn(SpawnRequest, PlacementDecision) abstract
Merged
#447 fleetd #444: test that pins the place()-to-spawn() window PlacementDecision closes
Merged
#445 fleetd #442: pin startup report calls
Merged
#443 fleetd #440: derive coordinator.heldDurable from queue durability + ack mode
Merged
#433 fleetd #425 rework: resolve acquireWithWorktree via real placement routing
Merged
#438 fleetd #421: let a lead peek its own held peer mail, primary-only
Merged
#436 fleetd #435: FixedPlacementPolicy now honors maxLoad
Merged
#434 fleetd #422 follow-up: make the model gate's own state observable
Merged
#432 fleetd #431: pin profileForSlot, isSlot, nameForSlot against a live reload
Merged
#429 fleetd #422: enforce model allow-list on/off at spawn, hot reload
Merged
#423 fleetd #415: split coverage() feature-state wording by pattern fallback semantics
Merged
#419 fleetd #418: barrier the throw-path push-loop test on state decide() reads
Merged
#420 fleetd #416: fleet_list must enumerate the STARTUP profile set
Merged
#414 fleetd #409: deterministic test for the #399 completion-stamp ordering race
Merged
#406 fleetd #404: report armed detection from startup map
Merged
#402 fleetd #391: refuse lead fleet replies
Merged
#405 fleetd #399: fix TTL test race — wait on completion stamp, not Phase.DONE
Merged
#363 #362: make the plugin visible, and fix the drift that made it unusable
Merged
#235 CB-612: suppress unneeded token warnings
Merged
#233 fleetd #176: fail-fast + safe UNKNOWN refinement in the spawn-readiness gate
Merged
#231 fleetd #175: check opencode's actual model against the profile, quarantine on a real mismatch
Merged
#229 fleetd#222: put the claude-code charter file where the member can read it
Merged
#228 fleetd #226: reserve architect slots before launch
Merged
#230 fleetd #224 / #225: share worktreeRoot with the group + fix a CWD-dependent group test helper
Merged
#223 CB-619 / fleetd #123: refuse an architect spawn with no matching slot
Merged
#221 fleetd#219: fix OpenCodeLauncher config/discovery roots under memberHerdrSocket
Merged
#198 #197: measure the async ticket TTL from completion, not from creation
Merged
#195 CB-189: cover every remote, both URLs, and any non-SSH scheme in the credential check
Merged
#196 CB-185: fix two blockers to switching on memberHerdrSocket
Merged
#194 CB-192: fix false credential-gap WARN under allow-list+zsh, split its log guard
Merged
#193 #168: audit current wiki snapshot
Merged
#186 CB-185: route members to separate herdr
Merged
#188 CB-185: fix connection-identity/status/health gaps a second herdr daemon exposes
Merged
#187 #185: refuse an unowned paneId when more than one herdr daemon could own it
188 Issues closed from 3 users
Closed
#621 contextNotice hardcodes "ask the operator" and ignores requireOperatorConfirm, so the knob cannot actually stop the asking
Closed
#615 A herdr throw during a lead roll leaves status(token) reporting IN_PROGRESS forever
Closed
#613 A role with no pool and no charter starts anyway, and nothing at boot says so
Closed
#608 Flaky: MessageServiceTest.anAlreadyCollectedTicketProducesNoNudge orders an async scheduler with Thread.sleep
Closed
#603 redeploy-fleetd.sh reports "no process appeared" on a deploy that fully succeeded
Closed
#582 AmqpReplyInbox / LeadMailbox: the pendingByMsgId half of the dual-map cleanup is unasserted at every error-path site (from the #577 sweep)
Closed
#562 LoopWatchdog.health() is public and nothing reads it — surface it without breaking the two scripts that parse /healthz
Closed
#581 CompletionResolver: 9 sites maintain the CAS-remove invariant, only 2 are asserted (from the #577 sweep)
Closed
#571 MessageService collapses the new ATTEMPTED state into TIMED_OUT_QUEUED, so a possibly-delivered message is reported as one that will never arrive
Closed
#577 Sweep: find every invariant kept at N sites but asserted at fewer than N
Closed
#575 The waiter cleanup in MessageService is maintained at three sites, and one of them is outside any finally
Closed
#567 LeadMailbox.inspect leaks a probe channel if the close is removed, and 15 contract tests stay green while the line runs 3 times
Closed
#551 Injector records the delivery AFTER the irreversible send, so a failure in the response window marks a delivered brief NOT_DELIVERED
Closed
#556 TurnListener implementations are the only thing maintaining the Injector's own invariant
Closed
#555 redeploy-fleetd.sh: the 67-test suite covers the functions and barely touches the main flow
Closed
#513 MessageService's TIMED_OUT_QUEUED javadoc says the message is still queued; the code 12 lines of behaviour later cancels it
Closed
#544 A dead StatusPoller or SessionReaper can now be restarted, but nothing restarts it and nothing notices
Closed
#552 redeploy-fleetd.sh exits non-zero after a SUCCESSFUL restart if the post-restart mktemp fails, reporting a working deployment as a failure
Closed
#553 A throwable from any listener callback in onStatus skips the delivered-future completion, so a caller waits forever on a message that WAS delivered
Closed
#550 shasum is macOS-only: on Linux redeploy-fleetd.sh reports an existing jar as "absent" with exit 0, and the shell suite dies at 127 looking green
Closed
#545 redeploy-fleetd.sh cannot run on Linux at all: every mktemp -t NAME fails under GNU coreutils, and the script blames the systemd bus for it
Closed
#546 Merging #543 made re-delivery reachable: an Error after Injector's send now loops and types the same brief again
Closed
#538 An Error in StatusPoller.loop or SessionReaper.loop kills the thread permanently and silently, and start() then refuses to restart it
Closed
#426 FleetHealthMonitor.coverage has no test at all, and it feeds fleet_list's healthCoverage field
Closed
#459 Five {@link} targets in main do not exist, and nothing in CI would ever say so
Closed
#537 CapturedLog.close() detaches the appender and no test proves it — the fix for #525/#529/#535 is itself unpinned
Closed
#535 FleetdLeadMailboxSelectionTest leaks three ListAppenders onto the shared Fleetd logger, and the appender-balance invariant catches it
Closed
#512 redeploy-fleetd.sh says "no ERROR lines since restart" while blind to the exact failure #493 is about
Closed
#529 Tests pin shared logger levels and never restore them: 9 files, 19 pins, one proven cross-class collision — share SessionManagerTest's CapturedLog
Closed
#528 "Extract the decision so the suite can call it" pins the decision and never the wiring — drain_gate_refusal's call site can be bypassed with the suite green
Closed
#525 A test pins the shared SessionManager logger to WARN and never restores it, so any later test asserting an INFO line silently sees nothing
Closed
#518 Nothing tests that FleetMcp uses the CallerResolver — forcing the legacy identity path leaves the whole suite green at 1696/0
Closed
#521 The swap guard can be disabled with the suite green — a "successful" redeploy that never puts the new jar in place
Closed
#519 probe-member-credentials.sh has no test harness at all, and its new arity message is off by one on an empty parse
Closed
#517 A source-text test pins what a message SAYS and never whether it is REACHED — the drain-gate abort branch can be disabled with the suite green
Closed
#500 probe-member-credentials.sh uses mapfile (bash 4+) with no set -e, so on bash 3.2 it silently reports an empty field list
Closed
#509 #505 follow-up: the two-client completeness fold is unpinned, and legacyPrincipal can still re-open the escalation
Closed
#511 #493 follow-up: the drain-gate abort message names a recovery that does not work, and jar_id()'s default is unpinned
Closed
#493 redeploy-fleetd.sh builds into the LIVE jar path, so the shutdown drain can die on a class it never loaded (#413)
Closed
#505 A herdr error during the pane scan makes a worker's connection resolve as the primary — #317's escalation, through the door #317 did not close
Closed
#492 redeploy-fleetd.sh is launchd-only: on fleet01 it kills the daemon and leaves TWO running
Closed
#501 Injector's readiness-grace warn prints the configured budget as if it were elapsed time — #494's defect, in the delivery loop
Closed
#498 awaitHerdr has two return-false paths and the caller reports both as "did not answer within 30s"
Closed
#494 LeadRollover's failure logs print the CONFIGURED budget, not the measured wait — the exact thing that hid #489
Closed
#474 A charter naming an unregistered tool is refused at startup but accepted on reload
Closed
#399 MessageServiceTest TTL test races the completedNanos stamp — passes on macOS, fails on Linux
Closed
#464 Nothing checks charter text against the tool surface the server registers
Closed
#400 SECURITY (receipt): the scrub reports a name as blanked when eval returned 0 without blanking it
Closed
#393 "skill seeding: N of M" is logged as success for opencode members, which never read .claude/skills/
Closed
#469 Charter text is still never checked against the live tool surface — one canonical tool-name set (follow-up to #464)
Closed
#424 Revoking an architect slot does not revoke it: MemberRegistry freezes fleet.architects, and the reload says it applied
Closed
#463 listFleet's compat overloads default callerIsPrimary to true, so the coordinator gate fails open
Closed
#446 The model gate can be turned off at runtime but its usage-limit detector cannot be turned on at runtime
Closed
#458 Canonical invariant 5 states a mechanism where it means a purpose — restate it (blocked on #455)
Closed
#439 fleet_list hands every worker the coordinator row: peer coord-ids and 80-char previews of lead-to-lead bodies
Closed
#455 Invariant 5 ("never drive the multiplexer directly, no socket") is unsatisfiable for a worker assigned to herdr code
Closed
#453 PeerLauncher.place(role) and defaultProfileFor(role) are default methods that silently ignore the role
Closed
#422 Model gating units 2+3: enforce the allow-list at spawn, and turn one model off at runtime without editing profiles
Closed
#449 The herdr contract tests were left behind by the protocol-19 port, and nothing runs them
Closed
#437 fleet_ack says "acknowledged <msgId>" for a held peer message it never touches
Closed
#450 PeerLauncher's default spawn(req, decision) is still the #425 defect, and nothing tests it — make it abstract
Closed
#444 PlacementDecision's javadoc claims the place()-to-spawn() window is closed, but no test pins it — and PeerLauncher's default spawn(req, decision) reopens it
Closed
#442 Deleting any of the four startup report calls from Fleetd.main leaves the suite green
Closed
#440 coordinatorView reports heldDurable as a hardcoded true, so it cannot go false when durability breaks
Closed
#425 fleet_profiles reports a frozen "default" profile while placement already moved to a new one
Closed
#441 Nothing warns that a profile has no usage-limit detection, so "off when the subscription limit is reached" silently covers only the profiles that opted in
Closed
#421 A lead can neither read nor safely keep its own held peer messages — the only drain is an injectable pane
Closed
#435 maxLoad says it is excluded from "every automatic policy" — fixed, the default, ignores it
Closed
#415 The startup coverage line reports "off" for errorPattern, but legacy classification is still running
Closed
#417 Sweep for the REVERSE config mismatch: a boot snapshot read where the key is hot
Closed
#418 Second instance of the #399 shape: a test uses Phase.ASKING as a barrier for push-loop state published later
Closed
#416 fleet_list advertises free capacity for a profile the spawn gate refuses — CapacitySource reads the live profile keySet
Closed
#409 The #399 completion-stamp race is verifiable deterministically — nowNanos is already injectable
Closed
#404 exhaustionDetectionArmed reads the LIVE config while detection reads the STARTUP snapshot — after a reload it reports armed when detection is still off
Closed
#391 fleet_reply cannot answer a peer lead — the charter documents a path that does not exist
Closed
#395 A profile with no exhaustedPattern has usage-limit detection silently OFF — 6 of 8 live profiles, including both Claude subscription ones
Closed
#394 SECURITY: the credential scrub aborts silently at the first read-only parameter, leaving every later name unscrubbed
Closed
#190 Process argv is a credential channel we never enumerated: 4 secrets are readable by any local user via ps
Closed
#386 Every fleetd timer freezes while the macOS host sleeps: System.nanoTime() stops, so the stall detector missed a member that was BUSY for 101 minutes
Closed
#384 "left no scrub report" has a third, benign cause it does not name — a group-shared ZDOTDIR cannot receive the receipt
Closed
#382 SpawnRequest has the arity trap's two halves, but not yet a colliding arity — CompositePeerLauncher:372 will drop the next component added
Closed
#381 Members cannot run shell commands: the classifier refuses every one, so they cannot build, test or open a PR (3 members, 2 profiles)
Closed
#385 A lead-coordination message can be delivered forever: ack() returns success when connection recovery cleared the held entry
Closed
#377 GITEA_HOST carries a scheme and a trailing slash, and every member builds URLs from it
Closed
#365 Three more tool results claim "delivered" for something weaker than delivery
Closed
#373 The hermetic env on production GitWorktrees instances is unpinned — stripping it stays green, poisoned or not
Closed
#358 Two more records can silently drop a field the way FleetConfig could (MemberSession, FleetConfig.Profile)
Closed
#376 A fast model trips the 2000ms crash floor: a real answer is reported as a failed turn
Closed
#352 CompositePeerLauncher.clearContext silently stops working for every peer that survives a restart
Closed
#375 The idle-sleep guard is a no-op on Linux, so fleet01 still sleeps under a live member
Closed
#369 GitWorktreesTest reads the machine's real git config, so it can pass or fail for reasons outside the repo
Closed
#368 A lead's delegation binding outlives the lead, so reply nudges keep going to a dead terminal
Closed
#360 deploy/fleetd.service silently disables caller identity and the credential scrub
Closed
#359 Dead lead tabs accumulate and make lead-to-lead coordination permanently undeliverable
Closed
#362 The plugin shipped in CB-527 is invisible and has drifted — and it cannot deliver worker skills
Closed
#361 Lead-to-lead coordination has a send half with tools and a receive half without
Closed
#334 MessageService: an async ticket is still stranded between clearAsyncQuestion and closeAsk
Closed
#348 The same text-match-drives-a-cooldown shape exists for ExhaustionSink, where the penalty is 30x longer
Closed
#335 MessageService: three more places a thrown exception has nowhere to go
Closed
#342 CompositePeerLauncher's single-daemon stop() shortcut can skip tab teardown, orphaning a worker tab
Closed
#345 MessageService's use of Cancellation.DELIVERED is unpinned — the seam is tested, the caller is not
Closed
#337 The deferred key set still has no reporting coverage — measured, a dropped comparison is invisible
Closed
#341 One flag guards two different memberCredentials WARNs, so a later leaking credential name is never reported
Closed
#339 A member's normal report that quotes "API Error:" is recorded as a real backend outage
Closed
#338 A TIMED_OUT_QUEUED send leaves its message in the injector queue, and it is delivered later as a stale brief
Closed
#333 fleet: is a third split key, sitting in the coverage checker's own escape hatch — and SPLIT_KEYS membership does not prove any reporting code exists
Closed
#329 sendAsync's executor catch swallows every exception thrown after the future completes, and that is why a stranded async ticket is invisible
Closed
#330 health: and coordinator: are split keys — add a fourth reload class that says so, then make the top-level list prove its own coverage
Closed
#326 The top-level deferred-key list has the same drift as #323, and two keys are read both from the startup snapshot and live — no single classification is right for them
Closed
#324 answer() mutates a Task under sessionLocks while ask()'s timeout path mutates the same Task under no lock — finishAsyncTask can throw NPE on a null turnId
Closed
#323 The reload classifier says "compares every component the launcher reads at spawn" and misses four keys, so those reloads report success and do nothing
Closed
#318 A reply delivered while release() is running is stranded unacked forever: #298 closed "already held", not "arrives during release"
Closed
#316 The dirty-worktree check runs before the worker is stopped, so work written during teardown is deleted with no preserve and no snapshot
Closed
#317 A failed peer-PID lookup promotes a worker to primary: the escalation PaneLocator's own javadoc names, through a trigger CB-161 did not close
Closed
#315 The fixed placement policy ignores the retry loop's unreachable set, so an unqualified spawn retries the same dead profile and never tries the healthy one
Closed
#307 A worker's report is stranded when its fleet_ask timed out, and the lead is told the worker never replied
Closed
#309 A timed-out git worktree add leaks the worktree it half-created: #274's cleanup does not cover the add itself
Closed
#308 A worker spawned during the shutdown drain is never drained: its pane is left running and its worktree is never preserved
Closed
#310 reapIdle can release a session that just became BUSY, contradicting its own documented invariant
Closed
#306 A worker that wedges in unknown during post-turn housekeeping is polled forever and can never be delivered to again
Closed
#305 A worker connecting from 127.0.0.2 is resolved as the primary, and gains spawn/stop/send/drain
Closed
#304 POST /members and DELETE /members/{paneId} turn a herdr failure into a bare 500, and a stop that already succeeded can never report success
Closed
#302 A REST reply with no content field silently becomes an empty reply, and the lead cannot tell it from a member that said nothing
Closed
#298 AmqpReplyInbox.release() drops an unacked reply that the broker still holds — a worker's report is lost silently
Closed
#297 REST is the weaker door, and it is the door a lead falls back to when MCP drops
Closed
#296 Two spawn-side exits leak a live pane, and the caller cannot clean up because it never learns the pane id
Closed
#293 A failing tab.close during teardown leaks both the ZDOTDIR and the worktree
Closed
#280 Follow-up to #275: a health-detected GONE member whose ask lapses may still strand its ticket
Closed
#290 Restore a test for reapIdle's per-session guard, which #283 left uncovered
Closed
#285 seedTrustDialog writes the trust seed to fleetd's own home when memberHerdrSocket is set, so the member never sees it
Closed
#284 A BACKEND_ERROR member holds its seat forever: counted as live, never reaped, never reclaimable
Closed
#283 Two teardown leaks in SessionManager: an unguarded worktree removal, and a branch the spawn-failure catch forgets
Closed
#282 A second fleet_ask in the same resumed turn kills its own async ticket: answer() completes it with QUESTION
Closed
#281 Pin the Authz.Action every handler chooses, so the next #272 fails a test instead of shipping
Closed
#275 Investigate: abandon() skips a task in ASKING, so a torn-down member mid-fleet_ask may strand its ticket forever
Closed
#274 GitWorktrees.add() leaks the worktree and branch when a step after git worktree add throws
Closed
#273 A malformed exhaustedPattern crashes startup unnamed — only errorPattern is validated at load
Closed
#272 fleet_poll{target} drains another session's inbox behind a READ gate — authorization fails open
Closed
#276 #269 follow-up: the "allowed N of M" INFO line still claims to describe the member's pane
Closed
#267 The #175 model-mismatch check never runs for opencode spawns without a provisioned worktree — the majority of them
Closed
#266 memberCredentials.sshAuthSock: rename the values block/allow to omit/inherit — "block" names something fleetd does not do
Closed
#103 CB-605: the systemd unit has launchd's login-shell secret gap, and does not mention the two tokens it drops
Closed
#257 fleet_list's free and the spawn gate now disagree by one on subscription profiles
Closed
#247 Seeding ~/.claude.json can still lose an external writer's change
Closed
#141 CB-631: four credential env vars reach every member pane unblocked — the CB-596 policy has drifted
Closed
#116 CB-613: one worker branch holds a teardown test that never reached main, and 55 stale branches hide it
Closed
#252 The REST surface is a supported operator fallback and is documented nowhere — 14 routes, 0 entries
Closed
#258 Two test fixtures write to the operator's real ~/.claude.json on every run (116 dead entries accumulated)
Closed
#114 CB-609: docs/MCP-Contract.md is a pre-build design doc that CLAUDE.md still points every session at
Closed
#113 CB-611: our checkers keep covering less than they look like they cover — three instances in one day
Closed
#111 CB-608: the credential probe hardcodes its own name list, so it silently under-reports the policy
Closed
#155 the member credential scrub is zsh-only: a member on any other login shell gets every credential
Closed
#176 capacity counts panes, not subscription seats, so a subscription profile reports a free slot that cannot be filled
Closed
#249 fleet_list can report another session's agentSessionId, and fleet_spawn{resumeSessionId} will resume it
Closed
#115 CB-612: the daemon warns at every boot about a token no profile needs, teaching operators to skim startup warnings
Closed
#149 A claude-code member spawned into a fresh worktree blocks forever on the workspace trust dialog
Closed
#248 Fleetd's composition root is untested: a feature can be silently unwired and every test stays green
Closed
#148 CB-634: the parity overlay copies .env/.envrc into worker worktrees — the same shape as CB-633
Closed
#234 #175's model read-back fires a FALSE POSITIVE on any member spawned without a worktree — and the quarantine it announces never happens
Closed
#227 A backend outage is invisible at fleet level: repeated backend errors are never correlated, never cool off a credential, and never reach the lead
Closed
#201 Backend-error classification needs a real mechanism, not hard-coded strings (#164 points 3 and 4)
Closed
#241 Completion fallback can hand the lead back its own brief as the member's "report"
Closed
#134 CB-628: a tracked file neutralized by the worktree overlay can never be edited by a worker, and nothing says so
Closed
#232 autoCompactWindow is set by flag and never read back — the #175 shape in a second place
Closed
#175 fleetd asks opencode for a model and never checks it got that model, so a withdrawn name silently bills a paid credential
Closed
#226 A member that loses the architect slot race still runs on the architect charter while the gate treats it as a worker
Closed
#222 ClaudeCodeLauncher writes the member's charter file into fleetd's own java.io.tmpdir — the #219 site-1 shape, third instance
Closed
#224 worktreeRoot itself is never made group-traversable, so a restrictive umask breaks every member under memberHerdrSocket
Closed
#225 Two memberHerdrSocket tests pass or fail depending on where the repo is checked out
Closed
#123 CB-619: a spawn asking for a role its profile has no slot for is silently demoted to worker, and the roster still says otherwise
Closed
#219 OpenCodeLauncher decides a member's config and discovery roots from fleetd's own filesystem — the #213 defect, twice
Closed
#140 CB-630: every subscription claude-code profile fails to spawn — the pane exits and the error names a symptom
Closed
#213 #185: the ZDOTDIR scrub decides on fleetd's own shell and writes to fleetd's own TMPDIR, so under memberHerdrSocket it silently protects nothing
Closed
#211 A backend-exhausted pane with no ⏺ marker is classified as an empty scrape, so the credential is never quarantined
Closed
#214 A plain claude-code spawn mints no session id, so the most common member on the fleet can never be resumed
Closed
#220 Launch command is typed into the pane and silently cut at 1024 bytes
Closed
#209 agentSessionId is resolved once at spawn and frozen, so an opencode member never reports one
Closed
#199 GET /members returns its rows under a "workers" key, so a caller reading "members" sees an empty fleet
Closed
#206 opencode session discovery never finds a member's record, so agentSessionId is always null and resume silently does nothing
Closed
#137 CB-629: after a bridge_ask round-trip, the worker's final reply orphans its ticket, which then reports a false failure
Closed
#172 LAVINMQ_URI reaches every member pane with the broker password inline — on neither credential list, and not credential-shaped
Closed
#161 a pane member's grandchild process may be resolved as primary — possible worker→primary escalation
Closed
#126 CB-622: rename the MCP tool namespace bridge_* -> fleet_*, with one dual-name release
Closed
#164 a member whose backend 400s dies in one second, and fleetd returns the lost turn as a successful empty reply
Closed
#168 CB-638: the wiki describes a system that no longer exists — audit index of pages to revise, build and retire
Closed
#197 Async ticket TTL runs from creation, so a long task's report is destroyed on arrival
Closed
#192 CB-633 follow-up: the credential-gap WARN claims names are inherited unblocked when the scrub does blank them, and one guard can hide the real report
Closed
#189 CB-157 follow-up: the remote-URL credential checks only see origin, only https, and never a push URL
218 Issues created by 2 users
Opened
#189 CB-157 follow-up: the remote-URL credential checks only see origin, only https, and never a push URL
Opened
#190 Process argv is a credential channel we never enumerated: 4 secrets are readable by any local user via ps
Opened
#192 CB-633 follow-up: the credential-gap WARN claims names are inherited unblocked when the scrub does blank them, and one guard can hide the real report
Opened
#197 Async ticket TTL runs from creation, so a long task's report is destroyed on arrival
Opened
#199 GET /members returns its rows under a "workers" key, so a caller reading "members" sees an empty fleet
Opened
#201 Backend-error classification needs a real mechanism, not hard-coded strings (#164 points 3 and 4)
Opened
#206 opencode session discovery never finds a member's record, so agentSessionId is always null and resume silently does nothing
Opened
#209 agentSessionId is resolved once at spawn and frozen, so an opencode member never reports one
Opened
#211 A backend-exhausted pane with no ⏺ marker is classified as an empty scrape, so the credential is never quarantined
Opened
#213 #185: the ZDOTDIR scrub decides on fleetd's own shell and writes to fleetd's own TMPDIR, so under memberHerdrSocket it silently protects nothing
Opened
#214 A plain claude-code spawn mints no session id, so the most common member on the fleet can never be resumed
Opened
#219 OpenCodeLauncher decides a member's config and discovery roots from fleetd's own filesystem — the #213 defect, twice
Opened
#220 Launch command is typed into the pane and silently cut at 1024 bytes
Opened
#222 ClaudeCodeLauncher writes the member's charter file into fleetd's own java.io.tmpdir — the #219 site-1 shape, third instance
Opened
#224 worktreeRoot itself is never made group-traversable, so a restrictive umask breaks every member under memberHerdrSocket
Opened
#225 Two memberHerdrSocket tests pass or fail depending on where the repo is checked out
Opened
#226 A member that loses the architect slot race still runs on the architect charter while the gate treats it as a worker
Opened
#227 A backend outage is invisible at fleet level: repeated backend errors are never correlated, never cool off a credential, and never reach the lead
Opened
#232 autoCompactWindow is set by flag and never read back — the #175 shape in a second place
Opened
#234 #175's model read-back fires a FALSE POSITIVE on any member spawned without a worktree — and the quarantine it announces never happens
Opened
#241 Completion fallback can hand the lead back its own brief as the member's "report"
Opened
#247 Seeding ~/.claude.json can still lose an external writer's change
Opened
#248 Fleetd's composition root is untested: a feature can be silently unwired and every test stays green
Opened
#249 fleet_list can report another session's agentSessionId, and fleet_spawn{resumeSessionId} will resume it
Opened
#252 The REST surface is a supported operator fallback and is documented nowhere — 14 routes, 0 entries
Opened
#257 fleet_list's free and the spawn gate now disagree by one on subscription profiles
Opened
#258 Two test fixtures write to the operator's real ~/.claude.json on every run (116 dead entries accumulated)
Opened
#266 memberCredentials.sshAuthSock: rename the values block/allow to omit/inherit — "block" names something fleetd does not do
Opened
#267 The #175 model-mismatch check never runs for opencode spawns without a provisioned worktree — the majority of them
Opened
#272 fleet_poll{target} drains another session's inbox behind a READ gate — authorization fails open
Opened
#273 A malformed exhaustedPattern crashes startup unnamed — only errorPattern is validated at load
Opened
#274 GitWorktrees.add() leaks the worktree and branch when a step after git worktree add throws
Opened
#275 Investigate: abandon() skips a task in ASKING, so a torn-down member mid-fleet_ask may strand its ticket forever
Opened
#276 #269 follow-up: the "allowed N of M" INFO line still claims to describe the member's pane
Opened
#280 Follow-up to #275: a health-detected GONE member whose ask lapses may still strand its ticket
Opened
#281 Pin the Authz.Action every handler chooses, so the next #272 fails a test instead of shipping
Opened
#282 A second fleet_ask in the same resumed turn kills its own async ticket: answer() completes it with QUESTION
Opened
#283 Two teardown leaks in SessionManager: an unguarded worktree removal, and a branch the spawn-failure catch forgets
Opened
#284 A BACKEND_ERROR member holds its seat forever: counted as live, never reaped, never reclaimable
Opened
#285 seedTrustDialog writes the trust seed to fleetd's own home when memberHerdrSocket is set, so the member never sees it
Opened
#290 Restore a test for reapIdle's per-session guard, which #283 left uncovered
Opened
#293 A failing tab.close during teardown leaks both the ZDOTDIR and the worktree
Opened
#296 Two spawn-side exits leak a live pane, and the caller cannot clean up because it never learns the pane id
Opened
#297 REST is the weaker door, and it is the door a lead falls back to when MCP drops
Opened
#298 AmqpReplyInbox.release() drops an unacked reply that the broker still holds — a worker's report is lost silently
Opened
#302 A REST reply with no content field silently becomes an empty reply, and the lead cannot tell it from a member that said nothing
Opened
#304 POST /members and DELETE /members/{paneId} turn a herdr failure into a bare 500, and a stop that already succeeded can never report success
Opened
#305 A worker connecting from 127.0.0.2 is resolved as the primary, and gains spawn/stop/send/drain
Opened
#306 A worker that wedges in unknown during post-turn housekeeping is polled forever and can never be delivered to again
Opened
#307 A worker's report is stranded when its fleet_ask timed out, and the lead is told the worker never replied
Opened
#308 A worker spawned during the shutdown drain is never drained: its pane is left running and its worktree is never preserved
Opened
#309 A timed-out git worktree add leaks the worktree it half-created: #274's cleanup does not cover the add itself
Opened
#310 reapIdle can release a session that just became BUSY, contradicting its own documented invariant
Opened
#315 The fixed placement policy ignores the retry loop's unreachable set, so an unqualified spawn retries the same dead profile and never tries the healthy one
Opened
#316 The dirty-worktree check runs before the worker is stopped, so work written during teardown is deleted with no preserve and no snapshot
Opened
#317 A failed peer-PID lookup promotes a worker to primary: the escalation PaneLocator's own javadoc names, through a trigger CB-161 did not close
Opened
#318 A reply delivered while release() is running is stranded unacked forever: #298 closed "already held", not "arrives during release"
Opened
#323 The reload classifier says "compares every component the launcher reads at spawn" and misses four keys, so those reloads report success and do nothing
Opened
#324 answer() mutates a Task under sessionLocks while ask()'s timeout path mutates the same Task under no lock — finishAsyncTask can throw NPE on a null turnId
Opened
#326 The top-level deferred-key list has the same drift as #323, and two keys are read both from the startup snapshot and live — no single classification is right for them
Opened
#329 sendAsync's executor catch swallows every exception thrown after the future completes, and that is why a stranded async ticket is invisible
Opened
#330 health: and coordinator: are split keys — add a fourth reload class that says so, then make the top-level list prove its own coverage
Opened
#333 fleet: is a third split key, sitting in the coverage checker's own escape hatch — and SPLIT_KEYS membership does not prove any reporting code exists
Opened
#334 MessageService: an async ticket is still stranded between clearAsyncQuestion and closeAsk
Opened
#335 MessageService: three more places a thrown exception has nowhere to go
Opened
#337 The deferred key set still has no reporting coverage — measured, a dropped comparison is invisible
Opened
#338 A TIMED_OUT_QUEUED send leaves its message in the injector queue, and it is delivered later as a stale brief
Opened
#339 A member's normal report that quotes "API Error:" is recorded as a real backend outage
Opened
#340 classify() can call an active pane IDLE — measured real, trigger never seen (half A closed)
Opened
#341 One flag guards two different memberCredentials WARNs, so a later leaking credential name is never reported
Opened
#342 CompositePeerLauncher's single-daemon stop() shortcut can skip tab teardown, orphaning a worker tab
Opened
#345 MessageService's use of Cancellation.DELIVERED is unpinned — the seam is tested, the caller is not
Opened
#348 The same text-match-drives-a-cooldown shape exists for ExhaustionSink, where the penalty is 30x longer
Opened
#352 CompositePeerLauncher.clearContext silently stops working for every peer that survives a restart
Opened
#358 Two more records can silently drop a field the way FleetConfig could (MemberSession, FleetConfig.Profile)
Opened
#359 Dead lead tabs accumulate and make lead-to-lead coordination permanently undeliverable
Opened
#360 deploy/fleetd.service silently disables caller identity and the credential scrub
Opened
#361 Lead-to-lead coordination has a send half with tools and a receive half without
Opened
#362 The plugin shipped in CB-527 is invisible and has drifted — and it cannot deliver worker skills
Opened
#365 Three more tool results claim "delivered" for something weaker than delivery
Opened
#368 A lead's delegation binding outlives the lead, so reply nudges keep going to a dead terminal
Opened
#369 GitWorktreesTest reads the machine's real git config, so it can pass or fail for reasons outside the repo
Opened
#373 The hermetic env on production GitWorktrees instances is unpinned — stripping it stays green, poisoned or not
Opened
#375 The idle-sleep guard is a no-op on Linux, so fleet01 still sleeps under a live member
Opened
#376 A fast model trips the 2000ms crash floor: a real answer is reported as a failed turn
Opened
#377 GITEA_HOST carries a scheme and a trailing slash, and every member builds URLs from it
Opened
#381 Members cannot run shell commands: the classifier refuses every one, so they cannot build, test or open a PR (3 members, 2 profiles)
Opened
#382 SpawnRequest has the arity trap's two halves, but not yet a colliding arity — CompositePeerLauncher:372 will drop the next component added
Opened
#383 FleetHealthMonitor detects TURN_BOUNDARY_LOST but no API surface exposes it — a stalled member reads as healthy on every field a lead can see
Opened
#384 "left no scrub report" has a third, benign cause it does not name — a group-shared ZDOTDIR cannot receive the receipt
Opened
#385 A lead-coordination message can be delivered forever: ack() returns success when connection recovery cleared the held entry
Opened
#386 Every fleetd timer freezes while the macOS host sleeps: System.nanoTime() stops, so the stall detector missed a member that was BUSY for 101 minutes
Opened
#388 The credential scrub does not run in a shell that is neither login nor interactive — .zshenv is the only file zsh always reads, and it is the one file the scrub is not in
Opened
#391 fleet_reply cannot answer a peer lead — the charter documents a path that does not exist
Opened
#392 notifications.mode: webhook reports healthCoverage "full" and sends nothing — the knob relabels the absence of a sink
Opened
#393 "skill seeding: N of M" is logged as success for opencode members, which never read .claude/skills/
Opened
#394 SECURITY: the credential scrub aborts silently at the first read-only parameter, leaving every later name unscrubbed
Opened
#395 A profile with no exhaustedPattern has usage-limit detection silently OFF — 6 of 8 live profiles, including both Claude subscription ones
Opened
#399 MessageServiceTest TTL test races the completedNanos stamp — passes on macOS, fails on Linux
Opened
#400 SECURITY (receipt): the scrub reports a name as blanked when eval returned 0 without blanking it
Opened
#404 exhaustionDetectionArmed reads the LIVE config while detection reads the STARTUP snapshot — after a reload it reports armed when detection is still off
Opened
#407 Five log-only startup reporters in Fleetd.main have unpinned call sites — deleting any one ships a green build
Opened
#408 Five places trust a command's exit code as proof of its effect and never read the value back — the #400 shape outside the scrub
Opened
#409 The #399 completion-stamp race is verifiable deterministically — nowNanos is already injectable
Opened
#410 liveStatus reported "working" for a member that did nothing for 76 minutes — the health snapshot has no progress signal
Opened
#411 Warn at load when two profiles share a tokenEnv but quarantine under different credential ids
Opened
#412 A daemon-side crash in the completion thread is silently misattributed to the worker
Opened
#413 Building in the fleetd tree disarms the running daemon, and nothing warns at either end
Opened
#415 The startup coverage line reports "off" for errorPattern, but legacy classification is still running
Opened
#416 fleet_list advertises free capacity for a profile the spawn gate refuses — CapacitySource reads the live profile keySet
Opened
#417 Sweep for the REVERSE config mismatch: a boot snapshot read where the key is hot
Opened
#418 Second instance of the #399 shape: a test uses Phase.ASKING as a barrier for push-loop state published later
Opened
#421 A lead can neither read nor safely keep its own held peer messages — the only drain is an injectable pane
Opened
#422 Model gating units 2+3: enforce the allow-list at spawn, and turn one model off at runtime without editing profiles
Opened
#424 Revoking an architect slot does not revoke it: MemberRegistry freezes fleet.architects, and the reload says it applied
Opened
#425 fleet_profiles reports a frozen "default" profile while placement already moved to a new one
Opened
#426 FleetHealthMonitor.coverage has no test at all, and it feeds fleet_list's healthCoverage field
Opened
#427 A config key's hot/deferred class is a claim about every consumer, and nothing checks it that way
Opened
#431 MemberRegistry: profileForSlot, nameForSlot and isSlot were made live with no test pinning any of them
Opened
#435 maxLoad says it is excluded from "every automatic policy" — fixed, the default, ignores it
Opened
#437 fleet_ack says "acknowledged <msgId>" for a held peer message it never touches
Opened
#439 fleet_list hands every worker the coordinator row: peer coord-ids and 80-char previews of lead-to-lead bodies
Opened
#440 coordinatorView reports heldDurable as a hardcoded true, so it cannot go false when durability breaks
Opened
#441 Nothing warns that a profile has no usage-limit detection, so "off when the subscription limit is reached" silently covers only the profiles that opted in
Opened
#442 Deleting any of the four startup report calls from Fleetd.main leaves the suite green
Opened
#444 PlacementDecision's javadoc claims the place()-to-spawn() window is closed, but no test pins it — and PeerLauncher's default spawn(req, decision) reopens it
Opened
#446 The model gate can be turned off at runtime but its usage-limit detector cannot be turned on at runtime
Opened
#449 The herdr contract tests were left behind by the protocol-19 port, and nothing runs them
Opened
#450 PeerLauncher's default spawn(req, decision) is still the #425 defect, and nothing tests it — make it abstract
Opened
#453 PeerLauncher.place(role) and defaultProfileFor(role) are default methods that silently ignore the role
Opened
#454 The herdr protocol number has no single home — 21 sites, no constant, and the fake and reality can disagree silently
Opened
#455 Invariant 5 ("never drive the multiplexer directly, no socket") is unsatisfiable for a worker assigned to herdr code
Opened
#458 Canonical invariant 5 states a mechanism where it means a purpose — restate it (blocked on #455)
Opened
#459 Five {@link} targets in main do not exist, and nothing in CI would ever say so
Opened
#460 The /mcp transport has no test: 57 endpoint tests all enter through REST, the door no agent uses
Opened
#463 listFleet's compat overloads default callerIsPrimary to true, so the coordinator gate fails open
Opened
#464 Nothing checks charter text against the tool surface the server registers
Opened
#466 Exhaustion quarantine is a flat 30 minutes, so a weekly subscription limit is retried ~336 times
Opened
#469 Charter text is still never checked against the live tool surface — one canonical tool-name set (follow-up to #464)
Opened
#474 A charter naming an unregistered tool is refused at startup but accepted on reload
Opened
#477 MessageServiceTest.anAlreadyCollectedTicketProducesNoNudge races the push loop's 300ms tick — reproduced on both hosts, under load
Opened
#478 A directory in parityOverlay is reported as copied and arrives empty — Files.copy does not recurse
Opened
#479 Exhaustion detection is armed on 2 of 8 profiles, and the default profile is not one of them
Opened
#480 Lead session rollover: the lead asks for a handover, fleetd verifies it, clears the pane and boots the next lead
Opened
#482 A context reset changes the peer's session id, and fleetd keeps reporting the dead one — so resumeSessionId would resume the conversation clearAfterTurn just discarded
Opened
#486 LeadRollover's settle poll has no bound of its own, so a regression shows up as a CI hang instead of a red test
Opened
#488 fleet.leaders.<name>.cwd is never validated — a relative value is silently accepted
Opened
#489 #480 lead rollover: /clear and bootstrapText concatenate into one line — the roll never bootstraps
Opened
#491 #480 a relative handoverPath lands in the LEAD'S repo, which is usually not the repo that ignores it
Opened
#492 redeploy-fleetd.sh is launchd-only: on fleet01 it kills the daemon and leaves TWO running
Opened
#493 redeploy-fleetd.sh builds into the LIVE jar path, so the shutdown drain can die on a class it never loaded (#413)
Opened
#494 LeadRollover's failure logs print the CONFIGURED budget, not the measured wait — the exact thing that hid #489
Opened
#497 A sentinel that conflates "measured: no" with "could not measure" — three instances, three subsystems
Opened
#498 awaitHerdr has two return-false paths and the caller reports both as "did not answer within 30s"
Opened
#500 probe-member-credentials.sh uses mapfile (bash 4+) with no set -e, so on bash 3.2 it silently reports an empty field list
Opened
#501 Injector's readiness-grace warn prints the configured budget as if it were elapsed time — #494's defect, in the delivery loop
Opened
#504 redeploy-fleetd.sh: four places where a failed command is reported as a clean result
Opened
#505 A herdr error during the pane scan makes a worker's connection resolve as the primary — #317's escalation, through the door #317 did not close
Opened
#506 Tests that stay green when the behaviour is absent: four sites, three mechanisms, one cheap discriminator
Opened
#507 MessageServiceTest's coalescing tests sleep because the only observation available collects the ticket — the instrument is missing, not the test
Opened
#509 #505 follow-up: the two-client completeness fold is unpinned, and legacyPrincipal can still re-open the escalation
Opened
#511 #493 follow-up: the drain-gate abort message names a recovery that does not work, and jar_id()'s default is unpinned
Opened
#512 redeploy-fleetd.sh says "no ERROR lines since restart" while blind to the exact failure #493 is about
Opened
#513 MessageService's TIMED_OUT_QUEUED javadoc says the message is still queued; the code 12 lines of behaviour later cancels it
Opened
#517 A source-text test pins what a message SAYS and never whether it is REACHED — the drain-gate abort branch can be disabled with the suite green
Opened
#518 Nothing tests that FleetMcp uses the CallerResolver — forcing the legacy identity path leaves the whole suite green at 1696/0
Opened
#519 probe-member-credentials.sh has no test harness at all, and its new arity message is off by one on an empty parse
Opened
#521 The swap guard can be disabled with the suite green — a "successful" redeploy that never puts the new jar in place
Opened
#525 A test pins the shared SessionManager logger to WARN and never restores it, so any later test asserting an INFO line silently sees nothing
Opened
#528 "Extract the decision so the suite can call it" pins the decision and never the wiring — drain_gate_refusal's call site can be bypassed with the suite green
Opened
#529 Tests pin shared logger levels and never restore them: 9 files, 19 pins, one proven cross-class collision — share SessionManagerTest's CapturedLog
Opened
#530 The idle reaper destroys a member's pane, worktree, ticket and report and logs nothing at info — while the refs sweep beside it reports its count
Opened
#535 FleetdLeadMailboxSelectionTest leaks three ListAppenders onto the shared Fleetd logger, and the appender-balance invariant catches it
Opened
#537 CapturedLog.close() detaches the appender and no test proves it — the fix for #525/#529/#535 is itself unpinned
Opened
#538 An Error in StatusPoller.loop or SessionReaper.loop kills the thread permanently and silently, and start() then refuses to restart it
Opened
#544 A dead StatusPoller or SessionReaper can now be restarted, but nothing restarts it and nothing notices
Opened
#545 redeploy-fleetd.sh cannot run on Linux at all: every mktemp -t NAME fails under GNU coreutils, and the script blames the systemd bus for it
Opened
#546 Merging #543 made re-delivery reachable: an Error after Injector's send now loops and types the same brief again
Opened
#550 shasum is macOS-only: on Linux redeploy-fleetd.sh reports an existing jar as "absent" with exit 0, and the shell suite dies at 127 looking green
Opened
#551 Injector records the delivery AFTER the irreversible send, so a failure in the response window marks a delivered brief NOT_DELIVERED
Opened
#552 redeploy-fleetd.sh exits non-zero after a SUCCESSFUL restart if the post-restart mktemp fails, reporting a working deployment as a failure
Opened
#553 A throwable from any listener callback in onStatus skips the delivered-future completion, so a caller waits forever on a message that WAS delivered
Opened
#555 redeploy-fleetd.sh: the 67-test suite covers the functions and barely touches the main flow
Opened
#556 TurnListener implementations are the only thing maintaining the Injector's own invariant
Opened
#561 Fleetd's TurnListener.onDelivered ordering is load-bearing and untested: swapping two lines makes a throw strand the caller for its full timeout
Opened
#562 LoopWatchdog.health() is public and nothing reads it — surface it without breaking the two scripts that parse /healthz
Opened
#563 A lead cannot ack its own held peer mail, so every lead-to-lead message it polls is delivered a second time and costs a whole lead turn
Opened
#567 LeadMailbox.inspect leaks a probe channel if the close is removed, and 15 contract tests stay green while the line runs 3 times
Opened
#568 The hunter skill has no MemberRole, so every sweep is spawned under a role whose agent file contradicts it
Opened
#571 MessageService collapses the new ATTEMPTED state into TIMED_OUT_QUEUED, so a possibly-delivered message is reported as one that will never arrive
Opened
#572 MessageService.answer() can lose its session-lock release and the whole suite stays green
Opened
#575 The waiter cleanup in MessageService is maintained at three sites, and one of them is outside any finally
Opened
#577 Sweep: find every invariant kept at N sites but asserted at fewer than N
Opened
#578 MessageService.Outcome leaks .name() onto the wire; its sibling ReplyOutcome pins a wireName 400 lines up
Opened
#581 CompletionResolver: 9 sites maintain the CAS-remove invariant, only 2 are asserted (from the #577 sweep)
Opened
#582 AmqpReplyInbox / LeadMailbox: the pendingByMsgId half of the dual-map cleanup is unasserted at every error-path site (from the #577 sweep)
Opened
#586 Sweep: enum.name().toLowerCase() is the house idiom for wire tokens at 15 sites across 5 files, and no test pins any of the long-form ones
Opened
#587 Fleetd.java: sweep every constructor-arg wiring site and report which ones no test would notice being rewired
Opened
#588 Every async ticket dies at a hardcoded 30 minutes, so any unit longer than that always reports failed while the member is still working — 41 measured at exactly 1800s
Opened
#589 Fleetd.main's wiring has 1 behavioural test across 44 sites: pin the 11 whose failure silently turns off a control, with runtime tests not source-text ones
Opened
#590 A broker unreachable at boot silently disables coordination for the process lifetime, and then reports itself as "not configured"
Opened
#591 A lead-side rule has no delivery surface: fleet.charters reaches only spawned members, and it is per-daemon — so an operator policy aimed at "the fleets" cannot reach a single lead
Opened
#592 An architect-settled decision leaves no trace the operator can find: there is no event for it, and #392 shows there is no sink to send one to
Opened
#593 redeploy-fleetd.sh knows which supervisor it is talking to (#492) but not which one it is watching: the log source and the pid count are both still macOS-shaped, and on systemd they fail in opposite directions
Opened
#594 LeadRollover refuses to touch a BLOCKED lead and documents why; LeadHeartbeatLoop:161 gates on injectable() and accepts one — the same question answered both ways in one codebase
Opened
#595 Should a lead be rolled on a timer to pick up instruction changes? Measured: the roll is the expensive answer, it contradicts a stated invariant, and it has never been observed working end-to-end
Opened
#603 redeploy-fleetd.sh reports "no process appeared" on a deploy that fully succeeded
Opened
#604 Charter comparison across hosts: fleet_list drops charterBytes, and a digest cannot tell you what changed
Opened
#608 Flaky: MessageServiceTest.anAlreadyCollectedTicketProducesNoNudge orders an async scheduler with Thread.sleep
Opened
#609 Act on a HIGH lead context: tell the idle lead to hand over
Opened
#612 Fleetd.main's injected wirings are unpinned: 16 call sites go inert with a fully green suite
Opened
#613 A role with no pool and no charter starts anyway, and nothing at boot says so
Opened
#615 A herdr throw during a lead roll leaves status(token) reporting IN_PROGRESS forever
Opened
#618 Measured: CLAUDE_CODE_AUTO_COMPACT_WINDOW beats --autocompact, so fleetd.yaml's comment is backwards
Opened
#621 contextNotice hardcodes "ask the operator" and ignores requireOperatorConfirm, so the knob cannot actually stop the asking
Opened
#625 The subscription guard's call site is unpinned: deleting assertPrimaryClean from Fleetd.main leaves the whole suite green
Opened
#629 FleetdAssembly's herdr boot wait bypasses ResourcePorts: every test with an unhealthy lead herdr pays 30 real seconds
Opened
#630 The assembly's requireOperatorConfirm wiring is unpinned: dropping it silently reverts fleetd #621 with a fully green suite
10 Unresolved Conversations
Open
#131
CB-627: enforce package boundaries with an ArchUnit test instead of splitting into Maven modules
Open
#159
rotate AI_GATEWAY_TOKEN and WORKER_GITEA_TOKEN — both were printed into an operator transcript
Open
#182
ROTATE: a git.ltms.dev token for user ltms was exposed in an agent session, and any credential-resolving command on the Mac will do it again
Open
#184
blocking SSH_AUTH_SOCK is not a control: the forge key is an unencrypted file the member can read
Open
#185
Run members on a second herdr under a dedicated user (optional; single-herdr stays the default)
Open
#110
CB-607: a member holds the operator's ssh-agent socket, which outranks the repo-scoped token CB-302 built
Open
#156
fleet01 as-built: the full second-host setup, and the gaps a fleet manager would have to own
Open
#144
CB-633: constrain a member's environment with an allow-list — the denylist misses a whole secret file
Open
#120
CB-617: deliver a member's role contract as an agent definition file, not as argv — and keep the model in config
Open
#145
CB-632: purge "bridge" from the code, the artefacts and the ops surface