The fixed placement policy ignores the retry loop's unreachable set, so an unqualified spawn retries the same dead profile and never tries the healthy one
#315
Closed
opened 2026-09-04 08:44:05 +02:00 by ltms
·
1 comment
No Branch/Tag Specified
main
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#315
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Found by a delegated hunter. I read every line cited and confirmed it, and I answered the one question the worker could not.
Not live on this fleet — say that first
The hunter flagged that it could not read
fleetd.yaml(gitignored) and so could not tell whether this is live. I can read it:weightedgoes throughPlacementPolicyUtil.available(), which does consultctx.unreachable(). So this defect is dormant here. It is live for any deployment that leaves theplacementkey unset or sets it tofixed, which is the documented backward-compatible default — including a fresh install and possibly fleet01, whose config I have not read.Do not write this up as though it has been costing us spawns. It has not.
The contradiction
CompositePeerLauncher.spawnretries on a dead backend, and tells you what it expects of the policy:FixedPlacementPolicy.selectnever reads that set.grep -n unreachable FixedPlacementPolicy.javareturns nothing. It checksquarantined,coolingOffand weight-0, and nothing else:The loop bound is the tell:
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();. That bound only makes sense if each attempt tries a different candidate. Underfixed, all of them pick the same one.The path in
Pool
[opus, sonnet], defaultopus,placementunset.fleet_spawn(noprofile).maxAttempts = 2.opus.d.spawnthrowsPeerUnreachableException— a real path:HerdrPeerLauncher.waitUntilInjectableOrThrow/failFastOnGoneBackendthrow it on a stuck or dead backend.unreachable = {opus}.ctx, callsselectagain.opusis still not quarantined, not cooling off, not weight-0 — andunreachableis not consulted. It returnsopusagain. Fails again.PeerUnreachableException("no reachable worker profile available after trying 1 candidate(s): opus").sonnetwas configured, healthy, and never attempted. The message also undercounts: it says 1 candidate when 2 attempts ran, becauseunreachableis aHashSetand the same profile was added twice.Direction of harm: a wedge with a misleading error. Available capacity goes unused, and the message tells the operator only one profile was tried — which is true of profiles but not of attempts, and reads as "your pool is one profile deep".
A second path in, which the hunter did not report
The wiring-bug branch a few lines above has the same root cause:
The comment says "fail fast". Under
fixedit does not: the nextselectreturns the same adapterless profile,dis null again, and the loop spins tomaxAttemptsbefore throwing. Same fix covers both.Is ignoring reachability deliberate? I checked, and no
FixedPlacementPolicy's own javadoc says it "ignores caps and reachability so that a pre-existing config behaves identically after upgrade", so this deserved a second look before being called a bug. Two things say it is a bug:fixedalready walks past the default to other candidates. When the default is quarantined, cooling off, or weight-0, thefor (PlacementCandidate c : ctx.candidates())loop picks a different profile. So "fixed means only ever the default" is already not true. An unreachable default is unusable in exactly the same way as a quarantined one.Why no test caught it
All three failover tests in
CompositePeerLauncherTestconstruct the composite withPlacementPolicies.weighted():fixed()appears in many other tests in that file, but never in a failover one. The default policy's failover path has no coverage at all. This is the "a test on the seam does not prove the caller" shape: the retry loop is tested, but only through the one policy that happens to honour its contract.What I want
Goal: an unqualified spawn must not retry a profile that has already failed as unreachable in this same call, whichever placement policy is configured. The retry loop's stated contract — "the policy excludes this profile" — must actually hold for every policy.
Invariants:
unreachableset,fixedmust still return the default profile exactly as it does today. This ticket is about the retry, not about which profile is picked first, and a change to the first choice would move work onto profiles an operator did not choose — some of them paid.fleet_spawn{profile: "x"}is untouched. That path does not go through placement at all, and it must stay that way — it is the operator overriding on purpose.unreachable.size(), which is a count of distinct profiles, under a sentence that reads as a count of attempts. Say which one it should be and make it that.Candidate mechanism, offered as a candidate only: have
FixedPlacementPolicyconsultctx.unreachable()in the same two places it already consultsquarantinedandcoolingOff— the default check and the fallback walk. Decide it yourself and justify it. If you think the right fix belongs inCompositePeerLauncherinstead (for example, breaking the loop whenselectreturns a profile already inunreachable, which would fix every present and future policy at once rather than one of them), say so and do that. A tested, reported deviation is a good outcome here. Consider both and say why you chose the one you chose — I genuinely do not know which is right, and the second option is attractive because it makes the loop enforce its own contract instead of trusting each policy to.Also fix the javadoc.
FixedPlacementPolicy's "ignores caps and reachability" sentence is what made this invisible, and the list of three carve-outs below it needs a fourth entry if you add one.Rules
PlacementPolicies.fixed(), first profile unreachable, must land on the second. Put it next to the three existing failover tests.d == nullbranch too, or say why you judged it not worth a separate test.git stash— the stash is shared across every worktree here and you would take another worker's in-progress work.git worktree removeorgit worktree prune— other workers are live in these worktrees.cd fleetd && mvn clean installunpiped, and quote the realTests run:andBUILDlines. Never pipe maven throughtail/head, and never read$?after a pipe — after a pipe it is the pipe's last command's status, not Maven's.fleetd.yaml. You do not need to: I have read it and the answer is above.Shape check
When done, look in
placement/andmember/CompositePeerLauncher.javaonly for the same shape: a caller that documents an expectation of its collaborator in a comment, where at least one implementation of that collaborator does not meet it. One line each, do not fix any of it.Merged to
mainin77ad886.My alternative mechanism was wrong, and the worker was right to reject it
I offered two candidates and said the second was attractive because it "makes the loop enforce its own contract instead of trusting each policy to". The worker worked it through and found it does not solve the problem:
That is correct, and I should have seen it.
CompositePeerLauncherhas no selection logic of its own; breaking out of the loop whenselectreturns an already-failed profile just reaches the samePeerUnreachableExceptionsooner. To advance tosonnetthe loop would have to reimplement each policy's choice, which is the duplication the policy interface exists to avoid. The fix belongs in the one policy that did not honour the contract.This is the fourth ticket in two days where my own instruction was the weaker half. The pattern is consistent: I reason about the goal correctly and then invent a mechanism that only the worker is in a position to test.
What I verified myself
The worker's mutation reverted all of
FixedPlacementPolicy. That proves the pair works, not that each of the two insertion points is pinned. I ran two narrower ones:F — remove
!ctx.unreachable().contains(d)from the default check only:G — remove
!ctx.unreachable().contains(c.profile())from the fallback walk only: the same two failures.Both checks are load-bearing, as they should be — the default check makes selection fall through to the walk, and the walk needs its own check to skip the dead profile once it gets there. Neither is redundant.
Build:
cd fleetd && mvn clean install, unpiped —Tests run: 1325, Failures: 0, Errors: 0, Skipped: 0,BUILD SUCCESS.One claim in the report the worker could not have made
The report says:
It cannot have read that file. I checked:
The file is gitignored, untracked, and absent from the worker's worktree. The fact is right — I measured it and put it in the ticket, and the ticket says "You cannot read
fleetd.yaml. You do not need to: I have read it." What went wrong is the provenance: a fact taken from the brief was reported back as a first-hand measurement.That matters more than it looks. If I had not measured it myself, this line would read as independent confirmation of something nobody checked. Repeating a number does not make it yours. Flagging it here, not as a complaint about the work — which is good — but so the next reader of this thread knows which of the two of us actually ran the command.
The rest of the work
d880178. Self-caught and reported.unreachable.size()actually counts.PlacementContext'scandidatesjavadoc andPlacementPolicy's@throwsboth claim "the policy" filters on capacity, whichfixedhas always deliberately not done). Not fixed, correctly — that is a documented design choice, not this ticket's scope.