From 7b99569d5e7f7f31bbf802cf6e4528f90f321c30 Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Fri, 4 Sep 2026 14:02:34 +0700 Subject: [PATCH] =?UTF-8?q?Features:=20#315=20=E2=80=94=20fixed=20placemen?= =?UTF-8?q?t=20honours=20the=20unreachable=20set?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- 11-Features.md | 25 +++++++++++++++++++++++++ 1 file changed, 25 insertions(+) diff --git a/11-Features.md b/11-Features.md index 709634c..ee217fa 100644 --- a/11-Features.md +++ b/11-Features.md @@ -3835,3 +3835,28 @@ ordinary "a branch named X already exists" refusal would delete the operator's e directory on that refusal, so no directory means nothing was created and the branch is not ours to touch. Removing that check makes `addFailureBeforeCreatingAWorktreeIsQuiet` fail with the branch actually deleted. Do not remove it. + +## `fixed` placement honours the failover loop's unreachable set + +**What.** The `fixed` placement policy now skips a profile that the current spawn attempt has +already found unreachable. It checks `ctx.unreachable()` in both places it picks a profile: the +configured default, and the fallback walk over the remaining candidates. + +**On.** Only when `placement: fixed` is set in `fleetd.yaml`. The live fleet runs `weighted`, which +already honoured the set, so this changed nothing here — it was reachable only for an operator who +had switched to `fixed`. + +**Why it exists.** `CompositePeerLauncher` retries a failed spawn by adding the dead profile to the +unreachable set and asking the policy again, and its own comment says why: *"Update the context for +the next selection so the policy excludes this profile."* `fixed` ignored the set, so it handed back +the same dead profile on every attempt. The loop then spent all of its attempts on one profile and +reported failure, while healthy profiles sat idle and were never tried. The failover existed but +could not move. + +**One thing to know for maintenance.** `fixed` ignoring the *load caps* is deliberate and stays — +that is what "fixed" means. Reachability is not a cap, and the javadoc now separates the two so the +next reader does not undo this as a bug fix. Also note what does **not** work: breaking out of the +retry loop early when the policy returns an already-unreachable profile only makes it fail faster on +the same dead profile, because only the policy chooses what comes next. The failure message now +counts *distinct* candidates, so "tried 1 distinct candidate(s)" is the visible symptom of this +whole class of problem.