CB-589: cost-first placement + a gateway that reports live capacity #74

Open
opened 2026-08-15 16:28:49 +02:00 by ltms · 1 comment
Owner

The rule this comes from

Operator, 2026-08-15: "maxLoad = 2 is good, as long as we always try to have the boxes busy... we will check the capacity later, expecting a gateway that can answer this question in realtime".

Two separate gaps sit behind that one sentence.

Gap 1 — no placement policy expresses "cheapest first"

WeightedRoundRobinPolicy is smooth weighted round-robin. It picks among every candidate that has a
free slot, in weight ratio. PlacementPolicyUtil.available(ctx) removes candidates that are at cap,
unreachable, quarantined or weight-0 — but it does not rank by cost. So with local: 10,
terra: 2, sonnet: 1 roughly a quarter of unqualified spawns went to a paid profile while gx10
still had a free slot
. That is the opposite of the rule above.

There is a second, sharper failure. The current score map lives for the daemon's whole life:

local at maxLoad  -> filtered out of available(), its score FREEZES
                  -> terra and sonnet keep accumulating and paying against total=3
local frees a slot-> comes back with a stale score (~ -6 + 10 = 4)
                  -> terra can be sitting at ~4 too, and win the pick
                  -> paid spawn while gx10 is idle

FixedPlacementPolicy cannot help: it ignores maxLoad entirely and throws instead of overflowing.
RoundRobinPlacementPolicy has no cost notion either.

Workaround already applied (bridged.yaml, not committed — it is gitignored): local.weight
10 -> 100. At that ratio the margin is decisive (~97 against terra's ceiling of ~4), so gx10 wins
every pick it is eligible for and paid profiles get only the genuine overflow. It works, but it
expresses a preference order through a ratio knob, and a future profile added at weight 150
would silently break it.

Proposed: a packed (or cost-first) policy — order candidates by an explicit cost tier, fill
the cheapest reachable tier to its cap, overflow to the next only when full. Weight then means
"share within a tier", which is what a ratio is actually good for.

Gap 2 — maxLoad is a static guess

maxLoad: 2 for gx10 is a number picked without knowing the box's headroom. It is both a throttle
and a cost lever, and today nothing can tell us whether 2 is conservative or already too high.
bridge_list reports capacity as maxLoad/live/free — all derived from our own bookkeeping,
never from the box.

Proposed: the per-host gateway from the §11 distributed-sandbox topology reports live capacity
(free VRAM, queue depth, in-flight requests, whatever it can honestly answer), and placement asks
it instead of reading a static maxLoad. maxLoad stays as a hard ceiling, not the primary signal.

Acceptance criteria

  1. A placement policy exists that fills the cheapest reachable profile to capacity before it places
    on a more expensive one, and a test proves a paid profile is never chosen while a cheaper one
    has a free slot.
  2. The stale-score hazard above is either impossible in the new policy or covered by a test that
    fails against the current WeightedRoundRobinPolicy.
  3. Cost order is explicit config, not inferred from weight. Adding a new profile without a cost tier
    must not silently reorder existing ones — see the silent-default trap this repo keeps hitting.
  4. local.weight in bridged.yaml can go back to a normal value, and the comment block explaining
    the 100 workaround is removed.
  5. Placement can consult a live-capacity source when one is configured, and falls back to maxLoad
    when it is absent or unreachable — an absent gateway must never block a spawn.
  6. bridge_list's capacity block distinguishes what came from the box from what came from our own
    bookkeeping, so a reader cannot mistake a guess for a measurement.
  7. wiki/11-Features.md gets an entry: what it does, the knob that turns it on, why it exists,
    the gotcha.

Notes

Criteria 1–4 (Gap 1) are self-contained and can ship without the gateway. 5–6 depend on the gateway
existing, so they should be split out if Gap 1 is picked up first.

## The rule this comes from Operator, 2026-08-15: *"maxLoad = 2 is good, as long as we always try to have the boxes busy... we will check the capacity later, expecting a gateway that can answer this question in realtime"*. Two separate gaps sit behind that one sentence. ## Gap 1 — no placement policy expresses "cheapest first" `WeightedRoundRobinPolicy` is smooth weighted round-robin. It picks among every candidate that has a free slot, in weight ratio. `PlacementPolicyUtil.available(ctx)` removes candidates that are at cap, unreachable, quarantined or weight-0 — but it does **not** rank by cost. So with `local: 10`, `terra: 2`, `sonnet: 1` roughly a quarter of unqualified spawns went to a paid profile **while gx10 still had a free slot**. That is the opposite of the rule above. There is a second, sharper failure. The `current` score map lives for the daemon's whole life: ``` local at maxLoad -> filtered out of available(), its score FREEZES -> terra and sonnet keep accumulating and paying against total=3 local frees a slot-> comes back with a stale score (~ -6 + 10 = 4) -> terra can be sitting at ~4 too, and win the pick -> paid spawn while gx10 is idle ``` `FixedPlacementPolicy` cannot help: it ignores `maxLoad` entirely and throws instead of overflowing. `RoundRobinPlacementPolicy` has no cost notion either. **Workaround already applied** (`bridged.yaml`, not committed — it is gitignored): `local.weight` 10 -> 100. At that ratio the margin is decisive (~97 against terra's ceiling of ~4), so gx10 wins every pick it is eligible for and paid profiles get only the genuine overflow. It works, but it expresses a *preference order* through a *ratio* knob, and a future profile added at weight 150 would silently break it. **Proposed:** a `packed` (or `cost-first`) policy — order candidates by an explicit cost tier, fill the cheapest reachable tier to its cap, overflow to the next only when full. Weight then means "share within a tier", which is what a ratio is actually good for. ## Gap 2 — `maxLoad` is a static guess `maxLoad: 2` for gx10 is a number picked without knowing the box's headroom. It is both a throttle and a cost lever, and today nothing can tell us whether 2 is conservative or already too high. `bridge_list` reports `capacity` as `maxLoad`/`live`/`free` — all derived from our own bookkeeping, never from the box. **Proposed:** the per-host gateway from the §11 distributed-sandbox topology reports live capacity (free VRAM, queue depth, in-flight requests, whatever it can honestly answer), and placement asks it instead of reading a static `maxLoad`. `maxLoad` stays as a hard ceiling, not the primary signal. ## Acceptance criteria 1. A placement policy exists that fills the cheapest reachable profile to capacity before it places on a more expensive one, and a test proves a paid profile is **never** chosen while a cheaper one has a free slot. 2. The stale-score hazard above is either impossible in the new policy or covered by a test that fails against the current `WeightedRoundRobinPolicy`. 3. Cost order is explicit config, not inferred from weight. Adding a new profile without a cost tier must not silently reorder existing ones — see the silent-default trap this repo keeps hitting. 4. `local.weight` in `bridged.yaml` can go back to a normal value, and the comment block explaining the 100 workaround is removed. 5. Placement can consult a live-capacity source when one is configured, and falls back to `maxLoad` when it is absent or unreachable — an absent gateway must never block a spawn. 6. `bridge_list`'s `capacity` block distinguishes what came from the box from what came from our own bookkeeping, so a reader cannot mistake a guess for a measurement. 7. `wiki/11-Features.md` gets an entry: what it does, the knob that turns it on, **why it exists**, the gotcha. ## Notes Criteria 1–4 (Gap 1) are self-contained and can ship without the gateway. 5–6 depend on the gateway existing, so they should be split out if Gap 1 is picked up first.
ltms added the ready-to-delegatesilent-default labels 2026-08-15 16:29:15 +02:00
ltms added this to the 1.1 — single-host close-out milestone 2026-08-16 16:49:37 +02:00
Author
Owner

Release cut decision: the policy work moves to 1.2. The documentation half is done now, in 1.1.

Splitting this the way it actually divides.

Shipped now (1.1) — 48d7841 on main and a matching Features entry. bridged.example.yaml now says next to the placement: key that weighted is not cheapest-first, explains the frozen-score effect, and documents the local.weight: 100 workaround together with its trap.

That part could not wait, and the reason is the sharp bit of this ticket. bridged.yaml is gitignored. The workaround exists only in an untracked file on this one host. A fresh host starts with proportional weights and quietly pays for spawns while a free box sits idle, with nothing anywhere to tell the operator why. That is a release problem, not a feature gap.

Deferred to 1.2 — the actual cost-first placement policy, and a gateway that reports live capacity.

The reasoning: 1.1 is the single-host close-out, and the admission rule is if it would still be broken with exactly one host, it belongs in 1.1. A missing placement policy is not broken — it is a working behaviour that costs more than it should, and it has a workaround that is decisive in practice (weight 100 against a paid ceiling of ~4). A missing explanation of that workaround genuinely was broken, and is now fixed.

What is left here when it is picked up:

  1. A placement policy that ranks by cost rather than expressing preference through a ratio.
  2. Fix the frozen-score effect — a profile filtered out at maxLoad should not return with a stale score and lose the next pick.
  3. The live-capacity question the operator asked for: "expecting a gateway that can answer this question in realtime".

Keeping this open on the 1.1 milestone would misrepresent what 1.1 contains, so I am moving it to 2.0 rather than leaving it to be quietly dropped at tag time.

**Release cut decision: the policy work moves to 1.2. The documentation half is done now, in 1.1.** Splitting this the way it actually divides. **Shipped now (1.1)** — `48d7841` on `main` and a matching [Features](https://git.ltms.dev/lms/claude-bridge/wiki/11-Features) entry. `bridged.example.yaml` now says next to the `placement:` key that `weighted` is **not** cheapest-first, explains the frozen-score effect, and documents the `local.weight: 100` workaround together with its trap. That part could not wait, and the reason is the sharp bit of this ticket. `bridged.yaml` is gitignored. The workaround exists **only** in an untracked file on this one host. A fresh host starts with proportional weights and quietly pays for spawns while a free box sits idle, with nothing anywhere to tell the operator why. That is a release problem, not a feature gap. **Deferred to 1.2** — the actual cost-first placement policy, and a gateway that reports live capacity. The reasoning: 1.1 is the single-host close-out, and the admission rule is *if it would still be broken with exactly one host, it belongs in 1.1*. A missing placement **policy** is not broken — it is a working behaviour that costs more than it should, and it has a workaround that is decisive in practice (weight 100 against a paid ceiling of ~4). A missing **explanation** of that workaround genuinely was broken, and is now fixed. What is left here when it is picked up: 1. A placement policy that ranks by cost rather than expressing preference through a ratio. 2. Fix the frozen-score effect — a profile filtered out at `maxLoad` should not return with a stale score and lose the next pick. 3. The live-capacity question the operator asked for: *\"expecting a gateway that can answer this question in realtime\"*. Keeping this open on the 1.1 milestone would misrepresent what 1.1 contains, so I am moving it to **2.0** rather than leaving it to be quietly dropped at tag time.
ltms modified the milestone from 1.1 — single-host close-out to 2.0 — one operation centre, many hosts 2026-08-16 18:27:43 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#74