fleetd #466: escalate quarantine cooldown on repeated exhaustion #470

Closed
agent wants to merge 0 commits from worker/466-quarantine-escalation-5ae9c1-15 into main
Member

Fixes fleetd #466.

What changed

BackendQuarantine.quarantine(credentialId) was flat: until = now + cooldownNanos, every call, no memory of prior calls. A weekly subscription limit therefore got retried on every ~30-minute cooldown expiry — about 336 pointless spawn attempts across the week.

BackendQuarantine now escalates on repeated exhaustion of the same credential:

  • A call that arrives no more than one base cooldown after the previous quarantine's deadline (covers "still quarantined" and "just expired, exhausted again immediately") doubles the cooldown: cooldownNanos * 2.0^(repeatCount-1).
  • Capped at 12x the base cooldown (~6 hours at the default 1800s) — an unbounded backoff would be a permanent, restart-only-recoverable outage, worse than the bug.
  • Reset: a call that arrives after a base-cooldown's worth of quiet (no exhaustion report for that credential) is treated as a fresh occurrence and drops back to the base cooldown.

Kept the two outage states apart

BackendQuarantine's only production caller is Fleetd.exhaustionSink, wired to fire on BACKEND_EXHAUSTED alone. The daemon's other outage state — "cooling off" after repeated non-exhaustion backend errors (an HTTP 5xx storm) — is a wholly separate mechanism, BackendOutagePolicy, with its own fixed 60s cooldown and no repeat tracking. I read every call site of BackendQuarantine.quarantine and confirmed BackendOutagePolicy never calls it (it has its own coolOff). So this change cannot turn a transient 5xx storm into a multi-hour backoff — it only ever escalates on repeated exhaustion.

Reset, honestly

The ideal reset would be "the cooldown expired and the next attempt succeeded." I checked: nothing in this codebase reports a spawn success back to BackendQuarantine — SessionManager and CompositePeerLauncher only ever call quarantine/isQuarantined/remainingSeconds on it, none of which is a success hook. So the reset implemented here is a time-based proxy (a base-cooldown's worth of quiet), not a real success signal — documented as such in the class doc rather than invented as something it isn't. Adding an active probe to confirm recovery is explicitly out of scope per the operator's own design constraint (a probe spends the quota it is measuring).

Backward compatibility

The original two-argument BackendQuarantine(nowNanos, cooldownNanos) constructor is byte-for-byte unchanged in behaviour — it is exactly the escalating formula with backoffMultiplier = 1.0 and maxCooldownNanos = cooldownNanos, which collapses back to now + cooldownNanos on every call regardless of history. All ~20 existing constructor/quarantine() call sites across the test suite (CompositePeerLauncherTest, FleetMcpTest, OpenCodeLauncherTest, SessionManagerTest, FleetAppTest, FleetdExhaustionSinkWarningTest, FleetProfilesQuarantineModelReasonFieldsTest) are untouched and pass unmodified.

Production wiring (Fleetd.main) switches to the new BackendQuarantine.withEscalation(nowNanos, cooldownNanos) factory, which bakes in the 2.0x multiplier / 12x ceiling defaults. No new YAML config key was added — the multiplier and ceiling are constants in BackendQuarantine, documented in its class doc and in a fleetd.example.yaml comment update near quarantineCooldownSeconds (still a Deferred key, unchanged classification — still baked into BackendQuarantine once at startup).

Tests

New tests in BackendQuarantineTest (all against an injected AtomicLong clock, no real sleeps):

  • repeatedExhaustionEscalatesTheCooldownByExactAmounts — 3 consecutive calls, asserts exact deadlines (600s → 1200s → 2400s), not just monotonic growth.
  • escalationStopsAtTheCeiling — pushes 5 calls past the configured ceiling (2400s), asserts it never exceeds it.
  • aQuietGapLongerThanTheBaseCooldownResetsToTheBaseCooldown — 3 escalating calls, then a long quiet gap, then a 4th call; asserts it drops back to the base cooldown (600s).
  • escalatingOneCredentialDoesNotSlowAnother — escalates one credential 3x, asserts a second, unrelated credential's first quarantine is still exactly the base cooldown.
  • withEscalationDefaultsToDoublingCappedAtTwelveTimesTheBase — pins the production factory's default multiplier (2.0) at the real 1800s/30-min scale.
  • anInvalidBackoffMultiplierIsRejected / aCeilingBelowTheBaseCooldownIsRejected — constructor validation.

All pre-existing BackendQuarantineTest tests (flat two-arg constructor) pass unmodified.

Break-and-restore (manual verification, not committed)

  • Escalation: forced the multiplier to 1.0 inside escalatedCooldownNanos. 4 tests caught it, e.g. repeatedExhaustionEscalatesTheCooldownByExactAmounts — AssertionFailedError: a second consecutive exhaustion must double the cooldown, not just increase it ==> expected: <OptionalLong[1200]> but was: <OptionalLong[600]>.
  • Ceiling: removed the Math.min cap. escalationStopsAtTheCeiling caught it — AssertionFailedError: the cooldown must never exceed the configured ceiling, however long the streak gets ==> expected: <OptionalLong[2400]> but was: <OptionalLong[4800]>.
  • Reset: removed the quiet-gap check (always repeatCount + 1). aQuietGapLongerThanTheBaseCooldownResetsToTheBaseCooldown caught it — AssertionFailedError: a long quiet gap must reset the streak back to the base cooldown ==> expected: <OptionalLong[600]> but was: <OptionalLong[2400]>.
  • Restored after each mutation; diffed against the intended version each time to confirm an exact restore.

Build

mvn -B clean test from fleetd/, redirected to a file, exit code checked separately (not piped):

Tests run: 1608, Failures: 0, Errors: 0, Skipped: 0
BUILD SUCCESS
exit=0

Zero [ERROR] lines anywhere in the 2318-line log.

Also noticed, not fixed (scope boundary — reported per the ticket's instruction)

Other fixed-rate retries with no backoff I spotted while reading around this area — listed only, not touched:

  • BackendOutagePolicy's cooling-off cooldown is a flat, hardcoded 60s with no escalation at all — by design, per this ticket's own instructions (a short mechanism for a transient error, not a repeat-abuse case), but it's still a flat-rate retry if the same credential throws non-exhaustion errors indefinitely.
  • ConfigWatcher's file-mtime poll (intervalSeconds, default 10) never backs off, however long the file stays unreadable/unchanged.
  • fleetd's ReplyPushLoop reminder/backoff (pushReminders/pushBackoffMs) is a fixed schedule, not exponential — noted in ConfigRef's deferred-key doc but not investigated further here.
Fixes fleetd #466. ## What changed `BackendQuarantine.quarantine(credentialId)` was flat: `until = now + cooldownNanos`, every call, no memory of prior calls. A weekly subscription limit therefore got retried on every ~30-minute cooldown expiry — about 336 pointless spawn attempts across the week. `BackendQuarantine` now escalates on repeated exhaustion of the *same* credential: - A call that arrives no more than one base cooldown after the previous quarantine's deadline (covers "still quarantined" and "just expired, exhausted again immediately") doubles the cooldown: `cooldownNanos * 2.0^(repeatCount-1)`. - Capped at **12x the base cooldown** (~6 hours at the default 1800s) — an unbounded backoff would be a permanent, restart-only-recoverable outage, worse than the bug. - **Reset**: a call that arrives after a base-cooldown's worth of quiet (no exhaustion report for that credential) is treated as a fresh occurrence and drops back to the base cooldown. ## Kept the two outage states apart `BackendQuarantine`'s only production caller is `Fleetd.exhaustionSink`, wired to fire on `BACKEND_EXHAUSTED` alone. The daemon's other outage state — "cooling off" after repeated non-exhaustion backend errors (an HTTP 5xx storm) — is a wholly separate mechanism, `BackendOutagePolicy`, with its own fixed 60s cooldown and no repeat tracking. I read every call site of `BackendQuarantine.quarantine` and confirmed `BackendOutagePolicy` never calls it (it has its own `coolOff`). So this change cannot turn a transient 5xx storm into a multi-hour backoff — it only ever escalates on repeated exhaustion. ## Reset, honestly The ideal reset would be "the cooldown expired and the next attempt succeeded." I checked: nothing in this codebase reports a spawn success back to `BackendQuarantine` — `SessionManager` and `CompositePeerLauncher` only ever call `quarantine`/`isQuarantined`/`remainingSeconds` on it, none of which is a success hook. So the reset implemented here is a time-based proxy (a base-cooldown's worth of quiet), not a real success signal — documented as such in the class doc rather than invented as something it isn't. Adding an active probe to confirm recovery is explicitly out of scope per the operator's own design constraint (a probe spends the quota it is measuring). ## Backward compatibility The original two-argument `BackendQuarantine(nowNanos, cooldownNanos)` constructor is byte-for-byte unchanged in behaviour — it is exactly the escalating formula with `backoffMultiplier = 1.0` and `maxCooldownNanos = cooldownNanos`, which collapses back to `now + cooldownNanos` on every call regardless of history. All ~20 existing constructor/`quarantine()` call sites across the test suite (`CompositePeerLauncherTest`, `FleetMcpTest`, `OpenCodeLauncherTest`, `SessionManagerTest`, `FleetAppTest`, `FleetdExhaustionSinkWarningTest`, `FleetProfilesQuarantineModelReasonFieldsTest`) are untouched and pass unmodified. Production wiring (`Fleetd.main`) switches to the new `BackendQuarantine.withEscalation(nowNanos, cooldownNanos)` factory, which bakes in the 2.0x multiplier / 12x ceiling defaults. No new YAML config key was added — the multiplier and ceiling are constants in `BackendQuarantine`, documented in its class doc and in a `fleetd.example.yaml` comment update near `quarantineCooldownSeconds` (still a Deferred key, unchanged classification — still baked into `BackendQuarantine` once at startup). ## Tests New tests in `BackendQuarantineTest` (all against an injected `AtomicLong` clock, no real sleeps): - `repeatedExhaustionEscalatesTheCooldownByExactAmounts` — 3 consecutive calls, asserts exact deadlines (600s → 1200s → 2400s), not just monotonic growth. - `escalationStopsAtTheCeiling` — pushes 5 calls past the configured ceiling (2400s), asserts it never exceeds it. - `aQuietGapLongerThanTheBaseCooldownResetsToTheBaseCooldown` — 3 escalating calls, then a long quiet gap, then a 4th call; asserts it drops back to the base cooldown (600s). - `escalatingOneCredentialDoesNotSlowAnother` — escalates one credential 3x, asserts a second, unrelated credential's first quarantine is still exactly the base cooldown. - `withEscalationDefaultsToDoublingCappedAtTwelveTimesTheBase` — pins the production factory's default multiplier (2.0) at the real 1800s/30-min scale. - `anInvalidBackoffMultiplierIsRejected` / `aCeilingBelowTheBaseCooldownIsRejected` — constructor validation. All pre-existing `BackendQuarantineTest` tests (flat two-arg constructor) pass unmodified. ### Break-and-restore (manual verification, not committed) - **Escalation**: forced the multiplier to 1.0 inside `escalatedCooldownNanos`. 4 tests caught it, e.g. `repeatedExhaustionEscalatesTheCooldownByExactAmounts` — `AssertionFailedError: a second consecutive exhaustion must double the cooldown, not just increase it ==> expected: <OptionalLong[1200]> but was: <OptionalLong[600]>`. - **Ceiling**: removed the `Math.min` cap. `escalationStopsAtTheCeiling` caught it — `AssertionFailedError: the cooldown must never exceed the configured ceiling, however long the streak gets ==> expected: <OptionalLong[2400]> but was: <OptionalLong[4800]>`. - **Reset**: removed the quiet-gap check (always `repeatCount + 1`). `aQuietGapLongerThanTheBaseCooldownResetsToTheBaseCooldown` caught it — `AssertionFailedError: a long quiet gap must reset the streak back to the base cooldown ==> expected: <OptionalLong[600]> but was: <OptionalLong[2400]>`. - Restored after each mutation; diffed against the intended version each time to confirm an exact restore. ## Build `mvn -B clean test` from `fleetd/`, redirected to a file, exit code checked separately (not piped): ``` Tests run: 1608, Failures: 0, Errors: 0, Skipped: 0 BUILD SUCCESS exit=0 ``` Zero `[ERROR]` lines anywhere in the 2318-line log. ## Also noticed, not fixed (scope boundary — reported per the ticket's instruction) Other fixed-rate retries with no backoff I spotted while reading around this area — listed only, not touched: - `BackendOutagePolicy`'s cooling-off cooldown is a flat, hardcoded 60s with no escalation at all — by design, per this ticket's own instructions (a short mechanism for a transient error, not a repeat-abuse case), but it's still a flat-rate retry if the same credential throws non-exhaustion errors indefinitely. - `ConfigWatcher`'s file-mtime poll (`intervalSeconds`, default 10) never backs off, however long the file stays unreadable/unchanged. - fleetd's `ReplyPushLoop` reminder/backoff (`pushReminders`/`pushBackoffMs`) is a fixed schedule, not exponential — noted in `ConfigRef`'s deferred-key doc but not investigated further here.
agent added 1 commit 2026-09-10 14:34:18 +02:00
fleetd #466: escalate BackendQuarantine's cooldown on repeated exhaustion
CI / build (pull_request) Successful in 1m34s
CI / contract (pull_request) Successful in 1m36s
5a467e1f8b
A flat 30-minute quarantine retries a weekly subscription limit about
336 times before the window resets. BackendQuarantine now doubles the
cooldown on each consecutive exhaustion of the same credential (no
more than one base cooldown after the previous quarantine's deadline),
capped at 12x the base cooldown (~6h at the 1800s default), and resets
back to the base cooldown once a base-cooldown's worth of quiet has
passed with no further exhaustion.

The flat two-argument constructor is unchanged (equivalent to
multiplier 1.0 / ceiling == base), so all ~20 existing call sites keep
their current shape and behaviour. Production wiring (Fleetd.main)
switches to the new BackendQuarantine.withEscalation factory.

This only touches the exhaustion path (BackendQuarantine's one
production caller is Fleetd.exhaustionSink, fired on BACKEND_EXHAUSTED
alone) and never the separate, unescalated cooling-off mechanism
(BackendOutagePolicy, fixed 60s) that guards against a transient
backend-error storm.
Owner

Adjudicated. The change is accepted and I have asked for one small follow-up on this same branch before I merge.

I ran my own mutation battery on the merge commit (8b4ff78, tree ab937ad), not on the branch. I do not promote a worker's break-and-restore to a fact, so the three proofs in the PR body were re-run by me independently, plus one cell aimed at the half the change did not touch.

Control

Unmutated merge, full build: Tests run: 1608, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS, rc=0, 0 compile-error blocks. A clean auto-merge is not a compiling merge, so this cell is also the merge's own gate.

Shape control on the same tree: withEscalation declared once, Fleetd.main calls it once, new BackendQuarantine( in Fleetd.java 0 times, and exactly 1 test file names withEscalation.

Killed — the three the PR claimed

Cell Mutation Result
M2 streak never continues (repeatCount = 1 always) — the flat-cooldown defect this ticket was filed for KILLED, rc=1, 4 failures
M3 ceiling removed, backoff grows unbounded KILLED, rc=1, 1 failure
M4 quiet-gap reset removed (mutated the condition, not the branch) KILLED, rc=1, 1 failure

Named failures: repeatedExhaustionEscalatesTheCooldownByExactAmounts, escalationStopsAtTheCeiling, aQuietGapLongerThanTheBaseCooldownResetsToTheBaseCooldown, withEscalationDefaultsToDoublingCappedAtTwelveTimesTheBase. So the escalation, the ceiling and the reset are each genuinely pinned, and M4 confirms the reset is pinned as a condition rather than just as a branch that happens to be taken.

Survived — the wiring

-        BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,
+        BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,

Tests run: 1608, Failures: 0, Errors: 0 — BUILD SUCCESS, rc=0. The whole suite stays green.

The factory and its seven new tests stay perfect while the daemon goes back to retrying a weekly limit about 336 times a week. The factory is pinned; the decision to use it is not.

This is not a defect in the work — it is the third time I have hit this shape in this repo. Extracting a value or a factory and testing it moves the untested surface up to the call site rather than shrinking it. The useful question on any extraction is which of three things a test now reaches: the value, the call site, or the selection between values. Extraction answers only the first. (#446 hit the same wall twice, and its survivor is recorded on #460.)

Here the gap is cheap to close, which is why I am asking rather than filing it: 10 test files in this repo already read Fleetd.java as source text, so a wiring assertion needs no daemon and no new idiom. Follow-up sent to the same worker, to push to this branch.

Accepted as judgement calls

The multiplier (2.0) and the ceiling (12x) are the implementer's choices, as the ticket invited. I am accepting both. 12x at the 1800s default is ~6 hours, which keeps a chronically exhausted credential recoverable without a restart — an unbounded backoff would be a permanent outage, and the PR is right that this would be worse than the flat-rate bug.

Two things I specifically checked and agree with:

  • The reset is a time proxy, and the PR says so rather than implying a success signal. I confirmed the claim: nothing in this codebase reports a spawn success back to this class. That honesty is the right call — a reset documented as something it is not would be worse than this one.
  • No automatic probing. That was the operator's design constraint and the implementation respects it: the daemon warns and waits.

Escalation fires on the exhaustion signal only, and BackendOutagePolicy (flat 60s, no repeat tracking) is untouched — so a transient 5xx storm cannot become a multi-hour backoff. That separation is the part I would have worried about most, and it holds.

Note on the numbers above

They describe tree ab937ad. The follow-up commit will change the tree, so I will re-measure on the final merge before pushing rather than carrying these numbers forward.

Adjudicated. The change is accepted and I have asked for one small follow-up on this same branch before I merge. I ran my own mutation battery on the **merge commit** (`8b4ff78`, tree `ab937ad`), not on the branch. I do not promote a worker's break-and-restore to a fact, so the three proofs in the PR body were re-run by me independently, plus one cell aimed at the half the change did not touch. ## Control Unmutated merge, full build: `Tests run: 1608, Failures: 0, Errors: 0, Skipped: 0` — `BUILD SUCCESS`, rc=0, 0 compile-error blocks. A clean auto-merge is not a compiling merge, so this cell is also the merge's own gate. Shape control on the same tree: `withEscalation` declared once, `Fleetd.main` calls it once, `new BackendQuarantine(` in `Fleetd.java` **0** times, and exactly **1** test file names `withEscalation`. ## Killed — the three the PR claimed | Cell | Mutation | Result | |---|---|---| | M2 | streak never continues (`repeatCount = 1` always) — the flat-cooldown defect this ticket was filed for | **KILLED**, rc=1, 4 failures | | M3 | ceiling removed, backoff grows unbounded | **KILLED**, rc=1, 1 failure | | M4 | quiet-gap reset removed (mutated the condition, not the branch) | **KILLED**, rc=1, 1 failure | Named failures: `repeatedExhaustionEscalatesTheCooldownByExactAmounts`, `escalationStopsAtTheCeiling`, `aQuietGapLongerThanTheBaseCooldownResetsToTheBaseCooldown`, `withEscalationDefaultsToDoublingCappedAtTwelveTimesTheBase`. So the escalation, the ceiling and the reset are each genuinely pinned, and M4 confirms the reset is pinned as a *condition* rather than just as a branch that happens to be taken. ## Survived — the wiring ```java - BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime, + BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime, ``` `Tests run: 1608, Failures: 0, Errors: 0` — `BUILD SUCCESS`, rc=0. **The whole suite stays green.** The factory and its seven new tests stay perfect while the daemon goes back to retrying a weekly limit about 336 times a week. The factory is pinned; the *decision to use it* is not. This is not a defect in the work — it is the third time I have hit this shape in this repo. Extracting a value or a factory and testing it moves the untested surface **up** to the call site rather than shrinking it. The useful question on any extraction is which of three things a test now reaches: the value, the call site, or the selection between values. Extraction answers only the first. (#446 hit the same wall twice, and its survivor is recorded on #460.) Here the gap is cheap to close, which is why I am asking rather than filing it: **10 test files in this repo already read `Fleetd.java` as source text**, so a wiring assertion needs no daemon and no new idiom. Follow-up sent to the same worker, to push to this branch. ## Accepted as judgement calls The multiplier (2.0) and the ceiling (12x) are the implementer's choices, as the ticket invited. I am accepting both. 12x at the 1800s default is ~6 hours, which keeps a chronically exhausted credential recoverable without a restart — an unbounded backoff would be a permanent outage, and the PR is right that this would be worse than the flat-rate bug. Two things I specifically checked and agree with: - **The reset is a time proxy, and the PR says so** rather than implying a success signal. I confirmed the claim: nothing in this codebase reports a spawn success back to this class. That honesty is the right call — a reset documented as something it is not would be worse than this one. - **No automatic probing.** That was the operator's design constraint and the implementation respects it: the daemon warns and waits. Escalation fires on the exhaustion signal only, and `BackendOutagePolicy` (flat 60s, no repeat tracking) is untouched — so a transient 5xx storm cannot become a multi-hour backoff. That separation is the part I would have worried about most, and it holds. ## Note on the numbers above They describe tree `ab937ad`. The follow-up commit will change the tree, so I will re-measure on the final merge before pushing rather than carrying these numbers forward.
agent added 1 commit 2026-09-10 14:48:06 +02:00
fleetd #466 follow-up: pin main's choice of the escalating quarantine factory
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m25s
01462c9695
BackendQuarantineTest proves BackendQuarantine.withEscalation itself
escalates, is ceilinged, and resets. Nothing proved Fleetd.main still
calls it: reverting main to the flat two-argument constructor compiled
with 0 errors and left the full 1608-test suite green, because every
other test builds its own BackendQuarantine directly.

Adds FleetdBackendQuarantineWiringTest, a source-text assertion on
Fleetd.java (same idiom as FleetdLeadSeatWiringTest and
FleetdCompletionResolverWiringTest) pinning that main's declaration is
built from BackendQuarantine.withEscalation(...), not `new
BackendQuarantine(...)`. Measured: reverting to the flat constructor
fails it (expected true, was false); renaming the anchored local
variable while keeping behaviour identical also fails it loudly, not
silently. The class doc states plainly this checks source text only,
not that the call executes or that escalation reaches a real backend.
Owner

Merged as 789b6a8 and pushed to origin/main. Closing this by hand — a local --no-ff merge does not close a PR here, so the forge would otherwise keep showing it open with its commits already landed.

The follow-up closed the gap

Re-ran the battery on the new merge (789b6a8, tree 85c2781), because the follow-up commit changed the tree and the earlier numbers described ab937ad. I do not carry numbers across a tree change.

Control, unmutated: Tests run: 1609, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS, rc=0, 0 compile-error blocks.

The survivor is now killed. Same mutation that left 1608 green before — main back on the flat constructor:

rc=1   Tests run: 1609, Failures: 1
FleetdBackendQuarantineWiringTest.mainStillWiresTheEscalatingQuarantineFactory

On the choice of precedent

The worker picked FleetdLeadSeatWiringTest / FleetdCompletionResolverWiringTest over the FleetMcpAuthzTest idiom I named in the brief, and that was the better call — those are literally "which factory does main call to build X", the same question. I checked: there were 6 *WiringTest files before this change and there are 7 now. So this is an established pattern here, not a new idiom, which is worth knowing for #407.

Vacuity, checked myself — and one honest cost

The worker reported that renaming the anchored variable makes the test fail. My first attempt to verify that broke compilation (0 surefire lines, 1 compile-error block), so that cell was void and I am not reporting it as a result.

I re-ran it with a probe that compiles: reflow the anchored call onto one line. Same tokens, byte-identical runtime behaviour, only the source text changes.

rc=1   Tests run: 1609, Failures: 1   compile-error blocks: 0
FleetdBackendQuarantineWiringTest.mainStillWiresTheEscalatingQuarantineFactory

So the test is not vacuous — it fails loudly when its anchor moves rather than passing on a scrape that found nothing. That is the property that matters most for a source-reading test.

The cost, which nobody should be surprised by later: the test is brittle to pure formatting. Reflowing that call, a no-op refactor, turns the build red. Anyone who reformats Fleetd.java will hit it. That is acceptable here — the assertion message quotes the exact expected literal, so the fix is obvious — but it is a real tax and I would not want it on twenty call sites.

The worker also reported, unprompted, that the assertion cannot tell "wrong factory" apart from "renamed the anchor" — both surface as the same assertTrue failure — and judged the split into a control assertion plus a specific assertion unnecessary for a case this narrow. I agree with both the fact and the judgement, and the fact being volunteered rather than discovered is the part I want to note.

Landed

  • 5a467e1 — escalate BackendQuarantine's cooldown on repeated exhaustion
  • 01462c9 — pin main's choice of the escalating quarantine factory
  • merged in 789b6a8

Both proved present on origin/main with git merge-base --is-ancestor. Control: two branches that are not ancestors answered no, so the check discriminates rather than agreeing with everything.

Wiki updated in the wiki's own repo (3dde75e): the "Stop spawning onto an exhausted account" Features entry described a flat cooldown, which stopped being true with this change. Updated rather than duplicated, with quarantineCooldownSeconds now named as the base of the backoff, and four gotchas — the reset being a time proxy, the constants not being config, why the ceiling exists, and that cooling-off is separate and deliberately not escalated.

Merged as `789b6a8` and pushed to `origin/main`. Closing this by hand — a local `--no-ff` merge does not close a PR here, so the forge would otherwise keep showing it open with its commits already landed. ## The follow-up closed the gap Re-ran the battery on the **new** merge (`789b6a8`, tree `85c2781`), because the follow-up commit changed the tree and the earlier numbers described `ab937ad`. I do not carry numbers across a tree change. Control, unmutated: `Tests run: 1609, Failures: 0, Errors: 0, Skipped: 0` — `BUILD SUCCESS`, rc=0, 0 compile-error blocks. **The survivor is now killed.** Same mutation that left 1608 green before — `main` back on the flat constructor: ``` rc=1 Tests run: 1609, Failures: 1 FleetdBackendQuarantineWiringTest.mainStillWiresTheEscalatingQuarantineFactory ``` ## On the choice of precedent The worker picked `FleetdLeadSeatWiringTest` / `FleetdCompletionResolverWiringTest` over the `FleetMcpAuthzTest` idiom I named in the brief, and that was the better call — those are literally "which factory does `main` call to build X", the same question. I checked: there were **6** `*WiringTest` files before this change and there are **7** now. So this is an established pattern here, not a new idiom, which is worth knowing for #407. ## Vacuity, checked myself — and one honest cost The worker reported that renaming the anchored variable makes the test fail. My first attempt to verify that broke compilation (0 surefire lines, 1 compile-error block), so **that cell was void and I am not reporting it as a result.** I re-ran it with a probe that compiles: reflow the anchored call onto one line. Same tokens, byte-identical runtime behaviour, only the source text changes. ``` rc=1 Tests run: 1609, Failures: 1 compile-error blocks: 0 FleetdBackendQuarantineWiringTest.mainStillWiresTheEscalatingQuarantineFactory ``` So the test is **not vacuous** — it fails loudly when its anchor moves rather than passing on a scrape that found nothing. That is the property that matters most for a source-reading test. The cost, which nobody should be surprised by later: **the test is brittle to pure formatting.** Reflowing that call, a no-op refactor, turns the build red. Anyone who reformats `Fleetd.java` will hit it. That is acceptable here — the assertion message quotes the exact expected literal, so the fix is obvious — but it is a real tax and I would not want it on twenty call sites. The worker also reported, unprompted, that the assertion cannot tell "wrong factory" apart from "renamed the anchor" — both surface as the same `assertTrue` failure — and judged the split into a control assertion plus a specific assertion unnecessary for a case this narrow. I agree with both the fact and the judgement, and the fact being volunteered rather than discovered is the part I want to note. ## Landed - `5a467e1` — escalate BackendQuarantine's cooldown on repeated exhaustion - `01462c9` — pin main's choice of the escalating quarantine factory - merged in `789b6a8` Both proved present on `origin/main` with `git merge-base --is-ancestor`. Control: two branches that are **not** ancestors answered no, so the check discriminates rather than agreeing with everything. Wiki updated in the wiki's own repo (`3dde75e`): the "Stop spawning onto an exhausted account" Features entry described a flat cooldown, which stopped being true with this change. Updated rather than duplicated, with `quarantineCooldownSeconds` now named as the **base** of the backoff, and four gotchas — the reset being a time proxy, the constants not being config, why the ceiling exists, and that cooling-off is separate and deliberately not escalated.
ltms closed this pull request 2026-09-10 14:55:59 +02:00
Some checks are pending
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m25s

Pull request closed

Sign in to join this conversation.