Exhaustion quarantine is a flat 30 minutes, so a weekly subscription limit is retried ~336 times #466
Open
opened 2026-09-10 13:55:17 +02:00 by ltms
·
2 comments
No Branch/Tag Specified
main
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#466
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
The operator's requirement, in their words: "runtime model limit monitoring, off when subscription limit reach and on when the limit lifted". This ticket is the gap between that and what the code does today.
What the code does, measured on
f5e02feBackendQuarantine.quarantine(String credentialId)is the whole mechanism:cooldownNanosis a constructor field. There is no repeat count and no escalation: every call sets the deadline to now plus the same constant. The default isquarantineCooldownSeconds: 1800(fleetd.example.yaml:459).Why that is wrong for the case it exists for
The two things it is asked to handle have completely different lengths:
A weekly limit at 1800s means about 336 retry cycles. Each cycle is a spawn that either fails outright or — the worse case already recorded in this repo — spawns a member that runs and produces nothing. So "off when the limit is reached" is true for 30 minutes at a time, and "on when the limit is lifted" happens on a timer that knows nothing about whether the limit lifted.
This is live for the operator right now: their opencode weekly allowance ran out.
What #446 does and does not solve
#446 (PR #457, in review) adds the detection half and a warning that names the fix:
That is a manual answer to a requirement stated as automatic, and it is the right first step — the operator gets told exactly which line to flip, and
models:is hot so it takes effect with no restart. What it does not do is stop the 336 retries in the window before anybody reads the log.The design question, and the part that does not depend on it
The open question — which I have asked and which is genuinely the operator's — is how the gate learns the limit lifted:
enabled: false, and a human puts it back. Deterministic, zero wasted spawns, needs a human.The part worth building before that is settled: the escalation in A is an improvement under every one of the three. Under B it is the fallback for when nobody flips the switch. Under C it is the fallback for a provider whose error carries no reset time. Under A it is the whole answer. Today's flat 30 minutes is the worst of all three, so the escalation can be built now and does not pre-empt the decision.
Scope
credentialId: double each time, fromquarantineCooldownSeconds, capped by a newquarantineMaxCooldownSeconds(suggest 43200, 12 hours). Reset the counter on a successful spawn on that credential.fleet_profiles/fleet_listalready carryquarantinedForSeconds; add the attempt count so an operator can see "this is the 5th time" rather than inferring it.quarantineMaxCooldownSecondsequal toquarantineCooldownSecondsmust give exactly today's flat cooldown, so an operator who liked the old shape can pin it.A second, separable finding
quarantineCooldownSecondsis classified DEFERRED —fleetd.example.yaml:481lists it under keys that need a restart, because CB-578 stage B bakes it once into the tracker built at startup.That is the wrong classification for this key. It is exactly the knob an operator wants to change during an outage, and the answer today is "restart the daemon", which drops every in-flight ticket and rendezvous. Note this is the mirror of #427, which was a live-hot key documented as needing a restart; this one is genuinely deferred and ought to be hot.
Whether to fix that here or separately is the implementer's call, but do not silently reclassify it in the docs without moving the code — that is the #427 defect in the other direction.
Acceptance criteria
BackendQuarantinealready takes aLongSupplier nowNanos, so a test can advance time exactly and there is no reason to sleep or to hunt a load level.quarantineMaxCooldownSeconds == quarantineCooldownSecondsreproduces today's flat behaviour.mvn -B clean testinfleetd/, output redirected to a file, exit code captured withrc=$?on its own line. Never pipemvnintotailorgrep, and never usemvn -q— it deletes theTests run:line.The core mechanism landed in
789b6a8(PR #470, closed). I closed this ticket with it and then reopened it a minute later, because I was wrong: two of the three scope items were not delivered, and closing on the strength of the first one would have buried them.Recording that plainly rather than filing a fresh ticket that hides the miss.
Item 1 — landed, with a deliberate substitution
The escalation is in and proven. Doubling per consecutive exhaustion of the same credential, capped at 12x the base, reset after a base cooldown of quiet. My own mutations on the merge killed the escalation, the ceiling and the reset separately, and a wiring test now pins that
Fleetd.mainactually chooses the escalating factory.But the ticket asked for a new
quarantineMaxCooldownSecondsconfig key, and the implementation made the multiplier (2.0) and the ceiling (12x) constants inBackendQuarantineinstead. That was a considered choice, stated openly in the PR: no new YAML surface means no newConfigRefhot/cold/deferred classification question. I accept the reasoning — but it is a substitution, not the thing asked for, and it has a consequence the next item names.Item 3 — no longer possible
With the multiplier and ceiling as constants, an operator cannot pin the flat shape at all. The flat behaviour still exists in the two-argument constructor, and the test suite exercises it — but nothing in
fleetd.yamlcan reach it. Production is now always escalating.I think that is probably the right default, and I am not asking for it to be undone. But the ticket promised an escape hatch and there is none, so an operator who wants the old behaviour has no move except editing Java. That needs a decision rather than silence.
Item 2 — not delivered at all
Nothing was added.
QuarantineStatetracksrepeatCountinternally, and it is not exposed anywhere an operator can see it.This is the item I care about most, and it is why the ticket stays open. The whole point of this work is an operator understanding an outage while it is happening. Right now they see
quarantinedForSeconds: 21600and have to work backwards through the doubling to learn this is the fifth consecutive exhaustion — which requires knowing the base, the multiplier and the ceiling, none of which are in the output. AquarantineRepeatCountbeside the seconds turns arithmetic into a fact.There is a trap here worth stating for whoever picks it up, because this repo has been bitten by it repeatedly: the reported count must be read from the same accessor the cooldown itself uses. A second, independently-derived count can disagree with the behaviour, and a receipt that disagrees with the thing it reports on is worse than no receipt.
CompositePeerLauncher.modelGateState()is the pattern to copy — its javadoc says explicitly that it reads the same accessor the gate enforces so the two can never diverge.The separable finding — still open, correctly
quarantineCooldownSecondsis still deferred. The implementation left it that way and said so, which is the honest outcome: the ticket warned against reclassifying it in the docs without moving the code, and that trap was avoided. The argument for making it hot stands unchanged — it is exactly the knob an operator wants to turn during an outage, and today the answer is a restart, which drops every in-flight ticket and rendezvous.What remains on this ticket
fleet_profiles/fleet_list), read off the same accessor the cooldown uses.quarantineCooldownSeconds, separable as the ticket said.My own error, for the record
This is the second time today that treating a ticket's state as the summary of its content cost something. Earlier I delegated #424, which was already fixed and merged, because I read "open" as "work remains". Here I nearly did the reverse: read "the main mechanism landed" as "the ticket is done". Both come from not re-reading the scope list against what actually shipped. The fix is the same in both directions — read the ticket body at the moment you change its state, not the moment you assign it.
Scope item 2 merged as
25ba7f1(treef22c392). PR #473. Verified landed withgit merge-base --is-ancestor, against an unmerged control branch that reportedno.fleet_profilesandfleet_listnow carryquarantineAttemptbesidequarantinedForSeconds.My own mutation battery, on the merge commit
Every cell a full
mvn -B clean teston treef22c392, the merge's own tree.CONTROL 0:
status()declared once, onequarantines.get()inside it, 2 call sites using it, 0 oldremainingSeconds(credentialId)call sites left, 2quarantineAttemptwrites.CONTROL 1, unmutated: 1629 tests, 0 failures, 0 errors, BUILD SUCCESS. This cell mattered more than usual —
FleetMcp.javawas auto-merged and #469 had also changed it, and a clean auto-merge is not a compiling merge.BackendQuarantineTest.aQuietGapResetsTheReportedAttemptCountToOneToofleet_profilesonlyFleetMcpTest.profilesReportsQuarantineAttemptBesideRemainingSecondsfleet_listonlyFleetMcpTest.capacityRowReportsQuarantineAttemptBesideRemainingSecondsM1 is the ticket's own warning, built in full rather than approximated. I gave the report an independent
ConcurrentHashMapcounter, incremented at the same place the real streak is written, and read the report off that. Both halves look correct in isolation; they disagree only after a quiet gap clears the real streak but not the shadow copy. One test caught exactly that, and it is the one the worker named. So the "one accessor, one read" claim is real, not a comment.M3 and M4 are the pair I care about most. Two call sites are two surfaces, and a test on one proves nothing about the other — that is the shape I have had to hand back three times recently (#446, #466 item 1, #469). Here each call site fails its own named test when its field alone is removed. Nothing had to go back to the worker.
The worker's report was lost, and it cost nothing
Its
fleet_replynever arrived andfleet_pollon its inbox returned empty.fleet_statussaiddone, which on its own looks like a member that produced nothing.It had pushed first and put the full report in the PR body, as the brief required. So the whole report survived. That is the second time that instruction has saved a turn, and it is why I keep it in every brief rather than only the long ones. I also checked the worktree before believing the silence:
git -C <worktree> logshowed the commit, and the branch it committed to (worker/466-quarantine-repeatcount-report) is not the one spawn provisioned (worker/466-quarantine-escalation-5ae9c1-15) — the usual trap.Where #466 now stands
789b6a8.25ba7f1. Done.quarantineCooldownSecondsshould become hot. Neither is urgent: nothing regresses while both stay as they are.I am leaving this ticket open on item 3 alone. Items 1 and 2 are complete and merged.