CB-598: per-item reminder counts so backoff-window work is never orphaned #94

Closed
agent wants to merge 0 commits from worker/cb598-6c7ba7-17 into main
Member

Reply/ticket reminder counts used to be one counter per lead per source, carried forward across scheduled ticks. Work that queued during the ~15s push_backoff_ms window between two ticks landed in the pending map before the next tick's own start-of-tick snapshot, so it read as ordinary stale backlog once the shared counter was already at cap - even though no nudge had ever named it.

Fix: ReplyPushLoop.tick() now recomputes each source's reminder count fresh every tick, as the minimum nudge count among that source's currently pending items (tracked per-item in pendingReplies/pendingTickets). A freshly-arrived item (count 0) keeps its source eligible no matter how depleted an older, still-undrained sibling's count is; that sibling keeps riding along in the combined nudge text without spending further budget. decide()'s signature and body are unchanged - only what tick() feeds it changed - so the existing cap-enforcement tests (decideAtCapIsStop, ticketNudgesSendUpToCapThenStop, etc.) stay green unmodified.

stopOrRestart's existing decision-to-release micro-race handling is untouched; it still covers the genuine race that per-item tracking cannot.

Adds two regression tests (aTargetArrivingDuringTheBackoffGetsNudgedDespiteAnAlreadyCappedSibling, aTicketArrivingDuringTheBackoffGetsNudgedDespiteAnAlreadyCappedSibling) that drive tick() directly (package-private, same rationale as stopOrRestart) to deterministically reproduce the backoff-window interleaving. Verified both fail against the old shared-counter logic (assertEquals expected 2 but was 1) before being fixed.

mvn -f bridged/pom.xml clean install: Tests run: 824, Failures: 0, Errors: 0, Skipped: 0 - BUILD SUCCESS, run unpiped. (AmqpReplyInboxRecoveryRaceTest is separately flaky, unrelated to this change - confirmed by rerunning it standalone.)

Reply/ticket reminder counts used to be one counter per lead per source, carried forward across scheduled ticks. Work that queued during the ~15s push_backoff_ms window between two ticks landed in the pending map before the next tick's own start-of-tick snapshot, so it read as ordinary stale backlog once the shared counter was already at cap - even though no nudge had ever named it. Fix: ReplyPushLoop.tick() now recomputes each source's reminder count fresh every tick, as the minimum nudge count among that source's currently pending items (tracked per-item in pendingReplies/pendingTickets). A freshly-arrived item (count 0) keeps its source eligible no matter how depleted an older, still-undrained sibling's count is; that sibling keeps riding along in the combined nudge text without spending further budget. decide()'s signature and body are unchanged - only what tick() feeds it changed - so the existing cap-enforcement tests (decideAtCapIsStop, ticketNudgesSendUpToCapThenStop, etc.) stay green unmodified. stopOrRestart's existing decision-to-release micro-race handling is untouched; it still covers the genuine race that per-item tracking cannot. Adds two regression tests (aTargetArrivingDuringTheBackoffGetsNudgedDespiteAnAlreadyCappedSibling, aTicketArrivingDuringTheBackoffGetsNudgedDespiteAnAlreadyCappedSibling) that drive tick() directly (package-private, same rationale as stopOrRestart) to deterministically reproduce the backoff-window interleaving. Verified both fail against the old shared-counter logic (assertEquals expected 2 but was 1) before being fixed. mvn -f bridged/pom.xml clean install: Tests run: 824, Failures: 0, Errors: 0, Skipped: 0 - BUILD SUCCESS, run unpiped. (AmqpReplyInboxRecoveryRaceTest is separately flaky, unrelated to this change - confirmed by rerunning it standalone.)
agent added 1 commit 2026-08-16 17:49:42 +02:00
CB-598: track reminder counts per pending item, not per lead per source
CI / build (pull_request) Successful in 1m12s
CI / contract (pull_request) Successful in 1m11s
81a0cf4710
Work that arrived during the ~15s push_backoff_ms window between two
ticks landed in the pending map before the next tick's start-of-tick
snapshot, so a shared per-lead-per-source counter (carried forward via
scheduleNext(lead, count+1, ...)) already treated it as exhausted
backlog even though no nudge had ever named it. ReplyPushLoop.tick now
recomputes each source's reminder count fresh every tick as the
minimum nudge count among that source's currently pending items, so a
freshly-arrived item (count 0) keeps its source eligible regardless of
how depleted an older, still-undrained sibling's count is. decide()
itself is unchanged.
Owner

Merged locally as 7a120b3 and pushed to main. Gitea cannot mark a locally merged PR as merged, so I am closing it by hand — merged, not rejected.

Verification is written up on #87. The per-item tracking is the right fix and the javadoc explaining why a carried-forward counter cannot work is the most useful part of the change.

Merged locally as 7a120b3 and pushed to `main`. Gitea cannot mark a locally merged PR as merged, so I am closing it by hand — **merged, not rejected**. Verification is written up on #87. The per-item tracking is the right fix and the javadoc explaining *why a carried-forward counter cannot work* is the most useful part of the change.
ltms closed this pull request 2026-08-16 18:13:58 +02:00
Some checks are pending
CI / build (pull_request) Successful in 1m12s
CI / contract (pull_request) Successful in 1m11s

Pull request closed

Sign in to join this conversation.