CB-588 round 3: close the STOP-vs-onTicketTerminal race, pin nudge polarity, fix comment
CI / build (pull_request) Successful in 1m19s
CI / contract (pull_request) Successful in 1m31s

- ReplyPushLoop.stopOrRestartTicketLoop: after releasing a lead's active-schedule
  slot on STOP, restart only if a ticket landed that the pre-decision snapshot
  did not already account for. A naive "restart on any pending ticket" version
  was tried first and reverted: it defeated the reminder cap by restarting
  forever on a stale, never-collected ticket (broke
  successfulTicketNudgeIncrementsDelivered and ticketNudgesSendUpToCapThenStop).
  Diffing against a pendingBefore snapshot distinguishes a genuine race arrival
  from stale cap-exhausted backlog.
- Two new ReplyPushLoopTest cases exercise stopOrRestartTicketLoop directly
  (now package-private) rather than forcing the underlying thread race:
  aTicketStillPendingWhenTheLoopStopsIsNotStranded (the race must restart) and
  aStaleUncollectedTicketAtCapDoesNotRestartTheLoop (the cap must still hold).
  Both were verified to fail against deliberately-reverted versions of the fix
  before being restored to green.
- MessageServiceTest: pin the success-path nudge test's negative direction too
  (must not contain "FAILED"), not just the failure-path test.
- MessageService: complete the whenComplete comment's list of completion paths
  (TIMED_OUT/BUSY/BACKEND_EXHAUSTED via finishAsyncTask, completeExceptionally
  on throw, and answer() -> finishAsyncTask(turnId, result)).
This commit is contained in:
Dai Ha
2026-08-15 17:02:05 +02:00
parent 6ebad2a91f
commit ac044e7573
4 changed files with 105 additions and 8 deletions
@@ -594,10 +594,14 @@ public final class MessageService {
if (pushLoop != null) {
// CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a
// worker paused in bridge_ask leaves it running, per finishAsyncTask's own contract — so
// this fires exactly once, from whichever path completes it: a worker's bridge_reply, the
// CB-106 completion fallback, a CB-109 wedge, or a CB-516 abandon() on teardown. Without
// this, MessageService.reply's rendezvous fast path (the one an async ticket always takes)
// never told the push loop anything happened — see the class javadoc on sendAsync/CB-107.
// this fires exactly once, from whichever path completes it: finishAsyncTask(task, result)
// below on any non-QUESTION outcome of send() — a worker's bridge_reply, the CB-106
// completion fallback, a CB-109 wedge, TIMED_OUT, BUSY, or BACKEND_EXHAUSTED — the same
// finishAsyncTask reached via answer()'s finishAsyncTask(turnId, result) once a QUESTION
// is resolved, completeExceptionally(t) just below when send() itself throws, or a CB-516
// abandon() on teardown. Without this, MessageService.reply's rendezvous fast path (the
// one an async ticket always takes) never told the push loop anything happened — see the
// class javadoc on sendAsync/CB-107.
task.future.whenComplete((reply, ex) -> {
boolean failed = ex != null || reply == null || !reply.completed();
pushLoop.onTicketTerminal(ticket, target, failed);
@@ -9,6 +9,7 @@ import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
@@ -279,6 +280,7 @@ public final class ReplyPushLoop {
/** Execute one ticket-loop tick — called on the scheduler thread. */
private void ticketTick(String lead, int reminderCount) {
Set<String> pendingBefore = pendingIdsFor(lead);
var action = decideTickets(lead, reminderCount);
switch (action) {
case INJECT -> {
@@ -286,13 +288,58 @@ public final class ReplyPushLoop {
scheduleTicketTick(lead, reminderCount + 1);
}
case WAIT_BUSY -> scheduleTicketTick(lead, reminderCount);
case STOP -> {
activeLeads.remove(lead);
log.debug("push: ticket reminder loop ended for lead {}", lead);
}
case STOP -> stopOrRestartTicketLoop(lead, pendingBefore);
}
}
/** Ticket IDs pending for {@code lead} right now, as a plain snapshot for race comparison. */
private Set<String> pendingIdsFor(String lead) {
return pendingFor(lead).stream().map(PendingTicket::ticket).collect(Collectors.toUnmodifiableSet());
}
/**
* Release {@code lead}'s active-schedule slot, then restart it only if a ticket landed that
* {@code pendingBefore} — the snapshot taken just before this tick's decision — did not already
* account for. {@code onTicketTerminal} reads {@code activeLeads} to decide whether to coalesce
* onto an existing schedule or start one, so a ticket that lands between {@code decideTickets}
* returning {@link Action#STOP} and this removal running sees the (soon-to-be-stale) slot as
* occupied, coalesces onto a schedule that is about to die, and gets no nudge scheduled at all —
* a lost nudge, the exact failure CB-588 exists to remove (found in review, gitea PR #73).
*
* <p>Restarting on ANY non-empty {@code pendingFor(lead)} would be wrong: when STOP is reached
* because the reminder cap was hit rather than the backlog draining, the same never-collected
* ticket is expected to still be sitting there — that is the cap doing its job — and restarting
* would nudge about it forever, defeating the bound (the original CB-307 bounded-reminder
* guarantee, carried into CB-588 by acceptance criterion #7 — this exact regression showed up as
* two existing tests failing once a naive "any pending ticket restarts" version of this fix went
* in: {@code successfulTicketNudgeIncrementsDelivered} and {@code ticketNudgesSendUpToCapThenStop}).
* Diffing the current pending set against {@code pendingBefore} tells the two cases apart: a
* ticket present before this tick's decision is stale backlog, not a race; only a ticket absent
* from {@code pendingBefore} can only have arrived during the decision-to-release window, which is
* exactly the race this method closes.
*
* <p>Package-private so a test can drive the interleaving directly rather than trying to force a
* genuine thread race: pass the exact {@code pendingBefore} snapshot a race requires (or does
* not) and call this to prove the recheck responds correctly either way.
*
* <p>Terminates rather than spinning: this method restarts the schedule at most once per call, and
* a fresh {@link #onTicketTerminal} racing the recheck below still terminates in one of two ways —
* either it observes the slot already vacated (by the {@code activeLeads.remove} above, which
* happens-before this recheck in program order) and claims it itself, or it lands first and this
* recheck then observes its ticket in {@code pendingTickets} and reclaims the slot instead. Exactly
* one side always wins; neither can miss the other, so this never loops on its own account.
*/
void stopOrRestartTicketLoop(String lead, Set<String> pendingBefore) {
activeLeads.remove(lead);
boolean ticketRacedIn = pendingFor(lead).stream().anyMatch(t -> !pendingBefore.contains(t.ticket()));
if (ticketRacedIn && activeLeads.putIfAbsent(lead, Boolean.TRUE) == null) {
log.debug("push: a ticket for lead {} raced the reminder loop's stop — restarting", lead);
scheduleTicketTick(lead, 0);
return;
}
log.debug("push: ticket reminder loop ended for lead {}", lead);
}
/** Send the coalesced ticket nudge and log the event. */
private void injectTicketNudge(String lead, int reminderCount) {
// Re-read rather than threading it down from decideTickets(): a ticket can be collected (or