CB-588 round 3: close the STOP-vs-onTicketTerminal race, pin nudge polarity, fix comment
- ReplyPushLoop.stopOrRestartTicketLoop: after releasing a lead's active-schedule slot on STOP, restart only if a ticket landed that the pre-decision snapshot did not already account for. A naive "restart on any pending ticket" version was tried first and reverted: it defeated the reminder cap by restarting forever on a stale, never-collected ticket (broke successfulTicketNudgeIncrementsDelivered and ticketNudgesSendUpToCapThenStop). Diffing against a pendingBefore snapshot distinguishes a genuine race arrival from stale cap-exhausted backlog. - Two new ReplyPushLoopTest cases exercise stopOrRestartTicketLoop directly (now package-private) rather than forcing the underlying thread race: aTicketStillPendingWhenTheLoopStopsIsNotStranded (the race must restart) and aStaleUncollectedTicketAtCapDoesNotRestartTheLoop (the cap must still hold). Both were verified to fail against deliberately-reverted versions of the fix before being restored to green. - MessageServiceTest: pin the success-path nudge test's negative direction too (must not contain "FAILED"), not just the failure-path test. - MessageService: complete the whenComplete comment's list of completion paths (TIMED_OUT/BUSY/BACKEND_EXHAUSTED via finishAsyncTask, completeExceptionally on throw, and answer() -> finishAsyncTask(turnId, result)).
This commit is contained in:
@@ -769,6 +769,8 @@ class MessageServiceTest {
|
||||
assertTrue(nudge.contains(ticket), "the nudge should name the ticket: " + nudge);
|
||||
assertTrue(nudge.contains("bridge_poll(ticket="),
|
||||
"the nudge should name the exact ticket-collecting call: " + nudge);
|
||||
assertFalse(nudge.toUpperCase().contains("FAILED"),
|
||||
"a successfully-replied ticket's nudge must not say it failed: " + nudge);
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -15,6 +15,7 @@ import java.util.ArrayList;
|
||||
import java.util.Collections;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.CountDownLatch;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
@@ -341,6 +342,49 @@ class ReplyPushLoopTest {
|
||||
assertFalse(loop.isActive(), "stopping clears the active ticket schedule too");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aTicketStillPendingWhenTheLoopStopsIsNotStranded() {
|
||||
// Regression for the race a reviewer found in gitea PR #73: onTicketTerminal's
|
||||
// activeLeads.putIfAbsent can see the lead's slot as still occupied a moment before
|
||||
// decideTickets' STOP releases it, so the ticket coalesces onto a schedule that is about to
|
||||
// die and nothing ever nudges about it. Forcing that exact thread interleaving is not
|
||||
// reliable, so this drives stopOrRestartTicketLoop — the STOP path's own release-and-recheck —
|
||||
// directly, arranging the state it must not lose a ticket in: a ticket pending for the lead
|
||||
// that was NOT part of the pre-decision snapshot (pendingBefore=empty), standing in for one
|
||||
// that races in during the decision-to-release window.
|
||||
var rec = recordingClient();
|
||||
agents = new AgentControl(rec);
|
||||
ReplyPushLoop loop = loop(1, 100_000); // long backoff — no natural tick fires during this test
|
||||
|
||||
loop.onTicketTerminal("task-1", WORKER, false); // pendingTickets={task-1}; activeLeads={PRIMARY}
|
||||
|
||||
// Stand in for the scheduler thread reaching decideTickets==STOP for this lead — with nothing
|
||||
// pending at decide time — while task-1 races in before the release below runs.
|
||||
loop.stopOrRestartTicketLoop(PRIMARY, Set.of());
|
||||
|
||||
assertTrue(loop.isActive(), "a ticket that raced the loop's stop must reclaim the schedule "
|
||||
+ "slot, not be stranded with no schedule left to ever nudge about it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aStaleUncollectedTicketAtCapDoesNotRestartTheLoop() {
|
||||
// The other direction of the same fix: restarting on ANY non-empty pendingFor(lead) would be
|
||||
// wrong. When STOP is reached because the reminder cap was hit, the same never-collected
|
||||
// ticket is expected to still be there — that is the cap doing its job (acceptance criterion
|
||||
// #7: nudges stay bounded). task-1 here was already accounted for at decide time (it is in
|
||||
// pendingBefore), so it must not restart the loop just because it is still sitting there.
|
||||
var rec = recordingClient();
|
||||
agents = new AgentControl(rec);
|
||||
ReplyPushLoop loop = loop(1, 100_000);
|
||||
|
||||
loop.onTicketTerminal("task-1", WORKER, false); // pendingTickets={task-1}; activeLeads={PRIMARY}
|
||||
|
||||
loop.stopOrRestartTicketLoop(PRIMARY, Set.of("task-1"));
|
||||
|
||||
assertFalse(loop.isActive(), "a stale ticket already accounted for at decide time must not "
|
||||
+ "restart the loop — that would defeat the reminder cap");
|
||||
}
|
||||
|
||||
@Test
|
||||
void ticketNudgeFormatIsCorrect() {
|
||||
String single = ReplyPushLoop.TICKET_NUDGE_FORMAT.formatted("task-1", "", "task-1");
|
||||
|
||||
Reference in New Issue
Block a user