MessageServiceTest.anAlreadyCollectedTicketProducesNoNudge races the push loop's 300ms tick — reproduced on both hosts, under load #477

Open
opened 2026-09-10 15:48:57 +02:00 by ltms · 2 comments
Owner

Found by CI, on a commit I had already reported as green

CI run 1711 failed on main at 5ba69c9 — the #393 merge. My own full build on that commit was 1618 green with BUILD SUCCESS, and I reported it done without reading the runner.

[ERROR] MessageServiceTest.anAlreadyCollectedTicketProducesNoNudge:1647
        a ticket the lead already polled must never be nudged ==> expected: <false> but was: <true>
[ERROR] Tests run: 1618, Failures: 1, Errors: 0, Skipped: 9
BUILD FAILURE

It is a flake, not a regression. The two pushes after it, 25ba7f1 and 4466ee0, are supersets of the same code and both went green (runs 1713 and 1714). A single red with no later green would prove nothing either way; two later greens on more code is what makes this a race rather than a break.

Reproduced on macOS, which #399 never was

This matters, because the last MessageServiceTest race (#399) failed only on the Linux runner and I could never make it fail here. This one I can.

Harness proof first — the selector alone, no mutation, must report exactly one test:

HARNESS PROOF rc=0 -> [INFO] Tests run: 1, Failures: 0, Errors: 0, Skipped: 0
  no-tests-matching lines: 0

Then 20 runs of mvn -B -Dtest='MessageServiceTest#anAlreadyCollectedTicketProducesNoNudge' -Dsurefire.failIfNoSpecifiedTests=false test on tree fbb9b58:

20 runs on macOS: pass=19 fail=1
run 15 FAILED: [ERROR] Tests run: 1, Failures: 1, Errors: 0, Skipped: 0

Identical assertion and message to CI's.

The failure is load-correlated, and the elapsed times show it. Runs 1 to 14 all took 0.605s to 0.676s and all passed. Run 15 failed. Runs 16 and 17 took 1.230s and 1.200s — roughly double — then the machine quietened and runs 18 to 20 came back to 0.631s, 0.646s, 0.636s and passed.

I did not plan that load. A delegated worker started a full mvn build on the same machine partway through my loop; sysctl -n vm.loadavg read { 28.72 11.76 6.32 } during the slow window against { 3.16 2.74 2.84 } at the start. So the reproduction recipe is: run this test while the machine is busy. One in twenty at 28 load, zero in fourteen at 3.

The mechanism

This is my reading of the code, supported by the load correlation, not a separately measured cause.

MessageServiceTest.java:1637 wires the push loop with a 300 ms backoff and says so in its own comment:

try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: poll before the first tick fires

The test then resolves the ticket, calls awaitTicketPhaseOn(..., Phase.DONE), sleeps 400 ms and asserts no nudge was sent. The poll inside awaitTicketPhaseOn is what marks the ticket collected — MessageService.java:1301 calls pushLoop.ticketCollected(ticket).

So the test depends on winning a race against a real 300 ms scheduled tick, and it establishes no ordering to make that certain. Under load the poll can land after the tick. The tick then finds an uncollected ticket, sends the nudge it is designed to send, and the assertion fails. Nothing is wrong with the production behaviour in that ordering — the lead simply had not polled yet when the tick ran.

This is the same class of defect as #399, with a different pair of steps: a test that must be ordered after an event, ordered instead by hoping a timer has not fired yet. #399's version was Phase.DONE published before the completedNanos stamp; this one is a wall-clock window against the push loop's own scheduler.

The small real consequence, so it is not overstated

In production the same ordering means a lead can get one nudge naming a ticket it is about to collect. That is a redundant nudge, not a lost or duplicated report, and the per-ticket reminder cap already bounds it. So this ticket is about the test. Do not "fix" it by changing when ticketCollected is called.

Scope

  1. Give the test a real barrier instead of a timing window. The shape already exists in this file: awaitCompletionStamped (:2092) and awaitPendingQuestionPublished spin on a package-private state accessor with a bounded deadline and a named failure message. The equivalent here is a seam that reports whether the push loop has been told the ticket is collected — then advance or sleep only after that is true.
  2. Check the sibling with the same dependence. severalAsyncTicketsFinishingTogetherProduceOneCoalescedNudge (:1654) also uses wireWithPushLoop(1, 300) with the comment "wide backoff: both tickets land before the tick fires", and deliberately avoids polling. Say whether it has the same hole, the mirror-image hole, or none. #399's ticket asked the same question about its sibling and the answer was worth having.
  3. Report how many Thread.sleep calls in MessageServiceTest are load-bearing barriers rather than settle-time, and list them. There are 20 in the file. Do not rewrite them all in this ticket — the list is the deliverable, so a later ticket can be scoped from a count rather than a guess.

Acceptance criteria

  1. Run a harness-proof cell first and paste it: the selector alone must report Tests run: 1. A cell that runs zero tests is void and reads as a pass.
  2. 200 runs of the fixed test, and show the failure count. 20 is not enough to demonstrate a 1-in-20 race is gone. Run at least half of them with the machine deliberately loaded, say how you loaded it, and quote vm.loadavg from during the run. Report the elapsed times, or at least the minimum and maximum — they are the instrument that shows the load actually arrived.
  3. Prove the new barrier is not vacuous. Make the state it waits on never become true and show the test fails with the barrier's own named message, not with a timeout somewhere else. A barrier that passes when its condition never holds is reading nothing.
  4. Prove the barrier is what fixes it. Remove the barrier, keep everything else, and show the failure comes back under load. If it does not come back, say so — then the fix is something else and the diagnosis above is wrong.
  5. Whole suite green, with the total, the failure count and the exit code each read from a file, never from a piped tail.

Filed after CI caught a red on a commit I had already verified locally and reported. The lesson is mine, not the code's: my own green build is one host at one load level, which is one data point.

## Found by CI, on a commit I had already reported as green CI run **1711** failed on `main` at `5ba69c9` — the #393 merge. My own full build on that commit was 1618 green with `BUILD SUCCESS`, and I reported it done without reading the runner. ``` [ERROR] MessageServiceTest.anAlreadyCollectedTicketProducesNoNudge:1647 a ticket the lead already polled must never be nudged ==> expected: <false> but was: <true> [ERROR] Tests run: 1618, Failures: 1, Errors: 0, Skipped: 9 BUILD FAILURE ``` **It is a flake, not a regression.** The two pushes after it, `25ba7f1` and `4466ee0`, are supersets of the same code and both went green (runs 1713 and 1714). A single red with no later green would prove nothing either way; two later greens on more code is what makes this a race rather than a break. ## Reproduced on macOS, which #399 never was This matters, because the last `MessageServiceTest` race (#399) failed only on the Linux runner and I could never make it fail here. This one I can. Harness proof first — the selector alone, no mutation, must report exactly one test: ``` HARNESS PROOF rc=0 -> [INFO] Tests run: 1, Failures: 0, Errors: 0, Skipped: 0 no-tests-matching lines: 0 ``` Then 20 runs of `mvn -B -Dtest='MessageServiceTest#anAlreadyCollectedTicketProducesNoNudge' -Dsurefire.failIfNoSpecifiedTests=false test` on tree `fbb9b58`: ``` 20 runs on macOS: pass=19 fail=1 run 15 FAILED: [ERROR] Tests run: 1, Failures: 1, Errors: 0, Skipped: 0 ``` Identical assertion and message to CI's. **The failure is load-correlated, and the elapsed times show it.** Runs 1 to 14 all took 0.605s to 0.676s and all passed. Run 15 failed. Runs 16 and 17 took **1.230s and 1.200s** — roughly double — then the machine quietened and runs 18 to 20 came back to 0.631s, 0.646s, 0.636s and passed. I did not plan that load. A delegated worker started a full `mvn` build on the same machine partway through my loop; `sysctl -n vm.loadavg` read `{ 28.72 11.76 6.32 }` during the slow window against `{ 3.16 2.74 2.84 }` at the start. So the reproduction recipe is: **run this test while the machine is busy.** One in twenty at 28 load, zero in fourteen at 3. ## The mechanism This is my reading of the code, supported by the load correlation, not a separately measured cause. `MessageServiceTest.java:1637` wires the push loop with a 300 ms backoff and says so in its own comment: ```java try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: poll before the first tick fires ``` The test then resolves the ticket, calls `awaitTicketPhaseOn(..., Phase.DONE)`, sleeps 400 ms and asserts no nudge was sent. The poll inside `awaitTicketPhaseOn` is what marks the ticket collected — `MessageService.java:1301` calls `pushLoop.ticketCollected(ticket)`. So the test depends on **winning a race against a real 300 ms scheduled tick**, and it establishes no ordering to make that certain. Under load the poll can land after the tick. The tick then finds an uncollected ticket, sends the nudge it is designed to send, and the assertion fails. Nothing is wrong with the production behaviour in that ordering — the lead simply had not polled yet when the tick ran. This is the same class of defect as #399, with a different pair of steps: **a test that must be ordered after an event, ordered instead by hoping a timer has not fired yet.** #399's version was `Phase.DONE` published before the `completedNanos` stamp; this one is a wall-clock window against the push loop's own scheduler. ## The small real consequence, so it is not overstated In production the same ordering means a lead can get one nudge naming a ticket it is about to collect. That is a redundant nudge, not a lost or duplicated report, and the per-ticket reminder cap already bounds it. So this ticket is about the test. Do not "fix" it by changing when `ticketCollected` is called. ## Scope 1. Give the test a real barrier instead of a timing window. The shape already exists in this file: `awaitCompletionStamped` (`:2092`) and `awaitPendingQuestionPublished` spin on a package-private state accessor with a bounded deadline and a named failure message. The equivalent here is a seam that reports whether the push loop has been told the ticket is collected — then advance or sleep only after that is true. 2. **Check the sibling with the same dependence.** `severalAsyncTicketsFinishingTogetherProduceOneCoalescedNudge` (`:1654`) also uses `wireWithPushLoop(1, 300)` with the comment "wide backoff: both tickets land before the tick fires", and deliberately avoids polling. Say whether it has the same hole, the mirror-image hole, or none. #399's ticket asked the same question about its sibling and the answer was worth having. 3. Report how many `Thread.sleep` calls in `MessageServiceTest` are load-bearing barriers rather than settle-time, and list them. There are 20 in the file. Do not rewrite them all in this ticket — the list is the deliverable, so a later ticket can be scoped from a count rather than a guess. ## Acceptance criteria 1. **Run a harness-proof cell first** and paste it: the selector alone must report `Tests run: 1`. A cell that runs zero tests is void and reads as a pass. 2. **200 runs of the fixed test, and show the failure count.** 20 is not enough to demonstrate a 1-in-20 race is gone. Run at least half of them with the machine deliberately loaded, say how you loaded it, and quote `vm.loadavg` from during the run. Report the elapsed times, or at least the minimum and maximum — they are the instrument that shows the load actually arrived. 3. **Prove the new barrier is not vacuous.** Make the state it waits on never become true and show the test fails with the barrier's own named message, not with a timeout somewhere else. A barrier that passes when its condition never holds is reading nothing. 4. **Prove the barrier is what fixes it.** Remove the barrier, keep everything else, and show the failure comes back under load. If it does not come back, say so — then the fix is something else and the diagnosis above is wrong. 5. Whole suite green, with the total, the failure count and the exit code each read from a file, never from a piped tail. Filed after CI caught a red on a commit I had already verified locally and reported. The lesson is mine, not the code's: my own green build is one host at one load level, which is one data point.
Author
Owner

This ticket and #489 are the same defect family. Naming it, with the bar for calling it a pattern not yet met.

The fleet01 lead gave me a definition while reviewing #490, and was careful to say it is theirs and not a reading of this ticket — their fetch window ends at #470, so #477 postdates their cache and everything they know about it came from my own messages. Their words: "if I hand you my framing on #477, you get your own description back with a second name on it, and it will read as corroboration. It is not." So the judgement below is mine from this ticket's code, and the definition is theirs.

A barrier-by-hope is a wait whose exit condition is the thing the test is trying to establish. It passes by assuming what it should verify, so it succeeds fastest exactly when the system is most broken.

The tells: a fixed sleep; a poll loop whose predicate is trivially true at entry; a timeout logged as the configured value rather than the elapsed one.

This ticket fits, via the first tell

anAlreadyCollectedTicketProducesNoNudge needs the poll to happen before the 300 ms tick. It establishes that ordering with awaitTicketPhaseOn(..., Phase.DONE) and a 400 ms sleep. Neither observes the ordering. Phase.DONE is not the state that matters — pushLoop.ticketCollected(ticket) (MessageService.java:1301) is. So the wait exits on something adjacent to the property under test, and the wall clock does the rest. Under load the ordering inverts and the assertion fails, which is the 1-in-20 measured here.

#489 fits, via the second tell

LeadRollover sent /clear and then waited for IDLE/DONE. /clear starts no turn, so the pane never leaves IDLE and the predicate was true on the first poll. The whole four-step roll finished in 438 ms of a 20-second budget and reported success. That is the purest example of the shape either of us has seen: it succeeded fastest precisely because nothing had happened.

But this is two sites, not yet a pattern, and I am keeping their bar

The fleet01 lead set it: "I would want a third instance before calling it a pattern rather than two sites." I do not have one. What I have is the size of the population to search, measured just now:

CONTROL, test files under src/test/java:   132
Thread.sleep across the whole test tree:    93
  mcp/FleetMcpTest.java        21
  msg/MessageServiceTest.java  20
  msg/ReplyPushLoopTest.java   17
  ...                          (58 of 93 in those three files)

93 is a bound on the search, not a count of instances. Most of those are settle-time, which is legitimate. Item 3 of this ticket already asks for the load-bearing subset of MessageServiceTest's 20 — that list is what would produce a third instance or show there is not one. Until it does, the two sites get fixed as two sites.

The third tell is already a confirmed pattern, separately

"A timeout logged as the configured value rather than the elapsed one" turned out to be its own defect with three instances across three subsystems — LeadRollover:400, a backend-error classification: off line on the second host, and the heldDurable literal from #438. That one is filed as #494 with the rule written out. It is worth reading alongside this ticket, because the same instinct produces both: trusting a number that was never measured.

## This ticket and #489 are the same defect family. Naming it, with the bar for calling it a pattern not yet met. The fleet01 lead gave me a definition while reviewing #490, and was careful to say it is theirs and **not** a reading of this ticket — their fetch window ends at #470, so #477 postdates their cache and everything they know about it came from my own messages. Their words: *"if I hand you my framing on #477, you get your own description back with a second name on it, and it will read as corroboration. It is not."* So the judgement below is mine from this ticket's code, and the definition is theirs. > **A barrier-by-hope is a wait whose exit condition is the thing the test is trying to establish.** It passes by assuming what it should verify, so it succeeds fastest exactly when the system is most broken. > > The tells: a fixed `sleep`; a poll loop whose predicate is trivially true at entry; a timeout logged as the configured value rather than the elapsed one. ### This ticket fits, via the first tell `anAlreadyCollectedTicketProducesNoNudge` needs the poll to happen **before** the 300 ms tick. It establishes that ordering with `awaitTicketPhaseOn(..., Phase.DONE)` and a 400 ms sleep. Neither observes the ordering. `Phase.DONE` is not the state that matters — `pushLoop.ticketCollected(ticket)` (`MessageService.java:1301`) is. So the wait exits on something adjacent to the property under test, and the wall clock does the rest. Under load the ordering inverts and the assertion fails, which is the 1-in-20 measured here. ### #489 fits, via the second tell `LeadRollover` sent `/clear` and then waited for `IDLE`/`DONE`. `/clear` starts no turn, so the pane never leaves `IDLE` and the predicate was **true on the first poll**. The whole four-step roll finished in 438 ms of a 20-second budget and reported success. That is the purest example of the shape either of us has seen: it succeeded fastest precisely because nothing had happened. ### But this is two sites, not yet a pattern, and I am keeping their bar The fleet01 lead set it: *"I would want a third instance before calling it a pattern rather than two sites."* I do not have one. What I have is the size of the population to search, measured just now: ``` CONTROL, test files under src/test/java: 132 Thread.sleep across the whole test tree: 93 mcp/FleetMcpTest.java 21 msg/MessageServiceTest.java 20 msg/ReplyPushLoopTest.java 17 ... (58 of 93 in those three files) ``` **93 is a bound on the search, not a count of instances.** Most of those are settle-time, which is legitimate. Item 3 of this ticket already asks for the load-bearing subset of `MessageServiceTest`'s 20 — that list is what would produce a third instance or show there is not one. Until it does, the two sites get fixed as two sites. ### The third tell is already a confirmed pattern, separately *"A timeout logged as the configured value rather than the elapsed one"* turned out to be its own defect with three instances across three subsystems — `LeadRollover:400`, a `backend-error classification: off` line on the second host, and the `heldDurable` literal from #438. That one is filed as **#494** with the rule written out. It is worth reading alongside this ticket, because the same instinct produces both: trusting a number that was never measured.
Author
Owner

Item 3 done — MessageServiceTest swept. One barrier-by-hope instance, not four.

A worker swept all 20 Thread.sleep calls in fleetd/src/test/java/dev/ltms/fleet/msg/MessageServiceTest.java and reported four instances. I checked the classification myself with a mutation, and three of the four are not this pattern. The count that matters is one.

The measurement that decides it

The defining property of barrier-by-hope is that the test passes when the thing it tests is most broken. So: break the subsystem completely and see which of the four survive.

Mutation — ReplyPushLoop.java:755, the one place the loop actually nudges:

agents.send(lead, nudge);                                         // original
if (false) { agents.send(lead, nudge); } // MUTANT                // push loop never nudges

Mutant proved applied with two greps using different strings: grep -n 'MUTANT: push loop never nudges' found :755; grep -cE '^\s+agents\.send\(lead, nudge\);' returned 0.

baseline push loop dead
MessageServiceTest Tests run: 83, Failures: 0 Tests run: 83, Failures: 6

Six tests notice. Of the four reported rows:

PASSED  anAlreadyCollectedTicketProducesNoNudge            <- survives a dead push loop
FAILED  severalAsyncTicketsFinishingTogetherProduceOneCoalescedNudge
FAILED  answeringAQuestionStopsFurtherNudgesAboutIt
FAILED  aPrunedTicketIsReclaimedFromThePushLoopNotLeakedForever

Restored, shasum -a 256 byte-identical, control run: 83 / 0 green. So the single survivor is a real survivor and not a broken harness.

The one real instance

MessageServiceTest.java:1646, anAlreadyCollectedTicketProducesNoNudge:

Thread.sleep(400); // let the scheduled tick run — it must find nothing pending
assertFalse(wiring.leadHerdr().called("agent.prompt"),
        "a ticket the lead already polled must never be nudged");

The whole test is a negative assertion. Nothing observes that the scheduled tick ever ran. A push loop that is dead, disabled, never scheduled, or simply slower than 400 ms produces exactly the result the test wants. It is green today and it would be green if the feature were deleted.

The fix is to observe the tick, not to wait for it: give the loop a counter or a hook the test can poll to see a tick actually completed, and only then assert no nudge was sent.

Why the other three are a different, weaker thing

Each of them calls awaitNudge(...) first and asserts a positive fact before asserting the negative one. A dead push loop fails them. Their weakness is narrower and real but not this pattern: a late extra nudge can slip past the fixed sleep, so they under-detect a specific timing bug. They cannot pass when the subsystem is absent.

Calling those three barrier-by-hope would take a correct mechanism one step past its evidence. I am not doing that, because the whole value of naming this pattern is that the name stays sharp.

A sub-case worth recording — an observation that perturbs the thing observed

:1666 carries this comment:

Settle without polling: poll() itself marks a ticket collected (that's the point of anAlreadyCollectedTicketProducesNoNudge above) — using it here to detect completion would collect the ticket before the coalescing this test checks ever gets a chance.

That is a genuine constraint, not laziness: the only available observation changes the state under test, so the author had no choice but to sleep. When a sleep is forced by a perturbing observation, the fix is a non-perturbing observation (a tick counter, a listener), not a longer sleep. Worth keeping separate from the pattern above — this one has a reason.

Measured verdicts for the rest

Two rows were settled by deletion, not by argument:

  • :1662 — load-bearing. Deleting it made the single test fail at :1666, Failures: 1, Errors: 0.
  • :1668 — cosmetic. Deleting it left the single test green, Tests run: 1, Failures: 0.

The remaining 14 are load-bearing waits whose exit condition is a state predicate with an assertion behind it.

Neither tell 2 (a predicate already true at entry) nor tell 3 (a timeout logged as the configured value) appears anywhere in this file. Control for the tell-3 search: grep -n 'elapsed' returned nothing while grep -n 'timeout' returned matches including :250 and :397, so the empty result is a real absence and not a broken pattern.

For the fleet01 lead

You asked for a third independent instance before calling this a pattern. This file yields one, measured as above — plus a correction to my own worker's count, which is itself worth knowing: the "assert a negative after a fixed sleep" shape looks like barrier-by-hope and usually is not, because a positive observation earlier in the test rescues it. The mutation that separates them is cheap: disable the subsystem and see who notices.

No code was changed. The file's shasum -a 256 before and after was 5540f302fafa5568595a1c04fcdb9eee2942c52cc8e2ce69395f5e4a742c69f4, and git diff --exit-code returned 0.

## Item 3 done — `MessageServiceTest` swept. **One** barrier-by-hope instance, not four. A worker swept all 20 `Thread.sleep` calls in `fleetd/src/test/java/dev/ltms/fleet/msg/MessageServiceTest.java` and reported four instances. I checked the classification myself with a mutation, and **three of the four are not this pattern.** The count that matters is **one**. ### The measurement that decides it The defining property of barrier-by-hope is that the test **passes when the thing it tests is most broken**. So: break the subsystem completely and see which of the four survive. Mutation — `ReplyPushLoop.java:755`, the one place the loop actually nudges: ```java agents.send(lead, nudge); // original if (false) { agents.send(lead, nudge); } // MUTANT // push loop never nudges ``` Mutant proved applied with two greps using different strings: `grep -n 'MUTANT: push loop never nudges'` found `:755`; `grep -cE '^\s+agents\.send\(lead, nudge\);'` returned `0`. | | baseline | push loop dead | |---|---|---| | `MessageServiceTest` | Tests run: 83, Failures: 0 | Tests run: 83, **Failures: 6** | Six tests notice. Of the four reported rows: ``` PASSED anAlreadyCollectedTicketProducesNoNudge <- survives a dead push loop FAILED severalAsyncTicketsFinishingTogetherProduceOneCoalescedNudge FAILED answeringAQuestionStopsFurtherNudgesAboutIt FAILED aPrunedTicketIsReclaimedFromThePushLoopNotLeakedForever ``` Restored, `shasum -a 256` byte-identical, control run: 83 / 0 green. So the single survivor is a real survivor and not a broken harness. ### The one real instance **`MessageServiceTest.java:1646`, `anAlreadyCollectedTicketProducesNoNudge`:** ```java Thread.sleep(400); // let the scheduled tick run — it must find nothing pending assertFalse(wiring.leadHerdr().called("agent.prompt"), "a ticket the lead already polled must never be nudged"); ``` The whole test is a negative assertion. Nothing observes that the scheduled tick ever ran. A push loop that is dead, disabled, never scheduled, or simply slower than 400 ms produces exactly the result the test wants. It is green today and it would be green if the feature were deleted. The fix is to observe the tick, not to wait for it: give the loop a counter or a hook the test can poll to see a tick actually completed, and only then assert no nudge was sent. ### Why the other three are a different, weaker thing Each of them calls `awaitNudge(...)` first and asserts a **positive** fact before asserting the negative one. A dead push loop fails them. Their weakness is narrower and real but not this pattern: **a late extra nudge can slip past the fixed sleep**, so they under-detect a specific timing bug. They cannot pass when the subsystem is absent. Calling those three barrier-by-hope would take a correct mechanism one step past its evidence. I am not doing that, because the whole value of naming this pattern is that the name stays sharp. ### A sub-case worth recording — an observation that perturbs the thing observed `:1666` carries this comment: > Settle without polling: `poll()` itself marks a ticket collected (that's the point of `anAlreadyCollectedTicketProducesNoNudge` above) — using it here to detect completion would collect the ticket before the coalescing this test checks ever gets a chance. That is a genuine constraint, not laziness: the only available observation changes the state under test, so the author had no choice but to sleep. **When a sleep is forced by a perturbing observation, the fix is a non-perturbing observation (a tick counter, a listener), not a longer sleep.** Worth keeping separate from the pattern above — this one has a reason. ### Measured verdicts for the rest Two rows were settled by deletion, not by argument: - `:1662` — **load-bearing**. Deleting it made the single test fail at `:1666`, `Failures: 1, Errors: 0`. - `:1668` — **cosmetic**. Deleting it left the single test green, `Tests run: 1, Failures: 0`. The remaining 14 are load-bearing waits whose exit condition is a state predicate with an assertion behind it. Neither tell 2 (a predicate already true at entry) nor tell 3 (a timeout logged as the configured value) appears anywhere in this file. Control for the tell-3 search: `grep -n 'elapsed'` returned nothing while `grep -n 'timeout'` returned matches including `:250` and `:397`, so the empty result is a real absence and not a broken pattern. ### For the fleet01 lead You asked for a third independent instance before calling this a pattern. This file yields **one**, measured as above — plus a correction to my own worker's count, which is itself worth knowing: the "assert a negative after a fixed sleep" shape looks like barrier-by-hope and usually is not, because a positive observation earlier in the test rescues it. The mutation that separates them is cheap: disable the subsystem and see who notices. No code was changed. The file's `shasum -a 256` before and after was `5540f302fafa5568595a1c04fcdb9eee2942c52cc8e2ce69395f5e4a742c69f4`, and `git diff --exit-code` returned 0.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#477