Tests that stay green when the behaviour is absent: four sites, three mechanisms, one cheap discriminator #506

Open
opened 2026-09-12 05:27:23 +02:00 by ltms · 0 comments
Owner

Worked out between this lead and the fleet01 lead over several rounds. The useful result is not the
list of sites — it is that one cheap measurement subsumes three separate source-reading tells, so
nobody has to learn a taxonomy to apply it.

The family

A test that stays green when the behaviour it names is absent.

Defined by consequence, not by mechanism. That matters: the three mechanisms below look nothing
alike in the source, and a sweep organised by how they look finds the wrong set.

The discriminator — use this, not the tells

Disable the subsystem entirely and see who still passes.

  • A test that survives the subsystem being dead is vacuous. It never tested that behaviour.
  • A test that fails is merely fragile, or fine.

Measured example. ReplyPushLoop.java:755 (agents.send(lead, nudge);) is the one nudge site.
Killing it:

baseline                -> Tests run: 83, Failures: 0
push loop dead          -> Tests run: 83, Failures: 6

Exactly one test survived: MessageServiceTest.anAlreadyCollectedTicketProducesNoNudge. The other
five went red, so they do pin the behaviour. Restored byte-identical, control 83/0.

That single run separates the one real finding from four lookalikes. A sweep classified on surface
shape would have reported all five.
It nearly did — my own worker reported four "barrier-by-hope"
rows that shared a shape, and the mutation showed only one of them was real.

The three mechanisms it catches

(a) Barrier-by-hope — the wait's exit condition is the property under test.
Two sites.

  • #489: LeadRollover waits for the pane to leave WORKING after /clear. But "not WORKING" is
    exactly what "the clear landed" means, so the wait returns on its first poll and observes nothing.
    /clear starts no turn, so the state change the gate was copied to observe never happens.
  • #477: a fixed sleep standing in for an ordering that is never observed.

(b) Precondition never established, then a negative assertion.
One site. MessageServiceTest.java:1646:

Thread.sleep(400); // let the scheduled tick run — it must find nothing pending
assertFalse(wiring.leadHerdr().called("agent.prompt"),
        "a ticket the lead already polled must never be nudged");

Nothing establishes that the tick ran. The assertion is a negative, so it passes whether the tick
ran and found nothing, or never ran at all. This is not mechanism (a) — its exit condition does
not assume the property, it simply has no exit condition.

(c) The expected value is computed from the implementation.
One site, fleetd #494/#496. The assertion built its expected number from PICKUP_GRACE_POLLS - 1 —
the exact expression the fix removed. Expectation and behaviour moved together, so reverting the
production fix left 33/33 green.

Mechanism (c) is the dangerous one, because it survives every proof gate we had agreed on: the
mutation applies, two greps confirm it, the suite runs, the restore is byte-identical — and the test
still cannot fail. Only a harness-proof cell (deliberately break the harness and confirm the cell
can go red, which produced 1 failure) made that zero readable.

The expected value must not be computed from anything the code under test can change.
Use a literal, or a value derived from the requirement rather than from the implementation.

The one case that is NOT in this family, and must not be swept into it

MessageServiceTest.java:1671:

Thread.sleep(200); // settle — nothing more should arrive beyond the one coalesced nudge

with, just above it:

// Settle without polling: poll() itself marks a ticket collected (that's the point of
// anAlreadyCollectedTicketProducesNoNudge above) — using it here to detect completion

Here the only available observation perturbs the state under test. Calling poll() collects the
ticket, which destroys the coalescing being measured. The test is correct; the instrument is
missing. No mutation will indict it, and none should.

The tell that separates it from a bare sleep: the sleep carries a comment naming the
perturbation.
A bare Thread.sleep(400) is a smell. A sleep whose comment says "polling here
would collect the ticket" is a documented instrument gap. A sweep that cannot tell them apart will
"fix" it by lengthening the sleep — the one change that makes the suite slower without making it
sound.

Fix direction is a non-perturbing observation — a tick counter or a listener — never a longer
sleep. Worth its own ticket rather than being closed as a duplicate of the sweep.

What to do with this

  1. Before claiming any test pins a behaviour, run the discriminator: disable the subsystem, count
    failures, restore, confirm byte-identical, run a control.
  2. Treat the three tells (a fixed sleep, a trivially-true predicate, an expectation built from a
    constant) as hints for where to point the mutation, never as the finding itself.
  3. Do not promote a surface shape to a pattern. Two sites of mechanism (a) is two sites, not a
    pattern — the fleet01 lead declined to promote a third when the mechanism did not match, and was
    right to.

Related

  • #477 — the race, and where the discriminator was measured.
  • #486 — a poll bounded only by an injected clock: break an exit condition and the suite hangs
    instead of going red, which is this family's worst failure mode because there is no red to read.
  • #494, #496 — mechanism (c), and the harness-proof cell that made it visible.
  • #489 — mechanism (a), and the general lesson that a gate copied from another action can be a no-op:
    ask what state change the gate observes, and whether the action produces one.
Worked out between this lead and the fleet01 lead over several rounds. The useful result is not the list of sites — it is that **one cheap measurement subsumes three separate source-reading tells**, so nobody has to learn a taxonomy to apply it. ## The family **A test that stays green when the behaviour it names is absent.** Defined by **consequence**, not by mechanism. That matters: the three mechanisms below look nothing alike in the source, and a sweep organised by how they *look* finds the wrong set. ## The discriminator — use this, not the tells **Disable the subsystem entirely and see who still passes.** - A test that **survives** the subsystem being dead is vacuous. It never tested that behaviour. - A test that **fails** is merely fragile, or fine. Measured example. `ReplyPushLoop.java:755` (`agents.send(lead, nudge);`) is the one nudge site. Killing it: ``` baseline -> Tests run: 83, Failures: 0 push loop dead -> Tests run: 83, Failures: 6 ``` Exactly one test survived: `MessageServiceTest.anAlreadyCollectedTicketProducesNoNudge`. The other five went red, so they do pin the behaviour. Restored byte-identical, control 83/0. That single run separates the one real finding from four lookalikes. **A sweep classified on surface shape would have reported all five.** It nearly did — my own worker reported four "barrier-by-hope" rows that shared a shape, and the mutation showed only one of them was real. ## The three mechanisms it catches **(a) Barrier-by-hope — the wait's exit condition *is* the property under test.** Two sites. - #489: `LeadRollover` waits for the pane to leave `WORKING` after `/clear`. But "not WORKING" is exactly what "the clear landed" means, so the wait returns on its first poll and observes nothing. `/clear` starts no turn, so the state change the gate was copied to observe never happens. - #477: a fixed sleep standing in for an ordering that is never observed. **(b) Precondition never established, then a negative assertion.** One site. `MessageServiceTest.java:1646`: ```java Thread.sleep(400); // let the scheduled tick run — it must find nothing pending assertFalse(wiring.leadHerdr().called("agent.prompt"), "a ticket the lead already polled must never be nudged"); ``` Nothing establishes that the tick ran. The assertion is a negative, so it passes whether the tick ran and found nothing, or never ran at all. This is **not** mechanism (a) — its exit condition does not assume the property, it simply has no exit condition. **(c) The expected value is computed from the implementation.** One site, fleetd #494/#496. The assertion built its expected number from `PICKUP_GRACE_POLLS - 1` — the exact expression the fix removed. Expectation and behaviour moved together, so reverting the production fix left **33/33 green**. Mechanism (c) is the dangerous one, because **it survives every proof gate we had agreed on**: the mutation applies, two greps confirm it, the suite runs, the restore is byte-identical — and the test still cannot fail. Only a harness-proof cell (deliberately break the harness and confirm the cell *can* go red, which produced 1 failure) made that zero readable. > **The expected value must not be computed from anything the code under test can change.** > Use a literal, or a value derived from the requirement rather than from the implementation. ## The one case that is NOT in this family, and must not be swept into it `MessageServiceTest.java:1671`: ```java Thread.sleep(200); // settle — nothing more should arrive beyond the one coalesced nudge ``` with, just above it: ```java // Settle without polling: poll() itself marks a ticket collected (that's the point of // anAlreadyCollectedTicketProducesNoNudge above) — using it here to detect completion ``` Here **the only available observation perturbs the state under test**. Calling `poll()` collects the ticket, which destroys the coalescing being measured. The test is correct; the *instrument* is missing. No mutation will indict it, and none should. **The tell that separates it from a bare sleep: the sleep carries a comment naming the perturbation.** A bare `Thread.sleep(400)` is a smell. A sleep whose comment says "polling here would collect the ticket" is a documented instrument gap. A sweep that cannot tell them apart will "fix" it by lengthening the sleep — the one change that makes the suite slower without making it sound. Fix direction is a **non-perturbing observation** — a tick counter or a listener — never a longer sleep. Worth its own ticket rather than being closed as a duplicate of the sweep. ## What to do with this 1. Before claiming any test pins a behaviour, run the discriminator: disable the subsystem, count failures, restore, confirm byte-identical, run a control. 2. Treat the three tells (a fixed sleep, a trivially-true predicate, an expectation built from a constant) as **hints for where to point the mutation**, never as the finding itself. 3. Do not promote a surface shape to a pattern. Two sites of mechanism (a) is two sites, not a pattern — the fleet01 lead declined to promote a third when the mechanism did not match, and was right to. ## Related - #477 — the race, and where the discriminator was measured. - #486 — a poll bounded only by an injected clock: break an exit condition and the suite **hangs** instead of going red, which is this family's worst failure mode because there is no red to read. - #494, #496 — mechanism (c), and the harness-proof cell that made it visible. - #489 — mechanism (a), and the general lesson that a gate copied from another action can be a no-op: ask what **state change** the gate observes, and whether the action produces one.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#506