fleetd #449: name the decisive cell and the load the claim was measured at
The previous comment named a cause with no cell behind it. fleet01's review made the point: a confident wrong mechanism gets copied, and a confident under-determined one gets copied the same way. So the comment now names the cell that settles it. Hold the old 1000ms write sleep and change only the read - 800ms fixed sleep becomes a 5s poll - and the test goes 0 of 3 to 3 of 3. The read deadline was the whole story. It also says where: a 12-core macOS host near idle (load 2.6 to 5.9). The loaded run agreed, but its load climbed from 7 to 50 while the cells ran and the old version ran last, so it is not clean evidence and the comment says so. Comment only. No test or production code changed.
This commit is contained in:
@@ -25,13 +25,23 @@ import static org.junit.jupiter.api.Assumptions.assumeTrue;
|
|||||||
* So this polls for a real signal instead of guessing a sleep length.
|
* So this polls for a real signal instead of guessing a sleep length.
|
||||||
*
|
*
|
||||||
* <p><strong>What was measured, and what was not.</strong> Polling fixes it: 5 standalone runs
|
* <p><strong>What was measured, and what was not.</strong> Polling fixes it: 5 standalone runs
|
||||||
* green. The load-bearing half is {@link #waitForText}. With {@link #SHELL_READY_TIMEOUT_MS}
|
* green. The load-bearing half is {@link #waitForText}, and one cell proves it. Keep the old
|
||||||
* set to 0 — so input is typed at once, with no settle wait at all — the test still passed 3 of
|
* 1000ms write sleep and change only the read — the 800ms fixed sleep becomes a 5s poll — and
|
||||||
* 3. So the proven cause is the 800ms READ deadline being too short, not the 1000ms write delay.
|
* the test goes from 0 of 3 passing to 3 of 3. Removing the write wait instead
|
||||||
* Note the direction, because it matters: typing at 0ms works where typing at 1000ms failed. The
|
* ({@link #SHELL_READY_TIMEOUT_MS} set to 0, so input is typed at once) also passes 3 of 3. So
|
||||||
* earlier explanation for this test — that input typed before the prompt is swallowed by the
|
* the cause is the 800ms READ deadline, not the 1000ms write delay. The old version fails every
|
||||||
* shell's startup — is therefore NOT supported by any measurement here. Please do not repeat it
|
* time, not sometimes, so "race" is the wrong word for it. The earlier explanation — that input
|
||||||
* as the reason; if it were true, 0ms would be worse than 1000ms, and it is better.
|
* typed before the prompt is swallowed by the shell's startup — is not supported by anything
|
||||||
|
* measured here. Please do not repeat it: if it were true, typing at 0ms would be worse than
|
||||||
|
* typing at 1000ms, and it is not.
|
||||||
|
*
|
||||||
|
* <p><strong>Where this was measured.</strong> A 12-core macOS host, load average 2.6 to 5.9,
|
||||||
|
* on commit 20c1094. The same four cells were also run under load and gave the same answer, but
|
||||||
|
* that run is not clean evidence: the load average climbed from 7 to 50 while the cells ran, and
|
||||||
|
* the old version ran last, at the top of that climb. Above about load 20 everything here is
|
||||||
|
* slow for reasons that have nothing to do with this seam. So read the claim as "measured near
|
||||||
|
* idle on a 12-core host", and nothing stronger. If this test fails on a smaller or busier
|
||||||
|
* machine, raise {@link #OUTPUT_TIMEOUT_MS} before you suspect the seam.
|
||||||
*
|
*
|
||||||
* <p>{@link #waitUntilSettled} is kept as cheap insurance against that swallow case, not because
|
* <p>{@link #waitUntilSettled} is kept as cheap insurance against that swallow case, not because
|
||||||
* anyone showed it was needed. If you want to delete it, the honest test is whether you can make
|
* anyone showed it was needed. If you want to delete it, the honest test is whether you can make
|
||||||
|
|||||||
Reference in New Issue
Block a user