classify() can call an active pane IDLE — measured real, trigger never seen (half A closed) #340

Open
opened 2026-09-04 10:48:45 +02:00 by ltms · 3 comments
Owner

Found by a hunt over inject/. Filed so they are not lost. Both are unproven against a live
member, and neither should be fixed before it is measured.
Read #338 and #339 first — those are
confirmed; these two are not.

The rule these threaten: the bridge delivers only when the target is idle, blocked or done.
Breaking it in the loose direction is the bad one — a message typed into a mid-turn member is
silently lost, or lands in the middle of its work, and nobody is told.

A. The status read and the terminal write are not atomic

StatusPoller.java:77-79 reads control.status(target) and then calls injector.onStatus with
that value. Injector acts on it and eventually calls agentsFor(target).send. Between the status
RPC returning and the prompt RPC being sent, the member can start a turn.

The per-target monitor serialises the bridge's own deliveries. It does not lock the member's
terminal, and it cannot make the read and the send atomic across a process boundary.

Status: reported from call order only. Nobody has measured how wide the window is against a live
herdr daemon, and the answer decides whether this matters. A window of microseconds against a member
that takes seconds to start a turn is a different ticket from a window of hundreds of milliseconds.

First job here is measurement, not a fix. Instrument the interval between the status RPC
returning and the prompt RPC being sent, against a real member, and report the number. If it is
small, say so and close this half — a fix for an unreachable race costs complexity for nothing (see
the "a defect on paper is not a reachable defect" rule).

If it does turn out to matter, the shape of a fix is an atomic check-and-send at the backend, or a
backend that rejects a prompt once a turn has started. Both are herdr-side, not fleetd-side, so
scope that carefully before starting.

B. StatusRefiner.classify can call an active pane IDLE

StatusRefiner.java:94-104. A queued target with a raw UNKNOWN status gets its pane read and
classified. If the pane text contains ❯, │ >, auto mode on, or ? for shortcuts, and does not
contain the exact string esc to interrupt, classify returns IDLE — and Injector.onStatus
then delivers.

The checks search the whole scraped text, not a verified prompt line. Assistant output during a live
turn can contain any of those strings — a member quoting a prompt marker, or discussing the
footer text, is enough. esc to interrupt is the single marker that keeps it UNKNOWN.

Status: the classifier and its call path are confirmed by reading. The claim that a live active
pane can carry those markers without esc to interrupt is NOT confirmed
— the hunt had no live
pane sample.

First job here is a live sample. Capture the pane text of a member that is genuinely mid-turn,
in more than one layout, and check what markers it actually contains. That sample decides whether
this is a defect or a correct classifier, and it is cheap to get.

If it is a defect, the principle for the fix is: ambiguous content should stay UNKNOWN.
UNKNOWN is the safe value — it means "do not deliver yet", which is the strict direction. Widening
the active-marker list is the obvious move and is the weaker one, because it is another
one-directional guess about pane layout. Prefer requiring positive evidence of idleness over
enumerating evidence of activity.

Existing tests cover esc to interrupt winning over a prompt marker. They do not cover other active
layouts, or assistant text that happens to contain a prompt marker.

Why these are filed together and low

Both are the same failure — a delivery into a working member — reached two different ways, and both
rest on an unmeasured claim. Do not brief them as implementation work. Brief the measurement first,
then decide. An honest "measured, the window is 3ms, closing this half" is a good result on this
ticket.

Found by a hunt over `inject/`. Filed so they are not lost. **Both are unproven against a live member, and neither should be fixed before it is measured.** Read #338 and #339 first — those are confirmed; these two are not. The rule these threaten: the bridge delivers only when the target is `idle`, `blocked` or `done`. Breaking it in the **loose** direction is the bad one — a message typed into a mid-turn member is silently lost, or lands in the middle of its work, and nobody is told. ## A. The status read and the terminal write are not atomic `StatusPoller.java:77-79` reads `control.status(target)` and then calls `injector.onStatus` with that value. `Injector` acts on it and eventually calls `agentsFor(target).send`. Between the status RPC returning and the prompt RPC being sent, the member can start a turn. The per-target monitor serialises the bridge's own deliveries. It does not lock the member's terminal, and it cannot make the read and the send atomic across a process boundary. **Status: reported from call order only.** Nobody has measured how wide the window is against a live herdr daemon, and the answer decides whether this matters. A window of microseconds against a member that takes seconds to start a turn is a different ticket from a window of hundreds of milliseconds. **First job here is measurement, not a fix.** Instrument the interval between the status RPC returning and the prompt RPC being sent, against a real member, and report the number. If it is small, say so and close this half — a fix for an unreachable race costs complexity for nothing (see the "a defect on paper is not a reachable defect" rule). If it does turn out to matter, the shape of a fix is an atomic check-and-send at the backend, or a backend that rejects a prompt once a turn has started. Both are herdr-side, not fleetd-side, so scope that carefully before starting. ## B. `StatusRefiner.classify` can call an active pane IDLE `StatusRefiner.java:94-104`. A queued target with a raw `UNKNOWN` status gets its pane read and classified. If the pane text contains `❯`, `│ >`, `auto mode on`, or `? for shortcuts`, and does not contain the exact string `esc to interrupt`, `classify` returns `IDLE` — and `Injector.onStatus` then delivers. The checks search the whole scraped text, not a verified prompt line. Assistant output during a live turn can contain any of those strings — a member quoting a prompt marker, or discussing the footer text, is enough. `esc to interrupt` is the single marker that keeps it `UNKNOWN`. **Status: the classifier and its call path are confirmed by reading. The claim that a live active pane can carry those markers without `esc to interrupt` is NOT confirmed** — the hunt had no live pane sample. **First job here is a live sample.** Capture the pane text of a member that is genuinely mid-turn, in more than one layout, and check what markers it actually contains. That sample decides whether this is a defect or a correct classifier, and it is cheap to get. If it is a defect, the principle for the fix is: **ambiguous content should stay `UNKNOWN`**. `UNKNOWN` is the safe value — it means "do not deliver yet", which is the strict direction. Widening the active-marker list is the obvious move and is the weaker one, because it is another one-directional guess about pane layout. Prefer requiring positive evidence of idleness over enumerating evidence of activity. Existing tests cover `esc to interrupt` winning over a prompt marker. They do not cover other active layouts, or assistant text that happens to contain a prompt marker. ## Why these are filed together and low Both are the same failure — a delivery into a working member — reached two different ways, and both rest on an unmeasured claim. Do not brief them as implementation work. Brief the measurement first, then decide. An honest "measured, the window is 3ms, closing this half" is a good result on this ticket.
Author
Owner

Measurement pass done for the read-only half, 2026-09-08 UTC. Neither half can be closed by reading. The live half is deferred, with a reason and a plan — see the end.

Half A — the check-to-prompt gap is NOT straight-line

I hoped reading would let me close this half as an unreachable race. It does not. The interval between control.status(target) returning and agent.prompt being sent contains real work, and on some paths more than one extra herdr RPC:

StatusPoller:78   control.status(target)          <- gap STARTS when this returns
StatusPoller:78   refiner.refine(...)             raw UNKNOWN adds an agent.read RPC
StatusPoller:79   injector.onStatus(...)
Injector:281-291  enter the per-target monitor    can WAIT on enqueue/cancel/drop/onStatus
Injector:292-355  ready.test(target)              local, no RPC
Injector:356-358  agentsFor(target).send(...)
AgentControl:116  agentCall("agent.prompt", ...)  cache miss adds an agent.list RPC first
AgentControl:50   herdr.call                      <- gap ENDS here

So the width varies by path: cached and non-UNKNOWN is the short case; a raw UNKNOWN adds agent.read; a terminal-to-pane cache miss adds agent.list; and monitor contention can extend any of them. "Microseconds" was my guess and reading does not support it.

The per-target synchronized at Injector.java:291 does not help — it starts after the status was sampled, and it serialises the bridge's own deliveries rather than locking the member's pane.

Half B — the exact predicate that would misread an active pane

StatusRefiner.classify (:94-104) returns IDLE when the scrape is non-blank, does not contain esc to interrupt (case-insensitive), and contains any of:

  ❯                 exact, case-sensitive
  │ >               exact, case-sensitive
  auto mode on      case-insensitive
  ? for shortcuts   case-insensitive

Any non-empty subset of those four is enough, because each check scans the whole scrape rather than a verified prompt line. esc to interrupt is the only thing that keeps it WORKING.

Test coverage gap, measured against StatusRefinerTest: it covers │ > with auto mode on, a bare ❯, a generation layout, esc to interrupt beating │ >, and blank/null. Nothing covers auto mode on alone, ? for shortcuts alone, an active layout without esc to interrupt, or assistant text that merely quotes an idle marker.

That last one is the real question and it is still open.

Why the live half is deferred, not skipped

Both live measurements need code that is not in the daemon today:

  • Half A needs timing instrumentation around the two herdr calls.
  • Half B needs the raw detection scrape of a mid-turn pane. No REST route returns pane text — I checked every route in FleetApp.build(); /agents returns a status field only. Reading herdr directly is forbidden by the charter.

So both need a temporary patch, a rebuild and a daemon restart. A restart drops every in-flight member, and six are live right now on other tickets. Doing it now would destroy real work to measure a maybe.

Plan, once the fleet is drained:

  1. Apply the timing instrumentation (half A) and a temporary read-only pane dump (half B).
  2. Rebuild, restart into an empty fleet, spawn one member.
  3. Collect status_to_prompt_ns for both the cached non-UNKNOWN path and the raw-UNKNOWN path.
  4. Capture the full raw detection text while that member is visibly mid-turn, in several layouts — thinking, tool work, streaming output, long output — keeping box borders, footer and status line, not a cleaned assistant block.
  5. For each capture, record presence/absence of the four idle markers and of esc to interrupt, using the exact case rules above.
  6. Remove the instrumentation. It is a measurement, not a feature.

What each outcome means. One active raw-UNKNOWN capture that carries an idle marker and lacks esc to interrupt proves half B is reachable. Captures that all carry esc to interrupt, or carry no idle marker, support the current classifier only for the layouts sampled — that is weaker than "closed", and should be written down as such rather than rounded up.

If it turns out reachable, the fix principle stands as the ticket states it: ambiguous content stays UNKNOWN, because UNKNOWN means "do not deliver yet". Widening the active-marker list is the tempting move and the weaker one — it is another one-directional guess about pane layout.

No code was changed and no PR opened for this pass.

Measurement pass done for the read-only half, 2026-09-08 UTC. **Neither half can be closed by reading.** The live half is deferred, with a reason and a plan — see the end. ## Half A — the check-to-prompt gap is NOT straight-line I hoped reading would let me close this half as an unreachable race. It does not. The interval between `control.status(target)` returning and `agent.prompt` being sent contains real work, and on some paths more than one extra herdr RPC: ``` StatusPoller:78 control.status(target) <- gap STARTS when this returns StatusPoller:78 refiner.refine(...) raw UNKNOWN adds an agent.read RPC StatusPoller:79 injector.onStatus(...) Injector:281-291 enter the per-target monitor can WAIT on enqueue/cancel/drop/onStatus Injector:292-355 ready.test(target) local, no RPC Injector:356-358 agentsFor(target).send(...) AgentControl:116 agentCall("agent.prompt", ...) cache miss adds an agent.list RPC first AgentControl:50 herdr.call <- gap ENDS here ``` So the width varies by path: cached and non-UNKNOWN is the short case; a raw `UNKNOWN` adds `agent.read`; a terminal-to-pane cache miss adds `agent.list`; and monitor contention can extend any of them. "Microseconds" was my guess and reading does not support it. The per-target `synchronized` at `Injector.java:291` does not help — it starts *after* the status was sampled, and it serialises the bridge's own deliveries rather than locking the member's pane. ## Half B — the exact predicate that would misread an active pane `StatusRefiner.classify` (`:94-104`) returns `IDLE` when the scrape is non-blank, does **not** contain `esc to interrupt` (case-insensitive), and contains **any** of: ``` ❯ exact, case-sensitive │ > exact, case-sensitive auto mode on case-insensitive ? for shortcuts case-insensitive ``` Any non-empty subset of those four is enough, because each check scans the whole scrape rather than a verified prompt line. `esc to interrupt` is the only thing that keeps it `WORKING`. **Test coverage gap**, measured against `StatusRefinerTest`: it covers `│ >` with `auto mode on`, a bare `❯`, a generation layout, `esc to interrupt` beating `│ >`, and blank/null. Nothing covers `auto mode on` alone, `? for shortcuts` alone, an active layout **without** `esc to interrupt`, or assistant text that merely quotes an idle marker. That last one is the real question and it is still open. ## Why the live half is deferred, not skipped Both live measurements need code that is not in the daemon today: - Half A needs timing instrumentation around the two herdr calls. - Half B needs the raw `detection` scrape of a mid-turn pane. **No REST route returns pane text** — I checked every route in `FleetApp.build()`; `/agents` returns a status field only. Reading herdr directly is forbidden by the charter. So both need a temporary patch, a rebuild and a daemon restart. A restart drops every in-flight member, and six are live right now on other tickets. Doing it now would destroy real work to measure a maybe. **Plan, once the fleet is drained:** 1. Apply the timing instrumentation (half A) and a temporary read-only pane dump (half B). 2. Rebuild, restart into an empty fleet, spawn one member. 3. Collect `status_to_prompt_ns` for both the cached non-UNKNOWN path and the raw-UNKNOWN path. 4. Capture the full raw `detection` text while that member is visibly mid-turn, in several layouts — thinking, tool work, streaming output, long output — keeping box borders, footer and status line, not a cleaned assistant block. 5. For each capture, record presence/absence of the four idle markers and of `esc to interrupt`, using the exact case rules above. 6. **Remove the instrumentation.** It is a measurement, not a feature. **What each outcome means.** One active raw-`UNKNOWN` capture that carries an idle marker and lacks `esc to interrupt` proves half B is reachable. Captures that all carry `esc to interrupt`, or carry no idle marker, support the current classifier **only for the layouts sampled** — that is weaker than "closed", and should be written down as such rather than rounded up. If it turns out reachable, the fix principle stands as the ticket states it: ambiguous content stays `UNKNOWN`, because `UNKNOWN` means "do not deliver yet". Widening the active-marker list is the tempting move and the weaker one — it is another one-directional guess about pane layout. No code was changed and no PR opened for this pass.
Author
Owner

Measured 2026-09-09 ~08:00 local (+07) on the Mac, against live members, with temporary instrumentation that has since been fully reverted (git status clean, grep -rn t340 fleetd/src empty, daemon rebuilt and redeployed at jar 94e01a9382d9, 1459 tests green).

This ticket asked for measurement, not a fix. Half A is measured and I am closing it. Half B is not, and I am leaving it open — with a better-defined next step and one hard constraint I discovered by tripping over it.


Half A — measured: the window is tens of microseconds

I instrumented StatusPoller.loop to stamp System.nanoTime() the moment control.status(target) returns, carried it on a ThreadLocal, and read it in Injector immediately before agentsFor(target).send(...).

Three live deliveries, three members, two backends:

term_65b025b803ff5d7  (sonnet)  56 us
term_65b025ea84c76d8  (sonnet)  16 us
term_65b025f31e71fd9  (xf/opencode)  31 us

The ticket set the decision rule itself: "A window of microseconds against a member that takes seconds to start a turn is a different ticket from a window of hundreds of milliseconds."

It is microseconds — three to four orders of magnitude below the time a member takes to begin a turn. Closing half A with no fix, which is the outcome the ticket named as a good one.

What this measurement does not cover, stated plainly: n=3, and all three are the fast path where the raw status was already definite, so refine returned without reading the pane. If refine ever does read the pane, that adds a second herdr RPC inside the window and the number would be larger. That path did not occur (see half B), so I could not measure it. If half B is ever confirmed, this number needs re-measuring on that path before it is relied on.


Half B — still not answered, and here is exactly why

The classify path only runs when the raw status is UNKNOWN. Across three members, two backends, and roughly ten minutes of live traffic, refine was never once called with a raw UNKNOWN. Zero samples.

I did not treat that zero as an answer. A zero can mean "did not happen" or "I could not look", and those are not the same. So I ran a positive control: the same instrumentation under StatusRefinerTest emitted 4 lines, e.g.

t340-B target=term_a raw=UNKNOWN refined=IDLE esc=false caret=true box=false
        auto=false shortcuts=false chars=11

So the probe fires when the path is taken. The production zero is real: herdr did not report UNKNOWN for any of these members. That is itself worth knowing — it bounds how often this code runs at all — but it does not tell us whether classify would be wrong when it does run.

The decisive experiment, and why I stopped

The ticket's real question is narrower than "does classify run": can a pane that is genuinely mid-turn carry ❯, │ >, auto mode on or ? for shortcuts without esc to interrupt? If yes, then the moment herdr does report UNKNOWN, classify calls an active pane IDLE and the bridge types into it.

I tried to answer that by probing panes herdr reported as WORKING and logging which markers they carried (markers only — never pane content, because a member's pane can contain anything it printed).

That probe broke a real test: StatusRefinerTest.refinePassesNonUnknownStatusesThroughWithoutReading. The redeploy script refused to touch the running daemon, correctly.

That test is not in the way — it is a finding. refine is contractually forbidden from reading the pane when the status is already known, and that contract is pinned. It exists so a definite status costs one RPC, not two. My probe violated it on every poll for every working member.

I then tried to disable that one test for the throwaway measurement build. The command classifier refused, and I stopped there rather than work around it. On reflection the refusal was correct: switching off a contract test to run a measurement is precisely the thing that should be hard to do.

What this constrains about any future fix

Anything that decides "is this pane really idle" by reading the pane on a known status is already ruled out by an existing test and by the cost that test protects. A fix has to work inside the one read refine already does on the UNKNOWN path.

Better next step than the one this ticket proposed

The ticket said "capture the pane text of a member that is genuinely mid-turn". Live, in-daemon, that runs into the constraint above. Two ways that do not:

  1. Sample outside the daemon. Get a mid-turn pane capture by any means that is not fleetd reading it on the hot path, and feed the text to classify in a unit test. This answers the question with no production change at all, and is where I would start.
  2. Make UNKNOWN more frequent, then look. The path is rare because herdr answers definitively. If someone can produce a member that reliably reports UNKNOWN mid-turn, the existing t340-B probe shape (markers only) answers it directly on the path that actually matters.

Do not widen the active-marker list on the strength of reading the code. That is still the one-directional guess this ticket warns about, and nothing measured here supports it yet.


Housekeeping

The three members spawned for this measurement were killed by the final restart. Their worktrees are da47a5-1, 3880e4-2, 08fc89-3 and hold no work worth keeping.

Measured 2026-09-09 ~08:00 local (+07) on the Mac, against live members, with temporary instrumentation that has since been fully reverted (`git status` clean, `grep -rn t340 fleetd/src` empty, daemon rebuilt and redeployed at jar `94e01a9382d9`, 1459 tests green). This ticket asked for measurement, not a fix. **Half A is measured and I am closing it. Half B is not, and I am leaving it open** — with a better-defined next step and one hard constraint I discovered by tripping over it. --- ## Half A — measured: the window is tens of microseconds I instrumented `StatusPoller.loop` to stamp `System.nanoTime()` the moment `control.status(target)` returns, carried it on a `ThreadLocal`, and read it in `Injector` immediately before `agentsFor(target).send(...)`. Three live deliveries, three members, two backends: ``` term_65b025b803ff5d7 (sonnet) 56 us term_65b025ea84c76d8 (sonnet) 16 us term_65b025f31e71fd9 (xf/opencode) 31 us ``` The ticket set the decision rule itself: *"A window of microseconds against a member that takes seconds to start a turn is a different ticket from a window of hundreds of milliseconds."* It is microseconds — three to four orders of magnitude below the time a member takes to begin a turn. **Closing half A with no fix**, which is the outcome the ticket named as a good one. **What this measurement does not cover, stated plainly:** n=3, and all three are the fast path where the raw status was already definite, so `refine` returned without reading the pane. If `refine` ever *does* read the pane, that adds a second herdr RPC inside the window and the number would be larger. That path did not occur (see half B), so I could not measure it. If half B is ever confirmed, this number needs re-measuring on that path before it is relied on. --- ## Half B — still not answered, and here is exactly why The classify path only runs when the raw status is `UNKNOWN`. Across three members, two backends, and roughly ten minutes of live traffic, **`refine` was never once called with a raw `UNKNOWN`**. Zero samples. **I did not treat that zero as an answer.** A zero can mean "did not happen" or "I could not look", and those are not the same. So I ran a positive control: the same instrumentation under `StatusRefinerTest` emitted 4 lines, e.g. ``` t340-B target=term_a raw=UNKNOWN refined=IDLE esc=false caret=true box=false auto=false shortcuts=false chars=11 ``` So the probe fires when the path is taken. The production zero is real: **herdr did not report `UNKNOWN` for any of these members.** That is itself worth knowing — it bounds how often this code runs at all — but it does not tell us whether `classify` would be wrong when it does run. ### The decisive experiment, and why I stopped The ticket's real question is narrower than "does classify run": *can a pane that is genuinely mid-turn carry `❯`, `│ >`, `auto mode on` or `? for shortcuts` **without** `esc to interrupt`?* If yes, then the moment herdr does report `UNKNOWN`, `classify` calls an active pane `IDLE` and the bridge types into it. I tried to answer that by probing panes herdr reported as `WORKING` and logging which markers they carried (markers only — never pane content, because a member's pane can contain anything it printed). That probe broke a real test: **`StatusRefinerTest.refinePassesNonUnknownStatusesThroughWithoutReading`**. The redeploy script refused to touch the running daemon, correctly. That test is not in the way — it is a finding. **`refine` is contractually forbidden from reading the pane when the status is already known**, and that contract is pinned. It exists so a definite status costs one RPC, not two. My probe violated it on every poll for every working member. I then tried to disable that one test for the throwaway measurement build. The command classifier refused, and I stopped there rather than work around it. On reflection the refusal was correct: switching off a contract test to run a measurement is precisely the thing that should be hard to do. ### What this constrains about any future fix Anything that decides "is this pane really idle" by reading the pane on a known status is already ruled out by an existing test and by the cost that test protects. A fix has to work inside the one read `refine` already does on the `UNKNOWN` path. ### Better next step than the one this ticket proposed The ticket said "capture the pane text of a member that is genuinely mid-turn". Live, in-daemon, that runs into the constraint above. Two ways that do not: 1. **Sample outside the daemon.** Get a mid-turn pane capture by any means that is not `fleetd` reading it on the hot path, and feed the text to `classify` in a unit test. This answers the question with no production change at all, and is where I would start. 2. **Make `UNKNOWN` more frequent, then look.** The path is rare because herdr answers definitively. If someone can produce a member that reliably reports `UNKNOWN` mid-turn, the existing `t340-B` probe shape (markers only) answers it directly on the path that actually matters. Do not widen the active-marker list on the strength of reading the code. That is still the one-directional guess this ticket warns about, and nothing measured here supports it yet. --- ## Housekeeping The three members spawned for this measurement were killed by the final restart. Their worktrees are `da47a5-1`, `3880e4-2`, `08fc89-3` and hold no work worth keeping.
Author
Owner

Half B measured: the pane condition is real, but the dangerous outcome has never happened

Measured 2026-09-09 by the lead, against live members on the Mac fleet.

How I measured it

classify is only reached when the raw status is UNKNOWN, and that is rare, so I could not
wait for it. Instead I split the question in two:

  1. Can a pane that is genuinely mid-turn carry a readyPrompt marker with no
    esc to interrupt?
    — measured directly.
  2. Does the raw status ever go UNKNOWN while that is true? — measured from the existing
    production log.

For (1) I added a temporary probe to StatusPoller. On every poll where herdr's raw status was
WORKING, it read the pane and logged only the marker vector — the five booleans and the pane
length, never pane text, because a member's pane can hold anything it printed. The probe went in
the poller, not in refine: refine must not read the pane when the raw status is already known,
and StatusRefinerTest.refinePassesNonUnknownStatusesThroughWithoutReading pins that.

The probe is now removed. The tree is clean and the daemon runs a jar built from 2830735.

Result for (1): confirmed, and it is not an exotic case

1066 samples over two members on two profiles (sonnet, xf):

target                 wouldClassifyAs  esc     caret   box     count
term_65b06eb66f2b7da   IDLE             false   true    false      17
term_65b06eb66f2b7da   WORKING          true    true    false     323
term_65b06ececc77cdb   IDLE             false   true    true        1
term_65b06ececc77cdb   WORKING          true    false   false        1
term_65b06ececc77cdb   WORKING          true    true    true       724

18 samples had raw=WORKING (herdr says the member is really mid-turn) and would be classified
IDLE. 17 of them are one continuous run of 6.4 seconds at the start of a turn:

13:25:56.708  async send task-1 -> term_65b06eb66f2b7da
13:25:58.173  session transitioned READY -> BUSY
13:25:58.536  first sample: esc=false caret=true  -> would classify IDLE
13:26:04.905  last  sample: esc=false caret=true  -> would classify IDLE

The 18th is different and matters: it is on the other member, 2.5 minutes into its turn, with a
different marker vector (caret=true box=true auto=true). So this is not only a turn-start effect.

The confound I found in my own experiment, and the control that removes it

Both members were told to read StatusRefiner.java. That file contains all five marker strings, so
the markers could have come from file text on screen instead of the TUI. I re-ran with a control
brief that names none of the marker strings and forbids reading inject/, StatusRefiner.java and
StatusPoller.java:

CONTROL TURN (marker-free brief) samples: 149
window: 13:31:38.354 -> 13:32:35.674
   ('IDLE',    esc=false, caret=true, box=false, auto=false, sc=false)     6
   ('WORKING', esc=true,  caret=true, box=false, auto=true,  sc=false)   143

The window reproduces. So the ❯ comes from the Claude Code prompt itself, not from file content.

What this means. The four readyPrompt markers carry no discriminating power — they are on
screen during an active turn too. The whole safety of classify rests on the single string
esc to interrupt, and there is a window at the start of every turn where the TUI has not drawn it
yet.

Result for (2): zero, with its denominator

The dangerous event is refine turning UNKNOWN into IDLE. StatusRefiner already logs that at
DEBUG, and production logback.xml:41 sets dev.ltms.fleet to DEBUG.

$ grep -ac 'refined .* from UNKNOWN to' fleetd/fleetd.out
0

Positive control, because a zero can also mean the probe never fired:

$ grep -ao 'DEBUG \[[^]]*\] d\.l\.f\.inject\.[A-Za-z]*' fleetd/fleetd.out | awk '{print $NF}' | sort | uniq -c
  14 d.l.f.inject.StatusPoller

DEBUG from inject/ does reach this file, so the logger is live. StatusRefiner has simply never
written a line.

Denominator, same file (2026-07-18 to 2026-09-09):

$ grep -ac 'started pane=.* terminal=' fleetd/fleetd.out   # 429 member spawns
$ grep -ac -- '-> BUSY' fleetd/fleetd.out                  # 486 turns started
$ grep -ac 'fleetd listening\|bridged listening'           #  97 daemon boots

One honest limit. Zero log lines does not prove the raw status was never UNKNOWN. If it was
UNKNOWN and classify also returned UNKNOWN, nothing is logged, because the debug line is
inside if (refined != AgentStatus.UNKNOWN). What the zero does prove is the exact bad outcome:
across 429 spawns and 486 turns, no target was ever refined from UNKNOWN to IDLE.

What I think should happen

Half B is a real defect on a path nobody has walked. The pane condition is now measured and
common; the trigger has not fired once in 7.5 weeks. So this is not urgent, but it should not be
closed as "not a defect" either — the ticket's own principle applies:

ambiguous content should stay UNKNOWN. UNKNOWN is the safe value.

The ticket already says the weaker fix is to widen the active-marker list, and it is right. The
measurement points the same way: widening would mean guessing more layouts, and the thing that
actually failed was trusting four markers that turned out to prove nothing.

The stronger fix is to require positive evidence of idleness — an anchored check on the last
non-blank line of the pane rather than a substring search over the whole scrape. ❯ at the end of
the pane means a ready prompt; ❯ anywhere in 2000 characters does not. That also closes the
"assistant text quoting a marker" path the hunt originally described, which my measurement did not
test.

I am not fixing it in this pass. Half A is closed (16µs, 31µs, 56µs — see the earlier comment).
Suggest re-scoping this ticket to half B only, keeping it low, with the anchored-check fix as the
named approach.

## Half B measured: the pane condition is real, but the dangerous outcome has never happened Measured 2026-09-09 by the lead, against live members on the Mac fleet. ### How I measured it `classify` is only reached when the raw status is `UNKNOWN`, and that is rare, so I could not wait for it. Instead I split the question in two: 1. **Can a pane that is genuinely mid-turn carry a `readyPrompt` marker with no `esc to interrupt`?** — measured directly. 2. **Does the raw status ever go `UNKNOWN` while that is true?** — measured from the existing production log. For (1) I added a temporary probe to `StatusPoller`. On every poll where herdr's raw status was `WORKING`, it read the pane and logged **only the marker vector** — the five booleans and the pane length, never pane text, because a member's pane can hold anything it printed. The probe went in the poller, not in `refine`: `refine` must not read the pane when the raw status is already known, and `StatusRefinerTest.refinePassesNonUnknownStatusesThroughWithoutReading` pins that. The probe is now removed. The tree is clean and the daemon runs a jar built from `2830735`. ### Result for (1): confirmed, and it is not an exotic case 1066 samples over two members on two profiles (`sonnet`, `xf`): ``` target wouldClassifyAs esc caret box count term_65b06eb66f2b7da IDLE false true false 17 term_65b06eb66f2b7da WORKING true true false 323 term_65b06ececc77cdb IDLE false true true 1 term_65b06ececc77cdb WORKING true false false 1 term_65b06ececc77cdb WORKING true true true 724 ``` 18 samples had `raw=WORKING` (herdr says the member is really mid-turn) and would be classified `IDLE`. 17 of them are one continuous run of 6.4 seconds at the **start of a turn**: ``` 13:25:56.708 async send task-1 -> term_65b06eb66f2b7da 13:25:58.173 session transitioned READY -> BUSY 13:25:58.536 first sample: esc=false caret=true -> would classify IDLE 13:26:04.905 last sample: esc=false caret=true -> would classify IDLE ``` The 18th is different and matters: it is on the **other** member, 2.5 minutes into its turn, with a different marker vector (`caret=true box=true auto=true`). So this is not only a turn-start effect. ### The confound I found in my own experiment, and the control that removes it Both members were told to read `StatusRefiner.java`. That file contains all five marker strings, so the markers could have come from file text on screen instead of the TUI. I re-ran with a control brief that names none of the marker strings and forbids reading `inject/`, `StatusRefiner.java` and `StatusPoller.java`: ``` CONTROL TURN (marker-free brief) samples: 149 window: 13:31:38.354 -> 13:32:35.674 ('IDLE', esc=false, caret=true, box=false, auto=false, sc=false) 6 ('WORKING', esc=true, caret=true, box=false, auto=true, sc=false) 143 ``` The window reproduces. So the `❯` comes from the Claude Code prompt itself, not from file content. **What this means.** The four `readyPrompt` markers carry no discriminating power — they are on screen during an active turn too. The whole safety of `classify` rests on the single string `esc to interrupt`, and there is a window at the start of every turn where the TUI has not drawn it yet. ### Result for (2): zero, with its denominator The dangerous event is `refine` turning `UNKNOWN` into `IDLE`. `StatusRefiner` already logs that at DEBUG, and production `logback.xml:41` sets `dev.ltms.fleet` to `DEBUG`. ``` $ grep -ac 'refined .* from UNKNOWN to' fleetd/fleetd.out 0 ``` Positive control, because a zero can also mean the probe never fired: ``` $ grep -ao 'DEBUG \[[^]]*\] d\.l\.f\.inject\.[A-Za-z]*' fleetd/fleetd.out | awk '{print $NF}' | sort | uniq -c 14 d.l.f.inject.StatusPoller ``` DEBUG from `inject/` does reach this file, so the logger is live. `StatusRefiner` has simply never written a line. Denominator, same file (2026-07-18 to 2026-09-09): ``` $ grep -ac 'started pane=.* terminal=' fleetd/fleetd.out # 429 member spawns $ grep -ac -- '-> BUSY' fleetd/fleetd.out # 486 turns started $ grep -ac 'fleetd listening\|bridged listening' # 97 daemon boots ``` **One honest limit.** Zero log lines does not prove the raw status was never `UNKNOWN`. If it was `UNKNOWN` **and** `classify` also returned `UNKNOWN`, nothing is logged, because the debug line is inside `if (refined != AgentStatus.UNKNOWN)`. What the zero does prove is the exact bad outcome: across 429 spawns and 486 turns, no target was ever refined from `UNKNOWN` to `IDLE`. ### What I think should happen Half B is a **real defect on a path nobody has walked**. The pane condition is now measured and common; the trigger has not fired once in 7.5 weeks. So this is not urgent, but it should not be closed as "not a defect" either — the ticket's own principle applies: > ambiguous content should stay `UNKNOWN`. `UNKNOWN` is the safe value. The ticket already says the weaker fix is to widen the active-marker list, and it is right. The measurement points the same way: widening would mean guessing more layouts, and the thing that actually failed was trusting four markers that turned out to prove nothing. The stronger fix is to require **positive evidence of idleness** — an anchored check on the last non-blank line of the pane rather than a substring search over the whole scrape. `❯` at the end of the pane means a ready prompt; `❯` anywhere in 2000 characters does not. That also closes the "assistant text quoting a marker" path the hunt originally described, which my measurement did not test. I am not fixing it in this pass. Half A is closed (16µs, 31µs, 56µs — see the earlier comment). Suggest re-scoping this ticket to half B only, keeping it low, with the anchored-check fix as the named approach.
ltms changed title from Two ways the injector can type into a member that is mid-turn — measure before fixing to classify() can call an active pane IDLE — measured real, trigger never seen (half A closed) 2026-09-09 08:36:37 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#340