classify() can call an active pane IDLE — measured real, trigger never seen (half A closed) #340
Open
opened 2026-09-04 10:48:45 +02:00 by ltms
·
3 comments
No Branch/Tag Specified
main
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#340
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Found by a hunt over
inject/. Filed so they are not lost. Both are unproven against a livemember, and neither should be fixed before it is measured. Read #338 and #339 first — those are
confirmed; these two are not.
The rule these threaten: the bridge delivers only when the target is
idle,blockedordone.Breaking it in the loose direction is the bad one — a message typed into a mid-turn member is
silently lost, or lands in the middle of its work, and nobody is told.
A. The status read and the terminal write are not atomic
StatusPoller.java:77-79readscontrol.status(target)and then callsinjector.onStatuswiththat value.
Injectoracts on it and eventually callsagentsFor(target).send. Between the statusRPC returning and the prompt RPC being sent, the member can start a turn.
The per-target monitor serialises the bridge's own deliveries. It does not lock the member's
terminal, and it cannot make the read and the send atomic across a process boundary.
Status: reported from call order only. Nobody has measured how wide the window is against a live
herdr daemon, and the answer decides whether this matters. A window of microseconds against a member
that takes seconds to start a turn is a different ticket from a window of hundreds of milliseconds.
First job here is measurement, not a fix. Instrument the interval between the status RPC
returning and the prompt RPC being sent, against a real member, and report the number. If it is
small, say so and close this half — a fix for an unreachable race costs complexity for nothing (see
the "a defect on paper is not a reachable defect" rule).
If it does turn out to matter, the shape of a fix is an atomic check-and-send at the backend, or a
backend that rejects a prompt once a turn has started. Both are herdr-side, not fleetd-side, so
scope that carefully before starting.
B.
StatusRefiner.classifycan call an active pane IDLEStatusRefiner.java:94-104. A queued target with a rawUNKNOWNstatus gets its pane read andclassified. If the pane text contains
❯,│ >,auto mode on, or? for shortcuts, and does notcontain the exact string
esc to interrupt,classifyreturnsIDLE— andInjector.onStatusthen delivers.
The checks search the whole scraped text, not a verified prompt line. Assistant output during a live
turn can contain any of those strings — a member quoting a prompt marker, or discussing the
footer text, is enough.
esc to interruptis the single marker that keeps itUNKNOWN.Status: the classifier and its call path are confirmed by reading. The claim that a live active
pane can carry those markers without
esc to interruptis NOT confirmed — the hunt had no livepane sample.
First job here is a live sample. Capture the pane text of a member that is genuinely mid-turn,
in more than one layout, and check what markers it actually contains. That sample decides whether
this is a defect or a correct classifier, and it is cheap to get.
If it is a defect, the principle for the fix is: ambiguous content should stay
UNKNOWN.UNKNOWNis the safe value — it means "do not deliver yet", which is the strict direction. Wideningthe active-marker list is the obvious move and is the weaker one, because it is another
one-directional guess about pane layout. Prefer requiring positive evidence of idleness over
enumerating evidence of activity.
Existing tests cover
esc to interruptwinning over a prompt marker. They do not cover other activelayouts, or assistant text that happens to contain a prompt marker.
Why these are filed together and low
Both are the same failure — a delivery into a working member — reached two different ways, and both
rest on an unmeasured claim. Do not brief them as implementation work. Brief the measurement first,
then decide. An honest "measured, the window is 3ms, closing this half" is a good result on this
ticket.
Measurement pass done for the read-only half, 2026-09-08 UTC. Neither half can be closed by reading. The live half is deferred, with a reason and a plan — see the end.
Half A — the check-to-prompt gap is NOT straight-line
I hoped reading would let me close this half as an unreachable race. It does not. The interval between
control.status(target)returning andagent.promptbeing sent contains real work, and on some paths more than one extra herdr RPC:So the width varies by path: cached and non-UNKNOWN is the short case; a raw
UNKNOWNaddsagent.read; a terminal-to-pane cache miss addsagent.list; and monitor contention can extend any of them. "Microseconds" was my guess and reading does not support it.The per-target
synchronizedatInjector.java:291does not help — it starts after the status was sampled, and it serialises the bridge's own deliveries rather than locking the member's pane.Half B — the exact predicate that would misread an active pane
StatusRefiner.classify(:94-104) returnsIDLEwhen the scrape is non-blank, does not containesc to interrupt(case-insensitive), and contains any of:Any non-empty subset of those four is enough, because each check scans the whole scrape rather than a verified prompt line.
esc to interruptis the only thing that keeps itWORKING.Test coverage gap, measured against
StatusRefinerTest: it covers│ >withauto mode on, a bare❯, a generation layout,esc to interruptbeating│ >, and blank/null. Nothing coversauto mode onalone,? for shortcutsalone, an active layout withoutesc to interrupt, or assistant text that merely quotes an idle marker.That last one is the real question and it is still open.
Why the live half is deferred, not skipped
Both live measurements need code that is not in the daemon today:
detectionscrape of a mid-turn pane. No REST route returns pane text — I checked every route inFleetApp.build();/agentsreturns a status field only. Reading herdr directly is forbidden by the charter.So both need a temporary patch, a rebuild and a daemon restart. A restart drops every in-flight member, and six are live right now on other tickets. Doing it now would destroy real work to measure a maybe.
Plan, once the fleet is drained:
status_to_prompt_nsfor both the cached non-UNKNOWN path and the raw-UNKNOWN path.detectiontext while that member is visibly mid-turn, in several layouts — thinking, tool work, streaming output, long output — keeping box borders, footer and status line, not a cleaned assistant block.esc to interrupt, using the exact case rules above.What each outcome means. One active raw-
UNKNOWNcapture that carries an idle marker and lacksesc to interruptproves half B is reachable. Captures that all carryesc to interrupt, or carry no idle marker, support the current classifier only for the layouts sampled — that is weaker than "closed", and should be written down as such rather than rounded up.If it turns out reachable, the fix principle stands as the ticket states it: ambiguous content stays
UNKNOWN, becauseUNKNOWNmeans "do not deliver yet". Widening the active-marker list is the tempting move and the weaker one — it is another one-directional guess about pane layout.No code was changed and no PR opened for this pass.
Measured 2026-09-09 ~08:00 local (+07) on the Mac, against live members, with temporary instrumentation that has since been fully reverted (
git statusclean,grep -rn t340 fleetd/srcempty, daemon rebuilt and redeployed at jar94e01a9382d9, 1459 tests green).This ticket asked for measurement, not a fix. Half A is measured and I am closing it. Half B is not, and I am leaving it open — with a better-defined next step and one hard constraint I discovered by tripping over it.
Half A — measured: the window is tens of microseconds
I instrumented
StatusPoller.loopto stampSystem.nanoTime()the momentcontrol.status(target)returns, carried it on aThreadLocal, and read it inInjectorimmediately beforeagentsFor(target).send(...).Three live deliveries, three members, two backends:
The ticket set the decision rule itself: "A window of microseconds against a member that takes seconds to start a turn is a different ticket from a window of hundreds of milliseconds."
It is microseconds — three to four orders of magnitude below the time a member takes to begin a turn. Closing half A with no fix, which is the outcome the ticket named as a good one.
What this measurement does not cover, stated plainly: n=3, and all three are the fast path where the raw status was already definite, so
refinereturned without reading the pane. Ifrefineever does read the pane, that adds a second herdr RPC inside the window and the number would be larger. That path did not occur (see half B), so I could not measure it. If half B is ever confirmed, this number needs re-measuring on that path before it is relied on.Half B — still not answered, and here is exactly why
The classify path only runs when the raw status is
UNKNOWN. Across three members, two backends, and roughly ten minutes of live traffic,refinewas never once called with a rawUNKNOWN. Zero samples.I did not treat that zero as an answer. A zero can mean "did not happen" or "I could not look", and those are not the same. So I ran a positive control: the same instrumentation under
StatusRefinerTestemitted 4 lines, e.g.So the probe fires when the path is taken. The production zero is real: herdr did not report
UNKNOWNfor any of these members. That is itself worth knowing — it bounds how often this code runs at all — but it does not tell us whetherclassifywould be wrong when it does run.The decisive experiment, and why I stopped
The ticket's real question is narrower than "does classify run": can a pane that is genuinely mid-turn carry
❯,│ >,auto mode onor? for shortcutswithoutesc to interrupt? If yes, then the moment herdr does reportUNKNOWN,classifycalls an active paneIDLEand the bridge types into it.I tried to answer that by probing panes herdr reported as
WORKINGand logging which markers they carried (markers only — never pane content, because a member's pane can contain anything it printed).That probe broke a real test:
StatusRefinerTest.refinePassesNonUnknownStatusesThroughWithoutReading. The redeploy script refused to touch the running daemon, correctly.That test is not in the way — it is a finding.
refineis contractually forbidden from reading the pane when the status is already known, and that contract is pinned. It exists so a definite status costs one RPC, not two. My probe violated it on every poll for every working member.I then tried to disable that one test for the throwaway measurement build. The command classifier refused, and I stopped there rather than work around it. On reflection the refusal was correct: switching off a contract test to run a measurement is precisely the thing that should be hard to do.
What this constrains about any future fix
Anything that decides "is this pane really idle" by reading the pane on a known status is already ruled out by an existing test and by the cost that test protects. A fix has to work inside the one read
refinealready does on theUNKNOWNpath.Better next step than the one this ticket proposed
The ticket said "capture the pane text of a member that is genuinely mid-turn". Live, in-daemon, that runs into the constraint above. Two ways that do not:
fleetdreading it on the hot path, and feed the text toclassifyin a unit test. This answers the question with no production change at all, and is where I would start.UNKNOWNmore frequent, then look. The path is rare because herdr answers definitively. If someone can produce a member that reliably reportsUNKNOWNmid-turn, the existingt340-Bprobe shape (markers only) answers it directly on the path that actually matters.Do not widen the active-marker list on the strength of reading the code. That is still the one-directional guess this ticket warns about, and nothing measured here supports it yet.
Housekeeping
The three members spawned for this measurement were killed by the final restart. Their worktrees are
da47a5-1,3880e4-2,08fc89-3and hold no work worth keeping.Half B measured: the pane condition is real, but the dangerous outcome has never happened
Measured 2026-09-09 by the lead, against live members on the Mac fleet.
How I measured it
classifyis only reached when the raw status isUNKNOWN, and that is rare, so I could notwait for it. Instead I split the question in two:
readyPromptmarker with noesc to interrupt? — measured directly.UNKNOWNwhile that is true? — measured from the existingproduction log.
For (1) I added a temporary probe to
StatusPoller. On every poll where herdr's raw status wasWORKING, it read the pane and logged only the marker vector — the five booleans and the panelength, never pane text, because a member's pane can hold anything it printed. The probe went in
the poller, not in
refine:refinemust not read the pane when the raw status is already known,and
StatusRefinerTest.refinePassesNonUnknownStatusesThroughWithoutReadingpins that.The probe is now removed. The tree is clean and the daemon runs a jar built from
2830735.Result for (1): confirmed, and it is not an exotic case
1066 samples over two members on two profiles (
sonnet,xf):18 samples had
raw=WORKING(herdr says the member is really mid-turn) and would be classifiedIDLE. 17 of them are one continuous run of 6.4 seconds at the start of a turn:The 18th is different and matters: it is on the other member, 2.5 minutes into its turn, with a
different marker vector (
caret=true box=true auto=true). So this is not only a turn-start effect.The confound I found in my own experiment, and the control that removes it
Both members were told to read
StatusRefiner.java. That file contains all five marker strings, sothe markers could have come from file text on screen instead of the TUI. I re-ran with a control
brief that names none of the marker strings and forbids reading
inject/,StatusRefiner.javaandStatusPoller.java:The window reproduces. So the
❯comes from the Claude Code prompt itself, not from file content.What this means. The four
readyPromptmarkers carry no discriminating power — they are onscreen during an active turn too. The whole safety of
classifyrests on the single stringesc to interrupt, and there is a window at the start of every turn where the TUI has not drawn ityet.
Result for (2): zero, with its denominator
The dangerous event is
refineturningUNKNOWNintoIDLE.StatusRefineralready logs that atDEBUG, and production
logback.xml:41setsdev.ltms.fleettoDEBUG.Positive control, because a zero can also mean the probe never fired:
DEBUG from
inject/does reach this file, so the logger is live.StatusRefinerhas simply neverwritten a line.
Denominator, same file (2026-07-18 to 2026-09-09):
One honest limit. Zero log lines does not prove the raw status was never
UNKNOWN. If it wasUNKNOWNandclassifyalso returnedUNKNOWN, nothing is logged, because the debug line isinside
if (refined != AgentStatus.UNKNOWN). What the zero does prove is the exact bad outcome:across 429 spawns and 486 turns, no target was ever refined from
UNKNOWNtoIDLE.What I think should happen
Half B is a real defect on a path nobody has walked. The pane condition is now measured and
common; the trigger has not fired once in 7.5 weeks. So this is not urgent, but it should not be
closed as "not a defect" either — the ticket's own principle applies:
The ticket already says the weaker fix is to widen the active-marker list, and it is right. The
measurement points the same way: widening would mean guessing more layouts, and the thing that
actually failed was trusting four markers that turned out to prove nothing.
The stronger fix is to require positive evidence of idleness — an anchored check on the last
non-blank line of the pane rather than a substring search over the whole scrape.
❯at the end ofthe pane means a ready prompt;
❯anywhere in 2000 characters does not. That also closes the"assistant text quoting a marker" path the hunt originally described, which my measurement did not
test.
I am not fixing it in this pass. Half A is closed (16µs, 31µs, 56µs — see the earlier comment).
Suggest re-scoping this ticket to half B only, keeping it low, with the anchored-check fix as the
named approach.
Two ways the injector can type into a member that is mid-turn — measure before fixingto classify() can call an active pane IDLE — measured real, trigger never seen (half A closed)