A lead nudge pastes AND submits, so it clobbers the operator's half-typed prompt — the injectable gate cannot see a draft #761

Closed
opened 2026-10-05 10:23:07 +02:00 by ltms · 2 comments
Owner

Reported by the operator: "fix the issue when I typing something in lead session and fleetd also try to inject its instructions at the same time". I have not reproduced it deliberately, but the mechanism below fully explains it and every step is read from the code or measured on the live daemon.

What happens

AgentControl.send is the only delivery call, and its own javadoc says what it does:

"Deliver text to an agent as its next prompt and submit it — herdr's agent.prompt pastes the text (embedded newlines preserved verbatim) and submits it in the same call"

So a nudge arriving while the operator is mid-sentence does not merely interleave. It pastes into a prompt box that already holds a partial line and then presses Enter. The operator's unfinished message is submitted, with the nudge text attached. That is the harm — not the stray text, the forced submit.

Why the existing gate does not stop it

Every nudge path gates on the same predicate, herdr/AgentStatus.java:36:

public boolean injectable() {
    return this == IDLE || this == BLOCKED || this == DONE;
}

That status describes the agent. An idle agent with the operator halfway through typing reports idle — byte-identical to an idle agent with an empty prompt box. There is no signal anywhere in this chain for "the human has uncommitted keystrokes", so the gate is blind to exactly the state it needs to see.

Three loops share the gate and all inject into a lead's own pane:

Loop Gate Fires when
msg/ReplyPushLoop.java:401 status.injectable() a reply, ticket, question or incident is uncollected
msg/LeadHeartbeatLoop.java:241 status.injectable() the lead has been idle past the quiet period
msg/LeadCoordLoop.java:147 status.injectable() lead-to-lead mail is held

Measured on the live daemon

ReplyPushLoop is the one hitting the operator. Counted in the current fleetd/fleetd.out:

$ grep -c "push: nudge sent to lead" fleetd.out
353
$ grep "push: nudge sent to lead" fleetd.out | tail -3
09:57:23.853 ... ReplyPushLoop - push: nudge sent to lead term_65d106559b02e1 (reply 1/5, ticket 1/5, question 1/5; 0 reply target(s), 1 ticket(s), 0 question(s))
10:10:17.104 ... ReplyPushLoop - push: nudge sent to lead term_65d106559b02e1 (reply 1/5, ticket 1/5, question 1/5; 0 reply target(s), 1 ticket(s), 0 question(s))
10:19:02.223 ... ReplyPushLoop - push: nudge sent to lead term_65d106559b02e1 (reply 1/5, ticket 1/5, question 1/5; 0 reply target(s), 1 ticket(s), 0 question(s))

leadHeartbeat is configured live but with quietNudgeCap: 0, so it never sends a content-free nudge and is not the main source.

Note ticket 1/5: a nudge repeats up to 5 times per ticket until it is collected. So an uncollected ticket multiplies the collision chances, and a lead running several async delegations is nudged often.

The signal already exists

inject/StatusRefiner.java:31-35 already reads the right region:

"detection is the region herdr itself uses for status detection (the prompt/footer tail), which is exactly what we need to tell 'idle at prompt' from 'mid-turn'."

And classify already inspects the input box — pane.contains("│ >"). It treats the box's presence as IDLE and does not distinguish an empty box from one holding the operator's draft. A Claude Code prompt mid-typing looks like │ > fix the issue when I typ…; empty it is │ > and nothing more.

But refine returns early on any non-UNKNOWN status (:67, if (raw != AgentStatus.UNKNOWN) return raw;). A healthy lead pane reports idle, so the refiner never runs for it. The fix cannot live in classify alone.

Proposed

Add a draft check to the lead-nudge gates only. Before injecting into a lead, read the detection region and treat a non-empty input box as not-injectable. The three loops already handle "not injectable → hold and retry", and ReplyPushLoop.injectNudge:754-761 deliberately re-reads its pending work on every attempt, so nothing is lost by waiting — the nudge simply lands after the operator presses Enter.

Scope it to leads. A spawned member has no human typing into it, and a pane read per member poll would add herdr traffic for a state that cannot occur.

Two things to decide, not assume:

  1. A draft that never resolves. An operator who types half a line and walks away would hold nudges indefinitely. ReplyPushLoop has a 5-reminder cap and backoff, so this degrades to "delayed", not "lost" — but it should be bounded and logged rather than silent.
  2. The residual race. Even a perfect gate leaves the window between the read and the paste. Because agent.prompt pastes and submits atomically, losing that race still submits the operator's line. Pasting without Enter would make the failure harmless, but it would also stop a nudge waking a genuinely idle lead, which is ReplyPushLoop's whole purpose. So the conditional is the better shape: empty box ⇒ paste and submit as now; draft present ⇒ do not inject at all. That keeps the wake behaviour and removes all but the race.

I am not proposing a knob. The current behaviour is wrong for every operator, not a preference.

Reported by the operator: *"fix the issue when I typing something in lead session and fleetd also try to inject its instructions at the same time"*. I have not reproduced it deliberately, but the mechanism below fully explains it and every step is read from the code or measured on the live daemon. ## What happens `AgentControl.send` is the only delivery call, and its own javadoc says what it does: > *"Deliver `text` to an agent as its next prompt **and submit it** — herdr's `agent.prompt` pastes the text (embedded newlines preserved verbatim) and submits it in the same call"* So a nudge arriving while the operator is mid-sentence does not merely interleave. It pastes into a prompt box that already holds a partial line and then presses Enter. **The operator's unfinished message is submitted**, with the nudge text attached. That is the harm — not the stray text, the forced submit. ## Why the existing gate does not stop it Every nudge path gates on the same predicate, `herdr/AgentStatus.java:36`: ```java public boolean injectable() { return this == IDLE || this == BLOCKED || this == DONE; } ``` That status describes the **agent**. An idle agent with the operator halfway through typing reports `idle` — byte-identical to an idle agent with an empty prompt box. There is no signal anywhere in this chain for "the human has uncommitted keystrokes", so the gate is blind to exactly the state it needs to see. Three loops share the gate and all inject into a lead's own pane: | Loop | Gate | Fires when | |---|---|---| | `msg/ReplyPushLoop.java:401` | `status.injectable()` | a reply, ticket, question or incident is uncollected | | `msg/LeadHeartbeatLoop.java:241` | `status.injectable()` | the lead has been idle past the quiet period | | `msg/LeadCoordLoop.java:147` | `status.injectable()` | lead-to-lead mail is held | ## Measured on the live daemon `ReplyPushLoop` is the one hitting the operator. Counted in the current `fleetd/fleetd.out`: ``` $ grep -c "push: nudge sent to lead" fleetd.out 353 $ grep "push: nudge sent to lead" fleetd.out | tail -3 09:57:23.853 ... ReplyPushLoop - push: nudge sent to lead term_65d106559b02e1 (reply 1/5, ticket 1/5, question 1/5; 0 reply target(s), 1 ticket(s), 0 question(s)) 10:10:17.104 ... ReplyPushLoop - push: nudge sent to lead term_65d106559b02e1 (reply 1/5, ticket 1/5, question 1/5; 0 reply target(s), 1 ticket(s), 0 question(s)) 10:19:02.223 ... ReplyPushLoop - push: nudge sent to lead term_65d106559b02e1 (reply 1/5, ticket 1/5, question 1/5; 0 reply target(s), 1 ticket(s), 0 question(s)) ``` `leadHeartbeat` is configured live but with `quietNudgeCap: 0`, so it never sends a content-free nudge and is not the main source. Note `ticket 1/5`: a nudge **repeats up to 5 times per ticket until it is collected**. So an uncollected ticket multiplies the collision chances, and a lead running several async delegations is nudged often. ## The signal already exists `inject/StatusRefiner.java:31-35` already reads the right region: > *"`detection` is the region herdr itself uses for status detection (the prompt/footer tail), which is exactly what we need to tell 'idle at prompt' from 'mid-turn'."* And `classify` already inspects the input box — `pane.contains("│ >")`. It treats the box's presence as `IDLE` and does not distinguish an **empty** box from one holding the operator's draft. A Claude Code prompt mid-typing looks like `│ > fix the issue when I typ…`; empty it is `│ >` and nothing more. But `refine` returns early on any non-`UNKNOWN` status (`:67`, `if (raw != AgentStatus.UNKNOWN) return raw;`). A healthy lead pane reports `idle`, so the refiner never runs for it. The fix cannot live in `classify` alone. ## Proposed **Add a draft check to the lead-nudge gates only.** Before injecting into a lead, read the `detection` region and treat a non-empty input box as not-injectable. The three loops already handle "not injectable → hold and retry", and `ReplyPushLoop.injectNudge:754-761` deliberately re-reads its pending work on every attempt, so nothing is lost by waiting — the nudge simply lands after the operator presses Enter. Scope it to leads. A spawned member has no human typing into it, and a pane read per member poll would add herdr traffic for a state that cannot occur. **Two things to decide, not assume:** 1. **A draft that never resolves.** An operator who types half a line and walks away would hold nudges indefinitely. `ReplyPushLoop` has a 5-reminder cap and backoff, so this degrades to "delayed", not "lost" — but it should be bounded and logged rather than silent. 2. **The residual race.** Even a perfect gate leaves the window between the read and the paste. Because `agent.prompt` pastes and submits atomically, losing that race still submits the operator's line. Pasting without Enter would make the failure harmless, but it would also stop a nudge waking a genuinely idle lead, which is `ReplyPushLoop`'s whole purpose. So the conditional is the better shape: empty box ⇒ paste and submit as now; draft present ⇒ do not inject at all. That keeps the wake behaviour and removes all but the race. I am not proposing a knob. The current behaviour is wrong for every operator, not a preference.
Author
Owner

Correction — PR #765 must not merge as it stands. The box marker is wrong.

This is the thing the PR itself asked me to re-measure ("the one thing I would most want re-measured against a live lead pane before merge"). I measured it. The gate can never clear on this host, so every lead nudge would be held forever.

The design, the seam choice and the fail-closed direction are all right. Only the marker string is wrong. Fixing it is small.

What I measured (2026-10-05, read-only, herdr agent read --source detection)

Five live Claude Code panes, counting both markers in each pane's detection region:

pane │ > ❯
wB:p1 (anki, idle) 0 2
w9:p1 (vms, idle) 0 3
w2:p1D (lead: opus, working) 0 1
wA:p1 (trinotes) 0 2
w1:p17 (adev) 0 2

│ > appears zero times on every pane. PromptBox.BOX_MARKER = "│ >" is the only thing that can produce EMPTY, so EMPTY is unreachable. Every nudge takes the UNREADABLE path and holds.

The live UI draws the input box as a line starting with ❯, between two ─── rules:

───────────────────────────────────────────────── lead: opus ─
❯
──────────────────────────────────────────────────────────────
  lead: opus · Opus 5 (1M context) · ~/LTMS/claude-bridge
  ⏵⏵ auto mode on (shift+tab to cycle) · ← 1 agent

That is my own lead pane with an empty box. It is exactly the state the gate must call EMPTY.

Why the tests did not catch it

FakeHerdr.IDLE_PROMPT_BOX draws an older Claude Code UI:

╭────────────────────────╮
│ >                      │
╰────────────────────────╯

So the fixture and the production code share the same wrong assumption, and they agree with each other. All 10 behavioural tests pass, the before-failure evidence is genuine, and none of it touches the real screen. A test can only disagree with the subject about things the fixture did not copy from it.

Where │ > came from, and the exact mistake

inject/StatusRefiner.java:100-102:

boolean readyPrompt = pane.contains("❯")
        || pane.contains("│ >")
        || lower.contains("auto mode on");

Three alternatives joined by ||. ❯ is checked first and is the one that fires here. The PR took the second disjunct alone and used it as the sole marker. So this is not a missing case: it is one branch of a working three-way test, lifted out of a context that no longer needed it to work on its own.

Two more things the measurement showed

1. The detection region is not a short tail — it is about 76 lines and carries scrollback. The PR's reasoning ("the prompt/footer tail") is not what the region actually is. anki has ❯ on line 36 (an earlier, already-submitted prompt) and on line 75 (the live box). So "take the last marker" is not an optimisation, it is required. Keep that part exactly as it is.

2. Match the marker at the start of the line. In my own pane the detection region contained ❯ twice inside quoted text at columns 8 and 9 — my own measurement command, echoed on my screen. The live box was at column 0. So the rule should be the last line whose first character is ❯, not the last occurrence anywhere. Otherwise a pane that prints the glyph in ordinary output will be read as a draft.

What to change

  1. Replace BOX_MARKER with the live marker, and keep │ > as an accepted alternative rather than deleting it — the same reason StatusRefiner keeps all three. A pane may be an older client. The empty test becomes: the last line starting with either marker has nothing after it but whitespace or a cursor glyph.
  2. Drop the rule that treats a bare ❯ as "a shell prompt where the box should be, therefore unreadable". That line is the box, and treating it as unreadable is what breaks this.
  3. Add a fixture taken from the real screen, as a second constant beside IDLE_PROMPT_BOX — do not edit the existing one, 5 tests in StatusPollerRoutingTest, ClaudeCodeLauncherTest and SessionManagerTest depend on its current meaning, as the PR found. Use the live capture above for the empty case and the draft case.
  4. Keep everything else: the seam in herdr, the three call sites, WAIT_BUSY / DRAFT_HELD, the hold-on-unreadable direction, the 20-hold warning that logs a count and never the text, and not wiring it into inject/Injector.

And the bug is real and common — two panes were holding a draft right now

While measuring I found unsubmitted text sitting in two idle panes:

  • anki: yes, send it to lead: opus
  • vms: draft the observer row for CLAUDE.md

Both agents report idle, so injectable() is true for both, and a nudge to either would have submitted the operator's half-finished line. That is the reported bug, caught in the wild, twice, in one sample of five panes. It also means the UNREADABLE-holds-everything behaviour would have hidden the real fix behind a feature that silently stopped working.

On the honesty of the report

The PR flagged this risk itself, said plainly it could not measure it from a worker, and named it as the one thing to check before merge. That is why it took one command to find instead of a week of missing nudges. The gate was right to be built fail-closed, and right to be reported as unverified.

## Correction — PR #765 must not merge as it stands. The box marker is wrong. This is the thing the PR itself asked me to re-measure ("the one thing I would most want re-measured against a live lead pane before merge"). I measured it. The gate can never clear on this host, so every lead nudge would be held forever. The design, the seam choice and the fail-closed direction are all right. Only the marker string is wrong. Fixing it is small. ### What I measured (2026-10-05, read-only, `herdr agent read --source detection`) Five live Claude Code panes, counting both markers in each pane's detection region: | pane | `│ >` | `❯` | |---|---|---| | `wB:p1` (anki, idle) | **0** | 2 | | `w9:p1` (vms, idle) | **0** | 3 | | `w2:p1D` (lead: opus, working) | **0** | 1 | | `wA:p1` (trinotes) | **0** | 2 | | `w1:p17` (adev) | **0** | 2 | `│ >` appears **zero times on every pane**. `PromptBox.BOX_MARKER = "│ >"` is the only thing that can produce `EMPTY`, so `EMPTY` is unreachable. Every nudge takes the `UNREADABLE` path and holds. The live UI draws the input box as a line starting with `❯`, between two `───` rules: ``` ───────────────────────────────────────────────── lead: opus ─ ❯ ────────────────────────────────────────────────────────────── lead: opus · Opus 5 (1M context) · ~/LTMS/claude-bridge ⏵⏵ auto mode on (shift+tab to cycle) · ← 1 agent ``` That is my own lead pane with an empty box. It is exactly the state the gate must call `EMPTY`. ### Why the tests did not catch it `FakeHerdr.IDLE_PROMPT_BOX` draws an **older** Claude Code UI: ``` ╭────────────────────────╮ │ > │ ╰────────────────────────╯ ``` So the fixture and the production code share the same wrong assumption, and they agree with each other. All 10 behavioural tests pass, the before-failure evidence is genuine, and none of it touches the real screen. A test can only disagree with the subject about things the fixture did not copy from it. ### Where `│ >` came from, and the exact mistake `inject/StatusRefiner.java:100-102`: ```java boolean readyPrompt = pane.contains("❯") || pane.contains("│ >") || lower.contains("auto mode on"); ``` Three alternatives joined by `||`. `❯` is checked **first** and is the one that fires here. The PR took the **second** disjunct alone and used it as the sole marker. So this is not a missing case: it is one branch of a working three-way test, lifted out of a context that no longer needed it to work on its own. ### Two more things the measurement showed **1. The detection region is not a short tail — it is about 76 lines and carries scrollback.** The PR's reasoning ("the prompt/footer tail") is not what the region actually is. `anki` has `❯` on line 36 (an earlier, already-submitted prompt) and on line 75 (the live box). So "take the last marker" is not an optimisation, it is required. Keep that part exactly as it is. **2. Match the marker at the start of the line.** In my own pane the detection region contained `❯` twice inside quoted text at columns 8 and 9 — my own measurement command, echoed on my screen. The live box was at column 0. So the rule should be *the last line whose first character is `❯`*, not the last occurrence anywhere. Otherwise a pane that prints the glyph in ordinary output will be read as a draft. ### What to change 1. Replace `BOX_MARKER` with the live marker, and keep `│ >` as an accepted alternative rather than deleting it — the same reason `StatusRefiner` keeps all three. A pane may be an older client. The empty test becomes: the last line starting with either marker has nothing after it but whitespace or a cursor glyph. 2. Drop the rule that treats a bare `❯` as "a shell prompt where the box should be, therefore unreadable". That line is the box, and treating it as unreadable is what breaks this. 3. Add a fixture taken from the real screen, as a second constant beside `IDLE_PROMPT_BOX` — do not edit the existing one, 5 tests in `StatusPollerRoutingTest`, `ClaudeCodeLauncherTest` and `SessionManagerTest` depend on its current meaning, as the PR found. Use the live capture above for the empty case and the draft case. 4. Keep everything else: the seam in `herdr`, the three call sites, `WAIT_BUSY` / `DRAFT_HELD`, the hold-on-unreadable direction, the 20-hold warning that logs a count and never the text, and not wiring it into `inject/Injector`. ### And the bug is real and common — two panes were holding a draft right now While measuring I found unsubmitted text sitting in two **idle** panes: - `anki`: `yes, send it to lead: opus` - `vms`: `draft the observer row for CLAUDE.md` Both agents report `idle`, so `injectable()` is true for both, and a nudge to either would have submitted the operator's half-finished line. That is the reported bug, caught in the wild, twice, in one sample of five panes. It also means the `UNREADABLE`-holds-everything behaviour would have hidden the real fix behind a feature that silently stopped working. ### On the honesty of the report The PR flagged this risk itself, said plainly it could not measure it from a worker, and named it as the one thing to check before merge. That is why it took one command to find instead of a week of missing nudges. The gate was right to be built fail-closed, and right to be reported as unverified.
Author
Owner

Merged to main as 7dad054 and live since 11:19:29 today.

What I checked myself.

  • Build in a throwaway worktree: MVN_EXIT=0, BUILD SUCCESS, 179 surefire XML files, 2167 tests, 0 failures. The redeploy build reported the same 2167.
  • Test count: +27 @Test, -0 against the worker's merge-base 02eff4c.
  • I ran the merged PromptBox.classify against real live pane captures: my own empty lead pane gave EMPTY chars=0, and three panes holding unsent text gave DRAFT with 32, 26 and 22 characters.
  • Deployed with scripts/redeploy-fleetd.sh --yes. Old pid 9796 exited, new pid 96638, jar b0b40b6ccd4b -> 1924721cdb40, fresh fleetd listening line at 11:19:29, no ERROR lines since restart, fleet_whoami still primary.
  • lsof -p 96638 shows the live process holds fleetd/run/fleetd.jar, and javap on that jar's PromptBox prints both ❯ and │ >. So the running daemon has the gate.

Marker census, re-measured today on 7 live panes (my own excluded, because my pane echoes my own command text):

marker hits
❯ at line start 1-3 on every pane
│ > 0 on all 7
esc to interrupt 0 on all 7

This is why the first version of the gate could never clear: it carried only │ >, lifted out of StatusRefiner's working three-way ||. All 10 unit tests still passed, because the fixture drew the same older UI the check looked for. A fixture copied from the production code cannot disagree with it about the real world.

The 1-3 hits per pane also confirm the "take the last box line" rule is needed, not a nicety: the detection region is about 73-76 lines and carries scrollback, including prompts already submitted.

Instruction surface updated, as the project rule requires for an injector/status-gating change:

  • CLAUDE.md invariant 4 now says a lead's own pane has a second gate - ac436ef.
  • wiki/11-Features.md has an entry: what it does, no knob (always on), why injectable() could not see this, and five gotchas - wiki 0df2d9b.
  • wiki/7-Use-Cases.md regenerated; the sync check prints in sync: True. Both pushed and verified by ref.

Two things I am NOT claiming. The esc to interrupt census above proves nothing, because none of those 7 panes was mid-turn. The separate evidence is that a lead pane that was generating showed ✢ Whirlpooling… (29s · ↓ 1.5k tokens) with that string absent, so the marker looks inert. A follow-up ticket covers it, together with StatusRefiner.java:100, which tests pane.contains("❯") with no start-of-line rule.

Merged to `main` as `7dad054` and **live** since 11:19:29 today. **What I checked myself.** - Build in a throwaway worktree: `MVN_EXIT=0`, `BUILD SUCCESS`, 179 surefire XML files, **2167 tests, 0 failures**. The redeploy build reported the same 2167. - Test count: `+27 @Test, -0` against the worker's merge-base `02eff4c`. - I ran the merged `PromptBox.classify` against **real live pane captures**: my own empty lead pane gave `EMPTY` chars=0, and three panes holding unsent text gave `DRAFT` with 32, 26 and 22 characters. - Deployed with `scripts/redeploy-fleetd.sh --yes`. Old pid 9796 exited, new pid 96638, jar `b0b40b6ccd4b` -> `1924721cdb40`, fresh `fleetd listening` line at 11:19:29, no ERROR lines since restart, `fleet_whoami` still `primary`. - `lsof -p 96638` shows the live process holds `fleetd/run/fleetd.jar`, and `javap` on that jar's `PromptBox` prints both `❯` and `│ >`. So the running daemon has the gate. **Marker census, re-measured today on 7 live panes** (my own excluded, because my pane echoes my own command text): | marker | hits | |---|---| | `❯` at line start | 1-3 on every pane | | `│ >` | **0 on all 7** | | `esc to interrupt` | 0 on all 7 | This is why the first version of the gate could never clear: it carried only `│ >`, lifted out of `StatusRefiner`'s working three-way `||`. All 10 unit tests still passed, because the fixture drew the same older UI the check looked for. A fixture copied from the production code cannot disagree with it about the real world. The 1-3 hits per pane also confirm the "take the last box line" rule is needed, not a nicety: the `detection` region is about 73-76 lines and carries scrollback, including prompts already submitted. **Instruction surface updated**, as the project rule requires for an injector/status-gating change: - `CLAUDE.md` invariant 4 now says a lead's own pane has a second gate - `ac436ef`. - `wiki/11-Features.md` has an entry: what it does, no knob (always on), why `injectable()` could not see this, and five gotchas - wiki `0df2d9b`. - `wiki/7-Use-Cases.md` regenerated; the sync check prints `in sync: True`. Both pushed and verified by ref. **Two things I am NOT claiming.** The `esc to interrupt` census above proves nothing, because none of those 7 panes was mid-turn. The separate evidence is that a lead pane that *was* generating showed `✢ Whirlpooling… (29s · ↓ 1.5k tokens)` with that string absent, so the marker looks inert. A follow-up ticket covers it, together with `StatusRefiner.java:100`, which tests `pane.contains("❯")` with no start-of-line rule.
ltms closed this issue 2026-10-05 11:21:00 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#761