a member whose backend 400s dies in one second, and fleetd returns the lost turn as a successful empty reply #164

Closed
opened 2026-08-24 16:39:22 +02:00 by agent · 2 comments
Member

Found during a communication dry run on fleet01 (2026-08-24). A reviewer member spawned on profile local never did any work, and fleet_send reported success anyway.

Symptom

A blocking POST /sessions/{id}/message returned HTTP 200 with:

{"reply":"","sessionId":"term_659cbd67f40da12","replySource":"transcript"}

No error anywhere in the caller's view. GET /sessions/{id}/replies was {"replies":[]} — nothing arrived late either. The turn was simply lost, and the caller was told it succeeded.

Two defects, one of them ours

1. Profile local is dead with Claude Code 2.1.231 — the gateway rejects thinking.type: adaptive

Read off the member's own pane (herdr agent read w1:pD --source recent):

● API Error: 400 1 validation error:
    {'type': 'literal_error', 'loc': ('body', 'thinking', 'type'),
     'msg': "Input should be 'enabled' or 'disabled'",
     'input': 'adaptive'}

Claude Code 2.1.231 sends thinking.type: "adaptive"; https://llm.ltms.dev/anthropic accepts only enabled / disabled, so it 400s the request before any work starts.

This is not a launcher, auth, or routing problem. The spawn is already correct — bare claude, no ccs wrapper, with ANTHROPIC_BASE_URL + ANTHROPIC_AUTH_TOKEN set exactly as ClaudeCodeLauncher.java:235-236 intends:

argv=[claude, --mcp-config, {...}, --append-system-prompt, <charter>, --model, deepseek-v4-flash]

What the gateway refuses is a request body field, so nothing in the env/argv layer can fix it. The fix belongs at the gateway (normalize or drop an unknown thinking.type instead of 400ing) or in a Claude Code setting that pins thinking to enabled/disabled.

Observed on every turn attempted on local today. The 06:09 local spawn is not a counterexample — it was released without ever being sent a message.

2. fleetd turns that hard failure into a silent success — the actual fleetd bug

From fleetd-run/fleetd.out:

14:31:46.634 DEBUG SessionManager   - session transitioned ... READY -> BUSY turn=1
14:31:47.668 DEBUG SessionManager   - session transitioned ... BUSY -> DONE turn=1
14:31:47.678 DEBUG CompletionResolver - resolved send to term_659cbd67f40da12
                                        via turn-completion fallback (0 chars scraped)

The member went BUSY -> DONE in 1.03 s — it cannot have read a file, let alone answered. CompletionResolver accepted that as a completed turn and resolved the caller's blocking send with an empty scrape.

A confirmed working -> idle transition is currently treated as sufficient evidence of a finished turn. It isn't: a turn that dies on a backend 400 produces the same transition, just much faster and with nothing on screen.

The same resolver misfired earlier in the session on a healthy gx member too, resolving a send after 20 s with 2452 chars of scraped TUI — box-drawing characters, the echoed prompt, and leftovers from the previous turn — delivered to the caller as if it were the reply. So the fallback fails in both directions: empty when the turn died, garbage when it had not finished.

Suggested direction

  1. Treat a suspiciously short turn as a failure, not a completion. A BUSY -> DONE inside some floor (a second or two) is a crash signature. Resolve the send as an error naming the member, not as a reply.
  2. Never resolve a send with an empty scrape. 0 chars scraped should be a hard failure — the caller can retry or escalate; it cannot do anything with "" that it was told is a success.
  3. Surface backend errors. The pane had a legible API Error: 400 on screen the whole time. If the scrape is going to be the fallback anyway, pattern-match obvious backend failures and return them as errors instead of as content.
  4. Guard the profile at spawn. A profile whose first turn 400s should be quarantined with the backend's message attached, rather than staying ready in /members and accepting more sends.

Point 2 is the one that matters most: today a lost turn and a successful empty answer are indistinguishable to every caller, including a lead deciding whether to delegate again.

Repro

curl -sX POST 'http://127.0.0.1:8765/members?role=reviewer&profile=local&cwd=<repo>&worktree=true&ticket=t'
curl -sX POST http://127.0.0.1:8765/sessions/<terminalId>/message \
  -H 'Content-Type: application/json' \
  -d '{"wait":true,"timeoutMs":120000,"content":"Read README.md and summarise it."}'
# -> HTTP 200 {"reply":"","replySource":"transcript"}
herdr agent read <paneId> --source recent   # the 400 is plainly visible here

Environment: fleet01, bridged.jar (started 06:07), herdr 0.8.0 protocol 19, Claude Code 2.1.231, profile local = kind: claude-code, baseUrl: https://llm.ltms.dev/anthropic, model: deepseek-v4-flash.

Found during a communication dry run on **fleet01** (2026-08-24). A `reviewer` member spawned on profile `local` never did any work, and `fleet_send` reported success anyway. ## Symptom A blocking `POST /sessions/{id}/message` returned **HTTP 200** with: ```json {"reply":"","sessionId":"term_659cbd67f40da12","replySource":"transcript"} ``` No error anywhere in the caller's view. `GET /sessions/{id}/replies` was `{"replies":[]}` — nothing arrived late either. The turn was simply lost, and the caller was told it succeeded. ## Two defects, one of them ours ### 1. Profile `local` is dead with Claude Code 2.1.231 — the gateway rejects `thinking.type: adaptive` Read off the member's own pane (`herdr agent read w1:pD --source recent`): ``` ● API Error: 400 1 validation error: {'type': 'literal_error', 'loc': ('body', 'thinking', 'type'), 'msg': "Input should be 'enabled' or 'disabled'", 'input': 'adaptive'} ``` Claude Code 2.1.231 sends `thinking.type: "adaptive"`; `https://llm.ltms.dev/anthropic` accepts only `enabled` / `disabled`, so it 400s the request before any work starts. **This is not a launcher, auth, or routing problem.** The spawn is already correct — bare `claude`, no `ccs` wrapper, with `ANTHROPIC_BASE_URL` + `ANTHROPIC_AUTH_TOKEN` set exactly as `ClaudeCodeLauncher.java:235-236` intends: ``` argv=[claude, --mcp-config, {...}, --append-system-prompt, <charter>, --model, deepseek-v4-flash] ``` What the gateway refuses is a **request body field**, so nothing in the env/argv layer can fix it. The fix belongs at the gateway (normalize or drop an unknown `thinking.type` instead of 400ing) or in a Claude Code setting that pins thinking to `enabled`/`disabled`. Observed on every turn attempted on `local` today. The 06:09 `local` spawn is not a counterexample — it was released without ever being sent a message. ### 2. fleetd turns that hard failure into a silent success — the actual fleetd bug From `fleetd-run/fleetd.out`: ``` 14:31:46.634 DEBUG SessionManager - session transitioned ... READY -> BUSY turn=1 14:31:47.668 DEBUG SessionManager - session transitioned ... BUSY -> DONE turn=1 14:31:47.678 DEBUG CompletionResolver - resolved send to term_659cbd67f40da12 via turn-completion fallback (0 chars scraped) ``` The member went `BUSY -> DONE` in **1.03 s** — it cannot have read a file, let alone answered. `CompletionResolver` accepted that as a completed turn and resolved the caller's blocking send with an empty scrape. A confirmed `working -> idle` transition is currently treated as sufficient evidence of a finished turn. It isn't: a turn that dies on a backend 400 produces the same transition, just much faster and with nothing on screen. The same resolver misfired earlier in the session on a healthy `gx` member too, resolving a send after 20 s with **2452 chars of scraped TUI** — box-drawing characters, the echoed prompt, and leftovers from the previous turn — delivered to the caller as if it were the reply. So the fallback fails in both directions: empty when the turn died, garbage when it had not finished. ## Suggested direction 1. **Treat a suspiciously short turn as a failure, not a completion.** A `BUSY -> DONE` inside some floor (a second or two) is a crash signature. Resolve the send as an error naming the member, not as a reply. 2. **Never resolve a send with an empty scrape.** `0 chars scraped` should be a hard failure — the caller can retry or escalate; it cannot do anything with `""` that it was told is a success. 3. **Surface backend errors.** The pane had a legible `API Error: 400` on screen the whole time. If the scrape is going to be the fallback anyway, pattern-match obvious backend failures and return them as errors instead of as content. 4. **Guard the profile at spawn.** A profile whose first turn 400s should be quarantined with the backend's message attached, rather than staying `ready` in `/members` and accepting more sends. Point 2 is the one that matters most: today a lost turn and a successful empty answer are indistinguishable to every caller, including a lead deciding whether to delegate again. ## Repro ```bash curl -sX POST 'http://127.0.0.1:8765/members?role=reviewer&profile=local&cwd=<repo>&worktree=true&ticket=t' curl -sX POST http://127.0.0.1:8765/sessions/<terminalId>/message \ -H 'Content-Type: application/json' \ -d '{"wait":true,"timeoutMs":120000,"content":"Read README.md and summarise it."}' # -> HTTP 200 {"reply":"","replySource":"transcript"} herdr agent read <paneId> --source recent # the 400 is plainly visible here ``` Environment: fleet01, `bridged.jar` (started 06:07), herdr 0.8.0 protocol 19, Claude Code 2.1.231, profile `local` = `kind: claude-code`, `baseUrl: https://llm.ltms.dev/anthropic`, `model: deepseek-v4-flash`.
Owner

Rescued: a partial fix for this existed but was never committed

While cleaning up worker worktrees today I found real work for this issue sitting uncommitted in an abandoned worktree (.bridged-worktrees/2445bd-3), untouched since 2026-08-28. It was one git worktree remove away from being lost.

It is now committed and pushed, so it is safe:

  • branch: worker/cb-164-empty-scrape-false-success-1a80af-3
  • commit: 851ebca
  • patch backup: ~/LTMS/.bridged-handover/cb-164-empty-scrape-false-success.patch

What it does

It covers points 1 and 3 of the suggested direction, and nothing else.

  • CompletionResolver.MIN_COMPLETED_TURN_NANOS = 2 seconds. The comment gives the reasoning: "A backend rejection can return the pane to idle in about one second. Two seconds is above the measured 1.03s crash while staying low enough not to reject ordinary short model answers." That matches the 1.03 s figure in this issue.
  • A BACKEND_ERROR pattern, deliberately narrow: (?i)\bAPI Error\s*:. Its comment says "One stable, explicit backend-failure marker; broader error lists would be brittle." That is the right instinct — a long list of error strings rots.
  • TurnListener.onTurnComplete and CompletionResolver.resolveBeforePostAction gain an elapsedNanos parameter, with the old no-arg forms kept and defaulting to Long.MAX_VALUE (so an unmeasured turn is never judged too short).

Diff: 6 files, +187 / −43, including tests in CompletionResolverTest and MessageServiceTest.

What is NOT done

  • Point 2 is not addressed. An empty scrape is still not, on its own, a hard failure. Point 2 is the one this issue calls most important, so this work does not close the issue.
  • Point 4 is not addressed — no spawn-time quarantine of a profile whose first turn 400s.
  • The work was never built and never reviewed at that commit.
  • The branch is 42 commits behind main, and a rebase conflicts in three files: CompletionResolver.java, CompletionResolverTest.java, and MessageServiceTest.java. Resolving those needs a real read of what main has done to CompletionResolver since 2026-08-28, not a mechanical merge.

Next step

Rebase it onto main, resolve the three conflicts, build, then add point 2 on top. Do not merge it as it stands.

There is also a process lesson worth writing down: a worker's work only survives if it is committed. This branch had zero commits, so the usual backup (git fetch <worktree> <branch>) would have captured nothing.

## Rescued: a partial fix for this existed but was never committed While cleaning up worker worktrees today I found real work for this issue sitting **uncommitted** in an abandoned worktree (`.bridged-worktrees/2445bd-3`), untouched since 2026-08-28. It was one `git worktree remove` away from being lost. It is now committed and pushed, so it is safe: - branch: `worker/cb-164-empty-scrape-false-success-1a80af-3` - commit: `851ebca` - patch backup: `~/LTMS/.bridged-handover/cb-164-empty-scrape-false-success.patch` ### What it does It covers **points 1 and 3** of the suggested direction, and nothing else. - `CompletionResolver.MIN_COMPLETED_TURN_NANOS` = 2 seconds. The comment gives the reasoning: *"A backend rejection can return the pane to idle in about one second. Two seconds is above the measured 1.03s crash while staying low enough not to reject ordinary short model answers."* That matches the 1.03 s figure in this issue. - A `BACKEND_ERROR` pattern, deliberately narrow: `(?i)\bAPI Error\s*:`. Its comment says *"One stable, explicit backend-failure marker; broader error lists would be brittle."* That is the right instinct — a long list of error strings rots. - `TurnListener.onTurnComplete` and `CompletionResolver.resolveBeforePostAction` gain an `elapsedNanos` parameter, with the old no-arg forms kept and defaulting to `Long.MAX_VALUE` (so an unmeasured turn is never judged too short). Diff: 6 files, +187 / −43, including tests in `CompletionResolverTest` and `MessageServiceTest`. ### What is NOT done - **Point 2 is not addressed.** An empty scrape is still not, on its own, a hard failure. Point 2 is the one this issue calls most important, so this work does not close the issue. - **Point 4 is not addressed** — no spawn-time quarantine of a profile whose first turn 400s. - The work was **never built and never reviewed** at that commit. - The branch is **42 commits behind `main`**, and a rebase conflicts in three files: `CompletionResolver.java`, `CompletionResolverTest.java`, and `MessageServiceTest.java`. Resolving those needs a real read of what `main` has done to `CompletionResolver` since 2026-08-28, not a mechanical merge. ### Next step Rebase it onto `main`, resolve the three conflicts, build, then add point 2 on top. Do not merge it as it stands. There is also a process lesson worth writing down: a worker's work only survives if it is **committed**. This branch had zero commits, so the usual backup (`git fetch <worktree> <branch>`) would have captured nothing.
Owner

Points 1 and 2 are done and on main. Closing.

Point State Where
1 — an unreadable scrape must not resolve as success done 3bfa828
1 — an empty scrape must not resolve as success done 3bfa828
2 — a suspiciously fast turn must fail done 3bfa828 (MIN_TURN_NANOS, 2s)
2 — a backend rejection must not read as a reply done 4ac688b (this close)
3 — broader backend-error surfacing open, see below —
4 — out of scope here open —

4ac688b adds a deliberately narrow BACKEND_ERROR pattern ((?i)\bAPI Error\s*:). A scrape that reads cleanly but carries the backend's own rejection is now WORKER_FAILED, not a completed reply.

The failure reason carries the whole pane tail, not just the matched line. The pattern is a heuristic — it also matches a member that forgot fleet_reply while reporting about a backend error, which happens here, since members do investigate gateway 400s. Failing that turn is still right; throwing its report away would have been this same bug wearing a failure label.

1037 tests, mvn clean install unpiped, exit 0. Each of the 5 new tests fails with the fix commented out — I ran that myself rather than reading the worker's transcript.

Two things I am deliberately leaving behind

A dead-code ordering question, for point 3. The rescued branch had a visibleTurn raw-screen fallback, so a TUI-hidden error line could still classify when lastAssistantBlock parses blank. On the current check order the empty-tail fail fires first, so that fallback would never run without also reordering the chain. Reordering is a real behaviour change — it moves which classification wins for a whole class of turns — and it belongs with point 3, not smuggled in here.

Why the pattern stays narrow. A growing list of backend error strings rots as backends reword them, and a stale pattern fails in the silent direction: it stops matching, and the scrape resolves as a success again. Point 3 should decide on a real mechanism (a per-profile configured pattern, like exhaustedPattern already is), not more hard-coded strings.

Process note

851ebca, the branch I rescued from an about-to-be-removed worktree, was mostly already on main. Someone had committed the same fix directly as 3bfa828 the same day. I briefed a worker to rebase and extend a branch that was largely redundant, and it cost a full worker run to find that out — the worker caught it, I had not.

That is the third time this week (#173, #174) I have acted on a branch without first running git merge-base --is-ancestor / git log <branch>..origin/main. The rescue case is the one I had not generalised to: I checked that the work was unsaved, and never checked whether it was needed. Recorded so the next rescue starts with that check.

Points 1 and 2 are done and on `main`. Closing. | Point | State | Where | |---|---|---| | 1 — an unreadable scrape must not resolve as success | done | `3bfa828` | | 1 — an empty scrape must not resolve as success | done | `3bfa828` | | 2 — a suspiciously fast turn must fail | done | `3bfa828` (`MIN_TURN_NANOS`, 2s) | | 2 — a backend rejection must not read as a reply | done | `4ac688b` (this close) | | 3 — broader backend-error surfacing | **open, see below** | — | | 4 — out of scope here | open | — | `4ac688b` adds a deliberately narrow `BACKEND_ERROR` pattern (`(?i)\bAPI Error\s*:`). A scrape that reads cleanly but carries the backend's own rejection is now `WORKER_FAILED`, not a completed reply. The failure reason carries the **whole pane tail**, not just the matched line. The pattern is a heuristic — it also matches a member that forgot `fleet_reply` while reporting *about* a backend error, which happens here, since members do investigate gateway 400s. Failing that turn is still right; throwing its report away would have been this same bug wearing a failure label. 1037 tests, `mvn clean install` unpiped, exit 0. Each of the 5 new tests fails with the fix commented out — I ran that myself rather than reading the worker's transcript. ## Two things I am deliberately leaving behind **A dead-code ordering question, for point 3.** The rescued branch had a `visibleTurn` raw-screen fallback, so a TUI-hidden error line could still classify when `lastAssistantBlock` parses blank. On the current check order the empty-tail fail fires first, so that fallback would never run without also reordering the chain. Reordering is a real behaviour change — it moves which classification wins for a whole class of turns — and it belongs with point 3, not smuggled in here. **Why the pattern stays narrow.** A growing list of backend error strings rots as backends reword them, and a stale pattern fails in the silent direction: it stops matching, and the scrape resolves as a success again. Point 3 should decide on a real mechanism (a per-profile configured pattern, like `exhaustedPattern` already is), not more hard-coded strings. ## Process note `851ebca`, the branch I rescued from an about-to-be-removed worktree, was mostly **already on `main`**. Someone had committed the same fix directly as `3bfa828` the same day. I briefed a worker to rebase and extend a branch that was largely redundant, and it cost a full worker run to find that out — the worker caught it, I had not. That is the third time this week (#173, #174) I have acted on a branch without first running `git merge-base --is-ancestor` / `git log <branch>..origin/main`. The rescue case is the one I had not generalised to: I checked that the *work* was unsaved, and never checked whether it was *needed*. Recorded so the next rescue starts with that check.
ltms closed this issue 2026-08-31 05:57:32 +02:00
Sign in to join this conversation.
2 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#164