a member whose backend 400s dies in one second, and fleetd returns the lost turn as a successful empty reply #164
Closed
opened 2026-08-24 16:39:22 +02:00 by agent
·
2 comments
No Branch/Tag Specified
main
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#164
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Found during a communication dry run on fleet01 (2026-08-24). A
reviewermember spawned on profilelocalnever did any work, andfleet_sendreported success anyway.Symptom
A blocking
POST /sessions/{id}/messagereturned HTTP 200 with:No error anywhere in the caller's view.
GET /sessions/{id}/replieswas{"replies":[]}— nothing arrived late either. The turn was simply lost, and the caller was told it succeeded.Two defects, one of them ours
1. Profile
localis dead with Claude Code 2.1.231 — the gateway rejectsthinking.type: adaptiveRead off the member's own pane (
herdr agent read w1:pD --source recent):Claude Code 2.1.231 sends
thinking.type: "adaptive";https://llm.ltms.dev/anthropicaccepts onlyenabled/disabled, so it 400s the request before any work starts.This is not a launcher, auth, or routing problem. The spawn is already correct — bare
claude, noccswrapper, withANTHROPIC_BASE_URL+ANTHROPIC_AUTH_TOKENset exactly asClaudeCodeLauncher.java:235-236intends:What the gateway refuses is a request body field, so nothing in the env/argv layer can fix it. The fix belongs at the gateway (normalize or drop an unknown
thinking.typeinstead of 400ing) or in a Claude Code setting that pins thinking toenabled/disabled.Observed on every turn attempted on
localtoday. The 06:09localspawn is not a counterexample — it was released without ever being sent a message.2. fleetd turns that hard failure into a silent success — the actual fleetd bug
From
fleetd-run/fleetd.out:The member went
BUSY -> DONEin 1.03 s — it cannot have read a file, let alone answered.CompletionResolveraccepted that as a completed turn and resolved the caller's blocking send with an empty scrape.A confirmed
working -> idletransition is currently treated as sufficient evidence of a finished turn. It isn't: a turn that dies on a backend 400 produces the same transition, just much faster and with nothing on screen.The same resolver misfired earlier in the session on a healthy
gxmember too, resolving a send after 20 s with 2452 chars of scraped TUI — box-drawing characters, the echoed prompt, and leftovers from the previous turn — delivered to the caller as if it were the reply. So the fallback fails in both directions: empty when the turn died, garbage when it had not finished.Suggested direction
BUSY -> DONEinside some floor (a second or two) is a crash signature. Resolve the send as an error naming the member, not as a reply.0 chars scrapedshould be a hard failure — the caller can retry or escalate; it cannot do anything with""that it was told is a success.API Error: 400on screen the whole time. If the scrape is going to be the fallback anyway, pattern-match obvious backend failures and return them as errors instead of as content.readyin/membersand accepting more sends.Point 2 is the one that matters most: today a lost turn and a successful empty answer are indistinguishable to every caller, including a lead deciding whether to delegate again.
Repro
Environment: fleet01,
bridged.jar(started 06:07), herdr 0.8.0 protocol 19, Claude Code 2.1.231, profilelocal=kind: claude-code,baseUrl: https://llm.ltms.dev/anthropic,model: deepseek-v4-flash.Rescued: a partial fix for this existed but was never committed
While cleaning up worker worktrees today I found real work for this issue sitting uncommitted in an abandoned worktree (
.bridged-worktrees/2445bd-3), untouched since 2026-08-28. It was onegit worktree removeaway from being lost.It is now committed and pushed, so it is safe:
worker/cb-164-empty-scrape-false-success-1a80af-3851ebca~/LTMS/.bridged-handover/cb-164-empty-scrape-false-success.patchWhat it does
It covers points 1 and 3 of the suggested direction, and nothing else.
CompletionResolver.MIN_COMPLETED_TURN_NANOS= 2 seconds. The comment gives the reasoning: "A backend rejection can return the pane to idle in about one second. Two seconds is above the measured 1.03s crash while staying low enough not to reject ordinary short model answers." That matches the 1.03 s figure in this issue.BACKEND_ERRORpattern, deliberately narrow:(?i)\bAPI Error\s*:. Its comment says "One stable, explicit backend-failure marker; broader error lists would be brittle." That is the right instinct — a long list of error strings rots.TurnListener.onTurnCompleteandCompletionResolver.resolveBeforePostActiongain anelapsedNanosparameter, with the old no-arg forms kept and defaulting toLong.MAX_VALUE(so an unmeasured turn is never judged too short).Diff: 6 files, +187 / −43, including tests in
CompletionResolverTestandMessageServiceTest.What is NOT done
main, and a rebase conflicts in three files:CompletionResolver.java,CompletionResolverTest.java, andMessageServiceTest.java. Resolving those needs a real read of whatmainhas done toCompletionResolversince 2026-08-28, not a mechanical merge.Next step
Rebase it onto
main, resolve the three conflicts, build, then add point 2 on top. Do not merge it as it stands.There is also a process lesson worth writing down: a worker's work only survives if it is committed. This branch had zero commits, so the usual backup (
git fetch <worktree> <branch>) would have captured nothing.Points 1 and 2 are done and on
main. Closing.3bfa8283bfa8283bfa828(MIN_TURN_NANOS, 2s)4ac688b(this close)4ac688badds a deliberately narrowBACKEND_ERRORpattern ((?i)\bAPI Error\s*:). A scrape that reads cleanly but carries the backend's own rejection is nowWORKER_FAILED, not a completed reply.The failure reason carries the whole pane tail, not just the matched line. The pattern is a heuristic — it also matches a member that forgot
fleet_replywhile reporting about a backend error, which happens here, since members do investigate gateway 400s. Failing that turn is still right; throwing its report away would have been this same bug wearing a failure label.1037 tests,
mvn clean installunpiped, exit 0. Each of the 5 new tests fails with the fix commented out — I ran that myself rather than reading the worker's transcript.Two things I am deliberately leaving behind
A dead-code ordering question, for point 3. The rescued branch had a
visibleTurnraw-screen fallback, so a TUI-hidden error line could still classify whenlastAssistantBlockparses blank. On the current check order the empty-tail fail fires first, so that fallback would never run without also reordering the chain. Reordering is a real behaviour change — it moves which classification wins for a whole class of turns — and it belongs with point 3, not smuggled in here.Why the pattern stays narrow. A growing list of backend error strings rots as backends reword them, and a stale pattern fails in the silent direction: it stops matching, and the scrape resolves as a success again. Point 3 should decide on a real mechanism (a per-profile configured pattern, like
exhaustedPatternalready is), not more hard-coded strings.Process note
851ebca, the branch I rescued from an about-to-be-removed worktree, was mostly already onmain. Someone had committed the same fix directly as3bfa828the same day. I briefed a worker to rebase and extend a branch that was largely redundant, and it cost a full worker run to find that out — the worker caught it, I had not.That is the third time this week (#173, #174) I have acted on a branch without first running
git merge-base --is-ancestor/git log <branch>..origin/main. The rescue case is the one I had not generalised to: I checked that the work was unsaved, and never checked whether it was needed. Recorded so the next rescue starts with that check.