A blocking fleet_send that times out returns no ticket id, so its own advice to "poll status" cannot be followed #801
Closed
opened 2026-10-06 19:33:58 +02:00 by ltms
·
4 comments
No Branch/Tag Specified
main
worker/808-0c71eb-2
worker/801-blocking-send-receipt-abe6ca-14
worker/803-ticket-read-gaps-31efc8-13
worker/797-disable-box-gate-5b5478-12
worker/799-a76336-9
worker/796-a7911a-5
worker/778-e301b0-1
worker/791-fleet-plugin-0-3-0-9e25b0-2
worker/790-observer-to-lead-send-523d45-1
worker/788-eb7dd0-1
worker/782-styling-probe-4ceb10-9
worker/778-12988a-4
worker/782-18f8bc-5
worker/780-faf58b-3
worker/759-authz-comment-and-role-list-e3f0e5-5
worker/756-758-observer-pane-discovery-7e6ffd-1
worker/759-role-model-comments-5d8409-3
worker/743-pane-discovery-ad5b75-5
worker/743-observer-send-4706db-6
worker/749-edge-baseline-28d1a0-3
worker/748-dead-comment-refs-f42ac5-4
worker/737-9c61d3-4
worker/737-a263f3-3
worker/737-20d1d9-1
worker/737-b038f7-2
worker/726-unit2-75cb13-4
worker/737-owner-key-ff061f-10
worker/736-presence-forget-f35144-9
worker/705-observer-14c258-6
worker/722-024c34-5
worker/726-ea34a0-2
worker/726-10cbf0-1
worker/729-5961c6-3
worker/727-ee14ed-3
worker/719-bdd95e-4
worker/702-4f5c7f-2
worker/715-5c43fc-1
worker/721-70f9ea-5
worker/718-99362b-2
worker/task-15-af0d10-12
worker/task-16-50a702-13
worker/task-12-4d0479-9
worker/task-13-823ce2-10
worker/705-ticket-owner-af9928-8
worker/703-list-collaborators-9c06c2-7
worker/669-example-truth-0b303d-6
worker/669-collab-deliverability-9ba859-3
worker/669-collab-reload-report-2a21bd-4
worker/669-7e80a6-1
worker/669-unit-d-efbbd7-1
worker/669-1b786a-1
worker/669-1d1d9f-1
worker/692-4afb9d-2
worker/689-02fced-13
worker/693-cf23fa-14
worker/677-fix-lead-collision-f69073-12
worker/638-fix-overmask-dbb1bf-11
worker/675-5b7478-4
worker/669-unit-a-70cc8f-3
worker/677-8cdaaf-5
worker/638-a7b391-1
worker/683-4536d6-2
worker/651-a75bbe-8
worker/680-20607d-7
worker/664-c12e95-3
worker/668-08534d-4
worker/672-0f2469-2
worker/670-7d1022-1
worker/661-ac7c28-2
worker/664-37fb9b-3
worker/663-remove-3arg-read-3f6783-1
worker/659-remove-dead-backcompat-ba5e6f-1
worker/637-revision-60a488-23
worker/656-redact-regression-tests-892903-19
worker/637-context-gauge-threshold-466eb5-16
worker/639-redact-line-numbers-de4ac4-17
worker/641-set-reformat-guard-6f96a4-18
worker/642-herdr-guard-scope-5de0e4-15
worker/650-javadoc-scope-95f3b3-14
worker/612-01e9f7-13
worker/612-a-r4-quarantine-outage-7ab0e8-5
worker/612-a-r9-r11-capacity-coverage-peers-cfcc79-7
worker/612-a-r10-loophealth-ccc872-8
worker/612-a-r12-turnregistrar-9e3bb7-9
worker/612-a-r5-leadconfigdir-9e70cf-6
lead/config-edit-redact-anchor-wording
worker/config-edit-seam-ca8dc1-1
worker/612-r67-630-lifecycle-290b8d-3
worker/629-625-ports-seams-da7d5d-4
worker/612-r12-exhaustion-f37cd7-1
worker/612-r38-amqp-24b083-2
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#801
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What happens
A blocking
fleet_send(the default,waitunset) to a busy pane returns:There is no ticket id and no msgId in that receipt. So the caller cannot poll, and the only action the message suggests — "retry" — creates a duplicate.
Measured today
I sent one message with the default
wait:truetoterm_65d106559db144at about 19:20:53 and got the receipt above. The daemon's own log shows it did create and resolve a ticket for that send:So a ticket existed, went terminal 16 seconds later, and was pruned 10 minutes after that (
TICKET_TTL_NANOS,MessageService.java:65). Its content is gone and I never had the id.For contrast,
wait:falseto the same pane one minute later returnedaccepted — task delegated. Poll fleet_poll with ticket=task-7da785-70immediately.Why it matters
Three separate costs, in order of how much they hurt:
CLAUDE.mdblock already warns a lead not to re-send because a call looks slow. The receipt advises the opposite.This is close in shape to the hazard in the block's step 5 — "a
fleet_sendto a working member is accepted and returns a ticket, and is then never delivered" — but it is the mirror image: here delivery is recorded asdelivered=trueand it is the handle that is lost, not the message.Suggested fix
Return the ticket id in the timeout receipt, so the two paths read the same way:
Also drop the bare word "retry" from the message, or qualify it, since on this channel a retry is a known way to produce a duplicate rather than a safe repair.
Not verified
Found while running a cross-session message-loss matrix; the box-gate half of that run is on #797.
This is worse than the missing ticket id: the message is lost
I filed this as a usability defect — the receipt withholds the ticket id. A cross-session loss-detection run has now shown the payload itself is discarded.
The measurement
Five messages to one pane (
term_65d106559db144), taggedL-A1/5…L-A5/5, sent in that order.L-A1/5wait:true[no reply within 25000ms — worker busy; retry or poll status], no ticketL-A2/5wait:falseaccepted … ticket=task-7da785-70L-A3/5wait:falseaccepted … ticket=task-7da785-73L-A4/5wait:falseaccepted … ticket=task-7da785-76L-A5/5wait:falseaccepted … ticket=task-7da785-79The receiving session reports L-A2 through L-A5 arrived, in order, no duplicates, each delivered by the fleet mod. For L-A1 it reports: never arrived by any route, not pasted, and
fleet_inboxreturnedcount: 0twice afterwards. It also checked its own session transcript for the 19:21 window and found no garbled or merged prompt — so the message did not arrive in some mangled form either.I did not retry L-A1, so this is a clean single-trial result rather than a duplicate hunt.
delivered=truein the log is falseFIFO settles it. At 19:21:18 that pane's queue still held the phase-1 brief, which was not handed over until 19:21:35. A later message cannot be delivered ahead of an earlier one in the same queue, so
L-A1was not delivered at 19:21:18 anddelivered=trueis wrong.The receiver initially proposed that the 2492-character message was
L-A1. It was not:L-A1's body began "RUN7 L-A1/5 — phase 2 of the loss-detection matrix, lead → anki" and ran about 1100 characters. Different text, different length, and it carried an instruction the receiver demonstrably did not have untilL-A5repeated it.Why this ranks high
On this channel, the blocking send is the unsafe one, and nothing says so. Every
wait:falsesend in the run arrived; the singlewait:truesend is the only lost message in the entire matrix — against three legs that were a perfect 5 of 5 in order with no duplicates.A caller therefore gets: no ticket, no msgId, a log line that claims delivery, and no message. There is no artefact anywhere that would let them notice. I only caught it because the receiver was counting numbered messages.
Revised fix
Beyond returning the ticket id:
wait:falsepath does.delivered=true. If the flag describes the mailbox rather than the pane, it must not use the word "delivered". This is the same confusion as the accepted-but-never-delivered hazard the canonicalCLAUDE.mdblock already warns leads about, and the same shape as the related note that a receipt about the mailbox is not a fact about the pane.wait:truenon-default, or refusing it to a pane whose status is notidle/blocked/done, since thewait:falsepath for the identical target is reliable.Still not verified
Correcting myself. The "Measured today" section above says "the daemon's own log shows it did create and resolve a ticket for that send" and "So a ticket existed, went terminal 16 seconds later, and was pruned". That is wrong, and so is the suggested fix that follows from it.
I have now read the code. A blocking
fleet_sendcreates no ticket at all, ever.new Task(...)appears exactly once inMessageService.java, at line 1346, inside thesendAsyncpath. The blocking path callssend(target, content, timeout, onAccepted, callerOwner), which passes noTask. The log line I quoted —— is
MessageService.java:1036, the blocking send's own debug line. It is not evidence of a ticket. I read a ticket into it because the async path logs look similar, and I filed that as a measurement. It was an inference.So "return the ticket id in the timeout receipt" is not implementable: there is no id to return.
What is actually wrong, after reading the code
1. The receipt names the wrong recovery route, and a real one exists. There are exactly two
inbox.publishcall sites inMessageService(lines 615 and 865). Line 615 is in thereply()path: when a worker callsfleet_replyand no open send or ticket matches it, the reply is published to that target's inbox and is drainable withfleet_poll{target=<sessionId>}. So a latefleet_replyafter a blocking timeout is recoverable — and the receipt never mentions the one call that recovers it. It says "retry or poll status" instead, where "retry" is the documented way to create a duplicate.2. A turn that ends with no
fleet_replyis not recoverable, and that is the real loss.CompletionResolver(lines 440–453) resolves the scrape withrendezvous.resolveCompletion(waiter, completion)and has no inbox fallback on that path. So when the member ends its turn withoutfleet_reply, the scrape goes into a rendezvous waiter and nowhere else. That is what happened in my case: the 301-character scrape was the receiver's answer, and the only route that would have saved it — the inbox — is not on that path.The asymmetry is the defect: a late structured reply survives a blocking timeout; a late unstructured completion does not.
What I have not verified
Whether the waiter
resolveCompletionfound at 19:21:34 was still connected to a caller. The call logged its success, so a waiter existed 16 seconds after the timeout; I did not check whether the blocking send'sinjector.cancel(delivery)atMessageService.java:1033is supposed to have removed it. If it leaves the waiter in place deliberately, that is the seam where an inbox publish belongs.Revised fix
fleet_poll{target=<sessionId>}to drain the inbox. Remove the bare word "retry", or qualify it — on this channel a retry produces a duplicate, and the canonical block already tells a lead not to re-send because a call looks slow.The 25000ms figure is
FleetMcp.DEFAULT_TIMEOUT_MS, clamped byMAX_TIMEOUT_MS = 120_000, andtimeoutMsis a per-call argument. It is not the same bound as the caller's own MCP client timeout, which is roughly 60s and outside fleetd's control — so the two numbers in the original report disagree because they measure different things, and that part of the report was simply two unrelated figures placed side by side.I asked Claude to read this code path; the line numbers and the grep results above are from this session, against
mainat9f4b736.Lead review of
24782fe— one change still neededBoth defects I raised are fixed, and extending the same conditional to the REST receipt was right (
GET /sessions/{id}/repliesis indeedDRAIN). I checkedFleetApp.allowmyself: it returnstruewhenauth == null(FleetApp.java:329-331, "legacy: authorization not enforced"), so theauth == null ||mirror is accurate.The tests are good.
aLateCompletionThatFailsToPublishIsLoggedNotLostassertshasStrandedReply(T)is false after the failed publish — that catches a flag that would otherwise lie.anArchitectsTimedOutSendDoesNotNameTheRepliesRouteItCannotDrainends with a real403on/replies, which is the positive control that proves the receipt was right to stay quiet.What must change: drop the two test-only overloads
FleetMcp.java:1036-1042and1074-1082add short overloads ofsendandanswerthat defaultmayDrainPolltotrue.Measured on your branch: neither has a production caller. The only calls are the overloads' own bodies and 23 call sites in
FleetMcpTest. The production handler always passes the real grant.Two reasons this has to go:
truemeans "namefleet_poll{target}" — the exact receipt this ticket is fixing. If a default has to exist at all it must befalse, because a receipt naming no route is harmless and one naming a refused route is the defect. Failing closed is whatCLAUDE.mdmeans by "fail toward the recoverable error".Do this:
send(..., String callerOwner, boolean mayDrainPoll)andanswer(..., String callerOwner, boolean mayDrainPoll).truewhere the test asserts the route is named,falsewhere it asserts it is not. The two sites already passingfalse(FleetMcpTest.java:488and515) stay as they are.Acceptance:
mvn clean installpasses in your worktree, andgit grep -c 'callerOwner)$'shows no remaining 7-argumentsendor 6-argumentanswerdeclaration. Report the test count and the real build result, including a failure if you get one.Nothing else on this PR needs changing — do not touch
MessageService.java,FleetApp.java, or the REST tests.Fixed and merged to
mainas11352d0. PR #807 is closed (merged locally, not through the forge).Measured on the merge result:
mvn clean installin a throwaway worktree, no[ERROR]lines from maven,Tests run: 2244, Failures: 0, Errors: 0, Skipped: 0tallied from 183 surefire XML files. The staged tree matched the tree I built exactly (517590e).What shipped:
detailnow carryMessageService.NO_TICKET_NO_RESEND— one constant, so the two surfaces cannot drift — and name the one route that recovers a late answer.Authz.Action.DRAINis granted.fleet_poll{target}andGET /sessions/{id}/repliesboth resolve toDRAIN, which is primary-only, whileSENDreaches an architect, a collaborator and an observer too. Everyone else is told the reply cannot be recovered on their channel.MessageService.strandLateResolutionpublishes a late turn-completion to the target's inbox. OnlyKind.COMPLETIONis routed: an explicit latefleet_replyalready lands there viareply()'s own lookup oncerendezvous.closehas run, so routing it here as well would publish it twice. I checked thatcloseruns insend()'sfinally(MessageService.java:1061-1064) before concluding it.whenCompleteaction is captured by the discarded dependent stage and reaches no caller.Documentation, per this repo's rule that the prompt ships with the code:
CLAUDE.mdand the wiki template now state that a blocking send creates no ticket (a2cbf50, byte-identical check passes), andwiki/11-Features.mdhas the entry (fleetd.wiki f812689).One correction to this ticket's own record
My earlier comment already retracted the claim that the daemon log showed a ticket being created for a blocking send. Worth stating why it was wrong, because the mistake is reusable: both the blocking and the async paths log
send to term_…, so their output is not a discriminator, and I read a ticket into a line written by the path that never creates one.grep -n 'new Task(' MessageService.javareturns exactly one hit, insidesendAsync. The original suggested fix — "return the ticket id in the timeout receipt" — was therefore unimplementable, and the real finding turned out to be better than the invented one.Closing.