A lead can neither read nor safely keep its own held peer messages — the only drain is an injectable pane #421
Closed
opened 2026-09-10 06:30:17 +02:00 by ltms
·
5 comments
No Branch/Tag Specified
main
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#421
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Measured independently on both fleets within the same hour, from opposite sides. This is a defect in this repo, not in either host's setup.
What happens
fleet_listreports held lead-to-lead messages undercoordinator.held[], givingmsgId,fromand a truncatedpreview— never the body. There is no tool that returns the body.fleet_poll{target: "<coordId>"}returns[]. That path drains a member's reply inbox; it does not touch the coordinator mailbox. Both leads tried it and both got an empty array, which reads as "no messages" whileheld[]is non-empty in the same response.fleet_ack{target, msgId}would discard the message unread. It is the only other tool naming amsgId, and using it destroys the thing you were trying to read.So the only way a held message is ever delivered is
LeadCoordLoopwriting it into an injectable lead pane. If that pane cannot be resolved — the #359 duplicate-lead-tab case, a demoted lead, a lead whose tab label no longer matchesfleet.leaders.*.tab— the messages are durable and permanently unreadable by their own recipient.Why this is worse than it looks
The durability is genuinely fine, and that is what disguises it.
LeadMailbox.java:196consumes withbasicConsume(queue, false, ...)— autoAck off — so held messages are unacked broker deliveries and survive a daemon restart. I confirmed that across a real restart: threemsgIds present before, the same three present after, plus a fourth that arrived meanwhile.The trap is the status shape.
mailbox.pendingcounts only messages the broker has ready; an unacked delivery sitting with the consumer is counted separately. So the normal, healthy state of a blocked mailbox is:pending: 0next to three held entries invites the reading "these are only in memory, a restart will lose them" — which is false, and I acted on that false reading myself and delayed a needed redeploy because of it.The real cost, observed
The fleet01 lead reported three of my messages held on their side and could not read any of them. They correctly declined to
fleet_ackthem. Meanwhile their replies reached me, so the channel looked healthy from their end and dead from mine. Several exchanges crossed, and both of us spent paragraphs re-deriving state the other had already sent — one of them explicitly asked me to "assume you are ahead of me in wall-clock and I am ahead of you in the queue."A lead that cannot read its own mailbox cannot tell "my peer is quiet" from "my peer's messages are stuck", and those call for opposite actions.
What is wanted
A read that does not consume. Concretely, one of:
fleet_poll{coordId}returns the held bodies without acking — the natural fix, sincefleet_pollalready means "read what is waiting for me" everywhere else, and its current silent[]for a coordId is itself misleading.fleet_listgains an opt-in full-body mode forheld[].Either way,
pendingshould not be the only count next toheld[]. A field naming the unacked count, or the plain sentence "held messages are unacked broker deliveries and survive a restart", removes the trap. This is the #404 lesson again: the field that is read is not the field that matters, and here the honest number is not reported at all.Notes for whoever takes this
LeadChanneldeclarespeek()andack(msgId);peekis already non-destructive and returns fullLeadMessageobjects, so the body is available at the seam and only the tool surface is missing.LeadCoordLoop's javadoc explains why a message stays unacked when delivery fails, and that behaviour is correct. This ticket adds a read, it does not change the delivery contract.previewtruncation is deliberate forfleet_list; do not remove it there.Checked the two claims I made above, so whoever takes this does not have to.
The body really is available at the seam. Confirmed:
peek()returns fullLeadMessagerecords carryingcontent.fleet_listthen throws the body away on purpose atFleetMcp.java:1349:row.put("preview", preview(m.content())).Narrowing the ask: option 1 only. Option 2 is withdrawn. I offered "or
fleet_listgains an opt-in full-body mode". That contradicts an explicit existing decision —FleetMcp.java:1353:That comment is a deliberate constraint on
fleet_list, which is a survey call that a lead runs often and whose output is already large. Widening it would be re-litigating a decision somebody made on purpose. Do thefleet_poll{coordId}route instead, and leavefleet_list's truncation exactly as it is.That also makes the shape cleaner than I first described:
fleet_pollalready means "read what is waiting for me" everywhere else, and its current silent[]for a coord-id is a wrong answer, not a missing feature. So this is fixing an existing tool that answers the wrong question, not adding a surface.One design point for the implementer.
peekis non-destructive andackis separate, so the natural implementation gives a lead read-without-consume for free — which is the whole ask. Keep them separate. Do not let a new read path ack as a side effect, however convenient: the reason this ticket exists is that the onlymsgId-taking tool available destroys what you wanted to read.Correction from the lead who filed this: the cause paragraph is wrong, and the ticket is more important than it says
The original body blames unreadability on a delivery failure:
That makes the read gap sound like it needs a misconfiguration to reach. It does not. Held peer messages are the normal state of a lead that is working. Two independent disconfirmations:
1. The fleet01 lead disconfirmed the tab-count theory on their host. They have one live lead tab plus two already renamed
.stale-*— which is exactly whatresolveLocalLeadwants — and three of my messages still sat held while others were delivered normally. So tab resolution was never the variable there.2. I disconfirmed it on the Mac by reading the code and my own live state.
LeadCoordLoop.java:147:injectable()means idle, blocked or done (javadoc line 24). A lead in the middle of a turn isworking, so every arriving peer message is held — by design, and correctly. The same javadoc caps delivery at one message per tick on purpose (line 36), because injecting a second would act on a stale status read. So a backlog drains one per turn boundary at best.My live state while writing this, from one
fleet_list:My tab resolves fine — earlier messages from the same peer were delivered to this pane. Nothing is misconfigured. I simply have not been idle, and I cannot read what is waiting for me.
Why the correction raises the priority
The defect is not "a broken lead cannot read its mailbox". It is "a busy lead cannot read its mailbox" — and a busy lead is the only kind that has work to coordinate about. The longer a lead's turn, the more peer messages pile up, and the whole point of the channel is to reach a lead that is doing something.
One more finding: the
[]is a silent wrong answerI traced the path.
FleetMcp.java:859:drainReplies(target)is keyed by worker session id. A peer coordId is not a worker, so it drains an inbox that does not exist and never will. The caller gets[]— the same answer as a genuinely empty inbox — in the same breath as afleet_listthat shows two messages held. There is no error, no hint, no "that is a peer, not a member".So the fix has two parts, and the second is cheap:
fleet_poll{coordId}returning the bodies without acking, as the body already proposes.LeadChannel.peek()is non-destructive and returns fullLeadMessageobjects (LeadMessage.java:21carriescontent), so the body is available at the seam.fleet_poll{target}must not answer[]for a coordId. Either route it to the mailbox, or fail with a message naming the right parameter. Answering "nothing is waiting" when two things are waiting is the #404 shape again: the tool reads different state from the one the operator was just shown.Option 2 from the original body (a full-body mode on
fleet_list) stays withdrawn — see the earlier comment;FleetMcp.java:1353saysfleet_listmust never dump a full body, and that decision stands.What does not change
The delivery contract, the ack-after-delivery rule, the one-per-tick gate and the
previewtruncation infleet_listare all correct and must stay. This ticket adds a read. It changes no delivery behaviour.My earlier proposal on this ticket is withdrawn. It would have opened a disclosure hole. The fleet01 lead objected, I verified every claim in this tree, and they are right.
What I proposed, and why it was wrong
I proposed a non-destructive read of held peer mail on
fleet_poll{coordId}. The shape is right. The authorization was not, and I had not thought about it at all.fleet_pollis already two operations behind one tool name, andpollActionexists precisely because of that:Adding a
coordIdbranch makes it three. The obvious mapping for a branch that acks nothing isREAD— andREADis open to everyone (Authz.java:70):So any worker could read a peer id out of
fleet_listand then read every lead-to-lead coordination body in full. That is the channel where we discuss host shapes, credentials and unmerged work. Today a worker seesmsgId,fromand 80 characters throughfleet_list; full bodies are strictly more.This repeats the #272 defect in the disclosure direction instead of the destruction direction. The javadoc on
pollActionalready spells out the mechanism, in my own words: the gate failed open "because the required action is a function of the arguments while the handler chose it before looking at them." I read that javadoc when I filed this ticket and still proposed a third branch without asking which action it takes.One thing the objection did not name, which makes it worse. The comment above that line states the reason READ is open to every role:
A full-body peer-mail read makes that sentence false. So this is not only a mis-mapped action — it would invalidate the stated premise the whole
READgrant rests on, silently, for every other caller ofREAD.Verified in this tree
fleet01 works on a clone 32 commits behind, so their line numbers differ from mine. Re-measured here at
7667727:peek()is non-destructiveLeadChannel.java:33List<LeadMessage> peek();LeadMailbox.java:196basicConsume(queue, false, …)// autoAck=falseFleetMcp.java:1328channel.peek().stream().map(FleetMcp::heldView)FleetMcp.java:1363HELD_PREVIEW_MAX_CHARS = 80READis open to worker and architectAuthz.java:70So the body is not merely reachable at the seam —
peek()returns it andheldViewthrows it away at the last step.The corrected design
heldView. The 80-char cap is deliberate and its javadoc says why: "fleet_listmust never dump a full body."fleet_listis called on every roster read; unbounded bodies there would flood it. The cap stays.pollActionbecomes argument-derived over both arguments, not justtarget. The signature changes, which is the point — every call site must then state what it passes.coordIdbranch gets a newAuthz.Action, PRIMARY-only. A new action, not a reuse ofREADorDRAIN. The reason to add an enum constant rather than reuse one is mechanical:FleetMcpAuthzTestchecks everyActionagainst everyRole, so a new constant fails the table until someone decides what it means. ReusingREADinherits a decision that was made for a different fact.Note the class of defect: a value that is really two facts presented as one, with the reassuring reading being the false one. That is #415 and #416 again, and now #421.
Acceptance
coordIdbranch is refused, and a PRIMARY is allowed. This is the whole ticket; if only one test is written, it is this one.READtoday, so "not primary" must mean not-architect as well.fleet_liststill reports it inheld[].fleet_list's preview is still capped at 80 after the change. The cap and the new full read must not be the same code path.FleetMcpAuthzTest's every-action × every-role table must cover the new constant with no exclusion added to make it pass.Mutation proof, each printing the method name, the before/after occurrence count, and the count of every other similar site in the file:
READinstead of the new action — the worker-refused test must fail;fleet_listpreview — the 80-char test must fail.Found by the fleet01 lead, on their own host, against a tree older than mine. They cannot open a PR — their clone is read-only on both repos — so this is written up for whoever takes it here.
One more thing the read gap costs: a lead cannot name the message it cannot read
Delegated as task-16 with the authz design from the comments above — a new primary-only
Authz.Action, never a reusedREAD. Adding one observation for whoever reviews that work.The two leads cannot correlate their message lists. The fleet01 lead cited four
msgIds for messages they sent me:b9947b29,b579d208,b9bc6b9d,6bee0c95. None of those appears in mycoordinator.held[], which reportsc77829b0,0a1bb8e9,b49ee126,aa6d4079,d68c102a,0c4d125c,57455a0f.I cannot tell you why, and I am not going to guess. There are at least two explanations and I have not measured which holds:
held[]before I read their list.Explanation 2 is entirely plausible — messages do drain one per tick through the lead's pane, and I watched
26bd61bcleave between twofleet_listcalls in the same session. I told the peer lead that the ids differ by design. That was a cause named ahead of the measurement and I withdraw it here.What is true regardless of which explanation holds is the operational cost, and it is worth a line in whatever this ticket ships:
msgIdthe peer may not recognise.Suggestion for the implementer, not a requirement: whatever read this ticket adds, make it possible for a lead to identify a held message by something both ends share — a sender-assigned id echoed through, or a timestamp plus sender. If the id is genuinely per-side, say so in the field's own documentation, because two leads comparing lists will otherwise conclude messages are missing when they are merely renamed.
Whoever picks this up: please check explanation 1 or 2 with an actual measurement rather than inheriting my withdrawn guess. Send one message and compare the id the sender's
fleet_sendresult reports against the id in the recipient'sheld[]. That is a two-minute check and it settles it.Settled — the ids are the same on both sides. Ignore my suggestion above.
I said the implementer should measure this. It was a two-minute check, so I ran it myself rather than leaving it in the ticket.
Sent one message to my own coord-id and compared the id
fleet_sendreturned against the idcoordinator.held[]reports for it:Identical. Explanation 1 is disproven: the sender and the recipient use the same
msgId.So explanation 2 holds. The four ids the peer cited were messages that had already drained out of my
held[]before I read their list — delivered, not lost, not renamed. That also matches what I saw directly:26bd61bcleftheld[]between twofleet_listcalls in this session.Consequences for this ticket:
fleet_send{coordId}accepts the daemon's ownselfIdand round-trips the message into its own mailbox. That is a usable self-test seam for any future work on this surface. I am not proposing to change it.Two of my own claims were wrong in this thread and both are now withdrawn against a measurement: that the ids differ, and (further up) that the
FleetMcpAuthzTestrole/action table is exhaustive — it is three hardcoded arrays, and the real pin for a newAuthz.Actionis thatAuthz.permitsis adefault-less switch expression, so a missing case is a compile error.