Review-shaped briefs keep ending with no fleet_reply, and the pane scrape hides it — the charter, the skill and an explicit brief line all failed to prevent it #698

Open
opened 2026-10-03 23:48:19 +02:00 by ltms · 1 comment
Owner

What I measured today

A reviewer member (term_65cf67175a84b55, profile sonnet, role reviewer, ticket #669) finished its turn without calling fleet_reply. fleet_poll{ticket:"task-3"} returned:

[done — worker finished without a structured fleet_reply; transcript tail follows]

followed by a scrape of its pane. The content was coherent and the finding was usable, so the fallback worked this time. That is the problem: it looks like success.

Three separate mechanisms were in place and all three failed to produce a reply:

  1. the launcher's REPLY_CHARTER, which exists for exactly this rule and reaches every spawned member;
  2. the reviewer skill, which specifies the finding format and the reply;
  3. an explicit line in my own brief — "End the turn with exactly one fleet_reply carrying your single most important real finding", plus "Anything only printed in your terminal reaches nobody."

So this is not a case of a brief forgetting to ask. The ask was present, at three layers, and the turn still ended without the call.

A second reviewer in the same batch (term_65cf671dffb3b56, profile terra, same brief shape) did reply correctly. Same role, same skill, same brief template, different profile — so whatever this is, it is not deterministic across backends.

Why the fallback is not a fix

Two reasons, from the project's own documented behaviour:

  • The scrape returns at most the last 4000 characters. A clipped scrape is marked partial, but the missing text is gone. Today's report was short enough to survive. A longer one would have reached me with its end cut off, and the part a reviewer puts last is usually the conclusion.
  • The scrape returns whatever is on the pane, not the structured finding. Today's arrived as prose with a numbered list rather than the skill's format, so a lead parsing it mechanically would get nothing.

There is also a quieter cost: the scrape makes the failure invisible in aggregate. Nothing counts it, so "members sometimes do not reply" never becomes a number anyone can act on.

Prior occurrences — cited, not measured by me

I did not measure these; I am recording where they are already written down so the count is not rebuilt from scratch:

  • CLAUDE.md's project addendum states this "has already cost three workers' turns: each wrote a good report to its terminal and ended the turn with no fleet_reply, and the scrape returned the tail of the brief instead." Those are attributed to naming reviewer for a multi-finding sweep, which is not what happened today — today's brief asked for exactly one finding, matching the skill.
  • My own session notes record 2 of 3 capped reviewer briefs ending with no reply, and one uncapped dev brief doing the same. The uncapped case matters: it means the capped-format theory does not cover every instance.

Today's case is a further data point against the capped-format explanation, because the brief and the skill agreed on one finding.

What I have NOT diagnosed

I did not find the cause, and I am not proposing a fix. I stopped both reviewer panes before thinking to check the one instrument that would have helped, so these are hypotheses with no evidence either way:

  • Did the charter actually reach the member? fleet_list reports charterSha256 per member. Nobody checked it for this session, and it is the cheapest first measurement: a missing or divergent digest would move this from a model-behaviour question to a delivery bug. Check it while the member is alive — I did not.
  • Did the member try and get refused? fleet_reply is only-as-itself; a refused call would be in the daemon log. I did not grep for one.
  • Did the turn end before the call? If something ends the turn early, the member never reaches its last step.

Please do not accept "the model forgot" without ruling out the first two. That explanation is the comfortable one and it is unfalsifiable as stated.

Suggested first step

Add a counter and a log line for the fallback path, so a turn resolved by scrape rather than by fleet_reply is countable. Right now the only trace is the string in the poll output, which a lead reads once and nobody aggregates. A rate would tell us whether this is rare and model-shaped or common and structural — and a count with no denominator cannot tell those apart.

Filed from #669's Unit B review round rather than buried in that ticket's comments, because it is about the bridge rather than about collaborators.

## What I measured today A `reviewer` member (`term_65cf67175a84b55`, profile `sonnet`, role `reviewer`, ticket #669) finished its turn without calling `fleet_reply`. `fleet_poll{ticket:"task-3"}` returned: ``` [done — worker finished without a structured fleet_reply; transcript tail follows] ``` followed by a scrape of its pane. The content was coherent and the finding was usable, so **the fallback worked this time**. That is the problem: it looks like success. Three separate mechanisms were in place and all three failed to produce a reply: 1. the launcher's `REPLY_CHARTER`, which exists for exactly this rule and reaches every spawned member; 2. the `reviewer` skill, which specifies the finding format and the reply; 3. **an explicit line in my own brief** — "End the turn with exactly one `fleet_reply` carrying your single most important real finding", plus "Anything only printed in your terminal reaches nobody." So this is not a case of a brief forgetting to ask. The ask was present, at three layers, and the turn still ended without the call. A second `reviewer` in the same batch (`term_65cf671dffb3b56`, profile `terra`, same brief shape) **did** reply correctly. Same role, same skill, same brief template, different profile — so whatever this is, it is not deterministic across backends. ## Why the fallback is not a fix Two reasons, from the project's own documented behaviour: - The scrape returns **at most the last 4000 characters**. A clipped scrape is marked partial, but the missing text is gone. Today's report was short enough to survive. A longer one would have reached me with its end cut off, and the part a reviewer puts last is usually the conclusion. - The scrape returns whatever is on the pane, not the structured finding. Today's arrived as prose with a numbered list rather than the skill's format, so a lead parsing it mechanically would get nothing. There is also a quieter cost: the scrape makes the failure invisible in aggregate. Nothing counts it, so "members sometimes do not reply" never becomes a number anyone can act on. ## Prior occurrences — cited, not measured by me I did not measure these; I am recording where they are already written down so the count is not rebuilt from scratch: - `CLAUDE.md`'s project addendum states this "has already cost three workers' turns: each wrote a good report to its terminal and ended the turn with no `fleet_reply`, and the scrape returned the tail of the brief instead." Those are attributed to naming `reviewer` for a multi-finding sweep, which is **not** what happened today — today's brief asked for exactly one finding, matching the skill. - My own session notes record 2 of 3 capped `reviewer` briefs ending with no reply, and one **uncapped** `dev` brief doing the same. The uncapped case matters: it means the capped-format theory does not cover every instance. Today's case is a further data point against the capped-format explanation, because the brief and the skill agreed on one finding. ## What I have NOT diagnosed I did not find the cause, and I am not proposing a fix. I stopped both reviewer panes before thinking to check the one instrument that would have helped, so these are hypotheses with no evidence either way: - **Did the charter actually reach the member?** `fleet_list` reports `charterSha256` per member. Nobody checked it for this session, and it is the cheapest first measurement: a missing or divergent digest would move this from a model-behaviour question to a delivery bug. Check it *while the member is alive* — I did not. - **Did the member try and get refused?** `fleet_reply` is only-as-itself; a refused call would be in the daemon log. I did not grep for one. - **Did the turn end before the call?** If something ends the turn early, the member never reaches its last step. Please do not accept "the model forgot" without ruling out the first two. That explanation is the comfortable one and it is unfalsifiable as stated. ## Suggested first step Add a counter and a log line for the fallback path, so a turn resolved by scrape rather than by `fleet_reply` is countable. Right now the only trace is the string in the poll output, which a lead reads once and nobody aggregates. A rate would tell us whether this is rare and model-shaped or common and structural — and a count with no denominator cannot tell those apart. Filed from #669's Unit B review round rather than buried in that ticket's comments, because it is about the bridge rather than about collaborators.
Author
Owner

Charter delivery is ruled out as the cause. Measured on two live reviewers, before stopping either pane.

The previous lead noted it had destroyed the best evidence by stopping both reviewer panes before reading charterSha256, and said not to accept "the model forgot" without ruling out charter delivery first. I had two reviewers running on #669 Unit C, so I took the reading while both were alive.

What I measured

Two reviewers, spawned for the same review, on two different profiles:

sessionId profile charterSource charterSha256 charterBytes
term_65cf72e1fecb858 sonnet fleet.charters.reviewer 38d3821e13a8aa1c0b10f09305c4d1e09ea07b31108509905eae58d0c093315f 1346
term_65cf72e94b39959 terra fleet.charters.reviewer 38d3821e13a8aa1c0b10f09305c4d1e09ea07b31108509905eae58d0c093315f 1346

Identical digest, identical byte count. The profile is not part of the digest, so these are two readings that differ in the axis being tested rather than one instrument used twice.

The digest covers the reply rule — checked, not assumed

A digest over the role charter alone would say nothing about REPLY_CHARTER, which is where the reply rule actually lives. HerdrPeerLauncher.java:500 composes them:

String replyCharter = cfg.hasMcp() ? REPLY_CHARTER : null;
String charter = roleCharter == null ? replyCharter
        : replyCharter == null ? roleCharter : roleCharter + "\n\n" + replyCharter;

So I measured the two pieces and checked them against the reported number:

reviewer roleCharter length  = 545
REPLY_CHARTER length         = 799
roleCharter + 2 + reply      = 1346      <- the "\n\n" join
fleet_list charterBytes      = 1346

The arithmetic matches exactly, so charterBytes/charterSha256 describe the composed charter, REPLY_CHARTER included. That was worth checking rather than assuming: a receipt that reads a different source than the behaviour is a defect shape this project has already hit.

I also confirmed the reply rule is not in any role charter, and that this is by design rather than an omission:

charter architect : fleet_reply mentions = 0
charter dev       : fleet_reply mentions = 0
charter reviewer  : fleet_reply mentions = 0
charter hunter    : fleet_reply mentions = 0
positive control  : "One finding" in the reviewer charter = 1

All four are zero and the control fires, so the pattern works. It matches the layering rule in CLAUDE.md: the reply charter reaches every spawned member from the launcher, not from the per-role config.

What this does and does not establish

Established: a reviewer that ends its turn without fleet_reply did receive the instruction. Both reviewers got a byte-identical 1346-byte charter containing the full 799-byte REPLY_CHARTER, which states the rule explicitly ("you MUST end EVERY turn by calling fleet_reply"). Charter delivery, charter divergence between profiles, and a missing reply rule are all off the list.

Not established: why the turn ends without the call. I have not reproduced the failure in this session — both of these reviewers were still running when I took the reading, so this is evidence about delivery only, not about the failing case. I did not inspect what the member's own context looks like at the end of a long turn.

One branch still worth ruling out, which I did not test: replyCharter is null when cfg.hasMcp() is false. For these two profiles it is plainly true, since the digest includes REPLY_CHARTER. But a profile with no MCP mount would get a role charter and no reply rule at all, and its charterBytes would equal the role charter length alone — 545 for a reviewer. That is a cheap check for anyone picking this up: if a member that failed to reply reported charterBytes of 545, the cause is that branch and not the model.

Suggested next step

Take charterSha256 and charterBytes from the roster while the member is still alive on the next occurrence, and record them on this ticket next to whether the reply arrived. The instrument is sound, so the remaining work is catching a failure with the reading already taken.

## Charter delivery is ruled out as the cause. Measured on two live reviewers, before stopping either pane. The previous lead noted it had destroyed the best evidence by stopping both reviewer panes before reading `charterSha256`, and said not to accept "the model forgot" without ruling out charter delivery first. I had two reviewers running on #669 Unit C, so I took the reading **while both were alive**. ### What I measured Two reviewers, spawned for the same review, on **two different profiles**: | sessionId | profile | charterSource | charterSha256 | charterBytes | |---|---|---|---|---| | `term_65cf72e1fecb858` | `sonnet` | `fleet.charters.reviewer` | `38d3821e13a8aa1c0b10f09305c4d1e09ea07b31108509905eae58d0c093315f` | 1346 | | `term_65cf72e94b39959` | `terra` | `fleet.charters.reviewer` | `38d3821e13a8aa1c0b10f09305c4d1e09ea07b31108509905eae58d0c093315f` | 1346 | Identical digest, identical byte count. The profile is not part of the digest, so these are two readings that differ in the axis being tested rather than one instrument used twice. ### The digest covers the reply rule — checked, not assumed A digest over the *role* charter alone would say nothing about `REPLY_CHARTER`, which is where the reply rule actually lives. `HerdrPeerLauncher.java:500` composes them: ```java String replyCharter = cfg.hasMcp() ? REPLY_CHARTER : null; String charter = roleCharter == null ? replyCharter : replyCharter == null ? roleCharter : roleCharter + "\n\n" + replyCharter; ``` So I measured the two pieces and checked them against the reported number: ``` reviewer roleCharter length = 545 REPLY_CHARTER length = 799 roleCharter + 2 + reply = 1346 <- the "\n\n" join fleet_list charterBytes = 1346 ``` The arithmetic matches exactly, so `charterBytes`/`charterSha256` describe the **composed** charter, `REPLY_CHARTER` included. That was worth checking rather than assuming: a receipt that reads a different source than the behaviour is a defect shape this project has already hit. I also confirmed the reply rule is **not** in any role charter, and that this is by design rather than an omission: ``` charter architect : fleet_reply mentions = 0 charter dev : fleet_reply mentions = 0 charter reviewer : fleet_reply mentions = 0 charter hunter : fleet_reply mentions = 0 positive control : "One finding" in the reviewer charter = 1 ``` All four are zero and the control fires, so the pattern works. It matches the layering rule in `CLAUDE.md`: the reply charter reaches every spawned member from the launcher, not from the per-role config. ### What this does and does not establish **Established:** a reviewer that ends its turn without `fleet_reply` *did* receive the instruction. Both reviewers got a byte-identical 1346-byte charter containing the full 799-byte `REPLY_CHARTER`, which states the rule explicitly ("you MUST end EVERY turn by calling fleet_reply"). Charter delivery, charter divergence between profiles, and a missing reply rule are all off the list. **Not established:** why the turn ends without the call. I have not reproduced the failure in this session — both of these reviewers were still running when I took the reading, so this is evidence about delivery only, not about the failing case. I did not inspect what the member's own context looks like at the end of a long turn. **One branch still worth ruling out, which I did not test:** `replyCharter` is `null` when `cfg.hasMcp()` is false. For these two profiles it is plainly true, since the digest includes `REPLY_CHARTER`. But a profile with no MCP mount would get a role charter and **no reply rule at all**, and its `charterBytes` would equal the role charter length alone — 545 for a reviewer. That is a cheap check for anyone picking this up: if a member that failed to reply reported `charterBytes` of 545, the cause is that branch and not the model. ### Suggested next step Take `charterSha256` and `charterBytes` from the roster **while the member is still alive** on the next occurrence, and record them on this ticket next to whether the reply arrived. The instrument is sound, so the remaining work is catching a failure with the reading already taken.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#698