Members cannot run shell commands: the classifier refuses every one, so they cannot build, test or open a PR (3 members, 2 profiles) #381

Closed
opened 2026-09-09 02:29:11 +02:00 by ltms · 3 comments
Owner

Measured 2026-09-08 UTC across one fan-out of six members on the Mac.

What happened

Two terra members reported that their shell runner refused every command they tried.

The #376 member, verbatim:

The shell runner refused every command with "classifier produced no valid verdict after 3 attempt(s)", including git, mvn clean install, and the commands needed to stage, commit, push, and create the PR.

The #352 member, verbatim:

Build: not run because no code changed; I attempted mvn clean install, but the terminal command classifier refused it before execution.

Both were briefed as implementers. Both wrote correct-looking source changes into their worktrees. Neither could compile, test, stage, commit, push, or open a pull request.

Why this matters more than a lost turn

A member in this state is not obviously broken. It reads files, edits them, and writes a fluent report — it simply cannot execute anything. The #376 member produced a real change and an articulate summary of tests it had "added", and none of it had ever been compiled or run.

I only found out because its report said so plainly. A less careful member would have reported success, and the implementer brief asks for a PR URL, which would have been missing — but a lead skimming a good-looking report could easily miss that. The failure is quiet in exactly the way that costs a merge.

It also silently converts a delegation into lead work. I had to read the change, build it, test it, find its defect, rewrite it and commit it myself. That is the opposite of what delegating is for.

What I have and have not established

Measured: two terra members, same fan-out, same refusal. Both briefed the same way as three sonnet members and one gx member spawned minutes apart from the same lead.

Not established:

  • Whether this is specific to the terra profile, to the OpenAI-backed adapter generally, or to something about that fan-out. A third terra member on #340 did read-only work and never needed a command, so it is not evidence either way.
  • Whether the sonnet members hit it. They were still running when I filed this.
  • What the classifier is actually rejecting. "produced no valid verdict after 3 attempt(s)" is the runner's own wording and I have not traced it to a component. It may not be a fleetd component at all.

I am filing the observation, not a diagnosis. Do not fix anything from this text alone.

Suggested first step — measure, do not fix

Spawn one terra member and one sonnet member with the same trivial brief: run git status and echo ok, and report the exact output or the exact refusal. That settles profile-specificity in one turn and costs almost nothing.

If it reproduces, the next question is where the refusal is raised, because that decides whose bug it is. Nothing here yet shows it is fleetd's.

Consequence for how the fleet is used, until this is understood

terra should not be given implementer work — it cannot complete the hand-off the implementer skill requires. It appears fine for read-only analysis: the #340 terra member produced the best report of the whole fan-out, and needed no commands to do it.

Related

  • #377 — the other failure in this family: a request that never ran, indistinguishable from one that ran and was refused.
Measured 2026-09-08 UTC across one fan-out of six members on the Mac. ## What happened Two `terra` members reported that their shell runner refused every command they tried. The #376 member, verbatim: > The shell runner refused every command with "classifier produced no valid verdict after 3 attempt(s)", including `git`, `mvn clean install`, and the commands needed to stage, commit, push, and create the PR. The #352 member, verbatim: > Build: not run because no code changed; I attempted `mvn clean install`, but the terminal command classifier refused it before execution. Both were briefed as implementers. Both wrote correct-looking source changes into their worktrees. Neither could compile, test, stage, commit, push, or open a pull request. ## Why this matters more than a lost turn A member in this state is not obviously broken. It reads files, edits them, and writes a fluent report — it simply cannot execute anything. The #376 member produced a real change and an articulate summary of tests it had "added", and none of it had ever been compiled or run. I only found out because its report said so plainly. **A less careful member would have reported success**, and the implementer brief asks for a PR URL, which would have been missing — but a lead skimming a good-looking report could easily miss that. The failure is quiet in exactly the way that costs a merge. It also silently converts a delegation into lead work. I had to read the change, build it, test it, find its defect, rewrite it and commit it myself. That is the opposite of what delegating is for. ## What I have and have not established **Measured:** two `terra` members, same fan-out, same refusal. Both briefed the same way as three `sonnet` members and one `gx` member spawned minutes apart from the same lead. **Not established:** - Whether this is specific to the `terra` profile, to the OpenAI-backed adapter generally, or to something about that fan-out. A third `terra` member on #340 did read-only work and never needed a command, so it is not evidence either way. - Whether the `sonnet` members hit it. They were still running when I filed this. - What the classifier is actually rejecting. "produced no valid verdict after 3 attempt(s)" is the runner's own wording and I have not traced it to a component. It may not be a `fleetd` component at all. I am filing the observation, not a diagnosis. Do not fix anything from this text alone. ## Suggested first step — measure, do not fix Spawn one `terra` member and one `sonnet` member with the same trivial brief: run `git status` and `echo ok`, and report the exact output or the exact refusal. That settles profile-specificity in one turn and costs almost nothing. If it reproduces, the next question is where the refusal is raised, because that decides whose bug it is. Nothing here yet shows it is `fleetd`'s. ## Consequence for how the fleet is used, until this is understood `terra` should not be given implementer work — it cannot complete the hand-off the implementer skill requires. It appears fine for read-only analysis: the #340 `terra` member produced the best report of the whole fan-out, and needed no commands to do it. ## Related - #377 — the other failure in this family: a request that never ran, indistinguishable from one that ran and was refused.
Author
Owner

The title is wrong. This is not a terra problem.

A gx member (#377, worker/t377-7f587b-7) hit the identical refusal about 20 minutes after I filed this. Its own words:

the host command classifier is down — every bash command I try (including trivial probes like pwd/date, and one via a subagent whose shell failed identically) is refused with "classifier produced no valid verdict after 3 attempt(s)". That blocks mvn clean install, git add/commit/push, and the PR-creation curl.

So the count is now three members across two profiles: terra ×2, gx ×1. My "suggested first step" above — spawn one terra and one sonnet and compare — was built on a theory this disproves. Do not run that experiment as written.

What the gx data point adds

It rules out two things the terra pair could not:

  • Not profile-specific. Two different profiles, two different backends.
  • Not the model failing to form a command. pwd and date were refused. There is nothing to get wrong about pwd. The refusal is upstream of whatever the member typed.

It also tells us the failure is total, not selective — every command, including a subagent's shell.

Still not established

  • Whether it is time-bound. Three sonnet members ran a full mvn clean install successfully earlier in the same fan-out, on the same host. So it is not "the classifier is down for everyone, always". Either it broke partway through the evening, or it does not affect every member. I did not capture timestamps precisely enough to separate those, and that is the single most useful thing the next person could measure.
  • Where the refusal is raised. Still untraced. "classifier produced no valid verdict after 3 attempt(s)" is the runner's wording. Nothing yet shows this is fleetd's bug rather than the host command classifier's.

Better first step, replacing the one above

Do not compare profiles. Compare time. Spawn two members on any profile, a few minutes apart, brief each to run pwd and report the exact output or exact refusal with a timestamp. If the first succeeds and the second is refused, this is a service that degrades, and the question becomes what recovers it.

If a member is refused, have it report immediately rather than retry — the gx member spent a long stretch probing, which bought nothing.

What it cost, and what stopped it costing more

The gx member had the whole change written and could not build, commit, push or open a PR. I told it to stop probing and hand over, then did the build myself.

Its code failed on the first real build — 7 of its 9 tests failed. Not because the production code was wrong (it was correct), but because logback-test.xml sets dev.ltms.fleet to WARN and the new lines log at INFO, so the test's appender captured an empty list. The member had copied the attach() helper from MemberTrustModelReportTest without the setLevel(INFO) its siblings do at each call site.

That is the real cost of this bug. A member that cannot run anything reports "code-complete" in good faith, and it is not — there is no way for it to know. The three sonnet members tonight all found and fixed real problems during their own build runs. This one could not, and its work needed a fix before it could merge.

Merged in 2830735 after I fixed the harness and proved the guard by mutation.

Consequence, updated

Not "terra should not be given implementer work". Any member that reports this refusal should be told to stop and hand over immediately, and its work must be built by the lead before it is trusted. Read-only analysis is unaffected.

## The title is wrong. This is not a terra problem. A **gx** member (#377, `worker/t377-7f587b-7`) hit the identical refusal about 20 minutes after I filed this. Its own words: > the host command classifier is down — every bash command I try (including trivial probes like `pwd`/`date`, and one via a subagent whose shell failed identically) is refused with "classifier produced no valid verdict after 3 attempt(s)". That blocks `mvn clean install`, `git add/commit/push`, and the PR-creation curl. So the count is now **three members across two profiles**: terra ×2, gx ×1. My "suggested first step" above — spawn one terra and one sonnet and compare — was built on a theory this disproves. Do not run that experiment as written. ## What the gx data point adds It rules out two things the terra pair could not: - **Not profile-specific.** Two different profiles, two different backends. - **Not the model failing to form a command.** `pwd` and `date` were refused. There is nothing to get wrong about `pwd`. The refusal is upstream of whatever the member typed. It also tells us the failure is **total, not selective** — every command, including a subagent's shell. ## Still not established - **Whether it is time-bound.** Three sonnet members ran a full `mvn clean install` successfully earlier in the same fan-out, on the same host. So it is not "the classifier is down for everyone, always". Either it broke partway through the evening, or it does not affect every member. I did not capture timestamps precisely enough to separate those, and that is the single most useful thing the next person could measure. - **Where the refusal is raised.** Still untraced. "classifier produced no valid verdict after 3 attempt(s)" is the runner's wording. Nothing yet shows this is `fleetd`'s bug rather than the host command classifier's. ## Better first step, replacing the one above Do not compare profiles. Compare **time**. Spawn two members on any profile, a few minutes apart, brief each to run `pwd` and report the exact output or exact refusal with a timestamp. If the first succeeds and the second is refused, this is a service that degrades, and the question becomes what recovers it. If a member is refused, have it report immediately rather than retry — the gx member spent a long stretch probing, which bought nothing. ## What it cost, and what stopped it costing more The gx member had the whole change written and could not build, commit, push or open a PR. I told it to stop probing and hand over, then did the build myself. **Its code failed on the first real build — 7 of its 9 tests failed.** Not because the production code was wrong (it was correct), but because `logback-test.xml` sets `dev.ltms.fleet` to `WARN` and the new lines log at `INFO`, so the test's appender captured an empty list. The member had copied the `attach()` helper from `MemberTrustModelReportTest` without the `setLevel(INFO)` its siblings do at each call site. That is the real cost of this bug. A member that cannot run anything reports "code-complete" in good faith, and it is not — there is no way for it to know. The three sonnet members tonight all found and fixed real problems during their own build runs. This one could not, and its work needed a fix before it could merge. Merged in 2830735 after I fixed the harness and proved the guard by mutation. ## Consequence, updated Not "terra should not be given implementer work". **Any member that reports this refusal should be told to stop and hand over immediately**, and its work must be built by the lead before it is trusted. Read-only analysis is unaffected.
ltms changed title from terra members cannot run shell commands: the classifier refuses every one, so they cannot build, test or open a PR to Members cannot run shell commands: the classifier refuses every one, so they cannot build, test or open a PR (3 members, 2 profiles) 2026-09-09 02:53:18 +02:00
Author
Owner

Root cause found, fixed and verified live on the Mac, 2026-09-10.

The axis was never time, and it was never the profile

My earlier comment said "compare time, not profile", and told the next person not to run the profile comparison. That was wrong, and the reason it was wrong is worth naming: I concluded "not profile-specific" from two profiles, terra and gx. Both are kind: opencode. Two profile names looked like two independent data points and were one.

The right experiment holds time constant and varies the backend kind. I spawned one terra and one sonnet in the same minute with the same three-command brief:

  • terra (kind: opencode) → classifier produced no valid verdict after 3 attempt(s):
  • sonnet (kind: claude-code) → all three ran; date returned Thu Sep 10 06:33:34 +07 2026

Same host, same minute, opposite results. Every refusal on record — 2 terra, 1 gx, plus a fourth terra earlier today — is an opencode member. Every success is a Claude Code member.

Where the refusal is raised

Not in fleetd. It is an opencode plugin:

~/.cache/opencode/packages/@common-creation/opencode-plugin-automode@latest/
    node_modules/@common-creation/opencode-plugin-automode/dist/classifier.js:143

throw new ClassificationError(
  `classifier produced no valid verdict after ${maxRetries + 1} attempt(s): ${lastError ?? "unknown error"}`);

The plugin asks a model to approve every bash command. Claude Code members never load it, which is the whole asymmetry.

Why the model call failed

classifier.js sends body.model from OPENCODE_AUTOMODE_MODEL when that variable is set. It is set on this host to openai/gpt-5.4-mini. That model is refused outright by the account:

$ opencode run -m openai/gpt-5.4-mini "..."
Error: Bad Request: {"detail":"The 'gpt-5.4-mini' model is not supported when using Codex with a ChatGPT account."}

So all three attempts failed and the plugin threw. Every command in every opencode member was refused, including pwd and date — which is exactly what the gx data point showed and what I could not explain at the time.

Note the reported error was empty: after 3 attempt(s): with nothing after the colon. The plugin catches the failure and keeps only text.slice(0, 200) of an empty body, so the one line that identifies this in five seconds is thrown away before anyone sees it. That reporting gap is why this took two days instead of ten minutes. It is the plugin's bug, not ours, and I am not filing it here.

The fix, and why it is in our config

fleetd.yaml was passing the broken pin straight into every member. memberCredentials.allow listed OPENCODE_AUTOMODE_MODEL, so the env scrub kept it. Removed. The name is already on the known: block list, so with the allow entry gone the scrub blanks it for members and the operator's own shell is untouched.

With the variable absent, the plugin's own fallback resolves the model from the caller session — the member's working model — and the classifier answers normally.

This keeps the command gate on. It does not disable automode, and it does not weaken what a member is allowed to run. It only stops pointing the gate at a model the account will not serve.

fleetd.yaml is gitignored, so there is no commit. The key is hot-read per spawn, so no restart was needed.

Verification

Same profile that failed, four minutes later, after the change:

pwd    -> /Users/dai.ha/LTMS/.bridged-worktrees/af5a95-3
date   -> Thu Sep 10 06:37:52 +07 2026
git status --porcelain -> (no output)
shell ran

The worktree path matches what fleet_spawn assigned, so this is real output and not a member describing what it thinks would happen. grep -ac 'produced no valid verdict' against the daemon log since the change: 0.

For fleet01

fleet01 runs its own fleetd.yaml. If it also allows OPENCODE_AUTOMODE_MODEL through to members, and the host has that variable set to a model its account refuses, its opencode members are in the same state. Cheap check: spawn one opencode member, brief it to run pwd, and see whether it comes back with output or with the same refusal string. I have not measured fleet01.

What this cost, restated

Four members across two profiles wrote correct-looking code they could never compile. Two of them said so plainly; the implementer brief's missing PR URL was the other signal. The consequence line from my earlier comment stands and is now explained rather than just observed: a member that reports this refusal should stop and hand over immediately, and its work must be built by the lead before it is trusted.

Root cause found, fixed and verified live on the Mac, 2026-09-10. ## The axis was never time, and it was never the profile My earlier comment said "compare time, not profile", and told the next person not to run the profile comparison. That was wrong, and the reason it was wrong is worth naming: I concluded "not profile-specific" from two profiles, `terra` and `gx`. Both are `kind: opencode`. Two profile names looked like two independent data points and were one. The right experiment holds time constant and varies the backend kind. I spawned one `terra` and one `sonnet` in the same minute with the same three-command brief: - `terra` (`kind: opencode`) → `classifier produced no valid verdict after 3 attempt(s):` - `sonnet` (`kind: claude-code`) → all three ran; `date` returned `Thu Sep 10 06:33:34 +07 2026` Same host, same minute, opposite results. Every refusal on record — 2 terra, 1 gx, plus a fourth terra earlier today — is an opencode member. Every success is a Claude Code member. ## Where the refusal is raised Not in `fleetd`. It is an **opencode plugin**: ``` ~/.cache/opencode/packages/@common-creation/opencode-plugin-automode@latest/ node_modules/@common-creation/opencode-plugin-automode/dist/classifier.js:143 throw new ClassificationError( `classifier produced no valid verdict after ${maxRetries + 1} attempt(s): ${lastError ?? "unknown error"}`); ``` The plugin asks a model to approve every bash command. Claude Code members never load it, which is the whole asymmetry. ## Why the model call failed `classifier.js` sends `body.model` from `OPENCODE_AUTOMODE_MODEL` when that variable is set. It is set on this host to `openai/gpt-5.4-mini`. That model is refused outright by the account: ``` $ opencode run -m openai/gpt-5.4-mini "..." Error: Bad Request: {"detail":"The 'gpt-5.4-mini' model is not supported when using Codex with a ChatGPT account."} ``` So all three attempts failed and the plugin threw. Every command in every opencode member was refused, including `pwd` and `date` — which is exactly what the gx data point showed and what I could not explain at the time. Note the reported error was empty: `after 3 attempt(s): ` with nothing after the colon. The plugin catches the failure and keeps only `text.slice(0, 200)` of an empty body, so the one line that identifies this in five seconds is thrown away before anyone sees it. That reporting gap is why this took two days instead of ten minutes. It is the plugin's bug, not ours, and I am not filing it here. ## The fix, and why it is in our config `fleetd.yaml` was passing the broken pin straight into every member. `memberCredentials.allow` listed `OPENCODE_AUTOMODE_MODEL`, so the env scrub kept it. Removed. The name is already on the `known:` block list, so with the allow entry gone the scrub blanks it for members and the operator's own shell is untouched. With the variable absent, the plugin's own fallback resolves the model from the caller session — the member's working model — and the classifier answers normally. **This keeps the command gate on.** It does not disable automode, and it does not weaken what a member is allowed to run. It only stops pointing the gate at a model the account will not serve. `fleetd.yaml` is gitignored, so there is no commit. The key is hot-read per spawn, so no restart was needed. ## Verification Same profile that failed, four minutes later, after the change: ``` pwd -> /Users/dai.ha/LTMS/.bridged-worktrees/af5a95-3 date -> Thu Sep 10 06:37:52 +07 2026 git status --porcelain -> (no output) shell ran ``` The worktree path matches what `fleet_spawn` assigned, so this is real output and not a member describing what it thinks would happen. `grep -ac 'produced no valid verdict'` against the daemon log since the change: **0**. ## For fleet01 fleet01 runs its own `fleetd.yaml`. If it also allows `OPENCODE_AUTOMODE_MODEL` through to members, and the host has that variable set to a model its account refuses, its opencode members are in the same state. Cheap check: spawn one opencode member, brief it to run `pwd`, and see whether it comes back with output or with the same refusal string. I have not measured fleet01. ## What this cost, restated Four members across two profiles wrote correct-looking code they could never compile. Two of them said so plainly; the implementer brief's missing PR URL was the other signal. The consequence line from my earlier comment stands and is now explained rather than just observed: **a member that reports this refusal should stop and hand over immediately, and its work must be built by the lead before it is trusted.**
ltms closed this issue 2026-09-10 01:38:55 +02:00
Author
Owner

Confirmed on a full implementer job, not just a probe. When I closed this I had verified the fix with a short spawn that ran real shell commands. That left the real question open: can a terra member now do a whole unit of work?

It can. A terra member took fleetd #384 end to end on 2026-09-10 and reported:

Tests: mvn -Dtest=EnvAllowListScrubTest test passed. mvn clean install passed:
"Tests run: 1462, Failures: 0, Errors: 0, Skipped: 0" and "BUILD SUCCESS".

It also ran a mutation on its own change, read the failure text, restored the code, committed, and pushed. I merged that work after checking it myself. Every one of those steps is a bash command that this ticket's defect refused for two days.

And the classifier is now doing its job rather than erroring. The same member hit a real denial:

PR: not created. The bridge safety classifier blocked the Gitea API request because it sends GITEA_TOKEN.

That is the difference worth recording. Before the fix, the classifier could not produce a verdict at all and refused pwd and date along with everything else. Now it allows ordinary commands and denies one that would put a token on the wire. The gate was never the problem — the model name was.

The member stopped and said so instead of working around it, which is the correct behaviour. It pushed its branch and put its report in the reply instead of a PR body, so nothing was lost.

Nothing to reopen here. Recording it because "the refusal stopped" and "the member can do a day's work" are two different claims, and I had only measured the first.

**Confirmed on a full implementer job, not just a probe.** When I closed this I had verified the fix with a short spawn that ran real shell commands. That left the real question open: can a `terra` member now do a whole unit of work? It can. A `terra` member took fleetd #384 end to end on 2026-09-10 and reported: ``` Tests: mvn -Dtest=EnvAllowListScrubTest test passed. mvn clean install passed: "Tests run: 1462, Failures: 0, Errors: 0, Skipped: 0" and "BUILD SUCCESS". ``` It also ran a mutation on its own change, read the failure text, restored the code, committed, and pushed. I merged that work after checking it myself. Every one of those steps is a bash command that this ticket's defect refused for two days. **And the classifier is now doing its job rather than erroring.** The same member hit a real denial: > PR: not created. The bridge safety classifier blocked the Gitea API request because it sends GITEA_TOKEN. That is the difference worth recording. Before the fix, the classifier could not produce a verdict at all and refused `pwd` and `date` along with everything else. Now it allows ordinary commands and denies one that would put a token on the wire. The gate was never the problem — the model name was. The member stopped and said so instead of working around it, which is the correct behaviour. It pushed its branch and put its report in the reply instead of a PR body, so nothing was lost. Nothing to reopen here. Recording it because "the refusal stopped" and "the member can do a day's work" are two different claims, and I had only measured the first.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#381