The same text-match-drives-a-cooldown shape exists for ExhaustionSink, where the penalty is 30x longer #348

Closed
opened 2026-09-04 11:13:50 +02:00 by ltms · 1 comment
Owner

Reported by the #339 implementer under that ticket's "find what else has this shape" tail, and
correctly left unfixed as out of scope.

The shape

#339 fixed this for backendErrorSink: a free-text pattern match against a member's pane was enough
to record a credential outage, so a member writing prose about an error could cool off a
credential it shared with other profiles. The fix requires the pattern at the start of its matched
line, ignoring leading terminal chrome.

CompletionResolver runs the same kind of classification for exhaustion:

// CompletionResolver.java:374 and :472
String matchedLine = exhausted == null ? null : firstMatchingLine(assistantBlock, exhausted);

and notifies ExhaustionSink on a match. That path did not get #339's start-of-line check.

Why this is worth its own ticket rather than a copy of #339

The penalty is much heavier. Per the project's own notes:

  • a backend-error cooldown is 60 seconds, fixed;
  • a quarantine from exhaustion is 1800 seconds by default.

So the same false positive costs thirty times as much here. A member that ends a turn without
fleet_reply while discussing a usage limit could take a credential out for half an hour.

Not verified

I have not measured this one. #339 was confirmed by reading the pattern and by a probe; this is
the analogous code path, reported by an implementer, and I have only read that the sink call exists
on it. Treat the claim as unproven until someone runs it.

First job on this ticket is the measurement, exactly as #339's was: drive a normal member report
that mentions an exhaustion phrase through resolve, and assert what reaches ExhaustionSink. If
the sink does not fire, say so and close this — a fix for a bug that is not there is worse than no
fix.

Goal and invariants

Goal: a member's own prose about running out of capacity must not quarantine a credential.

Invariants:

  1. A genuine exhaustion must still reach the sink. This direction is worse than the false
    positive, and more so here than in #339 — an unrecorded exhaustion means the fleet keeps spawning
    into a credential that has no capacity, and every one of those spawns fails.
  2. Keep failing the send on a match, and keep carrying the whole pane tail. Same reasoning as #339.

Candidate mechanism, as a candidate only: reuse startsWithBackendError — it is already written,
already skips leading terminal chrome, and is already tested against both a real error behind chrome
and a member's prose. Renaming it to something pattern-neutral would be part of that. Decide it
yourself; if the exhaustion patterns are shaped differently enough that the same check is wrong for
them, say why and propose something else.

One thing to get right, learned the hard way in #339

#339's first attempt used a bare lookingAt(), which rejected a genuine error line rendered as
"| 503 Service Unavailable: ..." — the send failed and the outage went unrecorded. It was caught
on merge by a probe, not by the test suite. Whatever check you use, test it against a real error
line carrying leading TUI chrome
, not only against a clean one. The raw-scrape path's own comment
says to expect chrome there.

Related: #339 (the same shape, fixed, with the chrome correction in 6a81417).

Reported by the #339 implementer under that ticket's "find what else has this shape" tail, and correctly left unfixed as out of scope. ## The shape #339 fixed this for `backendErrorSink`: a free-text pattern match against a member's pane was enough to record a credential outage, so a member writing prose *about* an error could cool off a credential it shared with other profiles. The fix requires the pattern at the start of its matched line, ignoring leading terminal chrome. `CompletionResolver` runs the *same* kind of classification for exhaustion: ```java // CompletionResolver.java:374 and :472 String matchedLine = exhausted == null ? null : firstMatchingLine(assistantBlock, exhausted); ``` and notifies `ExhaustionSink` on a match. That path did **not** get #339's start-of-line check. ## Why this is worth its own ticket rather than a copy of #339 The penalty is much heavier. Per the project's own notes: - a backend-error cooldown is **60 seconds**, fixed; - a quarantine from exhaustion is **1800 seconds** by default. So the same false positive costs thirty times as much here. A member that ends a turn without `fleet_reply` while discussing a usage limit could take a credential out for half an hour. ## Not verified I have **not** measured this one. #339 was confirmed by reading the pattern and by a probe; this is the analogous code path, reported by an implementer, and I have only read that the sink call exists on it. Treat the claim as unproven until someone runs it. **First job on this ticket is the measurement**, exactly as #339's was: drive a normal member report that mentions an exhaustion phrase through `resolve`, and assert what reaches `ExhaustionSink`. If the sink does not fire, say so and close this — a fix for a bug that is not there is worse than no fix. ## Goal and invariants **Goal:** a member's own prose about running out of capacity must not quarantine a credential. **Invariants:** 1. **A genuine exhaustion must still reach the sink.** This direction is worse than the false positive, and more so here than in #339 — an unrecorded exhaustion means the fleet keeps spawning into a credential that has no capacity, and every one of those spawns fails. 2. Keep failing the send on a match, and keep carrying the whole pane tail. Same reasoning as #339. **Candidate mechanism, as a candidate only:** reuse `startsWithBackendError` — it is already written, already skips leading terminal chrome, and is already tested against both a real error behind chrome and a member's prose. Renaming it to something pattern-neutral would be part of that. Decide it yourself; if the exhaustion patterns are shaped differently enough that the same check is wrong for them, say why and propose something else. ## One thing to get right, learned the hard way in #339 #339's first attempt used a bare `lookingAt()`, which rejected a genuine error line rendered as `"| 503 Service Unavailable: ..."` — the send failed and the outage went unrecorded. It was caught on merge by a probe, not by the test suite. **Whatever check you use, test it against a real error line carrying leading TUI chrome**, not only against a clean one. The raw-scrape path's own comment says to expect chrome there. Related: #339 (the same shape, fixed, with the chrome correction in `6a81417`).
Author
Owner

Merged as c1c8794 (--no-ff, the branch was behind main), plus my own follow-up 4a5030a.
main is green at Tests run: 1370, Failures: 0, Errors: 0, Skipped: 0.

The measurement confirmed it. A probe drove a normal member report through resolve and
ExhaustionSink fired:

completion for term_a classified BACKEND_EXHAUSTED ...
  I reviewed capacity handling. The usage limit has been reached means no more work can start.

So a member writing prose about a usage limit really could quarantine a credential for 1800
seconds. This ticket was filed unproven; it is proven now.

The worker rejected my candidate mechanism, and it was right. I suggested reusing
startsWithBackendError. It used a looser rule instead: the match must sit before the line's first
sentence ending, with no start-of-line requirement. Its reason was that an exhaustedPattern may
name only the decisive words — "usage limit has been reached" without the leading "The" — so a
start-of-line check would reject the genuine refusal.

I doubted that and tested it rather than argue. Mutation FF: replace the looser rule with
startsWithBackendError. Full suite, unpiped:

CompletionResolverTest.aRealExhaustionBehindTerminalChromeStillNotifiesTheSink:573        expected: <1> but was: <0>
CompletionResolverTest.aWinningBackendExhaustedClassificationNotifiesTheExhaustionSink:552 expected: <1> but was: <0>
CompletionResolverTest.anExhaustionLineWithNoMarkerAndLeadingChromeIsClassifiedFromTheRawScrapeAndNotifiesTheSink:751 expected: <1> but was: <0>
CompletionResolverTest.exhaustionKeepsWinningOverBackendErrorEvenWhenBothPatternsMatchTheSameLine:1011 expected: <1> but was: <0>
BUILD FAILURE

Four tests, three of them pre-existing. My candidate would have stopped recording genuine
exhaustions — invariant 1's direction, the one this ticket called worse. The looser rule also
accepts a strict superset of what the strict rule accepts, so adopting it cannot add a false
negative. That argument is now written into the method's javadoc.

What I changed on merge (4a5030a), two things.

  1. Deleted a step that cannot fire. The new method copied startsWithBackendError's
    leading-chrome loop. Mutation EE: delete that loop. All 1369 tests stayed green — and
    they must, because the scan only looks for ., ! and ?, and no terminal chrome character is
    one of those. A step that cannot change the result is worse than no step: the next reader takes
    it as evidence chrome was handled. The javadoc now records that measurement instead of claiming
    chrome is skipped.
  2. Put #339's history back. The change moved my startsWithBackendError javadoc onto the new
    method and left a one-line summary behind. That javadoc records the measurement that a bare
    lookingAt rejected a real | 503 Service Unavailable: line and the outage went unrecorded —
    the reason the method looks the way it does. Restored.

I also added anExhaustionPatternCarryingItsLeadingWordsStillNotifiesTheSink. Every test here uses
a pattern without the leading "The", but the live fleet configures
exhaustedPattern: "The usage limit has been reached" — the shape the worker could not see, because
fleetd.yaml is not in the repo. That is my gap, not the worker's: I should have put the deployed
pattern in the brief.

Known limits, stated rather than implied. It is still a heuristic. Prose whose first sentence
carries the pattern still notifies the sink. A genuine refusal behind an earlier full stop — a
hostname, a version number — still does not. Both are in the javadoc; neither is fixed here.

No PR to close. The worker reported that its authenticated POST to the forge was refused by the
safety classifier, so it pushed the branch and reported in fleet_reply instead. It said so plainly
rather than claiming a PR existed, which is the right call. The branch worker/fd348-f1ab27-4 is
merged.

Closing.

Merged as `c1c8794` (`--no-ff`, the branch was behind main), plus my own follow-up `4a5030a`. `main` is green at `Tests run: 1370, Failures: 0, Errors: 0, Skipped: 0`. **The measurement confirmed it.** A probe drove a normal member report through `resolve` and `ExhaustionSink` fired: ``` completion for term_a classified BACKEND_EXHAUSTED ... I reviewed capacity handling. The usage limit has been reached means no more work can start. ``` So a member writing prose about a usage limit really could quarantine a credential for 1800 seconds. This ticket was filed unproven; it is proven now. **The worker rejected my candidate mechanism, and it was right.** I suggested reusing `startsWithBackendError`. It used a looser rule instead: the match must sit before the line's first sentence ending, with no start-of-line requirement. Its reason was that an `exhaustedPattern` may name only the decisive words — `"usage limit has been reached"` without the leading `"The"` — so a start-of-line check would reject the genuine refusal. I doubted that and tested it rather than argue. **Mutation FF**: replace the looser rule with `startsWithBackendError`. Full suite, unpiped: ``` CompletionResolverTest.aRealExhaustionBehindTerminalChromeStillNotifiesTheSink:573 expected: <1> but was: <0> CompletionResolverTest.aWinningBackendExhaustedClassificationNotifiesTheExhaustionSink:552 expected: <1> but was: <0> CompletionResolverTest.anExhaustionLineWithNoMarkerAndLeadingChromeIsClassifiedFromTheRawScrapeAndNotifiesTheSink:751 expected: <1> but was: <0> CompletionResolverTest.exhaustionKeepsWinningOverBackendErrorEvenWhenBothPatternsMatchTheSameLine:1011 expected: <1> but was: <0> BUILD FAILURE ``` Four tests, three of them pre-existing. My candidate would have stopped recording genuine exhaustions — invariant 1's direction, the one this ticket called worse. The looser rule also accepts a strict superset of what the strict rule accepts, so adopting it cannot add a false negative. That argument is now written into the method's javadoc. **What I changed on merge (`4a5030a`), two things.** 1. **Deleted a step that cannot fire.** The new method copied `startsWithBackendError`'s leading-chrome loop. **Mutation EE**: delete that loop. All **1369** tests stayed green — and they must, because the scan only looks for `.`, `!` and `?`, and no terminal chrome character is one of those. A step that cannot change the result is worse than no step: the next reader takes it as evidence chrome was handled. The javadoc now records that measurement instead of claiming chrome is skipped. 2. **Put #339's history back.** The change moved my `startsWithBackendError` javadoc onto the new method and left a one-line summary behind. That javadoc records the measurement that a bare `lookingAt` rejected a real `| 503 Service Unavailable:` line and the outage went unrecorded — the reason the method looks the way it does. Restored. I also added `anExhaustionPatternCarryingItsLeadingWordsStillNotifiesTheSink`. Every test here uses a pattern *without* the leading `"The"`, but the live fleet configures `exhaustedPattern: "The usage limit has been reached"` — the shape the worker could not see, because `fleetd.yaml` is not in the repo. That is my gap, not the worker's: I should have put the deployed pattern in the brief. **Known limits, stated rather than implied.** It is still a heuristic. Prose whose *first* sentence carries the pattern still notifies the sink. A genuine refusal behind an earlier full stop — a hostname, a version number — still does not. Both are in the javadoc; neither is fixed here. **No PR to close.** The worker reported that its authenticated POST to the forge was refused by the safety classifier, so it pushed the branch and reported in `fleet_reply` instead. It said so plainly rather than claiming a PR existed, which is the right call. The branch `worker/fd348-f1ab27-4` is merged. Closing.
ltms closed this issue 2026-09-04 11:58:11 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#348