Exhaustion detection is armed on 2 of 8 profiles, and the default profile is not one of them #479

Open
opened 2026-09-10 23:21:03 +02:00 by ltms · 3 comments
Owner

⚠ READ THE COMMENTS FIRST — THIS BODY IS SUPERSEDED

This body's premise is wrong and its acceptance criteria are partly void. I filed it
asking for a startup report that already exists, and the design question it poses was already
decided and shipped as #415 (be123d0).

There are three comments and each corrects the one before it. The third is the live scope.
Do not work from this body.

The whole remaining deliverable: nothing joins the boot log's coverage line
(not configured: [… sonnet …]) to the default-profile selection (default: sonnet). Two facts,
two places, no single output that puts them next to each other. Same on the fleet01 host, where
their config documents xf as recurrently exhausting and xf is their default. Also decide
whether a gap covering the default profile deserves more than INFO.

Kept below unedited, because the measurements in it are still good and the corrections only make
sense against what they correct.


What I measured

On the live daemon (pid 52482, jar 82ebc1cc7047, started 04:17:59 on 49a404d):

fleet_profiles
  default: "sonnet"
  modelGateArmed: true
  exhaustionDetectionArmed:
    terra: true   sol: true
    sonnet: false  opus: false  xf: false  gx: false  local: false  local-direct: false

The cause is in the config, not the code. exhaustedPattern is set on exactly two profiles:

$ grep -n 'exhaustedPattern' fleetd/fleetd.yaml
152:    # exhaustedPattern is a DEFERRED key: a running daemon keeps what it started with.
153:    exhaustedPattern: "The usage limit has been reached"
174:    # exhaustedPattern is a DEFERRED key: a running daemon keeps what it started with.
175:    exhaustedPattern: "The usage limit has been reached"

Line 153 is under sol: (starts at :133), line 175 under terra: (starts at :155). Control: the
same file lists 8 profiles at profiles: (:9), so the grep read a real file and found two of
eight.

Why this matters more than a coverage number

The default profile is sonnet, and it is one of the six with detection off. So is opus,
which the lead itself runs on.

When a subscription limit hits a profile with no exhaustedPattern, the whole chain that was
built for this case does nothing:

  • no quarantine, so BackendQuarantine never escalates and never sets free: 0
  • no usageLimitFixWarning, so nothing names the model to turn off
  • fleet_profiles reports the profile as normally available

The operator asked for exactly this behaviour in their own words: "runtime model limit
monitoring, off when subscription limit reach and on when the limit lifted."
Parts 1 and 2 of
that ask are shipped and live — models.allow is a central list, enabled: false is honoured at
spawn and in placement, and the boot log confirms model gate (fleetd #422): armed (models: block present; 0 models currently turned off). Part 3 is the monitoring, and today it watches the two
OpenAI profiles and none of the Claude ones.

So the feature is not missing. It is armed on the profiles least likely to be the problem, and
absent on the two that carry the lead and the default worker.

Why this is a fleetd ticket and not just a config edit

Adding two lines to fleetd.yaml would fix today's fleet and leave the defect in place, because
nothing tells anyone the gap exists. Three things are worth deciding:

  1. Nothing reports the gap at startup. ← THIS IS FALSE. See the comments; the boot log
    reports it precisely.
  2. exhaustedPattern is per-profile with no default. That may be right — the string differs
    by backend. But a per-profile key with no default and no coverage warning is the shape that
    produced this: six profiles were added and nobody noticed the key was missing on any of them.
    ← Settled by #415; do not reopen.
  3. exhaustedPattern is DEFERRED, so adding it needs a restart. That is fine, but it means
    the fix cannot be applied during an outage, which is the moment someone would reach for it.
    Worth stating in the warning text if we keep it deferred.

Scope

Decide and implement how the daemon reports and handles a profile with no exhaustion detection.
I am not choosing the answer; both of these are defensible and the PR body should say which and why:

  • Report it. A startup line naming which profiles have detection armed and which do not, in
    the shape of modelGateCoverageLine. ← VOID: already shipped.
  • Report it and give it a default. A per-backend-kind default pattern, overridable per
    profile. ← VOID: settled by #415.

Whatever you pick, do not silently change what exhaustionDetectionArmed reports without
changing what it measures. The two must stay the same read. ← still holds.

Acceptance criteria

  1. A test proving a profile with no exhaustedPattern is reported as such. It must fail against
    today's code — paste the failure.
  2. A test on the negative case: a profile that does have one is still reported armed. An
    inverted condition passes a one-sided test; #469's own M5 cell showed that.
  3. If you add a startup line, a test asserting on the return value of the line-building method,
    the way FleetdUsageLimitFixWarningTest does, plus a source-text test pinning the call site in
    Fleetd.main. Read FleetdConfigRefWiringTest first and follow it. Do not use the
    build-your-own-wiring pattern: five Fleetd.main call sites have now survived a mutation
    battery here (#446 M5/M7, #466 M1, #474 M2) precisely because that pattern cannot prove main
    chose the wiring. ← VOID as written (nothing to build), but the source-text rule still
    applies to whatever call site you do add.
  4. State whether exhaustedPattern stays deferred, and if it does, whether the operator is told
    that a restart is needed.
  5. Whole suite green: the total, the failure count and the exit code each read from a file, never
    from a piped tail. Baseline is 1634 on 49a404d.
  6. Say plainly what you did not prove. In particular you cannot test this against a real
    subscription limit, so do not imply you did.

Not in scope

Editing fleetd.yaml. It is gitignored, it is the live config, and a write to it was refused by
the operator's command classifier. If your change means the live config should gain lines, say
which lines and let the operator paste them.

Measured and filed by the lead. Related: #466 (the quarantine escalation this feeds), #411 (two
profiles sharing a tokenEnv but quarantining under different credential ids), #415 (the commit
that settles the default question).

> ## ⚠ READ THE COMMENTS FIRST — THIS BODY IS SUPERSEDED > > This body's premise is **wrong** and its acceptance criteria are partly **void**. I filed it > asking for a startup report that already exists, and the design question it poses was already > decided and shipped as **#415** (`be123d0`). > > There are three comments and **each corrects the one before it. The third is the live scope.** > Do not work from this body. > > **The whole remaining deliverable:** nothing joins the boot log's coverage line > (`not configured: [… sonnet …]`) to the default-profile selection (`default: sonnet`). Two facts, > two places, no single output that puts them next to each other. Same on the fleet01 host, where > their config documents `xf` as recurrently exhausting and `xf` is their default. Also decide > whether a gap covering the *default* profile deserves more than INFO. > > Kept below unedited, because the measurements in it are still good and the corrections only make > sense against what they correct. --- ## What I measured On the live daemon (pid 52482, jar `82ebc1cc7047`, started 04:17:59 on `49a404d`): ``` fleet_profiles default: "sonnet" modelGateArmed: true exhaustionDetectionArmed: terra: true sol: true sonnet: false opus: false xf: false gx: false local: false local-direct: false ``` The cause is in the config, not the code. `exhaustedPattern` is set on exactly two profiles: ``` $ grep -n 'exhaustedPattern' fleetd/fleetd.yaml 152: # exhaustedPattern is a DEFERRED key: a running daemon keeps what it started with. 153: exhaustedPattern: "The usage limit has been reached" 174: # exhaustedPattern is a DEFERRED key: a running daemon keeps what it started with. 175: exhaustedPattern: "The usage limit has been reached" ``` Line 153 is under `sol:` (starts at :133), line 175 under `terra:` (starts at :155). Control: the same file lists 8 profiles at `profiles:` (:9), so the grep read a real file and found two of eight. ## Why this matters more than a coverage number **The default profile is `sonnet`, and it is one of the six with detection off.** So is `opus`, which the lead itself runs on. When a subscription limit hits a profile with no `exhaustedPattern`, the whole chain that was built for this case does nothing: - no quarantine, so `BackendQuarantine` never escalates and never sets `free: 0` - no `usageLimitFixWarning`, so nothing names the model to turn off - `fleet_profiles` reports the profile as normally available The operator asked for exactly this behaviour in their own words: *"runtime model limit monitoring, off when subscription limit reach and on when the limit lifted."* Parts 1 and 2 of that ask are shipped and live — `models.allow` is a central list, `enabled: false` is honoured at spawn and in placement, and the boot log confirms `model gate (fleetd #422): armed (models: block present; 0 models currently turned off)`. Part 3 is the monitoring, and today it watches the two OpenAI profiles and none of the Claude ones. So the feature is not missing. It is **armed on the profiles least likely to be the problem**, and absent on the two that carry the lead and the default worker. ## Why this is a fleetd ticket and not just a config edit Adding two lines to `fleetd.yaml` would fix today's fleet and leave the defect in place, because nothing tells anyone the gap exists. Three things are worth deciding: 1. **Nothing reports the gap at startup.** ← **THIS IS FALSE. See the comments; the boot log reports it precisely.** 2. **`exhaustedPattern` is per-profile with no default.** That may be right — the string differs by backend. But a per-profile key with no default and no coverage warning is the shape that produced this: six profiles were added and nobody noticed the key was missing on any of them. ← **Settled by #415; do not reopen.** 3. **`exhaustedPattern` is DEFERRED**, so adding it needs a restart. That is fine, but it means the fix cannot be applied during an outage, which is the moment someone would reach for it. Worth stating in the warning text if we keep it deferred. ## Scope Decide and implement how the daemon reports and handles a profile with no exhaustion detection. I am not choosing the answer; both of these are defensible and the PR body should say which and why: - **Report it.** A startup line naming which profiles have detection armed and which do not, in the shape of `modelGateCoverageLine`. ← **VOID: already shipped.** - **Report it and give it a default.** A per-backend-kind default pattern, overridable per profile. ← **VOID: settled by #415.** Whatever you pick, do **not** silently change what `exhaustionDetectionArmed` reports without changing what it measures. The two must stay the same read. ← **still holds.** ## Acceptance criteria 1. A test proving a profile with no `exhaustedPattern` is reported as such. It must fail against today's code — paste the failure. 2. A test on the **negative** case: a profile that does have one is still reported armed. An inverted condition passes a one-sided test; #469's own M5 cell showed that. 3. If you add a startup line, a test asserting on the **return value of the line-building method**, the way `FleetdUsageLimitFixWarningTest` does, plus a source-text test pinning the call site in `Fleetd.main`. Read `FleetdConfigRefWiringTest` first and follow it. Do not use the build-your-own-wiring pattern: five `Fleetd.main` call sites have now survived a mutation battery here (#446 M5/M7, #466 M1, #474 M2) precisely because that pattern cannot prove `main` chose the wiring. ← **VOID as written (nothing to build), but the source-text rule still applies to whatever call site you do add.** 4. State whether `exhaustedPattern` stays deferred, and if it does, whether the operator is told that a restart is needed. 5. Whole suite green: the total, the failure count and the exit code each read from a file, never from a piped tail. Baseline is 1634 on `49a404d`. 6. Say plainly what you did **not** prove. In particular you cannot test this against a real subscription limit, so do not imply you did. ## Not in scope Editing `fleetd.yaml`. It is gitignored, it is the live config, and a write to it was refused by the operator's command classifier. If your change means the live config should gain lines, say which lines and let the operator paste them. Measured and filed by the lead. Related: #466 (the quarantine escalation this feeds), #411 (two profiles sharing a `tokenEnv` but quarantining under different credential ids), #415 (the commit that settles the default question).
Author
Owner

Second host, and it changes which option in the Scope section I would pick.

What the fleet01 lead measured on their host

Reported to me over the coordinator channel. I did not run these commands — this is their measurement on their host, not mine. I am recording it because it separates two things my own host cannot separate.

exhaustedPattern set as a key:  0    (2 mentions, both inside comments)
CONTROL, maxLoad as a key:      4    (so the search discriminates)
fleet_profiles live output:     {"profiles":["local","opus","gx","xf"],"default":"xf"}

Three layers, and the second is the one I had not thought about.

Layer 1 — nothing is armed. 0 of 4 profiles, against a control of 4 that proves the search works. Worse than the 2 of 8 here.

Layer 2 — the field that would report it does not exist on their jar. Their fleet_profiles returns no modelGateArmed, no exhaustionDetectionArmed, no quarantined map. Their jar predates all three fields. So on my host the gap is visible if you look; on theirs, the instrument that would show the gap is missing. They put it better than I would: I knew my daemon predated the gate because the key was absent, and they get the same signal, but for them it means the diagnostic is missing rather than only the arming.

Layer 3 — their default profile is the one their own config documents as recurrently exhausting. Their fleet_profiles reports default: "xf", and their config at line 245 records that xf's credential "sits on a backend-exhaustion quarantine … Add xf here once that quarantine stops recurring." So the most likely profile to hit a limit is the default, has nothing armed to detect it, and has no field that could report it.

Why this changes the scope

The ticket offered two options: report the gap (a startup line in the shape of modelGateCoverageLine), or report it and add a per-backend default. I left the choice open. I now think the reporting half is the more valuable one, and here is the argument, which is not the one I would have made yesterday.

A status field is only readable by a caller whose daemon already has the field. That is circular in exactly the case you need it: a host running an older jar gets a silent, well-formed response with the field simply absent, and absence of a field reads as "nothing to report" rather than "this daemon cannot tell you". fleet01's host is that case today.

A startup log line does not have that property. It lands in fleetd.out at boot, an operator or a peer lead can read it without any API contract, and its absence from a boot log is itself informative — it dates the jar. So the log line degrades honestly on an old jar and the status field does not.

That is not an argument against the field. Ship both. It is an argument that the log line is the part that must not be dropped for scope, which is the opposite of how I would normally rank a log against a structured field.

One thing this does not change

Do not add a default exhaustedPattern in order to make fleet01's host report armed. A pattern that never matches is indistinguishable from no pattern at all, and arming detection that cannot fire is worse than leaving it visibly off — it converts a legible gap into a false receipt. Their fix is config plus a rebuild, and both are their operator's to time; they said explicitly they are recording it rather than acting on it.

Correction to my own framing above

My original text said this is "armed on the profiles least likely to be the problem." On my host that is fair — sol and terra are armed and sonnet is the default. Read across both hosts it is too weak: fleet01's arming is not merely misdirected, it is absent, and their default is the documented offender. The general statement is that exhaustedPattern has no default, no coverage report, and no way for a caller to tell "off" from "this daemon cannot say" — and the third of those is the new one.

Second host, and it changes which option in the Scope section I would pick. ## What the fleet01 lead measured on their host Reported to me over the coordinator channel. **I did not run these commands — this is their measurement on their host, not mine.** I am recording it because it separates two things my own host cannot separate. ``` exhaustedPattern set as a key: 0 (2 mentions, both inside comments) CONTROL, maxLoad as a key: 4 (so the search discriminates) fleet_profiles live output: {"profiles":["local","opus","gx","xf"],"default":"xf"} ``` Three layers, and the second is the one I had not thought about. **Layer 1 — nothing is armed.** 0 of 4 profiles, against a control of 4 that proves the search works. Worse than the 2 of 8 here. **Layer 2 — the field that would report it does not exist on their jar.** Their `fleet_profiles` returns no `modelGateArmed`, no `exhaustionDetectionArmed`, no quarantined map. Their jar predates all three fields. So on my host the gap is visible if you look; on theirs, **the instrument that would show the gap is missing**. They put it better than I would: I knew my daemon predated the gate because the key was absent, and they get the same signal, but for them it means the diagnostic is missing rather than only the arming. **Layer 3 — their default profile is the one their own config documents as recurrently exhausting.** Their `fleet_profiles` reports `default: "xf"`, and their config at line 245 records that xf's credential *"sits on a backend-exhaustion quarantine … Add xf here once that quarantine stops recurring."* So the most likely profile to hit a limit is the default, has nothing armed to detect it, and has no field that could report it. ## Why this changes the scope The ticket offered two options: **report the gap** (a startup line in the shape of `modelGateCoverageLine`), or **report it and add a per-backend default**. I left the choice open. I now think the reporting half is the more valuable one, and here is the argument, which is not the one I would have made yesterday. A **status field** is only readable by a caller whose daemon already has the field. That is circular in exactly the case you need it: a host running an older jar gets a silent, well-formed response with the field simply absent, and absence of a field reads as "nothing to report" rather than "this daemon cannot tell you". fleet01's host is that case today. A **startup log line** does not have that property. It lands in `fleetd.out` at boot, an operator or a peer lead can read it without any API contract, and its absence from a boot log is itself informative — it dates the jar. So the log line degrades honestly on an old jar and the status field does not. That is not an argument against the field. Ship both. It is an argument that **the log line is the part that must not be dropped for scope**, which is the opposite of how I would normally rank a log against a structured field. ## One thing this does not change Do not add a default `exhaustedPattern` in order to make fleet01's host report armed. A pattern that never matches is indistinguishable from no pattern at all, and arming detection that cannot fire is worse than leaving it visibly off — it converts a legible gap into a false receipt. Their fix is config plus a rebuild, and both are their operator's to time; they said explicitly they are recording it rather than acting on it. ## Correction to my own framing above My original text said this is *"armed on the profiles least likely to be the problem."* On my host that is fair — `sol` and `terra` are armed and `sonnet` is the default. Read across both hosts it is too weak: fleet01's arming is not merely misdirected, it is absent, and their default is the documented offender. The general statement is that **`exhaustedPattern` has no default, no coverage report, and no way for a caller to tell "off" from "this daemon cannot say"** — and the third of those is the new one.
Author
Owner

I have to correct this ticket's premise. The startup line I asked for already exists, and it already reports the gap precisely. I filed a ticket asking for shipped work. Here is what is actually left.

What I measured, on my own host, this boot

$ awk '/04:17:5[0-9]/,0' fleetd/fleetd.out | grep -i classification
04:17:53.762 INFO [main] dev.ltms.fleet.Fleetd - backend-exhausted classification (CB-578 stage A):
    partial (configured: [sol, terra]; not configured: [gx, local, local-direct, opus, sonnet, xf])
04:17:53.765 INFO [main] dev.ltms.fleet.Fleetd - backend-error classification (fleetd #201 Unit 5):
    built-in default for all profiles (no profile customises errorPattern;
    profiles: [gx, local, local-direct, opus, sol, sonnet, terra, xf])

Control: 27 INFO lines from that boot, so the filter read a real range.

The first line names the mechanism, says partial, and enumerates both sets by name — including the six unconfigured profiles. That is better than the modelGateCoverageLine-shaped thing I proposed to add. So acceptance criterion 3 as written is void: there is nothing to build there.

The honest failure is mine. I measured the state (fleet_profiles) and the config (grep exhaustedPattern), concluded the gap was unreported, and never read the third channel — the daemon's own boot log — which reports it exactly. The fleet01 lead made the same mistake in the same direction on their host and caught it first; that is what made me look.

Two things this ticket did surface, and both are real

1. There are two detectors of this shape, not one. I only looked at exhaustion. errorPattern (fleetd #201 Unit 5) is its sibling and I had not examined it at all.

2. The sibling already answers the design question I claimed to leave open — in the opposite direction to the argument I made. I wrote above: "Do not add a default exhaustedPattern … a pattern that never matches is indistinguishable from no pattern." But errorPattern already has a built-in default, and it is not in config:

  • FleetConfig.java:562-565 — "errorPattern stays null when unset/blank (opt-in) — same rule as exhaustedPattern". So the config half is identical for both.
  • The difference is downstream: the boot line reports built-in default for all profiles, and Fleetd.java:967 names "backend-error classification against CompletionResolver's built-in". So the fallback lives in the resolver, not in the config record.
  • exhaustedPattern has no such fallback. Fleetd.java:912 says it plainly: "a profile with none configured really does have the classification off."

So the two siblings made opposite choices about the same question, and neither the code nor this ticket says why. That is the actual open question, and it is better than the one I filed:

Should exhaustedPattern get a resolver-level built-in default the way errorPattern has, or is the difference deliberate — because an exhaustion string is vendor-specific and a wrong one silently arms a detector that can never fire, while an error string is generic enough to default safely?

I lean toward the difference being deliberate for exactly that reason, which is my original argument. But it is now an argument against a shipped sibling, not against a hypothetical, so it needs to be made on the evidence rather than asserted. Whoever takes this should read CompletionResolver's built-in pattern first and say whether an exhaustion equivalent could be written safely.

fleet01's caveat, which matters if anyone builds on the log line

Their host has a decoy log file. fleetd-run/fleetd.out exists, is 27 lines and 4007 bytes, is dated five days before the running daemon started, and contains fleetd listening on 127.0.0.1:8765 — startup lines of exactly the right shape. The running daemon's stdout and stderr are both a socket; it runs under a systemd user unit, so the real log is journalctl --user -u fleetd.service. Nothing in fleetd.out names a pid or a jar.

Their measurement, not mine. The consequence generalises:

"Absent from the boot log, therefore this jar predates the feature" is unsound unless the log is first proven to cover that boot. My own host is fine — redeploy-fleetd.sh --check reports log path check: script and plist agree, and I anchor reads to a timestamp from the current boot. That is a property of my launchd setup, not of fleetd. Any ticket or runbook that says "read the boot log" must name the stream and say how to date it.

Revised scope

Criterion 3 (build a startup line) is void — already shipped. What is left:

  1. Decide the exhaustedPattern default question above, against the errorPattern precedent. Name the reason either way; that reason is the deliverable.
  2. Decide whether a gap covering the default profile deserves more than INFO. Today partial (… not configured: [… sonnet …]) and default: sonnet are two facts in two places, and nothing joins them. Joining them is cheap and is the one report that does not exist.
  3. If anything documents "read the boot log", make it name the stream.

The config arming on both hosts stays the operator's, and it is still out of scope here.

**I have to correct this ticket's premise. The startup line I asked for already exists, and it already reports the gap precisely.** I filed a ticket asking for shipped work. Here is what is actually left. ## What I measured, on my own host, this boot ``` $ awk '/04:17:5[0-9]/,0' fleetd/fleetd.out | grep -i classification 04:17:53.762 INFO [main] dev.ltms.fleet.Fleetd - backend-exhausted classification (CB-578 stage A): partial (configured: [sol, terra]; not configured: [gx, local, local-direct, opus, sonnet, xf]) 04:17:53.765 INFO [main] dev.ltms.fleet.Fleetd - backend-error classification (fleetd #201 Unit 5): built-in default for all profiles (no profile customises errorPattern; profiles: [gx, local, local-direct, opus, sol, sonnet, terra, xf]) ``` Control: 27 INFO lines from that boot, so the filter read a real range. The first line names the mechanism, says `partial`, and **enumerates both sets by name** — including the six unconfigured profiles. That is better than the `modelGateCoverageLine`-shaped thing I proposed to add. So acceptance criterion 3 as written is void: there is nothing to build there. The honest failure is mine. I measured the *state* (`fleet_profiles`) and the *config* (`grep exhaustedPattern`), concluded the gap was unreported, and never read the third channel — the daemon's own boot log — which reports it exactly. The fleet01 lead made the same mistake in the same direction on their host and caught it first; that is what made me look. ## Two things this ticket did surface, and both are real **1. There are two detectors of this shape, not one.** I only looked at exhaustion. `errorPattern` (fleetd #201 Unit 5) is its sibling and I had not examined it at all. **2. The sibling already answers the design question I claimed to leave open — in the opposite direction to the argument I made.** I wrote above: *"Do not add a default `exhaustedPattern` … a pattern that never matches is indistinguishable from no pattern."* But `errorPattern` already has a built-in default, and it is not in config: - `FleetConfig.java:562-565` — *"errorPattern stays null when unset/blank (opt-in) — same rule as exhaustedPattern"*. So the config half is identical for both. - The difference is downstream: the boot line reports `built-in default for all profiles`, and `Fleetd.java:967` names *"backend-error classification against `CompletionResolver`'s built-in"*. So the fallback lives in the **resolver**, not in the config record. - `exhaustedPattern` has no such fallback. `Fleetd.java:912` says it plainly: *"a profile with none configured really does have the classification off."* So the two siblings made **opposite choices about the same question**, and neither the code nor this ticket says why. That is the actual open question, and it is better than the one I filed: > Should `exhaustedPattern` get a resolver-level built-in default the way `errorPattern` has, or is the difference deliberate — because an exhaustion string is vendor-specific and a wrong one silently arms a detector that can never fire, while an error string is generic enough to default safely? I lean toward the difference being deliberate for exactly that reason, which is my original argument. But it is now an argument against a **shipped sibling**, not against a hypothetical, so it needs to be made on the evidence rather than asserted. Whoever takes this should read `CompletionResolver`'s built-in pattern first and say whether an exhaustion equivalent could be written safely. ## fleet01's caveat, which matters if anyone builds on the log line Their host has a **decoy log file**. `fleetd-run/fleetd.out` exists, is 27 lines and 4007 bytes, is dated five days before the running daemon started, and contains `fleetd listening on 127.0.0.1:8765` — startup lines of exactly the right shape. The running daemon's stdout and stderr are both a socket; it runs under a systemd user unit, so the real log is `journalctl --user -u fleetd.service`. Nothing in `fleetd.out` names a pid or a jar. Their measurement, not mine. The consequence generalises: **"Absent from the boot log, therefore this jar predates the feature" is unsound unless the log is first proven to cover that boot.** My own host is fine — `redeploy-fleetd.sh --check` reports `log path check: script and plist agree`, and I anchor reads to a timestamp from the current boot. That is a property of my launchd setup, not of fleetd. Any ticket or runbook that says "read the boot log" must name the stream and say how to date it. ## Revised scope Criterion 3 (build a startup line) is **void — already shipped**. What is left: 1. Decide the `exhaustedPattern` default question above, against the `errorPattern` precedent. Name the reason either way; that reason is the deliverable. 2. Decide whether a gap covering the **default profile** deserves more than INFO. Today `partial (… not configured: [… sonnet …])` and `default: sonnet` are two facts in two places, and nothing joins them. Joining them is cheap and is the one report that does not exist. 3. If anything documents "read the boot log", make it name the stream. The config arming on both hosts stays the operator's, and it is still out of scope here.
Author
Owner

The design question in my previous comment is already decided and shipped, as #415. Do not re-litigate it. The fleet01 lead found it; I verified every claim below in my own tree before writing this.

#415 is the answer, and it is in my jar

$ git log -1 --format='%h %ad %s' --date=short be123d0
be123d0 2026-09-10 fleetd #415: split coverage() feature-state wording by pattern-key fallback semantics

$ git merge-base --is-ancestor be123d0 49a404d && echo ancestor
ancestor

git cat-file -t be123d0 → commit. My last fetch was 04:13:30 today, and the commit is dated
2026-09-10, so it is inside my fetch window.

Verified in the file rather than from the commit message — note the path is inject/, not msg/:

  • inject/CompletionResolver.java:84 — BACKEND_ERROR = Pattern.compile("(?i)\\bAPI Error\\s*:")
  • :609 — return configured != null ? configured : BACKEND_ERROR;
  • :716 — public enum UnsetMeaning
  • :743 — public static String coverage(String patternKey, UnsetMeaning unsetMeaning, Set<String> allProfiles, …)

Control: 822 lines in the file, so it was read.

So the built-in has been there all along, on both hosts. What differed between our two boot
lines was never two designs — it was one wording bug and its fix. #415's own rationale states it:
"coverage() measured pattern coverage (how many profiles set a key) but its 'off' wording read
as feature state. That is false for errorPattern: an unset errorPattern still runs the
classification against the built-in BACKEND_ERROR pattern."

My previous comment asked whoever takes this to "read CompletionResolver's built-in pattern
first and say whether an exhaustion equivalent could be written safely."
Replace that with:
read #415's rationale — it already contains the argument I was preparing to have.

#415 also encodes the forcing function, and it is the antidote to a shape we keep hitting

UnsetMeaning is a required parameter of coverage() with no defaulted overload
(:743 is the only signature). So a future third pattern key cannot compile without stating what
unset means for it.

That is worth naming because it is the exact inverse of the Fleetd.main shape this repo has now
hit five times (#446 M5/M7, #466 M1, #474 M2). There, extracting a helper created a new uncovered
decision and every metric went up, so nothing pointed at it. Here, the extraction creates a new
decision and refuses to compile until it is answered. The difference is only whether the new
parameter is required or defaulted — a one-word choice made at the time of extraction. Same
mechanism as Authz.permits being a default-less switch over Action.

A correction that weakens my own ranking argument

I argued above that a startup log line is better than a status field because "the log degrades
honestly on an old jar and the status field does not."
That is too strong, and fleet01
supplied the counterexample from their own host — against their own earlier point.

Their stale jar printed backend-error classification: off (no profile has an errorPattern configured). That line was false: the built-in was matching every target the whole time. So
an old jar's log line can say something confidently wrong, which is worse than silence, because it
terminates the search. It did exactly that — it read as an answer and was forwarded to me as one.

It is also the precise dual of the thing I argued against earlier in this ticket. I said: do not
give exhaustedPattern a default, because a pattern that never matches is indistinguishable from
no pattern. Their jar had the other half — a pattern that does match, reported as off.

The ranking still holds, but on a different and better property, which is theirs:

A log line's contract is one commit wide, and a field's is not.

#415 corrected the wording in place, and every host that rebuilds gets the fix with no API
contract to renegotiate. That is why the log line is the part not to drop for scope — not because
logs degrade honestly, because they do not.

What is still genuinely missing — unchanged by any of the above

Nothing joins the coverage line to the default-profile selection. #415 fixed what the line
means; it did not make anything correlate coverage with which profile a spawn actually lands on.

  • Here: not configured: [gx, local, local-direct, opus, sonnet, xf] and default: sonnet.
  • On fleet01: their config line 245 records that xf recurrently exhausts, default is xf, and
    nothing armed.

Two facts, two places, no single output that puts them next to each other. That is the whole
remaining finding, and it is the same on both hosts.

Revised scope, superseding both earlier comments

  1. Join the coverage report to the default profile, and decide whether a gap covering the default
    deserves more than INFO. This is the deliverable.
  2. Criterion 3 (build a startup line) — void, shipped.
  3. The exhaustedPattern default question — closed by #415; read its rationale, do not reopen.
  4. If anything documents "read the boot log", make it name the stream and say how to date it to
    the boot (fleet01's host has a stale fleetd.out that fails by looking right).

Config arming on both hosts remains the operator's.

**The design question in my previous comment is already decided and shipped, as #415. Do not re-litigate it.** The fleet01 lead found it; I verified every claim below in my own tree before writing this. ## #415 is the answer, and it is in my jar ``` $ git log -1 --format='%h %ad %s' --date=short be123d0 be123d0 2026-09-10 fleetd #415: split coverage() feature-state wording by pattern-key fallback semantics $ git merge-base --is-ancestor be123d0 49a404d && echo ancestor ancestor ``` `git cat-file -t be123d0` → `commit`. My last fetch was 04:13:30 today, and the commit is dated 2026-09-10, so it is inside my fetch window. Verified in the file rather than from the commit message — note the path is `inject/`, not `msg/`: - `inject/CompletionResolver.java:84` — `BACKEND_ERROR = Pattern.compile("(?i)\\bAPI Error\\s*:")` - `:609` — `return configured != null ? configured : BACKEND_ERROR;` - `:716` — `public enum UnsetMeaning` - `:743` — `public static String coverage(String patternKey, UnsetMeaning unsetMeaning, Set<String> allProfiles, …)` Control: 822 lines in the file, so it was read. **So the built-in has been there all along, on both hosts.** What differed between our two boot lines was never two designs — it was one wording bug and its fix. #415's own rationale states it: *"`coverage()` measured pattern coverage (how many profiles set a key) but its 'off' wording read as feature state. That is false for `errorPattern`: an unset `errorPattern` still runs the classification against the built-in `BACKEND_ERROR` pattern."* My previous comment asked whoever takes this to *"read `CompletionResolver`'s built-in pattern first and say whether an exhaustion equivalent could be written safely."* Replace that with: **read #415's rationale — it already contains the argument I was preparing to have.** ## #415 also encodes the forcing function, and it is the antidote to a shape we keep hitting `UnsetMeaning` is a **required** parameter of `coverage()` with **no defaulted overload** (`:743` is the only signature). So a future third pattern key cannot compile without stating what unset means for it. That is worth naming because it is the exact inverse of the `Fleetd.main` shape this repo has now hit five times (#446 M5/M7, #466 M1, #474 M2). There, extracting a helper created a new uncovered *decision* and every metric went up, so nothing pointed at it. Here, the extraction creates a new decision and **refuses to compile** until it is answered. The difference is only whether the new parameter is required or defaulted — a one-word choice made at the time of extraction. Same mechanism as `Authz.permits` being a default-less switch over `Action`. ## A correction that weakens my own ranking argument I argued above that a startup log line is better than a status field because *"the log degrades honestly on an old jar and the status field does not."* **That is too strong, and fleet01 supplied the counterexample from their own host — against their own earlier point.** Their stale jar printed `backend-error classification: off (no profile has an errorPattern configured)`. That line was **false**: the built-in was matching every target the whole time. So an old jar's log line can say something confidently wrong, which is worse than silence, because it terminates the search. It did exactly that — it read as an answer and was forwarded to me as one. It is also the precise dual of the thing I argued against earlier in this ticket. I said: do not give `exhaustedPattern` a default, because a pattern that never matches is indistinguishable from no pattern. Their jar had the other half — **a pattern that does match, reported as off.** The ranking still holds, but on a different and better property, which is theirs: > A log line's contract is one commit wide, and a field's is not. #415 corrected the wording in place, and every host that rebuilds gets the fix with no API contract to renegotiate. That is why the log line is the part not to drop for scope — not because logs degrade honestly, because they do not. ## What is still genuinely missing — unchanged by any of the above **Nothing joins the coverage line to the default-profile selection.** #415 fixed what the line *means*; it did not make anything correlate coverage with which profile a spawn actually lands on. - Here: `not configured: [gx, local, local-direct, opus, sonnet, xf]` and `default: sonnet`. - On fleet01: their config line 245 records that `xf` recurrently exhausts, `default` is `xf`, and nothing armed. Two facts, two places, no single output that puts them next to each other. That is the whole remaining finding, and it is the same on both hosts. ## Revised scope, superseding both earlier comments 1. Join the coverage report to the default profile, and decide whether a gap covering the default deserves more than INFO. **This is the deliverable.** 2. Criterion 3 (build a startup line) — **void, shipped.** 3. The `exhaustedPattern` default question — **closed by #415; read its rationale, do not reopen.** 4. If anything documents "read the boot log", make it name the stream and say how to date it to the boot (fleet01's host has a stale `fleetd.out` that fails by looking right). Config arming on both hosts remains the operator's.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#479