fleetd #466 scope item 2: report the quarantine repeat count, not only the seconds #473

Closed
agent wants to merge 0 commits from worker/466-quarantine-repeatcount-report into main
Member

Third and final unit of fleetd #466 (units 1-2 merged in 789b6a8 / PR #470, now closed).

Adds BackendQuarantine#status(credentialId) -> Optional<Status>: one quarantines.get(credentialId) read that produces both remainingSeconds and repeatCount together, the same one-accessor pattern documented on CompositePeerLauncher.modelGateState() -- so the reported count can never disagree with the cooldown it describes.

FleetMcp.profilesView (feeds fleet_profiles and, via FleetApp, GET /profiles) and FleetMcp.capacityView (feeds fleet_list's per-profile capacity rows, CB-583) now call status() instead of remainingSeconds() and add a quarantineAttempt field beside quarantinedForSeconds: 1 for a first occurrence, 2 for the second in a row, and so on.

none() reports no status for anything (it quarantines nothing). The flat two-argument constructor still reports a real, growing quarantineAttempt even though its cooldown stays flat.

No change to escalation, ceiling, or reset logic -- reporting only, per the ticket's explicit scope.

Tests added (BackendQuarantineTest, FleetMcpTest): first-occurrence value, growing count alongside the escalating cooldown, count still growing past the cooldown ceiling, a quiet-gap reset back to attempt 1, per-credential isolation, none()'s empty status, the flat instance's growing count, and fleet_profiles/fleet_list JSON assertions for the new field (present and absent).

Mutation proof (the acceptance criterion's strong form): decoupled status()'s repeatCount from the one QuarantineState entry quarantine() writes, reading instead from a second, independently-maintained, non-resetting counter -- the exact shape the ticket warned against. aQuietGapResetsTheReportedAttemptCountToOneToo caught it: expected: <1> but was: <4>, because the real streak resets on a quiet gap but the independent counter does not. Four more targeted mutations (first-occurrence off-by-one, hardcoded repeatCount=1, a fabricated non-empty status on none()/never-quarantined, and skipped tracking on the flat instance) were each caught by name and restored byte-identical afterward.

mvn -B clean test: Tests run: 1620, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS, exit 0.

Third and final unit of fleetd #466 (units 1-2 merged in 789b6a8 / PR #470, now closed). Adds `BackendQuarantine#status(credentialId) -> Optional<Status>`: one `quarantines.get(credentialId)` read that produces both `remainingSeconds` and `repeatCount` together, the same one-accessor pattern documented on `CompositePeerLauncher.modelGateState()` -- so the reported count can never disagree with the cooldown it describes. `FleetMcp.profilesView` (feeds `fleet_profiles` and, via `FleetApp`, `GET /profiles`) and `FleetMcp.capacityView` (feeds `fleet_list`'s per-profile capacity rows, CB-583) now call `status()` instead of `remainingSeconds()` and add a `quarantineAttempt` field beside `quarantinedForSeconds`: 1 for a first occurrence, 2 for the second in a row, and so on. `none()` reports no status for anything (it quarantines nothing). The flat two-argument constructor still reports a real, growing `quarantineAttempt` even though its cooldown stays flat. No change to escalation, ceiling, or reset logic -- reporting only, per the ticket's explicit scope. **Tests added** (`BackendQuarantineTest`, `FleetMcpTest`): first-occurrence value, growing count alongside the escalating cooldown, count still growing past the cooldown ceiling, a quiet-gap reset back to attempt 1, per-credential isolation, `none()`'s empty status, the flat instance's growing count, and `fleet_profiles`/`fleet_list` JSON assertions for the new field (present and absent). **Mutation proof (the acceptance criterion's strong form):** decoupled `status()`'s repeatCount from the one `QuarantineState` entry `quarantine()` writes, reading instead from a second, independently-maintained, non-resetting counter -- the exact shape the ticket warned against. `aQuietGapResetsTheReportedAttemptCountToOneToo` caught it: `expected: <1> but was: <4>`, because the real streak resets on a quiet gap but the independent counter does not. Four more targeted mutations (first-occurrence off-by-one, hardcoded repeatCount=1, a fabricated non-empty status on `none()`/never-quarantined, and skipped tracking on the flat instance) were each caught by name and restored byte-identical afterward. `mvn -B clean test`: Tests run: 1620, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS, exit 0.
agent added 1 commit 2026-09-10 15:09:33 +02:00
fleetd #466 scope item 2: report the quarantine repeat count, not only the seconds
CI / build (pull_request) Successful in 1m31s
CI / contract (pull_request) Successful in 1m33s
e95ed99bf7
Add BackendQuarantine#status(credentialId) -> Optional<Status>, a single
QuarantineState read that answers both remainingSeconds and repeatCount
together -- the same "one accessor" pattern CompositePeerLauncher.
modelGateState() already uses, so the two facts can never disagree.

fleet_profiles/REST GET /profiles and fleet_list's capacity rows
(FleetMcp.profilesView/capacityView) now call status() instead of
remainingSeconds() and add a "quarantineAttempt" field beside
"quarantinedForSeconds": 1 for a first occurrence, 2 for the second
in a row, and so on.

No change to the escalation, ceiling, or reset logic itself -- this
unit is reporting only.
Owner

Merged by hand as 25ba7f1 (--no-ff, tree f22c392) and pushed to main. A local --no-ff merge does not close a Gitea PR, so I am closing this one by hand.

I ran my own five-cell mutation battery on the merge commit's own tree, not on the branch. Every cell was a full mvn -B clean test.

  • CONTROL 1, unmutated: 1629 tests, 0 failures, 0 errors, BUILD SUCCESS.
  • M1, the ticket's own warning built in full — the report reads its own non-resetting counter: KILLED by BackendQuarantineTest.aQuietGapResetsTheReportedAttemptCountToOneToo, the exact test you named.
  • M2, hardcode the reported attempt to 1: KILLED, 6 tests.
  • M3, drop the field from fleet_profiles only: KILLED by FleetMcpTest.profilesReportsQuarantineAttemptBesideRemainingSeconds.
  • M4, drop it from fleet_list only: KILLED by FleetMcpTest.capacityRowReportsQuarantineAttemptBesideRemainingSeconds.
  • M5, off by one on a first occurrence: KILLED, 8 tests.

Two notes for the record.

M3 and M4 are the pair that usually fails here. Two call sites are two surfaces, and a test on one proves nothing about the other. Three recent tickets had to go back to a worker for exactly that. Yours did not: each call site fails its own named test when only its own field is removed.

Your fleet_reply never arrived. fleet_poll on your inbox came back empty and fleet_status said done, which on its own looks like a member that produced nothing. The report survived because you pushed first and put it in this PR body, which the brief asked for. Also worth knowing: you committed to worker/466-quarantine-repeatcount-report, not the branch spawn provisioned. I check the worktree before believing silence, so it cost nothing — but a lead that trusts fleet_list's branch name would have merged nothing and been told "Already up to date".

Full detail is on #466. Items 1 and 2 of that ticket are done; item 3 stays open and is mine to decide.

Merged by hand as `25ba7f1` (`--no-ff`, tree `f22c392`) and pushed to `main`. A local `--no-ff` merge does not close a Gitea PR, so I am closing this one by hand. I ran my own five-cell mutation battery on the merge commit's own tree, not on the branch. Every cell was a full `mvn -B clean test`. - CONTROL 1, unmutated: **1629 tests, 0 failures, 0 errors, BUILD SUCCESS**. - M1, the ticket's own warning built in full — the report reads its own non-resetting counter: **KILLED** by `BackendQuarantineTest.aQuietGapResetsTheReportedAttemptCountToOneToo`, the exact test you named. - M2, hardcode the reported attempt to 1: **KILLED**, 6 tests. - M3, drop the field from `fleet_profiles` only: **KILLED** by `FleetMcpTest.profilesReportsQuarantineAttemptBesideRemainingSeconds`. - M4, drop it from `fleet_list` only: **KILLED** by `FleetMcpTest.capacityRowReportsQuarantineAttemptBesideRemainingSeconds`. - M5, off by one on a first occurrence: **KILLED**, 8 tests. Two notes for the record. **M3 and M4 are the pair that usually fails here.** Two call sites are two surfaces, and a test on one proves nothing about the other. Three recent tickets had to go back to a worker for exactly that. Yours did not: each call site fails its own named test when only its own field is removed. **Your `fleet_reply` never arrived.** `fleet_poll` on your inbox came back empty and `fleet_status` said `done`, which on its own looks like a member that produced nothing. The report survived because you pushed first and put it in this PR body, which the brief asked for. Also worth knowing: you committed to `worker/466-quarantine-repeatcount-report`, not the branch spawn provisioned. I check the worktree before believing silence, so it cost nothing — but a lead that trusts `fleet_list`'s branch name would have merged nothing and been told "Already up to date". Full detail is on #466. Items 1 and 2 of that ticket are done; item 3 stays open and is mine to decide.
ltms closed this pull request 2026-09-10 15:31:13 +02:00
Some checks are pending
CI / build (pull_request) Successful in 1m31s
CI / contract (pull_request) Successful in 1m33s

Pull request closed

Sign in to join this conversation.