notifications.mode: webhook reports healthCoverage "full" and sends nothing — the knob relabels the absence of a sink #392

Open
opened 2026-09-10 02:51:38 +02:00 by ltms · 0 comments
Owner

The defect

Setting health.notifications.mode: webhook changes what fleet_list reports and changes no behaviour at all.

// FleetConfig.java:846
public record Notifications(String mode) {
    public boolean configured() { return "webhook".equalsIgnoreCase(mode); }
}
// FleetHealthMonitor.java:378
return !enabled ? "off" : notificationConfigured ? "full" : "detection-only";

So a lead reading fleet_list sees healthCoverage: "full" on the strength of one string comparison.

There is no sink to report

webhook occurs once in the whole of src/main/java, and it is the comparison above:

$ grep -rn "webhook" fleetd/src/main/java/
fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:846

No HTTP client, no delivery path, no retry, no test. The mode name promises a mechanism that was never built.

Why this is worth a ticket rather than a doc note

fleetd.example.yaml:101 is honest about it — it says the knob flips what fleet_list reports and does not make fleetd send anything. But the honest warning lives in the example file, while the misleading value lives in the running config, and fleet_list is what a lead actually reads.

The failure mode: an operator sets webhook intending to turn notifications on, healthCoverage goes full, and health events keep going nowhere. The label now says covered. Nothing is covered. That is strictly worse than disabled, which is at least true.

Measured consequence on the fleet01 host, which correctly left it disabled: over 24 h, STALL_SUSPECTED x3, TURN_BOUNDARY_LOST x2, REPLY_STRANDED x2, recovered x1 — every one detected, every one delivered to nobody. That host's lead is building a sink out of journalctl instead, because the config knob does not provide one.

Fix — pick one, do not leave it as is

  1. Refuse the value. mode: webhook fails at config load with "no webhook delivery is implemented; use disabled". Smallest change, and it can never lie.
  2. Report what is true. configured() returns true only when a sink is actually wired. mode alone never moves healthCoverage off detection-only.
  3. Build the sink. Larger, and it should be its own ticket rather than smuggled in behind a string compare.

I recommend 1 now and 3 separately. Option 2 alone leaves a config value that parses, does nothing, and reads as if it should.

Acceptance

  • No configuration can make healthCoverage report full unless health events actually reach something outside the JVM.
  • A test asserts that. Today no test covers the mapping from mode to healthCoverage at all.
  • fleetd.example.yaml's warning is either removed (because the value is refused) or still accurate.

Found by the fleet01 lead while acting on an audit recommendation of mine. I had recommended turning notifications on; they checked the code first and found the knob does not do that. Verification and filing: mac lead.

## The defect Setting `health.notifications.mode: webhook` changes what `fleet_list` **reports** and changes no behaviour at all. ```java // FleetConfig.java:846 public record Notifications(String mode) { public boolean configured() { return "webhook".equalsIgnoreCase(mode); } } ``` ```java // FleetHealthMonitor.java:378 return !enabled ? "off" : notificationConfigured ? "full" : "detection-only"; ``` So a lead reading `fleet_list` sees `healthCoverage: "full"` on the strength of one string comparison. ## There is no sink to report `webhook` occurs **once** in the whole of `src/main/java`, and it is the comparison above: ``` $ grep -rn "webhook" fleetd/src/main/java/ fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:846 ``` No HTTP client, no delivery path, no retry, no test. The mode name promises a mechanism that was never built. ## Why this is worth a ticket rather than a doc note `fleetd.example.yaml:101` is honest about it — it says the knob flips what `fleet_list` reports and does not make fleetd send anything. But the honest warning lives in the example file, while the misleading value lives in the running config, and `fleet_list` is what a lead actually reads. The failure mode: an operator sets `webhook` intending to turn notifications *on*, `healthCoverage` goes `full`, and health events keep going nowhere. The label now says covered. Nothing is covered. That is strictly worse than `disabled`, which is at least true. Measured consequence on the fleet01 host, which correctly left it `disabled`: over 24 h, `STALL_SUSPECTED` x3, `TURN_BOUNDARY_LOST` x2, `REPLY_STRANDED` x2, recovered x1 — every one detected, every one delivered to nobody. That host's lead is building a sink out of `journalctl` instead, because the config knob does not provide one. ## Fix — pick one, do not leave it as is 1. **Refuse the value.** `mode: webhook` fails at config load with "no webhook delivery is implemented; use `disabled`". Smallest change, and it can never lie. 2. **Report what is true.** `configured()` returns true only when a sink is actually wired. `mode` alone never moves `healthCoverage` off `detection-only`. 3. **Build the sink.** Larger, and it should be its own ticket rather than smuggled in behind a string compare. I recommend 1 now and 3 separately. Option 2 alone leaves a config value that parses, does nothing, and reads as if it should. ## Acceptance - No configuration can make `healthCoverage` report `full` unless health events actually reach something outside the JVM. - A test asserts that. Today no test covers the mapping from `mode` to `healthCoverage` at all. - `fleetd.example.yaml`'s warning is either removed (because the value is refused) or still accurate. Found by the fleet01 lead while acting on an audit recommendation of mine. I had recommended turning notifications on; they checked the code first and found the knob does not do that. Verification and filing: mac lead.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#392