An architect-settled decision leaves no trace the operator can find: there is no event for it, and #392 shows there is no sink to send one to #592

Open
opened 2026-09-19 09:52:29 +02:00 by ltms · 0 comments
Owner

Why this exists now

The operator changed the escalation rule on 2026-09-19:

in case of some decision need to be made -> it should consults with one or more architect
member and they are authorized to agreed on one decision to unblock and moving on, don't need
my input - maybe we need a channel to notify for such events

The first clause removes the operator from the loop. The second clause is the thing that makes
that safe. Without it, the fleet now makes binding decisions the operator never sees. That is
a bigger change than it looks: the operator's old notification channel was the block itself —
work stopped, and they found out. Remove the block and you remove the signal with it.

Measured state — 2026-09-19

There is no sink. health.notifications is one string:

// FleetConfig.java:900
public record Notifications(String mode) {

Its only consumer is a label. Fleetd.java:592-598:

String coverage = FleetHealthMonitor.coverage(true,
        cfg.health().notifications() != null && cfg.health().notifications().configured());
if ("detection-only".equals(coverage)) {
    log.warn("fleet health: {} (no notification sink configured)", coverage);
}

and the same expression again at Fleetd.java:1040 to feed fleet_list. Live fleet_list on
the Mac returns "healthCoverage":"detection-only" right now, and the live config has
health.notifications.mode: disabled.

#392 already owns that half: flipping the mode to webhook reports healthCoverage: "full"
and sends nothing — the knob relabels the absence of a sink. That issue should be fixed first or
alongside; this one must not grow a second, private sink beside it.

There is also no event. Nothing in the daemon knows that a decision was made. The architect
exchange is ordinary fleet_send traffic between two members, and the outcome arrives as a
fleet_reply body — free text that no code inspects. So even with #392 fixed, there would be
nothing to publish.

The recommendation: the record is a pull channel; the push is only a pointer

Make the ticket the place the decision lives, and let any webhook be a notification about it.

The reason is already written into the project's own rules, and it was measured the hard way.
CLAUDE.md says, of corrections to a member:

A push delivery needs the recipient free at send time; a pull channel needs only that they look
before acting. So put every correction on the ticket … all corrections go to the ticket.

An operator is the least available recipient in the fleet — they are asleep, or away, precisely
when autonomous decisions are most valuable. A push-only design repeats the failure this session
watched on fleet01: a message was delivered into a pane, the receipt was true, and nobody read it
for five days. A ticket comment has none of that dependence on timing. It is also durable,
timestamped, searchable, and already where the work is.

So: a webhook that says "decision D was settled, see ticket T" is useful. A webhook that is
the only record is not.

What the event must carry

A notification that cannot be audited later is theatre. Each record needs:

  • the question that blocked the lead, in one sentence;
  • which architects were consulted, with their profile and session id;
  • what they agreed, and whether it was agreement or the lead deciding after two rounds failed
    (the existing architect charter caps the exchange at two rounds);
  • the evidence each position rested on, with the unchecked claims marked as unchecked;
  • the ticket the decision applies to;
  • an explicit statement that the operator was not consulted, and the timestamp.

That last line matters most. It is what lets the operator scan a week of decisions and find the
one they would have made differently.

Acceptance criteria

Properties under a change, not the presence of a construct.

  1. Remove the emit and something goes red. Delete the call that records a settled decision and
    a test must fail. A test asserting only that the record type exists, or that a well-formed record
    serialises, stays green when nothing ever emits one — that is the #506 family and does not pass.
  2. A decision with no operator in it is findable without knowing it happened. Given a fleet
    that settled a decision while the operator was away, the operator must be able to list every such
    decision from one place, without being told a ticket number first. "It is in the logs" is not a
    pass; the daemon log was unreadable on fleet01 for 40 hours and nobody noticed.
  3. A failed notification is visible as a failure. If the sink is down or refuses, the surface
    must say so. Reporting full while sending nothing is exactly the #392 defect and must not be
    reproduced here.
  4. The record survives the session. Kill the lead session after it settles a decision, start a
    fresh one, and the decision must still be readable. A record that lives only in a pane or a
    conversation does not pass.

Scope note

This issue owns the event and the record. #392 owns the sink. Fixing this one by adding a
second notification path beside #392 would leave two knobs that both claim to notify, which is the
shape #497 catalogues. Coordinate the two.

Related: #591 (the policy that created this gap has no delivery surface for a lead), #383 (a
detected health state with no API surface — the same "detected but unreportable" shape).

## Why this exists now The operator changed the escalation rule on 2026-09-19: > in case of some decision need to be made -> it should consults with one or more architect > member and they are authorized to agreed on one decision to unblock and moving on, don't need > my input - maybe we need a channel to notify for such events The first clause removes the operator from the loop. The second clause is the thing that makes that safe. **Without it, the fleet now makes binding decisions the operator never sees.** That is a bigger change than it looks: the operator's old notification channel was the block itself — work stopped, and they found out. Remove the block and you remove the signal with it. ## Measured state — 2026-09-19 **There is no sink.** `health.notifications` is one string: ```java // FleetConfig.java:900 public record Notifications(String mode) { ``` Its only consumer is a label. `Fleetd.java:592-598`: ```java String coverage = FleetHealthMonitor.coverage(true, cfg.health().notifications() != null && cfg.health().notifications().configured()); if ("detection-only".equals(coverage)) { log.warn("fleet health: {} (no notification sink configured)", coverage); } ``` and the same expression again at `Fleetd.java:1040` to feed `fleet_list`. Live `fleet_list` on the Mac returns `"healthCoverage":"detection-only"` right now, and the live config has `health.notifications.mode: disabled`. **#392 already owns that half**: flipping the mode to `webhook` reports `healthCoverage: "full"` and sends nothing — the knob relabels the absence of a sink. That issue should be fixed first or alongside; this one must not grow a second, private sink beside it. **There is also no event.** Nothing in the daemon knows that a decision was made. The architect exchange is ordinary `fleet_send` traffic between two members, and the outcome arrives as a `fleet_reply` body — free text that no code inspects. So even with #392 fixed, there would be nothing to publish. ## The recommendation: the record is a pull channel; the push is only a pointer Make the ticket the place the decision lives, and let any webhook be a notification *about* it. The reason is already written into the project's own rules, and it was measured the hard way. `CLAUDE.md` says, of corrections to a member: > A push delivery needs the recipient free at send time; a pull channel needs only that they look > before acting. So put every correction on the **ticket** … **all corrections go to the ticket.** An operator is the least available recipient in the fleet — they are asleep, or away, precisely when autonomous decisions are most valuable. A push-only design repeats the failure this session watched on fleet01: a message was delivered into a pane, the receipt was true, and nobody read it for five days. A ticket comment has none of that dependence on timing. It is also durable, timestamped, searchable, and already where the work is. So: **a webhook that says "decision D was settled, see ticket T" is useful. A webhook that *is* the only record is not.** ## What the event must carry A notification that cannot be audited later is theatre. Each record needs: - the question that blocked the lead, in one sentence; - which architects were consulted, with their profile and session id; - what they agreed, and whether it was agreement or the lead deciding after two rounds failed (the existing `architect` charter caps the exchange at two rounds); - the evidence each position rested on, with the unchecked claims marked as unchecked; - the ticket the decision applies to; - an explicit statement that the operator was not consulted, and the timestamp. That last line matters most. It is what lets the operator scan a week of decisions and find the one they would have made differently. ## Acceptance criteria Properties under a change, not the presence of a construct. 1. **Remove the emit and something goes red.** Delete the call that records a settled decision and a test must fail. A test asserting only that the record type exists, or that a well-formed record serialises, stays green when nothing ever emits one — that is the #506 family and does not pass. 2. **A decision with no operator in it is findable without knowing it happened.** Given a fleet that settled a decision while the operator was away, the operator must be able to list every such decision from one place, without being told a ticket number first. "It is in the logs" is not a pass; the daemon log was unreadable on fleet01 for 40 hours and nobody noticed. 3. **A failed notification is visible as a failure.** If the sink is down or refuses, the surface must say so. Reporting `full` while sending nothing is exactly the #392 defect and must not be reproduced here. 4. **The record survives the session.** Kill the lead session after it settles a decision, start a fresh one, and the decision must still be readable. A record that lives only in a pane or a conversation does not pass. ## Scope note This issue owns **the event and the record**. #392 owns **the sink**. Fixing this one by adding a second notification path beside #392 would leave two knobs that both claim to notify, which is the shape #497 catalogues. Coordinate the two. Related: #591 (the policy that created this gap has no delivery surface for a lead), #383 (a detected health state with no API surface — the same "detected but unreportable" shape).
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#592