Files
fleetd/docs/M4-Fleet-Health.md
Dai Ha 2e138a199b CB-634: one shared "fleet" workspace + rename bridged -> fleetd cutover
Two changes ship together here.

1. One shared herdr workspace. The lead and every worker now live in one
   workspace called "fleet", so the operator sees one "session" with many
   windows, not two. Before, the lead sat in a "leads" workspace and workers
   in "bridged-workers", which read as two sessions. The lead is still told
   apart from workers by its exact tab label ("lead: <name>"), so putting them
   in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
   for split layouts; Fleetd now passes an empty exclude set.

2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
   launchd/systemd units, module dir, and MCP mount).
   - Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
   - Log line, comments, docs, and CLAUDE.md updated to say fleetd.
   - Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
     bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
   - Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
     bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
   - Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
     still read as a fallback, and still gitignored.
   - MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
     server name is "fleet". The mount name in the local .mcp.json becomes
     "fleet" (gitignored, not in this commit).
   - Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
     BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.

Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.

Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.

The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).

949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
2026-08-25 04:01:08 +02:00

51 KiB

M4 - Fleet health, recovery, routing, and capacity

Status: Design accepted on 2026-08-15. CB-573 part 1 has shipped the classification model and the fleet_list capacity view; the remaining M4 units are not yet shipped. See Unit 2 — what has landed so far before planning Unit 2 work: some of its criteria were met by separate CB tickets, and one of them contradicts the unit text. Scope: Fleet evidence, safe mechanical repair, lead routing, capacity reporting, and optional human notification. Grounded in: health/FleetHealth, health/PaneBudget, inject/StatusPoller, inject/StatusRefiner, inject/CompletionResolver, inject/Injector, session/SessionManager, msg/MessageService, msg/ReplyInbox, msg/ReplyPushLoop, msg/LeadHeartbeatLoop, mcp/PrimaryRegistry, and herdr/AgentControl.

1. Problem and decision boundary

The operator asked the bridge to detect idle agents, exceptions, stopped work, and broken communication. The bridge may read an agent pane from time to time. It must notify a person when the fleet cannot move forward.

The four operator terms are not four equal health states. IDLE is a normal mode. An exception is sometimes visible only as pane text. Stopped work may look the same as slow work. Broken communication can occur on several links.

M4 uses this boundary:

  • The bridge detects facts and joins evidence.
  • The bridge repairs only mechanical failures with no judgement.
  • The lead decides whether to stop, retry, replace, or reassign a member.
  • A human is notified only when no healthy lead can act.
  • n8n may route an outbound incident. It never classifies state or chooses recovery.

An inbound n8n decider would need bridge authority. No narrow machine-decider role exists. Giving a workflow engine lead authority is unsafe, while adding a new role is a separate authorization design. An outbound sink needs no bridge role.

The bridge must never replay a delivered task. That task may already have changed files, pushed a branch, opened a pull request, or changed external state. A replay can run those side effects twice. This rule must remain true even if later code stores delivered prompt text.

2. Evidence model

A health state is mainly a comparison between two views:

  • herdr view: current agents and raw live status from one AgentControl.list() call.
  • bridge view: session FSM, MCP presence, accepted turns, tasks, inbox state, and lead ownership.

A strong fault often appears as a disagreement between those views. For example, BUSY in the session FSM and DONE in herdr means the bridge missed a turn boundary. Pane reads support this model, but they are not the main monitor.

SessionManager.rosterView already joins session state and live status. AgentControl.list() already gets the whole live fleet in one call. M4 makes that join persistent and adds timers, accepted-turn state, and incident state.

2.1 Real traces behind the design

The first trace was an architect that stopped making progress:

profile=opus role=architect state=busy liveStatus=done

The session moved from DONE to BUSY for turn 2. Eighteen minutes later, the session still said BUSY, herdr still said DONE, the async task still said PENDING, and no completion fallback had run. This is TURN_BOUNDARY_LOST, not a general slow-turn guess.

The second trace had two async sends to the same pane, one second apart. The pane was then stopped. One ticket became failed. The other stayed pending - worker unknown. Current MessageService.abandon resolves only Rendezvous.currentWaiter(target), while async tasks live in a separate ticket map. CB-568 is intended to fix that bug. M4 still keeps an independent post-teardown invariant so a later regression becomes DELEGATION_ORPHANED.

2.2 Corrections made during design

The first state table missed BUSY in fleetd plus IDLE or DONE in herdr. It would have found the real trace only through a late, weak stall timer. The final model adds TURN_BOUNDARY_LOST as a strong disagreement state.

The first notification design also required a webhook before health.enabled could turn on. That removed useful local detection to avoid a narrower human-notification gap. The final design splits detection from notification. Missing human escalation is shown as partial coverage instead of disabling health.

3. Classification precedence

Evidence is applied in this order. A lower rule cannot hide a higher one.

  1. Control link: failed fleet list plus failed ping becomes CONTROL_LINK_DOWN.
  2. Definitive target loss: _not_found becomes GONE or LEAD_UNREACHABLE when the control link is healthy.
  3. Startup and teardown invariants: readiness expiry becomes NEVER_READY; surviving tasks after teardown become DELEGATION_ORPHANED.
  4. Bridge/live disagreement: BUSY plus stable raw IDLE or DONE becomes TURN_BOUNDARY_LOST.
  5. Known screen evidence: a tested fatal signature becomes ERROR_ON_SCREEN.
  6. Timed suspicion: unchanged sparse pane probes may become STALL_SUSPECTED.
  7. Communication quality: completion fallback becomes MUTE; an old inbox entry becomes REPLY_STRANDED.
  8. Normal mode: STARTING, IDLE, WORKING, WORK_PENDING, or BLOCKED_AMBIGUOUS.

The member flow in Figure 1 shows lifecycle states and the main fault exits. Fault states are reported beside the session FSM; most are not new FSM values.

flowchart TD
    Registered["Member registered"] --> Starting["STARTING"]
    Starting -->|"MCP presence"| Idle["IDLE"]
    Starting -->|"Readiness grace expires"| NeverReady["NEVER_READY"]
    Idle -->|"Accepted delivery"| Working["WORKING"]
    Working -->|"Trusted turn boundary"| Idle
    Working -->|"Bridge BUSY and herdr IDLE or DONE"| Lost["TURN_BOUNDARY_LOST"]
    Working -->|"Known fatal screen"| Error["ERROR_ON_SCREEN"]
    Working -->|"Long age and unchanged sparse probes"| Stall["STALL_SUSPECTED"]
    Working -->|"Target not found"| Gone["GONE"]
    Idle -->|"Inbox or queued delivery exists"| Pending["WORK_PENDING"]
    Pending -->|"Delivery or collection finishes"| Idle
    Idle -->|"Raw BLOCKED with an open turn"| Blocked["BLOCKED_AMBIGUOUS"]
    Lost -->|"Strict guarded repair"| Repaired["DONE with reconciled completion"]
    Lost -->|"Repair refused"| LeadDecision["Lead decision required"]

Figure 1. The member lifecycle and the main health exits. Pane-based states never authorise an automatic retry of the task.

4. State model

4.1 Normal and transitional member states

State Exact evidence Meaning and certainty
STARTING Session is SPAWNING; MCP presence is absent Normal inside the startup grace. MCP contact is the readiness signal.
IDLE Session is READY or DONE; live status is IDLE or DONE; no open turn or inbox item exists Normal. Idle is not a fault.
WORKING Session is BUSY; raw live status is WORKING; the accepted turn is open Certain that herdr sees work. It does not prove useful progress.
WORK_PENDING Queued delivery or inbox content exists while the target is injectable Transitional. Existing injector or push logic should move it.
BLOCKED_AMBIGUOUS An open turn exists and raw live status is BLOCKED The bridge cannot tell whether this is permission, input, or a settled screen.

Idle may drive configured resource cleanup. It never opens an incident and never pages a person.

4.2 Member fault and quality states

State Exact evidence Certainty and action
NEVER_READY SPAWNING, no MCP presence, and an accepted delivery waits through the existing readiness grace Delivery never became possible. The exact cause is unknown. Fail the send, stop the process, and preserve a provisioned worktree.
GONE Per-target herdr call returns _not_found while fleet list or ping works Certain target loss. Fail all target work. Do not replay it.
TURN_BOUNDARY_LOST Same session turn stays BUSY; same accepted task stays open; two raw snapshots show IDLE or DONE Strong disagreement. Strict reconciliation may repair it.
ERROR_ON_SCREEN Suspicious non-working state survives grace; detection matches a tested adapter-specific fatal signature Certain only for the matched signature. A bare word such as Exception is not enough.
STALL_SUSPECTED Open turn is older than the configured threshold; two normalised recent_unwrapped digests are unchanged; no boundary or reply occurs Not certain. A long valid API call can look the same. Lead decides.
MUTE Turn resolves through completion fallback instead of fleet_reply Certain that no structured reply won. It does not prove an MCP failure. A single event is a metric, not an incident.
REPLY_STRANDED Typed reply or health message remains after owning-lead push reaches its cap Collection failed. This does not explain whether the lead is busy, dead, or ignoring the nudge.
DELEGATION_ORPHANED Target is gone, failed, or released, but one or more tasks remain PENDING after reconciliation grace Certain bridge invariant failure. This is not an inbox-drain fault.
WORK_PRODUCT_AT_RISK Provisioned branch has commits after its recorded base; member is DONE, FAILED, or preserved after release; no turn or inbox item remains; long-idle threshold passed A warning, not proof of loss. Work may already have an open pull request or a squash merge.

MUTE opens an incident only after a small fixed rate threshold for one target or profile, or when it appears with another fault.

WORK_PRODUCT_AT_RISK must not become WORK_PRODUCT_UNCOLLECTED. The bridge does not know pull request or merge state. If committed work appears with REPLY_STRANDED or DELEGATION_ORPHANED, the existing incident gains committedWorkAtRisk: true.

State Exact evidence Certainty and action
CONTROL_LINK_DOWN Two full-fleet agent.list calls fail across the grace, and herdr ping also fails Certain for the fleetd-to-herdr link. Retry calls, record the incident, and use human escalation if no lead can be reached.

A failed fleet list alone is not a dead-member claim. A single _not_found with a healthy global link is a target fault, not a control-link fault.

4.4 Lead states

State Exact evidence Meaning and action
LEAD_IDLE Expected lead is present with raw injectable status; no actionable state waits Normal. Existing heartbeat may run under its own policy.
LEAD_WORKING Expected lead is present with raw WORKING; stall threshold is not met Reachable and busy. Never inject into the live turn.
LEAD_STATUS_UNKNOWN Expected lead is present with raw UNKNOWN Neither dead nor a healthy routing target. Retain evidence and retry.
LEAD_UNREACHABLE Expected lead is absent from two successful live-agent snapshots while ping works, or targeted lookup returns _not_found with a healthy control link Route to a healthy peer. If none exists, use human escalation.
LEAD_UNRESPONSIVE Actionable state waits; lead stays injectable; bounded nudges exhaust; inbox remains uncollected Route to a healthy peer or a person.
LEAD_STALL_SUSPECTED Lead stays WORKING past threshold; two sparse pane probes show no progress Not certain. Never kill or restart automatically. Route to peer or person.

The monitor retains the lead name and terminal, last successful sighting, raw status and age, consecutive list absences, targeted errors, pane-probe facts, pending incident age, and nudge outcomes. Current heartbeat and push loops discard much of this history.

Expected lead identity comes from the same supplier used by CallerResolver. It is not liveness evidence. LeadTabScanner keeps cached identity after a failed scan, so the health monitor compares that identity with a fresh successful agent list. A dynamic identity also survives a two-successful- snapshot retirement grace. This stops a dead lead from escaping health by disappearing from one map.

4.5 Evidence limits

M4 cannot tell these cases apart with current evidence:

  • A valid long call and a hung call may have the same status and pane digest.
  • BLOCKED does not explain which input is needed.
  • An idle prompt after failure may look like an idle prompt after success.
  • A missing structured reply does not prove a broken MCP connection.
  • An undrained inbox does not explain why the lead did not collect it.
  • Arbitrary pane text cannot safely classify arbitrary exceptions.
  • A branch ahead of its base does not prove that work was not collected.

Logs are outputs, not classifier inputs. The monitor never parses its own logs.

5. Automatic action and lead action

5.1 Actions the bridge may take

The bridge may:

  • retry transient herdr status, list, ping, and pane-read failures with bounded backoff;
  • re-submit Enter after the existing paste/submit race;
  • fail queued delivery after NEVER_READY;
  • stop a never-ready process while preserving its provisioned worktree;
  • fail all queued, accepted, and async tasks for a gone or released target;
  • reconcile one lost boundary when every strict gate in Section 8 passes;
  • hold typed messages, nudge the owning lead, and stop at the configured cap;
  • use the existing bounded idle-lead heartbeat;
  • deduplicate, route, update, and resolve incidents.

These actions do not choose new work and do not replay old work.

5.2 Decisions reserved for the lead

Only the lead may:

  • stop or continue BLOCKED_AMBIGUOUS;
  • stop, inspect, or wait on ERROR_ON_SCREEN;
  • kill or continue STALL_SUSPECTED;
  • spawn a replacement or reassign work;
  • retry a delivered task;
  • choose how to use partial work in a worktree;
  • restart herdr or change network, model, credentials, backend, or configuration.

Reports include literal safe tool calls such as fleet_status(sessionId="..."), fleet_poll(ticket="..."), fleet_list(), and optional fleet_stop(paneId="..."). A judgement state never presents stop as the only action.

5.3 Release causes and worktree safety

Release cause Process action Provisioned worktree
SPAWN_ROLLBACK before registration or delivery Stop and clean up Remove
COMPLETED for READY or DONE without pending work, idle TTL, or successful context-cap completion Stop Remove only if clean; preserve a dirty worktree (CB-576)
NEVER_READY Stop Preserve
GONE Best-effort stop Preserve
TURN_FAILED or lead abort while BUSY or FAILED Stop Preserve
RELEASE_WITH_PENDING_TASKS Stop Preserve
SHUTDOWN Stop Preserve

Explicit stop is state-aware. SPAWNING, BUSY, FAILED, or any target with pending tasks uses a preserving cause.

Before abnormal release removes the live session, M4 writes an atomic manifest under the worktree root. It records session identity, owner, role, profile, repository, path, branch, base commit, release cause, release time, state, and pending task ids. fleet_list.preservedWorktrees loads these manifests after restart. Stop output and WARN logs also name the path and cause. M4 never auto-deletes a preserved worktree.

6. Fleet health monitor

Add FleetHealthMonitor. Do not widen StatusPoller into a policy loop.

StatusPoller has a 250 ms delivery cadence and samples only injector targets with outstanding work. Health needs all sessions, all leads, task state, inbox age, and global control evidence. One loop cannot serve both cadences safely.

Build the monitor like LeadHeartbeatLoop:

  • pure decide(snapshot, priorState, now) logic;
  • a thin scheduler;
  • an injected clock;
  • edge-triggered state changes;
  • no network work in the pure function;
  • no sleeping in tests.

Each enabled fleet tick reads:

  • one AgentControl.list() result for the whole fleet;
  • one in-memory SessionManager.roster() snapshot;
  • accepted turns and async task state;
  • typed inbox depth, kind, and age;
  • push and heartbeat outcomes;
  • configured and discovered leads.

Existing failure paths publish structured evidence to the monitor. The monitor does not infer events from log text.

6.1 Pane budget

Healthy idle members, recent working members, and quiet leads cause no pane reads.

A pane is eligible only for a stable lost boundary, sustained BLOCKED or UNKNOWN, work older than the suspect threshold, or one final evidence read for a confirmed fault when the pane exists.

Compiled brakes apply even if config asks for more:

  • per-target pane cooldown is at least 60 seconds;
  • working age before the first progress probe is at least 300 seconds;
  • at most two pane reads occur in one fleet tick;
  • targets rotate fairly;
  • only a normalised digest and optional clipped local excerpt are stored;
  • no pane excerpt leaves fleetd in a human webhook.

Use detection for tested screen signatures. Use normalised recent_unwrapped only for progress comparison.

7. Typed inbox and routing

7.1 Semantic record

The typed inbox record carries:

schemaVersion
kind: reply | health
msgId, target, subjectTerminal, recipientLead
severity, state, evidence
createdAtEpochMillis, firstSeenEpochMillis, lastSeenEpochMillis
recoveryTried, suggestedToolCalls, content

A health message never calls Rendezvous.resolve. It cannot look like the member's task result.

Both inbox adapters share field preservation, first-id-wins dedup, FIFO among decoded messages, explicit ownership, ack, and release rules. The in-memory adapter stores typed records directly. It does not copy AMQP migration logic.

7.2 AMQP migration

The reader uses AMQP content_type, never body sniffing:

Legacy v0: text/plain
Typed family: application/vnd.ltms.fleet.inbox-message+json

A legacy reply may begin with {. It remains plain text because its media type is text/plain. Legacy text becomes kind=reply with exact UTF-8 content and absent typed metadata.

Typed JSON has required integer schemaVersion: 1. Version 1 ignores unknown optional fields. Missing required fields, invalid enums, malformed UTF-8 or JSON, and property/body identity mismatch are invalid data.

An unknown schema version is not partly decoded. It remains unacknowledged on the original queue and creates one operator-visible unsupported_version failure. A newer daemon may read it later.

Invalid known-format data is copied byte-for-byte to durable queue agent.<target>.inbox.quarantine. A dedicated confirm-mode publisher confirms the persistent copy before the original is acknowledged. A failed quarantine handoff leaves the original unacknowledged. The raw body never enters logs.

Decode failure creates a redacted WARN, metric, fleet_list summary, and routed health incident. One bad entry never escapes the consumer callback and never stops later valid messages.

Safe downgrade is not supported. The previous build ignores content_type and would show typed JSON as ordinary reply text. If drained, it would acknowledge the message and lose typed meaning. Typed queues must be drained or preserved before an old jar runs.

The existing contract suite uses RabbitMQ. Production uses LavinMQ. The migration and lead-key ownership cases must run once against production LavinMQ before release, or the release must state that LavinMQ was not checked.

7.3 Member routing

A member incident first goes to the exact lead that owns its accepted delegation. PrimaryRegistry needs a no-fallback delegatingLeadFor(memberTarget) query. Health routing must not use the old singular-primary fallback when several leads exist.

Publish the incident under the affected member target. Trigger the existing bounded push route. The push waits until the owning lead is injectable, so it does not interrupt a live lead turn.

7.4 Peer lead routing

Figure 2 shows the route from incident to lead, peer, or person.

flowchart TD
    Incident["Open incident"] --> Member{"Member incident?"}
    Member -->|"yes"| Known{"Exact delegation owner known?"}
    Known -->|"no"| Sink{"Human webhook enabled and healthy?"}
    Known -->|"yes"| Owner{"Owner lead healthy?"}
    Owner -->|"yes"| OwnerInbox["Publish to owner lead path"]
    Owner -->|"no"| PeerSet["Build healthy peer candidate set"]
    Member -->|"no, lead incident"| PeerSet
    PeerSet --> Peer{"Healthy peer exists?"}
    Peer -->|"yes"| Select["Choose fewest assigned incidents<br/>then stable name and terminal id"]
    Select --> PeerInbox["Publish to peer lead inbox<br/>and status-gated push"]
    Peer -->|"no"| Sink
    Sink -->|"yes"| Webhook["Send classified outbound incident"]
    Sink -->|"no"| Passive["Keep incident open<br/>show partial coverage on local surfaces"]

Figure 2. Routing keeps delegation ownership separate from temporary peer fallback.

Peer candidates exclude the incident subject, failed owner, absent leads, raw-unknown leads, and leads with an open unhealthy state. A reachable WORKING peer may be selected; its push waits for an injectable window.

Choose the candidate with the fewest assigned foreign incidents. Break ties by stable lead name, then terminal id. Pin the recipient. Reassign only if that peer becomes unhealthy or retires. A routing generation marks a reassignment, and old pending assignments become superseded.

fleet_list lead rows show health, health age, assigned foreign incident count, and a bounded list of incident id, subject, state, severity, age, and routing generation. The top-level view also shows owner, recipient, and routing reason.

A peer incident is published under the recipient lead's inbox key, not the failed subject's key. Its status-gated nudge names the failed lead and gives the exact fleet_poll(target="<recipient-terminal>") call.

7.5 Lead inbox ownership

Add LeadInboxRegistry, driven by the same expected-lead supplier as CallerResolver.

It calls replyInbox.own(leadTerminal) at startup for configured leads, after successful discovery, after config adds a lead, and before publication. Ownership is not an authorization side effect.

A missing lead keeps its key owned. Release happens only after confirmed retirement, all incidents are reassigned or resolved, typed health messages move or ack, and the queue is empty. Own a replacement terminal before moving messages from the old key. Never release a non-empty in-memory lead key, because in-memory release clears local data.

7.6 Single-lead deployment

One lead and no peer is a normal mode, not an edge case.

An idle, reachable lead may receive the existing bounded nudge. An unreachable or stalled sole lead has no safe in-loop recovery. The bridge must not restart or replace it. A new lead would not have the failed lead's plan or context, and an uncertain relaunch could create two orchestrators.

With no webhook, only fleet_list, /healthz, metrics, WARN logs, and the incident journal remain. These are passive surfaces. They are not a human notification.

8. Lost-boundary reconciliation

This is the only M4 path that reconstructs a result. It must prefer a visible stall over a fabricated reply.

8.1 Why normal completion rules are not enough

Current CompletionResolver.resolve has two fail-open rules. It resolves when the delivery baseline is missing. It also resolves an empty completion when the pane read fails. Those choices are valid after a trusted WORKING -> IDLE boundary because the bridge knows the turn ran. They are unsafe when health only guesses that a boundary was lost.

M4 gives each accepted send an internal TurnToken. It ties target, exact waiter, session turn, delivery baseline, and task outcome together.

8.2 Delivery baseline

Capture the baseline immediately after prompt send and before the delivery future completes. Store:

TurnToken
exact waiter identity
capture time and pane source
normalised assistant block clipped to MAX_SCRAPE_CHARS
whether a supported assistant marker was recognised
capture result: PRESENT | READ_FAILED | UNRECOGNISED

A failed or missing baseline never authorises repair. A late baseline is not valid evidence. After a daemon restart, the old waiter, task, token, and baseline are gone, so the old turn cannot be repaired.

Automatic repair is enabled only for agent kinds with tested assistant-block fixtures. Current extraction is Claude Code-specific and falls back to arbitrary raw text without ⏺. That raw fallback cannot authorise repair. OpenCode repair stays disabled until live pane fixtures exist.

8.3 Strict gates and resolver result

Figure 3 shows the repair gates. Any failed gate keeps the waiter unchanged.

The two raw snapshots must describe the same TurnToken and session turn. No WORKING, BLOCKED, UNKNOWN, missing-agent, reply, failure, or new-delivery observation may occur between them.

flowchart TD
    Candidate["TURN_BOUNDARY_LOST candidate"] --> Stable{"Same TurnToken and BUSY turn<br/>across two raw IDLE or DONE snapshots?"}
    Stable -->|"no"| Resnapshot["Take a fresh snapshot"]
    Stable -->|"yes"| Waiter{"Exact captured waiter<br/>still open by identity?"}
    Waiter -->|"no"| Stale["STALE_TURN or ALREADY_RESOLVED"]
    Waiter -->|"yes"| Baseline{"Successful recognised<br/>delivery baseline exists?"}
    Baseline -->|"no"| Refuse["Refuse repair<br/>leave ticket pending"]
    Baseline -->|"yes"| Read{"Fresh pane read succeeds?"}
    Read -->|"no"| Refuse
    Read -->|"yes"| Output{"Recognised non-blank assistant block<br/>differs from clipped baseline?"}
    Output -->|"no"| Refuse
    Output -->|"yes"| Resolve["Shared CompletionResolver guard core<br/>resolves exact waiter"]
    Resolve -->|"won race"| Repaired["RECONCILED_COMPLETION<br/>same turn becomes DONE"]
    Resolve -->|"lost race"| Resnapshot

Figure 3. Repair needs stronger evidence than a normal observed turn boundary.

Refactor the current resolver into one guard core with two policies:

resolveCaptured(target, inFlight, OBSERVED_BOUNDARY)
resolveCaptured(target, inFlight, LOST_BOUNDARY_REPAIR)

The health monitor calls only:

CompletionResolver.reconcileLostBoundary(target, expectedTurnToken)

It returns REPAIRED, ALREADY_RESOLVED, REFUSED_NO_CAPTURE, REFUSED_NO_BASELINE, REFUSED_UNREADABLE, REFUSED_UNCHANGED, REFUSED_AMBIGUOUS_OUTPUT, STALE_TURN, or RACE_LOST.

Only REPAIRED and same-turn ALREADY_RESOLVED may move that turn from BUSY to DONE. A per-target reconciliation gate stops a queued second send from being accepted between waiter resolution and the FSM transition.

8.4 Lead-visible marker and refusal

A repaired result uses distinct RECONCILED_COMPLETION values in Rendezvous, MessageService, task poll source, and metrics. The lead sees:

[repaired completion - fleetd detected a lost turn boundary. The member did not call
fleet_reply; pane-derived text follows and may be partial]

Clipped text also keeps the existing clipped-tail marker.

A refused repair leaves TURN_BOUNDARY_LOST open and the ticket pending. The report states that no reply was reconstructed and no task was replayed. UNCHANGED, UNREADABLE, and AMBIGUOUS_OUTPUT get at most one delayed retry for the same token. Missing capture or baseline gets no retry. After two refused scrapes, automatic repair stops for that token.

8.5 Target-wide teardown invariant

CB-568 owns the multi-ticket cancellation mechanism. M4 routes every terminal cause through that one idempotent operation and checks this independent invariant after teardown:

  • no injector entry exists for the target;
  • no accepted turn or completion record exists;
  • no rendezvous waiter or ask exists;
  • every async task is terminal or was already terminal;
  • no thread waiting for the target send lock can later accept it;
  • new sends fail immediately;
  • each old task has one terminal outcome and one metric count.

A violation becomes DELEGATION_ORPHANED. The monitor may call the same idempotent target-wide failure operation once. It never recreates the task.

9. Capacity and utilisation

Capacity is a view, not a health state.

fleet_list adds one block per profile:

profile, maxLoad, live, free, reclaimable

For an unlimited profile, maxLoad and free are null. free is max(0, maxLoad - live) for a capped profile.

The view must use the exact live-count function used by placement. A second calculation could show a free slot that placement then refuses. Member rows add idleForSeconds only when state is READY or DONE, no accepted turn exists, and the inbox is empty. reclaimable means only that the member holds capacity without open bridge work.

The existing idle-lead nudge gains a bounded capacity summary. It lists per-profile live, cap, free, and reclaimable counts, plus at most three long-idle members. Capacity does not make FleetState.hasPending() true. A changed capacity fingerprint may re-arm one capped heartbeat sequence. The fingerprint excludes changing idle durations, so a static idle fleet cannot reset the cap forever. Reply-push stand-down remains first.

The bridge must never:

  • spawn a member because a slot is free;
  • generate a task or acceptance criteria;
  • move queued work to another member or profile;
  • treat a free slot or idle member as an incident;
  • stop an idle member only to improve utilisation.

The bridge knows capacity facts but has no work list. Only the lead has the plan, task context, side-effect history, and acceptance criteria.

Capacity calculation is in memory and adds no pane reads. Work-product checks run on a terminal session edge, not every fleet tick.

This capacity design adds no automatic stop. The accepted NEVER_READY cleanup can still stop a very slow startup after the existing grace, which is a known risk. Free capacity and long idle time never trigger that path.

10. Human escalation and notification

10.1 Escalation rule

Notify a person only when no healthy lead can act:

  • CONTROL_LINK_DOWN survives grace;
  • a lead is unhealthy and no healthy peer can receive the incident;
  • a member incident has no known owning lead;
  • the only owning lead becomes unreachable, unresponsive, or stalled;
  • incident publication or routing itself fails.

Do not page a person for a member fault while a healthy owning lead exists. An uncollected member incident feeds lead-health evidence. If the lead then becomes unhealthy, peer or human routing starts.

10.2 Detection and notification switches

health.enabled controls detection and bridge-local reporting. It does not require a webhook.

health.notifications.mode is disabled or webhook. Disabled is valid and is the default. Webhook mode requires a resolved environment variable. Turning notification off stops outbound attempts but keeps incidents. Turning it back on resumes still-open human incidents.

Without a sink, fleet_list.healthCoverage states that human escalation is unavailable. /healthz keeps its existing HTTP liveness result and adds a nested fleetHealth.status=partial component. Metrics and one startup or reload WARN expose the same limit.

10.3 Incident and delivery deduplication

One open incident uses this key:

(scope, subjectStableId, state, causeFingerprint)

The cause fingerprint includes stable error codes, dependency names, signature ids, or invariant names. It excludes times, ages, retry counts, pane text, and changing digests. A later recurrence after resolution gets a new generation and incident id.

Each outbound event uses:

Idempotency-Key = hash(incidentId, eventType, eventRevision)

Event types are open, severity_changed, reminder, and resolved. Transport retries keep the same key.

An atomic owner-only journal beside the active config stores open incidents, routing, delivered revisions, retry state, and resolution state. It stores no pane or task content. Journal failure does not stop detection, but notification coverage becomes degraded.

10.4 Retry, reminder, and resolve

Send the first event immediately. Retry network errors, timeouts, HTTP 408, HTTP 429, and HTTP 5xx with full-jitter exponential backoff:

base: 5 seconds
factor: 3
maximum delay: 15 minutes
one outstanding attempt per event

Respect Retry-After up to 15 minutes. Other HTTP 4xx responses are permanent for that event until config changes or a person requests replay.

Transport retry is not an incident reminder. humanRepeatSeconds creates a new reminder revision for an unresolved critical incident after the last successful human event. Disabled mode does not build an unbounded reminder queue.

Send resolved only if at least one human event for that incident was delivered. If an incident resolves before its first successful delivery, cancel the pending open event and record local resolution.

10.5 Outbound payload boundary

An outbound payload may contain incident id and event type, severity, state, scope, stable bridge ids, role or profile, times, duration, structured evidence type and counts, recovery attempted, routing reason, coverage, and safe tool calls.

It must never contain:

  • raw pane text, pane excerpts, or pane digests;
  • task briefs, prompts, or member reply content;
  • source files, diffs, or worktree file content;
  • worktree paths;
  • environment values, tokens, credentials, headers, or webhook URL;
  • raw exception messages or stack traces;
  • arbitrary model output.

The sink response body is ignored. A webhook cannot direct recovery. n8n remains outbound-only.

10.6 Metrics

M4 adds bounded-label series:

fleet_health_incidents{scope,state,severity}
fleet_health_incidents_total{event}
fleet_health_notifications_total{event,outcome}
fleet_health_notification_queue_depth
fleet_health_notification_last_success_seconds
fleet_health_notification_capability{mode,status}
fleet_lead_health{lead,state}
fleet_lead_assigned_incidents{lead}

Metric labels never include terminal ids, incident ids, URLs, or error text.

11. Configuration

The optional health: block is absent or disabled by default. The dormant monitor scheduler does no herdr or pane work while disabled. Every listed key is hot because the monitor reads ConfigRef on each tick or notification.

Key Class Default and hard bound Purpose
health.enabled Hot false Enable detection and bridge-local reporting.
health.snapshotIntervalSeconds Hot default 30, minimum 15 Whole-fleet comparison cadence.
health.workingSuspectAfterSeconds Hot default 600, minimum 300 Age before working-pane probes.
health.paneProbeIntervalSeconds Hot default 60, minimum 60 Per-target pane cooldown.
health.leadUnresponsiveAfterSeconds Hot default 300, minimum 120 Delay after exhausted actionable nudges before lead fault.
health.humanRepeatSeconds Hot default 3600, minimum 900 Minimum repeat period for one open human incident.
health.capacityLongIdleAfterSeconds Hot default 900, minimum 300 Long-idle threshold for capacity summaries.
health.includePaneExcerpt Hot false Allow a clipped excerpt in local lead reports only. Human payloads still exclude it.
health.notifications.mode Hot disabled Select disabled or webhook.
health.notifications.webhookUrlEnv Hot required in webhook mode Name of the environment variable that holds the sink URL.
health.notifications.requestTimeoutMs Hot default 10000, range 1000-30000 Whole webhook request limit.

Two consecutive snapshots are compiled floors for lost boundary, lead disappearance, and control link failure. The two-pane-reads-per-tick limit is also compiled and cannot be weakened by config.

12. Delivery units and acceptance

Unit 1 - Evidence model and fleet snapshot

Scope: health state model, fleet join, clocks, evidence retention, and pane budget.

Acceptance criteria:

  1. One agent.list call covers one enabled fleet tick.
  2. Pure decision tests cover every state and every evidence limit in Section 4.
  3. BUSY plus stable raw DONE opens TURN_BOUNDARY_LOST after two snapshots.
  4. Healthy fleet snapshots perform zero pane reads.
  5. Pane cooldown, two-read fleet budget, and fair rotation cannot be disabled by config.
  6. Logs are outputs only; no log parsing exists.
  7. Fleet snapshots expose the same profile live-count calculation that placement uses.
  8. Capacity rows report cap, live, free, and reclaimable values without opening incidents.

Unit 2 - Lost boundary and task reconciliation

Scope: accepted-turn identity, guarded repair, target-wide teardown, release causes, and preserved worktree discovery.

Acceptance criteria:

  1. Every accepted send receives a stable TurnToken tied to target, exact waiter, and delivery baseline.

    Corrected during implementation (2026-08-15). This criterion first also required the session turn number and the task outcome. That is not implementable at this layer, and the implementer refused it three times rather than fabricate a value — correctly. The reason is an ordering fact that is invisible from any single class: MessageService owns acceptance and holds the waiter and the async Task, but it learns nothing about delivery, because the delivery event goes to CompletionResolver through TurnListener.onDelivered. And CompletionResolver.onDelivered runs before SessionManager.onDelivered, so the session turn number does not exist yet at the only point where the token could capture it.

    Two ways out were rejected. A shared registry keyed by target reintroduces exactly the "whichever send happens to be waiting" ambiguity the token exists to remove — the same weak claim Rendezvous.currentWaiter warns about. Injecting a turn counter into MessageService adds a required cross-layer dependency to populate a field that nothing in this slice reads, which is speculative coupling across a boundary already shown to be fragile.

    So the token identifies the accepted send, and SessionManager keeps verifying its own delivery separately. Repair (criterion 2) does need the session turn; binding it means resolving that acceptance-versus-delivery ordering first, and that work belongs to the repair unit, not here. The token record carries a comment saying the field is deliberately absent.

  2. Repair requires the same BUSY token, two raw IDLE or DONE snapshots, no conflicting observation, exact open waiter, successful baseline, and new recognised assistant output.

  3. Missing, failed, late, or post-restart baseline never authorises repair.

  4. Repair is enabled only for agent kinds with tested assistant-block extraction. Raw-text fallback without a recognised marker refuses repair.

  5. Normal completion and repair use one resolver guard core. Waiter, scrape, clipping, unchanged, and exact-turn guards are not duplicated.

  6. reconcileLostBoundary returns every typed result named in Section 8.3.

  7. Only REPAIRED and same-turn ALREADY_RESOLVED may move the same turn to DONE.

  8. A per-target reconciliation gate blocks a queued second send during repair and FSM update.

  9. Repaired completion has distinct rendezvous kind, message outcome, poll source, lead marker, and metric. Clipping keeps its extra marker.

  10. Unchanged, unreadable, or ambiguous evidence gets at most one delayed retry. Missing capture or baseline gets none.

  11. Refusal leaves the ticket pending and tells the lead that no result was rebuilt or replayed.

  12. Release, gone, never-ready, and abnormal stop use CB-568's one idempotent target-wide failure operation.

  13. The post-teardown invariant in Section 8.5 is tested independently of CB-568 internals.

  14. A violated teardown invariant creates DELEGATION_ORPHANED and retries only the idempotent failure operation.

  15. SPAWN_ROLLBACK and normal COMPLETED remove worktrees. Abnormal and shutdown causes preserve them.

  16. Explicit stop is state-aware. Any pending task or non-terminal state preserves the worktree.

  17. Atomic preserved-worktree manifests reload after restart and appear in lead-only fleet_list.preservedWorktrees.

  18. Manifest failure preserves the worktree and opens an operator-visible health failure.

  19. Provision records the base commit. Terminal, long-idle worktrees report WORK_PRODUCT_AT_RISK only under the evidence in Section 4.2 and never auto-delete work.

  20. No path replays a delivered task, rebuilds its brief, or retargets it, even when prompt text is available.

  21. Tests cover both real traces, all repair refusals, clipping, explicit-reply and next-turn races, restart without capture, concurrent send and release, and preserved discovery after restart.

Unit 2 - what has landed so far

Checked against main at e09cac6 on 2026-08-15. Unit 2 was written as one block, but parts of it have since been built by separate CB tickets. Read this before planning the rest, or that work gets done twice.

The check was a symbol survey of fleetd/src/main/java plus the merge history. It tells you whether the machinery exists at all. It is not a line-by-line audit of whether each criterion is fully met, and I did not run one.

Criterion Marker searched for Found in main source Reading
1 TurnToken 8 files Done — unit 2a, merged as fec284e. Criterion 1 was corrected first; see the note under it.
2-5, 9 REPAIRED 0 files Not started. The whole guarded-repair path is absent.
6, 7, 10 reconcileLostBoundary 0 files Not started.
12 CB-568 failure operation via CB-580 Partial. CB-580 (0af902e) routes GONE and NEVER_READY into the one idempotent target-wide failure. I did not check that release and abnormal stop go through the same call.
14 DELEGATION_ORPHANED 3 files Partial. The health state exists. The teardown-invariant check that creates it, and the retry rule, do not.
15 SPAWN_ROLLBACK 0 files Contradicted — see below.
16 — — Partial at best. CB-576 made release preserve a dirty worktree; whether explicit stop is state-aware is not checked.
17, 18 preservedWorktrees 0 files Not started. No manifest, and no lead-only fleet_list field.
19 WORK_PRODUCT_AT_RISK 0 files Not started.

Criterion 15 no longer matches the code, and the code is right. It says "normal COMPLETED remove worktrees". Since CB-576 (500bfa2) that is false on purpose: a COMPLETED release now preserves the worktree when it still holds uncommitted work, because deleting it destroys work nobody can get back. CB-576 was filed after exactly that loss. CB-581 goes further — if the dirty-check itself fails, the worktree is preserved rather than removed, since "we could not tell" must not be treated as "it is clean".

So criterion 15 should be rewritten as: SPAWN_ROLLBACK and a COMPLETED release with a clean worktree remove it; abnormal causes, shutdown, a dirty worktree, and a failed dirty-check all preserve it. SPAWN_ROLLBACK itself does not exist yet.

Unit 3 - Typed inbox and member routing

Scope: semantic record, AMQP migration, both adapters, member routing, polling, and member health in fleet_list.

Acceptance criteria:

  1. AMQP selects legacy or typed decoding only from content_type; it never sniffs the body.
  2. Persistent text/plain from the old build becomes kind=reply with exact UTF-8 content, including content beginning with {.
  3. New entries use the vendor media type, schemaVersion: 1, UTF-8, persistent delivery, and AMQP message ids.
  4. Version 1 ignores unknown optional fields but rejects missing fields and identity mismatch.
  5. Unknown versions are not decoded or acked. They remain on the original queue and create one deduplicated failure.
  6. Invalid known data never escapes the callback, appears as a reply, or blocks later valid messages.
  7. Invalid data reaches durable per-target quarantine before original ack. Failed handoff leaves the original unacked.
  8. Decode failures create redacted WARN, metric, fleet_list summary, and routed incident without raw content.
  9. Both adapters pass one semantic contract for fields, FIFO, dedup, ownership, ack, and release.
  10. Lead keys require explicit ownership. Publication never claims a queue.
  11. Unit codec tests cover legacy {, Unicode, malformed UTF-8, typed round trip, additive fields, malformed JSON, missing fields, identity mismatch, media type, version, and dedup.
  12. A live broker contract writes old wire data and reads it with the new adapter after reconnect.
  13. Live contract tests cover mixed entries, quarantine confirm-before-ack, unsupported redelivery, later progress past poison, property persistence, lead ownership, and ack removal.
  14. Safe downgrade is documented as unsupported.
  15. RabbitMQ contract tests pass with mvn test -Pcontract. The same cases run once on production LavinMQ, or the release states that LavinMQ was not checked.
  16. Member incidents route to the exact delegating lead and never resolve a task rendezvous.
  17. fleet_list shows compact member health and capacity without pane content. Member idleForSeconds is present only when no accepted turn or inbox item exists.

Unit 4 - Lead health and peer routing

Scope: lead evidence, exact ownership, peer selection, explicit-recipient push, and lead inbox lifecycle.

Acceptance criteria:

  1. Lead identity uses the CallerResolver supplier. Liveness uses successful current agent data.
  2. Two successful-list absences with healthy ping become LEAD_UNREACHABLE; global link failure does not.
  3. Raw WORKING, raw UNKNOWN, first-seen time, failures, last success, and error class persist across ticks.
  4. Heartbeat and push publish status and nudge outcomes before safe no-injection decisions.
  5. Dynamic lead identity survives a two-successful-snapshot retirement grace.
  6. Member incidents first use exact delegation ownership with no singular-primary fallback.
  7. Peer selection follows the exclusions, load rule, and stable tie break in Section 7.4.
  8. A selected working peer is not interrupted. Its push waits for an injectable window.
  9. Recipient assignment stays pinned. Reassignment increments generation and supersedes old pending assignment.
  10. fleet_list shows bounded foreign assignments, recipient, reason, and generation without pane content.
  11. LeadInboxRegistry owns configured and discovered lead keys before publication.
  12. Missing leads keep ownership. Retirement needs an empty queue and handled incidents.
  13. Replacement owns the new key before messages move. Non-empty in-memory keys are not released.
  14. Tests cover dead versus busy, unknown, global failure, stale scan cache, disappearance, one peer, several peers, reassignment, and no peer.
  15. Adapter tests cover lead ownership, restart re-ownership, retirement, and terminal replacement. LavinMQ is checked or named as unchecked.
  16. A sole unreachable or stalled lead is never restarted or replaced. Without a sink, only passive evidence remains and every coverage surface says so.

Unit 5 - Human sink, hot config, metrics, and operator coverage

Scope: generic webhook, config split, incident journal, retry, resolve, metrics, example config, and operator documentation.

Acceptance criteria:

  1. health.enabled works without a human sink.
  2. Notification mode is hot, defaults to disabled, and supports disabled or webhook.
  3. Webhook mode requires a resolved environment value. Bad notification config does not disable an already valid detector.
  4. Mode changes keep open incidents. Re-enable resumes eligible incidents.
  5. fleet_list, /healthz, metrics, and one WARN show partial coverage without a sink. HTTP liveness behavior stays unchanged.
  6. One-lead, no-sink coverage states that lead failure has no active notification or recovery.
  7. Incident and outbound dedupe use the stable keys in Section 10.3.
  8. The owner-only local journal survives restart and contains no pane or task content.
  9. Journal failure keeps detection running but marks notification coverage degraded.
  10. Retry tests cover network failure, timeout, 408, 429, Retry-After, 5xx, permanent 4xx, jitter, delay cap, config re-arm, and one outstanding attempt.
  11. Reminders and transport retries remain separate. Disabled mode does not build an unbounded queue.
  12. Resolve sends only after an earlier human event succeeded. Resolve-before-delivery cancels stale open delivery.
  13. Metrics use bounded labels and exclude ids, URLs, and error text.
  14. Payload tests reject every content type forbidden in Section 10.5.
  15. Webhook response bodies are ignored and cannot direct recovery.
  16. Tests cover disabled mode, one lead without sink, open/update/reminder/resolve, restart, dedup, reassignment, disable/re-enable, and sink failure while local health continues.
  17. fleetd.example.yaml documents all hot keys and compiled floors.
  18. The operator Features wiki is updated separately. The portable CLAUDE.md block is checked and changed only if shipped tool or inbox semantics make it untrue.
  19. mvn clean install passes.

13. Not checked and release gates

These limits are part of the design, not optional follow-up notes.

  • OpenCode pane status and assistant markers were not checked. OpenCode lost-boundary repair is disabled until live fixtures exist.
  • Permission-prompt status was not checked for Claude Code or OpenCode. BLOCKED remains ambiguous and has no automatic action.
  • recent_unwrapped stability was not checked across all supported agent kinds. If normalisation is not stable, STALL_SUSPECTED must say its evidence is weaker.
  • The real BUSY + DONE trace was not replayed against live herdr. The design uses the observed production trace and current poller behavior.
  • CB-568 was not present when Unit 2 was designed. Unit 2 must inspect the landed API and keep its independent teardown invariant.
  • Production LavinMQ was not checked. Existing durable-inbox contracts use RabbitMQ. Migration, quarantine, redelivery, lead ownership, and reassignment must run on LavinMQ before release or be recorded as unchecked.
  • Live multi-lead routing was not checked. Peer choice and reassignment are design rules backed by fake-clock and adapter tests until a live exercise runs.
  • A live sole-lead failure with a webhook was not checked. The no-peer path is a design result, not a tested recovery.
  • No n8n, Slack, PagerDuty, or other receiver was checked. The webhook remains generic and outbound-only.
  • Deployment supervisor behavior for nested /healthz fields was not checked. HTTP liveness status stays unchanged to reduce this risk.
  • Incident-journal crash behavior was not checked because the journal does not exist yet. Unit 5 must test atomic replacement and restart recovery.
  • Worktree merge state cannot be checked reliably without forge or explicit collection evidence. WORK_PRODUCT_AT_RISK stays a warning.

14. Locked exclusions

M4 does not expose agent.read as a bridge tool. It does not add a workflow engine, inbound n8n authority, automatic task assignment, task replay, automatic lead replacement, or automatic member spawn for free capacity.

The bridge remains a message bus with evidence and bounded mechanical repair. The lead remains the place where judgement and work planning happen.