Unit 3 renamed bridged.example.yaml to fleetd.example.yaml and taught Fleetd to read fleetd.yaml first. Two things it left behind. bridged/.gitignore still ignored only bridged.yaml. An operator who follows the new comment and copies the example to fleetd.yaml gets an untracked live config holding tokens, and git offers to commit it. Both names are ignored now, because both names work until the cutover. Three design docs still pointed readers at bridged.example.yaml, a file that no longer exists under that name.
51 KiB
M4 - Fleet health, recovery, routing, and capacity
Status: Design accepted on 2026-08-15. CB-573 part 1 has shipped the classification model and
the fleet_list capacity view; the remaining M4 units are not yet shipped. See
Unit 2 — what has landed so far before planning Unit 2 work:
some of its criteria were met by separate CB tickets, and one of them contradicts the unit text.
Scope: Fleet evidence, safe mechanical repair, lead routing, capacity reporting, and optional
human notification.
Grounded in: health/FleetHealth, health/PaneBudget, inject/StatusPoller,
inject/StatusRefiner, inject/CompletionResolver, inject/Injector, session/SessionManager,
msg/MessageService, msg/ReplyInbox, msg/ReplyPushLoop, msg/LeadHeartbeatLoop,
mcp/PrimaryRegistry, and herdr/AgentControl.
1. Problem and decision boundary
The operator asked the bridge to detect idle agents, exceptions, stopped work, and broken communication. The bridge may read an agent pane from time to time. It must notify a person when the fleet cannot move forward.
The four operator terms are not four equal health states. IDLE is a normal mode. An exception is
sometimes visible only as pane text. Stopped work may look the same as slow work. Broken
communication can occur on several links.
M4 uses this boundary:
- The bridge detects facts and joins evidence.
- The bridge repairs only mechanical failures with no judgement.
- The lead decides whether to stop, retry, replace, or reassign a member.
- A human is notified only when no healthy lead can act.
- n8n may route an outbound incident. It never classifies state or chooses recovery.
An inbound n8n decider would need bridge authority. No narrow machine-decider role exists. Giving a workflow engine lead authority is unsafe, while adding a new role is a separate authorization design. An outbound sink needs no bridge role.
The bridge must never replay a delivered task. That task may already have changed files, pushed a branch, opened a pull request, or changed external state. A replay can run those side effects twice. This rule must remain true even if later code stores delivered prompt text.
2. Evidence model
A health state is mainly a comparison between two views:
- herdr view: current agents and raw live status from one
AgentControl.list()call. - bridge view: session FSM, MCP presence, accepted turns, tasks, inbox state, and lead ownership.
A strong fault often appears as a disagreement between those views. For example, BUSY in the
session FSM and DONE in herdr means the bridge missed a turn boundary. Pane reads support this
model, but they are not the main monitor.
SessionManager.rosterView already joins session state and live status. AgentControl.list()
already gets the whole live fleet in one call. M4 makes that join persistent and adds timers,
accepted-turn state, and incident state.
2.1 Real traces behind the design
The first trace was an architect that stopped making progress:
profile=opus role=architect state=busy liveStatus=done
The session moved from DONE to BUSY for turn 2. Eighteen minutes later, the session still said
BUSY, herdr still said DONE, the async task still said PENDING, and no completion fallback had
run. This is TURN_BOUNDARY_LOST, not a general slow-turn guess.
The second trace had two async sends to the same pane, one second apart. The pane was then stopped.
One ticket became failed. The other stayed pending - worker unknown. Current
MessageService.abandon resolves only Rendezvous.currentWaiter(target), while async tasks live in
a separate ticket map. CB-568 is intended to fix that bug. M4 still keeps an independent
post-teardown invariant so a later regression becomes DELEGATION_ORPHANED.
2.2 Corrections made during design
The first state table missed BUSY in fleetd plus IDLE or DONE in herdr. It would have found
the real trace only through a late, weak stall timer. The final model adds
TURN_BOUNDARY_LOST as a strong disagreement state.
The first notification design also required a webhook before health.enabled could turn on. That
removed useful local detection to avoid a narrower human-notification gap. The final design splits
detection from notification. Missing human escalation is shown as partial coverage instead of
disabling health.
3. Classification precedence
Evidence is applied in this order. A lower rule cannot hide a higher one.
- Control link: failed fleet list plus failed ping becomes
CONTROL_LINK_DOWN. - Definitive target loss:
_not_foundbecomesGONEorLEAD_UNREACHABLEwhen the control link is healthy. - Startup and teardown invariants: readiness expiry becomes
NEVER_READY; surviving tasks after teardown becomeDELEGATION_ORPHANED. - Bridge/live disagreement:
BUSYplus stable rawIDLEorDONEbecomesTURN_BOUNDARY_LOST. - Known screen evidence: a tested fatal signature becomes
ERROR_ON_SCREEN. - Timed suspicion: unchanged sparse pane probes may become
STALL_SUSPECTED. - Communication quality: completion fallback becomes
MUTE; an old inbox entry becomesREPLY_STRANDED. - Normal mode:
STARTING,IDLE,WORKING,WORK_PENDING, orBLOCKED_AMBIGUOUS.
The member flow in Figure 1 shows lifecycle states and the main fault exits. Fault states are reported beside the session FSM; most are not new FSM values.
flowchart TD
Registered["Member registered"] --> Starting["STARTING"]
Starting -->|"MCP presence"| Idle["IDLE"]
Starting -->|"Readiness grace expires"| NeverReady["NEVER_READY"]
Idle -->|"Accepted delivery"| Working["WORKING"]
Working -->|"Trusted turn boundary"| Idle
Working -->|"Bridge BUSY and herdr IDLE or DONE"| Lost["TURN_BOUNDARY_LOST"]
Working -->|"Known fatal screen"| Error["ERROR_ON_SCREEN"]
Working -->|"Long age and unchanged sparse probes"| Stall["STALL_SUSPECTED"]
Working -->|"Target not found"| Gone["GONE"]
Idle -->|"Inbox or queued delivery exists"| Pending["WORK_PENDING"]
Pending -->|"Delivery or collection finishes"| Idle
Idle -->|"Raw BLOCKED with an open turn"| Blocked["BLOCKED_AMBIGUOUS"]
Lost -->|"Strict guarded repair"| Repaired["DONE with reconciled completion"]
Lost -->|"Repair refused"| LeadDecision["Lead decision required"]
Figure 1. The member lifecycle and the main health exits. Pane-based states never authorise an automatic retry of the task.
4. State model
4.1 Normal and transitional member states
| State | Exact evidence | Meaning and certainty |
|---|---|---|
STARTING |
Session is SPAWNING; MCP presence is absent |
Normal inside the startup grace. MCP contact is the readiness signal. |
IDLE |
Session is READY or DONE; live status is IDLE or DONE; no open turn or inbox item exists |
Normal. Idle is not a fault. |
WORKING |
Session is BUSY; raw live status is WORKING; the accepted turn is open |
Certain that herdr sees work. It does not prove useful progress. |
WORK_PENDING |
Queued delivery or inbox content exists while the target is injectable | Transitional. Existing injector or push logic should move it. |
BLOCKED_AMBIGUOUS |
An open turn exists and raw live status is BLOCKED |
The bridge cannot tell whether this is permission, input, or a settled screen. |
Idle may drive configured resource cleanup. It never opens an incident and never pages a person.
4.2 Member fault and quality states
| State | Exact evidence | Certainty and action |
|---|---|---|
NEVER_READY |
SPAWNING, no MCP presence, and an accepted delivery waits through the existing readiness grace |
Delivery never became possible. The exact cause is unknown. Fail the send, stop the process, and preserve a provisioned worktree. |
GONE |
Per-target herdr call returns _not_found while fleet list or ping works |
Certain target loss. Fail all target work. Do not replay it. |
TURN_BOUNDARY_LOST |
Same session turn stays BUSY; same accepted task stays open; two raw snapshots show IDLE or DONE |
Strong disagreement. Strict reconciliation may repair it. |
ERROR_ON_SCREEN |
Suspicious non-working state survives grace; detection matches a tested adapter-specific fatal signature |
Certain only for the matched signature. A bare word such as Exception is not enough. |
STALL_SUSPECTED |
Open turn is older than the configured threshold; two normalised recent_unwrapped digests are unchanged; no boundary or reply occurs |
Not certain. A long valid API call can look the same. Lead decides. |
MUTE |
Turn resolves through completion fallback instead of fleet_reply |
Certain that no structured reply won. It does not prove an MCP failure. A single event is a metric, not an incident. |
REPLY_STRANDED |
Typed reply or health message remains after owning-lead push reaches its cap | Collection failed. This does not explain whether the lead is busy, dead, or ignoring the nudge. |
DELEGATION_ORPHANED |
Target is gone, failed, or released, but one or more tasks remain PENDING after reconciliation grace |
Certain bridge invariant failure. This is not an inbox-drain fault. |
WORK_PRODUCT_AT_RISK |
Provisioned branch has commits after its recorded base; member is DONE, FAILED, or preserved after release; no turn or inbox item remains; long-idle threshold passed |
A warning, not proof of loss. Work may already have an open pull request or a squash merge. |
MUTE opens an incident only after a small fixed rate threshold for one target or profile, or when
it appears with another fault.
WORK_PRODUCT_AT_RISK must not become WORK_PRODUCT_UNCOLLECTED. The bridge does not know pull
request or merge state. If committed work appears with REPLY_STRANDED or
DELEGATION_ORPHANED, the existing incident gains committedWorkAtRisk: true.
4.3 Control-link state
| State | Exact evidence | Certainty and action |
|---|---|---|
CONTROL_LINK_DOWN |
Two full-fleet agent.list calls fail across the grace, and herdr ping also fails |
Certain for the fleetd-to-herdr link. Retry calls, record the incident, and use human escalation if no lead can be reached. |
A failed fleet list alone is not a dead-member claim. A single _not_found with a healthy global
link is a target fault, not a control-link fault.
4.4 Lead states
| State | Exact evidence | Meaning and action |
|---|---|---|
LEAD_IDLE |
Expected lead is present with raw injectable status; no actionable state waits | Normal. Existing heartbeat may run under its own policy. |
LEAD_WORKING |
Expected lead is present with raw WORKING; stall threshold is not met |
Reachable and busy. Never inject into the live turn. |
LEAD_STATUS_UNKNOWN |
Expected lead is present with raw UNKNOWN |
Neither dead nor a healthy routing target. Retain evidence and retry. |
LEAD_UNREACHABLE |
Expected lead is absent from two successful live-agent snapshots while ping works, or targeted lookup returns _not_found with a healthy control link |
Route to a healthy peer. If none exists, use human escalation. |
LEAD_UNRESPONSIVE |
Actionable state waits; lead stays injectable; bounded nudges exhaust; inbox remains uncollected | Route to a healthy peer or a person. |
LEAD_STALL_SUSPECTED |
Lead stays WORKING past threshold; two sparse pane probes show no progress |
Not certain. Never kill or restart automatically. Route to peer or person. |
The monitor retains the lead name and terminal, last successful sighting, raw status and age, consecutive list absences, targeted errors, pane-probe facts, pending incident age, and nudge outcomes. Current heartbeat and push loops discard much of this history.
Expected lead identity comes from the same supplier used by CallerResolver. It is not liveness
evidence. LeadTabScanner keeps cached identity after a failed scan, so the health monitor compares
that identity with a fresh successful agent list. A dynamic identity also survives a two-successful-
snapshot retirement grace. This stops a dead lead from escaping health by disappearing from one map.
4.5 Evidence limits
M4 cannot tell these cases apart with current evidence:
- A valid long call and a hung call may have the same status and pane digest.
BLOCKEDdoes not explain which input is needed.- An idle prompt after failure may look like an idle prompt after success.
- A missing structured reply does not prove a broken MCP connection.
- An undrained inbox does not explain why the lead did not collect it.
- Arbitrary pane text cannot safely classify arbitrary exceptions.
- A branch ahead of its base does not prove that work was not collected.
Logs are outputs, not classifier inputs. The monitor never parses its own logs.
5. Automatic action and lead action
5.1 Actions the bridge may take
The bridge may:
- retry transient herdr status, list, ping, and pane-read failures with bounded backoff;
- re-submit Enter after the existing paste/submit race;
- fail queued delivery after
NEVER_READY; - stop a never-ready process while preserving its provisioned worktree;
- fail all queued, accepted, and async tasks for a gone or released target;
- reconcile one lost boundary when every strict gate in Section 8 passes;
- hold typed messages, nudge the owning lead, and stop at the configured cap;
- use the existing bounded idle-lead heartbeat;
- deduplicate, route, update, and resolve incidents.
These actions do not choose new work and do not replay old work.
5.2 Decisions reserved for the lead
Only the lead may:
- stop or continue
BLOCKED_AMBIGUOUS; - stop, inspect, or wait on
ERROR_ON_SCREEN; - kill or continue
STALL_SUSPECTED; - spawn a replacement or reassign work;
- retry a delivered task;
- choose how to use partial work in a worktree;
- restart herdr or change network, model, credentials, backend, or configuration.
Reports include literal safe tool calls such as fleet_status(sessionId="..."),
fleet_poll(ticket="..."), fleet_list(), and optional fleet_stop(paneId="..."). A judgement
state never presents stop as the only action.
5.3 Release causes and worktree safety
| Release cause | Process action | Provisioned worktree |
|---|---|---|
SPAWN_ROLLBACK before registration or delivery |
Stop and clean up | Remove |
COMPLETED for READY or DONE without pending work, idle TTL, or successful context-cap completion |
Stop | Remove only if clean; preserve a dirty worktree (CB-576) |
NEVER_READY |
Stop | Preserve |
GONE |
Best-effort stop | Preserve |
TURN_FAILED or lead abort while BUSY or FAILED |
Stop | Preserve |
RELEASE_WITH_PENDING_TASKS |
Stop | Preserve |
SHUTDOWN |
Stop | Preserve |
Explicit stop is state-aware. SPAWNING, BUSY, FAILED, or any target with pending tasks uses a
preserving cause.
Before abnormal release removes the live session, M4 writes an atomic manifest under the worktree
root. It records session identity, owner, role, profile, repository, path, branch, base commit,
release cause, release time, state, and pending task ids. fleet_list.preservedWorktrees loads these
manifests after restart. Stop output and WARN logs also name the path and cause. M4 never
auto-deletes a preserved worktree.
6. Fleet health monitor
Add FleetHealthMonitor. Do not widen StatusPoller into a policy loop.
StatusPoller has a 250 ms delivery cadence and samples only injector targets with outstanding
work. Health needs all sessions, all leads, task state, inbox age, and global control evidence. One
loop cannot serve both cadences safely.
Build the monitor like LeadHeartbeatLoop:
- pure
decide(snapshot, priorState, now)logic; - a thin scheduler;
- an injected clock;
- edge-triggered state changes;
- no network work in the pure function;
- no sleeping in tests.
Each enabled fleet tick reads:
- one
AgentControl.list()result for the whole fleet; - one in-memory
SessionManager.roster()snapshot; - accepted turns and async task state;
- typed inbox depth, kind, and age;
- push and heartbeat outcomes;
- configured and discovered leads.
Existing failure paths publish structured evidence to the monitor. The monitor does not infer events from log text.
6.1 Pane budget
Healthy idle members, recent working members, and quiet leads cause no pane reads.
A pane is eligible only for a stable lost boundary, sustained BLOCKED or UNKNOWN, work older
than the suspect threshold, or one final evidence read for a confirmed fault when the pane exists.
Compiled brakes apply even if config asks for more:
- per-target pane cooldown is at least 60 seconds;
- working age before the first progress probe is at least 300 seconds;
- at most two pane reads occur in one fleet tick;
- targets rotate fairly;
- only a normalised digest and optional clipped local excerpt are stored;
- no pane excerpt leaves fleetd in a human webhook.
Use detection for tested screen signatures. Use normalised recent_unwrapped only for progress
comparison.
7. Typed inbox and routing
7.1 Semantic record
The typed inbox record carries:
schemaVersion
kind: reply | health
msgId, target, subjectTerminal, recipientLead
severity, state, evidence
createdAtEpochMillis, firstSeenEpochMillis, lastSeenEpochMillis
recoveryTried, suggestedToolCalls, content
A health message never calls Rendezvous.resolve. It cannot look like the member's task result.
Both inbox adapters share field preservation, first-id-wins dedup, FIFO among decoded messages, explicit ownership, ack, and release rules. The in-memory adapter stores typed records directly. It does not copy AMQP migration logic.
7.2 AMQP migration
The reader uses AMQP content_type, never body sniffing:
Legacy v0: text/plain
Typed family: application/vnd.ltms.fleet.inbox-message+json
A legacy reply may begin with {. It remains plain text because its media type is text/plain.
Legacy text becomes kind=reply with exact UTF-8 content and absent typed metadata.
Typed JSON has required integer schemaVersion: 1. Version 1 ignores unknown optional fields.
Missing required fields, invalid enums, malformed UTF-8 or JSON, and property/body identity mismatch
are invalid data.
An unknown schema version is not partly decoded. It remains unacknowledged on the original queue and
creates one operator-visible unsupported_version failure. A newer daemon may read it later.
Invalid known-format data is copied byte-for-byte to durable queue
agent.<target>.inbox.quarantine. A dedicated confirm-mode publisher confirms the persistent copy
before the original is acknowledged. A failed quarantine handoff leaves the original unacknowledged.
The raw body never enters logs.
Decode failure creates a redacted WARN, metric, fleet_list summary, and routed health incident.
One bad entry never escapes the consumer callback and never stops later valid messages.
Safe downgrade is not supported. The previous build ignores content_type and would show typed JSON
as ordinary reply text. If drained, it would acknowledge the message and lose typed meaning. Typed
queues must be drained or preserved before an old jar runs.
The existing contract suite uses RabbitMQ. Production uses LavinMQ. The migration and lead-key ownership cases must run once against production LavinMQ before release, or the release must state that LavinMQ was not checked.
7.3 Member routing
A member incident first goes to the exact lead that owns its accepted delegation.
PrimaryRegistry needs a no-fallback delegatingLeadFor(memberTarget) query. Health routing must not
use the old singular-primary fallback when several leads exist.
Publish the incident under the affected member target. Trigger the existing bounded push route. The push waits until the owning lead is injectable, so it does not interrupt a live lead turn.
7.4 Peer lead routing
Figure 2 shows the route from incident to lead, peer, or person.
flowchart TD
Incident["Open incident"] --> Member{"Member incident?"}
Member -->|"yes"| Known{"Exact delegation owner known?"}
Known -->|"no"| Sink{"Human webhook enabled and healthy?"}
Known -->|"yes"| Owner{"Owner lead healthy?"}
Owner -->|"yes"| OwnerInbox["Publish to owner lead path"]
Owner -->|"no"| PeerSet["Build healthy peer candidate set"]
Member -->|"no, lead incident"| PeerSet
PeerSet --> Peer{"Healthy peer exists?"}
Peer -->|"yes"| Select["Choose fewest assigned incidents<br/>then stable name and terminal id"]
Select --> PeerInbox["Publish to peer lead inbox<br/>and status-gated push"]
Peer -->|"no"| Sink
Sink -->|"yes"| Webhook["Send classified outbound incident"]
Sink -->|"no"| Passive["Keep incident open<br/>show partial coverage on local surfaces"]
Figure 2. Routing keeps delegation ownership separate from temporary peer fallback.
Peer candidates exclude the incident subject, failed owner, absent leads, raw-unknown leads, and
leads with an open unhealthy state. A reachable WORKING peer may be selected; its push waits for an
injectable window.
Choose the candidate with the fewest assigned foreign incidents. Break ties by stable lead name, then terminal id. Pin the recipient. Reassign only if that peer becomes unhealthy or retires. A routing generation marks a reassignment, and old pending assignments become superseded.
fleet_list lead rows show health, health age, assigned foreign incident count, and a bounded list
of incident id, subject, state, severity, age, and routing generation. The top-level view also shows
owner, recipient, and routing reason.
A peer incident is published under the recipient lead's inbox key, not the failed subject's key. Its
status-gated nudge names the failed lead and gives the exact
fleet_poll(target="<recipient-terminal>") call.
7.5 Lead inbox ownership
Add LeadInboxRegistry, driven by the same expected-lead supplier as CallerResolver.
It calls replyInbox.own(leadTerminal) at startup for configured leads, after successful discovery,
after config adds a lead, and before publication. Ownership is not an authorization side effect.
A missing lead keeps its key owned. Release happens only after confirmed retirement, all incidents are reassigned or resolved, typed health messages move or ack, and the queue is empty. Own a replacement terminal before moving messages from the old key. Never release a non-empty in-memory lead key, because in-memory release clears local data.
7.6 Single-lead deployment
One lead and no peer is a normal mode, not an edge case.
An idle, reachable lead may receive the existing bounded nudge. An unreachable or stalled sole lead has no safe in-loop recovery. The bridge must not restart or replace it. A new lead would not have the failed lead's plan or context, and an uncertain relaunch could create two orchestrators.
With no webhook, only fleet_list, /healthz, metrics, WARN logs, and the incident journal remain.
These are passive surfaces. They are not a human notification.
8. Lost-boundary reconciliation
This is the only M4 path that reconstructs a result. It must prefer a visible stall over a fabricated reply.
8.1 Why normal completion rules are not enough
Current CompletionResolver.resolve has two fail-open rules. It resolves when the delivery baseline
is missing. It also resolves an empty completion when the pane read fails. Those choices are valid
after a trusted WORKING -> IDLE boundary because the bridge knows the turn ran. They are unsafe
when health only guesses that a boundary was lost.
M4 gives each accepted send an internal TurnToken. It ties target, exact waiter, session turn,
delivery baseline, and task outcome together.
8.2 Delivery baseline
Capture the baseline immediately after prompt send and before the delivery future completes. Store:
TurnToken
exact waiter identity
capture time and pane source
normalised assistant block clipped to MAX_SCRAPE_CHARS
whether a supported assistant marker was recognised
capture result: PRESENT | READ_FAILED | UNRECOGNISED
A failed or missing baseline never authorises repair. A late baseline is not valid evidence. After a daemon restart, the old waiter, task, token, and baseline are gone, so the old turn cannot be repaired.
Automatic repair is enabled only for agent kinds with tested assistant-block fixtures. Current
extraction is Claude Code-specific and falls back to arbitrary raw text without ⏺. That raw fallback
cannot authorise repair. OpenCode repair stays disabled until live pane fixtures exist.
8.3 Strict gates and resolver result
Figure 3 shows the repair gates. Any failed gate keeps the waiter unchanged.
The two raw snapshots must describe the same TurnToken and session turn. No WORKING,
BLOCKED, UNKNOWN, missing-agent, reply, failure, or new-delivery observation may occur between
them.
flowchart TD
Candidate["TURN_BOUNDARY_LOST candidate"] --> Stable{"Same TurnToken and BUSY turn<br/>across two raw IDLE or DONE snapshots?"}
Stable -->|"no"| Resnapshot["Take a fresh snapshot"]
Stable -->|"yes"| Waiter{"Exact captured waiter<br/>still open by identity?"}
Waiter -->|"no"| Stale["STALE_TURN or ALREADY_RESOLVED"]
Waiter -->|"yes"| Baseline{"Successful recognised<br/>delivery baseline exists?"}
Baseline -->|"no"| Refuse["Refuse repair<br/>leave ticket pending"]
Baseline -->|"yes"| Read{"Fresh pane read succeeds?"}
Read -->|"no"| Refuse
Read -->|"yes"| Output{"Recognised non-blank assistant block<br/>differs from clipped baseline?"}
Output -->|"no"| Refuse
Output -->|"yes"| Resolve["Shared CompletionResolver guard core<br/>resolves exact waiter"]
Resolve -->|"won race"| Repaired["RECONCILED_COMPLETION<br/>same turn becomes DONE"]
Resolve -->|"lost race"| Resnapshot
Figure 3. Repair needs stronger evidence than a normal observed turn boundary.
Refactor the current resolver into one guard core with two policies:
resolveCaptured(target, inFlight, OBSERVED_BOUNDARY)
resolveCaptured(target, inFlight, LOST_BOUNDARY_REPAIR)
The health monitor calls only:
CompletionResolver.reconcileLostBoundary(target, expectedTurnToken)
It returns REPAIRED, ALREADY_RESOLVED, REFUSED_NO_CAPTURE, REFUSED_NO_BASELINE,
REFUSED_UNREADABLE, REFUSED_UNCHANGED, REFUSED_AMBIGUOUS_OUTPUT, STALE_TURN, or
RACE_LOST.
Only REPAIRED and same-turn ALREADY_RESOLVED may move that turn from BUSY to DONE. A
per-target reconciliation gate stops a queued second send from being accepted between waiter
resolution and the FSM transition.
8.4 Lead-visible marker and refusal
A repaired result uses distinct RECONCILED_COMPLETION values in Rendezvous, MessageService,
task poll source, and metrics. The lead sees:
[repaired completion - fleetd detected a lost turn boundary. The member did not call
fleet_reply; pane-derived text follows and may be partial]
Clipped text also keeps the existing clipped-tail marker.
A refused repair leaves TURN_BOUNDARY_LOST open and the ticket pending. The report states that no
reply was reconstructed and no task was replayed. UNCHANGED, UNREADABLE, and
AMBIGUOUS_OUTPUT get at most one delayed retry for the same token. Missing capture or baseline gets
no retry. After two refused scrapes, automatic repair stops for that token.
8.5 Target-wide teardown invariant
CB-568 owns the multi-ticket cancellation mechanism. M4 routes every terminal cause through that one idempotent operation and checks this independent invariant after teardown:
- no injector entry exists for the target;
- no accepted turn or completion record exists;
- no rendezvous waiter or ask exists;
- every async task is terminal or was already terminal;
- no thread waiting for the target send lock can later accept it;
- new sends fail immediately;
- each old task has one terminal outcome and one metric count.
A violation becomes DELEGATION_ORPHANED. The monitor may call the same idempotent target-wide
failure operation once. It never recreates the task.
9. Capacity and utilisation
Capacity is a view, not a health state.
fleet_list adds one block per profile:
profile, maxLoad, live, free, reclaimable
For an unlimited profile, maxLoad and free are null. free is
max(0, maxLoad - live) for a capped profile.
The view must use the exact live-count function used by placement. A second calculation could show a
free slot that placement then refuses. Member rows add idleForSeconds only when state is READY or
DONE, no accepted turn exists, and the inbox is empty. reclaimable means only that the member
holds capacity without open bridge work.
The existing idle-lead nudge gains a bounded capacity summary. It lists per-profile live, cap, free,
and reclaimable counts, plus at most three long-idle members. Capacity does not make
FleetState.hasPending() true. A changed capacity fingerprint may re-arm one capped heartbeat
sequence. The fingerprint excludes changing idle durations, so a static idle fleet cannot reset the
cap forever. Reply-push stand-down remains first.
The bridge must never:
- spawn a member because a slot is free;
- generate a task or acceptance criteria;
- move queued work to another member or profile;
- treat a free slot or idle member as an incident;
- stop an idle member only to improve utilisation.
The bridge knows capacity facts but has no work list. Only the lead has the plan, task context, side-effect history, and acceptance criteria.
Capacity calculation is in memory and adds no pane reads. Work-product checks run on a terminal session edge, not every fleet tick.
This capacity design adds no automatic stop. The accepted NEVER_READY cleanup can still stop a
very slow startup after the existing grace, which is a known risk. Free capacity and long idle time
never trigger that path.
10. Human escalation and notification
10.1 Escalation rule
Notify a person only when no healthy lead can act:
CONTROL_LINK_DOWNsurvives grace;- a lead is unhealthy and no healthy peer can receive the incident;
- a member incident has no known owning lead;
- the only owning lead becomes unreachable, unresponsive, or stalled;
- incident publication or routing itself fails.
Do not page a person for a member fault while a healthy owning lead exists. An uncollected member incident feeds lead-health evidence. If the lead then becomes unhealthy, peer or human routing starts.
10.2 Detection and notification switches
health.enabled controls detection and bridge-local reporting. It does not require a webhook.
health.notifications.mode is disabled or webhook. Disabled is valid and is the default.
Webhook mode requires a resolved environment variable. Turning notification off stops outbound
attempts but keeps incidents. Turning it back on resumes still-open human incidents.
Without a sink, fleet_list.healthCoverage states that human escalation is unavailable. /healthz
keeps its existing HTTP liveness result and adds a nested fleetHealth.status=partial component.
Metrics and one startup or reload WARN expose the same limit.
10.3 Incident and delivery deduplication
One open incident uses this key:
(scope, subjectStableId, state, causeFingerprint)
The cause fingerprint includes stable error codes, dependency names, signature ids, or invariant names. It excludes times, ages, retry counts, pane text, and changing digests. A later recurrence after resolution gets a new generation and incident id.
Each outbound event uses:
Idempotency-Key = hash(incidentId, eventType, eventRevision)
Event types are open, severity_changed, reminder, and resolved. Transport retries keep the
same key.
An atomic owner-only journal beside the active config stores open incidents, routing, delivered revisions, retry state, and resolution state. It stores no pane or task content. Journal failure does not stop detection, but notification coverage becomes degraded.
10.4 Retry, reminder, and resolve
Send the first event immediately. Retry network errors, timeouts, HTTP 408, HTTP 429, and HTTP 5xx with full-jitter exponential backoff:
base: 5 seconds
factor: 3
maximum delay: 15 minutes
one outstanding attempt per event
Respect Retry-After up to 15 minutes. Other HTTP 4xx responses are permanent for that event until
config changes or a person requests replay.
Transport retry is not an incident reminder. humanRepeatSeconds creates a new reminder revision
for an unresolved critical incident after the last successful human event. Disabled mode does not
build an unbounded reminder queue.
Send resolved only if at least one human event for that incident was delivered. If an incident
resolves before its first successful delivery, cancel the pending open event and record local
resolution.
10.5 Outbound payload boundary
An outbound payload may contain incident id and event type, severity, state, scope, stable bridge ids, role or profile, times, duration, structured evidence type and counts, recovery attempted, routing reason, coverage, and safe tool calls.
It must never contain:
- raw pane text, pane excerpts, or pane digests;
- task briefs, prompts, or member reply content;
- source files, diffs, or worktree file content;
- worktree paths;
- environment values, tokens, credentials, headers, or webhook URL;
- raw exception messages or stack traces;
- arbitrary model output.
The sink response body is ignored. A webhook cannot direct recovery. n8n remains outbound-only.
10.6 Metrics
M4 adds bounded-label series:
fleet_health_incidents{scope,state,severity}
fleet_health_incidents_total{event}
fleet_health_notifications_total{event,outcome}
fleet_health_notification_queue_depth
fleet_health_notification_last_success_seconds
fleet_health_notification_capability{mode,status}
fleet_lead_health{lead,state}
fleet_lead_assigned_incidents{lead}
Metric labels never include terminal ids, incident ids, URLs, or error text.
11. Configuration
The optional health: block is absent or disabled by default. The dormant monitor scheduler does no
herdr or pane work while disabled. Every listed key is hot because the monitor reads ConfigRef on
each tick or notification.
| Key | Class | Default and hard bound | Purpose |
|---|---|---|---|
health.enabled |
Hot | false |
Enable detection and bridge-local reporting. |
health.snapshotIntervalSeconds |
Hot | default 30, minimum 15 | Whole-fleet comparison cadence. |
health.workingSuspectAfterSeconds |
Hot | default 600, minimum 300 | Age before working-pane probes. |
health.paneProbeIntervalSeconds |
Hot | default 60, minimum 60 | Per-target pane cooldown. |
health.leadUnresponsiveAfterSeconds |
Hot | default 300, minimum 120 | Delay after exhausted actionable nudges before lead fault. |
health.humanRepeatSeconds |
Hot | default 3600, minimum 900 | Minimum repeat period for one open human incident. |
health.capacityLongIdleAfterSeconds |
Hot | default 900, minimum 300 | Long-idle threshold for capacity summaries. |
health.includePaneExcerpt |
Hot | false |
Allow a clipped excerpt in local lead reports only. Human payloads still exclude it. |
health.notifications.mode |
Hot | disabled |
Select disabled or webhook. |
health.notifications.webhookUrlEnv |
Hot | required in webhook mode | Name of the environment variable that holds the sink URL. |
health.notifications.requestTimeoutMs |
Hot | default 10000, range 1000-30000 | Whole webhook request limit. |
Two consecutive snapshots are compiled floors for lost boundary, lead disappearance, and control link failure. The two-pane-reads-per-tick limit is also compiled and cannot be weakened by config.
12. Delivery units and acceptance
Unit 1 - Evidence model and fleet snapshot
Scope: health state model, fleet join, clocks, evidence retention, and pane budget.
Acceptance criteria:
- One
agent.listcall covers one enabled fleet tick. - Pure decision tests cover every state and every evidence limit in Section 4.
BUSYplus stable rawDONEopensTURN_BOUNDARY_LOSTafter two snapshots.- Healthy fleet snapshots perform zero pane reads.
- Pane cooldown, two-read fleet budget, and fair rotation cannot be disabled by config.
- Logs are outputs only; no log parsing exists.
- Fleet snapshots expose the same profile live-count calculation that placement uses.
- Capacity rows report cap, live, free, and reclaimable values without opening incidents.
Unit 2 - Lost boundary and task reconciliation
Scope: accepted-turn identity, guarded repair, target-wide teardown, release causes, and preserved worktree discovery.
Acceptance criteria:
-
Every accepted send receives a stable
TurnTokentied to target, exact waiter, and delivery baseline.Corrected during implementation (2026-08-15). This criterion first also required the session turn number and the task outcome. That is not implementable at this layer, and the implementer refused it three times rather than fabricate a value — correctly. The reason is an ordering fact that is invisible from any single class:
MessageServiceowns acceptance and holds the waiter and the asyncTask, but it learns nothing about delivery, because the delivery event goes toCompletionResolverthroughTurnListener.onDelivered. AndCompletionResolver.onDeliveredruns beforeSessionManager.onDelivered, so the session turn number does not exist yet at the only point where the token could capture it.Two ways out were rejected. A shared registry keyed by target reintroduces exactly the "whichever send happens to be waiting" ambiguity the token exists to remove — the same weak claim
Rendezvous.currentWaiterwarns about. Injecting a turn counter intoMessageServiceadds a required cross-layer dependency to populate a field that nothing in this slice reads, which is speculative coupling across a boundary already shown to be fragile.So the token identifies the accepted send, and
SessionManagerkeeps verifying its own delivery separately. Repair (criterion 2) does need the session turn; binding it means resolving that acceptance-versus-delivery ordering first, and that work belongs to the repair unit, not here. The token record carries a comment saying the field is deliberately absent. -
Repair requires the same
BUSYtoken, two rawIDLEorDONEsnapshots, no conflicting observation, exact open waiter, successful baseline, and new recognised assistant output. -
Missing, failed, late, or post-restart baseline never authorises repair.
-
Repair is enabled only for agent kinds with tested assistant-block extraction. Raw-text fallback without a recognised marker refuses repair.
-
Normal completion and repair use one resolver guard core. Waiter, scrape, clipping, unchanged, and exact-turn guards are not duplicated.
-
reconcileLostBoundaryreturns every typed result named in Section 8.3. -
Only
REPAIREDand same-turnALREADY_RESOLVEDmay move the same turn toDONE. -
A per-target reconciliation gate blocks a queued second send during repair and FSM update.
-
Repaired completion has distinct rendezvous kind, message outcome, poll source, lead marker, and metric. Clipping keeps its extra marker.
-
Unchanged, unreadable, or ambiguous evidence gets at most one delayed retry. Missing capture or baseline gets none.
-
Refusal leaves the ticket pending and tells the lead that no result was rebuilt or replayed.
-
Release, gone, never-ready, and abnormal stop use CB-568's one idempotent target-wide failure operation.
-
The post-teardown invariant in Section 8.5 is tested independently of CB-568 internals.
-
A violated teardown invariant creates
DELEGATION_ORPHANEDand retries only the idempotent failure operation. -
SPAWN_ROLLBACKand normalCOMPLETEDremove worktrees. Abnormal and shutdown causes preserve them. -
Explicit stop is state-aware. Any pending task or non-terminal state preserves the worktree.
-
Atomic preserved-worktree manifests reload after restart and appear in lead-only
fleet_list.preservedWorktrees. -
Manifest failure preserves the worktree and opens an operator-visible health failure.
-
Provision records the base commit. Terminal, long-idle worktrees report
WORK_PRODUCT_AT_RISKonly under the evidence in Section 4.2 and never auto-delete work. -
No path replays a delivered task, rebuilds its brief, or retargets it, even when prompt text is available.
-
Tests cover both real traces, all repair refusals, clipping, explicit-reply and next-turn races, restart without capture, concurrent send and release, and preserved discovery after restart.
Unit 2 - what has landed so far
Checked against main at e09cac6 on 2026-08-15. Unit 2 was written as one block, but parts of it
have since been built by separate CB tickets. Read this before planning the rest, or that work gets
done twice.
The check was a symbol survey of bridged/src/main/java plus the merge history. It tells you whether
the machinery exists at all. It is not a line-by-line audit of whether each criterion is fully
met, and I did not run one.
| Criterion | Marker searched for | Found in main source | Reading |
|---|---|---|---|
| 1 | TurnToken |
8 files | Done — unit 2a, merged as fec284e. Criterion 1 was corrected first; see the note under it. |
| 2-5, 9 | REPAIRED |
0 files | Not started. The whole guarded-repair path is absent. |
| 6, 7, 10 | reconcileLostBoundary |
0 files | Not started. |
| 12 | CB-568 failure operation | via CB-580 | Partial. CB-580 (0af902e) routes GONE and NEVER_READY into the one idempotent target-wide failure. I did not check that release and abnormal stop go through the same call. |
| 14 | DELEGATION_ORPHANED |
3 files | Partial. The health state exists. The teardown-invariant check that creates it, and the retry rule, do not. |
| 15 | SPAWN_ROLLBACK |
0 files | Contradicted — see below. |
| 16 | — | — | Partial at best. CB-576 made release preserve a dirty worktree; whether explicit stop is state-aware is not checked. |
| 17, 18 | preservedWorktrees |
0 files | Not started. No manifest, and no lead-only fleet_list field. |
| 19 | WORK_PRODUCT_AT_RISK |
0 files | Not started. |
Criterion 15 no longer matches the code, and the code is right. It says "normal COMPLETED
remove worktrees". Since CB-576 (500bfa2) that is false on purpose: a COMPLETED release now
preserves the worktree when it still holds uncommitted work, because deleting it destroys work
nobody can get back. CB-576 was filed after exactly that loss. CB-581 goes further — if the
dirty-check itself fails, the worktree is preserved rather than removed, since "we could not tell"
must not be treated as "it is clean".
So criterion 15 should be rewritten as: SPAWN_ROLLBACK and a COMPLETED release with a clean
worktree remove it; abnormal causes, shutdown, a dirty worktree, and a failed dirty-check all
preserve it. SPAWN_ROLLBACK itself does not exist yet.
Unit 3 - Typed inbox and member routing
Scope: semantic record, AMQP migration, both adapters, member routing, polling, and member health in
fleet_list.
Acceptance criteria:
- AMQP selects legacy or typed decoding only from
content_type; it never sniffs the body. - Persistent
text/plainfrom the old build becomeskind=replywith exact UTF-8 content, including content beginning with{. - New entries use the vendor media type,
schemaVersion: 1, UTF-8, persistent delivery, and AMQP message ids. - Version 1 ignores unknown optional fields but rejects missing fields and identity mismatch.
- Unknown versions are not decoded or acked. They remain on the original queue and create one deduplicated failure.
- Invalid known data never escapes the callback, appears as a reply, or blocks later valid messages.
- Invalid data reaches durable per-target quarantine before original ack. Failed handoff leaves the original unacked.
- Decode failures create redacted WARN, metric,
fleet_listsummary, and routed incident without raw content. - Both adapters pass one semantic contract for fields, FIFO, dedup, ownership, ack, and release.
- Lead keys require explicit ownership. Publication never claims a queue.
- Unit codec tests cover legacy
{, Unicode, malformed UTF-8, typed round trip, additive fields, malformed JSON, missing fields, identity mismatch, media type, version, and dedup. - A live broker contract writes old wire data and reads it with the new adapter after reconnect.
- Live contract tests cover mixed entries, quarantine confirm-before-ack, unsupported redelivery, later progress past poison, property persistence, lead ownership, and ack removal.
- Safe downgrade is documented as unsupported.
- RabbitMQ contract tests pass with
mvn test -Pcontract. The same cases run once on production LavinMQ, or the release states that LavinMQ was not checked. - Member incidents route to the exact delegating lead and never resolve a task rendezvous.
fleet_listshows compact member health and capacity without pane content. MemberidleForSecondsis present only when no accepted turn or inbox item exists.
Unit 4 - Lead health and peer routing
Scope: lead evidence, exact ownership, peer selection, explicit-recipient push, and lead inbox lifecycle.
Acceptance criteria:
- Lead identity uses the
CallerResolversupplier. Liveness uses successful current agent data. - Two successful-list absences with healthy ping become
LEAD_UNREACHABLE; global link failure does not. - Raw
WORKING, rawUNKNOWN, first-seen time, failures, last success, and error class persist across ticks. - Heartbeat and push publish status and nudge outcomes before safe no-injection decisions.
- Dynamic lead identity survives a two-successful-snapshot retirement grace.
- Member incidents first use exact delegation ownership with no singular-primary fallback.
- Peer selection follows the exclusions, load rule, and stable tie break in Section 7.4.
- A selected working peer is not interrupted. Its push waits for an injectable window.
- Recipient assignment stays pinned. Reassignment increments generation and supersedes old pending assignment.
fleet_listshows bounded foreign assignments, recipient, reason, and generation without pane content.LeadInboxRegistryowns configured and discovered lead keys before publication.- Missing leads keep ownership. Retirement needs an empty queue and handled incidents.
- Replacement owns the new key before messages move. Non-empty in-memory keys are not released.
- Tests cover dead versus busy, unknown, global failure, stale scan cache, disappearance, one peer, several peers, reassignment, and no peer.
- Adapter tests cover lead ownership, restart re-ownership, retirement, and terminal replacement. LavinMQ is checked or named as unchecked.
- A sole unreachable or stalled lead is never restarted or replaced. Without a sink, only passive evidence remains and every coverage surface says so.
Unit 5 - Human sink, hot config, metrics, and operator coverage
Scope: generic webhook, config split, incident journal, retry, resolve, metrics, example config, and operator documentation.
Acceptance criteria:
health.enabledworks without a human sink.- Notification mode is hot, defaults to disabled, and supports disabled or webhook.
- Webhook mode requires a resolved environment value. Bad notification config does not disable an already valid detector.
- Mode changes keep open incidents. Re-enable resumes eligible incidents.
fleet_list,/healthz, metrics, and one WARN show partial coverage without a sink. HTTP liveness behavior stays unchanged.- One-lead, no-sink coverage states that lead failure has no active notification or recovery.
- Incident and outbound dedupe use the stable keys in Section 10.3.
- The owner-only local journal survives restart and contains no pane or task content.
- Journal failure keeps detection running but marks notification coverage degraded.
- Retry tests cover network failure, timeout, 408, 429,
Retry-After, 5xx, permanent 4xx, jitter, delay cap, config re-arm, and one outstanding attempt. - Reminders and transport retries remain separate. Disabled mode does not build an unbounded queue.
- Resolve sends only after an earlier human event succeeded. Resolve-before-delivery cancels stale open delivery.
- Metrics use bounded labels and exclude ids, URLs, and error text.
- Payload tests reject every content type forbidden in Section 10.5.
- Webhook response bodies are ignored and cannot direct recovery.
- Tests cover disabled mode, one lead without sink, open/update/reminder/resolve, restart, dedup, reassignment, disable/re-enable, and sink failure while local health continues.
fleetd.example.yamldocuments all hot keys and compiled floors.- The operator Features wiki is updated separately. The portable
CLAUDE.mdblock is checked and changed only if shipped tool or inbox semantics make it untrue. mvn clean installpasses.
13. Not checked and release gates
These limits are part of the design, not optional follow-up notes.
- OpenCode pane status and assistant markers were not checked. OpenCode lost-boundary repair is disabled until live fixtures exist.
- Permission-prompt status was not checked for Claude Code or OpenCode.
BLOCKEDremains ambiguous and has no automatic action. recent_unwrappedstability was not checked across all supported agent kinds. If normalisation is not stable,STALL_SUSPECTEDmust say its evidence is weaker.- The real
BUSY + DONEtrace was not replayed against live herdr. The design uses the observed production trace and current poller behavior. - CB-568 was not present when Unit 2 was designed. Unit 2 must inspect the landed API and keep its independent teardown invariant.
- Production LavinMQ was not checked. Existing durable-inbox contracts use RabbitMQ. Migration, quarantine, redelivery, lead ownership, and reassignment must run on LavinMQ before release or be recorded as unchecked.
- Live multi-lead routing was not checked. Peer choice and reassignment are design rules backed by fake-clock and adapter tests until a live exercise runs.
- A live sole-lead failure with a webhook was not checked. The no-peer path is a design result, not a tested recovery.
- No n8n, Slack, PagerDuty, or other receiver was checked. The webhook remains generic and outbound-only.
- Deployment supervisor behavior for nested
/healthzfields was not checked. HTTP liveness status stays unchanged to reduce this risk. - Incident-journal crash behavior was not checked because the journal does not exist yet. Unit 5 must test atomic replacement and restart recovery.
- Worktree merge state cannot be checked reliably without forge or explicit collection evidence.
WORK_PRODUCT_AT_RISKstays a warning.
14. Locked exclusions
M4 does not expose agent.read as a bridge tool. It does not add a workflow engine, inbound n8n
authority, automatic task assignment, task replay, automatic lead replacement, or automatic member
spawn for free capacity.
The bridge remains a message bus with evidence and bounded mechanical repair. The lead remains the place where judgement and work planning happen.