|
|
|
@@ -0,0 +1,917 @@
|
|
|
|
|
# M4 - Fleet health, recovery, routing, and capacity
|
|
|
|
|
|
|
|
|
|
**Status:** Design accepted on 2026-08-15. CB-573 part 1 has shipped the classification model and
|
|
|
|
|
the `bridge_list` capacity view; the remaining M4 units are not yet shipped.
|
|
|
|
|
**Scope:** Fleet evidence, safe mechanical repair, lead routing, capacity reporting, and optional
|
|
|
|
|
human notification.
|
|
|
|
|
**Grounded in:** `health/FleetHealth`, `health/PaneBudget`, `inject/StatusPoller`,
|
|
|
|
|
`inject/StatusRefiner`, `inject/CompletionResolver`, `inject/Injector`, `session/SessionManager`,
|
|
|
|
|
`msg/MessageService`, `msg/ReplyInbox`, `msg/ReplyPushLoop`, `msg/LeadHeartbeatLoop`,
|
|
|
|
|
`mcp/PrimaryRegistry`, and `herdr/AgentControl`.
|
|
|
|
|
|
|
|
|
|
## 1. Problem and decision boundary
|
|
|
|
|
|
|
|
|
|
The operator asked the bridge to detect idle agents, exceptions, stopped work, and broken
|
|
|
|
|
communication. The bridge may read an agent pane from time to time. It must notify a person when
|
|
|
|
|
the fleet cannot move forward.
|
|
|
|
|
|
|
|
|
|
The four operator terms are not four equal health states. `IDLE` is a normal mode. An exception is
|
|
|
|
|
sometimes visible only as pane text. Stopped work may look the same as slow work. Broken
|
|
|
|
|
communication can occur on several links.
|
|
|
|
|
|
|
|
|
|
M4 uses this boundary:
|
|
|
|
|
|
|
|
|
|
- The bridge detects facts and joins evidence.
|
|
|
|
|
- The bridge repairs only mechanical failures with no judgement.
|
|
|
|
|
- The lead decides whether to stop, retry, replace, or reassign a member.
|
|
|
|
|
- A human is notified only when no healthy lead can act.
|
|
|
|
|
- n8n may route an outbound incident. It never classifies state or chooses recovery.
|
|
|
|
|
|
|
|
|
|
An inbound n8n decider would need bridge authority. No narrow machine-decider role exists. Giving a
|
|
|
|
|
workflow engine lead authority is unsafe, while adding a new role is a separate authorization
|
|
|
|
|
design. An outbound sink needs no bridge role.
|
|
|
|
|
|
|
|
|
|
The bridge must never replay a delivered task. That task may already have changed files, pushed a
|
|
|
|
|
branch, opened a pull request, or changed external state. A replay can run those side effects twice.
|
|
|
|
|
This rule must remain true even if later code stores delivered prompt text.
|
|
|
|
|
|
|
|
|
|
## 2. Evidence model
|
|
|
|
|
|
|
|
|
|
A health state is mainly a comparison between two views:
|
|
|
|
|
|
|
|
|
|
- **herdr view:** current agents and raw live status from one `AgentControl.list()` call.
|
|
|
|
|
- **bridge view:** session FSM, MCP presence, accepted turns, tasks, inbox state, and lead ownership.
|
|
|
|
|
|
|
|
|
|
A strong fault often appears as a disagreement between those views. For example, `BUSY` in the
|
|
|
|
|
session FSM and `DONE` in herdr means the bridge missed a turn boundary. Pane reads support this
|
|
|
|
|
model, but they are not the main monitor.
|
|
|
|
|
|
|
|
|
|
`SessionManager.rosterView` already joins session state and live status. `AgentControl.list()`
|
|
|
|
|
already gets the whole live fleet in one call. M4 makes that join persistent and adds timers,
|
|
|
|
|
accepted-turn state, and incident state.
|
|
|
|
|
|
|
|
|
|
### 2.1 Real traces behind the design
|
|
|
|
|
|
|
|
|
|
The first trace was an architect that stopped making progress:
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
profile=opus role=architect state=busy liveStatus=done
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
The session moved from `DONE` to `BUSY` for turn 2. Eighteen minutes later, the session still said
|
|
|
|
|
`BUSY`, herdr still said `DONE`, the async task still said `PENDING`, and no completion fallback had
|
|
|
|
|
run. This is `TURN_BOUNDARY_LOST`, not a general slow-turn guess.
|
|
|
|
|
|
|
|
|
|
The second trace had two async sends to the same pane, one second apart. The pane was then stopped.
|
|
|
|
|
One ticket became failed. The other stayed `pending - worker unknown`. Current
|
|
|
|
|
`MessageService.abandon` resolves only `Rendezvous.currentWaiter(target)`, while async tasks live in
|
|
|
|
|
a separate ticket map. CB-568 is intended to fix that bug. M4 still keeps an independent
|
|
|
|
|
post-teardown invariant so a later regression becomes `DELEGATION_ORPHANED`.
|
|
|
|
|
|
|
|
|
|
### 2.2 Corrections made during design
|
|
|
|
|
|
|
|
|
|
The first state table missed `BUSY` in bridged plus `IDLE` or `DONE` in herdr. It would have found
|
|
|
|
|
the real trace only through a late, weak stall timer. The final model adds
|
|
|
|
|
`TURN_BOUNDARY_LOST` as a strong disagreement state.
|
|
|
|
|
|
|
|
|
|
The first notification design also required a webhook before `health.enabled` could turn on. That
|
|
|
|
|
removed useful local detection to avoid a narrower human-notification gap. The final design splits
|
|
|
|
|
detection from notification. Missing human escalation is shown as partial coverage instead of
|
|
|
|
|
disabling health.
|
|
|
|
|
|
|
|
|
|
## 3. Classification precedence
|
|
|
|
|
|
|
|
|
|
Evidence is applied in this order. A lower rule cannot hide a higher one.
|
|
|
|
|
|
|
|
|
|
1. **Control link:** failed fleet list plus failed ping becomes `CONTROL_LINK_DOWN`.
|
|
|
|
|
2. **Definitive target loss:** `_not_found` becomes `GONE` or `LEAD_UNREACHABLE` when the control
|
|
|
|
|
link is healthy.
|
|
|
|
|
3. **Startup and teardown invariants:** readiness expiry becomes `NEVER_READY`; surviving tasks
|
|
|
|
|
after teardown become `DELEGATION_ORPHANED`.
|
|
|
|
|
4. **Bridge/live disagreement:** `BUSY` plus stable raw `IDLE` or `DONE` becomes
|
|
|
|
|
`TURN_BOUNDARY_LOST`.
|
|
|
|
|
5. **Known screen evidence:** a tested fatal signature becomes `ERROR_ON_SCREEN`.
|
|
|
|
|
6. **Timed suspicion:** unchanged sparse pane probes may become `STALL_SUSPECTED`.
|
|
|
|
|
7. **Communication quality:** completion fallback becomes `MUTE`; an old inbox entry becomes
|
|
|
|
|
`REPLY_STRANDED`.
|
|
|
|
|
8. **Normal mode:** `STARTING`, `IDLE`, `WORKING`, `WORK_PENDING`, or `BLOCKED_AMBIGUOUS`.
|
|
|
|
|
|
|
|
|
|
The member flow in Figure 1 shows lifecycle states and the main fault exits. Fault states are
|
|
|
|
|
reported beside the session FSM; most are not new FSM values.
|
|
|
|
|
|
|
|
|
|
```mermaid
|
|
|
|
|
flowchart TD
|
|
|
|
|
Registered["Member registered"] --> Starting["STARTING"]
|
|
|
|
|
Starting -->|"MCP presence"| Idle["IDLE"]
|
|
|
|
|
Starting -->|"Readiness grace expires"| NeverReady["NEVER_READY"]
|
|
|
|
|
Idle -->|"Accepted delivery"| Working["WORKING"]
|
|
|
|
|
Working -->|"Trusted turn boundary"| Idle
|
|
|
|
|
Working -->|"Bridge BUSY and herdr IDLE or DONE"| Lost["TURN_BOUNDARY_LOST"]
|
|
|
|
|
Working -->|"Known fatal screen"| Error["ERROR_ON_SCREEN"]
|
|
|
|
|
Working -->|"Long age and unchanged sparse probes"| Stall["STALL_SUSPECTED"]
|
|
|
|
|
Working -->|"Target not found"| Gone["GONE"]
|
|
|
|
|
Idle -->|"Inbox or queued delivery exists"| Pending["WORK_PENDING"]
|
|
|
|
|
Pending -->|"Delivery or collection finishes"| Idle
|
|
|
|
|
Idle -->|"Raw BLOCKED with an open turn"| Blocked["BLOCKED_AMBIGUOUS"]
|
|
|
|
|
Lost -->|"Strict guarded repair"| Repaired["DONE with reconciled completion"]
|
|
|
|
|
Lost -->|"Repair refused"| LeadDecision["Lead decision required"]
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
*Figure 1. The member lifecycle and the main health exits. Pane-based states never authorise an
|
|
|
|
|
automatic retry of the task.*
|
|
|
|
|
|
|
|
|
|
## 4. State model
|
|
|
|
|
|
|
|
|
|
### 4.1 Normal and transitional member states
|
|
|
|
|
|
|
|
|
|
| State | Exact evidence | Meaning and certainty |
|
|
|
|
|
|---|---|---|
|
|
|
|
|
| `STARTING` | Session is `SPAWNING`; MCP presence is absent | Normal inside the startup grace. MCP contact is the readiness signal. |
|
|
|
|
|
| `IDLE` | Session is `READY` or `DONE`; live status is `IDLE` or `DONE`; no open turn or inbox item exists | Normal. Idle is not a fault. |
|
|
|
|
|
| `WORKING` | Session is `BUSY`; raw live status is `WORKING`; the accepted turn is open | Certain that herdr sees work. It does not prove useful progress. |
|
|
|
|
|
| `WORK_PENDING` | Queued delivery or inbox content exists while the target is injectable | Transitional. Existing injector or push logic should move it. |
|
|
|
|
|
| `BLOCKED_AMBIGUOUS` | An open turn exists and raw live status is `BLOCKED` | The bridge cannot tell whether this is permission, input, or a settled screen. |
|
|
|
|
|
|
|
|
|
|
Idle may drive configured resource cleanup. It never opens an incident and never pages a person.
|
|
|
|
|
|
|
|
|
|
### 4.2 Member fault and quality states
|
|
|
|
|
|
|
|
|
|
| State | Exact evidence | Certainty and action |
|
|
|
|
|
|---|---|---|
|
|
|
|
|
| `NEVER_READY` | `SPAWNING`, no MCP presence, and an accepted delivery waits through the existing readiness grace | Delivery never became possible. The exact cause is unknown. Fail the send, stop the process, and preserve a provisioned worktree. |
|
|
|
|
|
| `GONE` | Per-target herdr call returns `_not_found` while fleet list or ping works | Certain target loss. Fail all target work. Do not replay it. |
|
|
|
|
|
| `TURN_BOUNDARY_LOST` | Same session turn stays `BUSY`; same accepted task stays open; two raw snapshots show `IDLE` or `DONE` | Strong disagreement. Strict reconciliation may repair it. |
|
|
|
|
|
| `ERROR_ON_SCREEN` | Suspicious non-working state survives grace; `detection` matches a tested adapter-specific fatal signature | Certain only for the matched signature. A bare word such as `Exception` is not enough. |
|
|
|
|
|
| `STALL_SUSPECTED` | Open turn is older than the configured threshold; two normalised `recent_unwrapped` digests are unchanged; no boundary or reply occurs | Not certain. A long valid API call can look the same. Lead decides. |
|
|
|
|
|
| `MUTE` | Turn resolves through completion fallback instead of `bridge_reply` | Certain that no structured reply won. It does not prove an MCP failure. A single event is a metric, not an incident. |
|
|
|
|
|
| `REPLY_STRANDED` | Typed reply or health message remains after owning-lead push reaches its cap | Collection failed. This does not explain whether the lead is busy, dead, or ignoring the nudge. |
|
|
|
|
|
| `DELEGATION_ORPHANED` | Target is gone, failed, or released, but one or more tasks remain `PENDING` after reconciliation grace | Certain bridge invariant failure. This is not an inbox-drain fault. |
|
|
|
|
|
| `WORK_PRODUCT_AT_RISK` | Provisioned branch has commits after its recorded base; member is `DONE`, `FAILED`, or preserved after release; no turn or inbox item remains; long-idle threshold passed | A warning, not proof of loss. Work may already have an open pull request or a squash merge. |
|
|
|
|
|
|
|
|
|
|
`MUTE` opens an incident only after a small fixed rate threshold for one target or profile, or when
|
|
|
|
|
it appears with another fault.
|
|
|
|
|
|
|
|
|
|
`WORK_PRODUCT_AT_RISK` must not become `WORK_PRODUCT_UNCOLLECTED`. The bridge does not know pull
|
|
|
|
|
request or merge state. If committed work appears with `REPLY_STRANDED` or
|
|
|
|
|
`DELEGATION_ORPHANED`, the existing incident gains `committedWorkAtRisk: true`.
|
|
|
|
|
|
|
|
|
|
### 4.3 Control-link state
|
|
|
|
|
|
|
|
|
|
| State | Exact evidence | Certainty and action |
|
|
|
|
|
|---|---|---|
|
|
|
|
|
| `CONTROL_LINK_DOWN` | Two full-fleet `agent.list` calls fail across the grace, and herdr `ping` also fails | Certain for the bridged-to-herdr link. Retry calls, record the incident, and use human escalation if no lead can be reached. |
|
|
|
|
|
|
|
|
|
|
A failed fleet list alone is not a dead-member claim. A single `_not_found` with a healthy global
|
|
|
|
|
link is a target fault, not a control-link fault.
|
|
|
|
|
|
|
|
|
|
### 4.4 Lead states
|
|
|
|
|
|
|
|
|
|
| State | Exact evidence | Meaning and action |
|
|
|
|
|
|---|---|---|
|
|
|
|
|
| `LEAD_IDLE` | Expected lead is present with raw injectable status; no actionable state waits | Normal. Existing heartbeat may run under its own policy. |
|
|
|
|
|
| `LEAD_WORKING` | Expected lead is present with raw `WORKING`; stall threshold is not met | Reachable and busy. Never inject into the live turn. |
|
|
|
|
|
| `LEAD_STATUS_UNKNOWN` | Expected lead is present with raw `UNKNOWN` | Neither dead nor a healthy routing target. Retain evidence and retry. |
|
|
|
|
|
| `LEAD_UNREACHABLE` | Expected lead is absent from two successful live-agent snapshots while ping works, or targeted lookup returns `_not_found` with a healthy control link | Route to a healthy peer. If none exists, use human escalation. |
|
|
|
|
|
| `LEAD_UNRESPONSIVE` | Actionable state waits; lead stays injectable; bounded nudges exhaust; inbox remains uncollected | Route to a healthy peer or a person. |
|
|
|
|
|
| `LEAD_STALL_SUSPECTED` | Lead stays `WORKING` past threshold; two sparse pane probes show no progress | Not certain. Never kill or restart automatically. Route to peer or person. |
|
|
|
|
|
|
|
|
|
|
The monitor retains the lead name and terminal, last successful sighting, raw status and age,
|
|
|
|
|
consecutive list absences, targeted errors, pane-probe facts, pending incident age, and nudge
|
|
|
|
|
outcomes. Current heartbeat and push loops discard much of this history.
|
|
|
|
|
|
|
|
|
|
Expected lead identity comes from the same supplier used by `CallerResolver`. It is not liveness
|
|
|
|
|
evidence. `LeadTabScanner` keeps cached identity after a failed scan, so the health monitor compares
|
|
|
|
|
that identity with a fresh successful agent list. A dynamic identity also survives a two-successful-
|
|
|
|
|
snapshot retirement grace. This stops a dead lead from escaping health by disappearing from one map.
|
|
|
|
|
|
|
|
|
|
### 4.5 Evidence limits
|
|
|
|
|
|
|
|
|
|
M4 cannot tell these cases apart with current evidence:
|
|
|
|
|
|
|
|
|
|
- A valid long call and a hung call may have the same status and pane digest.
|
|
|
|
|
- `BLOCKED` does not explain which input is needed.
|
|
|
|
|
- An idle prompt after failure may look like an idle prompt after success.
|
|
|
|
|
- A missing structured reply does not prove a broken MCP connection.
|
|
|
|
|
- An undrained inbox does not explain why the lead did not collect it.
|
|
|
|
|
- Arbitrary pane text cannot safely classify arbitrary exceptions.
|
|
|
|
|
- A branch ahead of its base does not prove that work was not collected.
|
|
|
|
|
|
|
|
|
|
Logs are outputs, not classifier inputs. The monitor never parses its own logs.
|
|
|
|
|
|
|
|
|
|
## 5. Automatic action and lead action
|
|
|
|
|
|
|
|
|
|
### 5.1 Actions the bridge may take
|
|
|
|
|
|
|
|
|
|
The bridge may:
|
|
|
|
|
|
|
|
|
|
- retry transient herdr status, list, ping, and pane-read failures with bounded backoff;
|
|
|
|
|
- re-submit Enter after the existing paste/submit race;
|
|
|
|
|
- fail queued delivery after `NEVER_READY`;
|
|
|
|
|
- stop a never-ready process while preserving its provisioned worktree;
|
|
|
|
|
- fail all queued, accepted, and async tasks for a gone or released target;
|
|
|
|
|
- reconcile one lost boundary when every strict gate in Section 8 passes;
|
|
|
|
|
- hold typed messages, nudge the owning lead, and stop at the configured cap;
|
|
|
|
|
- use the existing bounded idle-lead heartbeat;
|
|
|
|
|
- deduplicate, route, update, and resolve incidents.
|
|
|
|
|
|
|
|
|
|
These actions do not choose new work and do not replay old work.
|
|
|
|
|
|
|
|
|
|
### 5.2 Decisions reserved for the lead
|
|
|
|
|
|
|
|
|
|
Only the lead may:
|
|
|
|
|
|
|
|
|
|
- stop or continue `BLOCKED_AMBIGUOUS`;
|
|
|
|
|
- stop, inspect, or wait on `ERROR_ON_SCREEN`;
|
|
|
|
|
- kill or continue `STALL_SUSPECTED`;
|
|
|
|
|
- spawn a replacement or reassign work;
|
|
|
|
|
- retry a delivered task;
|
|
|
|
|
- choose how to use partial work in a worktree;
|
|
|
|
|
- restart herdr or change network, model, credentials, backend, or configuration.
|
|
|
|
|
|
|
|
|
|
Reports include literal safe tool calls such as `bridge_status(sessionId="...")`,
|
|
|
|
|
`bridge_poll(ticket="...")`, `bridge_list()`, and optional `bridge_stop(paneId="...")`. A judgement
|
|
|
|
|
state never presents stop as the only action.
|
|
|
|
|
|
|
|
|
|
### 5.3 Release causes and worktree safety
|
|
|
|
|
|
|
|
|
|
| Release cause | Process action | Provisioned worktree |
|
|
|
|
|
|---|---|---|
|
|
|
|
|
| `SPAWN_ROLLBACK` before registration or delivery | Stop and clean up | Remove |
|
|
|
|
|
| `COMPLETED` for `READY` or `DONE` without pending work, idle TTL, or successful context-cap completion | Stop | Remove under completed policy |
|
|
|
|
|
| `NEVER_READY` | Stop | Preserve |
|
|
|
|
|
| `GONE` | Best-effort stop | Preserve |
|
|
|
|
|
| `TURN_FAILED` or lead abort while `BUSY` or `FAILED` | Stop | Preserve |
|
|
|
|
|
| `RELEASE_WITH_PENDING_TASKS` | Stop | Preserve |
|
|
|
|
|
| `SHUTDOWN` | Stop | Preserve |
|
|
|
|
|
|
|
|
|
|
Explicit stop is state-aware. `SPAWNING`, `BUSY`, `FAILED`, or any target with pending tasks uses a
|
|
|
|
|
preserving cause.
|
|
|
|
|
|
|
|
|
|
Before abnormal release removes the live session, M4 writes an atomic manifest under the worktree
|
|
|
|
|
root. It records session identity, owner, role, profile, repository, path, branch, base commit,
|
|
|
|
|
release cause, release time, state, and pending task ids. `bridge_list.preservedWorktrees` loads these
|
|
|
|
|
manifests after restart. Stop output and WARN logs also name the path and cause. M4 never
|
|
|
|
|
auto-deletes a preserved worktree.
|
|
|
|
|
|
|
|
|
|
## 6. Fleet health monitor
|
|
|
|
|
|
|
|
|
|
Add `FleetHealthMonitor`. Do not widen `StatusPoller` into a policy loop.
|
|
|
|
|
|
|
|
|
|
`StatusPoller` has a 250 ms delivery cadence and samples only injector targets with outstanding
|
|
|
|
|
work. Health needs all sessions, all leads, task state, inbox age, and global control evidence. One
|
|
|
|
|
loop cannot serve both cadences safely.
|
|
|
|
|
|
|
|
|
|
Build the monitor like `LeadHeartbeatLoop`:
|
|
|
|
|
|
|
|
|
|
- pure `decide(snapshot, priorState, now)` logic;
|
|
|
|
|
- a thin scheduler;
|
|
|
|
|
- an injected clock;
|
|
|
|
|
- edge-triggered state changes;
|
|
|
|
|
- no network work in the pure function;
|
|
|
|
|
- no sleeping in tests.
|
|
|
|
|
|
|
|
|
|
Each enabled fleet tick reads:
|
|
|
|
|
|
|
|
|
|
- one `AgentControl.list()` result for the whole fleet;
|
|
|
|
|
- one in-memory `SessionManager.roster()` snapshot;
|
|
|
|
|
- accepted turns and async task state;
|
|
|
|
|
- typed inbox depth, kind, and age;
|
|
|
|
|
- push and heartbeat outcomes;
|
|
|
|
|
- configured and discovered leads.
|
|
|
|
|
|
|
|
|
|
Existing failure paths publish structured evidence to the monitor. The monitor does not infer events
|
|
|
|
|
from log text.
|
|
|
|
|
|
|
|
|
|
### 6.1 Pane budget
|
|
|
|
|
|
|
|
|
|
Healthy idle members, recent working members, and quiet leads cause no pane reads.
|
|
|
|
|
|
|
|
|
|
A pane is eligible only for a stable lost boundary, sustained `BLOCKED` or `UNKNOWN`, work older
|
|
|
|
|
than the suspect threshold, or one final evidence read for a confirmed fault when the pane exists.
|
|
|
|
|
|
|
|
|
|
Compiled brakes apply even if config asks for more:
|
|
|
|
|
|
|
|
|
|
- per-target pane cooldown is at least 60 seconds;
|
|
|
|
|
- working age before the first progress probe is at least 300 seconds;
|
|
|
|
|
- at most two pane reads occur in one fleet tick;
|
|
|
|
|
- targets rotate fairly;
|
|
|
|
|
- only a normalised digest and optional clipped local excerpt are stored;
|
|
|
|
|
- no pane excerpt leaves bridged in a human webhook.
|
|
|
|
|
|
|
|
|
|
Use `detection` for tested screen signatures. Use normalised `recent_unwrapped` only for progress
|
|
|
|
|
comparison.
|
|
|
|
|
|
|
|
|
|
## 7. Typed inbox and routing
|
|
|
|
|
|
|
|
|
|
### 7.1 Semantic record
|
|
|
|
|
|
|
|
|
|
The typed inbox record carries:
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
schemaVersion
|
|
|
|
|
kind: reply | health
|
|
|
|
|
msgId, target, subjectTerminal, recipientLead
|
|
|
|
|
severity, state, evidence
|
|
|
|
|
createdAtEpochMillis, firstSeenEpochMillis, lastSeenEpochMillis
|
|
|
|
|
recoveryTried, suggestedToolCalls, content
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
A health message never calls `Rendezvous.resolve`. It cannot look like the member's task result.
|
|
|
|
|
|
|
|
|
|
Both inbox adapters share field preservation, first-id-wins dedup, FIFO among decoded messages,
|
|
|
|
|
explicit ownership, ack, and release rules. The in-memory adapter stores typed records directly. It
|
|
|
|
|
does not copy AMQP migration logic.
|
|
|
|
|
|
|
|
|
|
### 7.2 AMQP migration
|
|
|
|
|
|
|
|
|
|
The reader uses AMQP `content_type`, never body sniffing:
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
Legacy v0: text/plain
|
|
|
|
|
Typed family: application/vnd.ltms.bridged.inbox-message+json
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
A legacy reply may begin with `{`. It remains plain text because its media type is `text/plain`.
|
|
|
|
|
Legacy text becomes `kind=reply` with exact UTF-8 content and absent typed metadata.
|
|
|
|
|
|
|
|
|
|
Typed JSON has required integer `schemaVersion: 1`. Version 1 ignores unknown optional fields.
|
|
|
|
|
Missing required fields, invalid enums, malformed UTF-8 or JSON, and property/body identity mismatch
|
|
|
|
|
are invalid data.
|
|
|
|
|
|
|
|
|
|
An unknown schema version is not partly decoded. It remains unacknowledged on the original queue and
|
|
|
|
|
creates one operator-visible `unsupported_version` failure. A newer daemon may read it later.
|
|
|
|
|
|
|
|
|
|
Invalid known-format data is copied byte-for-byte to durable queue
|
|
|
|
|
`agent.<target>.inbox.quarantine`. A dedicated confirm-mode publisher confirms the persistent copy
|
|
|
|
|
before the original is acknowledged. A failed quarantine handoff leaves the original unacknowledged.
|
|
|
|
|
The raw body never enters logs.
|
|
|
|
|
|
|
|
|
|
Decode failure creates a redacted WARN, metric, `bridge_list` summary, and routed health incident.
|
|
|
|
|
One bad entry never escapes the consumer callback and never stops later valid messages.
|
|
|
|
|
|
|
|
|
|
Safe downgrade is not supported. The previous build ignores `content_type` and would show typed JSON
|
|
|
|
|
as ordinary reply text. If drained, it would acknowledge the message and lose typed meaning. Typed
|
|
|
|
|
queues must be drained or preserved before an old jar runs.
|
|
|
|
|
|
|
|
|
|
The existing contract suite uses RabbitMQ. Production uses LavinMQ. The migration and lead-key
|
|
|
|
|
ownership cases must run once against production LavinMQ before release, or the release must state
|
|
|
|
|
that LavinMQ was not checked.
|
|
|
|
|
|
|
|
|
|
### 7.3 Member routing
|
|
|
|
|
|
|
|
|
|
A member incident first goes to the exact lead that owns its accepted delegation.
|
|
|
|
|
`PrimaryRegistry` needs a no-fallback `delegatingLeadFor(memberTarget)` query. Health routing must not
|
|
|
|
|
use the old singular-primary fallback when several leads exist.
|
|
|
|
|
|
|
|
|
|
Publish the incident under the affected member target. Trigger the existing bounded push route. The
|
|
|
|
|
push waits until the owning lead is injectable, so it does not interrupt a live lead turn.
|
|
|
|
|
|
|
|
|
|
### 7.4 Peer lead routing
|
|
|
|
|
|
|
|
|
|
Figure 2 shows the route from incident to lead, peer, or person.
|
|
|
|
|
|
|
|
|
|
```mermaid
|
|
|
|
|
flowchart TD
|
|
|
|
|
Incident["Open incident"] --> Member{"Member incident?"}
|
|
|
|
|
Member -->|"yes"| Known{"Exact delegation owner known?"}
|
|
|
|
|
Known -->|"no"| Sink{"Human webhook enabled and healthy?"}
|
|
|
|
|
Known -->|"yes"| Owner{"Owner lead healthy?"}
|
|
|
|
|
Owner -->|"yes"| OwnerInbox["Publish to owner lead path"]
|
|
|
|
|
Owner -->|"no"| PeerSet["Build healthy peer candidate set"]
|
|
|
|
|
Member -->|"no, lead incident"| PeerSet
|
|
|
|
|
PeerSet --> Peer{"Healthy peer exists?"}
|
|
|
|
|
Peer -->|"yes"| Select["Choose fewest assigned incidents<br/>then stable name and terminal id"]
|
|
|
|
|
Select --> PeerInbox["Publish to peer lead inbox<br/>and status-gated push"]
|
|
|
|
|
Peer -->|"no"| Sink
|
|
|
|
|
Sink -->|"yes"| Webhook["Send classified outbound incident"]
|
|
|
|
|
Sink -->|"no"| Passive["Keep incident open<br/>show partial coverage on local surfaces"]
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
*Figure 2. Routing keeps delegation ownership separate from temporary peer fallback.*
|
|
|
|
|
|
|
|
|
|
Peer candidates exclude the incident subject, failed owner, absent leads, raw-unknown leads, and
|
|
|
|
|
leads with an open unhealthy state. A reachable `WORKING` peer may be selected; its push waits for an
|
|
|
|
|
injectable window.
|
|
|
|
|
|
|
|
|
|
Choose the candidate with the fewest assigned foreign incidents. Break ties by stable lead name,
|
|
|
|
|
then terminal id. Pin the recipient. Reassign only if that peer becomes unhealthy or retires. A
|
|
|
|
|
routing generation marks a reassignment, and old pending assignments become superseded.
|
|
|
|
|
|
|
|
|
|
`bridge_list` lead rows show health, health age, assigned foreign incident count, and a bounded list
|
|
|
|
|
of incident id, subject, state, severity, age, and routing generation. The top-level view also shows
|
|
|
|
|
owner, recipient, and routing reason.
|
|
|
|
|
|
|
|
|
|
A peer incident is published under the recipient lead's inbox key, not the failed subject's key. Its
|
|
|
|
|
status-gated nudge names the failed lead and gives the exact
|
|
|
|
|
`bridge_poll(target="<recipient-terminal>")` call.
|
|
|
|
|
|
|
|
|
|
### 7.5 Lead inbox ownership
|
|
|
|
|
|
|
|
|
|
Add `LeadInboxRegistry`, driven by the same expected-lead supplier as `CallerResolver`.
|
|
|
|
|
|
|
|
|
|
It calls `replyInbox.own(leadTerminal)` at startup for configured leads, after successful discovery,
|
|
|
|
|
after config adds a lead, and before publication. Ownership is not an authorization side effect.
|
|
|
|
|
|
|
|
|
|
A missing lead keeps its key owned. Release happens only after confirmed retirement, all incidents
|
|
|
|
|
are reassigned or resolved, typed health messages move or ack, and the queue is empty. Own a
|
|
|
|
|
replacement terminal before moving messages from the old key. Never release a non-empty in-memory
|
|
|
|
|
lead key, because in-memory release clears local data.
|
|
|
|
|
|
|
|
|
|
### 7.6 Single-lead deployment
|
|
|
|
|
|
|
|
|
|
One lead and no peer is a normal mode, not an edge case.
|
|
|
|
|
|
|
|
|
|
An idle, reachable lead may receive the existing bounded nudge. An unreachable or stalled sole lead
|
|
|
|
|
has no safe in-loop recovery. The bridge must not restart or replace it. A new lead would not have the
|
|
|
|
|
failed lead's plan or context, and an uncertain relaunch could create two orchestrators.
|
|
|
|
|
|
|
|
|
|
With no webhook, only `bridge_list`, `/healthz`, metrics, WARN logs, and the incident journal remain.
|
|
|
|
|
These are passive surfaces. They are not a human notification.
|
|
|
|
|
|
|
|
|
|
## 8. Lost-boundary reconciliation
|
|
|
|
|
|
|
|
|
|
This is the only M4 path that reconstructs a result. It must prefer a visible stall over a fabricated
|
|
|
|
|
reply.
|
|
|
|
|
|
|
|
|
|
### 8.1 Why normal completion rules are not enough
|
|
|
|
|
|
|
|
|
|
Current `CompletionResolver.resolve` has two fail-open rules. It resolves when the delivery baseline
|
|
|
|
|
is missing. It also resolves an empty completion when the pane read fails. Those choices are valid
|
|
|
|
|
after a trusted `WORKING -> IDLE` boundary because the bridge knows the turn ran. They are unsafe
|
|
|
|
|
when health only guesses that a boundary was lost.
|
|
|
|
|
|
|
|
|
|
M4 gives each accepted send an internal `TurnToken`. It ties target, exact waiter, session turn,
|
|
|
|
|
delivery baseline, and task outcome together.
|
|
|
|
|
|
|
|
|
|
### 8.2 Delivery baseline
|
|
|
|
|
|
|
|
|
|
Capture the baseline immediately after prompt send and before the delivery future completes. Store:
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
TurnToken
|
|
|
|
|
exact waiter identity
|
|
|
|
|
capture time and pane source
|
|
|
|
|
normalised assistant block clipped to MAX_SCRAPE_CHARS
|
|
|
|
|
whether a supported assistant marker was recognised
|
|
|
|
|
capture result: PRESENT | READ_FAILED | UNRECOGNISED
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
A failed or missing baseline never authorises repair. A late baseline is not valid evidence. After a
|
|
|
|
|
daemon restart, the old waiter, task, token, and baseline are gone, so the old turn cannot be
|
|
|
|
|
repaired.
|
|
|
|
|
|
|
|
|
|
Automatic repair is enabled only for agent kinds with tested assistant-block fixtures. Current
|
|
|
|
|
extraction is Claude Code-specific and falls back to arbitrary raw text without `⏺`. That raw fallback
|
|
|
|
|
cannot authorise repair. OpenCode repair stays disabled until live pane fixtures exist.
|
|
|
|
|
|
|
|
|
|
### 8.3 Strict gates and resolver result
|
|
|
|
|
|
|
|
|
|
Figure 3 shows the repair gates. Any failed gate keeps the waiter unchanged.
|
|
|
|
|
|
|
|
|
|
The two raw snapshots must describe the same `TurnToken` and session turn. No `WORKING`,
|
|
|
|
|
`BLOCKED`, `UNKNOWN`, missing-agent, reply, failure, or new-delivery observation may occur between
|
|
|
|
|
them.
|
|
|
|
|
|
|
|
|
|
```mermaid
|
|
|
|
|
flowchart TD
|
|
|
|
|
Candidate["TURN_BOUNDARY_LOST candidate"] --> Stable{"Same TurnToken and BUSY turn<br/>across two raw IDLE or DONE snapshots?"}
|
|
|
|
|
Stable -->|"no"| Resnapshot["Take a fresh snapshot"]
|
|
|
|
|
Stable -->|"yes"| Waiter{"Exact captured waiter<br/>still open by identity?"}
|
|
|
|
|
Waiter -->|"no"| Stale["STALE_TURN or ALREADY_RESOLVED"]
|
|
|
|
|
Waiter -->|"yes"| Baseline{"Successful recognised<br/>delivery baseline exists?"}
|
|
|
|
|
Baseline -->|"no"| Refuse["Refuse repair<br/>leave ticket pending"]
|
|
|
|
|
Baseline -->|"yes"| Read{"Fresh pane read succeeds?"}
|
|
|
|
|
Read -->|"no"| Refuse
|
|
|
|
|
Read -->|"yes"| Output{"Recognised non-blank assistant block<br/>differs from clipped baseline?"}
|
|
|
|
|
Output -->|"no"| Refuse
|
|
|
|
|
Output -->|"yes"| Resolve["Shared CompletionResolver guard core<br/>resolves exact waiter"]
|
|
|
|
|
Resolve -->|"won race"| Repaired["RECONCILED_COMPLETION<br/>same turn becomes DONE"]
|
|
|
|
|
Resolve -->|"lost race"| Resnapshot
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
*Figure 3. Repair needs stronger evidence than a normal observed turn boundary.*
|
|
|
|
|
|
|
|
|
|
Refactor the current resolver into one guard core with two policies:
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
resolveCaptured(target, inFlight, OBSERVED_BOUNDARY)
|
|
|
|
|
resolveCaptured(target, inFlight, LOST_BOUNDARY_REPAIR)
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
The health monitor calls only:
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
CompletionResolver.reconcileLostBoundary(target, expectedTurnToken)
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
It returns `REPAIRED`, `ALREADY_RESOLVED`, `REFUSED_NO_CAPTURE`, `REFUSED_NO_BASELINE`,
|
|
|
|
|
`REFUSED_UNREADABLE`, `REFUSED_UNCHANGED`, `REFUSED_AMBIGUOUS_OUTPUT`, `STALE_TURN`, or
|
|
|
|
|
`RACE_LOST`.
|
|
|
|
|
|
|
|
|
|
Only `REPAIRED` and same-turn `ALREADY_RESOLVED` may move that turn from `BUSY` to `DONE`. A
|
|
|
|
|
per-target reconciliation gate stops a queued second send from being accepted between waiter
|
|
|
|
|
resolution and the FSM transition.
|
|
|
|
|
|
|
|
|
|
### 8.4 Lead-visible marker and refusal
|
|
|
|
|
|
|
|
|
|
A repaired result uses distinct `RECONCILED_COMPLETION` values in `Rendezvous`, `MessageService`,
|
|
|
|
|
task poll source, and metrics. The lead sees:
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
[repaired completion - bridged detected a lost turn boundary. The member did not call
|
|
|
|
|
bridge_reply; pane-derived text follows and may be partial]
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Clipped text also keeps the existing clipped-tail marker.
|
|
|
|
|
|
|
|
|
|
A refused repair leaves `TURN_BOUNDARY_LOST` open and the ticket pending. The report states that no
|
|
|
|
|
reply was reconstructed and no task was replayed. `UNCHANGED`, `UNREADABLE`, and
|
|
|
|
|
`AMBIGUOUS_OUTPUT` get at most one delayed retry for the same token. Missing capture or baseline gets
|
|
|
|
|
no retry. After two refused scrapes, automatic repair stops for that token.
|
|
|
|
|
|
|
|
|
|
### 8.5 Target-wide teardown invariant
|
|
|
|
|
|
|
|
|
|
CB-568 owns the multi-ticket cancellation mechanism. M4 routes every terminal cause through that one
|
|
|
|
|
idempotent operation and checks this independent invariant after teardown:
|
|
|
|
|
|
|
|
|
|
- no injector entry exists for the target;
|
|
|
|
|
- no accepted turn or completion record exists;
|
|
|
|
|
- no rendezvous waiter or ask exists;
|
|
|
|
|
- every async task is terminal or was already terminal;
|
|
|
|
|
- no thread waiting for the target send lock can later accept it;
|
|
|
|
|
- new sends fail immediately;
|
|
|
|
|
- each old task has one terminal outcome and one metric count.
|
|
|
|
|
|
|
|
|
|
A violation becomes `DELEGATION_ORPHANED`. The monitor may call the same idempotent target-wide
|
|
|
|
|
failure operation once. It never recreates the task.
|
|
|
|
|
|
|
|
|
|
## 9. Capacity and utilisation
|
|
|
|
|
|
|
|
|
|
Capacity is a view, not a health state.
|
|
|
|
|
|
|
|
|
|
`bridge_list` adds one block per profile:
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
profile, maxLoad, live, free, reclaimable
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
For an unlimited profile, `maxLoad` and `free` are null. `free` is
|
|
|
|
|
`max(0, maxLoad - live)` for a capped profile.
|
|
|
|
|
|
|
|
|
|
The view must use the exact live-count function used by placement. A second calculation could show a
|
|
|
|
|
free slot that placement then refuses. Member rows add `idleForSeconds` only when state is `READY` or
|
|
|
|
|
`DONE`, no accepted turn exists, and the inbox is empty. `reclaimable` means only that the member
|
|
|
|
|
holds capacity without open bridge work.
|
|
|
|
|
|
|
|
|
|
The existing idle-lead nudge gains a bounded capacity summary. It lists per-profile live, cap, free,
|
|
|
|
|
and reclaimable counts, plus at most three long-idle members. Capacity does not make
|
|
|
|
|
`FleetState.hasPending()` true. A changed capacity fingerprint may re-arm one capped heartbeat
|
|
|
|
|
sequence. The fingerprint excludes changing idle durations, so a static idle fleet cannot reset the
|
|
|
|
|
cap forever. Reply-push stand-down remains first.
|
|
|
|
|
|
|
|
|
|
The bridge must never:
|
|
|
|
|
|
|
|
|
|
- spawn a member because a slot is free;
|
|
|
|
|
- generate a task or acceptance criteria;
|
|
|
|
|
- move queued work to another member or profile;
|
|
|
|
|
- treat a free slot or idle member as an incident;
|
|
|
|
|
- stop an idle member only to improve utilisation.
|
|
|
|
|
|
|
|
|
|
The bridge knows capacity facts but has no work list. Only the lead has the plan, task context,
|
|
|
|
|
side-effect history, and acceptance criteria.
|
|
|
|
|
|
|
|
|
|
Capacity calculation is in memory and adds no pane reads. Work-product checks run on a terminal
|
|
|
|
|
session edge, not every fleet tick.
|
|
|
|
|
|
|
|
|
|
This capacity design adds no automatic stop. The accepted `NEVER_READY` cleanup can still stop a
|
|
|
|
|
very slow startup after the existing grace, which is a known risk. Free capacity and long idle time
|
|
|
|
|
never trigger that path.
|
|
|
|
|
|
|
|
|
|
## 10. Human escalation and notification
|
|
|
|
|
|
|
|
|
|
### 10.1 Escalation rule
|
|
|
|
|
|
|
|
|
|
Notify a person only when no healthy lead can act:
|
|
|
|
|
|
|
|
|
|
- `CONTROL_LINK_DOWN` survives grace;
|
|
|
|
|
- a lead is unhealthy and no healthy peer can receive the incident;
|
|
|
|
|
- a member incident has no known owning lead;
|
|
|
|
|
- the only owning lead becomes unreachable, unresponsive, or stalled;
|
|
|
|
|
- incident publication or routing itself fails.
|
|
|
|
|
|
|
|
|
|
Do not page a person for a member fault while a healthy owning lead exists. An uncollected member
|
|
|
|
|
incident feeds lead-health evidence. If the lead then becomes unhealthy, peer or human routing starts.
|
|
|
|
|
|
|
|
|
|
### 10.2 Detection and notification switches
|
|
|
|
|
|
|
|
|
|
`health.enabled` controls detection and bridge-local reporting. It does not require a webhook.
|
|
|
|
|
|
|
|
|
|
`health.notifications.mode` is `disabled` or `webhook`. Disabled is valid and is the default.
|
|
|
|
|
Webhook mode requires a resolved environment variable. Turning notification off stops outbound
|
|
|
|
|
attempts but keeps incidents. Turning it back on resumes still-open human incidents.
|
|
|
|
|
|
|
|
|
|
Without a sink, `bridge_list.healthCoverage` states that human escalation is unavailable. `/healthz`
|
|
|
|
|
keeps its existing HTTP liveness result and adds a nested `fleetHealth.status=partial` component.
|
|
|
|
|
Metrics and one startup or reload WARN expose the same limit.
|
|
|
|
|
|
|
|
|
|
### 10.3 Incident and delivery deduplication
|
|
|
|
|
|
|
|
|
|
One open incident uses this key:
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
(scope, subjectStableId, state, causeFingerprint)
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
The cause fingerprint includes stable error codes, dependency names, signature ids, or invariant
|
|
|
|
|
names. It excludes times, ages, retry counts, pane text, and changing digests. A later recurrence
|
|
|
|
|
after resolution gets a new generation and incident id.
|
|
|
|
|
|
|
|
|
|
Each outbound event uses:
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
Idempotency-Key = hash(incidentId, eventType, eventRevision)
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Event types are `open`, `severity_changed`, `reminder`, and `resolved`. Transport retries keep the
|
|
|
|
|
same key.
|
|
|
|
|
|
|
|
|
|
An atomic owner-only journal beside the active config stores open incidents, routing, delivered
|
|
|
|
|
revisions, retry state, and resolution state. It stores no pane or task content. Journal failure does
|
|
|
|
|
not stop detection, but notification coverage becomes degraded.
|
|
|
|
|
|
|
|
|
|
### 10.4 Retry, reminder, and resolve
|
|
|
|
|
|
|
|
|
|
Send the first event immediately. Retry network errors, timeouts, HTTP 408, HTTP 429, and HTTP 5xx
|
|
|
|
|
with full-jitter exponential backoff:
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
base: 5 seconds
|
|
|
|
|
factor: 3
|
|
|
|
|
maximum delay: 15 minutes
|
|
|
|
|
one outstanding attempt per event
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Respect `Retry-After` up to 15 minutes. Other HTTP 4xx responses are permanent for that event until
|
|
|
|
|
config changes or a person requests replay.
|
|
|
|
|
|
|
|
|
|
Transport retry is not an incident reminder. `humanRepeatSeconds` creates a new reminder revision
|
|
|
|
|
for an unresolved critical incident after the last successful human event. Disabled mode does not
|
|
|
|
|
build an unbounded reminder queue.
|
|
|
|
|
|
|
|
|
|
Send `resolved` only if at least one human event for that incident was delivered. If an incident
|
|
|
|
|
resolves before its first successful delivery, cancel the pending open event and record local
|
|
|
|
|
resolution.
|
|
|
|
|
|
|
|
|
|
### 10.5 Outbound payload boundary
|
|
|
|
|
|
|
|
|
|
An outbound payload may contain incident id and event type, severity, state, scope, stable bridge
|
|
|
|
|
ids, role or profile, times, duration, structured evidence type and counts, recovery attempted,
|
|
|
|
|
routing reason, coverage, and safe tool calls.
|
|
|
|
|
|
|
|
|
|
It must never contain:
|
|
|
|
|
|
|
|
|
|
- raw pane text, pane excerpts, or pane digests;
|
|
|
|
|
- task briefs, prompts, or member reply content;
|
|
|
|
|
- source files, diffs, or worktree file content;
|
|
|
|
|
- worktree paths;
|
|
|
|
|
- environment values, tokens, credentials, headers, or webhook URL;
|
|
|
|
|
- raw exception messages or stack traces;
|
|
|
|
|
- arbitrary model output.
|
|
|
|
|
|
|
|
|
|
The sink response body is ignored. A webhook cannot direct recovery. n8n remains outbound-only.
|
|
|
|
|
|
|
|
|
|
### 10.6 Metrics
|
|
|
|
|
|
|
|
|
|
M4 adds bounded-label series:
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
bridged_health_incidents{scope,state,severity}
|
|
|
|
|
bridged_health_incidents_total{event}
|
|
|
|
|
bridged_health_notifications_total{event,outcome}
|
|
|
|
|
bridged_health_notification_queue_depth
|
|
|
|
|
bridged_health_notification_last_success_seconds
|
|
|
|
|
bridged_health_notification_capability{mode,status}
|
|
|
|
|
bridged_lead_health{lead,state}
|
|
|
|
|
bridged_lead_assigned_incidents{lead}
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Metric labels never include terminal ids, incident ids, URLs, or error text.
|
|
|
|
|
|
|
|
|
|
## 11. Configuration
|
|
|
|
|
|
|
|
|
|
The optional `health:` block is absent or disabled by default. The dormant monitor scheduler does no
|
|
|
|
|
herdr or pane work while disabled. Every listed key is hot because the monitor reads `ConfigRef` on
|
|
|
|
|
each tick or notification.
|
|
|
|
|
|
|
|
|
|
| Key | Class | Default and hard bound | Purpose |
|
|
|
|
|
|---|---|---|---|
|
|
|
|
|
| `health.enabled` | Hot | `false` | Enable detection and bridge-local reporting. |
|
|
|
|
|
| `health.snapshotIntervalSeconds` | Hot | default 30, minimum 15 | Whole-fleet comparison cadence. |
|
|
|
|
|
| `health.workingSuspectAfterSeconds` | Hot | default 600, minimum 300 | Age before working-pane probes. |
|
|
|
|
|
| `health.paneProbeIntervalSeconds` | Hot | default 60, minimum 60 | Per-target pane cooldown. |
|
|
|
|
|
| `health.leadUnresponsiveAfterSeconds` | Hot | default 300, minimum 120 | Delay after exhausted actionable nudges before lead fault. |
|
|
|
|
|
| `health.humanRepeatSeconds` | Hot | default 3600, minimum 900 | Minimum repeat period for one open human incident. |
|
|
|
|
|
| `health.capacityLongIdleAfterSeconds` | Hot | default 900, minimum 300 | Long-idle threshold for capacity summaries. |
|
|
|
|
|
| `health.includePaneExcerpt` | Hot | `false` | Allow a clipped excerpt in local lead reports only. Human payloads still exclude it. |
|
|
|
|
|
| `health.notifications.mode` | Hot | `disabled` | Select `disabled` or `webhook`. |
|
|
|
|
|
| `health.notifications.webhookUrlEnv` | Hot | required in webhook mode | Name of the environment variable that holds the sink URL. |
|
|
|
|
|
| `health.notifications.requestTimeoutMs` | Hot | default 10000, range 1000-30000 | Whole webhook request limit. |
|
|
|
|
|
|
|
|
|
|
Two consecutive snapshots are compiled floors for lost boundary, lead disappearance, and control
|
|
|
|
|
link failure. The two-pane-reads-per-tick limit is also compiled and cannot be weakened by config.
|
|
|
|
|
|
|
|
|
|
## 12. Delivery units and acceptance
|
|
|
|
|
|
|
|
|
|
### Unit 1 - Evidence model and fleet snapshot
|
|
|
|
|
|
|
|
|
|
Scope: health state model, fleet join, clocks, evidence retention, and pane budget.
|
|
|
|
|
|
|
|
|
|
Acceptance criteria:
|
|
|
|
|
|
|
|
|
|
1. One `agent.list` call covers one enabled fleet tick.
|
|
|
|
|
2. Pure decision tests cover every state and every evidence limit in Section 4.
|
|
|
|
|
3. `BUSY` plus stable raw `DONE` opens `TURN_BOUNDARY_LOST` after two snapshots.
|
|
|
|
|
4. Healthy fleet snapshots perform zero pane reads.
|
|
|
|
|
5. Pane cooldown, two-read fleet budget, and fair rotation cannot be disabled by config.
|
|
|
|
|
6. Logs are outputs only; no log parsing exists.
|
|
|
|
|
7. Fleet snapshots expose the same profile live-count calculation that placement uses.
|
|
|
|
|
8. Capacity rows report cap, live, free, and reclaimable values without opening incidents.
|
|
|
|
|
|
|
|
|
|
### Unit 2 - Lost boundary and task reconciliation
|
|
|
|
|
|
|
|
|
|
Scope: accepted-turn identity, guarded repair, target-wide teardown, release causes, and preserved
|
|
|
|
|
worktree discovery.
|
|
|
|
|
|
|
|
|
|
Acceptance criteria:
|
|
|
|
|
|
|
|
|
|
1. Every accepted send receives a stable `TurnToken` tied to target, exact waiter, session turn,
|
|
|
|
|
delivery baseline, and task outcome.
|
|
|
|
|
2. Repair requires the same `BUSY` token, two raw `IDLE` or `DONE` snapshots, no conflicting
|
|
|
|
|
observation, exact open waiter, successful baseline, and new recognised assistant output.
|
|
|
|
|
3. Missing, failed, late, or post-restart baseline never authorises repair.
|
|
|
|
|
4. Repair is enabled only for agent kinds with tested assistant-block extraction. Raw-text fallback
|
|
|
|
|
without a recognised marker refuses repair.
|
|
|
|
|
5. Normal completion and repair use one resolver guard core. Waiter, scrape, clipping, unchanged, and
|
|
|
|
|
exact-turn guards are not duplicated.
|
|
|
|
|
6. `reconcileLostBoundary` returns every typed result named in Section 8.3.
|
|
|
|
|
7. Only `REPAIRED` and same-turn `ALREADY_RESOLVED` may move the same turn to `DONE`.
|
|
|
|
|
8. A per-target reconciliation gate blocks a queued second send during repair and FSM update.
|
|
|
|
|
9. Repaired completion has distinct rendezvous kind, message outcome, poll source, lead marker, and
|
|
|
|
|
metric. Clipping keeps its extra marker.
|
|
|
|
|
10. Unchanged, unreadable, or ambiguous evidence gets at most one delayed retry. Missing capture or
|
|
|
|
|
baseline gets none.
|
|
|
|
|
11. Refusal leaves the ticket pending and tells the lead that no result was rebuilt or replayed.
|
|
|
|
|
12. Release, gone, never-ready, and abnormal stop use CB-568's one idempotent target-wide failure
|
|
|
|
|
operation.
|
|
|
|
|
13. The post-teardown invariant in Section 8.5 is tested independently of CB-568 internals.
|
|
|
|
|
14. A violated teardown invariant creates `DELEGATION_ORPHANED` and retries only the idempotent
|
|
|
|
|
failure operation.
|
|
|
|
|
15. `SPAWN_ROLLBACK` and normal `COMPLETED` remove worktrees. Abnormal and shutdown causes preserve
|
|
|
|
|
them.
|
|
|
|
|
16. Explicit stop is state-aware. Any pending task or non-terminal state preserves the worktree.
|
|
|
|
|
17. Atomic preserved-worktree manifests reload after restart and appear in lead-only
|
|
|
|
|
`bridge_list.preservedWorktrees`.
|
|
|
|
|
18. Manifest failure preserves the worktree and opens an operator-visible health failure.
|
|
|
|
|
19. Provision records the base commit. Terminal, long-idle worktrees report
|
|
|
|
|
`WORK_PRODUCT_AT_RISK` only under the evidence in Section 4.2 and never auto-delete work.
|
|
|
|
|
20. No path replays a delivered task, rebuilds its brief, or retargets it, even when prompt text is
|
|
|
|
|
available.
|
|
|
|
|
21. Tests cover both real traces, all repair refusals, clipping, explicit-reply and next-turn races,
|
|
|
|
|
restart without capture, concurrent send and release, and preserved discovery after restart.
|
|
|
|
|
|
|
|
|
|
### Unit 3 - Typed inbox and member routing
|
|
|
|
|
|
|
|
|
|
Scope: semantic record, AMQP migration, both adapters, member routing, polling, and member health in
|
|
|
|
|
`bridge_list`.
|
|
|
|
|
|
|
|
|
|
Acceptance criteria:
|
|
|
|
|
|
|
|
|
|
1. AMQP selects legacy or typed decoding only from `content_type`; it never sniffs the body.
|
|
|
|
|
2. Persistent `text/plain` from the old build becomes `kind=reply` with exact UTF-8 content,
|
|
|
|
|
including content beginning with `{`.
|
|
|
|
|
3. New entries use the vendor media type, `schemaVersion: 1`, UTF-8, persistent delivery, and AMQP
|
|
|
|
|
message ids.
|
|
|
|
|
4. Version 1 ignores unknown optional fields but rejects missing fields and identity mismatch.
|
|
|
|
|
5. Unknown versions are not decoded or acked. They remain on the original queue and create one
|
|
|
|
|
deduplicated failure.
|
|
|
|
|
6. Invalid known data never escapes the callback, appears as a reply, or blocks later valid messages.
|
|
|
|
|
7. Invalid data reaches durable per-target quarantine before original ack. Failed handoff leaves the
|
|
|
|
|
original unacked.
|
|
|
|
|
8. Decode failures create redacted WARN, metric, `bridge_list` summary, and routed incident without
|
|
|
|
|
raw content.
|
|
|
|
|
9. Both adapters pass one semantic contract for fields, FIFO, dedup, ownership, ack, and release.
|
|
|
|
|
10. Lead keys require explicit ownership. Publication never claims a queue.
|
|
|
|
|
11. Unit codec tests cover legacy `{`, Unicode, malformed UTF-8, typed round trip, additive fields,
|
|
|
|
|
malformed JSON, missing fields, identity mismatch, media type, version, and dedup.
|
|
|
|
|
12. A live broker contract writes old wire data and reads it with the new adapter after reconnect.
|
|
|
|
|
13. Live contract tests cover mixed entries, quarantine confirm-before-ack, unsupported redelivery,
|
|
|
|
|
later progress past poison, property persistence, lead ownership, and ack removal.
|
|
|
|
|
14. Safe downgrade is documented as unsupported.
|
|
|
|
|
15. RabbitMQ contract tests pass with `mvn test -Pcontract`. The same cases run once on production
|
|
|
|
|
LavinMQ, or the release states that LavinMQ was not checked.
|
|
|
|
|
16. Member incidents route to the exact delegating lead and never resolve a task rendezvous.
|
|
|
|
|
17. `bridge_list` shows compact member health and capacity without pane content. Member
|
|
|
|
|
`idleForSeconds` is present only when no accepted turn or inbox item exists.
|
|
|
|
|
|
|
|
|
|
### Unit 4 - Lead health and peer routing
|
|
|
|
|
|
|
|
|
|
Scope: lead evidence, exact ownership, peer selection, explicit-recipient push, and lead inbox
|
|
|
|
|
lifecycle.
|
|
|
|
|
|
|
|
|
|
Acceptance criteria:
|
|
|
|
|
|
|
|
|
|
1. Lead identity uses the `CallerResolver` supplier. Liveness uses successful current agent data.
|
|
|
|
|
2. Two successful-list absences with healthy ping become `LEAD_UNREACHABLE`; global link failure does
|
|
|
|
|
not.
|
|
|
|
|
3. Raw `WORKING`, raw `UNKNOWN`, first-seen time, failures, last success, and error class persist
|
|
|
|
|
across ticks.
|
|
|
|
|
4. Heartbeat and push publish status and nudge outcomes before safe no-injection decisions.
|
|
|
|
|
5. Dynamic lead identity survives a two-successful-snapshot retirement grace.
|
|
|
|
|
6. Member incidents first use exact delegation ownership with no singular-primary fallback.
|
|
|
|
|
7. Peer selection follows the exclusions, load rule, and stable tie break in Section 7.4.
|
|
|
|
|
8. A selected working peer is not interrupted. Its push waits for an injectable window.
|
|
|
|
|
9. Recipient assignment stays pinned. Reassignment increments generation and supersedes old pending
|
|
|
|
|
assignment.
|
|
|
|
|
10. `bridge_list` shows bounded foreign assignments, recipient, reason, and generation without pane
|
|
|
|
|
content.
|
|
|
|
|
11. `LeadInboxRegistry` owns configured and discovered lead keys before publication.
|
|
|
|
|
12. Missing leads keep ownership. Retirement needs an empty queue and handled incidents.
|
|
|
|
|
13. Replacement owns the new key before messages move. Non-empty in-memory keys are not released.
|
|
|
|
|
14. Tests cover dead versus busy, unknown, global failure, stale scan cache, disappearance, one peer,
|
|
|
|
|
several peers, reassignment, and no peer.
|
|
|
|
|
15. Adapter tests cover lead ownership, restart re-ownership, retirement, and terminal replacement.
|
|
|
|
|
LavinMQ is checked or named as unchecked.
|
|
|
|
|
16. A sole unreachable or stalled lead is never restarted or replaced. Without a sink, only passive
|
|
|
|
|
evidence remains and every coverage surface says so.
|
|
|
|
|
|
|
|
|
|
### Unit 5 - Human sink, hot config, metrics, and operator coverage
|
|
|
|
|
|
|
|
|
|
Scope: generic webhook, config split, incident journal, retry, resolve, metrics, example config, and
|
|
|
|
|
operator documentation.
|
|
|
|
|
|
|
|
|
|
Acceptance criteria:
|
|
|
|
|
|
|
|
|
|
1. `health.enabled` works without a human sink.
|
|
|
|
|
2. Notification mode is hot, defaults to disabled, and supports disabled or webhook.
|
|
|
|
|
3. Webhook mode requires a resolved environment value. Bad notification config does not disable an
|
|
|
|
|
already valid detector.
|
|
|
|
|
4. Mode changes keep open incidents. Re-enable resumes eligible incidents.
|
|
|
|
|
5. `bridge_list`, `/healthz`, metrics, and one WARN show partial coverage without a sink. HTTP
|
|
|
|
|
liveness behavior stays unchanged.
|
|
|
|
|
6. One-lead, no-sink coverage states that lead failure has no active notification or recovery.
|
|
|
|
|
7. Incident and outbound dedupe use the stable keys in Section 10.3.
|
|
|
|
|
8. The owner-only local journal survives restart and contains no pane or task content.
|
|
|
|
|
9. Journal failure keeps detection running but marks notification coverage degraded.
|
|
|
|
|
10. Retry tests cover network failure, timeout, 408, 429, `Retry-After`, 5xx, permanent 4xx, jitter,
|
|
|
|
|
delay cap, config re-arm, and one outstanding attempt.
|
|
|
|
|
11. Reminders and transport retries remain separate. Disabled mode does not build an unbounded queue.
|
|
|
|
|
12. Resolve sends only after an earlier human event succeeded. Resolve-before-delivery cancels stale
|
|
|
|
|
open delivery.
|
|
|
|
|
13. Metrics use bounded labels and exclude ids, URLs, and error text.
|
|
|
|
|
14. Payload tests reject every content type forbidden in Section 10.5.
|
|
|
|
|
15. Webhook response bodies are ignored and cannot direct recovery.
|
|
|
|
|
16. Tests cover disabled mode, one lead without sink, open/update/reminder/resolve, restart, dedup,
|
|
|
|
|
reassignment, disable/re-enable, and sink failure while local health continues.
|
|
|
|
|
17. `bridged.example.yaml` documents all hot keys and compiled floors.
|
|
|
|
|
18. The operator Features wiki is updated separately. The portable `CLAUDE.md` block is checked and
|
|
|
|
|
changed only if shipped tool or inbox semantics make it untrue.
|
|
|
|
|
19. `mvn clean install` passes.
|
|
|
|
|
|
|
|
|
|
## 13. Not checked and release gates
|
|
|
|
|
|
|
|
|
|
These limits are part of the design, not optional follow-up notes.
|
|
|
|
|
|
|
|
|
|
- **OpenCode pane status and assistant markers were not checked.** OpenCode lost-boundary repair is
|
|
|
|
|
disabled until live fixtures exist.
|
|
|
|
|
- **Permission-prompt status was not checked** for Claude Code or OpenCode. `BLOCKED` remains
|
|
|
|
|
ambiguous and has no automatic action.
|
|
|
|
|
- **`recent_unwrapped` stability was not checked** across all supported agent kinds. If normalisation
|
|
|
|
|
is not stable, `STALL_SUSPECTED` must say its evidence is weaker.
|
|
|
|
|
- **The real `BUSY + DONE` trace was not replayed against live herdr.** The design uses the observed
|
|
|
|
|
production trace and current poller behavior.
|
|
|
|
|
- **CB-568 was not present when Unit 2 was designed.** Unit 2 must inspect the landed API and keep its
|
|
|
|
|
independent teardown invariant.
|
|
|
|
|
- **Production LavinMQ was not checked.** Existing durable-inbox contracts use RabbitMQ. Migration,
|
|
|
|
|
quarantine, redelivery, lead ownership, and reassignment must run on LavinMQ before release or be
|
|
|
|
|
recorded as unchecked.
|
|
|
|
|
- **Live multi-lead routing was not checked.** Peer choice and reassignment are design rules backed by
|
|
|
|
|
fake-clock and adapter tests until a live exercise runs.
|
|
|
|
|
- **A live sole-lead failure with a webhook was not checked.** The no-peer path is a design result,
|
|
|
|
|
not a tested recovery.
|
|
|
|
|
- **No n8n, Slack, PagerDuty, or other receiver was checked.** The webhook remains generic and
|
|
|
|
|
outbound-only.
|
|
|
|
|
- **Deployment supervisor behavior for nested `/healthz` fields was not checked.** HTTP liveness
|
|
|
|
|
status stays unchanged to reduce this risk.
|
|
|
|
|
- **Incident-journal crash behavior was not checked** because the journal does not exist yet. Unit 5
|
|
|
|
|
must test atomic replacement and restart recovery.
|
|
|
|
|
- **Worktree merge state cannot be checked reliably** without forge or explicit collection evidence.
|
|
|
|
|
`WORK_PRODUCT_AT_RISK` stays a warning.
|
|
|
|
|
|
|
|
|
|
## 14. Locked exclusions
|
|
|
|
|
|
|
|
|
|
M4 does not expose `agent.read` as a bridge tool. It does not add a workflow engine, inbound n8n
|
|
|
|
|
authority, automatic task assignment, task replay, automatic lead replacement, or automatic member
|
|
|
|
|
spawn for free capacity.
|
|
|
|
|
|
|
|
|
|
The bridge remains a message bus with evidence and bounded mechanical repair. The lead remains the
|
|
|
|
|
place where judgement and work planning happen.
|