2e138a199b
Two changes ship together here.
1. One shared herdr workspace. The lead and every worker now live in one
workspace called "fleet", so the operator sees one "session" with many
windows, not two. Before, the lead sat in a "leads" workspace and workers
in "bridged-workers", which read as two sessions. The lead is still told
apart from workers by its exact tab label ("lead: <name>"), so putting them
in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
for split layouts; Fleetd now passes an empty exclude set.
2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
launchd/systemd units, module dir, and MCP mount).
- Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
- Log line, comments, docs, and CLAUDE.md updated to say fleetd.
- Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
- Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
- Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
still read as a fallback, and still gitignored.
- MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
server name is "fleet". The mount name in the local .mcp.json becomes
"fleet" (gitignored, not in this commit).
- Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.
Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.
Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.
The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).
949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
973 lines
51 KiB
Markdown
973 lines
51 KiB
Markdown
# M4 - Fleet health, recovery, routing, and capacity
|
|
|
|
**Status:** Design accepted on 2026-08-15. CB-573 part 1 has shipped the classification model and
|
|
the `fleet_list` capacity view; the remaining M4 units are not yet shipped. See
|
|
[Unit 2 — what has landed so far](#unit-2---what-has-landed-so-far) before planning Unit 2 work:
|
|
some of its criteria were met by separate CB tickets, and one of them contradicts the unit text.
|
|
**Scope:** Fleet evidence, safe mechanical repair, lead routing, capacity reporting, and optional
|
|
human notification.
|
|
**Grounded in:** `health/FleetHealth`, `health/PaneBudget`, `inject/StatusPoller`,
|
|
`inject/StatusRefiner`, `inject/CompletionResolver`, `inject/Injector`, `session/SessionManager`,
|
|
`msg/MessageService`, `msg/ReplyInbox`, `msg/ReplyPushLoop`, `msg/LeadHeartbeatLoop`,
|
|
`mcp/PrimaryRegistry`, and `herdr/AgentControl`.
|
|
|
|
## 1. Problem and decision boundary
|
|
|
|
The operator asked the bridge to detect idle agents, exceptions, stopped work, and broken
|
|
communication. The bridge may read an agent pane from time to time. It must notify a person when
|
|
the fleet cannot move forward.
|
|
|
|
The four operator terms are not four equal health states. `IDLE` is a normal mode. An exception is
|
|
sometimes visible only as pane text. Stopped work may look the same as slow work. Broken
|
|
communication can occur on several links.
|
|
|
|
M4 uses this boundary:
|
|
|
|
- The bridge detects facts and joins evidence.
|
|
- The bridge repairs only mechanical failures with no judgement.
|
|
- The lead decides whether to stop, retry, replace, or reassign a member.
|
|
- A human is notified only when no healthy lead can act.
|
|
- n8n may route an outbound incident. It never classifies state or chooses recovery.
|
|
|
|
An inbound n8n decider would need bridge authority. No narrow machine-decider role exists. Giving a
|
|
workflow engine lead authority is unsafe, while adding a new role is a separate authorization
|
|
design. An outbound sink needs no bridge role.
|
|
|
|
The bridge must never replay a delivered task. That task may already have changed files, pushed a
|
|
branch, opened a pull request, or changed external state. A replay can run those side effects twice.
|
|
This rule must remain true even if later code stores delivered prompt text.
|
|
|
|
## 2. Evidence model
|
|
|
|
A health state is mainly a comparison between two views:
|
|
|
|
- **herdr view:** current agents and raw live status from one `AgentControl.list()` call.
|
|
- **bridge view:** session FSM, MCP presence, accepted turns, tasks, inbox state, and lead ownership.
|
|
|
|
A strong fault often appears as a disagreement between those views. For example, `BUSY` in the
|
|
session FSM and `DONE` in herdr means the bridge missed a turn boundary. Pane reads support this
|
|
model, but they are not the main monitor.
|
|
|
|
`SessionManager.rosterView` already joins session state and live status. `AgentControl.list()`
|
|
already gets the whole live fleet in one call. M4 makes that join persistent and adds timers,
|
|
accepted-turn state, and incident state.
|
|
|
|
### 2.1 Real traces behind the design
|
|
|
|
The first trace was an architect that stopped making progress:
|
|
|
|
```text
|
|
profile=opus role=architect state=busy liveStatus=done
|
|
```
|
|
|
|
The session moved from `DONE` to `BUSY` for turn 2. Eighteen minutes later, the session still said
|
|
`BUSY`, herdr still said `DONE`, the async task still said `PENDING`, and no completion fallback had
|
|
run. This is `TURN_BOUNDARY_LOST`, not a general slow-turn guess.
|
|
|
|
The second trace had two async sends to the same pane, one second apart. The pane was then stopped.
|
|
One ticket became failed. The other stayed `pending - worker unknown`. Current
|
|
`MessageService.abandon` resolves only `Rendezvous.currentWaiter(target)`, while async tasks live in
|
|
a separate ticket map. CB-568 is intended to fix that bug. M4 still keeps an independent
|
|
post-teardown invariant so a later regression becomes `DELEGATION_ORPHANED`.
|
|
|
|
### 2.2 Corrections made during design
|
|
|
|
The first state table missed `BUSY` in fleetd plus `IDLE` or `DONE` in herdr. It would have found
|
|
the real trace only through a late, weak stall timer. The final model adds
|
|
`TURN_BOUNDARY_LOST` as a strong disagreement state.
|
|
|
|
The first notification design also required a webhook before `health.enabled` could turn on. That
|
|
removed useful local detection to avoid a narrower human-notification gap. The final design splits
|
|
detection from notification. Missing human escalation is shown as partial coverage instead of
|
|
disabling health.
|
|
|
|
## 3. Classification precedence
|
|
|
|
Evidence is applied in this order. A lower rule cannot hide a higher one.
|
|
|
|
1. **Control link:** failed fleet list plus failed ping becomes `CONTROL_LINK_DOWN`.
|
|
2. **Definitive target loss:** `_not_found` becomes `GONE` or `LEAD_UNREACHABLE` when the control
|
|
link is healthy.
|
|
3. **Startup and teardown invariants:** readiness expiry becomes `NEVER_READY`; surviving tasks
|
|
after teardown become `DELEGATION_ORPHANED`.
|
|
4. **Bridge/live disagreement:** `BUSY` plus stable raw `IDLE` or `DONE` becomes
|
|
`TURN_BOUNDARY_LOST`.
|
|
5. **Known screen evidence:** a tested fatal signature becomes `ERROR_ON_SCREEN`.
|
|
6. **Timed suspicion:** unchanged sparse pane probes may become `STALL_SUSPECTED`.
|
|
7. **Communication quality:** completion fallback becomes `MUTE`; an old inbox entry becomes
|
|
`REPLY_STRANDED`.
|
|
8. **Normal mode:** `STARTING`, `IDLE`, `WORKING`, `WORK_PENDING`, or `BLOCKED_AMBIGUOUS`.
|
|
|
|
The member flow in Figure 1 shows lifecycle states and the main fault exits. Fault states are
|
|
reported beside the session FSM; most are not new FSM values.
|
|
|
|
```mermaid
|
|
flowchart TD
|
|
Registered["Member registered"] --> Starting["STARTING"]
|
|
Starting -->|"MCP presence"| Idle["IDLE"]
|
|
Starting -->|"Readiness grace expires"| NeverReady["NEVER_READY"]
|
|
Idle -->|"Accepted delivery"| Working["WORKING"]
|
|
Working -->|"Trusted turn boundary"| Idle
|
|
Working -->|"Bridge BUSY and herdr IDLE or DONE"| Lost["TURN_BOUNDARY_LOST"]
|
|
Working -->|"Known fatal screen"| Error["ERROR_ON_SCREEN"]
|
|
Working -->|"Long age and unchanged sparse probes"| Stall["STALL_SUSPECTED"]
|
|
Working -->|"Target not found"| Gone["GONE"]
|
|
Idle -->|"Inbox or queued delivery exists"| Pending["WORK_PENDING"]
|
|
Pending -->|"Delivery or collection finishes"| Idle
|
|
Idle -->|"Raw BLOCKED with an open turn"| Blocked["BLOCKED_AMBIGUOUS"]
|
|
Lost -->|"Strict guarded repair"| Repaired["DONE with reconciled completion"]
|
|
Lost -->|"Repair refused"| LeadDecision["Lead decision required"]
|
|
```
|
|
|
|
*Figure 1. The member lifecycle and the main health exits. Pane-based states never authorise an
|
|
automatic retry of the task.*
|
|
|
|
## 4. State model
|
|
|
|
### 4.1 Normal and transitional member states
|
|
|
|
| State | Exact evidence | Meaning and certainty |
|
|
|---|---|---|
|
|
| `STARTING` | Session is `SPAWNING`; MCP presence is absent | Normal inside the startup grace. MCP contact is the readiness signal. |
|
|
| `IDLE` | Session is `READY` or `DONE`; live status is `IDLE` or `DONE`; no open turn or inbox item exists | Normal. Idle is not a fault. |
|
|
| `WORKING` | Session is `BUSY`; raw live status is `WORKING`; the accepted turn is open | Certain that herdr sees work. It does not prove useful progress. |
|
|
| `WORK_PENDING` | Queued delivery or inbox content exists while the target is injectable | Transitional. Existing injector or push logic should move it. |
|
|
| `BLOCKED_AMBIGUOUS` | An open turn exists and raw live status is `BLOCKED` | The bridge cannot tell whether this is permission, input, or a settled screen. |
|
|
|
|
Idle may drive configured resource cleanup. It never opens an incident and never pages a person.
|
|
|
|
### 4.2 Member fault and quality states
|
|
|
|
| State | Exact evidence | Certainty and action |
|
|
|---|---|---|
|
|
| `NEVER_READY` | `SPAWNING`, no MCP presence, and an accepted delivery waits through the existing readiness grace | Delivery never became possible. The exact cause is unknown. Fail the send, stop the process, and preserve a provisioned worktree. |
|
|
| `GONE` | Per-target herdr call returns `_not_found` while fleet list or ping works | Certain target loss. Fail all target work. Do not replay it. |
|
|
| `TURN_BOUNDARY_LOST` | Same session turn stays `BUSY`; same accepted task stays open; two raw snapshots show `IDLE` or `DONE` | Strong disagreement. Strict reconciliation may repair it. |
|
|
| `ERROR_ON_SCREEN` | Suspicious non-working state survives grace; `detection` matches a tested adapter-specific fatal signature | Certain only for the matched signature. A bare word such as `Exception` is not enough. |
|
|
| `STALL_SUSPECTED` | Open turn is older than the configured threshold; two normalised `recent_unwrapped` digests are unchanged; no boundary or reply occurs | Not certain. A long valid API call can look the same. Lead decides. |
|
|
| `MUTE` | Turn resolves through completion fallback instead of `fleet_reply` | Certain that no structured reply won. It does not prove an MCP failure. A single event is a metric, not an incident. |
|
|
| `REPLY_STRANDED` | Typed reply or health message remains after owning-lead push reaches its cap | Collection failed. This does not explain whether the lead is busy, dead, or ignoring the nudge. |
|
|
| `DELEGATION_ORPHANED` | Target is gone, failed, or released, but one or more tasks remain `PENDING` after reconciliation grace | Certain bridge invariant failure. This is not an inbox-drain fault. |
|
|
| `WORK_PRODUCT_AT_RISK` | Provisioned branch has commits after its recorded base; member is `DONE`, `FAILED`, or preserved after release; no turn or inbox item remains; long-idle threshold passed | A warning, not proof of loss. Work may already have an open pull request or a squash merge. |
|
|
|
|
`MUTE` opens an incident only after a small fixed rate threshold for one target or profile, or when
|
|
it appears with another fault.
|
|
|
|
`WORK_PRODUCT_AT_RISK` must not become `WORK_PRODUCT_UNCOLLECTED`. The bridge does not know pull
|
|
request or merge state. If committed work appears with `REPLY_STRANDED` or
|
|
`DELEGATION_ORPHANED`, the existing incident gains `committedWorkAtRisk: true`.
|
|
|
|
### 4.3 Control-link state
|
|
|
|
| State | Exact evidence | Certainty and action |
|
|
|---|---|---|
|
|
| `CONTROL_LINK_DOWN` | Two full-fleet `agent.list` calls fail across the grace, and herdr `ping` also fails | Certain for the fleetd-to-herdr link. Retry calls, record the incident, and use human escalation if no lead can be reached. |
|
|
|
|
A failed fleet list alone is not a dead-member claim. A single `_not_found` with a healthy global
|
|
link is a target fault, not a control-link fault.
|
|
|
|
### 4.4 Lead states
|
|
|
|
| State | Exact evidence | Meaning and action |
|
|
|---|---|---|
|
|
| `LEAD_IDLE` | Expected lead is present with raw injectable status; no actionable state waits | Normal. Existing heartbeat may run under its own policy. |
|
|
| `LEAD_WORKING` | Expected lead is present with raw `WORKING`; stall threshold is not met | Reachable and busy. Never inject into the live turn. |
|
|
| `LEAD_STATUS_UNKNOWN` | Expected lead is present with raw `UNKNOWN` | Neither dead nor a healthy routing target. Retain evidence and retry. |
|
|
| `LEAD_UNREACHABLE` | Expected lead is absent from two successful live-agent snapshots while ping works, or targeted lookup returns `_not_found` with a healthy control link | Route to a healthy peer. If none exists, use human escalation. |
|
|
| `LEAD_UNRESPONSIVE` | Actionable state waits; lead stays injectable; bounded nudges exhaust; inbox remains uncollected | Route to a healthy peer or a person. |
|
|
| `LEAD_STALL_SUSPECTED` | Lead stays `WORKING` past threshold; two sparse pane probes show no progress | Not certain. Never kill or restart automatically. Route to peer or person. |
|
|
|
|
The monitor retains the lead name and terminal, last successful sighting, raw status and age,
|
|
consecutive list absences, targeted errors, pane-probe facts, pending incident age, and nudge
|
|
outcomes. Current heartbeat and push loops discard much of this history.
|
|
|
|
Expected lead identity comes from the same supplier used by `CallerResolver`. It is not liveness
|
|
evidence. `LeadTabScanner` keeps cached identity after a failed scan, so the health monitor compares
|
|
that identity with a fresh successful agent list. A dynamic identity also survives a two-successful-
|
|
snapshot retirement grace. This stops a dead lead from escaping health by disappearing from one map.
|
|
|
|
### 4.5 Evidence limits
|
|
|
|
M4 cannot tell these cases apart with current evidence:
|
|
|
|
- A valid long call and a hung call may have the same status and pane digest.
|
|
- `BLOCKED` does not explain which input is needed.
|
|
- An idle prompt after failure may look like an idle prompt after success.
|
|
- A missing structured reply does not prove a broken MCP connection.
|
|
- An undrained inbox does not explain why the lead did not collect it.
|
|
- Arbitrary pane text cannot safely classify arbitrary exceptions.
|
|
- A branch ahead of its base does not prove that work was not collected.
|
|
|
|
Logs are outputs, not classifier inputs. The monitor never parses its own logs.
|
|
|
|
## 5. Automatic action and lead action
|
|
|
|
### 5.1 Actions the bridge may take
|
|
|
|
The bridge may:
|
|
|
|
- retry transient herdr status, list, ping, and pane-read failures with bounded backoff;
|
|
- re-submit Enter after the existing paste/submit race;
|
|
- fail queued delivery after `NEVER_READY`;
|
|
- stop a never-ready process while preserving its provisioned worktree;
|
|
- fail all queued, accepted, and async tasks for a gone or released target;
|
|
- reconcile one lost boundary when every strict gate in Section 8 passes;
|
|
- hold typed messages, nudge the owning lead, and stop at the configured cap;
|
|
- use the existing bounded idle-lead heartbeat;
|
|
- deduplicate, route, update, and resolve incidents.
|
|
|
|
These actions do not choose new work and do not replay old work.
|
|
|
|
### 5.2 Decisions reserved for the lead
|
|
|
|
Only the lead may:
|
|
|
|
- stop or continue `BLOCKED_AMBIGUOUS`;
|
|
- stop, inspect, or wait on `ERROR_ON_SCREEN`;
|
|
- kill or continue `STALL_SUSPECTED`;
|
|
- spawn a replacement or reassign work;
|
|
- retry a delivered task;
|
|
- choose how to use partial work in a worktree;
|
|
- restart herdr or change network, model, credentials, backend, or configuration.
|
|
|
|
Reports include literal safe tool calls such as `fleet_status(sessionId="...")`,
|
|
`fleet_poll(ticket="...")`, `fleet_list()`, and optional `fleet_stop(paneId="...")`. A judgement
|
|
state never presents stop as the only action.
|
|
|
|
### 5.3 Release causes and worktree safety
|
|
|
|
| Release cause | Process action | Provisioned worktree |
|
|
|---|---|---|
|
|
| `SPAWN_ROLLBACK` before registration or delivery | Stop and clean up | Remove |
|
|
| `COMPLETED` for `READY` or `DONE` without pending work, idle TTL, or successful context-cap completion | Stop | Remove only if clean; preserve a dirty worktree (CB-576) |
|
|
| `NEVER_READY` | Stop | Preserve |
|
|
| `GONE` | Best-effort stop | Preserve |
|
|
| `TURN_FAILED` or lead abort while `BUSY` or `FAILED` | Stop | Preserve |
|
|
| `RELEASE_WITH_PENDING_TASKS` | Stop | Preserve |
|
|
| `SHUTDOWN` | Stop | Preserve |
|
|
|
|
Explicit stop is state-aware. `SPAWNING`, `BUSY`, `FAILED`, or any target with pending tasks uses a
|
|
preserving cause.
|
|
|
|
Before abnormal release removes the live session, M4 writes an atomic manifest under the worktree
|
|
root. It records session identity, owner, role, profile, repository, path, branch, base commit,
|
|
release cause, release time, state, and pending task ids. `fleet_list.preservedWorktrees` loads these
|
|
manifests after restart. Stop output and WARN logs also name the path and cause. M4 never
|
|
auto-deletes a preserved worktree.
|
|
|
|
## 6. Fleet health monitor
|
|
|
|
Add `FleetHealthMonitor`. Do not widen `StatusPoller` into a policy loop.
|
|
|
|
`StatusPoller` has a 250 ms delivery cadence and samples only injector targets with outstanding
|
|
work. Health needs all sessions, all leads, task state, inbox age, and global control evidence. One
|
|
loop cannot serve both cadences safely.
|
|
|
|
Build the monitor like `LeadHeartbeatLoop`:
|
|
|
|
- pure `decide(snapshot, priorState, now)` logic;
|
|
- a thin scheduler;
|
|
- an injected clock;
|
|
- edge-triggered state changes;
|
|
- no network work in the pure function;
|
|
- no sleeping in tests.
|
|
|
|
Each enabled fleet tick reads:
|
|
|
|
- one `AgentControl.list()` result for the whole fleet;
|
|
- one in-memory `SessionManager.roster()` snapshot;
|
|
- accepted turns and async task state;
|
|
- typed inbox depth, kind, and age;
|
|
- push and heartbeat outcomes;
|
|
- configured and discovered leads.
|
|
|
|
Existing failure paths publish structured evidence to the monitor. The monitor does not infer events
|
|
from log text.
|
|
|
|
### 6.1 Pane budget
|
|
|
|
Healthy idle members, recent working members, and quiet leads cause no pane reads.
|
|
|
|
A pane is eligible only for a stable lost boundary, sustained `BLOCKED` or `UNKNOWN`, work older
|
|
than the suspect threshold, or one final evidence read for a confirmed fault when the pane exists.
|
|
|
|
Compiled brakes apply even if config asks for more:
|
|
|
|
- per-target pane cooldown is at least 60 seconds;
|
|
- working age before the first progress probe is at least 300 seconds;
|
|
- at most two pane reads occur in one fleet tick;
|
|
- targets rotate fairly;
|
|
- only a normalised digest and optional clipped local excerpt are stored;
|
|
- no pane excerpt leaves fleetd in a human webhook.
|
|
|
|
Use `detection` for tested screen signatures. Use normalised `recent_unwrapped` only for progress
|
|
comparison.
|
|
|
|
## 7. Typed inbox and routing
|
|
|
|
### 7.1 Semantic record
|
|
|
|
The typed inbox record carries:
|
|
|
|
```text
|
|
schemaVersion
|
|
kind: reply | health
|
|
msgId, target, subjectTerminal, recipientLead
|
|
severity, state, evidence
|
|
createdAtEpochMillis, firstSeenEpochMillis, lastSeenEpochMillis
|
|
recoveryTried, suggestedToolCalls, content
|
|
```
|
|
|
|
A health message never calls `Rendezvous.resolve`. It cannot look like the member's task result.
|
|
|
|
Both inbox adapters share field preservation, first-id-wins dedup, FIFO among decoded messages,
|
|
explicit ownership, ack, and release rules. The in-memory adapter stores typed records directly. It
|
|
does not copy AMQP migration logic.
|
|
|
|
### 7.2 AMQP migration
|
|
|
|
The reader uses AMQP `content_type`, never body sniffing:
|
|
|
|
```text
|
|
Legacy v0: text/plain
|
|
Typed family: application/vnd.ltms.fleet.inbox-message+json
|
|
```
|
|
|
|
A legacy reply may begin with `{`. It remains plain text because its media type is `text/plain`.
|
|
Legacy text becomes `kind=reply` with exact UTF-8 content and absent typed metadata.
|
|
|
|
Typed JSON has required integer `schemaVersion: 1`. Version 1 ignores unknown optional fields.
|
|
Missing required fields, invalid enums, malformed UTF-8 or JSON, and property/body identity mismatch
|
|
are invalid data.
|
|
|
|
An unknown schema version is not partly decoded. It remains unacknowledged on the original queue and
|
|
creates one operator-visible `unsupported_version` failure. A newer daemon may read it later.
|
|
|
|
Invalid known-format data is copied byte-for-byte to durable queue
|
|
`agent.<target>.inbox.quarantine`. A dedicated confirm-mode publisher confirms the persistent copy
|
|
before the original is acknowledged. A failed quarantine handoff leaves the original unacknowledged.
|
|
The raw body never enters logs.
|
|
|
|
Decode failure creates a redacted WARN, metric, `fleet_list` summary, and routed health incident.
|
|
One bad entry never escapes the consumer callback and never stops later valid messages.
|
|
|
|
Safe downgrade is not supported. The previous build ignores `content_type` and would show typed JSON
|
|
as ordinary reply text. If drained, it would acknowledge the message and lose typed meaning. Typed
|
|
queues must be drained or preserved before an old jar runs.
|
|
|
|
The existing contract suite uses RabbitMQ. Production uses LavinMQ. The migration and lead-key
|
|
ownership cases must run once against production LavinMQ before release, or the release must state
|
|
that LavinMQ was not checked.
|
|
|
|
### 7.3 Member routing
|
|
|
|
A member incident first goes to the exact lead that owns its accepted delegation.
|
|
`PrimaryRegistry` needs a no-fallback `delegatingLeadFor(memberTarget)` query. Health routing must not
|
|
use the old singular-primary fallback when several leads exist.
|
|
|
|
Publish the incident under the affected member target. Trigger the existing bounded push route. The
|
|
push waits until the owning lead is injectable, so it does not interrupt a live lead turn.
|
|
|
|
### 7.4 Peer lead routing
|
|
|
|
Figure 2 shows the route from incident to lead, peer, or person.
|
|
|
|
```mermaid
|
|
flowchart TD
|
|
Incident["Open incident"] --> Member{"Member incident?"}
|
|
Member -->|"yes"| Known{"Exact delegation owner known?"}
|
|
Known -->|"no"| Sink{"Human webhook enabled and healthy?"}
|
|
Known -->|"yes"| Owner{"Owner lead healthy?"}
|
|
Owner -->|"yes"| OwnerInbox["Publish to owner lead path"]
|
|
Owner -->|"no"| PeerSet["Build healthy peer candidate set"]
|
|
Member -->|"no, lead incident"| PeerSet
|
|
PeerSet --> Peer{"Healthy peer exists?"}
|
|
Peer -->|"yes"| Select["Choose fewest assigned incidents<br/>then stable name and terminal id"]
|
|
Select --> PeerInbox["Publish to peer lead inbox<br/>and status-gated push"]
|
|
Peer -->|"no"| Sink
|
|
Sink -->|"yes"| Webhook["Send classified outbound incident"]
|
|
Sink -->|"no"| Passive["Keep incident open<br/>show partial coverage on local surfaces"]
|
|
```
|
|
|
|
*Figure 2. Routing keeps delegation ownership separate from temporary peer fallback.*
|
|
|
|
Peer candidates exclude the incident subject, failed owner, absent leads, raw-unknown leads, and
|
|
leads with an open unhealthy state. A reachable `WORKING` peer may be selected; its push waits for an
|
|
injectable window.
|
|
|
|
Choose the candidate with the fewest assigned foreign incidents. Break ties by stable lead name,
|
|
then terminal id. Pin the recipient. Reassign only if that peer becomes unhealthy or retires. A
|
|
routing generation marks a reassignment, and old pending assignments become superseded.
|
|
|
|
`fleet_list` lead rows show health, health age, assigned foreign incident count, and a bounded list
|
|
of incident id, subject, state, severity, age, and routing generation. The top-level view also shows
|
|
owner, recipient, and routing reason.
|
|
|
|
A peer incident is published under the recipient lead's inbox key, not the failed subject's key. Its
|
|
status-gated nudge names the failed lead and gives the exact
|
|
`fleet_poll(target="<recipient-terminal>")` call.
|
|
|
|
### 7.5 Lead inbox ownership
|
|
|
|
Add `LeadInboxRegistry`, driven by the same expected-lead supplier as `CallerResolver`.
|
|
|
|
It calls `replyInbox.own(leadTerminal)` at startup for configured leads, after successful discovery,
|
|
after config adds a lead, and before publication. Ownership is not an authorization side effect.
|
|
|
|
A missing lead keeps its key owned. Release happens only after confirmed retirement, all incidents
|
|
are reassigned or resolved, typed health messages move or ack, and the queue is empty. Own a
|
|
replacement terminal before moving messages from the old key. Never release a non-empty in-memory
|
|
lead key, because in-memory release clears local data.
|
|
|
|
### 7.6 Single-lead deployment
|
|
|
|
One lead and no peer is a normal mode, not an edge case.
|
|
|
|
An idle, reachable lead may receive the existing bounded nudge. An unreachable or stalled sole lead
|
|
has no safe in-loop recovery. The bridge must not restart or replace it. A new lead would not have the
|
|
failed lead's plan or context, and an uncertain relaunch could create two orchestrators.
|
|
|
|
With no webhook, only `fleet_list`, `/healthz`, metrics, WARN logs, and the incident journal remain.
|
|
These are passive surfaces. They are not a human notification.
|
|
|
|
## 8. Lost-boundary reconciliation
|
|
|
|
This is the only M4 path that reconstructs a result. It must prefer a visible stall over a fabricated
|
|
reply.
|
|
|
|
### 8.1 Why normal completion rules are not enough
|
|
|
|
Current `CompletionResolver.resolve` has two fail-open rules. It resolves when the delivery baseline
|
|
is missing. It also resolves an empty completion when the pane read fails. Those choices are valid
|
|
after a trusted `WORKING -> IDLE` boundary because the bridge knows the turn ran. They are unsafe
|
|
when health only guesses that a boundary was lost.
|
|
|
|
M4 gives each accepted send an internal `TurnToken`. It ties target, exact waiter, session turn,
|
|
delivery baseline, and task outcome together.
|
|
|
|
### 8.2 Delivery baseline
|
|
|
|
Capture the baseline immediately after prompt send and before the delivery future completes. Store:
|
|
|
|
```text
|
|
TurnToken
|
|
exact waiter identity
|
|
capture time and pane source
|
|
normalised assistant block clipped to MAX_SCRAPE_CHARS
|
|
whether a supported assistant marker was recognised
|
|
capture result: PRESENT | READ_FAILED | UNRECOGNISED
|
|
```
|
|
|
|
A failed or missing baseline never authorises repair. A late baseline is not valid evidence. After a
|
|
daemon restart, the old waiter, task, token, and baseline are gone, so the old turn cannot be
|
|
repaired.
|
|
|
|
Automatic repair is enabled only for agent kinds with tested assistant-block fixtures. Current
|
|
extraction is Claude Code-specific and falls back to arbitrary raw text without `⏺`. That raw fallback
|
|
cannot authorise repair. OpenCode repair stays disabled until live pane fixtures exist.
|
|
|
|
### 8.3 Strict gates and resolver result
|
|
|
|
Figure 3 shows the repair gates. Any failed gate keeps the waiter unchanged.
|
|
|
|
The two raw snapshots must describe the same `TurnToken` and session turn. No `WORKING`,
|
|
`BLOCKED`, `UNKNOWN`, missing-agent, reply, failure, or new-delivery observation may occur between
|
|
them.
|
|
|
|
```mermaid
|
|
flowchart TD
|
|
Candidate["TURN_BOUNDARY_LOST candidate"] --> Stable{"Same TurnToken and BUSY turn<br/>across two raw IDLE or DONE snapshots?"}
|
|
Stable -->|"no"| Resnapshot["Take a fresh snapshot"]
|
|
Stable -->|"yes"| Waiter{"Exact captured waiter<br/>still open by identity?"}
|
|
Waiter -->|"no"| Stale["STALE_TURN or ALREADY_RESOLVED"]
|
|
Waiter -->|"yes"| Baseline{"Successful recognised<br/>delivery baseline exists?"}
|
|
Baseline -->|"no"| Refuse["Refuse repair<br/>leave ticket pending"]
|
|
Baseline -->|"yes"| Read{"Fresh pane read succeeds?"}
|
|
Read -->|"no"| Refuse
|
|
Read -->|"yes"| Output{"Recognised non-blank assistant block<br/>differs from clipped baseline?"}
|
|
Output -->|"no"| Refuse
|
|
Output -->|"yes"| Resolve["Shared CompletionResolver guard core<br/>resolves exact waiter"]
|
|
Resolve -->|"won race"| Repaired["RECONCILED_COMPLETION<br/>same turn becomes DONE"]
|
|
Resolve -->|"lost race"| Resnapshot
|
|
```
|
|
|
|
*Figure 3. Repair needs stronger evidence than a normal observed turn boundary.*
|
|
|
|
Refactor the current resolver into one guard core with two policies:
|
|
|
|
```text
|
|
resolveCaptured(target, inFlight, OBSERVED_BOUNDARY)
|
|
resolveCaptured(target, inFlight, LOST_BOUNDARY_REPAIR)
|
|
```
|
|
|
|
The health monitor calls only:
|
|
|
|
```text
|
|
CompletionResolver.reconcileLostBoundary(target, expectedTurnToken)
|
|
```
|
|
|
|
It returns `REPAIRED`, `ALREADY_RESOLVED`, `REFUSED_NO_CAPTURE`, `REFUSED_NO_BASELINE`,
|
|
`REFUSED_UNREADABLE`, `REFUSED_UNCHANGED`, `REFUSED_AMBIGUOUS_OUTPUT`, `STALE_TURN`, or
|
|
`RACE_LOST`.
|
|
|
|
Only `REPAIRED` and same-turn `ALREADY_RESOLVED` may move that turn from `BUSY` to `DONE`. A
|
|
per-target reconciliation gate stops a queued second send from being accepted between waiter
|
|
resolution and the FSM transition.
|
|
|
|
### 8.4 Lead-visible marker and refusal
|
|
|
|
A repaired result uses distinct `RECONCILED_COMPLETION` values in `Rendezvous`, `MessageService`,
|
|
task poll source, and metrics. The lead sees:
|
|
|
|
```text
|
|
[repaired completion - fleetd detected a lost turn boundary. The member did not call
|
|
fleet_reply; pane-derived text follows and may be partial]
|
|
```
|
|
|
|
Clipped text also keeps the existing clipped-tail marker.
|
|
|
|
A refused repair leaves `TURN_BOUNDARY_LOST` open and the ticket pending. The report states that no
|
|
reply was reconstructed and no task was replayed. `UNCHANGED`, `UNREADABLE`, and
|
|
`AMBIGUOUS_OUTPUT` get at most one delayed retry for the same token. Missing capture or baseline gets
|
|
no retry. After two refused scrapes, automatic repair stops for that token.
|
|
|
|
### 8.5 Target-wide teardown invariant
|
|
|
|
CB-568 owns the multi-ticket cancellation mechanism. M4 routes every terminal cause through that one
|
|
idempotent operation and checks this independent invariant after teardown:
|
|
|
|
- no injector entry exists for the target;
|
|
- no accepted turn or completion record exists;
|
|
- no rendezvous waiter or ask exists;
|
|
- every async task is terminal or was already terminal;
|
|
- no thread waiting for the target send lock can later accept it;
|
|
- new sends fail immediately;
|
|
- each old task has one terminal outcome and one metric count.
|
|
|
|
A violation becomes `DELEGATION_ORPHANED`. The monitor may call the same idempotent target-wide
|
|
failure operation once. It never recreates the task.
|
|
|
|
## 9. Capacity and utilisation
|
|
|
|
Capacity is a view, not a health state.
|
|
|
|
`fleet_list` adds one block per profile:
|
|
|
|
```text
|
|
profile, maxLoad, live, free, reclaimable
|
|
```
|
|
|
|
For an unlimited profile, `maxLoad` and `free` are null. `free` is
|
|
`max(0, maxLoad - live)` for a capped profile.
|
|
|
|
The view must use the exact live-count function used by placement. A second calculation could show a
|
|
free slot that placement then refuses. Member rows add `idleForSeconds` only when state is `READY` or
|
|
`DONE`, no accepted turn exists, and the inbox is empty. `reclaimable` means only that the member
|
|
holds capacity without open bridge work.
|
|
|
|
The existing idle-lead nudge gains a bounded capacity summary. It lists per-profile live, cap, free,
|
|
and reclaimable counts, plus at most three long-idle members. Capacity does not make
|
|
`FleetState.hasPending()` true. A changed capacity fingerprint may re-arm one capped heartbeat
|
|
sequence. The fingerprint excludes changing idle durations, so a static idle fleet cannot reset the
|
|
cap forever. Reply-push stand-down remains first.
|
|
|
|
The bridge must never:
|
|
|
|
- spawn a member because a slot is free;
|
|
- generate a task or acceptance criteria;
|
|
- move queued work to another member or profile;
|
|
- treat a free slot or idle member as an incident;
|
|
- stop an idle member only to improve utilisation.
|
|
|
|
The bridge knows capacity facts but has no work list. Only the lead has the plan, task context,
|
|
side-effect history, and acceptance criteria.
|
|
|
|
Capacity calculation is in memory and adds no pane reads. Work-product checks run on a terminal
|
|
session edge, not every fleet tick.
|
|
|
|
This capacity design adds no automatic stop. The accepted `NEVER_READY` cleanup can still stop a
|
|
very slow startup after the existing grace, which is a known risk. Free capacity and long idle time
|
|
never trigger that path.
|
|
|
|
## 10. Human escalation and notification
|
|
|
|
### 10.1 Escalation rule
|
|
|
|
Notify a person only when no healthy lead can act:
|
|
|
|
- `CONTROL_LINK_DOWN` survives grace;
|
|
- a lead is unhealthy and no healthy peer can receive the incident;
|
|
- a member incident has no known owning lead;
|
|
- the only owning lead becomes unreachable, unresponsive, or stalled;
|
|
- incident publication or routing itself fails.
|
|
|
|
Do not page a person for a member fault while a healthy owning lead exists. An uncollected member
|
|
incident feeds lead-health evidence. If the lead then becomes unhealthy, peer or human routing starts.
|
|
|
|
### 10.2 Detection and notification switches
|
|
|
|
`health.enabled` controls detection and bridge-local reporting. It does not require a webhook.
|
|
|
|
`health.notifications.mode` is `disabled` or `webhook`. Disabled is valid and is the default.
|
|
Webhook mode requires a resolved environment variable. Turning notification off stops outbound
|
|
attempts but keeps incidents. Turning it back on resumes still-open human incidents.
|
|
|
|
Without a sink, `fleet_list.healthCoverage` states that human escalation is unavailable. `/healthz`
|
|
keeps its existing HTTP liveness result and adds a nested `fleetHealth.status=partial` component.
|
|
Metrics and one startup or reload WARN expose the same limit.
|
|
|
|
### 10.3 Incident and delivery deduplication
|
|
|
|
One open incident uses this key:
|
|
|
|
```text
|
|
(scope, subjectStableId, state, causeFingerprint)
|
|
```
|
|
|
|
The cause fingerprint includes stable error codes, dependency names, signature ids, or invariant
|
|
names. It excludes times, ages, retry counts, pane text, and changing digests. A later recurrence
|
|
after resolution gets a new generation and incident id.
|
|
|
|
Each outbound event uses:
|
|
|
|
```text
|
|
Idempotency-Key = hash(incidentId, eventType, eventRevision)
|
|
```
|
|
|
|
Event types are `open`, `severity_changed`, `reminder`, and `resolved`. Transport retries keep the
|
|
same key.
|
|
|
|
An atomic owner-only journal beside the active config stores open incidents, routing, delivered
|
|
revisions, retry state, and resolution state. It stores no pane or task content. Journal failure does
|
|
not stop detection, but notification coverage becomes degraded.
|
|
|
|
### 10.4 Retry, reminder, and resolve
|
|
|
|
Send the first event immediately. Retry network errors, timeouts, HTTP 408, HTTP 429, and HTTP 5xx
|
|
with full-jitter exponential backoff:
|
|
|
|
```text
|
|
base: 5 seconds
|
|
factor: 3
|
|
maximum delay: 15 minutes
|
|
one outstanding attempt per event
|
|
```
|
|
|
|
Respect `Retry-After` up to 15 minutes. Other HTTP 4xx responses are permanent for that event until
|
|
config changes or a person requests replay.
|
|
|
|
Transport retry is not an incident reminder. `humanRepeatSeconds` creates a new reminder revision
|
|
for an unresolved critical incident after the last successful human event. Disabled mode does not
|
|
build an unbounded reminder queue.
|
|
|
|
Send `resolved` only if at least one human event for that incident was delivered. If an incident
|
|
resolves before its first successful delivery, cancel the pending open event and record local
|
|
resolution.
|
|
|
|
### 10.5 Outbound payload boundary
|
|
|
|
An outbound payload may contain incident id and event type, severity, state, scope, stable bridge
|
|
ids, role or profile, times, duration, structured evidence type and counts, recovery attempted,
|
|
routing reason, coverage, and safe tool calls.
|
|
|
|
It must never contain:
|
|
|
|
- raw pane text, pane excerpts, or pane digests;
|
|
- task briefs, prompts, or member reply content;
|
|
- source files, diffs, or worktree file content;
|
|
- worktree paths;
|
|
- environment values, tokens, credentials, headers, or webhook URL;
|
|
- raw exception messages or stack traces;
|
|
- arbitrary model output.
|
|
|
|
The sink response body is ignored. A webhook cannot direct recovery. n8n remains outbound-only.
|
|
|
|
### 10.6 Metrics
|
|
|
|
M4 adds bounded-label series:
|
|
|
|
```text
|
|
fleet_health_incidents{scope,state,severity}
|
|
fleet_health_incidents_total{event}
|
|
fleet_health_notifications_total{event,outcome}
|
|
fleet_health_notification_queue_depth
|
|
fleet_health_notification_last_success_seconds
|
|
fleet_health_notification_capability{mode,status}
|
|
fleet_lead_health{lead,state}
|
|
fleet_lead_assigned_incidents{lead}
|
|
```
|
|
|
|
Metric labels never include terminal ids, incident ids, URLs, or error text.
|
|
|
|
## 11. Configuration
|
|
|
|
The optional `health:` block is absent or disabled by default. The dormant monitor scheduler does no
|
|
herdr or pane work while disabled. Every listed key is hot because the monitor reads `ConfigRef` on
|
|
each tick or notification.
|
|
|
|
| Key | Class | Default and hard bound | Purpose |
|
|
|---|---|---|---|
|
|
| `health.enabled` | Hot | `false` | Enable detection and bridge-local reporting. |
|
|
| `health.snapshotIntervalSeconds` | Hot | default 30, minimum 15 | Whole-fleet comparison cadence. |
|
|
| `health.workingSuspectAfterSeconds` | Hot | default 600, minimum 300 | Age before working-pane probes. |
|
|
| `health.paneProbeIntervalSeconds` | Hot | default 60, minimum 60 | Per-target pane cooldown. |
|
|
| `health.leadUnresponsiveAfterSeconds` | Hot | default 300, minimum 120 | Delay after exhausted actionable nudges before lead fault. |
|
|
| `health.humanRepeatSeconds` | Hot | default 3600, minimum 900 | Minimum repeat period for one open human incident. |
|
|
| `health.capacityLongIdleAfterSeconds` | Hot | default 900, minimum 300 | Long-idle threshold for capacity summaries. |
|
|
| `health.includePaneExcerpt` | Hot | `false` | Allow a clipped excerpt in local lead reports only. Human payloads still exclude it. |
|
|
| `health.notifications.mode` | Hot | `disabled` | Select `disabled` or `webhook`. |
|
|
| `health.notifications.webhookUrlEnv` | Hot | required in webhook mode | Name of the environment variable that holds the sink URL. |
|
|
| `health.notifications.requestTimeoutMs` | Hot | default 10000, range 1000-30000 | Whole webhook request limit. |
|
|
|
|
Two consecutive snapshots are compiled floors for lost boundary, lead disappearance, and control
|
|
link failure. The two-pane-reads-per-tick limit is also compiled and cannot be weakened by config.
|
|
|
|
## 12. Delivery units and acceptance
|
|
|
|
### Unit 1 - Evidence model and fleet snapshot
|
|
|
|
Scope: health state model, fleet join, clocks, evidence retention, and pane budget.
|
|
|
|
Acceptance criteria:
|
|
|
|
1. One `agent.list` call covers one enabled fleet tick.
|
|
2. Pure decision tests cover every state and every evidence limit in Section 4.
|
|
3. `BUSY` plus stable raw `DONE` opens `TURN_BOUNDARY_LOST` after two snapshots.
|
|
4. Healthy fleet snapshots perform zero pane reads.
|
|
5. Pane cooldown, two-read fleet budget, and fair rotation cannot be disabled by config.
|
|
6. Logs are outputs only; no log parsing exists.
|
|
7. Fleet snapshots expose the same profile live-count calculation that placement uses.
|
|
8. Capacity rows report cap, live, free, and reclaimable values without opening incidents.
|
|
|
|
### Unit 2 - Lost boundary and task reconciliation
|
|
|
|
Scope: accepted-turn identity, guarded repair, target-wide teardown, release causes, and preserved
|
|
worktree discovery.
|
|
|
|
Acceptance criteria:
|
|
|
|
1. Every accepted send receives a stable `TurnToken` tied to target, exact waiter, and delivery
|
|
baseline.
|
|
|
|
**Corrected during implementation (2026-08-15).** This criterion first also required the session
|
|
turn number and the task outcome. That is not implementable at this layer, and the implementer
|
|
refused it three times rather than fabricate a value — correctly. The reason is an ordering fact
|
|
that is invisible from any single class: `MessageService` owns acceptance and holds the waiter and
|
|
the async `Task`, but it learns nothing about delivery, because the delivery event goes to
|
|
`CompletionResolver` through `TurnListener.onDelivered`. And `CompletionResolver.onDelivered` runs
|
|
*before* `SessionManager.onDelivered`, so the session turn number does not exist yet at the only
|
|
point where the token could capture it.
|
|
|
|
Two ways out were rejected. A shared registry keyed by target reintroduces exactly the "whichever
|
|
send happens to be waiting" ambiguity the token exists to remove — the same weak claim
|
|
`Rendezvous.currentWaiter` warns about. Injecting a turn counter into `MessageService` adds a
|
|
required cross-layer dependency to populate a field that nothing in this slice reads, which is
|
|
speculative coupling across a boundary already shown to be fragile.
|
|
|
|
So the token identifies the **accepted send**, and `SessionManager` keeps verifying its own
|
|
delivery separately. Repair (criterion 2) does need the session turn; binding it means resolving
|
|
that acceptance-versus-delivery ordering first, and that work belongs to the repair unit, not
|
|
here. The token record carries a comment saying the field is deliberately absent.
|
|
2. Repair requires the same `BUSY` token, two raw `IDLE` or `DONE` snapshots, no conflicting
|
|
observation, exact open waiter, successful baseline, and new recognised assistant output.
|
|
3. Missing, failed, late, or post-restart baseline never authorises repair.
|
|
4. Repair is enabled only for agent kinds with tested assistant-block extraction. Raw-text fallback
|
|
without a recognised marker refuses repair.
|
|
5. Normal completion and repair use one resolver guard core. Waiter, scrape, clipping, unchanged, and
|
|
exact-turn guards are not duplicated.
|
|
6. `reconcileLostBoundary` returns every typed result named in Section 8.3.
|
|
7. Only `REPAIRED` and same-turn `ALREADY_RESOLVED` may move the same turn to `DONE`.
|
|
8. A per-target reconciliation gate blocks a queued second send during repair and FSM update.
|
|
9. Repaired completion has distinct rendezvous kind, message outcome, poll source, lead marker, and
|
|
metric. Clipping keeps its extra marker.
|
|
10. Unchanged, unreadable, or ambiguous evidence gets at most one delayed retry. Missing capture or
|
|
baseline gets none.
|
|
11. Refusal leaves the ticket pending and tells the lead that no result was rebuilt or replayed.
|
|
12. Release, gone, never-ready, and abnormal stop use CB-568's one idempotent target-wide failure
|
|
operation.
|
|
13. The post-teardown invariant in Section 8.5 is tested independently of CB-568 internals.
|
|
14. A violated teardown invariant creates `DELEGATION_ORPHANED` and retries only the idempotent
|
|
failure operation.
|
|
15. `SPAWN_ROLLBACK` and normal `COMPLETED` remove worktrees. Abnormal and shutdown causes preserve
|
|
them.
|
|
16. Explicit stop is state-aware. Any pending task or non-terminal state preserves the worktree.
|
|
17. Atomic preserved-worktree manifests reload after restart and appear in lead-only
|
|
`fleet_list.preservedWorktrees`.
|
|
18. Manifest failure preserves the worktree and opens an operator-visible health failure.
|
|
19. Provision records the base commit. Terminal, long-idle worktrees report
|
|
`WORK_PRODUCT_AT_RISK` only under the evidence in Section 4.2 and never auto-delete work.
|
|
20. No path replays a delivered task, rebuilds its brief, or retargets it, even when prompt text is
|
|
available.
|
|
21. Tests cover both real traces, all repair refusals, clipping, explicit-reply and next-turn races,
|
|
restart without capture, concurrent send and release, and preserved discovery after restart.
|
|
|
|
#### Unit 2 - what has landed so far
|
|
|
|
Checked against `main` at `e09cac6` on 2026-08-15. Unit 2 was written as one block, but parts of it
|
|
have since been built by separate CB tickets. Read this before planning the rest, or that work gets
|
|
done twice.
|
|
|
|
The check was a symbol survey of `fleetd/src/main/java` plus the merge history. It tells you whether
|
|
the machinery exists at all. It is **not** a line-by-line audit of whether each criterion is fully
|
|
met, and I did not run one.
|
|
|
|
| Criterion | Marker searched for | Found in main source | Reading |
|
|
|---|---|---|---|
|
|
| 1 | `TurnToken` | 8 files | **Done** — unit 2a, merged as `fec284e`. Criterion 1 was corrected first; see the note under it. |
|
|
| 2-5, 9 | `REPAIRED` | 0 files | Not started. The whole guarded-repair path is absent. |
|
|
| 6, 7, 10 | `reconcileLostBoundary` | 0 files | Not started. |
|
|
| 12 | CB-568 failure operation | via CB-580 | **Partial.** CB-580 (`0af902e`) routes `GONE` and `NEVER_READY` into the one idempotent target-wide failure. I did not check that release and abnormal stop go through the same call. |
|
|
| 14 | `DELEGATION_ORPHANED` | 3 files | **Partial.** The health state exists. The teardown-invariant check that creates it, and the retry rule, do not. |
|
|
| 15 | `SPAWN_ROLLBACK` | 0 files | **Contradicted — see below.** |
|
|
| 16 | — | — | Partial at best. CB-576 made release preserve a dirty worktree; whether explicit stop is state-aware is not checked. |
|
|
| 17, 18 | `preservedWorktrees` | 0 files | Not started. No manifest, and no lead-only `fleet_list` field. |
|
|
| 19 | `WORK_PRODUCT_AT_RISK` | 0 files | Not started. |
|
|
|
|
**Criterion 15 no longer matches the code, and the code is right.** It says "normal `COMPLETED`
|
|
remove worktrees". Since CB-576 (`500bfa2`) that is false on purpose: a `COMPLETED` release now
|
|
preserves the worktree when it still holds uncommitted work, because deleting it destroys work
|
|
nobody can get back. CB-576 was filed after exactly that loss. CB-581 goes further — if the
|
|
dirty-check itself fails, the worktree is preserved rather than removed, since "we could not tell"
|
|
must not be treated as "it is clean".
|
|
|
|
So criterion 15 should be rewritten as: `SPAWN_ROLLBACK` and a `COMPLETED` release with a **clean**
|
|
worktree remove it; abnormal causes, shutdown, a dirty worktree, and a failed dirty-check all
|
|
preserve it. `SPAWN_ROLLBACK` itself does not exist yet.
|
|
|
|
### Unit 3 - Typed inbox and member routing
|
|
|
|
Scope: semantic record, AMQP migration, both adapters, member routing, polling, and member health in
|
|
`fleet_list`.
|
|
|
|
Acceptance criteria:
|
|
|
|
1. AMQP selects legacy or typed decoding only from `content_type`; it never sniffs the body.
|
|
2. Persistent `text/plain` from the old build becomes `kind=reply` with exact UTF-8 content,
|
|
including content beginning with `{`.
|
|
3. New entries use the vendor media type, `schemaVersion: 1`, UTF-8, persistent delivery, and AMQP
|
|
message ids.
|
|
4. Version 1 ignores unknown optional fields but rejects missing fields and identity mismatch.
|
|
5. Unknown versions are not decoded or acked. They remain on the original queue and create one
|
|
deduplicated failure.
|
|
6. Invalid known data never escapes the callback, appears as a reply, or blocks later valid messages.
|
|
7. Invalid data reaches durable per-target quarantine before original ack. Failed handoff leaves the
|
|
original unacked.
|
|
8. Decode failures create redacted WARN, metric, `fleet_list` summary, and routed incident without
|
|
raw content.
|
|
9. Both adapters pass one semantic contract for fields, FIFO, dedup, ownership, ack, and release.
|
|
10. Lead keys require explicit ownership. Publication never claims a queue.
|
|
11. Unit codec tests cover legacy `{`, Unicode, malformed UTF-8, typed round trip, additive fields,
|
|
malformed JSON, missing fields, identity mismatch, media type, version, and dedup.
|
|
12. A live broker contract writes old wire data and reads it with the new adapter after reconnect.
|
|
13. Live contract tests cover mixed entries, quarantine confirm-before-ack, unsupported redelivery,
|
|
later progress past poison, property persistence, lead ownership, and ack removal.
|
|
14. Safe downgrade is documented as unsupported.
|
|
15. RabbitMQ contract tests pass with `mvn test -Pcontract`. The same cases run once on production
|
|
LavinMQ, or the release states that LavinMQ was not checked.
|
|
16. Member incidents route to the exact delegating lead and never resolve a task rendezvous.
|
|
17. `fleet_list` shows compact member health and capacity without pane content. Member
|
|
`idleForSeconds` is present only when no accepted turn or inbox item exists.
|
|
|
|
### Unit 4 - Lead health and peer routing
|
|
|
|
Scope: lead evidence, exact ownership, peer selection, explicit-recipient push, and lead inbox
|
|
lifecycle.
|
|
|
|
Acceptance criteria:
|
|
|
|
1. Lead identity uses the `CallerResolver` supplier. Liveness uses successful current agent data.
|
|
2. Two successful-list absences with healthy ping become `LEAD_UNREACHABLE`; global link failure does
|
|
not.
|
|
3. Raw `WORKING`, raw `UNKNOWN`, first-seen time, failures, last success, and error class persist
|
|
across ticks.
|
|
4. Heartbeat and push publish status and nudge outcomes before safe no-injection decisions.
|
|
5. Dynamic lead identity survives a two-successful-snapshot retirement grace.
|
|
6. Member incidents first use exact delegation ownership with no singular-primary fallback.
|
|
7. Peer selection follows the exclusions, load rule, and stable tie break in Section 7.4.
|
|
8. A selected working peer is not interrupted. Its push waits for an injectable window.
|
|
9. Recipient assignment stays pinned. Reassignment increments generation and supersedes old pending
|
|
assignment.
|
|
10. `fleet_list` shows bounded foreign assignments, recipient, reason, and generation without pane
|
|
content.
|
|
11. `LeadInboxRegistry` owns configured and discovered lead keys before publication.
|
|
12. Missing leads keep ownership. Retirement needs an empty queue and handled incidents.
|
|
13. Replacement owns the new key before messages move. Non-empty in-memory keys are not released.
|
|
14. Tests cover dead versus busy, unknown, global failure, stale scan cache, disappearance, one peer,
|
|
several peers, reassignment, and no peer.
|
|
15. Adapter tests cover lead ownership, restart re-ownership, retirement, and terminal replacement.
|
|
LavinMQ is checked or named as unchecked.
|
|
16. A sole unreachable or stalled lead is never restarted or replaced. Without a sink, only passive
|
|
evidence remains and every coverage surface says so.
|
|
|
|
### Unit 5 - Human sink, hot config, metrics, and operator coverage
|
|
|
|
Scope: generic webhook, config split, incident journal, retry, resolve, metrics, example config, and
|
|
operator documentation.
|
|
|
|
Acceptance criteria:
|
|
|
|
1. `health.enabled` works without a human sink.
|
|
2. Notification mode is hot, defaults to disabled, and supports disabled or webhook.
|
|
3. Webhook mode requires a resolved environment value. Bad notification config does not disable an
|
|
already valid detector.
|
|
4. Mode changes keep open incidents. Re-enable resumes eligible incidents.
|
|
5. `fleet_list`, `/healthz`, metrics, and one WARN show partial coverage without a sink. HTTP
|
|
liveness behavior stays unchanged.
|
|
6. One-lead, no-sink coverage states that lead failure has no active notification or recovery.
|
|
7. Incident and outbound dedupe use the stable keys in Section 10.3.
|
|
8. The owner-only local journal survives restart and contains no pane or task content.
|
|
9. Journal failure keeps detection running but marks notification coverage degraded.
|
|
10. Retry tests cover network failure, timeout, 408, 429, `Retry-After`, 5xx, permanent 4xx, jitter,
|
|
delay cap, config re-arm, and one outstanding attempt.
|
|
11. Reminders and transport retries remain separate. Disabled mode does not build an unbounded queue.
|
|
12. Resolve sends only after an earlier human event succeeded. Resolve-before-delivery cancels stale
|
|
open delivery.
|
|
13. Metrics use bounded labels and exclude ids, URLs, and error text.
|
|
14. Payload tests reject every content type forbidden in Section 10.5.
|
|
15. Webhook response bodies are ignored and cannot direct recovery.
|
|
16. Tests cover disabled mode, one lead without sink, open/update/reminder/resolve, restart, dedup,
|
|
reassignment, disable/re-enable, and sink failure while local health continues.
|
|
17. `fleetd.example.yaml` documents all hot keys and compiled floors.
|
|
18. The operator Features wiki is updated separately. The portable `CLAUDE.md` block is checked and
|
|
changed only if shipped tool or inbox semantics make it untrue.
|
|
19. `mvn clean install` passes.
|
|
|
|
## 13. Not checked and release gates
|
|
|
|
These limits are part of the design, not optional follow-up notes.
|
|
|
|
- **OpenCode pane status and assistant markers were not checked.** OpenCode lost-boundary repair is
|
|
disabled until live fixtures exist.
|
|
- **Permission-prompt status was not checked** for Claude Code or OpenCode. `BLOCKED` remains
|
|
ambiguous and has no automatic action.
|
|
- **`recent_unwrapped` stability was not checked** across all supported agent kinds. If normalisation
|
|
is not stable, `STALL_SUSPECTED` must say its evidence is weaker.
|
|
- **The real `BUSY + DONE` trace was not replayed against live herdr.** The design uses the observed
|
|
production trace and current poller behavior.
|
|
- **CB-568 was not present when Unit 2 was designed.** Unit 2 must inspect the landed API and keep its
|
|
independent teardown invariant.
|
|
- **Production LavinMQ was not checked.** Existing durable-inbox contracts use RabbitMQ. Migration,
|
|
quarantine, redelivery, lead ownership, and reassignment must run on LavinMQ before release or be
|
|
recorded as unchecked.
|
|
- **Live multi-lead routing was not checked.** Peer choice and reassignment are design rules backed by
|
|
fake-clock and adapter tests until a live exercise runs.
|
|
- **A live sole-lead failure with a webhook was not checked.** The no-peer path is a design result,
|
|
not a tested recovery.
|
|
- **No n8n, Slack, PagerDuty, or other receiver was checked.** The webhook remains generic and
|
|
outbound-only.
|
|
- **Deployment supervisor behavior for nested `/healthz` fields was not checked.** HTTP liveness
|
|
status stays unchanged to reduce this risk.
|
|
- **Incident-journal crash behavior was not checked** because the journal does not exist yet. Unit 5
|
|
must test atomic replacement and restart recovery.
|
|
- **Worktree merge state cannot be checked reliably** without forge or explicit collection evidence.
|
|
`WORK_PRODUCT_AT_RISK` stays a warning.
|
|
|
|
## 14. Locked exclusions
|
|
|
|
M4 does not expose `agent.read` as a bridge tool. It does not add a workflow engine, inbound n8n
|
|
authority, automatic task assignment, task replay, automatic lead replacement, or automatic member
|
|
spawn for free capacity.
|
|
|
|
The bridge remains a message bus with evidence and bounded mechanical repair. The lead remains the
|
|
place where judgement and work planning happen.
|