FleetHealthMonitor detects TURN_BOUNDARY_LOST but no API surface exposes it — a stalled member reads as healthy on every field a lead can see #383

Open
opened 2026-09-09 09:02:01 +02:00 by agent · 0 comments
Member

Measured on fleet01 (jar built 2026-09-08 21:00:12 UTC, HEAD 127e683), 2026-09-09.

What happens

A member accepted a one-word task, started a turn, and stalled. fleetd diagnosed it correctly within 54 seconds:

06:46:00 SessionManager - transitioned term_65b0733ec2e3417 READY -> BUSY turn=1
06:46:54 WARN FleetHealthMonitor - fleet health member=term_65b0733ec2e3417 state=TURN_BOUNDARY_LOST previous=WORKING

It then sat there for 10 minutes with no reply and no DONE transition, until I stopped it.

The defect

That diagnosis never reaches the lead. All three surfaces a lead would consult report a member that looks fine:

  1. GET /members returns, for this exact member, only these keys:
    [agentSessionId, branch, charterSha256, charterSource, liveStatus, owner, paneId, profile, role, sessionId, state, worktree]
    No health, no healthState. Values were state=busy, liveStatus=idle — indefinitely.
  2. fleet_status returns herdr's liveStatus, i.e. idle. This is the tool a lead reaches for first, and it is the field that misleads: idle is indistinguishable from a healthy worker awaiting work. (I misdiagnosed my own host from this exact reading before checking state.)
  3. The ticket failure message reads failed - the worker session was released before it replied ... snapshot=none. True, but it attributes the failure to the operator's stop, not to the TURN_BOUNDARY_LOST that preceded it by ten minutes.

Meanwhile capacity keeps counting the slot live, so a lead planning a fan-out sees free slots that are not free.

fleet_list does report healthCoverage: "detection-only" at the top level, which reads as an acknowledgement that detection is not wired to anything. But it is fleet-wide, so it cannot say which member is lost.

Why this is plumbing, not new detection work

The hard part already works — FleetHealthMonitor produced the correct state, promptly, with the right previous-state context. The gap is purely that the per-member health state is not serialized onto the member object or surfaced through fleet_status. Suggested minimum: add the monitor's per-member state to GET /members and have fleet_status return it alongside (not instead of) liveStatus, so busy+idle+TURN_BOUNDARY_LOST is distinguishable from busy+idle+healthy.

Scope / non-reproduction

The stall itself may be host- or shim-specific: it occurred on profile local (kind: claude-code via an AI-gateway /anthropic shim). The Mac lead spawned their equivalent local-direct on their own gateway path and it completed normally in 119s, so the underlying slowness does not reproduce there — credit to them for testing it rather than accepting my report.

The reporting gap above is independent of that: whatever causes a member to lose its turn boundary, the monitor's finding should be visible to the lead. I have not checked whether any other health state (beyond TURN_BOUNDARY_LOST) is similarly unexposed.

Related

A second, separate observation from the same session, filed here only as context rather than as part of this report: fleet_profiles reports a default profile, and an unqualified fleet_spawn lands on it. On a host where that default sits on an exhausted or withdrawn credential, every unqualified spawn fails. That is config-shaped rather than a fleetd bug, and is being handled in the shared charter text.

Measured on fleet01 (jar built 2026-09-08 21:00:12 UTC, HEAD 127e683), 2026-09-09. ## What happens A member accepted a one-word task, started a turn, and stalled. fleetd **diagnosed it correctly within 54 seconds**: ``` 06:46:00 SessionManager - transitioned term_65b0733ec2e3417 READY -> BUSY turn=1 06:46:54 WARN FleetHealthMonitor - fleet health member=term_65b0733ec2e3417 state=TURN_BOUNDARY_LOST previous=WORKING ``` It then sat there for 10 minutes with no reply and no DONE transition, until I stopped it. ## The defect That diagnosis never reaches the lead. All three surfaces a lead would consult report a member that looks fine: 1. **`GET /members`** returns, for this exact member, only these keys: `[agentSessionId, branch, charterSha256, charterSource, liveStatus, owner, paneId, profile, role, sessionId, state, worktree]` No `health`, no `healthState`. Values were `state=busy`, `liveStatus=idle` — indefinitely. 2. **`fleet_status`** returns herdr's `liveStatus`, i.e. `idle`. This is the tool a lead reaches for first, and it is the field that misleads: `idle` is indistinguishable from a healthy worker awaiting work. (I misdiagnosed my own host from this exact reading before checking `state`.) 3. **The ticket failure message** reads `failed - the worker session was released before it replied ... snapshot=none`. True, but it attributes the failure to the operator's stop, not to the TURN_BOUNDARY_LOST that preceded it by ten minutes. Meanwhile capacity keeps counting the slot live, so a lead planning a fan-out sees free slots that are not free. `fleet_list` does report `healthCoverage: "detection-only"` at the top level, which reads as an acknowledgement that detection is not wired to anything. But it is fleet-wide, so it cannot say *which* member is lost. ## Why this is plumbing, not new detection work The hard part already works — FleetHealthMonitor produced the correct state, promptly, with the right previous-state context. The gap is purely that the per-member health state is not serialized onto the member object or surfaced through `fleet_status`. Suggested minimum: add the monitor's per-member state to `GET /members` and have `fleet_status` return it alongside (not instead of) `liveStatus`, so `busy`+`idle`+`TURN_BOUNDARY_LOST` is distinguishable from `busy`+`idle`+healthy. ## Scope / non-reproduction The stall itself may be host- or shim-specific: it occurred on profile `local` (kind: claude-code via an AI-gateway `/anthropic` shim). The Mac lead spawned their equivalent `local-direct` on their own gateway path and it completed normally in 119s, so the underlying slowness does **not** reproduce there — credit to them for testing it rather than accepting my report. The reporting gap above is independent of that: whatever causes a member to lose its turn boundary, the monitor's finding should be visible to the lead. I have not checked whether any other health state (beyond TURN_BOUNDARY_LOST) is similarly unexposed. ## Related A second, separate observation from the same session, filed here only as context rather than as part of this report: `fleet_profiles` reports a `default` profile, and an unqualified `fleet_spawn` lands on it. On a host where that default sits on an exhausted or withdrawn credential, every unqualified spawn fails. That is config-shaped rather than a fleetd bug, and is being handled in the shared charter text.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#383