CB-615: there is no way to SEE the fleet — a monitoring dashboard for bridged #118

Open
opened 2026-08-22 08:52:49 +02:00 by ltms · 0 comments
Owner

Why this exists

The lead's only view of the fleet is the bridge_* MCP tools. That view has three holes, and all
three showed up in one session on 2026-08-22:

  1. A daemon restart cuts the lead's own MCP mount. The mount does not reconnect. After a
    redeploy the lead cannot call bridge_list or bridge_whoami at all, so the moment you most
    need to check the fleet is the moment you cannot.
  2. Placement is invisible. To answer "where will the next dev spawn land?" the lead had to read
    bridged.yaml, read WeightedRoundRobinPolicy.java, and do the ratio by hand. Nothing reports
    the effective share, and CB-589 already records that weighted is not cheapest-first — so the
    number people assume is wrong.
  3. A deferred config key is visible only as one log line. Adding the xf profile printed
    config reloaded; these changes need a restart to take effect: profiles (added/removed: xf)
    and nothing else. Miss that line and the file on disk and the running daemon disagree, silently.

There is no page anywhere that answers "what is the fleet doing right now".

What exists today, and where it stops

Surface Gives you Stops at
GET /healthz daemon up, herdr version + protocol says nothing about members; green while every spawn fails
GET /members the member roster flat JSON, no history, no timers
GET /sessions herdr workspaces and pane counts terminal view, not fleet view
GET /agents every pane herdr knows includes panes that are not members
GET /sessions/{id}/status one member's state one at a time
GET /tasks/{ticket} one async ticket you must already know the ticket id
GET /metrics one gauge family: bridged_sessions{state=…} no placement, cost, ticket, config or health metric at all
GET /profiles profile names + the default no weight, no maxLoad, no live count, no quarantine

/metrics is the honest measure of the gap. Checked on a live daemon, it emits 6 lines total.

flowchart LR
    A["bridged daemon"] --> B["SessionManager roster"]
    A --> C["CompositePeerLauncher<br/>placement + quarantine"]
    A --> D["ConfigRef<br/>live vs deferred"]
    A --> E["MessageBus<br/>tickets + reply inbox"]
    A --> F["herdr<br/>pane liveStatus"]
    B --> G["dashboard"]
    C --> G
    D --> G
    E --> G
    F --> G

The five sources a dashboard has to join. Each already exists in the daemon; none is exposed as a
view.

Features to build

A. Fleet roster — who is alive

  • Every member in one table: role, profile, model, state, pane, tab, worktree, uptime.
  • The branch the member actually committed to, read from its worktree — not the branch spawn
    provisioned. Those differ, and trusting the provisioned one makes you push an empty ref.
  • Time since the pane last produced output. A member can report working for hours while doing
    nothing. Status alone cannot tell you; last-output time can.
  • Countdown to the idle reaper (lifecycle.idleTtlSeconds, currently 1800). When it fires, the
    member's ticket, pane and report are gone together and there is no scrape fallback. The operator
    needs to see that clock before it runs out, not after.
  • Which lead owns which member.

B. Work in flight — what is owed to whom

  • Open tickets from wait:false sends, each with its TTL countdown. A ticket older than ten
    minutes has expired and the lead must send again; nothing tells you that today.
  • Undrained replies per session, as a count only. Reading the reply inbox is destructive
    (GET /sessions/{id}/replies drains on first read), so the dashboard must show that something is
    waiting without consuming it. This is a hard requirement, not a preference.
  • Open bridge_ask questions: the question, the turnId that answers it, and the remaining part of
    its ~55 second window. An ask that times out is invisible in every current surface.

C. Placement and cost — where the next spawn goes

  • The effective spawn ratio per role pool, computed live from weight, maxLoad, live count and
    quarantine. Not the configured weights — the ratio that will actually apply on the next call.
  • A plain answer to "if I spawn a dev right now, which profile gets it, and why that one".
  • Live count against maxLoad per profile, so an approaching PlacementException is visible before
    it throws.
  • Quarantine state per credential with the remaining cooldown (CB-578 stage B). Profiles sharing
    a credentialId must be shown as one group — sol and terra share openai-shared, and an
    exhaustion on either locks both.
  • Spawn counts split by free / paid / subscription, over time. This is the number that tells an
    operator whether the weights are doing what they think.

D. Config truth — is the running daemon the file on disk?

  • The deferred set: which keys are in bridged.yaml but not live, and therefore need a restart.
    ConfigRef.reloadDiff already computes exactly this and then throws it away into a log line.
  • The running jar fingerprint and boot time, next to the git HEAD on disk. A merge is not a
    deployment, and today the only way to tell them apart is reading bridged.out by eye.
  • Charter receipt per member (CB-571): which charter bytes that member was actually launched
    with, by digest. A member launched with the wrong or empty charter looks normal from outside.
  • Which credentials a member inherited at spawn (CB-596, #82). The policy is deny-by-default in
    config, but the pane's login shell re-exports what it likes. Show what the member really got, as
    names only.

E. Health — is the plumbing sound?

  • herdr version and protocol number, next to the protocol the adapter expects. healthz green with
    a changed protocol still means every spawn fails.
  • Broker connection state. The live log carries AMQP … connection recovery errors that no surface
    reports.
  • ERROR and WARN counts since boot, with the boot marker, so old errors are not read as new ones.

F. Actions — small and explicit

Read-only by default. Stop a member, drain a session, poll a ticket. Every action goes through the
existing Authz table — the dashboard gets no privilege the MCP tools do not have.

Non-goals

  • Not a replacement for the bridge_* tools. Agents keep using the bridge; this is for the
    human operator.
  • Not a public surface. It stays behind bind: 127.0.0.1 like everything else.
  • Never renders a secret value. Credential names only, never values — the same rule the
    --check script already follows.
  • Not a log viewer. bridged.out is fine for that.

Acceptance criteria

  • One page answers all five of: who is alive, what is owed, where the next spawn goes, does the
    running daemon match the file on disk, is the plumbing sound.
  • The reply-inbox count is shown without draining the inbox. Proven by a test that reads the
    dashboard twice and then still drains a reply successfully.
  • The placement panel's predicted profile matches what an actual spawn picks. Proven against a real
    spawn, not against the policy in isolation — a test that calls the policy directly walks around
    the gate it is meant to prove.
  • Every new metric appears in GET /metrics too, so the page is a view over data and not a second
    source of truth.
  • An entry in wiki/11-Features.md: what it does, the knob that turns it on, why it exists, the
    gotcha.

Open questions

  1. Served by the existing Javalin app under /ui, or a separate process? Same process is simpler
    and shares Authz; a separate process survives a daemon restart, which is exactly when the
    operator wants to look.
  2. Server-rendered page, or JSON endpoints plus a small static app?
  3. Does it show one host only, or is it the seat the 2.0 federation line (#6, CB-308) later plugs
    many hosts into? That answer decides the shape of the roster model now.
## Why this exists The lead's only view of the fleet is the `bridge_*` MCP tools. That view has three holes, and all three showed up in one session on 2026-08-22: 1. **A daemon restart cuts the lead's own MCP mount.** The mount does not reconnect. After a redeploy the lead cannot call `bridge_list` or `bridge_whoami` at all, so the moment you most need to check the fleet is the moment you cannot. 2. **Placement is invisible.** To answer "where will the next dev spawn land?" the lead had to read `bridged.yaml`, read `WeightedRoundRobinPolicy.java`, and do the ratio by hand. Nothing reports the effective share, and CB-589 already records that `weighted` is not cheapest-first — so the number people assume is wrong. 3. **A deferred config key is visible only as one log line.** Adding the `xf` profile printed `config reloaded; these changes need a restart to take effect: profiles (added/removed: xf)` and nothing else. Miss that line and the file on disk and the running daemon disagree, silently. There is no page anywhere that answers "what is the fleet doing right now". ## What exists today, and where it stops | Surface | Gives you | Stops at | |---|---|---| | `GET /healthz` | daemon up, herdr version + protocol | says nothing about members; green while every spawn fails | | `GET /members` | the member roster | flat JSON, no history, no timers | | `GET /sessions` | herdr workspaces and pane counts | terminal view, not fleet view | | `GET /agents` | every pane herdr knows | includes panes that are not members | | `GET /sessions/{id}/status` | one member's state | one at a time | | `GET /tasks/{ticket}` | one async ticket | you must already know the ticket id | | `GET /metrics` | **one** gauge family: `bridged_sessions{state=…}` | no placement, cost, ticket, config or health metric at all | | `GET /profiles` | profile names + the default | no weight, no maxLoad, no live count, no quarantine | `/metrics` is the honest measure of the gap. Checked on a live daemon, it emits 6 lines total. ```mermaid flowchart LR A["bridged daemon"] --> B["SessionManager roster"] A --> C["CompositePeerLauncher<br/>placement + quarantine"] A --> D["ConfigRef<br/>live vs deferred"] A --> E["MessageBus<br/>tickets + reply inbox"] A --> F["herdr<br/>pane liveStatus"] B --> G["dashboard"] C --> G D --> G E --> G F --> G ``` The five sources a dashboard has to join. Each already exists in the daemon; none is exposed as a view. ## Features to build ### A. Fleet roster — who is alive - Every member in one table: role, profile, model, state, pane, tab, worktree, uptime. - **The branch the member actually committed to**, read from its worktree — not the branch spawn provisioned. Those differ, and trusting the provisioned one makes you push an empty ref. - **Time since the pane last produced output.** A member can report `working` for hours while doing nothing. Status alone cannot tell you; last-output time can. - **Countdown to the idle reaper** (`lifecycle.idleTtlSeconds`, currently 1800). When it fires, the member's ticket, pane and report are gone together and there is no scrape fallback. The operator needs to see that clock before it runs out, not after. - Which lead owns which member. ### B. Work in flight — what is owed to whom - Open tickets from `wait:false` sends, each with its **TTL countdown**. A ticket older than ten minutes has expired and the lead must send again; nothing tells you that today. - **Undrained replies per session**, as a count only. Reading the reply inbox is destructive (`GET /sessions/{id}/replies` drains on first read), so the dashboard must show that something is waiting without consuming it. This is a hard requirement, not a preference. - Open `bridge_ask` questions: the question, the `turnId` that answers it, and the remaining part of its ~55 second window. An ask that times out is invisible in every current surface. ### C. Placement and cost — where the next spawn goes - **The effective spawn ratio per role pool**, computed live from weight, maxLoad, live count and quarantine. Not the configured weights — the ratio that will actually apply on the next call. - A plain answer to "if I spawn a dev right now, which profile gets it, and why that one". - Live count against `maxLoad` per profile, so an approaching `PlacementException` is visible before it throws. - **Quarantine state per credential** with the remaining cooldown (CB-578 stage B). Profiles sharing a `credentialId` must be shown as one group — `sol` and `terra` share `openai-shared`, and an exhaustion on either locks both. - Spawn counts split by free / paid / subscription, over time. This is the number that tells an operator whether the weights are doing what they think. ### D. Config truth — is the running daemon the file on disk? - **The deferred set**: which keys are in `bridged.yaml` but not live, and therefore need a restart. `ConfigRef.reloadDiff` already computes exactly this and then throws it away into a log line. - The running jar fingerprint and boot time, next to the git HEAD on disk. A merge is not a deployment, and today the only way to tell them apart is reading `bridged.out` by eye. - **Charter receipt per member** (CB-571): which charter bytes that member was actually launched with, by digest. A member launched with the wrong or empty charter looks normal from outside. - **Which credentials a member inherited** at spawn (CB-596, #82). The policy is deny-by-default in config, but the pane's login shell re-exports what it likes. Show what the member really got, as names only. ### E. Health — is the plumbing sound? - herdr version and protocol number, next to the protocol the adapter expects. `healthz` green with a changed protocol still means every spawn fails. - Broker connection state. The live log carries `AMQP … connection recovery` errors that no surface reports. - ERROR and WARN counts since boot, with the boot marker, so old errors are not read as new ones. ### F. Actions — small and explicit Read-only by default. Stop a member, drain a session, poll a ticket. Every action goes through the existing `Authz` table — the dashboard gets no privilege the MCP tools do not have. ## Non-goals - **Not a replacement for the `bridge_*` tools.** Agents keep using the bridge; this is for the human operator. - **Not a public surface.** It stays behind `bind: 127.0.0.1` like everything else. - **Never renders a secret value.** Credential *names* only, never values — the same rule the `--check` script already follows. - Not a log viewer. `bridged.out` is fine for that. ## Acceptance criteria - One page answers all five of: who is alive, what is owed, where the next spawn goes, does the running daemon match the file on disk, is the plumbing sound. - The reply-inbox count is shown without draining the inbox. Proven by a test that reads the dashboard twice and then still drains a reply successfully. - The placement panel's predicted profile matches what an actual spawn picks. Proven against a real spawn, not against the policy in isolation — a test that calls the policy directly walks around the gate it is meant to prove. - Every new metric appears in `GET /metrics` too, so the page is a view over data and not a second source of truth. - An entry in `wiki/11-Features.md`: what it does, the knob that turns it on, why it exists, the gotcha. ## Open questions 1. Served by the existing Javalin app under `/ui`, or a separate process? Same process is simpler and shares `Authz`; a separate process survives a daemon restart, which is exactly when the operator wants to look. 2. Server-rendered page, or JSON endpoints plus a small static app? 3. Does it show one host only, or is it the seat the 2.0 federation line (#6, CB-308) later plugs many hosts into? That answer decides the shape of the roster model now.
ltms added this to the 2.0 — one operation centre, many hosts milestone 2026-08-22 14:01:54 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#118