CB-642: add shared fleets status skill
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m36s

This commit is contained in:
Dai Ha
2026-08-27 21:55:27 +07:00
parent 97d9cebc59
commit da2625acfb
2 changed files with 260 additions and 1 deletions
+258
View File
@@ -0,0 +1,258 @@
---
name: fleets-status
description: Report the status of every fleet that shares one LavinMQ instance. Use for local daemon health, broker-wide fleet presence, and cross-host lead coordination checks.
---
# Status of every fleet on the shared LavinMQ instance
**The headline: always report what is missing.** This skill starts with the local fleet, then adds
broker-wide facts when its read-only credential exists. A missing fleet must appear as `unknown` or
`not reachable`, with the reason and the fix. Never leave it out.
The known topology has one LavinMQ instance on `10.10.20.13` (`fleet01`). AMQP uses port `5672`,
and the management API uses port `15672`. The Mac fleet owns vhost `/mac`. The fleet01 fleet owns
vhost `/fleet01`.
## 1. Protect credentials before any probe
**Hard rule — never print `LAVINMQ_URI`.** It is an AMQP URI with its password inline. It only
resolves in a login shell because `${SHARED_ENV}/tools/secrets.sh` supplies it. A non-login shell
can make every broker probe look empty.
- Never run `echo "$LAVINMQ_URI"`.
- Never put `${LAVINMQ_URI:-something}` in output. That form expands to the secret value when set.
- Parse the user, host, and password into shell or Python variables. Use them without printing them.
- Prefer `resolves` or `does not resolve` over any part of the value.
- Every command that can read `LAVINMQ_URI` must send all output through this redaction before it
reaches the report:
```bash
sed -E 's#://[^@]*@#://<redacted>@#'
```
Keep `pipefail` on when applying that filter. Otherwise the filter can hide a failed probe. Apply
the same no-print rule to the management password below, even though it is not in an AMQP URI.
## 2. Tier 1 — this fleet (always run)
Start here even when the broker tier is blocked. Work from the local fleetd checkout.
First run the read-only deployment check. It already checks the daemon process, deployed jar versus
the checkout `HEAD`, launchd state, and whether each configured token resolves in a login shell.
Do not copy those checks into new shell code. The script reads `LAVINMQ_URI`, so redact all output:
```bash
set -o pipefail
scripts/redeploy-fleetd.sh --check 2>&1 \
| sed -E 's#://[^@]*@#://<redacted>@#'
git rev-parse HEAD
```
Treat jar drift as a top-level warning. A merge is not a deployment. State the running jar result
as `matches HEAD`, `drift`, or `unknown`; do not turn an unclear timestamp into a match.
Report the process identifier (PID) and uptime too:
```bash
PIDS="$(pgrep -f 'target/fleetd.jar' || true)"
if [ -z "$PIDS" ]; then
printf '%s\n' 'fleetd: not running'
else
for PID in $PIDS; do
ps -p "$PID" -o pid=,etime=,lstart=,command=
done
fi
```
Read the full health response. Keep the HTTP status because `503` means fleetd is running but herdr
is not reachable. Report both `herdr.version` and `herdr.protocol` when present:
```bash
curl -sS --max-time 5 -w '\nHTTP %{http_code}\n' http://127.0.0.1:8765/healthz
```
Call `fleet_whoami`, then call `fleet_list`. Preserve its sections in the report:
- `leads`, including which row is this lead;
- `members`, including state, role, profile, branch, and worktree when present;
- every per-profile `capacity` row, including `maxLoad`, `live`, `free`, and quarantine facts;
- the exact `healthCoverage` value.
Do not describe an empty `members` list as an empty fleet. It says only that no members are spawned.
Also do not hide a profile with `free: 0`; say whether load or credential quarantine caused it.
Show only WARN, ERROR, and SEVERE lines after the last `fleetd listening` line. This anchor stops an
old incident from looking current:
```bash
python3 - <<'PY'
from pathlib import Path
import re
path = Path("fleetd/fleetd.out")
if not path.exists():
print("cannot check current WARN/ERROR: fleetd/fleetd.out does not exist")
else:
lines = path.read_text(errors="replace").splitlines()
starts = [i for i, line in enumerate(lines) if "fleetd listening" in line]
if not starts:
print("cannot anchor WARN/ERROR: no 'fleetd listening' line exists")
else:
current = lines[starts[-1]:]
alerts = [line for line in current if re.search(r"\b(?:WARN|ERROR|SEVERE)\b", line)]
print(f"current WARN/ERROR/SEVERE count: {len(alerts)}")
for line in alerts[-50:]:
print(line)
PY
```
**What this tier cannot see:** it proves facts only about the Mac daemon at `127.0.0.1:8765`.
It cannot show the fleet01 daemon, broker queue depth, or broker consumers. The fleet01 REST service
at `10.10.20.13:8765` is not reachable from the Mac, and SSH as `dai.ha@10.10.20.13` is denied.
Say this in the report rather than omitting fleet01.
## 3. Tier 2 — the shared broker (run when management access exists)
**This tier is blocked today.** The AMQP user in `LAVINMQ_URI` can connect on port `5672`, but gets
HTTP `401` from the management API on port `15672`. An AMQP connection does not grant monitoring
access.
The operator must create a separate, read-only LavinMQ management user with the `monitoring` tag.
It needs access to inspect both `/mac` and `/fleet01`. Store its values as
`LAVINMQ_MANAGEMENT_USER` and `LAVINMQ_MANAGEMENT_PASSWORD` in
`${SHARED_ENV}/tools/secrets.sh`. Do not reuse or print the AMQP URI. Full multi-fleet status stays
blocked until this user exists.
When both variables resolve, run this from a login shell. It calls `GET /api/overview`,
`GET /api/vhosts`, `GET /api/queues`, and `GET /api/connections`. It prints selected status fields,
but never the user, password, Authorization header, or AMQP URI:
```bash
zsh -lc 'python3 - "$@"' -- <<'PY'
import base64
import json
import os
import sys
import urllib.error
import urllib.request
base = "http://10.10.20.13:15672"
user = os.environ.get("LAVINMQ_MANAGEMENT_USER", "")
password = os.environ.get("LAVINMQ_MANAGEMENT_PASSWORD", "")
if not user or not password:
print("broker tier: BLOCKED — management credential does not resolve in a login shell")
sys.exit(0)
token = base64.b64encode(f"{user}:{password}".encode()).decode()
def get(path):
request = urllib.request.Request(
base + path,
headers={"Authorization": "Basic " + token, "Accept": "application/json"},
)
with urllib.request.urlopen(request, timeout=5) as response:
return json.load(response)
try:
overview = get("/api/overview")
vhosts = get("/api/vhosts")
queues = get("/api/queues")
connections = get("/api/connections")
except urllib.error.HTTPError as error:
print(f"broker tier: BLOCKED — management API returned HTTP {error.code}")
sys.exit(0)
except Exception as error:
print(f"broker tier: BLOCKED — management API is not reachable: {type(error).__name__}")
sys.exit(0)
fleet_names = {"/mac": "Mac fleet", "/fleet01": "fleet01 fleet"}
print(json.dumps({
"overview": {
"lavinmq_version": overview.get("lavinmq_version"),
"rabbitmq_version": overview.get("rabbitmq_version"),
"queue_totals": overview.get("queue_totals", {}),
"object_totals": overview.get("object_totals", {}),
},
"fleets": [
{
"fleet": fleet_names.get(vhost.get("name"), "UNKNOWN FLEET"),
"vhost": vhost.get("name"),
"queues": [
{
"name": queue.get("name"),
"messages": queue.get("messages", 0),
"messages_ready": queue.get("messages_ready", 0),
"messages_unacknowledged": queue.get("messages_unacknowledged", 0),
"consumers": queue.get("consumers", 0),
}
for queue in queues if queue.get("vhost") == vhost.get("name")
],
"connections": [
{
"name": connection.get("name"),
"peer_host": connection.get("peer_host"),
"state": connection.get("state"),
}
for connection in connections if connection.get("vhost") == vhost.get("name")
],
}
for vhost in vhosts
],
}, indent=2, sort_keys=True))
PY
```
Map `/mac` to the Mac fleet and `/fleet01` to the fleet01 fleet. Keep any other vhost in the
report as `UNKNOWN FLEET`; do not drop it. For each vhost, total the ready, unacknowledged, and all
messages. Report every queue's consumer count and each live connection.
**A vhost with queues but zero consumers means that fleet's daemon is down while its durable state
survives. Call this out as a top-level warning.** This is the main reason to use the management API
instead of calling each remote daemon.
**What this tier cannot see:** without the new `monitoring` credential it cannot enumerate any
vhost, queue, depth, consumer, or connection. With the credential it still cannot report fleet01's
daemon PID, uptime, jar revision, `/healthz`, herdr version, or member capacity. Those need reachable
fleet01 REST or SSH access, which the Mac does not have today.
## 4. Tier 3 — cross-fleet lead coordination
Use the queue data from Tier 2. Select queues whose names match `lead.<coordId>.inbox`. Report each
queue's vhost, depth, consumer count, and the `coordId` between the prefix and suffix.
- A lead inbox with a consumer shows that a lead mailbox is live on that vhost.
- A durable lead inbox with zero consumers shows saved coordination state, but no live receiver.
- No lead inbox is not proof that coordination is disabled. The daemon may be down before declaring
its queue, or this account may not be allowed to see the vhost.
This Mac fleet currently sets both `broker.uriEnv` and `coordinator.uriEnv` to the same variable,
`LAVINMQ_URI`. Therefore its coordinator connects to `/mac`. Cross-host `fleet_send{coordId}` routes
only when both leads share the same coordinator vhost. If the fleet01 lead uses `/fleet01` for its
coordinator, the leads cannot see each other and the send will not route.
**Open question:** the fleet01 coordinator vhost has not been checked. Surface this question in
every report until a live `lead.<coordId>.inbox` consumer or fleet01's config proves the answer. Do
not claim that fleet01 uses `/fleet01` just because its member queues do.
Also compare these broker facts with the `leads` rows from local `fleet_list`. A missing remote lead
is `not visible from this coordinator`, not `down`, unless the broker consumer facts prove it.
**What this tier cannot see:** without Tier 2 management access it cannot list lead inboxes or their
consumers. Even with that access, a stopped fleet01 daemon leaves only durable queue history. That
history cannot prove which coordinator URI its current config would use after restart.
## 5. Report all fleets
Use one row per known or discovered fleet. Include blocked rows.
| Fleet | Daemon | Deployment | Herdr | Members/capacity | Queues/consumers | Lead coordination | Cannot check |
|---|---|---|---|---|---|---|---|
| Mac (`/mac`) | PID + uptime | jar vs `HEAD` | health + version + protocol | `fleet_list` + `healthCoverage` | facts or blocked reason | inbox facts or open question | exact missing facts |
| fleet01 (`/fleet01`) | reachable/down/unknown | value or `not reachable` | value or `not reachable` | value or `not reachable` | facts or blocked reason | inbox facts plus coordinator-vhost question | exact missing facts and fix |
Add rows for unknown vhosts. End with three short sections: `Current warnings`, `Checks that were
blocked`, and `Operator action`. Until the management user exists, `Operator action` must say:
> Create a read-only LavinMQ management user with the `monitoring` tag and access to `/mac` and
> `/fleet01`. Put its user and password in `${SHARED_ENV}/tools/secrets.sh` as
> `LAVINMQ_MANAGEMENT_USER` and `LAVINMQ_MANAGEMENT_PASSWORD`.
+2 -1
View File
@@ -191,7 +191,8 @@ must obey belongs in the charter, not here.
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
`reviewer` (scoped review → one structured finding). Name one in every delegation.
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
`port-to-opencode` (make an OpenCode session a participant in this workspace).
`port-to-opencode` (make an OpenCode session a participant in this workspace) and
`fleets-status` (report every fleet that shares one LavinMQ instance).
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **Flows and the error model** — rendezvous, `fleet_ask`, detached delivery, turn-done fallback —