diff --git a/.claude/skills/fleets-status/SKILL.md b/.claude/skills/fleets-status/SKILL.md new file mode 100644 index 0000000..d34ca7a --- /dev/null +++ b/.claude/skills/fleets-status/SKILL.md @@ -0,0 +1,258 @@ +--- +name: fleets-status +description: Report the status of every fleet that shares one LavinMQ instance. Use for local daemon health, broker-wide fleet presence, and cross-host lead coordination checks. +--- + +# Status of every fleet on the shared LavinMQ instance + +**The headline: always report what is missing.** This skill starts with the local fleet, then adds +broker-wide facts when its read-only credential exists. A missing fleet must appear as `unknown` or +`not reachable`, with the reason and the fix. Never leave it out. + +The known topology has one LavinMQ instance on `10.10.20.13` (`fleet01`). AMQP uses port `5672`, +and the management API uses port `15672`. The Mac fleet owns vhost `/mac`. The fleet01 fleet owns +vhost `/fleet01`. + +## 1. Protect credentials before any probe + +**Hard rule — never print `LAVINMQ_URI`.** It is an AMQP URI with its password inline. It only +resolves in a login shell because `${SHARED_ENV}/tools/secrets.sh` supplies it. A non-login shell +can make every broker probe look empty. + +- Never run `echo "$LAVINMQ_URI"`. +- Never put `${LAVINMQ_URI:-something}` in output. That form expands to the secret value when set. +- Parse the user, host, and password into shell or Python variables. Use them without printing them. +- Prefer `resolves` or `does not resolve` over any part of the value. +- Every command that can read `LAVINMQ_URI` must send all output through this redaction before it + reaches the report: + +```bash +sed -E 's#://[^@]*@#://@#' +``` + +Keep `pipefail` on when applying that filter. Otherwise the filter can hide a failed probe. Apply +the same no-print rule to the management password below, even though it is not in an AMQP URI. + +## 2. Tier 1 — this fleet (always run) + +Start here even when the broker tier is blocked. Work from the local fleetd checkout. + +First run the read-only deployment check. It already checks the daemon process, deployed jar versus +the checkout `HEAD`, launchd state, and whether each configured token resolves in a login shell. +Do not copy those checks into new shell code. The script reads `LAVINMQ_URI`, so redact all output: + +```bash +set -o pipefail +scripts/redeploy-fleetd.sh --check 2>&1 \ + | sed -E 's#://[^@]*@#://@#' +git rev-parse HEAD +``` + +Treat jar drift as a top-level warning. A merge is not a deployment. State the running jar result +as `matches HEAD`, `drift`, or `unknown`; do not turn an unclear timestamp into a match. + +Report the process identifier (PID) and uptime too: + +```bash +PIDS="$(pgrep -f 'target/fleetd.jar' || true)" +if [ -z "$PIDS" ]; then + printf '%s\n' 'fleetd: not running' +else + for PID in $PIDS; do + ps -p "$PID" -o pid=,etime=,lstart=,command= + done +fi +``` + +Read the full health response. Keep the HTTP status because `503` means fleetd is running but herdr +is not reachable. Report both `herdr.version` and `herdr.protocol` when present: + +```bash +curl -sS --max-time 5 -w '\nHTTP %{http_code}\n' http://127.0.0.1:8765/healthz +``` + +Call `fleet_whoami`, then call `fleet_list`. Preserve its sections in the report: + +- `leads`, including which row is this lead; +- `members`, including state, role, profile, branch, and worktree when present; +- every per-profile `capacity` row, including `maxLoad`, `live`, `free`, and quarantine facts; +- the exact `healthCoverage` value. + +Do not describe an empty `members` list as an empty fleet. It says only that no members are spawned. +Also do not hide a profile with `free: 0`; say whether load or credential quarantine caused it. + +Show only WARN, ERROR, and SEVERE lines after the last `fleetd listening` line. This anchor stops an +old incident from looking current: + +```bash +python3 - <<'PY' +from pathlib import Path +import re + +path = Path("fleetd/fleetd.out") +if not path.exists(): + print("cannot check current WARN/ERROR: fleetd/fleetd.out does not exist") +else: + lines = path.read_text(errors="replace").splitlines() + starts = [i for i, line in enumerate(lines) if "fleetd listening" in line] + if not starts: + print("cannot anchor WARN/ERROR: no 'fleetd listening' line exists") + else: + current = lines[starts[-1]:] + alerts = [line for line in current if re.search(r"\b(?:WARN|ERROR|SEVERE)\b", line)] + print(f"current WARN/ERROR/SEVERE count: {len(alerts)}") + for line in alerts[-50:]: + print(line) +PY +``` + +**What this tier cannot see:** it proves facts only about the Mac daemon at `127.0.0.1:8765`. +It cannot show the fleet01 daemon, broker queue depth, or broker consumers. The fleet01 REST service +at `10.10.20.13:8765` is not reachable from the Mac, and SSH as `dai.ha@10.10.20.13` is denied. +Say this in the report rather than omitting fleet01. + +## 3. Tier 2 — the shared broker (run when management access exists) + +**This tier is blocked today.** The AMQP user in `LAVINMQ_URI` can connect on port `5672`, but gets +HTTP `401` from the management API on port `15672`. An AMQP connection does not grant monitoring +access. + +The operator must create a separate, read-only LavinMQ management user with the `monitoring` tag. +It needs access to inspect both `/mac` and `/fleet01`. Store its values as +`LAVINMQ_MANAGEMENT_USER` and `LAVINMQ_MANAGEMENT_PASSWORD` in +`${SHARED_ENV}/tools/secrets.sh`. Do not reuse or print the AMQP URI. Full multi-fleet status stays +blocked until this user exists. + +When both variables resolve, run this from a login shell. It calls `GET /api/overview`, +`GET /api/vhosts`, `GET /api/queues`, and `GET /api/connections`. It prints selected status fields, +but never the user, password, Authorization header, or AMQP URI: + +```bash +zsh -lc 'python3 - "$@"' -- <<'PY' +import base64 +import json +import os +import sys +import urllib.error +import urllib.request + +base = "http://10.10.20.13:15672" +user = os.environ.get("LAVINMQ_MANAGEMENT_USER", "") +password = os.environ.get("LAVINMQ_MANAGEMENT_PASSWORD", "") +if not user or not password: + print("broker tier: BLOCKED — management credential does not resolve in a login shell") + sys.exit(0) + +token = base64.b64encode(f"{user}:{password}".encode()).decode() + +def get(path): + request = urllib.request.Request( + base + path, + headers={"Authorization": "Basic " + token, "Accept": "application/json"}, + ) + with urllib.request.urlopen(request, timeout=5) as response: + return json.load(response) + +try: + overview = get("/api/overview") + vhosts = get("/api/vhosts") + queues = get("/api/queues") + connections = get("/api/connections") +except urllib.error.HTTPError as error: + print(f"broker tier: BLOCKED — management API returned HTTP {error.code}") + sys.exit(0) +except Exception as error: + print(f"broker tier: BLOCKED — management API is not reachable: {type(error).__name__}") + sys.exit(0) + +fleet_names = {"/mac": "Mac fleet", "/fleet01": "fleet01 fleet"} +print(json.dumps({ + "overview": { + "lavinmq_version": overview.get("lavinmq_version"), + "rabbitmq_version": overview.get("rabbitmq_version"), + "queue_totals": overview.get("queue_totals", {}), + "object_totals": overview.get("object_totals", {}), + }, + "fleets": [ + { + "fleet": fleet_names.get(vhost.get("name"), "UNKNOWN FLEET"), + "vhost": vhost.get("name"), + "queues": [ + { + "name": queue.get("name"), + "messages": queue.get("messages", 0), + "messages_ready": queue.get("messages_ready", 0), + "messages_unacknowledged": queue.get("messages_unacknowledged", 0), + "consumers": queue.get("consumers", 0), + } + for queue in queues if queue.get("vhost") == vhost.get("name") + ], + "connections": [ + { + "name": connection.get("name"), + "peer_host": connection.get("peer_host"), + "state": connection.get("state"), + } + for connection in connections if connection.get("vhost") == vhost.get("name") + ], + } + for vhost in vhosts + ], +}, indent=2, sort_keys=True)) +PY +``` + +Map `/mac` to the Mac fleet and `/fleet01` to the fleet01 fleet. Keep any other vhost in the +report as `UNKNOWN FLEET`; do not drop it. For each vhost, total the ready, unacknowledged, and all +messages. Report every queue's consumer count and each live connection. + +**A vhost with queues but zero consumers means that fleet's daemon is down while its durable state +survives. Call this out as a top-level warning.** This is the main reason to use the management API +instead of calling each remote daemon. + +**What this tier cannot see:** without the new `monitoring` credential it cannot enumerate any +vhost, queue, depth, consumer, or connection. With the credential it still cannot report fleet01's +daemon PID, uptime, jar revision, `/healthz`, herdr version, or member capacity. Those need reachable +fleet01 REST or SSH access, which the Mac does not have today. + +## 4. Tier 3 — cross-fleet lead coordination + +Use the queue data from Tier 2. Select queues whose names match `lead..inbox`. Report each +queue's vhost, depth, consumer count, and the `coordId` between the prefix and suffix. + +- A lead inbox with a consumer shows that a lead mailbox is live on that vhost. +- A durable lead inbox with zero consumers shows saved coordination state, but no live receiver. +- No lead inbox is not proof that coordination is disabled. The daemon may be down before declaring + its queue, or this account may not be allowed to see the vhost. + +This Mac fleet currently sets both `broker.uriEnv` and `coordinator.uriEnv` to the same variable, +`LAVINMQ_URI`. Therefore its coordinator connects to `/mac`. Cross-host `fleet_send{coordId}` routes +only when both leads share the same coordinator vhost. If the fleet01 lead uses `/fleet01` for its +coordinator, the leads cannot see each other and the send will not route. + +**Open question:** the fleet01 coordinator vhost has not been checked. Surface this question in +every report until a live `lead..inbox` consumer or fleet01's config proves the answer. Do +not claim that fleet01 uses `/fleet01` just because its member queues do. + +Also compare these broker facts with the `leads` rows from local `fleet_list`. A missing remote lead +is `not visible from this coordinator`, not `down`, unless the broker consumer facts prove it. + +**What this tier cannot see:** without Tier 2 management access it cannot list lead inboxes or their +consumers. Even with that access, a stopped fleet01 daemon leaves only durable queue history. That +history cannot prove which coordinator URI its current config would use after restart. + +## 5. Report all fleets + +Use one row per known or discovered fleet. Include blocked rows. + +| Fleet | Daemon | Deployment | Herdr | Members/capacity | Queues/consumers | Lead coordination | Cannot check | +|---|---|---|---|---|---|---|---| +| Mac (`/mac`) | PID + uptime | jar vs `HEAD` | health + version + protocol | `fleet_list` + `healthCoverage` | facts or blocked reason | inbox facts or open question | exact missing facts | +| fleet01 (`/fleet01`) | reachable/down/unknown | value or `not reachable` | value or `not reachable` | value or `not reachable` | facts or blocked reason | inbox facts plus coordinator-vhost question | exact missing facts and fix | + +Add rows for unknown vhosts. End with three short sections: `Current warnings`, `Checks that were +blocked`, and `Operator action`. Until the management user exists, `Operator action` must say: + +> Create a read-only LavinMQ management user with the `monitoring` tag and access to `/mac` and +> `/fleet01`. Put its user and password in `${SHARED_ENV}/tools/secrets.sh` as +> `LAVINMQ_MANAGEMENT_USER` and `LAVINMQ_MANAGEMENT_PASSWORD`. diff --git a/CLAUDE.md b/CLAUDE.md index 1dd5178..202e993 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -191,7 +191,8 @@ must obey belongs in the charter, not here. - **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and `reviewer` (scoped review → one structured finding). Name one in every delegation. - **Primary-side skills** (not delegation playbooks — a worker cannot use them): - `port-to-opencode` (make an OpenCode session a participant in this workspace). + `port-to-opencode` (make an OpenCode session a participant in this workspace) and + `fleets-status` (report every fleet that shares one LavinMQ instance). - **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/` (a submodule with its own remote). - **Flows and the error model** — rendezvous, `fleet_ask`, detached delivery, turn-done fallback —