A broker unreachable at boot silently disables coordination for the process lifetime, and then reports itself as "not configured" #590

Open
opened 2026-09-14 08:32:30 +02:00 by ltms · 0 comments
Owner

Measured live on the Mac fleet today, 2026-09-14. Lead-to-lead messaging was off for 6 hours 35 minutes and nothing said so.

What happened

The daemon restarted at 06:55:29. The broker lives on another host, about 230 ms away. At boot the network was not up yet, so both AMQP connections timed out:

06:56:46  startup secret LAVINMQ_URI: set (broker uriEnv)
06:57:17  reply inbox: AMQP broker (durable) via env var LAVINMQ_URI (prefetch=32)
06:58:18  WARN cannot reach AMQP broker (amqp://10.10.20.13:5672/mac) — falling back to the
          in-memory reply inbox for this process lifetime. Reason: SocketTimeoutException: Connect timed out
06:59:18  WARN cannot reach the AMQP coordination broker (amqp://10.10.20.13:5672/coord) —
          lead-to-lead messaging is OFF for this process lifetime. Reason: SocketTimeoutException: Connect timed out

Both give up after one attempt, each costing a 60 s timeout, and the decision is final for the life of the process.

By the time I looked, the broker was fine. Measured before I touched anything:

nc -z 10.10.20.13 5672   ->  succeeded
ping                     ->  0% loss, avg 230 ms

A plain restart fixed it. coordinator came straight back in fleet_list with configured: true, mailbox.consumers: 1, and the fleet01 peer showing consumers: 1.

So nothing was wrong with the broker, the config, or the credential. The daemon simply asked once, at the worst possible moment, and never asked again.

Item 1 — the failure reports itself as the wrong thing

This is the half that cost the time, because it removes the evidence that anything failed.

leadChannel == null carries two states that need opposite responses:

  • no coordinator: block was ever configured
  • a coordinator: block is configured and its broker could not be reached at boot

FleetMcp.java:911 answers both with the same text:

if (leadChannel == null) {
    return error("lead coordination is not configured (no coordinator: block) — cannot send to "
            + "peer lead \"" + coordId + "\". Add a coordinator: block with a shared broker uri "
            + "and this daemon's selfId, then restart fleetd.");
}

In state 2 that advice is wrong in a way that wastes the reader's time: it tells a lead to add a block that is already there, and says nothing about the one action that actually works. The right advice in state 2 is "the broker was unreachable when this process started; check it is reachable now and restart."

fleet_list is worse, because it omits the coordinator row entirely. An absent row reads as "this fleet has no peers", which is a confident answer to a question the daemon cannot currently answer.

The code already knows. FleetMcp.java:319 documents the conflation in its own @param:

null whenever no coordinator: block is configured or its broker could not be reached at boot — cross-daemon lead messaging is simply off, and fleet_send{coordId} says so rather than failing obscurely

It does not say so. It says the opposite of one of the two cases.

Fleetd.java:1531 also predicts it, in the warning itself: "fleet_send{{coordId}} will report it as not configured". That line is a note that a future reader will be misled, written at the one place where the truth is still known and then thrown away.

Fix: a third state. Keep null for never-configured, and add a distinct value for configured-but-unreachable that carries the broker address and the failure reason. Then fleet_send{coordId} names the real cause, and fleet_list reports a coordinator row with configured: true, connected: false plus the reason, instead of vanishing.

Item 2 — one attempt, no retry, permanent

Independent of item 1 and worth fixing on its own. Fleetd.java:1525-1537 opens the mailbox once and returns null on IllegalStateException. Nothing retries for the life of the process.

A daemon under launchd with KeepAlive starts on boot or wake, which is exactly when a remote broker is least likely to answer. On this host that is not an unlucky edge case, it is the normal startup order.

Two options, and I have not measured which is better here:

  • retry with backoff in the background and upgrade the channel when it succeeds — best behaviour, most work, and it has to be safe against a fleet_send racing the upgrade
  • fail the boot instead of degrading, and let the supervisor restart — much smaller, and it turns a silent 6-hour outage into a visible restart loop that names the broker

The same single-attempt pattern applies to the durable reply inbox at Fleetd.java:1596, which fell back to in-memory in the same incident. That one is arguably worse: replies stop being durable across a restart, and the only notice is one WARN line at boot.

Why this matters beyond one morning

The whole point of coordinator.peers (fleetd #364) is to turn "is my peer actually consuming what I send?" into a fact you can read instead of guess. This defect breaks that guarantee in the one case where you most need it — the channel is down — and answers with a shape that looks like "you have no peers".

How to reproduce

Point coordinator.broker at an address that black-holes (a firewalled host, not a closed port — you want a timeout, not a refusal), start the daemon, then make it reachable. fleet_list has no coordinator row and fleet_send{coordId} tells you to add a block that is already in your config. Restart, and everything works with no config change.

Not in scope

Nothing here is about the broker, LavinMQ, or the credential. All three were correct throughout.

Measured live on the Mac fleet today, 2026-09-14. Lead-to-lead messaging was off for 6 hours 35 minutes and nothing said so. ## What happened The daemon restarted at 06:55:29. The broker lives on another host, about 230 ms away. At boot the network was not up yet, so both AMQP connections timed out: ``` 06:56:46 startup secret LAVINMQ_URI: set (broker uriEnv) 06:57:17 reply inbox: AMQP broker (durable) via env var LAVINMQ_URI (prefetch=32) 06:58:18 WARN cannot reach AMQP broker (amqp://10.10.20.13:5672/mac) — falling back to the in-memory reply inbox for this process lifetime. Reason: SocketTimeoutException: Connect timed out 06:59:18 WARN cannot reach the AMQP coordination broker (amqp://10.10.20.13:5672/coord) — lead-to-lead messaging is OFF for this process lifetime. Reason: SocketTimeoutException: Connect timed out ``` Both give up after one attempt, each costing a 60 s timeout, and the decision is final for the life of the process. By the time I looked, the broker was fine. Measured before I touched anything: ``` nc -z 10.10.20.13 5672 -> succeeded ping -> 0% loss, avg 230 ms ``` A plain restart fixed it. `coordinator` came straight back in `fleet_list` with `configured: true`, `mailbox.consumers: 1`, and the `fleet01` peer showing `consumers: 1`. So nothing was wrong with the broker, the config, or the credential. The daemon simply asked once, at the worst possible moment, and never asked again. ## Item 1 — the failure reports itself as the wrong thing This is the half that cost the time, because it removes the evidence that anything failed. `leadChannel == null` carries two states that need opposite responses: * no `coordinator:` block was ever configured * a `coordinator:` block **is** configured and its broker could not be reached at boot `FleetMcp.java:911` answers both with the same text: ```java if (leadChannel == null) { return error("lead coordination is not configured (no coordinator: block) — cannot send to " + "peer lead \"" + coordId + "\". Add a coordinator: block with a shared broker uri " + "and this daemon's selfId, then restart fleetd."); } ``` In state 2 that advice is wrong in a way that wastes the reader's time: it tells a lead to add a block that is already there, and says nothing about the one action that actually works. The right advice in state 2 is "the broker was unreachable when this process started; check it is reachable now and restart." `fleet_list` is worse, because it omits the `coordinator` row entirely. An absent row reads as "this fleet has no peers", which is a confident answer to a question the daemon cannot currently answer. The code already knows. `FleetMcp.java:319` documents the conflation in its own `@param`: > `null` whenever no `coordinator:` block is configured **or its broker could not be reached at boot** — cross-daemon lead messaging is simply off, and `fleet_send{coordId}` says so rather than failing obscurely It does not say so. It says the opposite of one of the two cases. `Fleetd.java:1531` also predicts it, in the warning itself: *"fleet_send{{coordId}} will report it as not configured"*. That line is a note that a future reader will be misled, written at the one place where the truth is still known and then thrown away. **Fix: a third state.** Keep `null` for never-configured, and add a distinct value for configured-but-unreachable that carries the broker address and the failure reason. Then `fleet_send{coordId}` names the real cause, and `fleet_list` reports a `coordinator` row with `configured: true, connected: false` plus the reason, instead of vanishing. ## Item 2 — one attempt, no retry, permanent Independent of item 1 and worth fixing on its own. `Fleetd.java:1525-1537` opens the mailbox once and returns `null` on `IllegalStateException`. Nothing retries for the life of the process. A daemon under `launchd` with `KeepAlive` starts on boot or wake, which is exactly when a remote broker is least likely to answer. On this host that is not an unlucky edge case, it is the normal startup order. Two options, and I have not measured which is better here: * retry with backoff in the background and upgrade the channel when it succeeds — best behaviour, most work, and it has to be safe against a `fleet_send` racing the upgrade * fail the boot instead of degrading, and let the supervisor restart — much smaller, and it turns a silent 6-hour outage into a visible restart loop that names the broker The same single-attempt pattern applies to the durable reply inbox at `Fleetd.java:1596`, which fell back to in-memory in the same incident. That one is arguably worse: replies stop being durable across a restart, and the only notice is one WARN line at boot. ## Why this matters beyond one morning The whole point of `coordinator.peers` (fleetd #364) is to turn "is my peer actually consuming what I send?" into a fact you can read instead of guess. This defect breaks that guarantee in the one case where you most need it — the channel is down — and answers with a shape that looks like "you have no peers". ## How to reproduce Point `coordinator.broker` at an address that black-holes (a firewalled host, not a closed port — you want a timeout, not a refusal), start the daemon, then make it reachable. `fleet_list` has no `coordinator` row and `fleet_send{coordId}` tells you to add a block that is already in your config. Restart, and everything works with no config change. ## Not in scope Nothing here is about the broker, LavinMQ, or the credential. All three were correct throughout.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#590