an unreachable broker at startup stops the daemon from booting at all #152

Closed
opened 2026-08-23 13:37:02 +02:00 by ltms · 1 comment
Owner

What happens

If the configured AMQP broker is unreachable when the daemon starts, the daemon does not start at all.

AmqpReplyInbox.open turns any connection failure into an unchecked throw:

} catch (Exception e) {
    throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
}

and Fleetd calls it without a catch:

if (cfg.broker() != null && cfg.broker().isConfigured()) {
    replyInbox = AmqpReplyInbox.open(cfg.broker().uri(), cfg.broker().prefetchOrDefault());

So the exception propagates out of main(). A broker that is merely slow to come up, or briefly down, takes the whole fleet with it — no leads, no members, no REST, no MCP.

Why this matters now

Today each fleet has its own loopback broker, so the broker is up whenever the host is. We are about to point two fleets at one shared LavinMQ instance. That turns a single broker outage into a total outage of every fleet at once, and it makes daemon start order significant across machines, which it never was before.

It is also inconsistent with how the same connection is treated afterwards. The factory sets setAutomaticRecoveryEnabled(true) and setTopologyRecoveryEnabled(true), so a broker that drops after startup is expected to come back on its own and the daemon rides it out. Only the boot path is fatal. The same blip is survivable at 10:00 and fatal at 09:59.

Suggested fix

Catch the failure at the call site and fall back to InMemoryReplyInbox, with a warning that says exactly what was lost:

  • durable, cross-restart reply delivery is off
  • replies are soft-state and will not survive a restart
  • the URI that failed, with credentials stripped, and the reason

Starting degraded and saying so loudly is better than not starting. A daemon that refuses to boot gives an operator no fleet and no tools to diagnose with; a daemon that boots without durability gives them a working fleet and a clear line in the log.

Do not silently fall back. This is the silent-default-disables-features shape — a defaulted dependency that compiles, passes tests, and quietly turns a feature off. The whole value here is in the loudness of the warning.

Worth deciding

Should a failed broker at boot be recoverable later, or does it stay in-memory until the next restart? Retrying in the background and upgrading to the durable inbox once the broker appears would be better behaviour, but it is a larger change and needs care around targets already owned on the in-memory adapter. Falling back for the process lifetime is the smaller, safer fix and I would start there.

Tests

  • broker configured and reachable ⇒ AMQP inbox, as today
  • broker configured and unreachable ⇒ daemon starts, in-memory inbox, warning logged
  • the warning never contains the broker password
  • no broker configured ⇒ in-memory, existing quiet path unchanged

The unreachable case needs a URI pointing at a closed port; it does not need a container.

## What happens If the configured AMQP broker is unreachable when the daemon starts, the daemon **does not start at all**. `AmqpReplyInbox.open` turns any connection failure into an unchecked throw: ```java } catch (Exception e) { throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e); } ``` and `Fleetd` calls it without a catch: ```java if (cfg.broker() != null && cfg.broker().isConfigured()) { replyInbox = AmqpReplyInbox.open(cfg.broker().uri(), cfg.broker().prefetchOrDefault()); ``` So the exception propagates out of `main()`. A broker that is merely slow to come up, or briefly down, takes the whole fleet with it — no leads, no members, no REST, no MCP. ## Why this matters now Today each fleet has its own loopback broker, so the broker is up whenever the host is. We are about to point **two** fleets at one shared LavinMQ instance. That turns a single broker outage into a total outage of every fleet at once, and it makes daemon start order significant across machines, which it never was before. It is also inconsistent with how the same connection is treated afterwards. The factory sets `setAutomaticRecoveryEnabled(true)` and `setTopologyRecoveryEnabled(true)`, so a broker that drops **after** startup is expected to come back on its own and the daemon rides it out. Only the boot path is fatal. The same blip is survivable at 10:00 and fatal at 09:59. ## Suggested fix Catch the failure at the call site and fall back to `InMemoryReplyInbox`, with a **warning that says exactly what was lost**: - durable, cross-restart reply delivery is off - replies are soft-state and will not survive a restart - the URI that failed, with credentials stripped, and the reason Starting degraded and saying so loudly is better than not starting. A daemon that refuses to boot gives an operator no fleet and no tools to diagnose with; a daemon that boots without durability gives them a working fleet and a clear line in the log. Do **not** silently fall back. This is the [[silent-default-disables-features]] shape — a defaulted dependency that compiles, passes tests, and quietly turns a feature off. The whole value here is in the loudness of the warning. ## Worth deciding Should a failed broker at boot be recoverable later, or does it stay in-memory until the next restart? Retrying in the background and upgrading to the durable inbox once the broker appears would be better behaviour, but it is a larger change and needs care around targets already owned on the in-memory adapter. Falling back for the process lifetime is the smaller, safer fix and I would start there. ## Tests - broker configured and reachable ⇒ AMQP inbox, as today - broker configured and unreachable ⇒ daemon **starts**, in-memory inbox, warning logged - the warning never contains the broker password - no broker configured ⇒ in-memory, existing quiet path unchanged The unreachable case needs a URI pointing at a closed port; it does not need a container.
Author
Owner

Shipped and live on both hosts. Merged as 4644359 (PR #153). 908 tests green on the Mac and on fleet01.

Fleetd.selectReplyInbox now catches the connect failure at boot and falls back to InMemoryReplyInbox for the process lifetime. The warning says durable cross-restart delivery is off, that replies are soft-state, which source failed (uri or uriEnv <NAME>), and the failing URI with credentials stripped. There is no background retry — a broker that drops after startup already self-heals through the AMQP client's automatic recovery, so only the boot path changed.

Why this matters more than it looks: before this, a broker that was down took the whole daemon with it, and with the daemon went every lead, every member and every pane. Losing durable replies is bad; losing the fleet because the reply store is down is worse.

On testing — the decision was extracted into a package-private selectReplyInbox(broker, env, opener) with the env map and the AMQP opener injected, so the tests drive the real decision path rather than a re-implementation of it. Seven cases, including one that calls the real AmqpReplyInbox::open against a guaranteed-closed port, and each asserts the password never appears in any log line. main() reaches it through a single direct call I read at the merge, so the caller is covered by inspection rather than by a test that walks around the gate (see #113).

Documented in wiki 11 Features → An unreachable broker does not stop the daemon.

Shipped and live on both hosts. Merged as `4644359` (PR #153). 908 tests green on the Mac and on fleet01. `Fleetd.selectReplyInbox` now catches the connect failure at boot and falls back to `InMemoryReplyInbox` for the process lifetime. The warning says durable cross-restart delivery is off, that replies are soft-state, which source failed (`uri` or `uriEnv <NAME>`), and the failing URI **with credentials stripped**. There is no background retry — a broker that drops *after* startup already self-heals through the AMQP client's automatic recovery, so only the boot path changed. Why this matters more than it looks: before this, a broker that was down took the whole daemon with it, and with the daemon went every lead, every member and every pane. Losing durable replies is bad; losing the fleet because the *reply store* is down is worse. On testing — the decision was extracted into a package-private `selectReplyInbox(broker, env, opener)` with the env map and the AMQP opener injected, so the tests drive the real decision path rather than a re-implementation of it. Seven cases, including one that calls the **real** `AmqpReplyInbox::open` against a guaranteed-closed port, and each asserts the password never appears in any log line. `main()` reaches it through a single direct call I read at the merge, so the caller is covered by inspection rather than by a test that walks around the gate (see #113). Documented in wiki *11 Features* → *An unreachable broker does not stop the daemon*.
ltms closed this issue 2026-08-23 14:04:16 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#152