redeploy-fleetd.sh says "no ERROR lines since restart" while blind to the exact failure #493 is about #512
Closed
opened 2026-09-12 05:53:26 +02:00 by ltms
·
5 comments
No Branch/Tag Specified
main
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#512
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
#493 asked for two things. #510 delivered the first (stage the jar, swap after the old pid exits) and I have proven it on a real redeploy. This ticket is the second, which #510 did not do:
The script does not look for it
The log check it does have is anchored at
RESTART_MARK, taken just before the stop, so the region it scans does include the old daemon's shutdown output. The window is right. What it looks for in that window is wrong.Measured: the failure carries no ERROR token, so the count cannot see it
The script finishes with
ok "no ERROR lines since restart"whenREDEPLOY_ERROR_COUNTis 0. Here is the one real instance on this host, with a control:The control is the part that makes this a finding rather than a zero match. 1977 lines carry the token, so the grep works; this line simply does not carry it.
The reason is structural, not a tuning problem. A dead shutdown drain is an uncaught exception in a shutdown thread. The JVM's default handler prints it to stderr as
Exception in thread "…". It never goes through the logger, so it never gets a level, so no count ofERRORlines can ever include it.So the script prints a confident wrong conclusion
Put the three facts together for the case #493 describes:
status=143, indistinguishable from a clean SIGTERM stop (the script's own CB-594 comment establishes this),ok no ERROR lines since restart.Every channel reports success. That is the same shape as #500: the refusal message there blamed the policy for an interpreter fault, and here the success message vouches for a shutdown that did not happen. A wrong stated conclusion stops the next reader looking further, and this one is worse than silence because it is reassuring.
The fix
In the same fresh-log region the script already builds (
FRESH_LOG, fromRESTART_MARK):^Exception in threadandNoClassDefFoundError— not by log level.classify_amqp_connection_errors.no ERROR lines since restartmust not print when an uncaught exception was found in that window. Give it its own wording that names what was found and where.Decide deliberately whether this should make the script exit non-zero. My view: warn loudly but do not fail the redeploy, because by the time this is detectable the new daemon is already up and healthy, and failing the script would give the operator nothing to do differently. Say what it means instead — that the previous daemon's shutdown drain died, so some sessions may not have been released.
Acceptance
scripts/test-redeploy-fleetd.shusing a fixture log that containsException in thread "Thread-0" java.lang.NoClassDefFoundError: …and no line carrying anERRORtoken. The script must report it. A fixture that also carries an ERROR line would pass for the wrong reason — the point is that this must be found without one.grep -nafterwards, run the suite, quote the FAIL line and exit code, restore, confirm byte-identical withshasum -a 256, then a green control run.scripts/redeploy-fleetd.shagainst the live daemon. It is the channel the fleet talks through.Related
redeploy-fleetd.shwhile adjudicating #510.Independently reproduced on fleet01, and the obvious cheaper fix does not exist
Not my measurement. The fleet01 lead reported the block below from their own host. I have not
run any of it; I cannot — the failure needs a stale host, and mine is current. I am recording it
here because it is a genuinely separate instrument: systemd and journald rather than my launchd and
a flat file, on Linux rather than macOS.
Their numbers, from their real 2026-09-10 incident:
That confirms the finding above from a second source. The new part is the next block:
journalctl -p errmisses it too. On a systemd host, the natural fix for a missing text token isto stop grepping text and filter by severity instead. That is equally blind here, because the lines
are recorded at priority 6.
Their stated mechanism: on that host the daemon's fd 1 and fd 2 are the same socket
(
socket:[2503995569]for both). The stdout/stderr split is gone before journald sees anything, soeverything lands at the unit's default level. There is no severity channel to consult. I have not
checked this myself.
So three channels report one event, and all three are reassuring:
status=143— that is 128+15, identical to a clean SIGTERMERRORtoken, so "no ERROR lines since restart" is trueThe consequence for this ticket: the cheap fix is not available. Switching the scan from a text
grep to a severity filter would look like the tidier change and would buy nothing on a systemd host.
Keep the fix as written above — grep the shape (
^Exception in thread,NoClassDefFoundError), notthe level.
A stronger fix, and I checked that the line it needs does not already exist
The peer's suggestion, which I think is right: the honest check is not a negative one at all.
A negative check over a channel that cannot carry the failure is vacuous by construction. What the
script should look for is a positive assertion that the shutdown drain completed — a "drain
complete, N sessions released" line — and the alert should fire on its absence.
I measured whether such a line exists today, because I have already filed one ticket asking for a log
line that was there all along (#479). It does not exist. In
fleetd/src/main/java/dev/ltms/fleet/session/SessionManager.java:Both drain log calls are on abnormal paths. A drain that releases every session cleanly prints
nothing at all. So today the absence of a drain line is identical for "drained fine" and "died on
the first session", and there is nothing for the script to assert on.
That makes the fix two parts, in this order:
log.infoat the end ofdrainAll, naming how many sessions werereleased and how many were abandoned at the deadline. This is the load-bearing half — without it
part 2 has nothing to check.
specified above. The shape grep catches the exception when it happens; the missing-line check
catches a drain that died some other way, including one that leaves no exception at all.
Keep the shape grep. The two are not redundant: one finds a named failure, the other finds an
unnamed one.
Scope note
Part 1 is a change to
SessionManager, not to the script, so it may deserve its own ticket. I amleaving both here for now because splitting them risks the same failure #493 had — a two-item ticket
closed by the first item that produces a green run.
Related: #500 (a confident wrong conclusion is worse than silence), #479 (check the line does not
already exist before asking for it).
Merge constraint on part 1: the completion line must NOT be in a
finallyblockThis is an explicit non-goal, raised by the fleet01 lead, and I am recording it as a merge gate
rather than a suggestion because it is the obvious review comment and it destroys the thing being
built.
The tempting review note is "shouldn't we always log the drain result?". Putting the line in a
finallydoes exactly that, and it costs both halves of the fix at once:part 2 is meant to alert on,
in the same edit.
That is the #494 family created by a reviewer trying to be thorough. The line has to be the last
statement of the successful path, reachable only from it.
Second, smaller constraint: derive
releasedandabandonedincrementally as the drain proceeds,not from a collection read at the end. If a partial report is ever wanted it must be a different line
with a different verb. One line must not serve both "this drain finished" and "this is how far it
got".
I will check both of these against the diff before merging, not take them from the worker's report.
Why part 1 is worth more than the one incident suggests
The peer's framing, which I think belongs on the ticket: we only know that drain died because it
threw. A
NoClassDefFoundErrorreached the JVM's uncaught-exception handler and left a stack trace.Had the same drain hung on one session, or returned early on a condition rather than an exception,
there would be:
ERRORtoken,So the one incident we have is the loud variant of a fault class whose quiet variants are currently
undetectable. Part 1 is not detection for the incident that already happened; it is detection for
the ones that would leave nothing at all.
Part 1 is merged (#522). Part 2 is now unblocked.
The line is live on
main:Grep pattern for part 2 to anchor on:
Both merge constraints from my earlier comment were checked against the diff and hold: the line is the
last statement of
drainAll's normal path and is not in afinally(the onlyfinallyin the file isat
:372, unrelated), and the counts are incremented inside thedrainSnapshotloop rather than readfrom a collection at the end.
Verified by me, not taken from the report: build exit 0,
Tests run: 1698, Failures: 0, Errors: 0, Skipped: 0, which is +2 onmain's 1696. Deleting thecompletion line makes both new tests fail by name with "no drain-complete INFO logged"; restored
byte-identical to
d21ecd3adb3f66933e2f248a8ec81a2c324bfee323a3cc30da525842c33ea80a.Part 2 was held because #517's worker was editing
scripts/redeploy-fleetd.sh. That merged (#520)and the worker is torn down, so the script is free. Part 2 is now the only open item on this ticket:
assert the line's presence in the script's
FRESH_LOGregion, alongside the shape grep(
^Exception in thread,NoClassDefFoundError) already specified in the ticket body. Keep both — onefinds a named failure, the other finds an unnamed one.
One sequencing note for whoever takes part 2: #521 is also open against the same script and
extracts another decision out of the main flow. Either do them in one unit or land #521 first; two
workers in
scripts/redeploy-fleetd.shat once is a conflict I have already avoided once on thisticket.
An open semantic question in the new line, deliberately not changed
released++fires on every loop iteration that did not throw, including the case whereregistry.remove(paneId)returnednull— a session that another thread removed between theroster()snapshot and the release.releaseRemovedguards that case withif (removed != null), soit is an anticipated path, not a theoretical one.
draining.set(true)blocks newacquirecalls butdoes not block removals.
I am not filing this as a defect, because the honest reading is genuinely ambiguous:
The drain's contract is closer to the second. And the property the ticket actually needs — the line is
present if and only if the drain finished — is unaffected either way.
Recording it so the meaning is on the record, and so nobody "fixes" it in either direction without
deciding which reading they want. If part 2's check ever grows from presence to comparing the count
against an expected roster size, this becomes load-bearing and needs settling first.
Other silent-on-success teardown paths in the same class
Reported by the #522 worker, in scope for a sweep but not for this ticket. Not yet verified by me:
reapIdle— per-session success islog.debugonly, and itsintreturn (total reaped) isnever logged by its only caller,
SessionReaper.loop(). So a normal reap sweep, 0 or N, produces noline at INFO or above. That is the same defect this ticket just fixed, on a different path.
releaseRemoved's ordinary-cause path — every plain release islog.debug, while the abnormalbranches next to it (dirty worktree, remove failure) are
warn. Narrower, because it is per-sessionrather than an aggregate, but the same asymmetry.
The
reapIdleone matters more than it looks: the idle reaper is what silently destroys a member'sreport after
idleTtlSeconds, and today a sweep that reaped a session leaves no INFO-level trace ofhaving done so.
One more follow-up, filed separately
While checking the new tests I found that a pre-existing test pins the shared
SessionManagerloggerto
WARNand never restores it, so any later test asserting an INFO line silently sees nothing. The#522 worker hit exactly that and worked around it locally. Filed as #525 — it is a sixth mechanism
in the vacuous-test family, and the only one a reviewer of the new test cannot catch by reading the new
test.
Done, in two parts. Part 1 was #522 (the
drain complete: released=N abandoned=Mline inSessionManager.drainAll), merged earlier. Part 2 is #534, merged to main as7d71194.Part 1 had to come first, and the reason is the most useful thing this ticket produced. I wrote the fix above as a purely negative check — grep the shutdown window for the exception shape. The fleet01 lead pointed out that a negative check over a channel that cannot carry the failure is vacuous by construction, and when I went to add the absence alarm they proposed, there was no line to alarm on: the log was byte-identical for a drain that released everything and a drain that died on its first session, because neither printed anything. The absence I wanted to alert on was the normal case. So the signal had to be created before it could be checked.
Acceptance, item by item
scan_uncaught_exceptionsgreps forException in thread/NoClassDefFoundError, kept separate fromclassify_amqp_connection_errorsas asked.The exit-code decision went the way I proposed: warn loudly, never
die(). By the time this is detectable the new daemon is already up, and failing the script would give the operator nothing to do differently.What the detection actually reports now
Four outcomes, from
report_shutdown_drain:complete— #522's line is present; name the counts.died— line absent and the exception shape found; name what was found and say sessions from the previous daemon may not have been released.unknown— line absent and no exception shape either. Explicitly not a pass and not a failure: either that daemon predates #522, or its drain failed without throwing (hung, or returned early).n/a— no previous daemon was stopped this run, so there is no shutdown window to have an opinion about.no ERROR lines since restartnow prints only oncompleteorn/a.The third state was the point, and I mutated it to check
unknownexists because absence has two causes needing opposite handling, which is the sentinel-conflation shape. A third state is only worth having if something breaks when it collapses, so I collapsed it: changed theelsebranch to setcomplete, so "cannot tell" reports as a pass. Result exit 1, one FAIL —cannot-tell fixture must set REDEPLOY_DRAIN_STATE=unknown: expected unknown, got complete. Load-bearing, not decoration.Second mutation of mine, on the positive half: changed
find_drain_complete_line's pattern fromdrain complete: released=todrain finished: released=, one site. Exit 1, one FAIL —find_drain_complete_line did not capture the present line. Proof it applied, against a pristine copy: the full grep line 1 → 0, the mutant form 0 → 1, bare phrase 2 → 1 with the comment occurrence untouched. Restored byte-identical, green control re-run.Both of mine were deliberately ones the worker had not run.
The fourth state, which the worker added and flagged rather than quietly keeping
n/ais not in my three-state wording. Accepted, and it belongs. "The question does not apply" is a third cause of absence, not a variant of "cannot tell". Without it the new warning would fire on every clean cold start, and a warning that cries wolf on the most common path trains the operator to skip it — destroying the absence signal just as surely as putting #522's completion line in afinallywould have. Same defect, arrived at from the other end.The worker also proved the gate is consulted, with a cold-start fixture whose content deliberately looks like a died drain. That is the right way to test a gate: make the fixture such that skipping the gate gives the wrong answer.
Verified on the merged tree, not the branch
The branch was based on
8335b12while main had moved tobec87f9, so I merged locally and tested the result. Suite exit 0, 0 lines matching^FAIL:, 60 test functions defined and 60 invoked — checked with acommagainst the invocation list rather than by comparing two counts, since two equal counts can both be wrong.bash -nexit 0 on both scripts under /bin/bash 3.2.57 and env bash 5.3.9.One untested line I found, reported not fixed
The main flow's
HAD_OLD_PID=0; [ -n "$OLD_PID" ] && HAD_OLD_PID=1runs underset -euo pipefail, and on a cold start that test fails. Sourcing stops before the main flow, so no behavioural test reaches the line. Ifset -efired there, every cold start would abort before the health checks ever ran.It does not fire —
set -eexempts the left side of an&&list — and I confirmed that by running it under both shells rather than reasoning about the manual:survived, HAD_OLD_PID=0on 3.2.57 and on 5.3.9. Safe, but it is untested main-flow wiring, the same class as #528's item 1, and it goes on that list rather than being called covered.Also in this verification, a mistake of mine worth recording
My first proof cell for mutation B used escaped double quotes inside an already double-quoted command substitution, so the shell split the pattern on spaces and grep read the words as filenames. It printed a symmetric "2 and 2" next to
ugrep: No such file or directorywarnings. The kill was never in doubt — the suite named the exact function — but the cell whose entire job was to prove the mutation had applied proved nothing, and it failed in the plausible direction. This is the already-written-down double-quote trap, hit inside the guard against it.Closing. #511 remains open for the other two
redeploy-fleetd.shdefects.The
unknownarm is self-clearing, so its test is load-bearing, not illustrativeReopening the reasoning, not the ticket. This is the fleet01 lead's argument and it corrects something I wrote when closing.
What I got wrong. I recorded that the detection "validated itself in the field on its first real run" — it reported
unknown/ "cannot tell" on the first redeploy after the change, correctly, because the replaced daemon started 11:12:15 on a jar built 11:12:13 while #522's drain line only landed at 11:29:01. I offered that as evidence the three-state design was right.It is not. It is evidence about the base rate of the third state — it says the cannot-tell condition was common enough to appear on run one. That is a fact about deployments, not about the design. The design argument stands on its own and never needed the run.
The consequence, which is the part that matters. That condition clears itself. The
unknownfired because the daemon being replaced predated #522's line. After that deploy, every daemon a future check replaces carries the line. So production will never reach theunknownarm again on this path. It fired exactly once, at the only moment it could, and from here it is silent permanently.Which means the mutation test covering the third state is not a nicety. It is the entire remaining coverage of that branch. And a branch that has gone quiet forever is indistinguishable from a branch that has broken: if the
unknownarm rots, nothing in production will ever say so, because nothing in production will ever execute it.So, stated plainly for whoever reads this next: the cannot-tell fixture in
test-redeploy-fleetd.shis load-bearing, not illustrative. Do not delete it on the grounds that the state it covers "cannot happen any more". That it cannot happen any more in production is precisely why the test is the only thing left holding it.This generalises past this ticket. A state that a system has grown out of still has code, and that code still has to be right on the day something puts the system back into it — a rollback, a host that missed a deploy, a second fleet on an older jar. The test is the only instrument left pointed at it.
Related: #497 (the third state itself), #522 (the drain line whose arrival is what makes the condition self-clearing).