redeploy-fleetd.sh reports "no process appeared" on a deploy that fully succeeded #603

Closed
opened 2026-09-20 11:08:11 +02:00 by ltms · 0 comments
Owner

What happened

I redeployed the live daemon with scripts/redeploy-fleetd.sh. The script printed:

FAIL no process appeared

The deploy had worked. I checked four ways afterwards, by hand:

  • /healthz answered 200
  • the daemon was running as pid 99966
  • fleetd/fleetd.out had a fresh fleetd listening line, dated after the restart
  • fleet_whoami still answered primary

(Those four numbers were measured earlier in this session, right after the deploy. The two source facts below I re-measured just now.)

Why it happens

scripts/redeploy-fleetd.sh:1288:

for _ in $(seq 10); do
  NEW_PID="$(running_pid)"
  [ -n "$NEW_PID" ] && break
  sleep 1
done
[ -n "${NEW_PID:-}" ] || die "no process appeared. ..."

That gives the process 10 seconds to appear, and then it is a hard die.

Under launchd, launchctl load returns as soon as launchd has accepted the job. The java process does not exist yet at that moment. On this host it takes longer than 10 seconds to show up.

Compare with line 82 in the same script:

HEALTH_WAIT=60    # seconds to wait for /healthz to answer after start

So the script is willing to wait 60 seconds for the daemon to answer, but only 10 seconds for it to exist. It dies at the shorter gate and never reaches the longer one — the check that would have told it the truth.

Why this one matters more than a cosmetic wrong message

It fails in the dangerous direction. It says the deploy failed when the deploy worked.

A lead reading FAIL no process appeared reasonably concludes the daemon is down and reaches for kill plus a manual start. That is exactly what this script exists to prevent, and CLAUDE.md and the redeploy-fleetd skill both ban it:

If a call is still refused, do not route around it by running the stop and start as separate commands — that is exactly the approval the script replaced.

So a false FAIL here pushes the operator or the lead straight into the banned path, at the moment they are most likely to take it.

This is the second time through the same door

Line 77 of the same file already records a fix for a false "no process appeared":

Anchoring on the absolute path alone was a real bug: the daemon restarted correctly and the script still reported "no process appeared", because it launched with …

That fix corrected how the process is matched. It did not touch how long the script waits for one. The failure mode came back through the other door, with a green build.

Worth looking for the same shape elsewhere in the script: any other place where a fixed short timeout decides a hard die, while a later and more truthful check is never reached.

Suggested direction, not a specification

Do not simply raise 10 to a bigger number. That trades a false FAIL for a slow one, and picks a new magic number that will be wrong on the next host.

Better: make "did it start?" ask the question the script actually cares about. /healthz answering 200 from a process whose start time is after the restart is direct proof the new daemon is up. The pid lookup is a means to that, not the goal. If the pid poll is kept, it should share the HEALTH_WAIT budget rather than hold its own shorter one, and a miss should fall through to the health check instead of killing the run.

Whatever the fix, the acceptance test should be a property: with the process made to appear later than the pid poll's budget, the script must still report success as long as the daemon really is up — and must still report failure when it really is not. Both halves are needed. A test that only proves the slow-start case passes can be satisfied by a script that never fails at all.

Not done

I have not fixed this. I hit it, verified the daemon by hand, and carried on.

## What happened I redeployed the live daemon with `scripts/redeploy-fleetd.sh`. The script printed: ``` FAIL no process appeared ``` The deploy had worked. I checked four ways afterwards, by hand: - `/healthz` answered 200 - the daemon was running as pid 99966 - `fleetd/fleetd.out` had a fresh `fleetd listening` line, dated after the restart - `fleet_whoami` still answered `primary` (Those four numbers were measured earlier in this session, right after the deploy. The two source facts below I re-measured just now.) ## Why it happens `scripts/redeploy-fleetd.sh:1288`: ```bash for _ in $(seq 10); do NEW_PID="$(running_pid)" [ -n "$NEW_PID" ] && break sleep 1 done [ -n "${NEW_PID:-}" ] || die "no process appeared. ..." ``` That gives the process **10 seconds** to appear, and then it is a hard `die`. Under launchd, `launchctl load` returns as soon as launchd has accepted the job. The java process does not exist yet at that moment. On this host it takes longer than 10 seconds to show up. Compare with line 82 in the same script: ```bash HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start ``` So the script is willing to wait 60 seconds for the daemon to answer, but only 10 seconds for it to exist. It dies at the shorter gate and never reaches the longer one — the check that would have told it the truth. ## Why this one matters more than a cosmetic wrong message **It fails in the dangerous direction.** It says the deploy failed when the deploy worked. A lead reading `FAIL no process appeared` reasonably concludes the daemon is down and reaches for `kill` plus a manual start. That is exactly what this script exists to prevent, and `CLAUDE.md` and the `redeploy-fleetd` skill both ban it: > If a call is still refused, do **not** route around it by running the stop and start as separate commands — that is exactly the approval the script replaced. So a false FAIL here pushes the operator or the lead straight into the banned path, at the moment they are most likely to take it. ## This is the second time through the same door Line 77 of the same file already records a fix for a false "no process appeared": > Anchoring on the absolute path alone was a real bug: the daemon restarted correctly and the script still reported "no process appeared", because it launched with … That fix corrected **how** the process is matched. It did not touch **how long** the script waits for one. The failure mode came back through the other door, with a green build. Worth looking for the same shape elsewhere in the script: any other place where a fixed short timeout decides a hard `die`, while a later and more truthful check is never reached. ## Suggested direction, not a specification Do not simply raise `10` to a bigger number. That trades a false FAIL for a slow one, and picks a new magic number that will be wrong on the next host. Better: make "did it start?" ask the question the script actually cares about. `/healthz` answering 200 from a process whose start time is after the restart is direct proof the new daemon is up. The pid lookup is a means to that, not the goal. If the pid poll is kept, it should share the `HEALTH_WAIT` budget rather than hold its own shorter one, and a miss should fall through to the health check instead of killing the run. Whatever the fix, the acceptance test should be a property: with the process made to appear later than the pid poll's budget, the script must still report success as long as the daemon really is up — and must still report failure when it really is not. Both halves are needed. A test that only proves the slow-start case passes can be satisfied by a script that never fails at all. ## Not done I have not fixed this. I hit it, verified the daemon by hand, and carried on.
ltms closed this issue 2026-09-20 11:30:10 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#603